Backtesting
The Step-by-Step Process of Backtesting a Strategy
Choose the problem that reflects your current situation.
Step-by-Step Process of Backtesting a Strategy
1. Defining the Strategy Rules Before You Start
👉 My backtests look brilliant, then fall apart as soon as I trade the same idea live
The Reality Check
Updated 2026
If the rules are still in your head, the backtest will invent them as you go. That is not research. That is a highlight reel with extra steps. Frozen rules are the only way the sample means anything.
❓ The Painful Question Traders Ask
“Why do my backtests look brilliant — and then fall apart as soon as I try to trade the same idea live?”
The Core Insight
Updated 2026
Define entry, invalidation, target or management, filters, session, and skip conditions in writing before candle one. If two traders could not mark the same trades from your page, the rules are not ready. Clarity is the test. Profit is later.
Related Reflection Questions
- Could I hand this page to someone else and get the same trades?
- What am I still deciding “when I see it”?
- Which filter do I add only after a loss in the test?
- Have I written when I must stand aside?
- Is the stop a location or a feeling?
⚠️ The Brutal Consequences of Avoiding This
- Hindsight entries that never existed in real time
- A strategy that only works when you already know the outcome
- Endless tweaks that destroy comparability
- False confidence, then live confusion
- Repeating the same “test” forever
✅ The Deep Solution
Continue to the Full Lesson
2. Identifying and Marking Every Trade Opportunity
👉 I only test the trades I like looking at
The Reality Check
Updated 2026
Skipping the ugly valid setups is how you fake a win rate. If it met the rules, it goes on the sheet — win, loss, or “I would have hesitated.” A backtest that only marks the pretty ones is a mood board.
❓ The Painful Question Traders Ask
“Am I testing the strategy — or only the trades I like looking at?”
The Core Insight
Updated 2026
Every opportunity that matches the frozen spec is a data point. Mark it at the moment of the trigger, with future price hidden or ignored. Missed marks are as important as taken ones if you are studying execution. Completeness is the integrity of the test.
Related Reflection Questions
- Do I skip trades that look messy but still fit?
- Do I mark at the signal bar or after I see the follow-through?
- How do I handle consecutive signals — each as a new opportunity or one campaign?
- Would I be embarrassed to show the skipped charts?
- Is my sample smaller because I was selective?
⚠️ The Brutal Consequences of Avoiding This
- Inflated expectancy
- Shock when live includes the “ugly” A’s
- No count of true frequency
- Arguments with yourself about what “counts”
- A strategy you cannot size because the sample is theatre
✅ The Deep Solution
Continue to the Full Lesson
3. Logging Entry, Exit, and Reasoning for Each Trade
👉 I cannot explain my backtest trades a week later, even though I tested them
The Reality Check
Updated 2026
A win/loss tick is not a log. Without entry, exit, R, and the reason at the time, you cannot audit whether you followed the spec or edited it. Memory will supply a better reason later. The sheet must capture the one you had at the click.
❓ The Painful Question Traders Ask
“Why can I not explain my backtest trades a week later — even though I ‘tested’ them?”
The Core Insight
Updated 2026
Log the facts and the sentence you would have said before knowing the outcome. Entry price, stop, exit, R, setup tag, and one line of reasoning. That row is how you catch rule drift. Without it, the summary is a rumour.
Related Reflection Questions
- Could I reconstruct this trade from the row alone?
- Did I write the reason before or after I saw the result?
- Are my exits following the spec or “because it felt done”?
- Do I log time of day and market condition?
- What field do I skip when I am tired — and what does that hide?
⚠️ The Brutal Consequences of Avoiding This
- You cannot compute expectancy honestly
- You cannot see which rule you break
- Reviews become stories
- You repeat the same vague test
- Live journaling feels optional because the test was optional too
✅ The Deep Solution
Continue to the Full Lesson
4. Measuring Win Rate, Risk-Reward, and Expectancy
👉 I don’t know which number tells me whether this backtest is a real edge
The Reality Check
Updated 2026
A high win rate with ugly R can still be a losing business. A low win rate with large winners can still pay. If you only quote one number, you are decorating. Expectancy is the sentence: average R per trade over the sample, after costs you actually pay.
❓ The Painful Question Traders Ask
“Which number tells me if this backtest is a real edge — and which numbers just make me feel better?”
The Core Insight
Updated 2026
Compute win rate, average win in R, average loss in R, then expectancy. Haircut for spread and slippage. Sample size must be large enough that one hero trade cannot own the story. Win rate is a headline. Expectancy is the job.
Related Reflection Questions
- Did I include costs?
- Is one trade carrying the average win?
- What happens to expectancy if I remove the best 5% of winners?
- Is my win rate high because I skipped losers in the log?
- Could I live with this win rate emotionally?
⚠️ The Brutal Consequences of Avoiding This
- Trading a “high win” system that loses money
- Abandoning a valid low-win-rate edge out of discomfort
- Sizing from the wrong metric
- Curve-fitting to protect a favourite number
- Live shock when costs show up
✅ The Deep Solution
Continue to the Full Lesson
5. Tagging and Categorizing Trades by Setup Type
👉 I don’t know if my edge is one thing or a mix where half the setups should be retired
The Reality Check
Updated 2026
A blended win rate can hide a setup that is a parasite. If you do not tag type, condition, and session, you will keep the dead sleeve because the living one pays for it. Categories are how you kill what does not earn its risk.
❓ The Painful Question Traders Ask
“Is my edge one thing — or a mix where half the tags should be retired?”
The Core Insight
Updated 2026
Tag before you know the result: setup family, trend vs range, session, news vs quiet. After the sample, sort expectancy by tag. Keep the tags with evidence. Sit out or redesign the rest. Mixing without tags is how a weak idea survives inside a strong average.
Related Reflection Questions
- How many setup names do I actually use — and can I define each?
- Do I tag after the result, when I already like the trade?
- Which tag would I be afraid to see isolated?
- Is “other” my largest bucket?
- Would I still take a tag if it were the only thing I traded?
⚠️ The Brutal Consequences of Avoiding This
- Carrying losing variants forever
- No idea when to stand aside
- Arguments that “the strategy works” while one sleeve bleeds
- Overfitting by adding tags after the fact
- Live confusion about which A you are in
✅ The Deep Solution
Continue to the Full Lesson
6. Calculating Drawdowns, Heat, and Equity Curve Behavior
👉 The backtest makes money, but I already know I would have quit in the middle
The Reality Check
Updated 2026
A positive expectancy with a drawdown you cannot sit through is not your edge. Heat is how much open or sequential pain the method produces. The equity curve tells you if the path is liveable — not just if the end is green.
❓ The Painful Question Traders Ask
“The backtest makes money — so why do I already know I would have quit in the middle?”
The Core Insight
Updated 2026
Measure max drawdown in R, longest losing streak, and typical heat to the stop. Plot the equity curve. If the valley would have changed your size or rules, the test is not complete until you admit that. Survivable path beats prettier total return.
Related Reflection Questions
- What is the worst peak-to-trough in this sample?
- How many losers in a row did I eat?
- Does the curve stair-step or cliff?
- Would my funded daily cap survive this heat?
- Am I only looking at the final number?
⚠️ The Brutal Consequences of Avoiding This
- Abandoning a valid method at the first real valley
- Sizing as if the curve is smooth
- Funded breaches from normal heat
- Shock that “winning systems” hurt
- Tweaking mid-drawdown in live because you never measured it
✅ The Deep Solution
Continue to the Full Lesson
7. Reviewing Missed Trades and Setup Violations
👉 I follow the rules in the backtest, then skip and break them when it is time to click
The Reality Check
Updated 2026
The taken trades are not the whole test. Missed valids tell you about hesitation. Violations tell you about integrity. If you only study the rows you took cleanly, you will be surprised by your live behaviour — because live is where misses and violations live.
❓ The Painful Question Traders Ask
“If I follow the rules in the backtest, why do I still skip and break them when it is time to click?”
The Core Insight
Updated 2026
Log misses (valid, not taken) and violations (taken, not valid, or rules broken) as first-class events. Review them as clusters: time of day, after losses, boredom. The method’s expectancy assumes a behaviour. If the behaviour is not in the test, the expectancy is not yours yet.
Related Reflection Questions
- Did I skip because of the spec — or because of discomfort?
- What violation do I still call “discretion”?
- Are misses clustered after red trades?
- Would including misses change how often I think this trades?
- What rule do I break the moment the test is over?
⚠️ The Brutal Consequences of Avoiding This
- Live frequency much lower or sloppier than the test
- Blame the strategy for your skips
- Hidden extra risk from violations
- No training target for execution
- A clean sheet that cannot survive contact with you
✅ The Deep Solution
Continue to the Full Lesson
8. Recording Psychological Reactions During Manual Tests
👉 I still feel impatient, scared, or tempted to cheat the next bar, even though it is only a test
The Reality Check
Updated 2026
Manual backtesting still has a nervous system. Boredom, urge to skip, urge to peek, excitement after a streak — if you do not record it, you will think live psychology is a new problem. It was already in the replay. You just did not log it.
❓ The Painful Question Traders Ask
“If this is only a test, why do I still feel impatient, scared, or tempted to cheat the next bar?”
The Core Insight
Updated 2026
The test is a rehearsal of attention. Log the reaction at the bar: impatient, FOMO, “just this once,” fatigue. Those notes predict live leaks. A backtest that ignores psychology is incomplete even if the R math is clean — because you are the operator.
Related Reflection Questions
- When did I want to skip ahead in replay?
- Did a winning streak in the test change how loosely I marked?
- What time in the sitting do my notes get sloppy?
- Did I feel the same body cues I feel live?
- Would I trust a test I rushed?
⚠️ The Brutal Consequences of Avoiding This
- Rushed samples that leak hindsight
- Surprise at live emotion you already displayed in replay
- No training plan for the real operator issues
- False belief that “sim is calm”
- Fatigue errors counted as strategy errors
✅ The Deep Solution
Continue to the Full Lesson
9. Creating a Summary Report That Tells the Real Story
👉 I don’t know if my backtest is permission to go live or just more paper
The Reality Check
Updated 2026
A pile of rows is not a conclusion. If you cannot write one page that a sceptical trader would respect — N, expectancy after costs, drawdown, tags, misses, limitations — you do not have a result. You have a folder. The real story includes what the test cannot say.
❓ The Painful Question Traders Ask
“How do I know if this backtest is permission to go live — or just more paper?”
The Core Insight
Updated 2026
The summary is a decision document: what was tested, what it earned in R, what it cost in pain, which tags survive, what remains untested (live friction, psychology, regime). Go live only if the story is complete enough to size small. If the story is a highlight, stay in testing.
Related Reflection Questions
- Could I defend this sample to someone who wants to find holes?
- What did I not test (costs, news, execution delay)?
- Which chart in the report would I rather hide?
- Is the recommendation “live small,” “forward test,” or “kill”?
- Did I write limitations — or only the green parts?
⚠️ The Brutal Consequences of Avoiding This
- Going live on a vibe
- Repeating tests because there was no decision
- Hiding drawdown in a spreadsheet tab
- No baseline to compare the next version against
- Arguing with yourself every week with no document
✅ The Deep Solution
Continue to the Full Lesson
Continue Learning
Next Module: Analyzing Your Backtest Results Like a Professional →