Quant Basics

Backtest Overfitting: 3 Checks Before Going Live

2026-10-06·10 min read

You've just swept the parameters, and the backtest curve looks so clean you want to hit live trading right now — 120% annualised, 8% max drawdown, 62% win rate. What's probably about to happen: the first month of live trading eats half of that three-year backtest profit. That's not bad luck. That's the receipt overfitting leaves behind.

This post isn't about which strategy logic you chose — the logic itself may be fine. The problem usually lives in how you picked the parameters, and on which slice of data. Here are three things you have to do after optimisation finishes and before the strategy touches a real account. None of them are optional.

TL;DR
TL;DR: Parameter optimisation isn't the finish line, it's the starting line. Once you've picked the "best" parameter set, do at least three things — a walk-forward run on unseen data, a parameter-sensitivity sweep, and a small live run at a third of your normal size. Only after all three pass does the strategy earn a full-size position. However pretty the backtest numbers look, if these three checks haven't run, the strategy hasn't actually been validated.

First, Be Clear About What Overfitting Actually Fits to What

The mechanism is simple. On the same slice of historical data, you're doing two things at once — picking the strategy and tuning the parameters. Whatever combination wins is, fundamentally, the combination that scored perfectly on this stretch of noise— not the combination that captured the market's structural behaviour.

Think of the data as an exam paper. If you read the paper first and then work backwards to pick your answering method, of course you get 100% — but that tells you nothing about how you'll do on the next paper. That is exactly what overfitting is: you use data you've already looked at to pick your parameters, and then you use the same data to verify the choice. Nowhere in that loop is there any information about whether those parameters work on data you haven't seen.

Mathematical signature: in-sample error tiny, out-of-sample error blows up
This is the most direct signal you have. On the optimisation slice (in-sample) you see win rate 62%, Sharpe 2.1; switch to a slice you never touched (out-of-sample), and the win rate falls to 48%, Sharpe 0.4. That gap isn't "bad luck" — it's direct evidence that the curve was bent to fit this stretch of history.

Check One: Split the Data First, and Only Optimise on In-Sample

The most basic step, and the one most often skipped: slice the data into two blocks before you do anything else. One block is your playground — tune parameters, add indicators, rewrite rules, whatever you want. The other block is locked away and only comes out the moment optimisation is done. It never participates in any decision. Look at it once and it's spent — you don't get to use it again.

70 / 30
Common split — first 70% for optimisation
80 / 20
When you don't have much data
1 look
OOS gets exactly one look — after that it's burned
6–12 months
OOS should cover at least one meaningful market regime

There's no single right split — it depends on how many years of data you have. (Common baselines gathered while writing this piece.)

Here's the trap that's very easy to fall into: you look at the out-of-sample block, don't like what you see, go back and tweak parameters, then look again.The moment you do that, the OOS is spent — it's been touched indirectly by the optimisation loop, so it's not fundamentally different from in-sample any more. The "second look" at OOS isn't out-of-sample; it's just another in-sample.

One way to defend against this is to carve off a third block called hold-out. It never gets touched during development — the first time you look at it is the day the strategy actually goes live. It can be small (15% to 20% is fine). Its job is to give you one honest look at what these parameters do on data that really wasn't part of the process.

Check Two: Walk-Forward — Make Optimisation a Rolling Event

A single in-sample / out-of-sample split has one weakness: you've only tested one cut point — "parameters chosen on the first 70% still work on the last 30%". If the market regime has shifted, the parameters you tuned three years ago on old data might have been quietly losing edge that entire time, and you just haven't noticed.

Walk-forward keeps the idea, but rolls it. Slice the data into many short segments; for each segment run "pick parameters on the previous N months, evaluate on the next M months"; then stitch the evaluation windows together into one continuous out-of-sample curve. That curve is a much better proxy for "what this process would actually have earned you over these years".

Single OOS splitWalk-forward
How many time points validated?1 cut pointMany rolling windows
Resistance to regime shiftsLow (parameters chosen on old chunk)High (each window re-picks)
Setup complexitySimpleHave to pick window length and step size
Compute cost1 optimisation runN optimisation runs
Similarity to live behaviourNot muchClose to how you'd actually re-tune

The tradeoff between the two. Walk-forward is more expensive but much closer to reality: strategies get re-tuned periodically.

Walk-forward has a byproduct that's just as important: what parameters did each window pick? If ten windows land on ten wildly different values, that isn't a "sensitive strategy" — it means the strategy has no real regularity, it's just fitting whichever noise happens to be in front of it. If the ten windows all settle around similar values, that's when you can start to believe the strategy has caught hold of something.

The tradeoff here you have to make yourself
How long should the window be? How large should the step? There's no official answer and the choice affects the result. A common starting point is "train window = 6 months, test window = 1–2 months, step forward one test window at a time", but the right numbers depend on whether you're on crypto vs equities, daily bars vs hourly. This is a guideline derived from the concept, not a black-and-white rule — sanity-check against your strategy's trade frequency before you commit.

Check Three: Parameter Sensitivity — Did You Pick a Peak or a Ridge?

The same Sharpe 2.1 backtest can be two very different things. A peak: you happened to land on the only high point, and shifting any parameter by one notch drops Sharpe to 0.5. Or a ridge: everything around your chosen point is Sharpe 1.8–2.1, and a small nudge barely moves anything. The first is textbook overfitting. The second is the kind of thing you can actually take live.

1
Sweep every dimension of the chosen parameter set by ±10%. How much does Sharpe drop?
Drops 50% or more →It's a peak. This set is almost certainly fitting noise on this data slice — do not take it live. Go back and simplify the logic or shrink the parameter space.
Drops 20–40% →It's a slope. You can keep going, but test at a small position size.
Drops under 10% →It's a ridge. This is the shape you're looking for. Move on to the next step.
2
Draw the range on each dimension where the parameters still clear your threshold. How many combinations survive?
Only that one combination clears →Back to the peak case — treat it the same way.
A continuous block of combinations all clear →Pick from the median of that block, not the highest point. The median is the location that can absorb a little drift.

Plotting a 3D heat-map of "parameter vs performance" is the most direct way to see this. A strategy with real regularity looks like a smooth slope; an overfit one looks like a handful of sharp spikes. You don't need any new tool for this — just visualise the optimisation results you already have.

Side Note: Do You Have Enough Data to Support That Many Parameters

More parameters, more overfitting — that's almost common sense, but rarely does anyone give you a concrete threshold. Here's a rough rule of thumb, taken from the general machine-learning principle (not something specific to trading): each additional parameter you tune roughly doubles the independent samples you need. Treat "number of trades in the backtest" as your sample count:

2 parameters
At least 100 trades
3 parameters
At least 200 trades
5 parameters
At least 500 trades
7 or more
Usually already beyond saving

These are rough health-check numbers, not precise thresholds. The more complex the strategy, the higher the numbers should be.

If your strategy fires 40 times in three years but you're tuning 5 parameters, the problem isn't that the strategy "only trades on rare setups" — the data simply doesn't contain enough independent events to distinguish "actually works" from "happened to fit". The only sensible move here is to switch to a higher-frequency timeframe, or cut the parameters down to two or fewer.

All Three Checks Passed: The Final Two Steps Before Going Live

Even if all three checks are green, don't jump straight to a full-size position. Two post-launch things matter just as much:

  • Start with a third of your normal position size or smaller for one to two months. Things a backtest silently omits — real fill latency, slippage, partial fills, the odd exchange rejection — only live trading tells you about.
  • Log every backtest vs live divergence. Same-bar signal, actual fill price, and end-of-trade PnL — is there systematic drift from what the backtest promised?
  • Set a stop line: if cumulative live loss exceeds X% or if N consecutive trades deviate from the backtest, pause and inspect. Don't wait until three months are gone before you notice something is off.
  • Before scaling the position up, re-run walk-forward using the one-to-two months of live data. That's the most honest out-of-sample slice you'll ever get.

How you actually scale a position up from small to normal size is a separate problem from the strategy logic itself — it belongs to money management. That part is covered in Kelly Criterion and Position Sizing — which is also where you'll find why "half-Kelly" is the practical compromise most people settle on.

What These Three Checks Cannot Save You From

Having laid out what these steps can do, it's worth being clear about what they cannot — this is the part of the article most easily misread.Walk-forward and parameter sensitivity are both validations within the same market regime. They can't defend you against the following:

These things only live trading tells you
Structural market change: the 2021 crypto bull run and the 2023 chop are two different markets. Parameters walk-forward-tuned on 2021 data will still fail on 2023.
Execution costs underestimated: most backtest tools default to idealised fees and slippage. Real market-maker inventory, taker fees, and partial fills all take a bite out of the theoretical net profit.
Signal sparsity: a strategy that fires ten times a year can pass walk-forward and still be meaningless — ten data points is too few for the result to carry statistical weight.
Whether you'll actually leave it alone: a backtest that survives a 40% drawdown on paper is not the same as you sitting through it live. At 20% you can't sleep any more and you switch it off. That's a behavioural problem; no amount of data testing measures it.

An Honest Section: Putting Backtest Numbers in Their Place

This post deliberately does not quote any specific backtest performance numbers. Any performance figure not attached to instrument, period, parameter set, and out-of-sample method — however pretty — is dangerous for a reader to make decisions from.Past performance does not indicate future results, and that's especially true of backtest results: they are simulations, not live-trade records.

If you plan to put backtest results in your own article, on Twitter, or in signals for paying subscribers, the minimum viable disclosure is: instrument (e.g. BTCUSDT.P), time range (e.g. 2023-01 to 2025-12), parameter set (not just "after optimisation"), whether you did walk-forward, and one sentence stating "these are simulated results and do not represent live-trading profits". Miss any one of these and no one else can decide whether your track record is worth following.

This is also why we suggest reading How Retail Traders Get Started in Quant first — there's a whole section in it on "500% annualised is either overfitting or hasn't blown up yet", which reads as a companion piece to this one.

FAQ

Every walk-forward window re-picks parameters — isn't that just continuously optimising?
Yes, and that's the point. Walk-forward is simulating the behaviour "you will periodically re-tune parameters" — it isn't trying to find one set of parameters that lasts forever. It's measuring whether "this strategy logic plus a periodic re-tuning process" is sustainable. In live trading, the actual routine is exactly that: every so often, re-optimise on the latest data.
Why ±10% for the sensitivity sweep, not ±20% or ±5%?
No official answer — 10% is a reasonable starting point. The real motivation is "how much drift can these parameters absorb in live trading before they stop being reliable". If your parameters are discrete steps (e.g. EMA can only be an integer), "±1 step" and "±2 steps" are more appropriate than ±10%. This is a sensitivity check by concept, not a precise formula.
If in-sample looks bad, can I just change the strategy logic and re-run the whole process?
You can, but every round counts. Each time you go back and rewrite the logic, the data has been "looked at" one more time, and the overfitting risk compounds. A practical discipline is to cap the number of iterations(say, three), and at the end bring out the hold-out block for a final check. If it still doesn't work after several attempts, it may be that this problem simply doesn't contain a regularity to find — the issue isn't the parameters.
Which strategies are most prone to overfitting?
A few high-risk symptoms: many parameters (five or more),sparse signals (fewer than a few dozen trades a year),compound conditions (three or more indicators must all agree before entry), short timeframe with only a few months of backtest data. If two or more of these apply, be extra suspicious — the prettier the backtest, the more suspect it is.
How do I actually run these three checks in tooling?
Data splitting and parameter sensitivity are supported by most backtest tools; the key is you have to remember to do the split yourself — don't just click "optimize" and stop when you see the best result. Walk-forward is built into some platforms and requires assembly on others via Python (backtesting.py, vectorbt). Look for "walk-forward" or "rolling window" in your tool's settings first; if it's not there, export the data and run it elsewhere.

Get started

Ready to ship what you just learned?

TVSBot is a non-custodial TradingView-to-exchange auto-execution platform. Take your validated strategy, wire the webhook, keep your own API key, and run dry-run or a small live position first — this is the step that comes after 'all three checks passed'.

Get started free