PaceAlgo
Field notes

Methodology

Why a perfect backtest means nothing: in-sample vs out-of-sample

July 13, 2026 · 6 min read

BACKTESTINGOVERFITGENERALISESPerfect fit ≠ predictiveILLUSTRATIVE — NOT A LIVE TRADE RECORD

A flawless backtest is one of the easiest things in the world to produce — and one of the least meaningful. The line between a backtest that actually predicts something and one that just describes the past comes down to a single distinction most marketing never mentions: in-sample versus out-of-sample.

In-sample: the exam you've already seen the answers to

In-sample data is the history a strategy was built and tuned on. The developer runs it, sees which trades lost, and adjusts the rules and settings until those losses shrink. Do that long enough and the equity curve becomes beautiful — not because the strategy found an edge, but because it was shaped to fit that exact stretch of history.

That's not prediction. It's memorization. Sitting an exam after you've already read the answer key guarantees a perfect score and tells you nothing about how you'd do on questions you haven't seen.

Out-of-sample: the only test that counts

Out-of-sample data is history the strategy was never allowed to see while it was being built. You hold it back, lock it away, and only run the finished strategy against it once — at the end. Because the rules were never tuned to this data, the result is an honest estimate of how the strategy behaves on conditions it doesn't already know.

This is the number that matters. It's the closest thing to a preview of live trading you can get before risking real money — and it's exactly the number that curve-fit products can't show you, because they used all of their history to make the picture look good (see how to spot a curve-fit indicator).

Why in-sample always looks amazing

Give a model enough knobs to turn and it can fit almost any past perfectly — including the random noise that will never repeat. Statisticians call it overfitting. The more parameters a strategy has and the more they were tuned, the more its backtest describes coincidences in one specific history rather than anything durable.

The uncomfortable truth: an amazing in-sample result isn't evidence of an edge. Often it's evidence of the opposite — that the strategy was tuned hard enough to fit noise.

How to tell an in-sample-only claim

  • A single backtest over “all available history,” with no mention of any data held back for testing.
  • Language like “optimized on the full dataset” — if everything was used to tune, nothing was left to validate.
  • Results that are suspiciously clean: tiny drawdowns, near-vertical equity curves, win rates in the 90s.
  • No mention of trading costs, and a long list of adjustable settings that “work best” at oddly specific values.

A good backtest is a hypothesis, not a promise. In-sample is where you form it; out-of-sample is what turns it into evidence — and even then, nothing guarantees the future. But if a tool can't or won't show you the out-of-sample number, you're not looking at an edge. You're looking at a very detailed memory of the past.

PaceAlgo is built on exactly this standard: validated out-of-sample, non-repainting, and every trade published — wins and losses alike.

PaceAlgo is an analytical and educational tool, not financial or investment advice. Trading involves substantial risk of loss. Past performance is not indicative of future results.