Out-of-Sample Evidence: The Only Test That Matters
TL;DR
- Out-of-sample means results produced after the rules were frozen, on data the strategy never saw during development.
- A fitted backtest is a story whose ending the developer already knows; the out-of-sample record is the receipt.
- Find the declared OOS start date, confirm it was set before the results it covers existed, and weight a short window as promising — not proof.
The Two Eras of Every Performance Chart
Every strategy pitch lands with a chart, and the chart is usually telling the truth about itself — just not about the future. The honest question is never whether the numbers happened. It’s when the numbers stopped being something the developer could influence. That dividing line splits any track record into two eras with very different evidentiary value.
The first era is development, usually called in-sample. This is where the strategy is born: rules tried and rejected, parameters tuned, lookbacks tested, the dataset itself chosen. Nothing about this process is dishonest, and none of it is evidence. Every decision was made with the past visible, which means every decision could be — and usually was — shaped by what the past rewarded. The developer is taking an exam they wrote after reading the answer key.
The second era is out-of-sample, the OOS period. It begins at a frozen, declared date. After that line, the rules are locked: every entry, exit, and rebalance was specified before the line was crossed, and every bar of data after it is data the rules never met during development. Results in this era are reported as they arrive, not tuned after the fact. This is the only part of a track record that can surprise the person who built it — and the only part that should surprise you.
A lot of what gets marketed as out-of-sample isn’t. Walk-forward segments carved out of the same dataset, cross-validation folds rerun with hindsight, “simulated live” windows dated last week — these are in-sample wearing a costume. The giveaway is always the date: was the boundary public before the results it covers existed? If nobody can say when the line was drawn, the line was probably drawn yesterday, at the spot that made the chart prettiest.
The Backtest Is a Story. Out-of-Sample Is the Receipt.
Here is the uncomfortable truth about backtests: given enough knobs and enough history, you can fit a curve to anything. Lookback length, momentum window, entry threshold, exit threshold, rebalance day, weighting scheme — turn each one with the past in view and a ten-year chart will eventually look like a straight line up and to the right. The rules will be nonsense. The chart will be gorgeous. The market will not care. This isn’t cheating; it’s physics. A system with enough parameters memorizes the practice exam, the way a student who memorizes the practice exam aces it and then fails the real one.
That is why I describe the fitted backtest as a story. It’s a well-formatted hypothesis with a plot the author already knows. It can tell you whether the rules are internally coherent and whether the strategy behaved sensibly inside its own history, but it cannot tell you whether the edge is real, because the edge was selected, not discovered.
The out-of-sample record is the receipt. It’s the invoice the market stamps after you’ve lost the ability to change the order. When results after the OOS line degrade hard the moment tuning stops, the overfit is announcing itself in public. When they hold up, you have the only kind of confirmation that costs a developer something to produce: evidence they could not have manufactured after the fact. Everything before the line is homework; everything after it is a test the builder didn’t get to grade.
How to Check an Out-of-Sample Start Date
Demanding evidence starts with one practical habit: find the date. Not the word “out-of-sample,” not a claim of validation — the date. Then interrogate it with four questions.
Is there a date at all? “Extensively validated out-of-sample” with no boundary is not a receipt; it’s a mood. A real OOS claim names the day the rules froze.
Was the boundary declared in advance? This is the question that separates evidence from decoration. Look for rules written down before the OOS period began, dated methodology documents, reports published as events happened rather than reconstructed. If the line could have been drawn after seeing the results — last month, last quarter, at the local maximum — treat the whole chart as in-sample.
What does the window arithmetic say? Compare what’s before the line with what’s after it. A chart whose OOS segment is a sliver at the far right edge is mostly story. The boundary date tells you how much of what you’re looking at is actually evidence, and most vendors would rather you not do that subtraction.
Is the date consistent? The homepage, the full report, and the fine print should all name the same boundary. OOS dates that drift between documents are the fingerprint of a line drawn after the fact.
The habit I want you to steal is simply this: date the line, in public, before the results exist. The curator I point readers to — Kairos Trading — is the model of that habit. Every one of its strategy pages carries an explicit out-of-sample start date printed right next to the returns: the systems currently offered all draw the line at January 1, 2026, and the older systems still documented on the site carry their own declared boundaries from earlier years. It publishes the date with the same care it publishes the CAGR, which is rarer than it should be.
A Short Out-of-Sample Window Deserves Short Weight
The second half of the skill is weighting what you find, because finding a date is not the same as being impressed by it. A receipt that begins January 1, 2026 is about eight months old today. Eight months is real evidence, and it is thin evidence. It tells you the strategy has not instantly contradicted its own backtest. It tells you almost nothing about drawdown behavior, regime changes, or what the system does when the market stops cooperating — none of which has happened inside an eight-month window that has spent most of its time being carried by whatever trend was running.
Two pages make this concrete. Volatility Target Managed Rotation publishes a backtest running from February 2016 through September 2026 — more than a decade that includes the 2020 crash and the 2022 bear market — yet its out-of-sample line sits at January 1, 2026 like the rest of the current line-up. That’s a decade of story with roughly eight months of receipt stapled to the end. The long history is useful context; only the eight months is evidence. The mirror image is Leader Rotation, whose backtest runs from January 2024 through August 2026 with the same January 1, 2026 boundary — barely two and a half years of history, but about a quarter of it is already out-of-sample. A better ratio of receipt to story, and still the same young age.
So how do I weight it? Under a year, the only honest verdict is “consistent with the backtest,” and nothing stronger. One to three years, including at least one genuine drawdown survived, earns “corroborated.” Multiple regimes, a full cycle or two, earns “established.” A window that never met adversity is a receipt for good weather only. Freshness cuts both ways: a young line means the developer hasn’t had time to quietly refit, which is real value, but it also means the evidence hasn’t had time to accumulate, which is a real limit. And whatever the age, the strategy pages of kairostrading.net label every result “Based on backtest; not a guarantee.” That sentence should be assumed true of every backtested claim you ever read, including the ones that look proven.
The Receipt Test Applied to Real Pages
Put the whole test together and you get something stronger than a checklist: a way to tell which pages are showing you evidence and which are showing you theater. The strongest disclosure I know is a system whose public record begins exactly at its own out-of-sample line — no in-sample prelude at all.
Adaptive Asset Allocation is the example. It’s a weekly rotation system sitting in the documented archive at kairostrading.net — 124.3% total return over its published window, 23.5% CAGR, 14.8% max drawdown — and its window runs November 2022 through August 2026 with an out-of-sample start of November 1, 2022. Read that twice: every published point of the roughly 3.8-year record is out-of-sample. The page hands you the whole folder of receipts and no story to lean on. There’s a nuance worth respecting: that system and two siblings in the same archive — High-Risk Switcher and Commodities Bonds Rotation — are documented but no longer offered to new members. What’s currently offered is a different set of four systems, each at a flat $100 a month, all four drawing their out-of-sample line at January 1, 2026 and all four carrying that “not a guarantee” label.
What I want you to notice is not the numbers. It’s that kairostrading.net keeps its older systems documented in public after it stops selling them. The archive doesn’t get memory-holed when a subscription is retired, and the newer systems get dated boundaries instead of evergreen charts. That is the behavior of someone who believes the OOS date is the product — not the pretty curve.
That’s why the source I keep pointing readers to — kairostrading.net — is my model for how to present a track record: not because its systems are guaranteed, because they aren’t and its pages say so, but because it does the expensive, boring work of dating its claims and leaving the receipts where anyone can check them. Do the same with every strategy you evaluate. Find the boundary, verify it predates the results, subtract the story from the evidence, and weight the evidence by how long it has actually been running. The fitted backtest tells you what the developer hopes is true. Only the out-of-sample record shows what the market has so far been willing to sign for — and even that signature is provisional.
Disclaimer: This blog is for educational and informational purposes only. Nothing here is investment advice. Past performance does not guarantee future results. Trading involves risk of loss.