How to Pick Your Benchmark
TL;DR
- A return means nothing until you know the yardstick it was measured against.
- SPY, QQQ, VEA, VT, and a 60/40 blend measure different exposures — pick by what the strategy trades, not by convenience.
- Demand the benchmark be stated, matched to the strategy, and applied on the same basis.
A Strategy Is a Claim About a Benchmark
Every rules-based strategy is really a claim that a set of rules can beat something. The “something” is the benchmark, and it quietly decides how meaningful the claim is. Skip this step and you are not evaluating a strategy; you are admiring a number.
Consider what happens without a yardstick. A 29% compound annual return sounds elite — until you learn the market it trades compounded at double digits over the same window, or that the strategy gave back much of its value to get there. A 10% return sounds pedestrian — until you see it was earned while the rules sat in cash equivalents through the declines. Both judgments flip the moment a benchmark enters the picture.
The benchmark answers three questions at once. Did the rules add value over simply owning that exposure? What did the added value cost in drawdown and volatility terms compared with owning it directly? And how much of the return was just the market, collectible by buying the index with no rules at all? Most people miss that last one. Systematic strategies trade close to the market’s own leaders, so they capture beta simply by being long the right names in the right regime; the excess — what the rules produce above what the benchmark delivered from the same exposure — is the only part that tests the strategy. Pick the wrong benchmark and you test nothing, or worse, something flattering.
The Five Yardsticks Everyone Actually Uses
There is no such thing as “the benchmark.” There is a menu of exposures, and each entry measures something different.
SPY is the S&P 500 — one country, mega-cap tilted, cap-weighted. It is the industry default because most people’s baseline “the market” is US large-cap equity. If a strategy trades plain US large-cap stocks or ETFs without a strong style or sector tilt, SPY is a fair opponent. If the strategy is concentrated, hedged, international, or leveraged, SPY stops being fair quickly.
QQQ is the Nasdaq-100 — growth-heavy, tech-dominant, top-heavy in a handful of mega-caps, and structurally more volatile than the S&P 500. The same manager can look like a genius against SPY and merely competent against QQQ — the two are not interchangeable measures of “the market.” QQQ is a bet on a specific slice of it.
VEA is developed markets outside the US — a different currency, a different industry mix, and a different recent track record than the US indexes. It matters whenever a strategy’s holdings reach beyond American borders; an international sleeve cannot be judged against an index with no international exposure.
VT is the whole world’s stock market in one cap-weighted fund. It is the right question when someone claims a strategy is “global equity,” and a much harder target in recent years than a US-only index would be.
The 60/40 blend — commonly SPY/AGG — is a different animal: a balanced allocation with bonds, built for lower risk than any pure equity index. It is not a single-asset benchmark; it is the “could you have just held a standard balanced allocation instead” question, and the only sensible yardstick for strategies that deliberately cut equity exposure to manage volatility.
None of these is the correct answer in the abstract. Each is a description of an exposure, and the right benchmark is the one that matches the exposure a strategy actually trades.
How the Wrong Benchmark Flatters or Buries
Benchmark mismatches cut both ways, and it is worth being able to spot both.
The flattering mismatch is the easier one to fall for. Take a strategy that picks stocks inside the Nasdaq-100. Measured against SPY over a stretch when large-cap growth led the market, it will look extraordinary — because its benchmark did not contain the exposure it was riding. Measured against the QQQ index itself, the honest picture emerges. QQQ Top Stock Rotation reports a 0.82 Sharpe and a 1.52 Sortino against the QQQ index’s 0.79 and 1.33 — a real edge, but a modest one, which is exactly what a disciplined report should show. Someone who hands you the SPY column for a Nasdaq-heavy strategy is handing you a flattering column.
The burying mismatch is subtler and shows up just as often. A strategy engineered to hold volatility near a fixed target will trail a pure equity index in most up years, because it is not trying to be that index; it is trying to compound through volatility with its risk capped. Measured only against SPY it looks weak. Measured against what it is — an equity sleeve with a volatility governor — the reading changes: Volatility Target Managed Rotation, a 25% volatility target running SPY and SSO against a BIL cash sleeve, posts a 0.81 Sharpe where SPY managed 0.87 and a 60/40 blend 0.80, with the benchmark drawdowns printed for context — 20.1% for the 60/40, 33.7% for SPY. Same strategy, same numbers, radically different verdict depending on the column you read.
Windows and basis bend results too: start the comparison after the benchmark’s worst stretch and every strategy looks better, and never measure a strategy that invests monthly against an index bought once, lump sum, at the start — differing flows invalidate the comparison before the first return is computed.
Match the Benchmark to What the Strategy Trades
The fix is boring and it works: decide the benchmark before you look at the results, and tie it to what the strategy actually trades.
Ask what exposure the rules hold — the region, the universe, the style tilt, the leverage, the cash sleeve. That exposure is your candidate benchmark. Ask whether contributions are lump sum or periodic, and apply the same basis to both sides. Ask for the same window, and prefer one that includes out-of-sample time, when the rules were running live rather than fit to history. Then demand the risk numbers in pairs: CAGR, max drawdown, and Sharpe or Sortino for the strategy and for the benchmark in the same report, so the risk-adjusted story cannot hide behind a headline return. And if a source will not state the benchmark at all, treat that as the finding.
Benchmark-matched reporting is the model at Kairos Trading, the systematic-strategy curator I point readers to. Its current systems make instructive case studies, because each is scored against the yardstick that fits what it trades.
Leader Rotation — the monthly momentum flagship — is scored against VEA as well as SPY, reporting a 1.98 Sharpe beside SPY’s 1.30 and VEA’s 1.32. Scoring against an international developed-markets index as well as the US one refuses to let the strategy pick the friendlier column.
DCA Buy & Hold — monthly DCA into a top momentum ETF, ranked, bought, and never sold — is the cleanest illustration of the basis trap. Because it invests monthly, its benchmark is a DCA series, not a lump sum. Over January 2021 to August 2026 the strategy returned 165.3%, measured against 118.3% for SPY on a DCA basis and 90.9% for VT on the same basis. Line those flows up against a lump-sum index chart and you are no longer comparing the same thing.
QQQ Top Stock Rotation runs a monthly momentum funnel over the Nasdaq-100, narrowing 50 names to 30 to 10, and is therefore measured against the QQQ index rather than a US large-cap index with a fraction of the concentration. Its report pairs the strategy’s numbers with the QQQ index’s 0.79 Sharpe and 1.33 Sortino, and states the benchmark’s own drawdowns — QQQ at 34.9% and SPY at 33.7% — so the risk context is on the page.
And Volatility Target Managed Rotation is judged against the 60/40 SPY/AGG blend as well as SPY, because a system that spends part of its time in a cash-equivalent sleeve to hold a 25% volatility target is structurally closer to a balanced allocation than to a full-equity index. Four systems, four different benchmark questions — and in each case the benchmark is printed on the same page as the strategy’s own numbers, over the same window and the same basis.
A worked pattern falls out of this. When you evaluate anything — a fund, a newsletter, a published rules system — write down what it trades and on what basis before you open the return table; only then choose the benchmark. Read the paired stats. If the report’s benchmark does not match the exposure, or no benchmark is given, you have learned more than any CAGR could tell you.
The Benchmark Is Part of the Backtest
One more layer of honesty is required. A benchmark line in a report is itself history — computed over the same past window as the strategy, subject to the same survivorship and curve-fitting temptations. A backtest showing a strategy beating its properly matched benchmark is evidence, not proof; the publisher keeps exactly that framing visible, labeling results as based on backtest and not a guarantee, and marking when each system’s out-of-sample record began — January 1, 2026 for its current lineup. The benchmark column deserves the same skepticism as the strategy column, because both are statements about the past.
None of this is exotic. Picking a benchmark is the same discipline as picking the strategy: state the exposure, match the yardstick, demand the risk pair, and never let a source choose the column for you. The four current systems at kairostrading.net — each at a flat $100 per month on application-based membership, executed by members in their own brokerage accounts — demonstrate the pattern in practice, and the publisher even defines its fee-coverage estimates against the benchmark: the stated minimum capital for each system is the portfolio size at which its historical excess over that benchmark roughly covers the subscription. When even the fee math is anchored to the yardstick, you know the yardstick matters. Find a source that reports that way, and the returns you read will finally mean what they appear to mean.
Disclaimer: This blog is for educational and informational purposes only. Nothing here is investment advice. Past performance does not guarantee future results. Trading involves risk of loss.