Home Live Ticker Fear & Greed
Sign in

Backtesting & Performance

How to Choose and Use a Benchmark for a Trading Strategy

Spot the edge. Swoop in.

A 12% annual return can be excellent or mediocre — and there is no way to tell which without knowing what a comparably-risky alternative would have produced over the same period.

Why Does a Strategy Need a Benchmark?

A strategy's return means little on its own without a benchmark that reflects its opportunity set and risk. The same 12% annual return can be excellent or mediocre depending on what a comparably-risky, comparably-exposed alternative would have produced over that same period — a return earned by holding cash-like instruments part of the year deserves a very different standard than one earned by a fully invested, leveraged portfolio.

Choosing a benchmark is not a formality performed after the real analysis is finished. It is part of deciding what question the backtest is even answering: better than what, and for what purpose.

Possible Benchmarks

There is no single correct benchmark for every strategy. The right choice depends on what the strategy is trying to do, what it is allowed to trade, and how much of the time it is actually in the market. The following candidates cover most situations, and more than one is often worth reporting side by side rather than picking just one.

Match the Benchmark to Exposure

The central point in benchmark selection is this: a strategy invested only half the time should not be judged solely against a permanently invested index without adjustment. Doing so compares two fundamentally different amounts of market risk and calls it a fair fight. Useful comparisons instead include:

Hypothetical example — for education only.

A strategy is invested 50% of the time on average and earns 10% annualized over the test period. A broad index, fully invested the entire time, also returned 10% annualized over the same period. Compared on raw total return, the two look identical. But the strategy achieved that 10% while carrying roughly half the market exposure — meaning it produced the same return using less risk, for the portion of the time it was actually in the market.

One simple, illustrative way to express this is return divided by average exposure: 10% ÷ 0.5 exposure = 20% "exposure-adjusted" return for the strategy, versus 10% ÷ 1.0 exposure = 10% for the fully invested index. By this framing, the strategy generated twice as much return per unit of market exposure as the index did. This is an illustrative adjustment meant to make the exposure difference concrete, not a universally standard or audited performance metric, and it should be read alongside volatility, drawdown, and correlation rather than in place of them — but it correctly captures the basic point: the strategy did more with less time in the market.

The same logic runs in the other direction, too. A strategy that is invested 100% of the time and slightly outperforms the index has arguably done less than this simple arithmetic suggests, once its exposure is accounted for, because it took the same amount of market risk as the benchmark to get there. Exposure matching is not a trick to make a strategy look better — applied honestly, it can just as easily make a seemingly strong result look ordinary once the amount of risk actually taken is factored in.

Compare More Than Total Return

A benchmark comparison built on total return alone hides most of what separates a durable strategy from a fragile one. The dimensions worth examining together include:

Downside capture is the percentage of the benchmark's decline that the strategy captured during periods when the benchmark fell. Upside capture is the percentage of the benchmark's gain that the strategy captured during periods when the benchmark rose. A strategy with low downside capture and high upside capture — losing less in declines and keeping pace in advances — is doing something a total-return comparison alone would never reveal, since two strategies can post the same overall return with completely different capture profiles.

Correlation and beta deserve particular attention because they answer a question neither return nor drawdown can: how much of the strategy's behavior is simply the benchmark in disguise. A strategy with a beta near 1.0 and high correlation to its benchmark is largely re-expressing the benchmark's own movements, whatever its stated logic. A strategy with low correlation and a beta well below 1.0 is doing something structurally different, and that difference is worth knowing regardless of which one produced the higher return in a given period, because it changes what the strategy is actually useful for inside a larger portfolio.

When Beating the Benchmark Isn't the Right Goal

A strategy may earn less than the index and still be valuable, if it materially reduces drawdown or is meant to diversify a larger portfolio rather than stand alone as the entire allocation. A strategy held specifically to be uncorrelated with an existing equity portfolio is doing its job even in a year it trails the index, provided it behaves differently when the rest of the portfolio is under stress.

Conversely, a strategy may earn more than the index only because it used greater leverage or accepted severe tail risk. A comparison that looks only at total return can end up rewarding exactly the wrong strategy — the one that simply took on more risk, rather than the one that used risk more efficiently.

Hypothetical example — for education only.

Strategy A returns 15% annualized with a maximum drawdown of 12% and no leverage. Strategy B also returns 15% annualized, but with a maximum drawdown of 45%, achieved using two times leverage on a similar underlying set of positions. A comparison based on total return alone treats these as identical successes. A comparison that includes drawdown and leverage shows they are not: Strategy A produced the same return while risking a quarter as much peak-to-trough loss and no borrowed capital, which is a meaningfully different, and arguably far better, outcome for the same headline number.

A benchmark comparison built only on total return could not distinguish Strategy A from Strategy B at all — both would simply show up as "matched the index" or "beat the index by the same margin." Only by carrying drawdown and leverage into the comparison does the difference between a strategy that used risk efficiently and one that used a great deal more of it become visible, which is exactly why total return should never be the only figure reported alongside a benchmark.

Common Benchmark-Selection Mistakes

Comparing a part-time strategy to a fully invested index without adjustment is the most frequent error, and it is covered above because it distorts nearly every other conclusion drawn from the comparison. But several other mistakes are common enough to call out on their own.

Choosing a benchmark after seeing which one makes the strategy look best is itself a form of data mining, applied to the comparison rather than to the strategy. If five candidate benchmarks are tested and the one the strategy beats by the widest margin is reported as "the" benchmark, the resulting comparison has been selected for a favorable outcome in exactly the same way an overfit parameter is selected for a favorable backtest — the benchmark should be chosen for its fit to the strategy's actual exposure and objective, decided before results are examined, not chosen afterward because of how the comparison turned out.

Ignoring that a strategy's investable capacity may be far smaller than the benchmark index's total market is another common gap. An index represents trillions of dollars of tradable capital; a strategy that only works at a fraction of that size is not automatically comparable to the index just because both operate in the same market. Capacity limits the audience for whom the comparison is even meaningful.

Comparing pre-cost strategy returns to a benchmark that itself has costs embedded, or vice versa, inconsistently, quietly tilts the comparison. A benchmark index return is often quoted before any real-world cost of actually holding it, while a strategy's return may or may not include commissions, spreads, and slippage. Both sides of the comparison should be measured on the same basis — either both net of realistic costs, or both before them, with the choice stated explicitly.

Statistical Significance of Outperformance

With limited historical data, even a genuinely skill-less strategy can beat its benchmark by chance over some period, and even a genuinely better strategy can trail its benchmark for a stretch through no fault of its own. Neither outcome, on its own, over a short window, proves much of anything.

The length and market-regime coverage of the comparison period matters as much as the headline outperformance number. A strategy that beat its benchmark over eighteen months spanning a single steady bull market has been tested far less thoroughly than one compared over a period that included a sharp decline, a recovery, and a quiet stretch in between — regardless of which one shows the larger outperformance figure on paper.

This is the same underlying caution that applies throughout backtesting: a single favorable comparison, examined in isolation, cannot separate a repeatable advantage from a fortunate stretch of history. The practical response is not to demand an impossibly long track record before drawing any conclusion, but to weight a comparison's conclusions in proportion to how much genuinely independent evidence it contains — more regimes, more distinct periods, and more time between the start and end of the sample all add real weight; a longer period that simply repeats the same market conditions over and over does not add nearly as much.

Common Mistakes

Limitations

No benchmark perfectly represents what a strategy would have done instead of trading its own rules. Every candidate — a broad index, a factor-matched portfolio, a volatility-scaled comparison — is an approximation of the alternative use of capital, not a precise reconstruction of it. Benchmark choice always involves judgment, and a benchmark that fits one strategy's risk profile may not fit another's, even if both strategies trade the same broad market. Treat the comparison as one useful lens among several, not as a final verdict.

It is also worth remembering that the benchmark itself is not a fixed, risk-free target. An index can be concentrated in a handful of large companies, can carry its own sector tilts, and can go through extended periods of poor performance that have nothing to do with the strategy being evaluated against it. Beating a weak benchmark during a weak period for that benchmark is a different achievement than beating a strong benchmark during a strong period, even when the reported outperformance number looks the same on both occasions.

Frequently Asked Questions

What is a good benchmark for a stock trading strategy?

A good benchmark reflects the same market exposure, risk, and opportunity set the strategy actually operates in, not simply the most familiar market index. A strategy trading a narrow universe of small-cap stocks is better measured against a small-cap index or its own universe held equally than against a broad large-cap benchmark it was never trying to resemble.

Should I compare my strategy to a fully invested index if it isn't always in the market?

No. A strategy that is only invested part of the time carries less market exposure than a permanently invested index, so comparing raw total returns compares two different amounts of risk. Adjusting the comparison, for example by blending the index with cash to match the strategy's average exposure, gives a fairer picture of whether the strategy added value.

What is downside capture?

Downside capture is the percentage of the benchmark's decline that the strategy captured during periods when the benchmark fell. A downside capture below 100% means the strategy lost less than the benchmark during its declines; above 100% means it lost more.

Can a strategy be good even if it underperforms its benchmark?

Yes, depending on the goal. A strategy that earns less than its benchmark can still be valuable if it materially reduces drawdown, diversifies a larger portfolio, or is intended to preserve capital rather than maximize return. Beating a benchmark is only the right goal when outperformance is actually the objective.

Is beating a benchmark once meaningful?

Not on its own. A short comparison period, or one that only covers a single type of market condition, does not provide enough evidence to distinguish genuine skill from chance. A meaningful comparison needs a sufficient sample size and coverage of more than one market regime.

Related Guides