skip to content

A trading rule's backtest Sharpe is 2.5 in-sample but 0.3 the following year. What happened?

level: seniorimportance: should knowfreq 45%

answer

  1. the number was chosen, not measured
  2. count what was tried
  3. maximum of noisy statistics
  4. search spends degrees of freedom
  5. regime change is the alibi, not the cause

basics

~20 s

The in-sample Sharpe was never a forecast: it is the best of many rules scored on the very history that selected them. A maximum over noisy statistics is biased upward, so the winner's number was inflated from the start.

solid answer

~40 s

This is the optimism of training error wearing a different metric. The 2.5 was computed on the same price history used to pick the rule, its thresholds, its universe and its holding period, so whichever variant was luckiest on that history is the one that got selected. A maximum over many noisy statistics is biased upward, and the more variants tried, the bigger the inflation — whether or not anyone counted them. Out of sample the luck does not repeat and only the real edge survives, here about 0.3. So ask how many rules, settings and filters were evaluated on that history: that count is the effective degrees of freedom of the search and it sets the optimism. An honest number needs history that played no part in building or choosing the rule.

code

python · 15 lines
python
import random, statistics

random.seed(0)
n, p = 200, 50
y_now = [random.gauss(0, 1) for _ in range(n)]    # target is pure noise: nothing to learn
y_next = [random.gauss(0, 1) for _ in range(n)]   # same rows, fresh outcomes
columns = [[random.gauss(0, 1) for _ in range(n)] for _ in range(p)]

# 'fitting' here is just selection: keep the column that fits the data we then score on
winner = max(columns, key=lambda c: abs(statistics.correlation(c, y_now)))

in_sample = statistics.correlation(winner, y_now) ** 2
out_sample = statistics.correlation(winner, y_next) ** 2
print('in-sample fit of the winner:', round(in_sample, 4))
print('same column, fresh outcomes:', round(out_sample, 4))

go deeper

for a junior

Remember the one-line version: a result measured on the data used to choose it is not a forecast, so ask which history the number came from before you believe it.

for a middle

Be ready to explain why picking the best of many candidates on one dataset inflates the winner's score, and that the effect is about selection rather than about the metric chosen.

for a senior

Show the diagnosis you would run: count the configurations evaluated, probe the result's fragility to small perturbations, and check whether the performance is concentrated in a few periods before you accept it.

for a principal

Own the process rule that prevents this — who is allowed to see which data, when a specification is frozen, and how much shrinkage is assumed by default before any number reaches a committee.

## In-sample risk versus out-of-sample risk The optimism of training error is usually taught with squared error and a fitted regression, which hides how general it is. It applies to **any performance number computed with the same data that shaped the thing being scored**, and to any metric: accuracy, AUC, lift, revenue per user, or a Sharpe ratio. The two quantities in play here are: - **In-sample risk**: how the rule performed over the history used to construct it. - **Out-of-sample risk**: how it performs on data that took no part in its construction — the only number that forecasts the future. A backtest Sharpe of 2.5 that becomes 0.3 is not a mystery or a betrayal; it is the expected behaviour of an in-sample statistic reported as if it were out-of-sample. ## Where the degrees of freedom came from The rule may have only two visible parameters — say a lookback window and an entry threshold. That is a misleading count. Everything that was tried and discarded on the same history spends freedom: - every value of the lookback that was swept, - every threshold, stop-loss and holding period tested, - every choice of universe, sector filter or de-listing rule, - every earlier idea abandoned because its backtest looked poor, - every re-run after "fixing" the data in a way that improved the curve. Each of these is an evaluation of a noisy statistic on one finite history. If you evaluate 300 variants, and the true edge of all of them is zero, the *best* of the 300 will still show a healthy positive Sharpe purely from sampling variation — and it is the best one you will report, because that is how you chose it. Selection is a fitting procedure. The reported number is a maximum, and the expectation of a maximum sits well above the expectation of any single draw. This is the same mechanism as adding free parameters to a regression: more freedom exercised on one dataset, more of that dataset's noise absorbed into the answer, larger gap between the measured number and the truth. ## The give-away symptoms When you inherit a result like this, look for the fingerprints of an in-sample number: - **Nobody can state how many variants were evaluated.** If the search is uncounted, the optimism is unbounded. - **The rule is fragile to small perturbations.** Shifting the entry threshold or the window by one step collapses the result — real edges are usually a plateau, fitted noise is usually a spike. - **Performance is concentrated in a few periods.** A handful of episodes carrying the whole result is what a noise-fitted rule looks like. - **The final specification includes clauses whose only justification is the backtest** — an excluded month, a filtered sector, a rounded threshold. ## What to do instead The fix is conceptual before it is procedural: *decide which data is allowed to influence the rule, and never score on that data*. In practice that means holding back a stretch of history that plays no part in construction, freezing the specification, and only then measuring — and if you go back and change the rule after seeing that number, it has become in-sample too and is spent. Forward performance after freezing is the strongest evidence, because it cannot have been searched over at all. It also means budgeting the search. If you know you will try many variants, expect the winner's score to be inflated and discount it before making any decision. The honest headline is not "this rule earns 2.5" but "the best of 300 variants earned 2.5 on the history that selected it, which is consistent with no edge at all". ## A second, separate explanation to acknowledge A real regime change in the following year is a distinct cause with a distinct signature, and a good answer mentions it without hiding behind it: if the world genuinely changed, other unrelated strategies degrade at the same time and the rule's assumptions can be named and checked. In practice the search-driven explanation is far more common, and "the market changed" is the standard way a fitted-noise result avoids a post-mortem. Rule out the optimism first, because it is the explanation you can quantify by counting what was tried. ## Interview register Do not lead with the market. Lead with the count of things tried on the same history, connect it to the optimism of an in-sample statistic, note that the effect applies to Sharpe exactly as it applies to squared error, and finish with the discipline you would impose: a slice of data that never touched construction, a frozen specification, and a stated expectation of shrinkage before anyone sees the number.

  • The author says the rule has only two parameters, so how can it be overfitted?
    Because the visible parameters are not the freedom that was spent. Every window, threshold, filter and abandoned idea evaluated on that history is a degree of freedom, even though none survives in the final specification. Ask how many configurations were run in total; that number, not the parameter count, sets the expected inflation of the reported Sharpe.
  • How would you get a number you could actually defend to a risk committee?
    Fix the specification, then measure it on a stretch of history that took no part in building or selecting it — and treat that number as spent the moment anyone tweaks the rule in response to it. Strongest of all is performance measured forward after the rule is frozen, since no search could have reached that data.
  • Does the same optimism affect a metric like AUC or precision, or only squared error?
    Any of them. The bias comes from computing the statistic on data that shaped the model or the choice, not from the loss function. An AUC selected as the best of forty candidate feature sets is inflated for exactly the reason the backtest Sharpe is, and it shrinks on data that played no part in the selection.

Run a hundred coin-flippers, keep the one who flipped twelve heads out of fifteen, and report her hit rate as a skill estimate. Next week she reverts to half, and nothing about the coin changed.

saying these in an interview costs you the question

  • Blames a regime change without counting the variants tried
  • Reports the winning backtest number as the expected return
  • Thinks optimism affects error metrics but not Sharpe or AUC
  • Adds parameters to repair the out-of-sample collapse
  • Counts only the two visible parameters as model complexity

context