If you have ever built a strategy that printed money on historical data and then quietly bled out in live trading, this article is for you. Not necessarily because the code was wrong. The zero-edge control measures how good a backtest can look when there is nothing there, which is the comparison a live result never gets.
Everyone in this corner of the internet tells you to run robustness tests. Split your data. Walk it forward. Run a Monte Carlo. Check your parameter stability. What almost nobody does is measure whether those tests work — whether a strategy that passes them actually does better afterwards than one that does not.
So I measured it. 5,472 parameter backtests across 8 markets and 3 strategy families, plus 5,000 randomly generated rules mined on SPY and re-mined on data engineered to contain no edge at all. Then I scored 11 robustness checks by how well each one predicted what happened in a slice of history the checks never saw.
Some of them are worth every minute. One of them is actively misleading. And the whole exercise selects for something different from what you think it selects for.
Then I did the thing that should come first, and it cost several of the conclusions below. Eight markets sounds like breadth. Measured properly, those eight were worth 2.87 independent bets (SPY and ES are one index in two wrappers), so I reran the headline experiments on 53 markets and 36,252 backtests. The main finding got stronger. Several of the individual numbers below were overstated by a factor of three, and two checks invert on some asset classes: the parameter-neighbour check turns negative in FX and rates, and the Sharpe-0.5 check in grains and softs. That section is near the end, and it is the most useful part of this article.
TL;DR — the answer box
- The noise floor is high, and almost nobody accounts for it. Optimising 384 ordinary RSI variants on three years of shuffled, drift-free data (where an edge cannot exist) produced a median best backtest of 9.2% a year at Sharpe 1.17. On twenty years of the same data it produced 3.1% at Sharpe 0.46. Your backtest is only as impressive as its margin over the floor for a search of the same family, length and size.
- Mining 5,000 rules on real SPY found an 11.2% CAGR strategy. Mining the same 5,000 rules on zero-edge data found a 6.4% one. The real effect is the difference — 4.8% a year — not the headline.
- Keeping the in-sample winner rarely blows up. It just rarely wins. Across 24 market-strategy grids the in-sample champion stayed profitable out-of-sample in 19 of 24 cases, but beat plain buy-and-hold in only 9 of 24.
- Walk-forward delivers about four-tenths of its promise. Median walk-forward efficiency across 471 rolling windows: 0.42. It beat buy-and-hold in 6 of 24 market-strategy combinations.
- On the original eight markets, robustness testing selected for survival, not outperformance. Variants passing at least 10 of 11 checks were profitable in the untouched final quarter 98.0% of the time versus a 63.7% base rate, while the share of them beating buy-and-hold was lower: 20.7% against a 27.8% base rate.
- The stack holds on 53 markets. The individual checks do not. Retested on 36,252 backtests, variants passing all 11 checks still survived 97.2% of the time (measured before the breadth fix described below), but the parameter-neighbour check drops from +0.598 lift to +0.166, and a development Sharpe above 0.5 collapses from +0.370 to +0.041.
- Where you test decides whether the checks work at all. The neighbour check is worth +0.366 on equity index markets and −0.210 on FX, where variants that pass it survive 29.5% of the time against a 45.6% base rate. On FX, a pass went with worse odds of survival, not better.
How I tested
Eight markets, all daily bars, all frozen local files: SPY, QQQ, IWM and DIA on the equity side, and ES, NQ, NG and JY as continuous back-adjusted futures. SPY runs from 1993-02-02 to 2026-06-12; the shortest series, the futures, start in January 2007. The last data point anywhere in the study is 2026-07-02.
Three strategy families, picked as a sample of what retail traders optimise rather than for being clever:
| Family | Rule | Parameters swept | Variants |
|---|---|---|---|
| RSI mean reversion | Buy when RSI drops below a threshold, hold N days | RSI period, threshold, hold | 384 |
| Moving-average trend | Long while the fast average is above the slow one | fast, slow, confirmation days | 132 |
| Donchian breakout | Buy an N-day high, exit on an M-day low | lookback, exit, trend filter | 168 |
684 variants per market × 8 markets = 5,472 backtests. Signals are read at the close and entered at the next open, so nothing sees a price it could not have traded on. Long only, one unit, no leverage, frictionless unless a cost is explicitly charged. Every variant is measured in-sample (the first 70% of each market's bars), out-of-sample (the last 30%), and over the full sample.
Two things make this more than another grid sweep.
A control group. Every experiment that measures luck is re-run on bootstrapped versions of SPY's own bars, with the order shuffled and the drift removed. Same fat tails, same daily return distribution, no sequence and no upward pull. Buy-and-hold expects zero there, and no rule can have an edge by construction. Whatever a backtest earns on that data is what luck alone produces.
A held-back final quarter. For the scorecard, every check is computed on the first 75% of each market's history and scored on the last 25%, which the checks never touched.
Every number below is computed by the study's own scripts (research/study.py, plus research/threshold_25.py for the 25% rule and research/wide_universe.py for the 53-market rerun) and written into the study's facts sheet.
Does keeping the in-sample winner actually ruin the strategy?
Not in 19 of the 24 cases tested, and this is the first place the folklore is wrong.
I optimised each of the 24 market-strategy grids on the in-sample period, kept the single best variant by return, and carried it unchanged into the out-of-sample period.
| What happened to the 24 in-sample winners | Result |
|---|---|
| Median in-sample CAGR | 6.3% |
| Median out-of-sample CAGR | 7.5% |
| Still profitable out-of-sample | 19 of 24 (79.2%) |
| Beat buy-and-hold out-of-sample | 9 of 24 (37.5%) |
| Median out-of-sample rank inside its own grid | 88th percentile |
| Median rank correlation, in-sample vs out-of-sample | 0.48 |
Read the third row and the fourth row together, because that is the whole story of this article. The optimised strategy usually keeps making money. It usually loses to owning the index.
The winners that did break, broke where you would expect: on markets with no underlying tailwind. Japanese yen futures, RSI mean reversion: 1.7% in-sample became -7.5% out-of-sample, landing at the 10th percentile of its own grid. Natural gas trend and breakout picks lost money in both halves.
There is a catch buried in the method, though. The 70/30 split is itself a choice, and it is not a stable one. Moving the in-sample share from 50% to 70% changes which variant you would have kept in 62.5% of the 24 cells. At 90/10 it changes 91.7% of them. You did not discover the best parameters. You discovered the best parameters for the split you happened to pick.
What does a great backtest look like when there is nothing there?
This is the test almost nobody runs, and it reframes everything else.
Take a bathroom scale that swings five pounds either way. Step on it, read three pounds down, and celebrate — that is a backtest without a noise floor. You are not measuring weight loss. You are measuring the scale.
So I built the scale. Shuffled SPY's own bars, removed the drift, and ran the identical parameter grids on the result. No sequence, no trend, nothing to find. Then I recorded the best variant each search produced.
| Backtest length | Variants tested | Median best CAGR | Median best Sharpe | 1-in-20 best Sharpe |
|---|---|---|---|---|
| 3 years | 384 | 9.2% | 1.17 | 1.78 |
| 5 years | 384 | 6.6% | 0.91 | 1.42 |
| 10 years | 384 | 4.6% | 0.67 | 1.02 |
| 20 years | 384 | 3.1% | 0.46 | 0.70 |
A Sharpe ratio of 1.17 over three years, from a strategy with zero edge, found by testing a completely ordinary 384-variant grid. That is the median result — half the searches did better, and one in twenty produced a Sharpe of 1.78.
The mechanism is just arithmetic. On the same zero-edge data, testing one RSI variant returns a median Sharpe of 0.00. Testing all 384 returns 0.36. Every extra combination is another lottery ticket, and you keep only the winning one.
This is not a StatOasis invention. It is the same problem Bailey, Borwein, López de Prado and Zhu formalised in The Probability of Backtest Overfitting, and the reason Harvey, Liu and Zhu argued that a newly discovered factor should clear a t-statistic of 3.0 rather than the usual 2.0. The numbers above are what that abstraction looks like on a chart of SPY.
What happens when you search 5,000 rules instead of 384?
A parameter grid is the mild version of searching. The real version — what a strategy generator, a spreadsheet marathon or an AI assistant does — is searching over rules.
So I built a pool of 221 ordinary technical conditions: oscillator levels, moving-average positions, breakouts, consecutive up and down days, volatility regimes, and seven deliberately meaningless calendar conditions like "it is a Tuesday". Then I combined them at random into 5,000 rules of one to three conditions each, gave each a fixed holding period, and backtested every one on SPY.
The winner was RSI(2) < 15, hold 5 days: 11.2% a year at a Sharpe of 0.79 on 339 trades. Out-of-sample it did 9.6%. That is a genuinely good result, and it survived.
Now the same 5,000 rules, mined on the zero-edge data:
| Real SPY | Zero-edge data | |
|---|---|---|
| Best in-sample CAGR of 5,000 | 11.2% | 6.4% (median across 10 datasets) |
| Best in-sample Sharpe | 0.79 | 0.65 |
| Best single result seen | — | 10.4% |
| What the winners did next | 9.6% | -1.2% (median across 10 datasets) |
Search hard enough on data with no edge and you will find a 23-year backtest returning 6.4% a year at a Sharpe of 0.65. One of the ten runs found 10.4%. If someone showed you that equity curve without telling you where it came from, you would fund it.
The honest read of the SPY result, then, is not "11.2%". It is 11.2% minus 6.4% — the real effect is worth about 4.8% a year of CAGR and 0.13 of Sharpe over what searching that hard produces from nothing.
Two more findings from the mining run, one reassuring and one brutal:
- The junk conditions did not win on real data. Calendar conditions made up 7.2% of the rule pool and 0.0% of the top 25 rules on SPY. Not one calendar condition made the top 25. On SPY, where a real mean-reversion effect is present, this 5,000-rule search put no junk calendar condition in its top 25. On the zero-edge data, junk conditions appeared in a median 4.0% of the top 25 — the search had nothing better to grab.
- Beating the market is the rare part. Of the 5,000 rules, 41.5% were profitable in-sample, and 84.5% of those stayed profitable out-of-sample. But only 3.9% beat buy-and-hold in-sample, and only 0.2% — 10 rules out of 5,000 — beat it in both periods. SPY itself returned 6.9% in the in-sample window and 13.4% out-of-sample, price only, no dividends.
Is your in-sample champion just the luckiest one?
There is a formal way to ask that question, and it does not need you to guess. Combinatorially symmetric cross-validation, from the Bailey et al. paper above, chops history into 10 chunks, forms every one of the 252 ways to split them into an in-sample half and an out-of-sample half, and asks how often the in-sample champion lands below the out-of-sample median. That share is the probability of backtest overfitting, or PBO. 50% means your selection carries no information whatsoever.
| Strategy family | Median PBO | Range |
|---|---|---|
| RSI mean reversion | 19.2% | 6.7% – 40.1% |
| Donchian breakout | 44.8% | 28.6% – 62.7% |
| Moving-average trend | 71.4% | 44.8% – 96.4% |
The median across all 24 grids was 43.7%, and 7 of 24 scored above 50% — selection that was actively worse than picking at random.
Look at the spread between families, because this is the finding I did not expect. The identical test, run with identical code on identical markets, says RSI mean reversion selection is broadly informative and moving-average trend selection was worse than random selection in this sample: a median PBO of 71.4%, where 50% means the selection carries no information. There is no such thing as "my process is robust". There is only "this effect held up, and that one did not".
What does walk-forward actually deliver?
Walk-forward analysis is the method every guide calls the gold standard: optimise on a window, trade the next window unchanged, roll forward, repeat. I ran it properly — 4 years in-sample, 1 year traded, stepped annually, 471 windows across all 24 market-strategy combinations.
| Walk-forward result | Value |
|---|---|
| Median walk-forward efficiency | 0.42 |
| Median share of windows profitable | 69.3% |
| Beat buy-and-hold on CAGR | 6 of 24 cells |
| Median share of windows that beat buy-and-hold | 28.6% |
| Distinct "best" parameter sets chosen on SPY | 17 across 29 annual windows |
Walk-forward efficiency is the traded year's return divided by the optimised return that justified choosing those parameters. The median is 0.42. You keep about four-tenths of what the optimiser promised — which is exactly what walk-forward is for, since it measures that on held-back history instead of after you have funded it.
The parameter churn is the underrated number. On SPY, the optimiser picked 17 different "best" parameter sets across 29 annual windows. Whatever the optimiser found each year, it did not stay found.
And walk-forward is not a way to beat the market. It beat buy-and-hold in 6 of 24 cells — and those wins are concentrated in the two markets that fell over the period, natural gas and the yen, where being out of the market most of the time was the entire edge.
Should you pick the peak or the middle of the plateau?
Every robustness guide tells you not to pick the spike on the parameter surface — pick the middle of the broad flat region, because a peak surrounded by cliffs is an accident. I have written that advice myself. So I tested it: for each grid, the outright in-sample best against the variant with the best average result across its immediate parameter neighbourhood.
It did not reliably improve out-of-sample results.
| Peak vs plateau | Result |
|---|---|
| Plateau pick beat peak pick out-of-sample | 9 of 24 cells |
| Median difference in out-of-sample CAGR | 0.00% |
| Median out-of-sample percentile rank | peak 88th, plateau 83rd |
| Worst out-of-sample outcome | peak -7.5%, plateau -8.8% |
I expected the plateau pick to win. The data said it does not, on this grid, on these markets, at this resolution. Report it and move on.
But the related check does work, and the distinction matters. Asking "is the whole neighbourhood profitable?" — not "which point in it is highest?" — was one of the strongest predictors in the entire study. More on that below.
The one hard threshold anybody quotes, tested
There is a rule that circulates in exactly one form: vary a parameter by 25%, and the metric should move by less than 25%. It is the only concrete robustness number most traders can name. So I ran it.
I went in expecting it to be folklore. It is not.
| On the testable population | Precision | Lift |
|---|---|---|
| No filter (base rate) | 90.5% | — |
| Passes the 25% rule | 96.3% | +5.8 pts |
| Every neighbour profitable (the binary version) | 93.0% | +2.5 pts |
Two things to read carefully here. First, that base rate is 90.5%, not 63.7%. The rule can only be tested where a percentage change means something and the grid can actually make the shift, so two conditions come first: a development CAGR of at least 1% a year, and a neighbouring grid value within 40% of the intended 25% shift. 2,862 of the 5,472 meet both, and inside that group most things survive anyway.
Second, and more interesting: the graded rule beats the yes/no version of the same idea. Asking how much the neighbourhood moves is worth more than twice the lift of asking whether the neighbours are profitable, on identical data.
Is 25% the right number? Close enough that it does not matter:
| Tolerance | Passes | Precision | Lift |
|---|---|---|---|
| 5% | 4.4% | 90.4% | -0.1 pts |
| 10% | 10.8% | 95.1% | +4.6 pts |
| 15% | 19.4% | 96.4% | +5.9 pts |
| 20% | 27.6% | 96.5% | +6.0 pts |
| 25% | 35.5% | 96.3% | +5.8 pts |
| 50% | 62.1% | 94.7% | +4.2 pts |
| 100% | 89.6% | 91.5% | +1.0 pts |
Anywhere from 15% to 25% behaves the same. What breaks it is tightening: at a 5% tolerance the lift is -0.001, no better than not filtering at all.
And most strategies fail the test. The median variant's worst metric move under a ~25% parameter shift is 35.4%, and 61.3% of them move more, proportionally, than the parameter you shifted. The typical variant tested here is more sensitive than the rule allows, which is the point of having the rule.
One caveat, and it is the same one PBO produced. By family, the lift is +0.105 for RSI mean reversion, +0.041 for MA trend, and -0.002 for Donchian breakout. On breakout the rule does nothing at all. A robustness test only reports something where there is an effect to be robust about.
How bad can the drawdown really get?
A backtest's worst drawdown comes from one ordering of its trades, and you got to see exactly one. For the strategy below, the same trades resampled with replacement drew a worse drawdown 29.5% of the time.
I took SPY's in-sample RSI winner (RSI(2) < 15, hold 5 days, 470 trades, 70% win rate, profit factor 2.13) and resampled its own trades with replacement 10,000 times.
| Maximum drawdown | Value |
|---|---|
| What the backtest showed | -27.4% |
| Typical run (median of 10,000) | -23.7% |
| 1 run in 20 | -37.6% |
| Worst of 10,000 | -65.1% |
| Percentile of the backtest's own figure | 29th |
The backtested drawdown sat at the 29th percentile of resampled histories, meaning 29.5% of the resamples drew a worse one. 2.4% of runs exceeded one and a half times the backtested drawdown, and 0.1% exceeded twice it.
Sizing to tolerate only the backtested drawdown would not have covered this strategy's 5th-percentile resampled drawdown. If -37.6% is more than you can sit through, size the position until you can, because this strategy's own trades produced it 1 run in 20. The study itself ran one unit and tested no sizing rule, so the cut to make is your own arithmetic on that figure, not a result measured here.
How many trades, how many markets, and how much cost?
Three practical thresholds, all measured rather than asserted.
Trades. Every in-sample-profitable variant, bucketed by how many trades produced the result:
| In-sample trades | Variants | Stayed profitable out-of-sample | Median out-of-sample rank |
|---|---|---|---|
| Fewer than 30 | 1,256 | 70.3% | 44th percentile |
| 30 to 99 | 1,408 | 85.2% | 51st percentile |
| 100 to 299 | 993 | 87.0% | 69th percentile |
| 300 to 999 | 225 | 95.6% | 82nd percentile |
The relationship is monotone and it is large: 70.3% at the thin end, 95.6% at the thick end.
Markets. Take an RSI parameter set and one of the 8 markets. Count how many of the 8 markets that parameter set was profitable on in-sample, using only bars dated before that market's own out-of-sample period begins. Then look at what the set did out-of-sample on that market:
| Markets profitable in-sample | Parameter set and market pairs | Share of out-of-sample results profitable |
|---|---|---|
| 0 | 288 | 0.0% |
| 3 | 57 | 56.1% |
| 5 | 370 | 60.3% |
| 7 | 759 | 73.0% |
For RSI on these eight markets, a parameter set profitable nowhere in-sample never made money out-of-sample, and one profitable on seven did so 73.0% of the time. The test needs no code you do not already have, and in the scorecard below it carries a lift of 54 percentage points in the share still profitable.
Costs. Frictionless is the repo default, so I charged costs explicitly to see what breaks:
| Holding period | Median in-sample trades | Profitable frictionless | Profitable at 10 bps round trip | Killed by costs |
|---|---|---|---|---|
| 1 day | 98 | 70.1% | 59.2% | 10.9% |
| 5 days | 67.5 | 67.8% | 62.9% | 4.9% |
| 10 days | 54 | 60.9% | 57.2% | 3.7% |
Across the whole RSI grid, going from zero to 20 basis points round trip cut the profitable share from 65.8% to 53.9%. For daily strategies holding several days, costs are a haircut, not an executioner. The faster you trade, the less that sentence applies.
Which robustness checks actually predict anything?
Here is the experiment the rest of the internet has not run.
I defined 11 checks, computed every one of them using only the first 75% of each market's history, and then scored each variant on the untouched final 25%. "Lift" below is the difference in outcome between variants that passed a check and variants that failed it. The base rate across all 5,472 variants: 63.7% were profitable in that final quarter, and 27.8% beat buy-and-hold there.
| Check | Passed | Profitable if passed | Profitable if failed | Lift |
|---|---|---|---|---|
| Profitable in the held-back validation slice | 69.7% | 84.1% | 16.7% | +67 pts |
| Profitable over the whole development sample | 71.1% | 82.2% | 18.1% | +64 pts |
| Still profitable at 5 bps per side | 66.5% | 83.9% | 23.5% | +60 pts |
| Parameter neighbours also profitable | 53.1% | 91.7% | 31.9% | +60 pts |
| Profitable on 5+ of the 8 markets | 87.4% | 70.5% | 16.4% | +54 pts |
| Profitable in all three thirds of the sample | 41.0% | 91.3% | 44.5% | +47 pts |
| Validation CAGR at least half the fitted CAGR | 73.9% | 73.9% | 34.9% | +39 pts |
| Development Sharpe of at least 0.5 | 16.7% | 94.5% | 57.5% | +37 pts |
| At least 100 trades | 28.7% | 81.2% | 56.7% | +25 pts |
| No single trade worth a quarter of the profit | 71.0% | 68.8% | 51.2% | +18 pts |
| Beats buy-and-hold in development | 28.2% | 41.2% | 72.5% | −31 pts |
Ten of the eleven checks carry a positive lift. The two largest are holding back a validation slice (+67 points) and profitability across the whole development sample (+64 points), and the parameter-neighbour check is close behind: variants whose neighbours were also profitable stayed profitable 91.7% of the time, against 31.9% for variants whose neighbours were not.
Then there is the last row, which is the most interesting number in the study. Variants that beat buy-and-hold during development were profitable afterwards only 41.2% of the time — well below the 72.5% of variants that never beat the market at all. The same check is the only one with a large lift on beating the market (at least 100 trades and cross-market breadth carry small positive lifts): those variants beat buy-and-hold again 76.9% of the time versus 8.5% for the rest.
The check that predicts beating the market is the one that predicts not surviving, and the reverse holds for eight of the ten other checks, with at least 100 trades and cross-market breadth the two exceptions. The check does not lie; it just answers a different question than the one you thought you asked.
Stack the checks and the pattern is unmistakable:
| Checks passed | Variants | Profitable in final quarter | Beat buy-and-hold | Median forward CAGR |
|---|---|---|---|---|
| 0 (everything) | 5,472 | 63.7% | 27.8% | 1.7% |
| At least 5 | 3,772 | 84.3% | 14.2% | 3.7% |
| At least 8 | 2,100 | 93.1% | 9.9% | 5.1% |
| At least 10 | 454 | 98.0% | 20.7% | 7.8% |
Among these 5,472 variants, the ones passing at least 10 of the 11 checks were profitable in the final quarter 98.0% of the time against a 63.7% base rate, with a median forward return of 7.8% a year against 1.7%. The same group beat buy-and-hold 20.7% of the time against 27.8%.
If staying profitable matters more to you than beating the index, that trade is worth making. It is just not the trade most people think they are making.
Do any of these checks work outside US equities?
Everything above rests on eight markets, and eight markets is not what it sounds like.
SPY and ES are the same index in two wrappers. Run the maths on the whole set and the eight are worth 2.87 independent bets, at a mean absolute pairwise correlation of 0.533 over their 4,890 common bars. Their common history starts in 2007, so every number above was fitted inside one macro era. That is a poor foundation for a claim about generality, and "these checks predict survival" is exactly that kind of claim.
So I rebuilt the universe: 53 markets spanning equity index, equity sector, FX, metals, energy, rates, grains, softs and livestock. 17.28 effective bets at mean correlation 0.231 — 6.0× the independent information — and 36,252 backtests under the identical protocol.
The thesis survived. Several of my numbers did not.
| Check | Lift on 8 markets | Lift on 53 markets |
|---|---|---|
| Profitable in the held-back validation slice | +0.675 | +0.293 |
| Profitable over the whole development sample | +0.641 | +0.205 |
| Parameter neighbours also profitable | +0.598 | +0.166 |
| Profitable in all three thirds | +0.469 | +0.182 |
| Development Sharpe of at least 0.5 | +0.370 | +0.041 |
| At least 100 trades | +0.245 | +0.170 |
The top two hold their places. The other four reorder, and every magnitude shrinks. Parameter neighbours fall from third to fifth and the Sharpe filter from fifth to last: a filter that looked worth 37 points of survival is worth 4. Every one of these was measured honestly the first time; they were just measured on a universe that was mostly one market.
What did hold is the part the whole article is built on. Variants passing all 11 checks still survived 97.2% of the time: 325 variants, 0.9% of the grid. The base rate itself is lower on the wider universe, 59.1% against 63.7%. The rerun measures that drop, not its cause: the 45 added markets differ from the original eight in asset-class mix, in their histories and in how they move together, and the study does not separate those. The stack is the finding. The individual checks were never meant to carry it alone, and now there is evidence they cannot.
One limit sits on that 97.2%. The 11 checks include cross-market breadth, and after this rerun I fixed a look-ahead in how breadth is counted: a market's breadth now reads the other markets only over bars dated before that market's own out-of-sample period. Every eight-market number in the text and tables of this article comes from the fixed code. The 53-market run was made before the fix, and I have not been able to repeat it, because the data for most of the 45 added markets is not on hand. So the 97.2%, the 325 variants and the 0.9% still carry the old breadth count. The lifts in the table above, the asset-class table below and the PBO figures do not use breadth.
The check that inverts
Here is the part eight markets could not have told me:
| Asset class | Markets | Base rate | Survival if neighbours pass | Neighbour lift | Sharpe-0.5 lift |
|---|---|---|---|---|---|
| Equity index | 9 | 77.4% | 89.7% | +0.366 | +0.201 |
| Metals | 6 | 74.7% | 83.7% | +0.155 | +0.241 |
| Equity sector | 9 | 68.9% | 80.4% | +0.241 | +0.188 |
| Livestock | 3 | 64.9% | 74.5% | +0.146 | +0.356 |
| Energy | 3 | 50.1% | 84.2% | +0.407 | −0.001 |
| Grains | 6 | 49.7% | 56.5% | +0.097 | −0.145 |
| Softs | 3 | 48.1% | 55.8% | +0.136 | −0.061 |
| FX | 11 | 45.6% | 29.5% | −0.210 | −0.399 |
| Rates | 3 | 25.3% | 18.1% | −0.183 | +0.037 |
Read the bottom two rows slowly. On FX, a variant whose parameter neighbours are also profitable survives 29.5% of the time against a 45.6% base rate. The check does not merely stop helping. It points the wrong way. Same on rates: 18.1% against 25.3%.
The parameter-neighbour check is not a general filter. Its lift is largest in energy at 0.407 and equity index at 0.366, and it points the wrong way in FX and rates.
I do not have a clean mechanism for it, and I would rather say that than invent one. The pattern is consistent with these markets being genuinely harder — rates has a 25.3% base rate against equity index's 77.4% — so a smooth, profitable-looking parameter neighbourhood in a market that mostly does not reward the strategy may be the signature of a fit rather than an edge. That is a hypothesis, not a result.
Mean reversion's advantage was the universe
The same correction hits the family comparison. PBO — the probability the in-sample champion lands below the out-of-sample median, where higher is worse:
| Family | Median PBO on 8 markets | Median PBO on 53 markets |
|---|---|---|
| RSI mean reversion | 0.192 | 0.444 |
| Donchian breakout | 0.448 | 0.532 |
| Moving-average trend | 0.714 | 0.603 |
Mean reversion looked like the safe family, at a median overfitting probability of 19.2% against trend's 71.4%. On a proper universe it is barely distinguishable from breakout. And trend improves: its median overfitting probability falls from 0.714 on the eight markets to 0.603 across 53. The study measures the fall without isolating what caused it.
The original eight are measurably the flattering subset: median PBO 0.391 against 0.492 for the 45 markets added. And the spread across asset classes — livestock 0.052 to softs 0.837 — is wider than the spread across strategy families.
Where you test matters more than what you test. That is not a caveat on the study. It is the largest single effect in it.
How many strategies survive everything?
Finally, the whole population through a seven-stage gauntlet, in order:
| Stage | Survivors | Share of 5,472 |
|---|---|---|
| Backtested | 5,472 | 100% |
| Profitable in-sample, 30+ trades | 2,626 | 48.0% |
| Beats buy-and-hold in-sample | 473 | 8.6% |
| Still profitable out-of-sample | 370 | 6.8% |
| Keeps half its fitted return | 297 | 5.4% |
| Survives 5 bps per side | 294 | 5.4% |
| Profitable in all four quarters of history | 262 | 4.8% |
| Same parameters work on other markets | 262 | 4.8% |
262 of 5,472 — 4.8% — survived all seven. Judge them on Sharpe instead of raw return and 592 survive (10.8%), so the choice of benchmark metric more than doubles the number of survivors.
The single biggest cut is not out-of-sample degradation. It is the buy-and-hold comparison: 2,626 variants were profitable in-sample on 30 or more trades and only 473 beat simply owning the thing, an 18.0% pass rate. Everything after that stage is comparatively gentle.
And look at what is left: 253 mean-reversion variants, 7 trend, 2 breakout. That is 7 of 1,056 trend variants and 2 of 1,344 breakout variants, against 253 of 3,072 for mean reversion. On these markets over these periods, the long-only trend and breakout grids almost never cleared the gauntlet.
The verdict — and the honest limits
The eleven-check stack works. It does not work the way robustness testing is sold.
Where the standard advice is right. Holding back data and looking at it was the most predictive of the 11 checks on eight markets, at +67 points, and at +0.293 on 53 it is still the largest lift of the six checks measured in both runs. Checking that neighbouring parameters also work is worth +0.166 across the wide universe. Trade count matters more than the usual "30 trades" folklore allows: among variants already profitable in-sample, those built on under 30 in-sample trades stayed profitable out-of-sample 70.3% of the time, and those built on 300 to 999 trades 95.6%.
Where it is wrong. Picking the middle of the plateau instead of the peak did not reliably improve out-of-sample results here (9 of 24, median difference 0.00%). For the daily RSI strategies tested here, at up to 20 basis points round trip, costs were a haircut, not a killer, for multi-day holds. Walk-forward measures an edge and does not create one: median efficiency 0.42, and 6 of 24 wins against buy-and-hold.
What eight markets could not tell me. "My process is robust" is not a claim a process can make about itself, and this study proved that on its own first pass. Eight markets put PBO at 19.2% on mean reversion against 71.4% on trend, which reads as mean reversion being the sturdier family. On 53 markets those become 0.444 and 0.603, so the gap largely closes and trend improves. The eight were the flattering subset (median PBO 0.391 against 0.492 for the 45 added). Every individual check lift above is smaller on a wide universe, several by a factor of three, and on FX and rates the neighbour check inverts outright. The stack still works at 97.2%, in the 53-market run made before the breadth fix. The pieces were oversold on a narrow universe, and only the width found it.
The part nobody says out loud. In the 53-market run made before the breadth fix, a variant that passed all 11 checks was profitable in the untouched final quarter of history 97.2% of the time, and that is the horizon the study measured. On the eight markets, where the buy-and-hold comparison was measured, it is also a variant selected against beating the index. If your goal is to beat buy-and-hold, this stack is not the tool: variants passing at least 10 of the 11 checks beat it 20.7% of the time against a 27.8% base rate. If your goal is a variant that stays profitable in a stretch of history it was never fitted on, this stack is what picked those variants out here.
Limits. This study is long only, one unit, no leverage, no shorting, no stops, no position sizing; adding any of those changes the numbers. The ETF series are price-only, so the buy-and-hold benchmark excludes dividends and is understated — a dividend-adjusted benchmark would make these strategies look worse, not better. Three strategy families and 221 mining conditions are a sample of retail practice, not a census of it. The zero-edge control removes serial structure and drift by construction; it is a floor for how good a backtest can look by luck alone, not a claim that markets are random. "Out-of-sample" still means history I happened to hold back, not the future.
The wide rerun has limits of its own, and they run the other way. Energy is thin: it rests on three markets, HO, NG and RB, and its NG is the original eight's own 2007 file. CL, BZ and the longer NG history are additively back-adjusted in a way that makes percentage returns meaningless, so they are not in the run. The bond and multi-asset ETFs are excluded because they arrived dividend-adjusted, and mixing total-return with price series would compare two different quantities. Four asset classes (energy, rates, softs and livestock) carry only three markets each, so their numbers are directional, not precise. And the breadth check could not be carried across literally: "profitable in 5 of 8 markets" is trivially easy across 53, so it is applied as the same fraction instead. That wide breadth check, and the stacked 97.2% that counts it, predate the look-ahead fix to breadth and have not been re-run, because the data for most of the added markets is not on hand.
Two independent facts worth putting beside all of this: McLean and Pontiff found that published anomaly returns are 26% lower out-of-sample and 58% lower after publication, in the Journal of Finance. Decay is the normal case even for effects that cleared academic review. Assume yours decays too.
What this means for you
- Compute your noise floor before you celebrate. Shuffle your own data, remove the drift, re-run your identical optimisation, and record the best result. Optimising 384 RSI mean-reversion variants over three years of shuffled, drift-free SPY produced a median best Sharpe of 1.17 from nothing. Judge your result by its margin over the floor your own re-run records, not by the result alone.
- Count your trials and say the number out loud. Testing 384 variants on zero-edge data lifts the best Sharpe from 0.00 to 0.36. If you tested 5,000 rules, subtract what the same search finds on zero-edge data: here that left 4.8% a year of an 11.2% headline.
- Hold back a slice and actually look at it. The strongest of the 11 checks on eight markets, at +67 points, and the largest lift of the six checks measured on both universes, at +0.293 on 53. And accept that the split is arbitrary: at 70/30 you keep a different variant than at 50/50 in 62.5% of cases, so treat the choice as a sensitivity, not a truth.
- Check the neighbours, not the peak — and know what you are trading. Ask whether every parameter set adjacent to yours is profitable. It is worth +0.366 on equity index markets and −0.210 on FX, so on currencies and rates this check is worse than useless. Do not bother relocating to the centre of the plateau — that part did not reliably help. Better still, grade it: hold the metric's move under 25% for a 25% parameter shift, which more than doubles the lift of the yes/no version on the same data.
- Demand trades and markets. Among variants already profitable in-sample, out-of-sample survival is 70.3% under 30 trades and 95.6% from 300 to 999. For RSI parameter sets profitable on 7 of the 8 markets, 73.0% of forward results held up, and for sets profitable on none, 0.0%.
- Count your markets the way you count your trials. Eight markets, six of them US equity proxies, bought me 2.87 independent bets and a conclusion about strategy families that did not survive contact with 53. Before you trust a cross-market result, check the correlation matrix.
- Read the drawdown off the Monte Carlo, not the backtest. For SPY's in-sample RSI winner, the backtested drawdown sat at the 29th percentile of resampled histories, and 1 run in 20 reached -37.6% against a backtested -27.4%.
- Benchmark against buy-and-hold every single time, and know what you are choosing. Of the 2,626 variants profitable in-sample on 30 or more trades, only 473 beat it in-sample. If a variant does beat it, that is the one check that predicts beating it again (76.9%), and the one that predicts a worse chance of just staying profitable (41.2%).
Four of those steps have a study of their own, each one built on this same population of strategies: how to configure a walk-forward (fit length is the only setting that matters), the shortest validation checklist that works (it is two items long), which metric to rank the survivors on (the plainest one), and what Monte Carlo gets wrong (the drawdown, badly).
If you want the worked examples, they are all on this site and they all went through some version of this treatment: the Better-RSI showdown as a test of a widely believed claim, 33,792 backtests of RSI against Stochastic and Williams %R, what survives when you backtest ICT and smart-money concepts, the Larry Williams COT filter, a hedge-fund trend rule that did survive, what happens when an AI generates the strategies, and measuring a market effect before trading it.
Running this stack against one strategy rather than a population of 5,472 is what AlgoChef does: you import a backtest you have already run, and it applies the Monte Carlo, the in-sample/out-of-sample split and the overfitting check to it.








