TL;DR — the answer box
- The z-score mean reversion strategy's famous 74% win rate is real — I reproduced it on both S&P 500 futures and SPY. It's also a trap. A high win rate is not an edge.
- I ran 2,400 variants — 1,200 on ES futures, 1,200 on SPY. Not one of the 1,942 that cleared the 50-trade reliability floor beat buy-and-hold of its own market: 0 of 934 on futures, 0 of 1,008 on SPY.
- On ES the advantage is a smaller drawdown that comes with time out of the market: the stable region is in cash 75% of the time and its CAR/MaxDD sits inside the random control's spread. On SPY the control points the other way, with the featured variant at 0.62 against a control at 0.114, which the study's exposure figures alone do not account for. Shorting overbought z-scores loses on both: 1 of 971 reliable short variants turned a profit.
- Leverage makes it worse, not better. Buy-and-hold one ES contract on a $35,000 account draws down $60,387 (114% of the account), and it gets wiped out. The tilt never takes that drawdown, and it is in the market only 25% of the time.
- The tradable answer is a stable region, not the best backtest. The single best variant is overfit. But at lookback 10, buying the dip, median R-expectancy is 0.15 on ES and 0.18 on SPY, with the full lookback-10 ranges running 0.09 to 0.45 and 0.14 to 0.28. Modest and consistent, and on SPY it clears the study's random control while on ES it does not. On futures you size it for the drawdown.
How we tested
I rebuilt the z-score mean reversion strategy and ran it 1,200 ways on each of two markets, same $35,000 account, next-day-open fills, no look-ahead, no compounding.
- ES — E-mini S&P 500 futures, my own trading instrument. $50 per index point, one contract per trade. Data from 2007 to 2026 (19.5 years, 4,917 days). This is the leveraged, how-I-actually-trade view.
- SPY — the ETF, bought in whole shares. Data from 1993 to 2026 (33.4 years, 8,398 days) — the long-history view.
A z-score measures how far price has stretched from its own recent average, in units of its own volatility: z = (close − moving average) ÷ standard deviation, over a lookback window. A z-score of −2 means "two standard deviations below average" — statistically cheap. I swept the lookback (10–50 days), the entry (z below −1 to −3), the exit (z back above 0 to 1), mean-reversion vs. breakout, long vs. short, and four time-exits (0–15 days).
One thing to hold onto: the percentage results — win rate, per-trade Sharpe, how often you're in the market — are the same whether you size in shares or contracts. What changes with futures is dollars against your account, and that's the whole point of leverage. I report 1,008 reliable SPY variants and 934 reliable ES variants (those clearing 50 trades). And I never show a win rate or a profit without a risk-adjusted number and the market exposure beside it.
Does a z-score mean reversion strategy beat buy-and-hold?
No — on either market, not a single variant did. And on futures, buy-and-hold itself doesn't survive.
| Market | Buy-and-hold profit | Buy-and-hold max drawdown | Survives on $35k? | Best strategy variant | Variants that beat hold |
|---|---|---|---|---|---|
| ES futures (1 contract) | $274,175 | $60,387 (114% of account) | No — wiped out | $234,250 | 0 of 934 |
| SPY ETF (791 shares) | $551,738 | $95,829 (56%) | Yes, brutally | $96,867 | 0 of 1,008 |
Buy-and-hold vs. the best z-score variant, same $35,000 account.
On SPY, buy-and-hold turned $35,000 into a $551,738 profit (8.8% a year). The best z-score variant made $96,867 — 18 cents on the dollar against doing nothing. It has one thing going for it: on a return-per-unit-of-drawdown basis it scores 0.75 vs. buy-and-hold's 0.16, because it never eats the full 56% drawdown. The best variant is invested only 32% of the time, and the study's random control is matched to the market's median reliable variant rather than to this one, so it cannot separate the timing from the exposure here. Lower risk, far lower return.
On ES futures, the leverage flips the picture in a way every futures trader needs to see. Buy-and-hold one contract made more in raw dollars terms over its shorter history, but it drew down $60,387, which is 114% of the account. At the March 2009 low the account hit −$5,737. It doesn't survive. A single naked contract held through a bear market is a blown account, full stop. The strategy never takes that drawdown: the featured ES variant is out of the market 71% of the time and its worst drawdown is 31.6%, the same low-exposure shape as SPY, now wearing a life jacket.
Is the famous 74% win rate real?
Yes, on both markets — and it still tells you almost nothing.
The strategy this rebuilds bought when z dropped below −1 and claimed a 74% win rate. I found it on both instruments. Buying the dip and exiting near the mean, the median win rate was 69.5% on ES and 69.7% on SPY; the best hit 81.0% and 80.0%. Twenty-two futures variants and thirty-four SPY variants cleared 74%. The claim is real.
Now put a risk number next to it.
| Market | Best win rate | Median profit | Median R-expectancy | Median per-trade Sharpe | Median return/drawdown | Exposure |
|---|---|---|---|---|---|---|
| ES futures | 81.0% | $96,587 | 0.16 | 0.130 | 0.05 | 22% |
| SPY ETF | 80.0% | $52,892 | 0.19 | 0.170 | 0.18 | 22% |
The dip-buy (mean-reversion long) high-win-rate variants, by market.
Two-thirds-plus win rates, and the risk-adjusted return is ordinary. R-expectancy, how much you make per trade in units of what you risked, sits at 0.16 on ES and 0.19 on SPY: sixteen to nineteen cents of edge per dollar risked. On leveraged futures the return-per-drawdown is actually worse (0.05 median), because the wins are small and the leveraged drawdowns are not. This is the batting-average trap: you can hit .750 by only swinging at the softest pitches, and here you are in the market only a fifth of the sessions, 22% of the time. The win rate on its own settles nothing: these high-win variants make a positive 0.16 R per trade on ES and 0.19 on SPY, and not one of them beat buy-and-hold.
Does trading it on leveraged futures help?
It makes bigger dollars and a worse strategy. That's the leverage trade in one line.
| Measure | ES futures | SPY ETF |
|---|---|---|
| Best net profit (reliable variant) | $234,250 | $96,867 |
| Best return-per-drawdown (CAR/MaxDD) | 0.32 | 0.75 |
| Buy-and-hold drawdown vs. account | 114% (blown) | 56% |
Same strategy, two instruments, best-of-grid.
The best futures variant made $234,250 — more than twice the best SPY dollar figure — and every one of those dollars is real leverage at work. But the best ES variant on risk-adjusted return scores 0.32 against SPY's 0.75, because leverage scales the drawdowns right along with the gains, and your account doesn't grow to match. Leverage amplified the pain, not the edge. If you trade this on futures, the lesson isn't "size up" — it's "size so a normal drawdown doesn't end you," because a single contract already draws down more than the account.
Which side and style actually work?
Long and patient works. Short is a graveyard on both markets.
| Style | Side | Variants | % profitable | Win rate | R-exp | Sharpe | Ret/DD | Exposure |
|---|---|---|---|---|---|---|---|---|
| Mean reversion | Long | 225 | 100% | 67.6% | 0.20 | 0.170 | 0.16 | 21% |
| Mean reversion | Short | 279 | 0% | 42.0% | −0.12 | −0.090 | −0.03 | 65% |
| Breakout | Long | 279 | 100% | 57.8% | 0.11 | 0.090 | 0.09 | 65% |
| Breakout | Short | 225 | 0% | 32.3% | −0.26 | −0.170 | −0.03 | 21% |
SPY, all four style × side combinations, reliable variants.
Buying oversold z-scores (mean-reversion long) was profitable in 100% of SPY variants and 99% on ES — a genuine, modest tilt. Following strength instead (breakout long) also made money, weaker on SPY (R-expectancy 0.11) but, on the shorter, trendier futures history, it actually produced the single best dollar figure.
The short side is where the fantasy dies, on both markets. 970 of the 971 reliable short variants — fading rallies, shorting overbought spikes — lost. That rounds to 0% profitable on each of the 504 SPY and 467 ES sets, median profit factor around 0.70. Shorting a long-term uptrend because a number says "overbought" is a reliable way to give money away. The short side loses.
So what can you actually trade? The stable region
Not the best backtest. The single highest R-expectancy anywhere in the grid was 0.76 on ES and 0.87 on SPY — but both came from a slow 50-day lookback that traded barely 50 times. That's one lucky cell. Trade it and you're trading an accident.
The tradable answer is the opposite of a peak: it's a plateau. Pick the lookback that trades the most — 10 days (median ~196 trades on ES, ~347 on SPY, up to 541) — so the result rests on hundreds of trades, then look at how R-expectancy behaves as you move the entry and exit around. It barely moves.
Every cell is positive and clustered: on SPY, R-expectancy holds 0.17 to 0.23 across the whole plane (median 0.18); on ES, 0.12 to 0.22. Nothing falls off a cliff when you nudge a threshold. On SPY it clears the luck bar the study ran: the featured variant scores CAR/MaxDD 0.62 against a random control at 0.114 with a spread of 0.076. The limit worth stating is that the control is matched to the median reliable SPY variant, 334 trades at a ten-bar hold, not to this variant's 522 trades. On ES it does not: the stable region's 0.05 sits inside the ES control's 0.083 with a spread of 0.095, which is the middle of what random entries already do. So the neighbourhood result is an edge on SPY and, on ES, a plateau the control does not distinguish from luck. Median win 71 to 72%, exposure ~25%. On SPY it's a modest, robust, buy-the-dip tilt you can actually deploy, around fifteen to eighteen cents of edge per dollar risked, on a few hundred trades. Just don't expect it to beat holding, and on futures, size it so the drawdown can't end you.
The one to trade, named
If you want a single configuration rather than a region, take the interior cell whose whole neighborhood is strongest and steadiest (boxed in the heatmap above): lookback 10, buy when z drops below −1, take profit when z climbs back above +0.5, with a 5-day time stop. The same setup wins on both markets — a good sign it isn't a fluke.
| Instrument | Trades | Win rate | Net profit | R-expectancy | Per-trade Sharpe | Return/drawdown | Max drawdown | Exposure |
|---|---|---|---|---|---|---|---|---|
| SPY ETF | 522 | 65.7% | $86,038 | 0.22 | 0.17 | 0.62 | 11.4% | 29% |
| ES futures | 301 | 65.1% | $140,800 | 0.14 | 0.14 | 0.13 | 31.6% | 29% |
On SPY that variant earns about 16% of buy-and-hold's money with roughly a fifth of the drawdown (11.4% vs 56%) — a genuine lower-risk tilt, in cash 71% of the time. On futures the same variant makes $140,800 against SPY's $86,038, at a 31.6% drawdown against 11.4%: tradable, but only if you size it to survive.
The verdict — and the honest limits
The z-score mean reversion strategy is neither a scam nor a money machine, and on SPY it is tradable, as long as you trade the stable region and not the headline. The famous 74% win rate is a vanity number; the single best backtest is an overfit peak; and on the leveraged futures I trade, buy-and-hold a single contract doesn't even survive. The plateau underneath, lookback 10 and buy the dip, holds a steady median R-expectancy of 0.18 on SPY and 0.15 on ES across hundreds of trades. Only one of those two clears the study's luck bar: on SPY the plateau's median CAR/MaxDD is 0.34 against a random control at 0.114 with a spread of 0.076; on ES it is 0.05, inside a control at 0.083 with a spread of 0.095. So the deployable edge is the SPY one, modest and sized for the drawdown, and on ES the plateau is a shape the control cannot separate from luck.
Where the hype is right: buying oversold z-scores on the S&P does make money, consistently, and it wins most of its trades, on the ETF and on futures. Where it's wrong: that win rate is a vanity number. Not one of the 1,942 variants that cleared the 50-trade reliability floor beat buy-and-hold. Its shallower drawdown on ES comes with sitting in cash 75% of the time. On SPY the random control does not account for the result on exposure alone. Shorting is a straight loss. And on the leveraged instrument I actually trade, buy-and-hold a single contract doesn't even survive, so the honest use of z-score there is survival sizing, not a profit engine.
The limits — read them before you act:
- Frictionless. No commissions or slippage. The busiest variants trade several hundred times; real costs hit an active, thin-edge strategy hardest — the live number is worse than shown, and worse still on futures.
- Futures data is back-adjusted continuous. That preserves the point-to-point moves the P&L needs, but the price levels aren't literal quotes — read the ES numbers as dollars-per-contract, not index prices.
- Different windows. ES history is 19.5 years (2007+); SPY is 33.4 (1993+). The two aren't a like-for-like race; they're two honest views of the same idea.
- In-sample. I swept 1,200 variants per market and then showed you the best — by definition the luckiest on that history. That's the overfitting trap; a winner needs out-of-sample and Monte-Carlo checks before a dollar rides on it, which is exactly what a validation tool like AlgoChef is for.
- Fixed 1 contract / fixed shares, no compounding, no margin model. Sizing is flat; that keeps the comparison clean and understates neither the leverage risk nor the flat return.
- Exposure is not isolated. The variants that dodge the buy-and-hold drawdown are also out of the market most of the time. I did not run the counterfactual that separates time in cash from where the entries fall, so low exposure travels with these results without the study showing it caused them.
What this means for you
- Stop grading strategies by win rate. Ask for the risk-adjusted return and the exposure in the same breath. A 74% win rate that's in cash 75% of the time is not what it sounds like.
- Benchmark against buy-and-hold — the survivable version. If a strategy can't beat holding, its job is lower risk, not more money. And on futures, "buy-and-hold" a single contract isn't even the safe option — it blew the account here.
- On futures, size for the drawdown, not the win rate. One ES contract already draws down more than a $35k account on a naked hold. The z-score tilt is out of the market 75% of the time, and its featured ES variant still drew down 31.6%. Size for that.
- Trade the stable region, not the best backtest. Lookback 10, buy the dip (z below −1 to −2), exit near or above the mean — the neighbourhood's median R-expectancy is 0.15 on ES and 0.18 on SPY, on hundreds of trades. Pick anywhere in that plateau; don't chase the single lucky variant. And pair it with something that uses the 75% of the time you're idle.
- Do not short overbought z-scores on these two uptrending indexes. One of 971 reliable short variants across both markets turned a profit. Don't.
- Validate before you trust the win rate. Any headline number from a single backtest — including this study's best variant — should survive out-of-sample and resampling first.






