ATMResearch
Join Overfit - free
Overfit cover card on dark navy, kicker 'Z-score mean reversion': the headline 'A 74% win rate that never beat buy-and-hold.' over the line '2,400 backtests on S&P futures and SPY. The win rate is real — the edge isn't.', with a corner badge reading 'SPY & ES futures · 33 years'.
  1. Overfit/
  2. Research/
  3. Z-Score Mean Reversion Strategy: 2,400 Backtests on Futures and SPY, and the 74% Win Rate Is a Trap

March 21, 2025

Z-Score Mean Reversion Strategy: 2,400 Backtests on Futures and SPY, and the 74% Win Rate Is a Trap

Share

8 min read

Written by Ali Casey, founder of StatOasis and AlgoChef, creator of the Algo Trading Masterclass (ATM), with over 10 years of experience building systematic trading tools - building algorithmic strategies, testing ideas with data, and teaching traders how to build structured, portfolio-based trading workflows.

Published March 21, 2025 · Updated September 8, 2026 · Method

← Back to Research
Table of contents▾
  • TL;DR — the answer box
  • How we tested
  • Does a z-score mean reversion strategy beat buy-and-hold?
  • Is the famous 74% win rate real?
  • Does trading it on leveraged futures help?
  • Which side and style actually work?
  • So what can you actually trade? The stable region
  • The verdict — and the honest limits
  • What this means for you
  • Methodology
  • FAQs

The short version

I tested 2,400 z-score mean reversion variants on S&P 500 futures and SPY. The famous 74% win rate reproduces on both markets, and it is still a trap. Not one reliable variant beat simply holding the market: 0 of 934 on futures, 0 of 1,008 on SPY. What is left is a stable region at the 10-day lookback, not the best backtest, and it clears the study's random control on SPY but not on ES.

TL;DR — the answer box

  • The z-score mean reversion strategy's famous 74% win rate is real — I reproduced it on both S&P 500 futures and SPY. It's also a trap. A high win rate is not an edge.
  • I ran 2,400 variants — 1,200 on ES futures, 1,200 on SPY. Not one of the 1,942 that cleared the 50-trade reliability floor beat buy-and-hold of its own market: 0 of 934 on futures, 0 of 1,008 on SPY.
  • On ES the advantage is a smaller drawdown that comes with time out of the market: the stable region is in cash 75% of the time and its CAR/MaxDD sits inside the random control's spread. On SPY the control points the other way, with the featured variant at 0.62 against a control at 0.114, which the study's exposure figures alone do not account for. Shorting overbought z-scores loses on both: 1 of 971 reliable short variants turned a profit.
  • Leverage makes it worse, not better. Buy-and-hold one ES contract on a $35,000 account draws down $60,387 (114% of the account), and it gets wiped out. The tilt never takes that drawdown, and it is in the market only 25% of the time.
  • The tradable answer is a stable region, not the best backtest. The single best variant is overfit. But at lookback 10, buying the dip, median R-expectancy is 0.15 on ES and 0.18 on SPY, with the full lookback-10 ranges running 0.09 to 0.45 and 0.14 to 0.28. Modest and consistent, and on SPY it clears the study's random control while on ES it does not. On futures you size it for the drawdown.

How we tested

I rebuilt the z-score mean reversion strategy and ran it 1,200 ways on each of two markets, same $35,000 account, next-day-open fills, no look-ahead, no compounding.

  • ES — E-mini S&P 500 futures, my own trading instrument. $50 per index point, one contract per trade. Data from 2007 to 2026 (19.5 years, 4,917 days). This is the leveraged, how-I-actually-trade view.
  • SPY — the ETF, bought in whole shares. Data from 1993 to 2026 (33.4 years, 8,398 days) — the long-history view.

A z-score measures how far price has stretched from its own recent average, in units of its own volatility: z = (close − moving average) ÷ standard deviation, over a lookback window. A z-score of −2 means "two standard deviations below average" — statistically cheap. I swept the lookback (10–50 days), the entry (z below −1 to −3), the exit (z back above 0 to 1), mean-reversion vs. breakout, long vs. short, and four time-exits (0–15 days).

One thing to hold onto: the percentage results — win rate, per-trade Sharpe, how often you're in the market — are the same whether you size in shares or contracts. What changes with futures is dollars against your account, and that's the whole point of leverage. I report 1,008 reliable SPY variants and 934 reliable ES variants (those clearing 50 trades). And I never show a win rate or a profit without a risk-adjusted number and the market exposure beside it.

One z-score trade on SPY: buy when the 10-day z-score drops below −1, sell when it snaps back above 0.5. Held six days through the pre-2020-crash dip for +2.2%.

Does a z-score mean reversion strategy beat buy-and-hold?

No — on either market, not a single variant did. And on futures, buy-and-hold itself doesn't survive.

MarketBuy-and-hold profitBuy-and-hold max drawdownSurvives on $35k?Best strategy variantVariants that beat hold
ES futures (1 contract)$274,175$60,387 (114% of account)No — wiped out$234,2500 of 934
SPY ETF (791 shares)$551,738$95,829 (56%)Yes, brutally$96,8670 of 1,008

Buy-and-hold vs. the best z-score variant, same $35,000 account.

On SPY, buy-and-hold turned $35,000 into a $551,738 profit (8.8% a year). The best z-score variant made $96,867 — 18 cents on the dollar against doing nothing. It has one thing going for it: on a return-per-unit-of-drawdown basis it scores 0.75 vs. buy-and-hold's 0.16, because it never eats the full 56% drawdown. The best variant is invested only 32% of the time, and the study's random control is matched to the market's median reliable variant rather than to this one, so it cannot separate the timing from the exposure here. Lower risk, far lower return.

On ES futures, the leverage flips the picture in a way every futures trader needs to see. Buy-and-hold one contract made more in raw dollars terms over its shorter history, but it drew down $60,387, which is 114% of the account. At the March 2009 low the account hit −$5,737. It doesn't survive. A single naked contract held through a bear market is a blown account, full stop. The strategy never takes that drawdown: the featured ES variant is out of the market 71% of the time and its worst drawdown is 31.6%, the same low-exposure shape as SPY, now wearing a life jacket.

Buy-and-hold max drawdown as a share of a $35,000 account: one ES contract draws down 114% — the account is wiped out — versus 56% for SPY shares.

Is the famous 74% win rate real?

Yes, on both markets — and it still tells you almost nothing.

The strategy this rebuilds bought when z dropped below −1 and claimed a 74% win rate. I found it on both instruments. Buying the dip and exiting near the mean, the median win rate was 69.5% on ES and 69.7% on SPY; the best hit 81.0% and 80.0%. Twenty-two futures variants and thirty-four SPY variants cleared 74%. The claim is real.

Now put a risk number next to it.

MarketBest win rateMedian profitMedian R-expectancyMedian per-trade SharpeMedian return/drawdownExposure
ES futures81.0%$96,5870.160.1300.0522%
SPY ETF80.0%$52,8920.190.1700.1822%

The dip-buy (mean-reversion long) high-win-rate variants, by market.

Two-thirds-plus win rates, and the risk-adjusted return is ordinary. R-expectancy, how much you make per trade in units of what you risked, sits at 0.16 on ES and 0.19 on SPY: sixteen to nineteen cents of edge per dollar risked. On leveraged futures the return-per-drawdown is actually worse (0.05 median), because the wins are small and the leveraged drawdowns are not. This is the batting-average trap: you can hit .750 by only swinging at the softest pitches, and here you are in the market only a fifth of the sessions, 22% of the time. The win rate on its own settles nothing: these high-win variants make a positive 0.16 R per trade on ES and 0.19 on SPY, and not one of them beat buy-and-hold.

Does trading it on leveraged futures help?

It makes bigger dollars and a worse strategy. That's the leverage trade in one line.

MeasureES futuresSPY ETF
Best net profit (reliable variant)$234,250$96,867
Best return-per-drawdown (CAR/MaxDD)0.320.75
Buy-and-hold drawdown vs. account114% (blown)56%

Same strategy, two instruments, best-of-grid.

The best futures variant made $234,250 — more than twice the best SPY dollar figure — and every one of those dollars is real leverage at work. But the best ES variant on risk-adjusted return scores 0.32 against SPY's 0.75, because leverage scales the drawdowns right along with the gains, and your account doesn't grow to match. Leverage amplified the pain, not the edge. If you trade this on futures, the lesson isn't "size up" — it's "size so a normal drawdown doesn't end you," because a single contract already draws down more than the account.

Which side and style actually work?

Long and patient works. Short is a graveyard on both markets.

StyleSideVariants% profitableWin rateR-expSharpeRet/DDExposure
Mean reversionLong225100%67.6%0.200.1700.1621%
Mean reversionShort2790%42.0%−0.12−0.090−0.0365%
BreakoutLong279100%57.8%0.110.0900.0965%
BreakoutShort2250%32.3%−0.26−0.170−0.0321%

SPY, all four style × side combinations, reliable variants.

Buying oversold z-scores (mean-reversion long) was profitable in 100% of SPY variants and 99% on ES — a genuine, modest tilt. Following strength instead (breakout long) also made money, weaker on SPY (R-expectancy 0.11) but, on the shorter, trendier futures history, it actually produced the single best dollar figure.

The short side is where the fantasy dies, on both markets. 970 of the 971 reliable short variants — fading rallies, shorting overbought spikes — lost. That rounds to 0% profitable on each of the 504 SPY and 467 ES sets, median profit factor around 0.70. Shorting a long-term uptrend because a number says "overbought" is a reliable way to give money away. The short side loses.

So what can you actually trade? The stable region

Not the best backtest. The single highest R-expectancy anywhere in the grid was 0.76 on ES and 0.87 on SPY — but both came from a slow 50-day lookback that traded barely 50 times. That's one lucky cell. Trade it and you're trading an accident.

The tradable answer is the opposite of a peak: it's a plateau. Pick the lookback that trades the most — 10 days (median ~196 trades on ES, ~347 on SPY, up to 541) — so the result rests on hundreds of trades, then look at how R-expectancy behaves as you move the entry and exit around. It barely moves.

Shorter lookbacks trade far more often — a median 347 trades at 10 days versus 98 at 50 — for nearly the same R-expectancy. The 10-day lookback is where a result you can trust lives.
R-expectancy across every entry and exit threshold at the 10-day lookback on SPY: a flat 0.17–0.23 plateau, median 0.18. Every cell is positive.

Every cell is positive and clustered: on SPY, R-expectancy holds 0.17 to 0.23 across the whole plane (median 0.18); on ES, 0.12 to 0.22. Nothing falls off a cliff when you nudge a threshold. On SPY it clears the luck bar the study ran: the featured variant scores CAR/MaxDD 0.62 against a random control at 0.114 with a spread of 0.076. The limit worth stating is that the control is matched to the median reliable SPY variant, 334 trades at a ten-bar hold, not to this variant's 522 trades. On ES it does not: the stable region's 0.05 sits inside the ES control's 0.083 with a spread of 0.095, which is the middle of what random entries already do. So the neighbourhood result is an edge on SPY and, on ES, a plateau the control does not distinguish from luck. Median win 71 to 72%, exposure ~25%. On SPY it's a modest, robust, buy-the-dip tilt you can actually deploy, around fifteen to eighteen cents of edge per dollar risked, on a few hundred trades. Just don't expect it to beat holding, and on futures, size it so the drawdown can't end you.

The one to trade, named

If you want a single configuration rather than a region, take the interior cell whose whole neighborhood is strongest and steadiest (boxed in the heatmap above): lookback 10, buy when z drops below −1, take profit when z climbs back above +0.5, with a 5-day time stop. The same setup wins on both markets — a good sign it isn't a fluke.

InstrumentTradesWin rateNet profitR-expectancyPer-trade SharpeReturn/drawdownMax drawdownExposure
SPY ETF52265.7%$86,0380.220.170.6211.4%29%
ES futures30165.1%$140,8000.140.140.1331.6%29%

On SPY that variant earns about 16% of buy-and-hold's money with roughly a fifth of the drawdown (11.4% vs 56%) — a genuine lower-risk tilt, in cash 71% of the time. On futures the same variant makes $140,800 against SPY's $86,038, at a 31.6% drawdown against 11.4%: tradable, but only if you size it to survive.

The verdict — and the honest limits

The z-score mean reversion strategy is neither a scam nor a money machine, and on SPY it is tradable, as long as you trade the stable region and not the headline. The famous 74% win rate is a vanity number; the single best backtest is an overfit peak; and on the leveraged futures I trade, buy-and-hold a single contract doesn't even survive. The plateau underneath, lookback 10 and buy the dip, holds a steady median R-expectancy of 0.18 on SPY and 0.15 on ES across hundreds of trades. Only one of those two clears the study's luck bar: on SPY the plateau's median CAR/MaxDD is 0.34 against a random control at 0.114 with a spread of 0.076; on ES it is 0.05, inside a control at 0.083 with a spread of 0.095. So the deployable edge is the SPY one, modest and sized for the drawdown, and on ES the plateau is a shape the control cannot separate from luck.

Where the hype is right: buying oversold z-scores on the S&P does make money, consistently, and it wins most of its trades, on the ETF and on futures. Where it's wrong: that win rate is a vanity number. Not one of the 1,942 variants that cleared the 50-trade reliability floor beat buy-and-hold. Its shallower drawdown on ES comes with sitting in cash 75% of the time. On SPY the random control does not account for the result on exposure alone. Shorting is a straight loss. And on the leveraged instrument I actually trade, buy-and-hold a single contract doesn't even survive, so the honest use of z-score there is survival sizing, not a profit engine.

The limits — read them before you act:

  • Frictionless. No commissions or slippage. The busiest variants trade several hundred times; real costs hit an active, thin-edge strategy hardest — the live number is worse than shown, and worse still on futures.
  • Futures data is back-adjusted continuous. That preserves the point-to-point moves the P&L needs, but the price levels aren't literal quotes — read the ES numbers as dollars-per-contract, not index prices.
  • Different windows. ES history is 19.5 years (2007+); SPY is 33.4 (1993+). The two aren't a like-for-like race; they're two honest views of the same idea.
  • In-sample. I swept 1,200 variants per market and then showed you the best — by definition the luckiest on that history. That's the overfitting trap; a winner needs out-of-sample and Monte-Carlo checks before a dollar rides on it, which is exactly what a validation tool like AlgoChef is for.
  • Fixed 1 contract / fixed shares, no compounding, no margin model. Sizing is flat; that keeps the comparison clean and understates neither the leverage risk nor the flat return.
  • Exposure is not isolated. The variants that dodge the buy-and-hold drawdown are also out of the market most of the time. I did not run the counterfactual that separates time in cash from where the entries fall, so low exposure travels with these results without the study showing it caused them.

What this means for you

  1. Stop grading strategies by win rate. Ask for the risk-adjusted return and the exposure in the same breath. A 74% win rate that's in cash 75% of the time is not what it sounds like.
  2. Benchmark against buy-and-hold — the survivable version. If a strategy can't beat holding, its job is lower risk, not more money. And on futures, "buy-and-hold" a single contract isn't even the safe option — it blew the account here.
  3. On futures, size for the drawdown, not the win rate. One ES contract already draws down more than a $35k account on a naked hold. The z-score tilt is out of the market 75% of the time, and its featured ES variant still drew down 31.6%. Size for that.
  4. Trade the stable region, not the best backtest. Lookback 10, buy the dip (z below −1 to −2), exit near or above the mean — the neighbourhood's median R-expectancy is 0.15 on ES and 0.18 on SPY, on hundreds of trades. Pick anywhere in that plateau; don't chase the single lucky variant. And pair it with something that uses the 75% of the time you're idle.
  5. Do not short overbought z-scores on these two uptrending indexes. One of 971 reliable short variants across both markets turned a profit. Don't.
  6. Validate before you trust the win rate. Any headline number from a single backtest — including this study's best variant — should survive out-of-sample and resampling first.

Methodology

Data source
SPY (S&P 500 ETF) daily OHLCV 1993-2026, and ES E-mini S&P 500 futures daily 2007-2026, from the StatOasis research dataset.
Date range
SPY 1993-02-02 to 2026-06-12 (33.4 years, 8,398 bars); ES 2007-01-03 to 2026-07-01 (19.5 years, 4,917 bars).
Entry / exit rules
Signal: z-score = (close minus the moving average over the lookback) divided by the population standard deviation. Mean-reversion long buys when z crosses below the entry threshold (swept -1 to -3) and exits when z climbs back above the exit threshold (0 to 1); breakout and short variants are swept too. Decision at the close, fill at the next bar's open — no look-ahead. A protective time exit of 0/5/10/15 days is also swept.
Sizing
$35,000 starting capital, no compounding, flat-only (one position at a time), frictionless. SPY sized in whole shares (floor of capital divided by price); ES sized as one contract at $50 per index point, which is the futures-leverage view.
Overlap mode
Signal-to-signal and flat-only — a new signal is ignored while a position is open. Reliability bar: only variants with 50 or more trades are reported.
Look-ahead
The z-score is computed on the close of bar t and the fill is the open of bar t+1, entries and exits alike. No number uses information unavailable at the decision.
Minimum sample
50 trades. 1,008 of the SPY variants and 934 of the ES variants clear it; only those are reported.
Buy-and-hold benchmark
Same bars, same sizing. SPY: $551,738.32 net, 8.82% CAGR, and a worst drawdown of 56.47% in March 2009 (CAR over that drawdown, 0.16). ES: $274,175.00 net, 11.83% CAGR, and a worst drawdown of 114.33% of the account in March 2009 - one contract against $35,000 is leveraged enough to lose more than the whole account on the way, which puts its CAGR over that worst drawdown at 0.103.
Random control
Frequency-matched seeded coin flip, 10 seeds from base seed 20260803, matched to each market's own median reliable variant. SPY (334 entries, 10-bar hold): $34,231.56 net (sd $15,933.68), a 26.79% worst drawdown, CAR over that drawdown 0.097. ES (212 entries, 10-bar hold): $98,056.00 net (sd $67,571.99), a 67.31% worst drawdown, CAR over that drawdown 0.034. Computed by the StatOasis control harness.
Parameter scopeParameters swept

The study searched the parameter space and reports the spread, not one tuned setting.

2,400 variants sweeping lookback, entry threshold (-1 to -3), exit threshold (0 to 1), direction, entry style and a protective time exit of 0/5/10/15 days, on both SPY and ES. The article reports the spread across that space, and the 74% win rate is shown as what it is rather than as a selected best.

Run to v1 of the StatOasis research standard - the rules every study here has to meet before it is published. The version is the study's own: a standard that gained a rule later never reaches back and claims this one met it.

Historical backtest results are not a guarantee of future returns. This content is for educational purposes only and is not investment advice. Hypothetical performance disclosure (CFTC Rule 4.41).

Frequently asked questions

What is a z-score in trading?⌄

A z-score measures how far price has moved from its recent average, in units of its own volatility: z = (close minus the moving average) divided by the standard deviation, over a chosen lookback. A reading of -2 means price is two standard deviations below average, unusually cheap.

Does a z-score mean reversion strategy actually work?⌄

It makes money but never beats buy-and-hold. Across 2,400 variants on S&P futures and SPY, oversold-buy configurations were profitable roughly 100% of the time, yet 0 of 934 futures variants and 0 of 1,008 SPY variants out-earned simply holding.

Does the z-score strategy work better on futures than SPY?⌄

No. On ES futures the best variant made more raw dollars ($234,250 vs $96,867) purely from leverage, but the best ES variant on return-per-drawdown scores 0.32 against SPY's 0.75. Leverage scales the losses too.

Is buy-and-hold safe on futures?⌄

Not on a single contract with a small account. Holding one ES contract through 2008 drew down $60,387, which is 114% of a $35,000 account, and the account went negative. One contract against this account did not survive this history.

Is a high win rate enough to make a strategy profitable?⌄

No, and this study is the cleanest proof. The dip-buy variants have a median win rate of 69.5% on ES and 69.7% on SPY, with the best reaching 81.0%, yet a median R-expectancy near 0.16 to 0.19, and they never beat buy-and-hold. Win rate without a risk-adjusted number is meaningless.

What is the tradable z-score configuration?⌄

Not the single best backtest, but the stable region around it: a 10-day lookback, buy when z drops below -1, exit when z climbs back above +0.5, with a 5-day time stop. On SPY that variant made $86,038 across 522 trades at a 65.7% win rate, a 0.62 return-to-drawdown, and only an 11.4% max drawdown.

What z-score means oversold or overbought?⌄

Conventionally, below -2 is oversold and above +2 is overbought. In this study the buying edge existed across entries from -1 to -3, and the tighter the entry, the fewer the trades.

Can you short overbought z-scores?⌄

The data says do not. 970 of 971 reliable short configurations lost, rounding to 0% profitable on each of the 504 SPY and 467 ES sets, at a median profit factor around 0.70. Shorting a long-term uptrend on an overbought signal is a losing bet.

Should you use z-score for mean reversion or momentum?⌄

Both made money long-only. Mean reversion (buying dips) was the steadier tilt; breakout (buying strength) topped the dollar table on the trendier futures history. Neither beat buy-and-hold.

Read the Strategies, Backtested hub
← Back to Research

Table of contents

  • TL;DR — the answer box
  • How we tested
  • Does a z-score mean reversion strategy beat buy-and-hold?
  • Is the famous 74% win rate real?
  • Does trading it on leveraged futures help?
  • Which side and style actually work?
  • So what can you actually trade? The stable region
  • The verdict — and the honest limits
  • What this means for you
  • Methodology
  • FAQs

Overfit - the newsletter

Skip the hype. Trust the data.One practical takeaway per issue.

Running a quick security check before this can be sent.

Delivered +2 times a month - when the work is ready, not on a calendar.I'll never sell your address.Unsubscribe in one click, any time.

Related Articles

Back to all research
#142

The Better-RSI Showdown: We Tested 4 RSI Upgrades on SPY, QQQ, IWM, and DIA

Sep 10, 2026 · 9 min read

1,856 backtests across four RSI families on SPY, QQQ, IWM and DIA: Connors RSI and Z-Score RSI modestly beat plain RSI on median risk-adjusted return, Laguerre RSI was the worst of the four despite its lag-free marketing, and the popular Triple RSI claim does not survive contact with the data.

Read more→
#141

I Backtested ICT / Smart Money Concepts — What Survives

Sep 3, 2026 · 8 min read

The four core ICT / Smart Money Concepts entries — order blocks, fair value gaps, liquidity sweeps and Optimal Trade Entry — codified into mechanical rules and run against three textbook entries and a coin flip across four markets. None showed a statistically significant forward-return edge on SPY.

Read more→

Overfit - the newsletter

Skip the hype. Trust the data.

One practical, evidence-driven takeaway per issue - strategy testing, portfolio construction, market structure, trading psychology, tactical asset allocation.

Running a quick security check before this can be sent.

Delivered +2 times a month - when the work is ready, not on a calendar.I'll never sell your address.Unsubscribe in one click, any time.

StatOasis is calm, evidence-based algorithmic-trading education, founded by Ali Casey. Ali builds systematic trading strategies and teaches the workflow behind them: research, build, test, combine, deploy. He writes the Overfit newsletter, published since 2024, and runs the Algo Trading Masterclass.

Socials

  • X ↗
  • YouTube ↗
  • Instagram ↗
  • LinkedIn ↗
  • GitHub ↗
  • Muck Rack ↗
  • LinkedIn SO ↗
  • GitHub SO ↗

Products

  • Overfit - the newsletter
  • Algo Trading Masterclass
  • StatOasis Community
  • 36 Ways to Buy the Dip
  • AlgoChef ↗

Reading & tools

  • Research
  • Methodology
  • Survive the Decade
  • Wall of Love

StatOasis

  • About Ali Casey
  • Contact
© 2026 StatOasis. Calm, evidence-based.
PrivacyTermsHypothetical resultsCalifornia