ATMResearch
Join Overfit - free
Overfit cover card on dark navy, kicker 'Z-score mean reversion': the headline 'A 74% win rate that never beat buy-and-hold.' over the line '2,400 backtests on S&P futures and SPY. The win rate is real — the edge isn't.', with a corner badge reading 'SPY & ES futures · 33 years'.
  1. Overfit/
  2. Research/
  3. Z-Score Mean Reversion Strategy: 2,400 Backtests on Futures and SPY, and the 74% Win Rate Is a Trap

March 21, 2025

Z-Score Mean Reversion Strategy: 2,400 Backtests on Futures and SPY, and the 74% Win Rate Is a Trap

Share

7 min read

Written by Ali Casey, founder of StatOasis and AlgoChef, creator of the Algo Trading Masterclass (ATM), with over 10 years of experience building systematic trading tools - building algorithmic strategies, testing ideas with data, and teaching traders how to build structured, portfolio-based trading workflows.

Published March 21, 2025 · Updated August 22, 2026 · Method

← Back to Research
Table of contents▾
  • TL;DR — the answer box
  • How we tested
  • Does a z-score mean reversion strategy beat buy-and-hold?
  • Is the famous 74% win rate real?
  • Does trading it on leveraged futures help?
  • Which side and style actually work?
  • So what can you actually trade? The stable region
  • The verdict — and the honest limits
  • What this means for you
  • Methodology
  • FAQs

The short version

I tested 2,400 z-score mean reversion variants on S&P 500 futures and SPY. The famous 74% win rate reproduces on both markets, and it is still a trap: not one reliable variant beat simply holding the market — 0 of 934 on futures, 0 of 1,008 on SPY. What is left worth trading is a stable region at the 10-day lookback, not the best backtest.

TL;DR — the answer box

  • The z-score mean reversion strategy's famous 74% win rate is real — I reproduced it on both S&P 500 futures and SPY. It's also a trap. A high win rate is not an edge.
  • I ran 2,400 variants — 1,200 on ES futures, 1,200 on SPY. Not one beat buy-and-hold of its own market: 0 of 934 reliable futures variants, 0 of 1,008 on SPY.
  • The strategy's only advantage is a smaller drawdown — not skill, just exposure (in cash 60–75% of the time). Shorting overbought z-scores loses outright: 0% of short variants were profitable on either market.
  • Leverage makes it worse, not better. Buy-and-hold one ES contract on a $35,000 account draws down $60,387 (114% of the account) — it gets wiped out. The tilt survives only because it's barely in the market.
  • The tradable answer is a stable region, not the best backtest. The single best variant is overfit. But at lookback 10, buying the dip, R-expectancy holds a steady ~0.15–0.18 across every entry and exit threshold on hundreds of trades. Modest, consistent, real — and on futures you size it for the drawdown.

How we tested

I rebuilt the z-score mean reversion strategy and ran it 1,200 ways on each of two markets, same $35,000 account, next-day-open fills, no look-ahead, no compounding.

  • ES — E-mini S&P 500 futures, my own trading instrument. $50 per index point, one contract per trade. Data from 2007 to 2026 (19.5 years, 4,917 days). This is the leveraged, how-I-actually-trade view.
  • SPY — the ETF, bought in whole shares. Data from 1993 to 2026 (33.4 years, 8,398 days) — the long-history view.

A z-score measures how far price has stretched from its own recent average, in units of its own volatility: z = (close − moving average) ÷ standard deviation, over a lookback window. A z-score of −2 means "two standard deviations below average" — statistically cheap. I swept the lookback (10–50 days), the entry (z below −1 to −3), the exit (z back above 0 to 1), mean-reversion vs. breakout, long vs. short, and four time-exits (0–15 days).

One thing to hold onto: the percentage results — win rate, per-trade Sharpe, how often you're in the market — are the same whether you size in shares or contracts. What changes with futures is dollars against your account, and that's the whole point of leverage. I report 1,008 reliable SPY variants and 934 reliable ES variants (those clearing 50 trades). And I never show a win rate or a profit without a risk-adjusted number and the market exposure beside it.

One z-score trade on SPY: buy when the 10-day z-score drops below −1, sell when it snaps back above 0.5. Held six days through the pre-2020-crash dip for +2.2%.

Does a z-score mean reversion strategy beat buy-and-hold?

No — on either market, not a single variant did. And on futures, buy-and-hold itself doesn't survive.

MarketBuy-and-hold profitBuy-and-hold max drawdownSurvives on $35k?Best strategy variantVariants that beat hold
ES futures (1 contract)$274,175$60,387 (114% of account)No — wiped out$234,2500 of 934
SPY ETF (791 shares)$551,738$95,829 (56%)Yes, brutally$96,8670 of 1,008

Buy-and-hold vs. the best z-score variant, same $35,000 account.

On SPY, buy-and-hold turned $35,000 into a $551,738 profit (8.8% a year). The best z-score variant made $96,867 — 18 cents on the dollar against doing nothing. It has one thing going for it: on a return-per-unit-of-drawdown basis it scores 0.75 vs. buy-and-hold's 0.16, because it never eats the full 56% drawdown. But that's not timing skill — it's that the best variant is invested only 38% of the time. Lower risk, far lower return.

On ES futures, the leverage flips the picture in a way every futures trader needs to see. Buy-and-hold one contract made more in raw dollars terms over its shorter history, but it drew down $60,387 — 114% of the account. At the 2008 low the account hit −$5,737. It doesn't survive. A single naked contract held through a bear market is a blown account, full stop. The strategy survives that only because it's mostly in cash — which is the same low-exposure story as SPY, now wearing a life jacket.

Buy-and-hold max drawdown as a share of a $35,000 account: one ES contract draws down 114% — the account is wiped out — versus 56% for SPY shares.

Is the famous 74% win rate real?

Yes, on both markets — and it still tells you almost nothing.

The strategy this rebuilds bought when z dropped below −1 and claimed a 74% win rate. I found it on both instruments. Buying the dip and exiting near the mean, the median win rate was 69.5% on ES and 69.7% on SPY; the best hit 81.0% and 80.0%. Twenty-two futures variants and thirty-four SPY variants cleared 74%. The claim is real.

Now put a risk number next to it.

MarketBest win rateMedian profitMedian R-expectancyMedian per-trade SharpeMedian return/drawdownExposure
ES futures81.0%$96,5870.160.1300.0525%
SPY ETF80.0%$52,8920.190.1700.1826%

The dip-buy (mean-reversion long) high-win-rate variants, by market.

Two-thirds-plus win rates, and the risk-adjusted return is ordinary. R-expectancy — how much you make per trade in units of what you risked — sits near 0.16–0.19: fifteen-odd cents of edge per dollar risked. On leveraged futures the return-per-drawdown is actually worse (0.05 median), because the wins are small and the leveraged drawdowns are not. This is the batting-average trap: you can hit .750 by only swinging at the softest pitches, but if you almost never step to the plate — 25% of the time here — you don't score enough to win the game. A high win rate on a strategy that's mostly in cash is a comfortable feeling, not an edge.

Does trading it on leveraged futures help?

It makes bigger dollars and a worse strategy. That's the leverage trade in one line.

MeasureES futuresSPY ETF
Best net profit (reliable variant)$234,250$96,867
Best return-per-drawdown (CAR/MaxDD)0.320.75
Buy-and-hold drawdown vs. account114% (blown)56%

Same strategy, two instruments, best-of-grid.

The best futures variant made $234,250 — more than twice the best SPY dollar figure — and every one of those dollars is real leverage at work. But its best risk-adjusted return is 0.32 against SPY's 0.75, because leverage scales the drawdowns right along with the gains, and your account doesn't grow to match. Leverage amplified the pain, not the edge. If you trade this on futures, the lesson isn't "size up" — it's "size so a normal drawdown doesn't end you," because a single contract already draws down more than the account.

Which side and style actually work?

Long and patient works. Short is a graveyard on both markets.

StyleSideVariants% profitableWin rateR-expSharpeRet/DDExposure
Mean reversionLong225100%67.6%0.200.1700.1624%
Mean reversionShort2790%42.0%−0.12−0.090−0.0371%
BreakoutLong279100%57.8%0.110.0900.0971%
BreakoutShort2250%32.3%−0.26−0.170−0.0324%

SPY, all four style × side combinations, reliable variants.

Buying oversold z-scores (mean-reversion long) was profitable in 100% of SPY variants and 99% on ES — a genuine, modest tilt. Following strength instead (breakout long) also made money, weaker on SPY (R-expectancy 0.11) but, on the shorter, trendier futures history, it actually produced the single best dollar figure.

The short side is where the fantasy dies, on both markets. Every short variant — fading rallies, shorting overbought spikes — lost. 0% profitable across 504 SPY and 467 ES reliable short variants, median profit factor around 0.70. Shorting a long-term uptrend because a number says "overbought" is a reliable way to give money away. The short side loses.

So what can you actually trade? The stable region

Not the best backtest. The single highest R-expectancy anywhere in the grid was 0.76 on ES and 0.87 on SPY — but both came from a slow 50-day lookback that traded barely 50 times. That's one lucky cell. Trade it and you're trading an accident.

The tradable answer is the opposite of a peak: it's a plateau. Pick the lookback that trades the most — 10 days (median ~196 trades on ES, ~347 on SPY, up to 541) — so the result rests on hundreds of trades, then look at how R-expectancy behaves as you move the entry and exit around. It barely moves.

Shorter lookbacks trade far more often — a median 347 trades at 10 days versus 98 at 50 — for nearly the same R-expectancy. The 10-day lookback is where a result you can trust lives.
R-expectancy across every entry and exit threshold at the 10-day lookback on SPY: a flat 0.17–0.23 plateau. The whole neighbourhood works — the mark of a real edge, not a curve-fit.

Every cell is positive and clustered: on SPY, R-expectancy holds 0.17–0.23 across the whole plane (median 0.18); on ES, 0.12–0.22. Nothing falls off a cliff when you nudge a threshold. That is the signature of a real edge rather than a curve-fit — you don't need the exact settings, because the whole neighborhood works. Median win 71–72%, exposure ~30%. It's a modest, robust, buy-the-dip tilt you can actually deploy — around fifteen to eighteen cents of edge per dollar risked, on a few hundred trades. Just don't expect it to beat holding, and on futures, size it so the drawdown can't end you.

The one to trade, named

If you want a single configuration rather than a region, take the interior cell whose whole neighborhood is strongest and steadiest (boxed in the heatmap above): lookback 10, buy when z drops below −1, take profit when z climbs back above +0.5, with a 5-day time stop. The same setup wins on both markets — a good sign it isn't a fluke.

InstrumentTradesWin rateNet profitR-expectancyPer-trade SharpeReturn/drawdownMax drawdownExposure
SPY ETF52265.7%$86,0380.220.170.6211.4%35%
ES futures30165.1%$140,8000.140.140.1331.6%35%

On SPY that variant earns about 16% of buy-and-hold's money with roughly a fifth of the drawdown (11.4% vs 56%) — a genuine lower-risk tilt, in cash 65% of the time. On futures the leverage doubles the dollars and triples the drawdown (31.6%): tradable, but only if you size it to survive.

The verdict — and the honest limits

The z-score mean reversion strategy is neither a scam nor a money machine — and it is tradable, as long as you trade the stable region, not the headline. The famous 74% win rate is a vanity number; the single best backtest is an overfit peak; and on the leveraged futures I trade, buy-and-hold a single contract doesn't even survive. But the plateau underneath — lookback 10, buy the dip, a steady ~0.15–0.18 R-expectancy on hundreds of trades — is a genuine, deployable edge. Modest, honest, and sized for the drawdown, not the fantasy.

Where the hype is right: buying oversold z-scores on the S&P does make money, consistently, and it wins most of its trades — on the ETF and on futures. Where it's wrong: that win rate is a vanity number. The strategy never beat buy-and-hold across 2,400 tries; its only edge (a shallower drawdown) is bought by sitting in cash 60–75% of the time; shorting is a straight loss; and on the leveraged instrument I actually trade, buy-and-hold a single contract doesn't even survive — so the honest use of z-score there is survival sizing, not a profit engine.

The limits — read them before you act:

  • Frictionless. No commissions or slippage. The busiest variants trade several hundred times; real costs hit an active, thin-edge strategy hardest — the live number is worse than shown, and worse still on futures.
  • Futures data is back-adjusted continuous. That preserves the point-to-point moves the P&L needs, but the price levels aren't literal quotes — read the ES numbers as dollars-per-contract, not index prices.
  • Different windows. ES history is 19.5 years (2007+); SPY is 33.4 (1993+). The two aren't a like-for-like race; they're two honest views of the same idea.
  • In-sample. I swept 1,200 variants per market and then showed you the best — by definition the luckiest on that history. That's the overfitting trap; a winner needs out-of-sample and Monte-Carlo checks before a dollar rides on it, which is exactly what a validation tool like AlgoChef is for.
  • Fixed 1 contract / fixed shares, no compounding, no margin model. Sizing is flat; that keeps the comparison clean and understates neither the leverage risk nor the flat return.

What this means for you

  1. Stop grading strategies by win rate. Ask for the risk-adjusted return and the exposure in the same breath. A 74% win rate that's in cash 75% of the time is not what it sounds like.
  2. Benchmark against buy-and-hold — the survivable version. If a strategy can't beat holding, its job is lower risk, not more money. And on futures, "buy-and-hold" a single contract isn't even the safe option — it blew the account here.
  3. On futures, size for the drawdown, not the win rate. One ES contract already draws down more than a $35k account on a naked hold. The z-score tilt survives only because it's mostly flat — respect that.
  4. Trade the stable region, not the best backtest. Lookback 10, buy the dip (z below −1 to −2), exit near or above the mean — the whole neighborhood holds ~0.15–0.18 R-expectancy on hundreds of trades. Pick anywhere in that plateau; don't chase the single lucky variant. And pair it with something that uses the 60%+ of the time you're idle.
  5. Never short overbought z-scores on an uptrending index. Zero of ~970 short variants across both markets worked. Don't.
  6. Validate before you trust the win rate. Any headline number from a single backtest — including this study's best variant — should survive out-of-sample and resampling first.

Methodology

Data source
SPY (S&P 500 ETF) daily OHLCV 1993-2026, and ES E-mini S&P 500 futures daily 2007-2026, from the StatOasis research dataset.
Date range
SPY 1993-02-02 to 2026-06-12 (33.4 years, 8,398 bars); ES 2007-01-03 to 2026-07-01 (19.5 years, 4,917 bars).
Entry / exit rules
Signal: z-score = (close minus the moving average over the lookback) divided by the population standard deviation. Mean-reversion long buys when z crosses below the entry threshold (swept -1 to -3) and exits when z climbs back above the exit threshold (0 to 1); breakout and short variants are swept too. Decision at the close, fill at the next bar's open — no look-ahead. A protective time exit of 0/5/10/15 days is also swept.
Sizing
$35,000 starting capital, no compounding, flat-only (one position at a time), frictionless. SPY sized in whole shares (floor of capital divided by price); ES sized as one contract at $50 per index point, which is the futures-leverage view.
Overlap mode
Signal-to-signal and flat-only — a new signal is ignored while a position is open. Reliability bar: only variants with 50 or more trades are reported.
Look-ahead
The z-score is computed on the close of bar t and the fill is the open of bar t+1, entries and exits alike. No number uses information unavailable at the decision.
Minimum sample
50 trades. 1,008 of the SPY variants and 934 of the ES variants clear it; only those are reported.
Buy-and-hold benchmark
Same bars, same sizing. SPY: $551,738.32 net, 8.82% CAGR, and a worst drawdown of 56.47% in March 2009 (CAR over that drawdown, 0.098). ES: $274,175.00 net, 11.83% CAGR, and a worst drawdown of 114.33% of the account in March 2009 - one contract against $35,000 is leveraged enough to lose more than the whole account on the way (CAR over that drawdown, 0.060).
Random control
Frequency-matched seeded coin flip, 10 seeds from base seed 20260803, matched to each market's own median reliable variant. SPY (334 entries, 10-bar hold): $34,231.56 net (sd $15,933.68), a 26.79% worst drawdown, CAR over that drawdown 0.097. ES (212 entries, 10-bar hold): $98,056.00 net (sd $67,571.99), a 67.31% worst drawdown, CAR over that drawdown 0.034. Computed by tools/controls_report.py.
Parameter scopeParameters swept

The study searched the parameter space and reports the spread, not one tuned setting.

2,400 variants sweeping lookback, entry threshold (-1 to -3), exit threshold (0 to 1), direction, entry style and a protective time exit of 0/5/10/15 days, on both SPY and ES. The article reports the spread across that space, and the 74% win rate is shown as what it is rather than as a selected best.

Run to v1 of the StatOasis research standard - the rules every study here has to meet before it is published. The version is the study's own: a standard that gained a rule later never reaches back and claims this one met it.

Historical backtest results are not a guarantee of future returns. This content is for educational purposes only and is not investment advice. Hypothetical performance disclosure (CFTC Rule 4.41).

Frequently asked questions

What is a z-score in trading?⌄

A z-score measures how far price has moved from its recent average, in units of its own volatility: z = (close minus the moving average) divided by the standard deviation, over a chosen lookback. A reading of -2 means price is two standard deviations below average, unusually cheap.

Does a z-score mean reversion strategy actually work?⌄

It makes money but never beats buy-and-hold. Across 2,400 variants on S&P futures and SPY, oversold-buy configurations were profitable roughly 100% of the time, yet 0 of 934 futures variants and 0 of 1,008 SPY variants out-earned simply holding.

Does the z-score strategy work better on futures than SPY?⌄

No. On ES futures the best variant made more raw dollars ($234,250 vs $96,867) purely from leverage, but its risk-adjusted return was worse (return-per-drawdown 0.32 vs 0.75). Leverage scales the losses too.

Is buy-and-hold safe on futures?⌄

Not on a single contract with a small account. Holding one ES contract through 2008 drew down $60,387, which is 114% of a $35,000 account, and the account went negative. Buy-and-hold is only safe when it is unleveraged.

Is a high win rate enough to make a strategy profitable?⌄

No, and this study is the cleanest proof. The dip-buy variants win 70 to 80% of trades yet have a median R-expectancy near 0.16 to 0.19 and never beat buy-and-hold. Win rate without a risk-adjusted number is meaningless.

What is the tradable z-score configuration?⌄

Not the single best backtest, but the stable region around it: a 10-day lookback, buy when z drops below -1, exit when z climbs back above +0.5, with a 5-day time stop. On SPY that variant made $86,038 across 522 trades at a 65.7% win rate, a 0.62 return-to-drawdown, and only an 11.4% max drawdown.

What z-score means oversold or overbought?⌄

Conventionally, below -2 is oversold and above +2 is overbought. In this study the buying edge existed across entries from -1 to -3; the tighter the entry, the fewer and higher-quality the trades.

Can you short overbought z-scores?⌄

The data says do not. Every short variant lost: 0% of 504 SPY and 467 ES reliable short configurations were profitable, at a median profit factor around 0.70. Shorting a long-term uptrend on an overbought signal is a losing bet.

Should you use z-score for mean reversion or momentum?⌄

Both made money long-only. Mean reversion (buying dips) was the steadier tilt; breakout (buying strength) topped the dollar table on the trendier futures history. Neither beat buy-and-hold.

↓Download the dataset (1.2 MB)

Read the Strategies, Backtested hub
← Back to Research

Table of contents

  • TL;DR — the answer box
  • How we tested
  • Does a z-score mean reversion strategy beat buy-and-hold?
  • Is the famous 74% win rate real?
  • Does trading it on leveraged futures help?
  • Which side and style actually work?
  • So what can you actually trade? The stable region
  • The verdict — and the honest limits
  • What this means for you
  • Methodology
  • FAQs

Overfit - the newsletter

Skip the hype. Trust the data.One practical takeaway per issue.

Running a quick security check before this can be sent.

Delivered +2 times a month - when the work is ready, not on a calendar.I'll never sell your address.Unsubscribe in one click, any time.

Related Articles

Back to all research
#137

Keltner Channels vs Bollinger Bands: 116,640 Backtests (and the Third Band I Built)

Aug 2, 2025 · 10 min read

116,640 backtests put Bollinger Bands, Keltner Channels and Casey Bands through one identical harness: the band you pick is worth 0.52 on return-to-drawdown, the side you trade is worth 90.4 points — and paying 0.05% a side drops Bollinger below buy-and-hold while the other two hold.

Read more→
#136

Larry Connors R3 Strategy — Rebuilt for Index Futures (With a Smarter Filter)

Jul 26, 2025 · 4 min read

Larry Connors’ R3 strategy still works, if you update it. Discover how volatility filters and index futures give more trades and a better edge.

Read more→

Overfit - the newsletter

Skip the hype. Trust the data.

One practical, evidence-driven takeaway per issue - strategy testing, portfolio construction, market structure, trading psychology, tactical asset allocation.

Running a quick security check before this can be sent.

Delivered +2 times a month - when the work is ready, not on a calendar.I'll never sell your address.Unsubscribe in one click, any time.

StatOasis is calm, evidence-based algorithmic-trading education, founded by Ali Casey. Ali builds systematic trading strategies and teaches the workflow behind them: research, build, test, combine, deploy. He writes the Overfit newsletter, published since 2024, and runs the Algo Trading Masterclass.

Socials

  • X ↗
  • YouTube ↗
  • Instagram ↗
  • LinkedIn ↗
  • GitHub ↗
  • Muck Rack ↗
  • LinkedIn SO ↗
  • GitHub SO ↗

Products

  • Overfit - the newsletter
  • Algo Trading Masterclass ↗
  • StatOasis Community
  • Digital Products
  • 36 Ways to Buy the Dip
  • AlgoChef ↗

Reading & tools

  • Research
  • Methodology
  • Survive the Decade
  • Wall of Love

StatOasis

  • About Ali Casey
  • Contact
© 2026 StatOasis. Calm, evidence-based.
PrivacyTermsHypothetical resultsCalifornia