Methodology & risk note: Backtested event study of SPY daily price data, 1993-2026 (8,397 opening gaps), frictionless. Results are hypothetical and not investment advice. Past patterns don't guarantee future results. Full method and disclaimer below.
TL;DR
- Direction is almost noise. Up-gaps and down-gaps drift up by nearly the same amount over the next five days (about +0.2% either way). Whether the market gapped up or down tells you very little about what comes next.
- Size is the signal. The bigger the gap, the bigger the forward drift, and it climbs in order from the 0.5% bucket upward: a gap over 2% averages +0.368% the next day and +0.734% over five, versus roughly +0.04% for the smallest gaps.
- Volatility-relative size is the sharpest read. Measured against the market's own recent range (ATR), the edge concentrates once a gap clears 1 ATR. Those gaps average +0.200% next-day, far above every smaller bucket.
- Gaps fill fast and almost always. Down-gaps fill 98% of the time and up-gaps 94%, usually the same day (median zero bars).
- Long ranks above short, and neither beats luck. Every one of the top ten variants by risk-adjusted return is long, and the median short variant loses under all three lenses. But all three long medians sit below a seeded random-entry control's CAR/MaxDD of 0.036, inside its spread. 348 of the 1,080 short variants still made money, against 732 of the 1,080 long ones.
You're reading the one number that doesn't matter
SPY gapped this morning, and your feed already told you what it means. Half of it says "gap up, that's momentum, get long." The other half says "it'll fill, fade it." A gap down flips both scripts to fear. You're expected to pick a side on direction before the first regular-session trade even prints.
Here's what that instinct costs you. We measured every SPY opening gap since 1993, all 8,397 of them in the ETF's daily history, and direction, the thing everyone argues about, is almost worthless. Over the next five days, up-gaps drift +0.206% and down-gaps +0.194%. That's a rounding error apart, and both point the same way: up. The single most-repeated belief about gaps, that the direction tells you where the market goes next, falls apart the moment you count.
So put down the direction and pick up the thing that actually carries information: the size of the gap. That's the whole study in one line. A gap is not one number, and how you measure its size decides whether you even see the edge. Measured three different ways across 33 years, the three views disagree, and the disagreement is the finding. Let's walk through it, one question at a time.
What counts as an "opening gap"?
A gap is just the distance between yesterday's close and today's open, the price jump that happens overnight, before the regular session does anything. Every trading day produces one (even a tiny one), so the question isn't whether a gap exists but how big it is and which way it points.
Here's the dataset, stated plainly so you know exactly what's behind every number below:
- Instrument: SPY, the S&P 500 ETF, daily bars.
- Window: February 1993 through June 2026, 33 years.
- Events: 8,397 opening gaps detected and measured.
- Costs: frictionless. No commission or slippage. The magnitudes here are relative, not what you'd net after trading costs.
Every gap gets tagged by direction (up, down, or flat) and by size, and the forward outcome (what the market did over the next 1 and 5 days) is recorded. A "fill" means price traded back to the prior close at some point (an intraday touch counts). This is measurement, not a trading system: it tells you the base rates, and you decide what to do with them.
Is a gap up bullish and a gap down bearish?
Short answer: not really. The direction of the gap barely moves the forward odds. Both up-gaps and down-gaps are followed by a small upward drift, and the two are close enough that direction alone barely separates the outcomes in this sample.
Here's the split:
| Gap direction | Events | Avg next-day return | Avg 5-day return |
|---|---|---|---|
| Down | 3,659 | +0.055% | +0.194% |
| Up | 4,625 | +0.040% | +0.206% |
| Flat | 113 | +0.008% | +0.134% |
Look at the five-day column. A down-gap is followed by +0.194% on average; an up-gap by +0.206%. That's a rounding error apart. Over a single day the down-gaps are very slightly hotter (+0.055% vs +0.040%), but again, the gap is tiny and points the same way (up) regardless of which direction the market gapped.
This is the myth worth killing. People treat a gap down as a warning and a gap up as a green light. The data says both just drift up modestly afterward. SPY has a structural upward tilt in this sample, and gap direction does not reverse it. If your plan is "short the gap down because it's weak," you're betting against the market's own drift on the basis of a signal that isn't there. For the longer history behind that tilt, see our count of every S&P 500 drawdown since 1871.
Another way to see how little direction matters: ask how often price traded at least 1% above the event-day close at some point within the next 20 sessions. After down-gaps, 84% of the time. After up-gaps, 83%. One percentage point apart. The market's upside availability over the following month is essentially indifferent to which way the day opened.
And the quiet oddity in the table is the row nobody trades: the flat open. It's rare (just 113 of 8,397 sessions opened exactly at the prior close) and it's also the sleepiest cohort on all three forward measures: the weakest next-day drift (+0.008%), the weakest five-day drift (+0.134%), and the lowest 20-day +1% touch rate (78%). Which is its own small confirmation of the study's theme: the presence of an overnight move coincided with larger subsequent moves than a flat open. A session that opens exactly where it closed was the weakest of the three direction cohorts on those three outcomes.
So if direction is a dead end, what isn't?
Size is the signal
This is the finding the whole study turns on. Forget which way the gap pointed and look at how big it was as a percent of price. The bigger the gap, the bigger the forward drift, from the half-percent bucket upward. Below that it does not climb in order: the 0.25 to 0.5% bucket sits below the smallest bucket and goes negative next-day.
| Gap size (% of price) | Events | Avg next-day return | Avg 5-day return |
|---|---|---|---|
| Under 0.25% | 3,839 | +0.043% | +0.190% |
| 0.25-0.5% | 2,233 | −0.007% | +0.116% |
| 0.5-1% | 1,582 | +0.059% | +0.243% |
| 1-2% | 594 | +0.147% | +0.337% |
| Over 2% | 149 | +0.368% | +0.734% |
Read the next-day column from the bottom up. A gap over 2% averages +0.368% the following day, roughly nine times the drift of the smallest gaps and about nine times the +0.0404% an ordinary session averaged across the whole window. Over five days, that biggest bucket averages +0.734%. The pattern is monotonic from the half-percent bucket on up: more size, more drift. The very smallest gaps (under a quarter percent) are basically noise, so small they barely qualify as events.
The intuition is simple. A big overnight gap is the market repricing hard on real news, and that repricing tends to have momentum behind it for a few days. A tiny gap is just the open landing a hair away from yesterday's close. That bucket averaged +0.043% the next day, less than every bucket at or above 0.5%. Size is doing the work that direction can't.
There's a catch, though. How you measure "big" matters a lot, and one common way of measuring it plants a false dead zone in the middle of the size buckets.
Why measure size three ways?
Because "big" isn't one thing. A gap of two points was enormous in 1995 when SPY averaged $54.31, and trivial at the $741.75 it closed at on the study's last bar. So we measured every gap's size three different ways and compared what each one revealed:
- As a percent of price, the simplest single number, shown above.
- In raw index points, the way a lot of people instinctively think about it.
- Relative to the market's own recent range (ATR), sizing the gap against how much the market has been moving lately.
Running all three side by side is the honest way to do this. If an edge only shows up under one definition, it's probably an artifact. If it holds across all three, it is not an artifact of one ruler. And the comparison turns up something genuinely useful: the raw-points lens breaks, while the volatility-relative lens sharpens the next-day read.
The points "dead zone"
Watch what happens when you measure gaps in raw index points instead of percent:
| Gap size (points) | Events | Avg next-day return | Avg 5-day return |
|---|---|---|---|
| Under 0.5 | 4,593 | +0.042% | +0.181% |
| 0.5-1 | 1,738 | +0.046% | +0.230% |
| 1-2 | 1,214 | −0.023% | +0.007% |
| 2-4 | 582 | +0.206% | +0.605% |
| Over 4 | 270 | +0.075% | +0.341% |
There's a hole right in the middle. The 1 to 2 point bucket goes negative next-day (−0.023%) and flat over five days (+0.007%), even though the buckets on either side of it are solidly positive. That's not a real pattern. It's a measurement artifact. The 1 to 2 point bucket is a blender: it mixes a 1.5-point gap from 1995 (a giant move, 2.8% of that year's average price) with a 1.5-point gap at the study's last close in 2026 (a rounding error, 0.20% of that $741.75 price). Throw those into the same bucket and they cancel out into a false "dead zone."
This is exactly why raw points mislead. They mash together completely different market eras under the same label. The percent-of-price lens doesn't have this problem, and the next lens fixes it even more cleanly.
The ATR tell: the edge lives past 1 ATR
The sharpest way to size a gap is against the market's own recent volatility: its average true range, or ATR, a standard gauge of how much the market has been moving day to day. A gap "worth 1 ATR" is as big as a typical full day's range. Here's what that reveals:
| Gap size (vs ATR) | Events | Avg next-day return | Avg 5-day return |
|---|---|---|---|
| Under 0.25 ATR | 4,445 | +0.040% | +0.190% |
| 0.25-0.5 ATR | 2,402 | +0.039% | +0.142% |
| 0.5-0.75 ATR | 950 | +0.047% | +0.345% |
| 0.75-1.0 ATR | 329 | +0.055% | +0.237% |
| Over 1.0 ATR | 251 | +0.200% | +0.333% |
Notice how flat the first four buckets are next-day, all clustered between +0.039% and +0.055%. Then the top bucket, gaps over 1 ATR, jumps to +0.200%, 3.6 to 5.1 times the buckets below it. The edge doesn't build gradually here; it concentrates. Below 1 ATR, the next-day return barely moves. Above it, the gap is genuinely outsized relative to how the market has been behaving, and that's where the next-day drift lives. The five-day column has no such step: the 0.5 to 0.75 ATR bucket's +0.345% edges out the +0.333% after gaps over 1 ATR.
This is the sharpest tell in the data. A gap of a given point-size or even percent-size means different things in a calm market versus a jumpy one. Sizing it against ATR normalizes for that automatically. In this sample the next-day return stays relatively flat below 1 ATR, then jumps to +0.200% once the gap clears that line. If measuring moves against volatility is new to you, our beginner's guide to the VIX and market volatility is a good companion read on why a move's meaning depends on the market's mood, not its raw size.
Do gaps actually fill?
Yes, overwhelmingly, and fast. A "fill" means price traded back to the prior close at some point. By direction, down-gaps fill 98% of the time and up-gaps 94%, and the typical gap filled the same day it opened (a median of zero bars).
| Gap size (% of price) | Fill rate | Median bars to fill |
|---|---|---|
| Under 0.25% | 95% | 0 (same day) |
| 0.25-0.5% | 96% | 0 (same day) |
| 0.5-1% | 92% | 1 |
| 1-2% | 87% | 2 |
| Over 2% | 83% | 2 |
Both are near-certainties. The pattern in the table is intuitive: small gaps almost always fill, and they fill instantly, because price barely moved away from the prior close in the first place. Bigger gaps fill a bit less often and take a couple of days: fill rates fall from 95% and 96% in the two smallest buckets to 83% above 2%, and median bars to fill rise from zero to two.
Here's the important nuance, and it's where the "gaps always fill, so fade them" crowd gets it wrong. The same big gaps that fill less reliably are the ones that carry the forward drift from the size section above. A gap over 2% fills 83% of the time, still likely, but it's also the bucket that averages +0.368% the next day and +0.734% over five. So the high fill rate and the upward drift aren't in conflict: both are totals over the same events, measured separately. Nothing here follows a single gap through a fill and then onward, so the ordering is not something these numbers establish. A high fill rate is not, by itself, a reason to bet against that bucket's measured +0.368% next-day average.
Long ranks above short, and neither beats luck
Everything above is measurement of what gaps do. The last question is what happens when you actually try to trade them, long versus short. To check, we ran a backtest sweep: 2,160 variants across the three size lenses, both directions, and a range of holding periods, then ranked them by risk-adjusted return (return relative to worst drawdown). One result dominates: the long side's medians beat the short side's under all three lenses. Neither side clears the matched random control's spread, so this ranks the two sides rather than showing either has beaten luck. All three long medians sit below that control's own CAR/MaxDD of 0.036, inside its seed-to-seed spread of 0.072.
- Every single one of the top 10 reliable variants (50 or more trades) by risk-adjusted return is Long. Not most. All of them.
- By the median across reliable variants, Long is positive and Short is negative under all three size lenses. Long's median risk-adjusted return runs around +0.010 to +0.020 depending on the lens; Short sits at roughly −0.020 across the board.
- The strongest pocket by median risk-adjusted return is long on a down-gap over 2%, profitable in 100% of its 8 reliable variants. Consistency is not confined to large gaps, though: long on the smallest down-gaps (under 0.25%) was profitable in 100% of 36, and long on a down-gap over 1 ATR in 85% of 20.
- The standard regime filters (is volatility rising, is the trend up) helped only marginally here, a median improvement of about +0.01. The filters added little on top of the size buckets they were applied to.
I've watched traders short a scary gap-down on pure instinct, younger me included, and hand the drift straight back to the market. Long variants outscored short on median CAR/MaxDD under all three size lenses, and SPY itself rose over this 1993-2026 window. Shorting gaps, even "obviously weak" gap-downs, sits on the losing side of that ranking. It's the same long-side ranking that fell out of our 33,792-backtest oscillator study, where 89.7% of long variants made money against 8.8% of short. On SPY, the median short variant keeps losing to the tape.
The verdict, and the honest caveats
Direction is the number everyone watches and the one that carries the least. Size carries the signal, ATR reads it most cleanly, gaps fill fast but that isn't a fade signal, and the long side outscores the short side under all three lenses. That's the study.
None of it is a finished trading system, though, and it would be dishonest to present it as one. Three things to keep front of mind:
- Frictionless. Every number here is computed with no commission and no slippage. The magnitudes are relative, useful for comparing one bucket against another, not what you'd actually net after costs. No cost level was tested here, so nothing here says which of these drifts survives friction and which does not.
- An event study, not a system. This measures base rates. The events overlap (we're cataloguing what gaps do, not trading them sequentially), and the backtest sweep is a flat-only measurement tool, not a tuned strategy. The job here is to replace gut-feel myths with measured odds. Turning those odds into a robust strategy is a separate, careful piece of work, the part where most retail strategies quietly die, which is why we treat robustness testing as its own discipline.
- Thin cells are flagged, not trusted. Any backtest variant with fewer than 50 trades is flagged as low-reliability, not dropped. The size, ATR and fill tables rest on buckets of at least 149 events each (the smallest size bucket above 1 ATR still holds 251 events), and the long-side medians use only variants at or above the 50-trade floor (1,808 of the 2,160 clear it, long and short together). But the more granular a cell gets, the more its exact number can wobble on fresh data. Lean on the broad patterns, not the third decimal place.
The findings at a glance
| Finding | The number | What it means for a trader |
|---|---|---|
| Direction | up-gaps +0.206% / down-gaps +0.194% over 5 days | Don't trade the gap's direction. Both just drift up. |
| Size | over-2% gaps average +0.368% next-day, climbing at every step above 0.5% | The bigger the gap, the bigger the forward drift. |
| ATR lens | edge concentrates past 1 ATR (+0.200% next-day) | Size the gap against recent volatility, not raw points. |
| Fills | down-gaps 98%, up-gaps 94%, usually same day | Gaps fill almost always, but the largest bucket still shows upward drift as a separate total. |
| Side | every top-ten variant by risk-adjusted return is Long; the median Short loses in all lenses | Long ranks above short, but neither side beats a random-entry control. Shorting fights the market's drift. |
What this means for you
- Stop trading the gap's direction on its own. Up or down barely changed the next five days (+0.206% vs +0.194%). If your rule starts with "gap down means weak," it's built on a signal that isn't there.
- Judge the gap by size, and size it against volatility. Percent-of-price works; ATR works best. The next-day drift concentrated once a gap cleared 1 ATR. A gap that's small relative to how the market's been moving is just noise.
- Don't fade a big gap just because "it'll fill." The biggest gaps fill a little less often (83% for 2%+ gaps) and still show the largest average forward drift. Those are separate totals, not a path this test followed through a fill.
- If you trade gaps, the long side ranks above the short. Every top-ten variant by risk-adjusted return was long, and the median short variant lost under all three size lenses. On SPY, the short side fights a 33-year tailwind. Even the long medians sit below a random-entry control's 0.036, so long is the better of two sides, not an edge over luck.
- Treat this as base rates, not a system. These are the odds. Building a tradeable, cost-aware, risk-managed strategy on top of them is separate work. These numbers carry no commission or slippage, so they cannot tell you which drift survives trading costs.
Methodology: how this data was generated
This is a backtested event study, not live trading results. These numbers come from our own event-study research engine, which we built in-house and run over the full SPY price history, not from third-party summaries or reproduced figures. Here's exactly how it was built, in plain terms.
- Data source: Daily OHLCV (open, high, low, close, volume) price data for SPY, the S&P 500 ETF.
- Date range: February 1993 through June 2026, 33 years of daily bars.
- What an "event" is: Any opening gap, the difference between the prior close and the current open. All 8,397 gaps in the history were detected and measured. Each was tagged by direction (up / down / flat) and by size under three lenses: percent of price, raw index points, and size relative to ATR (average true range, a standard volatility gauge).
- Forward-outcome measurement: For each gap, we recorded what happened over the following 1 and 5 trading days (average return), and whether and when price filled back to the prior close (an intraday touch counts; median bars-to-fill reported). This is measurement of base rates, not a trading system. Events are allowed to overlap.
- The trading sweep: Separately, we ran a flat-only backtest of 2,160 variants (the three size lenses x direction x holding periods), starting from $35,000 in capital, ranked by risk-adjusted return (return relative to maximum drawdown), to check which side actually pays.
- Frictionless assumption: Results are computed without commissions or slippage. Real-world trading costs would reduce any edge shown here. Treat these as indicative base rates, not net-of-cost returns.
- Reliability: Findings rest on large samples drawn from 33 years of data. Any backtest variant with fewer than 50 trades is flagged as low-reliability rather than trusted. The size, ATR and fill tables rest on buckets of at least 149 events each.
Whether it's worth trading a gap at all also depends on the regime you're in, which is a study of its own. See mastering market regimes: when to trade and when to stay out. For more on why even a clean backtest is not the same as a live edge, watch Why Most Profitable Backtests Fail in Live Trading on the StatOasis YouTube channel.
Disclaimer
All of these results are derived from historical backtesting using daily SPY price data and do not represent actual trading results. Backtested performance is hypothetical. Past performance of any pattern does not guarantee future results. This article is for educational and informational purposes only and does not constitute investment advice. StatOasis is not a registered investment advisor. Nothing here is a recommendation to buy or sell any security. Please consult a licensed financial professional before making any investment decision.
Transparency: StatOasis sells trading-education products. Our research is produced independently and is not altered to favor a sale.
Freshness: The data is current through 2026. We re-review these studies when the underlying dataset is extended.
Get the next study in your inbox
This is one study in an ongoing series. If you want the next one, the same kind of measured, myth-busting, frictionless-but-honest backtest, join The Overfit newsletter at StatOasis.com/Overfit. It's where we publish the data behind the trading ideas everyone argues about.
A gap opened. Now you know which number to read, and which one to ignore.






