Methodology & risk note: Two-layer backtested event study on daily OHLCV data: SPY (1993+), QQQ (1999+), DIA (1998+), IWM (2000+). Layer 1: forward-return edge (event mean minus all-bar baseline) at 1/2/3/5/10/20-day horizons with two-sample t-stats at 5 and 10 days. Layer 2: 7 entry families (4 ICT codified from published definitions + 3 simple) × parameter settings (18/18/3/6/3/3/3) × time exits {5, 10, 20} bars × 4 markets = 648 backtests, flat-only, next-open fills, $35,000 starting capital, no compounding, frictionless (no commissions or slippage). Baselines: buy-and-hold and a frequency-matched seeded random-entry strategy (50 seeds, mean). All detection uses only completed bars (no look-ahead). Results are historical and for educational purposes only. Past performance does not guarantee future results. Not investment advice.
TL;DR — the answer box
- I codified the four core ICT / Smart Money Concepts entries — order blocks, fair value gaps, liquidity sweeps, and Optimal Trade Entry — into explicit mechanical rules and backtested them on daily bars. One caveat up front: ICT is usually taught on fast intraday charts, so this tests whether ICT's concepts carry a measurable edge on the daily timeframe, not the intraday discretionary craft.
- Layer 1 asked "is the phenomenon even real?" On SPY, none of the four ICT concepts showed a statistically significant forward-return edge. The strongest was the order block: +0.121% extra return over 5 days across 736 events, t-stat +1.22 — below the standard significance bar of 2.
- Layer 2 ran 648 backtests: 7 entries (4 ICT + 3 simple) × parameter variants × three time exits × four markets (SPY, QQQ, DIA, IWM), against a seeded coin-flip and buy-and-hold. Nearly every variant was "profitable." Exactly zero of 648 beat buy-and-hold on net profit.
- The twist: on SPY, ICT's best entries were not worse than the simple textbook entries. Order blocks were the strongest family on SPY: 81.5% of its variants beat the random baseline, and its best variant, chosen after the fact on the same data it is scored on, made $110,039.85 net at 0.241 return per unit of drawdown. On DIA it holds only per unit of drawdown: the inside-bar breakout's best variant out-earned every ICT family's on net profit ($58,516.07 at 0.084 against $53,758.65 at 0.084 for the liquidity sweep), while the order block's best variant returned 0.219 per unit of drawdown against 0.130 for the best simple entry, RSI Oversold. On QQQ it holds on neither measure: the inside-bar breakout's best variant out-earned every ICT family's ($86,380.43 at 0.073 against $71,311.78 at 0.104 for order blocks), and RSI Oversold's best variant returned the most per unit of drawdown (0.121 against 0.118 for the best ICT family, fair value gaps). ICT's entries have some structure. They just don't have enough to beat doing nothing on net profit.
- No entry family, ICT or simple, survived all four markets against both baselines. The honest verdict: on daily bars, most mechanical ICT variants made money in a rising market (though only 55.6% of liquidity sweep variants did on QQQ) and none beat buy-and-hold on net profit. It is a story you can trade, not a concept that beat both the coin flip and buy-and-hold across all four markets.
How I tested this
Every existing "ICT backtest" you can find is one of three things: a concept explainer with hand-picked chart markups, a code library that detects the patterns but never publishes results, or an anecdote claiming a win rate with no methodology, no data file, and no baseline. Nobody defines the rules mechanically, tests whether the underlying phenomenon is statistically real, and then runs a controlled contest. So that is what I built.
The four ICT concepts, codified. Detection happens at the close of bar t; every fill is the next bar's open — no look-ahead anywhere. All rules are long-only (these are index ETFs, and all four rose over their tested windows; shorting is a different study).
- Order Block (OB): a down-close bar confirmed by an up-impulse within the next K bars (the move must clear M × ATR20). The bar's range becomes the zone; enter long on the first later bar that retraces into it. 18 parameter variants.
- Fair Value Gap (FVG): the classic bullish 3-bar imbalance — the high of bar i−2 sits below the low of bar i, leaving a gap. Enter long on the first later bar that retraces into the gap. 18 variants (minimum gap size, touch level, expiry).
- Liquidity Sweep (LS): price breaks below the N-bar low intrabar (the "stop hunt") but closes back above it — enter long. 3 variants.
- Optimal Trade Entry (OTE): after an upward structure shift, enter on the retrace into the 61.8%–78.6% Fibonacci band of the last swing leg. 6 variants.
The three simple comparison entries: RSI(14) crossing down through an oversold threshold, a close-back-above-SMA pullback, and an inside-bar breakout — 3 variants each. These are the "boring textbook" entries ICT claims to improve on.
The two honesty baselines: buy-and-hold (enter the first bar, hold to the last), and a seeded coin-flip that fires long entries at the same average frequency as the real signals, averaged over 50 random seeds — the "could a monkey do this?" bar.
The standardized exit. Every entry — ICT, simple, and random — uses the same protective time exit, swept across 5, 10, and 20 bars. Holding exits constant isolates exactly the thing ICT sells: entry quality.
Two layers of testing. Layer 1 is an event study: after each signal fires, does the market actually rise more than its baseline drift over the next 1–20 days? This asks whether the phenomenon is real before any trading logic touches it. Layer 2 is the contest: all 7 entries × all variants × 3 time exits × 4 markets (SPY, QQQ, DIA, IWM — daily bars, SPY back to 1993) through the same flat-only backtest engine used in every StatOasis study. $35,000 starting capital, frictionless, no compounding. 648 total backtests.
Are fair value gaps (and order blocks) real on daily bars?
Layer 1 measures the "forward edge": the average return in the 5 (or 10) days after a signal, minus the average return after any bar. If order blocks mark institutional footprints, the days after a retrace into one should beat the market's ordinary drift. Here is SPY:
| Concept | Events (n) | 5-day edge | t-stat | 10-day edge | t-stat |
|---|---|---|---|---|---|
| Order Block | 736 | +0.121% | +1.22 | +0.148% | +1.11 |
| Fair Value Gap | 1,122 | -0.005% | -0.08 | +0.014% | +0.16 |
| Liquidity Sweep | 547 | +0.119% | +0.94 | +0.125% | +0.81 |
| Optimal Trade Entry | 206 | -0.028% | -0.17 | +0.134% | +0.60 |
| RSI Oversold (simple) | 20 | +2.125% | +3.42 | +0.598% | +0.50 |
| SMA Pullback (simple) | 743 | +0.022% | +0.26 | +0.020% | +0.16 |
| Inside Bar Breakout (simple) | 663 | -0.119% | -1.41 | -0.192% | -1.70 |
A t-stat measures whether an edge is distinguishable from luck; the conventional bar is 2. No ICT concept clears it on SPY, QQQ or IWM. Across all four markets and both measured horizons, the ICT concepts produced 32 t-stats and exactly one crossed 2 (Optimal Trade Entry on DIA at the 10-day horizon, t = +2.07), and the other 31 did not. Running many tests and keeping the one that clears the bar is data dredging, which is why the count of tests is printed next to the count of passes.
Fair value gaps deserve a special mention because "gaps get filled" is the load-bearing claim of the whole framework. Gaps do get revisited: 1,122 bullish FVG retrace events fired on SPY alone, each one a price touch back to the top of the gap. But the measured return after that touch is -0.005% at 5 days (t = -0.08) and +0.014% at 10 days (t = +0.16), both far short of the t = 2 significance bar. The market wanders back into gaps constantly, and this sample detects no statistically significant forward-return edge that differs from ordinary drift. The touch is real. On SPY daily bars, no significant forward-return edge follows it.
The one signal that jumps off the table is not an ICT concept. RSI Oversold fired only 20 times on SPY (it needs a genuine washout), but those 20 events preceded an average 5-day return 2.125% above baseline, t = +3.42. It repeats on QQQ (+3.136%, t = +3.09) and DIA (+1.716%, t = +2.31) — though it flips negative on IWM (-1.502%), so treat it as suggestive, not proven. The irony writes itself: the statistically strongest phenomenon in an ICT study is the boring RSI.
Does any ICT entry beat a coin flip or buy-and-hold?
Layer 2 is where concepts become strategies. Every entry, every parameter variant, every time exit, run through the engine on SPY — then scored against the two baselines. Buy-and-hold SPY over the same window: $551,738.32 on $35,000. The frequency-matched coin flip averaged $46,270.44 across 50 seeds.
| Entry | Type | Variants | Profitable | Beat random | Beat buy-and-hold | Best-variant net profit | Best-variant return per unit of drawdown |
|---|---|---|---|---|---|---|---|
| Order Block | ICT | 54 | 100.0% | 81.5% | 0.0% | $110,039.85 | 0.241 |
| Fair Value Gap | ICT | 54 | 100.0% | 44.4% | 0.0% | $100,388.46 | 0.206 |
| Liquidity Sweep | ICT | 9 | 100.0% | 33.3% | 0.0% | $68,897.63 | 0.096 |
| Optimal Trade Entry | ICT | 18 | 94.4% | 0.0% | 0.0% | $39,990.73 | 0.171 |
| RSI Oversold | simple | 9 | 100.0% | 0.0% | 0.0% | $30,639.22 | 0.090 |
| SMA Pullback | simple | 9 | 100.0% | 33.3% | 0.0% | $57,145.41 | 0.120 |
| Inside Bar Breakout | simple | 9 | 100.0% | 33.3% | 0.0% | $89,036.77 | 0.148 |
| Buy-and-hold | baseline | — | — | — | — | $551,738.32 | 0.156 |
| Random entries (50-seed mean) | baseline | — | — | — | — | $46,270.44 | 0.086 |
("Best variant" throughout = the family's highest return per unit of drawdown, chosen after the fact on the same history it is scored on, so it is the top of each family's range and not a setting fixed before the test. The net profit shown is that same variant's, not the family's maximum net profit, because the ranking is risk-adjusted. The ratio is CAGR over the worst percentage drawdown, on all three series including both baselines, so the comparison is one definition throughout. It has to be: the engine's usual CAR/MaxDD ratio annualises by trade frequency, which is fine for ranking variants against each other but meaningless against a buy-and-hold that takes exactly one trade. A best-of-family number picked that way is the top of a search. It is labelled as one here for the reason Bailey, Borwein, Lopez de Prado and Zhu set out: search enough variants and the best of them looks good whether or not an edge exists.)
Three things are true at once, and all three matter.
First: almost everything is "profitable." 100.0% of order block variants made money. 100.0% of FVG variants. 100.0% of the simple entries. This is the trap in a long-only backtest on a rising index, and this SPY test shows it: SPY spent three decades going up, and even the frequency-matched coin flip, which has no entry logic at all, averaged $46,270.44 net on it. "My backtest is profitable" was close to meaningless for the long-only entries tested here, on SPY, QQQ, DIA and IWM daily bars.
Second: on SPY, the best order block and FVG variants outran every simple entry's best. I expected the simple entries to embarrass ICT. They didn't. Picked after the fact on the same data, the best order block variant ($110,039.85, 0.241 return per unit of drawdown over 296 trades) and the best FVG variant ($100,388.46, 0.206) outran the best simple entry (inside-bar breakout, $89,036.77, 0.148). The other two ICT families did not: the liquidity sweep's best variant ($68,897.63, 0.096) trailed the inside-bar breakout's on both measures, and Optimal Trade Entry's ($39,990.73, 0.171) scored higher per unit of drawdown but made less than half its net profit. Order blocks beat the coin flip in 81.5% of their variants, the strongest showing in the study, and that share counts every variant rather than a hindsight pick. Meanwhile Optimal Trade Entry and RSI Oversold never beat random even once on SPY (0.0% of variants). If you came here for "ICT is pure garbage," the data won't give you that. On SPY, defined mechanically, order blocks beat the coin flip in most of their variants and FVGs' hindsight-best variant outran every simple entry's, though neither family's forward-edge t-stat reached 2.
Third: nothing beat buy-and-hold on money. Nothing. Not one of the 648 backtests — no ICT variant, no simple variant, no market, no time exit — beat holding the index it trades on net profit. The best cherry-picked order block variant made $110,039.85 while buy-and-hold made $551,738.32 on the same data. That is not a gap any of the seven tested entries closed. Every tested entry is time-limited and buy-and-hold is not, and nothing here decomposes how much of that gap the difference in time spent in the market accounts for.
What survives across four markets?
One market can flatter any strategy, so this study sets its own survival rule: a family survives only if its best variant beats both buy-and-hold and the random baseline on all four tested markets. The mechanical rules sold as "smart money leaves footprints" are tested on every market in the sample, not on SPY alone, so the whole contest repeats on QQQ, DIA, and IWM. Best-variant return per unit of drawdown (CAGR over the worst percentage drawdown) per family:
| Family | SPY | QQQ | DIA | IWM |
|---|---|---|---|---|
| Order Block | 0.241 | 0.104 | 0.219 | 0.113 |
| Fair Value Gap | 0.206 | 0.118 | 0.158 | 0.174 |
| Liquidity Sweep | 0.096 | 0.065 | 0.084 | 0.085 |
| Optimal Trade Entry | 0.171 | 0.093 | 0.180 | 0.229 |
| RSI Oversold | 0.090 | 0.121 | 0.130 | 0.193 |
| SMA Pullback | 0.120 | 0.095 | 0.100 | 0.116 |
| Inside Bar Breakout | 0.148 | 0.073 | 0.084 | 0.119 |
The strong SPY numbers do not travel. Order Block's 0.241 on SPY, the highest family-best anywhere, drops to 0.104 on QQQ, well under that market's 0.122. FVG's best market is SPY (0.206), its worst is QQQ (0.118). The family that leads on one market is mid-pack on the next, and no family holds its rank across all four. A ranking read off SPY alone would have picked a leader that QQQ put third.
And the formal survival test — does any family's best variant beat both buy-and-hold and the random baseline on all four markets? — returns an empty set. Survivors: none. Not one of the seven entry families, ICT or simple. (For scale: buy-and-hold made $457,482.92 on QQQ, $199,069.84 on DIA, and $170,987.43 on IWM. Zero variants cleared those bars either — the 0.0% beat-buy-and-hold column repeats on every market.)
The closest thing to a survivor deserves its own sentence. Drop the buy-and-hold requirement and ask only which family beat the coin flip in a majority of its variants on every market, and exactly one name comes back: Order Block, with 81.5% of variants beating random on SPY, 66.7% on DIA, 59.3% on QQQ and 55.6% on IWM. No other family, ICT or simple, managed a majority on even three markets (FVG came nearest at 44.4% SPY / 37.0% DIA / 63.0% QQQ / 55.6% IWM). So if the study has one consistent, cross-market pocket, it is mechanical order blocks. That consistency is a count of parameter variants beating the average of 50 coin-flip runs, not a significance test, and the order block's own Layer 1 t-stats never reach 2 on any market. The family still captured only a fraction of the index's own drift everywhere it was tested. That is the fair ceiling on the framework's most defensible concept.
The verdict — and the honest limits
Mechanical ICT on daily bars is a middling entry system wearing an extraordinary story.
As standalone phenomena the concepts did not clear the bar: across four markets and two horizons, 31 of 32 ICT scores fell short of it and one crossed. The entries built from them are not losers: comparing hindsight-best variants, order blocks beat the simple textbook entries head-to-head on SPY, and most order block variants beat a coin flip on every market. But nothing in 648 backtests beat buy-and-hold on net profit, nothing survived all four markets, and the concept marketed as ICT's precision instrument (Optimal Trade Entry, run on the 61.8% to 78.6% Fibonacci band and a control band) beat random entries on SPY in none of its 18 variants. ICT teaches that these patterns mark institutional footprints. This study did not measure institutions. Whether large orders leave traces a chart can read belongs to market microstructure, the field that studies how orders turn into prices, and answering it takes order-level data rather than daily bars. What it did measure, forward returns and results against two benchmarks, shows no edge that beats both benchmarks on all four markets.
The honest limits, stated plainly:
- This is daily bars, tested mechanically. ICT is taught as an intraday, discretionary craft. This study proves nothing about what a skilled discretionary trader does at 9:47am on a 5-minute chart. What it does show: the concepts, reduced to their published mechanical definitions on the timeframe where anyone can verify them, cleared t = 2 in 1 of 32 tests, and no entry family built from them beat both benchmarks on all four markets. This test did not find ICT's value in the mechanical daily-bar core run under one shared exit ladder. Untested settings this test left out include intraday timeframes, exits tuned to each concept, and the discretion, which is the part you cannot backtest, audit, or learn from a win-rate screenshot.
- Frictionless results. No commissions or slippage, so no cost impact was measured. What is measured is the trade count: buy-and-hold pays one round turn while the best SPY variants alone took 76 to 310, so under the same per-turn cost assumption the strategies would absorb far more transaction-cost drag than buy-and-hold's single trade.
- RSI Oversold's Layer-1 edge is thin-sample. 20 events on SPY, 11 on QQQ. Statistically loud, but flag-and-verify territory — and it flipped negative on IWM.
- "Profitable" is a low bar in this sample. Even long-only entries with no logic made money on average here: in this frictionless SPY sample the frequency-matched coin flip averaged $46,270.44 net across 50 seeds (MAR 0.086), and six of the seven SPY families had every variant profitable. That is exactly why the baselines exist, and why "0 of 648 beat buy-and-hold on net profit" is the number to read instead.
What does this mean for you?
- "Profitable in a backtest" is the lowest bar in trading — demand baselines. 100.0% of order block variants were profitable, and 0.0% beat buy-and-hold on net profit. In this study, the buy-and-hold and random-entry comparisons are what changed the conclusion, and a result without them is missing the context that mattered here.
- Entry quality is worth less than you think. ICT's entire pitch is superior entries. Even the winning family's best variant, chosen after the fact (order blocks, $110,039.85 at 0.241 return per unit of drawdown), captured about a fifth of what holding the index captured. This study held exits and sizing fixed, so it cannot tell you how much a better exit or sizing rule adds. It does show that the best entry, picked with hindsight, did not close the gap to doing nothing. Test your entry against holding the index before you spend another month refining it.
- If a concept can't be written as a rule, you can't own it. Everything tested here was codified from ICT's own published definitions. The moment a framework retreats to "you had to be there, it's discretionary," you have no way to distinguish skill from survivorship in the reporting.
- Test across markets before you believe anything. Order Block's SPY numbers stood out at first glance: 81.5% of its variants beat the frequency-matched random-control mean. Then QQQ cut its best score, picked after the fact on each market, from 0.241 to 0.104. One-market results are auditions, not verdicts.
- The boring things won the headline measures. The strongest statistical signal in this entire ICT study was RSI dipping below 25. On net profit, the one rule with no entry logic at all, holding the index, beat every one of the 648 backtests. On return per unit of drawdown it did not: SPY's hindsight-best order block variant scored 0.241 against buy-and-hold's 0.156, which is why the two measures are reported side by side.
If you want to check whether a result like this holds up on your own setup — ICT or anything else — that is what AlgoChef is built for: you bring a backtest you have already run, and it grades the edge rather than re-running it. StatOasis.com/Overfit







