Methodology & risk note: Backtested event study on SPY daily OHLCV data, 1993-02-02 to 2026-06-12, totalling 8,398 bars. Layer A: 2,304 strategy variants (RSI, Stochastic, Williams %R sweep), of which 2,068 were eligible (at least 30 IS trades). Layer B: 5 LLM-generated rule-sets implemented verbatim. In-sample: 1993-02-02 to 2016-06-03 (5,878 bars). Out-of-sample: 2016-06-06 to 2026-06-12 (2,520 bars). $35,000 starting capital. Flat-only sequential backtest (no compounding). Frictionless (no commissions or slippage). Signals read on the close, entries filled at the next open. All results are historical and for educational purposes only. Past performance does not guarantee future results. Not investment advice.
TL;DR — the answer box
- I ran two experiments on SPY. First: a 2,304-variant optimizer sweep to simulate what AI does when it searches for "the best strategy." Second: 5 rules generated verbatim from an AI prompt, tested on the locked out-of-sample window.
- The optimizer found 25 top in-sample picks with a median risk-adjusted score (CAR/MaxDD) of 0.80. Out-of-sample, that score collapsed 85% to 0.12. All 25 stayed profitable — but 0 of 2,068 eligible strategies beat simply buying and holding SPY, which returned +$87,835 out-of-sample.
- Of 5 AI-generated rules, all 5 made money out-of-sample. 2 cleared the 95th Monte Carlo percentile: Rule 4 (+$31,954, CAR/MaxDD 0.1185, percentile 96.7) and Rule 5 (+$2,591, CAR/MaxDD 0.0740, percentile 95.5). Rule 1 fell short at 85.5, and random entries matched or beat Rules 2 and 3 more than half the time (percentiles 43.8 and 44.2).
- Even the stronger of the two, Rule 4, lost to buy-and-hold by more than $55,000.
- The verdict on these five rules: impressive backtests, not guaranteed edges. An out-of-sample test and a Monte Carlo screen are what separated the five AI rules here: the first measures whether a result holds up on data the rule was not built on, the second whether its entry timing beats random entries.
How I tested this
If you have ever typed "build me a profitable trading strategy for SPY" into an AI and stared at the output wondering whether it actually works — this study is for you.
I ran two separate experiments, each designed to answer a different version of the same question.
Experiment A — the optimizer sweep:
I ran a systematic sweep across 2,304 SPY daily strategy variants. Think of this as one brute-force optimizer-and-selection workflow: try a massive number of combinations, pick the best-looking one in the backtest window, and hand it to you as the answer. I used three momentum oscillators — RSI, Stochastic %K, and Williams %R — swept across multiple lookback lengths, entry conditions, trade directions (long and short), entry styles (mean-reversion and breakout), and time exits.
The data: SPY daily price history from 1993-02-02 to 2026-06-12, totalling 8,398 trading bars. I split this 70/30. In-sample (IS), used for building and selecting strategies: 1993-02-02 to 2016-06-03, covering 5,878 bars. Out-of-sample (OOS), the blind test window the strategies have never seen: 2016-06-06 to 2026-06-12, covering 2,520 bars. The OOS window was locked from the start and never touched during the sweep. Holding a slice back and scoring on it is the standard way to judge anything that was tuned on data.
All results are frictionless: no commissions, no slippage, next-open fills, $35,000 starting capital, flat-only (no compounding). Of 2,304 total variants, 2,068 (89.8%) qualified with at least 30 in-sample trades. The 236 variants below that eligibility floor are excluded from the eligible population. The eligible pool splits across the three indicator families as 532 RSI variants (25.7%), 768 Stochastic (37.1%), and 768 Williams %R (37.1%), and the three families landed in the same place out of sample: 50.0% of each family stayed profitable, every family's median return-to-drawdown was slightly negative, and none of the 2,068 beat buy-and-hold. That pattern (half staying profitable, a slightly negative median return-to-drawdown, none beating buy-and-hold) held separately within each of the three indicator families, not just in the pooled 2,068.
Experiment B — 5 AI-generated rules, verbatim:
I gave an AI a single prompt and asked it to generate 5 fully mechanical daily trading strategies for SPY using a fixed indicator set. I then implemented each rule exactly as written — no tweaking, no parameter adjustment — and ran it through the same engine and the same IS/OOS split.
Here is the exact prompt I used:
"You are a systematic trading strategy developer. Build me 5 profitable daily trading strategies for SPY (the S&P 500 ETF). Each strategy must use ONLY these indicators: RSI(14), Stochastic %K(14, smooth=3), Williams %R(14), SMA(100), ATR(20), 1-day prior close-to-close return, or 5-day prior return. For each strategy provide: (1) the exact entry condition with the indicator name and the precise numeric threshold, (2) the entry style — mean-reversion or momentum/breakout, (3) the direction — long or short, (4) a time-based exit in trading days — choose from 5, 10, or 15 bars. Every strategy must be fully mechanical with zero discretion so it can be coded and backtested immediately."
The validation layer — Monte Carlo:
For the 5 AI rules, I also ran a Monte Carlo test. Monte Carlo percentile answers one specific question: if a strategy entered trades on random bars with the same frequency as this rule, what share of those random simulations would do worse? Rule 4's percentile of 96.7 means random-entry runs matched or beat it 3.3% of the time. Rule 2's 43.8 means they matched or beat it 56.2% of the time, which is what picking trade entries out of a hat looks like. The bar a rule has to clear here is the 95th percentile.
Finding 1: Does the optimizer's single best in-sample pick survive out-of-sample?
The optimizer's top pick, the one this selection rule hands back as its best result, was a Williams %R(2) long mean-reversion strategy with a 10-bar time exit. In-sample, the numbers looked exceptional.
| Metric | In-Sample | Out-of-Sample |
|---|---|---|
| Net Profit | $98,354.68 | $23,365.85 |
| CAR/MaxDD | 1.4610 | 0.0980 |
| Win Rate | 73.58% | 70.62% |
| Trades | 405 | 160 |
CAR/MaxDD stands for Compound Annual Return divided by Maximum Drawdown, both in percent, so it measures return against worst-case loss. A score of 1.0 means the annual return, in percent, matched the maximum drawdown, in percent. The in-sample score of 1.46 was the highest of the 2,068 eligible variants, against a top-25 median of 0.80. The out-of-sample score of 0.10 is a fraction of that.
The win rate held (73.6% IS, 70.6% OOS). The strategy stayed profitable. But the risk-adjusted quality dropped sharply from 1.46 to 0.10 — and the strategy made $23,365.85 over the OOS window while simply buying and holding SPY made $87,835 over the exact same period. The optimizer's top pick trailed buy-and-hold by more than $64,000.
Finding 2: What happens to the top 25 "AI-selected" strategies when tested on new data?
One bad pick might be a fluke. What if you take the top 25 — the cream of 2,068 eligible strategies, ranked by in-sample risk-adjusted score? This is the shortlist the optimizer proxy in this study generates.
| Metric | Top 25 — In-Sample | Top 25 — Out-of-Sample |
|---|---|---|
| Median CAR/MaxDD | 0.8032 | 0.1242 |
| Still profitable | — | 25 of 25 (100%) |
| Beat buy-and-hold | — | 0 of 25 (0%) |
Median CAR/MaxDD dropped from 0.80 to 0.12 — an 85% collapse in the risk-adjusted score.
All 25 stayed profitable out-of-sample. They are not blown-up, money-losing strategies. But not one of them beat the baseline: buy-and-hold SPY returned $87,835 over the OOS window. The best of the best from a 2,304-variant sweep could not match a strategy with zero intelligence — just hold the index.
An in-sample backtest is like a job interview where the candidate designed the questions themselves. They score close to perfect every time. The out-of-sample test is the actual job. Acing the custom interview and excelling at the real work are two very different things.
Finding 3: Across all 2,068 eligible strategies, does any optimized variant beat buy-and-hold out-of-sample?
Not one.
| Metric | Whole Eligible Population (OOS) |
|---|---|
| Eligible variants | 2,068 |
| Profitable OOS (net > $0) | 1,034 (50.0%) |
| Beat buy-and-hold OOS | 0 (0.0%) |
| Median OOS net profit | $0.00 |
Exactly half of the 2,068 strategies made money out-of-sample. The other half did not. The median out-of-sample net profit was exactly $0.00. And none of them, across the full eligible population, crossed the $87,835 buy-and-hold line.
The optimizer is extremely good at one thing: finding parameters that fit the past data it was shown. It is not finding anything that beat buying and holding. The shapes it finds look like edges in the training window, and the top 25 by in-sample score lost a median 85% of that score in the test window.
This does not mean systematic trading is hopeless. The buy-and-hold result itself is a systematic strategy, and it had the highest out-of-sample net profit in this comparison. What it shows is that adding oscillator conditions to a basic long-index position is not automatically adding edge. The OOS window (2016 to 2026) was a rising decade for SPY. Whether that drift is what decided the comparison was never tested here.
The benchmark nobody optimized: a published public rule
There is one more comparison in the data, and it is the most deflating one for the optimizer. Alongside the sweep, I ran a fixed benchmark rule with no parameter search: the classic RSI(2) long mean-reversion strategy from Larry Connors and Cesar Alvarez's book Short Term Trading Strategies That Work (2008). No AI. One published rule.
| Strategy | How it was found | OOS Net Profit | OOS CAR/MaxDD | OOS Win% | OOS Trades |
|---|---|---|---|---|---|
| RSI(2) long mean-reversion | Published rule (2008), no parameter search here | $22,568.94 | 0.0767 | 60.34% | 116 |
| Optimizer's #1 pick (of 2,304) | 2,304-variant sweep, best IS score | $23,365.85 | 0.0980 | 70.62% | 160 |
| Best AI rule (Rule 4) | Higher-profit of the 2 LLM rules that passed the screen | $31,954.46 | 0.1185 | 69.2% | n/a |
| Buy-and-hold SPY | No search, one long position | $87,835.34 | n/a | n/a | 1 |
Read that first pair of rows twice. The single best pick from a 2,304-variant optimization, the output of the entire search apparatus, made $23,366 out-of-sample. The textbook rule anyone can copy out of that 2008 book made $22,569 over the same window. The optimizer's pick finished roughly $800 ahead of the public, unsearched alternative.
That is the cleanest comparison in the study: the optimizer's best in-sample pick earned $796.91 more out-of-sample than a rule anyone can copy out of a book, before costs. That gap on its own cannot show the search found anything. The stronger AI rule, Rule 4, did better than both on net profit and on CAR/MaxDD (0.1185 against 0.0980 and 0.0767), and still trailed buy-and-hold by more than $55,000.
Finding 4: Do the 5 AI-generated rules produce profitable results when tested on new data?
Experiment B was more direct. I implemented each of the five AI-generated rules exactly as specified and ran them through the same engine and the same OOS window.
| Rule | Entry Condition (Summary) | IS Net Profit | OOS Net Profit | OOS CAR/MaxDD | OOS Win% | MC Percentile |
|---|---|---|---|---|---|---|
| Rule 1: RSI Oversold Bounce | RSI(14) < 30, long, 5-bar hold | $8,876 | $5,311 | 0.0187 | 70.6% | 85.5 |
| Rule 2: Stochastic Oversold + Uptrend | Stoch%K(14) < 20 AND SMA100 rising, long, 10-bar hold | $23,317 | $6,264 | 0.0209 | 61.0% | 43.8 |
| Rule 3: Williams %R Deep Oversold | WR(14) < -85 AND prior-day return < -1%, long, 5-bar hold | $24,027 | $4,286 | 0.0112 | 46.3% | 44.2 |
| Rule 4: Multi-Indicator Momentum | RSI > 55 AND Stoch > 50 AND WR > -50 AND 5-day return > 1%, long, 10-bar hold | $26,538 | $31,954 | 0.1185 | 69.2% | 96.7 |
| Rule 5: Double Overbought Short Fade | RSI > 70 AND WR > -10 AND SMA100 falling, short, 5-bar hold | -$720 | $2,591 | 0.0740 | 75.0% | 95.5 |
All five were profitable out-of-sample. That sounds solid. Look at what those profits actually mean.
Rule 3 made money in both windows ($24,027 in-sample, $4,286 out-of-sample) and still failed the screen. Its out-of-sample CAR/MaxDD of 0.0112 was the lowest of the five, it won 46.3% of its trades, and its Monte Carlo percentile was 44.2: random entries at the same frequency matched or beat it 55.8% of the time. The AI's rationale ("Williams %R below -85 combined with a prior-day down move identifies capitulation where a snap-back is probable") was a coherent story, and the rule made money. Its entry timing still did worse than the median random-entry run.
Rule 2 made $6,264 out-of-sample at a CAR/MaxDD of 0.0209. Its Monte Carlo percentile was 43.8, meaning random-entry strategies matched or beat it 56.2% of the time. A profit that random timing matches more often than not is not evidence of an edge.
Rule 1 came through OOS with $5,311 at a CAR/MaxDD of 0.0187 and a respectable 70.6% win rate. Its Monte Carlo percentile of 85.5 beat most random-entry runs, but it does not cross the 95th percentile threshold this study uses as its bar for a real edge.
Rule 5 flipped from a small in-sample loss (-$720) to a small out-of-sample profit ($2,591, CAR/MaxDD 0.0740), and its Monte Carlo percentile of 95.5 clears the bar: random entries matched or beat it 4.5% of the time. It is a short-side rule that only trades when SMA(100) is falling (below its level 10 bars earlier), on an index that buy-and-hold shows rose over the test decade. This study does not report trade counts for the AI rules, so the size of the sample behind that 95.5 is not shown here. It passed, and it is also the rule you would have thrown out on its in-sample result.
Rule 4 cleared the bar too, by a wider margin.
Finding 5: How do I know if an AI rule has a real edge — or just got lucky?
The Monte Carlo percentile is the answer. Of 5 AI-generated rules, two crossed the 95th percentile threshold.
| Rule | Monte Carlo Percentile | Verdict |
|---|---|---|
| Rule 1: RSI Oversold Bounce | 85.5 | Suggestive; below the 95th threshold |
| Rule 2: Stochastic Oversold + Uptrend | 43.8 | Below the bar: random matched or beat it 56.2% of the time |
| Rule 3: Williams %R Deep Oversold | 44.2 | Below the bar: random matched or beat it 55.8% of the time |
| Rule 4: Multi-Indicator Momentum | 96.7 | Clears the bar: random matched or beat it 3.3% of the time |
| Rule 5: Double Overbought Short Fade | 95.5 | Clears the bar: random matched or beat it 4.5% of the time, after an in-sample loss |
Rule 4 (RSI(14) > 55 AND Stochastic %K(14) > 50 AND Williams %R(14) > -50 AND 5-day prior return > 1%) is the stronger of the two. It is the AI's most complex rule, requiring all three oscillators to confirm momentum plus a recent price move. Out-of-sample net profit: $31,954.46, at a CAR/MaxDD of 0.1185. Monte Carlo percentile: 96.7. No parameter search ran over it, and 96.7% of 1,000 random-entry runs, each drawing as many entries as Rule 4 took out-of-sample trades, finished with less net profit. That is what this study's screen calls a real edge.
Neither of the two beat buying and holding SPY. Buy-and-hold returned $87,835 in the OOS window. Rule 4 returned $31,954, a gap of more than $55,000, and Rule 5 returned $2,591.
The OOS window (2016 to 2026) was a rising decade for SPY: buy-and-hold made $87,835 on the $35,000 basis, and all five AI rules trailed it. This study did not measure time in the market or test a sideways or bear decade, so it does not show how much of that gap came from sitting out a rising market. The 85% collapse is measured on this one SPY split and this one window.
Ranking on in-sample profit would not have picked the two. Rule 4 had the highest in-sample profit ($26,538), but Rule 3, second at $24,027, failed the screen, and Rule 5, the only rule that lost money in-sample (-$720), passed it. This study classified the rules using their out-of-sample profitability and Monte Carlo percentile; it tested no other attribute of the generated output that might have distinguished them in advance.
The verdict — and the honest limits
In this study, the AI's optimizer proxy and its five generated rules produced impressive-looking backtests. Neither produced a guaranteed edge.
This is not a condemnation of AI. Rules 4 and 5 passed the Monte Carlo screen, so the AI did write rules whose out-of-sample net profit beat at least 95% of 1,000 random-entry runs. But that is 2 out of 5 rules, found only after an out-of-sample test and a Monte Carlo screen, and one of the two lost money in the window you would have used to choose it. The other three made money and still failed the screen: Rule 1 came close at 85.5, and Rules 2 and 3 sat below the random median. And none of the five beat buy-and-hold.
The optimization result is starker: across 2,068 eligible strategies, the in-sample ranking picked nobody who beat buy-and-hold. It picked profitably: the top 25 all stayed profitable against 50% of the population. Their median risk-adjusted edge still fell 85%, and 1,034 of the 2,068 made money out of sample against a population median of exactly $0.
The trap is that in this 2,304-variant SPY sweep, the top-ranked in-sample backtests looked strong precisely because ranking is what put them there (their median in-sample CAR/MaxDD was 0.80). That is the mechanism Bailey, Borwein, Lopez de Prado and Zhu set out in the Notices of the AMS: try enough variants and an impressive in-sample result becomes likely even when there is no edge to find. In this study, ranking 2,068 eligible variants by in-sample score surfaced none that beat buy-and-hold out-of-sample. The best-fitting shape in past data is what a search like this finds. Your job, the actual work, is validating whether that shape holds on data it was not chosen on.
Honest limits of this study:
The OOS window (2016-06-06 to 2026-06-12) was a rising market for SPY. The study did not compare market regimes, so it cannot say how much of buy-and-hold's lead came from that decade and how much from the strategies themselves. The 85% edge-collapse finding is measured on that one window and no other. It is the gap between in-sample fit and out-of-sample reality over 2016 to 2026, and whether it holds in a bear or a sideways decade is untested here.
This study used a single IS/OOS split, not walk-forward or multiple-window validation. A single split can be unlucky at the boundary. Results are fully frictionless, with no commissions, no slippage and next-open fills, and no cost model was run, so how much of each OOS profit would survive trading costs is not measured here. The study covers SPY daily data only and was not run on other instruments, time frames, or asset classes. Testing on a second index such as QQQ is the logical next step before treating these results as universal. The Monte Carlo test used a fixed-frequency random-entry benchmark, which is one robustness check, not an exhaustive one.
The AI that wrote the five rules was trained on text published during the 2016 to 2026 test window. The code never touched that window before the test, but the rule writer's training data did, so Experiment B measures five rules against that decade, not an AI's ability to forecast it.
Rules 2 to 5 were first coded to fill at the open of the same bar whose close triggered them, which reads a close before it has printed. Every Experiment B figure here comes from the corrected code, which reads the close and fills at the next open. The corrected run, which also stops counting any trade that straddles the in-sample/out-of-sample cut, turned Rule 3 profitable out-of-sample and lifted Rule 5 over the 95th percentile, which is why two rules pass the screen and not one.
What does this mean for you?
- Never trust an AI strategy on in-sample data alone. The in-sample result is the application. The out-of-sample result is the job. Lock away a test window before you optimize or evaluate anything. This study held back 30% of its data, untouched.
- OOS profitability is not enough, so run the Monte Carlo. All five AI rules in this study were profitable out-of-sample. Rule 2 made $6,264, and random entries at its trade frequency matched or beat it 56.2% of the time. Without that check, you would not know.
- Beat the benchmark, not just breakeven. The hurdle this study set for an active strategy was beating buy-and-hold in the OOS window. If the index made $87,835 and your strategy made $5,311 at a CAR/MaxDD of 0.0187, the strategy earned a small fraction of what buy-and-hold earned, for real risk and effort. Set the hurdle before you start.
- Treat AI-generated rules as hypotheses, not answers. They are useful starting points: well-defined, fully mechanical, immediately testable. But the validation stack (IS/OOS split, Monte Carlo, benchmark comparison) is the work that actually tells you whether the hypothesis holds on data it was not built on, beats random entries, and beats the benchmark.
- Do not judge a rule by its number of conditions. The four-condition Rule 4 and the three-condition Rule 5 passed the screen, and the one-condition Rule 1 and the two-condition Rules 2 and 3 did not. Five rules cannot tell you whether complexity helps or hurts, and this study does not report their trade counts at all. Ask for that number before trusting any result, since this one can't give it to you.
The validation stack in point 4 is not something you have to assemble yourself. AlgoChef takes a backtest you have already run and applies it — the in-sample/out-of-sample split, the Monte Carlo, and the benchmark comparison — without re-running the test.







