ATMResearch
Join Overfit - free
Overfit cover card on dark navy, kicker 'AI strategy validation': the headline 'AI wrote 5 strategies. 2 survived.' over the line 'Out-of-sample on SPY: 5 of 5 made money, 2 of 5 passed the full screen.', with a corner badge reading SPY, 1993 to 2026.
  1. Overfit/
  2. Research/
  3. Can AI Build a Profitable Trading Strategy? I Backtested 5 LLM-Generated Rules on SPY to Find Out

August 22, 2026

Can AI Build a Profitable Trading Strategy? I Backtested 5 LLM-Generated Rules on SPY to Find Out

Share

11 min read

Written by Ali Casey, founder of StatOasis and AlgoChef, creator of the Algo Trading Masterclass (ATM), with over 10 years of experience building systematic trading tools - building algorithmic strategies, testing ideas with data, and teaching traders how to build structured, portfolio-based trading workflows.

Published August 22, 2026 · Updated September 17, 2026 · Method

← Back to Research
Table of contents▾
  • TL;DR — the answer box
  • How I tested this
  • Finding 1: Does the optimizer's single best in-sample pick survive out-of-sample?
  • Finding 2: What happens to the top 25 "AI-selected" strategies when tested on new data?
  • Finding 3: Across all 2,068 eligible strategies, does any optimized variant beat buy-and-hold out-of-sample?
  • The benchmark nobody optimized: a published public rule
  • Finding 4: Do the 5 AI-generated rules produce profitable results when tested on new data?
  • Finding 5: How do I know if an AI rule has a real edge — or just got lucky?
  • The verdict — and the honest limits
  • What does this mean for you?
  • Methodology
  • FAQs

The short version

I ran two experiments on SPY: a 2,304-variant optimizer sweep, standing in for what an AI does when it hunts for the best strategy, and 5 AI-generated rules implemented exactly as written. The median risk-adjusted score of the optimizer's top 25 picks fell 85% the moment they touched data they had never seen, and 2 of the 5 AI rules cleared the Monte Carlo bar. The benchmark test: 0 of 2,068 eligible strategies beat simply buying and holding SPY out-of-sample, and none of the 5 AI rules did either.

Methodology & risk note: Backtested event study on SPY daily OHLCV data, 1993-02-02 to 2026-06-12, totalling 8,398 bars. Layer A: 2,304 strategy variants (RSI, Stochastic, Williams %R sweep), of which 2,068 were eligible (at least 30 IS trades). Layer B: 5 LLM-generated rule-sets implemented verbatim. In-sample: 1993-02-02 to 2016-06-03 (5,878 bars). Out-of-sample: 2016-06-06 to 2026-06-12 (2,520 bars). $35,000 starting capital. Flat-only sequential backtest (no compounding). Frictionless (no commissions or slippage). Signals read on the close, entries filled at the next open. All results are historical and for educational purposes only. Past performance does not guarantee future results. Not investment advice.

TL;DR — the answer box

  • I ran two experiments on SPY. First: a 2,304-variant optimizer sweep to simulate what AI does when it searches for "the best strategy." Second: 5 rules generated verbatim from an AI prompt, tested on the locked out-of-sample window.
  • The optimizer found 25 top in-sample picks with a median risk-adjusted score (CAR/MaxDD) of 0.80. Out-of-sample, that score collapsed 85% to 0.12. All 25 stayed profitable — but 0 of 2,068 eligible strategies beat simply buying and holding SPY, which returned +$87,835 out-of-sample.
  • Of 5 AI-generated rules, all 5 made money out-of-sample. 2 cleared the 95th Monte Carlo percentile: Rule 4 (+$31,954, CAR/MaxDD 0.1185, percentile 96.7) and Rule 5 (+$2,591, CAR/MaxDD 0.0740, percentile 95.5). Rule 1 fell short at 85.5, and random entries matched or beat Rules 2 and 3 more than half the time (percentiles 43.8 and 44.2).
  • Even the stronger of the two, Rule 4, lost to buy-and-hold by more than $55,000.
  • The verdict on these five rules: impressive backtests, not guaranteed edges. An out-of-sample test and a Monte Carlo screen are what separated the five AI rules here: the first measures whether a result holds up on data the rule was not built on, the second whether its entry timing beats random entries.

How I tested this

If you have ever typed "build me a profitable trading strategy for SPY" into an AI and stared at the output wondering whether it actually works — this study is for you.

I ran two separate experiments, each designed to answer a different version of the same question.

Experiment A — the optimizer sweep:

I ran a systematic sweep across 2,304 SPY daily strategy variants. Think of this as one brute-force optimizer-and-selection workflow: try a massive number of combinations, pick the best-looking one in the backtest window, and hand it to you as the answer. I used three momentum oscillators — RSI, Stochastic %K, and Williams %R — swept across multiple lookback lengths, entry conditions, trade directions (long and short), entry styles (mean-reversion and breakout), and time exits.

The data: SPY daily price history from 1993-02-02 to 2026-06-12, totalling 8,398 trading bars. I split this 70/30. In-sample (IS), used for building and selecting strategies: 1993-02-02 to 2016-06-03, covering 5,878 bars. Out-of-sample (OOS), the blind test window the strategies have never seen: 2016-06-06 to 2026-06-12, covering 2,520 bars. The OOS window was locked from the start and never touched during the sweep. Holding a slice back and scoring on it is the standard way to judge anything that was tuned on data.

All results are frictionless: no commissions, no slippage, next-open fills, $35,000 starting capital, flat-only (no compounding). Of 2,304 total variants, 2,068 (89.8%) qualified with at least 30 in-sample trades. The 236 variants below that eligibility floor are excluded from the eligible population. The eligible pool splits across the three indicator families as 532 RSI variants (25.7%), 768 Stochastic (37.1%), and 768 Williams %R (37.1%), and the three families landed in the same place out of sample: 50.0% of each family stayed profitable, every family's median return-to-drawdown was slightly negative, and none of the 2,068 beat buy-and-hold. That pattern (half staying profitable, a slightly negative median return-to-drawdown, none beating buy-and-hold) held separately within each of the three indicator families, not just in the pooled 2,068.

Experiment B — 5 AI-generated rules, verbatim:

I gave an AI a single prompt and asked it to generate 5 fully mechanical daily trading strategies for SPY using a fixed indicator set. I then implemented each rule exactly as written — no tweaking, no parameter adjustment — and ran it through the same engine and the same IS/OOS split.

Here is the exact prompt I used:

"You are a systematic trading strategy developer. Build me 5 profitable daily trading strategies for SPY (the S&P 500 ETF). Each strategy must use ONLY these indicators: RSI(14), Stochastic %K(14, smooth=3), Williams %R(14), SMA(100), ATR(20), 1-day prior close-to-close return, or 5-day prior return. For each strategy provide: (1) the exact entry condition with the indicator name and the precise numeric threshold, (2) the entry style — mean-reversion or momentum/breakout, (3) the direction — long or short, (4) a time-based exit in trading days — choose from 5, 10, or 15 bars. Every strategy must be fully mechanical with zero discretion so it can be coded and backtested immediately."

The validation layer — Monte Carlo:

For the 5 AI rules, I also ran a Monte Carlo test. Monte Carlo percentile answers one specific question: if a strategy entered trades on random bars with the same frequency as this rule, what share of those random simulations would do worse? Rule 4's percentile of 96.7 means random-entry runs matched or beat it 3.3% of the time. Rule 2's 43.8 means they matched or beat it 56.2% of the time, which is what picking trade entries out of a hat looks like. The bar a rule has to clear here is the 95th percentile.

Finding 1: Does the optimizer's single best in-sample pick survive out-of-sample?

The optimizer's top pick, the one this selection rule hands back as its best result, was a Williams %R(2) long mean-reversion strategy with a 10-bar time exit. In-sample, the numbers looked exceptional.

MetricIn-SampleOut-of-Sample
Net Profit$98,354.68$23,365.85
CAR/MaxDD1.46100.0980
Win Rate73.58%70.62%
Trades405160

CAR/MaxDD stands for Compound Annual Return divided by Maximum Drawdown, both in percent, so it measures return against worst-case loss. A score of 1.0 means the annual return, in percent, matched the maximum drawdown, in percent. The in-sample score of 1.46 was the highest of the 2,068 eligible variants, against a top-25 median of 0.80. The out-of-sample score of 0.10 is a fraction of that.

The win rate held (73.6% IS, 70.6% OOS). The strategy stayed profitable. But the risk-adjusted quality dropped sharply from 1.46 to 0.10 — and the strategy made $23,365.85 over the OOS window while simply buying and holding SPY made $87,835 over the exact same period. The optimizer's top pick trailed buy-and-hold by more than $64,000.

The win rate barely moved — 73.58% in-sample, 70.62% out-of-sample — while the risk-adjusted score fell from 1.4610 to 0.0980. For this pick, the number that survived was not the number that mattered.

Finding 2: What happens to the top 25 "AI-selected" strategies when tested on new data?

One bad pick might be a fluke. What if you take the top 25 — the cream of 2,068 eligible strategies, ranked by in-sample risk-adjusted score? This is the shortlist the optimizer proxy in this study generates.

MetricTop 25 — In-SampleTop 25 — Out-of-Sample
Median CAR/MaxDD0.80320.1242
Still profitable—25 of 25 (100%)
Beat buy-and-hold—0 of 25 (0%)

Median CAR/MaxDD dropped from 0.80 to 0.12 — an 85% collapse in the risk-adjusted score.

All 25 stayed profitable out-of-sample. They are not blown-up, money-losing strategies. But not one of them beat the baseline: buy-and-hold SPY returned $87,835 over the OOS window. The best of the best from a 2,304-variant sweep could not match a strategy with zero intelligence — just hold the index.

An in-sample backtest is like a job interview where the candidate designed the questions themselves. They score close to perfect every time. The out-of-sample test is the actual job. Acing the custom interview and excelling at the real work are two very different things.

The shortlist this optimizer proxy hands you: median CAR/MaxDD drops from 0.8032 to 0.1242, an 84.5% collapse. All 25 stayed profitable; none of them beat the index.

Finding 3: Across all 2,068 eligible strategies, does any optimized variant beat buy-and-hold out-of-sample?

Not one.

MetricWhole Eligible Population (OOS)
Eligible variants2,068
Profitable OOS (net > $0)1,034 (50.0%)
Beat buy-and-hold OOS0 (0.0%)
Median OOS net profit$0.00

Exactly half of the 2,068 strategies made money out-of-sample. The other half did not. The median out-of-sample net profit was exactly $0.00. And none of them, across the full eligible population, crossed the $87,835 buy-and-hold line.

The optimizer is extremely good at one thing: finding parameters that fit the past data it was shown. It is not finding anything that beat buying and holding. The shapes it finds look like edges in the training window, and the top 25 by in-sample score lost a median 85% of that score in the test window.

This does not mean systematic trading is hopeless. The buy-and-hold result itself is a systematic strategy, and it had the highest out-of-sample net profit in this comparison. What it shows is that adding oscillator conditions to a basic long-index position is not automatically adding edge. The OOS window (2016 to 2026) was a rising decade for SPY. Whether that drift is what decided the comparison was never tested here.

No in-sample score, however strong, came with a variant that beat buy-and-hold: of 2,068 eligible variants, 1,034 made money out-of-sample and not one crossed the $87,835.34 line.

The benchmark nobody optimized: a published public rule

There is one more comparison in the data, and it is the most deflating one for the optimizer. Alongside the sweep, I ran a fixed benchmark rule with no parameter search: the classic RSI(2) long mean-reversion strategy from Larry Connors and Cesar Alvarez's book Short Term Trading Strategies That Work (2008). No AI. One published rule.

StrategyHow it was foundOOS Net ProfitOOS CAR/MaxDDOOS Win%OOS Trades
RSI(2) long mean-reversionPublished rule (2008), no parameter search here$22,568.940.076760.34%116
Optimizer's #1 pick (of 2,304)2,304-variant sweep, best IS score$23,365.850.098070.62%160
Best AI rule (Rule 4)Higher-profit of the 2 LLM rules that passed the screen$31,954.460.118569.2%n/a
Buy-and-hold SPYNo search, one long position$87,835.34n/an/a1

Read that first pair of rows twice. The single best pick from a 2,304-variant optimization, the output of the entire search apparatus, made $23,366 out-of-sample. The textbook rule anyone can copy out of that 2008 book made $22,569 over the same window. The optimizer's pick finished roughly $800 ahead of the public, unsearched alternative.

That is the cleanest comparison in the study: the optimizer's best in-sample pick earned $796.91 more out-of-sample than a rule anyone can copy out of a book, before costs. That gap on its own cannot show the search found anything. The stronger AI rule, Rule 4, did better than both on net profit and on CAR/MaxDD (0.1185 against 0.0980 and 0.0767), and still trailed buy-and-hold by more than $55,000.

Finding 4: Do the 5 AI-generated rules produce profitable results when tested on new data?

Experiment B was more direct. I implemented each of the five AI-generated rules exactly as specified and ran them through the same engine and the same OOS window.

RuleEntry Condition (Summary)IS Net ProfitOOS Net ProfitOOS CAR/MaxDDOOS Win%MC Percentile
Rule 1: RSI Oversold BounceRSI(14) < 30, long, 5-bar hold$8,876$5,3110.018770.6%85.5
Rule 2: Stochastic Oversold + UptrendStoch%K(14) < 20 AND SMA100 rising, long, 10-bar hold$23,317$6,2640.020961.0%43.8
Rule 3: Williams %R Deep OversoldWR(14) < -85 AND prior-day return < -1%, long, 5-bar hold$24,027$4,2860.011246.3%44.2
Rule 4: Multi-Indicator MomentumRSI > 55 AND Stoch > 50 AND WR > -50 AND 5-day return > 1%, long, 10-bar hold$26,538$31,9540.118569.2%96.7
Rule 5: Double Overbought Short FadeRSI > 70 AND WR > -10 AND SMA100 falling, short, 5-bar hold-$720$2,5910.074075.0%95.5

All five were profitable out-of-sample. That sounds solid. Look at what those profits actually mean.

Rule 3 made money in both windows ($24,027 in-sample, $4,286 out-of-sample) and still failed the screen. Its out-of-sample CAR/MaxDD of 0.0112 was the lowest of the five, it won 46.3% of its trades, and its Monte Carlo percentile was 44.2: random entries at the same frequency matched or beat it 55.8% of the time. The AI's rationale ("Williams %R below -85 combined with a prior-day down move identifies capitulation where a snap-back is probable") was a coherent story, and the rule made money. Its entry timing still did worse than the median random-entry run.

Rule 2 made $6,264 out-of-sample at a CAR/MaxDD of 0.0209. Its Monte Carlo percentile was 43.8, meaning random-entry strategies matched or beat it 56.2% of the time. A profit that random timing matches more often than not is not evidence of an edge.

Rule 1 came through OOS with $5,311 at a CAR/MaxDD of 0.0187 and a respectable 70.6% win rate. Its Monte Carlo percentile of 85.5 beat most random-entry runs, but it does not cross the 95th percentile threshold this study uses as its bar for a real edge.

Rule 5 flipped from a small in-sample loss (-$720) to a small out-of-sample profit ($2,591, CAR/MaxDD 0.0740), and its Monte Carlo percentile of 95.5 clears the bar: random entries matched or beat it 4.5% of the time. It is a short-side rule that only trades when SMA(100) is falling (below its level 10 bars earlier), on an index that buy-and-hold shows rose over the test decade. This study does not report trade counts for the AI rules, so the size of the sample behind that 95.5 is not shown here. It passed, and it is also the rule you would have thrown out on its in-sample result.

Rule 4 cleared the bar too, by a wider margin.

Five rules from one prompt land anywhere from $2,591.31 (CAR/MaxDD 0.0740) to $31,954.46 (CAR/MaxDD 0.1185) out-of-sample, and even the best of them falls short of the $87,835.34 buy-and-hold made.

Finding 5: How do I know if an AI rule has a real edge — or just got lucky?

The Monte Carlo percentile is the answer. Of 5 AI-generated rules, two crossed the 95th percentile threshold.

RuleMonte Carlo PercentileVerdict
Rule 1: RSI Oversold Bounce85.5Suggestive; below the 95th threshold
Rule 2: Stochastic Oversold + Uptrend43.8Below the bar: random matched or beat it 56.2% of the time
Rule 3: Williams %R Deep Oversold44.2Below the bar: random matched or beat it 55.8% of the time
Rule 4: Multi-Indicator Momentum96.7Clears the bar: random matched or beat it 3.3% of the time
Rule 5: Double Overbought Short Fade95.5Clears the bar: random matched or beat it 4.5% of the time, after an in-sample loss

Rule 4 (RSI(14) > 55 AND Stochastic %K(14) > 50 AND Williams %R(14) > -50 AND 5-day prior return > 1%) is the stronger of the two. It is the AI's most complex rule, requiring all three oscillators to confirm momentum plus a recent price move. Out-of-sample net profit: $31,954.46, at a CAR/MaxDD of 0.1185. Monte Carlo percentile: 96.7. No parameter search ran over it, and 96.7% of 1,000 random-entry runs, each drawing as many entries as Rule 4 took out-of-sample trades, finished with less net profit. That is what this study's screen calls a real edge.

Neither of the two beat buying and holding SPY. Buy-and-hold returned $87,835 in the OOS window. Rule 4 returned $31,954, a gap of more than $55,000, and Rule 5 returned $2,591.

The OOS window (2016 to 2026) was a rising decade for SPY: buy-and-hold made $87,835 on the $35,000 basis, and all five AI rules trailed it. This study did not measure time in the market or test a sideways or bear decade, so it does not show how much of that gap came from sitting out a rising market. The 85% collapse is measured on this one SPY split and this one window.

Ranking on in-sample profit would not have picked the two. Rule 4 had the highest in-sample profit ($26,538), but Rule 3, second at $24,027, failed the screen, and Rule 5, the only rule that lost money in-sample (-$720), passed it. This study classified the rules using their out-of-sample profitability and Monte Carlo percentile; it tested no other attribute of the generated output that might have distinguished them in advance.

Rules 4 and 5 clear the 95th-percentile bar, at 96.7 and 95.5. Rule 1 reaches 85.5, and Rules 2 and 3 sit at 43.8 and 44.2, where random entries matched or beat the rule more than half the time.

The verdict — and the honest limits

In this study, the AI's optimizer proxy and its five generated rules produced impressive-looking backtests. Neither produced a guaranteed edge.

This is not a condemnation of AI. Rules 4 and 5 passed the Monte Carlo screen, so the AI did write rules whose out-of-sample net profit beat at least 95% of 1,000 random-entry runs. But that is 2 out of 5 rules, found only after an out-of-sample test and a Monte Carlo screen, and one of the two lost money in the window you would have used to choose it. The other three made money and still failed the screen: Rule 1 came close at 85.5, and Rules 2 and 3 sat below the random median. And none of the five beat buy-and-hold.

The optimization result is starker: across 2,068 eligible strategies, the in-sample ranking picked nobody who beat buy-and-hold. It picked profitably: the top 25 all stayed profitable against 50% of the population. Their median risk-adjusted edge still fell 85%, and 1,034 of the 2,068 made money out of sample against a population median of exactly $0.

The trap is that in this 2,304-variant SPY sweep, the top-ranked in-sample backtests looked strong precisely because ranking is what put them there (their median in-sample CAR/MaxDD was 0.80). That is the mechanism Bailey, Borwein, Lopez de Prado and Zhu set out in the Notices of the AMS: try enough variants and an impressive in-sample result becomes likely even when there is no edge to find. In this study, ranking 2,068 eligible variants by in-sample score surfaced none that beat buy-and-hold out-of-sample. The best-fitting shape in past data is what a search like this finds. Your job, the actual work, is validating whether that shape holds on data it was not chosen on.

Honest limits of this study:

The OOS window (2016-06-06 to 2026-06-12) was a rising market for SPY. The study did not compare market regimes, so it cannot say how much of buy-and-hold's lead came from that decade and how much from the strategies themselves. The 85% edge-collapse finding is measured on that one window and no other. It is the gap between in-sample fit and out-of-sample reality over 2016 to 2026, and whether it holds in a bear or a sideways decade is untested here.

This study used a single IS/OOS split, not walk-forward or multiple-window validation. A single split can be unlucky at the boundary. Results are fully frictionless, with no commissions, no slippage and next-open fills, and no cost model was run, so how much of each OOS profit would survive trading costs is not measured here. The study covers SPY daily data only and was not run on other instruments, time frames, or asset classes. Testing on a second index such as QQQ is the logical next step before treating these results as universal. The Monte Carlo test used a fixed-frequency random-entry benchmark, which is one robustness check, not an exhaustive one.

The AI that wrote the five rules was trained on text published during the 2016 to 2026 test window. The code never touched that window before the test, but the rule writer's training data did, so Experiment B measures five rules against that decade, not an AI's ability to forecast it.

Rules 2 to 5 were first coded to fill at the open of the same bar whose close triggered them, which reads a close before it has printed. Every Experiment B figure here comes from the corrected code, which reads the close and fills at the next open. The corrected run, which also stops counting any trade that straddles the in-sample/out-of-sample cut, turned Rule 3 profitable out-of-sample and lifted Rule 5 over the 95th percentile, which is why two rules pass the screen and not one.

What does this mean for you?

  1. Never trust an AI strategy on in-sample data alone. The in-sample result is the application. The out-of-sample result is the job. Lock away a test window before you optimize or evaluate anything. This study held back 30% of its data, untouched.
  2. OOS profitability is not enough, so run the Monte Carlo. All five AI rules in this study were profitable out-of-sample. Rule 2 made $6,264, and random entries at its trade frequency matched or beat it 56.2% of the time. Without that check, you would not know.
  3. Beat the benchmark, not just breakeven. The hurdle this study set for an active strategy was beating buy-and-hold in the OOS window. If the index made $87,835 and your strategy made $5,311 at a CAR/MaxDD of 0.0187, the strategy earned a small fraction of what buy-and-hold earned, for real risk and effort. Set the hurdle before you start.
  4. Treat AI-generated rules as hypotheses, not answers. They are useful starting points: well-defined, fully mechanical, immediately testable. But the validation stack (IS/OOS split, Monte Carlo, benchmark comparison) is the work that actually tells you whether the hypothesis holds on data it was not built on, beats random entries, and beats the benchmark.
  5. Do not judge a rule by its number of conditions. The four-condition Rule 4 and the three-condition Rule 5 passed the screen, and the one-condition Rule 1 and the two-condition Rules 2 and 3 did not. Five rules cannot tell you whether complexity helps or hurts, and this study does not report their trade counts at all. Ask for that number before trusting any result, since this one can't give it to you.

The validation stack in point 4 is not something you have to assemble yourself. AlgoChef takes a backtest you have already run and applies it — the in-sample/out-of-sample split, the Monte Carlo, and the benchmark comparison — without re-running the test.

Methodology

Data source
SPY daily price history from the StatOasis research dataset.
Date range
1993-02-02 to 2026-06-12 (8,398 trading bars), split 70/30 — in-sample 1993-02-02 to 2016-06-03 (5,878 bars), out-of-sample 2016-06-06 to 2026-06-12 (2,520 bars). The out-of-sample window was locked from the start and never touched during the sweep.
Entry / exit rules
Experiment A sweeps 2,304 SPY daily variants across three momentum oscillators (RSI, Stochastic %K, Williams %R) over multiple lookback lengths, entry conditions, directions (long/short), entry styles (mean-reversion/breakout) and time exits, as this study's proxy for the brute-force search an optimizer-driven AI strategy builder runs. Experiment B implements 5 AI-generated mechanical rules exactly as written, with no tweaking or parameter adjustment, through the same engine and the same in-sample/out-of-sample split. In both experiments a signal is read on the close of bar t and filled at the open of bar t+1.
Sizing
$35,000 starting capital, flat-only, no compounding. Frictionless: no commissions or slippage.
Overlap mode
Flat-only, one position at a time, so overlapping signals are skipped. Of 2,304 variants, 2,068 (89.8%) qualified with at least 30 in-sample trades, and the 236 below that eligibility floor are excluded from the eligible population.
Look-ahead
Signals are read on the close of bar t and filled at the open of bar t+1 in both experiments, and a trade that enters before the in-sample/out-of-sample cut and exits after it is counted in neither window. The out-of-sample window was locked from the start and never touched during the sweep, and the random control draws its entries from that window only. One limit the code cannot remove: the AI that wrote the five Experiment B rules was trained on text published during the out-of-sample window.
Minimum sample
30 in-sample trades, not the engine default of 50: the split leaves each half smaller, so the floor was lowered deliberately. 2,068 of the 2,304 variants qualify, and the 236 below it are excluded. No other floor was tested.
Buy-and-hold benchmark
Buy and hold SPY across the out-of-sample window only (2016-06-06 to 2026-06-12): $87,835.34 net on the $35,000 basis. That is the hurdle every surviving variant is measured against, and the article's headline is how few clear it.
Random control
Seeded random control on the same out-of-sample window: entries placed at random, frequency-matched to the median out-of-sample trade count of the eligible variants, same 5-bar hold, averaged over 10 seeds from a fixed base seed. $8,201.11 net, CAR/MaxDD 0.0494, 93.2 trades. It is the baseline this study most needs: a search across 2,304 variants produces a best-looking one whether or not an edge exists, so beating buy-and-hold or beating RSI(2) cannot on its own show the search found anything.
Parameter scopeParameters swept

The study searched the parameter space and reports the spread, not one tuned setting.

Experiment A sweeps all 2,304 variants across three oscillator families, lengths, thresholds, directions, entry styles and time exits, as this study's proxy for the brute-force search an optimizer-driven AI strategy builder runs. Experiment B then implements 5 AI-generated rules exactly as written, with no tuning at all.

Run to v1 of the StatOasis research standard - the rules every study here has to meet before it is published. The version is the study's own: a standard that gained a rule later never reaches back and claims this one met it.

Historical backtest results are not a guarantee of future returns. This content is for educational purposes only and is not investment advice. Hypothetical performance disclosure (CFTC Rule 4.41).

Frequently asked questions

Can ChatGPT build a profitable trading strategy?⌄

Sometimes. In this study the five rules could not be told apart until they were tested: all 5 AI-generated SPY rules made money out-of-sample, and 2 of the 5 cleared the 95th Monte Carlo percentile: Rule 4 (Multi-Indicator Momentum) at 96.7 with +$31,954 OOS and a CAR/MaxDD of 0.1185, and Rule 5 (Double Overbought Short Fade) at 95.5 with +$2,591 OOS and a CAR/MaxDD of 0.0740. The other three sat at 85.5, 44.2 and 43.8, and 0 of the 5 beat simply buying and holding SPY (+$87,835 OOS). The AI was trained on text published during that test window, so this measures its rules, not its foresight. The AI generates a strategy definition. Validation tells you whether it works.

How do you test if an AI-generated trading strategy actually works?⌄

Three steps, in order. First, split your data before you start: use 70% for building (in-sample) and hold 30% back as a blind test (out-of-sample), and never touch the OOS window until the strategy is fully specified and locked. Second, run the strategy on the OOS data and measure net profit, win rate, and CAR/MaxDD (risk-adjusted return). Third, run a Monte Carlo test: simulate thousands of random-entry strategies at the same trade frequency and check what percentile your rule lands in. The 95th percentile is the minimum bar for "likely a real edge." Below the 50th percentile, more than half of the random-entry runs matched or beat the rule.

What happens to AI trading strategies when tested out of sample?⌄

The risk-adjusted quality collapsed here. The top 25 in-sample strategies had a median CAR/MaxDD of 0.80. The same strategies out-of-sample had a median CAR/MaxDD of 0.12 — an 85% degradation. They stayed profitable (25 of 25 made money OOS), but 0 of 25 beat the buy-and-hold baseline (+$87,835). Across the full 2,068 eligible variants, exactly 50% were profitable OOS and the median net profit was exactly $0.

Do AI trading bots actually make money in live trading?⌄

This study covers backtests only, frictionless, on historical SPY data, with no live execution. What it shows is that even under ideal backtest conditions with no commissions or slippage, the median of the 2,068 eligible optimized strategies made $0 net profit out-of-sample and none of them beat buy-and-hold. Real trading adds friction (commissions, slippage, fill latency) and behavioral pressure (drawdown anxiety, execution hesitation) that backtests do not capture. Any live-trading claim from an AI strategy tool should be viewed through those caveats.

How do you avoid overfitting an AI-generated strategy?⌄

Three practices. (1) Lock your test window before you build anything, so the data you judge a rule on played no part in choosing it. (2) Demand a large trade sample, and do not judge a rule by how many conditions it has: in this study the four-condition Rule 4 and the three-condition Rule 5 passed the Monte Carlo screen, while the one-condition Rule 1 and the two-condition Rules 2 and 3 did not. (3) Run the Monte Carlo screen, which asks whether the rule's entry timing beats random entries placed at the same frequency.

What is the difference between in-sample and out-of-sample backtesting?⌄

In-sample is the historical data you use to build and optimize a strategy. Out-of-sample is data the strategy has never seen — a blind test of whether the rules that worked historically hold on genuinely new data. A strategy can score beautifully in-sample simply by fitting to noise in past prices. The OOS test is what separates a rule that survives new data from one that does not, though only a random-entry null can say whether the search found anything at all. In this study, in-sample ran from 1993-02-02 to 2016-06-03 (5,878 bars); out-of-sample ran from 2016-06-06 to 2026-06-12 (2,520 bars).

Can AI find trading edges that human traders miss?⌄

Possibly, as a source of hypotheses worth testing. Rule 4 (Multi-Indicator Momentum Confluence: RSI > 55, Stochastic > 50, Williams %R > -50, 5-day return > 1%) was the AI's most complex rule and one of the two that cleared the 95th Monte Carlo percentile in this study, at 96.7 with +$31,954 OOS and a CAR/MaxDD of 0.1185. The other was Rule 5, a short fade that lost money in-sample. But the AI was trained on text published during the 2016 to 2026 test window, so a rule it writes is not a blind forecast of that decade, and without the validation stack (OOS test, Monte Carlo, benchmark comparison) there was no way to tell the two rules that passed from the three that did not. The AI identifies candidates. The validation screens them.

Which AI tool is best for building trading strategies in 2026?⌄

This study does not benchmark AI tools against each other. What it shows is that nothing in the AI's output separated its rules: all five came from the same prompt to the same AI, and the difference between the two that passed the Monte Carlo screen and the three that did not showed up only in the backtest engine and the validation step. A tool that runs a rigorous IS/OOS split and Monte Carlo test tells you something a confident-sounding strategy description cannot.

Why do strategies that look great in a backtest fail in live trading?⌄

Two main reasons. First, curve-fitting: a strategy chosen by searching across many parameter combinations is chosen for how well it fit its training data, and that fit is no promise about new data. In this study, the median risk-adjusted score of the top-25 IS strategies fell 85% out-of-sample, before touching any live market. The study never separated curve-fitting from the change of decade: the in-sample window runs 1993 to 2016 and the out-of-sample window 2016 to 2026, so both sit inside that 85%. Second, real trading adds what a frictionless backtest leaves out: commissions, slippage, and the discipline a drawdown demands.

How many trades does a backtest need before you can trust it?⌄

This study required at least 30 in-sample trades and tested no other floor: any variant with fewer than 30 was excluded from the eligible population, 236 of 2,304. It tested no out-of-sample trade-count threshold either, and it reports no trade counts for the five AI rules, so it sets no number for how many trades are enough. More trades give a statistical test more resolution, which is why a percentile reported without its trade count is hard to judge.

Is a high win rate proof that a strategy works?⌄

No. For the optimizer's top pick, win rate stayed the most stable of the numbers checked here and told you the least about what mattered. It won 73.6% of trades in-sample and still won 70.6% out-of-sample, yet its risk-adjusted score collapsed from 1.46 to 0.10 and it trailed buy-and-hold by more than $64,000. Its 70.6% win rate did not establish a strong risk-adjusted score or a lead over buy-and-hold for this pick, and those are where it fell short. Judge a rule by CAR/MaxDD, net profit against buy-and-hold, and the Monte Carlo percentile, not by its win rate.

Did the AI or the optimizer beat a simple published rule like RSI(2)?⌄

Barely, or not at all, for the optimizer. The classic RSI(2) long mean-reversion rule from Larry Connors and Cesar Alvarez's book Short Term Trading Strategies That Work (2008), run here with no parameter search, made $22,569 out-of-sample at a CAR/MaxDD of 0.0767. The optimizer's #1 pick from 2,304 variants made $23,366 over the same window at 0.0980, roughly $800 more. Rule 4, the stronger of the two AI rules that passed the Monte Carlo screen, made $31,954 at 0.1185 and a 96.7 percentile, and it still trailed buy-and-hold by more than $55,000. The other AI rule to pass, Rule 5, made $2,591 at 0.0740, below RSI(2) on both counts. An $800 lead over a published textbook rule cannot, on its own, show the search found anything.

Read the Tools, Software & Tech Stack hub
← Back to Research

Table of contents

  • TL;DR — the answer box
  • How I tested this
  • Finding 1: Does the optimizer's single best in-sample pick survive out-of-sample?
  • Finding 2: What happens to the top 25 "AI-selected" strategies when tested on new data?
  • Finding 3: Across all 2,068 eligible strategies, does any optimized variant beat buy-and-hold out-of-sample?
  • The benchmark nobody optimized: a published public rule
  • Finding 4: Do the 5 AI-generated rules produce profitable results when tested on new data?
  • Finding 5: How do I know if an AI rule has a real edge — or just got lucky?
  • The verdict — and the honest limits
  • What does this mean for you?
  • Methodology
  • FAQs

Overfit - the newsletter

Skip the hype. Trust the data.One practical takeaway per issue.

Running a quick security check before this can be sent.

Delivered +2 times a month - when the work is ready, not on a calendar.I'll never sell your address.Unsubscribe in one click, any time.

Related Articles

Back to all research
#122

Robustness Testing: Why Most Traders Fail, and What 36,252 Backtests Say Actually Works

Mar 28, 2025 · 18 min read

Robustness testing is the step between a good backtest and deciding whether to trade it. I measured which checks actually predict what a strategy does in held-back history, then retested the whole thing on 53 markets — where the stack still works, several individual checks turn out to have been overstated threefold, and two of them point the wrong way on some asset classes.

Read more→
#142

The Better-RSI Showdown: We Tested 4 RSI Upgrades on SPY, QQQ, IWM, and DIA

Sep 10, 2026 · 10 min read

1,856 backtests across four RSI families on SPY, QQQ, IWM and DIA: Connors RSI and Z-Score RSI modestly beat plain RSI on median risk-adjusted return, Laguerre RSI was the worst of the four despite its lag-free marketing, and the simplified three-condition Triple RSI tested here does not reproduce the popular win-rate claim.

Read more→
#141

I Backtested ICT / Smart Money Concepts — What Survives

Sep 3, 2026 · 10 min read

The four core ICT / Smart Money Concepts entries — order blocks, fair value gaps, liquidity sweeps and Optimal Trade Entry — codified into mechanical rules and run against three textbook entries and a coin flip across four markets. None showed a statistically significant 5- or 10-day forward-return edge on SPY.

Read more→

Overfit - the newsletter

Skip the hype. Trust the data.

One practical, evidence-driven takeaway per issue - strategy testing, portfolio construction, market structure, trading psychology, tactical asset allocation.

Running a quick security check before this can be sent.

Delivered +2 times a month - when the work is ready, not on a calendar.I'll never sell your address.Unsubscribe in one click, any time.

StatOasis is calm, evidence-based algorithmic-trading education, founded by Ali Casey. Ali builds systematic trading strategies and teaches the workflow behind them: research, build, test, combine, deploy. He writes the Overfit newsletter, published since 2024, and runs the Algo Trading Masterclass.

Socials

  • X ↗
  • YouTube ↗
  • Instagram ↗
  • LinkedIn ↗
  • GitHub ↗
  • Muck Rack ↗
  • LinkedIn SO ↗
  • GitHub SO ↗

Products

  • Overfit - the newsletter
  • Algo Trading Masterclass
  • StatOasis Community
  • 36 Ways to Buy the Dip
  • AlgoChef ↗

Reading & tools

  • Research
  • Methodology
  • Survive the Decade
  • Wall of Love

StatOasis

  • About Ali Casey
  • Contact
© 2026 StatOasis. Calm, evidence-based.
PrivacyTermsHypothetical resultsCalifornia