TL;DR — the answer box
- Prompt in four narrow passes, not one broad one. Experience and criteria, then the market edge, then exits, then filters.
- Start from a market edge you can name. An unconstrained model invents behaviour that is not there, and sounds identical when it does.
- Ask for exits before you ask for profit. Drawdown is the thing you have to ask for. The workflow's third prompt does.
- The output is a candidate, never a conclusion. In the companion optimizer sweep, the median risk-adjusted score of the top 25 in-sample picks fell 85% out-of-sample.
- The validation step is not optional and it is not part of this page. It is the whole of the companion study.
What this guide is, and what it is not
This is the how. It is the prompting workflow, the order to run it in, and the four screenshots of the actual prompts.
It is deliberately not the evidence that AI-generated strategies work, because I tested that separately and the answer does not fit in a how-to. If what you want is the verdict — how many survive, by how much they degrade, what beats them — read Can AI Build a Profitable Trading Strategy? instead. That study swept 2,304 variants and implemented 5 AI-written rule-sets.
I am separating them on purpose. A guide that teaches a workflow and also grades it tends to grade it generously.
Step 1 — Start from a market edge, not from a prompt
A market edge is a specific, repeatable behaviour you have a reason to believe in. US index ETFs tending to bounce after several consecutive lower closes is a hypothesis stated that way. "Find me something profitable" is not.
This matters more than any prompt wording. An unconstrained search is data dredging with a chat window in front of it: test enough ideas against the same history and something comes back looking good. An AI asked to search everything will produce a confident, well-structured strategy for a behaviour that does not exist — and it reads exactly like one for a behaviour that does. Naming the edge first is the workflow used here: it turns the model's job into filling in the details of an idea you can already defend. No prompt-versus-prompt test was run to measure the difference.
Step 2 — Set the experience level and the criteria
This step's purpose is to change the density of the answer: naming the audience, then handing over the criteria — instrument, direction, holding period, what counts as acceptable — is designed to stop the model hedging across every possibility. Neither effect is something this page measured; it is the workflow's design, not a tested result.
Step 3 — Name the edge in the prompt itself
Here the edge from Step 1 goes in as a constraint: a mean-reversion, long-only approach on the S&P 500. Everything the model returns from this point sits inside that box. The box does not make the idea right: in the companion optimizer sweep, 1,034 of the 2,068 eligible variants (50%) were still profitable out of sample. Prompt wording was never tested against that number.
Step 4 — Ask for exits, and ask for several
This is the prompt most people skip. Neither this page nor the companion study compared entry prompts with and without an explicit drawdown request, but in this workflow's own experience, asking for entries alone got entries; asking for drawdown-reducing exits is what got exits aimed at that goal. Asking for several exits also gives you something to compare, which matters later — a strategy whose result changes wildly depending on which exit you pick was never robust to begin with.
Step 5 — Ask for filters that remove trades
A filter should take trades away. Framing the request that way produced a single meaningful condition in this example, an Internal Bar Strength threshold, rather than a stack of indicators that each remove a handful of trades and collectively fit the past. Sample size sets the floor: the companion study would not even score a variant on fewer than 30 in-sample trades, and 236 of its 2,304 variants failed that bar before any result was read.
What the output actually looks like
Running those four prompts on a defined edge, in this worked example, produced something concrete enough to test:
- Market: S&P 500
- Entry: 3 lower closes within 4 bars
- Exit 1: RSI(2) above 65
- Exit 2: after 4 bars
- Filter: Internal Bar Strength below 0.2
Tested over 18+ years of daily data, that produces this:
Read that chart carefully, because it is the exact shape that fools people. It rises. It has no benchmark line on it. It has no random-entry control on it. Nothing about it tells you whether this strategy beat simply owning the index over the same 18 years, and nothing about it tells you whether a search of similar rules would have produced an equally pretty curve by chance.
I am leaving the chart on this page because it is honest about what the workflow produces. What it is not is evidence that the workflow produces something that works.
The step this page does not cover, and you cannot skip
Everything above gets you a candidate. Whether any AI-produced candidate is real is a separate question, and this page does not answer it by assertion: the companion study below tested 2,068 optimizer-swept variants and five separately AI-generated rule-sets. This exact candidate was not itself run through that pipeline; here is what happened to the ones that were:
- 0 of 2,068 eligible optimized variants beat buy-and-hold out-of-sample. Not the best one. None of them.
- The median risk-adjusted score of the top 25 in-sample picks fell 85% when tested on data they had never seen, from 0.80 to 0.12.
- 1,034 of the 2,068 eligible optimized variants (50%) were still profitable out-of-sample, and the median net profit across all 2,068 was $0.00.
- Of 5 AI-written rule-sets implemented verbatim, 2 passed a Monte Carlo screen: each one's out-of-sample profit beat at least 95% of random-entry runs. 0 of the 5 beat buy-and-hold, and the AI that wrote them was trained on text published during the out-of-sample decade.
- Buy-and-hold on the same window returned $87,835 on a $35,000 account.
Those numbers are not measured on this page and they are not mine to summarise loosely — every one comes from the companion study, which is where the method, the controls and the full tables live.
The practical consequence for this workflow: treat the AI's output the way you would treat a strategy a stranger emailed you. Split your data before you start, keep a slice the strategy has never touched, and compare the result to owning the index rather than to zero. Scoring a rule on data it was not built from is the standard way to test one.
Where AI actually goes wrong
Hallucinated behaviour. The model will describe an indicator doing something the real calculation does not do, in fluent and specific language. Ask it to restate the rule in plain English, then re-derive the indicator yourself before trusting any curve.
Silent code errors. If you ask for code, expect logic bugs that do not throw — a lookahead in the fill, an off-by-one in the bar index. Both produce a backtest that runs cleanly and reports a number that never happened.
Neither of those two failures was measured, on this page or in the companion study. They are a checklist to run every time, not a finding.
Overfitting, which is not an AI problem. This is worth being precise about, because it is where this page used to overlap with the study. Curve-fitting is what any optimizer does when nothing stops it, human or machine. The canonical paper on it makes the point without mentioning AI at all: search enough variants and an impressive-looking backtest becomes likely even with no edge present. This page's own companion study shows the pattern: 0 of 2,068 eligible optimizer-swept variants beat buy-and-hold out-of-sample, and the top 25 in-sample picks lost a median 85% of their score once tested on data they had not seen. This study separately tested 2,304 optimizer-swept variants and five AI-generated rules, and an AI can write either kind up in prose confident enough to disguise what it is. The defence is the same as it has always been: test on data the rule never saw, and compare against a control that has no edge in it. The companion study measures how far the top in-sample picks fell once the data was new.
If you want a defined market edge to prompt from rather than inventing one, and a testing framework to check what comes back, that is what the Algo Trading Masterclass is built around. Every study behind this workflow — including the one that grades it — is published in full at StatOasis.com/Overfit
Methodology & risk note: This article runs no backtest of its own. The worked example is a single S&P 500 strategy produced by the prompts shown and tested on 18+ years of daily data in 2025; it carries no benchmark and no control, and is presented as an illustration of output shape rather than as a result. Every quantitative claim about whether AI-generated strategies survive testing is inherited from the companion study on SPY daily OHLCV, 1993-02-02 to 2026-06-12: 2,304 variants swept, 2,068 eligible at 30+ in-sample trades, in-sample to 2016-06-03 and out-of-sample from 2016-06-06, $35,000 starting capital, flat-only with no compounding, frictionless, next-open fills, measured against buy-and-hold and a seeded random-entry control. All results are historical and for educational purposes only. Past performance does not guarantee future results. Not investment advice.






