The research standard · v1.1
How a StatOasis study is run.
Published · Last updated
Studies designed, run and interpreted by Ali Casey · About
Every study on this site makes a claim you can check. Don't trust the prose - check the method: this page states the rules a study has to meet before it is published here, so you never have to take a number on trust.
It applies to every study published from August 3, 2026. Where a rule is newer than the practice, it says so - and where the standard deliberately promises nothing, it says that too.
Whole parameter spaces, not one tuned setting.
Any indicator can be made to look brilliant by reporting its best settings. So where the question is whether a rule is tuned or real, studies here sweep the space and report what the range did - because the range is what you would actually have got. Where a study holds one configuration still on purpose, rule 13 makes it say so.
Measure the effect. Then ask if it's tradable.
Different questions, different sets of rules. Collapsing them is how a real statistical effect gets sold as a trading system it was never tested as - and how a study that held its settings still for a reason gets read as one that simply did not check.
Layer 1
Measurement
Does the effect exist at all?
- Base rates over every occurrence in the history
- Overlaps kept - this is a measurement, not a system
- No capital, no entries, nothing traded
Layer 2
Strategy
Would it have survived as something you could trade?
- Swept across the parameter space, not tuned to one set
- Flat-only, uncompounded, costs stated
- Against buy-and-hold and a seeded random control
A study can be both, in layers - and that is the preferred shape. Measure the effect first; only then ask whether it survives contact with costs, overlaps and a benchmark.
A different question
Controlled comparison
Does a condition change the outcome?
- One configuration, held fixed on purpose - it is the instrument, not the subject
- The condition is what varies: regime, season, session, instrument
- Every trading rule still applies - plus rule 13: the fixed scope is declared, and what it costs
Asking whether market regime changes what a rule does means holding that rule still. Sweeping it as well would leave the difference unattributable to either. A fixed configuration here is the experiment, not a shortcut around one.
Every study, without exception.
Five rules that apply whether a study trades anything or not.
Step 1
Source named
The data source is named
Step 2
Range counted
The date range carries an observation count
Step 3
Overlap declared
Overlap handling is declared
Step 4
Capital stated
The capital basis is stated
Step 5
Rebuildable
The study is specified well enough to rebuild
The data source is named
The actual series and where it came from. “Our research engine” only counts with the instrument and bar type named alongside it.
See it: Shiller monthly + daily S&P 500 →The date range carries an observation count
“1993-02-02 to 2026-06-12 (8,398 bars)”. A range without a count hides how thin the sample is.
See it: 8,398 bars, 33.4 years →Overlap handling is declared
Flat-only, non-overlapping, or overlaps deliberately kept for a measurement layer - stated either way, never left to inference.
The capital basis is stated
Starting capital and position sizing - or an explicit note that the study measures something and trades nothing.
See it: $35,000, flat-only, no compounding →The study is specified well enough to rebuild
That is what the other rules are for - source, range, overlaps, capital, costs, fills, floors - stated exactly enough that someone with equivalent data gets the same answer. Where the output itself can be published, it is: the tables behind the headline number, not the vendor feed underneath it. A conclusion nobody could recompute is an opinion.
See it: Full drawdown table, free →
And if it claims to be tradable.
Eight more. The five marked new were adopted on August 2026 - most were real practice before, but not on every study, and a standard that only sometimes applies is not a standard.
Flat-only
One position at a time; overlapping signals are skipped. A backtest stacking unlimited concurrent positions is measuring leverage, not the rule.
No compounding
Fixed capital per position, so the result measures the rule rather than the compounding curve sitting on top of it.
Costs are stated, not hidden
Frictionless testing is allowed - most studies here are - but it is named as a limitation. A frictionless backtest flatters a strategy, and you are told so rather than finding out later.
Benchmarked against buy-and-holdNew · 2026-08-03
Every strategy reports what the same capital would have done doing nothing over the same window. A strategy that makes money and still loses to buy-and-hold has failed, and the study says so.
See it: 0 of 2,068 beat buy-and-hold →Tested against a seeded random controlNew · 2026-08-03
A coin-flip entry firing at the same average frequency as the real signal, averaged over multiple seeds - so “better than random” is measured instead of assumed.
No look-ahead, and it says soNew · 2026-08-03
Signals are detected at the close of bar t; every fill is the next bar's open. True of the engine already - now declared in every study, so you never have to take it on trust.
A minimum-sample floorNew · 2026-08-03
A variant with too few trades is reported as unreliable rather than celebrated. The default floor is 50 trades; a study using a different one names it.
The parameter scope is declaredNew · 2026-08-03
Either the study sweeps the space and reports the spread, or it holds one configuration fixed and says why - and what that costs: the finding is conditional on that setting and makes no claim to survive re-tuning. Like rule 8, this asks for a declaration rather than a method. Reporting one tuned setting as though it were the whole space is what it forbids.
And if it's a review of a product.
One more, added on August 19, 2026. A review runs no backtest, so rules 6 to 13 have nothing to bite on - and the thing that could quietly bias it is not in the data at all.
The commercial relationship is disclosed, before the verdictNew · 2026-08-19
A review says what relationship exists between StatOasis and the product it judges - an affiliate commission, a licence the vendor supplied, an appearance in the vendor's marketing - or says plainly that there is none. Every other rule here is about data; this one is not, because that is not where a review's honesty risk lives. None of those relationships makes a review dishonest. Concealing one does. “No commercial relationship” is a complete answer - silence is not - and you read it above the review rather than underneath it.
The limits
What this standard does not promise.
Stating the limits is what keeps the rest of it worth anything.
- Not every study is split in-sample / out-of-sample
- Sweeping a whole parameter space and reporting the spread answers overfitting a different way, and an IS/OOS split does not fit a measurement study at all. Where a study does split it says so - and reports the collapse.
- Not every study is Monte Carlo'd
- Resampling belongs to validating a strategy you intend to deploy, not to establishing whether an effect exists in the first place.
- Frictionless is not free
- Real commission and slippage would reduce every net figure published here, and by more for the higher-frequency variants. Rule 8 exists so that is never a surprise.
- This is a research publication, not a code repository
- The method is stated in full. The code that ran it is not published, and neither is price data licensed against redistribution. What ships is what the claim rests on: the specification, and the result tables where they can be published.
- None of this is a live-trading result
- These are historical measurements. A backtest that clears every rule above is still a backtest, and nothing here is investment advice.
What this covers.
Studies published from August 2026 onward. Articles published before that date predate the standard and are not retro-fitted or quietly covered by it - a standard backdated over work that never met it would be the first thing on this page worth disbelieving.
Earlier articles are being re-run against these rules and updated in place, at their existing URLs. Until one has been, it makes no claim to have met this standard.
This standard is versioned, and it is on v1.1. When a rule changes, the new version applies from its own date forward - a study published under an earlier version is never re-described as having met a rule that did not exist when it ran. Rule 14 is the current example: it applies from August 2026, and nothing published before it is described as having met it.
What produced the numbers.
Years of StatOasis research ran on TradeStation, MultiCharts and StrategyQuant X. They are serious professional environments - TradeStation and MultiCharts sit close enough to execution that a tested idea and a traded one share the same logic, and StrategyQuant X puts sophisticated strategy research within reach of traders who don't write code. The articles that came out of that work predate this standard, and each one says so.
New research runs in Python. Not because those platforms are worse at it - they aren't - but because of who can check the work afterwards. Python is free, mature, and standard equipment in statistics and scientific computing. Rebuilding a study run in Python costs a reader time. Rebuilding one run on licensed software costs them the price of the licence.
That is the whole argument. Different tools solve different problems, and for research published in public the question that matters is how expensive it is for someone else to check.
The historic corpus
TradeStation · MultiCharts · StrategyQuant X
Specified, but rebuilding one needs the same licence we used. These articles predate this standard and each one says so.
New research
Python
Free to rebuild. The reader's cost of checking a study is their time, not a software purchase.
What produced the number, and what shaped the sentence.
Two different questions. The research question, the tests, the data, the analysis and the conclusions come first, and they are ours. Editorial tools - including AI - help organise and tighten how a finding is presented. They do not produce the finding. No result on this site was arrived at by asking a language model what it thought, and where a study found nothing, no amount of editing turns that into something.
Questions about the method
The five asked most often about how these numbers are produced.
Why is buy-and-hold the benchmark?
Because it is the honest alternative. Doing nothing is available to every reader at zero effort and near-zero cost, so a strategy that takes work, risk and attention has to beat it before it is worth anything. Plenty of profitable-looking strategies do not: 0 of 648 ICT backtests and 0 of 2,068 optimized variants beat buy-and-hold out-of-sample.
What is a seeded random control?
A coin-flip entry that fires at the same average frequency as the real signal, run over many random seeds and averaged. It answers the question a profitable backtest cannot answer on its own: would a rule with no information in it have done just as well? Testing against random is how you tell an edge from a market that simply went up.
Why are the backtests frictionless?
Modelling commission and slippage accurately requires assumptions about broker, size and era that are themselves guesses, and a wrong cost model is worse than a stated absence of one. So costs are excluded and that exclusion is stated on every study. Real costs would reduce every net figure published here, more so for high-frequency variants.
Does this standard apply to older articles?
No. It applies to studies published from August 3, 2026 onward. Articles published before that date predate the standard and are not retro-fitted or silently covered by it; they are being brought up to standard as they are re-run, at their existing URLs.
Why publish studies that found nothing?
Because a research site that only publishes its wins is a marketing site. A negative result costs the same work as a positive one and usually says more about how markets behave - and the experiments that failed are the ones a reader cannot get anywhere else.
See it applied
Every study states its method in full - including the ones that found nothing. Browse the research archive.

