Test many strategies — without fooling yourself
Run a rule across hundreds of assets, or blend a dozen legs into a portfolio, and something will always look excellent. The catch is the more you test, the better the best result looks by chance alone — that's the multiple-comparisons problem, and it's how scale quietly manufactures edges that aren't there. The workbench is built around the correction for it.
More tests don't mean more edge — they mean a bigger haircut. When you search across many strategies and assets, the honest question isn't “what won?” but “would the winner still stand once we discount it for how many we tried?” That discount is the deflated Sharpe ratio, and it runs on every batch scan here.
The multiple-testing problem, plainly
Flip enough coins and one run of ten heads shows up — not because the coin is special, but because you ran a lot of coins. Backtesting at scale is the same: test 500 strategy-and-asset pairs and the top few will post strong numbers on the past even if none has a durable edge. A single impressive Sharpe ratio means one thing; the best Sharpe out of hundreds means much less, and the naive number hides that. The deflated Sharpe ratio corrects it — it lowers the bar in proportion to how many strategies were tried and how noisy their returns are, so a standout has to clear a higher, honest threshold before it counts as more than luck.
New to the terms? Real signal or just luck? and the glossary cover the Sharpe ratio and overfitting in plain language.
The workbench
Four tools for testing at scale — each one designed to make a result harder to believe, not easier. Winners and losers are shown; nothing here promises you an edge.
Batch scan
Run one strategy across many assets at once, ranked whole-glass — then read the deflated-Sharpe haircut that discounts the top result for how many you tried.
Portfolio
Blend several strategy-and-asset legs into one combined test, and check whether tuning the weights on the past actually beats a plain equal split out-of-sample.
Robustness
Perturb the parameters and the window and watch how much the result moves — a fragile edge that only works at one exact setting shows itself here.
Walk-forward
Fit on the past, score only on data that came after — the out-of-sample check that separates a rule that generalises from one that memorised its history.
Batch scan and portfolio are part of the System plan; robustness and walk-forward give every account a free monthly allowance and go unlimited on Prove; and the honest verdict on any single strategy is always free on the census.During the launch beta, all of it is free.
Run one rule across a whole asset class and read the results with the multiple-testing haircut applied, so the top of the list is honest about how much of its shine is luck.
Hypothetical / simulated results — past performance is not a reliable indicator of future results. These are educational, analytical tools, not investment advice and not a recommendation to buy or sell anything. A haircut-adjusted result is a more honest read of the past, never a prediction; you make all decisions and execute on your own broker.