Stratsemble
The quant workbench · testing at scale, honestly

Test many strategies — without fooling yourself

Run a rule across hundreds of assets, or blend a dozen legs into a portfolio, and something will always look excellent. The catch is the more you test, the better the best result looks by chance alone — that's the multiple-comparisons problem, and it's how scale quietly manufactures edges that aren't there. The workbench is built around the correction for it.

More tests don't mean more edge — they mean a bigger haircut. When you search across many strategies and assets, the honest question isn't “what won?” but “would the winner still stand once we discount it for how many we tried?” That discount is the deflated Sharpe ratio, and it runs on every batch scan here.

The multiple-testing problem, plainly

Flip enough coins and one run of ten heads shows up — not because the coin is special, but because you ran a lot of coins. Backtesting at scale is the same: test 500 strategy-and-asset pairs and the top few will post strong numbers on the past even if none has a durable edge. A single impressive Sharpe ratio means one thing; the best Sharpe out of hundreds means much less, and the naive number hides that. The deflated Sharpe ratio corrects it — it lowers the bar in proportion to how many strategies were tried and how noisy their returns are, so a standout has to clear a higher, honest threshold before it counts as more than luck.

New to the terms? Real signal or just luck? and the glossary cover the Sharpe ratio and overfitting in plain language.

The workbench

Four tools for testing at scale — each one designed to make a result harder to believe, not easier. Winners and losers are shown; nothing here promises you an edge.

Batch scan and portfolio are part of the System plan; robustness and walk-forward give every account a free monthly allowance and go unlimited on Prove; and the honest verdict on any single strategy is always free on the census.During the launch beta, all of it is free.

Start with a scan

Run one rule across a whole asset class and read the results with the multiple-testing haircut applied, so the top of the list is honest about how much of its shine is luck.

Hypothetical / simulated results — past performance is not a reliable indicator of future results. These are educational, analytical tools, not investment advice and not a recommendation to buy or sell anything. A haircut-adjusted result is a more honest read of the past, never a prediction; you make all decisions and execute on your own broker.