The honest math of finding a strategy that beats holding — and the trap that fakes it
Most famous strategies lag simply holding; a minority beat it. The honest hunt is finding which — and putting it forward on data it hasn't seen, before you risk a cent.
The honest answer is a search, not a secret. Most ideas don't beat simply holding — but a few do, and you raise your chance of finding one not by running more tests, but by holding every one to a bar fixed in advance, counting them all, and taking a bigger haircut the more you run. Breadth without that discipline is just more ways to fool yourself.
Re-running one idea until it looks good — the trap
Take a single idea and try it fifty ways — tweak the numbers, swap a filter, keep the best-looking run. You didn't find an edge; you found the luckiest version of one idea. The more ways you try the same idea, the better the best one looks by luck alone — and your real chance of a keeper hasn't moved an inch. That's why every build you try is counted, and the haircut gets harsher the more you try.
Trying distinct ideas, each held to the same test — the search
Now take fifty different mechanisms — not fifty tunings of one — and put each through the same gauntlet, counted the same way. Forward-testing on data an idea was never fitted to is the strongest single filter there is — but it isn't free of luck either: take the best of fifty forward survivors and you've run fifty comparisons, and one can pass its window by chance. So distinct ideas are counted too, and the more you run, the bigger the haircut the winner has to clear. Running more honest, distinct tests raises your chance of finding a keeper AND of landing a lucky fluke — the count and the haircut are what tell them apart.
The shape of it
Say a distinct, honestly-tested idea has some small chance of holding up — call it p, and p is small (most ideas wash out, and a forward-survivor is rarer still than the in-sample share the live census shows). Test one idea and your chance of finding a keeper is p. Test k independent ideas, each to the same fixed bar, and your chance of finding at least one is 1 − (1 − p)ᵏ — it climbs with k.
Illustrative — p is the live in-sample share that beat; the fluke line assumes even a strict bar lets roughly one lucky null idea in ten slip through. at the live in-sample rate — p ≈ 27%.
Four things keep that honest — and without all four it's a cheat:
- 1k counts independent ideas, each forward-tested once on data it was never fitted to — never k retunes of one idea on the same data, which doesn't raise p.
- 2The bar never moves — every idea clears the same out-of-sample, multiplicity-corrected test.
- 3The outcome is finding a forward-survivor, not a profit — the rare idea worth taking forward with your eyes open, never proof it will make money.
- 4The winner carries the count — one survivor among fifty tries is not fifty tries all pointing one way, so the more you run, the more a lucky pass can hide among them, and the survivor's edge is deflated for how many ideas you searched.
k raises your chance of a keeper AND of a fluke; only that haircut tells them apart.
The whole-glass census the odds are drawn from is public — every winner and every loser. See how few beat holding, then run your own idea and put it to the same honest test.
The base rate p is computed live from simulated backtests over roughly the last five years and updates as new bars close. These are famous names that still trade today. Names that delisted, went bankrupt or were wound down aren't in the set (survivorship bias) — and that history isn't available from the free public data this tool runs on — so this rate isn't a representative base rate for every stock or ETF that has ever traded. Hypothetical / simulated results — past performance is not a reliable indicator of future results. This is an educational, analytical tool, not investment advice and not a recommendation to buy or sell anything. You make all decisions and execute on your own broker.