What survivorship bias does to a backtest — and why ours has it too
Survivorship bias is why a clean backtest can look better than it will ever perform.
Our open, DOI'd census carries the same bias — it runs only on famous names that still trade, and we say so. On that survivor-only set, 721 of 839 judged backtests (86%) still lagged simply holding the same asset. Here is what the bias is, which way it actually cuts per metric, and how to spot it in your own backtesting.
Survivorship bias flatters a backtest's returns — not the question of whether a rule beat simply holding.
As of September 25, 2026 · frozen census release · DOI 10.5281/zenodo.22974116
A frozen, reproducible snapshot — not a live number and not a promise about the future. 25 famous strategies × 36 famous US stocks & ETFs = 900 strategy-asset combinations, of which 839 had enough trades to judge; each is measured against that asset's own buy & hold, net of modelled costs, over one ~five-year window, with no out-of-sample split, and is not significance-tested.
What survivorship bias is
Survivorship bias is what happens when you measure a group using only the members that lasted to today. The ones that failed — the stocks that delisted, the funds that closed, the coins that died — are gone from the sample, and failure is exactly the bad outcome, so the survivors that remain look better than the full group ever did. A backtest run only on names that still trade has quietly dropped every name that went to zero, so any number computed on the survivors — an average return, a worst drawdown, or a share that beat holding — belongs to the survivors, not to the full set an investor had to choose from at the start.
The classic illustration is Abraham Wald's in the Second World War: the returning bombers showed damage spread over the wings and tail, so the instinct was to armour those spots. Wald saw the opposite — the planes hit in the engines and cockpit were the ones that never came back, and so weren't in the data at all. The gaps in the sample were the whole story.
Survivorship bias, in action — on our own data
Our census is survivor-only, and we say so: 36 famous US stocks and ETFs that still trade today — the names that delisted, went bankrupt or were wound down aren't in the set. These are liquid, well-known names chosen for how famous they are — a hand-picked set of celebrated survivors, not a random or representative draw — and even so, 721 of 839 judged backtests still failed to beat simply holding the same asset over the same five-year window.
The survivors — liquid, well-known, measurable.
Delisted, bankrupt, wound down, or (crypto) dead — not in any free feed, so we can't even count them.
A survivor-only set can still give an honest answer here for one reason: we never score the strategy's return on its own. We score it against the SAME asset's own buy & hold — so whatever the survivor selection flatters, it flatters both sides of that comparison equally, and the question of which one won is left standing.
These are famous names that still trade today. Names that delisted, went bankrupt or were wound down aren't in the set (survivorship bias) — and that history isn't available from the free public data this tool runs on — so this rate isn't a representative base rate for every stock or ETF that has ever traded.
Which way the bias actually cuts
The beat rate. The beat rate is direction-free. Adding the delisted names back would be a near-total loss for buy & hold that an exit-based rule might step out of ahead of the fall (pushing the share that beat holding up) or that a buy-the-dip rule would ride all the way down (pushing it the other way). So we do not claim the real rate is higher or lower than what the survivors show.
Absolute returns. Absolute returns are flattered upward — but on both sides. On a survivor-only set the raw returns of the strategies AND of buy & hold both look better than the full universe would have delivered, and because that lift hits both columns, it does not cleanly move the relative question of whether the rule beat holding.
Drawdown. Drawdown is the one place the bias clearly points one way: the names that went to zero are the deepest holes of all, and they are missing, so the real-world range of drawdowns is wider than this slice and buy & hold's true risk is understated here. See the drawdown study →
None of that licenses a direction on the beat rate itself.
Where it bites — in your backtest, and in a record someone sells you
The failed names are missing from YOUR backtest: delisted and bankrupt stocks, and — far more so — dead coins and rug-pulled tokens. Run a rule on today's survivors and the worst outcomes an investor could actually have picked are simply absent. Test a rule where the bias is named →
The failed names are missing from a record someone SELLS you: closed and merged funds, the quietly-shut hedge funds whose numbers vanish with them, and 'backfilled' track records where only the runs that worked get published. A spotless multi-year curve is often a survivor of exactly this kind. How to spot a fake backtest →
The bias is larger for crypto, where dead coins and rug-pulled tokens vanish from the record entirely — which is one reason this study headlines the stocks-and-ETFs slice and treats crypto as context only.
These are well-known coins that still trade today. Dead coins, rug-pulls and delisted tokens aren't in the set (survivorship bias — larger here than for stocks) — and that history isn't available from the free public data this tool runs on — so this rate isn't a representative base rate for crypto as an asset class.
How to spot it in your own backtesting
An honest backtest runs on data that still contains the names that failed. A clean curve built only on today's survivors is a hypothesis, not a result — read it as an upper bound, and lean on what a rule does going forward rather than on how neatly it fits a set of known winners. We disclose our own survivor-only bias rather than hide it, and we have not solved it.
The limits here are real: this is about 36 hand-picked famous names, and we cannot put a number on our own survivorship gap — the data for the names that failed isn't available from the free public feeds this tool runs on. It is one roughly five-year window, measured with no out-of-sample split, and it is not significance-tested.
Test a rule on the asset you care about — in the free sandbox, the survivorship note travels with every result. No sign-up.
Sources
- Brown, S. J., Goetzmann, W. N., Ibbotson, R. G., & Ross, S. A. (1992). Survivorship Bias in Performance Studies. Review of Financial Studies, 5(4), 553–580.
- Malkiel, B. G. (1995). Returns from Investing in Equity Mutual Funds 1971 to 1991. Journal of Finance, 50(2), 549–572.
- The returning-bombers intuition is Abraham Wald's, from his WWII work on aircraft survivability.
- Our own data: The Honest Backtest Census (frozen release, September 25, 2026; DOI 10.5281/zenodo.22974116).
These are simulated backtests over roughly a five-year window on a survivorship-biased set of famous names, frozen as a citable release — hypothetical results, not live figures and not advice. Past performance is not a reliable indicator of future results. This is an educational, analytical tool — not investment advice, and not a recommendation to buy or sell anything. You make all decisions and execute on your own broker.