Stratsemble
Glossary

Trading & backtesting glossary

Plain-English definitions for every metric, check and indicator on Stratsemble — the same explanations that sit behind each on the site. Each one says what it means and its catch, because a definition that hides the downside is how people get fooled. It's a reference, not advice.

Reading a result

The headline numbers a backtest reports — and why the biggest one is rarely the one that matters.

Total return
How much the strategy grew your money over the whole period, after costs. It's the headline — but on its own it hides how much risk you took and whether just buying and holding would have done better.
vs Buy & hold
The strategy's return minus simply buying and holding the same asset. This is the honest scoreboard: a big return that still trails buy & hold means the strategy added work and risk for nothing.
Buy & hold
What the asset itself did over the same window if you'd just bought it and done nothing — the benchmark the strategy has to beat. If the strategy can't clear this line, all its trading added risk and cost for nothing.
CAGR
The smoothed yearly growth rate — the steady annual return that would produce the same end result. Easier to compare across different time spans than a raw total.
Win rate
The share of trades that made money. High isn't automatically good: a strategy can win often and still lose overall if its few losers are large.
Profit factor
Gross profit divided by gross loss. Above 1 means winners outweigh losers; below 1 loses money overall. Around 1.0 is a coin flip after costs.
Time in market
How much of the period the strategy actually held a position. A decent return with low exposure can mean less risk — or just luck in a few short windows.
Trades
How many round-trip trades the strategy made. Too few (under a handful) and any result is mostly luck — there isn't enough evidence to trust it.
Best trade
The single most profitable trade. If one trade carries most of the return, the strategy is fragile — remove that one lucky trade and little is left.
Worst trade
The single biggest losing trade — a reminder of how bad one position got, which averages and totals quietly smooth over.
Costs & slippage
Every backtest here charges realistic trading costs plus slippage — the small gap between the price you'd want and the price you'd actually get, which widens on fast or thin moves. Leaving these out is the most common way a backtest flatters itself; they're always on here.

Risk & drawdown

How much pain a result hid behind its average — the losses, the wait to recover, the ride.

Max drawdown
The worst peak-to-trough drop along the way — the deepest hole you'd have been sitting in. It's the pain you'd have had to stomach without giving up.
Worst drawdowns
The deepest peak-to-trough falls, with dates and timing. 'To trough' is how long the fall took; 'Underwater' is how long until you got back to the old high; 'Recovered' says whether you ever did. It's the pain a single max-drawdown number hides — long underwater stretches are where people give up.
Sharpe ratio
Return earned per unit of wobble (volatility). Higher means a smoother ride; roughly, above 1 is good. Its blind spot: it treats rare, huge losses too kindly.
Sortino ratio
Like Sharpe, but it only counts downside wobble — it doesn't punish a strategy for big up-moves. A fairer read of pain-adjusted return.
Calmar ratio
Annual return divided by the worst drawdown. It asks a blunt question — how much yearly gain did you get for the deepest hole you had to sit through? Higher is a better reward-for-pain trade.
Ulcer index
A single number for how much time you spent underwater and how deep — it punishes long, deep drawdowns more than brief dips. Lower is a calmer ride; two strategies with the same return can have very different ulcers.
Time underwater
The share of the period the strategy spent below a previous high — i.e. how often you'd have been staring at a losing position waiting to get back to even. Long stretches are where people give up.
Recovery factor
Total return divided by the worst drawdown — roughly, how many times over the strategy earned back its deepest pain. Above ~2 is healthy; near or below 1 means the gains barely covered the drawdowns.
Tail ratio
The size of the best days versus the worst days (95th vs 5th percentile of daily moves). Above 1 means the big up-days outweigh the big down-days; below 1 means the crashes are bigger than the pops.
Volatility
How much the portfolio's value bounced around, per year. More volatility means a rougher ride for the same destination.

Is the edge real? (reality checks)

The checks and views that interrogate a good backtest — the statistical ones can only rule an edge out, never prove it will pay forward.

Robustness
Whether an edge survives small changes to its settings. A result that only works at one exact number is probably curve-fit to the past; one that keeps working across a range of nearby settings is far more likely to be real.
Robustness heatmap
We re-run the strategy across a grid of nearby settings. The more of those settings that still beat buy & hold, the more robust the edge; if only one isolated setting worked, it was probably fit to the past. The colour follows whichever metric you pick, but the verdict is always judged on beating buy & hold.
Resampling your trades
We redraw your trades with replacement a thousand times to see how much the result leaned on a handful of specific (lucky) trades rather than a repeatable edge — the compounded return is the same whatever order they came in, so this is about WHICH trades you happened to get, not their sequence. A result that survives redrawing is sturdier than one that falls apart. (Different from the noise floor, which tests whether random TIMING would have done as well.)
Noise floor (vs random)
We take your trades and slide their entries to random dates a thousand times — same asset, same number of trades, same holding lengths, same sides — and see how often random timing did as well as your rules. If you can't clearly beat random entries with the same activity, the 'edge' is probably luck. On a strongly trending asset even random entries do well, so beating buy & hold still matters more. It does NOT correct for cherry-picking (that's the deflated Sharpe).
Deflated Sharpe (cherry-picking check)
When you scan one strategy across many assets and keep the best, some of that top result is luck — try enough assets and one looks good by chance. This takes the winner's risk-adjusted edge over buy & hold and discounts it for how many assets you tried, leaving the share that's more than selection luck. A high number means the best result probably isn't just the luckiest of the bunch — it does NOT mean the strategy will make money forward, only that this scan didn't manufacture the result by trying a lot. It counts only the assets in THIS scan: every other strategy, window and scan you've tried makes the real correction harsher, so read it as a best case, not a point estimate. Measured on daily returns vs buy & hold, adjusted for the fact that nearby days move together; a strategy that mostly sits in cash and wins by dodging a few crashes shows a small, choppy daily edge, so a low number there means 'not proven here', not 'disproven'. The winner shown is the one that beat buy & hold by the most; scoring it on its risk-adjusted edge is a deliberately conservative way to run the check. A different kind of luck from the noise floor (random timing on one result) and Monte-Carlo (which trades carried it) — these checks can only disprove a strategy, never prove one. Same scan, same number (no randomness).
Tuning-luck check
When you sweep a setting and keep the value that looks best, some of that is just luck — a value that happened to fit the past. This splits the history into slices every possible way, and for each split picks the setting that looked best on the 'training' slices, then checks where it lands on the slices it never saw. The number is how often that best-looking setting slips to the middle of the pack or worse out-of-sample: high means the winning setting was mostly fit to the past; low means the tuning didn't obviously overfit — which is NOT the same as 'the strategy works', and says nothing about profit going forward. It covers only the SETTINGS axis (it needs a real range of settings, not near-identical ones); if you also scanned many assets, that's a separate kind of luck the deflated Sharpe covers. The many train/test splits reuse the same slices of history in overlapping combinations, so their count reflects how thoroughly it's checked, not an independent sample size. Don't re-run it hunting for a lower number.
Walk-forward / out-of-sample
Tune the strategy on one slice of history, then test those settings — untouched — on a later slice it never saw; roll forward and stitch the unseen slices together. The gap between the hindsight-tuned result and the walk-forward result is the overfitting you can't actually use. The closest a backtest gets to honest — but it's still ONE out-of-sample path over relatively few trades, so pair it with the robustness heatmap and Monte-Carlo, and don't re-run it hunting for a good-looking result.
Folds that count
History is split into folds, each with an out-of-sample slice. A fold only counts toward the verdict if its out-of-sample slice had enough real trades (at least 3) to mean something; and the verdict needs at least 3 of the 4 folds to count — otherwise the sample is too thin to judge honestly.
Is it just leveraged beta?
Measured only on the bars you were invested: 'beta' is how much your position moved per unit of the asset's move, and 'R²' is how tightly it tracked. Beta near 1 with a high R² means you basically held the asset — so any excess is the TIMING of plain exposure, not a separate return source. Above ~1.15 is leverage or oversized bets: bigger swings both ways, more risk, not skill. This check can only rule an edge OUT (show the result is just exposure) — it can NEVER prove one; whether the timing itself is real is what the noise-floor and walk-forward checks are for. The benchmark includes dividends while your P&L is price-only, so beta reads a touch low on high-yield names.
Market regimes
How the strategy did when the market was rising (bull) or falling (bear) — set from the benchmark versus its own trailing average, using only past data (no hindsight). The high-volatility row is a separate, overlapping slice: the most volatile roughly-quarter of days measured across the whole period. A strategy that only makes money in bull markets is fragile; one that holds up across conditions is sturdier — the honest test of 'did this have an edge, or just ride the market up?'
Prop-firm rules check
Would this strategy have survived a funded-account challenge (FTMO/Topstep/Apex-style)? It overlays a firm's published risk limits — max daily loss, max or trailing drawdown, profit target, minimum trading days — on your strategy's own ABSOLUTE account equity (not the vs-buy&hold number shown elsewhere; a firm measures the account). Because firms enforce those limits INTRADAY and daily bars only have closes, we test every breach against the bar's worst intraday point: a computed BREACH is the read worth knowing (worst case, the account reached that level — we assume the position was held through the bar's range), while a clean run is only 'didn't breach on a daily basis' — never a promise you'll pass, since the intraday path and the firm's exact costs/instrument differ. It slides the challenge across every non-overlapping window and reports the distribution (not one cherry-picked start), and refuses to state a rate when there are too few windows. The honest next step is never 'buy the challenge' — it's to forward-test it on paper, then on your own broker.
Data quality
An automatic check of the underlying price data for gaps, stale feeds, or implausible jumps — problems in the data, not the strategy, that could quietly distort a result.
Bar-by-bar replay
Replays the strategy's history one bar at a time with the future hidden, so you can watch each mechanical trade commit at a bar's close before the next bar exists — a visceral demonstration that the engine acts only on past data. The guarantee lives in the closed-bar-safe engine (it fills at the next open, never on the bar it just used to decide); this replay just unveils those already-computed, causal results in time order. The chart's vertical scale only ever uses the bars revealed so far, so nothing about the future leaks in early. It describes what the rules did; it is not a prediction and not a game — there's nothing to 'call'.
Trades on the price
Your strategy's actual entries (●) and exits (○) drawn on the asset's own price, so you can see WHERE it bought and sold — not just an equity curve. It shows exactly what the engine did: fills at the modeled price (next open, costs and slippage on), no look-ahead. Green/red mark whether each trade made or lost money — that's data about what happened, not a verdict. It describes the past; it makes no claim about the future.

Technical indicators

What each indicator MEASURES — and its blind spot. None of them predicts the future; each is a gauge with a catch.

Simple moving average
The plain average of the last N closing prices, re-figured each bar — it smooths jumpy price into a single trend line. The catch is lag: it reacts late to turns, and it weights a price from N bars ago exactly the same as today's.
Exponential moving average
A moving average that leans on recent prices more heavily, so it turns faster than a simple average of the same length. Faster means it also reacts to more noise — so it whips back and forth more in choppy, directionless markets.
MACD
The gap between a fast and a slow moving average (the MACD line), plus a slower average of that gap (the signal line). It gauges momentum — the line pulling above its signal is read as momentum turning up — but because it's built from lagging averages it still turns after price does, and drifts around zero in sideways markets.
Bollinger bands
A moving average with an upper and lower band set a number of standard deviations away, so the band widens when price is volatile and pinches in when it's calm. A close outside a band means price is statistically stretched for its recent range — but stretched is not the same as about-to-reverse: in a strong trend price can 'walk the band' for a long time.
Keltner channel
A band around a moving average set a multiple of average true range wide — a volatility-based cousin of Bollinger Bands that measures range with the true bar range instead of the spread of closes, so it reacts differently to gaps. Same idea, same catch: 'outside the band' means stretched, not turning.
Donchian channel
The highest high and lowest low of the last N bars, drawn as a channel; a close above the top is a new N-bar high — the classic breakout trigger. Here the channel deliberately excludes the current bar, so a 'new high' can't peek at the very bar it's testing (no look-ahead). The catch: breakouts fail often in range-bound markets, buying the top just before a snap-back.
RSI
The Relative Strength Index measures how one-sided recent moves have been — the size of gains versus losses over the last N bars, on a 0–100 scale. Readings near the top (conventionally 'overbought') or bottom ('oversold') mean price has moved a lot one way; it's a gauge of stretch, not a prediction of a turn, and in a strong trend it can sit pinned at an extreme while price keeps running.
Stochastic oscillator
Shows where the latest close sits inside the recent high–low range — 0 at the bottom of the range, 100 at the top — with a smoothed second line (%D). Near the top is called 'overbought', near the bottom 'oversold', but like any range gauge it stays pinned near an extreme through a strong trend, so a stretched reading is not a reversal signal.
Stochastic RSI
The Stochastic formula applied to RSI instead of to price: it shows where RSI sits within its own recent range, on a 0–1 scale. Because it stretches an already-smoothed indicator, it's faster and noisier than plain RSI — it hits the extremes far more often, so it's more often wrong there. It gauges how stretched momentum is, not a prediction of a turn.
Commodity channel index (CCI)
The Commodity Channel Index measures how far the typical price (high, low and close averaged) has stretched from its recent average, scaled by its own average deviation. On very quiet ranges that deviation is tiny, so CCI can spike past ±100 on almost no real movement — it reads stretch, and calm markets exaggerate it.
Awesome oscillator
The gap between a fast (5-bar) and a slow (34-bar) average of each bar's midpoint — positive when the fast average leads the slow one. It's essentially a moving-average crossover in disguise, so it carries the same catch: it lags turns and whipsaws in sideways markets.
Rate of change (ROC)
The percent change in price versus N bars ago — the simplest momentum gauge: positive means today's price is above the price N bars back. It's raw and unsmoothed, so it reacts fast but is jumpy, and it says nothing about whether that move is likely to continue or reverse.
Time-series momentum
Measures the asset against ITSELF over time — its own trailing return over roughly a year, skipping the most recent month (the '12-1' rule, which drops the short-term bounce-back), scored against the average of its own history. A positive reading means recent momentum is above that historical average — it can read positive even while the raw return is negative — and it does not rank or compare different assets. A slow read with a long (~2-year) warm-up, prone to whipsaws and to being caught out by sudden crashes.
SuperTrend
A trailing line placed a multiple of average true range above or below price that flips from bearish to bullish as price crosses it. It adapts its distance to volatility and trails trends closely — but it flips late at turns and whipsaws in sideways markets, giving back part of a move at each flip.
Parabolic SAR
Wilder's 'stop and reverse' — a trailing dot that accelerates toward price as a trend runs and flips the trend when price crosses it. Each dot is projected from the prior bar, so it never peeks at the bar it's testing. It trails winners tightly, but it reverses often, so it whipsaws badly in sideways markets and gives back part of every move at each flip.
ADX / DMI
The Directional Movement system: +DI and −DI show which side — up or down — is winning, and ADX (0–100) shows how strong the trend is regardless of direction. A high ADX means a real trend is present, but ADX rises only AFTER a trend is underway and says nothing about which way — so it's a filter, not a signal.
Ichimoku cloud
The Ichimoku 'cloud' is a band built from midpoints of past highs and lows, shifted forward so it derives only from past data: price above the cloud is a bullish regime, below is bearish, inside is a neutral zone. It's a trend/regime filter that lags turns by design and whips around the cloud edge in choppy markets. (We deliberately leave out the one Ichimoku line that looks backward in time — using it would be look-ahead.)
Tenkan & Kijun lines
The two fast midlines inside Ichimoku: the Tenkan (conversion) is the midpoint — highest high and lowest low, averaged — of about the last 9 bars, and the Kijun (base) is the slower ~26-bar version. Here they mostly shape the cloud rather than fire signals of their own: this strategy trades only whether price is above or below the cloud, so these two lines are the raw material behind it, not triggers.
On-balance volume
A running total that adds a bar's volume on an up-close and subtracts it on a down-close, so the line climbs when volume flows in on up-days. Its absolute level is arbitrary — it depends on where your data happens to start — so only the crossing of its own moving averages (a difference, in which the arbitrary offset cancels) carries information, never the raw number, and on a thin or volume-less feed it says very little.
Money flow index
A volume-weighted cousin of RSI — it measures stretched buying versus selling pressure on a 0–100 scale, weighting each bar by its volume. Strip the volume out on thin or volume-less data and it behaves much like RSI (though it scores each bar by its money-flow level, not the size of the move), and like any stretch gauge it can stay pinned at an extreme through a strong trend.
Chaikin money flow
Measures where each bar's close finished inside its high–low range (a close near the high suggests buying pressure), weights that by volume, and averages it to a reading between −1 and +1 — above zero suggests accumulation, below zero distribution. It lags turns and whipsaws sideways. A bar with no range contributes a neutral zero rather than a fabricated ±1, but it doesn't blank the reading; only a volume-less feed (or the warm-up) leaves no reading at all.
Average true range
The typical size of a bar's move, gaps from the prior close included — a pure volatility gauge, not a direction one. It's the yardstick behind volatility-adaptive tools (SuperTrend, Keltner, ATR stops): bigger ATR means wider swings, so stops and bands set from it sit further from price on jumpy assets. The catch: it's in the asset's own price units, so an ATR of $5 is huge on a $20 stock and tiny on a $2,000 one — you can't compare raw ATR across assets, and a single gap or spike inflates it for many bars.
Heikin-Ashi candles
Each candle redrawn from an average of its own open/high/low/close and the prior candle's averaged values, so the chart looks smoother and trends look calmer. The smoothing is cosmetic — it adds lag, not information — and it hides the real intraday drawdown: a calm-looking Heikin-Ashi chart is not the account you'd have lived through. Every backtest here still fills and scores on the REAL price, never the smoothed candle.
Volume
How many shares or contracts changed hands in a bar — the fuel behind the volume indicators (OBV, MFI, CMF). It can be informative on liquid names, but noisy and patchy on thin ones, and free daily feeds sometimes report it split-adjusted or incomplete — so a volume signal is only as trustworthy as the feed under it.
Standard deviation (σ)
A measure of how spread out recent prices are around their average — the width behind Bollinger Bands. A bigger σ means a wider, more volatile range; but it treats an up-move and a down-move the same, and leans on a tidy bell-curve spread that real markets, with their fat tails and crashes, don't quite follow.

Risk controls

The exits and sizing rules you can add to a strategy — each one a trade-off, not a free lunch.

Stop-loss
An automatic exit if the price falls a set % below your entry, to cap how much one trade can lose. The catch: normal wobble can stop you out just before the price recovers, so a tight stop can quietly cut your returns.
Take-profit
An automatic exit once a trade is up a set %, to lock in a gain. The catch: it also caps your winners, so it can close a big trend early and leave money on the table.
Trailing stop
A stop that follows the price up and only ever moves in your favour, exiting if the price then falls a set % from its high. It lets winners run while protecting gains — but a sharp dip that later recovers will still stop you out.
ATR stop
A stop placed a multiple of recent volatility below your entry — ATR (average true range) is the typical size of a bar's move, so the stop sits wider on jumpy assets and tighter on calm ones. It adapts to the asset instead of a fixed %, but a bigger multiple means a deeper loss before it triggers.
Time stop
Exit a trade after a set number of bars whatever the price, so dead trades don't tie up your capital indefinitely. Handy for rules that should work quickly — but it can also close a slow winner before it pays off.
Position sizing
How much of your capital goes into each trade. 'Fixed %' always commits the same share; 'Risk %' sizes each trade so that hitting its stop costs a set % of capital — smaller when the stop is far away, larger when it's near. Risk sizing needs a stop to measure against.

Portfolios

What happens when you blend several strategies or assets — diversification, contribution, and whether optimizing the mix even helps.

Correlation
How much the legs moved together, measured only on days they both traded. Low or negative means they offset each other (real diversification); high means they're basically one bet wearing several hats.
Contribution
How much each leg added to (or subtracted from) the portfolio's total return, after its weight. The contributions add up exactly to the portfolio's return.
Stress the weights (does optimizing help?)
A sharper question than the blend above: does CHOOSING each leg's weight from past performance actually beat just splitting equally, out-of-sample? We apply walk-forward to the weights — pick the allocation that looked best on each past slice, then test it untouched on the next slice it never saw, and stitch those unseen slices together. Because maximizing past return over a buy-and-hold blend always concentrates 100% on the single best past leg (there is no diversified 'optimum' to find), that is exactly what it does — and the winner changing fold to fold IS the instability. The honest headline is the comparison: optimized-from-the-past vs a plain equal split (1/N), both measured against holding the same assets equally. Usually 1/N does as well or better — the well-known result, and the reason we do NOT sell an optimizer. It resets the weights each period (unlike the hold-once blend above), the numbers mean 'beat an equal hold', not 'made money', and it is one out-of-sample path — never a recommendation to use any particular weights.
Blended buy & hold
What you'd have made by just buying the same assets in the same weights and holding them — no strategy, no trading. It's the yardstick the blend has to clear; beating it by only a little isn't worth the extra work and risk.
vs Blended buy & hold
The blended strategies' return minus that same weighted buy & hold. This is the honest scoreboard for a portfolio: if combining strategies can't clear a plain passive split of the same assets, the extra complexity earned nothing.

Track record & forward testing

The evidence that actually accrues over time — a forward test on data the rules were never fit to, checked against your own real fills.

Your record
Every strategy you test, forward-prove, and check against your own real fills accretes into a track record that's yours — and one that can't be rebuilt overnight. It's a real switching cost, honestly earned: the value is the accumulated evidence you can trust your own rules, not a score.
Held up out-of-sample
A forward test that has accumulated enough out-of-sample trades to tell its edge from luck AND is still beating buy & hold at that point. It's the strongest evidence a test here can give — genuinely earned on data the rules were never fit to — but still not a promise the edge lasts.
Real fills checked
Real broker fills you've logged against a forward test, so you can see how far the honest simulation drifted from your actual execution. Once you've verified our numbers with your own trade — your broker, your money, we never touch a cent — no unchecked tool can claim the same.
Log reconciliation
A side-by-side of the trades you logged against the trades the sim's rules took over the same window. It only reconciles — it does NOT score your discipline, because it genuinely can't: when a sim signal has no matching fill in your log, that's either a trade you chose not to take OR one you took and never logged, and nothing in the data can tell those apart. So it names what your log doesn't cover, shows where your fills and the sim's timing differed (both ways, never just the misses), and asks you to log the rest — it never tells you a number you 'gave up'. The sim's own entry/exit is one rule seen with hindsight, not the right answer to grade yourself against.
Signals your log doesn't cover
Sim trades in your logged window with no matching entry in your log. Crucially this is NOT 'signals you skipped': it's either trades you didn't take, or trades you took and haven't logged — this tool can't distinguish those, so it never counts them as a discipline failure. Logging the rest is what completes the picture.
Proof clock
A forward test only means something once it has enough out-of-sample trades to tell a real edge from luck. This shows how far along that road it is — the target is estimated from this test's own trades (the same math as the trade sample-size calculator), and it refines as more trades close. You can't rush it, and that's the point: trades your rules couldn't have been fit to are the only ones that count. Reaching the target is genuine out-of-sample evidence — still never a promise the edge lasts. If the test isn't beating buy & hold yet, it says so plainly instead of showing a countdown.

Alerts & signals

Mechanical heads-ups when your own rule triggers — never advice, and you always place any trade yourself on your own broker.

Alerts
A mechanical heads-up that one of your own rules just triggered on the daily close — not advice, and not a signal to follow blindly. You place any trade yourself on your own broker; we never touch your funds or your account. The trigger is measured on the close, but a real fill happens at the next open and can differ. Best used only after you've forward-tested the strategy — a backtest alone isn't proof.
Notification channels
Get a fired alert somewhere other than email — a Discord channel in this version. It's a mechanical notification that your own rule triggered, never advice: I only ever post a message, and I never place, route, or size an order. You paste a Discord webhook URL you create in your own server; it's stored encrypted and shown back only masked. Email always fires too. If a webhook stops working (you deleted it), the channel quietly disables itself so it doesn't keep failing.
Copy this signal to your broker
A plain-text restatement of what your own rule signalled — the action (buy/sell), the asset, the signal date, and a reference-only price — that you copy and re-enter yourself on your own broker. It deliberately carries NO size: sizing a trade to your circumstances would be personalized advice, so you always decide the size. We never place, route, or size an order, and never touch a cent. The price is our data-feed reference, not an order price, and our symbol is a data key — confirm the exact instrument on your broker, since symbols differ across venues. After you trade, you can log what you actually got on your forward test to compare sim vs reality.
Alert expiry
An alert that's been sitting idle eventually expires, so abandoned alerts don't quietly eat your quota — hit Renew to reset the clock (the clock also resets whenever it fires). An alert that's currently holding a position never expires (the exit signal is still owed), and higher plans expire later or not at all. Expiring only stops the watching; it doesn't delete your signal history.
Go deeper

These are the pieces. The full recipe — the honesty rules the engine can't bend — is on the methodology page, and you can put any of these to work in the free backtest sandbox.