Every strategy eventually asks the same question: is this drawdown normal, or is the edge gone? Equity Control answers it with two tools: a health diagnosis (is the recent path inside the strategy's historical physiology?) and an overlay simulation (what would a disciplined On · Reduce · Off governance have cost or saved?).
You can run it on the whole portfolio or a single strategy. Active weekday filters are respected, and — like Monte Carlo — the series is built from operating days only (days with non-zero P/L). All windows, cooldowns and lags are counted in operating days, so a slow strategy is not penalised by calendar time.
Analysis, not signals
Equity Control is a diagnostic instrument. Its states are advisory readings on your data, not trading orders: VEEMAN does not connect to a broker and does not execute anything. Whether to act on a Stopped reading is — and remains — your decision.
The series is split in two:
At least 60 baseline days and 10 monitoring days are required. Reliability tiers: with 300+ baseline days thresholds are fully estimable; 150–299 days are indicative (tail thresholds carry wide uncertainty); below 150 the diagnosis should be read as descriptive only. A warning is shown for each tier below full.
Setting the split by date. The primary control is In-Sample end: the In-Sample (calibration) runs up to that date, and everything after it is the monitored live window. It takes precedence over the numeric live-tail length (kept under Advanced for those who prefer counting operating days). One of the two must be set — until you define the split there is no diagnosis (no auto-default is applied, so opening the panel does not pre-judge an In-Sample you did not choose). Where you put this boundary is a first-order decision — the same strategy on the same data can read healthy or stopped depending on whether a regime shift falls inside the calibration or inside the monitored window. Choose the In-Sample to match the backtest you actually validated.
After a successful run the parameter form folds into a one-line recipe strip (source · split · governance · α) so the page opens on the verdict; Edit parameters reopens the full form, and Collapse parameters folds it back without re-running.
Bounding the range by date. By default the whole history is analysed. Two optional calendar fields under Advanced restrict it: Monitor from drops baseline older than that date, and To (as-of) ends the analysis on a past date — the diagnosis then runs as if that day were today, which is how you replay the governance at an earlier point. The live tail is always the final stretch of the bounded range. The Align with session button (next to those fields) fills both with the current session's data span (the union of the visible strategies' dates, or the selected strategy's own range).
All three calendar fields (In-Sample end, Monitor from, To) are bounded by that same span, and each one also bounds the others. A date typed by hand outside the data — or an inverted from/to pair — is pulled back to the nearest valid value: the field is rewritten when you leave it and its hint turns into a warning, so the run never starts on a window that contains no days.
Before anything else, the module checks that the In-Sample is large enough to validate the edge — you cannot meaningfully monitor the decay of an edge you have not first established. The check follows Bailey and López de Prado (2012) and combines two hard floors with a power criterion:
1 / Sharpe²: a weak edge needs a much longer track record
to be told apart from zero.The gate passes when the floors are met and PSR ≥ 95% (equivalently, In-Sample days ≥ MinTRL). The Sharpe used is per-observation (MinTRL is in operating days); the panel also shows the annualised figure.
Hard block on insufficient data
If the In-Sample cannot validate the edge, the diagnosis is blocked — the module does not produce a verdict. The pre-flight semaphore under the form shows the outcome live as you move the In-Sample end date: green (edge validated, Run enabled — a quiet one-line confirmation with PSR and MinTRL, full detail in its tooltip) or red with the shortfall (e.g. "you have 285 days, you need ~801 — extend the In-Sample"). A non-positive In-Sample edge (mean P/L ≤ 0) is also blocked: there is no advantage to validate.
Thresholds are not calibrated on the raw baseline edge. A healthy strategy typically earns less
live than in its backtest (selection and overfitting haircuts), so calibrating on the raw mean would
flag perfectly healthy strategies. The "healthy" hypothesis H0 is the baseline series with its mean
shrunk to (1 − haircut) · μ (default haircut 25%), preserving volatility, skew and clustering.
Scenario paths are generated with the stationary bootstrap (Politis-Romano) using the
Politis-White automatic block length — the same implementation as the Monte Carlo module.
If the baseline mean is ≤ 0, or its t-statistic is below 1.5, the diagnosis warns you explicitly: you cannot meaningfully detect the decay of an edge that never showed up.
Monitoring is continuous, so pointwise percentile thresholds would fire far more often than their nominal level (checking a "95%" line every day for a year yields many times 5% false alarms). The corridor uses simultaneous bands instead:
q(t) = the mean drawdown curve under H0 at each operating day t (drawdown from the running
peak, starting flat at the window start — measured in the corridor unit set by the basis: dollars
or geometric percent, see Corridor basis);M = max_t DD(t) / q(t);c · q(t) with c the (1 − α) quantile of M — so the probability that
a healthy path crosses the band anywhere on the horizon is α, by construction.Three levels are drawn: Watch (α = 4×stop budget), Reduce (2×) and Stop (the α you set,
default 5%). This 4×/2×/1× spacing is the default anchored ladder, not a theorem — Stop is the
most consequential action so it gets the tightest (rarest) false-alarm budget, Watch the loosest
because it never touches capital. You can switch any level off or give each its own α (see
Customising the governance). Thresholds are calibrated over an
evaluation horizon of max(window, 252) operating days, so the α budget reads as "false alarms per
operating year"; a shorter live window observes only the first part of the boundary and is therefore
conservative.
The corridor can be measured two ways — the same additive / compound choice as Monte Carlo, and for the same reason:
r = ΔEquity / Equity), and the corridor is the
geometric drawdown — a percentage of the running equity. This is the credible model when position
size scales with the account (compounding / fixed-fractional): there, the dollar drawdown grows
mechanically as the account grows, so a dollar corridor calibrated on a mix of small- and
large-account periods drifts upward and loses meaning, while the percentage drawdown stays
stationary. Everything is calibrated in return space — the bands, the CUSUM (in σ-units of the
return series), the state machine and the overlay — so the diagnosis is internally consistent.The basis defaults to your workspace scaling (compounding on → percent, otherwise dollars) and can
be overridden per run. The result always carries both units: a $ / % toggle on the corridor
switches the chart, the drawdown reading and the Stop threshold between them, so you can read the
credible one for your sizing and still glance at the other.
If you change the workspace scaling (or any input that moves the aggregate equity — capital, weights, sizing, weekday filters) after a diagnosis and then return to Equity Control, the module re-runs automatically with the same recipe against the new equity, and the basis re-follows the workspace (turning compounding on switches the corridor to percent) — a stale manual override is dropped so the corridor stays coherent with the equity you see. Pin a basis explicitly only if you want it to persist within an unchanged scaling.
Percent is geometric, not just rescaled dollars
The percentage corridor is not the dollar corridor divided by a constant. Under compounding the bands are re-derived from bootstrapped returns, so the drawdown is measured against the running peak equity — the only construction that stays stationary as the account compounds. If the account equity would hit zero over the baseline, compounded returns are undefined and the module asks you to use the dollar basis instead.
The bands watch the path; Page's CUSUM watches the mean. In σ-units of the baseline:
S(t) = max(0, S(t−1) + (k − x(t)) / σ), k = μ_deflated / 2
k is the classic reference value for detecting a drift from the deflated healthy mean toward zero.
The statistic charges while the strategy runs below its deflated edge and discharges when it earns.
Alarm thresholds are the bootstrap quantiles of max S under H0 at the same α budgets.
For the two levels that touch capital (Reduce, Stop) the α budget is split Bonferroni-style between the DD band and the CUSUM (each calibrated at α/2), so the joint false-alarm rate stays at or below the budget — and the empirically measured joint rate is reported, not assumed.
Two additional signals can raise the Watch state only (they never touch capital, so they are kept loose): a rolling profit factor below the sup-corrected healthy band (the α-quantile of per-path minima under H0), and time under water beyond the 90th percentile of H0 maxima.
| State | Risk applied | Enter when |
|---|---|---|
| Operating | 100% | default |
| Watch | 100% | DD > watch band, CUSUM in watch zone, PF low or TUW high |
| Reduced | reduce factor (default 50%) | DD > reduce band or CUSUM ≥ reduce threshold |
| Stopped | 0% | DD > stop band or CUSUM ≥ stop threshold |
| Phased re-entry | reduce factor | after cooldown, with signals settled |
Discipline rules, all in operating days: escalation is immediate; de-escalation steps down one
level only after 5 consecutive quieter days (hysteresis); a stop starts a cooldown of
max(10, block length) days; re-entry requires the shadow drawdown back under the watch band and
the CUSUM discharged below 25% of its stop threshold; a failed re-entry (severity rising again)
re-stops immediately; at most one restart per 40 days (anti flip-flop lockout).
No lookahead: the risk factor applied to day t is decided by the state at the end of day t − 1. Health is always evaluated on the pure (shadow) P/L, never on the overlay-reduced one.
The three levels are not fixed. What is mathematically load-bearing stays fixed; what is a design convention is yours to change.
Fixed (there is a reason). The Bonferroni ½-split between the drawdown band and the CUSUM on the
capital-touching levels (so the joint false-alarm rate stays within budget); the CUSUM reference
value k = μ_deflated / 2 (Page-optimal for a drift toward zero); and calibrating every threshold as
a bootstrap sup-quantile of a false-alarm budget. Watch/Reduce/Stop are the same three levels seen
by two detectors (the band and the CUSUM), so they always move together.
Yours to change (it is a convention). Whether a level exists at all, and the α spacing between them.
α Stop ≤ α Reduce ≤ α Watch across the enabled levels — otherwise the
bands would cross and the diagnosis is rejected with a clear message.This changes governance, not the health verdict
Turning levels off changes what the overlay does and its cost/benefit simulation, not the underlying health reading: P(edge intact), the path anomaly and the CUSUM/drawdown statistics are computed the same way regardless of which levels are armed. With Stop off there is no "wrongful stop" to count, so detection under the breakage scenarios is measured on the first throttle instead.
An on/off overlay adds expected value only if the P/L is serially dependent — if losing periods cluster, so the recent past predicts the near future. On serially independent P/L with a positive edge, any drawdown rule has negative expected value: every skipped day skips a positive-mean day. It can still be rational as insurance (paying return to shorten the tail), but that is a different purchase and the module says so.
Four tests run on the full history: Ljung-Box on P/L (memory in returns), runs test on signs
(clustering of wins/losses), loss persistence P(loss | loss) − P(loss) with a permutation
p-value, and Ljung-Box on |P/L| (volatility clustering: predictable risk, not direction). Two or
more significant mean-tests → dependence present; one, or volatility clustering only → weak;
none → absent, and the verdict tells you to read the healthy-scenario cost as an insurance premium.
Thresholds are calibrated on one batch of H0 paths; all costs and benefits are measured on independent batches (a light double bootstrap), so the overlay never gets to grade its own homework:
−μ_deflated: outright breakage. Detection rate, median
lag, median capital saved, median max-DD reduction — this is where the overlay earns its keep.The healthy/dead/inverted ledgers sit behind a Counterfactual stress tests disclosure, collapsed by default — open it when you want the full cost/benefit accounting.
The replay is a scrubbable timeline, not a static chart. It opens on the finished window — both equity lines fully drawn, the slider parked at the far right. Drag the slider back toward the left to retrace the path day by day, or press Play to rebuild it forward from the start. The always-on and with-overlay equity curves grow together and visibly diverge wherever the overlay throttles risk (shaded spans). Each governance transition — Watch, Reduce, Stop, gradual re-entry — appears on the exact day it fires, marked on the axis, and listed in the transition log beneath the player with its cause and the numbers (e.g. "Reduced — CUSUM 5.24 ≥ 5.00", or "DD beyond the Reduce band"). The log is synced with the playhead — the current event is highlighted, entries past the cursor are dimmed — and clicking a row jumps the replay to that instant. Directly below the equity chart, the CUSUM detector shares the same playhead, so you watch the statistic load toward a threshold, cross it, trigger the intervention, then discharge and re-arm. A live readout follows the cursor: current date, state, applied risk, the overlay's running effect (Δ), and the CUSUM value.
A side-by-side metrics table compares the two versions, with the better value of each row highlighted. The numbers come from the same engine as the Metrics page — net PnL, return, CAGR, max drawdown ($ and %), MAR (= CAGR / max drawdown), Sharpe, Sortino, annual volatility, win rate, profit factor — with the delta per row. A toggle switches the scope:
The highlight marks the better value only where the direction is unambiguous; for the risk-adjusted ratios it is suppressed on a losing window (see below), and volatility is never flagged (lower is not inherently better). Below the table, the interventions strip lists the replay facts — stops, re-entries, time reduced/flat, losses avoided, gains missed, protection efficiency — always referred to the monitored window regardless of the scope toggle.
Ratios on a losing window
When the window is in loss, the risk-adjusted ratios (MAR, Sharpe, Sortino) invert and become misleading: dividing a negative return by a much smaller volatility makes the ratio look worse for the version that actually lost less. In that regime read PnL, drawdown and volatility — which are unambiguous — and treat the ratios as meaningful only when the return is positive. The table shows a caveat automatically whenever either version's return is negative.
How to decide
Read three numbers together: the feasibility verdict, the healthy-scenario cost, and the inverted-edge savings. If dependence is absent and the healthy cost is material, the overlay is expensive insurance for your strategy — knowing that is exactly the point of the module.
| Parameter | Default | Range | Meaning |
|---|---|---|---|
| In-Sample end | — (required) | any session date | date where the In-Sample ends; the rest is the live window (takes precedence over Live tail) |
| Corridor basis | workspace scaling | Dollars (additive) · Percent (compound) | unit the corridor is calibrated in — dollars (fixed size) or geometric percent (compounding); result carries both |
| Live tail | — (optional) | 10 – 2000 | final operating days classified as live; alternative to In-Sample end (under Advanced) |
| Monitor from / To | full history | any session date | calendar bounds of the analysed window (To = as-of date) |
| Edge haircut | 25% | 0 – 50% | deflation of the baseline mean under H0 |
| False-alarm budget α (Stop) | 5% | 1 – 20% | wrongful-stop budget per ~operating year; anchors the default ladder (Reduce 2α, Watch 4α) |
| Governance preset | Graduated | Graduated · Kill-switch · Reduce-only · Custom | which levels are armed (see Customising the governance) |
| Levels on/off | all on | per level | disable Watch, Reduce or Stop individually; at least one must stay on |
| Reduce factor | 50% | 10 – 90% | size kept in the Reduced state (50% = halve, 25% = cut by three quarters); also the phased re-entry risk |
| Per-level α | anchored ladder | 1 – 50% each | independent α per level; requires Extended threshold control in Settings; must stay monotonic (Stop ≤ Reduce ≤ Watch) |
| Bootstrap paths | 2000 | 500 – 20000 | calibration batch (evaluation batches are half, min 500) |
| Block length | auto (Politis-White) | 2 – 500 | stationary-bootstrap expected block |
| Rolling PF window | 20 | 10 – 60 | operating days of the monitor-only PF signal (Watch level only) |
The diagnosis runs as a background job like the other simulations: the POST returns a job_id,
the client polls, and the run survives tab switches, phone lock and even a backend restart
(see API reference).