METHOD · 6 OCTOBER 2026
Guard defaults study: the method
Every result comes from Guard’s own engine (zunder-risk) through the same replay this site’s backtest uses, run natively for speed and checked against the WebAssembly build on every input. The protocol was pre-registered and committed before any replay ran.
5,940 pre-registered rule sets tried on 262 accounts (train 136, validation 63, held-out test 63); none beat the defaults out of sample, so they stay. On the held-out test the defaults cut the median max drawdown by 42 points and let 10% of entries through at full size. All four studies.
Method (as run)
- Data path:
fetchAccountHistoryandbuildReplayInputfromweb/live/src/audit.ts. Window[AS_OF − 180 d, AS_OF], AS_OF 2026-10-06 01:00 UTC. - Population (262 accounts):
- Cohort M, ordinary active traders: M1 80, M2 63, M3 71. Drawn from the leaderboard in ROI terciles, as pre-registered. 596 accounts tried; M2 and M3 stopped at the 200-try cap.
- Cohort A, long-term winners from the compliance study: 11.
- Cohort B, liquidated accounts from the compliance study: 37.
- Excluded for fewer than 50 judged entries: 10 from A, 43 from B. Market makers and accounts with too few fills were dropped at the pre-screen.
- Splits (seed 20261008, within each group): train 136, validation 63, test 63. Time split: 68 train accounts in the early window, 158 train and validation accounts in the late window (at least 20 judged entries each).
- Requests: about 148,000 weight in all, with zero HTTP 429s.
- Cohort M was fetched at 200, then 450, 500 and 900 weight in any rolling minute, as the other jobs on the IP allowed.
- Ledgers for A and B were fetched at 300 a minute.
- Cohort M’s ledgers came from the survival study’s cache.
- Trials: 5,940 grid variants and 6 ablation variants (analysis only), each run on train, validation and the time windows. The test ran exactly 3 rule sets once. The 5 deviations below count as trials too.
Bias checks
-
Native against WebAssembly: identical on all 608 inputs.
-
Unseen stop touches (test accounts; 1-second candles from the bucket; every coin of every rested stop had candles). The share of attached stops that the replay never fired, but that the price touched between the fills it saw:
- defaults (2% stop): 16% of 1,676;
- “careful” (1.5% stop): 21% of 1,893;
- “active” (1.5% stop): 23% of 1,425.
The replay misses more stop-outs at 1.5% than at 2%. Under the pre-registered rule, the 1.5% stop result is unconfirmed, and no change of the default stop distance would be recommended even if one had passed test. All results flatter Guard somewhat: in about one rested stop in six, a real Guard would have closed the position, at a loss, and then missed any recovery.
-
What-if: each account’s later decisions are kept as traded.
-
Liquidation distance: not judged (past liquidation prices are unknown), so it was not tuned.
What could still be wrong
- Small held-out sample. 63 accounts, only 2 of them long-term winners, so group figures on test are anecdotes. Train and validation give 9 winners.
- Regime. Both windows sit inside one six-month market.
- Selection of the population. B is chosen for liquidations and A for survival. M is the fair sample, but it is restricted to accounts with at least 50 judged entries, so heavy, active traders dominate.
- The replay’s assumptions. Equity is realised, stops are seen only at fills (quantified above), and liquidation distance is not judged.
- Friction. Pass-through counts only entries taken at full size. A resized entry is a real trade at a smaller size and still counts as friction here.
Deviations (each counts as a trial; none was made after a selection result was seen)
- The engine and data path moved to main’s twice, before any selection result.
- At 05:40 UTC, main’s replay left HIP-3 and outcome markets out, marked both curves at the window start and booked settlements. Cohort M was judged again with it (
study.ts refilter).
- At 05:40 UTC, main’s replay left HIP-3 and outcome markets out, marked both curves at the window start and booked settlements. Cohort M was judged again with it (
- Requests beyond the compliance cache. The current path also asks for window-start candles, funding from the first served fill and the ledger. These were fetched (A and B ledgers at 300 a minute), and cohort M’s ledgers were read from the survival study’s cache.
- Sub-window starts. A sub-window may also start from the account’s first deposits when the account was empty at its start. A start backed out from today’s value is still refused.
- A cheaper pre-screen bound during the cohort M fetch: orders that open or add to a position after the oldest order served. It can only skip accounts that cannot pass the 50-entry filter.
- Analysis added at the coordinator’s request, outside the selection: the 6 ablation variants, the per-entry reasons and the flow estimate.
- Cohort M re-judged with the new engine. M1 then had 71 passes in 173 tries, so it was fetched further until 80 passed (196 tried), as the protocol’s stopping rule requires. M2 and M3 stayed at their 200-try caps. Members are the first 80 passing in draw order.