The sealed holdout is depleted by adaptivity, not by touches: roll-forward reseal with a length-scaled query budget
The question
"Pin the holdout-refresh policy for the discovery loop: touch-once-then-retire makes the sealed holdout a depleting resource — is a periodic roll-forward reseal the sanctioned replenishment, or a different policy?"
This resolves open-follow-up #4 from [[2026-06-26-strategy-discovery-loop-architecture]] and open question #4 in [[2026-05-29-strategy-pipeline-architecture-v0]]. It is a Phase-2 blocker: HoldoutManager does not exist yet, so the policy is still free to specify.
What we already know (from the vault)
- [[2026-05-29-strategy-pipeline-architecture-v0]] specs the sealed holdout as an import-level wall:
validation/holdout.pyis the only module with read access, discovery gets a feed handle that physically excludes holdout dates, and a test asserts the discovery path RAISES on any holdout bar. A candidate runs on holdout "EXACTLY ONCE" after clearing GATE 0/1/2 plus dual buy-and-hold, and "a second peek BURNS the holdout (logged, segment retired)." - The v0 spec is internally ambiguous, and that ambiguity is the actual gap. Line 94 says a second peek burns the segment; line 96 says the architecture routes discovery at large-sample surfaces "where holdout segments are long enough to amortize many candidates." Those only reconcile if "once" means once per candidate and the segment absorbs many candidates. v0 never says how many, never bounds it, and never says what a re-seal looks like. That unbounded count is the depletion the founder is asking about.
- v0 already names the exhaustion risk as "the red-team's sharpest flaw" and pre-answers it with scope discipline rather than a refresh rule: route the swarm at liquid large-sample surfaces, keep rare-state cyclicals (~5-6 trades in ~40 years) on the v2 leave-one-out harness with one thesis and one touch. That carve-out survives everything below.
- [[2026-06-01-cpcv-deflated-sharpe-autoinv-validation]] supplies the constraint that ends up dominating this decision: DSR and MinTRL are defined over T = return observations, and below roughly 30 near-independent observations the skew/kurtosis feeding the deflation term are estimation noise. MinTRL is the tool that says whether a block is long enough to conclude anything at all.
- [[2026-07-05-autoinv-cpcv-vs-walk-forward-confirm-stage]] locked the confirm stage as two CV regimes (anchored walk-forward + CPCV). That matters here because it means the holdout is not carrying the statistical inference load. CPCV supplies the OOS distribution feeding DSR/PBO. The holdout is the final arbiter over a pre-registered binary bar, which is a much cheaper query and is the lever this policy exploits.
- Verified against the repo, not the notes (2026-08-31):
/Users/ray/Projects/automated-investing/autoinv/validation.pystill wraps sklearnTimeSeriesSplitonly.grep -rniE "holdout|sealed"overautoinv/,tests/,scripts/returns two hits, both prose admissions inreport.pyandPHASE1-SPINE.mdthat there is "no multiple-testing penalty, no sealed holdout." The single cached series isSPY_2018-01-01_2023-12-31— about 1,500 daily bars total.
What the web says
- Dwork, Feldman, Hardt, Pitassi, Reingold and Roth (Science 2015) established that a holdout can be reused if the analyst is only ever answered through a differentially-private mechanism. Thresholdout adds Laplace noise to both a threshold and the reported value; when the noisy training answer is close to the noisy holdout answer it returns the training answer (leaking nothing), and only when they diverge beyond the threshold does it spend budget and return a noisy holdout answer. A budget B caps the number of divergence events; once B is exhausted the mechanism stops answering from the holdout (arXiv 1506.02629; Science 10.1126/science.aaa9375, paywalled).
- The headline claim is a scaling result, not a free lunch. A naive holdout safely answers on the order of n non-adaptive queries; Thresholdout's claim is that the number of adaptively chosen queries answerable within tolerance grows exponentially in holdout size n. The guarantee is asymptotic in n, and the paper is explicit that the mechanism reverts to training answers on budget exhaustion (arXiv 1506.02629). I did not verify the exact constants against the primary; treat the exponent as directional.
- Blum and Hardt's Ladder (ICML 2015) is the cheaper cousin and the better fit for a pass/fail gate. It answers a leaderboard query by revealing a new score only when the submission improves on the current best by more than a step size eta, and otherwise re-reports the old best. Coarsening the answer is what limits leakage: the analyst cannot read fine-grained gradient off the holdout, so leaderboard error grows roughly logarithmically in the number of submissions k rather than polynomially (arXiv 1502.04585; PMLR v37). No noise calibration and no privacy accounting required.
- The empirical record on holdout reuse is milder than the theory. A meta-analysis comparing public-to-private leaderboard movement across 100+ ML competitions found little evidence of substantial adaptive overfitting, and reported that multiclass problems are notably more robust to test-set reuse than binary ones (Roelofs et al., NeurIPS 2019, abstract-level only). This cuts against maximal paranoia, but the competitions studied had holdout sets orders of magnitude larger than 300 market bars, and a competition's private set is a fresh iid split rather than a non-stationary future.
- The quant canon's answer is the opposite move: stop relying on the holdout at all. Bailey and López de Prado hold that "standard statistical techniques designed to prevent regression overfitting, such as hold-out, tend to be unreliable and inaccurate in the context of investment backtests," and route the burden onto trial-count accounting instead. The prescription is to keep a cumulative ledger of every backtest run against the data so PBO can be estimated and the Sharpe properly deflated (SSRN 2460551; SSRN 2326253; LBL mirror).
- Walk-forward practice quietly already does roll-forward reseal. Anchored walk-forward advances the test window as data accrues and never re-tests a window twice, which is the same shape as a periodic reseal with a one-touch-per-block rule. The known weakness carries over: consecutive windows are one historical path, not independent draws (ScienceDirect controlled comparison, abstract-level only).
Convergences and contradictions
- Convergence on the diagnosis. Adaptive-data-analysis theory and the quant canon agree the resource being consumed is not "touches" but information fed back into candidate selection. A holdout query whose result changes what the next sweep generates costs far more than one whose result only routes a finished candidate to deploy-or-retire. Both literatures converge on the same two mitigations: coarsen the answer, and keep an honest cumulative count.
- Contradiction on the remedy, and the vault sides with the canon. The ML line says a holdout can be made reusable via noise; López de Prado says holdout is the wrong instrument for backtests regardless. The reconciliation is sample size. Thresholdout's guarantee is asymptotic in n and assumes an exchangeable sample; RDCO's holdout is deliberately the most recent market regime, which is neither iid nor exchangeable with the training distribution. The mechanism's own noise is O(1/sqrt(n)) scale, which at n around 300 daily bars is comparable to the entire effect being measured. The guarantee does not transfer. The design pattern does.
- The contradiction the arithmetic exposes, which outranks the whole reuse debate. Under the standard iid approximation (Lo 2002), the standard error of an annualized Sharpe estimated on T daily bars is about sqrt(252/T). On a 300-bar holdout that is roughly +/-0.92, so a true Sharpe of 1.0 lands in a 95% interval near [-0.8, 2.8]. Reaching a +/-0.35 standard error needs roughly 2,060 daily bars, about 8 years. RDCO's entire cached series is 6 years. The binding constraint today is not query leakage but raw estimator noise: the current holdout is too short for even one touch to yield a decisive Sharpe. That is my own arithmetic under an iid assumption financial returns violate, so treat it as an optimistic floor.
Synthesis for RDCO
Recommended policy: periodic roll-forward reseal, calendar-triggered, with a block-length-scaled query budget and a coarsened pass/fail answer. Not strict touch-once, and not Thresholdout. Call it reseal-with-budget. The critical property is that strict touch-once-then-retire is the B=1 corner of this policy rather than a rival to it, so RDCO can adopt the general rule now and it will correctly behave as touch-once for as long as the data is thin. Nothing about today's behavior changes; what changes is that the policy has a defined path out of depletion instead of a cliff.
Operationally, four parameters. (1) Reseal trigger: calendar, never exhaustion. The reseal date is written into the ledger at pipeline epoch, before any result exists. Exhaustion-triggered reseal is the failure mode to avoid: it lets a spent budget become the reason to unlock fresh data, which turns the reseal itself into a selection decision made with knowledge of the results. Default cadence is annual on the daily-bar surface, on the first trading day after each 12-month anniversary of the epoch. (2) How much new data a reseal requires: a reseal appends only genuinely newly-accrued bars and never re-labels previously-trained data as holdout. Minimum accrual increment is 252 new daily bars; below that no reseal fires and the pipeline reports "no holdout capacity" rather than sealing a stub. A block becomes queryable only once it reaches L_min, recommended at 1,000 daily bars (about 4 years), on MinTRL grounds. The compensating benefit is real: at reseal the retired block is promoted into TRAIN, since it is already spent for validation purposes, so the training surface grows every cycle. (3) Query budget against a sealed block: pre-registered as B = min(5, floor(L_block / 252)), one unit per candidate evaluation, where a re-run of a modified candidate consumes a fresh unit. Budget exhaustion retires the block early and does not trigger a reseal. On the current 6-year SPY surface this yields B=1, which is exactly today's touch-once rule, derived rather than asserted. (4) The answer is coarsened, Ladder-style: the holdout returns PASS/FAIL against the pre-registered three-bar wet-rehearsal criterion plus an interval, and never a rankable score. Candidates are never sorted, tuned, or shortlisted on holdout output. This is the single highest-value cheap borrowing from the literature, because a binary verdict leaks a fraction of what a Sharpe point estimate leaks and it is all the go/no-go bar ever needed.
What gets logged: an append-only holdout ledger, one row per query, that is auditable without the code. Per block: block_id, start and end dates, length in bars, seal timestamp, and the content hash of the sealed data snapshot. Per query: sequence number, remaining budget, candidate_id and full param vector, the cumulative trial-N at query time, the verbatim pre-registered bar with its pre-registration timestamp (which the ledger must assert precedes the query timestamp), the PASS/FAIL verdict plus interval, raw SR and DSR recorded for audit but flagged not-for-ranking, the disposition (DEPLOYED or PERMANENTLY-RETIRED, where a FAIL blacklists the candidate and its param neighborhood from re-entering the sweep), and an explicit adaptivity flag defaulting to NO. Per reseal: date, bars added, and the block promoted to TRAIN. Two ledger invariants matter more than the fields. First, cumulative N spans blocks: a candidate's DSR bar must account for every holdout query the project has ever made, not per-block, or a reseal silently launders the multiple-testing discount. Second, the PERMANENTLY-RETIRED disposition is what keeps queries near-non-adaptive, because a FAIL that cannot re-enter the sweep cannot steer the next generation.
Honest costs of each option. (a) Strict touch-once-then-retire buys the only regime where a reported interval means what it says, and costs throughput: roughly one verdict per data-accrual epoch, which is incompatible with a swarm emitting candidates continuously. It is right today and wrong at scale. (b) Periodic roll-forward reseal costs three things that must be priced, not waved at. Each reseal opens a new multiple-testing surface, so a candidate failing block 1 and passing block 2 has been tested twice and only a cross-block cumulative N keeps that honest. Consecutive blocks are one historical path rather than independent draws, so "passed three blocks" is materially weaker than three independent trials. And roll-forward always seals the most recent regime, structurally over-weighting current conditions. (c) Thresholdout or a noised DP budget costs calibration complexity and buys a guarantee that does not transfer at RDCO's sample size, for the reasons above. Do not implement it, and specifically do not cite its guarantee in any report. Borrow only its two design patterns, both already in the recommendation: coarsen the answer, and hard-stop on a pre-registered budget.
The replenishment nobody asked about, which is faster than waiting. Time is the most expensive axis to buy holdout capacity on; RDCO accrues 252 bars per year and cannot accelerate that. Cross-sectional sealing is cheaper. Sealing a set of symbols, sectors, or an adjacent market never touched in training creates additional holdout capacity immediately, without waiting for the calendar. The honest discount is that equities co-move hard within a regime, so effective independence across a sealed symbol set is far below its cardinality, and the DSR ledger must treat a cross-sectional block as substantially fewer than one fresh trial per symbol. Even discounted, it beats a one-year wait, and it is the concrete thing to spec next.
Why this is in the vault
This closes open-follow-up #4 of [[2026-06-26-strategy-discovery-loop-architecture]] and open question #4 of [[2026-05-29-strategy-pipeline-architecture-v0]] with a buildable spec, and it resolves a live internal contradiction in the v0 architecture doc, where line 94 ("a second peek BURNS the holdout") and line 96 ("segments long enough to amortize many candidates") cannot both be true without a bounded per-block query budget that v0 never defines. It is a Phase-2 build input for validation/holdout.py in /Users/ray/Projects/automated-investing, which as of 2026-08-31 contains no holdout code at all, so the parameters here can be pre-registered before any result exists rather than retrofitted after.
Open follow-ups
- Pin L_min and the reseal cadence per surface, not globally. The 1,000-bar / annual defaults were derived for a daily-bar equity surface; a weekly-bar cyclical surface and an intraday surface need their own numbers, and the intraday case may accrue enough bars to justify quarterly reseal.
- Decide the cross-block DSR accounting rule concretely. Does a candidate re-tested on a later block count as N+1 trials, or is a correlation-discounted increment applied because consecutive blocks share a regime? This changes every DSR the pipeline reports after the first reseal.
- Spec cross-sectional sealing: which symbols or sectors get held out, how the effective-independence discount is computed for the DSR ledger, and whether an adjacent market (international equity, rates) is a legitimately independent holdout surface or an unpriced regime bet.
- Verify the Lo (2002) Sharpe standard-error arithmetic against autocorrelated, fat-tailed daily returns rather than the iid approximation used here. The iid number is an optimistic floor and the true L_min may be materially longer than 1,000 bars.
- Read the Ladder primary to fix the exact error bound and the step-size eta selection rule, and read "Climbing a shaky ladder: Better adaptive risk estimation" (the follow-on work), which was surfaced but not read for this brief.
- Decide whether a holdout FAIL should blacklist only the candidate or its whole param neighborhood, and how "neighborhood" is defined in param space. The recommendation assumes neighborhood-level blacklisting to keep queries non-adaptive, which is stricter than v0 and costs search breadth.
Related
- [[2026-06-26-strategy-discovery-loop-architecture]] — the parent brief whose open-follow-up #4 this closes; defines the five-stage loop and the single sealed-holdout arbiter
- [[2026-05-29-strategy-pipeline-architecture-v0]] — the v0 architecture whose sealed-holdout section this brief disambiguates and bounds
- [[2026-06-01-cpcv-deflated-sharpe-autoinv-validation]] — DSR/MinTRL formulas and the "T = return observations" caveat that sets L_min
- [[2026-07-05-autoinv-cpcv-vs-walk-forward-confirm-stage]] — the two-CV-regime confirm stage that lets the holdout carry a binary verdict instead of the inference load
- [[2026-06-01-ensemble-admission-correlation-threshold]] — ensembles are re-scored on the holdout as their own candidates, so they consume query budget under this policy
- [[2026-05-31-ensembles-systematic-trading-overfitting]] — the "you can ensemble overfit signals into a confidently overfit ensemble" finding that motivates the gates-not-reports posture
Sources
Vault:
~/rdco-vault/06-reference/research/2026-06-26-strategy-discovery-loop-architecture.md— parent brief; five-stage control flow, open-follow-up #4~/rdco-vault/01-projects/investing/2026-05-29-strategy-pipeline-architecture-v0.md— sealed-holdout spec (lines 94, 96), Phase-2 build plan, open question #4~/rdco-vault/06-reference/research/2026-06-01-cpcv-deflated-sharpe-autoinv-validation.md— DSR / MinTRL / CPCV mechanics and the 30-observation floor~/rdco-vault/06-reference/research/2026-07-05-autoinv-cpcv-vs-walk-forward-confirm-stage.md— two-regime confirm stage decision~/rdco-vault/06-reference/research/2026-06-01-ensemble-admission-correlation-threshold.md— ensemble admission gate and holdout re-scoring~/rdco-vault/06-reference/research/2026-05-31-ensembles-systematic-trading-overfitting.md— ensembles do not cure spurious correlation
Repo (verified 2026-08-31):
/Users/ray/Projects/automated-investing/autoinv/validation.py—TimeSeriesSplitonly; no holdout, CPCV, purge, embargo, DSR or PBO/Users/ray/Projects/automated-investing/PHASE1-SPINE.mdandautoinv/report.py— both state "no multiple-testing penalty, sealed holdout, or ..." as a known Phase-1 gap/Users/ray/Projects/automated-investing/data_cache/SPY_2018-01-01_2023-12-31_*.csv— the only cached series, about 1,500 daily bars
Web:
- Dwork, Feldman, Hardt, Pitassi, Reingold, Roth — "Generalization in Adaptive Data Analysis and Holdout Reuse" (NIPS 2015) — https://arxiv.org/pdf/1506.02629 (Thresholdout mechanics, budget B, exponential-in-n query claim)
- Same authors — "The reusable holdout: Preserving validity in adaptive data analysis," Science 2015 — https://www.science.org/doi/10.1126/science.aaa9375 (PAYWALLED, not fetched; the arXiv companion above was used instead)
- Blum & Hardt — "The Ladder: A Reliable Leaderboard for Machine Learning Competitions" (ICML 2015) — https://arxiv.org/pdf/1502.04585 and https://proceedings.mlr.press/v37/blum15.pdf (coarsened answers limit leakage)
- Roelofs et al. — "A meta-analysis of overfitting in machine learning" (NeurIPS 2019) — https://dl.acm.org/doi/10.5555/3454287.3455110 (abstract-level only; little observed leaderboard overfitting across 100+ competitions)
- Bailey & López de Prado — "The Deflated Sharpe Ratio" — https://papers.ssrn.com/sol3/papers.cfm?abstract_id=2460551
- Bailey, Borwein, López de Prado, Zhu — "The Probability of Backtest Overfitting" — https://papers.ssrn.com/sol3/papers.cfm?abstract_id=2326253 and https://sdm.lbl.gov/oapapers/ssrn-id2507040-bailey.pdf ("hold-out tends to be unreliable... in the context of investment backtests"; keep a cumulative trial ledger)
- Backtest-overfitting comparison in a synthetic controlled environment — https://www.sciencedirect.com/science/article/abs/pii/S0950705124011110 (abstract-level only, paywalled body; walk-forward tests a single historical path)