06-reference/research

sealed holdout refresh policy

2026-08-31·research-brief·source: deep-research·by Ray Data Co (deep-research synthesis)
sealed-holdoutadaptive-data-analysisbacktest-overfittingstrategy-discoverythresholdout

The sealed holdout is depleted by adaptivity, not by touches: roll-forward reseal with a length-scaled query budget

The question

"Pin the holdout-refresh policy for the discovery loop: touch-once-then-retire makes the sealed holdout a depleting resource — is a periodic roll-forward reseal the sanctioned replenishment, or a different policy?"

This resolves open-follow-up #4 from [[2026-06-26-strategy-discovery-loop-architecture]] and open question #4 in [[2026-05-29-strategy-pipeline-architecture-v0]]. It is a Phase-2 blocker: HoldoutManager does not exist yet, so the policy is still free to specify.

What we already know (from the vault)

What the web says

Convergences and contradictions

Synthesis for RDCO

Recommended policy: periodic roll-forward reseal, calendar-triggered, with a block-length-scaled query budget and a coarsened pass/fail answer. Not strict touch-once, and not Thresholdout. Call it reseal-with-budget. The critical property is that strict touch-once-then-retire is the B=1 corner of this policy rather than a rival to it, so RDCO can adopt the general rule now and it will correctly behave as touch-once for as long as the data is thin. Nothing about today's behavior changes; what changes is that the policy has a defined path out of depletion instead of a cliff.

Operationally, four parameters. (1) Reseal trigger: calendar, never exhaustion. The reseal date is written into the ledger at pipeline epoch, before any result exists. Exhaustion-triggered reseal is the failure mode to avoid: it lets a spent budget become the reason to unlock fresh data, which turns the reseal itself into a selection decision made with knowledge of the results. Default cadence is annual on the daily-bar surface, on the first trading day after each 12-month anniversary of the epoch. (2) How much new data a reseal requires: a reseal appends only genuinely newly-accrued bars and never re-labels previously-trained data as holdout. Minimum accrual increment is 252 new daily bars; below that no reseal fires and the pipeline reports "no holdout capacity" rather than sealing a stub. A block becomes queryable only once it reaches L_min, recommended at 1,000 daily bars (about 4 years), on MinTRL grounds. The compensating benefit is real: at reseal the retired block is promoted into TRAIN, since it is already spent for validation purposes, so the training surface grows every cycle. (3) Query budget against a sealed block: pre-registered as B = min(5, floor(L_block / 252)), one unit per candidate evaluation, where a re-run of a modified candidate consumes a fresh unit. Budget exhaustion retires the block early and does not trigger a reseal. On the current 6-year SPY surface this yields B=1, which is exactly today's touch-once rule, derived rather than asserted. (4) The answer is coarsened, Ladder-style: the holdout returns PASS/FAIL against the pre-registered three-bar wet-rehearsal criterion plus an interval, and never a rankable score. Candidates are never sorted, tuned, or shortlisted on holdout output. This is the single highest-value cheap borrowing from the literature, because a binary verdict leaks a fraction of what a Sharpe point estimate leaks and it is all the go/no-go bar ever needed.

What gets logged: an append-only holdout ledger, one row per query, that is auditable without the code. Per block: block_id, start and end dates, length in bars, seal timestamp, and the content hash of the sealed data snapshot. Per query: sequence number, remaining budget, candidate_id and full param vector, the cumulative trial-N at query time, the verbatim pre-registered bar with its pre-registration timestamp (which the ledger must assert precedes the query timestamp), the PASS/FAIL verdict plus interval, raw SR and DSR recorded for audit but flagged not-for-ranking, the disposition (DEPLOYED or PERMANENTLY-RETIRED, where a FAIL blacklists the candidate and its param neighborhood from re-entering the sweep), and an explicit adaptivity flag defaulting to NO. Per reseal: date, bars added, and the block promoted to TRAIN. Two ledger invariants matter more than the fields. First, cumulative N spans blocks: a candidate's DSR bar must account for every holdout query the project has ever made, not per-block, or a reseal silently launders the multiple-testing discount. Second, the PERMANENTLY-RETIRED disposition is what keeps queries near-non-adaptive, because a FAIL that cannot re-enter the sweep cannot steer the next generation.

Honest costs of each option. (a) Strict touch-once-then-retire buys the only regime where a reported interval means what it says, and costs throughput: roughly one verdict per data-accrual epoch, which is incompatible with a swarm emitting candidates continuously. It is right today and wrong at scale. (b) Periodic roll-forward reseal costs three things that must be priced, not waved at. Each reseal opens a new multiple-testing surface, so a candidate failing block 1 and passing block 2 has been tested twice and only a cross-block cumulative N keeps that honest. Consecutive blocks are one historical path rather than independent draws, so "passed three blocks" is materially weaker than three independent trials. And roll-forward always seals the most recent regime, structurally over-weighting current conditions. (c) Thresholdout or a noised DP budget costs calibration complexity and buys a guarantee that does not transfer at RDCO's sample size, for the reasons above. Do not implement it, and specifically do not cite its guarantee in any report. Borrow only its two design patterns, both already in the recommendation: coarsen the answer, and hard-stop on a pre-registered budget.

The replenishment nobody asked about, which is faster than waiting. Time is the most expensive axis to buy holdout capacity on; RDCO accrues 252 bars per year and cannot accelerate that. Cross-sectional sealing is cheaper. Sealing a set of symbols, sectors, or an adjacent market never touched in training creates additional holdout capacity immediately, without waiting for the calendar. The honest discount is that equities co-move hard within a regime, so effective independence across a sealed symbol set is far below its cardinality, and the DSR ledger must treat a cross-sectional block as substantially fewer than one fresh trial per symbol. Even discounted, it beats a one-year wait, and it is the concrete thing to spec next.

Why this is in the vault

This closes open-follow-up #4 of [[2026-06-26-strategy-discovery-loop-architecture]] and open question #4 of [[2026-05-29-strategy-pipeline-architecture-v0]] with a buildable spec, and it resolves a live internal contradiction in the v0 architecture doc, where line 94 ("a second peek BURNS the holdout") and line 96 ("segments long enough to amortize many candidates") cannot both be true without a bounded per-block query budget that v0 never defines. It is a Phase-2 build input for validation/holdout.py in /Users/ray/Projects/automated-investing, which as of 2026-08-31 contains no holdout code at all, so the parameters here can be pre-registered before any result exists rather than retrofitted after.

Open follow-ups

Related

Sources

Vault:

Repo (verified 2026-08-31):

Web: