06-reference/research

cpcv partition sizing per autoinv universe

2026-09-05·research-brief·source: deep-research·by Ray Data Co (deep-research synthesis)
autoinvcpcvvalidationbacktestinginvesting

N=6, k=2 cannot be confirmed per-surface, because the surfaces were never spec'd — but the sizing rule is closed-form, and on the only real surface it holds to a 42-bar label horizon

The question

"Confirm N=6, k=2 CPCV against each autoinv swarm surface's train/validation span — does each of the N groups hold enough bars after purge + embargo per universe, or is a different (N,k) needed?"

Carried as open follow-up #2 from [[2026-06-26-strategy-discovery-loop-architecture]], re-raised in [[2026-07-05-autoinv-cpcv-vs-walk-forward-confirm-stage]]. It gates Phase-2 coding of combinatorial_purged_cv().

What we already know (from the vault)

What the web says

Convergences and contradictions

Synthesis for RDCO

The question as posed cannot be answered, and that is the finding. There is no per-universe train/validation span to check N=6, k=2 against, because no autoinv document assigns a span to any named surface. Two surfaces are named in passing; neither has a date range, a bar count, a rebalance frequency, or a label horizon. Producing that table is a prerequisite, not a nicety, and it has been deferred three times (2026-06-01, 2026-06-26, 2026-07-05) plus generalized once (2026-08-31: "pin L_min and the reseal cadence per surface, not globally"). What follows is the decision rule to apply the moment the table exists, plus the answer for the one surface where real data does exist.

The sizing arithmetic is closed-form, and the N-dependence is weaker than the question assumes. Let T = bars in the train/validation partition, N = groups, k = test groups, G = ⌊T/N⌋ = group size, H = max label horizon in bars, h = embargo fraction, E = ⌈h·T⌉. Each of the k test blocks costs the training set H bars on its left edge (purge: training labels overlapping the test window) and E bars on its right edge (embargo). Worst case, all k blocks are non-adjacent and interior:

train_net = (N − k)·G − k·(H + E)
train_net / T = (N − k)/N − k·H/T − k·h

Read the second line carefully: the purge+embargo cost is a function of k, H and T, and is completely independent of N. Raising N does not reduce contamination loss. It raises the raw training fraction from (N−k)/N, shrinks each group, and, for k=2, raises the path count to exactly N−1. It buys nothing in sample size — because each reassembled CPCV path is full-length T regardless of N (each of the N groups appears once per path). N=6 with 1,000 bars gives you five paths of 1,000 bars, not five paths of 167. The statistical power of a path Sharpe is set by T alone.

The real failure mode is a group being smaller than its own contamination bite. If H + E > G, purge plus embargo eats an entire adjacent training group, silently collapsing the effective (N−k) and breaking the partition's stated structure. That gives the floor the public literature does not supply:

Applied to the only surface that has data, N=6, k=2 passes — up to a 42-bar label horizon, and not beyond. The SPY cache is 1,509 daily bars; the three-partition discipline reserves the final segment as sealed holdout, and at the spec's own "~2 years ago" boundary that leaves T ≈ 1,005 bars for train/validation. At N=6, k=2, h=0.01: G=167, E=11.

H (label bars) H+E G/(H+E) train_net retention N_max at 3x verdict
5 (weekly) 16 10.44 636 95.2% 20 OK
21 (monthly) 32 5.22 604 90.4% 10 OK
42 (2-month) 53 3.15 562 84.1% 6 OK, at the line
63 (quarterly) 74 2.26 520 77.8% 4 WEAK
126 (semi-annual) 137 1.22 394 59.0% 2 WEAK, near structural failure
252 (annual) 263 0.63 142 21.3% 1 FAIL — group smaller than its bite

The two named surfaces plausibly land in different cells, and this is where I stop and label an inference as an inference: cross-sectional equity momentum on a monthly rebalance implies H≈21, which is comfortably OK; a Markov vol-regime overlay's label horizon is regime dwell time, which for equity vol regimes is plausibly 60-130 bars, which is exactly the WEAK band. Neither H is written down anywhere. That is the missing input, and it is a bigger lever than (N,k).

Recommendation: keep N=6, k=2 as the default and do not raise N. Three reasons. First, the contamination loss is N-independent, so raising N does not fix the problem the question is worried about — it makes it worse by shrinking G. At T=1,005 and H=21, N=10 drops G to 100 and G/(H+E) to 3.12, right at the working floor, in exchange for nine paths instead of five. Second, the [[2026-08-31-dsr-effective-n-estimator-precommit]] ruling removed path count from GATE-1's effective-N, so extra paths no longer buy calibration; they buy a slightly better-resolved path-Sharpe dispersion and 45 backtests instead of 15. Third, raising k is strictly worse: at N=6, k=3 buys ten paths but drops the training fraction to 50% and net training bars from 604 to 405. Instead, make the label horizon the gated quantity: write combinatorial_purged_cv to compute G/(H+E) from the actual label_spans and refuse to run below 1.0, warn below 3.0. That is a five-line assertion that converts an unanswerable per-universe audit into a runtime invariant, and it degrades gracefully when the surface table finally lands.

The binding constraint is not (N,k) at all — it is T. The standard error of an annualized Sharpe on T daily bars is ≈ √(252/T). At T=1,005 that is ±0.50, so a true Sharpe of 1.0 sits in a 95% interval near [0.02, 1.98]. Reaching ±0.35 needs ~2,060 bars, about 8.2 years; the entire cached series is 6. [[2026-08-31-sealed-holdout-refresh-policy]] :42 reached this conclusion independently for the holdout. It is equally true of the train/validation partition, and it means the honest framing is: tuning (N,k) is rearranging deck chairs on a series that is too short for any partition scheme to produce a decisive Sharpe. The (6,2) default is fine. Extending the data, and pinning per-surface label horizons, is where the leverage is.

Why this is in the vault

This closes open follow-up #2 from [[2026-06-26-strategy-discovery-loop-architecture]] with a verdict rather than another deferral, and it unblocks Phase-2 coding of combinatorial_purged_cv() in autoinv/validation.py by replacing an un-runnable per-universe audit with a runtime invariant (G ≥ 3(H+E), refuse below 1.0) that the function can enforce itself. It also reframes the swarm-surface enumeration from a documentation chore into a blocking prerequisite that three prior briefs have now deferred.

Open follow-ups

Related

Sources

Vault (RDCO-authored; all N=6/k=2, embargo_frac, φ-formula, SPY bar count, three-partition and Sharpe-SE figures come from here):

Repo (read directly, 2026-09-05):

Computed for this brief (reproducible script at /tmp/cpcv_sizing.py; every table above is my arithmetic, not a cited figure):

External: