N=6, k=2 cannot be confirmed per-surface, because the surfaces were never spec'd — but the sizing rule is closed-form, and on the only real surface it holds to a 42-bar label horizon
The question
"Confirm N=6, k=2 CPCV against each autoinv swarm surface's train/validation span — does each of the N groups hold enough bars after purge + embargo per universe, or is a different (N,k) needed?"
Carried as open follow-up #2 from [[2026-06-26-strategy-discovery-loop-architecture]], re-raised in [[2026-07-05-autoinv-cpcv-vs-walk-forward-confirm-stage]]. It gates Phase-2 coding of combinatorial_purged_cv().
What we already know (from the vault)
- The swarm surfaces are not enumerated with spans. Anywhere. The phrase "swarm surfaces" appears four times across the corpus and is never once accompanied by a start date, end date, bar count, rebalance frequency, or label horizon. The closest thing to an enumeration is a two-item parenthetical repeated verbatim across four docs: "large-sample liquid surfaces (cross-sectional equity momentum, the SPY vol-regime overlay)" ([[2026-05-29-strategy-pipeline-architecture-v0]]
:20,:149; [[2026-06-01-cpcv-deflated-sharpe-autoinv-validation]]:94; [[2026-07-05-autoinv-cpcv-vs-walk-forward-confirm-stage]]:22). Verified by grep over both the vault and the repo, plus a semantic qmd sweep. This question has been open since 2026-06-01 and has never been answered because its input was never produced. - The parameters under test are fully specified. N=6, k=2 → C(6,2)=15 splits, φ[6,2]=(2/6)·15=5 paths; each group appears in exactly 5 test splits ([[2026-06-01-cpcv-deflated-sharpe-autoinv-validation]]
:32). Proposed signature:combinatorial_purged_cv(n_obs, label_spans, n_groups=6, k_test=2, embargo_frac=0.01)([[2026-07-05-autoinv-cpcv-vs-walk-forward-confirm-stage]]:46). Order is embargo first, then purge; de Prado's embargo default is on the order of h ≈ 1% of total observations ([[2026-06-01-cpcv-deflated-sharpe-autoinv-validation]]:34). label_spansis itself unspecified. It is an input with no stated value for any surface, and specifying it is its own open item ([[2026-07-05-autoinv-cpcv-vs-walk-forward-confirm-stage]]:55). This matters more than (N,k): the arithmetic below shows the label horizon, not N, is what actually breaks the partition.- The only real data in the project is one cached series.
data_cache/SPY_2018-01-01_2023-12-31_5f6c0c96.csv, 1,509 daily bars (2018-01-02 → 2023-12-29). No other cached series exists;data.py/feed.pycarry no default start/end dates. Corroborated by the vault's own repo audit in [[2026-08-31-sealed-holdout-refresh-policy]]:27("about 1,500 daily bars total") and the "current 6-year SPY surface" language at:48. - None of this is built.
autoinv/validation.pyis 98 lines wrapping sklearnTimeSeriesSplitplus a four-boolean bias-audit dataclass. Zero occurrences of purge, embargo, CPCV, DSR, PBO, or MinTRL in the repo.PHASE1-SPINE.md:55says so plainly: "No multiple-testing penalty, sealed holdout, or walk-forward yet — those are Phase 2+." - The partition shape is specified only relatively. TRAIN+VALIDATION = "e.g. inception →
2 years ago"; SEALED HOLDOUT = "final segment + the rarest stress episodes" ([[2026-05-29-strategy-pipeline-architecture-v0]]".:92-94). Note the "e.g." and the "
What the web says
- skfolio's
CombinatorialPurgedCVships defaults ofn_folds=10, n_test_folds=8, purged_size=0, embargo_size=0with only two hard constraints stated: "n_folds ... Must be at least 3" and "n_test_folds ... Must be at least 2. For only one test fold, usesklearn.model_validation.KFold" (skfolio docs). Two things worth flagging: (a) purge and embargo are absolute observation counts, not fractions, unlike de Prado'sh; (b) the library's default N=10/k=8 leaves a 20% training fraction, which is a path-maximizing regime, not a fit-quality regime. - No mainstream implementation or reference publishes a minimum-fold-size rule. I checked skfolio's docs and QuantInsti's CPCV walkthrough directly; neither states a minimum observations-per-group requirement or warns that purge+embargo can consume an entire adjacent fold (skfolio; QuantInsti). This is a genuine gap in the public literature, not a gap in my search. The floor proposed below is RDCO-derived, not borrowed.
- QuantInsti's worked example is exactly N=6, k=2 — reported as "n_sim: 15" and "n_paths: 5" — run on a dataset of 6,685 observations (QuantInsti). That is ~4.4x the length of RDCO's entire SPY cache and ~6.6x its usable train/validation partition. The canonical (6,2) demo is a large-sample demo.
- Embargo sizing guidance is qualitative and keys to feature lookback, not to a universal fraction. QuantInsti's only concrete instance is removing "the first 63 or so days from fold (5)" when using a 63-day lookback indicator. That is a lookback-matched embargo, which for a long-lookback feature is far larger than de Prado's ~1% default (QuantInsti).
- Search results surfaced a claim that path counts below ~100 produce "unstable and noisy" performance distributions (attributed to quantbeckman). I did not fetch that source and am not treating the number as established. If it holds, 5 paths is far below it — but the claim conflicts with de Prado's own canonical (6,2) worked example, so it needs its own verification before it moves any decision.
Convergences and contradictions
- Convergence: N=6, k=2 is genuinely the canonical default. RDCO's spec, de Prado's text, and QuantInsti's tutorial all use it. Nothing in the external literature argues against it on principle.
- Contradiction that matters: every published (6,2) demonstration runs on thousands of observations (QuantInsti: 6,685). RDCO's train/validation partition on its only real surface is ~1,005 bars (derived below). The default is being inherited from a data regime RDCO is not in.
- Contradiction inside RDCO's own docs: [[2026-07-05-autoinv-cpcv-vs-walk-forward-confirm-stage]]
:46pre-registers effective-N and V[SR_n] as coming from CPCV path Sharpes; [[2026-08-31-dsr-effective-n-estimator-precommit]]:56explicitly reverses this ("GATE-1's N isN_eff... It is NOT the CPCV path count"). The August ruling is later and governs. Consequence for this brief: path count is no longer load-bearing for GATE-1, which materially weakens the case for raising N to buy paths.
Synthesis for RDCO
The question as posed cannot be answered, and that is the finding. There is no per-universe train/validation span to check N=6, k=2 against, because no autoinv document assigns a span to any named surface. Two surfaces are named in passing; neither has a date range, a bar count, a rebalance frequency, or a label horizon. Producing that table is a prerequisite, not a nicety, and it has been deferred three times (2026-06-01, 2026-06-26, 2026-07-05) plus generalized once (2026-08-31: "pin L_min and the reseal cadence per surface, not globally"). What follows is the decision rule to apply the moment the table exists, plus the answer for the one surface where real data does exist.
The sizing arithmetic is closed-form, and the N-dependence is weaker than the question assumes. Let T = bars in the train/validation partition, N = groups, k = test groups, G = ⌊T/N⌋ = group size, H = max label horizon in bars, h = embargo fraction, E = ⌈h·T⌉. Each of the k test blocks costs the training set H bars on its left edge (purge: training labels overlapping the test window) and E bars on its right edge (embargo). Worst case, all k blocks are non-adjacent and interior:
train_net = (N − k)·G − k·(H + E)
train_net / T = (N − k)/N − k·H/T − k·h
Read the second line carefully: the purge+embargo cost is a function of k, H and T, and is completely independent of N. Raising N does not reduce contamination loss. It raises the raw training fraction from (N−k)/N, shrinks each group, and, for k=2, raises the path count to exactly N−1. It buys nothing in sample size — because each reassembled CPCV path is full-length T regardless of N (each of the N groups appears once per path). N=6 with 1,000 bars gives you five paths of 1,000 bars, not five paths of 167. The statistical power of a path Sharpe is set by T alone.
The real failure mode is a group being smaller than its own contamination bite. If H + E > G, purge plus embargo eats an entire adjacent training group, silently collapsing the effective (N−k) and breaking the partition's stated structure. That gives the floor the public literature does not supply:
- Hard floor:
G > H + E. Below this the split is structurally invalid, not merely noisy. - Working floor (RDCO-proposed):
G ≥ 3·(H + E), so no more than a third of any adjacent group is consumed. Equivalently, N_max = ⌊T / (3·(H + E))⌋. - Sanity floor:
G ≥ 60daily bars, so a group is at least a quarter-length contiguous regime chunk rather than a slice of one market state.
Applied to the only surface that has data, N=6, k=2 passes — up to a 42-bar label horizon, and not beyond. The SPY cache is 1,509 daily bars; the three-partition discipline reserves the final segment as sealed holdout, and at the spec's own "~2 years ago" boundary that leaves T ≈ 1,005 bars for train/validation. At N=6, k=2, h=0.01: G=167, E=11.
| H (label bars) | H+E | G/(H+E) | train_net | retention | N_max at 3x | verdict |
|---|---|---|---|---|---|---|
| 5 (weekly) | 16 | 10.44 | 636 | 95.2% | 20 | OK |
| 21 (monthly) | 32 | 5.22 | 604 | 90.4% | 10 | OK |
| 42 (2-month) | 53 | 3.15 | 562 | 84.1% | 6 | OK, at the line |
| 63 (quarterly) | 74 | 2.26 | 520 | 77.8% | 4 | WEAK |
| 126 (semi-annual) | 137 | 1.22 | 394 | 59.0% | 2 | WEAK, near structural failure |
| 252 (annual) | 263 | 0.63 | 142 | 21.3% | 1 | FAIL — group smaller than its bite |
The two named surfaces plausibly land in different cells, and this is where I stop and label an inference as an inference: cross-sectional equity momentum on a monthly rebalance implies H≈21, which is comfortably OK; a Markov vol-regime overlay's label horizon is regime dwell time, which for equity vol regimes is plausibly 60-130 bars, which is exactly the WEAK band. Neither H is written down anywhere. That is the missing input, and it is a bigger lever than (N,k).
Recommendation: keep N=6, k=2 as the default and do not raise N. Three reasons. First, the contamination loss is N-independent, so raising N does not fix the problem the question is worried about — it makes it worse by shrinking G. At T=1,005 and H=21, N=10 drops G to 100 and G/(H+E) to 3.12, right at the working floor, in exchange for nine paths instead of five. Second, the [[2026-08-31-dsr-effective-n-estimator-precommit]] ruling removed path count from GATE-1's effective-N, so extra paths no longer buy calibration; they buy a slightly better-resolved path-Sharpe dispersion and 45 backtests instead of 15. Third, raising k is strictly worse: at N=6, k=3 buys ten paths but drops the training fraction to 50% and net training bars from 604 to 405. Instead, make the label horizon the gated quantity: write combinatorial_purged_cv to compute G/(H+E) from the actual label_spans and refuse to run below 1.0, warn below 3.0. That is a five-line assertion that converts an unanswerable per-universe audit into a runtime invariant, and it degrades gracefully when the surface table finally lands.
The binding constraint is not (N,k) at all — it is T. The standard error of an annualized Sharpe on T daily bars is ≈ √(252/T). At T=1,005 that is ±0.50, so a true Sharpe of 1.0 sits in a 95% interval near [0.02, 1.98]. Reaching ±0.35 needs ~2,060 bars, about 8.2 years; the entire cached series is 6. [[2026-08-31-sealed-holdout-refresh-policy]] :42 reached this conclusion independently for the holdout. It is equally true of the train/validation partition, and it means the honest framing is: tuning (N,k) is rearranging deck chairs on a series that is too short for any partition scheme to produce a decisive Sharpe. The (6,2) default is fine. Extending the data, and pinning per-surface label horizons, is where the leverage is.
Why this is in the vault
This closes open follow-up #2 from [[2026-06-26-strategy-discovery-loop-architecture]] with a verdict rather than another deferral, and it unblocks Phase-2 coding of combinatorial_purged_cv() in autoinv/validation.py by replacing an un-runnable per-universe audit with a runtime invariant (G ≥ 3(H+E), refuse below 1.0) that the function can enforce itself. It also reframes the swarm-surface enumeration from a documentation chore into a blocking prerequisite that three prior briefs have now deferred.
Open follow-ups
- Produce the swarm-surface table. This is the missing input, and it blocks four separate open items. For each surface the swarm will actually run on: name, universe construction rule, bar frequency, inception date, train/validation end boundary, sealed-holdout boundary, resulting bar count, rebalance frequency, and max label horizon in bars. Two surfaces are currently named in passing ("cross-sectional equity momentum", "SPY vol-regime overlay"); confirm whether that is the closed set.
- Pin
label_spansper surface, and per strategy family within a surface. The arithmetic above shows H is the dominant term — it moves N=6,k=2 from 95% training retention to structural failure across a plausible H range.label_spansis currently an unvalued input parameter. Specifically: what is the label horizon of a Markov vol-regime strategy, given that regime dwell time is stochastic and not known at fit time? - Decide whether the embargo is a fraction of T (de Prado's h≈1%) or lookback-matched (QuantInsti's practice). These diverge badly on short series: at T=1,005, h=0.01 gives E=11 bars, while a 63-day lookback indicator would demand E=63. The spec currently hardcodes
embargo_frac=0.01, which may be under-embargoing any feature with a lookback longer than ~11 bars. - Verify or discard the "path counts below ~100 are unstable" claim. Surfaced in search but not fetched at source, and it conflicts with de Prado's own (6,2) worked example. If it holds it argues for N≫6; if it does not, the 5-path default stands unchallenged. This is cheap to resolve and currently sits as the only external argument against the recommendation above.
- Extend the cached data before tuning the partition. At 1,509 bars total the series cannot support a decisive Sharpe under any (N,k). What would it cost to pull SPY back to 1993 (~8,300 bars) and to build the cross-sectional momentum panel survivorship-free? Note that [[2026-06-01-cpcv-deflated-sharpe-autoinv-validation]]
:92already gates DSR reporting on trade counts; the same logic should gate whether CPCV is worth running at all on a given surface. - Reconcile whether the sealed-holdout boundary is "~2 years ago" or something surface-specific. T≈1,005 above is my derivation from the spec's illustrative "e.g. inception → ~2 years ago"; the actual boundary has never been pinned, and every number in the table above scales with it.
Related
- [[2026-06-26-strategy-discovery-loop-architecture]]
- [[2026-06-01-cpcv-deflated-sharpe-autoinv-validation]]
- [[2026-07-05-autoinv-cpcv-vs-walk-forward-confirm-stage]]
- [[2026-05-29-strategy-pipeline-architecture-v0]]
- [[2026-08-31-sealed-holdout-refresh-policy]]
- [[2026-08-31-dsr-effective-n-estimator-precommit]]
- [[2026-09-04-dsr-denominator-sr0-vs-srobs-resolution]]
Sources
Vault (RDCO-authored; all N=6/k=2, embargo_frac, φ-formula, SPY bar count, three-partition and Sharpe-SE figures come from here):
~/rdco-vault/06-reference/research/2026-06-26-strategy-discovery-loop-architecture.md— open follow-up #2, the question's origin~/rdco-vault/06-reference/research/2026-06-01-cpcv-deflated-sharpe-autoinv-validation.md— CPCV/DSR math,combinatorial_purged_cvsketch, trade-count precision gates~/rdco-vault/06-reference/research/2026-07-05-autoinv-cpcv-vs-walk-forward-confirm-stage.md— proposed signature with defaults, two-regime confirm stage~/rdco-vault/01-projects/investing/2026-05-29-strategy-pipeline-architecture-v0.md— the surface parenthetical, three-partition discipline, DSR gates~/rdco-vault/06-reference/research/2026-08-31-sealed-holdout-refresh-policy.md— repo audit, Sharpe-SE arithmetic, per-surface L_min~/rdco-vault/06-reference/research/2026-08-31-dsr-effective-n-estimator-precommit.md— N_eff ruling that supersedes CPCV-paths-as-N~/rdco-vault/06-reference/research/2026-09-04-dsr-denominator-sr0-vs-srobs-resolution.md
Repo (read directly, 2026-09-05):
/Users/ray/Projects/automated-investing/autoinv/validation.py— 98 lines,TimeSeriesSplit+BiasAuditonly; no purge/embargo/CPCV/Users/ray/Projects/automated-investing/data_cache/SPY_2018-01-01_2023-12-31_5f6c0c96.csv— 1,509 daily bars, the only cached series/Users/ray/Projects/automated-investing/PHASE1-SPINE.md— explicit "Phase 2+" deferral of walk-forward and holdout/Users/ray/Projects/automated-investing/autoinv/data.py,autoinv/feed.py— no default date spans anywhere
Computed for this brief (reproducible script at /tmp/cpcv_sizing.py; every table above is my arithmetic, not a cited figure):
train_net = (N−k)·⌊T/N⌋ − k·(⌈h·T⌉ + H);N_max = ⌊T / (3·(H+E))⌋;paths = C(N−1, k−1);SE(SR_ann) ≈ √(252/T)
External:
- skfolio — CombinatorialPurgedCV — defaults
n_folds=10, n_test_folds=8, purged_size=0, embargo_size=0; constraintsn_folds ≥ 3,n_test_folds ≥ 2 - QuantInsti — Cross Validation in Finance: Purging, Embargoing, Combinatorial — worked (6,2) example on 6,685 observations; lookback-matched embargo practice
- Not fetched, flagged not established: a search-result claim (attributed to quantbeckman.com) that path counts below ~100 yield unstable performance distributions