06-reference/research

afml embargo cscv primary source verification

2026-09-05·research-brief·source: deep-research·by Ray Data Co (deep-research synthesis)
cscvcpcvpbodeflated-sharpeprimary-source-verification

CSCV's S = 16 is primary-confirmed (and the paper misprints its own combination count); the 1% embargo is still blog-sourced; and "lower PBO / higher DSR vs walk-forward" is not a de Prado claim at all

The question

Read AFML primary (de Prado ch. 7 + ch. 12) and SSRN 2460551 to verify the embargo-fraction default and the exact CSCV partition count behind the "lower PBO / higher DSR" claim (cited from secondary sources in the discovery-loop brief, not deep-read).

Open follow-up #1 of [[2026-06-26-strategy-discovery-loop-architecture]], carried forward unresolved through [[2026-06-01-cpcv-deflated-sharpe-autoinv-validation]] and [[2026-07-05-autoinv-cpcv-vs-walk-forward-confirm-stage]]. The whole value of this brief is provenance: every number below carries an explicit tier label — (a) primary text I actually read, (b) a source that quotes/cites the primary, or (c) still only blog/secondary.

What we already know (from the vault)

What the web says

Finding 0 — the question's own citation is misrouted. [tier (a), settled] SSRN 2460551 is The Deflated Sharpe Ratio (Bailey & López de Prado, 2014). It contains no CSCV partition count and no embargo. The CSCV/PBO paper is SSRN 2326253: Bailey, Borwein, López de Prado & Zhu, The Probability of Backtest Overfitting, revised 27 Feb 2015, free full text at davidhbailey.com/dhbpapers/backtest-prob.pdf. I retrieved that PDF and read the extracted text directly; every quote below is from it.

Finding 1 — CSCV's partition count S. [tier (a), PRIMARY CONFIRMED] Algorithm 2.3, step two (p.11): "we partition M across rows, into an even number S of disjoint submatrices of equal dimensions. Each of these submatrices M_s, with s = 1, …, S, is of order (T/S × N)." Step three forms all combinations taken in groups of S/2, giving C(S, S/2) (Eq. 2.3). The paper's recommendation is explicit (p.21): "if M contains 4 years of daily data, S = 16 would equate to quarterly partitions, and the serial correlation structure would be preserved. For these two reasons, we believe that S = 16 is a reasonable value to use in most cases." So S = 16 is a stated recommendation, not a hard default, and it is derived from two data-dependent constraints: enough logits to populate the left tail, and partitions long enough to preserve serial-correlation structure.

Finding 2 — the paper misprints C(16,8), twice. [tier (a), arithmetic checked] Both p.11 and p.21 read "if S = 16, we will form 12, 780 combinations" / "12, 780 logits." The correct value is C(16,8) = 12,870. The misprint is a digit transposition, and the paper's other two worked examples are exact: S = 12 → "924 logits" = C(12,6) = 924 ✓; S = 24 → "2,704,156 logits" = C(24,12) = 2,704,156 ✓. Its downstream inference survives either number (σ[f(λ)] = √(1/4N) = 0.00441 at 12,870 vs 0.00442 at 12,780, both under the stated 0.0045 bound). Use 12,870. Any RDCO unit test asserting the paper's printed figure would be asserting a typo.

Finding 3 — the PBO definition and its rejection threshold. [tier (a), PRIMARY CONFIRMED] Definition 2.2 (p.9), verbatim: "A strategy with optimal performance IS is not necessarily optimal OOS. Moreover, there is a non-null probability that this strategy with optimal performance IS ranks below the median OOS. This is what we define as the probability of backtest overfit (PBO)," formalized as PBO = Σ_n Prob[r_n < N/2 | r ∈ Ω*_n]·Prob[r ∈ Ω*n] (Eq. 2.2). Estimated under CSCV as φ = ∫{−∞}^{0} f(λ)dλ over the logits λ_c = ln(ω̄_c/(1−ω̄_c)). And a threshold the vault does not currently carry (p.13): "a customary approach would be to reject models for which PBO is estimated to be greater than 0.05."

Finding 4 — the embargo fraction is NOT in this paper at all. [tier (a) negative result] A full-text search of the 2015 CSCV paper returns zero occurrences of "embargo", "purge", or "purging". CSCV as published is a symmetric combinatorial resampling of the trials matrix with no leakage treatment. Purging and embargoing are AFML (2018) ch. 7 constructs; "combinatorial purged CV" in ch. 12 is the later composition of ch. 7's leakage treatment onto this 2015 algorithm. This is a meaningful negative: the paper the vault pointed at for the embargo default could never have contained it.

Finding 5 — the embargo fraction h ≈ 1% remains UNVERIFIED. [tier (c), still secondary] AFML is a copyrighted Wiley book with no legitimate free full text. My reachable-source attempts failed: the mlfinlab implementation docs 404'd; the arXiv paper surfaced by search ("Confronting Machine Learning With Financial Research", arXiv 2103.00366) turns out to be by Lommers, El Harzli & Kim — not de Prado — and contains no numeric embargo fraction; targeted search for de Prado's own free lecture-slide version of ch. 7 returned nothing usable. No copy of AFML exists on local disk. The "~1% of observations" figure therefore still rests only on quantinsti/Wikipedia-tier summaries, exactly where it stood in June. I am not upgrading it.

Finding 6 — "lower PBO / higher DSR vs walk-forward" is a secondary extrapolation, not a de Prado claim. [tier (a) negative result] The CSCV paper contains zero occurrences of "walk-forward." Its comparison baseline throughout is the hold-out method, argued against on four grounds (it consumes sample, ignores the number of trials, has high estimation variance, and is inadequate for small samples); k-fold and LOOCV are mentioned once in passing as related literature. The paper reports no Sharpe ratios and never computes a DSR — DSR is a different paper (SSRN 2460551) published the following year. So the comparative claim is a composition across three documents, none of which makes it. Its actual origin is a 2024 Knowledge-Based Systems controlled-environment study (ScienceDirect S0950705124011110) — a third-party empirical comparison, paywalled, skipped, and not primary de Prado.

Convergences and contradictions

Synthesis for RDCO

Two of the three target claims are now resolved to primary, and the third is honestly left open. The CSCV partition count is settled: S = 16, C(16,8) = 12,870, presented as "a reasonable value to use in most cases" and justified by two constraints — enough logits for a stable left-tail estimate, and partitions long enough to preserve serial correlation (quarterly blocks on four years of daily data). That framing matters more than the number. S is not a constant to hardcode; it is a function of how much history the surface has. On a swarm surface with two years of daily bars, S = 16 gives ~six-week partitions and starts shredding the monthly structure the paper warns about, and the honest move is either to halve S (S = 12 → 924 logits, still adequate for a 0.05 gate) or to lengthen T. The PBO rejection threshold is likewise now primary-sourced at φ > 0.05, which the vault's GATE-2 spec never carried; that is a free, citable gate value that autoinv Phase-2 should adopt rather than invent.

The load-bearing correction is architectural, not numeric. CSCV is not a consumer of CPCV output. It resamples the (T × N) matrix of per-period returns for all N candidate configurations — the raw trials matrix from the vectorbt sweep, not the five reassembled CPCV paths from the confirm stage. This changes the validation.py signature and, more importantly, changes where GATE-2 sits in the pipeline: PBO is a property of the search process over a family of candidates, computable only when you retain the full per-period return series for every cell you tried, and it is meaningless on a single survivor. If the sweep stage discards the losing cells' return series and forwards only the winner's parameters, GATE-2 becomes uncomputable downstream. That is a data-retention requirement on the sweep, and it needs to be in the Phase-2 spec before code. It also vindicates the vault's own instinct that a high PBO should "reject the whole family, not just the cell" — the primary's Definition 2.2 is literally a statement about the selection procedure, not about a strategy.

The embargo default has to be marked as unsourced in the build, and the comparative claim has to be demoted. embargo_frac = 0.01 currently sits in the combinatorial_purged_cv signature with a provenance chain that terminates in a quant blog. That is not a reason to change the value — 1% is plausible and conservative on daily data — but it is a reason to (i) write # NOT primary-sourced: quantinsti/Wikipedia tier; AFML ch.7 not read next to it, and (ii) treat it as a parameter to sensitivity-test rather than a constant to trust. The right empirical substitute costs nothing: sweep h over {0, 0.005, 0.01, 0.02, 0.05} on a known-null strategy and pick the smallest h at which the CV score stops being inflated. That is a better answer than a citation would have been. Separately, the "CPCV → lower PBO / higher DSR than walk-forward" line should be rewritten wherever it appears as "a 2024 controlled-environment study reports…" and never as de Prado canon. The de Prado corpus argues CSCV beats hold-out, and argues CPCV yields a distribution where walk-forward yields a point — a structural argument, not a measured performance comparison. The 07-05 brief's conclusion (carry both regimes) is unaffected, because it rests on the structural argument, which is primary-supported.

Net effect on the discovery-loop spec: no gate is removed, one gate's input is corrected, one gate gains a primary threshold, and one constant is downgraded to a labelled guess. The stack was directionally right and is now more precisely sourced in two places and more honestly hedged in a third.

Why this is in the vault

It closes open-follow-up #1 of [[2026-06-26-strategy-discovery-loop-architecture]] (carried unresolved in [[2026-06-01-cpcv-deflated-sharpe-autoinv-validation]] and [[2026-07-05-autoinv-cpcv-vs-walk-forward-confirm-stage]]) and hands autoinv Phase-2 three concrete build inputs before validation.py's GATE-2 is coded: the corrected (T × N) input matrix for probability_of_backtest_overfitting(), the primary-sourced rejection threshold φ > 0.05, and the correct combination count 12,870 (the paper's own printed 12,780 is a typo and must not become a unit-test constant).

Open follow-ups

Related

Sources

Primary, read directly (tier a):

Attempted and failed / not primary:

Still secondary (tier c), carried forward unchanged:

Vault:

Research budget used: 3 WebSearch, 3 WebFetch (1 of which returned a PDF extracted locally), 1 qmd query, 7 vault docs consulted via targeted grep.