GATE-1's N is a property of the search, not of the candidate: pre-registering the effective-N estimator
The question
"Pre-register the effective-N estimator for the discovery-loop DSR gate: does GATE-1's N come from the cumulative trial counter, the CPCV path count, or a correlation-adjusted effective-N? (Canon says CPCV paths; v0's sweep counts grid cells.) Resolve before coding Phase 2."
Context: this is open-follow-up #1 carried from the 2026-06-26 discovery-loop architecture brief and re-filed 2026-07-05 as blocking. Every Deflated Sharpe Ratio (DSR) the autoinv pipeline will ever report is a function of this choice, so it is a pre-registration, not a report.
What we already know (from the vault)
- v0 counts grid cells, and already knows that is wrong. [[2026-05-29-strategy-pipeline-architecture-v0]] specs the discovery sweep as "auto-incrementing a cumulative trial counter N," where "every cell is a counted trial that increments N," cumulative across sessions so the data-mining discount stays honest over the project's life. It also carries the red-team's own caveat verbatim: "a param grid produces highly-correlated neighboring cells, so effective N is below counted N and naive DSR will UNDER-deflate." Its only proposed mitigation is blunt: cap grid granularity so cells are meaningfully distinct.
- The formulas are already pinned. [[2026-06-01-cpcv-deflated-sharpe-autoinv-validation]] records the Bailey and López de Prado machinery: the benchmark SR₀ = √V[SR_n] · ((1−γ)·Φ⁻¹[1−1/N] + γ·Φ⁻¹[1−1/(N·e)]) with γ ≈ 0.5772; DSR = Φ((SR* − SR₀)·√(T−1) / √(1 − γ̂₃·SR + ((γ̂₄−1)/4)·SR²)); and MinTRL as the refusal gate. It also states the scope rule that governs T: DSR is defined over T = return observations, never round-trip trades.
- The vault contains the conflation this brief has to break. The same 2026-06-01 brief's
validation.pysketch annotates the two arguments inconsistently:n_trials: int, # effective N (cumulative trial counter, not grid-cell count)alongsidetrial_sharpe_var: float, # V[SR_n] across trials -> from the CPCV path Sharpes. [[2026-07-05-autoinv-cpcv-vs-walk-forward-confirm-stage]] then hardens the second half into a recommendation: "DSR needs V[SR_n] ... and an effective-N," and "its near-independent path Sharpes are a cleaner V[SR_n] and effective-N than any grid-cell count." That is the claim under test here. - RDCO already owns a correlation-to-independence estimator, one layer down the stack. [[2026-06-01-ensemble-admission-correlation-threshold]] pre-registers the Effective Number of Bets (ENB) computed via Minimum-Linear-Torsion, with the explicit finding that the Principal-Components route produces false negatives (ENB ≈ 1.0 for three near-uncorrelated assets) and a pre-registered floor of ENB ≥ 0.7 × k. v0 runs the same collapse-correlated-things logic at the position layer ("MU+SMH+SNDK+INTC = ~one semi-memory bet"). Effective-N for the trial ledger is the third instance of one mechanism RDCO has already committed to twice.
- The False-Strategy framing is already in the vault via [[2026-05-31-ensembles-systematic-trading-overfitting]]: E[max SR] ≈ √(2·ln N)·σ(trial Sharpes), so best-of-N looks good by construction. The variable being logged there is σ of the trial Sharpes.
What the web says
- N is unambiguously the trial count of the search process, and the literature's own fix for correlation is clustering, not cross-validation. The DSR reference article states that "N is the effective number of independent trials that have been run. It is not always equal to the literal count of backtests if many are highly correlated," and that "to estimate the effective number of independent trials N, López de Prado (2018) proposes three techniques to clustering similar strategies using unsupervised learning techniques: The Optimal Number of Clusters (ONC) algorithm" (Wikipedia, Deflated Sharpe ratio).
- The dedicated primary on effective-N is a clustering paper, and it clusters trials. López de Prado and Lewis "introduce an unsupervised learning algorithm that determines the number of effectively uncorrelated trials carried out in the context of a discovery, which is critical for computing the familywise false positive probability and for filtering out false investment strategies." The named techniques are ONC, hierarchical clustering (a conservative lower bound), and spectral methods on the eigenvalue distribution of the correlation matrix (Detection of false investment strategies using unsupervised learning methods, Quantitative Finance 19(9), 2019; SSRN 3167017).
- The DSR primary defines the two inputs as properties of the trial set. N is "the count of independent strategy trials evaluated," and V[SR_n] "reflects dispersion across the N trial Sharpe ratios generated" (Bailey and López de Prado, The Deflated Sharpe Ratio, SSRN 2460551). The paper separately treats T as "the number of return observations/periods in the backtest," which "directly affects the standard error of the Sharpe ratio estimate" — a different quantity in a different slot of the formula.
- A working implementation confirms the composed shape, and names the estimator menu. The ml4t-diagnostic package notes that "when candidate strategies are closely related, raw K can overstate the true breadth of the search," and "therefore also supports correlation-adjusted K_eff via
effective_rank,marchenko_pastur, andclustering" (ml4trading.io, Deflated Sharpe Ratio diagnostic). Same source on magnitude: "the expected spurious Sharpe ratio grows logarithmically with the number of strategies tested," and for K=50 it is already ~0.4–0.8. - Nothing in the DSR or CPCV literature proposes cross-validation paths as N. CPCV's stated product is "a distribution of out-of-sample performance estimates, enabling robust statistical inference" for one configuration (Towards AI, CPCV method), and the controlled comparison credits CPCV with lower PBO and higher DSR as a validation regime, not as the source of DSR's N (ScienceDirect, backtest overfitting comparison).
Convergences and contradictions
- Convergence: vault and web agree on the disease. Counted grid cells overstate independence, the correction is downward, and the correction should be estimated from the correlation structure of the trials rather than asserted by capping grid granularity.
- The contradiction, and it is the point of this brief. The vault's expected answer (carried in the backlog note and stated in the 2026-07-05 synthesis) is that effective-N and V[SR_n] come from CPCV path Sharpes. That conflates two different populations. DSR's multiple-testing deflation is about how many strategy configurations were tried — the breadth of the search, which is what creates selection bias. CPCV paths are how many reassembled backtest paths exist for one configuration — the precision of a single estimate. Substituting CPCV paths for N replaces something on the order of 2,000 with 5 (φ[6,2]), and since SR₀ grows in ln N, that lowers the hurdle by roughly a factor of √(ln 2000 / ln 5) ≈ 2.2 in the dispersion multiplier. It is not a conservative correction. It is the largest possible under-deflation, applied in the name of fixing under-deflation.
- A second, smaller conflation in the same slot. V[SR_n] is the cross-sectional variance across trials. The dispersion of a single strategy's CPCV path Sharpes is the sampling variability of one SR estimate, which the DSR formula already handles in the denominator via T, skew and kurtosis. Feeding path dispersion into V[SR_n] double-books the sampling error into the hurdle and leaves the selection-bias term unpopulated.
- They are not mutually exclusive, but they are not interchangeable. All three candidates belong in the pipeline; they occupy three different slots. The cumulative counter is the ledger cardinality feeding the effective-N estimator; the correlation-adjusted effective-N is what actually enters SR₀; the CPCV paths supply SR_obs, the empirical standard-error cross-check, and the CSCV partitions for GATE-2 (PBO). Composed formula below.
- A formula-level discrepancy worth fixing while the file is open. The 2026-06-01 brief's prose writes the DSR denominator with SR₀ inside it (
1 − γ̂₃·SR₀ + ((γ̂₄−1)/4)·SR₀²) while its own code comment and its MinTRL expression both use the observed SR. The denominator is the standard error of the Sharpe estimator (Lo/Mertens), so it must be evaluated at the observed SR-hat. High confidence, but verify against the primary before it is coded, since the two versions give different DSRs on every candidate.
Synthesis for RDCO
The canon does not say "CPCV paths." Read closely, it says the opposite. The one paper in the López de Prado corpus written specifically to answer "what is the effective N" is a clustering paper, and what it clusters is the set of trials. Wikipedia's DSR article, which the 2026-06-01 brief already cited, points at ONC for exactly this. So the "canon says CPCV paths" premise in the backlog question is a vault-internal artifact: it entered as a code comment on 2026-06-01, hardened into a recommendation on 2026-07-05, and was never checked against the DSR literature. It should not be ratified. The distinction it obscures is the whole reason GATE-1 exists. DSR asks a question about the search: given that we looked at a large number of configurations and kept the best one, how much of that best Sharpe is the expected maximum of a set of lucky draws? A CPCV path count carries no information about how many configurations were examined, so it cannot answer that question at any value of N. Meanwhile the thing the search does know, the cumulative trial ledger, is exactly the object the literature says to correct for correlation.
The correction is real, and it must be built adversarially, because it only ever weakens the gate. Every effective-N estimator reduces N, which reduces SR₀, which makes GATE-1 easier to pass. This is a knob that points in one direction, on a gate whose entire purpose is to be hard to pass, in a pipeline whose pre-registered expectation is that most candidates FAIL. That asymmetry earns four mechanical guardrails, all pre-registered here: (1) the estimator is fixed before Phase-2 code is written and cannot be re-selected after seeing a candidate's DSR; (2) when multiple estimators disagree, take the maximum N_eff, since larger N is the harsher hurdle; (3) N_eff is clamped to [1, N_literal] and can never exceed the ledger cardinality; (4) if the trial return matrix is unavailable, degenerate, or fewer than ~30 trials, the estimator refuses and falls back to N_literal. The ledger stays cumulative across sessions per v0, and it must record abandoned and negative-result trials, because a trial you ran and discarded still consumed a draw from the null.
Where CPCV legitimately enters, and one T-counting trap that will silently inflate every DSR if it is missed. CPCV feeds three things, none of them N. First, SR_obs: the candidate's honest point estimate should be the mean of its CPCV path Sharpes from the event-driven autoinv confirm stage with the real cost model, not the optimistic vectorbt sweep Sharpe. Pairing a cost-realistic SR_obs against a sweep-derived SR₀ hurdle is conservative in the direction the discipline wants. Second, an empirical cross-check on the DSR denominator: the analytic standard error √(1 − γ̂₃·SR + ((γ̂₄−1)/4)·SR²)/√(T−1) assumes away serial correlation and regime dependence, so compute the standard deviation of the CPCV path Sharpes as well and, if it exceeds the analytic term, use the larger. That gives CPCV a principled role without letting it touch N. Third, GATE-2's CSCV partitions, which is a separate count entirely. The trap: CPCV paths reuse observations, since every bar appears in multiple test folds. Pooling all path returns and setting T to the concatenated length would inflate T by roughly φ[N,k] (5× at N=6, k=2) and DSR scales in √(T−1), so a 5× T inflation is a ~2.2× inflation of the DSR z-score. T must be the count of distinct return observations in the OOS union, each counted once. This is the operational teeth on the existing "T = return observations, not trades" rule, and it belongs in a unit test.
Units consistency is the last thing to nail down before code. SR₀ is a property of the search and is computed in sweep space; SR_obs is a property of the candidate and is computed in confirm space. For the subtraction (SR_obs − SR₀) to mean anything, both must use the same convention: per-return-observation (non-annualized) Sharpe, on the same bar frequency, with the same nominal cost model applied. Concretely that means running the vectorbt sweep with the same fee and slippage assumptions as the confirm engine, even though the sweep's fills are idealized, so the trial Sharpe dispersion is measured in comparable units. Record the assumption in the multiple-testing ledger. One proposal for the GATE-0 hook, offered as a build choice rather than a literature finding: implement prior_strength as a pre-registered multiplier on N_eff (a none-prior candidate is effectively drawn from a far larger implicit search space, so ×10 its trial count) rather than as a second DSR threshold. That keeps one threshold, 0.5 reject / 1.0 confidence, and expresses the story-less penalty in the units it actually lives in.
PRE-REGISTERED ESTIMATOR — this is the thing to code.
GATE-1's N is
N_eff, the correlation-adjusted effective number of trials, computed from the cumulative trial ledger. It is NOT the CPCV path count, and it is NOT the raw grid-cell count.Artifact: the multiple-testing ledger produced by the vectorbt discovery sweep on the TRAIN/VALIDATION partition, cumulative across sessions. For every trial
iin1..N_literalit storestrial_id, the params, the per-observation validation-partition return seriesr_i, and the non-annualized validation SharpeSR_i. Holdout is import-walled out of this object.Computation:
N_literal= ledger cardinality, cumulative, including abandoned and negative-result trials.C= correlation matrix of the trial return series{r_i}on the validation partition.K_onc= number of clusters from ONC onC(de Prado/Lewis primary).K_rank=effective_rank(C).K_mp= count of eigenvalues ofCabove the Marchenko-Pastur noise bound. Compute all three.N_eff = clamp( ceil( max(K_onc, K_rank, K_mp) ), 1, N_literal )— maximum, because larger N is the harsher hurdle. Fall back toN_eff = N_literalif the ledger has fewer than 30 trials orCis degenerate.V[SR_n]= cross-sectional variance of the cluster-representative Sharpes (one per ONC cluster, the cluster's aggregated return series), consistent with step 4's clustering. NOT the CPCV path Sharpes, and not the sampling variance of any single strategy.SR_0 = sqrt(V[SR_n]) * ((1 - EM) * Phi_inv(1 - 1/N_eff) + EM * Phi_inv(1 - 1/(N_eff * e))),EM = 0.5772156649.SR_obs= mean of the candidate's CPCV path Sharpes from theautoinvevent-driven confirm stage, same per-observation convention and cost model as step 1.T= count of distinct return observations in the CPCV out-of-sample union. Never the concatenated path length.se = max( sqrt(1 - skew*SR_obs + ((kurt - 1)/4) * SR_obs**2) / sqrt(T - 1), stdev(cpcv_path_sharpes) ).DSR = Phi( (SR_obs - SR_0) / se ). Gate at DSR < 0.5 reject, > 1.0 confidence. Refuse to emit at all ifT < MinTRL(SR_obs, skew, kurt, SR_0, confidence).Report
N_literal,N_eff, all three cluster estimates,V[SR_n],SR_0,SR_obs,T, and both standard-error terms side by side in the multiple-testing ledger. IfN_eff / N_literal < 0.05, flag the run for review rather than silently accepting a 20× hurdle reduction.
Why this is in the vault
This pre-registers the single numeric input that sets the pass bar for GATE-1 in the autoinv Phase-2 build (validation.py: deflated_sharpe(...) and the multiple-testing ledger), which is the specific coding task blocked since 2026-07-05. It also corrects a wrong instruction already sitting in two vault briefs, so anyone who codes from [[2026-06-01-cpcv-deflated-sharpe-autoinv-validation]] or [[2026-07-05-autoinv-cpcv-vs-walk-forward-confirm-stage]] without reading this one will build a DSR gate that under-deflates by roughly 2x in the dispersion multiplier and passes noise.
Open follow-ups
- Verify the within-cluster aggregation rule against the López de Prado and Lewis primary. Step 5 assumes
V[SR_n]is taken over cluster-representative Sharpes formed from aggregated cluster return series. Both PDF fetches for that paper failed to extract text, so this specific mechanic is asserted from secondary sources. The alternative (variance over allN_literaltrial Sharpes while N usesN_eff) gives a different, generally largerV[SR_n]and therefore a harsher hurdle. - Resolve the DSR-denominator discrepancy in [[2026-06-01-cpcv-deflated-sharpe-autoinv-validation]] — prose writes SR₀ inside the standard-error term, the code comment and MinTRL both write SR_obs. Confirm against SSRN 2460551 and correct the older brief, since the two give different DSRs on every candidate.
- Calibrate the
N_eff / N_literal < 0.05review flag and the ≥30-trial fallback floor against a real sweep. Both are pre-registered defaults with no empirical basis; run the Phase-2 pure-noise 100-cell V-model test and a real momentum sweep and see what ratio a healthy grid actually produces. - Decide whether
N_effis recomputed per candidate or once per ledger state. Clustering the full cumulative ledger every time a candidate is gated is expensive and makes a candidate's DSR depend on trials run after it. A frozen-at-gate-time snapshot is cheaper and reproducible but drifts from the true cumulative breadth. - Test whether ENB-via-MLT from [[2026-06-01-ensemble-admission-correlation-threshold]] should replace or join the three estimators in step 3. RDCO has already pre-registered MLT over PCA at the ensemble layer with a documented PCA false-negative; using a different independence estimator two layers apart in the same pipeline needs either justification or unification.
- Specify how the
prior_strengthmultiplier onN_effis calibrated. The ×10 for anone-prior candidate is a proposal in this brief, not a literature finding, and it changes the gate for exactly the candidates most likely to be pattern-mined artifacts.
Related
- [[2026-06-26-strategy-discovery-loop-architecture]] — the parent control-flow brief; this resolves its open-follow-up #1, and corrects its "source effective-N from CPCV path Sharpes" line
- [[2026-07-05-autoinv-cpcv-vs-walk-forward-confirm-stage]] — the two-CV-regime ruling that stands; its effective-N/V[SR_n] recommendation is the one superseded here
- [[2026-06-01-cpcv-deflated-sharpe-autoinv-validation]] — DSR/SR₀/MinTRL formulas, the φ[N,k] path math, and the
validation.pysketch whosetrial_sharpe_varcomment introduced the conflation - [[2026-05-29-strategy-pipeline-architecture-v0]] — GATE-0/1/2 design, the cumulative auto-N trial ledger, and the red-team's original effective-N-below-counted-N flag
- [[2026-06-01-ensemble-admission-correlation-threshold]] — ENB-via-MLT, the PCA false-negative finding, and RDCO's existing correlation-to-independence machinery
- [[2026-05-31-ensembles-systematic-trading-overfitting]] — the False-Strategy framing E[max SR] ≈ √(2·ln N)·σ(trial Sharpes) and the purge/embargo/CPCV definitions
Sources
Vault:
~/rdco-vault/06-reference/research/2026-06-26-strategy-discovery-loop-architecture.md— open-follow-up #1 verbatim; the CPCV-paths-for-effective-N line under test~/rdco-vault/06-reference/research/2026-07-05-autoinv-cpcv-vs-walk-forward-confirm-stage.md— the hardened "CPCV path Sharpes are a cleaner V[SR_n] and effective-N" recommendation, superseded here~/rdco-vault/06-reference/research/2026-06-01-cpcv-deflated-sharpe-autoinv-validation.md— SR₀/DSR/MinTRL formulas, φ[N,k]=(k/N)·C(N,k), thevalidation.pysignatures, "T = return observations not trades"~/rdco-vault/01-projects/investing/2026-05-29-strategy-pipeline-architecture-v0.md— "every cell is a counted trial that increments N," cumulative-across-sessions rule, DSR 0.5/1.0 thresholds, effective-N-below-counted-N caveat, Phase-2 V-model noise test~/rdco-vault/06-reference/research/2026-06-01-ensemble-admission-correlation-threshold.md— ENB via Minimum-Linear-Torsion, PCA false-negative, ENB ≥ 0.7k floor~/rdco-vault/06-reference/research/2026-05-31-ensembles-systematic-trading-overfitting.md— False-Strategy Theorem framing, purge/embargo/CPCV definitions
Web:
- Deflated Sharpe ratio — https://en.wikipedia.org/wiki/Deflated_Sharpe_ratio ("N is the effective number of independent trials"; ONC named as the estimator). Direct fetch returned HTTP 404; content quoted via search-result extract, not deep-read.
- Bailey and López de Prado, The Deflated Sharpe Ratio — https://papers.ssrn.com/sol3/papers.cfm?abstract_id=2460551 (N = strategy configurations attempted; V[SR_n] = dispersion across the N trial Sharpes; T = return observations). Fetched via the davidhbailey.com PDF mirror; PDF text extraction was partial, so quotations are paraphrased and flagged rather than verbatim.
- López de Prado and Lewis, Detection of false investment strategies using unsupervised learning methods, Quantitative Finance 19(9) 2019 — https://www.tandfonline.com/doi/abs/10.1080/14697688.2019.1622311 and https://papers.ssrn.com/sol3/papers.cfm?abstract_id=3167017 (number of effectively uncorrelated trials; ONC, hierarchical clustering as conservative lower bound, spectral/eigenvalue methods). Tandfonline is paywalled; the codemacher.com PDF mirror fetched but failed text extraction. Cited from search-tier summary — the primary read is an open follow-up.
- ml4t-diagnostic, Deflated Sharpe Ratio — https://www.ml4trading.io/docs/diagnostic/methods/deflated-sharpe-ratio/ ("raw K can overstate the true breadth of the search"; correlation-adjusted K_eff via effective_rank, marchenko_pastur, clustering; spurious max SR ~0.4-0.8 at K=50)
- The Combinatorial Purged Cross-Validation method (Towards AI) — https://towardsai.com/p/l/the-combinatorial-purged-cross-validation-method (CPCV's product is a distribution of OOS estimates for one configuration)
- Backtest overfitting comparison, controlled environment — https://www.sciencedirect.com/science/article/abs/pii/S0950705124011110 (CPCV as a validation regime yielding lower PBO / higher DSR; not a source of DSR's N). Abstract-only, full text paywalled.