06-reference/research

dsr effective n estimator precommit

2026-08-31·research-brief·source: deep-research·by Ray Data Co (deep-research synthesis)
deflated-sharpeeffective-nmultiple-testingcpcvautomated-investingpre-registration

GATE-1's N is a property of the search, not of the candidate: pre-registering the effective-N estimator

The question

"Pre-register the effective-N estimator for the discovery-loop DSR gate: does GATE-1's N come from the cumulative trial counter, the CPCV path count, or a correlation-adjusted effective-N? (Canon says CPCV paths; v0's sweep counts grid cells.) Resolve before coding Phase 2."

Context: this is open-follow-up #1 carried from the 2026-06-26 discovery-loop architecture brief and re-filed 2026-07-05 as blocking. Every Deflated Sharpe Ratio (DSR) the autoinv pipeline will ever report is a function of this choice, so it is a pre-registration, not a report.

What we already know (from the vault)

What the web says

Convergences and contradictions

Synthesis for RDCO

The canon does not say "CPCV paths." Read closely, it says the opposite. The one paper in the López de Prado corpus written specifically to answer "what is the effective N" is a clustering paper, and what it clusters is the set of trials. Wikipedia's DSR article, which the 2026-06-01 brief already cited, points at ONC for exactly this. So the "canon says CPCV paths" premise in the backlog question is a vault-internal artifact: it entered as a code comment on 2026-06-01, hardened into a recommendation on 2026-07-05, and was never checked against the DSR literature. It should not be ratified. The distinction it obscures is the whole reason GATE-1 exists. DSR asks a question about the search: given that we looked at a large number of configurations and kept the best one, how much of that best Sharpe is the expected maximum of a set of lucky draws? A CPCV path count carries no information about how many configurations were examined, so it cannot answer that question at any value of N. Meanwhile the thing the search does know, the cumulative trial ledger, is exactly the object the literature says to correct for correlation.

The correction is real, and it must be built adversarially, because it only ever weakens the gate. Every effective-N estimator reduces N, which reduces SR₀, which makes GATE-1 easier to pass. This is a knob that points in one direction, on a gate whose entire purpose is to be hard to pass, in a pipeline whose pre-registered expectation is that most candidates FAIL. That asymmetry earns four mechanical guardrails, all pre-registered here: (1) the estimator is fixed before Phase-2 code is written and cannot be re-selected after seeing a candidate's DSR; (2) when multiple estimators disagree, take the maximum N_eff, since larger N is the harsher hurdle; (3) N_eff is clamped to [1, N_literal] and can never exceed the ledger cardinality; (4) if the trial return matrix is unavailable, degenerate, or fewer than ~30 trials, the estimator refuses and falls back to N_literal. The ledger stays cumulative across sessions per v0, and it must record abandoned and negative-result trials, because a trial you ran and discarded still consumed a draw from the null.

Where CPCV legitimately enters, and one T-counting trap that will silently inflate every DSR if it is missed. CPCV feeds three things, none of them N. First, SR_obs: the candidate's honest point estimate should be the mean of its CPCV path Sharpes from the event-driven autoinv confirm stage with the real cost model, not the optimistic vectorbt sweep Sharpe. Pairing a cost-realistic SR_obs against a sweep-derived SR₀ hurdle is conservative in the direction the discipline wants. Second, an empirical cross-check on the DSR denominator: the analytic standard error √(1 − γ̂₃·SR + ((γ̂₄−1)/4)·SR²)/√(T−1) assumes away serial correlation and regime dependence, so compute the standard deviation of the CPCV path Sharpes as well and, if it exceeds the analytic term, use the larger. That gives CPCV a principled role without letting it touch N. Third, GATE-2's CSCV partitions, which is a separate count entirely. The trap: CPCV paths reuse observations, since every bar appears in multiple test folds. Pooling all path returns and setting T to the concatenated length would inflate T by roughly φ[N,k] (5× at N=6, k=2) and DSR scales in √(T−1), so a 5× T inflation is a ~2.2× inflation of the DSR z-score. T must be the count of distinct return observations in the OOS union, each counted once. This is the operational teeth on the existing "T = return observations, not trades" rule, and it belongs in a unit test.

Units consistency is the last thing to nail down before code. SR₀ is a property of the search and is computed in sweep space; SR_obs is a property of the candidate and is computed in confirm space. For the subtraction (SR_obs − SR₀) to mean anything, both must use the same convention: per-return-observation (non-annualized) Sharpe, on the same bar frequency, with the same nominal cost model applied. Concretely that means running the vectorbt sweep with the same fee and slippage assumptions as the confirm engine, even though the sweep's fills are idealized, so the trial Sharpe dispersion is measured in comparable units. Record the assumption in the multiple-testing ledger. One proposal for the GATE-0 hook, offered as a build choice rather than a literature finding: implement prior_strength as a pre-registered multiplier on N_eff (a none-prior candidate is effectively drawn from a far larger implicit search space, so ×10 its trial count) rather than as a second DSR threshold. That keeps one threshold, 0.5 reject / 1.0 confidence, and expresses the story-less penalty in the units it actually lives in.

PRE-REGISTERED ESTIMATOR — this is the thing to code.

GATE-1's N is N_eff, the correlation-adjusted effective number of trials, computed from the cumulative trial ledger. It is NOT the CPCV path count, and it is NOT the raw grid-cell count.

Artifact: the multiple-testing ledger produced by the vectorbt discovery sweep on the TRAIN/VALIDATION partition, cumulative across sessions. For every trial i in 1..N_literal it stores trial_id, the params, the per-observation validation-partition return series r_i, and the non-annualized validation Sharpe SR_i. Holdout is import-walled out of this object.

Computation:

  1. N_literal = ledger cardinality, cumulative, including abandoned and negative-result trials.
  2. C = correlation matrix of the trial return series {r_i} on the validation partition.
  3. K_onc = number of clusters from ONC on C (de Prado/Lewis primary). K_rank = effective_rank(C). K_mp = count of eigenvalues of C above the Marchenko-Pastur noise bound. Compute all three.
  4. N_eff = clamp( ceil( max(K_onc, K_rank, K_mp) ), 1, N_literal ) — maximum, because larger N is the harsher hurdle. Fall back to N_eff = N_literal if the ledger has fewer than 30 trials or C is degenerate.
  5. V[SR_n] = cross-sectional variance of the cluster-representative Sharpes (one per ONC cluster, the cluster's aggregated return series), consistent with step 4's clustering. NOT the CPCV path Sharpes, and not the sampling variance of any single strategy.
  6. SR_0 = sqrt(V[SR_n]) * ((1 - EM) * Phi_inv(1 - 1/N_eff) + EM * Phi_inv(1 - 1/(N_eff * e))), EM = 0.5772156649.
  7. SR_obs = mean of the candidate's CPCV path Sharpes from the autoinv event-driven confirm stage, same per-observation convention and cost model as step 1.
  8. T = count of distinct return observations in the CPCV out-of-sample union. Never the concatenated path length.
  9. se = max( sqrt(1 - skew*SR_obs + ((kurt - 1)/4) * SR_obs**2) / sqrt(T - 1), stdev(cpcv_path_sharpes) ).
  10. DSR = Phi( (SR_obs - SR_0) / se ). Gate at DSR < 0.5 reject, > 1.0 confidence. Refuse to emit at all if T < MinTRL(SR_obs, skew, kurt, SR_0, confidence).

Report N_literal, N_eff, all three cluster estimates, V[SR_n], SR_0, SR_obs, T, and both standard-error terms side by side in the multiple-testing ledger. If N_eff / N_literal < 0.05, flag the run for review rather than silently accepting a 20× hurdle reduction.

Why this is in the vault

This pre-registers the single numeric input that sets the pass bar for GATE-1 in the autoinv Phase-2 build (validation.py: deflated_sharpe(...) and the multiple-testing ledger), which is the specific coding task blocked since 2026-07-05. It also corrects a wrong instruction already sitting in two vault briefs, so anyone who codes from [[2026-06-01-cpcv-deflated-sharpe-autoinv-validation]] or [[2026-07-05-autoinv-cpcv-vs-walk-forward-confirm-stage]] without reading this one will build a DSR gate that under-deflates by roughly 2x in the dispersion multiplier and passes noise.

Open follow-ups

Related

Sources

Vault:

Web: