The DSR denominator is the standard error of the OBSERVED Sharpe: resolving the SR₀-vs-SR_obs defect in the 2026-06-01 brief
The question
"Resolve the DSR-denominator discrepancy: the 2026-06-01 CPCV brief writes SR₀ inside the standard-error term in prose but SR_obs in its own code comment (and in MinTRL). Confirm against SSRN 2460551 and correct the older brief."
Context: this is a live correctness defect in vault guidance that autoinv Phase 2 is being coded against. The two forms give a different Deflated Sharpe Ratio (DSR) on every candidate, so this is not a documentation nit — it is the gate arithmetic.
What we already know (from the vault)
- The defect is real and it is internal to one file. [[2026-06-01-cpcv-deflated-sharpe-autoinv-validation]] states the DSR in prose as
DSR = Φ( (SR*−SR₀)·√(T−1) / √(1 − γ̂₃·SR₀ + ((γ̂₄−1)/4)·SR₀²) )— benchmark Sharpe inside the radical — and labels that line "(Bailey & López de Prado 2014, verbatim)". - Its own
validation.pysketch, eight lines later, contradicts the prose. The code comment reads# DSR = Phi( (SR_obs - SR0)*sqrt(T-1) / sqrt(1 - skew*SR_obs + (kurt-1)/4 * SR_obs**2) )— observed Sharpe inside the radical. - MinTRL in the same file agrees with the code, not the prose.
MinTRL = 1 + (1 − γ̂₃·SR* + ((γ̂₄−1)/4)·SR*²) · ( Φ⁻¹(confidence) / (SR*−SR₀) )², where the file definesSR* = observed (non-annualized) Sharpe. So the file uses the observed Sharpe in that bracket in two places out of three, and the odd one out is the line carrying the "verbatim" claim. - The same file admits it never read the primary. Its Sources section lists the DSR paper as "(primary, cited not deep-read)". The "verbatim" label on line 33 was therefore unearned, which is the root cause: the formula was transcribed from a secondary summary.
- [[2026-08-31-dsr-effective-n-estimator-precommit]] flagged it at high confidence but stopped short. It called the correct answer ("the denominator is the standard error of the Sharpe estimator (Lo/Mertens), so it must be evaluated at the observed SR-hat") and explicitly deferred: "verify against the primary before it is coded." This brief closes that.
- Nothing else in the vault carries the wrong form. A full-vault grep for the DSR/MinTRL formulas returns only the 2026-06-01 file, [[2026-08-31-dsr-effective-n-estimator-precommit]] (which already writes it correctly,
1 - skew*SR_obs + ...), and prose references in [[2026-07-05-autoinv-cpcv-vs-walk-forward-confirm-stage]], [[2026-08-31-sealed-holdout-refresh-policy]] and [[2026-05-29-strategy-pipeline-architecture-v0]] that do not restate the denominator. The blast radius is one line.
What the web says
The primary settles it unambiguously. Bailey & López de Prado, Eq. (2), page 8, reproduced exactly as printed:
D̂SR ≡ P̂SR(ŜR₀) = Z[ (ŜR − ŜR₀)·√(T−1) / √( 1 − γ̂₃·ŜR + ((γ̂₄−1)/4)·ŜR² ) ]
The Sharpe inside the radical is ŜR, the observed/estimated Sharpe of the selected strategy — not ŜR₀. The paper's variable list on the same page is equally explicit: "ŜR₀ = √V{ŜR_n}Z⁻¹[1−1/N] + γZ⁻¹[1−(1/N)e⁻¹]), V[{ŜR_n}] is the variance across the trials' estimated SR and N is the number of independent trials. We also use information concerning the selected strategy: ŜR is its estimated SR, T is the sample length, γ̂₃ is the skewness of the returns distribution and γ̂₄ is the kurtosis of the returns distribution for the selected strategy." (Bailey & López de Prado 2014, p.8; SSRN 2460551). Read from the full PDF, rendered page images — not a search-result extract.
The paper's own numerical example plugs the observed Sharpe into the denominator, in print. Page 10 shows the fully-substituted expression:
DSR ≈ Z[ ((2.5/√250 − 0.1132)·√1249) / √(1 − (−3)·(2.5/√250) + ((10−1)/4)·(2.5/√250)²) ] = 0.9004. Every occurrence in the radical is2.5/√250— the observed Sharpe. Nowhere does0.1132appear below the line (ibid., p.10).The structural reason, stated by the paper and by an independent technical source. DSR is defined as
PSR(ŜR₀)— the Probabilistic Sharpe Ratio evaluated at threshold ŜR₀. In PSR the threshold enters the numerator only; the denominator is a fixed property of the estimator's sampling distribution. Portfolio Optimizer states the denominator of PSR is "the estimator of the standard error of ŜR", denoting itSE(ŜR), and attributes it to the asymptotic distribution of ŜR (Portfolio Optimizer, The Probabilistic Sharpe Ratio). Bailey & López de Prado cite Lo [2002] and Mertens [2002] for exactly this correction (p.8).Implementations agree with the paper. The widely-forked reference implementation computes
sr_std = sqrt((1 + 0.5*sr**2 - skew*sr + ((kurtosis-3)/4)*sr**2)/(n-1))withsrthe observed Sharpe (rubenbriones/Probabilistic-Sharpe-Ratio). That is algebraically identical to the paper:0.5 + (κ−3)/4 = (κ−1)/4. No implementation found anywhere writes the threshold into the denominator. There is no dissenting camp to adjudicate.Secondary sources are unreliable on the kurtosis convention, and the paper is not. Portfolio Optimizer's page was read as asserting excess kurtosis (Normal = 0). The primary contradicts that directly: "If the strategy had exhibited Normal returns (γ̂₃ = 0, γ̂₄ = 3)" (p.10). γ̂₄ is raw kurtosis. This is also forced analytically — at γ̂₄ = 3 the radicand collapses to
1 + SR²/2, the classic Lo (2002) normal-returns standard error; plugging excess kurtosis (0) would give1 − SR²/4, which is not it.
Convergences and contradictions
- Total convergence, zero genuine dispute. Paper equation, paper worked example, paper variable list, PSR structure, and every implementation checked all put the observed Sharpe in the denominator. The 2026-06-01 prose line is simply wrong; it is a transcription error from a secondary summary, not a defensible alternative convention.
- Three independent numerical reproductions confirm the reading, to four decimals. Using ŜR in the denominator: (1) the headline example returns 0.9003 against the paper's printed 0.9004 (the 0.0001 gap is the paper's own 4-dp rounding of ŜR₀); (2) the paper's claim that N=46 would have given 0.9505 reproduces as 0.9505 exactly; (3) the claim that Normal returns (γ̂₃=0, γ̂₄=3) give 0.9505 at N=88 reproduces as 0.9505 exactly. Reconstructing ŜR₀ from the paper's stated N=100 and V[{ŜR_n}]=1/(2·250) returns 0.1132, matching print. A misread equation does not reproduce three separate printed constants.
- The one thing the primary does NOT settle is the thing the 08-31 brief already settled. The paper says N is "the number of independent trials" and V[{ŜR_n}] is "the variance across the trials' estimated SR" — consistent with [[2026-08-31-dsr-effective-n-estimator-precommit]], and still silent on how to estimate effective-N. That supersession is out of scope here and is not touched by this correction.
Synthesis for RDCO
The answer is determinate, and the reason it is determinate matters more than the answer. The denominator of DSR is not a modelling choice, a house convention, or a conservatism dial. It is the estimated standard error of ŜR — the spread of the sampling distribution of the statistic we actually computed. A standard error is a property of an estimator, so it is evaluated at the estimate. ŜR₀ is a threshold: a number we are testing against, drawn from a different population entirely (the cross-section of trial Sharpes), and it has no business describing how noisy our one candidate's Sharpe is. The paper makes this structurally obvious by defining DSR ≡ PSR(ŜR₀): DSR is just the Probabilistic Sharpe Ratio with the user-chosen threshold set to the multiple-testing hurdle. In PSR the threshold slides freely in the numerator while the denominator stays put. Writing ŜR₀ into the denominator would make the standard error of a strategy's Sharpe depend on how many other strategies you happened to test, which is incoherent. So: prose line 33 of the 2026-06-01 brief is wrong, and the code comment and MinTRL line in the same file were right all along.
The error is small in magnitude, systematically anti-conservative, and lands exactly where it hurts. Direction first: with negative skew (γ̂₃ < 0), the radicand is 1 + |γ̂₃|·SR + …, which grows with the Sharpe plugged in. For any candidate that clears the hurdle (ŜR > ŜR₀), substituting the smaller ŜR₀ shrinks the denominator, inflates the z-score, and raises DSR. The wrong form is therefore biased toward passing candidates on precisely the return profile RDCO's surfaces exhibit — momentum and vol-overlay strategies with negative skew and fat tails. It flips sign only under positive skew, which is the case we care least about. On the paper's own example the inflation is +0.0123 (0.9003 → 0.9126). Across a realistic RDCO parameter box (N_eff 20–2000, T 500–2500, annualized trial-Sharpe dispersion 0.1–0.5, γ̂₃ down to −3, γ̂₄ up to 12) the maximum absolute DSR distortion is 0.0205. In DSR units that looks negligible. Stated in the units that actually govern a gate — the implied false-discovery probability 1 − DSR — it is a 12–14% understatement across every negative-skew scenario tested (e.g. N_eff=40, γ̂₃=−2.5, γ̂₄=10, T=1250: p = 0.0845 correct vs 0.0737 wrong, a 12.8% understatement; γ̂₄=12 pushes it to 14.2%).
And it does flip decisions, in a band you can hit. Holding N_eff=40, annualized trial-Sharpe variance 0.25, T=1250, γ̂₃=−2.5, γ̂₄=10 (so ŜR₀ ≈ 1.095 annualized), there is a live flip band at a 0.95 gate between annualized ŜR 1.900 and 1.945: at ŜR=1.925 the correct form returns DSR = 0.9457 (REJECT) and the wrong form returns 0.9510 (PASS). That is a ~2.4%-wide window of observed Sharpe in which the vault's prose formula admits a candidate the paper's formula refuses. GATE-1's entire job is to be hard to pass; a defect whose only effect is to make it easier to pass is the worst possible sign for the error, even at this magnitude. Coded once and run over hundreds of candidates, it is a slow leak of false positives, not a rounding artifact.
Two build consequences, both cheap to lock in now. First, validation.py should evaluate the bracket (1 − γ̂₃·ŜR + ((γ̂₄−1)/4)·ŜR²) once, at SR_obs, and reuse the identical expression in both deflated_sharpe and mintrl — the paper uses the same quantity in both, and sharing one helper makes the two lines incapable of drifting apart the way the vault's prose and code did. Second, and separately worth catching before it bites: γ̂₄ is RAW kurtosis, Normal = 3, not excess kurtosis. The 2026-06-01 build note says to source skew/kurtosis from stats.normality_report; if that function returns Fisher/excess kurtosis (SciPy's scipy.stats.kurtosis default is fisher=True), feeding it straight in silently subtracts 3 from γ̂₄ and shrinks the denominator by 0.75·SR² — again anti-conservative, again small (≈ +0.0014 DSR on the paper's example), again free to fix. Both belong in the same unit test: assert that the function reproduces the paper's published DSR = 0.9004 from (ŜR=2.5/√250, ŜR₀=0.1132, T=1250, γ̂₃=−3, γ̂₄=10), and that (γ̂₃=0, γ̂₄=3, N=88) reproduces 0.9505. Those two published constants are a free, primary-sourced regression test for the whole gate, and any of the errors described above breaks at least one of them.
Why this is in the vault
This closes open-follow-up #2 of [[2026-08-31-dsr-effective-n-estimator-precommit]] and repairs the specific line of [[2026-06-01-cpcv-deflated-sharpe-autoinv-validation]] that autoinv Phase-2's validation.py::deflated_sharpe is being coded from, so GATE-1 is not built on an anti-conservative denominator. It also supplies two published constants (DSR = 0.9004 and 0.9505) that serve as the primary-sourced unit test for that function.
Open follow-ups
- Does
autoinv'sstats.normality_reportreturn raw or excess kurtosis? The whole γ̂₄ convention risk resolves to one line of existing source in~/Projects/automated-investing/autoinv/stats.py, which this brief did not open. If it returns excess, either the caller adds 3 or the function grows afisher=flag. - Does the Lo/Mertens standard error in Eq. (2) assume iid returns, and if so what is the serial-correlation correction? Lo (2002) gives an autocorrelation-adjusted Sharpe standard error that Eq. (2) does not incorporate. For daily momentum and vol-overlay returns with non-trivial autocorrelation, the analytic denominator may be understated by more than the SR₀/SR_obs defect this brief fixes — which would make the 08-31 brief's
max(analytic, CPCV-path-dispersion)rule load-bearing rather than belt-and-braces. Worth reading Lo [2002] and Opdyke [2007] directly. - Is
V[{ŜR_n}]in ŜR₀ a variance of annualized or per-observation trial Sharpes? The paper's example carriesV[{ŜR_n}] = 1/(2·250), i.e. an annualized dispersion divided by observations-per-year, and then calls ŜR₀ "non-annualized (with 250 observations per year)". The unit convention is inferable from the example but never stated in words, and getting it wrong scales ŜR₀ by √250.
Related
- [[2026-06-01-cpcv-deflated-sharpe-autoinv-validation]] — the file corrected by this brief; DSR/SR₀/MinTRL formulas and the
validation.pysketch - [[2026-08-31-dsr-effective-n-estimator-precommit]] — flagged this discrepancy as its open-follow-up #2 and pre-registered the effective-N estimator; its
SR_obsusage is confirmed correct here - [[2026-07-05-autoinv-cpcv-vs-walk-forward-confirm-stage]] — the two-CV-regime ruling that supplies
SR_obsfrom the confirm stage - [[2026-05-29-strategy-pipeline-architecture-v0]] — GATE-0/1/2 design and the Phase-2 build this correction unblocks
- [[2026-08-31-sealed-holdout-refresh-policy]] — sibling Phase-2 pre-registration citing the same DSR machinery
- [[2026-05-31-ensembles-systematic-trading-overfitting]] — the False-Strategy Theorem framing behind ŜR₀
Sources
Vault:
~/rdco-vault/06-reference/research/2026-06-01-cpcv-deflated-sharpe-autoinv-validation.md— the defective prose line 33, the correct code comment line 83, the correct MinTRL line 34, and the "primary, cited not deep-read" admission~/rdco-vault/06-reference/research/2026-08-31-dsr-effective-n-estimator-precommit.md— the flag, the correct diagnosis, and the deferral this brief closes~/rdco-vault/06-reference/research/2026-07-05-autoinv-cpcv-vs-walk-forward-confirm-stage.md— CPCV confirm-stage role~/rdco-vault/01-projects/investing/2026-05-29-strategy-pipeline-architecture-v0.md— GATE-1 spec~/rdco-vault/06-reference/research/2026-08-31-sealed-holdout-refresh-policy.md— sibling Phase-2 pre-registration~/rdco-vault/06-reference/research/2026-05-31-ensembles-systematic-trading-overfitting.md— False-Strategy Theorem framing
Web:
- Bailey, D. H. & López de Prado, M., "The Deflated Sharpe Ratio: Correcting for Selection Bias, Backtest Overfitting and Non-Normality," Journal of Portfolio Management, 2014 — https://papers.ssrn.com/sol3/papers.cfm?abstract_id=2460551, PDF mirror https://www.davidhbailey.com/dhbpapers/deflated-sharpe.pdf. Primary, deep-read this time: full 22-page PDF retrieved and pages 8–10 read as rendered images to recover the math layout that text extraction drops. Eq. (2) p.8; variable definitions p.8–9; numerical example p.9–10 (N=100, V[{ŜR_n}]=1/(2·250), T=1250, γ̂₃=−3, γ̂₄=10, ŜR₀≈0.1132, DSR=0.9004; N=46→0.9505; γ̂₃=0, γ̂₄=3, N=88→0.9505).
- Portfolio Optimizer, "The Probabilistic Sharpe Ratio: Bias-Adjustment, Confidence Intervals, Hypothesis Testing and Minimum Track Record Length" — https://portfoliooptimizer.io/blog/the-probabilistic-sharpe-ratio-bias-adjustment-confidence-intervals-hypothesis-testing-and-minimum-track-record-length/ (denominator named as "the estimator of the standard error of ŜR"). Corroborates the SR_obs finding; its kurtosis-convention statement conflicts with the primary and the primary governs.
- rubenbriones/Probabilistic-Sharpe-Ratio — https://github.com/rubenbriones/Probabilistic-Sharpe-Ratio/blob/master/src/sharpe_ratio_stats.py (
sr_stdcomputed at the observedsr; algebraically identical to Eq. (2)). Cited from search-result extract, not opened. - Numerical reproductions in this brief were computed locally in Python from the paper's printed inputs; all three published constants reproduce to four decimal places.