06-reference/research

dsr denominator sr0 vs srobs resolution

2026-09-04·research-brief·source: deep-research·by Ray Data Co (deep-research synthesis)
deflated-sharpeprobabilistic-sharpebacktest-validationautomated-investingvault-correction

The DSR denominator is the standard error of the OBSERVED Sharpe: resolving the SR₀-vs-SR_obs defect in the 2026-06-01 brief

The question

"Resolve the DSR-denominator discrepancy: the 2026-06-01 CPCV brief writes SR₀ inside the standard-error term in prose but SR_obs in its own code comment (and in MinTRL). Confirm against SSRN 2460551 and correct the older brief."

Context: this is a live correctness defect in vault guidance that autoinv Phase 2 is being coded against. The two forms give a different Deflated Sharpe Ratio (DSR) on every candidate, so this is not a documentation nit — it is the gate arithmetic.

What we already know (from the vault)

What the web says

Convergences and contradictions

Synthesis for RDCO

The answer is determinate, and the reason it is determinate matters more than the answer. The denominator of DSR is not a modelling choice, a house convention, or a conservatism dial. It is the estimated standard error of ŜR — the spread of the sampling distribution of the statistic we actually computed. A standard error is a property of an estimator, so it is evaluated at the estimate. ŜR₀ is a threshold: a number we are testing against, drawn from a different population entirely (the cross-section of trial Sharpes), and it has no business describing how noisy our one candidate's Sharpe is. The paper makes this structurally obvious by defining DSR ≡ PSR(ŜR₀): DSR is just the Probabilistic Sharpe Ratio with the user-chosen threshold set to the multiple-testing hurdle. In PSR the threshold slides freely in the numerator while the denominator stays put. Writing ŜR₀ into the denominator would make the standard error of a strategy's Sharpe depend on how many other strategies you happened to test, which is incoherent. So: prose line 33 of the 2026-06-01 brief is wrong, and the code comment and MinTRL line in the same file were right all along.

The error is small in magnitude, systematically anti-conservative, and lands exactly where it hurts. Direction first: with negative skew (γ̂₃ < 0), the radicand is 1 + |γ̂₃|·SR + …, which grows with the Sharpe plugged in. For any candidate that clears the hurdle (ŜR > ŜR₀), substituting the smaller ŜR₀ shrinks the denominator, inflates the z-score, and raises DSR. The wrong form is therefore biased toward passing candidates on precisely the return profile RDCO's surfaces exhibit — momentum and vol-overlay strategies with negative skew and fat tails. It flips sign only under positive skew, which is the case we care least about. On the paper's own example the inflation is +0.0123 (0.9003 → 0.9126). Across a realistic RDCO parameter box (N_eff 20–2000, T 500–2500, annualized trial-Sharpe dispersion 0.1–0.5, γ̂₃ down to −3, γ̂₄ up to 12) the maximum absolute DSR distortion is 0.0205. In DSR units that looks negligible. Stated in the units that actually govern a gate — the implied false-discovery probability 1 − DSR — it is a 12–14% understatement across every negative-skew scenario tested (e.g. N_eff=40, γ̂₃=−2.5, γ̂₄=10, T=1250: p = 0.0845 correct vs 0.0737 wrong, a 12.8% understatement; γ̂₄=12 pushes it to 14.2%).

And it does flip decisions, in a band you can hit. Holding N_eff=40, annualized trial-Sharpe variance 0.25, T=1250, γ̂₃=−2.5, γ̂₄=10 (so ŜR₀ ≈ 1.095 annualized), there is a live flip band at a 0.95 gate between annualized ŜR 1.900 and 1.945: at ŜR=1.925 the correct form returns DSR = 0.9457 (REJECT) and the wrong form returns 0.9510 (PASS). That is a ~2.4%-wide window of observed Sharpe in which the vault's prose formula admits a candidate the paper's formula refuses. GATE-1's entire job is to be hard to pass; a defect whose only effect is to make it easier to pass is the worst possible sign for the error, even at this magnitude. Coded once and run over hundreds of candidates, it is a slow leak of false positives, not a rounding artifact.

Two build consequences, both cheap to lock in now. First, validation.py should evaluate the bracket (1 − γ̂₃·ŜR + ((γ̂₄−1)/4)·ŜR²) once, at SR_obs, and reuse the identical expression in both deflated_sharpe and mintrl — the paper uses the same quantity in both, and sharing one helper makes the two lines incapable of drifting apart the way the vault's prose and code did. Second, and separately worth catching before it bites: γ̂₄ is RAW kurtosis, Normal = 3, not excess kurtosis. The 2026-06-01 build note says to source skew/kurtosis from stats.normality_report; if that function returns Fisher/excess kurtosis (SciPy's scipy.stats.kurtosis default is fisher=True), feeding it straight in silently subtracts 3 from γ̂₄ and shrinks the denominator by 0.75·SR² — again anti-conservative, again small (≈ +0.0014 DSR on the paper's example), again free to fix. Both belong in the same unit test: assert that the function reproduces the paper's published DSR = 0.9004 from (ŜR=2.5/√250, ŜR₀=0.1132, T=1250, γ̂₃=−3, γ̂₄=10), and that (γ̂₃=0, γ̂₄=3, N=88) reproduces 0.9505. Those two published constants are a free, primary-sourced regression test for the whole gate, and any of the errors described above breaks at least one of them.

Why this is in the vault

This closes open-follow-up #2 of [[2026-08-31-dsr-effective-n-estimator-precommit]] and repairs the specific line of [[2026-06-01-cpcv-deflated-sharpe-autoinv-validation]] that autoinv Phase-2's validation.py::deflated_sharpe is being coded from, so GATE-1 is not built on an anti-conservative denominator. It also supplies two published constants (DSR = 0.9004 and 0.9505) that serve as the primary-sourced unit test for that function.

Open follow-ups

Related

Sources

Vault:

Web: