No Published Benchmark Tests the Over-Prescription Claim — the Nearest Evidence Is Mixed, and Two Studies Point the Other Way
The question
Is there published benchmarking that prompt over-specification degrades capable-model (Fable/Opus-tier) output versus objective-oriented prompts, and at what specificity threshold does it flip?
Context: the CAF Move 4 collapse thesis (106 prescriptive skills → ~15-20 objective skills) currently rests on a single vendor assertion in Anthropic's Fable 5 prompting guide. This brief tests whether independent evidence supports it.
What we already know (from the vault)
- The thesis and its single load-bearing citation. [[2026-06-09-caf-restructure-proposal]] Move 4 states the diagnosis verbatim: "106 multi-hundred-line micro-skills is the 'written for older models, too prescriptive' pattern. Fable's prompting guidance is that over-prescription degrades capable-model output." The entire evidentiary chain terminates at one vendor doc.
- The memo that carried it is explicit about its own sourcing. [[2026-06-09-fable5-workflow-optimization-memo]] rec #2 cites the Anthropic Fable 5 prompting guide ("skills written for older models are often too prescriptive and degrade Fable output") and Ethan Mollick's hands-on report ("I no longer steer; I commission"). The memo labels Mollick "credible-practitioner tier, mild Anthropic-positive bias" — i.e. it never claimed measured benchmarking. Notably, the memo's own recommendation was to test default-vs-scaffolded per skill and "keep simpler only if genuinely as good." That A/B was never run.
- We have already caught ourselves extrapolating a vendor heuristic once on this exact architecture. [[2026-07-07-claude-skill-count-degradation-skill-packs]] found Anthropic's "20-50 simultaneous skills degrades quality" warning was a rule of thumb borrowed from the tool-count literature, not a measured skills threshold. Same guide, same architecture, same failure mode. That precedent is directly cautionary here.
- An independent-of-Anthropic strand does support periodic de-scaffolding. [[2026-04-12-cross-check-agent-architecture]] recommends a quarterly "scaffolding audit" testing whether removing skill steps improves or degrades output — framed as a hypothesis to test, not a settled result.
- The collapse has a second, unrelated justification already in the vault. [[2026-06-25-productize-framework-armstrong-vecteris]] argues the 106-skill CAF is "service trapped as tacit expertise" and should be collapsed for productization reasons. This does not depend on the over-prescription claim at all.
What the web says
- (a) Measured — instruction count degrades adherence, but frontier models hold near-perfect far past CAF-relevant densities. IFScale (arXiv 2507.11538) benchmarks 20 models on 10-500 keyword-inclusion instructions. Reasoning models show "threshold decay" — near-perfect until ~150-250 instructions, then steep decline (o3: 97.8% at 250 → 62.8% at 500). At 10 instructions: 98-100%. At 50: 98.8-99.6%. The authors caveat that keyword-inclusion "differs structurally from typical agent prompts requiring reasoning, tool use, or conditional logic."
- (a) Measured — corroborating degradation-with-count from two independent groups. ManyIFEval/StyleMBPP (arXiv 2509.21051, 10 LLMs, up to 10 instructions) and SCALEDIF (arXiv 2510.14842) both find performance degrades monotonically as instructions are added, with inter-instruction conflict named as a key driver.
- (a) Measured — and this one CONTRADICTS the thesis. The DETAIL framework (arXiv 2512.02246) explicitly manipulates prompt specificity (vague / moderate / detailed) on GPT-4 and o3-mini. Result: monotonic improvement with more specificity, no degradation threshold found. GPT-4 baseline 0.60 (vague) → 0.83 (detailed). Crucially for our question, the more capable model degraded less from vagueness (GPT-4 0.60 vs o3-mini 0.34), which is the opposite of "capable models want less specification." Limitations are serious: 30 tasks, 2 models, no confidence intervals.
- (a) Measured — also cuts against, from the code-generation side. arXiv 2604.24712 mutates prompts across 10 models spanning three capability tiers (incl. Claude Sonnet 4). Under-specification cost -11.8% Pass@1 on minimally-specified HumanEval but only -0.9% on richly-specified LiveCodeBench — the authors read this as structural redundancy in multi-layered specs providing robustness. They do note the near-parity masks opposing forces (rich specs can "silently mislead toward memorized patterns"), but find no flip threshold. Self-labeled "hypothesis-generating rather than hypothesis-confirming."
- (b) Vendor — the origin of our claim, and it stands alone. Anthropic's Fable 5 prompting guide asserts skills written for older models are "often too prescriptive" and degrade output. No public methodology, eval set, or effect size accompanies it. This is the same document family that produced the 20-50 skills heuristic we already debunked.
- (c) Practitioner — directionally supportive, non-quantified. Mollick's "I no longer steer; I commission" report and a DEV Community writeup (dev.to) describing an internal three-style comparison where a concise prompt beat an over-specified one. No public data, n, or task set in either.
- (b/c) Countervailing framing from the 2026 agent-skill-eval literature. Survey work (arXiv 2606.11435) frames the 2026 question as "not can the agent solve this task, but does it solve it the way I want" — with opinionated, prescriptive instructions in skills as the mechanism for method control. On this framing prescription is a feature for compliance-shaped work, not legacy debt.
Convergences and contradictions
- Convergence: vault and web agree that instruction count is a real degradation axis, and that inter-instruction conflict (not verbosity per se) is the mechanism. The vault's [[2026-07-07-claude-skill-count-degradation-skill-packs]] finding that the binding constraint is selection quality over an overlapping menu is the same mechanism one level up.
- Contradiction — and it is the headline. The vault asserts (via Anthropic) that more specification degrades capable-model output. The two studies that directly manipulate specificity as the independent variable (DETAIL; the under-specification code study) find the opposite or find null. No published benchmark reproduces Anthropic's claim.
- Convergence on the meta-pattern: we have now twice found that a confident quantitative-sounding assertion in this vendor guide is a heuristic without published measurement behind it. The 20-50 skills precedent and this one are the same error class.
Synthesis for RDCO
The direct answer is no, with a specificity threshold of "none published." There is no benchmark I can find that tests prescriptive-procedural prompts against objective-oriented prompts on Fable/Opus-tier models and reports a crossover point. I am not going to synthesize a number; anyone quoting one is extrapolating. The closest thing to a threshold in the literature is IFScale's ~150-250-instruction inflection for reasoning models, and that is measuring something materially different (keyword-constraint satisfaction) from what we mean.
The critical distinction the literature does not isolate — and this is why the evidence looks contradictory. Every measured study above varies task specificity: how much detail about what to produce. CAF Move 4 is about method prescription: how much detail about how to get there. Those are different variables, and the published work conflates or ignores the second. So the DETAIL and code-generation results do not straightforwardly refute the over-prescription thesis — they refute a neighboring claim ("more detail hurts capable models," which is false as stated). But they also mean we cannot cite them for us. The honest position is that the specific mechanism CAF Move 4 depends on — that a verbatim step-checklist amputates a capable model's ability to find a better route — is untested in public literature. It is plausible, it is vendor-asserted, it is corroborated by one named practitioner, and it is unmeasured.
Would this survive a skeptic asking "does this justify collapsing 106 skills to 15-20?" On over-prescription grounds alone: no. A skeptic would note (1) the sole primary source is the vendor whose guide we already caught over-claiming on this exact architecture, (2) the only studies isolating specificity found monotonic improvement, (3) IFScale shows frontier models at 98-100% adherence at 10-50 instructions, which is roughly the constraint load a single CAF leaf skill carries — so the instruction-density evidence does not bite at the per-skill level, only across a fully-loaded 106-skill chain, and (4) the 2026 skill-eval literature actively argues prescription is the tool for method conformance, which is exactly what a governance-bearing consulting framework needs. That fourth point is the sharpest one against us, and it is already half-conceded in the proposal's own DO-NOT-COLLAPSE list.
But the collapse itself is still the right call — on better-evidenced grounds we should switch to leading with. Three independent justifications survive this brief untouched: selection-surface quality (trigger collision across 106 similar descriptions is the measured constraint per [[2026-07-07-claude-skill-count-degradation-skill-packs]], and is the strongest evidence-backed argument we have), productization ([[2026-06-25-productize-framework-armstrong-vecteris]]), and maintainability (the v3→v4 migration debt — corrupted thresholds, stale phase numbers — scales with skill count and is documented fact, not inference). I recommend the CAF proposal and the Fable memo both be amended to demote over-prescription from the diagnosis to a hypothesis, and promote selection-surface + maintainability to the load-bearing argument. The collapse target of ~15-20 does not change; its justification gets more defensible.
Confidence: high that no published benchmark supports the specific claim (the search was targeted and the two most on-point papers both landed on the other side). Moderate that the claim is nonetheless directionally true for method-prescription specifically — the mechanism is plausible and Mollick's report is real evidence, just not measured evidence. High that the collapse decision is robust to this brief because it has independent support. The actionable gap: the Fable memo's own recommended per-skill default-vs-scaffolded A/B was never run, and it remains the cheapest way to convert this from vendor assertion to RDCO measurement. Running it on 2-3 CAF skills before the Phase-E full collapse would cost little and would be a genuinely publishable RDCO artifact, since nobody else has published it.
Why this is in the vault
This brief is the evidence audit for CAF Move 4 — the single largest architectural change in the phData CAF restructure ([[2026-06-09-caf-restructure-proposal]]) and the founder's main-bet client tooling. It establishes that the collapse's stated justification is vendor-only and unreplicated, identifies the better-evidenced justifications to lead with instead, and prevents the proposal from being defended in front of Andrew's team on a claim that does not survive a citation check.
Open follow-ups
- Run the never-executed A/B from [[2026-06-09-fable5-workflow-optimization-memo]] rec #2: take 2-3 real CAF leaf skills, run prescriptive-original vs objective-collapsed on the same engagement input, score with a fresh-eyes critic. This is the measurement that closes the gap — and it is publishable.
- Does method prescription (step-checklists) behave differently from task specification (requirement detail) on frontier models? No published study isolates this variable; it is the actual research gap and the actual thing CAF depends on.
- Does Anthropic publish any eval methodology behind the "too prescriptive" guidance, or is it practitioner distillation like the 20-50 skills number turned out to be? Worth asking directly via the partner-portal channel tied to the Claude Certified Architect track.
- At what constraint count per skill does a collapsed CAF objective skill start losing load-bearing rules? IFScale suggests frontier models hold to ~150; a collapsed phase skill aggregating a whole phase's PROHIBITED-from rules may approach that, which would be an argument for collapsing less aggressively than 15.
- Is there a measurable difference between "objective + acceptance test" and "objective + acceptance test + steps" — i.e. does keeping the steps as non-binding reference cost anything vs deleting them? This is the cheap middle path the proposal never considered.
- Does the 2026 agent-skill-eval framing ("does it solve it the way I want") imply CAF's governance/compliance skills should stay prescriptive by design rather than as a grudging DO-NOT-COLLAPSE exception?
Related
- [[2026-06-09-caf-restructure-proposal]] — Move 4 is the decision this brief audits; its over-prescription diagnosis needs amending
- [[2026-06-09-fable5-workflow-optimization-memo]] — the source of the vendor claim, and the home of the never-run A/B this brief revives
- [[2026-07-07-claude-skill-count-degradation-skill-packs]] — direct precedent: the same vendor guide's 20-50 heuristic was also unmeasured; also supplies the better-evidenced selection-surface argument
- [[2026-04-12-cross-check-agent-architecture]] — the independent "quarterly scaffolding audit" recommendation, framed correctly as a hypothesis to test
- [[2026-06-25-productize-framework-armstrong-vecteris]] — the productization justification for the collapse that survives independent of over-prescription
- [[2026-04-11-garry-tan-thin-harness-fat-skills]] — the thin-harness/fat-skills frame the collapse sits inside
- [[2026-04-15-thariq-claude-code-session-management-1m-context]] — context-rot as the parent mechanism for count-driven degradation
Sources
Vault
~/rdco-vault/01-projects/phdata/2026-06-09-caf-restructure-proposal.md~/rdco-vault/08-tooling/2026-06-09-fable5-workflow-optimization-memo.md~/rdco-vault/06-reference/research/2026-07-07-claude-skill-count-degradation-skill-packs.md~/rdco-vault/06-reference/cross-checks/2026-04-12-cross-check-agent-architecture.md~/rdco-vault/06-reference/2026-06-25-productize-framework-armstrong-vecteris.md
Web — tier (a), measured
- How Many Instructions Can LLMs Follow at Once? (IFScale) — https://arxiv.org/html/2507.11538v1
- When Instructions Multiply: Measuring and Estimating LLM Capabilities of Multiple Instructions Following — https://arxiv.org/abs/2509.21051
- Boosting Instruction Following at Scale (SCALEDIF) — https://arxiv.org/html/2510.14842v1
- DETAIL Matters: Measuring the Impact of Prompt Specificity on Reasoning in LLMs — https://arxiv.org/html/2512.02246v1
- When Prompt Under-Specification Improves Code Correctness — https://arxiv.org/html/2604.24712v1
Web — tier (b), vendor / framing
- Anthropic Fable 5 prompting guide (platform.claude.com) — cited via vault memo, not re-fetched
- Agent Skill Evaluation and Evolution: Frameworks and Benchmarks — https://arxiv.org/html/2606.11435v1
Web — tier (c), practitioner
- Prompt Complexity vs Output Quality: When More Instructions Hurt Performance — https://dev.to/jasrandhawa/prompt-complexity-vs-output-quality-when-more-instructions-hurt-performance-2hi5
- Ethan Mollick, oneusefulthing.org — cited via vault memo, not re-fetched