06-reference/research

prompt over specification capable model degradation

2026-07-20·research-brief·source: deep-research·by Ray Data Co (deep-research synthesis)
promptingagent-harnesscafskillsevidence-audit

No Published Benchmark Tests the Over-Prescription Claim — the Nearest Evidence Is Mixed, and Two Studies Point the Other Way

The question

Is there published benchmarking that prompt over-specification degrades capable-model (Fable/Opus-tier) output versus objective-oriented prompts, and at what specificity threshold does it flip?

Context: the CAF Move 4 collapse thesis (106 prescriptive skills → ~15-20 objective skills) currently rests on a single vendor assertion in Anthropic's Fable 5 prompting guide. This brief tests whether independent evidence supports it.

What we already know (from the vault)

What the web says

Convergences and contradictions

Synthesis for RDCO

The direct answer is no, with a specificity threshold of "none published." There is no benchmark I can find that tests prescriptive-procedural prompts against objective-oriented prompts on Fable/Opus-tier models and reports a crossover point. I am not going to synthesize a number; anyone quoting one is extrapolating. The closest thing to a threshold in the literature is IFScale's ~150-250-instruction inflection for reasoning models, and that is measuring something materially different (keyword-constraint satisfaction) from what we mean.

The critical distinction the literature does not isolate — and this is why the evidence looks contradictory. Every measured study above varies task specificity: how much detail about what to produce. CAF Move 4 is about method prescription: how much detail about how to get there. Those are different variables, and the published work conflates or ignores the second. So the DETAIL and code-generation results do not straightforwardly refute the over-prescription thesis — they refute a neighboring claim ("more detail hurts capable models," which is false as stated). But they also mean we cannot cite them for us. The honest position is that the specific mechanism CAF Move 4 depends on — that a verbatim step-checklist amputates a capable model's ability to find a better route — is untested in public literature. It is plausible, it is vendor-asserted, it is corroborated by one named practitioner, and it is unmeasured.

Would this survive a skeptic asking "does this justify collapsing 106 skills to 15-20?" On over-prescription grounds alone: no. A skeptic would note (1) the sole primary source is the vendor whose guide we already caught over-claiming on this exact architecture, (2) the only studies isolating specificity found monotonic improvement, (3) IFScale shows frontier models at 98-100% adherence at 10-50 instructions, which is roughly the constraint load a single CAF leaf skill carries — so the instruction-density evidence does not bite at the per-skill level, only across a fully-loaded 106-skill chain, and (4) the 2026 skill-eval literature actively argues prescription is the tool for method conformance, which is exactly what a governance-bearing consulting framework needs. That fourth point is the sharpest one against us, and it is already half-conceded in the proposal's own DO-NOT-COLLAPSE list.

But the collapse itself is still the right call — on better-evidenced grounds we should switch to leading with. Three independent justifications survive this brief untouched: selection-surface quality (trigger collision across 106 similar descriptions is the measured constraint per [[2026-07-07-claude-skill-count-degradation-skill-packs]], and is the strongest evidence-backed argument we have), productization ([[2026-06-25-productize-framework-armstrong-vecteris]]), and maintainability (the v3→v4 migration debt — corrupted thresholds, stale phase numbers — scales with skill count and is documented fact, not inference). I recommend the CAF proposal and the Fable memo both be amended to demote over-prescription from the diagnosis to a hypothesis, and promote selection-surface + maintainability to the load-bearing argument. The collapse target of ~15-20 does not change; its justification gets more defensible.

Confidence: high that no published benchmark supports the specific claim (the search was targeted and the two most on-point papers both landed on the other side). Moderate that the claim is nonetheless directionally true for method-prescription specifically — the mechanism is plausible and Mollick's report is real evidence, just not measured evidence. High that the collapse decision is robust to this brief because it has independent support. The actionable gap: the Fable memo's own recommended per-skill default-vs-scaffolded A/B was never run, and it remains the cheapest way to convert this from vendor assertion to RDCO measurement. Running it on 2-3 CAF skills before the Phase-E full collapse would cost little and would be a genuinely publishable RDCO artifact, since nobody else has published it.

Why this is in the vault

This brief is the evidence audit for CAF Move 4 — the single largest architectural change in the phData CAF restructure ([[2026-06-09-caf-restructure-proposal]]) and the founder's main-bet client tooling. It establishes that the collapse's stated justification is vendor-only and unreplicated, identifies the better-evidenced justifications to lead with instead, and prevents the proposal from being defended in front of Andrew's team on a claim that does not survive a citation check.

Open follow-ups

  1. Run the never-executed A/B from [[2026-06-09-fable5-workflow-optimization-memo]] rec #2: take 2-3 real CAF leaf skills, run prescriptive-original vs objective-collapsed on the same engagement input, score with a fresh-eyes critic. This is the measurement that closes the gap — and it is publishable.
  2. Does method prescription (step-checklists) behave differently from task specification (requirement detail) on frontier models? No published study isolates this variable; it is the actual research gap and the actual thing CAF depends on.
  3. Does Anthropic publish any eval methodology behind the "too prescriptive" guidance, or is it practitioner distillation like the 20-50 skills number turned out to be? Worth asking directly via the partner-portal channel tied to the Claude Certified Architect track.
  4. At what constraint count per skill does a collapsed CAF objective skill start losing load-bearing rules? IFScale suggests frontier models hold to ~150; a collapsed phase skill aggregating a whole phase's PROHIBITED-from rules may approach that, which would be an argument for collapsing less aggressively than 15.
  5. Is there a measurable difference between "objective + acceptance test" and "objective + acceptance test + steps" — i.e. does keeping the steps as non-binding reference cost anything vs deleting them? This is the cheap middle path the proposal never considered.
  6. Does the 2026 agent-skill-eval framing ("does it solve it the way I want") imply CAF's governance/compliance skills should stay prescriptive by design rather than as a grudging DO-NOT-COLLAPSE exception?

Related

Sources

Vault

Web — tier (a), measured

Web — tier (b), vendor / framing

Web — tier (c), practitioner