06-reference

alphasignal agent skills 8100 trial study

2026-08-23·reference·source: AlphaSignal·by Ben Dickson
agent-skillsharness-engineeringprocedural-anchoringskill-retrievalempirical-study

AlphaSignal — What 8,100 trial records reveal about AI agent skills (Aug 23 2026)

Why this is in the vault

A multi-university study (8,135 trial records across Terminal-Bench 2.0 and SkillsBench) gives empirical backing to claims RDCO has operated on by intuition: skills work as procedural anchors, not knowledge injection, and skill-catalog scale creates a retrieval precision cliff RDCO hasn't yet gated against.

Mapping against Ray Data Co

This directly validates the "Skills over commands" standing rule ([[feedback_skills_over_commands]]) and the entire ~/.claude/skills/ architecture — but it also exposes a live gap: RDCO's skill catalog has grown past 80 named skills (visible in this session's own tool listing), and the study measures that as exactly the zone where retrieval precision collapses from 29.6% to 3.3% as a pool grows from 5 to 100 options. RDCO has no domain-bucket-first, strict-trigger-second gating layer — skills are flat-listed and matched by description alone, the precise "semantic confusability" failure mode the researchers warn against. The study's reassurance (task success held near 36-39% even as precision cratered, because near-miss skills still supply partial procedural support) is not a reason to ignore this — it's a reason RDCO hasn't noticed the degradation yet, per [[research/2026-07-07-claude-skill-count-degradation-skill-packs]], which already flagged trigger collision (not token budget) as the real ceiling at 60+ skills. This new data adds a concrete number to that open question.

Second connection: the "runbook vs. tutorial" finding — skills should be step-by-step procedural checklists, not fact repositories, with procedural anchoring accounting for 65.7% of successful skill cases vs. 4.5% for knowledge injection — is the implicit design principle behind /skillify ([[2026-04-22-garry-tan-skillify-it-workflow]]) and the implementation-notes sub-agent pattern ([[feedback_implementation_notes_sub_agent_pattern]]). Both already produce procedural runbooks rather than explanatory docs; this study is the first hard evidence that instinct was correct, not just stylistically preferred.

Third: the "no-hint ablation" finding (74.6% success with success/failure annotations visible during skill distillation vs. 40.0% without) is a direct argument for annotating outcomes when any RDCO workflow gets skillified from raw run logs — currently an unstated convention, not an enforced one.

The core argument

Researchers ran controlled comparisons (raw execution vs. raw workflow memory vs. distilled skill) on the same prior experience across Terminal-Bench 2.0 and SkillsBench. Distilled, standardized skills beat raw workflow-memory injection by 6.06 percentage points — format of the experience matters as much as having it. Skills reduced environment-infrastructure failures from 5.3% (raw execution) to 0.2% (distilled skill), by acting as strict execution anchors rather than teaching new facts: one case example converted serialized await calls to Promise.all per the skill's runbook and passed a latency check the raw agent had failed.

The retrieval-scale finding is the sharper warning: as a skill pool grows 5→100, ground-truth retrieval precision falls from 29.6% to 3.3%, yet downstream task success stays roughly flat (~36-39%) because related, non-ground-truth skills still supply partial guidance — "exact ground-truth invocation is neither sufficient nor necessary for success." The real danger named is semantic confusability among skills that look identical in embedding space. Recommended fix: two-level gating — an LLM router assigns a domain bucket first, then a rigid trigger condition (e.g., an exact error string) selects within the bucket, rather than a flat vector search across the whole catalog.

⚠️ Sponsorship

OpenRouter (unified multi-provider LLM routing API) sponsored this issue via a "From OpenRouter" callout box plus a closing "One API for every LLM" CTA. Standard AlphaSignal single-sponsor-slot placement; no evidence the sponsor influenced the editorial content or study selection — the deep-dive is sourced to an independent academic study, not an OpenRouter-commissioned one. Flagging per whitelist discipline: no prior AlphaSignal sponsor has been logged in this file's audit trail before, so this is the first entry.

Related