Why this is in the vault
This is Anthropic's own write-up of a build-an-eval-then-climb-it loop. The founder tied it to the "targeting system" in [[book-solve-everything-ch6-the-engine-2026-04-13]]: measured targets plus a feedback loop, at the scale of a single agent. It also fits the "agents in production" credibility lane ([[project_credibility_for_phdata_sales]]).
What it says
The method is two stages, packaged as two new claude-api skill commands.
- build-eval gathers test cases in priority order: production transcripts, then bug reports and support tickets, then 5-10 hand-written cases, then synthetic cases generated from the codebase. Grading is programmatic (exact match or schema) or large-language-model-as-judge (LLM-as-judge) against checkable rubrics, not numeric scales.
- hillclimb has Claude propose one targeted patch per round. A fixed model and harness score each patch on a train set and a held-out test set. The loop keeps a patch only if both sets improve. It reverts on a train-only gain or any regression, and it flags suspected overfitting. It stops after 2-3 stalled rounds (which triggers root-cause analysis). Separately, before the first round, it checks whether the eval's noise floor is smaller than the smallest improvement worth acting on, and suggests more repetitions or cases if not.
Reported results:
- Customer-support benchmark: on the 14 held-out tickets never used for hillclimbing, accuracy went from 78.6% to 90.5% at about one-fifth of the original cost. On the 30 training tickets, accuracy went from 74.4% (4.6¢ per ticket), through an Opus 5.5 step at 87.8% (1.9¢), to a final train-set score of 98.9%.
- Claude-API skill task: pass rate went from 66% to about 88% after 24 rounds.
The authors give no confidence-interval values or sample sizes for the headline numbers, though the method includes an explicit noise-floor check before the hillclimb starts.
Failure modes the authors name, with their fixes (only Overfitting is its own section heading; the rest come from the article's "Validating the grader" and "Diagnostic checks" sections):
- Overfitting to eval quirks. Use a train/test split, and never paste failures into prompts.
- Reward hacking. Keep answers structurally out of the model's reach.
- Grader drift. Grade each output twice, and validate against sample transcripts.
- Ambiguous tasks. Require two domain experts to agree before trusting a case.
- Also flagged: infrastructure noise (timeouts, API errors, cut-off answers) and environmental contamination (leftover state from an earlier trial).
⚠️ Bias
The post is vendor marketing: it promotes /claude-api build-eval and /claude-api hillclimb. Both reported results are single illustrative examples from Anthropic's own use, not a benchmark suite, and there is no independent replication. Read the support numbers carefully: there is a train-set track (74.4% → 98.9%) and a held-out track (78.6% → 90.5%). Mixing the two overstates generalization, which an earlier draft of this note did.
Mapping against Ray Data Co
- Targeting system, agent scale. In Solve Everything ch 6, the targeting system is blinded evals plus decision records plus red-teaming. It turns "is the AI good?" into a falsifiable claim. build-eval is the target, and the keep/revert loop is the feedback. The accuracy-up, cost-down result has the same shape as ch 6's Return on Cognitive Spend.
- Our critic chain already does keep/revert by hand. The studio gates, verify-* skills and critic caps are exactly this loop. What's missing is a held-out set and a stall rule. Open question: does the studio charter's 3-iteration critic cap (and the 2-round escalation rule in verify-vault-write) stop at the right point, or should the stop rule be "no gain on held-out cases"?
- phData lane. This loop pairs with Snowflake AI Observability and the Retrieval-Augmented Generation (RAG) triad (context relevance, groundedness, answer relevance), which the founder asked about after passing GES-C02 ([[founder-exam-experience-2026-09-28]]). The observability side is the grader, and this is the climber. Candidate post and client talk: "how to hillclimb a Cortex agent."
- Offered 2026-09-28 (awaiting founder): a real trial of build-eval + hillclimb on one of our own agents, for example the community-watch drafter.
Related
- [[book-solve-everything-ch6-the-engine-2026-04-13]]
- [[2026-09-28-data-engineering-weekly-289-jev-pipelines-data-agents]]
- [[project_credibility_for_phdata_sales]]