06-reference

Automating eval design and hillclimbing with Claude

2026-09-28·reference·status: gate-fixes-applied (verify-vault-write ITERATE → 7 fixes)·source: https://claude.dev/blog/automating-eval-design-and-hillclimbing/·by Lance Martin (claude.dev blog)
evalshillclimbingharness-engineeringagents-in-productionanthropic

Why this is in the vault

This is Anthropic's own write-up of a build-an-eval-then-climb-it loop. The founder tied it to the "targeting system" in [[book-solve-everything-ch6-the-engine-2026-04-13]]: measured targets plus a feedback loop, at the scale of a single agent. It also fits the "agents in production" credibility lane ([[project_credibility_for_phdata_sales]]).

What it says

The method is two stages, packaged as two new claude-api skill commands.

  1. build-eval gathers test cases in priority order: production transcripts, then bug reports and support tickets, then 5-10 hand-written cases, then synthetic cases generated from the codebase. Grading is programmatic (exact match or schema) or large-language-model-as-judge (LLM-as-judge) against checkable rubrics, not numeric scales.
  2. hillclimb has Claude propose one targeted patch per round. A fixed model and harness score each patch on a train set and a held-out test set. The loop keeps a patch only if both sets improve. It reverts on a train-only gain or any regression, and it flags suspected overfitting. It stops after 2-3 stalled rounds (which triggers root-cause analysis). Separately, before the first round, it checks whether the eval's noise floor is smaller than the smallest improvement worth acting on, and suggests more repetitions or cases if not.

Reported results:

The authors give no confidence-interval values or sample sizes for the headline numbers, though the method includes an explicit noise-floor check before the hillclimb starts.

Failure modes the authors name, with their fixes (only Overfitting is its own section heading; the rest come from the article's "Validating the grader" and "Diagnostic checks" sections):

⚠️ Bias

The post is vendor marketing: it promotes /claude-api build-eval and /claude-api hillclimb. Both reported results are single illustrative examples from Anthropic's own use, not a benchmark suite, and there is no independent replication. Read the support numbers carefully: there is a train-set track (74.4% → 98.9%) and a held-out track (78.6% → 90.5%). Mixing the two overstates generalization, which an earlier draft of this note did.

Mapping against Ray Data Co

Related