06-reference

alphasignal sonnet55 terminal bench beats opus5

2026-09-29·reference·source: AlphaSignal·by Lior Alexander
anthropicclaudesonnet-5-5terminal-benchopus5agentic-codingmodel-pricing

Why this is in the vault

AlphaSignal's Top News lead quantifies what Every's 2026-09-28 Sonnet 5.5 vibe check argued qualitatively: on Terminal-Bench 4.0, Sonnet 5.5 jumps from Sonnet 5's 10.3% to 70.6% — ahead of Opus 5.5's 66.4% — at the same price as Sonnet 5, using fewer tokens per task (up to 30% cheaper, 30% faster).

Mapping against Ray Data Co

Concrete connection: this very pipeline — the /process-newsletter sub-agent writing this note, the brigade stations, the deep-research one-sub-agent-per-question fan-out — runs on Claude Code agents doing exactly the class of work Terminal-Bench measures (autonomous command-line engineering tasks). AlphaSignal's headline number (Sonnet 5.5 beating Sonnet 5's best score "at low effort, for about one-tenth the cost") is a concrete, dated data point that the model underneath RDCO's own agentic-coding workloads just got materially cheaper and more capable without a price change — worth a deliberate check of whether any RDCO agent configs still pin to Sonnet 5 rather than picking up 5.5 automatically. It also sharpens the calibration point Every's vibe check already raised (2026-09-28-every-vibe-check-sonnet-5-5): that piece found Sonnet 5.5 earns its keep at low/medium effort under a steering hand but overbuilds at high effort or left unattended — AlphaSignal's Terminal-Bench framing is pure autonomous-task benchmark, no mention of that overbuild failure mode, so the vendor number and the hands-on report disagree on how much supervision the model actually needs. Second-order relevance: RDCO's phData credibility project (project_credibility_for_phdata_sales) is explicitly staked on "agents in production" as the domain — a same-price, lower-token, higher-Terminal-Bench-score model is exactly the kind of concrete capability delta worth citing there, not the marketing framing around it.

Curation section

⚠️ Sponsorship

Three distinct paid placements this issue:

Masthead slot this issue was not the usual unresolved "In Partnership with" image — it resolved to AlphaSignal's own promotion of attending a16z Speedrun (SF, 10/7/2026), framed as a community partnership (50 discounted tickets for AlphaSignal subscribers) rather than a paid third-party ad block. Treating this as house/event self-promotion, not a new pool sponsor, since no advertiser paid for placement — flagging the ambiguity rather than silently dropping it.

Bias implication: standard AlphaSignal rotating-pool pattern — treat all three sponsor blocks as commercially motivated by default. The Sonnet 5.5, GPT-6 patch, and Grok Team Bots items are first-party vendor news, not sponsored, but AlphaSignal's own Terminal-Bench framing should be read alongside Every's more skeptical hands-on take (see mapping) rather than taken at face value alone.

Related