Why this is in the vault
AlphaSignal's Top News lead quantifies what Every's 2026-09-28 Sonnet 5.5 vibe check argued qualitatively: on Terminal-Bench 4.0, Sonnet 5.5 jumps from Sonnet 5's 10.3% to 70.6% — ahead of Opus 5.5's 66.4% — at the same price as Sonnet 5, using fewer tokens per task (up to 30% cheaper, 30% faster).
Mapping against Ray Data Co
Concrete connection: this very pipeline — the /process-newsletter sub-agent writing this note, the brigade stations, the deep-research one-sub-agent-per-question fan-out — runs on Claude Code agents doing exactly the class of work Terminal-Bench measures (autonomous command-line engineering tasks). AlphaSignal's headline number (Sonnet 5.5 beating Sonnet 5's best score "at low effort, for about one-tenth the cost") is a concrete, dated data point that the model underneath RDCO's own agentic-coding workloads just got materially cheaper and more capable without a price change — worth a deliberate check of whether any RDCO agent configs still pin to Sonnet 5 rather than picking up 5.5 automatically. It also sharpens the calibration point Every's vibe check already raised (2026-09-28-every-vibe-check-sonnet-5-5): that piece found Sonnet 5.5 earns its keep at low/medium effort under a steering hand but overbuilds at high effort or left unattended — AlphaSignal's Terminal-Bench framing is pure autonomous-task benchmark, no mention of that overbuild failure mode, so the vendor number and the hands-on report disagree on how much supervision the model actually needs. Second-order relevance: RDCO's phData credibility project (project_credibility_for_phdata_sales) is explicitly staked on "agents in production" as the domain — a same-price, lower-token, higher-Terminal-Bench-score model is exactly the kind of concrete capability delta worth citing there, not the marketing framing around it.
Curation section
- Anthropic releases Claude Sonnet 5.5 (66,454 likes) — same price as Sonnet 5, up to 30% cheaper per task and 30% faster from lower token usage. Terminal-Bench 4.0: 70.6% (Sonnet 5 was 10.3%, Opus 5.5 is 66.4%). 1M token context with 128K output. Positioned for agentic coding (fewer steps/tool calls), knowledge work (slides/docs/sheets needing minimal editing), and drop-in replacement for Sonnet 5 everywhere it's already deployed. See mapping above.
- OpenAI patches a GPT-6 image-understanding bug (6,198 likes) — silent fix to GPT-6 Sol and Luna, which were misreading visual inputs across the API, Codex, and computer-use tasks. No new feature; AlphaSignal's advice is to rerun any image-pipeline tests since prior "worse" results may have been this bug, not the model.
- xAI launches Grok Team Bots (4,766 likes) — shared (not per-user) bots configured once and used by a whole team, each person keeps a private conversation thread, shared skills/knowledge across the group. Slack or Grok Bot integration, pre-built role templates (sales/marketing/product/data), Teams/Enterprise plans only.
- Signals list (lower-signal, not individually mapped): Typesafe AI's CLI claims 40% lower coding-agent cost on SWE-bench; a free 4-month DataTalks Club ML engineering course; a Google autonomous research agent reportedly outperforming human papers on 86 tasks by 25%; IST-DASLab pruning a 354GB coding model down to 58GB by cutting half its experts; AMD's acquisition of World Labs, making Fei-Fei Li its chief scientist.
⚠️ Sponsorship
Three distinct paid placements this issue:
- AgentField AI — standalone "Presented by" block for CodeAF, an open-source coding harness (Apache 2.0, one Go binary) built for open models (DeepSeek, Qwen, GLM, Kimi), claiming nearly 4x Claude Code's solve rate on DeepSWE at roughly half the per-solve cost of the next-best harness on the same open model. NEW — not in the previously tracked 27+ member pool (OpenRouter, Teleport, Granola, Unblocked, Datadog, Vanta, Finest, Google Cloud, Tiger Data, Redis, Attio, Flint AI, QA.tech, Ory, LaunchDarkly, Voices, Span, Wiz, WorkOS, Encord, AI Conference, Nyra, pre.dev, Prior Labs, Origin Technology, Sentry, The Linux Foundation). Flagging for the README's running pool count — rotating pool is now 28+.
- Prior Labs — standalone "Presented by" block for TabPFN-3.5 (tabular prediction API). Recurring — first logged 2026-09-25, confirmed pool member.
- Google Cloud — native ad inside the Signals list (item #2: NVIDIA RTX PRO GPU-accelerated agents, multi-agent drug-discovery demo). Recurring — confirmed pool member since 2026-09-04.
Masthead slot this issue was not the usual unresolved "In Partnership with" image — it resolved to AlphaSignal's own promotion of attending a16z Speedrun (SF, 10/7/2026), framed as a community partnership (50 discounted tickets for AlphaSignal subscribers) rather than a paid third-party ad block. Treating this as house/event self-promotion, not a new pool sponsor, since no advertiser paid for placement — flagging the ambiguity rather than silently dropping it.
Bias implication: standard AlphaSignal rotating-pool pattern — treat all three sponsor blocks as commercially motivated by default. The Sonnet 5.5, GPT-6 patch, and Grok Team Bots items are first-party vendor news, not sponsored, but AlphaSignal's own Terminal-Bench framing should be read alongside Every's more skeptical hands-on take (see mapping) rather than taken at face value alone.
Related
- [[2026-09-28-every-vibe-check-sonnet-5-5]]
- [[2026-09-23-alphasignal-opus55-gpt6-sol-luna-pricing]]
- [[project_credibility_for_phdata_sales]]
- [[project_l5_north_star_strategic_direction]]