Why this is in the vault
Every's same-day head-to-head of OpenAI's GPT-6 Sol and Anthropic's Claude Opus 5.5 — Dan Shipper's verdict is a split by task type, not a single winner, and it's the latest data point in Every's running Vibe Check series tracking which model wins which job.
The core argument
The newsletter itself was a paywalled preview (only the intro rendered — "Two big models dropped today," teasing the comparison without the verdict). The actual assessment came through cleanly in Shipper's companion X thread, which carries the same content: Sol is now his "new daily driver in Codex... faster, and 50% cheaper than 5.6 Sol," while Opus 5.5 was "the bigger surprise" of the two releases.
Shipper's per-category breakdown:
- Writing: Sol scored close to Astra (still Every's top writing model) on a real paragraph-writing task, producing clean prose that leads with the point. Opus 5.5 was "pleasant to work with" but its drafts "tend to bury the point."
- Computer use: Sol tracks Astra closely — an earlier preview matched Astra on 17 of 18 attempts across simpler Hands tasks, at much lower token cost.
- Coding: Sol improves on GPT-5.6, but Opus 5.5 "has the higher ceiling for long, autonomous builds."
- Friction: Codex's security classifier kept interrupting already-authorized long runs to re-ask for approval — a harness issue, not a Sol capability gap, but one that made autonomous runs harder to leave alone.
Overall framing: Sol is "an S-class iPhone release" — most of Astra's power at roughly a fifth of the price, best for people who live in Codex reading, writing, and shipping day to day. Opus 5.5 is the one worth handing a hard, open-ended coding or visual project to see how far it can run — strong enough that some of Every's Codex converts are wobbling back toward Claude.
Mapping against Ray Data Co
This directly validates RDCO's existing per-station model-effort pairing pattern (the skill-agent-brigade's station-code-author/sw-builder split, and the Delegation memory rule that Fable delegations run at high/xhigh effort for meaningfully large tasks only): the same segmentation Shipper describes — a fast/cheap model for high-volume daily read-write work, a higher-ceiling model reserved for long autonomous builds — is exactly the logic behind routing fanout-style extraction to a cheap low-effort model while reserving Fable-high/xhigh for real build tasks. It's also a live argument against picking one "house model" for everything: RDCO's own harness (Scribble Works' AI Gateway model routing, Claude Code station dispatch) is built around task-conditional model choice, not a single default, which this Vibe Check reinforces from the outside.
Related
- [[2026-09-03-every-gpt6-astra-vibe-check]]
- [[2026-07-09-every-vibe-check-gpt56-sol]]
- [[2026-07-24-every-vibe-check-opus5]]