Why this is in the vault
Every's Dan Shipper spent days testing the biggest of OpenAI's 20+ DevDay 2026 launches — his verdict is that OpenAI's plan is now plain (ChatGPT as the operating system for work, with agents, documents, and other companies' apps all running inside it), but the flagship always-on agent still has rough reliability edges.
The core argument
OpenAI announced more than 20 products and features at DevDay 2026 — by Shipper's count, the biggest DevDay since the first one in 2023. Four launches anchor the review (full comparison, plugin extensions, the $500 Pro plan, and Sol 6.1 detail sit behind Every's paywall):
- Dots, OpenAI's always-on agent, became Shipper's main way to use ChatGPT — it caught a flight change that would clash with a meeting request buried in an unread Slack thread. But it also dropped messages, tripped over permissions, and often couldn't reach its own browser.
- Space, OpenAI's new home for docs/slides/spreadsheets, showed edits appearing faster than an agent working in Google Docs, and lets a user tag their dot directly in a comment — Shipper thinks it makes prying a team off Google or Notion harder to resist.
- The Decisions API, pitched as OpenAI's rival to TypeSafe's Jev, split Every's testers: on senior editor Jack Cheng's replay of computer-use tasks it picked correctly on 76 of 78 steps (Jev: 73) in 230ms (Jev: 500ms); on Cora GM Kieran Klaassen's thread-sorting test the two tied on accuracy, with Jev about twice as fast. No pricing announced yet.
- Plus/Pro plans now work inside roughly 15 partner apps, including Devin, with usage counting against existing plan limits.
Shipper's bottom line: try Dots now if you like new tools with rough edges, otherwise give OpenAI a week or two to fix the paper cuts; stay put on Grok Bots, Muse, or Instinct unless already living in ChatGPT or Codex.
Mapping against Ray Data Co
Concrete connection: Dots' failure signature — dropped messages, permission trip-ups, and a browser the agent "often couldn't reach" — is the exact failure class project_channels_agent_setup (RDCO's own always-on Mac Mini/iMessage agent) and feedback_fresh_session_registry_not_settled ("missing MCP tools ≠ broken; read the log before escalating") already exist to manage. OpenAI shipping a v1 always-on agent with this same reliability shape, on far more resources than RDCO has, is useful outside confirmation: unattended-agent reliability (tool reachability, permission handling, dropped output) is an unsolved industry problem right now, not a symptom of RDCO's specific implementation being unusually rough — which is the same "factory failure mode" territory feedback_workflow_agent_output_integrity catalogs (pointer returns, false "verified" stamps) for RDCO's own agents.
- Second, sharper data point: the Decisions API is explicitly framed as a rival to Jev (TypeSafe), which RDCO already tracks via the
typesafe-aiskill and the2026-09-15-every-typesafe-jev-vibe-checknote. The mixed benchmark (OpenAI faster and slightly more accurate on one replay, tied-but-slower on another) is a live competitive signal worth re-checking before leaning further on Jev for inline judgment calls. - Caveat: free-preview email only, tagged paywall-truncated — the plugin-extension, $500 Pro plan, and Sol 6.1 detail sit behind Every's paywall and aren't reflected here. This issue carries no third-party sponsor block, only Every's standard house cross-promo footer (Sparkle/Cora/Spiral/Monologue, Every All Access) — treat as unsponsored.
Related
- [[2026-09-15-every-typesafe-jev-vibe-check]]
- [[2026-09-22-every-gpt6-sol-vs-opus-5-5-vibe-check]]
- [[2026-09-03-every-gpt6-astra-vibe-check]]
- [[project_channels_agent_setup]]
- [[feedback_workflow_agent_output_integrity]]
- [[feedback_fresh_session_registry_not_settled]]