06-reference

every vibe check devday 2026

2026-09-29·reference·source: Every·by Dan Shipper
openaidevday-2026always-on-agentsdecisions-apichatgptevery-vibe-check

Why this is in the vault

Every's Dan Shipper spent days testing the biggest of OpenAI's 20+ DevDay 2026 launches — his verdict is that OpenAI's plan is now plain (ChatGPT as the operating system for work, with agents, documents, and other companies' apps all running inside it), but the flagship always-on agent still has rough reliability edges.

The core argument

OpenAI announced more than 20 products and features at DevDay 2026 — by Shipper's count, the biggest DevDay since the first one in 2023. Four launches anchor the review (full comparison, plugin extensions, the $500 Pro plan, and Sol 6.1 detail sit behind Every's paywall):

Shipper's bottom line: try Dots now if you like new tools with rough edges, otherwise give OpenAI a week or two to fix the paper cuts; stay put on Grok Bots, Muse, or Instinct unless already living in ChatGPT or Codex.

Mapping against Ray Data Co

Concrete connection: Dots' failure signature — dropped messages, permission trip-ups, and a browser the agent "often couldn't reach" — is the exact failure class project_channels_agent_setup (RDCO's own always-on Mac Mini/iMessage agent) and feedback_fresh_session_registry_not_settled ("missing MCP tools ≠ broken; read the log before escalating") already exist to manage. OpenAI shipping a v1 always-on agent with this same reliability shape, on far more resources than RDCO has, is useful outside confirmation: unattended-agent reliability (tool reachability, permission handling, dropped output) is an unsolved industry problem right now, not a symptom of RDCO's specific implementation being unusually rough — which is the same "factory failure mode" territory feedback_workflow_agent_output_integrity catalogs (pointer returns, false "verified" stamps) for RDCO's own agents.

Related