You're Probably Sleeping on Computer Use
Why this is in the vault
Concrete field evidence — from Every's own staff, not vendor marketing — that computer-use agents (clicking, typing, scrolling real apps with no API) have crossed from clunky demo to daily-driver chore automation; a direct data point for RDCO's agent-deployer/harness thesis.
The core argument
Every's head of evals, Mike Taylor, rated OpenAI's Astra "yellow" (skippable) in the September 3 Vibe Check, then flipped within two weeks to calling it borderline AGI. The cause wasn't a capability jump — it was computer use. He started with a low-stakes test (filling out his daughter's school forms), watched it succeed, and escalated to having Astra edit six full Google Slides decks and audit every hyperlink in the PDF proofs of his new book (opening each link, checking the page loaded, and confirming it matched the surrounding text). The mechanism: computer use removes the integration tax on chores that don't have a clean API — the things people "should probably do" but keep deferring become things you just hand to an agent.
Curation section
- Inside Every (staffer computer-use log): video lead runs Codex by voice while cooking and tags task titles with a screen emoji to track which one currently controls his machine; biz-dev lead has Codex maintain his calendar, including kids' school/camp events ("the highest-leverage thing I've done in my adult life"); social lead has Codex browse iPhone photos to list marketplace items and clear WhatsApp storage via iPhone Mirroring, and prefers Codex's computer use running quietly in the background over Claude Code, which "always brings the window up front"; an engineer uses computer use for Blender/Unity 3D-asset work — "the better it gets, the more it makes possible for me, and I have less of a life on weekends"; consulting lead had Codex negotiate a cheaper Verizon plan live inside the support chat; ops lead defaults to computer use for legacy software with no reliable connector, and to diagnose why her computer is "slow AF."
- Steal this workflow (slide decks): Every's head of marketing hand-builds two or three reference slides, has Codex draft the copy, then has it open Google Slides directly and build the full deck via computer use using the references as templates — reviewing the first pass (Codex once crammed 30 examples onto a single slide) and feeding corrections back as durable "writing skill" rules for next time.
- Thesis Statements: six new one-line contestable predictions about AI and the future of work (e.g., "good design will signal mediocrity," "technical competence will yield to people skills"), tied to Every's November 5, 2026 Thesis: 2027 conference.
- Links worth a click: Google engineers adopting Claude Code; a $24M AI-soldier startup; Trump weighing in on the AI safety debate; OpenAI buying a smartphone-camera startup for $300M+; Anthropic "barreling toward" its IPO.
Standard Every footer self-promo also appears (Sparkle/Cora/Spiral/Monologue bundle, All Access membership, Thesis:2027 conference plug) — boilerplate cross-sell, not part of the core reporting, and no paid third-party sponsor in this issue.
Mapping against Ray Data Co
This is field validation of the agent-deployer / harness-engineering thesis: the gating factor for daily agent use wasn't model capability — Astra scored only "yellow" on raw capability in the September 3 Vibe Check (see [[2026-09-03-every-gpt6-astra-vibe-check]]) — it was closing the last-mile gap between "the model can reason about this" and "the model can actually operate the tool with no API." That's exactly the harness/interface layer RDCO's agent-deployer positioning is built around. The staffer list is a live inventory of the un-glamorous, low-ROI-per-task chores (calendar entry, dead-link QA, marketplace listing, support-chat negotiation) that computer use unlocks precisely because they don't justify bespoke API integration — the same "integration tax" argument behind the bet that harness quality, not model quality, is the remaining bottleneck for agent deployment in SMB/ops contexts (see [[project_phdata_cert_escalator_path]]). One gap worth flagging: every example here is single-operator, ad hoc, and unverified — there's no QA/eval layer on top of the computer-use actions (Astra checks its own link-audit spreadsheet, with no second agent verifying it). That's the exact risk pattern [[feedback_workflow_agent_output_integrity]] calls out. Computer use raises capability without raising verification, so this piece is evidence for automating the chore and evidence for keeping a verification gate in front of anything that leaves the sandbox.
Related
- [[2026-09-03-every-gpt6-astra-vibe-check]]
- [[2026-08-06-technically-computer-use-agents]]
- [[2026-04-01-write-with-ai-claude-computer-use]]
- [[2026-09-06-every-fable-vs-astra-context-window]]
- [[feedback_workflow_agent_output_integrity]]
- [[project_phdata_cert_escalator_path]]