06-reference

every computer use agent chores

2026-09-15·reference·source: Every·by Laura Entis
computer-useagentic-automationastracodexagent-deployerharness-engineering

You're Probably Sleeping on Computer Use

Why this is in the vault

Concrete field evidence — from Every's own staff, not vendor marketing — that computer-use agents (clicking, typing, scrolling real apps with no API) have crossed from clunky demo to daily-driver chore automation; a direct data point for RDCO's agent-deployer/harness thesis.

The core argument

Every's head of evals, Mike Taylor, rated OpenAI's Astra "yellow" (skippable) in the September 3 Vibe Check, then flipped within two weeks to calling it borderline AGI. The cause wasn't a capability jump — it was computer use. He started with a low-stakes test (filling out his daughter's school forms), watched it succeed, and escalated to having Astra edit six full Google Slides decks and audit every hyperlink in the PDF proofs of his new book (opening each link, checking the page loaded, and confirming it matched the surrounding text). The mechanism: computer use removes the integration tax on chores that don't have a clean API — the things people "should probably do" but keep deferring become things you just hand to an agent.

Curation section

Standard Every footer self-promo also appears (Sparkle/Cora/Spiral/Monologue bundle, All Access membership, Thesis:2027 conference plug) — boilerplate cross-sell, not part of the core reporting, and no paid third-party sponsor in this issue.

Mapping against Ray Data Co

This is field validation of the agent-deployer / harness-engineering thesis: the gating factor for daily agent use wasn't model capability — Astra scored only "yellow" on raw capability in the September 3 Vibe Check (see [[2026-09-03-every-gpt6-astra-vibe-check]]) — it was closing the last-mile gap between "the model can reason about this" and "the model can actually operate the tool with no API." That's exactly the harness/interface layer RDCO's agent-deployer positioning is built around. The staffer list is a live inventory of the un-glamorous, low-ROI-per-task chores (calendar entry, dead-link QA, marketplace listing, support-chat negotiation) that computer use unlocks precisely because they don't justify bespoke API integration — the same "integration tax" argument behind the bet that harness quality, not model quality, is the remaining bottleneck for agent deployment in SMB/ops contexts (see [[project_phdata_cert_escalator_path]]). One gap worth flagging: every example here is single-operator, ad hoc, and unverified — there's no QA/eval layer on top of the computer-use actions (Astra checks its own link-audit spreadsheet, with no second agent verifying it). That's the exact risk pattern [[feedback_workflow_agent_output_integrity]] calls out. Computer use raises capability without raising verification, so this piece is evidence for automating the chore and evidence for keeping a verification gate in front of anything that leaves the sandbox.

Related