06-reference

every copilot org chart autopilot

2026-09-25·reference·source: Every·by Ryan Sloan
microsoftcopilotautopilotagentic-org-chartevalsenterprise-agents

Why this is in the vault

First on-the-ground report (Ryan Sloan, at Microsoft's Redmond launch event) of Copilot's new "Autopilot" — a persistent, named agent with its own permissions and a place in the org chart that reports to a human owner — plus Microsoft's own account of how it evaluates whether such an agent is trustworthy (16,000 internal evals, customer-specific eval-building).

Mapping against Ray Data Co

Autopilot is Microsoft trying to productize the exact thing RDCO already runs: a persistent, named agent (Dot, Apollo) with delegated permissions, an owner it reports to, and a place in the org's structure — which is a description of Ray (Claude Code as COO) with different branding. The piece's most useful detail isn't the feature list, it's Lamanna's operating bar: Autopilot has to "clear a higher bar" than an on-demand tool because it spends tokens and takes action around the clock, and the way customers earn trust in it is by having "the person who does the job every day build the eval," measuring the human baseline first, then "embracing the red" as tests fail and harden. That is a vendor-scale restatement of RDCO's own verification discipline (feedback_workflow_agent_output_integrity, feedback_verification_independent_worker_pattern) — evals built by the domain expert, a baseline to beat, failing tests treated as signal, not noise. It's a useful proof point that this isn't an RDCO-specific eccentricity; a company Microsoft's size is converging on the same shape at 16,000x the scale.

Second, more tactical connection: Sloan's demo-floor failure (Copilot couldn't move a table from Word to Excel — "lacked permission") is the same enterprise-IT-permissions friction already logged in [[2026-08-03-every-microsoft-copilot-studio-trap]] (Mike Taylor's Copilot Studio onboarding gauntlet) and directly relevant to Ben's phData DSA/TAL seat: most phData clients are enterprise Snowflake/Azure shops that default to Copilot, and "the controls meant to make Copilot trustworthy could also stop it from doing its job" is a reusable, now twice-sourced Every observation for framing why a technically capable Microsoft AI tool underdelivers on the ground relative to Claude Code.

Third: "Autopilot runs on usage-based credits and is off by default" continues the metered-intelligence thread from [[2026-06-05-every-microsoft-metered-intelligence]] — always-on agentic work has to justify its own token spend, the same cost logic behind RDCO's own model-mixing and sub-agent context-isolation practice.

The core argument

At a Redmond launch event, Microsoft unified its separate consumer and work Copilot apps into one, built around three tabs: Home (chat + Office context), Code (no-code app building/sharing), and Autopilot (a persistent agent you assign a goal and permissions to). Autopilot agents get their own identity, permissions, and a visible place in the org chart, reporting to the human who owns them — Nadella framed the AI agent itself as "the new massive insider risk," which is why Autopilot is scoped this way rather than given blanket access. In the demo, Sloan could reformat a PowerPoint slide via chat but couldn't move a table from Word into Excel ("lacked permission"), and a Chat-to-Code handoff also failed — live evidence that IT permissioning, not model capability, is the binding constraint. Charles Lamanna (EVP, Copilot/agents/platform) named his own Autopilot "Apollo," gave it his EA's Outlook access, and talks to it in Teams like a colleague; it runs "Routines" — standing jobs like meeting prep — around the clock on usage-based credits. Microsoft's answer to "is this actually good" is scale-evals: ~16,000 internal evals built from real usage data by domain experts, plus customer-specific evals built jointly with clients, with the operating discipline that the person who does the job daily builds the eval, a human baseline gets measured first, and failing tests ("embrace the red") get fixed and hardened rather than avoided.

⚠️ Sponsorship

No third-party paid sponsor in this issue. Every runs its standard house footer — a "Want to sponsor Every?" CTA plus the bundled self-promo block for its own products (Sparkle, Cora, Spiral, Monologue) and Every All Access. Disclosed for completeness (sponsor_entity: house-promo); it doesn't bias Sloan's reported observations from the launch event.

Related