06-reference

every ai employee security framework

2026-08-14·reference·source: Every·by Nityesh Agarwal and Claude

Securing an Always-on AI Employee - Nityesh Agarwal and Claude

Re-read at full length 2026-10-03 (paid).

Why this is in the vault

This is the only published security framework written for an agent shaped like Ray: Claude Code, 24/7 on a dedicated Mac mini, reached through a chat channel, with its own email and logged-in browser sessions. The earlier version of this note was filed from a truncated free preview plus a web reconstruction (about 250 words). It named the three threats but marked the "four layers" and the self-audit prompt as paywalled. This rewrite covers the full guide.

The core argument

Claudie (Every Consulting's chief-of-staff agent, run through Slack) was deliberately given broad access first, to learn what she could do, then pared back. The trigger was the March 2026 npm compromises (Axios, then TanStack in May). An always-on agent has no human denying suspicious tool calls in real time, and "anyone who can influence that context" can steer it.

Three threat vectors:

  1. Supply chain - compromised dependencies run with the agent's full permissions.
  2. Prompt injection - inbound content posing as instructions. This is not hypothetical: Claudie's non-public address already gets CEO-impersonation phishing ("URGENT RESPONSE!!!").
  3. Internal leakage - the most likely day-to-day failure. No attacker is needed; a helpful agent answers someone above their clearance.

Four layers, ordered by falling reliability and rising flexibility:

Layer What Strength / weakness
Least access Agent has its own identity and accounts; data shared selectively Most reliable; hard to change later
Programmatic Permission modes, PreToolUse hooks, identity gates Can't be talked out of it; binary
Prompt-based Ring-based access doc in context Handles nuance; improves with each model release
Observability JSONL session logs, conversation viewer, thinking-token review, forensic skill Prevents nothing, "catches everything"

Per-vector controls:

Gap-classification rubric: Is the risk caught by any programmatic control? Then it is a non-threat. Caught only by prompt plus observability? A known risk, acceptable if non-catastrophic. Caught by nothing? A true gap: fix it or explicitly accept it. Examples Every found in its own setup: session inheritance (a non-admin continuing an admin thread inherits elevated mode; known risk, 30-minute idle timeout, fix = re-check identity per message) and browser actions (prompt-only block; tolerable).

Audit prompt: pick the five highest-stakes actions, then run four parallel subagents, one per layer, and compile the results as non-threats / known risks / true gaps.

Mapping against Ray Data Co

Hard mapping of Ray's always-on Mac mini against the four layers (checked against ~/.claude/settings.json, ~/.npmrc, ~/.config/uv/uv.toml on 2026-10-03):

The earlier version of this note proposed consolidating Ray's scattered protections into one security-posture doc. With the full guide, that doc is the four-layer table above, plus the five-action audit.

Why it matters for RDCO / The Denominator

⚠️ Sponsorship

No third-party sponsor. House promo: written by Every Consulting's engineer about Every's own agent, and it doubles as a capability showcase for the consulting arm. The guide links to an external Claudie security briefing page. Bias implication: the framework is presented as proven, but the evidence is one team's agent and two self-reported phishing anecdotes. The rubric and controls are standard defense-in-depth and hold up on their own merits.

Related