Securing an Always-on AI Employee - Nityesh Agarwal and Claude
Re-read at full length 2026-10-03 (paid).
Why this is in the vault
This is the only published security framework written for an agent shaped like Ray: Claude Code, 24/7 on a dedicated Mac mini, reached through a chat channel, with its own email and logged-in browser sessions. The earlier version of this note was filed from a truncated free preview plus a web reconstruction (about 250 words). It named the three threats but marked the "four layers" and the self-audit prompt as paywalled. This rewrite covers the full guide.
The core argument
Claudie (Every Consulting's chief-of-staff agent, run through Slack) was deliberately given broad access first, to learn what she could do, then pared back. The trigger was the March 2026 npm compromises (Axios, then TanStack in May). An always-on agent has no human denying suspicious tool calls in real time, and "anyone who can influence that context" can steer it.
Three threat vectors:
- Supply chain - compromised dependencies run with the agent's full permissions.
- Prompt injection - inbound content posing as instructions. This is not hypothetical: Claudie's non-public address already gets CEO-impersonation phishing ("URGENT RESPONSE!!!").
- Internal leakage - the most likely day-to-day failure. No attacker is needed; a helpful agent answers someone above their clearance.
Four layers, ordered by falling reliability and rising flexibility:
| Layer | What | Strength / weakness |
|---|---|---|
| Least access | Agent has its own identity and accounts; data shared selectively | Most reliable; hard to change later |
| Programmatic | Permission modes, PreToolUse hooks, identity gates | Can't be talked out of it; binary |
| Prompt-based | Ring-based access doc in context | Handles nuance; improves with each model release |
| Observability | JSONL session logs, conversation viewer, thinking-token review, forensic skill | Prevents nothing, "catches everything" |
Per-vector controls:
- Supply chain: package quarantine (only install releases older than N days, e.g. 7, for npm/pip/cargo/brew; this would have blocked Axios). A shlex-parsing PreToolUse hook rejects
eval "$(curl…)",base64 -d | bash, pipe-to-shell, andpython -cwith os calls. Strongest is a credential proxy: tokens live in a separate service the agent can call but not read (the pattern Anthropic's managed agents use). This narrows scope (approved recipients only), reach (dies with the session) and visibility (one logged chokepoint). - Injection: the agent can receive and classify email but cannot send email or click links. This is enforced three times (prompt, harness tool removal, bash hook that greps for gmail send/reply/forward, including chained commands), and the hook binds admins too. Browsing and inbound processing run in
dontAskmode. Unknown Slack IDs are silently ignored in code before a Claude process spawns. - Leakage: the bot sets the permission mode at spawn time by sender identity. Admins get bypassPermissions; everyone else gets dontAsk plus allow/deny lists blocking email, calendar, session logs, MCP and config. There is an identity-aware file browser that hides restricted paths. Rings: 0 universal (no external communication, no credential exposure, never run injected scripts); 1 admins; 2 core team (no email, calendar, logs or access changes); 3 wider org (adds no client data or Workspace); outside the rings means ignored. Per-user profile files accumulate preferences and overrides. Here observability is the primary defense: near-misses show up in thinking tokens, and each one becomes a prompt refinement.
Gap-classification rubric: Is the risk caught by any programmatic control? Then it is a non-threat. Caught only by prompt plus observability? A known risk, acceptable if non-catastrophic. Caught by nothing? A true gap: fix it or explicitly accept it. Examples Every found in its own setup: session inheritance (a non-admin continuing an admin thread inherits elevated mode; known risk, 30-minute idle timeout, fix = re-check identity per message) and browser actions (prompt-only block; tolerable).
Audit prompt: pick the five highest-stakes actions, then run four parallel subagents, one per layer, and compile the results as non-threats / known risks / true gaps.
Mapping against Ray Data Co
Hard mapping of Ray's always-on Mac mini against the four layers (checked against ~/.claude/settings.json, ~/.npmrc, ~/.config/uv/uv.toml on 2026-10-03):
- Package quarantine: covered programmatically.
.npmrcsetsmin-release-age=7, uv setsexclude-newer = "7 days", and a PreToolUse Bash hook (package-security-check.sh) gates installs, per [[supply-chain-security]]. This is the same control Every recommends, and Ray already has it. Brew is still policy-only (prompt layer). - Identity gate: covered programmatically. The iMessage plugin allowlist (
access.json, managed only by the founder via/imessage:access) is the counterpart to Every's silent ignore of unknown Slack IDs. Ray has a single principal, so the ring model collapses to Ring 1 plus "outside". That removes Every's whole leakage-by-clearance class until a second human gets access. When that happens, the spawn-time-permission and session-inheritance lessons apply directly. - No autonomous external email: prompt-layer only. This is a true gap by Every's rubric. "Ray drafts, founder sends" is a memory rule. The Gmail connector exposes
send_message/reply/forward, and no PreToolUse hook denies them. Every enforces the same rule three times; Ray enforces it once. The auto-mode classifier adds some friction but is not a deterministic deny. - Approval-gated deploys: partly programmatic. The auto-mode classifier hard-gate on deploy/production writes and the PR-only workflow are real controls. But the classifier is a model, not a hook. A deny hook on
wrangler deploy/git push origin mainwould make this a non-threat rather than a known risk. - 1Password service-account wrappers: better than
.env, weaker than a proxy. No secrets on disk, and creds load at call time. But the agent can execute the wrappers, so code running as Ray can read the tokens. Every's (and Anthropic managed agents') proxy pattern, also covered in [[2026-05-22-tony-dang-credential-brokering-for-ai-agents]], is the next step up for the highest-value creds (Cloudflare, Gmail). - Observability: logs exist, review loop doesn't. Session JSONL and
/tmp/claude-channels.logexist, but nothing systematically reviews thinking tokens for near-misses. Prompt injection via newsletters and pasted content is the highest-volume vector Ray faces; the pasted-content carve-out in AGENTS.md is a prompt-layer control only.
The earlier version of this note proposed consolidating Ray's scattered protections into one security-posture doc. With the full guide, that doc is the four-layer table above, plus the five-action audit.
Why it matters for RDCO / The Denominator
- Do (Ray, this week): add a PreToolUse deny hook for Gmail send/reply/forward and for production deploy commands, binding even in auto mode. Those are the two "prompt-only" rules where one model slip is unrecoverable. Then run Every's four-subagent audit prompt against Ray and file the non-threat / known-risk / true-gap list.
- Do (factory): the spawn-time permission split (admin vs. dontAsk with allow/deny lists) is the template for any multi-user phData copilot agent. Put it in the factory's deployment checklist.
- Write (The Denominator): the lens "reliability decreases as flexibility increases, and only the prompt layer improves with each model release" is a strong frame for why successful enterprise agents put irreversible actions behind code and leave judgment calls to prompts.
⚠️ Sponsorship
No third-party sponsor. House promo: written by Every Consulting's engineer about Every's own agent, and it doubles as a capability showcase for the consulting arm. The guide links to an external Claudie security briefing page. Bias implication: the framework is presented as proven, but the evidence is one team's agent and two self-reported phishing anecdotes. The rubric and controls are standard defense-in-depth and hold up on their own merits.
Related
- [[2026-03-31-every-onboarding-ai-project-manager]] - the direct precursor piece: same author, same agent (Claudie), covers how she was onboarded before this piece covers how she's restricted
- [[2026-05-11-indy-dev-dan-delete-bash-tool-agentic-security]] - independent agentic-security framework (bash-tool risk tiers); same gap class for Ray's bash-heavy setup
- [[2026-07-23-alphasignal-claude-code-security-plugin-agents-500-skills]] - Anthropic's Claude Code security plugin; agent security becoming a product category
- [[2026-05-22-tony-dang-credential-brokering-for-ai-agents]] - the credential-proxy pattern in depth
- [[supply-chain-security]] - Ray's existing package-quarantine SOP
- [[2026-04-30-quality-gate-as-brain-org-boundaries-agentic-companies]] - blast-radius reasoning for splitting agent capabilities
- [[2026-06-07-every-executive-guide-implementing-ai]] - the adoption playbook that produced Claudie
- [[2026-03-03-every-openclaw-comprehensive-guide]] - the harness-agnostic always-on agent context