06-reference

every voice ai workflow

2026-07-31·reference·source: Every·by Naveen Naidu (with Laura Entis; GPT co-credit)
voice-aiai-writing-workflowagentic-workflowpromptingevery

Every — "Build Faster With Voice"

Subject line: The Definitive Guide to Using Voice With AI Byline: Naveen Naidu (GM of Monologue, Every's own voice-to-text product), with Laura Entis, GPT co-credited

Why this is in the vault

A concrete, reusable prompt template ("the agentic voice loop") for turning raw spoken input into finished agent output — directly transferable to how the founder already works with Ray.

The core argument

Naidu's thesis: once AI agents can absorb raw, unstructured input and turn it into structured output, the old requirement to "translate everything into polished prose or code" before acting disappears. Voice removes that translation step entirely — you can speak a half-formed idea, point an agent at supporting context, name the outcome you want, and let the agent do the structuring. He frames this as a five-step "agentic voice loop": capture raw material → add context → define the outcome → agent acts → review and redirect. He also gives a fill-in-the-blank spoken brief template: "Here is what's happening: [situation]. Retrieve more context from: [sources]. What I want you to generate: [artifact] for [destination]. Constraints: [rules]." Naidu distinguishes "active collaboration" (live back-and-forth voice sessions with an agent) from "passive capture" (recording a meeting or ramble and handing the transcript to an agent later) as the two modes voice enters a workflow. Tools referenced: Monologue and Monologue Notes (Every's own products), plus Codex and Claude Code as the agent side of the loop.

Mapping against Ray Data Co

Medium-strong. The most concrete connection: the founder's actual operating pattern with Ray — dictate or type a rough ask via iMessage/Discord, Ray pulls context from the vault/Notion/Gmail, produces a defined artifact — is already an instance of Naidu's "agentic voice loop," just without the literal speech-to-text step. The spoken-brief template ("here's what's happening / pull context from X / generate Y for Z / constraints") is a usable structuring device the founder could apply explicitly when dictating asks on the go (e.g., voice-memo-to-iMessage), rather than something net-new to adopt. It's a smaller connection to CLAUDE.md's own hard-rule structure (situation → context sources → required artifact → constraints) than to a genuinely new capability — this is a naming/discipline upgrade to an existing practice, not a new tool RDCO lacks. No direct bearing on Sanity Check craft voice, the Ray mascot's ElevenLabs voice (that's output TTS, not input dictation), or multi-brand voice consistency — those are a different sense of "voice" than this piece addresses.

⚠️ Sponsorship

Two layers of self-promotion, no independent third-party sponsor:

  1. The guide itself is written by the GM of Monologue, Every's own dictation product, and uses the piece to promote Monologue and Monologue Notes as the reference tools for the workflow it describes — a structural conflict of interest (the author is pitching the product he runs), though the underlying technique (the voice-loop template) is usable independent of that product.
  2. The same newsletter issue separately promotes "Paper and Mobbin join Every's Builder Pack" — Builder Pack (builderpack.ai) is Every's own bundled-deals product aggregating credits from partner tools, not a single paid sponsor placement bought for this issue. Treat both as house cross-promo: read the technique, discount the tool endorsement.

Related