06-reference

every jev usage guide

2026-09-23·reference·source: Every·by Laura Entis
typesafejevclassification-modelsllm-judgeagent-verificationprompt-engineeringintent-based-software

How to Get the Most Out of Jev

Why this is in the vault

Third Every piece on TypeSafe's Jev in nine days — this one is a hands-on usage guide (install steps, prompt templates, a refinement loop) rather than a review, and it supplies the adoption evidence and concrete workflow recipe the prior two notes flagged as still missing.

The core argument

Jev — TypeSafe's classification-only "System One" model — has gone viral since its September 15 release: demos of it playing video games, re-ranking search results, and inferring intent from speech-plus-gesture (Jack Cheng's tldraw demo, 1M+ X views) are everywhere, and TypeSafe's release spawned an X leaderboard tracking who burns the most Jev tokens. Every's staff writer Laura Entis packages three internal use cases into a "steal this workflow" recipe: (1) install the official TypeSafe skill from github.com/typesafe-ai/skills so a coding agent knows how to structure Jev questions; (2) use a coding agent to interrogate a fuzzy judgment ("is this urgent," "does this sound like AI") into a plain-English working definition by walking through examples and edge cases; (3) break that definition into discrete, independently-scored checks — Jack Cheng did this for an email filter (real person? someone waiting on him? costs him something if ignored?) run across 90 days of his inbox; (4) validate by running checks on a sample the human has already judged, without feeding those judgments to Jev, comparing misses, and revising. Douglas Brundage's news-triage Grok bot went from reviewing <20 posts per run to 500+ after bolting on Jev. Pricing: 4.2 cents per million input tokens, output free — vs. Gemini 3.7 Flash ($0.75/$3.75), Haiku 4.5 ($1/$5), Sol ($4/$20), Astra and Fable 5.1 ($10/$50 each).

Mapping against Ray Data Co

This directly answers the open item both prior Jev notes left hanging. The 2026-09-15 vibe-check note flagged Jev as "the kind of cheap judge that could make a real benchmark-driven selector affordable" but had no concrete build recipe; the 2026-09-20 Innermost Loop note flagged Jev as "a tool already in RDCO's own stack (typesafe:typesafe-ai skill)" worth "a beat of due diligence" rather than passive filing. This issue supplies both the recipe and the due-diligence signal: TypeSafe ships an official skill (github.com/typesafe-ai/skills, separate from and possibly more current than RDCO's installed typesafe:typesafe-ai skill — worth diffing) plus a four-step pattern (define → decompose into scored checks → validate against a held-out human-judged sample → revise) that maps almost exactly onto how RDCO's fresh-eyes critic family (verify-vault-write, verify-dispatch, station-critic) already works, except those critics run once, post-hoc, on a full-reasoning model. The concrete next step this justifies: prototype a Jev pre-filter for one narrow, high-volume check inside an existing critic (e.g., a single binary sub-check inside verify-vault-write's rubric) using Entis's exact refinement loop, and measure it against the 4.2-cents-per-million-tokens number here rather than assuming the economics. Separately, the adoption signal itself (viral demos, a public token-burn leaderboard, Douglas's 20-to-500-posts before/after) is a real escalation from the 09-20 note's one-line curation mention — this is no longer "a vendor named in a roundup," it's a vendor with a public usage surge and an official agent-skill distribution channel, which raises the priority of the due-diligence beat those two prior notes opened.

Related

Curation section