06-reference

every typesafe jev vibe check

2026-09-15·reference·source: Every·by Mike Taylor
model-evaluationllm-judgestructured-outputsharness-engineeringagent-verificationtypesafecalibration

Mini-Vibe Check: TypeSafe's Jev Judged Everything I've Written in 0.7 Seconds

Why this is in the vault

TypeSafe launched Jev, a non-autoregressive "System One" model trained via Reinforcement Learning for Calibrated Decisions (RLCD) to answer fuzzy yes/no or categorical questions with calibrated probabilities instead of chat text — Every's head of evals, Mike Taylor, ran it against his own writing and against Every CEO Dan Shipper's writing-defect benchmark, and it's directly relevant to how RDCO thinks about cheap inline verification of agent output.

The core argument

Jev isn't a chatbot — it's purpose-built to sit inside a software pipeline and return a structured probability ("is this customer angry? 0.9") rather than prose, so calling code never has to parse an LLM's flowery text into a boolean. Because it skips token-by-token generation, it's fast (Taylor ran 777 judgments across 37 documents in 0.7 seconds) and TypeSafe prices it at $42/billion tokens with free output tokens — TypeSafe says it's "too cheap to meter." In Dan Shipper's head-to-head against Fable 5.1 on four writing-defect checks across 12 synthetic passages, Jev was ~25x faster (0.35s vs. 8.83s median) and ~580x cheaper, but caught 6 of 7 intended defects to Fable's 7 of 7 — missing an "unexplained action" check Fable caught. Taylor's framing: Jev is "a code linter for knowledge work" — cheap and fast enough to run checks while an agent is still working, not just once at the end, catching an early-warning-system tier of accuracy rather than a courtroom-grade one.

Mapping against Ray Data Co

The concrete gap this exposes: every fresh-eyes critic RDCO runs today — verify-vault-write, verify-dispatch, verify-strategic-output, station-critic in the skill-agent-brigade convergence loop — is a post-hoc gate, invoked once, after the artifact is fully built, using a full-reasoning model (Fable at max effort per the sw-critic agent definition). Jev's pitch is the same judgment shape (does this artifact meet a bar?) at a price and latency (sub-cent, sub-second for hundreds of judgments) that make running it continuously during generation plausible — e.g., checking each paragraph of a Sanity Check draft, or each step of a /loop build, rather than waiting for the single end-of-chain critic pass. Taylor's own accuracy caveat (6/7 vs. Fable's 7/7, missing exactly the kind of subtle "unexplained mechanism" defect draft-review already screens for) is the reason this isn't a replacement for the fresh-eyes critic family — it reads as a cheap pre-filter that could catch obvious misses early and let the expensive Fable-based critic focus on what a $42/billion-token model can't reliably catch. Also notable for the investing-adjacent evals thesis: TypeSafe's cofounder Diogo Almeida co-authored OpenAI's 2022 InstructGPT paper, another data point for RDCO's "evals-as-moat" tracking of where former frontier-lab researchers spin out into evaluation-infrastructure startups (see feedback_delegation_model_effort_pairing on RDCO's own model/effort pairing, which — as 2026-09-10-every-evals-for-everyone already flagged — is set by task-size heuristic today, not by a benchmark score; Jev is the kind of cheap judge that could make a real benchmark-driven selector affordable).

Related