Mini-Vibe Check: TypeSafe's Jev Judged Everything I've Written in 0.7 Seconds
Why this is in the vault
TypeSafe launched Jev, a non-autoregressive "System One" model trained via Reinforcement Learning for Calibrated Decisions (RLCD) to answer fuzzy yes/no or categorical questions with calibrated probabilities instead of chat text — Every's head of evals, Mike Taylor, ran it against his own writing and against Every CEO Dan Shipper's writing-defect benchmark, and it's directly relevant to how RDCO thinks about cheap inline verification of agent output.
The core argument
Jev isn't a chatbot — it's purpose-built to sit inside a software pipeline and return a structured probability ("is this customer angry? 0.9") rather than prose, so calling code never has to parse an LLM's flowery text into a boolean. Because it skips token-by-token generation, it's fast (Taylor ran 777 judgments across 37 documents in 0.7 seconds) and TypeSafe prices it at $42/billion tokens with free output tokens — TypeSafe says it's "too cheap to meter." In Dan Shipper's head-to-head against Fable 5.1 on four writing-defect checks across 12 synthetic passages, Jev was ~25x faster (0.35s vs. 8.83s median) and ~580x cheaper, but caught 6 of 7 intended defects to Fable's 7 of 7 — missing an "unexplained action" check Fable caught. Taylor's framing: Jev is "a code linter for knowledge work" — cheap and fast enough to run checks while an agent is still working, not just once at the end, catching an early-warning-system tier of accuracy rather than a courtroom-grade one.
Mapping against Ray Data Co
The concrete gap this exposes: every fresh-eyes critic RDCO runs today — verify-vault-write, verify-dispatch, verify-strategic-output, station-critic in the skill-agent-brigade convergence loop — is a post-hoc gate, invoked once, after the artifact is fully built, using a full-reasoning model (Fable at max effort per the sw-critic agent definition). Jev's pitch is the same judgment shape (does this artifact meet a bar?) at a price and latency (sub-cent, sub-second for hundreds of judgments) that make running it continuously during generation plausible — e.g., checking each paragraph of a Sanity Check draft, or each step of a /loop build, rather than waiting for the single end-of-chain critic pass. Taylor's own accuracy caveat (6/7 vs. Fable's 7/7, missing exactly the kind of subtle "unexplained mechanism" defect draft-review already screens for) is the reason this isn't a replacement for the fresh-eyes critic family — it reads as a cheap pre-filter that could catch obvious misses early and let the expensive Fable-based critic focus on what a $42/billion-token model can't reliably catch. Also notable for the investing-adjacent evals thesis: TypeSafe's cofounder Diogo Almeida co-authored OpenAI's 2022 InstructGPT paper, another data point for RDCO's "evals-as-moat" tracking of where former frontier-lab researchers spin out into evaluation-infrastructure startups (see feedback_delegation_model_effort_pairing on RDCO's own model/effort pairing, which — as 2026-09-10-every-evals-for-everyone already flagged — is set by task-size heuristic today, not by a benchmark score; Jev is the kind of cheap judge that could make a real benchmark-driven selector affordable).
Related
- [[2026-09-10-every-evals-for-everyone]] — same publication, same "turn judgment into a repeatable check" thesis; Taylor's Jev experiment is the infrastructure-layer counterpart to Laura Entis's personal-benchmark piece
- [[2026-09-03-every-gpt6-astra-vibe-check]] — Taylor used GPT-6 Astra (via Codex) to generate his test scenarios in this same piece; also the precedent for
single-thread-deep-diveas the format for Every's single-model reviews - [[2026-05-20-lotte-verheyden-evals-explained-langfuse-academy]] — foundational evals-methodology note this extends: calibrated-probability judges vs. the eval patterns already in the vault
- [[feedback_delegation_model_effort_pairing]] — RDCO's current model/effort selection is heuristic-based, not benchmark-scored; Jev is exactly the class of cheap judge that could close that gap
- [[feedback_fresh_eyes_subagent_for_own_artifacts]] — the standing rationale for why RDCO's critics are separate fresh-eyes passes rather than self-grading; Jev doesn't change this, but changes what's affordable to run inline before that pass
- [[2026-09-23-every-jev-usage-guide]] — follow-up issue with the concrete build recipe (install official skill, decompose into scored checks, validate against a judged sample) this note's "cheap pre-filter" idea was missing