Why this is in the vault
A hands-on review of Deft's DFT v1 (a model post-trained specifically to fix "AI smell" in writing via a batch-distribution method rather than per-response grading), worth keeping as a data point on whether algorithmic AI-writing-quality fixes are production-ready or still lose to human editorial judgment.
Mapping against Ray Data Co
The direct connection is to the /draft-review skill's core job: catching exactly the tells Parrott names — patterns of three, "not X, but Y," a short punchy sentence tacked onto a paragraph's end to signal importance. Deft's pitch is that these tells are a distribution problem (a large batch of AI outputs collapses onto the same handful of moves) and that comparing batches of model output against batches of human writing during post-training, rather than grading one response at a time, produces more range. That's a more rigorous framing of the same diagnosis Sanity Check's CopyThat pattern-checking already runs on informally — useful vocabulary to borrow, but not evidence to buy or build a Deft-equivalent. Parrott's verdict is that Deft trades one failure mode for another: sentences got "more varied and surprising" but also "dense, hard to parse, and poorly sequenced," and in strict mode (with a detailed brief) it still invented an unsupported claim, then in the API/Codex-relay path it fabricated dates, dialogue, and product history despite strict mode being set — for $0.49 and no usable draft. That's a concrete caution against treating any AI-writing-quality layer as a drop-in replacement for the human review gate Sanity Check already runs (founder-run analysis + AI-assisted prose + human review before publish, per the IC-mode-vs-production-mode split) — the piece reinforces that editorial judgment on hierarchy, sequencing, and fact-boundary discipline isn't yet something a model can own unsupervised, which is exactly why that gate stays in place rather than getting automated away.
The core argument
Deft (cofounded by Justin Murphy and researcher "Rosmine") built DFT v1 around "distribution fine-tuning": instead of scoring individual outputs, it compares batches of model output against batches of human writing and closes the distributional gap, aiming for range rather than per-response acceptability. Parrott tested it across a from-scratch essay, an SEO article, a rewrite of an existing GPT-5.6 Sol draft, the web console, and the API (including piping it through Codex, which turned out to be a relay to Deft's own system rather than direct/iterative model access). Prose was more varied than typical AI writing but harder to parse and poorly sequenced — Deft's own reading-score tool rated its SEO output at an 11th-grade level. Even in "strict mode" (use only the prompt's details), it invented unsupported specifics; the API path fabricated dates, dialogue, and product history outright. Her framing, via the 2021 "stochastic parrot" critique (Bender, Gebru, Mitchell, McMillan-Major): Deft addresses the "stochastic" half (distributional sameness) but not the "parrot" half (meaning/grounding, sequencing, reader theory of mind). Verdict: "a promising demonstration," "a strong start," but "I wouldn't yet use Deft for day-to-day work."
Related
- [[2026-08-20-every-defense-of-ai-writing]] — same publication, adjacent AI-authorship-and-editing theme; that piece's Mike Taylor authorship framing (human-run analysis, AI-assisted expression) is the production model this note's mapping section leans on.
- [[2026-03-12-every-ai-writing-style-science]] — Marcus Moretti's piece on why AI writing stays detectable names the same tells (patterns of three, "delve," etc.) that Deft is explicitly trying to engineer away; read together they bracket the diagnosis (Moretti) and a first attempted algorithmic fix (this note).
- [[2026-03-18-every-editing-ai-writing]] — direct precedent on editing-as-differentiator, the same craft layer this note's RDCO mapping ties to
/draft-review.