06-reference/research

sw vision evaluator cost per customize run

2026-09-22·research-brief·source: deep-research·by Ray Data Co (deep-research synthesis)
scribble-worksunit-economicsvision-evaluatorsonnet-5customize

What the mandatory Sonnet 5 vision screen costs per Customize run

The question

"What does the mandatory per-image Sonnet 5 vision-evaluator call actually cost per Customize run — real input tokens for the image plus the structured-audit rubric, and how many fire including retries on a failed verdict?"

Context: [[2026-09-16-scribble-works-model-cost-ceiling-refresh]] flagged the evaluator as "the one genuinely unbounded term" in Scribble Works unit cost, because no cost model prices it. This brief prices it from the code on origin/main (commit 5ddf745, 2026-09-21, read via git show in ~/Projects/sw-pr-clearance).

Headline: each evaluator call costs about $0.010-0.020 (expected ~$0.013). That is 50-100% of the $0.02 Grok image it screens. Per Customize run the whole picture step (generations plus screens) costs $0.03-0.08 for a scene page and $0.06-0.16 for a two-icon page. The cost is not unbounded. Code bounds it per call, per run, per month, and per week.

What we already know (from the vault)

What the web says

The code trace (all verified on origin/main unless labeled)

Where it fires. src/lib/customize/art-rail.js:121-128 runEvaluator POSTs to the gateway's Anthropic /v1/messages. The comment says: "The vision screen is MANDATORY for every model on both rails: no verdict, no picture." Two callers:

Request shape (src/lib/customize/art.js:243-273 scene, :479-497 icon):

Rubric size (measured by building the real request objects in Node):

Request System prompt User text JSON schema Total chars
Scene (EVAL_SYSTEM) 4,578 218 1,814 6,610
Icon (ICON_EVAL_SYSTEM) 3,323 158 1,820 5,301
Original 2026-09-01 scene (commit 0bed8a1) 1,504 218 664 2,386

Between the first version and today the rubric grew ~2.8x. The growth came from the "MANDATORY STRUCTURE AUDIT" extremity inventory (added 2026-09-06, #107) and the clarity check. Effort also went from low to medium and max_tokens from 600 to 1200.

Real calibration point (verified). studio/notes/customize-art/smoke-2026-09-02T02-10-10.json records one live evaluator call on the old rubric: 2,332 input tokens, 23 output tokens, 0 thinking. It screened an 864x1152 JPEG (dimensions read with sips from the vault copy at 01-projects/printables-product/studio-notes-2026-09/customize-art/).

Input tokens now (inference, two methods, calibrated on that point):

Method How Scene Icon
A: proportional Scale the 1,030 tokens by the character count ~2,850 ~2,290
B: fixed overhead chars/3.08 plus a ~255-token fixed overhead implied by the old call ~2,400 ~1,980

Totals per call:

Output tokens (inference). The schema now requires:

Medium-effort adaptive thinking also counts as output. Estimates:

Latency hints that the stricter screen really does think. The icon evaluator measured 8.6 s before the audit existed (customize-icons.js:68). The anatomy regression probe measured 17.0 s (customize-art.js:96). No post-audit usage log exists anywhere in the repo or the vault, so the output figure is the weakest number in this brief.

Cost per evaluator call at $2/$10:

Case Scene Icon
Low (~250 output tokens) $0.0099 $0.0092
Expected (~500 output tokens) $0.0129 $0.0120
Cap (1,200 output tokens) $0.0203 $0.0193
Absolute ceiling (4,784-token image + cap) $0.0273 $0.0262

How many fire (verified).

Per Customize run, picture step only (Grok $0.02/image + screen):

Run Best (pass first try) Expected (first-pass fail rate p = 0.3) p = 0.6 Worst (both attempts, output cap)
Scene page $0.030 $0.043 $0.053 $0.081
Two-icon page $0.058 $0.083 $0.102 $0.157
Of which: evaluator only $0.010 / $0.018 $0.017 / $0.031 $0.021 / $0.038 $0.041 / $0.077

The p values are illustrative. Nobody has measured the fail rate under the current rubric.

Aggregate bounds (verified). The evaluator is metered, just not modeled:

Convergences and contradictions

Synthesis for RDCO

The number the parent brief was missing is small, but it is not noise. A vision screen costs about $0.013. That is roughly the same as the entire Stage 1 words call and about two-thirds of the Grok image it screens. The evaluator is the largest per-image cost after generation, and it doubles on every failed verdict because each retry pays for a new image and a new screen. Adding it to the unit model, the picture step of a Customize run is ~$0.03-0.08 (scene) or ~$0.06-0.16 (two icons). The parent brief's "$0.06 for up to 3 images" is too high on image count (at most 2 slots per page today) and too low on multipliers (retries plus screens).

For pricing decisions this is a correction, not a rethink. The worst credible picture-heavy Customize adds ~$0.16 to generation cost. That moves the charter's credit break-even from $0.74 to about $0.87 per pack if every pack triggers an icon redraw at the worst case. The expected case ($0.04-0.08) moves it by a few cents. The dominant uncertainty is not the rubric token count, which is now pinned to about ±10%. It is the first-pass fail rate under the anatomy audit, which sets the expected column, and the thinking-token spend at medium effort, which sets the output line. Both can be read from the D1 ledger (charged_micros per Anthropic call on use_case='customization') without making any new call.

Two smaller design facts matter operationally:

Why this is in the vault

It prices the evaluator line item that the Scribble Works credit tiers and break-even bands ([[2026-08-31-studio-charter]] §4, [[2026-09-05-kids-subscription-box-comparables-scribble-works]]) omit. It also corrects the "unbounded" label in [[2026-09-16-scribble-works-model-cost-ceiling-refresh]] before anyone uses it to argue for or against keeping the mandatory screen.

Open follow-ups

Related

Sources