What the mandatory Sonnet 5 vision screen costs per Customize run
The question
"What does the mandatory per-image Sonnet 5 vision-evaluator call actually cost per Customize run — real input tokens for the image plus the structured-audit rubric, and how many fire including retries on a failed verdict?"
Context: [[2026-09-16-scribble-works-model-cost-ceiling-refresh]] flagged the evaluator as "the one genuinely unbounded term" in Scribble Works unit cost, because no cost model prices it. This brief prices it from the code on origin/main (commit 5ddf745, 2026-09-21, read via git show in ~/Projects/sw-pr-clearance).
Headline: each evaluator call costs about $0.010-0.020 (expected ~$0.013). That is 50-100% of the $0.02 Grok image it screens. Per Customize run the whole picture step (generations plus screens) costs $0.03-0.08 for a scene page and $0.06-0.16 for a two-icon page. The cost is not unbounded. Code bounds it per call, per run, per month, and per week.
What we already know (from the vault)
- The charter's $0.71/pack ceiling is text-only. The engine spec called image cost "noise" and priced images with FLUX at ~$0.001 each ([[2026-09-01-engine-build-spec]] §6). Neither doc has an evaluator line. The engine spec's "evaluator" was a fresh-eyes subagent at authoring time, not a paid API call per image.
- The engine-spec spike saw 2.5 generations per passed slot (5 generations for 2 slots, n=2, no retry cap, FLUX era). This is the only prior we have for the fail rate, and it is weak.
- The Stage 1 words call is priced at ~$0.010/call (~2,800 in / ~350 out on Sonnet 5) in [[2026-09-01-customize-spec]]. Each vision screen costs about the same as, or more than, the call that rewrites the whole page.
- The parent brief put the image line at "up to 3 images for $0.06" ([[2026-09-16-scribble-works-model-cost-ceiling-refresh]]). Two things in the code contradict that: no game has both a scene slot and icon slots today, and retries plus screens multiply the line (see below).
What the web says
- Image token formula (verified): cost =
ceil(width/28) x ceil(height/28)visual tokens. "Claude 4.7 and later models" are on the high-resolution tier: 2576 px max long edge, 4,784 max visual tokens. Larger images are downscaled, not rejected (Anthropic vision docs). Sonnet 5 falls under "4.7 and later". That is an inference from the version number, not a named row in the table. - Sonnet 5 price (verified): $2 input / $10 output per MTok. The $2/$10 introductory price is now the standard price, and the planned 1 Sep 2026 rise to $3/$15 "will not occur" (Anthropic pricing). Output tokens include thinking tokens.
- Tokenizer (verified): Claude 4.7+ tokenizer "produces approximately 30% more tokens for the same text" (same pricing page). So plain chars/4 undercounts. I used chars/3.08 as one bound.
- Grok output size (partly verified): xAI docs describe a "1k" tier as 1024 px on the long edge (xAI Imagine docs). But the one real Customize smoke image in the vault is 864x1152 at aspect 3:4. That is 1152 px on the long edge, so trust the measured file over the doc for scene art. The size of 1:1 icons was not measured. I assumed 1024x1024.
The code trace (all verified on origin/main unless labeled)
Where it fires. src/lib/customize/art-rail.js:121-128 runEvaluator POSTs to the gateway's Anthropic /v1/messages. The comment says: "The vision screen is MANDATORY for every model on both rails: no verdict, no picture." Two callers:
- Scene art:
functions/api/customize-art.js:189-216. - Icon art:
functions/api/customize-icons.js:115-140.
Request shape (src/lib/customize/art.js:243-273 scene, :479-497 icon):
model: 'claude-sonnet-5'(art.js:23)max_tokens: 1200thinking: {type:'adaptive'}output_config: {effort:'medium', format: json_schema}- The image goes in as raw base64 at the size Grok returned. There is no resize before the call.
- There is no
cache_control, so the static rubric is billed at full price on every call.
Rubric size (measured by building the real request objects in Node):
| Request | System prompt | User text | JSON schema | Total chars |
|---|---|---|---|---|
Scene (EVAL_SYSTEM) |
4,578 | 218 | 1,814 | 6,610 |
Icon (ICON_EVAL_SYSTEM) |
3,323 | 158 | 1,820 | 5,301 |
Original 2026-09-01 scene (commit 0bed8a1) |
1,504 | 218 | 664 | 2,386 |
Between the first version and today the rubric grew ~2.8x. The growth came from the "MANDATORY STRUCTURE AUDIT" extremity inventory (added 2026-09-06, #107) and the clarity check. Effort also went from low to medium and max_tokens from 600 to 1200.
Real calibration point (verified). studio/notes/customize-art/smoke-2026-09-02T02-10-10.json records one live evaluator call on the old rubric: 2,332 input tokens, 23 output tokens, 0 thinking. It screened an 864x1152 JPEG (dimensions read with sips from the vault copy at 01-projects/printables-product/studio-notes-2026-09/customize-art/).
- The image costs 31 x 42 = 1,302 visual tokens.
- That leaves 1,030 tokens for the old rubric, schema, and structured-output framing.
Input tokens now (inference, two methods, calibrated on that point):
| Method | How | Scene | Icon |
|---|---|---|---|
| A: proportional | Scale the 1,030 tokens by the character count | ~2,850 | ~2,290 |
| B: fixed overhead | chars/3.08 plus a ~255-token fixed overhead implied by the old call | ~2,400 | ~1,980 |
Totals per call:
- Scene: 1,302 image + 2,400-2,850 text = ~3,700-4,150 input tokens.
- Icon: ~1,369 image (1024x1024 assumed = 37 x 37) + 1,980-2,290 text = ~3,350-3,650 input tokens.
- Absolute ceiling: if Grok ever returns a larger image, Anthropic caps it at 4,784 visual tokens. That gives about 7,100-7,600 input tokens.
Output tokens (inference). The schema now requires:
extremity_inventory: three region strings plus a total.anatomy_check: the parser keeps up to 600 chars.rewriteon a fail.
Medium-effort adaptive thinking also counts as output. Estimates:
- Low case: ~250 tokens (a clean pass with little thinking).
- Expected: ~500.
- Hard cap: 1,200 (
max_tokens).
Latency hints that the stricter screen really does think. The icon evaluator measured 8.6 s before the audit existed (customize-icons.js:68). The anatomy regression probe measured 17.0 s (customize-art.js:96). No post-audit usage log exists anywhere in the repo or the vault, so the output figure is the weakest number in this brief.
Cost per evaluator call at $2/$10:
| Case | Scene | Icon |
|---|---|---|
| Low (~250 output tokens) | $0.0099 | $0.0092 |
| Expected (~500 output tokens) | $0.0129 | $0.0120 |
| Cap (1,200 output tokens) | $0.0203 | $0.0193 |
| Absolute ceiling (4,784-token image + cap) | $0.0273 | $0.0262 |
How many fire (verified).
- Retry rule. Each image slot gets at most 2 attempts (
for (let attempt = 0; attempt < 2; ...)). Every attempt is a new Grok generation and a new evaluation. A failed verdict feeds itsrewriteinto the next prompt (customize-art.js:216,customize-icons.js:140). - When the retry does not fire. An evaluator error (timeout, 4xx, or a parse failure) breaks the loop with no retry (
customize-art.js:214). An exhausted time budget skips the evaluation, so the generation is paid for and never screened (:211-212). - Scene page. 1 slot, so 1-2 evaluator calls and 1-2 Grok images.
- Icon page. Up to
MAX_ICONS = 2(art.js:43), drawn in parallel, so 2-4 evaluator calls and 2-4 Grok images. The code comment says: "worst case is 2 icons x 2 attempts = 4". - Which pages qualify. Checked across all
content/games/*/slots.jsonon main:- Scene only: color-the-dino, color-the-dragon, color-the-soccer-star, colorea-al-unicornio.
- Icons only: bunnys-snack-maze.
- None has both. The client would fire both jobs if a page did (
src/scripts/customize.js:488-490), for a structural maximum of 6 generations and 6 screens.
- Words-only runs. A run where the parent never asks for a new picture (
brief.changefalse, emptyicon_briefs) fires zero evaluator calls.
Per Customize run, picture step only (Grok $0.02/image + screen):
| Run | Best (pass first try) | Expected (first-pass fail rate p = 0.3) | p = 0.6 | Worst (both attempts, output cap) |
|---|---|---|---|---|
| Scene page | $0.030 | $0.043 | $0.053 | $0.081 |
| Two-icon page | $0.058 | $0.083 | $0.102 | $0.157 |
| Of which: evaluator only | $0.010 / $0.018 | $0.017 / $0.031 | $0.021 / $0.038 | $0.041 / $0.077 |
The p values are illustrative. Nobody has measured the fail rate under the current rubric.
Aggregate bounds (verified). The evaluator is metered, just not modeled:
- Monthly. Every Grok attempt takes a slot from
CUSTOMIZE_ART_MONTHLY_CEILING(default 300,art-rail.js:16). The retry takes one too (customize-art.js:192). Screens can never outnumber generations, so there are at most300 screens a month: **$3.90 expected, ~$8.20 absolute ceiling**. - Per browser. Each browser is limited to 3 picture runs a day (
art-rail.js:14). - Weekly. Both endpoints route through
budgetedFetch(src/lib/generation-budget.js). It reserves a quote before each Anthropic call and settles on actualusageagainst the shared $19 rolling-7-day cap (LIMIT_MICROS,:5).- The reserve formula at
:73quotes 63,118 micros (~$0.063) per scene screen and 59,823 per icon screen. Settled actual is ~$0.013, so the reservation is ~5x too high. - "Uncertain calls keep their full reservation indefinitely" (
:3). A timed-out screen therefore holds ~$0.063 of the $19 week, and still that full ~$0.063 once the Grok call it screened is settled.
- The reserve formula at
Convergences and contradictions
- Contradiction, parent brief vs code: the evaluator is not unbounded. It is bounded per call (1,200 output tokens, 4,784 visual tokens), per run (2 attempts per slot, 2 icons), per month (300 generations), and per week ($19 ledger). What is true is that it is unmodeled: the charter, the engine spec, and the break-even tables all omit it.
- Contradiction, "image cost is noise" ([[2026-09-01-engine-build-spec]]):
- That was true for FLUX at $0.001 with an agent-side reviewer.
- Under Grok plus a paid Sonnet screen, the worst-case picture step is $0.08-0.16. That is 11-22% of the $0.71 text ceiling.
- The evaluator alone is ~32-49% of the whole picture line.
- Convergence:
generation-budget.js:19("Sonnet $2/$10 ... Grok $0.02/image") and the Anthropic pricing page agree. The runtime ledger has always been priced correctly. Only the planning docs are missing the line.
Synthesis for RDCO
The number the parent brief was missing is small, but it is not noise. A vision screen costs about $0.013. That is roughly the same as the entire Stage 1 words call and about two-thirds of the Grok image it screens. The evaluator is the largest per-image cost after generation, and it doubles on every failed verdict because each retry pays for a new image and a new screen. Adding it to the unit model, the picture step of a Customize run is ~$0.03-0.08 (scene) or ~$0.06-0.16 (two icons). The parent brief's "$0.06 for up to 3 images" is too high on image count (at most 2 slots per page today) and too low on multipliers (retries plus screens).
For pricing decisions this is a correction, not a rethink. The worst credible picture-heavy Customize adds ~$0.16 to generation cost. That moves the charter's credit break-even from $0.74 to about $0.87 per pack if every pack triggers an icon redraw at the worst case. The expected case ($0.04-0.08) moves it by a few cents. The dominant uncertainty is not the rubric token count, which is now pinned to about ±10%. It is the first-pass fail rate under the anatomy audit, which sets the expected column, and the thinking-token spend at medium effort, which sets the output line. Both can be read from the D1 ledger (charged_micros per Anthropic call on use_case='customization') without making any new call.
Two smaller design facts matter operationally:
- The rubric (~2,400-2,850 tokens) is re-sent uncached on every screen. At 300 screens a month, caching is not worth the engineering; this is not a build recommendation.
- The ledger's ~5x over-reservation, together with the "uncertain calls keep their full reservation" rule, means evaluator timeouts consume the $19 weekly budget far faster than real spend. Scene screens run close to their budget: a 17 s probe against a 25 s timeout. If the weekly cap ever trips before real spend explains it, look at evaluator timeouts first.
Why this is in the vault
It prices the evaluator line item that the Scribble Works credit tiers and break-even bands ([[2026-08-31-studio-charter]] §4, [[2026-09-05-kids-subscription-box-comparables-scribble-works]]) omit. It also corrects the "unbounded" label in [[2026-09-16-scribble-works-model-cost-ceiling-refresh]] before anyone uses it to argue for or against keeping the mandatory screen.
Open follow-ups
- What is the observed first-pass fail rate and mean settled cost per
customizationevaluator call in the D1 generation ledger since the 2026-09-06 anatomy audit landed? This one question turns the p=0.3 column from illustrative into measured. - How many output tokens (thinking plus JSON) does Sonnet 5 actually spend at
effort: mediumon the audit rubric, and how often does it hitmax_tokens: 1200? A cap hit truncates the JSON and turns a paid screen intoevaluator_unavailable. - What share of Customize runs request a picture at all (
brief.changetrue or non-emptyicon_briefs)? That fraction sets the per-pack, not per-run, evaluator cost.
Related
- [[2026-09-16-scribble-works-model-cost-ceiling-refresh]]
- [[2026-09-01-engine-build-spec]]
- [[2026-09-01-customize-spec]]
- [[2026-08-31-studio-charter]]
- [[2026-09-05-kids-subscription-box-comparables-scribble-works]]
Sources
- Vault:
06-reference/research/2026-09-16-scribble-works-model-cost-ceiling-refresh.md,01-projects/printables-product/2026-09-01-engine-build-spec.md(§6, spike retry multiplier),01-projects/printables-product/2026-09-01-customize-spec.md(Stage 1 cost),01-projects/printables-product/2026-08-31-studio-charter.md,06-reference/research/2026-09-05-kids-subscription-box-comparables-scribble-works.md,01-projects/printables-product/studio-notes-2026-09/customize-art/smoke-2026-09-02T02-10-10-attempt1.jpg(864x1152). - Code (scribble-works
origin/main@5ddf745):src/lib/customize/art-rail.js:14-16,121-128;src/lib/customize/art.js:23,43,185-298,438-497;functions/api/customize-art.js:21-28,88-98,183-237;functions/api/customize-icons.js:25-27,58-87,115-140;src/lib/generation-budget.js:3-5,19-22,73,98-112;src/scripts/customize.js:476-490;studio/notes/customize-art/smoke-2026-09-02T02-10-10.json;content/games/*/slots.json; historicalart.jsat commit0bed8a1. - Web: Anthropic vision docs (image token formula, high-res tier), Anthropic pricing (Sonnet 5 $2/$10, tokenizer +30%), xAI Imagine docs (1k/2k resolution tiers).