01-projects/printables-product

Icon-set bake-off: Quiver Arrow 2 (sheet vs singles) vs Grok Imagine sheet

2026-09-19·decision-memo·status: gated-pass·source: Founder ask by iMessage 2026-09-19 08:43 ET; runs in scratchpad quiver-bakeoff/ and hunt-icons/; PR #321
scribble-worksprint-artquivergrokbake-off

Question

Founder (2026-09-19 08:43 ET): is Quiver easier to work with for a set of icons in one API call, and is the artwork better, versus Grok producing individual images that are cheaper and faster in total?

Fixture

The 12 Halloween Scavenger Hunt subjects (PR #321): jack-o'-lantern, skeleton, ghost, spider web, bat, witch, black cat, tombstone (RIP), scarecrow, inflatable, string lights, "Boo!" sign. Same line-art brief for all three arms.

Arms and results

Arm Calls Wall time Output tokens / cost Result
Quiver Arrow 2, ONE-call 3x4 sheet 1 ≈105 s (curl elapsed, hand-read, not logged to disk; lower bound 62 s from created to file write) 21,093 tokens ≈ $0.42 (≈3.5¢/icon) All 12 present, in the requested order, clean outlines. RIP misdrawn ("R:P"). SVG has no groups: 243 flat elements (205 paths, 38 circles/ellipses), so splitting needs bounding-box (bbox) clustering.
Quiver Arrow 2, 12 singles (parallel) 12 53 s 66,893 tokens ≈ $1.34 (≈11¢/icon) Best quality of the three. RIP and BOO! lettered correctly, inflatable has tie-downs, one icon per file, zero post-processing. Used in the game.
Grok Imagine via AI Gateway, two 2x3 sheets (gen-icon-sheet.mjs pipeline; both accepted sheets needed the uncommitted scratch sheet-runner.mjs: evaluator max_tokens 4000, letter-free descriptions) 2 accepted (sheet A via an eval-only re-run) + 1 declined run of 3 attempts n/a (sequential tool) $0.02 per generation call (provider cost_in_usd_ticks 200000000 on all 5 calls, 4 in _icon-sheet-gen-log.json, 1 in gen-b2.log), $0.10 total Clean, consistent raster. Refused lettering: the evaluator's text rule fails any letters, so BOO! and RIP came back blank; the cat is a plain sitting cat. Needs crop + evaluator.

Cost basis: Arrow 2 billing is token_usage, output rate 2,000,000 millicents per 1M tokens = $20 per 1M output tokens, input $4 per 1M (from GET /v1/models/arrow-2). Quiver figures are output tokens only; input adds $0.003 (sheet) and $0.034 (singles), so true totals are $0.42 and $1.37, ratio 3.2x. Grok figure is generation only and excludes the six mandatory Sonnet evaluator calls (5 attempts + 1 eval-only re-run, ~25k input tokens).

Findings

Recommendation

Authoring-time carve-out, not a change to the approved V1 engine (engine-build-spec: Grok Imagine via AI Gateway): for line-art icon slots use Quiver singles when the subject list is short (≤12); use the Quiver one-call sheet plus bbox split when a game needs many icons. Whether model lettering is allowed at all is the open ruling above. Keep Grok for scene-class art (gen-game-art.mjs), not icons. Extends the PR #311 spike, which favored Arrow on quality. Route Arrow spend through the ledger from the next batch.

Related