Question
Founder (2026-09-19 08:43 ET): is Quiver easier to work with for a set of icons in one API call, and is the artwork better, versus Grok producing individual images that are cheaper and faster in total?
Fixture
The 12 Halloween Scavenger Hunt subjects (PR #321): jack-o'-lantern, skeleton, ghost, spider web, bat, witch, black cat, tombstone (RIP), scarecrow, inflatable, string lights, "Boo!" sign. Same line-art brief for all three arms.
Arms and results
| Arm | Calls | Wall time | Output tokens / cost | Result |
|---|---|---|---|---|
| Quiver Arrow 2, ONE-call 3x4 sheet | 1 | ≈105 s (curl elapsed, hand-read, not logged to disk; lower bound 62 s from created to file write) |
21,093 tokens ≈ $0.42 (≈3.5¢/icon) | All 12 present, in the requested order, clean outlines. RIP misdrawn ("R:P"). SVG has no groups: 243 flat elements (205 paths, 38 circles/ellipses), so splitting needs bounding-box (bbox) clustering. |
| Quiver Arrow 2, 12 singles (parallel) | 12 | 53 s | 66,893 tokens ≈ $1.34 (≈11¢/icon) | Best quality of the three. RIP and BOO! lettered correctly, inflatable has tie-downs, one icon per file, zero post-processing. Used in the game. |
Grok Imagine via AI Gateway, two 2x3 sheets (gen-icon-sheet.mjs pipeline; both accepted sheets needed the uncommitted scratch sheet-runner.mjs: evaluator max_tokens 4000, letter-free descriptions) |
2 accepted (sheet A via an eval-only re-run) + 1 declined run of 3 attempts | n/a (sequential tool) | $0.02 per generation call (provider cost_in_usd_ticks 200000000 on all 5 calls, 4 in _icon-sheet-gen-log.json, 1 in gen-b2.log), $0.10 total |
Clean, consistent raster. Refused lettering: the evaluator's text rule fails any letters, so BOO! and RIP came back blank; the cat is a plain sitting cat. Needs crop + evaluator. |
Cost basis: Arrow 2 billing is token_usage, output rate 2,000,000 millicents per 1M tokens = $20 per 1M output tokens, input $4 per 1M (from GET /v1/models/arrow-2). Quiver figures are output tokens only; input adds $0.003 (sheet) and $0.034 (singles), so true totals are $0.42 and $1.37, ratio 3.2x. Grok figure is generation only and excludes the six mandatory Sonnet evaluator calls (5 attempts + 1 eval-only re-run, ~25k input tokens).
Findings
The founder's sheet bet holds on price: one call is ~3x cheaper per icon than singles. It loses on handling: Arrow returns flat elements, not groups, so a sheet needs a splitting step.
Singles win on quality and are drop-in SVGs. For a 12-icon game that is ~$1.37, still cheap against the review cost of a bad icon.
Grok stays the cheapest per call but is raster and cannot letter subjects under the current evaluator. Two pipeline gaps recorded in the PR #321 notes: evaluator
max_tokenstoo small for 6-cell sheets with--describe, and thetextrule contradicts lettered subjects. The committed tool cannot reproduce the "2 accepted" result today without the scratch runner's changes.Second Arrow batch, $1.80 ($1.76 output-only), billed to Quiver credits, not ledgered. The cost-basis review (2026-09-11, line 50) asked for ledger routing before a second batch; that did not happen here.
Letterforms are code-drawn per the customize spec (2026-09-01, "stays code-drawn, permanently: letterforms, numerals"). PR #321 shipped model-lettered BOO! and RIP inside the Quiver vectors. That needs a ruling before lettering counts as an engine criterion (escalation filed:
~/.claude/state/studio/escalations.md, 2026-09-19 item 4).The Quiver Model Context Protocol (MCP) server returned
model_errortheninsufficient_creditsin the ≈08:40-08:48 ET window (error text not captured beyond the tool's one-line codes) while the REST API with the 1PasswordQUIVERAI_API_KEYworked and the account had $22.66 (dashboard screenshotquiver-500.png, 08:48 ET). Use the REST path from scripts; treat the MCP as a convenience only.Founder read of the one-call sheet (09:57 ET): spider web, skeleton and tombstone weaker; the rest "passable for kids drawings"; "we don't want it to look too overdrawn." Follow-up run: the same three as singles with "as few lines as possible, no interior detail, something a 4-year-old could copy" appended. Result: visibly simpler on all three and 13,914 output tokens for the set (≈4.6k/icon vs 5.9k/icon on the original singles, ≈25% cheaper). Prompt language, not engine choice, controls the detail level. Now a contract rule: [[DESIGN-scribble-works]] amendment 2026-09-19 and design-critic Step 2.6. Files:
quiver-bakeoff/simple/.
Recommendation
Authoring-time carve-out, not a change to the approved V1 engine (engine-build-spec: Grok Imagine via AI Gateway): for line-art icon slots use Quiver singles when the subject list is short (≤12); use the Quiver one-call sheet plus bbox split when a game needs many icons. Whether model lettering is allowed at all is the open ruling above. Keep Grok for scene-class art (gen-game-art.mjs), not icons. Extends the PR #311 spike, which favored Arrow on quality. Route Arrow spend through the ledger from the next batch.
Related
scribble-works/studio/notes/2026-09-19-halloween-scavenger-hunt.md(PR #321 branch, not a vault note)- PR #311 Arrow 2 bake-off spike (2026-09-18)
- [[2026-09-01-engine-build-spec]] · [[2026-09-01-customize-spec]] · [[2026-09-11-unit-economics-cost-basis]] · [[DESIGN-scribble-works]] v2