Scribble Works generative engine — production build spec (from the $0 spike)
Source: engine spike run 2026-08-31 evening on branch feat/engine-spike (repo
~/Projects/scribble-works); filed under the 2026-09-01 dateline the dispatch pinned.
Durability note: the evidence artifacts (scorecard, loop log, corpus) live on that
unmerged branch — if the branch is ever deleted unmerged, re-export them to this folder
first. Everything below is grounded in that run's artifacts: 24-image
bake-off matrix, blind scorecard, a closed evaluator kickback loop on one real game, and a
working corpus v1. Spike total spend: $0 (Cloudflare Free plan before and after; free-tier
models only; partner-billed models excluded).
1. Winning image model + scorecard summary
Winner: @cf/black-forest-labs/flux-1-schnell. Blind fresh-eyes scoring (model
identities withheld, print-scale 1275x1650 renders, design-critic axes vs
[[../../02-sops/DESIGN-scribble-works]]):
| model | mean /20 | range | violations |
|---|---|---|---|
| flux-1-schnell | 18.0 | 16-20 | 1 embedded-text (garbled watermark on line-art) |
| stable-diffusion-xl-base-1.0 | 12.3 | 9-16 | 3 non-white bg, 1 embedded-text |
| stable-diffusion-xl-lightning | 10.8 | 8-14 | 4 non-white bg, 1 embedded-text |
| dreamshaper-8-lcm | DQ (disqualified) | — | 4/6 prompts returned pure-black frames (suspected safety filter on kid-adjacent prompts) + 2 catastrophic renders |
flux-1-schnell was the ONLY model that consistently produced white-background, crayon-
textured, print-composed art — the scorer (blind) called its corgi "Ship this." Full
detail: content/engine-spike/bakeoff/scorecard.md on the branch. Enumeration was live
(/accounts/<acct>/ai/models/search?task=Text-to-Image, 10 models); the 5 partner-billed
models (flux-2 family, Leonardo) were cut under the spike's $0 rule and remain untested.
Known flux weaknesses the evaluator must hold: occasional garbled watermark text (caught once on line-art), and multi-constraint compositions (the kids-spelling-OHIO body-letters spec) exceeded every model's ability — such slots need decomposed generation, a different tool, or human art.
2. Pipeline contract (3-pass, proven end-to-end)
- Text pass (Claude): writes game content + per-slot image specs. A slot spec =
{subject, composition, slot_type: color|line_art}; the locked style block (sw-style-v1, inscripts/engine-spike/specs.json) is appended at generation time and versioned separately so a style change invalidates the corpus cleanly. Generation prompts are POSITIVE-ONLY (purple-elephant rule) — all prohibitions live in the evaluator. When the evaluator suggests prohibition-phrased fixes, the text pass translates them to positive phrasing before regeneration (worked twice in the spike). - Image pass (Workers AI): corpus lookup FIRST (retrieval-before-generation); on miss, generate with the slot's configured model (default flux-1-schnell, 8 steps, 1024px). Engine contract facts from the spike: flux returns base64 JPEG inside JSON (not PNG); HTTP 400 code 8007 "NSFW content" fires false-positive on innocuous kid prompts and must be handled as a retryable prompt-rewrite signal, not an error.
- Evaluator pass (fresh-eyes subagent, zero context): strict per-slot contract check; on FAIL → kickback to pass 1 (prompt rewrite) and/or model switch on retry; on PASS → corpus add. Then assembly: HTML template → headless Chrome → PDF.
Spike proof: game blast-off-count-and-color — both slots FAILED ORGANICALLY in round
1 (6 stars where the counting answer required 5; stray objects on a "moon only" coloring
slot), two prompt-rewrite kickbacks later both PASSED (evaluator independently counted 5
stars), final 1-page letter PDF assembled with the generated art placed. Loop log:
content/engine-spike/game/loop-log.md.
3. Evaluator contract (v1, exercised)
Prohibitions (any breach = FAIL): photorealistic children (illustrated/cartoon kids are FINE — [[../../02-sops/DESIGN-scribble-works]] Photography rule) · any legible text/letters/numbers/watermark baked into art (scene-text judgment slots excepted per-slot) · non-white/paper background · line-art slots must be open colorable line art (no shading/fills/color) · print-quality at target size. Plus per-slot LOAD-BEARING facts (e.g., a countable-objects count) — the evaluator must verify them by observation, not trust the prompt. The spike's rounds show the count check catching 6-then-4-then-5.
4. Corpus scheme (v1, working)
Key = `sha256(sha256(canonical slot spec {subject, composition, slot_type}) + style_version
- model)
. Store = image file named by key +manifest.json(entry: spec hash, style version, model, slot type, subject, evaluator=PASS, timestamps). Only evaluator-PASSED images enter. Lookup precedes every generation. Spike demo: HIT (moon spec re-run retrieved, zero regeneration) and MISS both in the loop log. Production home: R2 under the existing bucket; the spike corpus lives on the branch atcontent/engine-spike/corpus/`. Exact-match keying means any wording change misses — a normalization or semantic-similarity layer is a later optimization, not v1.
5. The Anthropic text door — FOUNDER-FACING FINDING
Founder ruling was "Anthropic-bias for text, billed through Cloudflare." Verified live 2026-08-31:
- The Workers AI catalog contains zero Anthropic models (live search: "anthropic" → 0, "claude" → 0). Anthropic text cannot run as a native Workers AI model.
- AI Gateway supports Anthropic two ways
(developers.cloudflare.com/ai-gateway/usage/providers/anthropic/, read 2026-08-31):
- BYO key — the bill goes to Anthropic; Cloudflare is proxy/observability only.
- Unified Billing (developers.cloudflare.com/ai-gateway/features/unified-billing/, read 2026-08-31) — Cloudflare supplies the key and usage lands on the Cloudflare bill via prepaid credits. "Inference pricing from providers is passed through with no markup," but "A 5% fee is applied to all credits purchased through Unified Billing." Credits are purchased in the dashboard; auto-replenish available.
FOUNDER DECISION: the one-bill intent IS satisfiable via Unified Billing — at a 5% credit-purchase fee and a prepaid-credits mechanic (loading credits is a spend action behind the founder's billing gate). Alternative: BYO Anthropic key (two bills, no fee). Neither is needed while generation stays on the Max subscription (current ruling).
6. Unit economics (image side; every number cited)
Rates read live from developers.cloudflare.com/workers-ai/platform/pricing/ on 2026-08-31: free allocation 10,000 neurons/day; overage $0.011 per 1,000 neurons; flux-1-schnell 4.80 neurons per 512x512 tile + 9.60 neurons per step (equivalently $0.0000528/tile + $0.0001056/step).
- One flux image as the spike ran it (1024x1024 = 4 tiles, 8 steps): 4x4.80 + 8x9.60 = 96 neurons = $0.001056.
- Spike-observed evaluator retry multiplier: 5 generations for 2 passed slots = 2.5x (small n=2 slots; the corpus pulls this down over time by eliminating regeneration of known-good slots).
- Per customization, modeled on the locked playset shape (6 games + 1 parent title page, the founder's "lucky 7" ruling — [[2026-08-31-marketplace-taxonomy-proposal]]), assuming ~2 image slots/game = 12 slots: 12 x 2.5 x 96 = 2,880 neurons ≈ $0.032 at paid rates — or $0 inside the free allocation, which covers ~3 such customizations/day (10,000 / 2,880).
- What a credit must price at: image cost is noise. The binding generation cost is TEXT — the studio charter's ceiling estimate is $0.71/pack on Sonnet ([[2026-08-31-studio-charter]] §4, which itself flags its model rates as cached 2026-06-24, not live; that caveat carries through here). Image adds ~$0.03. A credit is margin-positive above roughly $0.74 generation cost + payment overhead — i.e. the charter's standing design rule (full-burn contribution must not fall as price rises) is unchanged by the image engine; the image side moves the floor by ~4%.
7. What production needs that the spike didn't build
- Rung-2 machinery (charter §3): Worker front door, Queue, Durable Object job state, D1 credits ledger — and the $5/mo Workers Paid flip (founder billing gate).
- Evaluator as a dispatchable production gate (the spike used ad-hoc fresh-eyes subagents; production wants the design-critic-style skill wrapper + the NSFW-400 retry handling).
- Corpus on R2 + eviction/versioning when
sw-style-v1bumps; semantic/normalized keying. - The designed page template (spike template is bare; footer/speech-bubble collision noted) and the parent-overview page integration.
- Multi-constraint slots (body-letters class): decomposed generation or human art — unsolved by every model tested.
- Partner-model lane (flux-2/Leonardo): untested under the $0 rule; a paid bake-off is a founder call if flux-1-schnell quality ever caps out.
FOUNDER DECISION items in this doc: (1) Anthropic door — Unified Billing (5% credit fee, one bill) vs BYO key vs stay-on-Max; (2) the $5 Workers Paid flip timing (rung 2); (3) whether a paid partner-model bake-off is ever worth it.
Related
- [[2026-08-31-studio-charter]] — org, rungs, unit-cost basis this doc extends
- [[2026-08-31-marketplace-taxonomy-proposal]] — the playset-shape ("lucky 7") ruling the per-customization model uses
- [[../../02-sops/DESIGN-scribble-works]] — the style contract the evaluator enforces
- Branch artifacts:
content/engine-spike/(bakeoff + scorecard + game + corpus),scripts/engine-spike/(specs, bakeoff.mjs, pipeline.mjs, corpus.mjs)
RESOLUTION ADDENDUM — founder rulings, 2026-08-31 ~21:39 ET (closes decision item 1; supersedes conflicting text above)
- Unified Billing RULED IN: AI Gateway with prepaid credits (no auto-refill = the spend gate), 5% markup accepted at current scale. Watch-item: revisit direct provider integration once a specific model/family proves sticky.
- V1 engine configuration (founder-approved): text = Anthropic via AI Gateway · image = Grok Imagine via AI Gateway (supplementary blind bake-off 2026-08-31 late: Grok 24.5 vs FLUX 20.5 mean, wins 5/6; scorecard in session scratchpad frontier-bakeoff/, key finding = FLUX hallucinated a "Disney" watermark on the dino coloring page — unshippable class) · FLUX schnell = free fallback · evaluator watermark/embedded-text screen is a MANDATORY gate regardless of model.
- Both models failed the body-letters (OHIO-arms) slot class — prompt engineering (per-child pose specs) or code-drawn hybrid remains the answer there; carried as a production note.
- Not yet tested: Nano Banana (no Google key wired — honest skip), gpt-image-2 (wrapper wired, classifier-blocked; one founder approval runs it if ever wanted).
- Setup queued for production build: $5 Workers Paid flip + AI Gateway creation + credit prepay — inside the founder's <$20/mo standing authority (2026-08-31), announce-on-execution.