01-projects/printables-product

engine build spec

·build-spec·status: draft-for-founder
scribble-worksgenerative-engineworkers-aievaluatorcorpusunit-economics

Scribble Works generative engine — production build spec (from the $0 spike)

Source: engine spike run 2026-08-31 evening on branch feat/engine-spike (repo ~/Projects/scribble-works); filed under the 2026-09-01 dateline the dispatch pinned. Durability note: the evidence artifacts (scorecard, loop log, corpus) live on that unmerged branch — if the branch is ever deleted unmerged, re-export them to this folder first. Everything below is grounded in that run's artifacts: 24-image bake-off matrix, blind scorecard, a closed evaluator kickback loop on one real game, and a working corpus v1. Spike total spend: $0 (Cloudflare Free plan before and after; free-tier models only; partner-billed models excluded).

1. Winning image model + scorecard summary

Winner: @cf/black-forest-labs/flux-1-schnell. Blind fresh-eyes scoring (model identities withheld, print-scale 1275x1650 renders, design-critic axes vs [[../../02-sops/DESIGN-scribble-works]]):

model mean /20 range violations
flux-1-schnell 18.0 16-20 1 embedded-text (garbled watermark on line-art)
stable-diffusion-xl-base-1.0 12.3 9-16 3 non-white bg, 1 embedded-text
stable-diffusion-xl-lightning 10.8 8-14 4 non-white bg, 1 embedded-text
dreamshaper-8-lcm DQ (disqualified) — 4/6 prompts returned pure-black frames (suspected safety filter on kid-adjacent prompts) + 2 catastrophic renders

flux-1-schnell was the ONLY model that consistently produced white-background, crayon- textured, print-composed art — the scorer (blind) called its corgi "Ship this." Full detail: content/engine-spike/bakeoff/scorecard.md on the branch. Enumeration was live (/accounts/<acct>/ai/models/search?task=Text-to-Image, 10 models); the 5 partner-billed models (flux-2 family, Leonardo) were cut under the spike's $0 rule and remain untested.

Known flux weaknesses the evaluator must hold: occasional garbled watermark text (caught once on line-art), and multi-constraint compositions (the kids-spelling-OHIO body-letters spec) exceeded every model's ability — such slots need decomposed generation, a different tool, or human art.

2. Pipeline contract (3-pass, proven end-to-end)

  1. Text pass (Claude): writes game content + per-slot image specs. A slot spec = {subject, composition, slot_type: color|line_art}; the locked style block (sw-style-v1, in scripts/engine-spike/specs.json) is appended at generation time and versioned separately so a style change invalidates the corpus cleanly. Generation prompts are POSITIVE-ONLY (purple-elephant rule) — all prohibitions live in the evaluator. When the evaluator suggests prohibition-phrased fixes, the text pass translates them to positive phrasing before regeneration (worked twice in the spike).
  2. Image pass (Workers AI): corpus lookup FIRST (retrieval-before-generation); on miss, generate with the slot's configured model (default flux-1-schnell, 8 steps, 1024px). Engine contract facts from the spike: flux returns base64 JPEG inside JSON (not PNG); HTTP 400 code 8007 "NSFW content" fires false-positive on innocuous kid prompts and must be handled as a retryable prompt-rewrite signal, not an error.
  3. Evaluator pass (fresh-eyes subagent, zero context): strict per-slot contract check; on FAIL → kickback to pass 1 (prompt rewrite) and/or model switch on retry; on PASS → corpus add. Then assembly: HTML template → headless Chrome → PDF.

Spike proof: game blast-off-count-and-color — both slots FAILED ORGANICALLY in round 1 (6 stars where the counting answer required 5; stray objects on a "moon only" coloring slot), two prompt-rewrite kickbacks later both PASSED (evaluator independently counted 5 stars), final 1-page letter PDF assembled with the generated art placed. Loop log: content/engine-spike/game/loop-log.md.

3. Evaluator contract (v1, exercised)

Prohibitions (any breach = FAIL): photorealistic children (illustrated/cartoon kids are FINE — [[../../02-sops/DESIGN-scribble-works]] Photography rule) · any legible text/letters/numbers/watermark baked into art (scene-text judgment slots excepted per-slot) · non-white/paper background · line-art slots must be open colorable line art (no shading/fills/color) · print-quality at target size. Plus per-slot LOAD-BEARING facts (e.g., a countable-objects count) — the evaluator must verify them by observation, not trust the prompt. The spike's rounds show the count check catching 6-then-4-then-5.

4. Corpus scheme (v1, working)

Key = `sha256(sha256(canonical slot spec {subject, composition, slot_type}) + style_version

5. The Anthropic text door — FOUNDER-FACING FINDING

Founder ruling was "Anthropic-bias for text, billed through Cloudflare." Verified live 2026-08-31:

FOUNDER DECISION: the one-bill intent IS satisfiable via Unified Billing — at a 5% credit-purchase fee and a prepaid-credits mechanic (loading credits is a spend action behind the founder's billing gate). Alternative: BYO Anthropic key (two bills, no fee). Neither is needed while generation stays on the Max subscription (current ruling).

6. Unit economics (image side; every number cited)

Rates read live from developers.cloudflare.com/workers-ai/platform/pricing/ on 2026-08-31: free allocation 10,000 neurons/day; overage $0.011 per 1,000 neurons; flux-1-schnell 4.80 neurons per 512x512 tile + 9.60 neurons per step (equivalently $0.0000528/tile + $0.0001056/step).

7. What production needs that the spike didn't build

FOUNDER DECISION items in this doc: (1) Anthropic door — Unified Billing (5% credit fee, one bill) vs BYO key vs stay-on-Max; (2) the $5 Workers Paid flip timing (rung 2); (3) whether a paid partner-model bake-off is ever worth it.

Related

RESOLUTION ADDENDUM — founder rulings, 2026-08-31 ~21:39 ET (closes decision item 1; supersedes conflicting text above)