01-projects/printables-product

Scribble Works personal shopper — build spec (the magic moment, new front door)

2026-09-01·build-spec·status: spec
scribble-worksproductengine

Personal shopper — build spec

0. Why this exists

Founder, tonight: "Instead of the full browsing and filtering experience you type out what you want for the kid and it brings together a recommended 6 games to make a playset... a simple chat box and then the games float into the 6 outlined spots and text fades in below the completed playset explaining why these 6 were selected." He also said there are "no magic moment experiences so far" and no cohort gets the site until there is one. This is that moment and the new Home front door. Browse/filter stays as the power path.

Trust ladder. Rungs: manual → Customize → magic-wand auto-complete → fully generated. The shopper is rung-3 shaped and lands before rung 2. Fine: it carries rung-1 risk, selecting six existing human-made games and generating nothing.

1. The moment, in 6 beats

Home hero: six outlined slots (dashed paper frames, fanned 1-2°), a one-line text box under a Ray speech bubble: "Tell me about your kid and I'll set the table."

  1. Empty (0s). Six empty frames, the box, five chips (§2), placeholder cycling three example requests every 4s. Below: "or browse all {n} games".
  2. Enter (0-0.2s). Box locks, chips dim, bubble: "Reading your note...". An in-flight request disables the box. If the tray holds games not seeded by the shopper, the existing window.confirm(REPLACE_PROMPT) guard fires first.
  3. Loading (0.2-4s). Each empty frame draws its own scribble border (SVG stroke-dashoffset loop, 1.2s, 100ms stagger). Bubble advances at 1.3s ("Picking games...") and 2.6s ("Almost..."); at 4s "Still picking, one sec". At 11s the client gives up (failure state).
  4. Six games float in (response +0 to +1.0s). Cards enter slot 1→6, 110ms stagger, 420ms each, from translateY(24px) scale(.96) rotate(±3deg) opacity 0 to the slot's fanned tilt, easing cubic-bezier(.2,.8,.2,1). Each one-liner fades in 150ms after its card lands.
  5. "Why these six" (+1.25s, 500ms). Two sentences below the set, translateY(8px)→0 + opacity. Bubble: "Swap any card, or print the set."
  6. Tray + download. seedPlayset(slugs, title, {slug: 'shopper', title, blurb: rationale, games: slugs}); seed.slug is the reserved value shopper (shopper-fallback on fallback) so normalizeSeed accepts it and seedIntact() keeps the rationale on the PDF title page. PR3 renders <PlaysetTray/> on Home (today Home has no tray, so init() wires no download handler); "Print this playset" is the tray's unchanged path (7-page PDF, existing beacon). On pageshow/bfcache restore, re-render from the tray.

Failure state. Model down, validation fails twice, cap hit, kill switch, limiter unavailable, or client timeout → the fallback set (§3) floats in with the same animation and, in the rationale's place: "My picker is napping, so here's a set we built by hand for {age band} kids. Swap anything."

Thin-band state. Only 2 games are eligible for 7-8 or 9-10 (snack-shop-sums, treasure-boxes, the latter band via stretch), so 7_8+ always takes this path. If the eligible pool (band ∪ stretch) is under 6, no model call: in-band games first (slug alphabetical), then the nearest band, es games excluded unless Spanish was asked, low_ink first if that chip is set, then alphabetical; fill-ins wear a "younger pick" tag. Rationale: "We only have {n} games for {band} so far; the rest are our best picks from {adjacent band}."

2. Input contract

3. Selection engine: /api/shopper

Pages Function functions/api/shopper.js, same export shape as downloads.js (onRequestPost, catch-all 405); every response is a full playset JSON, never a 204. Catalog facets are embedded at build time from meta.yaml (slug = directory name, facets under taxonomy:), filtered to status: live with a printable (count = whatever survives; Home reads it from the same manifest): {slug, title, age_band, age_band_stretch, skill, skill_secondary, activity, theme, flags, language} ≈ 2.9 KB minified. PR1 adds taxonomy.language (en default, es on the four Spanish games).

Model. Claude Sonnet 5 via AI Gateway Unified Billing (RESOLUTION ADDENDUM). Structured outputs (output_config.format, schema below), thinking: {type: "adaptive"}, output_config.effort: "low", max_tokens: 2000 (thinking counts against it). PR1 verifies structured outputs and adaptive thinking work together on Sonnet 5; if not, drop thinking. No cache breakpoint in v1: the ~1,700-token prefix may sit under the minimum cacheable size, so the estimate assumes no cache hits.

Timing. First call timeout 5s; one retry at 3s; function ceiling 9s; client gives up at 11s.

Prompt shape. System: role ("you pick six of these games for one child; you never invent games"), the catalog JSON, rules: exactly 6 distinct slugs; each in the chosen band via age_band or age_band_stretch; ≥3 distinct activity types unless the parent asks for one kind; honor low_ink, two_player, and language if named; rationale = 2 warm sentences to the parent, kid's name if given, no outcome promises; one-liners ≤ 14 words; title ≤ 40 chars. User message: {text, chips} as JSON, labelled as data from a parent.

Response schema (function → client; the model emits the same minus source/note): {"source":"model|fallback|thin_band","age_band":"3_4","title":"Maya's dinosaur morning","rationale":"...","games":[{"slug":"color-the-dino","why":"..."} ×6],"note":null}

Validation (server): exactly 6; all slugs live; no duplicates; age_band in vocabulary and equal to the age chip if sent; every game in band ∪ stretch; rationale non-empty ≤ 400 chars; each why ≤ 120; title ≤ 40. stop_reason: max_tokens, a refusal, or non-JSON count as failures. Any failure → one retry with the validator's error appended to the user message; second failure → fallback.

Fallback rule (deterministic). Band = age chip, else a digit in the text, else 3_4. If a curated set in recommended-playsets.yaml matches the band (only 3_4 and 5_6 sets exist): language: es → primeros-pasos, else greatest overlap between the set's tags and words in the text, else first-day-of-school. No curated match → the thin-band rule and copy. The client ships the same sets and rule, so an unreachable function still produces a set.

Cost per call. Input ≈ 2,000 tokens (system + catalog + rules ≈ 1,700, parent text + chips ≈ 100, output schema ≈ 200). Output ≈ 600 tokens (300-400 real plus low-effort thinking). Sonnet 5 ($2 / $10 per MTok, claude-api table cached 2026-06-24): $0.004 + $0.006 = $0.010, $0.0105 per model request with the 5% Unified Billing fee; a retry is a second request. Haiku 4.5 ($1 / $5) ≈ $0.00525.

Recommendation: Sonnet 5. The pick is easy for either; the rationale is what the parent reads, and that copy is the product. Haiku 4.5 is a prior generation without effort control. Revisit only if latency, not cost, misses the 4s bar.

4. Client

src/scripts/shopper.js, loaded on Home. Posts to /api/shopper; on success seeds the tray (§1 beat 6) and plays the reveal via CSS classes toggled per slot with animation-delay from the stagger. prefers-reduced-motion: no transforms, no scribble loop (static dashed borders + bubble copy), opacity-only 200ms fades, all six within 300ms.

Swap one. Never re-asks the model: from the eligible pool (band ∪ stretch; language matched if the set has any es game), pick the highest-scoring game not in the set: same skill +2, same activity +1, low_ink +1 if the chip was set; tiebreak slug alphabetical. New card's why = first sentence of its meta.yaml blurb; the rationale shows "(you swapped one)" display-only, seed.blurb unchanged. PR2 exports replaceSlot(index, slug) from playset.js (today seedPlayset is the only export); it rewrites state.seed.games so seedIntact() stays true; tray unit tests re-run.

5. Guardrails

6. Build plan: 3 PRs, one daytime block

  1. PR1 function + fallback. functions/api/shopper.js, taxonomy.language, live-only catalog build step, validator + fallback + strip + thin-band as named exports with unit tests. Accept: 5 scripted prompts (dinosaur 4yo; Spanish 5yo; 7yo low-ink; "Leo is 3"; text containing an email) each return 6 live, age-honest slugs; the 7yo returns thin_band with exactly snack-shop-sums, treasure-boxes, big-bigger-biggest, build-the-word, count-school-supplies, finish-the-leaf-pattern; no email reaches the request; SHOPPER_ENABLED=false and unbound KV both return fallback; Gateway logging verified off.
  2. PR2 client + animation. shopper.js, slot markup, reduced-motion, replaceSlot, swap, bfcache, confirm guard. Accept: Playwright frames at 0, 0.5, 1.0, 1.5, 2.0s after response show order 1→6 with the rationale after the last card; reduced-motion shows no transform; double-Enter fires one request.
  3. PR3 Home front-door swap + copy. Hero becomes the shopper; the fanned spread becomes the six slots; <PlaysetTray/> on Home; "browse all {n} games" secondary; privacy sentence; FAQ entry. Accept: design-critic PASS; print works from Home; Browse in one click; no analytics script on Home.

Critic gates. behavior-critic exercises the live preview with the same 5 prompts plus the kill switch, source withheld. design-critic scores the PR2 frames as a strip against DESIGN-scribble-works.

7. Open founder rulings (Ray's default in bold)

  1. Freemium says generative features need a free account; the shopper calls a model with none. Default: allow it, no account; it selects rather than generates and the rate limit is the cap.
  2. The charter keeps generation on the Max subscription "until volume forces the API"; the shopper is metered API traffic from day one. Default: proceed; a ~$0.01 selection call is not the generation spend that ruling guards, and Unified Billing was ruled in for this door.
  3. Parent text leaves Cloudflare for Anthropic (30-day upstream retention). Default: yes, with the strip and one privacy-page sentence; the alternative is keyword matching and no magic.
  4. Over the monthly ceiling: silent fallback or a visible "picker resting" banner. Default: silent fallback with the honest one-liner.

Home layout is a Ray call: the shopper IS the hero.