Scribble Works — strategy and spec audit (fresh eyes, 2026-09-01)
Scope: six documents only (charter [[2026-08-31-studio-charter]], [[2026-08-31-product-model-rulings]], [[2026-09-01-engine-build-spec]], [[2026-09-01-recommended-playsets-spec]], [[2026-08-31-marketplace-taxonomy-proposal]], studio queue). No UI audit. Blunt by instruction.
Headline: the body of work has built a credits ledger, a generative engine, a corpus scheme, a freemium ladder, and a paid week-view plan on top of exactly one observed behavior: one mother hand-making six games a day for one three-year-old. No document contains a mechanism to learn whether any second household prints twice. The charter's own §7 said this, the founder never read it, and everything written since has moved further away from it.
1. Coherence check — contradictions and unstated load-bearing assumptions
1.1 The charter is officially OPEN and unread, yet every later document treats it as law
- Charter, Resolution attempt: "founder clarified he had NOT read the charter and wants to talk it through... All three decisions back to OPEN." Frontmatter still says
status: open. - Same document, §8 amendment one day later: "the §8c circuit breaker's batch size moves from ONE artifact per round to up to FOUR" — an amendment to a circuit breaker on a cron that is, per the same page, held.
- Queue header: "Manifest + charter + session cron all updated." Queue P1: "Round 1, 2026-09-01:
count-school-suppliesshipped." The nightly studio is running. - Playsets spec §5: "Author: product role (charter §2 ...)"; engine spec §7: "Rung-2 machinery (charter §3)".
The founder's 09:36 message ("trying to up our build time without my intervention, especially overnight") is a de facto greenlight, but nobody wrote the resolution. Every role boundary, escalation line, and the QA gate map are being enforced from a document the founder has not accepted. When he does read it, anything he amends has already been built against.
1.2 Child profile: birthday vs age band — direct contradiction, and the portal round is ungated
- Charter §6: "Birthday, not an age band." and "Why birthday, reversing the earlier sketch... Founder specified birthday."
- Product-model ruling 11 (2026-09-01 14:45): "one profile per kid (name, age band, interests, focus skills)".
- Taxonomy proposal, Facet 1: "the parent-portal child profile, which per the charter keys on birthday."
- Queue item 18: "UNGATED 2026-09-01 14:45... portal round — household account + per-kid profiles".
The design round that instantiates the profile is ungated with the field unresolved. Also notice the charter's languages + dominance and per-skill progress (introduced / practicing / secure) fields vanished from ruling 11 entirely. Nobody noted the drop.
1.3 The second loop the founder wants is prohibited by the privacy stance he confirmed
- Charter §6: "Child data drives generation and never leaves household scope... Not aggregated across households, not used to train anything."
- Product-model ruling 13: "(b) 'upping the quality across the whole service.' He doesn't yet know how (b) manifests."
(b) is, by definition, cross-household aggregation of upload data. Uploads are photos of marked-up pages: the child's handwriting, frequently the child's name (kids write their name on worksheets), the parent's handwriting, whatever is on the table. This is the most sensitive artifact in the whole system and it is the one earmarked for aggregation. Neither document acknowledges the collision. See §4 for how to thread it; the point here is that it is unstated.
1.4 The paid plan consumes ~30 games per kid per week; the catalog has 14
- Product-model ruling 12: "week view of 5 day-cards, each a playset". Playsets spec §1: "EXACTLY 6 slugs" (validator fails otherwise). 5 × 6 = 30 games/week/kid, 35 printed pages.
- Playsets spec §6: "14-game catalog as of 2026-09-01". Launch candidates reuse
trace-abc,trace-the-shapes,star-dot-to-dotin 4 of 5 sets. - Queue P1 round 1 observed throughput: one game per night, not four.
A paid subscriber exhausts the entire catalog inside the first week. The only way the week view is fillable is rung 4 (fully generated) by default — which the trust ladder says parents must "slowly build trust" toward. The paid product as specified skips the ladder it claims to climb. Nobody has written the number 30.
1.5 Age range 3-10 vs a catalog and taxonomy that are all preschool
- Taxonomy ruling 4: "Copy must not market toddler-only... not the audience ceiling." Queue item 13: "de-toddler the marketing copy sitewide."
- Every slug in the playsets spec (
trace-abc,first-sounds,color-the-dino,bunnys-snack-maze,who-eats-what) is a 3-5 activity. Recommended playsets carry no age band in their data shape (tags: themes, skillsonly) despite Facet 1 being the first thing a parent needs to know. - Product-model ruling 11 puts
age bandon the kid; the playset that will be recommended to that kid has no age band to match against.
De-toddlering the copy while the product is 100% toddler is exactly the mismatch a 7-year-old's parent will notice in ten seconds.
1.6 "Work self-arrives" vs a queue that is entirely founder iMessage rulings
- Charter §1: "Work self-arrives... The founder is not a task queue."
- Queue: virtually every item cites a founder message to the minute (08:57, 09:36, 15:24, 15:28, 15:55, 15:57, 16:07, 16:37, 17:42, 19:35, 20:10, 20:20, 20:34, 21:26, 21:39, 14:45). Three of the charter's own falsifiers are arguably already tripped: "Any nightly round requires a morning correction the founder did not ask for" — 17:42 "I don't see a way to add to playset" and 08:57 "game meta carries NO pack/page concepts... Never reintroduce" are morning corrections. Nobody is scoring the studio frame against its own pre-registered failure signals.
1.7 Instrumentation is on the critical path and absent from the queue
- Charter §2 growth row: "Discovery + the instrumentation the kill criteria depend on." §2 named gap: "Growth's instrumentation is the only feedback closing that loop, which is why it is on the critical path." §4: "ship free and instrumented and not to price until there is a non-zero number in the dashboard."
- Queue growth items: 3 (USPTO check), 4 (domain price pulls), 9 (trademark memo, DEFERRED). Zero instrumentation items. There is no dashboard. There is no download counter. The kill criterion ("ten non-family downloads in ninety days") cannot currently be measured.
1.8 Build order is inverted: credits before accounts before any Customize user
- Queue P4: "Accounts/login — AFTER P3 defines requirements. P5 billing — last, post-Customize + cohort signal."
- Queue item 23 (same day): "Production engine build: AI Gateway + Unified Billing prepaid credits... $5 Workers Paid flip... This is the week's flagship build." Engine spec §7 puts "D1 credits ledger" in the same build.
- Product-model: "Free account: unlocks generative features with a capped number of free generative uses" — the cap is never specified anywhere.
"Billing last, post-cohort signal" and "prepaid credits this week" are in the same file. The engine is a supply-side build with zero demand-side evidence; the founder's own ruling at 19:35 ordered it the other way.
1.9 Charter tier-2 text is superseded but not amended
- Charter §3: "Generation stays local on the founder's Max subscription until volume forces the API" and the "check back in a few minutes" consequence.
- Queue 20:10: "any external-user-triggered generation = API + credits." Engine spec addendum: Unified Billing ruled in. The charter's most-discussed product consequence is dead and still printed as live.
1.10 Cost model measures the wrong unit
- Charter §4 prices a 7-page pack, 4 critic rounds at $0.71. Engine spec §6 leans on it: "The binding generation cost is TEXT — the studio charter's ceiling estimate is $0.71/pack."
- Customize (product-model, taxonomy) regenerates one game. Wand fills empty slots. Neither is a 7-page pack. The spike measured image retries (2.5x, n=2) and recorded no text token count for the actual pipeline. Every credit-pricing sentence rests on a number for a product shape that no longer exists.
1.11 "Launched" is undefined, so the kill clock may never start
- Charter: "Kill criteria DORMANT until 'launched' is declared - no clock during development."
- Stated strategy: natural discovery, tiny cohort, distribution deferred. Under that strategy nothing ever gets "declared launched," and the only pre-registered discipline in the body of work never activates. Someone must define launched = "first non-household print," or the clock is decorative.
1.12 Smaller items
- Taxonomy open question 1 ("exactly 6, or 'up to 6'?") was never answered in the resolution log; the playsets validator hard-codes 6.
- Product-model ruling 11: "Per-kid pricing = open pricing lever" directly conflicts with the charter's household-is-the-billing-unit design and the 240-970 subscriber break-even, which was computed per household.
- Charter §2: no role owns the cohort. Product escalates "any commitment to a person outside the household"; growth escalates "every outbound email to anyone but the founder." The stated strategy is four outside households and there is no queue item to onboard them.
2. The riskiest unvalidated assumption, and the cheapest test
Assumption: that the daily-playset rhythm is a parent need and not Michelle's idiosyncrasy — i.e., that a parent other than the founder's wife will print a second playset without being asked.
The trust ladder's whole engine is stated in the product-model: "Otherwise the parents will burn out on it. Same way Michelle runs the risk of burning out handcrafting these games each day." Most parents never handcraft; there is no habit to burn out on, and therefore no felt need for push. Every downstream artifact (freemium caps, credits, week-view plan, push email, upload loop, corpus economics) assumes repeat printing exists. n=1, and that one is the co-founder.
Cheapest experiment (this week, zero build):
- Hand each cohort household (sister ×2 kids, teacher, neighbor) one printed 7-page playset and the browse URL. Do not explain the product. Do not create accounts.
- Say one sentence: "Text me a photo of any page when it's done." That is the upload loop, run over iMessage.
- Measure three things over 7 days: (a) did they print a second playset unprompted (ask on day 7 if no signal), (b) did any photo arrive, (c) did anyone mention the title page. Also note which kid ages are in the cohort; if none are 6+, the 3-10 claim stays untested.
- Pass bar: 2 of 4 households print again within a week. Fail: the plan/push/credits stack gets deferred until the first-print experience itself is the work.
Second-cheapest and worth running in parallel: on the same four printers, does the low-ink claim hold and do the PDFs render clean (footers, cut lines)? Nobody has printed on a non-household printer.
3. Queue critique — re-ranked by leverage toward "prints, comes back, pays"
| Rank | Item | Why here |
|---|---|---|
| 1 | (MISSING) Cohort onboarding + return instrumentation | Nothing else can be learned without it. First-party download counting on R2/Pages (no child profile exists at tier 1, so the privacy rule permits it) + the §2 hand-delivery test. Owner: growth. |
| 2 | P1 new-game production, re-aimed | Catalog depth is the hard cap on a second visit (14 games = 2 playsets before repeats). Bias to ages 5-8 and to the "unrollable sets" the playsets spec identified. Four-per-night is the plan; one-per-night is the observed rate — fix the rate before adding surfaces that consume games. |
| 3 | Item D — mobile nav visibility bug at 375px | Parents arrive on phones. A nav that "reportedly not VISIBLE at 375px" is a first-print blocker, not polish. Verify on a device today. |
| 4 | Item 2 — playset title page template | The charter's own trust artifact ("a straight answer to 'why these pages, in this order'"). It is the one thing that distinguishes this from free supply, and it ships on every print. |
| 5 | Item 20 — footer-free PDFs / clean "N of 7" | The object in the parent's hand. Small, and it compounds across every print. |
| 6 | R2/R3 — recommended playsets, but two sets not five | Cuts choice paralysis on first print. Ship "First Day of School" and "Learning to Read"; the other three are 60% the same games and make the shelf look thin. Add age band to the data shape first. |
| 7 | Round B layout + item 13 copy sweep | First-visit comprehension: what is this, for what age, what does my kid get. Fold Home game-type highlights in. |
| 8 | Item 17 Customize, as concierge | The differentiator, but run it Wizard-of-Oz for the cohort: parent texts "make the dino one about trucks for Maya," Ray generates on Max, sends the PDF. That validates whether anyone asks before the engine is productionized. |
| 9 | Item 24 — deterministic pre-check layer | Pure throughput: catches dumb failures before a critic burns capacity. Directly raises the games-per-night rate. |
| 10 | Items 3 + 4 — USPTO check + domain prices | Cheap, promised, and "clearance precedes visibility" is correct. Keep them cheap. |
Over-invested (relative to evidence):
- Item 23 — production engine with AI Gateway, Unified Billing, prepaid credits, $5 flip, D1 ledger, "the week's flagship build." Zero external Customize requests exist. Everything built here is a two-way door, but it also consumes the week's studio capacity that games and cohort learning need.
- Item 18 portal round — designs a paid week-view plan for a product with no accounts, no repeat-print evidence, and a catalog that cannot fill it (§1.4).
- Engine corpus scheme, semantic keying, style-version invalidation — capital accumulation for a factory with no orders.
- Item 25 frozen-scenario eval harness for DESIGN-*.md edits — process infrastructure for a 14-game catalog.
- Item 12 Ray situational poses, Round A palette variants, item 22 wordmark B2 — taste work is fine overnight, but three design items are queued ahead of a nav that may not render on phones.
- Taxonomy anti-slop machinery — correct design, but a 5-facet filter rail over 14 games returns empty for most combinations. Do not build the filter UI (item 15) until ~40 games.
Missing entirely:
- Any download/print counter or dashboard (§1.7).
- Cohort onboarding owner and script (§1.12).
- Age 6-10 content lane (§1.5).
- A definition of "launched" (§1.11).
- The free-tier generative cap number.
- Real-printer QA outside the household.
- A written resolution of the charter's open status (§1.1).
- Any tracking of the studio's own three falsifiers (§1.6).
4. The second loop — how uploads could raise quality across the whole service
Precondition for all three: an explicit opt-in on the upload page ("help us improve the games — we keep what we learn, not the photo"), storage of derived facts only (game slug, age band, outcome fields, note text with names stripped), and the photo deleted after extraction. That is the only shape that survives the charter's "not aggregated across households" line; amend the charter to say so rather than pretending the loop doesn't need it.
Mechanism A — Field critic: parent margin notes become studio findings
- What: the vision pass reads the parent's handwriting ("too easy," "she didn't get the instructions," "loved the dino") and the free-text field, classifies it against a closed set (too easy / too hard / unclear instructions / loved it / print problem / other), and files it to the studio queue as a critic finding against that game slug. A game with two "unclear" notes goes back through the gate with the notes as the spec. The queue already has the shape: "Failed-gate artifacts return here WITH findings."
- Data needed: slug, note text, one classification label, age band. No image retained.
- Useful from: upload number one. A single "instructions unclear" from a real parent outranks any fresh-eyes critic.
- Cohort scale: Realistic and the best fit. Under 10 kids is exactly the sample where qualitative notes carry more information than counts. This is the mechanism to build first, and the iMessage photo experiment in §2 is its manual prototype.
Mechanism B — Broken-game detector / completion calibration
- What: per upload, the evaluator extracts {completed?, correctness on load-bearing items, evidence of struggle (erasures, adult handwriting doing the work)}. Aggregate per slug across households: completion rate and error pattern. Feeds game meta: a corrected age band, a difficulty tag, and a "commonly missed" note. Also flags games nobody ever uploads (never finished, or never printed).
- Data needed: slug, age band, completed flag, correct/total, struggle flag. Numeric only.
- Useful from: ~5 uploads per game to say "this one is broken"; ~30+ per game to recalibrate an age band with any confidence.
- Cohort scale: Partially realistic. With 14 games and 5 kids you get broken-game detection within a month; you do not get calibration. Honest framing: at cohort scale this is a smoke alarm, not a thermostat.
Mechanism C — Progression priors for the wand and recommended sets
- What: sequences of (playset → outcomes → next playset → outcomes) across many kids reveal which orderings move a skill from introduced to secure fastest. Those priors become the wand's balance logic and the default order inside recommended playsets, and eventually the "gap detector" the playsets spec describes gets driven by demand rather than by curation instinct.
- Data needed: per-kid longitudinal sequences with skill-state transitions, age band, hundreds of kid-weeks minimum; de-identified household IDs.
- Useful from: low hundreds of active kids over 2-3 months.
- Cohort scale: Not realistic. Anything claimed here under 10 kids is anecdote dressed as data. Design the schema so the transitions are recorded from day one (cheap), but do not build any logic on it.
Order: A now (manual via iMessage, then the upload page), B's schema now and its aggregation when uploads per game pass five, C schema only.
5. Three things to kill or defer
- Defer item 23 (production engine + Unified Billing + credits + $5 flip) until a non-household parent has asked for a customized game; run Customize as a concierge on Max for the cohort instead.
- Defer item 18 (portal / paid week-view plan) until repeat printing is observed and the catalog can fill more than one week without repeats; the design will otherwise encode the birthday-vs-age-band contradiction and a 30-games-per-week promise the product cannot keep.
- Kill item 25 (frozen-scenario eval harness for design-contract edits) and shelve items 12 and 15; process and filter infrastructure for a 14-game catalog is the "machinery becomes the work" failure the charter's §7 predicted.