01-projects/printables-product/audit-2026-09-02

Scribble Works — execution vs vision (audit D, 2026-09-02)

2026-09-02·audit·status: live
scribble-worksauditvisionexecutioninstrumentation

Execution vs vision — Scribble Works, 2026-09-02

Grade: Wide but shallow.

Not a criticism of craft. In roughly 48 hours the studio put 22 games, two magic moments, an account system, legal pages and a feedback rail into production, and the printed artifacts are good enough that the one outside human who has seen them sent back a single cosmetic note. The shallowness is not in the building. It is that every surface is one layer deep, none of it has been used by a stranger, and the counters that were built to notice a stranger cannot tell one from us.

0. Method and evidence quality

Everything below is read from the sources named in the frontmatter plus live infrastructure. Two corrections to my own working, recorded because they change what can be claimed:

1. Vision → shipped map

The left column is the founder's phrasing, quoted where it is his.

Vision element (his words) Shipped Pending review Specced only Absent
"Games are our unit" (2026-09-01 09:36) 22 live games, packs scrapped from the catalog, browse = 22 individual games with previews 5 games + 1 UGC variant merged but pending-review, awaiting one founder tap
Magic moment 1 — the Planner ("Plan today's playset") Live on / and /planner/. Chips + sentence both submit. 13 model calls this month Paid fill-on-demand generation when curation finds fewer than 6
Magic moment 2 — Customize ("Make it theirs") Live on every game page. Words (Sonnet), art (Grok), icons, and a code-GENERATED maze. 17 runs, 3 art, 1 icon Stage 2 server-rendered PDF Blocked on the $5 Workers Paid flip, still unclicked. Customize output prints from the browser, not a real PDF
Trust ladder rung 1 — "manual version to build the playsets" Persistent tray, 6 slots, playset PDF with parent page on top. 6 playset downloads recorded
Trust ladder rung 2 — "then customize a game" Live (above)
Trust ladder rung 3 — "then auto-complete a playset" (the magic wand) Queue item 19, marked UNGATED Not built. No spec written
Trust ladder rung 4 — "then completely generate playsets" Not built
"That turns into a need for push" — email delivery, the ladder's payoff Not built. No scheduled delivery, no reminder, nothing on the site
Marketplace grown by UGC variants (ruling 21) Server-side customization log live (D1 scribble-works-ugc, PII-stripped, anonymous per the 08:22 addendum). First variant hand-built colorea-al-unicornio merged pending-review, not live Nightly classifier written, "not armed until you nod" Zero variants have reached the marketplace. Corpus = 1 logged customization row
Households first (ruling 27: "start with parents/households") Accounts v1 live in production 13:16 ET today: magic link via Resend, __Host-sw_session, D1, /account, generative rail gated on a verified household, Turnstile v2: household adults, kid profiles, consent
"expand to teachers/schools", "treat them as a household for now" (ruling 27) Terms permit classroom printing. Teachers are ordinary households v3 designed in full: Organization, Org membership, Classroom, Seat, Enrollment, Link invite. Plus a five-state student-privacy law comparison No teacher or classroom surface on the site, correctly
Classroom tier later ("that would likely be a different tier… For now, I'm happy with anyone using it") Designed Correctly absent from the product
Freemium: no account browses and prints; free account unlocks capped generative; paid gets calendar + higher limits Anonymous browse and print live. Generative gated behind a verified account with "Three a day keeps this free" Paid tier: members-area spec, Stripe behind a flag No pricing, no credits, no checkout, no paid tier anywhere on the live site. Consistent with the step-1 memo's "ship free and instrumented"
Parent page on the front of every pack (charter §3) Footer and copy promise it; playset assembly builds it Not independently verified per playset in this audit
Feedback motion ("This is the sort of feedback motion I'd like to develop") Feedback-intake Worker built and deployed inert (no route, no secrets, no public URL). Email Routing + Resend live on scribbleworks.co Text door needs Twilio, founder must open it 0 rows in the feedback table. The one real feedback loop that has run went over the founder's iMessage, by hand
The site itself 22 games, browse with 6 facet groups, 4 ready-made playsets, FAQ, privacy, terms, account. No trackers, no ad pixels, one cookie No robots.txt, no sitemap.xml, no 404 page. Every unknown path returns HTTP 200 with the home page. scribbleworks.co is not attached to the Pages project

Charter items still unresolved. The charter is still status: open and still carries line 464: the founder "had NOT read the charter." The name was adopted anyway and now carries a domain, a live site, a privacy policy and an operating entity. "Launched" is still undefined, so the kill criteria remain dormant with no trigger condition. USPTO clearance was "promised to founder, this week" and has not run.

2. Delta since the [[2026-09-01-fresh-eyes-audit-synthesis]]

Nine of fifteen ranked targets addressed, three partial, three not. The three misses are the three that were about knowing rather than building.

# 09-01 target Status today Evidence
1 Apex R2 CORS DONE paper-playground allowed_origins now https://scribble-works.pages.dev, https://*.scribble-works.pages.dev, https://scribbleworks.co, https://www.scribbleworks.co
2 Download counter + 4-household print test HALF DONE Counter live and writing (21 keys, real values). Cohort PARKED by founder 09-01 18:22: "no magic moment experiences so far. I wouldn't want to give this to a cohort group yet." Never run
3 Publication gate, live-only DONE src/lib/game-status.js + games.js:127 filter on LIVE_STATUS (PR #13)
4 Honest age bands DONE, with a new hole Browse facets now read 3-4: 12 · 5-6: 12 · 7-8: 6 · 9-10: 4. But Ages 2-3 shows 0 games while the tagline and drawer both promise 2-10. Same species of dishonesty, moved one band down
5 Footer strip on every sheet DONE PR #44 wordmark footer + bubble audit across all 22; PR #18 answer/ask strips
6 Fix passive and unusable pages DONE Cut-and-glue v2 (PR #32), full regen
7 Defer engine/credits (23) and portal (18) ADOPTED, THEN OVERTAKEN Queue still reads "DEFER pending cohort read." Escalation #1 still logged OPEN. But ruling 25 (10:56) pulled accounts forward and ruling 28 (13:22) commissioned the members area, which is the portal. The defer was never lifted by decision; it was outrun
8 Honest CTAs, real FAQ PARTIAL Login CTA is now real. FAQ has six real answers under a heading that says "Five quick answers," and its first answer says "ages 3-4" against a 2-10 site. Terms says the headline is "no accounts, no trackers" while accounts are live
9 Single-game print + free sample DONE "Print this game (PDF)" on every game page
10 npm test owning publication truth + CI NOT DONE No test script. No .github/ at all. 20+ test:* scripts and a prebuild validator chain exist, but nothing gates a PR. PR #49 shipped on mocked tests alone and broke Customize in production for 6 minutes — precisely the failure this item predicted
11 Mobile craft at 375-390px PARTIAL Queue item D still open; a founder-reported planner overflow at ~1060px was dispatched 13:21 today
12 Track game source, pinned render PARTIAL render-games.mjs and bundled Andika shipped. Pinned Chromium not evidenced
13 Re-aim P1 at the empty lanes DONE, over-delivered 7-10 lane went from 2 games to 10
14 Charter resolution, define "launched", settle birthday vs age band NOT DONE Charter status: open, unread. "Launched" undefined. Birthday vs age band now has a third answer proposed (month + year, relationship-model Decision 2) which is itself pending-founder
15 Pre-publish checklist, machine tells MOSTLY DONE PR #44 bubble audit; queue item 24 (deterministic pre-check in prebuild) still open

The pattern in the misses. Items 2 (the cohort half), 10 (CI) and 14 (charter, "launched", kill criteria) are the only three that would have told the studio whether it was right. All three are open. Everything that produced an artifact got done.

3. Where execution ran ahead of the vision, and where it lags

Ahead

Accounts shipped to production before any non-founder had one. The production households table holds exactly 1 row, created 13:18 ET today, which per escalation #9 is the founder's own sign-in. A sign-in system, session store, Turnstile integration, magic-link mailer and generative access gate are live for a population of one, and that one is the person who built it.

A 13-entity relationship model for a phase the founder deferred. He said at 11:22 ET: "start with parents/households and we can expand to teachers/schools." What came back designs Organization, Org membership, Classroom, Seat, Enrollment and Link invite in full, plus five consent scopes, an append-only consent event log, a seven-actor by seven-field visibility matrix, a seven-row retention schedule, and a query-layer field-mask rule. The document's own bear case is the fairest verdict available: the consent layer "is the expensive half and it buys nothing a customer can see… plausibly as much work as everything else in v2 combined, built for a school phase just deferred and a legal regime that may not apply." Alongside it sits a five-state student-privacy law comparison for a school phase that, by the doc's own words, "None of it gates v2."

The members area went wider on a catalog that shrank underneath it. Yesterday's audit computed that ruling 12's five-day week view needed roughly 30 games per kid per week against a 14-game catalog, and recommended deferring it. The response was a seven-day calendar (ruling 28 addendum, correctly founder-driven) specced across 7 routes with 9 new decisions and a six-rule auto-fill, which is roughly 42 slots per kid per week against 22 live games. The spec acknowledges the shortfall and resolves it with paid generate-to-fill, which is rung 4 of a ladder whose rung 3 does not exist yet.

Feedback rails built, then superseded 74 minutes later. The feedback-intake Worker was dispatched 09:42 ET under ruling 24. At 10:56 ET ruling 25 said "before we develop that feedback feature more we should work on the account creation." The Worker is merged and deployed inert with zero rows.

A UGC promotion pipeline on a corpus of one. PR #55 logs customizations to D1 for a nightly classifier to promote themes into marketplace variants. The table holds 1 row (bunnys-snack-maze, English, 09:03 ET today). The classifier's demand threshold is a placeholder of ≥1. The one variant that exists was built by hand and is not live.

Legal surfaces published ahead of counsel. /privacy and /terms are live, name Ray Data LLC as operator, set Florida governing law, and were written by Ray. Decision 0 of the relationship model asks permission to source counsel and is an open blocking escalation. The privacy page states "It is not directed at children" and "We hold no child profiles of any kind" — both true today, and both things the next planned feature is designed to change. COPPA is never named on the page.

Lags

Rung 3 of the founder's own ladder does not exist. "then auto-complete a playset" — the magic wand — is queue item 19, marked UNGATED, with no spec and no build. Rung 4 and push email likewise. The ladder's stated purpose is burnout prevention ("Otherwise the parents will burn out on it… Same way Michelle runs the risk of burning out"). Nothing shipped yet reduces a parent's daily effort; the Planner reduces choosing, not doing. The studio climbed rungs 1 and 2 and then turned sideways into account infrastructure.

Discovery has had zero work. The [[2026-08-30-step1-memo|step-1 memo]] named discovery twice as the whole problem. Today the site has no robots.txt, no sitemap.xml, no 404 handling (every bad path returns 200 with the home page, which is actively bad for indexing), no external links, no distribution and no invited audience. scribbleworks.co is registered but not attached to the Pages project. The growth lane's three items (USPTO clearance, domain prices, trademark packet) are respectively not-done, not-done and founder-deferred.

Six finished artifacts are idle. Five pending-review games and the unicorn variant passed every gate and sit unflipped on one founder tap (owed since 09:00 ET).

Ages 2-3 is empty against a 2-10 promise in the tagline, the footer and the drawer.

No CI, and the cost of not having it has already been paid once in a production outage.

4. The "one real user" test

How many non-founder humans have used each surface, from the evidence:

Surface Non-founder humans Hard evidence
Printed sheet, in a child's hands 1 — [[2026-09-02-grammy-feedback-log|Grammy]], 2 sheets Feedback log; photos via founder iMessage 09:33 ET. She was handed paper. She has never touched the site
Live site, any page 0 confirmed No analytics by design; counters carry no unique or session dimension
Playset download 6 downloads total, 0 attributable to a stranger KV total:playset = 6; per-game adds spread across 17 slugs (first-sounds 4, color-the-dino 3, bunnys-snack-maze 3, count-school-supplies 3…)
Planner 13 model calls this month, 0 attributable KV month:2026-09 = 13
Customize 17 word runs, 3 art, 1 icon, 0 attributable KV total:customize = 17, total:customize-art = 3, total:customize-icons = 1
Account 1 household, created 13:18 ET today = the founder Production D1: 1 household, 1 state row
Feedback intake 0 D1 feedback = 0 rows, feedback_contacts = 0
UGC / marketplace variants 0 promoted D1 customizations = 1 row
Cohort test 0 of 4 households Parked 09-01 18:22, never run. [[2026-09-01-cohort-week1-kit|Kit]] frontmatter still says status: ready

The strategy stack rests on n=2, and one of them lives in the house. Michelle hand-making six games a day is the origin observation. Grammy is the only outside human on record, contributed one cosmetic note, and reached the product on paper.

The instrumentation is built and blind. The counters are bare integers. There is no unique-visitor, session, or first-versus-repeat dimension anywhere. The step-1 kill criterion asks for ten non-family downloads in ninety days. With what shipped, that number cannot be computed — not because traffic is low, but because nothing distinguishes a household from a stranger, or a second print from a first. The counter satisfies the letter of yesterday's audit item 2 and none of its purpose.

What this implies. The build queue is not the constraint and has not been for two days; 24 PRs merged and 13 production deploys in a single day proves it. The constraint is that nobody outside the house has tried any of it, and the founder's stated reason for parking the only test that would fix that has now expired: he parked the cohort because "no magic moment experiences so far." Both magic moments have been live since last night. The park is stale.

So: less should be built next, and more should be measured next. The single exception worth building is rung 3, because it is the founder's own next rung and it is the first thing in the ladder that reduces a parent's work rather than their choosing.

5. Risks the current path creates

Legal — the sharpest one. A real Delaware LLC's name is on a live privacy policy and terms of service written without counsel, asserting a specific posture ("not directed at children," "no child profiles"). The very next planned feature — kid profiles in the members area — falsifies both assertions. Decision 0 (authorize counsel) is blocking and open. The school-phase research already found that Florida's operator test is disjunctive and could be tripped by school-facing marketing alone, which is a live constraint on growth, not a v3 concern. Building v2 before the counsel answer is the one genuinely expensive-to-reverse move on the board.

Complexity. Thirteen entity types, five consent scopes and a field-mask enforcement layer for one household. Every one of those is code that must be maintained, migrated and reasoned about before a single stranger has printed anything.

Scope creep with a founder signature on it. Rulings 25 through 28 arrived within two and a half hours and each one enlarged the surface. That is legitimate — they are his rulings. But the effect is that the audit's two defer recommendations were neutralised without anyone deciding to lift them, and the escalation asking him to confirm the defer (#1) is still sitting open and now answers a question the world has moved past.

Quality outrunning the gate. Thirteen production deploys and one rollback in a day, with no CI, a broken mocked-test assumption already demonstrated, and 20+ test scripts that nothing runs automatically.

Founder time — and two charter falsifiers nobody is scoring. Charter §7 pre-registered three signals that mean the studio is the wrong shape. Signal 2 is "Founder time goes UP after the studio starts." He issued nine rulings today (20-28), answered three overnight escalations at 06:16, and has ten or more taps owed. Signal 3 is "Any nightly round requires a morning correction the founder did not ask for. Once is a bug; twice means unattended operation is not earned." The 06:16 ET font swap and Ray-mark rulings were exactly morning corrections on overnight work, and the wordmark size default is a second standing unanswered correction. The queue diligently tracks signal 1 ("0 consecutive no-pass rounds") and tracks neither of the other two. The two-week cron review is due ~2026-09-14 and should score all three, not one.

Status drift as a compounding risk. The cohort kit reads status: ready after being parked. Customize and shopper specs read status: spec while live in production. Terms tells visitors there are "no accounts." FAQ says ages 3-4. The charter says the founder has not read it. Anyone reading one document gets the wrong picture in both directions, and this audit needed the live infrastructure to correct two claims the documents alone would have produced.

6. Plain answer: how is execution compared to the vision?

Wide but shallow.

The vision has been honoured on its two hardest, most specific asks. Games are the unit, and there are 22 of them with packs properly scrapped. Both magic moments are live, and the second one has a code-generated maze behind it because the founder pushed back on an artificial limit and was right. The printed artifacts survived contact with a real grandmother and needed one redrawn carrot.

But the product is one layer deep everywhere and nobody outside the household has been in it. The ladder stops at rung 2 while the org built accounts, a consent model, a school-phase legal appendix and a seven-route members area — the last three for phases the founder deferred in his own words. The one test that would tell anyone whether the paper is worth printing twice was parked for a reason that no longer holds, and the counter built to replace it cannot tell a stranger from us. The charter's own falsifiers, two of which are plausibly tripped, are not being scored.

The studio is not failing. It is doing the thing the charter's §7 bear case predicted word for word: "Build the org, build the machine that feeds the org, let the machinery become the work… Output is not the constraint." Output has been extraordinary. It is still not the constraint.

Three moves for the next week, phrased so they can be answered yes

1. "Yes — flip the five games and the unicorn live, then run the four-household print test this week." You parked the cohort on 09-01 because "no magic moment experiences so far." Both magic moments have been live since that night, so the reason has expired. Six gate-passed artifacts unlock on one tap, and the test is the only thing on the board that produces evidence rather than output. Ask Ray for one change first: make the counter record first-visit versus repeat, so the answer is a number and not a vibe.

2. "Yes — freeze accounts at v1 and authorize counsel today." No kid profiles, no members-area build, no consent layer until a children's-privacy attorney answers the two questions in Decision 0. Ray sources three names this week and you pick or decline. This is the only expensive-to-reverse item on the board: a live privacy policy already asserts "no child profiles," and v2 is designed to make that false. Everything else can wait a week at no cost.

3. "Yes — build the wand next, not the members area." Rung 3, "auto-complete a playset", is your own next rung and the first feature that reduces a parent's work instead of their choosing. It runs on the catalog you already have. Ship it behind the same account gate, and put one round into a npm test plus GitHub Actions on every PR while it builds — the six-minute Customize outage on 09-02 was the bill for not having that, and it will not be the last one.

Related