What "accepted" means for a Scribble Works game (spec v0)
Why this exists
The founder (08:06 ET, 2026-10-01): "I think we've been working on the pedestal. Account login, marketing copy, transactional emails, subscription payment plans, brand design. Those are all solved areas... if we can nail the AI educational assistant first and creation of the worksheets, then we have trained our monkey." Ray's reading: the "monkey" is an AI that makes a good tailored sheet for one child, cheaply, with no human fixing it. We can't hill-climb game CREATION until "good" is a yes/no definition. This spec is that definition, version 0. Founder, 09:45: "I'd prioritize the game creation hill climbing before the game curation hill climb. I think it's easier to quantify than the curation." So creation is climbed first; curation second.
The two numbers we drive
- Accept rate: the share of generated games that pass the rubric below. We climb on the 16 per-check pass rates; the accept rate is the result they add up to. While it sits near zero, it doesn't show which check to fix.
- Cost per accepted tailored sheet: total generation spend for a batch, including any automatic revisions, divided by the number of ACCEPTs in that batch. A game a human fixed does not count as an ACCEPT. Founder, 08:34: "we could give ourselves some more wiggle room and say a coloring book costs $10, but an 8 or 5 cent target sounds great." Target 5-8 cents (founder). Ceiling about 14 cents (Ray's arithmetic): $10 a month for one child, less card fees of 3.6% + $0.30, is $9.34; at a 70% margin, 30% of that ($2.80) is left for costs; $2.80 over 20 sheets is $0.14. The 70% margin and 20 sheets are Ray's assumptions; the allowance is still open in PR #332, which calls 20 "unapproved illustrative". This is an upper bound: it gives the whole 30% to generation, but curation, email, storage, rendering, infrastructure and refunds come out of the same 30% (MONETIZATION-PLAN.md lines 138-148, PR #332), and a household with a free third child brings in less per child. Library games are excluded: they are made once and reused by every household.
Measured today: nothing for the thing this metric counts. No original-creation sample exists (GENERATION-COSTS.md lines 5, 10), and no game has been graded with this rubric yet. The nearest data are revisions of library games: one successful preview revision at $0.145144 and one rejected trial at $0.161792 (GENERATION-COSTS.md line 9); an unattended production revision of the Knight maze at $0.112, rejected by the final critic and the founder (issue #304); and six #304 debugging runs on the Knight maze, $0.478432 in total, each on a different Worker build, none able to pass because of the page's locked layout (PR #316). None of these is a tailored sheet. GENERATION-COSTS.md's illustrative cost for one original attempt is $0.15 (lines 28-29), above the 14-cent ceiling even if every attempt passed. Both levers are open: the cost of one attempt has to fall, and the accept rate has to rise. (Correction sent to the founder 2026-10-01 ~10:2x ET; Ray had earlier called the $0.145 run "accepted" and the six debug runs a 0-accepted batch.)
The rubric (v0.1, 16 yes/no checks)
Printable version: rubric.pdf (sent to the founder 2026-10-01). Every check is a yes/no question about the page, not a 1-5 score (blog: "checkable claims"). Graders may circle "?" for can't-tell. A "?" counts as not-yes for ACCEPT, and in the human-versus-judge comparison it is counted separately, not as agreement or disagreement.
A. Works on paper (each becomes a code check where possible) A1 fits the page · A2 readable · A3 pictures identifiable without labels · A4 clear finish · A5 solvable, key correct · A6 targets big enough for the age · A7 works in black and white.
B. Teaches what it claims (LLM judge, calibrated against human grades) B1 most time goes to the claimed skill · B2 right for the age · B3 directions explainable in one sentence · B4 every item correct · B5 builds or self-checks.
C. Fun (human or real-kid signal; the judge may guess, kids overrule) C1 a reason to play · C2 fits the attention span · C3 the kid chooses or makes something · C4 they'd ask for another.
Verdict: ACCEPT = all A yes, all B yes, at least 2 of C yes. REVISE = would pass with a fix you can name. REJECT = the idea fails.
How the gold set works
- The founder (and Michelle if she's willing) grades games on paper. Ray grades the same games separately and holds his grades back until the human grades arrive, so the humans aren't anchored.
- Per check, we compare the human and Ray answers. Disagreements are the useful output: each one either fixes the rubric wording or teaches the judge.
- An LLM judge counts as trustworthy only when all three hold on human-graded games it was not tuned on: (a) its ACCEPT or not-ACCEPT verdict matches the human verdict on at least 9 of 10 games, with both outcomes among them; (b) on no single A or B check does it say N where the human said Y on more than 1 game in 10; (c) run twice on the same game, it gives the same verdict (blog: "runs the grader twice on the same output"). The judge does not use the generator's model (blog: "it should not be the model you are testing"). This needs at least 10 held-back games, so batch 1 alone can't meet it. Ray proposes these thresholds; they await the founder's read. The bar is high because ACCEPT needs all 12 A and B checks to pass: at a 5% chance per check of saying N where a human said Y, the judge would fail about 46% of the games the humans accept.
- Two sets are held out before any tuning. (a) Generator inputs: the requests the generator is climbed on (age band, skill, interests) are split at random into train and test (blog: "splits the evaluation set at random into test and train"). A generator change is kept only if the judge-graded accept rate on the test requests holds or improves. (b) Judge calibration: some human-graded games are kept back and used only to score the judge, never to tune its prompt.
- Real games come first; synthetic cases are only variations seeded from real ones.
- Batch 1 (sent 2026-10-01): baby-animal-mamas (2-3, Science & Nature), count-the-garden (3-4, Numbers & Counting), first-sounds (3-4, Letters & Phonics), bedtime-routine-sequence (5-6, Logic & Puzzles), making-change (9-10, Math). Bands and skills are from each game's meta.yaml; the printed pages show neither, so the list was sent to the founder separately (~10:2x ET). All five are live library games, so batch 1 will be mostly Y. Batch 2 adds known failures, starting with the founder-rejected Knight renders (scribble-works docs/evidence/304/before-*.png), so the judge is tested on games that should fail; the blog puts production samples and bug reports first. Store-bought games can be graded with the same sheet; each one costs grading time, the scarcest input.
Proposed: a sixth taxonomy facet, learning objective (founder decision)
taxonomy.yaml has five frozen facets (age band, skill, activity, theme, materials/flags). "Skill" is coarse ("Math"). Today B1 ("teaches what it claims") and curation can check only the coarse skill (for example, "Math"). A learning objective lets them check the specific target, for example "counts objects to 10" or "identifies the first sound of a word". Proposal: add learning_objective as facet 6, with values taken from real frameworks rather than invented ones: the Head Start Early Learning Outcomes Framework (ELOF) for ages 2-5, and the Common Core State Standards (math and English language arts) and the Next Generation Science Standards (NGSS, science) for K-5. Gap: none of these has a strand for Logic & Puzzles, the largest skill among live games for ages 5-10 (10 of 27), or for Art & Drawing (1 of 27). Those need a stated mapping (for example, to the Common Core Standards for Mathematical Practice) or a fourth source; otherwise their objectives would be invented. A sixth facet is a founder decision under the taxonomy's own guardrail.
What would make this spec wrong
- Cost binds before accept rate does: GENERATION-COSTS.md's illustrative original attempt ($0.15) is above the 14-cent ceiling even at a 100% accept rate. If the first measured creation attempts cost more than 14 cents each, raising the accept rate alone can't reach the target.
- The judge can't be trusted: ACCEPT needs 12 checks to pass, so a judge that is right on most checks can still be wrong on many verdicts. Until it meets step 3, the accept rate it reports is not evidence.
- The gold set teaches the wrong lesson: if it holds only shipped library games, the judge never sees the failures the generator makes.
Open questions for the founder
- Is ACCEPT's bar right (every B must be yes)? Batch 1 tests it: for each of these live games that the founder would keep in the library but the sheet does not ACCEPT, the check that failed it gets reworded or dropped in v0.2.
- Does Michelle grade too (a second expert means we can measure human-to-human agreement)?
- Approve facet 6, and the ELOF / Common Core / NGSS sources?
Rubric v0.2 (2026-10-01 evening, from the founder's batch-1 grades)
Files: rubric/accepted-rubric-v0.2.{html,pdf} (v0.1 kept beside it). Changes, each traced to his sheets (gold-set/2026-10-01-founder-grades.md):
- A1 now names pictures, numbers, borders and labels (he found clipped badges, kite tail and border on making-change).
- A6 adds room for safety scissors and no words near cut lines (bedtime kid play: "Words were cut. Safety scissors do not work well in tight spaces").
- A8 NEW art quality (his top complaint, on every sheet: "Poor drawing quality all around!").
- A9 NEW kid page holds only the game (he raised moving Answers / Ray says / Ask to a parent page on 3 of 5 sheets).
- B2 split into B2 (can start after one explanation) and B3 (not trivially easy); old B3-B4 renumbered B4-B5.
- B6 (was B5) eased for ages 2-4: a matching picture or count is enough to "tell they got it right" (Ray marked N where he marked Y on 2 of 2 young sheets).
- Header adds "Skill it claims" and "Kid who played". 19 checks; ACCEPT rule unchanged (all A, all B, at least 2 C). Michelle's and his dad's packets use v0.1: map v0.1 B3→B4, B4→B5, B5→B6; v0.1 B2 counts for both B2 and B3; A8/A9 blank.