01-projects/printables-product

Scribble Works — fresh-eyes onboarding review (astra)

2026-09-06·review·status: draft

Founder-requested onboarding review by astra (gpt-6-astra, high reasoning effort), run via ~/.claude/scripts/ask-model.sh, the third independent fresh-eyes pass on this repo in two days (after codex/gpt-5.6-sol on release vs main, and grok/grok-4.6-build on main directly). Read-only: cloned the repo, read the requested documentation/notes/migrations/configurations/core flows, examined the last 30 commits, ran npm ci + npm test -- --no-bail in a disposable clone with PP_LIBRARY pointed at an empty directory. Reviewed main at ca40a91a026b986975fb1022c20cc565da724c4a, committed 2026-09-05 21:57:32 EDT — the same commit grok reviewed, and the post-merge state of the commit codex reviewed as release.

No garbling. Unlike the grok run, this raw output (codex exec with -o writing a clean final payload, not an interleaved thinking stream) required no cleanup — read the full 277-line raw capture end to end and it is coherent throughout, single well-structured document, no duplicated/interleaved fragments. Reproduced below verbatim (only the internal /tmp/rdco-deleg.rP5rCB/... and clone-relative source links are kept as astra wrote them — they resolve on astra's disposable clone, not this vault, but preserve the exact file:line citations for verification).

Verification note (Ray, 2026-09-06): astra correctly identifies both sw-daily-playset and sw-retention-sweeper as "Independently deployed" in its own runtime table below — it does not make grok's "undeployed" error. Astra does pose two open questions about operational status ("Is retention deployed and running successfully? … last successful run and orphan-photo inventory?" and "Is daily delivery enabled remotely?") — those are legitimately unanswerable from a read-only repo clone with no Cloudflare account access, and are answered here from live ops knowledge rather than corrected as errors: sw-retention-sweeper has been deployed since 2026-09-02 (cron 20 4 * * * UTC); sw-daily-playset is deployed and actively sending — this session sent a real household email through it earlier today, 2026-09-06. No astra claim below is factually wrong about deployment state; the caveat is scoped to those two open questions, not to a correction.


Astra's review (verbatim)

The main concern is household authorization. I reproduced a removed adult regaining access to their former household as its lead. I also found cross-household tray leakage, delivery to removed adults, and retention failures. I would address those before expanding account features or unattended delivery.

Reviewed main at ca40a91a026b986975fb1022c20cc565da724c4a, committed September 5, 2026, 21:57:32 EDT. I cloned the public repository directly, read the requested documentation, notes, migrations, configurations, and core flows, and examined the last 30 commits.

The review clone is unchanged. Installation and tests ran in a disposable copy with PP_LIBRARY pointing to an empty directory. I did not access the founder vault, change source files, push, open a PR, deploy, or send messages.

Verification. npm ci succeeded, followed by npm test -- --no-bail: 21 gates passed, 0 failed, approximately 49 seconds. The suites reported 1,762 passing checks, 0 failures.

Suite Passed
Auth / accounts / account APIs 281 / 243 / 160
Customize / Shopper / UGC / maze 404 / 92 / 52 / 53
Playset curation / daily delivery / storage 174 / 124 / 10
Downloads / slots 33 / 19
Feedback / retention 63 / 18
Headers / taxonomy / playset validation 23 / 8 / 5

The build produced 41 Astro routes, plus 31 standalone game HTML documents; all 31 game PDFs were reachable and 62 preview images were generated. Both copies remained clean in Git.

This ran on Node 25.9.0; CI specifies Node 22. I did not run live model requests or deployed account flows. The additional reproductions below used production modules with in-memory SQL or controlled browser-state/network fixtures.

System mental model. This is a static publishing system with several independent services attached to it — not one application deployment.

Runtime Responsibilities and dependencies
Build and studio tools Validate taxonomy, playsets and customization slots; derive bundled catalogs; fetch published PDFs from R2; render previews; build Astro; generate CSP hashes. Operator scripts classify UGC, triage feedback, prepare replies and change plans. Several operational dependencies live outside the repository.
Browser Browse/filter the static catalog, assemble a six-game tray, run the Planner UI, customize a worksheet, print/download, and manage accounts. Anonymous state is localStorage; signed-in trays synchronize to D1.
Cloudflare Pages Serves Astro output and standalone game HTML. Pages Functions provide auth, account APIs, tray storage, Planner, customization, generated-art retrieval, PDF proxying, download counts and client-error intake.
Daily-playset Worker Independently deployed. Hourly cron at :30 selects due Plus households, deterministically curates games, packages one PDF per child, sends through Resend and records deliveries. A token-protected /run supports operator execution.
Feedback-intake Worker Receives Cloudflare Email Routing messages and Twilio SMS/MMS webhooks. Writes feedback/contact rows to D1, photographs to private R2, and inbox pointers to KV. Optional Discord notification. Replies are a separate operator script.
Retention-sweeper Worker Independently deployed nightly job at 04:20 UTC. Removes expired UGC, feedback/photos, orphan contacts, and spent/expired auth records. No destructive HTTP endpoint.

The root configuration is active: it contains pages_build_output_dir = "dist". Top-level bindings describe preview; production bindings are under env.production. Preview has its own accounts database but shares the production UGC database and all three KV namespaces. (wrangler.toml:55, wrangler.toml:68)

The two logical D1 databases have these intended schemas:

Migration Intended database Main contents
0001_customizations.sql UGC Customization outputs, stripped request text, hashed household proxy, classification/promotion metadata and retention deadline.
0002_feedback.sql UGC Feedback, photo-key references, and a separate hash-to-raw-contact lookup for replies.
0003_accounts.sql Accounts Households, provider identities, magic links, sessions and synchronized household state.
0004_counters_and_retention.sql UGC Atomic spending counters and retention indexes.
0005_accounts_v2_household.sql Accounts Adults, identity links, children, household events/waitlist, consent records/events, deletion requests, and reserved classrooms/seats.
0006_plans_and_deliveries.sql Accounts Free/Plus plan and delivery preferences; delivery history unique by household/date/kind.
0007_session_adults.sql Accounts Session-to-acting-adult mapping.
0008_session_identities.sql Accounts Session-to-verified-identity mapping.
0008_client_errors.sql UGC Client diagnostic events.

There are nine migration files, including two numbered 0008. Most relationships are enforced in application code rather than foreign keys; household_plans → households explicitly uses a cascading foreign key.

The identity model is the hardest part to reason about:

Those overlapping representations are directly involved in the highest-severity findings.

The core flows work as follows:

The documented trust ladder is "catalog-local → website JIT → MCP → scheduled → turnkey, with turnkey out of scope," not the exact manual/customize/auto-complete/generated sequence suggested in the prompt. I found studio/customize provenance and deterministic generated slots, but no implemented public system for arbitrary full-game generation or an MCP runtime in this repository. (README.md:15)

Recent velocity. The last 30 commits span September 2–5. September 5 alone includes adult invitations, multi-household membership, packaged daily PDFs, PDF proxying, telemetry, and a security/privacy repair. Earlier commits add games, repair printable layouts, and revise privacy disclosures. The newest security commit, 78f6091, adds important checks, but its actor-resolution check runs after the fast path exploited below.

This is rapid product and security development occurring together. Tests and explanatory notes are extensive, but lifecycle transitions and interactions between separately completed features receive less coverage.

Hard constraints found in the repository. These are documented or coded requirements; inclusion here does not mean every path satisfies them.

Constraint Source
Games are standalone; pack/pack_page provenance must not return to the game schema. README.md:4
Printable changes require a one-page Letter render with approved fonts, followed by human visual inspection. PRE-PR-CHECKLIST.md:39
Raw auth tokens must never be persisted; only their hashes belong in D1. auth/tokens.js:5
State-changing requests require same-origin checks in addition to the session cookie. auth/tokens.js:147
Only active leads manage membership, children and consent; maximum three adults and eight children. accounts/rules.js:9, accounts/rules.js:67
A household must retain an active lead; adults cannot remove themselves. accounts/store.js:485
Parents cannot change their plan through the account API. accounts/plans.js:5
Accounts migrations must never be applied to the UGC database. 0003_accounts.sql:6, 0005_accounts_v2_household.sql:9
The household migration deliberately avoids ALTER TABLE and foreign-key references. 0005_accounts_v2_household.sql:16
Dependent processing must check live consent; ugc_theme_log is described as parent-switchable. 0005_accounts_v2_household.sql:149
Models must not write answer keys or geometry; generated-slot parameters are validated and rendered by code. customize/engine.js:46, customize/engine.js:462
Every generated image requires a successful vision screen. customize/art-rail.js:136
Twilio fails closed without its token; fetched media must use allowed HTTPS hosts and respect a streaming size cap. twilio.js:54, twilio.js:70
Retention deletes must be bounded and use explicitly selected IDs. sweep.js:25
Feature PRs target release; the founder ships release → main. Changes to model requests require a live smoke. PRE-PR-CHECKLIST.md:60, PRE-PR-CHECKLIST.md:114

Code quality findings, ranked. "Reproduced" below means exercised locally against the reviewed code, not against customer data.

  1. High — removed adult can become the former household's lead. Reproduced.

    When an existing identity accepts an invitation to another household, the handler "repairs" a missing home identity link by linking it to ensureLeadAdult(home). That helper returns an existing lead without matching the invitee's email. Removing an adult previously deleted precisely that identity link.

    A subsequent magic-link login stamps the wrongly linked lead. resolveActor() accepts that stamp before checking the session's verified identity.

    My fixture followed: join A → removed from A → accept invitation to B → sign in again. The resulting session acted as A's original lead and successfully created a child in A through the real handler, returning 201.

    Sources: accept-invite.js:145, accounts/store.js:120, auth/store.js:184, accounts/store.js:443.

    Remedy: Resolve authorization from the verified identity and matching active membership on every request. Never repair an identity link to a different email. Audit existing inconsistent links and add removal/reinvitation lifecycle tests. The legacy fallback that selects a lead also needs elimination or strict containment.

  2. High — household switching can copy one household's private tray into another. Reproduced at the client-module boundary.

    The browser caches its household storage scope for the page's lifetime. Switching household changes the server-side session shared by every tab, while only the switching tab reloads. An older tab then uploads its cached A tray to /api/state; the server selects B from the now-switched cookie. The request contains no expected household.

    A controlled fixture produced exactly that upload. Besides disclosure to B's other adults, it can overwrite B's tray.

    Sources: playset-storage.js:28, account-sync.js:51, account-households.js:48, state.js:35.

    Remedy: Include explicit household context in requests, authorize it server-side, and reject stale context. Broadcast sign-in/switch/sign-out changes across tabs and cancel pending synchronization. Other account forms should receive the same review.

  3. High — daily playsets can continue going to a removed founding adult. Reproduced recipient selection.

    The Worker sends to households.email, without requiring that address to remain an active member. A founding lead can promote another adult and then be removed. Removal leaves the contact address unchanged.

    My fixture ended with only current-lead@example.com active, while the daily recipient remained former-founder@example.com. The delivered material includes children's names and personalized information.

    Sources: deliver.js:52, deliver.js:119, accounts/store.js:524.

    Remedy: Model delivery recipients explicitly and require current, verified membership. Removal must disable or replace affected recipients; revalidate before sending.

  4. High — retention loses photo references and exceeds D1's parameter limit. Both reproduced.

    sweepFeedback() catches failed R2 deletions, then deletes all selected feedback rows anyway. Missing R2 bindings behave similarly. A failed-delete fixture returned one deleted row and zero deleted photos, leaving an object outside normal cleanup tracking.

    Separately, delete batches contain up to 500 bound parameters, while D1 documents a 100-parameter limit. A fixture enforcing that limit left 101 expired rows undeleted. This can stall cleanup as the backlog grows.

    Sources: sweep.js:32, sweep.js:58, sweep.js:89.

    The passing test named "failed R2 delete keeps the row" never asserts that the row remains. (retention test:118)

    Remedy: Delete only rows whose dependent objects were successfully removed, or retain a durable cleanup task. Use D1-compatible batch sizes and test actual deletion postconditions and boundary sizes.

  5. High operational risk — migration configuration does not separate the databases.

    The preview accounts runner explicitly uses migrations/; the UGC binding uses the same directory. Other bindings do not override Wrangler's default migration directory. Comments saying "ACCOUNTS ONLY" do not filter SQL files.

    Applying all nine files alphabetically to an empty SQLite database succeeded and created both concerns together. Therefore the documented migration commands do not enforce the separation they promise.

    Sources: wrangler-accounts-preview.toml:17, wrangler.toml:122, wrangler.toml:266.

    Remedy: Give Accounts and UGC separate migration directories and explicit configuration everywhere. First inspect each deployed database's schema and migration ledger; reorganizing files without reconciling existing history could cause a second problem. Add a routing test that asserts forbidden tables never appear in either database.

  6. High privacy-readiness gap — consent controls and account deletion are largely dormant.

    The schema says dependent reads check live grants, but normal child creation does not record consent, and customization/daily processing does not consult it. Consent helpers exist chiefly as storage functions and tests. UGC request retention is controlled by a deployment variable rather than the documented per-parent scope.

    Household deletion records a deadline and revokes sessions, but there is no exposed deletion endpoint or sweeper implementation for that deadline. The public page honestly says deletion is manual; whether that manual process covers all stores cannot be established here.

    Sources: children/index.js:27, customize.js:191, accounts/store.js:409, sweep.js:152.

    Remedy: Define the actual product consent/deletion lifecycle, then implement it across accounts, saved HTML, deliveries, UGC and object storage. Do not treat the presence of consent tables as evidence that consent enforcement exists.

  7. Medium — saved customization has incompatible schemas, limits and asset lifetimes.

    Client entries retain values, language, answers, name, art and creation time. Server normalization keeps only slug, title, HTML and an optional differently named timestamp. A round trip drops metadata and defaults restored language to English.

    The aggregate server limit is 64 KB, while an individual customized page may be roughly 400 KB. Six existing source documents already exceeded the aggregate limit in my fixture. Sync failures are swallowed.

    Finally, serialized HTML references generated images that expire after seven days; restoring on another device after expiration can yield missing artwork.

    Sources: auth/state.js:38, playset-custom.js:18, account-sync.js:51, customize/art.js:25.

    Remedy: Share a versioned persistence contract between browser and server, choose a realistic storage strategy, expose sync failure, and make saved-art lifetime match the saved-playset promise.

  8. Medium — name redaction does not satisfy its stated behavior. Reproduced.

    redactName() drops name components shorter than two characters and does not implement its claimed accent-insensitive matching:

    • Name V, description V loves horses → unchanged.
    • Name José, description Jose loves horses → unchanged.

    The resulting description is used in the model request and customization log.

    Sources: customize/engine.js:100, customize.js:242.

    Remedy: Fix these specific cases, test Unicode normalization and name variants, and narrow absolute privacy copy to guarantees the system can actually enforce. Regex stripping cannot reliably identify every name embedded in arbitrary prose.

  9. Medium — anonymous telemetry permits unbounded durable writes and trusts client scrubbing.

    /api/client-errors accepts unauthenticated inserts without rate limits or retention. Server validation accepts arbitrary message text and caller-selected timestamps. The browser scrubber therefore is not a storage privacy boundary: a direct caller bypasses it.

    Sources: client-errors.js:16, client-errors.js:34, 0008_client_errors.sql:16.

    Remedy: Prefer allowlisted diagnostic codes and server-generated messages/timestamps; add bounded intake and expiration. Since this shares UGC D1, telemetry abuse can affect unrelated product services.

  10. Medium — the "global" generative brake does not stop the Planner. Reproduced helper behavior.

    GENERATIVE_ENABLED=false disables customization but is ignored by shopper.isEnabled(). The code describing a global switch over four endpoints is incorrect. Anonymous Planner access is explicitly intentional; this finding concerns incident shutdown, not its access policy.

    Sources: auth/guard.js:36, shopper.js:61.

    Remedy: Apply the global cost/incident brake independently of authentication, with one test proving every paid rail stops making provider calls.

  11. Medium — daily email is not idempotent at the external side effect.

    Delivery history is checked before processing, but the message is sent before the unique delivery row is inserted. Concurrent invocations—or successful sending followed by a database failure—can send duplicates. Conversely, a recorded transient failure suppresses retry for the entire day.

    Sources: deliver.js:155, deliver.js:177, send.js:20.

    Remedy: Use a durable claim/outbox, provider idempotency key, explicit attempt states and bounded retry. A uniqueness constraint written after sending cannot enforce "one email."

  12. Medium product risk — daily curation quickly produces short sets with the real catalog. Reproduced.

    For one English-only five-year-old, five consecutive simulated days yielded 6, 4, 5, 4, 4 games. Repeat candidates exclude stretch games, and activity caps further constrain the remaining native pool. Tests deliberately accept short sets in synthetic constrained pools; the email has honest "a few" wording.

    Thus this is a product-rule conflict, not merely a missing length assertion.

    Sources: curate.js:231, curate.js:240, test-playset-curate.mjs:175.

    Remedy: Decide which rule yields when the catalog cannot satisfy all constraints, and simulate several weeks against every supported age/language/sibling combination.

  13. Medium — CSP coverage and the PDF fallback disagree with runtime behavior.

    The packager's emergency fallback fetches public R2 after a proxy 5xx, but the shipped connect-src policy permits only self and Turnstile. The browser will block that fallback.

    Separately, _headers describes itself as covering every Pages response, but Pages Functions must set their own headers. Auth interstitial responses omit CSP/framing headers.

    Sources: package.js:200, public/_headers:27, callback.js:49.

    Remedy: Keep recovery server-side or explicitly support and test the necessary browser policy. Apply shared security headers to Function-generated HTML and test real response types.

  14. Medium maintainability risk — present-tense documentation contradicts executable configuration.

    The architecture runbook says Wrangler is ignored because pages_build_output_dir is absent; it is present. It says feedback workers_dev=true; the configuration says false. Its known-debt list says headers and retention code do not exist; they do. README/checklist counts are stale, and CI comments incorrectly say previews skip every game.

    Sources: ARCHITECTURE.md:417, ARCHITECTURE.md:534, ci.yml:195.

    Remedy: Rebuild the operational reference around executable configuration and dated deployment evidence. Preserve historical notes, but move accepted current decisions into a concise reference. This drift is dangerous for future delegated operations because following the documentation can target the wrong infrastructure.

Additional findings worth tracking:

There is also substantial good engineering to retain: pure shared engines, injectable clocks/transports, real migration SQL in account tests, conditional single-use token consumption, server-side output validation, deterministic answer computation, and hardened Twilio media fetching. The central weakness is testing isolated features more thoroughly than their combined lifecycle.

Open questions the repository cannot answer.

Pushback.

I disagree with the blanket "no ALTER" and near-blanket "no foreign keys" conventions. Retained consent evidence may reasonably outlive a child record; that does not justify leaving unrelated relationships unenforced. Avoiding ordinary migrations has produced several overlapping identity maps and security-sensitive repair paths. A simpler principal/membership/session model would be easier to defend.

I would also change household selection from a mutable property shared across browser tabs to explicit request context. Authentication should establish the person; each operation should identify and authorize its household. Client-side cache invalidation alone cannot make the current model reliable.

Preview should not share production mutation stores. The separate accounts database is useful, but shared UGC, limits, art and telemetry still allow preview activity to affect production. Cost and privacy boundaries deserve actual environmental separation.

"Always acknowledge feedback" needs durable intake behind it. Failing open can preserve usability, but acknowledging while storage is unavailable loses the feedback the product promises to act on. A durable queue or recoverable intake record should precede acknowledgment.

I would revisit the standing Pages/R2 prohibition as a documented incident workaround, not a permanent architecture law. Likewise, public r2.dev is a development endpoint with variable rate limits, not an ideal production PDF origin.

Finally, more passing assertions are not the immediate answer. The highest-value additions are lifecycle tests: removal followed by reinvitation, stale tabs after switching identities, concurrent delivery, failed object deletion, deployed D1 limits, and multi-day curation with the real catalog. Those are the cases this review found ordinary feature tests systematically missed.


Three-way comparison: codex-sol, grok, astra

Migrations_dir P0 — three-way convergence, the strongest signal in the whole exercise. All three independently-run models, on three separate clones, hand-tracing the config, land on the exact same defect: a single shared migrations/ directory means the documented wrangler d1 migrations apply invocations can (and per astra's fixture, do) apply accounts-only SQL to the UGC database and vice versa. Codex called it P0-1 ("hold release"). Grok called it P0 #1. Astra calls it finding #5, "High operational risk," and is the only one of the three to actually prove it by applying all nine files to an empty SQLite DB and showing both schemas land together. Three independent models, three independent runs, same root cause, same file citations (wrangler-accounts-preview.toml, wrangler.toml) — this is as close to ground truth as a fresh-eyes review gets.

Severity framing: astra splits the difference, but leans toward codex on the headline finding. Codex's "hold release" verdict rested on four P0s, top of which was the removed-adult privilege escalation (P0-2). Grok, reviewing the same commit post-merge, independently found the same bug but downgraded it to P1 and reframed it primarily as a lockout, not an escalation. Astra reproduces codex's original framing almost exactly — "removed adult can become the former household's lead," ranked #1, "High," with a fixture proving a 201-response child creation under the wrong household's session. On this specific finding, astra and codex agree with each other over grok. On ARCHITECTURE.md staleness, the reverse happens: grok made it a P0 (its second one), while astra ranks it #14 out of 14, "Medium maintainability risk" — closer to codex's original treatment (docs debt, not a release blocker). So astra doesn't cleanly side with either predecessor on severity taxonomy; it agrees with codex on the privilege-escalation bug's weight and with codex (against grok) on docs debt's weight, while structurally using grok's numbered-severity style rather than codex's single "hold release" verdict.

What astra found that neither codex nor grok flagged:

Where all three converge (the real signal): the shared-migrations/-directory cross-contamination risk, and the removed-adult identity/session model producing unsafe authorization fallbacks (all three independently traced the resolveActor/ensureLeadAdult/session-stamp interaction, even though they disagree on how severe to call it). Everything else is genuine independent-finding diversity — each model's fixture-based reproduction surfaced a different slice of the same underlying pattern: feature-complete, lifecycle-untested. All three explicitly say some version of "individual features are well tested; combinations and removals/transitions are not" — that meta-finding is itself a three-way convergence, arguably more actionable than any single bug.

What none of the three could verify: live deployment/migration-application state, live consent/retention backlog, and provider-account training/retention settings. All three explicitly flag this as outside read-only-clone scope. Codex and astra state it plainly; grok stated it too but then contradicted itself by asserting specific workers were "undeployed" — an error this file's sibling note corrects. Astra avoids that error entirely: its runtime table correctly labels both sw-daily-playset and sw-retention-sweeper as "Independently deployed," and its open questions about them are scoped to operational status (last successful run, backlog), not deployment existence.