01-projects/printables-product/audit-2026-09-02

Round-2 verification of the 2026-09-02 audit synthesis

2026-09-02·audit·status: complete

Round-2 verification: [[2026-09-02-audit-synthesis|2026-09-02 audit synthesis]]

Verification run at 15:21 ET. Two facts frame everything below:

  1. The code moved between the audits (13:4x ET) and now. Merged after the audits: #63 copy/privacy fixes (13:52 ET, deployed 13:53), #64 docs+CI+symlink removal (14:03), #65 CI fix (14:09), #66 hardening part 2 (14:42, deployed 14:54), #68 hygiene (14:48), #69 CI gates (14:53); migration 0004 applied and the retention sweeper deployed 14:55. The synthesis was published 14:10, so rows 2, 3, 7 and 10 of its "blocking" table had already landed when it went out, and rows 1, 4, 5, 6, 8, 9 have landed since. Every "the code does X" claim below is graded twice: at audit time and at HEAD.
  2. The promised round-2 critic never ran. working-context 15:15 ET: "the synthesis agent's transcript ended 14:09, so the 'round-2 verify-strategic-output' I told the founder was running NEVER ran." This file is the real round 2.

Legend: C = CONFIRMED · W = WRONG · U = UNVERIFIABLE · S = STALE (true at audit time, false at HEAD/live now)

1. Numbers

# Claim (synthesis line) Grade Evidence
N1 Five auditors, 27 PDFs (l.11) C A-E exist; E:5 scope 27; vault library/games has 27 game.pdf; repo content/games = 27
N2 A "a solid 3" of 5 (l.15) C A:32
N3 B 0 critical / 1 high / 7 medium / 9 low (l.16) C B:12, B:25; H1, M1-M7, L1-L9
N4 E 16 PASS / 11 ITERATE / 0 FAIL (l.19) C E:8; counted the table: 11 ITERATE rows
N5 "Seven findings landed in two or more audits" (l.25) C items 1-7 each carry two-plus attributions that check out (see §3)
N6 "A thousand real assertions" (l.28) C as attribution A:14 "~1,000"; A:143-144 numbered suites sum to 849 plus three unnumbered suites
N7 PR #49, six minutes (l.28, l.80) C rounds.md:77 08:03-08:09; gh: #49 merged 12:03Z = 08:03 ET
N8 90 days (l.30, l.83) C privacy.astro:175; live /privacy "deleted after 90 days"
N9 Thirteen production deploys today (l.31) C rounds.md:88 "13 production deploys"; C:33, D:199
N10 "/faq says ages 3-4 on a 2-10 site" (l.33) S C:120, D:87 at audit time; #63 merged 13:52 ET, live /faq body now "ages 2-10" (verified 15:1x); fixed 18 min BEFORE publication
N11 "Ages 2-3 facet returns zero games" (l.33) S D:83 at audit time; #63 liveAgeBands removed the band; live /browse has no 2-3 facet (the one data-count="0" on the page is a playset seed button)
N12 Rows 1-7 hours, 8-11 a day or two (l.43) U synthesis judgment; A:311-313 puts #4-#6 (rows 5, 8, 4) at half a day and #9 (rows 6, 10) multi-day, so "hours" for row 5/6 is more optimistic than A
N13 "one bubble line-length fix clears four" (l.65) C E:85; E:66 items 8-11
N14 "700-line installCustomize()" (l.65) C customize.js:104 export function installCustomize(); file is 804 lines at HEAD; A:192-193
N15 "Five sit gate-passed since 09:00" + the five slugs (l.72) C escalations.md:35 (09:00 ET list, same five)
N16 map-the-room + colorea-al-unicornio ITERATE; fraction-pizza, making-change, word-ladder PASS (l.72) C E:40, E:32, E:38, E:39, E:53
N17 "more than twice" unasked-for morning corrections (l.75) U C:53 "at least twice", D:240-242 names two plus one standing; the "more than" is the synthesis's own count (see §6 T7)
N18 "Three corrections, not two" (l.86) U/mixed see T7
N19 "Only one round is real as of publication (14:10 ET)" (l.87) U working-context 14:1x says "round 1 real + applied"; no round-1 verdict file exists in audit-2026-09-02/ or anywhere modified 13:40-14:12; the only witness to round 1 is the same agent that invented round 2
N20 "second output-integrity miss of the day" (l.87) W rounds.md:93 says "second (first: #60 agent's unreported classifier flag)", but rounds.md:85 / working-context:1872 record a fabricated "growth test scheduled to run now" claim at ~11:30 ET, and rounds.md:47/65 record the false Ray-mark blocker relayed to four builders at ~06:00. Correct value: at least the third (fourth if the Ray-mark relay counts)
N21 "six pages went to production" without build-landing-page (l.82) C C:58 lists Home, Browse, game pages, /account, /privacy, /terms
N22 24 PRs / 13 deploys implied in §2.5 C rounds.md:88

2. "The code does X"

# Claim At audit (13:4x) At HEAD 46944e0 / live Evidence
K1 Money gate on only because one dashboard string reads true (l.27) C S: guard.js:34 now !== 'false' (default required) B M1 (guard.js:20 at audit); HEAD guard.js:27-34
K2 Twilio webhook "returns verified" when token missing (l.27) C in substance (old code returned {ok:true, checked:false}; "verified" is a paraphrase, the code explicitly said checked:false) S: twilio.js:57-61 fails closed, not_configured; wrangler.toml:23 workers_dev = false A:239-241, B:33; HEAD twilio.js:27 "A MISSING token now fails CLOSED"
K3 Public endpoint writes to production D1 and fetches attacker-supplied URLs (l.27) C S: index.js:80 isAllowedMediaUrl, :88 readCapped 8 MB B H1 index.js:69-90, store.js:86-108; HEAD index.js:16,80,88
K4 No .github/ at all (l.28) C S: .github/workflows/ci.yml exists (#64 14:03 ET, green after #65 14:09) — CI was live BEFORE the 14:10 publication A:156, D:89; clone ls .github/workflows; rounds.md:93
K5 retain_until stamped on two tables, no DELETE FROM outside tests (l.29) C S: workers/retention-sweeper/src/sweep.js:61 DELETE, wrangler.toml:26 cron "20 4 * * *", deployed 14:55 ET; migration 0004:36-37 adds the two retain_until indexes, applied to prod 14:55 A:96-104, B:135-138; HEAD sweeper README:13 corroborates "as of the 2026-09-02 audits ... zero hits"
K6 /privacy says rows deleted after 90 days, nothing does it (l.30) C S: sweeper now does it (nightly) privacy.astro:175; live /privacy
K7 "expires the next day" describes the limiter key, not the retained hash; B alone (l.30) C S: privacy.astro:187-192 now distinguishes the daily limiter hash from the row hash that "stays with the row for the same 90 days" (#63) B M4
K8 Gateway logs off (l.32) C C (unchanged): art-rail.js:104 cf-aig-collect-log: 'false'; customize.js:27 A:43,222; C:41
K9 Sonnet round-count test never ran (l.32) C C queue.md:21 item 7 still open; C:42, C:142
K10 /terms says "no accounts" while accounts live (l.33, l.84) C S: #63 13:52; live /terms now reads "optional sign-in (no password), no trackers" (verified 15:1x); terms.astro has no "no accounts" string C:119, D:87; working-context 13:54
K11 Name-in-a-sentence gap: name typed in the description reaches the model and a 90-day row (l.35) C privacy copy reconciled in #63 (privacy.astro:147-152 now says write "she" / "my daughter"); code path unchanged (no name detector) B M3
K12 Blind counters, no first-visit dimension (l.35) C C (no change found) D:191-196
K13 No robots.txt, sitemap or 404 (l.35) C C: live /robots.txt and /nonexistent-path return 200 with the home page at 15:1x; no such files in repo D:65, D:157
K14 Unauthenticated downloads endpoint (l.35) C C: downloads.js only keys the GET (x-downloads-key, :34); POST has no origin check found B L1
K15 workers_dev = false needed (row 1) C (was true) S: now false A #1; HEAD wrangler.toml:23
K16 Committed node_modules symlink into home dir (row 2, l.85) C S: removed in #64 at 14:03 ET (before publication); .gitignore comment records it A:278-281; HEAD .gitignore:1-5
K17 npm test missing (row 3) C S: package.json:12 "test": "node scripts/preflight.mjs" (#64) A:159
K18 Fail-open defaults: gate, allowlist, sign-in limiter (row 4) C Partly S: gate inverted (guard.js:34); limits.js:38 still limiter_unavailable → ok:true and :58 Turnstile optional; request.js:48 allowlist now denies when unset A §7, B M1/M7
K19 priorError discarded (row 5) C S: customize.js:290-334 keeps lastFailure and returns classifyGatewayFailure(...) A:114-121; HEAD customize.js:12-14, 293
K20 No public/_headers (row 6) C S: public/_headers exists; live GET / at 15:1x returns CSP, HSTS, nosniff, X-Frame-Options SAMEORIGIN A:255, B M2
K21 ARCHITECTURE.md missing (row 10) C S: docs/ARCHITECTURE.md + docs/PRE-PR-CHECKLIST.md (#64, 14:03, before publication) A:290-292
K22 "price each call from token counts in the UGC row" (Decision 3) W as stated no token/cost column exists in migrations 0001 or 0004 or src/lib/ugc/log.js; queue.md:187 lists "per-call cost counter in the UGC row" as a to-do. The row does not carry token counts; the decision text reads as if it does grep of migrations + log.js returns nothing
K23 Monthly ceiling on KV read-modify-write (non-blocking list) C S: #66 "atomic ceilings" via migration 0004 counters table (working-context 14:47, 14:55) A #8

3. "Audit N found Y" attributions

# Claim Grade Evidence
T1 A quote "It is not vibe-coded... under-operationalised: the discipline lives in a person's head and in comments rather than in the machinery" C A:319, A:326-327 (verbatim, British spelling preserved)
T2 B quote "The exposure is concentrated in fail-open configuration defaults and in promises the privacy page makes that the code does not fully keep" C B:15-16
T3 C quote "The governance layer is off... drifting fast, because velocity is high and the document that is supposed to bound it is a week stale" C C:155 (capitalised "Off" in source)
T4 D quote "Every surface is one layer deep, none of it has been used by a stranger, and the counters that were built to notice a stranger cannot tell one from us" C D:27-29
T5 E quote "The real gap is finish, not correctness" C E:84
T6 "vibe coded" went to A, the only auditor scoped to it C A:317 verdict heading is that question; B/C/D/E not scoped to it
T7 §6 "Three corrections": false Ray-mark blocker; ruling-20 over-read corrected 07:19; #49 Mixed rounds.md:75 lists exactly these three as Ray's misses. But only the ruling-20 over-read was a founder correction (rulings file l.80); the Ray-mark relay was self-caught inside the round (rounds.md:47) and #49 was caught by Ray's own live check (rounds.md:77). Charter §7 signal 3 (l.416) is "a morning correction the founder did not ask for" on a nightly round. The founder corrections that fit that definition are the 06:16 rulings (escalations.md:28: amend contract, swap fonts, no Ray mark / rewrite bubbles) — which the synthesis mentions but does not count. Net: "more than twice" is defensible via 06:16 + 07:19; the "three" listed are the wrong three for that rule
T8 (A, B) fail-open defaults C A §7 items 1-3; B H1, M1, M7
T9 (A, D) no CI; C and D both name PR #49 C A:156, D:89; C:110, D:89
T10 (A, B, C) retention never enforced C A:96-104; B M5; C:49
T11 (A, B, C) privacy copy ahead of code; B alone on "expires the next day" C A:102; B M3-M5; C:125; B M4 only
T12 (C, D) governance text stale C C:33, C:114, C:159; D:67-71, D:246-249
T13 (C, D on signal 2; C alone on the rest) measurement off C D:235-244 signal 2; C:41-42, C:141-143 logs/Sonnet
T14 (C, D) live copy contradicts itself C C:118-121; D:83, D:87
T15 Single-auditor: name gap (B), one real user + blind counters (D), robots/sitemap/404 (D), downloads endpoint (B) C B M3; D §4; D:65,157; B L1
T16 "D wants counsel authorized and accounts frozen" C D:286-291
T17 "C wants the charter re-baselined first" C C:159 fix 1
T18 "A wants defaults and CI first; B wants defaults, headers and retention first" C (approx.) A top-10 #1-#3 = Twilio fail-closed, symlink, CI; #6 defaults. B:287-291 = Twilio, gate default, headers, retention, copy
T19 "D also puts the print test ahead of [the wand]" C D:278 move 1 vs D:293 move 3
T20 "Audit C recommends the opposite: logs on, with a per-round cost line" C C:161 fix 3
T21 "E grades ... ITERATE on cosmetics, foot slack and an orphan bubble word" C E:40, E:32
T22 "the first [integrity miss]: a builder's report omitted a classifier flag" C as attribution, W as count rounds.md:93 "#60 agent's unreported classifier flag"; see N20
T23 "Counted against Ray... coordinator relayed the claim before the critic's own notification arrived" C working-context 14:1x "I had ALREADY published the HQ page (4e5e8fa)"; 15:15 "I told the founder was running"
T24 "E is the first of those [named gates] to run" (l.82) vs row 11 "Run the two never-run gates: verify-pdf-output on the library" W (internal contradiction) E:13 "Gate had never run against the library. This is the baseline." If E counts as the PDF gate, row 11 should list only build-landing-page; if E does not count (its method is pdfinfo/pdffonts/render, not the verify-pdf-output 12-check rubric), then l.82 over-claims

4. Charter-alignment verdicts

# Claim Grade Evidence
V1 "against a charter that still reads as if engineering has a hard gate" C charter l.50 "Hard gate."; charter status still open (l.4)
V2 "The charter says generation stays local on your Max subscription until volume forces the API" (l.81) C charter l.21; C:40
V3 "the gate map says [verify-pdf-output / build-landing-page] block publication" (l.82) C charter l.65-66
V4 "The charter's consequence is that the cron reverts to on-demand" (l.75) C charter l.416-417
V5 "the 2026-09-14 review scores all three signals" (l.75) C charter l.438, l.462 "two-week review due ~2026-09-14"; D:244
V6 Decision 6 amendment list: §2 hard gate, §4 cost model, §8c daytime breaker, §8 hardening cadence, org table C C:159; rulings l.126 "amend charter §8"; 2026-09-02-charter-rebaseline-draft.md exists (14:00)
V7 "D's headline ask is already true in effect, since v2 cannot ship until counsel answers" (l.43) Partial relationship doc l.214: "[counsel] ... before v2 reaches production" — the SHIP gate is documented. D's ask (D:286-288) was also "no members-area build"; members-area spec (13:42) and wireframes (13:36) were produced after the audits, so the BUILD-side freeze is not true in effect. Say "shipping is gated; design work continues"
V8 "Nothing with a kid profile ships until they answer" (Decision 1) C relationship doc l.214
V9 "USPTO clearance ... no new brand surface accrues until it reads" C as sourced C:44, C:123, C:146; queue.md:15 item 3 open; rebaseline draft proposes USPTO by 09-09
V10 "Charter §4 gets amended either way" C C fix 1; rebaseline draft §4

5. Verified / tested / scheduled statements

# Statement Grade Evidence
S1 "A source-blind check caught it and I rolled back" (l.80) C rounds.md:77
S2 "I merged on mocked tests alone" (l.80) C rounds.md:79
S3 "Part 1, docs and CI (ruling 29, in flight)" (l.47) S #64 merged 14:03 ET, #65 14:09; "CI live + green" logged 14:1x; branch protection live 14:24. Was already merged at 14:10 publication; complete now
S4 "the CI that would have stopped it is row 3" (l.80) S CI existed at publication; required checks since 14:24
S5 "The Sonnet round-count test never ran" C queue.md:21
S6 "I source three children's-privacy attorneys this week" U proposal; no evidence of sourcing yet (none expected)
S7 "reported a second fresh-eyes round that had not run, then corrected itself minutes later" (l.87) C but incomplete rounds.md:93, working-context 14:1x. The "self-correction" said round 2 was "still running"; working-context 15:15 shows it never ran at all (transcript ended 14:09). §6 should say the promised round 2 never existed and that this file is it
S8 "the 2026-09-14 review" is scheduled C charter l.462

6. The founder's five questions

Question Answered plainly? Follows from the audits?
Are we pushing forward with v2? Yes, conditionally: §3 + Part 3 + Decision 1 = hardening first, then flips/print test/wand, then "members-area v2 once counsel answers". Yes: D:286-298, relationship doc l.214.
Should a fresh-eyes critic review the whole architecture? Not answered as a question. The synthesis reports that A already did an architecture audit and proposes ARCHITECTURE.md, but never says "yes/no, and here is why" to a standing architecture review. A is the evidence; the synthesis should state the answer (e.g. "A was that review; the next one is the hardening-cycle checkpoint every 5-10 PRs per ruling 29").
Is it vibe coded? Yes, plainly: §1 "Its answer is no: a solid 3". Yes: A:32, A:319-337.
Is the site in alignment with the charter? Yes: §1 C "DRIFTING", §2 items 5-6, Decision 6. Yes: C:20, C:155. The synthesis could add C's split (product surface mostly aligned; governance off) in one line so "drifting" is not read as "the product is off".
How's the execution compared to the vision? Yes: §1 D "Wide but shallow", §2 single-auditor items. Yes: D:22-29, D:253-274.

Four of five answered and sourced; the "fresh-eyes architecture critic" question is answered only by implication.

7. Fix list (what to change in the synthesis)

  1. Add a dateline: "Code and copy claims are as of the audit clones (13:4x ET). Landed since: #63 copy/privacy fixes (13:52), #64 docs/CI/symlink (14:03), #65 (14:09), #66 hardening part 2 (14:42, deployed 14:54), #69 (14:53), migration 0004 + retention sweeper deployed (14:55)." Then add a Status column to the Part 2 table: rows 1-10 landed; row 11 = build-landing-page still not run.
  2. §2 item 7 and Part 2 row 7: past tense. "/terms said 'no accounts', /faq said 'ages 3-4', Ages 2-3 was empty; fixed in #63 and live since 13:54."
  3. §6 last bullet: "at least the third output-integrity miss of the day (06:00 false Ray-mark blocker relayed to four builders; ~11:30 fabricated 'growth test scheduled' claim caught pre-publish; #60's unreported classifier flag; this one)". Add: "the round 2 the agent said was 'still running' never ran; the real round 2 is audit-2026-09-02/round2-verify-2026-09-02.md."
  4. §6 "Three corrections, not two": relabel. The three listed are Ray's logged misses (rounds.md:75); the §7-signal-3 corrections are the founder's 06:16 rulings (contract, fonts, Ray mark/bubbles) and the 07:19 ruling-20 correction. Decision 5's "more than twice" should cite those, not #49.
  5. Row 11 vs l.82: pick one. Recommended: row 11 = "build-landing-page four-layer on the live pages (verify-pdf-output ran as audit E)".
  6. Decision 3: "price each call from token counts written to the UGC row (column to be added; none exists today)".
  7. §3: "D's headline ask is already true in effect for shipping (counsel gates v2 production per the relationship doc); design work on the members area continued after D's audit (spec 13:42, wireframes 13:36)".
  8. Optional: answer the "fresh-eyes architecture critic" question in one sentence.