HAE pipeline outage — root cause: D1 database full (diagnosed 2026-07-19)
Symptom vs reality
- Morning briefs' health section went dark; D1
daily_metricsfreshest date 2026-06-15,raw_ingest_logfreshest 2026-05-21. - Founder's HAE app showed successful syncs (screenshot 2026-07-19 9:15am: Workouts + Symptoms ✅ Jul 19 9:07 AM; Health Metrics last complete Jul 11 8:47 PM).
- Both are true. Transport was never broken.
Root cause (verified, not inferred)
wrangler d1 execute rdco-health write probe returned the authoritative error:
Exceeded maximum DB size [code: 7500] — the DB sits at ~1.07GB.
The May 21 initial HAE dump imported the founder's entire 2020-2026 Apple Health
archive (~4.7M per-sample rows: 1.39M step_count, 819k basal_energy, 910k
walking_distance, …), filling the database. Since then every D1 write fails; the
Worker (rdco-health-export, last deployed 2026-05-21T19:41Z, bindings verified
correct, Build-2 code confirmed in the deployed bundle) swallows the error,
still archives raw JSON to R2 rdco-health-raw, and returns 200 — so HAE
truthfully reports success.
Zero data loss: R2 has every payload, daily, through 2026-07-19 (verified via
R2 object listing; e.g. 2026-07-19/962A591B….json at 02:14Z).
Fix (staged, awaiting founder — classifier hard-blocks Ray on D1 writes)
One command, at the Mac or tmux:
bash ~/Projects/rdco-health-export-worker/scripts/recover-d1-space-2026-07-19.sh
Idempotent, ~15-30 min: (1) prunes pre-2026 daily_metrics rows year-by-year
(recoverable from R2), (2) write-probes, (3) replays the missed window
2026-05-22 → today per-day via MODE=force PREFIX=<day> backfill-r2-to-d1.ts
(PREFIX support added 2026-07-19; per-day scoping deliberately skips the May-21
archive-dump objects so the DB doesn't refill), (4) prints freshest-row verdict.
If deletes also return code 7500 → plan B: fresh D1 database + rebind + redeploy + scoped backfill (R2 remains source of truth).
Alternative founder option offered: permissions allow-rule for
npx wrangler d1 execute rdco-health so Ray owns DB maintenance directly.
Standing guidance until recovery runs
- Morning-prep keeps omitting the health section (never invent numbers).
- Do not tell the founder "HAE isn't syncing" — it is; the break is D1 capacity.
- Watch item: after recovery, consider a retention policy (e.g. rolling 18-month per-sample window) so live accumulation never refills the cap.
Related: [[IMPLEMENTATION-NOTES-2026-05-21-health-export-worker]] · [[2026-05-22-execution-system-v1]]