01-projects/printables-product

creation prioritization algorithm v2

Scribble Works creation prioritization — v2 (post-critique revision)

v1 was a weighted-sum scoring system. Both Sol and Grok, working independently, converged on the same verdict: scrap the weighted sum, use hard constraints plus a deterministic priority waterfall. Both also independently caught the identical bug (pending-review games counted as "coverage," masking real gaps after a rejection) and the identical inconsistency (the coverage matrix is 3D — age × activity × theme — but v1's scoring silently collapsed to 2D). That level of independent agreement is the signal to actually change the design, not tune the numbers.

What v1 got wrong (credited to the critiques, not softened)

  1. The weights were product decisions laundered as arithmetic. 40/30/20/-15 quietly encoded "catalog completeness beats the founder telling us to stop" and "an empty cell nobody wants beats a live seasonal deadline." Neither tradeoff was defended in v1; both are wrong on their face for a business whose demand is calendar-driven search.
  2. No demand signal at all. v1 optimized for an evenly-filled grid, not for what a parent actually searches for or buys. Grok's framing: the objective function was catalog completeness, not commercial value. That's the single biggest miss.
  3. Pending-review counted as coverage. Four rejected-but-still-pending games make a lane look "full" and suppress further work on it — then the Bads land, the lane is actually empty, and the algorithm doesn't notice because nothing re-triggered it.
  4. The feedback loop contradicted itself. "Boost remediation" and "down-weight repeated failures" pull opposite directions — the design could learn to avoid a weak lane instead of fixing it.
  5. The diversity penalty couldn't diversify. -15 inside one day's batch is too weak to stop four similar games from winning anyway, and there's no cross-day penalty at all.
  6. The backlog throttle was a blunt, unproven cliff (10 → halve) with no basis in actual review throughput, and it throttled build volume rather than the actual constraint (review capacity / stale queue items), which can starve time-boxed seasonal work while the queue is full of unrelated stale items.

v2 design: constraints first, then a waterfall, no weighted sum

1. Legal cells, not all cells

Maintain an explicit table of which (age_band × activity) combinations are legal targets at all. Thinness inside an illegal cell is not a signal — a cell can be empty because it should be. This replaces "inversely proportional to coverage" with "only rank inside cells that are allowed to exist."

2. Waterfall, evaluated in this fixed order, first match wins

  1. Explicit founder ask (use case #2) — always wins, doesn't go through this system at all.
  2. Open remediation — a named Iterate/Bad pattern with no shipped fix yet. This is a stop-and-fix, not a competing score: while a lane has an open remediation, no new candidate in that same lane is proposed until the fix ships (or the founder explicitly says otherwise).
  3. Seasonal quota — only inside that holiday's own defined window (each holiday gets its own window, not one universal 45/21-7 curve — Grok's point that Halloween's planning horizon and Thanksgiving's are not the same length is correct), and only while live count for that theme is below a target K. This is a quota to fill, not a point bonus — while inside-window-and-under-K, seasonal candidates are the only candidates proposed.
  4. Oldest-unserved legal cell, round-robin by age band so no single band monopolizes build capacity. "Oldest-unserved" means measured against live games only.
  5. One exploration/demand-driven slot, reserved for whatever real signal exists (see Open gap below) — deliberately last, not competing with 1-4.

3. Coverage counts LIVE games only

Pending-review items are tracked separately and never suppress further building in a cell. A cell isn't "covered" until something in it actually shipped.

4. Backlog handling is a gate, not a volume knob

If the review queue is deeper than a threshold grounded in actual founder review throughput (not a guessed constant — this needs a real number from his review cadence, which we don't have yet), skip ordinary waterfall tiers 4-5 for the day; tiers 1-3 (explicit asks, remediation, seasonal deadline) still run, because those are the time-boxed or founder-driven items that shouldn't wait on an arbitrary queue-depth cliff. This targets the actual constraint (his attention) instead of uniformly throttling volume.

5. Feedback loop, scoped correctly this time

Track pass/iterate/bad outcomes per cell to inform tier-4 ordering (an oldest-unserved cell that keeps failing review might need a different activity type, not just another attempt) — but this never overrides tier 1-3, and it is explicitly logged as "this cell has a pattern" for a human to look at, not as a silent down-weight that makes the system quietly avoid a weak spot forever.

Open gap, named plainly, not solved here

There is no demand signal in this design, and both critiques are right that this is the real risk: without downloads, search terms, or founder-declared strategic importance as an input, the waterfall will happily fill legal-but-commercially-dead cells just because they're the oldest-unserved. Fixing this needs real data (what gets downloaded, what parents search, what converts) that doesn't exist yet. Until it does, tier 5's "exploration/demand-driven slot" is a placeholder, not a real mechanism — flagging that honestly rather than pretending a made-up demand score would help.

What this does NOT do (unchanged from v1's own scope boundary)