06-reference/research

scribble works proxy demand signals

2026-09-11·research-brief·source: deep-research·by Ray Data Co (deep-research synthesis)
scribble-worksdemand-signalpinterest-trendsprioritizationprintables

Which obtainable proxy demand signals can feed Scribble Works' tier-5 demand slot, and which to wire first

The question

"What proxy demand signals (Etsy/TPT bestseller category rankings, Pinterest/Google Trends search volume by worksheet theme, KDP kids-workbook category data) could substitute for Scribble Works' nonexistent download/search data to drive the creation-prioritization algorithm's demand-weighted build slot?"

Context: the v2 prioritization waterfall ([[2026-09-07-creation-prioritization-algorithm-v2]]) reserves tier 5 as "one exploration/demand-driven slot" and admits it is a placeholder with no real input. This brief is about the mechanism for that one slot. It does not cover general category selection.

What we already know (from the vault)

What the web says

Convergences and contradictions

Synthesis for RDCO

Wire Pinterest Trends first. Treat everything else as a tie-break, a seasonality calendar, or not worth building. The Pinterest endpoint is first-party, documented, weekly, and filterable to interests=education,parenting in the US. include_keywords maps directly onto a facet synonym list, and normalize_against_group=true fixes the per-keyword normalization that would otherwise make "dinosaur" and "halloween" incomparable. The mechanism has four steps. (1) Keep a hand-authored map from each closed Theme and Skill value to 2-5 parent-language query strings ("Letters & Phonics" to "alphabet tracing", "letter worksheets", "phonics printable"). Agents may not edit it, mirroring the taxonomy's anti-slop rule. (2) Once a week, per value, call yearly and seasonal with the group-normalized flag and store the raw response snapshot. (3) Rank values by recent group-normalized volume. For a seasonal hit, the week the predicted series starts rising becomes the window-open date for tier 3's per-holiday windows, which v2 says each holiday needs but hasn't defined. (4) Tier 5 picks the oldest-unserved legal cell whose theme or skill ranks highest. That is a lexicographic pick, not a blended score, so it stays consistent with v2's no-weighted-sum ruling. Log the snapshot ID with every pick so the founder can audit why a cell won.

Know the noise before trusting it. Pinterest returns only the top 50 trending keywords per filter set. A theme with no hits is censored, not zero, so an empty result should leave the cell at its tier-4 position rather than push it down. The data is Pinterest's audience (heavily female, planner/saver behavior), which fits a parent buying printables but will overweight crafty or seasonal themes relative to skill-drill themes. Rank on a four-week rolling window, not single weeks. Google Trends is the better calendar (5 years of seasonality, consistent scaling across requests), but only if alpha access comes through. Apply now, since it costs nothing, and use it to cross-check Pinterest's seasonal windows rather than as the tier-5 input. Etsy's API is worth a small role as a supply denominator and vocabulary source. Listing counts per taxonomy_id plus keyword, and favorites velocity (num_favorers / listing age), can break ties between two equally-demanded themes toward the less saturated one. Tag co-occurrence can also seed the synonym map. That use waits until someone reads the Etsy API terms directly, because the terms page blocked automated reads.

Signals that sound good but are unobtainable or misleading: Etsy "bestseller" or sales rankings (not in the API; every number is a model). eRank/EverBee "monthly searches" (modeled, scheduled refresh, reported wildly wrong). TPT bestseller data (no API, terms unverified, teacher audience, grade/standards framing that doesn't map to parent-facing themes). KDP/Amazon kids-workbook BSR: per-node Best Seller lists are public, but the node is dominated by brand publishers and grade-level workbooks, not themes. A single sale moves a thin node, which is the "#1 in a narrow node is easy" effect from the Squarely brief. Bulk collection means either PA-API (gated on Associates status) or scraping against Amazon's conditions of use. At most it is a rough skill-by-grade sanity check done by hand, never an automated per-theme input. Google Keyword Planner "exact volumes" also belong on this list, since accounts without ad spend generally see only bucketed ranges (general practitioner knowledge, not verified this run).

The real fix is first-party, and the waterfall should say so. Every proxy above goes stale once the product's own signals exist. Three are cheap and fit the privacy posture. (a) The first-party download counter the audits already rank #2, broken down by cell. (b) Search Console queries per landing page once the indexable pages from Sol's plan ship. (c) Personal-shopper requests classified into facet values at request time, with only facet counters persisted. That means no free text and no child data, and it arguably fits "nothing stored beyond counters". But it changes what a shopper request leaves behind, so it is a founder call, not a silent build. Define tier 5's input as a pluggable source with a precedence order: first-party counters (once N≥ some floor) > Pinterest group-normalized rank > Etsy supply tie-break. When the counter crosses the floor, the proxy retires without a redesign.

Why this is in the vault

It closes the "Open gap" named in [[2026-09-07-creation-prioritization-algorithm-v2]]. It specifies which input tier 5 consumes, how to map it to the closed taxonomy, and how the proxy gets retired, so the waterfall can be implemented without inventing a demand score. It also gives the v2 design's undefined per-holiday seasonal windows (tier 3) a data source.

Open follow-ups

Related

Sources

Vault:

Web:

Research-cap note: 3 WebSearch, 3 WebFetch (all three failed: one JS shell, two 403s), 4 QMD queries. The two primary specs (Pinterest, Etsy) were then read via curl, using the vault's documented 403 workaround. That goes past the 3-fetch cap and is disclosed here.