The ranked Portfolio Scoring scorecard — and why the ranking inverts the moment you add the one criterion the founder already called non-negotiable
The question
Verbatim: "Produce the ranked, weighted Portfolio Scoring scorecard (Squarely / MAC / Sanity Check / CAF) with shipped weights — the open checkbox in the Productize framework doc, now calibratable against named case benchmarks."
Filed as a follow-up off [[2026-06-28-productized-consulting-scalable-anchor-transition]], which named the checkbox but did not fill it. The question is stale in two ways and this brief answers the corrected version, not the filed one (both corrections are handled explicitly in the next section).
What we already know (from the vault)
- The shipped weights are real and I read them off the primary artifact, not the vault summary. Vecteris tool #4 (
Product Portfolio Scoring Template.pptx, watermarked internal-use, stored locally 2026-06-25): Benefits — Strategic Alignment ×50, Customer Need/Value ×30, Distinctiveness ×25, ROI ×20; Costs — Ease of Sales ×25, Operational Costs ×20; Risks — Distance from current organizational capabilities ×30. ([[2026-06-25-productize-framework-armstrong-vecteris]] has the weights; the polarity correction below is new.) - Correction to the vault note. [[2026-06-25-productize-framework-armstrong-vecteris]] records the method as "score 0-5, apply weights, sum." The template itself says something different: "1 is low and 5 is high for both Benefit and Cost… an item with 5's on the Benefit side and 1's on the Cost side will calculate the maximum score." The Cost and Risk columns are inverted and subtracted (5 = Hard to sell / High operational cost / Many capability gaps). Every score in this brief uses the template's polarity. Max attainable = 550, worst = −250, range 800. That range matters later, because it is how you tell a real gap from noise.
- "CAF" was retired 2026-08-10. The framing is now Organizational Intelligence (OI): Organizational Platform → Org Map, Intelligence Platform → Intelligence Maturity Assessment, plus the OIP process and Pulse. Old CAF was taken over internally and rebranded "Forge Blueprint." The question's fourth bet is scored below as OI, not CAF.
- The bet portfolio changed after the question was filed. [[2026-08-31-studio-charter]] stood up Scribble Works (kids' educational printables engine, scribbleworks.co, live at scribble-works.pages.dev) on 2026-08-31, with a four-agent studio, critic gates, a nightly cron, and $0.710/pack marginal cost. It did not exist on 2026-06-28 and it is now the most agent-native bet in the stack.
- The binding constraint is demand generation, not capital. Established 2026-07-31 and re-confirmed by [[2026-09-08-37signals-services-to-product-cashflow-bridge]]: the demand assets are dormant, and on the 37signals timeline "RDCO is at 2001-2002, not 2004." That brief also reframes MAC and Sanity Check as the gap-filler tier (37signals' workshops-and-$99-report layer), not as candidate anchor products — which is a direct instruction about how to score them.
- The move-1 screen already ran on OI. [[2026-09-06-move1-anchor-screen-oi-skill-catalog]] scored the 106-skill catalog on Urgent / Expensive / Recognized / Standardizable and found exactly five clusters clear the bar, all of them verification offers rather than assessment offers, and roughly 95 of 106 skills answer a question the consultant has rather than one the client has. OI's Customer Need score below is entirely contingent on that isolation actually happening.
- The four-walls spec is the founder's own instrument ([[user-career-commitment-shape]]): feelable mission / whole lane / build-then-TEACH / real upside stake. Wall #4 is the wall that ruled CAF out as a mission — his words, "make lots of sales for an entity I don't have any upside stake in."
What the web says
- The shipped weights are the vendor's default and the vendor says to change them. The template's own instruction: "Adjust the weight of each column to reflect your priorities. Weighting is flexible and can add up to any number you choose. Add, edit, and remove columns to suit your needs." Running the shipped weights unmodified is a choice, not obedience.
- The known failure mode of weighted scoring is criterion selection, not weight calibration. The product-management literature is blunt: the criteria you pick dictate the answer, and weighted scoring offers no guidance on picking criteria (Product School; airfocus). This is the single most load-bearing external finding in the brief and it is what produces Variant C below.
- Digital printables benchmark (calibrates Scribble Works). TPT paid out ~$253M to sellers in 2024; the distribution is extremely top-heavy — most active sellers average roughly $27/month, the top 1% roughly $6,300/month, and 300+ sellers have cleared $1M lifetime. Full-time earners "almost all started before 2020" (Medium/Write & Rise; SEOLumina). Confidence: MEDIUM. These are SEO-content aggregations of a platform payout figure, not audited data; use the shape of the distribution, not the decimals.
- Paid-newsletter benchmark, primary source (calibrates Sanity Check). beehiiv's State of Paid Newsletters 2026, thousands of paid publications across 2021-2026: median free-to-paid conversion 0.62% (~6 paying per 1,000 free), top quartile 2-5%, Technology top-decile 5.55%; median price $10/mo and $100/yr, stable since early 2024; monthly churn 5.06% (Food & Drink) to 16.67% (Money); median subscriber LTV $83-230 (beehiiv). Confidence: HIGH on the numbers as beehiiv measured them, MEDIUM on transfer — it is one platform's book and excludes Substack and Ghost.
- What that means arithmetically for Sanity Check. At beehiiv's median conversion and median price, $1,000/mo of subscription revenue needs roughly 16,000 free subscribers. At the Technology top-decile rate (5.55%) it needs roughly 1,800. Sanity Check currently has 21 archived issues, a finished unpublished relaunch essay, and zero paid subscribers. Both numbers are distant; the gap between them is the entire value of being good at this.
- Physical-goods price anchors already in the vault ([[2026-09-05-kids-subscription-box-comparables-scribble-works]]): Lovevery ~$40/mo equivalent, KiwiCo ~$24/mo, Osmo $79 one-time hardware with no subscription. The brief's own correction is the important part — those are physical-goods prices, and the real anchor for a digital kids' PDF is "$3-8 one-time and a large free tier on Pinterest."
- Productized-consulting outcome benchmarks carried forward from [[2026-06-28-productized-consulting-scalable-anchor-transition]]: Pointerpro ~$287K MRR from a tool built for one client; Jeff Sauer's $1,997/yr expert membership; Mike Gammarino's $2,500 audit converting to $25k+ implementations; Jane Portman's three fixed-scope offers; ConvertKit $5k bootstrap → ~$29M ARR; Basecamp funded entirely by agency cashflow. Timeline across all of them: 12-24 months to traction, 3-5 years to full benefit.
Convergences and contradictions
- Convergence: every external benchmark prices distribution, not product. TPT's top-heavy curve, beehiiv's 0.62% median, and the 12-24 month productization timeline all say the same thing the vault says independently — the scarce input is an audience, not an artifact. Armstrong's rubric encodes this as a single ×25 cost criterion ("Ease of Sales"), which is the correct concept at what I read as the wrong weight for a firm with no sales function.
- Contradiction: the rubric has no owner. Armstrong wrote this for a consulting firm scoring its own products, where "who captures the upside" is not a variable. RDCO's portfolio contains one bet the founder does not own. The shipped rubric cannot see that, and the product-management literature's warning above ("the criteria dictate the answer") predicts exactly this failure. Variant C tests it.
- Contradiction: the ROI column is being scored on stale instrumentation. All three P&L ledgers in
09-bet-stacks/were last updated 2026-05-18 and all three read $0. Scribble Works has no ledger at all. Every ROI score below is therefore a judgment against known-stale data, and the honest read is that the portfolio was asked for a scorecard before it had a P&L. Labeled guess, flagged rather than smoothed over.
The scorecard
Portfolio scored (Ray-assembled, awaiting the founder's read). Versus the question's four: added Scribble Works; renamed CAF → Organizational Intelligence; kept Squarely, MAC, Sanity Check. Not scored, deliberately: Automated Investing (no customer, so Customer Need / Ease of Sales / Distinctiveness are undefined — it is own-capital compounding, not a productization candidate), the acquisition thesis (an option, gated on the wife conversation and the step-away number, with no asset to score yet), and the COO agent / HQ (infrastructure — it is the thing the other bets are scored with, and scoring it separately double-counts it).
Scoring key: Benefits 5 = best. Costs and Risks 5 = worst and are subtracted, per the template.
Variant A — shipped weights, unmodified (this is the literal answer to the question)
| Criterion | W | OI | Scribble Works | Sanity Check | MAC | Squarely |
|---|---|---|---|---|---|---|
| Strategic Alignment | 50 | 5 | 3 | 5 | 3 | 2 |
| Customer Need/Value | 30 | 4 | 3 | 2 | 3 | 2 |
| Distinctiveness | 25 | 2 | 4 | 4 | 3 | 3 |
| ROI (36mo) | 20 | 3 | 2 | 1 | 1 | 1 |
| Ease of Sales (5=hard) | −25 | 3 | 4 | 3 | 4 | 4 |
| Operational Cost (5=high) | −20 | 4 | 2 | 3 | 2 | 2 |
| Capability distance (5=far) | −30 | 2 | 3 | 2 | 3 | 3 |
| Weighted total | 265 | 150 | 235 | 105 | 25 |
Rank: 1. OI (265) · 2. Sanity Check (235) · 3. Scribble Works (150) · 4. MAC (105) · 5. Squarely (25).
Variant B — RDCO-adjusted weights (three changes, each defended below)
Changes: Strategic Alignment 50 → 40; Ease of Sales 25 → 45; new criterion Agent-Routability ×35. Everything else untouched, deliberately — fewer changes is more defensible.
| Criterion | W | OI | SW | SC | MAC | Squarely |
|---|---|---|---|---|---|---|
| Strategic Alignment (four-walls + L5 fit) | 40 | 5 | 3 | 5 | 3 | 2 |
| Customer Need/Value | 30 | 4 | 3 | 2 | 3 | 2 |
| Distinctiveness | 25 | 2 | 4 | 4 | 3 | 3 |
| ROI (36mo) | 20 | 3 | 2 | 1 | 1 | 1 |
| Agent-Routability (move-3 readiness) | 35 | 2 | 5 | 2 | 3 | 2 |
| Ease of Sales (5=hard) | −45 | 3 | 4 | 3 | 4 | 4 |
| Operational Cost (5=high) | −20 | 4 | 2 | 3 | 2 | 2 |
| Capability distance (5=far) | −30 | 2 | 3 | 2 | 3 | 3 |
| Weighted total | 225 | 215 | 195 | 100 | −5 |
Rank: 1. OI (225) · 2. Scribble Works (215) · 3. Sanity Check (195) · 4. MAC (100) · 5. Squarely (−5).
Variant C — add Upside Capture ×40 (the criterion the rubric is missing)
Score = share of the value the founder actually keeps. OI 1 (salary plus $10k cert escalators, no equity). Squarely 4 (co-owned, the IP is his dad's). Scribble Works / Sanity Check / MAC 5 (wholly owned).
| Bet | Variant B | + Upside Capture ×40 | Total |
|---|---|---|---|
| Scribble Works | 215 | +200 | 415 |
| Sanity Check | 195 | +200 | 395 |
| MAC | 100 | +200 | 300 |
| OI | 225 | +40 | 265 |
| Squarely | −5 | +160 | 155 |
Rank: 1. Scribble Works · 2. Sanity Check · 3. MAC · 4. OI · 5. Squarely. The inversion is robust to the weight, not an artifact of picking 40 — the 1-versus-5 score spread does the work. Re-run at ×20 and OI recovers third (SW 315 · SC 295 · OI 245 · MAC 200 · Squarely 75), but it stays below both wholly-owned front-runners at either weight, which is the finding. The OI-versus-MAC ordering is weight-sensitive; the OI-below-SW-and-SC result is not.
Why each weight is what it is
- Strategic Alignment 50 → 40. Armstrong defines it as "fits our core capabilities," a firm-fit test. At ×50 it is 40% of all benefit weight and it lets "adjacent to data engineering" dominate everything else. For a solo founder the real fit test is the four-walls spec, which is stricter and narrower. I lowered rather than removed it because alignment genuinely is the top benefit criterion; I just do not think it should be worth more than the next two combined.
- Ease of Sales 25 → 45. This is the change I hold with the most confidence and it is the only one derived from a verified fact rather than a judgment. The rubric assumes a sales organization exists; RDCO's binding constraint, established 2026-07-31 and re-confirmed by [[2026-09-08-37signals-services-to-product-cashflow-bridge]], is demand generation. A criterion that names the binding constraint cannot sit at ×25 while capability-fit sits at ×50. Every external benchmark in this brief agrees: TPT's curve, beehiiv's 0.62%, the 12-24 month timeline.
- Agent-Routability ×35, new. [[2026-06-28-productized-consulting-scalable-anchor-transition]] concluded that move 3 (decouple delivery from founder hours) is RDCO's actual bottleneck and told this scorecard to weight it heavily. The 37signals brief sharpened it into the one recommendation it held with real confidence: the standing allocation should be denominated in agent-hours with a defended floor. Armstrong's rubric has no column for it because her readers have staff. ×35 puts it above Distinctiveness and below Strategic Alignment, which is where I think it belongs; labeled guess.
- ROI held at 20, deliberately. Every bet in this portfolio has $0 revenue and a stale ledger. Raising ROI weight rewards the best forecast rather than the best bet, and forecast quality is exactly what a solo founder scoring his own bets cannot audit. Holding it low is a bias control, not an economic claim.
- Distinctiveness ×25, Operational Cost ×20, Capability distance ×30 unchanged. No RDCO-specific reason to move them, and unnecessary edits make the whole instrument look tuned to a conclusion.
Sensitivity — read this before the ranking
On an 800-point range, Variant A separates first from second by 30 points (3.75%) and Variant B separates first from third by 30 points. That is not a ranking at the top, it is a tie. Three single-cell flips, each of them a judgment call I could defend either way, reorder it:
- OI Customer Need 4 → 2. I scored 4 for the five verified anchor clusters from [[2026-09-06-move1-anchor-screen-oi-skill-catalog]], not for the catalog as it ships — where roughly 95 of 106 skills answer a consultant's question, not a client's. Score the catalog as-is and OI drops 60 points to 205 / 165, third in both variants. OI's first-place finish is entirely contingent on the isolation actually happening.
- Scribble Works Customer Need 3 → 4. If parents recognize the problem the way the trust ladder claims (and Michelle's handcrafting burnout is a real N=1 signal), Scribble Works takes first place in Variant B at 245.
- Sanity Check ROI 1 → 2. One notch, and it ties Scribble Works.
The one result robust across all three variants is Squarely last, by 80-145 points, with no single flip closing it. MAC is less stable than it looks: fourth in A and B, but third in Variant C, because it is wholly owned and cheap to run. That is worth reading carefully rather than as noise — it says MAC's problem is neither ownership nor operating cost, it is that a finished-ish product has sat pre-launch since a 2026-05-25 target with no distribution attached. Variant C is also the only variant with a clean separation at the top, and it gets there by adding a criterion rather than by tuning a weight, which is precisely the failure mode the product-management literature warns about running in the useful direction.
Synthesis for RDCO
The scorecard's real output is not an order, it is a diagnosis of the instrument. Run with Armstrong's shipped weights, this portfolio ranks OI first, Sanity Check second, and Scribble Works third, and the top three are separated by less than 9% of the scale. Nudge the weights toward RDCO's verified constraints and Scribble Works and Sanity Check swap. Add one criterion — who keeps the money — and the ranking inverts completely, with OI falling from first to fourth. A ranking that survives none of those perturbations should not be used to allocate anything. What it can do, and what I think it does well, is tell you which three judgment calls are actually carrying the decision: whether OI's five anchors get isolated, whether parents recognize the Scribble Works problem, and whether the founder is scoring bets he owns against a bet he does not.
The upside-capture finding is the one I would keep if you kept nothing else. Armstrong's rubric was built for a consulting firm scoring its own products, so ownership is a constant and needs no column. In this portfolio it is a variable with the widest spread of any input, and its absence structurally over-ranks the bet where someone else captures the value. That is not a subtle bias, it is a 150-point swing. It also converges with something the founder already settled in his own words on 2026-07-24: wall #4, real upside stake, is the wall that ruled the old CAF out as a mission. So Variant C is not me discovering a preference for him. It is the shipped rubric being blind to a constraint he had already declared, and the scorecard agreeing with him once it can see it. Worth naming the distinction plainly: OI ranking first on productization-readiness and fourth on ownership-adjusted value is not a contradiction — those are two different questions, and the shipped rubric silently answers only the first.
Where the benchmarks bite hardest is ROI, and they bite in the same direction for every bet. beehiiv's median 0.62% conversion at a $10 median price means Sanity Check needs roughly 16,000 free subscribers for $1,000/mo, or roughly 1,800 if it performs at the Technology top decile. TPT's distribution says most active sellers make about $27/month while the top 1% make about $6,300 — a curve, not a market, and the full-time earners on it almost all started before 2020. The productized-consulting cases run 12-24 months to traction. None of that says any bet is bad. It says every one of them is a distribution problem wearing a product costume, which is the same conclusion the 37signals brief reached from a completely different direction, and the same one the demand-generation finding reached in July. Three independent paths, one answer. That convergence is stronger evidence than any cell in the tables above.
What I would actually do with this, stated as a recommendation and not a decision. The scorecard supports a small, reversible allocation move and nothing larger. Scribble Works scores 5 on Agent-Routability and is the only bet that does; it is already running the studio pattern nightly. Sanity Check scores 5 on Strategic Alignment because it is the demand asset for a portfolio whose binding constraint is demand, and it has been finished-but-unpublished for roughly 160 days. Those two are the standing agent-hour allocation the 37signals brief argued for, and the case for them does not depend on which of them ranks second. Squarely is the clean result — last in all three variants by 80-145 points — and the honest version of that is dormancy with a named revisit trigger, not a kill. MAC is the ambiguous one and should not be filed alongside Squarely: it rises to third in Variant C, which says its low scores come from four months of non-launch rather than from anything structural, and that is a fixable condition rather than a verdict. OI needs no allocation decision at all — it is the funding leg and the medium, it is where the founder's hours already go, and the only thing this scorecard asks of it is the isolation move that [[2026-09-06-move1-anchor-screen-oi-skill-catalog]] already specified.
Why this is in the vault
It closes the open checkbox left in [[2026-06-25-productize-framework-armstrong-vecteris]] and re-opened as a follow-up by [[2026-06-28-productized-consulting-scalable-anchor-transition]], and it does three things beyond that: it corrects the vault's recorded scoring method (Cost and Risk columns are inverted and subtracted, not summed), it registers the post-2026-08-10 OI rename and the addition of Scribble Works in a document future bet-allocation questions will read, and it puts a number on the agent-hour allocation argument that [[2026-09-08-37signals-services-to-product-cashflow-bridge]] made qualitatively.
Open follow-ups
- Should Upside Capture be added permanently to RDCO's copy of the Portfolio Scoring rubric, and at what weight? Founder call, since it encodes his own four-walls wall #4. This brief proposes ×40 and shows the result is robust from ×20 up.
- The three P&L ledgers in
09-bet-stacks/have been stale since 2026-05-18 and Scribble Works has none. Every ROI score here is a judgment against known-stale data. What is the minimum viable per-bet revenue instrumentation, and does Scribble Works need a ledger before or after its first dollar? - The Scribble Works charter still prices founder time at a $150/hr placeholder, which sets a break-even band of 4 to 971 subscribers. That single unresolved input dominates its ROI cell. What replaces it?
- Sanity Check's ROI score of 1 measures it as a product. Its actual role is a funnel for everything else. What is the right way to score a bet whose value is entirely an externality on the other bets, given that Armstrong's rubric has no column for it?
- Does the five-anchor isolation from the move-1 screen have an owner and a date inside phData, or is OI's top ranking resting on a move nobody has been assigned?
- Squarely scores lowest in every variant but has an iOS app in the deploy pipeline whose launch is not reflected in any score here. Does a shipped app change the Distinctiveness or Ease-of-Sales cells enough to matter, or does it confirm the dormancy read?
Related
- [[2026-06-25-productize-framework-armstrong-vecteris]]
- [[2026-06-28-productized-consulting-scalable-anchor-transition]]
- [[2026-09-08-37signals-services-to-product-cashflow-bridge]]
- [[2026-09-06-move1-anchor-screen-oi-skill-catalog]]
- [[2026-09-05-kids-subscription-box-comparables-scribble-works]]
- [[2026-08-31-studio-charter]]
- [[2026-05-20-services-pricing-model-for-rdco-future]]
- [[2026-05-29-spine-as-a-service-productized-conversion-playbook]]
- [[user-career-commitment-shape]]
Sources
Vault:
- [[2026-06-25-productize-framework-armstrong-vecteris]] — the Productize Pathway, shipped Portfolio Scoring weights, Seven Deadly Mistakes
06-reference/assets/2026-06-25-productize-tools/extracted/4. Product Portfolio Scoring Template.pptx— primary source, read directly for this brief; supplied the weights AND the inverted cost/risk polarity that the vault note had recorded incorrectly. Vecteris copyright, internal use only, do not republish- [[2026-06-28-productized-consulting-scalable-anchor-transition]] — the parent brief; move-1/2/3 framing, the productized-consulting case benchmarks, the open checkbox this closes
- [[2026-09-08-37signals-services-to-product-cashflow-bridge]] — agent-hour "third client" allocation; MAC and Sanity Check as the gap-filler tier; "RDCO is at 2001-2002"
- [[2026-09-06-move1-anchor-screen-oi-skill-catalog]] — the U/E/R/S screen on the 106-skill OI catalog; the five clearing clusters; 95-of-106 consultant-facing finding
- [[2026-09-05-kids-subscription-box-comparables-scribble-works]] — Lovevery / KiwiCo / Osmo price anchors; churn and LTV correction; the acquisition-model gap
- [[2026-08-31-studio-charter]] — Scribble Works org design, $0.710/pack marginal cost, break-even band, kill-criteria dormancy
- [[2026-05-20-services-pricing-model-for-rdco-future]] — services-leg + productized-leg architecture; founder attention as the binding input
- [[2026-05-29-spine-as-a-service-productized-conversion-playbook]] — why bespoke per-client conversion engagements were rejected
- [[user-career-commitment-shape]] — the four-walls spec; wall #4 (upside stake) and the CAF resolution
09-bet-stacks/{squarely,mac,sanity-check}-pnl-ledger.md— all $0, all last updated 2026-05-18 (the staleness caveat)01-projects/squarely-puzzles/STRATEGY.md,01-projects/newsletter/STRATEGY.md,01-projects/mac/— per-bet current state
Web:
- https://www.beehiiv.com/blog/the-state-of-paid-newsletters-2026 — primary; median 0.62% free-to-paid, top quartile 2-5%, Technology top decile 5.55%, $10/mo and $100/yr medians, 5.06-16.67% monthly churn, $83-230 median LTV
- https://medium.com/write-rise/tpt-sellers-made-253-million-last-year-most-people-have-no-idea-5b552ea873e5 — TPT $253M seller payouts, top-heavy distribution, 300+ million-dollar sellers
- https://seolumina.com/blog/how-much-do-tpt-sellers-make-in-2026-real-data-income-breakdown — ~$27/mo median active seller vs ~$6,300/mo top 1%; pre-2020 cohort effect
- https://productschool.com/blog/product-fundamentals/weighted-scoring-model — weighted scoring mechanics and the criterion-selection pitfall
- https://airfocus.com/blog/weighted-decision-matrix-prioritization/ — weighted decision matrix; "the criteria you use dictate what you choose"