Only five things in the 106-skill catalog are problems a client would name unprompted — and the July refactor cut two of them
The question
Verbatim: "Test the move-1 'urgent + recognized problem' anchor screen against CAF's 106 skills — which ~5 clear the bar?"
Context: this is the third open follow-up from [[2026-06-28-productized-consulting-scalable-anchor-transition]], and the companion open checkbox in [[2026-06-25-productize-framework-armstrong-vecteris]]. Naming convention note for traceability: "CAF" was retired as a name on 2026-08-10; the work now runs under the Organizational Intelligence (OI) umbrella (Organizational Platform → Organizational Map · Intelligence Platform → Intelligence Maturity Assessment · OIP + Pulse). The 106-skill artifact this question asks about is the caf-engine skill library, which predates the rename.
What we already know (from the vault)
- The screen, as Armstrong actually states it. The load-bearing filter is "the urgent and expensive problem the customer knows they have." Verbatim from the vault: "A merely frequent problem is not enough — it has to be urgent AND expensive AND already-recognized by the customer" ([[2026-06-25-productize-framework-armstrong-vecteris]]). It sits at Pathway stage 3, Define the Urgent and Expensive Problem.
- The parent brief's move-1, verbatim: "Standardize one urgent-and-recognized recurring problem into a fixed-scope, fixed-price anchor offer — the thing you keep re-solving, packaged so it sells identically every time" ([[2026-06-28-productized-consulting-scalable-anchor-transition]]).
- The parent's own instruction on how to use it: "don't catalog all 106; isolate the handful tied to an urgent+expensive client problem (Portfolio Scoring weight: Customer Need ×30, Strategic Alignment ×50) and productize those first."
- The catalog was never enumerated in the vault before this brief. Every prior vault reference cites the number 106 and the phase taxonomy, never the skill list ([[2026-06-10-caf-current-state-teardown]], [[2026-06-09-caf-restructure-proposal]]). The list exists only in the code checkout.
- A commoditization constraint the screen alone will not catch: big-4 and boutique competitors all sell the readiness diagnostic; the differentiated ground is coupling the diagnostic to a build ([[2026-07-15-agentic-assessment-framework-competitive-landscape]]).
Verification of the count — 106 is real, but it is a July snapshot
Counted directly against the local checkout at ~/Projects/phdata-private/phdata-ai-wf-plugins-dd673d8e215f/plugins/caf/skills (dated 2026-07-03): exactly 106 skill directories. The backlog row's number is correct as of that tree. Three corrections ride on top of it:
- The live surface is 70, not 106. Founder, 2026-07-27: "We have also refactored CAF. We cut out phases 5-8" ([[2026-06-28-caf-8-phase-structure-and-skill-pipeline-mapping]], correction block; confirmed in [[2026-07-27-caf-technical-architecture-and-backlog]]). The phase-5–8 prefixes are
6a(5) +6b(6) +6c(5) +6d(5) +s5(7) +s6(5) +s7(3) = 36 skills cut. Phases 1–4 plus the Meta-Council leave1+2+3+4a+4b+4c+4d+4g+s4+m= 70. - A second, non-interoperable tree exists. The same repo carries a project-scoped
.claude/skillsimplementation with 12 entries, where the ~106 dimensions live asreference/*.mdinside 8 phase skills ([[2026-07-01-caf-ecosystem-map-and-brigade-restructure-read]]). "The catalog" is ambiguous unless you say which tree. - The current OI spec carries no catalog of this shape at all.
~/Documents/phdata-projects/organizational-intelligence/spec/is a product book plus registries: the Capability Registry is 6 layers → 28 sub-capabilities → 159 capability details, alongside a Role Catalog and Value Registry. Whethercaf-enginestill executes underneath the OI deliverables is not verified here [assumption flagged] — the OI repo does not reference it.
What the web says
- The 2026 productized-services literature converges on a triad that is effectively my fourth gate: a successful productized offer needs a clear starting state, a defined transformation, and a specific deliverable (ManyRequests; Assembly). No source I found names Armstrong's screen by that label, so the screen stays a vault-internal construct.
- Buyers now recognize the ROI problem out loud. Companies expect ~171% ROI from agentic AI but only 39% can point to any EBIT impact, and Gartner projects 40%+ of agentic AI projects cancelled by 2027 on unclear ROI, escalating cost and inadequate controls (Futurum; The AI Index).
- Buyers now recognize the adoption problem out loud. 80% of enterprise applications shipped in Q1 2026 embed at least one AI agent, but only 31% of organizations have an agent running in production; 95% of pilots produce no measurable P&L impact and 88% of POCs never reach production (Digital Applied; Paul Okhrem).
- Buyers now recognize the verification problem out loud. "Agent washing" has made buyer due diligence on actual agentic capability a named purchasing step, with deflection claims above 80% flagged for cross-check against accuracy data (Digital Applied).
Those three are the only problem statements in this category with independent, quantified evidence of buyer-side recognition in 2026. That is not decoration — it is the evidence base for gate R below.
Convergences and contradictions
- Convergence: the market recognizes exactly the problems that live in the cut half of the catalog. The web's three recognized problems (ROI unproven, nothing gets adopted, answers can't be trusted) are post-build problems. CAF/OI's surviving phases 1–4 stop at the C4 Build Manifest — a plan. The catalog's answers to the recognized problems sit disproportionately in the 36 skills cut on 2026-07-27.
- Contradiction: the backlog row drops a criterion. The row says "urgent + recognized." Armstrong's screen is urgent AND expensive AND recognized. Dropping "expensive" would let cheap-but-annoying problems through. I have tested against all three; if the founder intended the two-criterion version, results 4 and 5 below are unaffected and nothing new clears.
- Contradiction: passing the screen is not sufficient.
2-1-agentic-maturity-assessorscores 7/8 on the screen and is still the worst anchor candidate in the catalog, because every competitor sells it ([[2026-07-15-agentic-assessment-framework-competitive-landscape]]). The screen measures demand, not defensibility. It needs a distinctiveness gate bolted on before it is used to pick anything.
Synthesis for RDCO
The scoring rule (my operationalization, not the founder's, not Armstrong's). Armstrong states the screen prosaically; it is not a rubric anywhere in the vault. I turned it into four gates scored 0 / 1 / 2, and I am labeling it as mine:
| Gate | Passes at 2 when… |
|---|---|
| U — Urgent | the client feels it on a clock: a budget cycle, a board question, a stalled go-live, a regulator |
| E — Expensive | not solving it costs six figures or more in wasted build, wasted licence, rework, or penalty |
| R — Recognized | the buyer states the problem in their own words before we arrive. The killer gate |
| S — Standardizable | sellable as fixed-scope, fixed-price, identical deliverable every time, without the rest of the pipeline running first |
Clears the bar = U, E and R all at 2, plus S ≥ 1 (total ≥ 7/8). One methodological finding first: no single skill clears S at 2 on its own. These are sub-atomic consultant micro-steps, not offers. A sellable anchor is a tight cluster of 2–3, so each result below names a cluster and its lead skill.
The five that clear (Ray-selected, awaiting the founder's read).
- Adoption Risk Screen — lead
4a-2-adoption-modeler, with4a-6-workflow-intrusion-analystand6a-2-adoption-monitor. 8/8. "We bought it and nobody uses it" is the single most-stated buyer complaint of 2026 (80% embed vs 31% in production), it is urgent at renewal, and the sunk licence plus build cost is the expense. Fully standardizable: score a proposed feature for behavioral friction, workflow disruption, champion presence, install barriers, return a Confidence number. - AI Portfolio Value Screen — lead
2-7-value-screener, with4a-3-time-value-quantifierand4b-6-risk-adjusted-value. 8/8. The skill's own description is "prevent innovation theater," which is the CFO's question verbatim in a year when Gartner forecasts 40%+ cancellations for unclear ROI. Standardizable as a fixed-scope pass over an existing initiative list. - Agent Correctness & Eval Harness — lead
6d-8-adversarial-correctness-validator, with6b-6-evaluation-harness. 8/8. Independent re-execution, magnitude check, lineage cross-reference, freshness check on generated SQL, code and narrative. "How do we know the answer is right" now blocks go-live and is the concrete form of the agent-washing due-diligence step. Both skills are in the cut 36. - AI Oversight & Accountability Boundary — lead
2-5-hitl-boundary, with3-9-compliance-scannerand4c-1-accountability-mapper. 7/8 (S=1). Who is liable, what is explainable, what do we show the auditor. Urgent because it gates production in regulated industries, expensive because a stop-ship late is the worst-cost outcome. S is 1 because the deliverable varies materially by regulator, so it will not sell perfectly identically. - Domain Rule Extraction —
3-9c-domain-rule-extractor, standalone. 7/8 (U=1). Clients say "all the rules live in Dave's head" unprompted, so R is a clean 2, and the skill's own text calls it "the single highest-ROI addition in v3.0 for correctness" — silently-wrong outputs are the expense. U is 1 because urgency is event-driven: a retirement, an attrition, a migration. No event, no clock.
The next-closest, and the gate each one fails. 2-1-agentic-maturity-assessor (7/8, fails the unwritten distinctiveness gate — everyone sells it). 3-6-value-leakage (U=1: "quantifies loss without implying feasibility," and a diagnostic with no action attached does not create urgency to buy). 2-3-data-readiness-interface-steward (U=0: "our data isn't ready" is the most-recognized and least-urgent problem in the category, perennially deferred and thoroughly commoditized). 3-4-exception-miner (R=0 by construction — it surfaces hidden complexity, and a client cannot recognize a problem defined as hidden). 6c-4-drift-monitor and 6b-4-failure-analyst (R arrives only after deployment, making these expansion offers rather than anchors).
What the screen actually revealed, which is not the list of five. The screen does not sort the catalog by phase, by quality, or by effort. It sorts by direction of gaze. Roughly 95 of the 106 answer a question the consultant has — classify this input, propagate this tag, harmonize these metrics, validate this contract. The handful that clear name a failure the client has already lived through: money spent, nobody used it, the answer was wrong, the auditor asked. That is the whole finding, and it is a positioning result rather than a prioritization result. It also explains the adoption number in [[2026-07-01-caf-ecosystem-map-and-brigade-restructure-read]] — six hours per company against colleagues' thirty minutes on the old monolith. A pipeline built almost entirely of consultant-facing micro-steps optimizes for consultant rigor, and rigor is not what the buyer is paying to remove.
The uncomfortable structural consequence. The 2026-07-27 refactor cut 36 of 106 skills, and that cut is not neutral against this screen: two of my five anchors sit wholly or partly inside the cut set, and three of the six next-closest do too. The reason is structural rather than accidental. Urgent, expensive, recognized problems in this category are overwhelmingly post-build problems, and phases 1–4 terminate at a plan. Truncating at the C4 Build Manifest optimizes the framework for the pre-sales motion, which is a defensible call for a sales-facing marketplace, and it simultaneously removes the part of the catalog with the strongest independent buyer demand. Both things are true. The productization-honest read is that OI's front half is a lead product and its anchor-offer candidates live in the half that was cut, so if anchor revenue is the goal, the cut half needs a home rather than a deletion.
Where this lands against the parent brief's move-3 conclusion. The parent argued RDCO's binding constraint is move 3, decoupling delivery from founder hours, and that the transition is the 37signals shape rather than a firm conversion. This test does not disturb that. It sharpens the move-1 input: the five anchors above are the only defensible candidates for the "isolate the handful" instruction, and four of the five are outcome verification offers (did it get adopted, did it pay, is it right, is it defensible) rather than assessment offers. Assessment is what the catalog is mostly made of and it is what the market has commoditized. Verification is what the market has started to recognize and has not yet standardized. If there is a productizable anchor anywhere in this catalog, that is the shape of it.
Why this is in the vault
This closes the third open follow-up in [[2026-06-28-productized-consulting-scalable-anchor-transition]] and the "which of the 106 skills clear the urgent + expensive bar" checkbox in [[2026-06-25-productize-framework-armstrong-vecteris]], and it supplies the Customer Need (×30) input that the still-unbuilt weighted Portfolio Scoring scorecard needs before it can rank the OI bet against Squarely, MAC and Sanity Check. It also corrects a live factual drift the founder will otherwise carry into that scorecard: the catalog's live surface is 70 skills post-refactor, not 106.
Open follow-ups
- Does
caf-enginestill execute under the OI umbrella, or has the 106/70 catalog been superseded by the Capability Registry (6 layers / 28 sub-capabilities / 159 details)? Everything above assumes the catalog is still live machinery. If it is not, the screen should be re-run against the Capability Registry instead, and this brief becomes a historical scoring of a retired artifact. - Where did the 36 cut skills go? Deleted, archived, or moved into the delivery-side framework? This determines whether anchors 3 and 5 are recoverable assets or need rebuilding.
- Add a distinctiveness gate to the screen.
2-1-agentic-maturity-assessorpassing at 7/8 while being the most commoditized item in the catalog is a live false positive. What does the D gate look like, and does adding it change the ranking of the five? - Do the five anchors survive contact with a real buyer? All recognition evidence here is market-level survey data, not a phData buyer saying it. Armstrong's mistake #5 is "designing in a vacuum," and this brief was written in one.
- Which of the five is closest to already-existing artifact? The Adoption Risk Screen and the Portfolio Value Screen both look like they could be run today against the Kwik Trip use-case portfolio as a zero-cost pilot. Worth checking before any of this is built.
Related
- [[2026-06-28-productized-consulting-scalable-anchor-transition]]
- [[2026-06-25-productize-framework-armstrong-vecteris]]
- [[2026-06-10-caf-current-state-teardown]]
- [[2026-06-09-caf-restructure-proposal]]
- [[2026-06-28-caf-8-phase-structure-and-skill-pipeline-mapping]]
- [[2026-07-27-caf-technical-architecture-and-backlog]]
- [[2026-07-01-caf-ecosystem-map-and-brigade-restructure-read]]
- [[2026-07-15-agentic-assessment-framework-competitive-landscape]]
Sources
Vault:
- [[2026-06-28-productized-consulting-scalable-anchor-transition]] — the parent brief; move-1 verbatim and the "don't catalog all 106" instruction
- [[2026-06-25-productize-framework-armstrong-vecteris]] — the "urgent and expensive problem the customer knows they have" screen verbatim; Portfolio Scoring weights; Seven Deadly Mistakes
- [[2026-06-10-caf-current-state-teardown]] — 106 count, 7-phase taxonomy, Meta-Council, why the count is 106
- [[2026-06-09-caf-restructure-proposal]] — the 106 → ~15-20 collapse proposal
- [[2026-06-28-caf-8-phase-structure-and-skill-pipeline-mapping]] — the 2026-07-27 correction block recording the phases 5-8 cut
- [[2026-07-27-caf-technical-architecture-and-backlog]] — post-refactor four-phase structure; three-repo naming
- [[2026-07-01-caf-ecosystem-map-and-brigade-restructure-read]] — the two coexisting implementations; 6 hours vs 30 minutes runtime; per-skill evals
- [[2026-07-15-agentic-assessment-framework-competitive-landscape]] — everyone sells the diagnostic; the distinctiveness constraint
Primary artifacts (local, phData-internal, not vault docs):
~/Projects/phdata-private/phdata-ai-wf-plugins-dd673d8e215f/plugins/caf/skills— the 106 skill directories and their frontmatter descriptions, counted and read 2026-09-06 (checkout dated 2026-07-03)~/Documents/phdata-projects/organizational-intelligence/spec/01-index.mdandspec/A1-capability-registry.md— current OI structure and the 6/28/159 Capability Registry
Web:
- https://futurumgroup.com/press-release/enterprise-ai-roi-shifts-as-agentic-priorities-surge/ — 171% expected ROI vs 39% able to show EBIT impact
- https://report-ai.org/indexes/enterprise-ai/enterprise-ai-statistics-2026/ — Gartner 40%+ agentic project cancellation forecast
- https://www.digitalapplied.com/blog/ai-agent-adoption-2026-enterprise-data-points — 80% embed vs 31% in production; 95% of pilots no P&L impact; 88% of POCs never reach production
- https://www.digitalapplied.com/blog/ai-customer-support-statistics-2026-adoption-roi-data — agent washing and buyer due diligence
- https://paul-okhrem.com/enterprise-ai-agents-statistics-2026/ — enterprise agent deployment gap
- https://www.manyrequests.com/blog/productized-consulting — starting state / transformation / deliverable triad
- https://assembly.com/blog/productized-services — productized-services packaging patterns, 2026 edition