06-reference/research

correctness criticality cannibalization threshold

2026-09-23·research-brief·source: deep-research·by Ray Data Co (deep-research synthesis)
snowflakecortex-senseorganizational-intelligencecorrectness-criticalityevidence-discipline

There Is No Stakes Threshold, Because 83% Is a Ceiling on Snowflake's Own Data, Not a Floor on the Client's — and the Cannibalization Variable Is Detection Cost, Not Accuracy

The question

Verbatim, as filed 2026-07-08: "At what use-case stakes / correctness-criticality level does Cortex Sense's free 83%-accuracy floor cannibalize a paid CAF engagement — what does a segmentation cut of CAF's use-case ledger by correctness-criticality reveal?"

Context: this was open follow-up #4 in [[2026-07-08-cortex-sense-semantic-layer-wedge-caf]], auto-promoted 2026-07-09. Naming note, once: "CAF" was retired as a program name on 2026-08-10; the work now runs under the Organizational Intelligence (OI) umbrella (Organizational Platform to Org Map, Intelligence Platform to Intelligence Maturity Assessment, plus OIP and Pulse). The question text keeps the old name because it was filed in July; the rest of this brief uses current naming.

Scope, stated up front. The second clause asks for a cut of an internal engagement ledger. That cut is not performable and this brief does not attempt it, estimate it, or reconstruct it. Section "Synthesis" says why, and names the specific inputs that would make it performable. Only the framework half is researched here.

What we already know (from the vault)

What the web says

Convergences and contradictions

Synthesis for RDCO

The first half of the question dissolves on inspection, and the dissolution is the finding. Three of the premise's four load-bearing words fail. It is not free (consumption-priced in AI Credits; Snowflake's own figure has the grounded query cheaper, not costless). It is not a floor (it is a ceiling measured on the single most favourable estate in existence for it - Snowflake's own instrumented internal analytics data with semantic views on hand - so a mid-market estate should be expected to land below it, not at it). It is not reliably 83 (the same benchmark publishes as 21/23/24.1/~25 to 83/86.3 across the vendor's own surfaces, with undisclosed n, undisclosed graders, and undisclosed correctness criteria). And as of the last GA/preview audit it is not purchasable. Inverting "ceiling" to "floor" is what made this question feel urgent: it silently assumed the client gets Snowflake's best case for nothing. Nothing in the evidence supports that, and a DSA who argues from a stakes threshold derived from it is arguing from a number the client can look up and dispute in the room.

The variable that actually governs cannibalization is detection cost, and the accuracy framing hides it. A wrong answer is only cheap when someone finds out. Snowflake's own material describes no refusal or confidence-surfacing behaviour, so the failure mode is a confidently wrong number with no flag. That makes the decisive question not "how often is it wrong" but "who would notice, how fast, and at what cost." This reframes the erosion risk in a way that is more defensible than the parent brief's version: a use case is genuinely cannibalizable only when consequence is low and detection is cheap - the analyst eyeballs the number, sees it is off, and moves on. Where detection is expensive or deferred (a metric that only reconciles at quarter-close, a definition that only surfaces as wrong when two departments compare decks), low stakes do not protect you, because silent error accumulates unpriced. That is the segment the free-tier narrative is worst at describing, and it is where an OI engagement is easiest to justify without quoting any percentage at all.

A correctness-criticality taxonomy for analytics and text-to-SQL use cases should run on five axes, not one stakes dial. Synthesising SR 11-7's materiality drivers, the translation ladder's assignment criteria, and the vault's own gate-placement rule: C1 consequence severity (cosmetic / internal decision input / external commitment / financial-statement or regulatory); C2 detection latency and cost (self-evidently wrong / caught at reconciliation / undetectable without a rebuild); C3 reversibility and blast radius (retractable before propagation, or not - the vault's existing gate rule keyed directly to this); C4 definitional ambiguity density (how many contested definitions the answer traverses: which revenue, whose fiscal calendar, what nets against what); C5 action coupling (a human reads the number, versus an agent acts on it unattended - SR 11-7's "degree of manual intervention," and the axis GARP flags as where the old framework strains). Then assurance scales with tier, in ITAAC three-part form per [[2026-06-11-high-reliability-acceptance-gates-agent-contracts]]: per use case, the commitment, the verification method, and the acceptance criterion. The non-obvious payoff of splitting the axes is that cannibalization risk and engagement value load onto different ones. Risk of being skipped is roughly C1 low and C2 low. Value of the governed layer is roughly C4 high and C5 automated - because a context-retrieval service can retrieve competing definitions but cannot adjudicate between two legitimate ones, which is authorship, not retrieval. A use case can therefore be low-stakes and still un-helped by Sense; it just will not be worth paying to fix. Those cells are the honest "do not sell into this" list, and they are narrower than a single stakes cut would suggest.

On the threshold itself, the calibrated answer is a shape, not a number, and I am reasoning by analogy here rather than from evidence about analytics. The translation market is the closest thing to a natural experiment, and its outcome was re-stratification with rate compression, not displacement of the paid tier: the free tier took the volume, standards bodies codified the middle (ISO 18587), and the priced expertise migrated from production toward post-editing and tier assignment itself. Transferring that to OI: expect the commoditized half to be context assembly (which [[2026-07-08-cortex-sense-semantic-layer-wedge-caf]] already concluded), expect the surviving paid work to concentrate in adjudicating definitions and proving correctness, and expect deciding which use case belongs in which tier to become a deliverable in its own right rather than a pre-sales judgment call. That transfer is an analogy between two markets with different buyers, different liability structures and different standards regimes, so treat it as a hypothesis that shapes the delivery method, not as a prediction with a date on it.

For the founder's actual seat - DSA and TAL, delivery and architecture, not sales engineering - the useful output is a method asset, not a talk-track. Two concrete moves follow. First, make correctness-criticality a typed field set (C1-C5) on the Build Manifest, with a verification method bound to each tier. That merges tiering with gate placement into one mechanism and completes the ITAAC three-part upgrade the June brief already recommended, which is a practice-level asset that scales across engagements rather than a per-deal argument. Second, make the Confidently-Wrong Probe from [[2026-09-20-cowork-plugin-accuracy-unmodeled-midmarket-data]] tier-aware: sample its question battery across C1-C5 cells rather than just "hard questions," so the deliverable is a per-tier error profile on the client's own estate instead of one alarming aggregate percentage. A per-tier profile is also the only instrument that could eventually answer this question empirically, because it measures the client's required accuracy and the platform's delivered accuracy in the same run.

Why the ledger cut is not performable, in order of severity. (1) The field probably does not exist. The ledger in question is the governed use-case portfolio and typed Build Manifest emitted by Assess (schema v1.1 plus validator, per [[2026-06-09-caf-restructure-proposal]] and the OI roadmap). The vault holds no record of any correctness-criticality attribute in that schema. If it is absent, there is nothing to cut - the fields have to be specified and back-filled first, which is a schema question, not a data pull. (2) n is far too small. As of the roadmap, two engagements had run end-to-end. Two cases cannot support a threshold claim regardless of how clean the fields are. (3) Access and boundary. The repo is not on this machine ([[2026-06-15-caf-architecture-decisions-and-meta-council]]), and phData engagement data must not be persisted to the vault in any case, so even a clean cut belongs inside phData and not here. The specific inputs that would make it performable: C1-C5 scored per use case with a written rubric; a recorded disposition per use case (built / deferred / declined) with the client's stated reason, since "we think the platform covers this" is the only direct observation of cannibalization and it lives in the sales conversation rather than in the manifest's technical fields; and engagements in the tens rather than the ones.

Why this is in the vault

This brief closes open follow-up #4 of [[2026-07-08-cortex-sense-semantic-layer-wedge-caf]] by rejecting its premise rather than answering it, and it stops a stakes-threshold argument built on the already-retired 83% figure from reaching an OI proposal or an Intelligence Maturity Assessment deliverable. It also supplies the concrete C1-C5 axis set that the Build Manifest schema and the Confidently-Wrong Probe would need in order to make the question empirically answerable later, and it records - so it is not re-asked - that the ledger cut is blocked on a missing schema field and n=2, not merely on repo access.

Open follow-ups

Related

Sources

Vault:

Web:

Research caps used: 4 of 5 qmd queries, 3 of 3 WebSearch, 3 of 3 WebFetch, 6 vault docs. No paywalls encountered.