There Is No Stakes Threshold, Because 83% Is a Ceiling on Snowflake's Own Data, Not a Floor on the Client's — and the Cannibalization Variable Is Detection Cost, Not Accuracy
The question
Verbatim, as filed 2026-07-08: "At what use-case stakes / correctness-criticality level does Cortex Sense's free 83%-accuracy floor cannibalize a paid CAF engagement — what does a segmentation cut of CAF's use-case ledger by correctness-criticality reveal?"
Context: this was open follow-up #4 in [[2026-07-08-cortex-sense-semantic-layer-wedge-caf]], auto-promoted 2026-07-09. Naming note, once: "CAF" was retired as a program name on 2026-08-10; the work now runs under the Organizational Intelligence (OI) umbrella (Organizational Platform to Org Map, Intelligence Platform to Intelligence Maturity Assessment, plus OIP and Pulse). The question text keeps the old name because it was filed in July; the rest of this brief uses current naming.
Scope, stated up front. The second clause asks for a cut of an internal engagement ledger. That cut is not performable and this brief does not attempt it, estimate it, or reconstruct it. Section "Synthesis" says why, and names the specific inputs that would make it performable. Only the framework half is researched here.
What we already know (from the vault)
- The 83% is a vendor self-report on the vendor's own data, and the vault has already retired it. [[2026-09-20-cowork-plugin-accuracy-unmodeled-midmarket-data]] traced it to primary source: Snowflake's own post reports 24.1% to 86.3%, comparing a frontier coding agent over MCP, vanilla CoCo, and CoCo grounded by Cortex Sense, on "product analytics questions on Snowflake's internal data." Query count, gold-answer authorship, and the grading rubric are all undisclosed. The same baseline appears as 21%, 23%, 24.1% and ~25% across Snowflake's own surfaces, and the ceiling as 83% or 86.3%. That brief's verdict: record it as derived and unverified for our use, and do not build a client artifact on it.
- It is a ceiling, not a floor, and the question inverts the direction of the uncertainty. The benchmark holds the data estate constant at Snowflake's own well-instrumented internal analytics data (with semantic views available to the grounded configuration) and varies only the agent harness. A mid-market estate with no modeled gold layer is the less favorable case, so the expected number there is below 83, not at it. The parent brief's risk sentence ("the free 83% floor may be good enough to skip a paid engagement") inherits the inversion and therefore overstates the threat.
- "Free" does not survive either. The vault records no separate SKU for Cortex Sense, but consumption is billed in AI Credits at $2.00 global / $2.20 regional, and Snowflake's own figure has the grounded query at $0.59 vs $1.76 ungrounded ([[2026-09-12-pupius-four-axis-snowflake-data-ai-strategy]], [[2026-07-08-cortex-sense-semantic-layer-wedge-caf]]). Sense is cheaper, not free. The competitive mechanism is price compression on one line item, not a zero-price entrant.
- And it is not purchasable. As of 2026-09-14 Cortex Sense and the domain plugins were private preview at most ([[2026-09-14-snowflake-cowork-cortex-ga-vs-preview-matrix]] via [[2026-09-20-cowork-plugin-accuracy-unmodeled-midmarket-data]]). A capability a buyer cannot switch on at contract date cannot cannibalize that contract today.
- The vault's own graded-assurance rule already answers half the taxonomy question. [[2026-06-11-high-reliability-acceptance-gates-agent-contracts]]: "place a HITL gate wherever the next phase multiplies the blast radius of an undetected error or makes it irreversible; everywhere else, an autonomous verifier suffices," plus Mark 7, "scale verification to failure cost," and the ITAAC three-part format (design commitment / verification method / acceptance criterion).
- The one vault benchmark with a real modeled-vs-unmodeled delta points the other way. [[2026-04-07-dbt-semantic-layer-vs-text-to-sql-benchmark]]: semantic layer 98.2-100%, raw text-to-SQL 84.1-90% on an 11-question insurance set. A 10-15 point gap, not 60. Also a vendor benchmark.
What the web says
- The displacement threshold in the automation literature is not a capability number; it is the required accuracy of the task, and required accuracy is set by consequence. Li, Aboutorabi, Lyu, Qian, Fleming, Goehring and Thompson (arXiv:2603.29121, preprint, 2026-03-31) model a required-accuracy threshold per task, elicited from "domain-expert surveys eliciting acceptable error rates" rather than read off the tool. Automation is full only when delivered accuracy meets that task-specific requirement.
- Accuracy cost is convex, which makes partial automation the long-run equilibrium rather than a way station. Same paper: "the cost of achieving 'good' accuracy on a task may be relatively inexpensive, but pushing from good to near-perfect accuracy...can be orders of magnitude more costly," and "when the marginal cost of further accuracy improvement exceeds the marginal labor saving it enables, firms optimally stop short of full automation." Partial automation "occupies the largest share of the task space." Their computer-vision estimate: only ~11% of exposed labour compensation is economically attractive to automate, most of it partially.
- The comparable market that already ran this experiment is translation, and it re-stratified rather than collapsed. Machine translation crossed "good enough" for comprehension years ago. The industry response was a formal tier ladder with standards attached: raw MT for "internal gisting" and preliminary research only; light post-editing for internal/informational use where "minor stylistic imperfections are acceptable"; full post-editing to ISO 18587 for publication-grade high-volume content; and full human translation to ISO 17100 with dual review for legally significant, safety-critical, patient-facing and brand-voice content (Tomedes, a translation vendor with a commercial interest in tier selection, so treat the framing as motivated).
- Tier assignment in that market turns on audience, publication-vs-internal, regulatory exposure and reversibility — not on the tool's score — and mis-tiering is framed as cost transfer, not cost saving. Same source: "choosing a lower tier to save money on a high-risk document does not reduce cost - it transfers cost to rework, liability, or reputational damage downstream." It reports post-editing share rising from 26% in 2022 to nearly 46% by late 2024 (vendor-cited, unverified at source).
- Banking already has the canonical correctness-criticality tiering framework, and its central principle is proportionality. SR 11-7 tiers models by materiality drivers — financial impact, decision criticality, regulatory relevance, complexity, and degree of manual intervention — and expects validation depth, monitoring cadence and attestation to scale with tier (Umbrex SR 11-7 framework summary; RiskTemplates, AI model risk tiering, which adds autonomy level, explainability gap, data sensitivity and velocity of change as AI-specific drivers).
- But the framework strains at exactly the place this question lives, and nobody has published the fix. Krishan Sharma (Citigroup Model Risk Management), GARP, 2026-02-27: "sound governance, independent validation, and effective challenge" still hold, but SR 11-7 assumed "bounded scope, stable parameters, and decision paths that could be reconstructed ex post" and systems are now "dynamic rather than static, probabilistic rather than deterministic." Notably, the article names the strain and proposes no materiality thresholds, tiers, or escalation triggers, and says nothing about silent or unflagged wrong outputs (GARP).
Convergences and contradictions
- Convergence across three unrelated domains on the same structural answer: tier by consequence, then scale assurance to the tier. SR 11-7 (models), the translation quality ladder (content), and the vault's own high-reliability gate work (agent contracts) independently arrive at proportional assurance keyed to failure cost. None of the three keys the tier to the tool's accuracy score. That is a genuinely strong convergence and it is the part of this brief to lean on.
- Contradiction with the question's own framing. The question treats one number as the independent variable ("at what stakes level does 83% become good enough"). Every framework surveyed treats stakes as the first variable and capability as the second, and the automation paper makes the direction explicit: required accuracy is elicited per task, then compared to delivered accuracy. Asking where 83% is good enough is asking the question backwards, and it also assumes an 83% that the vault has already retired.
- Contradiction on magnitude, still unresolved and now twice-noted. Snowflake implies un-grounded retrieval near 23%; dbt Labs measured raw text-to-SQL at 84.1-90%. Both vendor benchmarks, ~60 points apart, pointing opposite directions. Any threshold claim built on either is cherry-picked, which is the second reason no number appears in this brief.
- A gap, and it looks real rather than a search miss. Three searches surfaced tiering frameworks for models (SR 11-7) and for content (translation), and nothing that tiers the answer to a business question as the unit of analysis. No published work was found that measures how long a wrong analytics number survives inside an organization before detection, which is the variable the next section argues matters most.
Synthesis for RDCO
The first half of the question dissolves on inspection, and the dissolution is the finding. Three of the premise's four load-bearing words fail. It is not free (consumption-priced in AI Credits; Snowflake's own figure has the grounded query cheaper, not costless). It is not a floor (it is a ceiling measured on the single most favourable estate in existence for it - Snowflake's own instrumented internal analytics data with semantic views on hand - so a mid-market estate should be expected to land below it, not at it). It is not reliably 83 (the same benchmark publishes as 21/23/24.1/~25 to 83/86.3 across the vendor's own surfaces, with undisclosed n, undisclosed graders, and undisclosed correctness criteria). And as of the last GA/preview audit it is not purchasable. Inverting "ceiling" to "floor" is what made this question feel urgent: it silently assumed the client gets Snowflake's best case for nothing. Nothing in the evidence supports that, and a DSA who argues from a stakes threshold derived from it is arguing from a number the client can look up and dispute in the room.
The variable that actually governs cannibalization is detection cost, and the accuracy framing hides it. A wrong answer is only cheap when someone finds out. Snowflake's own material describes no refusal or confidence-surfacing behaviour, so the failure mode is a confidently wrong number with no flag. That makes the decisive question not "how often is it wrong" but "who would notice, how fast, and at what cost." This reframes the erosion risk in a way that is more defensible than the parent brief's version: a use case is genuinely cannibalizable only when consequence is low and detection is cheap - the analyst eyeballs the number, sees it is off, and moves on. Where detection is expensive or deferred (a metric that only reconciles at quarter-close, a definition that only surfaces as wrong when two departments compare decks), low stakes do not protect you, because silent error accumulates unpriced. That is the segment the free-tier narrative is worst at describing, and it is where an OI engagement is easiest to justify without quoting any percentage at all.
A correctness-criticality taxonomy for analytics and text-to-SQL use cases should run on five axes, not one stakes dial. Synthesising SR 11-7's materiality drivers, the translation ladder's assignment criteria, and the vault's own gate-placement rule: C1 consequence severity (cosmetic / internal decision input / external commitment / financial-statement or regulatory); C2 detection latency and cost (self-evidently wrong / caught at reconciliation / undetectable without a rebuild); C3 reversibility and blast radius (retractable before propagation, or not - the vault's existing gate rule keyed directly to this); C4 definitional ambiguity density (how many contested definitions the answer traverses: which revenue, whose fiscal calendar, what nets against what); C5 action coupling (a human reads the number, versus an agent acts on it unattended - SR 11-7's "degree of manual intervention," and the axis GARP flags as where the old framework strains). Then assurance scales with tier, in ITAAC three-part form per [[2026-06-11-high-reliability-acceptance-gates-agent-contracts]]: per use case, the commitment, the verification method, and the acceptance criterion. The non-obvious payoff of splitting the axes is that cannibalization risk and engagement value load onto different ones. Risk of being skipped is roughly C1 low and C2 low. Value of the governed layer is roughly C4 high and C5 automated - because a context-retrieval service can retrieve competing definitions but cannot adjudicate between two legitimate ones, which is authorship, not retrieval. A use case can therefore be low-stakes and still un-helped by Sense; it just will not be worth paying to fix. Those cells are the honest "do not sell into this" list, and they are narrower than a single stakes cut would suggest.
On the threshold itself, the calibrated answer is a shape, not a number, and I am reasoning by analogy here rather than from evidence about analytics. The translation market is the closest thing to a natural experiment, and its outcome was re-stratification with rate compression, not displacement of the paid tier: the free tier took the volume, standards bodies codified the middle (ISO 18587), and the priced expertise migrated from production toward post-editing and tier assignment itself. Transferring that to OI: expect the commoditized half to be context assembly (which [[2026-07-08-cortex-sense-semantic-layer-wedge-caf]] already concluded), expect the surviving paid work to concentrate in adjudicating definitions and proving correctness, and expect deciding which use case belongs in which tier to become a deliverable in its own right rather than a pre-sales judgment call. That transfer is an analogy between two markets with different buyers, different liability structures and different standards regimes, so treat it as a hypothesis that shapes the delivery method, not as a prediction with a date on it.
For the founder's actual seat - DSA and TAL, delivery and architecture, not sales engineering - the useful output is a method asset, not a talk-track. Two concrete moves follow. First, make correctness-criticality a typed field set (C1-C5) on the Build Manifest, with a verification method bound to each tier. That merges tiering with gate placement into one mechanism and completes the ITAAC three-part upgrade the June brief already recommended, which is a practice-level asset that scales across engagements rather than a per-deal argument. Second, make the Confidently-Wrong Probe from [[2026-09-20-cowork-plugin-accuracy-unmodeled-midmarket-data]] tier-aware: sample its question battery across C1-C5 cells rather than just "hard questions," so the deliverable is a per-tier error profile on the client's own estate instead of one alarming aggregate percentage. A per-tier profile is also the only instrument that could eventually answer this question empirically, because it measures the client's required accuracy and the platform's delivered accuracy in the same run.
Why the ledger cut is not performable, in order of severity. (1) The field probably does not exist. The ledger in question is the governed use-case portfolio and typed Build Manifest emitted by Assess (schema v1.1 plus validator, per [[2026-06-09-caf-restructure-proposal]] and the OI roadmap). The vault holds no record of any correctness-criticality attribute in that schema. If it is absent, there is nothing to cut - the fields have to be specified and back-filled first, which is a schema question, not a data pull. (2) n is far too small. As of the roadmap, two engagements had run end-to-end. Two cases cannot support a threshold claim regardless of how clean the fields are. (3) Access and boundary. The repo is not on this machine ([[2026-06-15-caf-architecture-decisions-and-meta-council]]), and phData engagement data must not be persisted to the vault in any case, so even a clean cut belongs inside phData and not here. The specific inputs that would make it performable: C1-C5 scored per use case with a written rubric; a recorded disposition per use case (built / deferred / declined) with the client's stated reason, since "we think the platform covers this" is the only direct observation of cannibalization and it lives in the sales conversation rather than in the manifest's technical fields; and engagements in the tens rather than the ones.
Why this is in the vault
This brief closes open follow-up #4 of [[2026-07-08-cortex-sense-semantic-layer-wedge-caf]] by rejecting its premise rather than answering it, and it stops a stakes-threshold argument built on the already-retired 83% figure from reaching an OI proposal or an Intelligence Maturity Assessment deliverable. It also supplies the concrete C1-C5 axis set that the Build Manifest schema and the Confidently-Wrong Probe would need in order to make the question empirically answerable later, and it records - so it is not re-asked - that the ledger cut is blocked on a missing schema field and n=2, not merely on repo access.
Open follow-ups
- Has anyone published a criticality taxonomy where the answer to a business question is the unit of analysis? Searches found tiering for models (SR 11-7) and for content (translation standards) but nothing for BI/analytics answers, and the absence looks structural rather than a search miss.
- Is there any measurement of detection cost for wrong analytics numbers - how long a wrong figure survives in an organization before someone catches it, by report type? This is axis C2 and nothing was found. It is the highest-value unknown in the framework.
- What is the non-vendor evidence on rate and composition change in translation after MT crossed "good enough"? The 26%-to-46% post-editing figure is vendor-cited; an industry survey or academic source would establish whether the re-stratification pattern is real and how much of the paid tier actually compressed.
- Has any supervisor (OCC, Bank of England SS1/23, ECB) issued tiering guidance that handles non-deterministic, silently wrong outputs? GARP names the strain and proposes nothing; if a regulator has published a tiering scheme for this, it is a ready-made backbone for C1-C5.
- Does the bundled-good-enough-tier pattern generalize beyond translation - tax software vs CPAs, robo-advice vs advisors, technology-assisted review vs document review - and in those markets did the adjacent paid service lose volume, lose price, or change composition? Three cases pointing the same way would upgrade the analogy above from hypothesis toward evidence.
Related
- [[2026-07-08-cortex-sense-semantic-layer-wedge-caf]] - the parent brief; source of this question and of the "free 83% floor" phrasing this brief corrects
- [[2026-09-20-cowork-plugin-accuracy-unmodeled-midmarket-data]] - the provenance audit that retired the 23-to-83 figure and proposed the Confidently-Wrong Probe
- [[2026-09-12-pupius-four-axis-snowflake-data-ai-strategy]] - AI Credit pricing, the unpublished decomposition of the accuracy gap, and the "all four bets hinge on one number" framing
- [[2026-09-14-snowflake-cowork-cortex-ga-vs-preview-matrix]] - the GA-vs-preview status that makes the free tier unpurchasable at contract date
- [[2026-06-11-high-reliability-acceptance-gates-agent-contracts]] - "scale verification to failure cost," gate placement by blast radius, and the ITAAC three-part contract format the C1-C5 ladder plugs into
- [[2026-04-07-dbt-semantic-layer-vs-text-to-sql-benchmark]] - the competing modeled-vs-unmodeled benchmark that contradicts Snowflake's implied floor
- [[2026-06-09-caf-restructure-proposal]] - the manifest-as-state-owner design that would carry the C1-C5 fields
- [[2026-06-15-caf-architecture-decisions-and-meta-council]] - records that the assessment repo is not on this machine
- [[2026-07-08-cowork-industry-plugins-vs-caf-delivery]] - sibling brief where the same accuracy sentence first propagated
Sources
Vault:
- [[2026-07-08-cortex-sense-semantic-layer-wedge-caf]] -
~/rdco-vault/06-reference/research/2026-07-08-cortex-sense-semantic-layer-wedge-caf.md - [[2026-09-20-cowork-plugin-accuracy-unmodeled-midmarket-data]] -
~/rdco-vault/06-reference/research/2026-09-20-cowork-plugin-accuracy-unmodeled-midmarket-data.md - [[2026-09-12-pupius-four-axis-snowflake-data-ai-strategy]] -
~/rdco-vault/06-reference/research/2026-09-12-pupius-four-axis-snowflake-data-ai-strategy.md - [[2026-06-11-high-reliability-acceptance-gates-agent-contracts]] -
~/rdco-vault/06-reference/research/2026-06-11-high-reliability-acceptance-gates-agent-contracts.md - [[2026-06-09-caf-restructure-proposal]] -
~/rdco-vault/01-projects/phdata/2026-06-09-caf-restructure-proposal.md - [[2026-06-15-caf-architecture-decisions-and-meta-council]] -
~/rdco-vault/01-projects/phdata/2026-06-15-caf-architecture-decisions-and-meta-council.md - OI roadmap (Build Manifest v1.1, two engagements end-to-end) -
~/rdco-vault/01-projects/phdata/roadmap-v1-work/caf-die-roadmap-v1.md - [[2026-09-14-snowflake-cowork-cortex-ga-vs-preview-matrix]] and [[2026-04-07-dbt-semantic-layer-vs-text-to-sql-benchmark]] - cited at second hand via the 2026-09-20 provenance audit, not re-read for this brief
Web:
- Li, Aboutorabi, Lyu, Qian, Fleming, Goehring, Thompson - "Economics of Human and AI Collaboration: When is Partial Automation More Attractive than Full Automation?" (arXiv:2603.29121v1 [econ.GN], 2026-03-31, preprint, not peer-reviewed): https://arxiv.org/html/2603.29121
- Tomedes - "Human translation, MTPE, or raw AI output: which quality tier does your content actually need?" (translation vendor, commercial interest in tier selection): https://www.tomedes.com/translator-hub/human-translation-vs-mtpe-vs-raw-ai
- Krishan Sharma (Citigroup Model Risk Management), GARP - "SR 11-7 in the Age of Agentic AI: Where the Framework Holds and Where It Strains" (2026-02-27): https://www.garp.org/risk-intelligence/operational/sr-11-7-age-agentic-ai-260227
- Umbrex - SR 11-7 Model Risk Management framework summary (consulting-network reference page, secondary): https://umbrex.com/resources/frameworks/regulatory-frameworks/model-risk-management-sr-11-7-framework/
- RiskTemplates - "AI Model Risk Tiering: How to Classify AI Models by Risk Level" (2026-03-31, vendor content, secondary): https://risktemplate.com/blog/2026-03-31-ai-model-risk-tiering-classification-guide/
Research caps used: 4 of 5 qmd queries, 3 of 3 WebSearch, 3 of 3 WebFetch, 6 vault docs. No paywalls encountered.