No One Has Measured a CoWork Plugin on Un-Modeled Mid-Market Data. The "~23% Floor" Is Snowflake's Own Baseline on Snowflake's Own Data.
The question
"What is the measured accuracy of a Snowflake CoWork finance/sales plugin on un-modeled mid-market data (the ~23% floor)? Quantifying the 'confidently wrong' risk turns it into a concrete sales artifact for the foundation-first CAF pitch."
Context: this was open follow-up #5 in [[2026-07-08-cowork-industry-plugins-vs-caf-delivery]], auto-promoted to the research queue on 2026-07-22. Naming note, once: "CAF" was retired as a program name on 2026-08-10. The work now runs under the Organizational Intelligence (OI) umbrella (Organizational Platform to Organizational Map, Intelligence Platform to Intelligence Maturity Assessment, plus OIP and Pulse). The rest of this brief uses the current naming.
Headline: no such measurement exists, and the ~23% is not what the question assumes it is. It is Snowflake's own reported baseline for a frontier coding agent reaching Snowflake over MCP, evaluated on Snowflake's internal product-analytics data. It is not a finance/sales plugin, not a mid-market estate, and not an un-modeled schema. Treat it as unverified for the use the question wants, and do not build a sales artifact on it.
What we already know (from the vault)
- The number entered the vault as a secondary citation, not a measurement. [[2026-07-08-cowork-industry-plugins-vs-caf-delivery]] wrote "a frontier agent over Snowflake MCP alone hits 23% accuracy on complex enterprise queries; with Cortex Sense context it hits 83%," sourced to the Snowflake CoWork blog. The sibling brief [[2026-07-08-cortex-sense-semantic-layer-wedge-caf]] carried a three-point version (about 23-24 percent general agent, about 47 percent CoWork/CoCo without Sense, about 83 percent with Sense) and already flagged Snowflake's own caveat that "context is not correctness."
- The vault has already warned once against quoting this line to a client. [[2026-09-14-snowflake-cowork-cortex-ga-vs-preview-matrix]] states: "A DSA who uses the 23%-to-83% accuracy line in a proposal is quoting a private-preview (or not-yet-private) capability." Cortex Sense was private preview at most as of 2026-09-14, so the 83% describes a future state, not a shipping product.
- The vault has already identified the unpublished split that matters. [[2026-09-12-pupius-four-axis-snowflake-data-ai-strategy]] notes that part of the 23-to-83 gap is permissions, lineage and business definitions (organizational knowledge a better model cannot infer) and part is inferable from schema and query history, and that "neither Snowflake nor anyone else publishes that split."
- The one vault-held benchmark with a real modeled-vs-unmodeled delta disagrees sharply. [[2026-04-07-dbt-semantic-layer-vs-text-to-sql-benchmark]]: on an 11-question ACME Insurance set, semantic layer scored 98.2-100%, raw text-to-SQL scored 84.1-90%. That is a 10-15 point gap, not a 60 point one, and it is the closest thing the vault holds to "how bad is un-modeled."
- The plugins themselves are a Cortex Sense capability in private preview. Same parent brief. There is no GA product to measure, which is the structural reason no measurement exists.
What the web says
- Snowflake's own Cortex Sense post reports a different pair of numbers than the one the vault carries, and names the dataset. It states Cortex Sense "improved accuracy from 24.1% to 86.3%," comparing "a frontier coding agent with direct access to SQL execution via model context protocol (MCP), vanilla CoCo (which had the ability to retrieve relevant semantic views if it needed) and CoCo grounded by Cortex Sense." The evaluation set is described as "product analytics questions on Snowflake's internal data requiring cross-table joins, metric formula lookups and filter convention knowledge." The post does not disclose query count, who authored the gold answers, or what counts as correct. It also cites a ~25% internal baseline and a ~21% figure it attributes to Anthropic (Snowflake, Cortex Sense for Enterprise AI Agents).
- So the "23%" is a rounded, chart-level restatement of a floating baseline. Across Snowflake's own surfaces the same baseline appears as 21%, 23%, 24.1% and ~25%, and the ceiling as 83% or 86.3%. A number that moves 4 points depending on which vendor page you read is not a measurement you put in a client deck.
- The baseline is a coding agent over MCP, which is the opposite of an un-modeled-data test. The comparison isolates the value of Snowflake's context-retrieval layer against an external agent reaching in through a connector. It holds the data estate constant (Snowflake's own, presumably well-instrumented internal analytics data) and varies the agent harness. The question assumes the opposite design: hold the harness constant and vary the data quality. No published Snowflake number does that.
- No finance or sales plugin has been measured separately, by Snowflake or anyone else. The plugins are described as bundles of "skills, business logic and MCP connectors" for finance and sales, in private preview. Searches surfaced no third-party evaluation of a packaged CoWork domain plugin, and no benchmark of any CoWork configuration against un-curated mid-market schemas.
- Snowflake's most methodologically transparent accuracy post covers Cortex Analyst, not plugins, and explicitly scopes itself away from messy schemas. It claims "90%+ SQL accuracy on real-world use cases" against "an internal benchmark suite of 150 questions" where "human evaluators manually select all the correct SQL queries," and states "our initial focus was on ensuring high accuracy across a wide variety of customer use cases on a single view with pre-joined data." It publishes no with-model versus without-model delta, no un-curated-schema number, and describes no refusal or uncertainty-flagging behavior (Snowflake engineering blog, Cortex Analyst text-to-SQL accuracy).
- The nearest independent-looking artifact measures something adjacent and scores low. AvalancheBench evaluates whether data agents recover "the segments, drivers, temporal events, and relationships that explain the data" in a controlled synthetic setting, and reports that "the strongest configuration of a leading coding agent recovers only 26% of the rubric." It names no vendor product, does not evaluate Cortex Sense, Cortex Analyst, CoWork or any plugin, and does not contrast modeled against un-modeled schemas. Several of its authors are Snowflake-affiliated, so it is not a clean third party (arXiv 2605.24183).
- Vendor self-benchmarking on vendor-built harnesses is a documented reliability problem, and an analyst has already flagged it for this exact vendor. HFS Research notes that no independent replication of Snowflake's ADE-Bench run had been published, that "a governance-native agent evaluated on the same harness that provides its claimed advantage is structurally positioned to score higher," and cites research finding vendor-run scores systematically exceed independently replicated ones (HFS Research).
Convergences and contradictions
- Contradiction, and it is the load-bearing one: the vault's own two benchmarks disagree by roughly 60 points about how bad un-modeled retrieval is. Snowflake implies a floor near 23%. dbt Labs measured raw text-to-SQL at 84.1-90% on a real (if small) insurance schema. Both are vendor benchmarks with an interest in the answer, pointing in opposite directions. Any artifact that quotes one and not the other is cherry-picked.
- Convergence on the qualitative claim, and only the qualitative claim. Snowflake, dbt Labs, typedef.ai and the vault all agree that governed semantic definitions raise answer quality and that the failure mode is silent rather than flagged. Snowflake's own posts describe no refusal or confidence-surfacing behavior, and dbt's benchmark is explicit that "wrong answers look plausible." That is the defensible claim. The magnitude is not.
- Convergence on the absence. Three separate searches and the most detailed vendor accuracy posts all fail to produce a plugin-specific number, an un-modeled-schema number, or a mid-market estate. This looks like a genuine gap rather than a search miss, and it has a structural cause: the plugins are private preview, so there is no GA artifact for anyone to benchmark.
Synthesis for RDCO
The provenance check fails, and that is the finding. The question asked for "the measured accuracy of a finance/sales plugin on un-modeled mid-market data" and treated "~23%" as that measurement. It is not. Snowflake's 23% (also published as 21%, 24.1% and ~25%) is the score of an external frontier coding agent reaching Snowflake through MCP, on Snowflake's internal product-analytics data, against an undisclosed question set with undisclosed grading. Three of the four nouns in the question are wrong: it is not a plugin, not un-modeled data, and not a mid-market estate. The one thing it does measure well is the value of Snowflake's own context layer versus a generic connector, which is exactly the comparison Snowflake wanted to publish. The vault should record the ~23% as derived and unverified for our use, not as a measured floor, and [[2026-07-08-cowork-industry-plugins-vs-caf-delivery]] should carry a correction pointer to this brief. This is the second time a number from that brief has needed narrowing; [[2026-09-14-snowflake-cowork-cortex-ga-vs-preview-matrix]] already narrowed the same sentence on GA grounds.
So the sales artifact does not get built in its proposed form. A one-pager titled "packaged plugins score 23% on your data" would be a fabricated claim wearing a citation. It would also be fragile in the room in three separate ways: the client can read the same Snowflake blog and see the number is about a coding agent, not a plugin; the client can note that Cortex Sense and the plugins are private preview so neither the floor nor the ceiling is purchasable today; and a competitor can produce dbt's 84-90% un-semantic-layer number and make the OI pitch look like it was shopping for the scariest statistic. The founder's calibration rule applies directly here: higher conviction than the evidence supports on a claim a client can check is the failure mode, not the win.
What we can claim in front of a client, stated precisely. Three things survive the evidence. First, a qualitative and vendor-admitted claim: Snowflake itself says the semantic view "stays the gold standard," that Cortex Sense "does not execute or validate queries, check whether answers are mathematically correct, or replace the need for semantic modeling," and its own posts describe no mechanism by which a wrong answer gets flagged. Second, a status claim, which is now the stronger card: as of 2026-09-14 the plugins and Cortex Sense are private preview at most, so "just turn on the plugin" is not an option a mid-market buyer has on their contract date, and the governed semantic views we build on GA objects become Sense's highest-authority input when it ships. Third, an architectural claim backed by two independent vendor benchmarks pointing the same qualitative direction: the data layer, not the model, determines answer quality, and the failure mode is a confident wrong number rather than an error. None of those three needs a percentage.
The better artifact is an instrument, not a statistic - and we can build it. The reason no one has published plugin-on-un-modeled-data accuracy is that the measurement is client-specific by construction: "un-modeled" is a property of their estate, and the answer is different for every buyer. That is not a research dead end, it is a product. The Intelligence Maturity Assessment should carry a Confidently-Wrong Probe: a short, fixed battery of ten to fifteen finance and sales questions with genuine definitional ambiguity (which revenue, whose fiscal calendar, what counts as a closed opportunity, how returns and credits net), run against the client's estate with whatever agent surface they have today, scored by their own controller or RevOps lead rather than by us. The deliverable is their number on their data, with the wrong answers printed next to the right ones. That converts an un-citable vendor statistic into a first-party measurement we own, it produces the emotional moment the sales artifact was reaching for (a CFO reading a confident wrong revenue figure from their own system), and it is a repeatable OIP instrument rather than a one-off slide. It also quietly generates the dataset nobody has: run across ten mid-market engagements, we would hold the only real distribution of un-modeled-estate agent accuracy in the market, which is a publishable asset and a Sanity Check piece with an original re-frame rather than a borrowed one.
Why this is in the vault
This brief retires a specific number that had already propagated into four vault briefs and was one step from a client-facing artifact, and it replaces the proposed artifact with a buildable Intelligence Maturity Assessment instrument (the Confidently-Wrong Probe). It directly governs what a phData DSA may and may not say about CoWork plugin accuracy in a proposal, and it is the correction pointer for the accuracy sentence in [[2026-07-08-cowork-industry-plugins-vs-caf-delivery]].
Open follow-ups
- What does the Confidently-Wrong Probe actually contain? Draft the ten to fifteen question battery and the scoring rubric (correct / wrong / refused / ambiguous-but-unflagged), and decide whether it runs pre-sales as a free assessment hook or inside a paid Intelligence Maturity Assessment.
- Does the ~23-to-83 gap decompose into organizational knowledge versus schema-inferable signal, and in what proportion? [[2026-09-12-pupius-four-axis-snowflake-data-ai-strategy]] named this as the most useful unknown and it is still unpublished. Our own probe data would eventually answer it.
- Has any third party replicated any Snowflake agent benchmark since HFS flagged the absence? A single independent replication would change how much weight the qualitative claim can carry.
- When the finance/sales plugins reach public preview, what does Snowflake publish about their evaluation, if anything? Set a watch on the Cortex Sense release notes rather than re-running this question blind.
- Should the vault get a standing convention for numbers inherited across briefs (a "derived, unverified at source" tag) so a chart-level statistic cannot reach a client artifact without one gate hitting the primary source?
Related
- [[2026-07-08-cowork-industry-plugins-vs-caf-delivery]] - the parent brief; source of both the question and the ~23% figure this brief narrows
- [[2026-07-08-cortex-sense-semantic-layer-wedge-caf]] - origin of the 23/47/83 three-point version and the "context is not correctness" caveat
- [[2026-09-14-snowflake-cowork-cortex-ga-vs-preview-matrix]] - the GA-versus-preview narrowing of the same sentence; plugins and Cortex Sense are not GA
- [[2026-09-12-pupius-four-axis-snowflake-data-ai-strategy]] - the unpublished decomposition of the accuracy gap
- [[2026-04-07-dbt-semantic-layer-vs-text-to-sql-benchmark]] - the contradicting vendor benchmark; raw text-to-SQL at 84.1-90%, semantic layer at 98.2-100%
- [[2026-09-19-cowork-group-plugin-assignment-enforcement]] - adjacent September CoWork plugin research; admin distribution mechanics
- [[2026-09-15-snowflake-intelligence-vs-cowork-naming-standard]] - which product name to use with clients in H2 2026
- [[2026-05-20-phdata-cortex-agents-practice]] - the golden-set / TruLens eval spine the Confidently-Wrong Probe would extend
- [[2026-06-03-semantic-layer-validation-controls-rdco]] - definition versus safe-computation contract, the mechanism behind the qualitative claim
Sources
Vault:
- [[2026-07-08-cowork-industry-plugins-vs-caf-delivery]] -
~/rdco-vault/06-reference/research/2026-07-08-cowork-industry-plugins-vs-caf-delivery.md - [[2026-07-08-cortex-sense-semantic-layer-wedge-caf]] -
~/rdco-vault/06-reference/research/2026-07-08-cortex-sense-semantic-layer-wedge-caf.md - [[2026-09-14-snowflake-cowork-cortex-ga-vs-preview-matrix]] -
~/rdco-vault/06-reference/research/2026-09-14-snowflake-cowork-cortex-ga-vs-preview-matrix.md - [[2026-09-12-pupius-four-axis-snowflake-data-ai-strategy]] -
~/rdco-vault/06-reference/research/2026-09-12-pupius-four-axis-snowflake-data-ai-strategy.md - [[2026-04-07-dbt-semantic-layer-vs-text-to-sql-benchmark]] -
~/rdco-vault/06-reference/2026-04-07-dbt-semantic-layer-vs-text-to-sql-benchmark.md - [[2026-09-19-cowork-group-plugin-assignment-enforcement]] -
~/rdco-vault/06-reference/research/2026-09-19-cowork-group-plugin-assignment-enforcement.md - [[2026-09-15-snowflake-intelligence-vs-cowork-naming-standard]] -
~/rdco-vault/06-reference/research/2026-09-15-snowflake-intelligence-vs-cowork-naming-standard.md - [[2026-05-20-phdata-cortex-agents-practice]] -
~/rdco-vault/06-reference/research/2026-05-20-phdata-cortex-agents-practice.md - [[2026-06-03-semantic-layer-validation-controls-rdco]] -
~/rdco-vault/08-tooling/2026-06-03-semantic-layer-validation-controls-rdco.md
Web (all fetched or searched 2026-09-20):
- Snowflake, "Cortex Sense for Enterprise AI Agents" (24.1% to 86.3%; Snowflake-internal product-analytics questions; no query count or grading disclosed): https://www.snowflake.com/en/blog/enterprise-ai-agents-grounded-context/
- Snowflake engineering blog, Cortex Analyst text-to-SQL accuracy (90%+ on a 150-question internal suite; scoped to a single pre-joined view; no un-curated-schema number): https://www.snowflake.com/en/engineering-blog/cortex-analyst-text-to-sql-accuracy-bi/
- Snowflake CoWork blog (origin of the 23/47/83 chart-level restatement; plugins private preview soon) - referenced via the parent brief and search results, not re-fetched: https://www.snowflake.com/en/blog/snowflake-cowork-personal-work-agent/
- arXiv 2605.24183, "AvalancheBench: Evaluating Enterprise Data Agents Through Latent World Recovery" (26% rubric recovery; synthetic controlled setting; Snowflake-affiliated authors; no plugin or un-modeled-schema contrast): https://arxiv.org/abs/2605.24183
- HFS Research, "Snowflake Agentic AI Beats Claude Code on Its Own Benchmark" (no independent replication; harness-advantage critique; vendor-versus-replicated score gap): https://www.hfsresearch.com/news/snowflake-agentic-ai-beats-claude-code-on-its-own-benchmark-what-that-means/
- typedef.ai, "What Is Cortex Sense? Snowflake's Runtime Context Layer, Explained" (source of the 23/47/83 three-point version and "context is not correctness"; cited here via the vault brief that fetched it): https://www.typedef.ai/blog/what-is-cortex-sense-snowflakes-runtime-context-layer-explained