06-reference/research

cowork plugin accuracy unmodeled midmarket data

2026-09-20·research-brief·source: deep-research·by Ray Data Co (deep-research synthesis)
snowflakecoworkcortex-senseevidence-disciplineorganizational-intelligence

No One Has Measured a CoWork Plugin on Un-Modeled Mid-Market Data. The "~23% Floor" Is Snowflake's Own Baseline on Snowflake's Own Data.

The question

"What is the measured accuracy of a Snowflake CoWork finance/sales plugin on un-modeled mid-market data (the ~23% floor)? Quantifying the 'confidently wrong' risk turns it into a concrete sales artifact for the foundation-first CAF pitch."

Context: this was open follow-up #5 in [[2026-07-08-cowork-industry-plugins-vs-caf-delivery]], auto-promoted to the research queue on 2026-07-22. Naming note, once: "CAF" was retired as a program name on 2026-08-10. The work now runs under the Organizational Intelligence (OI) umbrella (Organizational Platform to Organizational Map, Intelligence Platform to Intelligence Maturity Assessment, plus OIP and Pulse). The rest of this brief uses the current naming.

Headline: no such measurement exists, and the ~23% is not what the question assumes it is. It is Snowflake's own reported baseline for a frontier coding agent reaching Snowflake over MCP, evaluated on Snowflake's internal product-analytics data. It is not a finance/sales plugin, not a mid-market estate, and not an un-modeled schema. Treat it as unverified for the use the question wants, and do not build a sales artifact on it.

What we already know (from the vault)

What the web says

Convergences and contradictions

Synthesis for RDCO

The provenance check fails, and that is the finding. The question asked for "the measured accuracy of a finance/sales plugin on un-modeled mid-market data" and treated "~23%" as that measurement. It is not. Snowflake's 23% (also published as 21%, 24.1% and ~25%) is the score of an external frontier coding agent reaching Snowflake through MCP, on Snowflake's internal product-analytics data, against an undisclosed question set with undisclosed grading. Three of the four nouns in the question are wrong: it is not a plugin, not un-modeled data, and not a mid-market estate. The one thing it does measure well is the value of Snowflake's own context layer versus a generic connector, which is exactly the comparison Snowflake wanted to publish. The vault should record the ~23% as derived and unverified for our use, not as a measured floor, and [[2026-07-08-cowork-industry-plugins-vs-caf-delivery]] should carry a correction pointer to this brief. This is the second time a number from that brief has needed narrowing; [[2026-09-14-snowflake-cowork-cortex-ga-vs-preview-matrix]] already narrowed the same sentence on GA grounds.

So the sales artifact does not get built in its proposed form. A one-pager titled "packaged plugins score 23% on your data" would be a fabricated claim wearing a citation. It would also be fragile in the room in three separate ways: the client can read the same Snowflake blog and see the number is about a coding agent, not a plugin; the client can note that Cortex Sense and the plugins are private preview so neither the floor nor the ceiling is purchasable today; and a competitor can produce dbt's 84-90% un-semantic-layer number and make the OI pitch look like it was shopping for the scariest statistic. The founder's calibration rule applies directly here: higher conviction than the evidence supports on a claim a client can check is the failure mode, not the win.

What we can claim in front of a client, stated precisely. Three things survive the evidence. First, a qualitative and vendor-admitted claim: Snowflake itself says the semantic view "stays the gold standard," that Cortex Sense "does not execute or validate queries, check whether answers are mathematically correct, or replace the need for semantic modeling," and its own posts describe no mechanism by which a wrong answer gets flagged. Second, a status claim, which is now the stronger card: as of 2026-09-14 the plugins and Cortex Sense are private preview at most, so "just turn on the plugin" is not an option a mid-market buyer has on their contract date, and the governed semantic views we build on GA objects become Sense's highest-authority input when it ships. Third, an architectural claim backed by two independent vendor benchmarks pointing the same qualitative direction: the data layer, not the model, determines answer quality, and the failure mode is a confident wrong number rather than an error. None of those three needs a percentage.

The better artifact is an instrument, not a statistic - and we can build it. The reason no one has published plugin-on-un-modeled-data accuracy is that the measurement is client-specific by construction: "un-modeled" is a property of their estate, and the answer is different for every buyer. That is not a research dead end, it is a product. The Intelligence Maturity Assessment should carry a Confidently-Wrong Probe: a short, fixed battery of ten to fifteen finance and sales questions with genuine definitional ambiguity (which revenue, whose fiscal calendar, what counts as a closed opportunity, how returns and credits net), run against the client's estate with whatever agent surface they have today, scored by their own controller or RevOps lead rather than by us. The deliverable is their number on their data, with the wrong answers printed next to the right ones. That converts an un-citable vendor statistic into a first-party measurement we own, it produces the emotional moment the sales artifact was reaching for (a CFO reading a confident wrong revenue figure from their own system), and it is a repeatable OIP instrument rather than a one-off slide. It also quietly generates the dataset nobody has: run across ten mid-market engagements, we would hold the only real distribution of un-modeled-estate agent accuracy in the market, which is a publishable asset and a Sanity Check piece with an original re-frame rather than a borrowed one.

Why this is in the vault

This brief retires a specific number that had already propagated into four vault briefs and was one step from a client-facing artifact, and it replaces the proposed artifact with a buildable Intelligence Maturity Assessment instrument (the Confidently-Wrong Probe). It directly governs what a phData DSA may and may not say about CoWork plugin accuracy in a proposal, and it is the correction pointer for the accuracy sentence in [[2026-07-08-cowork-industry-plugins-vs-caf-delivery]].

Open follow-ups

Related

Sources

Vault:

Web (all fetched or searched 2026-09-20):