06-reference/research

agent eval frameworks snowflake cortex

2026-09-06·research-brief·source: deep-research·by Ray Data Co (deep-research synthesis)

Agent eval frameworks for Cortex-native pipelines: Snowflake already shipped the runtime layer, so the open slot is the behavior spec

The question

"Which agent evaluation frameworks (LangSmith, RAGAS, DeepEval, Microsoft ASSERT) are production-ready for Snowflake/Cortex-native agentic pipelines, and which best fits phData's CAF catalog architecture?"

Context: the "CAF" name was retired 2026-08-10; the work continues under the Organizational Intelligence (OI) umbrella (Org Platform → Org Map · Intelligence Platform → IMA · OIP + Pulse). The question is really about the OI catalog: what evaluation layer should attach to catalog entries so their "sandbox-proven" status is computed rather than asserted.

What we already know (from the vault)

What the web says

Convergences and contradictions

Synthesis for RDCO

The question as posed has a false premise, and correcting it is the finding. LangSmith, RAGAS, DeepEval and ASSERT are not four candidates for one slot. Evaluation of a Cortex-native pipeline decomposes into three layers, and the four named tools sort cleanly across them once you notice that. Layer one is runtime telemetry and scoring inside the client's account, and Snowflake has already taken that slot: Cortex Agent evaluations plus AI Observability are partly GA as of the September 2026 docs, run server-side with SQL-queryable results, and can be scheduled in Tasks. Nothing in the third-party field beats "does not leave the governance boundary" for enterprise Snowflake delivery. Layer two is the behavior specification, the artifact that says what the agent is supposed to do before anyone writes a metric. Layer three is the CI gate that runs on RDCO's own skills, in RDCO's own repos, where no Snowflake account is involved.

The real gap is layer two, and it is the one none of the vault's existing material covers. The catalog's whole economic argument in [[2026-06-14-caf-restructure-organizing-brief]] rests on an entry carrying a "sandbox-proven" status. Today that status is asserted by whoever last touched the entry. A behavior spec per catalog entry turns it into a computed status with a date and a score, which is the difference between an actuarial table and a spreadsheet of opinions. ASSERT is the closest shipped match to that shape: it starts from a policy written in YAML, compiles it into executable tests, scores with judge rationales, and traces failures over OpenTelemetry. Critically it is MIT-licensed and framework-agnostic, so it composes with the Snowflake layer instead of competing with it. The behavior spec is the durable catalog artifact; the tests it generates can execute against a Cortex Agent and land their telemetry in the account's observability tables. Two hard caveats before this goes anywhere near a client: ASSERT is roughly three months old as of this writing, its quality evidence is entirely Microsoft first-party, and its origin will read as an Azure play in a Snowflake room. Pilot it internally on an RDCO skill first, and if it does not survive the pilot, the fallback is not another vendor - it is writing the behavior-spec schema ourselves and generating cases with the brigade's existing spec-author station, which is arguably where this was going anyway.

For layer three, DeepEval, and the reason is structural rather than featural. The Notion backlog row that generated this question named the actual pain: the Markov tracker and pipeline-code-author skills have no behavior spec and no eval layer, so quality regressions are invisible. DeepEval is Apache 2.0 and pytest-shaped, which means it drops into the same GitHub Actions surface the brigade already uses and maps one-to-one onto station-critic's per-axis fan-out with no new SaaS dependency and no new spend under the RDCO autonomous-spend threshold. RAGAS is worth importing as a metric library where retrieval faithfulness genuinely matters, not adopted as a platform. LangSmith should be read for method and not bought: the ideal-trajectory efficiency ratios from [[2026-04-13-langchain-evals-deep-agents]] are the single most reusable idea in that article, they are tool-agnostic, and they answer a question Snowflake's metric set does not - not "was the answer right" but "did it cost five times what it should have," which is exactly the number a fixed-bid catalog entry needs to carry.

The commercial read. [[2026-07-25-evals-as-competitive-moat]] argued the eval score should ship as a deliverable artifact and flagged that the format contract does not exist. This brief supplies the missing half of that: the format should be behavior-spec-plus-score, the score should carry method, N and agreement statistic, and the agreement bar should be Airbnb's high-80s kappa rather than an internal vibe. A catalog entry that ships as "recipe + behavior spec + last eval score + judge agreement" is a materially different sales object from one that ships as "recipe + we tried it once." That is a positioning asset for the OI catalog, and it is buildable now, because the runtime layer it depends on is already GA in the platform the clients are already on.

Why this is in the vault

This closes the eval-layer gap named in the Research Backlog row against two specific live surfaces: the OI catalog entry schema (which currently asserts "sandbox-proven" with no computed backing) and the RDCO skills that [[2026-07-25-evals-as-competitive-moat]] said should ship with lift scores but do not - the Markov tracker and pipeline-code-author. It also supplies the format contract that concept note explicitly left unresolved.

Open follow-ups

Related

Sources

Vault:

Web (all accessed 2026-09-06):