Agent eval frameworks for Cortex-native pipelines: Snowflake already shipped the runtime layer, so the open slot is the behavior spec
The question
"Which agent evaluation frameworks (LangSmith, RAGAS, DeepEval, Microsoft ASSERT) are production-ready for Snowflake/Cortex-native agentic pipelines, and which best fits phData's CAF catalog architecture?"
Context: the "CAF" name was retired 2026-08-10; the work continues under the Organizational Intelligence (OI) umbrella (Org Platform → Org Map · Intelligence Platform → IMA · OIP + Pulse). The question is really about the OI catalog: what evaluation layer should attach to catalog entries so their "sandbox-proven" status is computed rather than asserted.
What we already know (from the vault)
- The catalog is the asset, and it is the fixed-bid actuarial table. [[2026-06-14-caf-restructure-organizing-brief]] centers the whole architecture on a use-case/agent catalog where each entry carries proven deployment recipes, sandbox-proven status, and actual delivery effort. Fixed bid is only profitable when scope, cost, and reuse are all knowable, and all three are downstream of that one object. An entry whose quality can silently regress between engagements is an actuarial table with a hole in it.
- Evals are already framed as the commercial differentiator, not a QA preference. [[2026-07-25-evals-as-competitive-moat]] argues the eval score should ship as a first-class artifact alongside every skill, and names the unresolved gap: the brigade's eval station emits pass/fail, not a lift score with method, N, and confidence interval.
- The vault's only framework-specific eval source is vendor-authored. [[2026-04-13-langchain-evals-deep-agents]] (LangChain, published 2026-03-26) is genuinely useful on method (four-source eval pipeline; ideal-trajectory step / tool-call / latency ratios; pytest on GitHub Actions), and the note already flags the bias: every recommendation routes through LangSmith.
- The calibration bar is already set from a non-vendor source. [[2026-08-03-airbnb-eval-driven-development]] gives hard numbers we can reuse: 50-100 example golden set seeded with bad examples, Cohen's kappa / Krippendorff's alpha agreement in the high 80s-90s, one judge per dimension (explicitly no "God evaluator"), trajectory evals scoring intermediate tool calls, and the same judges running in production on a ~5% traffic sample.
- The OI method sits above the Snowflake stack, not inside it. [[2026-07-07-snowflake-intelligence-vs-cortex-ai-boundary]] places the assessment method above the whole stack: it classifies use cases, routes by autonomy, and emits a build manifest deciding which Cortex Agents get built. That positioning is what makes an eval layer attachable at the catalog level rather than per-project.
What the web says
- Snowflake shipped native Cortex Agent evaluations, and they are partly GA. Per Snowflake docs (fetched 2026-09-06): ground-truth metrics are Answer Correctness (GA), Tool Selection Accuracy (Public Preview) and Tool Execution Accuracy (Public Preview); the reference-free metric is Logical Consistency (GA); custom LLM-judged metrics are supported via prompt plus scoring rubric. Runs are configured in YAML and executed from Snowsight, from SQL via
EXECUTE_AI_EVALUATION, or from the Cortex Code CLI. Results are queryable withGET_AI_EVALUATION_DATAandGET_AI_RECORD_TRACE, and runs can be scheduled in Tasks for CI/CD. - The Cortex eval limitations are specific and load-bearing. Same doc, same date: no MCP server tool support, code-execution tool outputs are not persisted during evaluation, no session attribute or row-access-policy support, and ground-truth staleness on time-relative queries. LLM judges run via cross-region inference on AWS and Azure regions but not Google Cloud. The v1 default judge model
claude-4-sonnetentered legacy state 2026-08-12, so accounts must pin v2 or v3 metrics. - AI Observability is the surrounding telemetry layer, split native vs custom. Snowflake docs (fetched 2026-09-06): Cortex Agents and CoWork expose live thread/trace monitoring plus batch evaluations directly in Snowsight, with events logged to
SNOWFLAKE.LOCAL.AI_OBSERVABILITY_EVENTS; Cortex Analyst and Cortex Search have their own request telemetry. For custom applications you own end to end (a standalone agent, a graph of Cortex Agents, a RAG pipeline combining Cortex Search with AI_COMPLETE), the documented path is streaming telemetry in via TruLens. - Microsoft ASSERT is a behavior-spec compiler, MIT-licensed, and roughly three months old. Announced at Build 2026, reported 2026-06-03: the workflow is behavior specification → generated test cases → execution against the agent → scoring with judge rationales → trace-based debugging, with YAML policy config and OpenTelemetry tracing. It integrates with LangChain, CrewAI, LiteLLM, OpenAI and 100+ model endpoints, and explicitly does not require moving the application into Microsoft Foundry. Maturity caveat, stated in the same report: the 80-90% judge/human agreement figure is Microsoft first-party evidence, and the "1.2x intended behavior space" coverage claim is preliminary.
- Azure Foundry's own agent evaluators are the platform-locked sibling. Microsoft Learn (accessed 2026-09-06) documents Intent Resolution, Task Adherence and Tool Call Accuracy as built-in agent evaluators that behave like unit tests, emitting binary pass/fail or thresholded scores. Useful as a metric taxonomy to steal; not a candidate for a Snowflake-resident pipeline.
- The 2026 comparison consensus is "two tools, not one." Across the framework round-ups surfaced 2026-09-06 (Morph, MLflow, Confident AI, DeepEval's own comparisons): DeepEval is Apache 2.0, pytest-style, 50+ metrics, built for code-first CI gating with Confident AI as the hosted collaboration layer; RAGAS is open source, RAG-origin, now carrying agent metrics including goal accuracy and tool-call accuracy, and is best used as a lightweight metric library; LangSmith is a hosted SaaS whose distinguishing property is automatic tracing of every LangChain and LangGraph run from a single env var. The recurring recommendation is to pair a lightweight CI gate (DeepEval, RAGAS, Promptfoo) with a platform for annotation, regression tracking and dashboards (LangSmith, Braintrust, Arize). The round-ups also report that offline suites alone under-deliver, and production needs a per-turn layer.
- Caveat on the comparison sources. Several of the round-ups are published by vendors comparing themselves (DeepEval/Confident AI in particular). Treat the per-framework capability lists as directionally right and the rankings as marketing.
Convergences and contradictions
- Convergence: the vault called the shape before the market shipped it. The Airbnb protocol (one judge per dimension, trajectory evals over intermediate tool calls) and Snowflake's metric set (Logical Consistency across instructions/planning/tool calls, Tool Selection and Tool Execution Accuracy as separate metrics) are the same design arrived at independently. The station-critic per-axis fan-out is the same shape a third time. This is a good sign that the axis decomposition is right, not a coincidence worth re-litigating.
- Contradiction: the "pick a framework" framing is wrong for Snowflake-resident work. The round-ups assume the eval harness is a free choice, because they assume traces can leave for a SaaS. In a client's Snowflake account, trace egress is a security-review question, not a preference. Snowflake's native path keeps evaluation data inside the governance boundary and queryable in SQL, which removes LangSmith from serious contention for client-resident pipelines regardless of how good its tracing is. This is a constraint the web sources do not model.
- Contradiction to hold open: Snowflake's native eval cannot see MCP tools. The documented "no MCP server tool support" limitation collides directly with where agent tooling is heading. Any catalog entry whose recipe involves MCP-served tools is currently un-evaluable by the native path, and that is not a small carve-out.
Synthesis for RDCO
The question as posed has a false premise, and correcting it is the finding. LangSmith, RAGAS, DeepEval and ASSERT are not four candidates for one slot. Evaluation of a Cortex-native pipeline decomposes into three layers, and the four named tools sort cleanly across them once you notice that. Layer one is runtime telemetry and scoring inside the client's account, and Snowflake has already taken that slot: Cortex Agent evaluations plus AI Observability are partly GA as of the September 2026 docs, run server-side with SQL-queryable results, and can be scheduled in Tasks. Nothing in the third-party field beats "does not leave the governance boundary" for enterprise Snowflake delivery. Layer two is the behavior specification, the artifact that says what the agent is supposed to do before anyone writes a metric. Layer three is the CI gate that runs on RDCO's own skills, in RDCO's own repos, where no Snowflake account is involved.
The real gap is layer two, and it is the one none of the vault's existing material covers. The catalog's whole economic argument in [[2026-06-14-caf-restructure-organizing-brief]] rests on an entry carrying a "sandbox-proven" status. Today that status is asserted by whoever last touched the entry. A behavior spec per catalog entry turns it into a computed status with a date and a score, which is the difference between an actuarial table and a spreadsheet of opinions. ASSERT is the closest shipped match to that shape: it starts from a policy written in YAML, compiles it into executable tests, scores with judge rationales, and traces failures over OpenTelemetry. Critically it is MIT-licensed and framework-agnostic, so it composes with the Snowflake layer instead of competing with it. The behavior spec is the durable catalog artifact; the tests it generates can execute against a Cortex Agent and land their telemetry in the account's observability tables. Two hard caveats before this goes anywhere near a client: ASSERT is roughly three months old as of this writing, its quality evidence is entirely Microsoft first-party, and its origin will read as an Azure play in a Snowflake room. Pilot it internally on an RDCO skill first, and if it does not survive the pilot, the fallback is not another vendor - it is writing the behavior-spec schema ourselves and generating cases with the brigade's existing spec-author station, which is arguably where this was going anyway.
For layer three, DeepEval, and the reason is structural rather than featural. The Notion backlog row that generated this question named the actual pain: the Markov tracker and pipeline-code-author skills have no behavior spec and no eval layer, so quality regressions are invisible. DeepEval is Apache 2.0 and pytest-shaped, which means it drops into the same GitHub Actions surface the brigade already uses and maps one-to-one onto station-critic's per-axis fan-out with no new SaaS dependency and no new spend under the RDCO autonomous-spend threshold. RAGAS is worth importing as a metric library where retrieval faithfulness genuinely matters, not adopted as a platform. LangSmith should be read for method and not bought: the ideal-trajectory efficiency ratios from [[2026-04-13-langchain-evals-deep-agents]] are the single most reusable idea in that article, they are tool-agnostic, and they answer a question Snowflake's metric set does not - not "was the answer right" but "did it cost five times what it should have," which is exactly the number a fixed-bid catalog entry needs to carry.
The commercial read. [[2026-07-25-evals-as-competitive-moat]] argued the eval score should ship as a deliverable artifact and flagged that the format contract does not exist. This brief supplies the missing half of that: the format should be behavior-spec-plus-score, the score should carry method, N and agreement statistic, and the agreement bar should be Airbnb's high-80s kappa rather than an internal vibe. A catalog entry that ships as "recipe + behavior spec + last eval score + judge agreement" is a materially different sales object from one that ships as "recipe + we tried it once." That is a positioning asset for the OI catalog, and it is buildable now, because the runtime layer it depends on is already GA in the platform the clients are already on.
Why this is in the vault
This closes the eval-layer gap named in the Research Backlog row against two specific live surfaces: the OI catalog entry schema (which currently asserts "sandbox-proven" with no computed backing) and the RDCO skills that [[2026-07-25-evals-as-competitive-moat]] said should ship with lift scores but do not - the Markov tracker and pipeline-code-author. It also supplies the format contract that concept note explicitly left unresolved.
Open follow-ups
- Does ASSERT's YAML policy schema survive contact with a real catalog entry, or does the OI catalog need its own behavior-spec schema? One internal pilot on an existing RDCO skill answers this.
- What is the practical cost of a Cortex Agent evaluation run? The docs describe judge invocation via cross-region inference but no pricing model surfaced; fixed-bid scoping needs this number before evals go into a delivery estimate.
- How do we evaluate catalog entries whose recipes use MCP-served tools, given Snowflake's documented "no MCP server tool support" limitation in Cortex Agent evaluations?
- Can ASSERT (or any OTel-emitting harness) write into
SNOWFLAKE.LOCAL.AI_OBSERVABILITY_EVENTS, or does the account boundary force two separate trace stores and a reconciliation problem? - What is the judge-drift policy? The
claude-4-sonnetlegacy transition on 2026-08-12 means judge models move underneath scored results; a catalog entry's eval score needs a judge-version stamp or it silently expires. - Do the LangChain efficiency ratios (step / tool-call / latency vs ideal trajectory) reproduce as a Cortex custom metric, and would that number be defensible enough to price a fixed bid against?
Related
- [[2026-06-14-caf-restructure-organizing-brief]] — the catalog-as-actuarial-table reframe this brief attaches an eval layer to
- [[2026-07-25-evals-as-competitive-moat]] — the positioning argument; this brief supplies its missing format contract
- [[2026-04-13-langchain-evals-deep-agents]] — the vault's existing LangSmith-adjacent eval methodology, vendor bias flagged
- [[2026-08-03-airbnb-eval-driven-development]] — non-vendor calibration protocol and the kappa bar
- [[2026-07-07-snowflake-intelligence-vs-cortex-ai-boundary]] — where the OI method sits relative to the Cortex stack
- [[2026-09-03-etom-lifecycle-first-org-map-spine]] — current Org Map spine work under the OI umbrella
- [[2026-05-10-neural-avb-design-experiments-evaluate-agentic-harness]] — eval harness as RL environment; the layer-three framing
- [[2026-06-04-agent-workflow-patterns-catalog]] — evaluator-optimizer pattern and the brigade station shape
Sources
Vault:
~/rdco-vault/01-projects/phdata/2026-06-14-caf-restructure-organizing-brief.md~/rdco-vault/06-reference/concepts/2026-07-25-evals-as-competitive-moat.md~/rdco-vault/06-reference/2026-04-13-langchain-evals-deep-agents.md~/rdco-vault/06-reference/2026-08-03-airbnb-eval-driven-development.md~/rdco-vault/06-reference/research/2026-07-07-snowflake-intelligence-vs-cortex-ai-boundary.md~/rdco-vault/06-reference/research/2026-09-03-etom-lifecycle-first-org-map-spine.md~/rdco-vault/06-reference/2026-05-10-neural-avb-design-experiments-evaluate-agentic-harness.md~/rdco-vault/06-reference/2026-06-04-agent-workflow-patterns-catalog.md
Web (all accessed 2026-09-06):
- Snowflake Documentation, "Cortex Agent evaluations" — https://docs.snowflake.com/en/user-guide/snowflake-cortex/cortex-agents-evaluations (fetched; metric list, GA/preview split, privileges, limitations)
- Snowflake Documentation, "AI Observability with Snowflake Cortex" — https://docs.snowflake.com/en/user-guide/snowflake-cortex/ai-observability (fetched; native vs TruLens split, event tables. No publication date visible on the page.)
- WinBuzzer, "Microsoft ASSERT Framework Turns AI-Agent Policies Into Executable Tests," 2026-06-03 — https://winbuzzer.com/2026/06/03/microsoft-assert-turns-ai-agent-policies-into-tests-xcxwbn/ (fetched; secondary reporting on a Build 2026 announcement, not a primary Microsoft source)
- Microsoft Learn, "Agent Evaluators for Generative AI" — https://learn.microsoft.com/en-us/azure/foundry/concepts/evaluation-evaluators/agent-evaluators (search summary only, not fetched)
- Snowflake Documentation, "Evaluate applications with TruLens" — https://docs.snowflake.com/en/user-guide/snowflake-cortex/ai-observability/evaluate-applications-trulens (search summary only, not fetched)
- Morph, "AI Agent Evaluation Frameworks (2026): 7 Compared" — https://www.morphllm.com/ai-agent-evaluation-frameworks (fetch failed: HTTP 429, not retried; search-result summary only)
- DeepEval, "DeepEval vs Ragas" and "Top 5 LLM Evaluation Frameworks in 2026" — https://deepeval.com/blog/deepeval-vs-ragas · https://deepeval.com/blog/top-5-llm-evaluation-frameworks (search summaries only; vendor comparing itself, rankings discounted)
- MLflow, "Top 5 Agent Evaluation Tools in 2026" — https://mlflow.org/top-5-agent-evaluation-frameworks/ (search summary only)
- Confident AI, "Top 9 LLM Evaluation Tools in 2026" — https://www.confident-ai.com/knowledge-base/compare/best-llm-evaluation-tools (search summary only; vendor-authored)