Indonesia Singapore ไทย Pilipinas Việt Nam Malaysia မြန်မာ ລາວ
← Back to Blog

AI Agents in Production: Why Live Data Beats Lab Scores

A passing grade on curated test data means nothing if your AI agent fails on real customer conversations — continuous live evaluation is non-negotiable.

An AI agent evaluation dashboard showing live traffic signals diverging from a perfect lab score
Illustrated by Mikael Venne

Snowflake's Cortex agent evals give you a grade. Live traffic gives you the truth. Here's how to build evaluation discipline that scales in production.

A curated test dataset is a controlled environment. Your customers are not.

Snowflake’s Cortex Agent Evaluations offer something genuinely useful: a structured, GPA-scored framework for assessing AI agent performance before it ships. But as Monte Carlo’s analysis of the platform makes clear, a clean score on a golden dataset is a starting point, not a finish line. The gap between lab performance and live production behaviour is where trust in AI — and in the data programmes that feed it — actually gets built or broken.

The Evaluation Illusion: What a Good Score Actually Tells You

Cortex Agent Evaluations work by running an agent against a curated set of test conversations, scoring outputs across dimensions like relevance, groundedness, and task completion. Think of it as a controlled exam under ideal conditions.

The problem is that real customer conversations don’t behave like exam questions. They’re ambiguous, multilingual, context-dependent, and shaped by the specific data your agent has access to at that moment. In Southeast Asian markets — where a single customer might switch between Bahasa Indonesia and English mid-conversation, or arrive via a Shopee product page with highly specific purchase intent — the distance between a golden test set and live traffic can be enormous.

Monte Carlo’s approach is to run the same evaluation logic continuously against live conversation data, not just at deployment. That shift — from point-in-time testing to always-on monitoring — is the difference between knowing your agent was good at launch and knowing it is good right now.

First-Party Data Quality Is the Upstream Problem Nobody Wants to Own

Here’s what often goes unsaid in conversations about agent evaluation: the quality of an AI agent’s outputs is a direct downstream consequence of the quality of the data it reasons over. Garbage in, plausible-sounding garbage out — and plausible-sounding is the dangerous kind.

For brands running first-party data programmes in Southeast Asia, this creates a specific accountability gap. You may have consent-compliant data collection in place. You may have clean CRM records. But if the data pipeline feeding your AI layer has silent drift — schema changes, missing values, stale segments — your agent will confidently produce answers based on a reality that no longer exists.

This is where the dbt community’s growing focus on data contracts becomes relevant. The dbt Summit 2026 programme signals that analytics engineering is maturing toward explicit, testable agreements between data producers and consumers. Treating the data layer as a trusted supplier — with SLAs, not just good intentions — is the foundation any serious AI evaluation programme needs.

The implementation implication: before investing in agent evaluation tooling, audit whether your first-party data infrastructure can surface data quality failures automatically and alert teams before those failures propagate into AI outputs.


What Byzantine Fault Tolerance Has to Do With Your Data Stack

Towards Data Science recently revisited Byzantine Fault Tolerance — the computer science concept of building systems that can reach correct conclusions even when some of their inputs are unreliable or actively wrong. The framing is deliberately provocative: how do you make decisions when you can’t trust everyone in the room?

It’s a useful mental model for modern data teams. In a production AI environment, you’re rarely working with a single clean data source. You have CRM data, behavioural signals, consent management outputs, campaign attribution feeds — each with its own freshness profile and failure modes. An agent that assumes all of those inputs are trustworthy simultaneously is brittle by design.

The practical answer isn’t paranoia — it’s redundancy and explicitness. Build evaluation pipelines that can flag when an agent’s reasoning relies heavily on a data source that’s showing quality degradation. Design your data architecture so that the agent can signal uncertainty rather than manufacture confidence. For teams using Snowflake’s Cortex, pairing the native evaluation framework with an observability layer like Monte Carlo creates exactly that kind of fault-tolerant feedback loop.

For Southeast Asian brands specifically, this matters because data collection environments here are genuinely more complex — multiple platforms, multiple consent jurisdictions, multiple languages. The tolerance for silent data failures needs to be lower, not higher.

From Evaluation to Outcomes: The ROI Accountability Shift

Airship’s recently released commissioned study claims a 476% ROI for a composite organisation using its customer engagement platform — and while commissioned studies carry obvious caveats, the broader trend it represents is real: enterprise buyers are increasingly demanding outcome guarantees, not capability promises.

Airship’s results guarantee is a commercial bet that their platform’s performance is measurable enough to stake revenue on. That’s a meaningful signal for how the market is evolving. AI agent deployments are heading in the same direction. Brands won’t accept “our evaluations look good” as a sufficient answer when agent-driven experiences are influencing purchase decisions at scale.

This puts evaluation infrastructure — the ability to continuously measure agent performance against business outcomes, not just technical benchmarks — at the centre of the AI investment conversation. The teams that will win aren’t those with the most sophisticated models. They’re the ones who can demonstrate, in near real-time, that their AI is performing as intended against the outcomes that actually matter: conversion rates, resolution rates, customer satisfaction scores.

Building that capability requires treating evaluation as a product, not a pre-launch checklist.


Key Takeaways

  • Run agent evaluations continuously against live traffic, not just against curated golden datasets before deployment — production behaviour is the only behaviour that matters.
  • Invest in first-party data quality infrastructure before AI evaluation tooling; downstream agent failures are almost always upstream data problems in disguise.
  • As enterprise buyers demand outcome guarantees from AI platforms, teams that can instrument and report on real business metrics — not just technical scores — will own the internal credibility to scale.

The shift from lab-grade AI evaluation to production-grade observability is not a technical upgrade — it’s a governance decision. It requires someone to own the question: what does good actually look like, in the real conversations our customers are having right now? As AI agents take on more consequential roles in customer journeys across Southeast Asia, that question deserves a named owner, a defined methodology, and a live dashboard — not a spreadsheet that gets reviewed quarterly.

At grzzly, we help brands across Southeast Asia build first-party data programmes that are architected for exactly this kind of accountability — from consent infrastructure through to the data pipelines that feed AI-driven experiences. If you’re thinking about how to move from point-in-time evaluation to always-on data confidence, we’d enjoy that conversation. Let’s talk

Lavender Grizzly

Written by

Lavender Grizzly

Turning privacy constraints into competitive advantage. Builds first-party data programmes that are compliant by design, valuable by intent, and trusted by the people whose data they hold.

Enjoyed this?
Let's talk.

Start a conversation