Writing ·
How to validate agent simulations against production outcomes
Screen a candidate under a fixed simulator and evaluator contract, check it against matched production versions, then confirm selected changes in a bounded live test.
By Youssef Hemimy · agent evaluation · agent reliability · AgentOps
Use simulated customers and tool replies to screen agent changes, not to authorize deployment. Compare a candidate with the incumbent under one fixed scenario mix and evaluator contract. Characterize that simulator against matched production versions. Advance a passing candidate to a bounded live experiment with real outcome measures. A simulated win alone cannot establish how customers, backends, or durable state will behave.
Make the screening decision explicit
The useful question is whether simulation helps reject weak changes while preserving changes worth testing live. Plausible conversations, candidate rankings, and observed customer outcomes are different evidence. Specify one candidate change and its expected improvement before generating trajectories.
| Record before screening | Decision it supports |
|---|---|
| Incumbent, candidate, and falsifiable hypothesis | Keeps interventions separable and names the improvement and regressions to inspect. |
| Scenario distribution and simulator version | Exposes incumbent and candidate to comparable simulated users and tool conditions. |
| Standing and hypothesis-specific evaluators | Checks routine quality as well as the proposed benefit; a new judge needs its own calibration. |
| Prespecified screening criterion | Defines eligibility for a live test before favorable scores are visible. |
| Failed cases and trace IDs | Makes defects inspectable and permits comparison with later production evidence. |
Nubank's reported workflow fixes simulator configuration for a screening round, scores incumbent and candidate with the same suite, and makes a candidate meeting its criterion eligible for live A/B testing. A new evaluator also scores baseline trajectories. Its standing Card Delivery judges were calibrated against human annotations; hypothesis-specific judges require their own calibration report. These are local test contracts, not a threshold the paper prescribes for other teams.
Repeated edits against the same synthetic cases turn those cases into tuning material. Reserve held-out evidence for promotion; the harness-change evaluation guide explains that boundary. This separation is an engineering recommendation, not an experimental result of the Nubank study.
Characterize the simulator against production versions
A screen need not reproduce every live metric exactly, but its decision-relevant behavior needs a measured comparison. The Nubank study compares four deployed Card Delivery versions using 1,000 held-out production conversations and 250 simulated trajectories per version. It inspects conversation length, transcript-embedding proximity, version-level evaluator scores, and blinded expert judgments about whether conversations are real or simulated.
| Diagnostic | What it can—and cannot—tell you |
|---|---|
| User behavior and goal adherence | Compare turn counts, message length, interaction style, and whether synthetic users pursue their assigned goals. Plausible language alone is insufficient. |
| Version ordering | Ask whether the same judges distinguish known stronger and weaker deployed versions. Four version-level points are too few to promise prediction of the next candidate. |
| Judge calibration | Check agreement with human review for the specific failure labels that drive the decision. A realistic simulator cannot rescue a judge that rewards wrong resolutions. |
| Tool-state coherence | A repeated order or account read should stay consistent unless an intervening action changes it. Schema-compatible replies are not proof that the real backend works. |
The study reports aggregate failure-score association across four versions (Pearson r = 0.74; Kendall τ = 0.67), but the middle two versions swap order. Simulated customer messages were also substantially longer than real ones in the pooled comparison. These are useful diagnostics, not proof of reliable candidate-level prediction.
An independent EMNLP 2025 study, SimulatorArena, compares simulators with 909 annotated human–assistant conversations in math tutoring and document creation. Generic simulated users differed from humans; profile conditioning improved alignment with human assistant ratings. It supports checking behavior and rating agreement separately, but does not replicate the Nubank customer-service result.
Mark the boundary where simulation stops
In the reported setup, synthetic customers interact with the agent while selected tool calls receive scenario-conditioned replies instead of production backend responses. This can expose conversation and tool-sequencing failures before user exposure. It does not test backend implementation, actual side effects, persistent state, production latency, or every integration boundary. Some read-only dependencies may remain live, so document exactly which tools are mocked and which are not.
Write down schemas, identifiers, mocked and live tools, and outcomes left unverified. For workflows that write, send, or change state, follow the screen with durable-outcome evaluation. A transcript saying “resolved” is not proof that a real case was resolved. Treat a synthetic user that abandons its goal, an incoherent tool reply, or a miscalibrated judge as broken test infrastructure—not an agent win.
Escalate only to a bounded live comparison
A passing screen earns a live test plan, not a release badge. Define the online outcome, guardrails, population, comparator, and rollback path before exposure; use the agent rollout checklist for that boundary. Retain the screening hypothesis so the live result can show whether the expected behavior translated.
The Nubank paper describes two distinct live comparisons: a Card Management arm improved self-service and transactional NPS against Card Delivery; later, a screened open-weight model improved self-service against its own incumbent while transactional NPS was statistically unchanged. The later model search tested 29 configurations across more than 16,000 simulated conversations. Different interventions and controls mean neither live result isolates simulation as the cause of improvement.
What the evidence cannot establish
One organization's customer-experience setting and four deployed versions do not establish transfer to other agents. The independent simulator study covers different domains. Neither supplies a universal fidelity score, pass threshold, or substitute for backend and customer-outcome validation. The defensible passing-screen claim is narrow: under this versioned simulator, scenario mix, and evaluator contract, the candidate was strong enough to test live.
Sources
FAQ
Can a good simulation score approve an agent release?
No. A simulation screen can qualify a candidate for a bounded live test. It does not validate real backends, persistent effects, or customer outcomes.
How do we know whether our simulated users are useful?
Compare behavior and evaluator results with version-aligned production conversations. Check goal adherence, message patterns, version ordering, judge calibration, and tool-state coherence separately.
What should happen when live and simulated results disagree?
Keep the original screening record, investigate the mismatched scenarios or judges, revise the simulator or evaluation contract, and recheck earlier candidates before relying on it again.
Building something that has to hold up?
We do this work for teams — agent reliability hardening, custom MCP servers, and full-stack AI systems built to survive production.