Uploaded July 2026 | Updated September 2026, 2 weeks ago
Learn how to write a deterministic code eval that catches real production bugs, no API calls and no cost. Not every eval needs an LLM: the simplest checks are often the most reliable.
A ticker check sounds trivial, but "all passed" doesn't mean the bug is gone. It just means you haven't found where this eval stops being enough.
Watch this to learn:
• How to write a deterministic code eval (a ticker check) in about 5 lines of Python
• How to connect the AX client from your notebook and register an evaluator
• Why "all passed" doesn't guarantee the bug is gone, and where code evals fit vs. LLM judges
Chapters:
00:00 Writing your first eval
00:13 The simplest useful eval: a ticker check
00:37 Connecting the AX client in your notebook
01:03 Building the code evaluator
01:40 Running it, all pass but bugs still lurk
02:30 Real failures this catches
03:34 When to reach for a code eval
04:26 Next: LLM-as-a-judge
👉 Sign up for free: app.arize.com/auth/login?utm_source=youtube&utm_medium=organic_social&utm_campaign=arize_ax_getting_started
🔗 Learn more about Arize AX: arize.com
📓 Colab notebook for the series: colab.research.google.com/drive/1dViThD0kJjbqtDGE7-ciBYIRwV-HBN1T
📚 Docs: docs.arize.com
🔔 Subscribe for weekly content on LLMs, agents, and evaluation: youtube.com/@arizeai?sub_confirmation=1
#ArizeAX #CodeEvals #LLMEvaluation
Learn how to write a deterministic code eval that catches real production bugs, no API calls and no cost. Not every eval needs an LLM: the simplest checks are often the most reliable.
A ticker check sounds trivial, but "all passed" doesn't mean the bug is gone. It just means you haven't found where this eval stops being enough.
Watch this to learn:
• How to write a deterministic code eval (a ticker check) in about 5 lines of Python
• How to connect the AX client from your notebook and register an evaluator
• Why "all passed" doesn't guarantee the bug is gone, and where code evals fit vs. LLM judges
Chapters:
00:00 Writing your first eval
00:13 The simplest useful eval: a ticker check
00:37 Connecting the AX client in your notebook
01:03 Building the code evaluator
01:40 Running it, all pass but bugs still lurk
02:30 Real failures this catches
03:34 When to reach for a code eval
04:26 Next: LLM-as-a-judge
👉 Sign up for free: app.arize.com/auth/login?utm_source=youtube&utm_medium=organic_social&utm_campaign=arize_ax_getting_started
🔗 Learn more about Arize AX: arize.com
📓 Colab notebook for the series: colab.research.google.com/drive/1dViThD0kJjbqtDGE7-ciBYIRwV-HBN1T
📚 Docs: docs.arize.com
🔔 Subscribe for weekly content on LLMs, agents, and evaluation: youtube.com/@arizeai?sub_confirmation=1
#ArizeAX #CodeEvals #LLMEvaluation










![How Uber Evaluates AI Agents at Production Scale | Arize Observe 2026
Running evals is one thing. Building an evaluation system that uncovers problems, changes the roadmap, and continuously improves an AI agent is much harder.
In this Arize:Observe 2026 session, Aayush Agrawal, Senior AI Product Manager at Uber, explains what Uber learned while building an evaluation platform for production AI agents.
Uber’s agent platform supports teams ranging from first-time agent builders to engineers shipping customer-facing agents at global scale. Across those teams, Uber repeatedly found that access to evaluation tools was not enough. Teams needed tracing by default, automatically generated evaluators, continuously updated datasets, broader ownership, and a development process built around learning from production.
Aayush covers:
• Why teams tend to build their agents first and add evals later
• How Uber provides tracing automatically when an agent is deployed
• Why complete agent trajectories matter more than input-output logs
• How Uber generates agent-specific evaluators from configurations and traces
• Turning production failures into continuously updated evaluation datasets
• Using real conversations to simulate multi-turn agent behavior
• Giving product managers, designers, and operations teams ownership of evals
• Why optimizing for a single launch score can produce misleading results
• The questions Uber uses to assess whether an evaluation system is useful
• How a voice-booking agent exposed a failure that offline evals missed
• Using changes in session length to detect unexpected production behavior
• How production traces can power agent insights, experiments, and improved versions
One example came from Uber’s voice-booking agent. When a child mentioned wanting pizza during a ride request, the agent interpreted the background speech as a new destination. Offline evals had not anticipated the scenario, but a spike in the average number of conversation turns surfaced the problem. A conversational designer then used that insight to update the agent, evaluators, and dataset.
The larger lesson is that evals should help teams learn what to build next. The most effective systems connect production traces, failure analysis, datasets, experiments, and agent improvements in one continuous loop. :contentReference[oaicite:0]{index=0}
Chapters:
00:00 Why having evals is not enough
00:48 The stakes of running AI agents at Uber
01:40 Inside Uber’s agent platform
02:41 Supporting every type of agent builder
03:38 Why teams add evals too late
04:25 Making tracing the default
05:48 Automatically generating useful evaluators
07:03 Building datasets from production failures
08:20 Bringing product and design teams into evals
09:37 Why launch-gate metrics fail
10:09 Better questions for evaluating your evals
12:01 What Uber’s voice-booking agent taught the team
13:12 From eval afterthought to product insight engine
13:36 The future agent improvement loop
15:01 The hardest part was never the tooling
🔗 Learn more about Arize: https://arize.com
🔗 Explore Arize:Observe: https://arize.com/observe
🔔 Subscribe for more talks about AI agents, LLM evaluation, observability, and production AI:
https://www.youtube.com/@arizeai?sub_confirmation=1
#AIAgentEvals #Uber #AIEngineering How Uber Evaluates AI Agents at Production Scale | Arize Observe 2026](https://i.ytimg.com/vi/vJh126DQzEc/mqdefault.jpg)