Uploaded July 2026 | Updated September 2026, 2 weeks ago
Learn how to feed eval explanations back into a coding agent so your system keeps getting better on its own, at scale. This is the finale: an agent that improves itself instead of waiting on you to notice it's broken.
Two prompts and a demo isn't a real system. The eval explanations you've been collecting all series are the highest-leverage input you have for fixing it for good.
Watch this to learn:
• Why eval explanations, not just labels, are the highest-leverage output you have
• How to hand hundreds of failing traces to a coding agent like Claude Code to find themes and propose fixes
• How to keep the loop honest: feed it your requirements, and verify you haven't regressed what already worked
• Where to start if you only have 15 minutes
This is the final video of Arize AX: Getting Started, the full self-improving AI lifecycle.
Chapters:
00:00 The finale: self-improving agents
00:26 Why two prompts isn't a real system
01:17 Eval explanations are the gold
02:23 Hand failing traces to a coding agent
03:00 The improvement workflow with Claude Code
04:02 Keep the loop honest, feed it requirements
04:41 Themes vs. individual failures
05:17 Worked example: a week in production
06:32 Verify you didn't regress what worked
06:47 The full software development lifecycle
07:53 Where to start, 15 minutes
08:43 Wrap-up and resources
👉 Sign up for free: app.arize.com/auth/login?utm_source=youtube&utm_medium=organic_social&utm_campaign=arize_ax_getting_started
🔗 Learn more about Arize AX: arize.com
📓 Colab notebook for the series: colab.research.google.com/drive/1dViThD0kJjbqtDGE7-ciBYIRwV-HBN1T
📚 Docs: docs.arize.com
🔔 Subscribe for weekly content on LLMs, agents, and evaluation: youtube.com/@arizeai?sub_confirmation=1
#ArizeAX #SelfImprovingAgents #AIAgents
Learn how to feed eval explanations back into a coding agent so your system keeps getting better on its own, at scale. This is the finale: an agent that improves itself instead of waiting on you to notice it's broken.
Two prompts and a demo isn't a real system. The eval explanations you've been collecting all series are the highest-leverage input you have for fixing it for good.
Watch this to learn:
• Why eval explanations, not just labels, are the highest-leverage output you have
• How to hand hundreds of failing traces to a coding agent like Claude Code to find themes and propose fixes
• How to keep the loop honest: feed it your requirements, and verify you haven't regressed what already worked
• Where to start if you only have 15 minutes
This is the final video of Arize AX: Getting Started, the full self-improving AI lifecycle.
Chapters:
00:00 The finale: self-improving agents
00:26 Why two prompts isn't a real system
01:17 Eval explanations are the gold
02:23 Hand failing traces to a coding agent
03:00 The improvement workflow with Claude Code
04:02 Keep the loop honest, feed it requirements
04:41 Themes vs. individual failures
05:17 Worked example: a week in production
06:32 Verify you didn't regress what worked
06:47 The full software development lifecycle
07:53 Where to start, 15 minutes
08:43 Wrap-up and resources
👉 Sign up for free: app.arize.com/auth/login?utm_source=youtube&utm_medium=organic_social&utm_campaign=arize_ax_getting_started
🔗 Learn more about Arize AX: arize.com
📓 Colab notebook for the series: colab.research.google.com/drive/1dViThD0kJjbqtDGE7-ciBYIRwV-HBN1T
📚 Docs: docs.arize.com
🔔 Subscribe for weekly content on LLMs, agents, and evaluation: youtube.com/@arizeai?sub_confirmation=1
#ArizeAX #SelfImprovingAgents #AIAgents
![How Uber Evaluates AI Agents at Production Scale | Arize Observe 2026
Running evals is one thing. Building an evaluation system that uncovers problems, changes the roadmap, and continuously improves an AI agent is much harder.
In this Arize:Observe 2026 session, Aayush Agrawal, Senior AI Product Manager at Uber, explains what Uber learned while building an evaluation platform for production AI agents.
Uber’s agent platform supports teams ranging from first-time agent builders to engineers shipping customer-facing agents at global scale. Across those teams, Uber repeatedly found that access to evaluation tools was not enough. Teams needed tracing by default, automatically generated evaluators, continuously updated datasets, broader ownership, and a development process built around learning from production.
Aayush covers:
• Why teams tend to build their agents first and add evals later
• How Uber provides tracing automatically when an agent is deployed
• Why complete agent trajectories matter more than input-output logs
• How Uber generates agent-specific evaluators from configurations and traces
• Turning production failures into continuously updated evaluation datasets
• Using real conversations to simulate multi-turn agent behavior
• Giving product managers, designers, and operations teams ownership of evals
• Why optimizing for a single launch score can produce misleading results
• The questions Uber uses to assess whether an evaluation system is useful
• How a voice-booking agent exposed a failure that offline evals missed
• Using changes in session length to detect unexpected production behavior
• How production traces can power agent insights, experiments, and improved versions
One example came from Uber’s voice-booking agent. When a child mentioned wanting pizza during a ride request, the agent interpreted the background speech as a new destination. Offline evals had not anticipated the scenario, but a spike in the average number of conversation turns surfaced the problem. A conversational designer then used that insight to update the agent, evaluators, and dataset.
The larger lesson is that evals should help teams learn what to build next. The most effective systems connect production traces, failure analysis, datasets, experiments, and agent improvements in one continuous loop. :contentReference[oaicite:0]{index=0}
Chapters:
00:00 Why having evals is not enough
00:48 The stakes of running AI agents at Uber
01:40 Inside Uber’s agent platform
02:41 Supporting every type of agent builder
03:38 Why teams add evals too late
04:25 Making tracing the default
05:48 Automatically generating useful evaluators
07:03 Building datasets from production failures
08:20 Bringing product and design teams into evals
09:37 Why launch-gate metrics fail
10:09 Better questions for evaluating your evals
12:01 What Uber’s voice-booking agent taught the team
13:12 From eval afterthought to product insight engine
13:36 The future agent improvement loop
15:01 The hardest part was never the tooling
🔗 Learn more about Arize: https://arize.com
🔗 Explore Arize:Observe: https://arize.com/observe
🔔 Subscribe for more talks about AI agents, LLM evaluation, observability, and production AI:
https://www.youtube.com/@arizeai?sub_confirmation=1
#AIAgentEvals #Uber #AIEngineering How Uber Evaluates AI Agents at Production Scale | Arize Observe 2026](https://i.ytimg.com/vi/vJh126DQzEc/mqdefault.jpg)






