Uploaded June 2026 | Updated September 2026, 2 weeks ago
As AI agents become more autonomous, traditional observability approaches are reaching their limits. The challenge is no longer just collecting traces and logs—it’s understanding why agents behave the way they do.
In this session, Nate Slater from AWS explores how agentic AI is fundamentally changing observability. Drawing on his experience building AI-powered root cause analysis systems, Nate explains why traditional monitoring tools struggle with probabilistic systems, multi-agent workflows, temporal drift, and the explosion of telemetry generated by AI agents. He argues that the future of observability requires agents that can reason about other agents, transforming observability from passive data collection into active analysis and diagnosis.
The talk covers the evolution from traditional application monitoring to AI-native observability, the challenges of debugging agentic systems at scale, and how technologies like OpenTelemetry, AWS Bedrock Agent Core, and Arize observability tools can work together to provide visibility into increasingly autonomous systems. Learn why recording telemetry is no longer enough—and why reasoning over observability data is becoming essential for enterprise AI adoption.
Timestamps
00:00 Why Observability Gets Harder with AI Agents
00:52 Traditional Tracing and Debugging Challenges
02:16 Using AI to Understand AI Systems
02:50 Introduction and Background
04:21 Correlation vs. Causation in Observability
06:18 New Challenges in Agentic Systems
06:46 Probabilistic Behavior and Error Propagation
07:30 Temporal Drift and Agent Memory
08:23 The Scale Problem: Thousands of Agents
08:56 From Recording Data to Reasoning About It
09:46 Why Traditional Observability Falls Short
10:23 Causal Reasoning and Intent-Aware Detection
12:19 Why Observability Matters More in the Agent Era
13:44 Enterprise Risk and Agent Liability
14:48 Building Agent Observability with AWS and Arize
16:34 From Observability to Reasoning
17:40 AI Agents Debugging Other AI Agents
19:37 The Future of Agent Observability
20:04 Key Takeaways
Key Takeaways
• AI agents introduce new observability challenges including probabilistic behavior, compounding errors, temporal drift, and massive increases in telemetry volume.
• Traditional observability tools focus on collecting data, but future systems must be able to reason about that data and identify causal relationships.
• Enterprises will need agent-based observability systems to understand, monitor, and govern increasingly autonomous AI workflows.
• Observability is becoming a prerequisite for safely deploying large-scale agentic systems in production environments.
• The future of AI operations involves agents observing, evaluating, and debugging other agents.
--
#AIAgents #Observability #AWS #LLMOps #AgenticAI #OpenTelemetry #AIEngineering #PlatformEngineering #AIInfrastructure #ArizeObserve
🔗 Try Arize AX & Phoenix OSS: arize.com
🔔 Subscribe for weekly content on LLMs, agents, and evaluation: youtube.com/@arizeai?sub_confirmation=1
As AI agents become more autonomous, traditional observability approaches are reaching their limits. The challenge is no longer just collecting traces and logs—it’s understanding why agents behave the way they do.
In this session, Nate Slater from AWS explores how agentic AI is fundamentally changing observability. Drawing on his experience building AI-powered root cause analysis systems, Nate explains why traditional monitoring tools struggle with probabilistic systems, multi-agent workflows, temporal drift, and the explosion of telemetry generated by AI agents. He argues that the future of observability requires agents that can reason about other agents, transforming observability from passive data collection into active analysis and diagnosis.
The talk covers the evolution from traditional application monitoring to AI-native observability, the challenges of debugging agentic systems at scale, and how technologies like OpenTelemetry, AWS Bedrock Agent Core, and Arize observability tools can work together to provide visibility into increasingly autonomous systems. Learn why recording telemetry is no longer enough—and why reasoning over observability data is becoming essential for enterprise AI adoption.
Timestamps
00:00 Why Observability Gets Harder with AI Agents
00:52 Traditional Tracing and Debugging Challenges
02:16 Using AI to Understand AI Systems
02:50 Introduction and Background
04:21 Correlation vs. Causation in Observability
06:18 New Challenges in Agentic Systems
06:46 Probabilistic Behavior and Error Propagation
07:30 Temporal Drift and Agent Memory
08:23 The Scale Problem: Thousands of Agents
08:56 From Recording Data to Reasoning About It
09:46 Why Traditional Observability Falls Short
10:23 Causal Reasoning and Intent-Aware Detection
12:19 Why Observability Matters More in the Agent Era
13:44 Enterprise Risk and Agent Liability
14:48 Building Agent Observability with AWS and Arize
16:34 From Observability to Reasoning
17:40 AI Agents Debugging Other AI Agents
19:37 The Future of Agent Observability
20:04 Key Takeaways
Key Takeaways
• AI agents introduce new observability challenges including probabilistic behavior, compounding errors, temporal drift, and massive increases in telemetry volume.
• Traditional observability tools focus on collecting data, but future systems must be able to reason about that data and identify causal relationships.
• Enterprises will need agent-based observability systems to understand, monitor, and govern increasingly autonomous AI workflows.
• Observability is becoming a prerequisite for safely deploying large-scale agentic systems in production environments.
• The future of AI operations involves agents observing, evaluating, and debugging other agents.
--
#AIAgents #Observability #AWS #LLMOps #AgenticAI #OpenTelemetry #AIEngineering #PlatformEngineering #AIInfrastructure #ArizeObserve
🔗 Try Arize AX & Phoenix OSS: arize.com
🔔 Subscribe for weekly content on LLMs, agents, and evaluation: youtube.com/@arizeai?sub_confirmation=1




![How Uber Evaluates AI Agents at Production Scale | Arize Observe 2026
Running evals is one thing. Building an evaluation system that uncovers problems, changes the roadmap, and continuously improves an AI agent is much harder.
In this Arize:Observe 2026 session, Aayush Agrawal, Senior AI Product Manager at Uber, explains what Uber learned while building an evaluation platform for production AI agents.
Uber’s agent platform supports teams ranging from first-time agent builders to engineers shipping customer-facing agents at global scale. Across those teams, Uber repeatedly found that access to evaluation tools was not enough. Teams needed tracing by default, automatically generated evaluators, continuously updated datasets, broader ownership, and a development process built around learning from production.
Aayush covers:
• Why teams tend to build their agents first and add evals later
• How Uber provides tracing automatically when an agent is deployed
• Why complete agent trajectories matter more than input-output logs
• How Uber generates agent-specific evaluators from configurations and traces
• Turning production failures into continuously updated evaluation datasets
• Using real conversations to simulate multi-turn agent behavior
• Giving product managers, designers, and operations teams ownership of evals
• Why optimizing for a single launch score can produce misleading results
• The questions Uber uses to assess whether an evaluation system is useful
• How a voice-booking agent exposed a failure that offline evals missed
• Using changes in session length to detect unexpected production behavior
• How production traces can power agent insights, experiments, and improved versions
One example came from Uber’s voice-booking agent. When a child mentioned wanting pizza during a ride request, the agent interpreted the background speech as a new destination. Offline evals had not anticipated the scenario, but a spike in the average number of conversation turns surfaced the problem. A conversational designer then used that insight to update the agent, evaluators, and dataset.
The larger lesson is that evals should help teams learn what to build next. The most effective systems connect production traces, failure analysis, datasets, experiments, and agent improvements in one continuous loop. :contentReference[oaicite:0]{index=0}
Chapters:
00:00 Why having evals is not enough
00:48 The stakes of running AI agents at Uber
01:40 Inside Uber’s agent platform
02:41 Supporting every type of agent builder
03:38 Why teams add evals too late
04:25 Making tracing the default
05:48 Automatically generating useful evaluators
07:03 Building datasets from production failures
08:20 Bringing product and design teams into evals
09:37 Why launch-gate metrics fail
10:09 Better questions for evaluating your evals
12:01 What Uber’s voice-booking agent taught the team
13:12 From eval afterthought to product insight engine
13:36 The future agent improvement loop
15:01 The hardest part was never the tooling
🔗 Learn more about Arize: https://arize.com
🔗 Explore Arize:Observe: https://arize.com/observe
🔔 Subscribe for more talks about AI agents, LLM evaluation, observability, and production AI:
https://www.youtube.com/@arizeai?sub_confirmation=1
#AIAgentEvals #Uber #AIEngineering How Uber Evaluates AI Agents at Production Scale | Arize Observe 2026](https://i.ytimg.com/vi/vJh126DQzEc/mqdefault.jpg)





