When AI Agents Fail in Production: Oracle, CA DMV & Tripadvisor @arizeai
When AI Agents Fail in Production: Oracle, CA DMV & Tripadvisor  @arizeai
Uploaded July 2026 | Updated September 2026, 2 weeks ago
AI agents are moving into production across healthcare, public services, and travel, where trust, privacy, safety, and user experience are all on the line.

In this Observe 26 customer panel, Anu Trivedi from Oracle Health, Ajay Gupta from California DMV, Rahul Todkar from Tripadvisor, and Clay Miner from Arize discuss what it takes to build reliable AI systems in high-stakes environments.

They cover how teams validate AI agents before launch, where human-in-the-loop review still matters, why compliance is only the starting point for trust, how traces and reasoning help users understand AI outputs, and how observability, evals, KPIs, guardrails, staged rollouts, and customer feedback help teams detect issues in production.

The panel also explores the next set of challenges for multi-agent systems, including shared memory, goal alignment, long-running workflows, and coordination across agents.

Chapters:
00:00 Panel intro: When AI fails
00:51 Production AI needs higher reliability
01:38 Ajay Gupta on AI at California DMV
03:20 DMV use cases: identity, license plates, documents, and monitoring
05:13 Rahul Todkar on Tripadvisor’s AI platform
05:55 Agentic commerce and context graphs for travel
06:50 Anu Trivedi on Oracle Health
07:46 Healthcare AI and common memory fabric
08:09 Trust and accountability in production AI
10:09 The missing SDLC for AI agents
11:27 Public sector AI: privacy, accountability, and human review
15:09 Compliance vs trust in healthcare AI
17:16 Why reasoning and traces matter to clinicians
18:27 Detecting AI failures with KPIs, monitors, and feedback
22:35 What-if analysis, staged rollouts, and signal vs noise
25:26 Multi-agent failures and shared memory
26:13 Long-running agents and memory management

Subscribe for more conversations on AI agents, LLM evaluation, AI observability, tracing, and production AI reliability.

🔗 Try Arize AX & Phoenix OSS: arize.com
🔔 Subscribe for weekly content on LLMs, agents, and evaluation: youtube.com/@arizeai?sub_confirmation=1
When AI Agents Fail in Production: Oracle, CA DMV & TripadvisorAG2 - Agents for Production EngineeringTriaging Agent Errors with Phoenix and PXIYour Next User Is Not a HumanWhy Most AI Agents Fail—and How Anthropic Builds Reliable Ones | Arize Observe 2026Prompt Optimization TechniquesHow to Build the Right Evals for AI Agents | Arize PhoenixMulti-Agent Frameworks: Building & Debugging with Groq and LlamaIndexHow to test AI agents with traces, evals, and CI/CDThe AI Agent That Bypassed Our SecurityI Told It to Pass the Tests... So It Deleted Them.Introducing the New Arize Phoenix Open Source LLM Evals Library
Arize AI |

When AI Agents Fail in Production: Oracle, CA DMV & Tripadvisor

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER