How we debug AI agents using AI agents (real trace debugging workflows) @arizeai
How we debug AI agents using AI agents (real trace debugging workflows)  @arizeai
Uploaded May 2026 | Updated September 2026, 2 weeks ago
We use our AI engineering Alyx to debug Alyx.

Read the full blog: arize.com/blog/ai-agent-feedback-loop-arize-alyx

In this video, we walk through the real workflows our engineering team uses to analyze traces, triage failures, debug prompts, and shorten the AI engineering feedback loop.

Modern AI agent traces are too dense to inspect manually. A single trace can contain:

• dozens of spans
• long prompts
• tool calls
• retrieved context
• nested JSON
• token and latency metadata
• exception traces

Instead of manually combing through traces, we use Alyx to:

• search across traces
• analyze failures
• categorize errors semantically
• aggregate patterns across sessions
• identify prompt bugs
• turn production failures into evaluation datasets

We also show how we used Alyx during a company-wide dogfooding session to triage 68 production failures without pulling engineers away from their actual work.

Topics covered:

• AI agent debugging
• LLM observability
• trace analysis
• AI agent evals
• prompt debugging
• semantic error categorization
• dogfooding AI systems
• production AI workflows

Chapters:
00:00 Why we use Alyx to build Alyx
00:42 The problem with debugging AI agents
01:27 Why manual trace inspection breaks down
02:10 Debugging prompts from production traces
03:25 Finding hidden prompt failures
04:20 Using Alyx to fix prompt bugs faster
05:02 Triaging 68 production failures automatically
05:45 Semantic error categorization across traces
06:12 What AI agent debugging workflows look like

Deep dive series:
Part 1 (planning): arize.com/blog/how-to-build-planning-into-your-agent

Part 2 (context management):
arize.com/blog/how-to-manage-llm-context-windows-for-ai-agents

Part 3 (testing and evals):
arize.com/blog/why-testing-ai-agents-is-non-negotiable

#AIEngineering #AIAgents #LLMObservability #AgenticAI #AIInfrastructure #MLOps #promptengineering
How we debug AI agents using AI agents (real trace debugging workflows)How Phoenix became the standard for AI observabilityLLMs Are Leverage (And Why Most Developers Misuse Them)Why AI Agents Break in Production (and Why You Need Evals) | Ep. 1In a World Where Everyone Can Code, What Are You Worth?How to Measure AI Coding Agent ROI with Claude Code Tracing | Arize AXLLM-as-a-Judge 101Improving Agents in Production with Online Evals - Arize AXTypeScript Agents: How To Build and EvaluateOne AI Question - where do agents fail in production, with Fuad AliArize Skills: Add Instrumentation & Tracing to Your AI App with Claude Code, Copilot, or CursorIntroduction To Arize AX Evals
Arize AI |

How we debug AI agents using AI agents (real trace debugging workflows)

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER