Uploaded May 2026 | Updated September 2026, 2 weeks ago
We use our AI engineering Alyx to debug Alyx.
Read the full blog: arize.com/blog/ai-agent-feedback-loop-arize-alyx
In this video, we walk through the real workflows our engineering team uses to analyze traces, triage failures, debug prompts, and shorten the AI engineering feedback loop.
Modern AI agent traces are too dense to inspect manually. A single trace can contain:
• dozens of spans
• long prompts
• tool calls
• retrieved context
• nested JSON
• token and latency metadata
• exception traces
Instead of manually combing through traces, we use Alyx to:
• search across traces
• analyze failures
• categorize errors semantically
• aggregate patterns across sessions
• identify prompt bugs
• turn production failures into evaluation datasets
We also show how we used Alyx during a company-wide dogfooding session to triage 68 production failures without pulling engineers away from their actual work.
Topics covered:
• AI agent debugging
• LLM observability
• trace analysis
• AI agent evals
• prompt debugging
• semantic error categorization
• dogfooding AI systems
• production AI workflows
Chapters:
00:00 Why we use Alyx to build Alyx
00:42 The problem with debugging AI agents
01:27 Why manual trace inspection breaks down
02:10 Debugging prompts from production traces
03:25 Finding hidden prompt failures
04:20 Using Alyx to fix prompt bugs faster
05:02 Triaging 68 production failures automatically
05:45 Semantic error categorization across traces
06:12 What AI agent debugging workflows look like
Deep dive series:
Part 1 (planning): arize.com/blog/how-to-build-planning-into-your-agent
Part 2 (context management):
arize.com/blog/how-to-manage-llm-context-windows-for-ai-agents
Part 3 (testing and evals):
arize.com/blog/why-testing-ai-agents-is-non-negotiable
#AIEngineering #AIAgents #LLMObservability #AgenticAI #AIInfrastructure #MLOps #promptengineering
We use our AI engineering Alyx to debug Alyx.
Read the full blog: arize.com/blog/ai-agent-feedback-loop-arize-alyx
In this video, we walk through the real workflows our engineering team uses to analyze traces, triage failures, debug prompts, and shorten the AI engineering feedback loop.
Modern AI agent traces are too dense to inspect manually. A single trace can contain:
• dozens of spans
• long prompts
• tool calls
• retrieved context
• nested JSON
• token and latency metadata
• exception traces
Instead of manually combing through traces, we use Alyx to:
• search across traces
• analyze failures
• categorize errors semantically
• aggregate patterns across sessions
• identify prompt bugs
• turn production failures into evaluation datasets
We also show how we used Alyx during a company-wide dogfooding session to triage 68 production failures without pulling engineers away from their actual work.
Topics covered:
• AI agent debugging
• LLM observability
• trace analysis
• AI agent evals
• prompt debugging
• semantic error categorization
• dogfooding AI systems
• production AI workflows
Chapters:
00:00 Why we use Alyx to build Alyx
00:42 The problem with debugging AI agents
01:27 Why manual trace inspection breaks down
02:10 Debugging prompts from production traces
03:25 Finding hidden prompt failures
04:20 Using Alyx to fix prompt bugs faster
05:02 Triaging 68 production failures automatically
05:45 Semantic error categorization across traces
06:12 What AI agent debugging workflows look like
Deep dive series:
Part 1 (planning): arize.com/blog/how-to-build-planning-into-your-agent
Part 2 (context management):
arize.com/blog/how-to-manage-llm-context-windows-for-ai-agents
Part 3 (testing and evals):
arize.com/blog/why-testing-ai-agents-is-non-negotiable
#AIEngineering #AIAgents #LLMObservability #AgenticAI #AIInfrastructure #MLOps #promptengineering










