Triaging Agent Errors with Phoenix and PXI @arizeai
Triaging Agent Errors with Phoenix and PXI  @arizeai
Uploaded July 2026 | Updated September 2026, 2 weeks ago
Staring at a pile of agent errors? Most of them aren't yours to fix. Here's how to find the ones that are.

Watch this video to learn:

- How to sort agent failures by who actually owns the fix
- Why good eval categories have yes-or-no answers
- How to trace a bug back to the system prompt that caused it

Chapters
00:00 What can we actually fix?
00:21 Filtering to error spans
00:30 Most errors aren't in your control
01:07 Asking PXI to suggest categories
01:28 Root spans β†’ all spans: now there are 17
01:41 Reading PXI's observation journal
02:58 Why categories need yes/no answers
03:38 Pushing back: narrow these down
04:33 Annotating all 17 error spans
05:04 Tip: bypass approvals
06:27 An injection attempt that was a feature
07:19 Purchasing zero items β€” this one's on us
07:44 Diagnosing the system prompt
09:17 A fix is a hypothesis until you measure it

πŸ”— Try Phoenix: arize.com/phoenix/?utm_source=youtube&utm_medium=video&utm_content=rnabors
πŸ”” Subscribe for weekly content on LLMs, agents, and evaluation: youtube.com/@arizeai?sub_confirmation=1
Triaging Agent Errors with Phoenix and PXIYour Next User Is Not a HumanWhy Most AI Agents Failβ€”and How Anthropic Builds Reliable Ones | Arize Observe 2026Prompt Optimization TechniquesHow to Build the Right Evals for AI Agents | Arize PhoenixMulti-Agent Frameworks: Building & Debugging with Groq and LlamaIndexHow to test AI agents with traces, evals, and CI/CDThe AI Agent That Bypassed Our SecurityI Told It to Pass the Tests... So It Deleted Them.Introducing the New Arize Phoenix Open Source LLM Evals LibraryTracing Agents and Running Evals in TypeScriptBefore You Write Evals for Agents, Read Your Data | Ep. 5
Arize AI |

Triaging Agent Errors with Phoenix and PXI

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER