How to Build the Right Evals for AI Agents | Arize Phoenix @arizeai
How to Build the Right Evals for AI Agents | Arize Phoenix  @arizeai
Uploaded July 2026 | Updated September 2026, 2 weeks ago
AI agent evaluation can feel overwhelming. Agents use tools, skills, memory, files, and multi-turn conversations, creating far more failure points than a traditional LLM application.

In this Arize:Observe 2026 session, Elizabeth Hutton, Senior Software Engineer for Evals at Arize, shares what the Phoenix team learned while building and evaluating Pixie, an AI engineering agent inside the open source Arize Phoenix platform.

Rather than trying to design a complete evaluation suite upfront, the team built evals gradually as the agent gained new capabilities. Elizabeth explains how tracing, targeted capability evals, regression testing, production feedback, and trace analysis work together to create a practical agent improvement loop.

The session covers:

• Why tracing should come before formal evaluations
• How traces become the source of truth for debugging agents
• How to create small capability evals for tools, skills, and output formatting
• Using synthetic data to cover edge cases, typos, and negative examples
• Storing eval datasets and harnesses alongside application code
• Running agent evals and experiments with coding agents
• Turning targeted evals into regression checks in CI
• Combining narrow evals with end-to-end Playwright tests
• Why synthetic first-turn queries fail to represent real agent behavior
• Building realistic multi-turn datasets from production traces
• Using user feedback, LLM-as-a-judge, code evaluators, and trajectory evals
• Reviewing traces to discover failure modes you did not anticipate
• Moving production failures into development and regression datasets
• Closing the loop between tracing, experiments, implementation, and evals

The central lesson: evaluation is a development practice, not a one-time project. Start with tracing, add narrow evals as capabilities emerge, and expand the suite using failures discovered in production. :contentReference[oaicite:0]{index=0}

Chapters:
00:00 Why AI agent evals feel overwhelming
01:09 Meet Pixie, Phoenix’s AI engineering agent
02:33 Agent architecture and the growing failure surface
03:34 Start with tracing
04:53 Add narrow capability evals
06:52 Turn capability evals into regression tests
08:53 Why synthetic evaluation data falls short
11:03 Discovery evals for production agents
12:04 User feedback and automated evaluations
12:58 Trace review and agent error analysis
14:33 Closing the agent improvement loop
14:59 Lessons learned from building Pixie
16:11 Open source evals in Arize Phoenix

🔗 Explore Arize Phoenix: phoenix.arize.com
🔗 View Phoenix on GitHub: github.com/Arize-ai/phoenix
🔗 Read the documentation: arize.com/docs/phoenix
🔔 Subscribe for more videos about AI agents, LLM evaluation, observability, and open source AI:
youtube.com/@arizeai?sub_confirmation=1

#ArizePhoenix #LLMEvals #AIAgents
How to Build the Right Evals for AI Agents | Arize PhoenixMulti-Agent Frameworks: Building & Debugging with Groq and LlamaIndexHow to test AI agents with traces, evals, and CI/CDThe AI Agent That Bypassed Our SecurityI Told It to Pass the Tests... So It Deleted Them.Introducing the New Arize Phoenix Open Source LLM Evals LibraryTracing Agents and Running Evals in TypeScriptBefore You Write Evals for Agents, Read Your Data | Ep. 5AI Builders Meetup - San FranciscoBinary Versus Score LLM Evals: What the Research SaysStop Manually Debugging AI Agents: Build an Automated PR Loop | AI BuildersMeta AI Researcher explains ARE: scaling up agent environments and evaluations
Arize AI |

How to Build the Right Evals for AI Agents | Arize Phoenix

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER