Before You Write Evals for Agents, Read Your Data | Ep. 5 @arizeai
Before You Write Evals for Agents, Read Your Data | Ep. 5  @arizeai
Uploaded July 2026 | Updated September 2026, 2 weeks ago
Learn the most-skipped step in AI evaluation, and the most important: actually reading your agent's outputs before you measure anything. Error analysis tells you what's worth measuring in the first place.

Skip this step and you end up measuring what's easy to check instead of what's actually breaking, like tone or formatting, while real bugs like hallucinated numbers slip through.

Watch this to learn:
• Why reading traces first stops you from measuring what's easy instead of what matters
• How to define success criteria and generate synthetic test data
• How to categorize failures (hallucinations, wrong tickers, scope errors) to find real patterns
• The "Swiss cheese" model of stacked evals

Chapters:
00:00 Read your data before writing a single eval
00:26 Measure what matters, not what's easy
01:29 Define what "good" means, your success criteria
02:33 Generating synthetic test data
03:11 Spotting failures in the traces
04:32 Hallucinations and wrong tickers
05:22 Categorize failures, patterns emerge
07:00 The Swiss cheese model
07:26 Next: your first code eval

👉 Sign up for free: app.arize.com/auth/login?utm_source=youtube&utm_medium=organic_social&utm_campaign=arize_ax_getting_started
🔗 Learn more about Arize AX: arize.com
📓 Colab notebook for the series: colab.research.google.com/drive/1dViThD0kJjbqtDGE7-ciBYIRwV-HBN1T
📚 Docs: docs.arize.com
🔔 Subscribe for weekly content on LLMs, agents, and evaluation: youtube.com/@arizeai?sub_confirmation=1
#ArizeAX #ErrorAnalysis #LLMEvaluation
Before You Write Evals for Agents, Read Your Data | Ep. 5AI Builders Meetup - San FranciscoBinary Versus Score LLM Evals: What the Research SaysStop Manually Debugging AI Agents: Build an Automated PR Loop | AI BuildersMeta AI Researcher explains ARE: scaling up agent environments and evaluationsWhat Production AI Agent Teams Are Building Today | Mastra | Arize Observe 2026The Hard Truth About Being an AI EngineerOne AI Question - whats your spicy take on AI with Jim BennettLLM as a Judge 102:  Meta EvaluationTrace-Level Evals for a Recommendation AgentHow To Debug AI Agents: Tracing, Observability & EvalsLLM Evaluation Using Prompt Learning
Arize AI |

Before You Write Evals for Agents, Read Your Data | Ep. 5

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER