Uploaded July 2026 | Updated September 2026, 2 weeks ago
Learn the most-skipped step in AI evaluation, and the most important: actually reading your agent's outputs before you measure anything. Error analysis tells you what's worth measuring in the first place.
Skip this step and you end up measuring what's easy to check instead of what's actually breaking, like tone or formatting, while real bugs like hallucinated numbers slip through.
Watch this to learn:
• Why reading traces first stops you from measuring what's easy instead of what matters
• How to define success criteria and generate synthetic test data
• How to categorize failures (hallucinations, wrong tickers, scope errors) to find real patterns
• The "Swiss cheese" model of stacked evals
Chapters:
00:00 Read your data before writing a single eval
00:26 Measure what matters, not what's easy
01:29 Define what "good" means, your success criteria
02:33 Generating synthetic test data
03:11 Spotting failures in the traces
04:32 Hallucinations and wrong tickers
05:22 Categorize failures, patterns emerge
07:00 The Swiss cheese model
07:26 Next: your first code eval
👉 Sign up for free: app.arize.com/auth/login?utm_source=youtube&utm_medium=organic_social&utm_campaign=arize_ax_getting_started
🔗 Learn more about Arize AX: arize.com
📓 Colab notebook for the series: colab.research.google.com/drive/1dViThD0kJjbqtDGE7-ciBYIRwV-HBN1T
📚 Docs: docs.arize.com
🔔 Subscribe for weekly content on LLMs, agents, and evaluation: youtube.com/@arizeai?sub_confirmation=1
#ArizeAX #ErrorAnalysis #LLMEvaluation
Learn the most-skipped step in AI evaluation, and the most important: actually reading your agent's outputs before you measure anything. Error analysis tells you what's worth measuring in the first place.
Skip this step and you end up measuring what's easy to check instead of what's actually breaking, like tone or formatting, while real bugs like hallucinated numbers slip through.
Watch this to learn:
• Why reading traces first stops you from measuring what's easy instead of what matters
• How to define success criteria and generate synthetic test data
• How to categorize failures (hallucinations, wrong tickers, scope errors) to find real patterns
• The "Swiss cheese" model of stacked evals
Chapters:
00:00 Read your data before writing a single eval
00:26 Measure what matters, not what's easy
01:29 Define what "good" means, your success criteria
02:33 Generating synthetic test data
03:11 Spotting failures in the traces
04:32 Hallucinations and wrong tickers
05:22 Categorize failures, patterns emerge
07:00 The Swiss cheese model
07:26 Next: your first code eval
👉 Sign up for free: app.arize.com/auth/login?utm_source=youtube&utm_medium=organic_social&utm_campaign=arize_ax_getting_started
🔗 Learn more about Arize AX: arize.com
📓 Colab notebook for the series: colab.research.google.com/drive/1dViThD0kJjbqtDGE7-ciBYIRwV-HBN1T
📚 Docs: docs.arize.com
🔔 Subscribe for weekly content on LLMs, agents, and evaluation: youtube.com/@arizeai?sub_confirmation=1
#ArizeAX #ErrorAnalysis #LLMEvaluation










