Uploaded July 2026 | Updated September 2026, 2 weeks ago
Learn the two primitives every AI evaluation strategy is built on, and why agents make both of them harder. Before you can fix an agent, you need a shared vocabulary for what's actually going wrong.
Multi-step agents fail in ways a single LLM call never will: picking the wrong tool, compounding small mistakes across steps, or breaking down entirely in multi-agent setups.
Watch this to learn:
β’ Traces and spans: how to see every LLM call, tool call, and agent step
β’ Code evals vs. LLM-as-a-judge, and when to use each
β’ Why agents raise the difficulty (tool selection, multi-step non-determinism, multi-agent systems)
β’ Capability evals vs. regression evals
Chapters:
00:00 Recap: why shipping AI is different
00:29 What evals actually are
00:42 Traces and spans, explained
01:23 Code evals vs. LLM-as-a-judge
03:04 Why agents make evaluation harder
05:11 Capability evals vs. regression evals
05:37 Recap and what's next
π Sign up for free: app.arize.com/auth/login?utm_source=youtube&utm_medium=organic_social&utm_campaign=arize_ax_getting_started
π Learn more about Arize AX: arize.com
π Colab notebook for the series: colab.research.google.com/drive/1dViThD0kJjbqtDGE7-ciBYIRwV-HBN1T
π Docs: docs.arize.com
π Subscribe for weekly content on LLMs, agents, and evaluation: youtube.com/@arizeai?sub_confirmation=1
#ArizeAX #AgentObservability #LLMEvaluation
Learn the two primitives every AI evaluation strategy is built on, and why agents make both of them harder. Before you can fix an agent, you need a shared vocabulary for what's actually going wrong.
Multi-step agents fail in ways a single LLM call never will: picking the wrong tool, compounding small mistakes across steps, or breaking down entirely in multi-agent setups.
Watch this to learn:
β’ Traces and spans: how to see every LLM call, tool call, and agent step
β’ Code evals vs. LLM-as-a-judge, and when to use each
β’ Why agents raise the difficulty (tool selection, multi-step non-determinism, multi-agent systems)
β’ Capability evals vs. regression evals
Chapters:
00:00 Recap: why shipping AI is different
00:29 What evals actually are
00:42 Traces and spans, explained
01:23 Code evals vs. LLM-as-a-judge
03:04 Why agents make evaluation harder
05:11 Capability evals vs. regression evals
05:37 Recap and what's next
π Sign up for free: app.arize.com/auth/login?utm_source=youtube&utm_medium=organic_social&utm_campaign=arize_ax_getting_started
π Learn more about Arize AX: arize.com
π Colab notebook for the series: colab.research.google.com/drive/1dViThD0kJjbqtDGE7-ciBYIRwV-HBN1T
π Docs: docs.arize.com
π Subscribe for weekly content on LLMs, agents, and evaluation: youtube.com/@arizeai?sub_confirmation=1
#ArizeAX #AgentObservability #LLMEvaluation



