Uploaded July 2026 | Updated September 2026, 3 weeks ago
Learn how to grade your agent on real production traffic, not just the test cases you thought of. Real users do things your test data never imagined, and today's live failure becomes tomorrow's test case.
Online evals reuse the same evaluators you already wrote. You don't need to grade every single request either: the right sampling gives you representative coverage without the cost.
Watch this to learn:
β’ What online evals are and how they reuse the evaluators you already wrote
β’ Evaluation scopes (span, trace, session) and how sampling gives representative coverage
β’ How to describe evaluators in plain English, and set one up as an online eval
β’ Why your eval suite becomes a competitive advantage over time
Chapters:
00:00 From test data to real users
00:41 What online evals are
00:55 Scopes: span, trace, session
01:07 Sampling for representative coverage
01:34 Setting up an online eval
02:44 Describe evaluators in plain English
03:11 The loop that closes
04:04 Your eval suite as a competitive advantage
π Sign up for free: app.arize.com/auth/login?utm_source=youtube&utm_medium=organic_social&utm_campaign=arize_ax_getting_started
π Learn more about Arize AX: arize.com
π Colab notebook for the series: colab.research.google.com/drive/1dViThD0kJjbqtDGE7-ciBYIRwV-HBN1T
π Docs: docs.arize.com
π Subscribe for weekly content on LLMs, agents, and evaluation: youtube.com/@arizeai?sub_confirmation=1
#ArizeAX #OnlineEvals #AgentObservability
Learn how to grade your agent on real production traffic, not just the test cases you thought of. Real users do things your test data never imagined, and today's live failure becomes tomorrow's test case.
Online evals reuse the same evaluators you already wrote. You don't need to grade every single request either: the right sampling gives you representative coverage without the cost.
Watch this to learn:
β’ What online evals are and how they reuse the evaluators you already wrote
β’ Evaluation scopes (span, trace, session) and how sampling gives representative coverage
β’ How to describe evaluators in plain English, and set one up as an online eval
β’ Why your eval suite becomes a competitive advantage over time
Chapters:
00:00 From test data to real users
00:41 What online evals are
00:55 Scopes: span, trace, session
01:07 Sampling for representative coverage
01:34 Setting up an online eval
02:44 Describe evaluators in plain English
03:11 The loop that closes
04:04 Your eval suite as a competitive advantage
π Sign up for free: app.arize.com/auth/login?utm_source=youtube&utm_medium=organic_social&utm_campaign=arize_ax_getting_started
π Learn more about Arize AX: arize.com
π Colab notebook for the series: colab.research.google.com/drive/1dViThD0kJjbqtDGE7-ciBYIRwV-HBN1T
π Docs: docs.arize.com
π Subscribe for weekly content on LLMs, agents, and evaluation: youtube.com/@arizeai?sub_confirmation=1
#ArizeAX #OnlineEvals #AgentObservability










