How to Run LLM Evals on Live Production Traffic with your Agents | Ep. 11 @arizeai
How to Run LLM Evals on Live Production Traffic with your Agents | Ep. 11  @arizeai
Uploaded July 2026 | Updated September 2026, 3 weeks ago
Learn how to grade your agent on real production traffic, not just the test cases you thought of. Real users do things your test data never imagined, and today's live failure becomes tomorrow's test case.

Online evals reuse the same evaluators you already wrote. You don't need to grade every single request either: the right sampling gives you representative coverage without the cost.

Watch this to learn:
β€’ What online evals are and how they reuse the evaluators you already wrote
β€’ Evaluation scopes (span, trace, session) and how sampling gives representative coverage
β€’ How to describe evaluators in plain English, and set one up as an online eval
β€’ Why your eval suite becomes a competitive advantage over time

Chapters:
00:00 From test data to real users
00:41 What online evals are
00:55 Scopes: span, trace, session
01:07 Sampling for representative coverage
01:34 Setting up an online eval
02:44 Describe evaluators in plain English
03:11 The loop that closes
04:04 Your eval suite as a competitive advantage

πŸ‘‰ Sign up for free: app.arize.com/auth/login?utm_source=youtube&utm_medium=organic_social&utm_campaign=arize_ax_getting_started
πŸ”— Learn more about Arize AX: arize.com
πŸ““ Colab notebook for the series: colab.research.google.com/drive/1dViThD0kJjbqtDGE7-ciBYIRwV-HBN1T
πŸ“š Docs: docs.arize.com
πŸ”” Subscribe for weekly content on LLMs, agents, and evaluation: youtube.com/@arizeai?sub_confirmation=1

#ArizeAX #OnlineEvals #AgentObservability
How to Run LLM Evals on Live Production Traffic with your Agents | Ep. 11How to Set Up LLM Tracing in Minutes with OpenTelemetry | Ep. 3How TripAdvisor Debugs AI Agents at ScaleLeveling Up AI Agents with LLM Evaluations, Feedback Loops and Context EngineeringFrom User Feedback to Code Pull Requests: The AI Flywheel5 LLM and Agent Eval Mistakes That Turn Metrics Into Noise | Ep. 8Kubernetes Is Not Your Sandbox: Building Infrastructure for AI Agents | Daytona | Arize Observe 2026AI Agent Mastery Certification Course: Module 5 – RAG & Agentic RAGAI’s Next Wave: What VCs Are Betting On in 2026 | Jaya Gupta | Arize Observe 2026How DeepSeek is Pushing the Boundaries of AI DevelopmentUsing Annotations to Build an Eval-Driven LLM Development PipelineServiceNow’s AgentArch: Benchmarking AI Agents for Enterprise Workflows
Arize AI |

How to Run LLM Evals on Live Production Traffic with your Agents | Ep. 11

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER