Uploaded July 2026 | Updated September 2026, 2 weeks ago
Join us for the final session of our Agentic AI webinar series, where you'll learn how leading AI teams evaluate, test, and monitor AI agents before deploying them to production. Discover the frameworks, tools, and best practices that help transform promising demos into reliable production systems.
๐ก What we'll cover:
Why traditional software testing, single LLM evaluations, and multi-agent evaluations require different strategies
The four essential evaluator types: rule-based, LLM-as-a-judge, trajectory evaluation, and recovery-from-failure
Building evaluation scorecards and regression testing workflows for every prompt, model, tool, or architecture update
State-of-the-art agent evaluation workflows using LangSmith, plus open-source alternatives with Langfuse and OpenTelemetry
Online evaluation, telemetry, and pass^k reliability for production-ready AI agents
Best practices for continuously monitoring and improving agent performance after deployment
๐ Hands-on demonstration included:
Watch a live multi-agent system run through real production traces and telemetry while learning how modern AI teams evaluate agent behavior, identify failures, and validate performance using production-grade evaluation frameworks.
Perfect for AI engineers, ML engineers, developers, platform teams, and technical leaders building, deploying, and scaling AI agents in production.
-----------
๐ Learn more about Data Science Dojo here:
datasciencedojo.com
๐ Watch the latest video tutorials here:
datasciencedojo.com/tutorials
๐ See what our past attendees are saying here:
https://datasciencedojo.com/data-scie...
--
At Data Science Dojo, we believe data science is for everyone. Our in-person data science training has been attended by more than 8000+ employees from over 2000+ companies globally, including many leaders in tech like Microsoft, Apple, and Facebook.
--
๐ Subscribe to our newsletter for data science content & infographics: datasciencedojo.com/newsletter
Join us for the final session of our Agentic AI webinar series, where you'll learn how leading AI teams evaluate, test, and monitor AI agents before deploying them to production. Discover the frameworks, tools, and best practices that help transform promising demos into reliable production systems.
๐ก What we'll cover:
Why traditional software testing, single LLM evaluations, and multi-agent evaluations require different strategies
The four essential evaluator types: rule-based, LLM-as-a-judge, trajectory evaluation, and recovery-from-failure
Building evaluation scorecards and regression testing workflows for every prompt, model, tool, or architecture update
State-of-the-art agent evaluation workflows using LangSmith, plus open-source alternatives with Langfuse and OpenTelemetry
Online evaluation, telemetry, and pass^k reliability for production-ready AI agents
Best practices for continuously monitoring and improving agent performance after deployment
๐ Hands-on demonstration included:
Watch a live multi-agent system run through real production traces and telemetry while learning how modern AI teams evaluate agent behavior, identify failures, and validate performance using production-grade evaluation frameworks.
Perfect for AI engineers, ML engineers, developers, platform teams, and technical leaders building, deploying, and scaling AI agents in production.
-----------
๐ Learn more about Data Science Dojo here:
datasciencedojo.com
๐ Watch the latest video tutorials here:
datasciencedojo.com/tutorials
๐ See what our past attendees are saying here:
https://datasciencedojo.com/data-scie...
--
At Data Science Dojo, we believe data science is for everyone. Our in-person data science training has been attended by more than 8000+ employees from over 2000+ companies globally, including many leaders in tech like Microsoft, Apple, and Facebook.
--
๐ Subscribe to our newsletter for data science content & infographics: datasciencedojo.com/newsletter










