Uploaded February 2026 | Updated September 2026, 2 weeks ago
This workshop is an evaluation 101 primer for teams who want the latest, research-driven best practices for doing evals well across the agent lifecycle. We cover evaluation fundamentals, dig into the practical tradeoffs of LLM-as-a-Judge, and walk through a concrete code-based eval example. It wraps with an Arize AX demo highlighting how evals connect to other important workflows.
This workshop is an evaluation 101 primer for teams who want the latest, research-driven best practices for doing evals well across the agent lifecycle. We cover evaluation fundamentals, dig into the practical tradeoffs of LLM-as-a-Judge, and walk through a concrete code-based eval example. It wraps with an Arize AX demo highlighting how evals connect to other important workflows.










