LLM-as-a-Judge 101 @arizeai
LLM-as-a-Judge 101  @arizeai
Uploaded November 2025 | Updated September 2026, 2 weeks ago
Curious about AI evals, but not sure where to start? In this hands-on, beginner-friendly session, we walk you through the core building blocks of LLM-as-a-judge evaluations.

You’ll learn how to design your first evaluation from scratch, including:

➡️ What to measure: Understand the key qualities of a good metric and identify the specific criteria that will provide the most actionable insights into your application.
➡️ ​Which model to use: Learn how to choose the right judge model for your needs—whether you're optimizing for cost and speed or maximum quality.
➡️ ​How to prompt effectively: See examples of prompt formats that yield consistent, interpretable results, with tips on avoiding common pitfalls.
➡️ ​How to improve your eval: Learn how to perform meta-evaluation, conduct error analysis, and iteratively refine your prompts for stronger insights.

​This session is led by industry experts who have hands-on experience evaluating real-world AI applications and are deeply familiar with the latest research. You'll walk away with practical guidelines and a clear mental model for how to structure evaluations.

Follow Elizabeth Hutton: linkedin.com/in/elizabeth-hutton
Follow Sri Chavali: linkedin.com/in/srilakshmi-chavali

Learn more about LLM-as-a-judge evaluation: arize.com/llm-as-a-judge
LLM-as-a-Judge 101Improving Agents in Production with Online Evals - Arize AXTypeScript Agents: How To Build and EvaluateOne AI Question - where do agents fail in production, with Fuad AliArize Skills: Add Instrumentation & Tracing to Your AI App with Claude Code, Copilot, or CursorIntroduction To Arize AX EvalsAI Agent Mastery Certification Course: Module 2 – Agent Engineering & ObservabilityYour First Code Eval for Agents: Catch Bugs in 5 Lines of Python | Ep. 6The Flaw in Most AI Evaluation VendorsHow Microsoft’s Azure AI Foundry Builds Trustworthy AI Agents, with PhoenixAI Agent Mastery Certification Course: Lab 4 – Tools & MCPLLM-as-a-Judge for Agents: How to Build a Custom Eval Rubric That Works | Ep. 7
Arize AI |

LLM-as-a-Judge 101

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER