LLM-as-a-Judge for Agents: How to Build a Custom Eval Rubric That Works | Ep. 7 @arizeai
LLM-as-a-Judge for Agents: How to Build a Custom Eval Rubric That Works | Ep. 7  @arizeai
Uploaded July 2026 | Updated September 2026, 2 weeks ago
Learn how to build an LLM-as-a-judge eval that gives you signal you can actually trust. When there's no single right answer, you need an LLM to grade your LLM, and a rubric that's specific enough to be useful.

A vague rubric gets you a vague judge. The fix is concrete criteria, real examples, and a judge that reasons before it labels.

Watch this to learn:
• The anatomy of an LLM judge: judge model, rubric, and criteria
• When built-in evals like faithfulness work, and when they can't
• How to write a custom "actionability" judge, with specific criteria, concrete examples, and labels and scores defined in config
• Why the judge should reason before it gives a label

Chapters:
00:00 From code evals to LLM-as-a-judge
00:24 Anatomy of an LLM judge
01:02 Trying the built-in faithfulness eval
01:53 Setting up the judge, Sonnet grading Haiku
02:55 Writing a custom actionability judge
04:00 Rubrics: be specific, add examples
06:10 Define labels and scores in configuration
07:00 Make the judge reason before it labels
07:39 Reading the failing explanations

👉 Sign up for free: app.arize.com/auth/login?utm_source=youtube&utm_medium=organic_social&utm_campaign=arize_ax_getting_started
🔗 Learn more about Arize AX: arize.com
📓 Colab notebook for the series: colab.research.google.com/drive/1dViThD0kJjbqtDGE7-ciBYIRwV-HBN1T
📚 Docs: docs.arize.com
🔔 Subscribe for weekly content on LLMs, agents, and evaluation: youtube.com/@arizeai?sub_confirmation=1
#ArizeAX #LLMasaJudge #LLMEvaluation
LLM-as-a-Judge for Agents: How to Build a Custom Eval Rubric That Works | Ep. 7Inside Typeforms AI Agent StackWhy AI Agents Need Their Own Observability Layer | AWS | Arize Observe 2026Testing Self-Evaluation Bias of LLMsUpstart’s First AI Voice Bot: Lessons From Production | Shiv Indap | Arize Observe 2026Analyzing LLM Evaluations of Customer Reviews Using Repetitions FeatureHow to Build Self-Improving AI Agents with Coding Agents | Ep. 13How Uber Evaluates AI Agents at Production Scale | Arize Observe 2026Identity, Permissions, and Security for AI Agents | WorkOS | Arize Observe 2026How My AI Agent Rewrites Itself Overnight | Chi Wang, AG2 | Arize Observe 2026Traces and Evals Explained: The Building Blocks of AI and Agent Testing | Ep. 2Nebulocks Ron Cahlon on Building AI for Cybersecurity
Arize AI |

LLM-as-a-Judge for Agents: How to Build a Custom Eval Rubric That Works | Ep. 7

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER