Is Your LLM Judge Right? Calibrate with Meta-Evaluation | Ep. 9 @arizeai
Is Your LLM Judge Right? Calibrate with Meta-Evaluation | Ep. 9  @arizeai
Uploaded July 2026 | Updated September 2026, 3 weeks ago
Learn how to check that your LLM judge is actually trustworthy, not just confident. Your judge grades your agent, but if nobody grades the judge, you're trusting a number you haven't verified.

Judges have their own failure modes: position bias, length bias, self-preference. The fix isn't a perfect judge, it's one that agrees with humans often enough to trust.

Watch this to learn:
โ€ข How to build a golden set of human annotations and split it into dev and test sets
โ€ข How to compare judge vs. human, and why you often prioritize recall over precision here
โ€ข LLM judge biases to watch for (position, length, self-preference)
โ€ข Why human agreement, not perfection, is the real benchmark

Chapters:
00:00 Meta-evaluation: is your judge right?
00:52 Adding human annotations in AX
01:31 Building a golden set
02:34 Dev and test splits
03:00 Comparing judge vs. human
04:15 Precision and recall, why recall wins here
05:19 LLM judge biases to watch for
06:23 Benchmark against human agreement
07:12 When the eval is broken, not the agent
08:03 Next: improving the agent

๐Ÿ‘‰ Sign up for free: app.arize.com/auth/login?utm_source=youtube&utm_medium=organic_social&utm_campaign=arize_ax_getting_started
๐Ÿ”— Learn more about Arize AX: arize.com
๐Ÿ““ Colab notebook for the series: colab.research.google.com/drive/1dViThD0kJjbqtDGE7-ciBYIRwV-HBN1T
๐Ÿ“š Docs: docs.arize.com
๐Ÿ”” Subscribe for weekly content on LLMs, agents, and evaluation: youtube.com/@arizeai?sub_confirmation=1

#ArizeAX #MetaEvaluation #LLMasaJudge
Is Your LLM Judge Right? Calibrate with Meta-Evaluation | Ep. 9Building and Scaling ProductsElastic AI - Walking Your Way to PhoenixMeet PXI: the AI engineering agent inside PhoenixHow to Build a Real AI Agent (Financial Analyst) with the Claude Agent SDK | Ep. 4Stop Vibe-Testing Your AI Agents: How to Actually Run Evals (in 25 Minutes)Stop Blaming the Model: Fixing the AI Product Bottleneck | Rise of the AI Engineer | Hamel HusainFrom Build to Production: Engineering Reliable AI Agents with Google and ArizeHow Cursor Uses AI Agents to Build Cursor | Arize Observe 2026Arc Prize  - Measuring AGIHow Tripadvisor Runs AI Agents in Production with Arize AXTrunk Tools - Rise of the Agent Engineer
Arize AI |

Is Your LLM Judge Right? Calibrate with Meta-Evaluation | Ep. 9

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER