Uploaded July 2026 | Updated September 2026, 3 weeks ago
Learn how to check that your LLM judge is actually trustworthy, not just confident. Your judge grades your agent, but if nobody grades the judge, you're trusting a number you haven't verified.
Judges have their own failure modes: position bias, length bias, self-preference. The fix isn't a perfect judge, it's one that agrees with humans often enough to trust.
Watch this to learn:
โข How to build a golden set of human annotations and split it into dev and test sets
โข How to compare judge vs. human, and why you often prioritize recall over precision here
โข LLM judge biases to watch for (position, length, self-preference)
โข Why human agreement, not perfection, is the real benchmark
Chapters:
00:00 Meta-evaluation: is your judge right?
00:52 Adding human annotations in AX
01:31 Building a golden set
02:34 Dev and test splits
03:00 Comparing judge vs. human
04:15 Precision and recall, why recall wins here
05:19 LLM judge biases to watch for
06:23 Benchmark against human agreement
07:12 When the eval is broken, not the agent
08:03 Next: improving the agent
๐ Sign up for free: app.arize.com/auth/login?utm_source=youtube&utm_medium=organic_social&utm_campaign=arize_ax_getting_started
๐ Learn more about Arize AX: arize.com
๐ Colab notebook for the series: colab.research.google.com/drive/1dViThD0kJjbqtDGE7-ciBYIRwV-HBN1T
๐ Docs: docs.arize.com
๐ Subscribe for weekly content on LLMs, agents, and evaluation: youtube.com/@arizeai?sub_confirmation=1
#ArizeAX #MetaEvaluation #LLMasaJudge
Learn how to check that your LLM judge is actually trustworthy, not just confident. Your judge grades your agent, but if nobody grades the judge, you're trusting a number you haven't verified.
Judges have their own failure modes: position bias, length bias, self-preference. The fix isn't a perfect judge, it's one that agrees with humans often enough to trust.
Watch this to learn:
โข How to build a golden set of human annotations and split it into dev and test sets
โข How to compare judge vs. human, and why you often prioritize recall over precision here
โข LLM judge biases to watch for (position, length, self-preference)
โข Why human agreement, not perfection, is the real benchmark
Chapters:
00:00 Meta-evaluation: is your judge right?
00:52 Adding human annotations in AX
01:31 Building a golden set
02:34 Dev and test splits
03:00 Comparing judge vs. human
04:15 Precision and recall, why recall wins here
05:19 LLM judge biases to watch for
06:23 Benchmark against human agreement
07:12 When the eval is broken, not the agent
08:03 Next: improving the agent
๐ Sign up for free: app.arize.com/auth/login?utm_source=youtube&utm_medium=organic_social&utm_campaign=arize_ax_getting_started
๐ Learn more about Arize AX: arize.com
๐ Colab notebook for the series: colab.research.google.com/drive/1dViThD0kJjbqtDGE7-ciBYIRwV-HBN1T
๐ Docs: docs.arize.com
๐ Subscribe for weekly content on LLMs, agents, and evaluation: youtube.com/@arizeai?sub_confirmation=1
#ArizeAX #MetaEvaluation #LLMasaJudge










