Uploaded October 2025 | Updated September 2026, 2 weeks ago
About a year ago, Arize AI released some early research on how reliable foundational models were at LLM-as-a-Judge when the output was a binary vs score eval. The results were very clear - binary evals were the way to go. It’s been over a year and the models are getting better. Does the research still hold?
In this session, Elizabeth Hutton (Senior AI Engineer) and Srilakshmi Chavali (AI Engineer) dive into findings from newly released research.
See the full writeup on the research here: arize.com/blog/testing-binary-vs-score-llm-evals-on-the-latest-models
Code to replicate the tests yourself: github.com/Arize-ai/LLMTest_NeedleInAHaystack/blob/main/ScoreEvalsTest.py
About a year ago, Arize AI released some early research on how reliable foundational models were at LLM-as-a-Judge when the output was a binary vs score eval. The results were very clear - binary evals were the way to go. It’s been over a year and the models are getting better. Does the research still hold?
In this session, Elizabeth Hutton (Senior AI Engineer) and Srilakshmi Chavali (AI Engineer) dive into findings from newly released research.
See the full writeup on the research here: arize.com/blog/testing-binary-vs-score-llm-evals-on-the-latest-models
Code to replicate the tests yourself: github.com/Arize-ai/LLMTest_NeedleInAHaystack/blob/main/ScoreEvalsTest.py










