The Future of Evals: From LLM as a Judge to Agent as a Judge — Aparna Dhinakaran, Arize AI @aiDotEngineer
The Future of Evals: From LLM as a Judge to Agent as a Judge — Aparna Dhinakaran, Arize AI  @aiDotEngineer
Uploaded July 2026 | Updated September 2026, 3 weeks ago
Across a dozen eval jobs Arize watches the top teams run, one pattern holds: the eval has to change as fast as the agent it grades. In 2023 an agent was barely more than a prompt; since then reasoning, tool calls, and long multi step loops piled on, and every jump in capability quietly broke the eval that came before. So the evals evolved with them. Deterministic checks catch what you can define up front, LLM as a judge adds the analysis a fixed rule cannot, and the newest step, agent as a judge, hunts for failure modes you would never think to write a check for and can open a pull request to fix what it finds. Aparna Dhinakaran's argument is that this arc, from static checks to an agent grading another agent, is where evals go next.

Speaker info:
- https://x.com/aparnadhinak
- linkedin.com/in/aparnadhinakaran

Timestamps:
0:00 - Opening: the future of the Evals track
2:06 - Why evals got harder as agents evolved
3:45 - From deterministic checks to LLM as a judge
4:36 - Agent as a judge, and where evals go next
The Future of Evals: From LLM as a Judge to Agent as a Judge — Aparna Dhinakaran, Arize AIWhen AI Agents Pay and Sellers Monetize: Building x402 Apps on AWS — Anil Nadiminti, AWSHow to Generate Mergeable Code with a Context Engine — Peter Werry, UnblockedHow Anthropic Builds: Lessons from Labs — Mike Krieger, AnthropicGive the Agent a Budget, Not a Token — Sachin Malhotra, AnthropicAutomating Large Scale Refactors with Parallel Agents - Robert Brennan, OpenHandsThe Last Human Code Review: Building Trust in AI-Generated Code — Itamar Friedman, QodoMCP Tasks (async): Why Arent Any Agents Supporting Them? — Cornelia Davis, TemporalWearing the Agent: From Group Chats to Glasses — Sai Krishna RallabandiThe Half Life of Agent Infrastructure — Ben Kus, BoxCodex, Behind the Harness — Dominik Kundel, OpenAIThe Agentic Web and the Bazaar Era of AI - Ramesh Raskar, MIT Media Lab
AI Engineer |

The Future of Evals: From LLM as a Judge to Agent as a Judge — Aparna Dhinakaran, Arize AI

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER