Stop Vibe-Testing Your AI Agents: How to Actually Run Evals (in 25 Minutes) @arizeai
Stop Vibe-Testing Your AI Agents: How to Actually Run Evals (in 25 Minutes)  @arizeai
Uploaded May 2026 | Updated September 2026, 3 weeks ago
Most AI features ship on vibes: run it three times, looks fine, push to prod. Then it breaks on inputs you didn't test. In this 25-minute talk, Laurie Voss (Head of DevRel at Arize AI) lays out a complete, practical framework for evaluating LLM agents, from your first trace to a production-grade eval suite.

You'll learn:
- Why traditional unit tests don't work for non-deterministic AI outputs
- The three types of evals: code, LLM-as-a-judge, and human — and when to use each
- Capability evals vs. regression evals (and why this is the most useful mental model in AI testing)
- Eval-driven development: TDD for AI agents
- The 5 parts of an LLM-as-a-judge prompt that actually works
- Why you must validate your judge against a golden dataset — and how to measure precision and recall
- Common LLM judge pitfalls: length bias, self-preference, and how to avoid them
- How to turn eval results into a flywheel that compounds into a competitive moat

Whether you're building your first agent or trying to bolt evals onto a system already in production, this is the roadmap.

🔗 Arize AX: arize.com
🎟️ Arize Observe Conference — June 4, San Francisco. Use the code on screen for $50 tickets!
These slides: slides.com/seldo/stop-vibe-testing
Stop Vibe-Testing Your AI Agents: How to Actually Run Evals (in 25 Minutes)Stop Blaming the Model: Fixing the AI Product Bottleneck | Rise of the AI Engineer | Hamel HusainFrom Build to Production: Engineering Reliable AI Agents with Google and ArizeHow Cursor Uses AI Agents to Build Cursor | Arize Observe 2026Arc Prize  - Measuring AGIHow Tripadvisor Runs AI Agents in Production with Arize AXTrunk Tools - Rise of the Agent EngineerAtropos Health’s Arjun Mukerji, PhD, Explains RWESummaryHow Salesforce Evaluates Multi-Agent AI Systems | Arize Observe 2026Why Software Needs to Be Redesigned for Non-Human UsersHow LG U+ Scales AI Agents for 30M+ Users (Evaluation-Driven Dev)AI Agent for AI Engineers: Alyx Full Demo
Arize AI |

Stop Vibe-Testing Your AI Agents: How to Actually Run Evals (in 25 Minutes)

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER