Uploaded May 2026 | Updated September 2026, 3 weeks ago
Most AI features ship on vibes: run it three times, looks fine, push to prod. Then it breaks on inputs you didn't test. In this 25-minute talk, Laurie Voss (Head of DevRel at Arize AI) lays out a complete, practical framework for evaluating LLM agents, from your first trace to a production-grade eval suite.
You'll learn:
- Why traditional unit tests don't work for non-deterministic AI outputs
- The three types of evals: code, LLM-as-a-judge, and human — and when to use each
- Capability evals vs. regression evals (and why this is the most useful mental model in AI testing)
- Eval-driven development: TDD for AI agents
- The 5 parts of an LLM-as-a-judge prompt that actually works
- Why you must validate your judge against a golden dataset — and how to measure precision and recall
- Common LLM judge pitfalls: length bias, self-preference, and how to avoid them
- How to turn eval results into a flywheel that compounds into a competitive moat
Whether you're building your first agent or trying to bolt evals onto a system already in production, this is the roadmap.
🔗 Arize AX: arize.com
🎟️ Arize Observe Conference — June 4, San Francisco. Use the code on screen for $50 tickets!
These slides: slides.com/seldo/stop-vibe-testing
Most AI features ship on vibes: run it three times, looks fine, push to prod. Then it breaks on inputs you didn't test. In this 25-minute talk, Laurie Voss (Head of DevRel at Arize AI) lays out a complete, practical framework for evaluating LLM agents, from your first trace to a production-grade eval suite.
You'll learn:
- Why traditional unit tests don't work for non-deterministic AI outputs
- The three types of evals: code, LLM-as-a-judge, and human — and when to use each
- Capability evals vs. regression evals (and why this is the most useful mental model in AI testing)
- Eval-driven development: TDD for AI agents
- The 5 parts of an LLM-as-a-judge prompt that actually works
- Why you must validate your judge against a golden dataset — and how to measure precision and recall
- Common LLM judge pitfalls: length bias, self-preference, and how to avoid them
- How to turn eval results into a flywheel that compounds into a competitive moat
Whether you're building your first agent or trying to bolt evals onto a system already in production, this is the roadmap.
🔗 Arize AX: arize.com
🎟️ Arize Observe Conference — June 4, San Francisco. Use the code on screen for $50 tickets!
These slides: slides.com/seldo/stop-vibe-testing










