The Great Evals Debate — Ankur Goyal & Malte Ubl @LatentSpacePod
The Great Evals Debate — Ankur Goyal & Malte Ubl  @LatentSpacePod
Uploaded December 2025 | Updated September 2026, 2 weeks ago
Ankur Goyal and Malte Ubl, co-founders of Braintrust and Vercel respectively, join swyx for a spirited debate about the role of evaluations (evals) in building AI coding agents. Sparked by a viral clip of Anthropic's Boris Power stating that Claude's $500M+ coding agent business was built largely on "vibes" rather than traditional evals, this conversation digs into whether offline evals are essential infrastructure or premature optimization—and why the best teams are deliberate about investing in multiple feedback loops.

We discuss:

* Why *feedback loops matter more than the term "evals"*—from offline evals to A/B tests to pure vibe checks, and how the best teams deliberately invest in all three
* The evolution from *golden datasets to production-driven evals:* why manufacturing test cases upfront is a waste, and how top teams now pull real user failures from logs into their eval suites daily
* How *evals enable velocity:* knowing your "first derivative" (whether a change improves things) lets you ship aggressively without fear of regression, just like unit tests in traditional software
* Why *coding evals are uniquely verifiable* yet still underutilized: from "does it compile?" to "does it render without errors?" and how Vercel uses these signals in RL pipelines to fine-tune models that fix trivial errors 100x faster than agentic loops
* *Evals as product management:* how rubrics and LLM-as-judge scoring functions let product managers encode domain expertise (finance, healthcare) more precisely than 50-page PRDs, and why PMs are now deeply involved in eval design
* The *privilege of AI labs* that build evals in-house versus the rest of the world that must make do with frontier models, and why proprietary evals are competitive moats while public benchmarks are marketing
* *RL environments as the next frontier:* why they're powerful for computer-use agents and decoupling from expensive human labeling, but require specialized expertise to avoid reward hacking
* An *inversion of control* for evals: why companies like Vercel should publish Next.js evals so model labs can optimize for their frameworks, creating a new marketplace where eval creators aren't the same entities training models
* How *vibes are also evals*—just extraordinarily accurate, expensive scoring functions—and why the goal isn't choosing between evals and vibes but building complementary feedback loops at different speeds and costs
* Practical wins from Braintrust + Vercel integration: one-click AI trace logging from Vercel apps to Braintrust, and how Vercel's composite model architecture (frontier draft + fine-tuned fix) cuts latency by orders of magnitude



Ankur Goyal

* X: https://x.com/ankrgyl
* LinkedIn: linkedin.com/in/ankurgoyal

Malte Ubl

* X: https://x.com/cramforce
* LinkedIn: linkedin.com/in/malteubl

Where to find Latent Space

* X: https://x.com/latentspacepod
* Substack: https://www.latent.space/

00:00:00 Introduction: The Great Evals Debate
00:00:59 Background and Context: From Google Search to AI Engineering
00:02:35 The Effort-Efficiency Tradeoff in Evals
00:03:45 Why Coding is Different: Verifiability and Privilege
00:05:19 Public Benchmarks vs Internal Evals
00:08:45 Open-Endedness and the Limits of Offline Evals
00:11:26 The Workflow: From Production Logs to Offline Testing
00:12:05 When Vibes and Data Disagree
00:18:08 Evals as Product Management Tool
00:22:45 RL Environments and the Future of Evals
00:26:46 Inversion of Control: Who Should Write Evals?
00:33:36 Wrap-up and Vercel-BrainTrust Integration
The Great Evals Debate — Ankur Goyal & Malte UblScaling Simulations for AI ImprovementAI Studio: Simplifying Model SelectionClaude Code for Finance + The Global Memory Shortage: Doug OLaughlin, SemiAnalysisUnlocking the Immune System A Deep Dive into BiologyAI Life Coach  A Simple LLM ModelData to Insight: The Full JourneyWhy RL Won — Kyle Corbitt, OpenPipe (acq. CoreWeave)Neural Nets My AI Pill Moment with Chris Olah⚡️ How to turn Documents into Knowledge: Graphs in Modern AI — Emil Eifrem, CEO Neo4JOne Year of MCP — with David Soria Parria and AAIF leads from OpenAI, Goose, Linux Foundation⚡️Context Graphs: according to the authors — Jaya Gupta, Ashu Garg, Foundation Capital
Latent Space |

The Great Evals Debate — Ankur Goyal & Malte Ubl

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER