Uploaded December 2025 | Updated September 2026, 2 weeks ago
Ankur Goyal and Malte Ubl, co-founders of Braintrust and Vercel respectively, join swyx for a spirited debate about the role of evaluations (evals) in building AI coding agents. Sparked by a viral clip of Anthropic's Boris Power stating that Claude's $500M+ coding agent business was built largely on "vibes" rather than traditional evals, this conversation digs into whether offline evals are essential infrastructure or premature optimization—and why the best teams are deliberate about investing in multiple feedback loops.
We discuss:
* Why *feedback loops matter more than the term "evals"*—from offline evals to A/B tests to pure vibe checks, and how the best teams deliberately invest in all three
* The evolution from *golden datasets to production-driven evals:* why manufacturing test cases upfront is a waste, and how top teams now pull real user failures from logs into their eval suites daily
* How *evals enable velocity:* knowing your "first derivative" (whether a change improves things) lets you ship aggressively without fear of regression, just like unit tests in traditional software
* Why *coding evals are uniquely verifiable* yet still underutilized: from "does it compile?" to "does it render without errors?" and how Vercel uses these signals in RL pipelines to fine-tune models that fix trivial errors 100x faster than agentic loops
* *Evals as product management:* how rubrics and LLM-as-judge scoring functions let product managers encode domain expertise (finance, healthcare) more precisely than 50-page PRDs, and why PMs are now deeply involved in eval design
* The *privilege of AI labs* that build evals in-house versus the rest of the world that must make do with frontier models, and why proprietary evals are competitive moats while public benchmarks are marketing
* *RL environments as the next frontier:* why they're powerful for computer-use agents and decoupling from expensive human labeling, but require specialized expertise to avoid reward hacking
* An *inversion of control* for evals: why companies like Vercel should publish Next.js evals so model labs can optimize for their frameworks, creating a new marketplace where eval creators aren't the same entities training models
* How *vibes are also evals*—just extraordinarily accurate, expensive scoring functions—and why the goal isn't choosing between evals and vibes but building complementary feedback loops at different speeds and costs
* Practical wins from Braintrust + Vercel integration: one-click AI trace logging from Vercel apps to Braintrust, and how Vercel's composite model architecture (frontier draft + fine-tuned fix) cuts latency by orders of magnitude
—
Ankur Goyal
* X: https://x.com/ankrgyl
* LinkedIn: linkedin.com/in/ankurgoyal
Malte Ubl
* X: https://x.com/cramforce
* LinkedIn: linkedin.com/in/malteubl
Where to find Latent Space
* X: https://x.com/latentspacepod
* Substack: https://www.latent.space/
00:00:00 Introduction: The Great Evals Debate
00:00:59 Background and Context: From Google Search to AI Engineering
00:02:35 The Effort-Efficiency Tradeoff in Evals
00:03:45 Why Coding is Different: Verifiability and Privilege
00:05:19 Public Benchmarks vs Internal Evals
00:08:45 Open-Endedness and the Limits of Offline Evals
00:11:26 The Workflow: From Production Logs to Offline Testing
00:12:05 When Vibes and Data Disagree
00:18:08 Evals as Product Management Tool
00:22:45 RL Environments and the Future of Evals
00:26:46 Inversion of Control: Who Should Write Evals?
00:33:36 Wrap-up and Vercel-BrainTrust Integration
Ankur Goyal and Malte Ubl, co-founders of Braintrust and Vercel respectively, join swyx for a spirited debate about the role of evaluations (evals) in building AI coding agents. Sparked by a viral clip of Anthropic's Boris Power stating that Claude's $500M+ coding agent business was built largely on "vibes" rather than traditional evals, this conversation digs into whether offline evals are essential infrastructure or premature optimization—and why the best teams are deliberate about investing in multiple feedback loops.
We discuss:
* Why *feedback loops matter more than the term "evals"*—from offline evals to A/B tests to pure vibe checks, and how the best teams deliberately invest in all three
* The evolution from *golden datasets to production-driven evals:* why manufacturing test cases upfront is a waste, and how top teams now pull real user failures from logs into their eval suites daily
* How *evals enable velocity:* knowing your "first derivative" (whether a change improves things) lets you ship aggressively without fear of regression, just like unit tests in traditional software
* Why *coding evals are uniquely verifiable* yet still underutilized: from "does it compile?" to "does it render without errors?" and how Vercel uses these signals in RL pipelines to fine-tune models that fix trivial errors 100x faster than agentic loops
* *Evals as product management:* how rubrics and LLM-as-judge scoring functions let product managers encode domain expertise (finance, healthcare) more precisely than 50-page PRDs, and why PMs are now deeply involved in eval design
* The *privilege of AI labs* that build evals in-house versus the rest of the world that must make do with frontier models, and why proprietary evals are competitive moats while public benchmarks are marketing
* *RL environments as the next frontier:* why they're powerful for computer-use agents and decoupling from expensive human labeling, but require specialized expertise to avoid reward hacking
* An *inversion of control* for evals: why companies like Vercel should publish Next.js evals so model labs can optimize for their frameworks, creating a new marketplace where eval creators aren't the same entities training models
* How *vibes are also evals*—just extraordinarily accurate, expensive scoring functions—and why the goal isn't choosing between evals and vibes but building complementary feedback loops at different speeds and costs
* Practical wins from Braintrust + Vercel integration: one-click AI trace logging from Vercel apps to Braintrust, and how Vercel's composite model architecture (frontier draft + fine-tuned fix) cuts latency by orders of magnitude
—
Ankur Goyal
* X: https://x.com/ankrgyl
* LinkedIn: linkedin.com/in/ankurgoyal
Malte Ubl
* X: https://x.com/cramforce
* LinkedIn: linkedin.com/in/malteubl
Where to find Latent Space
* X: https://x.com/latentspacepod
* Substack: https://www.latent.space/
00:00:00 Introduction: The Great Evals Debate
00:00:59 Background and Context: From Google Search to AI Engineering
00:02:35 The Effort-Efficiency Tradeoff in Evals
00:03:45 Why Coding is Different: Verifiability and Privilege
00:05:19 Public Benchmarks vs Internal Evals
00:08:45 Open-Endedness and the Limits of Offline Evals
00:11:26 The Workflow: From Production Logs to Offline Testing
00:12:05 When Vibes and Data Disagree
00:18:08 Evals as Product Management Tool
00:22:45 RL Environments and the Future of Evals
00:26:46 Inversion of Control: Who Should Write Evals?
00:33:36 Wrap-up and Vercel-BrainTrust Integration










![⚡️Context Graphs: according to the authors — Jaya Gupta, Ashu Garg, Foundation Capital
In this Lightning pod, swyx hosts Jaya Gupta and Ashu Garg from Foundation Capital to discuss the emergence of context graphs. They define this new framework as the institutional memory of the why behind business decisions, captured through decision traces—the sequence of steps and human reasoning that models often miss. The conversation explores how these graphs will become the defensible moat for the next generation of applied AI companies and systems of agents.
Section Timestamps
[00:03] – Introductions and the early vibe of AI hackathons post-ChatGPT.
[02:00] – The origin story of the Context Graph thesis at Foundation Capital.
[04:59] – Defining the Context Graph and the Decision Trace.
[07:32] – Who is building this today? Examples like Player Zero and Glean.
[09:37] – Technical implementation: Is there an ideal data structure?
[12:09] – Explaining Systems of Agents vs. standard chatbots.
[15:26] – The importance of the Right Path (operational) vs. the Read Path (analytical).
[18:46] – Why these will be new platforms rather than features in Slack or GitHub.
[21:48] – Addressing pushbacks: Can you truly capture the Why or just the How?
[24:04] – Privacy, data governance, and Metadata 3.0.
[26:43] – Context Graphs vs. Data Mesh: Why universal graphs are unlikely.
[31:18] – 2026 Predictions: The Context Graph Stack and production scale.
show notes
https://foundationcapital.com/context-graphs-ais-trillion-dollar-opportunity/
https://x.com/JayaGup10/status/2003525933534179480
https://simple.ai/p/what-are-context-graphs ⚡️Context Graphs: according to the authors — Jaya Gupta, Ashu Garg, Foundation Capital](https://i.ytimg.com/vi/zP8P7hJXwE0/mqdefault.jpg)