Latent Space
[LIVE] Anthropic Distillation & How Models Cheat (SWE-Bench Dead) | Nathan Lambert & Sebastian Rasc
updated
Instead of juggling API keys for every service, imagine a unified payment layer where AI agents automatically handle transactions for each tool call
Stripe is exploring MCP payment abstractions, some are looking at stablecoins, others at HTTP 402 payment requests
Still early days but the infrastructure is taking shape. The future of AI commerce might be more autonomous than we think
What's your take on self-paying agents?
Word is they're holding back their coding model until it can beat a specific competitor. Strategic patience or missed opportunity?
The real story: how X built a competitive AI from basically nothing. Data sourcing beyond Twitter, massive compute investment, and the right team.
Most interesting part? Some AGI labs are frustrated with performance while others are celebrating. Same resources, different execution.
The AI race just got more interesting 🚀
Companies are literally cannibalizing their own paid products by offering free MCP integrations
Example: Sentry launched a paid AI agent for issue resolution, but their free MCP lets you do the same thing through Claude/other coding agents
This raises huge questions about monetization strategies in the AI era
Are we heading toward MCP subscription models? Will companies fracture their pricing across different access methods?
The economics don't make sense yet, but the technical capabilities are incredible
Early adopters who understood MCP's potential are now seeing this play out in real-time
His take on AGI benchmarks is fascinating - Arc Prize doesn't aim for PhD-level impossible problems. Their constraint? Can humans do this?
As long as they can create problems humans solve but AI can't, we don't have AGI yet.
Arc 2 is still uncracked. Arc 3 coming 2026.
Grok 4 hit 16% on Arc - everyone's celebrating but we're still far from the finish line.
The path to AGI isn't about the hardest possible problems - it's about human-solvable problems that stump our best AI systems.
The real test? Interactive benchmarks where AI has to intuit goals, make plans, and adapt to novel situations - just like humans do.
That's why static benchmarks won't declare AGI. We need environments that test true generalization, not memorized patterns.
The shift from static to interactive evaluation is where the real intelligence emerges.
Key insights from this clip:
• Model costs trending toward zero (Gemini going free?)
• Price transparency ) obfuscation
• Developers need visibility into costs & models
• ROI on coding agents is "almost hard to calculate"
• Usage-based plans work when users trust the pricing
The real value isn't in running the models - it's in the tooling, transparency, and enabling developers to build things they never would have attempted before.
Thoughts on inference commoditization? 👇
Akshay Agarwal just blew my mind with this Marimo demo. Draw something in MS Paint, it renders in real-time, then gets fed to a multimodal AI model that converts your drawing into a mermaid diagram.
This is the kind of reactive notebook experience that makes you completely rethink what's possible with data workflows. The fact that you can chain together drawing → AI analysis → code generation seamlessly is wild.
Marimo's execution model based on variable dependencies (like Excel) + wide screen layouts = notebooks that actually feel modern.
The AI-native capabilities for model calls are just the cherry on top.
With Marimo, you can literally drag-select data points in a scatter plot and instantly see the underlying images. No complex setup, no painful workarounds - just pure interactivity.
This is what modern data exploration should look like. The notebook reacts to your UI interactions in real-time, sending selected data back to Python automatically.
Best part? Converting from Jupyter is as simple as copy-paste or using their CLI tool.
Data scientists are about to have their minds blown 🚀
His take: "Agent memories are mostly useless"
Why? Because forcing agents to remember specific project quirks or team workflows doesn't actually work the way we think it does. People don't want to manually manage what their AI remembers.
The real insight: Those coding rules and team preferences? Better to put them in guidelines files than rely on agent memory.
But here's the kicker - how do we capture that "tribal knowledge" that agents pick up naturally without making users think about it?
This is the kind of nuanced AI engineering insight that separates the builders from the hype followers 💡
While most AI engineers are building complex agent frameworks with memory systems, compression layers, and elaborate harnesses, Brown believes we're heading toward "thin agents" that won't need all this scaffolding.
But here's the thing - the jury's still out. Current agents need scaffolds to handle memory, context compression, and multi-step reasoning. Will future models just absorb all of this into pure neural computation?
The debate between thick vs thin agents is heating up, and honestly, I could argue both sides.
What's your take? Are we overengineering agent architectures or building the foundation for true AGI?
Cline's founder breaks down the brutal reality of maintaining VS Code forks - Microsoft makes it "notoriously difficult" with constant updates, API changes, and backend improvements that require reverse engineering
Meanwhile, building as an extension gives you:
• Better distribution across editors
• Less maintenance overhead
• Focus on core product vs infrastructure
Sometimes the "obvious" path is a trap. Smart founders know when to build on top vs rebuild from scratch.
@akshayagarwal shows how Marimo lets you "vibe-code" your entire analysis without ever leaving your notebook. The AI sees your variables, database connections, AND column names in real-time.
No more copy-pasting to ChatGPT. No more context switching. Just pure flow state coding.
The future of data analysis is conversational and contextual. This is it.
People won't even practice pronunciation in front of teachers they're paying
But here's the thing - communication and pronunciation are almost completely separate skills
You can have terrible pronunciation and still communicate effectively
The key? Just move your mouth and make sounds. Don't worry about perfection
AI tutors are removing that human judgment barrier that stops people from practicing
Interesting insight from Andrew Struggs on language learning psychology
Official website: three.arcprize.org
X: https://x.com/latentspacepod
Substack: https://latent.space
Chapters
00:00:00 Introduction
00:00:47 Background on ARC AGI and the Benchmark
00:02:16 Defining Intelligence and Benchmark Philosophy
00:04:20 Challenges in Measuring Intelligence & Generalization
00:07:05 Introducing ARC AGI V3 and Interactive Games
00:08:26 Game Demo: Locksmith & Mechanics
00:12:30 Game Design, API, and Agent Competition
00:23:32 Team Structure and Building the Benchmark
00:24:43 Future Roadmap for ARC AGI Benchmarks
00:27:05 Durability, Timelines & Benchmark Evolution
00:29:17 The Role of ARC Prize in AGI Progress
00:32:05 Grok 4 Livestream & Model Evaluation
00:36:10 OpenAI, Competition, and the State of AI Labs
00:38:47 Closing Remarks & Call to Action
0:00 Introduction
0:46 Overview of Marimo and Its Features
2:33 Origin Story and Motivation Behind Marimo
4:26 Demo: Classical Machine Learning with MNIST in Marimo
6:52 Notebook Compatibility and Conversion from Jupyter
7:42 Demo: Interactive Notebook with Custom UI and Layout
10:08 AI-Native Utilities and Code Generation with Language Models
11:36 Dependency Management and Integration with UV Package Manager
13:00 Demo: Data Annotation Workflow Using a PS5 Controller
15:51 Starting from Scratch: Blank Canvas AI Use Cases
18:27 Context Formatting for AI Code Generation
19:54 Chat Interface and Local/Remote Model Support
21:01 WebAssembly Support and MoLab Cloud-Hosted Notebooks
23:21 Future Plans and Breaking Out of Old Notebook Habits
25:40 Running Marimo Notebooks as Scripts or Data Apps
26:44 Exploring AI Agents and Community Contributions
26:56 Call to Action: How to Get Started and Contribute
Andrew Hsu describes the exact moment his team realized speech AI had become superhuman. They played audio of a Korean English learner speaking. Four humans in the room closed their eyes and couldn't understand a word.
Whisper got it perfectly.
This was 2022. The pieces were finally clicking together - advanced speech models, ChatGPT launching, GPT-3.5 turbo dropping. Everything they'd predicted years earlier was suddenly real.
The most interesting part? They saw it coming but were still shocked when it actually happened.
Fast-apply models? Dead in 3 months, maybe less
Here's what's happening: The founders of companies built entirely around fine-tuning fast-apply models are having candid conversations about their own expiration date
Cursor opened this window in July, but it's closing fast
Why? Models are getting dramatically better at large context and search/replace operations. RAG and FastApply were bandaids for when models couldn't handle these tasks well
Now they're just extra ingredients that can make things go wrong
The writing is on the wall - we're moving beyond these intermediate solutions faster than anyone expected
What does this mean for AI engineers? Time to adapt or get left behind
The real answer? Enterprise adoption through open source trust
When companies can control their data, choose their providers, and avoid sending code to unknown servers - that's when they're willing to pay
Data privacy isn't just a feature anymore, it's the business model
#AI #OpenSource #Enterprise #DataPrivacy #Startups
"Outside the Bay Area, GPT-4 changed *nothing*" - and he's not wrong. While we're all talking about AI revolution, most people's daily lives remain unchanged.
The real insight? We need more builders creating actual applications, not just talking about potential. Real world inertia is massive.
This is exactly what we want though - slow takeoff gives us time to adapt. The tech is here, but the transformation takes time.
What's your take? Are you seeing AI impact in your daily work or is it still mostly hype where you are?
Full writeup: https://www.latent.space/p/cline
X: https://x.com/latentspacepod
Chapters:
00:00 - Introductions
01:35 - Plan and Act Paradigm
05:37 - Model Evaluation and Early Development of Cline
08:14 - Use Cases of Cline Beyond Coding
09:09 - Why Cline is a VS Code Extension and Not a Fork
12:07 - Economic Value of Programming Agents
16:07 - Early Adoption for MCPs
19:35 - Local vs Remote MCP Servers
22:10 - Anthropic's Role in MCP Registry
22:49 - Most Popular MCPs and Their Use Cases
25:26 - Challenges and Future of MCP Monetization
27:32 - Security and Trust Issues with MCPs
28:56 - Alternative History Without MCP
29:43 - Market Positioning of Coding Agents and IDE Integration Matrix
32:57 - Visibility and Autonomy in Coding Agents
35:21 - Evolving Definition of Complexity in Programming Tasks
38:16 - Forks of Cline and Open Source Regrets
40:07 - Simplicity vs Complexity in Agent Design
46:33 - How Fast Apply Got Bitter Lesson'd
49:12 - Cline's Business Model and Bring-Your-Own-API-Key Approach
54:18 - Integration with OpenRouter and Enterprise Infrastructure
55:32 - Impact of Declining Model Costs
57:48 - Background Agents and Multi-Agent Systems
1:00:42 - Vision and Multi-Modalities
1:01:07 - State of Context Engineering
1:07:37 - Memory Systems in Coding Agents
1:10:14 - Standardizing Rules Files Across Agent Tools
1:11:16 - Cline's Personality and Anthropomorphization
1:12:55 - Hiring at Cline and Team Culture
Chapters
00:00:00 Introduction and Guest Intros
00:00:29 What is Klein? Product Overview
00:01:42 Plan and Act Paradigm
00:05:22 Model Evolution and Building Klein
00:07:40 Beyond Coding: Klein as a General Agent
00:09:12 Why Focus on VS Code Extension?
00:11:26 The Future of Programming and Agentic Paradigm
00:12:34 Economic Value: Programming vs. Other Use Cases
00:16:04 MCP Ecosystem: Growth and Marketplace
00:21:30 Security, Discoverability, and Trust in MCPs
00:22:55 Popular MCPs and Workflow Automation
00:25:30 Monetization and Payments for MCPs
00:37:53 Competition, Forks, and Open Source Philosophy
00:40:39 RAG, Fast Apply, and Agentic Simplicity
00:50:11 Business Model and Enterprise Adoption
00:57:04 Background Agents, Multi-Agent Systems, and CLI
01:00:41 Context Engineering and Memory
01:12:39 Team, Culture, and Closing Thoughts
Gen 1: Rosetta Stone CDs at airports
Gen 2: Mobile apps like Duolingo (gamified, casual)
Gen 3: AI-native apps focused on functional fluency
Andrew Struggs breaks down how AI is transforming language learning - instead of vocab drills, it's about practicing real conversations with your Uber driver until speaking becomes automatic.
The shift from gamification to practical fluency is huge. Language learning isn't just about points anymore - it's about actually being able to communicate naturally.
What's your take on AI-powered language learning vs traditional methods?
Prateek spent months finding where models actually break in real scenarios. Turns out they're still making critical mistakes that benchmarks miss.
Time to recalibrate expectations. Models aren't as reliable as the numbers suggest.
Speak CTO Andrew Struggs drops a reality check: everyone obsesses over how fast the first audio bytes come back, but they're missing the real bottleneck.
The actual killer? Voice Activity Detection (VAD) - figuring out when someone stops talking. That can easily add another second if done poorly.
Even worse for language learners who pause mid-sentence for 10+ seconds while thinking.
The metric that actually matters: user stops talking → model starts responding.
This is why building real-time AI feels so much harder than the benchmarks suggest. The devil is always in the details nobody talks about.
We recorded this pod with Pratik Bhavsar to discuss one of the leading approachest to Agent evals: huggingface.co/spaces/galileo-ai/agent-leaderboard
read more:
Chapters
00:00:00 Introduction & Guest Welcome
00:00:39 Shift from Model to Agent Evaluations
00:02:02 Designing the Agent Leaderboard
00:03:23 Methodology: Data Sets and Metrics
00:06:42 Key Findings and Surprising Results
00:11:42 Deep Dive on Benchmarks: TSQ, BFCO, XLAN, and Tau
00:13:44 Tau Bench: Domain-Specific Evaluation
00:16:38 LLM-as-Judge Metrics and Prompting
00:23:59 Preview of Agent Leaderboard V2
00:26:41 Complexity in Realistic Evaluations
00:30:33 Action Completion Metrics & Closing Thoughts
Andrew Struggs from Speak drops some fascinating insights:
• Technical barriers are real - German verbs come at the end of sentences, creating unavoidable latency
• But more importantly: people don't just want translation, they want CONNECTION
Asian users aren't looking for a translator - they want to look you in the eye and speak the same language. It's about becoming a better person and connecting with others.
The human element always wins. Whether it's dating someone from Romania or learning Italian for your wife - language learning is deeply personal.
Real-time translation will complement learning, not replace it.


