Terminal-Bench: Pushing Claude Code, OpenAI Codex, Factory Droid, et al to the limits @LatentSpacePod
Terminal-Bench: Pushing Claude Code, OpenAI Codex, Factory Droid, et al to the limits  @LatentSpacePod
Uploaded October 2025 | Updated September 2026, 2 weeks ago
We are joined by *Alex Shaw* and *Mike Merrill,* the creators of Terminal-Bench (tbench.ai/leaderboard). It is a coding agent benchmark that seemingly came out of nowhere to become the industry standard, adopted by all major frontier labs including Anthropic, OpenAI, and leading agent companies. They explain why they bet on terminal-based interaction over GUI computer use and how Nicholas Carlini at Anthropic became an early champion who helped get it featured prominently in Claude's model card. We also talked about the minimal agent "Terminus" designed to isolate model capabilities from agent optimizations.
00:00:00 Introduction and Terminal Bench Origins
00:03:57 The Anthropic Surprise and Industry Adoption
00:06:14 Terminal vs GUI: The Design Philosophy
00:07:53 Task Design and Benchmark Diversity
00:10:41 Example Deep Dive: Train FastText Task
00:15:50 Building a Meta-Benchmark Framework
00:21:03 Terminus Agent: Separating Model from Harness
00:25:34 The Future of Agent Evaluation
00:26:51 Roadmap: Cloud Hosting and Framework Evolution
00:29:18 Beyond Accuracy: Multi-Dimensional Evaluation
Terminal-Bench: Pushing Claude Code, OpenAI Codex, Factory Droid, et al to the limitsSenior Dev: This Grill Me Prompt Is Going Viral Among Top EngineersThe $15B Physical AI Company: Simulation, Autonomy OS, Neural Sim, & 1K Engineers—Applied IntuitionJamba Mini A 3B Dense Model for Long Context & Hybrid ArchitectureSynthetic data + tool use for LLM improvements 🦙SAM 3: The Eyes for AI  — Nikhila & Pengchuan (Meta Superintelligence), ft. Joseph Nelson (Roboflow)Terminal Bench: A Framework for Diverse BenchmarksCursors Third Era: Cloud Agents — ft. Sam Whitmore, Jonas Nelle, Cursor[State of Context Engineering] Agentic RAG, Context Rot, MCP, Subagents — Nina Lopatina, ContextualAI to AEs: Grit, Glean, and Kleiner Perkins next Enterprise AI hit — Joubin Mirzadegan, RoadrunnerJ3 Efficient MOE for InferenceData Science: The Art of Tasteful Decisions
Latent Space |

Terminal-Bench: Pushing Claude Code, OpenAI Codex, Factory Droid, et al to the limits

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER