The AI Progress Chart Everyone Is Misreading — Beth Barnes & David Rein @MachineLearningStreetTalk
The AI Progress Chart Everyone Is Misreading — Beth Barnes & David Rein  @MachineLearningStreetTalk
Uploaded May 2026 | Updated September 2026, 1 week ago
Beth Barnes and David Rein on the one graph that ate the AI timelines discourse, and why the two people who built it are the most careful about how you read it.

**SPONSOR**
Prolific - Quality data. From real people. For faster breakthroughs.
prolific.com/?utm_source=mlst
Interview: youtu.be/cnxZZTl1tkk
---

Beth Barnes and David Rein from METR on the one graph that ate the AI timelines discourse, and why the people who built it are the most careful about how it gets read.

Beth founded METR after leaving OpenAI alignment. David is first author on GPQA and co-author on HCAST and the METR Time Horizons paper. Together they built the measurement Daniel Kokotajlo called the single most important piece of evidence on AI timelines: the log-linear line of "how long a task a frontier model can complete at 50% reliability" vs release date.

The conversation opens on reward hacking. Current models can articulate in chat why a behaviour is undesired and then execute it anyway as agents. From there: construct validity, Melanie Mitchell's four-problem taxonomy, and the ARC-AGI 1-to-2 collapse as a worked example of adversarially-selected benchmarks regressing once labs target them. Beth's counter: METR deliberately does not adversarially select. David's: models do not have to do the right thing for the right reasons.

Methodology, then specification — David's compiler analogy, Beth on four-month tasks as expensive to evaluate rather than unspecifiable. Then the SWE-bench reality check, the METR finding that half of passing PRs would not be merged, and Beth's horses-versus-bank-tellers analogy for the labour market.

The close: monitorability, the coin-spinning boat, two-year recursive self-improvement, and Beth's line that "overhyped now" and "big deal later" are not correlated claims.

---
TIMESTAMPS:
00:00:00 Intro
00:02:06 Sponsor break: Prolific human-feedback infrastructure
00:02:33 Welcome and the scalable oversight motivation
00:06:02 Construct validity, benchmark pathologies and the Chollet worry
00:15:45 Time Horizons: human time, HCAST tasks and the 50% logistic
00:24:50 Is human difficulty really one variable?
00:33:05 Agent harness evolution and the inference-compute dividend
00:40:00 Scaffolding bells, token budgets and the credit-assignment problem
00:44:15 Look at the damn graph: regularisation bug and reliability nuance
00:50:00 Why 50%? Reliability, reward hacking and pizza-party transcripts
00:55:20 Extrapolation risk and straight lines on graphs
00:59:25 Software engineering as a specification acquisition problem
01:07:40 Compilers also made ugly code: vibe-coding quality and Claude on METR Slack
01:15:15 Strongest defensible claim, Carlini's compiler swarm and AI 2027
01:23:45 SWE-bench merge rates, the bank-teller analogy and horses
01:31:45 Scheming, alignment faking and the mentalistic vocabulary problem
01:40:45 Reward hacking, monitorability and chain-of-thought faithfulness
01:45:25 Recursive self-improvement, knowledge vs intelligence and closing

See top comment for references!
The AI Progress Chart Everyone Is Misreading — Beth Barnes & David ReinGenuine understanding is the worlds most valuable commodity right nowThere are monsters in your LLM. (Murray Shanahan)We need AIs with PHYSICAL experience (Jeff Beck)
Machine Learning Street Talk |

The AI Progress Chart Everyone Is Misreading — Beth Barnes & David Rein

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER