Uploaded August 2026 | Updated September 2026, 3 weeks ago
Every leaderboard you have seen was built by asking a model to do one task, wiping its memory, and asking it another. Parth Asawa's objection is that this quietly assumes learning across instances does not count. His benchmark measures what that assumption hides, using a metric called gain: run a system with state, then run the identical system reset between every single instance, and take the difference. Cumulative reward cannot show you this, because a stronger base model can post a higher total while learning less than a weaker one that genuinely improves.
Building tasks that can measure learning turns out to be the hard part, and he sets three requirements. Headroom, so the task is not already solved by pretraining. Shared latent structure across instances, since standard benchmarks are deliberately independent and therefore offer nothing to improve on, which is why chaining existing benchmarks together does not work. And a learning signal in the environment, whether reward, error messages, or plain text. Continual Learning Bench 1.0 spans six domains including database exploration, where a system should need fewer SQL queries by the tenth question, after which a schema migration tests whether it can throw away stale knowledge without throwing away the useful kind. The headline result is uncomfortable: plain in context learning tops the leaderboard, beating the more elaborate context management systems on reward, on gain, and on cost. Failure modes land on either side of stability and plasticity, including a forecasting model that overpredicts, is corrected, underpredicts, is corrected again, and then jumps straight back to its original overprediction instead of splitting the difference.
Speaker info:
- https://x.com/pgasawa
- linkedin.com/in/pgasawa
- pgasawa.github.io
Timestamps:
0:00 - How we evaluate models today
1:28 - Imagine forgetting everything after every task
2:05 - What continual learning actually means
2:42 - In context, external memory, or parametric
3:18 - The case that we are not measuring it at all
3:58 - What the existing literature does
4:37 - Why those evaluations are not enough
5:13 - Why you cannot chain existing benchmarks
5:51 - Design criterion one: headroom
7:06 - Shared structure and a learning mechanism
7:42 - Reward, and why cumulative reward misleads
8:58 - Gain: the same system with memory wiped
10:15 - Isolating learning from base capability
10:54 - The database exploration task
12:06 - Adding concept drift with a migration
13:20 - Six domains in the benchmark
13:57 - Results, and the in context learning surprise
15:12 - Failure modes on stability and plasticity
15:51 - A forecast that forgets its own correction
16:28 - A notepad that refuses to update
17:07 - Why the training stack was never built for this
17:47 - The sunk cost fallacy in continual learning
19:02 - Rethinking third party AI research
19:38 - Roadmap
Every leaderboard you have seen was built by asking a model to do one task, wiping its memory, and asking it another. Parth Asawa's objection is that this quietly assumes learning across instances does not count. His benchmark measures what that assumption hides, using a metric called gain: run a system with state, then run the identical system reset between every single instance, and take the difference. Cumulative reward cannot show you this, because a stronger base model can post a higher total while learning less than a weaker one that genuinely improves.
Building tasks that can measure learning turns out to be the hard part, and he sets three requirements. Headroom, so the task is not already solved by pretraining. Shared latent structure across instances, since standard benchmarks are deliberately independent and therefore offer nothing to improve on, which is why chaining existing benchmarks together does not work. And a learning signal in the environment, whether reward, error messages, or plain text. Continual Learning Bench 1.0 spans six domains including database exploration, where a system should need fewer SQL queries by the tenth question, after which a schema migration tests whether it can throw away stale knowledge without throwing away the useful kind. The headline result is uncomfortable: plain in context learning tops the leaderboard, beating the more elaborate context management systems on reward, on gain, and on cost. Failure modes land on either side of stability and plasticity, including a forecasting model that overpredicts, is corrected, underpredicts, is corrected again, and then jumps straight back to its original overprediction instead of splitting the difference.
Speaker info:
- https://x.com/pgasawa
- linkedin.com/in/pgasawa
- pgasawa.github.io
Timestamps:
0:00 - How we evaluate models today
1:28 - Imagine forgetting everything after every task
2:05 - What continual learning actually means
2:42 - In context, external memory, or parametric
3:18 - The case that we are not measuring it at all
3:58 - What the existing literature does
4:37 - Why those evaluations are not enough
5:13 - Why you cannot chain existing benchmarks
5:51 - Design criterion one: headroom
7:06 - Shared structure and a learning mechanism
7:42 - Reward, and why cumulative reward misleads
8:58 - Gain: the same system with memory wiped
10:15 - Isolating learning from base capability
10:54 - The database exploration task
12:06 - Adding concept drift with a migration
13:20 - Six domains in the benchmark
13:57 - Results, and the in context learning surprise
15:12 - Failure modes on stability and plasticity
15:51 - A forecast that forgets its own correction
16:28 - A notepad that refuses to update
17:07 - Why the training stack was never built for this
17:47 - The sunk cost fallacy in continual learning
19:02 - Rethinking third party AI research
19:38 - Roadmap

![Building Turbopuffer: Gergely Orosz (@pragmaticengineer ) × Simon Eskildsen (CEO)
This fireside chat between Gergely Orosz and Simon Eskildsen explores the technical journey and engineering philosophy behind the database company Turbopuffer.
Video Timestamps
0:00 Introduction and Simon’s early history with computers
3:02 The International Olympiad in Informatics and early competitive programming
4:13 How Simon was recruited by Shopify while still in high school
8:46 Engineering challenges and scaling infrastructure at Shopify
14:56 Decision to leave Shopify and the creation of the napkin math project
20:40 The origin and technical motivations behind Turbopuffer
24:46 Design challenges of building a database on top of S3
28:41 Cursor becoming the first major customer
35:36 The meeting with Jensen Huang and Nvidia’s push for GPUs
39:01 The competitive reality of cloud infrastructure and CPU scarcity
43:06 Philosophical perspective on venture capital and funding
51:45 Building a remote-first culture with the campfire concept
Quotes
(19:49) Because you batch. So an f-sync happens on usually a 4K... its not intuitive. Its actually—I got caught—I just got obsessed with this question.
(30:33) Yeah, you could do a million vectors for a dollar. And before that, I think the cheapest was maybe $100 per million for something that actually worked.
(36:56) [Jensen Huang] said, Judging by your slide, maybe you should [pivot into vapes].
(43:53) I promised Cursor that Justine and I could get their bill to 4K a month... thats the pricing we ship with.
(49:54) The third reason to raise capital is for the founders ego... I wish that it was more talked about because youre diluting all of your employees when you do it.
## Speakers
### Gergely Orosz
Author / Founder, The Pragmatic Engineer · The Pragmatic Engineer
[X/Twitter](https://twitter.com/gergelyorosz) · [LinkedIn](https://www.linkedin.com/in/gergelyorosz/) · [Website](https://pragmaticengineer.com)
Software engineer, engineering leader, and author of The Software Engineers Guidebook; best known for The Pragmatic Engineer newsletter and blog covering software engineering practices, engineering leadership, and the tech industry. Previously held engineering leadership roles at Uber and worked at companies including Skype and Skyscanner.
### Simon Eskildsen
CEO and co-founder · turbopuffer
[X/Twitter](https://x.com/Sirupsen) · [LinkedIn](https://www.linkedin.com/in/sirupsen/) · [Website](https://sirupsen.com) · [Blog](https://sirupsen.com/napkin)
Co-founder and CEO at turbopuffer. Formerly Principal Engineer at Shopify, where he helped scale infra from 1K → 1M RPS.
— [View on the schedule](https://www.ai.engineer/worldsfair/schedule?session=asn_slot_2026_06_30_main_stage_1230_2026_06_25t07_57_06_000z) Building Turbopuffer: Gergely Orosz (@pragmaticengineer ) × Simon Eskildsen (CEO)](https://i.ytimg.com/vi/jQDXzEVHMSE/mqdefault.jpg)








![[Full Workshop] Building Metrics that actually work — David Karam, Pi Labs (fmr Google Search)
[Full Workshop] Building Metrics that actually work — David Karam, Pi Labs (fmr Google Search) [Full Workshop] Building Metrics that actually work — David Karam, Pi Labs (fmr Google Search)](https://i.ytimg.com/vi/jxrGodnopHo/mqdefault.jpg)