Uploaded August 2026 | Updated September 2026, 3 weeks ago
An agent calls a refund tool and the request times out. Did the customer get their money? Salman Munaf uses that to make his central point, which is that a timeout has never meant failure, it means unknown, and an agent's first instinct on any failure is to try again. Without request identifiers, idempotency keys and a status lookup, that instinct refunds someone twice. He works in site reliability at TikTok, and his argument is that the moment a model started calling external services it stopped being a model problem and became a distributed systems problem, complete with every failure mode that field spent decades naming.
The reframing he keeps returning to is that an agent is a probabilistic coordinator. Older systems coordinated multi step workflows too, but they followed a decision tree somebody drew. This one does not, so the determinism has to live in the controls around it: circuit breakers, spend and turn ceilings, compensating actions defined per step, and credentials scoped to separate reads from writes rather than handed over wholesale. He is good on two things teams get wrong. Context that can influence an action is state, so it goes stale and needs invalidation and provenance like any cache. And human approval has to bind to an action, an actor and an expiry, or approving a 30 dollar refund quietly becomes approval for a 300 dollar one.
Speaker info:
- linkedin.com/in/salman96
Timestamps:
0:00 - Two incidents that systems thinking would have caught
2:33 - When the architectural boundary left the model
3:46 - The agent as a probabilistic coordinator
4:57 - Every step of the loop crosses a boundary
7:19 - A timeout means unknown, not failure
8:32 - Idempotency keys and status lookups
9:42 - Retry storms, backoff and budgets
10:57 - Context that influences action is state
12:08 - Treating memory as a cache
13:19 - Compensating actions across systems
14:36 - Circuit breakers, rate limits and ceilings
15:43 - Scoped credentials over blanket permissions
16:51 - Why logs are not enough
19:20 - What the system lets it do when it is wrong
An agent calls a refund tool and the request times out. Did the customer get their money? Salman Munaf uses that to make his central point, which is that a timeout has never meant failure, it means unknown, and an agent's first instinct on any failure is to try again. Without request identifiers, idempotency keys and a status lookup, that instinct refunds someone twice. He works in site reliability at TikTok, and his argument is that the moment a model started calling external services it stopped being a model problem and became a distributed systems problem, complete with every failure mode that field spent decades naming.
The reframing he keeps returning to is that an agent is a probabilistic coordinator. Older systems coordinated multi step workflows too, but they followed a decision tree somebody drew. This one does not, so the determinism has to live in the controls around it: circuit breakers, spend and turn ceilings, compensating actions defined per step, and credentials scoped to separate reads from writes rather than handed over wholesale. He is good on two things teams get wrong. Context that can influence an action is state, so it goes stale and needs invalidation and provenance like any cache. And human approval has to bind to an action, an actor and an expiry, or approving a 30 dollar refund quietly becomes approval for a 300 dollar one.
Speaker info:
- linkedin.com/in/salman96
Timestamps:
0:00 - Two incidents that systems thinking would have caught
2:33 - When the architectural boundary left the model
3:46 - The agent as a probabilistic coordinator
4:57 - Every step of the loop crosses a boundary
7:19 - A timeout means unknown, not failure
8:32 - Idempotency keys and status lookups
9:42 - Retry storms, backoff and budgets
10:57 - Context that influences action is state
12:08 - Treating memory as a cache
13:19 - Compensating actions across systems
14:36 - Circuit breakers, rate limits and ceilings
15:43 - Scoped credentials over blanket permissions
16:51 - Why logs are not enough
19:20 - What the system lets it do when it is wrong








![Building Turbopuffer: Gergely Orosz (@pragmaticengineer ) × Simon Eskildsen (CEO)
This fireside chat between Gergely Orosz and Simon Eskildsen explores the technical journey and engineering philosophy behind the database company Turbopuffer.
Video Timestamps
0:00 Introduction and Simon’s early history with computers
3:02 The International Olympiad in Informatics and early competitive programming
4:13 How Simon was recruited by Shopify while still in high school
8:46 Engineering challenges and scaling infrastructure at Shopify
14:56 Decision to leave Shopify and the creation of the napkin math project
20:40 The origin and technical motivations behind Turbopuffer
24:46 Design challenges of building a database on top of S3
28:41 Cursor becoming the first major customer
35:36 The meeting with Jensen Huang and Nvidia’s push for GPUs
39:01 The competitive reality of cloud infrastructure and CPU scarcity
43:06 Philosophical perspective on venture capital and funding
51:45 Building a remote-first culture with the campfire concept
Quotes
(19:49) Because you batch. So an f-sync happens on usually a 4K... its not intuitive. Its actually—I got caught—I just got obsessed with this question.
(30:33) Yeah, you could do a million vectors for a dollar. And before that, I think the cheapest was maybe $100 per million for something that actually worked.
(36:56) [Jensen Huang] said, Judging by your slide, maybe you should [pivot into vapes].
(43:53) I promised Cursor that Justine and I could get their bill to 4K a month... thats the pricing we ship with.
(49:54) The third reason to raise capital is for the founders ego... I wish that it was more talked about because youre diluting all of your employees when you do it.
## Speakers
### Gergely Orosz
Author / Founder, The Pragmatic Engineer · The Pragmatic Engineer
[X/Twitter](https://twitter.com/gergelyorosz) · [LinkedIn](https://www.linkedin.com/in/gergelyorosz/) · [Website](https://pragmaticengineer.com)
Software engineer, engineering leader, and author of The Software Engineers Guidebook; best known for The Pragmatic Engineer newsletter and blog covering software engineering practices, engineering leadership, and the tech industry. Previously held engineering leadership roles at Uber and worked at companies including Skype and Skyscanner.
### Simon Eskildsen
CEO and co-founder · turbopuffer
[X/Twitter](https://x.com/Sirupsen) · [LinkedIn](https://www.linkedin.com/in/sirupsen/) · [Website](https://sirupsen.com) · [Blog](https://sirupsen.com/napkin)
Co-founder and CEO at turbopuffer. Formerly Principal Engineer at Shopify, where he helped scale infra from 1K → 1M RPS.
— [View on the schedule](https://www.ai.engineer/worldsfair/schedule?session=asn_slot_2026_06_30_main_stage_1230_2026_06_25t07_57_06_000z) Building Turbopuffer: Gergely Orosz (@pragmaticengineer ) × Simon Eskildsen (CEO)](https://i.ytimg.com/vi/jQDXzEVHMSE/mqdefault.jpg)

