AI Agents Are Just Distributed Systems Now — Salman Munaf, TikTok @aiDotEngineer
AI Agents Are Just Distributed Systems Now — Salman Munaf, TikTok  @aiDotEngineer
Uploaded August 2026 | Updated September 2026, 3 weeks ago
An agent calls a refund tool and the request times out. Did the customer get their money? Salman Munaf uses that to make his central point, which is that a timeout has never meant failure, it means unknown, and an agent's first instinct on any failure is to try again. Without request identifiers, idempotency keys and a status lookup, that instinct refunds someone twice. He works in site reliability at TikTok, and his argument is that the moment a model started calling external services it stopped being a model problem and became a distributed systems problem, complete with every failure mode that field spent decades naming.

The reframing he keeps returning to is that an agent is a probabilistic coordinator. Older systems coordinated multi step workflows too, but they followed a decision tree somebody drew. This one does not, so the determinism has to live in the controls around it: circuit breakers, spend and turn ceilings, compensating actions defined per step, and credentials scoped to separate reads from writes rather than handed over wholesale. He is good on two things teams get wrong. Context that can influence an action is state, so it goes stale and needs invalidation and provenance like any cache. And human approval has to bind to an action, an actor and an expiry, or approving a 30 dollar refund quietly becomes approval for a 300 dollar one.

Speaker info:
- linkedin.com/in/salman96

Timestamps:
0:00 - Two incidents that systems thinking would have caught
2:33 - When the architectural boundary left the model
3:46 - The agent as a probabilistic coordinator
4:57 - Every step of the loop crosses a boundary
7:19 - A timeout means unknown, not failure
8:32 - Idempotency keys and status lookups
9:42 - Retry storms, backoff and budgets
10:57 - Context that influences action is state
12:08 - Treating memory as a cache
13:19 - Compensating actions across systems
14:36 - Circuit breakers, rate limits and ceilings
15:43 - Scoped credentials over blanket permissions
16:51 - Why logs are not enough
19:20 - What the system lets it do when it is wrong
AI Agents Are Just Distributed Systems Now — Salman Munaf, TikTokWhy Large? Tiny LMs & Agents on Edge/Robotics — Cormac Brick, GoogleVoiceVision RAG - Integrating Visual Document Intelligence with Voice Response — Suman Debnath, AWSWe Vetted 2000 AI Skills Before They Reached Developers — Lucas Palma, NubankRealtime multiplayer, automation, and you! — Idan Gazit, GitHubShip Production Software in Minutes, Not Months — Eno Reyes, FactoryFull Workshop: Setting Yourself Up for Success —Jason Liu, OpenAI CodexBeyond Static Intelligence: Evaluating Continual Learning — Parth Asawa, UC BerkeleyDesigning Agents (The Floor Is the Frontier) — Ben Hylak, RaindropBuilding Turbopuffer: Gergely Orosz (@pragmaticengineer ) × Simon Eskildsen (CEO)A Taxonomy for Next-gen Reasoning — Nathan Lambert, Allen Institute (AI2) & Interconnects.aiEverything Is a Rollout — Alex Shaw + Ryan Marten, Terminal-Bench, Harbor, Laude Institute
AI Engineer |

AI Agents Are Just Distributed Systems Now — Salman Munaf, TikTok

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER