How Stripe Moves Petabytes of Data with 5.5 Nines of Reliability @infoq
How Stripe Moves Petabytes of Data with 5.5 Nines of Reliability  @infoq
Uploaded May 2026 | Updated September 2026, 2 weeks ago
πŸ“© Subscribe to the InfoQ Weekly newsletter!

No hype. No fluff. Just the signals senior engineers actually care about - from #AI, #DevOps, and #Java to #CloudComputing & #SoftwareArchitecture.

πŸ—“οΈ Delivered every Tuesday, it's your quick round-up of innovator & early-adopter technologies.

Join 250,000+ developers who read it to stay ahead πŸ‘‰ infoq.com/news/#infoq-nl
************************************************
How does Stripe process $1.4 trillion in payments annually with 5.5 nines of availability? Stripe Staff Software Engineer Jimmy Morzaria breaks down the custom zero-downtime data movement platform that powers their critical database tier.

Scaling database infrastructure for global commerce requires moving from treating shards like "pets" to an automated "herd." In this InfoQ, discover how Stripe handles 5 million database queries per second across 2,000+ MongoDB shards. Jimmy pulls back the curtain on why Stripe bypassed off-the-shelf solutions like MongoDB Atlas/mongos to build "DocDB" - their in-house Database-as-a-Service.

You’ll learn the exact blueprint for their zero-downtime horizontal data migration platform, including how they achieved a 10x write throughput boost using B-tree insertion ordering, orchestrated bidirectional replication for safe rollbacks, and implemented custom version gating for seamless traffic switching.

⏱️ Video Timestamps (For Navigation)
00:00 β€” The Hartsfield-Jackson Airport Engineering Analogy
01:30 β€” Stripe’s Scale: $1.4 Trillion & 5.5 Nines Reliability
02:45 β€” The Evolution of Stripe’s Database Infrastructure (2011–2020)
04:15 β€” Moving Beyond the Physical Limits of Vertical Scaling
05:30 β€” Architecture Deep Dive: Why Stripe Built DocDB In-House
07:45 β€” Blueprint for Zero-Downtime Data Movement (First Principles)
09:15 β€” Achieving 10x Write Throughput via B-Tree Optimization
10:45 β€” Bidirectional Replication & Ensuring Idempotency via the Oplog
12:30 β€” The Traffic Switch: Custom Version Gating Protocol
14:50 β€” Beyond Sharding: Black Friday prep & Skip-Major-Version Upgrades
17:15 β€” Q&A: Handling Split-Brain, Fencing Proxies, & Migration Speeds

πŸ”— Transcript & slides available on InfoQ: bit.ly/4dK2RId

#SystemDesign #SoftwareArchitecture #Stripe #DatabaseScaling #MongoDB #DistributedSystems
How Stripe Moves Petabytes of Data with 5.5 Nines of ReliabilityMandy Gu on Generative AI (GenAI) Implementation, User Profiles and Adoption of LLMsHow eBPF Empowers Developers to Observe Inside the Linux Kernel in a Safe and Unintrusive WayThe MCP Megalith: Surviving the 1,000-Tool Explosion and Auth FragmentationPatrick Debois: Why Your Job Is Shifting from Coding to Managing AgentsGreat Architects Facilitate, Not Dictate Software Decisions – Insights from Andrew Harmel-LawThe Truth About RAG & vLLM: Why Your Multimodal System Fails at ScaleThe Craft of Software Architecture in the Age of AI ToolsVibecoding your Own Multi-Agent WorkstationArchitecting APIs in Regulated Industries: From 6 Months to 2 HoursBeyond Copilots: How LinkedIn Scales Multi-Agent SystemsFounders, Friction, and Focus: Building Engineering Teams at Early-Stage Startups
InfoQ |

How Stripe Moves Petabytes of Data with 5.5 Nines of Reliability

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER