The $2B Company Cutting AI Costs By 60% | Tuhin Srivastava @WeightsBiases
The $2B Company Cutting AI Costs By 60% | Tuhin Srivastava  @WeightsBiases
Uploaded November 2025 | Updated September 2026, 2 weeks ago
In this episode of Gradient Dissent, Lukas Biewald talks with Tuhin Srivastava, CEO and founder of Baseten, one of the fastest-growing companies in the AI inference ecosystem. Tuhin shares the real story behind Baseten’s rise and how the market finally aligned with the infrastructure they’d spent years building.

They get into the core challenges of modern inference, including why dedicated deployments matter, how runtime and infrastructure bottlenecks stack up, and what makes serving large models fundamentally different from smaller ones.

Tuhin also explains how vLLM, TensorRT-LLM, and SGLang differ in practice, what it takes to tune workloads for new chips like the B200, and why reliability becomes harder as systems scale.

The conversation dives into company-building, from killing product lines to avoiding premature scaling while navigating a market that shifts every few weeks.

Timestamps:

00:00 Intro
02:13 The Journey of Baseten
08:36 The Impact of ChatGPT and Stable Diffusion
19:11 Operational Discipline and Company Culture
20:31 Differentiating in the Inference Market
30:18 Infrastructure and Runtime Challenges
32:14 Optimizing Inference Performance
37:39 The Role of Hardware in AI Inference
45:47 Market Dynamics and Future Predictions
50:52 The Importance of Inference in AI
58:52 Conclusion and Final Thoughts


Connect with us here:

Tuhin Srivastva: linkedin.com/in/tuhin-srivastava
Lukas Biewald: linkedin.com/in/lbiewald
Weights & Biases: linkedin.com/company/wandb
The $2B Company Cutting AI Costs By 60% | Tuhin SrivastavaShe Raised $64M to Build an AI Math Prodigy | Carina Hong, CEO of AxiomUsing MCP to analyse your experiments & create reports in natural languageAtlassian’s Most Controversial Growth Decision | Mike Cannon-BrookesInside the $41B AI Cloud Challenging Big Tech | CoreWeave SVPAccelerate LLM post training with W&B Serverless SFTMost AI Startups Are Scaling Into Bankruptcy | Lin Qiao, CEO of FireworksUnified stack of Kubernetes, Ray, PyTorch, and vLLM.Rapid prototyping for reinforcement learning in industry - Festo @ Fully Connected London 25Beyond RAG: Production-ready AI agents powered by enterprise-scale dataHow Woven by Toyota builds video AI agents for automated driving with W&B WeaveFully Connected Tokyo 2025: Opening Keynote with W&B Cofounders Lukas Biewald & Chris Van Pelt
Weights & Biases |

The $2B Company Cutting AI Costs By 60% | Tuhin Srivastava

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER