Uploaded November 2025 | Updated September 2026, 2 weeks ago
This is the clearest explanation of modern inference we’ve heard all year.
We brought Baseten CEO Tuhin Srivastava onto Gradient Dissent, and he breaks down what actually happens when teams move from demos to real production workloads.
Tuhin explains what teams run into once they move past demos and into real production:
• Scaling across large GPU fleets without breaking workloads or spiking latency
• Pushing runtimes to their limits as you choose between vLLM, TensorRT-LLM, and SGLang on new hardware
He also shares why many teams eventually move from closed models to open source once cost and control become critical.
If you want a clear snapshot of how large models run in the real world, this episode is worth your time.
This is the clearest explanation of modern inference we’ve heard all year.
We brought Baseten CEO Tuhin Srivastava onto Gradient Dissent, and he breaks down what actually happens when teams move from demos to real production workloads.
Tuhin explains what teams run into once they move past demos and into real production:
• Scaling across large GPU fleets without breaking workloads or spiking latency
• Pushing runtimes to their limits as you choose between vLLM, TensorRT-LLM, and SGLang on new hardware
He also shares why many teams eventually move from closed models to open source once cost and control become critical.
If you want a clear snapshot of how large models run in the real world, this episode is worth your time.










