The CEO Behind the Fastest-Growing AI Inference Company @WeightsBiases
The CEO Behind the Fastest-Growing AI Inference Company  @WeightsBiases
Uploaded November 2025 | Updated September 2026, 2 weeks ago
This is the clearest explanation of modern inference we’ve heard all year.

We brought Baseten CEO Tuhin Srivastava onto Gradient Dissent, and he breaks down what actually happens when teams move from demos to real production workloads.

Tuhin explains what teams run into once they move past demos and into real production: 

• Scaling across large GPU fleets without breaking workloads or spiking latency
• Pushing runtimes to their limits as you choose between vLLM, TensorRT-LLM, and SGLang on new hardware

He also shares why many teams eventually move from closed models to open source once cost and control become critical.

If you want a clear snapshot of how large models run in the real world, this episode is worth your time.
The CEO Behind the Fastest-Growing AI Inference CompanyW&B Inference: test open-source LLMs in SECONDSW&B Models: AlertsHe Built A $15B Company Without A Sales TeamWhy I would never touch a peptideLarge-scale agentic quant research with Weights & BiasesWhy Big Tech Buys GPUs From CoreWeave | Corey SandersRun the model, not the risk: Powering private inference for enterprise AI anywhereW&B Inference: Access CoreWeave-hosted open-source models in W&B WeaveOne size doesn’t fit all: Building AI agents specialized for your enterpriseWhy Pharma’s AI Bet Might Be Wrong | Martin ShkreliDefining factors for enterprise AI agents - JetBrains @ FC London 25
Weights & Biases |

The CEO Behind the Fastest-Growing AI Inference Company

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER