Uploaded November 2025 | Updated September 2026, 2 weeks ago
Baseten CEO, Tuhin Srivastava explains why serious LLM workloads can’t rely on simple shared endpoints. Each model brings its own constraints, and the entire system has to adapt around them.
A real-time view into the complexity behind production-scale inference.
Full episode on the channel.
Baseten CEO, Tuhin Srivastava explains why serious LLM workloads can’t rely on simple shared endpoints. Each model brings its own constraints, and the entire system has to adapt around them.
A real-time view into the complexity behind production-scale inference.
Full episode on the channel.










