Uploaded August 2026 | Updated September 2026, 3 weeks ago
Don't miss out! Join us at our next KubeCon + CloudNativeCon events in Shanghai, China (8-9 September, 2026) and Salt Lake City, United States (Nov 9–12, 2026). Connect with our current graduated, incubating, and sandbox projects as the community gathers to further the education and advancement of cloud native computing. Learn more at kubecon.io
Shared Yet Isolated at Scale: Building Multi-Tenant Inference Platform on Kubernetes - Yuto Hiraki, SoftBank Corp. & Yusuke Tanaka, ITOCHU Techno-Solutions Corporation
As AI Agent adoption grows, so does LLM usage and the cost of running it. A self-hosted LLM inference platform is one approach enterprises are exploring to regain control over costs, security, and governance. But it comes with a real challenge: teams have different use cases, access patterns, and security requirements. Some only need an API endpoint, while others need full administrative control. Traditional namespace isolation falls short of meeting both. At the same time, high GPU utilization is critical to keeping cost per token low. Shared yet isolated, at scale - how do you resolve that contradiction?
In this talk, we share practical insights from building an LLM inference platform on Kubernetes - covering tenant isolation, dynamic and granular GPU allocation, and the trade-offs in between. We'll show how we resolved that contradiction, with practical patterns.
Don't miss out! Join us at our next KubeCon + CloudNativeCon events in Shanghai, China (8-9 September, 2026) and Salt Lake City, United States (Nov 9–12, 2026). Connect with our current graduated, incubating, and sandbox projects as the community gathers to further the education and advancement of cloud native computing. Learn more at kubecon.io
Shared Yet Isolated at Scale: Building Multi-Tenant Inference Platform on Kubernetes - Yuto Hiraki, SoftBank Corp. & Yusuke Tanaka, ITOCHU Techno-Solutions Corporation
As AI Agent adoption grows, so does LLM usage and the cost of running it. A self-hosted LLM inference platform is one approach enterprises are exploring to regain control over costs, security, and governance. But it comes with a real challenge: teams have different use cases, access patterns, and security requirements. Some only need an API endpoint, while others need full administrative control. Traditional namespace isolation falls short of meeting both. At the same time, high GPU utilization is critical to keeping cost per token low. Shared yet isolated, at scale - how do you resolve that contradiction?
In this talk, we share practical insights from building an LLM inference platform on Kubernetes - covering tenant isolation, dynamic and granular GPU allocation, and the trade-offs in between. We'll show how we resolved that contradiction, with practical patterns.










