Uploaded August 2026 | Updated September 2026, 2 weeks ago
Don't miss out! Join us at our next KubeCon + CloudNativeCon events in Shanghai, China (8-9 September, 2026) and Salt Lake City, United States (Nov 9–12, 2026). Connect with our current graduated, incubating, and sandbox projects as the community gathers to further the education and advancement of cloud native computing. Learn more at kubecon.io
How To Evolve Your LLM Self-Hosting Platform: A Practical Guide To Adopting Advanced Optimizations - Shingo Omura & Yiyang Zhan, LY Corporation
When self-hosting LLMs, adopting advanced inference technologies prematurely drastically increases deployment complexity. How should engineers evaluate these tools and determine the right time to introduce them?
Through a real-world case study, the speakers will share their journey of evolving a minimal vLLM architecture into a robust multi-tenant platform using Envoy AI Gateway and Athenz. They will highlight how establishing this foundation—enabling authentication, rate limiting, and crucial tenant-level usage statistics—is a vital use case for safely supporting multiple teams.
Building on this operational foundation, they will discuss their ongoing evaluation and incremental adoption of advanced optimizations like P/D disaggregation and intelligent routing using actual production bottlenecks.
Attendees will learn what each optimization improves, the operational trade-offs, and the optimal timing to introduce them with manageable complexity.
Don't miss out! Join us at our next KubeCon + CloudNativeCon events in Shanghai, China (8-9 September, 2026) and Salt Lake City, United States (Nov 9–12, 2026). Connect with our current graduated, incubating, and sandbox projects as the community gathers to further the education and advancement of cloud native computing. Learn more at kubecon.io
How To Evolve Your LLM Self-Hosting Platform: A Practical Guide To Adopting Advanced Optimizations - Shingo Omura & Yiyang Zhan, LY Corporation
When self-hosting LLMs, adopting advanced inference technologies prematurely drastically increases deployment complexity. How should engineers evaluate these tools and determine the right time to introduce them?
Through a real-world case study, the speakers will share their journey of evolving a minimal vLLM architecture into a robust multi-tenant platform using Envoy AI Gateway and Athenz. They will highlight how establishing this foundation—enabling authentication, rate limiting, and crucial tenant-level usage statistics—is a vital use case for safely supporting multiple teams.
Building on this operational foundation, they will discuss their ongoing evaluation and incremental adoption of advanced optimizations like P/D disaggregation and intelligent routing using actual production bottlenecks.
Attendees will learn what each optimization improves, the operational trade-offs, and the optimal timing to introduce them with manageable complexity.


