Uploaded August 2026 | Updated September 2026, 2 weeks ago
Don't miss out! Join us at our next KubeCon + CloudNativeCon events in Shanghai, China (8-9 September, 2026) and Salt Lake City, United States (Nov 9–12, 2026). Connect with our current graduated, incubating, and sandbox projects as the community gathers to further the education and advancement of cloud native computing. Learn more at kubecon.io
Open Source Meets Production MaaS: GMI Inference Engine and Karmada for Global AI Inference - Xiaokang Wang, GMI Cloud
Global AI inference across 40+ regions and 5+ GPU vendors needs more than GPU capacity. It also needs a practical way to govern clusters, distribute control-plane artifacts, and maintain consistent operations at scale. GMI Inference Engine is GMI Cloud’s Kubernetes-native inference platform, with intelligent routing, heterogeneous resource abstraction, GPU page fault, and entropy-aware placement. This talk shows how GMI uses Karmada, a CNCF-incubating project, to improve multi-cluster governance, CRD distribution, failover readiness, and platform delivery for global AI services.
Agenda
1. Business challenges in running global AI inference across regions and GPU vendors
2. Where Karmada helps: multi-cluster governance, CRD distribution, failover, and consistency
3. Where GMI Inference Engine helps: routing, GPU awareness, elasticity, and placement
4. How open-source Karmada supports GMI’s production platform
5. Lessons learned and collaboration opportunities for the community
Don't miss out! Join us at our next KubeCon + CloudNativeCon events in Shanghai, China (8-9 September, 2026) and Salt Lake City, United States (Nov 9–12, 2026). Connect with our current graduated, incubating, and sandbox projects as the community gathers to further the education and advancement of cloud native computing. Learn more at kubecon.io
Open Source Meets Production MaaS: GMI Inference Engine and Karmada for Global AI Inference - Xiaokang Wang, GMI Cloud
Global AI inference across 40+ regions and 5+ GPU vendors needs more than GPU capacity. It also needs a practical way to govern clusters, distribute control-plane artifacts, and maintain consistent operations at scale. GMI Inference Engine is GMI Cloud’s Kubernetes-native inference platform, with intelligent routing, heterogeneous resource abstraction, GPU page fault, and entropy-aware placement. This talk shows how GMI uses Karmada, a CNCF-incubating project, to improve multi-cluster governance, CRD distribution, failover readiness, and platform delivery for global AI services.
Agenda
1. Business challenges in running global AI inference across regions and GPU vendors
2. Where Karmada helps: multi-cluster governance, CRD distribution, failover, and consistency
3. Where GMI Inference Engine helps: routing, GPU awareness, elasticity, and placement
4. How open-source Karmada supports GMI’s production platform
5. Lessons learned and collaboration opportunities for the community










