Uploaded November 2025 | Updated September 2026, 2 weeks ago
At Ray Summit 2025, Ding Ke and Chendi Xue from Intel share the latest advancements in bringing high-performance vLLM inference to Intel’s full hardware portfolio—including Intel GPUs (XPU), Gaudi Accelerators (HPU), and Intel CPUs.
They begin by highlighting Intel’s commitment to ensuring that vLLM delivers top-tier performance across diverse hardware backends. The session provides a comprehensive update on the state of vLLM enablement on Intel platforms, covering four key areas:
Feature Parity & Performance: A deep dive into support for the new vLLM v1 architecture—including KV connector, data parallelism, multi-token prediction—and how these features perform on Intel GPUs, HPUs, and CPUs.
Model Support: Updates on Intel-optimized models such as DeepSeek and GPT-OSS, and how the model ecosystem continues to expand.
Bridging the Ecosystem Gap: Insights from Intel’s efforts to migrate vLLM capabilities from CUDA to non-CUDA environments. The speakers detail strategies for minimizing developer friction by aligning APIs with torch.cuda behavior, and highlight open-sourced kernels (Cutlass, Triton) upstreamed into vLLM, BitsAndBytes, and other libraries.
Future Directions: A forward look at Intel’s roadmap, including upcoming optimizations and capabilities that will further enhance performance and developer experience across all Intel hardware.
Attendees will walk away with a clear understanding of how to deploy performant LLM inference using vLLM on Intel platforms, practical lessons from real-world migration challenges, and insights into Intel’s broader vision for a unified, developer-friendly AI ecosystem.
Subscribe to our YouTube channel to stay up-to-date on the future of AI! youtube.com/c/anyscale
🔗 Connect with us:
LinkedIn: linkedin.com/company/joinanyscale
X: https://x.com/anyscalecompute
Website: anyscale.com
At Ray Summit 2025, Ding Ke and Chendi Xue from Intel share the latest advancements in bringing high-performance vLLM inference to Intel’s full hardware portfolio—including Intel GPUs (XPU), Gaudi Accelerators (HPU), and Intel CPUs.
They begin by highlighting Intel’s commitment to ensuring that vLLM delivers top-tier performance across diverse hardware backends. The session provides a comprehensive update on the state of vLLM enablement on Intel platforms, covering four key areas:
Feature Parity & Performance: A deep dive into support for the new vLLM v1 architecture—including KV connector, data parallelism, multi-token prediction—and how these features perform on Intel GPUs, HPUs, and CPUs.
Model Support: Updates on Intel-optimized models such as DeepSeek and GPT-OSS, and how the model ecosystem continues to expand.
Bridging the Ecosystem Gap: Insights from Intel’s efforts to migrate vLLM capabilities from CUDA to non-CUDA environments. The speakers detail strategies for minimizing developer friction by aligning APIs with torch.cuda behavior, and highlight open-sourced kernels (Cutlass, Triton) upstreamed into vLLM, BitsAndBytes, and other libraries.
Future Directions: A forward look at Intel’s roadmap, including upcoming optimizations and capabilities that will further enhance performance and developer experience across all Intel hardware.
Attendees will walk away with a clear understanding of how to deploy performant LLM inference using vLLM on Intel platforms, practical lessons from real-world migration challenges, and insights into Intel’s broader vision for a unified, developer-friendly AI ecosystem.
Subscribe to our YouTube channel to stay up-to-date on the future of AI! youtube.com/c/anyscale
🔗 Connect with us:
LinkedIn: linkedin.com/company/joinanyscale
X: https://x.com/anyscalecompute
Website: anyscale.com










