vLLM TPU: A new unified-backend supporting Pytorch and JAX natively on TPU | Ray Summit 2025 @anyscale
vLLM TPU: A new unified-backend supporting Pytorch and JAX natively on TPU | Ray Summit 2025  @anyscale
Uploaded November 2025 | Updated September 2026, 2 weeks ago
At Ray Summit 2025, Manoj Krishnan and Brittany Rockwell from Google share an in-depth look at the new optimized TPU backend in vLLM, designed to unify and accelerate large-scale inference across both PyTorch and JAX models under a single, consolidated codepath.

They begin by highlighting how this new backend preserves vLLM’s hallmark ease-of-use and portability—allowing developers to move seamlessly between hardware types—while introducing a suite of next-generation TPU capabilities purpose-built for XL-scale model deployments. These include:

Disaggregated serving for more flexible resource allocation

Advanced parallelism strategies for Mixture-of-Experts (MoE) models

Highly optimized Pallas kernels for maximized TPU performance

Enhanced multimodal support tailored for large, heterogeneous model architectures

Manoj and Brittany walk through architectural details, performance optimizations, and practical deployment patterns that make the new TPU backend a powerful option for teams running frontier-scale models.

Subscribe to our YouTube channel to stay up-to-date on the future of AI! youtube.com/c/anyscale

🔗 Connect with us:
LinkedIn: linkedin.com/company/joinanyscale
X: https://x.com/anyscalecompute
Website: anyscale.com
vLLM TPU: A new unified-backend supporting Pytorch and JAX natively on TPU | Ray Summit 2025Hugging Face + vLLM: One Model Definition to Rule Them All | Ray Summit 2025Matrix: Reliable Framework for Data-Centric Experimentation at Scale  | Ray Summit 2025From Spark to Ray: CSSs Data Revolution with Daft | Ray Summit 2024How Rubrik Unlocked AI at Scale with Ray Serve | Ray Summit 2024How Workday Achieved 50x Cheaper Model Serving with Ray Serve | Ray Summit 2025ByteDances Platform for Reinforcement Learning from Human Feedback | Ray Summit 2024Dynamic Scheduling for Large Language Model Serving | Ray Summit 2024How Prime Intellect Builds Scalable Infrastructure for Agentic RL | Ray Summit 2025BentoML or RayServe, You Can Choose Both with BentoRayHow KubeRay Is Evolving for Massive AI Workloads  | Ray Summit 2025RDMA P2P Deep Dive: KvCache Transfer, Weight Updates & MoE Routing at Perplexity | Ray Summit 2025
Anyscale |

vLLM TPU: A new unified-backend supporting Pytorch and JAX natively on TPU | Ray Summit 2025

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER