Hugging Face + vLLM: One Model Definition to Rule Them All | Ray Summit 2025 @anyscale
Hugging Face + vLLM: One Model Definition to Rule Them All | Ray Summit 2025  @anyscale
Uploaded December 2025 | Updated September 2026, 2 weeks ago
At Ray Summit 2025, Harry Mellor from Hugging Face shares how the new Transformers backend for vLLM enables teams to use the exact same model implementation for both training and inference—eliminating duplication, simplifying maintenance, and dramatically accelerating deployment workflows.

He begins by introducing the core idea: with the Transformers backend, vLLM can now run any Transformers-compatible model—whether officially merged into the library or fully custom—directly from its original model definition, but at vLLM-grade performance. This means your modeling code becomes the single source of truth, while vLLM handles high-speed inference, attention optimizations, and distributed scaling.

Harry then walks through:

What the Transformers backend is and how it works under the hood

How to make a model compatible, including tips for adapting custom or experimental architectures

How this unlocks unified codepaths, allowing one modeling codebase to serve both research experimentation and production-grade inference

The session highlights how this backend bridges the gap between training frameworks and high-performance inference engines—streamlining development workflows and reducing operational complexity.

Attendees will leave with a clear understanding of how to adopt the Transformers backend, how to bring custom models into vLLM, and how this new capability enables seamless transitions from model training to large-scale production serving.

Liked this video? Check out other Ray Summit breakout session recordings youtube.com/playlist?list=PLzTswPQNepXllnU0C36WtkC0dqkAoDulh

Subscribe to our YouTube channel to stay up-to-date on the future of AI! youtube.com/c/anyscale

🔗 Connect with us:
LinkedIn: linkedin.com/company/joinanyscale
Hugging Face + vLLM: One Model Definition to Rule Them All | Ray Summit 2025Matrix: Reliable Framework for Data-Centric Experimentation at Scale  | Ray Summit 2025From Spark to Ray: CSSs Data Revolution with Daft | Ray Summit 2024How Rubrik Unlocked AI at Scale with Ray Serve | Ray Summit 2024How Workday Achieved 50x Cheaper Model Serving with Ray Serve | Ray Summit 2025ByteDances Platform for Reinforcement Learning from Human Feedback | Ray Summit 2024Dynamic Scheduling for Large Language Model Serving | Ray Summit 2024How Prime Intellect Builds Scalable Infrastructure for Agentic RL | Ray Summit 2025BentoML or RayServe, You Can Choose Both with BentoRayHow KubeRay Is Evolving for Massive AI Workloads  | Ray Summit 2025RDMA P2P Deep Dive: KvCache Transfer, Weight Updates & MoE Routing at Perplexity | Ray Summit 2025Ray Direct Transport: RDMA Support in Ray Core | Ray Summit 2025
Anyscale |

Hugging Face + vLLM: One Model Definition to Rule Them All | Ray Summit 2025

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER