Uploaded December 2025 | Updated September 2026, 2 weeks ago
At Ray Summit 2025, Harry Mellor from Hugging Face shares how the new Transformers backend for vLLM enables teams to use the exact same model implementation for both training and inference—eliminating duplication, simplifying maintenance, and dramatically accelerating deployment workflows.
He begins by introducing the core idea: with the Transformers backend, vLLM can now run any Transformers-compatible model—whether officially merged into the library or fully custom—directly from its original model definition, but at vLLM-grade performance. This means your modeling code becomes the single source of truth, while vLLM handles high-speed inference, attention optimizations, and distributed scaling.
Harry then walks through:
What the Transformers backend is and how it works under the hood
How to make a model compatible, including tips for adapting custom or experimental architectures
How this unlocks unified codepaths, allowing one modeling codebase to serve both research experimentation and production-grade inference
The session highlights how this backend bridges the gap between training frameworks and high-performance inference engines—streamlining development workflows and reducing operational complexity.
Attendees will leave with a clear understanding of how to adopt the Transformers backend, how to bring custom models into vLLM, and how this new capability enables seamless transitions from model training to large-scale production serving.
Liked this video? Check out other Ray Summit breakout session recordings youtube.com/playlist?list=PLzTswPQNepXllnU0C36WtkC0dqkAoDulh
Subscribe to our YouTube channel to stay up-to-date on the future of AI! youtube.com/c/anyscale
🔗 Connect with us:
LinkedIn: linkedin.com/company/joinanyscale
At Ray Summit 2025, Harry Mellor from Hugging Face shares how the new Transformers backend for vLLM enables teams to use the exact same model implementation for both training and inference—eliminating duplication, simplifying maintenance, and dramatically accelerating deployment workflows.
He begins by introducing the core idea: with the Transformers backend, vLLM can now run any Transformers-compatible model—whether officially merged into the library or fully custom—directly from its original model definition, but at vLLM-grade performance. This means your modeling code becomes the single source of truth, while vLLM handles high-speed inference, attention optimizations, and distributed scaling.
Harry then walks through:
What the Transformers backend is and how it works under the hood
How to make a model compatible, including tips for adapting custom or experimental architectures
How this unlocks unified codepaths, allowing one modeling codebase to serve both research experimentation and production-grade inference
The session highlights how this backend bridges the gap between training frameworks and high-performance inference engines—streamlining development workflows and reducing operational complexity.
Attendees will leave with a clear understanding of how to adopt the Transformers backend, how to bring custom models into vLLM, and how this new capability enables seamless transitions from model training to large-scale production serving.
Liked this video? Check out other Ray Summit breakout session recordings youtube.com/playlist?list=PLzTswPQNepXllnU0C36WtkC0dqkAoDulh
Subscribe to our YouTube channel to stay up-to-date on the future of AI! youtube.com/c/anyscale
🔗 Connect with us:
LinkedIn: linkedin.com/company/joinanyscale










