Uploaded November 2025 | Updated September 2026, 3 weeks ago
At Ray Summit 2025, Chen Xia from Cohere shares how the company is building a scalable, efficient, and enterprise-ready AI platform—combining vLLM with state-of-the-art optimizations in quantization, kernel performance, and data communication to deliver low-latency, high-throughput inference at minimal compute cost.
Chen begins by outlining Cohere’s mission: providing private, secure, and high-performance AI solutions for enterprises. Achieving this requires an inference stack that preserves model accuracy while dramatically reducing hardware requirements—especially as context lengths grow and enterprise workloads demand predictable, low-latency behavior.
The talk dives into the core innovations behind Cohere’s serving infrastructure, including:
Accuracy-preserving low-bit quantization techniques that cut memory footprint and compute overhead without degrading model output quality
Extensive kernel optimizations built on top of vLLM to accelerate attention, sampling, and IO-heavy inference operations
High-efficiency data communication paths that reduce inter-GPU overhead and latency for large-context inference
A serving pipeline engineered for both low cost and high reliability, tailored for enterprise environments
Chen highlights real-world impact through Cohere’s Command A model series, which can be served on a single H100 GPU while supporting context lengths beyond 128K tokens—all while maintaining low-latency performance suitable for production applications like retrieval-augmented generation, agentic workflows, and enterprise assistants.
Attendees will gain a detailed understanding of how Cohere combines quantization, kernel engineering, and vLLM-based optimization to deliver secure, cost-efficient, production-grade LLM inference at scale.
Liked this video? Check out other Ray Summit breakout session recordings youtube.com/playlist?list=PLzTswPQNepXllnU0C36WtkC0dqkAoDulh
Subscribe to our YouTube channel to stay up-to-date on the future of AI! youtube.com/c/anyscale
🔗 Connect with us:
LinkedIn: linkedin.com/company/joinanyscale
X: https://x.com/anyscalecompute
Website: anyscale.com
At Ray Summit 2025, Chen Xia from Cohere shares how the company is building a scalable, efficient, and enterprise-ready AI platform—combining vLLM with state-of-the-art optimizations in quantization, kernel performance, and data communication to deliver low-latency, high-throughput inference at minimal compute cost.
Chen begins by outlining Cohere’s mission: providing private, secure, and high-performance AI solutions for enterprises. Achieving this requires an inference stack that preserves model accuracy while dramatically reducing hardware requirements—especially as context lengths grow and enterprise workloads demand predictable, low-latency behavior.
The talk dives into the core innovations behind Cohere’s serving infrastructure, including:
Accuracy-preserving low-bit quantization techniques that cut memory footprint and compute overhead without degrading model output quality
Extensive kernel optimizations built on top of vLLM to accelerate attention, sampling, and IO-heavy inference operations
High-efficiency data communication paths that reduce inter-GPU overhead and latency for large-context inference
A serving pipeline engineered for both low cost and high reliability, tailored for enterprise environments
Chen highlights real-world impact through Cohere’s Command A model series, which can be served on a single H100 GPU while supporting context lengths beyond 128K tokens—all while maintaining low-latency performance suitable for production applications like retrieval-augmented generation, agentic workflows, and enterprise assistants.
Attendees will gain a detailed understanding of how Cohere combines quantization, kernel engineering, and vLLM-based optimization to deliver secure, cost-efficient, production-grade LLM inference at scale.
Liked this video? Check out other Ray Summit breakout session recordings youtube.com/playlist?list=PLzTswPQNepXllnU0C36WtkC0dqkAoDulh
Subscribe to our YouTube channel to stay up-to-date on the future of AI! youtube.com/c/anyscale
🔗 Connect with us:
LinkedIn: linkedin.com/company/joinanyscale
X: https://x.com/anyscalecompute
Website: anyscale.com










