CoServe: Max Performance, Minimal Compute | Ray Summit 2025 @anyscale
CoServe: Max Performance, Minimal Compute | Ray Summit 2025  @anyscale
Uploaded November 2025 | Updated September 2026, 3 weeks ago
At Ray Summit 2025, Chen Xia from Cohere shares how the company is building a scalable, efficient, and enterprise-ready AI platform—combining vLLM with state-of-the-art optimizations in quantization, kernel performance, and data communication to deliver low-latency, high-throughput inference at minimal compute cost.

Chen begins by outlining Cohere’s mission: providing private, secure, and high-performance AI solutions for enterprises. Achieving this requires an inference stack that preserves model accuracy while dramatically reducing hardware requirements—especially as context lengths grow and enterprise workloads demand predictable, low-latency behavior.

The talk dives into the core innovations behind Cohere’s serving infrastructure, including:

Accuracy-preserving low-bit quantization techniques that cut memory footprint and compute overhead without degrading model output quality

Extensive kernel optimizations built on top of vLLM to accelerate attention, sampling, and IO-heavy inference operations

High-efficiency data communication paths that reduce inter-GPU overhead and latency for large-context inference

A serving pipeline engineered for both low cost and high reliability, tailored for enterprise environments

Chen highlights real-world impact through Cohere’s Command A model series, which can be served on a single H100 GPU while supporting context lengths beyond 128K tokens—all while maintaining low-latency performance suitable for production applications like retrieval-augmented generation, agentic workflows, and enterprise assistants.

Attendees will gain a detailed understanding of how Cohere combines quantization, kernel engineering, and vLLM-based optimization to deliver secure, cost-efficient, production-grade LLM inference at scale.

Liked this video? Check out other Ray Summit breakout session recordings youtube.com/playlist?list=PLzTswPQNepXllnU0C36WtkC0dqkAoDulh

Subscribe to our YouTube channel to stay up-to-date on the future of AI! youtube.com/c/anyscale

🔗 Connect with us:
LinkedIn: linkedin.com/company/joinanyscale
X: https://x.com/anyscalecompute
Website: anyscale.com
CoServe: Max Performance, Minimal Compute | Ray Summit 2025How Ray Data Powers Scalable AI Workloads | Ray Summit 2025Inside Uber: Scaling Model Training with Ray | Ray Summit 2025Transforming Multimodal Data Management with LanceDB-Ray | Ray Summit 2024Why Ray Became a Distributed Computing Engine for Modern AIRLlib: Lessons from the V2 Stack and Road Ahead | Ray Summit 2025Pinterests ML Evolution: Distributed Training with Ray | Ray Summit 2024How xAI Scales Image & Video Processing with Ray | Ray Summit 2025Ion Stoica on Agentic Systems and AI Reliability | Ray on the Road – NYC 2025Anyscale on Azure: Build and deploy AI at scale in your own tenantAn Overview of CloudKitchenss Ray-Powered ML Platform | Ray Summit 2024Prompt Learning: A Reinforcement Learning-Inspired Approach to AI Optimization | Ray Summit 2025
Anyscale |

CoServe: Max Performance, Minimal Compute | Ray Summit 2025

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER