Wisedocs’ Journey: Rebuilding & Accelerating ML with KubeRay | Ray Summit 2025 @anyscale
Wisedocs’ Journey: Rebuilding & Accelerating ML with KubeRay | Ray Summit 2025  @anyscale
Uploaded November 2025 | Updated September 2026, 2 weeks ago
At Ray Summit 2025, Denys Linkov from Wisedocs shares how the company rebuilt its ML and AI serving layer on KubeRay—unlocking major gains in performance, cost efficiency, and deployment velocity.

He begins by outlining the challenges and opportunities involved in redesigning a production-grade serving stack. Wisedocs migrated 10 models to KubeRay to power both real-time and batch workloads, achieving a 50% reduction in cost while improving throughput by 10×. Denys walks through the architectural decisions that made this possible, from compute orchestration to workload isolation and scaling strategies.

The talk also covers the information and technical architecture that enabled Wisedocs to cut its time-to-production dramatically—from one month to just two days. Denys highlights how internal abstractions, standardized deployment patterns, and Ray’s distributed execution model streamlined the entire development-to-production cycle.

Finally, he discusses the tradeoffs of serving GenAI models versus encoder-based models within an internal Kubernetes environment, sharing lessons learned on performance tuning, resource management, and operational complexity.

Attendees will gain practical insights into modernizing ML serving stacks with Ray and KubeRay, balancing efficiency with reliability, and accelerating production deployment at scale.

Subscribe to our YouTube channel to stay up-to-date on the future of AI! youtube.com/c/anyscale

🔗 Connect with us:
LinkedIn: linkedin.com/company/joinanyscale
X: https://x.com/anyscalecompute
Website: anyscale.com
Wisedocs’ Journey: Rebuilding & Accelerating ML with KubeRay | Ray Summit 2025[Ray Meetup] Ray + vLLM in Action: Lessons from Pinterest and Large Scale Distributed InferenceHow Zoox Built a Reliable, High-Velocity Model Serving Platform with Ray Serve | Ray Summit 2025Scaling Machine Learning at Tripadvisor: Our Journey with Ray and Anyscale | Ray Summit 2025AWS + vLLM: Building the Future of Open, Fast LLM Serving | Ray Summit 2025Ray + vLLM  Efficient Multi Node Orchestration for Sparse MoE Model Serving | Ray Summit 2025Hybrid RL + Imitation Learning for Robotics with Ray at RAI InstituteHow Runhouse Orchestrates Multi-Cluster Ray Workloads | Ray Summit 2025How vLLM and Ray Work TogetherCoinbases ML Training Evolution: From Sagemaker to Ray | Ray Summit 2024Secure & Scalable AI on Ray + Kubernetes: Google’s Decoupled Agent Pattern | Ray Summit 2025vLLM TPU: A new unified-backend supporting Pytorch and JAX natively on TPU | Ray Summit 2025
Anyscale |

Wisedocs’ Journey: Rebuilding & Accelerating ML with KubeRay | Ray Summit 2025

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER