How Alibaba Cloud Accelerates AI Pipelines with AnalyticDB Ray | Ray Summit 2025 @anyscale
How Alibaba Cloud Accelerates AI Pipelines with AnalyticDB Ray | Ray Summit 2025  @anyscale
Uploaded November 2025 | Updated September 2026, 2 weeks ago
At Ray Summit 2025, Liang Lin and Fei Xue from Alibaba Cloud share how AnalyticDB Ray is enabling high-performance, multi-modal AI pipelines directly inside the data warehouse—accelerating everything from ETL to large-scale inference and distributed fine-tuning.

They begin by outlining why efficient multi-modal data processing is essential for modern AI systems, and how AnalyticDB Ray brings Ray’s distributed compute capabilities into a cloud-native data warehouse environment. The session then highlights three real-world applications demonstrating this integration in action:

Optimizing Advertising Recommendation Inference
Learn how AnalyticDB Ray powers offline CTR estimation pipelines using heterogeneous compute. CPU and GPU resources are auto-scaled independently to maximize utilization—boosting GPU usage from 5% to 40%—while dynamic storage scaling accelerates data processing by 2–3X.

Accelerating LLM Offline Batch Inference & Data Distillation
Discover how Ray Data and vLLM/SGLang are used to distill datasets from large models such as Qwen and DeepSeek for downstream training. Benefits include 2–3× faster data loading through caching, efficient scheduling of 40,000 fine-grained tasks in a single Ray cluster, and a 50% performance gain for DeepSeek INT8 quantization versus FP8 in offline distillation.

Efficient Distributed Fine-tuning of Multi-modal Models
See how AnalyticDB Ray integrates with Lance and Ray Data to process large-scale image–text datasets for personalized multimodal scenarios. With seamless integration into LLaMA-Factory, teams achieve 3–5× faster distributed fine-tuning for Qwen-VL models—offering an end-to-end workflow from data labeling to training.

Liang and Fei show how these use cases reveal the power of in-warehouse AI pipelines, where multi-modal ETL and machine learning run side-by-side to shorten the path from raw data to intelligent decision-making.

Liked this video? Check out other Ray Summit breakout session recordings
youtube.com/playlist?list=PLzTswPQNepXllnU0C36WtkC0dqkAoDulh

Subscribe to our YouTube channel to stay up-to-date on the future of AI!
youtube.com/c/anyscale

🔗 Connect with us:
LinkedIn: linkedin.com/company/joinanyscale
X: https://x.com/anyscalecompute
Website: anyscale.com
How Alibaba Cloud Accelerates AI Pipelines with AnalyticDB Ray | Ray Summit 2025How Vivix Scales Video Ad Classification with Ray | Ray Summit 2024Ray Summit 2025 Keynote: Building Cursor Composer with Sasha RushRay Summit 2025: Bringing AI to the Physical World with Chelsea Finn from Physical IntelligenceFrom Batch to Personalized: Attentives Journey with Ray and Anyscale | Ray on the Road – NYC 2025High-Performance LLM Serving on Intel: vLLM for XPU, HPU & CPU | Ray Summit 2025Ray Meetup: LLMs on Ray + LanceDBStreamlining AI Workflows with Apache Airflow and Ray | Ray Summit 2024How Nubank Uses Foundation Models for Financial Data | Ray on the Road – NYC 2025Inside Netflix’s Mako: The Next-Gen ML Training Platform | Ray Summit 2025The Evolution of Multi-GPU Inference in vLLM | Ray Summit 2024The LLM-Cloud Synergy: NebiusAIs Insider Perspective | Ray Summit 2024
Anyscale |

How Alibaba Cloud Accelerates AI Pipelines with AnalyticDB Ray | Ray Summit 2025

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER