Uploaded November 2025 | Updated September 2026, 2 weeks ago
Slides: drive.google.com/file/d/1G3DPYUd9i5dxsGwI9QjAd7NJe0jMDah4/view?usp=sharing
At Ray Summit 2025, Balaji Veeramani from Anyscale shares how Ray Data has evolved into one of the most widely used libraries in the Ray ecosystem—purpose-built for the new generation of AI workloads.
Unlike traditional data processing engines, Ray Data is designed from the ground up for multimodal, accelerator-native, and AI-centric pipelines. In this talk, the speakers provide an overview of Ray Data’s core capabilities and highlight the major features added over the past year to support:
Large-scale batch inference across GPUs and clusters
Distributed training data preparation and ingestion for massive models
High-performance multimodal data processing, spanning images, video, text, and more
Whether you're building LLM pipelines, multimodal training workflows, or high-throughput inference systems, this session provides a clear look at how Ray Data powers modern AI at scale.
Liked this video? Check out other Ray Summit breakout session recordings youtube.com/playlist?list=PLzTswPQNepXllnU0C36WtkC0dqkAoDulh
Subscribe to our YouTube channel to stay up-to-date on the future of AI! youtube.com/c/anyscale
🔗 Connect with us:
LinkedIn: linkedin.com/company/joinanyscale
X: https://x.com/anyscalecompute
Website: anyscale.com
Slides: drive.google.com/file/d/1G3DPYUd9i5dxsGwI9QjAd7NJe0jMDah4/view?usp=sharing
At Ray Summit 2025, Balaji Veeramani from Anyscale shares how Ray Data has evolved into one of the most widely used libraries in the Ray ecosystem—purpose-built for the new generation of AI workloads.
Unlike traditional data processing engines, Ray Data is designed from the ground up for multimodal, accelerator-native, and AI-centric pipelines. In this talk, the speakers provide an overview of Ray Data’s core capabilities and highlight the major features added over the past year to support:
Large-scale batch inference across GPUs and clusters
Distributed training data preparation and ingestion for massive models
High-performance multimodal data processing, spanning images, video, text, and more
Whether you're building LLM pipelines, multimodal training workflows, or high-throughput inference systems, this session provides a clear look at how Ray Data powers modern AI at scale.
Liked this video? Check out other Ray Summit breakout session recordings youtube.com/playlist?list=PLzTswPQNepXllnU0C36WtkC0dqkAoDulh
Subscribe to our YouTube channel to stay up-to-date on the future of AI! youtube.com/c/anyscale
🔗 Connect with us:
LinkedIn: linkedin.com/company/joinanyscale
X: https://x.com/anyscalecompute
Website: anyscale.com










