Distributed Embeddings at Scale: Processing 10M+ Rows/ Day with Ray, GPUs & Qdrant | Ray Summit 2025 @anyscale
Distributed Embeddings at Scale: Processing 10M+ Rows/ Day with Ray, GPUs & Qdrant | Ray Summit 2025  @anyscale
Uploaded December 2025 | Updated September 2026, 1 week ago
At Ray Summit 2025, Justin Miller from ZEFR shares how his team built a production-grade, multi-platform NLP pipeline using Ray and GPU acceleration to process millions of social media posts across TikTok, YouTube, and Instagram.

He begins by describing the challenges of handling massive, fast-changing content streams across multiple platforms—each with unique data formats, ingestion patterns, and quality constraints. To meet these demands, ZEFR engineered a robust distributed pipeline that uses Ray to orchestrate scalable embedding generation, GPU-heavy processing, and high-throughput vector search ingestion.

Justin walks through the architecture step-by-step:

Snowflake → Ray ingestion: Retrieve rows for each platform with consistent batch scheduling

Cleaning, chunking, and preprocessing: Normalize and prepare multimodal content at scale

Distributed embedding generation: Use Ray Actors to shard GPU inference tasks across the cluster

High-throughput writes: Send results to Google Cloud Storage (GCS), Qdrant for vector search, and back to Snowflake for analytics and pipeline tracking

Shard lifecycle management: Delete stale shards, manage multi-platform ingestion, and maintain healthy storage footprints

He also shares practical, real-world guidance for operating Ray in production—covering deployment patterns, debugging tips, failure recovery, throughput tuning, and cost management.

Whether you’re processing large multi-source datasets, running GPU-heavy inference pipelines, or building modern vector-search–backed systems, this talk provides both code-level insights and actionable advice for running Ray at scale.

Liked this video? Check out other Ray Summit breakout session recordings youtube.com/playlist?list=PLzTswPQNepXllnU0C36WtkC0dqkAoDulh

Subscribe to our YouTube channel to stay up-to-date on the future of AI! youtube.com/c/anyscale

🔗 Connect with us:
LinkedIn: linkedin.com/company/joinanyscale
Distributed Embeddings at Scale: Processing 10M+ Rows/ Day with Ray, GPUs & Qdrant | Ray Summit 2025Applied Intuition’s Blueprint for Scalable RL + Batch Inference | Ray Summit 2025Building Scalable Cross-Modal Search with Ray | Ray Summit 2024Ray at IBM: Transforming Large-Scale Data Processing for AI and Science | Ray Summit 2024Elastic Expert Parallelism for vLLM | Ray Summit 2025Apple’s Approach to Scalable Machine Learning Infrastructure on Ray | Ray Summit 2025How Uber Optimize Marketplaces with Ray | Ray Summit 2024Improved Scheduling Flexibility with Label Selectors in Ray | Ray Summit 2025Fighting Fire with Algorithms: Lockheeds RL-Based Wildfire Solution | Ray Summit 2024Meet verl: An RL Framework for LLM Reasoning & Tool Use | Ray Summit 2025How BMW Scales Automotive AI Workloads with the Ray Framework | Ray Summit 2025Introduction to Anyscale and Ray AI Libraries
Anyscale |

Distributed Embeddings at Scale: Processing 10M+ Rows/ Day with Ray, GPUs & Qdrant | Ray Summit 2025

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER