Uploaded December 2025 | Updated September 2026, 1 week ago
At Ray Summit 2025, Justin Miller from ZEFR shares how his team built a production-grade, multi-platform NLP pipeline using Ray and GPU acceleration to process millions of social media posts across TikTok, YouTube, and Instagram.
He begins by describing the challenges of handling massive, fast-changing content streams across multiple platforms—each with unique data formats, ingestion patterns, and quality constraints. To meet these demands, ZEFR engineered a robust distributed pipeline that uses Ray to orchestrate scalable embedding generation, GPU-heavy processing, and high-throughput vector search ingestion.
Justin walks through the architecture step-by-step:
Snowflake → Ray ingestion: Retrieve rows for each platform with consistent batch scheduling
Cleaning, chunking, and preprocessing: Normalize and prepare multimodal content at scale
Distributed embedding generation: Use Ray Actors to shard GPU inference tasks across the cluster
High-throughput writes: Send results to Google Cloud Storage (GCS), Qdrant for vector search, and back to Snowflake for analytics and pipeline tracking
Shard lifecycle management: Delete stale shards, manage multi-platform ingestion, and maintain healthy storage footprints
He also shares practical, real-world guidance for operating Ray in production—covering deployment patterns, debugging tips, failure recovery, throughput tuning, and cost management.
Whether you’re processing large multi-source datasets, running GPU-heavy inference pipelines, or building modern vector-search–backed systems, this talk provides both code-level insights and actionable advice for running Ray at scale.
Liked this video? Check out other Ray Summit breakout session recordings youtube.com/playlist?list=PLzTswPQNepXllnU0C36WtkC0dqkAoDulh
Subscribe to our YouTube channel to stay up-to-date on the future of AI! youtube.com/c/anyscale
🔗 Connect with us:
LinkedIn: linkedin.com/company/joinanyscale
At Ray Summit 2025, Justin Miller from ZEFR shares how his team built a production-grade, multi-platform NLP pipeline using Ray and GPU acceleration to process millions of social media posts across TikTok, YouTube, and Instagram.
He begins by describing the challenges of handling massive, fast-changing content streams across multiple platforms—each with unique data formats, ingestion patterns, and quality constraints. To meet these demands, ZEFR engineered a robust distributed pipeline that uses Ray to orchestrate scalable embedding generation, GPU-heavy processing, and high-throughput vector search ingestion.
Justin walks through the architecture step-by-step:
Snowflake → Ray ingestion: Retrieve rows for each platform with consistent batch scheduling
Cleaning, chunking, and preprocessing: Normalize and prepare multimodal content at scale
Distributed embedding generation: Use Ray Actors to shard GPU inference tasks across the cluster
High-throughput writes: Send results to Google Cloud Storage (GCS), Qdrant for vector search, and back to Snowflake for analytics and pipeline tracking
Shard lifecycle management: Delete stale shards, manage multi-platform ingestion, and maintain healthy storage footprints
He also shares practical, real-world guidance for operating Ray in production—covering deployment patterns, debugging tips, failure recovery, throughput tuning, and cost management.
Whether you’re processing large multi-source datasets, running GPU-heavy inference pipelines, or building modern vector-search–backed systems, this talk provides both code-level insights and actionable advice for running Ray at scale.
Liked this video? Check out other Ray Summit breakout session recordings youtube.com/playlist?list=PLzTswPQNepXllnU0C36WtkC0dqkAoDulh
Subscribe to our YouTube channel to stay up-to-date on the future of AI! youtube.com/c/anyscale
🔗 Connect with us:
LinkedIn: linkedin.com/company/joinanyscale










