Uploaded April 2026 | Updated September 2026, 1 week ago
Talk by Justin Miller
socallinuxexpo.org/scale/23x/presentations/distributed-embeddings-scale-processing-10-million-rows-day-ray-and-gpus
In this talk, we’ll describe a production-grade NLP pipeline that processes millions of pieces of social media content across TikTok, YouTube, and Instagram using Ray and GPU acceleration. Learn how we use Ray's distributed computing model to orchestrate scalable embedding generation, sharded batch writes to Qdrant for vector search, and end-to-end pipeline tracking with Snowflake. We'll also talk about selecting a vector store and how to best evaluate the many options available.
Talk by Justin Miller
socallinuxexpo.org/scale/23x/presentations/distributed-embeddings-scale-processing-10-million-rows-day-ray-and-gpus
In this talk, we’ll describe a production-grade NLP pipeline that processes millions of pieces of social media content across TikTok, YouTube, and Instagram using Ray and GPU acceleration. Learn how we use Ray's distributed computing model to orchestrate scalable embedding generation, sharded batch writes to Qdrant for vector search, and end-to-end pipeline tracking with Snowflake. We'll also talk about selecting a vector store and how to best evaluate the many options available.










