Uploaded May 2025 | Updated September 2026, 1 week ago
Scaling RAG isn’t just about the model – it’s about managing the messy data pipeline. In this session, we’ll break down where most RAG pipelines fail and show how to build a scalable, production-ready workflow from document to embedding using Ray and Anyscale. Includes a live demo, prompt tips, and product updates for better batch and real-time inference.
What you’ll learn:
✅ Common RAG pipeline pitfalls and how to avoid them
✅ Scalable strategies for ingestion, chunking, and embedding
✅ Prompting best practices for guardrails and citations
✅ New tools to boost offline and real-time LLM performance
Perfect for AI engineers, architects, and decision makers building enterprise-grade RAG systems.
Scaling RAG isn’t just about the model – it’s about managing the messy data pipeline. In this session, we’ll break down where most RAG pipelines fail and show how to build a scalable, production-ready workflow from document to embedding using Ray and Anyscale. Includes a live demo, prompt tips, and product updates for better batch and real-time inference.
What you’ll learn:
✅ Common RAG pipeline pitfalls and how to avoid them
✅ Scalable strategies for ingestion, chunking, and embedding
✅ Prompting best practices for guardrails and citations
✅ New tools to boost offline and real-time LLM performance
Perfect for AI engineers, architects, and decision makers building enterprise-grade RAG systems.










