Uploaded November 2025 | Updated September 2026, 2 weeks ago
This video shows how to retrieve diverse, high-value facts for RAG pipelines in a practical way. We walk through the classic MMR (lambda knob) approach that balances relevance vs. diversity, its limits (duplicates, tuning headaches), and a more principled view based on information gain and uncertainty. Using a dartboard analogy, we explain why K-nearest neighbors can produce redundant blobs of evidence, how modeling embedding uncertainty (sigma) changes your selection strategy, and why a small combinatorial optimization can pick a set that maximizes the chance the LLM actually has the right fact in context.
Perfect for ML engineers and prompt engineers building production retrieval systems. Drop a comment with your RAG pain points.
#RAG #MMR #RetrievalAugmentedGeneration #Embeddings #CrossEncoder #LLM #InformationRetrieval #NLP #PromptEngineering #MachineLearning #DataScience
This video shows how to retrieve diverse, high-value facts for RAG pipelines in a practical way. We walk through the classic MMR (lambda knob) approach that balances relevance vs. diversity, its limits (duplicates, tuning headaches), and a more principled view based on information gain and uncertainty. Using a dartboard analogy, we explain why K-nearest neighbors can produce redundant blobs of evidence, how modeling embedding uncertainty (sigma) changes your selection strategy, and why a small combinatorial optimization can pick a set that maximizes the chance the LLM actually has the right fact in context.
Perfect for ML engineers and prompt engineers building production retrieval systems. Drop a comment with your RAG pain points.
#RAG #MMR #RetrievalAugmentedGeneration #Embeddings #CrossEncoder #LLM #InformationRetrieval #NLP #PromptEngineering #MachineLearning #DataScience





