Late chunking improves context recall in RAG pipelines @Weaviate
Late chunking improves context recall in RAG pipelines  @Weaviate
Uploaded October 2024 | Updated September 2026, 52 minutes ago
Optimizing your chunking techniques is one of the top places to improve performance in your RAG pipelines, but what’s the best one?

Jina AI just released a new method called late chunking that takes the same amount of storage space as naive chunking, but solves the problem of lost context, similarly to ColBERT.

You can implement it super easily with just a few extra lines in your embedding step!

Blog: weaviate.io/blog/late-chunking
Notebook: github.com/weaviate/recipes/blob/main/weaviate-features/services-research/late_chunking_berlin.ipynb

📄 Papers
Late Chunking: arxiv.org/pdf/2409.04701
ColBERT: arxiv.org/pdf/2004.12832

▬▬▬▬▬▬▬▬▬▬▬▬ CONNECT WITH US ▬▬▬▬▬▬▬▬▬▬▬▬
- Visit weaviate.io
- Star us on GitHub github.com/weaviate/weaviate
- Stay updated and subscribe to our newsletter: newsletter.weaviate.io
- Try out Weaviate Cloud for free here: https://console.weaviate.cloud/

Got a question?
- Forum: forum.weaviate.io
- Slack: weaviate.io/slack

Connect with us on
- Twitter: twitter.com/weaviate_io
- LinkedIn: linkedin.com/company/weaviate-io
Late chunking improves context recall in RAG pipelinesWhat is MCP (Model Context Protocol)?Humans and AI with John Maeda: AI-Native Databases #3How can AI and teachers revolutionize education together?Scaling Test-Time Compute in Search ModeRAG for Veterinarians (VetRec) with David de Matheu - Weaviate Podcast #92!3. The History of AI AgentsNew RAFT Consensus in 1.25Text-to-SQL is dead: The next generation of querying is AgenticGEPA Explained!1. An Introduction to AI AgentsLate Interaction combines the best of Keyword and Semantic Search
Weaviate vector database |

Late chunking improves context recall in RAG pipelines

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER