Uploaded November 2024 | Updated September 2026, 3 hours ago
Traditional document embeddings have a significant limitation: they encode documents independently, without considering their context or neighboring documents.
This means they have to choose a single global weighting for terms, potentially missing important contextual nuances, or overweighting terms that might occur a lot in the dataset. This can be problematic when embedding in different domains or contexts.
✨ The Solution: Contextual Document Embeddings (CDE) ✨
CDE operates in two stages:
1️⃣ Adversarial contrastive learning: batch and embed related context from neighboring documents
2️⃣ Embed the target document while considering the contextual embeddings of the related document batch
CDE can:
- Improve performance in domain-specific scenarios
- Better handle of out-of-domain queries
but also has the benefits of:
- No additional storage requirements during retrieval
- Maintains fast search capabilities
The approach has achieved state-of-the-art results on the MTEB benchmark: huggingface.co/spaces/mteb/leaderboard
Want to dive deeper? Check out the full research paper: arxiv.org/abs/2410.02525
Or try it out with this notebook: github.com/weaviate/recipes/blob/main/weaviate-features/services-research/contextual_document_embeddings.ipynb
▬▬▬▬▬▬▬▬▬▬▬▬ CONNECT WITH US ▬▬▬▬▬▬▬▬▬▬▬▬
- Visit weaviate.io
- Star us on GitHub github.com/weaviate/weaviate
- Stay updated and subscribe to our newsletter: newsletter.weaviate.io
- Try out Weaviate Cloud for free here: https://console.weaviate.cloud/
Got a question?
- Forum: forum.weaviate.io
- Slack: weaviate.io/slack
Connect with us on
- Twitter: twitter.com/weaviate_io
- LinkedIn: linkedin.com/company/weaviate-io
Traditional document embeddings have a significant limitation: they encode documents independently, without considering their context or neighboring documents.
This means they have to choose a single global weighting for terms, potentially missing important contextual nuances, or overweighting terms that might occur a lot in the dataset. This can be problematic when embedding in different domains or contexts.
✨ The Solution: Contextual Document Embeddings (CDE) ✨
CDE operates in two stages:
1️⃣ Adversarial contrastive learning: batch and embed related context from neighboring documents
2️⃣ Embed the target document while considering the contextual embeddings of the related document batch
CDE can:
- Improve performance in domain-specific scenarios
- Better handle of out-of-domain queries
but also has the benefits of:
- No additional storage requirements during retrieval
- Maintains fast search capabilities
The approach has achieved state-of-the-art results on the MTEB benchmark: huggingface.co/spaces/mteb/leaderboard
Want to dive deeper? Check out the full research paper: arxiv.org/abs/2410.02525
Or try it out with this notebook: github.com/weaviate/recipes/blob/main/weaviate-features/services-research/contextual_document_embeddings.ipynb
▬▬▬▬▬▬▬▬▬▬▬▬ CONNECT WITH US ▬▬▬▬▬▬▬▬▬▬▬▬
- Visit weaviate.io
- Star us on GitHub github.com/weaviate/weaviate
- Stay updated and subscribe to our newsletter: newsletter.weaviate.io
- Try out Weaviate Cloud for free here: https://console.weaviate.cloud/
Got a question?
- Forum: forum.weaviate.io
- Slack: weaviate.io/slack
Connect with us on
- Twitter: twitter.com/weaviate_io
- LinkedIn: linkedin.com/company/weaviate-io











- Star us on GitHub https://github.com/weaviate/weaviate
- Stay updated and subscribe to our newsletter: https://newsletter.weaviate.io/
- Try out Weaviate Cloud Services for free here: https://console.weaviate.cloud/
Got a question?
- Forum: https://forum.weaviate.io/
- Slack: https://weaviate.io/slack
Connect with us on
- Twitter: https://twitter.com/weaviate_io
- LinkedIn: https://www.linkedin.com/company/weaviate-io/ AI Agents Explained: Making AI Actually WORK For You](https://i.ytimg.com/vi/lzlvi6oxQnA/mqdefault.jpg)