New embedding model: Contextual Document Embeddings @Weaviate
New embedding model: Contextual Document Embeddings  @Weaviate
Uploaded November 2024 | Updated September 2026, 3 hours ago
Traditional document embeddings have a significant limitation: they encode documents independently, without considering their context or neighboring documents.

This means they have to choose a single global weighting for terms, potentially missing important contextual nuances, or overweighting terms that might occur a lot in the dataset. This can be problematic when embedding in different domains or contexts.

✨ The Solution: Contextual Document Embeddings (CDE) ✨

CDE operates in two stages:
1️⃣ Adversarial contrastive learning: batch and embed related context from neighboring documents
2️⃣ Embed the target document while considering the contextual embeddings of the related document batch

CDE can:
- Improve performance in domain-specific scenarios
- Better handle of out-of-domain queries

but also has the benefits of:
- No additional storage requirements during retrieval
- Maintains fast search capabilities

The approach has achieved state-of-the-art results on the MTEB benchmark: huggingface.co/spaces/mteb/leaderboard

Want to dive deeper? Check out the full research paper: arxiv.org/abs/2410.02525
Or try it out with this notebook: github.com/weaviate/recipes/blob/main/weaviate-features/services-research/contextual_document_embeddings.ipynb


▬▬▬▬▬▬▬▬▬▬▬▬ CONNECT WITH US ▬▬▬▬▬▬▬▬▬▬▬▬

- Visit weaviate.io
- Star us on GitHub github.com/weaviate/weaviate
- Stay updated and subscribe to our newsletter: newsletter.weaviate.io
- Try out Weaviate Cloud for free here: https://console.weaviate.cloud/

Got a question?
- Forum: forum.weaviate.io
- Slack: weaviate.io/slack

Connect with us on
- Twitter: twitter.com/weaviate_io
- LinkedIn: linkedin.com/company/weaviate-io
New embedding model: Contextual Document EmbeddingsWhat is a Vector Database?Recursive Language Models with Alex Zhang - Weaviate Podcast #142!AI-Native Development with Guy Podjarny and Bob van Luijt - Weaviate Podcast #102!Stop dumping chat history into your context window5. AI Agents Explained: LLMsSemantic Query Engines with Matthew Russo - Weaviate Podcast #131!OCR vs. Image Embeddings for PDF RAG: Which One is Better?Building AI Agents with TARS & WeaviateZain and JP chat about: Vector embedding models for AIAI Renaissance Berlin - The Grid editionAI Agents Explained: Making AI Actually WORK For You
Weaviate vector database |

New embedding model: Contextual Document Embeddings

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER