Drowning in Documents with Mathew Jacob - Weaviate Podcast #141! @Weaviate
Drowning in Documents with Mathew Jacob - Weaviate Podcast #141!  @Weaviate
Uploaded August 2026 | Updated September 2026, 2 hours ago
Mathew Jacob, lead author of "Drowning in Documents: Consequences of Scaling Reranker Inference" and now a PhD student in ML systems at the University of Washington, joins the Weaviate Podcast to unpack one of the most surprising results in modern search: cross-encoder rerankers get worse as you give them more documents. The paper began during his Databricks internship, where scaling reranking past roughly 100 documents sent recall@10 plummeting, a result so counterintuitive he assumed it was a bug.

The conversation digs into why this happens, reframing rerankers through the lens of boosting, rather than being strictly stronger than first-stage retrievers. Cross-encoders are very good at correcting retriever errors within the distribution they were trained on. Full-scoring experiments over 10,000 randomly sampled documents drive the point home, with BM25 beating state-of-the-art cross-encoders. From there, the discussion moves into phantom hits, cases where wildly irrelevant documents scored highly. For example, a dishwasher document surfacing for a query about disease in Gabonese children. We also discuss whether ensembling rerankers can patch these false positives.

The second half explores what comes next for reranking: prompt-based listwise reranking with sliding windows, which proved far more robust than pointwise scoring; RankZephyr-style fine-tuning versus encoding learning signal in prompts with GEPA and DSPy, reasoning rerankers like Rank1 and their latency trade-offs, hard negative mining behind ZeroEntropy's zELO, and pairwise and setwise designs that sit between cross-encoders and full listwise ranking. Adaptive retrieval comes into focus through Natural Language Query to Configuration for Retrieval Agents, predicting per query whether to run simple retrieval, multi-hop, or full agentic search to push the cost-quality frontier.

The conversation lands on TraceLab, from Mathew's lab at UW: 40,000 real traces harvested from Claude Code and Codex usage, revealing how coding agents actually behave, prefix cache patterns, long-tailed tool calls, and how understanding these workloads unlocks the next generation of serving optimizations.

Links:
Drowning in Documents: arxiv.org/pdf/2411.11767
SyFI TraceLab: https://tracelab.cs.washington.edu/
Natural Language Query to Configuration for Retrieval Agents: arxiv.org/pdf/2605.27361
Optimizing Compound Retrieval Systems: arxiv.org/pdf/2504.12063

Chapters
0:00 Welcome Mathew!
0:59 An Overview of Drowning in Documents
7:29 Reranking Deeper Pools
15:22 Phantom Hits from Cross Encoders
21:59 Listwise Rerankers
34:12 Next-Generation Cross Encoders
40:29 Ranking Cascades
48:39 Exciting Directions for AI
51:12 Understand Coding Agents with SyFI TraceLab
Drowning in Documents with Mathew Jacob - Weaviate Podcast #141!10. Use Cases of AI AgentsAI Agents That Matter with Sayash Kapoor and Benedikt Stroebl - Weaviate Podcast #104!ACORN for 10x faster filtered vector searchesWeaviate 1.25 | Release AnnouncementInstructor with Jason Liu - Weaviate Podcast #88!Can we win Europes biggest Hackathon?Self-Discover in DSPy with Chris Dossman - Weaviate Podcast #90!Tanmay Chopra on Emissary - Weaviate Podcast #75!DSPy + Weaviate for the Next Generation of LLM AppsRudy Lai on Tactic Generate - Weaviate Podcast #78!New embedding model: Contextual Document Embeddings
Weaviate vector database |

Drowning in Documents with Mathew Jacob - Weaviate Podcast #141!

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER