Search Agents with Nandan Thakur - Weaviate Podcast #137! @Weaviate
Search Agents with Nandan Thakur - Weaviate Podcast #137!  @Weaviate
Uploaded May 2026 | Updated September 2026, 2 hours ago
Dr. Nandan Thakur returns to the Weaviate Podcast fresh off defending his dissertation to discuss the evolution from neural retrieval to agentic search and his new work on Orbit, a synthetic training data pipeline for search agents. The conversation opens with reflections on his PhD journey, tracing the field's shift from ColBERT-style models and sparse retrievers through RAG and into today's agentic search paradigm where LLMs iteratively search, reason, and refine.

The discussion dives deep into how Orbit generates multi-hop, riddle-style training queries using DeepSeek's API on a personal laptop over four to six months, making high-quality search agent training data accessible without massive compute budgets. Thakur draws a sharp distinction between deep research (broad, multi-tool report generation) and search agents (focused on search and browse tools to answer specific questions), then connects Orbit's multi-hop queries to BrowseComp's filter-style riddles where each clue narrows the answer space like a funnel. The conversation explores the design of deep research harnesses, chunking strategies, Anthropic's contextual retrieval for entity disambiguation, context compaction to manage bloated agent contexts, and memory services like Weaviate's Engram for compressing search results between reasoning rounds.

From there, the episode tackles sequential versus parallel search trajectories, the pass@K approach to rollouts in GRPO training, and whether isolated trajectories should share progress through message passing. Thakur makes a compelling case for training search agents to produce keyword-focused queries optimized for BM25 versus semantic queries for dense retrieval: the idea that one query does not fit all search engines. The conversation closes on future directions: efficiency-focused Pareto frontiers for search agents, long-form report generation evaluation through TREC RAG, and the coming wave of multilingual and multimodal search benchmarks.

Chapters:
0:00 Welcome Nandan!
3:52 Nandan’s Ph.D. Journey
7:56 What are Search Agents?
16:24 Search Agent Benchmarks
23:02 Training Data for Search Agents
27:40 Deep Research Harness Design
40:52 ORBIT Data Generation
47:02 Search Agent Trajectories
57:40 Exciting Directions for AI

Links:
Benchmarks, Data, and Evaluation for Robust Retrieval and Retrieval-Augmented Generation on Heterogeneous Domains and Languages, Nandan Thakur Ph.D. Dissertation: uwspace.uwaterloo.ca/items/dc44caa2-f092-49d1-b527-da3a30097367
ORBIT: Scalable and Verifiable Data Generation for Search Agents on a Tight Budget: arxiv.org/pdf/2604.01195
Search Agents with Nandan Thakur - Weaviate Podcast #137!Weaviate 1.29 Release Highlights & Outlook9. The Future of AI AgentsIRPAPERS Explained!Agentic RAG vs Vanilla RAGRAGAS with Jithin James, Shahul Es, and Erika Cardenas - Weaviate Podcast #77!Data Agents with Shreya Shankar - Weaviate Podcast #135!DSPy and ColBERT with Omar Khattab! - Weaviate Podcast #85Advanced Chunking Techniques: Semantic & LLM-Based Chunking (Simply!) ExplainedThe Future of Search with Nils Reimers and Erika Cardenas - Weaviate Podcast #97!AI Assistant for Cyclists: Meet BAIKWhat brings you to the Hack Night in Berlin?
Weaviate vector database |

Search Agents with Nandan Thakur - Weaviate Podcast #137!

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER