Weaviate • Vector DatabaseBen walks us through Box's three-layer infrastructure puzzle: First, the mind-boggling base infrastructure (think millions of interactions per second and trillions of files). Second, their unique multi-tenant security challenge - unlike most SaaS platforms, Box users share content across company boundaries, making traditional tenant isolation impossible. And third, ensuring AI respects all these complex permissions while still delivering value. The podcast then dives further into how vector embeddings can balloon file sizes - a few hundred bytes of text can require 4-6KB of vector data storage! We also dig into why RAG remains essential despite growing context windows, and how Box is developing AI agents that transform painful enterprise processes like RFP responses.
A fun, insightful look at what happens when cutting-edge AI meets enterprise-grade security and scale requirements!
Chapters 0:00 Weaviate Podcast #120! 0:20 Welcome Ben! 1:10 Exabyte Scale at Box 8:10 Infrastructure Rent vs. Buy 12:30 Founder-Led AI Adoption 19:37 Embeddings, RAG, and The AI Tipping Point 28:05 How many Vectors per Data Object? 37:20 Storage Tiers and Dynamic Indices in Retrieval 45:00 Agents 51:15 Agents for RFPs
Box AI with Ben Kus and Bob van Luijt - Weaviate Podcast #120!Weaviate • Vector Database2025-05-07 | Ben walks us through Box's three-layer infrastructure puzzle: First, the mind-boggling base infrastructure (think millions of interactions per second and trillions of files). Second, their unique multi-tenant security challenge - unlike most SaaS platforms, Box users share content across company boundaries, making traditional tenant isolation impossible. And third, ensuring AI respects all these complex permissions while still delivering value. The podcast then dives further into how vector embeddings can balloon file sizes - a few hundred bytes of text can require 4-6KB of vector data storage! We also dig into why RAG remains essential despite growing context windows, and how Box is developing AI agents that transform painful enterprise processes like RFP responses.
A fun, insightful look at what happens when cutting-edge AI meets enterprise-grade security and scale requirements!
Chapters 0:00 Weaviate Podcast #120! 0:20 Welcome Ben! 1:10 Exabyte Scale at Box 8:10 Infrastructure Rent vs. Buy 12:30 Founder-Led AI Adoption 19:37 Embeddings, RAG, and The AI Tipping Point 28:05 How many Vectors per Data Object? 37:20 Storage Tiers and Dynamic Indices in Retrieval 45:00 Agents 51:15 Agents for RFPsWeaviate TECH Hands-On: Personalization AgentWeaviate • Vector Database2025-07-11 | In this hands-on workshop session JP dives into building personalized recommendation systems and user personalized queries for the Weaviate vector database using the Weaviate Personalization Agent.
00:00:00 Intro 00:02:14 Environment Setup 00:12:35 Intro to Personalization Agent 00:17:38 Hands-On part: How to use the Personalization Agent 00:50:07 Summary and further resources
▬▬▬▬▬▬▬▬▬▬▬▬ CONNECT WITH US ▬▬▬▬▬▬▬▬▬▬▬▬
- Visit [weaviate.io](http://weaviate.io/) - Star us on GitHub github.com/weaviate/weaviate - Stay updated and subscribe to our newsletter: newsletter.weaviate.io - Try out Weaviate Cloud Services for free here: https://console.weaviate.cloud/
- Twitter: twitter.com/weaviate_io - LinkedIn: linkedin.com/company/weaviate-ioAgentic Topic Modeling with Maarten Grootendorst - Weaviate Podcast #126!Weaviate • Vector Database2025-07-09 | Maarten Grootendorst is a psychologist turned AI engineer who has created BERTopic and authored "Hands-On Large Language Models" with Jay Alammar. The rise of LLMs and Agents are transforming many areas of software! This podcast dives deep into their impact on Topic Modeling! Maarten designed BERTopic from the start with modularity in mind -- letting you ablate embedding models, dimensionality reduction, clustering algorithms, and more. This early insight to prioritize modularity makes BERTopic incredibly well structured to become more "Agentic". An "Agentic" Topic Modeling algorithm can use LLMs to generate topics or topic descriptions, as well as contrast them with other topics. It can decide which topics to subdivide, and it can integrate human feedback and evaluate topics in novel ways... I hope you find the podcast interesting!
Chapters 0:00 Welcome Maarten! 1:57 Hands-On Large Language Models 7:34 An Overview of Topic Modeling 10:45 LLM Topic Generation 17:13 Topic Modeling with Human Feedback 21:33 Topic Granularity 26:00 Visualizing Topics 31:18 Contrastive Topics 33:24 LLM-as-Judge for Topics 39:06 Separating Generation from Assignment 44:20 Applications of Topic Modeling 55:14 Semi-Supervised BERTopic 1:01:06 Future Directions for AIWeaviate TECH Hands-On: Query Agent in JavaScriptWeaviate • Vector Database2025-07-07 | In this session, Daniel walks you through our Weaviate Query Agent and uses the agent to build an Agentic RAG system in JavaScript
00:00:00 Intro & Overview 00:01:15 What is the Weaviate Query Agent 00:07:37 Live Coding with Weaviate Query Agent 00:21:00 Further Resources
Connect with us on - Twitter: twitter.com/weaviate_io - LinkedIn: linkedin.com/company/weaviate-ioSufficient Context with Hailey Joren - Weaviate Podcast #125!Weaviate • Vector Database2025-07-02 | Hailey Joren is a Ph.D. student at UCSD! Hailey and collaborators at Duke University and Google have recently published Sufficient Context: A New Lens on Retrieval Augmented Generation Systems in ICLR 2025! There are so many interesting nuggets to this work! Firstly, it really helped me understand the difference between *relevant* search results and sufficient context for answering the question. Armed with this lens of looking at retrieved context, Hailey and collaborators make all sorts of interesting observations about the current state of Hallucination. RAG unfortunately makes the models far less likely to hallucinate, and the existing RAG benchmarks unfortunately do not emphasize retrieval adaptation well enough -- indicated by LLMs outputting correct answers despite insufficient context 35-62% of the time! However, reason for optimism! Hailey and team develop an autorater that can detect insufficient context 93% of the time! There are all sorts of interesting ideas around this paper! I really hope you find the podcast useful!
Links: Sufficient Context: A New Lens on Retrieval Augmented Generation Systems - arxiv.org/pdf/2411.06037
Chapters 0:00 Welcome Hailey 1:40 Definition of Sufficient Context 4:47 Context Engineering 8:42 AutoRater 13:08 Self-Rated Confidence 23:45 RAG and Long Context LLMs 29:35 Search Relevance vs. Sufficient Context 35:00 Retrieval-Aware Fine-Tuning 44:05 Directions for the future of AIRAG Benchmarks with Nandan Thakur - Weaviate Podcast #124!Weaviate • Vector Database2025-06-25 | Nandan Thakur is a Ph.D. student at the University of Waterloo! Nandan has worked on many of the most impactful works in Retrieval-Augmented Generation (RAG) and Information Retrieval. His work ranges from benchmarks such as BEIR, MIRACLE, TREC, and FreshStack, to improving the training of embedding models and re-rankings, and more!
Chapters 0:00 Welcome Nandan! 1:15 The BEIR Benchmarks 8:25 Evolution of RAG Benchmarks 12:25 Diversity in Search Results 18:20 Reasoning and Query Writing 23:03 Search Result Summarization 26:10 Looping Searches 29:45 High Recall Search 33:10 Paginating Search Results 38:05 Searching with Filters 44:35 Mixture of Retrievers 51:08 Creating FreshStack Weaviate 1:00:35 Exciting Directions for AIRAG is the Backbone of Enterprise AIWeaviate • Vector Database2025-06-20 | “You don’t need RAG, context windows are huge now…” Sounds great, right?
But the reality is: 𝗯𝗶𝗴𝗴𝗲𝗿 𝗶𝘀𝗻’𝘁 𝗮𝗹𝘄𝗮𝘆𝘀 𝗯𝗲𝘁𝘁𝗲𝗿, especially if it’s slower, costlier, and oblivious to updates.
Join Box’s CTO Ben Kus, along with Bob van Lujit and Connor Shorten, in the latest Weaviate Podcast to learn more about enterprise-scale AI solutions, Box's three-layer infrastructure puzzle, embeddings, retrieval, and much more!
• LLM reasoning included: Custom instructions for specific use cases
Whether you're building a clothing brand or recipe platform (recipes below 🧑🍳), imagine search results that actually match what each user wants to see.
The coolest part? You can choose classic ML or go full LLM with agent ranking that even explains its reasoning.
- Visit [weaviate.io](http://weaviate.io/) - Star us on GitHub github.com/weaviate/weaviate - Stay updated and subscribe to our newsletter: newsletter.weaviate.io - Try out Weaviate Cloud Services for free here: https://console.weaviate.cloud/
- Twitter: twitter.com/weaviate_io - LinkedIn: linkedin.com/company/weaviate-ioMUVERA with Rajesh Jayaram and Roberto Esposito - Weaviate Podcast #123!Weaviate • Vector Database2025-05-28 | Multi-vector retrieval offers richer, more nuanced search, but often comes with a significant cost in storage and computational overhead. How can we harness the power of multi-vector representations without breaking the bank? Rajesh Jayaram, the first author of the groundbreaking MUVERA algorithm from Google, and Roberto Esposito from Weaviate, who spearheaded its implementation, reveal how MUVERA tackles this critical challenge.
Dive deep into MUVERA, a novel compression technique specifically designed for multi-vector retrieval. Rajesh and Roberto explain how it leverages contextualized token embeddings and innovative fixed dimensional encodings to dramatically reduce storage requirements while maintaining high retrieval accuracy. Discover the intricacies of quantization within MUVERA, the interpretability benefits of this approach, and how LSH clustering can play a role in topic modeling with these compressed representations.
This conversation explores the core mechanics of efficient multi-vector retrieval, the challenges of benchmarking these advanced systems, and the evolving landscape of vector database schemas designed to handle such complex data. Rajesh and Roberto also share their insights on the future directions in artificial intelligence where efficient, high-dimensional data representation is paramount. Whether you're an AI researcher grappling with the scalability of vector search, an engineer building advanced retrieval systems, or fascinated by the cutting edge of information retrieval and AI frameworks, this episode delivers unparalleled insights directly from the source. You'll gain a fundamental understanding of MUVERA, practical considerations for its application in making multi-vector retrieval feasible, and a clear view of future directions in AI.
Chapters 0:00 Welcome Rajesh and Roberto 2:10 Intro to Multi-Vector Retrieval 7:53 Contextualized Token Embeddings 14:04 Interpretability of Multi-Vector Retrieval 17:46 Multi-Vector Storage Cost 20:10 MUVERA Deep Dive 32:30 Fixed Dimensional Encodings 55:32 Quantization in MUVERA 1:00:04 LSH Clustering for Topic Modeling 1:02:44 Benchmarks for Multi-Vector Retrieval 1:06:33 Future of Vector Database Schemas 1:09:28 Directions for the Future of AIPatronus AI with Anand Kannappan - Weaviate Podcast #122!Weaviate • Vector Database2025-05-15 | AI agents are getting more complex and harder to debug. How do you know what's happening when your agent makes 20+ function calls? What if you have a Multi-Agent System orchestrating several Agents? Anand Kannappan, co-founder of Patronus AI, reveals how their groundbreaking tool Percival transforms agent debugging and evaluation. Percival can instantly analyze complex agent traces, it pinpoints failures across 60 different modes, and it automatically suggests prompt fixes to improve performance. Anand unpacks several of these common failure modes. This includes the critical challenges of "context explosion" where agents process millions of tokens. He also explains domain adaptation for specific use cases, and the complex challenge of multi-agent orchestration. The paradigm of AI Evals is shifting from static evaluation to dynamic oversight! Also learn how Percival's memory architecture leverages both episodic and semantic knowledge with Weaviate!
This conversation explores powerful concepts like process vs. outcome rewards and LLM-as-judge approaches. Anand shares his vision for "agentic supervision" where equally capable AI systems provide oversight for complex agent workflows. Whether you're building AI agents, evaluating LLM systems, or interested in how debugging autonomous systems will evolve, this episode delivers concrete techniques. You'll gain philosophical insights on evaluation and a roadmap for how evaluation must transform to keep pace with increasingly autonomous AI systems.
Chapters 0:00 Welcome Anand! 1:15 Percival! 17:20 Online and Offline Agent Tracing 20:40 Complex Agent Traces 23:05 Quick Insights and Deep Research 24:47 Automated Agent Tuning 31:19 LLM-as-Judge and Scalable Oversight 42:24 Agent Inbox for Evals 45:49 Causal Inference and AI 51:24 Percival and Weaviate 56:04 Exciting Directions for AIEmbedding model evaluation & selection guideWeaviate • Vector Database2025-05-14 | Selecting the right embedding model can make or break your AI application's performance. In this guide, JP from Weaviate walks you through a practical 4-step framework to navigate the complex landscape of embedding models.
Learn how embedding models transform content into numerical vectors that power search, recommendations, and more. Discover why your model choice impacts not just performance, but also resource requirements and operational costs.
0:00 The cookie recipe challenge 0:41 Introduction 1:05 Why embedding model selection matters 2:10 The 4-step selection framework 4:40 Practical tips for better decision-making 5:36 Recap & resources
🔑 KEY TAKEAWAYS:
- How to identify your specific needs across data characteristics, performance requirements, operational factors, and business constraints - Strategies for narrowing down hundreds of models to a manageable shortlist - Methods for rigorously evaluating models using both standard benchmarks and your own data - When and how to reassess your model selection as the field evolves
But what if we told you could simply tell your database what changes you want to make to your data, in plain language, and have it handle all the technical details for you.
That's exactly where Weaviate's new Transformation Agent comes in!
The Transformation Agent lets you modify and enhance your data using natural language instructions. It can:
- Add new properties to objects based on existing data - Update existing properties with enhanced information - Process multiple operations in parallel - Apply transformations across entire collections
This agent handles the tedious task of designing database updates and lets you focus on what information you need, not how to get it.
- Twitter: twitter.com/weaviate_io - LinkedIn: linkedin.com/company/weaviate-ioHaize Labs with Leonard Tang - Weaviate Podcast #121!Weaviate • Vector Database2025-05-12 | How do you ensure your AI systems actually do what you expect them to do? Leonard Tang takes us deep into the revolutionary world of AI evaluation with concrete techniques you can apply today. Learn how Haize Labs is transforming AI testing through "scaling judge-time compute" - stacking weaker models to effectively evaluate stronger ones. Leonard unpacks the game-changing Verdict library that outperforms frontier models by 10-20% while dramatically reducing costs. Discover practical insights on creating contrastive evaluation sets that extract maximum signal from human feedback, implementing debate-based judging systems, and building custom reward models that align with enterprise needs. The conversation reveals powerful nuggets like using randomized agent debates to achieve consensus and lightweight guardrail models that run alongside inference. Whether you're developing AI applications or simply fascinated by how we'll ensure increasingly powerful AI systems perform as expected, this episode delivers immediate value with techniques you can implement right away, philosophical perspectives on AI safety, and a glimpse into the future of evaluation that will fundamentally shape how AI evolves.
Chapters 0:00 Weaviate Podcast #121! 0:46 Welcome Leonard! 1:16 Founding Haize Labs 8:31 UX for Evals 17:01 Scaling Judge-Time Compute with Verdict 23:26 Debate Judges 26:06 Compute Scaling 28:50 Declarative Judge Pipelines 31:13 Custom Reward Models 37:20 Reasoning in Reward Models 39:20 Mechanistic Interpretability 45:30 Guardrails in Inference Pipelines 47:35 Can we control Superintelligence? 52:08 Exciting Directions for AIReimagine Data Workflows with Weaviate AgentsWeaviate • Vector Database2025-04-30 | Calling all AI devs, tech leads, and novices experimenting with agentic AI!
Join us for a hands-on walkthrough of Weaviate Agents — a powerful suite of services designed to simplify and automate your AI data workflows.
In this live session, we'll showcase how the 𝐐𝐮𝐞𝐫𝐲 𝐀𝐠𝐞𝐧𝐭, 𝐓𝐫𝐚𝐧𝐬𝐟𝐨𝐫𝐦𝐚𝐭𝐢𝐨𝐧 𝐀𝐠𝐞𝐧𝐭, 𝐚𝐧𝐝 𝐏𝐞𝐫𝐬𝐨𝐧𝐚𝐥𝐢𝐳𝐚𝐭𝐢𝐨𝐧 𝐀𝐠𝐞𝐧𝐭 work under the hood to enable natural language querying, real-time data transformation, and context-aware personalization without heavy lifting.
This session will include live demos, example use cases, and implementation tips.
[𝐈𝐦𝐩𝐨𝐫𝐭𝐚𝐧𝐭] 𝐈𝐟 𝐲𝐨𝐮 𝐥𝐢𝐤𝐞 𝐭𝐨 𝐫𝐞𝐜𝐞𝐢𝐯𝐞 𝐦𝐚𝐭𝐞𝐫𝐢𝐚𝐥𝐬 𝐮𝐬𝐞𝐝 𝐝𝐮𝐫𝐢𝐧𝐠 𝐭𝐡𝐞 𝐬𝐞𝐬𝐬𝐢𝐨𝐧 𝐚𝐧𝐝 𝐟𝐨𝐥𝐥𝐨𝐰-𝐮𝐩 𝐢𝐧𝐟𝐨𝐫𝐦𝐚𝐭𝐢𝐨𝐧 𝐦𝐚𝐤𝐞 𝐬𝐮𝐫𝐞 𝐭𝐨 𝐚𝐠𝐫𝐞𝐞 𝐭𝐨 𝐖𝐞𝐚𝐯𝐢𝐚𝐭𝐞𝐬 𝐏𝐫𝐢𝐯𝐚𝐜𝐲 𝐏𝐨𝐥𝐢𝐜𝐲Stateful Agents self-optimize their Context WindowWeaviate • Vector Database2025-04-25 | Have you tried asking ChatGPT what it knows about you? This clip from Sarah Wooders describes how Letta is pioneering a new wave of Stateful Agents that self-optimize their context windows based on your message history and external documents.
Since recording this podcast, Letta has continued to push the frontier of Stateful Agents with the introduction of "Sleep-time Compute"! I think it is well worth your time to take a look at their announcement resources below, and I also hope this clip inspires your interest in the podcast with Sarah!
- Visit [weaviate.io](http://weaviate.io/) - Star us on GitHub github.com/weaviate/weaviate - Stay updated and subscribe to our newsletter: newsletter.weaviate.io - Try out Weaviate Cloud Services for free here: https://console.weaviate.cloud/
Join Matt Biilmann, Sebastian Witalec, Charles Pierse, and Connor Shorten on the Weaviate Podcast to learn all about Agent Experience (AX) - and why it’s becoming so important alongside the user and developer experience.10. Use Cases of AI AgentsWeaviate • Vector Database2025-04-14 | The future of work isn’t human OR AI.
It’s human WITH AI Agents. Let’s have a look at the industries that AI is transforming right NOW:9. The Future of AI AgentsWeaviate • Vector Database2025-04-11 | How much power should we give AI Agents? Here’s what the future really looks like:AI Agents Explained: Making AI Actually WORK For YouWeaviate • Vector Database2025-04-10 | How can you make AI agents work for you? At Weaviate, we've developed three specialized agents to make working with your data easier:
1️⃣ Query Agent: Finds exactly what you need from your database when you ask in natural language 2️⃣ Transformation Agent: Lets you define database operations conversationally (e.g., "create a summary of this text") 3️⃣ Personalization Agent (coming soon): Gives smarter recommendations by incorporating user behavior and preferences
🔔 Stay TUNED for in-depth videos on each Weaviate Agent + Demos 🔔
TIMESTAMPS 00:00 Introduction: Beyond Basic Language Models 00:17 What Is an AI Agent? 00:35 How do AI Agents Actually Work? 01:27 Memory & Context Retention 01:45 How Can We Build AI Agents With Weaviate? 01:50 Query Agent 01:57 Transformation Agent 02:10 Personalization Agent 02:28 How to Try Weaviate Agents Today
- Visit [weaviate.io](http://weaviate.io/) - Star us on GitHub github.com/weaviate/weaviate - Stay updated and subscribe to our newsletter: newsletter.weaviate.io - Try out Weaviate Cloud Services for free here: https://console.weaviate.cloud/
Connect with us on - Twitter: twitter.com/weaviate_io - LinkedIn: linkedin.com/company/weaviate-ioWeaviate Tech Hands-On: Query AgentWeaviate • Vector Database2025-04-10 | In this session Tuana demonstrates in a hands-on technical walkthrough how to use the new Weaviate Query Agent. She uses natural language in order to query data spread across multiple collections and derive meaningfull insights from it while the Query Agent handles the search logic for her.
00:00:00 Intro 00:01:40 What is the Weaviate Query Agent 00:04:30 Search vs. Aggregation operations 00:05:17 setup eCommerce Dataset in Weaviate 00:10:40 Setup the Weaviate Query Agent 00:11:30 Run the Weaviate Query Agent 00:21:30 Run context-aware follow-up questions 00:25:59 Search across multiple collections 00:40:20 Change system prompts for the Weaviate Query Agent
▬▬▬▬▬▬▬▬▬▬▬▬ CONNECT WITH US ▬▬▬▬▬▬▬▬▬▬▬▬
- Visit [weaviate.io](http://weaviate.io/) - Star us on GitHub github.com/weaviate/weaviate - Stay updated and subscribe to our newsletter: newsletter.weaviate.io - Try out Weaviate Cloud Services for free here: https://console.weaviate.cloud/
Tuana Socials - Twitter/X: https://x.com/tuanacelik - LinkedIn: linkedin.com/in/tuanacelikWhat is Vibe Coding?Weaviate • Vector Database2025-04-10 | Vibe coding is shifting the development focus from specific implementation details and knowledge to overall architecture and wisdom. Domain knowledge, architectural thinking, and system design principles provided by humans will still be necessary, but just expressed through prompts and edits rather than direct implementation.
Check out Etienne’s experiment: https://x.com/etiennedi/status/1899843351030964535Structured Outputs with Will Kurt and Cameron Pfiffer - Weaviate Podcast #119!Weaviate • Vector Database2025-04-09 | Hey everyone! Thanks so much for watching another episode of the Weaviate Podcast! Dive into the fascinating world of structured outputs with Will Kurt and Cameron Pfeiffer, the brilliant minds behind Outlines, the revolutionary open-source library from .txt.ai that's changing how we interact with LLMs. In this episode, we explore how constrained decoding enables predictable, reliable outputs from language models—unlocking everything from perfect JSON generation to guided reasoning processes.
Will and Cameron share their journey to founding .txt.ai, explain the technical magic behind Outlines (hint: it involves finite state machines!), and debunk misconceptions around structured generation performance. You'll discover practical applications like knowledge graph construction, metadata extraction, and report generation that simply weren't possible before this technology.
Whether you're building AI systems or curious about where the field is heading, you'll gain valuable insights on how structured outputs integrate with inference engines like vLLM, why multi-task inference outperforms single-task approaches, and how this technology enables scalable agent systems that could transform software architecture forever. Join us for this mind-expanding conversation about one of AI's most important but underappreciated innovations—and discover why the future might belong to systems that combine freedom with structure.
Chapters 0:00 Welcome Will and Cameron! 1:42 What lead you to dottxt? 7:20 Structured Outputs for Beginners 9:00 Metadata Extraction 17:36 Structured Reasoning 23:00 Report Generation 28:25 Multi-Task Inference 30:55 How does Outlines work? 35:55 Integration with vLLM and Inference Engines 43:55 Let Me Speak Freely: A Rebuttal 56:35 Distribution Alignment 1:04:10 Exciting Directions for AI8. AI Agents: Tools & APIsWeaviate • Vector Database2025-04-09 | An AI agent's power lies in the tools it can use!
Tools and APIs enable agents to:What is MCP (Model Context Protocol)?Weaviate • Vector Database2025-04-08 | 𝗠𝗼𝗱𝗲𝗹 𝗖𝗼𝗻𝘁𝗲𝘅𝘁 𝗣𝗿𝗼𝘁𝗼𝗰𝗼𝗹 (𝗠𝗖𝗣) is turning into the universal connector for AI systems.
Think of MCP as the USB-C cable of AI integrations. Instead of building separate connections for each data source or tool, you create one MCP server that exposes specific capabilities (like searching a database or querying weather data).
Topics: 00:00 Intro 00:22 What is the Weaviate Transformation Agent 05:00 Setup a test dataset inside Weaviate 15:10 Challenges when working with data 19:00 Using agents to add properties 26:50 Using agents to perform multiple operations simultaniously 33:30 Using agents to update existing properties 41:35 Bonus: Using the Weaviate Query Agent 44:30 Bonus: Follow Up Queries with the Query Agent 47:00 Bonus: Current limitations 51:20 Additional Resources
Learn how to significantly reduce the memory footprint of large language models and embedding models while preserving their functionality - even on constrained devices like Raspberry Pi 5!
🔑 Key Topics Covered: - LLM quantization techniques (from FP16/FP8 to 4-bit precision) - The GGUF format and LLAMA.cpp framework - Why feed-forward layer parameters are more sensitive than attention layers - Embedding model quantization using ONNX - Vector database quantization methods (Product, Binary, and Scalar) - Running vector databases and AI models on edge devices
This technical deep dive is perfect for developers looking to optimize AI models for memory-constrained environments or deploy vector search capabilities on edge devices.
Learn more from Weaviate at https://weaviate.io.Build a No-Code Agentic Workflow in Under 5 MinutesWeaviate • Vector Database2025-03-25 | Developers often spend weeks building agentic workflows.
Femke just built one in 5 minutes, without writing a single line of code.
• Route general questions to Google search • Direct document-specific queries to your knowledge base
Using Stack AI's no-code platform, I built a workflow that automatically chooses between these routes based on the query type.
From searching the web to processing 200+ page documents, building secure, robust, customizable AI applications and pipelines on your own data is way easier than most companies think it is.
And of course, powered by Weaviate's vector database for efficient document processing and search 💚
Learn more about Weaviate: weaviate.io Explore StackAI: StackAI: stack-ai.comSynthetic Data with David Berenstein and Ben Burtenshaw - Weaviate Podcast #118!Weaviate • Vector Database2025-03-25 | Synthetic Data: The Building Bocks of AI's Future!
Hey everyone! I am SUPER EXCITED to publish the 118th episode of the Weaviate Podcast featuring David Berenstein and Ben Burtenshaw from HuggingFace! This podcast explores the intricacies of synthetic data generation, detailing methodologies such as data augmentation, distillation, and instruction refinement. The conversation delves into persona-driven synthetic data, highlighting applications like Persona Hub, and discusses algorithms to enhance diversity, complexity, and quality of generated data. Additionally, they cover integration with Hugging Face’s ecosystem, including Argilla for annotation, AutoTrain for fine-tuning, and advanced data exploration tools like the Data Studio and SQL console. The podcast also touches upon the potential for synthetic image data generation and the exciting future of AI education and accessibility.
0:00 Welcome David and Ben! 0:56 The Synthetic Data Generator 3:25 Data Generation Algorithms 11:32 PersonaHub 15:15 Exploration in Data Generation 22:05 Systems for Data Generation 27:20 DSPy and Distilabel 33:30 Diversity in Data Generation 39:10 Synthetic Image Generation 44:50 Evolutionary Algorithms for Synthetic Data 50:48 Prompt to Model 59:25 Exciting Directions for AIText-to-SQL is dead: The next generation of querying is AgenticWeaviate • Vector Database2025-03-12 | This paper (arxiv.org/abs/2502.00032) introduces a type of agentic querying called Function Calling that uses an LLM to structure queries using predefined function calls in JSON format, with optional arguments for search, filters, aggregation, and grouping.
Along with that, it also tests out a bunch of different models with a new dataset, DBGorilla, designed to evaluate agentic querying techniques on real-world use cases.
Weaviate also just released a Query Agent, designed based on some of the work in this paper, to handle advanced agentic querying out of the box, find out more here: weaviate.io/blog/query-agent
We dive into features like: • Multi-valued Vectors aka multi-vector embeddings based on colBERT, enabling you to increase the accuracy of your search results • A new B25 implementation based on BlockMax WAND, speeding up search performance, • General Availability of role-based access control (RBAC) • A new NVIDIA model integration • and much more!
Connect with us on - Twitter: twitter.com/weaviate_io - LinkedIn: linkedin.com/company/weaviate-ioLetta AI with Sarah Wooders - Weaviate Podcast #117!Weaviate • Vector Database2025-03-03 | Hey everyone! Thank you so much for watching the 117th episode of the Weaviate podcast! In this episode, we dive deep into the cutting edge of AI agent development with Sarah Wooders, co-founder and CTO of Letta AI. Emerging from Berkeley's Sky Computing Lab, Sarah and her team have pioneered a revolutionary approach to stateful agents - AI systems that genuinely remember both you and themselves across extended conversations. The conversation explores how the groundbreaking MemGPT project evolved into Letta's comprehensive Agent Development Environment (ADE), which empowers developers to build truly persistent AI experiences. Sarah shares powerful insights on context management, memory prioritization, and the critical role of databases in agent architecture. Whether you're building AI systems or simply curious about where conversational AI is heading, this episode illuminates how the future of agents depends not just on their reasoning capabilities, but on their ability to maintain coherent identity and memory over time.
LLMs Can Easily Learn to Reason from Demonstrations Structure, not content, is what matters! - arxiv.org/pdf/2502.07374
0:00 Welcome Sarah! 4:03 Stateful Agents 8:28 Context Window Optimization 9:13 Self-Optimizing Agents 12:43 Agent Personas 14:48 Gradient Descent and Prompts 19:58 Tools and Function Calling 21:53 Agent Development Environment (ADE) 31:54 Databases for Agents 36:28 LLM-Tool Use Examples 37:03 Multi-Agent Systems 41:28 Model Routers 43:26 Reasoning Models and Context Windows 47:53 Inspiration for ADE 50:52 Bring your own Tools 57:00 Future Directions for AIAgent Experience with Matt Biilmann, Sebastian Witalec, and Charles Pierse - Weaviate Podcast #116!Weaviate • Vector Database2025-02-27 | Hey everyone! Thank you so much for watching another episode of the Weaviate Podcast! I am SUPER excited to welcome Matt Biilmann, Co-Founder and CEO of Netlify, as well as Sebastian Witalec and Charles Pierse from Weaviate to discuss Agent Experience! You have probably heard about how you can connect LLMs to external software tools. This supercharges the capabilities of AI systems and what they can do. So what does that mean for you as a software developer?
This podcast explores different ideas around designing software user experiences for Agents as well as Humans. How do we write documentation for Agents differently than Humans? How do we design REST or gRPC APIs, or programming languages clients, for Agents differently than Humans? llms.txt, JSON tool definitions, agents to agent, breaking changes, … there were so many interesting topics explored in this podcast! I really hope you find it useful! As always more than happy to discuss these ideas further with you!
Read the blog here! - https://biilmann.blog/articles/introducing-ax/
Chapters 0:00 Weaviate Podcast #116! 0:20 Welcome Matt, Sebastian, Charles 1:45 Agent Experience 10:20 Docs for Agents 13:12 llms.txt 17:20 JSON Tool Definitions 21:16 Have your Agent talk to my Agent 26:27 gRPC APIs 29:56 Agentic Workflows 33:04 LLM Bias in Software Vendors 42:52 New APIs and Breaking ChangesHow to pick the best LLM in 2025Weaviate • Vector Database2025-02-20 | Stop relying on OpenAI models for everything!
Whether you need speed, quality, or performance, there's a perfect language model out there for your needs.
But with so many choices, how do you pick the best one?
Here are the 4 key factors to consider:
• 𝗤𝘂𝗮𝗹𝗶𝘁𝘆: Look for an LLM that performs consistently across tasks like chatbot interactions, language understanding, and coding. Best options: DeepSeek R1, OpenAI's o1, OpenAI's o3 Mini
• 𝗣𝗿𝗶𝗰𝗲: Consider the cost per token for both input and output. Best options: Gemini 2.0 Flash, GPT-4o Mini, and Mistral Small 3
• 𝗦𝗽𝗲𝗲𝗱/𝗧𝗵𝗿𝗼𝘂𝗴𝗵𝗽𝘂𝘁: Generally, smaller models like Mistral Small 3 and GPT-4o Mini are faster than larger, high-quality models. Best options: Gemini 2.0 Flash, and GPT-4o Mini
• 𝗢𝗽𝗲𝗻 𝗦𝗼𝘂𝗿𝗰𝗲: Models that allow for private use Best options: Llama 3, Microsoft Phi-4, Deepseek's latest models
• 𝗖𝗼𝗱𝗶𝗻𝗴: Models that excel at code generation Best options: Claude 3.5 Sonnet, OpenAI's o3 Mini, Deepseek R1 & V3
Try it in @weaviate.io: weaviate.io/developers/weaviate/search/generativeOptimizing Retrieval Agents with Shirley Wu - Weaviate Podcast #115!Weaviate • Vector Database2025-02-19 | Hey everyone! Thank you so much for watching the 115th episode of the Weaviate Podcast featuring Shirley Wu from Stanford University! We explore the innovative Avatar Optimizer—a novel framework that leverages contrastive reasoning to refine LLM agent prompts for optimal tool usage. Shirley explains how this self-improving system evolves through iterative feedback by contrasting positive and negative examples, enabling agents to handle complex tasks more effectively.
We also dive into the STaRK Benchmark, a comprehensive testbed designed to evaluate retrieval systems on semi-structured knowledge bases. The discussion highlights the challenges of unifying textual and relational retrieval, exploring concepts such as multi-vector embeddings, relational graphs, and dynamic data modeling. Learn how these approaches help overcome information loss, enhance precision, and enable scalable, context-aware retrieval in diverse domains—from product recommendations to precision medicine.
Whether you’re interested in advanced prompt optimization, multi-agent system design, or the future of human-centered language models, this episode offers a wealth of insights and a forward-looking perspective on integrating sophisticated AI techniques into real-world applications.
Chapters 0:00 Welcome Shirely! 0:30 What is the state of AI? 2:18 Graph Data Models and GraphRAG 8:22 Unifying Textual and Relational Retrieval 10:50 Multi-Vector Retrieval 15:55 Embeddings and Properties 21:58 PrimeKG 27:25 AvaTaR: LLM Tool Use Optimization 41:40 Complex Retrieval 46:28 Agent Memory and DSPy 55:00 Future Directions for AIContextual AI with Amanpreet Singh - Weaviate Podcast #114!Weaviate • Vector Database2025-02-12 | Hey everyone! Thank you so much for watching the 114th episode of the Weaviate Podcast featuring Amanpreet Singh, Co-Founder and CTO of Contextual AI! Contextual AI is at the forefront of production-grade RAG agents! I learned so much from this conversation! We began by discussing the vision of RAG 2.0, jointly optimizing generative and retrieval models! This then lead us to discuss Agentic RAG and how the RAG 2.0 roadmap is evolving with emerging perspectives on tool use. Amanpreet continues to further motivate the importance of continual learning of the model and the prompt / few-shot examples -- discussing the limits of prompt engineering. Personally I have to admit I think I have been a bit too bullish on only tuning instructions / examples, Amanpreet made an excellent case for updating the weights of the models as well -- citing issues such as parametric knowledge conflicts, and later on discussing how Mechanistic Interpretability is used to audit models and their updates in enterprise settings. We then discussed Contextual AI's LMUnit for evaluating these systems. This then lead us into my favorite part of the podcast, a deep dive into RL algorithms for LLMs. I highly recommend checking out the links below to learn more about Contextual's innovations on APO and KTO! We then discuss the importance of domain specific data, Mechanistic Interpretability, return to another question on RAG 2.0, and conclude with Amanpreet's most exciting future directions for AI! I hope you enjoy the podcast!
Connect with us on - Twitter: twitter.com/weaviate_io - LinkedIn: linkedin.com/company/weaviate-ioCartesia AI with Karan Goel - Weaviate Podcast #113!Weaviate • Vector Database2025-01-28 | Hey everyone! Thank you so much for watching the 113th episode of the Weaviate Podcast with Karan Goel from Cartesia AI! Cartesia AI is leading the AI world in text-to-speech models! As exciting as these new applications in speech generation are, Cartesia is also building around an incredibly exciting new neural network architecture that cuts across all of AI -- State Space Models. State Space Models (SSMs) present a new approach to modeling long sequences circumventing the quadratic attention bottlenecks of transformers. In the podcast, we discuss Karan's perspectives around end-to-end modeling, long context and Multimodal processing, building and deploying a new kind of model, and more! I hope you find the podcast interesting and useful! As always more than happy to discuss these ideas further with you! Thank you for listening!
Chapters: 0:00 Welcome Karan! 0:37 Founding Cartesia AI 2:33 State Space Models (SSMs) 9:04 Audio Data Deep Dive 14:45 The launch of Sonic 23:30 Inference vs. Agent APIs 29:58 Learning, Search, and Reasoning 34:35 RETRO and Fusion-in-Decoder RAG 38:00 State in Sequence Models 42:00 CUDA for SSMs 49:00 Many Shot In-Context Learning 50:35 AI-Native Text-to-Speech ApplicationsHow to choose an embedding modelWeaviate • Vector Database2025-01-28 | How do you chose the best embedding model for your use case? (and how do they even work, anyways?)