Weaviate vector database
Charles Pierse on Tactic Generate - Weaviate Podcast #69!
updated
00:00:00 Intro
00:02:14 Environment Setup
00:12:35 Intro to Personalization Agent
00:17:38 Hands-On part: How to use the Personalization Agent
00:50:07 Summary and further resources
▬▬▬▬▬▬▬▬▬▬▬▬ CONNECT WITH US ▬▬▬▬▬▬▬▬▬▬▬▬
- Visit [weaviate.io](http://weaviate.io/)
- Star us on GitHub github.com/weaviate/weaviate
- Stay updated and subscribe to our newsletter: newsletter.weaviate.io
- Try out Weaviate Cloud Services for free here: https://console.weaviate.cloud/
Got a question?
- Forum: forum.weaviate.io
- Slack: weaviate.io/slack
Connect with us on
- Twitter: twitter.com/weaviate_io
- LinkedIn: linkedin.com/company/weaviate-io
Links:
Hands-On Large Language Models: oreilly.com/library/view/hands-on-large-language/9781098150952
BERTopic: github.com/MaartenGr/BERTopic
BERTopic (paper): arxiv.org/abs/2203.05794
Learn more about Maarten Grootendorst: maartengrootendorst.com
TopicGPT: arxiv.org/abs/2311.01449
TnT-LLM: arxiv.org/abs/2403.12173
Chapters
0:00 Welcome Maarten!
1:57 Hands-On Large Language Models
7:34 An Overview of Topic Modeling
10:45 LLM Topic Generation
17:13 Topic Modeling with Human Feedback
21:33 Topic Granularity
26:00 Visualizing Topics
31:18 Contrastive Topics
33:24 LLM-as-Judge for Topics
39:06 Separating Generation from Assignment
44:20 Applications of Topic Modeling
55:14 Semi-Supervised BERTopic
1:01:06 Future Directions for AI
00:00:00 Intro & Overview
00:01:15 What is the Weaviate Query Agent
00:07:37 Live Coding with Weaviate Query Agent
00:21:00 Further Resources
▬▬▬▬▬▬▬▬▬▬▬▬ CONNECT WITH US ▬▬▬▬▬▬▬▬▬▬▬▬
- Visit http://weaviate.io
- Star us on GitHub github.com/weaviate/weaviate
- Stay updated and subscribe to our newsletter: newsletter.weaviate.io
- Try out Weaviate Cloud for free here: https://console.weaviate.cloud/
Got a question?
- Forum: forum.weaviate.io
- Slack: weaviate.io/slack
Connect with us on
- Twitter: twitter.com/weaviate_io
- LinkedIn: linkedin.com/company/weaviate-io
Links:
Sufficient Context: A New Lens on Retrieval Augmented Generation Systems - arxiv.org/pdf/2411.06037
Chapters
0:00 Welcome Hailey
1:40 Definition of Sufficient Context
4:47 Context Engineering
8:42 AutoRater
13:08 Self-Rated Confidence
23:45 RAG and Long Context LLMs
29:35 Search Relevance vs. Sufficient Context
35:00 Retrieval-Aware Fine-Tuning
44:05 Directions for the future of AI
Links:
Nandan Thakur: thakur-nandan.github.io
FreshStack Benchmarks: fresh-stack.github.io
BEIR Benchmarks: arxiv.org/abs/2104.08663
BrowseComp: arxiv.org/abs/2504.12516
Sufficient Context: arxiv.org/abs/2411.06037
Check out the Weaviate Query Agent! weaviate.io/developers/agents/query
Chapters
0:00 Welcome Nandan!
1:15 The BEIR Benchmarks
8:25 Evolution of RAG Benchmarks
12:25 Diversity in Search Results
18:20 Reasoning and Query Writing
23:03 Search Result Summarization
26:10 Looping Searches
29:45 High Recall Search
33:10 Paginating Search Results
38:05 Searching with Filters
44:35 Mixture of Retrievers
51:08 Creating FreshStack Weaviate
1:00:35 Exciting Directions for AI
Sounds great, right?
But the reality is: 𝗯𝗶𝗴𝗴𝗲𝗿 𝗶𝘀𝗻’𝘁 𝗮𝗹𝘄𝗮𝘆𝘀 𝗯𝗲𝘁𝘁𝗲𝗿, especially if it’s slower, costlier, and oblivious to updates.
Join Box’s CTO Ben Kus, along with Bob van Lujit and Connor Shorten, in the latest Weaviate Podcast to learn more about enterprise-scale AI solutions, Box's three-layer infrastructure puzzle, embeddings, retrieval, and much more!
➡️ youtube.com/watch?v=pPvSur8iEXY
00:00:00 Intro & Overview
00:03:00 MUVERA
00:15:50 Flexible Vectorizer Changes
00:21:00 Shard Movements - comming in 1.32
00:27:00 Weaviate Cloud - Import Data Tool
00:44:30 Shard Movement - Replication Factor
00:50:30 HNSW Snapshots
▬▬▬▬▬▬▬▬▬▬▬▬ CONNECT WITH US ▬▬▬▬▬▬▬▬▬▬▬▬
- Visit: http://weaviate.io
- Star us on GitHub github.com/weaviate/weaviate
- Stay updated and subscribe to our newsletter: newsletter.weaviate.io
- Try out Weaviate Cloud Services for free here: https://console.weaviate.cloud/
Got a question?
- Forum: forum.weaviate.io
- Slack: weaviate.io/slack
Connect with us on
- Twitter: twitter.com/weaviate_io
- LinkedIn: linkedin.com/company/weaviate-io
This isn't just another ranking tool. It's the first step toward truly personalized search that understands:
• User profiles (preferences, interests, characteristics)
• Real-time interactions (likes, dislikes, engagement patterns)
• LLM reasoning included: Custom instructions for specific use cases
Whether you're building a clothing brand or recipe platform (recipes below 🧑🍳), imagine search results that actually match what each user wants to see.
The coolest part? You can choose classic ML or go full LLM with agent ranking that even explains its reasoning.
Blog + Recipes: weaviate.io/blog/personalization-agent?utm_source=channels&utm_medium=w_social&utm_campaign=agents&utm_content=video_post_268065893
Tuana Socials
- Twitter/X: https://x.com/tuanacelik
- LinkedIn: linkedin.com/in/tuanacelik
Femke Socials
- Twitter/X: https://x.com/femke_plantinga
- LinkedIn: linkedin.com/in/femke-plantinga
- TikTok: tiktok.com/@femkeplant.ai
- Instagram: instagram.com/femkeplant.ai
▬▬▬▬▬▬▬▬▬▬▬▬ CONNECT WITH US ▬▬▬▬▬▬▬▬▬▬▬▬
- Visit [weaviate.io](http://weaviate.io/)
- Star us on GitHub github.com/weaviate/weaviate
- Stay updated and subscribe to our newsletter: newsletter.weaviate.io
- Try out Weaviate Cloud Services for free here: https://console.weaviate.cloud/
Got a question?
- Forum: forum.weaviate.io
- Slack: weaviate.io/slack
Connect with us on
- Twitter: twitter.com/weaviate_io
- LinkedIn: linkedin.com/company/weaviate-io
Dive deep into MUVERA, a novel compression technique specifically designed for multi-vector retrieval. Rajesh and Roberto explain how it leverages contextualized token embeddings and innovative fixed dimensional encodings to dramatically reduce storage requirements while maintaining high retrieval accuracy. Discover the intricacies of quantization within MUVERA, the interpretability benefits of this approach, and how LSH clustering can play a role in topic modeling with these compressed representations.
This conversation explores the core mechanics of efficient multi-vector retrieval, the challenges of benchmarking these advanced systems, and the evolving landscape of vector database schemas designed to handle such complex data. Rajesh and Roberto also share their insights on the future directions in artificial intelligence where efficient, high-dimensional data representation is paramount.
Whether you're an AI researcher grappling with the scalability of vector search, an engineer building advanced retrieval systems, or fascinated by the cutting edge of information retrieval and AI frameworks, this episode delivers unparalleled insights directly from the source. You'll gain a fundamental understanding of MUVERA, practical considerations for its application in making multi-vector retrieval feasible, and a clear view of future directions in AI.
Links:
MUVERA: arxiv.org/abs/2405.19504
CRISP: arxiv.org/pdf/2505.11471
ColBERT: arxiv.org/abs/2004.12832
ColPali: arxiv.org/abs/2407.01449
Multi-Vector Embeddings with Weaviate (Tutorial): weaviate.io/developers/weaviate/tutorials/multi-vector-embeddings
Chapters
0:00 Welcome Rajesh and Roberto
2:10 Intro to Multi-Vector Retrieval
7:53 Contextualized Token Embeddings
14:04 Interpretability of Multi-Vector Retrieval
17:46 Multi-Vector Storage Cost
20:10 MUVERA Deep Dive
32:30 Fixed Dimensional Encodings
55:32 Quantization in MUVERA
1:00:04 LSH Clustering for Topic Modeling
1:02:44 Benchmarks for Multi-Vector Retrieval
1:06:33 Future of Vector Database Schemas
1:09:28 Directions for the Future of AI
This conversation explores powerful concepts like process vs. outcome rewards and LLM-as-judge approaches. Anand shares his vision for "agentic supervision" where equally capable AI systems provide oversight for complex agent workflows. Whether you're building AI agents, evaluating LLM systems, or interested in how debugging autonomous systems will evolve, this episode delivers concrete techniques. You'll gain philosophical insights on evaluation and a roadmap for how evaluation must transform to keep pace with increasingly autonomous AI systems.
Links:
Percival Launch: patronus.ai/percival
Docs: docs.patronus.ai/docs/percival
Paper: arxiv.org/abs/2505.08638
Chapters
0:00 Welcome Anand!
1:15 Percival!
17:20 Online and Offline Agent Tracing
20:40 Complex Agent Traces
23:05 Quick Insights and Deep Research
24:47 Automated Agent Tuning
31:19 LLM-as-Judge and Scalable Oversight
42:24 Agent Inbox for Evals
45:49 Causal Inference and AI
51:24 Percival and Weaviate
56:04 Exciting Directions for AI
Learn how embedding models transform content into numerical vectors that power search, recommendations, and more. Discover why your model choice impacts not just performance, but also resource requirements and operational costs.
This video is a summary of our in-depth guide on Weaviate Academy. You can find the guide here: weaviate.io/developers/academy/theory/embedding_model_selection?utm_source=youtube&utm_medium=w_social&utm_campaign=dev_education&utm_content=video_post_680568036
⏱️ TIMESTAMPS:
0:00 The cookie recipe challenge
0:41 Introduction
1:05 Why embedding model selection matters
2:10 The 4-step selection framework
4:40 Practical tips for better decision-making
5:36 Recap & resources
🔑 KEY TAKEAWAYS:
- How to identify your specific needs across data characteristics, performance requirements, operational factors, and business constraints
- Strategies for narrowing down hundreds of models to a manageable shortlist
- Methods for rigorously evaluating models using both standard benchmarks and your own data
- When and how to reassess your model selection as the field evolves
👨💻 RESOURCES:
In-depth guidance, code examples, and evaluation metrics, visit our comprehensive guide at Weaviate Academy: weaviate.io/developers/academy/theory/embedding_model_selection
Connect with JP on LinkedIn: linkedin.com/in/jphwang
MTEB leaderboard: huggingface.co/spaces/mteb/leaderboard
▬▬▬▬▬▬▬▬▬▬▬▬ CONNECT WITH US ▬▬▬▬▬▬▬▬▬▬▬▬
- Visit http://weaviate.io
- Star us on GitHub: github.com/weaviate/weaviate
- Stay updated and subscribe to our newsletter: newsletter.weaviate.io
- Try out Weaviate Cloud Services for free here: https://console.weaviate.cloud/
Got a questions?
- Forum: forum.weaviate.io
- Slack: weaviate.io/slack
Connect with us on
- Twitter: twitter.com/weaviate_io
- LinkedIn: linkedin.com/company/weaviate-io
But what if we told you could simply tell your database what changes you want to make to your data, in plain language, and have it handle all the technical details for you.
That's exactly where Weaviate's new Transformation Agent comes in!
The Transformation Agent lets you modify and enhance your data using natural language instructions. It can:
- Add new properties to objects based on existing data
- Update existing properties with enhanced information
- Process multiple operations in parallel
- Apply transformations across entire collections
This agent handles the tedious task of designing database updates and lets you focus on what information you need, not how to get it.
Blog + Recipe: weaviate.io/blog/transformation-agent
Tuana Socials
- Twitter/X: https://x.com/tuanacelik
- LinkedIn: linkedin.com/in/tuanacelik
Femke Socials
- Twitter/X: https://x.com/femke_plantinga
- LinkedIn: linkedin.com/in/femke-plantinga
- TikTok: tiktok.com/@femkeplant.ai
- Instagram: instagram.com/femkeplant.ai
▬▬▬▬▬▬▬▬▬▬▬▬ CONNECT WITH US ▬▬▬▬▬▬▬▬▬▬▬▬
- Visit: http://weaviate.io
- Star us on GitHub github.com/weaviate/weaviate
- Stay updated and subscribe to our newsletter: newsletter.weaviate.io
- Try out Weaviate Cloud Services for free here: https://console.weaviate.cloud/
Got a question?
- Forum: forum.weaviate.io
- Slack: weaviate.io/slack
Connect with us on
- Twitter: twitter.com/weaviate_io
- LinkedIn: linkedin.com/company/weaviate-io
Learn more about Haize Labs! - haizelabs.com
Check out Verdict on GitHub - github.com/haizelabs/verdict
Chapters
0:00 Weaviate Podcast #121!
0:46 Welcome Leonard!
1:16 Founding Haize Labs
8:31 UX for Evals
17:01 Scaling Judge-Time Compute with Verdict
23:26 Debate Judges
26:06 Compute Scaling
28:50 Declarative Judge Pipelines
31:13 Custom Reward Models
37:20 Reasoning in Reward Models
39:20 Mechanistic Interpretability
45:30 Guardrails in Inference Pipelines
47:35 Can we control Superintelligence?
52:08 Exciting Directions for AI
The podcast then dives further into how vector embeddings can balloon file sizes - a few hundred bytes of text can require 4-6KB of vector data storage! We also dig into why RAG remains essential despite growing context windows, and how Box is developing AI agents that transform painful enterprise processes like RFP responses.
A fun, insightful look at what happens when cutting-edge AI meets enterprise-grade security and scale requirements!
Chapters
0:00 Weaviate Podcast #120!
0:20 Welcome Ben!
1:10 Exabyte Scale at Box
8:10 Infrastructure Rent vs. Buy
12:30 Founder-Led AI Adoption
19:37 Embeddings, RAG, and The AI Tipping Point
28:05 How many Vectors per Data Object?
37:20 Storage Tiers and Dynamic Indices in Retrieval
45:00 Agents
51:15 Agents for RFPs
Join us for a hands-on walkthrough of Weaviate Agents — a powerful suite of services designed to simplify and automate your AI data workflows.
In this live session, we'll showcase how the 𝐐𝐮𝐞𝐫𝐲 𝐀𝐠𝐞𝐧𝐭, 𝐓𝐫𝐚𝐧𝐬𝐟𝐨𝐫𝐦𝐚𝐭𝐢𝐨𝐧 𝐀𝐠𝐞𝐧𝐭, 𝐚𝐧𝐝 𝐏𝐞𝐫𝐬𝐨𝐧𝐚𝐥𝐢𝐳𝐚𝐭𝐢𝐨𝐧 𝐀𝐠𝐞𝐧𝐭 work under the hood to enable natural language querying, real-time data transformation, and context-aware personalization without heavy lifting.
This session will include live demos, example use cases, and implementation tips.
[𝐈𝐦𝐩𝐨𝐫𝐭𝐚𝐧𝐭] 𝐈𝐟 𝐲𝐨𝐮 𝐥𝐢𝐤𝐞 𝐭𝐨 𝐫𝐞𝐜𝐞𝐢𝐯𝐞 𝐦𝐚𝐭𝐞𝐫𝐢𝐚𝐥𝐬 𝐮𝐬𝐞𝐝 𝐝𝐮𝐫𝐢𝐧𝐠 𝐭𝐡𝐞 𝐬𝐞𝐬𝐬𝐢𝐨𝐧 𝐚𝐧𝐝 𝐟𝐨𝐥𝐥𝐨𝐰-𝐮𝐩 𝐢𝐧𝐟𝐨𝐫𝐦𝐚𝐭𝐢𝐨𝐧 𝐦𝐚𝐤𝐞 𝐬𝐮𝐫𝐞 𝐭𝐨 𝐚𝐠𝐫𝐞𝐞 𝐭𝐨 𝐖𝐞𝐚𝐯𝐢𝐚𝐭𝐞𝐬 𝐏𝐫𝐢𝐯𝐚𝐜𝐲 𝐏𝐨𝐥𝐢𝐜𝐲
Since recording this podcast, Letta has continued to push the frontier of Stateful Agents with the introduction of "Sleep-time Compute"! I think it is well worth your time to take a look at their announcement resources below, and I also hope this clip inspires your interest in the podcast with Sarah!
letta.com/blog/sleep-time-compute
(But guess what, you don’t have to anymore!)
We’re introducing the Query Agent - an AI assistant that seamlessly searches across your Weaviate collections!
Using the power of LLMs, Query Agent automatically:
- Decides which collections to search
- Combines information from multiple sources
- Provides comprehensive answers to complex questions
✨ Available now for all Weaviate Cloud & free sandbox users!
Try it today: weaviate.io/developers/wcs?utm_source=channels&utm_medium=w_social&utm_campaign=agents&utm_content=video_post_680409745
00:00 Introduction: The Query Agent
00:25 Example Clothing Brand (Femke)
01:35 Query Agent Demo (Tuana)
03:00 How to Try The Query Agent
Tuana Socials
- Twitter/X: https://x.com/tuanacelik
- LinkedIn: linkedin.com/in/tuanacelik
Femke Socials
- Twitter/X: https://x.com/femke_plantinga
- LinkedIn: linkedin.com/in/femke-plantinga
- TikTok: tiktok.com/@femkeplant.ai
- Instagram: instagram.com/femkeplant.ai
▬▬▬▬▬▬▬▬▬▬▬▬ CONNECT WITH US ▬▬▬▬▬▬▬▬▬▬▬▬
- Visit [weaviate.io](http://weaviate.io/)
- Star us on GitHub github.com/weaviate/weaviate
- Stay updated and subscribe to our newsletter: newsletter.weaviate.io
- Try out Weaviate Cloud Services for free here: https://console.weaviate.cloud/
Got a question?
- Forum: forum.weaviate.io
- Slack: weaviate.io/slack
Connect with us on
- Twitter: twitter.com/weaviate_io
- LinkedIn: linkedin.com/company/weaviate-io
Too human-friendly = agents break.
So… who are you really designing for?
Join Matt Biilmann, Sebastian Witalec, Charles Pierse, and Connor Shorten on the Weaviate Podcast to learn all about Agent Experience (AX) - and why it’s becoming so important alongside the user and developer experience.
It’s human WITH AI Agents. Let’s have a look at the industries that AI is transforming right NOW:
1️⃣ Query Agent: Finds exactly what you need from your database when you ask in natural language
2️⃣ Transformation Agent: Lets you define database operations conversationally (e.g., "create a summary of this text")
3️⃣ Personalization Agent (coming soon): Gives smarter recommendations by incorporating user behavior and preferences
🔔 Stay TUNED for in-depth videos on each Weaviate Agent + Demos 🔔
TIMESTAMPS
00:00 Introduction: Beyond Basic Language Models
00:17 What Is an AI Agent?
00:35 How do AI Agents Actually Work?
01:27 Memory & Context Retention
01:45 How Can We Build AI Agents With Weaviate?
01:50 Query Agent
01:57 Transformation Agent
02:10 Personalization Agent
02:28 How to Try Weaviate Agents Today
Tuana Socials
- Twitter/X: https://x.com/tuanacelik
- LinkedIn: linkedin.com/in/tuanacelik
Femke Socials
- Twitter/X: https://x.com/femke_plantinga
- LinkedIn: linkedin.com/in/femke-plantinga
- TikTok: tiktok.com/@femkeplant.ai
- Instagram: instagram.com/femkeplant.ai
▬▬▬▬▬▬▬▬▬▬▬▬ CONNECT WITH US ▬▬▬▬▬▬▬▬▬▬▬▬
- Visit [weaviate.io](http://weaviate.io/)
- Star us on GitHub github.com/weaviate/weaviate
- Stay updated and subscribe to our newsletter: newsletter.weaviate.io
- Try out Weaviate Cloud Services for free here: https://console.weaviate.cloud/
Got a question?
- Forum: forum.weaviate.io
- Slack: weaviate.io/slack
Connect with us on
- Twitter: twitter.com/weaviate_io
- LinkedIn: linkedin.com/company/weaviate-io
00:00:00 Intro
00:01:30 RBAC Overview & Demo
00:36:10 Async Indexing
00:40:00 Conflict resolution for Multi-node clusters
00:43:18 Japanese Tokenizer
00:45:55 BlockMax WAND implementation for BM25
00:48:07 Voyage AI Multimodal Model
00:49:50 Weaviate embeddings
▬▬▬▬▬▬▬▬▬▬▬▬ CONNECT WITH US ▬▬▬▬▬▬▬▬▬▬▬▬
- Visit http://weaviate.io
- Star us on GitHub github.com/weaviate/weaviate
- Stay updated and subscribe to our newsletter: newsletter.weaviate.io
- Try out Weaviate Cloud for free here: https://console.weaviate.cloud/
Got a question?
- Forum: forum.weaviate.io
- Slack: weaviate.io/slack
Connect with us on
- Twitter: twitter.com/weaviate_io
- LinkedIn: linkedin.com/company/weaviate-io
00:00:00 Intro
00:01:40 What is the Weaviate Query Agent
00:04:30 Search vs. Aggregation operations
00:05:17 setup eCommerce Dataset in Weaviate
00:10:40 Setup the Weaviate Query Agent
00:11:30 Run the Weaviate Query Agent
00:21:30 Run context-aware follow-up questions
00:25:59 Search across multiple collections
00:40:20 Change system prompts for the Weaviate Query Agent
▬▬▬▬▬▬▬▬▬▬▬▬ CONNECT WITH US ▬▬▬▬▬▬▬▬▬▬▬▬
- Visit [weaviate.io](http://weaviate.io/)
- Star us on GitHub github.com/weaviate/weaviate
- Stay updated and subscribe to our newsletter: newsletter.weaviate.io
- Try out Weaviate Cloud Services for free here: https://console.weaviate.cloud/
Got a question?
- Forum: forum.weaviate.io
- Slack: weaviate.io/slack
Connect with us on
- Twitter: twitter.com/weaviate_io
- LinkedIn: linkedin.com/company/weaviate-io
Tuana Socials
- Twitter/X: https://x.com/tuanacelik
- LinkedIn: linkedin.com/in/tuanacelik
Check out Etienne’s experiment: https://x.com/etiennedi/status/1899843351030964535
Will and Cameron share their journey to founding .txt.ai, explain the technical magic behind Outlines (hint: it involves finite state machines!), and debunk misconceptions around structured generation performance. You'll discover practical applications like knowledge graph construction, metadata extraction, and report generation that simply weren't possible before this technology.
Whether you're building AI systems or curious about where the field is heading, you'll gain valuable insights on how structured outputs integrate with inference engines like vLLM, why multi-task inference outperforms single-task approaches, and how this technology enables scalable agent systems that could transform software architecture forever. Join us for this mind-expanding conversation about one of AI's most important but underappreciated innovations—and discover why the future might belong to systems that combine freedom with structure.
Links:
dottxt AI: dottxt.co
Making LLMs Reliable: Building an LLM-powered Web App to Generate Gift Ideas: blog.dottxt.co/gifter.html
Say What You Mean: A Response to 'Let Me Speak Freely': blog.dottxt.co/say-what-you-mean.html
Coalescence: making LLM inference 5x faster: blog.dottxt.co/coalescence.html
StructuredRAG: JSON Response Formatting with Large Language Models: arxiv.org/abs/2408.11061
Chapters
0:00 Welcome Will and Cameron!
1:42 What lead you to dottxt?
7:20 Structured Outputs for Beginners
9:00 Metadata Extraction
17:36 Structured Reasoning
23:00 Report Generation
28:25 Multi-Task Inference
30:55 How does Outlines work?
35:55 Integration with vLLM and Inference Engines
43:55 Let Me Speak Freely: A Rebuttal
56:35 Distribution Alignment
1:04:10 Exciting Directions for AI
Tools and APIs enable agents to:
Think of MCP as the USB-C cable of AI integrations. Instead of building separate connections for each data source or tool, you create one MCP server that exposes specific capabilities (like searching a database or querying weather data).
Check out the Weaviate MCP: github.com/weaviate/mcp-server-weaviate
These 6 frameworks will save you hundreds of development hours.
More in this series: youtube.com/playlist?list=PLTL2JUbrY6tXYK3n_bIYZWr52l1qBITBE
Download the FREE Ebook: weaviate.io/ebooks/agentic-architectures?utm_source=youtube&utm_medium=w_social&utm_campaign=agents&utm_content=videoshort_post_680256845
Let’s break it down in 75 seconds!
Watch more videos in this series: youtube.com/playlist?list=PLTL2JUbrY6tXYK3n_bIYZWr52l1qBITBE
Download the FREE Ebook: weaviate.io/ebooks/agentic-architectures?utm_source=youtube&utm_medium=w_social&utm_campaign=agents&utm_content=videoshort_post_680318483
But for AI agents, they need to include these three essential features:
More about AI Agents: youtube.com/playlist?list=PLTL2JUbrY6tXYK3n_bIYZWr52l1qBITBE
Download the FREE Ebook: weaviate.io/ebooks/agentic-architectures?utm_source=youtube&utm_medium=w_social&utm_campaign=agents&utm_content=videoshort_post_680283937
These buzzwords get mixed up a lot, but they're actually quite different!
See the other videos of this series: youtube.com/playlist?list=PLTL2JUbrY6tXYK3n_bIYZWr52l1qBITBE
Download the FREE Ebook: weaviate.io/ebooks/agentic-architectures?utm_source=youtube&utm_medium=w_social&utm_campaign=agents&utm_content=videoshort_post_680531700
(longer than you might think!)
Full playlist AI Agents Explained: youtube.com/playlist?list=PLTL2JUbrY6tXYK3n_bIYZWr52l1qBITBE
Download the FREE Ebook: weaviate.io/ebooks/agentic-architectures?utm_source=youtube&utm_medium=w_social&utm_campaign=agents&utm_content=videoshort_post_680296508
AI agents rely on these key building blocks:
AI Agents Explained series: youtube.com/playlist?list=PLTL2JUbrY6tXYK3n_bIYZWr52l1qBITBE
Download the FREE Ebook: weaviate.io/ebooks/agentic-architectures?utm_source=youtube&utm_medium=w_social&utm_campaign=agents&utm_content=videoshort_post_680496173
(They have actually been around for quite some time)
In this 10-part series, Femke shares AI Agent fundamentals.
Full series: youtube.com/playlist?list=PLTL2JUbrY6tXYK3n_bIYZWr52l1qBITBE
Download the FREE Ebook: weaviate.io/ebooks/agentic-architectures?utm_source=youtube&utm_medium=w_social&utm_campaign=agents&utm_content=videoshort_post_680656377
This no-code AI solution takes 5 minutes to build.
Sometimes it pays to work smarter, not harder - watch Victoira build this company knowledge Q&A system with zero coding, using Stack AI's Platform.
Check it this webinar to learn more! events.weaviate.io/agents-stackai-webinar-2025?utm_source=linkedin&utm_medium=vs_social&utm_campaign=agents&utm_content=videoshort_post_680821688
Topics:
00:00 Intro
00:22 What is the Weaviate Transformation Agent
05:00 Setup a test dataset inside Weaviate
15:10 Challenges when working with data
19:00 Using agents to add properties
26:50 Using agents to perform multiple operations simultaniously
33:30 Using agents to update existing properties
41:35 Bonus: Using the Weaviate Query Agent
44:30 Bonus: Follow Up Queries with the Query Agent
47:00 Bonus: Current limitations
51:20 Additional Resources
▬▬▬▬▬▬▬▬▬▬▬▬ CONNECT WITH US ▬▬▬▬▬▬▬▬▬▬▬▬
- Visit http://weaviate.io
- Star us on GitHub github.com/weaviate/weaviate
- Stay updated and subscribe to our newsletter: newsletter.weaviate.io
- Try out Weaviate Cloud for free here: https://console.weaviate.cloud/
Got a question?
- Forum: forum.weaviate.io
- Slack: weaviate.io/slack
Connect with us on
- Twitter: twitter.com/weaviate_io
- LinkedIn: linkedin.com/company/weaviate-io
Learn how to significantly reduce the memory footprint of large language models and embedding models while preserving their functionality - even on constrained devices like Raspberry Pi 5!
🔑 Key Topics Covered:
- LLM quantization techniques (from FP16/FP8 to 4-bit precision)
- The GGUF format and LLAMA.cpp framework
- Why feed-forward layer parameters are more sensitive than attention layers
- Embedding model quantization using ONNX
- Vector database quantization methods (Product, Binary, and Scalar)
- Running vector databases and AI models on edge devices
This technical deep dive is perfect for developers looking to optimize AI models for memory-constrained environments or deploy vector search capabilities on edge devices.
Learn more from Weaviate at https://weaviate.io.
Femke just built one in 5 minutes, without writing a single line of code.
• Route general questions to Google search
• Direct document-specific queries to your knowledge base
Using Stack AI's no-code platform, I built a workflow that automatically chooses between these routes based on the query type.
From searching the web to processing 200+ page documents, building secure, robust, customizable AI applications and pipelines on your own data is way easier than most companies think it is.
And of course, powered by Weaviate's vector database for efficient document processing and search 💚
📽️ Watch the recording: "How AI Agents Can Transform Your Business" and see a live demo of building intelligent routing agents with Stack AI: events.weaviate.io/agents-stackai-webinar-2025?utm_source=linkedin&utm_medium=fp_social&utm_campaign=stack_ai&utm_content=video_post_806872823
Learn more about Weaviate: weaviate.io
Explore StackAI: StackAI: stack-ai.com
Hey everyone! I am SUPER EXCITED to publish the 118th episode of the Weaviate Podcast featuring David Berenstein and Ben Burtenshaw from HuggingFace! This podcast explores the intricacies of synthetic data generation, detailing methodologies such as data augmentation, distillation, and instruction refinement. The conversation delves into persona-driven synthetic data, highlighting applications like Persona Hub, and discusses algorithms to enhance diversity, complexity, and quality of generated data. Additionally, they cover integration with Hugging Face’s ecosystem, including Argilla for annotation, AutoTrain for fine-tuning, and advanced data exploration tools like the Data Studio and SQL console. The podcast also touches upon the potential for synthetic image data generation and the exciting future of AI education and accessibility.
Further Resources
Argilla: github.com/argilla-io/argilla
Distilabel: github.com/argilla-io/distilabel
Synthetic Data Generator: github.com/argilla-io/synthetic-data-generator
SelfInstruct: arxiv.org/abs/2212.10560
Magpie: arxiv.org/abs/2406.08464
TextClassification: arxiv.org/abs/2401.00368
RAG: arxiv.org/abs/2401.00368
DEITA: arxiv.org/abs/2312.15685
PersonaHub: arxiv.org/abs/2406.20094
PersonaHub Exploration: youtube.com/watch?v=timmCn8Nr6g
Fine Personas: huggingface.co/datasets/argilla/FinePersonas-v0.1#how-it-was-built
Fine-tune for function calling: huggingface.co/learn/agents-course/en/bonus-unit1/fine-tuning
Agents Course: huggingface.co/learn/agents-course/unit0/introduction
HuggingFace Reasoning Course: huggingface.co/reasoning-course
Text2SQL on HuggingFace Datasets: youtube.com/watch?v=5LUZq7MHolA
0:00 Welcome David and Ben!
0:56 The Synthetic Data Generator
3:25 Data Generation Algorithms
11:32 PersonaHub
15:15 Exploration in Data Generation
22:05 Systems for Data Generation
27:20 DSPy and Distilabel
33:30 Diversity in Data Generation
39:10 Synthetic Image Generation
44:50 Evolutionary Algorithms for Synthetic Data
50:48 Prompt to Model
59:25 Exciting Directions for AI
Along with that, it also tests out a bunch of different models with a new dataset, DBGorilla, designed to evaluate agentic querying techniques on real-world use cases.
Weaviate also just released a Query Agent, designed based on some of the work in this paper, to handle advanced agentic querying out of the box, find out more here: weaviate.io/blog/query-agent
▬▬▬▬▬▬▬▬▬▬▬▬ CONNECT WITH US ▬▬▬▬▬▬▬▬▬▬▬▬
- Visit http://weaviate.io
- Star us on GitHub github.com/weaviate/weaviate
- Stay updated and subscribe to our newsletter: newsletter.weaviate.io
- Try out Weaviate Cloud for free here: https://console.weaviate.cloud/
Got a question?
- Forum: forum.weaviate.io
- Slack: weaviate.io/slack
Connect with us on
- Twitter: twitter.com/weaviate_io
- LinkedIn: linkedin.com/company/weaviate-io
We dive into features like:
• Multi-valued Vectors aka multi-vector embeddings based on colBERT, enabling you to increase the accuracy of your search results
• A new B25 implementation based on BlockMax WAND, speeding up search performance,
• General Availability of role-based access control (RBAC)
• A new NVIDIA model integration
• and much more!
Read the full blog post to learn more: weaviate.io/blog/weaviate-1-29-release
00:00 Introduction
00:50 Agenda
01:07 Multi-vector embeddings
15:50 RBAC
33:06 Acync Replication
36:14 NVIDIA model integration
38:35 BlockMax WAND
41:40 2025 Weaviate Roadmap & Outlook
▬▬▬▬▬▬▬▬▬▬▬▬ CONNECT WITH US ▬▬▬▬▬▬▬▬▬▬▬▬
- Visit http://weaviate.io
- Star us on GitHub github.com/weaviate/weaviate
- Stay updated and subscribe to our newsletter: newsletter.weaviate.io
- Try out Weaviate Cloud for free here: https://console.weaviate.cloud/
Got a question?
- Forum: forum.weaviate.io
- Slack: weaviate.io/slack
Connect with us on
- Twitter: twitter.com/weaviate_io
- LinkedIn: linkedin.com/company/weaviate-io
Links:
Get started with Letta Desktop: docs.letta.com/quickstart/desktop
Stateful Agents: letta.com/blog/stateful-agents
MemGPT: arxiv.org/abs/2310.08560
MemGPT Explained! - youtube.com/watch?v=nQmZmFERmrg
LLMs Can Easily Learn to Reason from Demonstrations Structure, not content, is what matters! - arxiv.org/pdf/2502.07374
0:00 Welcome Sarah!
4:03 Stateful Agents
8:28 Context Window Optimization
9:13 Self-Optimizing Agents
12:43 Agent Personas
14:48 Gradient Descent and Prompts
19:58 Tools and Function Calling
21:53 Agent Development Environment (ADE)
31:54 Databases for Agents
36:28 LLM-Tool Use Examples
37:03 Multi-Agent Systems
41:28 Model Routers
43:26 Reasoning Models and Context Windows
47:53 Inspiration for ADE
50:52 Bring your own Tools
57:00 Future Directions for AI
This podcast explores different ideas around designing software user experiences for Agents as well as Humans. How do we write documentation for Agents differently than Humans? How do we design REST or gRPC APIs, or programming languages clients, for Agents differently than Humans? llms.txt, JSON tool definitions, agents to agent, breaking changes, … there were so many interesting topics explored in this podcast! I really hope you find it useful! As always more than happy to discuss these ideas further with you!
Read the blog here! - https://biilmann.blog/articles/introducing-ax/
Chapters
0:00 Weaviate Podcast #116!
0:20 Welcome Matt, Sebastian, Charles
1:45 Agent Experience
10:20 Docs for Agents
13:12 llms.txt
17:20 JSON Tool Definitions
21:16 Have your Agent talk to my Agent
26:27 gRPC APIs
29:56 Agentic Workflows
33:04 LLM Bias in Software Vendors
42:52 New APIs and Breaking Changes
Whether you need speed, quality, or performance, there's a perfect language model out there for your needs.
But with so many choices, how do you pick the best one?
Here are the 4 key factors to consider:
• 𝗤𝘂𝗮𝗹𝗶𝘁𝘆: Look for an LLM that performs consistently across tasks like chatbot interactions, language understanding, and coding.
Best options: DeepSeek R1, OpenAI's o1, OpenAI's o3 Mini
• 𝗣𝗿𝗶𝗰𝗲: Consider the cost per token for both input and output.
Best options: Gemini 2.0 Flash, GPT-4o Mini, and Mistral Small 3
• 𝗦𝗽𝗲𝗲𝗱/𝗧𝗵𝗿𝗼𝘂𝗴𝗵𝗽𝘂𝘁: Generally, smaller models like Mistral Small 3 and GPT-4o Mini are faster than larger, high-quality models.
Best options: Gemini 2.0 Flash, and GPT-4o Mini
• 𝗢𝗽𝗲𝗻 𝗦𝗼𝘂𝗿𝗰𝗲: Models that allow for private use
Best options: Llama 3, Microsoft Phi-4, Deepseek's latest models
• 𝗖𝗼𝗱𝗶𝗻𝗴: Models that excel at code generation
Best options: Claude 3.5 Sonnet, OpenAI's o3 Mini, Deepseek R1 & V3
In this video, I used the Quality vs. Throughput and Price Diagram from artificialanalysis.ai/models
Try it in @weaviate.io: weaviate.io/developers/weaviate/search/generative
We also dive into the STaRK Benchmark, a comprehensive testbed designed to evaluate retrieval systems on semi-structured knowledge bases. The discussion highlights the challenges of unifying textual and relational retrieval, exploring concepts such as multi-vector embeddings, relational graphs, and dynamic data modeling. Learn how these approaches help overcome information loss, enhance precision, and enable scalable, context-aware retrieval in diverse domains—from product recommendations to precision medicine.
Whether you’re interested in advanced prompt optimization, multi-agent system design, or the future of human-centered language models, this episode offers a wealth of insights and a forward-looking perspective on integrating sophisticated AI techniques into real-world applications.
Links:
AvaTaR: arxiv.org/abs/2406.11200
STaRK: arxiv.org/abs/2404.13207
Learn more about Shirley Wu: https://cs.stanford.edu/~shirwu/
Chapters
0:00 Welcome Shirely!
0:30 What is the state of AI?
2:18 Graph Data Models and GraphRAG
8:22 Unifying Textual and Relational Retrieval
10:50 Multi-Vector Retrieval
15:55 Embeddings and Properties
21:58 PrimeKG
27:25 AvaTaR: LLM Tool Use Optimization
41:40 Complex Retrieval
46:28 Agent Memory and DSPy
55:00 Future Directions for AI
Learn more about Contextual AI! - contextual.ai
Contextual AI Platform: contextual.ai/blog/contextual-ai-platform-generally-available
Links:
Anchored Preference Optimization: arxiv.org/abs/2408.06266
KTO: Model Alignment as Prospect Theoretic Optimization: arxiv.org/pdf/2402.01306
RETRO: arxiv.org/abs/2112.04426
Fusion-in-Decoder: arxiv.org/pdf/2007.01282
RAG: arxiv.org/pdf/2005.11401
LMUnit: contextual.ai/lmunit
Chapters
0:00 Welcome Amanpreet!
0:28 Founding Story and RAG 2.0
5:00 Agentic RAG
10:25 The Limits of Prompt Engineering
19:55 LMUnit Evals
28:35 RL for LLMs
40:00 Model Specialization
46:12 Mechanistic Interpretability
49:10 Back to RAG 2.0
55:15 Future Directions for AI
Paper: arxiv.org/abs/2501.12948
▬▬▬▬▬▬▬▬▬▬▬▬ CONNECT WITH US ▬▬▬▬▬▬▬▬▬▬▬▬
- Visit http://weaviate.io
- Star us on GitHub github.com/weaviate/weaviate
- Stay updated and subscribe to our newsletter: newsletter.weaviate.io
- Try out Weaviate Cloud for free here: https://console.weaviate.cloud/
Got a question?
- Forum: forum.weaviate.io
- Slack: weaviate.io/slack
Connect with us on
- Twitter: twitter.com/weaviate_io
- LinkedIn: linkedin.com/company/weaviate-io
Cartesia AI: cartesia.ai
Introduction to State Space Models (SSMs): huggingface.co/blog/lbourdois/get-on-the-ssm-train
Efficiently modeling long sequences with structured state spaces by Albert Gu, Karan Goel, and Christopher Re: arxiv.org/pdf/2111.00396
Karan Goel Google Scholar: scholar.google.com/citations?user=1i3X2GgAAAAJ
Chapters:
0:00 Welcome Karan!
0:37 Founding Cartesia AI
2:33 State Space Models (SSMs)
9:04 Audio Data Deep Dive
14:45 The launch of Sonic
23:30 Inference vs. Agent APIs
29:58 Learning, Search, and Reasoning
34:35 RETRO and Fusion-in-Decoder RAG
38:00 State in Sequence Models
42:00 CUDA for SSMs
49:00 Many Shot In-Context Learning
50:35 AI-Native Text-to-Speech Applications
(and how do they even work, anyways?)
- Learn more in this upcoming webinar: https://hubs.la/Q031qM_h0
- Weaviate Embeddings: weaviate.io/blog/introducing-weaviate-embeddings
- MTEB Leaderboard: huggingface.co/spaces/mteb/leaderboard
- How to chose an embedding model, on demand webinar: webinars.techstronglearning.com/simplify-building-ai-native-embedding-models-and-vector-databases
- Word2Vec paper: arxiv.org/abs/1301.3781
- Colorful vectors: huggingface.co/spaces/jphwang/colorful_vectors
▬▬▬▬▬▬▬▬▬▬▬▬ CONNECT WITH US ▬▬▬▬▬▬▬▬▬▬▬▬
- Visit http://weaviate.io
- Star us on GitHub github.com/weaviate/weaviate
- Stay updated and subscribe to our newsletter: newsletter.weaviate.io
- Try out Weaviate Cloud for free here: https://console.weaviate.cloud/
Got a question?
- Forum: forum.weaviate.io
- Slack: weaviate.io/slack
Connect with us on
- Twitter: twitter.com/weaviate_io
- LinkedIn: linkedin.com/company/weaviate-io
But, agentic RAG?
It's the strategist — planning steps, reasoning, and fetching exactly what's needed, even from the web or multiple tools!
Watch the full podcast: youtube.com/watch?v=Eh4uQq43jA4


