Uploaded January 2025 | Updated September 2026, 1 hour ago
Hey everyone! Thank you so much for watching the 111th episode of the Weaviate Podcast with Aravind Kesiraju! Aravind is a Principal Software Engineer at Morningstar where he has led the effort behind the Morningstar Intelligence Engine! The podcast begins by describing the Morningstar Intelligence Engine, an API-based AI platform that helps asset managers, wealth managers, and other financial-services firms conduct investment research faster than ever. We then dive into Aravind’s early journey with Weaviate and the wave of RAG, Agents, and Vector Databases. The podcast then covers a series of deep dives into Data Pipeline tooling for RAG with topics such as embedding queues with Amazon SQS and Chunking strategies such as Anthropic’s Contextual Retrieval and Doc2query data ingestion. The synthetic data topic then pivots in evals and using LLMs for content generation. We then cover Agents and Function Calling, digging deeper into the Morningstar Product Action Catalog. We then cover Autogen! Autogen is a Multi-Agent framework that has been around for a little while but is getting a lot of buzz lately with the pivot to AG2 and spin off from Microsoft to Google DeepMind. We then transition into the topic of Hallucination Guardrails, which I think is particularly interesting for things like financial / healthcare Agents. We then discuss Aravind's work on Text-to-SQL and exciting directions for the future of AI!
Links:
Morningstar Intelligence Engine: morningstar.com/business/brands/data-analytics/products/direct-web-services/features/intelligence-engine
Morningstar on Weaviate Case Studies: weaviate.io/case-studies/morningstar
Doc2Query: arxiv.org/abs/1904.08375
Contextual Retrieval: anthropic.com/news/contextual-retrieval
Chapters
0:00 Welcome Aravind
0:32 Morningstar Intelligence Engine
2:35 Early Journey with LLMs and RAG
6:50 Evolution of RAG Pipelines
9:25 Data Pipeline Tooling
12:15 Chunking Strategies
16:25 Evals
18:28 State of LLMs
19:55 Fine-Tuning LLMs for Content Writing
23:35 Agents and Function Calling
27:50 AutoGen
31:15 Background Task Agents
33:00 Hallucination Guardrails
41:24 Text-to-SQL
48:30 What future directions for AI excite you the most?
Hey everyone! Thank you so much for watching the 111th episode of the Weaviate Podcast with Aravind Kesiraju! Aravind is a Principal Software Engineer at Morningstar where he has led the effort behind the Morningstar Intelligence Engine! The podcast begins by describing the Morningstar Intelligence Engine, an API-based AI platform that helps asset managers, wealth managers, and other financial-services firms conduct investment research faster than ever. We then dive into Aravind’s early journey with Weaviate and the wave of RAG, Agents, and Vector Databases. The podcast then covers a series of deep dives into Data Pipeline tooling for RAG with topics such as embedding queues with Amazon SQS and Chunking strategies such as Anthropic’s Contextual Retrieval and Doc2query data ingestion. The synthetic data topic then pivots in evals and using LLMs for content generation. We then cover Agents and Function Calling, digging deeper into the Morningstar Product Action Catalog. We then cover Autogen! Autogen is a Multi-Agent framework that has been around for a little while but is getting a lot of buzz lately with the pivot to AG2 and spin off from Microsoft to Google DeepMind. We then transition into the topic of Hallucination Guardrails, which I think is particularly interesting for things like financial / healthcare Agents. We then discuss Aravind's work on Text-to-SQL and exciting directions for the future of AI!
Links:
Morningstar Intelligence Engine: morningstar.com/business/brands/data-analytics/products/direct-web-services/features/intelligence-engine
Morningstar on Weaviate Case Studies: weaviate.io/case-studies/morningstar
Doc2Query: arxiv.org/abs/1904.08375
Contextual Retrieval: anthropic.com/news/contextual-retrieval
Chapters
0:00 Welcome Aravind
0:32 Morningstar Intelligence Engine
2:35 Early Journey with LLMs and RAG
6:50 Evolution of RAG Pipelines
9:25 Data Pipeline Tooling
12:15 Chunking Strategies
16:25 Evals
18:28 State of LLMs
19:55 Fine-Tuning LLMs for Content Writing
23:35 Agents and Function Calling
27:50 AutoGen
31:15 Background Task Agents
33:00 Hallucination Guardrails
41:24 Text-to-SQL
48:30 What future directions for AI excite you the most?










