Uploaded October 2024 | Updated September 2026, 1 hour ago
Hey everyone! Thanks so much for watching the 105th episode of the Weaviate Podcast with Philip Kiely! This one dives into all sorts of apsects related to Compound AI Systems! We are now seeing far better results with AI models by breaking up tasks into multiple stages and inferences. Philip explains the work they are doing at Baseten to optimize and scale deployments of these emerging systems and all sorts of aspects about them from Structured Generation to their distinction with Agents! I hope you find it useful!
References:
Learn more about Baseten: docs.baseten.co/welcome
Let Me Speak Freely? - arxiv.org/abs/2408.02442
StructuredRAG - github.com/weaviate/structured-rag
No more bad outputs with structured generation: Remi Louf - youtube.com/watch?v=aNmfvN6S_n4
Coalescence: making LLM inference 5x faster - blog.dottxt.co/coalescence.html
Cameron Pfiffer (.txt) - https://x.com/cameron_pfiffer
Remi Louf (.txt) - https://x.com/remilouf
OPRO - arxiv.org/abs/2309.03409
ALTO - An Efficient Network Orchestrator for Compound AI Systems - arxiv.org/pdf/2403.04311
Chapters:
0:00 Welcome Philip!
0:25 The state of AI models
2:30 Multimodal Models
3:35 Structured Outputs in Multimodal Generation
7:48 Generative Feedback Loops and Structured Outputs
10:30 Structured Outputs in Compound AI Systems
13:35 Deploying and Scaling Compound AI Systems
22:30 Transformers, Mixture-of-Experts, SSMs
27:52 vLLM
33:00 Examples of Compound AI Systems
Hey everyone! Thanks so much for watching the 105th episode of the Weaviate Podcast with Philip Kiely! This one dives into all sorts of apsects related to Compound AI Systems! We are now seeing far better results with AI models by breaking up tasks into multiple stages and inferences. Philip explains the work they are doing at Baseten to optimize and scale deployments of these emerging systems and all sorts of aspects about them from Structured Generation to their distinction with Agents! I hope you find it useful!
References:
Learn more about Baseten: docs.baseten.co/welcome
Let Me Speak Freely? - arxiv.org/abs/2408.02442
StructuredRAG - github.com/weaviate/structured-rag
No more bad outputs with structured generation: Remi Louf - youtube.com/watch?v=aNmfvN6S_n4
Coalescence: making LLM inference 5x faster - blog.dottxt.co/coalescence.html
Cameron Pfiffer (.txt) - https://x.com/cameron_pfiffer
Remi Louf (.txt) - https://x.com/remilouf
OPRO - arxiv.org/abs/2309.03409
ALTO - An Efficient Network Orchestrator for Compound AI Systems - arxiv.org/pdf/2403.04311
Chapters:
0:00 Welcome Philip!
0:25 The state of AI models
2:30 Multimodal Models
3:35 Structured Outputs in Multimodal Generation
7:48 Generative Feedback Loops and Structured Outputs
10:30 Structured Outputs in Compound AI Systems
13:35 Deploying and Scaling Compound AI Systems
22:30 Transformers, Mixture-of-Experts, SSMs
27:52 vLLM
33:00 Examples of Compound AI Systems






- Star us on GitHub https://github.com/weaviate/weaviate
- Stay updated and subscribe to our newsletter: https://newsletter.weaviate.io/
- Try out Weaviate Cloud Services for free here: https://console.weaviate.cloud/
Got a question?
- Forum: https://forum.weaviate.io/
- Slack: https://weaviate.io/slack
Connect with us on
- Twitter: https://twitter.com/weaviate_io
- LinkedIn: https://www.linkedin.com/company/weaviate-io/
Tuana Socials
- Twitter/X: https://x.com/tuanacelik
- LinkedIn: https://www.linkedin.com/in/tuanacelik/ Weaviate Tech Hands-On: Query Agent](https://i.ytimg.com/vi/tqTFTS_E4i0/mqdefault.jpg)




