Uploaded September 2024 | Updated September 2026, 1 hour ago
AI Researchers have overfit to maximizing state-of-the-art accuracy at the expense of the cost to run these AI systems! We need to account for cost during optimization. Even if a chatbot can produce an amazing answer, it isn't that valuable if it costs, say $5 per response!
I am beyond excited to present the 104th Weaviate Podcast with Sayash Kapoor and Benedikt Stroebl from Princeton Language and Intelligence! Sayash and Benedikt are co-first authors of "AI Agents That Matter"! This is one of my favorite papers I've studied recently which introduces Pareto Optimal optimization to DSPy and really tames the chaos of Agent benchmarking!
This was such a fun conversation! I am beyond grateful to have met them both and to feature their research on the Weaviate Podcast! I hope you find it interesting and useful!
Chapters
0:00 Welcome Sayash and Benedikt!
1:04 Inspiration for AI Agents That Matter
4:28 Pareto Optimal Agent Optimization
21:40 Tool Use Considerations
24:54 Generative Feedback Loops
27:15 RAG versus Long Context LLM Benchmarking
30:05 Problems with Agent Benchmarks
38:45 Measuring Intelligence
44:05 Agent Architectures
49:10 Human Feedback in Agent Optimization
53:55 What directions for the future of AI excite you the most?
Links:
Sayash Kapoor on X: https://x.com/sayashk
Benedikt Stroebl on X: https://x.com/benediktstroebl
AI Agents That Matter: arxiv.org/pdf/2407.01502
Accuracy-Cost Tradeoff Visual: benediktstroebl.github.io/agent-eval-webapp
AI Snake Oil: https://press.princeton.edu/books/hardcover/9780691249131/ai-snake-oil
AlphaCode: arxiv.org/abs/2203.07814
More Agents are all you need: arxiv.org/pdf/2402.05120
Are More LM Calls All You Need? Towards the Scaling Properties of Compound AI Systems: arxiv.org/pdf/2403.02419
Inference Scaling Laws (Google): arxiv.org/pdf/2407.21787
Exploring Inference Scaling with Modal by Charles Frye: modal.com/blog/llama-human-eval
Efficiently Estimating Pareto Frontiers from Jacob Portes (referenced as "from MosaicML" in the podcast): databricks.com/blog/efficiently-estimating-pareto-frontiers
Automated Design of Agentic Systems: arxiv.org/abs/2408.08435
SWE-bench: arxiv.org/abs/2310.06770
Generative Feedback Loops: youtube.com/watch?v=H7d-2nn63vU
MemGPT: arxiv.org/abs/2310.08560
On the Measure of Intelligence: arxiv.org/abs/1911.01547
Please subscribe to the channel to see more episodes of the Weaviate Podcast!
AI Researchers have overfit to maximizing state-of-the-art accuracy at the expense of the cost to run these AI systems! We need to account for cost during optimization. Even if a chatbot can produce an amazing answer, it isn't that valuable if it costs, say $5 per response!
I am beyond excited to present the 104th Weaviate Podcast with Sayash Kapoor and Benedikt Stroebl from Princeton Language and Intelligence! Sayash and Benedikt are co-first authors of "AI Agents That Matter"! This is one of my favorite papers I've studied recently which introduces Pareto Optimal optimization to DSPy and really tames the chaos of Agent benchmarking!
This was such a fun conversation! I am beyond grateful to have met them both and to feature their research on the Weaviate Podcast! I hope you find it interesting and useful!
Chapters
0:00 Welcome Sayash and Benedikt!
1:04 Inspiration for AI Agents That Matter
4:28 Pareto Optimal Agent Optimization
21:40 Tool Use Considerations
24:54 Generative Feedback Loops
27:15 RAG versus Long Context LLM Benchmarking
30:05 Problems with Agent Benchmarks
38:45 Measuring Intelligence
44:05 Agent Architectures
49:10 Human Feedback in Agent Optimization
53:55 What directions for the future of AI excite you the most?
Links:
Sayash Kapoor on X: https://x.com/sayashk
Benedikt Stroebl on X: https://x.com/benediktstroebl
AI Agents That Matter: arxiv.org/pdf/2407.01502
Accuracy-Cost Tradeoff Visual: benediktstroebl.github.io/agent-eval-webapp
AI Snake Oil: https://press.princeton.edu/books/hardcover/9780691249131/ai-snake-oil
AlphaCode: arxiv.org/abs/2203.07814
More Agents are all you need: arxiv.org/pdf/2402.05120
Are More LM Calls All You Need? Towards the Scaling Properties of Compound AI Systems: arxiv.org/pdf/2403.02419
Inference Scaling Laws (Google): arxiv.org/pdf/2407.21787
Exploring Inference Scaling with Modal by Charles Frye: modal.com/blog/llama-human-eval
Efficiently Estimating Pareto Frontiers from Jacob Portes (referenced as "from MosaicML" in the podcast): databricks.com/blog/efficiently-estimating-pareto-frontiers
Automated Design of Agentic Systems: arxiv.org/abs/2408.08435
SWE-bench: arxiv.org/abs/2310.06770
Generative Feedback Loops: youtube.com/watch?v=H7d-2nn63vU
MemGPT: arxiv.org/abs/2310.08560
On the Measure of Intelligence: arxiv.org/abs/1911.01547
Please subscribe to the channel to see more episodes of the Weaviate Podcast!










