Uploaded July 2025 | Updated September 2026, 2 weeks ago
In this practical and engaging talk, Ben McHone and Vicky Bang from Source Allies walk through their journey building a multi-agent system to support enterprise decision-making. They share real-world lessons in:
Structuring agents for specific responsibilities
Using evaluations to track and improve performance
Architecting for complexity while reducing token load
Building trust with business users through transparent observability
Chapters:
00:00 - Intro: Jack of All Trades or Master of One?
00:55 - Real-world enterprise use case
02:30 - Why traditional BI dashboards weren’t enough
03:40 - Introducing the GenAI-powered data analyst agent
05:00 - Using evaluations to improve reliability
06:15 - The challenge of large context windows
07:15 - Managing complexity in prompts and examples
08:30 - Techniques for taming context: tools, RAG, and summarization
09:45 - Evaluation results: From 30% to 95% accuracy
10:25 - Handling business-specific terminology
11:00 - Transition to multi-agent system: business term and supervisor agents
12:10 - Modular multi-agent benefits: context trimming and specialization
13:30 - Architecture insights: permissioning and reuse
14:00 - Why evaluations + observability build trust
15:00 - Final takeaways and live Q&A
Watch to learn why sometimes the best AI agent isn't a jack-of-all-trades—but a master of one (or a team of specialists working together).
🛠️ Tools and techniques covered include:
Dynamic system prompts
Agentic RAG
OpenTelemetry + Arize Phoenix
OAuth-based permission handling
Practical eval-driven development
👥 Follow the speakers and their work:
🔗 Ben McHone — linkedin.com/in/benjamin-mchone-4b983abb
🔗 Vicky Bang — linkedin.com/in/vickyyunqi
📢 Explore Arize Phoenix. github.com/Arize-ai/phoenix
📅 More community events: arize.com/community
In this practical and engaging talk, Ben McHone and Vicky Bang from Source Allies walk through their journey building a multi-agent system to support enterprise decision-making. They share real-world lessons in:
Structuring agents for specific responsibilities
Using evaluations to track and improve performance
Architecting for complexity while reducing token load
Building trust with business users through transparent observability
Chapters:
00:00 - Intro: Jack of All Trades or Master of One?
00:55 - Real-world enterprise use case
02:30 - Why traditional BI dashboards weren’t enough
03:40 - Introducing the GenAI-powered data analyst agent
05:00 - Using evaluations to improve reliability
06:15 - The challenge of large context windows
07:15 - Managing complexity in prompts and examples
08:30 - Techniques for taming context: tools, RAG, and summarization
09:45 - Evaluation results: From 30% to 95% accuracy
10:25 - Handling business-specific terminology
11:00 - Transition to multi-agent system: business term and supervisor agents
12:10 - Modular multi-agent benefits: context trimming and specialization
13:30 - Architecture insights: permissioning and reuse
14:00 - Why evaluations + observability build trust
15:00 - Final takeaways and live Q&A
Watch to learn why sometimes the best AI agent isn't a jack-of-all-trades—but a master of one (or a team of specialists working together).
🛠️ Tools and techniques covered include:
Dynamic system prompts
Agentic RAG
OpenTelemetry + Arize Phoenix
OAuth-based permission handling
Practical eval-driven development
👥 Follow the speakers and their work:
🔗 Ben McHone — linkedin.com/in/benjamin-mchone-4b983abb
🔗 Vicky Bang — linkedin.com/in/vickyyunqi
📢 Explore Arize Phoenix. github.com/Arize-ai/phoenix
📅 More community events: arize.com/community










