Uploaded June 2026 | Updated September 2026, 2 weeks ago
[2026 - DAY 2 - AI ENGINEERING] Agents have evolved. We've moved from orchestration frameworks — where agents operate through definite steps and turns you defined upfront — to harnesses, where the LLM uses skills and tools to chart its own trajectory. Modern agents run for hours or days, producing hundreds to thousands of steps in a single session. This calls for a fundamentally different methodology for monitoring and evaluating them in production.
This talk shares what we've learned building observability infrastructure for agent harnesses at HoneyHive. We'll start with why the harness — not the model — has become the hardest engineering problem in production AI, and why traditional APM breaks down when traces are 10,000 spans deep and failures happen four tool calls deep. We'll walk through a live trajectory view to see what long-running agent traces actually look like at scale, and the specific challenges they create: context rot, semantic failure modes, and the needle-in-a-haystack problem of finding the moment that mattered.
Then we'll dig into skills as the new unit of behavior and the dual role of clustering in agent development: unsupervised clustering for discovering emergent patterns and identifying where guardrails are needed, and supervised classifiers for production evaluation at scale. We'll close on what comes next — swarm observability for multi-agent systems.
SPEAKER:
Sunny Bakhda - Founding Engineer, HoneyHive
👉 Sign up for our "No BS" Newsletter to get the latest technical data & AI content: aicouncil.com/newsletter
ABOUT AI COUNCIL:
AI Council brings together the brightest minds in data to share industry knowledge, technical architectures and best practices in building cutting edge data & AI systems and tools.
FIND US:
Website: aicouncil.com
LinkedIn: linkedin.com/company/aicouncilconf
X: https://x.com/aicouncilconf
[2026 - DAY 2 - AI ENGINEERING] Agents have evolved. We've moved from orchestration frameworks — where agents operate through definite steps and turns you defined upfront — to harnesses, where the LLM uses skills and tools to chart its own trajectory. Modern agents run for hours or days, producing hundreds to thousands of steps in a single session. This calls for a fundamentally different methodology for monitoring and evaluating them in production.
This talk shares what we've learned building observability infrastructure for agent harnesses at HoneyHive. We'll start with why the harness — not the model — has become the hardest engineering problem in production AI, and why traditional APM breaks down when traces are 10,000 spans deep and failures happen four tool calls deep. We'll walk through a live trajectory view to see what long-running agent traces actually look like at scale, and the specific challenges they create: context rot, semantic failure modes, and the needle-in-a-haystack problem of finding the moment that mattered.
Then we'll dig into skills as the new unit of behavior and the dual role of clustering in agent development: unsupervised clustering for discovering emergent patterns and identifying where guardrails are needed, and supervised classifiers for production evaluation at scale. We'll close on what comes next — swarm observability for multi-agent systems.
SPEAKER:
Sunny Bakhda - Founding Engineer, HoneyHive
👉 Sign up for our "No BS" Newsletter to get the latest technical data & AI content: aicouncil.com/newsletter
ABOUT AI COUNCIL:
AI Council brings together the brightest minds in data to share industry knowledge, technical architectures and best practices in building cutting edge data & AI systems and tools.
FIND US:
Website: aicouncil.com
LinkedIn: linkedin.com/company/aicouncilconf
X: https://x.com/aicouncilconf
![RLVR in Practice: From Synthetic Data to GRPO | NVIDIA
[2026 - DAY 3 - MODEL SYSTEMS] Reinforcement Learning from Verifiable Rewards (RLVR) is increasingly common in post-training pipelines, but the practical details are often glossed over. How do you design reward functions that programmatically verify model outputs? What makes synthetic training data effective? How do you build a custom RL environment that doesnt silently break your training?
SPEAKER: Chris Alexiuk - Product Research Engineer, NVIDIA
👉 Sign up for our No BS Newsletter to get the latest technical data & AI content: https://aicouncil.com/newsletter
ABOUT AI COUNCIL:
AI Council brings together the brightest minds in data to share industry knowledge, technical architectures and best practices in building cutting edge data & AI systems and tools.
FIND US:
Website: https://aicouncil.com/
LinkedIn: https://www.linkedin.com/company/aicouncilconf/
X: https://x.com/aicouncilconf RLVR in Practice: From Synthetic Data to GRPO | NVIDIA](https://i.ytimg.com/vi/sVyZVtnygD8/mqdefault.jpg)
![Optimizing Model Training End-to-End: A Tiny MoE Case Study Lambda
[2026 - DAY 3 - MODEL SYSTEMS] Cloud compute is expensive, and wasting runs on the guise of a just scale will fix any problems leaves you with less time to fix errors, and less compute to train the model you want. In this talk, I will discuss what are the easy optimizatiosn you might miss (minimizing communications, using the most effective algorithms, ensuring youre getting the most FLOPs possible) at the small scale, before ensuring that when you do scale up nothing is going to waste. In this particular talk, Ill be focusing on what worked at home, that then let me scale it further onto the cloud.
SPEAKER: Zach Mueller - Head of Developer Relations, Lambda
👉 Sign up for our No BS Newsletter to get the latest technical data & AI content: https://aicouncil.com/newsletter
ABOUT AI COUNCIL:
AI Council brings together the brightest minds in data to share industry knowledge, technical architectures and best practices in building cutting edge data & AI systems and tools.
FIND US:
Website: https://aicouncil.com/
LinkedIn: https://www.linkedin.com/company/aicouncilconf/
X: https://x.com/aicouncilconf Optimizing Model Training End-to-End: A Tiny MoE Case Study Lambda](https://i.ytimg.com/vi/s_hSPBYQ3BA/mqdefault.jpg)





![Agentic AI: From Risk Awareness to Practical Control | Noma Security
[2026 - DAY 3 - AI SECURITY & SAFETY] If you would not hand an intern your credentials, payment data, and production access, do not hand them to an AI agent without understanding the risk. Agentic systems do more than generate content. They take action across tools, data stores, and workflows with delegated authority. That shifts the trust boundary, expands identity and data risk, and stretches governance. This session explores practical controls security teams can apply now.
SPEAKER:
Diana Kelley - CISO, Noma Security
👉 Sign up for our No BS Newsletter to get the latest technical data & AI content: https://aicouncil.com/newsletter
ABOUT AI COUNCIL:
AI Council brings together the brightest minds in data to share industry knowledge, technical architectures and best practices in building cutting edge data & AI systems and tools.
FIND US:
Website: https://aicouncil.com/
LinkedIn: https://www.linkedin.com/company/aicouncilconf/
X: https://x.com/aicouncilconf Agentic AI: From Risk Awareness to Practical Control | Noma Security](https://i.ytimg.com/vi/uo_C7rh01GY/mqdefault.jpg)

![Guardrails for the Future AI Safety and Responsible AI in Practice
[2025 - Day 2 - Keynote] Jake Brill, Rachad Alao, Krishnaram Kenthapadi, and Daniel Olmedilla share insights from implementing responsible AI safeguards at scale, moving beyond theoretical discussions to explore the practical realities of deploying ethical AI systems. This panel offers candid perspectives on the complex trade-offs and technical challenges of responsible AI deployment, essential for anyone building trust and safety protocols or managing AI governance.
ABOUT THE SPEAKERS:
Jake Brill, Head of Product - Integrity, OpenAI
Rachad Alao, Senior Engineering Director, Meta
Krishnaram Kenthapadi, Chief Scientist - Clinical AI, Oracle Health
Daniel Olmedilla, Distinguished Engineer - AI & Trust, LinkedIn (Moderator) -
🎟️ GET YOUR TICKET TO AI COUNCIL 2026 🎟️
Meet the worlds top AI infrastructure minds where architects of AI share what works. Three days of high-quality technical talks and meaningful interactions.
→ https://aicouncil.com/sf-2026
⚡ FIND US:
X: https://x.com/AICouncilConf
LinkedIn: https://www.linkedin.com/company/aicouncilconf/
Website: https://aicouncil.com/ Guardrails for the Future AI Safety and Responsible AI in Practice](https://i.ytimg.com/vi/vW4VK-X2CKY/mqdefault.jpg)
![AI Launchpad 2025: Mooncake
[2025 - Day 1 - AI Launchpad] Pranav Aurora and Zhou Sun share insights from Mooncake, a real-time search and analytics system built on object-store for GenAI applications, exploring open development principles on open table formats. Like the delicacy its named after, this system offers valuable perspectives on building shared, open-source solutions that everyone can enjoy, particularly relevant for teams developing search and analytics capabilities for AI-powered applications.
ABOUT THE SPEAKERS:
Pranav Aurora, Co-Founder, Mooncake
Zhou Sun, Co-Founder, Mooncake -
🎟️ GET YOUR TICKET TO AI COUNCIL 2026 🎟️
Meet the worlds top AI infrastructure minds where architects of AI share what works. Three days of high-quality technical talks and meaningful interactions.
→ https://aicouncil.com/sf-2026
⚡ FIND US:
X: https://x.com/AICouncilConf
LinkedIn: https://www.linkedin.com/company/aicouncilconf/
Website: https://aicouncil.com/ AI Launchpad 2025: Mooncake](https://i.ytimg.com/vi/vWo2DD9P3UA/mqdefault.jpg)