From Spans to Trajectories: Observability for Long-Running Agents | HoneyHive @aicouncilconf
From Spans to Trajectories: Observability for Long-Running Agents | HoneyHive  @aicouncilconf
Uploaded June 2026 | Updated September 2026, 2 weeks ago
[2026 - DAY 2 - AI ENGINEERING] Agents have evolved. We've moved from orchestration frameworks — where agents operate through definite steps and turns you defined upfront — to harnesses, where the LLM uses skills and tools to chart its own trajectory. Modern agents run for hours or days, producing hundreds to thousands of steps in a single session. This calls for a fundamentally different methodology for monitoring and evaluating them in production.

This talk shares what we've learned building observability infrastructure for agent harnesses at HoneyHive. We'll start with why the harness — not the model — has become the hardest engineering problem in production AI, and why traditional APM breaks down when traces are 10,000 spans deep and failures happen four tool calls deep. We'll walk through a live trajectory view to see what long-running agent traces actually look like at scale, and the specific challenges they create: context rot, semantic failure modes, and the needle-in-a-haystack problem of finding the moment that mattered.

Then we'll dig into skills as the new unit of behavior and the dual role of clustering in agent development: unsupervised clustering for discovering emergent patterns and identifying where guardrails are needed, and supervised classifiers for production evaluation at scale. We'll close on what comes next — swarm observability for multi-agent systems.

SPEAKER:
Sunny Bakhda - Founding Engineer, HoneyHive

👉 Sign up for our "No BS" Newsletter to get the latest technical data & AI content: aicouncil.com/newsletter

ABOUT AI COUNCIL:
AI Council brings together the brightest minds in data to share industry knowledge, technical architectures and best practices in building cutting edge data & AI systems and tools.

FIND US:
Website: aicouncil.com
LinkedIn: linkedin.com/company/aicouncilconf
X: https://x.com/aicouncilconf
From Spans to Trajectories: Observability for Long-Running Agents | HoneyHiveRLVR in Practice: From Synthetic Data to GRPO | NVIDIAOptimizing Model Training End-to-End: A Tiny MoE Case Study LambdaAI Launchpad 2026: Golden AnalyticsChang She on Why He Walked Away from Parquet to Build LanceDBThe 2% gains that 10x your inference | Lessons from AWS on optimizing VLMsQ&A with Scott Breitenother, Kilo: Engineers need to be the CEOs of agents. Are they ready?Beyond MLOps: Building AI systems with MetaflowAgentic AI: From Risk Awareness to Practical Control | Noma SecurityShould agents be durable? | RenderGuardrails for the Future AI Safety and Responsible AI in PracticeAI Launchpad 2025: Mooncake
AI Council |

From Spans to Trajectories: Observability for Long-Running Agents | HoneyHive

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER