Uploaded August 2026 | Updated September 2026, 2 weeks ago
How do you know whether AI coding agents are actually making your engineering team more productive?
In this Arize AX walkthrough, Doug shows how code harness tracing can turn Claude Code sessions into measurable signals about engineering productivity, agent efficiency, cost, and team performance.
Instead of looking only at token consumption, you can trace complete coding-agent sessions and connect what the agent did to outcomes such as commits, PRs, points shipped, retry loops, and reusable skill opportunities.
The demo covers:
• Tracing complete multi-turn Claude Code sessions
• Inspecting prompts, responses, Bash commands, file operations, and other tool calls
• Tracking token usage, latency, users, sessions, and agent behavior
• Detecting whether a coding session resulted in a PR or commit
• Identifying retry loops and inefficient agent exploration
• Finding repeated workflows that could become reusable agent skills
• Scoring coding-agent efficiency with an agentic evaluator
• Using explanations to understand why a session was efficient or inefficient
• Using Signal to surface recurring problems across coding-agent traces
• Detecting sessions with excessive prompt tokens, retries, or missing skills
• Syncing trace data to BigQuery for organization and team analytics
• Comparing coding-agent spend with points and commits shipped
• Measuring cost per point across teams and individual users
• Identifying power users and developers who may be getting stuck in agent loops
• Turning recurring workflows into skills that can improve the entire team
The result is a way to evaluate coding agents at both the individual-session and organizational level.
A single trace can tell you why an agent struggled. Across a team, the same telemetry can help answer broader questions: Which teams are getting the most value from coding agents? Where is spend increasing without corresponding output? Which workflows should become shared skills? Where are agents repeatedly getting stuck?
Arize AX combines tracing, evaluations, Signal, and downstream analytics to help teams connect coding-agent behavior to measurable engineering outcomes.
Chapters:
00:00 Measuring coding agent productivity with tracing
01:58 Did the session ship a PR or commit?
04:22 Evaluating coding agent efficiency
06:14 Measuring cost, points shipped, and team ROI
08:22 Comparing coding agent performance by user
10:22 Finding reusable agent skills
🔗 Explore Arize AX: arize.com/ax
🔗 Learn more about Arize: arize.com
🔔 Subscribe for more videos on AI agents, coding agents, evaluation, observability, and AI engineering:
youtube.com/@arizeai?sub_confirmation=1
#CodingAgents #ClaudeCode #AIObservability
How do you know whether AI coding agents are actually making your engineering team more productive?
In this Arize AX walkthrough, Doug shows how code harness tracing can turn Claude Code sessions into measurable signals about engineering productivity, agent efficiency, cost, and team performance.
Instead of looking only at token consumption, you can trace complete coding-agent sessions and connect what the agent did to outcomes such as commits, PRs, points shipped, retry loops, and reusable skill opportunities.
The demo covers:
• Tracing complete multi-turn Claude Code sessions
• Inspecting prompts, responses, Bash commands, file operations, and other tool calls
• Tracking token usage, latency, users, sessions, and agent behavior
• Detecting whether a coding session resulted in a PR or commit
• Identifying retry loops and inefficient agent exploration
• Finding repeated workflows that could become reusable agent skills
• Scoring coding-agent efficiency with an agentic evaluator
• Using explanations to understand why a session was efficient or inefficient
• Using Signal to surface recurring problems across coding-agent traces
• Detecting sessions with excessive prompt tokens, retries, or missing skills
• Syncing trace data to BigQuery for organization and team analytics
• Comparing coding-agent spend with points and commits shipped
• Measuring cost per point across teams and individual users
• Identifying power users and developers who may be getting stuck in agent loops
• Turning recurring workflows into skills that can improve the entire team
The result is a way to evaluate coding agents at both the individual-session and organizational level.
A single trace can tell you why an agent struggled. Across a team, the same telemetry can help answer broader questions: Which teams are getting the most value from coding agents? Where is spend increasing without corresponding output? Which workflows should become shared skills? Where are agents repeatedly getting stuck?
Arize AX combines tracing, evaluations, Signal, and downstream analytics to help teams connect coding-agent behavior to measurable engineering outcomes.
Chapters:
00:00 Measuring coding agent productivity with tracing
01:58 Did the session ship a PR or commit?
04:22 Evaluating coding agent efficiency
06:14 Measuring cost, points shipped, and team ROI
08:22 Comparing coding agent performance by user
10:22 Finding reusable agent skills
🔗 Explore Arize AX: arize.com/ax
🔗 Learn more about Arize: arize.com
🔔 Subscribe for more videos on AI agents, coding agents, evaluation, observability, and AI engineering:
youtube.com/@arizeai?sub_confirmation=1
#CodingAgents #ClaudeCode #AIObservability










