Uploaded August 2026 | Updated September 2026, 3 weeks ago
Where is your enterprise AI budget actually going, and is that spend translating into measurable quality? As developer teams adopt coding agents like Claude Code, Cursor, and GitHub Copilot alongside custom production LLM applications, tracking inference spend across models, token types, and execution spans becomes critical to proving ROI.
In this episode of AI Builders, we detail how to use Arize AX to gain complete visibility into your AI bills, connect spend directly to eval quality metrics, and automate cost optimization workflows.
Chapters:
00:00 Where is your AI budget going?
00:39 Intro: Cost alongside quality
01:37 Coding agents vs. production agent spend
03:07 The two drivers of AI cost: model price and token usage
05:05 Why cheaper models do not always mean lower cost
06:32 Connecting cost to quality and ROI
08:27 What does a “better” AI result actually mean?
11:52 The cost-to-quality decision framework in Arize AX
14:11 Tracing cost down to individual model calls
15:40 Monitoring AI cost spikes in production
17:04 Using managed agents to investigate cost
19:17 Demo: Setting up the Arize AX Cost Agent
22:15 How managed agents work
24:07 Finding the biggest sources of wasted spend
27:07 Automatically proposing cost-saving code changes
29:15 Why you still need to validate cost reductions
31:07 Common ways to reduce AI costs
33:30 When your evals become too expensive
35:01 Testing cheaper models without sacrificing quality
37:45 Protecting the AI spend that creates value
39:18 Demo: Tracing Claude Code cost and efficiency
41:19 Using Signal to find coding agent inefficiencies
43:22 Turning trace insights into reusable agent skills
45:09 Five takeaways for reducing AI cost without hurting quality
46:17 Q&A: The cost and quality decision framework
47:15 Closing
Resources:
🔬 Phoenix (open source): phoenix.arize.com
🔗 Arize AX: arize.com
📖 OpenInference: github.com/Arize-ai/openinference
📖 Phoenix docs: docs.arize.com/phoenix
What techniques is your team using to keep LLM bills under control as you scale agents? Share your setup in the comments below!
What techniques is your team using to keep LLM bills under control as you scale agents? Share your setup in the comments below!
Subscribe and hit the notification bell to catch every episode of AI Builders! youtube.com/@arizeai?sub_confirmation=1
#AIEngineering #LLMCost #AIObservability
Where is your enterprise AI budget actually going, and is that spend translating into measurable quality? As developer teams adopt coding agents like Claude Code, Cursor, and GitHub Copilot alongside custom production LLM applications, tracking inference spend across models, token types, and execution spans becomes critical to proving ROI.
In this episode of AI Builders, we detail how to use Arize AX to gain complete visibility into your AI bills, connect spend directly to eval quality metrics, and automate cost optimization workflows.
Chapters:
00:00 Where is your AI budget going?
00:39 Intro: Cost alongside quality
01:37 Coding agents vs. production agent spend
03:07 The two drivers of AI cost: model price and token usage
05:05 Why cheaper models do not always mean lower cost
06:32 Connecting cost to quality and ROI
08:27 What does a “better” AI result actually mean?
11:52 The cost-to-quality decision framework in Arize AX
14:11 Tracing cost down to individual model calls
15:40 Monitoring AI cost spikes in production
17:04 Using managed agents to investigate cost
19:17 Demo: Setting up the Arize AX Cost Agent
22:15 How managed agents work
24:07 Finding the biggest sources of wasted spend
27:07 Automatically proposing cost-saving code changes
29:15 Why you still need to validate cost reductions
31:07 Common ways to reduce AI costs
33:30 When your evals become too expensive
35:01 Testing cheaper models without sacrificing quality
37:45 Protecting the AI spend that creates value
39:18 Demo: Tracing Claude Code cost and efficiency
41:19 Using Signal to find coding agent inefficiencies
43:22 Turning trace insights into reusable agent skills
45:09 Five takeaways for reducing AI cost without hurting quality
46:17 Q&A: The cost and quality decision framework
47:15 Closing
Resources:
🔬 Phoenix (open source): phoenix.arize.com
🔗 Arize AX: arize.com
📖 OpenInference: github.com/Arize-ai/openinference
📖 Phoenix docs: docs.arize.com/phoenix
What techniques is your team using to keep LLM bills under control as you scale agents? Share your setup in the comments below!
What techniques is your team using to keep LLM bills under control as you scale agents? Share your setup in the comments below!
Subscribe and hit the notification bell to catch every episode of AI Builders! youtube.com/@arizeai?sub_confirmation=1
#AIEngineering #LLMCost #AIObservability










![When AI Can Write Code, What Are Software Engineers Worth? | Citadel
When AI can generate code in seconds, what still makes a software engineer valuable?
In this Arize:Observe session, Craig Owenby of Citadel explores how agentic coding tools are changing software engineering, and why the profession still requires far more than producing code.
Craig compares large language models to the printing press. The printing press replaced the manual work of copying books, but it did not replace authors. In the same way, AI coding agents can automate the mechanics of writing code without replacing the judgment, vision, empathy, and experience required to build useful software.
The session covers:
• Why “coder” and “software engineer” are increasingly different roles
• How tools like Claude Code, Codex, Copilot, and Cursor remove traditional barriers to building software
• What the printing press teaches us about AI-assisted development
• Why engineers should avoid competing with AI on raw code generation
• How Sears lost its advantage by competing with e-commerce on the wrong terms
• Why human experience, intuition, and empathy remain essential
• How constraints can improve product and engineering decisions
• Why shipping more features can increase volatility and reduce user trust
• How AI acts as leverage for strong and weak engineering decisions
• Why product direction and problem selection matter more as implementation gets easier
The central lesson: software engineering is not primarily about writing code. It is about deciding what should be built, understanding why it matters, and applying technology with judgment.
When everyone can code, engineers differentiate themselves through their standards, instincts, product sense, and ability to understand the people using what they build. :contentReference[oaicite:0]{index=0}
Chapters:
00:00 When everyone can code, what are engineers worth?
00:46 LLMs and the printing press
01:30 The limits that shaped software engineering
02:42 We finally live in a world where everyone can code
03:45 The existential question for experienced engineers
04:45 Coders versus software engineers
06:05 Don’t make the same mistake as Sears
07:38 Human judgment, experience, and empathy
09:02 Why constraints can produce better software
10:25 Product volatility and the Sharpe ratio
11:49 Software engineering is about solving problems
12:38 LLMs as leverage for engineers
13:45 What sets engineers apart
14:25 Your humanity is the key
🔗 Learn more about Arize: https://arize.com
🔗 Explore Arize:Observe: https://arize.com/observe
🔔 Subscribe for more talks on AI engineering, coding agents, evaluation, observability, and the future of software development:
https://www.youtube.com/@arizeai?sub_confirmation=1
#SoftwareEngineering #CodingAgents #AIEngineering When AI Can Write Code, What Are Software Engineers Worth? | Citadel](https://i.ytimg.com/vi/TllPOmVWF8s/mqdefault.jpg)