Uploaded July 2026 | Updated September 2026, 3 weeks ago
Most AI demos look promising. Far fewer make it into production with measurable ROI.
In this Observe 2026 session, CVS Health shares what it takes to move AI initiatives from demos and stalled pilots into durable, production-grade systems. The core message: model capability is not the bottleneck. The real work is building the discipline around the model — evaluation, governance, deployment, observability, economics, and measurable business outcomes.
You’ll learn:
- Why better models do not automatically create durable AI ROI
- How AI initiatives get stuck in “pilot purgatory”
- The three layers of an AI-native SDLC: personal productivity, collaboration/process, and validation/observability
- Why teams need shared tools, reusable practices, clear specs, and documented tribal knowledge
- Why evals are to GenAI what unit tests are to software engineering
- Why cost per outcome matters more than cost per token
- How to tie AI spend to business KPIs like cycle time, throughput, error rate, and time to resolution
- Why explainability, audit trails, and guardrails are essential in regulated environments
- How to build AI systems that can be trusted, measured, governed, and stopped when they are not working
Chapters:
0:00 From demos to durable ROI
0:58 CVS Health’s AI production journey
1:30 Real gains from AI in production
2:12 The bottleneck isn’t the model
3:04 Pilot purgatory is a discipline problem
3:28 Three layers of the AI-native SDLC
4:54 Layer 1: personal productivity
6:05 Where AI coding tools fail
7:23 Standardizing tools so practices compound
7:50 Layer 2: collaboration and process
9:18 When individual speed creates team bottlenecks
10:03 Layer 3: validation and observability
11:26 User validation, autonomy, and compliance risk
12:50 Three disciplines for durable ROI
13:43 Evals are unit tests for GenAI
14:24 Cost per outcome beats cost per token
15:00 Tie AI spend to business KPIs
15:37 Data hygiene and durable infrastructure
16:26 Automate the system, not the step
17:32 Explainability and governance
18:30 Guardrails in regulated environments
20:53 Final takeaways: measuring durable AI ROI
Presented at Observe 2026.
#AIAgents #EnterpriseAI #CVSHealth
🔗 Try Arize AX & Phoenix OSS: arize.com
🔔 Subscribe for weekly content on LLMs, agents, and evaluation: youtube.com/@arizeai?sub_confirmation=1
Most AI demos look promising. Far fewer make it into production with measurable ROI.
In this Observe 2026 session, CVS Health shares what it takes to move AI initiatives from demos and stalled pilots into durable, production-grade systems. The core message: model capability is not the bottleneck. The real work is building the discipline around the model — evaluation, governance, deployment, observability, economics, and measurable business outcomes.
You’ll learn:
- Why better models do not automatically create durable AI ROI
- How AI initiatives get stuck in “pilot purgatory”
- The three layers of an AI-native SDLC: personal productivity, collaboration/process, and validation/observability
- Why teams need shared tools, reusable practices, clear specs, and documented tribal knowledge
- Why evals are to GenAI what unit tests are to software engineering
- Why cost per outcome matters more than cost per token
- How to tie AI spend to business KPIs like cycle time, throughput, error rate, and time to resolution
- Why explainability, audit trails, and guardrails are essential in regulated environments
- How to build AI systems that can be trusted, measured, governed, and stopped when they are not working
Chapters:
0:00 From demos to durable ROI
0:58 CVS Health’s AI production journey
1:30 Real gains from AI in production
2:12 The bottleneck isn’t the model
3:04 Pilot purgatory is a discipline problem
3:28 Three layers of the AI-native SDLC
4:54 Layer 1: personal productivity
6:05 Where AI coding tools fail
7:23 Standardizing tools so practices compound
7:50 Layer 2: collaboration and process
9:18 When individual speed creates team bottlenecks
10:03 Layer 3: validation and observability
11:26 User validation, autonomy, and compliance risk
12:50 Three disciplines for durable ROI
13:43 Evals are unit tests for GenAI
14:24 Cost per outcome beats cost per token
15:00 Tie AI spend to business KPIs
15:37 Data hygiene and durable infrastructure
16:26 Automate the system, not the step
17:32 Explainability and governance
18:30 Guardrails in regulated environments
20:53 Final takeaways: measuring durable AI ROI
Presented at Observe 2026.
#AIAgents #EnterpriseAI #CVSHealth
🔗 Try Arize AX & Phoenix OSS: arize.com
🔔 Subscribe for weekly content on LLMs, agents, and evaluation: youtube.com/@arizeai?sub_confirmation=1








![When AI Can Write Code, What Are Software Engineers Worth? | Citadel
When AI can generate code in seconds, what still makes a software engineer valuable?
In this Arize:Observe session, Craig Owenby of Citadel explores how agentic coding tools are changing software engineering, and why the profession still requires far more than producing code.
Craig compares large language models to the printing press. The printing press replaced the manual work of copying books, but it did not replace authors. In the same way, AI coding agents can automate the mechanics of writing code without replacing the judgment, vision, empathy, and experience required to build useful software.
The session covers:
• Why “coder” and “software engineer” are increasingly different roles
• How tools like Claude Code, Codex, Copilot, and Cursor remove traditional barriers to building software
• What the printing press teaches us about AI-assisted development
• Why engineers should avoid competing with AI on raw code generation
• How Sears lost its advantage by competing with e-commerce on the wrong terms
• Why human experience, intuition, and empathy remain essential
• How constraints can improve product and engineering decisions
• Why shipping more features can increase volatility and reduce user trust
• How AI acts as leverage for strong and weak engineering decisions
• Why product direction and problem selection matter more as implementation gets easier
The central lesson: software engineering is not primarily about writing code. It is about deciding what should be built, understanding why it matters, and applying technology with judgment.
When everyone can code, engineers differentiate themselves through their standards, instincts, product sense, and ability to understand the people using what they build. :contentReference[oaicite:0]{index=0}
Chapters:
00:00 When everyone can code, what are engineers worth?
00:46 LLMs and the printing press
01:30 The limits that shaped software engineering
02:42 We finally live in a world where everyone can code
03:45 The existential question for experienced engineers
04:45 Coders versus software engineers
06:05 Don’t make the same mistake as Sears
07:38 Human judgment, experience, and empathy
09:02 Why constraints can produce better software
10:25 Product volatility and the Sharpe ratio
11:49 Software engineering is about solving problems
12:38 LLMs as leverage for engineers
13:45 What sets engineers apart
14:25 Your humanity is the key
🔗 Learn more about Arize: https://arize.com
🔗 Explore Arize:Observe: https://arize.com/observe
🔔 Subscribe for more talks on AI engineering, coding agents, evaluation, observability, and the future of software development:
https://www.youtube.com/@arizeai?sub_confirmation=1
#SoftwareEngineering #CodingAgents #AIEngineering When AI Can Write Code, What Are Software Engineers Worth? | Citadel](https://i.ytimg.com/vi/TllPOmVWF8s/mqdefault.jpg)

