Uploaded July 2026 | Updated September 2026, 3 weeks ago
When an AI feature fails in production, developers almost always blame model intelligence. But the reality of building production LLM applications is very different: the model isn't the bottleneck; your product specification and evaluation harness are.
In this episode, Hamel Husain from Parlance Labs breaks down why traditional software paradigms fail when applied to generative AI. You cannot specify "quality" upfront in a traditional test suite; you have to find what actually matters by observing live system traces and bringing domain experts into the evaluation loop early.
Chapters:
00:00 Introduction
00:19 When was the model not the problem?
00:36 Why can't you specify quality upfront?
01:08 How do you find what's most important?
01:28 Why is it important to get domain experts as soon as possible?
02:00 What's the most important thing to learn when it comes to evals?
Resources:
🔬 Phoenix (open source): phoenix.arize.com
đź”— Arize AX: arize.com
đź“– OpenInference: github.com/Arize-ai/openinference
đź“– Phoenix docs: docs.arize.com/phoenix
Subscribe to Arize AI to watch the full "Rise of the AI Engineer" interview series! 🚀: youtube.com/@arizeai?sub_confirmation=1
#AIEngineering #ProductEngineering #ArizeAI
When an AI feature fails in production, developers almost always blame model intelligence. But the reality of building production LLM applications is very different: the model isn't the bottleneck; your product specification and evaluation harness are.
In this episode, Hamel Husain from Parlance Labs breaks down why traditional software paradigms fail when applied to generative AI. You cannot specify "quality" upfront in a traditional test suite; you have to find what actually matters by observing live system traces and bringing domain experts into the evaluation loop early.
Chapters:
00:00 Introduction
00:19 When was the model not the problem?
00:36 Why can't you specify quality upfront?
01:08 How do you find what's most important?
01:28 Why is it important to get domain experts as soon as possible?
02:00 What's the most important thing to learn when it comes to evals?
Resources:
🔬 Phoenix (open source): phoenix.arize.com
đź”— Arize AX: arize.com
đź“– OpenInference: github.com/Arize-ai/openinference
đź“– Phoenix docs: docs.arize.com/phoenix
Subscribe to Arize AI to watch the full "Rise of the AI Engineer" interview series! 🚀: youtube.com/@arizeai?sub_confirmation=1
#AIEngineering #ProductEngineering #ArizeAI










