Glean on AI Agent Evals, Permissions, and Production Trust | Arize Observe 2026 @arizeai
Glean on AI Agent Evals, Permissions, and Production Trust | Arize Observe 2026  @arizeai
Uploaded July 2026 | Updated September 2026, 3 weeks ago
How do you build enterprise AI agents that people can actually trust?

In this Observe 2026 fireside chat, Eddie Zhou, founding engineer at Glean, shares what Glean has learned from building production AI systems for enterprise knowledge, work AI, agents, and company context.

Eddie discusses Glean’s early production lessons, including a real customer permissions scare, why AI systems expose human error, how Glean closes the demo-to-production gap through dogfooding, and why evals, observability, telemetry, and debugging infrastructure matter from day zero.

The conversation also covers why there is no single “silver bullet” eval, how LLM-as-judge evolves into full evaluation harnesses, what makes multi-turn agent evals so hard, how Glean debugs production quality issues, and how engineering teams should think about trust, accountability, permissions, and AI-generated code.

You’ll learn:
- What Glean is and how it thinks about work AI, agents, and enterprise context
- Why enterprise AI systems need strict permissioning and trust boundaries
- How AI can expose human error in source systems
- Why production dogfooding helps close the demo gap
- Why evals require a suite of methods, not one perfect metric
- How Glean uses real work artifacts, synthetic evals, online feedback, and dogfooding
- Why LLM-as-judge is evolving into evaluation harnesses
- Why multi-turn agent evals are still an open challenge
- How Glean approaches debugging, trace analysis, and production quality escalations
- Why AI engineers still need to understand the systems they are building
- How to avoid “slop cannon” engineering culture as AI-generated work increases

Chapters:
0:00 What is Glean?
0:47 Glean as work AI and enterprise context
1:22 The first production permissions scare
2:35 When AI exposes human error
3:45 Closing the demo gap
4:26 Dogfooding the real production system
5:27 Telemetry, traces, and debugging search
6:54 Why there’s no silver-bullet eval
7:50 Grounding evals in real work outcomes
9:05 Why online evals and dogfooding still matter
10:41 From LLM-as-judge to eval harnesses
12:37 The hard problem of multi-turn evals
14:41 Debugging AI agents in production
15:56 Context failures: too much or too little
17:15 Automating trace analysis
18:25 Why AI engineers still need to understand the system
20:07 Detecting subtle regressions
22:34 Permissions, actions, and agent trust
23:20 How agents change accountability
25:24 Building a culture of AI quality
26:17 Avoiding the “slop cannon” problem

Presented at Observe 2026.

#Glean #AIAgents #EnterpriseAI

🔗 Try Arize AX & Phoenix OSS: arize.com
🔔 Subscribe for weekly content on LLMs, agents, and evaluation: youtube.com/@arizeai?sub_confirmation=1
Glean on AI Agent Evals, Permissions, and Production Trust | Arize Observe 2026Build Your First Eval: Creating a Custom LLM Evaluator with a Golden DatasetCUGA Agent: From Benchmarks to Business Impact of IBMs Generalist AgentOne AI Question - whats a hot take on evals, with Cam YoungAI Agent Mastery Certification Course: Module 7 – Post-Deployment & MonitoringDataDog CEO Olivier Pomel On the Future of AI and Agent EngineeringOne AI Question - when should I start doing evals, with Aparna DhinakaranAlyx: Cursor-Like AI Agent for AI Engineering (Short Demo)How to Track & Cut Coding Agent Spend (Claude Code & Cursor) | AI BuildersMulti-Agent Observability: Debugging Agent-to-Agent Communication | Band | Arize Observe 2026AI Agents in Production: From Demos to Durable ROI | CVS Health | Arize Observe 2026AI Agent Mastery Certification Course: Module 3 – Agent Architectures & Frameworks
Arize AI |

Glean on AI Agent Evals, Permissions, and Production Trust | Arize Observe 2026

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER