Uploaded June 2026 | Updated September 2026, 3 weeks ago
What does it take to trust AI agents before they go to production?
In this Observe 2026 session, Manjit Singh from Salesforce walks through a practical framework for evaluating AI agents across the agent development lifecycle — from single-agent evals to multi-agent systems, orchestrators, handoffs, and end-to-end workflow behavior.
The core idea: you can live code, but you can’t live operate. To scale AI agents, teams need evals, test datasets, metrics, observability, and a clear strategy for catching failures before and after production.
You’ll learn:
- Why evals are essential for building trust in AI agents
- What layer of the stack teams should evaluate
- The three core components of an eval strategy: framework, test dataset, and metrics
- How to choose between human review, LLM-as-judge, and code-based checks
- How to evaluate single-agent and multi-turn systems
- Why multi-agent systems introduce new failure modes
- How to test individual agents, handoffs, and end-to-end workflows
- What can go wrong with routing, reasoning loops, retries, memory, and context handoffs
Chapters:
0:00 Intro: a practical guide to agent evals
0:32 Session roadmap: from eval basics to multi-agent systems
1:37 Agent development lifecycle at Salesforce
2:01 “You can live code, but you can’t live operate”
2:46 What evals are and why they build trust
3:46 What layer of the stack should you evaluate?
5:48 Evaluation frameworks, test datasets, and metrics
6:29 Offline vs. online evals
8:23 Choosing humans, LLM judges, and code-based checks
10:30 Why you have to look at your logs
11:04 Single-agent evals
11:58 Evaluating multi-turn conversations
13:11 The multi-agent evaluation challenge
13:52 Super agents, orchestrators, and decomposed tasks
15:21 Common failure modes in multi-agent systems
16:09 The T × N failure-probability problem
16:35 Example: router, task agent, and research agent
17:07 Trajectory vs. end-to-end workflow evaluation
17:50 Three layers of multi-agent evals
18:07 Layer 1: testing each individual agent
20:46 Metrics to track: cost, latency, and quality
21:20 Layer 2: testing handoffs and context preservation
21:38 Travel-booking example: when handoffs go wrong
Presented at Observe 2026.
#AIAgents #Salesforce #AIEvals
🔗 Try Arize AX & Phoenix OSS: arize.com
🔔 Subscribe for weekly content on LLMs, agents, and evaluation: youtube.com/@arizeai?sub_confirmation=1
What does it take to trust AI agents before they go to production?
In this Observe 2026 session, Manjit Singh from Salesforce walks through a practical framework for evaluating AI agents across the agent development lifecycle — from single-agent evals to multi-agent systems, orchestrators, handoffs, and end-to-end workflow behavior.
The core idea: you can live code, but you can’t live operate. To scale AI agents, teams need evals, test datasets, metrics, observability, and a clear strategy for catching failures before and after production.
You’ll learn:
- Why evals are essential for building trust in AI agents
- What layer of the stack teams should evaluate
- The three core components of an eval strategy: framework, test dataset, and metrics
- How to choose between human review, LLM-as-judge, and code-based checks
- How to evaluate single-agent and multi-turn systems
- Why multi-agent systems introduce new failure modes
- How to test individual agents, handoffs, and end-to-end workflows
- What can go wrong with routing, reasoning loops, retries, memory, and context handoffs
Chapters:
0:00 Intro: a practical guide to agent evals
0:32 Session roadmap: from eval basics to multi-agent systems
1:37 Agent development lifecycle at Salesforce
2:01 “You can live code, but you can’t live operate”
2:46 What evals are and why they build trust
3:46 What layer of the stack should you evaluate?
5:48 Evaluation frameworks, test datasets, and metrics
6:29 Offline vs. online evals
8:23 Choosing humans, LLM judges, and code-based checks
10:30 Why you have to look at your logs
11:04 Single-agent evals
11:58 Evaluating multi-turn conversations
13:11 The multi-agent evaluation challenge
13:52 Super agents, orchestrators, and decomposed tasks
15:21 Common failure modes in multi-agent systems
16:09 The T × N failure-probability problem
16:35 Example: router, task agent, and research agent
17:07 Trajectory vs. end-to-end workflow evaluation
17:50 Three layers of multi-agent evals
18:07 Layer 1: testing each individual agent
20:46 Metrics to track: cost, latency, and quality
21:20 Layer 2: testing handoffs and context preservation
21:38 Travel-booking example: when handoffs go wrong
Presented at Observe 2026.
#AIAgents #Salesforce #AIEvals
🔗 Try Arize AX & Phoenix OSS: arize.com
🔔 Subscribe for weekly content on LLMs, agents, and evaluation: youtube.com/@arizeai?sub_confirmation=1










