Uploaded June 2026 | Updated September 2026, 2 weeks ago
In this episode of "AWS Show and Tell – Build Agents That Self-Improve: Evaluations, Insights, and Optimization with Amazon Bedrock AgentCore", we will show you how to move beyond “it works on my machine” and build agentic applications you can actually trust in real-world environments. Using Amazon Bedrock AgentCore, you will learn how to treat agents as observable systems, with full visibility into every interaction so you can see exactly what your agent did, why it did it, and where things went wrong. We will walk through best practices for evaluating agents at each stage of their lifecycle, from early development and pre-production testing to continuous monitoring and drift detection in production. Then we will demonstrate a practical observe, evaluate, optimize loop: using AgentCore’s evaluation, optimization and insights, to surface hallucinations and failure modes, understand root causes, and iteratively improve agent prompts, tools, and policies over time. By the end, you will have a concrete playbook for building agents that are debuggable, measurable, and continuously optimized in production.
- AgentCore Evaluations overview: docs.aws.amazon.com/bedrock-agentcore/latest/devguide/evaluations.html
- AgentCore Optimization “how it works” (recommendations, A/B tests, bundles): docs.aws.amazon.com/bedrock-agentcore/latest/devguide/optimization-how-it-works.html
- Insights “how it works”: docs.aws.amazon.com/bedrock-agentcore/latest/devguide/insights-how-it-works.html
- Feature code samples: github.com/awslabs/agentcore-samples/tree/main/01-features/06-observe-evaluate-optimize-your-agent
- Use case code sample: github.com/awslabs/agentcore-samples/tree/main/02-use-cases/01-conversational-agents/market-trends-agent
In this episode of "AWS Show and Tell – Build Agents That Self-Improve: Evaluations, Insights, and Optimization with Amazon Bedrock AgentCore", we will show you how to move beyond “it works on my machine” and build agentic applications you can actually trust in real-world environments. Using Amazon Bedrock AgentCore, you will learn how to treat agents as observable systems, with full visibility into every interaction so you can see exactly what your agent did, why it did it, and where things went wrong. We will walk through best practices for evaluating agents at each stage of their lifecycle, from early development and pre-production testing to continuous monitoring and drift detection in production. Then we will demonstrate a practical observe, evaluate, optimize loop: using AgentCore’s evaluation, optimization and insights, to surface hallucinations and failure modes, understand root causes, and iteratively improve agent prompts, tools, and policies over time. By the end, you will have a concrete playbook for building agents that are debuggable, measurable, and continuously optimized in production.
- AgentCore Evaluations overview: docs.aws.amazon.com/bedrock-agentcore/latest/devguide/evaluations.html
- AgentCore Optimization “how it works” (recommendations, A/B tests, bundles): docs.aws.amazon.com/bedrock-agentcore/latest/devguide/optimization-how-it-works.html
- Insights “how it works”: docs.aws.amazon.com/bedrock-agentcore/latest/devguide/insights-how-it-works.html
- Feature code samples: github.com/awslabs/agentcore-samples/tree/main/01-features/06-observe-evaluate-optimize-your-agent
- Use case code sample: github.com/awslabs/agentcore-samples/tree/main/02-use-cases/01-conversational-agents/market-trends-agent



