WWDC26: Create robust evaluations for agentic apps | Apple @AppleDeveloper
WWDC26: Create robust evaluations for agentic apps | Apple  @AppleDeveloper
Uploaded June 2026 | Updated September 2026, 1 week ago
Learn how to leverage advanced features of the Evaluations framework to build robust evaluations for your app. Explore evaluating flows with tool calling and dynamic conditions, and how to define what correct behavior means for your use case. Discover how to generate synthetic data, use judges effectively, and validate your datasets for reliable results.


Explore related documentation, sample code, and more:
Generating synthetic datasets: developer.apple.com/documentation/Evaluations/generating-synthetic-evaluation-datasets
Evaluating tool-calling behavior: developer.apple.com/documentation/Evaluations/evaluating-tool-calling-behavior
Scoring with model-as-judge evaluators: developer.apple.com/documentation/Evaluations/scoring-with-model-as-judge-evaluators
Book Tracker: Using Evaluations to evaluate an intelligent feature: developer.apple.com/documentation/Evaluations/book-tracker-using-evaluations-to-evaluate-an-intelligent-feature

More Apple Developer resources:
Video sessions: apple.co/VideoSessions
Documentation: apple.co/DeveloperDocs
Forums: apple.co/DeveloperForums
App: apple.co/DeveloperApp
WWDC26: Create robust evaluations for agentic apps | AppleWWDC26: Build real-time neural rendering pipelines with Metal | AppleSpeedrun your game port to MacWWDC26: What’s new in image understanding | AppleSwiftUI drag and drop in just 2 stepsWWDC26: Best practices for integrating visual intelligence in your app | AppleSmooth gameplay, no tweaks! Cyberpunk 2077 on MacControl every device from your Mac with Device HubShip AI with confidenceWWDC26: Get the most out of Device Hub | AppleLearning SwiftUI in a world of agentic AIThe 2026 Apple Design Award winners are here!
Apple Developer |

WWDC26: Create robust evaluations for agentic apps | Apple

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER