Uploaded May 2026 | Updated September 2026, 2 weeks ago
Trace every agent run end-to-end, generate synthetic datasets to stress-test on demand, fire automated red team attacks at your own agents, and pin down exactly why evaluations fail — all from the Microsoft Foundry control plane.
Microsoft Foundry turns a working coding agent into production-ready software with built-in tracing, automatic and human evaluations, red teaming, and runtime guardrails that inspect every tool call. Mohammad Abuomar, Responsible AI Principal Architect, walks through the testing, evaluation, and controls that raise an agent's bar for quality, safety, performance, and cost.
👥 Who it's for: AI engineers and developers building agents, ML and AI platform teams, responsible AI and security leads, and IT decision-makers evaluating how to govern and deploy production AI agents on Azure.
⏱️ Chapters:
00:00 Why testing and controls beat model choice
00:33 See the finished production-ready coding agent
02:30 Where the agent started in the Foundry SDK
03:19 Trace every run with OTel + Azure Monitor
04:04 Built-in monitoring: cost, tokens, error rate
04:34 Evaluation types and synthetic datasets
05:51 Automated red team attacks on your agent
07:08 Evaluation results and AI cluster analysis
08:14 Fix failures with runtime guardrails
09:12 Fleet-wide visibility and enforced policies
The model or framework you pick is only part of the story — the testing, evaluation, and controls surrounding your agent are what make it production-ready. This walkthrough follows a coding agent that takes a simple prompt and builds a working app, then shows exactly where the Microsoft Foundry control plane comes in to make it better.
Start in Visual Studio Code with an agent built on the Foundry SDK, wired up with tools like web search and code interpreter. In the Foundry portal, review OTel traces of every run, backed by Azure Monitor, and drill into any Trace ID to read the input and output turns, system message, and tool outputs — far easier to parse than standard logs. Built-in monitoring rolls up estimated cost, token usage, agent runs, tool calls, and error rate over time.
Then move into evaluations: automatic evaluation powered by AI, human evaluation with surveys, and red teaming that runs automated attacks to expose vulnerabilities. Save time by generating a synthetic dataset from a prompt, choose built-in evaluators for agents, quality, and safety, and launch red team runs using attack strategies like AsciiSmuggler, Base64, Jailbreak, StringJoin, UnicodeSubstitution, and IndirectJailbreak. Use AI-driven cluster analysis to see why a TaskCompletion score falls short, then fix it with a runtime guardrail that checks task adherence on every tool call. Foundry also delivers fleet-wide visibility and centrally enforced policies to keep every agent compliant.
► Link References
Get everything you need in Microsoft Foundry at ai.azure.com
► Unfamiliar with Microsoft Mechanics? Microsoft's Official Video Series for IT
- Subscribe youtube.com/c/MicrosoftMechanicsSeries
- Microsoft Tech Community: techcommunity.microsoft.com/t5/microsoft-mechanics-blog/bg-p/MicrosoftMechanicsBlog
- Podcast: microsoftmechanics.libsyn.com/podcast
► Join us on social:
- twitter.com/MSFTMechanics
- linkedin.com/company/microsoft-mechanics
- instagram.com/msftmechanics
- tiktok.com/@msftmechanics
#MicrosoftFoundry #AIAgents #AgentOps #RedTeaming #ResponsibleAI
Trace every agent run end-to-end, generate synthetic datasets to stress-test on demand, fire automated red team attacks at your own agents, and pin down exactly why evaluations fail — all from the Microsoft Foundry control plane.
Microsoft Foundry turns a working coding agent into production-ready software with built-in tracing, automatic and human evaluations, red teaming, and runtime guardrails that inspect every tool call. Mohammad Abuomar, Responsible AI Principal Architect, walks through the testing, evaluation, and controls that raise an agent's bar for quality, safety, performance, and cost.
👥 Who it's for: AI engineers and developers building agents, ML and AI platform teams, responsible AI and security leads, and IT decision-makers evaluating how to govern and deploy production AI agents on Azure.
⏱️ Chapters:
00:00 Why testing and controls beat model choice
00:33 See the finished production-ready coding agent
02:30 Where the agent started in the Foundry SDK
03:19 Trace every run with OTel + Azure Monitor
04:04 Built-in monitoring: cost, tokens, error rate
04:34 Evaluation types and synthetic datasets
05:51 Automated red team attacks on your agent
07:08 Evaluation results and AI cluster analysis
08:14 Fix failures with runtime guardrails
09:12 Fleet-wide visibility and enforced policies
The model or framework you pick is only part of the story — the testing, evaluation, and controls surrounding your agent are what make it production-ready. This walkthrough follows a coding agent that takes a simple prompt and builds a working app, then shows exactly where the Microsoft Foundry control plane comes in to make it better.
Start in Visual Studio Code with an agent built on the Foundry SDK, wired up with tools like web search and code interpreter. In the Foundry portal, review OTel traces of every run, backed by Azure Monitor, and drill into any Trace ID to read the input and output turns, system message, and tool outputs — far easier to parse than standard logs. Built-in monitoring rolls up estimated cost, token usage, agent runs, tool calls, and error rate over time.
Then move into evaluations: automatic evaluation powered by AI, human evaluation with surveys, and red teaming that runs automated attacks to expose vulnerabilities. Save time by generating a synthetic dataset from a prompt, choose built-in evaluators for agents, quality, and safety, and launch red team runs using attack strategies like AsciiSmuggler, Base64, Jailbreak, StringJoin, UnicodeSubstitution, and IndirectJailbreak. Use AI-driven cluster analysis to see why a TaskCompletion score falls short, then fix it with a runtime guardrail that checks task adherence on every tool call. Foundry also delivers fleet-wide visibility and centrally enforced policies to keep every agent compliant.
► Link References
Get everything you need in Microsoft Foundry at ai.azure.com
► Unfamiliar with Microsoft Mechanics? Microsoft's Official Video Series for IT
- Subscribe youtube.com/c/MicrosoftMechanicsSeries
- Microsoft Tech Community: techcommunity.microsoft.com/t5/microsoft-mechanics-blog/bg-p/MicrosoftMechanicsBlog
- Podcast: microsoftmechanics.libsyn.com/podcast
► Join us on social:
- twitter.com/MSFTMechanics
- linkedin.com/company/microsoft-mechanics
- instagram.com/msftmechanics
- tiktok.com/@msftmechanics
#MicrosoftFoundry #AIAgents #AgentOps #RedTeaming #ResponsibleAI










