How to Build Self-Improving AI Agents with Coding Agents | Ep. 13 @arizeai
How to Build Self-Improving AI Agents with Coding Agents | Ep. 13  @arizeai
Uploaded July 2026 | Updated September 2026, 2 weeks ago
Learn how to feed eval explanations back into a coding agent so your system keeps getting better on its own, at scale. This is the finale: an agent that improves itself instead of waiting on you to notice it's broken.

Two prompts and a demo isn't a real system. The eval explanations you've been collecting all series are the highest-leverage input you have for fixing it for good.

Watch this to learn:
• Why eval explanations, not just labels, are the highest-leverage output you have
• How to hand hundreds of failing traces to a coding agent like Claude Code to find themes and propose fixes
• How to keep the loop honest: feed it your requirements, and verify you haven't regressed what already worked
• Where to start if you only have 15 minutes
This is the final video of Arize AX: Getting Started, the full self-improving AI lifecycle.

Chapters:
00:00 The finale: self-improving agents
00:26 Why two prompts isn't a real system
01:17 Eval explanations are the gold
02:23 Hand failing traces to a coding agent
03:00 The improvement workflow with Claude Code
04:02 Keep the loop honest, feed it requirements
04:41 Themes vs. individual failures
05:17 Worked example: a week in production
06:32 Verify you didn't regress what worked
06:47 The full software development lifecycle
07:53 Where to start, 15 minutes
08:43 Wrap-up and resources

👉 Sign up for free: app.arize.com/auth/login?utm_source=youtube&utm_medium=organic_social&utm_campaign=arize_ax_getting_started
🔗 Learn more about Arize AX: arize.com
📓 Colab notebook for the series: colab.research.google.com/drive/1dViThD0kJjbqtDGE7-ciBYIRwV-HBN1T
📚 Docs: docs.arize.com
🔔 Subscribe for weekly content on LLMs, agents, and evaluation: youtube.com/@arizeai?sub_confirmation=1

#ArizeAX #SelfImprovingAgents #AIAgents
How to Build Self-Improving AI Agents with Coding Agents | Ep. 13How Uber Evaluates AI Agents at Production Scale | Arize Observe 2026Identity, Permissions, and Security for AI Agents | WorkOS | Arize Observe 2026How My AI Agent Rewrites Itself Overnight | Chi Wang, AG2 | Arize Observe 2026Traces and Evals Explained: The Building Blocks of AI and Agent Testing | Ep. 2Nebulocks Ron Cahlon on Building AI for CybersecurityAI Agent Mastery Certification Course: Lab 6 – Agent EvalsHarnessing Splits in your Dataset with Arize PhoenixHomework 3 for AI Evals Course: LLM-as-a-Judge
Arize AI |

How to Build Self-Improving AI Agents with Coding Agents | Ep. 13

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER