5 Lessons from the Classroom for Evaluating Agents | Dev Interrupted @aicouncilconf
5 Lessons from the Classroom for Evaluating Agents | Dev Interrupted  @aicouncilconf
Uploaded June 2026 | Updated September 2026, 3 weeks ago
[2026 - DAY 2 - ANALYTICS & DATA SCI] Education researchers have spent a century figuring out how to build environments where complex, unpredictable systems self-correct through structured feedback. Software engineers building AI agents are solving the exact same problem from scratch and ignoring all of it. This talk bridges that gap. I'm a former classroom teacher turned AI engineer, and I'll walk you through five pedagogical frameworks that map directly to eval design patterns for AI agents: backward design (define success criteria before you build), formative assessment (eval continuously, not just at the end), rubric design (multi-dimensional scoring instead of pass/fail), error analysis (categorize failure modes because same symptom doesn't mean same cause), and differentiated feedback (the agent, the user, and the knowledge base each need their own signal channel). What ties them together is back pressure: each framework is a way to capture signal from problems and route it to where it drives change. That's what makes a system self-correcting instead of just self-reporting. Most AI observability is still about watching systems after the fact. This talk is about designing systems where the eval layer captures back pressure from failures and feeds it back upstream, so the system iterates on itself. I've built production agents and won hackathons with this approach, and the core insight is simple: the best eval systems aren't tests, they're environments. And nobody knows more about designing those environments than teachers.

SPEAKER:
Andrew Zigler - GTM Engineer, LinearB & Podcast Host, Dev Interrupted

๐Ÿ‘‰ Sign up for our "No BS" Newsletter to get the latest technical data & AI content: aicouncil.com/newsletter

ABOUT AI COUNCIL:
AI Council brings together the brightest minds in data to share industry knowledge, technical architectures and best practices in building cutting edge data & AI systems and tools.

FIND US:
Website: aicouncil.com
LinkedIn: linkedin.com/company/aicouncilconf
X: https://x.com/aicouncilconf
5 Lessons from the Classroom for Evaluating Agents | Dev InterruptedHow Vercel Builds Dozens of Metrics from One Heterogenous TableFrom Scaling to Observability Solving Key Challenges for Distributed ML with RayFrom Playgrounds to Production: The Evolution of AI Evaluation at CodaAI: too good to be true, too bad to be useful | TypeSafe AIDuckDB Co-Creator Hannes Mรผhleisen on Why Single-Node Beats DistributedWhy dbt Acquired SDF Building true SQL ComprehensionAGI is Already Here (But Its Not What You Think)AI Launchpad 2025: NAOReal-Time Analytics for Small Data TeamsBuilding a Data Native Agent: Cortex Code | SnowflakeCalvin French-Owen on the Future of Agentic Coding
AI Council |

5 Lessons from the Classroom for Evaluating Agents | Dev Interrupted

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER