Constructing and Judging Modern Agentic Workflows | Real Python Podcast #302 @realpython
Constructing and Judging Modern Agentic Workflows | Real Python Podcast #302  @realpython
Uploaded July 2026 | Updated September 2026, 2 weeks ago
How can you improve your LLM agent systems through specification enrichment? What are the advantages of having an LLM act as a judge within an agent system? This week on the show, Senior IEEE Member and Quality Engineer Suneet Malhotra joins us to discuss building and evaluating agentic architecture.

πŸ‘‰ Links from the show: realpython.com/podcasts/rpp/302

Suneet Malhotra is an independent practitioner-researcher with 18 years of experience in Quality Engineering (QE) and test automation for consumer-scale platforms. He discusses building specification-enrichment loops, monitoring performance, and using Cohen's kappa to measure agreement between LLM judgments.

Suneet is currently publishing multiple papers that are under peer review on these topics. He also provides links to his work and GitHub projects if you want to experiment with these concepts and methods yourself.

Topics:

- 00:00:00 -- Introduction
- 00:00:56 -- Survey: RP Podcast show notes
- 00:02:11 -- How did you get into testing?
- 00:05:04 -- Has working for large public-facing corporations changed how you approach testing?
- 00:07:06 -- Writing a paper on LLM-as-Judge
- 00:09:22 -- Looking across the Software Development Lifecycle
- 00:14:46 -- Agentic AI: theater vs methodology
- 00:17:23 -- Specification enrichment
- 00:27:52 -- Video Course Spotlight
- 00:29:18 -- Saving the specifications
- 00:31:27 -- Using the LLM as a judge & Cohen's kappa
- 00:39:31 -- How can people try out the project?
- 00:43:35 -- What are some of the failure modes you've seen?
- 00:50:26 -- What's an inexpensive way to try these ideas out?
- 00:54:04 -- What are you excited about in the world of Python?
- 00:56:14 -- What do you want to learn next?
- 00:57:14 -- How can people follow your work online?
- 00:57:32 -- Thanks and goodbye

πŸ‘‰ Links from the show: realpython.com/podcasts/rpp/302

Download your free Python Cheat Sheet here: realpython.com/cheatsheet
Free Python Skill Test with instant level + learning plan: realpython.com/skill-test
Want to learn faster? Become a Python Expert with unlimited access to 5,000+ tutorials, videos, and exercises: realpython.com/start

🐍 Become a Python expert with real-world tutorials, on-demand courses, interactive quizzes, and 24/7 access to a community of experts at realpython.com

β–°β–°β–°β–°β–°β–°β–°β–°β–°β–°β–°β–°β–°β–°β–°β–°β–°β–°β–°β–°
🐍 Start Here β†’ realpython.com/start
πŸ—ΊοΈ Guided Learning Paths β†’ realpython.com/learning-paths
🎧 Real Python Podcast β†’ realpython.com/podcast

πŸ“š Python Books β†’ realpython.com/books
πŸ“– Python Reference β†’ realpython.com/ref
πŸ§‘β€πŸ’» Quizzes & Exercises β†’ realpython.com/quizzes

πŸŽ“ Live Courses: realpython.com/live
⭐️ Reviews & Learner Stories: realpython.com/learner-stories
β–°β–°β–°β–°β–°β–°β–°β–°β–°β–°β–°β–°β–°β–°β–°β–°β–°β–°β–°β–°
Constructing and Judging Modern Agentic Workflows | Real Python Podcast #302Declarative vs Imperative Prompting for AI Agentsthe Complainer Who Poisons Your Work LifeShould You Understand Your Entire Python Codebase? | Real Python Podcast #305Communicate Uncomfortably MuchSmall, Specialized Open Source AI ModelsLazy Imports: Faster Startup!Custom Python List Comprehensions`except: pass` Is a Python Anti-PatternAI Slop Is Ruining Open Source PRsRunning Python Locally in a Sandbox | Real Python Podcast #301Secure Python Installs? Why You Should Always Use Wheels
Real Python |

Constructing and Judging Modern Agentic Workflows | Real Python Podcast #302

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER