Uploaded July 2026 | Updated September 2026, 3 weeks ago
KodeKloud AI tutor β https://goo.gle/44Zrg8C
Welcome back to The Agent Factory! In this episode, host Luke Schlangen sits down with Mumshad Mannambeth, the founder of KodeKloud, to explore the future of DevOps, cloud upskilling, and building production-ready AI agents.
Mumshad shares an exclusive, behind-the-scenes look at how KodeKloud built their cutting-edge AI Tutorβa hands-on, AI-assisted learning platform that has doubled student retention! We dive deep into the multi-agent architecture under the hood, showing how they orchestrate five specialized agents using LangGraph and cost-effective Gemini Flash models on Google Cloud. Mumshad also walks us through their implementation of Model Armor on Cloud Run to secure LLM calls against prompt injections and malicious inputs, and how they synchronize AI-generated voice with dynamic React code to generate video lessons on the fly.
We also tackle the shifting landscape of tech interviews, why foundational knowledge in Kubernetes and Linux is more critical than ever, and how to avoid developer burnout in the age of AI. Stick around for a rapid-fire round where Mumshad shares his thoughts on whether certifications are overrated, if Python should be your first language, and how to navigate the overwhelming CNCF landscape.
Chapters:
0:00 - Intro
0:44 - Do we still need Kubernetes?
1:32 - CKA vs CKAD: Admin vs Developer Kubernetes certs
3:35 - Luke's personal certification journey
4:22 - Running Large Language Models (LLMs) on Kubernetes
5:10 - Do we need humans in DevOps if AI can code?
7:27 - The future of coding interviews & system design
11:25 - The AI productivity paradox: Why are we working more?
12:26 - Navigating the massive CNCF landscape
14:53 - How to learn effectively in the AI era
17:05 - Demystifying KodeKloud's AI Tutor
18:31 - Hands-on Demo: Kubernetes Services Lab
24:07 - Facing the Kubernetes Quiz & Live Terminal
30:31 - Interactive debugging with KodeKloud's Live Assist
35:06 - Sneak peek behind the curtain: AI Tutor under the hood
37:03 - Hosting & scaling 30k virtual labs on Google Cloud
39:06 - How the AI assistant doubled student retention
42:52 - Multi-Agent orchestration with LangGraph
46:14 - Syncing 11 Labs audio with custom React UI
48:11 - Leveraging cost-effective Gemini Flash models
51:25 - Personalization and student memory
53:17 - Google Cloud Run deployment & architecture
54:15 - Protecting LLMs from prompt injection with Model Armor
55:55 - Combatting crypto miners in free sandboxes
58:23 - Rapid Fire: Kubernetes, certifications, and developer burnout
1:03:05 - Mumshad's advice on career transitions & lifelong learning
1:04:32 - Outro
More resources:
How to deploy an AI agent to Cloud Run (step by step) β https://goo.gle/4haPa8w
Watch more of The Agent Factory β youtube.com/playlist?list=PLIivdWyY5sqLXR1eSkiM5bE6pFlXC-OSs
π Subscribe to Google Cloud Tech β https://goo.gle/GoogleCloudTech
#DevOps #Kubernetes #AIAgents #GoogleCloud #ModelContextProtocol #GenerativeAI #KodeKloud #TheAgentFactory #LangGraph
Speakers: Luke Schlangen, Mumshad Mannambeth
Products Mentioned: Google Cloud Run, Google Kubernetes Engine, Google Cloud Build, Google Artifact Registry, Google Secret Manager, Google Model Armor, Gemini Flash, LangGraph, 11 Labs, Prometheus, Grafana, MongoDB, Docker, Podman
KodeKloud AI tutor β https://goo.gle/44Zrg8C
Welcome back to The Agent Factory! In this episode, host Luke Schlangen sits down with Mumshad Mannambeth, the founder of KodeKloud, to explore the future of DevOps, cloud upskilling, and building production-ready AI agents.
Mumshad shares an exclusive, behind-the-scenes look at how KodeKloud built their cutting-edge AI Tutorβa hands-on, AI-assisted learning platform that has doubled student retention! We dive deep into the multi-agent architecture under the hood, showing how they orchestrate five specialized agents using LangGraph and cost-effective Gemini Flash models on Google Cloud. Mumshad also walks us through their implementation of Model Armor on Cloud Run to secure LLM calls against prompt injections and malicious inputs, and how they synchronize AI-generated voice with dynamic React code to generate video lessons on the fly.
We also tackle the shifting landscape of tech interviews, why foundational knowledge in Kubernetes and Linux is more critical than ever, and how to avoid developer burnout in the age of AI. Stick around for a rapid-fire round where Mumshad shares his thoughts on whether certifications are overrated, if Python should be your first language, and how to navigate the overwhelming CNCF landscape.
Chapters:
0:00 - Intro
0:44 - Do we still need Kubernetes?
1:32 - CKA vs CKAD: Admin vs Developer Kubernetes certs
3:35 - Luke's personal certification journey
4:22 - Running Large Language Models (LLMs) on Kubernetes
5:10 - Do we need humans in DevOps if AI can code?
7:27 - The future of coding interviews & system design
11:25 - The AI productivity paradox: Why are we working more?
12:26 - Navigating the massive CNCF landscape
14:53 - How to learn effectively in the AI era
17:05 - Demystifying KodeKloud's AI Tutor
18:31 - Hands-on Demo: Kubernetes Services Lab
24:07 - Facing the Kubernetes Quiz & Live Terminal
30:31 - Interactive debugging with KodeKloud's Live Assist
35:06 - Sneak peek behind the curtain: AI Tutor under the hood
37:03 - Hosting & scaling 30k virtual labs on Google Cloud
39:06 - How the AI assistant doubled student retention
42:52 - Multi-Agent orchestration with LangGraph
46:14 - Syncing 11 Labs audio with custom React UI
48:11 - Leveraging cost-effective Gemini Flash models
51:25 - Personalization and student memory
53:17 - Google Cloud Run deployment & architecture
54:15 - Protecting LLMs from prompt injection with Model Armor
55:55 - Combatting crypto miners in free sandboxes
58:23 - Rapid Fire: Kubernetes, certifications, and developer burnout
1:03:05 - Mumshad's advice on career transitions & lifelong learning
1:04:32 - Outro
More resources:
How to deploy an AI agent to Cloud Run (step by step) β https://goo.gle/4haPa8w
Watch more of The Agent Factory β youtube.com/playlist?list=PLIivdWyY5sqLXR1eSkiM5bE6pFlXC-OSs
π Subscribe to Google Cloud Tech β https://goo.gle/GoogleCloudTech
#DevOps #Kubernetes #AIAgents #GoogleCloud #ModelContextProtocol #GenerativeAI #KodeKloud #TheAgentFactory #LangGraph
Speakers: Luke Schlangen, Mumshad Mannambeth
Products Mentioned: Google Cloud Run, Google Kubernetes Engine, Google Cloud Build, Google Artifact Registry, Google Secret Manager, Google Model Armor, Gemini Flash, LangGraph, 11 Labs, Prometheus, Grafana, MongoDB, Docker, Podman


![How to design a multi-agent system that skips the LLM
Github repo β https://goo.gle/race-condition
Previous episode β https://goo.gle/marathonagent
A thousand AI agents run a marathon, and almost none of them ever call the LLM.
In this multi-agent system deep dive, Casey West breaks down the one architectural decision behind Race Condition: a 1000-agent system built on Googles Agent Development Kit (ADK).
The question every AI engineer is wrestling with: when do you let an LLM decide, and when do you just write the code in a multiagent system? We trace one decision end to end, planning a marathon route, then show how the same idea (skip the LLM where you dont need it) scales to a thousand agents running on deterministic code.
What youll learn:
* When to use an LLM vs deterministic logic
* The before_model_callback trick, keep the agent, skip the model
* Why route planning is deterministic (NP-hard + the Spine & Sprout algorithm)
* How 1,000 autopilot runners make 0 LLM calls
* Where the tokens actually go (the AI decides, the code runs)
* Scaling 1,000 stateless sessions with Redis
Chapters
00:00 - Intro: 1,000 AI agents that dont call the LLM
00:41 - When should an agent use an LLM?
01:02 - [Demo] Planning a marathon route
01:59 - Why Google Maps cant route a marathon
05:08 - Why the LLM Is the wrong tool (NP-hard)
05:40 - The deterministic spine & sprout algorithm
06:58 - Using AI Studio to choose the algorithm
09:00 - The trick: Skip the LLM with a callback
12:26 - before_model_callback β the reveal
17:50 - Autopilot runners: 1,000 agents, 0 LLM calls
21:31 - How many tokens? Where they actually go
23:28 - The second cost: Session state & redis
29:05 - Wrap up
More resources:
Google Agent Development Kit (ADK) β https://goo.gle/3PItVzL
Google ADK Community (Redis session service) β https://goo.gle/4ugzmUw
Agent Runtime β https://goo.gle/4nXDhnX
Google Cloud Memory Store β https://goo.gle/4nXxBtT
Agent2Agent Protocol (A2A) protocol β https://goo.gle/4u5x8HF
Casey West on LinkedIn β https://goo.gle/4dXnsJr
Annie Wang on LinkedIn β https://goo.gle/43GCXAo
Watch more Hands on AI β https://www.youtube.com/playlist?list=PLIivdWyY5sqKnJOvP89yF8t9mWuzMTcbM
π Subscribe to Google Cloud Tech β https://goo.gle/GoogleCloudTech
#AIAgent #GoogleADK #Gemini #MultiAgentSystem #AgenticAI #GoogleCloud
Speakers: Casey West, Annie Wang
Products Mentioned: Google Agent Development Kit, Gemini API, Agent Runtime, Google Cloud Pub/Sub, AlloyDB, Agent2Agent Protocol How to design a multi-agent system that skips the LLM](https://i.ytimg.com/vi/Fzd0BWMH65s/mqdefault.jpg)







