Uploaded July 2026 | Updated September 2026, 1 week ago
Code Materials: docs.google.com/document/d/1wMPQL2NJTzT70GLBVYr3hKrObCmYrwhwvTgoEb0PLWk/edit?tab=t.0
Join our Advanced Route-Production AI And LLM Engineering-Frontier AI From Research To Production Bootcamp starting from July 19th 2026
krishnaik.in/liveclass2/Advanded-Route?id=12
Timestamps
00:00:04 — Stream Introduction & Agenda Setup
00:09:05 — Advanced Production AI Boot Camp Announcement
Main sections
00:11:19 — Core Agentic RAG Codebase Review
00:17:53 — LangGraph & FlashRank Re-ranking Mechanics
00:23:02 — Qdrant Vector DB Configuration
00:25:58 — Pydantic Logfire Observability & Tracing
00:33:33 — Off-Topic Input Problem & Security Goals
00:35:07 — LLM Security: Guardrails & Gateways Overview
00:38:20 — Nemo Guardrails & Colang Rule Architecture
00:41:57 — Portkey Gateway Load Balancing & Virtual Keys
01:17:41 — Implementing Nemo Guardrails in Fast API Backend
01:36:26 — Configuring Portkey Multi-Model Fallback Routine
01:54:59 — Integrating Gateway Caching into Query Module
02:29:29 — Live Demo: Guardrails Blocking Off-Topic Requests
02:35:46 — Introduction to LLM Evaluation Pipelines
02:41:53 — Walkthrough of Evaluation Streaming Dashboard
02:53:55 — Core Eval Terms: Ground Truth & Goldens
03:04:16 — Student-Teacher Examination Framework Analogy
03:21:52 — Scalability: Automated LLM-as-a-Judge Reasoning
03:36:25 — Integrating Evals into CI/CD Pipelines
03:44:44 — RAG Metrics Deep Dive: Faithfulness
03:52:56 — RAG Metrics Deep Dive: Answer Relevancy
04:00:02 — RAG Metrics Deep Dive: Context Recall
04:21:55 — Coding Ragas & Deepeval Into the Module
05:06:13 — Moving Beyond Local: Enterprise Cloud Architecture
05:20:40 — Transitioning to Jina AI Embedding Models
05:33:33 — Persisting LangGraph State to Neon PostGreSQL
05:37:23 — Adding Serverless Semantic Cache with Upstash Redis
05:44:26 — Optimizing Dockerfiles Using UV Package Manager
06:21:52 — AWS Console Access & IAM Configuration
06:24:15 — AWS ECS Fargate, VPC, & Load Balancer Setup
06:41:15 — Structuring GitHub Actions for Automatic ECR Push
06:49:45 — Managing Environment Secrets in AWS Secret Manager
07:32:46 — Introduction to Multi-Modal Parsing Challenges
07:36:58 — Three Multi-Modal Document Parsing Paradigms
07:43:58 — Analyzing ColPali Multi-Vector Projections
07:50:51 — Code Execution: ColQwen 2.5 on Cloud GPU
08:04:09 — Single-Stage Models: Neotron Parse & Unlimited OCR
08:11:07 — Dual-Stage Pipelines: PP-DocLayout & GLM-OCR
08:24:30 — Live Bounding Box Comparison Using Neotron Parse
08:40:35 — Serving the Unlimited OCR Model on an L4 GPU
08:52:41 — Benchmark Analysis & Final Output Review
Code Materials: docs.google.com/document/d/1wMPQL2NJTzT70GLBVYr3hKrObCmYrwhwvTgoEb0PLWk/edit?tab=t.0
Join our Advanced Route-Production AI And LLM Engineering-Frontier AI From Research To Production Bootcamp starting from July 19th 2026
krishnaik.in/liveclass2/Advanded-Route?id=12
Timestamps
00:00:04 — Stream Introduction & Agenda Setup
00:09:05 — Advanced Production AI Boot Camp Announcement
Main sections
00:11:19 — Core Agentic RAG Codebase Review
00:17:53 — LangGraph & FlashRank Re-ranking Mechanics
00:23:02 — Qdrant Vector DB Configuration
00:25:58 — Pydantic Logfire Observability & Tracing
00:33:33 — Off-Topic Input Problem & Security Goals
00:35:07 — LLM Security: Guardrails & Gateways Overview
00:38:20 — Nemo Guardrails & Colang Rule Architecture
00:41:57 — Portkey Gateway Load Balancing & Virtual Keys
01:17:41 — Implementing Nemo Guardrails in Fast API Backend
01:36:26 — Configuring Portkey Multi-Model Fallback Routine
01:54:59 — Integrating Gateway Caching into Query Module
02:29:29 — Live Demo: Guardrails Blocking Off-Topic Requests
02:35:46 — Introduction to LLM Evaluation Pipelines
02:41:53 — Walkthrough of Evaluation Streaming Dashboard
02:53:55 — Core Eval Terms: Ground Truth & Goldens
03:04:16 — Student-Teacher Examination Framework Analogy
03:21:52 — Scalability: Automated LLM-as-a-Judge Reasoning
03:36:25 — Integrating Evals into CI/CD Pipelines
03:44:44 — RAG Metrics Deep Dive: Faithfulness
03:52:56 — RAG Metrics Deep Dive: Answer Relevancy
04:00:02 — RAG Metrics Deep Dive: Context Recall
04:21:55 — Coding Ragas & Deepeval Into the Module
05:06:13 — Moving Beyond Local: Enterprise Cloud Architecture
05:20:40 — Transitioning to Jina AI Embedding Models
05:33:33 — Persisting LangGraph State to Neon PostGreSQL
05:37:23 — Adding Serverless Semantic Cache with Upstash Redis
05:44:26 — Optimizing Dockerfiles Using UV Package Manager
06:21:52 — AWS Console Access & IAM Configuration
06:24:15 — AWS ECS Fargate, VPC, & Load Balancer Setup
06:41:15 — Structuring GitHub Actions for Automatic ECR Push
06:49:45 — Managing Environment Secrets in AWS Secret Manager
07:32:46 — Introduction to Multi-Modal Parsing Challenges
07:36:58 — Three Multi-Modal Document Parsing Paradigms
07:43:58 — Analyzing ColPali Multi-Vector Projections
07:50:51 — Code Execution: ColQwen 2.5 on Cloud GPU
08:04:09 — Single-Stage Models: Neotron Parse & Unlimited OCR
08:11:07 — Dual-Stage Pipelines: PP-DocLayout & GLM-OCR
08:24:30 — Live Bounding Box Comparison Using Neotron Parse
08:40:35 — Serving the Unlimited OCR Model on an L4 GPU
08:52:41 — Benchmark Analysis & Final Output Review










