Uploaded December 2025 | Updated September 2026, 2 weeks ago
Nitin Kanukolanu, Applied AI Engineer at Redis, focused on semantic caching during AI Dev 25 x NYC.
As LLMs drive the next wave of applications, compute bottlenecks are becoming a critical challenge. Semantic caching has emerged as a practical strategy to cut costs, reduce latency, and improve consistency in agentic systems. This session covered real-world use cases, explained how semantic caching works, highlighted what to measure in production, and shared strategies for boosting both performance and quality at scale.
Take our course with Redis on this topic: deeplearning.ai/short-courses/semantic-caching-for-ai-agents
--------------
Join us at AI Dev 26 x San Francisco! Tickets: ai-dev.deeplearning.ai
Nitin Kanukolanu, Applied AI Engineer at Redis, focused on semantic caching during AI Dev 25 x NYC.
As LLMs drive the next wave of applications, compute bottlenecks are becoming a critical challenge. Semantic caching has emerged as a practical strategy to cut costs, reduce latency, and improve consistency in agentic systems. This session covered real-world use cases, explained how semantic caching works, highlighted what to measure in production, and shared strategies for boosting both performance and quality at scale.
Take our course with Redis on this topic: deeplearning.ai/short-courses/semantic-caching-for-ai-agents
--------------
Join us at AI Dev 26 x San Francisco! Tickets: ai-dev.deeplearning.ai










