Uploaded August 2026 | Updated September 2026, 2 weeks ago
Don't miss out! Join us at our next KubeCon + CloudNativeCon events in Shanghai, China (8-9 September, 2026) and Salt Lake City, United States (Nov 9–12, 2026). Connect with our current graduated, incubating, and sandbox projects as the community gathers to further the education and advancement of cloud native computing. Learn more at kubecon.io
From Tool Calls To Context Fabric: Building AI-Native Observability for Platform Engineering - Deepak Choudhary & Dheeraj Kapur, NVIDIA
When an engineer asks "Why is this service failing?", the answer is rarely in one place. Metrics live in Prometheus, logs in Loki, anomaly signals elsewhere, and operational knowledge in runbooks. Existing AI approaches—RAG, Text-to-PromQL, single-tool copilots—each work well in one domain but struggle to connect signals across them. We present an AI context fabric for platform observability that links services, metrics, logs, anomalies, and documentation ahead of time using knowledge graphs, metadata caches, and cross-domain relationship modeling, all exposed to agents via Model Context Protocol (MCP). Instead of forcing agents to piece together raw, fragmented data, the fabric hands them grounded, cross-domain context. Our architecture combines a schema cache for instant metadata discovery, a documentation knowledge graph with semantic edges, and a metrics catalog—unified through MCP. The result: faster root cause analysis, sharper reasoning, and significantly lower token cost.
Don't miss out! Join us at our next KubeCon + CloudNativeCon events in Shanghai, China (8-9 September, 2026) and Salt Lake City, United States (Nov 9–12, 2026). Connect with our current graduated, incubating, and sandbox projects as the community gathers to further the education and advancement of cloud native computing. Learn more at kubecon.io
From Tool Calls To Context Fabric: Building AI-Native Observability for Platform Engineering - Deepak Choudhary & Dheeraj Kapur, NVIDIA
When an engineer asks "Why is this service failing?", the answer is rarely in one place. Metrics live in Prometheus, logs in Loki, anomaly signals elsewhere, and operational knowledge in runbooks. Existing AI approaches—RAG, Text-to-PromQL, single-tool copilots—each work well in one domain but struggle to connect signals across them. We present an AI context fabric for platform observability that links services, metrics, logs, anomalies, and documentation ahead of time using knowledge graphs, metadata caches, and cross-domain relationship modeling, all exposed to agents via Model Context Protocol (MCP). Instead of forcing agents to piece together raw, fragmented data, the fabric hands them grounded, cross-domain context. Our architecture combines a schema cache for instant metadata discovery, a documentation knowledge graph with semantic edges, and a metrics catalog—unified through MCP. The result: faster root cause analysis, sharper reasoning, and significantly lower token cost.










