Uploaded April 2026 | Updated September 2026, 2 weeks ago
In this podcast, we're talking all about AI observability: why is it difficult to observe AI? How can you debug AI when something goes wrong? What are evaluations, and how do you use them to improve your confidence in the quality of your AI app?
The Context Window is a regular video podcast about engineering AI at Grafana Labs. The hosts today are Senior Developer Advocates Tiffany Jernigan and Nicole van der Hoeven, and the guests are Staff Software Engineer Alexander Sniffin and Senior Software Engineer Jack Gordley.
TIMESTAMPS:
00:00:00 Introductions
00:02:00 The last month in AI news
00:11:09 What is AI Observability in Grafana Cloud?
00:19:20 What is an evaluation?
00:21:05 The origin of AI Observability
00:25:09 Demo: Setting up AI Observability and evaluators
00:32:04 What is LLM as judge?
00:38:43 Demo: AI Observability Analytics
00:40:52 AI O11y is based on OpenTelemetry
00:42:17 Demo: Instrumenting a local coding agent
00:47:18 Potential future agentic use cases
00:52:00 Evaluators that catch the most bugs
01:02:23 Demo: System prompt analysis
01:05:11 Guess the prompt
---
Links/resources:
NEWS
Opus 4.7 is out: anthropic.com/news/claude-opus-4-7
Gemma 4: https://deepmind.google/models/gemma/gemma-4/
Qwen 3.6: qwen.ai/blog?id=qwen3.6
o11y-bench announcement: grafana.com/blog/o11y-bench-open-benchmark-for-observability-agents
Opt out of GitHub Copilot training: https://github.blog/news-insights/company-news/updates-to-github-copilot-interaction-data-usage-policy/
All the GrafanaCon announcements: https://gra.fan/gcon26
OpenCode: opencode.ai
(docs) Online evals on Grafana Cloud: grafana.com/docs/grafana-cloud/machine-learning/ai-observability/introduction/#online-evaluation
(docs) OpenTelemetry integration with AI Observability: grafana.com/docs/grafana-cloud/machine-learning/ai-observability/introduction/#opentelemetry-integration
The Sigil SDK: github.com/grafana/sigil-sdk
(blog) Anthropic: Emotion concepts and their function in a large language model: anthropic.com/research/emotion-concepts-function
OpenClaw: openclaw.ai
Learn about Grafana Assistant: grafana.com/docs/grafana-cloud/machine-learning/assistant
Check out our AI blog, Context Horizon: https://gra.fan/ch
Learn how we handle your privacy and security for Grafana Assistant: grafana.com/docs/grafana-cloud/machine-learning/assistant/privacy-and-security
Check out the pricing page - Assistant is included in the free tier too!: grafana.com/pricing
Get started with the Grafana Cloud forever-free tier: grafana.com/g/cloud
Have a question? Ask Grot, your AI helper: grafana.com/grot
Reach out in our community forums: https://gra.fan/communityyf
---
Thanks for watching!
👍 Was this video helpful? Like and subscribe to our channel for more videos.
Connect with Grafana Labs:
X: (twitter.com/grafana)
LinkedIn: (linkedin.com/company/grafana-labs/)
Facebook: (facebook.com/grafana)
#Grafana #Observability #assistant #ai #actuallyusefulai
In this podcast, we're talking all about AI observability: why is it difficult to observe AI? How can you debug AI when something goes wrong? What are evaluations, and how do you use them to improve your confidence in the quality of your AI app?
The Context Window is a regular video podcast about engineering AI at Grafana Labs. The hosts today are Senior Developer Advocates Tiffany Jernigan and Nicole van der Hoeven, and the guests are Staff Software Engineer Alexander Sniffin and Senior Software Engineer Jack Gordley.
TIMESTAMPS:
00:00:00 Introductions
00:02:00 The last month in AI news
00:11:09 What is AI Observability in Grafana Cloud?
00:19:20 What is an evaluation?
00:21:05 The origin of AI Observability
00:25:09 Demo: Setting up AI Observability and evaluators
00:32:04 What is LLM as judge?
00:38:43 Demo: AI Observability Analytics
00:40:52 AI O11y is based on OpenTelemetry
00:42:17 Demo: Instrumenting a local coding agent
00:47:18 Potential future agentic use cases
00:52:00 Evaluators that catch the most bugs
01:02:23 Demo: System prompt analysis
01:05:11 Guess the prompt
---
Links/resources:
NEWS
Opus 4.7 is out: anthropic.com/news/claude-opus-4-7
Gemma 4: https://deepmind.google/models/gemma/gemma-4/
Qwen 3.6: qwen.ai/blog?id=qwen3.6
o11y-bench announcement: grafana.com/blog/o11y-bench-open-benchmark-for-observability-agents
Opt out of GitHub Copilot training: https://github.blog/news-insights/company-news/updates-to-github-copilot-interaction-data-usage-policy/
All the GrafanaCon announcements: https://gra.fan/gcon26
OpenCode: opencode.ai
(docs) Online evals on Grafana Cloud: grafana.com/docs/grafana-cloud/machine-learning/ai-observability/introduction/#online-evaluation
(docs) OpenTelemetry integration with AI Observability: grafana.com/docs/grafana-cloud/machine-learning/ai-observability/introduction/#opentelemetry-integration
The Sigil SDK: github.com/grafana/sigil-sdk
(blog) Anthropic: Emotion concepts and their function in a large language model: anthropic.com/research/emotion-concepts-function
OpenClaw: openclaw.ai
Learn about Grafana Assistant: grafana.com/docs/grafana-cloud/machine-learning/assistant
Check out our AI blog, Context Horizon: https://gra.fan/ch
Learn how we handle your privacy and security for Grafana Assistant: grafana.com/docs/grafana-cloud/machine-learning/assistant/privacy-and-security
Check out the pricing page - Assistant is included in the free tier too!: grafana.com/pricing
Get started with the Grafana Cloud forever-free tier: grafana.com/g/cloud
Have a question? Ask Grot, your AI helper: grafana.com/grot
Reach out in our community forums: https://gra.fan/communityyf
---
Thanks for watching!
👍 Was this video helpful? Like and subscribe to our channel for more videos.
Connect with Grafana Labs:
X: (twitter.com/grafana)
LinkedIn: (linkedin.com/company/grafana-labs/)
Facebook: (facebook.com/grafana)
#Grafana #Observability #assistant #ai #actuallyusefulai







