Session Evaluation On An AI Tutor Chatbot @arizeai
Session Evaluation On An AI Tutor Chatbot  @arizeai
Uploaded August 2025 | Updated September 2026, 2 weeks ago
This tutorial shows you how to run session-level evaluations on conversations with an AI tutor using Arize. Session-level evaluations provide a holistic view of entire interactions, enabling you to assess broader patterns and answer high-level questions about user experience and system performance.

Resources:
Notebook: arize.com/docs/ax/cookbooks/evaluation/session-level-evaluations-for-an-ai-tutor
Make a free Arize account: arize.com/sign-up
Arize Community Slack: arize.com/community
Arize Docs: arize.com/docs/ax

See our explainer on sessions, traces, spans in LLM observability: arize.com/blog/llm-observability-for-ai-agents-and-applications

This tutorial covers how to:
Agent tracing for multi-turn AI tutor conversations
Aggregate spans into structured sessions with truncation support
Evaluate sessions across multiple dimensions (Correctness, Goal Completion, Frustration)
Format evaluation outputs to match Arize's schema
Log results back to Arize for monitoring and analysis

Background:
A session captures the entire dialogue between a user and your app—the full journey, not just isolated spans. It reflects how people actually use your product.

Evaluating at the session level unlocks a broader view. You can measure things like frustration, context retention, and whether the user’s goal was achieved—insights you simply can’t get from span-level evaluations alone.

We built an AI tutor and turned its traces into complete sessions. Then, using Arize, we ran session-level evaluations that assess the whole conversation, not just snippets.
Session Evaluation On An AI Tutor ChatbotMeet Quiet-STaR and Minimo: Understanding Self Discovered Reasoning EnvironmentsRise of the Agent Engineer: Booking.coms Chana RossHow LG Uplus Built an AI Contact Center Serving 30 Million Customers | Arize Observe 202The Math Humans Physically Cant DoKeynote | The Future of AI Agents | Arize Observe 2026AI Agent Got the Right Answer the Wrong Way | Rise of the AI Engineer | Michael Grinich, WorkOSUltimate OpenTelemetry Guide for Tracing AI ApplicationsOne AI Question - how should you use AI in the legal field, with Typer NiederwerderSynthetic Data Generation for LLM Evaluators and AgentsUsing Code Evaluators in PhoenixFrameworks for Building Agents Panel
Arize AI |

Session Evaluation On An AI Tutor Chatbot

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER