Arize AI
Building Agentic RAG Systems
updated
github.com/Arize-ai/phoenix/tree/main/tutorials/ai_evals_course
Covers homework 1 (System Prompt Engineering with Phoenix) and homework 2 (Recipe Bot Error Analysis).
First in a series of five homeworks by the Arize AI team for the AI Evals For Engineers & PMs Maven course by Parlance Labs.
Learn more about Phoenix: arize.com/docs/phoenix
github.com/Arize-ai/phoenix/tree/main/tutorials/ai_evals_course
This alternate third extra credit homework featuring Arize Phoenix is part of a series developed by the Arize AI team for the AI Evals For Engineers & PMs course by Parlance Labs.
Build an LLM-as-a-Judge in Homework 3 of our AI Evaluations module using Phoenix (open source) to assess a recipe bot’s outputs against dietary restrictions (vegan, Whole30, diabetic-friendly, gluten-free, paleo). We spin up Phoenix, trace OpenAI calls, create & compare ground truths, run experiments (accuracy, TP/TN, confusion matrix), analyze failure cases (e.g., honey/whole-wheat/peas), and iteratively refine the evaluator with definitions + few-shot examples—ending with strong test accuracy and end-to-end trace evaluations.
github.com/Arize-ai/phoenix/tree/main/tutorials/ai_evals_course
github.com/Arize-ai/phoenix/tree/main/tutorials/ai_evals_course
Reinforcement Learning Gyms — Ankit Jasuja (Turing) walked through how RL gyms are reshaping the training of AI agents, offering safer, more dynamic environments for real-world readiness.
Learn more about the difference between Cursor and Claude Code -- and how to be come a power user of the latter: arize.com/blog/claude-code-vs-cursor-a-power-users-playbook
Learn more about Arize's tool for tracing Claude Code: arize.com/blog/claude-code-observability-and-tracing-introducing-dev-agent-lens
More on the difference between Claude Code and Cursor & power user tips: arize.com/blog/claude-code-vs-cursor-a-power-users-playbook
⏱️ Timestamps
00:00 Intro & talk overview
00:13 What is Claude Code? (agent loop, tool calls) + Cursor vs. Claude Code
01:30 Plan-first workflow (iterate plan → save → clear context → execute)
03:08 Speeding up the loop (tool use, fast tests/lint/build, rules, TDD, helper scripts)
04:15 Advanced tips (git worktrees, automate plan/review, MCP + Playwright, Whisper)
05:04 Recap & resources
💡 Key takeaways
- Iterate on the plan, not the code.
- Manage context windows; start fresh for execution.
- Tighten the dev loop so the agent can verify changes fast.
- Use git (rollbacks, worktrees) to move safely and in parallel.
Connect with Alec:
linkedin.com/in/alec-swanson
Join the Arize AI community for events, office hours, and tutorials: arize.com/community
Notebook: arize.com/docs/phoenix/cookbook/tracing-and-annotations/generating-synthetic-datasets-using-llms
Arize Community Slack: arize.com/community
Make a free Phoenix account: app.phoenix.arize.com
Arize Phoenix docs: arize.com/docs/phoenix
agents – then perfect them in production. Arize AX feautres Alyx, a built-in co-pilot with agent mode.
Request a demo or lunch-and-learn workshop for your team: arize.com/request-a-demo
Free signup: app.arize.com/auth/join
Key Arize AX features include:
GenAI Tracing
Unlock instant, end-to-end visibility into your AI applications and agents with seamless OTEL instrumentation for both development and production environments. Automate observability with built-in support for leading AI frameworks, eliminating complex setup. Gain granular insights with detailed tracing of prompts, variables, tool calls, agents, and assistants—empowering you to pinpoint issues faster.
LLM Evaluation & LLM as a Judge
Evaluation-driven development assesses your AI’s performance with automated evaluations at every stage. Trigger checks as you push code, with LLM- as-a-Judge powered insights explaining issues and code-based tests catching errors early. Run evaluations at scale in production to ensure reliable, high-performing AI systems
Real-Time Monitoring
Real-time visibility with AI-powered monitoring that detects anomalies, forecasts failures, and automates root cause analysis. Stay ahead with auto-thresholding, integrated alerts via PagerDuty and Slack, and fully customizable metrics for complex systems. Scaled monitoring, analytical dashboards, and AI-assisted metric creation keep your performance optimized and reliable.
Annotations
Combine human expertise with automated workflows to generate high-quality labels and annotations. Quickly identify edge cases, refine datasets, and enhance your AI applications with smarter, more reliable data inputs.
Prompt & Eval IDE
Prompt playground and eval hub to design, test, and evaluate prompts in a unified workspace built for iteration. Leverage built-in evaluation tools and feedback loops to optimize functionality, address edge cases, and continuously improve your AI applications across development and production.
Speakers showcase workflows for tracing, LLM-as-a-judge evaluations, and feedback loops, with live demos highlighting how to detect, debug, and improve AI systems in both development and production.
Wayfair shares real-world successes using these tools for catalog enrichment, supply chain optimization, and multi-agent frameworks, while Google Cloud outlines Gemini’s capabilities and responsible AI approach.
Timestamps
0:00 – Intro & Agenda – Jason (CEO, Arize) opens event, outlines focus on AI observability & evaluation, agenda with Google Cloud + Wayfair.
2:07 – What is Evaluation? – LLM-as-judge, moving beyond CSVs, 3 pillars: tracing, evals, experimentation.
4:48 – Arize Evolution – From observability to evals & prompt IDE, intro to “Alex” AI assistant.
6:21 – Platform Architecture – Dev ↔ Prod feedback loop, tracing, eval spectrum (LLM-judge, code checks, human annotations).
12:24 – Eval Workflows – Dev playground (CSV) vs. online production evals.
14:10 – Live Demo – Tracing, span-level evals, fixing hallucinations, prompt changes, Prompt Hub, AI Assist.
26:06 – Google Cloud (Cindy) – Responsible AI, Gemini’s long context, reasoning, multimodality, low hallucination.
39:09 – Customer Stories – Wayfair catalog tagging, Viral Nation brand safety.
42:25 – Keys to Agents – Must be autonomous & optimized; tracing, evals, feedback loops.
43:03 – Dylan: Feedback-Driven Dev – Agent challenges, eval points (router, tools, decision path, sessions), English feedback, Prompt Learning SDK.
1:01:33 – Wayfair Panel – LLMOps early, agent examples, multi-agent use, LangChain, Google A2A, MCP security, LLM-as-a-jury.
1:20:00 – Closing – Community & customer support importance, networking.
Relevant links
Arize Enterprise Platform – arize.com
Arize Phoenix (Open-Source) – github.com/Arize-ai/phoenix
Google Cloud Vertex AI – cloud.google.com/vertex-ai
Model Context Protocol (MCP) – modelcontextprotocol.io
Notebook: arize.com/docs/ax/cookbooks/evaluation/trace-level-evaluations-for-a-recommendation-agent
Arize AX Docs: arize.com/docs/ax/evaluate/trace-level-evaluations
Make a free Arize AX account: arize.com/sign-up
Community: arize.com/community
Handshake, renowned for its innovative career and job matching platform, has developed an orchestration framework enabling product and engineering teams to rapidly spin up tailored, company-specific LLM use cases. This adaptable infrastructure provides plug-and-play integration of diverse new models and frameworks, accelerating deployment and enhancing agility. Complementing this, Arize AI provides full-stack observability, empowering teams to quickly iterate, validate model outputs, proactively manage hallucination risks, and maintain consistency across multiple frameworks.
Together, this approach provides a model for organizations to deploy scalable, reliable LLM pipelines capable of adapting swiftly to new innovations.
This session covers:
→ How To Build Adaptable Technical Infrastructure: Strategies for building flexible, future-proof LLM stacks for rapid deployment.
→ Scalable Orchestration: Best practices from Handshake's approach to minimize code and accelerate deployment of custom use cases.
→ Multi-Provider Integration: Techniques to efficiently manage multiple integrations with Claude, Gemini, OpenAI, and DeepSeek through structured YAML outputs.
→ Prompt Engineering & Evaluation: Collaborative optimization strategies using Arize's real-time tracing and evaluation capabilities.
Resources:
Notebook: arize.com/docs/ax/cookbooks/evaluation/session-level-evaluations-for-an-ai-tutor
Make a free Arize account: arize.com/sign-up
Arize Community Slack: arize.com/community
Arize Docs: arize.com/docs/ax
See our explainer on sessions, traces, spans in LLM observability: arize.com/blog/llm-observability-for-ai-agents-and-applications
This tutorial covers how to:
Agent tracing for multi-turn AI tutor conversations
Aggregate spans into structured sessions with truncation support
Evaluate sessions across multiple dimensions (Correctness, Goal Completion, Frustration)
Format evaluation outputs to match Arize's schema
Log results back to Arize for monitoring and analysis
Background:
A session captures the entire dialogue between a user and your app—the full journey, not just isolated spans. It reflects how people actually use your product.
Evaluating at the session level unlocks a broader view. You can measure things like frustration, context retention, and whether the user’s goal was achieved—insights you simply can’t get from span-level evaluations alone.
We built an AI tutor and turned its traces into complete sessions. Then, using Arize, we ran session-level evaluations that assess the whole conversation, not just snippets.
The paper’s lead author John Kirchenbauer walks through the research and its implications.
Key takeaways on the research: arize.com/blog/a-watermark-for-large-language-models
Read the paper: arxiv.org/pdf/2301.10226
Check out the repo: github.com/jwkirchenbauer/lm-watermarking
Notebook: arize.com/docs/phoenix/cookbook/evaluation/creating-a-custom-llm-evaluator-with-a-benchmark-dataset
Arize Community Slack: arize.com/community
Make a free Phoenix account: app.phoenix.arize.com
Arize Phoenix docs: arize.com/docs/phoenix
Custom Annotations Example: youtu.be/JK2JQUqpcqM?si=HHh4qPuVJMstPnp8
More about LLM as a Judge: arize.com/llm-as-a-judge
A custom annotation UI makes it easy to collect structured human feedback on traces directly in Phoenix, enabling faster iteration and improvement of your LLM systems. By establishing this feedback loop, you can effectively monitor and enhance your application’s performance.
Find the notebook here: arize.com/docs/phoenix/cookbook/tracing-and-annotations/using-human-annotations-for-eval-driven-development
Join our community slack: arize.com/community
Get started with Phoenix for free: app.arize.com/auth/phoenix/signup
More on Phoenix annotations: arize.com/docs/phoenix/tracing/features-tracing/how-to-annotate-traces
🎮 What if measuring AGI could be more like playing a game?
Greg Camerad, President of the ARC Prize Foundation, breaks down how their groundbreaking third generation benchmark — ARC AGI 3 — will transform how we measure general intelligence.
🔍 Highlights:
-Why classic static benchmarks fall short at capturing real generalization
-How ARC’s new approach uses interactive reasoning benchmarks with 100+ novel games
-The surprising philosophy of measuring skill acquisition efficiency (Francois Chollet’s elegant idea)
-What made the Atari benchmark phase flawed — and how this is different
-The roadmap for launching this new benchmark (with APIs & contests!)
-Why these tests could finally tell us if an AI is truly learning on the fly
🚀 Whether you're an ML researcher, an AI skeptic, or just fascinated by the path to AGI, this talk is a must-watch.
⏱ Timestamps
00:03 - Intro: Why measuring general intelligence is about to get more fun
0:10 - Claude & Gemini playing Pokemon: Why it’s NOT AGI yet
0:56 - Who is Greg Camerad & ARC Prize: A northstar for open AGI
1:35 - ARC’s unique benchmark philosophy: Humans as the target, because humans = only known proof of general intelligence
2:23 - Defining intelligence: McCarthy & Chollet’s views
3:48 - Skill acquisition efficiency: Learn new things & show it — that’s the real test
4:43 - How ARC AGI 1 & 2 worked: Learning novel transformations, 1000+ tiny skills
5:55 - Why static benchmarks can’t test long-term, open-ended learning
6:54 - Rich Sutton’s “era of experience”: Why agents need to collect their own data
7:52 - ARC’s solution: Interactive reasoning benchmarks — long horizon, multi-turn
8:47 - Why games are the perfect medium for testing intelligence
9:46 - The problems with Atari benchmarks: Dense rewards, unlimited sampling, overfit developers
10:47 - The dream benchmark: 100 games, novel to AI & developers, sparse rewards, no instructions
12:36 - Ensuring zero cultural knowledge: Only core knowledge priors (objects, geometry, agency)
13:50 - Humans still uniquely good at interrogating failures & improving — can AI close this gap?
14:19 - Measuring skill acquisition efficiency in actions, not static scores
14:53 - Timeline: Preview games & API coming July 17, full 100-game launch Q1 next year
15:43 - How you can help: Philanthropy, building agents, or designing novel games
16:52 - Rules for games: Fun for humans, no cultural symbols, no overlapping skills
17:55 - Deep question: But what about human dopamine loops — why do agents even act?
18:23 - Setting rewards & alignment: A question for upstream model builders
18:55 - How to define clear rewards even in chaotic or competitive games
#AGI, #artificialintelligence, #ARCPrize
What really is an AI agent? Are multi-agent systems worth it — or overhyped? Will one framework rule them all, or will developers stay close to the metal?
🔥 In this super candid, spicy panel, top founders from Crew AI, LlamaIndex, Mastra and Leta dive into the evolving world of agent frameworks. They debate:
What defines an agent vs basic LLM plumbing?
-Are most “agent frameworks” just glorified abstractions?
-Where do protocols like MCP and ADA actually fit (or fail)?
-Is multi-agent orchestration practical, or a demo trap?
-Why agent evaluation (evals) might be overrated — and when it matters.
-What’s next: perpetual agents, no-code builders, dependable systems.
-Perfect for developers, architects, or anyone deep into the future of agentic AI.
⏱ Timestamps
00:02 - Intro: Top agent framework founders on stage (Crew AI, LlamaIndex, Mastra, Leta)
0:26 - The age-old question: What IS an agent? Each founder defines it
1:46 - From LLMs with agency, to closed-loop systems, to control & memory
2:48 - How the term "agent" evolved from 2015 robotics RL to modern LLMs
5:16 - What’s an agent framework? Types in the ecosystem: workflows, orchestration, SDK vs service
6:48 - Why many frameworks are just middleware — or libraries vs servers
8:06 - Spicy take: “LangChainJS sucked — that’s why we built Mastra.”
9:50 - On abstractions: conventions, feature overload, and moving fast
12:50 - Will there be one framework to rule them all, or language-specific stacks?
14:55 - Library vs service trade-offs; easy to swap libraries, harder with hosted state
17:48 - Companies move from playful prototyping to secure production — frameworks commoditize
19:40 - Multi-agent debate: prompt engineering, decomposition, microservices parallels
23:00 - When multi-agent is roleplay vs serious system design
24:35 - Most production use cases today still resemble controlled pipelines, not unconstrained agent swarms
25:40 - MCP vs ADA: why MCP solves tool calling, but ADA is a solution looking for a problem
28:00 - The politics of protocols & vendor lock-in, plus open standards
30:50 - On evals: are they a moat or hype? When customers actually start caring
34:35 - Parallel to research: benchmarking vs method building — vibes first, evals later
36:45 - The next year: stable, perpetual agents; non-techs building agents
#AIagents, #agentframeworks, #CrewAI, #LlamaIndex, #Mastra, #Letta, #MCP
Whether you’re an AI PM, platform engineer, or builder launching LLM features, Aman covers:
The three archetypes of AI PMs (AI Product, AI Platform, AI Powered)
When to think about evals—and how to write one
How evals align with business metrics and human labels
Best practices for turning evals into product requirements
Why close collaboration between engineering and product matters more than ever
⏱️ Chapters:
0:00 Intro
0:28 Who This Talk Is For (AI PMs & Engineers)
1:16 Three AI PM Archetypes
2:28 Workflow 1: Vibe Coding (No Eval)
3:54 Why Evals Matter
4:40 What an Eval Looks Like
6:01 Where Eval Fits in the Workflow
7:00 Who Writes the Eval?
8:28 Evals as AI Product Requirements
9:30 Thinking About Evals as a Funnel
10:57 Troubleshooting: When Evals & Labels Disagree
12:01 Aligning Evals With Business Metrics
12:41 Workflow 3: Evals in Production
14:48 FAQs on Evals Best Practices
15:47 Recap: Workflows, Roles & Questions to Ask
16:28 Looking Ahead: PMs & Engineers Working Together
🎉 Join upcoming Arize events:
arize.com/community
Learn more about AI product management:
arize.com/ai-product-manager
Build your own eval:
arize.com/docs/ax/evaluate/llm-as-a-judge/custom-evaluators
In this Arize Observe 2025 talk, Manjeet Singh -- Sr. Director of Product Management at Salesforce (Agentforce) -- breaks down the agent lifecycle across build, test, and production, offering practical insights for:
Testing agents with evals and simulation frameworks
Tracing multi-agent workflows across handoffs
Using observability to diagnose failure points in retrieval, generation, or orchestration
The emerging role of supervisor agents as control planes for digital labor
🧪 Learn how agent testing and observability parallels DevOps in complexity—and what new tools are needed to ensure reliability at scale.
👉 Join Arize community events: arize.com/community
In this fast-paced session, Aditi walks through live code and shows how to:
Choose and customize models like GPT-4o
Combine Azure AI Search with Foundry Agent Services
Run hybrid and semantic search with RRF for accurate results
Trace agents using Phoenix, an open-source observability library
Chapters:
0:00 Introduction by Aditi Maheshwari
0:29 Why Agentic AI is the next wave
1:01 Azure AI Foundry overview
1:47 Choosing the right model (OpenAI, Meta, etc.)
2:20 Customization with Azure AI Search and Blob Storage
3:45 Vector search vs hybrid search vs semantic re-ranking
5:12 Integrating Phoenix for agent observability
6:08 Tool definitions and multi-agent setup
7:10 Defining agents and prompt-based behavior
8:00 Live agent demo with Phoenix traces
8:40 Comparing Arize and Phoenix for observability
9:13 Closing and additional resources
🎯 Use cases covered: Enterprise RAG, knowledge management, and evaluation of agent workflows
📦 Built-in support for custom tools, interoperability, and secure deployment
📘 Learn more about Arize Phoenix: docs.arize.com/phoenix
🎟️ Attend upcoming Arize AI events: arize.com/community
Drawing on his experience working on AutoML at Amazon and teaching AI at Stanford, Yujian explores real-world agent design patterns, common use cases, and architectural tradeoffs. Whether you're a builder looking to ship agentic systems or just exploring the hype, this talk offers an accessible framework for understanding where agents are today—and where they're going.
Chapters:
0:00 - Intro: Meet Yujian Tang
0:54 - What is an AI agent?
1:39 - The spectrum of autonomy in agents
2:44 - Human-centric agents: tools for task support
3:16 - Human-in-the-loop agents: collaborative workflows
4:46 - Fully autonomous agents: goals, tools, decisions
5:24 - Basic agent architecture
5:45 - Orchestrator agents and sub-agent hierarchies
6:19 - Agentic swarms: recursive, cooperative systems
6:40 - Cautions with recursion, complexity, and cost
7:01 - Recap and contact info
Stay current on events:
arize.com/community
Chapters:
0:00 - Intro
0:06 - Rahul Todkar on agent use cases at Tripadvisor
1:40 - Anu Trivedi on Oracle Health’s healthcare AI strategy
4:36 - Models vs. applications: what should do the heavy lifting?
6:00 - Domain-specific challenges in healthcare
7:07 - Vertical AI and patient-level customization
8:46 - Agent studios for healthcare personalization
9:36 - Reimagining UX in travel with LLMs
11:00 - Conversational design and multimodal inputs
13:05 - Oracle Health: customer outcomes over tech fads
15:04 - AI-powered digital caregivers
17:00 - Tooling and framework decisions
18:00 - LangGraph, Arize, and custom infra choices
21:05 - Insights on agent production success
22:26 - Defining agent inputs, outputs, and integrations
23:46 - Measuring and evaluating agent quality
24:40 - When not to use an agent
Explore Arize AI community events and stay up-to-date on the future of AI: arize.com/community
🧠 This talk covers:
Evolving workload types: multimodal data, agentic inference, post-training with RL
Real-world examples from Uber, Pinterest, Roblox, and DeepSeek
Why traditional stacks fail, and how Ray enables new patterns
What’s next for distributed AI systems
⏱️ Chapters
0:00 – Introduction & Background at UC Berkeley
0:51 – Why Ray: Scaling Was the Research Bottleneck
1:45 – Early Adoption: Uber, Pinterest, Ant Group
2:30 – GenAI Inflection Point for Distributed Compute
3:25 – Workload Shift #1: GPU + Multimodal Data Processing
5:10 – From SQL on CPUs to Inference on GPUs
6:35 – Workload Shift #2: Agentic AI & System Complexity
8:15 – Deployment Complexity: APIs, Models, Accelerators
9:10 – Workload Shift #3: Rise of Post-Training & RL
10:55 – A New Stack: Inference, Environment, and Training Loops
13:00 – Hardware Bottlenecks and Scaling Decisions
14:10 – Layer 1: Training & Inference Frameworks (PyTorch, vLLM)
15:30 – Transformer-Specific Optimizations (Speculative Decoding)
16:45 – Layer 2: Distributed Compute Engines (Ray, Spark)
18:00 – Layer 3: Container Orchestration (Kubernetes)
19:15 – Dynamic Interactions Between Layers (e.g. Autoscaling)
20:00 – Infra Case Studies: Amazon, Pinterest, Instacart
20:30 – Wrap-up
Designing agents with contextual relevance, adaptability, and memory
Using Strands Agents to go from prototype to production in hours
Tying agent outcomes directly to customer lifetime value
Leveraging feedback loops, AB testing, and continuous evaluation
Visualizing agent behavior with Arize Phoenix for maximum transparency
⏱️ Chapters
0:00 – Intro: The Experience Is the Only Feature That Matters
0:54 – Empathy Over Capability: What Makes Agents Great
2:00 – AWS Agent Stack: Q, Bedrock Agents, and Strands
3:40 – Why Strands Agents? SDK Design Philosophy
5:10 – Key Benefits: Open Source, Pre-Built Tools, Deployment Agnostic
6:30 – Integrating With AWS Services: Bedrock, S3, OpenSearch
8:00 – Observability Built-In: Traces, Logs, Guardrails
9:00 – New Integration: Arize + Bedrock + Strands Agents
9:40 – Building Trust: Contextual Relevance and Agent Memory
11:05 – Bring Back Data Science: Closing the Experience Gap
12:00 – Loops, Churn, and Connecting to Business Metrics
13:00 – Feedback as a Tool for Agents
13:40 – Continuous Evaluation & Alarming in Arize
14:30 – How To Get Started: 30 Lines of Code & Docs
15:10 – Closing: Documentation, Samples, and Getting Involved
🔗 Explore Strands Agents
strandsagents.com
👉 Want to stay up to date on trustworthy GenAI? Check out upcoming events at:
arize.com/community
Structuring agents for specific responsibilities
Using evaluations to track and improve performance
Architecting for complexity while reducing token load
Building trust with business users through transparent observability
Chapters:
00:00 - Intro: Jack of All Trades or Master of One?
00:55 - Real-world enterprise use case
02:30 - Why traditional BI dashboards weren’t enough
03:40 - Introducing the GenAI-powered data analyst agent
05:00 - Using evaluations to improve reliability
06:15 - The challenge of large context windows
07:15 - Managing complexity in prompts and examples
08:30 - Techniques for taming context: tools, RAG, and summarization
09:45 - Evaluation results: From 30% to 95% accuracy
10:25 - Handling business-specific terminology
11:00 - Transition to multi-agent system: business term and supervisor agents
12:10 - Modular multi-agent benefits: context trimming and specialization
13:30 - Architecture insights: permissioning and reuse
14:00 - Why evaluations + observability build trust
15:00 - Final takeaways and live Q&A
Watch to learn why sometimes the best AI agent isn't a jack-of-all-trades—but a master of one (or a team of specialists working together).
🛠️ Tools and techniques covered include:
Dynamic system prompts
Agentic RAG
OpenTelemetry + Arize Phoenix
OAuth-based permission handling
Practical eval-driven development
👥 Follow the speakers and their work:
🔗 Ben McHone — linkedin.com/in/benjamin-mchone-4b983abb
🔗 Vicky Bang — linkedin.com/in/vickyyunqi
📢 Explore Arize Phoenix. github.com/Arize-ai/phoenix
📅 More community events: arize.com/community
Speakers:
Nachiket Karajagi is Global Senior Director of Data & AI at PepsiCo
Shobhit Varshney is Head of Data & AI at IBM Consulting
⏱ Timestamps
00:02 - Welcome & introductions: PepsiCo, IBM Consulting, scale of AI initiatives
1:06 - PepsiCo’s value chain & surprising AI use cases: precision agriculture, soil testing
1:52 - The staggering data scale: 60 petabytes doubling yearly + IoT data
2:11 - Hyper-personalized Gatorade bottles: driving loyalty with AI-generated designs
3:43 - The impact on sales & ensuring brand-safe outputs
4:34 - Why observability is critical when scaling AI to billions of impressions
5:07 - How PepsiCo’s AI stack mirrors human intelligence: memory, attention, skills
5:15 - Inside Pep GenX: central platform enabling 100+ GenAI use cases
6:12 - Centralizing observability & evaluations with PepVigil
7:32 - Benefits of an abstracted, tech-agnostic platform for flexibility & speed
8:12 - Federating innovation across global teams while maintaining control
9:25 - Pep Vigil observing not just GenAI, but conversational, vision & IoT AI
10:28 - Example: planogram compliance via mobile edge AI — from manual surveys to real-time image processing
11:47 - Tackling scale: processing 10M+ images globally, managing regional nuances
12:26 - Using embeddings & clustering via Arize to detect drift & issues
13:32 - Computer vision for potato quality on conveyor belts
14:13 - Observability for forecasting models across markets
14:35 - Conversational AI: “chatting with data” with Pepsi’s FP&A bot
15:26 - Multi-layer filters for personalized Gatorade images (tying back to responsible AI)
15:43 - Enforcing responsible AI: upfront risk assessments, closed-loop policy execution
16:54 - The future: orchestrating multi-vendor agents, connecting ROI to business value
17:48 - The power of a “single pane of glass” for all AI across PepsiCo
18:48 - Closing thoughts & lessons learned: enabling fast time-to-market & trusted AI
Find other AI engineering events from Arize AI: arize.com/community
⏱️ Chapters
00:00 – Intro: Why RAG Evaluation Is So Hard
00:50 – The Golden Answers Problem
01:45 – What Is open-rag-eval?
02:50 – Live Demo: Side-by-Side RAG Comparison Dashboard
04:20 – UMBRELA: Reference-Free Relevance Scoring
05:30 – AutoNuggetizer: Atomic Fact Detection Without Labels
06:40 – Metrics: Citation Faithfulness & Hallucination Detection
08:10 – How It Works Without Golden Chunks or Answers
09:10 – Flexible Connectors: LangChain, LlamaIndex & More
10:15 – In-Progress: Response Consistency Metric (Preview)
11:30 – How To Get Started With open-rag-eval
12:30 – Q&A: Chunking Strategies & Multimodal RAG
🔗 More on open-rag-eval: github.com/vectara/open-rag-eval
📌 Learn more about Arize community events: arize.com/community
Chapters
00:00 – Intro: Traditional vs GenAI Observability
01:45 – Debugging LLMs: Start With Your Laptop
03:00 – It’s Just a Web Service: Demystifying LLM APIs
05:20 – Observability Tools: Using mitmproxy & OpenTelemetry
07:30 – Integration Testing for LLMs: Handling Flaky Responses
10:05 – Recording & Replaying LLM Calls (VCR, Knock)
12:15 – Offline Evals With Phoenix + Unit Test Runners
14:00 – Writing Custom Metrics: Example With Hallucination
16:00 – Human-in-the-Loop Feedback + Eval Backfilling
18:00 – Combining Elastic & Phoenix: Traces Across Systems
20:00 – Tips on Guardrails, Redaction, and Secrets
22:10 – Q&A: Prompt Eval, Redaction, and Guardrail Systems
Learn more about Phoenix:
arize.com/docs/phoenix
Chapters
00:00 - Why agents fail at long-horizon tasks
01:15 - Introducing Plan-and-Act
03:40 - LPUs: Treating LLMs like CPUs
06:00 - The LLM Compiler (ICML 2024)
08:10 - Planning remains hard for LLMs
09:50 - Using synthetic data to train better planners
11:25 - Step-by-step: how Plan-and-Act is trained
13:00 - Benchmark results: WebArena, WebVoyager
14:30 - Real-world demo: Narada AI for enterprise workflows
17:00 - Automating tools like Concur, Salesforce, Hubspot
19:20 - Vision for agent abstraction beyond SaaS
21:00 - Summary and enterprise invite
🧠 Related Papers:
Plan-and-Act: arxiv.org/abs/2503.09572
LLMCompiler: arxiv.org/abs/2312.04511
LLM2LLM: arxiv.org/abs/2403.15042
💡 Want to join the conversation? Attend upcoming events at the Arize AI community:
👉 arize.com/community
In this talk, Nick Haber (Stanford University) explores two research breakthroughs from the Autonomous Agents Lab:
Quiet-STaR: A method for continued pretraining that inserts and rewards self-generated “thoughts” in natural text, improving downstream reasoning performance.
Minimo: A fully self-supervised theorem discovery and proving system, where a transformer trained from scratch invents its own math conjectures and learns to prove them—no human data required.
These projects redefine how we train LLMs for general reasoning, pushing beyond the limitations of human-curated benchmarks and toward truly open-ended cognitive capabilities.
🔬 Quiet-STaR Paper: Language Models Can Teach Themselves to Think Before Speaking – arxiv.org/abs/2403.09629
📘 Minimo Paper: Learning Formal Mathematics from Intrinsic Motivation – arxiv.org/abs/2407.00695
📍 Talk recorded at Arize:Observe 2025
🌐 Join Arize AI community events: arize.com/community
Full talk descripition from the event:
Traditional reinforcement learning for LLM reasoning relies on curated problem sets—mathematics benchmarks, coding challenges, and other structured tasks. This approach faces fundamental scalability limits: human-designed problems are finite, expensive to create, and may not capture the full breadth of reasoning required in real applications. We examine two of our works that expand beyond this paradigm: Quiet-STaR discovers reasoning opportunities in ordinary pretraining text, while Minimo generates its own mathematical conjectures for exploration. Both approaches transform the training environment itself—from solving predefined problems to finding and creating reasoning challenges autonomously. This shift suggests a path toward LLMs that can continuously discover new reasoning domains rather than being confined to human-curated problem distributions.
⏱️ Chapters
0:00 – Intro: Helping 100s of AI startups scale
0:40 – Autonomy spectrum: Manual → Assistive → Agentive → Autonomous
2:15 – Why autonomy needs observability
3:35 – McDonald's IVR failure & lessons from voice AI
4:50 – Good Enough Revolution vs Real Autonomy
5:45 – Why chatbots are often just retry systems
6:30 – Observability for agentic systems
7:50 – Beyond logs: traces, metrics, and emergent behavior
8:45 – Economics of startup AI maturity: $5K → $500K/month
10:10 – Compression revolution: using AI to reduce boilerplate
11:40 – From Clippy to co-intelligence: the new product design loop
13:00 – 5 questions to unlock better AI system design
13:45 – Microsoft for Startups: credits, tools, and technical support
Learn more about credits, support, and tools for founders at startups.microsoft.com
See more AI engineering talks from Arize AI at arize.com/community
In this session, one of the creators of VLAA-Thinker -- Haoqin Tu, PhD Student, UC Santa Cruz --- share details on dataset curation, insights from model training, and performance results.
⏱️ Chapters
0:00 - Introduction: Vision Language Models & Multimodal AI
1:00 - VLAA Thinker Overview: Training Data, Recipes, and Reasoning
2:30 - Prompting Evolution: From CoT to Inherent Reasoning
4:10 - Inheriting Reasoning: DeepSeek-R1 and Training Implications
5:30 - Research Questions: Is Two-Stage Training Necessary?
6:30 - SFT Pipeline and Data Generation Using DeepSeek
8:00 - How Rewriting and Verifiers Clean Reasoning Traces
9:00 - Pseudo Reasoning Problems from SFT
10:30 - RL-Only Training and Reward Design (GRPO)
12:00 - Mixed Reward Module: Math, IOU, and Open Reasoning
13:30 - SFT vs RL: Which Works Best?
15:00 - Leaderboard Results on Vision Reasoning Benchmarks
16:00 - Conclusion: Toward Transparent, Generalizable Multimodal Reasoning
More about VLAA Thinking family: github.com/UCSC-VLAA/VLAA-Thinking
Follow Haoqin Tu: https://x.com/haoqint
🚀 Highlights include:
Automatically detecting and correcting errors in ML datasets
Ensuring LLM outputs are accurate and safe
Real demos showing how issues get flagged and fixed
Integration examples with tools like Jira for end-to-end remediation
Chapters
00:03 - Introduction & background
00:27 - Challenges with noisy labels in cheating detection
00:54 - Developing confident learning to fix noisy data
1:39 - Shift from ML training to LLMs & RAG systems
1:53 - Cleanlab’s unique capabilities: detecting incorrect outputs and controlling LLM agents
2:32 - Real-world failures from Air Canada, NYC chatbot, and others
4:47 - Cursor’s issue reproduced & how Cleanlab would prevent it
6:00 - Types of common chatbot errors: missing answers, hallucinations, wrong context
7:03 - Analogy to network security & controlling agent outputs
8:14 - Live demo: blocking and remediating problematic responses
10:22 - Jira integration example for fixing root causes
11:01 - Q&A and wrap-up
🤝 See other AI Engineering events from Arize AI:
arize.com/community
Learn more about open source Cleanlab LLM evaluation: arize.com/docs/phoenix/integrations/evaluation-integrations/cleanlab
#AI safety, #machine learning, #noisy labels, #confident learning, #Cleanlab, #LLM, #GPT, #RAG systems, #AI hallucinations, #chatbot
💬 “Make it go pretty — then go have a cocktail.”
⏱️ Chapters:
00:00 Intro – There Are Two Ideas in This Talk. Maybe.
01:00 Who’s Todd, and Why Is He Here Again?
02:00 The Audience (That’s You) and the AI Hype Hangover
03:25 Yes, the Models Are Getting Better — Really Better
05:00 No, That Doesn’t Mean They Work Yet
06:00 What Is Production Engineering, Anyway?
07:10 Why Reliability Needs to Be Organizational, Not Bolted On
08:30 AI Ops Has Been Around — It Just Didn’t Work
09:30 Unstructured Logs and Docs? GenAI Is Actually Useful Here
10:30 Anomaly Detection: Still Hard. Still Mostly Broken.
11:30 Natural Language Interfaces: SQL for the Rest of Us
12:30 What Works in AI Ops — And Why It’s So Narrow
13:45 The Dream Prompt: “Make My Systems Work Better, Please”
14:30 Agents Might Help — But Not the Way You Think
15:30 Near-Term Agent Wins: Alert Context, Retros, Follow-Ups
17:00 Can Agents Triage Incident Impact Intelligently? Maybe.
18:00 Next Level: Agents in Change Management & Rollouts
19:45 Could They Prevent the Next Outage? Possibly.
21:00 From Alerting to Infrastructure Planning — What Could Be Automated
22:30 The Big Leap: Systems That Architect and Deploy Themselves
24:00 What’s Holding It Back? Legibility for the Meatbags (Us)
25:00 This Future Is Inevitable — and That Shouldn’t Scare You
26:00 Final Thoughts – Just Don’t Be the Person Who Ships It Too Early
🧠 Themes Covered:
Why AI Ops doesn’t solve anything (yet)
What agents might realistically help with
How observability, alerting, and retros could improve
What’s plausible vs hype in AI + infra
Why the future of systems engineering might include agents — and when
Follow Todd on LinkedIn: linkedin.com/in/toddunder
See other AI engineering events from Arize AI: arize.com/community
Learn how the framework:
Separates step-finding and conversation generation
Uses vector similarity scoring for more deterministic next-step selection
Supports dynamic yet linear workflows typical of production IVR environments
Is benchmarked across multiple simulated customer scenarios
Read the paper: arxiv.org/abs/2503.06410
🔗 For upcoming AI events: arize.com/community
The rise of Agentic applications and automation in the Voice AI industry has led to an increased reliance on Large Language Models (LLMs) to navigate graph-based logic workflows composed of nodes and edges. However, existing methods face challenges such as alignment errors in complex workflows and hallucinations caused by excessive context size. To address these limitations, we introduce the Performant Agentic Framework (PAF), a novel system that assists LLMs in selecting appropriate nodes and executing actions in order when traversing complex graphs. PAF combines LLM-based reasoning with a mathematically grounded vector scoring mechanism, achieving both higher accuracy and reduced latency. Our approach dynamically balances strict adherence to predefined paths with flexible node jumps to handle various user inputs efficiently. Experiments demonstrate that PAF significantly outperforms baseline methods, paving the way for scalable, real-time Conversational AI systems in complex business environments.
Discover why inference isn’t just about running a single LLM call anymore, but about powering complex multi-step workflows, memory, and tool orchestration at massive scale.
You’ll also see how Friendly AI’s suite, from their high-performance endpoints to upcoming agent orchestration layer, makes this easy, fast, and cost-effective — all with a live demo on how to deploy 400K+ Hugging Face models in one click.
🚀 Key highlights:
What makes inference for agentic AI fundamentally harder than standard GenAI
How Friendly AI’s stack optimizes every layer: quantization, caching, scheduling, autoscaling
Achieving unmatched speed (500+ tokens/sec on LLaMA 3) and huge GPU savings
Friendly Agent: simplifying workflows, memory, and API integration for agentic AI
$10K free inference credit program for scaling teams
⏱ Chapters
00:03 - Introduction & background: from continuous batching to agentic AI
0:29 - Entering the era of agentic AI: beyond single-shot LLMs
0:56 - How agentic AI works: workflows, memory, control, tool use
1:33 - Why inference is the real bottleneck: low latency, throughput, GPU cost
2:13 - What teams care about: cost efficiency, performance, reliability, security, debuggability
2:45 - Introducing Friendly Suite: inference stack + upcoming Friendly Agent
3:23 - Inside Friendly Inference: quantization, caching, routing, autoscaling
3:54 - Why customers love it: 400K Hugging Face models, unmatched speed & cost savings
4:55 - Live demo: deploying a Hugging Face model to Friendly in one click
6:01 - Friendly Agent: orchestrating workflows, memory, APIs at the agent layer
6:28 - Wrap-up: letting builders focus on AI products, not infrastructure
6:55 - Announcing the $10K inference credit program
7:15 - Closing & invite to connect
See other AI engineering events from Arize AI: arize.com/community
In this session from Arize:Observe, Jay Rodge—Senior Developer Advocate for LLMs at NVIDIA—demonstrates how to combine NVIDIA’s inference microservices with Phoenix’s evaluation and tracing capabilities. From orchestrating multi-agent systems to identifying agent reasoning failures and optimizing retrieval, Jay shares a real-world blueprint for modern AI agent development.
Chapters:
00:00 - Intro: GPU-Accelerated LLM Workflows
00:34 - From Transformers to Agentic AI
01:15 - Real-World Agents in Production
02:16 - How Agents Work: Perceive, Think, Act
03:06 - Key Enablers: Reasoning Models, CrewAI, Observability
04:00 - Challenges in Scaling Agents
04:35 - The Data Flywheel for Self-Improving Agents
05:25 - What Is NVIDIA NIM?
06:03 - Overview of the NVIDIA NeMo Framework
07:00 - Building the Full Stack: Curator, Customizer, Guardrails
07:28 - Deploying LLMs with NVIDIA NIM in One Command
08:15 - Model Inventory and Partner Ecosystem
08:58 - Connecting Your Stack With NIM + Phoenix
09:45 - Unified Multi-Agent Observability with Arize Phoenix
10:30 - Demo: CrewAI, LlamaIndex, LangChain Agents Traced in Phoenix
12:00 - YAML Config: Defining Agents and Tools
13:10 - Launching the Unified Workflow with IQ Run
13:45 - Visualizing Traces and Debugging in Phoenix
14:45 - How NIM Compares to A2A and MCP
15:30 - Summary: Enterprise-Grade Agentic Stack from NVIDIA
🚀 Topics Covered:
Building multi-agent systems using NIMs and open-source frameworks
Unifying LangChain, CrewAI, LlamaIndex, and more into a single traceable workflow
How Arize Phoenix captures detailed telemetry across diverse agents
Using Phoenix to debug, evaluate, and optimize multi-agent reasoning
Real-world examples of production agent deployment
🔍 Jay walks through a fully observable workflow that combines:
A RAG agent (LlamaIndex)
A chat agent (Haystack)
A research agent with web search (LangChain)
A supervisor agent coordinating them all
🛠️ Get the Tools:
NVIDIA NIM: nvidia.com/en-us/ai-data-science/products/nim-microservices
Phoenix: arize.com/docs/phoenix
Learn more about how NVIDIA and Arize automate LLM performance optimization:
arize.com/blog/arize-nvidia-nemo-integration
Learn more about how Arize AI accelerates enterprise AI adoption on-premises with NVIDIA:
arize.com/blog/arize-ai-accelerates-enterprise-ai-adoption-on-premises-with-nvidia
Session description from the Arize:Observe site:
This session demonstrates how to create high-performance agentic workflows using NVIDIA NIM Microservices with integrated observability through Arize Phoenix. Jay Rodge -- Senior Developer Advocate, LLMs, NVIDIA -- showcases practical implementations where developers can leverage NIMs to build sophisticated AI agents that reason, plan, and execute complex tasks while maintaining comprehensive visibility into their performance. Learn how to instrument multi-agent systems that coordinate specialized NIM microservices—from LLM reasoning to embedding and reranking—with Arize Phoenix's evaluation capabilities. The presentation includes real-world examples of how this integration helps AI teams rapidly detect issues in agent reasoning, measure accuracy improvements, and optimize retrieval systems. By combining NVIDIA's GPU-accelerated inference capabilities with Arize's observability tools, developers can build more reliable, transparent AI applications that deliver measurable business value while maintaining visibility into every step of the agentic workflow.
Follow the story of Jane—a developer caught in a 3-hour production nightmare—and see how AI observability could turn that into a 5-minute resolution.
🚀 Highlights:
Why AI-written code & faster dev cycles demand a new approach to observability
How natural language, conversational, & generative UIs change debugging
The future of alerting: alert-on-everything, dynamic thresholds, AI auto-triage
The new frontier of “agent experience” (AX)—building observability for agents, not just humans
Perfect for anyone working in software engineering, devops, or SRE who wants to understand what’s coming next.
⏱ Chapters
00:03 - Introduction & why observability matters in the AI era
00:24 - About Sherwood: from building AI sales reps to observability at scale
1:24 - The story of Jane: a painful 3-hour production incident
3:26 - Why troubleshooting is so hard: complex systems, clunky tools, steep expertise needed
4:23 - Enter AI: code is now written by AI, shipping is faster, humans can’t keep up
6:00 - The rise of “vibe coding” & less review—leading to more unknowns in production
7:32 - Why this mirrors the 2010s: cloud, microservices, containers increased complexity
9:01 - Observability history: from Apollo missions to Twitter, Zipkin & Jaeger
10:40 - We need a new AI observability for this era: not just code gen, but faster debugging
11:11 - Defining “AI observability”: using AI to monitor traditional software
12:00 - Three big areas of impact: user interfaces, alerting, & agent experience (AX)
12:20 - Smarter UIs: from where’s-Waldo dashboards to natural language search & conversational telemetry
13:32 - Generative UIs that adapt for each user, each session
16:18 - Reimagining alerting: alert on everything, zero config, dynamic thresholds
20:22 - AI auto-triage: reducing human toil, LLMs pulling context & summarizing
20:56 - The rise of agent experience (AX): building APIs & docs for agents to debug systems
23:05 - Building agent-friendly interfaces, MCP APIs, webhooks, LLM.ext for documentation
25:04 - Conclusion: from hours to minutes—Jane’s story in an AI observability future
26:30 - Wrap-up & how to reach out
See more
👉 See other AI engineering events from Arize AI: arize.com/community
#aiobservability, #softwareengineering, #debugging, #AImonitoring
See how combining real-time telemetry with AI agents creates trustworthy, autonomous systems that know not just what was, but what is — preventing catastrophic failures before they happen.
🚀 Highlights:
Why static AI safety mechanisms break in production
The move from outdated “maps” to live real-time data
How New Relic’s architecture gives AI situational awareness
Example with an SRE agent detecting latency, then preventing a cascading failure by checking live telemetry
How this pattern can apply beyond software: in finance, healthcare, and more
Chapters
00:00 - Introduction: high-stakes AI agents & challenges
00:34 - About New Relic & focus on AI observability
1:16 - The problem with static safety checks in dynamic systems
2:43 - Moving from static “maps” to live data safety
2:57 - Introducing the proposer-critic architecture
4:55 - Why a static critic still fails: needs live context
5:58 - Evolving to “live fire auditing” at New Relic
6:42 - Example: SRE agent avoiding a cascading failure by rejecting a risky restart
8:04 - Providing actionable recommendations & graceful human fallback
9:13 - From brittle scripts to trusted autonomy with real-time situational awareness
10:08 - Broader applications in finance, medical, and any high-stakes domain
10:44 - Closing: AI needs to see the world as it is, not just as it was
👉 See other AI engineering events from Arize AI: arize.com/community
#AIsafety, #AIagents, #NewRelic, #observability
In this talk from Arize:Observe, Stanford researcher Ayush Chakravarthy presents findings from his recent paper Data Properties for Self-Improving Reasoning, diving into how cognitive behaviors like backtracking, subgoal setting, and verification influence a model’s ability to improve via online reinforcement learning.
The paper proposes that it's not just additional compute or chain-of-thought length that matters—it's the presence of specific behaviors in training data that determine success.
Ayush compares Quen and LLaMA across fine-tuning regimes on the Countdown arithmetic task and explores how data curation can turn otherwise static models into self-improving agents.
👉 Chapters
00:00 - Intro: From Reasoning to Agency
00:38 - What Makes Reasoning Hard for LLMs?
01:53 - Claude's Failure and Human Success on Countdown
03:10 - Cognitive Behaviors in Reasoning Models
04:25 - The Game of Countdown: Setup and Significance
05:51 - Quen vs. LLaMA on Reasoning Tasks
07:09 - Four Key Behaviors: Backtracking, Verification & More
08:39 - Do These Behaviors Exist in Base Models?
10:01 - Fine-Tuning on Curated Behaviors (SFT Results)
11:30 - Why Backtracking Matters Most
12:33 - Addressing Reviewer Skepticism: Controls and Baselines
14:00 - Can We Reverse Engineer Quen’s Success?
15:03 - Pretraining on Behavior-Rich vs. Behavior-Minimized Data
16:00 - Generalization: Do Behaviors Transfer?
17:16 - Open Questions: Eliciting Discovery and Tool Use
🔗 Read the paper: arxiv.org/pdf/2503.01307
🎓 Speaker: Ayush Chakravarthy, Stanford AI Lab
🧠 Research Focus: Cognitive behaviors in LLMs, online RL, and reasoning emergence
At Observe 2025, Arize's Mikyo King and John Gilhuly walk through the full lifecycle of a production-grade multi-agent system. Learn how to inspect traces, label outcomes, scale evaluations with LLMs-as-judges, and improve agent behavior iteratively using Arize Phoenix.
Whether you're debugging memory, tool calls, or experimenting with routing logic, this talk offers a playbook for taking agents from prototype to production with measurable impact.
🎤 Speaker
Mikyo King (Head of Open Source)
https://x.com/mikeldking
🛠 Supporting Integration
Get started with Phoenix: Open-Source Agent Evaluation
docs.arize.com/phoenix
🔗 Find more live AI engineering talks: arize.com/community
Chapters:
00:00 - Intro to João Moura and CrewAI
01:05 - Scaling to 60 Million AI Agents a Month
02:20 - From Framework to Ecosystem at CrewAI
04:00 - How Companies Mature in Agent Adoption
06:30 - The Agentic Stack: From Data to UI
08:00 - Enterprise Needs: Interoperability and Governance
10:00 - CrewAI Platform Demo and Capabilities
12:30 - From Zero to Production with Observability
14:00 - Agent Deployment Patterns in Enterprises
15:00 - Case Study: Fortune 500 Automation With CrewAI
16:30 - Top Use Cases for CrewAI Agents
18:30 - Common Pitfalls in Deploying AI Agents
19:45 - Final Thoughts and Key Takeaways
Speaker: João Moura, CEO of CrewAI
linkedin.com/in/joaomdmoura
Learn more about CrewAI tracing and observability with Arize AX: arize.com/docs/ax/integrations/frameworks-and-platforms/crewai/crewai-tracing
See other AI engineering events from Arize AI: arize.com/community
⏱️ Chapters
00:00 - Intro: Using LLMs for Customer-Centric Recommendations
01:03 - The Merchant Recommendation Problem
02:00 - Why Card Transaction Data Is Challenging
03:12 - Limited Metadata in Merchant Descriptions
04:27 - Using LLMs to Generate Merchant Metadata
05:15 - Pipeline Overview: From Names to Knowledge Graph
06:20 - Name Reconciliation for Messy Transaction Labels
07:00 - Prompted Metadata Generation in Regulated Environments
08:03 - From JSON to Knowledge Graph Triples
09:10 - Metadata Verification With Claims & LLMs
10:08 - Triaging with LLM-as-a-Judge and Human-in-the-Loop
11:12 - Graph Schema Design and Triple Ingestion
12:03 - Visualizing Merchant Knowledge Graph Connections
13:01 - Evaluation: Embedding Merchants With meta-ada vs TransE
14:12 - Retrieval & LLM Judgment for Recommendations
15:30 - Results: Precision Evaluation and Violin Plots
16:25 - Qualitative Merchant Matches: Meta vs Baselines
17:11 - Lessons Learned and Future Work
17:58 - Takeaways: Responsible Metadata Pipelines at Scale
🔍 Learn how her team:
Reconcilies merchant names to unify messy transaction data
Generates merchant metadata using regulated, structured prompts
Uses LLMs-as-a-judge to verify claims against external evidence
Embeds the knowledge graph using TransE and evaluates with LLM-based judgments
Achieves high-performance retrieval of relevant merchants for cardholder offer recommendations
Whether you're building LLM pipelines, evaluating embeddings, or deploying recommendation systems in high-stakes domains—this is a must-watch example of responsible, real-world AI.
Session Description from the Arize:Observe site:
In this session, Brenda Ng -- Applied AI/ML Director at JPMorganChase -- discusses the challenge of using credit card transactions to curate relevant features for credit card offers recommendation. In any customer-facing recommendation problem, it is customary to build a customer profile based on that customer's transaction history. Unlike in e-commerce (e.g., Amazon) where a transaction itemizes the products or services transacted by the customer, no such granular information is generally given in credit card transactions. Instead, using the merchant name and category alone, Ng's team uses LLMs to generate merchant metadata, apply open-source information to vet the generated content, and compile the vetted content into a knowledge graph linking merchants to their products, services, industries, competitors and target customers.
📅 See More AI Engineering Events from Arize
👉 arize.com/community
Learn more about LLM-as-a-judge:
arize.com/llm-as-a-judge
Learn how enterprise AI teams can move beyond spreadsheets and slow iteration to rapidly prototype intelligent chatbots that handle real-world customer care scenarios—like damaged packages, wrong items, or product comparisons.
🚀 You’ll see:
Why traditional agent development is broken
How to structure a support agent using modular LLM components
Building prompts, test datasets, and evaluators with Arize AX
Performing A/B/C testing on prompts and models
Evaluating results with tools like LLM-as-a-Judge and trajectory metrics
This is a hands-on guide for AI teams looking to automate customer service with production-ready, traceable agents.
Chapters
00:00 Why AI Customer Service Agents Are Hard To Build
01:00 Common Challenges: Spreadsheets, Silos, and Slow Iteration
01:50 The 3 Pillars: Observability, Evaluation, and Experimentation
02:30 Goal: Build a Multimodal Customer Support Agent
03:30 Best Practices Before Writing Code
04:30 Designing an Automated Customer Support Agent Architecture
05:30 Focus on the Router Agent Step
06:30 Building a Minimal Testing Dataset (Text + Images)
07:40 Uploading Dataset and Adding Ground Truth in Arize AX
08:30 Writing the First Prompt in the Prompt Playground
09:30 Testing With a Real Example and Image Inputs
10:20 Running the Prompt Across the Full Dataset
11:20 Creating Evaluators: LLM-as-a-Judge and Ground Truth Match
12:30 Measuring Prompt and Model Quality Automatically
13:10 Rapid Prototyping With A/B/C Testing
14:30 Observing Results and Making Data-Driven Decisions
15:30 What’s Next: Expanding to Other Agent Steps
16:00 Advanced Metrics: Trajectory, Reflection, and Image Quality
16:40 Final Thoughts: Iterate Fast, Build Better Agents
🔗 Follow Hakan Tekgul: linkedin.com/in/hakantekgul
👉 See other AI engineering events from Arize AI: arize.com/community
Chapters:
00:00 Intro – Why TypeScript for AI Agents?
00:34 The Problem With Python-First AI Frameworks
01:18 Meet Mastra: A TypeScript Agent Framework
01:58 Building Agents With MRA: The Basics
03:00 Using GitHub Tools via MCP Servers
04:40 Turning Agents Into APIs With Mastra CLI
05:40 Adding Memory to Agents
06:02 Building a Hacker News Agent
07:00 Creating Your Own MCP Server (Tailwind Example)
08:54 RAG Isn’t Dead: How Mastra Does Retrieval
10:28 Live RAG Demo With Vector Search
11:50 Writing Custom Tools in TypeScript
12:40 Showcasing the Ghibli Buddy
13:50 Workflows With LLMs – Weather-Based Activity Planner
15:00 Running Live Evals Inside Mastra
16:40 Multi-Agent Networks in Action
18:00 Final Demo – Unified Network of Agents
🔗 Explore Mastra: github.com/mrkaran/mastra
🔗 Learn how to trace Mastra agents: arize.com/docs/ax/integrations/mastra/mastra-tracing
👉 Attend other AI engineering events from Arize AI: arize.com/community
Learn why traditional QA breaks down for AI, and get a framework to identify failure modes before they reach users. We’ll cover error analysis that reveals what’s truly broken, how to involve domain experts without bottlenecks, and when to build custom tools versus using existing ones.
The speakers dive into GitHub Copilot’s HarnessLib system, which used unit tests to evaluate thousands of code completions—bridging the gap between algorithmic checks and human judgment. You’ll see how they focused on three key product metrics (acceptance rate, character retention, latency) while using guardrail metrics to monitor regressions and ship with confidence.
Walk away with a repeatable process and real-world examples to help your AI product scale beyond the prototype stage.
Chapters:
00:00 Why AI Products Fail Without Evaluation
00:34 What Is Error Analysis? (Grounded Theory Method)
01:45 Sampling Traces & Finding Patterns
03:00 Who Should Do Error Analysis? (Hint: Not Engineers)
04:40 Building Annotation Tools That Fit Your Product
06:00 The Truth About Generic Metrics
07:45 Top Mistakes Teams Make With Evals
09:00 Error Analysis Builds Team Alignment
09:45 Intro to Evals Beyond Error Analysis
10:00 Shawn Simister: Lessons From GitHub Copilot
11:00 HarnessLib: Evaluating Code Gen With Unit Tests
13:00 From R&D to Regression Testing
14:10 Rethinking Evaluation for LLM Chat Interfaces
15:00 Data-Driven Eval Design Using Label Studio
16:30 LLM-as-a-Judge: Self-Eval With Checklists
17:30 Closing Thoughts: Evaluation Is Product Work
🤝 See other AI Engineering events from Arize AI:
arize.com/community
Learn more about AI Product Management:
arize.com/ai-product-manager
Learn more about LLM evals:
arize.com/llm-evaluation
The system breaks the loan evaluation process into specialized agents—such as credit analysis and risk assessment—each running independently as a service and orchestrated via LangGraph. Using Arize's tracing, we track every decision step, input-output transformation, and agent execution in real time. This provides a transparent audit trail of how each loan decision is made, including latency and field-level reasoning across agents.
The talk includes a live example showing how a single loan application flows through the agent graph and appears in Arize as a fully visualized trace. By combining agentic AI with cloud-native infrastructure and open observability standards, this approach demonstrates how next-gen systems can be both intelligent and inspectable.
📌 Chapters
0:00 – Introduction: Welcome and overview by Vive and Surya from AWS SageMaker team.
0:32 – AI Agents in Production: Tools vs. Orchestration: Why agentic systems benefit from structured tool use and how orchestration evolves in enterprise use cases.
1:33 – MCP: Modular, Scalable Agent Interfaces: What MCP (Model Connector Protocol) is and why it’s ideal for standardizing agent-tool interaction in a scalable, framework-agnostic way.
3:05 – Benefits of MCP Servers in Enterprise AI: Scalability, dynamic discovery, flexible tooling, and compliance observability.
5:00 – Observability & Tracing With Arize AI: Why tracing agent steps is crucial in real-world systems and how Arize enables that via OpenInference.
6:50 – Agentic Loan Underwriting Architecture: Breakdown of the demo system with modular MCP servers for credit, risk, and decision-making.
8:40 – Agent Flow & LLM Inputs: Loan inputs, agent responsibilities, and how outputs pass between MCP servers.
10:00 – Demo: Loan Decision Pipeline in Action: Live walkthrough of the pipeline execution with OpenInference logs and credit decision outputs.
13:35 – Observability Setup With Langchain & Arize: How to easily integrate OpenTelemetry and Arize into Langchain-based systems.
15:10 – Real-Time Results in Arize: View actual traces, MCP port logging, summaries, and credit decisions for a loan applicant.
16:20 – Future of Agentic RAG & Memory: Q&A and discussion on expanding workflows with retrieval, memory types, and compliance readiness.
🔗 Learn about Amazon Bedrock Agents observability using Arize AI
aws.amazon.com/blogs/machine-learning/amazon-bedrock-agents-observability-using-arize-ai
🤝 See other AI Engineering events from Arize AI:
arize.com/community
This talk walks through a principled framework for building and optimizing agent-guided systems using AG2, an open-source system for building AI agents. Kamradt outlines core dimensions of agent design (natural interface, strong capabilities, composable architecture), and introduces patterns for multi-agent orchestration, evaluation, and optimization — from early design to hill climbing and failure attribution. Also covered: techniques for structuring agent workflows, examples of community use cases in industry (e.g., NVIDIA, Better Future Labs), and what it means to support agent-native design across domains.
⏱️ Chapters:
00:00 – Introduction: Three dimensions of strong agent systems
01:30 – Design tradeoffs: Quality, latency, cost, monitoring
02:00 – A thinking framework: Design, evaluate, optimize
03:00 – AG2 OS: Unified abstraction for agent composition
04:00 – Multi-agent orchestration patterns & message-passing
05:30 – Example: Customer service system with nested agents
07:00 – Structuring design space: simplifying & scaling composition
08:00 – Evaluation: Failure attribution in multi-agent systems
09:30 – AgentArena & other tools for comparative evaluation
10:30 – Optimization techniques: Captain agent, hill climbing
12:00 – Automatic team generation, plan iteration & reflection
13:30 – Using AG2 with no-code UIs like Wadis
14:30 – Case studies: Deep research, analyst workflows, due diligence
16:00 – Enterprise domains: Cybersecurity, finance, healthcare, chip design
18:00 – Planning and control: Structured outputs, dynamic updates
19:00 – Building reusable, general-purpose research agents
20:00 – Invitation to join the AG2 open source community
🔗 Follow AG2 Community:
ag2.ai
🤝 See other AI Engineering events from Arize AI:
arize.com/community
Learn more about AI agent evaluation:
arize.com/ai-agents/agent-evaluation


