How to Build Reliable AI Agents (Context + Evals Explained) | Tobias Leong, Axium @arizeai
How to Build Reliable AI Agents (Context + Evals Explained) | Tobias Leong, Axium  @arizeai
Uploaded April 2026 | Updated September 2026, 2 weeks ago
Welcome to Rise of the AI Engineer — APJ Edition, a series from RISE AI spotlighting the thought leaders building agentic AI systems across Asia Pacific.

In this episode, Arize's Patrick Kelly sits down with Tobias Leong, CTO and co-founder of Axium Industries, to talk about what it really takes to go from demo to production with AI agents in supply chain.

Tobias shares his unconventional path from Singapore diplomat to GovTech data engineer to Databricks solutions architect, and why he believes the APJ market is one of the most exciting places to build right now.

🔗 Read the full blog post here: arize.com/blog/ai-agents-in-production-context-evaluation

What you’ll learn
• Why swapping to a “better model” often does nothing
• What “confident irrelevance” looks like in production systems
• How to design agents using retrieval + reasoning separation
• Why context matters more than model choice
• How Axium builds golden datasets with domain experts
• What evaluation looks like for real-world agent systems
• When to build vs. buy (Axium’s experience with Arize Phoenix)
• What the AI Engineer role actually requires

About Axium Industries

Axium builds agentic AI systems for supply chain across:

• logistics
• manufacturing
• maritime
• energy

Their platform (Supply OS) runs agents in real-world environments where decisions have real cost.

#AIEngineering #AIAgents #LLMEvaluation #AgenticAI #AIInProduction #Arize #SupplyChainAI #APAC #AIObservability

Chapters

0:00 Intro — Why AI agents fail in production
1:10 Tobias’s background (diplomat → GovTech → Databricks → founder)
5:00 Why he started Axium Industries
7:30 What Axium builds (agentic supply chain systems)
10:30 The shift in software engineering (AI writing code)
13:30 Demos vs production: where agents break
16:00 Lesson #1: Domain context matters more than models
19:30 “Confident irrelevance” explained
22:00 Why better models don’t fix bad systems
24:30 Build vs buy (Axium’s experience with Arize Phoenix)
28:00 How Axium approaches evaluations
31:30 Golden datasets + domain experts
35:00 The rise of the AI Engineer role
39:00 Designing reliable agents (retrieval vs reasoning)
42:30 What Axium would do differently
45:00 Why APAC is a major opportunity
46:30 What “great” looks like for Axium / Outro
How to Build Reliable AI Agents (Context + Evals Explained) | Tobias Leong, AxiumHow we debug AI agents using AI agents (real trace debugging workflows)How Phoenix became the standard for AI observabilityLLMs Are Leverage (And Why Most Developers Misuse Them)Why AI Agents Break in Production (and Why You Need Evals) | Ep. 1In a World Where Everyone Can Code, What Are You Worth?How to Measure AI Coding Agent ROI with Claude Code Tracing | Arize AXLLM-as-a-Judge 101Improving Agents in Production with Online Evals - Arize AXTypeScript Agents: How To Build and EvaluateOne AI Question - where do agents fail in production, with Fuad AliArize Skills: Add Instrumentation & Tracing to Your AI App with Claude Code, Copilot, or Cursor
Arize AI |

How to Build Reliable AI Agents (Context + Evals Explained) | Tobias Leong, Axium

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER