Benchmarking LLM Costs: GPT-5.5, Kimi K3, DeepSeek, and 8 More Models | AI Builders @arizeai
Benchmarking LLM Costs: GPT-5.5, Kimi K3, DeepSeek, and 8 More Models | AI Builders  @arizeai
Uploaded August 2026 | Updated September 2026, 2 weeks ago
When choosing a foundation model for production AI agents, quoting price per million tokens from a pricing page is misleading. A model with a low per-token cost that fails two-thirds of its attempts leaves you paying a massive "retry tax" on failed runs, malformed outputs, and token budget timeouts.

In this episode of AI Builders, Arize AI and Fireworks AI present the findings of a 2,400-run benchmark study across 10 open and closed models (including GPT-5.5, Kimi K3, DeepSeek V4 Pro, Gemini 3.1 Flash-Lite, Gemini 3.5 Flash, GLM-5.2, and gpt-oss-120b) evaluated on 40 real-world Terminal-Bench tasks.

We introduce a critical metric: Cost Per Successful Task, and break down how to design deployable model escalation ladders that beat frontier models on both cost and task reliability.

Hashtags: #ArizeAI #LLMEvaluation #AgentObservability

Chapters:

00:00 Why cost per token is the wrong metric
00:32 Introducing Arize and Fireworks
02:41 Why LLM pricing pages can mislead you
04:19 Cost per successful task and the retry tax
05:48 Comparing 10 models by price
06:26 Ranking models by cost per successful task
07:35 How we benchmarked 2,400 agent runs
11:50 Easy vs. hard tasks change which model wins
13:34 GPT-5.5 vs. Kimi K3: same average, different strengths
15:25 Fireworks benchmark: Kimi K3 vs. Fable
17:55 Why top-level benchmark scores hide important differences
21:21 Oracle routing: picking the right model for each task
24:21 Cost vs. coverage: why the cheapest model isn't enough
25:00 Open vs. closed models: what actually predicts performance?
25:49 Why model routing can beat any single model
27:10 Building an escalation ladder
29:01 How to detect failure and route to another model
32:47 Deterministic verifiers, distress signals, and LLM judges
34:32 Using traces to understand expensive agent failures
36:13 When a model spirals and burns its token budget
38:19 Benchmark caveats and how to run this yourself
39:38 Key takeaways: optimize for cost per successful task
40:58 Q&A
41:34 Can you really swap models inside the same agent harness?
43:26 Open models, security, and data privacy
47:44 Does the same open model perform differently across providers?

đź”— Try Arize AX for free: arize.com/products/ax/?utm_source=youtube&utm_medium=social&utm_campaign=q32026-webinar-real-cost-of-ai-na&utm_content=youtube-description

Resources:
🔬 Arize Phoenix (open source): https://phoenix.arize.com?utm_source=youtube&utm_medium=video&utm_campaign=otel-genai-semantic-conventions&utm_content=phoenix
đź“– OpenInference: github.com/Arize-ai/openinference?utm_source=youtube&utm_medium=video&utm_campaign=otel-genai-semantic-conventions&utm_content=openinference
đź“– Phoenix docs: docs.arize.com/phoenix?utm_source=youtube&utm_medium=video&utm_campaign=otel-genai-semantic-conventions&utm_content=phoenix-docs

đź”” Subscribe for weekly content on LLMs, agents, and evaluation: youtube.com/@arizeai?sub_confirmation=1

Music by Anthony Powell, an open source engineer at Arize AI.
Benchmarking LLM Costs: GPT-5.5, Kimi K3, DeepSeek, and 8 More Models | AI BuildersA Watermark for Large Language ModelsGoogle TUMIX AI Agent Paper, Explained By Its AuthorOpenClaw vs Hermes: The Future of Open-Source AI Agents | Arize Observe 2026One AI Question - what do you do at night,  doom prompting with Matt WilsonAI Enablement At Enterprise Scale1.4 Billion Smiles: How PepsiCo Scales AI with PurposeProving a Prompt Fix Works in Production with Phoenixs PXIHow PromptQL Built a Self-Updating Company Brain for AI Agents | Arize Observe 2026Harnessing User Feedback at ChatGPT Scale | OpenAI | Arize Observe 2026Building, Deploying, and Optimizing AI Agents with Microsoft Foundry | Arize Observe 2026Session Evaluation On An AI Tutor Chatbot
Arize AI |

Benchmarking LLM Costs: GPT-5.5, Kimi K3, DeepSeek, and 8 More Models | AI Builders

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER