Upstart’s First AI Voice Bot: Lessons From Production | Shiv Indap | Arize Observe 2026 @arizeai
Upstart’s First AI Voice Bot: Lessons From Production | Shiv Indap | Arize Observe 2026  @arizeai
Uploaded July 2026 | Updated September 2026, 2 weeks ago
Voice AI demos are easy. Production voice bots are not.

In this Observe 2026 session, Shiv Indap, Senior Engineering Manager at Upstart, shares the hard-won lessons from building Upstart’s first production voice bot: a customer-support automation system for loan application verification.

The use case sounded simple: ask three verification questions, reduce a backlog of thousands of calls, and help customers get through the process faster. But once the bot moved toward production, the team ran into the real challenges of voice AI: latency, background noise, WebRTC reliability, multimodal vs. cascading architectures, premature call endings, off-topic conversations, transcription errors, model regressions, and monitoring production quality.

Shiv walks through the architecture decisions, vendor tradeoffs, production bugs, and evaluation strategy behind the system — including why Upstart moved from early multimodal approaches to a cascading architecture using speech-to-text, an LLM inference layer, and text-to-speech.

You’ll learn:
- Why Upstart chose a narrow verification-call workflow for its first voice bot
- How to think about multimodal vs. cascading voice AI architecture
- Why latency, cost, and customization tradeoffs matter in production
- What went wrong with early AWS Nova Sonic and OpenAI experiments
- How Upstart landed on a Google + Gemini Flash Lite + TTS stack
- Why background noise and first-message interruptions are harder than they look
- What the “phantom yes” problem is in voice AI
- Why WebRTC infrastructure is not something most teams should build themselves
- How model regressions can create unexpected production behavior
- How simulations, monitoring, traces, and voice evals help keep a bot reliable
- What Upstart learned after reaching thousands of calls and 96% accuracy

Chapters:
0:00 Why Upstart built a voice bot
0:33 Picking the right first voice AI use case
1:24 The verification-call backlog
2:49 Why voice AI fit this workflow
3:22 Real call demo: user friction in voice AI
5:34 Multimodal vs. cascading architecture
7:10 Why flexibility mattered more than simplicity
8:22 Attempt 1: AWS Nova Sonic
9:32 Attempt 2: OpenAI latency issues
11:05 Final stack: Google, Gemini Flash Lite, and Rime
11:51 The first production bugs
12:16 The “phantom yes” problem
13:21 Off-topic conversations and time limits
14:10 Interruptions, WebRTC, and connectivity
15:28 Background noise and audio cleanup
15:57 The Gen Z personality regression
16:55 Transcription errors, monitoring, and evals
18:38 Voice observability with Arize and Coval
19:00 Results: 4,000 calls and 96% accuracy

Presented at Observe 2026.

#VoiceAI #AIAgents #Upstart

🔗 Try Arize AX & Phoenix OSS: arize.com
🔔 Subscribe for weekly content on LLMs, agents, and evaluation: youtube.com/@arizeai?sub_confirmation=1
Upstart’s First AI Voice Bot: Lessons From Production | Shiv Indap | Arize Observe 2026Analyzing LLM Evaluations of Customer Reviews Using Repetitions FeatureHow to Build Self-Improving AI Agents with Coding Agents | Ep. 13How Uber Evaluates AI Agents at Production Scale | Arize Observe 2026Identity, Permissions, and Security for AI Agents | WorkOS | Arize Observe 2026How My AI Agent Rewrites Itself Overnight | Chi Wang, AG2 | Arize Observe 2026Traces and Evals Explained: The Building Blocks of AI and Agent Testing | Ep. 2Nebulocks Ron Cahlon on Building AI for CybersecurityAI Agent Mastery Certification Course: Lab 6 – Agent EvalsHarnessing Splits in your Dataset with Arize PhoenixEvaluating TypeScript Agents - Mastra and Arize PhoenixHomework 3 for AI Evals Course: LLM-as-a-Judge
Arize AI |

Upstart’s First AI Voice Bot: Lessons From Production | Shiv Indap | Arize Observe 2026

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER