Speech to Text Is Harder Than You Think @WhatsAI
Speech to Text Is Harder Than You Think  @WhatsAI
Uploaded December 2025 | Updated September 2026, 2 hours ago
Most people think speech to text is just “audio in, words out.”
That’s fine… until you build a real voice agent.

Then one misheard city name breaks your CRM.
One wrong digit fires the wrong automation.
And suddenly your “great WER” means nothing.

What actually matters is entity accuracy, latency measured in milliseconds, and handling messy human speech. Accents. Code switching. Half finished sentences. Real conversations don’t wait for perfect transcripts.

This is why partial transcripts, fast turn taking, and strong NER matter more than leaderboard metrics. And why Montreal is the ultimate stress test for STT systems.

This is also why I like how Gladia approaches the problem. No hype. Just engineering for how people actually speak.

If you’re building voice agents, ask a better question than “what’s the best STT?”
Ask whether it understands humans in the real world. 🎙️

I’m Louis-François, PhD dropout, now CTO & co-founder at Towards AI. Follow me for tomorrow’s no-BS AI roundup 🚀

#AIEngineering #VoiceAI #SpeechToText #short
Speech to Text Is Harder Than You ThinkFine-Tuning Explained in 60 Seconds (No Math!)How Prompt Injection Attacks RAG and AI AgentsDeepSeek-OCR beats 70B-param giants with 256 vision tokens 🤯Why the US Government Blacklisted AnthropicWhy AI Agents Need Context CompactionYou’re Not Training ChatGPT By Pasting DataHow to Control Randomness in ChatGPT and ClaudeDay 4/42: How AI understands meaningThe hidden cost of waitingIs Synthetic Data Ruining LLMs?4x faster coding with AI? Meet Composer by Cursor
Whats AI by Louis-François Bouchard |

Speech to Text Is Harder Than You Think

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER