Uploaded August 2026 | Updated September 2026, 2 weeks ago
💡 Dan Maloney, CEO of Landing AI, says it's 80/20 — and most of the hard work is vision. Classic OCR reads a page top-left to bottom-right. But what if the document has three columns? Or reads right to left? Or has tables, images, and mixed layouts? OCR loses all that context. That's a vision problem, not a language one.
The real breakthrough came when vision transformers replaced classic OCR and agentic systems started looking at documents the way humans do — multiple passes, multiple angles. Once you crack the vision side, the language part becomes the easy part.
🎧 Watch the full episode: youtu.be/uBfOQt4Ls4U
đź”— View all podcasts: datasciencedojo.com/podcast
#VisualAI #DocumentAI #OCR #ComputerVision #LandingAI #FutureofDataAndAI
💡 Dan Maloney, CEO of Landing AI, says it's 80/20 — and most of the hard work is vision. Classic OCR reads a page top-left to bottom-right. But what if the document has three columns? Or reads right to left? Or has tables, images, and mixed layouts? OCR loses all that context. That's a vision problem, not a language one.
The real breakthrough came when vision transformers replaced classic OCR and agentic systems started looking at documents the way humans do — multiple passes, multiple angles. Once you crack the vision side, the language part becomes the easy part.
🎧 Watch the full episode: youtu.be/uBfOQt4Ls4U
đź”— View all podcasts: datasciencedojo.com/podcast
#VisualAI #DocumentAI #OCR #ComputerVision #LandingAI #FutureofDataAndAI










