Data Quality Is the Compute Multiplier — Ari Morcos, DatologyAI @aiDotEngineer
Data Quality Is the Compute Multiplier — Ari Morcos, DatologyAI  @aiDotEngineer
Uploaded July 2026 | Updated September 2026, 3 weeks ago
Swap compute for data on the scaling curve and the same money buys a better model, which is why Ari Morcos calls data quality the compute multiplier and the most underinvested part of training. His frame is an oil refinery for data rather than a firehose: clean, curate, create, and compose, with quality classifiers, deduplication, and synthetic generation each earning their place, and the sequencing across stages mattering as much as any single step. The scarce resource now is not tokens but signal per token, and finding data that is optimal for a given target is where the leverage hides.

The proof points are concrete. Better curated data lets a small multilingual model beat far larger ones trained on many more tokens, and it buys real inference efficiency because a model reaches the same quality with less. Morcos points to DatologyAI's customer results, from Thomson Reuters gaining on proprietary legal data in mid-training to Arcee's Trinity reaching the open frontier on public data alone. The closing argument is blunt: it is cheaper to manufacture high quality data than to buy more compute, so data curation is quietly shaping the future of model training.

Speaker info:
- https://x.com/arimorcos
- linkedin.com/in/arimorcos
- arimorcos.com

Timestamps:
0:00 - Data is all we think about
0:52 - Why good data became scarce
2:19 - Swapping compute for data on the curve
3:48 - An oil refinery for data
5:52 - Curation work at DatologyAI
6:54 - Proof: small models beating bigger ones
8:58 - Inference efficiency from better data
9:24 - Multilingual gains
12:16 - Synthetic data done right
14:10 - Thomson Reuters and Arcee results
17:43 - Cheaper than buying compute
Data Quality Is the Compute Multiplier — Ari Morcos, DatologyAIThe Era of Compound Engineering — Kieran Klaassen, Every/CoraAgents, codebases, and teams — Aditya Khandelwal, Amazon AGI LabAI Evals for Cross-Functional Teams — Nachiket Paranjape & Swaroop Chitlur Haridas, DoorDashHow Windsurf writes 90% of your code with an Agentic IDE - Kevin Hou, WindsurfEvaling Video Slop — Maor Bril, Character.aiPrototyping as Leadership: How a CTO Ships with AI Agents — Hursh Agrawal, The Browser CompanyInfra behind Krea 2: How to train and serve at scale — Gabriel Jorge Menezes, Krea.aiWhats Next After RLHF? — Diogo Almeida, TypeSafe AIVending-Bench: Long-Horizon Agent Evals — Lukas Petersson, Andon LabsAgentic Development Security — Ezra Tanzer, SnykTribal Dungeons of Global Shipping: AI Agents at Global Scale — Dmitry Buykin, Maersk
AI Engineer |

Data Quality Is the Compute Multiplier — Ari Morcos, DatologyAI

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER