Don’t be data poor — Anuj Iravane, Anterior @aiDotEngineer
Don’t be data poor — Anuj Iravane, Anterior  @aiDotEngineer
Uploaded August 2026 | Updated September 2026, 3 weeks ago
Roughly 70% of medical communication still moves by fax. What reaches Anterior is scanned fax bundles that can run past 300 pages, carrying handwriting, checkboxes, tables and images across one patient's entire clinical trajectory. Anuj Iravane calls it an observation through a fuzzy lens over a lifespan. It is exactly the data his evals need, and the data he is least allowed to keep: their contracts rule out retaining it, deriving from it, or holding redacted or anonymized copies. Nothing survives into a dataset. In a domain where 95% accuracy is not good enough, that is a real problem.

So they generate it, by running the inference workflow backwards. The forward task takes unstructured data plus a policy, follows a reasoning trace and arrives at a label. Reversed, you sample a label, sample a reasoning trace, then build the record that would have produced it. That works because Anterior already models policies explicitly as decision trees, so traces come from a far more uniform distribution than a model asked to invent variety, which tends to collapse onto the same few cases. A coarse to fine pipeline layers patient invariants into a journey of provider encounters, then fans out into documents, with a consistency eval catching contradictions between documents written in parallel. Because generation starts from the label, labels are correct by construction and ground truthing disappears. Clinicians own the pipeline as skills rather than code. Roughly 90% of their datasets are now synthetic, and in a blind review clinicians separated synthetic from real only about 60% of the time.

Speaker info:
- https://x.com/anujiravane
- linkedin.com/in/anujiravane
- anterior.com

Timestamps:
0:00 - Policy guided decisions over highly unstructured data
1:05 - Most medical communication still arrives by fax
2:11 - Why 95% is not good enough
2:37 - The data you need most is the data you cannot keep
3:05 - Betting on generating it instead
3:55 - Why one shotting a 300 page record fails
5:00 - Reversing the forward task
5:51 - Policies as decision trees you can sample from
7:19 - Testing the edge cases production data never had
8:09 - Building a record coarse to fine
9:54 - The refinement loop and the round trip check
11:09 - Why it never becomes a PDF
11:34 - Giving clinicians the keys through skills
14:12 - Results, and datasets built just in time
Don’t be data poor — Anuj Iravane, AnteriorAdaption Labs: Gradient-Free Continual Learning — Sara Hooker, AdaptionAutonomous Agents for Scientific Tasks - Sina Shahandeh, RadicaitGenerative Video at the Speed of Light — Keegan McCallum, uRunThe Missing Layer in Agentic AI — Giedrius Šteimantas, OxylabsGuardrails First: Engineering Member-Facing Health AI — Rashi Agrawal, Hinge HealthKV Cache-Aware Routing and P/D Disaggregation on Kubernetes — Yuchen Fama & Ashish Kamra, Red HatHow to Kill the Code Review — Ankit Jain, AviatorDeepSWE: A Contamination-Resistant Coding Benchmark — James Shi, DatacurvePersona Engineering: A Field Guide to AI Synthetic Personas — Ishan Anand, InsightSciences.aiTrading Desks to Clinical Trials: Parallels in Applied Vertical AI — Ayush Bhardwaj, Allos AIAnthropics CCA Exam as a Field-Guide for Agentic Engineering — Frank Coyle, UC Berkeley
AI Engineer |

Don’t be data poor — Anuj Iravane, Anterior

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER