Inside 847 Production Clinical AI Notes — Sebastian Fox, Composo @aiDotEngineer
Inside 847 Production Clinical AI Notes — Sebastian Fox, Composo  @aiDotEngineer
Uploaded August 2026 | Updated September 2026, 3 weeks ago
A clinical note from a real consultation reads like a routine tension headache, and nothing in it is wrong. What never reached the page is that the patient also mentioned her jaw aches when she chews, which alongside a new headache over 50 is a red flag for a condition that can take her sight within days. Sebastian Fox pulled that error, and every other failure here, out of three leading production ambient scribes in one afternoon. In the largest real world study of these notes, roughly one in 20 carried an error serious enough to cause significant harm, nearly one in five had an important omission, and more than one in 10 contained a hallucination. Ambient scribes now run in about a third of US practices, and almost none of this surfaces as a reported incident.

The obvious fix is a checker after the generator, and Fox built the best version he had seen: a frontier model, a faithfulness rubric with worked examples, automatic rubric optimization, deterministic concept counting. One in five of the notes it waved through still hid a serious error. Verification is only cheap for the easy half, spotting what changed between transcript and note. Deciding which differences matter is tacit, contextual and always moving, so it was never written down anywhere a rubric could read. Two notes drop the same holiday detail, and France is noise while Lake Malawi is the diagnosis. His answer is to keep the standard as examples rather than specifications, discovered from real outputs and assembled per note.

Speaker info:
- linkedin.com/in/seb--fox
- composo.ai

Timestamps:
0:00 - The note that looks completely fine
1:29 - The obvious errors, and how common they are
4:02 - Mapping every failure across three production scribes
5:30 - Mishearings, additions, changes, omissions
7:03 - The hard part is knowing what matters
7:56 - Put a checker after the generator
9:35 - The best judge waved a fifth of them through
12:06 - France versus Lake Malawi
13:49 - Discover, capture, calibrate
17:13 - Three judges on the same notes
18:06 - Beyond healthcare
Inside 847 Production Clinical AI Notes — Sebastian Fox, ComposoYour Finance Agents Bottleneck Is You — Ramana Siddanth Emani, Auditoria AIVoice agents with Realtime Video — Sidney Primas, LemonSliceCoding Agents Dont Scale Themselves. Neither Do Your Teams. — Patrick Debois, TesslScaling up Continual Learning — Ronak Malde, TrajectoryAgents Need Feature Flags - Sachin GuptaUnlock Agent Autonomy: The Runtime for AI-Native Systems — Tushar Jain, DockerEmulated: The Data for Fully Autonomous Software Engineers and Companies — Joseph WangProductionizing LLM Gateways: Architecture, Tradeoffs and Hard Lessons — Kanish Manuja, Twilio
AI Engineer |

Inside 847 Production Clinical AI Notes — Sebastian Fox, Composo

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER