The Gap Between Humans and Machines Is   [Dr. Max Bartolo] @MachineLearningStreetTalk
The Gap Between Humans and Machines Is   [Dr. Max Bartolo]  @MachineLearningStreetTalk
Uploaded March 2025 | Updated September 2026, 1 week ago
SPONSOR MESSAGES:
***
Tufa AI Labs is a brand new research lab in Zurich started by Benjamin Crouzier focussed on o-series style reasoning and AGI. They are hiring a Chief Engineer and ML engineers. Events in Zurich.

Dr. Max Bartolo from Cohere discusses the gap between model capabilities and genuine robustness: why next-token prediction can produce impressive results yet still fail on slightly reformulated…

---
TIMESTAMPS:
00:00:00 Model Reasoning and Consistency Verification
00:03:25 Influence Functions and Distributed Knowledge Analysis
00:10:28 AI Application Development and Model Deployment
00:14:24 AI Alignment and Human Feedback Limitations
00:20:15 Human Evaluation Challenges and Factuality Assessment
00:27:15 Cultural and Demographic Influences on Model Behavior
00:32:43 Adversarial Examples and Model Robustness
00:41:54 DynaBench and Dynamic Benchmarking Approaches
00:50:02 Benchmarking Challenges and Data-Centric Evaluation
00:55:15 Cohere Command A Development Process
01:00:26 Model Quantization and Performance Evaluation
01:05:18 Reasoning Capabilities and Training Progression
01:13:48 Context Windows and Enterprise Applications

---
REFERENCES:
person:
[00:00:00] Max Bartolo Website
maxbartolo.com
company:
[00:00:00] Cohere
cohere.com/command
[00:12:10] Command A Model
huggingface.co/CohereForAI/c4ai-command-a-03-2025
paper:
[00:03:25] Procedural Knowledge in Pretraining Drives Reasoning in LLMs
cohere.com/research/papers/procedural-knowledge-in-pretraining-drives-reasoning-in-large-language-models-2024-11-20
[00:04:15] Influence Functions in Machine Learning
arxiv.org/abs/1703.04730
[00:08:05] Studying Large Language Model Generalization with Influence Functions
arxiv.org/abs/2308.03296
[00:16:15] Human Feedback is not Gold Standard
arxiv.org/abs/2309.16349
[00:27:15] The PRISM Alignment Dataset
arxiv.org/abs/2404.16019
[00:32:50] Adversarial Examples Are Not Bugs, They Are Features
arxiv.org/abs/1905.02175
[00:43:00] DynaBench: Rethinking Benchmarking in NLP
aclanthology.org/2021.naacl-main.324.pdf
[00:50:15] Sara Hooker on Compute Limitations
arxiv.org/html/2407.05694v1
[00:53:25] DataPerf: Benchmarks for Data-Centric AI
arxiv.org/abs/2207.10062
[01:04:35] DROP: A Reading Comprehension Benchmark
arxiv.org/abs/1903.00161
[01:07:05] GSM8k
paperswithcode.com/sota/arithmetic-reasoning-on-gsm8k
[01:09:30] ARC-AGI Challenge
github.com/fchollet/ARC-AGI

---
LINKS:
Full Transcript: app.rescript.info/share/163bf5e7338685f635fc0b8d6920005c
Download PDF transcript: app.rescript.info/api/public/sessions/7167a679366f97f8/pdf

REFS:
[00:03:10] Research at Cohere with Laura Ruis et al., Max Bartolo, Laura Ruis et al.
cohere.com/research/papers/procedural-knowledge-in-pretraining-drives-reasoning-in-large-language-models-2024-11-20
[00:04:15] Influence functions in machine learning, Koh & Liang
arxiv.org/abs/1703.04730
[00:08:05] Studying Large Language Model Generalization with Influence Functions, Roger Grosse et al.
storage.prod.researchhub.com/uploads/papers/2023/08/08/2308.03296.pdf
[00:11:10] The LLM ARChitect: Solving ARC-AGI Is A Matter of Perspective, Daniel Franzen, Jan Disselhoff, and David Hartmann
github.com/da-fr/arc-prize-2024/blob/main/the_architects.pdf
[00:12:10] Hugging Face model repo for C4AI Command A, Cohere and Cohere For AI
huggingface.co/CohereForAI/c4ai-command-a-03-2025
[00:13:30] OpenInterpreter
github.com/KillianLucas/open-interpreter
[00:16:15] Human Feedback is not Gold Standard, Tom Hosking, Max Bartolo, Phil Blunsom
arxiv.org/abs/2309.16349
[00:27:15] The PRISM Alignment Dataset, Hannah Kirk et al.
arxiv.org/abs/2404.16019
[00:32:50] How adversarial examples arise, Andrew Ilyas, Shibani Santurkar, Dimitris Tsipras, Logan Engstrom, Brandon Tran, Aleksander Madry
arxiv.org/abs/1905.02175
[00:43:00] DynaBench platform paper, Douwe Kiela et al.
aclanthology.org/2021.naacl-main.324.pdf
[00:50:15] Sara Hooker's work on compute limitations, Sara Hooker
arxiv.org/html/2407.05694v1
[00:53:25] DataPerf: Community-led benchmark suite, Mazumder et al.
arxiv.org/abs/2207.10062
[01:04:35] DROP, Dheeru Dua et al.
arxiv.org/abs/1903.00161
[01:07:05] GSM8k, Cobbe et al.
paperswithcode.com/sota/arithmetic-reasoning-on-gsm8k
[01:09:30] ARC, François Chollet
github.com/fchollet/ARC-AGI
[01:15:50] Command A, Cohere
cohere.com/blog/command-a
[01:22:55] Enterprise search using LLMs, Cohere
cohere.com/blog/commonly-asked-questions-about-search-from-coheres-enterprise-customers
The Gap Between Humans and Machines Is   [Dr. Max Bartolo]The Man Who Invented Modern AI (Before Everyone Else) — Jürgen SchmidhuberAI IN THE DOCK (Sam Altman, Gary Marcus)Manhattan Project for AI Safety [Connor Leahy]He Co-Invented the Transformer. Now: Continuous Thought Machines [Llion Jones / Luke Darlow]SchmidhuberedDo AI chatbots behave WEIRDLY?When AI Discovers the Next Transformer — Robert LangeTransformers Need Glasses! [Federico Barbero]is integrated information theory pseudoscience? Prof. Friston explains why it isnt #consciousnessAI is stuck in Plato’s Cave (Maxwell Ramstead)29.4% ARC-AGI-2 🤯 (TOP SCORE!) - Jeremy Berman
Machine Learning Street Talk |

The Gap Between Humans and Machines Is ___ [Dr. Max Bartolo]

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER