AI Sandbagging - Computerphile @Computerphile
AI Sandbagging - Computerphile  @Computerphile
Uploaded May 2025 | Updated September 2026, 3 weeks ago
Following the theme of AI research and safety, Aric Floyd talks about how some Large Language Models might follow the all too human trait of sandbagging - "lying" about their true capabilities.

AI Sandbagging Paper: apolloresearch.ai/research/scheming-reasoning-evaluations

Computerphile is supported by Jane Street. Learn more about them (and exciting career opportunities) at: jane-st.co/computerphile

This video was filmed and edited by Sean Riley.

Computerphile is a sister project to Brady Haran's Numberphile. More at bradyharanblog.com
AI Sandbagging - ComputerphileWorld Foundation Models - ComputerphileQuantum Simulation & Nature - ComputerphileAutomated Mathematical Proofs - ComputerphileL Systems : Creating Plants from Simple Rules - ComputerphileAlternative Uses for Blockchain - ComputerphileFoundations of Data Visualisation - ComputerphileDefining Harm for Ai Systems - ComputerphileGenerative AIs Greatest Flaw - ComputerphileDo Computer Scientists Prefer Tea or Coffee? (Microphone Sound Check Question 2025)  - ComputerphileMalware and Machine Learning - ComputerphileJavascript Card Trick - Computerphile
Computerphile |

AI Sandbagging - Computerphile

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER