Mutually Assured AI Malfunction [Dan Hendrycks] @MachineLearningStreetTalk
Mutually Assured AI Malfunction [Dan Hendrycks]  @MachineLearningStreetTalk
Uploaded August 2025 | Updated September 2026, 1 week ago
Deep dive with Dan Hendrycks, a leading AI safety researcher and co-author of the "Superintelligence Strategy" paper with former Google CEO Eric Schmidt and Scale AI CEO Alexandr Wang.

*** SPONSOR MESSAGES
Gemini CLI is an open-source AI agent that brings the power of Gemini directly into your terminal - github.com/google-gemini/gemini-cli

Prolific: Quality data. From real people. For faster breakthroughs.
prolific.com/mlst?utm_campaign=98404559-MLST&utm_source=youtube&utm_medium=podcast&utm_content=script-gen
***

Hendrycks argues that society is making a fundamental mistake in how it views artificial intelligence. We often compare AI to transformative but ultimately manageable technologies like electricity or the internet. He contends a far better and more realistic analogy is nuclear technology. Like nuclear power, AI has the potential for immense good, but it is also a dual-use technology that carries the risk of unprecedented catastrophe.

The Problem with an AI "Manhattan Project":

A popular idea is for the U.S. to launch a "Manhattan Project" for AI—a secret, all-out government race to build a superintelligence before rivals like China. Hendrycks argues this strategy is deeply flawed and dangerous for several reasons:

- It wouldn’t be secret. You cannot hide a massive, heat-generating data center from satellite surveillance.

- It would be destabilizing. A public race would alarm rivals, causing them to start their own desperate, corner-cutting projects, dramatically increasing global risk.

- It’s vulnerable to sabotage. An AI project can be crippled in many ways, from cyberattacks that poison its training data to physical attacks on its power plants. This is what the paper refers to as a "maiming attack."

This vulnerability leads to the paper's central concept: Mutual Assured AI Malfunction (MAIM). This is the AI-era version of the nuclear-era's Mutual Assured Destruction (MAD). In this dynamic, any nation that makes an aggressive, destabilizing bid for a world-dominating AI must expect its rivals to sabotage the project to ensure their own survival.

This deterrence, Hendrycks argues, is already the default reality we live in.

A Better Strategy: The Three Pillars
Instead of a reckless race, the paper proposes a more stable, three-part strategy modeled on Cold War principles:

- Deterrence: Acknowledge the reality of MAIM. The goal should not be to "win" the race to superintelligence, but to deter anyone from starting such a race in the first place through the credible threat of sabotage.

- Nonproliferation: Just as we work to keep fissile materials for nuclear bombs out of the hands of terrorists and rogue states, we must control the key inputs for catastrophic AI. The most critical input is advanced AI chips (GPUs). Hendrycks makes the powerful claim that building cutting-edge GPUs is now more difficult than enriching uranium, making this strategy viable.

- Competitiveness: The race between nations like the U.S. and China should not be about who builds superintelligence first. Instead, it should be about who can best use existing AI to build a stronger economy, a more effective military, and more resilient supply chains (for example, by manufacturing more chips domestically).

Dan says the stakes are high if we fail to manage this transition:

- Erosion of Control: Society becomes so dependent on AI systems for its economy and military that we can no longer turn them off without risking total collapse. We become "passengers in an autonomous economy" where humans are no longer in the driver's seat.
- Intelligence Recursion: This is the scenario where AI becomes capable of improving itself, kicking off a rapid "intelligence explosion" that could race ahead of any human attempts at control.
- Worthless Labor: When AI can perform most human cognitive tasks, the economic value of human labor could plummet, leading to massive societal instability and stripping people of their bargaining power.

Hendrycks maintains that while the risks are existential, the future is not set.

TOC:
1 Measuring the Beast [00:00:00]
2 Defining the Beast [00:11:34]
3 The Core Strategy [00:38:20]
4 Ideological Battlegrounds [00:53:12]
5 Mechanisms of Control [01:34:45]

TRANSCRIPT:
app.rescript.info/public/share/cOKcz4pWRPjh7BTIgybd7PUr_vChUaY6VQW64No8XMs

REFS:
Superintelligence Strategy
arxiv.org/abs/2503.05628

Humanity's Last Exam
arxiv.org/abs/2501.14249

Enigma Eval
arxiv.org/abs/2502.08859

Natural Selection Favors AIs Over Humans
arxiv.org/abs/2303.16200

Utility Engineering
arxiv.org/abs/2502.08640

Unsolved Problems in ML Safety [Emergence ref]
arxiv.org/abs/2109.13916

"Situational Awareness" by Leopold Aschenbrenner
situational-awareness.ai

Large Language Models and Emergence [Krakauer]
arxiv.org/abs/2506.11135

Fractured Entangled Representations [Stanley/Kumar]
arxiv.org/pdf/2505.11581
Mutually Assured AI Malfunction [Dan Hendrycks]Your Brain Is a Prediction Machine, Not a Processor — Karl FristonChatGPT is overconfident (Anil Ananthaswamy)People wanted faster horses - Jeff CluneThe Mathematical Foundations of Intelligence [Professor Yi Ma]Why Humans Are Still Powering AI [Sponsored] - Phelim BradleyWhen We Asked GPT-4 to Solve a Problem, It Chose Blackmail — Sara Saab & Enzo BlindowOptimize GPU performance for AI - Prof. Gennady PekhimenkoYour Brain Doesnt Command Your Body. It Predicts It. [Max Bennett]Impostor Intelligence?Language is a set of pointers to internal simulationsBuild Specialist LLMs Like It’s 2019 (Randall Balestriero)
Machine Learning Street Talk |

Mutually Assured AI Malfunction [Dan Hendrycks]

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER