Uploaded October 2024 | Updated September 2026, 1 week ago
Ben Goertzel discusses AGI development, transhumanism, and the potential societal impacts of superintelligent AI. He predicts human-level AGI by 2029 and argues that the transition to superintelligence could happen within a few years after. Goertzel explores the challenges of AI regulation, the limitations of current language models, and the need for neuro-symbolic approaches in AGI research. He also addresses concerns about resource allocation and cultural perspectives on transhumanism.
TOC:
[00:00:00] AGI Timeline Predictions and Development Speed
[00:00:45] Limitations of Language Models in AGI Development
[00:02:18] Current State and Trends in AI Research and Development
[00:09:02] Emergent Reasoning Capabilities and Limitations of LLMs
[00:18:15] Neuro-Symbolic Approaches and the Future of AI Systems
[00:20:00] Evolutionary Algorithms and LLMs in Creative Tasks
[00:21:25] Symbolic vs. Sub-Symbolic Approaches in AI
[00:28:05] Language as Internal Thought and External Communication
[00:30:20] AGI Development and Goal-Directed Behavior
[00:35:51] Consciousness and AI: Expanding States of Experience
[00:48:50] AI Regulation: Challenges and Approaches
[00:55:35] Challenges in AI Regulation
[00:59:20] AI Alignment and Ethical Considerations
[01:09:15] AGI Development Timeline Predictions
[01:12:40] OpenCog Hyperon and AGI Progress
[01:17:48] Transhumanism and Resource Allocation Debate
[01:20:12] Cultural Perspectives on Transhumanism
[01:23:54] AGI and Post-Scarcity Society
[01:31:35] Challenges and Implications of AGI Development
New! PDF Show notes: dropbox.com/scl/fi/fyetzwgoaf70gpovyfc4x/BenGoertzel.pdf?rlkey=pze5dt9vgf01tf2wip32p5hk5&st=svbcofm3&dl=0
Refs:
00:00:15 Ray Kurzweil's AGI timeline prediction, Ray Kurzweil, en.wikipedia.org/wiki/Technological_singularity
00:01:45 Ben Goertzel: SingularityNET founder, Ben Goertzel, singularitynet.io
00:02:35 AGI Conference series, AGI Conference Organizers, agi-conf.org/2024
00:03:55 Ben Goertzel's contributions to AGI, Wikipedia contributors, en.wikipedia.org/wiki/Ben_Goertzel
00:11:05 Chain-of-Thought prompting, Subbarao Kambhampati, arxiv.org/abs/2405.04776
00:11:35 Algorithmic information content, Pieter Adriaans, https://plato.stanford.edu/entries/information-entropy/
00:12:10 Turing completeness in neural networks, Various contributors, https://plato.stanford.edu/entries/turing-machine/
00:16:15 AlphaGeometry: AI for geometry problems, Trieu, Li, et al., nature.com/articles/s41586-023-06747-5
00:18:25 Shane Legg and Ben Goertzel's collaboration, Shane Legg, en.wikipedia.org/wiki/Shane_Legg
00:20:00 Evolutionary algorithms in music generation, Yanxu Chen, arxiv.org/html/2409.03715v1
00:22:00 Peirce's theory of semiotics, Charles Sanders Peirce, https://plato.stanford.edu/entries/peirce-semiotics/
00:28:10 Chomsky's view on language, Noam Chomsky, chomsky.info/1983____
00:34:05 Greg Egan's 'Diaspora', Greg Egan, amazon.co.uk/Diaspora-post-apocalyptic-thriller-perfect-MIRROR/dp/0575082097
00:40:35 'The Consciousness Explosion', Ben Goertzel & Gabriel Axel Montes, amazon.com/Consciousness-Explosion-Technological-Experiential-Singularity/dp/B0D8C7QYZD
00:41:55 Ray Kurzweil's books on singularity, Ray Kurzweil, amazon.com/Singularity-Near-Humans-Transcend-Biology/dp/0143037889
00:50:50 California AI regulation bills, California State Senate, sd18.senate.ca.gov/news/senate-unanimously-approves-senator-padillas-artificial-intelligence-package
00:56:40 Limitations of Compute Thresholds, Sara Hooker, arxiv.org/abs/2407.05694
00:56:55 'Taming Silicon Valley', Gary F. Marcus, penguinrandomhouse.com/books/768076/taming-silicon-valley-by-gary-f-marcus
01:09:15 Kurzweil's AGI prediction update, Ray Kurzweil, theguardian.com/technology/article/2024/jun/29/ray-kurzweil-google-ai-the-singularity-is-nearer
01:14:45 OpenCog Hyperon framework, Ben Goertzel et al., arxiv.org/abs/2310.18318
01:18:25 Malnutrition in Ethiopia, Abriham Shiferaw Areba, frontiersin.org/journals/nutrition/articles/10.3389/fnut.2024.1403591/full
01:18:40 Transhumanism ethical debate, Nick Bostrom, nickbostrom.com/papers/history.pdf
Ben Goertzel discusses AGI development, transhumanism, and the potential societal impacts of superintelligent AI. He predicts human-level AGI by 2029 and argues that the transition to superintelligence could happen within a few years after. Goertzel explores the challenges of AI regulation, the limitations of current language models, and the need for neuro-symbolic approaches in AGI research. He also addresses concerns about resource allocation and cultural perspectives on transhumanism.
TOC:
[00:00:00] AGI Timeline Predictions and Development Speed
[00:00:45] Limitations of Language Models in AGI Development
[00:02:18] Current State and Trends in AI Research and Development
[00:09:02] Emergent Reasoning Capabilities and Limitations of LLMs
[00:18:15] Neuro-Symbolic Approaches and the Future of AI Systems
[00:20:00] Evolutionary Algorithms and LLMs in Creative Tasks
[00:21:25] Symbolic vs. Sub-Symbolic Approaches in AI
[00:28:05] Language as Internal Thought and External Communication
[00:30:20] AGI Development and Goal-Directed Behavior
[00:35:51] Consciousness and AI: Expanding States of Experience
[00:48:50] AI Regulation: Challenges and Approaches
[00:55:35] Challenges in AI Regulation
[00:59:20] AI Alignment and Ethical Considerations
[01:09:15] AGI Development Timeline Predictions
[01:12:40] OpenCog Hyperon and AGI Progress
[01:17:48] Transhumanism and Resource Allocation Debate
[01:20:12] Cultural Perspectives on Transhumanism
[01:23:54] AGI and Post-Scarcity Society
[01:31:35] Challenges and Implications of AGI Development
New! PDF Show notes: dropbox.com/scl/fi/fyetzwgoaf70gpovyfc4x/BenGoertzel.pdf?rlkey=pze5dt9vgf01tf2wip32p5hk5&st=svbcofm3&dl=0
Refs:
00:00:15 Ray Kurzweil's AGI timeline prediction, Ray Kurzweil, en.wikipedia.org/wiki/Technological_singularity
00:01:45 Ben Goertzel: SingularityNET founder, Ben Goertzel, singularitynet.io
00:02:35 AGI Conference series, AGI Conference Organizers, agi-conf.org/2024
00:03:55 Ben Goertzel's contributions to AGI, Wikipedia contributors, en.wikipedia.org/wiki/Ben_Goertzel
00:11:05 Chain-of-Thought prompting, Subbarao Kambhampati, arxiv.org/abs/2405.04776
00:11:35 Algorithmic information content, Pieter Adriaans, https://plato.stanford.edu/entries/information-entropy/
00:12:10 Turing completeness in neural networks, Various contributors, https://plato.stanford.edu/entries/turing-machine/
00:16:15 AlphaGeometry: AI for geometry problems, Trieu, Li, et al., nature.com/articles/s41586-023-06747-5
00:18:25 Shane Legg and Ben Goertzel's collaboration, Shane Legg, en.wikipedia.org/wiki/Shane_Legg
00:20:00 Evolutionary algorithms in music generation, Yanxu Chen, arxiv.org/html/2409.03715v1
00:22:00 Peirce's theory of semiotics, Charles Sanders Peirce, https://plato.stanford.edu/entries/peirce-semiotics/
00:28:10 Chomsky's view on language, Noam Chomsky, chomsky.info/1983____
00:34:05 Greg Egan's 'Diaspora', Greg Egan, amazon.co.uk/Diaspora-post-apocalyptic-thriller-perfect-MIRROR/dp/0575082097
00:40:35 'The Consciousness Explosion', Ben Goertzel & Gabriel Axel Montes, amazon.com/Consciousness-Explosion-Technological-Experiential-Singularity/dp/B0D8C7QYZD
00:41:55 Ray Kurzweil's books on singularity, Ray Kurzweil, amazon.com/Singularity-Near-Humans-Transcend-Biology/dp/0143037889
00:50:50 California AI regulation bills, California State Senate, sd18.senate.ca.gov/news/senate-unanimously-approves-senator-padillas-artificial-intelligence-package
00:56:40 Limitations of Compute Thresholds, Sara Hooker, arxiv.org/abs/2407.05694
00:56:55 'Taming Silicon Valley', Gary F. Marcus, penguinrandomhouse.com/books/768076/taming-silicon-valley-by-gary-f-marcus
01:09:15 Kurzweil's AGI prediction update, Ray Kurzweil, theguardian.com/technology/article/2024/jun/29/ray-kurzweil-google-ai-the-singularity-is-nearer
01:14:45 OpenCog Hyperon framework, Ben Goertzel et al., arxiv.org/abs/2310.18318
01:18:25 Malnutrition in Ethiopia, Abriham Shiferaw Areba, frontiersin.org/journals/nutrition/articles/10.3389/fnut.2024.1403591/full
01:18:40 Transhumanism ethical debate, Nick Bostrom, nickbostrom.com/papers/history.pdf
![Can Outsourcing Thinking Make Us Dumber? [Prof. David Krakauer]
Prof. David Krakauer, President of the Santa Fe Institute argues that we are fundamentally confusing knowledge with intelligence, especially when it comes to AI.
He defines true intelligence as the ability to do more with less—to solve novel problems with limited information. This is contrasted with current AI models, which he describes as doing less with more; they require astounding amounts of data to perform tasks that dont necessarily demonstrate true understanding or adaptation. He humorously calls this really shit programming.
David challenges the popular notion of emergence in Large Language Models (LLMs). He explains that the tech communitys definition—seeing a sudden jump in a models ability to perform a task like three-digit math—is superficial. True emergence, from a complex systems perspective, involves a fundamental change in the systems internal organization, allowing for a new, simpler, and more powerful level of description. He gives the example of moving from tracking individual water molecules to using the elegant laws of fluid dynamics. For LLMs to be truly emergent, wed need to see them develop new, efficient internal representations, not just get better at memorizing patterns as they scale.
Drawing on his background in evolutionary theory, David explains that systems like brains, and later, culture, evolved to process information that changes too quickly for genetic evolution to keep up. He calls culture evolution at light speed because it allows us to store our accumulated knowledge externally (in books, tools, etc.) and build upon it without corrupting the original.
This leads to his concept of exbodiment, where we outsource our cognitive load to the world through things like maps, abacuses, or even language itself.
We create these external tools, internalize the skills they teach us, improve them, and create a feedback loop that enhances our collective intelligence.
However, he ends with a warning. While technology has historically complemented our deficient abilities, modern AI presents a new danger. Because we have an evolutionary drive to conserve energy, we will inevitably outsource our thinking to AI if we can. He fears this is already leading to a diminution and dilution of human thought and creativity. Just as our muscles atrophy without use, he argues our brains will too, and we risk becoming mentally dependent on these systems.
RESCRIPT LINK (interactive transcript):
https://app.rescript.info/public/share/nCL8fdE_m3J6fA3SHlreCADpkWpbXaEF1JQ14z6N7Y8
TOC:
[00:00:00] Intelligence: Doing more with less
[00:02:10] Why brains evolved: The limits of evolution
[00:05:18] Culture as evolution at light speed
[00:08:11] True meaning of emergence: More is Different
[00:10:41] Why LLM capabilities are not true emergence
[00:15:10] What real emergence would look like in AI
[00:19:24] Symmetry breaking: Physics vs. Life
[00:23:30] Two types of emergence: Knowledge In vs. Out
[00:26:46] Causality, agency, and coarse-graining
[00:32:24] Exbodiment: Outsourcing thought to objects
[00:35:05] Collective intelligence & the boundary of the mind
[00:39:45] Mortal vs. Immortal forms of computation
[00:42:13] The risk of AI: Atrophy of human thought
David Krakauer
President and William H. Miller Professor of Complex Systems
https://www.santafe.edu/people/profile/david-krakauer
REFS:
Large Language Models and Emergence: A Complex Systems Perspective
David C. Krakauer, John W. Krakauer, Melanie Mitchell
https://arxiv.org/abs/2506.11135
Filmed at the Diverse Intelligences Summer Institute:
https://disi.org/ Can Outsourcing Thinking Make Us Dumber? [Prof. David Krakauer]](https://i.ytimg.com/vi/jXa8dHzgV8U/mqdefault.jpg)
![AI HAS A BODY PROBLEM... [Dr. Maxwell Ramstead]
This episode features Dr. Maxwell Ramstead and Jason Fox both from Noumenal discussing why current AI approaches fall short for real-world applications and whats needed for true physical AI.
The guests argue that todays AI systems, including large language models, are fundamentally stuck in data space - they only process patterns in data rather than understanding the physical world that generates that data.
Maxwell uses Platos Cave as a powerful metaphor: like prisoners seeing only shadows on a wall, LLMs interact with representations of reality (text, images) rather than reality itself.
Rather than building monolithic models, Noumenal is creating a compositional system - essentially a marketplace of models where specialized AI components can be dynamically combined and deployed to robots.
Sponsor messages:
Tufa AI Labs are hiring for ML Engineers and a Chief Scientist in Zurich/SF. They are top of the ARCv2 leaderboard!
https://tufalabs.ai/
TRANSCRIPT:
http://app.rescript.info/public/share/DVaCBhYF3Y-1Q2kA2N5-YkV4kOGBNZc5KLrDwntLznU
https://www.noumenal.ai/
https://x.com/mjdramstead
https://scholar.google.ca/citations?user=ILpGOMkAAAAJ&hl=fr
https://x.com/jasongfox?lang=en-GB
TOC
Opening & Context
00:00:00 - Opening Hook: Why Create a Physical AI Company?
00:01:59 - Sponsor: Tufa AI Labs
00:02:30 - Guest Introductions: Maxwell Ramstead & Jason Fox
00:05:18 - Noumenal Background
Core Problems with Current AI
00:09:30 - The Embodiment Problem: Why Bodies Matter
00:10:15 - LLMs Lack Physical Grounding
00:12:00 - AI Stuck in Platos Cave
00:16:15 - Language as Wrong Compression for Physics
00:17:22 - The Exhaustion of Static Datasets
00:19:54 - Humans as the Grounding for LLMs
Philosophical Foundations
00:28:00 - Fractured vs. Deep Understanding
00:32:15 - Defining Real: When You Bump Into Things
00:37:00 - Emergence: Weak vs. Strong Causal Power
00:41:45 - The Free Energy Principle Explained
00:44:15 - Constraints: How the Universe Builds Things
Objects, Intelligence & Grounding
00:46:15 - What Is an Object? From Data to Physics
00:51:00 - Learning Primitives & Predictive Grip
00:55:58 - There Is No General Intelligence
01:00:15 - The Human-AI Feedback Loop
01:03:08 - The Irony of LLM Specialization
01:06:05 - LLMs as Tools vs. Autonomous Agents
01:08:45 - Hallucinating Capabilities: The Third Leg Problem
The Noumenal Solution
01:09:00 - A Marketplace of Specialized Models
01:13:45 - Dynamic Skill Loading: Phone a Friend
01:16:15 - Learning from Brain Evolution
01:18:00 - Business Model Critique: Why OpenAI Wont Work
01:22:30 - The Physical Dataset Problem
Implementation & Future
01:22:30 - Community-Driven Data Collection
01:24:45 - Jim Fans Physical Turing Test
01:26:30 - Enterprise vs. Consumer Models
01:27:22 - Docker for Robotics: The Technical Architecture
01:30:12 - Reproducibility in Learning Systems
01:32:00 - Closing Thoughts AI HAS A BODY PROBLEM... [Dr. Maxwell Ramstead]](https://i.ytimg.com/vi/jsIt2sTB_vo/mqdefault.jpg)
![The Fabric of Knowledge - David Spivak
MLST is sponsored by Brave:
The Brave Search API covers over 20 billion webpages, built from scratch without Big Tech biases or the recent extortionate price hikes on search API access. Perfect for AI model training and retrieval augmentated generation. Try it now - get 2,000 free queries monthly at http://brave.com/api.
David Spivak, a mathematician at MITs Topos Institute known for his work in applied category theory, talks with Tim Scarfe about the nature of intelligence, creativity, and knowledge itself.
Spivak explains category theory in surprisingly concrete terms. Categories are about systems of relationships not just collections of things, but how those things relate to each other. Functors map one system of relationships to another, like how counting connects the world of sets to the world of numbers. He argues category theorys value lies in making the obvious things mathematically precise, which sounds trivial until you realize how much of mathematics depends on shared but unspoken assumptions.
The conversation moves to collective intelligence and sense-making. Drawing on Mike Levins claim that all intelligence is collective intelligence, Spivak describes sense-making as a process of accounting like balancing a checkbook, where different perspectives contribute until the books settle and understanding stabilizes. This applies at every scale, from neurons communicating in a shared language to two people trying to agree on what category theory means for machine learning.
Where things get genuinely interesting is Spivaks take on creativity and open-endedness. He pushes back on the idea that Karl Fristons prediction error minimization framework captures everything interesting about intelligence. His counterexample: a kid shoveling sand in a sandbox who cries when pulled away. Theres something about care and engagement that doesnt obviously reduce to prediction error. Questions, he argues, are more important than answers the act of questioning creates a spaciousness where real insight can arise.
On AI, Spivak is measured but candid. He thinks current approaches are kicking the ball really hard without thinking about where the ball needs to go. He worries about optimization without understanding what were optimizing for. The discussion covers embodiment and how physical experience shapes abstract thought, the role of written language in transmitting knowledge across generations, and whether the intelligence explosion is an extension of evolutionary processes or something genuinely new.
TIMESTAMPS:
00:00:00 Introduction to Category Theory
00:04:40 Collective Intelligence and Sense-Making
00:09:54 Embodiment and Physical Concepts in Knowledge
00:16:23 Creativity, Open-Endedness, and Care
00:25:46 Modeling Creativity and the Role of Questioning
00:36:04 Evolution, Optimization, and AI
00:44:14 Written Language and Knowledge Transmission
REFERENCES:
person:
[00:00:00] David Spivak - Personal Page
http://www.dspivak.net/
[00:04:40] Mike Levin - Google Scholar
https://scholar.google.com/citations?user=luouyakAAAAJ&hl=en
reference:
[00:00:00] MIT Category Theory Lectures by David Spivak
https://www.youtube.com/watch?v=UusLtx9fIjs
[00:00:00] Spotify Podcast Version
https://podcasters.spotify.com/pod/show/machinelearningstreettalk/episodes/The-Fabric-of-Knowledge David-Spivak-e2o220h
[00:25:46] Herbert Simon - Satisficing and Bounded Rationality
https://plato.stanford.edu/entries/bounded-rationality/
[00:36:04] Eric Smith - Complexity and Early Life
https://www.youtube.com/watch?v=SpJZw-68QyE
book:
[00:09:54] Karl Friston - Active Inference
https://direct.mit.edu/books/oa-monograph/5299/Active-InferenceThe-Free-Energy-Principle-in-Mind
[00:36:04] Richard Dawkins - The Selfish Gene
https://amzn.to/3X73X8w
[00:36:04] Carl Sagan - The Cosmos Knowing Itself
https://amzn.to/3XhPruK
paper:
[00:36:04] DeepMind - Open-Ended Systems Paper
https://arxiv.org/abs/2406.04268
LINKS:
Full Transcript: https://app.rescript.info/share/3e4986b488817a53549bf134e2a5695a
Download PDF transcript: https://app.rescript.info/api/public/sessions/069a33878cbdd032/pdf The Fabric of Knowledge - David Spivak](https://i.ytimg.com/vi/ju17bM9p2RU/mqdefault.jpg)
![One Step Closer to the Star Trek Voice AI Assistant!
Will Williams is CTO of Speechmatics in Cambridge. In this sponsored episode - he shares deep technical insights into modern speech recognition technology and system architecture. The episode covers several key technical areas:
Will Williams is CTO of Speechmatics, the Cambridge-based speech recognition company. Tim has used their API for years to caption MLST episodes, and this conversation happened in their offices which explains the live demo at the start where an AI moderates a political debate clip in real time.
The technical meat covers how Speechmatics builds production ASR systems. Their approach is hybrid: self-supervised pre-training on unlabeled audio gets them comparable accuracy to fully supervised systems like Whisper, but with roughly 100x less labeled data. Williams explains why this matters for scaling to low-resource languages where you simply dont have thousands of hours of human-transcribed speech.
The architecture discussion is detailed. Their system runs multiple operating points with different latency-accuracy tradeoffs. They pad latency up to 1.8 seconds to keep the user experience consistent rather than optimizing for raw speed. Decoding uses lattices with language model integration, which lets them rescore hypotheses and handle things like proper nouns and domain-specific vocabulary without retraining the acoustic model.
Diarization figuring out who said what comes up repeatedly. Williams calls it harder than ASR itself, partly because speaker embeddings get corrupted by acoustic environments and partly because cross-talk creates genuinely ambiguous boundaries. Theyre pushing hard on implicit source separation but the problem remains open.
The conversation also covers their testing infrastructure (mirrored production traffic catches edge cases that unit tests miss), why they resist customer-specific fine-tuning (it fragments the model and makes global improvements harder), and Williams critique of PyTorch memory management in production settings. He argues for more direct memory allocation rather than letting the framework handle it, which is a practical concern when youre serving models at scale.
Featuring: Will Williams (CTO, Speechmatics) and Tim Scarfe.
TIMESTAMPS:
00:00:00 ASR and diarization fundamentals
00:05:25 Real-time conversational AI architecture
00:09:21 Neural network streaming and multi-modal integration
00:12:49 Enterprise voice AI and real-time translation
00:20:00 Production deployment and testing infrastructure
00:29:38 Model architecture and latency-accuracy tradeoffs
00:35:40 Lattice-based decoding and language model integration
00:44:00 ASR performance metrics and real-world evaluation
00:51:30 Ethics and privacy in speech technology
01:00:50 Self-supervised learning and low-resource languages
01:11:00 Feature engineering to automated ML
01:21:00 Infrastructure scaling and PyTorch critique
01:35:00 Future of conversational AI and Ursa 2
REFERENCES:
paper:
[00:00:05] Speechmatics PDF shownotes
https://www.dropbox.com/scl/fi/d94b1jcgph9o8au8shdym/Speechmatics.pdf?rlkey=bi55wvktzomzx0y5sic6jz99y&st=6qwofv8t&dl=0
[00:10:09] GFlowNets
https://arxiv.org/abs/2106.04399
[01:35:00] Ursa 2 model
https://www.speechmatics.com/company/articles-and-news/ursa-2-elevating-speech-recognition-across-52-languages
company:
[00:01:15] Speechmatics
https://www.speechmatics.com/
person:
[00:01:32] Will Williams
https://x.com/wjwwilliams
LINKS:
Full Transcript: https://app.rescript.info/share/c6887b6d7b214f93daad1c18d70e2eb6
Download PDF transcript: https://app.rescript.info/api/public/sessions/abeef42b31287680/pdf
Will Williams, CTO, Speechmatics
https://x.com/wjwwilliams One Step Closer to the Star Trek Voice AI Assistant!](https://i.ytimg.com/vi/k6eXkBtYIHg/mqdefault.jpg)
![AIs can now imagine video games in real-time
Ashley Edwards, who was working at DeepMind when she co-authored the Genie paper and is now at Runway, covered several key aspects of the Genie AI system and its applications in video generation, robotics, and game creation.
MLST is sponsored by Brave:
The Brave Search API covers over 20 billion webpages, built from scratch without Big Tech biases or the recent extortionate price hikes on search API access. Perfect for AI model training and retrieval augmentated generation. Try it now - get 2,000 free queries monthly at http://brave.com/api.
Genies approach to learning interactive environments, balancing compression and fidelity.
The use of latent action models and VQE models for video processing and tokenization.
Challenges in maintaining action consistency across frames and integrating text-to-image models.
Evaluation metrics for AI-generated content, such as FID and PS&R diff metrics.
The discussion also explored broader implications and applications:
The potential impact of AI video generation on content creation jobs.
Applications of Genie in game generation and robotics.
The use of foundation models in robotics and the differences between internet video data and specialized robotics data.
Challenges in mapping AI-generated actions to real-world robotic actions.
Ashley Edwards: https://ashedwards.github.io/
TOC (*) are best bits
00:00:00 1. Intro to Genie & Brave Search API: Trade-offs & limitations *
00:02:26 2. Genies Architecture: Latent action, VQE, video processing *
00:05:06 3. Genies Constraints: Frame consistency & image model integration
00:07:26 4. Evaluation: FID, PS&R diff metrics & latent induction methods
00:09:44 5. AI Video Gen: Content creation impact, depth & parallax effects
00:11:39 6. Model Scaling: Training data impact & computational trade-offs
00:13:50 7. Game & Robotics Apps: Gamification & action mapping challenges *
00:16:16 8. Robotics Foundation Models: Action space & data considerations *
00:19:18 9. Mask-GPT & Video Frames: Real-time optimization, RL from videos
00:20:34 10. Research Challenges: AI value, efficiency vs. quality, safety
00:24:20 11. Future Dev: Efficiency improvements & fine-tuning strategies
Refs:
1. Genie (learning interactive environments from videos) / Ashley and DM collegues [00:01]
https://arxiv.org/abs/2402.15391
2. VQ-VAE (Vector Quantized Variational Autoencoder) / Aaron van den Oord, Oriol Vinyals, Koray Kavukcuoglu [02:43]
https://arxiv.org/abs/1711.00937
3. FID (Fréchet Inception Distance) metric / Martin Heusel et al. [07:37]
https://arxiv.org/abs/1706.08500
4. PS&R (Precision and Recall) metric / Mehdi S. M. Sajjadi et al. [08:02]
https://arxiv.org/abs/1806.00035
5. Vision Transformer (ViT) architecture / Alexey Dosovitskiy et al. [12:14]
https://arxiv.org/abs/2010.11929
6. Genie (robotics foundation models) / Google DeepMind [17:34]
https://deepmind.google/research/publications/60474/
7. Chelsea Finns lab work on robotics datasets / Chelsea Finn [17:38]
https://ai.stanford.edu/~cbfinn/
8. Imitation from observation in reinforcement learning / YuXuan Liu [20:58]
https://arxiv.org/abs/1707.03374
9. Waymos autonomous driving technology / Waymo [22:38]
https://waymo.com/
10. Gen3 model release by Runway / Runway [23:48]
https://runwayml.com/
11. Classifier-free guidance technique / Jonathan Ho and Tim Salimans [24:43]
https://arxiv.org/abs/2207.12598 AIs can now imagine video games in real-time](https://i.ytimg.com/vi/kbt0ZFoI2Hc/mqdefault.jpg)

![Neural Networks Are Elastic Origami! [Prof. Randall Balestriero]
Professor Randall Balestriero joins us to discuss neural network geometry, spline theory, and emerging phenomena in deep learning, based on research presented at ICML. Topics include the delayed emergence of adversarial robustness in neural networks (grokking), geometric interpretations of neural networks via spline theory, and challenges in reconstruction learning. We also cover geometric analysis of Large Language Models (LLMs) for toxicity detection and the relationship between intrinsic dimensionality and model control in RLHF.
SPONSOR MESSAGES:
***
CentML offers competitive pricing for GenAI model deployment, with flexible options to suit a wide range of models, from small to large-scale deployments.
https://centml.ai/pricing/
Tufa AI Labs is a brand new research lab in Zurich started by Benjamin Crouzier focussed on o-series style reasoning and AGI. Are you interested in working on reasoning, or getting involved in their events?
Goto https://tufalabs.ai/
***
Show notes and transcript: https://www.dropbox.com/scl/fi/3lufge4upq5gy0ug75j4a/RANDALLSHOW.pdf?rlkey=nbemgpa0jhawt1e86rx7372e4&dl=0
TOC:
[00:00:00] Introduction
1. Neural Network Geometry and Spline Theory
[00:01:41] 1.1 Neural Network Geometry and Spline Theory
[00:07:41] 1.2 Deep Networks Always Grok
[00:11:39] 1.3 Grokking and Adversarial Robustness
[00:16:09] 1.4 Double Descent and Catastrophic Forgetting
2. Reconstruction Learning
[00:18:49] 2.1 Reconstruction Learning
[00:24:15] 2.2 Frequency Bias in Neural Networks
3. Geometric Analysis of Neural Networks
[00:29:02] 3.1 Geometric Analysis of Neural Networks
[00:34:41] 3.2 Adversarial Examples and Region Concentration
4. LLM Safety and Geometric Analysis
[00:40:05] 4.1 LLM Safety and Geometric Analysis
[00:46:11] 4.2 Toxicity Detection in LLMs
[00:52:24] 4.3 Intrinsic Dimensionality and Model Control
[00:58:07] 4.4 RLHF and High-Dimensional Spaces
5. Conclusion
[01:02:13] 5.1 Neural Tangent Kernel
[01:08:07] 5.2 Conclusion
REFS:
[00:01:35] Balestriero/Humayun – Deep network geometry & input space partitioning
https://arxiv.org/html/2408.04809v1
[00:03:55] Balestriero & Paris – Linking deep networks to adaptive spline operators
https://proceedings.mlr.press/v80/balestriero18b/balestriero18b.pdf
[00:13:55] Song et al. – Gradient-based white-box adversarial attacks
https://arxiv.org/abs/2012.14965
[00:16:05] Humayun, Balestriero & Baraniuk – Grokking phenomenon & emergent robustness
https://arxiv.org/abs/2402.15555
[00:18:25] Humayun – Training dynamics & double descent via linear region evolution
https://arxiv.org/abs/2310.12977
[00:20:15] Balestriero – Power diagram partitions in DNN decision boundaries
https://arxiv.org/abs/1905.08443
[00:23:00] Frankle & Carbin – Lottery Ticket Hypothesis for network pruning
https://arxiv.org/abs/1803.03635
[00:24:00] Belkin et al. – Double descent phenomenon in modern ML
https://arxiv.org/abs/1812.11118
[00:25:55] Balestriero et al. – Batch normalization’s regularization effects
https://arxiv.org/pdf/2209.14778
[00:29:35] EU – EU AI Act 2024 with compute restrictions
https://www.lw.com/admin/upload/SiteAttachments/EU-AI-Act-Navigating-a-Brave-New-World.pdf
[00:39:30] Humayun, Balestriero & Baraniuk – SplineCam: Visualizing deep network geometry
https://openaccess.thecvf.com/content/CVPR2023/papers/Humayun_SplineCam_Exact_Visualization_and_Characterization_of_Deep_Network_Geometry_and_CVPR_2023_paper.pdf
[00:40:40] Carlini – Trade-offs between adversarial robustness and accuracy
https://arxiv.org/abs/1902.06705
[00:44:55] Balestriero & LeCun – Limitations of reconstruction-based learning methods
https://raw.githubusercontent.com/mlresearch/v235/main/assets/balestriero24b/balestriero24b.pdf
[00:47:20] Balestriero & LeCun – Spectral analysis of neural network learning
https://proceedings.neurips.cc/paper_files/paper/2022/file/aa56c74513a5e35768a11f4e82dd7ffb-Paper-Conference.pdf
[00:49:45] He et al. – MAE: Masked Autoencoders for self-supervised learning
https://arxiv.org/abs/2111.06377
[00:54:50] Balestriero et al. – Geometric analysis of LLM layers for toxicity detection
https://arxiv.org/abs/2309.12312
[00:59:35] Balestriero et al. – Superior toxicity detection via geometric features
https://arxiv.org/html/2312.01648v2
[01:04:45] UofT ML – Self-attention control & context length effects
https://arxiv.org/abs/2310.04444
[01:11:55] Roberts – Foundations of deep learning theory
https://arxiv.org/abs/2106.10165
[01:15:40] Balestriero & Cha – Kolmogorov GAM Networks via spline partition theory
https://arxiv.org/pdf/2501.00704
[01:16:40] Various – Graph Kolmogorov-Arnold Networks (GKAN) extension
https://www.nature.com/articles/s41598-024-85083-8 Neural Networks Are Elastic Origami! [Prof. Randall Balestriero]](https://i.ytimg.com/vi/l3O2J3LMxqI/mqdefault.jpg)



