Uploaded April 2026 | Updated September 2026, 2 weeks ago
We’ve been on a bit of a mini World Models series over the last quarter: from introducing the topic with Yi Tay, to exploring Marble with World Labs’ Fei-Fei Li and Justin Johnson, to previewing World Models learned from massive1 gaming datasets with General Intuition’s Pim de Witte (who has now written down their approach to World Models with Not Boring), to discussing the Cosmos World Model with with Andrew White of Edison Scientific on our new Science pod, to writing up our own theses on Adversarial World Models. Meanwhile Nvidia, Waymo and Tesla have published their own approaches, Google has released Genie 3, and Yann LeCun has raised $1B for AMI and published LeWorldModel.
Today’s guests have a radically different approach to World Modeling to every player we just mentioned — while Genie 3 is impressive, its many flaws demonstrate the issues with their approach - terrain clipping, noninteractivity (single player, no physics/no objects other than the player move), and maximum of 60 second immersion.
Moonlake AI (inspired by the Dreamworks logo) is the diametric opposite - immediately multiplayer, incredibly interactive, indefinite lifetime, capable of MANY different kinds of world models by simulating environments, predicting outcomes, and planning over long horizons. This is enabled by bootstrapping from game engines and training custom agents:
https://www.latent.space/p/moonlake
moonlakeai.com/blog/building-interactive-worlds
moonlakeai.com/blog/why-world-models-need-structure-not-just-scale
Timestamps
00:00 Benchmarking Gets Hard
00:47 Meet Moonlake Founders
01:26 Why Build World Models
03:12 Structure Not Just Scale
05:37 Defining Action Conditioned Worlds
07:32 Abstraction Versus Bitter Lesson
14:39 Language Versus JEPA Debate
20:27 Reasoning Traces And Rendering Layer
37:00 Gameplay Over Graphics
38:02 Fiction Rules And World Tweaks
39:15 Code Engines Beat Learned Priors
41:10 Diffusion Scaling Limits
43:23 Symbolic Versus Diffusion Boundary
46:14 Platform Vision Beyond Games
50:24 Spatial Audio And Multimodal Latents
54:23 NLP Roots Hiring And Moon Lake Name
We’ve been on a bit of a mini World Models series over the last quarter: from introducing the topic with Yi Tay, to exploring Marble with World Labs’ Fei-Fei Li and Justin Johnson, to previewing World Models learned from massive1 gaming datasets with General Intuition’s Pim de Witte (who has now written down their approach to World Models with Not Boring), to discussing the Cosmos World Model with with Andrew White of Edison Scientific on our new Science pod, to writing up our own theses on Adversarial World Models. Meanwhile Nvidia, Waymo and Tesla have published their own approaches, Google has released Genie 3, and Yann LeCun has raised $1B for AMI and published LeWorldModel.
Today’s guests have a radically different approach to World Modeling to every player we just mentioned — while Genie 3 is impressive, its many flaws demonstrate the issues with their approach - terrain clipping, noninteractivity (single player, no physics/no objects other than the player move), and maximum of 60 second immersion.
Moonlake AI (inspired by the Dreamworks logo) is the diametric opposite - immediately multiplayer, incredibly interactive, indefinite lifetime, capable of MANY different kinds of world models by simulating environments, predicting outcomes, and planning over long horizons. This is enabled by bootstrapping from game engines and training custom agents:
https://www.latent.space/p/moonlake
moonlakeai.com/blog/building-interactive-worlds
moonlakeai.com/blog/why-world-models-need-structure-not-just-scale
Timestamps
00:00 Benchmarking Gets Hard
00:47 Meet Moonlake Founders
01:26 Why Build World Models
03:12 Structure Not Just Scale
05:37 Defining Action Conditioned Worlds
07:32 Abstraction Versus Bitter Lesson
14:39 Language Versus JEPA Debate
20:27 Reasoning Traces And Rendering Layer
37:00 Gameplay Over Graphics
38:02 Fiction Rules And World Tweaks
39:15 Code Engines Beat Learned Priors
41:10 Diffusion Scaling Limits
43:23 Symbolic Versus Diffusion Boundary
46:14 Platform Vision Beyond Games
50:24 Spatial Audio And Multimodal Latents
54:23 NLP Roots Hiring And Moon Lake Name










