ImageNet Moment for Reinforcement Learning? [Prof. Jakob Foerster] @MachineLearningStreetTalk
ImageNet Moment for Reinforcement Learning? [Prof. Jakob Foerster]  @MachineLearningStreetTalk
Uploaded February 2025 | Updated September 2026, 1 week ago
SPONSOR MESSAGES:
***
CentML offers competitive pricing for GenAI model deployment, with flexible options to suit a wide range of models, from small to large-scale deployments. Check out their super fast DeepSeek R1 hosting!
centml.ai/pricing

Prof. Jakob Foerster (FLAIR lab, Oxford / Meta AI) and his PhD student Chris Lu make the case that deep reinforcement learning is finally winning the hardware lottery. The core thesis: RL has underperformed not because the ideas are wrong, but because running environments on CPUs while training agents on GPUs created a computational bottleneck that made experimentation slow, expensive, and brittle. JAX-based GPU-native environments now deliver ~4000x speedups, enabling the kind of rapid iteration that made supervised deep learning successful.

Chris Lu explains the technical foundation -- how JAX's JIT compilation and vmap (vectorized map) allow writing a single environment instance in NumPy-like code and scaling it to millions of parallel copies on GPU. This started from necessity: their lab initially had only Google Colab free-tier compute. The constraint forced them to put environments on GPU for their Model-Free Opponent Shaping paper, and the results were surprisingly effective.

The conversation then shifts to discovered policy optimization. Foerster's mirror learning framework provides theoretical grounding for why PPO works, and crucially, lets you parameterize the drift function as a neural network and meta-learn it. Evolution strategies, not gradient-based meta-learning, turned out to be the better optimizer for this -- a vindication of the bitter lesson. The learned policy optimization function shows intriguing "too good to be true" behavior: when advantages are very high, it clips more aggressively, as if it has learned skepticism.

The second half covers multi-agent systems, emergent communication, and AI governance. Foerster argues forcefully for open-source AI development and democratic control, drawing analogies to CERN. His position: the biggest alignment challenge is not between AI and humans, but between those who control AI systems and the rest of the population. Concentrated AI development creates fragile single points of failure; distributed development is both safer and more innovative.

---
REFERENCES:
paper:
[00:03:05] Deep RL Doesn't Work Yet
alexirpan.com/2018/02/14/rl-hard.html
[00:06:10] JaxMARL
arxiv.org/html/2311.10090v5
[00:08:50] M-FOS: Model-Free Opponent Shaping
arxiv.org/abs/2205.01447
[00:12:10] Kinetix Physics Simulator
arxiv.org/abs/2410.23208
[00:14:42] Mirror Learning Framework
arxiv.org/abs/2208.01682
[00:16:30] Discovered Policy Optimisation
arxiv.org/abs/2210.05639
[00:28:55] AlphaGo
arxiv.org/abs/1712.01815
[00:41:00] Open Source Generative AI
arxiv.org/abs/2405.08597
tool:
[00:09:45] JAX Library
github.com/jax-ml/jax
[00:49:51] Llama 3
ai.meta.com/blog/meta-llama-3
concept:
[00:24:10] Goodhart's Law
en.wikipedia.org/wiki/Goodhart%27s_law

---
LINKS:
Full Transcript: app.rescript.info/share/04498ba49b081dbcc254c40ba7b56035
Download PDF transcript: app.rescript.info/api/public/sessions/8fe06defb6d2216d/pdf

Prof. Jakob Foerster
https://x.com/j_foerst
jakobfoerster.com
University of Oxford Profile:
eng.ox.ac.uk/people/jakob-foerster

REFS
[[00:00:05] ARC Benchmark, Chollet
github.com/fchollet/ARC-AGI

[00:09:45] JAX Library, Google Research
github.com/jax-ml/jax

[00:25:15] LLM ARChitect, Franzen et al.
github.com/da-fr/arc-prize-2024/blob/main/the_architects.pdf
ImageNet Moment for Reinforcement Learning? [Prof. Jakob Foerster]Every Definition of Intelligence Is Wrong. Heres Why — Michael BennettAI Isnt Creative [Prof. Kenneth Stanley]Jay Alammar on LLMs, RAG, and AI EngineeringChatGPT will beat you at chess nowDavid Hansons Vision for Sentient RobotsModel quantisation leads to decoherence - Federico BarberoThe Real Reason Huge AI Models Actually Work [Prof. Andrew Wilson]What If Intelligence Didnt Evolve? It Was There From the Start! - Blaise Agüera y ArcasARC Prize Version 2 Launch Video! [Francois Chollet, Mike Knoop]Sara Hooker on language and reasoning #aiYou dont fine-tune your way to AGI - Heres why. [Eiso Kant]
Machine Learning Street Talk |

ImageNet Moment for Reinforcement Learning? [Prof. Jakob Foerster]

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER