Uploaded January 2026 | Updated September 2026, 2 weeks ago
From shipping *Gemini Deep Think* and *IMO Gold* to launching the *Reasoning and AGI team in Singapore,* *Yi Tay* has spent the last 18 months living through the full arc of Google DeepMind's pivot from architecture research to RL-driven reasoning—watching his team go from a dozen researchers to 300+, training models that solve International Math Olympiad problems in a live competition, and building the infrastructure to scale deep thinking across every domain, and driving Gemini to the top of the leaderboards across every category. Yi Returns to dig into the inside story of the IMO effort and more!
We discuss:
* Yi's path: *Brain → Reka → Google DeepMind → Reasoning and AGI team Singapore,* leading model training for Gemini Deep Think and IMO Gold
* The *IMO Gold story:* four co-captains (Yi in Singapore, Jonathan in London, Jordan in Mountain View, and Tong leading the overall effort), training the checkpoint in ~1 week, live competition in Australia with professors punching in problems as they came out, and the tension of not knowing if they'd hit Gold until the human scores came in (because the Gold threshold is a percentile, not a fixed number)
* Why they *threw away AlphaProof:* "If one model can't do it, can we get to AGI?" The decision to abandon symbolic systems and bet on end-to-end Gemini with RL was bold and non-consensus
* *On-policy vs. off-policy RL:* off-policy is imitation learning (copying someone else's trajectory), on-policy is the model generating its own outputs, getting rewarded, and training on its own experience—"humans learn by making mistakes, not by copying"
* Why *self-consistency and parallel thinking* are fundamental: sampling multiple times, majority voting, LM judges, and internal verification are all forms of self-consistency that unlock reasoning beyond single-shot inference
* The *data efficiency frontier:* humans learn from 8 orders of magnitude less data than models, so where's the bug? Is it the architecture, the learning algorithm, backprop, off-policyness, or something else?
* Three schools of thought on *world models:* (1) Genie/spatial intelligence (video-based world models), (2) Yann LeCun's JEPA + FAIR's code world models (modeling internal execution state), (3) the amorphous "resolution of possible worlds" paradigm (curve-fitting to find the world model that best explains the data)
* Why *AI coding crossed the threshold:* Yi now runs a job, gets a bug, pastes it into Gemini, and relaunches without even reading the fix—"the model is better than me at this"
* The *Pokémon benchmark:* can models complete Pokédex by searching the web, synthesizing guides, and applying knowledge in a visual game state? "Efficient search of novel idea space is interesting, but we're not even at the point where models can consistently apply knowledge they look up"
* *DSI and generative retrieval:* re-imagining search as predicting document identifiers with semantic tokens, now deployed at YouTube (symmetric IDs for RecSys) and Spotify
* Why *RecSys and IR feel like a different universe:* "modeling dynamics are strange, like gravity is different—you hit the shuttlecock and hear glass shatter, cause and effect are too far apart"
* The *closed lab advantage is increasing:* the gap between frontier labs and open source is growing because ideas compound over time, and researchers keep finding new tricks that play well with everything built before
* Why *ideas still matter:* "the last five years weren't just blind scaling—transformers, pre-training, RL, self-consistency, all had to play well together to get us here"
* *Gemini Singapore:* hiring for RL and reasoning researchers, looking for track record in RL or exceptional achievement in coding competitions, and building a small, talent-dense team close to the frontier
—
Yi Tay
* Google DeepMind: https://deepmind.google
* X: https://x.com/YiTayML
00:00:00 Introduction: Returning to Google DeepMind and the Singapore AGI Team
00:04:52 The Philosophy of On-Policy RL: Learning from Your Own Mistakes
00:12:00 IMO Gold Medal: The Journey from AlphaProof to End-to-End Gemini
00:21:33 Training IMO Cat: Four Captains Across Three Time Zones
00:26:19 Pokemon and Long-Horizon Reasoning: Beyond Academic Benchmarks
00:36:29 AI Coding Assistants: From Lazy to Actually Useful
00:32:59 Reasoning, Chain of Thought, and Latent Thinking
00:44:46 Is Attention All You Need? Architecture, Learning, and the Local Minima
00:55:04 Data Efficiency and World Models: The Next Frontier
01:08:12 DSI and Generative Retrieval: Reimagining Search with Semantic IDs
01:17:59 Building GDM Singapore: Geography, Talent, and the Symposium
01:24:18 Hiring Philosophy: High Stats, Research Taste, and Student Budgets
01:28:49 Health, HRV, and Research Performance: The 23kg Journey
From shipping *Gemini Deep Think* and *IMO Gold* to launching the *Reasoning and AGI team in Singapore,* *Yi Tay* has spent the last 18 months living through the full arc of Google DeepMind's pivot from architecture research to RL-driven reasoning—watching his team go from a dozen researchers to 300+, training models that solve International Math Olympiad problems in a live competition, and building the infrastructure to scale deep thinking across every domain, and driving Gemini to the top of the leaderboards across every category. Yi Returns to dig into the inside story of the IMO effort and more!
We discuss:
* Yi's path: *Brain → Reka → Google DeepMind → Reasoning and AGI team Singapore,* leading model training for Gemini Deep Think and IMO Gold
* The *IMO Gold story:* four co-captains (Yi in Singapore, Jonathan in London, Jordan in Mountain View, and Tong leading the overall effort), training the checkpoint in ~1 week, live competition in Australia with professors punching in problems as they came out, and the tension of not knowing if they'd hit Gold until the human scores came in (because the Gold threshold is a percentile, not a fixed number)
* Why they *threw away AlphaProof:* "If one model can't do it, can we get to AGI?" The decision to abandon symbolic systems and bet on end-to-end Gemini with RL was bold and non-consensus
* *On-policy vs. off-policy RL:* off-policy is imitation learning (copying someone else's trajectory), on-policy is the model generating its own outputs, getting rewarded, and training on its own experience—"humans learn by making mistakes, not by copying"
* Why *self-consistency and parallel thinking* are fundamental: sampling multiple times, majority voting, LM judges, and internal verification are all forms of self-consistency that unlock reasoning beyond single-shot inference
* The *data efficiency frontier:* humans learn from 8 orders of magnitude less data than models, so where's the bug? Is it the architecture, the learning algorithm, backprop, off-policyness, or something else?
* Three schools of thought on *world models:* (1) Genie/spatial intelligence (video-based world models), (2) Yann LeCun's JEPA + FAIR's code world models (modeling internal execution state), (3) the amorphous "resolution of possible worlds" paradigm (curve-fitting to find the world model that best explains the data)
* Why *AI coding crossed the threshold:* Yi now runs a job, gets a bug, pastes it into Gemini, and relaunches without even reading the fix—"the model is better than me at this"
* The *Pokémon benchmark:* can models complete Pokédex by searching the web, synthesizing guides, and applying knowledge in a visual game state? "Efficient search of novel idea space is interesting, but we're not even at the point where models can consistently apply knowledge they look up"
* *DSI and generative retrieval:* re-imagining search as predicting document identifiers with semantic tokens, now deployed at YouTube (symmetric IDs for RecSys) and Spotify
* Why *RecSys and IR feel like a different universe:* "modeling dynamics are strange, like gravity is different—you hit the shuttlecock and hear glass shatter, cause and effect are too far apart"
* The *closed lab advantage is increasing:* the gap between frontier labs and open source is growing because ideas compound over time, and researchers keep finding new tricks that play well with everything built before
* Why *ideas still matter:* "the last five years weren't just blind scaling—transformers, pre-training, RL, self-consistency, all had to play well together to get us here"
* *Gemini Singapore:* hiring for RL and reasoning researchers, looking for track record in RL or exceptional achievement in coding competitions, and building a small, talent-dense team close to the frontier
—
Yi Tay
* Google DeepMind: https://deepmind.google
* X: https://x.com/YiTayML
00:00:00 Introduction: Returning to Google DeepMind and the Singapore AGI Team
00:04:52 The Philosophy of On-Policy RL: Learning from Your Own Mistakes
00:12:00 IMO Gold Medal: The Journey from AlphaProof to End-to-End Gemini
00:21:33 Training IMO Cat: Four Captains Across Three Time Zones
00:26:19 Pokemon and Long-Horizon Reasoning: Beyond Academic Benchmarks
00:36:29 AI Coding Assistants: From Lazy to Actually Useful
00:32:59 Reasoning, Chain of Thought, and Latent Thinking
00:44:46 Is Attention All You Need? Architecture, Learning, and the Local Minima
00:55:04 Data Efficiency and World Models: The Next Frontier
01:08:12 DSI and Generative Retrieval: Reimagining Search with Semantic IDs
01:17:59 Building GDM Singapore: Geography, Talent, and the Symposium
01:24:18 Hiring Philosophy: High Stats, Research Taste, and Student Budgets
01:28:49 Health, HRV, and Research Performance: The 23kg Journey
 and Isomorphic for example), it is starting to look like the appetite of Pharma for biotech tools has finally started to grow. Why the sudden interest?
Timestamps:
(0:00) The challenges of starting a biotech lab and generating data from scratch.
(0:55) Introduction of Ron Alfa and Dan Bear from Noetik.
(4:09) The complexity of cancer: Why curing cancer is a misleading concept and the need for new, multimodal data.
(8:24) Identifying therapeutically relevant cancer subtypes to improve clinical trial success rates.
(11:27) The importance of intentional, high-quality data generation in AI biotech.
(17:09) Lessons learned from Recursion Pharmaceuticals regarding batch effects and data design.
(20:14) Introduction to Noetiks core data modalities: Pathology (H&E), spatial transcriptomics, and genomic alterations.
(30:15) The philosophy of self-supervised learning and avoiding bias from electronic health records.
(36:01) Translating latent space embeddings and patient clusters into actionable insights for pharma.
(41:40) Using PerturbMap and in-vivo mouse models to validate human AI predictions.
(53:38) Technical deep dive: The Tario transformer-based model and auto-regressive training objectives.
(1:00:26) The GSK partnership: Licensing OctoVC and the shift toward platform-based biotech deals.
(1:13:55) Advice for small biotech AI startups: Scaling, data conviction, and lessons from scientific history. 🔬 Training Transformers to solve 95% failure rate of Cancer Trials — Ron Alfa & Daniel Bear, Noetik](https://i.ytimg.com/vi/uqM8qjbLRHA/mqdefault.jpg)









