The RL Irony in LLMs (and its insane new meta) @bycloudAI
The RL Irony in LLMs (and its insane new meta)  @bycloudAI
Uploaded January 2026 | Updated September 2026, 1 week ago
Start learning cyber security with TryHackMe: tryhackme.com/bycloud Use my code "BYCLOUD25" to get 25% off on annual subscription!

This video breaks down what's wrong with scaling RL for LLMs, especially in the direction of reaching AGI, but why RL still matters. As RL is noisy and can hurt generalization, yet it enables exploration and self-correction that pretraining can’t, we are stuck between a rock and a hard place with this direction. We’ll also look at why LoRA is becoming the practical way to do RL cheaply, swappable adapters that can match full fine-tuning on reasoning and make personalized agents easier to deploy, which might look like a promising future direction to apply RL on a massive scale.


my latest project: Intuitive AI Academy
https://intuitiveai.academy/
code "NYNM" for 50% off forever (limited to 50)


Dwarkesh Podcast w/ AK
[YouTube] youtu.be/lXUZvyajciY

Dwarkesh Podcast w/ Ilya
[YouTube] youtu.be/aR20FWCCjAs

Beyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective Reinforcement Learning for LLM Reasoning
[Paper] arxiv.org/abs/2506.01939

The Path Not Taken: RLVR Provably Learns Off the Principals
[Paper] arxiv.org/abs/2511.08567

LoRA Without Regret
[Blog] thinkingmachines.ai/blog/lora

Tina: Tiny Reasoning Models via LoRA
[Paper] arxiv.org/abs/2504.15777

Tinker
[Website] thinkingmachines.ai/tinker


My Newsletter
mail.bycloud.ai

My Patreon
patreon.com/c/bycloud


Try out my new fav place to learn how to code scrimba.com/?via=bycloudAI

This video is supported by the kind Patrons & YouTube Members:
🙏Spam Maj, Alex, Chris LeDoux, DX Research Group, Poof N' Inu, Deagan, Robert Zawiasa, Ryszard Warzocha, Tobe2d, Louis Muk, Akkusativ, Kevin Tai, Mark Buckler, NO U, Tony Jimenez, Ângelo Fonseca, jiye, Anushka, Asad Dhamani, Binnie Yiu, Calvin Yan, Clayton Ford, Diego Silva, Etrotta, Gonzalo Fidalgo, Handenon, Hector, Jake Disco very, Michael Brenner, Nilly K, OlegWock, Daddy Wen, Shuhong Chen, Sid_Cipher, Stefan Lorenz, Sup, tantan assawade, Thipok Tham, Thomas Di Martino, Thomas Lin, Richárd Nagyfi, Paperboy, mika, Leo, Berhane-Meskel, Kadhai Pesalam, mayssam, Bill Mangrum, nyaa, Toru Mon, Lame Plane, Matej Macak


[Discord] discord.gg/NhJZGtH
[Twitter] twitter.com/bycloudai
[Patreon] patreon.com/bycloud
[Business Inquiries] bycloud@smoothmedia.co
[Profile & Banner Art] twitter.com/pygm7
[Video Editor] Abhay and @Booga04
[Ko-fi] ko-fi.com/bycloudai
The RL Irony in LLMs (and its insane new meta)First Look At Metas Emu Edit & Emu VideoAI Generated Videos Are Getting Out of HandOpenAIs Deep Research Is MidGemini 2.5 Pro is just the best choice for AI right nowLLMs Are Better At Jailbreaking Themselves Than Us...New AI Paradigm?! Energy-Based Transformers ExplainedThese TikTokers Are AI Generated [SD Multi-Frame Rendering][VOD] Picking RTX4080 Super Giveaway Winner & ChillLlama 3.2 Deep Dive - Tiny LM & NEW VLM Unleashed By MetaSam Altman Sells Crypto, SDXL 1.0 Release, ChatGPT Custom Instructions & More [The AI Timeline #11]Reading AIs Mind - Mechanistic Interpretability Explained [Anthropic Research]
bycloud |

The RL Irony in LLMs (and its insane new meta)

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER