Have a great week! πHow to contribute to open source AI with ZERO experience | feat. DanAdvantageDeep Learning with Yacine2025-12-04 | today Iβm going to show you exactly how to get involved in open source the right way, using the popular reinforcement learning library PufferLib, with insights from one of its contributors, DanAdvantage.
# Table of Content - Intro: 0:00 - Good mindset for open source contributions: 1:56 - Use community as a driver for learning: 3:06 - Become a power user to contribute: 4:40 - DanAdvantage contribution story time: 5:55 - Conclusion: 14:25
## Shout Out - check out pufferlib and star the repo to support: github.com/PufferAI/PufferLib - check out danAdvantage: https://x.com/DanAdvantage - check out joseph tutorial: https://x.com/jsuarez5341/articles
Have a great week! πExploring Understanding R1-Zero-Like Training (Dr. GRPO) | Deep Learning Study SessionDeep Learning with Yacine2025-11-21 | November 20 session where we are diving into the paper "Understanding R1-Zero-Like Training: A Critical Perspective" by the Sea AI Lab which culminate in the Dr. GRPO algorithm.
come chat with me and ask any questions related to research deep learning in general! :) Enjoy! πΉ
If you want to continue the discussion holla at me on twitter: π https://x.com/yacinelearningExploring Understanding R1-Zero-Like Training (Dr. GRPO) | Deep Learning Study SessionDeep Learning with Yacine2025-11-20 | November 20 session where we are diving into the paper "Understanding R1-Zero-Like Training: A Critical Perspective" by the Sea AI Lab which culminate in the Dr. GRPO algorithm.
come chat with me and ask any questions related to research deep learning in general! :) Enjoy! πΉ
If you want to continue the discussion holla at me on twitter: π https://x.com/yacinelearningHow to build runescape for RL Agent? | Neural MMO v3 with Joseph SuarezDeep Learning with Yacine2025-10-31 | October 29 session where we are going to go over the pufferlib 2.0 library and talk with Dr. Joseph Suarez on his intense work on the Neural MMO v3.0 environment π joseph suarez twitter: https://x.com/jsuarez5341 π contribute to pufferlib over here: github.com/PufferAI/PufferLib
paper links : π PufferLib 2.0: Reinforcement Learning at 1M steps/s: openreview.net/pdf?id=qRyteMTgn0 π about neural mmo: https://x.com/jsuarez5341/status/1866127102627438866
come chat with the author and ask any questions related to RL and environment design :) Enjoy! πΉ
If you want to continue the discussion holla at me on twitter: π https://x.com/yacinelearningHow to Finetune 8 Billion Parameters with Evolution Strategies? | with Yulu GanDeep Learning with Yacine2025-10-24 | October 23 session where we have the pleasure of having a chat with Yulu Gan one of the author of the paper βEvolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learningβ π yulu gan twitter: https://x.com/yule_gan π his website: yulugan.com/about.html
come chat with the author and ask any questions related to research + evolutionary strategies! :) Enjoy! πΉ
If you want to continue the discussion holla at me on twitter: π https://x.com/yacinelearningexploring Evolution Strategies at Scale for LLM Finetuning | Deep Learning Study SessionDeep Learning with Yacine2025-10-22 | october 21 session where we review the ES paper about finetuning LLM with billions of parameter!
Come study with me and ask any questions related to deep learning! :)
Enjoy! πΉ
Do consider hopping up into our Discord for more discussions: π discord.gg/QpkxRbQBpfexploring TRM Less is More: Recursive Reasoning with Tiny Networks | Deep Learning Study SessionDeep Learning with Yacine2025-10-18 | october 17 session where we review the TRM paper yes more latent reasoning stuff!
Reinforcement learning is becoming the defining ingredient behind the most capable AI agents. From OpenAIβs Deep Research to Anthropicβs Claude Code, RL is used to specialize models for reasoning, coding, and tool use.
In this video we'll do a beginner friendly overview of Reinforcement Learning with Verifiable Rewards (RLVR) environment and how to build them using the verifiers library!
π also, if you are a beginner: learn to code from full-stack to AI with Scrimba scrimba.com/?via=yacineMahdid (extra 20% off pro with my link, great resource, I love the team)
# Table of Content 00:00 - Introduction: RLβs growing role in agentic AI 01:10 - The RLVR loop: dataset, policy, rollouts, rewards, updates 02:13 - Overview of the state of RLVR 03:50 - Small-model RLVR: performance, latency, and cost benefits 06:00 - RLVR vs RLHF: key conceptual differences 07:32 - Open-source frameworks: ReasoningGym, ART, TRL and Verifiers 08:12 - deep dive into the verifiers 7 steps with math-python env 08:25 - deep dive into the verifiers | step 1 : data 09:09 - deep dive into the verifiers | step 2 : interaction style 09:40 - deep dive into the verifiers | step 3 : environment logic 10:05 - deep dive into the verifiers | step 4 : rewards function (rubric) 11:23 - deep dive into the verifiers | step 5 : parser (optional) 11:46 - deep dive into the verifiers | step 6 : package environment 12:07 - deep dive into the verifiers | step 7 : run eval or training 12:30 - a few community environments 13:25 - Case study: Building a Vision-Language RLVR environment feat alexine 13:56 - vision SR1 - overview 16:46 - vision SR1 - environment 1 18:29 - vision SR1 - environment 2 20:03 - Interview with prime Will Brown, creator of Verifiers 20:18 - Interview with prime Will Brown - verifiers development story 23:16 - Interview with prime Will Brown - what's the vision for environment hub? 24:17 - Interview with prime Will Brown - what future is there for RL environment? 26:27 - πΊπ¦πΊπ¦πΊπ¦
# Shout Out πΊ big thanks for alexine for her envrionment and for hoping on the video, check her out folks: https://x.com/alexinexxx πΊ thanks will for taking the time to have come down from gpu heaven to chat with us about verifiers: https://x.com/willccbb
Have a great week! πexploring emergent hierarchical reasoning in LLMs with RL (2025) | Deep Learning Study SessionDeep Learning with Yacine2025-10-01 | September 30 session where we review the paper "exploring emergent hierarchical reasoning in LLMs with RL" yes we are going back to hierarchical reasoning!
Come study with me and ask any questions related to deep learning! :)
Enjoy! πΉ
Do consider hopping up into our Discord for more discussions: π discord.gg/QpkxRbQBpfadagrad, adadelta, adam, adamax, adamw, aaaahhhhhh | Deep Learning Study SessionDeep Learning with Yacine2025-09-20 | September 19 session, where we are going to study optimizers some more and untangle this big adamess
Come study with me and ask any questions related to deep learning! :)
Enjoy! πΉ
Do consider hopping up into our Discord for more discussions: π discord.gg/QpkxRbQBpfOptimizers, ARC-AGI and Evolutionary Algorithm | Deep Learning Study SessionDeep Learning with Yacine2025-09-17 | September 17 session, where we are going to study optimizers and also check a new result on arc-agi!
Come study with me and ask any questions related to deep learning! :)
Enjoy! πΉ
Do consider hopping up into our Discord for more discussions: π discord.gg/QpkxRbQBpfNo BS Advices for Beginner Deep Learning ResearchersDeep Learning with Yacine2025-09-15 | π learn to code from full-stack to AI with Scrimba scrimba.com/?via=yacineMahdid (extra 20% off pro with my link, great resource, I love the team)
I found that surprisingly, a lot of students and engineers are interested in doing research.
so here are 3 kinds of kinda blunt advice I gave recently in one of our paper review live streams.
π also, if you are a beginner: learn to code from full-stack to AI with Scrimba scrimba.com/?via=yacineMahdid (extra 20% off pro with my link, great resource, I love the team)
# Table of Content - Introduction: 0:00 - following research passion with no career reward?: 1:34 - why it's a good idea to follow that research passion?: 4:00 - where to find research idea? : 6:23
Have a great week! πNo BS advice for deep learning beginnersDeep Learning with Yacine2025-09-04 | check out deep-ml here: https://www.deep-ml.com?ref=yacinelearning
I receive a lot of DM from stressed-out students who want to get beginner guidance on the deep learning field.
Often, people seeking this type of guidance are learning deep learning with the hope of achieving something else, which is not properly stated.
So, here Iβll give two general guidelines to help you structure your goal-setting and your job trajectory.
π also, if you are a beginner: learn to code from full-stack to AI with Scrimba scrimba.com/?via=yacineMahdid (extra 20% off pro with my link, great resource, I love the team)
Table of Content: - Introduction: 0:00 - beginner advice number 1: 2:01 - beginner advice number 2: 5:46
Have a great week! πLearning the periodic table by heart in 2 days | deep learning study sessionsDeep Learning with Yacine2025-08-30 | This friday, we'll learn the full periodic table by heart and in order.
buckle up folks we're going to learn about how memory work in the human mind and what that means for deep learning models.
Don't hesitate to hop on to ask questions about anything machine learning or deep learning related!
Shout out to all my members, really nice to hang out with you all!
π [SDK to integrate into your app] https://tldraw.dev/?utm_source=youtube&utm_medium=socials&utm_campaign=standard&utm_term=yacinemahdid
In this tutorial, we will review the muon optimizer, which is currently making waves in large language modelling by powering Kimi K2.
You will see it's not that complicated.
π also, if you are a beginner: learn to code from full-stack to AI with Scrimba scrimba.com/?via=yacineMahdid (extra 20% off pro with my link, great resource, I love the team)
# Table of Contents - introduction: 0:00 - why muon is useful?: 2:04 - adam overview: 3:30 - adamw overview: 4:32 - what muon is doing?: 7:31 - muon authors overview: 8:26 - muon results: 10:39 - kimi k2 performance with muon-clip: 12:29 - what does muon do?: 13:54 - deep dive in newton schulz: 16:52 - coding muon in numpy: 27:59
π [SDK to integrate into your app] https://tldraw.dev/?utm_source=youtube&utm_medium=socials&utm_campaign=standard&utm_term=yacinemahdid
In this tutorial, we will review the AdamW optimizer, which is currently the state of the art in most deep learning training regimens.
AdamW is a variant of Adam which apply regularization in the form of weight decay in a very specific way that makes it much more stable than the regular Adam optimizer.
We'll break down where this regularization helps and how it differs from the traditional Adam + L2 regularization.
π also, if you are a beginner: learn to code from full-stack to AI with Scrimba scrimba.com/?via=yacineMahdid (extra 20% off pro with my link, great resource, I love the team)
# Table of Content - Introduction: 0:00 - Where does AdamW fit?: 2:30 - What type of regularization does AdamW apply?: 3:25 - Why Adam with L2 sucked?: 4:49 - Isn't L2 and Weight Decay the same?: 5:48 - AdamW formula breakdown: 6:19 - AdamW code implementation (lol): 14:32 - AdamW Recap: 19:34
Have a great week! πAdam Optimizer from Scratch in PythonDeep Learning with Yacine2025-07-22 | In this tutorial, we will review the Adam optimizer, which is used during the training of deep neural networks.
Have a great week! πCoding backprop through time in Python | Mathematics for Machine Learning Study SessionDeep Learning with Yacine2025-06-26 | We'll be going through some live coding exercises in Python, and I'll do my best to code the backdrop through time algo!
Have a great week! πCoding a RNN in Python | Mathematics for Machine Learning Study SessionDeep Learning with Yacine2025-06-24 | We'll be going through some live coding exercises in Python, and I'll do my best to create a simple recurrent neural network!
Have a great week! πWhat the heck is Artificial Immune System??? | AI Paper ReviewDeep Learning with Yacine2025-06-16 | Buckle up folks, we are going to figure out what the heck the field of artificial immune system is and if it's still releveant in this day and age.
I've discovered that field last week while trying to figure out what hypervector and hyperdimensional computing was while working on this deep-ml problem: π deep-ml.com/problems/74?ref=yacinelearning
Have a great week! πCoding a Composite Hypervector in Python | Mathematics for Machine Learning Study SessionDeep Learning with Yacine2025-06-09 | We'll be going through some live coding exercises in Python, and I'll do my best to create a composite hypervector for a dataset row!
Have a great week! πBackpropagation with Automatic Differentiation from Scratch in PythonDeep Learning with Yacine2025-06-05 | In this tutorial, we will review the automatic differentiation algorithm which is at the core of the autograd library in Pytorch.
Have a great week! πCoding Gaussian Elimination in Python | Mathematics for Machine Learning Study SessionDeep Learning with Yacine2025-06-03 | We'll be going through some live coding exercises in Python, and I'll do my best to implement the Gaussian Elimination for Solving Linear Systems!
Don't hesitate to hop on to ask questions about anything machine learning or deep learning related! Big thanks to all my sponsors, really nice to hang out with you al!
Have a great week! πMasked Self-Attention from Scratch in PythonDeep Learning with Yacine2025-05-28 | In this tutorial, we will review the algorithm for the masked self-attention and code it in numpy.
The algorithm is very commonly used in pretraining large language models, and being familiar with its quirks will allow you to better understand pre-training in LLMs.
Have a great week! πReduced Row Echelon Form (RREF) Algorithm From Scratch in PythonDeep Learning with Yacine2025-05-20 | In this tutorial, we will review the algorithm for transforming an augmented matrix into its reduced row echelon form in Python.
Have a great week! πCoding Reduced Row Echelon Form in Python | Mathematics for Machine Learning Study SessionDeep Learning with Yacine2025-05-18 | We'll be going through some live coding exercises in Python, and I'll do my best to implement the Reduced Row Echelon Form (RREF) function!
Have a great week! πHow to train LLMs with long context?Deep Learning with Yacine2025-05-09 | In today's video, I wanted to cover context windows in the transformer's architecture and how to make them BIG.
# Table of Content - Introduction: 0:00 - Why more context is good: 0:33 - R1 longer context: 1:06 - A little retrieval test: 1:56 - Needle-in-a-haystack: 2:40 - Multi-Round Needle-in-a-haystack: 3:38 - Machine Translation from One Book MTOB: 4:52 - Attention Calculation Recap: 6:16 - How to encode positions: 8:51 - Issue with increasing context: 10:07 - How to extend context: 11:26 - Fixing positional encoding: 11:45 - Fixing Attention Calculation: 13:21 - Flash Attention: 13:55 - Sparse Attention: 14:52 - Low-Rank Decomposition: 18:14 - Chunking: 19:51 - Other type of strategy using linear components: 21:44 - LLama 4 changes: 24:12 - Google Long Context Team (Nikolay Savinov): 25:33 - see you folks! : 26:50
This is an especially interesting topic, at least for me, to dig into because we are starting to see models with quite large context windows.
For example, Gemini 2.5 has a context window of 1M tokens and Llama4 scout boasts 10M.
Have a great week! πCoding Gauss-Seidel Method in Python | Mathematics for Machine Learning Study SessionDeep Learning with Yacine2025-05-07 | We'll be going through some live coding exercises in Python, and I'll do my best to implement the Gauss-Seidel method for solving linear systems!
Have a great week! πCoding Conjugate Gradient Method in Python | Mathematics for Machine Learning Study SessionDeep Learning with Yacine2025-04-28 | We'll be going through some live coding exercises in Python, and I'll do my best to implement the conjugate gradient method for solving linear systems!
Multi-Agent Systems sounds super cool, but you might be confused about what it is.
In this tutorial, we will dive into the theory of two multi-agent system papers and a mini-project to showcase how these types of systems work:
# Table of Content - Introduction: 0:00 - Software Project Overview: 1:31 - Theory of Multi-Agent Systems: 2:18 - Overview of MetaGPT with Pong: 13:05 - Building a Software Project: 19:16 - Closing Words: 27:12 - Bloopers: 27:49
I had a blast working with that system. I don't know why but having these poor digital employees trying to figure out my vague requirements was absolutely hilarious!
Here are more links for you to continue exploring: π MetaGPT: Meta Programming for A Multi-Agent Collaborative Framework: arxiv.org/abs/2308.00352 π AFlow: Automating Agentic Workflow Generation: arxiv.org/abs/2410.10762 π CAMEL-AI: camel-ai.org π Generative Agents: Interactive Simulacra of Human Behavior: arxiv.org/abs/2304.03442
#MetaGPT #DeepWisdom #MGX ---- Join the newsletter for weekly AI content: yacinemahdid.com Join the Discord for general discussion: discord.gg/QpkxRbQBpf
AI Engineering is very different from other data-related skills. It's much closer to software engineering. In today's video, I'll show you my preferred resources for adding that skill to your toolbelt.
# Table of Contents - Introduction: 0:00 - Step 0 - Learn the tools: 1:13 - Sneakpeek of Scrimba UI: 2:52 - AI Engineering Path modules overview: 4:41 - Step 1 - Expand skills through projects: 6:42 - LLM's Engineering Handbook Project: 7:43 - Step 2 - Learn foundational models in depth: 8:58 - AI Engineering book by Chip: 9:24 - Final words: 10:48
Don't go too much into analysis paralysis; get started and learn!
PS: Try to count how many times I'm saying "actually" lol. ---- Join the newsletter for weekly AI content: yacinemahdid.com Join the Discord for general discussion: discord.gg/QpkxRbQBpf
Have a great week! πHow to Code Singular Value Decomposition? | Mathematics for Machine Learning Study SessionDeep Learning with Yacine2025-04-04 | We'll be going through some live coding exercises in Python, and I'll do my best to implement Singular Value Decomposition using the Jacobian method.
Have a great week! πMiniMax-01 Theory | 1M Context + Lightning Attention + GPU OptimizationDeep Learning with Yacine2025-03-31 | Here's an overview of the new open source models from MiniMax (MIT), minimax-text-01, and minimax-vl-01.
This new model leverage lightning attention (a type of linear attention) in order to remove the quadratic bottleneck on the traditional softmax attention in transformer models.
Table of contents: - Introduction: 0:00 - Model Overview: 3:04 - Main Result Overview: 8:14 - Background Information on Linear Attention: 11:00 - Lightning Attention Overview: 16:07 - I/O Optimization: 22:20 - Pre-training recipe: 25:10 - Post-training recipe: 26:31 - Full Results: 30:42 - Vision Modality for MiniMax-VL-01: 37:24 - Demo of MiniMax-text-01: 41:20 - Final Words: 45:04
Key Takeaways: - 456 billion parameters and a 4-million-token context window, made for tasks requiring long-context understanding. - The Lightning Attention mechanism makes MiniMax-01 faster and more memory-efficient than traditional transformers. - MiniMax 01 surpasses Llama 3.1 in various benchmarks and challenges Claude, while also matching Deepseek V3 in specific tests.
Come study with me and ask any questions related to machine or deep learning! :)
Enjoy! πΉ
Do consider hopping up into our Discord for more discussions: π discord.gg/QpkxRbQBpfLinear Algebra Study Session - Sarrus Rule and Laplace Expansion | Mathematics for Deep LearningDeep Learning with Yacine2025-03-17 | This Sunday, we will continue our beginner study session on linear algebra topics with the determinant!
Come study with me and ask any questions related to machine or deep learning! :)
You might wonder what agents are and be a bit confused when you found no clear cut definition. In todayβs tutorial, weβre going to clarify that with a bit of theory and several code walkthroughs.
# Table of Contents - Introduction: 0:00 - What is an Agent: 1:50 - HuggingFace agents definitions overview: 4:01 - Agent vs workflows: 8:20 - Prompt Chaining: 10:25 - Routing: 12:28 - Parallelization: 16:03 - Orchestrator-Workers: 18:13 - Evaluator-Optimizer: 24:18 - Why you shouldn't build autonomous agents: 29:47 - Focus on simple agentic systems: 34:22
Have a great week! πDont work in LLMs. - Yan Lecun Meta #deeplearning #llmDeep Learning with Yacine2025-03-04 | Still, LLMs are nice and should be improved.Why I would still learn to program in 2025 (even with AI).Deep Learning with Yacine2025-02-26 | 30 days free trial of Brilliant + 20% off premium subscription over here: brilliant.org/deeplearningwithyacine
Contrary to what Jensen, CEO of NVIDIA says, I still think you should learn to program in 2025. In this tutorial, I outline the three main reasons why I believe that:
# Table of Content - Introduction: 0:00 - Programming is not about coding: 0:58 - What Francois Chollet thinks: 3:33 - Programming concepts help with problem-solving: 4:27 - Programming makes you better at using AI: 12:42 - Conclusion: 15:14
FTC disclaimer: This video was sponsored by Brilliant ---- Join the newsletter for weekly AI content: yacinemahdid.com Join the Discord for general discussion: discord.gg/QpkxRbQBpf
Have a great week! πMachine learning sucks. - Yann Lecun Meta #deeplearning #llmDeep Learning with Yacine2025-02-25 | Highlighting the shortcomings of LLM is important.AGI by 2026? - Dario Amodei CEO Anthropic #deeplearning #llm #machinelearningDeep Learning with Yacine2025-02-20 | AGI between 2026 and 2126 confirmed!LLMs are conscious? - Geoffrey Hinton reaction #machinelearning #deeplearning #llmDeep Learning with Yacine2025-02-18 | How many neurons do you need to remove before consciousness goes away? π€ͺKL Divergence in DeepSeek R1 | Implementation Walk-throughDeep Learning with Yacine2025-02-13 | Sometimes, you read a deep learning formula and you have no idea where it comes from.
In this tutorial we are going to dive (too) deep into the KL divergence implementation of GRPO in DeepSeek R1.
## Table of Content: - Introduction: 0:00 - KL Divergence in GRPO vs PPO: 1:00 - KL Divergence refresher: 2:30 - Monte Carlo estimation of KL divergence: 6:42 - Schulman blog: 7:58 - k1 = log(q/p): 8:55 - k2 = 0.5*log(p/q)^2: 11:23 - k3 = (p/q - 1) - log(p/q): 13:35 - benchmarking: 15:58 - takeaways: 18:43
Have a great week! πGroup Relative Policy Optimization (GRPO) - Formula and CodeDeep Learning with Yacine2025-02-05 | The GRPO algorithm is at the heart of the newest DeepSeek R1 architecture. In this tutorial, we will discuss the details of the formula along with a code implementation walkthrough from the HuggingFace post-training team!
# Table of Content - Introduction: 0:00 - PPO vs GRPO: 1:18 - PPO formula overview: 4:24 - GRPO formula overview: 7:49 - GRPO pseudo code: 11:11 - GRPO Trainer code: 13:21 - Conclusion: 23:48
Have a great week! πDeepSeek R1 Theory Overview | GRPO + RL + SFTDeep Learning with Yacine2025-01-31 | Here's an overview of the DeepSeek R1 paper. I read the paper this week and I was fascinated by the methods, however it was a bit difficult to follow what was going on with all the models being used.
I found a neat map of the methodology which I'll be using in this tutorial to walk you through the paper.
I strongly recommend you to still read the paper over here: π PAPER: arxiv.org/pdf/2501.12948
Have a great week! πWhat are Foundational Models in Deep Learning? #deeplearning #machinelearning #aiDeep Learning with Yacine2025-01-30 | LxM are the way to go!Great course for learning deep learning #deeplearning #programmingDeep Learning with Yacine2025-01-28 | π course.fast.aiLearning to program in 2025? - Jensen Huang NVIDIA CEO reaction #deeplearning #programming #aiDeep Learning with Yacine2025-01-25 | ... Yeah no, learn to program folks.