LLM’s Billion Dollar Problem @bycloudAI
LLM’s Billion Dollar Problem  @bycloudAI
Uploaded February 2026 | Updated September 2026, 1 week ago
Check out Inngest and let your AI agents wear a harness now https://innge.st/yt-bycl-1

This video originally was going to be about the Linear Attention Saga that happened between June and November last year, but it turned out I needed quite some build up to explain what the significance of Linear Attention, 1 million context window, and compute scaling are. So it accidentally became a 17 mins video...


my latest project: Intuitive AI Academy
https://intuitiveai.academy/
limited time code "NYNM" for 50% off forever (only a 25 spots left!)

My Newsletter
mail.bycloud.ai

my project: find, discover & explain AI research semantically
findmypapers.ai

My Patreon
patreon.com/c/bycloud

Sauces (will try to be in order of apperance)

GPT-OSS
[Code] github.com/openai/gpt-oss
[Paper] arxiv.org/abs/2508.10925
(they did not mention it is sliding window attention explicitly in its paper)

DeepSeek V3.2
[Paper] arxiv.org/pdf/2512.02556

DeepSeek V2 (Multi-head Latent Attention)
[Paper] arxiv.org/abs/2405.04434

Kimi K2.5
[Paper] arxiv.org/abs/2602.02276

MiniMax
[Text-01] arxiv.org/abs/2501.08313
[M1] arxiv.org/abs/2506.13585
[M2] huggingface.co/MiniMaxAI/MiniMax-M2
[Zhihu Blog] zhihu.com/question/1965302088260104295/answer/1966810157473335067

Qwen-3 Next
[Project Page] qwen.ai/blog?id=4074cca80393150c248e508aa62983f9cb7d27cd

Hunyuan T1
[Project Page] tencent.github.io/llm.hunyuan.T1/README_EN.html

Kimi Linear
[Paper] arxiv.org/abs/2510.26692

Context Arena (used for comparing long context performance)
[Project Page] contextarena.ai

Gemini 3 Flash
[Blog] https://blog.google/products-and-platforms/products/gemini/gemini-3-flash/

Claude 4.6
[Blog] anthropic.com/news/claude-opus-4-6


Try out my new fav place to learn how to code scrimba.com/?via=bycloudAI

This video is supported by the kind Patrons & YouTube Members:
🙏Spam Maj, Alex, Chris LeDoux, DX Research Group, Poof N' Inu, Deagan, Robert Zawiasa, Ryszard Warzocha, Midwstmakr, Tobe2d, Louis Muk, Akkusativ, Kevin Tai, Mark Buckler, NO U, Tony Jimenez, Ângelo Fonseca, jiye, Anushka, Asad Dhamani, Binnie Yiu, Calvin Yan, Clayton Ford, Diego Silva, Etrotta, Gonzalo Fidalgo, Handenon, Hector, Jake Disco very, Michael Brenner, Nilly K, OlegWock, Daddy Wen, Shuhong Chen, Sid_Cipher, Stefan Lorenz, Sup, tantan assawade, Thipok Tham, Thomas Di Martino, Thomas Lin, Richárd Nagyfi, Paperboy, mika, Leo, Berhane-Meskel, Kadhai Pesalam, mayssam, Bill Mangrum, nyaa, Toru Mon, Lame Plane, Matej Macak, thechoephix


[Discord] discord.gg/NhJZGtH
[Twitter] twitter.com/bycloudai
[Patreon] patreon.com/bycloud
[Business Inquiries] bycloud@smoothmedia.co
[Music] @IraStoria
[Profile & Banner Art] twitter.com/pygm7
[Video Editor] Abhay and @Booga04
[Ko-fi] ko-fi.com/bycloudai
LLM’s Billion Dollar Problem5 New AI Scams You Should Tell Your Parents AboutAttention Residuals: Kimi AIs Elegant LLM Architecture BreakthroughCode Interpreter, Animate Diff, SDXL 1.0 & More [The AI Timeline #9]GPT-4: The Best AI (Gaslighter) Has Been Born!LLM Attention That Expands At Inference? Test Time Training ExplainedWhen AI Tries To Reason With Itself [AutoGPT & More]A new way to fine-tune LLMs just droppedFREE Text Generated Music Is Getting Too Good... [Metas MusicGen]How AI Photogrammetry Is Changing 3D ForeverAnthropic found a terrifying consequence of adding reasoning to AIFirst Look At GPT-4 With Vision
bycloud |

LLM’s Billion Dollar Problem

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER