Cheating LLM Benchmarks Is Easier Than You Think… @bycloudAI
Cheating LLM Benchmarks Is Easier Than You Think…  @bycloudAI
Uploaded March 2025 | Updated September 2026, 2 weeks ago
Sign up for NVIDIA GTC2025 here!
nvda.ws/48s4tmc

Join The RTX4080 SUPER Giveaway (enter between March 17-21st)
https://forms.gle/TbGgoD5obn1zRz5r7

In this video, you will learn about the fascinating ways of how AI companies can rig LLM benchmarks, for educational purposes of course... Plus the reason why people don’t trust “chatbot arena” as much as more.

My Newletter
mail.bycloud.ai

My Patreon
patreon.com/c/bycloud


I was really vague about the sources in this video to fit it in the narrative, but most facts presented are backed by the following papers

Changing Answer Order Can Decrease MMLU Accuracy
[Paper] arxiv.org/abs/2406.19470

Catch me if you can! How to beat GPT-4 with a 13B model
[Blog] lmsys.org/blog/2023-11-14-llm-decontaminator

Idiosyncrasies in Large Language Models
[Paper] arxiv.org/abs/2502.12150

Improving Your Model Ranking on Chatbot Arena by Vote Rigging
[Paper] arxiv.org/abs/2501.17858


This video is supported by the kind Patrons & YouTube Members:
🙏Andrew Lescelius, Ben Shaener, Chris LeDoux, Miguilim, Deagan, FiFaŁ, Robert Zawiasa, Marcelo Ferreira, Owen Ingraham, Daddy Wen, Tony Jimenez, Panther Modern, Jake Disco, Demilson Quintao, Penumbraa, Shuhong Chen, Hongbo Men, happi nyuu nyaa, Carol Lo, Mose Sakashita, Miguel, Bandera, Gennaro Schiano, gunwoo, Ravid Freedman, Mert Seftali, Mrityunjay, Richárd Nagyfi, Timo Steiner, Henrik G Sundt, projectAnthony, Brigham Hall, Kyle Hudson, Kalila, Jef Come, Jvari Williams, Tien Tien, BIll Mangrum, owned, Janne Kytölä, SO, Richárd Nagyfi, Hector, Drexon, Claxvii 177th, Inferencer, Michael Brenner, Akkusativ, Oleg Wock, FantomBloth, Thipok Tham, Clayton Ford, Theo, Handenon, Diego Silva, mayssam, Kadhai Pesalam, Tim Schulz, jiye, Anushka, Henrik Sundt, Julian Aßmann, Thomas Lin, Sid_Cypher, Mark Buckler, Kevin Tai, NO U, Gonzalo Fidalgo, Igor Alvarez, Alon Pluda, Clément Veyssière, Sander Zwaenepoel, etrotta, Binnie Yiu, Matej Macak, c zhou, Berhane-Meskel, sai sandeep mandava, Leo, Asad Dhamani, Charlie C, tantan assawade, Ângelo Fonseca, Stefan Lorenz, Paperboy, mika, Leo, Utsav Soi


[Discord] discord.gg/NhJZGtH
[Twitter] twitter.com/bycloudai
[Patreon] patreon.com/bycloud
[Business Inquiries] bycloudai@gmail.com
[Music] Massobeats - Glimmer
[Music] Massobeats - Honey jam
[Music] Massobeats - Lush
[Profile & Banner Art] twitter.com/pygm7
[Video Editor] @Booga04

[Bitcoin (BTC)] 3JFMJQVGXNA2HJE5V9qCwLiqy6wHY9Vhdx
[Ethereum (ETH)] 0x3d784F55E0bE5f35c1566B2E014598C0f354f190
[Litecoin (LTC)] MGHnqALjyU2W6NuJSSW9fTWV4dcHfwHZd7
[Bitcoin Cash (BCH)] 1LkyGfzHxnSfqMF8tN7ZGDwUTyBB6vcii9
[Solana (SOL)] 6XyMCEdVhtxJQRjMKgUJaySL8cGoBPzzA2NPDMPfVkKN
[Ko-fi] ko-fi.com/bycloudai
Cheating LLM Benchmarks Is Easier Than You Think…The RL Irony in LLMs (and its insane new meta)First Look At Metas Emu Edit & Emu VideoAI Generated Videos Are Getting Out of HandOpenAIs Deep Research Is MidGemini 2.5 Pro is just the best choice for AI right nowLLMs Are Better At Jailbreaking Themselves Than Us...New AI Paradigm?! Energy-Based Transformers ExplainedThese TikTokers Are AI Generated [SD Multi-Frame Rendering][VOD] Picking RTX4080 Super Giveaway Winner & ChillLlama 3.2 Deep Dive - Tiny LM & NEW VLM Unleashed By MetaSam Altman Sells Crypto, SDXL 1.0 Release, ChatGPT Custom Instructions & More [The AI Timeline #11]
bycloud |

Cheating LLM Benchmarks Is Easier Than You Think…

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER