Uploaded July 2026 | Updated September 2026, 2 weeks ago
In this video, we break down Poolside’s Laguna S2.1, an open-weights 118B MoE coding model (8B active per token) with a 1M-token context window, and why it performs above its size on agentic coding benchmarks like Terminal Bench 2.1. I cover how it was trained with reinforcement learning in FP8 on ~4,000 NVIDIA H200s in under nine weeks, plus how they tackled reward hacking on SWE-bench using an external LLM judge, prompt amendments, and network-blocked sandboxes.
Thanks to @NVIDIADeveloper for DGX Spark.
Laguna: poolside.ai/blog/introducing-laguna-s-2-1
DGX Spark: nvda.ws/3XIkwsh
Try it out: chat.poolside.ai
Pool Agent Harness: poolside.ai/get-started
vLLM Serving: github.com/MiaAI-Lab/Laguna-S-2.1-DGX-Spark-RTX-6000-PRO
DSpark video: youtu.be/eFgknPFK-g0
MoE Quantization paper: arxiv.org/pdf/2606.00206
My voice to text App: whryte.com
Website: engineerprompt.ai
RAG Beyond Basics Course:
prompt-s-site.thinkific.com/courses/rag
Signup for Newsletter, localgpt:
https://tally.so/r/3y9bb0
Let's Connect:
🦾 Discord: discord.com/invite/t4eYQRUcXB
☕ Buy me a Coffee: ko-fi.com/promptengineering
|🔴 Patreon: patreon.com/PromptEngineering
💼Consulting: calendly.com/engineerprompt/consulting-call
📧 Business Contact: engineerprompt@gmail.com
Become Member: tinyurl.com/y5h28s6h
💻 Pre-configured localGPT VM: bit.ly/localGPT (use Code: PromptEngineering for 50% off).
Signup for Newsletter, localgpt:
https://tally.so/r/3y9bb0
00:00 Laguna S2.1 Overview
01:17 Benchmarks and Harness
02:13 Reinforcement Learning
03:34 Reward Hacking Fixes
05:16 Running on DGX Spark
06:55 NVFP4 Quantization
08:03 Speculative Decoding Speed
10:09 Pool Harness Demo
12:07 Verbose Reasoning Loops
In this video, we break down Poolside’s Laguna S2.1, an open-weights 118B MoE coding model (8B active per token) with a 1M-token context window, and why it performs above its size on agentic coding benchmarks like Terminal Bench 2.1. I cover how it was trained with reinforcement learning in FP8 on ~4,000 NVIDIA H200s in under nine weeks, plus how they tackled reward hacking on SWE-bench using an external LLM judge, prompt amendments, and network-blocked sandboxes.
Thanks to @NVIDIADeveloper for DGX Spark.
Laguna: poolside.ai/blog/introducing-laguna-s-2-1
DGX Spark: nvda.ws/3XIkwsh
Try it out: chat.poolside.ai
Pool Agent Harness: poolside.ai/get-started
vLLM Serving: github.com/MiaAI-Lab/Laguna-S-2.1-DGX-Spark-RTX-6000-PRO
DSpark video: youtu.be/eFgknPFK-g0
MoE Quantization paper: arxiv.org/pdf/2606.00206
My voice to text App: whryte.com
Website: engineerprompt.ai
RAG Beyond Basics Course:
prompt-s-site.thinkific.com/courses/rag
Signup for Newsletter, localgpt:
https://tally.so/r/3y9bb0
Let's Connect:
🦾 Discord: discord.com/invite/t4eYQRUcXB
☕ Buy me a Coffee: ko-fi.com/promptengineering
|🔴 Patreon: patreon.com/PromptEngineering
💼Consulting: calendly.com/engineerprompt/consulting-call
📧 Business Contact: engineerprompt@gmail.com
Become Member: tinyurl.com/y5h28s6h
💻 Pre-configured localGPT VM: bit.ly/localGPT (use Code: PromptEngineering for 50% off).
Signup for Newsletter, localgpt:
https://tally.so/r/3y9bb0
00:00 Laguna S2.1 Overview
01:17 Benchmarks and Harness
02:13 Reinforcement Learning
03:34 Reward Hacking Fixes
05:16 Running on DGX Spark
06:55 NVFP4 Quantization
08:03 Speculative Decoding Speed
10:09 Pool Harness Demo
12:07 Verbose Reasoning Loops










