Infra behind Krea 2: How to train and serve at scale — Gabriel Jorge Menezes, Krea.ai @aiDotEngineer
Infra behind Krea 2: How to train and serve at scale — Gabriel Jorge Menezes, Krea.ai  @aiDotEngineer
Uploaded August 2026 | Updated September 2026, 3 weeks ago
GPU utilization is a lie. It read 100% straight through pretraining while the cluster was nowhere near well used, so Gabriel Jorge Menezes tracks tensor core utilization instead, and watched it climb as training resolution stepped from 128 pixels up to 1024. That is one of several numbers he argues you cannot train at this scale without. InfiniBand counters are exported by nothing off the shelf, and most of their failures turned out to be cross node communication, so they built that collection themselves. Any GPU running hotter than 78 degrees gets pulled rather than debugged, because one warm card throttles and destabilizes the entire run.

This is the infrastructure half of Krea 2, the model trained from scratch on thousands of GPUs. Crashes scaled with the cluster and often failed silently, with communication timing out while every dashboard stayed green, and the practical answer was to stop treating each one as a mystery. Let it crash, and the same nodes running the same code will frequently go 24 hours on the next attempt. What made that survivable was checkpointing aggressively against a filesystem quick enough to write a terabyte in under 30 seconds. Production and training then share one cluster, with training holding priority and inference evicted to outside providers through a fake Kubernetes node, migrated back gradually rather than all at once so the site never drops.

Speaker info:
- linkedin.com/in/gabriel-jorge-menezes
- gab-menezes.github.io

Timestamps:
0:00 - Krea 2, trained from scratch, and two open checkpoints
3:26 - Crashes at scale, and the silent ones
4:18 - Metrics are everything, starting with temperature
5:58 - GPU utilization is a lie, use tensor cores
6:48 - InfiniBand and NVLink metrics you have to build yourself
8:29 - Checkpointing hard against a fast filesystem
9:21 - Gang scheduling, and training outranking production
11:01 - Flipping inference out through a fake node
14:23 - Taints that stop you wasting GPUs
16:03 - Inference runs on almost any GPU
Infra behind Krea 2: How to train and serve at scale — Gabriel Jorge Menezes, Krea.aiWhats Next After RLHF? — Diogo Almeida, TypeSafe AIVending-Bench: Long-Horizon Agent Evals — Lukas Petersson, Andon LabsAgentic Development Security — Ezra Tanzer, SnykTribal Dungeons of Global Shipping: AI Agents at Global Scale — Dmitry Buykin, MaerskFrom Tokenmaxxing to Trusted Throughput — Mingsheng Hong, IroncladAI Consulting in Practice – NLW, Superintelligent, @AIDailyBrief⁩Data and Environment Curation for Post-Training LLMs — Mahesh Sathiamoorthy, Bespoke LabsPerceptual Evaluations: Evals for Aesthetics — Diego Rodriguez, Krea.aix402 isn’t good (yet) — Jan Curn, ApifyAI Agents Are Just Distributed Systems Now — Salman Munaf, TikTokWhy Large? Tiny LMs & Agents on Edge/Robotics — Cormac Brick, Google
AI Engineer |

Infra behind Krea 2: How to train and serve at scale — Gabriel Jorge Menezes, Krea.ai

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER