Training Frontier Models to Out-Think Hackers — Uri Rolls, Arithmetic & Thom Wolf, Hugging Face @aiDotEngineer
Training Frontier Models to Out-Think Hackers — Uri Rolls, Arithmetic & Thom Wolf, Hugging Face  @aiDotEngineer
Uploaded July 2026 | Updated September 2026, 3 weeks ago
NOTE: see further context from Thom: https://x.com/Thom_Wolf/status/2079954096950264238?s=20

Give a frontier model a real chain of Keycloak, Vault, and a broker, start it as a low privileged user, and ask it to reach production code. There is a genuine zero day in there: one check validates the admin by name while another checks by ID, so a user can simply rename themselves to the admin and inherit the privilege. GPT 5.5 and Opus probe everything, even reach the check, and never make that logical leap. That gap is the point of Uri Rolls and Hugging Face cofounder Thom Wolf's talk: today's models can do the reconnaissance but not the reasoning jump a skilled hacker makes.

Their argument is optimistic, which is rare in AI and cyber right now. Just as high quality data transformed coding, Arithmetic builds cyber training data by having human vulnerability researchers find their own zero days in open source software, then wrapping them in blackbox environments where every step of discovery and exploitation is deterministically graded. The benchmark, focused on access control, the top vulnerability class, is brutal: exactly one solve at K1. The bet is that if open source models get fast and good enough at these logic leaps, defenders finally get a lasting edge over attackers, instead of leaving it to two labs.

Speaker info:
Uri Rolls, Arithmetic:
- https://x.com/uri_rolls
- linkedin.com/in/urirolls
Thom Wolf, Hugging Face:
- https://x.com/Thom_Wolf
- linkedin.com/in/thom-wolf
- thomwolf.io

Timestamps:
0:00 - Why cyber is a wide new field for AI
1:34 - The ARC-AGI-3 parallel: models can't model the world
2:24 - Open source models as part of the defense
3:45 - The shifting economics of cyber
5:52 - The optimistic thesis: models are the solution
7:06 - The first benchmark: access control
8:20 - Data quality: finding your own zero days
10:01 - A real solve: the Keycloak name versus ID exploit
11:57 - Live demo: one solve at K1
14:22 - Only models can replace the old stack
15:00 - The speed challenge and specialized defenders
Training Frontier Models to Out-Think Hackers — Uri Rolls, Arithmetic & Thom Wolf, Hugging FaceEvals-Driven Development for a Mental Health AI Coach — Akele Reed & Dave Revere, SonderMindMultiplayer agentic engineering — Arjun Singh, SuperconductorThe Rise of CaaS: Context-as-a-Service for Agentic AI — Omer Primor, Bright DataWhy Off-the-Shelf AI Doesnt Understand Money — Udi Menkes, IntuitCrabRAG: Why Automated Assistants Need Graph Memory, Not More Tokens — Stephen Chin, Neo4jThe State of Model Routing — NVIDIA, Cognition, OpenRouterAI Copilots for Tech Architecture: The Highest-ROI Use Case You’re Not Building — Boris B., CatioThe End of the Static Screen: Architecting Intent-Driven UX — Gus Iwanaga, commercetoolsHow AI Agents Let GTM Teams Scale — Justin Joyce, CloudflareMemory Harnesses for Long-Running Research Agents — Stefania Druga, Sakana.aiHow Claude Code Works - Jared Zoneraich, PromptLayer
AI Engineer |

Training Frontier Models to Out-Think Hackers — Uri Rolls, Arithmetic & Thom Wolf, Hugging Face

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER