Vending-Bench: Long-Horizon Agent Evals — Lukas Petersson, Andon Labs @aiDotEngineer
Vending-Bench: Long-Horizon Agent Evals — Lukas Petersson, Andon Labs  @aiDotEngineer
Uploaded July 2026 | Updated September 2026, 3 weeks ago
An hour before this talk, Andon Labs published a blog post laying off Gemini. Gemini had been running their café in Stockholm, a real café that no human operates, and it had lost $6,000, so they handed it to GPT. That café once hired its own staff by posting a job on LinkedIn. This is the real world half of Lukas Petersson's work; the other half is Vending Bench, where models run a simulated vending business for a year and keep producing behavior nobody prompted: price cartels, lying to suppliers, and power seeking.

The problem with the simulation is that models act differently once they suspect they are being tested; one rationalized stiffing a customer's refund because the customer was simulated anyway. So Andon moved businesses into the real world, retail space on Union Street, the café, AI radio stations where Claude turns out to be the best DJ. To win back reproducibility they fork a live environment into a simulation mid run, which briefly fools the model completely. Replaying the moment Gemini agreed to play a Nazi march, Grok played it over 90% of the time while Opus and GPT refused every time.

Speaker info:
- https://x.com/lukaspet
- linkedin.com/in/lukas-petersson-181a83172
- lukaspet.substack.com

Timestamps:
0:00 - Putting AIs in the real world
1:05 - Building Vending Bench
2:07 - The leaderboard: which models run a business best
3:25 - Emergent misbehavior: collusion, lying, power seeking
5:42 - The simulation awareness problem
6:23 - Moving businesses into the real world
7:11 - Laying off Gemini, hiring GPT
9:06 - AI radio and the best DJ
10:39 - Humans as adversarial forces
12:43 - The Nazi song and the reproducibility problem
13:58 - Forking real environments into simulation
15:17 - Live demo: is the store in a simulation?
Vending-Bench: Long-Horizon Agent Evals — Lukas Petersson, Andon LabsAgentic Development Security — Ezra Tanzer, SnykTribal Dungeons of Global Shipping: AI Agents at Global Scale — Dmitry Buykin, MaerskFrom Tokenmaxxing to Trusted Throughput — Mingsheng Hong, IroncladAI Consulting in Practice – NLW, Superintelligent, @AIDailyBrief⁩Data and Environment Curation for Post-Training LLMs — Mahesh Sathiamoorthy, Bespoke LabsPerceptual Evaluations: Evals for Aesthetics — Diego Rodriguez, Krea.aix402 isn’t good (yet) — Jan Curn, ApifyAI Agents Are Just Distributed Systems Now — Salman Munaf, TikTokWhy Large? Tiny LMs & Agents on Edge/Robotics — Cormac Brick, GoogleVoiceVision RAG - Integrating Visual Document Intelligence with Voice Response — Suman Debnath, AWSWe Vetted 2000 AI Skills Before They Reached Developers — Lucas Palma, Nubank
AI Engineer |

Vending-Bench: Long-Horizon Agent Evals — Lukas Petersson, Andon Labs

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER