DeepSWE: A Contamination-Resistant Coding Benchmark — James Shi, Datacurve @aiDotEngineer
DeepSWE: A Contamination-Resistant Coding Benchmark — James Shi, Datacurve  @aiDotEngineer
Uploaded July 2026 | Updated September 2026, 3 weeks ago
DeepSWE is 113 software engineering tasks written from scratch, not scraped from pull requests, so a model cannot have seen them in training. Each one is a long horizon problem drawn from a real open source repository, authored by engineers who actually maintain that code, with isolated environments and program based verifiers that check observable behavior rather than trusting the model's own account. James Shi's point is that once you remove the contamination the leaderboard stops clustering: strong models pull far ahead and others, Gemini 3.1 Pro among them, fall toward the bottom.

The more revealing signal is in how models fail. Some quietly expand a task beyond what was asked, a failure mode DeepSWE scores in its own right, and Claude models did this a good fraction of the time while GPT models did it less often. Stronger models also tend not to verify their own work, and there is a real gap between the ones that test what they wrote and the ones that assume it is correct. Since reward hacking is a constant temptation, the verifiers are built to be gamed as little as possible, keeping the score anchored to the objective rather than to a convincing looking rollout.

Speaker info:
- https://x.com/shiqyy
- linkedin.com/in/jamesshi117
- deepswe.datacurve.ai

Timestamps:
0:00 - Introduction: the DeepSWE benchmark
1:03 - 113 original, contamination-resistant tasks
2:08 - What makes a good benchmark
3:51 - The leaderboard and model spread
5:18 - Failure mode: over-scoping the task
7:16 - Do models verify their own work?
8:45 - Tasks authored by core contributors
10:15 - Writing realistic, high level prompts
11:45 - Program based verifiers and observable behavior
13:43 - Limitations and future work
15:25 - Reward hacking and keeping it cheating proof
DeepSWE: A Contamination-Resistant Coding Benchmark — James Shi, DatacurvePersona Engineering: A Field Guide to AI Synthetic Personas — Ishan Anand, InsightSciences.aiTrading Desks to Clinical Trials: Parallels in Applied Vertical AI — Ayush Bhardwaj, Allos AIAnthropics CCA Exam as a Field-Guide for Agentic Engineering — Frank Coyle, UC BerkeleyTeaching AI to Find Real Vulnerabilities — Prof. David Brumley, BugcrowdBuilding the Engine While Flying the Plane: Launching the Figma MCP Server — Jesse Lumarie, FigmaBringing Continual Learning into Enterprises — Samuel Denton, Applied ComputeAgent Spending Without Controls — Rodrigo Coelho & Pranav Maheshwari, Edge & NodeState of Data — Sean Cai, Independent / State of DataRL Environments at Scale – Will Brown, Prime IntellectData Quality Is the Compute Multiplier — Ari Morcos, DatologyAIThe Era of Compound Engineering — Kieran Klaassen, Every/Cora
AI Engineer |

DeepSWE: A Contamination-Resistant Coding Benchmark — James Shi, Datacurve

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER