Uploaded September 2025 | Updated September 2026, 1 hour ago
AI benchmarks are finally stepping into the real world.
OpenAI just dropped GDPval, a new benchmark measuring how well AI can perform actual economic tasks, not just toy problems. Think 1,320 tasks across 44 jobs in finance, law, engineering, media, and more—where outputs aren’t multiple-choice answers but briefs, spreadsheets, CAD diagrams, even social posts.
Here’s the wild part:
✅ Average expert needs ~7 hours per task, worth ~$398 in wages
⚡ Models finish 100x faster and cheaper (execution only, no human review)
📂 68% of tasks require working with files—photos, videos, audio, documents—making this truly multi-modal
Why it matters: this shifts AI evaluation from abstract leaderboards to tangible economic impact. It shows how far models have come in accuracy and instruction-following, while also exposing their limits (no collaboration, no interpersonal skills, no iterative drafts).
And yes, OpenAI open-sourced a 220-task gold set so anyone can build on it.
AI isn’t replacing workers yet, but benchmarks like this change how we measure progress. From toy puzzles to real productivity.
I’m Louis-François, PhD dropout, now CTO & co-founder at Towards AI. Follow me for tomorrow’s no-BS AI roundup 🚀
#AInews #OpenAI #ArtificialIntelligence #short
AI benchmarks are finally stepping into the real world.
OpenAI just dropped GDPval, a new benchmark measuring how well AI can perform actual economic tasks, not just toy problems. Think 1,320 tasks across 44 jobs in finance, law, engineering, media, and more—where outputs aren’t multiple-choice answers but briefs, spreadsheets, CAD diagrams, even social posts.
Here’s the wild part:
✅ Average expert needs ~7 hours per task, worth ~$398 in wages
⚡ Models finish 100x faster and cheaper (execution only, no human review)
📂 68% of tasks require working with files—photos, videos, audio, documents—making this truly multi-modal
Why it matters: this shifts AI evaluation from abstract leaderboards to tangible economic impact. It shows how far models have come in accuracy and instruction-following, while also exposing their limits (no collaboration, no interpersonal skills, no iterative drafts).
And yes, OpenAI open-sourced a 220-task gold set so anyone can build on it.
AI isn’t replacing workers yet, but benchmarks like this change how we measure progress. From toy puzzles to real productivity.
I’m Louis-François, PhD dropout, now CTO & co-founder at Towards AI. Follow me for tomorrow’s no-BS AI roundup 🚀
#AInews #OpenAI #ArtificialIntelligence #short










