Benchmarking AI Agents Against Realistic Analytical Tasks with ADE-bench @aicouncilconf
Benchmarking AI Agents Against Realistic Analytical Tasks with ADE-bench  @aicouncilconf
Uploaded June 2026 | Updated September 2026, 2 weeks ago
[2026 - DAY 2 - CODING AGENTS] There are many benchmarks that attempt to measure how well LLMs and AI agents can write SQL queries or do complicated statistical analysis. But as most practitioners know, this is only a small part of our job. Before we can write a query, we have to figure out the business context behind the question. We have figure out which tables to use in a messy database. We have to make subjective decisions about vaguely defined problems. All of this makes benchmarking analytical agents difficult.

We built a new benchmark—ADE-bench—that aspires to do exactly that. It gives agents complex analytical environments to work in and ambiguous tasks to solve, and measures how well they perform.

In this talk, we'll share how we built the benchmark, the results of our tests, a bunch of things we learned along the way, and what we think is coming next.

The benchmark harness is open source, and can be found here: github.com/dbt-labs/ade-bench

SPEAKERS:
Benn Stancil - Founder, Mode (acquired by ThoughtSpot)
Jason Ganz - Director, DX + AI, dbt Labs

👉 Sign up for our "No BS" Newsletter to get the latest technical data & AI content: aicouncil.com/newsletter

ABOUT AI COUNCIL:
AI Council brings together the brightest minds in data to share industry knowledge, technical architectures and best practices in building cutting edge data & AI systems and tools.

FIND US:
Website: aicouncil.com
LinkedIn: linkedin.com/company/aicouncilconf
X: https://x.com/aicouncilconf
Benchmarking AI Agents Against Realistic Analytical Tasks with ADE-benchRevolutionize AI Engineering with AutoGenCausal Inference Methods for Bridging Experiments and Strategic ImpactThe Modern Data Stack Lost the War: Stop Building more DataFrame APIs | OpenAIMore Than Query  Future Directions of Query Languages, from SQL to MorelA 101 in Time Series Analytics with Apache Arrow, Pandas and ParquetAn Opinionated Blueprint for Developing Production AI ApplicationsTrillion is the New Billion: Managing Really Large Multimodal Datasets for AI | LanceDBHow Product and Research Build Together at the Frontier | HexAI Launchpad 2026: CocoIndexOramaCore: A Search Database with LLMs Built In
AI Council |

Benchmarking AI Agents Against Realistic Analytical Tasks with ADE-bench

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER