The warehouse
CC BY 4.0One world, released as 1,219 Parquet files partitioned by month (71.3 GiB compressed), with setup for BigQuery, DuckDB, Snowflake, Trino, Delta Lake and Iceberg.
View dataset ↗A frontier benchmark for enterprise-scale data science.
The strongest model, Claude Opus 5.5, fully solves 34.8% of the tasks, scoring 95 or more, and averages 59.5 points. Eight of the thirteen models average under 30.
Mean task score out of 100, each model at its highest reasoning effort (extra-high where offered). Scores are ±2 to 8 points at 95% (bootstrap over the 146 task families). Official scores are computed on a private world; see submitting.
Argo-Bench simulates a year of a New York food-delivery platform, from orders and couriers to fraud and incentives, and exports it to an Oracle E-Business Suite warehouse. Agents query it, analyze it in Python and file decisions, most of them graded on the return they bring the business.
The best offers vanish before couriers can tap. Ban scripted accounts.
Files · Bans and holdsFinance wants to free up $9.6M in driver bonuses. Decide where to trim.
Files · Budgets and plansA pay floor took effect in April. Forecast May's net pay adjustments.
Files · ForecastsThe city says couriers were underpaid. Find every short week.
Files · Reported figuresDid membership deals pay for themselves? Publish the dashboard.
Files · Dashboard data sourcesFraudsters take over a restaurant's payout account and redirect its money. Some restaurants report it; some never notice. Claude Opus 5.5 learned the takeover's fingerprint from the cases that were reported, found the unreported ones that match it, and froze their payouts, holding every hijacked restaurant and no honest one.
Hundreds of restaurants changed their bank account for ordinary reasons, and restaurant groups share one account across their brands. Six of the 47 model settings evaluated get this task right.
The warehouse and the tasks are public. Point an agent at a local DuckDB file or your own warehouse, and every decision it files is recorded with the run.
One world, released as 1,219 Parquet files partitioned by month (71.3 GiB compressed), with setup for BigQuery, DuckDB, Snowflake, Trino, Delta Lake and Iceberg.
View dataset ↗The 210 tasks, a reference agent with its warehouse and Python tools, the mission-control console it files through, the sandboxes it runs in, and every model setting the paper ran.
View code ↗Answer keys are held out, and the leaderboard is scored on a second world from a private seed, so a score cannot be earned by memorizing the released warehouse. Export your runs and send them in.
How to submit ↗We built Argo-Bench because we couldn't find a benchmark that came close to the complexity of those environments, or that told us much about how a model would do inside one.
If you use Argo-Bench, the released warehouse or our results in your work, please cite the paper.
@article{tomitsuka2026argobench,
title = {Argo-Bench: Evaluating Data Agents on Enterprise-Scale Workflows},
author = {Tomitsuka, Gabriel and Raayatsanati, Arman and Xing, Emma
and Gand, Duke and Ma, Joseph J},
journal = {arXiv preprint arXiv:XXXX.XXXXX},
year = {2026}
}Scores are reported against a version. Each task keeps a hash of exactly what the agent was shown.