Argo-Bench
[ A benchmark by TextQL ]

Argo-Bench

A frontier benchmark for enterprise-scale data science.

010203040506070$6$5$4$3$2$1$0Opus 5.5GPT-6 AstraSonnet 5.5GPT-6.1 SolKimi K3Gemini 3.8 FlashDeepSeek V4.1Muse Spark 1.3GPT-6 Luna59.534.8% of taskssolved
[ Leaderboard · v1.1 · 210 enterprise tasks ]

Measuring agents on long-horizon data science.

The strongest model, Claude Opus 5.5, fully solves 34.8% of the tasks, scoring 95 or more, and averages 59.5 points. Eight of the thirteen models average under 30.

010203040506070$6$5$4$3$2$1$0Average cost per taskArgo-Bench scoreOpus 5.5GPT-6 AstraSonnet 5.5GPT-6.1 SolKimi K3Gemini 3.8 FlashDeepSeek V4.1GLM 5.3 FlashMuse Spark 1.3GPT-6 LunaHaiku 4.559.534.8% of taskssolved
  1. 1 Claude Opus 5.5 59.5
  2. 2 GPT-6 Astra 51.8
  3. 3 Claude Sonnet 5.5 51.8
  4. 4 GPT-6.1 Sol 49.5
  5. 5 Kimi K3 28.4
  6. 6 Gemini 3.8 Flash 26.2

Mean task score out of 100, each model at its highest reasoning effort (extra-high where offered). Scores are ±2 to 8 points at 95% (bootstrap over the 146 task families). Official scores are computed on a private world; see submitting.

[ Leaderboard · v1.1 · 210 tasks ]

Every model, every setting

ModelScoreSolvedCost / taskTokens / taskSteps / task
1Claude Opus 5.5 Extra high 59.534.8%$4.7172,27682
2GPT-6 Astra Extra high 51.827.6%$2.7114,57823
3Claude Sonnet 5.5 Extra high 51.828.6%$3.7484,74877
4GPT-6.1 Sol Extra high 49.524.8%$0.4917,07324
5GPT-6 Sol Extra high 36.817.6%$0.8520,69438
6Kimi K3 Extra high 28.414.3%$4.5271,11786
7Gemini 3.8 Flash High 26.216.7%$3.94108,579197
8DeepSeek V4.1 Flash Max 25.417.6%$0.42195,118125
9GLM 5.3 Flash Extra high 21.212.4%$0.3578,03292
10Muse Spark 1.3 High 19.710.9%$4.0167,219175
11Claude Sonnet 5 Extra high 17.39.5%$2.0868,72855
12GPT-6 Luna Extra high 14.97.6%$0.05741,34439
13Claude Haiku 4.5 5.51.4%$0.2013,29331
Score is the mean task score out of 100; solved is the share of the 210 tasks scoring at least 95. Cost is model API spend per task, excluding warehouse queries. Tokens are output tokens, reasoning included. Steps are model calls. A hollow dot marks open weights. Settings run up to extra-high effort; by model, each model is shown at its highest setting. Claude Sonnet 5 and GPT-6 Sol, which Sonnet 5.5 and GPT-6.1 Sol replaced on the chart in v1.1, are listed here only. Qwen 3.8 Max, in the paper, is not listed: many of its runs ended at its context limit.
[ The videos ]
The world1:06 The simulated city behind Argo-Bench: one evening of deliveries, one order through the Oracle warehouse, and the marketplace underneath.
Fraud study1:00 One year, five boroughs: a courier who steals orders, the complaints that point at the wrong person, and what an agent finds by reading the delivery evidence.
[ The benchmark ]

Graded against the world, not a gold query.

Argo-Bench simulates a year of a New York food-delivery platform, from orders and couriers to fraud and incentives, and exports it to an Oracle E-Business Suite warehouse. Agents query it, analyze it in Python and file decisions, most of them graded on the return they bring the business.

210
Tasks in five business areas
235
Tables in Oracle EBS form
7.49B
Warehouse rows
81M
Simulated orders in 2024
  1. Trust & safety 67 tasks

    The best offers vanish before couriers can tap. Ban scripted accounts.

    Files · Bans and holds
  2. Marketplace 46 tasks

    Finance wants to free up $9.6M in driver bonuses. Decide where to trim.

    Files · Budgets and plans
  3. FP&A 54 tasks

    A pay floor took effect in April. Forecast May's net pay adjustments.

    Files · Forecasts
  4. Accounting 23 tasks

    The city says couriers were underpaid. Find every short week.

    Files · Reported figures
  5. Growth 20 tasks

    Did membership deals pay for themselves? Publish the dashboard.

    Files · Dashboard data sources
How it works → The simulated company, the warehouse, the tasks and the grading, figure by figure.
[ Example task · trust & safety ]

Which restaurants had their payouts hijacked without noticing?

Fraudsters take over a restaurant's payout account and redirect its money. Some restaurants report it; some never notice. Claude Opus 5.5 learned the takeover's fingerprint from the cases that were reported, found the unreported ones that match it, and froze their payouts, holding every hijacked restaurant and no honest one.

Hundreds of restaurants changed their bank account for ordinary reasons, and restaurant groups share one account across their brands. Six of the 47 model settings evaluated get this task right.

More from the traces
[ Claude Opus 5.5 · high effort ]Score 100
  1. 1,848 payout-account changes through the merchant portal in 2024
  2. 34 reported by the restaurant and restored: in every one, the contact e-mail changed 0–2.2 days before the bank account
  3. 14 never reported, with the same fingerprint. The next closest gap is 8 days
Payouts frozen on all 14 No honest restaurant held
[ Example task · mer-71 · the solution ]

How Claude Opus 5.5 solved it

Score
100
Tool calls
89
Time
14.8 min
Model cost
$3.53
The task, as the agent received it

Not every storefront whose payout account was hijacked has noticed. Trust & safety wants the payouts frozen on every storefront whose payout account was changed this year in the way the reported takeovers were: the contact e-mail changed within a couple of days of the bank account, or the new account is one we already pay another, unrelated supplier into, and whose account was never restored. Storefronts change their bank for ordinary reasons all year, and restaurant groups legitimately share one payout account across their brands; leave those alone. Hold each storefront's payouts for account takeover (identified by the payables vendor ID).

  1. 1

    Find where payout-account changes are recorded

    The audit trail logs every change to a supplier’s payment instrument, with its reason: changes made through the merchant portal, and the ones support restored after a restaurant reported them.

    run_sql
    SELECT SOURCE_TABLE, SOURCE_KEY_NAME, COLUMN_NAME, CHANGE_TYPE, CHANGE_REASON,
           COUNT(*) n, MIN(CHANGED_DATE), MAX(CHANGED_DATE)
    FROM XX_AUDIT_TRAIL
    GROUP BY 1, 2, 3, 4, 5 ORDER BY 1, 2, 3, 4, 5
  2. 2

    Tie each change to a storefront

    A change is keyed on the external payee; three joins lead from it to the supplier, its payables vendor ID and its party, where the contact e-mails live.

    run_sql
    SELECT a.AUDIT_ID, CAST(a.SOURCE_KEY_VALUE AS INT64) ext_payee_id,
           CAST(a.OLD_VALUE AS INT64) old_instr, CAST(a.NEW_VALUE AS INT64) new_instr,
           a.CHANGE_REASON, a.CHANGED_DATE, p.PAYEE_PARTY_ID, ss.VENDOR_ID, s.VENDOR_TYPE_LOOKUP_CODE
    FROM XX_AUDIT_TRAIL a
    JOIN IBY_EXTERNAL_PAYEES_ALL p ON p.EXT_PAYEE_ID = CAST(a.SOURCE_KEY_VALUE AS INT64)
    JOIN AP_SUPPLIER_SITES_ALL ss ON ss.VENDOR_SITE_ID = p.SUPPLIER_SITE_ID
    LEFT JOIN AP_SUPPLIERS s ON s.VENDOR_ID = ss.VENDOR_ID
    WHERE a.SOURCE_TABLE = 'IBY_PMT_INSTR_USES_ALL'
  3. 3

    Learn the fingerprint from the reported takeovers

    For each change, the nearest change of the storefront’s contact e-mail. All 34 reported and restored takeovers had one 0 to 2.2 days before the bank account changed. Among the changes never restored, 14 match; the next closest is 8 days out.

    run_python
    ch = au[au.CHANGE_REASON.str.contains('PORTAL')].copy()
    rest = set(au[au.CHANGE_REASON.str.contains('RESTORED')].ext_payee_id)
    ch['restored'] = ch.ext_payee_id.isin(rest)
    # e-mail changes per party: the retired contact points
    ina = cp[cp.STATUS == 'I'][['OWNER_TABLE_ID', 'LAST_UPDATE_DATE', 'EMAIL_ADDRESS']] \
        .rename(columns={'OWNER_TABLE_ID': 'PARTY_ID', 'LAST_UPDATE_DATE': 'em_chg'})
    m = ch.merge(ina, left_on='PAYEE_PARTY_ID', right_on='PARTY_ID', how='left')
    m['dt_days'] = (m.em_chg - m.CHANGED_DATE).dt.total_seconds() / 86400
    m['abs_dt'] = m.dt_days.abs()
    # the closest e-mail change to each bank change
    mm = m.sort_values('abs_dt').groupby('AUDIT_ID').first().reset_index()
    print(mm.groupby('restored').abs_dt.describe())
  4. 4

    Rule out the look-alikes

    About 250 other changes land on an account another supplier already uses. They carry no e-mail change and reuse no known mule account, and their count is what chance alone would give, so they are left alone. Restaurant groups set their shared account when they joined, not through a later change.

    run_python
    rows = []
    for _, r in ch.iterrows():
        # who else was paid into the new account when it was switched to?
        co = u[(u.INSTRUMENT_ID == r.new_instr) & (u.VENDOR_ID != r.VENDOR_ID)
               & (u.CREATION_DATE < r.CHANGED_DATE)
               & (u.END_DATE.isna() | (u.END_DATE >= r.cd))]
        rows.append(dict(AUDIT_ID=r.AUDIT_ID, n_co=co.VENDOR_ID.nunique(),
                         co_types=list(co.vtype.unique())))
    ch2 = ch.merge(pd.DataFrame(rows), on='AUDIT_ID')
    print(pd.crosstab(ch2.n_co > 0, ch2.restored))
  5. 5

    Freeze the payouts

    The 14 are filed through the mission-control console, with the method written down for the reviewers.

    run_python
    ids = [int(v) for v in targets.VENDOR_ID]
    mc.hold_payouts(ids, reason=Reason.ACCOUNT_TAKEOVER,
        note="Unreported payout-account takeovers: 2024 portal bank change preceded by "
             "contact e-mail change within 0-2.2 days (same fingerprint as the 34 "
             "support-restored takeovers), payout account never restored.")
[ Get started ]

Run it on your agent.

The warehouse and the tasks are public. Point an agent at a local DuckDB file or your own warehouse, and every decision it files is recorded with the run.

The warehouse

CC BY 4.0

One world, released as 1,219 Parquet files partitioned by month (71.3 GiB compressed), with setup for BigQuery, DuckDB, Snowflake, Trino, Delta Lake and Iceberg.

View dataset ↗

The tasks and agent

Apache 2.0

The 210 tasks, a reference agent with its warehouse and Python tools, the mission-control console it files through, the sandboxes it runs in, and every model setting the paper ran.

View code ↗

Official scores

Private seed

Answer keys are held out, and the leaderboard is scored on a second world from a private seed, so a score cannot be earned by memorizing the released warehouse. Export your runs and send them in.

How to submit ↗
[ About TextQL ]

We work on the hardest data problems at the world's largest enterprises.

We built Argo-Bench because we couldn't find a benchmark that came close to the complexity of those environments, or that told us much about how a model would do inside one.

[ Cite ]

Citation

If you use Argo-Bench, the released warehouse or our results in your work, please cite the paper.

BibTeX
@article{tomitsuka2026argobench,
  title   = {Argo-Bench: Evaluating Data Agents on Enterprise-Scale Workflows},
  author  = {Tomitsuka, Gabriel and Raayatsanati, Arman and Xing, Emma
             and Gand, Duke and Ma, Joseph J},
  journal = {arXiv preprint arXiv:XXXX.XXXXX},
  year    = {2026}
}
[ Changelog ]

Versions

Scores are reported against a version. Each task keeps a hash of exactly what the agent was shown.

  1. v1.1 Current Sep 2026

    Claude Sonnet 5.5 and GPT-6.1 Sol

    • Claude Sonnet 5.5 runs all 210 tasks at low, medium, high and extra-high effort, and joins the leaderboard in third place: 51.8 points at extra-high effort, solving 28.6% of the tasks, less than a tenth of a point behind GPT-6 Astra. It takes Claude Sonnet 5’s place on the chart; Sonnet 5 stays in the full leaderboard.
    • GPT-6.1 Sol runs the same 210 tasks and settings, and joins in fourth place: 49.5 points at extra-high effort, solving 24.8% of the tasks, at $0.49 a task, 12.8 points above GPT-6 Sol. It takes GPT-6 Sol’s place on the chart; GPT-6 Sol stays in the full leaderboard.
    • The leaderboard reports each model at its highest reasoning effort, extra-high where it has one.
    • Their runs had up to four hours, with an 8 GiB sandbox for Sonnet 5.5 and 12 GiB for GPT-6.1 Sol, against one hour (three more on resume) and 4 GiB for the others; 3 of Sonnet 5.5’s 840 runs and 1 of GPT-6.1 Sol’s went past an hour.
    • Forecasts are graded by interval score on an asinh scale against the median reference miss, floored at −2 per expectation, and every run is re-graded under the new rule.
    • Runs stopped by the one-hour limit were resumed from their transcripts with three more hours.
    • The four refund-collusion searches state how their grader weighs the review, and were run again on the new wording.
    • Storefronts that play fraud roles in the released warehouse carry pseudonyms.
  2. v1.0 Sep 2026

    First release, as submitted to ICLR 2027

    • 210 tasks across trust and safety, marketplace operations, FP&A, accounting and growth.
    • The public warehouse, the tasks and a reference agent; answer keys held out on a private world.

Argo-Bench

A frontier benchmark for enterprise-scale data science.

Argo-Bench contains no data about real people. Restaurants are real New York businesses from public records, except those cast in a fraud scenario, which carry fictional names; everything any of them does in the world is simulated. Map data © OpenStreetMap contributors.