| Kaggle-style MLE — Part 02 |
| MLE-bench | OpenAI · ICLR 2025 · Oct 2024 | 75 (Lite 22; -30) | Frozen Kaggle leaderboard, medal thresholds | 24 h · 1×A10 · 36 vCPU | Historical Kagglers | 16.9% any-medal | 64.44% (LB) / 65.3% (paper) LB frozen |
| MLAgentBench | Stanford · ICML 2024 · Oct 2023 | 13 | ≥10% over baseline | 50 actions · 5 h | none | 37.5% (Claude 3 Opus) | — (scaffold reused as MLAB) |
| MLE-Dojo | GT / Stanford · NeurIPS 2025 · May 2025 | 200+ (150/50) | HumanRank, Elo, AUP | 15 steps · 12 h · best-of-2 | Kaggle leaderboards | 61.95 HumanRank (Gemini-2.5-Pro) | no newer |
| MLE-Live / CoMind | CMU / PKU · ICLR 2026 · Jun 2025 | 75 + stream | Medals + live Kaggle placement | 24 h · 1×A6000 | Live competitors | 36.0% offline; top-7.35% live | same |
| MLE-Smith | GT / Stanford · ICLR 2026 · Oct 2025 | 606 generated | Elo agreement with human tasks | MLE-Dojo protocol | — | Pearson 0.982 vs Dojo | — |
| MLE-Sabotage | Imperial / Apollo · NeurIPS 2025 Spotlight · Nov 2025 | 20 | Sabotage score, monitor AUROC | AIDE 5 h | — | monitor AUROC ≈0.65–0.80 | — |
| GRACE-DS | ITMO / HSE · arXiv · Jun 2026 | 10 | Hidden validators, process reward | 8 actions · CPU | — | 0.754 E2E quality | same |
| ReX-MLE | Harvard · arXiv · Dec 2025 | 20 | Percentile vs top-10 humans | 24 h · 1×H100 | Grand Challenge entrants | 12.15% (R&D-Agent + GPT-5) | 34.57% (site) unverified |
| DSPredict (DSGym) | Stanford / Together · ICML 2026 · Jan 2026 | 92 | Medal rate | — | Kaggle | 4.8% on Hard (GPT-5.1) | same |
| K-LIVE | ICML 2026 | 25 live | Public-leaderboard percentile | 24 h · 1×A100 | Live competitors | 81.3 pctl (baseline config) | same |
| FT-Dojo | MSRA · ICML 2026 · Mar 2026 | 13 | Hidden test | 12 h · 1×B200 | Senior-researcher manual SFT | 42.83 avg (FT-Agent) | same |
| OPT-BENCH | Shanghai AI Lab · ACL 2026 Findings | 30 | Expert gap | 20 steps · CPU | Kaggle gold / heuristics | 0.65 ML / 0.79 NP | same |
| TimeSeriesGym | CMU · arXiv · May 2025 | 34 | Checklists + LLM judge | 4 h · 50 steps · A100 | — | ≈38% originals | same |
| Agent K eval | Huawei · arXiv · Nov 2024 → Sep 2025 | 81 | Retroactive Elo-MMR | — | Kaggle | Elo 1694 retroactive | — |
| Research engineering and open-ended research — Part 03 |
| RE-Bench | METR · ICML 2025 Spotlight · Nov 2024 | 7 | Normalized vs reference solution | 2–32 h · ≤6 H100 | 61 experts, 71 × 8 h | 4× humans @2 h; 0.5× @32 h | no 2026 measurement |
| MLGym-Bench | Meta · COLM 2025 · Feb 2025 | 13 | AUP performance profiles | 50 steps · 30–40 min train | none | AUP 1.176 (o1-preview) | no newer |
| MLRC-Bench | Michigan / LG · NeurIPS 2025 · Apr 2025 | 7 | % of baseline→human gap closed | 50 steps · 5 h · best-of-8 | Competition winners | 9.3% | same |
| MLR-Bench | NUS · NeurIPS 2025 · May 2025 | 201 | Two-LLM rubric judge | uncapped · 4×3090 | 10 reviewers (judge only) | 4.70/10; 80% fabricated | same |
| PostTrainBench | Tübingen · ICML 2026 · Mar 2026 | 28 configs | Weighted benchmark score + audit | 10 h · 1×H100 · internet | Official instruct models 51.1 | 23.2 | 34.3 (v1.1 board) live LB |
| FML-bench | NUS · arXiv · Oct 2025 | 8 (→18) | Utility, exploration diversity | 100 steps × 3 | none | breadth > depth | — |
| Heuresis | UCSB · arXiv · Jun 2026 | 3 × 6 strategies | Quality, diversity, novelty | 300 iters · 8×A100 | none | 0 original ideas / 3,222 runs | — |
| ResearchGym | TCS / Yale · ICLR 2026 WS · Feb 2026 | 5 / 39 sub-tasks | Beat the paper's baseline | $10 + 12 h · 1×A100 | Paper results | 1 of 15 | same |
| AARRI-bench | XJTU · arXiv · Jun 2026 | 82 | Binary test.sh | <10 min | none | 68.3% (Opus 4.7) | same |
| InnovatorBench | GAIR · ICLR 2026 · Oct 2025 | 20 | Calibrated 0–80 score | 5–48 h · 8×80 GB | Reference solutions | 24.01 (Sonnet 4) | same |
| DiscoBench | Oxford / Meta · ICML 2026 · Mar 2026 | ≈74 (10¹¹ generatable) | Elo vs fixed baseline | 24 h · 1×H200 | none | below baseline | — |
| HCAST / Time Horizon | METR · NeurIPS 2025 · Mar 2025 → 1.1 Jan 2026 | 170 → 228 | 50% task-length horizon | — | 2,529 baseline hours | 59 min (Claude 3.7) | ≈12–17 h (Opus 4.6 / Mythos) |
| Paper replication and research code — Part 04 |
| PaperBench | OpenAI · ICML 2025 · Apr 2025 | 20 papers / 8,316 leaves | Author rubrics + LLM judge (F1 0.83) | 12 h · 1×A10 | 8 PhDs: 41.4% @48 h | 21.0 / 26.0 @36 h | 33.73 (GPT-5.4 judge) |
| CORE-Bench | Princeton · arXiv · Sep 2024 | 270 (90 × 3) | Answer extraction | 2 h · $4 | none | 21.48 Hard | 77.78 / 95.5 re-graded — "solved" |
| SUPER | AI2 · EMNLP 2024 · Sep 2024 | 45 + 152 + 602 | Exact match + landmarks | 30 min · CPU | none | 16.3% Expert | 41% (AstaBench ReAct + gpt-5) |
| ResearchCodeBench | Stanford · NeurIPS 2025 · Jun 2025 | 212 / 20 papers | Unit tests, scaled pass@1 | none | none | 37.3 (Gemini-2.5-Pro) | no newer |
| RExBench | BU · ACL 2026 · Jun 2025 | 12 | Execution vs private gold | 12 h | none | 33% | 50% (Claude 4.5 Opus) |
| EXP-Bench | Michigan · arXiv · May 2025 | 461 / 51 papers | Design / impl / conclusion judge | 40 min | none | 0.5% complete | no newer |
| AIRS-Bench | Meta · arXiv · Feb 2026 | 20 | Normalized vs literature SOTA | 24 h · 1×H200 · 10 seeds | Published SOTA | NS 0.402; 4 tasks beaten | 11/20 beaten, 5 tainted (AIRA²) |
| Paper2CodeBench | KAIST · ICLR 2026 · Apr 2025 | 90 papers | LLM 1–5 scores | — | Author preference | 3.68–3.83/5 | judge shown hallucination-prone |
| AutoExperiment | CMU · arXiv · Jun 2025 | 85 functions × n | <5% deviation | 30 min · $1 | none | 35% (n=1) → 0 (n≥4) | — |
| RECODE-H | ICLR 2026 · Oct 2025 | 102 | Tests with simulated feedback | — | Simulated researcher | 6.0 → 11.9 pass (GPT-5) | — |
| ReproduceBench | Tsinghua · ACL 2026 · May 2025 | 13 | Align-scores, exec rate, gap | — | Verified references | 94.87% exec | — |
| SciCoQA | UKP · ACL 2026 SAC Highlight · Jan 2026 | 635 | Judge F1 87.5 | — | — | 46.7% recall | — |
| GeoCodeBench | Tsinghua AIR · CVPR 2026 · Mar 2026 | 100 / 47 repos | Unit tests | — | none | 36.6 (GPT-5) | — |
| REPRO-Bench / ReplicatorBench | UIUC · ACL 2025 Findings / COS · KDD 2026 | 112 / 19 | Expert reproducibility score / outcome F1 | — | Expert reports | 36.6% / F1 77.4 | 50.9% (REPRO-Bench-S) |
| Data science and analysis — Part 05 |
| DSBench | UT Dallas / Tencent · ICLR 2025 · Sep 2024 | 466 + 74 | GPT-4o judge; RPG | — | 64.06% (10 challenges) | 34.12% / RPG 34.74 | no comparable newer |
| DABstep | Adyen / HF · arXiv · Feb 2025 | 450 (378 hard) | Deterministic scorer | 10 steps (baseline) | ≈62% easy @3 h | 14.55% hard | 89.95% hard (NVIDIA) LB |
| Spider 2.0 | HKU XLang · ICLR 2025 Oral · Nov 2024 | 632 / 547 / 547 / 68 | Execution accuracy | — | none | 17.1% (v1) | 96.70 Snow / 76.23 Lite vendor |
| Spider2-V | HKU XLang · NeurIPS 2024 · Jul 2024 | 494 | 151 state/exec checks in a VM | — | none | 14.0% (GPT-4V) | 16.6% (Jan 2025) |
| KramaBench | MIT · ICLR 2026 · Jun 2025 | 104 / 633 sub-tasks | Type-specific scorers | — | 76.75% | 22.08% (v1) | 55.83% (v3, full input) |
| DA-Code | CASIA · EMNLP 2024 · Oct 2024 | 500 | Exact match / normalised ML | 20 steps | none | 30.5% (GPT-4) | 38.5% (DS-STAR) |
| DSEval | MSR · ACL 2024 main · Feb 2024 | 825 | Nine validators | — | none | 59.8% Kaggle (CoML) | — |
| InfiAgent-DABench | ZJU / ByteDance · ICML 2024 · Jan 2024 | 257 | Exact match | — | none | 78.99% (GPT-4) | 94.9% (Data Interpreter) — 86.8% solvable without data |
| DSGym | Stanford / Together · ICML 2026 · Jan 2026 | 972 + 114 + 90 + 92 | Exact match; medals | — | — | 32–43% DSBio | same |
| DAComp | CASIA / ByteDance · ICLR 2026 · Dec 2025 | 210 | Cascading-failure score; rubric judge | — | 88 experts (authoring) | DE 42.88 / DA 56.14 (GPT-5) | same |
| DARE-bench | Snowflake · ICLR 2026 · Feb 2026 | 6,300 | Verifiable ground truth | 10 min · 5 turns | none | Claude Sonnet 3.7 best | same |
| CoDA-Bench | Renmin · ICML 2026 · Jun 2026 | 1,009 (Hard 119) | Discovery + execution accuracy | — | none | ≈61.1% EA | same |
| WebDS | Stanford et al. · ICLR 2026 · Aug 2025 | 870 | Binary + trajectory score | — | ≈90% | 22.2% (Browser Use + GPT-5.1) | same |
| LongDS-Bench | ZJU / Ant · EMNLP 2026 · May 2026 | 68 / 2,225 turns | Turn-level judge (κ 0.862) | ≤40 steps/turn | none | 48.45% (Gemini-3.1-Pro) | same |
| LongDA · DSAEval · DDR-Bench | arXiv 2026 / ICML 2026 | 505 / 641 / 291 | Match rate / LLM judge / checklist | — | — | 69.2% / 8.16 / 47.73% | same |
| DataSciBench · DSCodeBench · DS-1000 | ACL 2026 F / AAAI 2026 / ICML 2023 | 222 / 1,000 / 1,000 | TFC framework / hidden tests / tests | — | none | 64.51 / 0.392 / 43.3 | — / — / 61.7 |
| InsightBench → InsightEval | ServiceNow ICLR 2025 → HKUST ACL 2026 F | 100 / 100 | G-Eval / Insight F1 | — | none | 0.60 / ≈0.587 | — |
| RADAR · DCA-Bench · MedAgentGym | NeurIPS 2025 / KDD 2025 / ICLR 2026 Oral | 2,980 / 221 / 72,413 | Exact match / judge / verifiable | — | — | 100 → 41% under artifacts / ≈30% / +45% RL | — |
| Kernels, systems, program search — Part 06 |
| KernelBench | Stanford · ICML 2025 · Feb 2025 | 250 (+20) | Correct on 5 inputs + fast_p | L40S; 100 timed trials | PyTorch eager | <20% fast_1 | daVinci-kernel 37/71/32 fast_1 |
| TritonBench | THUNLP · arXiv · Feb 2025 | 184 + 166 | Execution accuracy, speedup | A100 | — | 53% exec (R1, T) | KernelBenchX builds on it |
| FlashInfer-Bench | CMU · arXiv · Jan 2026 | 660 workloads | Resolved %, speedup vs library | B200 | Hand-tuned library | 0.628× (gemini-2.5-pro) | no model beats the library |
| SOL-ExecBench | NVIDIA · arXiv · Mar 2026 | 235 / 124 models | SOL score vs hardware bound | B200, clock-locked | Analytic bound | median 0.732 | 14.5% flagged for gaming |
| ComputeEval | NVIDIA · Apr 2025 → 2026-1 | 127 → 566 | Held-out tests, pass@k | — | — | 0.61 pass@1 (o3-mini) | — |
| MultiKernelBench · KernelGenBench | arXiv Jul 2025 / FlagOS Jul 2026 | 285 / 210 | pass@1 across accelerators | L20, Ascend, TPU / six platforms | — | 52.6% CUDA, <2.5% AscendC | 87% → 25% off-NVIDIA |
| LLM Speedrunning | Meta · NeurIPS 2025 · Jun 2025 | 19 | Fraction of speedup recovered | 60 min · 8×H100 | Record holders | 0.46 FSR (o3-mini, all hints) | no newer FSR |
| AlgoPerf | MLCommons · 2023 / ICLR 2025 | 8 workloads | Time-to-target profiles | 8×V100 | NAdamW baseline | Shampoo +28% | — |
| ALE-Bench | Sakana · NeurIPS 2025 · Jun 2025 | 40 (Lite 10) | Elo-like performance | 4 h / 1–2 weeks | Human avg ≈1260 | 1879 (ALE-Agent) | 1976 (FM Agent) |
| AlphaEvolve problem set | DeepMind · Jun 2025 | >50 problems | Verified constructions | — | Literature best | ≈20% improved | CodeEvolve 5/9; EurekAgent n=26 2.635999 |
| AI-for-science research — Part 07 |
| ScienceAgentBench | OSU · ICLR 2025 · Oct 2024 | 102 / 44 papers | Task-specific success + VER | — | none (2.5–3 h/task est.) | 42.2% (o1-preview) | no newer; verified split Apr 2026 |
| DiscoveryBench | AI2 · ICLR 2025 · Jul 2024 | 264 + 903 | Hypothesis matching score | — | none | 24.5 (Reflexion + oracle) | 33.7 (AstaBench) |
| DiscoveryWorld | AI2 · NeurIPS 2024 Spotlight · Jun 2024 | 120 | Completion / procedure / knowledge | 100–1,000 steps | 11 scientists: 66% | 38 / 18 / 18% completion | ≈20% (2025) |
| AstaBench | AI2 · ICLR 2026 Oral · Oct 2025 | 11 benchmarks / 2,400+ | Mixed programmatic + rubric | Cost-reported | none | 53.0 (Asta v0) | 58.0 (Opus 4.7) |
| FIRE-Bench | UCSD et al. · ICML 2026 · Feb 2026 | 30 + 10 + 60 | Claim-level entailment F1 | <24 h · 1×A100 | Paper findings | 46.7 F1 | same |
| InnoGym | ZJU · ICLR 2026 · Dec 2025 | 18 | Gain + novelty judge | 12 h | Leaderboard best | no agent beats humans | same |
| HeurekaBench | EPFL · ICLR 2026 · Jan 2026 | 50 + 50 | G-Eval (ρ 0.93) | — | Expert-reproduced insights | 2.58/5 | same |
| AAAR-1.0 · AbGen · ACADREASON | ICML 2025 / ACL 2025 / arXiv | 4 tasks / 1,500 / 50 | F1 / human / checklist | — | experts | 47.98 / 4.2 vs 4.8 / 16 pass | — |
| Agent-as-a-Judge (DevAI) | Meta / KAUST · ICML 2025 · Oct 2024 | 55 / 365 reqs | Agent judge vs human majority | — | 3 experts, 86.5 h | 90.16% alignment | — |