02 · Principles
Sutton, The Bitter Lesson (2019, essay) · Anthony, Tian & Barber, Thinking Fast and Slow with Deep Learning and Tree Search, 1705.08439 · Silver et al., AlphaZero, 1712.01815 · Jones, Scaling Scaling Laws with Board Games, 2104.03113 · Gandhi et al., Stream of Search, 2404.03683 · Yue et al., Does RLVR incentivize reasoning capacity, 2504.13837 · Liu et al., ProRL, 2505.24864 · Wu et al., The Invisible Leash, 2507.14843 · Chen et al., Pass@k Training, 2508.10751 · Brown et al., Large Language Monkeys, 2407.21787 · Snell et al., Scaling LLM Test-Time Compute Optimally, 2408.03314 · Wu et al., Inference Scaling Laws, 2408.00724 · Stroebl, Kapoor & Narayanan, The Limits of Inference Scaling Through Resampling, 2411.17501 · Gao, Schulman & Hilton, Scaling Laws for Reward Model Overoptimization, 2210.10760 · Karwowski et al., Goodhart's Law in Reinforcement Learning, 2310.09144 · Beirami et al., Theoretical guarantees on best-of-n, 2401.01879 · Wen et al., Rethinking Reward Model Evaluation, 2410.05584 · Jinnai et al., Regularized Best-of-N with MBR, 2404.01054 · Huang et al., LLMs Cannot Self-Correct Reasoning Yet, 2310.01798 · Shinn et al., Reflexion, 2303.11366 · Madaan et al., Self-Refine, 2303.17651 · Auer, Cesa-Bianchi & Fischer, Finite-time Analysis of the Multiarmed Bandit Problem (2002) · Kocsis & Szepesvári, Bandit based Monte-Carlo Planning (2006) · Yao et al., Tree of Thoughts, 2305.10601 · Besta et al., Graph of Thoughts, 2308.09687 · Hao et al., RAP, 2305.14992 · Zhou et al., LATS, 2310.04406 · Inoue et al., AB-MCTS, 2503.04412 · Parashar et al., Sys2Bench, 2502.12521 · Zhao, Awasthi & Gollapudi, Sample, Scrutinize and Scale, 2502.01839 · Wang et al., Don't Get Lost in the Trees, 2502.11183 · Koh et al., Tree Search for Language Model Agents, 2407.01476 · Srinivas et al., GP-UCB, 0912.3995 · Jamieson & Talwalkar, Non-stochastic Best Arm Identification, 1502.07943 · Li et al., Hyperband, 1603.06560 · Falkner, Klein & Hutter, BOHB, 1807.01774 · Swersky, Snoek & Adams, Freeze-Thaw, 1406.3896 · Liu et al., LLAMBO, 2402.03921 · Lehman & Stanley, Abandoning Objectives (2011) · Mouret & Clune, MAP-Elites, 1504.04909 · Ecoffet et al., Go-Explore, 1901.10995 · Wang et al., POET, 1901.01753 · Zhang et al., OMNI, 2306.01711 · Faldor et al., OMNI-EPIC, 2405.15568 · Hughes et al., Open-Endedness is Essential for ASI, 2406.04268 · Schmidhuber, Driven by Compression Progress, 0812.4360 · Schmidhuber, Gödel Machines, cs/0309048 · Hutter, The Fastest and Shortest Algorithm, cs/0206022 · Ellis et al., DreamCoder, 2006.08381 · Lehman et al., Evolution through Large Models, 2206.08896 · Bradley et al., QDAIF, 2310.13032 · Antoniades et al., Heuresis, 2606.25198 · Agarwal et al., AutoDiscovery, 2507.00310 · Wolpert & Macready, No Free Lunch (1997) · Lindley, On a Measure of the Information Provided by an Experiment (1956) · Kirk et al., Effects of RLHF on Generalisation and Diversity, 2310.06452 · Mohammadi, Creativity Has Left the Chat, 2406.05587 · Zhang et al., Verbalized Sampling, 2510.01171 · Cui et al., The Entropy Mechanism of RL, 2505.22617 · Song, Kempe & Munos, Outcome-based Exploration, 2509.06941 · Li et al., DARLING, 2509.02534 · Li et al., The Choice of Divergence, 2509.07430 · Yuan et al., Diversity Collapse via Overtraining, 2606.15455 · Zhao et al., Echo Chamber, 2504.07912 · Sha, Tan & Simchi-Levi, Diversity Collapse in LLM Game Play, 2607.19523
03 · Search inside ML-engineering agents
Jiang et al., AIDE, 2502.13138 and the WecoAI/aideml repository · Toledo et al., AI Research Agents for Machine Learning (AIRA-dojo), 2507.02554 · AIRA², Overcoming Bottlenecks in AI Research Agents, 2603.26499 · Liu et al., ML-Master, 2506.16499 · Zhu et al., ML-Master 2.0 / Ultra-Long-Horizon Agentic Science, 2601.10402 · Chi et al., SELA, 2410.17238 · Liang et al., I-MCTS, 2502.14693 · Du et al., AutoMLGen, 2510.08511 · Du et al., MLEvolve, 2606.06473 · Chen et al., MARS, 2602.02660 · Nam et al., MLE-STAR, 2506.15692 · R&D-Agent, 2505.14738 and the microsoft/RD-Agent repository · Zhang et al., Gome, 2603.01692 · Qiang et al., Matryoshka Agent, 2607.25090 · Kim, Talebirad & Zaïane, HASTE, 2606.30911 · Ou et al., AutoMind, 2506.10974 · Kulibaba et al., KompeteAI, 2508.10177 · Li et al., CoMind, 2506.20640 · Sahney et al., Operand Quant, 2510.11694 · FM Agent, 2510.26144 · Guo et al., DS-Agent, 2402.17453 · AutoKaggle, 2410.20424 · Agent K, 2411.03562 · MLZero, 2505.13941 · DS-STAR, 2509.21825 · Zheng et al., FORE-AGENT, 2601.05930 · Audran-Reiss et al., Studying the Role of Ideation Diversity, 2511.15593 · Chan et al., MLE-bench, 2410.07095 and the openai/mle-bench leaderboard
04 · Evolutionary program search
Real et al., AutoML-Zero, 2003.03384 · Romera-Paredes et al., FunSearch, Nature 625:468–475 with supplementary information · Novikov et al., AlphaEvolve, 2506.13131 and the DeepMind blog of 14 May 2025 · Georgiev, Gómez-Serrano, Tao & Wagner, Mathematical exploration and discovery at scale, 2511.02864 · Lange, Imajuku & Cetin, ShinkaEvolve, 2509.19349 · Assumpção et al., CodeEvolve, 2510.14150 · the OpenEvolve repository · Wang et al., ThetaEvolve, 2511.23473 · Yuksekgonul et al., TTT-Discover, 2601.16175 · Surina et al., EvoTune, 2504.05108 · Liu et al., Evolution of Heuristics, 2401.02051 · Liu et al., EoH-S, 2508.03082 · Ye et al., ReEvo, 2402.01145 · Shojaee et al., LLM-SR, 2404.18400 · Ma et al., Eureka, 2310.12931 · Chen, Dohan & So, EvoPrompting, 2302.14838 · Huang et al., CALM, 2505.12285 · Chen et al., A²DEPT, 2604.24043 · Pang et al., Deliberate Evolution, 2606.04360 · Real et al., AutoNumerics-Zero, 2312.08472 · van Stein & Bäck, LLaMEA, 2405.20132 · Su et al., Helix, 2603.07642 · Imajuku et al., ALE-Bench, 2506.09050 · Xin et al., EurekAgent, 2606.13662 · Agrawal et al., optimize_anything, 2605.19633 · Chung, Du & Wesley, Station, 2608.23691 · Jeddi et al., GEAR, 2605.13874 · Tan, Chin & Zhang, AgentGA, 2604.14655 · Jiang, Ding & Zhu, DeltaEvolve, 2602.02919 · Liu et al., ASI-Arch, 2507.18074 · Pelleriti et al., What Do Evolutionary Coding Agents Evolve?, 2605.20086 · Ishibashi et al., Effective Harness Engineering for Algorithm Discovery, 2605.15221 · Bahlous-Boldi et al., VPO, 2605.22817 · Nagda, Raghavan & Thakurta, Reinforced Generation of Combinatorial Structures, 2603.09172 · Dupont et al., Improving the matrix multiplication exponent, 2608.16884 · Chen et al., Magellan, 2601.21096 · Ananda et al., AI-PROPELLER, 2606.00131 · Bäuerle et al., Intentmaking and Sensemaking, 2605.05921
05 · Generating research ideas
Si, Yang & Hashimoto, Can LLMs Generate Novel Research Ideas?, 2409.04109 · Si, Hashimoto & Yang, The Ideation–Execution Gap, 2506.20803 · Gupta & Pruthi, All That Glitters is Not Novel, 2502.16487 · Zhang et al., NoveltyBench, 2504.05228 · Chen et al., Diversity Collapse in Multi-Agent LLM Systems, 2604.18005 · Wang et al., SciMON, 2305.14259 · Baek et al., ResearchAgent, 2404.07738 · Li et al., Chain of Ideas, 2410.13185 · Nova, 2410.14255 · Radensky et al., Scideator, 2409.14634 · Pu et al., IdeaSynth, 2410.04025 · Li et al., Learning to Generate Research Idea with Dynamic Control, 2412.14626 · Zhou et al., HypoGeniC, 2404.04326 · Liu et al., Literature Meets Data, 2410.17309 · Yang et al., MOOSE-Chem, 2410.07076 · MOOSE-Chem2, 2505.19209 · ResearchBench, 2503.21248 · Su et al., VirSci, 2410.09403 · LiveIdeaBench, 2412.17596 · AI Idea Bench 2025, 2504.14191 · IdeaBench, 2411.02429 · GraphEval, 2503.12600 · Wen et al., Predicting Empirical AI Research Outcomes, 2506.00794 · Mule, Garikaparthi & Patwardhan, Teaching LMs to Forecast Research Success, 2605.21491 · Lou et al., AAAR-1.0, 2410.22394 · Xu et al., LimitGen, 2507.02694 · Lu et al., The AI Scientist, 2408.06292 · Yamada et al., The AI Scientist-v2, 2504.08066 and Nature s41586-026-10265-5 · Schmidgall et al., Agent Laboratory, 2501.04227 · Gottweis et al., AI co-scientist, 2502.18864 and Nature s41586-026-10644-y · Jansen et al., CodeScientist, 2503.22708 · Weng et al., DeepScientist, 2509.26603 · InternAgent / NovelSeek, 2505.16938 · InternAgent-1.5, 2602.08990 · Dolphin, 2501.03916 · CycleResearcher, 2411.00816 · AI-Researcher, 2505.18705 · Kosmos, 2511.02824 · Robin, 2505.13400 · Jiang et al., BadScientist, 2510.18003 · Beel, Kan & Baumgart, Evaluating Sakana's AI Scientist, 2502.14297 · Trehan & Chopra, Why LLMs Aren't Scientists Yet, 2601.03315 · Bianchi et al., Agents4Science post-mortem, 2511.15534 · the ICLR 2026 Author Guide and the NeurIPS 2026 Main Track Handbook
06 · The empirical record
Zhang et al., Learning to Ideate for MLE Agents, 2601.17596 · Kim et al., Investigating Component Contributions in Multi-Agent ML Systems (K-LIVE), ICML 2026 · Zhao et al., Demystify the Role of Memory in MLE Agents, Findings of ACL 2026 · Meta, MLGym, 2502.14499 · Chen et al., MLR-Bench, 2505.19955 · Zou et al., FML-bench, 2510.10472 · Garikaparthi, Patwardhan & Cohan, ResearchGym, 2602.15112 · InnovatorBench, 2510.27598 · InnoGym, 2512.01822 · EXP-Bench, 2505.24785 · AARRI-Bench, 2606.07462 · Zhu et al., AI Scientists Fail Without Strong Implementation Capability, 2506.01372 · Lin et al., Can Language Models Discover Scaling Laws? (SLDAgent), 2507.21184 · Towards Execution-Grounded Automated AI Research, 2601.14525 · Yang, He-Yueya & Liang, duration-aware asynchronous RL, 2509.01684 · ML-Agent, 2505.23723 · AceGRPO, 2602.07906 · AgentNAS, 2607.07984 · MLE-Dojo, 2505.07782 · SERA, 2601.20789 · SWE-Effi, 2509.09853 · Wijk et al., RE-Bench, 2411.15114
07 · Searching over the searcher
Yang et al., OPRO, 2309.03409 · Fernando et al., Promptbreeder, 2309.16797 · Opsahl-Ong et al., MIPROv2, 2406.11695 · Yuksekgonul et al., TextGrad, 2406.07496 · Cheng, Nie & Swaminathan, Trace, 2406.16218 · Agrawal et al., GEPA, 2507.19457 · Zhang et al., ACE, 2510.04618 · Suzgun et al., Dynamic Cheatsheet, 2504.07952 · Hu, Lu & Clune, ADAS, 2408.08435 · Zhang et al., AFlow, 2410.10762 · Shang et al., AgentSquare, 2410.06153 · Li et al., AgentSwift, 2506.06017 · Zhang et al., MaAS, 2502.04180 · Ye et al., MAS-GPT, 2503.03686 · Zhuge et al., GPTSwarm, 2402.16823 · Saad-Falcon et al., Archon, 2409.15254 · Hu et al., EvoMAS, 2602.06511 · Zhang et al., Darwin Gödel Machine, 2505.22954 · Wang et al., Huxley-Gödel Machine, 2510.21614 · SICA, 2504.15228 · Yin et al., Gödel Agent, 2410.04444 · Lee et al., Meta-Harness, 2603.28052 · AHE, 2604.25850 · HarnessCompass, 2608.01918 · Self-Harness, 2606.09498 · RHO, 2606.05922 · DemoEvolve, 2605.24539 · AutoSaddler, 2608.23041 · Lin et al., Harness Updating Is Not Harness Benefit, 2605.30621 · Wang et al., Rethinking the Evaluation of Harness Evolution, 2607.12227 · DarwinX, 2608.07545 · Red Queen Gödel Machine, 2606.26294 · EvoTest, 2510.13220 · EVOTOOL, 2603.04900 · CoMAS, 2510.08529 · Shao et al., Your Agent May Misevolve, 2509.26354 · Chen, Wang & Qu, Recursive Self-Improvement survey, 2607.07663 · Zhuge et al., Agent-as-a-Judge, 2410.10934 · Tan et al., JudgeBench, 2410.12784 · Wang et al., LLMs are not Fair Evaluators, 2305.17926 · Panickssery, Bowman & Feng, self-recognition and self-preference, 2404.13076 · Kim et al., On the limits and opportunities of AI reviewers, 2605.20668 · Li et al., LLM-as-a-Reviewer, 2605.25415 · Bjarnason, Silva & Monperrus, On Randomness in Agentic Evaluations, 2602.07150 · AgentLens, 2605.12925 · When Agents Disagree With Themselves, 2602.11619 · Are "Solved Issues" Really Solved Correctly?, 2503.15223 · Chasing the Public Score, 2604.20200 · Task-CoEvolve, 2608.20169 · Claw-SWE-Bench, 2606.12344 · The Scaffold Effect, 2607.22585
08 · The record
Fawzi et al., AlphaTensor (Nature 2022) · Mankowitz et al., AlphaDev (Nature 2023) · Szymanski et al., A-Lab (Nature 2023, corrected January 2026) · the DeepMind IMO 2024 and IMO 2025 posts · the ICPC 2025 posts from DeepMind and OpenAI · Bubeck et al., Early science acceleration experiments with GPT-5, 2511.16072 · the erdosproblems.com forum and its maintainer's response of October 2025 · the AtCoder World Tour Finals 2025 and 2026 results · the Sakana AI CUDA Engineer post and its walk-back, and Towards Robust Agentic CUDA Kernel Benchmarking, 2509.14279 · the NVIDIA DeepSeek-R1 kernel post (February 2025) · the karpathy/autoresearch repository and the SkyPilot scale-up write-up · the modded-nanoGPT speedrun ledger · Weco AI's blog · the MLE-bench leaderboard · Markov, The False Dawn, 2306.09633 · Cheng et al., An Updated Assessment of RL for Macro Placement, 2302.11014 · ERA, An AI system to help scientists write expert-level empirical software, 2509.06503 and Nature s41586-026-10658-6 · Anthropic's Vibe Physics (March 2026) and protein-design (August 2026) posts · Microsoft Discovery (May 2025) · Google's Gemini for Science (May 2026) and AlphaEvolve availability posts (July 2026)
09 · The pre-LLM lineage
Bergstra & Bengio, Random Search for Hyper-Parameter Optimization, JMLR 13 (2012) · Snoek, Larochelle & Adams, Practical Bayesian Optimization, 1206.2944 · Hutter, Hoos & Leyton-Brown, SMAC (LION 2011) · Bergstra et al., TPE (NeurIPS 2011) · Li et al., ASHA, 1810.05934 · Jaderberg et al., Population Based Training, 1711.09846 · Probst, Boulesteix & Bischl, Tunability, 1802.09596 · Eggensperger et al., HPOBench, 2109.06716 · Mallik et al., PriorBand, 2306.12370 · Zhang et al., Using LLMs for HPO, 2312.04528 · Kristiadi et al., A Sober Look at LLMs for Material Discovery, 2402.05015 · Zoph & Le, Neural Architecture Search with RL, 1611.01578 · Zoph et al., NASNet, 1707.07012 · Pham et al., ENAS, 1802.03268 · Liu, Simonyan & Yang, DARTS, 1806.09055 · Real et al., Regularized Evolution, 1802.01548 · Li & Talwalkar, Random Search and Reproducibility for NAS, 1902.07638 · Yu et al., Evaluating the Search Phase of NAS, 1902.08142 · Yang, Esperança & Carlucci, NAS evaluation is frustratingly hard, 1912.12522 · Wan et al., On Redundancy and Diversity in Cell-based NAS, 2203.08887 · Ying et al., NAS-Bench-101, 1902.09635 · Dong & Yang, NAS-Bench-201, 2001.00326 · White, Nolen & Savani, Local Search is State of the Art for NAS Benchmarks, 2005.02960 · Mellor et al., NASWOT, 2006.04647 · Lee & Ham, AZ-NAS, 2403.19232 · Zheng et al., GENIUS, 2304.10970 · Zhou et al., Design Principle Transfer in NAS via LLMs, 2408.11330 · Tan & Le, EfficientNet, 1905.11946 · Thornton et al., Auto-WEKA, 1208.3719 · Feurer et al., auto-sklearn (NeurIPS 2015) and auto-sklearn 2.0, 2007.04074 · Olson et al., TPOT, 1603.06212 · Erickson et al., AutoGluon-Tabular, 2003.06505 · Gijsbers et al., AMLB, 2207.12560 · Guyon et al., ChaLearn AutoML challenge analysis (2019) · Drori et al., AlphaD3M, 2111.02508 · Li et al., DivBO, 2302.03255 · Hollmann, Müller & Hutter, CAAFE, 2305.03403 · Tornede et al., AutoML in the Age of LLMs, 2306.08107 · Andrychowicz et al., Learning to learn by gradient descent by gradient descent, 1606.04474 · Metz et al., VeLO, 2211.09760 · Dahl et al., AlgoPerf, 2306.07179 and Kasimbeg et al., AlgoPerf results, 2502.15015 · Li et al., AutoLoss-Zero, 2103.14026 · So, Liang & Le, The Evolved Transformer, 1901.11117 · Chen et al., Symbolic Discovery of Optimization Algorithms (Lion), 2302.06675 · La Cava et al., SRBench, 2107.14351