A Auto-Research / AI-Scientist agents 255
LLM agents and studies that automate parts of the research process itself: ideation, hypothesis generation, literature work, experiment execution, paper writing and peer review.
A1 End-to-end research agents & discovery systems 19
- AgentExpt: Automating AI Experiment Design with LLM-based Resource Retrieval Agent
- MARS: Modular Agent with Reflective Search for Automated AI Research
- Principle-Evolvable Scientific Discovery via Uncertainty Minimization
- Towards Execution-Grounded Automated AI Research
- Training AI Co-Scientists Using Rubric Rewards
- Can Language Models Discover Scaling Laws?
- DeepScientist: Advancing Frontier-Pushing Scientific Findings Progressively
- Towards Multimodal Data-Driven Scientific Discovery Powered by LLM Agents
- TusoAI: Agentic Optimization for Scientific Methods
- EvoSci: A Bio-Inspired Multi-Agent Framework for the Evolution of Scientific Discovery
- Large Language Model (LLM) as an Excellent Reinforcement Learning Researcher in both Single-Agent and Multi-Agent Scenarios
- AI-Researcher: Autonomous Scientific Innovation
- Language Modeling by Language Models
- AutoDiscovery: Open-ended Scientific Discovery via Bayesian Surprise
- AutoSciDACT: Automated Scientific Discovery through Contrastive Embedding and Hypothesis Testing
- ResearchTown: Simulator of Human Research Community
- CycleResearcher: Improving Automated Research via Automated Review
- Dolphin: Moving Towards Closed-loop Auto-research through Thinking, Practice, and Feedback
- CodeScientist: End-to-End Semi-Automated Scientific Discovery with Code-based Experimentation
A2 Research idea & hypothesis generation, evaluation, validation 41
- A Probabilistic Framework for LLM-Based Model Discovery
- HypoSpace: A Diagnostic Benchmark for Set-Valued Hypothesis Generation under Underdetermination and Sublinear Coverage Bounds
- InnoEval: On Research Idea Evaluation as a Knowledge-Grounded, Multi-Perspective Reasoning Problem
- MindFlow: Mind Supernet Powered Thinking Flows for Research Idea Innovation
- MOOSE-Star: Unlocking Tractable Training for Scientific Discovery by Breaking the Complexity Barrier
- Towards Diverse Scientific Hypothesis Search with Large Language Models
- InnoGym: Benchmarking the Innovation Potential of AI Agents
- MetaMuse: Algorithm Generation via Creative Ideation
- The Ideation-Execution Gap: Execution Outcomes of LLM-Generated versus Human Research Ideas
- CHIMERA: A Knowledge Base of Scientific Idea Recombinations for Research Analysis and Ideation
- EvoNarrator: Modeling Scientific Evolution for Feasible Hypothesis Generation
- Experiments or Outcomes? Probing Scientific Feasibility in Large Language Models
- Generating Literature-Driven Scientific Theories at Scale
- GUIDE: Towards Scalable Advising for Research Ideas
- MoRI: Learning Motivation-Grounded Reasoning for Scientific Ideation in Large Language Models
- Diversity Collapse in Multi-Agent LLM Systems: Structural Coupling and Collective Failure in Open-Ended Idea Generation
- NovBench: Evaluating Large Language Models on Academic Paper Novelty Assessment
- RoadMapper: A Multi-Agent System for Roadmap Generation of Solving Complex Research Problems
- Teaching Language Models to Forecast Research Success Through Comparative Idea Evaluation
- GraphMind: Unveiling Scientific Reasoning through Contextual Graphs for Novelty Assessment
- Curiosity-Driven Questioning for Engine-Agnostic LLM Research Ideation
- InfRL: Inference-time Reinforcement Learning for Research Idea Optimization
- Exploiting LLMs for Automatic Hypothesis Assessment via a Logit-Based Calibrated Prior
- MOOSE-Chem2: Exploring LLM Limits in Fine-Grained Scientific Hypothesis Discovery via Hierarchical Search
- Automated Hypothesis Validation with Agentic Sequential Falsifications
- Sparse Autoencoders for Hypothesis Generation
- Can LLMs Generate Novel Research Ideas? A Large-Scale Human Study with 100+ NLP Researchers
- GraphEval: A Lightweight Graph-Based LLM Framework for Idea Evaluation
- MOOSE-Chem: Large Language Models for Rediscovering Unseen Chemistry Scientific Hypotheses
- All That Glitters is Not Novel: Plagiarism in AI Generated Research
- Literature Meets Data: A Synergistic Approach to Hypothesis Generation
- Many Heads Are Better Than One: Improved Scientific Idea Generation by A LLM-Based Multi-Agent System
- MIR: Methodology Inspiration Retrieval for Scientific Research Problems
- FRAME: Feedback-Refined Agent Methodology for Enhancing Medical Research Insights
- NOVA: An Iterative Planning Framework for Enhancing Scientific Innovation with Large Language Models
- IdeaBench: Benchmarking Large Language Models for Research Idea Generation
- Position: Data-driven Discovery with Large Generative Models
- SciMON: Scientific Inspiration Machines Optimized for Novelty
- Large Language Models for Automated Open-domain Scientific Hypotheses Discovery
- Exploring and Verbalizing Academic Ideas by Concept Co-occurrence
- Less Likely Brainstorming: Using Language Models to Generate Alternative Hypotheses
A3 Literature, papers & scientific writing 89
- LECTOR: Joint Learning of Scientific Reasoning Graphs and Introduction Generation
- LitReview Arena: Evaluating Literature Review Agents with Battle-style Peer Review Platform
- LiveFigure: Generating Editable Scientific Illustration with VLM Agents
- PaperBanana: Automating Academic Illustration for AI Scientists
- PosterAgent: Agentic Poster Generation via Stage-Aware Reinforcement Learning
- SciNet: Evaluating AI Agents in Relation-Aware Scientific Literature Retrieval
- AirQA: A Comprehensive QA Dataset for AI Research with Instance-Level Evaluation
- AutoFigure: Generating and Refining Publication-Ready Scientific Illustrations
- Not Search, But Scan: Benchmarking MLLMs on Scan-Oriented Academic Paper Reasoning
- Presenting a Paper is an Art: Self-Improvement Aesthetic Agents for Academic Presentations
- Paper2Figure: A Multi-Agent Collaborative System for Figure Generation Towards Academic Research Paper
- Docora: A System for Interactive Knowledge Extraction and Visualization from Scientific PDFs
- SlideTailor: Personalized Presentation Slide Generation for Scientific Papers
- arXiv2Table: Toward Realistic Benchmarking and Evaluation for LLM-Based Literature-Review Table Generation
- Automatic and Reliable Evaluation for Academic Caption-to-Figure Generation with LMMs
- Bloom-Eval: A Hierarchical Evaluation Benchmark for Automatic Survey Generation Based on Bloom’s Taxonomy
- IntrAgent: An LLM Agent for Content-Grounded Information Retrieval through Literature Review
- Paper Circle: An Open-source Multi-agent Research Discovery and Analysis Framework
- PaperRegister: Boosting Flexible-grained Paper Search via Hierarchical Register Indexing
- PosterForest: Hierarchical Multi-Agent Collaboration for Scientific Poster Generation
- Reward Modeling for Scientific Writing Evaluation
- RPC-Bench: A Fine-grained Benchmark for Research Paper Comprehension
- ScholaWrite: A Dataset of End-to-End Scholarly Writing
- SciFlow-Bench: Evaluating Structure-Aware Scientific Diagram Generation via Inverse Parsing
- SciMDR: Advancing Scientific Multimodal Document Reasoning
- Text2Tabular – Reconstructing Tabular Research Data from Scientific Publications
- XtraGPT: Context-Aware and Controllable Academic Paper Revision via Human-AI Collaboration
- Automatic Paper Analysis and Categorisation for Systematic Reviews with Combined Reasoning-Augmented SFT and DAPO RL
- Can AI Revise Research Papers with Human Review Feedback? An Empirical Study and Benchmark
- Datasets for Scientific Literature Understanding: A Survey
- DeepSynth-Eval: Objectively Evaluating Information Consolidation in Deep Survey Writing
- Feedback Is The Key for Automated Survey Generation
- GRASP: Graph-Reasoning Aided Survey Planning for High-Fidelity Related Work Generation
- Human-Agent Collaborative Paper-to-Page Crafting
- PAPERMIND: Benchmarking Agentic Reasoning and Critique over Scientific Papers in Multimodal LLMs
- PaperScope: A Multi-Modal Multi-Document Benchmark for Agentic Deep Research Across Massive Scientific Papers
- LitBench: A Graph-Centric Large Language Model Benchmarking Tool For Literature Tasks
- AISE-Bench: A Full-Cycle Curated Benchmark for Information Seeking on Academic Knowledge Graphs
- Benchmarking Table Extraction from Heterogeneous Scientific PDF Documents
- CoRank: LLM-Based Compact Reranking with Document Features for Scientific Retrieval
- DECKBench: Benchmarking Multi-Agent Frameworks for Academic Slide Generation and Editing
- INDUS-SDE: A Language Model for Scientific Content Curation and Discovery
- SurveyReview: A Reviewer-Aligned Benchmark for Survey Evaluators
- VILLA: Versatile Information Retrieval from Scientific Literature Using Large Language Models
- SciArena: An Open Evaluation Platform for Non-Verifiable Scientific Literature-Grounded Tasks
- Paper2Poster: Towards Multimodal Poster Automation from Scientific Papers
- SciLitLLM: How to Adapt LLMs for Scientific Literature Understanding
- From Words to Structured Visuals: A Benchmark and Framework for Text-to-Diagram Generation and Editing
- TikZero: Zero-Shot Text-Guided Graphics Program Synthesis
- MMCR: Benchmarking Cross-Source Reasoning in Scientific Papers
- Preacher: Paper-to-Video Agentic System
- Mixture of Knowledge Minigraph Agents for Literature Review Generation
- Can LLMs Help Uncover Insights about LLMs? A Large-Scale, Evolving Literature Analysis of Frontier LLMs
- Completing A Systematic Review in Hours instead of Months with Interactive AI Agents
- PaSa: An LLM Agent for Comprehensive Academic Paper Search
- SciVer: Evaluating Foundation Models for Multimodal Scientific Claim Verification
- SURVEYFORGE : On the Outline Heuristics, Memory-Driven Generation, and Multi-dimensional Evaluation for Automated Survey Writing
- Can Multimodal Foundation Models Understand Schematic Diagrams? An Empirical Study on Information-Seeking QA over Scientific Papers
- LGAR: Zero-Shot LLM-Guided Neural Ranking for Abstract Screening in Systematic Literature Reviews
- Select, Read, and Write: A Multi-Agent Framework of Full-Text-based Related Work Generation
- ChatPD: An LLM-driven Paper-Dataset Networking System
- ScIRGen: Synthesize Realistic and Large-Scale RAG Dataset for Scientific Research
- AutoSurvey: Large Language Models Can Automatically Write Surveys
- CiteME: Can Language Models Accurately Cite Scientific Claims?
- cPAPERS: A Dataset of Situated and Multimodal Interactive Conversations in Scientific Papers
- HLM-Cite: Hybrid Language Model Workflow for Text-based Scientific Citation Prediction
- SciFIBench: Benchmarking Large Multimodal Models for Scientific Figure Interpretation
- SPIQA: A Dataset for Multimodal Question Answering on Scientific Papers
- Nougat: Neural Optical Understanding for Academic Documents
- SciSpace Copilot: Empowering Researchers through Intelligent Reading Assistance
- Multimodal ArXiv: A Dataset for Improving Scientific Comprehension of Large Vision-Language Models
- Systematic Task Exploration with LLMs: A Study in Citation Text Generation
- Automated Focused Feedback Generation for Scientific Writing Assistance
- CHIME: LLM-Assisted Hierarchical Organization of Scientific Studies for Literature Review Support
- SciMMIR: Benchmarking Scientific Multi-modal Information Retrieval
- SumSurvey: An Abstractive Dataset of Scientific Survey Papers for Long Document Summarization
- SymTax: Symbiotic Relationship and Taxonomy Fusion for Effective Citation Recommendation
- A Hierarchical Context Augmentation Method to Improve Retrieval-Augmented LLMs on Scientific Papers
- LitFM: A Retrieval Augmented Structure-aware Foundation Model For Citation Graphs
- SoAy: A Solution-based LLM API-using Methodology for Academic Information Seeking
- CSMeD: Bridging the Dataset Gap in Automated Citation Screening for Systematic Literature Reviews
- Scientific Document Retrieval using Multi-level Aspect-based Queries
- QASA: Advanced Question Answering on Scientific Articles
- DataFinder: Scientific Dataset Recommendation from Natural Language Descriptions
- SciReviewGen: A Large-scale Dataset for Automatic Literature Review Generation
- A Search Engine for Discovery of Scientific Challenges and Directions
- DisenCite: Graph-Based Disentangled Representation Learning for Context-Specific Citation Generation
- DOC2PPT: Automatic Presentation Slides Generation from Scientific Documents
- OAG-BERT: Towards a Unified Backbone Language Model for Academic Knowledge Services
A4 Peer review & meta-science 49
- Position: Stop Automating Peer Review Without Rigorous Evaluation
- CoCoReviewBench: A Completeness- and Correctness-Oriented Benchmark for AI Reviewers
- Does AI Reviewer See the Full Picture? Attacking and Defending Multimodal Peer Review
- Policies Permitting LLM Use for Polishing Peer Reviews Are Currently Not Enforceable
- Position: Peer Review Should Be Calibrated via LLM Scoring
- Position: Prompting Intent Should Be Audited in LLM-Assisted Peer Review
- Position: The AI Imperative: Scaling High-Quality Peer Review in Machine Learning
- Sem-Detect: Semantic Level Detection of AI Generated Peer-Reviews
- Dancing in Chains: Strategic Persuasion in Academic Rebuttal via Theory of Mind
- Is Your Paper Being Reviewed by an LLM? Benchmarking AI Text Detection in Peer Review
- Paper Copilot: Tracking the Evolution of Peer Review in AI Conferences
- PRISMM-Bench: A Benchmark of Peer-Review Grounded Multimodal Inconsistencies
- THEMIS: Towards Holistic Evaluation of MLLMs for Scientific Paper Fraud Forensics
- Automatic Paper Reviewing with Heterogeneous Graph Reasoning over LLM-Simulated Reviewer-Author Debates
- Format Matters: The Robustness of Multimodal LLMs in Reviewing Evidence from Tables and Charts
- Navigating Through Paper Flood: Advancing LLM-Based Paper Evaluation Through Domain-Aware Retrieval and Latent Reasoning
- Rescind: Countering Image Misconduct in Biomedical Publications with Vision-Language and State-Space Modeling
- AEGIS: A Holistic Benchmark for Evaluating Forensic Analysis of AI-Generated Academic Images
- Author-in-the-Loop Response Generation and Evaluation: Integrating Author Expertise and Intent in Responses to Peer Review
- BadScientist: Can a Research Agent Write Convincing but Unsound Papers that Fool LLM Reviewers?
- Can AI Be a Good Peer Reviewer? A Survey of Peer Review Process, Evaluation, and the Future
- CoCoNUTS: Concentrating on Content while Neglecting Uninformative Textual Styles for AI-Generated Peer Review Detection
- HalluCitation Matters: Revealing the Impact of Hallucinated References with 300 Hallucinated Papers in ACL Conferences
- Paper2Rebuttal: A Multi-Agent Framework for Transparent Author Response Assistance
- RATE: Reviewer Profiling and Annotation-free Training for Expertise Ranking in Peer Review Systems
- ReviewGrounder: Improving Review Substantiveness with Rubric-Guided, Tool-Integrated Agents
- DRPG (Decompose, Retrieve, Plan, Generate): An Agentic Framework for Academic Rebuttal
- Evaluating the Impact of Reviewer Guideline Design on LLM-Based Automated Peer Review
- From Isolated Scoring to Collaborative Ranking: A Comparison-Native Framework for LLM-Based Paper Evaluation
- Justice in Judgment: Unveiling (Hidden) Bias in LLM-assisted Peer Reviews
- PeerCheck: Enhancing LLM-Generated Academic Reviews Towards Human-Level Quality
- RbtAct: Rebuttal as Supervision for Actionable Review Feedback Generation
- Towards Reliable Paper Contributions Annotation in the ACL Rolling Review
- When Reviews Disagree: Fine-Grained Contradiction Analysis in Scientific Peer Reviews
- From Replication to Redesign: Exploring Pairwise Comparisons for LLM-Based Peer Review
- Agent Reviewers: Domain-specific Multimodal Agents with Shared Memory for Paper Review
- Benchmarking LLMs' Judgments with No Gold Standard
- Are Key-Phrases All That Reviewers Care About? A Comprehensive Benchmarking of Reviewer Matchmaking Systems
- Can LLMs Identify Critical Limitations within Scientific Research? A Systematic Evaluation on AI Research Papers
- DeepReview: Improving LLM-based Paper Review with Human-like Deep Thinking Process
- LazyReview: A Dataset for Uncovering Lazy Thinking in NLP Peer Reviews
- STRICTA: Structured Reasoning in Critical Text Assessment for Peer Review and Beyond
- Monitoring AI-Modified Content at Scale: A Case Study on the Impact of ChatGPT on AI Conference Peer Reviews
- A Sentiment Consolidation Framework for Meta-Review Generation
- ARIES: A Corpus of Scientific Paper Edits Made in Response to Peer Reviews
- GLIMPSE: Pragmatically Informative Multi-Document Summarization for Scholarly Reviews
- NLPeer: A Unified Resource for the Computational Study of Peer Review
- KID-Review: Knowledge-Guided Scientific Review Generation with Oracle Pre-training
- MReD: A Meta-Review Dataset for Structure-Controllable Text Generation
A5 Benchmarks, evaluation & position papers 30
- Position: Preregister Experiments with AI Agents
- CauSciBench: Can LLMs Automate Causal Inference in Real-World Scientific Research?
- FIRE-Bench: Evaluating AI Agents on the Rediscovery of Scientific Insights
- Position: Academic Conferences are Potentially Facing Denominator Gaming Caused by Fully Automated Scientific Agents
- Position: The Age of AI Agents Demands A New Scientific Paradigm To Sustain Trustworthy Science
- AstaBench: Rigorous Benchmarking of AI Agents with a Scientific Research Suite
- ACADREASON: Exploring the Limits of Reasoning Models with Academic Research Problems
- EXP-Bench: Can AI Conduct AI Research Experiments?
- HeurekaBench: A Benchmarking Framework for AI Co-scientist
- InnovatorBench: Evaluating Agents’ Ability to Conduct Innovative AI Research
- In-depth Research Impact Summarization through Fine-Grained Temporal Citation Analysis
- AI Agents for the Science of Science: A Survey of Tasks, Architectures, Evaluations, and Challenges
- ResearchBench: Benchmarking LLMs in Scientific Discovery via Inspiration-Based Task Decomposition
- SciExplore: Evaluating Autonomous Agents from Scientific Navigation to Information Integration
- SciImpact: A Multi-Dimensional, Multi-Field Benchmark for Scientific Impact Prediction
- The Autonomous AI Scientist of 2030 Needs Inferential Arbitrage
- Stop DDoS Attacking the Research Community with AI-Generated Survey Papers
- Foundation Models for Scientific Discovery: From Paradigm Enhancement to Paradigm Transition
- MLR-Bench: Evaluating AI Agents on Open-Ended Machine Learning Research
- Predicting Empirical AI Research Outcomes with Language Models
- AAAR-1.0: Assessing AI’s Potential to Assist Research
- DiscoveryBench: Towards Data-Driven Discovery with Large Language Models
- ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery
- From Words to Worth: Newborn Article Impact Prediction with LLM
- Towards Scientific Discovery with Generative AI: Progress, Opportunities, and Challenges
- AbGen: Evaluating Large Language Models in Ablation Study Design and Evaluation for Scientific Research
- DiscoveryWorld: A Virtual Environment for Developing and Evaluating Automated Scientific Discovery Agents
- Integrated Systems for Computational Scientific Discovery
- The Automatic Computer Scientist
- Interpretable Research Replication Prediction via Variational Contextual Consistency Sentence Masking
A6 Adjacent: web "deep research" agents (report-writing over the web, not scientific research) 27
- DEER: A Benchmark for Evaluating Deep Research Agents on Expert Report Generation
- Characterizing Deep Research: A Benchmark and Formal Definition
- DeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents
- DeepTRACE: Auditing Deep Research AI Systems for Tracking Reliability Across Citations and Evidence
- DRBench: A Realistic Benchmark for Enterprise Deep Research
- ResearchRubrics: A Benchmark of Prompts and Rubrics For Evaluating Deep Research Agents
- Towards Personalized Deep Research: Benchmarks and Evaluations
- WebWeaver: Structuring Web-Scale Evidence with Dynamic Outlines for Open-Ended Deep Research
- Deep Research Arena: The First Exam of LLMs’ Research Abilities via Seminar-Grounded Tasks
- Multimodal DeepResearcher: Generating Text-Chart Interleaved Reports from Scratch with Agentic Framework
- Beyond Single-shot Writing: Deep Research Agents are Unreliable at Multi-turn Report Revision
- Cognitive Scaffold: From Fluid Context to Crystallized Memory for Long-Horizon DeepResearch Agents
- Deep Research with Open-Domain Evaluation and Multi-Stage Guardrails for Safety
- Deep-Reporter: Deep Research for Grounded Multimodal Long-Form Generation
- DeepFact: Co-Evolving Benchmarks and Agents for Deep Research Factuality
- DR-Arena: an Automated Evaluation Framework for Deep Research Agents
- DREAM: Deep Research Evaluation with Agentic Metrics
- FlowSearch: Advancing Deep Research with Dynamic Structured Knowledge Flow
- FS-Researcher: Test-Time Scaling for Long-Horizon Research Tasks with File-System-Based Agents
- Language Models Don’t Know What You Want: Evaluating Personalization in Deep Research Needs Real Users
- ReportLogic: Evaluating Logical Quality in Deep Research Reports
- WebAggregator: Enhancing Compositional Reasoning Capabilities of Deep Research Agent Foundation Models
- CogGen: A Cognitively Inspired Recursive Framework for Deep Research Report Generation
- DeepPlanner: Scaling Planning Capability for Deep Research Agents via Advantage Shaping
- Inference-Time Scaling of Verification: Self-Evolving Deep Research Agents via Test-Time Rubric-Guided Verification
- MASS: Deep Research for Social Sciences with Memory-Augmented Social Simulation
- WebThinker: Empowering Large Reasoning Models with Deep Research Capability
B Auto-MLE / data-science agents 101
Agents that do machine-learning engineering, Kaggle-style modelling, data analysis and AI R&D, plus the benchmarks and environments that evaluate them.
B1 Agents & systems 38
- $R^3$DAO: Reactive Recovery and Reconstruction for Long-horizon Data Agent Orchestration
- DeepAnalyze: Agentic Large Language Models for Autonomous Data Science
- ML-Agent: Reinforcing LLM Agents for Autonomous Machine Learning Engineering
- CoDA: Agentic Systems for Collaborative Data Visualization
- CoMind: Towards Community-Driven Agents for Machine Learning Engineering
- Reinforcement Learning for Machine Learning Engineering Agents
- Scaling Generalist Data-Analytic Agents
- VisCoder2: Building Multi-Language Visualization Coding Agents
- Simple Agents Outperform Experts in Biomedical Imaging Workflow Optimization
- Jupiter: Enhancing LLM Data Analysis Capabilities via Notebook and Inference-Time Value-Guided Search
- Can We Predict Before Executing Machine Learning Agents?
- SPIO: Ensemble and Selective Strategies via LLM-Based Multi-Agent Planning in Automated Data Science
- DataSage: Multi-agent Collaboration for Insight Discovery with External Knowledge Retrieval, Multi-role Debating, and Multi-path Reasoning
- DataSeer: A Manager-Centric Collaborative Multi-Agent Framework with Multi-Branch Reasoning for Automated Insight Discovery
- Demystify the Role of Memory in Machine Learning Engineering Agents
- Reasoning as Gradient: Scaling MLE Agents Beyond Tree Search
- Tree-Notebook: A Context-Aware Agent with Tree Search and Entropy-Aware Data Shadow for Interactive Data Science
- EvoDS: Self-Evolving Autonomous Data Science Agent with Skill Learning and Context Management
- PACE: Unleashing the Power of Code Embeddings to Boost AutoML Agents
- ProfiliTable: Profiling-Driven Tabular Data Processing via Agentic Workflows
- Rewarding the Scientific Process: Process-Level Reward Modeling for Agentic Data Analysis
- AI Research Agents for Machine Learning: Search, Exploration, and Generalization in MLE-bench
- MLE-STAR: Machine Learning Engineering Agent via Search and Targeted Refinement
- MLZero: A Multi-Agent System for End-to-end Machine Learning Automation
- R&D-Agent-Quant: A Multi-Agent Framework for Data-Centric Factors and Model Joint Optimization
- Unlocking SLM Potential for Data Analysis Code Generation via Non-Parametric Knowledge Distillation
- Adaptive Self-improvement LLM Agentic System for ML Library Development
- AutoML-Agent: A Multi-Agent LLM Framework for Full-Pipeline AutoML
- DataEnvGym: Data Generation Agents in Teacher Environments with Student Feedback
- nvAgent: Automated Data Visualization from Natural Language via Collaborative Agent Workflow
- Data Interpreter: An LLM Agent for Data Science
- Automated Statistical Model Discovery with Language Models
- DS-Agent: Automated Data Science by Empowering Large Language Models with Case-Based Reasoning
- Modeling Collaborator: Enabling Subjective Vision Classification With Minimal Human Effort via LLM Tool-Use
- MatPlotAgent: Method and Evaluation for LLM-Based Agentic Scientific Data Visualization
- Efficient Mixture of Experts based on Large Language Models for Low-Resource Data Preprocessing
- Large Language Models for Automated Data Science: Introducing CAAFE for Context-Aware Automated Feature Engineering
- Reinforced Approximate Exploratory Data Analysis
B2 Benchmarks & environments 47
- CoDA-Bench: Can Code Agents Handle Data-Intensive Tasks?
- DSGym: A Standardized and Holistic Framework for Evaluating and Training Data Science Agents
- DV-World: Benchmarking Data Visualization Agents in Real-World Scenarios
- FT-Dojo: Towards Autonomous LLM Fine-Tuning with Language Agents
- Hunt Instead of Wait: Evaluating Deep Data Research on Large Language Models
- Investigating Component Contributions in Multi-Agent ML Systems
- PostTrainBench: Can LLM Agents Automate LLM Post-Training?
- Procedural Generation Of Algorithm Discovery Tasks in Machine Learning
- MedAgentGym: A Scalable Agentic Training Environment for Code-Centric Reasoning in Biomedical Data Science
- DAComp: Benchmarking Data Agents across the Full Data Intelligence Lifecycle
- DARE-bench: Evaluating Modeling and Instruction Fidelity of LLMs in Data Science
- KRAMABENCH: A Benchmark for AI Systems on Data-to-Insight Pipelines over Data Lakes
- MLE-Smith: Scaling MLE Tasks with Automated Multi-agent Pipeline
- WebDS: An End-to-End Benchmark for Web-based Data Science
- DSCodeBench: A Realistic Benchmark for Data Science Code Generation
- Why Do Open-Source LLMs Struggle with Data Analysis? A Systematic Empirical Study
- UniDataBench: Evaluating Data Analytics Agents Across Structured and Unstructured Data
- DataSciBench: An LLM Agent Benchmark for Data Science
- InsightEval: An Expert-Curated Benchmark for Assessing Insight Discovery in LLM-Driven Data Agents
- OPT-BENCH: Evaluating the Iterative Self-Optimization of LLM Agents in Large-Scale Search Spaces
- Benchmarking LLM Agents on Real-World Biological Database Curation for Data-Driven Scientific Discovery
- FDABench: A Benchmark for Data Agents on Analytical Queries over Heterogeneous Data
- TemporalBench: A Benchmark for Evaluating LLM-Based Agents on Contextual and Event-Informed Time Series Tasks
- CTRL-ALT-DECEIT Sabotage Evaluations for Automated AI R&D
- Measuring AI Ability to Complete Long Software Tasks
- MLE-Dojo: Interactive Environments for Empowering LLM Agents in Machine Learning Engineering
- MLRC-Bench: Can Language Agents Solve Machine Learning Research Challenges?
- RADAR: Benchmarking Language Models on Imperfect Tabular Data
- RE-Bench: Evaluating Frontier AI R&D Capabilities of Language Model Agents against Human Experts
- Agent-as-a-Judge: Evaluate Agents with Agents
- Are Large Language Models Ready for Multi-Turn Tabular Data Analysis?
- MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering
- DSBench: How Far Are Data Science Agents from Becoming Data Science Experts?
- InsightBench: Evaluating Business Analytics Agents Through Multi-Step Insight Generation
- ComfyBench: Benchmarking LLM-based Agents in ComfyUI for Autonomously Designing Collaborative AI Systems
- Spider2-V: How Far Are Multimodal Agents From Automating Data Science and Engineering Workflows?
- Are Large Language Models Good Statisticians?
- Can LLMs Implicitly Learn Numeric Parameter Constraints in Data Science APIs?
- DACO: Towards Application-Driven and Comprehensive Data Analysis via Code Generation
- InfiAgent-DABench: Evaluating Agents on Data Analysis Tasks
- MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation
- Text2Analysis: A Benchmark of Table Question Answering with Advanced Data Analysis and Unclear Queries
- Benchmarking Data Science Agents
- Are LLMs Capable of Data-based Statistical and Causal Reasoning? Benchmarking Advanced Quantitative Reasoning with Data
- DCA-Bench: A Benchmark for Dataset Curation Agents
- DS-1000: A Natural and Reliable Benchmark for Data Science Code Generation
- Natural Language to Code Generation in Interactive Data Science Notebooks
B3 Paper reproduction & research-code generation 16
- From Reproduction to Replication: Evaluating Research Agents with Progressive Code Masking
- Paper2Code: Automating Code Generation from Scientific Papers in Machine Learning
- RECODE-H: A Benchmark for Research Code Development with Interactive Human Feedback
- Benchmarking PhD-Level Coding in 3D Geometric Computer Vision
- NERFIFY: A Multi-Agent Framework for Turning NeRF Papers into Code
- AutoReproduce: Automatic AI Experiment Reproduction with Paper Lineage
- RExBench: Can coding agents autonomously implement AI research extensions?
- SciCoQA: Quality Assurance for Scientific Paper–Code Alignment
- What Makes AI Research Replicable? Executable Knowledge Graphs as Scientific Knowledge Representations
- HiRAS: A Hierarchical Multi-Agent Framework for Paper-to-Code Generation and Execution
- PaperRepro: Automated Computational Reproducibility Assessment for Social Science Papers
- ReplicatorBench: Benchmarking LLM Agents for Replicability in Social and Behavioral Sciences
- ResearchCodeBench: Benchmarking LLMs on Implementing Novel Machine Learning Research Code
- The Automated LLM Speedrunning Benchmark: Reproducing NanoGPT Improvements
- PaperBench: Evaluating AI’s Ability to Replicate AI Research
- REPRO-Bench: Can Agentic AI Systems Assess the Reproducibility of Social Science Research?
C Self-improving LLMs 458
Models and agents that improve from their own outputs: self-training, self-play, self-rewarding, self-correction, iterative preference optimisation, self-evolving agents, and the theory and failure modes of these loops.
C1 Self-training & bootstrapped reasoning (STaR family, self-generated data) 105
- Better, Faster: Harnessing Self-Improvement in Large Reasoning Models
- D²Evo: Dual Difficulty-Aware Self-Evolution for Data-Efficient Reinforcement Learning
- Learn to Think: Improving Multimodal Reasoning through Vision-Aware Self-Improvement Training
- One-Way Policy Optimization for Self-Evolving LLMs
- Reinforcement Learning via Self-Distillation
- Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language Models
- TMS: Trajectory-Mixed Supervision for On-Policy Self Distillation
- Through the Lens of Contrast: Self-Improving Visual Reasoning in VLMs
- NFT: Bridging Supervised Learning and Reinforcement Learning in Math Reasoning
- ProofOptimizer: Training Language Models to Simplify Proofs without Human Demonstrations
- RESTRAIN: From Spurious Votes to Signals — Self-Training RL with Self-Penalization
- SkillFactory: Self-Distillation for Learning Cognitive Behaviors
- Synthetic Bootstrapped Pretraining
- ViPER: Empowering the Self-Evolution of Visual Perception Abilities in Vision-Language Models
- Decouple to Generalize: Context-First Self-Evolving Learning for Data-Scarce Vision-Language Reasoning
- OSPO: Object-Centric Self-Improving Preference Optimization for Text-to-Image Generation
- Incorporating Self-Rewriting into Large Language Model Reasoning Reinforcement
- K-STaR: Knowledge-Aware Self-Taught Reasoner
- MathSE: Improving Multimodal Mathematical Reasoning via Self-Evolving Iterative Reflection and Reward-Guided Fine-Tuning
- MedS³: Towards Medical Slow Thinking with Self-Evolved Soft Dual-sided Process Supervision
- SAPO: Self-Adaptive Process Optimization Makes Small Reasoners Stronger
- Beyond Meta-Reasoning: Metacognitive Consolidation for Self-Improving LLM Reasoning
- Beyond Ranking: Fine-Grained Diagnostics and Self-Improvement for MLLMs
- Learn Like Humans: Use Meta-cognitive Reflection for Efficient Self-Improvement
- UniCorn: Towards Self-Improving Unified Multimodal Models through Self-Generated Supervision
- A Dual-Phase Self-Evolution Framework for Large Language Models
- CoTEvol: Self-Evolving Chain-of-Thoughts for Data Synthesis in Mathematical Reasoning
- Easy Samples Are All You Need: Self-Evolving LLMs via Data-Efficient Reinforcement Learning
- GALA: Geometric Data Selection with Strategic Prospecting for Large Language Model Self-training
- IREASONER: Trajectory-Aware Intrinsic Reasoning Supervision for Self-Evolving Large Multimodal Models
- Many-Shot Scaling of In-Context Learning with Self-Generated Demonstrations
- Language Models can Self-Improve at State-Value Estimation for Better Search
- SoTA with Less: MCTS-Guided Sample Selection for Data-Efficient Visual Reasoning Self-Improvement
- AdaSTaR: Adaptive Data Sampling for Training Self-Taught Reasoners
- Enhancing the Outcome Reward-based RL Training of MLLMs with Self-Consistency Sampling
- ExPO: Unlocking Hard Reasoning with Self-Explanation-Guided Reinforcement Learning
- Learning to Better Search with Language Models via Guided Reinforced Self-Training
- OpenVLThinker: Complex Vision-Language Reasoning via Iterative SFT-RL Cycles
- Sparta Alignment: Collectively Aligning Multiple Language Models through Combat
- Structural Entropy Guided Agent for Detecting and Repairing Knowledge Deficiencies in LLMs
- The First Few Tokens Are All You Need: An Efficient and Effective Unsupervised Prefix Fine-Tuning Method for Reasoning Models
- rStar-Math: Small LLMs Can Master Math Reasoning with Self-Evolved Deep Thinking
- BRiTE: Bootstrapping Reinforced Thinking Process to Enhance Language Model Reasoning
- Diving into Self-Evolving Training for Multimodal Reasoning
- Rethinking Chain-of-Thought from the Perspective of Self-Training
- Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search
- Self-Improving Transformers Overcome Easy-to-Hard and Length Generalization Challenges
- ReGenesis: LLMs can Grow into Reasoning Generalists via Self-Improvement
- Lean-STaR: Learning to Interleave Thinking and Proving
- Automatic Curriculum Expert Iteration for Reliable LLM Reasoning
- B-STaR: Monitoring and Balancing Exploration and Exploitation in Self-Taught Reasoners
- Enhancing Cognition and Explainability of Multimodal Foundation Models with Self-Synthesized Data
- From Few to Many: Self-Improving Many-Shot Reasoners Through Iterative Optimization and Generation
- Learning to Clarify: Multi-turn Conversations with Action-Based Contrastive Self-Training
- Learning to Plan Before Answering: Self-Teaching LLMs to Learn Abstract Plans for Problem Solving
- Multiagent Finetuning: Self Improvement with Diverse Reasoning Chains
- Self-Boosting Large Language Models with Synthetic Preference Data
- Self-MoE: Towards Compositional Large Language Models with Self-Specialized Experts
- Smaller, Weaker, Yet Better: Training LLM Reasoners via Compute-Optimal Sampling
- Video-STaR: Self-Training Enables Video Instruction Tuning with Any Supervision
- SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation
- STEP: Enhancing Video-LLMs’ Compositional Reasoning by Spatio-Temporal Graph-guided Self-Training
- Bootstrapping Grounded Chain-of-Thought in Multimodal LLMs for Data-Efficient Model Adaptation
- ILLUME: Illuminating Your LLMs to See, Draw, and Self-Enhance
- OURO: A Self-Bootstrapped Framework for Enhancing Multimodal Scene Understanding
- Predict-Optimize-Distill: A Self-Improving Cycle for 4D Object Understanding
- Beyond Human Data: Aligning Multimodal Large Language Models by Iterative Self-Evolution
- Empowering Self-Learning of LLMs: Inner Knowledge Explicitation as a Catalyst
- EvoChart: A Benchmark and a Self-Training Approach Towards Real-World Chart Understanding
- Importance Weighting Can Help Large Language Models Self-Improve
- InverseCoder: Self-improving Instruction-Tuned Code LLMs with Inverse-Instruct
- Genius: A Generalizable and Purely Unsupervised Self-Training Framework For Advanced Reasoning
- Interactive Evolution: A Neural-Symbolic Self-Training Framework For Large Language Models
- R2-MultiOmnia: Leading Multilingual Multimodal Reasoning via Self-Training
- Self-Error-Instruct: Generalizing from Errors for LLMs Mathematical Reasoning
- Self-Taught Agentic Long Context Understanding
- STaR-SQL: Self-Taught Reasoner for Text-to-SQL
- Breaking the Reasoning Barrier A Survey on LLM Complex Reasoning through the Lens of Self-Evolution
- CARE-STaR: Constraint-aware Self-taught Reasoner
- Generate, Discriminate, Evolve: Enhancing Context Faithfulness via Fine-Grained Sentence-Level Self-Evolution
- Let’s Be Self-generated via Step by Step: A Curriculum Learning Approach to Automated Reasoning with Large Language Models
- RISE: Reasoning Enhancement via Iterative Self-Exploration in Multi-hop Question Answering
- Self-Training Elicits Concise Reasoning in Large Language Models
- Self-Tuning: Instructing LLMs to Effectively Acquire New Knowledge through Self-Teaching
- SIKeD: Self-guided Iterative Knowledge Distillation for Mathematical Reasoning
- Sparse Rewards Can Self-Train Dialogue Agents
- Unlocking LLMs’ Self-Improvement Capacity with Autonomous Learning for Domain Adaptation
- AlphaMath Almost Zero: Process Supervision without Process
- Enhancing Large Vision Language Models with Self-Training on Image Comprehension
- ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search
- RL on Incorrect Synthetic Data Scales the Efficiency of LLM Math Reasoning by Eight-Fold
- RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedback
- Enabling Lanuguage Models to Implicitly Learn Self-Improvement
- Language Model Self-improvement by Reinforcement Learning Contemplation
- Self-Training Large Language Models for Improved Visual Program Synthesis With Visual Reinforcement
- ItD: Large Language Models Can Teach Themselves Induction through Deduction
- Self-Training with Direct Preference Optimization Improves Chain-of-Thought Reasoning
- Bootstrapping LLM-based Task-Oriented Dialogue Agents via Self-Talk
- Toolformer: Language Models Can Teach Themselves to Use Tools
- Training Chain-of-Thought via Latent-Variable Inference
- Formal Mathematics Statement Curriculum Learning
- Q: How To Specialize Large Vision-Language Models to Data-Scarce VQA Tasks? A: Self-Train on Unlabeled Images!
- The Dialog Must Go On: Improving Visual Dialog via Generative Self-Training
- Better Language Models of Code through Self-Improvement
- STaR: Bootstrapping Reasoning With Reasoning
C2 Self-play & zero-data RL 54
- $\textit{S}$-SPPO: Semantic-Calibrated Self-Play Preference Optimization
- Anchoring Self-Play for Code Repair
- Chasing Moving Targets with Online Self-Play Reinforcement Learning for Safer Language Models
- CPMöbius: Iterative Coach–Player Reasoning for Data-Free Reinforcement Learning
- Disentangling Intent from Role: Adversarial Self-Play for Persona-Invariant Safety Alignment
- Evolving Quantitative Reasoning through Self-Play in Digital Twin Markets
- Learn from Your Mistakes: Tree-like Self-Play for Secure Code LLMs
- OCNR: Stabilizing Self-Play by Mitigating Iteration-Collapse With One-Class Novelty Rewards
- Propose, Solve, Verify: Self-Play Through Formal Verification
- R-Diverse: Mitigating Diversity Illusion in Self-Play LLM Training
- RSPO: Regularized Self-Play Alignment of Large Language Models
- Teaching Models to Teach Themselves: Reasoning at the Edge of Learnability
- Toward Training Superintelligent Software Agents through Self-Play SWE-RL
- Beyond Pass@ 1: Self-Play with Variational Problem Synthesis Sustains RLVR
- MARSHAL: Incentivizing Multi-Agent Reasoning via Self-Play with Strategic LLMs
- R-Zero: Self-Evolving Reasoning LLM from Zero Data
- Search Self-Play: Pushing the Frontier of Agent Capability without Supervision
- SPELL: Self-Play Reinforcement Learning for Evolving Long-Context Language Models
- SPIRAL: Self-Play on Zero-Sum Games Incentivizes Reasoning via Multi-Agent Multi-Turn Reinforcement Learning
- Vision-Zero: Scalable VLM Self-Evolution via Multi-Agent Self-Play
- VisPlay: Self-Evolving Vision-Language Models
- Hide and Seek with LLMs: An Adversarial Game for Sneaky Error Generation and Self-Improving Diagnosis
- FormulaSPIN: Self-Play Fine-Tuning for Natural Language to Spreadsheet Formula Generation
- Stratagem: Learning Transferable Reasoning via Trajectory-Modulated Game Self-Play
- Team-Based Self-Play With Dual Adaptive Weighting for Fine-Tuning LLMs
- TriPlay-RL: Tri-Role Self-Play Reinforcement Learning for LLM Safety Alignment
- WIST: Web-Grounded Iterative Self-Play Tree for Domain-Targeted Reasoning Improvement
- Be Your Own Red Teamer: Safety Alignment via Self-Play and Reflective Experience Replay
- Improving LLM Code Reasoning via Semantic Equivalence Self-Play with Formal Verification
- Revisiting Self-Play Preference Optimization: On the Role of Prompt Difficulty
- Absolute Zero: Reinforced Self-play Reasoning with Zero Data
- AceSearcher: Bootstrapping Reasoning and Search for LLMs via Reinforced Self-Play
- Self-Challenging Language Model Agents
- SeRL: Self-play Reinforcement Learning for Large Language Models with Limited Data
- SPACE: Noise Contrastive Estimation Stabilizes Self-Play Fine-Tuning for Large Language Models
- SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning
- Token-Level Self-Play with Importance-Aware Guidance for Large Language Models
- Triplets Better Than Pairs: Towards Stable and Effective Self-Play Fine-Tuning for LLMs
- AMPO: Active Multi Preference Optimization for Self-play Preference Selection
- Improving Rationality in the Reasoning Process of Language Models through Self-playing Game
- STP: Self-play LLM Theorem Provers with Iterative Conjecturing and Proving
- Iterative Nash Policy Optimization: Aligning LLMs with General Preferences via No-Regret Learning
- Self-play with Execution Feedback: Improving Instruction-following Capabilities of Large Language Models
- Self-Play Preference Optimization for Language Model Alignment
- SPaR: Self-Play with Tree-Search Refinement to Improve Instruction-Following in Large Language Models
- SEAS: Self-Evolving Adversarial Safety Optimization for Large Language Models
- VERSE: Verification-based Self-Play for Code Instructions
- Fixing Distribution Shifts of LLM Self-Critique via On-Policy Self-Play Training
- Self-play through Computational Runtimes improves Chart Reasoning
- TRANS-ZERO: Self-Play Incentivizes Large Language Models for Multilingual Translation Without Parallel Data
- Learning Formal Mathematics From Intrinsic Motivation
- Coevolving with the Other You: Fine-Tuning LLM with Sequential Cooperative Multi-Agent Reinforcement Learning
- Self-playing Adversarial Language Game Enhances LLM Reasoning
- Self-Play Fine-Tuning Converts Weak Language Models to Strong Language Models
C3 Self-rewarding, self-verification & unsupervised / test-time RL 70
- Are LLM Evaluators Really Narcissists? Sanity Checking Self-Preference Evaluations
- Beyond Majority Voting: Self-Reflective Test-Time Reinforcement Learning for LLM Reasoning
- Breaking the Self-Confirming Loop: Diagnosing and Mitigating Systemic Reward Bias in Self-Rewarding RL
- Conversation for Non-verifiable Learning: Self-Evolving Large Language Models through Meta-Evaluation
- ECHO: Entropy-Confidence Hybrid Optimization for Test-Time Reinforcement Learning
- Future-Gain Guided Test-Time Learning for Large Language Models
- Learning to Self-Verify Makes Language Models Better Reasoners
- One-shot Entropy Minimization for Language Model Reasoning
- Reasoning on the Manifold: Bidirectional Consistency for Self-Verification in Diffusion Language Models
- Rubric Curriculum RL: Exploiting the Generation-Verification Gap in Non-Verifiable Domains
- Spurious Rewards: Rethinking Training Signals in RLVR
- Temporal Self-Rewarding Language Models: Decoupling Chosen-Rejected via Past-Future
- V1: Unifying Generation and Self-Verification for Parallel Reasoners
- Co-rewarding: Stable Self-supervised RL for Eliciting Reasoning in Large Language Models
- DuPO: Enabling Reliable Self-Verification via Dual Preference Optimization
- How Far Can Unsupervised RLVR Scale LLM Training?
- LaSeR: Reinforcement Learning with Last-Token Self-Rewarding
- Learning to Reason without External Rewards
- Reinforcing General Reasoning Without Verifiers
- Self-Aligned Reward: Towards Effective and Efficient Reasoners
- SELF-HARMONY: LEARNING TO HARMONIZE SELF-SUPERVISION AND SELF-PLAY IN TEST-TIME REINFORCEMENT LEARNING
- Turning Internal Gap into Self-Improvement: Promoting the Generation-Understanding Unification in MLLMs
- Vision-SR1: Self-Rewarding Vision-Language Model via Reasoning Decomposition and Multi-Reward Policy Optimization
- Let VLMs Grade Their Own Thoughts: A Self-Quantification Approach to Reasoning-Aware Reward Modeling
- SafeGRPO: Self-Rewarded Multimodal Safety Alignment via Rule-Governed Policy Optimization
- TTRV: Test-Time Reinforcement Learning for Vision Language Models
- Unified Generation and Self-Verification for Vision-Language Models via Advantage Decoupled Preference Optimization
- A Rolling Stone Gathers No Moss: Adaptive Policy Optimization for Stable Self-Evaluation in Large Multimodal Models
- CEC-Zero: Zero-Supervision Character Error Correction with Self-Generated Rewards
- GRAM-R²: Self-Training Generative Foundation Reward Models for Reward Reasoning
- In-Token Rationality Optimization: Towards Accurate and Concise LLM Reasoning via Self-Feedback
- OR-R1: Automating Modeling and Solving of Operations Research Optimization Problem via Test-Time Reinforcement Learning
- SERL: Self-Examining Reinforcement Learning on Open-Domain
- Test-Time Reinforcement Learning for GUI Grounding via Region Consistency
- Beyond Majority Voting: Towards Fine-grained and More Reliable Reward Signal for Test-Time Reinforcement Learning
- CURE: Critique-Driven Unified Reinforcement Learning for Test-Time Self-Improvement
- Free Energy-Driven Reinforcement Learning with Adaptive Advantage Shaping for Unsupervised Reasoning in LLMs
- Instructions are all you need: Self-supervised Reinforcement Learning for Instruction Following
- SSR-Zero: Simple Self-Rewarding Reinforcement Learning for Machine Translation
- Understanding and Mitigating Spurious Signal Amplification in Test-Time Reinforcement Learning for Math Reasoning
- LC-ERD: Mining Latent Logic for Self-Evolving Reasoning via Consistency-Regulated Reward Decomposition
- SUDER: Self-Improving Unified Large Multimodal Models for Understanding and Generation with Dual Self-rewards
- Right Question is Already Half the Answer: Fully Unsupervised LLM Reasoning Incentivization
- Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning
- First SFT, Second RL, Third UPT: Continual Improving Multi-Modal LLM Reasoning via Unsupervised Post-Training
- Incentivizing LLMs to Self-Verify Their Answers
- Self-Verifying Reflection Helps Transformers with CoT Reasoning
- Stackelberg Self-Annotation: A Robust Approach to Data-Efficient LLM Alignment
- The Unreasonable Effectiveness of Entropy Minimization in LLM Reasoning
- Trust, But Verify: A Self-Verification Approach to Reinforcement Learning with Verifiable Rewards
- TTRL: Test-Time Reinforcement Learning
- ReVISE: Learning to Refine at Test-Time via Intrinsic Self-Verification
- Self-Consistency Preference Optimization
- CREAM: Consistency Regularized Self-Rewarding Language Models
- On the self-verification limitations of large language models on reasoning and planning tasks
- SELF-EVOLVED REWARD LEARNING FOR LLMS
- LLaVA-Critic: Learning to Evaluate Multimodal Models
- Improving Large Vision and Language Models by Learning from a Panel of Peers
- Unsupervised Visual Chain-of-Thought Reasoning via Preference Optimization
- Process-based Self-Rewarding Language Models
- Self-Rewarding Large Vision-Language Models for Optimizing Prompts in Text-to-Image Generation
- LLM Evaluators Recognize and Favor Their Own Generations
- Time-Reversal Provides Unsupervised Feedback to LLMs
- Calibrated Self-Rewarding Vision Language Models
- Self-Rewarding Language Models
- SelfCheck: Using LLMs to Zero-Shot Check Their Own Step-by-Step Reasoning
- Solving Challenging Math Word Problems Using GPT-4 Code Interpreter with Code-based Self-Verification
- Aligning Large Language Models by On-Policy Self-Judgment
- Direct Large Language Model Alignment Through Self-Rewarding Contrastive Prompt Distillation
- Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations
C4 Self-correction, self-refinement & self-debugging 69
- Learning Self-Correction in Vision–Language Models via Rollout Augmentation
- ReSeek: A Self-Correcting Framework for Search Agents with Instructive Rewards
- A Stitch in Time Saves Nine: Proactive Self-Refinement for Language Models
- Goedel-Prover-V2: Scaling Formal Theorem Proving with Scaffolded Data Synthesis and Self-Correction
- Once-More: Continuous Self-Correction for Large Language Models via Perplexity-Guided Intervention
- RefineBench: Evaluating Refinement Capability of Language Models via Checklists
- Selection, Reflection and Self-Refinement: Revisit Reasoning Tasks via a Causal Lens
- Beyond Step Pruning: Information Theory Based Step-level Optimization for Self-Refining Large Language Models
- Step Back to Leap Forward: Self-Backtracking for Symbolic Reasoning and Planning in Language Models
- Self-Reflective Generation at Test Time
- CodeRise: Bootstrapping LLMs for Ultra Low-Resource Programming Languages via Progressive Self-Refinement Curriculum
- Learning to Refine: Self-Refinement of Parallel Reasoning in LLMs
- Meeseeks: A Feedback-Driven, Iterative Self-Correction Benchmark evaluating LLMs’ Instruction Following Capability
- CURE: Co-Evolving Coders and Unit Testers via Reinforcement Learning
- Afterburner: Reinforcement Learning Facilitates Self-Improving Code Efficiency Optimization
- Can LLMs Correct Themselves? A Benchmark of Self-Correction in LLMs
- FEEDBACK FRICTION: LLMs Struggle to Fully Incorporate External Feedback
- Sherlock: Self-Correcting Reasoning in Vision-Language Models
- Towards Self-Refinement of Vision-Language Models with Triangular Consistency
- AlphaVerus: Bootstrapping Formally Verified Code Generation through Self-Improving Translation and Treefinement
- Reinforce LLM Reasoning through Multi-Agent Reflection
- Self-Improving Language Models for Evolutionary Program Synthesis: A Case Study on ARC-AGI
- Teaching Language Models to Critique via Reinforcement Learning
- Training Language Models to Self-Correct via Reinforcement Learning
- Automated Proof Generation for Rust Code via Self-Evolution
- SuperCorrect: Advancing Small LLM Reasoning with Thought Template Distillation and Self-Correction
- Think Thrice Before You Act: Progressive Thought Refinement in Large Language Models
- Can Large Vision-Language Models Correct Semantic Grounding Errors By Themselves?
- Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning
- VISCO: Benchmarking Fine-Grained Critique and Correction Towards Self-Improvement in Visual Reasoning
- ReCoT: Reflective Self-Correction Training for Mitigating Confirmation Bias in Large Vision-Language Models
- SC-Captioner: Improving Image Captioning with Self-Correction by Reinforcement Learning
- MAGIC: Generating Self-Correction Guideline for In-Context Text-to-SQL
- S^3cMath: Spontaneous Step-Level Self-Correction Makes Large Language Models Better Mathematical Reasoners
- Confidence v.s. Critique: A Decomposition of Self-Correction Capability for LLMs
- ProgCo: Program Helps Self-Correction of Large Language Models
- Revisit Self-Debugging with Self-Generated Tests for Code Generation
- S^2R: Teaching LLMs to Self-verify and Self-correct via Reinforcement Learning
- Understanding the Dark Side of LLMs’ Intrinsic Self-Correction
- Critic-CoT: Boosting the Reasoning Abilities of Large Language Model via Chain-of-Thought Critic
- Entrospect: Information-Theoretic Self-Reflection Elicits Better Response Refinement of Small Language Models
- ReflectEvo: Improving Meta Introspection of Small LLMs by Learning Self-Reflection
- Self-Correction is More than Refinement: A Learning Framework for Visual and Language Reasoning Tasks
- Self-Critique Guided Iterative Reasoning for Multi-hop Question Answering
- Training Language Model to Critique for Better Refinement
- Unlocking Recursive Thinking of LLMs: Alignment via Refinement
- Advancing Tool-Augmented Large Language Models via Meta-Verification and Reflection Learning
- A Theoretical Understanding of Self-Correction through In-context Alignment
- EffiLearner: Enhancing Efficiency of Generated Code via Self-Optimization
- LeDex: Training LLMs to Better Self-Debug and Explain Code
- Recursive Introspection: Teaching Language Model Agents How to Self-Improve
- SelfCodeAlign: Self-Alignment for Code Generation
- CodeIt: Self-Improving Language Models with Prioritized Hindsight Replay
- CRITIC: Large Language Models Can Self-Correct with Tool-Interactive Critiquing
- Is Self-Repair a Silver Bullet for Code Generation?
- Large Language Models Cannot Self-Correct Reasoning Yet
- RAIN: Your Language Models Can Align Themselves without Finetuning
- Teaching Large Language Models to Self-Debug
- Self-correcting LLM-controlled Diffusion Models
- Small Language Model Can Self-Correct
- Self-Contrast: Better Reflection Through Inconsistent Solving Perspectives
- CriticBench: Benchmarking LLMs for Critique-Correct Reasoning
- Small Language Models Need Strong Verifiers to Self-Correct Reasoning
- Teaching Language Models to Self-Improve by Learning from Language Feedback
- Reflexion: language agents with verbal reinforcement learning
- Self-Refine: Iterative Refinement with Self-Feedback
- Generating Sequences by Learning to Self-Correct
- Language Models Can Teach Themselves to Program Better
- Self-Edit: Fault-Aware Code Editor for Code Generation
C5 Iterative preference optimisation & self-alignment 30
- Semantic Voting: A Self-Evaluation-Free Approach for Efficient LLM Self-Improvement on Unverifiable Open-ended Tasks
- Bootstrapping LLMs via Preference-Based Policy Optimization
- SPA: Achieving Consensus in LLM Alignment via Self-Priority Optimization
- Aligning Large Language Models via Fully Self-Synthetic Data
- Latent Principle Discovery for Language Model Self-Improvement
- Self-alignment of Large Video Language Models with Refined Regularized Preference Optimization
- SelfCite: Self-Supervised Alignment for Context Attribution in Large Language Models
- Preference Optimization for Reasoning with Pseudo Feedback
- Bootstrapping Language Models with DPO Implicit Rewards
- Building Math Agents with Multi-Turn Iterative Preference Learning
- Language Imbalance Driven Rewarding for Multilingual Self-improving
- LongPO: Long Context Self-Evolution of Large Language Models through Short-to-Long Preference Optimization
- Self-Improving Robust Preference Optimization
- Mitigating Hallucinations in Large Vision-Language Models via DPO: On-Policy Data Hold the Key
- ISR-DPO: Aligning Large Multimodal Models for Videos by Iterative Self-Retrospective DPO
- Self-Evolutionary Large Language Models Through Uncertainty-Enhanced Preference Optimization
- Subtle Errors in Reasoning: Preference Learning via Error-injected Self-editing
- APT: Improving Specialist LLM Performance with Weakness Case Acquisition and Iterative Preference Training
- Self-Improvement Towards Pareto Optimality: Mitigating Preference Conflicts in Multi-Objective Alignment
- Chain of Preference Optimization: Improving Chain-of-Thought Reasoning in LLMs
- Iterative Reasoning Preference Optimization
- Self-Supervised Alignment with Mutual Information: Learning to Follow Principles without Preference Labels
- Self-Alignment of Large Language Models via Monopolylogue-based Social Scene Simulation
- Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-constraint
- Self-Alignment with Instruction Backtranslation
- RLCD: Reinforcement Learning from Contrastive Distillation for LM Alignment
- SALMON: Self-Alignment with Instructable Reward Models
- Self-Alignment for Factuality: Mitigating Hallucinations in LLMs via Self-Evaluation
- Principle-Driven Self-Alignment of Language Models from Scratch with Minimal Human Supervision
- Self-Instruct: Aligning Language Models with Self-Generated Instructions
C6 Self-evolving agents & systems 83
- Agent0-VL: Exploring Self-Evolving Agent for Tool-Integrated Vision-Language Reasoning
- Darwinian Memory: A Training-Free Self-Regulating Memory System for GUI Agent Evolution
- EVOLVING ROLLOUTS: Harnessing Historical Experience for Web Agent Evolution in Reinforcement Learning
- From Interactions to Principles: Experience-Driven Self-Distillation for Evolving LLM Agents
- From Player to Master: Enhancing Test-Time Learning of LLM Agents via Reinforcement Learning over Memory
- MemEvolve: Meta-Evolution of Agent Memory Systems
- SE-GA: Memory-Augmented Self-Evolution for GUI Agents
- SEAgent: Self-Evolving Computer Use Agent with Autonomous Learning from Experience
- Self-evolving LLM agents with in-distribution Optimization
- ThetaEvolve: Test-time Learning on Open Problems
- Towards Feedback-to-Plan Decisions for Self-Evolving LLM Agents in CUDA Kernel Generation
- Huxley-G\"odel Machine: Human-Level Coding Agent Development by an Approximation of the Optimal Self-Improving Machine
- Agentic Context Engineering: Evolving Contexts for Self-Improving Language Models
- CoMAS: Co-Evolving Multi-Agent Systems via Interaction Rewards
- Darwin Gödel Machine: Open-Ended Evolution of Self-Improving Agents
- EvoTest: Evolutionary Test-Time Learning for Self-Improving Agentic Systems
- Learn the Ropes, Then Trust the Wins: Self-imitation with Progressive Exploration for Agentic Reinforcement Learning
- MAS$^2$: Self-Generative, Self-Configuring, Self-Rectifying Multi-Agent Systems
- MemGen: Weaving Generative Latent Memory for Self-Evolving Agents
- ReasoningBank: Scaling Agent Self-Evolving with Reasoning Memory
- ReVeal: Self-Evolving Code Agents via Reliable Self-Verification
- StepORLM: A Self-Evolving Framework With Generative Process Supervision For Operations Research Language Models
- Test-Time Adaptation for LLM Agents via Environment Interaction
- Toward Effective Tool-Integrated Reasoning via Self-Evolved Preference Learning
- Towards Self-Evolving Agent Benchmarks : Validatable Agent Trajectory via Test-Time Exploration
- Your Agent May Misevolve: Emergent Risks in Self-evolving LLM Agents
- EvoGraph-R1: Self-Evolving Multimodal Knowledge Hypergraphs for Agentic Retrieval
- JarvisEvo: Towards a Self-Evolving Photo Editing Agent with Synergistic Editor-Evaluator Optimization
- Learning to Adapt: Self-Improving Web Agent via Cognitive-Aware Exploration
- Co-EPG: A Framework for Co-Evolution of Planning and Grounding in Autonomous GUI Agents
- Evolving Generalist Virtual Agents with Generative and Associative Memory
- HiVA: Self-organized Hierarchical Variable Agent via Goal-driven Semantic-Topological Evolution
- EVOTOOL: Self-Evolving Tool-Use Policy Optimization in LLM Agents via Blame-Aware Mutation and Diversity-Aware Selection
- GUI0: Self-Evolving Foundational GUI Agents in Super App Ecosystems
- Mem^2Evolve: Towards Self-Evolving Agents via Co-Evolutionary Capability Expansion and Experience Distillation
- Reinforcement Learning for Self-Improving Agent with Skill Library
- SEARL: Joint Optimization of Policy and Tool Graph Memory for Self-Evolving Agents
- Learning to Evolve: A Self-Improving Framework for Multi-Agent Systems via Textual Parameter Graph Optimization
- On Safety Risks in Experience-Driven Self-Evolving Agents
- POLARIS: A Gödel Agent Framework for Small Language Models through Experience-Abstracted Policy Repair
- Remember Me, Refine Me: A Dynamic Procedural Memory Framework for Experience-Driven Agent Evolution
- Self-Evolving Multi-Agent Systems via Textual Backpropagation
- Towards Self-Evolving Agents: Enabling Autonomy through Interactive Experience Refinement
- Towards Self-Improving Error Diagnosis in Multi-Agent Systems
- TT-SI: Self-Improving LLM Agents with Test-Time Training
- AlphaOPT: Formulating Optimization Programs with Self-Improving LLM Experience Library
- Self-Evolutionary Reinforced Knowledge Distillation for Multi-Modal Tool-Use Agents
- AgentBreeder: Mitigating the AI Safety Risks of Multi-Agent Scaffolds via Self-Improvement
- Collaborative Reasoner: Self-Improving Social Agents with Synthetic Conversations
- Enhancing GUI Agent with Uncertainty-Aware Self-Trained Evaluator
- Iterative Self-Incentivization Empowers Large Language Models as Agentic Searchers
- MindGYM: What Matters in Question Synthesis for Thinking-Centric Fine-Tuning?
- SE-Agent: Self-Evolution Trajectory Optimization in Multi-Step Reasoning with LLM-Based Agents
- SE-GUI: Enhancing Visual Grounding for GUI Agents via Self-Evolutionary Reinforcement Learning
- SEEA-R1: Tree-Structured Reinforcement Fine-Tuning for Self-Evolving Embodied Agents
- Self-Adapting Language Models
- Self-Generated In-Context Examples Improve LLM Agents for Sequential Decision-Making Tasks
- Self-Improving Embodied Foundation Models
- SiriuS: Self-improving Multi-agent Systems via Bootstrapped Reasoning
- UI-Genie: A Self-Improving Approach for Iteratively Boosting MLLM-based Mobile GUI Agents
- Agent Workflow Memory
- Bootstrapping Self-Improvement of Language Model Programs for Zero-Shot Schema Matching
- Position: Truly Self-Improving Agents Require Intrinsic Metacognitive Learning
- Test-Time Learning for Large Language Models
- From Exploration to Mastery: Enabling LLMs to Master Tools via Self-Driven Interactions
- Learn-by-interact: A Data-Centric Framework For Self-Adaptive Agents in Realistic Environments
- Self-Evolving Multi-Agent Collaboration Networks for Software Development
- Strategist: Self-improvement of LLM Decision Making via Bi-Level Tree Search
- WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning
- Self-Evolving Visual Concept Library using Vision-Language Critics
- Contextual Experience Replay for Self-Improvement of Language Agents
- Gödel Agent: A Self-Referential Agent Framework for Recursively Self-Improvement
- MobileSteward: Integrating Multiple App-Oriented Agents with Self-Evolution to Automate Cross-App Instructions
- VLM Agents Generate Their Own Memories: Distilling Experience into Embodied Programs of Thought
- Richelieu: Self-Evolving LLM-Based Agents for AI Diplomacy
- StreamBench: Towards Benchmarking Continuous Improvement of Language Agents
- Toward Self-Improvement of LLMs via Imagination, Searching, and Criticizing
- BAGEL: Bootstrapping Agents by Guiding Exploration with Language
- Promptbreeder: Self-Referential Self-Improvement via Prompt Evolution
- CLOVA: A Closed-LOop Visual Assistant with Tool Usage and Update
- Hypothesis, Verification, and Induction: Grounding Large Language Models with Self-Driven Skill Learning
- Self-Evolving GPT: A Lifelong Autonomous Experiential Learner
- Way to Specialist: Closing Loop Between Specialized LLM and Evolving Domain Knowledge Graph
C7 Theory, limits & synthetic-data loops / model collapse 47
- A Task-centric Theory for Iterative Self-Improvement with Easy-to-Hard Curricula
- Curated Synthetic Data Doesn’t Have to Collapse: A Theoretical Study of Generative Retraining with Pluralistic Preferences
- Differential Smoothing Mitigates Sharpening and Improves LLM Reasoning
- Language Generation with Replay: A Learning-Theoretic View of Model Collapse
- Large Language Model Agents Are Not Always Faithful Self-Evolvers
- On the Generalization Gap in Self-Evolving Language Model Reasoning
- Position: Multiple Definitions & Unrealistic Assumptions of Model Collapse Distract from Real World Threats
- Position: Self-Play Only Evolves When Self-Synthetic Pipeline Ensures Learnable Information Gain
- Position: the Stochastic Parrot in the Coal Mine. Model Collapse is a Threat to Low-Resource Communities
- When and How Human Curation Backfires: Preference Alignment under Multi-Model Self-Consuming Loop
- When Sample Selection Bias Precipitates Model Collapse
- Escaping Model Collapse via Synthetic Data Verification: Near-term Improvements and Long-term Convergence
- Preventing Model Collapse Under Overparametrization: Optimal Mixing Ratios for Interpolation Learning and Ridge Regression
- Provable and Practical In-Context Policy Optimization for Self-Improvement
- Theoretical Modeling of Large Language Model Self-Improvement Training Dynamics Through Solver-Verifier Gap
- The Alignment Game: A Theory of Long-Horizon Alignment Through Recursive Curation
- Counteracting the Matthew Effect in Self-Improvement of LVLMs through Head-Tail Re-balancing
- Data Pollination: An Emergent Ecological Process Driving AI Population Evolution
- Escaping the Echo Trap: On Credit Assignment Failure in Multi-turn LLM Self-Reflection
- Observations and Remedies for Large Language Model Bias in Self-Consuming Performative Loop
- A Closer Look at Model Collapse: From a Generalization-to-Memorization Perspective
- Spend Wisely: Maximizing Post-Training Gains in Iterative Synthetic Data Bootstrapping
- Escaping Collapse: The Strength of Weak Data for Large Language Model Training
- Self-Verification Provably Prevents Model Collapse in Recursive Synthetic Training
- Weaver: Shrinking the Generation-Verification Gap by Scaling Compute for Verification
- When Models Don’t Collapse: On the Consistency of Iterative MLE
- Scaling Test-Time Compute Without Verification or RL is Suboptimal
- Collapse or Thrive: Perils and Promises of Synthetic Data in a Self-Generating World
- How to Synthesize Text Data without Model Collapse?
- Self-Consuming Generative Models with Adversarially Curated Data
- Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models
- Self-Improvement in Language Models: The Sharpening Mechanism
- Strong Model Collapse
- A Theoretical Perspective: How to Prevent Model Collapse in Self-consuming Training Loops
- Beyond Model Collapse: Scaling Up with Synthesized Data Requires Verification
- Progress or Regress? Self-Improvement Reversal in Post-training
- SELF-[IN]CORRECT: LLMs Struggle with Discriminating Self-Generated Responses
- The Self-Improvement Paradox: Can Language Models Bootstrap Reasoning Capabilities without External Scaffolding?
- Self-Consuming Generative Models with Curated Data Provably Optimize Human Preferences
- Bias Amplification in Language Model Evolution: An Iterated Learning Perspective
- Model Collapse Demystified: The Case of Regression
- A Tale of Tails: Model Collapse as a Change of Scaling Laws
- Self-Correcting Self-Consuming Loops for Generative Model Training
- Towards Theoretical Understandings of Self-Consuming Generative Models
- On the Stability of Iterative Retraining of Generative Models on their own Data
- Self-Consuming Generative Models Go MAD
- Pride and Prejudice: LLM Amplifies Self-Bias in Self-Refinement
D AutoML 499
Neural architecture search, hyperparameter optimisation, AutoML systems, automated feature engineering, prior-data fitted networks, and the newer LLM-driven AutoML and agent/workflow search.
D1 Neural architecture search (methods, proxies, benchmarks) 196
- Axiomatic Atlas: A Prescriptive Framework for Neural Architecture Design
- Beyond Model Base Retrieval: Weaving Knowledge to Master Fine-grained Neural Network Design
- BIOARC: Discovering Optimal Neural Architectures for Biological Foundation Models
- MFH-NAS:A Hybrid Neural Architecture Search Framework for Multimodal Fusion Object Detection
- Optimizing Network Simulation: Enhancing Performance Prediction Accuracy via Neural Architecture Search
- pTNAS: Progressive Neural Architecture Search for Tabular Data
- Search Space Synthesis for Parametric Functions
- Composer: A Search Framework for Hybrid Neural Architecture Design
- Constraint-guided Hardware-aware NAS through Gradient Modification
- RF-DETR: Neural Architecture Search for Real-Time Detection Transformers
- Training-Free Determination of Network Width via Neural Tangent Kernel
- HyperNAS: Enhancing Architecture Representation for NAS Predictor via Hypernetwork
- InTrain: Intrinsic Trainability for Zero-Cost Neural Architecture Search
- Progressive Neural Architecture Generation
- Progressive Supernet Training for Efficient Visual Autoregressive Modeling
- TAS-LoRA: Transformer Architecture Search with Mixture-of-LoRA Experts
- Vision-Oriented Lightweight Neural Architecture Search with Budget-Adaptive Evaluation
- DA-DFGAS:Differentiable Federated Graph Neural Architecture Search with Distribution-Aware Attentive Aggregation
- Towards Robust Edge Model Adaptation via Elastic Architecture Search
- Understanding and Enhancing Differentiable Architecture Search from Information Bottleneck Perspective
- Breaking Robustness Barriers in Cognitive Diagnosis: A One-Shot Neural Architecture Search Perspective
- Jet-Nemotron: Efficient Language Model with Post Neural Architecture Search
- Learning to Flow from Generative Pretext Tasks for Neural Architecture Encoding
- Per-Architecture Training-Free Metric Optimization for Neural Architecture Search
- Searching Efficient Semantic Segmentation Architectures via Dynamic Path Selection
- TF-MAS: Training-free Mamba2 Architecture Search
- Prior Knowledge Guided Neural Architecture Generation
- Puzzle: Distillation-Based NAS for Inference-Optimized LLMs
- Revisiting Neural Networks for Few-Shot Learning: A Zero-Cost NAS Perspective
- Runtime Analysis of Evolutionary NAS for Multiclass Classification
- STAR: Synthesis of Tailored Architectures
- Entropy-based Activation Function Optimization: A Method on Searching Better Activation Functions
- Multi-objective Differentiable Neural Architecture Search
- NEAR: A Training-Free Pre-Estimator of Machine Learning Model Performance
- W-PCA Based Gradient-Free Proxy for Efficient Search of Lightweight Language Models
- Zero-cost Proxy for Adversarial Robustness Evaluation
- L-SWAG: Layer-Sample Wise Activation with Gradients Information for Zero-Shot NAS on Vision Transformers
- NN-Former: Rethinking Graph Structure in Neural Architecture Representation
- Subnet-Aware Dynamic Supernet Training for Neural Architecture Search
- Training-free Neural Architecture Search through Variance of Knowledge of Deep Network Weights
- Beyond the Limits: Overcoming Negative Correlation of Activation-Based Training-Free NAS
- CARL: Causality-guided Architecture Representation Learning for an Interpretable Performance Predictor
- Loss Functions for Predictor-based Neural Architecture Search
- Neural Architecture Search Driven by Locally Guided Diffusion for Personalized Federated Learning
- TRNAS: A Training-Free Robust Neural Architecture Search
- AutoSGNN: Automatic Propagation Mechanism Discovery for Spectral Graph Neural Networks
- Behavior Importance-Aware Graph Neural Architecture Search for Cross-Domain Recommendation
- Efficient Few-Shot Neural Architecture Search by Counting the Number of Nonlinear Functions
- HEP-NAS: Towards Efficient Few-shot Neural Architecture Search via Hierarchical Edge Partitioning
- ParZC: Parametric Zero-Cost Proxies for Efficient NAS
- Causal-aware Graph Neural Architecture Search under Distribution Shifts
- Towards Efficient Few-shot Graph Neural Architecture Search via Partitioning Gradient Contribution
- CE-NAS: An End-to-End Carbon-Efficient Neural Architecture Search Framework
- einspace: Searching for Neural Architectures from Fundamental Operations
- HW-GPT-Bench: Hardware-Aware Architecture Benchmark for Language Models
- MOTE-NAS: Multi-Objective Training-based Estimate for Efficient Neural Architecture Search
- Vision Transformer Neural Architecture Search for Out-of-Distribution Generalization: Benchmark and Insights
- Disentangled Continual Graph Neural Architecture Search with Invariant Modular Supernet
- Encodings for Prediction-based Neural Architecture Search
- Surprisingly Strong Performance Prediction with Neural Graph Features
- Towards Neural Architecture Search through Hierarchical Generative Modeling
- SWAP-NAS: Sample-Wise Activation Patterns for Ultra-fast NAS
- Aux-NAS: Exploiting Auxiliary Labels with Negligibly Extra Inference Cost
- DiffusionNAG: Predictor-guided Neural Architecture Generation with Diffusion Models
- Masked Distillation Advances Self-Supervised Transformer Architecture Search
- Robust NAS under adversarial training: benchmark, theory, and beyond
- Robustifying and Boosting Training-Free Neural Architecture Search
- AZ-NAS: Assembling Zero-Cost Proxies for Network Architecture Search
- Boosting Order-Preserving and Transferability for Neural Architecture Search: a Joint Architecture Refined Search and Fine-tuning Approach
- Building Optimal Neural Architectures using Interpretable Knowledge
- FlowerFormer: Empowering Neural Architecture Encoding using a Flow-aware Graph Transformer
- Insights from the Use of Previously Unseen Neural Architecture Search Datasets
- SNED: Superposition Network Architecture Search for Efficient Video Diffusion Model
- Towards Accurate and Robust Architectures via Neural Architecture Search
- Auto-Prox: Training-Free Vision Transformer Architecture Search via Automatic Proxy Discovery
- Data-Augmented Curriculum Graph Neural Architecture Search under Distribution Shifts
- DC-NAS: Divide-and-Conquer Neural Architecture Search for Multi-Modal Classification
- DCLP: Neural Architecture Predictor with Curriculum Contrastive Learning
- EG-NAS: Neural Architecture Search with Fast Evolutionary Exploration
- G-NAS: Generalizable Neural Architecture Search for Single Domain Generalization Object Detection
- Hypergraph Neural Architecture Search
- IS-DARTS: Stabilizing DARTS through Precise Measurement on Candidate Importance
- Multimodal Graph Neural Architecture Search under Distribution Shifts
- PerFedRLNAS: One-for-All Personalized Federated Neural Architecture Search
- SasWOT: Real-Time Semantic Segmentation Architecture Search WithOut Training
- UniADS: Universal Architecture-Distiller Search for Distillation Gap
- ETAS: Zero-Shot Transformer Architecture Search via Network Trainability and Expressivity
- Mixture-of-Supernets: Improving Weight-Sharing Supernet Training with Architecture-Routed Mixture-of-Experts
- Automatic Multi-Task Learning Framework with Neural Architecture Search in Recommendations
- AutoSTF: Decoupled Neural Architecture Search for Cost-Effective Automated Spatio-Temporal Forecasting
- Reinforced Compressive Neural Architecture Search for Versatile Adversarial Robustness
- SiGeo: Sub-One-Shot NAS via Geometry of Loss Landscape
- Towards Lightweight Graph Neural Network Search with Curriculum Graph Sparsification
- Rethinking Bias Mitigation: Fairer Architectures Make for Fairer Face Recognition
- MeCo: Zero-Shot NAS with One Data and Single Forward Pass via Minimum Eigenvalue of Correlation
- AutoGO: Automated Computation Graph Optimization for Neural Network Evolution
- Construction of Hierarchical Neural Architecture Search Spaces based on Context-free Grammars
- Efficient Activation Function Optimization through Surrogate Modeling
- Evolutionary Neural Architecture Search for Transformer in Knowledge Tracing
- Generalizable Lightweight Proxy for Robust NAS against Diverse Perturbations
- MathNAS: If Blocks Have a Role in Mathematical Architecture Design
- Multi-task Graph Neural Architecture Search with Task-aware Collaboration and Curriculum
- Operation-Level Early Stopping for Robustifying Differentiable NAS
- Unsupervised Graph Neural Architecture Search with Disentangled Self-Supervision
- Do Not Train It: A Linear Neural Architecture Search of Graph Neural Networks
- PreNAS: Preferred One-Shot Learning Towards Efficient Neural Architecture Search
- Rethink DARTS Search Space and Renovate a New Benchmark
- Shortest Edit Path Crossover: A Theory-driven Solution to the Permutation Problem in Evolutionary Neural Architecture Search
- AutoGT: Automated Graph Transformer Architecture Search
- Meta-prediction Model for Distillation-Aware NAS on Unseen Datasets
- Transfer NAS with Meta-learned Bayesian Surrogates
- ZiCo: Zero-shot NAS via inverse Coefficient of Variation on Gradients
- $\Lambda$-DARTS: Mitigating Performance Collapse by Harmonizing Operation Selection among Cells
- Improving Differentiable Neural Architecture Search by Encouraging Transferability
- Neural Architecture Design and Robustness: A Dataset
- Adversarially Robust Neural Architecture Search for Graph Neural Networks
- DeepMAD: Mathematical Architecture Design for Deep Convolutional Neural Network
- Differentiable Architecture Search With Random Features
- DisWOT: Student Architecture Search for Distillation WithOut Training
- EMT-NAS:Transferring Architectural Knowledge Between Tasks From Different Datasets
- HOTNAS: Hierarchical Optimal Transport for Neural Architecture Search
- MDL-NAS: A Joint Multi-Domain Learning Framework for Vision Transformer
- NAR-Former: Neural Architecture Representation Learning Towards Holistic Attributes Prediction
- PA&DA: Jointly Sampling Path and Data for Consistent NAS
- ElasticViT: Conflict-aware Supernet Training for Deploying Fast Vision Transformer on Diverse Mobile Devices
- EMQ: Evolving Training-free Proxies for Automated Mixed Precision Quantization
- Extensible and Efficient Proxy for Neural Architecture Search
- MixPath: A Unified Approach for One-shot Neural Architecture Search
- ROME: Robustifying Memory-Efficient NAS via Topology Disentanglement and Gradient Accumulation
- ShiftNAS: Improving One-shot NAS via Probability Shift
- SpaceEvo: Hardware-Friendly Search Space Design for Efficient INT8 Inference
- Unleashing the Power of Gradient Signal-to-Noise Ratio for Zero-Shot NAS
- AIO-P: Expanding Neural Performance Predictors beyond Image Classification
- AutoNF: Automated Architecture Optimization of Normalizing Flows with Unconstrained Continuous Relaxation Admitting Optimal Discrete Solution
- Dynamic Ensemble of Low-Fidelity Experts: Mitigating NAS “Cold-Start”
- Dynamic Heterogeneous Graph Attention Neural Architecture Search
- GENNAPE: Towards Generalized Neural Architecture Performance Estimators
- NAS-LID: Efficient Neural Architecture Search with Local Intrinsic Dimension
- Neural Architecture Search for Wide Spectrum Adversarial Robustness
- PatchNAS: Repairing DNNs in Deployment with Patched Network Architecture Search
- PINAT: A Permutation INvariance Augmented Transformer for NAS Predictor
- ProxyBO: Accelerating Neural Architecture Search via Bayesian Optimization with Zero-Cost Proxies
- Zero-Cost Operation Scoring in Differentiable Architecture Search
- Training-free Neural Architecture Search for RNNs and Transformers
- AutoMoE: Heterogeneous Mixture-of-Experts with Adaptive Computation for Efficient Neural Machine Translation
- Neural Architecture Search for Parameter-Efficient Fine-tuning of Large Pre-trained Language Models
- Efficient and Joint Hyperparameter and Architecture Search for Collaborative Filtering
- BLOX: Macro Neural Architecture Search Benchmark and Algorithms
- LiteTransformerSearch: Training-free Neural Architecture Search for Efficient Language Models
- Bridge the Gap Between Architecture Spaces via A Cross-Domain Predictor
- Efficient Architecture Search for Diverse Tasks
- EZNAS: Evolving Zero-Cost Proxies For Neural Architecture Scoring
- Few-shot Task-agnostic Neural Architecture Search for Distilling Large Language Models
- Generalization Properties of NAS under Activation and Skip Connection Search
- Interpreting Operation Selection in Differentiable Architecture Search: A Perspective from Influence-Directed Explanations
- JAHS-Bench-201: A Foundation For Research On Joint Architecture And Hyperparameter Search
- NAS-Bench-360: Benchmarking Neural Architecture Search on Diverse Tasks
- NAS-Bench-Graph: Benchmarking Graph Neural Architecture Search
- NAS-Bench-Suite-Zero: Accelerating Research on Zero Cost Proxies
- Saliency-Aware Neural Architecture Search
- TabNAS: Rejection Sampling for Neural Architecture Search on Tabular Datasets
- Unifying and Boosting Gradient-Based Training-Free Neural Architecture Search
- ZARTS: On Zero-order Optimization for Neural Architecture Search
- Deep and Flexible Graph Neural Architecture Search
- MAE-DET: Revisiting Maximum Entropy Principle in Zero-Shot NAS for Efficient Object Detection
- AGNAS: Attention-Guided Micro- and Macro-Architecture Search
- Analyzing and Mitigating Interference in Neural Architecture Search
- Graph Neural Architecture Search Under Distribution Shifts
- Large-Scale Graph Neural Architecture Search
- Generalizing Few-Shot NAS with Gradient Matching
- Learning Versatile Neural Architectures by Propagating Network Codes
- NAS-Bench-Suite: NAS Evaluation is (Now) Surprisingly Easy
- NASI: Label- and Data-agnostic Neural Architecture Search at Initialization
- NASViT: Neural Architecture Search for Efficient Vision Transformers with Gradient Conflict aware Supernet Training
- On Redundancy and Diversity in Cell-based Neural Architecture Search
- SUMNAS: Supernet with Unbiased Meta-Features for Neural Architecture Search
- Surrogate NAS Benchmarks: Going Beyond the Limited Search Spaces of Tabular NAS Benchmarks
- Arch-Graph: Acyclic Architecture Relation Predictor for Task-Transferable Neural Architecture Search
- b-DARTS: Beta-Decay Regularization for Differentiable Architecture Search
- BaLeNAS: Differentiable Architecture Search via the Bayesian Learning Rule
- Demystifying the Neural Tangent Kernel From a Practical Perspective: Can It Be Trusted for Neural Architecture Search Without Training?
- Distribution Consistent Neural Architecture Search
- Global Convergence of MAML and Theory-Inspired Neural Architecture Search for Few-Shot Learning
- HyperSegNAS: Bridging One-Shot Neural Architecture Search With 3D Medical Image Segmentation Using HyperNet
- ISNAS-DIP: Image-Specific Neural Architecture Search for Deep Image Prior
- Learning To Learn by Jointly Optimizing Neural Architecture and Weights
- Neural Architecture Search With Representation Mutual Information
- Performance-Aware Mutual Knowledge Distillation for Improving Neural Architecture Search
- Shapley-NAS: Discovering Operation Contribution for Neural Architecture Search
- Training-Free Transformer Architecture Search
- AutoBERT-Zero: Evolving BERT Backbone from Scratch
- BM-NAS: Bilevel Multimodal Neural Architecture Search
- DPNAS: Neural Architecture Search for Deep Learning with Differential Privacy
- Learning from Mistakes – a Framework for Neural Architecture Search
- AutoFAS: Automatic Feature and Architecture Selection for Pre-Ranking System
- Graph Neural Networks with Node-wise Architecture
D2 Hyperparameter optimisation (multi-fidelity, meta / in-context BO, transfer) 89
- $\alpha$-PFN: Fast Entropy Search via In-Context Learning
- $\mu$pscaling small models: Principled warm starts and hyperparameter transfer
- Cost-aware Stopping for Bayesian Optimization
- Hyperparameter Transfer Laws for Non-Recurrent Multi-Path Neural Networks
- Hyperparameter Transfer with Mixture-of-Expert Layers
- Provably Data-driven Multiple Hyper-parameter Tuning with Structured Loss Function
- TabPack: Efficient Hyperparameter Ensembles for Tabular Deep Learning
- Completed Hyperparameter Transfer across Modules, Width, Depth, Batch and Duration
- GIT-BO: High-Dimensional Bayesian Optimization with Tabular Foundation Models
- Understanding the Mechanisms of Fast Hyperparameter Transfer
- Weight Decay may matter more than µP for Learning Rate Transfer in Practice
- DABO: Difficulty-Aware Bayesian Optimization with Diffusion-Learned Priors
- Difficulty-Aware Learning Curve Extrapolation
- HyperSHAP: Shapley Values and Interactions for Explaining Hyperparameter Optimization
- LAMDA: Two-Phase HPO via Learning Prior from Low-Fidelity Data
- PSEO: Optimizing Post-hoc Stacking Ensemble Through Hyperparameter Tuning
- Efficient Hyperparameter Optimization for LLM Reinforcement Learning
- Conditional PED-ANOVA: Hyperparameter Importance in Hierarchical & Dynamic Search Spaces
- Private Hyperparameter Tuning with Ex-Post Guarantee
- CAMO: Convergence-Aware Multi-Fidelity Bayesian Optimization
- Cost-Sensitive Freeze-thaw Bayesian Optimization for Efficient Hyperparameter Tuning
- Data Mixture Optimization: A Multi-fidelity Multi-scale Bayesian Framework
- LCDB 1.1: A Database Illustrating Learning Curves Are More Ill-Behaved Than Previously Thought
- Multi-Objective Hyperparameter Selection via Hypothesis Testing on Reliability Graphs
- Adaptive Learn-then-Test: Statistically Valid and Efficient Hyperparameter Selection
- AutoUAD: Hyper-parameter Optimization for Unsupervised Anomaly Detection
- Meta-Learning Hyperparameters for Parameter Efficient Fine-Tuning
- ULTHO: Ultra-Lightweight yet Efficient Hyperparameter Optimization in Deep Reinforcement Learning
- A Stochastic Approach to Bi-Level Optimization for Hyperparameter Optimization and Meta Learning
- Architecture-Aware Learning Curve Extrapolation via Graph Ordinary Differential Equation
- FedPop: Federated Population-based Hyperparameter Tuning
- Modeling All Response Surfaces in One for Conditional Search Spaces
- Bayesian Stream Tuner: Dynamic Hyperparameter Optimization for Real-Time Data Streams
- A Method for Evaluating Hyperparameter Sensitivity in Reinforcement Learning
- Lower Bounds of Uniform Stability in Gradient-Based Bilevel Algorithms for Hyperparameter Optimization
- Reshuffling Resampling Splits Can Improve Generalization of Hyperparameter Optimization
- UQ-Guided Hyperparameter Optimization for Iterative Learners
- MALIBO: Meta-learning for Likelihood-free Bayesian Optimization
- A New Linear Scaling Rule for Private Adaptive Hyperparameter Optimization
- In-Context Freeze-Thaw Bayesian Optimization for Hyperparameter Optimization
- Depthwise Hyperparameter Transfer in Residual Networks: Dynamics and Scaling Limit
- Principled Architecture-aware Scaling of Hyperparameters
- Efficient Hyperparameter Optimization with Adaptive Fidelity Identification
- Interactive Hyperparameter Optimization in Multi-Objective Problems via Preference Learning
- BTTackler: A Diagnosis-based Framework for Efficient Deep Learning Hyperparameter Optimization
- DP-HyPO: An Adaptive Private Framework for Hyperparameter Optimization
- Efficient Bayesian Learning Curve Extrapolation using Prior-Data Fitted Networks
- Efficient Hyper-parameter Optimization with Cubic Regularization
- End-to-End Meta-Bayesian Optimisation with Transformer Neural Processes
- New Bounds for Hyperparameter Tuning of Regression Problems Across Instances
- Practical Differentially Private Hyperparameter Tuning with Subsampling
- PriorBand: Practical Hyperparameter Optimization in the Age of Deep Learning
- Scaling Laws for Hyperparameter Optimization
- “Why Not Looking backward?” A Robust Two-Step Method to Automatically Terminate Bayesian Optimization
- FedHPO-Bench: A Benchmark Suite for Federated Hyperparameter Optimization
- Hyperparameters in Reinforcement Learning and How To Tune Them
- Optimizing Hyperparameters with Conformal Quantile Regression
- PFNs4BO: In-Context Learning for Bayesian Optimization
- EA-HAS-Bench: Energy-aware Hyperparameter and Architecture Search Benchmark
- Single-shot General Hyper-parameter Optimization for Federated Learning
- Targeted Hyperparameter Optimization with Lexicographic Preferences Over Multiple Objectives
- Deep Ranking Ensembles for Hyperparameter Optimization
- Gray-Box Gaussian Processes for Automated Reinforcement Learning
- Hyperparameter Optimization through Neural Network Partitioning
- PASHA: Efficient HPO and NAS with Progressive Resource Allocation
- LiDAR-in-the-Loop Hyperparameter Optimization
- Bayesian Optimization Meets Self-Distillation
- Code-Aware Cross-Program Transfer Hyperparameter Optimization
- HyperJump: Accelerating HyperBand via Risk Modelling
- Online Hyperparameter Optimization for Class-Incremental Learning
- Multi-armed bandits for resource efficient, online optimization of language model pre-training: the use case of dynamic masking
- On the Hyperparameter Loss Landscapes of Machine Learning Models: An Exploratory Study
- Self-Tuning Self-Supervised Image Anomaly Detection
- Supervising the Multi-Fidelity Race of Hyperparameter Configurations
- AUTOMATA: Gradient Based Data Subset Selection for Compute-Efficient Hyper-parameter Tuning
- Hyperparameter Sensitivity in Deep Outlier Detection: Analysis and a Scalable Hyper-Ensemble Solution
- Towards Learning Universal Hyperparameter Optimizers with Transformers
- Value Function based Difference-of-Convex Algorithm for Bilevel Hyperparameter Selection Problems
- Hyperparameter Tuning with Renyi Differential Privacy
- Online Hyperparameter Meta-Learning with Hypergradient Distillation
- Scalable One-Pass Optimisation of High-Dimensional Weight-Update Hyperparameters by Implicit Differentiation
- $\pi$BO: Augmenting Acquisition Functions with User Beliefs for Bayesian Optimization
- AME: Attention and Memory Enhancement in Hyper-Parameter Optimization
- The Role of Adaptive Optimizers for Honest Private Hyperparameter Selection
- Efficient Hyper-parameter Search for Knowledge Graph Embedding
- Demystify Hyperparameters for Stochastic Optimization with Transferable Representations
- Hierarchical Proxy Modeling for Improved HPO in Time Series Forecasting
- TransBO: Hyperparameter Optimization via Two-Phase Transfer Learning
- Transfer Learning based Search Space Design for Hyperparameter Tuning
D3 AutoML systems, CASH, algorithm selection & configuration, model selection 66
- Automatic Unsupervised Ensemble Outlier Model Selection
- CUPID in the Model Zoo: Online Matchmaking for Selecting Your Dream LLM
- DMCO: Budget-Aware Co-Optimization of Data Cleaning and AutoML
- Landmark-Guided Policy Optimization for Multi-Objective Language Model Selection
- Optimal Pricing for Data-Augmented AutoML Marketplaces
- Mordal: Automated Pretrained Model Selection for Vision Language Models
- Relatron: Automating Relational Machine Learning over Relational Databases
- Rethinking Model Selection in VLM Through the Lens of Gromov-Wasserstein Distance
- HAMLET4Fairness: Enhancing Fairness in AI Pipelines Through Human-Centered AutoML and Argumentation
- Neural Architecture and Hyperparameter Selection Through Meta-Learning on Time Series
- Practical, Utilitarian Algorithm Configuration
- Generalization or Memorization? Multi-Agent vs. Baseline LLMs and AutoML Models for Tabular Classification
- Test-Time Search for Automated GFM Fine-Tuning
- Generalization Bounds for Model-based Algorithm Configuration
- Implicit Modeling for Transferability Estimation of Vision Foundation Models
- Put CASH on Bandits: A Max K-Armed Problem for Automated Machine Learning
- Sequential Multi-Agent Dynamic Algorithm Configuration
- Unified Transferability Metrics for Time Series Foundation Models
- Towards Robustness and Explainability of Automatic Algorithm Selection
- Graph-Supported Dynamic Algorithm Configuration for Multi-Objective Combinatorial Optimization
- OOD-Chameleon: Is Algorithm Selection for OOD Generalization Learnable?
- Reliable Algorithm Selection for Machine Learning-Guided Design
- Vision-Language Model Selection and Reuse for Downstream Adaptation
- X-Hacking: The Threat of Misguided AutoML
- MetaOOD: Automatic Selection of OOD Detection Models
- Risk-Controlling Model Selection via Guided Bayesian Optimization
- Utilitarian Algorithm Configuration for Infinite Parameter Spaces
- Consensus-Driven Active Model Selection
- ADELA: Accelerating Evolutionary Design of Machine Learning Pipelines with the Accompanying Surrogate Model
- ConfigX: Modular Configuration for Evolutionary Algorithms via Multitask Reinforcement Learning
- Bayesian Optimization for Simultaneous Selection of Machine Learning Algorithms and Hyperparameters on Shared Latent Space
- HyperZero: A Customized End-to-End Auto-Tuning System for Recommendation with Hourly Feedback
- Position: A Call to Action for a Human-Centered AutoML Paradigm
- Quick-Tune: Quickly Learning Which Pretrained Model to Finetune and How
- LEAD: Exploring Logit Space Evolution for Model Selection
- Towards Reproducible, Automated, and Scalable Anomaly Detection
- CASH via Optimal Diversity for Ensemble Learning
- Algorithm Selection for Deep Active Learning with Imbalanced Datasets
- GLEMOS: Benchmark for Instantaneous Graph Learning Model Selection
- How to Select Which Active Learning Strategy is Best Suited for Your Specific Problem and Budget
- LOVM: Language-Only Vision Model Selection
- Utilitarian Algorithm Configuration
- AANG : Automating Auxiliary Learning
- Unsupervised Model Selection for Time Series Anomaly Detection
- AutoTransfer: AutoML with Knowledge Transfer - An Application to Graph Neural Networks
- MetaGL: Evaluation-Free Selection of Graph Learning Models via Meta-Learning
- The Dark Side of AutoML: Towards Architectural Backdoor Search
- Multi-Agent Automated Machine Learning
- AC-Band: A Combinatorial Bandit-Based Approach to Algorithm Configuration
- AutoSTL: Automated Spatio-Temporal Multi-Task Learning
- AutoXPCR: Automated Multi-Objective Model Selection for Time Series Forecasting
- Deep Pipeline Embeddings for AutoML
- Fast Unsupervised Deep Outlier Model Selection with Hypernetworks
- Test Accuracy vs. Generalization Gap: Model Selection in NLP without Accessing Training or Testing Data
- Multi-agent Dynamic Algorithm Configuration
- AutoML Two-Sample Test
- DivBO: Diversity-aware CASH for Ensemble Learning
- Efficient Non-Parametric Optimizer Search for Diverse Tasks
- HyperImpute: Generalized Iterative Imputation with Automatic Model Selection
- Zero-shot AutoML with Pretrained Models
- Learning meta-features for AutoML
- AutoLoss-GMS: Searching Generalized Margin-Based Softmax Loss Function for Person Re-Identification
- AutoLoss-Zero: Searching Loss Functions From Scratch for Generic Tasks
- Privacy-Preserving Online AutoML for Domain-Specific Face Detection
- Bandit Limited Discrepancy Search and Application to Machine Learning Pipeline Optimization
- Machine Learning for Online Algorithm Selection under Censored Feedback
D4 Automated feature engineering & data augmentation 27
- AutoDA-Timeseries: Automated Data Augmentation for Time Series
- Human-LLM Collaborative Feature Engineering for Tabular Data
- Optimizing Data Augmentation through Bayesian Model Selection
- Memory-Augmented LLM-based Multi-Agent System for Automated Feature Generation on Tabular Data
- DAGPipe: Differentiable DAG Learning for Automated Data Preparation
- UCB-Based Feature Engineering for Cold-Start in Recommenders
- Boost the Performance of Tabular Data Models with GPU Accelerated Feature Engineering
- Continuous Optimization for Feature Selection with Permutation-Invariant Embedding and Policy-Guided Search
- Heterogeneous Multi-Agent Reinforcement Learning with Attention for Cooperative and Scalable Feature Transformation
- Optimized Feature Generation for Tabular Data via LLMs with Decision Tree Reasoning
- Large Language Models Can Automatically Engineer Features for Few-Shot Tabular Learning
- Your Image is My Video: Reshaping the Receptive Field via Image-To-Video Differentiable AutoAugmentation and Fusion
- ERASE: Benchmarking Feature Selection Methods for Deep Recommender Systems
- Unsupervised Generative Feature Transformation via Graph Contrastive Pre-training and Multi-objective Fine-tuning
- A Performance-Driven Benchmark for Feature Selection in Tabular Deep Learning
- OpenFE: Automated Feature Generation with Expert-level Performance
- Learning a Data-Driven Policy Network for Pre-Training Automated Feature Engineering
- Learning Fair Graph Representations via Automated Data Augmentations
- Automated Data Augmentations for Graph Classification
- SLACK: Stable Learning of Augmentations With Cold-Start and KL Regularization
- Cognitive Evolutionary Search to Select Feature Interactions for Click-Through Rate Prediction
- Adversarial Auto-Augment with Label Preservation: A Representation Learning Principle Guided Approach
- Deep AutoAugment
- AIM: An Auto-Augmenter for Images and Meshes
- AutoGCL: Automated Graph Contrastive Learning via Learnable View Generators
- AdaFS: Adaptive Feature Selection in Deep Recommender System
- Group-wise Reinforcement Feature Generation for Optimal and Explainable Representation Space Reconstruction
D5 Prior-data fitted networks & tabular foundation models 39
- End-to-End Compression for Tabular Foundation Models
- FIRE: Multi-fidelity Regression with Distribution-conditioned In-context Learning using Tabular Foundation Models
- Frequentist Consistency of Prior-Data Fitted Networks for Causal Inference
- GOTabPFN: From Feature Ordering to Compact Tokenization for Tabular Foundation Models on High-Dimensional Data
- GraphPFN: A Prior-Data Fitted Graph Foundation Model
- Is One Layer Enough? Understanding Inference Dynamics in Tabular Foundation Models
- LimiX-2M: Mitigating Low-Rank Collapse and Attention Bottlenecks in Tabular Foundation Models
- Mitigating Label Shift in Tabular In-Context Learning via Test-Time Posterior Adjustment
- SwiftPFN: Revisiting Row-Wise Attention–Only Tabular Foundation Models with Adaptive Early Exit
- TabICLv2: A Better, Faster, Scalable, and Open Tabular Foundation Model
- TabMGP: Martingale Posterior with TabPFN
- Unveiling Prior-Data Fitted Networks on Causal Effect Estimation: Pre-Training or Fine-Tuning?
- ChunkTabPFN: Training-free Long Context
- Foundation Models for Causal Inference via Prior-Data Fitted Networks
- MultiModalPFN: Extending Prior-Data Fitted Networks for Multimodal Tabular Learning
- ConTextTab: A Semantics-Aware Tabular In-Context Learner
- TabArena: A Living Benchmark for Machine Learning on Tabular Data
- A Closer Look at TabPFN v2: Understanding Its Strengths and Extending Its Capabilities
- Effortless, Simulation-Efficient Bayesian Inference using Tabular Foundation Models
- EquiTabPFN: A Target-Permutation Equivariant Prior Fitted Network
- Mitra: Mixed Synthetic Priors for Enhancing Tabular Foundation Models
- TabDPT: Scaling Tabular Foundation Models on Real Data
- TabSTAR: A Tabular Foundation Model for Tabular Data with Text Fields
- Bayesian Neural Scaling Law Extrapolation with Prior-Data Fitted Networks
- FairPFN: A Tabular Foundation Model for Causal Fairness
- TabICL: A Tabular Foundation Model for In-Context Learning on Large Data
- TabPFN Unleashed: A Scalable and Effective Solution to Tabular Classification Problems
- Zero-shot Meta-learning for Tabular Prediction Tasks with Adversarially Pre-trained Transformer
- Mixture of In-Context Prompters for Tabular PFNs
- MotherNet: Fast Training and Inference via Hyper-Network Transformers
- Drift-Resilient TabPFN: In-Context Learning Temporal Distribution Shifts on Tabular Data
- Retrieval & Fine-Tuning for In-Context Tabular Models
- TuneTables: Context Optimization for Scalable Prior-Data Fitted Networks
- Position: Why Tabular Foundation Models Should Be a Research Priority
- HyperFast: Instant Classification for Tabular Data
- When Do Neural Nets Outperform Boosted Trees on Tabular Data?
- Statistical Foundations of Prior-Data Fitted Networks
- TabPFN: A Transformer That Solves Small Tabular Classification Problems in a Second
- Transformers Can Do Bayesian Inference
D6 LLM-driven AutoML & automated algorithm discovery 54
- $A_2$DEPT: Large Language Model–Driven Automated Algorithm Design via Evolutionary Program Trees
- A Language-Guided Bayesian Optimization for Efficient LoRA Hyperparameter Search
- LABO: LLM-Accelerated Bayesian Optimization through Broad Exploration and Selective Experimentation
- LILO: Bayesian Optimization with Natural Language Feedback
- PathWise: Planning through World Model for Automated Heuristic Design via Self-Evolving LLMs
- SAGE-NAS: Synergizing LLM-Based Semantic Agent with Graph-Based Evaluator for Neural Architecture Search
- Structured Progressive Knowledge Activation for LLM-Driven Neural Architecture Search
- AutoEP: LLMs-Driven Automation of Hyperparameter Evolution for Metaheuristic Algorithms
- Adaptive Acquisition Selection for Bayesian Optimization with Large Language Models
- CALM: Co-evolution of Algorithms and Language Model for Automatic Heuristic Design
- HiFo-Prompt: Prompting with Hindsight and Foresight for LLM-based Automatic Heuristic Design
- LLM-Guided Evolutionary Program Synthesis for Quasi-Monte Carlo Design
- Rethinking Code Similarity for Automated Algorithm Design with LLMs
- Scaling Multi-Task Bayesian Optimization with Large Language Models
- ShinkaEvolve: Towards Open-Ended and Sample-Efficient Program Evolution
- AutoTuneX: Interactive Automated Fine-Tuning for Large Language Models
- CO-Bench: Benchmarking Language Model Agents in Algorithm Search for Combinatorial Optimization
- CoEvo: Continual Evolution of Symbolic Solutions Using Large Language Models
- EoH-S: Evolution of Heuristic Set Using LLMs for Automated Heuristic Design
- TrajEvo: Trajectory Prediction Heuristics Design via LLM-driven Evolution
- Experience-Driven Reflective Co-Evolution of Prompts and Heuristics for Autonomous Algorithm Design
- What Makes an LLM a Good Optimizer? A Trajectory Analysis of LLM-Guided Evolutionary Search
- CoFE: Collaborative Feature Engineering via Semantically-Guided Exploration and Diagnostic-Driven Refinement
- CoFEH: LLM-driven Feature Engineering Empowered by Collaborative Bayesian Hyperparameter Optimization
- MORE-FE: Multi-Operator and Reinforcement Learning-Enhanced Evolution for LLM Feature Engineering
- Adaptive Kernel Design for Bayesian Optimization Is a Piece of CAKE with LLMs
- Automated Model Discovery via Multi-modal & Multi-step Pipeline
- DesignX: Human-Competitive Algorithm Designer for Black-Box Optimization
- Partition to Evolve: Niching-enhanced Evolution with LLMs for Automated Algorithm Discovery
- Revolutionizing Training-Free NAS: Towards Efficient Automatic Proxy Discovery via Large Language Models
- FunBO: Discovering Acquisition Functions for Bayesian Optimization with FunSearch
- Hyperband-based Bayesian Optimization for Black-box Prompt Selection
- Monte Carlo Tree Search for Comprehensive Exploration in LLM-Based Automatic Heuristic Design
- RZ-NAS: Enhancing LLM-guided Neural Architecture Search via Reflective Zero-Cost Strategy
- Decision Tree Induction Through LLMs via Semantically-Aware Evolution
- Searching for Optimal Solutions with LLMs via Bayesian Optimization
- NADER: Neural Architecture Design via Multi-Agent Collaboration
- Multimodal Large Language Model-Guided ISP Hyperparameter Optimization with Dynamic Preference Learning
- AutoMMLab: Automatically Generating Deployable Models from Language Instructions for Computer Vision Tasks
- Design Principle Transfer in Neural Architecture Search via Large Language Models
- Evolutionary Large Language Model for Automated Feature Transformation
- HSEvo: Elevating Automatic Heuristic Design with Diversity-Driven Harmony Search and Genetic Algorithm Using LLMs
- Large Language Models Enhanced Personalized Graph Neural Architecture Search in Federated Learning
- Multi-Objective Evolution of Heuristic Using Large Language Model
- Co-Evolution of Large Language Models and Configuration Strategies to Enhance Surrogate-Assisted Evolutionary Algorithm
- Efficient Heuristics Generation for Solving Combinatorial Optimization Problems Using Large Language Models
- ReEvo: Large Language Models as Hyper-Heuristics with Reflective Evolution
- Evolution of Heuristics: Towards Efficient Automatic Algorithm Design Using Large Language Model
- Large Language Models to Enhance Bayesian Optimization
- LLM can Achieve Self-Regulation via Hyperparameter Aware Generation
- LLM Performance Predictors are good initializers for Architecture Search
- 'Oh LLM, I'm Asking Thee, Please Give Me a Decision Tree': Zero-Shot Decision Tree Induction and Embedding with Large Language Models
- Large Language Model-driven Meta-structure Discovery in Heterogeneous Information Network
- EvoPrompting: Language Models for Code-Level Neural Architecture Search
D7 AutoML for LLM systems: agent, workflow & pipeline search 28
- GEPA: Reflective Prompt Evolution Can Outperform Reinforcement Learning
- RAAS: LLM Agentic System Architecture Search with GRPO
- AgentSwift: Efficient LLM Agent Design via Value-Guided Hierarchical Search
- A²Flow: Automating Agentic Workflow Generation via Self-Adaptive Abstraction Operators
- DAWN: Distributed LLM Multi-Agent Workflow Synthesis
- HiveMind: Contribution-Guided Online Prompt Optimization of LLM Multi-Agent Systems
- PASS: Probabilistic Agentic Supernet Sampling for Interpretable and Adaptive Chest X-Ray Reasoning
- FusionFlow: Enabling Deep Structural Exploration for Automated Agentic Workflow Generation
- Grammar Search for Multi-Agent Systems
- Hetero-Designer: Automated Design of Multi-Agent Systems with Heterogeneous LLMs
- Attribution-Based Analysis and Optimization of Modular Agentic Workflows
- Do We Always Need Query-Level Workflows? Rethinking Agentic Workflow Generation for Multi-Agent Systems
- Evolving Agentic Workflow Driven by Human-Agent Collaboration
- MedDCR: Learning to Design Agentic Workflows for Medical Coding
- SCOPE: Cost-Efficient Model Selection for Compound AI Systems under Quality Constraints
- Multi-agent Architecture Search via Agentic Supernet
- An Architecture Search Framework for Inference-Time Techniques
- MAS-GPT: Training LLMs to Build LLM-based Multi-Agent Systems
- AFlow: Automating Agentic Workflow Generation
- AgentSquare: Automatic LLM Agent Search in Modular Design Space
- Automated Design of Agentic Systems
- Benchmarking Agentic Workflow Generation
- Flow: Modularized Agentic Workflow Automation
- PromptWizard: Optimizing Prompts via Task-Aware, Feedback-Driven Self-Evolution
- Cognify: Supercharging Gen-AI Workflows With Hierarchical Autotuning
- Trace is the Next AutoDiff: Generative Optimization with Rich Feedback, Execution Traces, and LLMs
- GPTSwarm: Language Agents as Optimizable Graphs
- DSPy: Compiling Declarative Language Model Calls into State-of-the-Art Pipelines
E AI for Science 315
Scoped to LLM/agent-driven science: agents for scientific domains and lab automation, scientific reasoning benchmarks, LLM-driven equation/law/program discovery, and a selective set of scientific LLMs and foundation models. The long tail of domain-specific modelling papers (individual property predictors, PDE solvers, molecule generators) is deliberately excluded.
E1 LLM / agent systems for scientific domains & lab automation 92
- AutoMat: Physics-Guided Agentic Reasoning for Solving Ill-Posed Inverse Microscopy Problems
- AutoMS: Multi-Agent Evolutionary Search for Cross-Physics Inverse Microstructure Design
- Grounding LLMs in Scientific Discovery via Embodied Actions
- HVR-Met: A Hypothesis-Verification-Replanning Agentic System for Extreme Weather Diagnosis
- LatentChem: From Textual CoT to Latent Thinking in Chemical Reasoning
- PDAgent: An LLM-Driven Autonomous Agent Framework Towards *In Silico* Protein Design via Directed Mutation
- Protein Design with Agent Rosetta: A Case Study for Specialized Scientific Agents
- Retro-Expert: Collaborative Reasoning for Interpretable Retrosynthesis
- RetrOrchestrator: A Multi-Step Retrosynthesis Agent Dynamically Orchestrating Single-Step Transition Models
- Sim2Reason: Solving Physics Olympiad via Reinforcement Learning on Physics Simulators
- SimulCost: A Cost-Aware Benchmark and Toolkit for Automating Physics Simulations with LLMs
- SP-Mind: An Autonomous Reasoning Agent for Spatial Proteomics Analysis
- Steering Large Language Models through the DMTA Cycle: Structure-Based Drug Design via Knowledge-Driven Bi-Level Thompson Sampling
- AI-for-Science Low-code Platform with Bayesian Adversarial Multi-Agent Framework
- AutoBio: A Simulation and Benchmark for Robotic Automation in Digital Biology Laboratory
- BED-LLM: Intelligent Information Gathering with LLMs and Bayesian Experimental Design
- CellAgent: LLM-Driven Multi-Agent Framework for Natural Language-Based Single-Cell Analysis
- CP-Agent: Context‑Aware Multimodal Reasoning for Cellular Morphological Profiling under Chemical Perturbations
- Earth-Agent: Unlocking the Full Landscape of Earth Observation with Agents
- Eigen-Agent: Adaptive Multi-Agent Scientific Reasoning with Monitor-Based RAG
- From What to Why: A Multi-Agent System for Evidence-based Chemical Reaction Condition Reasoning
- LLEMA: Evolutionary Search with LLMs for Multi-Objective Materials Discovery
- MAC-AMP: A Closed-Loop Multi-Agent Collaboration System for Multi-Objective Antimicrobial Peptide Design
- Reference-guided Policy Optimization for Molecular Optimization via LLM Reasoning
- ScienceBoard: Evaluating Multimodal Autonomous Agents in Realistic Scientific Workflows
- SciNav: A General Agent Framework for Scientific Coding Tasks
- Towards Knowledge‑and‑Data‑Driven Organic Reaction Prediction: RAG‑Enhanced and Reasoning‑Powered Hybrid System with LLMs
- Unleashing Scientific Reasoning for Bio-experimental Protocol Generation via Structured Component-based Reward Mechanism
- Zephyrus: An Agentic Framework for Weather Science
- Assessing LLMs for Serendipity Discovery in Knowledge Graphs: A Case for Drug Repurposing
- Expert-Inspired Multi-Agent Coordination for Multi-Objective Molecular Optimization
- From Text to Simulation: A Multi-Agent LLM Workflow for Automated Chemical Process Design
- RAG-Enhanced Collaborative LLM Agents for Drug Discovery
- "Excuse me, may I say something..." CoLabScience, A Proactive AI Assistant for Biomedical Discovery and LLM-Expert Collaborations
- BioProAgent: Neuro-Symbolic Grounding for Constrained Scientific Planning
- BioTool: A Comprehensive Tool-Calling Dataset for Enhancing Biomedical Capabilities of Large Language Models
- FormalScience: Scalable Human-in-the-Loop Autoformalisation of Science with Agentic Code Generation in Lean
- Interleaved Tool-Call Reasoning for Protein Function Understanding
- LLM4Cell: Taxonomy and Evaluation of LLM and Agentic Models for Single-Cell Biology
- MolMem: Memory-Augmented Agentic Reinforcement Learning for Sample-Efficient Molecular Optimization
- Spatial-Agent: Agentic Geo-spatial Reasoning with Scientific Core Concepts
- ChemAmp: Amplified Chemistry Tools via Composable Agents
- CheMM-R1: Enhancing Chemical Structure Recognition and Elucidation with Reasoning Multimodal Large Language Models
- ClimAgent: LLM as Agents for Autonomous Open-ended Climate Science Analysis
- Feedback to Reasoning: LLM-Assisted Molecular Optimization with Domain Feedback and Historical Reasoning
- MotifAgent: Learning Molecular Assembly through Multi-Agent Collaboration for Chemical Language Understanding
- ProtoCycle: Reflective Tool-Augmented Planning for Text-Guided Protein Design
- SocraticChem: Physics-Grounded Socratic Inquiry for Safety-Critical Experimental Science
- ARIA: A Causal-Aware Framework for Rescuing LLM Reasoning in Trustworthy Materials Discovery
- Battery-Sim-Agent: Leveraging LLM-Agent for Inverse Battery Parameter Estimation
- CEAgent-GSL: Code-level Evolutionary Agent for Interpretable Graph Structure Learning on Omics
- Compass: Navigating Global Marine Lead Data Integration through Expert-Guided LLM Agent
- Enhancing Spatial Reasoning in Large Language Models for Metal-Organic Frameworks Structure Prediction
- Molecular Lead Optimization via Agentic Tool Planning
- PeroMAS: A Multi-agent System of Perovskite Material Discovery
- SpaCellAgent: A Self-Evolving LLM-Based Multi-Agent Framework for Trajectory Analysis
- VCAgent: A Mutation-Guided Self-Reflective Agent Framework for Virtual Cell Modeling
- CIDD: Collaborative Intelligence for Structure-Based Drug Design Empowered by LLMs
- LabUtopia: High-Fidelity Simulation and Hierarchical Benchmark for Scientific Embodied Agents
- LLM Meets Diffusion: A Hybrid Framework for Crystal Material Generation
- Retro-R1: LLM-based Agentic Retrosynthesis
- scPilot: Large Language Model Reasoning Toward Automated Single-Cell Analysis and Discovery
- Adapting While Learning: Grounding LLMs for Scientific Problems with Tool Usage Adaptation
- Code-Generated Graph Representations Using Multiple LLM Agents for Material Properties Prediction
- LLM-Augmented Chemical Synthesis and Design Decision Programs
- PINNsAgent: Automated PDE Surrogation with Large Language Models
- OSDA Agent: Leveraging Large Language Models for De Novo Design of Organic Structure Directing Agents
- BioDiscoveryAgent: An AI Agent for Designing Genetic Perturbation Experiments
- ChemAgent: Self-updating Memories in Large Language Models Improves Chemical Reasoning
- Efficient Evolutionary Search Over Chemical Space with Large Language Models
- Hierarchically Encapsulated Representation for Protocol Design in Self-Driving Labs
- LICO: Large Language Models for In-Context Molecular Optimization
- MatExpert: Decomposing Materials Discovery By Mimicking Human Experts
- Multimodal Large Language Models for Inverse Molecular Design with Retrosynthetic Planning
- Physics Context Builders: A Modular Framework for Physical Reasoning in Vision-Language Models
- Adaptive Experimental Design to Accelerate Scientific Discovery and Engineering Design
- AutoSciLab: A Self-Driving Laboratory for Interpretable Scientific Discovery
- ReactGPT: Understanding of Chemical Reactions via In-Context Tuning
- Boosting LLM’s Molecular Structure Elucidation with Knowledge Enhanced Tree Search Reasoning
- From Awareness to Adaptability: Enhancing Tool Utilization for Scientific Reasoning
- RL-Guider: Leveraging Historical Decisions and Feedback for Drug Editing with Large Language Models
- An Instructible Chemist-AI Alignment Framework for Generating Quaternary Ammonium Compound Structures
- Construction and Application of Materials Knowledge Graph in Multidisciplinary Materials Science via Large Language Model
- Expert-level protocol translation for self-driving labs
- A Sober Look at LLMs for Material Discovery: Are They Actually Good for Bayesian Optimization Over Molecules?
- CHEMREASONER: Heuristic Search over a Large Language Model’s Knowledge Space using Quantum-Chemical Feedback
- Entropy-Reinforced Planning with Large Language Models for Drug Discovery
- Knowledge-aware Reinforced Language Models for Protein Directed Evolution
- Fine-Tuned Language Models Generate Stable Inorganic Materials as Text
- Generating Novel Leads for Drug Discovery Using LLMs with Logical Feedback
- FoodPuzzle: Toward Developing Large Language Model Agents as Autonomous Flavor Scientists
- De novo Drug Design using Reinforcement Learning with Multiple GPT Agents
E2 Scientific reasoning & knowledge benchmarks 119
- ABC-Bench: An Agentic Bio-Capabilities Benchmark for Biosecurity
- BioAgent Bench: An AI Agent Evaluation Suite for Bioinformatics
- BioProBench: A Corpus and Benchmark for Biological Protocol Reasoning in Autonomous Science
- Demystifying Scientific Problem-Solving in LLMs by Probing Knowledge and Reasoning
- HiPhO: How Far Are (M)LLMs from Humans in the Latest High School Physics Olympiad Benchmark?
- IDRBench: Understanding the Capability of Large Language Models on Interdisciplinary Research
- InteractScience: Programmatic and Visually-Grounded Evaluation of Interactive Scientific Demonstration Code Generation
- Judging What We Cannot Solve: A Consequence-Based Approach for Oracle-Free Evaluation of Research-Level Math
- MADE: Benchmark Environments for Closed-Loop Materials Discovery
- MMClima: A Framework for Multimodal Climate Science Data and Evaluation
- SciAgentGym: Benchmarking Multi-Step Scientific Tool-Use in LLM Agents
- SciPredict: Can LLMs Predict the Outcomes of Scientific Experiments in Natural Sciences?
- SciVideoBench: Benchmarking Scientific Video Reasoning in Large Multimodal Models
- TadA-Bench: A Million-Variant Benchmark for Future-Round Discovery Toward Agentic Protein Engineering
- When Single Answer Is Not Enough: Rethinking Single-Step Retrosynthesis Benchmarks for LLMs
- XDomainBench: Diagnosing Reasoning Collapse in High-Dimensional Scientific Knowledge Composition
- CatalystBench: A Comprehensive Multi-Task Benchmark for Advancing Language Models in Catalysis Science
- ChemEval: A Multi-level and Fine-grained Chemical Capability Evaluation for Large Language Models
- CMPhysBench: A Benchmark for Evaluating Large Language Models in Condensed Matter Physics
- CMT-Benchmark: A Benchmark for Condensed Matter Theory Built by Expert Researchers
- EarthSE: A Benchmark Evaluating Earth Scientific Exploration Capability for Large Language Models
- ExpVid: A Benchmark for Experiment Video Understanding & Reasoning
- MolecularIQ: Characterizing Chemical Reasoning Capabilities Through Symbolic Verification on Molecular Graphs
- PRISM-Physics: Causal DAG-Based Process Evaluation for Physics Reasoning
- SC-Arena: A Natural Language Benchmark for Single-Cell Reasoning with Knowledge-Augmented Evaluation
- SCI-Verifier: Scientific Verifier with Thinking
- SciTS: Scientific Time Series Understanding and Generation with LLMs
- SpectraLLM: Uncovering the Ability of LLMs for Molecule Structure Elucidation from Multi-Spectra
- GeoMMBench and GeoMMAgent: Toward Expert-Level Multimodal Intelligence in Geoscience and Remote Sensing
- QUANTIPHY: A Quantitative Benchmark Evaluating Physical Reasoning Abilities of Vision-Language Models
- SCIEval: Evaluating and Benchmarking the Faithfulness of Scientific Image Generation and Interpretation with Large Multimodal Models
- DeepPhy: Benchmarking Agentic VLMs on Physical Reasoning
- MDK12-Bench: A Multi-Discipline Benchmark for Evaluating Reasoning in Multimodal Large Language Models
- MME-SCI: A Comprehensive and Challenging Science Benchmark for Multimodal Large Language Models
- ChemReason-Bench: Benchmarking Large Language Models for Procedural Reasoning in Experimental Chemistry
- Decoding Scientific Experimental Images: The SPUR Benchmark for Perception, Understanding, and Reasoning
- GenomeQA: Benchmarking General Large Language Models for Genome Sequence Understanding
- MMSciCode: Real-world Evaluation of Multilingual Multi-Discipline Scientific Research Coding
- MSEarth: A Multimodal Benchmark for Earth Science Phenomenon Discovery with MLLMs
- QuantumQA: Enhancing Scientific Reasoning via Physics-Consistent Dataset and Verification-Aware Reinforcement Learning
- SciCustom: A Framework for Custom Evaluation of Scientific Capabilities in Large Language Models
- ChemVLR: Prioritizing Reasoning in Perception for Chemical Vision-Language Understanding
- K-MetBench: A Multi-Dimensional Benchmark for Fine-Grained Evaluation of Expert Reasoning, Locality, and Multimodality in Meteorology
- LLMs as Lab Engineers: A Benchmark for Analytical Method Lifecycle Management
- PhageBench: Can LLMs Understand Raw Bacteriophage Genomes?
- Position: Multimodal Large Language Models Can Significantly Advance Scientific Reasoning
- SciVQR: A Multidisciplinary Multimodal Benchmark for Advanced Scientific Reasoning Evaluation
- Seeing Beyond Words: MatVQA for Challenging Visual-Scientific Reasoning in Materials Science
- ToxReason: A Benchmark for Mechanistic Chemical Toxicity Reasoning via Adverse Outcome Pathway
- WildSci: Advancing Scientific Reasoning from In-the-Wild Literature
- Benchmarking Multimodal LLMs on Recognition and Understanding over Chemical Tables
- Can Large Language Models Derive New Knowledge? A Dynamic Benchmark for Biological Knowledge Discovery
- MatSciBench: Benchmarking the Reasoning Ability of Large Language Models in Materials Science
- NMRGym: A Comprehensive Benchmark for Nuclear Magnetic Resonance Based Molecular Structure Elucidation
- OmniScience: A Large-scale Multi-modal Dataset for Scientific Image Understanding
- Sci-PRM: A Tool Aware Process Reward Model for Scientific Reasoning Verification
- SciChart: Visual Question Answering and Reasoning for Scientific Spectral Chart
- SciHorizon-DataEVA: An Agentic System to Scalable AI-Readiness Evaluation of Heterogeneous Scientific Data
- SciHorizon-Gene: Benchmarking LLM for Life Sciences Inference from Gene Knowledge to Functional Understanding
- Speak-to-Structure: Evaluating LLMs in Open-domain Natural Language-Driven Molecule Generation
- Who's Adam? Benchmarking Hallucinations in Scientific Dialogue
- ConnectomeBench: Can LLMs proofread the connectome?
- AstroVisBench: A Code Benchmark for Scientific Computing and Visualization in Astronomy
- AtmosSci-Bench: Evaluating the Recent Advance of Large Language Model for Atmospheric Science
- Beyond Chemical QA: Evaluating LLM's Chemical Reasoning with Modular Chemical Operations
- CellVerse: Do Large Language Models Really Understand Cell Biology?
- CGBench: Benchmarking Language Model Scientific Reasoning for Clinical Genetics Research
- Common Task Framework For a Critical Evaluation of Scientific Machine Learning Algorithms
- Mars-Bench: A Benchmark for Evaluating Foundation Models for Mars Science Tasks
- Measuring Scientific Capabilities of Language Models with a Systems Biology Dry Lab
- PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models
- RealMath: A Continuous Benchmark for Evaluating Language Models on Research-Level Mathematics
- Scaling Physical Reasoning with the PHYSICS Dataset
- Scientists' First Exam: Probing Cognitive Abilities of MLLM via Perception, Understanding, and Reasoning
- SeePhys: Does Seeing Help Thinking? – Benchmarking Vision-Based Physics Reasoning
- SuperGPQA: Scaling LLM Evaluation across 285 Graduate Disciplines
- Machine Learning meets Algebraic Combinatorics: A Suite of Datasets Capturing Research-level Conjecturing Ability in Pure Mathematics
- Generalists vs. Specialists: Evaluating LLMs on Highly-Constrained Biophysical Sequence Optimization Tasks
- RBench: Graduate-level Multi-disciplinary Benchmarks for LLM & MLLM Complex Reasoning Evaluation
- UGPhysics: A Comprehensive Benchmark for Undergraduate Physics Reasoning with Large Language Models
- ClimaQA: An Automated Evaluation Framework for Climate Question Answering Models
- Contextualizing biological perturbation experiments through language
- CURIE: Evaluating LLMs on Multitask Scientific Long-Context Understanding and Reasoning
- HARDMath: A Benchmark Dataset for Challenging Problems in Applied Mathematics
- LiveXiv - A Multi-Modal live benchmark based on Arxiv papers content
- MAPS: Advancing Multi-Modal Reasoning in Expert-Level Physical Science
- MicroVQA: A Multimodal Reasoning Benchmark for Microscopy-Based Scientific Research
- MMVU: Measuring Expert-Level Multi-Discipline Video Understanding
- Science-T2I: Addressing Scientific Illusions in Image Synthesis
- ProJudge: A Multi-Modal Multi-Discipline Benchmark and Instruction-Tuning Dataset for MLLM-based Process Judges
- SciVid: Cross-Domain Evaluation of Video Models in Scientific Applications
- PhysReason: A Comprehensive Benchmark towards Physics-Based Reasoning
- SeedBench: A Multi-task Benchmark for Evaluating Large Language Models in Seed Science
- YESciEval: Robust LLM-as-a-Judge for Scientific Question Answering
- MMSciBench: Benchmarking Language Models on Chinese Multimodal Scientific Problems
- Physics: Benchmarking Foundation Models on University-Level Physics Problem Solving
- SciVerse: Unveiling the Knowledge Comprehension and Visual Reasoning of LMMs on Multi-modal Scientific Problems
- SciHorizon: Benchmarking AI-for-Science Readiness from Scientific Data to Large Language Models
- MetamatBench: Integrating Heterogeneous Data, Computational Tools, and Visual Interface for Metamaterial Discovery
- Can LLMs Solve Molecule Puzzles? A Multimodal Benchmark for Molecular Structure Elucidation
- Empowering and Assessing the Utility of Large Language Models in Crop Science
- Micro-Bench: A Microscopy Benchmark for Vision-Language Understanding
- OlympicArena: Benchmarking Multi-discipline Cognitive Reasoning for Superintelligent AI
- SciCode: A Research Coding Benchmark Curated by Scientists
- VLM4Bio: A Benchmark Dataset to Evaluate Pretrained Vision-Language Models for Trait Discovery from Biological Images
- Assessing Large Language Models on Climate Information
- Language Models as Science Tutors
- SciBench: Evaluating College-Level Scientific Problem-Solving Abilities of Large Language Models
- MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI
- SciEval: A Multi-Level Large Language Model Evaluation Benchmark for Scientific Research
- T-SciQ: Teaching Multimodal Chain-of-Thought Reasoning via Large Language Model Signals for Science Question Answering
- OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems
- SceMQA: A Scientific College Entrance Level Multimodal Question Answering Benchmark
- ClimateIQA: A New Dataset and Benchmark to Advance Vision-Language Models in Meteorology Anomalies Analysis
- Mathematical Capabilities of ChatGPT
- ProBio: A Protocol-guided Multimodal Dataset for Molecular Biology Lab
- What can Large Language Models do in chemistry? A comprehensive benchmark on eight tasks
- MatSci-NLP: Evaluating Scientific Language Models on Materials Science Language Tasks Using Text-to-Schema Modeling
- Learn to Explain: Multimodal Reasoning via Thought Chains for Science Question Answering
E3 LLM-driven discovery: equations, laws, programs, mathematics 29
- AutoNumerics-Zero: Automated Discovery of State-of-the-Art Mathematical Functions
- DecAEvolve: Decompose, Adapt, and Evolve for Effective LLM-based Scientific Equation Discovery
- Deliberate Evolution: Agentic Reasoning for Sample-Efficient Symbolic Regression with LLMs
- Influence-Guided Symbolic Regression: Scientific Discovery via LLM-Driven Equation Search with Granular Feedback
- Towards Solving the Gilbert-Pollak Conjecture via Large Language Models
- Helix: Evolutionary Reinforcement Learning for Open-Ended Scientific Problem Solving
- NewtonBench: Benchmarking Generalizable Scientific Law Discovery in LLM Agents
- Robust Equation Structure Learning with Adaptive Refinement
- SR-Scientist: Scientific Equation Discovery With Agentic AI
- Unleashing LLMs in Bayesian Optimization: Preference-Guided Framework for Scientific Discovery
- Scientifically-Interpretable Reasoning Network (ScIReN): Discovering Hidden Relationships in the Carbon Cycle and Beyond
- SciText2Eq: Assessing LLMs for Explainable Equation Generation for Scientific Creativity
- LLM-Based Scientific Equation Discovery via Physics-Informed Token-Regularized Policy Optimization
- Rethinking Scientific Modeling: Toward Physically Consistent and Simulation-Executable Programmatic Generation
- Learning Interestingness in Automated Mathematical Theory Formation
- PhysGym: Benchmarking LLMs in Interactive Physics Discovery with Controlled Priors
- LLM-SRBench: A New Benchmark for Scientific Equation Discovery with Large Language Models
- Neural Discovery in Mathematics: Do Machines Dream of Colored Planes?
- Gravity-Bench-v1: A Benchmark on Gravitational Physics Discovery for Agents
- LLM-SR: Scientific Equation Discovery via Programming with Large Language Models
- PhysPDE: Rethinking PDE Discovery and a Physical HYpothesis Selection Benchmark
- DiSciPLE: Learning Interpretable Programs for Scientific Visual Discovery
- FIND: A Framework for Discovering Formulas in Data
- Data-Driven Discovery of Dynamical Systems in Pharmacology using Large Language Models
- Global Lyapunov functions: a long-standing open problem in mathematics, with symbolic transformers
- Symbolic Regression with a Learned Concept Library
- Unsupervised Discovery of Formulas for Mathematical Constants
- LLM and Simulation as Bilevel Optimizers: A New Paradigm to Advance Physical Scientific Discovery
- Automated Search for Conjectures on Mathematical Constants using Analysis of Integer Sequences
E4 Scientific LLMs & landmark scientific foundation models (selective) 75
- Origo: Interpretable Multi-physics PDE Foundation Model through Neural Operator Splitting
- PLaID++: A Preference Aligned Language Model for Targeted Inorganic Materials Design
- Proteo-R1: Reasoning Foundation Models for De Novo Protein Design
- Scaling Laws and Architectural Frontiers in Metagenomic Foundation Models
- mCLM: A Modular Chemical Language Model that Generates Functional and Makeable Molecules
- CellDuality: Unlocking Biological Reasoning in LLMs with Self-Supervised RLVR
- Context parroting: A simple but tough-to-beat baseline for foundation models in scientific machine learning
- CoT-Evo: Evolutionary Distillation of Chain-of-Thought for Scientific Reasoning
- FM4NPP: A Scaling Foundation Model for Nuclear and Particle Physics
- NESTOR: A Nested MOE-based Neural Operator for Large-Scale PDE Pre-Training
- OlmoEarth: Stable Latent Image Modeling for Multimodal Earth Observation
- SciPedia: Unlocking the Value of Scientific Data for Pre-training
- GLA: Grounding Large Language Models in Molecular Hierarchy for Chemical Understanding
- STELLA: A Multimodal LLM for Protein Functional Annotation via Unified Sequence-Structure Encoding
- Caduceus: MoE Foundation Models for Unifying Biological and Natural Language
- Scaling Unlocks Broader Generation and Deeper Functional Understanding of Proteins
- UMA: A Family of Universal Models for Atoms
- AION-1: Omnimodal Foundation Model for Astronomical Sciences
- BioReason: Incentivizing Multimodal Biological Reasoning within a DNA-LLM Model
- ChemOrch: Empowering LLMs with Chemical Intelligence via Groundbreaking Synthetic Instructions
- ChemPile: A 250 GB Diverse and Curated Dataset for Chemical Foundation Models
- KnowMol: Advancing Molecular Large Language Models with Multi-Level Chemical Knowledge
- Mol-LLaMA: Towards General Understanding of Molecules in Large Molecular Language Model
- Training a Scientific Reasoning Model for Chemistry
- All-atom Diffusion Transformers: Unified generative modelling of molecules and materials
- OmniArch: Building Foundation Model for Scientific Computing
- OneForecast: A Universal Framework for Global and Regional Weather Forecasting
- Unisolver: PDE-Conditional Transformers Towards Universal Neural PDE Solvers
- Bio-xLSTM: Generative modeling, representation and in-context learning of biological and chemical sequences
- DPLM-2: A Multimodal Diffusion Protein Language Model
- ProteinBench: A Holistic Evaluation of Protein Foundation Models
- WeatherGFM: Learning a Weather Generalist Foundation Model via In-context Learning
- AnySat: One Earth Observation Model for Many Resolutions, Scales, and Modalities
- BIOMEDICA: An Open Biomedical Image-Caption Archive, Dataset, and Vision-Language Models Derived from Scientific Literature
- Galaxy Walker: Geometry-aware VLMs For Galaxy-scale Understanding
- Towards a Unified Copernicus Foundation Model for Earth Vision
- TerraMind: Large-Scale Generative Multimodality for Earth Observation
- Bridging Molecular Graphs and Large Language Models
- ChemVLM: Exploring the Power of Multimodal Large Language Models in Chemistry Area
- Knowledge-driven Scientific Large Language Models
- Training Compute-Optimal Protein Language Models
- Multiple Physics Pretraining for Spatiotemporal Surrogate Models
- Poseidon: Efficient Foundation Models for PDEs
- Scaling transformer neural networks for skillful and reliable medium-range weather forecasting
- SciInstruct: a Self-Reflective Instruction Annotated Dataset for Training Scientific Language Models
- Universal Physics Transformers: A Framework For Efficiently Scaling Neural Operators
- Caduceus: Bi-Directional Equivariant Long-Range DNA Sequence Modeling
- Cell2Sentence: Teaching Large Language Models the Language of Biology
- Diffusion Language Models Are Versatile Protein Learners
- DPOT: Auto-Regressive Denoising Operator Transformer for Large-Scale PDE Pre-Training
- ESM All-Atom: Multi-Scale Protein Language Model for Unified Molecular Modeling
- SaProt: Protein Language Modeling with Structure-aware Vocabulary
- DNABERT-2: Efficient Foundation Model and Benchmark For Multi-Species Genomes
- From Molecules to Materials: Pre-training Large Generalizable Models for Atomic Property Prediction
- Llemma: An Open Language Model for Mathematics
- Mol-Instructions: A Large-Scale Biomolecular Instruction Dataset for Large Language Models
- Towards 3D Molecule-Text Interpretation in Language Models
- Towards Foundational Models for Molecular Learning on Large-Scale Multi-Task Datasets
- BioCLIP: A Vision Foundation Model for the Tree of Life
- Masked Autoencoders for Microscopy are Scalable Learners of Cellular Biology
- SkySense: A Multi-Modal Remote Sensing Foundation Model Towards Universal Interpretation for Earth Observation Imagery
- InstructProtein: Aligning Human and Protein Language via Knowledge Instruction
- OceanGPT: A Large Language Model for Ocean Science Tasks
- ProtLLM: An Interleaved Protein-Language LLM with Protein-as-Word Pre-Training
- ProtT3: Protein-to-Text Generation for Text-based Protein Understanding
- BioT5+: Towards Generalized Biological Understanding with IUPAC Integration and Multi-task Tuning
- Structure-Enhanced Protein Instruction Tuning: Towards General-Purpose Protein Understanding with LLMs
- HyenaDNA: Long-Range Genomic Sequence Modeling at Single Nucleotide Resolution
- GIMLET: A Unified Graph-Text Model for Instruction-Based Molecule Zero-Shot Learning
- Towards Foundation Models for Scientific Machine Learning: Characterizing Scaling and Transfer Behavior
- ClimaX: A foundation model for weather and climate
- Unifying Molecular and Textual Representations via Multi-task Language Modelling
- Uni-Mol: A Universal 3D Molecular Representation Learning Framework
- Foundation Model for Material Science
- MolXPT: Wrapping Molecules with Text for Generative Pre-training