ICLR 2027 Workshop · San Francisco · April 2027
Automatic Agents for Computer Science
A half-day workshop on autoresearch agents: AI systems that formulate questions and hypotheses, plan and run investigations, generate and validate data, interpret evidence, and produce results that others can reproduce and scrutinize.
Workshop proposal for ICLR 2027. Program and dates are tentative until ICLR's decision on November 29, 2026.
Format
Half-day, in-person, non-archival
Venue
ICLR 2027, San Francisco
Date
April 29 or 30, 2027
Submission deadline
February 1, 2027, tentative
About the workshop
AI agents can increasingly write code, search the technical literature, call tools, and run long experimental workflows. Research asks for more than tool use. A research agent must identify useful questions, formulate testable hypotheses, choose informative experiments, allocate limited resources, interpret noisy evidence, revise its plan, and report findings with enough provenance for others to scrutinize and reproduce them. It must also fail legibly: an agent that produces an impressive number while hiding invalid assumptions, uncontrolled variables, or irreproducible steps is not yet a research system.
This workshop convenes researchers studying the methods, environments, evidence standards, evaluation protocols, and interaction patterns that turn capable agents into reliable research collaborators. We center the scientific method itself, whatever the research substrate. Progress may come from a new idea, algorithm, proof, measurement, experiment, system, dataset, or human-agent workflow. Data synthesis sits at the center of this agenda, because research agents increasingly construct examples, curricula, simulations, and evaluation cases, then decide how those artifacts are filtered, validated, and used.
The scope spans algorithms, systems, programming languages, theory, human-computer interaction, software engineering, and other computer-science domains, together with interfaces to scientific discovery beyond machine learning.
Goals
- Define autoresearch as an end-to-end scientific process, from question formation through experimentation, verification, evaluation, and human collaboration.
- Develop credible evidence of research capability: reproducible artifacts, carefully specified environments, meaningful baselines, and evaluations that separate genuine progress from brittle benchmark optimization.
- Understand which research strategies succeed and which fail: how agents search, allocate compute, choose experiments, synthesize and validate data, respond to evidence, and divide labor with people.
- Build a durable community with shared terminology, evaluation principles, reusable artifacts, and an open-problem agenda for autoresearch across computer science.
Open questions
- What should count as a new, valid, and reproducible research result produced under a fixed compute budget?
- How should an agent generate, select, and validate data without creating self-confirming evidence or optimizing a misleading proxy?
- Which records of hypotheses, code changes, experiments, failures, and human interventions are needed to audit and reproduce an agent's result?
- Which capabilities transfer across research problems, and which gains come from problem-specific engineering or benchmark leakage?
- How should people supervise, collaborate with, evaluate, and share credit with autoresearch agents?
Call for contributions
The workshop is non-archival. We welcome empirical, methodological, systems, benchmark, and position work on the topics below. Submission instructions, page limits, and the OpenReview link will be posted on this page.
Topics
- End-to-end autoresearch agents and multi-agent research systems
- Hypothesis generation, research planning, experiment design, and adaptive experimentation
- Search, exploration, optimization, and compute allocation for research agents
- Coding agents and execution environments for research across computer science
- Data synthesis, curriculum construction, data selection, simulation, and validation inside research loops
- Automated analysis, interpretation, visualization, and scientific communication
- Verification, falsification, uncertainty estimation, and detection of spurious or irreproducible results
- Benchmarks, environments, and evaluation protocols for research agents
- Reproducible research artifacts and provenance of human and agent contributions
- Human-AI collaboration, interfaces, oversight, and division of labor
- Memory, reflection, self-improvement, and learning from research trajectories
- Negative results, failure analysis, reward hacking, safety, security, attribution, and governance
- Applications across algorithms, machine learning, scientific discovery, search, programming languages, software engineering, systems, theory, and human-computer interaction
Formats
Research papers and extended abstracts
Empirical, methodological, systems, benchmark, or position work, presented as posters, with selected contributed talks in the oral session.
Challenge reports and artifacts
Concise reports on competition systems, findings, negative results, and research trajectories.
Tiny and short papers
Focused implementations, modest self-contained results, replications, re-analyses, and early findings. This format follows the ICLR prohibition on AI-generated tiny papers.
AI-participation tracks
Fully AI-written
An agent is the primary author. A human submitter remains accountable for compliance, provenance, safety, and rights.
AI-written with human review
Agent-written work that people reviewed and revised before submission.
Completely human-written
Human-authored work, with AI assistance only as ICLR policy allows.
Every submission discloses its track and gives an auditable account of human and agent contributions. Tiny and short papers remain human-authored. Submissions are evaluated for relevance, technical soundness, clarity, quality of evidence or argument, reproducibility where applicable, and potential to generate useful discussion. Human reviewers and organizers make all acceptance decisions. Organizers do not review submissions from their own institutions or with other conflicts, and organizer-authored submissions are not permitted.
Important dates
All deadlines 11:59 p.m. AoE- Feb 1, 2027Submission deadline, tentative
- Feb 26, 2027Author notification. Accepted papers are posted on OpenReview.
- Late Feb 2027Competition opens, immediately after acceptance decisions
- Apr 2027Competition closes, the day before ICLR 2027 begins
- Apr 29 or 30, 2027Workshop day in San Francisco. ICLR assigns the exact day.
Schedule
Tentative morning program| Time | Session |
|---|---|
| 08:30–08:40 | Opening remarks and workshop questions |
| 08:40–09:10 | Invited talk 1. Speaker to be announced. |
| 09:10–09:40 | Invited talk 2. Speaker to be announced. |
| 09:40–10:20 | Oral paper session: selected contributed work and discussion |
| 10:20–10:50 | Coffee break with the poster and artifact session |
| 10:50–11:20 | Invited talk 3. Speaker to be announced. |
| 11:20–12:20 | Competition results, reproducibility summary, and winning-team talk |
| 12:20–12:30 | Awards, synthesis, and closing remarks |
Clock times follow the half-day block that ICLR assigns.
Invited speakers
Three invited talks anchor the program. Speakers will be announced here once they confirm.
Competition
Opens after paper decisionsThe workshop hosts a competition centered on one fixed, research-oriented problem. Participants receive a standardized environment and aim to produce a reproducible new research result under constrained compute. An automatic metric scores the final artifact. Entries and reports use the same three participation tracks as the workshop papers.
The competition opens immediately after workshop-paper decisions are released and closes the day before ICLR 2027 begins. Results, awards, and the winning-team talk are presented at the workshop. The problem statement, baseline, environment, metric, compute limits, and leaderboard will be posted on this page when the competition launches.
The competition gives the workshop a shared empirical case study: how research agents allocate experiments, synthesize and validate data, respond to evidence, and produce reproducible artifacts. The workshop remains open to autoresearch work that does not use the challenge environment.
Organizers
Dingning Cao
MIT
Undergraduate at MIT advised by William T. Freeman, working on human-computer interaction, design intelligence, and generative models. Organizer of the MIT Informatics Tournament and HackMIT.
Wenhao Chai
Princeton University · Google DeepMind
Ph.D. student at Princeton advised by Karthik Narasimhan and student researcher at Google DeepMind. Leads MovieChat, co-leads LiveCodeBench Pro, and has organized workshops and competitions at CVPR 2024, 2025, and 2026.
Xianbang Wang
MIT
Undergraduate in mathematics and computer science at MIT advised by Kaiming He, working on efficient visual generation and multi-agent systems for automated research. Leads MiniT2I. Gold medalist at the 65th International Mathematical Olympiad.
Qiuyang Mang
UC Berkeley
Ph.D. student in the Sky Computing Lab at UC Berkeley advised by Alvin Cheung. Leads Frontier-CS and FrontierSmith, which study open-ended, verifiable computer-science challenges and data synthesis for LLM-driven algorithm evolution.
Hanchen Li
UC Berkeley
Ph.D. student in the Sky Computing Lab at UC Berkeley advised by Ion Stoica and Joseph E. Gonzalez, working on agent infrastructure, inference, and context-engineering automation. EuroSys 2025 Best Paper Award.
Questions about the workshop or the competition: write to Wenhao Chai at wc9403@princeton.edu or Qiuyang Mang at qmang@berkeley.edu.