Submitted:
03 August 2026
Posted:
05 August 2026
You are already at the latest version
The Evolving Roles of Humans and AI in Research
Abstract
Large language model (LLM)-based agents can increasingly support and empower human scientists in their work by reading literature, generating hypotheses, designing experiments, analyzing data, interpreting results, and drafting manuscripts. In parallel, robotics and programmable, cloud-connected laboratories have the potential to transform wet-lab operations into software-addressable infrastructure. Yet biological experiments face a persistent reality gap in which silent failures, temporal drift, contamination, sample mix-ups, and ambiguous protocol intent can invalidate an otherwise sound experimental plan. Closing this gap is essential for bringing the speed of modern computational discovery to experimental biology and requires audit-grade execution and verification. Here, we distinguish scripted automation from flexible, AI-enabled autonomy and outline a staged roadmap toward human-led and trustworthy “self-driving” labs for biology. We define laboratory autonomy in the context of a human-led scientific endeavor as a bounded control loop in which scientists set goals and constraints, agents help plan and coordinate experiments, instruments execute experiments and report machine-checkable evidence, and humans verify and interpret results and govern high-consequence decisions. We discuss where such systems are most appropriate and relevant today and where caution or deferral is needed. We map key lab-side potential failure modes and argue for verifiability-first autonomy. If developed responsibly, such systems could enable scientists to more efficiently decipher combinatorially large perturbation, design, and optimization spaces that are currently impractical to test experimentally, while empowering, rather than replacing, human creativity and judgment in experimental biological discovery.
Keywords:
autonomous laboratories
; artificial intelligence
; agentic AI
; human-led science
; verifiability
Introduction
As the scale of biological experiments and the data they generate continue to expand, biomedical research has the potential to shift from mostly manual benchwork and ad hoc analysis toward agentic science, where AI agents can support and enhance human scientists in the context of reading literature, forming hypotheses, designing experiments, executing relevant protocols through robotics or cloud labs, interpreting results, and iterating. Recent advances in large language model (LLM)-based agents for science[1,2,3,4,5,6,7,8] have highlighted the enormous potential of human-led agentic autonomy in biological science: human scientists can explore a larger set of hypotheses, informed by broader prior knowledge (encoded by foundation models and/or accessible through analytical tools), and investigate them more deeply. In these settings, the human scientist is empowered, not supplanted, and their scope is expanded to match the enormous scale of biological systems. Early examples suggest the potential of agentic science in physical laboratories as well. These include a multi-agent “Virtual Lab”[9] that, with human feedback, designed 92 SARS-CoV-2 nanobodies with experimental validation; a ‘self-driving lab’ that optimized LNPs[10]; and reports[11] of GPT-5 iterating on cloning protocols and coupling with an automated lab to optimize cell-free protein synthesis[12]. These examples are promising but still narrow demonstrations, typically in constrained workflows with substantial human feedback, workflow engineering, and experimental validation. If expanded to broader workflows, autonomous lab science could help biologists explore design spaces that are currently impractical to test experimentally.
However, unlike lab automation, full lab autonomy remains elusive in life science research. Lab automation uses robotics and software to execute repetitive, well-specified operations, such as particular types of sample preparation, liquid handling, cell incubation, imaging, and sequencing library preparation[13,14]. Cloud labs are automated, remotely accessible facilities where users submit digitally specified experiments for execution on standardized instruments with structured logging[15]. However, even highly automated labs are rarely autonomous because humans still set goals, choose experiments, diagnose failures, and decide how to adapt. Careful automation engineering is also required to ensure reproducible execution, and experimental flexibility is typically traded off for regimented automation. As a result, the cost and engineering burden of automation, together with its reduced flexibility, are ill-suited to the substantial variation of biological systems, and remain major barriers to broad, impactful adoption. Thus, while some experimental workflows have become highly automated, most types of biological experiments have remained manual. By contrast, in human-led, AI-driven lab autonomy, an end-to-end scientific control loop, from ideation to wet-lab validation, could operate with increasing independence while remaining motivated, guided, verified, and governed by human scientists, who define its goals and constraints. AI agents could help translate intent into executable protocols, monitor experiments, learn from prior results, and propose subsequent experiments within explicit safety and governance constraints. Currently, however, autonomy in experimental biology is constrained by a persistent reality gap between in silico plans and the biological and physical realities of execution, alongside unresolved challenges in robustness, interpretability, governance, safety, and biosecurity. As a result, computational models now iterate far faster than experimental validation, and they do so under very different standards of observability and trust. Closing this gap requires audit-grade execution and verification at the agent-automation interface, so that agentic reasoning can become verifiable wet-lab action. If successful, greater lab autonomy in broader areas could make experimentation more adaptive, reusable, and scalable, enabling scientific questions that are too complex or labor-intensive to pursue efficiently today, while enabling human researchers to focus more deeply on scientific judgment, creative hypothesis generation, unexpected observations, and methodological innovation in the lab.
In this Perspective, we present a staged roadmap to trustworthy, human-led “self-driving” labs and lay out their potential impact on biological discovery and experimental practice.
What Lab Autonomy Could Bring to Life Science Researchers
Life science research is an iterative discovery loop in which ideas are repeatedly turned into evidence, prompting a new cycle of ideas (Figure 1a). This loop currently has two major bottlenecks. First, cycle time and handoffs across stages, including translating hypotheses into runnable protocols, coordinating instruments and samples, diagnosing failures, and deciding what to do next, slow down every iteration. Second, the scale of ideation and hypothesis generation is often many orders of magnitude smaller than that of the biological systems or questions at hand, narrowing exploration. These bottlenecks are especially acute in biology, where reductionist workflows (one gene, one pathway, or one condition at a time) collide with systems that are high-dimensional, combinatorial, and context-dependent.
Lab autonomy would help address this mismatch by creating a verifiable closed loop for relevant experimental workflows in which models propose experiments, the autonomous lab executes them, the data are analyzed and quality checked, and the next decision is made. For example, when design spaces are combinatorially large, with 20,000 genes, hundreds of thousands of disease-associated variants, and tens of thousands of cell types and states (and their numerous combinations), simply scaling screens is insufficient[16]. Progress instead depends on adaptive experimentation that learns from intermediate results and selects the next most informative test under real constraints to address the question of interest or learn a more performant model. Importantly, this does not mean every experiment should be automated or autonomously run. Autonomy is likely to be most useful for relatively standardized and iterative loops, while increasing human effort toward deep, complex investigations, assay innovation, and the scientific “long tail” that requires hands-on creativity and judgment. This could accelerate human experimentation while changing which biological questions are tractable.
This paradigm is already becoming relevant in model-guided perturbation biology and high-content phenotypic screening[17], automated manufacturing and QC[18,19], autonomous microbial optimization for biomanufacturing and biofuels[20,21], and closed-loop small-molecule discovery[22]. For example, in high-content imaging assays that mapped which drugs perturb which proteins[23], active learning selected informative perturbations, liquid-handling robotics and automated microscopy executed them, and the results updated the model, reducing the number of experiments needed. Similarly, in combinatorial genetic perturbation atlases, where interaction spaces are vast, “lab-in-the-loop” cycles are needed to iteratively select, execute, and learn from the next perturbations[17]. Autonomy makes these loops practical by reducing coordination overhead, accelerating recovery and reruns, and keeping each iteration traceable from computation to bench[17].
Lab autonomy therefore has the potential to shift scientists’ efforts from “executing the loop manually” to “steering the loop”, and augmenting it. Scientists could focus on hypothesis framing, constraints (including safety), high-level principles of experimental design, supervision, and interpretation, along with additional lab experiments that are currently outside the realm of autonomy. At the same time, the autonomous system maintains traceability, verification, and governance across design, execution, analysis, and reporting. Crucially, lab autonomy can also lower the adoption barrier by translating intent into executable workflows and instrument-control code, thus giving each individual scientist a broader experimental toolbox. In practical terms, autonomy could reduce coordination and recovery labor, accelerate iteration under budget and throughput constraints, prioritize informative experiments in large design spaces, and tighten the computation-to-bench interface (Figure 1b).
Biology has navigated similar transitions in experimental scale before. For example, automated first-generation[24] and next-generation[25,26] sequencing, high-throughput screening[27], or pooled perturbation assays such as Perturb-seq[28] each automated labor-intensive tasks and enabled experiments at previously inaccessible scales, leading to many discoveries. Moreover, each such change democratized access, shifting methods once available to a few scientists to broad adoption across the field. Thus, rather than diminishing biological judgment, critical thinking, or problem-solving, each technology changed how biologists work and expanded the range of questions they addressed experimentally. Lab autonomy should be viewed similarly: not as a replacement for scientists, but as a tool within human-led science that extends their capacity to design, execute, and interpret increasingly complex experiments and tackle questions that would otherwise be too large, combinatorial, or labor-intensive.
Figure 1.
Framework for laboratory autonomy across the discovery loop. a, AI-driven discovery and wet-lab experimentation form coupled loops across hypothesis generation, experimental design, execution, analysis, results interpretation, and writing/publication. The reality gap arises when AI-generated experimental intent must be converted into executable protocols and wet-lab outcomes must be returned as verified, audit-grade evidence. b, Advantages of laboratory autonomy. c, Laboratory autonomy ladder from L0 manual science to L5 general-purpose autonomous laboratory, with the corresponding distribution of human and agent responsibility across lifecycle stages. The ladder integrates scientific and physical autonomy, from instrument-level automation (L1) and integrated workcells (L2) to domain-specific (L4) and general-purpose autonomous labs (L5). Heatmap indicates for each class of activity (column) at each level (row) whether humans do, agents assist, humans approve, or humans audit (color bar, bottom). Scientific and physical autonomy can advance asynchronously; the numbered levels (left) describe integrated laboratory configurations (matrix rows, center; scatter plot, right), whereas individual agents or automation platforms may occupy off-diagonal positions. d, Selected milestones and enabling technologies toward self-driving labs, highlighting representative milestones and enabling trends, including cloud labs, AI-enabled laboratory robots, the rise of LLMs[29,30,31,32], and LLM-based biomedical agents[33,34,35,36,37].
Figure 1.
Framework for laboratory autonomy across the discovery loop. a, AI-driven discovery and wet-lab experimentation form coupled loops across hypothesis generation, experimental design, execution, analysis, results interpretation, and writing/publication. The reality gap arises when AI-generated experimental intent must be converted into executable protocols and wet-lab outcomes must be returned as verified, audit-grade evidence. b, Advantages of laboratory autonomy. c, Laboratory autonomy ladder from L0 manual science to L5 general-purpose autonomous laboratory, with the corresponding distribution of human and agent responsibility across lifecycle stages. The ladder integrates scientific and physical autonomy, from instrument-level automation (L1) and integrated workcells (L2) to domain-specific (L4) and general-purpose autonomous labs (L5). Heatmap indicates for each class of activity (column) at each level (row) whether humans do, agents assist, humans approve, or humans audit (color bar, bottom). Scientific and physical autonomy can advance asynchronously; the numbered levels (left) describe integrated laboratory configurations (matrix rows, center; scatter plot, right), whereas individual agents or automation platforms may occupy off-diagonal positions. d, Selected milestones and enabling technologies toward self-driving labs, highlighting representative milestones and enabling trends, including cloud labs, AI-enabled laboratory robots, the rise of LLMs[29,30,31,32], and LLM-based biomedical agents[33,34,35,36,37].

From Automation to Autonomy Across the Lifecycle
Most automation and AI still target individual lifecycle stages rather than an end-to-end, iterative loop. Automation accelerates execution and data collection in the laboratory, while LLM-based agentic systems and AI co-scientists[7] provide scientific autonomy in settings where execution is deterministic, failures are explicit, and safety risks are minimal, leaving wet-lab execution or validation scientist-mediated[4,5].
By contrast, in current wet-lab biology, the assumptions of computational autonomy no longer hold (Figure 1a,b). Execution is not deterministic because biological state and sample identity can drift. Failures are not always explicit. Safety risks are higher when agents act on physical samples with hazardous reagents or dual-use workflows. Measurements can also be ambiguous due to batch effects or assay artifacts. Together, these define the wet-lab reality gap. Thus, while co-scientist agents may serve as a decision-control layer for self-driving labs, they are not sufficient. True self-driving labs require both scientific and physical autonomy. Agents must plan and analyze, execute through an instrumented, programmable layer, validate outcomes through machine-checkable QC and provenance, and iterate within defined constraints. Achieving this shift requires gradual progress rather than a single step change.
The Autonomy Ladder and the Evolving Human Role
The physical transition toward autonomy unfolds in stages. Classical lab automation handles structured, repetitive tasks such as liquid handling, plate movement, incubation, and imaging under rigid scenarios. Instrumental autonomy adds self-reporting devices, multimodal sensing, and standardized interfaces so the system can verify what happened and recover from common faults. Embodied AI and more general robotic systems may eventually extend autonomy to less structured tasks such as loading, transport, setup, and troubleshooting, but this frontier remains limited by sparse interaction data, generalization across lab layouts, and stricter safety requirements.
We organize agent-driven laboratory systems using a staged laboratory autonomy ladder (Figure 1c), inspired by the Society of Automotive Engineers automation levels for self-driving vehicles[38]. Prior work has proposed autonomy scales for LLM agents[1] and, separately, levels of physical automation in laboratories[39]. Our ladder integrates scientific reasoning, physical robotics, programmable execution, and human-agent roles across the research loop. However, scientific and physical autonomy can advance asynchronously; these levels describe integrated laboratory systems (Figure 1c, scatter plot, right), whereas individual agents or automation platforms may occupy off-diagonal positions. The convergence of cloud laboratories, AI-enabled laboratory robots, and LLM-based agents is making higher levels of laboratory autonomy increasingly plausible (Figure 1d).
- Level 0 (Manual science): Humans perform the full research and verification lifecycle, with no meaningful autonomous reasoning or physical automation.
- Level 1 (Instrument-level automation): Individual instruments or fixed software pipelines execute predefined tasks under scripted control. Scientific reasoning remains human-led, and physical automation is limited to single devices or tasks.
- Level 2 (Integrated workcell): Multiple instruments are linked into a coordinated workcell, which executes a defined multi-step workflow under human-designed, largely scripted control. This reduces manual handoffs, but humans still design workflows, supervise runs, handle errors, and make decisions across stages. Stage-specific AI may assist individual tasks, but no agent autonomously coordinates or adapts the workflow across stages.
- Level 3 (AI-assisted automation): Agents add cross-stage reasoning and adaptive coordination to integrated workcells. They can support hypothesis generation, protocol compilation, data analysis, and next-step recommendations, while sensing and monitoring enable error detection, diagnosis, and bounded recovery. Humans still approve major decisions, handle non-routine failures, and judge scientific meaning and safety.
- Level 4 (Domain-specific autonomous laboratory): Agents and automation jointly run an end-to-end control loop across cycles within a defined scientific realm. Within human-specified goals and constraints, the system can propose experiments, launch runs, monitor outcomes, analyze results, and recommend or execute the next iteration with limited routine human intervention. Autonomy remains bounded to one domain or workflow family, and humans retain governance, exception handling, and high-consequence decisions.
- Level 5 (General-purpose autonomous laboratory): Agents and automation operate a coordinated laboratory environment that can support multiple experimental domains. The defining feature is cross-domain transfer across assays, instruments, workflows, and laboratory settings. Humans still retain overall steering, governance, and high-consequence decisions, as well as integration with scientist-led work outside the lab’s autonomous scope.
Along this ladder, the role of the human lab scientist evolves, but it neither vanishes nor diminishes. Tasks that are structured, repetitive, low-risk, and machine-verifiable are the most amenable to autonomy. By contrast, scientists will spend more time on less standardized, context-dependent, and discovery-driven work, such as method development, assay design, novel workflows, complex biochemistry, in vivo studies, and high-novelty troubleshooting. Lab scientists will have access to a broader toolbox of methods and workflows and will need the scientific context to use them effectively. Laboratory autonomy does not remove scientists from the loop, but shifts human effort away from routine execution and monitoring and toward scientific judgment, experimental innovation, and interpretation. This positions AI within a human-led scientific loop, rather than humans as residual safeguards inside an AI-controlled loop. Even in highly autonomous settings, human scientists remain crucial for defining goals, setting constraints, taking responsibility for safety and security, driving creativity, and judging whether outputs are biologically meaningful and ethically acceptable. Machine verification can establish whether a protocol ran as specified; human verification remains necessary to judge whether the result is biologically meaningful, artefactual, surprising, or worth pursuing.
We reserve “lab autonomy” (Level 4 and Level 5) for systems that sustain closed-loop operation across multiple stages and cycles with failure recovery, verifiability, safety constraints, and audit trails. Autonomy emerges only when the stages are integrated into a coherent, self-adapting control loop. Existing systems often illustrate different components of the ladder rather than occupying clean, single rungs. Virtual Lab[9], SpatialAgent[36], The AI Scientist[7], Co-Scientist[4], and Robin[5] demonstrate substantial scientific or computational autonomy, but wet-lab execution is absent or remains outside their agentic control loops. Adam[40] and Boiko et al. (2023)[41] couple algorithmic reasoning to automated experimentation and show Level 3 characteristics. “SAMPLE”[42], LUMI-lab[10], and GPT-5-Ginkgo[12] more closely approach Level 4 within narrow, standardized domains. However, the extent of autonomous recovery, machine-verifiable state, provenance, and sustained closed-loop operation varies across these systems. Fully demonstrated Level 4 autonomy therefore remains an open frontier. On the biological science side, the challenge is sustained, domain-specific reasoning without continuous human correction. On the physical side, the challenge is coordination at scale across workcells, scheduling, inventory, calibration, and facility-level fault tolerance. Level 5 remains more futuristic because it requires autonomy that generalizes across domains and physical laboratory environments.
Foundations for Self-Driving Labs
In a self-driving laboratory, experiments must be executable, stateful (able to retain prior context and run history), and auditable so agents can specify intent, observe what happened, and recover when reality deviates from the plan. These capabilities emerge from three converging pillars: infrastructure, algorithms, and human and artificial intelligence (Figure 2a).
Infrastructure provides the execution layer for autonomy (Figure 2a, bottom left). It combines robotic systems and specialized instruments[43,44,45] with a software stack that makes lab state machine-readable and experiments executable[46], while supporting closed-loop reasoning[47]. Three capabilities are especially important. Protocol-as-code paradigms must be validated and compiled[48,49,50], lab APIs and standard device interfaces must provide consistent commands and error reporting[51,52,53], and state monitoring must link electronic lab notebooks (ELN) and laboratory information management systems (LIMS)[54,55] with telemetry and standardized metadata/logging[56,57]. Together, this stack provides the provenance, verification, and observability needed for autonomous decision-making. At the hardware level, actuation enables physical work, instrumentation makes assay state observable, and telemetry makes that observability actionable for recovery, attribution, and auditing. Level 4 or 5 autonomy therefore depends not only on better agents, but also on instruments that can report what they are doing and sensing.
Algorithms form the decision core for closed-loop experimentation. Design of Experiments[58] (DoE) establishes baselines and controls. Bayesian optimization and active learning support adaptive selection under cost, throughput, and uncertainty constraints[59,60,61] by choosing the next most informative experiment. Uncertainty-aware models can help decide which variants, perturbations, or gene combinations to test next[62,63]. Reinforcement learning (RL), by contrast, is useful when the system must make a sequence of dependent decisions during execution[64]. Notably, in life sciences, “optimal” must be defined in the presence of delayed readouts, batch effects, and execution uncertainty, motivating verifiability-first autonomy rather than objective-only optimization[65].
The human and artificial intelligence pillar combines human scientists who provide motivation, intent, framing, creativity, domain judgment, and governance, with agents who make autonomy scalable by orchestrating infrastructure and algorithms across stages. Human scientists decide which questions are meaningful and appropriate for autonomous loops, define experimental campaigns, connect autonomous runs to non-autonomous or scientist-led experiments, and interrupt or redirect the loop when results are surprising, unsafe, or biologically ambiguous. LLM-based agents support tool orchestration, protocol compilation, and execution planning[33,35,36,66,67,68,69,70], while physical AI systems extend this orchestration into physical interaction by enabling embodied agents to perceive, reason, and act in a laboratory through visual, force, pressure, and other telemetry signals[42,71]. Rather than a simple extension of LLM planning, this is part of the physical-control problem, in which systems must translate high-level intent into reliable low-level action in variable lab environments under sparse data and strict safety constraints[72]. Agents also provide a natural layer for role decomposition, e.g., planner, executor, critic, and safety officer, and for governance functions such as permissioning, audit trails, escalation rules, and safe termination behaviors[53]. As automation increases, humans shift from hands-on operation to scientific steering, exception review, governance, and complementary scientist-executed experiments.
Figure 2.
System composition, safety governance, and evaluation of self-driving labs. a, Foundations organized into three pillars. The infrastructure pillar (bottom left, red) provides the execution layer, including lab instruments and hardware, software, protocol-as-code, standard device interfaces, telemetry, and state monitoring. The algorithmic pillar (bottom right, yellow) provides the decision core, including experimental design, machine learning, reasoning, and planning. The intelligence pillar (top, grey) coordinates the loop, combining agentic systems, physical AI controllers, and human scientific steering for discovery, execution, reflection, monitoring, approval, and feedback. b, Layered safety architecture. From left: An agent planner (blue) compiles scientific intent into executable protocols that pass through pre-execution validation (Gate 1), execution-time safeguards (Gate 2), and post-execution auditing (Gate 3) all engaging with a human governor (top), and producing provenance logs and QC artifacts (right) while routing anomaly feedback to update memory and future plans (bottom). c, Evaluation framework and deployment cycle. Benchmarks assess agent capability, system performance, and safety and governance under stress and adversarial conditions (left). Deployment requires independent verification and licensure, with continuous change management and re-verification after updates to models, tools, or hardware (right).
Figure 2.
System composition, safety governance, and evaluation of self-driving labs. a, Foundations organized into three pillars. The infrastructure pillar (bottom left, red) provides the execution layer, including lab instruments and hardware, software, protocol-as-code, standard device interfaces, telemetry, and state monitoring. The algorithmic pillar (bottom right, yellow) provides the decision core, including experimental design, machine learning, reasoning, and planning. The intelligence pillar (top, grey) coordinates the loop, combining agentic systems, physical AI controllers, and human scientific steering for discovery, execution, reflection, monitoring, approval, and feedback. b, Layered safety architecture. From left: An agent planner (blue) compiles scientific intent into executable protocols that pass through pre-execution validation (Gate 1), execution-time safeguards (Gate 2), and post-execution auditing (Gate 3) all engaging with a human governor (top), and producing provenance logs and QC artifacts (right) while routing anomaly feedback to update memory and future plans (bottom). c, Evaluation framework and deployment cycle. Benchmarks assess agent capability, system performance, and safety and governance under stress and adversarial conditions (left). Deployment requires independent verification and licensure, with continuous change management and re-verification after updates to models, tools, or hardware (right).

The Reality Gap
The transition from Level 3 (AI-assisted automation) to Level 4 (domain-specific autonomous laboratory) is constrained by the mismatch between digital models and the dynamic physical-biological world[11,73]. In wet labs, drift (time-dependent changes in instruments or assay conditions) and silent failures (hidden execution errors that appear successful) are inevitable, so even a computationally “correct” plan can yield invalid results. Biological systems vary substantially, including in unpredictable ways. Both model-side limitations and lab-side failure modes are expected, just as human-run experiments frequently fail due to technical errors, lack of tacit know-how, biological differences, operator variability, or incomplete troubleshooting. The goal is to make these failures more detectable, attributable, recoverable, and comparable by establishing closed-loop trust through continuous verification of what was executed, what was observed, and whether the outcome remains scientifically interpretable[74]. This makes instrumentation, observability, and recovery central to autonomy.
Model-side limitations. LLM-based agents remain fragile in the context of long-running, stateful experiments. They can hallucinate steps or outcomes, propagate biases when evidence is incomplete, and lose track of state across multi-day experiments with queues, dependencies, and timing constraints[1]. For instance, LLMAgent4Bio[75] showed that agents could design workflows like PCR protocols, yet miss temporal constraints, such as enzyme degradation, when robots were queued. Related failures also appear in recent GPT-5-driven cell-free protein synthesis workflows, where agents violated fixed volume constraints or introduced subtle unit conversion errors that silently invalidated runs[12]. Addressing these failures will require external memory tied to run histories and reagent lifetimes, constraint-aware planning with explicit unit and volume checks, retrieval from protocol and instrument documentation, and abstention or escalation when uncertainty is high. As errors are unavoidable, the standard should be agents that expose uncertainty, check constraints, preserve reasoning and tool traces, and know when to defer to human judgment. Emerging embodied AI systems face a related bottleneck. Training physical policies requires extensive real-world interaction data[76]. A deeper barrier is that much laboratory work depends on tacit, continuous human reasoning throughout lab work, which is rarely explicitly acknowledged or documented in lab notebooks, research papers, or even stated out loud[77]. The interaction and decision data needed to capture this reasoning are therefore scarce relative to other aspects of scientific language corpora, limiting translation of high-level plans into reliable control across diverse laboratory layouts[78]. Progress will therefore also depend on simulation-to-real transfer, reusable manipulation primitives, richer telemetry, and conservative low-level controllers that keep actions within validated safety envelopes. Interpretability and calibration are also limited, making it difficult to know when the agent is uncertain or why it chose a particular action[79]. Successful-looking outputs may still lack proper use of evidence, belief revision, and convergent testing[80]. Finally, agents must reason under cost, throughput, and latency limits, where decisions often depend on delayed or partial readouts rather than immediate feedback. Debugging, reflection, and digital-twin resimulation can further delay intervention, which may compromise time-sensitive cells, tissues, reagents, or samples. These limitations highlight that linguistic fluency does not equate to operational reliability in a high-stakes wet lab[81,82].
Lab-side failures and detection needs. A dominant failure mode in physical labs involves silent biological errors, where the robot executes the motion but the experimental state is compromised. A core problem underlying such silent errors is insufficient sensing, diagnosis, and recovery to verify that the robot actually produced the experiment that the science required. These failures span a spectrum that requires explicit observability strategies[74] (Figure 3). First, execution-layer failures include misplaced or mis-sealed plates, liquid handling artifacts (bubbles, clogs, volume inconsistencies), unreported instrument errors[22,83], and lack of real-time sensing for variables, including temperature, liquid viscosity, and oxygenation. These can mask failures of an apparently successful automated run. For example, a viscous sample can partially clog a tip and under-deliver reagents. A verifiable system detects the anomaly from pressure traces or camera-based models, triggers recovery, and flags the affected wells for QC. Instrument malfunctions, such as an incomplete door closure or a crooked plate placement, can compromise experimental results and further lead to physical hardware damage. Autonomy requires instrument state detection, problem identification, and automatic execution of corrective actions to resume operation. Preventing such failures requires sensing, instrument telemetry, online calibration, structured error reporting, and robust exception handling. Second, state and identity failures include contamination, carryover, reagent degradation, and sample identity or chain-of-custody failures (swaps, barcode errors, mislabeling, etc.). For example, liquid dripping from pipettes during robot movement can lead to contamination issues that range from invalidation of assay results to biohazardous surfaces. At Level 3 autonomy, vision-based models could continuously process data from cameras to identify such anomalies, tag the locations, help assess the potential scientific impact, and articulate corrective actions for human recovery; at Level 4, the platform actively resolves and verifies such issues. Reliable autonomy requires chain-of-custody tracking, explicit controls, and auditable logs tied to sample and reagent lineage[53,84]. Finally, in measurement and data failures, time-dependent readout drift, batch effects, and missing metadata silently corrupt downstream analyses. Here, the key countermeasures are standardized metadata and data formats, automated QC, and provenance that links each measurement to instrument settings, calibration state, and run context. Aberrations in assay data may implicate execution-layer or state and identity failures. In autonomy, models that integrate historical context could stream data and help localize root causes in real time. Across these failure classes, explicit thresholds should trigger human escalation when anomalies cannot be resolved safely. Digital-twin style models can further operationalize verifiability by comparing predicted and observed telemetry, emitting anomaly signals, and prompting preventive maintenance[85,86].
Overall, bridging the reality gap requires a verifiability-first engineering mindset, where autonomy should be bounded by what can be machine-verified, not by what can be verbally specified or flawlessly executed. Protocols must be compiled into machine-checkable steps where critical actions emit structured telemetry, controls are explicit, and failures trigger safe recovery or escalation rather than being silently accepted. Progress should be measured by whether systems can detect, localize, and recover from these failure classes while preserving audit-grade provenance and safety constraints[46,87,88].
Safety, Governance, and Evaluation as System Requirements for Trustworthy Autonomy
Achieving Level 4 or 5 autonomy shifts the focus from agent capability — what a model can reason about or propose — to system reliability: what the laboratory can execute safely, verifiably, and recoverably. Safety in autonomous laboratories cannot be retrofitted through prompt engineering or post hoc review. Once autonomy extends into experimental design and execution, lab safety, biosecurity, governance, and verifiability are core system requirements that must constrain agent behavior at every stage of the scientific control loop[89,90].
As agents aggregate large bodies of literature, synthesize cross-domain knowledge, and design protocols, system risk can expand beyond obvious misuse to include information hazards, physically unsafe actions arising from overgeneralization or unverified assumptions, and inadvertent dual-use discovery. Tool-using agents can fail through delegated authority, privacy leaks, destructive actions, and false completion reports[91]. Trustworthy autonomy therefore requires safeguards that operate independently of the agent’s internal reasoning and confidence, such as uncertainty reporting[92,93] or confidence thresholds[94].
We argue that agent-driven laboratories need layered enforcement spanning pre-execution validation, execution-time constraints, and post-execution auditing (Figure 2b). Before any command is dispatched, a decoupled safety module should screen compiled protocols and parameterized actions against curated constraints of regulated pathogens, hazardous sequences, chemical precursors, and restricted experimental classes. During execution, sandboxed action spaces, hard bounds on volumes, concentrations, temperatures, durations, and reagent combinations, and real-time telemetry should constrain what is physically possible. Controllers should pause or abort runs automatically when deviations exceed validated thresholds. After execution, systems should produce machine-actionable provenance, QC artifacts, and exception reports, with anomalies routed to human governance review and, where appropriate, agent learning pipelines. These safeguards must be complemented by cybersecurity and AI-security practices.
Evaluation must evolve in parallel. Existing benchmarks for biomedical agents largely assess isolated, disembodied tasks such as literature retrieval and question answering[95,96] or analysis and code generation[97,98,99,100]. More recent agentic benchmarks, such as CompBioBench[101], BioMysteryBench[102], GeneBench[103], and AssayBench[104], test multi-step computational biology tasks but remain poor abstractions for wet-lab autonomy, which is sequential, physical, multi-stage, and failure-prone. We propose a three-layer evaluation framework aligned with increasing autonomy (Figure 2c, left). Agent Capability measures functions such as protocol drafting, tool calling, and logical consistency. System Performance evaluates end-to-end behavior across lifecycle stages. Safety and Governance stress-tests the system under adversarial and unexpected conditions. A useful end-to-end benchmark would include protocol compilation to executable steps, induced execution or identity faults, and scoring of detection, safe abort, and recovery, while preserving complete provenance. For scientific claims, however, machine-checkable provenance is necessary but not sufficient. Human reviewers must still assess whether the evidence supports the biological interpretation.
Crucially, these evaluations must measure both successes and failures, anomalies, refusals of unsafe actions, and appropriate triggering of recovery mechanisms. Because autonomous labs evolve physically over time, autonomy claims must also be governed by a continuous change-management loop (Figure 2c, right). Before deployment, and after updates to models, control software, or lab hardware, systems should undergo independent verification, regression testing, and safety recalibration.
Roadmap to Incremental Autonomy
Achieving Level 4 or 5 autonomy will involve a phased co-evolution across five coupled dimensions (Figure 4a). Software must move from prose protocols, post hoc logging, and vendor-specific formats toward executable specifications, real-time telemetry, and shared standards[105]. Hardware must move from human-supervised automation toward self-diagnosing, telemetry-rich instruments that can validate success of experimental steps, reagent identity, and assay state, and recover from common faults. Algorithms must learn from both successful and failed runs, optimize under constraints of budget, throughput, delayed readouts, and execution uncertainty, and provide interpretable rationales for experiment selection. Agents must move from isolated copilots toward longer-horizon systems with memory, reflection, modular roles, and architecture-level safety constraints. Finally, human scientists must define biological questions, risks, constraints, stopping rules, and points for interruption, so that autonomy strengthens scientific steering rather than replacing it.
Near-term deployment may differ by laboratory setting (Figure 4b). In smaller research laboratories, including many academic laboratories and early-stage discovery settings, shared workcells, core facilities, or cloud-lab use can provide access to automation for standardized assays, such as perturbation screens, imaging, sequencing library preparation, and routine optimization. These resources can improve reproducibility, enable greater scale, and relieve coordination burden without replacing exploratory bench science. At the same time, because research labs continuously develop new assays, modify protocols, work with unusual samples, and revise objectives as results emerge, they provide a demanding test of whether autonomous systems can move beyond rigid, predefined workflows. Larger-scale laboratory settings, including industrial biotechnology labs and more centralized research platforms, such as genomics centers, are often better positioned for closed-loop optimization (for example, of strains, enzymes, proteins, media, and bioprocesses), because goals, metrics, controls, and throughput constraints can be specified explicitly. Their repeated design-build-test-learn cycles make them a particularly strong environment for developing genuinely closed-loop autonomy. Finally, pharmaceutical or biomanufacturing laboratories can use autonomy in standardized drug screening, formulation, cell-line engineering, process development, and QC-rich manufacturing workflows, but face the distinct challenge of integrating these capabilities across long, complex pipelines aimed at translating biological insight into safe and effective medicines or biological reagents. This requires high standards of robustness, traceability, safety, and human oversight, with key steps, such as target selection, molecule optimization, in vivo studies, translational interpretation, portfolio decisions, and clinical-risk judgments all remaining human-led. Across all settings, autonomy is most appropriate for more repetitive, instrumented, iterative, and machine-verifiable workflows, and least appropriate for poorly standardized, high-risk, irreversible, or any ethically-sensitive work.
Figure 4.
Roadmap to autonomous lab. a, Co-evolution required for autonomous laboratories across five dimensions. Software (blue): protocols shift from human-readable to executable, logging from post hoc summaries to real-time telemetry, and interoperability from custom outputs to shared standards. Hardware: operation shifts from human operation to robotic autonomy, sensing from open-loop execution to closed-loop vigilance, and manual recovery to automatic recovery. Algorithms: decision-making progresses from handcrafted rules and one-time training toward learning from experience, with increased interpretability. Agents: agents move from single-pass, short-context copilots to long-horizon systems with persistent memory, iterative reflection, and dynamic multi-agent architectures. Human scientist partnering layer: human roles shift from hands-on operation, failure interpretation, and routine execution toward system governance, novel, scientist-led investigation, method innovation, auditing, and decision-making. b, Human-centered autonomy across laboratory settings. Human scientists remain at the center of the plan-execute-verify-reflect loop, while the balance of intelligence, infrastructure, and algorithms and the division of human-led and AI-led activities can vary across small research laboratories, large instrumented platforms, and large-scale process-optimization settings. c, Domain-progressive deployment. Early feasible domains combine high assay standardization and automation readiness with lower safety or ethical risk and irreversibility (e.g., enzyme kinetics, synthetic sequence generation, protein engineering, mammalian cell line screens). Deferred domains include settings with limited standardization or higher ethical and safety risk (e.g., in vivo animal studies, and pathogen research). Positions are illustrative and depend on the organism, target, experimental scale, intended application, and safeguards.
Figure 4.
Roadmap to autonomous lab. a, Co-evolution required for autonomous laboratories across five dimensions. Software (blue): protocols shift from human-readable to executable, logging from post hoc summaries to real-time telemetry, and interoperability from custom outputs to shared standards. Hardware: operation shifts from human operation to robotic autonomy, sensing from open-loop execution to closed-loop vigilance, and manual recovery to automatic recovery. Algorithms: decision-making progresses from handcrafted rules and one-time training toward learning from experience, with increased interpretability. Agents: agents move from single-pass, short-context copilots to long-horizon systems with persistent memory, iterative reflection, and dynamic multi-agent architectures. Human scientist partnering layer: human roles shift from hands-on operation, failure interpretation, and routine execution toward system governance, novel, scientist-led investigation, method innovation, auditing, and decision-making. b, Human-centered autonomy across laboratory settings. Human scientists remain at the center of the plan-execute-verify-reflect loop, while the balance of intelligence, infrastructure, and algorithms and the division of human-led and AI-led activities can vary across small research laboratories, large instrumented platforms, and large-scale process-optimization settings. c, Domain-progressive deployment. Early feasible domains combine high assay standardization and automation readiness with lower safety or ethical risk and irreversibility (e.g., enzyme kinetics, synthetic sequence generation, protein engineering, mammalian cell line screens). Deferred domains include settings with limited standardization or higher ethical and safety risk (e.g., in vivo animal studies, and pathogen research). Positions are illustrative and depend on the organism, target, experimental scale, intended application, and safeguards.

Deployment should also proceed domain by domain (Figure 4c). Early demonstrations such as “SAMPLE”[42] or the OpenAI/Ginkgo work[12] show characteristics spanning Levels 3 and 4 in narrow, highly standardized spaces, like cell-free protein engineering. More broadly, domains should be prioritized by two axes: assay and workflow readiness (based on current reproducibility, instrumentation, and maturity) versus ethical risk and irreversibility (the potential for dual-use misuse, safety hazards, or irreversible harm). Early feasible domains include protein engineering, enzyme kinetics, synthetic sequence generation, and some mammalian cell line work, where protocols are more established, instrumentation is mature, and failure modes are containable. In contrast, in vivo animal studies should be deferred because timescales are long, standardization is limited, and risks are higher. Autonomous operation in pathogen research should be indefinitely and explicitly forbidden, given the substantially higher ethical or biosecurity consequences. Even with technological maturity, some domains will require permanent human execution as a matter of principle rather than capability.
Economic viability is another bottleneck. Adoption is limited not only by hardware costs, but by the operational burden of protocol translation, integration across vendor APIs, calibration, troubleshooting, and continual workflow adaptation. This burden is especially acute in research settings compared with manufacturing, where assays change quickly and scripts rarely transfer cleanly across contexts. Agentic systems can reduce this expertise bottleneck by generating executable protocols from natural language, debugging common failures, and adapting workflows across instruments with less manual recoding. Hardware costs will remain substantial, but as agentic capabilities mature and cloud labs achieve economies of scale, autonomous laboratories could broaden access to large-scale experimental campaigns that are now largely confined to industrial-scale facilities, while also making autonomy claims comparable through shared testbeds and benchmarks.
A Call to Action for Biologists
Autonomous labs should be designed by and for biologists, not imposed from outside. Their development should be guided by the scientists who understand which assays are reliable, which readouts are ambiguous, which failures are common, and which questions matter. Biologists should define the domains where autonomy is useful, the controls and metadata that make runs interpretable, the thresholds for human escalation, and the standards for reporting machine-executed experiments. This is especially important because machine verification can establish whether a protocol ran as specified, but human judgment remains essential for deciding whether an unexpected result is a clear artefact, a discovery, or something that requires further investigation.
The opportunity is timely. Biology has repeatedly integrated new technologies by building shared standards, quality controls, and communities of practice that expanded their reach and impact. Agent-driven laboratories should follow the same path. The immediate priorities are to define benchmark experiments and representative failure modes; establish reporting standards for machine-executed experiments; develop open, well-documented interfaces and standardized representations for protocols, metadata, telemetry, and outcomes; build shared abstraction layers and cross-laboratory testbeds that allow agents to interact with heterogeneous instruments without being redesigned for each vendor-specific system; and establish change-management approaches that trigger re-verification after updates to models, software, hardware, or workflows. Without these foundations, autonomous laboratories risk remaining isolated, non-transferable demonstrations rather than broadly usable scientific infrastructure. Their value will depend on whether they are made interoperable, legible, governable, and useful to working scientists across academia, biotechnological research, and pharmaceutical settings. Academic laboratories will need accessible shared infrastructure, core facilities, and transparent reporting standards; industrial research groups will need robust closed-loop optimization under explicit constraints; and pharmaceutical laboratories will need auditable systems that connect discovery, process development, QC, and translational judgment. Across these settings, the goal is to have AI that gives lab biologists greater experimental reach while preserving human responsibility for evidence, interpretation, safety, and impact.
Conclusions
An autonomous lab is defined not by classical automation of isolated steps, but by integrating design, execution, analysis, interpretation, and reporting into a verifiable, failure-recovering closed loop. That future lab is not a human-free room of robots, but a human-led scientific enterprise in which autonomous systems take on repetitive, time-sensitive, and failure-prone work such as protocol translation, scheduling, monitoring, QC, recovery, and routine optimization. This shift should give scientists more time to frame questions, set constraints, leverage a broader set of experimental approaches, interpret ambiguous results, integrate evidence across scales, innovate methods, and decide which directions matter biologically and clinically. Autonomy should therefore empower human scientists and amplify human-led science, not replace it.
In biomedical experimental science, these autonomous laboratories will be intertwined with the broader research ecosystem rather than operating in isolation. They will connect computational models, shared instrumentation, programmable cloud labs, biofoundries, data platforms, and downstream translational workflows into a more continuous path from hypothesis to evidence. Realizing this future will require better models, richer sensing, greater API standardization, and verifiability-first design that detects silent failures before they become misleading science. It will also require layered safety guardrails, from pre-execution screening and execution-time constraints to post-run auditing and continuous revalidation, so that autonomy claims remain trustworthy as systems evolve. The next phase should therefore prioritize system integration, observability, logging standards, and wet-lab-grounded evaluation infrastructure, so that autonomous labs become reliable parts of the scientific enterprise rather than isolated demonstrations. If developed responsibly and governed by the biological community, these systems could help scientists probe combinatorial biology more deeply, iterate faster through hypothesis and evidence, and translate insight into interventions more efficiently, with profound implications for biomedicine.
Author Contributions
Conceptualization, W.C., A.R., J.M. Visualization, A.H., Y.Z. Writing, W.C., M.M., S.S., J.E.R., G.A., C.B., M.A.L., A.R., J.M.
Acknowledgements
We thank the members of the Ma laboratory for helpful discussions. We thank the Medra team and Stacie Calad-Thomson for helpful discussions on the different levels of scientific and physical autonomy. This work was supported, in part, by the National Institutes of Health through grants UM1HG011593 (J.M.), UH3CA268202 (J.M.), R01HG007352 (J.M.), R01HG012303 (J.M.), R03OD039980 (J.M.), R21DA061481 (J.M.), and U24HG012070 (J.M.). J.M. was additionally supported by the Ray and Stephanie Lane Professorship, a Guggenheim Fellowship from the John Simon Guggenheim Memorial Foundation, and a Google Research Award. The funders had no role in the conception, writing, or decision to submit this manuscript.
Competing Interests
A.H., J.E.R., C.B., and A.R. are employees of Genentech, members of the Roche Group, and have equity in Roche. A.R. is a co-founder and equity holder in Immunitas and an inventor of multiple patents held by the Broad Institute for single cell genomics, spatial genomics and perturbation screens. M.A.L. is the founder and equity holder in Medra AI. G.A. is an employee of and equity holder in Medra AI. The remaining authors declare no competing interests.
Table 1.
Key AI and agentic-system terms.
| Term | Description |
|---|---|
| Large language model (LLM) | A foundation model trained on large-scale text, code, or scientific corpora that can understand, generate, and support reasoning over natural language or code-like inputs |
| AI agent | An AI system that can pursue a goal by planning actions, using tools, observing outcomes, updating context (memory), and adapting future steps |
| Multi-agent system | A system in which multiple specialized agents coordinate through roles, shared memory, or communication protocols to solve complex tasks |
| Agent tool | An external function, software package, database, instrument interface, or API that an agent can call to retrieve information, run analyses, or control software or instruments |
| Agent memory | Stored information that helps an agent maintain context across steps, experiments, or campaigns, such as previous plans, run histories, failures, decisions, and observations |
| Agent reflection | A process in which an agent reviews its prior actions, outcomes, or errors and revises its plan or strategy accordingly |
| Co-scientist agents / AI scientists | AI systems that assist with scientific discovery tasks such as literature review, hypothesis generation, experimental design, data analysis, interpretation, and manuscript drafting. The term describes scientific reasoning or assistance and does not by itself imply control of physical laboratory execution |
| Active learning | A machine-learning strategy in which a model selects the most informative next experiment or data point to measure, reducing the need for exhaustive testing |
| Bayesian optimization | An adaptive optimization method that uses a probabilistic model to choose the next experiment by balancing exploration of uncertain options with exploitation of promising candidates |
| Reinforcement learning | A machine learning framework in which an agent makes sequential decisions and improves based on feedback, rewards, or outcomes, especially when earlier actions affect later options |
| Embodied AI / Physical AI | AI systems that perceive, reason, and act in the physical world through sensors, robotics, and control policies |
| In-context learning | The ability of a model to adapt its behavior using examples, instructions, or context provided in the prompt, without updating its underlying parameters |
Table 2.
Key concepts related to autonomous laboratories.
| Term | Description |
|---|---|
| Cloud lab | A remotely accessible automated laboratory where users submit digitally specified experiments that run on standardized instruments with structured logging |
| Workcell | A physically and digitally connected set of instruments that executes a defined multi-step workflow, typically under predefined, human-designed control |
| Human-led autonomy | A model of autonomy in which scientists define goals and constraints, interpret scientific meaning, and retain responsibility for governance and high-consequence decisions, while agents and automation execute bounded workflows autonomously |
| Scientific autonomy | The ability of a system to reason across the research cycle by formulating or refining hypotheses, designing experiments, interpreting results, and selecting meaningful next steps within human-defined goals and constraints |
| Physical autonomy | The ability of a system to execute, monitor, coordinate, and recover physical experiments across instruments and workcells |
| Verifiability-first autonomy | An approach in which autonomous operation is bounded by what can be observed and machine-verified, with explicit controls, telemetry, provenance, recovery, and human escalation for unresolved cases. |
| Digital twin | A computational representation that is dynamically linked to a physical system, instrument, workflow, or laboratory state through observed data and can be used to compare expected and actual behavior, detect anomalies, and support control or maintenance |
| Protocol-as-code | A machine-executable representation of an experimental protocol that can be validated, versioned, compiled, and executed by laboratory software or instruments |
| Telemetry | Structured data emitted by instruments or software during execution, such as pressure traces, timing, temperature, images, device status, or error signals |
| Electronic lab notebooks (ELN) | Digital systems for recording experimental plans, observations, results, metadata, and researcher notes |
| Laboratory information management systems (LIMS) | Software systems for tracking samples, reagents, workflows, metadata, storage, processing steps, and experimental results across a laboratory |
| Design of Experiments (DoE) | A statistical framework for systematically planning experiments, controls, variables, and comparisons to obtain reliable information efficiently |
| Chain-of-custody | A record that tracks the identity, location, handling, and lineage of samples, reagents, and data throughout an experiment |
| Silent failure | A hidden execution or assay failure that appears successful but compromises the biological state or downstream data |
| Observability | The ability to determine and reconstruct the state of an experiment or laboratory system from telemetry, logs, metadata, and intermediate measurements, enabling anomaly detection, diagnosis, verification, and auditing |
References
- Gao, S.; et al. Empowering biomedical discovery with AI agents. Cell 2024, 187, 6125–6151. [Google Scholar] [CrossRef] [PubMed]
- MacKnight, R.; et al. Rethinking chemical research in the age of large language models. Nat. Comput. Sci. 2025, 5, 715–726. [Google Scholar] [CrossRef] [PubMed]
- Wei, J.; et al. From AI for science to Agentic Science: A survey on autonomous scientific discovery. arXiv [cs.LG] 2025. [Google Scholar] [CrossRef]
- Gottweis, J.; et al. Accelerating scientific discovery with Co-Scientist. Nature 2026, 655, 487–496. [Google Scholar] [CrossRef] [PubMed]
- Ghareeb, A. E.; et al. A multi-agent system for automating scientific discovery. Nature 2026, 655, 497–505. [Google Scholar] [CrossRef] [PubMed]
- Kitano, H. Nobel Turing Challenge: creating the engine for scientific discovery. npj Syst. Biol. Appl. 2021, 7, 29. [Google Scholar] [CrossRef] [PubMed]
- Lu, C.; et al. Towards end-to-end automation of AI research. Nature 2026, 651, 914–919. [Google Scholar] [CrossRef] [PubMed]
- Wang, H.; et al. Scientific discovery in the age of artificial intelligence. Nature 2023, 620, 47–60. [Google Scholar] [CrossRef] [PubMed]
- Swanson, K.; Wu, W.; Bulaong, N. L.; Pak, J. E.; Zou, J. The Virtual Lab of AI agents designs new SARS-CoV-2 nanobodies. Nature 2025, 646, 716–723. [Google Scholar] [CrossRef] [PubMed]
- Xu, Y.; et al. LUMI-lab: A foundation model-driven autonomous platform enabling discovery of ionizable lipid designs for mRNA delivery. Cell 2026, 189, 1620–1635.e25. [Google Scholar] [CrossRef] [PubMed]
- Measuring AI’s capability to accelerate biological research in the wet lab. Available online: https://openai.com/index/accelerating-biological-research-in-the-wet-lab/.
- Smith, A. A.; et al. Using a GPT-5-driven autonomous lab to optimize the cost and titer of cell-free protein synthesis. bioRxiv 2026.02.05.703998 2026. [Google Scholar] [CrossRef]
- Chapman, T. Lab automation and robotics: Automation on the move. Nature 2003, 421, 661. [Google Scholar] [CrossRef] [PubMed]
- Boyd, J. Tech.Sight. Robotic laboratory automation. Science 2002, 295, 517–518. [Google Scholar] [CrossRef] [PubMed]
- Arnold, C. Cloud labs: where robots do the research. Nature 2022, 606, 612–613. [Google Scholar] [CrossRef] [PubMed]
- Yao, D.; et al. Scalable genetic screening for regulatory circuits using compressed Perturb-seq. Nat. Biotechnol. 2024, 42, 1282–1295. [Google Scholar] [PubMed]
- Rood, J. E.; Hupalowska, A.; Regev, A. Toward a foundation model of causal cell and tissue biology with a Perturbation Cell and Tissue Atlas. Cell 2024, 187, 4520–4545. [Google Scholar] [CrossRef] [PubMed]
- Melocchi, A.; et al. Automated manufacturing of cell therapies. J. Control. Release 2025, 381, 113561. [Google Scholar] [CrossRef] [PubMed]
- Abou-El-Enein, M.; et al. Scalable manufacturing of CAR T cells for cancer immunotherapy. Blood Cancer Discov. 2021, 2, 408–422. [Google Scholar] [CrossRef] [PubMed]
- Carruthers, D. N.; et al. Automation and machine learning drive rapid optimization of isoprenol production in Pseudomonas putida. Nat. Commun. 2025, 16, 11489. [Google Scholar] [CrossRef] [PubMed]
- Martin, H. G.; et al. Perspectives for self-driving labs in synthetic biology. Curr. Opin. Biotechnol. 2023, 79, 102881. [Google Scholar] [CrossRef] [PubMed]
- Koscher, B. A.; et al. Autonomous, multiproperty-driven molecular discovery: From predictions to measurements and back. Science 2023, 382, eadi1407. [Google Scholar] [CrossRef] [PubMed]
- Naik, A. W.; Kangas, J. D.; Sullivan, D. P.; Murphy, R. F. Active machine learning-driven experimentation to determine compound effects on protein patterns. Elife 2016, 5, e10047. [Google Scholar] [CrossRef] [PubMed]
- Smith, L. M.; et al. Fluorescence detection in automated DNA sequence analysis. Nature 1986, 321, 674–679. [Google Scholar] [CrossRef] [PubMed]
- Margulies, M.; et al. Genome sequencing in microfabricated high-density picolitre reactors. Nature 2005, 437, 376–380. [Google Scholar] [CrossRef] [PubMed]
- Bentley, D. R.; et al. Accurate whole human genome sequencing using reversible terminator chemistry. Nature 2008, 456, 53–59. [Google Scholar] [CrossRef] [PubMed]
- Inglese, J.; et al. Quantitative high-throughput screening: a titration-based approach that efficiently identifies biological activities in large chemical libraries. Proc. Natl. Acad. Sci. U. S. A. 2006, 103, 11473–11478. [Google Scholar] [CrossRef] [PubMed]
- Dixit, A.; et al. Perturb-seq: Dissecting molecular circuits with scalable single-cell RNA profiling of pooled genetic screens. Cell 2016, 167, 1853–1866.e17. [Google Scholar] [CrossRef] [PubMed]
- Vaswani, A.; et al. Attention Is All You Need. NIPS 2017. [Google Scholar] [CrossRef]
- Guo, D.; et al. DeepSeek-R1 incentivizes reasoning in LLMs through reinforcement learning. Nature 2025, 645, 633–638. [Google Scholar] [CrossRef] [PubMed]
- Touvron, H.; et al. LLaMA: Open and efficient foundation language models. arXiv 2023. [Google Scholar] [CrossRef]
- Brown, T. B.; et al. Language Models are Few-Shot Learners. arXiv 2020. [Google Scholar] [CrossRef]
- Huang, K.; et al. Autonomous biomedical research with an artificial intelligence agent. Science 2026, eadz4351. [Google Scholar] [PubMed]
- Wang, Z.; et al. GeneAgent: self-verification language agent for gene-set analysis using domain databases. Nat. Methods 2025, 22, 1677–1685. [Google Scholar] [CrossRef] [PubMed]
- Qu, Y.; et al. CRISPR-GPT for agentic automation of gene-editing experiments. Nat. Biomed. Eng. 2025, 10, 245–258. [Google Scholar] [CrossRef] [PubMed]
- Wang, H.; et al. SpatialAgent: An Autonomous AI Agent for Spatial Biology. bioRxiv 2025. [Google Scholar] [CrossRef]
- Gao, S.; et al. Democratizing AI scientists using ToolUniverse. arXiv 2025. [Google Scholar] [CrossRef]
- SAE International Recommended Practice, Taxonomy and Definitions for Terms Related to Driving Automation Systems for On-Road Motor Vehicles. (SAE Standard J3016_202104, 2021).
- Angelopoulos, A.; Cahoon, J. F.; Alterovitz, R. Transforming science labs into automated factories of discovery. Sci. Robot. 2024, 9, eadm6991. [Google Scholar] [CrossRef] [PubMed]
- King, R. D.; et al. The automation of science. Science 2009, 324, 85–89. [Google Scholar] [CrossRef] [PubMed]
- Boiko, D. A.; MacKnight, R.; Kline, B.; Gomes, G. Autonomous chemical research with large language models. Nature 2023, 624, 570–578. [Google Scholar] [CrossRef] [PubMed]
- Rapp, J. T.; Bremer, B. J.; Romero, P. A. Self-driving laboratories to autonomously navigate the protein fitness landscape. Nat. Chem. Eng. 2024, 1, 97–107. [Google Scholar] [CrossRef] [PubMed]
- Dettinger, P.; et al. Open-source personal pipetting robots with live-cell incubation and microscopy compatibility. Nat. Commun. 2022, 13, 2999. [Google Scholar] [CrossRef] [PubMed]
- Huang, J.; Liu, H.; Junginger, S.; Thurow, K. Mobile robots in automated laboratory workflows. SLAS Technol. 2025, 30, 100240. [Google Scholar] [CrossRef]
- Jia, Y.; et al. Robot-assisted mapping of chemical reaction hyperspaces and networks. Nature 2025, 645, 922–931. [Google Scholar] [CrossRef] [PubMed]
- Canty, R. B.; et al. Science acceleration and accessibility with self-driving labs. Nat. Commun. 2025, 16, 3856. [Google Scholar] [CrossRef] [PubMed]
- Zhang, W.; et al. IvoryOS: an interoperable web interface for orchestrating Python-based self-driving laboratories. Nat. Commun. 2025, 16, 5182. [Google Scholar] [CrossRef] [PubMed]
- Mehr, S. H. M.; Craven, M.; Leonov, A. I.; Keenan, G.; Cronin, L. A universal system for digitization and automatic execution of the chemical synthesis literature. Science 2020, 370, 101–108. [Google Scholar] [CrossRef] [PubMed]
- Šiaučiulis, M.; Knittl-Frank, C.; M Mehr, S. H.; Clarke, E.; Cronin, L. Reaction blueprints and logical control flow for parallelized chiral synthesis in the Chemputer. Nat. Commun. 2024, 15, 10261. [Google Scholar] [CrossRef] [PubMed]
- Steiner, S.; et al. Organic synthesis in a modular robotic system driven by a chemical programming language. Science 2019, 363, eaav2211. [Google Scholar] [CrossRef] [PubMed]
- Dave, A.; et al. Autonomous optimization of non-aqueous Li-ion battery electrolytes via robotic experimentation and machine learning coupling. Nat. Commun. 2022, 13, 5454. [Google Scholar] [CrossRef] [PubMed]
- MacLeod, B. P.; et al. Self-driving laboratory for accelerated discovery of thin-film materials. Sci. Adv. 2020, 6, eaaz8867. [Google Scholar] [CrossRef] [PubMed]
- Wilkinson, M. D.; et al. The FAIR Guiding Principles for scientific data management and stewardship. Sci. Data 2016, 3, 160018. [Google Scholar] [CrossRef] [PubMed]
- Plass, F.; et al. Using OpenBIS As virtual research environment: An ELN-LIMS open-source database tool as a framework within the CRC 1411 Design of Particulate Products. Data Sci. J. 2023, 22, 44. [Google Scholar] [CrossRef]
- Edfeldt, K.; et al. A data science roadmap for open science organizations engaged in early-stage drug discovery. Nat. Commun. 2024, 15, 5640. [Google Scholar] [CrossRef] [PubMed]
- Rauschen, R.; Guy, M.; Hein, J. E.; Cronin, L. Universal chemical programming language for robotic synthesis repeatability. Nat. Synth. 2024, 3, 488–496. [Google Scholar] [CrossRef]
- Strieth-Kalthoff, F.; et al. Delocalized, asynchronous, closed-loop discovery of organic laser emitters. Science 2024, 384, eadk9227. [Google Scholar] [CrossRef] [PubMed]
- Dean, A.; Voss, D. Fractional Factorial Experiments. In Springer Texts in Statistics; Springer-Verlag: New York, 2006; pp. 483–545. [Google Scholar]
- Häse, F.; Roch, L. M.; Aspuru-Guzik, A. Next-generation experimentation with self-driving laboratories. Trends Chem. 2019, 1, 282–291. [Google Scholar] [CrossRef]
- Yang, K. K.; Wu, Z.; Arnold, F. H. Machine-learning-guided directed evolution for protein engineering. Nat. Methods 2019, 16, 687–694. [Google Scholar] [CrossRef] [PubMed]
- Kusne, A. G.; et al. On-the-fly closed-loop materials discovery via Bayesian active learning. Nat. Commun. 2020, 11, 5966. [Google Scholar] [CrossRef] [PubMed]
- Hie, B.; Bryson, B. D.; Berger, B. Leveraging uncertainty in machine learning accelerates biological discovery and design. Cell Syst. 2020, 11, 461–477.e9. [Google Scholar] [CrossRef] [PubMed]
- Zhang, R.; Bafna, M.; Ma, J.; Ma, J. Dango: Predicting higher-order genetic interactions. Cell Syst. 2026, 17, 101593. [Google Scholar] [CrossRef] [PubMed]
- Segler, M. H. S.; Preuss, M.; Waller, M. P. Planning chemical syntheses with deep neural networks and symbolic AI. Nature 2018, 555, 604–610. [Google Scholar] [CrossRef] [PubMed]
- Leek, J. T.; et al. Tackling the widespread and critical impact of batch effects in high-throughput data. Nat. Rev. Genet. 2010, 11, 733–739. [Google Scholar] [CrossRef] [PubMed]
- Coley, C. W.; et al. A robotic platform for flow synthesis of organic compounds informed by AI planning. Science 2019, 365, eaax1566. [Google Scholar] [CrossRef] [PubMed]
- Burger, B.; et al. A mobile robotic chemist. Nature 2020, 583, 237–241. [Google Scholar] [CrossRef] [PubMed]
- Bran, A. M.; et al. Augmenting large language models with chemistry tools. Nat. Mach. Intell. 2024, 6, 525–535. [Google Scholar] [CrossRef] [PubMed]
- Yao, S.; et al. ReAct: Synergizing reasoning and acting in language models. arXiv 2022. [Google Scholar]
- Schick, T.; et al. Toolformer: Language models can teach themselves to use tools. NeurIPS 2023. [Google Scholar] [CrossRef]
- Gregor, B. W.; et al. Automated human induced pluripotent stem cell culture and sample preparation for 3D live-cell microscopy. Nat. Protoc. 2024, 19, 565–594. [Google Scholar] [CrossRef] [PubMed]
- Wang, G.; et al. Voyager: An open-ended embodied agent with large language models. Trans. Mach. Learn. Res. 2024. [Google Scholar] [CrossRef]
- Seifrid, M.; et al. Autonomous Chemical Experiments: Challenges and Perspectives on Establishing a Self-Driving Lab. Acc. Chem. Res. 2022, 55, 2454–2466. [Google Scholar] [CrossRef] [PubMed]
- Szymanski, N. J.; et al. An autonomous laboratory for the accelerated synthesis of novel materials. Nature 2023, 624, 86–91. [Google Scholar] [CrossRef] [PubMed]
- Dip, S. A.; et al. Large language model agents for biological intelligence across genomics, proteomics, spatial biology, and biomedicine. Brief. Bioinform. 2026, 27, bbag110. [Google Scholar] [CrossRef] [PubMed]
- O’Neill, A.; et al. Open X-Embodiment: Robotic Learning Datasets and RT-X Models. In 2024 IEEE International Conference on Robotics and Automation (ICRA); IEEE, 2024; vol. 6, pp. 6892–6903. [Google Scholar]
- Rohrbach, S.; et al. Digitization and validation of a chemical synthesis literature database in the ChemPU. Science 2022, 377, 172–180. [Google Scholar] [CrossRef] [PubMed]
- Yuan, K.; Sajid, N.; Friston, K.; Li, Z. Hierarchical generative modelling for autonomous robots. Nat. Mach. Intell. 2023, 5, 1402–1414. [Google Scholar] [CrossRef]
- Mandal, I.; et al. Evaluating large language model agents for automation of atomic force microscopy. Nat. Commun. 2025, 16, 9104. [Google Scholar] [CrossRef] [PubMed]
- Ríos-García, M.; et al. AI scientists produce results without reasoning scientifically. arXiv 2026. [Google Scholar] [CrossRef]
- Kambhampati, S.; et al. LLMs can’t plan, but can help planning in LLM-Modulo Frameworks. arXiv 2024. [Google Scholar]
- Stechly, K.; Valmeekam, K.; Kambhampati, S. On the self-verification limitations of large language models on reasoning and planning tasks. In The Thirteenth International Conference on Learning Representations; 2025. [Google Scholar]
- Dai, T.; et al. Autonomous mobile robots for exploratory synthetic chemistry. Nature 2024, 635, 890–897. [Google Scholar] [CrossRef] [PubMed]
- Gostev, M.; et al. The BioSample Database (BioSD) at the European bioinformatics institute. Nucleic Acids Res. 2012, 40, D64–70. [Google Scholar] [CrossRef] [PubMed]
- Qian, W.; et al. Explainable mechanism for production process anomalies based on digital twin. Nat. Commun. 2026, 17, 1566. [Google Scholar] [CrossRef] [PubMed]
- Cooper, A. I.; et al. Accelerating discovery in natural science laboratories with AI and robotics: Perspectives and challenges. Sci. Robot. 2025, 10, eadv7932. [Google Scholar] [CrossRef] [PubMed]
- Bai, J.; et al. A dynamic knowledge graph approach to distributed self-driving laboratories. Nat. Commun. 2024, 15, 462. [Google Scholar] [CrossRef] [PubMed]
- Volk, A. A.; Abolhasani, M. Performance metrics to unleash the power of self-driving labs in chemistry and materials science. Nat. Commun. 2024, 15, 1378. [Google Scholar] [CrossRef] [PubMed]
- Tang, X.; et al. Risks of AI scientists: prioritizing safeguarding over autonomy. Nat. Commun. 2025, 16, 8317. [Google Scholar] [CrossRef] [PubMed]
- Zhu, K.; et al. SafeScientist: Enhancing AI scientist safety for risk-aware scientific discovery. In in Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing; Stroudsburg, PA, USA, Christodoulopoulos, C., Chakraborty, T., Rose, C., Peng, V., Eds.; Association for Computational Linguistics, 2025; pp. 2289–2317. [Google Scholar]
- Shapira, N.; et al. Agents of Chaos. arXiv 2026. [Google Scholar] [CrossRef]
- Zhao, Q.; et al. Uncertainty Propagation on LLM Agent. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics; Che, W., Nabende, J., Shutova, E., Pilehvar, M. T., Eds.; Association for Computational Linguistics: Stroudsburg, PA, USA, 2025; Volume 1, pp. 6064–6073. [Google Scholar]
- Duan, J.; et al. UProp: Investigating the uncertainty propagation of LLMs in multi-step agentic decision-making. arXiv 2025. [Google Scholar]
- Song, Y.; Jeong, M.; Sung, M. Trustworthy agents for Electronic Health Records through confidence estimation. arXiv 2025. [Google Scholar]
- Hendrycks, D.; et al. Measuring massive multitask language understanding. In International Conference on Learning Representations; 2021. [Google Scholar]
- Jin, Q.; Dhingra, B.; Liu, Z.; Cohen, W.; Lu, X. PubMedQA: A Dataset for Biomedical Research Question Answering. EMNLP 2019. [Google Scholar] [CrossRef]
- Bragg, J.; et al. AstaBench: Rigorous benchmarking of AI agents with a scientific research suite. arXiv 2025. [Google Scholar] [CrossRef]
- Laurent, J. M.; et al. LAB-bench: Measuring capabilities of language models for biology research. arXiv 2024. [Google Scholar]
- Mitchener, L.; et al. BixBench: A comprehensive benchmark for LLM-based agents in computational biology. arXiv 2025. [Google Scholar]
- Chen, Z.; et al. ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery. In The International Conference on Learning Representations; 2025. [Google Scholar]
- Nair, S.; et al. Agentic systems are adept at solving well-scoped, verifiable problems in computational biology. bioRxiv 2026. [Google Scholar] [CrossRef]
- Evaluating Claude’s bioinformatics research capabilities with BioMysteryBench. Available online: https://www.anthropic.com/research/Evaluating-Claude-For-Bioinformatics-With-BioMysteryBench.
- Li, J.; Ho, A. GeneBench: Assessing AI agents for multi-stage inference problems in genomics and quantitative biology. bioRxiv 2026. [Google Scholar] [CrossRef]
- De Brouwer, E.; et al. AssayBench: An assay-level virtual cell benchmark for LLMs and agents. arXiv 2026. [Google Scholar] [CrossRef]
- Ferreira da Silva, R.; et al. A grassroots network and community roadmap for interconnected autonomous science laboratories for accelerated discovery. In Workshop Proceedings of the 54th International Conference on Parallel Processing; ACM: New York, NY, USA, 2025; pp. 142–150. [Google Scholar]
Figure 3.
Failure spectrum and recovery in autonomous laboratories. Autonomous labs must detect failures (third column) and recover from them (rightmost column) across execution (top), sample state and identity (middle), and measurement/data layers (bottom). Telemetry, chain-of-custody tracking, digital twins, and automated QC can prevent silent failures from propagating into invalid analyses by triggering reruns, quarantine, recalibration, human escalation, or safe aborts.
Figure 3.
Failure spectrum and recovery in autonomous laboratories. Autonomous labs must detect failures (third column) and recover from them (rightmost column) across execution (top), sample state and identity (middle), and measurement/data layers (bottom). Telemetry, chain-of-custody tracking, digital twins, and automated QC can prevent silent failures from propagating into invalid analyses by triggering reruns, quarantine, recalibration, human escalation, or safe aborts.

Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.