Submitted:
04 August 2026
Posted:
12 August 2026
You are already at the latest version
Abstract
While CRISPR-Cas systems have emerged as transformative gene-editing technologies, conventional design strategies reliant on empirical rules and trial-and-error remain inefficient and cost-prohibitive. Artificial intelligence (AI) presents novel opportunities to enhance the precision, efficiency, and automation of CRISPR design. This review provides a systematic survey of AI-driven tools and methodologies utilized in this domain. We categorize these approaches by learning paradigm, discussing how five distinct families—traditional supervised learning, deep learning, attention-based models, generative AI and foundation models, and reinforcement learning—facilitate CRISPR design. Functionally, we classify these tools into eight categories: database search, structure and function prediction, sequence and structure generation, virtual screening, DNA synthesis, and experimental design. These tools have yielded significant outcomes across therapeutic, agricultural, and basic research settings. Despite persisting challenges regarding data quality, interpretability, experimental validation, and safety, advances in multimodal AI and personalized design are poised to significantly expand the impact of AI on precision medicine. Furthermore, this review offers practical guidelines for tool selection tailored to researchers with varying computational expertise, serving as an essential resource for the gene editing, computational biology, and translational medicine communities.
Keywords:
CRISPR-cas systems
; artificial intelligence
; bioinformatics tools
1. Introduction
1.1. Overview of CRISPR Systems
First identified in 1987 as atypical repetitive sequences in Escherichia coli, CRISPR-Cas systems were subsequently characterized as adaptive immune mechanisms in bacteria and archaea [1], heralding a new era in gene-editing technology. Following the identification and nomenclature of Cas genes in 2002 [2], the field progressed rapidly from mechanistic elucidation to engineering applications. Key milestones included the revelation of the homology between spacer sequences and phage or plasmid DNA in 2005 [3] and the experimental validation of CRISPR-mediated adaptive immunity in 2007 [4]. Notably, the repurposing of CRISPR-Cas9 in 2012 marked a pivotal transition from fundamental research to practical application [5].
Mediated by RNA-guided DNA recognition and cleavage, CRISPR-Cas systems facilitate genome editing with unprecedented precision. The available toolkit has evolved significantly from initial double-strand break (DSB)–dependent methods to encompass diverse editing modalities. Conventional Cas nucleases induce DSBs, relying on cellular repair pathways—specifically non-homologous end joining (NHEJ) or homology-directed repair (HDR)—to effectuate gene knockouts or insertions, thereby underpinning functional genomics and therapeutic strategies. However, because DSBs can result in unpredictable repair outcomes and genomic instability, there has been a drive toward higher-fidelity strategies [6,7]. Base editors (BEs) address the challenge of single-nucleotide editing by fusing a Cas9 nickase with deaminases, enabling C→T or A→G conversions without generating DSBss [8,9]. Prime editors (PEs) further expand the editing repertoire by combining a Cas9 nickase with reverse transcriptase; utilizing a prime-editing guide RNA (pegRNA) as a template, they achieve precise insertions, deletions, and substitutions in a DSB-independent manner [10]. Additionally, epigenetic editors (EEs) fuse catalytically dead Cas9 (dCas9) with epigenetic modifiers to modulate gene expression, offering robust tools for epigenetic research [11,12,13].
Figure 1.
Gene-editing modalities. Major CRISPR editing modes: Cas nuclease, base editor, prime editor, and epigenetic editor.
Figure 1.
Gene-editing modalities. Major CRISPR editing modes: Cas nuclease, base editor, prime editor, and epigenetic editor.

1.2. Challenges in CRISPR Design and Opportunities for AI
Despite their transformative potential, the practical application of CRISPR technologies encounters significant hurdles. Traditional design paradigms, often predicated on empirical heuristics and trial-and-error [14,15], are inherently inefficient, resource-intensive, and time-consuming. The design of single-guide RNAs (sgRNAs) necessitates a delicate balance among on-target efficiency, off-target risks, editing window positioning, and protospacer adjacent motif (PAM) constraints [16,17,18]. Similarly, engineering Cas proteins requires profound structure–function insights; however, conventional methodologies struggle to systematically navigate the vast sequence space [19,20,21]. Delivery optimization presents comparable complexities, involving intricate trade-offs among lipid nanoparticle (LNP) formulations, mRNA structural integrity, and tissue specificity [22,23,24].
As experimental scales and complexities expand, traditional methods prove inadequate for high-throughput design requirements [25]. Furthermore, clinical translation imposes rigorous standards for efficiency, specificity, and safety, necessitating robust multi-objective optimization [26,27]. This complexity delineates a critical niche for artificial intelligence (AI). It is important to note, however, that for straightforward targets, scenarios demanding high interpretability, or cases with extremely scarce labeled data, rule- and experience-based designs remain prevalent [28,29]. Consequently, AI does not wholly supplant traditional methods; rather, the two approaches often function synergistically [17,30,31].
Machine learning (ML) and deep learning (DL) excel at processing high-dimensional data, deciphering complex patterns, and executing multi-objective optimization. Within the context of CRISPR design, AI algorithms can elucidate the relationships between sgRNA efficiency and sequence features by mining large-scale experimental datasets [28,29], accurately predict off-target sites [30], optimize Cas protein structure and function [32,33], and even facilitate the de novo design of novel editing systems [34,35]. Moreover, generative models—such as diffusion models and variational autoencoders—offer innovative avenues for exploring sequence and structural landscapes. These computational approaches promise to accelerate the design cycle while enhancing success rates and predictability.
1.3. Aims and Organization of This Review
This review provides a comprehensive survey of AI-driven tools and methodologies for CRISPR system design, with a specific emphasis on applications spanning sgRNA design, off-target prediction, Cas protein engineering, delivery optimization, and gene-editing experimental planning. To ensure conceptual clarity, we distinguish between methods—defined here as algorithmic and modeling paradigms (e.g., supervised learning, deep learning)—and tools, which refer to the software, platforms, or databases that implement them. The mapping between these two categories is detailed in Section 4 and Section 5.
Figure 3.
Evolution of CRISPR design tools. Progression from the foundational (1987–2015) and computational-aided (2014–2021) phases to the AI-driven phase (2017–present), with representative tools, milestones, and technical features aligned to the toolkit classification in Table 2.
Figure 3.
Evolution of CRISPR design tools. Progression from the foundational (1987–2015) and computational-aided (2014–2021) phases to the AI-driven phase (2017–present), with representative tools, milestones, and technical features aligned to the toolkit classification in Table 2.

Reader Guide and Search Scope:
For an in-depth analysis of methodologies and algorithmic organization, please refer to Section 4 and Figure 4. For guidance on tool selection and functional taxonomy, see Section 5 (specifically Section 5.9) and Table 2. Clinical applications and case studies are discussed in Section 6, while Section 7 addresses current challenges and future outlooks. The literature and tools surveyed herein were identified through PubMed, Web of Science, Google Scholar, and bioRxiv using keywords such as CRISPR, artificial intelligence, sgRNA design, off-target prediction, and gene editing, covering publications through February 2026.
Scope of Coverage:
This review synthesizes five families of AI methods and eight categories of tools (encompassing dozens of representative examples listed in Table 2), highlighting their utility in therapy, agriculture, and basic research. For task-oriented navigation:
2. Technical Pipeline for AI-Driven CRISPR System Design
AI-driven CRISPR design follows a multi-stage, iteratively refined closed loop. Given an editing objective—target gene, site(s), and desired edit type—the workflow typically comprises four stages: component design, functional prediction, scheme optimization, and experimental validation.
Component design encompasses delivery systems, RNA sequences, and Cas proteins. Delivery design includes optimization of LNP or adeno-associated virus (AAV) vectors with respect to encapsulation efficiency, tissue specificity, and immunogenicity. RNA design covers mRNA and guide RNA (gRNA) optimization for stability, translation efficiency, and targeting specificity. Cas acquisition and optimization proceed along three paths: (i) AI-driven discovery of novel Cas systems using protein language models such as ESM [36]) and deep learning on metagenomic data; (ii) selection of known Cas variants from existing libraries according to PAM range, efficiency, and specificity; (iii) engineering of Cas via AI-assisted [32] protein design to extend PAM recognition, improve efficiency, or reduce size for delivery.
Functional prediction is where AI is central. Deep learning, graph neural networks, and related models predict sgRNA on-target efficiency, off-target risk, protein structural stability, and editing and repair outcomes [37,38], informing downstream optimization.
Scheme optimization uses predictions to tune parameters and components for specific editor types (BE, PE, EE), balancing efficiency, specificity, and safety—e.g., sgRNA adjustment, Cas variant choice, and editing window selection.
Experimental validation employs wet-lab methods such as GUIDE-seq [39] and CHANGE-seq [40] and AI-assisted analysis. Validation data feed model training and parameter updates, closing the loop. This data-driven iteration improves success rates and predictability and enables more systematic exploration of design space than rule-only approaches.
3. Evolution of CRISPR Design Tools
CRISPR design tools have evolved from foundational discovery [1,34] through computational assistance to AI-driven design. Chronologically and technically, three phases are distinguished: foundations (1987–2015), computational-aided design (2014–2021) [17], and AI-driven design (2017–present) [28], with 2017–2021 as an overlapping transition in which deep learning gained prominence alongside rule-based tools.
3.1. Foundational Phase (1987–2015)
Understanding CRISPR-Cas rested on nearly three decades of basic research. In 1987, Japanese researchers first observed unusual repeats in the E. coli genome [1]; the finding received limited attention but laid groundwork. In 2002, Cas genes were identified and named, linking repeats to Cas proteins [2]. In 2005, spacer–phage/plasmid DNA matches implied a role in adaptive immunity [3]. In 2007, CRISPR was shown experimentally to provide adaptive immunity by recognizing and clearing foreign DNA—a milestone for later applications [4].
In 2010, the Cascade complex (type I-E) clarified how multi-protein assemblies recognize and cleave DNA [41]. In 2012, engineered CRISPR-Cas9 began the modern gene-editing era [5]. The same year, Cpf1 [15] (later Cas12a) and Cas13 [42] broadened the toolkit—Cas12a with distinct PAM requirements and Cas13 for RNA targeting. In 2015, anti-CRISPR proteins revealed regulatory mechanisms and tools for temporal control of editing [43,44].
During this phase, effort focused on mechanism; design tools were not systematized and relied on experience and simple sequence rules [45].
3.2. Computational-Aided Phase (2014–2021)
As CRISPR spread, systematic computational tools for sgRNA design and analysis became essential. In 2014, early sgRNA tools appeared; CHOPCHOP v1 [46] marked the start of computational-aided design. MAGECK [47] and related tools enabled pooled CRISPR screen analysis. In 2015, GUIDE-seq [39] and similar methods linked experiments to predictive models and supplied high-quality training data. CRISPRscan [48] systematized rule-based scoring. From 2016 to 2017, Azimuth and CFD (cutting frequency determination) [17]standardized efficiency scoring; CRISPResso [49] standardized outcome analysis; CRISPOR [50,51] integrated multiple tools into a one-stop platform.
From 2020 to 2021, tools targeted specific editors: PrimeDesign [51] and BE-Hive [52] for prime and base editors; CHANGE-seq [40] expanded off-target data; SpG and SpRY [53] approached PAM-free targeting; CRISPick [54,55] standardized library design. This phase emphasized standardization and rule-based or shallow machine learning, setting the stage for AI-driven design.
3.3. AI-Driven Phase (2017–Present)
Transformers [56] in 2017 marked growing AI use in CRISPR design; attention mechanisms capture long-range sequence dependencies relevant to sgRNA efficiency.
From 2018 to 2020, deep learning advanced across tasks. DeepCRISPR [28] and DeepSpCas9 [57] demonstrated strong Cas9 sgRNA efficiency prediction from large datasets; tools diverged by effector—e.g., DeepCpf1 [57] for Cas12a and TIGER [58] for Cas13d alongside Cas9-oriented tools such as CRISPOR [50,51] (Table 2, C1d). Deep learning also improved off-target prediction [59].
Since 2021, generative AI and foundation models have created new opportunities. AlphaFold2 [33] and AlphaFold3 [60] enable accurate Cas structure prediction for engineering. Protein language models such as ESM [36] improve sequence representation and function prediction. Diffusion models [35] support de novo protein and variant design, shifting emphasis from prediction alone to generative design [34].
Current trends include multimodal integration, end-to-end pipelines, and automation—together indicating a growing role for AI in CRISPR system design [61].
4. AI Algorithms and Models in CRISPR Applications
AI toolkits have substantially reshaped protein and CRISPR system design. This section focuses on deep-learning–driven tools that infer sequence–structure–function relationships from data and capture patterns that traditional feature engineering often misses. Rule-based or shallow machine-learning tools such as Rule-set 2 [17] are noted but not emphasized. Deep learning here comprises two pillars (Figure 4): learning paradigms (left column)—how models learn sequence–structure–function mappings; model architectures (middle column)—how geometric and biological priors encode biophysical complexity. Together they support CRISPR tasks (right column, C1–C8). Section 4.1 and Section 4.2 summarize paradigms and architectures; representative methods and tools appear in Section 5 and Table 2.
4.1. Learning Paradigms
Generalizing patterns from large protein and sequence databases is central to intelligent CRISPR design; models typically implement one of three paradigms—supervised, unsupervised, or reinforcement learning (Figure 4, left column).
L1. Supervised learning trains models on labeled input–output pairs (Figure 4, L1).
Standard supervision. CRISPR AI tools often use assay-labeled data: sequence–function predictors take protein or sgRNA sequence as input and activity or efficiency as output (e.g., Rule Set 2 [17]); structure predictors such as AlphaFold2 [33] take sequence and PDB structures as labels. Deep learning and attention models (CNN [59], RNN [62], Transformer [56], DNABERT [63]) belong here and support end-to-end off-target prediction such as DeepCRISPR [28], Indel/repair outcome prediction, base-editor efficiency and window prediction [52], and multitask models integrating chromatin and cross–cell-type gRNA efficiency.
Label-efficient supervision. When labels are scarce or costly, transfer learning[65], zero-/few-shot learning [62], and active learning can improve data efficiency—relevant when CRISPR training data are limited [65].
L2. Unsupervised learning learns patterns from unlabeled data to obtain useful representations (Figure 4, L2). Common themes include the following (not exhaustive).
Language models use next-token prediction or masked language modeling (MLM); the former generates sequence stepwise [66], the latter infers masked tokens from context. Protein and nucleic acid language models (ESM [36], ProGen [67], DNABERT [63]) support Cas representation, function prediction, and sequence generation in CRISPR contexts.
Diffusion models learn a reverse noising process to recover structure or sequence [68]; training on unlabeled PDB structures, for example, enables unconditional or conditional generation of physically plausible proteins for Cas and editor design [69].
Variational autoencoders (VAEs) encode data into a distribution with a prior (often Gaussian) [70]; VQ-VAE uses discrete codes [71]. Together with diffusion and GANs [72], VAEs support exploration of sequence space for Cas design [73].
Contrastive learning distinguishes positive pairs from negatives (e.g., crops from the same protein image as positives) [74], aiding representation and multimodal alignment [75].
L3. Reinforcement learning (RL) learns long-horizon policies via environment interaction, reward [76], and state transitions (Figure 4, L3). Components may include:Policy: maps states to actions (stochastic).Value function: estimates expected return from a state.Model: predicts next state and reward for planning.
In protein design, EvoPlay [77] treats sequence as state and point mutations as actions; surrogate models (e.g., AlphaFold2 [33]) can approximate assay rewards. The agent combines policy, value, and sequence–function models to optimize predicted reward [78,79]. In CRISPR, RL suits gRNA library design, protocol optimization, and ordering of targets in multi-gene editing, but reward specification and sample complexity remain challenging [80].
4.2. Model Architectures
Beyond learning paradigms, model architecture encodes geometric and biological priors so that predictions respect intrinsic constraints (Figure 4, middle column) [81]. For example, rigid motions of a protein should not change biological behavior—geometric priors are built into the architecture [33]; biological priors (residue interactions, folding principles) encourage physically plausible outputs [82]. Without such priors, models may overfit superficial patterns and generalize poorly. Architectures are chosen for sequences, graphs, or 3D structure [83,84]; CRISPR-related tools often combine several [85]. Below we briefly summarize common architectures for protein and nucleic acid modeling.
M1. Convolutional neural networks (CNNs) detect local spatial patterns with approximate translation invariance [86] (Figure 4, M1). In vision and 1D sequence tasks they extract local motifs and window features [87]; in CRISPR they support sgRNA efficiency, editing windows, and end-to-end off-target prediction.
M2. Recurrent neural networks (RNNs) and LSTMs process sequences token by token while retaining memory of past tokens [89] (Figure 4, M2), suiting early sequence-to-function modeling.In protein bioinformatics they are employed to predict secondary structure and subcellular localization [90]; in genomics they capture long-range dependencies and regulatory motifs [91], such as the hybrid DanQ model for quantifying the function of DNA sequences.
M3. Transformers use self-attention to relate all token pairs in parallel rather than strictly left-to-right [57], improving long-range dependency modeling and coupling sequence termini (Figure 4, M3). In CRISPR they are widely used for gRNA/pegRNA efficiency, joint on/off-target prediction, and multitask models (e.g., DNABERT [63], TIGER [59], CRISP-RCNN [92], TransLNP [93]).
M4. Graph neural networks (GNNs) represent entities as nodes and relations as edges [94], where proteins are often modeled as contact graphs with residues as nodes [95] (Figure 4, M4). They support molecular property prediction [96], LNP lipid discovery [97], and protein–complex interface modeling [98] (Table 2: C2/C3/C6).
M5. Geometric and 3D networks operate on three-dimensional structure with explicit rotation and translation invariance [82], yielding orientation-independent predictions (Figure 4, M5). They underpin structure prediction and design (e.g., AlphaFold2/3 [33,61] , RFDiffusion [35], RF-AllAtom [99]) for Cas modeling and de novo design (Section 5: C2/C5).
In practice, researchers first choose a learning paradigm: interpretability or limited labels favor supervised learning (L1) and classical feature models [100]; exploration or de novo generation favors unsupervised and generative models (L2) or RL (L3) [101]. Architecture choice follows data type and compute [102]: long, context-sensitive tasks often use Transformers (M3); local motif and efficiency tasks often use CNNs (M1); molecules and complexes may use GNNs (M4) or geometric networks (M5). Mappings to specific tools appear in Section 5 and Table 2.
5. CRISPR Toolkit Classification and Applications
The AI-driven CRISPR toolkit has evolved from simple sequence-matching algorithms to complex generative systems. Based on the technical pipeline (Figure 2) and the mapping of learning paradigms (Figure 4), these tools are categorized into four functional clusters that bridge the gap between in silico design and experimental execution.
In this section, we categorize AI-driven CRISPR tools into eight functional modules (C1–C8), following the logical workflow of genome editing. A comprehensive summary of these tools, including their specific tasks, algorithmic foundations, and primary references, is provided in Table 2.
5.1. CRISPR-Cas Database Search and Component Identification(C1)
The primary challenge in CRISPR design is balancing on-target efficiency with genomic specificity. AI has redefined this stage by transitioning from rule-based heuristics to multi-dimensional, context-aware models.
- Component Discovery and PAM Characterization: Tools like CRISPRCasFinder [103] and Protein2PAM [107] have automated the identification of novel Cas systems and their PAM requirements from metagenomic data. By leveraging protein language models (PLMs), these tools can predict the targeting range of uncharacterized orthologs, significantly expanding the targetable genome.
- On-target Efficiency Prediction: Modern design platforms, such as DeepCRISPR [28] and TIGER [59], utilize CNNs and Transformers to capture complex sequence-activity relationships. Unlike early tools, these models integrate “contextual” features—including chromatin accessibility, DNA torsion, and epigenetic marks—to provide high-precision scoring for Cas9, Cas12a, and Cas13d systems.
- Specificity and Off-target Assessment: To mitigate the risk of unintended mutations, AI-driven specificity tools have moved beyond simple mismatch counting. Models like Elevation [30] and Cas-Offinder [109] employ machine learning to predict off-target landscapes across the entire genome. These computational predictions are increasingly integrated with high-throughput experimental data from assays such as GUIDE-seq [39] and CHANGE-seq [40], enabling a more rigorous safety profile for therapeutic applications.
5.2. Structural Elucidation and Functional Mapping (C2–C3)
Understanding the Cas-gRNA-DNA ternary complex is the prerequisite for rational engineering. AI has transformed this field from static crystallography to a multi-scale predictive pipeline, moving from individual components to functional ensembles.
- From Monomer Folding to Rapid Inference: The foundation of CRISPR structural modeling lies in predicting the three-dimensional fold of novel Cas proteins from their primary sequences. While AlphaFold2 [33] set the benchmark for atomic accuracy, protein language models like ESMFold [36] have enabled the structural annotation of millions of metagenomic Cas orthologs at unprecedented speeds, facilitating the discovery of compact and thermostable variants.
- Multi-entity Cofolding and RNP Assembly: The functional unit of CRISPR is the ribonucleoprotein (RNP) complex. The current “all-atom” frontier, led by AlphaFold3 [61] and RoseTTAFold All-Atom [99], allows for the simultaneous co-folding of proteins, gRNA, and target DNA. These models accurately position essential metal-ion cofactors (e.g., Mg2+) and small-molecule ligands, providing a “cleavage-competent” snapshot that is vital for designing chimeric editors like Base and Prime editors.
- Functional Hotspots and Regulatory Mapping: Beyond static scaffolds, geometric deep learning tools such as NucleicNet [122] and Metal3D [125] map the “hotspots” of molecular interaction. These models pinpoint DNA-binding interfaces and catalytic centers, enabling the rational design of high-fidelity mutants. Furthermore, AI predictors for post-translational modifications (PTMs), such as DeepMVP [127] and MusiteDeep [128], identify regulatory sites that govern Cas stability and nuclear localization in eukaryotic cells, offering a roadmap for optimizing editor performance in vivo.
5.3. Generative Synthesis and De Novo Design (C4–C5)
The most significant paradigm shift in CRISPR engineering is the move from predicting existing systems to generating novel, synthetic effectors [34]. This generative frontier explores a sequence-structure space far beyond natural evolution.
- Evolution-Guided Sequence Generation: Protein language models (PLMs) such as ESM-3 [133] and ProGen [129,130] treat amino acid sequences as a biological language, learning the underlying “grammar” of protein fitness. By sampling from these learned distributions, researchers can generate synthetic nucleases that maintain high functional activity while remaining evolutionarily distant from natural Cas9. Advanced genomic models like Evo [131,132] further extend this by co-designing the Cas protein alongside its associated non-coding RNA (gRNA) arrays.
- De Novo Structural Scaffolding: Diffusion-based models, led by RFDiffusion [138] and Chroma [139], have revolutionized the design of protein backbones from scratch. These tools can generate diverse and programmable architectures tailored to specific DNA targets. The recent advancement of RFDiffusionAA (All-Atom) [99] allows for the de novo design of protein pockets specifically coordinated around gRNA and metal-ion cofactors, enabling the creation of miniaturized or hyper-specific editors.
- Inverse Folding and Sequence-Structure Co-design: To bridge the gap between a designed 3D scaffold and a realizable sequence, “inverse folding” tools like ProteinMPNN [32], and LigandMPNN [135] identify amino acid sequences that fold into the target geometry with high stability. Integrated pipelines now support the simultaneous optimization of both sequence and structure, ensuring that the generated editors are physically stable and energetically favorable for ribonucleoprotein (RNP) assembly.
5.4. Developability Screening and Agentic Orchestration (C6–C8)
The final stage of the technical pipeline addresses the “delivery gap” and the complexity of experimental execution, ensuring that in silico designs translate effectively into in vivo outcomes.
- Virtual Screening for Function and Delivery: Before physical synthesis, candidates undergo multi-objective screening. Tools like EVOLVEpro [144] and AiCE [143] predict the functional impact of mutations on editing efficiency. Simultaneously, developability predictors such as LNP_ML [146] and COMET [147] navigate the vast chemical space of ionizable lipids to optimize mRNA delivery vehicles. For AAV-mediated delivery, Bryant et al. [152] utilized deep learning to design over 110,000 functional AAV2 capsid variants.
- Nucleic Acid Synthesis and Stability Optimization: To ensure robust expression of the editing machinery, DNA/RNA sequences must be optimized for synthesis and translation. LinearDesign [151] utilizes computational linguistics algorithms to identify mRNA sequences with optimal codon usage and high structural stability (Minimum Free Energy). This co-optimization significantly extends the half-life of the mRNA cargo, which is a critical factor for high-efficiency LNP-mediated delivery.
- Agentic Orchestration of Gene-Editing Experiments: The integration of the entire design-to-validation workflow is now facilitated by agentic frameworks like CRISPR-GPT [62]. By leveraging large language models (LLMs) and specialized toolkits (C1–C7), these systems can map natural-language biological goals into structured, multi-step experimental plans. This end-to-end orchestration reduces human error and democratizes access to complex genome engineering by providing automated guidance from target selection to protocol execution.
5.5. Current Bottlenecks
Despite these advances, critical challenges remain: (1) Data Bias: Most C1 tools are biased toward mammalian data, limiting generalization to non-model species. (2) Dynamics: Static models struggle to capture the conformational plasticity essential for the Cas catalytic cycle. (3) Validation: Generative models still produce “biological hallucinations,” necessitating extensive wet-lab validation to filter non-functional variants.
6. Case Studies: Quantifying the AI-Driven Paradigm Shift
The integration of AI into CRISPR workflows has fundamentally transitioned genome editing from stochastic discovery to deterministic engineering. This section analyzes three landmark cases where computational frameworks resolved long-standing bottlenecks in nuclease potency, vector capacity, and systemic delivery.
6.1. Generative Design of Novel Nucleases: The OpenCRISPR-1 Framework
Natural CRISPR-Cas systems are constrained by the evolutionary trajectories of their host prokaryotes. Ruffolo et al. [34] circumvented these biological limits by employing a Protein Language Model (PLM)-centric pipeline to design OpenCRISPR-1, a high-performance de novo editor.
- Computational Innovation: By operating in a latent sequence space, the generative model explored a landscape ~ 4.8-fold more diverse than extant metagenomic databases. This allowed for the identification of functional motifs that are evolutionarily distant from wild-type SpCas9.
- Performance Benchmarking: OpenCRISPR-1 incorporates ~ 400 mutations relative to its nearest natural orthologs. Despite this high structural divergence, it achieved a 95% reduction in off-target activity (0.32% vs. 6.1% for SpCas9) while maintaining a robust on-target indel rate of 56%.
- Bioinformatics Significance: This case demonstrates that PLMs can “hard-code” high fidelity into the primary sequence, bypassing the need for labor-intensive, iterative rational design or directed evolution.
6.2. Structure-Informed Optimization: Engineering Compact Fanzor Systems
The clinical utility of CRISPR is often restricted by the <4.7 kb cargo capacity of Adeno-Associated Virus (AAV) vectors. While eukaryotic Fanzor (Fz) proteins are naturally compact, their innate editing efficiency in mammalian cells is suboptimal. Li et al. [153] utilized AlphaFold3 and the EVOLVEpro framework to bridge this fitness gap.
- Structure-to-Function Mapping: By leveraging high-resolution structural predictions of the ribonucleoprotein (RNP) complex, the AI-guided evolution loop identified critical residues for stabilizing the ωRNA-DNA interface.
- Quantified Improvements: The resulting MmeFz2–ωRNA system exhibited a 32-fold mean increase in editing activity across 38 endogenous loci. Furthermore, AI-driven truncation reduced the ωRNA size by 30% while simultaneously enhancing total system activity by 20-fold.
- Translational Impact: This optimization enables the entire editing machinery to be packaged within a single AAV vector, resolving the efficiency loss and complex stoichiometry associated with dual-vector delivery systems.
6.3. Predictive Modeling for Delivery: Virtual Screening of Ionizable Lipids
The chemical space for ionizable lipids—the critical component of Lipid Nanoparticles (LNPs)—is astronomically large (>1010 combinations), rendering exhaustive physical screening computationally and experimentally intractable. Li et al. [154] integrated combinatorial chemistry with supervised machine learning to navigate this space.
- Algorithmic Efficiency: The team performed a virtual screen of ~ 40,000 candidates, but the high predictive power of the model allowed them to synthesize only 16 top-tier leads (0.04% of the virtual library).
- Validation & Hit Rate: The model achieved a 100% functional success rate in experimental assays. The lead lipid (119-23) demonstrated superior tissue-specific mRNA delivery compared to current industry standards (e.g., MC3 or SM-102).
- R&D Acceleration: This data-driven approach compressed the discovery timeline from the traditional 18–24 months to less than 6 months, illustrating how predictive modeling can “de-risk” the development of CRISPR delivery vehicles.
7. Challenges and Future Directions
7.1. Current Challenges
AI performance depends on data quality and quantity; CRISPR applications suffer from scarce high-quality labels, noisy high-throughput data, limited cross-experiment standardization, and modest cross-cell-type and cross-species generalization. Deep models are often opaque, complicating mechanism insight and regulatory review. Clinical translation demands experimental validation of predictions—costly and slow—with frequent prediction–experiment gaps. Ethics and regulation (e.g., germline editing, enhancement) and jurisdictional differences require robust safety assessment including off-target analysis.
7.2. Outlook
Progress will likely emphasize multimodal integration (sequence, structure, chromatin accessibility, histone marks, single-cell expression), real-time optimization (predict → validate → update models; online learning), and personalized design (patient genotype, cell type, tissue context). Research priorities include Cas engineering (PAM expansion, efficiency, specificity, miniaturization), improved off-target prediction, and smarter LNP/AAV design. Gaps remain for dedicated AI tools for prime and epigenetic editors, single-cell editing heterogeneity, and harmonized regulatory evidence standards. Interdisciplinary collaboration, data sharing, and ethical frameworks will be needed to translate AI-driven CRISPR safely from bench to clinic.
8. Limitations of This Review
Our literature and tool selection have limitations: search strategy (databases, keywords, time window) is summarized in Section 1.3 and may omit some recent work; taxonomy and representative tools involve subjective choices, and generative and RL tools evolve rapidly—readers should consult official documentation and primary literature. We do not comprehensively treat ethics and regulation beyond brief notes. Tool and dataset access: official sites or appendices as required by the target journal. Competing interests: The authors declare no competing interests.
9. Conclusion
AI contributions to CRISPR design can be summarized in three phases: 2014–2017 computational standardization and integration; 2017–2020 deep learning breakthroughs in sgRNA efficiency and off-target prediction; and from 2021 generative AI and foundation models enabling a shift from prediction to design. Five AI method families and eight tool categories span design through validation and experiment planning, with concrete applications in genetic disease, cancer, crops, and functional genomics. Challenges include data quality and generalization, interpretability, validation, and ethics; multimodal AI, real-time optimization, and personalized design are promising directions. Continued collaboration and governance will be essential to realize safe, effective translation. Deep integration of AI and CRISPR is reshaping discovery and development from target identification through clinical development and merits sustained attention from academia and industry.
Key Points
· Overview of artificial intelligence methods applied in CRISPR design.
· Evaluation of current computational tools for CRISPR-Cas systems.
· Discussion of diverse applications of AI-driven CRISPR technologies.
· Identification of critical computational challenges and future directions.
Data Availability
The comprehensive list of AI-driven CRISPR tools, related literature, and web server links discussed in this review is curated and continuously updated in our public GitHub repository, freely accessible at https://github.com/SamYangBio/papers_for_CRISPR-Cas_design_using_AI.
Biographical Note
Zhengxia Yang is a Bioinformatics Engineer at Reforgene Medicine, Guangzhou, China. His research primarily focuses on computational biology, CRISPR-based technologies, and the development of gene-editing therapeutics. Cheng Liu is a researcher at the Department of Cardiology, Guangzhou First People’s Hospital, South China University of Technology. His research interests include cardiology, cardio-oncology, and the molecular mechanisms of cardiovascular diseases.
Author Contributions
All authors contributed equally.
Funding
This research received no external funding.
Conflicts of Interest
The authors declare no competing interests.
References
- Ishino, Y.; Shinagawa, H.; Makino, K.; et al. Nucleotide sequence of the iap gene, responsible for alkaline phosphatase isozyme conversion in Escherichia coli, and identification of the gene product. J. Bacteriol. 1987, 169, 5429–33. [Google Scholar] [CrossRef] [PubMed]
- Ruud, Jansen; JanDAV, Embden; Wim, Gaastra; et al. Identification of genes that are associated with DNA repeats in prokaryotes. Mol. Microbiol. 2002, 43, 1565–75. [Google Scholar] [CrossRef] [PubMed]
- Mojica, F.J.M.; Díez-Villaseñor; García-Martínez; et al. Intervening Sequences of Regularly Spaced Prokaryotic Repeats Derive from Foreign Genetic Elements. J. Mol. Evol. 2005, 60, 174–82. [Google Scholar] [CrossRef] [PubMed]
- Barrangou, R.; Fremaux, C.; Deveau, H.; et al. CRISPR Provides Acquired Resistance Against Viruses in Prokaryotes. Science 2007, 315, 1709–12. [Google Scholar] [CrossRef] [PubMed]
- Jinek, M.; Chylinski, K.; Fonfara, I.; et al. A Programmable Dual-RNA–Guided DNA Endonuclease in Adaptive Bacterial Immunity. Science 2012, 337, 816–21. [Google Scholar] [CrossRef] [PubMed]
- Cong, L.; Ran, F.A.; Cox, D.; et al. Multiplex Genome Engineering Using CRISPR/Cas Systems. Science 2013, 339, 819–23. [Google Scholar] [CrossRef] [PubMed]
- Mali, P.; Yang, L.; Esvelt, K.M.; et al. RNA-Guided Human Genome Engineering via Cas9. Science 2013, 339, 823–6. [Google Scholar] [CrossRef] [PubMed]
- Komor, A.C.; Kim, Y.B.; Packer, M.S.; et al. Programmable editing of a target base in genomic DNA without double-stranded DNA cleavage. Nature 2016, 533, 420–4. [Google Scholar] [CrossRef] [PubMed]
- Gaudelli, N.M.; Komor, A.C.; Rees, H.A.; et al. Programmable base editing of A•T to G•C in genomic DNA without DNA cleavage. Nature 2017, 551, 464–71. [Google Scholar] [CrossRef] [PubMed]
- Anzalone, A.V.; Randolph, P.B.; Davis, J.R.; et al. Search-and-replace genome editing without double-strand breaks or donor DNA. Nature 2019, 576, 149–57. [Google Scholar] [CrossRef] [PubMed]
- Gilbert, L.A.; Larson, M.H.; Morsut, L.; et al. CRISPR-Mediated Modular RNA-Guided Regulation of Transcription in Eukaryotes. Cell 2013, 154, 442–51. [Google Scholar] [CrossRef] [PubMed]
- Hilton, I.B.; D’Ippolito, A.M.; Vockley, C.M.; et al. Epigenome editing by a CRISPR-Cas9-based acetyltransferase activates genes from promoters and enhancers. Nat. Biotechnol. 2015, 33, 510–7. [Google Scholar] [CrossRef] [PubMed]
- Liu, X.S.; Wu, H.; Ji, X.; et al. Editing DNA Methylation in the Mammalian Genome. Cell 2016, 167, 233–247.e17. [Google Scholar] [CrossRef] [PubMed]
- Hsu, P.D.; Lander, E.S.; Zhang, F. Development and Applications of CRISPR-Cas9 for Genome Engineering. Cell 2014, 157, 1262–78. [Google Scholar] [CrossRef] [PubMed]
- Zetsche, B.; Gootenberg, J.S.; Abudayyeh, O.O.; et al. Cpf1 Is a Single RNA-Guided Endonuclease of a Class 2 CRISPR-Cas System. Cell 2015, 163, 759–71. [Google Scholar] [CrossRef] [PubMed]
- Fu, Y.; Foden, J.A.; Khayter, C.; et al. High-frequency off-target mutagenesis induced by CRISPR-Cas nucleases in human cells. Nat. Biotechnol. 2013, 31, 822–6. [Google Scholar] [CrossRef] [PubMed]
- Doench, J.G.; Fusi, N.; Sullender, M.; et al. Optimized sgRNA design to maximize activity and minimize off-target effects of CRISPR-Cas9. Nat. Biotechnol. 2016, 34, 184–91. [Google Scholar] [CrossRef] [PubMed]
- Hsu, P.D.; Scott, D.A.; Weinstein, J.A.; et al. DNA targeting specificity of RNA-guided Cas9 nucleases. Nat. Biotechnol. 2013, 31, 827–32. [Google Scholar] [CrossRef] [PubMed]
- Kleinstiver, B.P.; Prew, M.S.; Tsai, S.Q.; et al. Engineered CRISPR-Cas9 nucleases with altered PAM specificities. Nature 2015, 523, 481–5. [Google Scholar] [CrossRef] [PubMed]
- Nishimasu, H.; Ran, F.A.; Hsu, P.D.; et al. Crystal Structure of Cas9 in Complex with Guide RNA and Target DNA. Cell 2014, 156, 935–49. [Google Scholar] [CrossRef] [PubMed]
- Slaymaker, I.M.; Gao, L.; Zetsche, B.; et al. Rationally engineered Cas9 nucleases with improved specificity. Science 2016, 351, 84–8. [Google Scholar] [CrossRef] [PubMed]
- Kauffman, K.J.; Dorkin, J.R.; Yang, J.H.; et al. Optimization of Lipid Nanoparticle Formulations for mRNA Delivery in Vivo with Fractional Factorial and Definitive Screening Designs. Nano Lett. 2015, 15, 7300–6. [Google Scholar] [CrossRef] [PubMed]
- Liu, C.; Zhang, L.; Liu, H.; et al. Delivery strategies of the CRISPR-Cas9 gene-editing system for therapeutic applications. J. Control. Release 2017, 266, 17–26. [Google Scholar] [CrossRef] [PubMed]
- Finn, J.D.; Smith, A.R.; Patel, M.C.; et al. A Single Administration of CRISPR/Cas9 Lipid Nanoparticles Achieves Robust and Persistent In Vivo Genome Editing. Cell Rep. 2018, 22, 2227–35. [Google Scholar] [CrossRef] [PubMed]
- Shalem, O.; Sanjana, N.E.; Hartenian, E.; et al. Genome-Scale CRISPR-Cas9 Knockout Screening in Human Cells. Science 2014, 343, 84–7. [Google Scholar] [CrossRef] [PubMed]
- Gillmore, J.D.; Gane, E.; Taubel, J.; et al. CRISPR-Cas9 In Vivo Gene Editing for Transthyretin Amyloidosis. N Engl. J. Med. 2021, 385, 493–502. [Google Scholar] [CrossRef] [PubMed]
- Doudna, J.A. The promise and challenge of therapeutic genome editing. Nature 2020, 578, 229–36. [Google Scholar] [CrossRef] [PubMed]
- Chuai, G.; Ma, H.; Yan, J.; et al. DeepCRISPR: optimized CRISPR guide RNA design by deep learning. Genome Biol. 2018, 19, 80. [Google Scholar] [CrossRef] [PubMed]
- Wang, D.; Zhang, C.; Wang, B.; et al. Optimized CRISPR guide RNA design for two high-fidelity Cas9 variants by deep learning. Nat. Commun. 2019, 10, 4284. [Google Scholar] [CrossRef] [PubMed]
- Listgarten, J.; Weinstein, M.; Kleinstiver, B.P.; et al. Prediction of off-target activities for the end-to-end design of CRISPR guide RNAs. Nat. BioMed Eng. 2018, 2, 38–47. [Google Scholar] [CrossRef] [PubMed]
- Lee, M. Deep learning in CRISPR-Cas systems: a review of recent studies. Front Bioeng. Biotechnol. 2023, 11, 1226182. [Google Scholar] [CrossRef] [PubMed]
- Dauparas, J.; Anishchenko, I.; Bennett, N.; et al. Robust deep learning–based protein sequence design using ProteinMPNN. Science 2022, 378, 49–56. [Google Scholar] [CrossRef] [PubMed]
- Jumper, J.; Evans, R.; Pritzel, A.; et al. Highly accurate protein structure prediction with AlphaFold. Nature 2021, 596, 583–9. [Google Scholar] [CrossRef] [PubMed]
- Ruffolo, J.A.; Nayfach, S.; Gallagher, J.; et al. Design of highly functional genome editors by modelling CRISPR–Cas sequences. Nature 2025, 645, 518–25. [Google Scholar] [CrossRef] [PubMed]
- Watson, J.L.; Juergens, D.; Bennett, N.R.; et al. De novo design of protein structure and function with RFdiffusion. Nature 2023, 620, 1089–100. [Google Scholar] [CrossRef] [PubMed]
- Lin, Z.; Akin, H.; Rao, R.; et al. Evolutionary-scale prediction of atomic-level protein structure with a language model. Science 2023, 379, 1123–30. [Google Scholar] [CrossRef] [PubMed]
- Shen, M.W.; Arbab, M.; Hsu, J.Y.; et al. Predictable and precise template-free CRISPR editing of pathogenic variants. Nature 2018, 563, 646–51. [Google Scholar] [CrossRef] [PubMed]
- Leenay, R.T.; Maksimchuk, K.R.; Slotkowski, R.A.; et al. Identifying and Visualizing Functional PAM Diversity across CRISPR-Cas Systems. Mol. Cell 2016, 62, 137–47. [Google Scholar] [CrossRef] [PubMed]
- Tsai, S.Q.; Zheng, Z.; Nguyen, N.T.; et al. GUIDE-seq enables genome-wide profiling of off-target cleavage by CRISPR-Cas nucleases. Nat. Biotechnol. 2015, 33, 187–97. [Google Scholar] [CrossRef] [PubMed]
- Lazzarotto, C.R.; Malinin, N.L.; Li, Y.; et al. CHANGE-seq reveals genetic and epigenetic effects on CRISPR–Cas9 genome-wide activity. Nat. Biotechnol. 2020, 38, 1317–27. [Google Scholar] [CrossRef] [PubMed]
- Jore, M.M.; Lundgren, M.; Van Duijn, E.; et al. Structural basis for CRISPR RNA-guided DNA recognition by Cascade. Nat. Struct. Mol. Biol. 2011, 18, 529–36. [Google Scholar] [CrossRef] [PubMed]
- Abudayyeh, O.O.; Gootenberg, J.S.; Konermann, S.; et al. C2c2 is a single-component programmable RNA-guided RNA-targeting CRISPR effector. Science 2016, 353, aaf5573. [Google Scholar] [CrossRef] [PubMed]
- Bondy-Denomy, J.; Pawluk, A.; Maxwell, K.L.; et al. Bacteriophage genes that inactivate the CRISPR/Cas bacterial immune system. Nature 2013, 493, 429–32. [Google Scholar] [CrossRef] [PubMed]
- Pawluk, A.; Amrani, N.; Zhang, Y.; et al. Naturally Occurring Off-Switches for CRISPR-Cas9. Cell 2016, 167, 1829–1838.e9. [Google Scholar] [CrossRef] [PubMed]
- Doench, J.G.; Hartenian, E.; Graham, D.B.; et al. Rational design of highly active sgRNAs for CRISPR-Cas9–mediated gene inactivation. Nat. Biotechnol. 2014, 32, 1262–7. [Google Scholar] [CrossRef] [PubMed]
- Montague, T.G.; Cruz, J.M.; Gagnon, J.A.; et al. CHOPCHOP: a CRISPR/Cas9 and TALEN web tool for genome editing. Nucleic Acids Res. 2014, 42, W401–7. [Google Scholar] [CrossRef] [PubMed]
- Li, W.; Xu, H.; Xiao, T.; et al. MAGeCK enables robust identification of essential genes from genome-scale CRISPR/Cas9 knockout screens. Genome Biol. 2014, 15, 554. [Google Scholar] [CrossRef] [PubMed]
- Moreno-Mateos, M.A.; Vejnar, C.E.; Beaudoin, J.-D.; et al. CRISPRscan: designing highly efficient sgRNAs for CRISPR-Cas9 targeting in vivo. Nat. Methods 2015, 12, 982–8. [Google Scholar] [CrossRef] [PubMed]
- Pinello, L.; Canver, M.C.; Hoban, M.D.; et al. Analyzing CRISPR genome-editing experiments with CRISPResso. Nat. Biotechnol. 2016, 34, 695–7. [Google Scholar] [CrossRef] [PubMed]
- Haeussler, M.; Schönig, K.; Eckert, H.; et al. Evaluation of off-target and on-target scoring algorithms and integration into the guide RNA selection tool CRISPOR. Genome Biol. 2016, 17, 148. [Google Scholar] [CrossRef] [PubMed]
- Concordet, J.-P.; Haeussler, M. CRISPOR: intuitive guide selection for CRISPR/Cas9 genome editing experiments and screens. Nucleic Acids Res. 2018, 46, W242–5. [Google Scholar] [CrossRef] [PubMed]
- Hsu, J.Y.; Grünewald, J.; Szalay, R.; et al. PrimeDesign software for rapid and simplified design of prime editing guide RNAs. Nat. Commun. 2021, 12, 1034. [Google Scholar] [CrossRef] [PubMed]
- Arbab, M.; Shen, M.W.; Mok, B.; et al. Determinants of Base Editing Outcomes from Target Library Analysis and Machine Learning. Cell 2020, 182, 463–480.e30. [Google Scholar] [CrossRef] [PubMed]
- Walton, R.T.; Christie, K.A.; Whittaker, M.N.; et al. Unconstrained genome targeting with near-PAMless engineered CRISPR-Cas9 variants. Science 2020, 368, 290–6. [Google Scholar] [CrossRef] [PubMed]
- Drepanos, L.M.; Srikanth, S.; Kaplan, E.G.; et al. Balancing off-target and on-target considerations for optimized Cas9 CRISPR knockout library design. Prepr. Genom. 2025. [Google Scholar] [CrossRef]
- Srikanth, S.; Zheng, F.; Drepanos, L.M.; et al. Optimized parameters for Cas9 CRISPR interference library design. Prepr. Bioeng. 2026. [Google Scholar] [CrossRef]
- Vaswani, A.; Shazeer, N.; Parmar, N.; et al. Attention Is All You Need, version 7. Prepr. arXiv 2017. [Google Scholar] [CrossRef]
- Kim, H.K.; Min, S.; Song, M.; et al. Deep learning improves prediction of CRISPR–Cpf1 guide RNA activity. Nat. Biotechnol. 2018, 36, 239–41. [Google Scholar] [CrossRef] [PubMed]
- Wessels, H.-H.; Stirn, A.; Méndez-Mancilla, A.; et al. Prediction of on-target and off-target activity of CRISPR–Cas13d guide RNAs using deep learning. Nat. Biotechnol. 2024, 42, 628–37. [Google Scholar] [CrossRef] [PubMed]
- Lin, J.; Wong, K.-C. Off-target predictions in CRISPR-Cas9 gene editing using deep learning. Bioinformatics 2018, 34, i656–63. [Google Scholar] [CrossRef] [PubMed]
- Abramson, J.; Adler, J.; Dunger, J.; et al. Accurate structure prediction of biomolecular interactions with AlphaFold 3. Nature 2024, 630, 493–500. [Google Scholar] [CrossRef] [PubMed]
- Qu, Y.; Huang, K.; Yin, M.; et al. CRISPR-GPT for agentic automation of gene-editing experiments. Nat. BioMed Eng. 2025, 10, 245–58. [Google Scholar] [CrossRef] [PubMed]
- Toufikuzzaman, M.; Hassan Samee, M.A.; Sohel Rahman, M. CRISPR-DIPOFF: an interpretable deep learning approach for CRISPR Cas-9 off-target prediction. Brief. Bioinform. 2024, 25, bbad530. [Google Scholar] [CrossRef] [PubMed]
- Ji, Y.; Zhou, Z.; Liu, H.; et al. DNABERT: pre-trained Bidirectional Encoder Representations from Transformers model for DNA-language in genome. Bioinformatics 2021, 37, 2112–20. [Google Scholar] [CrossRef] [PubMed]
- Elkayam, S.; Tziony, I.; Orenstein, Y. DeepCRISTL: deep transfer learning to predict CRISPR/Cas9 on-target editing efficiency in specific cellular contexts. Bioinformatics 2024, 40, btae481. [Google Scholar] [CrossRef] [PubMed]
- Du, Q.; Wang, H.; Jiang, B.; et al. Advancing genetic engineering with active learning: theory, implementations and potential opportunities. Brief. Bioinform. 2025, 26, bbaf286. [Google Scholar] [CrossRef] [PubMed]
- Ofer, D.; Brandes, N.; Linial, M. The language of proteins: NLP, machine learning & protein sequences. Comput. Struct. Biotechnol. J. 2021, 19, 1750–8. [Google Scholar] [CrossRef] [PubMed]
- Madani, A.; Krause, B.; Greene, E.R.; et al. Large language models generate functional protein sequences across diverse families. Nat. Biotechnol. 2023, 41, 1099–106. [Google Scholar] [CrossRef] [PubMed]
- Guo, Z.; Liu, J.; Wang, Y.; et al. Diffusion models in bioinformatics and computational biology. Nat. Rev. Bioeng. 2023, 2, 136–54. [Google Scholar] [CrossRef] [PubMed]
- Pindi, C.; Palermo, G. Computation and deep-learning-driven advances in CRISPR genome editing. Nat. Struct. Mol. Biol. 2026, 33, 203–14. [Google Scholar] [CrossRef] [PubMed]
- Kingma, D.P.; Welling, M. Auto-Encoding Variational Bayes, version 11. Prepr. arXiv 2013. [Google Scholar] [CrossRef]
- van den, Oord A; Vinyals, O.; Kavukcuoglu, K. Neural Discrete Representation Learning, version 2. Preprint; arXiv, 2017. [Google Scholar] [CrossRef]
- Repecka, D.; Jauniskis, V.; Karpus, L.; et al. Expanding functional protein sequence space using generative adversarial networks. Prepr. Synth. Biol. 2019. [Google Scholar] [CrossRef]
- Riesselman, A.J.; Ingraham, J.B.; Marks, D.S. Deep generative models of genetic variation capture the effects of mutations. Nat. Methods 2018, 15, 816–22. [Google Scholar] [CrossRef] [PubMed]
- Chen, T.; Kornblith, S.; Norouzi, M.; et al. A Simple Framework for Contrastive Learning of Visual Representations, version 3. Preprint. arXiv 2020. [Google Scholar] [CrossRef]
- Zhang, Z.; Xu, M.; Jamasb, A.; et al. Protein Representation Learning by Geometric Structure Pretraining, version 5. In Preprint, arXiv; 2022. [Google Scholar] [CrossRef]
- Sutton, R.S.; Barto, A.G. Reinforcement Learning: An Introduction, 2nd ed.; MIT Press: Cambridge, MA, 2018. [Google Scholar]
- Wang, Y.; Tang, H.; Huang, L.; et al. Self-play reinforcement learning guides protein engineering. Nat. Mach. Intell. 2023, 5, 845–60. [Google Scholar] [CrossRef]
- Angermueller, C.; Dohan, D.; Belanger, D.; et al. Model-based reinforcement learning for biological sequence design, paper delivered at ICLR 2020. International Conference on Learning Representations 2020, 27 Mar. 2026, date last accessed; Available online: https://openreview.net/forum?id=HklxbgBKvr.
- Arora, D.; Mishra, D.C.; Budhlakoti, N.; et al. Introduction of Reinforcement Learning in Bioinformatics. Biot. Today 2018, 8, 25. [Google Scholar] [CrossRef]
- Kim, M.; Go, M.; Kang, S.-H.; et al. Revolutionizing CRISPR technology with artificial intelligence. Exp. Mol. Med. 2025, 57, 1419–31. [Google Scholar] [CrossRef] [PubMed]
- Bronstein, M.M.; Bruna, J.; Cohen, T.; et al. Geometric Deep Learning: Grids, Groups, Graphs, Geodesics, and Gauges, version 2. Preprint; arXiv, 2021. [Google Scholar] [CrossRef]
- Baek, M.; DiMaio, F.; Anishchenko, I.; et al. Accurate prediction of protein structures and interactions using a three-track neural network. Science 2021, 373, 871–6. [Google Scholar] [CrossRef] [PubMed]
- Jing, B.; Eismann, S.; Suriana, P.; et al. Learning from Protein Structure with Geometric Vector Perceptrons, version 3. Preprint; arXiv, 2020. [Google Scholar] [CrossRef]
- Ingraham, J.; Garg, V.; Barzilay, R.; et al. Generative Models for Graph-Based Protein Design. In Advances in Neural Information Processing Systems; Wallach, H., Larochelle, H., Beygelzimer, A., d’Alché-Buc, F., Fox, E., Garnett, R., Eds.; Curran Associates, Inc.: n.p., 2019; vol. 32, Available online: https://proceedings.neurips.cc/paper_files/paper/2019/file/f3a4ff4839c56a5f460c88cce3666a2b-Paper.pdf.
- Lee, M. Deep learning in CRISPR-Cas systems: a review of recent studies. Front Bioeng. Biotechnol. 2023, 11, 1226182. [Google Scholar] [CrossRef] [PubMed]
- LeCun, Y.; Bengio, Y.; Hinton, G. Deep learning. Nature 2015, 521, 436–44. [Google Scholar] [CrossRef] [PubMed]
- Alipanahi, B.; Delong, A.; Weirauch, M.T.; et al. Predicting the sequence specificities of DNA- and RNA-binding proteins by deep learning. Nat. Biotechnol. 2015, 33, 831–8. [Google Scholar] [CrossRef] [PubMed]
- Hochreiter, S.; Schmidhuber, J. Long Short-Term Memory. Neural Comput. 1997, 9, 1735–80. [Google Scholar] [CrossRef] [PubMed]
- Almagro Armenteros, J.J.; Sønderby, C.K.; Sønderby, S.K.; et al. DeepLoc: prediction of protein subcellular localization using deep learning. Bioinformatics 2017, 33, 3387–95. [Google Scholar] [CrossRef] [PubMed]
- Quang, D.; Xie, X. DanQ: a hybrid convolutional and recurrent deep neural network for quantifying the function of DNA sequences. Nucleic Acids Res. 2016, 44, e107–e107. [Google Scholar] [CrossRef] [PubMed]
- Vora, D.S.; Yadav, S.; Sundar, D. Hybrid Multitask Learning Reveals Sequence Features Driving Specificity in the CRISPR/Cas9 System. Biomolecules 2023, 13, 641. [Google Scholar] [CrossRef] [PubMed]
- Wu, K.; Yang, X.; Wang, Z.; et al. Data-balanced transformer for accelerated ionizable lipid nanoparticles screening in mRNA delivery. Brief. Bioinform. 2024, 25, bbae186. [Google Scholar] [CrossRef] [PubMed]
- Scarselli, F.; Gori, M.; Tsoi, Ah Chung; et al. The Graph Neural Network Model. IEEE Trans. Neural Netw. 2009, 20, 61–80. [Google Scholar] [CrossRef] [PubMed]
- Fout, A.; Byrd, J.; Shariat, B.; et al. Protein Interface Prediction using Graph Convolutional Networks. In Advances in Neural Information Processing Systems; Guyon, I., Von Luxburg, U., Bengio, S., Wallach, H., Fergus, R., Vishwanathan, S., Garnett, R., Eds.; Curran Associates, Inc.: n.p., 2017; vol. 30, Available online: https://proceedings.neurips.cc/paper_files/paper/2017/file/f507783927f2ec2737ba40afbd17efb5-Paper.pdf.
- Duvenaud, D.; Maclaurin, D.; Aguilera-Iparraguirre, J.; et al. Convolutional Networks on Graphs for Learning Molecular Fingerprints, version 2. Preprint; arXiv, 2015. [Google Scholar] [CrossRef]
- Sela, M.; Chen, G.; Kadosh, H.; et al. AI-Validated Brain Targeted mRNA Lipid Nanoparticles with Neuronal Tropism. ACS Nano 2025, 19, 36106–28. [Google Scholar] [CrossRef] [PubMed]
- Gainza, P.; Sverrisson, F.; Monti, F.; et al. Deciphering interaction fingerprints from protein molecular surfaces using geometric deep learning. Nat. Methods 2020, 17, 184–92. [Google Scholar] [CrossRef] [PubMed]
- Krishna, R.; Wang, J.; Ahern, W.; et al. Generalized biomolecular modeling and design with RoseTTAFold All-Atom. Science 2024, 384, eadl2528. [Google Scholar] [CrossRef]
- Eraslan, G.; Avsec, Ž.; Gagneur, J.; et al. Deep learning: new computational modelling techniques for genomics. Nat. Rev. Genet 2019, 20, 389–403. [Google Scholar] [CrossRef] [PubMed]
- Ching, T.; Himmelstein, D.S.; Beaulieu-Jones, B.K.; et al. Opportunities and obstacles for deep learning in biology and medicine. J. R Soc. Interface 2018, 15, 20170387. [Google Scholar] [CrossRef] [PubMed]
- Sapoval, N.; Aghazadeh, A.; Nute, M.G.; et al. Current progress and open challenges for applying deep learning across the biosciences. Nat. Commun. 2022, 13, 1728. [Google Scholar] [CrossRef] [PubMed]
- Couvin, D.; Bernheim, A.; Toffano-Nioche, C.; et al. CRISPRCasFinder, an update of CRISRFinder, includes a portable version, enhanced performance and integrates search for Cas proteins. Nucleic Acids Res. 2018, 46, W246–51. [Google Scholar] [CrossRef] [PubMed]
- Pourcel, C.; Touchon, M.; Villeriot, N.; et al. CRISPRCasdb a successor of CRISPRdb containing CRISPR arrays and cas genes from complete genome sequences, and tools to download and query lists of repeats and spacers. Nucleic Acids Res. 2019, gkz915. [Google Scholar] [CrossRef] [PubMed]
- Zhang, F.; Zhao, S.; Ren, C.; et al. CRISPRminer is a knowledge base for exploring CRISPR-Cas systems in microbe and phage interactions. Commun. Biol. 2018, 1, 180. [Google Scholar] [CrossRef] [PubMed]
- Li, W.; Jiang, X.; Wang, W.; et al. Discovering CRISPR-Cas system with self-processing pre-crRNA capability by foundation models. Nat. Commun. 2024, 15, 10024. [Google Scholar] [CrossRef] [PubMed]
- Nayfach, S.; Bhatnagar, A.; Novichkov, A.; et al. Customizing CRISPR–Cas PAM specificity with protein language models. Nat. Biotechnol. published online. 2026. [Google Scholar] [CrossRef] [PubMed]
- Qi, C.; Shen, X.; Li, B.; et al. PAMPHLET: PAM Prediction HomoLogous-Enhancement Toolkit for precise PAM prediction in CRISPR-Cas systems. J. Genet. Genom. 2025, 52, 258–68. [Google Scholar] [CrossRef] [PubMed]
- Bae, S.; Park, J.; Kim, J.-S. Cas-OFFinder: a fast and versatile algorithm that searches for potential off-target sites of Cas9 RNA-guided endonucleases. Bioinformatics 2014, 30, 1473–5. [Google Scholar] [CrossRef] [PubMed]
- Lei, Z.; Meng, H.; Lv, Z.; et al. Detect-seq reveals out-of-protospacer editing and target-strand editing by cytosine base editors. Nat. Methods 2021, 18, 643–51. [Google Scholar] [CrossRef] [PubMed]
- Kim, D.; Bae, S.; Park, J.; et al. Digenome-seq: genome-wide profiling of CRISPR-Cas9 off-target effects in human cells. Nat. Methods 2015, 12, 237–43. [Google Scholar] [CrossRef] [PubMed]
- Evans, R.; O’Neill, M.; Pritzel, A.; et al. Protein complex prediction with AlphaFold-Multimer. Prepr. Bioinform. 2021. [Google Scholar] [CrossRef]
- Varadi, M.; Bertoni, D.; Magana, P.; et al. AlphaFold Protein Structure Database in 2024: providing structure coverage for over 214 million protein sequences. Nucleic Acids Res. 2024, 52, D368–75. [Google Scholar] [CrossRef] [PubMed]
- Passaro, S.; Corso, G.; Wohlwend, J.; et al. Boltz-2: Towards Accurate and Efficient Binding Affinity Prediction. Prepr. Mol. Biol. 2025. [Google Scholar] [CrossRef] [PubMed]
- Chai Discovery Team; Boitreaud, J.; Dent, J.; et al. Zero-shot antibody design in a 24-well plate. Prepr. Synth. Biol. 2025. [Google Scholar] [CrossRef]
- Delgado, J.; Reche, R.; Cianferoni, D.; et al. FoldX force field revisited, an improved version. Bioinformatics 2025, 41, btaf064. [Google Scholar] [CrossRef] [PubMed]
- Dieckhaus, H.; Brocidiacono, M.; Randolph, N.Z.; et al. Transfer learning to leverage larger datasets for improved prediction of protein stability changes. Proc. Natl. Acad. Sci. USA 2024, 121, e2314853121. [Google Scholar] [CrossRef] [PubMed]
- Pronk, S.; Páll, S.; Schulz, R.; et al. GROMACS 4.5: a high-throughput and highly parallel open source molecular simulation toolkit. Bioinformatics 2013, 29, 845–54. [Google Scholar] [CrossRef] [PubMed]
- Case, D.A.; Aktulga, H.M.; Belfon, K.; et al. AmberTools. J. Chem. Inf. Model 2023, 63, 6183–91. [Google Scholar] [CrossRef] [PubMed]
- Case, D.A.; Cerutti, D.S.; Cruzeiro, V.W.D.; et al. Recent Developments in Amber Biomolecular Simulations. J. Chem. Inf. Model 2025, 65, 7835–43. [Google Scholar] [CrossRef] [PubMed]
- Yan, J.; Kurgan, L. DRNApred, fast sequence-based method that accurately predicts and discriminates DNA- and RNA-binding residues. Nucleic Acids Res. 2017, gkx059. [Google Scholar] [CrossRef] [PubMed]
- Lam, J.H.; Li, Y.; Zhu, L.; et al. A deep learning framework to predict binding preference of RNA constituents on protein surface. Nat. Commun. 2019, 10, 4941. [Google Scholar] [CrossRef] [PubMed]
- Lu, C.-H.; Chen, C.-C.; Yu, C.-S.; et al. MIB2: metal ion-binding site prediction and modeling server. Bioinformatics 2022, 38, 4428–9. [Google Scholar] [CrossRef] [PubMed]
- Hu, X.; Dong, Q.; Yang, J.; et al. Recognizing metal and acid radical ion-binding sites by integrating ab initio modeling with template-based transferals. Bioinformatics 2016, 32, 3260–9. [Google Scholar] [CrossRef] [PubMed]
- Dürr, S.L.; Levy, A.; Rothlisberger, U. Metal3D: a general deep learning framework for accurate metal ion location prediction in proteins. Nat. Commun. 2023, 14, 2713. [Google Scholar] [CrossRef] [PubMed]
- Lin, X.; Su, Z.; Liu, Y.; et al. SuperMetal: A Generative AI Framework for Rapid and Precise Metal Ion Location Prediction in Proteins. Prepr. Bioinform. 2025. [Google Scholar] [CrossRef] [PubMed]
- Wen, B.; Wang, C.; Li, K.; et al. DeepMVP: deep learning models trained on high-quality data accurately predict PTM sites and variant-induced alterations. Nat. Methods 2025, 22, 1857–67. [Google Scholar] [CrossRef] [PubMed]
- Wang, D.; Liu, D.; Yuchi, J.; et al. MusiteDeep: a deep-learning based webserver for protein post-translational modification site prediction and visualization. Nucleic Acids Res. 2020, 48, W140–6. [Google Scholar] [CrossRef] [PubMed]
- Nijkamp, E.; Ruffolo, J.A.; Weinstein, E.N.; et al. ProGen2: Exploring the boundaries of protein language models. Cell Syst. 2023, 14, 968–978.e3. [Google Scholar] [CrossRef] [PubMed]
- Bhatnagar, A.; Jain, S.; Beazer, J.; et al. Scaling Unlocks Broader Generation and Deeper Functional Understanding of Proteins. Prepr. Synth. Biol. 2025. [Google Scholar] [CrossRef]
- Nguyen, E.; Poli, M.; Durrant, M.G.; et al. Sequence modeling and design from molecular to genome scale with Evo. Science 2024, 386, eado9336. [Google Scholar] [CrossRef] [PubMed]
- Merchant, A.T.; King, S.H.; Nguyen, E.; et al. Semantic design of functional de novo genes from a genomic language model. Nature 2026, 649, 749–58. [Google Scholar] [CrossRef] [PubMed]
- Hayes, T.; Rao, R.; Akin, H.; et al. Simulating 500 million years of evolution with a language model. Science 2025, 387, 850–8. [Google Scholar] [CrossRef] [PubMed]
- Hsu, C.; Verkuil, R.; Liu, J.; et al. Learning inverse folding from millions of predicted structures. Prepr. Syst. Biol. 2022. [Google Scholar] [CrossRef]
- Dauparas, J.; Lee, G.R.; Pecoraro, R.; et al. Atomic context-conditioned protein sequence design using LigandMPNN. Nat. Methods 2025, 22, 717–23. [Google Scholar] [CrossRef] [PubMed]
- Wei, G.-W. Protein structure prediction beyond AlphaFold. Nat. Mach. Intell. 2019, 1, 336–7. [Google Scholar] [CrossRef] [PubMed]
- Gainza, P.; Wehrle, S.; Van Hall-Beauvais, A.; et al. De novo design of protein interactions with learned surface fingerprints. Nature 2023, 617, 176–84. [Google Scholar] [CrossRef] [PubMed]
- Watson, J.L.; Juergens, D.; Bennett, N.R.; et al. De novo design of protein structure and function with RFdiffusion. Nature 2023, 620, 1089–100. [Google Scholar] [CrossRef] [PubMed]
- Ingraham, J.B.; Baranov, M.; Costello, Z.; et al. Illuminating protein space with a programmable generative model. Nature 2023, 623, 1070–8. [Google Scholar] [CrossRef] [PubMed]
- Pacesa, M.; Nickel, L.; Schellhaas, C.; et al. One-shot design of functional protein binders with BindCraft. Nature 2025, 646, 483–92. [Google Scholar] [CrossRef] [PubMed]
- Butcher, J.; Krishna, R.; Mitra, R.; et al. De novo Design of All-atom Biomolecular Interactions with RFdiffusion3. Prepr. Biochem. 2025. [Google Scholar] [CrossRef] [PubMed]
- Lisanza, S.L.; Gershon, J.M.; Tipps, S.W.K.; et al. Multistate and functional protein design using RoseTTAFold sequence space diffusion. Nat. Biotechnol. 2025, 43, 1288–98. [Google Scholar] [CrossRef] [PubMed]
- Fei, H.; Li, Y.; Liu, Y.; et al. Advancing protein evolution with inverse folding models integrating structural and evolutionary constraints. Cell 2025, 188, 4674–4692.e19. [Google Scholar] [CrossRef] [PubMed]
- Jiang, K.; Yan, Z.; Di Bernardo, M.; et al. Rapid in silico directed evolution by a protein language model with EVOLVEpro. Science 2025, 387, eadr6006. [Google Scholar] [CrossRef] [PubMed]
- Wolf, B.; Shehu, P.; Brenker, L.; et al. Rational engineering of allosteric protein switches by in silico prediction of domain insertion sites. Nat. Methods 2025, 22, 1698–706. [Google Scholar] [CrossRef] [PubMed]
- Witten, J.; Raji, I.; Manan, R.S.; et al. Artificial intelligence-guided design of lipid nanoparticles for pulmonary gene therapy. Nat. Biotechnol. 2025, 43, 1790–9. [Google Scholar] [CrossRef] [PubMed]
- Chan, A.; Kirtane, A.R.; Qu, Q.R.; et al. Designing lipid nanoparticles using a transformer-based neural network. Nat. Nanotechnol. 2025, 20, 1491–501. [Google Scholar] [CrossRef] [PubMed]
- Xu, Y.; Ma, S.; Cui, H.; et al. AGILE platform: a deep learning powered approach to accelerate LNP development for mRNA delivery. Nat. Commun. 2024, 15, 6305. [Google Scholar] [CrossRef] [PubMed]
- Li, H.; Sarkar, S.; Lu, W.; et al. Collective intelligence for AI-assisted chemical synthesis. Nature 2026, 651, 107–15. [Google Scholar] [CrossRef] [PubMed]
- Zhang, H.; Liu, H.; Xu, Y.; et al. Deep generative models design mRNA sequences with enhanced translational capacity and stability. Science 2025, 390, eadr8470. [Google Scholar] [CrossRef] [PubMed]
- Zhang, H.; Zhang, L.; Lin, A.; et al. Algorithm for optimized mRNA design improves stability and immunogenicity. Nature 2023, 621, 396–403. [Google Scholar] [CrossRef] [PubMed]
- Bryant, D.H.; Bashir, A.; Sinai, S.; et al. Deep diversification of an AAV capsid protein by machine learning. Nat. Biotechnol. 2021, 39, 691–6. [Google Scholar] [CrossRef] [PubMed]
- Li, S.; Xu, K.; Li, G.; et al. Engineering the MmeFz2-ωRNA system for efficient genome editing through an integrated computational-experimental framework. Nat. Commun. 2026, 17, 1867. [Google Scholar] [CrossRef] [PubMed]
- Li, B.; Raji, I.O.; Gordon, A.G.R.; et al. Accelerating ionizable lipid discovery for mRNA delivery using machine learning and combinatorial chemistry. Nat. Mater. 2024, 23, 1002–8. [Google Scholar] [CrossRef] [PubMed]
Figure 2.
Technical pipeline for AI-driven CRISPR system design. Closed loop from component design through experimental validation, with feedback among design, prediction, optimization, and validation.
Figure 2.
Technical pipeline for AI-driven CRISPR system design. Closed loop from component design through experimental validation, with feedback among design, prediction, optimization, and validation.

Figure 4.
Learning paradigms, model architectures, and CRISPR tasks. Sankey diagram with three columns: learning paradigms (L1 supervised, L2 unsupervised, L3 reinforcement learning), model architectures (M1 CNN, M2 RNN/LSTM, M3 Transformer, M4 GNN, M5 geometric/3D networks), and CRISPR task categories (C1–C8; see Table 2). Ribbons indicate support and correspondence from paradigms to architectures and from architectures to tasks. Representative tools and examples: Section 5 and Table 2.
Figure 4.
Learning paradigms, model architectures, and CRISPR tasks. Sankey diagram with three columns: learning paradigms (L1 supervised, L2 unsupervised, L3 reinforcement learning), model architectures (M1 CNN, M2 RNN/LSTM, M3 Transformer, M4 GNN, M5 geometric/3D networks), and CRISPR task categories (C1–C8; see Table 2). Ribbons indicate support and correspondence from paradigms to architectures and from architectures to tasks. Representative tools and examples: Section 5 and Table 2.

Table 2.
Summary of CRISPR toolkit classification.
| Toolkit | Sub-toolkit category | Applicable tasks | Representative tools |
|---|---|---|---|
| C1. CRISPR-Cas database searchand Component Identification | C1a. System Discovery and Annotation | Identify and classify CRISPR-Cas systems and loci in genomes | CRISPRCasFinder [103], CRISPRCasdb [104], CRISPRminer [105] |
| C1b. Pre-crRNA Processing Prediction | Predict precursor crRNA processing and mature gRNA sequences | CHOOSER [106] | |
| C1c. PAM Motif Characterization | Predict PAM sequence requirements for Cas variants | Protein2PAM [107], PAMPHLET [108] | |
| C1d. Precision sgRNA Design and Scoring | Design and score sgRNA sequences for target sites | Cas9: CRISPOR [50,51]; Cas12a: DeepCpf1 [58]; Cas13d: TIGER [59] | |
| C1e. Specificity and Off-target Assessment | Identify and assess off-target sites (in silico, in vivo, in vitro) | In silico: Cas-Offinder [109], CRISPOR [50,51], Elevation [30]; in vivo: GUIDE-seq [39], Detect-seq [110]; in vitro: CHANGE-seq [40], Digenome-seq [111] | |
| C2. CRISPR-Cas structure prediction | C2a. Single-chain Cas folding | Predict 3D structure of single-chain Cas proteins | Single-chain: AlphaFold2 [33], ESMFold [36] |
| C2b. Multi-entity Cofolding and RNP Assembly | Predict protein-nucleic acid and complex structures | Complex: RoseTTAFold [83], AF-Multimer [112]; Database: AlphaFoldDB [113], ESM Metagenomic Atlas [36]; RF-AllAtom [99], AlphaFold3 [61], BOLTZ-2 [114], Chai-2 [115] | |
| C2c. Structure stability | Predict and optimize protein structural stability | FlodX [116], ThermoMPNN [117] | |
| C2d. Conformational dynamics modeling | Model molecular dynamics and conformational changes | GROMACS [118], AMBER [119,120] | |
| C3. CRISPR-Cas function prediction | C3a. Binding site identification | Predict nucleic acid and metal ion binding sites | Nucleic acid site: DRNApred [121], NucleicNet [122]; metal ion site: MIB2 [123], IonCom [124], Metal3D [125], SuperMetal [126] |
| C3b. Post-translational modification | Predict post-translational modification sites | DeepMVP [127], MusiteDeep [128] | |
| C4. CRISPR-Cas sequence generation | C4a. Evolution-guided generation | Generate sequences guided by evolutionary information | ProGene2 [129], ProGene3 [130], ESM2 [36], Evo [131,132] |
| C4b. Function-to-sequence generation | Generate sequences from functional specifications | ProGen [68], ESM3 [133] | |
| C4c. Structure-to-sequence generation | Design sequences from target structures | ProteinMPNN [32], ESM-IF [134], LigandMPNN [135] | |
| C5. CRISPR-Cas structure generation | C5a. Template-based structure design | Design structures from known structural templates | DeepFragLib [136], MaSIF-search [137] |
| C5b. Generative structure design | De novo generation of protein structures | RFDiffusion [138], Chroma [139], RFDiffusionAA [99] | |
| C5c. Sequence-structure co-design | Co-design sequence and structure for binding or function | BinderCraft [140], RFDiffusion3 [141], ProteinGenerator [142] | |
| C6. Virtual screening | C6a. Binding and functional activity prediction | Mutation effect prediction; insertion site prediction | Mutation prediction: AiCE [143], EVOLVEpro [144]; insertion site prediction: ProDomino [145] |
| C6b. Developability assessment | Assess developability of delivery systems (e.g. LNP, AAV) | LNP_ML [146], COMET [147], AGILE [148], MOSAIC [149] | |
| C7. DNA synthesis | C7. Back translation | Optimize DNA/RNA sequence for expression or synthesis | GEMORNA [150], LinearDesign [151] |
| C8. Gene editing experiment | C8. Designing gene-editing experiments | Design and plan gene-editing experiments | CRISPR-GPT [62] |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.