Submitted:
02 October 2026
Posted:
05 October 2026
You are already at the latest version
Abstract
Ribonucleic acid (RNA) foundation models (FMs) are reshaping computational biology, using self-supervised pretraining on massive transcriptomic corpora to acquire generalisable molecular representations. Developed initially for sequence annotation and expression prediction, they are increasingly applied to plant science, supporting tasks from gene function inference to regulatory element discovery and stress-response modelling in both model and orphan crop species. This review surveys RNA FMs with a focus on plant-relevant applications. We catalogue the major model families, including DNABERT-derived RNA adaptations, PlantRNA-FM, and Nucleotide Transformer variants, comparing their pretraining objectives, input modalities, and evaluation protocols. We examine how these models are deployed across plant genomics, transcriptomics, and functional annotation pipelines, highlighting the opportunities and the constraints arising from species-specific data gaps, limited experimental validation, and interpretability challenges. Across applications, the decisive constraint is no longer architecture but data: corpus breadth, condition-resolved benchmarks, and experimentally anchored validation now dominate the marginal gains available from model design. We close with the key open questions and a roadmap for integrating RNA FMs into the plant systems biology toolkit, emphasising benchmark design, cross-species transfer, and experimentally grounded validation. Throughout, every claim about what a model can do is held to three questions of evidence: what data the model was trained on, which benchmark supports the claim, and whether experimental validation in a plant system exists.
Keywords:
RNA foundation models
; deep learning
; plant genomics
; pretraining
; non-coding RNA
; gene regulation
; transcriptomics
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.