Submitted:
23 September 2026
Posted:
24 September 2026
You are already at the latest version
Abstract
Cross-referencing an Arabidopsis thaliana locus against UniProt accessions, Gene Ontology (GO) terms, TAIR10/Araport11 genomic coordinates, InterPro domain annotations, Plant Ontology terms, and PubMed literature currently requires the integration of several independently formatted flat files. In many laboratories, these tasks are performed using ad hoc scripts that are developed for individual analyses and are rarely parallelized. Here, I present araseq (https://codeberg.org/gsablok/araseq), a single Rust binary that implements thirteen cross-referencing operations as clap subcommands and, optionally, exposes the same functionality through an axum-based web server. All operations execute within a user-configurable Rayon thread pool and follow a common read-index-filter paradigm: identifier lists are normalized where appropriate, reference files are streamed into in-memory hash maps, and the resulting maps are filtered according to the requested identifiers. I describe the implementation and independently reproduce two representative algorithms in Python using the TAIR10/Araport11 sample files distributed with the repository to assess algorithmic correctness. The evaluation confirms the expected behavior of the normalized matching workflow while also identifying an inconsistency in identifier normalization between modules. These findings indicate that araseq provides a compact and reusable framework for routine cross-referencing of A. thaliana annotation resources.

Keywords:
Arabidopsis thaliana
; genome annotation
; UNIPROT
; TAIR
1. Introduction
Arabidopsis thaliana remains a principal reference model for land-plant genomics. Its major annotation resources—including TAIR10 and Araport11 gene models [1,2], AGI-to-UniProt and TAIR-to-UniProt identifier mappings [3], GAF-style gene association files containing Gene Ontology terms [4], InterProScan domain annotations based on the InterPro signature database [6,7], and the Plant Ontology (PO) [8]—are distributed across large and heterogeneous flat files rather than a single queryable database. A routine task in Arabidopsis genomics is to convert a list of gene loci, such as AGI identifiers (e.g., AT5G27680), into corresponding UniProt accessions, GO terms, genomic coordinates, literature cross-references, and protein-domain annotations. These resources are distributed across files with distinct formats, column structures, and parsing requirements. Common complications include comment lines prefixed with !, splice-variant suffixes such as .1 that may need to be removed before matching, and semicolon- or pipe-delimited subfields. Consequently, such analyses are frequently performed using ad hoc, single-purpose scripts that are rewritten for individual projects and are rarely parallelized. This approach can become limiting when either the input identifier list or the reference files contain tens of thousands of entries, as is common in whole-genome analyses.
Here, I present araseq, a single Rust binary [9] that consolidates these recurring cross-referencing tasks into a unified tool with two interchangeable interfaces: a clap-based command-line interface, with one subcommand for each operation, and an axum-based web server exposing the same functionality through Bootstrap-styled HTML forms. The choice of Rust over a conventional scripting-language implementation has three principal motivations. First, memory safety and the absence of a garbage collector are advantageous for long-running batch and server workloads. Second, Rayon provides straightforward, user-configurable data parallelism for file-scanning and filtering operations that constitute a substantial proportion of execution time. Third, compilation into a single binary simplifies deployment on shared institutional computing infrastructure, where installing and maintaining a correctly versioned scripting-language environment can itself become a practical bottleneck. In this study, I describe the implementation of araseq, evaluate two representative algorithms against the TAIR10/Araport11 sample data distributed with the repository, and discuss the strengths and current limitations of the design.
2. Materials and Methods
araseq is implemented in Rust (2021 edition) and distributed as an MIT-licensed binary crate. Command-line parsing is performed using the clap crate and its derive API, while parallel execution is implemented using Rayon. Tabular data are parsed using the csv crate together with manual BufReader- and delimiter-based parsing. JSON output generated by InterProScan is parsed using serde and serde_json. Regular-expression-based extraction of Plant Ontology identifiers uses the regex crate. HTTP retrieval and HTML parsing of PubMed abstracts are implemented using the blocking client provided by reqwest and the scraper crate, respectively. The optional web interface is implemented using axum with a multi-threaded Tokio runtime. All dependency versions specified in Cargo.toml are exact-pinned.
Reference Data
araseq operates on standard TAIR and Araport distribution files. These files are not bundled with the compiled binary; however, their expected formats are documented in the source code, and representative samples are provided in the repository. The principal reference files include AGI2uniprot-Jul2023.txt and Uniprot2AGI-Jul2023.txt, which provide AGI–UniProt identifier mappings [3]; TAIR2UniprotMapping-Jul2023.txt, which provides TAIR locus–UniProt mappings; and gene_association.tair, a GAF-style file used to retrieve GO terms [4], TAIR/PubMed cross-reference numbers, and free-text gene descriptions.
For genomic coordinates, araseq uses TAIR10_GFF3_genes.gff [1] and Araport11 GFF3 gene models [2]. InterProScan JSON output is used to retrieve protein-domain and signature matches [6,7], while an InterPro hierarchy text dump provides parent–child IPR relationships. The Plant Ontology is represented by po.owl [8]. In addition, arbitrary FASTA files can be supplied for motif and k-mer frequency analysis.
Command Architecture and Concurrency Model
araseq exposes thirteen analytical subcommands and a Serve command that starts the optional web interface (Table 1). Each analytical subcommand accepts an explicit thread-count argument. The corresponding handler constructs a dedicated rayon::ThreadPoolBuilder with the requested number of threads and executes the analysis within that pool. Consequently, parallelism is configured independently for each invocation rather than through a global process-wide setting.The serve command instead initializes a Tokio runtime.Table 1 summarizes the thirteen analytical subcommands, their corresponding source modules, and their core algorithms.
Discussion
The implementation addresses its primary objective of consolidating recurring Arabidopsis annotation lookups into a single reusable binary that can be operated either from a command-line pipeline or through a web browser without duplicating the underlying computational logic. The CLI and Axum routes invoke the same core functions, thereby reducing the possibility of divergence between the two interfaces. The per-command, user-configurable Rayon thread pool provides explicit control over computational parallelism while avoiding reliance on a single global configuration. In the web-server implementation, execution of the synchronous analysis functions through tokio::task::spawn_blocking separates file and network operations from the asynchronous reactor and prevents blocking operations from interfering with Tokio’s event loop. Overall, araseq provides a compact framework for consolidating recurrent, script-based cross-referencing of A. thaliana identifiers across TAIR, UniProt, Gene Ontology, GFF3, InterPro, and Plant Ontology resources. Its Rust implementation combines memory safety with user-configurable parallel execution, while the optional web interface provides an alternative to command-line workflows without requiring a separate implementation of the underlying algorithms.
Availability
araseq is freely available at:https://codeberg.org/gsablok/araseq.
Author Contributions
GS conceived the idea, wrote the software and paper.
References
- Lamesch, P.; Berardini, T.Z.; Li, D.; Swarbreck, D.; Wilks, C.; Sasidharan, R.; Muller, R.; Dreher, K.; Alexander, D.L.; Garcia-Hernandez, M.; Karthikeyan, A.S.; Lee, C.H.; Nelson, W.D.; Ploetz, L.; Singh, S.; Wensel, A.; Huala, E. The Arabidopsis Information Resource (TAIR): improved gene annotation and new tools. Nucleic Acids Res. 2012, 40(D1), D1202–D1210. [Google Scholar] [CrossRef] [PubMed]
- Cheng, C.Y.; Krishnakumar, V.; Chan, A.P.; Thibaud-Nissen, F.; Schobel, S.; Town, C.D. Araport11: a complete reannotation of the Arabidopsis thaliana reference genome. Plant J. 2017, 89(4), 789–804. [Google Scholar] [CrossRef] [PubMed]
- The UniProt Consortium. UniProt: the Universal Protein Knowledgebase in 2023. Nucleic Acids Res. 2023, 51(D1), D523–D531. [Google Scholar] [CrossRef] [PubMed]
- Ashburner, M.; Ball, C.A.; Blake, J.A.; Botstein, D.; Butler, H.; Cherry, J.M.; Davis, A.P.; Dolinski, K.; Dwight, S.S.; Eppig, J.T.; Harris, M.A.; Hill, D.P.; Issel-Tarver, L.; Kasarskis, A.; Lewis, S.; Matese, J.C.; Richardson, J.E.; Ringwald, M.; Rubin, G.M.; Sherlock, G. Gene Ontology: tool for the unification of biology. Nat. Genet. 2000, 25(1), 25–29. [Google Scholar] [CrossRef] [PubMed]
- Eilbeck, K.; Lewis, S.E.; Mungall, C.J.; Yandell, M.; Stein, L.; Durbin, R.; Ashburner, M. The Sequence Ontology: a tool for the unification of genome annotations. Genome Biol. 2005, 6(5), R44. [Google Scholar] [CrossRef] [PubMed]
- Blum, M.; Chang, H.Y.; Chuguransky, S.; Grego, T.; Kandasaamy, S.; Mitchell, A.; Nuka, G.; Paysan-Lafosse, T.; Qureshi, M.; Raj, S.; Richardson, L.; Salazar, G.A.; Williams, L.; Bork, P.; Bridge, A.; Gough, J.; Haft, D.H.; Letunic, I.; Marchler-Bauer, A.; Mi, H.; Natale, D.A.; Necci, M.; Orengo, C.A.; Pandurangan, A.P.; Rivoire, C.; Sigrist, C.J.A.; Sillitoe, I.; Thanki, N.; Thomas, P.D.; Tosatto, S.C.E.; Wu, C.H.; Bateman, A.; Finn, R.D. The InterPro protein families and domains database: 20 years on. Nucleic Acids Res. 2021, 49(D1), D344–D354. [Google Scholar] [CrossRef] [PubMed]
- Jones, P.; Binns, D.; Chang, H.Y.; Fraser, M.; Li, W.; McAnulla, C.; McWilliam, H.; Maslen, J.; Mitchell, A.; Nuka, G.; Pesseat, S.; Quinn, A.F.; Sangrador-Vegas, A.; Scheremetjew, M.; Yong, S.Y.; Lopez, R.; Hunter, S. InterProScan 5: genome-scale protein function classification. Bioinformatics 2014, 30(9), 1236–1240. [Google Scholar] [CrossRef] [PubMed]
- Cooper, L.; Walls, R.L.; Elser, J.; Gandolfo, M.A.; Stevenson, D.W.; Smith, B.; Preece, J.; Athreya, B.; Mungall, C.J.; Rensing, S.; Hiss, M.; Lang, D.; Reski, R.; Berardini, T.Z.; Li, D.; Huala, E.; Schaeffer, M.; Menda, N.; Arnaud, E.; Shrestha, R.; Yamazaki, Y.; Jaiswal, P. The Plant Ontology as a tool for comparative plant anatomy and genomic analyses. Plant Cell Physiol. 2013, 54(2), e1. [Google Scholar] [CrossRef] [PubMed]
- Matsakis, N.D.; Klock, F.S., II. The Rust language. ACM SIGAda Ada Lett. 2014, 34(3), 103–104. [Google Scholar] [CrossRef]
Table 1.
Subcommand inventory of araseq.
| Subcommand(s) | Module | Core algorithm |
|---|---|---|
| UNIPROTAGI / AGIUNIPROT | uniprot.rs, agi.rs | Cross-map AGI ↔ UniProt identifiers using the AGI2uniprot/Uniprot2AGI tables without identifier normalization. |
| AGIGO | agigo.rs | Extract GO terms (column 5) from the cleaned GAF file, keyed by normalized AGI loci (column 2). |
| AGITAIR | tairagi.rs | Extract TAIR communication/PubMed cross-reference numbers (column 6) using normalized identifiers. |
| AGIDescription | agiassoc.rs | Extract free-text gene descriptions (column 11) using normalized identifiers. |
| TAIRUNIPROT | tair.rs | Match TAIR loci to UniProt accessions using the second field of the mapping table, with identifier normalization. |
| AGICoordinate | agigff.rs | Stream a GFF3 file, group records by feature type, and return coordinates for requested genes using normalized identifiers. |
| GFFEncoder | gffencoder.rs | Parse and group GFF3 features using the csv crate, including genes, exons, 3′ UTRs, and 5′ UTRs. |
| AGIPUBMDED | api.rs | Collect PMIDs associated with a locus from the GAF file and retrieve PubMed abstract text using reqwest and scraper. |
| PlantOnto | plantoml.rs | Extract PO_\d+ identifiers referenced in a Plant Ontology file using regular expressions. |
| DomainAnalyzer | domain.rs | Extract a selected field (e.g., signature, E-value, score, or accession) from InterProScan JSON matches. |
| IPRAnalyzer | ipr.rs | Parse an InterPro hierarchy dump to construct parent–child IPR and category lookups. |
| FrequencyMotif | frequency.rs | Count exact k-mer-window matches to a query motif across a FASTA file. |
| Serve | web.rs | Provide an Axum-based web interface exposing the analytical commands through Bootstrap 5 HTML forms. |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.