Preprint
Article

This version is not peer-reviewed.

GenomicsOps: An Offline-First Desktop Application for ACMG/AMP Variant Classification with a Strength-Aware Combining Engine and Hash-Chained Auditing

Submitted:

03 October 2026

Posted:

08 October 2026

You are already at the latest version

Abstract
Background: ACMG/AMP 2015 variant classification requires combining criteria of varying strength into a single five-tier class. The combining rules are specified, but in many tools the combining step is treated as a black box, and the resulting classifications cannot be inspected or reconstructed. This is a problem for clinical laboratories, which must defend classifications to accreditation bodies and ethics committees, and it is a more universal gap than the network-access constraints that motivate offline alternatives. Results: GenomicsOps is built around a strength-aware combining engine that implements the ACMG/AMP 2015 rules together with the ClinGen VCEP strength modifiers and the Tavtigian points-based extension. The combining engine achieves 95.8% concordance with curator classifications on 12,644 ClinGen ERepo variants when curator-supplied criteria are provided. This figure describes the combining step in isolation, not end-to-end accuracy. The criterion-extraction path that precedes combining is a partial automation: on average it recovers 68.9% of the criteria a human curator would report, with 92.5% agreement on a well-annotated subset of the evaluation corpus. Curators should expect to supplement its output. The pipeline runs fully offline and writes every classification to a SHA-256 hash-chained audit log. End-to-end accuracy on a user-supplied VCF is bounded by the criterion-extraction recall reported above and will be lower than the combining-only figure for variants carrying less structured annotation than the evaluation corpus. Conclusions: The central contribution is a reproducible implementation of the ACMG/AMP combining logic, evaluated on a large expert-curated corpus. Transparency of the classification record is an architectural property of the design, not an empirically validated benefit; a user study measuring the effect of criterion-level reporting on review behaviour is left as future work. Stratified analysis of the two combining paths shows that aggregate agreement (95.7% points-based vs 96.2% generic) conceals substantial subgroup differences: the choice of path affects per-class and per-criterion-count agreement by up to 10 percentage points in opposite directions, and the generic aggregate is dominated by a single panel.
Keywords: 
;  ;  ;  ;  ;  ;  
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.