Preprint
Article

This version is not peer-reviewed.

Multi-Platform Transcriptomic Classification of NSCLC Driver Mutation Subtypes via Heterogeneous Ensemble Learning

Submitted:

04 August 2026

Posted:

04 August 2026

You are already at the latest version

Abstract
Background: Accurate identification of EGFR-mutant, KRAS-mutant, and triple-negative (TN) NSCLC driver mutation subtypes is essential for guiding targeted therapy decisions, yet standard molecular testing remains invasive and infeasible in a clinically relevant subset of patients. Gene expressionbased machine learning approaches offer a non-invasive alternative; however, existing models are largely restricted to binary classification tasks and lack cross-platform generalizability. Methods: We developed a multi-cohort, harmonized classification pipeline integrating ComBat batch correction, SHAP-driven feature selection, and a Bayesian optimized heterogeneous ensemble comprising five complementary learners, further extended with NSGA-II Pareto-optimal ensemble weight optimization. A strict leak-free cross-validation framework was applied throughout model development. Generalizability was assessed through independent external validation on an RNA-seq cohort with reference-batch cross-platform harmonization and class-specific probability threshold calibration on an internal validation subset. Results: The Pareto-optimized ensemble (ParetoVoting) achieved the best internal performance, with F1-macro= 0.736 and AUROC= 0.878 (5-fold cross-validation mean); this improvement over the SoftVoting ensemble (F1-macro= 0.735, AUROC= 0.877) was negligible (Cliff’s δ = 0.00). Given this equivalence, the simpler SoftVoting ensemble was carried forward for independent external validation. On external validation (TCGA-LUAD, n = 508), SoftVoting attained F1-macro= 0.680 and AUROC= 0.853. Ablation experiments confirmed that each pipeline component contributed independently to overall performance. SHAP interpretability analysis identified biologically coherent discriminative signatures consistent with known NSCLC oncogenic signaling. Conclusions: This study establishes a reproducible, cross-platform generalizable framework for simultaneous three-class driver mutation subtyping from transcriptomic data, providing methodological benchmarks for future work in this underexplored classification setting. The approach holds potential as a complementary decision-support tool in clinical contexts where tissue-based molecular profiling is not feasible.
Keywords: 
;  ;  ;  ;  ;  
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.
Prerpints.org logo

Preprints.org is a free preprint server supported by MDPI in Basel, Switzerland.

Subscribe

© 2026 MDPI (Basel, Switzerland) unless otherwise stated

Accessibility

Disclaimer

Terms of Use

Privacy Policy

Privacy Settings