Submitted:
09 September 2026
Posted:
09 September 2026
You are already at the latest version
Abstract
Multimodal molecular profiles are often assumed to improve cancer drug-response prediction simply by adding information. We developed an evidence-bounded DepMap workflow to test that assumption across expression, copy-number, and mutation features. The declared evaluation universe contained 15 compounds from GDSC1, GDSC2, and CTD². Models used fold-local feature selection and Ridge regression, with drug-held-out and leave-one-dataset-out validation. In full-universe drug-held-out validation across five repeated splits, fusion had mean RMSE 0.2322 versus 0.2328 for expression alone (mean fusion-minus-expression RMSE −0.0006; seed-level SD 0.0013). Fusion improved in three of five splits and was slightly worse in two. In cross-study validation, fusion was slightly worse in all three held-out datasets: CTD² 0.3238 versus 0.3231, GDSC1 0.2339 versus 0.2322, and GDSC2 0.2387 versus 0.2347. Structured missingness and weak-to-moderate row-level modality correlations accompanied dataset-sensitive performance. These findings support evaluating multimodal models by incremental value, transferability, and failure modes rather than assuming that more omics is always better.

Keywords:
information fusion
; drug response
; multi-omics
; cancer
; missing data
; cross-study validation
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.