Submitted:
25 August 2026
Posted:
26 August 2026
You are already at the latest version
Abstract
Cancer-associated alterations perturb molecular systems whose organization is only partially captured by protein–protein interaction topology. We asked whether Gene Ontology (GO)- derived semantic information can improve reconstruction of mutation-defined cancer modules and, critically, whether the magnitude and ontology source of this improvement vary across tumour types. We constructed recurrent non-silent somatic-mutation modules for 31 non-haematological The Cancer Genome Atlas (TCGA) projects and embedded them in a frozen human interaction network enriched with a nine-dimensional GO semantic representation. Topology-only random walk with restart (RWR) was compared with semantic-RWR under symmetric leave-one-out module reconstruction. The benchmark comprised 7,440 evaluations spanning 31 cancers, module sizes of 50, 200 and 400 proteins, intact or 20% edge-depleted interactomes, two graph perturbation seeds, five module-subsampling seeds, and four semantic configurations (combined, Biological Process, Molecular Function and Cellular Component). Combined semantic diffusion significantly improved NDCG@100 in 27 of 31 cancers after Holm correction. Semantic gain increased strongly with module size and was largely preserved after removal of 20% of interactome edges. The optimal ontology branch differed among cancers: Molecular Function was optimal in 13, Biological Process in 12, Cellular Component in 4, and the combined representation in 2. These findings indicate that cancer modules are not organized by topology alone and support the existence of tumour-specific semantic network architectures: systems-level patterns describing whether functional coherence is expressed primarily through biological processes, biochemical activities, cellular compartments, or combinations thereof. Biologically, this heterogeneity is consistent with the known convergence of diverse cancer mutations onto shared pathways, complexes and cellular programs. PanCancerSemNet therefore provides a framework for interpreting cancer-associated alterations as context-dependent perturbations of functionally annotated molecular systems rather than as isolated mutated genes.
Keywords:
cancer networks
; network medicine
; Gene Ontology
; semantic similarity
; random walk with restart
; TCGA
; pan-cancer analysis
; disease modules
1. Introduction
Cancer is a disease of molecular systems rather than isolated genes. Large-scale sequencing has revealed a characteristic landscape in which a limited number of recurrent drivers coexist with a long tail of less frequent alterations, while different tumours can reach related malignant phenotypes through distinct combinations of genomic events (Bailey et al. 2018; Vogelstein et al. 2013; Lawrence et al. 2013). This diversity is accompanied by convergence at higher biological levels: alterations affecting different genes frequently impinge on common signalling pathways, cellular programs and hallmark capabilities such as sustained proliferation, evasion of growth control, genome instability, phenotypic plasticity and interactions with the tumour microenvironment (Hanahan 2022; Sanchez-Vega et al. 2018). Pan-cancer analyses have further shown that tissue of origin remains a dominant determinant of molecular identity while substantial functional commonalities cut across histological boundaries (Hoadley et al. 2018; The Cancer Genome Atlas Research Network et al. 2013).
These observations provide a strong rationale for network-based representations of cancer. Network medicine models disease-associated genes as perturbations of an interconnected molecular system, where disease modules occupy specific regions of the interactome and where relationships among genes can emerge from connectivity even when they are not apparent from mutation frequency alone (Barabasi et al. 2011; Menche et al. 2015). In cancer, network propagation and network-based stratification have demonstrated that sparse, heterogeneous mutations can be transformed into more coherent representations by projecting them onto molecular interaction maps (Cowen et al. 2017; Leiserson et al. 2015; Hofree et al. 2013). Such approaches are especially attractive in a pan-cancer setting because they can reveal pathway-level convergence among tumours that share few individual mutations.
Topology, however, captures only one dimension of biological organization. A protein–protein interaction indicates that two proteins are related in a molecular network, but it does not specify how they are related. The same topological connection may participate in signal transduction, DNA repair, metabolism, transcriptional regulation, a multiprotein complex or a particular subcellular compartment. Conversely, functionally related proteins may be separated in an incomplete interaction map. The Gene Ontology (GO) provides a structured vocabulary for Biological Process (BP), Molecular Function (MF) and Cellular Component (CC), enabling quantitative comparison of gene products at the level of biological meaning (Gene Ontology Consortium 2023; Pesquita et al. 2009; Wang et al. 2007). Semantic similarity has consequently been used to support functional inference, interaction analysis and disease-gene prioritization, and earlier work has shown the value of combining interactome information with GO annotations when searching for functionally associated proteins (Cho et al. 2013).
The biological motivation for separating BP, MF and CC is particularly strong in cancer. BP captures coordinated programmes such as cell-cycle control, DNA-damage response, apoptosis, differentiation, immune regulation and migration; MF captures biochemical capabilities such as kinase activity, receptor binding, transcription-factor activity and enzymatic catalysis; and CC captures the spatial and complex-level organization in which these activities occur, including membranes, chromatin, organelles and macromolecular assemblies. These dimensions need not contribute equally in every tumour type. Pan-cancer studies have already shown substantial heterogeneity in oncogenic pathway usage, immune state and DNA-repair deficiency (Sanchez-Vega et al. 2018; Thorsson et al. 2018; Knijnenburg et al. 2018). It is therefore plausible that the ontology branch that best organizes a mutation-defined module is itself informative about the dominant level of biological coordination in that cancer.
Here we introduce PanCancerSemNet, an ontology-aware network-propagation framework designed to test this hypothesis across 31 TCGA cancers. For each cancer, recurrent non-silent somatic mutations define a query-centred module that is embedded in a common frozen human interactome. Topology-only RWR is compared with semantic-RWR under symmetric leave-one-out reconstruction, and semantic gain is quantified with mean reciprocal rank (MRR), Recall@100, NDCG@100 and a bounded Semantic Dependency Index (SDI). We evaluate a combined semantic representation and the three GO branches separately, allowing each cancer to be described by a branch-specific semantic profile rather than assuming a universal ontology weighting.
The study is organized as a pan-cancer benchmark designed to quantify four complementary properties of ontology-aware propagation: the overall semantic advantage, its cancer-to-cancer heterogeneity, its dependence on module size, and its robustness to controlled interactome perturbation. Branch-specific analyses further determine whether BP, MF or CC contributes most strongly within each cancer, while the same endpoints and statistical framework are applied across all 31 tumour types.
2. Methods
2.1. Study Design and Cancer Cohort
The benchmark comprised 31 non-haematological TCGA cancer projects. Acute myeloid leukaemia (LAML) and diffuse large B-cell lymphoma (DLBC) were excluded from the solid-tumour benchmark to reduce major differences in tissue architecture, sample composition and mutational context. The analysis contained 7,440 evaluation runs: 31 cancers × 3 module sizes × 2 interactome conditions × 2 graph perturbation seeds × 5 module-subsampling seeds × 4 semantic configurations. All intended strata were present in the completed analysis.
2.2. TCGA Mutation Modules
Open-access masked somatic mutation files were obtained from the NCI Genomic Data Commons. For each cancer, variants classified as silent or non-coding categories not expected to alter protein sequence were removed. Tumour sample barcodes were collapsed to TCGA case identifiers, and genes were ranked by the number of unique tumour cases carrying at least one retained mutation. Mutation prevalence was used as a secondary score. The top 500 genes formed the candidate mutation module before graph mapping.
For cancer c, the ranked module was
where ordering reflects recurrence. Across the 31 cancers, approximately 91–96% of the top 500 genes mapped to the frozen human graph, allowing all cancers to support module sizes up to 400 proteins. Mutation recurrence was chosen as the first universal pan-cancer module definition because it is consistently available across all projects and does not depend on the availability of matched solid-tissue normal RNA samples. Transcriptomic, copy-number, proteomic and integrated multi-omic modules are reserved for subsequent analyses and will be treated as separate module sources.
2.3. Human Interaction Graph
The frozen human protein graph was derived from curated interaction and complex resources including BioGRID and Complex Portal (Oughtred et al. 2021; Balu et al. 2025). The graph used here contained 17,997 proteins and 925,977 protein–protein interactions. Protein identifiers were normalized to canonical UniProt accessions. The same graph was used across all cancer types, ensuring that differences in performance were attributable to disease modules, semantics and controlled perturbations rather than cancer-specific graph reconstruction.
2.4. GO Semantic Representation
Each interaction carried a nine-dimensional semantic representation partitioned into BP, MF and CC triplets. The representation was designed to preserve the three complementary biological views encoded by GO rather than collapsing them a priori. Within each branch, two precomputed semantic-similarity components were averaged and modulated by an availability indicator. This follows the broader principle of GO semantic similarity, in which relatedness is quantified from ontology structure and/or annotation information rather than from lexical identity alone (Pesquita et al. 2009; Wang et al. 2007). For branch b,
For the combined configuration,
The availability term prevents missing annotation from being treated as evidence of dissimilarity. The semantic weight therefore represents prior functional coherence conditional on available GO knowledge. It is not interpreted as evidence of causal regulation, direct biochemical interaction or tumour-specific activity; rather, it changes the probability with which information propagates across an already defined molecular interaction edge.
2.5. Random Walk with Restart
For topology-only RWR, the undirected adjacency matrix A was row-normalized to transition matrix P. With restart vector , scores were iterated as
with . Iteration stopped when the difference between successive vectors fell below or after 500 iterations. Semantic-RWR used the same update equation but constructed P from semantic edge weights. ALL, BP, MF and CC were evaluated separately.
2.6. Leave-One-Out Module Reconstruction
Evaluation was symmetric and query-centred. For a cancer module M and held-out target , the remaining proteins were used as restart seeds. The target was then ranked among non-seed graph nodes. The procedure was repeated for every selected module member.
The implementation exploited linearity of personalized PageRank. If is the personalized RWR solution from individual seed i, then for uniform seed set S,
Personalized vectors were therefore cached once for each graph/semantic perturbation and reused across held-out targets, module sizes and subsampling replicates.
2.7. Benchmark Design
The benchmark used module sizes of 50, 200 and 400 proteins to sample increasingly broad mutation-defined programmes. Interactome robustness was evaluated on the intact graph and after random deletion of 20% of edges. Two independent graph-perturbation seeds and five module-subsampling seeds were used. Each condition was evaluated for ALL, BP, MF and CC semantic propagation against the matched topology-only RWR baseline. The complete design comprised 31 cancers × 3 module sizes × 2 interactome conditions × 2 graph perturbation seeds × 5 module-subsampling seeds × 4 semantic configurations, yielding 7,440 evaluations. All intended strata were present in the completed analysis. Evaluations were performed independently and aggregated only at the pre-specified replicate level used for statistical inference.
2.8. Ranking Metrics and Semantic Dependency Index
Performance was measured using mean reciprocal rank (MRR), Recall@100 and normalized discounted cumulative gain at rank 100 (NDCG@100). The primary pan-cancer endpoint was
Analogous differences were calculated for MRR and Recall@100.
The bounded Semantic Dependency Index was defined as
with NDCG@100 for the primary analysis. Positive SDI denotes a relative semantic advantage; values near zero denote comparable performance.
2.9. Statistical Analysis
Individual experimental rows were not treated as fully independent because module-size and edge-dropout conditions share graph and module seeds. The primary replicate unit was therefore . Within each cancer and semantic configuration, gains were averaged across module-size and dropout conditions within each replicate unit. One-sided Wilcoxon signed-rank tests assessed whether replicate-level semantic gains were greater than zero. Holm correction was applied across 31 cancer tests separately for each semantic configuration and metric.
Confidence intervals for mean gain were estimated with 10,000 cluster-bootstrap resamples using the same graph-seed × module-seed unit as the resampling cluster. Ontology-branch heterogeneity was evaluated by Friedman tests across matched ALL/BP/MF/CC conditions, followed by Holm correction across cancers. Module-size effects were evaluated by Friedman tests across the three matched sizes and paired Wilcoxon comparisons of 400 versus 50. Intact versus 20% edge-depleted conditions were compared by paired Wilcoxon tests across cancer-level summaries. Statistical analyses were implemented in Python using NumPy, pandas and SciPy.
3. Results
3.1. High Graph Coverage Enabled a Uniform 31-Cancer Benchmark
All 31 mutation-defined cancer modules were mapped to the same frozen human graph. Mapping coverage of the top 500 recurrently altered genes was consistently high, with mapped counts in the approximate range 453–478 proteins per cancer. We therefore selected 50, 200 and 400 proteins as common module sizes, avoiding cancer-specific truncation while spanning compact to distributed molecular programmes. The completed benchmark contained exactly 7,440 evaluations covering all intended combinations of cancer, module size, semantic branch, edge condition and random seed.
3.2. Ontology-Aware Diffusion Improves Mutation-Module Reconstruction Across Most Cancers
Combined semantic propagation produced a positive and statistically supported NDCG@100 gain in 27 of 31 cancers after Holm correction (Figure 1b; Supplementary Figure S2). The forest plot shows a strongly non-uniform effect-size distribution rather than a small pan-cancer shift of similar magnitude. The largest mean gains were observed in SKCM (NDCG@100=0.0320), TGCT (NDCG@100=0.0277), PRAD (NDCG@100=0.0178), LUSC (NDCG@100=0.0178) and READ (NDCG@100=0.0159). SKCM showed the strongest average effect (NDCG@100=0.0320, bootstrap 95% CI 0.0287–0.0355; mean SDI=0.575), followed by TGCT (NDCG@100=0.0277, mean SDI=0.426).
The confidence-interval structure is biologically and statistically informative. Strongly positive cancers such as SKCM and TGCT are separated from zero with relatively narrow intervals, whereas the estimates for KICH, UCS, BRCA and SARC are close to the null and/or include zero (Figure 1b; Supplementary Figure S2). This pattern argues against a trivial global reweighting effect. Instead, ontology-derived information contributes differently according to the organization of each mutation-defined cancer module. The absolute changes in NDCG@100 are modest, as expected for a ranking task on a large interactome, but their reproducibility across matched perturbations and their concentration in specific cancers support a structured rather than stochastic effect.
3.3. Semantic Advantage Strengthens as Cancer Modules Become Larger
Semantic gain increased markedly with module size in every semantic configuration (Figure 1c; Supplementary Figure S3). For ALL, mean NDCG@100 increased from 0.0027 at to 0.0077 at and 0.0183 at . The overall size effect was highly significant (Friedman ), and the paired 400-versus-50 comparison remained strongly significant (). BP, MF and CC showed the same qualitative pattern. At , BP and MF reached the largest average improvements, while CC remained positive but lower; the combined configuration occupied an intermediate position.
The monotonic increase is important because the three module sizes correspond to progressively broader mutation-defined programmes. At , the semantic advantage is small and the error bars overlap substantially across branches. At , the separation becomes clearer, and by semantic propagation provides a substantially larger ranking advantage (Supplementary Figure S3). Thus, the utility of semantic information is not restricted to a compact core of highly recurrent genes. It becomes more pronounced as the module incorporates a broader and more heterogeneous set of recurrently altered proteins.
3.4. Semantic Gain Is Retained After 20% Interactome Edge Deletion
Random deletion of one fifth of interactome edges produced only a modest decrease in semantic advantage (Figure 1d; Supplementary Figure S4). For ALL, mean NDCG@100 changed from 0.0100 on the intact graph to 0.0091 after 20% edge deletion, with no significant paired difference (). The corresponding comparisons were also non-significant for BP (), MF (), and CC ().
The ordering of the semantic branches is also largely preserved under perturbation: BP and MF show the largest average gains in both the intact and edge-depleted interactomes, followed by the combined representation and CC (Supplementary Figure S4). Although mean performance decreases slightly for all four configurations, the overlapping uncertainty intervals and non-significant paired tests indicate that the main semantic signal is not driven by a small, fragile subset of interaction edges. These data therefore support complementarity between semantic priors and graph topology rather than dependence on a particular realization of the interactome.
3.5. Cancers Exhibit Distinct Ontology-Branch Dependencies
The ontology branch yielding the highest mean NDCG@100 varied substantially across cancers (Figure 1a; Supplementary Figure S1). MF was optimal in 13 cancers, BP in 12, CC in 4, and the combined representation in only 2. Matched branch differences were significant after Holm correction in 27 of 31 cancers. We therefore represent each cancer by
The heatmap emphasizes that this branch heterogeneity is not simply a consequence of the overall magnitude of semantic gain. SKCM displays a strong positive dependency across all four semantic configurations, with especially high values for the combined and BP representations. TGCT also shows a broad positive profile across branches. In contrast, several cancers show more selective patterns: some exhibit stronger BP or MF dependence, whereas PAAD and SARC show a visibly weaker or negative MF component despite non-identical behaviour in the other branches; BRCA lies close to a neutral combined effect while retaining branch-specific variation (Supplementary Figure S1). These patterns indicate that a single pooled semantic representation can conceal biologically distinct modes of functional coherence.
Hierarchical ordering of the SDI profiles further reveals groups of cancers with broadly positive semantic dependence and others in which one branch contributes little or negatively. We do not interpret these groups as a new cancer taxonomy at this stage. Rather, the heatmap provides a quantitative phenotype that can now be related to independent biological variables such as tissue of origin, mutational burden, pathway activity, immune state and DNA-repair phenotype.
3.6. Integrated Interpretation of the Pan-Cancer Benchmark
Taken together, the four panels of Figure 1 describe three separable properties of the method. First, the forest plot demonstrates that semantic benefit is cancer dependent rather than uniform. Second, the heatmap shows that the source of that benefit is itself cancer dependent, because BP, MF and CC do not contribute equally. Third, the module-size and edge-dropout analyses show that the effect becomes stronger as mutation modules become more distributed while remaining comparatively stable to substantial random topological loss. The combination of these observations is the basis for the concept of a tumour-specific semantic network architecture: semantic dependence is an effect-size profile, not merely a binary indication that ontology information is useful.
4. Discussion
PanCancerSemNet provides evidence that the molecular organization of cancer modules cannot be fully described by interaction topology alone. Across 31 TCGA solid-tumour projects, ontology-aware diffusion improved held-out reconstruction of mutation-defined modules in most cancers, but the magnitude, uncertainty and ontology source of the improvement varied substantially (Figure 1). The four complementary visualizations are important because they separate three questions that would otherwise be conflated: whether semantics improves reconstruction, which semantic axis contributes most strongly, and under which structural conditions the benefit is retained. Modern cancer genomics has established that tumours are driven by diverse combinations of alterations but often converge on a more limited set of pathways and cellular capabilities (Bailey et al. 2018; Hanahan 2022; Vogelstein et al. 2013; Sanchez-Vega et al. 2018). Our results extend that principle from pathway membership to network information flow: the biological knowledge that best organizes propagation through a cancer module differs among tumour types.
4.1. Semantic Architecture Provides a Functional Layer Above Topology
The central concept emerging from this analysis is the semantic network architecture, represented by the branch-specific SDI vector . This vector does not describe which genes are mutated; instead, it summarizes which type of functional information most improves reconstruction of the mutation-defined module once the same interactome is held fixed. In this sense, semantic architecture is orthogonal to conventional classifications based on tissue of origin, gene expression, copy number, immune state or driver frequency (Hoadley et al. 2018; Liu et al. 2018; Thorsson et al. 2018). It captures a property of the relationship between a tumour’s altered genes and the structured biological knowledge surrounding them.
The observation that MF and BP were most frequently optimal, with CC preferred in a smaller but non-negligible subset and the combined representation optimal in only two cancers, argues against the assumption that more semantic information is always better. Averaging BP, MF and CC can dilute a branch-specific signal when one aspect of biology is disproportionately informative. Branch identity should therefore be treated as biological information rather than as a nuisance parameter. This point is relevant beyond the present method: knowledge-guided machine-learning and graph models may benefit from preserving ontology axes instead of concatenating them into a single undifferentiated feature space.
4.2. The Pan-Cancer Effect-Size Landscape Is Biologically Heterogeneous
The forest plot provides a second level of biological information beyond statistical significance. SKCM and TGCT occupy the extreme positive end of the distribution, while a broad middle group—including PRAD, LUSC, READ, THCA, THYM, COAD, UCEC and others—shows smaller but reproducible gains. At the opposite end, BRCA and SARC are close to the null, and KICH and UCS show uncertainty compatible with little or no combined benefit (Figure 1b; Supplementary Figure S2). This continuum is more informative than dividing cancers into “significant” and “non-significant” groups. It suggests that the extent to which functional annotations add information beyond topology is itself a tumour-level property.
Several biological mechanisms could generate such a gradient. Cancers with highly distributed mutational landscapes may contain many altered genes that are not immediate neighbours in the interactome but converge on related functions, making semantic weighting especially useful. Conversely, cancers in which the most recurrent alterations are already concentrated in dense interaction neighbourhoods may leave less room for semantic information to improve ranking. Molecular heterogeneity can also attenuate a combined signal: when distinct subtypes perturb different processes, averaging BP, MF and CC across all tumours may blur subtype-specific coherence. These explanations remain hypotheses, but the effect-size landscape provides a concrete framework for testing them against independent pan-cancer variables.
4.3. Biological Interpretation of BP, MF and CC Dependencies
A preferential BP signal is naturally interpreted as evidence that the mutation module is organized around coordinated biological programmes rather than around a small set of locally connected proteins. Cancer-driving alterations frequently converge on processes such as cell-cycle control, apoptosis, DNA repair, differentiation, invasion and immune modulation even when the affected genes differ across patients or tumour types (Hanahan 2022; Sanchez-Vega et al. 2018; Knijnenburg et al. 2018). In such settings, BP semantics can connect proteins participating in the same programme despite separation in the physical interaction graph. A strong BP contribution may therefore be a systems-level signature of pathway convergence.
MF captures a different biological axis. Two proteins can perform related biochemical roles—for example kinase activity, DNA binding, receptor activity, ubiquitin ligase activity or catalytic functions—without belonging to the same immediate interaction neighbourhood. The frequent optimality of MF is compatible with the idea that cancers can preserve or repeatedly perturb biochemical capabilities through alternative molecular implementations. This is particularly relevant to oncogenic signalling, where different members of receptor, kinase or transcriptional-regulatory families may produce related downstream consequences. The TCGA pathway landscape demonstrates that diverse genomic mechanisms repeatedly affect a relatively constrained set of signalling systems (Sanchez-Vega et al. 2018); MF-oriented propagation offers a complementary way to capture that functional redundancy.
CC dependence emphasizes spatial and macromolecular organization. Many cancer-relevant events are compartment-specific: receptor signalling is initiated at membranes, genome maintenance and transcription occur in nuclear and chromatin contexts, mitochondrial processes influence metabolism and apoptosis, and multiprotein complexes impose stoichiometric and spatial constraints on function. Proteogenomic analyses show that genomic alterations can be propagated through protein abundance, phosphorylation and signalling states rather than remaining confined to the mutated gene itself (Mertins et al. 2016). Likewise, experimentally defined cancer-associated protein complexes, including hormone-responsive nuclear receptor interactors, illustrate how cellular localization and complex membership can organize disease-relevant function (Nassa et al. 2011). A preferential CC signal may therefore indicate that the disease module is particularly coherent at the level of compartment, complex or subcellular machinery.
These interpretations are hypotheses generated from the structure of GO and from established cancer biology; they should not be mistaken for direct measurements of pathway activity. A high BP SDI does not prove that a specific biological process is activated, and a high CC SDI does not prove altered localization. The semantic architecture instead identifies the level of biological annotation that most effectively organizes the observed mutation module. Mechanistic interpretation requires downstream enrichment, expression, phosphoproteomic, spatial or experimental validation.
4.4. Module-Size Dependence Is Consistent with Pathway Convergence and Mutational Heterogeneity
Semantic gain increased monotonically as modules expanded from 50 to 400 proteins (Figure 1c; Supplementary Figure S3). This is one of the most biologically informative results of the study. At , the gain is small for every semantic configuration, suggesting that topology already captures much of the local structure surrounding the most recurrently altered genes. At the semantic contribution becomes clearer, and at BP and MF reach mean gains above 0.02 NDCG@100, with the combined and CC representations also increasing substantially. The widening error bars at larger sizes indicate cross-cancer heterogeneity in the magnitude of this effect, but not a reversal of its overall direction.
Small mutation modules are enriched for the most recurrently altered genes, which often occupy recognizable local neighbourhoods and can therefore be recovered effectively from topology alone. Larger modules incorporate lower-frequency events and a broader spectrum of molecular alterations. Cancer genome studies have repeatedly shown that this long tail is extensive and that mutation frequencies vary markedly across and within tumour types (Vogelstein et al. 2013; Lawrence et al. 2013). Network analyses have further shown that rare alterations become more interpretable when considered as combinations affecting common pathways and complexes (Leiserson et al. 2015; Hofree et al. 2013).
The size effect is therefore consistent with pathway-level convergence. As the module broadens, topological density is expected to decrease because functionally convergent genes need not be direct or near-direct interactors. Semantic weighting can recover part of this lost coherence by favouring interaction routes whose endpoints share process, function or compartment information. Importantly, the strongest increase is observed for BP and MF, suggesting that the biological coherence of larger mutation sets may be expressed particularly through shared programmes and biochemical activities rather than through physical proximity alone. This interpretation remains mechanistic rather than causal, but it generates a testable prediction: semantic-rescued proteins from the 400-gene modules should show stronger enrichment for independent pathway and functional-class annotations than topology-only rescues.
4.5. Robustness to Edge Loss Supports Complementarity Between Semantics and the Interactome
The semantic advantage was largely retained after random deletion of 20% of interactome edges (Figure 1d; Supplementary Figure S4). All four semantic configurations show only a small downward shift after perturbation, and the relative ordering of BP, MF, combined and CC remains similar. This behaviour is relevant because the human interactome is incomplete, unevenly sampled and biased toward well-studied proteins and experimental contexts (Oughtred et al. 2021; Menche et al. 2015). A method that relies on a small number of specific edges can produce unstable biological conclusions when the interaction map changes. In PanCancerSemNet, ontology information acts as a complementary prior on existing edges and appears to preserve functional coherence when a fraction of topology is removed.
The robustness result also helps interpret the larger module-size effect. If semantic gain arose primarily from exploiting a few high-information edges, increasing module size while removing 20% of interactions would be expected to destabilize performance. Instead, semantic improvement persists, consistent with a distributed signal encoded across many functionally coherent routes. This does not prove robustness to all interactome uncertainty, because random edge deletion is an intentionally simple perturbation model. Missing interactions are not random, and high-degree hubs, tissue-specific interactions and low-confidence edges may be affected differently. Accordingly, non-random perturbation schemes and replication on alternative interactomes would provide stronger tests of biological robustness than additional random edge deletions alone.
4.6. Pan-Cancer Biological Implications
The cross-cancer heterogeneity of semantic gain suggests several biological uses for the framework. First, semantic architecture may provide a systems-level descriptor for comparing cancers that are genomically distinct but functionally convergent. TCGA studies have shown that tissue of origin, oncogenic pathway activity, immune state and DNA-repair deficiency define partially overlapping axes of tumour diversity (Hoadley et al. 2018; Sanchez-Vega et al. 2018; Thorsson et al. 2018; Knijnenburg et al. 2018). The SDI profile offers an additional axis: not which pathway is altered, but which level of functional organization best reconstructs the altered network.
Second, branch-specific propagation may improve candidate-gene prioritization. A topology-only method favours proteins close to known cancer genes, whereas semantic-RWR can elevate proteins connected through biologically coherent functions or compartments. This could be valuable for long-tail cancer alterations, where individual genes are rarely mutated but participate in recurrent disease programmes. Network-based cancer analyses have already demonstrated that rare mutations can become interpretable when aggregated into pathways and complexes (Leiserson et al. 2015); semantic weighting adds a principled way to distinguish functionally meaningful from merely topological proximity.
Third, semantic architecture could inform the design of downstream multi-omic analyses. A cancer with strong BP dependence may motivate pathway-activity and transcriptional analyses; strong MF dependence may motivate examination of enzyme, kinase or receptor classes; strong CC dependence may motivate proteomic, phosphoproteomic, complex-level or spatial analyses. Proteogenomic studies demonstrate that the consequences of somatic alterations can be more visible at protein and phosphorylation levels than at DNA level alone (Mertins et al. 2016). Accordingly, the most compelling next step is not to assign a biological label from GO alone, but to test whether branch preference predicts concordant evidence in RNA, protein, phosphoprotein or spatial data.
Finally, the framework has potential translational relevance but is not yet a clinical predictor. If semantic architectures prove stable across independent cohorts and molecular modalities, they could help prioritize pathway-consistent biomarkers, suggest mechanistically coherent combination targets or identify cancers with shared functional organization despite different driver genes. Such applications would require external validation, prospective association with therapeutic response and careful control for annotation bias before any clinical interpretation.
4.7. Cancer-Specific Signals Should Be Interpreted as Hypotheses, not Labels
The strongest combined gains were observed in SKCM and TGCT, and the heatmap shows that both cancers also retain positive branch-specific signals (Figure 1a,b). This broad positivity suggests that, in these modules, semantic coherence is not confined to a single GO axis. In contrast, BRCA, SARC, KICH and UCS did not show corrected evidence of a combined semantic advantage, and the heatmap reveals that their branch-specific profiles are not identical: a near-null combined value can coexist with positive or negative contributions from individual ontology branches. This distinction is important because failure of the combined representation should not be interpreted as absence of functional organization; it may instead indicate that averaging semantic axes dilutes a more selective signal.
These cancer-to-cancer differences are unlikely to be explained by a single biological variable. Tumours differ in mutation burden, mutational processes, lineage constraints, pathway redundancy, immune composition and the extent to which recurrent alterations map onto well-annotated proteins (Lawrence et al. 2013; Thorsson et al. 2018). In a mutation-rich cancer such as melanoma, semantic weighting may help organize a dispersed set of recurrent alterations into coherent functional neighbourhoods. In molecularly heterogeneous entities, a tumour-type aggregate may instead combine distinct subtype-specific programmes. The present data support these as testable interpretations rather than established mechanisms. Cancer-specific follow-up should therefore relate SDI profiles to mutation burden, pathway scores, histological or molecular subtype, immune subtype, DNA-repair phenotype and independent proteogenomic measurements.
4.8. Methodological Limitations
Several limitations define the scope of the conclusions. First, mutation recurrence is a reproducible pan-cancer module source but is not equivalent to driver probability, pathway activity or protein abundance. Large genes, tumour-specific background mutation rates and mutational signatures can influence recurrence (Lawrence et al. 2013). Second, GO annotations are incomplete and non-uniform. Semantic similarity depends on ontology structure and annotation density, both of which reflect historical research effort as well as biology (Pesquita et al. 2009). Third, the interactome is incomplete and biased toward well-studied proteins (Menche et al. 2015). Fourth, the benchmark contains two graph-perturbation seeds and five module-subsampling seeds. Replicate-aware statistics reduce pseudo-replication but do not replace deeper perturbation sampling.
Fifth, leave-one-out reconstruction evaluates internal module coherence rather than prospective discovery of previously unknown cancer genes. Strong held-out recovery demonstrates that semantic information organizes known mutation-defined modules, but independent candidate validation is required to establish discovery utility. Stronger external validation will require independent cohorts, alternative interactomes and orthogonal molecular modalities. CPTAC-style proteogenomic data are especially attractive because they can test whether semantic branch preference is reflected in protein abundance and signalling states (Mertins et al. 2016).
4.9. Future Directions
A direct next step is external validation of the branch-specific SDI profiles in independent cohorts and across mutation, expression, copy-number and proteomic module definitions. This would determine whether semantic architecture is stable to cohort composition and molecular modality rather than being specific to the present mutation-recurrence benchmark. Complementary analyses on alternative interactomes and non-random graph perturbations could further test sensitivity to network construction.
A second direction is to move from descriptive branch preference to mechanistic annotation. For each cancer, the proteins whose rank changes most strongly under BP, MF or CC weighting can be analysed for pathway enrichment, complexes, essentiality, druggability and cancer-specific expression. This would connect a global semantic architecture to specific biological mechanisms. A third direction is to compare diffusion with learned ontology-conditioned graph models. Such models may capture nonlinear interactions among topology and semantic dimensions, but they should be benchmarked against the transparent RWR formulation rather than replacing it without evidence.
More broadly, the results suggest a general disease-systems hypothesis: diseases may differ not only in the location of their molecular modules in the interactome, but also in the semantic dimension that best organizes those modules. Testing this idea across cancer and non-cancer diseases could lead to a semantic atlas of disease modules in which topology, process, function and cellular context are treated as complementary components of molecular pathophysiology.
5. Conclusions
Ontology-aware network propagation improves reconstruction of mutation-defined molecular programmes across most of 31 TCGA cancers, with the strongest gains emerging as modules become larger and more distributed and with much of the advantage retained after substantial random interactome perturbation. The contribution of BP, MF and CC differs among cancers, supporting tumour-specific semantic network architectures rather than a universal semantic effect. Biologically, these architectures provide a systems-level view of whether cancer-associated alterations are most coherently organized as coordinated processes, recurring biochemical capabilities or compartment- and complex-level structures. If validated across independent cohorts and molecular modalities, semantic architecture could become a useful bridge between mutation-centric cancer genomics and functional network biology, enabling more biologically informed prioritization of pathways, candidate genes and multi-omic follow-up studies.
Supplementary Materials
The following supporting information can be downloaded at the website of this paper posted on Preprints.org.
Author Contributions
P.H.G., P.V. and T.M.: Conceptualization, Methodology, Software, Formal analysis, Investigation, Visualization, Writing–original draft, Writing–review and editing, and Supervision. All authors read and approved the final manuscript.
Data Availability Statement
The analysis uses open-access TCGA masked somatic mutation data from the NCI Genomic Data Commons and public molecular-interaction and ontology resources. The PanCancerSemNet source code, analysis outputs, statistical-analysis scripts, figures and machine-readable result tables will be deposited in a public archival repository. The repository will document the software environment, source-data provenance and commands required to reproduce the analyses reported here. Large third-party source datasets will not be redistributed when their licenses prohibit redistribution; download manifests, identifiers and provenance records will be provided instead.
Conflicts of Interest
The authors declare no competing interests.
References
- Barabasi, A-L, N Gulbahce, and J. Loscalzo. 2011. Network medicine: a network-based approach to human disease. Nature Reviews Genetics 12: 56–68. [Google Scholar] [CrossRef] [PubMed]
- Cowen, L, T Ideker, BJ Raphael, and R. Sharan. 2017. Network propagation: a universal amplifier of genetic associations. Nature Reviews Genetics 18: 551–562. [Google Scholar] [CrossRef] [PubMed]
- Nassa, G, R Tarallo, C Ambrosino, A Bamundo, L Ferraro, O Paris, M Ravo, PH Guzzi, M Cannataro, M Baumann, and et al. 2011. A large set of estrogen receptor β-interacting proteins identified by tandem affinity purification in hormone-responsive human breast cancer cell nuclei. Proteomics 11, 1: 159–165. [Google Scholar] [CrossRef] [PubMed]
- Cho, YR, M Mina, Y Lu, N Kwon, and PH Guzzi. 2013. M-finder: Uncovering functionally associated proteins from interactome data integrated with GO annotations. Proteome Science 11 Suppl 1: S3. [Google Scholar] [CrossRef] [PubMed]
- Hoadley, KA, and et al. 2018. Cell-of-Origin Patterns Dominate the Molecular Classification of 10,000 Tumors from 33 Types of Cancer. Cell 173, 2: 291–304.e6. [Google Scholar] [CrossRef] [PubMed]
- Bailey, MH, and et al. 2018. Comprehensive Characterization of Cancer Driver Genes and Mutations. Cell 173, 2: 371–385.e18. [Google Scholar] [CrossRef] [PubMed]
- Liu, J, and et al. 2018. An Integrated TCGA Pan-Cancer Clinical Data Resource to Drive High-Quality Survival Outcome Analytics. Cell 173, 2: 400–416.e11. [Google Scholar] [CrossRef] [PubMed]
- Gene Ontology Consortium. 2023. The Gene Ontology knowledgebase in 2023. Genetics 224, 1: iyad031. [Google Scholar] [CrossRef] [PubMed]
- Oughtred, R, and et al. 2021. The BioGRID database: A comprehensive biomedical resource of curated protein, genetic, and chemical interactions. Protein Science 30, 1: 187–200. [Google Scholar] [CrossRef] [PubMed]
- Balu, S, and et al. 2025. Complex Portal 2025: predicted human complexes and enhanced visualisation tools for the comparison of orthologous and paralogous complexes. Nucleic Acids Research 53, D1: D644–D650. [Google Scholar] [CrossRef] [PubMed]
- Hanahan, D. 2022. Hallmarks of Cancer: New Dimensions. Cancer Discovery 12, 1: 31–46. [Google Scholar] [CrossRef] [PubMed]
- Vogelstein, B, N Papadopoulos, VE Velculescu, S Zhou, LA Diaz, Jr., and KW Kinzler. 2013. Cancer genome landscapes. Science 339, 6127: 1546–1558. [Google Scholar] [CrossRef] [PubMed]
- Lawrence, MS, P Stojanov, P Polak, and et al. 2013. Mutational heterogeneity in cancer and the search for new cancer-associated genes. Nature 499: 214–218. [Google Scholar] [CrossRef] [PubMed]
- The Cancer Genome Atlas Research Network, JN Weinstein, EA Collisson, GB Mills, and et al. 2013. The Cancer Genome Atlas Pan-Cancer analysis project. Nature Genetics 45: 1113–1120. [Google Scholar] [CrossRef] [PubMed]
- Sanchez-Vega, F, M Mina, J Armenia, and et al. 2018. Oncogenic signaling pathways in The Cancer Genome Atlas. Cell 173, 2: 321–337.e10. [Google Scholar] [CrossRef] [PubMed]
- Leiserson, MDM, F Vandin, HT Wu, and et al. 2015. Pan-cancer network analysis identifies combinations of rare somatic mutations across pathways and protein complexes. Nature Genetics 47: 106–114. [Google Scholar] [CrossRef] [PubMed]
- Hofree, M, JP Shen, H Carter, A Gross, and T. Ideker. 2013. Network-based stratification of tumor mutations. Nature Methods 10: 1108–1115. [Google Scholar] [CrossRef] [PubMed]
- Menche, J, A Sharma, M Kitsak, SD Ghiassian, M Vidal, J Loscalzo, and A-L. Barabasi. 2015. Uncovering disease–disease relationships through the incomplete human interactome. Science 347, 6224: 1257601. [Google Scholar] [CrossRef] [PubMed]
- Thorsson, V, DL Gibbs, SD Brown, and et al. 2018. The immune landscape of cancer. Immunity 48, 4: 812–830.e14. [Google Scholar] [CrossRef] [PubMed]
- Knijnenburg, TA, L Wang, MT Zimmermann, and et al. 2018. Genomic and molecular landscape of DNA damage repair deficiency across The Cancer Genome Atlas. Cell Reports 23, 1: 239–254.e6. [Google Scholar] [CrossRef] [PubMed]
- Mertins, P, DR Mani, KV Ruggles, and et al. 2016. Proteogenomics connects somatic mutations to signalling in breast cancer. Nature 534: 55–62. [Google Scholar] [CrossRef] [PubMed]
- Pesquita, C, D Faria, AO Falcao, P Lord, and FM Couto. 2009. Semantic similarity in biomedical ontologies. PLoS Computational Biology 5, 7: e1000443. [Google Scholar] [CrossRef] [PubMed]
- Wang, JZ, Z Du, R Payattakool, PS Yu, and C-F. Chen. 2007. A new method to measure the semantic similarity of GO terms. Bioinformatics 23, 10: 1274–1281. [Google Scholar] [CrossRef] [PubMed]
Figure 1.
Pan-cancer benchmark results. a, Semantic Dependency Index (SDI) heatmap for 31 cancers across combined, BP, MF and CC semantic configurations. Red denotes positive semantic dependency and blue denotes conditions in which topology-only propagation performs better. b, Mean combined semantic gain in NDCG@100 for each cancer with cluster-bootstrap 95% confidence intervals; the vertical line at zero denotes no difference between semantic-RWR and topology-only RWR. c, Mean semantic gain as a function of cancer-module size (50, 200 and 400 proteins), showing a marked increase in semantic advantage for larger modules. d, Mean semantic gain on the intact interactome and after random removal of 20% of edges, demonstrating that the semantic advantage is largely retained under graph perturbation. Error bars in panels c and d summarize variation across cancers. Standalone high-resolution versions of panels a–d are provided as Supplementary Figures S1–S4.
Figure 1.
Pan-cancer benchmark results. a, Semantic Dependency Index (SDI) heatmap for 31 cancers across combined, BP, MF and CC semantic configurations. Red denotes positive semantic dependency and blue denotes conditions in which topology-only propagation performs better. b, Mean combined semantic gain in NDCG@100 for each cancer with cluster-bootstrap 95% confidence intervals; the vertical line at zero denotes no difference between semantic-RWR and topology-only RWR. c, Mean semantic gain as a function of cancer-module size (50, 200 and 400 proteins), showing a marked increase in semantic advantage for larger modules. d, Mean semantic gain on the intact interactome and after random removal of 20% of edges, demonstrating that the semantic advantage is largely retained under graph perturbation. Error bars in panels c and d summarize variation across cancers. Standalone high-resolution versions of panels a–d are provided as Supplementary Figures S1–S4.

Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.