Computer Science and Mathematics

Sort by

Article
Computer Science and Mathematics
Mathematical and Computational Biology

Francesco Canonaco

,

Enzo Acerbi

,

Fabio Stella

Abstract: As microbiome research increasingly seeks to identify true ecological shifts, transitioning from associational to causal approaches is essential. However, detecting structural changes across independent networks remains challenging due to the absence of established biological ground truths and the small, imbalanced sample sizes typical of microbiome cohorts. To address this, we extend an existing network comparison framework to enable node-level mechanism-shift detection under the direct linear non-Gaussian acyclic model (DirectLiNGAM). We evaluate the Naive, Bootstrap, and Relative sample size Bootstrap Stability (RSBS) estimators across extensive synthetic discovery runs and a semi-synthetic, batch-corrected human gut microbiome cohort. Our results demonstrate that resampling-based estimation consistently outperforms a single-fit Naive baseline by trading marginal recall for substantial precision gains. On both semi-synthetic and synthetic data, standard Bootstrap is optimal for comparing datasets of equal size, whereas RSBS is the only estimator that reliably handles imbalanced cohorts. This advantage strengthens as network dimensionality increases and persists under authentic compositional noise. Navigating this complex and emerging research area is currently constrained by limitations in data quantity and quality, scarce biological knowledge, and a lack of dedicated software. To address these critical gaps, we provide the complete benchmark pipeline and synthetic data generators as an open-access Python package, causal-comparator, to support node-level mechanism-shift detection across systems biology applications.

Brief Report
Computer Science and Mathematics
Mathematical and Computational Biology

Pietro Hiram Guzzi

Abstract: Motivation: Multiomics data integration is essential for understanding complex biological systems, yet existing network-based approaches operate at a single level—either patient similarity networks or molecular interaction networks—without bridging the two. We address the gap between patient-level integration and molecular-level network analysis by proposing a hierarchical framework that unifies both. Results: We introduce HiNoN (Hierarchical Network-of-Networks), a two-level framework in which a fused patient-similarity network is linked to subgroup-specific microbial association and host-gene co-expression networks. Across 200 simulation runs, subgroup recovery was strongest for large cohorts (mean ARI 0.95 at n=500) and two-group settings (mean ARI 0.88), but declined with increasing noise and subgroup number. We then performed leakage-free repeated nested cross-validation on 50 paired coronary artery disease samples. Under the prespecified 16-component representation, early concatenation outperformed HiNoN-path2-SGC in ROC–AUC (0.923 versus 0.911), PR–AUC (0.958 versus 0.949), and calibration. Sensitivity analysis showed a different pattern under aggressive compression: with eight components, HiNoN-path2-SGC improved ROC--AUC by 0.073 and also improved balanced accuracy, Brier score, and calibration error. Unsupervised clusters had negligible agreement with disease labels (ARI =-0.006; NMI =0.016) and are therefore treated as exploratory. These results characterize HiNoN as a topology-aware regularizer under representation scarcity, rather than a uniformly superior classifier.

Article
Computer Science and Mathematics
Mathematical and Computational Biology

J. Wirsam

,

B. Wirsam

,

C. Leitzmann

Abstract: Since 1995, the intake of individual nutrients has been evaluated using fuzzy logic, with aggregation of the resulting fuzzy sets enabling the weighing of one nutrient against another, and the evaluation and optimization of overall nutrient intake. The aggregation operator used in earlier work, defined as the product of the minimum and harmonic mean operators (MIN×H), performed well in practice but had undesirable mathematical properties: it is not associative, model extensions can shift results even when all conditions are fully satisfied, and the resulting evaluation field contains discontinuities inconsistent with biological behavior. We therefore reviewed the family of parametrized T-norms for an alternative and identified the Schweizer (3) operator, fitting its parameter to best match the existing operator using 254 real nutrition protocols from a sustainability study; strong agreement was achieved at a parameter setting of p = 6. The new operator satisfies all required mathematical properties and produces a smooth, simply structured evaluation field that reveals synergistic and substitution effects among foodstuffs, as well as the essentiality of certain foodstuffs under specific conditions. The resulting evaluation field shows a wide, high-dimensional plateau rather than a sharp peak, interpretable as stability, error tolerance, and a high degree of diversity, offering a new, mathematically robust perspective on healthy nutrition and the interactions among foodstuffs.

Review
Computer Science and Mathematics
Mathematical and Computational Biology

Raed I. Seetan

,

Omar Darwish

Abstract: Machine learning has become a central framework for analyzing high-dimensional biological data generated by genomics, transcriptomics, proteomics, single-cell sequencing, and multi-omics technologies. However, predictive modeling is frequently limited by severe class imbalance, where biologically important populations including rare disease subtypes, uncommon molecular states, and low-frequency cell populations are underrepresented. Synthetic oversampling addresses this challenge by increasing minority-class representation, with the Synthetic Minority Oversampling Technique (SMOTE) being the most widely used approach. However, conventional SMOTE assumes that geometric proximity reflects biological similarity, an assumption often violated in high-dimensional, nonlinear, and heterogeneous biological data. This review examines clustering-guided synthetic generation methods that integrate unsupervised learning before oversampling to preserve underlying biological structure. We evaluate partition-based, density-aware, fuzzy, cluster-filtering, and representation-based approaches, emphasizing their mathematical foundations, computational assumptions, and applications in molecular classification, rare population analysis, single-cell modeling, and multi-omics integration. Key challenges including cluster instability, preservation of rare biological populations, synthetic data validation, and distinguishing statistical rarity from biological significance are discussed. Finally, we outline emerging directions involving biological foundation models, graph-based learning, multi-omics representation spaces, and biologically constrained generative frameworks. These advances represent a shift from conventional data balancing toward structure-aware synthetic modeling that preserves meaningful biological organization while improving predictive performance.

Brief Report
Computer Science and Mathematics
Mathematical and Computational Biology

Pietro Hiram Guzzi

,

Tommaso Mazza

,

Pierangelo Veltri

Abstract: Cancer-associated alterations perturb molecular systems whose organization is only partially captured by protein–protein interaction topology. We asked whether Gene Ontology (GO)- derived semantic information can improve reconstruction of mutation-defined cancer modules and, critically, whether the magnitude and ontology source of this improvement vary across tumour types. We constructed recurrent non-silent somatic-mutation modules for 31 non-haematological The Cancer Genome Atlas (TCGA) projects and embedded them in a frozen human interaction network enriched with a nine-dimensional GO semantic representation. Topology-only random walk with restart (RWR) was compared with semantic-RWR under symmetric leave-one-out module reconstruction. The benchmark comprised 7,440 evaluations spanning 31 cancers, module sizes of 50, 200 and 400 proteins, intact or 20% edge-depleted interactomes, two graph perturbation seeds, five module-subsampling seeds, and four semantic configurations (combined, Biological Process, Molecular Function and Cellular Component). Combined semantic diffusion significantly improved NDCG@100 in 27 of 31 cancers after Holm correction. Semantic gain increased strongly with module size and was largely preserved after removal of 20% of interactome edges. The optimal ontology branch differed among cancers: Molecular Function was optimal in 13, Biological Process in 12, Cellular Component in 4, and the combined representation in 2. These findings indicate that cancer modules are not organized by topology alone and support the existence of tumour-specific semantic network architectures: systems-level patterns describing whether functional coherence is expressed primarily through biological processes, biochemical activities, cellular compartments, or combinations thereof. Biologically, this heterogeneity is consistent with the known convergence of diverse cancer mutations onto shared pathways, complexes and cellular programs. PanCancerSemNet therefore provides a framework for interpreting cancer-associated alterations as context-dependent perturbations of functionally annotated molecular systems rather than as isolated mutated genes.

Article
Computer Science and Mathematics
Mathematical and Computational Biology

Changsoo Shin

Abstract: We propose a mathematical model for the structure of consciousness, and we do not treat the feeling itself. We build the model on a phase-augmented Recursive Heaviside Sequence Function, a cascade of nested smooth step functions in which each layer is one episodic memory carrying a time threshold \(\tau_i\), a sharpness \(s_i\), and an emotional phase \(\theta_i\). We propose that the objective of consciousness is survival, the drive to keep the remaining distance to a fixed endpoint short. Our central claim is that the partial derivatives of the objective functional \(E=|1-u_k|\) are thought, optimization, and decision, formed together at the present instant and without iteration. The state satisfies an advection equation whose right-hand side is a dummy source that the structure generates by itself, and substituting \(t=\tau_k\) removes the time variable and lowers the dimension by one. Each new experience adds an increment that decays exponentially with the delay, which we take to be the mechanism of forgetting and of childhood amnesia. In the second half, planning becomes a finite Taylor expansion about the present point, so proving a theorem, writing fiction, and telling a lie are one operation under different objectives. A perfect lie would need infinitely many partial derivatives, and no finite agent can supply them.

Article
Computer Science and Mathematics
Mathematical and Computational Biology

Jianghui Xiong

,

Qianchen Xia

Abstract: Virtual patients need a useful coordinate system before they can model patient-specific dynamics. Most models in AI for science operate at molecular, cellular, or organ-specific scales. However, treatment decisions are made across a whole person and across different forms of intervention. Here, we use drug repurposing as a human-scale testbed for a central virtual-patient question: can a structured representation of biological directions organize intervention-relevant knowledge well enough to prioritize known drug-disease relationships?We introduce SteeraMed Bench, a framework for evaluating module panels built from a 332-module atlas. The atlas combines extended aging hallmarks, traditional Chinese medicine syndrome proxies, nutraceutical targets, and food-as-medicine targets. Rather than assuming that all modules should form one universal model, the framework compares panels as alternative representations for each disease task. Across 1,916 DrugBank small molecules, five chronic disease tasks, and an exploratory extension to 23 disease categories, the panels carried useful within-benchmark ranking signal. The nutraceutical and nutraceutical-extension panel (NUT+NUTX, 117 modules) achieved a mean recall@20 of 0.494 across five diseases, close to 0.524 for the full atlas. The full atlas was the strictly highest observed configuration in only 9 of 23 disease categories. Different panels were most useful for different tasks, including extended aging hallmarks for type 2 diabetes and osteoporosis, food-as-medicine for depression, and nutraceutical modules for the atherosclerosis/hyperlipidemia task.The framework also evaluates newly proposed gene sets for incremental value and redundancy. In an exploratory LLM-assisted workflow, two refined candidates showed nominal positive increments, but neither remained significant after correction for multiple testing. Performance decreased under target-family-separated evaluation and approached chance in leave-one-disease-out evaluation. SteeraMed Bench lays the coordinate and evaluation foundation on which future patient-specific dynamic models and causal intervention simulations can be built.

Article
Computer Science and Mathematics
Mathematical and Computational Biology

Florin Avram

Abstract: Generalized Lotk–Volterra (GLV) systems, with roots in ecology, constitute one of the most studied classes of positive ODEs. Recently, a reaction-network perspective for a generalization useful in mathematical epidemiology, called block GLV systems, was offered by Adenane, Avram and Halanay (2026). These authors offer a "siphon calculus" in which the boundary stability of block GLV equations is determined by studying (i) invariant faces associated to minimal siphons, (ii) transversal Jacobians, (iii) invasibility expressed via R-invasion functions, (iv) relay graphs (v) exclusion partitions and (vi) Lyapunov functions, without leaving the original state space. Another reaction-network perspective for GLV systems was offered by Rojas La Luz, Yu and Craciun, who developed a global stability theory for positive equilibria by introducing associated poly-exponential systems obtained through the logarithmic change of variables \(x_i=e^{\xi_i}\). In these logarithmic coordinates, compatibility classes become affine subspaces and remarkably simple quadratic Lyapunov functions establish global convergence of complex-balanced systems. The purpose of the present paper is to combine the two perspectives. We revisit examples considered by Rojas La Luz, Yu and Craciun, where the use of siphon calculus can be avoided, but we still find it useful, at least for suggesting open problems for block GLV systems.

Article
Computer Science and Mathematics
Mathematical and Computational Biology

Siphokazi Princess Gatyeni

Abstract: Epidemic meningococcal meningitis in the African meningitis belt recurs on two superimposed timescales: a sharp annual dry-season peak and irregular multi-annual epidemics separated by five to twelve years. Because invasive disease is a rare, epidemiologically dead-end outcome of asymptomatic nasopharyngeal carriage, transmission models built on the standard susceptible-infectious template misrepresent the driving process. We formulate a deterministic carriage-structured model in which only carriers transmit, immunity against carriage re-acquisition is leaky, and a conjugate vaccine protects imperfectly and wanes. We show analytically that the basic reproduction number \( R_0 \) is a property of carriage transmission and is decoupled, to first order, from disease incidence, so that reproduction numbers inferred from case notifications estimate the wrong quantity. Using the Castillo-Chavez-Song centre manifold method, we derive, in closed form, the condition for backward bifurcation and prove that it is governed by a threshold \( \varepsilon^\ast \) on carriage-blocking immunity, not by vaccine leakiness or by case-management capacity: the invasive-disease compartment is provably absent from the bifurcation condition. Under seasonal forcing we characterize, through Floquet analysis and two-parameter continuation, the region of immunity-waning and seasonal-amplitude space in which multi-annual recurrence arises, confirming that annual forcing alone cannot generate it. A scenario analysis calibrated to published belt parameter ranges quantifies the burden averted by routine infant immunisation, catch-up campaigns, and improved carriage efficacy, and shows that whether sustained immunisation eliminates epidemics or merely postpones them depends on the same carriage-efficacy threshold that governs bistability.

Review
Computer Science and Mathematics
Mathematical and Computational Biology

kazi Hafiz Md Asad

,

Rafi Majid

,

Md Tanjeelur Rahman Labib

,

Ahsanur Rahman

Abstract: Group-level molecular-interaction evidence is often reduced to pairwise protein-protein interaction graphs before protein-complex detection, potentially discarding experiment membership and over-weighting large groups. Yet prior higher-order studies typically changed the representation, objective, output granularity, and evaluation protocol simultaneously. We isolate the operator-level contribution of higher-order representation by comparing an inverse-size-weighted clique graph, the normalized Zhou hypergraph operator, and a degree-aware penalized hypergraph operator on identical experiment-derived groups. Hyperedges were split by size into training, validation, and test sets; the training split alone defined the protein universe and operators; regularization was selected on validation data; and recovery was evaluated against held-out test hyperedges. Across nine IntAct-derived thematic PSI–MITAB collections and 25 paired initializations per dataset, the penalized hypergraph increased mean held-out symmetric best-match F1 on six datasets after within-dataset Benjamini–Hochberg correction, was indistinguishable on two, and was lower by 0.0015 on Cancer. Positive mean differences ranged from 0.0098 to 0.0356. However, the across-dataset Friedman test did not reach significance (χ2 = 5.20, p = 0.074), and the graph–penalized-hypergraph Nemenyi comparison was non-significant (p = 0.111). The penalized operator had the best mean rank (1.39), but secondary metrics and eigengap-selected cluster counts exposed dataset-dependent trade-offs. Observed recovery exceeded degree-preserving and hyperedge-size-preserving nulls on eight of nine datasets. Gene Ontology coherence was high for both representations, with no aspect-level difference surviving multiplicity correction and substantial pooled term overlap (Jaccard 0.78–0.86). Higher-order modeling is therefore conditionally beneficial rather than universally superior: it is most defensible when group membership is retained, regularization is validation-supported, and output granularity is controlled

Article
Computer Science and Mathematics
Mathematical and Computational Biology

Rim Adenane

,

Florin Avram

,

Andrei-Dan Halanay

Abstract: Persistence, coexistence, and boundary transcritical relays are usually studied through model-specific analyses in mathematical epidemiology, ecology, population dynamics, and chemical reaction network theory. Although these fields address closely related questions, they have developed largely independently. This separation is reflected, for example, in the limited mentions of the multi-strain epidemiologic models in ecology’s chemostats and gradostats literature, despite the fact that these are revealed to be very similar once the concept of siphons from chemical reaction network theory is integrated. Conversely, the next-generation matrices and invasion graphs from eco-epidemiology are not mentioned in chemical reaction network theory. Our contribution is firstly conceptual, terminological and definitional: we propose a common framework for the study of boundary phenomena in all positive ODEs subfields. We introduce and formalize notions like reproduction and invasion functions attached to siphon faces, relay graphs, relay tables, and boundary transcritical relays. Some of these concepts are known in one of the above fields but largely absent from the others, while others appear to be new; taken together, they suggest a common language for the analysis of boundary phenomena in positive dynamical systems. The usefulness of the framework is illustrated on multi-strain epidemic models like the Feng-Gavish model, for which we derive explicit, testable conditions. For example, coexistence requires the less fit strain to invade the fitter strain’s equilibrium (for this model, mutual invasibility also ensures coexistence, but explicit further assumptions under which one or the other criterion works for a larger class of models are still unknown). Our approach rests on four pillars: (i) Siphon (a CRN concept) geometry, namely the fact that forward-invariant coordinate faces correspond to siphons, with the disease-free face being the intersection of minimal siphons. (ii) The recently established fact that a transversal Jacobian block on a siphon face is Metzler, which puts under spotlight the roles of its Perron eigenvectors. (iii) A bifurcation theorem linking eigenvalue crossing at a boundary transcritical invasion relay to the emergence of a positive branch on an adjacent face. (iv) Next-generation matrices (NGMs), an MEconcept: on siphon faces, NGMs may be defined via regular splittings, and invasibility may be determined by comparing their spectral radii to > 1.

Article
Computer Science and Mathematics
Mathematical and Computational Biology

Jiajun Hong

,

Xiong You

,

Zihao Ji

,

Hengmin Lv

Abstract: Existing models of the circadian clock in Arabidopsis thaliana are conventionally formulated by integer-order ordinary differential equations (ODEs), which inherently lack the capacity to adequately capture non-local history-dependent regulatory dynamics that arise from sequential biochemical processes such as transcription, translation and protein degradation. Here we construct a Caputo fractional-order model of the Arabidopsis core circadian system and establish local existence and uniqueness, non-negativity, and boundedness of solutions under non-negative initial conditions. Model parameters are fitted to wild-type mRNA expression profiles collected under a standard 12 h light/12 h dark (12L12D) photoperiod. Without any subsequent refitting of parameters, the predictive performance of the fractional-order model is validated on two independent test datasets: wild-type expression time series under three additional photoperiod regimes, and publicly available expression data for major circadian clock loss-of-function mutants. Compared with the original ODE counterpart, the fractional-order formulation exhibits substantially improved performance in reproducing the post-peak decay kinetics and extended tough phase of the PRR5/TOC1 regulatory module. Quantitative error evaluation confirms that the fractional-order model achieves consistently lower mean squared errors across all four photoperiod conditions, and outperforms the integer-order counterpart in two of the four tested mutant backgrounds, indicating that performance gains are not uniformly distributed across all genetic perturbations. Through Matignon-type stability analysis and extensive numerical simulations, we identify a well-defined critical fractional-order threshold. When the fractional order exceeds this critical value, the system’s unique positive equilibrium loses its stability, giving rise to sustained oscillations. From a systems biology perspective, the introduction of fractional-order operators provides a compact phenomenological representation of aggregated historical memory effects and may influence the amplitude, phase, and long-term robustness of the core circadian-clock oscillations. This work offers a new mathematical framework for refining the dynamical characterization of eukaryotic circadian pacemakers.

Article
Computer Science and Mathematics
Mathematical and Computational Biology

Zhangchi Xu

,

Rong Pan

,

Tianzhou Ma

Abstract: Non-coding RNAs (ncRNAs) play important roles in various biological processes primarily via regulating gene expression at multiple levels. Nevertheless, compared to protein-coding genes, ncRNAs are relatively understudied, characterized by low expression, large variation, and more dynamic expression across different tissues and cell types, which creates unique challenges in differential expression (DE) analysis. We propose a new DE method specifically for ncRNAs based on a composite quantile regression model for count data, in which we combine multiple quantiles and select the best subset of quantiles that most differentiate the expression levels between conditions. The proposed method is flexible, robust and capable of capturing the typically low-count, multi-modal and wide-spread distribution of ncRNA count data. We showed in simulations that our method improved the detection power of ncRNAs while controlling the false positive rates as compared to existing DE tools especially when counts are low. Critical ncRNAs and biologically interpretable targeted pathways were identified when we further applied our method to a Smart-seq-total single-cell RNA-seq dataset and a human organ developmental RNA-seq dataset. An R package to implement CQRM and the codes for this study are publicly available: https://github.com/iamverywell/CQRM.

Case Report
Computer Science and Mathematics
Mathematical and Computational Biology

Karen Capano

,

Valentina Carbonari

,

Pierangelo Veltri

,

Pietro Hiram Guzzi

Abstract: Nowadays, the complexity of electronic health records (EHRs) requires tools capable of efficiently and accurately extracting and interpreting clinically relevant information to support clinicians. This study explores the use of the Cheshire Cat AI framework, configured with Ollama and using LLaMA3 as a language model, with the main purpose of performing automatic analysis of synthetic EHRs from Kaggle. Through specific structured queries, the model was able to successfully reconstruct patients’ clinical histories and extracted useful data such as diagnoses, treatments, visits, comorbidities and demographic data. A validation process through repeated queries was then performed, which confirmed a high level of accuracy. To preserve data privacy, only synthetic datasets were used in this work. Beyond the simple retrieval of information by means of queries, the study highlights the great potential of language models in clinical decision support. Their ability to interpret large and heterogeneous datasets certainly offers new opportunities to improve diagnostic accuracy, simplify workflows and personalise treatments. Specifically, natural language queries by tools such as Cheshire Cat AI can be used for intelligent support systems that can, for instance, integrate multimodal and real-time data to provide medical recommendations. These results represent a first step towards the exploitation of large language models not only for EHR analysis, but also to assist in clinical decision-making processes in different medical fields and, above all, for the study of specific complex diseases such as rare diseases.

Article
Computer Science and Mathematics
Mathematical and Computational Biology

Andrea Lomagno

,

Saleh Hamed

,

Ishak Yusuf

,

Pierluigi Luigi Mauri

,

Dario Di Silvestre

Abstract: The demand for user-friendly applications to support biologists in analyzing high-throughput proteomics data remains a pressing challenge. Given the complexity and the multiple intermediate steps involved, this process is time-consuming and often requires specialized computational skills. To simplify and accelerate the exploration of proteomics data, we present PiProteline, an R package designed to operate on high-dimensional data matrices assembled from the output of any search engine commonly used in bottom-up proteomics experiments. In addition to data preprocessing and descriptive statistics, PiProteline enables label-free quantitation, functional enrichment, and systems biology analyses. Notably, it supports both unweighted and weighted protein-protein interaction (PPI) network topological analyses for the identification of critical nodes, such as hubs and bottlenecks. By integrating multiple analytical approaches, PiProteline accelerates the selection of potentially relevant protein signatures, providing insights into the molecular mechanisms characterizing the investigated systems. This information facilitates the formulation of new hypotheses and the design of targeted experiments, ultimately reducing costs and advancing translational medicine. PiProteline is available as an open-source R package via GitHub (https://github.com/lomi95/PiProteline) and as a Shiny application (https://github.com/salehnhd/PiProteline-Shiny).

Article
Computer Science and Mathematics
Mathematical and Computational Biology

Mustapha Olawale Abdulyekeen

,

Blessing Oluwafikayo Adisa

,

Saheed Babatunde Oyetoro

Abstract: Eye diseases such as trachoma, allergic conjunctivitis, and dry eye syndrome have shown increasing prevalence in regions experiencing adverse environmental and climatic changes. Factors such as air pollution, dust exposure, humidity variations, and ultraviolet (UV) radiation directly impact ocular health, especially among vulnerable populations. In this study, we develop a deterministic compartmental model to explore the dynamics of environmentally-driven eye disease transmission and progression. The model integrates climate-sensitive variables, such as dust concentration and humidity, into the transmission and recovery rates of the disease. We analyse the model's equilibria, investigate the basic reproduction number R0, and assess the influence of environmental mitigation Strategies on disease control. Numerical Simulations are provided to illustrate how seasonal and anthropogenic changes in environmental conditions affect disease prevalence over time.

Article
Computer Science and Mathematics
Mathematical and Computational Biology

Prateek Mittal

,

Ayush Srivastava

,

Joohi Chauhan

Abstract: Whole-slide image (WSI) analysis is limited by a familiar mismatch: each slide contains tens of thousands of candidate tissue patches, while supervision is usually available only at slide level. Existing bag-construction strategies tend to address only one side of this problem. Uniform extraction and handcrafted heuristics do not control redundancy, attention-based multiple-instance models couple patch importance to a particular downstream classifier, and coreset methods optimise embedding-space coverage without modelling task-relevant patch quality. We introduce InfoDPP-PAC, a principled patch-selection framework that combines teacher-seeded Gaussian process relevance modelling, determinantal log-determinant diversity, submodular greedy optimisation, and a concentration-based adaptive stopping rule. The main theoretical result shows that the log-determinant diversity term used in DPP-style selection is the Gaussian process mutual information between a selected subset and the latent relevance function. This yields a monotone submodular objective with standard greedy approximation guarantees at fixed budgets. We further derive a PAC-style certificate for residual information gain, allowing the number of retained patches to vary by slide rather than being fixed a priori. The empirical study is deliberately scoped to selection-quality validation: it evaluates whether the selected subset is diverse, spatially and morphologically covering, non-redundant, and enriched for the teacher-derived relevance signal. It does not claim end- to-end diagnostic improvement after retraining a downstream MIL model. On 202 HISTAI gastrointestinal whole-slide images, the adaptive rule uses 83.7% fewer patches on average than a fixed full budget while retaining 97.9% of full-budget composite selection quality. At a matched budget, InfoDPP-PAC achieves the highest mean teacher-derived relevance score among fourteen baselines, with diversity and composite scores close to the strongest coreset methods. The results support InfoDPP-PAC as a controlled quality-diversity-cardinality selection framework, rather than as a downstream clinical predictor.

Article
Computer Science and Mathematics
Mathematical and Computational Biology

Edwin Barrios-Rivera

,

Daiver Cardona-Salgado

,

Ilya Dikariev

,

Carmen A. Ramirez-Bernate

,

Olga Vasilieva

,

Mikhail Svinin

Abstract: This study presents a mathematical modeling framework to analyze the impact of integrating sterilizing treatment into tuberculosis (TB) control strategies, particularly in resource-limited settings. Our findings highlight that while sterilizing treatment alone is highly effective in reducing TB incidence and mortality, its widespread implementation requires significant financial investment. The optimal control approach demonstrates that a mixed strategy, combining sterilizing and non-sterilizing treatments, can achieve comparable public health benefits at a lower cost, especially in the medium-term planning. From a long-term perspective, however, our results suggest that exclusively sterilizing treatment ultimately leads to greater reductions in TB incidence, prevalence, and mortality, justifying its higher initial cost. This is primarily due to the ability of sterilizing drugs to eliminate latent TB infections, prevent future active infections and disease-induced deaths, and avoid a considerable number of treatments. Additionally, under scenarios where the cost of sterilizing drugs decreases over time, a swift transition to a solely sterilizing treatment could result in both epidemiological and economic advantages for healthcare systems.

Review
Computer Science and Mathematics
Mathematical and Computational Biology

Amandeep Jast

,

Gourab Das

Abstract: Genome sequence information is the primordial need for studying species genetics, evolutionary history, disease mechanisms, risk prediction, adaptation and many more. Eventually large-scale initiatives are underway to sequence unknown genomes from various species including humans with various phenotypic states to gain insights into gene function and genetic diversity. However, assembling large genomes remains a computational challenge due to several factors including sequence complexity, continuous growth in sequencing throughput, lack of suitable benchmarking for the selection of optimal combination of tools etc. Additionally, its quality evaluation is another crucial step to perform but varying methods often lead to arbitrary comparisons. Moreover, high sequencing error rates necessitate the error correction and consensus sequence generation steps in genome assembly. To rectify these sequencing errors several polishing tools are already in use but their in-depth survey is still lacking. Appropriate selection of tools can enhance consensus quality and can generate precise assembly. Hence, a comprehensive survey is always beneficial to build a pipeline incorporating multiple evaluation indicators, including contiguity, accuracy, completeness, and contamination along with a proper guidance to select optimal tool and follow the right steps to achieve more accurate and complete genome assembly.

Article
Computer Science and Mathematics
Mathematical and Computational Biology

Azhar Jaan

,

Mohamed Abdella

,

Mohamed Basseem Abdullah Hilal

Abstract: Nonlocal circumstances in genetic engineering are crucial as they pertain to the understanding of genetic material. When these conditions are associated with differential-integral equations, particularly concerning the time variable, they yield comprehensive insights into the material's temporal memory, which can be advantageous for understanding all material properties (including chronic conditions or behavioral characteristics), thereby assisting specialists in managing its future evolution. This study investigates fractional nonlinear mixed integro-differential equations (FrNMIo-DE) with nonlocal circumstances, a category of mathematical problems prevalent in many domains including physics, engineering, and biological systems. Fractional calculus, which generalizes classical differentiation and integration to non-integer orders, offers a robust foundation for modelling memory and hereditary characteristics in complex systems. We examine the existence and uniqueness of solutions to FrNMIo-DE under nonlocal restrictions, using a discontinuous kernel dependent on location and time-space L2[−1,1] × C[0,T], where T < 1, via analytical methods. According to the features of fractional integrals, the FrNMIo-DE adheres to the second-kind Volterra-Hammerstein integral equation (V-HIE), characterized by a discontinuous kernel in position for the Hammerstein integral term and a continuous kernel in time for the Volterra integral (VI) term. Subsequently, we use a separation approach technique to produce HIE with time-dependent physical coefficients. Following an analysis of the system's convergence, a nonlinear algebraic system (NAS) is constructed using the Toeplitz matrix technique (TMT) and related methodologies. The numerical data and associated errors are shown via the Maple 2022 software.

of 15