Submitted:
12 June 2026
Posted:
12 June 2026
You are already at the latest version
Abstract

Keywords:
1. Introduction
1.1. Background and Motivation
1.2. Related Works
1.3. Research Gap
1.4. Main Contributions
- Framework Proposal: We introduce GINet-DGC, an end-to-end framework tailored with structural inductive biases to resolve the strict sample constraints in biomedical HDLSS tasks.
- Weight Generation Mechanism: We develop a structure-constrained generator that fuses multi-view statistical, topological, and sparse priors to model high-dimensional gene-expression features robustly.
- Adaptive Early Stopping: We implement an Overfitting-aware Index (OFI) alongside an adaptive training strategy to monitor generalization behavior and prevent overfitting during optimization.
- Extensive Empirical Validation: We comprehensively evaluate the proposed framework on eight public gene-expression datasets using repeated stratified cross-validation against 17 baseline models, demonstrating competitive stability.
2. Materials and Methods
2.1. Model Overview
2.2. Structure-Constrained Weight Generator
- 1.
- Latent Semantic Prior: It constructs basic subspace constraints.
- 2.
- Global Distributional Prior: It captures dataset-wide statistical stability.
- 3.
- Local Topological Prior: It preserves manifold geometry among features.
- 4.
- Hierarchical Semantic Prior: It provides multi-scale nonlinear abstractions.
2.2.1. Latent Semantic Prior
2.2.2. Global Distributional Prior
2.2.3. Local Topological Prior
2.2.4. Hierarchical Semantic Prior
2.2.5. Adaptive Prior Integration and Weight Generation
2.3. Dynamic Generalization Control
2.4. Training Process

2.5. Experimental Setup
2.5.1. Experimental Data
2.5.2. Evaluation Metrics
2.5.3. GINet-DGC Settings
3. Results
3.1. Baselines
3.2. Benchmark Comparison
3.3. Ablation Study
3.3.1. Multi-View Structural Priors
3.3.2. Adaptive Prior Integration
3.4. Analysis of Dynamic Generalization Control
3.5. Training Dynamics
4. Discussion
5. Conclusions
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Acknowledgments
Conflicts of Interest
Abbreviations
| AI | Artificial Intelligence |
| BA | Balanced Accuracy |
| DGC | Dynamic Generalization Control |
| DNN | Deep Neural Network |
| HDLSS | High-Dimensional Low-Sample Size |
| NMF | Non-negative Matrix Factorization |
| OFI | Overfitting-aware Index |
| SCWG | Structure-Constrained Weight Generator |
Appendix A. Detailed Dataset Description
| Name | Size | Features |
|---|---|---|
| Allaml | 72 | 7129 |
| CLL | 111 | 11340 |
| Gli | 85 | 22283 |
| Glioma | 50 | 4434 |
| Lung | 203 | 3312 |
| Prostate | 102 | 5966 |
| Smk | 187 | 19993 |
| Toxicity | 171 | 5748 |
Appendix B. Explanation of Ablation Effectiveness
| Initialization | Gli | Smk | Allaml | CLL | Glioma | Prostate | Toxicity | Lung |
|---|---|---|---|---|---|---|---|---|
| Kaiming | 83.45 | 70.78 | 88.06 | 73.32 | 65.50 | 90.47 | 83.51 | 95.20 |
| NMF | 82.43 | 66.14 | 90.15 | 70.50 | 64.67 | 88.67 | 84.71 | 95.20 |
| PCA | 79.53 | 60.66 | 93.78 | 77.91 | 63.83 | 89.27 | 93.97 | 95.95 |
| Xavier | 87.33 | 70.67 | 93.78 | 81.47 | 63.83 | 92.27 | 83.75 | 96.88 |
| GINet-DGC | 90.80 | 72.06 | 98.00 | 84.02 | 79.99 | 91.43 | 96.04 | 98.77 |

References
- Baldi, P.; Sadowski, P.; Whiteson, D. Searching for Exotic Particles in High-Energy Physics with Deep Learning. Nat. Commun. 2014, 5, 4308. [Google Scholar] [CrossRef]
- Keith, J.A.; Vassilev-Galindo, V.; Cheng, B.; Chmiela, S.; Gastegger, M.; Müller, K.R.; Tkatchenko, A. Combining Machine Learning and Computational Chemistry for Predictive Insights into Chemical Systems. Chem. Rev. 2021, 121, 9816–9872. [Google Scholar] [CrossRef] [PubMed]
- Feng, F.; He, X.; Wang, X.; Luo, C.; Liu, Y.; Chua, T.S. Temporal Relational Ranking for Stock Prediction. ACM Trans. Inf. Syst. 2019, 37, 1–30. [Google Scholar] [CrossRef]
- Wang, Z.; Gao, C.; Xiao, C.; Sun, J. MediTab: Scaling Medical Tabular Data Predictors via Data Consolidation, Enrichment, and Refinement. arXiv 2024, 2305.12081. [Google Scholar] [CrossRef]
- Ruan, Y.; Lan, X.; Tan, D.J.; Abdullah, H.R.; Feng, M. P-Transformer: A Prompt-Based Multimodal Transformer Architecture for Medical Tabular Data. arXiv 2025, 2303.17408. [Google Scholar] [CrossRef]
- Golling, T.; Heinrich, L.; Kagan, M.; Klein, S.; Leigh, M.; Osadchy, M.; Raine, J.A. Masked Particle Modeling on Sets: Towards Self-Supervised High Energy Physics Foundation Models. arXiv 2024, 2401.13537. [Google Scholar] [CrossRef]
- Zou, J.; Huss, M.; Abid, A.; Mohammadi, P.; Torkamani, A.; Telenti, A. A Primer on Deep Learning in Genomics. Nat. Genet. 2019, 51, 12–18. [Google Scholar] [CrossRef]
- Shwartz-Ziv, R.; Armon, A. Tabular Data: Deep Learning Is Not All You Need. arXiv 2021, 2106.03253. [Google Scholar] [CrossRef]
- Gorishniy, Y.; Rubachev, I.; Khrulkov, V.; Babenko, A. Revisiting Deep Learning Models for Tabular Data. Adv. Neural Inf. Process. Syst. 2021, 34, 18932–18943. [Google Scholar]
- Chen, T.; Guestrin, C. XGBoost: A Scalable Tree Boosting System. In Proceedings of the Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, 2016; pp. 785–794. [Google Scholar] [CrossRef]
- Ke, G.; Meng, Q.; Finley, T.; Wang, T.; Chen, W.; Ma, W.; Ye, Q.; Liu, T.Y. LightGBM: A Highly Efficient Gradient Boosting Decision Tree. In Proceedings of the Advances in Neural Information Processing Systems, 2017. [Google Scholar]
- Song, W.; Shi, C.; Xiao, Z.; Duan, Z.; Tang, J.; Xu, Y.; Zhang, M. AutoInt: Automatic Feature Interaction Learning via Self-Attentive Neural Networks. In Proceedings of the Proceedings of the 28th ACM International Conference on Information and Knowledge Management, 2019; pp. 1161–1170. [Google Scholar] [CrossRef]
- Arik, S.O.; Pfister, T. TabNet: Attentive Interpretable Tabular Learning. Proc. AAAI Conf. Artif. Intell. 2021, 35, 6679–6687. [Google Scholar] [CrossRef]
- Huang, X.; Khetan, A.; Cvitkovic, M.; Karnin, Z. TabTransformer: Tabular Data Modeling Using Contextual Embeddings. arXiv 2020, 2012.06678. [Google Scholar] [CrossRef]
- Somepalli, G.; Goldblum, M.; Schwarzschild, A.; Bruss, C.B.; Goldstein, T. SAINT: Improved Neural Networks for Tabular Data via Row Attention and Contrastive Pre-Training. arXiv 2021, 2106.01342. [Google Scholar] [CrossRef]
- Singh, D.; Yamada, M. FsNet: Feature Selection Network on High-Dimensional Biological Data; Manuscript, 2020. [Google Scholar]
- Ha, D.; Dai, A.; Le, Q.V. HyperNetworks. In Proceedings of the International Conference on Learning Representations, 2017. [Google Scholar]
- Wydmanski, W.; Bulenok, V. HyperTab: Hypernetwork Approach for Deep Learning on Small Tabular Datasets. arXiv 2023, 2303.00923. [Google Scholar] [CrossRef]
- Romero, A.; Carrier, P.L.; Erraqabi, A.; Sylvain, T.; Auvolat, A.; Dejoie, E.; Legault, M.A.; Dubé, M.P.; Hussin, J.G.; Bengio, Y. Diet Networks: Thin Parameters for Fat Genomics. In Proceedings of the International Conference on Learning Representations, 2017. [Google Scholar]
- Margeloiu, A.; Simidjievski, N.; Liò, P.; Jamnik, M. Weight Predictor Network with Feature Selection for Small Sample Tabular Biomedical Data. arXiv 2022, 2211.15616. [Google Scholar] [CrossRef]
- Margeloiu, A.; Simidjievski, N.; Liò, P.; Jamnik, M. GCondNet: A Novel Method for Improving Neural Networks on Small High-Dimensional Tabular Data. arXiv 2024, 2211.06302. [Google Scholar] [CrossRef]
- Chawla, N.V.; Bowyer, K.W.; Hall, L.O.; Kegelmeyer, W.P. SMOTE: Synthetic Minority Over-Sampling Technique. J. Artif. Intell. Res. 2002, 16, 321–357. [Google Scholar] [CrossRef]
- He, H.; Bai, Y.; Garcia, E.A.; Li, S. ADASYN: Adaptive Synthetic Sampling Approach for Imbalanced Learning. In Proceedings of the IEEE International Joint Conference on Neural Networks, 2008; pp. 1322–1328. [Google Scholar] [CrossRef]
- Xu, L.; Skoularidou, M.; Cuesta-Infante, A.; Veeramachaneni, K. Modeling Tabular Data Using Conditional GAN. In Proceedings of the Advances in Neural Information Processing Systems, 2019. [Google Scholar]
- Kotelnikov, A.; Baranchuk, D.; Rubachev, I.; Babenko, A. TabDDPM: Modelling Tabular Data with Diffusion Models. In Proceedings of the International Conference on Machine Learning, 2023; pp. 17564–17579. [Google Scholar]
- Hollmann, N.; Müller, S.; Purucker, L.; Krishnakumar, A.; Körfer, M.; Hoo, S.B.; Schirrmeister, R.T.; Hutter, F. Accurate Predictions on Small Data with a Tabular Foundation Model. Nature 2025, 637, 319–326. [Google Scholar] [CrossRef]
- Gardner, J.; Perdomo, J.C.; Schmidt, L. Large Scale Transfer Learning for Tabular Data via Language Modeling. arXiv 2024, 2406.12031. [Google Scholar] [CrossRef]
- Liu, B.; Wei, Y.; Zhang, Y.; Yang, Q. Deep Neural Networks for High Dimension, Low Sample Size Data. In Proceedings of the Proceedings of the Twenty-Sixth International Joint Conference on Artificial Intelligence, 2017; pp. 2287–2293. [Google Scholar] [CrossRef]
- Lee, D.D.; Seung, H.S. Learning the Parts of Objects by Non-Negative Matrix Factorization. Nature 1999, 401, 788–791. [Google Scholar] [CrossRef] [PubMed]
- Xie, J.; Girshick, R.; Farhadi, A. Unsupervised Deep Embedding for Clustering Analysis. In Proceedings of the International Conference on Machine Learning, 2016; pp. 478–487. [Google Scholar]
- Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A.N.; Kaiser, L.; Polosukhin, I. Attention Is All You Need. arXiv 2017, 1706.03762. [Google Scholar] [CrossRef]
- Prechelt, L. Early Stopping – But When? In Neural Networks: Tricks of the Trade; Springer, 1998; pp. 55–69. [Google Scholar] [CrossRef]
- LeCun, Y.; Bengio, Y.; Hinton, G. Deep Learning. Nature 2015, 521, 436–444. [Google Scholar] [CrossRef]
- Ioffe, S.; Szegedy, C. Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift. In Proceedings of the International Conference on Machine Learning, 2015; pp. 448–456. [Google Scholar]
- Srivastava, N.; Hinton, G.; Krizhevsky, A.; Sutskever, I.; Salakhutdinov, R. Dropout: A Simple Way to Prevent Neural Networks from Overfitting. J. Mach. Learn. Res. 2014, 15, 1929–1958. [Google Scholar]
- Loshchilov, I.; Hutter, F. Decoupled Weight Decay Regularization. In Proceedings of the International Conference on Learning Representations, 2019. [Google Scholar]
- Breiman, L. Random Forests. Mach. Learn. 2001, 45, 5–32. [Google Scholar] [CrossRef]
- Lemhadri, I.; Ruan, F.; Abraham, L.; Tibshirani, R. LassoNet: A Neural Network with Feature Sparsity. arXiv 2021, 1907.12207. [Google Scholar] [CrossRef]
- Liu, B.; Wei, Y.; Zhang, Y.; Yang, Q. Deep Neural Networks for High Dimension, Low Sample Size Data. In Proceedings of the Proceedings of the Twenty-Sixth International Joint Conference on Artificial Intelligence, 2017; pp. 2287–2293. [Google Scholar] [CrossRef]
- Feng, J.; Simon, N. Sparse-Input Neural Networks for High-Dimensional Nonparametric Regression and Classification. arXiv 2019, 1711.07592. [Google Scholar] [CrossRef]
- Kipf, T.N.; Welling, M. Semi-Supervised Classification with Graph Convolutional Networks. In Proceedings of the International Conference on Learning Representations, 2017. [Google Scholar]
- Brody, S.; Alon, U.; Yahav, E. How Attentive Are Graph Attention Networks? arXiv 2022, 2105.14491. [Google Scholar] [CrossRef]
- Cohen, J.M.; Li, Y.; Zhang, Z. Early Training Dynamics and Oscillations in Neural Networks, 2024. Manuscript.








| Method | gli | smk | allaml | CLL | glioma | prostate | toxicity | lung | Avg-Rank |
|---|---|---|---|---|---|---|---|---|---|
| N/D | 0.004 | 0.009 | 0.01 | 0.01 | 0.011 | 0.017 | 0.03 | 0.059 | |
| DietNetworks | 76.42 | 62.71 | 92.00 | 68.84 | 68.00 | 81.71 | 82.13 | 90.43 | 11.50 |
| FsNet | 74.52 | 56.27 | 78.00 | 66.38 | 53.17 | 84.74 | 60.26 | 91.75 | 14.13 |
| DNP | 83.17 | 66.61 | 96.18 | 85.13 | 75.00 | 88.71 | 93.49 | 92.81 | 6.00 |
| SPINN | 83.39 | 65.91 | 96.78 | 85.35 | 75.00 | 88.66 | 93.50 | 94.20 | 4.25 |
| WPFS | 83.86 | 66.89 | 96.42 | 79.14 | 73.83 | 89.15 | 88.29 | 98.43 | 5.50 |
| TabNet | 64.49 | 61.16 | 69.66 | 50.87 | 45.99 | 65.48 | 41.59 | 70.92 | 15.13 |
| TabTransformer | 78.82 | 64.00 | 88.38 | 76.81 | 63.50 | 85.96 | 87.67 | 94.03 | 10.00 |
| Hypertab | 50.50 | 50.00 | 85.43 | 44.67 | 51.48 | 58.09 | 24.42 | 89.59 | 16.25 |
| CAE | 74.18 | 59.96 | 89.80 | 71.94 | 67.83 | 87.60 | 60.36 | 85.00 | 12.75 |
| LassoNet | 53.91 | 51.04 | 50.80 | 30.63 | 29.17 | 54.78 | 26.67 | 25.11 | 17.50 |
| MLP | 77.72 | 64.62 | 91.30 | 78.30 | 73.00 | 88.76 | 93.21 | 94.20 | 8.00 |
| Random Forest | 83.83 | 67.06 | 96.00 | 79.29 | 75.50 | 91.36 | 80.33 | 91.05 | 7.25 |
| LightGBM | 80.49 | 69.51 | 93.00 | 80.37 | 75.50 | 91.91 | 82.54 | 93.47 | 6.75 |
| GCN | 84.30 | 62.22 | 85.50 | 70.64 | 67.18 | 84.73 | 76.50 | 96.33 | 10.25 |
| GATv2 | 76.01 | 57.35 | 74.17 | 55.52 | 54.80 | 76.63 | 76.65 | 92.63 | 13.75 |
| GCondNet | 86.36 | 68.08 | 97.56 | 80.70 | 77.67 | 90.38 | 95.25 | 96.64 | 3.00 |
| TabPFNTransformer | 75.50 | 71.51 | 95.56 | 75.98 | 73.85 | 93.37 | 90.26 | 95.05 | 5.88 |
| GINet-DGC (ours) | 90.80 | 72.06 | 98.00 | 84.02 | 79.99 | 91.43 | 96.04 | 98.77 | 1.50 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).