Submitted:
07 July 2026
Posted:
08 July 2026
You are already at the latest version
Abstract
Keywords:
1. Introduction
2. Related Work
2.1. Patch Selection and Bag Construction in Computational Pathology
2.2. Submodular Optimisation and GP Information Gain
2.3. Determinantal Point Processes in Machine Learning
2.4. PAC-Style Concentration for Stopping Rules
3. Notation and Preliminaries
Gaussian Processes.
Determinantal Point Processes.
Submodularity.
Notation.
4. The InfoDPP-PAC Framework
4.1. Problem Formulation
4.2. Teacher-Student GP Initialisation
Seed phase.
Inference phase.
4.3. Greedy Subset Selection
4.4. PAC Stopping Criterion
4.5. Residual Information Vector
4.6. Algorithm
| Algorithm 1 InfoDPP-PAC |
![]() |
5. Theoretical Analysis
5.1. Log-Determinant as Mutual Information
5.2. Submodularity and Approximation Guarantees
5.3. Redundancy Penalisation
5.4. PAC Stopping Guarantee
5.5. Approximation Fidelity Under Nyström Kernels
6. Experiments
6.1. Experimental Setup
Scope of validation.
Data.
Baselines.
Evaluation metrics.





6.2. Hyperparameter Selection
6.3. Adaptive Stopping in Practice
6.4. Main Comparison
| Method | Family | LogDet ↑ | Cl. Cov. ↑ | Sp. Cov. ↑ | Quality ↑ | Comp. ↑ |
|---|---|---|---|---|---|---|
| Uniform grid | naive | 16.74 | 0.99 | 0.266 | 1.370 | 0.570 |
| Random tissue | naive | 16.52 | 0.97 | 0.266 | 1.142 | 0.551 |
| Entropy | heuristic | 13.74 | 0.57 | 0.163 | 1.372 | 0.438 |
| Edge density | heuristic | 13.66 | 0.57 | 0.148 | 1.353 | 0.436 |
| Colour variance | heuristic | 15.25 | 0.67 | 0.178 | 1.367 | 0.473 |
| Texture energy | heuristic | 13.53 | 0.57 | 0.129 | 1.356 | 0.426 |
| Combined | heuristic | 14.03 | 0.60 | 0.141 | 1.360 | 0.442 |
| k-center greedy | coreset | 20.77 | 0.96 | 0.268 | 1.366 | 0.596 |
| Farthest point | coreset | 20.82 | 0.99 | 0.270 | 1.366 | 0.604 |
| k-means | coreset | 16.68 | 1.00 | 0.263 | 1.356 | 0.573 |
| ABMIL top-k | learned | 13.11 | 0.44 | 0.219 | 1.371 | 0.411 |
| CLAM top-k | learned | 13.11 | 0.44 | 0.203 | 1.368 | 0.405 |
| TransMIL top-k | learned | 13.25 | 0.45 | 0.205 | 1.375 | 0.410 |
| EvoPS | learned | 16.92 | 1.00 | 0.254 | 1.359 | 0.573 |
| InfoDPP-PAC (tuned) | proposed | 12.60 | 0.79 | 0.169 | 1.518 | 0.588 |
6.5. Qualitative Spatial Selection Analysis
6.6. Ablation Study
| Variant | LogDet ↑ | Sp. Cov. ↑ | Redundancy ↓ | Quality ↑ |
|---|---|---|---|---|
| Full InfoDPP-PAC | ||||
| No DPP (quality-only greedy) | ||||
| No GP (uniform pseudo-labels) | ||||
| No PAC (fixed k; identical to Full at this fixed k) | ||||
| Quality: only () | ||||
| Quality: only () | ||||
| Nyström DPP (rank 50) | ||||
| Matérn- kernel |
6.7. Preliminary Cross-Organ Generalisation
| Organ | Method | n | Mean k | Quality | Composite |
|---|---|---|---|---|---|
| Gastrointestinal (TMLR test) | InfoDPP-PAC (adaptive) | 22 | 47.1 | 1.513 | 0.592 |
| Farthest point () | 23 | 50 | 1.366 | 0.604 | |
| k-center greedy () | 23 | 50 | 1.366 | 0.596 | |
| Uniform grid () | 23 | 50 | 1.370 | 0.570 | |
| Breast (pilot) | InfoDPP-PAC (adaptive) | 7 | 8.7 | 1.560 | 0.550 |
| Farthest point () | 7 | 50 | 1.356 | 0.588 | |
| k-center greedy () | 7 | 50 | 1.361 | 0.589 | |
| Uniform grid () | 7 | 50 | 1.411 | 0.551 | |
| Colorectal (pilot) | InfoDPP-PAC (adaptive) | 7 | 7.7 | 1.555 | 0.553 |
| Farthest point () | 7 | 50 | 1.396 | 0.602 | |
| k-center greedy () | 7 | 50 | 1.395 | 0.607 | |
| Uniform grid () | 7 | 50 | 1.396 | 0.576 |
7. Discussion
Theoretical-empirical correspondence.
The role of PAC stopping.
Limitations and scope of claims.
Future directions.
8. Conclusion
Acknowledgments
References
- Gabriele Campanella, Matthew G Hanna, Luke Geneslaw, et al. Clinical-grade computational pathology using weakly supervised deep learning on whole slide images. Nat. Med. 2019, 25(8), 1301–1309. [CrossRef] [PubMed]
- Olivier Catoni. PAC-Bayesian Supervised Classification: The Thermodynamics of Statistical Learning, volume 56 of Lecture Notes–Monograph Series. Institute of Mathematical Statistics, 2007.
- Laming Chen, Guoxin Zhang, and Eric Zhou. Fast greedy MAP inference for determinantal point process to improve recommendation diversity. In Advances in Neural Information Processing Systems, volume 31, pp. 5622–5633, 2018.
- Richard J Chen, Chengkuan Chen, Yicong Li, Tiffany Y Chen, Andrew D Trister, Rahul G Krishnan, and Faisal Mahmood. Scaling vision transformers to gigapixel images via hierarchical self-supervised learning. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 16144–16155, 2022.
- Richard J Chen, Tong Ding, Ming Y Lu, Drew FK Williamson, Guillaume Jaume, Andrew H Song, Bowen Chen, Andrew Zhang, Daniel Shao, Muhammad Shaban, et al. Towards a general-purpose foundation model for computational pathology. Nat. Med. 2024, 30(3), 850–862. [CrossRef] [PubMed]
- Gintare Karolina Dziugaite and Daniel M Roy. Computing nonvacuous generalization bounds for deep (stochastic) neural networks with many parameters. In Conference on Uncertainty in Artificial Intelligence, 2017.
- Alexandre Filiot, Ridouane Ghermi, Antoine Olivier, Paul Jacob, Lucas Fidon, Alice Mac Kain, Charlie Saillard, and Jean-Baptiste Schiratti. Scaling self-supervised learning for histopathology with masked image modeling. medRxiv, 2023.
- Alex Gittens and Michael W Mahoney. Revisiting the Nyström method for improved large-scale machine learning. J. Mach. Learn. Res. 2016, 17, 1–65. [PubMed]
- Daniel Golovin and Andreas Krause. Adaptive submodularity: Theory and applications in active learning and stochastic optimization. J. Artif. Intell. Res. 2011, 42, 427–486.
- Boqing Gong, Wei-Lun Chao, Kristen Grauman, and Fei Sha. Diverse sequential subset selection for supervised video summarization. In Advances in Neural Information Processing Systems, volume 27, pp. 2069–2077, 2014.
- Saya Hashemian and Azam Asilian Bidgoli. EvoPS: Evolutionary patch selection for whole slide image analysis in computational pathology. arXiv 2025, arXiv:2511.07560.
- Roger A Horn and Charles R Johnson. Matrix Analysis. Cambridge University Press, 2nd edition, 2012.
- Maximilian Ilse, Jakub Tomczak, and Max Welling. Attention-based deep multiple instance learning. In International Conference on Machine Learning, pp. 2127–2136. PMLR, 2018.
- Andreas Krause, Ajit Singh, and Carlos Guestrin. Near-optimal sensor placements in Gaussian processes: Theory, efficient algorithms and empirical studies. J. Mach. Learn. Res. 2008, 9, 235–284.
- Alex Kulesza and Ben Taskar. Determinantal point processes for machine learning. Found. Trends Mach. Learn. 2012, 5(2–3), 123–286. [CrossRef]
- Bin Li, Yin Li, and Kevin W Eliceiri. Dual-stream multiple instance learning network for whole slide image classification with self-supervised contrastive learning. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 14318–14328, 2021.
- Ming Y Lu, Drew FK Williamson, Tiffany Y Chen, Richard J Chen, Matteo Barbieri, and Faisal Mahmood. Data-efficient and weakly supervised computational pathology on whole-slide images. Nature Biomedical Engineering 2021, 5, 555–570. [CrossRef] [PubMed]
- Zelda Mariet and Suvrit Sra. Diversity networks: Neural network compression using determinantal point processes. In International Conference on Learning Representations, 2016.
- David A McAllester. PAC-Bayesian model averaging. In Conference on Learning Theory, pp. 164–170, 1999.
- Baharan Mirzasoleiman, Jeff Bilmes, and Jure Leskovec. Coresets for data-efficient training of machine learning models. In International Conference on Machine Learning, pp. 6950–6960. PMLR, 2020.
- Dmitry Nechaev, Alexey Pchelnikov, and Ekaterina Ivanova. HISTAI: An open-source, large-scale whole slide image dataset for computational pathology. arXiv 2025, arXiv:2505.12120.
- George L Nemhauser, Laurence A Wolsey, and Marshall L Fisher. An analysis of approximations for maximizing submodular set functions—I. Math. Program. 1978, 14(1), 265–294. [CrossRef]
- Carl Edward Rasmussen and Christopher K I Williams. Gaussian Processes for Machine Learning. MIT Press, 2006.
- Manahil Raza, Ruqayya Awan, Raja Muhammad Saad Bashir, Talha Qaiser, and Nasir M Rajpoot. Dual attention model with reinforcement learning for classification of histology whole-slide images. arXiv 2023, arXiv:2302.09682.
- Ozan Sener and Silvio Savarese. Active learning for convolutional neural networks: A core-set approach. In International Conference on Learning Representations, 2018.
- Zhuchen Shao, Hao Bian, Yang Chen, Yifeng Wang, Jian Zhang, Xiangyang Ji, and Yongbing Zhang. TransMIL: Transformer based correlated multiple instance learning for whole slide image classification. In Advances in Neural Information Processing Systems, volume 34, pp. 2136–2147, 2021.
- Maxim Sviridenko. A note on maximizing a submodular set function subject to a knapsack constraint. Oper. Res. Lett. 2004, 32(1), 41–43. [CrossRef]
- Hanwen Xu, Naoto Usuyama, Jaspreet Bagga, Sheng Zhang, Rajesh Rao, Tristan Naumann, Cliff Wong, Zelalem Gero, Javier González, Yu Gu, et al. A whole-slide foundation model for digital pathology from real-world data. Nature 2024, 630, 181–188. [CrossRef] [PubMed]
- Junde Xu, Zikai Lin, Donghao Zhou, Yaodong Yang, Xiangyun Liao, Bian Wu, Guangyong Chen, and Pheng-Ann Heng. DPPMask: Masked image modeling with determinantal point processes. arXiv 2023, arXiv:2303.12736.
- Hongrun Zhang, Yanda Meng, Yitian Zhao, Yihong Qiao, Xiaoyun Yang, Sarah E Coupland, and Yalin Zheng. DTFD-MIL: Double-tier feature distillation multiple instance learning for histopathology whole slide image classification. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 18802–18812, 2022.


| Quantity | Adaptive () | Fixed () |
|---|---|---|
| Patches selected | 300 (fixed) | |
| Patch-count reduction | average | |
| Composite score | (mean) | (mean); retained |
| Mean quality | (mean) | (mean); retained |
| Log-det diversity | (mean) | (mean); retained |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).
