Submitted:
15 June 2025
Posted:
16 June 2025
You are already at the latest version
Abstract
Keywords:
1. Introduction
- I investigate the integration of Squeeze-and-Excitation (SE) blocks [11] into the original CNN-based encoder of HoVerNet, aiming to enhance channel-wise feature recalibration. This modification leads to measurable improvements in segmentation accuracy.
- I explore the replacement of HoVerNet’s residual CNN encoder with Transformer-based modules, including Vision Transformer (ViT) and Swin Transformer. A unified architecture is proposed to preserve compatibility with HoVerNet’s three-branch output structure.
- I design an enhanced decoder architecture that leverages dense connections and dropout regularization, facilitating improved information flow and mitigating overfitting risks during training.
- I conduct comparative experiments against the current state-of-the-art model, CellViT, and perform ablation studies to assess the individual contributions of each architectural component.
- All experiments are conducted on the PanNuke dataset, following consistent training protocols—such as fixed epoch count, learning rate, pretrained models, optimizer, and train-validation splits—to ensure fair evaluation. While not all proposed models surpass the original HoVerNet in every metric, the findings emphasize the value of attention mechanisms and architectural enhancements in advancing segmentation and classification performance.
2. Materials and Methods
2.1. Dataset
2.2. Preprocessing and Augmentation
2.3. Model Variants Based on HoVerNet Architecture
- HoVerNet + SE (HoverNetEnhanced): This variant augments the encoder with Squeeze-and-Excitation (SE) blocks, which are integrated within the residual units. SE blocks perform adaptive channel-wise recalibration, emphasizing informative features while suppressing less useful ones, thereby enhancing representational capacity [11].
- HoVerNet + Multi-head Attention (Multihead-HoverNet): In this model, multi-head self-attention (MHSA) modules [13] are embedded into the encoder to capture long-range dependencies and global contextual cues. Inspired by prior analysis of attention heads [14], this design seeks to improve performance on complex spatial structures.
- HoVerNet + SE + MHSA + Enhanced Decoder (MSDHV-Net): This is the most comprehensive CNN-based modification. It combines both SE and MHSA modules in the encoder and introduces a redesigned decoder featuring deeper DenseBlock structures, additional skip connections, and convolutional refinement layers. This design aims to facilitate robust feature propagation and enhanced spatial resolution restoration [15].
- HoVerNet + ViT Encoder (HoverViTNet): Here, the conventional CNN encoder is replaced with a Vision Transformer (ViT) [7], enabling the model to extract patch-wise global representations using self-attention. These transformer-derived features are decoded via a CNN-based decoder, allowing comparative evaluation of attention-driven global context modeling [16].
- HoVerNet + Custom SwinViT Encoder (HoVerIT): This architecture integrates a custom Swin Transformer encoder into the HoVerNet pipeline. The hierarchical design and shifted window self-attention in SwinViT capture both local and global dependencies more effectively than vanilla ViT [17]. The idea draws inspiration from the Swin-UNETR architecture [18], while maintaining compatibility with HoVerNet’s three-branch outputs.
- HoVerNet + SwinUNETR from MONAI (HoverSwinNet): In this variant, the encoder is directly replaced with the SwinUNETR backbone from MONAI [18]. Transformer-derived multi-scale features are routed through HoVerNet’s original three-branch decoders, serving as a strong baseline to assess the integration feasibility and performance of prebuilt transformer encoders.
2.4. Training Configuration
2.5. Evaluation Metrics
3. Results
3.1. Baseline Performance
3.2. CNN-Based Architectural Enhancements
3.3. Transformer-Based Architectural Variants
3.4. Summary of Comparative Performance
3.5. Formatting of Mathematical Components
4. Discussion


5. Conclusions
Author Contributions
Funding
Data Availability Statement
Acknowledgments
Conflicts of Interest
Abbreviations
| MDPI | Multidisciplinary Digital Publishing Institute |
| SE | Squeeze-and-Excitation |
| CNN | Convolutional Neural Network |
| ViT | Vision Transformer |
Appendix A. Additional Materials and Resources
- Trained model weights for all variants (e.g., HoverNetEnhenced, MSDHV-Net, HoverViTNet, HoVerIT) can be accessed at: https://drive.google.com/drive/folders/1fh0fiiGwIpPOafaoSF52WDrP4vAIxKY8?usp=drive_link.
- More overlay images of segmentation and classification results for representative samples are available at: https://drive.google.com/drive/folders/1urDlgA4QnI_25vAllJV2ouKhwg0Xdp5X?usp=sharing.
Appendix B. Model Structure and Implementation Code
- HoverSwinNet https://github.com/davidqu921/HoverSwinNet.
- HoverViTNet https://github.com/davidqu921/HoverViTNet.
Appendix C. TensorBoard Training Logs and Visualizations

References
- Sirinukunwattana, K.; Snead, D.; Epstein, D.; Aftab, Z.; Mujeeb, I.; Tsang, Y.W.; Cree, I.; Rajpoot, N. Novel digital signatures of tissue phenotypes for predicting distant metastasis in colorectal cancer. Scientific reports 2018, 8, 13692. [Google Scholar] [CrossRef] [PubMed]
- Javed, S.; Fraz, M.M.; Epstein, D.; Snead, D.; Rajpoot, N.M. Cellular community detection for tissue phenotyping in histology images. In Computational Pathology and Ophthalmic Medical Image Analysis; Springer: Cham, Switzerland, 2018; pp. 120–129. [Google Scholar] [CrossRef]
- Ronneberger, O.; Fischer, P.; Brox, T. U-Net: Convolutional Networks for Biomedical Image Segmentation. In Lecture Notes in Computer Science; Springer: Cham, Switzerland, 2015; pp. 234–241. [Google Scholar] [CrossRef]
- Isensee, F.; Jaeger, P.F.; Kohl, S.A.A.; Petersen, J.; Maier-Hein, K.H. nnU-Net: A Self-Configuring Method for Deep Learning-Based Biomedical Image Segmentation. Nat. Methods 2021, 18, 203–211. [Google Scholar] [CrossRef] [PubMed]
- Graham, S.; Vu, Q.D.; Raza, S.E.; Azam, A.; Tsang, Y.W.; Kwak, J.T.; Rajpoot, N. Hover-net: Simultaneous segmentation and classification of nuclei in multi-tissue histology images. Medical Image Analysis 2019, 58, 101563. [Google Scholar] [CrossRef] [PubMed]
- Gamper, J.; Koohbanani, N.A.; Graham, S.; Jahanifar, M.; Khurram, S.A.; Azam, A.; Hewitt, K.; Rajpoot, N. PanNuke Dataset Extension, Insights and Baselines. arXiv 2020, arXiv:2003.10778. https://arxiv.org/abs/2003.10778. [Google Scholar]
- Dosovitskiy, A.; Beyer, L.; Kolesnikov, A.; Weissenborn, D.; Zhai, X.; Unterthiner, T. An Image Is Worth 16×16 Words: Transformers for Image Recognition at Scale. arXiv 2020, arXiv:2010.11929. https://arxiv.org/abs/2010.11929. [Google Scholar]
- Liu, Z.; Lin, Y.; Cao, Y.; Hu, H.; Wei, Y.; Zhang, Z.; ...; Guo, B. Swin Transformer: Hierarchical Vision Transformer Using Shifted Windows. In Proc. IEEE/CVF Int. Conf. Comput. Vis.; 2021; pp. 10012–10022. [CrossRef]
- Hörst, F.; Rempe, M.; Heine, L.; Seibold, C.; Keyl, J.; Baldini, G.; Kleesiek, J. CellViT: Vision Transformers for Precise Cell Segmentation and Classification. Medical Image Analysis 2024, 94, 103143. [Google Scholar] [CrossRef]
- H. Xu et al., "Vision Transformers for Computational Histopathology," IEEE Reviews in Biomedical Engineering, vol. 17, pp. 63–79, 2024. [CrossRef]
- J. Hu, L. Shen, and G. Sun, "Squeeze-and-excitation networks," in Proc. IEEE Conf. Comput. Vis. Pattern Recognit. (CVPR), 2018, pp. 7132–7141. [CrossRef]
- Zhang, Z. Enhancing Distributed Machine Learning through Data Shuffling: Techniques, Challenges, and Implications. In ITM Web of Conferences 2025, 73, 03018. [Google Scholar] [CrossRef]
- Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A.N.; Kaiser, Ł.; Polosukhin, I. Attention is All You Need. Advances in Neural Information Processing Systems 2017, 30. [Google Scholar] [CrossRef]
- Voita, E.; Talbot, D.; Moiseev, F.; Sennrich, R.; Titov, I. Analyzing multi-head self-attention: Specialized heads do the heavy lifting, the rest can be pruned. arXiv 2020, arXiv:1905.09418. https://aclanthology.org/P19-1580/. [Google Scholar]
- Mao, X.; Shen, C.; Yang, Y.B. Image Restoration Using Very Deep Convolutional Encoder-Decoder Networks with Symmetric Skip Connections. Advances in Neural Information Processing Systems 2016, 29, 2802–2810. [Google Scholar] [CrossRef]
- Wang, Z.; Li, T.; Zheng, J.Q.; Huang, B. When CNN Meet with ViT: Towards Semi-Supervised Learning for Multi-Class Medical Image Semantic Segmentation. In Proceedings of the European Conference on Computer Vision (ECCV), Cham, Switzerland, October 2022; Springer Nature Switzerland: Cham, 2022; pp. 424–441. [Google Scholar] [CrossRef]
- Yang, J.; Li, C.; Zhang, P.; Dai, X.; Xiao, B.; Yuan, L.; Gao, J. Focal Self-Attention for Local-Global Interactions in Vision Transformers. arXiv preprint 2021, arXiv:2107.00641. [Google Scholar] [CrossRef]
- Hatamizadeh, A.; Nath, V.; Tang, Y.; Yang, D.; Roth, H.R.; Xu, D. Swin UNETR: Swin Transformers for Semantic Segmentation of Brain Tumors in MRI Images. In International MICCAI Brainlesion Workshop; Springer: Cham, Switzerland, 2021; pp. 272–284. [Google Scholar] [CrossRef]
- D. P. Kingma and J. Ba, "Adam: A method for stochastic optimization," arXiv preprint arXiv:1412.6980, 2014. https://doi.org/10.48550/arXiv.1412.6980. [CrossRef]
- K. He, X. Zhang, S. Ren, and J. Sun, "Deep residual learning for image recognition," in Proc. IEEE Conf. Comput. Vis. Pattern Recognit. (CVPR), 2016, pp. 770–778. https://www.cv-foundation.org/openaccess/content_cvpr_2016/papers/He_Deep_Residual_Learning_CVPR_2016_paper.pdf.
- I. Loshchilov and F. Hutter, "Decoupled weight decay regularization," arXiv preprint arXiv:1711.05101, 2017. https://www.semanticscholar.org/paper/Decoupled-Weight-Decay-Regularization-Loshchilov-Hutter/d07284a6811f1b2745d91bdb06b040b57f226882.
- A. Kirillov, K. He, R. Girshick, C. Rother, and P. Dollár, "Panoptic segmentation," in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2019, pp. 9404–9413. https://openaccess.thecvf.com/content_CVPR_2019/papers/Kirillov_Panoptic_Segmentation_CVPR_2019_paper.pdf.



| Model | NP-Dice | TP-Dice-0 | TP-Dice-1 | TP-Dice-2 | TP-Dice-3 |
|---|---|---|---|---|---|
| HoverNet | 0.8686 | 0.9703 | 0.7978 | 0.6728 | 0.6939 |
| HoverNetEnhanced | 0.8696 | 0.9697 | 0.8119 | 0.7175 | 0.7129 |
| Multihead-HoverNet | 0.8682 | 0.9698 | 0.7933 | 0.6643 | 0.6904 |
| MSDHV-Net | 0.8696 | 0.9697 | 0.8105 | 0.7162 | 0.7241 |
| HoverSwinNet | 0.8365 | 0.9643 | 0.7381 | 0.6221 | 0.6380 |
| HoverViTNet | 0.8564 | 0.9662 | 0.7672 | 0.6783 | 0.6816 |
| HoverIT | 0.8534 | 0.9676 | 0.7433 | 0.6297 | 0.6455 |
| CellViT | 0.8000 | 0.9729 | 0.9764 | 0.9632 | 0.9331 |
| Model | PQ-1 | Recall-1 | Precision-1 | PQ-2 | Recall-2 | Precision-2 |
|---|---|---|---|---|---|---|
| HoverNet | 0.4005 | 0.4192 | 0.4004 | 0.2830 | 0.2935 | 0.3206 |
| HoverNetEnhanced | 0.3843 | 0.3971 | 0.3924 | 0.2820 | 0.3119 | 0.2952 |
| Multihead-HoverNet | 0.3947 | 0.4070 | 0.4018 | 0.2813 | 0.2972 | 0.3150 |
| MSDHV-Net | 0.3825 | 0.3986 | 0.3872 | 0.2891 | 0.3179 | 0.3046 |
| HoverSwinNet | 0.3744 | 0.3824 | 0.3913 | 0.2400 | 0.2342 | 0.2945 |
| HoverViTNet | 0.3788 | 0.3984 | 0.3984 | 0.2793 | 0.3108 | 0.2976 |
| HoverIT | 0.3744 | 0.3984 | 0.3823 | 0.2400 | 0.3108 | 0.2976 |
| CellViT | 0.5606 | 0.6900 | 0.7200 | 0.4316 | 0.5700 | 0.5900 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2025 by the author. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).