Submitted:
30 June 2026
Posted:
01 July 2026
You are already at the latest version
Abstract
Keywords:
1. Introduction
- We propose a prompt-preserving MedSAM enhancement framework that keeps the original point- and box-prompt interface unchanged and primarily targets image-encoder representations.
- We design HBF to reuse historical bottleneck features from previous enhanced blocks, improving the preservation of boundary and small-target cues during deep token propagation.
- We design BCER as a soft pixel-wise convolutional expert-weighting module, enabling adaptive local enhancement for heterogeneous medical image regions while avoiding claims of sparse top-k MoE routing.
- We provide matched prompt-based comparisons and ablation analyses on fundus, dermoscopic, and thyroid ultrasound datasets, demonstrating consistent improvements over the matched MedSAM baseline under identical simulated-prompt settings.
2. Related Work
2.1. Promptable Medical Image Segmentation
2.2. Structural Preservation in Medical Segmentation
2.3. Expert Routing and Adaptive Local Enhancement
3. Materials and Methods
3.1. Overall Framework
3.2. Lightweight Decoder Adapter and Adaptation Scope
3.3. Historical Branch Fusion
3.4. Balanced Convolutional Expert Routing
3.5. Optimization Objective
4. Experiments and Results
4.1. Datasets and Implementation Details
| Setting | Image Encoder | Prompt Encoder | Mask Decoder | Decoder Adapter | HBF | BCER | Main Purpose |
|---|---|---|---|---|---|---|---|
| Original MedSAM | Pretrained | Pretrained | Original | No | No | No | Reference model definition |
| Fully fine-tuned MedSAM | Trainable | Task dependent | Trainable | No | No | No | Trainable-parameter efficiency reference |
| Frozen MedSAM + Adapter | Frozen | Frozen | Adapter-enhanced | Yes | No | No | Adapter-control concept |
| HBF-only variant | Frozen | Frozen | Adapter-enhanced | Yes | Yes | No | Ablation control |
| BCER-only variant | Frozen | Frozen | Adapter-enhanced | Yes | No | Yes | Ablation control |
| HBF-BCER | Frozen | Frozen | Adapter-enhanced | Yes | Yes | Yes | Proposed framework |
4.2. Prompt Protocol and Metrics
4.3. Matched Prompt-Based Results
4.4. Literature-Reported SAM-Style Reference Results
4.5. Non-Prompted Structural References
4.6. Ablation Study
4.7. Qualitative, Convergence, and Efficiency Analysis
5. Discussion
6. Conclusions
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Acknowledgments
Conflicts of Interest
Abbreviations
| SAM | Segment Anything Model |
| MedSAM | Medical Segment Anything Model |
| HBF | Historical Branch Fusion |
| BCER | Balanced Convolutional Expert Routing |
| PEFT | Parameter-efficient fine-tuning |
| IoU | Intersection over Union |
| HD95 | 95th percentile Hausdorff distance |
References
- Kirillov, A.; Mintun, E.; Ravi, N.; Mao, H.; Rolland, C.; Gustafson, L.; Xiao, T.; Whitehead, S.; Berg, A.C.; Lo, W.Y.; et al. Segment Anything. In Proceedings of the Proceedings of the IEEE/CVF International Conference on Computer Vision, 2023; pp. 4015–4026. [Google Scholar]
- Ma, J.; He, Y.; Li, F.; Han, L.; You, C.; Wang, B. Segment Anything in Medical Images. Nat. Commun. 2024, 15, 654. [Google Scholar] [CrossRef] [PubMed]
- Orlando, J.I.; Fu, H.; Breda, J.B.; van Keer, K.; Bathula, D.R.; Diaz-Pinto, A.; Fang, R.; Heng, P.A.; Kim, J.; Lee, J.; et al. REFUGE Challenge: A Unified Framework for Evaluating Automated Methods for Glaucoma Assessment from Fundus Photographs. Med. Image Anal. 2020, 59, 101570. [Google Scholar] [CrossRef] [PubMed]
- Bazi, Y.; Rahhal, M.M.A.; Elgibreen, H.; Zuair, M. Vision Transformers for Segmentation of Disc and Cup in Retinal Fundus Images. Biomed. Signal Process. Control 2024, 91, 105915. [Google Scholar] [CrossRef]
- Cassidy, B.; Kendrick, C.; Brodzicki, A.; Jaworek-Korjakowska, J.; Yap, M.H. Analysis of the ISIC Image Datasets: Usage, Benchmarks and Recommendations. Med. Image Anal. 2022, 75, 102305. [Google Scholar] [CrossRef] [PubMed]
- Mirikharaji, Z.; Abhishek, K.; Bissoto, A.; Barata, C.; Avila, S.; Valle, E.; Celebi, M.E.; Hamarneh, G. A Survey on Deep Learning for Skin Lesion Segmentation. Med. Image Anal. 2023, 88, 102863. [Google Scholar] [CrossRef] [PubMed]
- Chen, J.; You, H.; Li, K. A Review of Thyroid Gland Segmentation and Thyroid Nodule Segmentation Methods for Medical Ultrasound Images. Comput. Methods Programs Biomed. 2020, 185, 105329. [Google Scholar] [CrossRef] [PubMed]
- Kang, Q.; Lao, Q.; Li, Y.; Jiang, Z.; Qiu, Y.; Zhang, S.; Li, K. Thyroid Nodule Segmentation and Classification in Ultrasound Images Through Intra- and Inter-Task Consistent Learning. Med. Image Anal. 2022, 79, 102443. [Google Scholar] [CrossRef] [PubMed]
- Gong, H.; Chen, J.; Chen, G.; Li, H.; Li, G.; Chen, F. Thyroid Region Prior Guided Attention for Ultrasound Segmentation of Thyroid Nodules. Comput. Biol. Med. 2023, 155, 106389. [Google Scholar] [CrossRef] [PubMed]
- Cheng, J.; Ye, J.; Deng, Z.; Chen, J.; Li, T.; Wang, H.; Su, Y.; Huang, Z.; Chen, J.; Jiang, L.; et al. SAM-Med2D. arXiv 2023, arXiv:2308.16184. [Google Scholar]
- Wu, J.; Ji, W.; Liu, Y.; Fu, H.; Xu, M.; Xu, Y.; Jin, Y. Medical SAM Adapter: Adapting Segment Anything Model for Medical Image Segmentation. arXiv 2023, arXiv:2304.12620. [Google Scholar]
- Mazurowski, M.A.; Dong, H.; Gu, H.; Yang, J.; Konz, N.; Zhang, Y. Segment Anything Model for Medical Image Analysis: An Experimental Study. Med. Image Anal. 2023, 89, 102918. [Google Scholar] [CrossRef] [PubMed]
- Huang, Y.; Yang, X.; Liu, L.; Zhou, H.; Chang, A.; Zhou, X.; Chen, R.; Yu, J.; Chen, J.; Chen, C. Segment Anything Model for Medical Images? Med. Image Anal. 2024, 92, 103061. [Google Scholar] [CrossRef] [PubMed]
- Xiong, X.; Wu, Z.; Tan, S.; Li, W.; Tang, F.; Chen, Y.; Li, S.; Ma, J.; Li, G. SAM2-UNet: Segment Anything 2 Makes Strong Encoder for Natural and Medical Image Segmentation. Vis. Intell. 2026. [Google Scholar] [CrossRef]
- Zhu, J.; Hamdi, A.; Qi, Y.; Jin, Y.; Wu, J. Medical SAM 2: Segment Medical Images as Video via Segment Anything Model 2. arXiv 2024, arXiv:2408.00874. [Google Scholar]
- Ma, J.; Yang, Z.; Kim, S.; Chen, B.; Baharoon, M.; Fallahpour, A.; Asakereh, R.; Lyu, H.; Wang, B. MedSAM2: Segment Anything in 3D Medical Images and Videos. arXiv 2025, arXiv:2504.03600. [Google Scholar]
- Zhou, L.; Hu, J.; Zhang, S.; Du, X.; Song, M.; Zhang, X.; Feng, Z. DenseSAM: Semantic Enhance SAM for Efficient Dense Object Segmentation. Proceedings of the Proceedings of the Thirty-Fourth International Joint Conference on Artificial Intelligence, IJCAI-25. International Joint Conferences on Artificial Intelligence Organization 2025, 7994–8002. [Google Scholar] [CrossRef] [PubMed]
- Ronneberger, O.; Fischer, P.; Brox, T. U-Net: Convolutional Networks for Biomedical Image Segmentation. In Proceedings of the Medical Image Computing and Computer-Assisted Intervention – MICCAI 2015; Springer, 2015; pp. 234–241. [Google Scholar] [CrossRef]
- Zhou, Z.; Siddiquee, M.M.R.; Tajbakhsh, N.; Liang, J. UNet++: Redesigning Skip Connections to Exploit Multiscale Features in Image Segmentation. IEEE Trans. Med. Imaging 2020, 39, 1856–1867. [Google Scholar] [CrossRef] [PubMed]
- Isensee, F.; Jaeger, P.F.; Kohl, S.A.A.; Petersen, J.; Maier-Hein, K.H. nnU-Net: A Self-Configuring Method for Deep Learning-Based Biomedical Image Segmentation. Nat. Methods 2021, 18, 203–211. [Google Scholar] [CrossRef] [PubMed]
- Chen, J.; Mei, J.; Li, X.; Lu, Y.; Yu, Q.; Wei, Q.; Luo, X.; Xie, Y.; Adeli, E.; Wang, Y.; et al. TransUNet: Rethinking the U-Net Architecture Design for Medical Image Segmentation Through the Lens of Transformers. Med. Image Anal. 2024, 97, 103280. [Google Scholar] [CrossRef] [PubMed]
- Zou, Y.; Ge, Y.; Zhao, L.; Li, W. MR-Trans: MultiResolution Transformer for Medical Image Segmentation. Comput. Biol. Med. 2023, 165, 107456. [Google Scholar] [CrossRef] [PubMed]
- Liang, Z.; Zhao, K.; Liang, G.; Li, S.; Wu, Y.; Zhou, Y. MAXFormer: Enhanced Transformer for Medical Image Segmentation with Multi-Attention and Multi-Scale Features Fusion. Knowl.-Based Syst. 2023, 280, 110987. [Google Scholar] [CrossRef]
- Wu, H.; Min, W.; Gai, D.; Huang, Z.; Geng, Y.; Wang, Q.; Chen, R. HD-Former: A Hierarchical Dependency Transformer for Medical Image Segmentation. Comput. Biol. Med. 2024, 178, 108671. [Google Scholar] [CrossRef] [PubMed]
- Yang, R.; Liu, K.; Xu, S.; Yin, J.; Zhang, Z. ViT-UperNet: A Hybrid Vision Transformer with Unified-Perceptual-Parsing Network for Medical Image Segmentation. Complex Intell. Syst. 2024, 10, 3819–3831. [Google Scholar] [CrossRef]
- Ji, Z.; Chen, Z.; Ma, X. Grouped Multi-Scale Vision Transformer for Medical Image Segmentation. Sci. Rep. 2025, 15, 11122. [Google Scholar] [CrossRef] [PubMed]
- Fedus, W.; Zoph, B.; Shazeer, N. Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity. J. Mach. Learn. Res. 2022, 23, 1–39. [Google Scholar]
- Li, R.; Wu, L.; Gu, J.; Xu, Q.; Chen, W.; Cai, X.; Bu, J. MoE-SAM: Enhancing SAM for Medical Image Segmentation with Mixture-of-Experts. In Proceedings of the Medical Image Computing and Computer Assisted Intervention – MICCAI 2025; Springer, 2025; pp. 367–377. [Google Scholar] [CrossRef]
- Wei, J.; Zhao, X.; Woo, J.; Ouyang, J.; Fakhri, G.E.; Chen, Q.; Liu, X. Mixture-of-Shape-Experts (MoSE): End-to-End Shape Dictionary Framework to Prompt SAM for Generalizable Medical Segmentation. arXiv 2025, arXiv:2504.09601. [Google Scholar]
- Liu, J.; Desrosiers, C.; Zhou, Y. Att-MoE: Attention-Based Mixture of Experts for Nuclear and Cytoplasmic Segmentation. Neurocomputing 2020, 411, 139–148. [Google Scholar] [CrossRef]
- Zhang, Z.; Li, Y.; Shin, B.S. Learning Generalizable Visual Representation via Adaptive Spectral Random Convolution for Medical Image Segmentation. Comput. Biol. Med. 2023, 167, 107580. [Google Scholar] [CrossRef] [PubMed]
- Gao, J.; Zhou, S.; Yu, H.; Li, C.; Hu, X. SCESS-Net: Semantic Consistency Enhancement and Segment Selection Network for Audio–Visual Event Localization. Comput. Vis. Image Underst. 2025, 262, 104551. [Google Scholar] [CrossRef]
- Sinkhorn, R.; Knopp, P. Concerning Nonnegative Matrices and Doubly Stochastic Matrices. Pac. J. Math. 1967, 21, 343–348. [Google Scholar] [CrossRef]
- Kervadec, H.; Bouchtiba, J.; Desrosiers, C.; Granger, E.; Dolz, J.; Ayed, I.B. Boundary Loss for Highly Unbalanced Segmentation. Med. Image Anal. 2021, 67, 101851. [Google Scholar] [CrossRef] [PubMed]
- Ma, J.; Chen, J.; Ng, M.; Huang, R.; Li, Y.; Li, C.; Yang, X.; Martel, A.L. Loss Odyssey in Medical Image Segmentation. Med. Image Anal. 2021, 71, 102035. [Google Scholar] [CrossRef] [PubMed]
- Dice, L.R. Measures of the Amount of Ecologic Association Between Species. Ecology 1945, 26, 297–302. [Google Scholar] [CrossRef]
- Jaccard, P. The Distribution of the Flora in the Alpine Zone. New Phytol. 1912, 11, 37–50. [Google Scholar] [CrossRef]
- Huttenlocher, D.P.; Klanderman, G.A.; Rucklidge, W.J. Comparing Images Using the Hausdorff Distance. IEEE Trans. Pattern Anal. Mach. Intell. 1993, 15, 850–863. [Google Scholar] [CrossRef]
- Taha, A.A.; Hanbury, A. Metrics for Evaluating 3D Medical Image Segmentation: Analysis, Selection, and Tool. BMC Med. Imaging 2015, 15, 29. [Google Scholar] [CrossRef] [PubMed]
- Zhang, K.; Liu, X.; Shen, T.; Wu, J.; Dong, Y.; Zhang, R.; Huang, Z.; Wang, Q.; Zheng, Y.; Tian, J.; et al. Customize Segment Anything Model for Medical Image Segmentation. arXiv 2023, arXiv:2304.13785. [Google Scholar]
- Deng, G.; Zou, K.; Ren, K.; Wang, M.; Yang, X.; Xu, P.; Fu, H. SAM-U: Multi-Box Prompts Triggered Uncertainty Estimation for Reliable SAM in Medical Image. arXiv 2023, arXiv:2307.04973. [Google Scholar]
- Wu, H.; Chen, S.; Chen, G.; Wang, W.; Lei, B.; Wen, Z. FAT-Net: Feature Adaptive Transformers for Automated Skin Lesion Segmentation. Med. Image Anal. 2022, 76, 102327. [Google Scholar] [CrossRef] [PubMed]
- Alom, M.Z.; Yakopcic, C.; Hasan, M.; Taha, T.M.; Asari, V.K. Recurrent Residual U-Net for Medical Image Segmentation. J. Med. Imaging 2019, 6, 014006. [Google Scholar] [CrossRef] [PubMed]
- Hatamizadeh, A.; Nath, V.; Tang, Y.; Yang, D.; Roth, H.R.; Xu, D. Swin UNETR: Swin Transformers for Semantic Segmentation of Brain Tumors in MRI Images. In Proceedings of the Brainlesion: Glioma, Multiple Sclerosis, Stroke and Traumatic Brain Injuries; Springer, 2022; pp. 272–284. [Google Scholar] [CrossRef]
- Wang, W.; Chen, C.; Ding, M.; Yu, H.; Zha, S.; Li, J. TransBTS: Multimodal Brain Tumor Segmentation Using Transformer. In Proceedings of the Medical Image Computing and Computer Assisted Intervention – MICCAI 2021; Springer, 2021; pp. 109–119. [Google Scholar] [CrossRef] [PubMed]








| Method | REFUGE2 | ISIC2016 | TNMIX | |||||
| Disc Dice | Disc IoU | Cup Dice | Cup IoU | Dice | IoU | Dice | IoU | |
| MedSAM [2] | 93.8 | 86.2 | 82.1 | 73.8 | 87.5 | 78.6 | 81.6 | 75.1 |
| HBF-BCER | 95.4 | 91.3 | 91.0 | 83.7 | 95.9 | 92.4 | 91.4 | 85.1 |
| Method | REFUGE2 Disc | REFUGE2 Cup | TNMIX | ISIC† | ||||
| Dice | IoU | Dice | IoU | Dice | IoU | Dice | IoU | |
| SAMed [40] | 89.9 | 81.8 | 80.7 | 70.8 | 78.9 | 71.2 | 87.4 | 78.9 |
| SAM-Med2D [10] | 92.1 | 83.7 | 82.0 | 75.3 | 80.3 | 73.6 | 87.8 | 78.3 |
| SAM-U [41] | 91.2 | 82.4 | 81.5 | 73.2 | 79.8 | 74.0 | 88.7 | 79.6 |
| VMN | 92.5 | 83.9 | 82.8 | 76.1 | 81.4 | 74.2 | 88.3 | 79.1 |
| FCFI | 95.5 | 88.3 | 85.7 | 77.7 | 84.3 | 75.7 | 89.6 | 81.5 |
| MedSAM [2] | 92.9 | 85.5 | 82.1 | 73.8 | 81.3 | 74.7 | 86.8 | 77.5 |
| Medical SAM Adapter (Med-SA) [11] | 97.1 | 89.2 | 86.2 | 78.5 | 85.4 | 77.5 | 91.8 | 83.0 |
| Medical SAM 2 [15] | 97.8 | 88.8 | 87.6 | 81.2 | 87.6 | 80.1 | 92.7 | 84.6 |
| Method | REFUGE2 | ISIC2016 | TNMIX | |||||
| Disc Dice | Disc IoU | Cup Dice | Cup IoU | Dice | IoU | Dice | IoU | |
| FAT-Net [42] | 91.8 | 84.8 | 80.9 | 71.5 | 90.7 | 83.9 | 80.8 | 73.4 |
| ResUNet-style [43] | 92.9 | 85.5 | 80.1 | 72.3 | 87.3 | 78.2 | 78.3 | 70.7 |
| 2D Swin-UNETR-style [44] | 95.3 | 87.9 | 84.3 | 74.5 | 90.4 | 83.3 | 83.5 | 74.8 |
| 2D TransBTS-style [45] | 94.1 | 87.2 | 85.4 | 75.7 | 89.7 | 81.2 | 83.5 | 75.1 |
| nnU-Net [20] | 94.7 | 87.3 | 84.9 | 75.1 | 94.6 | 83.6 | 84.2 | 76.2 |
| REFUGE2 | ISIC2016 | TNMIX | ||||||||
| Method | Disc Dice | Disc IoU | HD95 | Cup Dice | Cup IoU | Dice | IoU | HD95 | Dice | IoU |
| Without HBF and BCER | 93.8 | 86.2 | 35.9 | 82.1 | 73.8 | 87.5 | 78.6 | 31.3 | 81.6 | 75.1 |
| Without HBF | 94.6 | 89.2 | 23.5 | 86.0 | 76.1 | 92.4 | 87.4 | 41.6 | 86.4 | 78.0 |
| Without BCER | 95.0 | 90.9 | 13.9 | 85.6 | 76.9 | 92.9 | 87.9 | 42.5 | 87.1 | 78.5 |
| HBF-BCER | 95.4 | 91.3 | 6.5 | 91.0 | 83.7 | 95.9 | 92.4 | 28.0 | 91.4 | 85.1 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).