Submitted:
13 December 2024
Posted:
13 December 2024
You are already at the latest version
Abstract
Keywords:
1. Introduction
2. Materials and Methods
2.1. Materials
2.2. Methods
2.2.1. Experimental Data Preparation
2.2.2. Detection of Parotid Glands
2.2.3. Classification Methods and Training Procedure

2.2.4. Lesion Segmentation
3. Experimantal Results and Discussion
3.1. Detection Results of Parotid Glands
3.2. Results of Disease Classification
3.3. Results of Lesion Segmentation
4. Conclusion
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Conflicts of Interest
References
- Lyer, J., Hariharan, A., Cao, U.M.H., Wang, C.T.T., , Khayambashi M. P., Nuguyen L., and Tran S. D. (2021). An overview on the histogenesis and morphogenesis of parotid gland neoplasms and evolving diagnostic approaches. Malignants, 13(15), 3910.
- ARAUJO, A.L.D., DASILVA V.M., ROLDÁN D.G., LOPES M.A., MORAES M.C., KOWALSKI L.P., SANTOS-SILVA A.G. (2024). Deep Learning for salivary gland tumor classification. Oral Surgery, Oral Medicine, Oral Pathology and Oral Radiology. 137(6), 297.
- Speight P.M., & A William Barrett. Salivary gland tumours: diagnostic challenges and an update on the latest WHO classification. Diagnostic Histopathology. 26(4), 147-158.
- Zhang H.B., Lai H.C.,. Wang Y, et al. (2021). Research on the Classification of Benign and Malignant Parotid Tumors Based on Transfer Learning and a Convolutional Neural Network. IEEE Access, 9, 40360-40371.
- Yuan J., Fan Y., Lv X. et al. (2020). Research on the practical classification and privacy protection of CT images of parotid tumors based on ResNet50 model. Journal of Physics: Conference Series, 1576, 012040.
- Hu Z., Wang B., Pan X., Cao D., Gao A., Yang X., Chen Y., Lin Z. (2022). Using deep learning to distinguish malugnant from begin parotid tumors on plain computered tomography images. Fronties in Oncology, August. [CrossRef]
- Onder M., Cengiz Evli, Ezgi Türk, Orhan Kazan, İbrahim Şevki Bayrakdar, Özer Çelik, Andre Luiz Ferreira Costa, João Pedro Perez Gomes, Celso Massahiro Ogawa, Rohan Jagtap and Kaan Orhan, (2023). Deep-learning-based automatic segmentation of parotid gland on computed tomography images. Diagnosis, 13, 581.
- Xu Z., Dai Y. Liu F., Li S., Liu S., Shi L., Fu J. (2022). Parotid gland MRI segmentation based on Swin-Unet and multimodal images. arXiv preprint, arXiv:2206.03336v2.
- Kawahara D., Tsuneda M., Ozawa S., Okamoto., Nakamura M., Nishio T., Saito A., Nagata Y. (2022) Stepwise deep neural network for head and neck suto-wegmentation on CT images. Computers in Biology and Medicine, 143, 105295.
- Wang C.Y., Bochkovskiy A., Liao H.-Y. M. (2022). YOLOv7: Trainable bag-of-freebies sets new state-of-the-art for real-time object detectors. arXiv preprint, arXiv:2207.02696v1.
- Russell, B. C., Torralba, A., Murphy, K. P., & Freeman, W. T. (2008). LabelMe: A database and web-based tool for image annotation. International Journal of Computer Vision, 77(1), 157-173.
- Zuiderveld, K. (1994). Contrast limited adaptive histogram equalization. In Graphics Gems IV (pp. 474–485). Academic Press Professional, Inc.
- Duan K., Bai S., Xie L., Qi H., Huang Q., Tian Q.. (2019). CenterNet: Keypoint Triplets for Object Detection. arXiv preprint, arXiv:1904.08189.
- Tan M., Pang K., Le Q.V. (2019). EfficientDet: Scalable and Efficient Object Detection. arXiv preprint, arXiv:1911.09070.
- Ren S., He K., Girshick R., Sun J. (2015). Faster R-CNN: Towards Real-time Object Detection with Region Proposal Networks. arXiv:1506.01497.
- Bochkovskiy A., Wang C.Y., Liao H. -Y. M. (2020). YOLOv4: Optimal Speed and Accuracy of Object Detection. arXiv preprint, arXiv:1506.01497.
- Liu W., Anguelov D., Erhan D., Szegedy C., Reed C., Fu C.Y., Berg A.C. (2016). SSD: Single Shot MultiBox Detector. arXiv preprint, arXiv:1512.02325v5.
- Woo S., Park J., Lee J.Y., Kweon, I. S. (2018). CBAM: Convolutional Block Attention Module. In Proceedings of the European Conference on Computer Vision (ECCV) (pp. 3–19).
- Tan M., Pang R., Le Q.V. (2020). EfficientDet: Scalable and efficient object detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (pp. 10781–10790).
- Dosovitskiy A., Beyer L., Kolesnikov A., et al. (2021). An image is worth 16x16 words: Transformers for image recognition at scale. In Proceedings of the International Conference on Learning Representations (ICLR).
- Liu Z., Lin Y., Cao Y., et al. (2021). Swin Transformer: Hierarchical vision transformer using shifted windows. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), (pp.10012–10022).
- Wu B., Xu C., Dai X., et al. (2021). CvT: Introducing convolutions to vision transformers. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), (pp. 22–31).
- Touvron H., Cord, M. Douze, M., et al. (2021). Training data-efficient image transformers & distillation through attention. In Proceedings of the International Conference on Machine Learning (ICML), (pp. 10347–10357).
- He K., Zhang X., Ren S., Sun J. (2020). Deep Residual Learning for Image Recognition. arXiv preprint, arXiv:1512.03385.
- Ronneberger O., Fischer P., Brox T. (2015). U-Net: Convolutional N eworks for Biomedical Image Segmentation. arXiv preprint, arXiv:1505.04597.
- Zhou Z., Siddiquee M.M.R., Tajbakhsh N., Liang J. (2018). UNet++: A Nested U-Net Architecture for Medical Image Segmentation, arXiv preprint, arXiv:1807.10165.
- Cao H., Wang Y., Chen J., Jiang D., Zhang X., Tian Q., Wang M. (2021). Swin-Unet: Unet-like Pure Transformer for Medical Image Segmentation. arXiv preprint, arXiv:2105.05537.
- Chen J., Lu Y., Yu Q., Luo X., Adeli E., Wang Y., Lu L., Yuille A.L., Zhou Y. (2021). TransUNet: Transformers Make Strong Encoders for Medical Image Segmentation, arXiv preprint, arXiv:2102.04306.
- Shen X.M., Mao Liang, Yang Z.Y. Chai Z.K., Sun T.G., Xu Y., Sun Z.J. (2022). Deep learning-assisted diagnosis of parotid gland tumors by using contrast-enhanced CT imaging. Oral Disease, 29, 3325-3336. [CrossRef]
- Wang W, Chen C., Ding M., Li J., Yu H., Transbts Z.S. (2021). Multimodal brain tumor segmentation using transformer. arXiv preprint, arXiv:2103.04430.
- Chang C.C., Horng M.H., Jiang J.Y. Deep learning-based computerized tomographic imaging for differentiation and segmentation of parotid gland neoplasm, (2024), ECEI. [CrossRef]





| Methods | AP50 (mean±S.D) |
| YOLOv4 | 0.9645±0.12 |
| VOLOv7 | 0.9801±0.13 |
| CenterNet | 0.9595±0.03 |
| EfficientDet | 0.9478±3.13 |
| Faster RCNN | 0.9575±0.02 |
| SSD | 0.8920±0.02 |
| Tumor classification (Methods) | Accuracy (%) (mean±s.d) |
| ResNet | 84.6±2.18 |
| ResNet+BiFPN | 87.2±0.76 |
| ResNet+CBAM | 86.7±2.88 |
| VIT | 78.9±1.2 |
| Swin_Transformer | 91.9±1.3 |
| CVT | 89.7±1.0 |
| DEIT | 92.3±1.2 |
| Classification of Warith and Malignant/Mixed tumors (method) | Accuracy (%) (mean±s.d) |
|---|---|
| ResNet | 92.1±1.8 |
| ResNet+BiFPN | 92.9±1.2 |
| ResNet+CBAM | 92.9±1.95 |
| VIT | 78.9±1.3 |
| Swin_Transform | 94.2±1.4 |
| CVT | 89.0±1.1 |
| DEIT | 94.7±0.74 |
| Malignant and Mixed Tumor classification (method) | Accuracy (%) (mean±s.d) |
| ResNet | 83.5±3.6 |
| ResNet+BiFPN | 82.6±2.4 |
| ResNet+CBAM | 82.0±3.6 |
| VIT | 66.1±4.2 |
| Swin_Transform | 62.9±4.2 |
| CVT | 74.0±2.5 |
| DEIT | 84.4±2.3 |
| Dice similarity coefficient (mean±S.D) | malignant | Mixed | Warthin |
| Unet | 0.82±0.03 | 0.84±0.08 | 0.89±0.09 |
| Unet++ | 0.83±0.05 | 0.85±0.06 | 0.91±0.08 |
| Swin-Unet | 0.84±0.09 | 0.83±0.03 | 0.91±0.03 |
| TransUNet | 0.89±0.09 | 0.93±0.05 | 0.94±0.07 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2024 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).