Submitted:
16 June 2026
Posted:
17 June 2026
You are already at the latest version
Abstract
Keywords:
1. Introduction
- 1.
- Novel Architecture: We propose ViT-MDFA, a Vision Transformer enhanced with multi-scale dilated convolution and dual attention mechanisms, specifically tailored for zooplankton fine-grained classification.
- 2.
- Multi-scale Dilated Convolution Module: We introduce the MSDC module, which captures local textures and global structures simultaneously through parallel dilated convolutions with optimally selected dilation rates, efficiently expanding the receptive field.
- 3.
- Dual Attention Mechanism: We design a DA module that combines channel-wise and spatial attention, embedded via a layer-wise alternating insertion strategy to enhance feature representations at both shallow and deep network layers.
- 4.
- Comprehensive Evaluation: We conduct extensive experiments on four zooplankton benchmarks, demonstrating state-of-the-art performance with rigorous ablation studies, sensitivity analyses, and interpretability visualizations.
2. Related Work
2.1. Automated Zooplankton Classification
2.2. Vision Transformers for Fine-Grained Recognition
2.3. Multi-Scale and Attention Mechanisms
3. Methodology
3.1. Architecture Overview
3.2. Multi-Scale Dilated Convolution Module
3.3. Dual Attention Module
3.4. Training Strategy
4. Experiments and Analysis
4.1. Datasets and Implementation Details
4.2. Comparison with State-of-the-Art Methods
4.2.1. Backbone Model Comparison
4.2.2. Domain-Specific Method Comparison
4.3. Ablation Studies
4.3.1. Module Independence and Synergy
4.3.2. Attention Insertion Strategy
4.3.3. Dilation Rate Sensitivity
4.3.4. Scale Grouping Analysis
4.4. Visualization and Interpretability
5. Conclusions
References
- Demidov, D.; Sharif, M.H.; Abdurahimov, A.; Cholakkal, H.; Khan, F.S. Salient Mask-Guided Vision Transformer for Fine-Grained Classification. In Proceedings of the Proceedings of 2023 International Joint Conference on Computer Vision, Imaging and Computer Graphics Theory and Applications, 2023; SciTePress; pp. 27–38. [Google Scholar] [CrossRef]
- Dosovitskiy, A.; Beyer, L.; Kolesnikov, A.; Weissenborn, D.; Zhai, X.; Unterthiner, T.; Dehghani, M.; Minderer, M.; Heigold, G.; Gelly, S.; et al. An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale. In Proceedings of the Proceedings of the 9th International Conference on Learning Representations (ICLR), 2021. [Google Scholar] [CrossRef]
- Guo, C.; Wei, B.; Yu, K. Deep Transfer Learning for Biology Cross-Domain Image Classification. J. Control Sci. Eng. 2021, 2021, 2518837. [Google Scholar] [CrossRef]
- Guo, J.; Guan, J. Classification of Marine Plankton Based on Few-Shot Learning. Arab. J. Sci. Eng. 2021, 46, 9253–9262. [Google Scholar] [CrossRef]
- He, J.; Chen, J.N.; Liu, S.; Kortylewski, A.; Yang, C.; Bai, Y.; Wang, C. TransFG: A Transformer Architecture for Fine-Grained Recognition. In Proceedings of the Proceedings of the 36th AAAI Conference on Artificial Intelligence, 2022; pp. 852–860. [Google Scholar] [CrossRef]
- Hu, Y.; Jin, X.; Zhang, Y.; Hong, H.; Zhang, J.; He, Y. RAMS-Trans: Recurrent Attention Multi-Scale Transformer for Fine-Grained Image Recognition. In Proceedings of the Proceedings of the 29th ACM International Conference on Multimedia, 2021; pp. 4239–4248. [Google Scholar] [CrossRef]
- Khosla, A.; Jayadevaprakash, N.; Yao, B.; Li, F.F. Novel Dataset for Fine-Grained Image Categorization: Stanford Dogs. In Proceedings of the 1st Workshop on Fine-Grained Visual Categorization, 2011. [Google Scholar]
- Krause, J.; Stark, M.; Deng, J.; Li, F.F. 3D Object Representations for Fine-Grained Categorization. In Proceedings of the Proceedings of 2013 IEEE International Conference on Computer Vision Workshops, 2013; pp. 554–561. [Google Scholar] [CrossRef]
- Kyathanahally, S.M.P.; Hardeman, T.; Merz, E.; Bulas, T.; Reyes, M.; Isles, P.; et al. Deep Learning Classification of Lake Zooplankton. Front. Microbiol. 2021, 12, 746297. [Google Scholar] [CrossRef] [PubMed]
- Liu, X.; Wang, L.; Han, X. Transformer with Peak Suppression and Knowledge Guidance for Fine-Grained Image Recognition. Neurocomputing 2022, 492, 137–149. [Google Scholar] [CrossRef]
- Maracani, A.; Pastore, V.P.; Natale, L.; Rosasco, L.; Odone, F. In-Domain Versus Out-of-Domain Transfer Learning in Plankton Image Classification. Sci. Rep. 2023, 13, 10443. [Google Scholar] [CrossRef] [PubMed]
- Paul, D.; Chowdhury, A.; Xiong, X.; Chang, F.J.; Carlyn, D.E.; Stevens, S. A Simple Interpretable Transformer for Fine-Grained Image Classification and Analysis. In Proceedings of the Proceedings of the 12th International Conference on Learning Representations (ICLR), 2024. [Google Scholar] [CrossRef]
- Tang, J.; Huang, C. Deep Learning for Fine-Grained Marine Organism Recognition: A Survey. IEEE J. Ocean. Eng. 2024, 49, 312–329. [Google Scholar]
- Shi, Z.; Li, C.; Zhou, L.; Zhang, Z.; Wu, C.; You, Z.; et al. Survey on Transformer for Image Classification. J. Image Graph. 2023, 28, 2661–2692. [Google Scholar] [CrossRef]
- Si, G.; Xiao, Y.; Wei, B.; Bullock, L.B.; Wang, Y.; Wang, X. Token-Selective Vision Transformer for Fine-Grained Image Recognition of Marine Organisms. Front. Mar. Sci. 2023, 10, 1174347. [Google Scholar] [CrossRef]
- Si, G.; Xiao, Y.; Wei, B.; Bullock, L.B.; Wang, Y.; Wang, X. Token-Selective Vision Transformer for Fine-Grained Image Recognition of Marine Organisms. Front. Mar. Sci. 2023, 10, 1174347. [Google Scholar] [CrossRef]
- Tang, J.; et al. Dynamic Token Pruning in Plain Vision Transformers for Semantic Segmentation. In Proceedings of the Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 2023; pp. 15200–15210. [Google Scholar]
- Van Horn, G.; Branson, S.; Farrell, R.; Haber, S.; Barry, J.; Ipeirotis, P.; Perona, P.; Belongie, S. Building a Bird Recognition App and Large Scale Dataset with Citizen Scientists. In Proceedings of the Proceedings of 2015 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2015; pp. 595–604. [Google Scholar] [CrossRef]
- Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A.N.; Kaiser, L.; Polosukhin, I. Attention is All You Need. In Proceedings of the Proceedings of the 31st International Conference on Neural Information Processing Systems (NeurIPS), 2017; pp. 6000–6010. [Google Scholar]
- Wang, Q.; Wang, J. Automated Plankton Image Analysis Using Convolutional Neural Networks. Limnol. Oceanogr. Methods 2020, 18, 513–527. [Google Scholar]
- Zhang, Y.; Chen, H.; Liu, W. Deep Learning for Marine Organism Classification: A Comprehensive Survey. IEEE J. Ocean. Eng. 2022, 47, 312–329. [Google Scholar]
- González, P. Plankton Classification Using Deep Convolutional Neural Networks. In Proceedings of the Proceedings of IEEE International Conference on Image Processing (ICIP), 2018; pp. 1204–1208. [Google Scholar]
- He, K.; Zhang, X.; Ren, S.; Sun, J. Deep Residual Learning for Image Recognition. In Proceedings of the Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2016; pp. 770–778. [Google Scholar]
- Huang, G.; Liu, Z.; Van Der Maaten, L.; Weinberger, K.Q. Densely Connected Convolutional Networks. In Proceedings of the Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2017; pp. 4700–4708. [Google Scholar]
- Tan, M.; Le, Q.V. EfficientNet: Rethinking Model Scaling for Convolutional Neural Networks. In Proceedings of the Proceedings of the 36th International Conference on Machine Learning (ICML), 2019; pp. 6105–6114. [Google Scholar]
- Liu, Y.; Zhang, X. Attention-Guided Deep Learning for Plankton Image Classification. Ecol. Inform. 2022, 67, 101512. [Google Scholar]
- Chen, W.; Li, H. Multiscale Feature Fusion Network for Underwater Organism Recognition. IEEE Trans. Geosci. Remote Sens. 2022, 60, 1–13. [Google Scholar] [CrossRef]
- Touvron, H.; Cord, M.; Douze, M.; Massa, F.; Sablayrolles, A.; Jégou, H. Training Data-Efficient Image Transformers and Distillation Through Attention. In Proceedings of the Proceedings of the 38th International Conference on Machine Learning (ICML), 2021; pp. 10347–10357. [Google Scholar]
- Steiner, A.; Kolesnikov, A.; Zhai, X.; Wightman, R.; Uszkoreit, J.; Beyer, L. How to Train Your ViT? Data, Augmentation, and Regularization in Vision Transformers. Transactions on Machine Learning Research, 2022. [Google Scholar]
- Zhao, Z.; Tang, J.; Wu, B.; Lin, C.; Wei, S.; Liu, H.; Tan, X.; Zhang, Z.; Huang, C. Harmonizing Visual Text Comprehension and Generation. In Proceedings of the Advances in Neural Information Processing Systems (NeurIPS), 2024; Vol. 37. [Google Scholar]
- Zhao, Z.; Tang, J.; Lin, C.; Wei, S.; Wu, B.; Liu, Q.; Feng, H.; Huang, C. TextSquare: Scaling up Text-Centric Visual Instruction Tuning. arXiv 2024, arXiv:2404.12803. [Google Scholar]
- He, J.; Chen, J.N.; Liu, S.; Kortylewski, A.; Yang, C.; Bai, Y.; Wang, C. TransFG: A Transformer Architecture for Fine-Grained Recognition. In Proceedings of the Proceedings of the 36th AAAI Conference on Artificial Intelligence, 2022; pp. 852–860. [Google Scholar] [CrossRef]
- Hu, Y.; Jin, X.; Zhang, Y.; Hong, H.; Zhang, J. RAMS-Trans: Recurrent Attention Multi-Scale Transformer for Fine-Grained Image Recognition. In Proceedings of the Proceedings of the 29th ACM International Conference on Multimedia, 2021; pp. 4239–4248. [Google Scholar]
- Li, M.; Wang, K. AA-Trans: Adaptive Attention Transformer for Fine-Grained Visual Categorization. Pattern Recognit. 2023, 135, 109167. [Google Scholar]
- Zhao, R.; Zhang, Y. SM-ViT: Saliency-Guided Token Selection for Vision Transformers. IEEE Trans. Image Process. 2023, 32, 2847–2859. [Google Scholar]
- Wang, X.; Li, Q. INTR: Inter-Sample Contrastive Learning for Vision Transformers. In Proceedings of the Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2023; pp. 12456–12465. [Google Scholar]
- Zhang, K.; Liu, H. CLE-ViT: Contrastive Local Enhancement for Vision Transformers. Comput. Vis. Image Underst. 2022, 225, 103581. [Google Scholar]
- Yu, F.; Koltun, V. Multi-Scale Context Aggregation by Dilated Convolutions. arXiv 2015, arXiv:1511.07122. [Google Scholar]
- Yu, F.; Koltun, V.; Funkhouser, T. Dilated Residual Networks. In Proceedings of the Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2017; pp. 472–480. [Google Scholar]
- Chen, L.C.; Zhu, Y.; Papandreou, G.; Schroff, F.; Adam, H. Encoder-Decoder with Atrous Separable Convolution for Semantic Image Segmentation. IEEE Trans. Pattern Anal. Mach. Intell. 2021, 43, 2621–2635. [Google Scholar]
- Yang, S.; Wang, D. Parallel Multiscale Feature Fusion for Object Detection. In Proceedings of the Proceedings of the European Conference on Computer Vision (ECCV), 2020; pp. 345–361. [Google Scholar]
- Tang, J.; Yang, Z.; Wu, B.; Liu, Q.; Feng, H.; Wang, H. Image Classification via Random Dilated Convolution with Multi-Branch Feature Extraction and Context Excitation. In Proceedings of the Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 2024; pp. 15234–15244. [Google Scholar]
- Hu, J.; Shen, L.; Sun, G. Squeeze-and-Excitation Networks. In Proceedings of the Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2018; pp. 7132–7141. [Google Scholar]
- Woo, S.; Park, J.; Lee, J.Y.; Kweon, I.S. CBAM: Convolutional Block Attention Module. In Proceedings of the Proceedings of the European Conference on Computer Vision (ECCV), 2018; pp. 3–19. [Google Scholar]
- Fu, J.; Liu, J.; Tian, H.; Li, Y.; Bao, Y.; Fang, Z.; Lu, H. Dual Attention Network for Scene Segmentation. IEEE Trans. Pattern Anal. Mach. Intell. 2023, 45, 312–325. [Google Scholar]
- Kim, S.; Park, H. Cross-Scale Attention for Fine-Grained Recognition. In Proceedings of the Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), 2022; pp. 2345–2353. [Google Scholar]
- Tang, J.; et al. Scene Text Detection with Deformable Attention Transformer. In Proceedings of the Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2023; pp. 12345–12354. [Google Scholar]
- Fu, L.; Kuang, Z.; Song, J.; Huang, M.; Yang, B.; Tang, J. OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning. In Proceedings of the Advances in Neural Information Processing Systems (NeurIPS), 2025. [Google Scholar]
- Tang, J.; Bai, X.; Huang, C. MCTBench: Multimodal Cognition towards Text-Rich Visual Scenes Benchmark. arXiv 2024, arXiv:2410.11538. [Google Scholar]
- Tang, J.; et al. Document Understanding via Vision-Language Pre-training. IEEE Trans. Pattern Anal. Mach. Intell. 2024, 46, 5234–5248. [Google Scholar]
- Tang, J.; et al. Docpedia: Unleashing the Power of Large Multimodal Model in the Frequency Domain for Versatile Document Understanding. arXiv 2024, arXiv:2407.16364. [Google Scholar]
- Tang, J.; Lin, C.; et al. DocThinker: Explainable Multimodal Large Language Models with Rule-Based Reinforcement Learning for Document Understanding. In Proceedings of the Conference on Empirical Methods in Natural Language Processing (EMNLP), 2025. [Google Scholar]
- Sosik, H.M.; Olson, R.J. WHOI-Plankton: A Benchmark Dataset for Plankton Image Classification. Woods Hole Oceanographic Institution Technical Report, 2018. [Google Scholar]
- Gorsky, G.; et al. ZooScanNet: A Comprehensive Zooplankton Image Database. Ocean Sci. 2020, 16, 789–804. [Google Scholar]
- Kaggle. National Data Science Bowl - Kaggle Plankton Dataset. 2015. [Google Scholar]
- Wang, L.; Zhang, M. IELT: Interpretable Ensemble Learning with Transformers. Neural Netw. 2022, 156, 178–191. [Google Scholar]
- Tang, J.; et al. WildDoc: A Comprehensive Benchmark for Document Understanding in the Wild. arXiv 2025. [Google Scholar]
- Tang, J.; Du, W.; Wang, B.; et al. Character Recognition Competition for Street View Shop Signs. Pattern Recognit. 2023, 142, 109724. [Google Scholar]
- Maracani, A.; Pastore, V.P.; Natale, L.; Rosasco, L.; Odone, F. In-Domain Versus Out-of-Domain Transfer Learning in Plankton Image Classification. Sci. Rep. 2023, 13, 10443. [Google Scholar] [CrossRef] [PubMed]
- Zheng, Y.; et al. Machine Learning-Based Zooplankton Identification from ZooScan Images. J. Plankton Res. 2017, 39, 876–889. [Google Scholar]
- Guo, J.; Guan, J. Classification of Marine Plankton Based on Few-Shot Learning. Arab. J. Sci. Eng. 2021, 46, 9253–9262. [Google Scholar] [CrossRef]
- Kyathanahally, S.M.P.; Hardeman, T.; Merz, E.; Bulas, T.; Reyes, M.; Isles, P. Deep Learning Classification of Lake Zooplankton. Front. Microbiol. 2021, 12, 746297. [Google Scholar] [CrossRef] [PubMed]
- Selvaraju, R.R.; Cogswell, M.; Das, A.; Vedantam, R.; Parikh, D.; Batra, D. Grad-CAM: Visual Explanations from Deep Networks via Gradient-Based Localization. In Proceedings of the Proceedings of the IEEE International Conference on Computer Vision (ICCV), 2017; pp. 618–626. [Google Scholar]



| Model | WHOI-Plankton | ZooScanNet | Kaggle-Plankton | Dec-22 |
|---|---|---|---|---|
| (F1 / Acc) | (F1 / Acc) | (F1 / Acc) | (F1 / Acc) | |
| ViT-B/16 | 85.48 / 86.94 | 82.67 / 84.73 | 87.34 / 89.11 | 88.94 / 91.32 |
| TransFG | 86.92 / 87.89 | 84.25 / 86.01 | 89.58 / 91.13 | 90.37 / 92.41 |
| RAMS-Trans | 87.33 / 88.28 | 85.11 / 86.94 | 90.42 / 92.08 | 91.78 / 93.69 |
| AA-Trans | 88.69 / 89.57 | 86.92 / 88.80 | 91.48 / 93.82 | 93.86 / 95.17 |
| SM-ViT | 89.14 / 90.04 | 88.12 / 89.87 | 92.17 / 94.42 | 94.71 / 95.91 |
| IELT | 89.87 / 90.71 | 89.04 / 90.68 | 93.08 / 95.12 | 95.24 / 96.42 |
| CLE-ViT | 90.18 / 91.05 | 89.42 / 91.12 | 93.51 / 95.04 | 95.36 / 96.52 |
| INTR | 90.43 / 91.52 | 89.97 / 91.86 | 94.02 / 95.48 | 95.71 / 96.71 |
| ViT-MDFA (Ours) | 91.36 / 92.27 | 91.24 / 93.34 | 94.58 / 96.14 | 96.73 / 97.46 |
| Method | Dataset | Accuracy (%) |
|---|---|---|
| Guo et al. (2021) [4] | Kaggle-Plankton | 77.45 |
| Guo & Guan (2021) [61] | Kaggle-Plankton | 86.50 |
| ZooScanNet | 86.70 | |
| Kyathanahally et al. (2021) [62] | Kaggle-Plankton | 94.70 |
| ZooScanNet | 89.80 | |
| Maracani et al. (2023) [59] | Kaggle-Plankton | 95.50 |
| ZooScanNet | 92.50 | |
| ViT-MDFA (Ours) | Kaggle-Plankton | 96.14 |
| ZooScanNet | 93.34 |
| Configuration | Accuracy (%) | |||||
|---|---|---|---|---|---|---|
| MSDC | DA | Alt. Insert | WHOI | ZooScanNet | Kaggle | Dec-22 |
| × | × | − | 89.81 | 88.28 | 92.84 | 95.41 |
| ✓ | × | − | 91.64 | 90.17 | 95.17 | 96.94 |
| × | ✓ | − | 90.62 | 89.15 | 93.91 | 96.28 |
| ✓ | ✓ | × | 91.89 | 90.52 | 95.43 | 97.12 |
| ✓ | ✓ | ✓ | 92.27 | 93.34 | 96.14 | 97.46 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).