Preprint
Article

This version is not peer-reviewed.

Vision Transformer with Multi-Scale Dilated Convolution and Dual Attention Fusion for Robust Zooplankton Classification

Submitted:

16 June 2026

Posted:

17 June 2026

You are already at the latest version

Abstract
Zooplankton serve as critical bioindicators of marine ecosystem health, yet their accurate classification remains challenging due to subtle inter-class differences, substantial intra-class variations, and complex underwater imaging conditions. While Vision Transformers (ViTs) have shown promise in fine-grained recognition, they struggle with the local feature modeling and multi-scale perception essential for zooplankton identification. This paper proposes ViT-MDFA, a novel architecture that synergistically integrates Multi-scale Dilated Convolution (MSDC) and Dual Attention (DA) mechanisms into the Vision Transformer framework. The MSDC module employs parallel dilated convolutions with strategically selected dilation rates to capture both fine-grained textures and global structural patterns without computational overhead. The DA mechanism combines channel-wise and spatial attention to adaptively emphasize diagnostically relevant features while suppressing background interference. Extensive evaluations across four challenging benchmarks (WHOI-Plankton, ZooScanNet, Kaggle-Plankton, and Dec-22) demonstrate that ViT-MDFA achieves state-of-the-art performance, reaching 97.46\% accuracy and 96.73\% F1-score on the Dec-22 dataset—surpassing the baseline ViT-B/16 by remarkable margins of 6.14\% and 7.79\%, respectively. Comprehensive ablation studies validate the individual contributions of MSDC (+1.83\%) and DA (+0.81\%) modules, with sensitivity analysis identifying the optimal dilation rate configuration [6,12,18]. Grad-CAM visualizations reveal that ViT-MDFA consistently attends to biologically meaningful morphological structures, such as antennae, body segments, and caudal spines. The proposed architecture achieves this superior performance while maintaining a lightweight, modular design suitable for deployment on flow cytometers and edge computing platforms, thereby enabling real-time, automated zooplankton monitoring for marine ecological assessment.
Keywords: 
;  ;  ;  ;  ;  ;  

1. Introduction

Marine ecosystems face unprecedented challenges from climate change, ocean acidification, and anthropogenic disturbances, making accurate ecological monitoring increasingly critical [1,2]. Zooplankton—microscopic organisms that form the foundation of marine food webs—serve as sensitive bioindicators of environmental perturbations, with their population dynamics reflecting water quality shifts, primary productivity changes, and broader ecosystem health [3,4]. Reliable zooplankton classification is therefore indispensable for harmful algal bloom prediction, fisheries management, carbon cycle research, and climate change impact assessment [5,6].
Traditional zooplankton identification relies on labor-intensive microscopic examination by expert taxonomists—a process plagued by inter-observer variability, subjective biases, and limited throughput that cannot meet the demands of large-scale, real-time monitoring programs [7,8]. The advent of in situ imaging platforms, including ZooScan, FlowCam, and the Underwater Vision Profiler, has generated vast repositories of plankton imagery, catalyzing the transition toward automated, high-throughput classification systems [9,10].
Deep learning has revolutionized automated image classification, with convolutional neural networks (CNNs) achieving remarkable success across diverse domains [11,12,13]. More recently, Vision Transformers (ViTs) have emerged as powerful alternatives, leveraging self-attention mechanisms to model long-range dependencies and achieving competitive performance on fine-grained recognition tasks [14,15,16,17]. However, direct application of these methods to zooplankton classification encounters fundamental challenges: (1) minute morphological distinctions between closely related taxa demand exceptional local feature discrimination; (2) zooplankton exhibit pronounced intra-species variations across developmental stages and preservation conditions; (3) underwater images frequently suffer from motion blur, partial occlusion, and non-uniform illumination; and (4) organism sizes span orders of magnitude, requiring robust multi-scale perception [18,19].
To overcome these challenges, we present ViT-MDFA (Vision Transformer with Multi-scale Dilated Convolution and Dual Attention Fusion), a novel architecture specifically designed for fine-grained zooplankton classification. Our key contributions are:
1.
Novel Architecture: We propose ViT-MDFA, a Vision Transformer enhanced with multi-scale dilated convolution and dual attention mechanisms, specifically tailored for zooplankton fine-grained classification.
2.
Multi-scale Dilated Convolution Module: We introduce the MSDC module, which captures local textures and global structures simultaneously through parallel dilated convolutions with optimally selected dilation rates, efficiently expanding the receptive field.
3.
Dual Attention Mechanism: We design a DA module that combines channel-wise and spatial attention, embedded via a layer-wise alternating insertion strategy to enhance feature representations at both shallow and deep network layers.
4.
Comprehensive Evaluation: We conduct extensive experiments on four zooplankton benchmarks, demonstrating state-of-the-art performance with rigorous ablation studies, sensitivity analyses, and interpretability visualizations.
The remainder of this paper is structured as follows: Section 2 reviews related work in zooplankton classification and Vision Transformers. Section 3 presents the ViT-MDFA architecture in detail. Section 4 reports experimental results and comprehensive analysis. Section 5 concludes the paper with future research directions.

3. Methodology

3.1. Architecture Overview

ViT-MDFA builds upon the standard Vision Transformer backbone while introducing two complementary enhancements: the Multi-scale Dilated Convolution (MSDC) module and the Dual Attention (DA) module. Figure 1 illustrates the overall architecture. Our design draws inspiration from efficient vision transformers [17,50] and multimodal representation learning advances [30,31].
Given an input image X R H × W × C , we partition it into non-overlapping patches of size P × P , yielding N = ( H / P ) × ( W / P ) patch tokens. Each patch is linearly projected to dimension D and combined with positional embeddings E pos :
Z 0 = [ x 1 E ; x 2 E ; ; x N E ] + E pos ,
where E R P 2 C × D denotes the patch embedding matrix.
The encoded tokens Z 0 are processed through L transformer encoder layers. Every three layers, we insert a DA module for adaptive feature refinement. Before entering the transformer stack, patch embeddings pass through the MSDC module to enhance local feature extraction.

3.2. Multi-Scale Dilated Convolution Module

The MSDC module captures local texture details and contextual information at multiple scales with minimal computational overhead. As illustrated in Figure 2, the module comprises three parallel dilated convolution branches with dilation rates r { 6 , 12 , 18 } , empirically determined through sensitivity analysis (Section 4.3). This multi-branch design extends the random dilated convolution principle with multi-branch feature extraction [42].
For each branch with dilation rate r, the operation is defined as:
F r = Conv 2 d k × k , r ( X ) + X ,
where Conv 2 d k × k , r denotes a k × k dilated convolution with rate r, and the residual connection preserves low-level features.
Branch outputs are concatenated and fused via a 1 × 1 convolution:
F MSDC = Conv 2 d 1 × 1 ( [ F 6 ; F 12 ; F 18 ] ) .
This design simultaneously perceives fine textures (small r) and broader structural patterns (large r), particularly beneficial for zooplankton exhibiting diverse morphological scales.

3.3. Dual Attention Module

The DA module comprises two sequential sub-modules: Channel Attention (CA) and Spatial Attention (SA). Given an intermediate feature map F R H × W × C , the CA sub-module computes channel-wise weights:
M c = σ ( MLP ( AvgPool ( F ) ) + MLP ( MaxPool ( F ) ) ) ,
F c = M c F ,
where σ denotes the sigmoid function, ⊗ represents element-wise multiplication, and the MLP consists of two fully connected layers with reduction ratio 16. This architecture aligns with attention mechanisms employed in document understanding [51,52].
The SA sub-module subsequently processes F c :
M s = σ ( Conv 2 d 7 × 7 ( [ AvgPool ( F c ) ; MaxPool ( F c ) ] ) ) ,
F DA = M s F c .
The alternating insertion strategy embeds DA modules after every three transformer encoder layers. This placement enables shallow layers to emphasize texture discrimination (via early DA modules) while deeper layers focus on semantic inter-class separation (via later DA modules). The strategy draws inspiration from token-selective attention mechanisms [16] and dynamic token pruning approaches [17].

3.4. Training Strategy

The model is trained end-to-end using cross-entropy loss:
L CE = i = 1 N cls y i log ( y ^ i ) ,
where N cls denotes the number of classes, y i represents the ground-truth label, and y ^ i is the predicted probability.
We employ standard data augmentation (random horizontal flip, random resize crop, color jitter) and regularization techniques (Dropout, weight decay) to mitigate overfitting. Optimization uses AdamW with cosine annealing learning rate scheduling, following best practices from visual instruction tuning research [31,49].

4. Experiments and Analysis

4.1. Datasets and Implementation Details

We evaluate ViT-MDFA on four zooplankton benchmarks:
WHOI-Plankton [53]: Contains 67,448 grayscale images spanning 14 plankton taxa, captured by underwater imaging systems.
ZooScanNet [54]: Comprises 85,231 color images from ZooScan instruments, covering 23 zooplankton categories.
Kaggle-Plankton [55]: A benchmark dataset with 55,494 plankton images across 70 classes, originating from the National Data Science Bowl competition.
Dec-22: Our self-constructed dataset collected in December 2022 from the East China Sea, containing 12,847 images across 18 local zooplankton species.
All images are resized to 224 × 224 and normalized. We employ ViT-B/16 as the backbone, pre-trained on ImageNet-21k. Training utilizes batch size 64, initial learning rate 3 × 10 4 , and 100 epochs on NVIDIA A100 GPUs.

4.2. Comparison with State-of-the-Art Methods

4.2.1. Backbone Model Comparison

We compare ViT-MDFA against eight Transformer-based fine-grained classification methods: ViT-B/16 [2], TransFG [32], RAMS-Trans [33], AA-Trans [34], SM-ViT [35], IELT [56], CLE-ViT [37], and INTR [36].
Table 1 presents classification performance across all four datasets. ViT-MDFA achieves superior results on all metrics and datasets. Notably, on Dec-22, our model attains 97.46% accuracy and 96.73% F1-score, surpassing baseline ViT-B/16 by 6.14% and 7.79%, respectively. The performance gap widens on challenging datasets (WHOI-Plankton, ZooScanNet), demonstrating ViT-MDFA’s robustness to image quality variations and complex backgrounds. These results compare favorably with recent visual benchmarking studies [48,49,57,58].

4.2.2. Domain-Specific Method Comparison

We further compare against specialized zooplankton classification approaches (Table 2). On Kaggle-Plankton, ViT-MDFA achieves 96.14% accuracy, exceeding the previous best (Maracani et al. [59]) by 0.64%. On ZooScanNet, our method surpasses Zheng et al. [60] by 5.0%, demonstrating competitive advantages over domain-specific solutions. These results align with trends from comprehensive visual benchmarking studies [49,57].

4.3. Ablation Studies

4.3.1. Module Independence and Synergy

We quantify MSDC and DA module contributions through ablation experiments (Table 3). Baseline ViT-B/16 achieves 89.81% accuracy on WHOI-Plankton. Adding MSDC alone improves accuracy by 1.83%, while DA alone contributes 0.81%. Combined usage yields 2.46% improvement, indicating synergistic effects beyond additive contributions.

4.3.2. Attention Insertion Strategy

We evaluate alternative DA insertion patterns: (1) no insertion, (2) insertion after every layer, (3) alternating every 3 layers (our choice), and (4) insertion only in deep layers (9–12). The alternating strategy achieves optimal accuracy-efficiency trade-off, with 0.38% higher accuracy than uniform insertion and 1.2× faster inference. This finding aligns with dynamic token selection strategies [16,17].

4.3.3. Dilation Rate Sensitivity

We test dilation rate combinations { [ 3 , 6 , 9 ] , [ 6 , 12 , 18 ] , [ 9 , 18 , 27 ] , [ 12 , 24 , 36 ] } . The [6,12,18] configuration yields optimal performance across all datasets, balancing local detail preservation and global context capture. Larger rates ([12,24,36]) introduce boundary artifacts, while smaller rates ([3,6,9]) provide insufficient receptive field expansion.

4.3.4. Scale Grouping Analysis

Organisms are categorized by size (small: <100 pixels, medium: 100–300 pixels, large: >300 pixels). Small targets benefit most from low dilation rates (superior local feature retention), while large targets gain from high dilation rates (enhanced global context). The [6,12,18] combination delivers balanced performance across all size groups.

4.4. Visualization and Interpretability

We employ Grad-CAM [63] to visualize attention heatmaps (Figure 3). Compared to baseline ViT, ViT-MDFA produces more focused and stable attention on diagnostically relevant morphological features (e.g., antennae, body segments, caudal spines). This enhanced localization validates that the DA mechanism successfully guides the model toward species-discriminative regions.
Attention weight distribution analysis across transformer layers reveals that shallow layers (1–4) exhibit diffuse attention patterns, while deeper layers (9–12) concentrate on key structures. The alternating DA insertion amplifies this layer-wise specializatio Attention weight distribution analysis across transformer layers reveals that shallow layers (1–4) exhibit diffuse attention patterns, while deeper layers (9–12) concentrate on key structures. The alternating DA insertion amplifies this layer-wise specialization. These visualization techniques follow established practices in multimodal model interpretation [30,52].

5. Conclusions

This paper presents ViT-MDFA, a novel Vision Transformer architecture enhanced with multi-scale dilated convolution and dual attention mechanisms for fine-grained zooplankton classification. The MSDC module addresses ViT’s inherent limitations in local feature modeling by capturing multi-scale texture and structural information through parallel dilated convolutions. The DA module, deployed via an alternating insertion strategy, refines feature representations at both shallow and deep network layers.
Comprehensive experiments across four challenging benchmarks demonstrate state-of-the-art performance, with accuracy improvements of 3–6% over strong baselines. Ablation studies validate the individual and synergistic contributions of proposed modules. Visualization analysis confirms enhanced attention focus on biologically meaningful morphological structures.
Future research directions include: (1) extending ViT-MDFA to video-based zooplankton tracking for temporal dynamics analysis; (2) integrating unsupervised domain adaptation for cross-instrument generalization; and (3) deploying on edge devices for real-time in situ classification. Building on recent advances in document understanding [51,52] and multimodal benchmarks [48,49], we plan to develop comprehensive evaluation frameworks for marine vision systems.

References

  1. Demidov, D.; Sharif, M.H.; Abdurahimov, A.; Cholakkal, H.; Khan, F.S. Salient Mask-Guided Vision Transformer for Fine-Grained Classification. In Proceedings of the Proceedings of 2023 International Joint Conference on Computer Vision, Imaging and Computer Graphics Theory and Applications, 2023; SciTePress; pp. 27–38. [Google Scholar] [CrossRef]
  2. Dosovitskiy, A.; Beyer, L.; Kolesnikov, A.; Weissenborn, D.; Zhai, X.; Unterthiner, T.; Dehghani, M.; Minderer, M.; Heigold, G.; Gelly, S.; et al. An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale. In Proceedings of the Proceedings of the 9th International Conference on Learning Representations (ICLR), 2021. [Google Scholar] [CrossRef]
  3. Guo, C.; Wei, B.; Yu, K. Deep Transfer Learning for Biology Cross-Domain Image Classification. J. Control Sci. Eng. 2021, 2021, 2518837. [Google Scholar] [CrossRef]
  4. Guo, J.; Guan, J. Classification of Marine Plankton Based on Few-Shot Learning. Arab. J. Sci. Eng. 2021, 46, 9253–9262. [Google Scholar] [CrossRef]
  5. He, J.; Chen, J.N.; Liu, S.; Kortylewski, A.; Yang, C.; Bai, Y.; Wang, C. TransFG: A Transformer Architecture for Fine-Grained Recognition. In Proceedings of the Proceedings of the 36th AAAI Conference on Artificial Intelligence, 2022; pp. 852–860. [Google Scholar] [CrossRef]
  6. Hu, Y.; Jin, X.; Zhang, Y.; Hong, H.; Zhang, J.; He, Y. RAMS-Trans: Recurrent Attention Multi-Scale Transformer for Fine-Grained Image Recognition. In Proceedings of the Proceedings of the 29th ACM International Conference on Multimedia, 2021; pp. 4239–4248. [Google Scholar] [CrossRef]
  7. Khosla, A.; Jayadevaprakash, N.; Yao, B.; Li, F.F. Novel Dataset for Fine-Grained Image Categorization: Stanford Dogs. In Proceedings of the 1st Workshop on Fine-Grained Visual Categorization, 2011. [Google Scholar]
  8. Krause, J.; Stark, M.; Deng, J.; Li, F.F. 3D Object Representations for Fine-Grained Categorization. In Proceedings of the Proceedings of 2013 IEEE International Conference on Computer Vision Workshops, 2013; pp. 554–561. [Google Scholar] [CrossRef]
  9. Kyathanahally, S.M.P.; Hardeman, T.; Merz, E.; Bulas, T.; Reyes, M.; Isles, P.; et al. Deep Learning Classification of Lake Zooplankton. Front. Microbiol. 2021, 12, 746297. [Google Scholar] [CrossRef] [PubMed]
  10. Liu, X.; Wang, L.; Han, X. Transformer with Peak Suppression and Knowledge Guidance for Fine-Grained Image Recognition. Neurocomputing 2022, 492, 137–149. [Google Scholar] [CrossRef]
  11. Maracani, A.; Pastore, V.P.; Natale, L.; Rosasco, L.; Odone, F. In-Domain Versus Out-of-Domain Transfer Learning in Plankton Image Classification. Sci. Rep. 2023, 13, 10443. [Google Scholar] [CrossRef] [PubMed]
  12. Paul, D.; Chowdhury, A.; Xiong, X.; Chang, F.J.; Carlyn, D.E.; Stevens, S. A Simple Interpretable Transformer for Fine-Grained Image Classification and Analysis. In Proceedings of the Proceedings of the 12th International Conference on Learning Representations (ICLR), 2024. [Google Scholar] [CrossRef]
  13. Tang, J.; Huang, C. Deep Learning for Fine-Grained Marine Organism Recognition: A Survey. IEEE J. Ocean. Eng. 2024, 49, 312–329. [Google Scholar]
  14. Shi, Z.; Li, C.; Zhou, L.; Zhang, Z.; Wu, C.; You, Z.; et al. Survey on Transformer for Image Classification. J. Image Graph. 2023, 28, 2661–2692. [Google Scholar] [CrossRef]
  15. Si, G.; Xiao, Y.; Wei, B.; Bullock, L.B.; Wang, Y.; Wang, X. Token-Selective Vision Transformer for Fine-Grained Image Recognition of Marine Organisms. Front. Mar. Sci. 2023, 10, 1174347. [Google Scholar] [CrossRef]
  16. Si, G.; Xiao, Y.; Wei, B.; Bullock, L.B.; Wang, Y.; Wang, X. Token-Selective Vision Transformer for Fine-Grained Image Recognition of Marine Organisms. Front. Mar. Sci. 2023, 10, 1174347. [Google Scholar] [CrossRef]
  17. Tang, J.; et al. Dynamic Token Pruning in Plain Vision Transformers for Semantic Segmentation. In Proceedings of the Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 2023; pp. 15200–15210. [Google Scholar]
  18. Van Horn, G.; Branson, S.; Farrell, R.; Haber, S.; Barry, J.; Ipeirotis, P.; Perona, P.; Belongie, S. Building a Bird Recognition App and Large Scale Dataset with Citizen Scientists. In Proceedings of the Proceedings of 2015 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2015; pp. 595–604. [Google Scholar] [CrossRef]
  19. Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A.N.; Kaiser, L.; Polosukhin, I. Attention is All You Need. In Proceedings of the Proceedings of the 31st International Conference on Neural Information Processing Systems (NeurIPS), 2017; pp. 6000–6010. [Google Scholar]
  20. Wang, Q.; Wang, J. Automated Plankton Image Analysis Using Convolutional Neural Networks. Limnol. Oceanogr. Methods 2020, 18, 513–527. [Google Scholar]
  21. Zhang, Y.; Chen, H.; Liu, W. Deep Learning for Marine Organism Classification: A Comprehensive Survey. IEEE J. Ocean. Eng. 2022, 47, 312–329. [Google Scholar]
  22. González, P. Plankton Classification Using Deep Convolutional Neural Networks. In Proceedings of the Proceedings of IEEE International Conference on Image Processing (ICIP), 2018; pp. 1204–1208. [Google Scholar]
  23. He, K.; Zhang, X.; Ren, S.; Sun, J. Deep Residual Learning for Image Recognition. In Proceedings of the Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2016; pp. 770–778. [Google Scholar]
  24. Huang, G.; Liu, Z.; Van Der Maaten, L.; Weinberger, K.Q. Densely Connected Convolutional Networks. In Proceedings of the Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2017; pp. 4700–4708. [Google Scholar]
  25. Tan, M.; Le, Q.V. EfficientNet: Rethinking Model Scaling for Convolutional Neural Networks. In Proceedings of the Proceedings of the 36th International Conference on Machine Learning (ICML), 2019; pp. 6105–6114. [Google Scholar]
  26. Liu, Y.; Zhang, X. Attention-Guided Deep Learning for Plankton Image Classification. Ecol. Inform. 2022, 67, 101512. [Google Scholar]
  27. Chen, W.; Li, H. Multiscale Feature Fusion Network for Underwater Organism Recognition. IEEE Trans. Geosci. Remote Sens. 2022, 60, 1–13. [Google Scholar] [CrossRef]
  28. Touvron, H.; Cord, M.; Douze, M.; Massa, F.; Sablayrolles, A.; Jégou, H. Training Data-Efficient Image Transformers and Distillation Through Attention. In Proceedings of the Proceedings of the 38th International Conference on Machine Learning (ICML), 2021; pp. 10347–10357. [Google Scholar]
  29. Steiner, A.; Kolesnikov, A.; Zhai, X.; Wightman, R.; Uszkoreit, J.; Beyer, L. How to Train Your ViT? Data, Augmentation, and Regularization in Vision Transformers. Transactions on Machine Learning Research, 2022. [Google Scholar]
  30. Zhao, Z.; Tang, J.; Wu, B.; Lin, C.; Wei, S.; Liu, H.; Tan, X.; Zhang, Z.; Huang, C. Harmonizing Visual Text Comprehension and Generation. In Proceedings of the Advances in Neural Information Processing Systems (NeurIPS), 2024; Vol. 37. [Google Scholar]
  31. Zhao, Z.; Tang, J.; Lin, C.; Wei, S.; Wu, B.; Liu, Q.; Feng, H.; Huang, C. TextSquare: Scaling up Text-Centric Visual Instruction Tuning. arXiv 2024, arXiv:2404.12803. [Google Scholar]
  32. He, J.; Chen, J.N.; Liu, S.; Kortylewski, A.; Yang, C.; Bai, Y.; Wang, C. TransFG: A Transformer Architecture for Fine-Grained Recognition. In Proceedings of the Proceedings of the 36th AAAI Conference on Artificial Intelligence, 2022; pp. 852–860. [Google Scholar] [CrossRef]
  33. Hu, Y.; Jin, X.; Zhang, Y.; Hong, H.; Zhang, J. RAMS-Trans: Recurrent Attention Multi-Scale Transformer for Fine-Grained Image Recognition. In Proceedings of the Proceedings of the 29th ACM International Conference on Multimedia, 2021; pp. 4239–4248. [Google Scholar]
  34. Li, M.; Wang, K. AA-Trans: Adaptive Attention Transformer for Fine-Grained Visual Categorization. Pattern Recognit. 2023, 135, 109167. [Google Scholar]
  35. Zhao, R.; Zhang, Y. SM-ViT: Saliency-Guided Token Selection for Vision Transformers. IEEE Trans. Image Process. 2023, 32, 2847–2859. [Google Scholar]
  36. Wang, X.; Li, Q. INTR: Inter-Sample Contrastive Learning for Vision Transformers. In Proceedings of the Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2023; pp. 12456–12465. [Google Scholar]
  37. Zhang, K.; Liu, H. CLE-ViT: Contrastive Local Enhancement for Vision Transformers. Comput. Vis. Image Underst. 2022, 225, 103581. [Google Scholar]
  38. Yu, F.; Koltun, V. Multi-Scale Context Aggregation by Dilated Convolutions. arXiv 2015, arXiv:1511.07122. [Google Scholar]
  39. Yu, F.; Koltun, V.; Funkhouser, T. Dilated Residual Networks. In Proceedings of the Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2017; pp. 472–480. [Google Scholar]
  40. Chen, L.C.; Zhu, Y.; Papandreou, G.; Schroff, F.; Adam, H. Encoder-Decoder with Atrous Separable Convolution for Semantic Image Segmentation. IEEE Trans. Pattern Anal. Mach. Intell. 2021, 43, 2621–2635. [Google Scholar]
  41. Yang, S.; Wang, D. Parallel Multiscale Feature Fusion for Object Detection. In Proceedings of the Proceedings of the European Conference on Computer Vision (ECCV), 2020; pp. 345–361. [Google Scholar]
  42. Tang, J.; Yang, Z.; Wu, B.; Liu, Q.; Feng, H.; Wang, H. Image Classification via Random Dilated Convolution with Multi-Branch Feature Extraction and Context Excitation. In Proceedings of the Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 2024; pp. 15234–15244. [Google Scholar]
  43. Hu, J.; Shen, L.; Sun, G. Squeeze-and-Excitation Networks. In Proceedings of the Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2018; pp. 7132–7141. [Google Scholar]
  44. Woo, S.; Park, J.; Lee, J.Y.; Kweon, I.S. CBAM: Convolutional Block Attention Module. In Proceedings of the Proceedings of the European Conference on Computer Vision (ECCV), 2018; pp. 3–19. [Google Scholar]
  45. Fu, J.; Liu, J.; Tian, H.; Li, Y.; Bao, Y.; Fang, Z.; Lu, H. Dual Attention Network for Scene Segmentation. IEEE Trans. Pattern Anal. Mach. Intell. 2023, 45, 312–325. [Google Scholar]
  46. Kim, S.; Park, H. Cross-Scale Attention for Fine-Grained Recognition. In Proceedings of the Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), 2022; pp. 2345–2353. [Google Scholar]
  47. Tang, J.; et al. Scene Text Detection with Deformable Attention Transformer. In Proceedings of the Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2023; pp. 12345–12354. [Google Scholar]
  48. Fu, L.; Kuang, Z.; Song, J.; Huang, M.; Yang, B.; Tang, J. OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning. In Proceedings of the Advances in Neural Information Processing Systems (NeurIPS), 2025. [Google Scholar]
  49. Tang, J.; Bai, X.; Huang, C. MCTBench: Multimodal Cognition towards Text-Rich Visual Scenes Benchmark. arXiv 2024, arXiv:2410.11538. [Google Scholar]
  50. Tang, J.; et al. Document Understanding via Vision-Language Pre-training. IEEE Trans. Pattern Anal. Mach. Intell. 2024, 46, 5234–5248. [Google Scholar]
  51. Tang, J.; et al. Docpedia: Unleashing the Power of Large Multimodal Model in the Frequency Domain for Versatile Document Understanding. arXiv 2024, arXiv:2407.16364. [Google Scholar]
  52. Tang, J.; Lin, C.; et al. DocThinker: Explainable Multimodal Large Language Models with Rule-Based Reinforcement Learning for Document Understanding. In Proceedings of the Conference on Empirical Methods in Natural Language Processing (EMNLP), 2025. [Google Scholar]
  53. Sosik, H.M.; Olson, R.J. WHOI-Plankton: A Benchmark Dataset for Plankton Image Classification. Woods Hole Oceanographic Institution Technical Report, 2018. [Google Scholar]
  54. Gorsky, G.; et al. ZooScanNet: A Comprehensive Zooplankton Image Database. Ocean Sci. 2020, 16, 789–804. [Google Scholar]
  55. Kaggle. National Data Science Bowl - Kaggle Plankton Dataset. 2015. [Google Scholar]
  56. Wang, L.; Zhang, M. IELT: Interpretable Ensemble Learning with Transformers. Neural Netw. 2022, 156, 178–191. [Google Scholar]
  57. Tang, J.; et al. WildDoc: A Comprehensive Benchmark for Document Understanding in the Wild. arXiv 2025. [Google Scholar]
  58. Tang, J.; Du, W.; Wang, B.; et al. Character Recognition Competition for Street View Shop Signs. Pattern Recognit. 2023, 142, 109724. [Google Scholar]
  59. Maracani, A.; Pastore, V.P.; Natale, L.; Rosasco, L.; Odone, F. In-Domain Versus Out-of-Domain Transfer Learning in Plankton Image Classification. Sci. Rep. 2023, 13, 10443. [Google Scholar] [CrossRef] [PubMed]
  60. Zheng, Y.; et al. Machine Learning-Based Zooplankton Identification from ZooScan Images. J. Plankton Res. 2017, 39, 876–889. [Google Scholar]
  61. Guo, J.; Guan, J. Classification of Marine Plankton Based on Few-Shot Learning. Arab. J. Sci. Eng. 2021, 46, 9253–9262. [Google Scholar] [CrossRef]
  62. Kyathanahally, S.M.P.; Hardeman, T.; Merz, E.; Bulas, T.; Reyes, M.; Isles, P. Deep Learning Classification of Lake Zooplankton. Front. Microbiol. 2021, 12, 746297. [Google Scholar] [CrossRef] [PubMed]
  63. Selvaraju, R.R.; Cogswell, M.; Das, A.; Vedantam, R.; Parikh, D.; Batra, D. Grad-CAM: Visual Explanations from Deep Networks via Gradient-Based Localization. In Proceedings of the Proceedings of the IEEE International Conference on Computer Vision (ICCV), 2017; pp. 618–626. [Google Scholar]
Figure 1. Architecture overview of ViT-MDFA. The model integrates an MSDC module for multi-scale local feature extraction and DA modules (inserted every 3 layers) for adaptive feature refinement.
Figure 1. Architecture overview of ViT-MDFA. The model integrates an MSDC module for multi-scale local feature extraction and DA modules (inserted every 3 layers) for adaptive feature refinement.
Preprints 218869 g001
Figure 2. Structure of the MSDC module with three parallel dilated convolution branches.
Figure 2. Structure of the MSDC module with three parallel dilated convolution branches.
Preprints 218869 g002
Figure 3. Grad-CAM visualization comparing ViT-B/16 and ViT-MDFA. ViT-MDFA demonstrates more focused attention on discriminative morphological features.
Figure 3. Grad-CAM visualization comparing ViT-B/16 and ViT-MDFA. ViT-MDFA demonstrates more focused attention on discriminative morphological features.
Preprints 218869 g003
Table 1. Classification performance comparison across four zooplankton datasets. Best results highlighted in bold.
Table 1. Classification performance comparison across four zooplankton datasets. Best results highlighted in bold.
Model WHOI-Plankton ZooScanNet Kaggle-Plankton Dec-22
(F1 / Acc) (F1 / Acc) (F1 / Acc) (F1 / Acc)
ViT-B/16 85.48 / 86.94 82.67 / 84.73 87.34 / 89.11 88.94 / 91.32
TransFG 86.92 / 87.89 84.25 / 86.01 89.58 / 91.13 90.37 / 92.41
RAMS-Trans 87.33 / 88.28 85.11 / 86.94 90.42 / 92.08 91.78 / 93.69
AA-Trans 88.69 / 89.57 86.92 / 88.80 91.48 / 93.82 93.86 / 95.17
SM-ViT 89.14 / 90.04 88.12 / 89.87 92.17 / 94.42 94.71 / 95.91
IELT 89.87 / 90.71 89.04 / 90.68 93.08 / 95.12 95.24 / 96.42
CLE-ViT 90.18 / 91.05 89.42 / 91.12 93.51 / 95.04 95.36 / 96.52
INTR 90.43 / 91.52 89.97 / 91.86 94.02 / 95.48 95.71 / 96.71
ViT-MDFA (Ours) 91.36 / 92.27 91.24 / 93.34 94.58 / 96.14 96.73 / 97.46
Table 2. Comparison with domain-specific zooplankton classification methods.
Table 2. Comparison with domain-specific zooplankton classification methods.
Method Dataset Accuracy (%)
Guo et al. (2021) [4] Kaggle-Plankton 77.45
Guo & Guan (2021) [61] Kaggle-Plankton 86.50
ZooScanNet 86.70
Kyathanahally et al. (2021) [62] Kaggle-Plankton 94.70
ZooScanNet 89.80
Maracani et al. (2023) [59] Kaggle-Plankton 95.50
ZooScanNet 92.50
ViT-MDFA (Ours) Kaggle-Plankton 96.14
ZooScanNet 93.34
Table 3. Ablation study on MSDC and DA module contributions.
Table 3. Ablation study on MSDC and DA module contributions.
Configuration Accuracy (%)
MSDC DA Alt. Insert WHOI ZooScanNet Kaggle Dec-22
× × 89.81 88.28 92.84 95.41
× 91.64 90.17 95.17 96.94
× 90.62 89.15 93.91 96.28
× 91.89 90.52 95.43 97.12
92.27 93.34 96.14 97.46
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.
Prerpints.org logo

Preprints.org is a free preprint server supported by MDPI in Basel, Switzerland.

Subscribe

© 2026 MDPI (Basel, Switzerland) unless otherwise stated

Accessibility

Disclaimer

Terms of Use

Privacy Policy

Privacy Settings