Submitted:
19 June 2024
Posted:
20 June 2024
You are already at the latest version
Abstract
Keywords:
1. Introduction
- Enhancing few-shot crop leaf disease classification accuracy by multimodal integration: By fine-tuning CLIP, we integrate image and text information, providing a paradigm for the development of multimodality in the agricultural field. Experimental results show that our method exhibits excellent classification performance in few-shot scenarios, effectively solving the challenge of data scarcity.
- Fine-grained disease feature description driven by VLM: We innovatively leverage Qwen-VL’s fine-grained recognition to generate detailed textual descriptions of crop diseases as prompt texts from a set of infected crop leaf images to assist in generating discriminative classifier weights. This enhances the model’s sensitivity and accuracy in identifying complex disease from the images of infected crop leaves.
- Enhancing key textual features by cross-attention and SE attention: In the process of processing prompt texts, we use cross-attention and SE attention respectively in the training-free and training-required modes to guide the model’s attention to important textual features. By dynamically adjusting the weights of crucial prompt texts’ features, we effectively improve the quality of the model’s classification weights.
2. Related Work
2.1. Development of Crop Disease Image Classification Technology
2.2. Vision-Language Models and Fine-tuning
2.3. Automatic Generation of Prompt Text
2.4. Attention Mechanism
3. Method
3.1. Background
3.1.1. Zero-Shot CLIP
3.1.2. Cache Model
3.1.3. APE
- The relationship between and , determined using Equation (1), represents the cosine similarity between the test image and the prompt texts.
- The relationship between and can be calculated using a similar method as described in Equation (2):
- The relationship between and involves APE’s zero-shot CLIP prediction on training data, denoted as . To evaluate CLIP’s downstream recognition capability, the KL-divergence, , between and is calculated:where is a smoothing factor. represents the score contribution of each training feature to the final prediction.
3.1.4. APE-T
- For in Equation (1), APE-T first pads the E-channel into D-channels as by filling the extra channels with zero. The padded , denoted as , is added to , updating CLIP’s zero-shot prediction by the optimized textual features, formulated as:
- For in Equation (4), APE-T first broadcasts the C-embedding into as by repeating it for each category. Next, APE-T adds the expanded to element-wise, improving the cache model’s few-shot prediction by optimizing training-set features. This process is formulated as:
- For in Equation (5), APE-T makes it learnable during training, allowing adaptive learning of optimal cache scores for different training-set features, determining their contribution to predictions.
3.2. Image-Driven Prompt Text Generation
- Representative image selection: For simplicity, we extract traditional image features for clustering. Recognizing the importance of color and texture features in crop disease recognition—characterized by high stability and intuitiveness, particularly in distinguishing subtle and complex disease types. We utilize these features to cluster each category separately, and select M representative images for each of the C categories to form a collection , where denotes the set of representative images belonging to class i.
- Prompt text generation: We sequentially input the selected representative images from each category into Qwen-VL, employing the following unified template command for querying: “Can you help me describe this [CLASS] leaf?”, where [CLASS] is replaced with the specific disease category. This operation aims at generating detailed descriptions that are closely aligned with the image content. As shown in Figure 3, we show three different categories of crop diseases, the prompt text generated by Qwen-VL with image-driven guidance, and the prompt text generated by the language model GPT3.5 [49] and Qwen-VL without image-driven guidance. For instance, in the image-driven prompt text generated for the “grape leaf blight” category, “dark brown spots” are identified as crucial indicators of the disease, and the generated text also describes additional information about the blade surface. In contrast, without image-driven guidance, although the text generated by GPT3.5 and Qwen-VL also includes relevant disease features, it is not as detailed as the prompt text generated with image-driven input.
- Integrate text information: We consolidate and organize the collection of textual descriptions corresponding to representative images of all categories. The generated prompt texts will serve as an important basis for subsequent classification weights.
3.3. Text Feature Fusion in Training-Free (VLCD) Mode
3.4. Text Feature Enhancement in Training-Required (VLCD-T) Mode
4. Experiment
4.1. Settings
4.1.1. Dataset
4.1.2. Implementation Details
4.2. Performance Analysis
4.3. Ablation Study
4.3.1. Different Prompt Texts
4.3.2. Representative Image Selection Strategy
4.3.3. Effectiveness of Attention Mechanisms
4.3.4. Different Network Backbones
5. Visualization
5.1. Performance Comparison of Different Prompt Text Generation Strategies
5.2. Dynamic Changes in Accuracy under the SE Attention Module
6. Summary
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Acknowledgments
Conflicts of Interest
References
- Li, L.; Zhang, S.; Wang, B. Plant disease detection and classification by deep learning—a review. IEEE Access 2021, 9, 56683–56698. [Google Scholar] [CrossRef]
- Xi, C.; Yunzhi, W.; Youhua, Z.; et al. Image recognition of stored grain pests: based on deep convolutional neural network. Chinese Agricultural Science Bulletin 2018, 34, 154–158. [Google Scholar]
- Radford, A.; Kim, J. W.; Hallacy, C.; et al. Learning transferable visual models from natural language supervision. In International Conference on Machine Learning; PMLR, 2021, pp. 8748–8763.
- Bai, J.; Bai, S.; Yang, S.; et al. Qwen-vl: A versatile vision-language model for understanding, localization, text reading, and beyond. arXiv preprint, arXiv:2308.12966, 2023.
- Zhang, R.; Fang, R.; Zhang, W.; et al. Tip-adapter: Training-free clip-adapter for better vision-language modeling. arXiv preprint, arXiv:2111.03930, 2021.
- Zhu, X.; Zhang, R.; He, B.; et al. Not all features matter: Enhancing few-shot clip with adaptive prior refinement. In Proceedings of the IEEE/CVF International Conference on Computer Vision; 2023; pp. 2605–2615. [Google Scholar]
- IRMAK, G.; SAYGILI, A. A Novel Approach for Tomato Leaf Disease Classification with Deep Convolutional Neural Networks. Journal of Agricultural Sciences 2024, 30, 367–385. [Google Scholar] [CrossRef]
- Ferentinos, K. P. Deep learning models for plant disease detection and diagnosis. Computers and Electronics in Agriculture 2018, 145, 311–318. [Google Scholar] [CrossRef]
- Guo, P.; Liu, T.; Li, N. Design of automatic recognition of cucumber disease image. Information Technology Journal 2014, 13, 2129. [Google Scholar] [CrossRef]
- Zhang, S.; Wu, X.; You, Z.; et al. Leaf image based cucumber disease recognition using sparse representation classification. Computers and Electronics in Agriculture 2017, 134, 135–141. [Google Scholar] [CrossRef]
- Kaya, A.; Keceli, A. S.; Catal, C.; et al. Analysis of transfer learning for deep neural network based plant classification models. Computers and Electronics in Agriculture 2019, 158, 20–29. [Google Scholar] [CrossRef]
- Bai, Y.; Hou, F.; Fan, X.; et al. An interpretable high-accuracy method for rice disease detection based on multi-source data and transfer learning. Agriculture 2023, 13, 1–23. [Google Scholar]
- Li, Y.; Chao, X. Semi-supervised few-shot learning approach for plant diseases recognition. Plant Methods 2021, 17, 1–10. [Google Scholar] [CrossRef] [PubMed]
- Nuthalapati, S. V.; Tunga, A. Multi-domain few-shot learning and dataset for agricultural applications. In Proceedings of the IEEE/CVF International Conference on Computer Vision; 2021.
- Zhang, J.; Huang, J.; Jin, S.; et al. Vision-language models for vision tasks: A survey. arXiv preprint, arXiv:2304.00685, 2023.
- Bossard, L.; Guillaumin, M.; Van Gool, L. Food-101–mining discriminative components with random forests. In European Conference on Computer Vision, 2014.
- Kiela, D.; Firooz, H.; Mohan, A.; et al. The hateful memes challenge: Detecting hate speech in multimodal memes. In Advances in Neural Information Processing Systems, 2020, pp. 2611–2624.
- Zhou, K.; Yang, J.; Loy, C. C.; et al. Learning to prompt for vision-language models. International Journal of Computer Vision 2022, 130, 2337–2348. [Google Scholar]
- Gao, P.; Geng, S.; Zhang, R.; et al. Clip-adapter: Better vision-language models with feature adapters. International Journal of Computer Vision 2023, 1–15. [Google Scholar] [CrossRef]
- Zhou, K.; Yang, J.; Loy, C. C.; et al. Conditional prompt learning for vision-language models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition; 2022; pp. 16816–16825. [Google Scholar]
- Yao, H.; Zhang, R.; Xu, C. Visual-language prompt tuning with knowledge-guided context optimization. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition; 2023; pp. 6757–6767. [Google Scholar]
- Xu, M.; Park, J. E.; Lee, J.; et al. Plant Disease Recognition Datasets in the Age of Deep Learning: Challenges and Opportunities. arXiv preprint, arXiv:2312.07905, 2023.
- Fuentes, A.; Yoon, S.; Park, D. S. Deep learning-based phenotyping system with glocal description of plant anomalies and symptoms. Frontiers in Plant Science 2019, 10, 1321. [Google Scholar] [CrossRef] [PubMed]
- Wang, C.; Zhou, J.; Zhang, Y.; et al. A plant disease recognition method based on fusion of images and graph structure text. Frontiers in Plant Science 2022, 12, 731688. [Google Scholar] [CrossRef]
- Cao, Y.; Chen, L.; Yuan, Y.; et al. Cucumber disease recognition with small samples using image-text-label-based multi-modal language model. Computers and Electronics in Agriculture 2023, 211, 107993. [Google Scholar] [CrossRef]
- Bender, A.; Whelan, B.; Sukkarieh, S. A high-resolution, multimodal data set for agricultural robotics: A Ladybird’s-eye view of Brassica. Journal of Field Robotics 2020, 37, 73–96. [Google Scholar] [CrossRef]
- Cao, Y.; Chen, L.; Yuan, Y.; et al. Cucumber disease recognition with small samples using image-text-label-based multi-modal language model. Computers and Electronics in Agriculture 2023, 211, 107993. [Google Scholar] [CrossRef]
- Zhou, J.; Li, J.; Wang, C.; et al. Crop disease identification and interpretation method based on multimodal deep learning. Computers and Electronics in Agriculture 2021, 189, 106408. [Google Scholar] [CrossRef]
- Trong, V. H.; Gwang-hyun, Y.; Vu, D. T.; Jin-young, K. Late fusion of multimodal deep neural networks for weeds classification. Computers and Electronics in Agriculture 2020, 175, 105506. [Google Scholar] [CrossRef]
- Lewis, K. M.; Mu, E.; Dalca, A. V.; et al. Gist: Generating image-specific text for fine-grained object classification. arXiv e-prints, arXiv:2307.11315, 2023.
- Ba, J.; Mnih, V.; Kavukcuoglu, K. Multiple object recognition with visual attention. 2020.
- Martins, A. F. T.; Astudillo, R. F. From Softmax to Sparsemax: A Sparse Model of Attention and Multi-Label Classification. In Proceedings of the International Conference on Machine Learning, New York, USA, 2016.
- Yang, Z.; Yang, D.; Dyer, C.; et al. Hierarchical Attention Networks for Document Classification. In Proceedings of the Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, San Diego, USA, 2016.
- Xie, E.; Wang, W.; Yu, Z.; et al. SegFormer: Simple and efficient design for semantic segmentation with transformers. Advances in Neural Information Processing Systems 2021, 34. [Google Scholar]
- Li, K.; Wang, Y.; Zhang, J.; et al. UniFormer: Unifying Convolution and Selfattention for Visual Recognition. arXiv preprint, arXiv:2201.09450, 2022.
- Lei, Z.; Zhang, G.; Wu, L.; et al. A Multi-level Mesh Mutual Attention Model for Visual Question Answering. Data Science and Engineering 2022, 1–15. [Google Scholar] [CrossRef]
- Meinhardt, T.; Kirillov, A.; Leal-Taixe, L.; et al. Trackformer: Multi-object tracking with transformers. arXiv preprint, arXiv:2101.02702, 2021.
- Maniparambil, M.; Vorster, C.; Molloy, D.; et al. Enhancing clip with gpt-4: Harnessing visual descriptions as prompts. In Proceedings of the IEEE/CVF International Conference on Computer Vision; 2023; pp. 262–271. [Google Scholar]
- Vaswani, A.; Shazeer, N.; Parmar, N.; et al. Attention is all you need. In Advances in neural information processing systems. 2017; 30. [Google Scholar]
- Hu, J.; Shen, L.; Sun, G.; et al. Squeeze-and-excitation networks. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Los Alamitos, USA, 2018, pp. 7132–7141.
- Singh, D.; Jain, N.; Jain, P.; et al. PlantDoc:A dataset for visual plant disease detection. In Proceedings of the 7th ACM IKDD CoDS and 25th COMAD; 2020; pp. 249–253. [Google Scholar]
- He, K.; Zhang, X.; Ren, S.; et al. Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition; 2016; pp. 770–778. [Google Scholar]
- Dosovitskiy, A.; Beyer, L.; Kolesnikov, A.; et al. An image is worth 16x16 words: Transformers for image recognition at scale. arXiv preprint, arXiv:2010.11929, 2020.
- Kingma, D. P.; Ba, J. Adam: A method for stochastic optimization. arXiv preprint, arXiv:1412.6980, 2014.
- Dosovitskiy, A.; Beyer, L.; Kolesnikov, A.; et al. An image is worth 16x16 words: Transformers for image recognition at scale. arXiv preprint, arXiv:2010.11929, 2020.
- Selvaraju, R. R.; Cogswell, M.; Das, A.; et al. Grad-cam: Visual explanations from deep networks via gradient-based localization. In Proceedings of the IEEE international conference on computer vision; 2017; pp. 618–626. [Google Scholar]
- Shen, S.; Li, L. H.; Tan, H.; et al. How much can clip benefit vision-and-language tasks? arXiv preprint, arXiv:2107.06383, 2021.
- Radford, A.; Narasimhan, K.; Salimans, T.; et al. Improving language understanding by generative pre-training. Journal Name 2018. [Google Scholar]
- Ouyang, L.; Wu, J.; Jiang, X.; et al. Training language models to follow instructions with human feedback. Advances in Neural Information Processing Systems 2022, 35, 27730–27744. [Google Scholar]









| Disease Category | Number of Images | Disease Category | Number of Images |
| apple black rot | 621 | strawberry leaf scorch | 1109 |
| apple cedar apple rust | 275 | tomato bacterial spot | 2127 |
| apple healthy | 1645 | tomato early blight | 1000 |
| apple scab | 630 | tomato healthy | 1591 |
| blueberry healthy | 1502 | tomato late blight | 1915 |
| cherry healthy | 854 | tomato leaf mold | 952 |
| cherry powdery mildew | 1052 | tomato mosaic virus | 373 |
| corn cercospora leaf spot gray leaf spot | 513 | tomato septoria leaf spot | 1771 |
| corn common rust | 1192 | tomato spider mites | 1676 |
| corn healthy | 1162 | tomato target spot | 1404 |
| corn northern blight | 985 | tomato yellow leaf curl virus | 3357 |
| grape black rot | 1180 | coffee healthy | 282 |
| grape black measles | 1383 | coffee red spider mite | 136 |
| grape healthy | 423 | coffee rust | 282 |
| grape leaf blight | 1076 | cotton diseased | 288 |
| orange huanglongbing | 5507 | cotton healthy | 427 |
| peach bacterial spot | 2297 | cucumber diseased | 227 |
| peach healthy | 360 | cucumber healthy | 241 |
| pepper bell bacterial spot | 997 | lemon diseased | 67 |
| pepper bell healthy | 1491 | lemon healthy | 149 |
| potato early blight | 1000 | mango diseased | 255 |
| potato healthy | 152 | mango healthy | 159 |
| potato late blight | 1000 | pomegranate diseased | 261 |
| raspberry healthy | 371 | pomegranate healthy | 277 |
| soybean healthy | 5090 | rice bacterial leaf blight | 40 |
| squash powdery mildew | 1835 | rice brown spot | 40 |
| strawberry healthy | 456 | rice leaf smut | 40 |
| Few-shot Setup | 1 | 2 | 4 | 8 | 16 |
| Zero-shot CLIP [3]: 13.72 | |||||
| Training-free | |||||
| Tip-Adapter [5] | 25.42 | 38.31 | 53.95 | 67.54 | 73.46 |
| APE [6] | 48.06 | 61.13 | 68.34 | 74.70 | 77.20 |
| VLCD | 49.31 | 62.55 | 69.20 | 75.14 | 77.44 |
| Training-required | |||||
| CoOp [17] | 43.44 | 42.01 | 64.14 | 70.79 | 85.59 |
| KgCoOp [20] | 25.49 | 27.52 | 32.26 | 25.49 | 55.24 |
| CLIP-Adapter [18] | 19.38 | 20.76 | 20.76 | 33.53 | 52.86 |
| Tip-Adapter-F [5] | 37.89 | 42.62 | 56.02 | 75.14 | 83.76 |
| APE-T [6] | 53.43 | 62.18 | 71.38 | 79.68 | 85.99 |
| VLCD-T | 58.16 | 66.07 | 72.32 | 80.95 | 87.31 |
| Few-shot Setup | 0 | 1 | 2 | 4 | 8 | 16 |
| Training-free | ||||||
| Random selection | 21.62 | 48.96 | 61.41 | 67.98 | 75.08 | 77.39 |
| Cluster selection | 22.59 | 49.31 | 62.55 | 69.20 | 75.14 | 77.44 |
| Training-required | ||||||
| Random selection | - | 57.83 | 65.72 | 72.19 | 80.65 | 87.08 |
| Cluster selection | - | 58.16 | 66.07 | 72.32 | 80.95 | 87.31 |
| Set up | Intra-class prompt texts processing | Inter-class prompt texts processing | Accuracy in different few-shot setup | |||||||
| Average | Cross-attention | Without SE | With SE | 0 | 1 | 2 | 4 | 8 | 16 | |
| Training-free | ✓ | - | - | - | 20.80 | 48.06 | 61.13 | 68.34 | 74.70 | 77.20 |
| - | ✓ | - | - | 22.59 | 49.31 | 62.55 | 69.20 | 75.14 | 77.44 | |
| Training-required | ✓ | - | ✓ | - | - | 53.43 | 62.18 | 71.38 | 79.68 | 85.96 |
| - | ✓ | ✓ | - | - | 54.04 | 62.35 | 71.47 | 79.88 | 86.26 | |
| ✓ | - | - | ✓ | - | 56.61 | 65.14 | 72.06 | 79.96 | 86.71 | |
| - | ✓ | - | ✓ | - | 58.16 | 66.07 | 72.32 | 80.95 | 87.31 | |
| Few-shot Setup | 1 | 2 | 4 | 8 | 16 |
| r=2 | 56.36 | 64.57 | 70.02 | 78.58 | 84.91 |
| r=4 | 56.89 | 65.03 | 71.08 | 79.33 | 85.18 |
| r=8 | 58.13 | 66.04 | 70.68 | 78.51 | 84.98 |
| r=16 | 55.56 | 65.20 | 69.97 | 79.90 | 84.91 |
| r=32 | 58.16 | 66.07 | 72.32 | 80.95 | 87.31 |
| Few-shot Setup | 1 | 2 | 4 | 8 | 16 | |
| epoch=10 | without SE | 41.44 | 55.40 | 59.46 | 70.70 | 78.30 |
| with SE | 50.41 | 62.37 | 67.35 | 76.06 | 82.75 | |
| +8.97 | +6.97 | +7.89 | +5.36 | +4.45 | ||
| epoch=20 | without SE | 52.49 | 60.76 | 68.73 | 76.06 | 82.49 |
| with SE | 58.01 | 66.06 | 71.87 | 79.88 | 86.14 | |
| +5.52 | +5.30 | +3.14 | +3.73 | +2.68 | ||
| epoch=30 | without SE | 54.04 | 62.35 | 71.07 | 79.68 | 86.26 |
| with SE | 58.16 | 66.07 | 72.32 | 80.95 | 87.31 | |
| +4.12 | +3.72 | +1.25 | +1.27 | +1.05 | ||
| epoch=40 | without SE | 54.48 | 62.35 | 71.35 | 80.36 | 86.76 |
| with SE | 58.36 | 65.62 | 74.45 | 82.10 | 88.00 | |
| +3.88 | +3.13 | +3.10 | +1.74 | +1.24 | ||
| epoch=50 | without SE | 54.44 | 63.00 | 72.39 | 81.28 | 87.60 |
| with SE | 58.16 | 65.88 | 74.89 | 82.49 | 88.16 | |
| +3.72 | +2.88 | +2.50 | +1.21 | +0.56 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2024 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).