Submitted:
16 December 2024
Posted:
16 December 2024
You are already at the latest version
Abstract
Keywords:
1. Introduction
- We introduce a novel dual-stage training framework for LVLMs that enhances spatial sensitivity using dynamic region masking and prompt-based fine-tuning.
- Our approach leverages adaptive region-aware prompts to guide large language models in generating precise and contextually relevant descriptions for localized regions.
- We demonstrate significant improvements in region understanding on multiple datasets, achieving state-of-the-art results in fine-grained tasks such as region description and spatial reasoning.
2. Related Work
2.1. Large Vision-Language Models
2.2. Vision Region Understanding
3. Method
3.1. Dual-Stage Framework Overview
3.2. Dynamic Region Masking with Vision Encoder
3.3. Adaptive Prompt-Based Fine-Tuning with Language Model
3.4. Training
- : Ensures that the model focuses on the correct regions by penalizing incorrect attention maps.
- : A cross-entropy loss between the generated and ground truth captions.
- : A self-supervised consistency loss to ensure that similar regions produce consistent descriptions.
3.5. Inference
4. Experiments
4.1. Experimental Setup
- CLIP-based Model: A strong baseline model that uses global image features.
- BLIP: A transformer-based model designed for vision-language pre-training.
- Region-VLP: A model focusing on region-aware pre-training.
4.2. Quantitative Results
4.3. Ablation Study
4.4. Human Evaluation
4.5. Analysis and Discussion
5. Conclusion
References
- Radford, A.; Kim, J.W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; Mishkin, P.; Clark, J.; Krueger, G.; Sutskever, I. Learning Transferable Visual Models From Natural Language Supervision. Proceedings of the 38th International Conference on Machine Learning, ICML 2021, 18-24 July 2021, Virtual Event; Meila, M.; Zhang, T., Eds. PMLR, 2021, Vol. 139, Proceedings of Machine Learning Research, pp. 8748–8763.
- Li, J.; Li, D.; Xiong, C.; Hoi, S.C.H. BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation. International Conference on Machine Learning, ICML 2022, 17-23 July 2022, Baltimore, Maryland, USA; Chaudhuri, K.; Jegelka, S.; Song, L.; Szepesvári, C.; Niu, G.; Sabato, S., Eds. PMLR, 2022, Vol. 162, Proceedings of Machine Learning Research, pp. 12888–12900.
- Zhou, Y.; Shen, T.; Geng, X.; Tao, C.; Xu, C.; Long, G.; Jiao, B.; Jiang, D. Towards Robust Ranker for Text Retrieval. Findings of the Association for Computational Linguistics: ACL 2023, 2023, pp. 5387–5401. [Google Scholar]
- Zhou, Y.; Shen, T.; Geng, X.; Tao, C.; Shen, J.; Long, G.; Xu, C.; Jiang, D. Fine-grained distillation for long document retrieval. Proceedings of the AAAI Conference on Artificial Intelligence, 2024, Vol. 38, pp. 19732–19740.
- Zhou, Y.; Long, G. Style-Aware Contrastive Learning for Multi-Style Image Captioning. Findings of the Association for Computational Linguistics: EACL 2023, 2023, pp. 2257–2267. [Google Scholar]
- Yu, Y.Q.; Liao, M.; Zhang, J.; Wu, J. TextHawk2: A Large Vision-Language Model Excels in Bilingual OCR and Grounding with 16x Fewer Tokens. arXiv preprint, arXiv:2410.05261 2024.
- Lin, B.; Tang, Z.; Ye, Y.; Cui, J.; Zhu, B.; Jin, P.; Zhang, J.; Ning, M.; Yuan, L. MoE-LLaVA: Mixture of Experts for Large Vision-Language Models. CoRR, 2024; abs/2401.15947, abs/2401.15947. [Google Scholar] [CrossRef]
- Yu, R.; Yu, W.; Wang, X. Attention Prompting on Image for Large Vision-Language Models. Computer Vision - ECCV 2024 - 18th European Conference, Milan, Italy, September 29-October 4, 2024, Proceedings, Part XXX; Leonardis, A.; Ricci, E.; Roth, S.; Russakovsky, O.; Sattler, T.; Varol, G., Eds. Springer, 2024, Vol. 15088, Lecture Notes in Computer Science, pp. 251–268. [CrossRef]
- Zhou, Y.; Li, X.; Wang, Q.; Shen, J. Visual In-Context Learning for Large Vision-Language Models. Findings of the Association for Computational Linguistics, ACL 2024, Bangkok, Thailand and virtual meeting, August 11-16, 2024. Association for Computational Linguistics, 2024, pp. 15890–15902.
- Zhou, Y.; Rao, Z.; Wan, J.; Shen, J. Rethinking Visual Dependency in Long-Context Reasoning for Large Vision-Language Models. arXiv preprint, arXiv:2410.19732 2024.
- Zhou, Y.; Geng, X.; Shen, T.; Tao, C.; Long, G.; Lou, J.G.; Shen, J. Thread of thought unraveling chaotic contexts. arXiv preprint, arXiv:2311.08734 2023.
- Qi, J.; Ding, M.; Wang, W.; Bai, Y.; Lv, Q.; Hong, W.; Xu, B.; Hou, L.; Li, J.; Dong, Y.; Tang, J. CogCoM: Train Large Vision-Language Models Diving into Details through Chain of Manipulations. CoRR, 2024; abs/2402.04236, [2402.04236]. [Google Scholar] [CrossRef]
- Chen, L.; Li, J.; Dong, X.; Zhang, P.; Zang, Y.; Chen, Z.; Duan, H.; Wang, J.; Qiao, Y.; Lin, D.; Zhao, F. Are We on the Right Way for Evaluating Large Vision-Language Models? CoRR, 2024; abs/2403.20330, [2403.20330]. [Google Scholar] [CrossRef]
- Zhu, T.; Liu, Q.; Wang, F.; Tu, Z.; Chen, M. Unraveling Cross-Modality Knowledge Conflicts in Large Vision-Language Models. CoRR, 2024; abs/2410.03659, [2410.03659]. [Google Scholar] [CrossRef]
- Deng, Y.; Lu, P.; Yin, F.; Hu, Z.; Shen, S.; Zou, J.; Chang, K.; Wang, W. Enhancing Large Vision Language Models with Self-Training on Image Comprehension. CoRR, 2024; abs/2405.19716, [2405.19716]. [Google Scholar] [CrossRef]
- Wu, J.; Zhong, M.; Xing, S.; Lai, Z.; Liu, Z.; Wang, W.; Chen, Z.; Zhu, X.; Lu, L.; Lu, T.; Luo, P.; Qiao, Y.; Dai, J. VisionLLM v2: An End-to-End Generalist Multimodal Large Language Model for Hundreds of Vision-Language Tasks. CoRR 2024, abs/2406.08394, [2406.08394]. [Google Scholar] [CrossRef]
- Zhou, Y.; Long, G. Improving Cross-modal Alignment for Text-Guided Image Inpainting. Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics, 2023, pp. 3445–3456.
- Zhou, Y.; Long, G. Multimodal Event Transformer for Image-guided Story Ending Generation. Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics, 2023, pp. 3434–3444.
- Guo, Q.; Mello, S.D.; Yin, H.; Byeon, W.; Cheung, K.C.; Yu, Y.; Luo, P.; Liu, S. RegionGPT: Towards Region Understanding Vision Language Model. IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2024, Seattle, WA, USA, June 16-22, 2024. IEEE, 2024, pp. 13796–13806. [CrossRef]
- Cheng, A.; Yin, H.; Fu, Y.; Guo, Q.; Yang, R.; Kautz, J.; Wang, X.; Liu, S. SpatialRGPT: Grounded Spatial Reasoning in Vision Language Model. CoRR, 2024; abs/2406.01584, [2406.01584]. [Google Scholar] [CrossRef]
- Chen, C.; Panda, R.; Fan, Q. RegionViT: Regional-to-Local Attention for Vision Transformers. The Tenth International Conference on Learning Representations, ICLR 2022, Virtual Event, April 25-29, 2022. OpenReview.net, 2022.
- Zhang, S.; Sun, P.; Chen, S.; Xiao, M.; Shao, W.; Zhang, W.; Liu, Y.; Chen, K.; Luo, P. Gpt4roi: Instruction tuning large language model on region-of-interest. arXiv preprint, arXiv:2307.03601 2023.
- Ma, C.; Jiang, Y.; Wu, J.; Yuan, Z.; Qi, X. Groma: Localized Visual Tokenization for Grounding Multimodal Large Language Models. Computer Vision - ECCV 2024 - 18th European Conference, Milan, Italy, September 29-October 4, 2024, Proceedings, Part VI; Leonardis, A.; Ricci, E.; Roth, S.; Russakovsky, O.; Sattler, T.; Varol, G., Eds. Springer, 2024, Vol. 15064, Lecture Notes in Computer Science, pp. 417–435. [CrossRef]
- Zhou, Y.; Geng, X.; Shen, T.; Zhang, W.; Jiang, D. Improving zero-shot cross-lingual transfer for multilingual question answering over knowledge graph. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, 2021, pp. 5822–5834.
- Zhou, Y.; Geng, X.; Shen, T.; Pei, J.; Zhang, W.; Jiang, D. Modeling event-pair relations in external knowledge graphs for script reasoning. Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021, 2021. [Google Scholar]
- Zhang, H.; Zhang, P.; Hu, X.; Chen, Y.; Li, L.H.; Dai, X.; Wang, L.; Yuan, L.; Hwang, J.; Gao, J. GLIPv2: Unifying Localization and Vision-Language Understanding. Advances in Neural Information Processing Systems 35: Annual Conference on Neural Information Processing Systems 2022, NeurIPS 2022, New Orleans, LA, USA, November 28 - December 9, 2022; Koyejo, S.; Mohamed, S.; Agarwal, A.; Belgrave, D.; Cho, K.; Oh, A., Eds., 2022.
- Bendale, A.; Boult, T. Towards open world recognition. Proceedings of the IEEE conference on computer vision and pattern recognition, 2015, pp. 1893–1902.
- Chen, C.; Qin, R.; Luo, F.; Mi, X.; Li, P.; Sun, M.; Liu, Y. Position-Enhanced Visual Instruction Tuning for Multimodal Large Language Models. CoRR, 2023; abs/2308.13437, [2308.13437]. [Google Scholar] [CrossRef]
| Model | Accuracy (%) | F1 Score | mIoU |
|---|---|---|---|
| CLIP-based Model | 74.5 | 0.68 | 0.55 |
| BLIP | 76.3 | 0.72 | 0.57 |
| Region-VLP | 78.2 | 0.75 | 0.60 |
| FineRegion-LM (Ours) | 82.1 | 0.79 | 0.64 |
| Model Variation | Accuracy (%) | F1 Score | mIoU |
|---|---|---|---|
| No Dynamic Masking | 78.3 | 0.74 | 0.61 |
| No Adaptive Prompts | 79.0 | 0.76 | 0.62 |
| Full Model (FineRegion-LM) | 82.1 | 0.79 | 0.64 |
| Model | Accuracy | Relevance | Fluency |
|---|---|---|---|
| CLIP-based Model | 3.8 | 3.7 | 4.0 |
| BLIP | 4.0 | 3.9 | 4.2 |
| Region-VLP | 4.3 | 4.2 | 4.4 |
| FineRegion-LM (Ours) | 4.7 | 4.5 | 4.8 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2024 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).