Submitted:
18 July 2026
Posted:
21 July 2026
You are already at the latest version
Abstract
Keywords:
1. Introduction
| What common process underlies MLLMs post-training, and how does it steer pretrained MLLMs toward desired multimodal behaviors? |
- We conceptualize MLLMs post-training as a process of multimodal behavior shaping, offering a unified perspective on how pretrained MLLMs acquire reliable and versatile behavioral capabilities.
- We present a taxonomy that organizes MLLMs post-training methods into five major families: multimodal instruction following, preference calibration, reasoning enhancement, domain adaptation, and scalable training.
- We systematically examine the datasets and evaluations used in MLLMs post-training, revealing how benchmarks and metrics define and measure desirable multimodal behavior.
- We identify key challenges and outline promising research directions, offering a roadmap for advancing MLLMs post-training toward dependable multimodal intelligence.
2. Overview of Post-Training for MLLMs
2.1. What Is MLLMs Post-Training?
2.2. Why MLLMs Post-Training: Progressive Behavior Shaping
2.3. Post-Training MLLMs as the Next Frontier
3. MMPoT of Instruction Following
3.1. Preliminary
3.2. Visual Instruction Tuning
3.3. Instruction Data Mixtures
4. MMPoT of Preference Calibration
4.1. Multimodal Reinforcement Learning with Human Feedback
4.1.1. Reward Mechanisms
4.1.2. Policy Learning of RLHF
4.2. Multimodal Reinforcement Learning with AI Feedback
4.2.1. RLHF vs. RLAIF
4.2.2. RLAIF Training Pipeline
4.3. Multimodal Direct Preference Optimization
4.3.1. Preliminary
4.3.2. The Evolution of Multimodal DPO
5. MMPoT of Reason Enhancement
5.1. R1-Based Multimodal Reasoning
5.1.1. From LLM-R1 to MLLM-R1
5.1.2. R1 Training Paradigm for MLLMs
5.2. Thinking with Images
5.3. Self-Evolution for Multimodal Reasoning
5.3.1. Self-Generated Data Learning
5.3.2. Reflection and Critique-Based Learning
5.3.3. Verifier-Guided Self-Improvement
5.4. Efficient Reasoning
5.4.1. Knowledge Distillation
5.4.2. On-Policy Distillation
6. MMPoT of Domain Adaptation
7. MMPoT of Scalable Training
7.1. Parameter-Efficient Post-Training
7.1.1. Low-Rank Adaptation
7.1.2. Mixture-of-Experts Adaptation
7.2. Compute-Efficient Post-Training
7.2.1. Efficient Visual Processing
7.2.2. Token Compression
7.2.3. Long-Context Optimization
8. MLLMs Post-Training Benchmarks
8.1. Datasets and Benchmarks
8.2. Evaluation Metrics
8.2.1. Reference-Based Metrics
8.2.2. Judge-Based Metrics
9. Future Directions
9.1. Grounded Behavior Shaping
9.2. Reliability-Aware Evaluation
9.3. Scaling for Generalization
10. Conclusion
References
- Liu, H.; Li, C.; Wu, Q.; Lee, Y.J. Visual instruction tuning. NeurIPS 2023, 34892–34916. [Google Scholar] [CrossRef]
- Dai, W.; Li, J.; Li, D.; Tiong, A.; Zhao, J.; Wang, W.; Li, B.; Fung, P.N.; Hoi, S. Instructblip: Towards general-purpose vision-language models with instruction tuning. NeurIPS 2023, 49250–49267. [Google Scholar] [CrossRef]
- Zhang, H.; Zeng, P.; Gao, L.; Song, J.; Duan, Y.; Lyu, X.; Shen, H.T. Text-video retrieval with global-local semantic consistent learning. TIP; 2025. [Google Scholar]
- Liu, Y.; Zhang, Y.; Cai, J.; Jiang, X.; Hu, Y.; Yao, J.; Wang, Y.; Xie, W. Lamra: Large multimodal model as your advanced retrieval assistant. In Proceedings of the CVPR, 2025. [Google Scholar]
- Jia, F.; Mao, W.; Liu, Y.; Zhao, Y.; Wen, Y.; Zhang, C.; Zhang, X.; Wang, T. Adriver-i: A general world model for autonomous driving. arXiv 2023, arXiv:2311.13549. [Google Scholar]
- Xu, Z.; Zhang, Y.; Xie, E.; Zhao, Z.; Guo, Y.; Wong, K.Y.K.; Li, Z.; Zhao, H. Drivegpt4: Interpretable end-to-end autonomous driving via large language model. RA-L 2024, 8186–8193. [Google Scholar] [CrossRef]
- Driess, D.; Xia, F.; Sajjadi, M.S.; Lynch, C.; Chowdhery, A.; Ichter, B.; Wahid, A.; Tompson, J.; Vuong, Q.; Yu, T.; et al. PaLM-E: an embodied multimodal language model. In Proceedings of the ICML, 2023; pp. 8469–8488. [Google Scholar]
- Kim, M.J.; Pertsch, K.; Karamcheti, S.; Xiao, T.; Balakrishna, A.; Nair, S.; Rafailov, R.; Foster, E.P.; Sanketi, P.R.; Vuong, Q.; et al. OpenVLA: An Open-Source Vision-Language-Action Model. In Proceedings of the CoRL.
- Bai, J.; Bai, S.; Chu, Y.; Cui, Z.; Dang, K.; Deng, X.; Fan, Y.; Ge, W.; Han, Y.; Huang, F.; et al. Qwen technical report. arXiv 2023, arXiv:2309.16609. [Google Scholar]
- Touvron, H.; Lavril, T.; Izacard, G.; Martinet, X.; Lachaux, M.A.; Lacroix, T.; Rozière, B.; Goyal, N.; Hambro, E.; Azhar, F.; et al. Llama: Open and efficient foundation language models. arXiv 2023, arXiv:2302.13971. [Google Scholar]
- Khayatkhoei, M.; Chhikara, P.; Ilievski, F.; et al. Mllms know where to look: Training-free perception of small visual details with multimodal llms. In Proceedings of the ICLR, 2025. [Google Scholar]
- Marino, K.; Rastegari, M.; Farhadi, A.; Mottaghi, R. Ok-vqa: A visual question answering benchmark requiring external knowledge. In Proceedings of the CVPR, 2019. [Google Scholar]
- Sarto, S.; Cornia, M.; Cucchiara, R. Image captioning evaluation in the age of multimodal llms: Challenges and future perspectives. arXiv 2025, arXiv:2503.14604. [Google Scholar]
- Chen, X.; Shukla, S.N.; Azab, M.; Singh, A.; Wang, Q.; Yang, D.; Peng, S.; Yu, H.; Yan, S.; Zhang, X.; et al. Compcap: Improving multimodal large language models with composite captions. In Proceedings of the ICCV, 2025. [Google Scholar]
- Wang, H.; Hu, K.; Gao, L. Docvideoqa: Towards comprehensive understanding of document-centric videos through question answering. In Proceedings of the ICASSP, 2025. [Google Scholar]
- Zhang, J.; Fan, Q.; Zhang, Y. DocAssistant: Integrating Key-region Reading and Step-wise Reasoning for Robust Document Visual Question Answering. In Proceedings of the EMNLP, 2025; pp. 3496–3511. [Google Scholar]
- Dong, Y.; Liu, Z.; Sun, H.L.; Yang, J.; Hu, W.; Rao, Y.; Liu, Z. Insight-v: Exploring long-chain visual reasoning with multimodal large language models. In Proceedings of the CVPR, 2025; pp. 9062–9072. [Google Scholar]
- Yang, S.; Niu, Y.; Liu, Y.; Ye, Y.; Lin, B.; Yuan, L. Look-back: Implicit visual re-focusing in mllm reasoning. In Proceedings of the AAAI, 2026; pp. 11694–11702. [Google Scholar]
- Zhang, H.; Luo, R.; Liu, X.; Wu, Y.; Lin, T.E.; Zeng, P.; Qu, Q.; Fang, F.; Yang, M.; Gao, L.; et al. Omnicharacter: Towards immersive role-playing agents with seamless speech-language personality interaction. In Proceedings of the ACL, 2025; pp. 26318–26331. [Google Scholar]
- Saab, K.; Tu, T.; Weng, W.H.; Tanno, R.; Stutz, D.; Wulczyn, E.; Zhang, F.; Strother, T.; Park, C.; Vedadi, E.; et al. Capabilities of gemini models in medicine. arXiv 2024, arXiv:2404.18416. [Google Scholar]
- Zhang, H.; Zeng, P.; Zhang, J.; Song, J.; Sebe, N.; Shen, H.T.; Gao, L. OmniCharacter++: Towards Comprehensive Benchmark for Realistic Role-Playing Agents. TPAMI, 2026. [Google Scholar]
- Wang, P.; Yang, A.; Men, R.; Lin, J.; Bai, S.; Li, Z.; Ma, J.; Zhou, C.; Zhou, J.; Yang, H. Ofa: Unifying architectures, tasks, and modalities through a simple sequence-to-sequence learning framework. In Proceedings of the ICML, 2022; pp. 23318–23340. [Google Scholar]
- Li, J.; Jiang, L.; Zhang, H.; Sebe, N. Token reduction via local and global contexts optimization for efficient video large language models. In Proceedings of the CVPR, 2026; pp. 10451–10461. [Google Scholar]
- Zhu, D.; Shen, X.; Li, X.; Elhoseiny, M.; et al. Minigpt-4: Enhancing vision-language understanding with advanced large language models. In Proceedings of the ICLR, 2024. [Google Scholar]
- Liu, H.; Li, C.; Li, Y.; Lee, Y.J. Improved baselines with visual instruction tuning. In Proceedings of the CVPR, 2024. [Google Scholar]
- Stephan, M.; Khazatsky, A.; Mitchell, E.; Chen, A.S.; Hsu, S.; Sharma, A.; Finn, C. RLVF: learning from verbal feedback without overgeneralization. In Proceedings of the ICML, 2024; pp. 46625–46656. [Google Scholar]
- Luo, J.; Dong, P.; Zhai, Y.; Ma, Y.; Levine, S. Rlif: Interactive imitation learning as reinforcement learning. In Proceedings of the ICLR, 2024; pp. 36329–36351. [Google Scholar]
- Rafailov, R.; Sharma, A.; Mitchell, E.; Manning, C.D.; Ermon, S.; Finn, C. Direct preference optimization: Your language model is secretly a reward model. NeurIPS 2023, 53728–53741. [Google Scholar] [CrossRef]
- Shao, Z.; Wang, P.; Zhu, Q.; Xu, R.; Song, J.; Bi, X.; Zhang, H.; Zhang, M.; Li, Y.; Wu, Y.; et al. Deepseekmath: Pushing the limits of mathematical reasoning in open language models. arXiv 2024, arXiv:2402.03300. [Google Scholar]
- Yu, Q.; Zhang, Z.; Zhu, R.; Yuan, Y.; Zuo, X.; Yue, Y.; Dai, W.; Fan, T.; Liu, G.; Liu, L.; et al. Dapo: An open-source llm reinforcement learning system at scale. NeurIPS 2026, 113222–113244. [Google Scholar]
- Jaech, A.; Kalai, A.; Lerer, A.; Richardson, A.; El-Kishky, A.; Low, A.; Helyar, A.; Madry, A.; Beutel, A.; Carney, A.; et al. Openai o1 system card. arXiv 2024, arXiv:2412.16720. [Google Scholar]
- Guo, D.; Yang, D.; Zhang, H.; Song, J.; Wang, P.; Zhu, Q.; Xu, R.; Zhang, R.; Ma, S.; Bi, X.; et al. Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning. arXiv 2025, arXiv:2501.12948. [Google Scholar]
- Yu, T.; Zhang, Y.F.; Fu, C.; Wu, J.; Lu, J.; Wang, K.; Lu, X.; Shen, Y.; Zhang, G.; Song, D.; et al. Aligning multimodal llm with human preference: A survey. arXiv 2025, arXiv:2503.14504. [Google Scholar]
- Zhou, G.; Qiu, P.; Chen, C.; Wang, J.; Yang, Z.; Xu, J.; Qiu, M. Reinforced mllm: A survey on rl-based reasoning in multimodal large language models. arXiv 2025, arXiv:2504.21277. [Google Scholar]
- Zhang, S.; Dong, L.; Li, X.; Zhang, S.; Sun, X.; Wang, S.; Li, J.; Hu, R.; Zhang, T.; Wang, G.; et al. Instruction tuning for large language models: A survey. ACM Comput. Surv. 2026, 1–36. [Google Scholar] [CrossRef]
- Sun, Z.; Shen, S.; Cao, S.; Liu, H.; Li, C.; Shen, Y.; Gan, C.; Gui, L.; Wang, Y.X.; Yang, Y.; et al. Aligning large multimodal models with factually augmented rlhf. In Proceedings of the ACL, Finding’g’s, 2024; pp. 13088–13110. [Google Scholar]
- Yu, T.; Yao, Y.; Zhang, H.; He, T.; Han, Y.; Cui, G.; Hu, J.; Liu, Z.; Zheng, H.T.; Sun, M.; et al. Rlhf-v: Towards trustworthy mllms via behavior alignment from fine-grained correctional human feedback. In Proceedings of the CVPR, 2024. [Google Scholar]
- Ouali, Y.; Bulat, A.; Martinez, B.; Tzimiropoulos, G. Clip-dpo: Vision-language models as a source of preference for fixing hallucinations in lvlms. In Proceedings of the ECCV, 2024; pp. 395–413. [Google Scholar]
- Huang, W.; Jia, B.; Zhai, Z.; Cao, S.; Ye, Z.; Zhao, F.; Xu, Z.; Tang, X.; Hu, Y.; Lin, S. Vision-r1: Incentivizing reasoning capability in multimodal large language models. arXiv 2025, arXiv:2503.06749. [Google Scholar]
- Hou, W.; Peng, S.; Wang, W.; Ruan, Z.; Zhang, Y.; Zhou, Z.; Gao, M.; Chen, Y.; Wang, K.; Yang, H.; et al. Uni-OPD: Unifying on-policy distillation with a dual-perspective recipe. arXiv 2026, arXiv:2605.03677. [Google Scholar]
- Li, J.; Yin, H.; Xu, H.; Xu, B.; Tan, W.; He, Z.; Ju, J.; Luo, Z.; Luan, J. Video-OPD: Efficient Post-Training of Multimodal Large Language Models for Temporal Video Grounding via On-Policy Distillation. arXiv 2026, arXiv:2602.02994. [Google Scholar]
- Ye, Q.; Xu, H.; Xu, G.; Ye, J.; Yan, M.; Zhou, Y.; Wang, J.; Hu, A.; Shi, P.; Shi, Y.; et al. mplug-owl: Modularization empowers large language models with multimodality. arXiv 2023, arXiv:2304.14178. [Google Scholar]
- Gao, P.; Han, J.; Zhang, R.; Lin, Z.; Geng, S.; Zhou, A.; Zhang, W.; Lu, P.; He, C.; Yue, X.; et al. Llama-adapter v2: Parameter-efficient visual instruction model. arXiv 2023, arXiv:2304.15010. [Google Scholar]
- Wang, W.; Lv, Q.; Yu, W.; Hong, W.; Qi, J.; Wang, Y.; Ji, J.; Yang, Z.; Zhao, L.; Song, X.; et al. Cogvlm: Visual expert for pretrained language models. NeurIPS 2024. [Google Scholar] [CrossRef]
- Liu, H.; Li, C.; Li, Y.; Li, B.; Zhang, Y.; Shen, S.; Lee, Y.J. Llavanext: Improved reasoning, ocr, and world knowledge; 2024. [Google Scholar]
- Wang, P.; Bai, S.; Tan, S.; Wang, S.; Fan, Z.; Bai, J.; Chen, K.; Liu, X.; Wang, J.; Ge, W.; et al. Qwen2-vl: Enhancing vision-language model’s perception of the world at any resolution. arXiv 2024, arXiv:2409.12191. [Google Scholar]
- Li, B.; Zhang, Y.; Guo, D.; Zhang, R.; Li, F.; Zhang, H.; Zhang, K.; Zhang, P.; Li, Y.; Liu, Z.; et al. Llava-onevision: Easy visual task transfer. arXiv 2024, arXiv:2408.03326. [Google Scholar]
- Chen, Z.; Wang, W.; Cao, Y.; Liu, Y.; Gao, Z.; Cui, E.; Zhu, J.; Ye, S.; Tian, H.; Liu, Z.; et al. Expanding performance boundaries of open-source multimodal models with model, data, and test-time scaling. arXiv 2024, arXiv:2412.05271. [Google Scholar]
- Lin, J.; Yin, H.; Ping, W.; Molchanov, P.; Shoeybi, M.; Han, S. Vila: On pre-training for visual language models. In Proceedings of the CVPR, 2024; pp. 26689–26699. [Google Scholar]
- Chu, X.; Qiao, L.; Lin, X.; Xu, S.; Yang, Y.; Hu, Y.; Wei, F.; Zhang, X.; Zhang, B.; Wei, X.; et al. Mobilevlm: A fast, strong and open vision language assistant for mobile devices. arXiv 2023, arXiv:2312.16886. [Google Scholar]
- Lu, H.; Liu, W.; Zhang, B.; Wang, B.; Dong, K.; Liu, B.; Sun, J.; Ren, T.; Li, Z.; Yang, H.; et al. Deepseek-vl: towards real-world vision-language understanding. arXiv 2024, arXiv:2403.05525. [Google Scholar]
- Tong, S.; Brown, E.; Wu, P.; Woo, S.; Middepogu, M.; Akula, S.C.; Yang, J.; Yang, S.; Iyer, A.; Pan, X.; et al. Cambrian-1: A fully open, vision-centric exploration of multimodal llms. NeurIPS 2024, 87310–87356. [Google Scholar] [CrossRef]
- Yao, Y.; Yu, T.; Zhang, A.; Wang, C.; Cui, J.; Zhu, H.; Cai, T.; Li, H.; Zhao, W.; He, Z.; et al. Minicpm-v: A gpt-4v level mllm on your phone. arXiv 2024, arXiv:2408.01800. [Google Scholar]
- Deitke, M.; Clark, C.; Lee, S.; Tripathi, R.; Yang, Y.; Park, J.S.; Salehi, M.; Muennighoff, N.; Lo, K.; Soldaini, L.; et al. Molmo and pixmo: Open weights and open data for state-of-the-art vision-language models. In Proceedings of the CVPR, 2025; pp. 91–104. [Google Scholar]
- Agrawal, P.; Antoniak, S.; Hanna, E.B.; Bout, B.; Chaplot, D.; Chudnovsky, J.; Costa, D.; De Monicault, B.; Garg, S.; Gervet, T.; et al. Pixtral 12B. arXiv 2024, arXiv:2410.07073. [Google Scholar]
- Li, B.; Zhang, Y.; Chen, L.; Wang, J.; Pu, F.; Cahyono, J.A.; Yang, J.; Li, C.; Liu, Z. Otter: A multi-modal model with in-context instruction tuning. TPAMI, 2025. [Google Scholar]
- Luo, G.; Zhou, Y.; Ren, T.; Chen, S.; Sun, X.; Ji, R. Cheap and quick: Efficient vision-language instruction tuning for large language models. NeurIPS 2023, 29615–29627. [Google Scholar] [CrossRef]
- Luo, R.; Zhao, Z.; Yang, M.; Yang, Z.; Qiu, M.; Wei, Z.; Wang, Y.; Chen, C. Valley: Video assistant with large language model enhanced ability. TOMM, 2026. [Google Scholar]
- Wu, S.; Fei, H.; Qu, L.; Ji, W.; Chua, T.S. Next-gpt: Any-to-any multimodal llm. arXiv 2023, arXiv:2309.05519. [Google Scholar]
- Chen, Z.; Wu, J.; Wang, W.; Su, W.; Chen, G.; Xing, S.; Zhong, M.; Zhang, Q.; Zhu, X.; Lu, L.; et al. Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks. In Proceedings of the CVPR, 2024. [Google Scholar]
- Hernandez, J.; Villegas, R.; Ordonez, V. Generative visual instruction tuning. arXiv 2024, arXiv:2406.11262. [Google Scholar]
- Tu, J.; Ni, Z.; Crispino, N.; Yu, Z.; Bendersky, M.; Gunel, B.; Jia, R.; Liu, X.; Lyu, L.; Song, D.; et al. MLAN: Language-Based Instruction Tuning Preserves and Transfers Knowledge in Multimodal Language Models. In Proceedings of the KnowFM, 2025; pp. 59–74. [Google Scholar]
- Peng, W.; Meng, L.; Chen, Y.; Xie, Y.; Liu, Y.; Gui, T.; Xu, H.; Qiu, X.; Wu, Z.; Jiang, Y.G. Inst-it: Boosting instance understanding via explicit visual prompt instruction tuning. NeurIPS 2026, 50062–50092. [Google Scholar]
- Oh, C.; Li, J.; Im, S.; Li, S. Visual instruction bottleneck tuning. NeurIPS 2026, 129164–129204. [Google Scholar]
- You, Z.; Nie, S.; Zhang, X.; ZHOU, J.; Lu, Z.; Wen, J.R.; Li, C. Llada-v: Large language diffusion models with visual instruction tuning. In Proceedings of the CVPR, 2026. [Google Scholar]
- Jiang, D.; He, X.; Zeng, H.; Wei, C.; Ku, M.; Liu, Q.; Chen, W. Mantis: Interleaved multi-image instruction tuning. arXiv 2024, arXiv:2405.01483. [Google Scholar]
- Han, J.; Zhang, R.; Shao, W.; Gao, P.; Xu, P.; Xiao, H.; Zhang, K.; Liu, C.; Wen, S.; Guo, Z.; et al. Imagebind-llm: Multi-modality instruction tuning. arXiv 2023, arXiv:2309.03905. [Google Scholar]
- Chen, L.; Li, J.; Dong, X.; Zhang, P.; He, C.; Wang, J.; Zhao, F.; Lin, D. Sharegpt4v: Improving large multi-modal models with better captions. In Proceedings of the ECCV, 2024; pp. 370–387. [Google Scholar]
- Radford, A.; Kim, J.W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; Mishkin, P.; Clark, J.; et al. Learning transferable visual models from natural language supervision. In Proceedings of the ICML, 2021. [Google Scholar]
- Chiang, W.L.; Li, Z.; Lin, Z.; Sheng, Y.; Wu, Z.; Zhang, H.; Zheng, L.; Zhuang, S.; Zhuang, Y.; Gonzalez, J.E.; et al. Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality. 6, 2023). Available online: https://vicuna. (accessed on 14 April 2023). [PubMed]
- Gong, T.; Lyu, C.; Zhang, S.; Wang, Y.; Zheng, M.; Zhao, Q.; Liu, K.; Zhang, W.; Luo, P.; Chen, K. Multimodal-gpt: A vision and language model for dialogue with humans. arXiv 2023, arXiv:2305.04790. [Google Scholar]
- Li, J.; Li, D.; Xiong, C.; Hoi, S. Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation. In Proceedings of the ICML, 2022; pp. 12888–12900. [Google Scholar]
- Zhang, S.; Sun, P.; Chen, S.; Xiao, M.; Shao, W.; Zhang, W.; Liu, Y.; Chen, K.; Luo, P. Gpt4roi: Instruction tuning large language model on region-of-interest. In Proceedings of the ECCV, 2025; pp. 52–70. [Google Scholar]
- Chen, K.; Zhang, Z.; Zeng, W.; Zhang, R.; Zhu, F.; Zhao, R. Shikra: Unleashing multimodal llm’s referential dialogue magic. arXiv 2023, arXiv:2306.15195. [Google Scholar]
- You, H.; Zhang, H.; Gan, Z.; Du, X.; Zhang, B.; Wang, Z.; Cao, L.; Chang, S.F.; Yang, Y. Ferret: Refer and ground anything anywhere at any granularity. In Proceedings of the ICLR, 2024; pp. 57153–57180. [Google Scholar]
- Zhang, H.; You, H.; Dufter, P.; Zhang, B.; Chen, C.; Chen, H.Y.; Fu, T.J.; Wang, W.Y.; Chang, S.F.; Gan, Z.; et al. Ferret-v2: An improved baseline for referring and grounding with large language models. arXiv 2024, arXiv:2404.07973. [Google Scholar]
- Peng, Z.; Wang, W.; Dong, L.; Hao, Y.; Huang, S.; Ma, S.; Ye, Q.; Wei, F. Grounding multimodal large language models to the world. In Proceedings of the ICLR, 2024. [Google Scholar]
- Zhang, H.; Li, X.; Bing, L. Video-llama: An instruction-tuned audio-visual language model for video understanding. In Proceedings of the EMNLP, 2023; pp. 543–553. [Google Scholar]
- Li, K.; He, Y.; Wang, Y.; Li, Y.; Wang, W.; Luo, P.; Wang, Y.; Wang, L.; Qiao, Y. Videochat: Chat-centric video understanding. Sci. China Inf. Sci. 2025, 200102. [Google Scholar]
- Luo, R.; Zhao, Z.; Yang, M.; Yang, Z.; Qiu, M.; Wei, Z.; Wang, Y.; Chen, C. Valley: Video assistant with large language model enhanced ability. TOMM, 2023. [Google Scholar]
- Bai, S.; Chen, K.; Liu, X.; Wang, J.; Ge, W.; Song, S.; Dang, K.; Wang, P.; Wang, S.; Tang, J.; et al. Qwen2.5-VL Technical Report. arXiv 2025, arXiv:2502.13923. [Google Scholar]
- Zhu, J.; Wang, W.; Chen, Z.; Liu, Z.; Ye, S.; Gu, L.; Tian, H.; Duan, Y.; Su, W.; Shao, J.; et al. Internvl3: Exploring advanced training and test-time recipes for open-source multimodal models. arXiv 2025, arXiv:2504.10479. [Google Scholar]
- Xu, Z.; Shen, Y.; Huang, L. Multiinstruct: Improving multi-modal zero-shot learning via instruction tuning. In Proceedings of the ACL, 2023; pp. 11445–11465. [Google Scholar]
- Wang, W.; Gao, Z.; Chen, L.; Chen, Z.; Zhu, J.; Zhao, X.; Liu, Y.; Cao, Y.; Ye, S.; Zhu, X.; et al. Visualprm: An effective process reward model for multimodal reasoning. arXiv 2025, arXiv:2503.10291. [Google Scholar]
- Team, G.; Anil, R.; Borgeaud, S.; Alayrac, J.B.; Yu, J.; Soricut, R.; Schalkwyk, J.; Dai, A.M.; Hauth, A.; Millican, K.; et al. Gemini: a family of highly capable multimodal models. arXiv 2023, arXiv:2312.11805. [Google Scholar]
- Guo, D.; Wu, F.; Zhu, F.; Leng, F.; Shi, G.; Chen, H.; Fan, H.; Wang, J.; Jiang, J.; Wang, J.; et al. Seed1. 5-vl technical report. arXiv 2025, arXiv:2505.07062. [Google Scholar]
- Yue, Z.; Lin, Z.; Song, Y.; Wang, W.; Ren, S.; Gu, S.; Li, S.; Li, P.; Zhao, L.; Li, L.; et al. MiMo-VL technical report. arXiv 2025, arXiv:2506.03569. [Google Scholar]
- Ahn, D.; Choi, Y.; Yu, Y.; Kang, D.; Choi, J. Tuning large multimodal models for videos using reinforcement learning from ai feedback. In Proceedings of the ACL, 2024. [Google Scholar]
- Yu, T.; Zhang, H.; Li, Q.; Xu, Q.; Yao, Y.; Chen, D.; Lu, X.; Cui, G.; Dang, Y.; He, T.; et al. Rlaif-v: Open-source ai feedback leads to super gpt-4v trustworthiness. Proc. CVPR 2025, 19985–19995. [Google Scholar] [CrossRef]
- Shi, D.; Glatt, R.; Klymko, C.; Mohole, S.; Choi, H.; Kushwaha, S.; Sakla, S.; da Silva, F.L. Oracle-RLAIF: An Improved Fine-Tuning Framework for Multi-modal Video Models through Reinforcement Learning from Ranking Feedback. arXiv 2025, arXiv:2510.02561. [Google Scholar]
- Zhang, R.; Gui, L.; Sun, Z.; Feng, Y.; Xu, K.; Zhang, Y.; Fu, D.; Li, C.; Hauptmann, A.G.; Bisk, Y.; et al. Direct preference optimization of video large multimodal models from language model reward. In Proceedings of the ACL, 2025. [Google Scholar]
- Zhao, Z.; Wang, B.; Ouyang, L.; Dong, X.; Wang, J.; He, C. Beyond hallucinations: Enhancing lvlms through hallucination-aware direct preference optimization. arXiv 2023, arXiv:2311.16839. [Google Scholar]
- Xie, Y.; Li, G.; Xu, X.; Kan, M.Y. V-dpo: Mitigating hallucination in large vision language models via vision-guided direct preference optimization. In Proceedings of the EMNLP, 2024; pp. 13258–13273. [Google Scholar]
- Compagnoni, A.; Caffagni, D.; Moratelli, N.; Baraldi, L.; Cornia, M.; Cucchiara, R.; et al. Mitigating Hallucinations in Multimodal LLMs via Object-aware Preference Optimization. In Proceedings of the BMVC, 2025. [Google Scholar]
- Wang, F.; Zhou, W.; Huang, J.Y.; Xu, N.; Zhang, S.; Poon, H.; Chen, M. mdpo: Conditional preference optimization for multimodal large language models. In Proceedings of the EMNLP, 2024; pp. 8078–8088. [Google Scholar]
- Chaubey, A.; Pang, J.; Soleymani, M. MoD-DPO: Towards Mitigating Cross-modal Hallucinations in Omni LLMs using Modality Decoupled Preference Optimization. Proc. CVPR 2026, 18284–18294. [Google Scholar]
- Chen, J.; Zhang, T.; Huang, S.; Niu, Y.; Sun, C.; Zhang, R.; Zhou, G.; Wen, L. OmniDPO: A Preference Optimization Framework to Address Omni-Modal Hallucination. Proc. AAAI 2026, 24, 20172–20180. [Google Scholar] [CrossRef]
- Qiu, L.; Ning, S.; Zhang, C.; Sun, J.; He, X. DA-DPO: Cost-efficient Difficulty-aware Preference Optimization for Reducing MLLM Hallucinations. arXiv 2026, arXiv:2601.00623. [Google Scholar]
- Zadeh, F.P.; Oh, Y.; Kim, G. Lpoi: Listwise preference optimization for vision language models. In Proceedings of the ACL, 2025; pp. 26830–26844. [Google Scholar]
- Zhang, Y.; Yu, T.; Tian, H.; Fu, C.; Li, P.; Zeng, J.; Xie, W.; Shi, Y.; Zhang, H.; Wu, J.; et al. MM-RLHF: The Next Step Forward in Multimodal LLM Alignment. In Proceedings of the ICML, 2025. [Google Scholar]
- Xiaomi, L.C.T. MiMo-VL Technical Report. arXiv 2025, arXiv:cs. [Google Scholar]
- Shen, H.; Liu, P.; Li, J.; Fang, C.; Ma, Y.; Liao, J.; Shen, Q.; Zhang, Z.; Zhao, K.; Zhang, Q.; et al. Vlm-r1: A stable and generalizable r1-style large vision-language model. arXiv 2025, arXiv:2504.07615. [Google Scholar]
- Liu, Z.; Sun, Z.; Zang, Y.; Dong, X.; Cao, Y.; Duan, H.; Lin, D.; Wang, J. Visual-rft: Visual reinforcement fine-tuning. Proc. ICCV 2025, 2034–2044. [Google Scholar] [CrossRef]
- Feng, K.; Gong, K.; Li, B.; Guo, Z.; Wang, Y.; Peng, T.; Wu, J.; Zhang, X.; Wang, B.; Yue, X. Video-r1: Reinforcing video reasoning in mllms. NeurIPS 2026. [Google Scholar] [CrossRef]
- Zhou, H.; Li, X.; Wang, R.; Cheng, M.; Zhou, T.; Hsieh, C.J. R1-Zero’s" Aha Moment" in Visual Reasoning on a 2B Non-SFT Model. arXiv 2025, arXiv:2503.05132. [Google Scholar]
- Meng, F.; Du, L.; Liu, Z.; Zhou, Z.; Lu, Q.; Fu, D.; Han, T.; Shi, B.; Wang, W.; He, J.; et al. Mm-eureka: Exploring the frontiers of multimodal reasoning with rule-based reinforcement learning. arXiv 2025, arXiv:2503.07365. [Google Scholar]
- Yang, Y.; He, X.; Pan, H.; Jiang, X.; Deng, Y.; Yang, X.; Lu, H.; Yin, D.; Rao, F.; Zhu, M.; et al. R1-onevision: Advancing generalized multimodal reasoning through cross-modal formalization. In Proceedings of the ICCV, 2025. [Google Scholar]
- Zhao, J.; Wei, X.; Bo, L. R1-omni: Explainable omni-multimodal emotion recognition with reinforcement learning. arXiv 2025, arXiv:2503.05379. [Google Scholar]
- Zhu, L.; Ji, D.; Chen, T.; Wu, H.; Wang, S. Retrv-r1: A reasoning-driven mllm framework for universal and efficient multimodal retrieval. NeurIPS, 2026. [Google Scholar]
- Fan, Y.; He, X.; Yang, D.; Zheng, K.; Kuo, C.C.; Zheng, Y.; Guan, X.; Wang, X. Grit: Teaching mllms to think with images. NeurIPS 2026, 116522–116543. [Google Scholar]
- Ni, M.; Yang, Z.; Li, L.; Lin, C.C.; Lin, K.; Zuo, W.; Wang, L. Point-rft: Improving multimodal reasoning with visually grounded reinforcement finetuning. NeurIPS 2026, 20538–20559. [Google Scholar]
- Su, Z.; Li, L.; Song, M.; Hao, Y.; Yang, Z.; Zhang, J.; Chen, G.; Gu, J.; Li, J.; Qu, X.; et al. Openthinkimg: Learning to think with images via visual tool reinforcement learning. arXiv 2025, arXiv:2505.08617. [Google Scholar]
- Liu, Y.; Qu, T.; Zhong, Z.; Peng, B.; Liu, S.; Yu, B.; Jia, J. VisionReasoner: Unified reasoning-integrated visual perception via reinforcement learning. arXiv 2025, arXiv:2505.12081. [Google Scholar]
- Zheng, Z.; Yang, M.; Hong, J.; Zhao, C.; Xu, G.; Yang, L.; Shen, C.; Yu, X. Deepeyes: Incentivizing" thinking with images" via reinforcement learning. arXiv 2025, arXiv:2505.14362. [Google Scholar]
- Wu, M.; Yang, J.; Jiang, J.; Li, M.; Yan, K.; Yu, H.; Zhang, M.; Zhai, C.; Nahrstedt, K. Vtool-r1: Vlms learn to think with images via reinforcement learning on multimodal tool use. arXiv 2025, arXiv:2505.19255. [Google Scholar]
- Viveiros, A.G.; Gonçalves, N.; Lindemann, M.; Martins, A. LanteRn: Latent Visual Structured Reasoning. arXiv 2026, arXiv:2603.25629. [Google Scholar]
- Wang, B.; Wu, F.; Han, X.; Peng, J.; Zhong, H.; Zhang, P.; Dong, X.; Li, W.; Li, W.; Wang, J.; et al. Vigc: Visual instruction generation and correction. Proc. AAAI 2024, 6, 5309–5317. [Google Scholar] [CrossRef]
- Xu, Z.; Chen, D.; Ling, Z.; Li, Y.; Shen, Y. MindGYM: What Matters in Question Synthesis for Thinking-Centric Fine-Tuning? NeurIPS 2026. [Google Scholar] [CrossRef]
- Wan, Z.; Dou, Z.; Liu, C.; Zhang, Y.; Cui, D.; Zhao, Q.; Shen, H.; Xiong, J.; Xin, Y.; Jiang, Y.; et al. Srpo: Enhancing multimodal llm reasoning via reflection-aware reinforcement learning. NeurIPS 2026. [Google Scholar] [CrossRef]
- Xiong, T.; Wang, X.; Guo, D.; Ye, Q.; Fan, H.; Gu, Q.; Huang, H.; Li, C. Llava-critic: Learning to evaluate multimodal models. In Proceedings of the CVPR, 2025. [Google Scholar]
- Wei, L.; Li, Y.; Wang, C.; Wang, Y.; Kong, L.; Huang, W.; Sun, L. First SFT, Second RL, Third UPT: Continual Improving Multi-Modal LLM Reasoning via Unsupervised Post-Training. NeurIPS 2026, 62293–62318. [Google Scholar]
- Wang, X.; Li, C.; Yang, J.; Zhang, K.; Liu, B.; Xiong, T.; Huang, F. Llava-critic-r1: Your critic model is secretly a strong policy model. arXiv 2025, arXiv:2509.00676. [Google Scholar]
- Cai, Y.; Zhang, J.; He, H.; He, X.; Tong, A.; Gan, Z.; Wang, C.; Xue, Z.; Liu, Y.; Bai, X. Llava-kd: A framework of distilling multimodal large language models. In Proceedings of the ICCV, 2025; pp. 239–249. [Google Scholar]
- Xu, S.; Li, X.; Yuan, H.; Qi, L.; Tong, Y.; Yang, M.H. Llavadi: What matters for multimodal large language models distillation. arXiv 2024, arXiv:2407.19409. [Google Scholar]
- Shu, F.; Liao, Y.; Zhang, L.; Zhuo, L.; Xu, C.; Zhang, G.; Shi, H.; Dai, W.; Yu, Z.; He, W.; et al. Llava-mod: Making llava tiny via moe-knowledge distillation. In Proceedings of the ICLR, 2025; pp. 9386–9404. [Google Scholar]
- Cao, D.; Fu, D.; Yu, H.; Zheng, S.; Tan, X.; Jin, T. X-opd: Cross-modal on-policy distillation for capability alignment in speech llms. arXiv 2026, arXiv:2603.24596. [Google Scholar]
- Yuan, Q.; Lou, J.; Yu, X.; Lin, H.; Sun, L.; Han, X.; Lu, Y. Vision-OPD: Learning to See Fine Details for Multimodal LLMs via On-Policy Self-Distillation. arXiv 2026, arXiv:2605.18740. [Google Scholar]
- Liu, R.; Lv, X.; Li, G.; Zhu, X.; Wang, Z.; Zhang, Z.; Chen, J.; Li, Z.; Li, B.; Gao, J.; et al. Visual-Advantage On-Policy Distillation for Vision-Language Models. arXiv 2026, arXiv:2605.21924. [Google Scholar]
- Li, L.; Xie, Z.; Li, M.; Chen, S.; Wang, P.; Chen, L.; Yang, Y.; Wang, B.; Kong, L. Silkie: Preference distillation for large visual language models. arXiv 2023, arXiv:2312.10665. [Google Scholar]
- Ahn, D.; Choi, Y.; Kim, S.; Yu, Y.; Kang, D.; Choi, J. Isr-dpo: Aligning large multimodal models for videos by iterative self-retrospective dpo. In Proceedings of the AAAI, 2025; pp. 1728–1736. [Google Scholar]
- Zhang, H.; Mao, Z.; Zhang, L.; Zhang, Y. Uncertainty-Aware Exploratory Direct Preference Optimization for Multimodal Large Language Models. In Proceedings of the CVPR, 2026; pp. 37831–37841. [Google Scholar]
- Yang, Z.; Luo, X.; Han, D.; Xu, Y.; Li, D. Mitigating hallucinations in large vision-language models via dpo: On-policy data hold the key. In Proceedings of the CVPR, 2025. [Google Scholar]
- Lambert, N.; Morrison, J.; Pyatkin, V.; Huang, S.; Ivison, H.; Brahman, F.; Miranda, L.J.V.; Liu, A.; Dziri, N.; Lyu, S.; et al. Tulu 3: Pushing frontiers in open language model post-training. arXiv 2024, arXiv:2411.15124. [Google Scholar]
- Wang, X.; Yang, Z.; Feng, C.; Lu, H.; Li, L.; Lin, C.C.; Lin, K.; Huang, F.; Wang, L. Sota with less: Mcts-guided sample selection for data-efficient visual reasoning self-improvement. NeurIPS 2026. [Google Scholar] [CrossRef]
- Peng, Y.; Wang, P.; Wang, X.; Wei, Y.; Pei, J.; Qiu, W.; Jian, A.; Hao, Y.; Pan, J.; Xie, T.; et al. Skywork r1v: Pioneering multimodal reasoning with chain-of-thought. arXiv 2025, arXiv:2504.05599. [Google Scholar]
- Hong, J.; Zhao, C.; Zhu, C.; Lu, W.; Xu, G.; Yu, X. Deepeyesv2: Toward agentic multimodal model. arXiv 2025, arXiv:2511.05271. [Google Scholar]
- Li, B.; Sun, X.; Liu, J.; Wang, Z.; Wu, J.; Yu, X.; Chen, H.; Barsoum, E.; Chen, M.; Liu, Z. Latent visual reasoning. arXiv 2025, arXiv:2509.24251. [Google Scholar]
- Xu, Y.; Li, C.; Zhou, H.; Wan, X.; Zhang, C.; Korhonen, A.; Vulić, I. Visual Planning: Let’s Think Only with Images. arXiv 2025, arXiv:2505.11409. [Google Scholar]
- Duan, C.; Fang, R.; Wang, Y.; Wang, K.; Huang, L.; Zeng, X.; Li, H.; Liu, X. Got-r1: Unleashing reasoning capability of mllm for visual generation with reinforcement learning. arXiv 2025, arXiv:2505.17022. [Google Scholar]
- Liu, J.; Huang, X.; Zheng, J.; Liu, B.; Wang, J.; Yoshie, O.; Liu, Y.; Li, H. Mm-instruct: Generated visual instructions for large multimodal model alignment. arXiv 2024, arXiv:2406.19736. [Google Scholar]
- Wang, J.; Xu, H.; Ye, J.; Yan, M.; Shen, W.; Zhang, J.; Huang, F.; Sang, J. Mobile-agent: Autonomous multi-modal mobile device agent with visual perception. arXiv 2024, arXiv:2401.16158. [Google Scholar]
- Luo, R.; Wang, L.; He, W.; Chen, L.; Li, J.; Xia, X. Gui-r1: A generalist r1-style vision-language action model for gui agents. arXiv 2025, arXiv:2504.10458. [Google Scholar]
- Hu, A.; Xu, H.; Ye, J.; Yan, M.; Zhang, L.; Zhang, B.; Zhang, J.; Jin, Q.; Huang, F.; Zhou, J. mplug-docowl 1.5: Unified structure learning for ocr-free document understanding. In Proceedings of the EMNLP, 2024; pp. 3096–3120. [Google Scholar]
- Guo, Z.; Xu, R.; Yao, Y.; Cui, J.; Ni, Z.; Ge, C.; Chua, T.S.; Liu, Z.; Huang, G. Llava-uhd: an lmm perceiving any aspect ratio and high-resolution images. In Proceedings of the ECCV, 2024; pp. 390–406. [Google Scholar]
- Cheng, D.; Huang, S.; Zhu, Z.; Zhang, X.; Zhao, W.X.; Luan, Z.; Dai, B.; Zhang, Z. On domain-adaptive post-training for multimodal large language models. arXiv 2024, arXiv:2411.19930. [Google Scholar]
- Cheng, K.; YanTao, L.; Xu, F.; Zhang, J.; Zhou, H.; Liu, Y. Vision-language models can self-improve reasoning via reflection. In Proceedings of the ACL, 2025; pp. 8876–8892. [Google Scholar]
- Luo, R.; Zhang, H.; Chen, L.; Lin, T.E.; Liu, X.; Wu, Y.; Yang, M.; Li, Y.; Wang, M.; Zeng, P.; et al. Mmevol: Empowering multimodal large language models with evol-instruct. In Proceedings of the ACL, Findings; 2025, pp. 19655–19682.
- Zhou, J.; Chen, Y.; Li, H.; Jiang, Q.; Zhou, H.; Chen, Y.C.; Zhang, L. V-Reflection: Transforming MLLMs from Passive Observers to Active Interrogators. arXiv 2026, arXiv:2604.03307. [Google Scholar]
- Wu, Z.; Shi, K.; Zhang, C.; Liao, Z.; Yang, J.; Yang, N.; Peng, Q.; Zhang, L.; Xu, H.; Su, T.; et al. When Models Judge Themselves: Unsupervised Self-Evolution for Multimodal Reasoning. arXiv 2026, arXiv:2603.21289. [Google Scholar]
- Zeng, Y.; Huang, W.; Huang, S.; Bao, X.; Qi, Y.; Zhao, Y.; Wang, Q.; Chen, L.; Chen, Z.; Chen, H.; et al. Agentic Jigsaw Interaction Learning for Enhancing Visual Perception and Reasoning in Vision-Language Models. arXiv 2025, arXiv:2510.01304. [Google Scholar]
- Wang, Z.; Zhu, J.; Tang, B.; Li, Z.; Xiong, F.; Yu, J.; Blaschko, M.B. Jigsaw-r1: A study of rule-based visual reinforcement learning with jigsaw puzzles. arXiv 2025, arXiv:2505.23590. [Google Scholar]
- Chen, S.; Jie, Z.; Ma, L. Llava-mole: Sparse mixture of lora experts for mitigating data conflicts in instruction finetuning mllms. arXiv 2024, arXiv:2401.16160. [Google Scholar]
- Shen, Y.; Xu, Z.; Wang, Q.; Cheng, Y.; Yin, W.; Huang, L. Multimodal instruction tuning with conditional mixture of lora. In Proceedings of the ACL, 2024; pp. 637–648. [Google Scholar]
- Wei, Y.; Miao, Y.; Zhou, D.; Hu, D. Moka: Multimodal low-rank adaptation for mllms. NeurIPS 2026, 137470–137492. [Google Scholar]
- Che, C.; Wang, Z.; Yang, P.; Wang, C.; Ma, H.; Shi, Z. LoRA in LoRA: Towards parameter-efficient architecture expansion for continual visual instruction tuning. Proc. AAAI 2026, 19978–19986. [Google Scholar] [CrossRef]
- Lin, B.; Tang, Z.; Ye, Y.; Huang, J.; Zhang, J.; Pang, Y.; Jin, P.; Ning, M.; Luo, J.; Yuan, L. Moe-llava: Mixture of experts for large vision-language models. TMM, 2026. [Google Scholar]
- Zhong, S.; Gao, S.; Huang, Z.; Wen, W.; Žitnik, M.; Zhou, P. Moextend: Tuning new experts for modality and task extension. In Proceedings of the ACL, 2024; pp. 494–505. [Google Scholar]
- Xu, J.; Guo, Z.; Hu, H.; Chu, Y.; Wang, X.; He, J.; Wang, Y.; Shi, X.; He, T.; Zhu, X.; et al. Qwen3-omni technical report. arXiv 2025, arXiv:2509.17765. [Google Scholar]
- Team, Q. Qwen3-VL Technical Report. 2025. [Google Scholar] [CrossRef]
- MiniMax. MiniMax-01: Scaling Foundation Models with Lightning Attention. 2025. [Google Scholar] [CrossRef] [PubMed]
- Xu, Z.; Nguyen, K.D.; Mukherjee, P.; Bagchi, S.; Chaterji, S.; Liang, Y.; Li, Y. Learning to inference adaptively for multimodal large language models. In Proceedings of the ICCV, 2025; pp. 3552–3563. [Google Scholar]
- Chen, L.; Zhao, H.; Liu, T.; Bai, S.; Lin, J.; Zhou, C.; Chang, B. An image is worth 1/2 tokens after layer 2: Plug-and-play inference acceleration for large vision-language models. In Proceedings of the ECCV, 2024; pp. 19–35. [Google Scholar]
- Yang, S.; Chen, Y.; Tian, Z.; Wang, C.; Li, J.; Yu, B.; Jia, J. Visionzip: Longer is better but not necessary in vision language models. In Proceedings of the CVPR, 2025. [Google Scholar]
- Zhang, Y.; Fan, C.K.; Ma, J.; Zheng, W.; Huang, T.; Cheng, K.; Gudovskiy, D.; Okuno, T.; Nakata, Y.; Keutzer, K.; et al. Sparsevlm: Visual token sparsification for efficient vision-language model inference. arXiv 2024, arXiv:2410.04417. [Google Scholar]
- Li, W.; Yuan, Y.; Liu, J.; Tang, D.; Wang, S.; Qin, J.; Zhu, J.; Zhang, L. Tokenpacker: Efficient visual projector for multimodal llm. IJCV 2025, 6794–6812. [Google Scholar] [CrossRef]
- Shen, X.; Xiong, Y.; Zhao, C.; Wu, L.; Chen, J.; Zhu, C.; Liu, Z.; Xiao, F.; Varadarajan, B.; Bordes, F.; et al. Longvu: Spatiotemporal adaptive compression for long video-language understanding. arXiv 2024, arXiv:2410.17434. [Google Scholar]
- Chen, Y.; Xue, F.; Li, D.; Hu, Q.; Zhu, L.; Li, X.; Fang, Y.; Tang, H.; Yang, S.; Liu, Z.; et al. Longvila: Scaling long-context visual language models for long videos. Proc. ICLR 2025, 18227–18246. [Google Scholar]
- Zhang, P.; Zhang, K.; Li, B.; Zeng, G.; Yang, J.; Zhang, Y.; Wang, Z.; Tan, H.; Li, C.; Liu, Z. Long context transfer from language to vision. arXiv 2024, arXiv:2406.16852. [Google Scholar]
- Li, X.; Wang, Y.; Yu, J.; Zeng, X.; Zhu, Y.; Huang, H.; Gao, J.; Li, K.; He, Y.; Wang, C.; et al. Videochat-flash: Hierarchical compression for long-context video modeling. arXiv 2024, arXiv:2501.00574. [Google Scholar]
- Song, D.; Wang, W.; Chen, S.; Wang, X.; Guan, M.X.; Wang, B. Less is more: A simple yet effective token reduction method for efficient multi-modal llms. Proc. ACL 2025, 7614–7623. [Google Scholar]
- Li, B.; Wang, R.; Wang, G.; Ge, Y.; Ge, Y.; Shan, Y. Seed-bench: Benchmarking multimodal llms with generative comprehension. arXiv 2023, arXiv:2307.16125. [Google Scholar]
- Yu, W.; Yang, Z.; Li, L.; Wang, J.; Lin, K.; Liu, Z.; Wang, X.; Wang, L. Mm-vet: Evaluating large multimodal models for integrated capabilities. arXiv 2023, arXiv:2308.02490. [Google Scholar]
- Liu, Y.; Duan, H.; Zhang, Y.; Li, B.; Zhang, S.; Zhao, W.; Yuan, Y.; Wang, J.; He, C.; Liu, Z.; et al. Mmbench: Is your multi-modal model an all-around player? In Proceedings of the ECCV, 2024; pp. 216–233. [Google Scholar]
- Ding, S.; Wu, S.; Zhao, X.; Zang, Y.; Duan, H.; Dong, X.; Zhang, P.; Cao, Y.; Lin, D.; Wang, J. Mm-ifengine: Towards multimodal instruction following. In Proceedings of the ICCV, 2025; pp. 1099–1109. [Google Scholar]
- Fu, C.; Chen, P.; Shen, Y.; Qin, Y.; Zhang, M.; Lin, X.; Yang, J.; Zheng, X.; Li, K.; Sun, X.; et al. Mme: A comprehensive evaluation benchmark for multimodal large language models. NeurIPS, 2026. [Google Scholar]
- He, W.; Ju, F.; Fan, Z.; Min, R.; Cheng, M. Empowering reliable visual-centric instruction following in mllms. In Proceedings of the ACL, 2026; pp. 9460–9482. [Google Scholar]
- Li, Y.; Du, Y.; Zhou, K.; Wang, J.; Zhao, X.; Wen, J.R. Evaluating object hallucination in large vision-language models. In Proceedings of the EMNLP, 2023; pp. 292–305. [Google Scholar]
- Sun, Z.; Shen, S.; Cao, S.; Liu, H.; Li, C.; Shen, Y.; Gan, C.; Gui, L.; Wang, Y.X.; Yang, Y.; et al. Aligning large multimodal models with factually augmented rlhf. In Proceedings of the ACL, 2024; pp. 13088–13110. [Google Scholar]
- Guan, T.; Liu, F.; Wu, X.; Xian, R.; Li, Z.; Liu, X.; Wang, X.; Chen, L.; Huang, F.; Yacoob, Y.; et al. Hallusionbench: an advanced diagnostic suite for entangled language hallucination and visual illusion in large vision-language models. In Proceedings of the CVPR, 2024. [Google Scholar]
- Shi, E.; Shao, P.; Zhang, Y.; Cui, C.; Lyu, J.; Xia, X.; Shen, F.; Chua, T.S. Lingua-safetybench: A benchmark for safety evaluation of multilingual vision-language models. arXiv 2026, arXiv:2601.22737. [Google Scholar]
- Lu, P.; Mishra, S.; Xia, T.; Qiu, L.; Chang, K.W.; Zhu, S.C.; Tafjord, O.; Clark, P.; Kalyan, A. Learn to explain: Multimodal reasoning via thought chains for science question answering. NeurIPS 2022, 2507–2521. [Google Scholar] [CrossRef]
- Yue, X.; Ni, Y.; Zhang, K.; Zheng, T.; Liu, R.; Zhang, G.; Stevens, S.; Jiang, D.; Ren, W.; Sun, Y.; et al. Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi. In Proceedings of the CVPR, 2024; pp. 9556–9567. [Google Scholar]
- Lu, P.; Bansal, H.; Xia, T.; Liu, J.; Li, C.; Hajishirzi, H.; Cheng, H.; Chang, K.W.; Galley, M.; Gao, J. Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts. In Proceedings of the ICLR, 2024. [Google Scholar]
- Zhang, Z.; Chen, Z.; Zhang, Z.; Sun, Y.; Tian, Y.; Jia, Z.; Li, C.; Liu, X.; Min, X.; Zhai, G. PuzzleBench: A Fully Dynamic Evaluation Framework for Large Multimodal Models on Puzzle Solving. arXiv 2025, arXiv:2504.10885. [Google Scholar]
- Jiang, D.; Zhang, R.; Guo, Z.; Li, Y.; Qi, Y.; Chen, X.; Wang, L.; Jin, J.; Guo, C.; Yan, S.; et al. Mme-cot: Benchmarking chain-of-thought in large multimodal models for reasoning quality, robustness, and efficiency. arXiv 2025, arXiv:2502.09621. [Google Scholar]
- Yuan, J.; Peng, T.; Jiang, Y.; Lu, Y.; Zhang, R.; Feng, K.; Fu, C.; Chen, T.; Bai, L.; Zhang, B.; et al. Mme-reasoning: A comprehensive benchmark for logical reasoning in mllms. arXiv 2025, arXiv:2505.21327. [Google Scholar]
- Mathew, M.; Karatzas, D.; Jawahar, C. Docvqa: A dataset for vqa on document images. In Proceedings of the WACV, 2021; pp. 2200–2209. [Google Scholar]
- Deng, X.; Gu, Y.; Zheng, B.; Chen, S.; Stevens, S.; Wang, B.; Sun, H.; Su, Y. Mind2web: Towards a generalist agent for the web. NeurIPS 2023, 28091–28114. [Google Scholar] [CrossRef]
- Liu, Y.; Li, Z.; Huang, M.; Yang, B.; Yu, W.; Li, C.; Yin, X.C.; Liu, C.L.; Jin, L.; Bai, X. Ocrbench: on the hidden mystery of ocr in large multimodal models. Sci. China Inf. Sci. 2024, 220102. [Google Scholar]
- Xia, R.; Ye, H.; Yan, X.; Liu, Q.; Zhou, H.; Chen, Z.; Shi, B.; Yan, J.; Zhang, B. Chartx & chartvlm: A versatile benchmark and foundation model for complicated chart reasoning. TIP; 2025. [Google Scholar]
- Cheng, K.; Sun, Q.; Chu, Y.; Xu, F.; YanTao, L.; Zhang, J.; Wu, Z. Seeclick: Harnessing gui grounding for advanced visual gui agents. In Proceedings of the ACL, 2024; pp. 9313–9332. [Google Scholar]
- Liu, X.; Zhu, Y.; Gu, J.; Lan, Y.; Yang, C.; Qiao, Y. Mm-safetybench: A benchmark for safety evaluation of multimodal large language models. In Proceedings of the ECCV, 2024; pp. 386–403. [Google Scholar]
- Gong, Y.; Ran, D.; Liu, J.; Wang, C.; Cong, T.; Wang, A.; Duan, S.; Wang, X. Figstep: Jailbreaking large vision-language models via typographic visual prompts. In Proceedings of the AAAI, 2025; pp. 23951–23959. [Google Scholar]
- Luo, W.; Ma, S.; Liu, X.; Guo, X.; Xiao, C. Jailbreakv: A benchmark for assessing the robustness of multimodal large language models against jailbreak attacks. arXiv 2024, arXiv:2404.03027. [Google Scholar]
- Kembhavi, A.; Salvato, M.; Kolve, E.; Seo, M.; Hajishirzi, H.; Farhadi, A. A diagram is worth a dozen images. In Proceedings of the ECCV, 2016; pp. 235–251. [Google Scholar]
- Johnson, J.; Hariharan, B.; Van Der Maaten, L.; Fei-Fei, L.; Lawrence Zitnick, C.; Girshick, R. Clevr: A diagnostic dataset for compositional language and elementary visual reasoning. In Proceedings of the CVP, 2017; pp. 2901–2910. [Google Scholar]
- Wang, K.; Pan, J.; Shi, W.; Lu, Z.; Ren, H.; Zhou, A.; Zhan, M.; Li, H. Measuring multimodal mathematical reasoning with math-vision dataset. NeurIPS 2024, 95095–95169. [Google Scholar] [CrossRef]
- He, C.; Luo, R.; Bai, Y.; Hu, S.; Thai, Z.; Shen, J.; Hu, J.; Han, X.; Huang, Y.; Zhang, Y.; et al. Olympiadbench: A challenging benchmark for promoting agi with olympiad-level bilingual multimodal scientific problems. In Proceedings of the ACL, 2024; pp. 3828–3850. [Google Scholar]
- Masry, A.; Tan, J.Q.; Joty, S.; Hoque, E.; et al. Chartqa: A benchmark for question answering about charts with visual and logical reasoning. In Proceedings of the ACL, Findings, 2022; pp. 2263–2279. [Google Scholar]












| Method | Base LLM | Visual Encoder | # Params. | Modality | Training Data Scale | Source |
|---|---|---|---|---|---|---|
| LLaVA [1] | Vicuna | CLIP ViT-L/14 | ∼7.3B / ∼13.3B | I+T→T | 595K PD + 158K ID | Paper / GitHub / Project |
| MiniGPT-4 [24] | Vicuna-7B / 13B | EVA-CLIP ViT-G/14 | ∼7B / 13B | I+T→T | ∼5M PD+ 3.5K ID | Paper / GitHub / Project |
| InstructBLIP [2] | Flan-T5 / Vicuna | EVA-CLIP ViT-g | 3B–13B variants | I+T→T | 129M PD | Paper / GitHub |
| mPLUG-Owl [42] | LLaMA-7B | CLIP ViT-L/14 | ∼7B | I+T→T | 392k ID | Paper / GitHub |
| LLaMA-Adapter V2 [43] | LLaMA-7B / 13B | CLIP ViT | 7B / 13B | I+T→T | 567K PD + 132K ID | Paper / GitHub |
| LLaVA-1.5 [25] | Vicuna-7B / 13B | CLIP ViT-L/14-336 | ∼7B / ∼13B | I+T→T | 558K PD + 665K ID | Paper / GitHub |
| CogVLM [44] | Vicuna-7B / LLaMA | EVA-CLIP ViT | ∼7B | I+T→T | 1.54B PD | Paper / GitHub |
| LLaVA-NeXT [45] | Vicuna / Mistral / Yi | CLIP ViT-L/14-336 | 7B / 13B / 34B variants | I+T→T | 558K PD + 760K ID | GitHub / Project |
| Qwen2-VL-Instruct [46] | Qwen2 | ViT | 2B / 7B / 72B | I+T+V→T | - | Paper / GitHub |
| LLaVA-OneVision [47] | Qwen2 | SigLIP | 0.5B / 7B / 72B | I+T+V→T | ∼4.8M PD + 4.8M ID | Paper / Project |
| InternVL2.5 [48] | InternLM2.5 / Qwen variants | InternViT | 1B–78B variants | I+T+V→T | ∼120B tokens + 16.3M ID | Paper / GitHub / Project |
| Qwen2.5-VL-Instruct | Qwen2.5 | ViT | 3B / 7B / 72B | I+T+V→T | ∼4.1T tokens + ∼2M ID | Paper / GitHub |
| VILA [49] | LLaMA-2 / Vicuna variants | CLIP ViT-L/14 | 2.7B / 7B / 13B / 40B | I+T+V→T | ∼50M PD | Paper / GitHub |
| MobileVLM [50] | MobileLLaMA-1.4B / 2.7B | CLIP ViT-L/14 | 1.7B / 3B | I+T→T | 558K PD + 665K ID | Paper / GitHub |
| DeepSeek-VL [51] | DeepSeek-LLM | SigLIP+SAM | 1.3B / 7B | I+T→T | - | Paper / GitHub |
| Cambrian-1 [52] | LLaMA-3 / Vicuna variants | ViT / ConvNeXT series | 8B / 13B / 34B variants | I+T→T | ∼7M ID | Paper / GitHub / Project |
| MiniCPM-V 2.6 [53] | Qwen2-7B | SigLIP-400M | ∼8B | I+T+V→T | 570M PD + ∼8.8M ID | Paper / GitHub / Model |
| Molmo [54] | OLMo | ViT | 1B / 7B / 72B variants | I+T→T | 712K PD + ∼31.6M ID | Paper / GitHub / Project |
| Pixtral-12B [55] | Mistral-based decoder | Pixtral-ViT | ∼12B | I+T→T | - | Paper / Project / Model |
| Otter [56] | OpenFlamingo / LLaMA-style LM | CLIP ViT-L/14 | ∼9B | I+T+V→T | 2.8M ID | Paper / GitHub |
| LaVIN [57] | LLaMA-7B / 13B | CLIP ViT | 7B / 13B | I+T→T | 210K ID | Paper / GitHub |
| Valley [58] | Stable-Vicuna | CLIP ViT-L/14 | 7B | I+T+V→T | 1.297M PD + 234K ID | Paper / GitHub |
| NExT-GPT [59] | Vicuna | ImageBind | 7B | I+T+V+A→I+T+V+A | 15K PD + 5K ID | Paper / GitHub / Project |
| InternVL-1.0 [60] | InternLM / Vicuna variants | InternViT | 6B–26B variants | I+T→T | 6.01B PD + ∼4M ID | Paper / GitHub |
| GenLLaVA [61] | Mistral-7B | SigLIP | 7B | I+T→I+T | 558K PD + ∼2.08M ID | Paper |
| MLAN [62] | LLaVA-style MLLMs | CLIP ViT-L/14 | 7B variants | I+T→T | 558K PD + 186K ID | Paper |
| Inst-IT [63] | Qwen2-7B* | SigLIP-SO400* | ∼7B | I+T+V→T | 558K PD + 243K ID | Paper / GitHub |
| Vittle [64] | LLaVA-style MLLMs | CLIP ViT-L/14-336px | Backbone-dependent | I+T→T | 558K PD + 665K ID | Paper |
| LLaDA-V [65] | LLaDA | SigLIP 2-so400m-patch14-384 | 8B | I+T→T | 558K PD + 12M ID | Paper / Project |
| LLaVA-NeXT-Interleave [47] | Vicuna / Mistral / Qwen variants | CLIP ViT-L/14-336 | 7B / 13B | I+T+V→T | 1.18M interleaved ID | Paper / Project |
| Mantis-Idefics2 [66] | LLaVA / Idefics-style backbones | SiGLIP | 7B / 8B | I+T→T | 143M PD + 721K ID | Paper / Project |
| ImageBind-LLM [67] | LLaMA | - | 7B | I+T+V+A+3D→T | 940M PD + 205.5K ID | Paper / GitHub |
| ShareGPT4V [68] | Vicuna-7B | CLIP ViT-L/14 | 7B | I+T→T | 1.2M PD + 665K ID | Paper / GitHub / Project |
| Method | Feedback Granularity | #Params. | Modality | Optimization Paradigm | Source |
|---|---|---|---|---|---|
| Multimodal RLHF | |||||
| LLaVA-RLHF [36] | Pair-level | 7B / 13B | I+T→T | PPO | Paper / GitHub / Project |
| RLHF-V [37] | Span-level | 13B | I+T→T | DDPO | Paper / GitHub / Project |
| VisualPRM [84] | Step-level | 8B | I+T→T | PRM + BoN | Paper / Project |
| Gemini [85] | Mixed-level | – | I+V+A+T→T | SFT + RLHF | Paper / Project |
| Seed1.5-VL [86] | Mixed-level | 20B | I/V+T→T | RLHF + RLVR + RS | Paper / GitHub / Project |
| MiMo-VL [87] | Mixed-level | 7B | I/V+T→T | MORL | Paper / GitHub |
| Multimodal RLAIF | |||||
| VLM-RLAIF [88] | Pair-level | 7B | V+T→T | Context-aware RM + PPO | Paper / GitHub / Project |
| RLAIF-V [89] | Mixed-level | 7B / 12B | I+T→T | Iterative DPO + BoN | Paper / GitHub |
| Oracle-RLAIF [90] | Listwise-level | 7B | V+T→T | Oracle ranker + rank-aware GRPO | Paper |
| Multimodal DPO | |||||
| Video-DPO [91] | Response-level | 7B | V+T→T | DPO | Paper / GitHub |
| HA-DPO [92] | Response-level | 7B / 13B | I+T→T | DPO | Paper / GitHub / Project |
| V-DPO [93] | Response-level | 7B | I+T→T | CFG + DPO | Paper / GitHub |
| CLIP-DPO [38] | Response-level | 1.7B / 3B / 7B | I+T→T | DPO + CLIP-ranked preferences | Paper |
| CHAIR-DPO [94] | Object-level | 7B / 8B | I+T→T | DPO + CHAIR-based preferences | Paper / GitHub |
| mDPO [95] | Response-level | 3B / 7B | I+T→T | conditional and anchored DPO | Paper / GitHub / Project |
| MoD-DPO [96] | Modality-level | – | A+V+T→T | DPO | Paper |
| OmniDPO [97] | Modality-level | – | A+V+T→T | DPO | Paper |
| DA-DPO [98] | Pair-level | 7B / 13B | I+T→T | DPO | Paper / GitHub / Project |
| LPOI [99] | Listwise-level | – | I+T→T | Anchored DPO | Paper / GitHub / Project |
| Method | Base Model | # Param. | Modality | SFT | Algorithm | Source |
|---|---|---|---|---|---|---|
| R1-based Multimodal Reasoning | ||||||
| VLM-R1 [102] | Qwen2.5-VL | ∼4B / 7B | I+T→T | ✗ | GRPO | Paper / GitHub |
| Visual-RFT [103] | Qwen2-VL | 2B | I+T→T | ✗ | GRPO | Paper |
| Vision-R1 [39] | Qwen2.5-VL | 7B / 32B / 72B | I+T→T | ✓ | GRPO | Paper / GitHub |
| Video-R1 [104] | Qwen2.5-VL-Instruct | 7B | I/V+T→T | ✓ | T-GRPO | Paper / GitHub |
| VisualPRM [84] | InternVL-style MLLMs | 8B | I+T→T | ✗ | – | Paper / Project |
| VisualThinker-R1-Zero [105] | Qwen2-VL | 2B | I+T→T | ✗ | GRPO | Paper / GitHub |
| MM-Eureka [106] | Qwen2.5-VL-Instruct | 7B / 32B | I+T→T | ✗ | GRPO | Paper / GitHub |
| R1-Onevision [107] | Qwen2.5-VL | 3B / 7B | I+T→T | ✓ | GRPO | Paper |
| R1-Omni [108] | HumanOmni | ∼0.5B | A+V+T→T | ✓ | GRPO | Paper |
| Retrv-R1 [109] | Qwen2.5-VL | 3B / 7B | I/T→I/T | ✓ | GRPO | Paper |
| Thinking with Images | ||||||
| GRIT [110] | Qwen2.5-VL / InternVL3 | 3B / 2B | I+T→T | ✗ | GRPO-GR | Paper |
| Point-RFT [111] | Qwen2.5-VL | 7B | I+T→T | ✓ | GRPO | Paper |
| OpenThinkIMG [112] | Qwen2-VL | 2B | I+T→T | ✓ | V-ToolRL | Paper |
| VisionReasoner [113] | Qwen2.5-VL | 7B | I+T→T | ✗ | GRPO | Paper |
| DeepEyes [114] | Qwen2.5-VL | 7B | I+T→T | ✗ | GRPO | Paper / GitHub |
| VTool-R1 [115] | Qwen2.5-VL | 3B / 7B / 32B | I+T→T | ✗ | GRPO | Paper / GitHub |
| LanteRn [116] | Qwen2.5-VL-Instruct | 3B | I+T→T | ✓ | Latent-aware GRPO | Paper |
| Multimodal Self-Evolving | ||||||
| VIGC [117] | Vicuna+ViT-G/14 | – | I+T→T | ✗ | – | Paper / Project |
| MindGYM [118] | Qwen2.5-VL / InternVL3 | 7B / 14B / 32B / 38B | I/T→T | ✗ | – | Paper |
| SRPO [119] | Qwen2.5-VL | 7B / 32B | I+T→T | ✓ | GRPO | Paper |
| LLaVA-Critic [120] | LLaVA-OneVision | 7B / 72B | I+T→T | ✗ | – | Paper |
| MM-UPT [121] | Qwen2.5-VL | 7B | I+T→T | ✓ | GRPO | Paper / GitHub |
| LLaVA-Critic-R1 [122] | Qwen2.5-VL / ThinkLite-VL | 7B | I/V+T→T | ✗ | GRPO | Paper |
| Multimodal Distillation | ||||||
| LLaVA-KD [123] | Qwen-series | 1B / 2B | I+T→T | ✗ | – | Paper / GitHub |
| LLAVADI [124] | LLaVA-v1.5 / MobileLLaMA | 13B / ∼2.7B | I+T→T | ✗ | – | Paper |
| LLaVA-MoD [125] | Qwen-1.5 | 7B / 2B | I+T→T | ✗ | – | Paper / GitHub |
| Video-OPD [41] | Qwen3-VL-Instruct | 32B / 8B | V+T→T | ✗ | OPD | Paper |
| X-OPD [126] | Qwen3-Omni / Qwen3 | ∼3B | A+T→T | ✗ | OPD | Paper |
| Uni-OPD [40] | Qwen3-VL-Instruct | 4B | I+T→T | ✗ | OPD | Paper |
| Vision-OPD [127] | Qwen3.5-VL | 4B / 9B | I+T→T | ✗ | OPD | Paper |
| VA-OPD [128] | Qwen3-VL | 2B / 4B / 8B / 32B | I+T→T | ✗ | OPD | Paper |
| Method | Domain | Input | Adapted Cap. | Source |
|---|---|---|---|---|
| Mobile-Agent [141] | GUI | Screenshot, Inst. | Perception+Action | Paper / GitHub |
| GUI-R1 [142] | GUI | Screenshot, Traj. | Reason+Action | Paper / GitHub |
| mPLUG-DocOwl1.5 [143] | Doc. | Doc., Chart | Layout+Reason | Paper / GitHub |
| LLaVA-UHD [144] | HRV | HRV | Perception | Paper / GitHub |
| Med-Gemini [20] | Med. | Med., EHR | Clinical Reason | Paper |
| AdaMLLM [145] | Med., Food, RS | Image | Efficiency | Paper |
| Method | Base | Trainable Params. | Rank | # GPUs | Source |
|---|---|---|---|---|---|
| LLaVA-MoLE [152] | Vicuna-7B-v1.5 | 0.3B/7B | 32 | 64*A100 | Paper/GitHub |
| MixLoRA [153] | Vicuna-7B v1.3 | ∼10M/7B | 4 | 4*A100 | Paper/GitHub |
| MokA [154] | LLaMA2 | ∼0.1B/7B | 4 | 8† | Paper/GitHub |
| LiLoRA [155] | LLaVA-1.5-7B | ∼0.25B/7B | 128 / 64 | 4† | Paper/GitHub |
| Method | #Exp./Act. | Act. Params | #Params. | Tuned | Source |
|---|---|---|---|---|---|
| MoE-LLaVA [156] | 4/2 | 3.6B | 5.3B | FFN-MoE | Paper/GitHub |
| MoExtend [157] | -/- | 13B | - | New Exp. | Paper/GitHub |
| Qwen3-Omni [158] | 128*/8* | 3B | 30B | T–T MoE | Paper/GitHub |
| Qwen3-VL [159] | -/- | 3B / 22B | 30B / 235B | VL MoE | Paper/GitHub |
| MiniMax-01 [160] | 32/- | 45.9B | 456B | MoE LLM | Paper/GitHub |
| Seed1.5-VL [86] | -/- | 20B | - | MoE LLM | Paper/GitHub |
| Type | Method | Input | Granularity | Core Technique | Main Goal | Source |
|---|---|---|---|---|---|---|
| EVP | LLaVA-UHD [144] | HR image | Region / Token | Image modularization, token compression, spatial schema | Fine-grained HR perception | Paper/GitHub |
| AdaMLLM/AdaLLaVA [161] | Image | Instance / Budget | Dynamic inference reconfiguration | Accuracy–latency trade-off | Paper/GitHub | |
| InternVL2 [60] | HR image | Tile | Dynamic image tiling | Dense visual perception | Project/GitHub | |
| TC | FastV [162] | Image / Video | Layer / Token | Attention-guided visual token pruning | Reduce prefilling and attention cost | Paper/GitHub |
| VisionZip [163] | Image / Video | Token | Informative token selection | Remove visual redundancy | Paper/GitHub | |
| SparseVLM [164] | Image / Video | Token / Layer | Text-guided sparsification, token recycling | Reduce visual FLOPs | Paper/GitHub | |
| TokenPacker [165] | Image | Projector / Token | Coarse-to-fine token packing | Compress while preserving details | Paper/GitHub | |
| LCO | LongVU [166] | Video | Frame / Token | Spatiotemporal adaptive compression | Long-video compression | Paper/GitHub |
| LongVILA [167] | Video | Sequence / System | Long-context extension, SFT, sequence parallelism | Scalable long-video training | Paper/GitHub | |
| LongVA [168] | Video | Context | Language-to-vision context transfer | Long-context video understanding | Paper/GitHub | |
| VideoChat-Flash [169] | Video | Hierarchical token | HiCo, short-to-long training | Extremely long-video modeling | Paper/GitHub |
for datasets mainly used in training,
for benchmarks designed for evaluation, and
+
for resources containing both training and evaluation splits. The “Origin” column identifies the primary organization using common abbreviations.
for datasets mainly used in training,
for benchmarks designed for evaluation, and
+
for resources containing both training and evaluation splits. The “Origin” column identifies the primary organization using common abbreviations.| Name | Year | Scale | Role | Origin | Source |
|---|---|---|---|---|---|
| Instruction Following | |||||
| LLaVA-Instruct-150K [1] | 2023 | ∼80K images / 158K instruction samples | ![]() |
UW–Madison, Columbia | Paper/GitHub/Data |
| SEED-Bench [171] | 2023 | ∼19K samples | ![]() |
Tencent AI Lab, ARC Lab | Paper/GitHub/Data |
| MM-Vet [172] | 2023 | 200 images / 218 samples | ![]() |
NUS, Microsoft Azure AI | Paper/GitHub/Data |
| MMBench [173] | 2024 | 3,217 data samples | ![]() |
SHLAB, ZJU | Paper/GitHub/Data |
| ShareGPT4V [68] | 2024 | 1,346K samples | ![]() |
USTC, SHLAB | Paper/GitHub/Data |
| MM-IFInstruct [174] | 2025 | 23K samples | ![]() |
FDU, SII | Paper/GitHub/Data |
| MME [175] | 2026 | 1,187 images / 2,374 samples | ![]() |
SKL-NST, CASIA | Paper/GitHub/Data |
| VC-IFInstruct [176] | 2026 | 10k samples | ![]() |
HKUST, PSU | Paper |
| Preference Calibration | |||||
| POPE [177] | 2023 | ∼9K question samples | ![]() |
RUC, Meituan Group | Paper/GitHub/Data |
| MMHal-Bench [178] | 2023 | ∼96 image-question pairs | ![]() |
UCB, MIT-IBM AI Lab | Paper/GitHub/Data |
| HallusionBench [179] | 2024 | ∼346 images / ∼1.1K question samples | ![]() |
UMD | Paper/GitHub/Data |
| LLaVA-RLHF [36] | 2024 | 10K image-based conversations | ![]() |
UCB, MIT-IBM AI Lab | Paper/GitHub/Data |
| RLHF-V [37] | 2024 | 1.4K image-based conversations | ![]() |
THU, NUS | Paper/GitHub/Data |
| Lingua-SafetyBench [180] | 2026 | ∼100K image-question pairs | ![]() |
HKU, ZJU | Paper/GitHub/Data |
| Reason Enhancement | |||||
| ScienceQA [181] | 2022 | ∼10.3K image-question pairs / ∼21.2K questions |
+
|
UCLA | Paper/GitHub/Data |
| MMMU [182] | 2024 | ∼11.5K questions | ![]() |
Waterloo, OSU, CMU | Paper/GitHub/Data |
| MathVista [183] | 2024 | 6,141 samples | ![]() |
UCLA, MSR | Paper/GitHub/Data |
| PuzzleBench [184] | 2025 | 11,840 samples | ![]() |
SJTU | Paper |
| MME-CoT [185] | 2025 | 1,130 questions | ![]() |
CUHK, ByteDance, NEU | Paper/GitHub/Data |
| MME-Reasoning [186] | 2025 | 1,188 questions | ![]() |
Fudan, CUHK, Shanghai AI Lab | Paper/GitHub/Data |
| Domain Adaptation | |||||
| DocVQA [187] | 2020 | ∼12.8K images + ∼50K questions |
+
|
UB, CVC | Paper/Data |
| Mind2Web [188] | 2023 | ∼137 websites + ∼2.4K tasks + ∼137K annotated steps |
+
|
OSU, CMU | Paper/GitHub/Data |
| OCRBench [189] | 2023 | ∼1K questions | ![]() |
SAIL | Paper/GitHub |
| ChartX [190] | 2024 | ∼48K images / 6K validation+test samples |
+
|
Shanghai AI Lab, SJTU | Paper/GitHub |
| ScreenSpot [191] | 2024 | ∼610 screenshots + ∼1.3K GUI instructions | ![]() |
CUHK, SAIL | Paper/GitHub/Data |
Short Biography of Authors
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.