Submitted:
30 October 2025
Posted:
03 November 2025
You are already at the latest version
Abstract
Keywords:
1. Introduction
2. What is a Language Model?
2.1. The Evolution of Language Models
2.2. Attention Mechanisms
3. Proprietary vs. Open Source LLMs
4. Key Large Language Models
4.1. GPT
4.2. Claude
4.3. Gemini
4.4. LLaMA
4.5. LLaMA-2 and LLaMA-2 Chat
4.6. MedAlpaca
4.7. Mistral 7B
4.8. Falcon-7B and Falcon-40B
4.9. Falcon-180B
4.10. Grok-1
| Model | Parameters | Commercial Use | License | Attention | Pre-training Token Length | VRAM / RAM Required | Open Source | Fine-tuneable |
| LLaMA | 7B | No | LLaMA License | MHA | 1T | 6GB VRAM | Yes | Yes |
| LLaMA | 13B | No | LLaMA License | MHA | 1.5T | 10GB VRAM | Yes | Yes |
| LLaMA | 65B | No | LLaMA License | MHA | 1.5T | 40GB VRAM | Yes | Yes |
| LLaMA-2 | 7B | Yes | LLaMA-2 License | GQA | 2T | 6GB VRAM | Yes | Yes |
| LLaMA-2 | 13B | Yes | LLaMA-2 License | GQA | 2T | 10GB VRAM | Yes | Yes |
| LLaMA-2 | 70B | Yes | LLaMA-2 License | GQA | 2T | 40GB VRAM | Yes | Yes |
| Mistral | 7B | Yes | Apache 2.0 | GQA | - | 6GB VRAM | Yes | Yes |
| Falcon | 7B | Yes | Apache 2.0 | MQA | 1.5T | 15GB RAM | Yes | Yes |
| Falcon | 40B | Yes | Apache 2.0 | MQA | 1T | 40GB RAM | Yes | Yes |
| Falcon | 180B | Yes | Apache 2.0 | MQA | 3.5T | 320GB RAM | Yes | Yes |
| GPT-3 | 175B | Yes | OpenAI License | MHA | 300B | Via API | No | Limited |
| GPT-3.5 turbo | 175B | Yes | OpenAI License | Not disclosed | Not disclosed | Via API | No | Yes |
| GPT-4 | Not disclosed | Yes | OpenAI License | Not disclosed | Not disclosed | Via API | No | No |
| Gemini | 137B | Yes | Gemini Pro License | MQA | Not disclosed | Via API | No | No |
| Claude | 93B | Yes | Claude Pro License | Unknown | Unknown | Via API | No | No |
| Claude 2 | 137B | Yes | Claude Pro License | Unknown | Unknown | Via API | No | No |
| Claude 3 | Unknown | Yes | Claude Pro License | Unknown | Unknown | Via API | No | No |
| Grok-1 | 314B | Yes | Apache 2.0 for code and Grok-1 weights | 48 attention heads for queries, 8 for keys/values | Unspecified | Unspecified | Yes | No |
5. Vision Models and Multi-Modal Large Language Models
5.1. Vision Models
5.1.1. BLIP-2
5.1.2. Vision Transformer (ViT)
5.1.3. Contrastive Language–Image Pretraining (CLIP)
5.2. Early Approaches to Multi-Modal Processing
5.3. Multi-Modal Large Language Models (MM-LLMs)
5.3.1. LLaVA (Large Language and Vision Assistant)
5.3.2. Kosmos-1 and Kosmos-2
5.3.3. MiniGPT-4
5.3.4. mPLUG-OWL
5.3.5. Summary and Comparison of Selected MM-LLMs
6. Model Tuning
6.1. Full Fine-Tuning
6.2. Parameter-Efficient Fine-Tuning (PEFT)
6.2.1. Low-Rank Adaptation (LoRA)
6.2.2. Quantised Low-Rank Adaptation (QLoRA)
6.2.3. Supervised Fine-Tuning (SFT)
6.3. Prompt Engineering
- Few-shot prompting, where multiple examples are provided;
- One-shot prompting, where only one example is given;
- Zero-shot prompting, where only the task description is supplied.
6.4. Reinforcement Learning with Human Feedback (RLHF)
7. Model Evaluation and Benchmarking
8. Conclusions
References
- Partha Pratim Ray, “ChatGPT: A comprehensive review on background, applications, key challenges, bias, ethics, limitations and future scope,” In Internet of Things and Cyber-Physical Systems, Elsevier, 2023.
- Alec Radford, Karthik Narasimhan, Tim Salimans, and Ilya Sutskever, “Improving language understanding with unsupervised learning,” Technical report, OpenAI, 2018.
- Leigang Qu et al., “LayoutLLM-T2I: Eliciting Layout Guidance from LLM for Text-to-Image Generation,” In Proceedings of the 31st ACM International Conference on Multimedia (MM ’23), Ottawa, Canada: Association for Computing Machinery, 2023, pp. 643–654. DOI: 10.1145/3581783.3612012.
- Long Lian et al., “LLM-grounded Video Diffusion Models,” In The Twelfth International Conference on Learning Representations, 2024. URL: https://openreview.net/forum?id=exKHibougU.
- Qingyao Ai et al., “Information Retrieval meets Large Language Models: A strategic report from Chinese IR community,” In AI Open, vol. 4, 2023, pp. 80–90. DOI: 10.1016/j.aiopen.2023.08.001.
- Yasmin Moslem, Rejwanul Haque, John D. Kelleher, and Andy Way, “Adaptive Machine Translation with Large Language Models,” In Proceedings of the 24th Annual Conference of the European Association for Machine Translation, Tampere, Finland, 2023, pp. 227–237. URL: https://aclanthology.org/2023.eamt-1.22.
- Priyanka Gupta, Bosheng Ding, Chong Guan, and Ding Ding, “Generative AI: A systematic review using topic modelling techniques,” In Data and Information Management, 2024, pp. 100066. DOI: 10.1016/j.dim.2024.100066.
- Jingfeng Yang et al., “Harnessing the Power of LLMs in Practice: A Survey on ChatGPT and Beyond,” In ACM Trans. Knowl. Discov. Data, New York, NY, USA: Association for Computing Machinery, 2024. DOI: 10.1145/3649506.
- Leigang Qu et al., “LayoutLLM-T2I: Eliciting Layout Guidance from LLM for Text-to-Image Generation,” In Proceedings of the 31st ACM International Conference on Multimedia, 2023.
- Long Lian et al., “LLM-grounded Video Diffusion Models,” In ICLR 2024, URL: https://openreview.net/forum?id=exKHibougU.
- Rohan Anil et al., “Gemini: a family of highly capable multimodal models,” arXiv preprint arXiv:2312.11805, 2023.
- “Claude 2: a guide to Anthropic’s AI model and Chatbot,” Accessed: 22-Feb-2024, https://www.zapier.com/blog/claude-ai/.
- Hugo Touvron et al., “LLaMA: Open and efficient foundation language models,” arXiv preprint arXiv:2302.13971, 2023.
- Mohaimenul Azam Khan Raiaan et al., “A Review on Large Language Models: Architectures, Applications, Taxonomies, Open Issues and Challenges,” In IEEE Access, vol. 12, 2024, pp. 26839–26874. DOI: 10.1109/ACCESS.2024.3365742.
- Kassym-Jomart Tokayev, “Ethical implications of large language models: a multidimensional exploration of societal, economic, and technical concerns,” In International Journal of Social Analytics, vol. 8, no. 9, 2023, pp. 17–33.
- Siddharth Samsi et al., “From words to watts: Benchmarking the energy costs of large language model inference,” In 2023 IEEE High Performance Extreme Computing Conference (HPEC), 2023, pp. 1–9. IEEE.
- Matthias C. Rillig et al., “Risks and benefits of large language models for the environment,” In Environmental Science & Technology, vol. 57, no. 9, 2023, pp. 3464–3466.
- Nick Barney, “What Is Language Modeling?—Definition from TechTarget,” [Accessed 22-Feb-2024], https://www.techtarget.com/searchenterpriseai/definition/language-modeling.
- Dorian Drost, “A brief history of language models—towardsdatascience.com,” [Accessed 22-Feb-2024], https://towardsdatascience.com/a-brief-history-of-language-models-d9e4620e025b, 2023.
- Amrita Anandika, Smita Prava Mishra, and Madhusmita Das, “Review on Usage of Hidden Markov Model in Natural Language Processing,” In Intelligent and Cloud Computing: Proceedings of ICICC 2019, Volume 1, 2021, pp. 415–423. Springer.
- Sheila Castilho et al., “Is neural machine translation the new state of the art?” In The Prague Bulletin of Mathematical Linguistics PBML, 2017.
- Rajvardhan Patil, Sorio Boit, Venkat Gudivada, and Jagadeesh Nandigam, “A Survey of Text Representation and Embedding Techniques in NLP,” In IEEE Access, 2023.
- Can Cui et al., “A Survey on Multimodal Large Language Models for Autonomous Driving,” 2023, arXiv:2311.12320 [cs.AI].
- Benjamin McCloskey, “Choosing Neural Networks over N-Gram Models for Natural Language Processing—towardsdatascience.com,” [Accessed 22-Feb-2024].
- Ashish Vaswani et al., “Attention Is All You Need,” 2017, arXiv:1706.03762.
- Shaohan Huang et al., “Language is not all you need: Aligning perception with language models,” arXiv preprint arXiv:2302.14045, 2023.
- Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova, “BERT: Pre-training of deep bidirectional transformers for language understanding,” arXiv preprint arXiv:1810.04805, 2018.
- Bernard Marr, “A Short History Of ChatGPT: How We Got To Where We Are Today—forbes.com,” [Accessed 22-Feb-2024], https://www.forbes.com/sites/bernardmarr/2023/05/19/a-short-history-of-chatgpt-how-we-got-to-where-we-are-today/, 2023.
- Gaudenz Boesch, “Llama 2: The Next Revolution in AI Language Models - Complete 2024 Guide - viso.ai,” [Accessed 07-Mar-2024], https://viso.ai/deep-learning/llama-2/.
- Shobhit Agarwal, “Navigating the Attention Landscape: MHA, MQA, and GQA Decoded,” [Accessed 23-Feb-2024], https://iamshobhitagarwal.medium.com/navigating-the-attention-landscape-mha-mqa-and-gqa-decoded-288217d0a7d1, 2024.
- Joshua Ainslie et al., “GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints,” arXiv preprint arXiv:2305.13245, 2023.
- Ari Chanen, “What I learned from Bloomberg’s experience of building their own LLM,” [Accessed 21-Feb-2024], linkedin.com, 2023.
- Sara Guaglione, “The case for and against open-source large language models for use in newsrooms,” [Accessed 22-Feb-2024], digiday.com, 2023.
- IBM Data and AI Team, “Open source large language models: Benefits, risks and types,” [Accessed 22-Feb-2024], ibm.com.
- Nikita Khudov, “The Future of LLMs: Proprietary versus Open-Source,” [Accessed 22-Feb-2024], linkedin.com, 2023.
- Dylan Patel, “Google ’We Have No Moat, And Neither Does OpenAI’,” [Accessed 22-Feb-2024], semianalysis.com, 2023.
- “EU AI Act: first regulation on artificial intelligence,” [Accessed 15-Mar-2024], europarl.europa.eu.
- Indumathi Pandiyan, “Open Source or Proprietary LLMs,” [Accessed 21-Feb-2024], medium.com, 2023.
- Tom Brown et al., “Language models are few-shot learners,” In Advances in Neural Information Processing Systems, vol. 33, 2020, pp. 1877–1901.
- Alec Radford et al., “Language models are unsupervised multitask learners,” In OpenAI Blog, 1.8, 2019, pp. 9.
- Josh Achiam et al., “GPT-4 technical report,” In arXiv preprint arXiv:2303.08774, 2023.
- Anthropic, “Introducing Claude,” [Accessed 26-Mar-2024], anthropic.com.
- Xiang Yue et al., “MMMU: A massive multi-discipline multimodal understanding and reasoning benchmark for expert AGI,” In arXiv preprint arXiv:2311.16502, 2023.
- Hugo Touvron et al., “LLaMA: Open and Efficient Foundation Language Models,” In arXiv preprint arXiv:2302.13971 [cs.CL], 2023.
- Jack W. Rae et al., “Scaling Language Models: Methods, Analysis & Insights from Training Gopher,” In arXiv preprint arXiv:2112.11446 [cs.CL], 2022.
- Sik-Ho Tsang, “Review: LLaMA: Open and Efficient Foundation Language Models,” [Accessed 05-Mar-2024], medium.com.
- Jordan Hoffmann et al., “Training Compute-Optimal Large Language Models,” In arXiv preprint arXiv:2203.15556 [cs.CL], 2022.
- Andrew Johnson, “Understanding Rotary Position Embedding: A Key Concept in Transformer Models,” [Accessed 05-Mar-2024], medium.com, 2023.
- Hugo Touvron et al., “LLaMA 2: Open Foundation and Fine-Tuned Chat Models”, 2023, arXiv:2307.09288 [cs.CL].
- Ankur A. Patel, “LLaMA 1 vs LLaMA 2: A Deep Dive into Meta’s LLMs—ankursnewsletter.com” [Accessed 07-Mar-2024], https://www.ankursnewsletter.com/p/llama-1-vs-llama-2-a-deep-dive-into, 2023.
- Sebastian Streng, “Game-Changer 2024: Meta’s LLAMA 2.0—sebastianstreng96” [Accessed 07-Mar-2024], https://medium.com/@sebastianstreng96/game-changer-2024-metas-llama-2-0-4ab1316b6aa4, 2023.
- Stephen M. Walker, “What is Grouped Query Attention (GQA)?” [Accessed 07-Mar-2024], https://klu.ai/glossary/grouped-query-attention.
- Rafael Pierre, “Parameter-Efficient Fine-Tuning (PEFT): Enhancing Large Language Models with Minimal Costs—mlopshowto.com” [Accessed 08-Mar-2024].
- Edward J. Hu et al., “LoRA: Low-Rank Adaptation of Large Language Models”, 2021, arXiv:2106.09685 [cs.CL].
- Tianyu Han et al., “MedAlpaca – An Open-Source Collection of Medical Conversational AI Models and Training Data”, 2023, arXiv:2304.08247 [cs.CL].
- Rewon Child, Scott Gray, Alec Radford and Ilya Sutskever, “Generating long sequences with sparse transformers”, 2019, arXiv:1904.10509 [cs.CL].
- Technology Innovation Institute (TII), “UAE’s Technology Innovation Institute Launches Open-Source Falcon 40B Large Language Model for Research & Commercial Utilization—tii.ae” [Accessed 23-Feb-2024].
- Minhajul Hoque, “Exploring the Falcon LLM: The New King of The Jungle—minh.hoque” [Accessed 23-Feb-2024], https://medium.com/@minh.hoque/exploring-the-falcon-llm-the-new-king-of-the-jungle-5c6a15b91159, 2023.
- Leandro Werra et al., “The Falcon has landed in the Hugging Face ecosystem—huggingface.co” [Accessed 23-Feb-2024], https://huggingface.co/blog/falcon, 2023.
- Ebtesam Almazrouei et al., “The Falcon Series of Open Language Models”, 2023, arXiv:2311.16867 [cs.CL].
- Guilherme Penedo et al., “The RefinedWeb Dataset for Falcon LLM: Outperforming Curated Corpora with Web Data, and Web Data Only”, 2023, arXiv:2306.01116 [cs.CL].
- Shaoni Mukherjee, “Introducing Falcon 180b: A Comprehensive Guide with a Hands-On Demo of the Falcon 40B—blog.paperspace.com” [Accessed 23-Feb-2024], https://blog.paperspace.com/introducing-falcon/, 2023.
- Femiloye Oyerinde, “BLIP-2: A Breakthrough Approach in Vision-Language Pre-training—femiloyeseun,” Medium, 2023. [Accessed: 20-Feb-2024]. Available: https://medium.com/@femiloyeseun/blip-2-a-breakthrough-approach-in-vision-language-pre-training-1de47b54f13a.
- Junnan Li, Dongxu Li, Silvio Savarese, and Steven Hoi, “BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models,” arXiv preprint arXiv:2301.12597, 2023.
- Alexey Dosovitskiy et al., “An image is worth 16x16 words: Transformers for image recognition at scale,” arXiv preprint arXiv:2010.11929, 2020.
- Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun, “Deep residual learning for image recognition,” in Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR), 2016, pp. 770–778.
- Alec Radford et al., “Learning transferable visual models from natural language supervision,” in International Conference on Machine Learning, 2021, pp. 8748–8763, PMLR.
- Yifan Li et al., “Evaluating object hallucination in large vision-language models,” arXiv preprint arXiv:2305.10355, 2023.
- Chen-Wei Xie et al., “RA-CLIP: Retrieval Augmented Contrastive Language-Image Pre-Training,” in IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2023, pp. 19265–19274. DOI: 10.1109/CVPR52729.2023.01846.
- Taraneh Ghandi, Hamidreza Pourreza, and Hamidreza Mahyar, “Deep learning approaches on image captioning: A review,” ACM Computing Surveys, vol. 56, no. 3, pp. 1–39, 2023.
- Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee, “Visual instruction tuning,” arXiv preprint arXiv:2304.08485, 2023.
- Wei-Lin Chiang et al., “Vicuna: An Open-Source Chatbot Impressing GPT-4 with 90% ChatGPT Quality,” LMSYS.org Blog, 2023. Available: https://lmsys.org/blog/2023-03-30-vicuna/.
- Tsung-Yi Lin et al., “Microsoft COCO: Common objects in context,” in ECCV 2014: 13th European Conference on Computer Vision, Zurich, Switzerland, 2014, pp. 740–755. Springer.
- Haotian Liu, Chunyuan Li, Yuheng Li, and Yong Jae Lee, “Improved baselines with visual instruction tuning,” arXiv preprint arXiv:2310.03744, 2023.
- Leo Gao et al., “The Pile: An 800GB dataset of diverse text for language modeling,” arXiv preprint arXiv:2101.00027, 2020.
- Common Crawl Foundation, “Common Crawl,” [Accessed: 22-Feb-2024]. Available: https://commoncrawl.org/.
- Zhiliang Peng et al., “Kosmos-2: Grounding Multimodal Large Language Models to the World,” arXiv preprint arXiv:2306.14824, 2023. Available: https://api.semanticscholar.org/CorpusID:259262263.
- Deyao Zhu et al., “MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models,” arXiv preprint arXiv:2304.10592, 2023.
- Wei-Lin Chiang et al., “Vicuna: An Open-Source Chatbot Impressing GPT-4 with 90% ChatGPT Quality,” LMSYS Blog, 2023. [Online]. Available: https://lmsys.org/blog/2023-03-30-vicuna/. [Accessed: Feb. 15, 2024].
- Qinghao Ye et al., “mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality,” arXiv preprint arXiv:2304.14178 [cs.CL], 2023.
- Alexey Dosovitskiy et al., “An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale,” arXiv preprint arXiv:2010.11929 [cs.CV], 2021.
- Najeeb Nabwani, “Full Fine-Tuning, PEFT, Prompt Engineering, or RAG?—deci.ai,” 2023. [Online]. Available: https://deci.ai/blog/fine-tuning-peft-prompt-engineering-and-rag-which-one-is-right-for-you. [Accessed: 17-Feb-2024].
- Lingling Xu et al., “Parameter-Efficient Fine-Tuning Methods for Pretrained Language Models: A Critical Review and Assessment,” 2023, arXiv:2312.12148 [cs.CL].
- Tim Dettmers, Artidoro Pagnoni, Ari Holtzman, Luke Zettlemoyer, “QLoRA: Efficient Finetuning of Quantized LLMs,” 2023, arXiv:2305.14314 [cs.LG].
- Jose J. Martinez, “Supervised Fine-tuning: customizing LLMs—medium.com,” 2023. [Online]. Available: https://medium.com/mantisnlp/supervised-fine-tuning-customizing-llms-a2c1edbf22c3. [Accessed: 02-Mar-2024].
- Armin Norouzi, “The Ultimate Guide to LLM Fine Tuning: Best Practices & Tools—Lakera,” 2023. [Online]. Available: https://www.lakera.ai/blog/llm-fine-tuning-guide. [Accessed: 02-Mar-2024].
- Pengfei Liu et al., “Pre-Train, Prompt, and Predict: A Systematic Survey of Prompting Methods in Natural Language Processing,” ACM Comput. Surv., vol. 55, no. 9, 2023, doi:10.1145/3560815.
- Shivam Garg, Dimitris Tsipras, Percy S. Liang, Gregory Valiant, “What can transformers learn in-context? A case study of simple function classes,” in Advances in Neural Information Processing Systems, vol. 35, 2022, pp. 30583–30598.
- Laria Reynolds, Kyle McDonell, “Prompt programming for large language models: Beyond the few-shot paradigm,” in Extended Abstracts of the 2021 CHI Conference on Human Factors in Computing Systems, 2021, pp. 1–7.
- Li Sun, Liuan Wang, Jun Sun, Takayuki Okatani, “Prompt Prototype Learning Based on Ranking Instruction For Few-Shot Visual Tasks,” in 2023 IEEE International Conference on Image Processing (ICIP), 2023, pp. 3235–3239, doi:10.1109/ICIP49359.2023.10222039.
- Now Next Later AI, “What is RLHF: Reinforcement Learning from Human Feedback—medium.com,” 2023. [Online]. Available: https://medium.com/generative-ai-insights-for-business-leaders-and/what-is-rlhf-reinforcement-learning-from-human-feedback-876da930bf16. [Accessed: 02-Mar-2024].
- Samuel Gehman et al., “RealToxicityPrompts: Evaluating Neural Toxic Degeneration in Language Models,” in Findings of the Association for Computational Linguistics: EMNLP, 2020, pp. 3356–3369.
- Alyssa Lees et al., “A New Generation of Perspective API: Efficient Multilingual Character-level Transformers,” in Proceedings of the 28th ACM SIGKDD Conference on Knowledge Discovery and Data Mining, KDD ’22, Washington DC, USA: ACM, 2022, pp. 3197–3207, doi:10.1145/3534678.3539147.
- Daniel Nest, “LLM Benchmarks: What Do They All Mean?—whytryai.com,” 2023. [Online]. Available: https://www.whytryai.com/p/llm-benchmarks. [Accessed: 02-Mar-2024].
- Peter Clark et al., “Think you have solved question answering? try ARC, the AI2 reasoning challenge,” 2018, arXiv:1803.05457.
- Rowan Zellers et al., “Hellaswag: Can a machine really finish your sentence?” 2019, arXiv:1905.07830.
- Christopher Clark et al., “BoolQ: Exploring the surprising difficulty of natural yes/no questions,” 2019, arXiv:1905.10044.
- Todor Mihaylov, Peter Clark, Tushar Khot, Ashish Sabharwal, “Can a Suit of Armor Conduct Electricity? A New Dataset for Open Book Question Answering,” in Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, 2018, pp. 2381–2391.
- Yonatan Bisk, Rowan Zellers, Jianfeng Gao, Yejin Choi, “PIQA: Reasoning about physical commonsense in natural language,” in Proceedings of the AAAI Conference on Artificial Intelligence, vol. 34.05, 2020, pp. 7432–7439.
- Dan Hendrycks et al., “Measuring massive multitask language understanding,” 2020, arXiv:2009.03300.
- Stephanie Lin, Jacob Hilton, Owain Evans, “TruthfulQA: Measuring How Models Mimic Human Falsehoods,” in Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2022, pp. 3214–3252.
- Anisha Gunjal, Jihan Yin, Erhan Bas, “Detecting and preventing hallucinations in large vision language models,” 2023, arXiv:2308.06394.
- Nagesh Mashette, “LLM Benchmarks (Introduction to Benchmarks Techniques),” 2023. [Online]. Available: https://medium.com/@nageshmashette32/llm-benchmarks-introduction-to-benchmarks-techniques-6518527620eb. [Accessed: 02-Mar-2024].
| Model | Open Source | Fine-Tuneable | LLM Used | Vision Model Used |
|---|---|---|---|---|
| LLaVA | Yes | Yes | Vicuna | CLIP ViT-L/14 |
| Kosmos-1 and -2 | Yes | Yes | Grounded image–text pairs to train an integrated model | – |
| MiniGPT-4 | Yes | Yes | Vicuna | Q-Former and CLIP ViT-G/14 |
| mPLUG-OWL | Yes | Yes | LLaMA-7B | CLIP ViT-L/14 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2025 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).