Submitted:
24 February 2024
Posted:
26 February 2024
You are already at the latest version
Abstract
Keywords:
1. Introduction
2. Related Works
Trends
Existing Arabic LLMs
AraNizer’s Impact
ArabianGPT’s Positioning
3. Datasets
3.1. Dataset D1
3.2. Dataset D2: Composite Linguistic Resource
4. Tokenization
4.1. Overview
- Word root identification: Separating prefixes, suffixes, and inflections from the core root morpheme poses challenges for standard tokenizers.
- Diacritic handling: Diacritics play a crucial role in Arabic morphology and semantics, yet their inconsistent representation and omission in some datasets can lead to ambiguities.
- Ligature recognition: The presence of ligatures, where multiple individual characters combine into a single glyph, further complicates word segmentation.
- Fine-grained tokenization: Utilizing the subword-level capabilities of SentencePiece, AraNizer achieves a deeper granularity than simple word segmentation, preserving morphological and semantic information within tokens.
- Diacritic-aware processing: AraNizer incorporates diacritics into its tokenization process, ensuring their contribution to word meaning is preserved and reflected in LLM representations.
- Ligature handling: AraNizer is trained to recognize and appropriately segment ligatures, avoiding misinterpretations and maintaining accurate morpheme boundaries.
4.2. Description
5. Small-scale model: ArabianGPT-0.1B
5.1. Model Overview
5.2. Description
- Text and Position Embedding Layer: Adds information about the individual words and their positions in the sentence, which helps the model to better understand the meaning of the text.
- Masked Multi-Head Self-Attention: The core component of the GPT-1 architecture. It allows the model to attend to different parts of the input sentence simultaneously, using multiple heads that focus on different aspects of the words. This gives the model the ability to capture long-range dependencies and context in the text.
- Feed-Forward Sub-Layer: Adds non-linearity to the model, allowing it to learn more complex relationships between words.
- Normalization Layer: Helps to stabilize the training process and improve the model’s overall performance.
- Unidirectional Language Modeling: ArabianGPT, like GPT-2, predicts the probability of a sequence of words in a text, learning to generate coherent and contextually relevant text based on the input it receives.
- Transformer-based Architecture [18]: The use of 12 Transformer layers aids in capturing long-range dependencies in text, a crucial aspect for Arabic text which often involves complex sentence structures.
- Attention Mechanism: Each Model Attention Layer in ArabianGPT is designed to focus on different parts of the input text, allowing the model to generate more nuanced and accurate Arabic text.
- Adaptive Learning Rate: ArabianGPT leverages an adaptive learning rate during training, which facilitates efficient and effective learning from the large-scale Arabic corpus.
5.3. Specifications
5.4. Training procedure
5.5. Testing Examples

6. Medium-scale model: ArabianGPT-0.3B
6.1. Model Overview
6.2. Description
6.3. Specifications
6.4. Training Procedure
6.5. Testing examples

7. Few-shot evaluation
- ARC (25-shot) normalized: Assesses scientific reasoning capabilities using normalized metrics.
- HellaSwag (10-shot): Evaluates common-sense reasoning in narrative contexts.
- MMLU (5-shot): Measures understanding across various subjects.
- TruthfulQA (0-shot) with mc2: Gauges the ability to provide truthful and accurate responses.
8. Fine Tuning
8.1. Summarization
8.1.1. Hyperparameters
8.1.2. Results


| Model | Language | Average | ARC (25-shot) |
HellaSwag (10-shot) |
MMLU (5-shot) |
TruthfulQA (0-shot) |
|---|---|---|---|---|---|---|
| ArabianGPT-0.1B-FT | Arabic | 31.7 | 24.5 | 28.0 | 25.0 | 49.3 |
| ArabianGPT-0.3B-FT | Arabic | 31.8 | 27.3 | 31.5 | 25.1 | 43.2 |
8.2. Sentiment Analysis
8.3. Question Answering
9. Conclusion
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Acknowledgments
References
- Radford, A.; Narasimhan, K.; Salimans, T.; Sutskever, I. Improving language understanding by generative pre-training. Advances in neural information processing systems, 2018, Vol. 30.
- Radford, A.; Wu, J.; Child, R.; Luan, D.; Amodei, D.; Sutskever, I. Language models are unsupervised multitask learners. OpenAI blog 2019, 1, 9. [Google Scholar]
- ElJundi, O.; Antoun, W.; El Droubi, N.; Hajj, H.; El-Hajj, W.; Shaban, K. hulmona: The universal language model in Arabic. Proceedings of the fourth Arabic natural language processing workshop, 2019, pp. 68–77.
- Antoun, W.; Baly, F.; Hajj, H. Arabert: Transformer-based model for Arabic language understanding. arXiv preprint arXiv:2003.00104 2020.
- Brown, T.; Mann, B.; Ryder, N.; Subbiah, M.; Kaplan, J.D.; Dhariwal, P.; Amodei, D. Language models are few-shot learners. Advances in neural information processing systems, 2020, Vol. 33, pp. 1877–1901.
- Sengupta, N.; Sahu, S.K.; Jia, B.; Katipomu, S.; Li, H.; Koto, F.; Xing, E. Jais and jais-chat: Arabic-centric foundation and instruction-tuned open generative large language models. arXiv preprint arXiv:2308.16149 2023.
- Huang, H.; Yu, F.; Zhu, J.; Sun, X.; Cheng, H.; Song, D.; Chen, Z.; Alharthi, A.; An, B.; Liu, Z.; others. AceGPT, Localizing Large Language Models in Arabic. arXiv preprint arXiv:2309.12053 2023.
- Lan, W.; Chen, Y.; Xu, W.; Ritter, A. An empirical study of pre-trained transformers for Arabic information extraction. arXiv preprint arXiv:2004.14519 2020.
- Abdul-Mageed, M.; Elmadany, A.; Nagoudi, E.M.B. ARBERT & MARBERT: deep bidirectional transformers for Arabic. arXiv preprint arXiv:2101.01785 2020.
- Antoun, W.; Baly, F.; Hajj, H. AraELECTRA: Pre-training text discriminators for Arabic language understanding. arXiv preprint arXiv:2012.15516 2020.
- Inoue, G.; Alhafni, B.; Baimukan, N.; Bouamor, H.; Habash, N. The interplay of variant, size, and task type in Arabic pre-trained language models. arXiv preprint arXiv:2103.06678 2021.
- Nagoudi, E.M.B.; Elmadany, A.; Abdul-Mageed, M. AraT5: Text-to-text transformers for Arabic language generation. arXiv preprint arXiv:2109.12068 2021.
- Abdelali, A.; Hassan, S.; Mubarak, H.; Darwish, K.; Samih, Y. Pre-training bert on arabic tweets: Practical considerations. arXiv preprint arXiv:2102.10684 2021.
- Ghaddar, A.; Wu, Y.; Bagga, S.; Rashid, A.; Bibi, K.; Rezagholizadeh, M.; Xing, C.; Wang, Y.; Duan, X.; Wang, Z.; others. Revisiting Pre-trained Language Models and their Evaluation for Arabic Natural Language Processing. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, 2022, pp. 3135–3151.
- Najar, O.; Sibaee, S.; Ghouti, L.; Koubaa, A. AraNizer 0.1.8. https://pypi.org/project/aranizer, 2023.
- Kudo, T.; Richardson, J. Sentencepiece: A simple and language independent subword tokenizer and detokenizer for neural text processing. arXiv preprint arXiv:1808.06226 2018.
- Shibata, Y.; Kida, T.; Fukamachi, S.; Takeda, M.; Shinohara, A.; Shinohara, T.; Arikawa, S. Byte Pair encoding: A text compression scheme that accelerates pattern matching 1999.
- Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A.N.; Polosukhin, I. Attention is all you need. Advances in neural information processing systems, 2017, Vol. 30.






| Attribute | Value |
|---|---|
| Model Name | ArabianGPT-0.1 |
| Architecture | GPT-2 (GPT2LMHeadModel) |
| Number of Layers | 12 |
| Number of Model Attention Layers (MALs) | 12 |
| Number of Heads | 12 |
| Vocabulary Size | 64,000 |
| Tokenizer | Aranizer |
| Model Size | 134M parameters |
| Initial Learning Rate | |
| Final Learning Rate | |
| Context Window Size | 768 tokens |
| Dropout Ratio for Attention (attn_pdrop) | 0.1 |
| Dropout Ratio for Embeddings (embd_pdrop) | 0.1 |
| Dropout Probability for all Fully Connected Layers (resid_pdrop) | 0.1 |
| Parameter | Value |
|---|---|
| Model | ArabianGPT-0.1B (Base Model) |
| Hardware | 2 NVIDIA A100 GPU (80GB each) |
| Number of Examples | 7.5 million sequences (sequence length = 768 tokens) |
| Batch Size | 512 |
| Number of Steps | 313,500 |
| Number of Epochs | 6 |
| Training Duration | 3 days |
| Final Loss | 3.97 |
| Attribute | Value |
|---|---|
| Model Name | ArabianGPT-MID |
| Architecture | GPT-2 (GPT2LMHeadModel) |
| Number of Layers | 24 |
| Number of Model Attention Layers (MALs) | 16 |
| Number of Heads | 16 |
| Vocabulary Size | 64,000 |
| Model Size | 345M parameters |
| Initial Learning Rate | |
| Final Learning Rate | |
| Context Window Size | 1024 tokens |
| Dropout Ratio for Attention (attn_pdrop) | 0.1 |
| Dropout Ratio for Embeddings (embd_pdrop) | 0.1 |
| Dropout Probability for all Fully Connected Layers (resid_pdrop) | 0.1 |
| Parameter | Value |
|---|---|
| Model | ArabianGPT-0.3B (Medium Model) |
| Hardware | 4 NVIDIA A100 GPU (80GB each) |
| Number of Examples | 70 million sequences (sequence length = 1024) |
| Batch Size | 512 |
| Number of Steps | 1,009,500 |
| Number of Epochs | 4.97 |
| Training Duration | 4.1 days |
| Final Loss | 3.82 |
| Model | Language | Average | ARC (25-shot) |
HellaSwag (10-shot) |
MMLU (5-shot) |
TruthfulQA (0-shot) |
|---|---|---|---|---|---|---|
| Bloom-7b1 | Multilingual | 36.2 | 31.4 | 43.3 | 27.5 | 42.6 |
| Llama-7B | Multilingual | 32.1 | 24.6 | 30.9 | 28.0 | 45.1 |
| ArabianGPT-0.3B | Arabic | 32.7 | 24.3 | 28.4 | 25.7 | 52.5 |
| ArabianGPT-0.1B | Arabic | 31.9 | 24.0 | 26.6 | 25.4 | 51.8 |
| AraGPT-Base | Arabic | 31.7 | 24.6 | 27.5 | 25.1 | 49.5 |
| AraGPT-Medium | Arabic | 32.2 | 23.9 | 28.5 | 26.3 | 50.0 |
| Parameter | Value |
|---|---|
| Epochs | 3 |
| Learning rate | to |
| Batch | 8 |
| GPU | A100 |
| Global Step | 1500 |
| Train Loss | 0.203 |
| Training Samples | 375 |
| Testing Samples | 375 |
| Training Percentage | 70% |
| Testing Percentage | 30% |
| Tokenizer | 64K aranizer |
| Model | Accuracy |
|---|---|
| LLM-0.1B-Base | 56% |
| LLM-0.1B-Base (fine-tuned) | 95% |

| Model | Language | Average | ARC | |||
|---|---|---|---|---|---|---|
| (25-shot) | HellaSwag | |||||
| (10-shot) | MMLU | |||||
| (5-shot) | TruthfulQA | |||||
| (0-shot) | ||||||
| ArabianGPT-0.1B-FT | Arabic | 31.2 | 25.2 | 25.0 | 24.9 | 49.6 |
| ArabianGPT-0.3B-FT | Arabic | 31.4 | 24.8 | 27.6 | 24.9 | 48.2 |

Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2024 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).