Submitted:
19 October 2025
Posted:
20 October 2025
You are already at the latest version
Abstract
Keywords:
1. Introduction
1.1. Research Problem
- Assessing the sensitivity of model performance to different hyperparameter settings
- Identifying the key components of the fine-tuning process that contribute most to classification accuracy
- Determining the minimum and optimal dataset sizes required for effective fine-tuning
1.2. Research Questions
- RQ1: Which hyperparameters have the most significant impact on BERT fine-tuning effectiveness for the SRC task?
- RQ2: How does dataset size affect BERT performance in the SRC task?
- RQ3: What is the optimal combination of dataset size and hyperparameter settings for maximizing performance in SRC tasks?
1.3. Research Contribution
- Optimally tuned BERT model for SRC
- 2.
- Empirical analysis of fine-tuning factors
2. Related Work
3. Material and method
3.1. Datasets
3.2. Model Architecture
- 12 Transformer layers (blocks),
- 768-dimensional hidden states,
- 12 self-attention heads,
- approximately 110 million parameters (Devlin et al., 2019; Gardazi et al., 2025).
3.3. Fine-Tuning Process
4. Results & Discussion
4.1. Experimental Results
4.2. Fine-Tuned BERT Model Results for the SRC Task
4.3. Hyperparameter Ablation
4.3.1. Effect of Training Epochs on BERT Fine-Tuning Performance
4.3.2. Effect of Batch Size Selection on BERT Fine-Tuning Performance
4.3.3. Effect of Learning Rate on BERT Fine-Tuning Performance
5. Limitations and Future Work
- Scaling the model to handle larger software requirements datasets while maintaining efficiency and accuracy.
- Developing specialized models for fine-grained classification of individual NFR types.
- Creating adaptive tuning mechanisms that automatically adjust hyperparameters based on dataset characteristics.
6. Conclusions
Author Contributions
Conflicts of Interest
| 1 | Source Code found at: https://github.com/omercomail/Requirements-BERT
|
References
- Abad, Z.S.H. , Karras, O., Ghazi, P., Glinz, M., Ruhe, G., Schneider, K., 2017. What Works Better? A Study of Classifying Requirements, in: 2017 IEEE 25th International Requirements Engineering Conference (RE). Presented at the 2017 IEEE 25th International Requirements Engineering Conference (RE), pp. 496–501. [CrossRef]
- Airlangga, G. , 2024. Enhancing Software Requirements Classification with Semisupervised GAN-BERT Technique. J. Electr. Comput. Eng. 2024, 4955691. [CrossRef]
- Baker, C. , Deng, L., Chakraborty, S., Dehlinger, J., 2019. Automatic Multi-class Non-Functional Software Requirements Classification Using Neural Networks, in: 2019 IEEE 43rd Annual Computer Software and Applications Conference (COMPSAC). Presented at the 2019 IEEE 43rd Annual Computer Software and Applications Conference (COMPSAC), pp. 610–615. [CrossRef]
- Devlin, J. , Chang, M.-W., Lee, K., Toutanova, K., 2019. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding, in: Burstein, J., Doran, C., Solorio, T. (Eds.), Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers). Presented at the NAACL-HLT 2019, Association for Computational Linguistics, Minneapolis, Minnesota, pp. 4171–4186. [CrossRef]
- EzzatiKarami, M. , Madhavji, N.H., 2021. Automatically Classifying Non-functional Requirements with Feature Extraction and Supervised Machine Learning Techniques: A Research Preview, in: Dalpiaz, F., Spoletini, P. (Eds.), Requirements Engineering: Foundation for Software Quality. Springer International Publishing, Cham, pp. 71–78. [CrossRef]
- Gardazi, N.M. , Daud, A., Malik, M.K., Bukhari, A., Alsahfi, T., Alshemaimri, B., 2025. BERT applications in natural language processing: a review. Artif. Intell. Rev. 58, 166. [CrossRef]
- Gkouti, N. , Malakasiotis, P., Toumpis, S., Androutsopoulos, I., 2024. Should I try multiple optimizers when fine-tuning pre-trained Transformers for NLP tasks? Should I tune their hyperparameters? [CrossRef]
- Haque, Md.A. , Abdur Rahman, Md., Siddik, M.S., 2019. Non-Functional Requirements Classification with Feature Extraction and Machine Learning: An Empirical Study, in: 2019 1st International Conference on Advances in Science, Engineering and Robotics Technology (ICASERT). Presented at the 2019 1st International Conference on Advances in Science, Engineering and Robotics Technology (ICASERT), pp. 1–5. [CrossRef]
- Hey, T. , Keim, J., Koziolek, A., Tichy, W.F., 2020a. NoRBERT: Transfer Learning for Requirements Classification, in: 2020 IEEE 28th International Requirements Engineering Conference (RE). Presented at the 2020 IEEE 28th International Requirements Engineering Conference (RE), IEEE, Zurich, Switzerland, pp. 169–179. [CrossRef]
- Hey, T. , Keim, J., Koziolek, A., Tichy, W.F., 2020b. NoRBERT: Transfer Learning for Requirements Classification, in: 2020 IEEE 28th International Requirements Engineering Conference (RE). Presented at the 2020 IEEE 28th International Requirements Engineering Conference (RE), IEEE, Zurich, Switzerland, pp. 169–179. [CrossRef]
- Jindal, R. , Malhotra, R., Jain, A., Bansal, A., 2021. Mining Non-Functional Requirements using Machine Learning Techniques. E-Inform. Softw. Eng. J. 15. [CrossRef]
- Kaur, K. , Kaur, P., 2024. The application of AI techniques in requirements classification: a systematic mapping. Artif. Intell. Rev. 57. [CrossRef]
- Kaur, K. , Kaur, P., 2023. BERT-CNN: Improving BERT for Requirements Classification using CNN. Procedia Comput. Sci. 218, 2604–2611. [CrossRef]
- Kici, D. , Malik, G., Cevik, M., Parikh, D., Başar, A., 2021. A BERT-based transfer learning approach to text classification on software requirements specifications. Proc. Can. Conf. Artif. Intell. [CrossRef]
- Kingma, D.P. , Ba, J., 2017. Adam: A Method for Stochastic Optimization. [CrossRef]
- Kurtanović, Z. , Maalej, W., 2017. Automatically Classifying Functional and Non-functional Requirements Using Supervised Machine Learning, in: 2017 IEEE 25th International Requirements Engineering Conference (RE). Presented at the 2017 IEEE 25th International Requirements Engineering Conference (RE), pp. 490–495. [CrossRef]
- Li, G. , Zheng, C., Li, M., Wang, H., 2022. Automatic Requirements Classification Based on Graph Attention Network. IEEE Access 10, 30080–30090. [CrossRef]
- Luo, X. , Xue, Y., Xing, Z., Sun, J., 2022. PRCBERT: Prompt Learning for Requirement Classification using BERT-based Pretrained Language Models, in: Proceedings of the 37th IEEE/ACM International Conference on Automated Software Engineering. Presented at the ASE ’22: 37th IEEE/ACM International Conference on Automated Software Engineering, ACM, Rochester MI USA, pp. 1–13. [CrossRef]
- Navarro-Almanza, R. , Juarez-Ramirez, R., Licea, G., 2017. Towards Supporting Software Engineering Using Deep Learning: A Case of Software Requirements Classification, in: 2017 5th International Conference in Software Engineering Research and Innovation (CONISOFT). Presented at the 2017 5th International Conference in Software Engineering Research and Innovation (CONISOFT), pp. 116–120. [CrossRef]
- PROMISE NFR, 2005. https://github.com/AleksandarMitrevski/se-requirements-classification/blob/master/0-datasets/PROMISE_exp/PROMISE_exp.arff.
- PURE: A Dataset of Public Requirements Documents [WWW Document], 2017. URL https://ieeexplore.ieee.org/document/8049173 (accessed 9.16.24).
- Shreda, Q.A. , Hanani, A.A., 2025. Identifying Non-Functional Requirements From Unconstrained Documents Using Natural Language Processing and Machine Learning Approaches. IEEE Access 13, 124159–124179. [CrossRef]
- Sommerville, I. , Sawyer, P., 1999. Requirements Engineering: A Good Practice Guide. Wiley, Chichester, Eng. ; New York.
- Sonali, S. , Thamada, S., 2024. FR_NFR_dataset. [CrossRef]
- Taj, S. , Daudpota, S.M., Imran, A.S., Kastrati, Z., 2025. Aspect-based sentiment analysis for software requirements elicitation using fine-tuned Bidirectional Encoder Representations from Transformers and Explainable Artificial Intelligence. Eng. Appl. Artif. Intell. 151, 110632. [CrossRef]
- Wiegers, K.E. , Beatty, J., Wiegers, K.E., 2013. Software requirements, 3. ed. [fully updated and expanded]. ed, Best practices. Microsoft Press, Redmond, Wash.
- Wolf, T. , Debut, L., Sanh, V., Chaumond, J., Delangue, C., Moi, A., Cistac, P., Rault, T., Louf, R., Funtowicz, M., Davison, J., Shleifer, S., Platen, P. von, Ma, C., Jernite, Y., Plu, J., Xu, C., Scao, T.L., Gugger, S., Drame, M., Lhoest, Q., Rush, A.M., 2020. HuggingFace’s Transformers: State-of-the-art Natural Language Processing. [CrossRef]






| Requirement | Class | Dataset |
| “The system shall refresh the display every 60 seconds.” | NFR |
PROMISE |
| “The system shall filter data by: Venues and Key Events.” | FR | |
| “The app shall run on a smart phone with Android OS least version 2.3 .” | NFR |
FR_NFR_dataset |
| “User shall be able to add, modify, or remove user data, with changes reflected successfully.” | FR |
| Parameter | PROMISE NFR dataset | FR NFR dataset |
|---|---|---|
| Optimizer | Adam | Adam |
| Max Seq Length | 256 | 256 |
| Dropout Rate | 0.2 | 0.2 |
| Batch Size | 8 | 16 |
| Learning Rate | 1e−5 | 1e−5 |
| Epochs | 8 | 4 |
| Study | Technique | FR | NFR | F1 Avg | ||||
| P | R | F | P | R | F | |||
| (Kurtanović and Maalej, 2017) | SVM (word features) | 0.92 | 0.93 | 0.93 | 0.93 | 0.92 | 0.92 | 0.93 |
| (Li et al., 2022) | Naïve Bayes (TF-IDF) | 0.80 | 0.93 | 0.86 | 0.95 | 0.85 | 0.90 | 0.88 |
| (Hey et al., 2020b) | BERT classfier | 0.92 | 0.88 | 0.90 | 0.92 | 0.95 | 0.93 | 0.92 |
| (Li et al., 2022) | Bert+ GAT | 0.94 | 0.90 | 0.92 | 0.94 | 0.98 | 0.96 | 0.94 |
| (Abad et al., 2017) | Processed data | 0.90 | 0.97 | 0.93 | 0.98 | 0.93 | 0.95 | 0.94 |
| (Luo et al., 2022) | PRCBERT with RoBERTa-large | 0.92 | 0.95 | 0.93 | 0.94 | 0.96 | 0.95 | 0.94 |
| Our model | Bert (fine-tuned) | 0.98 | 0.98 | 0.98 | 0.99 | 0.99 | 0.99 |
0.99 |
| (a) | ||||
| Precision | Recall | F1-Score | Support | |
| FR | 0.98 | 0.98 | 0.98 | 255 |
| NFR | 0.99 | 0.99 | 0.99 | 370 |
| Accuracy | 0.98 | 625 | ||
| Macro Avg | 0.98 | 0.98 | 0.98 | 625 |
| Weighted Avg | 0.98 | 0.98 | 0.98 | 625 |
| (b) | ||||
| Precision | Recall | F1-Score | Support | |
| FR | 0.97 | 0.98 | 0.98 | 3964 |
| NFR | 0.96 | 0.95 | 0.96 | 2153 |
| Accuracy | 0.97 | 6117 | ||
| Macro Avg | 0.97 | 0.97 | 0.97 | 6117 |
| Weighted Avg | 0.97 | 0.97 | 0.97 | 6117 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2025 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).