Submitted:
22 January 2025
Posted:
23 January 2025
You are already at the latest version
Abstract
Keywords:
1. Introduction
2. Related Work
3. Methodology
Data Collection
Data Preprocessing
- Removal of HTML Tags: HTML elements present in the scraped job descriptions were removed to retain only the core textual content.
- Text Standardization: Text was converted to lowercase to ensure uniformity across the dataset.
- Stopword Removal: Common stopwords were removed using the Natural Language Toolkit (NLTK) [23]. Additionally, domain-specific stopwords were included to improve relevance.
- Special Characters and Non-ASCII Removal: Non-ASCII characters and special symbols were filtered out to maintain compatibility and readability.
- Handling Numbers: While irrelevant digits (e.g., phone numbers) were removed, meaningful numeric data such as percentages, monetary values, and years were preserved.
- Email Address Removal: To maintain privacy and reduce noise, email addresses were stripped from the data.
- Punctuation Removal: The punctuation was removed using tokenization techniques to focus on alphanumeric content.
Phase 1: Qualitative Human Evaluation
Phase 2: Quantitative Evaluation and Model Selection
4. Results and Evaluation
5. Discussion
6. Conclusions
Funding
Institutional Review Board Statement
Informed Consent Statement
References
- Albaroudi, E., Mansouri, T., & Alameer, A. (2024). A Comprehensive Review of AI Techniques for Addressing Algorithmic Bias in Job Hiring. AI, 5(1), 383-404.
- Strohmeier, S., & Piazza, F. (2015). Artificial intelligence techniques in human resource management—a conceptual exploration. Intelligent Techniques in Engineering Management: Theory and Applications, 149-172.
- Le, Q., & Mikolov, T. (2014, June). Distributed representations of sentences and documents. In International conference on machine learning (pp. 1188-1196). PMLR.
- AL-Qassem, A. H., Agha, K., Vij, M., Elrehail, H., & Agarwal, R. (2023, March). Leading Talent Management: Empirical investigation on Applicant Tracking System (ATS) on e-Recruitment Performance. In 2023 International Conference on Business Analytics for Technology and Security (ICBATS) (pp. 1-5). IEEE.
- Li, C., Fisher, E., Thomas, R., Pittard, S., Hertzberg, V., & Choi, J. D. (2020). Competence-level prediction and resume & job description matching using context-aware transformer models. arXiv. arXiv:2011.02998.
- Parasurama, P., & Sedoc, J. (2021). Degendering Resumes for Fair Algorithmic Resume Screening. arXiv. arXiv:2112.08910.
- Deshmukh, A., & Raut, A. (2024). Enhanced Resume Screening for Smart Hiring Using Sentence-Bidirectional Encoder Representations from Transformers (S-BERT). International Journal of Advanced Computer Science & Applications, 15(8).
- James, V., Kulkarni, A., & Agarwal, R. (2022, December). Resume Shortlisting and Ranking with Transformers. In International Conference on Intelligent Systems and Machine Learning (pp. 99-108). Cham: Springer Nature Switzerland.
- Heakl, A., Mohamed, Y., Mohamed, N., Elsharkawy, A., & Zaky, A. (2024). ResuméAtlas: Revisiting Resume Classification with Large-Scale Datasets and Large Language Models. Procedia Computer Science, 244, 158-165.
- Vishaline, A. R., Kumar, R. K. P., VVNS, S. P., Vignesh, K. V. K., & Sudheesh, P. (2024, July). An ML-based Resume Screening and Ranking System. In 2024 International Conference on Signal Processing, Computation, Electronics, Power and Telecommunication (IConSCEPT) (pp. 1-6). IEEE.
- Minvielle, L. (2025, January 8). How AI is shaping applicant tracking systems. Geekflare. https://geekflare.com/guide/ats-and-ai/.
- Lavi, D., Medentsiy, V., & Graus, D. (2021). consultantbert: Fine-tuned siamese sentence-bert for matching jobs and job seekers. arXiv. arXiv:2109.06501.
- Skondras, P., Zervas, P., & Tzimas, G. (2023). Generating Synthetic Resume Data with Large Language Models for Enhanced Job Description Classification. Future Internet, 15(11), 363.
- Pendhari, H., Rodricks, S., Patel, M., Emmatty, S., & Pereira, A. (2023, December). Resume Screening using Machine Learning. In 2023 International Conference on Data Science, Agents & Artificial Intelligence (ICDSAAI) (pp. 1-5). IEEE.
- Bharadwaj, S., Varun, R., Aditya, P. S., Nikhil, M., & Babu, G. C. (2022, July). Resume Screening using NLP and LSTM. In 2022 International Conference on Inventive Computation Technologies (ICICT) (pp. 238-241). IEEE.
- Kavas, H., Serra-Vidal, M., & Wanner, L. (2024, August). Using Large Language Models and Recruiter Expertise for Optimized Multilingual Job Offer–Applicant CV Matching. In Proceedings of the Thirty-Third International Joint Conference on Artificial Intelligence (pp. 8696-8699).
- Ransing, R., Mohan, A., Emberi, N. B., & Mahavarkar, K. (2021, December). Screening and Ranking Resumes using Stacked Model. In 2021 5th International Conference on Electrical, Electronics, Communication, Computer Technologies and Optimization Techniques (ICEECCOT) (pp. 643-648). IEEE.
- Zaroor, A., Maree, M., & Sabha, M. (2017, November). JRC: a job post and resume classification system for online recruitment. In 2017 IEEE 29th International Conference on Tools with Artificial Intelligence (ICTAI) (pp. 780-787). IEEE.
- Nasser, S., Sreejith, C., & Irshad, M. (2018, July). Convolutional Neural Network with Word Embedding Based Approach for Resume Classification. In 2018 International Conference on Emerging Trends and Innovations In Engineering And Technological Research (ICETIETR) (pp. 1-6). IEEE.
- Koh, N. H., Plata, J., & Chai, J. (2023). BAD: BiAs Detection for Large Language Models in the context of candidate screening. arXiv preprint arXiv:2305.10407. arXiv:2305.10407.
- Gan, C., Zhang, Q., & Mori, T. (2024). Application of LLM agents in recruitment: A novel framework for resume screening. arXiv preprint arXiv:2401.08315. arXiv:2401.08315.
- Bevara, R. V. K. (2024). Scraping LinkedIn for Job Descriptions: A Data Scientist’s Guide. Medium. Retrieved from https://medium.com/@ravivarmakumarbevara/scraping-linkedin-for-job-descriptions-a-data-scientists-guide-337dc91de618.
- Bird, S., Klein, E., & Loper, E. (2009). Natural Language Processing with Python. O’Reilly Media, Inc.
- Willis, K. L., Holmes, B., & Burwell, N. (2022). Improving Graduate Student Recruitment, Retention, and Professional Development During COVID-19. American Journal of Educational Research, 10(2), 81–84. [CrossRef]
- Hung, S.-P., & Huang, H.-Y. (2022). Forced-Choice Ranking Models for Raters’ Ranking Data. Journal of Educational and Behavioral Statistics, 47(5), 603–634. [CrossRef]
- Shahriar, S., Lund, B. D., Mannuru, N. R., Arshad, M. A., Hayawi, K., Bevara, R. V. K., ... & Batool, L. (2024). Putting GPT-4.0 to the Sword: A Comprehensive Evaluation of Language, Vision, Speech, and Multimodal Proficiency. Applied Sciences, 14(17), 7782.
- Bevara, R. V. K., Mannuru, N. R., Karedla, S. P., & Xiao, T. (2024). Scaling Implicit Bias Analysis Across Transformer-Based Language Models Through Embedding Association Test and Prompt Engineering. Applied Sciences, 14(8), 3483.
- Bevara, R. V. K., Wagenvoord, I., Hosseini, F., Sharma, H., Nunna, V., & Xiao, T. (2024). Census2Vec: Enhancing Socioeconomic Predictive Models with Geo-Embedded Data. Intelligent Systems Conference, pp. 626-640. Springer Nature Switzerland.
- Bevara, R. V. K., Xiao, T., Hosseini, F., & Ding, J. (2023, October). Bias Analysis in Language Models using An Association Test and Prompt Engineering. In 2023 IEEE 23rd International Conference on Software Quality, Reliability, and Security Companion (QRS-C) (pp. 356-363). IEEE.






| Model | SVM | Random Forest | Gradient Boosting | Logistic Regression |
|---|---|---|---|---|
| BERT | 78.5% | 94.5% | 91.3% | 95.2% |
| RoBERTa | 70.5% | 93.5% | 92.5% | 94.5% |
| DistilBERT | 83.0% | 94.5% | 89.5% | 95.5% |
| GPT-4.0 | 90.5% | 91.0% | 55.0% | 89.5% |
| Gemini | 91.0% | 94.0% | 73.5% | 90.5% |
| Llama | 83.5% | 95.5% | 95.5% | 95.5% |
| Domain | nDCG (ATS) | nDCG (Proposed) |
|---|---|---|
| Data Science | 0.88 | 0.89 |
| Health and Fitness | 0.83 | 0.87 |
| Mechanical Engineering | 0.82 | 0.95 |
| Operations Manager | 0.90 | 0.86 |
| Software Testing | 0.96 | 0.90 |
| Domain | RBO (ATS) | RBO (Proposed) |
|---|---|---|
| Data Science | 0.84 | 0.92 |
| Health and Fitness | 0.69 | 0.80 |
| Mechanical Engineering | 0.88 | 1.00 |
| Operations Manager | 0.86 | 0.97 |
| Software Testing | 1.00 | 0.96 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2025 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).