Submitted:
14 March 2025
Posted:
14 March 2025
You are already at the latest version
Abstract
The Question Generation System (QGS) for Information Technology (IT) education, designed to create, evaluate, and improve Multiple-Choice Questions (MCQs) using Knowledge graphs (KGs) and Large Language Models (LLMs), encounters three major needs: ensuring the generation of contextually relevant and accurate distractors, enhancing the diversity of generated questions, and balancing the higher-order thinking of questions to match various learning levels. To address these needs, we proposed a multi-agent system named Multi-Examiner, which integrates knowledge graphs, domain-specific search tools, and local knowledge bases, categorized according to Bloom’s taxonomy, to enhance the contextual relevance, diversity, and higher-order thinking of automatically generated information technology multiple-choice questions. We designed a multidimensional evaluation rubric to assess the semantic coherence, answer correctness, question validity, distractor relevance, question diversity, and higher-order thinking, and applied it to questions generated for six knowledge points from the second chapter of the "Information Systems and Society" textbook using both the Multi-Examiner system and GPT-4, alongside real exam questions, evaluated by 30 high school IT teachers. The results demonstrated that: (i) overall, questions generated by the Multi-Examiner system outperformed those generated by GPT-4 across all dimensions and closely matched the quality of human-crafted questions in several dimensions; (ii) domain-specific search tools significantly enhanced the diversity of questions generated by Multi-Examiner; (iii) GPT-4 generated better questions for knowledge points at the "remembering" and "understanding" levels, while Multi-Examiner significantly improved the higher-order thinking of questions for "evaluating" and "creating" levels. This study highlights the potential of multi-agent systems in advancing question generation.
Keywords:
1. Introduction
2. Related Work
2.1. QGS Based on Knowledge Graphs and Knowledge Bases
2.2. QGS Based on LLMs and Intelligent Agents
2.3. Application of Educational Objective Taxonomies in QGS
2.4. Exam Question Evaluation Scale
2.5. Synthesis and Research Gaps
3. Methodology
3.1. Research Design
3.2. System Development
3.2.1. Knowledge Graph Construction
3.2.2. Knowledge Base Construction
3.2.3. Multi-Examiner System Design
| Algorithm 1 Question Validation and Finalization Process by Revisor |
![]() |
3.3. Experimental Design
3.3.1. Participants
3.3.2. Experimental Materials and Procedures
3.3.3. Measures and Instruments
3.3.4. Data Analysis
3.4. Ethical Considerations
4. Results
4.1. Analysis of the Contextual Relevance of Distractors (RQ1)
4.1.1. Descriptive Statistical Analysis
4.1.2. Two-Way Analysis of Variance (ANOVA)
4.1.3. Performance Differences by Generation Method Across Different Knowledge Types
4.2. Analysis of Enhancing Question Diversity (RQ2)
4.2.1. Descriptive Statistical Analysis
4.2.2. Multivariate Analysis of Variance (MANOVA)
4.2.3. Univariate Analysis of Variance (ANOVA) Follow-up Tests
4.2.4. Performance Differences by Generation Method Across Evaluation Dimensions
4.3. Analysis of the Effectiveness of the Assessment System in Generating Higher-Order Thinking Questions (RQ3)
4.3.1. Descriptive Statistical Analysis
4.3.2. Two-Way Analysis of Variance (ANOVA)
4.3.3. Post-Hoc Test Analysis
4.3.4. Differences in Performance Across Cognitive Levels by Generation Method
4.3.5. Quality Analysis of Higher-Order Thinking Questions
5. Discussions
5.1. Discussion of Distractor Contextual Relevance and Generation Method Effectiveness (RQ1)
5.2. Discussion on Enhancing Question Diversity and Cognitive Challenge through Automated Generation Methods (RQ2)
5.3. Evaluating the Effectiveness of Automated Systems in Generating Higher-Order Thinking Questions in K-12 IT Education (RQ3)
6. Conclusion
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Conflicts of Interest
References
- Abd-Alrazaq, A.; AlSaad, R.; Alhuwail, D.; Ahmed, A.; Healy, P. M.; Latifi, S.; Aziz, S.; Damseh, R.; Alabed, A. S.; Sheikh, J. Large language models in medical education: opportunities, challenges, and future directions; JMIR Medical Education, 2023, 9, 1.
- Shoufan, A. Can students without prior knowledge use ChatGPT to answer test questions? An empirical study; ACM Transactions on Computing Education, 2023. [CrossRef]
- Abulibdeh, A.; Zaidan, E.; Abulibdeh, R. Navigating the confluence of artificial intelligence and education for sustainable development in the era of industry 4.0: Challenges, opportunities, and ethical dimensions; Journal of Cleaner Production, 2024, 140527.
- Almufarreh, A.; Mohammed, N. K.; Saeed, M. N. Academic teaching quality framework and performance evaluation using machine learning; Applied Sciences, 2023, 13, 5.
- Anderson, L. W.; Krathwohl, D. A taxonomy for learning, teaching, and assessing: A revision of Bloom’s; Addison Wesley Longman, Inc.: USA, 2001.
- Bahroun, Z.; Anane, C.; Ahmed, V.; Zacca, A. Transforming education: A comprehensive review of generative artificial intelligence in educational settings through bibliometric and content analysis; Sustainability, 2023, 15, 17.
- Chang, Y.; Wang, X.; Wang, J.; Wu, Y.; Yang, L.; Zhu, K.; Chen, H.; Yi, X.; Wang, C.; Wang, Y. A survey on evaluation of large language models; ACM Transactions on Intelligent Systems and Technology, 2024, 15, 3.
- Chen, G.; Yang, J.; Hauff, C.; Houben, G. LearningQ: A LargeScale Dataset for Educational Question Generation; ICWSM, 2018, 12(1). [CrossRef]
- Chen, J.; Liu, Z.; Huang, X.; Wu, C.; Liu, Q.; Jiang, G.; Pu, Y.; Lei, Y.; Chen, X.; Wang, X. When large language models meet personalization: Perspectives of challenges and opportunities; World Wide Web, 2024, 27, 4.
- Conklin, J.; Anderson, L. W.; Krathwohl, D.; Airasian, P.; Cruikshank, K. A.; Mayer, R. E.; Pintrich, P.; Raths, J.; Wittrock, M. C. Educational Horizons, 2005, 83(3), 154–159. Available online: http://www.jstor.org/stable/42926529.
- Dienichieva, O. I.; Komogorova, M. I.; Lukianchuk, S. F.; Teletska, L. I.; Yankovska, I. M. From reflection to self-assessment: Methods of developing critical thinking in students; International Journal of Computer Science & Network Security, 2024, 24, 7.
- Espartinez, A. S. Exploring student and teacher perceptions of ChatGPT use in higher education: A Q-Methodology study; Computers and Education: Artificial Intelligence, 2024, 7, 100264.
- Folk, A.; Blocksidge, K.; Hammons, J.; Primeau, H. Building a bridge between skills and thresholds: Using Bloom’s to develop an information literacy taxonomy; Journal of Information Literacy, 2024, 18, 1.
- Gezer, M.; Oner Sunkur, M.; Sahin, I. F. An evaluation of the exam questions of social studies course according to revised Bloom’s taxonomy; Education Sciences & Psychology, 2014, 28(2).
- Goyal, M.; Mahmoud, Q. H. A systematic review of synthetic data generation techniques using generative AI; Electronics, 2024, 13, 17.
- Graesser, A. C.; Lu, S.; Jackson, G. T.; Mitchell, H. H.; Ventura, M.; Olney, A.; Louwerse, M. M. AutoTutor: A tutor with dialogue in natural language; Behavior Research Methods, Instruments, & Computers, 2004, 36(2), 180–192. [CrossRef]
- Guo, T.; Chen, X.; Wang, Y.; Chang, R.; Pei, S.; Chawla, N. V.; Wiest, O.; Zhang, X. Large language model based multi-agents: A survey of progress and challenges; ArXiv.org, 2024. [CrossRef]
- Hadi, M. U.; Tashi, A.; Shah, A.; Qureshi, R.; Muneer, A.; Irfan, M.; Zafar, A.; Shaikh, M. B.; Akhtar, N.; Wu, J. Large language models: A comprehensive survey of its applications, challenges, limitations, and future prospects; Authorea Preprints, 2024.
- Haladyna, T. M.; Downing, S. M. Validity of a taxonomy of multiple-choice item-writing rules; Applied Measurement in Education, 1989, 2(1), 51–78. [CrossRef]
- Halkiopoulos, C.; Gkintoni, E. Leveraging AI in e-learning: Personalized learning and adaptive assessment through cognitive neuropsychology—A systematic analysis; Electronics, 2024, 13, 18.
- Han, K.; Gardent, C. Generating and Answering Simple and Complex Questions from Text and from Knowledge Graphs; Hal.science, 2023. Available online: https://hal.science/hal-04369868.
- Hang, C. N.; Tan, C. W.; Yu, P.-D. Mcqgen: A large language model-driven mcq generator for personalized learning; IEEE Access, 2024; IEEE.
- Heilman, M.; Smith, N. A. Good Question! Statistical Ranking for Question Generation; In Proceedings of Human Language Technologies: The 2010 Annual Conference of the North American Chapter of the Association for Computational Linguistics, 2010, pp. 609–617.
- Hofer, M.; Obraczka, D.; Saeedi, A.; Köpcke, H.; Rahm, E. Construction of knowledge graphs: Current state and challenges; Information, 2024, 15, 8.
- Hwang, K.; Challagundla, S.; Alomair, M.; Chen, L. K.; Choa, F. S. Towards AI-assisted multiple choice question generation and quality evaluation at scale: Aligning with Bloom’s Taxonomy; Workshop on Generative AI for Education, 2023.
- Hwang, K.; Wang, K.; Alomair, M.; Choa, F.; Chen, L. K. Towards Automated Multiple Choice Question Generation and Evaluation: Aligning with Bloom’s Taxonomy; In A. M. Olney, I. Chounta, Z. Liu, O. C. Santos, B. I. Ibert (Eds.), Artificial Intelligence in Education, Springer Nature Switzerland, 2024, pp. 389–396.
- Jia, Z.; Pramanik, S.; Saha Roy, R.; Weikum, G. Complex Temporal Question Answering on Knowledge Graphs; In Proceedings of the 30th ACM International Conference on Information & Knowledge Management, 2021, pp. 792–802. [CrossRef]
- Koenig, N.; Tonidandel, S.; Thompson, I.; Albritton, B.; Koohifar, F.; Yankov, G.; Speer, A.; Jay, Gibson, C.; Frost, C. Improving measurement and prediction in personnel selection through the application of machine learning; Personnel Psychology, 2023, 76, 4.
- Kong, S.-C.; Yang, Y. A human-centred learning and teaching framework using generative artificial intelligence for self-regulated learning development through domain knowledge learning in K–12 settings; IEEE Transactions on Learning Technologies, 2024; IEEE.
- Krathwohl, D. R. A Revision of Bloom’s Taxonomy: An Overview; Theory into Practice, 2002, 41(4), 212–218. [CrossRef]
- Kurdi, G.; Leo, J.; Parsia, B.; Sattler, U.; AlEmari, S. A Systematic Review of Automatic Question Generation for Educational Purposes; International Journal of Artificial Intelligence in Education, 2020, 30(1), 121–204. [CrossRef]
- Lai, H.; Nissim, M. A survey on automatic generation of figurative language: From rule-based systems to large language models; ACM Computing Surveys, 2024, 56, 10.
- Leite, B.; Cardoso, H. L. Do rules still rule? Comprehensive evaluation of a rule-based question generation system; 2023, pp. 27–38.
- Li, W.; Li, L.; Xiang, T.; Liu, X.; Deng, W.; Garcia, N. Can multiple-choice questions really be useful in detecting the abilities of LLMs?; ArXiv (Cornell University), 2024. [CrossRef]
- Liang, K.; Meng, L.; Liu, M.; Liu, Y.; Tu, W.; Wang, S.; Zhou, S.; Liu, X.; Sun, F.; He, K. A survey of knowledge graph reasoning on graph types: Static, dynamic, and multi-modal; IEEE Transactions on Pattern Analysis and Machine Intelligence, 2024; IEEE.
- Liang, W.; Meo, P. D.; Tang, Y.; Zhu, J. A survey of multi-modal knowledge graphs: Technologies and trends; ACM Computing Surveys, 2024, 56, 11.
- Liu, Q.; Han, S.; Cambria, E.; Li, Y.; Kwok, K. PrimeNet: A framework for commonsense knowledge representation and reasoning based on conceptual primitives; In Cognitive Computation, Springer, 2024, pp. 1–28.
- Liu, Y.; Yao, Y.; Ton, J.-F.; Zhang, X.; Guo, R.; Cheng, H.; Klochkov, Y.; Taufiq, M. F.; Li, H. Trustworthy LLMs: A survey and guideline for evaluating large language models’ alignment; ArXiv, 2023. Available online: https://arxiv.org/abs/2308.05374.
- Moore, S.; Schmucker, R.; Mitchell, T.; Stamper, J. Automated generation and tagging of knowledge components from multiple-choice questions; 2024, pp. 122–133.
- Mulla, N.; Gharpure, P. Automatic question generation: A review of methodologies, datasets, evaluation metrics, and applications; Progress in Artificial Intelligence, 2023, 12(1), 1–32. [CrossRef]
- Naseer, F.; Khalid, M. U.; Ayub, N.; Rasool, A.; Abbas, T.; Afzal, M. W. Automated assessment and feedback in higher education using generative AI; IGI Global, 2024, pp. 433–461.
- Pan, S.; Luo, L.; Wang, Y.; Chen, C.; Wang, J.; Wu, X. Unifying large language models and knowledge graphs: A roadmap; IEEE Transactions on Knowledge and Data Engineering, 2024; IEEE.
- Pan, X.; Li, X.; Li, Q.; Hu, Z.; Bao, J. Evolving to multi-modal knowledge graphs for engineering design: State-of-the-art and future challenges; Journal of Engineering Design, 2024, pp. 1–40; Taylor & Francis.
- Paulheim, H. Knowledge Graph Refinement: A Survey of Approaches and Evaluation Methods; Semantic Web, 2016, 8(3), 489–508. [CrossRef]
- Rashid, M.; Torchiano, M.; Rizzo, G.; Mihindukulasooriya, N.; Corcho, O. A quality assessment approach for evolving knowledge bases; Semantic Web, 2019, 10, 2.
- Rodrigues, L.; Pereira, F. D.; Cabral, L.; Gašević, D.; Ramalho, G.; Mello, R. F. Assessing the quality of automatic-generated short answers using GPT-4; Computers and Education: Artificial Intelligence, 2024, 7, 100248.
- Rodriguez, M. C. Three Options Are Optimal for Multiple-Choice Items: A Meta-Analysis of 80 Years of Research; Educational Measurement: Issues and Practice, 2005, 24(2), 3–13. [CrossRef]
- Shahriar, S.; Lund, B. D.; Reddy, M. N.; Arshad, M. A.; Hayawi, K.; Ravi, B.; Mannuru, A.; Batool, L. Putting GPT-4o to the sword: A comprehensive evaluation of language, vision, speech, and multimodal proficiency; Applied Sciences, 2024, 14, 17.
- Shuraiqi, A.; Abdulrahman, A. A.; Masters, K.; Zidoum, H.; AlZaabi, A. Automatic generation of medical case-based multiple-choice questions (MCQs): A review of methodologies, applications, evaluation, and future directions; Big Data and Cognitive Computing, 2024, 8, 10.
- Singh, M.; Patvardhan, C.; Vasantha, L. C. Does ChatGPT spell the end of automatic question generation research?; IEEE, 2023, pp. 1–6.
- Sun, Y.; Yang, Y.; Fu, W. Exploring synergies between causal models and Large-Language models for enhanced understanding and inference; 2024, pp. 1–8.
- Tahri, C. Leveraging modern information seeking on research papers for real-world knowledge integration applications: An empirical study; 2023.
- Tao, Y.; Viberg, O.; Baker, R. S.; Kizilcec, R. F. Cultural bias and cultural alignment of large language models; PNAS Nexus, 2024, 3, 9.
- Vistorte, A. O. R.; Deroncele-Acosta, A.; Ayala, J. L. M.; Barrasa, A.; López-Granero, C.; Martí-González, M. Integrating artificial intelligence to assess emotions in learning environments: A systematic literature review; Frontiers in Psychology, 2024, 15, 1387089.
- Wang, X.; Yang, Q.; Qiu, Y.; Liang, J.; He, Q.; Gu, Z.; Xiao, Y.; Wang, W. Knowledgpt: Enhancing large language models with retrieval and storage access on knowledge bases; ArXiv, 2023. Available online: https://arxiv.org/abs/2308.11761.
- Wei, X. Evaluating ChatGPT-4 and ChatGPT-4o: Performance insights from NAEP mathematics problem solving; Frontiers Media SA, 2024, 9, 1452570.
- Wong, J. T.; Richland, L. E.; Hughes, B. S. Immediate versus delayed low-stakes questioning: Encouraging the testing effect through embedded video questions to support students’ knowledge outcomes, self-regulation, and critical thinking; Technology, Knowledge and Learning, 2024, pp. 1–36; Springer.
- Wu, S.; Cao, Y.; Cui, J.; Li, R.; Qian, H.; Jiang, B.; Zhang, W. A comprehensive exploration of personalized learning in smart education: From student modeling to personalized recommendations; ArXiv, 2024. Available online: https://arxiv.org/abs/2402.01666.
- Yenduri, G.; Ramalingam, M.; Chemmalar, S. G.; Supriya, Y.; Srivastava, G.; Kumar, P.; Deepti, R. G.; Jhaveri, R. H.; Prabadevi, B.; Wang, W. GPT (Generative Pre-trained Transformer)–A comprehensive review on enabling technologies, potential applications, emerging challenges, and future directions; IEEE Access, 2024; IEEE.
- Yu, F.-Y.; Kuo, C.-W. A systematic review of published student question-generation systems: Supporting functionalities and design features; Journal of Research on Technology in Education, 2024, 56, 2.
- Yu, T.; Fu, K.; Wang, S.; Huang, Q.; Yu, J. Prompting video-language foundation models with domain-specific fine-grained heuristics for video question answering; IEEE Transactions on Circuits and Systems for Video Technology, 2024; IEEE.
- Zhang, L.; Jr, C.; Greene, J. A.; Bernacki, M. L. Unraveling challenges with the implementation of universal design for learning: A systematic literature review; Educational Psychology Review, 2024, 36, 1.
- Zhao, R.; Tang, J.; Zeng, W.; Chen, Z.; Zhao, X. Zero-shot knowledge graph question generation via multi-agent LLMs and small models synthesis; 2024, pp. 3341–3351.
- Zong, C.; Yan, Y.; Lu, W.; Huang, E.; Shao, J.; Zhuang, Y. Triad: A Framework Leveraging a Multi-Role LLM-based Agent to Solve Knowledge Base Question Answering; ArXiv, 2024. [CrossRef]







| Generation Method | Knowledge Type | Mean | Standard Deviation | N |
|---|---|---|---|---|
| Multi-Examiner | Factual | 4.03 | 1.00 | 30 |
| Conceptual | 3.00 | 0.91 | 30 | |
| Procedural | 3.57 | 0.94 | 30 | |
| Metacognitive | 3.73 | 0.94 | 30 | |
| GPT-4 | Factual | 3.33 | 0.99 | 30 |
| Conceptual | 2.53 | 0.97 | 30 | |
| Procedural | 2.40 | 1.10 | 30 | |
| Metacognitive | 3.20 | 0.99 | 30 | |
| Human | Factual | 3.07 | 1.14 | 30 |
| Conceptual | 3.63 | 0.81 | 30 | |
| Procedural | 3.77 | 0.82 | 30 | |
| Metacognitive | 3.60 | 1.00 | 30 |
| Source of Variation |
Sum of Squares |
Degrees of Freedom (DF) |
Mean Square | F-value | p-value | Partial |
|---|---|---|---|---|---|---|
| Generation Method | 37.62 | 2 | 18.81 | 19.85 | 0.08 | |
| Knowledge Type | 12.33 | 3 | 4.11 | 4.34 | .005 | 0.03 |
| Interaction | 32.93 | 6 | 5.49 | 5.79 | 0.07 | |
| Error | 329.733 | 348 | 0.95 |
| Comparison | Mean Difference | Standard Error | p-value | 95% Confidence Interval |
|---|---|---|---|---|
| Multi-Examiner vs. GPT-4 | 0.71 | 0.16 | [0.41, 1.02] | |
| Multi-Examiner vs. Human | -0.07 | 0.16 | .870 | [-0.38, 0.25] |
| GPT-4 vs. Human | 0.65 | 0.16 | [0.34, 0.96] | |
| Factual vs. Conceptual | -0.42 | 0.21 | .039 | [-0.83, -0.01] |
| Factual vs. Procedural | -0.23 | 0.21 | .453 | [-0.64, 0.18] |
| Factual vs. Metacognitive | 0.03 | 0.21 | .997 | [-0.38, 0.44] |
| Generation Method | Mean | Standard Deviation | N |
|---|---|---|---|
| Multi-Examiner | 4.23 | 0.57 | 30 |
| GPT-4 | 3.40 | 1.13 | 30 |
| Human | 4.43 | 0.57 | 30 |
| Effect | F-value | Hypothesis DF | Error DF | p-value | Partial |
|---|---|---|---|---|---|
| Generation Method | 14.016 | 2 | 87 | 0.244 |
| Dependent Variable | Sum of Squares | DF | Mean Square | F-value | p-value | Partial |
|---|---|---|---|---|---|---|
| Diversity | 8.022 | 2 | 9.011 | 14.016 | 0.244 |
| Comparison | Mean Difference | Standard Error | p-value | 95% CI |
|---|---|---|---|---|
| Multi-Examiner vs. GPT-4 | 1.03 | 0.91 | [0.34, 1.33] | |
| Multi-Examiner vs. Human | -0.20 | 0.91 | .600 | [-0.69, 0.29] |
| GPT-4 vs. Human | 0.83 | 0.91 | [0.54, 1.53] | |
| GPT-4 vs. Human | 1.03 | 0.91 | [0.34, 1.33] |
| Generation Method | Cognitive Level | Mean | Standard Deviation | N |
|---|---|---|---|---|
| Multi-Examiner | Memory | 2.13 | 1.01 | 90 |
| Understanding | 2.27 | 1.11 | 30 | |
| Application | 2.07 | 0.95 | 60 | |
| Analysis | 1.97 | 0.95 | 90 | |
| Evaluation | 1.77 | 0.94 | 30 | |
| Creation | 2.12 | 1.08 | 60 | |
| GPT-4 | Memory | 2.97 | 0.99 | 90 |
| Understanding | 2.80 | 0.96 | 30 | |
| Application | 3.05 | 1.00 | 60 | |
| Analysis | 2.77 | 0.99 | 90 | |
| Evaluation | 2.20 | 0.71 | 30 | |
| Creation | 3.09 | 0.94 | 60 | |
| Human | Memory | 3.35 | 0.91 | 90 |
| Understanding | 3.67 | 1.03 | 30 | |
| Application | 3.18 | 0.81 | 60 | |
| Analysis | 2.88 | 1.01 | 90 | |
| Evaluation | 2.27 | 1.11 | 30 | |
| Creation | 3.43 | 0.96 | 60 |
| Source of Variation |
Sum of Squares |
Degrees of Freedom |
Mean Square | F-value | p-value | Partial |
|---|---|---|---|---|---|---|
| Generation Method | 232.67 | 2 | 116.34 | 122.48 | 0.09 | |
| Cognitive Level | 56.04 | 5 | 11.21 | 11.80 | 0.02 | |
| Interaction | 15.74 | 10 | 1.57 | 1.66 | 0.086 | 0.01 |
| Error | 1008.71 | 1062 | 0.95 |
| Comparison | Mean Difference | Standard Error | p-value | 95% CI |
|---|---|---|---|---|
| Multi-Examiner vs. GPT-4 | 0.81 | 0.09 | [0.64, 0.99] | |
| Multi-Examiner vs. Human | 0.28 | 0.09 | 0.0015 | [0.11, 0.46] |
| GPT-4 vs. Human | 1.09 | 0.09 | [0.92, 1.27] | |
| Evaluation vs. Creation | -0.74 | 0.19 | [-1.11, -0.36] | |
| Analysis vs. Evaluation | 0.83 | 0.23 | [-1.29, -0.37] | |
| Application vs. Analysis | 0.69 | 0.20 | [-1.09, -0.29] |
| Generation Method | Analysis (M ± SD) | Evaluation (M ± SD) | Creation (M ± SD) |
|---|---|---|---|
| Multi-Examiner | 2.97 ± 0.99 | 3.08 ± 0.94 | 2.80 ± 0.96 |
| GPT-4 | 2.13 ± 1.01 | 2.12 ± 1.08 | 2.27 ± 1.11 |
| Human | 3.34 ± 0.91 | 3.43 ± 0.96 | 3.67 ± 1.03 |
| Cognitive Level | F-level | p-level | Partial |
|---|---|---|---|
| Analysis | 36.67 | 0.14 | |
| Evaluation | 28.14 | 0.17 | |
| Creation | 13.96 | 0.12 |
| Cognitive Level | Comparison | Mean Difference | p-value | 95% CI |
|---|---|---|---|---|
| Evaluation | Multi-Examiner vs. GPT-4 | 0.97 | [0.54, 1.40] | |
| Multi-Examiner vs. Human | 1.32 | [0.89, 1.75] | ||
| GPT-4 vs. Human | 0.35 | 0.135 | [-0.08, 0.78] | |
| Creation | Multi-Examiner vs. GPT-4 | 0.53 | 0.120 | [-0.10, 1.17] |
| Multi-Examiner vs. Human | 1.40 | [0.76, 2.03] | ||
| GPT-4 vs. Human | 0.87 | 0.005 | [0.23, 1.50] |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2025 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).
