Preprint
Article

This version is not peer-reviewed.

Evaluation of Systematic Review Methodological Quality with Large Language Models Versus Human Reviewers

Submitted:

21 September 2026

Posted:

22 September 2026

You are already at the latest version

Abstract
Background: There continues to be a growing number of systematic reviews with variable methodological quality. Using a systematic review to inform decision making relies on the quality of methods used to synthesize evidence. Tools that have been developed, such as AMSTAR 2, aim to assess the methodological quality and validity of systematic reviews of healthcare interventions. As this task takes time, use of large language models (LLMs) may help expedite the process. Objectives: We aim to test the ability of two LLMs, Claude Opus 5 and ChatGPT 5.5, to perform AMSTAR 2 assessments on a set of systematic reviews and to compare the accuracy to standard methods using human reviewers. As a secondary objective, we aim to explore the temporal consistency of the AMSTAR 2 assessments from each LLM over time (baseline and one month later).Methods: Our study sample includes open-access systematic reviews of interventions in pediatric populations. Our primary outcome is accuracy of each LLM’s AMSTAR 2 assessments as compared to a human reviewer process as the reference standard. Accuracy will be calculated as the total number of assessments of AMSTAR 2 items that correctly match the human reviewer team. Two humans will independently perform AMSTAR 2 assessments and come to consensus through discussion, or involvement with a third senior reviewer when required. A randomly selected subset of the systematic reviews will be used for the secondary objective exploring temporal consistency of the same LLM model versions over time. Results: AMSTAR 2 assessments will be performed on all 45 systematic reviews by each LLM and two human reviewers. Accuracy of the LLMs versus the human reference standard will be recorded for each AMSTAR 2 item. We will also identify whether there are common discrepancies for specific AMSTAR 2 items, and possible reasons (e.g., items requiring methodological judgement, vague methods reported in the systematic reviews, etc.). The temporal consistency for each model over time will be summarized narratively with descriptive statistics.Conclusion: Evaluating the methodological quality of systematic reviews with LLMs has the potential to more efficiently respond to the needs of decision-makers, by offsetting time spent performing AMSTAR 2 assessments, while maintaining accuracy. Our study will provide evidence to support or reject the use of LLMs to replace human reviewers in this process based on the accuracy of assessments, as well as whether responses are consistent over time. We will also identify AMSTAR 2 items that appear consistently problematic, where decision rules or adequate prompts are needed.
Keywords: 
;  ;  ;  ;  
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.