Preprint
Review

This version is not peer-reviewed.

Annotations to Structure a Systematic Review: Methodological Update and Key Principles Using the S0–S3 Flow

Submitted:

29 July 2026

Posted:

30 July 2026

You are already at the latest version

Abstract
Systematic reviews and meta-analyses are essential methods for synthesising evidence, informing clinical decisions and defining research priorities. This manuscript updates and expands a previous teaching approach to the structure of a systematic review. Since that publication, reporting guidelines, risk-of-bias assessment, protocol management, digital tools and expectations regarding reproducibility have evolved substantially. This update is therefore not limited to adding recent references; it transforms an introductory teaching resource into an academic methodological document for researchers and authors of systematic reviews. The main improvements are the replacement of the PRISMA 2009 framework with PRISMA 2020 and its extensions, the incorporation of PRISMA-S and PRISMA-DTA where applicable, prospective protocol registration, the distinction between risk of bias and certainty of evidence, the use of design-specific appraisal tools, the transition from RevMan 5 to RevMan Web, the integration of GRADEpro GDT, and transparent documentation of artificial intelligence use. The organising axis is a four-stage mnemonic system: S0 identification, S1 screening, S2 eligibility and S3 final inclusion. Each stage is defined by its inputs, processes, decisions, traceable CSV files and deliverables. In this document, CSV refers to the structured management of bibliographic citation files exported in comma-separated value format, which can be used in spreadsheets and processed by bibliographic managers, statistical software or supervised AI agents to identify duplicates, classify records, prioritise full-text retrieval, document decisions and generate counts for the PRISMA flow diagram. CSV files are complementary to and interoperable with bibliographic formats such as RIS or BibTeX, commonly used in Zotero, EndNote, Mendeley and other reference managers. This enables reviewers to move from reference management to the tabular documentation of decisions throughout the review process. Figures 1-7 illustrate the S0-S3 flow as a reproducible and auditable working architecture.
Keywords: 
;  ;  ;  ;  ;  ;  ;  

1. Introduction

This document builds on a previous teaching resource on how to structure a systematic review [1]. That work provided guidance for researchers on organising a review around a clear question, reproducible searching, study selection, data extraction, assessment of bias and coherent reporting. Its pedagogical premise remains valid: a review should not be improvised. However, the methodological context has changed substantially. PRISMA 2020 has replaced PRISMA 2009; risk-of-bias instruments have become more consolidated; RevMan Web has progressively replaced the teaching use of RevMan 5; GRADE has developed a more precise distinction between study-level limitations and certainty in the body of evidence; and the use of automation and artificial intelligence now requires explicit transparency standards [2,3,8,9,14,15,19,20].
From this perspective, this paper is aimed at researchers, clinicians, academic teams, reviewers, and authors who need to transform a research question into a documented and reproducible systematic review in IMRD format.
Systematic literature research has evolved from a preparatory task, often considered secondary or documentary, into a research method in its own right. Its importance lies in bringing order, transparency and reproducibility to the relationship between the published literature and decision-making [2,6]. In this sense, it helps to reduce bias, identify knowledge gaps, guide new studies and support informed decisions by professionals, patients and managers [38]. Unlike narrative reviews, which depend largely on the author’s judgement and are vulnerable to selection bias, systematic reviews require explicit questions, declared sources, recorded strategies, eligibility criteria and justification for exclusions [4,17,21]. This shift has made it possible to transform literature review into a verifiable and cumulative procedure that shows not only what is known, but also how that knowledge has been obtained [2,3].
At present, the systematic review occupies a central place in evidence synthesis. It does not merely collect articles: it identifies knowledge gaps, compares findings, assesses the risk of bias in available studies, estimates the certainty of the evidence and helps determine whether further research is warranted [6,14,24]. Therefore, before designing a clinical trial, an observational study, a scoping review, a doctoral thesis, an academic project or a competitive grant proposal, researchers need to know what has been published, which questions remain open, which outcomes have been used, which populations have been studied, which biases limit the existing evidence and what contribution a new study can make [2,17,21].
In the context of a doctoral thesis, systematic literature searching has additional value. It synthesises available knowledge, identifies gaps and opportunities, clarifies variables for subsequent study and helps determine whether sufficient data exist to support preliminary analyses or pilot studies [24,25,26]. Such pilot work, derived from the review itself or from initial screening, can support the subsequent empirical phase of the thesis by anticipating problems in variable definition, data availability, methodological heterogeneity, outcome selection, reviewer consistency and design feasibility [27,28]. In addition, when conducted through a transparent and reproducible method, the review can be published as a stand-alone article and subsequently cited in the introduction of the thesis or related publications as part of the scientific context, justification and formulation of the research question [26,29].
From this perspective, systematic literature searching should be understood as the starting point of a research project, rather than as a final task used merely to justify the introduction of a manuscript. A well-formulated question arises from critical knowledge of previous evidence and must be translated into inclusion criteria, outcomes and coherent synthesis methods [6,21]. If this exploration is not conducted methodically, researchers may duplicate work that has already been resolved, define irrelevant objectives, choose non-comparable outcomes, underestimate feasibility problems or design studies that do not address a genuine scientific gap. Conversely, a documented systematic review guides the hypothesis, delimits the population, selects variables, anticipates biases, supports the theoretical framework and strengthens the ethical and scientific relevance of the project [2,4,17].
This update is grounded in that premise: reviewing the literature systematically is not simply a search for prior references, but the construction of an evidence map on which subsequent research can be based. The S0-S3 flow is intended to make this map visible from the first identified record to the final synthesis, so that the reference list becomes not merely a set of citations but a traceable, assessable and reusable knowledge base aligned with PRISMA 2020, PRISMA-S and prospective protocol registration [2,3,4,17].
This version provides a methodological update rather than a conventional literature review. Its aim is no longer merely to list steps, but to support the planning, conduct, auditing, synthesis and reporting of systematic reviews and meta-analyses according to current international standards [2,4,6]. Compared with the previous document, the main advance is the transition from an indicative teaching sequence to a traceable working architecture. Its components are introduced in section 2 and summarised in Table 1; the S0-S3 operational axis is presented in section 5.2.
The concepts presented in this document are part of a line of work developed by the first author in two areas: teaching and scientific publishing linked to Revista ORL, and academic collaboration in the environment of the SEORL-CCC in materials to support bibliographic research. This trajectory is taken up here as the basis of the updated approach, the structure of which is summarized in section 2 and whose operational development is organized through the S0-S3 flow [1,38].

2. Key Changes from the Previous Approach

The previous document offered teaching guidance for structuring a systematic review: formulating a question, searching the literature, selecting studies, extracting data, assessing study limitations and writing the report [1]. This update retains that pedagogical purpose, but develops it into a traceable methodology aligned with current standards for evidence synthesis [2,4,6]. The main difference is that the review is no longer presented as a succession of general steps, but as a documented chain of verifiable decisions.
The improvements are grouped into seven changes. First, the reporting framework is updated from PRISMA 2009 to PRISMA 2020 and its extensions, requiring more precise reporting of the search strategy, study selection, use of automation, data availability and certainty of evidence [2,3]. Second, the protocol is no longer treated as a generic recommendation, but as a prospective and registrable document that may require justified amendments [17]. Third, the search process is documented using PRISMA-S, with strategies by source, dates, filters, limits, original exports and explicit deduplication [4]. Fourth, methodological appraisal moves away from the simplified notion of an overall “quality” score and adopts design-specific instruments and bias domains [8,9,10,11,12,13]. Fifth, GRADE conceptually separates certainty of evidence from risk of bias, so that confidence is assessed by outcome and across the body of evidence [14]. Sixth, practical work shifts from RevMan 5 and auxiliary tables to RevMan Web, GRADEpro GDT, reference managers, screening platforms and interoperable CSV files [15,16]. Seventh, artificial intelligence is incorporated as supervised and auditable support, not as a substitute for reviewer judgement [19,20].
The main contribution of this version is the S0-S3 flow. Its operational definition is developed in section 5.2, mapped to PRISMA 2020 in Table 6 and related to IMRD writing in Table 7. The system helps avoid confusing records with reports, reports with studies, exclusions by title and abstract with exclusions after full-text assessment, and studies included in narrative synthesis with studies included in a specific meta-analysis [2,3]. This update therefore does more than incorporate recent standards; it proposes a working architecture that enables systematic reviews to be taught, conducted and audited with greater clarity [4,6].
The S0-S3 flow also acts as a bridge between two products that are often finalised late in the process: the PRISMA diagram and the manuscript in IMRD format. The distribution of counts by stage is shown in Table 7 and summarised in Figure 1. In this way, the diagram is no longer reconstructed retrospectively from scattered notes or recollections, but generated from versioned working files [2,3].
The same logic favors the writing of the manuscript. The relationship between S0-S3 and the IMRD sections is summarized in Table 7 and developed in section 3. Thus, S0-S3 is not only an aid to select articles, but a documentary skeleton that connects the execution of the review with the report required by PRISMA and with the orderly writing of the scientific manuscript [2,4,6,14].
Overall, the update moves the document from an introductory orientation to a working methodology. The new contribution is not only the incorporation of recent standards and tools, but the creation of a traceability logic that connects question, protocol, search, selection, eligibility, inclusion, synthesis, reporting and audit [2,4]. The specific elements of this update are summarized in Table 1 and developed in the following sections, especially in sections 5.2, 14 and 17.

3. Strategy for Writing the Final Report in IMRD Format

When preparing the final report of a systematic review—understood as a research article structured according to IMRD: Introduction, Methods, Results and Discussion—it is useful to distinguish the logical order of the manuscript from the practical order of writing. The article will be read in IMRD sequence, but it does not necessarily need to be written in that order. The recommended strategy is to begin the entire process with a clear, well-formulated and justified research question, derived from preliminary literature exploration [17,21]. This question determines the title, keywords, eligibility criteria, search strategy, extraction variables, analysis plan, PRISMA diagram and interpretation [2,4].
Once the question has been defined, drafting can begin with the sections that depend most directly on the protocol and the data: objectives, methods, analysis and results. The interpretation, discussion and conclusions can then be developed. Only after these elements are clear should the introduction be finalised, explaining concisely the context, the current state of knowledge, the concepts required to understand the review, what is already known and what remains uncertain, including the gaps or limitations in previous research that justify the review. The abstract, whether in one language or more than one, should be written at the end, once the objective, methods, results, interpretation and conclusions have been established [2,24].
The distinct role of the references used in each part of the manuscript should also be made explicit. The references used to write the introduction are not the same as the studies obtained through the systematic review. The former contextualise the problem, define concepts, describe the state of knowledge and justify the question; they may include preliminary readings, methodological documents, previous reviews, clinical guidelines, reports or key studies selected for their contextual value. By contrast, methods references correspond to the standards, guidelines and instruments used to design and report the review, whereas the references in the results and discussion derive mainly from included studies, excluded texts when they need to be cited, and sources required to interpret the findings [2,4,6]. Confusing these reference sets can produce overextended introductions, results contaminated by non-included citations or discussions that fail to distinguish included evidence from general background.
In practical terms, the S0-S3 flow helps maintain this separation. S0 contributes to formulating and justifying the question, S1 and S2 document which records and reports do or do not become part of the reviewed evidence, and S3 defines the core evidence base for the results and discussion. This correspondence is summarised in Table 7. The separation clarifies the IMRD manuscript: the introduction explains why the review was needed; the methods explain how it was conducted; the results describe what was found; and the discussion interprets those findings in light of the included evidence and the broader scientific context [2,6,14].
Table 2 summarises this writing strategy: first, the question and its preliminary bibliographic justification are defined; next, the sections dependent on the protocol and data are drafted; finally, the contextual introduction and abstract are completed.

4. Rationale for a Systematic Review

A systematic review is secondary research that uses explicit and reproducible methods to identify, select, evaluate and synthesise studies that answer a clearly formulated question [2,6]. Within the broad concept of literature research, distinctions can be made between scoping reviews, systematic reviews and umbrella reviews, each with different objectives and sources depending on the degree of development of the topic under investigation [38]. A systematic review may or may not include meta-analysis. Meta-analysis is a statistical technique for combining quantitative results; it does not replace the systematic review or, by itself, turn a literature search into reproducible research.
The design should be chosen according to the review question. For interventions, PICO or PICOS are usually used; for exposure and aetiology, PECO or PEO; for diagnostic accuracy, PIRD; for scoping reviews, PCC; and for prevalence, CoCoPop. The choice of framework shapes the eligibility criteria, search strategy, extracted variables and synthesis [5,7,21]. Table 3 emphasises that the question is not a preliminary formality, but the methodological blueprint for the entire review.

5. Literature Searching in Two Phases: Preliminary Phase and Development Phase

5.1. Preliminary Phase: Design, Pilot Search and Registrable Protocol

The literature research leading to a systematic review can be organized in two phases. The first is the preliminary phase, whose objective is to design the research before starting the definitive systematic search and to reduce the probability of rectifications during the subsequent development. In this stage, the team defines the problem, conducts exploratory readings, identifies key concepts and knowledge gaps, formulates an initial question, and assesses whether there is sufficient literature to justify a review [17,21,24,36]. It is not yet equivalent to systematic review nor does it generate its results: it is a stage of design, field learning and feasibility testing, aimed at arriving at the protocol with a question, objectives and stable criteria [2,6,37]. AI agents can support this planning by structuring the question (PICO, PECO, PCC, PIRD or CoCoPop), the conceptual organization of the problem and the initial proposal of terms, sources, criteria and variables; its function is auxiliary and does not replace the bibliographic search, the consultation of the original sources or the methodological judgment of the team [19,20]. The reusable prompt for this first scan is shown in Annex 3, Table 16.
A pilot search should be conducted within this preliminary phase. Its role is not to produce the results of the review, but to check if the question is suitable and if the search strategy will be viable. This test allows us to detect if the question is broad, narrow, ambiguous or not very recoverable; identify synonyms, acronyms, language variants, authors’ keywords and controlled vocabulary – MeSH, DeCS, Emtree or CINAHL Headings, depending on the sources; selecting potentially relevant databases and grey literature; locate sentinel studios; check whether the inclusion and exclusion criteria are applicable; anticipate extraction variables; and to estimate the volume of records that the work will generate [27,33,34,35]. An AI agent can facilitate the generation and comparison of these proposals, as well as build a pilot strategy for PubMed and suggest its preliminary adaptation to other platforms, but each controlled term, syntax, filter, limit, reference, and identifier must be checked in the corresponding source. To guide this task, the design prompt, pilot search and registrable protocol included in Annex 3, Table 16 can be used, adapted to the specific question. Based on the actual results of the pilot search, the team can reframe the question, reformulate objectives, adjust the provisional title, redefine eligibility criteria, and improve strategies before committing to the final protocol [4,21,33].
The use of AI in this phase should be subject to a rule of thumb: it can help design the review, but it should not be used as a single source to identify evidence, generate references, decide on eligibility, or produce PRISMA counts [19,20]. The input, prompt, tool and version when available, date, output, and human validation should be filed with the rest of the methodological documentation. The counts that subsequently feed PRISMA will come exclusively from executed searches, actual bibliographic exports, documented deduplication and verified screening decisions [2,3,4].
The preliminary phase culminates with the drafting of the research protocol. Only when the question, objectives, eligibility criteria, sources, verified preliminary strategies, outcomes, extraction variables, bias risk assessment plan, synthesis plan, and documentation of any AI support have been stabilized should the registrable protocol be drafted, e.g., in PROSPERO (https://www.crd.york.ac.uk/prospero/) or other appropriate registry. The protocol should not be born from an improvised search or an AI output accepted without verification, but from a documented and consensual preparation; Its role is to reduce bias, increase transparency, anticipate decisions, and provide a basis for the final report [2,3,17,30,37]. The prompt in Table 16 of Annex 3 can be used as a support list to prepare these components, but the team must manually review each proposal before incorporating it into the protocol. Some useful web resources for this phase are summarized in Annex 4. The recommended sequence is: preliminary design and pilot search; methodological validation and stabilization; protocol and registration; and, then, the development phase of systematic bibliographic research.
This distinction is essential for researchers who are starting or wishing to systematize a review. In the preliminary phase, the question, objectives, terms, criteria, sources, extraction variables or even the focus of the work can be modified, because its purpose is to design the review before compromising the protocol. The better this stage is recorded—including pilot searches, sentinel studies, strategy versions, and, if AI was used, prompts, inputs, outputs, and validations—the lower the likelihood of changes during development, when such adjustments may affect the validity, transparency, and credibility of the study [2,3,4,17,19,20,34]. On the other hand, once the protocol has been registered, substantial changes in the question, criteria, sources or outcomes must be recorded as amendments, dated and justified [17,30,37]. The preliminary phase ends when the team has a stable and justified question, clear objectives, agreed criteria, verified preliminary strategies and a protocol that allows S0 to be initiated.
Table 4. Sequence of the preliminary phase before starting the development phase of the bibliographic research.
Table 4. Sequence of the preliminary phase before starting the development phase of the bibliographic research.
Preliminary phase step What needs to be done Product to move to development
Delimitation of the problem Review exploratory literature, identify gaps and structure the question using PICO, PECO, PCC, PIRD or CoCoPop. Define population, intervention/test/exposure, comparator or reference, outcomes, context, and possible designs [21,24]. Structured initial question, working title, essential concepts, preliminary keywords and justification of scientific interest.
Pilot search Test free terms, synonyms, linguistic variants and controlled vocabularies; select sources and grey literature; check sentinel studies, volume of results and applicability of the criteria. An AI agent can propose these elements using the prompt in Annex 3, Table 16, but terms, syntax, filters, references, and identifiers must be manually validated [19,20,33,34,35]. Refined question and goals; verified preliminary strategy; adjusted sources and eligibility criteria; sentinel studies identified; and anticipated extraction variables.
Registrable protocol Stabilize question, objectives, criteria, sources, strategies, outcomes, variables, risk of bias and synthesis plan. Date and archive pilot searches and versions of the strategy; if AI was used, retain tool and version, date, prompt, input, output, and human validation [2,4,17,19,20,37]. Dated and, if applicable, recorded in PROSPERO or another registry; reproducible methodological file; and validated plan to start S0. Any subsequent substantial changes will need to be documented as an amendment.
Table 5 summarizes the complete logic: the preliminary phase designs and stabilizes the review; the development phase executes systematic bibliographic research using S0-S3 and produces the materials needed for PRISMA, the synthesis and the IMRD manuscript.

5.2. Development Phase: the S0-S3 Sequence

The second phase is the development phase of the review. It begins when the question has been refined, the protocol has been drafted and registered or archived, and the team is ready to run the definitive systematic search. This relationship with the preliminary phase is summarised in Table 5. At this stage, the S0-S3 flow organises the work into four steps: S0 identifies the universe of records; S1 screens titles, abstracts and keywords; S2 evaluates full texts; and S3 includes studies, extracts data, assesses methodological quality, synthesises evidence and supports reporting. This sequence is compatible with SALSA logic—Search, AppraisaL, Synthesis and Analysis—but translates it into a documentary architecture connected to PRISMA and IMRD [38]. From this point onwards, the review should function as a chain of traceable decisions consistent with PRISMA 2020, PRISMA-S, prospective protocol registration and current recommendations for searching, selection, extraction and synthesis [2,3,4,6,17].
Although it takes the PRISMA flow diagram as its starting point—identification, screening, eligibility and inclusion—the S0-S3 classification is proposed here as a practical tool for converting those blocks into a traceable work system. Its value lies in distinguishing decisions that are often conflated, separating records, reports and studies, assigning a CSV file to each phase and ensuring that PRISMA counts, decision logs, supervised AI use and IMRD writing derive from the same documentary architecture. This relationship is developed in Table 6, summarised in Figure 1 and connected to the IMRD structure in Table 7.
The S0-S3 system organises development as a chain of selection, documentation and writing. The letter S refers to “selection”, understood as the progressive selection of records, reports and studies; it also serves as a reminder of three core principles: sequential selection, documentary support and reproducible synthesis. Each phase has an operational objective, a master file, a relationship with PRISMA 2020 and a role in IMRD writing [2,3,4,24,25]. The operational correspondence is shown in Table 6 and summarised in Figure 1. Numbering begins at S0 because identification is not yet screening: it is the zero point at which the initial universe of records is generated and the traceability of the review is established.
The CSV file for each stage should not be regarded merely as an auxiliary table, but as a decision matrix. It enables reviewers to compare bibliographic fields, record decisions, generate PRISMA counts, prepare annexes, construct tables and maintain an auditable trail of each instruction and output when AI agents are used. At each stage, the prompt—the textual instruction given to the AI agent—should be archived in the same way as search strategies, with date, version, tool, input file, criteria applied, output obtained and human validation. This practice is consistent with transparency requirements for search strategies, study selection, automation and data availability [2,4,19,20]. Table 7 shows how S0-S3 feeds the PRISMA diagram and IMRD writing; Table 14 details the supplementary documentation file; and Annex 3, Table 16, provides reusable prompts for stages S0-S3.
The pedagogical and operational value of the system can be remembered through four stations within the development phase: S0 identifies; S1 screens; S2 assesses eligibility; and S3 includes and synthesises. The relationship with the PRISMA 2020 flow diagram is shown in Table 6 and Figure 1; therefore, this section emphasises its usefulness for avoiding the mixing of different types of decisions, such as exclusion by title and abstract, exclusion after full-text assessment or exclusion from a quantitative synthesis because compatible data are unavailable [2,3,4].
In addition to supporting the PRISMA diagram, S0-S3 organises manuscript writing. The correspondence with Introduction, Methods, Results, Discussion and Conclusions is presented in Table 7 and is connected to the writing strategy described in section 3. This relationship turns the S0-S3 flow into a template for moving from documentary work to the IMRD manuscript without losing traceability [2,6,14,25].

6. S0: Identification

S0 marks the beginning of the review development phase. Its purpose is no longer to explore whether the question is feasible, but to execute the systematic identification process specified in the protocol. It includes selecting sources, running searches, exporting results, preserving complete strategies and building a traceable raw record base. The question, objectives and criteria must have been stabilised during the preliminary phase; if a substantial change becomes necessary during S0, it should be recorded as a justified amendment to the protocol before proceeding [2,4,17].
The main output of S0 is the S0 CSV: a raw database of records, before final deduplication, that preserves source, search date, strategy used, title, authors, year, DOI, PMID or other identifiers, abstract and available metadata. This database enables reviewers to audit the quality of the executed strategy, verify whether sentinel studies defined in the preliminary phase have been retrieved, detect unexpected terms, identify noise, assess which databases provide unique records and prepare the initial counts for the PRISMA diagram [4,33,34]. These checks should be used to verify the search, not to redesign the question after registration.
The S0 CSV brings together database exports and prepares the material for identifying duplicates through exact or approximate matching of title, authors, year, DOI, PMID or other identifiers. At the same time, the researcher should create a complementary file containing search strategies, dates, filters, limits and the number of records retrieved, as detailed in section 14 and Table 14. In practice, the CSV functions as the working matrix, whereas RIS or BibTeX files preserve reference structure for exchange between bibliographic managers [2,4].
If an AI agent is used in S0, the prompt must specify the parsed bibliographic fields and the requested task. The general AI documentation requirements are described in section 17, while the minimum fields for recording prompts, inputs, outputs, and validation are summarized in Table 14 and exemplified in Annex 3 [19,20].
For intervention reviews, PICO should define population, intervention, comparator and outcomes. For diagnostic accuracy reviews, PIRD should define population, index test, reference standard and target condition. This distinction is crucial: a diagnostic review is not assessed with the same outcomes, bias tools or statistical models as an intervention review [5,7,21,22,23]. Figure 2 summarises S0 as the stage of identification and construction of the raw record base.
Before records are downloaded for the final review, the team should have completed the preliminary phase, tested at least one pilot search, adjusted the question and objectives if necessary, and defined the protocol, eligibility criteria and sources. The S0 output is not a clean list, but a snapshot of the universe retrieved from each source. For this reason, duplicates are initially preserved with their provenance and deduplication is documented later [4,33]. The minimum recommended fields for the S0 CSV are presented in Table 8.

7. S1: Screening

S1 transforms the raw record set into a set of candidates. It includes deduplication, title-and-abstract screening, preliminary classification and discrepancy resolution. PRISMA 2020 requires authors to report how many reviewers screened each record, whether reviewers worked independently and whether automation tools were used [2,3,30]. When artificial intelligence is used, the final decision must remain under human responsibility and its use must be reported transparently [19,20].
Deduplication should not be treated as a black box. The software used, the matching criteria and the number of records removed should be recorded [4,33]. Screening can be supported by Rayyan, Covidence, EPPI-Reviewer, DistillerSR, ASReview or other platforms, but none of these replaces the protocol. To avoid repeating the description of each resource here, their function within S0-S3 is set out in Annex 4, Table 19. In reviews with a high volume of records, AI may prioritise or classify records provided that its use is declared and validated, as detailed in section 17 [19,20].
In S1, the CSV enables an AI agent to identify duplicates, suggest potentially eligible records and prioritise each reference for full-text retrieval. The general conditions for AI use and prompt archiving are developed in section 17 and Annex 3; at this stage, classification is not equivalent to definitive inclusion, but serves as an aid for ordering the retrieval and reading of full texts. Figure 3 depicts this transition from the raw set to potentially eligible records.
The S1 stage helps to separate two actions that are often confused: removing duplicates and excluding records due to lack of relevance. The first is a bibliographic operation; the second, a preliminary eligibility decision. Both must be recorded with counts that will feed the PRISMA 2020 diagram [2,33]. Table 9 summarizes the practical recommendation for using AI responsibly during screening.

8. S2: Eligibility

S2 corresponds to full-text eligibility assessment. In this phase, each report is read in full, inclusion and exclusion criteria are applied and a primary reason for exclusion is recorded for each rejected text. PRISMA 2020 recommends citing studies that appeared to meet the criteria but were excluded, explaining why [2,3]. This phase produces the S2 CSV and the full-text exclusion table. Where AI is used to support this assessment, the documentation criteria described in section 17 and the prompt models in Annex 3 should be applied, without replacing verification by human reviewers [19,20].
The transition from S1 to S2 should not be understood as final selection. Many records pass title-and-abstract screening but fail at full text because of population, design, intervention, outcomes, extractable data, duplicate populations, incomplete text or discordance with the question. Eligibility requires critical reading and, in complex reviews, consensus or adjudication by a third reviewer [2,3]. Figure 4 shows how full-text assessment and reasons for exclusion should be documented.
Each full-text exclusion should be assigned to a predefined and sufficiently specific category: ineligible population, ineligible intervention or test, inadequate comparator, non-relevant outcome, ineligible design, absence of extractable data, duplication, non-retrievable text or another justified cause. Vague reasons such as “not relevant” should be avoided at full-text stage because PRISMA 2020 requires explicit reasons for exclusion [2,3]. Table 10 proposes operational exclusion categories for this phase.

9. S3: Final Inclusion

S3 is the inclusion and synthesis phase. It includes data extraction, risk of bias assessment, qualitative analysis, meta-analysis where appropriate, assessment of certainty with GRADE, and manuscript writing. The CSV S3 is the evidence base: it contains included studies, characteristics, variables, outcomes, bias judgments, numerical data, synthesis decisions, and cross-referencing with tables and figures from the manuscript [6,8,9,10,11,12,14].
In addition, the S3 CSV consolidates the decision for each article, prepares the counts for the PRISMA diagram and enables preliminary statistical or bibliometric exploration. AI-assisted extraction should support structured data capture and verifiable proposals, rather than replace methodological judgement; the fields and prompts recommended for this phase are detailed in Annex 1, Table 15, and Annex 3, Table 16.
Not all studies included in the qualitative review are included in all meta-analyses. A review may include studies for narrative synthesis, but exclude some from quantitative synthesis due to clinical heterogeneity, lack of comparable data, or non-combinable outcomes. That decision must be pre-specified or justified [6,18]. Figure 5 places inclusion as a bridge between selection, extraction, methodological evaluation, synthesis and reporting.
From S3, characteristic tables, risk of bias matrices, forest plots, SROC curves, Summary of Findings tables and conclusions are constructed. Final inclusion should not be confused with “accept all”: each synthesis may have its own subset of contributing studies [6,14,25].

10. Risk of Bias: Tools by Design

Risk-of-bias assessment should be conducted with design-specific tools and, where required by the instrument, by outcome. Assigning a single overall quality score can be insufficient and sometimes misleading. The current approach assesses domains such as randomisation, deviations from the intended intervention, missing data, outcome measurement, selection of the reported result, confounding, participant selection, exposure classification, reference standard and timing, depending on the study design [8,9,10,11,12,13]. Table 11 summarises recommended instruments according to the design of the included studies.
The phrase “the quality of the articles was assessed” should be avoided if the review actually assessed risk of bias. A more precise report should specify the tool, domains, number of reviewers, independence, discrepancy resolution, unit evaluated and how judgements were used in interpretation or sensitivity analysis [8,9].

11. Certainty of the Evidence: GRADE

GRADE assesses not only individual studies, but the certainty of the body of evidence for each outcome. Its main domains are risk of bias, inconsistency, indirectness, imprecision and publication bias. In observational studies, certainty may be upgraded for large effect size, dose-response gradient or plausible residual confounding that would reduce the observed effect. Summary of Findings tables should report critical outcomes, number of participants and studies, effect estimates, certainty ratings and explanations of judgements [14].
This update recommends integrating GRADE from the protocol: defining critical outcomes, clinically important thresholds, assumptions about missing data, and plans for heterogeneity. GRADEpro GDT can import or synchronise data with RevMan Web in Cochrane reviews, and can be used independently in non-Cochrane reviews using data packs or manual entry [15,16].

12. Qualitative Synthesis and Meta-Analysis

The synthesis must answer the review question and respect clinical, methodological and statistical comparability. Narrative synthesis is not free discussion: it should be organised by outcomes, subgroups, designs, risk of bias and direction of effect. When meta-analysis is not appropriate, SWiM guidance can be used to report synthesis without meta-analysis of effect estimates [18].
In intervention meta-analyses, effect measures should be specified: risk ratio, odds ratio, risk difference, mean difference, standardised mean difference or hazard ratio, depending on the outcome. The fixed-effect model assumes a common effect, whereas the random-effects model assumes a distribution of effects across studies. Heterogeneity should be interpreted using I2, Ta2², prediction intervals where appropriate and, above all, clinical plausibility [6].
In diagnostic accuracy reviews, data are usually organised in 2×2 tables: true positives, false positives, false negatives and true negatives. From these data, reviewers calculate sensitivity, specificity, positive and negative likelihood ratios, diagnostic odds ratios and SROC curves. Sensitivity expresses the proportion of patients correctly identified; specificity expresses the proportion of non-patients correctly classified; LR+ indicates how much the probability of disease increases after a positive result; LR− indicates how much it decreases after a negative result; and DOR summarises discrimination, although it may be clinically unintuitive. Bivariate hierarchical models or HSROC models are preferable to univariate averages when sensitivity and specificity are combined, because they preserve the correlation induced by diagnostic thresholds [7,22,23].

13. PRISMA 2020 & Extensions

PRISMA 2020 is not a manual for conducting reviews, but a guideline for reporting them completely and transparently. However, it is useful during planning because its items require the capture of information that must later be reported: eligibility criteria, sources, search strategies, selection, extraction, risk of bias, effect measures, synthesis methods, reporting bias, certainty of evidence, protocol, funding, competing interests and data availability [2,3,30]. Table 12 summarises selected PRISMA guidelines and extensions that may be used depending on the type of review.

14. Document Flow, Reproducibility and Auditing

Reproducibility depends on a third party being able to reconstruct the review without relying on oral explanations, team recollections, or scattered files. Therefore, the documentation of the study should not be limited to the main manuscript: it should allow the search, selection, extraction, analysis and synthesis to be reproduced. The structure of folders and materials is developed in section 15 and Table 13; the minimum content of the supplementary file is summarized in Table 14. In this section, the general principle is retained: the S0-S3 system turns each stage into an audit point and the combined use of CSV, RIS and BibTeX allows the decision matrix to be connected to the bibliographic library without losing traceability [2,4,20]. Figure 6 summarizes this chain of documentary audit from S0 to the final manuscript.
The audit principle can be summarised as follows: every decision should be documented, every result traceable to its source and every synthesis supported by verifiable calculations. This approach is useful for research teams, doctoral theses and reviews that may be updated in the future [2,4,26].

15. Practical Organization for Researchers

The project organization must allow the process to be rebuilt from the protocol to the final repository. In teaching and supervising academic work, it is useful to recommend four coordinated products: an Excel file or spreadsheet, a library in Zotero or another bibliographic manager, the Word document of the research report and a supplementary document. Its minimum content is already summarized in Table 14, while the folder structure is proposed in Table 13. Figure 7 shows this organization and its relationship with traceability materials.
Each file should be named with an ISO date and version number, for example: 2026-07-26_S1_screening_v03.csv. Substantial modifications should be recorded in a logbook. When spreadsheets are used, headers should be locked, categories validated, merged cells avoided in databases and a variable dictionary maintained. The final results table should be exportable to CSV without loss of information. At the same time, a bibliographic copy should be kept in RIS or BibTeX to ensure interoperability with Zotero, EndNote, Mendeley or other managers. In this way, the team maintains two connected levels: the reference library and the tabular matrix of methodological decisions [4,33]. Annex 4, Table 17, Table 18 and Table 19, lists these managers and other web resources useful for organising the review.
The supplementary documentation file can be organized as a text document, a spreadsheet or a mixed annex. Its purpose is to gather, independently of the main manuscript, the information necessary to reproduce the study. To avoid duplication, its minimum blocks are detailed in Table 14; where AI has been used, the Annex 2 declaration should be added and the corresponding Annex 3 prompts filed , along with inputs, outputs and human validation [2,4,19,20].
Table 14 summarizes the minimum elements that this supplementary documentation file should contain and its usefulness within the S0-S3 flow.
Table 14. Content of the supplementary documentation file of a systematic review.
Table 14. Content of the supplementary documentation file of a systematic review.
Documentary block Minimum content Methodological utility
Search strategies Search engine or platform, date, complete equation, filters, limits, language, time period, number of results obtained in S0, captures if applicable and exported file. It allows you to reproduce the search, audit S0, verify the search engines used and obtain the initial counts of the PRISMA diagram.
Bibliographic exports Original files in RIS, BibTeX, CSV, or other formats, with normalized name, download date, source, and preservation without loss of metadata. It preserves the original bibliographic library, facilitates exchange with managers and spreadsheets, and allows you to re-export records without loss of metadata.
Deduplication Tool used, matching criteria, fields compared, number of duplicates removed, and version of the resulting file. It documents deleted records before screening and prevents literature purge from being a black box.
AI Prompts Phase S0-S3, agent used, version if applicable, date, objective, full prompt used or adapted from Annex 3, input file, output generated in CSV or other documented format and responsible for validation. It allows you to reconstruct the instructions given to the agent and verify how deduplications, priorities, inclusions, exclusions, or extractions were proposed.
Selection decisions Decision by record or report: include, perhaps, exclude, high/medium/low priority, reason for exclusion, reviewer and resolution of discrepancies. It connects S1 and S2 with reviewer traceability and exclusion justification.
PRISMA Counts Records identified, duplicates deleted, records screened, records excluded, texts searched, texts evaluated, exclusions on grounds and studies included. It feeds the PRISMA 2020 diagram and avoids inconsistencies between search, screening, eligibility and final inclusion.
Versioning and auditing File name, ISO date, version, responsible, changes made, link to CSV/RIS/BibTeX and location of supplementary material. It facilitates future updating of the revision and allows a third party to rebuild the entire process.
Reproducible supplement Supplementary document with search strategies, search engines, dates, S0 counts, S1-S3 selections, tables of results, data matrices, statement on AI and automation adapted from Annex 2, full prompts used or adapted from Annex 3, and materials not included in the main manuscript. It should contain everything necessary for other researchers to reproduce the study, verify methodological decisions not detailed in the main text, and reconstruct the use of AI using the archived statement and prompts .
Table 14 defines what information must be kept in the supplementary file to ensure traceability of searches, exports, deduplication, prompts, selection decisions, and PRISMA counts. In order not to repeat this information in different sections, Annex 1 offers the template of CSV fields in stages S0-S3, Annex 2 proposes the narrative statement on AI and automation, and Annex 3 brings together reusable prompts that must be archived in their entirety or adapted.

16. Using RevMan Web, GRADEpro, and Companion Tools

RevMan Web is currently Cochrane’s main platform for study data management, analysis and reporting. It enables collaborative work, predefined PICOs and analysis criteria, meta-analysis and figure generation. GRADEpro GDT is used for Summary of Findings tables and certainty assessment; in Cochrane reviews it can be integrated with RevMan Web, and in non-Cochrane reviews it can be used through data transfer or manual entry [15,16].
For the rest of the workflow, teams can combine tools: Zotero, EndNote or Mendeley for reference management; Rayyan, Covidence or EPPI-Reviewer for screening; Covidence, EPPI-Reviewer or DistillerSR for extraction; R packages such as metafor, meta, mada and diagmeta for analysis; and repositories such as OSF, Zenodo or Figshare for materials. The choice of tool should be based on traceability, feasibility, team expertise, cost and reporting requirements, rather than on novelty alone [2,4]. Annex 4, Table 17, Table 18 and Table 19, provides an organised list of web resources and their function within S0-S3.

17. Artificial Intelligence and Automation Considerations

Automation can improve efficiency, but it may also introduce opacity, bias, errors and technological dependency. Recent recommendations from evidence synthesis organisations allow the use of AI provided that methodological rigour is not compromised and human oversight is maintained. Generative AI systems can support question formulation, protocol development, search strategy refinement, title-and-abstract screening, information extraction, data analysis, report writing and formal self-assessment of the manuscript, provided that instructions are clear and verifiable [38]. Any tool that suggests judgements about eligibility, extraction, risk of bias, synthesis or certainty should be declared in the methods or supplementary material, and the report should make explicit the human-AI interaction, delegated tasks, limitations and validation performed [19,20].
This document recommends documenting AI use according to the same traceability principles applied to search strategies: tool, version, date, task, input data, output generated, prompts or parameters, human validation, discrepancies and limitations. To avoid repetition, the minimum content is referred to Table 14 and a suggested statement is provided in Annex 2 [4,19,20].
The prompt used to select references should be treated as a methodological document. Just as search equations from PubMed, Embase, Scopus or Web of Science are archived, prompts used to deduplicate, classify, prioritise, exclude, select full texts or extract data should be preserved. Their general components are summarised in Table 14, and reusable models by stage are presented in Annex 3, Table 16. Annex 4, Table 18 and Table 19, also brings together databases, platforms and web tools useful for searching, screening, analysis and depositing materials [19,20].
AI-assisted extraction can incorporate visible information and subtle cues. The former include title, abstract, keywords, MeSH/DeCS (https://decs.bvsalud.org/) terms, document type, journal, year, DOI, PMID, essay identifiers, country, language, and affiliations; terminological standardization with DeCS/MeSH and VHL resources (https://bvsalud.org/) can facilitate matching between searches in English, Spanish, and Portuguese [31,32]. Annex 4, Table 18, includes terminological and bibliographic web resources that can support this standardization. The latter include methodological expressions of the abstract, study design, data source, sampling method, population or subpopulation, terms related to exposure, intervention, index test or reference standard, outcomes measured, instruments used, statistical analyses, literature cited, previous cohorts, clinical records, databases used, and possible sample overlaps. These signals should not produce a definitive automatic decision, but they can help prioritize work and make more transparent why a record is kept, discarded, or passed for human review [4,19].

18. Discussion

Updating the previous approach requires shifting the emphasis from a simple teaching sequence to an auditable methodology [1]. The most important conceptual refinement is to distinguish accurately between searching, screening, eligibility, qualitative inclusion, quantitative inclusion, risk of bias and certainty of evidence. As these elements have already been developed in sections 5 to 13, the discussion highlights their consequence: the review is organised as a chain of verifiable decisions, consistent with PRISMA 2020, PRISMA-S, prospective protocol registration, risk-of-bias tools and GRADE [2,3,4,8,9,10,11,12,14,15,16,17,34].
The S0-S3 system offers both pedagogical and operational advantages: it organises a complex process without oversimplifying it. Its structure, relationship with PRISMA, connection with IMRD and application to CSV files, AI and traceability have been presented in sections 5.2, 14 and 17. Its value does not depend merely on naming four phases, but on linking each phase to decisions, files, counts, human validation and reconstructable documentary evidence [2,4,6,18,19,20,26]. In this sense, the prompt ceases to be an informal instruction and becomes part of the study traceability, in the same way as a search strategy, extraction form or analysis file [4,19].

19. Limitations of this Methodological Proposal

This document has a general scope. It does not replace specialised manuals, statistical advice or consultation of specific PRISMA extensions. Some areas are evolving rapidly, particularly regulation of AI use and tools such as QUADAS-3 or ROBINS-I V2 [10,12,19,20]. Teams should therefore verify the current version of each instrument before initiating a review. Nor does this document, by itself, resolve problems of clinical heterogeneity, missing data or low quality in primary studies; rather, it provides a framework for documenting and addressing them transparently.

20. Conclusions

A systematic review should begin with a preliminary phase before the final search is run. At this stage, the problem is defined, the literature is explored, a viable question is formulated, terms and sources are tested, objectives are refined and the protocol is stabilised. This preparation reduces the likelihood of changes during the review and helps ensure that search, selection, extraction, synthesis and reporting answer an explicit and reproducible question. Understood in this way, literature searching is not an auxiliary task for the introduction, but the starting point of a research project and a process for identifying, selecting, appraising and synthesising evidence.
This update reframes a previous teaching document as a methodological guide aligned with current standards. Its contributions have been developed in Table 1 and in the sections on PRISMA, risk of bias, GRADE, RevMan Web, GRADEpro GDT, CSV files, traceability and AI.
The S0-S3 flow is the operational contribution of this approach. Its definition is provided in section 5.2, its correspondence with PRISMA in Table 6, its relationship with IMRD in Table 7 and its documentary application in sections 14 and 15. In summary, it separates records, reports and studies; maintains one CSV matrix per phase; produces verifiable counts for the PRISMA diagram; supports IMRD drafting; and enables the process to be audited from protocol to final manuscript.
Artificial intelligence can be incorporated into the S0-S3 flow as supervised support, provided that it does not replace reviewer judgement. Documentation requirements have been defined in section 17 and set out in Table 14, Annex 2 and Annex 3. Ultimately, the value of S0-S3 lies in turning the review into a process that can be taught, reproduced, audited and transferred into the IMRD report.

Author contributions

José Luis Pardal-Refoyo and Beatriz Pardal-Peláez contributed to conceptualisation, methodological design, writing, critical revision and final approval of the manuscript. Both authors take responsibility for the content and academic integrity of the document.

Funding

This work received no specific funding from public, commercial or not-for-profit sector agencies.

Ethical approval

As it is a methodological guide based on published literature and documentary resources, no approval by a research ethics committee or informed consent was required.

Data Availability Statement

No primary data were generated. The methodological materials, templates and prompts included in the manuscript may be reused and adapted with citation of the original source, unless the journal or repository establishes specific conditions of use.

Acknowledgments

Not applicable.

Conflicts of Interest

The authors declare no competing interests related to the content of this manuscript.

Use of artificial intelligence and automation

During manuscript preparation, writing assistance and methodological review tools were used to organise content, propose wording, review internal coherence and formulate example prompts. All methodological decisions, content selection, critical revision and approval of the final version remained the responsibility of the authors.

Annex 1. Minimum S0-S3 Staged CSV Template

Table 15 provides a minimum template of fields to consistently document the S0, S1, S2, and S3 stages in CSV files.
Table 15. S0-S3 staged CSV field template. 
Table 15. S0-S3 staged CSV field template. 
Stage Minimum fields
S0 id_s0, source, strategy, search_date, source_format, source_file_RIS_BibTeX_CSV, title, authors, year, DOI, PMID, abstract, keywords, MeSH_DeCS, language, document_type, annotation_tags, AI_prompt_S0, AI_output_S0
S1 id_s0, id_s1, duplicate, deduplication_method, title_abstract_decision, AI_priority, AI_justification, AI_prompt_S1, reexportable_RIS_BibTeX, reviewer1, reviewer2, conflict, final_decision
S2 id_s1, id_s2, full_text_retrieved, full_text_decision, exclusion_reason, AI_extracted_data, AI_prompt_S2, reexportable_RIS_BibTeX, observations, excluded_citation
S3 id_s2, study_id, RIS_BibTeX_reference, included_narrative, included_meta_analysis, outcome, extracted_data, AI_extracted_data, AI_prompt_S3, risk_of_bias, GRADE_certainty, human_review, reexportable_by_synthesis, table_figure

Annex 2. Suggested Statement on AI and Automation

When using AI or automation tools that make or suggest judgments, the following text should be adapted and incorporated: “[Tool name, version, date] was used for [specific task] during [phase S0/S1/S2/S3]. The tool received as input bibliographic citation files in CSV format with [fields used: title, authors, year, DOI, PMID, abstract, keywords, MeSH/DeCS terms, source, tags, annotations, previous decision or others]. The prompt used included the review question, inclusion and exclusion criteria, operational definitions of population, intervention/test/exposure, comparator, outcomes, design, fields to be analyzed, decision rules, exit categories, and uncertainty management. Results were [deduplication, classification by inclusion/exclusion criteria, priority to locate full text, preliminary decision to include/maybe/exclude, data extraction, counts for PRISMA or statistical/bibliometric exploration] and were verified by [number] reviewers. Discrepancies were resolved through [procedure]. The search strategies by search engine, dates, number of results, original exports, full prompts, date of use, AI agent used, input file, outputs generated, CSV versions and human validation were preserved in a file of complementary documentation. The authors maintain full responsibility for the methodological decisions and the results of the review.”

Annex 3. Reusable prompts for AI agents in S0-S3

The following prompts are written so that they can be copied and pasted into an AI agent. They must be adapted to the review question, the protocol and the inclusion and exclusion criteria of each project. Each prompt used should be archived in the supplemental documentation file, along with the date, AI agent, version if available, input file, output generated, and human validation. All prompts must explicitly include the inclusion and exclusion criteria corresponding to the methodological stage. These criteria should be shown in square brackets and supplemented by the specific requirements of each systematic review.
Table 16. Reusable prompts for AI agents in stages S0-S3.
Table 16. Reusable prompts for AI agents in stages S0-S3.
Stage Prompt Objective Copyable prompt
PRELIMINARY
Design, pilot search and registrable protocol Translate a research question into a preliminary systematic review plan by defining the methodological framework, key concepts, controlled vocabulary, pilot search strategies, eligibility criteria and PRISMA planning before the development phase begins. Act as a specialist in systematic review methodology and bibliographic searching in the health sciences. I am in the PLANNING PHASE of a systematic review. My research question is: [INSERT QUESTION]. I need you to: (1) formulate the question using the most appropriate framework—PICO, PECO, PCC, PIRD or CoCoPop; (2) identify the population, intervention or exposure, comparator and outcomes; (3) propose free-text terms, synonyms and linguistic variants; (4) suggest MeSH, DeCS, Emtree and other relevant controlled vocabulary terms; (5) design a pilot search strategy for PubMed and provide preliminary adaptations for Europe PMC, Web of Science, Embase, CINAHL, the Cochrane Library, the Virtual Health Library and other relevant databases; (6) propose inclusion and exclusion criteria; (7) identify sentinel studies that the search should retrieve; (8) identify potential methodological limitations; (9) propose data-extraction variables; (10) develop a preliminary PRISMA plan specifying the information that should be recorded throughout the review; and (11) propose a bibliographic export structure in CSV format containing the minimum recommended fields—database, authors, year, title, journal, DOI, PMID, abstract, keywords, language, document type, selection decision and notes—to support the S0–S3 stages. Do not fabricate references, DOIs, PMIDs or registration numbers. Clearly distinguish preliminary planning from the formal conduct of the review. Specify which aspects require human validation before the protocol is drafted and explain how the bibliographic records should subsequently be exported to a CSV file for management and traceability. Generate a CSV file containing your initial selection of sentinel studies and preliminary background sources.
DEVELOPMENT
S0
Normalization of Records
Review a CSV/RIS/BibTeX file exported from databases and prepare a homogeneous matrix of work. Act as a methodological assistant for a systematic review. Analyse the reference file that I provide and normalise the available bibliographic fields. Retain title, authors, year, journal, DOI, PMID, abstract, keywords, MeSH/DeCS terms, source, search date, document type, language and URL if available. Do not delete records at this stage. Return a CSV-exportable table, with one row per record, and add an observations column for fields that are missing, inconsistent or require human review. If the input file is in RIS or BibTeX, retain the identifiers and metadata necessary to re-export selected, excluded or questionable records in RIS or BibTeX without loss of information. Do not infer or generate information that is not present in the metadata. [INCLUSION CRITERIA: specify the criteria applicable to S0, derived from the review question, population, intervention/test/exposure, comparator, outcomes, design, sources and limits defined in the protocol.] [EXCLUSION CRITERIA: specify records clearly irrelevant to the question, excluded document types, excluded languages, dates outside the defined period, inadmissible sources or other limits set in the protocol.]
S0
Duplicate Identification
Detect exact or probable duplicates prior to screening. Act as a bibliographic deduplication assistant. Compare records by DOI, PMID, title, authors, year, journal and title similarity. Classify each possible duplicate as exact duplicate, probable duplicate or non-duplicate. Always retain the record with the most complete metadata and preserve the source of all duplicate records. Return a CSV-exportable table with id_s0, duplicate_group, proposed_decision, retained_record, removable_records, matching_criterion and justification. Retain the identifiers and metadata necessary to re-export preserved, removable, selected, excluded or doubtful records in RIS or BibTeX without loss of information. Do not delete records; mark all decisions as proposed and pending human validation. [INCLUSION CRITERIA: retain unique records or duplicate records that provide complementary metadata useful for S0 traceability.] [EXCLUSION CRITERIA: mark as removable duplicate records with exact or probable matches according to DOI, PMID, title, authors, year or journal, without losing the original source or relevant metadata.]
S0
Initial Thematic Exploration
Identify terms, thematic families, and useful signals for screening. Act as a bibliographic exploration assistant. From titles, abstracts, keywords and MeSH/DeCS terms, identify common terms, synonyms, linguistic variants, related concepts, populations, interventions, exposures, diagnostic tests, comparators, outcomes, study designs and statistical methods mentioned. Return a CSV-exportable matrix with term or concept, approximate frequency, field where it appears, possible relationship with inclusion criteria and observations for reviewing the search strategy or screening. If the matrix is derived from bibliographic references, retain the identifiers necessary to link each thematic signal to the original records and enable re-export in RIS or BibTeX when necessary. Do not modify the protocol criteria; only suggest useful signals for human review. [INCLUSION CRITERIA: identify thematic signals compatible with the review question and with the preliminary criteria for population, intervention/test/exposure, comparator, outcomes and design.] [EXCLUSION CRITERIA: indicate terms, concepts, document types, languages, populations or thematic areas clearly unrelated to the protocol or expressly excluded.]
S1
Screening by Title and Abstract
Classify records according to inclusion and exclusion criteria. Act as the second screening assessor of a systematic review. Use only the available title, abstract, keywords, MeSH/DeCS terms, tags and annotations. Review question: [insert question]. For each record, classify the preliminary decision as include, maybe or exclude. In addition, assign priority for full-text retrieval: high, medium, low or exclude. Justify each decision in one sentence using verifiable metadata. If the information is insufficient, use “maybe” and do not exclude definitively. Return a CSV-exportable table with columns: id_s0, AI_decision, full_text_priority, criterion_applied, justification, missing_data and human_review_required. Retain the identifiers and metadata necessary to re-export included, excluded or doubtful records in RIS or BibTeX without loss of information. [INCLUSION CRITERIA: specify population, intervention/test/exposure, comparator, outcomes, design, setting, period, language and any other criteria applicable to title-and-abstract screening.] [EXCLUSION CRITERIA: specify animal studies, letters, editorials, ineligible reviews, non-relevant population, intervention/test/exposure not relevant, non-relevant outcomes, excluded designs, excluded languages or other protocol criteria.]
S1
Full-text prioritization
Sort the retrieval of full texts by probability of inclusion. Act as a full-text retrieval prioritisation assistant. Review the records that have not been excluded after title-and-abstract screening. Classify each record as high, medium or low priority for full-text retrieval. Consider population match, intervention/test/exposure, comparator, outcomes, design, methods, keywords, MeSH/DeCS terms, information cited in the abstract, cohort or registry names and signals of extractable data. Return a CSV-exportable table with columns: id_s1, priority, primary_reason, potentially_extractable_variables_or_data, methodological_uncertainties and recommendation_for_human_reviewer. Retain the identifiers and metadata necessary to re-export prioritised, excluded or doubtful records in RIS or BibTeX without loss of information. [INCLUSION CRITERIA: prioritise records with clear or probable correspondence to the population, intervention/test/exposure, comparator, outcomes and design defined for S1.] [EXCLUSION CRITERIA: assign low priority or exclude records with clear signals of an ineligible population, excluded document type, inadmissible design, no relationship to the question or insufficient information to justify immediate retrieval.]
S1
Traceability Report
Generate summary for the supplementary file. Act as a methodological documentation assistant. From the S1 screening file, generate a traceability summary with the number of records assessed, records proposed for inclusion, records classified as maybe, records proposed for exclusion, records by high, medium or low full-text priority, main reasons for exclusion, fields used for the decision, records with insufficient information and recommendations for human review. Return the summary in a format suitable for the supplementary documentation file and include a CSV-exportable matrix with the main counts and categories. If the file contains bibliographic references, retain the identifiers and metadata necessary to re-export included, excluded or doubtful records in RIS or BibTeX without loss of information. Add a final statement that all AI decisions are proposals and must be validated by human reviewers. [INCLUSION CRITERIA: summarise as included or potentially included the records that meet or may meet the S1 criteria based on available information.] [EXCLUSION CRITERIA: summarise records excluded or proposed for exclusion using the primary reason applied, avoiding vague categories and distinguishing title-and-abstract exclusion from full-text exclusion.]
S2
Full-Text Eligibility Assessment
Apply definitive eligibility criteria on full-text articles. Act as the second eligibility assessor for a systematic review. Review the full text of the report and explicitly compare the population, intervention/test/exposure, comparator or reference standard, outcomes, design, period, setting, extractable data and possible sample overlap with the protocol. Classify the preliminary decision as include, maybe or exclude. If exclusion is proposed, assign a single main reason and justify it with verifiable information from the text, indicating the section, table, figure, page or supplementary material whenever possible. Do not definitively exclude a report if essential information is missing; in that case, classify it as maybe and mark it for mandatory human review. Return a CSV-exportable table with columns: id_s2, AI_full_text_decision, criterion_applied, primary_exclusion_reason, supporting_excerpt_or_data, location_in_text, missing_data, possible_overlap and human_review_required. Retain the bibliographic identifiers necessary to re-export included, excluded or questionable reports in RIS or BibTeX without loss of metadata. [INCLUSION CRITERIA: specify the definitive eligibility criteria applicable to S2, including population, intervention/test/exposure, comparator or reference standard, outcomes, design, period, setting, data availability and methodological requirements defined in the protocol.] [EXCLUSION CRITERIA: specify the single reason for exclusion to be recorded and justified, such as ineligible population, non-relevant intervention/test/exposure, inadequate comparator or reference standard, non-relevant outcomes, excluded design, absence of extractable data, non-retrievable full text, duplicate population or unresolved overlap.]
S2
Full-Text Exclusions Table
Standardize reasons for exclusion and prepare the PRISMA annex. Act as a methodological documentation assistant for S2. From the full-text eligibility file, review all excluded or questionable reports and construct an exclusion table with abbreviated citation, id_s2, exclusion_stage, proposed_final_decision, primary_exclusion_reason, PRISMA_compatible_category, brief_justification, location_of_evidence_in_text and observations_for_human_review. Do not use vague reasons such as “not relevant” when a methodological reason can be specified. If more than one reason applies, record only the main reason and mention secondary reasons in the observations column. Return a CSV-exportable matrix and a summary with the number of full texts assessed, excluded by category, pending and proposed for inclusion. Retain the bibliographic identifiers necessary to re-export included, excluded or questionable reports in RIS or BibTeX without loss of metadata. [INCLUSION CRITERIA: consider eligible the full texts that meet the final S2 criteria and provide sufficient information for the intended qualitative or quantitative synthesis.] [EXCLUSION CRITERIA: record a single main reason per excluded text, justified with an explicit and traceable category: population, intervention/test/exposure, comparator or reference standard, outcomes, design, data, duplication, non-retrievable text or another predefined cause.]
S3
Structured Data Extraction
Extract variables, outcomes and numerical data for synthesis. Act as a data extraction assistant for a systematic review. From the full text of the included studies, extract only information present in the article or in its supplementary materials. Record study identification, country, design, period, data source, population, sample size, intervention/test/exposure, comparator or reference standard, outcomes, operational definitions, numerical data, units of measurement, corresponding group or arm, effect measures, confidence intervals, 2×2 tables if applicable, events, denominators, means, standard deviations, losses to follow-up, funding, conflicts of interest and methodological notes. For each item, indicate whether it comes from the main text or the supplement and specify the section, table, figure, page or source material. Always separate original data, transformed or calculated data and the justification for any transformation. Do not compute derived data unless expressly requested, and explain any transformation performed. Return a CSV-exportable table with one row per study and outcome, including columns for original_data, transformed_data, unit_of_measurement, comparator_group, exact_source, location_in_text, methodological_uncertainty and human_review. Retain bibliographic identifiers that allow each study to be linked to its original reference and, where appropriate, re-export included references in RIS or BibTeX without loss of metadata. [INCLUSION CRITERIA: extract data only from studies definitively included in S3 for narrative, quantitative or outcome synthesis, based on predefined variables in the protocol and extraction form.] [EXCLUSION CRITERIA: do not extract data from excluded records, duplicate studies, reports with unresolved overlapping populations, unpredefined outcomes, unverifiable data or inferred information that is not present in the text.]
S3
Preparation of synthesis and meta-analysis
Identify studies and outcomes suitable for narrative or quantitative synthesis. Act as a synthesis preparation assistant. Review the S3 matrix and group studies by question, comparison, outcome, design, population, intervention/test/exposure, comparator or reference standard, effect measure and data availability. Indicate for each study whether it can contribute to narrative synthesis, meta-analysis, subgroup analysis, sensitivity analysis or tabular description only. Do not recommend meta-analysis solely because data are available; also assess clinical, methodological and statistical compatibility, consistency of definitions, unit of analysis, direction of effect and risk of overlap. Justify exclusions from quantitative synthesis on the basis of clinical heterogeneity, methodological or statistical incompatibility, absence of comparable data, sample overlap or non-combinable outcomes. Return a CSV-exportable table with columns: study_id, outcome, possible_synthesis_type, effect_measure, available_data, clinical_compatibility, methodological_compatibility, statistical_compatibility, reason_no_meta_analysis and human_review. Retain bibliographic identifiers so that each study can be linked to its reference and so that references included in or excluded from each synthesis can be re-exported in RIS or BibTeX when necessary. [INCLUSION CRITERIA: include in each synthesis definitively included studies that share population, comparison, outcome, design, effect measure and clinical, methodological and statistical conditions sufficiently compatible for the intended synthesis.] [EXCLUSION CRITERIA: exclude from a specific synthesis studies included in the review that do not provide combinable data, present unjustifiable heterogeneity, have incompatible outcomes, unresolved overlapping populations or insufficient information for analysis.]
S3
Statistical calculation and verification of the meta-analysis
Execute or verify statistical calculations with traceable and reproducible tools. Act as a statistical verification assistant for a systematic review with meta-analysis. From the validated S3 matrix, identify for each outcome the appropriate effect measure, the type of data required, the intended statistical model and the software or analytical tool used for the calculation, for example Copilot Analyst, R with the meta, metafor, mada, diagmeta or netmeta packages, Stata with meta, metan, metareg, midas or metandi, RevMan Web, JASP, Jamovi, Comprehensive Meta-Analysis, MedCalc or another specialised tool. Verify that the input data are sufficient and consistent: events and denominators, means and standard deviations, 2×2 tables, hazard ratios, odds ratios, risk ratios, mean differences, confidence intervals, standard errors or diagnostic matrices, as appropriate. Do not invent missing data or replace statistical decisions not specified in the protocol. Return a CSV-exportable table with columns: study_id, outcome, data_type, effect_measure, planned_model, software_or_agent, package_or_command, required_input_data, available_data, issues_detected, reproducible_calculation, script_or_output_file, main_result and human_review. Retain bibliographic identifiers to link calculations to original references and to allow re-export in RIS or BibTeX of studies included in each analysis when necessary. If code is generated, separate it from the results and state that it must be run and verified in the corresponding software. [INCLUSION CRITERIA: include in the statistical calculation only S3 studies and outcomes with sufficient, comparable and validated data for the intended quantitative synthesis, with a defined or justified effect measure, model and statistical tool.] [EXCLUSION CRITERIA: exclude from the calculation studies without sufficient numerical data, with non-combinable populations or outcomes, unresolved duplicate or overlapping data, incompatible effect measures, uncorrected errors, lack of correspondence with the protocol or lack of human validation of the S3 matrix.]
S3
Final Traceability Report S0-S3
Consolidate counts, decisions, and human validation for supplemental material. Act as a final traceability assistant for a systematic review. Integrate the S0, S1, S2 and S3 files and generate a methodological summary with PRISMA counts, distinguishing records, reports/full texts, studies and studies included in each synthesis. Report the number of records identified, duplicates removed, records screened, records excluded by title and abstract, reports sought for retrieval, reports not retrieved, full-text reports assessed, full-text exclusions by reason, studies included in narrative synthesis and studies included in each meta-analysis. Return a CSV-exportable matrix with counts, decisions and correspondence between phases, and add a table of prompts used with phase, objective, input file, output generated and human validation. Retain the bibliographic identifiers necessary to re-export included, excluded or doubtful records or studies by phase and by synthesis in RIS or BibTeX without loss of metadata. Identify inconsistencies between phases, residual duplicates, records without a final decision, reports without an exclusion reason, included studies without extracted data or PRISMA counts that do not correspond to the source files. Return the report in a format suitable for supplementary material. [INCLUSION CRITERIA: integrate only validated decisions or pending decisions clearly marked in S0-S3, distinguishing identified records, assessed reports, studies included in the review, studies included in narrative synthesis and studies included in each meta-analysis.] [EXCLUSION CRITERIA: do not mix previously excluded records with included studies, do not count duplicates as independent studies, do not confuse reports with studies, do not assign undocumented exclusion reasons and do not generate PRISMA counts without correspondence to the source files.]

References

  1. Pardal-Refoyo JL, Pardal-Peláez B. Anotaciones para estructurar una revisión sistemática. Rev ORL. 2020;11(2):155-60. [CrossRef]
  2. Page MJ, McKenzie JE, Bossuyt PM, Boutron I, Hoffmann TC, Mulrow CD, et al. The PRISMA 2020 statement: an updated guideline for reporting systematic reviews. BMJ. 2021;372:n71. [CrossRef]
  3. Page MJ, Moher D, Bossuyt PM, Boutron I, Hoffmann TC, Mulrow CD, et al. PRISMA 2020 explanation and elaboration: updated guidance and exemplars for reporting systematic reviews. BMJ. 2021;372:n160. [CrossRef]
  4. Rethlefsen ML, Kirtley S, Waffenschmidt S, Ayala AP, Moher D, Page MJ, et al. PRISMA-S: an extension to the PRISMA Statement for Reporting Literature Searches in Systematic Reviews. Syst Rev. 2021;10(1):39. [CrossRef]
  5. McInnes MDF, Moher D, Thombs BD, McGrath TA, Bossuyt PM, Clifford T, et al. Preferred Reporting Items for a Systematic Review and Meta-analysis of Diagnostic Test Accuracy Studies: the PRISMA-DTA Statement. JAMA. 2018;319(4):388-96. [CrossRef]
  6. Higgins JPT, Thomas J, Chandler J, Cumpston M, Li T, Page MJ, et al., editors. Cochrane Handbook for Systematic Reviews of Interventions. Version 6.5. Cochrane; 2024. Available from: www.cochrane.org/handbook.
  7. Deeks JJ, Bossuyt PM, Leeflang MM, Takwoingi Y, editors. Cochrane Handbook for Systematic Reviews of Diagnostic Test Accuracy. Version 2.0. Cochrane; 2023. Available from: https://training.cochrane.org/handbook-diagnostic-test-accuracy.
  8. Sterne JAC, Savović J, Page MJ, Elbers RG, Blencowe NS, Boutron I, et al. RoB 2: a revised tool for assessing risk of bias in randomised trials. BMJ. 2019;366:l4898. [CrossRef]
  9. Sterne JAC, Hernán MA, Reeves BC, Savović J, Berkman ND, Viswanathan M, et al. ROBINS-I: a tool for assessing risk of bias in non-randomised studies of interventions. BMJ. 2016;355:i4919. [CrossRef]
  10. Risk of Bias Tools. ROBINS-I V2 tool: risk of bias in non-randomized studies of interventions. Draft version. Bristol: University of Bristol; 2025. Available from: https://www.riskofbias.info/.
  11. Whiting PF, Rutjes AWS, Westwood ME, Mallett S, Deeks JJ, Reitsma JB, et al. QUADAS-2: a revised tool for the quality assessment of diagnostic accuracy studies. Ann Intern Med. 2011;155(8):529-36. [CrossRef]
  12. Whiting PF, Tomlinson E, Rutjes AWS, Davenport CF, Yang B, Westwood ME, et al. QUADAS-3: a revised tool for the quality assessment of diagnostic test accuracy studies. Ann Intern Med. 2026;179(4):548-55. [CrossRef]
  13. Wells GA, Shea B, O’Connell D, Peterson J, Welch V, Losos M, et al. The Newcastle-Ottawa Scale (NOS) for assessing the quality of nonrandomised studies in meta-analyses. Ottawa: Ottawa Hospital Research Institute; 2021. Available from: https://www.ohri.ca/programs/clinical_epidemiology/oxford.asp.
  14. Schünemann H, Brożek J, Guyatt G, Oxman A, editors. GRADE Handbook for grading quality of evidence and strength of recommendations. Updated October 2013. GRADE Working Group; 2013. Available from: https://gdt.gradepro.org/app/handbook/handbook.html.
  15. Cochrane. RevMan: review-writing software. London: Cochrane; 2026. Available from: https://revman.cochrane.org/.
  16. GRADEpro GDT. GRADEpro guideline development tool: online integration with RevMan Web. Hamilton: Evidence Prime; 2026. Available from: https://www.gradepro.org/.
  17. Booth A, Clarke M, Dooley G, Ghersi D, Moher D, Petticrew M, et al. The nuts and bolts of PROSPERO: an international prospective register of systematic reviews. Syst Rev. 2012;1:2. [CrossRef]
  18. Campbell M, McKenzie JE, Sowden A, Katikireddi SV, Brennan SE, Ellis S, et al. Synthesis without meta-analysis (SWiM) in systematic reviews: reporting guideline. BMJ. 2020;368:l6890. [CrossRef]
  19. Flemyng E, Noel-Storr A, Macura B, Gartlehner G, Thomas J, Meerpohl JJ, et al. Position statement on artificial intelligence use in evidence synthesis across Cochrane, the Campbell Collaboration, JBI and the Collaboration for Environmental Evidence. Cochrane Database Syst Rev. 2025;2025:ED000178. [CrossRef]
  20. Holst D, Moenck K, Koch J, Schmedemann O, Schüppstuhl T. Transparent reporting of AI in systematic literature reviews: development of the PRISMA-trAIce checklist. JMIR AI. 2025;4:e80247. [CrossRef]
  21. Stern C, Jordan Z, McArthur A. Developing the review question and inclusion criteria. Am J Nurs. 2014;114(4):53-56. [CrossRef]
  22. Reitsma JB, Glas AS, Rutjes AWS, Scholten RJPM, Bossuyt PM, Zwinderman AH. Bivariate analysis of sensitivity and specificity produces informative summary measures in diagnostic reviews. J Clin Epidemiol. 2005;58(10):982-90. [CrossRef]
  23. Rutter CM, Gatsonis CA. A hierarchical regression approach to meta-analysis of diagnostic test accuracy evaluations. Stat Med. 2001;20(19):2865-84. [CrossRef]
  24. Brignardello-Petersen R, Santesso N, Guyatt GH. Systematic reviews of the literature: an introduction to current methods. Am J Epidemiol. 2025;194(2):536-542. [CrossRef]
  25. McKenzie JE, Brennan SE, Ryan RE, Thomson HJ, Johnston RV. Chapter 9: Summarizing study characteristics and preparing for synthesis. In: Higgins JPT, Thomas J, Chandler J, Cumpston M, Li T, Page MJ, et al., editors. Cochrane Handbook for Systematic Reviews of Interventions. Version 6.5. Cochrane; 2024. Available from: https://www.cochrane.org/authors/handbooks-and-manuals/handbook/current/chapter-09.
  26. von Pressentin KB, Shabani JS, Young T. Integrating evidence synthesis into doctoral research: a guide for family medicine and primary care. Afr J Prim Health Care Fam Med. 2025;17(2):a5198. [CrossRef]
  27. Whaley P, Garside R, Eales JF. A General Protocol for Pilot-Testing the Screening Stage of a Systematic Review. protocols.io. 2020. [CrossRef]
  28. Ruangsomboon O, Lima JP, Eltorki M, Worster A. Methodological standards in the design and reporting of pilot and feasibility studies in emergency medicine literature: a systematic review. BMJ Open. 2024;14(11):e082648. [CrossRef]
  29. Ringsten M, Styrmisdottir L, Naesström M, Johansson M, Bruschettini M, Wallerstedt SM. Systematic Reviews as Part of Doctoral Theses and for the Promotion to Associate Professor: A Descriptive Study of University Policies in Sweden. Cochrane Evid Synth Methods. 2026;4(1):e70069. [CrossRef]
  30. PRISMA Statement. PRISMA 2020 statement. Oxford: PRISMA; 2026. Available from: https://www.prisma-statement.org/prisma-2020-statement.
  31. BIREME/PAHO/WHO. DeCS – Health Sciences Descriptors. São Paulo: Virtual Health Library; 2026. Available from: https://decs.bvsalud.org/.
  32. Pan American Health Organization. DeCS 2025 edition is available. Washington, DC: PAHO/WHO; 2025. Available from: https://www.paho.org/es/noticias/30-4-2025-decs-edicion-2025-esta-disponible.
  33. Bramer WM, de Jonge GB, Rethlefsen ML, Mast F, Kleijnen J. A systematic approach to searching: an efficient and complete method to develop literature searches. J Med Libr Assoc. 2018;106(4):531-541. [CrossRef]
  34. McGowan J, Sampson M, Salzwedel DM, Cogo E, Foerster V, Lefebvre C. PRESS Peer Review of Electronic Search Strategies: 2015 guideline statement. J Clin Epidemiol. 2016;75:40-46. [CrossRef]
  35. Atkinson LZ, Cipriani A. How to carry out a literature search for a systematic review: a practical guide. BJPsych Adv. 2018;24(2):74-82. [CrossRef]
  36. Calderon Martinez E, Flores Valdés JR, Castillo JL, Castillo JV, Blanco Montecino RM, Morin Jimenez JE, et al. Ten steps to conduct a systematic review. Cureus. 2023;15(12):e51422. [CrossRef]
  37. Covidence. A practical guide: protocol development for systematic reviews. Melbourne: Veritas Health Innovation; 2024. Available from: https://www.covidence.org/.
  38. Pardal Refoyo JL, Soto Varela A, López Poveda E, Caballero Borrego M. Estrategias y herramientas de inteligencia artificial para mejorar la investigación bibliográfica en Otorrinolaringología. Libro Blanco SEORL-CCC 2025. Madrid: Sociedad Española de Otorrinolaringología y Cirugía de Cabeza y Cuello; 2025.
  39. Pardal Refoyo JL, Soto Varela A, López Poveda EA, Caballero Borrego M. Metodología para realizar una revisión bibliográfica. En: Actualización SEORL 2026. Capítulo 110. Madrid: Sociedad Española de Otorrinolaringología y Cirugía de Cabeza y Cuello; 2026. En prensa. Disponible en: https://www.actualizacion.seorl.net/article?id=6a27d7e8-da6c-4be1-a076-74100aca0133&columns=1.
Figure 1. Correspondence between the S0-S3 flow and the PRISMA 2020 flow diagram. S0 identifies all records; S1 documents screening by title, abstract and keywords; S2 records full-text eligibility and reasons for exclusion; and S3 identifies the studies included in narrative synthesis and/or meta-analysis.
Figure 1. Correspondence between the S0-S3 flow and the PRISMA 2020 flow diagram. S0 identifies all records; S1 documents screening by title, abstract and keywords; S2 records full-text eligibility and reasons for exclusion; and S3 identifies the studies included in narrative synthesis and/or meta-analysis.
Preprints 225679 g001
Figure 2. Stage S0: identification. The figure summarises the inputs, processes, decisions and deliverables involved in running the systematic search and constructing the raw record base.
Figure 2. Stage S0: identification. The figure summarises the inputs, processes, decisions and deliverables involved in running the systematic search and constructing the raw record base.
Preprints 225679 g002
Figure 3. Stage S1: screening. The figure shows the transition from raw records to potentially eligible records, distinguishing deduplication, human title-and-abstract screening and AI-assisted prioritisation.
Figure 3. Stage S1: screening. The figure shows the transition from raw records to potentially eligible records, distinguishing deduplication, human title-and-abstract screening and AI-assisted prioritisation.
Preprints 225679 g003
Figure 4. Stage S2: eligibility. The figure represents full-text assessment, application of eligibility criteria and documentation of explicit reasons for exclusion.
Figure 4. Stage S2: eligibility. The figure represents full-text assessment, application of eligibility criteria and documentation of explicit reasons for exclusion.
Preprints 225679 g004
Figure 5. Stage S3: final inclusion. The figure shows how eligible studies proceed to data extraction, risk-of-bias assessment, evidence synthesis and reporting.
Figure 5. Stage S3: final inclusion. The figure shows how eligible studies proceed to data extraction, risk-of-bias assessment, evidence synthesis and reporting.
Preprints 225679 g005
Figure 6. Document audit and traceability from S0 to the manuscript. The figure shows how the protocol, search strategies, selection decisions, extracted data, analyses and reporting outputs should be connected in a verifiable chain.
Figure 6. Document audit and traceability from S0 to the manuscript. The figure shows how the protocol, search strategies, selection decisions, extracted data, analyses and reporting outputs should be connected in a verifiable chain.
Preprints 225679 g006
Figure 7. Practical organisation of folders, files and traceability materials for a systematic review.
Figure 7. Practical organisation of folders, files and traceability materials for a systematic review.
Preprints 225679 g007
Table 1. Changes incorporated with respect to the previous approach.
Table 1. Changes incorporated with respect to the previous approach.
Area Change from the previous approach Practical justification
PRISMA Replace PRISMA 2009 with PRISMA 2020 and its relevant extensions [2,3]. PRISMA 2020 includes a list of 27 items, updated diagram, automation report, data availability and certainty of evidence.
Protocol Make the protocol and prospective registration explicit. Registration, access to the protocol, amendments, and justification must be indicated if not registered.
Search Document complete strategies by source. PRISMA-S requires bases, platforms, dates, limits, filters, deduplication, and peer review of searches where appropriate [4].
Risk of bias Abandon the generic idea of “quality” as a total score. Use tools by design: RoB 2, ROBINS-I, QUADAS-2/3, Newcastle-Ottawa when warranted [8,9,10,11,12,13].
Certainty Distinguish risk of bias from global certainty. GRADE assesses confidence in the body of evidence by outcome, not the isolated quality of each article [14].
Software Upgrade RevMan 5 to RevMan Web and companion tools. RevMan Web, GRADEpro GDT, Rayyan, Covidence, EPPI-Reviewer, Zotero/EndNote/Mendeley, and open repositories [15,16].
AI Add transparency and human oversight. Any automation that suggests decisions must be justified, validated, monitored, and reported [19,20].
Table 2. Writing strategy for the report in IMRD format. The table separates the starting point, the order of writing and the function of the bibliography in introduction, methods, results and discussion.
Table 2. Writing strategy for the report in IMRD format. The table separates the starting point, the order of writing and the function of the bibliography in introduction, methods, results and discussion.
Practical Strategy for Writing the IMRD Final Report
1. Starting point
A clear, well-argued question derived from a preliminary bibliographic exploration. From it come the title, the keywords, the criteria, the search, the analysis and the interpretation.
2. Drafting order
Objectives → methods → analysis/results → interpretation/discussion → conclusions → introduction → abstract in Spanish and English. The manuscript reads like IMRD, but is written from the most verifiable elements.
3. Differentiated bibliography
The introduction uses contextual and preliminary bibliography. The methods cite guides and tools. Results and discussion are supported by the included studies and the evidence needed to interpret the findings.
S0-S3 maintains traceability: S0 formulates and justifies the question; S1-S2 document the selection; S3 defines the core evidence of results and discussion.
Table 3. Principle for formulating the review question.
Table 3. Principle for formulating the review question.
Principle
The review question functions as the methodological blueprint for the entire review. If the question changes after examining the results, the change must be recorded as an amendment and justified.
Table 5. Relationship between the preliminary phase and the development phase of bibliographic research.
Table 5. Relationship between the preliminary phase and the development phase of bibliographic research.
Literature research for a systematic review: two connected phases
Preliminary phase
Design the research before the definitive systematic search: delimit the problem, review exploratory literature, identify gaps, formulate initial question, conduct pilot search, adjust objectives, criteria, terms, and strategy, and write the registrable protocol [17,21,37].

Stabilized question and protocol
Development phase
Execute the review already designed using S0-S3: identify records, screen titles and abstracts, evaluate full texts, include studies, extract data, assess risk of bias, synthesize evidence, and prepare the report [2,3,4,6].
1. Define
problems, concepts and gaps
2. Pilot
terms, sources and criteria
3. Record
protocol and plan
4. Run S0-S3
Identify, Screen, Choose, and Include
5. Report
PRISMA, synthesis and IMRD
Rule of thumb: The preliminary phase allows the question to be adjusted and subsequent changes to be reduced; the development phase executes the protocol and documents any substantial changes as an amendment [2,17,30].
Table 6. Correspondence between the S0-S3 flow and PRISMA 2020.
Table 6. Correspondence between the S0-S3 flow and PRISMA 2020.
Schematic S0-S3 Correspondence with PRISMA 2020 Documented decision
S0 Identification
All records
Records identified in databases, records, gray literature, or other sources [2,4]. Total number retrieved by source and records deleted prior to screening, including duplicates.

S1 Screening
Title, abstract and keywords
Screened records and excluded records after reading the title and abstract, with reviewers, independence and automation documented where appropriate [2,3]. Preliminary decision: high priority, medium priority, low priority, or exclude.

S2 Eligibility
Full text and extractable data
Reports sought for recovery, reports evaluated for eligibility, and reports excluded for explicit reason [2,3]. Full-text decision: include, perhaps, or exclude, with primary reason for exclusion.

S3 Final inclusion
Qualitative synthesis and/or meta-analysis
Studies included in the review and, if applicable, studies included in each meta-analysis, narrative synthesis, or outcome synthesis [6,18,25]. Final inclusion for narrative synthesis, quantitative synthesis, outcome analysis, and evidence tables.
Table 7. Relationship between the S0-S3 flow, the PRISMA diagram and the IMRD wording.
Table 7. Relationship between the S0-S3 flow, the PRISMA diagram and the IMRD wording.
Stage Contributes to the PRISMA diagram Contribute to IMRD manuscript
S0 Identification Records identified by source and records deleted prior to screening, including duplicates [2,4,33]. Introduction: justification and question. Methods: protocol, sources, search strategies, limits, and record management.
S1 Screening Screened records and excluded records by title, abstract and keywords, with a transparent and reproducible selection process [2,3]. Methods: screening process, reviewers, discrepancies and automation. Results: number of excluded records.
S2 Eligibility Reports searched, not retrieved, evaluated in full text and excluded with explicit reason [2,3]. Methods: Eligibility criteria applied to full text. Results: exclusions with reason and table/annex of excluded texts.
S3 Final Inclusion Studies included in the review, in the narrative synthesis and, if appropriate, in each meta-analysis or synthesis by outcome [6,18,25]. Results: characteristics, risk of bias, synthesis, and GRADE. Discussion and conclusions: interpretation, limitations and implications.
Table 8. Fields for the S0 CSV file.
Table 8. Fields for the S0 CSV file.
CSV S0 Block Recommended fields
Identifier id_s0; source; search number; Date
Reference authors; year; title; journal; volume; pages; DOI/PMID/URL
Summary abstract; keywords; language; Document type
Search search equation; platform; filters; limits; number of results
Traceability original exported file; RIS/BibTeX/CSV export format; download date; responsible; prompt_IA_S0 if applicable; salida_IA_S0 if applicable
Table 9. Recommendation for the use of AI in S1.
Table 9. Recommendation for the use of AI in S1.
Responsible use of AI in S1
AI can suggest or prioritize decisions, but it cannot replace the reviewer. If you classify records, inputs, outputs, version, date, prompts or parameters, internal validation, and human review must be preserved. The prompt should be archived with the search strategies and include parsed fields, inclusion/exclusion rules, priority categories, output format, and justification for each recommendation with verifiable metadata.
Table 10. Exclusion categories in S2.
Table 10. Exclusion categories in S2.
Exclusion category in S2 Operational definition
Population Does not match PICO/PIRD; age, disease, context, or ineligible subgroup.
Intervention/Test/Exposure The intervention, index test or exposure does not correspond to the review question.
Comparator or reference Inadmissible comparator or inadequate reference standard.
Outcomes Does not report the predefined outcomes or does not allow construction of effect or accuracy measures.
Design Type of study excluded by protocol.
Data Insufficient, non-extractable data or non-resolvable overlap.
Table 11. Tools for assessing risk of bias according to study design.
Table 11. Tools for assessing risk of bias according to study design.
Design Recommended tool Methodological use
Randomized trials RoB 2 Evaluates specific results from randomized trials; Includes versions for cluster and crossover assays.
Non-randomised studies of interventions ROBINS-I / ROBINS-I V2 Domain-by-domain approach against a hypothetical target trial; attention to confounding, selection, classification and missing data.
Diagnostic accuracy QUADAS-2; QUADAS-3 QUADAS-2 retains extended use; QUADAS-3 appears as a revised tool focused on estimates of accuracy and applicability.
Observational cohorts/case-controls Newcastle-Ottawa Scale Star-based tool for selection, comparability and exposure/outcome assessment; useful when its adaptation and limitations are declared.
Previous systematic reviews AMSTAR 2 / ROBIS To rate reviews included in umbrella reviews or overviews.
Predictive models PROBAST / TRIPOD When the question evaluates prognostic models or predictive diagnoses.
Table 12. PRISMA guides and extensions according to the type of review.
Table 12. PRISMA guides and extensions according to the type of review.
Guide/Extension When to use it
PRISMA 2020 General report of systematic reviews and meta-analyses.
PRISMA-S Detailed report of bibliographic searches.
PRISMA-DTA Diagnostic accuracy reviews.
PRISMA-P Systematic review protocols.
PRISMA-ScR Scoping reviews.
PRISMA-NMA Network meta-analysis.
PRISMA-Harms Harms or adverse events.
SWiM Synthesis without meta-analysis of effect estimates.
Table 13. Organization of folders and materials of a systematic review.
Table 13. Organization of folders and materials of a systematic review.
Folder Recommended content
/00_protocol Protocol, registration, amendments and PRISMA-P checklist.
/01_searches_S0 Complete strategies by search engine, dates, filters, limits, number of results obtained in S0, captures, original RIS/BibTeX/CSV exports, searches stored in Zotero or another manager, CSV S0 and complementary documentation file.
/02_screening_S1 Deduplication, title/abstract decisions, conflicts, CSV S1, and, if applicable, RIS/BibTeX of included, excluded, or questionable records.
/03_full_texts_S2 PDFs, table of exclusions, reasons, CSV S2, and, if applicable, RIS/BibTeX of included, excluded, or questionable reports.
/04_extraction_S3 Forms, variable dictionary, extraction, risk of bias, CSV S3, Excel file with result sheets and data matrices, and RIS/BibTeX references of studies included or excluded by synthesis.
/05_analysis R/Stata/RevMan Web scripts, forest plots, SROC curves and sensitivity analyses.
/06_grade SoF tables, evidence profiles, GRADE justifications.
/07_manuscript Word document of the research report, main text, tables, figures, annexes, final checklist and versions sent to the journal or repository.
/08_repository Shareable data, code, data matrices, supplementary materials, reproducible supplementary document, DOI if applicable.
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.
Prerpints.org logo

Preprints.org is a free preprint server supported by MDPI in Basel, Switzerland.

Subscribe

© 2026 MDPI (Basel, Switzerland) unless otherwise stated

Accessibility

Disclaimer

Terms of Use

Privacy Policy

Privacy Settings