Preprint
Essay

This version is not peer-reviewed.

Citation Auditability in AI-Assisted Research: From Reference Verification to Research Integrity

Submitted:

20 September 2026

Posted:

21 September 2026

You are already at the latest version

Abstract
Generative artificial intelligence (GenAI) is transforming how scholarly references are discovered, synthesised, and incorporated into research. Yet a reference can exist, be bibliographically accurate, and appear relevant while remaining inadequately examined as evidence for the claim it supports. Existing approaches address reference verification, contextual support, AI disclosure, provenance, and human accountability, but give less explicit attention to whether the evidentiary pathway connecting an individual citation to a scholarly claim remains reconstructable. This conceptual article defines citation auditability as the degree to which the provenance, verification, evidentiary use, and human responsibility associated with a scholarly citation can be reconstructed and independently assessed. The framework comprises four interrelated dimensions: citation provenance, verification, evidentiary traceability, and human accountability. A 2 × 2 matrix further distinguishes citation auditability from citation correctness, demonstrating that a citation may be correct yet poorly auditable, or incorrect yet auditable. The article argues that citation integrity in AI-assisted scholarship should extend beyond verifying bibliographic objects toward ensuring that the evidentiary pathway from source to evidence to claim remains reconstructable and attributable to responsible human judgment. The article also considers the limits of this framework, including the risk that auditability claims may themselves be difficult to verify independently.
Keywords: 
;  ;  ;  ;  ;  
Subject: 
Social Sciences  -   Education

1. Introduction

A reference can be real but still be wrong. An article exists with correct authors, publication details, and a DOI that resolves correctly. However, these do not confirm it supports the cited proposition. This distinction matters more as GenAI takes over literature searching, synthesis, drafting, and referencing.
Early concerns about AI-assisted referencing mainly centred on hallucinated and inaccurate citations and large language models that produce completely fabricated references and incorrect bibliographic details for real publications (Walters & Wilder, 2023). More recent research shifted focus from reference existence to evidentiary validity. Jung et al. (2026), for instance, differentiate between verifying a reference’s existence, confirming its bibliographic information, and assessing whether the publication genuinely supports the cited claim. This is crucial because a true, correctly formatted reference can still be misused as evidence.
Around the same time, discussions regarding the responsible utilisation of AI in academic publishing are progressively extending beyond mere disclosure to encompass human accountability (Van Zoonen et al., 2026; Frimpong, 2026). Van Zoonen et al. (2026) introduce the concept of claim accountability, asserting that an identifiable human researcher must retain the ability to reconstruct and substantiate scholarly claims generated through AI-assisted methodologies. While these advancements strengthen the governance framework for AI-assisted research, they raise a pointed question: what constitutes the auditable nature of an individual citation?
Imagine two researchers citing the same genuine article for the same claim. One directly retrieved the original publication, carefully examined the evidence, evaluated whether it supports the claim, and took responsibility for that interpretation. The other accepted the citation from an AI-generated summary, confirming only that the article exists. Although their citations may look identical, the underlying evidence and verification processes differ. Checking citation correctness alone does not reveal these differences.
The concept of citation auditability has been introduced. It is defined as the degree to which the provenance, verification, evidentiary application, and human accountability associated with a scholarly citation can be reconstructed and independently assessed. This concept does not replace existing concerns regarding hallucinated references, miscitations, citation verification, or claim accountability. Instead, it shifts the focus from perceiving a citation as a final bibliographic entity to scrutinising the scholarly process that converts a source into evidence.
The framework comprises four interrelated dimensions: citation provenance, concerning how a reference enters the research process; verification, concerning what has actually been checked; evidentiary traceability, concerning whether the relationship between source and claim can be reconstructed; and human accountability, concerning who assumes responsibility for accepting the source as evidence.
The argument does not posit that GenAI has engendered unreliable citation practices. Researchers have historically misinterpreted, miscited, or inadequately examined sources. Instead, GenAI changes the conditions under which such deficiencies may manifest by expediting literature identification and synthesis, while potentially increasing the distance between researchers and original sources (Kitchenham et al., 2026; Clark et al., 2025). Consequently, a plausible AI-generated citation gains scholarly credibility before its evidentiary foundation is independently scrutinised (Walters & Wilder, 2023; Zhao et al., 2026).
This article argues that citation verification in AI scholarship should include citation auditability. The key issue is not just whether a reference is real, accurate, or appropriate, but whether its evidentiary pathway can be reconstructed and linked to responsible human judgment.

2. From Citation Error to Citation Auditability

Citation problems predate generative AI, including inaccurate quoting, improper sourcing, and the propagation of unsupported claims (Jergas & Baethge, 2015; Mogull, 2017). Greenberg (2009) showed that citation bias and distortion can turn weak claims into accepted knowledge. These findings underscore persistent concerns about research integrity.
Generative AI (GenAI) does not create this problem; rather, it alters the conditions under which it arises. Large language models produce fabricated references and inaccurate bibliographic information for authentic publications, often in forms that appear credible enough to evade superficial scrutiny (Walters & Wilder, 2023). This situation has understandably prompted increased scrutiny regarding reference verification. Nevertheless, establishing the existence of a publication and verifying the accuracy of its metadata does not inherently confirm that it supports the proposition for which it is cited. A genuine source may pertain to a different population, report an association presented as causal, yield a more qualified conclusion than the citing text implies, or simply lack sufficient evidence to substantiate the claim (Mogull, 2017).
Studies have started to clarify this distinction more clearly. Jung et al. (2026) differentiate between reference existence, metadata accuracy, and support for contextual claims, thereby separating bibliographic validity from evidentiary validity. Similarly, work on AI-driven literature searches has incorporated auditability during retrieval. León and Kudelka (2026), for instance, highlight an auditable evidence trail based on DOI traceability, metadata correctness, retrieval reliability, and reproducibility. Their approach, however, mainly targets the search and retrieval process, not the later process by which a source becomes linked to a specific scholarly claim.
This distinction is significant because a legitimate, yet evidentially inappropriate citation is more challenging to detect than a fabricated reference. Such citations may pass through DOI resolution, metadata verification, reference-management software, and standard editorial review (Greenberg, 2009; Aronsky, 2004). Bibliographic authenticity generates a perception of evidentiary validity even when the relationship between source and claim has not been thoroughly scrutinised (Greenberg, 2009; Gasparyan et al., 2015). Generative AI exacerbates this issue, as source discovery, summarisation, interpretation, and citation recommendation can all occur within a single mediated interaction.
The governance question consequently extends beyond verification. Current discussions of responsible AI utilisation in scholarly publishing increasingly scrutinise whether disclosure of AI involvement alone suffices to establish research integrity (Frimpong, 2026; Rajkovic & Wang, 2026). Van Zoonen et al. (2026) develop the concept of claim accountability, asserting that an identifiable human researcher should still be able to reconstruct and defend scholarly claims generated through AI-assisted processes. This appropriately shifts responsibility from the technological instrument to the human author.
Citation auditability extends these developments but targets a more specific aspect of analysis. While claim accountability focuses on whether a researcher can reconstruct and justify a scholarly assertion, citation auditability examines if the evidentiary link between a cited source and that claim can be traced and evaluated. Similarly, retrieval auditability pertains to whether the process of identifying literature is traceable (León & Kudelka, 2026). In contrast, citation auditability looks beyond retrieval, assessing how the source was verified, interpreted, and ultimately accepted as evidence.
The distinction also highlights the limitations of AI disclosure. While disclosure can confirm that GenAI took part in research, it does not show how researchers verified its input or assigned responsibility for the scholarly work. Frimpong (2026) identifies this gap—the AI Disclosure Integrity Gap—and recommends shifting from mere disclosure to better traceability, auditability, and accountability in AI-supported research. Extending this idea, citation auditability focuses on a specific aspect: the individual scholarly citation. Simply knowing that GenAI was used to find or synthesise literature does not reveal whether the references were thoroughly checked, why they were considered valid evidence, or how they relate to the claims. AI involvement does not automatically make a citation unreliable; an AI-suggested source can still be retrieved, read, and carefully assessed. The key issue is not just AI participation but whether the entire pathway from source discovery to its evidentiary use can be reconstructed.
Citation auditability tackles this inquiry by redirecting focus from the citation as a bibliographic entity to the evidentiary process that supports it. It does not constitute an additional classification of miscitation, a retrieval procedure, or a substitute for bibliographic and contextual validation. Instead, it pertains to whether the origin, verification, evidentiary application, and human accountability linked to a citation can be reconstructed and evaluated independently.
This also separates citation correctness from citation auditability. Correctness concerns the outcome: whether the reference and its evidentiary use are accurate and appropriate. Auditability concerns the reconstructability of the process underlying that outcome. A citation may be correct but poorly auditable, while an incorrect citation may remain auditable where its evidentiary pathway is sufficiently transparent to reveal how the error occurred. Auditability does not guarantee epistemic correctness; it makes citation decisions more open to scrutiny, challenge, and correction (Hashem et al., 2026; León & Kudelka, 2026).
The literature shows a progression from reference existence to bibliographic accuracy, to evidentiary support, and traceability and accountability (Walters & Wilder, 2023; León & Kudelka, 2026; Jung et al., 2026; van Zoonen et al., 2026). Citation auditability addresses these concerns at the citation level by asking whether the source-to-evidence pathway remains traceable and attributable to responsible judgment.

3. Citation Auditability: A Four-Dimensional Framework

Citation auditability refers to the ability to trace and evaluate the evidence trail that supports how a reference is integrated into a scholarly argument. This involves more than verifying a reference’s existence, its correct metadata, or its general alignment with the related claim. Studies show these are separate questions of integrity: references may be bibliographically authentic but contain metadata errors, and even correctly described publications might not adequately support the claims they underpin (Walters & Wilder, 2023; Jung et al., 2026). Citation auditability emphasises the process linking source discovery, verification, evidentiary use, and human accountability.
This article conceptualises citation auditability through four interconnected dimensions: citation provenance, verification, evidentiary traceability, and human accountability. These dimensions should not be treated as a linear checklist in which completing one aspect automatically leads to the next. Instead, they embody complementary attributes of an auditable citation practice. This distinction is significant because current layered approaches already frame citation verification as progressing from source existence and bibliographic accuracy to contextual support (Jung et al., 2026). Citation auditability, however, addresses a different inquiry: whether the scholarly process by which a reference is identified, examined, linked to a claim, and accepted as evidence remains reconstructible.

3.1. Citation Provenance

The first aspect involves citation provenance: the source and route through which a reference becomes part of the research. Traditionally, references come from bibliographic databases, reference lists, systematic searches, colleagues’ recommendations, citation chaining, or disciplinary knowledge. With GenAI, new avenues emerge: researchers ask an AI to suggest literature, find studies supporting a claim, summarise research, generate references for drafts, or propose citations during revisions. However, these pathways have reliability issues: LLM-generated bibliographies can include fabricated references and inaccurate details about real publications (Walters & Wilder, 2023).
Provenance does not create a hierarchy in which one discovery method is deemed more valid than another. An AI-identified reference might be scrutinised more thoroughly than one taken from a peer-reviewed article’s bibliography. Similarly, a source from a reputable academic database can still be cited without proper engagement with the original material (Liu, 2025; Chen et al., 2025). Earlier research on miscitation, before widespread use of GenAI, shows that misrepresentation of sources is not solely due to the discovery technology (Jergas & Baethge, 2015). Human citation practices have long been vulnerable to distortion, selective interpretation, and the spread of false claims (Jergas & Baethge, 2015). Provenance does not determine whether a citation is valid; it simply provides the initial point from which to understand the subsequent evidentiary relationship (Johns et al., 2023).
This distinction gains heightened significance, as GenAI can obscure the boundary between source discovery and evidence interpretation. When a researcher searches a traditional bibliographic database, the resulting record typically identifies potentially pertinent literature that must then be evaluated. Conversely, a GenAI system may present a reference with a summary of its purported findings and an explanation of how it supports a specific argument. As a result, the researcher encounters an interpretation of the source concurrently with the source itself. This situation engenders the possibility that AI-generated representations of scholarly literature are accepted without adequate engagement with the underlying publications, a concern echoed in broader discussions concerning verification and human responsibility in AI-assisted scholarship (Cleland et al., 2026).
Citation provenance thus extends beyond asking ‘Where did this article originate?’ to inquire ‘How did this source come to be regarded as evidence for this specific claim?’ The latter inquiry becomes particularly significant when Generative AI has mediated not only the discovery process but also the initial interpretation of the publication.
Citation auditability does not obligate authors to disclose the discovery process for every conventional reference within a published manuscript. Such a requirement would be disproportionate and challenging to enforce. Provenance assumes material importance when the manner in which a source was incorporated influences the capacity to reconstruct its subsequent verification and evidentiary utilisation. Its primary aim, therefore, is scholarly reconstructability rather than comprehensive oversight of the research procedure.

3.2. Verification

The second dimension is verification. Verification concerns what the researcher has independently established about a source after identifying it. This includes confirming that the publication exists, that its bibliographic details are correct, that the retrieved document matches the cited reference, and, crucially, that the researcher has examined its substantive content (Walters & Wilder, 2023; O’Driscoll & McGreal, 2026)
The importance of distinguishing these activities is increasingly recognised. Walters and Wilder (2023) demonstrate that fabricated references and substantive bibliographic errors constitute separate problems in AI-generated citations. More recently, Jung et al. (2026) distinguish bibliographic validity from evidentiary validity and propose verification across reference existence, metadata accuracy, and contextual citation support. These findings confirm that locating a publication alone cannot establish appropriate use (Walters & Wilder, 2023; Sarol et al., 2024).
Verification within citation auditability pertains to the substance of the verification conducted, rather than merely the occurrence of some form of checking. A statement such as “all AI-generated references were verified” remains ambiguous unless the specific nature of that verification is clearly articulated. This include checking DOIs, comparing titles and authors with database records, reading abstracts, consulting full texts, or examining the specific findings that underpin the manuscript’s claims. These activities entail different levels of assurance and should not be regarded as equivalent (Jung et al., 2026).
This distinction additionally ensures that verification is not solely a technological solution. Automated systems can help identify nonexistent references, metadata inconsistencies, and other bibliographic issues. Furthermore, increasingly, computational methods aim to assess whether citations genuinely support the claims they accompany. However, scholarly use of evidence often involves interpretive questions about methodological scope, qualifications, conflicting findings, population boundaries, causal inference, and the legitimacy of generalisations. Citation verification remains inherently linked to scholarly judgment, even when technology supports substantial parts of the process.
Verification and evidentiary traceability are easily conflated because both ultimately depend on the researcher having engaged with a source’s substantive content. The distinction lies in what that engagement establishes. Verification is source-directed: it asks whether the source is what it purports to be and whether its content has actually been read and understood, independent of any particular claim it might be cited for. Evidentiary traceability, addressed in the following subsection, is relationship-directed: it asks whether a specific portion of that content can be pointed to as warranting a specific proposition in the citing manuscript. A source can therefore be thoroughly verified without its evidentiary use being traceable, where a researcher has read a publication in full but cannot subsequently identify which finding underwrites a particular sentence in the manuscript. On the other hand, a narrowly verified source still yield a traceable citation if the researcher can point to the specific passage relied upon, even where broader verification of the publication was limited. Distinguishing these questions allows citation auditability to diagnose different failure points rather than treating “the source was checked” as a single undifferentiated condition.

3.3. Evidentiary Traceability

The third dimension, evidentiary traceability, concerns the relationship between source and claim rather than the source itself. Where verification asks what is known about the source, evidentiary traceability asks what the source contributes to this particular claim, and whether that evidentiary basis can be reconstructed. This builds on the distinction between bibliographic and evidentiary validity (Jung et al., 2026) but emphasises the ongoing traceability of that relationship rather than its correctness at a single point of verification.
This is narrower than the reproducibility sought by the broader open-science movement (Munafò et al., 2017): it does not require another researcher to reproduce the entire literature search or analysis, only to answer a specific evidentiary question—what in this source warrants this citation here?
The distinction is clearest where a genuine publication is used to support a proposition that exceeds its evidence: an association observed in one population is cited as a general causal relationship, or attitudes toward AI are used to substantiate a claim about behavioural dependence on it. A cautious, institutionally bounded finding is thereby presented as definitive and universal. Such problems are consistent with the broader phenomenon of miscitation, in which citing authors misrepresent, distort, or inadequately reproduce what a source establishes (Jergas & Baethge, 2015), and with longstanding concerns that bounded findings are carried forward as more definitive than the evidence supports (Ioannidis, 2005).
Evidentiary traceability therefore requires more than thematic relevance: a paper being “about” the same subject does not establish that it supports the specific proposition advanced. What matters is whether the relevant finding, argument, or conclusion within the cited work connects defensibly to the manuscript’s claim, consistent with the growing recognition that citation integrity requires contextual support, not only bibliographic authenticity (Jung et al., 2026).
This dimension is especially important in AI-assisted synthesis, where GenAI can compress multiple publications into fluent statements that blur distinctions among findings, qualifications, and sources. A sentence may read as a settled conclusion across several studies when each source in fact supports only part of it, and the greater the distance between original evidence and generated synthesis, the weaker the traceability back to each source’s actual contribution.
Evidentiary traceability thus safeguards against plausibility substituting for support. It does not require that a citation admit only one interpretation, only that the evidentiary relationship be transparent enough for the researcher’s interpretation to be examined and, where appropriate, contested.

3.4. Human Accountability

The fourth dimension is human accountability. Citation auditability depends on an identifiable human researcher taking responsibility for choosing sources as evidence. This aligns with broader guidance on responsible GenAI use in scholarship, which emphasises that accountability for accuracy and integrity remains with human authors, not AI systems (Cleland et al., 2026). More precisely, van Zoonen et al. (2026) introduce claim accountability, suggesting that governance should ensure an identifiable human can reconstruct and justify claims in the scholarly record. This concept is especially pertinent to citation auditability, as it highlights reconstructability as a key aspect of responsible AI-supported research. Citation auditability narrows this focus to the evidentiary link between a claim and the source that supports it.
At the citation level, accountability means the researcher must explain not only why a claim is reasonable but also why the cited source is appropriate evidence for that claim. Phrases such as “the AI provided the reference,” “the AI summarised the article,” or “the citation appeared in another publication” serve to delineate avenues of discovery; however, they do not absolve the researcher of scholarly responsibility for the resulting citation.
Human accountability does not mean researchers must manually perform every mechanical aspect of citation verification. Instead, they legitimately use databases, reference-management software, automated verification systems, and GenAI to support this process. The key boundary concern is delegating scholarly judgment, not using technological assistance. A researcher delegate the mechanical task of locating a DOI without relinquishing responsibility for determining whether the article substantively supports a particular proposition. This distinction aligns with the overarching principle that human authors remain responsible for AI-assisted scholarly outputs (Cleland et al., 2026; van Zoonen et al., 2026).
This boundary is important because an AI cannot assume responsibility for an erroneous citation. Human accountability closes the auditability chain by assigning responsibility to a specific scholarly actor.

3.5. Interaction Among the Four Dimensions

The four dimensions—provenance, verification, evidentiary traceability, and human accountability—are distinct but interconnected. Provenance shows how a citation entered research; verification checks what the researcher confirmed; traceability links sources to claims; and accountability assigns responsibility. They expand on existing concerns with bibliographic and contextual verification (Jung et al., 2026) and human responsibility for AI-assisted claims (van Zoonen et al., 2026), forming an integrated citation governance framework.
The dimensions can vary independently, justifying their treatment as four distinct dimensions rather than sequential stages of a single process. Provenance and verification are autonomous: a reference generated via an AI recommendation, and thus uncertain in provenance, may undergo comprehensive verification by a researcher who subsequently reads the full text. Equally, a reference obtained through a traditional database search—and therefore secure in provenance—never be substantively examined beyond its mere existence. Verification and evidentiary traceability operate independently as previously detailed: meticulous verification does not inherently guarantee the reconstruction of the specific source–claim link, and a narrowly verified source may still produce a traceable citation. Human accountability remains autonomous of all three aspects: a researcher assumes full responsibility for a citation whose provenance, verification, and evidentiary foundation are thoroughly documented, or equally claims responsibility for one where none of these elements can be reconstructed. Because accepting responsibility does not inherently certify the soundness of the underlying process, this independence gives the four-dimensional framework diagnostic significance: a citation’s overall auditability is not a single attribute but a profile across multiple dimensions, each of which may fail or succeed independently.
Their relationship can be represented conceptually as:
Source discovery → Citation provenance → Verification → Evidentiary relationship → Human acceptance → Scholarly citation
This depiction should not be understood as an inflexible workflow. Researchers often navigate recursively among sources and claims: a source may prompt a revision of a claim, supplementary evidence might modify an interpretation, or verification could reveal that an initially promising reference is inappropriate. Such revisions indicate genuine scholarly engagement rather than mere mechanical citation insertion.
The framework identifies four questions to answer when scrutinising a citation.
  • Provenance: How did this source enter the research process?
  • Verification: What was actually checked against the source?
  • Evidentiary traceability: What does the source contribute to the claim for which it is cited?
  • Human accountability: Who made and can defend the decision to use it as evidence?
A weakness in one dimension need not automatically invalidate a citation. Unknown provenance, for example, does not demonstrate that a citation is incorrect. Similarly, disagreement over interpretation does not establish that verification failed. Citation auditability should not be understood as a scoring system for determining whether individual citations are “good” or “bad.” It is a governance framework for identifying what should remain reconstructable if citation practices are to withstand scholarly scrutiny.

3.6. Citation Correctness and Citation Auditability (Figure 1)

A final distinction must be made. Citation auditability is not synonymous with citation correctness. Correctness pertains to the outcome: whether the reference and its evidentiary use are accurate and appropriate. Auditability pertains to whether the process underlying that outcome can be reconstructed and assessed.
Given that auditability is defined across four dimensions, a citation’s position along the horizontal axis does not represent a sum or average of four distinct scores. Instead, it indicates whether the evidentiary pathway remains coherent from start to finish: a citation is considered auditable to the extent that provenance, verification, evidentiary traceability, and human accountability can each be confirmed without any break in the chain linking source to claim. This approach employs a weakest-link logic rather than an additive model. A citation that possesses clear provenance, comprehensive verification, and an identifiable accountable author may still be situated toward the lower end of the axis if the specific relationship between source and claim, as well as evidentiary traceability, cannot be reconstructed, because this relationship is the ultimate focus of the audit. Conversely, limited provenance documentation does not necessarily place a citation at the low end if the evidentiary relationship and accountability remain unambiguous. Consequently, evidentiary traceability holds particular significance in determining a citation’s placement along the horizontal axis, as it is the dimension closest to the claim-level unit of analysis that citation auditability addresses; the other three dimensions establish the conditions under which that relationship can be trusted, examined, and defended. This shows why the matrix functions primarily as a conceptual framework rather than a scoring system: assigning a citation to a particular quadrant involves judgment about which link in the pathway is weakest, not calculation based on a formula.
As Figure 1 shows, treating citation correctness and citation auditability as analytically distinct produces four possible citation conditions. The matrix is not intended as a scoring or classification instrument, but as a conceptual device for demonstrating why correctness alone provides an incomplete account of citation integrity. A citation may be substantively correct while the process through which it was identified, verified, and accepted remains difficult to reconstruct. Conversely, an incorrect citation may nevertheless be highly auditable where its provenance, verification, and evidentiary reasoning are sufficiently transparent to reveal how the error occurred.
The distinction arises from the broader separation between bibliographic and evidentiary validity (Jung et al., 2026), adding the aspect of reconstructability. The upper-right quadrant shows the strongest alignment: proper source use and a reconstructable evidentiary pathway. The upper-left quadrant explains why correctness alone isn’t enough for auditability; a citation may be appropriate, but uncertainty about its discovery, verification, or interpretation limits process examination. This is especially relevant with AI-driven literature synthesis that produces correct citations without allowing reconstruction of source connections.
The lower quadrants demonstrate a different function of auditability. An incorrect and poorly auditable citation presents the greatest integrity concern because both the evidentiary outcome and the process behind it are deficient. An incorrect but auditable citation, by contrast, remains an error, but its reconstructable pathway makes that error more amenable to identification and correction. Auditability should therefore not be interpreted as evidence of citation quality in itself. Its value lies in making scholarly citation decisions open to scrutiny.
The matrix highlights that auditability does not prevent citation error but makes it more discoverable and correctable. This is crucial in AI-assisted research, where a citation’s plausibility can hide weaknesses in how it links to a scholarly claim.

4. Implications for AI-Assisted Scholarship

Citation auditability affects how researchers, reviewers, and journals approach GenAI-assisted scholarship. Its main point is that a citation’s integrity is not only about bibliographic accuracy or AI disclosure. What’s important is whether the link between source and claim remains clear enough for scrutiny.

4.1. Implications for Researchers

For researchers, citation auditability emphasises that AI aids in source discovery but shouldn’t replace engagement with evidence. A GenAI-suggested reference is a candidate, not verified proof. Confirming a publication’s existence is necessary but not enough; researchers must also verify if it supports their claims. This aligns with guidance demanding accurate citations and supporting sources (ICMJE, 2026).
Suppose a manuscript claims that GenAI use reduces cognitive load among knowledge workers, and a GenAI literature-search tool suggests a survey study as supporting evidence. Provenance is established by noting that the citation originated from an AI-assisted search rather than a manual database query. Verification involves retrieving the full text, confirming the publication and its authors, and reading the study beyond its abstract. Evidentiary traceability requires identifying precisely what the study measured: if it reports perceived time savings rather than cognitive load, the researcher must either qualify the claim accordingly or select a more directly supporting source, rather than treating thematic proximity as sufficient. Human accountability is discharged when the researcher can, if asked, explain why this particular study was judged to support this particular formulation of the claim, rather than deferring to the fact that an AI system proposed it. None of this requires a formal audit log attached to the manuscript. It requires only that the researcher could, if a reviewer or colleague raised the question, reconstruct these four answers from memory or brief notes. Where a researcher cannot do so, for example, when a citation was accepted from an AI-generated synthesis without the underlying study being retrieved and read, the citation should be treated as unverified until that engagement occurs.
Importantly, auditability does not obligate researchers to preserve every prompt, search query, or intermediate AI interaction. Such a requirement would impose an excessive procedural burden. Instead, the relevant standard is more specific: when a citation significantly supports a scholarly claim, the researcher should be able to reconstruct why the source was chosen, what was verified, and how it substantiates the claim. The focus is thus on reconstructing evidence, rather than on comprehensive documentation.

4.2. Implications for Reviewers and Editors

For reviewers and editors, the framework recommends that citation assessment go beyond spotting obviously fabricated or malformed references. A citation may be perfect bibliographically but weak in evidence. Attention should be given when a source is loosely connected to a claim, a strong assertion is supported by a qualified study, or multiple citations are used in an AI-generated synthesis without clear source contributions. This does not mean reviewers must reconstruct every citation’s provenance but suggests citation auditability helps target scrutiny where source–claim links are unclear. Reviewers may ask authors to clarify how a study supports a statement or if the source was consulted. These questions align with the principle that authors are responsible for the accuracy of AI-assisted content (ICMJE, 2026).

4.3. Implications for Journal Governance

At the journal level, citation auditability highlights a distinction between AI disclosure and evidentiary assurance. Disclosure clarifies whether and how AI was involved in preparing the manuscript, but it does not mean AI-generated references were independently verified. Current guidelines rightly emphasise transparency about AI use and maintain that human authors retain responsibility for accuracy (ICMJE, 2026). Citation auditability builds on this by clarifying what this responsibility entails specifically at the citation level.
Journals are not required to mandate detailed citation-audit logs for standard submissions. A proportionate approach would instead explicitly articulate expectations: authors are responsible not only for verifying the existence and accuracy of references but also for substantiating the connection between cited sources and the claims they support. When Generative AI has substantially contributed to literature identification or synthesis, this responsibility becomes particularly significant.
Citation auditability does not necessitate an additional layer of procedural compliance. Its significance lies in elucidating the baseline scholarly responsibilities that should endure despite technological mediation. As artificial intelligence becomes increasingly integrated into literature discovery and manuscript preparation, the goal should not be to eliminate AI from citation practices, but to ensure that the pathway from source to evidence to claim remains reconstructable and accountable to human judgment.

4.4. The Risk of Post-Hoc Rationalisation and the Limits of Self-Reported Auditability

Two objections deserve direct engagement rather than passing acknowledgement. The first is that formalising expectations around auditability could discourage legitimate use of GenAI for literature discovery, for example by encouraging researchers to under-report AI assistance or avoid AI-assisted search entirely rather than document it. This risk is real, but it applies with comparable force to disclosure requirements generally and is not a reason to abandon either disclosure or auditability. It is instead a reason to keep the threshold proportionate, as argued throughout this article, and to frame auditability as evidence of responsible AI use rather than as a penalty for using it.
The second objection is more fundamentally important. Since auditability, as defined herein, relies on a researcher’s ability to reconstruct provenance, verification, evidentiary traceability, and accountability, it remains susceptible to post-hoc rationalisation: a researcher who accepted a citation without genuine engagement with the source may, retrospectively, construct a plausible narrative of how it was verified and why it was deemed relevant. Self-reported auditability is not directly falsifiable in the same manner as, for example, a data-availability statement referencing a publicly accessible repository. In most cases, a reviewer or editor cannot independently verify that the verification process was conducted as described.
This represents a genuine limitation inherent to the framework as a self-assessment instrument, and it warrants explicit acknowledgement rather than unexamined assumption. Citation auditability does not resolve dishonest self-reporting; no purely narrative accountability mechanism can. What it can accomplish is to alter the cost and detectability of rationalisation at the margin through three distinct mechanisms. First, specificity is harder to fabricate convincingly than generality: asking for the specific passage supporting a claim, rather than merely affirming that “the source was checked,” raises the costs of fabricated accounts and makes inconsistencies more likely to be revealed under scrutiny. Second, claims of auditability become verifiable in the most critical cases, as reviewers, editors, or institutional investigations can request that a researcher produce the exact passage or reasoning underpinning a disputed citation; a previously undefined demand now has a concrete target. Third, an accountability structure that mandates a named, identifiable researcher to be responsible for a citation introduces reputational and professional repercussions for rationalisation, which are largely absent when responsibility for an AI-assisted citation is diffuse or unstated.
None of these rules out a confident, false account. Citation auditability raises the evidentiary bar for scholarly practice and enables scrutiny, but it cannot detect dishonesty on its own. This aligns with the article’s view that auditability is a governance principle, not a compliance tool: it makes citations accountable to challenge, not foolproof against false reporting.

5. Discussion and Conclusions

GenAI is changing not only how researchers discover references but also the distance between researchers and the evidence they cite: a researcher now encounters a publication alongside an AI-generated summary and suggested argument before examining the source itself. This creates efficiencies in literature discovery and synthesis but also risks incorporating a plausible representation of evidence into scholarship without sufficient engagement with the evidence itself. Existing concerns with fabricated references, bibliographic errors, and contextual support remain essential (Walters & Wilder, 2023; Jung et al., 2026), but do not fully capture whether the process by which a citation acquired evidentiary status can be reconstructed.
This article develops citation auditability to address that question, shifting attention from the citation as a finished bibliographic object to the process connecting source discovery, verification, evidentiary interpretation, and human acceptance. Its contribution is not another taxonomy of citation errors or verification protocol, but a research-integrity property of the citation process itself: the degree to which the pathway from source to scholarly claim remains reconstructable and open to assessment.
This perspective complements rather than replaces existing approaches: bibliographic verification establishes whether a reference exists and is accurately represented, and contextual verification examines whether the source supports the associated claim (Jung et al., 2026). Claim accountability places responsibility for scholarly claims with identifiable human researchers (van Zoonen et al., 2026). Citation auditability sits between these concerns, asking whether the evidentiary use of an individual citation can be traced from source to claim and attributed to responsible human judgment — a distinction that matters most where GenAI participates not only in finding literature but in summarising, interpreting, and recommending how it should be used.
The Citation Correctness × Citation Auditability Matrix demonstrates why these properties should not be conflated: a correct citation is not necessarily auditable, and an auditable citation is not necessarily correct. This prevents auditability from becoming another proxy for accuracy; its value is procedural. When an evidentiary pathway is reconstructable, questionable citation decisions can be identified, examined, challenged, and corrected; auditability thus contributes to research integrity not by eliminating error, but by increasing scholarship’s capacity to expose and respond to it.
Several boundaries should nevertheless be recognised. Citation auditability does not require exhaustive recording of researchers’ search histories, prompts, or routine AI interactions that would turn a research-integrity principle into an impractical documentation requirement. Nor does it imply that every citation warrants identical scrutiny: the appropriate level of reconstructability likely depends on the significance of the claim, the role of the citation in supporting it, and the extent to which AI mediated the evidentiary process. Future research could examine how citation auditability might be operationalised proportionately across research designs, disciplines, and publishing contexts.
The framework developed here is conceptual and requires empirical examination. Future studies could investigate how researchers currently verify AI-suggested references, whether they can reconstruct source–claim relationships in their own manuscripts, and whether greater auditability improves the detection or correction of citation errors. Experimental work could compare citations from conventional literature searching against AI-assisted workflows, while editorial studies could examine whether auditability-oriented review catches evidentiary problems that bibliographic checks miss.
GenAI does not remove the researcher from the chain of scholarly responsibility; if anything, its capacity to produce convincing references, summaries, and syntheses makes locating that responsibility more consequential. The critical question is no longer only “Is this reference real?” or “Does this source support the claim?” but also “Can we reconstruct how this source became evidence for this claim, and who accepted responsibility for that decision?” Citation auditability offers a framework for answering it. In AI-assisted scholarship, maintaining a reconstructable pathway from source to evidence to claim to responsible researcher becomes an increasingly important condition of research integrity.

References

  1. Aronsky, D. Accuracy of references in five biomedical Informatics journals. Journal of the American Medical Informatics Association 2004, 12(2), 225–228. [Google Scholar] [CrossRef] [PubMed]
  2. Chen, H.; Teplitskiy, M.; Jurgens, D. The Noisy Path from Source to Citation: Measuring How Scholars Engage with Past Research 2025, 31786–31802. [CrossRef]
  3. Clark, J.; Barton, B.; Albarqouni, L.; Byambasuren, O.; Jowsey, T.; Keogh, J.; Liang, T.; Moro, C.; O’Neill, H.; Jones, M. Generative artificial intelligence use in evidence synthesis: A systematic review. Research Synthesis Methods 2025, 16(4), 601–619. [Google Scholar] [CrossRef] [PubMed]
  4. Cleland, J.; Driessen, E.; Masters, K.; Lingard, L.; Maggio, L. A. When and how to disclose AI use in academic publishing: AMEE Guide No.192. Medical Teacher 2026, 48(4), 542–553. [Google Scholar] [CrossRef] [PubMed]
  5. Frimpong, V. AI Disclosure without Accountability: Paper Compliance and the Governance Limits of Transparency in Scientific Research. International Journal of Social Science Studies 2026, 14(3), 15. [Google Scholar] [CrossRef]
  6. Gasparyan, A. Y.; Yessirkepov, M.; Voronov, A. A.; Gerasimov, A. N.; Kostyukova, E. I.; Kitas, G. D. Preserving the integrity of citations and references by all stakeholders of science communication. Journal of Korean Medical Science 2015, 30(11), 1545. [Google Scholar] [CrossRef] [PubMed]
  7. Greenberg, S. A. How citation distortions create unfounded authority: analysis of a citation network. BMJ 2009, 339, b2680. [Google Scholar] [CrossRef] [PubMed]
  8. Hashem, R.; Fidalgo, P.; Alramamneh, Y.; Elkaleh, E.; Zein, F. E.; Hussien, A. From plausibility to auditability: generative systems, synthetic scholarship, and the epistemic risks of AI-Mediated research. Education Sciences 2026, 16(7), 1122. [Google Scholar] [CrossRef]
  9. International Committee of Medical Journal Editors. Recommendations for the conduct, reporting, editing, and publication of scholarly work in medical journals. ICMJE. 2026. Available online: https://www.icmje.org/recommendations/.
  10. Ioannidis, J. P. A. Why most published research findings are false. PLoS Medicine 2005, 2(8), e124. [Google Scholar] [CrossRef] [PubMed]
  11. Jergas, H.; Baethge, C. Quotation accuracy in medical journal articles —a systematic review and meta-analysis. PeerJ 2015, 3, e1364. [Google Scholar] [CrossRef] [PubMed]
  12. Johns, M.; Meurers, T.; Wirth, F. N.; Haber, A. C.; Müller, A.; Halilovic, M.; Balzer, F.; Prasser, F. Data Provenance in Biomedical Research: Scoping Review. Journal of Medical Internet Research 2023, 25(1), e42289. [Google Scholar] [CrossRef] [PubMed]
  13. Jung, E.; Yoo, J.; Kim, S. G.; Kim, Y. S. Citation integrity in the AI era: a focused evidence map and layered verification framework for gastroenterology and hepatology publishing. Gastroenterology Report 2026, 14, goag098. [Google Scholar] [CrossRef] [PubMed]
  14. Kitchenham, B.; Pizard, S.; Shepperd, M.; Souza Santos, R.; Madeyski, L.; Budgen, D. Preliminary Guidelines for Using and Evaluating GenAI Tools to Support Systematic Literature Reviews. arXiv; Cornell University, 2026. Available online: https://arxiv.org/abs/2607.24991.
  15. León, C.; Kudelka, M. Auditing GenAI Literature Search Workflows: A Replicable Protocol for Traceable, Accountable Retrieval in Student-Facing Inquiry. AI in Education 2026, 2(2), 8. [Google Scholar] [CrossRef]
  16. Liu, H. Fabricated citations in the age of AI: A wake-up call for editors, reviewers, and authors. Journal of Dental Sciences 2025, 21(1), 679–680. [Google Scholar] [CrossRef] [PubMed]
  17. Mogull, S. A. Accuracy of cited “facts” in medical research articles: A review of study methodology and recalculation of quotation error rate. PLoS ONE 2017, 12(9), e0184727. [Google Scholar] [CrossRef] [PubMed]
  18. Munafò, M. R.; Nosek, B. A.; Bishop, D. V. M.; Button, K. S.; Chambers, C. D.; Percie du Sert, N.; Simonsohn, U.; Wagenmakers, E.-J.; Ware, J. J.; Ioannidis, J. P. A. A manifesto for reproducible science. Nature Human Behaviour 2017, 1(1), 0021. [Google Scholar] [CrossRef] [PubMed]
  19. O’Driscoll, J.; McGreal, R. Determining Accountability in AI-Assisted Scholarship: A Report and Recommendations. The International Review of Research in Open and Distributed Learning 2026, 27(3), 1–10. [Google Scholar] [CrossRef]
  20. Rajkovic, A.; Wang, J. Responsible use of artificial intelligence in manuscript preparation guidance from the editors of Biology of Reproduction. In Biology of Reproduction; 2026. [Google Scholar] [CrossRef] [PubMed]
  21. Sarol, M. J.; Ming, S.; Radhakrishna, S.; Schneider, J.; Kilicoglu, H. Assessing citation integrity in biomedical publications: corpus annotation and NLP models. Bioinformatics 2024, 40(7). [Google Scholar] [CrossRef] [PubMed]
  22. Van Zoonen, W.; Morgan-Thomas, A.; Tursunbayeva, A. Beyond AI disclosure: Claim accountability and responsible research in scholarly publishing. European Management Journal 2026, 44(4), 550–556. [Google Scholar] [CrossRef]
  23. Walters, W. H.; Wilder, E. I. Fabrication and errors in the bibliographic citations generated by ChatGPT. Scientific Reports 2023, 13, 14045. [Google Scholar] [CrossRef] [PubMed]
  24. Zhao, Z.; Wang, Y.; Stuart, T.; Mathijs, D. V.; Ginsparg, P.; Yin, Y. LLM hallucinations in the wild: Large-scale evidence from non-existent citations. In arXiv; Cornell University, 2026. [Google Scholar] [CrossRef]
Figure 1. Citation Correctness × Citation Auditability Matrix. Note. The vertical axis represents citation correctness, ranging from incorrect (inaccurate or inappropriate) to correct (accurate and appropriate). The horizontal axis represents citation auditability, ranging from low (not reconstructable) to high (reconstructable), based on whether the evidentiary pathway can be reconstructed and assessed. Position along this axis reflects the weakest link across the four dimensions, not an aggregate score. The four quadrants illustrate that correctness and auditability are analytically distinct. A citation may therefore be correct but not auditable, or incorrect but auditable. Auditability does not guarantee correctness; rather, it enables scrutiny and helps identify and correct citation errors. Source: The Author, 2026.
Figure 1. Citation Correctness × Citation Auditability Matrix. Note. The vertical axis represents citation correctness, ranging from incorrect (inaccurate or inappropriate) to correct (accurate and appropriate). The horizontal axis represents citation auditability, ranging from low (not reconstructable) to high (reconstructable), based on whether the evidentiary pathway can be reconstructed and assessed. Position along this axis reflects the weakest link across the four dimensions, not an aggregate score. The four quadrants illustrate that correctness and auditability are analytically distinct. A citation may therefore be correct but not auditable, or incorrect but auditable. Auditability does not guarantee correctness; rather, it enables scrutiny and helps identify and correct citation errors. Source: The Author, 2026.
Preprints 234233 g001
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.