Preprint
Article

This version is not peer-reviewed.

Evidence-Bound Factual Repair in Retrieval-Augmented LLM Answers: Separating Semantic Localization from Deterministic Realization

Submitted:

04 September 2026

Posted:

07 September 2026

You are already at the latest version

Abstract
Retrieval-augmented generation (RAG) reduces reliance on parametric knowledge but does not prevent unsupported or contradictory statements. We evaluate a repair pipeline that separates semantic localization from factual realization: a language model selects evidence sentences, while a deterministic program substitutes them verbatim without authority to author replacement text. One frozen pipeline is evaluated against RAGTruth hallucination labels and ExpertQA evidence-grounding labels. Among 889 held-out RAGTruth responses, repair achieved a net reduction of 12 of 32 human-labelled hallucinations (37.50%; exact paired p = 0.0042; 95% CI [18.3%, 56.7%]). Among 731 ExpertQA responses, 26 of 110 reviewed grounding failures were repaired (23.64%; exact paired p = 4.17 × 10⁻⁷). Rationale-blinded second-author endpoint verification found 100% agreement on RAGTruth and 98.84% on ExpertQA (Cohen's κ = 1.000 and 0.977); two ExpertQA disagreements were resolved by consensus. Major completeness loss, defined as a change to the direct answer or protected meaning, occurred in 2.3–2.8% of responses across the two datasets. A zero-additional-cost provenance tool identifies evidence-copied, flagged, and unchecked text.
Keywords: 
;  ;  ;  ;  ;  ;  ;  ;  

1. Introduction

Retrieval-augmented generation combines a parametric language model with retrieved external evidence and has become a standard approach for knowledge-intensive generation [1]. Retrieval improves access to current or task-specific information, but it does not guarantee that the generated answer is entailed by, or even consistent with, the retrieved evidence. RAGTruth was introduced specifically to quantify this residual failure mode using manually annotated hallucination spans in naturally generated RAG responses [2].
The dominant mitigation strategies span retrieval improvements, prompt constraints, decoding-time interventions, self-reflection, verification, and post-generation correction. Recent reviews emphasize that hallucinations can enter both the retrieval and generation stages of RAG systems and therefore require controls across the pipeline [5]. Other work changes the decoding distribution to favor factual continuations [6], while layered deployment frameworks combine retrieval, prompting, monitoring, and escalation [7]. These approaches improve factuality but generally retain the language model as the author of the final factual wording, which reopens the door to exactly the failure mode being corrected.
This paper is scoped around one question: can a RAG hallucination-repair mechanism reduce unsupported or contradicted content without authoring any new factual content of its own? The model is given semantic control over localization and evidence binding, while its authority to write the accepted factual replacement is removed: the model points to evidence, and a deterministic program performs the substitution. This does not eliminate every error channel: the model may bind a claim to the wrong evidence, a source sentence may carry unresolved context, unsupported claims can remain open, and exact source replacement can reduce completeness or fluency. These residual risks are exposed directly to a reviewer, response by response, through a free provenance signal described in Section 4.
Two datasets are evaluated under a shared evidence-support construct. RAGTruth supplies span-level, human-annotated hallucination labels, whereas ExpertQA supplies per-claim evidence-grounding labels that indicate whether a claim is supported by the retrieved evidence, independently of whether it happens to be true. The study treats ExpertQA grounding failure as a hallucination-equivalent operational endpoint while preserving its native label structure and reporting it separately from RAGTruth’s hallucination endpoint.
Every headline quantity reported here reflects completed human review and adjudication for the relevant evaluation panel. The RAGTruth hallucination endpoint and ExpertQA grounding endpoint were additionally subjected to rationale-blinded second-author endpoint verification on their prespecified review panels (Section 5 and Section 6).
The study makes five contributions. First, it evaluates deterministic factual realization on two independently annotated datasets under one frozen pipeline, with both primary human-labelled endpoints independently checked through rationale-blinded second-author endpoint verification. Second, it isolates the mechanism from an unconstrained single-shot regeneration alternative on a matched sample, showing that removing the model’s authority to author the factual replacement measurably suppresses new-hallucination introduction at an equivalent removal rate. Third, it reports the coverage and rendering diagnostics that determine how much of the located hallucination surface is actually repairable, and at what rendering-quality and fluency cost. Fourth, it measures whole-response completeness cost on all officially clean held-out responses in both datasets, counting only a change to the direct answer as a cost. Fifth, it introduces a deterministic, zero-additional-cost evidence-provenance coloring tool for RAGTruth that gives a reviewer an incident-level, per-response provenance view without a second model call. The remainder of the paper keeps hallucination reduction, repair coverage, rendering quality, grammar, completeness, and provenance signalling as separate outcomes rather than collapsing them into one score.

2. Literature Review

2.1. Retrieval-Augmented Generation and Residual Hallucination

Lewis et al. formalized RAG as a combination of parametric and non-parametric memory for knowledge-intensive NLP [1]. Later work such as Self-RAG made retrieval adaptive and added model-generated critique signals, demonstrating that retrieval quality and self-reflection can improve factuality and citation behavior [4]. However, retrieval itself does not ensure that every generated proposition remains faithful to the retrieved context. RAGTruth addresses this evaluation gap with manual case- and word-level annotations of hallucinations in RAG outputs across tasks and model families [2]. ExpertQA provides an independently constructed grounding-evaluation surface: expert-curated, long-form questions with retrieve-and-read GPT-4 responses and per-claim human support and attribution labels [13], extending the long-form question-answering task family introduced by ASQA [12]. Its grounding endpoint is reported separately from RAGTruth’s hallucination endpoint.
Both benchmarks sit inside a broader, still-active problem rather than a solved one. General surveys of hallucination in natural language generation [14] and, more recently, specifically in large language models [15] catalogue causes, detection methods, and mitigation strategies across tasks and model families; Gao et al. separately survey the retrieval-augmented generation literature this paper’s intervention is downstream of, spanning naive, advanced, and modular RAG architectures [21]. FaithBench shows the problem persists even in newer, stronger generators by benchmarking hallucination in modern-LLM summarization specifically [9], reinforcing that hallucination reduction remains open rather than resolved by scaling generator quality alone.
The MDPI review by Zhang and Zhang organizes RAG hallucination causes across retrieval failure and generation deficiency and surveys mitigation methods spanning retrieval, prompt engineering, detection, and correction [5]. This framing is directly relevant here because the intervention studied is intentionally located after retrieval: the evidence is treated as fixed and its retrieval accuracy is not evaluated; the question is only whether hallucination can be reduced, without adding new unsupported content, once evidence selection is already done.

2.2. Factuality Evaluation and Claim-Level Decomposition

Fine-grained evaluation is important because a single answer may mix supported and unsupported propositions. FActScore operationalizes this problem by decomposing long-form generations into atomic facts and measuring the proportion supported by a knowledge source [3]. RAGTruth similarly provides span-level human annotations [2]. These ideas motivate this paper’s claim-level cascade analysis: repair coverage is reported per factual claim, while response-level hallucination-reduction results remain separate so that unresolved claims are not hidden by successfully repaired neighboring claims.
Claim-level verification against evidence has an older lineage in fact-checking: FEVER pairs claims with cited evidence and a supported/refuted/not-enough-info label using an extraction-and-verification pipeline structurally similar to the locate-then-check designs used throughout this paper [25]. Maynez et al. show that reference-overlap metrics correlate poorly with human faithfulness judgments in abstractive summarization, motivating this paper’s use of claim-level, evidence-grounded checking rather than a surface-similarity score [23]; RAGAS instead automates several such checks, including a faithfulness score, without a human reference, at the cost of relying on an LLM judge for every constituent score, a dependency this paper’s central provenance contribution (Section 4.2.4) is designed to reduce rather than replace outright [27].
RefChecker represents response content as claim triplets and evaluates each claim against a reference using separate extraction and checking stages [10]. RAGChecker extends fine-grained claim-level entailment into a diagnostic framework for retrieval and generation components [11]. The RAGTruth hallucination evaluation and the unconstrained-regeneration baseline evaluation in this paper both adopt the RefChecker-style extractor/checker separation and the RAGChecker principle of claim-level evidence diagnosis, but neither executes nor claims equivalence to either official software package; the ExpertQA grounding evaluation uses a separate, two-stage classification design suited to ExpertQA’s own per-claim support taxonomy rather than the RefChecker-style triplet extraction (Section 5).

2.3. Decoding-Time and Layered Mitigation

Lv et al. propose contrastive decoding with factual and hallucination prompts, showing that inference-time decoding can shift factuality without additional training [6]. The method illustrates a broad class of approaches in which the model still generates the corrected content, but the probability distribution is altered to favor factual continuations. In contrast, the primary system in this study does not ask the model to regenerate accepted factual payloads after evidence binding; a deterministic program copies source text directly; Section 6’s unconstrained-regeneration baseline is a direct empirical test of what is given up by relaxing that restriction.
Hiriyanna and Zhao present a multi-layered hallucination-mitigation framework for high-stakes applications that combines prompting, RAG, fine-tuning, monitoring, and escalation [7]. Their emphasis on operational trade-offs motivates this paper’s explicit accounting of latency, cost, unresolved claims, and reviewer-facing signaling. This architecture differs in scope: it is not a full deployment stack, but a controlled factual-repair mechanism designed to isolate the effect of deterministic factual realization.
A separate line of post-hoc correction work edits or revises model output after generation rather than altering decoding or training: RARR retrieves evidence for an existing generation and edits only the spans that evidence contradicts, preserving the rest [18]; Peng et al. use a similar retrieve-critique-revise loop to iteratively check and correct a response against external knowledge and automated feedback [28]. Both differ from the mechanism evaluated here in one specific respect this paper’s own ablation tests directly (Section 6.6): the correction is still authored by a model, whereas the primary mechanism here restricts the replacement text to a verbatim copy of the cited evidence sentence, with no model authorship of the corrected span at all. This restriction has an older precedent in text generation: pointer-generator networks let a decoder choose, token by token, between generating from a fixed vocabulary and copying a token directly from the source, improving factual reproduction of source content in abstractive summarization [26]; the Renderer in this paper (Section 4.2.3) can be read as pushing that copy mechanism to its deterministic extreme, copying whole evidence sentences verbatim or leaving the span unchanged, with no generative option in between. A different family of correction methods edits a model’s parameters directly rather than its output text: ROME locates and edits individual factual associations inside a GPT-style model’s feed-forward weights [24]. This paper’s intervention operates entirely at the text level, after generation, and makes no claim about what would happen if the same problem were instead addressed by editing model weights; the two approaches are complementary rather than competing.

2.4. Mechanism-Native Provenance Signalling Without an Additional Judge

Prior work on attributed generation has improved verifiability by attaching document citations, supporting passages, or fine-grained quotations to generated answers [17,30,31]. Related RAG-evaluation frameworks assess citation faithfulness or claim support using an additional model-based evaluator [11,27], while research on LLM-based evaluators shows that explicit calibration can improve alignment with expert judgments [32].
This study takes a complementary approach. Its provenance coloring is neither a newly generated citation nor an additional judge score. Instead, it projects records already created during evidence-bound repair back onto the final response, distinguishing evidence-copied spans, copied spans associated with preservation flags, and text that remained outside the checked repair path. Because this projection reuses the frozen repair trace, it requires no additional model inference. The contribution is therefore mechanism-native, incident-level provenance: a reviewer can inspect how each part of an individual response was produced rather than receiving only a response-level confidence score. Reliability of the RAGTruth hallucination and ExpertQA grounding endpoints is addressed through rationale-blinded second-author endpoint verification reported in Section 5 and Section 6.

2.5. Selective Prediction, Abstention, and Citation Attribution

Two adjacent literatures inform the coloring tool’s design without being directly used by it. Selective prediction asks a system to abstain rather than answer when it is likely wrong: Kamath et al. train a calibrator to decide when a question-answering model should abstain under domain shift [19], and Xin et al. survey and unify selective-prediction and error-regularization approaches for natural language processing more broadly [20]. Kadavath et al. show that a large language model’s own stated confidence, when elicited appropriately, can be informative about answer correctness, suggesting a model-internal alternative to an external judge [22]. The coloring tool’s confidence tier (Section 4.2.4) differs from all three in its source signal: it is not a learned calibrator, a survey-recommended abstention rule, or a model’s self-reported confidence, but a deterministic function of which claims the locator and Coverage Gate left open or flagged, computed with no additional model call.
Separately, attribution asks whether a specific piece of generated text is actually supported by a specific cited source: Rashkin et al. formalize Attributable to Identified Sources (AIS) as a human-evaluable standard for this question [16], and Bohnet et al. extend attribution into an end-to-end evaluated and modeled task for long-form question-answering systems [17]. The evidence-provenance coloring in this paper answers a narrower version of the same question without a judge: because every clean-rendered span is constructed by copying an identified evidence sentence verbatim, attribution for that span is true by construction rather than estimated after the fact, at the cost of only ever being available for the subset of text the pipeline actually rewrote.

2.6. Positioning and Research Gap

Prior work shows that retrieval, critique, verification, decoding constraints, selective abstention, attribution modeling, and citation generation can each improve a dimension of trustworthy generation (Section 2.1, Section 2.2, Section 2.3, Section 2.4 and Section 2.5), yet the accepted correction commonly remains model-authored. Exact copying, extractive answer construction, and evidence citation are not themselves new, and this paper does not claim otherwise.
The research gap addressed here has two linked and testable components. First, does removing the model’s authority to author an accepted factual replacement, while retaining model-based claim localization and evidence selection, reduce hallucination and newly unsupported content relative to unconstrained regeneration using the same evidence? This comparison is tested directly on RAGTruth. Second, can the repair mechanism’s existing execution trace be projected onto the final response as span-level provenance without invoking another verifier? The resulting coloring tool exposes evidence-copied, flagged, and unchecked text at zero additional inference cost. The repair mechanism is evaluated against RAGTruth’s hallucination endpoint and ExpertQA’s evidence-grounding endpoint, with rationale-blinded second-author endpoint verification of the corresponding human-evaluation decisions.

3. Datasets

This section describes the two datasets and the specific sub-populations evaluated; the repair architecture, its terminology, and the evaluation protocol applied to these datasets are described separately in Section 4 and Section 5. Retrieval quality and retrieval accuracy are outside the scope of this paper throughout: the evidence associated with each response is accepted as given, and no claim is made about how well it was retrieved.

3.1. RAGTruth Corpus and the GPT-4-0613 QA Arm

RAGTruth is a human-annotated corpus for studying hallucinations in RAG responses [2]. The local released response file used in this study contains 17,790 response records across the benchmark. This paper restricts the RAGTruth arm to the QA task and the GPT-4-0613 generator, which integrity checks in the frozen manifest identify as exactly 989 QA responses, of which 42 (4.25%) contain at least one released hallucination annotation and 947 (95.75%) are released as clean (the 42 hallucinated responses contain 51 annotated spans). This subset is deliberately chosen because it is comparatively clean at the source: with roughly 1 in 24 responses flagged, it lets the mechanism be tested on the harder, more realistic case of catching a small residual hallucination rate inside an already largely accurate generator, rather than on a noisier task/generator combination where hallucination is common enough that a coarser signal could look effective by coincidence. Repair’s locator also uses the same GPT-4-0613 snapshot as this generator (Section 4), so the QA/GPT-4-0613 restriction additionally keeps the generation and localization model family aligned.
The held-out set retains the complete remainder rather than quality-filtering it: 861 responses are marked good, 27 incorrect_refusal, and 1 truncated. It contains 755 train-split and 134 test-split records from the released benchmark metadata. Human labels are withheld from every locator call and are used only for sampling the development set and post-hoc evaluation.

3.2. Development/Held-Out Separation and Freeze Policy

A 100-response Wave-2 development set was fixed before the primary held-out experiment and excluded from it, identified as 10 officially hallucinated and 90 officially clean quality=“good” responses, leaving exactly 889 responses (32 human-positive, 857 clean) as the held-out cohort reported from Section 5 onward. The semantic-localization prompt, model snapshot, complete-sentence evidence representation, deterministic pointer validation, cascade-gate definitions, and final held-out rendering rules were all frozen before this 889-response run; development outputs were used only for method design and ablation reasoning. Section 7 gives the full independent verification of this split, including the manifest arithmetic and response-ID cross-check, and its role in the paper’s out-of-sample claims.

3.3. ExpertQA Corpus and the GPT-4 Retrieve-and-Read Arm

ExpertQA is used here as the second frozen dataset under the identical pipeline described in Section 4, evaluated in full rather than held in reserve. The held-out cohort comprises 731 GPT-4 retrieve-and-read responses, verified identically at every pipeline stage. Of these, 129/731 (17.65%) are hallucination-proxy positive at the source label level (at least one claim independently labelled Likely or Definitely incorrect by the original ExpertQA annotation) and 602/731 (82.35%) are hallucination-proxy clean; 622/731 (85.09%) carry at least one grounding-failure-proxy claim (Incomplete, Partial, or Missing support), covering 1,746 claims, and 211 total human factual-error claims are recorded across the cohort. As with RAGTruth, human labels are withheld from every locator call.
ExpertQA’s two axes. Unlike RAGTruth’s single hallucination/no-hallucination label, ExpertQA’s original per-claim annotation separates two distinct questions about a claim: factual correctness (is the claim actually true, checked against the annotator’s own world knowledge) and grounding (is the claim supported by the specific evidence passages retrieved for that response, independent of whether it happens to be true). A claim can be true but ungrounded (correct by luck or by outside knowledge, but not backed by the retrieved evidence), or grounded but false (the evidence itself is wrong or was misread upstream). This paper evaluates only the grounding axis and treats it as the ExpertQA counterpart to RAGTruth’s hallucination label, because the repair mechanism studied here can only ever make an answer match its cited evidence; it has no access to outside truth and cannot correct a claim that is wrong even though the evidence agrees with it. Factual correctness independent of grounding is a different, broader claim this architecture does not attempt and is explicitly out of scope. Section 6 reports the grounding axis, which the evidence-bound replacement directly targets.

3.4. Claim-Extraction Exploration Preceding the Final Architecture

Before the final evidence-bound architecture, an earlier, structurally different claim-extraction approach was tried; its corpus statistics, selector design, and results are reported as part of the design narrative in Section 4.1.

4. Materials and Methods

At a high level, the pipeline reads an original answer together with its evidence passages and passes it through four stages, each of which either checks or rewrites the answer without doing both: a locator step identifies which specific evidence sentence, if any, supports each factual claim; a coverage check determines which of those identifications are usable; a renderer performs the actual text substitution, copying evidence text verbatim rather than letting the model write new wording; and an optional fluency pass (evaluated only as a comparison, not used by default) lets a model re-polish the rendered text for grammar. For the RAGTruth analysis, a fifth, deterministic step turns the renderer’s own internal records into a color-coded, response-level view for a human reviewer. The four crisp terms used for these stages throughout the rest of this paper are: the original, pre-repair answer is B0; the locator-and-substitution step is Repair; the check on which of Repair’s claim mappings are safe to render is the Coverage Gate; and the deterministic text substitution itself, together with its internal completeness check, is the Renderer. The optional grammar re-polish comparison is the Fluency Ablation. Appendix A maps these five names to their internal program identifiers for code-level reproducibility; the main text uses only the names above.

4.1. Different Design Experiments: How the Final Architecture Was Chosen

Before settling on deterministic, evidence-bound realization, two different design experiments were run to find out whether a cheaper, judge-free way to flag risky output could work, and one further alternative to constrained rendering was tested directly rather than only argued against. Together they tell one coherent story about how the final architecture was chosen, not two unrelated negative results.
The first experiment asked whether a model even needs to check every claim in an answer, or whether risk can instead be predicted from surface form alone. An earlier claim-verification pipeline extracted atomic claim candidates from each answer using a frozen Stanza dependency parser [8], then applied a deterministic selector that flagged only claims containing a fragile value, a negation, a conditional, or a requirement, on the theory that this surface pattern would concentrate where real hallucinations occur; a companion program converted the selector’s routing into per-claim verification jobs and planning metrics. Tested directly against real hallucination-span ground truth on 600 RAGTruth responses, the result was the opposite of the theory: selected-ready claims overlapped a real human hallucination span at only 4.80% (29/604), while the claims the selector left out (nonselected-ready) overlapped at 6.33% (155/2,448) — a selected/nonselected relative risk of 0.758. Selective claim verification built this way does not cover the full hallucination signal; if anything it mildly anti-correlates with it, since the claims it skips are hallucinated more often than the claims it keeps. Predicting risk from surface form, in short, does not work, so the final architecture does not use claim-level selection to decide what gets checked at all: every claim the locator can identify is checked.
The second question was not about what to check, but about who is allowed to write the replacement text once a claim is checked. A natural alternative to copying evidence verbatim is to let the model regenerate the whole answer freely, evidence-consistent in instruction but not mechanically constrained to it; rather than argue against this on principle, Section 6 tests it directly, on a matched sample, against the constrained design (Section 6.6). Free regeneration removes the originally flagged hallucination at a comparable or even nominally higher rate, but it introduces new, evidence-unsupported content at roughly 4.5 times the rate of the constrained design at that equivalent removal rate. Letting the model author its own replacement text, in other words, trades a predictable, auditable failure mode (leaving an unsupported claim unchanged) for a several-fold higher rate of an unpredictable one (inventing a new, unsupported claim).
Put together, these two findings explain the final architecture rather than merely motivating it. Neither a cheap heuristic for deciding what to check, nor giving the model freedom in how to write the replacement, survives direct testing: the heuristic does not concentrate risk where the risk actually is, and free regeneration trades away its predictability advantage for a several-fold higher rate of newly invented content. The only remaining option that avoids both failure modes is the one this paper evaluates as primary: check every claim the locator can identify, and realize every accepted replacement by copying evidence text verbatim rather than letting the model write it, so that what changed is always exactly traceable to a specific evidence sentence, with a transparency signal built on top of that guarantee (Section 4.2.4). RAGTruth’s positive cohort is small, which limits the precision of its point estimate in isolation. ExpertQA’s independently annotated 110-response grounding-failure sample evaluates the deterministic evidence-bound repair mechanism under a different domain and annotation protocol.

4.2. Current Pipeline

The four subsections below name and specify each stage of the frozen pipeline outlined at the start of Section 4 — Repair, Coverage Gate, Renderer, and the deterministic coloring tool — exactly as implemented and run to produce every result reported from Section 5 onward, not as further design alternatives.

4.2.1. Repair: Semantic Localization with Deterministic Factual Realization

Repair operates on an existing answer (B0) and its retrieved source passages; it does not regenerate the answer as a whole, and the same program (same prompt contract, same validation code) is applied unchanged to both RAGTruth and ExpertQA. Before each model call, dataset navigation labels such as “passage 1:“ are masked with equal-length spaces so source offsets remain stable, and the source is then deterministically segmented into complete candidate sentences; incomplete or structurally ambiguous units are rejected and are never exposed as selectable evidence. Each accepted complete evidence sentence receives a stable identifier (E1, E2, …) while its exact source substring and character offsets are retained.
The locator model receives the question, B0, and the numbered evidence sentences. It identifies every externally verifiable factual proposition in the answer and returns an exact answer span plus either one or two evidence-sentence identifiers, or a signal that no exact replacement exists. The model returns identifiers only; it is not permitted to write, paraphrase, normalize, shorten, combine, or otherwise author replacement factual text. A deterministic validation step then checks the answer pointer, the selected sentence identifiers, a sentence-count limit, overlap safety, and a hard replacement-length guard; for a valid mapping, the selected complete source sentence(s) are copied into only the factual claim span. Unsupported, non-extractable, invalid, or unresolved claims remain unchanged; this fail-closed behavior is what the coloring tool (Section 4.2.4) turns into a reviewer-facing signal rather than leaving it implicit.
For RAGTruth, the locator uses the same model snapshot as the RAGTruth GPT-4 QA generator (GPT-4-0613), temperature 0, and a maximum completion budget of 2,200 tokens, so that generation and localization share a model family and no model-family confound is introduced. For ExpertQA, the identical prompt contract, validation code, and replacement mechanism are applied; only population loading and evidence-passage parsing differ between the two dataset adapters. Human hallucination/factual-error labels are never included in either locator prompt.

4.2.2. Coverage Gate: Cascade Coverage Analysis

Repair has several dependent stages: claim localization, answer-pointer validation, evidence-ID selection, deterministic validation, and replacement. An early failure can propagate because fail-closed behavior leaves the affected text unchanged. The Coverage Gate makes these cascade states observable instead of treating all unrepaired text as one undifferentiated outcome. It performs no model calls, does not alter answer text, and does not reject complete responses.
A claim mapping is classified resolved only when its answer pointer is valid, its selected evidence IDs are valid, no overlap or guard failure is present, and the replacement was actually applied; all other claim mappings are unresolved. Only resolved mappings are eligible for rendering; unresolved mappings retain their Repair-stage text. Every response falls into exactly one of five mutually exclusive categories, ranging from no usable evidence at all to full resolution of every located claim (Section 6).

4.2.3. Renderer: Deterministic Rendering and completeness audit

The Renderer replays only the mappings the Coverage Gate classified as resolved. It makes zero model calls, does not rerun claim localization, and does not change evidence selection; unresolved mappings are not rendered and retain their Repair-stage text. Rendering behaviors include: same-source-identity duplicate suppression; conservative structural-marker stripping that protects factual quantities such as 1 mg, 1 hour, 1%, and $1; seam-punctuation and boundary-word cleanup; exact-source-backed retention of an original span only when Python proves it occurs verbatim in the source; and structural line breaks for multi-sentence replacements.

4.2.4. Deterministic Evidence-Provenance Coloring

The fourth pipeline stage adds no new model call and consumes only the RAGTruth Repair and Renderer outputs already produced for each response. It is this paper’s central provenance contribution: an incident-level, per-response signal a reviewer can read directly, at zero marginal cost, for a specific response they are examining, not a dataset-level statistic. Section 6.7 reports its confidence-tier distribution over the complete 889-response RAGTruth held-out cohort.
Coloring. The Renderer records, in document order, every span it actually substituted with verbatim evidence text, each already carrying its own completeness-check result. A response’s final answer is tiled by sequentially locating each substituted span and marking it clean-rendered (verbatim evidence, no flag) or flagged-rendered (verbatim evidence, flagged); everything else is untouched original prose: text the pipeline never checked as a claim, neither confirmed nor contradicted by this signal. For a reviewer, this changes the task from re-reading an entire answer to scanning three colors: green text does not need to be checked again because it is a verified copy of cited evidence, amber text is exactly where a second look is warranted, and everything else was never touched by the repair mechanism and carries no claim about its accuracy either way. This turns a whole-response review into a targeted, incident-level task and is the practical reviewer-experience benefit this tool is designed around.
Worked example. Figure 1 reproduces one real RAGTruth response’s coloring-tool output unedited: the Renderer performed two substitutions, one copied verbatim from evidence with no flag (clean-rendered) and one also copied verbatim from evidence but flagged by the completeness-risk audit (flagged-rendered), alongside untouched original prose the pipeline never checked. A reviewer opening this one response can go straight to the flagged span rather than re-reading the entire answer. Section 6.7 evaluates this mechanism at full scale on RAGTruth.
Confidence tier. A four-way tier is computed from simple counts already available in the Renderer’s output: how many claims the locator found, how many were actually replaced, whether any claim was left open, and whether any substitution was flagged. A response with nothing to check gets its own category rather than being folded into a “resolved” count; a response with an unresolved claim is marked low-confidence; a response fully resolved with no flags is marked high-confidence; everything in between is medium. A companion coverage ratio (the fraction of the final answer’s characters that are verbatim-evidence-rendered) is reported alongside every tier so a high-confidence tier over a small fraction of the answer’s text cannot be mistaken for full-answer verification.
This tier is explicitly not a calibrated probability of correctness and is bounded by the locator’s own recall: a hallucinated span the locator never selected as a claim can never appear as an open claim, so the high-confidence label means “nothing the pipeline checked was left open or flagged,” not “this answer has no remaining error.” Green, clean-rendered text is evidence-bound, not independently verified against the real world. Section 6 reports the tier’s actual output distribution over the full RAGTruth held-out cohort together with its cost/latency profile.

5. Evaluation Protocol

This section describes how each result in Section 6 is measured. In outline: RAGTruth’s hallucination axis is measured by comparing B0 against the repaired answer with a human reference plus an LLM-assisted claim checker (Section 5.2); ExpertQA’s grounding axis is measured with a separate two-stage machine classification plus author review, since ExpertQA’s own claim-support taxonomy does not match RAGTruth’s binary label (Section 5.3); rendering quality, grammar, and completeness are each measured with their own independent protocol so that none of these outcomes is folded into another (Section 5.4, Section 5.5 and Section 5.6); the coloring tool is evaluated by running it once over the full RAGTruth held-out cohort with no further judgment involved (Section 5.7); an unconstrained-regeneration baseline is evaluated with the same claim-checking design as the primary RAGTruth result, on a separate matched sample (Section 5.8); and both primary human-labelled endpoints are checked through rationale-blinded second-author endpoint verification (Section 5.9). Table 1 summarizes the populations involved.
Both the RAGTruth 132-response panel and the ExpertQA 179-response panel underwent human/author review before their figures were treated as final; review depth and design differ between the two datasets because their underlying label schemas differ (Section 3.3), not because only one of the two received review.

5.1. Frozen Contrasts and Analysis Populations

Each ablation changes one stage while reusing its frozen parent outputs, run identically on both datasets except where a dataset’s label schema requires it. The response is the primary unit of analysis. Claim-level checks diagnose why a response receives a hallucination or grounding label, but completeness is judged only at whole-response level. The 100-response RAGTruth Wave-2 development cohort (Section 3.2) is excluded before formation of the 889-response held-out cohort and used only for method development.

5.2. RAGTruth Hallucination Measurement: Human Reference Plus Fine-Grained LLM Checking

B0’s label is never re-judged by the evaluator; RAGTruth human annotation is final for B0 (32 held-out responses human-positive, 857 human-clean). Only changed or repair-specific material is eligible to create a new repair error, preventing an LLM judge from inventing new B0 hallucinations or double-counting unchanged text. For changed material, the program follows a RefChecker-style design [10]: an extraction stage produces atomic subject–predicate–object claim triplets without seeing the evidence, and an independent checking stage evaluates each claim against the supplied evidence as entailment, contradiction, or neutral, recording the exact proposition, evidence quote, and reason, consistent with RAGChecker’s fine-grained claim-entailment principle [11] using a custom program rather than the official package.
Response transitions are mutually exclusive: removed (a human-labelled B0 hallucination is absent after repair), persisted (it remains), new (a repair-specific proposition is unsupported or contradicted after adjudication), clean on both (neither), and not applicable (the evaluator produced no valid machine-readable result). Stage 1 evaluates the full 132-response panel (32 human-positive plus 100 clean controls), which underwent the author review described above. Stage 2 evaluates the remaining 757 clean-reference responses; a second pass rechecks provisional new-hallucination cases to distinguish mere addition from evidence conflict.

5.3. ExpertQA Grounding Measurement: Two-Stage Machine Classification Plus Author Override

The ExpertQA arm evaluates the grounding axis over the full 731-response held-out cohort in a two-stage design: Stage 1 covers a 179-response sample and Stage 2 covers the remaining 552 responses. Machine response labels are: both grounded, grounding repaired, grounding failure persisted, new grounding failure, or retry required. A confirmation pass independently re-checks every Stage-2 new-grounding-failure candidate before it is accepted, so a provisional new-failure call is not taken at face value.
The Stage-1 sample additionally underwent author review via a per-claim override recorded against the machine’s provisional call. Three Stage-1 responses (expertqa_rr_gs_gpt4_1cb6ae7cbbafb526e1b4, expertqa_rr_gs_gpt4_690cefd6815e1b7fce3c, expertqa_rr_sphere_gpt4_9fc0dbd83ca1ff251a6d) carry an explicit author override reclassifying them from grounding-failure-persisted to grounding-repaired; in each case the machine’s own claim-level tracking had already marked the claim as unclear rather than confidently persisted, and the author’s review resolved that ambiguity toward repaired after inspecting the response directly. All figures in Section 6.1b use this author-reviewed Stage-1 count. Two further Stage-1 responses were reclassified from grounding-failure-persisted to grounding-repaired following rationale-blinded second-author endpoint verification and subsequent author consensus described in Section 5.9; see Section 6.8 for that review and Section 6.1b for its effect on the headline count.

5.4. Rendering-Quality Measurement

The B0-to-Renderer comparison uses no LLM and no readability score. Each deterministic action is matched to a predefined rendering contract: the intended defect must be present before, the expected correction after, and protected content must remain traceable. Primary outcomes are action-level validation and response-level elimination of screened defects. Paired safety screens separately record substantive numeric loss, unexpected numeric mutation, negation or modality loss, capitalized-entity candidates, and untraceable additions; the capitalized-phrase result is a candidate screen, not a named-entity error rate.

5.5. Grammar Benefit and Semantic-Preservation Measurement (Renderer vs. Fluency Ablation, 200-Response Sample)

For 200 responses deterministically sampled from the 821/889 held-out responses the Fluency Ablation changes, a blinded judge presents the Renderer’s output and the Fluency Ablation’s output in randomized order, chooses a grammar preference, and separately lists propositions present in only one version. Evidence-supported Fluency-Ablation-only propositions are scope expansions; unsupported or contradicted ones are new hallucinations; supported Renderer-only propositions absent from the Fluency Ablation are completeness loss. A net-safe result requires both a Fluency-Ablation grammar preference and no added or lost proposition. The 200-response results reported in Section 6.4 provide a judge-based screening estimate, with the 15.0% figure used to compare the relative safety of the Renderer and Fluency Ablation.

5.6. Whole-Response Completeness

Completeness is assessed on the complete B0 and rendered answers for all officially clean B0 responses (857 for RAGTruth, 602 for ExpertQA); no claim or hallucination-span decomposition is used. The forced categories are mutually exclusive: no loss retains all unique answer-bearing information; acceptable loss removes only redundant, boilerplate, or off-question material; minor loss removes useful context without changing the direct answer; major loss removes or changes the direct answer, decision-relevant content, or protected meaning. Only major loss is counted as a completeness cost in the headline trade-off.

5.7. Deterministic Provenance Coloring Evaluation (RAGTruth)

The tool described in Section 4.2.4 is run once, non-interactively, over RAGTruth’s complete held-out rendered-response file (889 responses), with zero additional model calls. Section 6 reports the resulting tier distribution, the fraction of responses whose rendered blocks could not be exactly relocated in the final answer text (a tiling-robustness diagnostic, not a factuality metric), and the mean evidence-coverage ratio per tier.

5.8. Baseline A: Unconstrained Single-Shot Regeneration Ablation (RAGTruth)

On the same 132-response matched RAGTruth panel used for the Stage-1 hallucination evaluation (32 human-positive responses plus 100 clean controls; Section 5.2), a second, independent baseline (Baseline A) generates one free single-shot rewrite of B0 using the same evidence sentences, the same model family (GPT-4-0613), and one model call per response. Unlike Repair, the model is free to reword, reorganize, and introduce its own phrasing anywhere in the answer while instructed to remain evidence-consistent; it is never restricted to copying evidence-sentence text verbatim, and no sentence-bound extraction step is applied. Holding evidence, model family, and sample fixed while removing only the copy-only constraint isolates the specific question Section 4.1’s design history was built around: what happens to hallucination removal and new-hallucination introduction when the model authors the factual replacement itself?
An independent LLM-judge evaluation, using the same claim-level, evidence-grounded design as Section 5.2 but a separate judge model, classifies each regenerated response’s relationship to B0 and the evidence into a mutually exclusive taxonomy: removed (an original human-labelled hallucination is absent from the regenerated answer), persisted (it remains), new unsupported claim (the regenerated answer introduces a claim not supported by the evidence, evaluated with precedence over removed, so a response that both fully repairs its original hallucination and introduces a new one is counted under new unsupported claim, since introducing new unsupported content is itself the adverse outcome this baseline is designed to surface), and clean (neither, for responses that started with no released hallucination). A deterministic verification layer independently re-checks every literal-span and evidence-citation claim the judge makes before accepting a shared-with-B0 or grounded classification, so paraphrased text that only superficially appears to carry B0’s content over is not accepted on the judge model’s word alone (Appendix A).

5.9. Rationale-Blinded Second-Author Endpoint Verification

Both primary human-evaluation endpoints (RAGTruth’s Stage-1 hallucination panel (Section 5.2) and ExpertQA’s Stage-1 grounding panel (Section 5.3)) were independently checked by a second author. The review sheets supplied the response identifier, B0, Repair output, relevant evidence, and the proposed binary and multicategory endpoints; first-author rationales and source labels were concealed. The second author recorded one classification per case: agree, partially agree, or disagree. For RAGTruth, the binary endpoint was hallucination present or absent and the transition endpoint was removed, persisted, new, or clean both. For ExpertQA, the binary endpoint was grounding failure present or absent and the transition endpoint was grounding repaired, grounding failure persisted, new grounding failure, or grounded both. Agreement was calculated before consensus. The two ExpertQA disagreements were then discussed, and the first author adopted the second author’s assessment in both cases; the consensus labels are used in the final 26/110 headline result.
Agreement statistics are computed separately per dataset, not pooled, using raw percent agreement with a Wilson 95% confidence interval, Cohen’s κ, Krippendorff’s nominal α, Gwet’s AC1 (bootstrapped, 10,000 replicates, seed 42), and an exact two-sided McNemar test on the binary-endpoint discordant pairs. A four-category κ/α is additionally computed where the second author’s multi-category label is identifiable from the review sheet; on ExpertQA, the reversing option identifies that the binary endpoint should flip but does not itself identify which of the two opposite-binary multi-category labels the second author would select, so the ExpertQA multi-category κ/α is reported as not estimable rather than approximated. Results appear in Section 6.8.

5.10. Statistical Analysis

Absolute paired counts are primary throughout. For RAGTruth’s B0-to-repaired comparison, an exact two-sided paired sign/binomial test on 14 improved and 2 worsened discordant outcomes (n = 16; ties excluded) gives p = 0.0042, equivalent to the exact McNemar formulation. For the B0-to-Renderer step, McNemar’s exact test compares paired response-level screened-defect presence. For grammar preference, an exact two-sided sign/binomial test excludes ties. Completeness outcomes are reported descriptively over each dataset’s full frozen clean-response population. ExpertQA’s grounding-repair axis uses the same exact two-sided paired sign test on 27 discordant response pairs (26 grounding repaired and 1 new grounding failure), giving p = 4.17 × 10−7. ExpertQA’s B0 state is fixed by its human support annotations, while the repaired state is evaluated separately and reduced to a response-level grounding endpoint. The 95% confidence interval for RAGTruth’s headline 37.50% net-reduction rate is calculated by delta-method variance propagation. The defined net effect is (14 − 2)/32; its variance is propagated from 14/32 removals in the human-positive cohort and 2/857 new failures in the human-clean cohort as independent binomial counts, yielding a 95% CI of [18.3%, 56.7%].

6. Results

6.1. RAGTruth: Hallucination Reduction, B0 Versus Repair

Among the 32 RAGTruth human-positive responses, Repair removed all released hallucinated content in 14 and left it persisting in 18. Two repair-specific adverse additions were counted conservatively: response 13812 was a Stage-1 new-claim case whose checker label was neutral, and response 17448 was a Stage-2 case containing mutually inconsistent source directions about whether a controller’s Connect button is on the back or front edge; both are reported as author-adjudicated adverse additions rather than asserted as a single raw model label. The Stage-2 evaluator completed 753/757 clean-reference cases: responses 14682, 16494, and 16716 failed after repeated malformed/truncated JSON outputs, and response 14652 failed deterministic literal-match validation; all four remain not-applicable, with no label imputed.
The conservative adverse-response count therefore changed from 32 at B0 to 20 after Repair among 885 evaluable outputs (865 evaluable responses clean after repair). The exact paired sign test on the 14 improved vs. 2 worsened discordant pairs gives p = 0.0042. This supports a directional reduction, but the small positive set (n = 32), LLM-assisted checking, author adjudication, and four not-applicable outputs limit the strength of the point estimate.
Table 2. RAGTruth: B0-to-Repair outcome counts on the 32-response human-positive panel.
Table 2. RAGTruth: B0-to-Repair outcome counts on the 32-response human-positive panel.
Outcome Count Denominator
Removed (hallucination absent after Repair) 14 32 human-positive
Persisted 18 32 human-positive
Adverse new/conflicted addition (author-adjudicated) 2 885 evaluable
Conservative adverse count after Repair 20 885 evaluable
Net reduction 12/32 (37.50%; 95% CI [18.3%, 56.7%]) exact paired p = 0.0042
Stage-2 evaluator not-applicable (excluded, not imputed) 4 889 held-out
The 32-response denominator is small; Section 10 records this as a limit on the precision of this point estimate, and Section 6.1 b reports the larger ExpertQA sample as an independent check. A 95% CI on the 37.50% net-reduction rate, computed by delta-method independent-binomial variance propagation, is [18.3%, 56.7%].

6.1b. ExpertQA: Evidence-Grounding Repair

The full 731-response held-out cohort resolves to: 610 both grounded, 26 grounding repaired, 84 grounding failure persisted, 10 retry required, and 1 new grounding failure (610+26+84+10+1 = 731). Restricting to responses with an original grounding problem (repaired or persisted; both-grounded and retry-required responses are, by construction, not counted here) gives 110 responses, of which Repair resolved 26 (23.64%) and left 84 (76.36%) persisting: the figure used throughout this paper incorporates the Stage-1 author overrides (Section 5.3) and, for two further responses, rationale-blinded second-author endpoint verification and author consensus (Section 6.8). The 1 new-grounding-failure response arises entirely outside this 110-response repair-eligible set (it is a Stage-2 response that started with no original grounding problem), so it is a separate, additional side effect of repair, not a subtraction from the 26/110 repair count.
This grounding-repair effect is statistically significant under the same exact paired sign-test design used for RAGTruth. Among 27 discordant response pairs, 26 favored repair and 1 favored a new grounding failure (p = 4.17 × 10−7; 95% Clopper–Pearson CI for the improvement share, 26/27: [81.0%, 99.9%]).
An approximate net-reduction figure, mirroring RAGTruth’s construction: RAGTruth’s headline 12/32 (37.50%) figure is computed as repaired minus adverse new additions over the same accounting base (Section 6.1). Applying the same subtraction to ExpertQA: 26 grounding-repaired minus the 1 new-grounding-failure case, over the 110-response repair-eligible denominator, gives 25/110 (22.73%). This is reported as an approximate, descriptive parallel only: unlike RAGTruth’s adverse additions, ExpertQA’s 1 new-grounding-failure case did not arise within the 110-response repair-eligible set itself (it is a Stage-2 response that started both-grounded), so the subtraction mixes two different populations rather than netting outcomes drawn from one matched group. The paper’s primary reported ExpertQA figure therefore remains the un-adjusted gross repair rate, 26/110 (23.64%); the 25/110 (22.73%) net figure is offered alongside it for comparability with RAGTruth’s net-reduction framing, not as a replacement for it.
The Stage-2 confirmation pass was rerun under a cleared cache, so every one of the six raw Stage-2 new-grounding-failure candidates was re-evaluated fresh rather than reusing any earlier cached judgment. Five of the six were demoted to both-grounded (the independent recheck found their flagged content was in fact evidence-supported). The remaining one could not be confirmed either way (its confirmation call failed deterministic validation on every retry) and is conservatively retained as a new grounding failure rather than dropped, consistent with this paper’s fail-closed treatment of unresolved cases elsewhere.
Table 3. ExpertQA: response-level grounding outcomes, full 731-response held-out cohort.
Table 3. ExpertQA: response-level grounding outcomes, full 731-response held-out cohort.
Outcome Stage 1 (179) Stage 2 (552) Full 731
Both grounded 63 547 610
Grounding repaired 26 0 26
Grounding failure persisted 84 0 84
Retry required 6 4 10
New grounding failure (1 unconfirmed, fail-closed) 0 1 1
Total 179 552 731
An exact two-sided paired sign test on the 27 discordant pairs gives p = 4.17 × 10−7; the 95% Clopper–Pearson CI for the discordant-pair improvement share is [81.0%, 99.9%].

6.2. Coverage Gate: Cascade Coverage (RAGTruth)

The zero-API Coverage Gate classified 3,740 of 4,559 located claim mappings as resolved and eligible for rendering, giving 82.036% claim-resolution coverage. The remaining 819 unresolved mappings were blocked from rendering and retained their Repair-stage text; they were not treated as 819 rejected responses. At response level, 821/889 contained at least one resolved mapping. The five response categories sum exactly to the 889-response cohort: 3 with no usable evidence at all, 54 with no extracted claims, 11 with unresolved claims only, 468 partially resolved, and 353 fully resolved. No complete response was rejected by the Coverage Gate.

6.3. RAGTruth: B0 Versus Renderer Output

The Renderer changed 622/889 responses (69.97%) without a model call, reflecting a strict minimality invariant (it never introduces a source sentence Repair did not select). It validated 1,522/1,524 predefined targeted actions in changed responses (99.87%); 594/595 responses containing targeted actions passed every action contract (99.83%). Across all 889 responses, paired screened rendering defects fell from 209 (23.51%) to 24 (2.70%). Among the 622 changed responses, content-preservation screens found 0 substantive numeric-value loss, 1 unexpected numeric-mutation action, 0 negation loss, 0 modality loss, 70 capitalized-entity-candidate losses (11.25%, a candidate screen rather than a named-entity error rate), and 0 untraceable additions.
The Renderer’s own completeness check flags at least one substitution in 516/622 changed responses (82.96%), a lost-negation preservation signal in 85/622 (13.67%), and more than four flags in 20/622 (3.22%). These substitution-level records feed the deterministic provenance-coloring tool described in Section 4.2.4.
The valid conclusion is improved deterministic rendering quality, not improved human comprehension; readability score is not the primary endpoint because punctuation, list markers, seams, and source-block structure are the intervention targets.

6.4. RAGTruth: Renderer Versus Fluency Ablation (200-Response Sample)

Across the full 889-response held-out cohort, the Fluency Ablation changes 821/889 responses (92.35%); 200 of these were deterministically sampled for blinded judging. The judge preferred the Fluency Ablation for grammar in 132/200 (66.0%), preferred the Renderer in 60/200 (30.0%), and tied in 8/200 (4.0%); the exact sign test excluding ties gives p = 2.18 × 10−7. Only 49/200 (24.5%) met the stricter net-safe endpoint (Fluency Ablation preferred and no proposition added or lost). Semantic preservation (no one-sided proposition at all) held in 68/200 (34.0%). The judge’s raw, not-yet-author-adjudicated output flags a new hallucination in 30/200 (15.0%), a source-supported scope expansion in 101/200 (50.5%), and a supported-fact loss in 84/200 (42.0%); a deterministic protected-content screen fires on 188/200 (94.0%) but is diagnostic only.
Table 4. Renderer versus Fluency Ablation, 200-response blinded-judge sample.
Table 4. Renderer versus Fluency Ablation, 200-response blinded-judge sample.
Outcome Count / 200 Rate
Fluency Ablation grammar preferred 132 66.0%
Renderer grammar preferred 60 30.0%
Grammar tie 8 4.0%
Net safe grammar improvement 49 24.5%
Semantic preservation (no one-sided proposition) 68 34.0%
New hallucination (judge-raw, not author-adjudicated) 30 15.0%
Source-supported scope expansion 101 50.5%
Supported fact loss 84 42.0%

6.5. Whole-Response Completeness: Major Loss as the Completeness Cost

Only major loss is treated as a real completeness cost; minor loss and acceptable loss are reported for completeness of the picture but are not summed into the cost figure used in Section 6.10. For RAGTruth (857 officially clean B0 responses): 599 no loss (69.89%), 11 acceptable loss (1.28%), 223 minor loss (26.02%), 24 major loss (2.80%). For ExpertQA (602 officially clean B0 responses): 295 no loss (49.00%), 11 acceptable loss (1.83%), 282 minor loss (46.84%), 14 major loss (2.33%).
Table 5. Whole-response completeness outcomes on officially clean B0 responses.
Table 5. Whole-response completeness outcomes on officially clean B0 responses.
Category RAGTruth (n = 857) ExpertQA (n = 602)
No information loss 599 (69.89%) 295 (49.00%)
Acceptable loss 11 (1.28%) 11 (1.83%)
Minor loss 223 (26.02%) 282 (46.84%)
Major loss (the counted cost) 24 (2.80%) 14 (2.33%)
The two datasets agree closely on the major-loss cost rate (2.80% vs. 2.33%) despite disagreeing sharply on the minor-loss rate (26.02% vs. 46.84%). ExpertQA’s long-form expert answers apparently carry substantially more replaceable supporting/contextual detail than RAGTruth’s shorter QA responses, so deterministic evidence-bound substitution more often trims useful-but-non-decisive context on ExpertQA, a genuine, dataset-dependent cost that the minor/major split makes visible instead of hiding inside one aggregate score. Because only major loss changes or removes the direct answer, the strict cost figure both datasets converge on for the headline trade-off is approximately 2.3–2.8% of officially clean responses.

6.6. Baseline A: Unconstrained Regeneration Versus Deterministic Repair (RAGTruth, Matched 132-Response Sample)

130/132 sampled responses were successfully regenerated under Baseline A (2 exhausted retries at generation and are excluded, both from the clean-control controls); of these, 129/130 were successfully evaluated (1 exhausted retries at evaluation, response 14754, from the human-positive panel). The matched sample decomposes into a 32-response positive cohort (an original released hallucination present) and a 98-response evaluated clean-control cohort (no original hallucination; 2 of the nominal 100 controls were lost to the generation-stage retries above).
Table 6. Baseline A (unconstrained regeneration): outcomes on the matched 132-response sample.
Table 6. Baseline A (unconstrained regeneration): outcomes on the matched 132-response sample.
Cohort Outcome Count Rate
Positive (n = 32) Removed 15 46.88%
Positive (n = 32) Persisted 15 46.88%
Positive (n = 32) New unsupported claim 1 3.13%
Positive (n = 32) Retry-exhausted (excluded) 1 3.13%
Clean-control (n = 98 evaluated) Clean 97 98.98%
Clean-control (n = 98 evaluated) New unsupported claim 1 1.02%
“New unsupported claim” in the positive cohort is not a third, independent outcome alongside removed/persisted: it is the one positive-cohort response whose original hallucination was removed but whose regenerated answer introduced a different, new evidence-unsupported claim in the same rewrite. The taxonomy’s precedence rule (Section 5.8) assigns this case to new-unsupported-claim rather than removed, since introducing new unsupported content is the adverse outcome this baseline is designed to surface; it is not persisted, because the original released hallucination is in fact gone. Removed, persisted, and new-unsupported-claim are mutually exclusive per response by this precedence rule, so the four positive-cohort rows still sum to 32.
Among the 30 positive-cohort responses that resolved to a removed-or-persisted verdict, exactly half were removed (15/30 = 50.00%). One positive-cohort response is worth naming individually: its regenerated answer fully repaired the original hallucination and introduced a new unsupported claim in the same rewrite, and the taxonomy’s precedence rule (Section 5.8) correctly assigns it new-unsupported-claim rather than removed, a concrete illustration of why “hallucination removed” and “no new hallucination introduced” must be tracked as separate outcomes rather than treated as fungible, the same principle Section 9.2 applies to Repair’s own cascade accounting. Pooled across all 129 evaluated responses, 2 (1.55%) received a new-unsupported-claim label.
The comparison against Repair is the point of this ablation. Baseline A’s gross hallucination-removal rate (15/32, 46.88%) is comparable to, and nominally higher than, Repair’s gross removal rate on the same 32-response positive cohort (14/32, 43.75%, Section 6.1); free regeneration is not worse at removing the originally flagged hallucination. But Baseline A’s adverse-new-content rate on clean-control responses (1/98, 1.02%) is approximately 4.5 times Repair’s adverse-addition rate measured over its much larger 885-response evaluable held-out cohort (2/885, 0.226%, Section 6.1), despite Baseline A’s clean-control exposure (98 responses) being roughly nine times smaller than Repair’s. Holding evidence, model, and sample fixed, removing the model’s authority to author the factual replacement (the single design choice separating Repair from Baseline A) is associated with a substantially lower rate of newly introduced, evidence-unsupported content at an equivalent hallucination-removal rate.
Table 7. Repair versus Baseline A: hallucination removal and new-content rate.
Table 7. Repair versus Baseline A: hallucination removal and new-content rate.
Metric Repair (Section 6.1) Baseline A (unconstrained regen.)
Gross hallucination removal, positive cohort 14/32 (43.75%) 15/32 (46.88%)
Net reduction, positive cohort (adjusted for adverse additions) 12/32 (37.50%), p = 0.0042 not computed (see remark)
New/adverse addition rate, clean-control responses 2/885 (0.226%) 1/98 (1.02%)
Repair’s net-reduction figure and its adverse-addition rate are computed over different, non-matched denominators (a 32-response positive cohort and an 885-response held-out cohort respectively); Baseline A’s smaller, fully matched 132-response sample does not support an equally powered net-reduction significance test. The comparison is therefore reported descriptively: it shows comparable hallucination removal and a several-fold difference in new-content introduction without asserting an inferentially matched effect.

6.7. Deterministic Provenance Coloring: RAGTruth Results

The tool described in Section 4.2.4 was run once over the complete RAGTruth rendered-response file (889 responses), with zero additional model calls. Every rendered block in every response was successfully relocated in its final answer text (0 tiling failures across 889 responses), confirming that the coloring mechanism is robust at full scale, not only on the synthetic self-test fixtures used during development.
Table 8. RAGTruth: confidence-tier distribution, full 889-response held-out cohort.
Table 8. RAGTruth: confidence-tier distribution, full 889-response held-out cohort.
Confidence tier RAGTruth (n = 889)
High confidence (fully resolved, no flags) 72 (8.10%)
Medium confidence (resolved but flagged) 281 (31.61%)
Low confidence (unresolved claim present) 479 (53.88%)
No claims detected 57 (6.41%)
Figure 2. Confidence-tier distribution produced by the deterministic evidence-provenance coloring tool, full RAGTruth held-out cohort (n = 889). Bars are ordered High to Low confidence; “No claims detected” is a distinct, non-ranked category shown in gray.
Figure 2. Confidence-tier distribution produced by the deterministic evidence-provenance coloring tool, full RAGTruth held-out cohort (n = 889). Bars are ordered High to Low confidence; “No claims detected” is a distinct, non-ranked category shown in gray.
Preprints 231785 g002
High plus medium confidence together cover 39.71% of the RAGTruth held-out cohort, and “high confidence” is a genuinely rare, conservative label (8.10%) rather than a default outcome the tool is tuned to award liberally: an honest, load-bearing finding rather than one to minimize.
What this means for a reviewer. For roughly two in five RAGTruth held-out responses, a reviewer can be shown the response with a majority of its evidence-bound content already colored green or amber and a defensible tier label, at zero marginal cost, and can go directly to any amber span rather than re-reading the whole answer. The underlying per-span coloring remains available for every RAGTruth response regardless of tier. Instead of deciding where to look across the whole response, the reviewer is directed to the specific incident-level spans requiring attention.

6.8. Second-Author Agreement on the Primary Human-Evaluation Endpoints

Both primary human-evaluation endpoints (Section 5.9) underwent rationale-blinded second-author endpoint verification. On RAGTruth, all 144 eligible cases were completed with 0 excluded and 100% raw binary agreement (Wilson 95% CI [0.974, 1.0]); binary Cohen’s κ = 1.000 (bootstrap 95% CI [1.0, 1.0], 10,000 replicates), Krippendorff’s nominal α = 1.000, Gwet’s AC1 = 1.000, and the exact McNemar test is uninformative by construction (0 discordant pairs in either direction, p = 1.000). The four-category endpoint is fully identifiable on RAGTruth and also agrees perfectly (multi-category κ = α = 1.000): all 14 removed, 18 persisted, 1 new, and 111 clean-on-both first-author labels were independently confirmed by the second author.
On ExpertQA, of 179 Stage-1 cases, 173 were eligible (6 excluded for a machine-readable output error, not disagreements), with 171 agreements and 2 disagreements; raw binary agreement 98.84% (Wilson 95% CI [0.9588, 0.9968]); binary Cohen’s κ = 0.977 (bootstrap 95% CI [0.942, 1.0]), Krippendorff’s nominal α = 0.977, Gwet’s AC1 = 0.977 (bootstrap 95% CI [0.942, 1.0]), and an exact two-sided McNemar test on the 2 discordant pairs (both first-author-present/second-author-absent, 0 in the reverse direction) gives p = 0.500. Both disagreements were Stage-1 responses where the first author’s endpoint was grounding-failure-persisted and the second author selected the reversing option; the authors subsequently reached consensus adopting the second author’s assessment for both, reclassifying them to grounding-repaired; Section 6.1b’s headline ExpertQA figure (26/110) reflects this resolution. Because the reversing option identifies only that the binary endpoint should flip, not which four-category label replaces it, the ExpertQA four-category κ/α remains not estimable for the 173-case agreement calculation itself, reported honestly as such rather than approximated; the subsequent consensus step that resolved these two cases to a specific label was a separate author discussion, not a recomputation of this statistic.
Table 9. Rationale-blinded second-author agreement on the primary human-evaluation endpoints.
Table 9. Rationale-blinded second-author agreement on the primary human-evaluation endpoints.
Statistic RAGTruth ExpertQA
Eligible cases 144 173 (of 179; 6 excluded)
Raw binary agreement 100.00% 98.84%
Wilson 95% CI [0.974, 1.0] [0.9588, 0.9968]
Cohen’s κ (binary) 1.000 0.977
Krippendorff’s α (binary) 1.000 0.977
Gwet’s AC1 1.000 0.977
McNemar exact two-sided p 1.000 (0 discordant) 0.500 (2 discordant)
Four-category κ/α 1.000 / 1.000 not estimable (see text)
Figure 3. Rationale-blinded second-author agreement on the primary human-evaluation endpoints: raw agreement, Cohen’s κ, and Gwet’s AC1 for RAGTruth (n = 144) and ExpertQA (n = 173 eligible cases), all on a 0-1 scale.
Figure 3. Rationale-blinded second-author agreement on the primary human-evaluation endpoints: raw agreement, Cohen’s κ, and Gwet’s AC1 for RAGTruth (n = 144) and ExpertQA (n = 173 eligible cases), all on a 0-1 scale.
Preprints 231785 g003
These results support the reliability of the paper’s primary human-evaluation endpoints: perfect agreement on the RAGTruth hallucination endpoint and agreement in the high-0.97 range on the ExpertQA grounding endpoint, with its two disagreements resolved by author consensus and reflected in the final counts used throughout Section 6.1b, 6.10, and 8.

6.9. Efficiency: Cost and Latency, and the Case for a Second Judge Pass

The held-out locator call is the pipeline’s only per-response model cost. On RAGTruth (889 responses) it used 1,327,794 total tokens for USD 48.1718 (USD 0.054187/response mean) at a mean latency of 8.3345 s/response. On ExpertQA (731 responses) it used 1,424,674 total tokens for USD 51.6100 (USD 0.070602/response mean) at a mean latency of 9.3906 s/response. Every stage after the locator call is free in the marginal sense that matters for a production deployment: the Coverage Gate performs zero model calls; the Renderer’s deterministic render plus its completeness check runs in a mean of 0.00207 s/response on RAGTruth and 0.00190 s/response on ExpertQA, with $0 marginal cost; and RAGTruth’s coloring pass (Section 4.2.4) adds a third deterministic step over the same output row, again with zero additional model calls and negligible latency (889 responses, 0 tiling failures, single non-interactive batch).
Table 10. Per-response cost and latency by pipeline stage.
Table 10. Per-response cost and latency by pipeline stage.
Stage Model calls / response Mean $ / response Mean latency / response
Locator (both datasets, required) 1 USD 0.054 (RAGTruth) / 0.071 (ExpertQA) 8.33 s / 9.39 s
Coverage Gate 0 $0 negligible (no model call)
Renderer + completeness check 0 $0 0.0021 s / 0.0019 s
Coloring pass, RAGTruth (Section 4.2.4) 0 $0 negligible (batch, 0 tiling failures / 889)
Added second-judge pass (grammar/semantic, Section 6.4) +1 per response ≈USD 0.059 (estimated, see text) ≈9.03 s/response (estimated, see text)
The source manifest records no separate $ rate or latency figure for the second-judge call, so both are estimated rather than metered, using the same scaling logic in both cases: applying the primary RAGTruth locator’s own realized per-token rate (USD 48.1718 / 1,327,794 tokens ≈ USD 0.0000363/token) to the second pass’s ≈1,619 tokens/response gives an estimated USD 0.059/response, about 108% of the locator’s own USD 0.054/response, for a combined estimated cost of roughly USD 0.113/response when both calls run. Applying the same 1,619/1,493.6 ≈ 1.084 token-ratio scaling to the locator’s own mean latency (8.3345 s/response) gives an estimated 9.03 s/response for the second-judge call, for a combined estimated latency of roughly 17.4 s/response when both calls run. Both estimates assume the second-judge call runs at a comparable per-token rate and per-token latency to the locator call, which is not independently verified, so they are reported as order-of-magnitude approximations rather than metered figures.
Beyond this cost, a permanent second-judge stage introduces a structural problem that the deterministic pipeline avoids by construction: if the second judge is added to check the first judge’s (the locator’s) output, the two can disagree, and disagreement between two judges is not self-resolving. Adjudicating that disagreement requires either an arbitrary tie-break rule (which silently discards one judge’s signal) or a third verifier to break the tie, at which point the same problem recurs one level up: a third judge can also disagree with the first two, motivating a fourth, and so on without a principled stopping point. The Renderer’s deterministic, evidence-bound realization step is preferred precisely because it is not a second opinion competing with the locator’s judgment; it is a fixed, auditable transformation of the locator’s own output, so its result is reproducible across runs and does not introduce a fresh source of inter-judge disagreement that would itself need adjudicating. Where a second, judge-based check is still useful, as in the grammar/semantic ablation of Section 6.4, it is reported as a bounded, disclosed evaluation exercise (Section 6.4, Section 9.3) rather than adopted as a permanent pipeline stage.

6.10. Integrated Interpretation

Four results, taken together, support a consistent interpretation. First, evidence-bound Repair produces a directional net reduction in adverse responses on both datasets: 12/32 (37.50%, p = 0.0042) on RAGTruth’s human-labelled hallucination axis, and 26/110 (23.64%) repaired on ExpertQA’s grounding axis, both independently supported by near-perfect second-author agreement on the underlying endpoints (Section 6.8). The Coverage Gate then shows why the RAGTruth repair rate is not higher: only 82.036% of located claim mappings are eligible for deterministic rendering at all (Section 6.2), so repair coverage, not repair correctness, is a primary limiting factor.
Second, the Baseline A ablation (Section 6.6) shows this is not merely a design preference: at a comparable or nominally higher hallucination-removal rate, unconstrained regeneration introduces new, evidence-unsupported content at roughly 4.5 times the rate of deterministic realization. Hallucination reduction without adding new unsupported content (this paper’s central positioning) is measurably easier to achieve when the model is not the author of the factual replacement.
Third, deterministic rendering materially improves RAGTruth’s specific rendering defects (23.51% → 2.70% screened-defect rate) with almost no content-preservation flags, while reopening generation with the Fluency Ablation trades a real grammar preference (66.0%) for a substantial and not-yet-fully-adjudicated new-hallucination rate (15.0% raw) at held-out scale, and hallucination reduction is achieved at a quantified and honestly bounded completeness cost: 2.80% of RAGTruth’s and 2.33% of ExpertQA’s officially clean responses, with a much larger and dataset-dependent minor-loss share (26.02% vs. 46.84%). The two datasets’ close agreement on the major-loss rate, despite their sharply different domain and answer length, is itself evidence that this cost is a structural property of evidence-bound substitution rather than an artifact of one dataset’s annotation style.
Fourth, the provenance layer foregrounded in this paper (coloring, evaluated at incident level in Section 4.2.4 and across the full RAGTruth held-out population in Section 6.7) gives a reviewer a judge-free, zero-marginal-cost view into an individual response, at the honestly reported cost of its most reassuring label being rare (8.10% high confidence) rather than a default outcome.
Hallucination reduction, grounding repair, repair coverage, rendering quality, grammar, completeness, and provenance signaling are therefore reported as separate outcomes, validated against an unconstrained-regeneration alternative and rationale-blinded second-author endpoint verification. The two distinct primary endpoints move in the same favorable direction, while major-loss costs are similar and minor-loss rates differ meaningfully across datasets.

7. Development Separation and Robustness Design

Section 3.2 introduced the 100-response Wave-2 development cohort and the resulting 889-response held-out cohort; this section gives the independent verification behind that split and the robustness reasoning that depends on it. The RAGTruth GPT-4-0613 QA arm contains 989 responses: 42 human-positive and 947 human-clean. Manifest arithmetic (989 − 889 = 100 responses; 42 − 32 = 10 hallucinated; 947 − 857 = 90 clean) and an independent direct count against the raw development-claims file (counting unique response IDs with human_hallucination_overlap = True) agree exactly on 10 human-positive and 90 clean development responses, not the approximately 20 hallucinated responses referenced in early working notes, an estimate this cross-check supersedes. The 100-response development cohort was used to freeze prompts, sentence-bound evidence representation, validation rules, rendering behavior, and the initial Fluency Ablation design; it is not pooled into the 889-response effectiveness denominator, and the supplied 889-response held-out output has 0 ID overlap with the development checkpoints.

8. Cross-Dataset Replication

An earlier, single-dataset scope of this study did not evaluate ExpertQA. In this paper, ExpertQA’s GPT-4 retrieve-and-read arm (731 held-out responses; 129 factual-error-proxy positive, 602 clean; 622 grounding-failure-proxy positive with 1,746 claims; 211 total human factual-error claims, of which 26/110 reviewed Stage-1 grounding failures are repaired and 84/110 persist, Section 6.1b) is fully evaluated under the pipeline described in Section 4, with schema-only adaptation for ExpertQA’s answer/claim/evidence field names. This section compares mechanism-level behavior while keeping RAGTruth’s hallucination endpoint and ExpertQA’s grounding endpoint distinct.
Figure 4. Cross-dataset mechanism comparison: RAGTruth hallucination-reduction rate (Section 6.1), ExpertQA grounding-repair rate (Section 6.1b), and major-loss completeness cost (Section 6.5). The two primary rates describe distinct endpoints and are not pooled.
Figure 4. Cross-dataset mechanism comparison: RAGTruth hallucination-reduction rate (Section 6.1), ExpertQA grounding-repair rate (Section 6.1b), and major-loss completeness cost (Section 6.5). The two primary rates describe distinct endpoints and are not pooled.
Preprints 231785 g004
Where the two datasets agree: both show a net reduction in adverse responses under Repair (RAGTruth 12/32 = 37.50%; ExpertQA 26/110 = 23.64% grounding-repaired); both show the Renderer succeeding on the large majority of eligible claim mappings with very low residual defect rates; both converge on an almost identical major-loss completeness cost (2.80% RAGTruth vs. 2.33% ExpertQA, Section 6.5); and both have their primary human-labelled endpoint independently confirmed by rationale-blinded second-author endpoint verification at agreement levels of κ ≥ 0.97 (Section 6.8).
Where the two datasets diverge: Major loss changes the direct answer, a decision-relevant fact, or other protected meaning, whereas minor loss removes useful supporting context without changing the direct answer (Section 6.5). ExpertQA’s minor-loss rate is higher than RAGTruth’s (46.84% vs. 26.02%), while their major-loss rates are similar (2.33% vs. 2.80%). This difference is consistent with ExpertQA’s longer answers containing more replaceable supporting context around a comparatively compact direct answer.
ExpertQA’s grounding-repair axis is independently significance-tested (p = 4.17 × 10−7) using the same response-level discordant-pair sign-test structure as the RAGTruth headline result. The datasets retain different raw annotation schemas as a deliberate measurement-validity choice: each endpoint is tested against its own fixed human reference rather than pooled into a shared schema that neither dataset natively provides.
The favorable direction of change in each dataset’s distinct primary endpoint, together with the similar major-loss cost rate, provides cross-dataset evidence for the mechanism across different domains, generator behavior, and annotation protocols. The outcomes remain conceptually and numerically separate. Divergence on minor-loss rate is reported as genuine dataset dependence rather than treated as a discrepancy to explain away or retune the frozen method against.

9. Discussion

9.1. What Deterministic Factual Realization Changes

The core intervention reduces factual-generation freedom rather than attempting unrestricted correction, on both datasets. On RAGTruth, the paired ablation supports this mechanism directly: 14 originally hallucinated responses were removed, two adverse additions were counted, and the net decrease was 12 responses (exact paired p = 0.0042). On ExpertQA, the analogous grounding-repair figure is 26/110 (23.64%; exact paired sign test, p = 4.17 × 10−7) after completed author adjudication and rationale-blinded second-author endpoint verification. The Baseline A comparison (Section 6.6) shows this restriction is doing real work, not merely adding process overhead: removing the model’s authority to author the factual replacement is associated with a roughly 4.5-times-lower new-hallucination rate than letting the same model regenerate freely, at an equivalent removal rate.

9.2. Cascade and Claim-Closure Coverage Are Part of the Result, Not an Implementation Detail

The 82.036% RAGTruth claim-resolution coverage (Section 6.2) shows why cascade accounting is necessary rather than optional: a fail-closed system can appear safe if unrepaired claims are simply omitted from analysis. Reporting coverage and correctness together prevents conditional success on one axis from being mistaken for end-to-end success.

9.3. Completeness and Fluency Are Independent Safety Dimensions from Hallucination Reduction

The ablations show that different quality dimensions move independently, on both datasets. Deterministic rendering removed the specific rendering defects targeted by the correction contracts (residual screened-defect rate 2.70% on RAGTruth), yet a real, quantified major-loss cost remains on both datasets (2.80% RAGTruth, 2.33% ExpertQA) alongside a much larger and dataset-dependent minor-loss share. Conversely, the Fluency Ablation is preferred for fluency 66% of the time on the full 200-response held-out sample, yet introduces new, evidence-unsupported propositions in a judge-raw 15% of changed responses, a rate that has not yet been author-adjudicated at this scale. A hallucination-reduction paper should therefore continue to avoid collapsing hallucination reduction, grammar preference, and completeness cost into a single quality score.

9.4. Operational Implications: Provenance Signaling Without a Second Judge

The architecture is operationally attractive specifically because the only mandatory per-response model cost is the single locator call; every subsequent stage (the Coverage Gate, the Renderer and its completeness check, and the RAGTruth coloring tool (Section 4.2.4)) is deterministic and free. A genuinely judge-free, zero-marginal-cost review signal is available today for RAGTruth, but its most reassuring label is rare (8.10% high-confidence) rather than a default outcome. A production deployment adopting this architecture should treat the coloring tool’s output as an always-available but conservative reviewer aid that turns whole-response review into targeted, incident-level review, rather than as a substitute for periodic human or judge-based sampling where that budget exists.

10. Limitations

Both headline tests are based on small numbers of discordant response pairs (RAGTruth n = 16; ExpertQA n = 27). Accordingly, RAGTruth’s 37.50% net-reduction estimate has a wide 95% CI [18.3%, 56.7%]. ExpertQA’s exact paired p = 4.17 × 10−7 reflects a strongly directional result but still arises from only 27 discordant pairs.
The datasets retain measurement designs matched to their native annotations. RAGTruth provides span-level hallucination annotations from which a response-level binary endpoint is derived and assessed using RefChecker-style claim extraction and evidence entailment (Section 5.2). ExpertQA provides a multi-category per-claim support taxonomy and is assessed through machine classification followed by recorded author review (Section 5.3). In both datasets, the original human reference is fixed and the repaired response is evaluated separately; the prespecified primary endpoint decisions then undergo rationale-blinded second-author endpoint verification. This dataset-specific design preserves annotation meaning while yielding the same response-level discordant-pair sign-test structure. The results provide independent corroboration but are not numerically interchangeable.
The same GPT-4 model family generated the original responses and performed semantic localization on both datasets, which controls model-family variation but may share systematic errors. Unresolved claims and unclosed evidence bindings remain unchanged from the source text; fail-closed repair, and a low-confidence coloring tier, are not equivalent to a verified-safe final answer. The rendering-defect screens (Section 5.4) are deterministic proxies for rendering quality and do not measure human comprehension.
The Baseline A comparison (Section 6.6) is descriptive: it uses a 132-response sample, whereas Repair is evaluated on an 885-response held-out cohort. The observed 4.5-fold difference in new-content introduction provides complementary evidence for the proposed mechanism, but the unequal sample sizes do not support a directly powered inferential comparison. This result is therefore interpreted as supporting evidence rather than a definitive estimate of relative performance.
Section 4.1 reports a claim-selection heuristic that did not concentrate hallucination risk; this negative result motivated the final design choice to check every claim the locator can identify. The RAGTruth coloring tool (Section 4.2.4 and Section 6.7) remains a deterministic provenance aid rather than a verified-safe-answer certificate.
A source-code audit of the frozen ablation programs used in this study found no data leakage, no overlap between the samples used to derive rules and the samples used to test them, and no hardcoded or fabricated results; every reported figure traces to a live, per-response computation. One caching gap was found in a small confirmation step of the ExpertQA grounding ablation: its cache did not track which prompt version produced a cached result. That step was rerun under a cleared cache; the rerun left the headline repair figure unchanged (Section 6.1b).
No third dataset or generator family has been evaluated; generalization beyond GPT-4-family RAG generation on RAGTruth-style and ExpertQA-style tasks remains untested.

11. Conclusions

This study separates semantic localization from deterministic factual realization and evaluates the resulting pipeline against two distinct, independently annotated endpoints: hallucination transitions on RAGTruth and evidence-grounding transitions on ExpertQA. Every quantity is reported after the completed human-review step and rationale-blinded second-author endpoint verification. On 889 held-out RAGTruth GPT-4-0613 QA responses, Repair removed 14 of 32 human-labelled hallucinated responses, 18 persisted, and two author-adjudicated adverse additions were counted; among 885 evaluable outputs, the conservative net reduction was 12/32 (37.50%; exact paired p = 0.0042, 95% CI [18.3%, 56.7%]). The non-ablation Coverage Gate resolved 3,740/4,559 claim mappings (82.036%) and rejected no complete response. The zero-API Renderer then reduced the residual screened rendering-defect rate to 2.70% of the held-out cohort.
On 731 held-out ExpertQA responses, after completed author adjudication and rationale-blinded second-author endpoint verification, Repair resolved 26 of 110 reviewed grounding failures (23.64%), with 84 persisting. The corresponding paired effect was independently significant (exact paired sign test, p = 4.17 × 10−7). This is an independently annotated evidence-grounding result, not a RAGTruth-style hallucination label. Both endpoints underwent rationale-blinded second-author endpoint verification, with 100% raw agreement on RAGTruth (κ = 1.000, n = 144) and 98.84% raw agreement on ExpertQA (κ = 0.977, n = 173).
A direct unconstrained-regeneration baseline on a matched 132-response RAGTruth sample shows why the deterministic mechanism matters: at a comparable or nominally higher hallucination-removal rate (46.88% vs. 43.75%), free regeneration introduces new, evidence-unsupported content at roughly 4.5 times Repair’s rate (1.02% vs. 0.226%). Removing the model’s authority to author the accepted factual replacement is therefore not only a design principle but also a measured descriptive safety advantage within the evaluated RAGTruth sample.
The safety costs are equally important and are quantified consistently on both datasets, counting only a change to the direct answer as the cost: 2.80% of RAGTruth’s and 2.33% of ExpertQA’s officially clean responses, alongside a larger and dataset-dependent minor-loss share (26.02% and 46.84%). A full 200-response held-out ablation shows the Fluency Ablation is fluency-preferred 66% of the time but introduces new, evidence-unsupported propositions in a judge-raw 15% of changed responses, reinforcing the Renderer over the Fluency Ablation as the safe default.
This paper introduces a deterministic, zero-additional-cost evidence-provenance coloring tool as its practical, judge-free contribution, evaluated at full scale on 889 held-out RAGTruth responses with zero tiling failures. It provides an incident-level view of response provenance, turning whole-response review into targeted review. Across the two evaluated datasets, the proposed mechanism improved factual reliability while incurring a major-completeness cost of 2.3–2.8% among already-clean responses. The coloring tool is positioned as a reviewer aid rather than a certificate of factual correctness.

Author Contributions

Conceptualization, S.R. and D.S.; methodology, D.S.; software, S.R.; validation, D.S.; formal analysis, S.R.; investigation, S.R.; resources, S.R.; data curation, S.R.; writing—original draft preparation, S.R.; writing—review and editing, D.S.; visualization, S.R.; supervision, D.S.; project administration, S.R. All authors have read and agreed to the published version of the manuscript.

Funding

This research received no external funding.

Institutional Review Board Statement

Not applicable. This study did not involve human participants, human data, human tissue, or animals, and no clinical or interventional procedure was conducted. It is a computational analysis of two publicly available, previously de-identified benchmark datasets (RAGTruth and ExpertQA); the “human review” and “second-author endpoint verification” described in this paper refer to the authors’ own inspection of automated pipeline outputs against existing dataset annotations, not to research conducted on human subjects, and therefore did not require Institutional Review Board approval or exemption.

Data Availability Statement

RAGTruth is publicly available from its original authors [2], and ExpertQA is publicly available from its original authors [13]. The nine frozen replication files used in this study are publicly archived in Zenodo [29]. These comprise the eight pipeline scripts and the full-sample RAGTruth deterministic-coloring workbook enumerated in Appendix A.

Acknowledgments

During preparation of this manuscript, the authors used Claude (Anthropic; accessed August 2026) to assist with drafting and revising manuscript text. The authors reviewed and edited all outputs and take full responsibility for the content of this publication.

Conflicts of Interest

The authors declare no conflict of interest.

Abbreviations

The following abbreviations are used in this manuscript:
  • RAG: Retrieval-Augmented Generation
  • LLM: Large Language Model
  • QA: Question Answering
  • IRR: Inter-Rater Reliability
  • CI: Confidence Interval
  • CRediT: Contributor Roles Taxonomy
  • DAS: Data Availability Statement
  • DOI: Digital Object Identifier
  • GenAI: Generative Artificial Intelligence
  • IRB: Institutional Review Board
  • MAKE: Machine Learning and Knowledge Extraction (journal)
  • ORCID: Open Researcher and Contributor ID

Appendix A. Zenodo Deposit Contents

This appendix lists the nine files deposited with the paper: eight frozen pipeline scripts (Repair, Coverage Gate, Renderer, Fluency Ablation, and Baseline A, as applicable by dataset) and the full-sample RAGTruth deterministic-coloring workbook described in Section 4.2.4 and used in Section 6.7.
Table A1. Files included in the Zenodo deposit.
Table A1. Files included in the Zenodo deposit.
File Dataset Contents
ragtruth_repair.py RAGTruth Repair: semantic localization + deterministic factual realization (Section 4.2.1)
expertqa_repair.py ExpertQA Repair, schema-adapted for ExpertQA (Section 4.2.1)
ragtruth_renderer.py RAGTruth Renderer: deterministic rendering + completeness audit (Section 4.2.3)
expertqa_renderer.py ExpertQA Renderer, ExpertQA-side copy of the same script (Section 4.2.3)
ragtruth_cascade_coverage.py RAGTruth Coverage Gate: cascade coverage diagnostic (Section 4.2.2)
expertqa_cascade_coverage.py ExpertQA Coverage Gate, ExpertQA adaptation (Section 4.2.2)
ragtruth_grammar_polish_rendered.py RAGTruth Fluency Ablation: held-out grammar-only reopening comparison (Section 4.1 and Section 6.4)
ragtruth_baseline_a_regenerate.py RAGTruth Unconstrained single-shot regeneration ablation (Section 4.1, Section 5.8 and Section 6.6)
ragtruth_deterministic_colored_answers_and_confidence_tiers.xlsx RAGTruth Full-sample deterministic evidence-provenance coloring + confidence tiers (Section 4.2.4 and Section 6.7)

References

  1. Lewis, P.; Perez, E.; Piktus, A.; Petroni, F.; Karpukhin, V.; Goyal, N.; Küttler, H.; Lewis, M.; Yih, W.-t.; Rocktäschel, T.; Riedel, S.; Kiela, D. Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks. Adv. Neural Inf. Process. Syst. 2020, 33, 9459–9474. [Google Scholar]
  2. Niu, C.; Wu, Y.; Zhu, J.; Xu, S.; Shum, K.S.; Zhong, R.; Song, J.; Zhang, T. RAGTruth: A Hallucination Corpus for Developing Trustworthy Retrieval-Augmented Language Models. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics, Bangkok, Thailand, 2024; Volume 1, pp. 10862–10878. [Google Scholar] [CrossRef]
  3. Min, S.; Krishna, K.; Lyu, X.; Lewis, M.; Yih, W.-t.; Koh, P.W.; Iyyer, M.; Zettlemoyer, L.; Hajishirzi, H. FActScore: Fine-Grained Atomic Evaluation of Factual Precision in Long Form Text Generation. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, Singapore, 2023; pp. 12076–12100. [Google Scholar] [CrossRef]
  4. Asai, A.; Wu, Z.; Wang, Y.; Sil, A.; Hajishirzi, H. Self-RAG: Learning to Retrieve, Generate, and Critique through Self-Reflection. In Proceedings of the Twelfth International Conference on Learning Representations, 2024. [Google Scholar]
  5. Zhang, W.; Zhang, J. Hallucination Mitigation for Retrieval-Augmented Large Language Models: A Review. Mathematics 2025, 13, 856. [Google Scholar] [CrossRef]
  6. Lv, B.; Feng, A.; Xie, C. Improving Factuality by Contrastive Decoding with Factual and Hallucination Prompts. Sensors 2024, 24, 7097. [Google Scholar] [CrossRef] [PubMed]
  7. Hiriyanna, S.; Zhao, W. Multi-Layered Framework for LLM Hallucination Mitigation in High-Stakes Applications: A Tutorial. Computers 2025, 14, 332. [Google Scholar] [CrossRef]
  8. Qi, P.; Zhang, Y.; Zhang, Y.; Bolton, J.; Manning, C.D. Stanza: A Python Natural Language Processing Toolkit for Many Human Languages. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics: System Demonstrations, Online, 2020; pp. 101–108. [Google Scholar] [CrossRef]
  9. Bao, F.S.; Li, M.; Qu, R.; Luo, G.; Wan, E.; Tang, Y.; Fan, W.; Tamber, M.S.; Kazi, S.; Sourabh, V.; et al. FaithBench: A Diverse Hallucination Benchmark for Summarization by Modern LLMs. In Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies, Albuquerque, NM, USA, 2025; Volume 2, pp. 448–461. [Google Scholar] [CrossRef]
  10. Hu, X.; Ru, D.; Qiu, L.; Guo, Q.; Zhang, T.; Xu, Y.; Luo, Y.; Liu, P.; Zhang, Y.; Zhang, Z. Knowledge-Centric Hallucination Detection. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, Miami, FL, USA, 2024; pp. 6953–6975. [Google Scholar] [CrossRef]
  11. Ru, D.; Qiu, L.; Hu, X.; Zhang, T.; Shi, P.; Chang, S.; Cheng, J.; Wang, C.; Sun, S.; Li, H.; et al. RAGChecker: A Fine-Grained Framework for Diagnosing Retrieval-Augmented Generation. arXiv 2024, arXiv:2408.08067. [Google Scholar] [CrossRef]
  12. Stelmakh, I.; Luan, Y.; Dhingra, B.; Chang, M.-W. ASQA: Factoid Questions Meet Long-Form Answers. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, Abu Dhabi, United Arab Emirates, 2022; pp. 8273–8288. [Google Scholar] [CrossRef]
  13. Malaviya, C.; Lee, S.; Chen, S.; Sieber, E.; Yatskar, M.; Roth, D. ExpertQA: Expert-Curated Questions and Attributed Answers. In Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Mexico City, Mexico, 2024; Volume 1, pp. 3025–3045. [Google Scholar] [CrossRef]
  14. Ji, Z.; Lee, N.; Frieske, R.; Yu, T.; Su, D.; Xu, Y.; Ishii, E.; Bang, Y.; Chen, D.; Dai, W.; et al. Survey of Hallucination in Natural Language Generation. ACM Comput. Surv. 2023, 55, 248. [Google Scholar] [CrossRef]
  15. Zhang, Y.; Li, Y.; Cui, L.; Cai, D.; Liu, L.; Fu, T.; Huang, X.; Zhao, E.; Zhang, Y.; Chen, Y.; et al. Siren’s Song in the AI Ocean: A Survey on Hallucination in Large Language Models. arXiv 2023, arXiv:2309.01219. [Google Scholar] [CrossRef]
  16. Rashkin, H.; Nikolaev, V.; Lamm, M.; Aroyo, L.; Collins, M.; Das, D.; Petrov, S.; Tomar, G.S.; Turc, I.; Reitter, D. Measuring Attribution in Natural Language Generation Models. Comput. Linguist. 2023, 49, 777–840. [Google Scholar] [CrossRef]
  17. Bohnet, B.; Tran, V.Q.; Verga, P.; Aharoni, R.; Andor, D.; Soares, L.B.; Ciaramita, M.; Eisenstein, J.; Ganchev, K.; Herzig, J.; et al. Attributed Question Answering: Evaluation and Modeling for Attributed Large Language Models. arXiv 2022, arXiv:2212.08037. [Google Scholar] [CrossRef]
  18. Gao, L.; Dai, Z.; Pasupat, P.; Chen, A.; Chaganty, A.T.; Fan, Y.; Zhao, V.; Lao, N.; Lee, H.; Juan, D.-C.; Guu, K. RARR: Researching and Revising What Language Models Say, Using Language Models. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics, Toronto, Canada, 2023; Volume 1, pp. 16477–16508. [Google Scholar] [CrossRef]
  19. Kamath, A.; Jia, R.; Liang, P. Selective Question Answering under Domain Shift. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, Online, 2020; pp. 5684–5696. [Google Scholar] [CrossRef]
  20. Xin, J.; Tang, R.; Yu, Y.; Lin, J. The Art of Abstention: Selective Prediction and Error Regularization for Natural Language Processing. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing, Online, 2021; Volume 1, pp. 1040–1051. [Google Scholar] [CrossRef]
  21. Gao, Y.; Xiong, Y.; Gao, X.; Jia, K.; Pan, J.; Bi, Y.; Dai, Y.; Sun, J.; Wang, M.; Wang, H. Retrieval-Augmented Generation for Large Language Models: A Survey. arXiv 2023, arXiv:2312.10997. [Google Scholar] [CrossRef]
  22. Kadavath, S.; Conerly, T.; Askell, A.; Henighan, T.; Drain, D.; Perez, E.; Schiefer, N.; Hatfield-Dodds, Z.; DasSarma, N.; Tran-Johnson, E.; et al. Language Models (Mostly) Know What They Know. arXiv 2022, arXiv:2207.05221. [Google Scholar] [CrossRef]
  23. Maynez, J.; Narayan, S.; Bohnet, B.; McDonald, R. On Faithfulness and Factuality in Abstractive Summarization. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, Online, 2020; pp. 1906–1919. [Google Scholar] [CrossRef]
  24. Meng, K.; Bau, D.; Andonian, A.; Belinkov, Y. Locating and Editing Factual Associations in GPT. Adv. Neural Inf. Process. Syst. 2022, 35, 17359–17372. [Google Scholar] [CrossRef]
  25. Thorne, J.; Vlachos, A.; Christodoulopoulos, C.; Mittal, A. FEVER: A Large-Scale Dataset for Fact Extraction and VERification. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, New Orleans, LA, USA, 2018; Volume 1, pp. 809–819. [Google Scholar] [CrossRef]
  26. See, A.; Liu, P.J.; Manning, C.D. Get to the Point: Summarization with Pointer-Generator Networks. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics, Vancouver, Canada, 2017; Volume 1, pp. 1073–1083. [Google Scholar] [CrossRef]
  27. Es, S.; James, J.; Espinosa Anke, L.; Schockaert, S. RAGAs: Automated Evaluation of Retrieval Augmented Generation. In Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics: System Demonstrations, St. Julians, Malta, 2024; pp. 150–158. [Google Scholar] [CrossRef]
  28. Peng, B.; Galley, M.; He, P.; Cheng, H.; Xie, Y.; Hu, Y.; Huang, Q.; Liden, L.; Yu, Z.; Chen, W.; Gao, J. Check Your Facts and Try Again: Improving Large Language Models with External Knowledge and Automated Feedback. arXiv 2023, arXiv:2302.12813. [Google Scholar] [CrossRef]
  29. Rajendran, S.; Singaravelu, D. Evidence-Bound Factual Repair in Retrieval-Augmented LLM Answers: Code and Replication Materials. Zenodo 2026. [Google Scholar] [CrossRef]
  30. Gao, T.; Yen, H.; Yu, J.; Chen, D. Enabling Large Language Models to Generate Text with Citations. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, Singapore, 2023; pp. 6465–6488. [Google Scholar] [CrossRef]
  31. Huang, L.; Feng, X.; Ma, W.; Gu, Y.; Zhong, W.; Feng, X.; Yu, W.; Peng, W.; Tang, D.; Tu, D.; et al. Learning Fine-Grained Grounded Citations for Attributed Large Language Models. Find. Assoc. Comput. Linguist. ACL 2024 2024, 14095–14113. [Google Scholar] [CrossRef]
  32. Liu, Y.; Yang, T.; Huang, S.; Zhang, Z.; Huang, H.; Wei, F.; Deng, W.; Sun, F.; Zhang, Q. Calibrating LLM-Based Evaluator. In Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024), Torino, Italy, 2024; pp. 2638–2656. Available online: https://aclanthology.org/2024.lrec-main.237/.
Figure 1. A real RAGTruth coloring-tool output (response 12750, question “examples of instantaneous speed”), reproduced unedited from the pipeline’s own xlsx output. Dark green is clean-rendered (verbatim evidence, no flag); amber is flagged-rendered (verbatim evidence, completeness-risk flag); black is untouched original prose.
Figure 1. A real RAGTruth coloring-tool output (response 12750, question “examples of instantaneous speed”), reproduced unedited from the pipeline’s own xlsx output. Dark green is clean-rendered (verbatim evidence, no flag); amber is flagged-rendered (verbatim evidence, completeness-risk flag); black is untouched original prose.
Preprints 231785 g001
Table 1. Frozen evaluation populations used in Section 6.
Table 1. Frozen evaluation populations used in Section 6.
Population Dataset n Used for
Held-out cohort RAGTruth 889 Primary hallucination-reduction result (Section 6.1)
Held-out cohort ExpertQA 731 Primary grounding-reduction result (Section 6.1b)
Hallucination Stage-1 sample RAGTruth 132 (32 positive + 100 clean) Human-reviewed panel (Section 6.1); also, the Baseline A matched sample (Section 6.6)
Grounding Stage-1 sample ExpertQA 179 Human-reviewed panel (Section 6.1b)
Completeness evaluation populations both 857 RAGTruth + 602 ExpertQA Full descriptive completeness result (Section 6.5)
Renderer-vs-Fluency-Ablation sample RAGTruth 200 Grammar/semantic judge ablation (Section 6.4)
Second-author IRR sample both 144 RAGTruth + 173 eligible ExpertQA Rationale-blinded endpoint verification (Section 6.8)
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.