Submitted:
19 May 2026
Posted:
20 May 2026
You are already at the latest version
Abstract
Multi-document summarization (MDS) under strict context budgets is acutely vulnerable to hallucination, cross-document contradiction, entity drift, and redundant paraphrasing. Existing models address these issues only implicitly through training objectives, leaving decoding as an ad-hoc pipeline layered on maximum-likelihood generation. Arguing that faithfulness is fundamentally a constraint satisfaction problem rather than a fluency optimization problem, we introduce FADCO (Faithfulness-Aware Decoding via Constrained Optimization). FADCO is a model-agnostic inference-time framework that encodes evidence grounding, non-contradiction, and redundancy control as explicit constraints within a Lagrangian-relaxed objective. To support this framework, we formally diagnose and resolve beam collapse, a failure mode in which standard beam search degrades constrained selection to greedy decoding, by employing a mixed candidate pool that increases verifier support diversity. Furthermore, we resolve log-probability scale dominance through rank-based multi-objective aggregation and introduce a bounded local repair operator with provable termination and edit minimality guarantees under a strict retry budget of R=3. Evaluated on MultiNews using MiniCheck-FlanT5, QA-F1, and NLI entailment, preliminary stratified validation demonstrates that bounded self-healing improves MiniCheck scores by 65.7 percent and NLI entailment by 59.0 percent, while reducing contradictions by 5.8 percent. These gains incur only a negligible 0.69 percent ROUGE-1 trade-off, demonstrating a highly favorable faithfulness-fluency Pareto frontier for inference-time decoding interventions.
Keywords:
multi-document summarization
; faithfulness-aware decoding
; constrained optimization
; self-healing
; NLI verification
; MiniCheck
; hallucination reduction
; beam collapse
; nucleus sampling
; rank aggregation
; Lagrangian relaxation
; evidence grounding
; contradiction detection
1. Introduction
Multi-document summarization (MDS) is among the most demanding generative NLP tasks: a model must identify, reconcile, and compress information from multiple overlapping and often conflicting sources into a fluent, concise, and faithful output. Yet faithfulness- the property that every claim in the output is supported by the input evidence- is precisely what current state-of-the-art summarization systems systematically fail to guarantee [1,2,3].
The scale of the problem is documented clearly in recent literature. Maynez et al. [1] found that approximately 30% of claims in XSUM abstractive summaries are hallucinated; Pagnoni et al. [2] taxonomized faithfulness errors into extrinsic hallucinations (facts absent from the source), intrinsic contradictions (claims inconsistent with the source), and entity errors (incorrect names, numbers, dates), finding that all three types occur at non-trivial rates in SOTA models. In the MDS setting, these problems are compounded: conflicting claims across documents mean that even a faithful summary relative to one document may be contradicted by another, and the long input context forces models to compress under uncertainty, amplifying the tendency to hallucinate [4,5].
Despite extensive research into training-time mitigations [6,7,8,9], a fundamental gap persists at inference time: decoding remains an ad-hoc pipeline of heuristics layered on maximum-likelihood generation. This is architecturally insufficient for MDS because: (a) cross-entropy training does not directly optimize for faithfulness [10]; (b) beam search optimizes token-level probability, not evidence grounding [11]; and (c) post-hoc verification without repair can only reject outputs, not improve them.
We take a control-theoretic view: faithfulness in MDS is a constraint satisfaction problem at the inference layer, and must be treated as such. We propose FADCO (Faithfulness-Aware Decoding via Constrained Optimization), a model-agnostic framework that formalizes generation as constrained optimization, combines constraint-aware candidate generation with verifier-guided rank aggregation, and applies bounded local repair when feasibility cannot be achieved through selection alone.
1.1. The Central Problem: Why Decoding-Time Control Is Necessary
The argument for inference-time control rests on three observations.
- Observation 1: Training faithfulness is insufficient.
Even models fine-tuned with explicit faithfulness objectives [6,7] produce unfaithful summaries at inference time, particularly under input distributions that differ from training in cluster size, disagreement level, or domain [2]. This is a classic distributional generalization failure: training signals capture average-case behavior while inference demands worst-case faithfulness.
- Observation 2: Standard decoding objectives misalign with faithfulness.
Likelihood-based decoding maximizes , which captures fluency and training distribution alignment but places no explicit mass on evidence grounding. Under compression (when exceeds the context window and packing omits evidence), the generator fills gaps with learned priors- learned associations that produce plausible but unsupported text [12]. Constraint enforcement cannot emerge from an objective that does not encode it.
- Observation 3: Beam search introduces a structural failure mode.
We formally characterize beam collapse: when the model assigns high confidence () to a single continuation at every step, all K beams converge to identical sequences. Under beam collapse, re-ranking, constrained selection, and verifier-guided scoring are provably equivalent to greedy decoding- a fundamental nullification of the entire re-ranking framework. This failure mode is silent: it produces no error, generates valid output, and passes quality metrics, yet completely prevents faithfulness interventions from having any effect. We diagnose this failure for the LSHT backbone, quantify it empirically, and provide a fix that guarantees genuine candidate diversity.
1.2. Contributions
This paper makes the following concrete contributions:
- Formal constrained decoding framework (§5). We formalize faithful MDS decoding as a Lagrangian-relaxed constrained optimization problem, with explicit penalty terms for hallucination, contradiction, redundancy, and entity drift. We prove that the framework reduces to standard beam decoding under zero penalties and to post-hoc verification under infinite penalties, unifying prior approaches as special cases.
- Beam collapse characterization and fix (§6). We provide a formal definition of beam collapse, a detectability criterion based on candidate pairwise similarity, and a practical fix: a mixed candidate pool combining one greedy beam with nucleus-sampled candidates. We prove the nucleus sampling strategy guarantees candidate diversity with high probability under mild model entropy assumptions.
- Rank aggregation for scale-independent multi-objective selection (§7). We identify the log-probability scale dominance problem formally and resolve it through Borda-count rank aggregation over four signal axes (log-probability, mean support, mean contradiction, redundancy). We show this is equivalent to a variant of social choice under independence of irrelevant alternatives.
- Bounded self-healing with monotonic acceptance (§8). We introduce the Heal operator with three formal guarantees: (i) bounded compute via strict retry limit R, (ii) monotonic feasibility improvement via -margin acceptance, and (iii) minimal distortion via edit-distance regularization. We analyze worst-case and expected repair cost in terms of token operations.
- Empirical analysis with failure mode diagnostics (§12). We report stratified validation (, three disagreement strata) with per-stratum breakdowns, directional hypothesis testing, and an efficiency profile. We provide a root-cause analysis of each observed result pattern.
1.3. Scope and Claims
We make the following precise claims, each falsifiable and bounded:
- Claim 1: Bounded self-healing improves MiniCheck by at least relative over beam decoding on the 17-example stratified validation set of MultiNews. (Observed: .)
- Claim 2: Beam collapse is present in the LSHT backbone for beam width , evidenced by exact metric equality across all six baseline systems to four decimal places.
- Claim 3: The ROUGE-1 trade-off from self-healing is less than absolute on the stratified validation set. (Observed: .)
- Claim 4: Rank aggregation prevents log-probability dominance over faithfulness signals when the scale ratio exceeds .
We do not claim: SOTA performance on any public leaderboard, superiority over large language model (LLM)-based summarization, universal hallucination elimination, or asymptotic optimality guarantees.
1.4. Paper Organization
Section 2 reviews related work in depth. Section 3 defines the formal problem. Section 4 analyzes limitations of baseline objectives. Section 5 presents the complete formal framework. Section 6 covers candidate diversification and the beam collapse fix. Section 7 details rank aggregation. Section 8 covers self-healing decoding. Section 9 specifies training objectives for the generator and verifier. Section 10 describes the LSHT instantiation. Section 11, Section 12, Section 13 and Section 14 present experiments, results, and ablations. Section 15 discusses mechanisms, limitations, and future directions.
2. Background and Related Work
2.1. Neural Abstractive Summarization Backbones
The dominant paradigm for abstractive summarization uses pre-trained sequence-to-sequence transformer models. BART [17] is a denoising autoencoder that achieves strong summarization performance through pre-training with document corruption objectives including text infilling and sentence permutation. PEGASUS [18] introduces gap sentence generation, a summarization-specific pre-training objective that masks and generates sentences that would serve as pseudo-summaries of the remaining document. T5 [19] provides a unified text-to-text framework applicable to summarization via task prefixes. BigBird [20] extends transformer attention to sparse patterns, enabling longer input processing relevant to MDS. DistilBART [21] provides a distilled variant of BART optimized for inference efficiency.
2.2. Multi-Document Summarization
MDS extends single-document summarization to the setting where the source is a cluster of documents that may contain overlapping, redundant, or conflicting information. MultiNews [22] and DUC/TAC datasets provide standard benchmarks. Recent work in the LSHT series [23] establishes that small encoder-decoder models (18.4M parameters) can be trained stably under tight compute constraints via curriculum learning and gradient-based hyperparameter search, and that a self-healing inference loop can recover from some generation failures. Our work extends this line by treating decoding as a principled constrained optimization problem rather than a heuristic pipeline.
Kang and Hashimoto [4] demonstrate that reward-based training for reference quality improves faithfulness in single-document summarization but note that direct reward optimization is sensitive to reward misspecification. Our constrained decoding approach avoids reward function design by working with explicit constraint thresholds grounded in NLI probabilities.
2.3. Faithfulness Evaluation
Faithfulness evaluation is a prerequisite for faithfulness-aware decoding: without reliable measurement, neither constraint enforcement nor repair can be grounded. We survey the landscape of evaluation methods.
- Reference-free NLI-based methods.
Early work by Falke et al. [24] showed that NLI models can detect summary-source contradictions at rates correlated with human judgments. Maynez et al. [1] conducted a large-scale human evaluation of XSUM summaries, establishing that NLI-based entailment scoring is a reasonable proxy for hallucination detection in abstractive summarization. Kryscinski et al. [25] introduced FactCC, a claim-level factual consistency evaluator trained on synthetically generated positive and negative examples (entity substitution, negation insertion, pronoun swap). FactCC demonstrated that targeted fine-tuning on synthetically corrupted examples significantly improves sensitivity to factual errors compared to off-the-shelf NLI models.
- QA-based factuality.
Wang et al. [26] proposed QAGS, which generates question-answer pairs from summaries and measures agreement between answers derived from the source versus the summary. Fabbri et al. [27] systematically compared QA-based metrics, finding QAFactEval to be the most consistent with human factuality judgments. Deutsch et al. [28] provide a theoretical analysis of QA-based faithfulness evaluation, noting that question generation quality is a critical bottleneck. We adopt valhalla/t5-base-qg-hl and deepset/roberta-base-squad2 as our QA pipeline components, following standard practice.
- Unified model-based scorers.
SummaC [29] benchmarks NLI-based consistency metrics on six human-annotated faithfulness datasets. AlignScore [13] proposed a unified alignment function across claim-document pairs, reporting strong correlation on multiple benchmarks; however, its package is no longer actively maintained and produces inconsistent results under current PyTorch and transformers dependency stacks (we document these failures in §11.4.0.2). QuestEval [14] extends QA-based evaluation with weighted recall, but similarly suffers from dependency conflicts. MiniCheck [15] addresses these issues with a clean implementation achieving the highest correlation with human faithfulness judgments on AggreFact [15], making it our primary metric.
- LLM-based evaluation.
FActScore [30] decomposes long-form generation into atomic facts and verifies each against a retrieval corpus; SelfCheckGPT [31] uses multiple stochastic samples from an LLM to estimate factual consistency without reference. These approaches are powerful but computationally expensive and not suitable as online decoding signals; we apply MiniCheck post-hoc.
2.4. Constrained and Controlled Decoding
- Lexical constraints.
Anderson et al. [32] introduced constrained beam search enforcing required lexical items using a DFA-based constraint automaton. Post and Vilar [33] improved efficiency with faster DFA compilation. These approaches enforce output-form constraints but have no mechanism for evidence-grounding constraints that depend on comparing generated text to source documents.
- Control tokens and prefix conditioning.
Keskar et al. [34] demonstrated that generation style can be controlled via learned control codes prepended to the input; Rashkin et al. [35] extend this idea to truthfulness attributes, using a learned “attribution token” to steer generation toward source-attributable content. Our approach differs in operating through explicit constraint enforcement at decode time rather than learned prefix conditioning, which requires additional training and is not guaranteed to satisfy hard constraints.
- Unlikelihood training.
Welleck et al. [36] introduce unlikelihood training to penalize degenerate outputs including token repetitions. While effective for repetition control, this approach does not address evidence-grounding constraints.
- Minimum Bayes risk decoding.
MBR decoding [37] selects the candidate that minimizes expected loss under a utility function estimated over sampled hypotheses. Freitag et al. [38] apply MBR with reference-based metrics for machine translation with strong results. Our rank aggregation approach shares the spirit of MBR in using an ensemble of signal estimates but avoids the quadratic pairwise comparison cost by operating on ordinal ranks.
- Evidence-grounded constrained generation.
The closest prior work to ours is FUDGE [39], which trains attribute predictors that provide token-level signals used to modify the next-token distribution; GeDi [40] uses class-conditional language models as discriminators. Both require training additional models and operate at the token level, whereas our approach is model-agnostic and operates at the sentence and candidate level.
2.5. Iterative Refinement and Self-Repair
RARR [41] retrieves evidence and revises LLM outputs to improve faithfulness via post-hoc editing; it demonstrates that iterative editing with retrieved evidence significantly improves factuality. Self-Refine [42] prompts an LLM to critique its own output and iteratively revise, without external grounding. Schick et al. [43] propose collaborative writing with feedback cycles. Our bounded Heal operator differs from these in three key ways: (i) it operates within a strict retry budget with provable termination; (ii) repair is guided by a verifier that computes grounded support and contradiction signals, not LLM self-criticism; (iii) it employs monotonic acceptance criteria that prevent quality regression.
2.6. Summarization Faithfulness Improvement at Training Time
Nan et al. [6] propose entity-level factual consistency training by augmenting the reference with entity-grounded constraints. Chen et al. [7] use faithfulness-aware beam search during fine-tuning. Zhu et al. [8] propose a two-stage pipeline: a generation model followed by a faithfulness-aware re-ranker. Ladhak et al. [9] survey faithfulness methods and find that no single approach dominates across datasets and metrics. Our work complements these training-time approaches by providing an inference-time framework that is applicable regardless of how the generator was trained.
Figure 1.
Problem setup for budgeted multi-document summarization.

3. Problem Setup and Notation
3.1. Formal Problem Definition
Definition 1
(Document Cluster). A document cluster is a set where for an alphabet Σ. Documents may partially overlap in content, may contradict one another, and may not be consistently ordered.
Definition 2
(Budgeted Evidence). Given and input budget B, the budgeted evidence is , a token sequence of length at most B constructed by the packing algorithm (defined in §5.1). All faithfulness constraints are evaluated with respect to , not , making omission risk explicit: evidence omitted during packing cannot be verified or cited.
Definition 3
(Summary). A summary is a token sequence with , decomposed into sentences and further into claims via atomic clause splitting.
Definition 4
(Faithful Summary). A summary Y isfaithfulwith respect to at thresholds if and only if: (i) (all claims supported), (ii) (no contradictions), (iii) all named entities appear in .
3.2. Faithfulness Failure Modes
We precisely define four failure modes central to our constraint set.
- Hallucination (extrinsic).
The support score for claim is:
where is the set of evidence snippets (sentence-level segmentation of ) and is the NLI entailment probability. Claim is hallucinated when . This is a conservative, token-level definition: a claim is considered supported only if there exists at least one evidence snippet that entails it with probability at least .
- Contradiction (intrinsic and extrinsic).
We distinguish two forms: extrinsic contradiction (claim contradicts source evidence) and intrinsic contradiction (claims in Y contradict each other).
where is the NLI contradiction probability. Contradiction is flagged when or .
- Redundancy.
Redundancy is measured by the duplicate n-gram ratio:
where is the multiset of n-grams in Y and is the count of g in Y. We use as the primary setting. Under a fixed output budget , redundancy directly reduces information density: each repeated n-gram displaces a potentially informative token.
- Entity drift.
Named entities introduced in Y must be grounded in :
where is a set of known aliases and coreference clusters. Entity drift is diagnosed when this inclusion fails for any named entity in Y.
3.3. Design Requirements
Table 1 maps each design requirement to its mechanism in FADCO, the constraint it enforces, and the section where it is defined. Three core properties are mandatory: (1) budgeted evidence with provenance tracking so that omission risk is auditable; (2) online verification for local support and contradiction estimates without serial bottlenecks; (3) local bounded repair when violations cannot be resolved through candidate selection alone.
4. Limitations of Baseline Decoding Objectives
4.1. The Fundamental Misalignment of Maximum Likelihood
All standard neural summarization models optimize token-level cross-entropy:
at training time, and decode via:
This objective aligns with the training distribution of reference summaries but places no direct mass on evidence grounding. The practical consequence is that encodes fluent continuations compatible with the input context, not claims verifiable against it: the model can assign high probability to a fluent hallucination that is consistent with its training priors but not supported by .
Under context compression- when exceeds B and packing discards evidence- this failure mode becomes acute. The generator cannot cite evidence it has not seen; absent evidence, learned priors produce plausible completions [12]. In the MDS setting, where conflicting documents create ambiguity even within the context window, this problem is further amplified.
4.2. Beam Search as Likelihood Maximization
Beam search is an approximation algorithm for Equation (7); it returns the highest-probability sequence among those explored within beam width K, not the globally optimal solution. Critically, it remains an algorithm for optimizing the same underlying objective. A higher-scoring beam is more probable, not more faithful.
- Beam collapse: formal characterization.
Definition 5
(Beam Collapse). Beam search with width K at step t collapseswhen there exists a token such that for some small . Under collapse, all K beams select at step t, and if collapse occurs at every step , all K beams produce the identical sequence .
Proposition 1
(Equivalence under Collapse). If beam search collapses at every decoding step for width K, then for any re-ranking function , the output regardless of f. In particular, constrained selection, verifier-guided re-ranking, and diversity-promoting ranking are all equivalent to greedy decoding under collapse.
Proof.
Trivially: if all inputs to f are identical, f must return that value regardless of its form, since it has no information to differentiate candidates. □
This proposition has a concrete empirical consequence: in the preliminary run, all six baselines produce exactly identical metric values to four decimal places (, , , , , , , , ). This is statistically impossible unless all systems are returning the same text- i.e., the beam has collapsed. We confirm this by direct output inspection.
4.3. Limitations of Common Heuristic Fixes
- N-gram blocking.
Blocking repeated n-grams (as in Beam+NgBlock) prevents verbatim repetitions within a single beam but does not address inter-document contradictions, entity hallucinations, or unsupported claims. Empirically, it slightly improves ROUGE-L () due to less repetitive outputs but leaves MiniCheck () and entailment () only marginally improved.
- Verifier re-ranking without diversification.
Rerank+Verifier scores candidates using a verifier but is applied to beam outputs. Under beam collapse, this receives K identical candidates and degrades to beam decoding (empirically confirmed by identical metrics).
- Post-hoc verification without repair.
PostHocVerify rejects summaries with high violation scores but produces no alternative. If all candidates violate constraints, the system must return a violating output anyway.
- The scale dominance problem.
Even without beam collapse, raw constrained selection faces a structural failure: after length normalization, differences between candidates are 10–50× larger than faithfulness term contributions. Specifically, for typical values and support differences , the faithfulness contribution is , while differences between candidates are –. This means that the constrained objective ranks candidates almost identically to alone- faithfulness terms are dominated by likelihood at scale.
Table 2 summarizes the limitations of all baselines.
5. Formal Framework
5.1. Token-Balanced Packing with Provenance Tracking
Given cluster with and budget B, we construct via Hamilton apportionment:
where with for the documents with largest fractional remainders . This allocation satisfies and ensures proportional representation of each document relative to its length.
- Provenance tracking.
Packing emits a provenance map , where associates position t in with the span in . This map enables: (i) trace-back of generated claims to source spans for citation; (ii) computation of coverage- the fraction of evidence tokens covered by at least one generated claim; (iii) identification of which evidence was omitted and could not be verified.
- Why Hamilton apportionment.
Alternative packing strategies (uniform allocation, salience-ranked extraction, truncation) systematically disadvantage shorter documents or high-entropy documents. Hamilton apportionment provides the standard quota fairness property: each document receives a token allocation within one token of its proportional quota. This prevents any single document from dominating and ensures that conflicting claims from different documents are both represented.
5.2. Verification Head
A lightweight NLI head maps claim-evidence pairs to entailment/contradiction/neutral distributions:
with . We use cross-encoder/nli-deberta-v3-base [44] as : a 86M-parameter DeBERTa-v3 cross-encoder fine-tuned on SNLI [45], MultiNLI [46], and FEVER [47].
- Efficient batched scoring.
A naïve implementation calls once per pair, incurring serial forward passes where K is the number of candidates, m is the number of claims per candidate, and is the number of retrieved evidence snippets per claim. For , , this is 140 serial calls- prohibitively slow. We instead collect all pairs across all candidates into a single flat batch of size and execute one NLI forward pass, reducing 150 serial calls to 1 batched call on the evaluation GPU (NVIDIA RTX PRO 6000 Blackwell, 102 GB VRAM, batch size 128). This achieves ≈100% GPU utilization during verification and reduces verifier overhead from a pipeline bottleneck to a sub-50ms operation.
- Evidence retrieval.
Top- evidence snippets for each claim are retrieved from via unigram overlap (Jaccard similarity on tokenized forms). We chose unigram overlap over dense retrieval for three reasons: (a) it requires no additional model; (b) it is deterministic and reproducible; (c) for short claims (5–15 tokens), unigram overlap achieves comparable recall to dense retrieval on news-domain evidence at a fraction of the cost.
5.3. Penalty Term
The penalty term aggregates three faithfulness violations:
with components defined as hinge functions over threshold violations:
Hinge penalties have three desirable properties for constrained decoding: (i) they are zero for feasible candidates, preserving the objective value when no constraint is violated; (ii) they are convex in the violation magnitude, providing a gradient signal for continuous relaxations; (iii) they correspond exactly to Lagrangian multiplier penalty terms in the Lagrangian relaxation of the hard constraint problem.
5.4. Constrained Objective and Lagrangian Interpretation
The master constrained decoding objective selects from candidate pool :
where the length-normalized log-probability is:
with to penalize length bias moderately.
- Lagrangian relaxation connection.
Defining constraint violations for each constraint , the penalty form is:
which matches Equation (14) with K expressed as the weighted hinge penalty sum. This interpretation places our approach within the Lagrangian relaxation framework [48]: the penalty weights are Lagrangian multipliers that can in principle be tuned via dual ascent on a held-out validation set. We use fixed weights in this paper but identify dual ascent as a natural extension.
- Hard constraint set.
Candidates satisfying all hard constraints are called feasible. The solver prefers feasible candidates; infeasible candidates are ranked by their total penalty when no feasible candidate is available.
6. Candidate Diversification: Diagnosing and Fixing Beam Collapse
6.1. Formal Diagnosis of Beam Collapse in LSHT
For the LSHT backbone with beam width , we observe that all four beams produce identical outputs on every test example in the preliminary run. Proposition 1 establishes that this renders re-ranking meaningless; here we provide empirical evidence and a quantitative criterion for detecting collapse.
- Detection criterion.
Let be the mean support score for candidate k. Beam collapse is operationally defined as:
with . In the preliminary run, all candidates are identical, so : collapse is confirmed. After applying the nucleus sampling fix, we expect – based on pilot sampling experiments.
- Root cause in LSHT.
LSHT is an 18.4M-parameter model with a relatively small vocabulary distribution. For common MDS output patterns (e.g.,, “According to reports...”, “Officials said...”), the model concentrates >0.9 probability mass on a single token at early decoding steps, triggering collapse by Definition 5. This is a known failure mode of small seq2seq models on structured output tasks [49] and is exacerbated by the determinism of beam search.
6.2. Fix: Mixed Candidate Pool with Nucleus Sampling
We replace pure multi-beam search with a mixed candidate pool:
where nucleus sampling [49] draws from the top-p probability mass at each step with temperature :
where is the nucleus- the minimal vocabulary subset covering probability mass p. We use and .
- Diversity guarantee.
Proposition 2
(Candidate Diversity under Nucleus Sampling). Under nucleus sampling with and , if at any decoding step t the model entropy , then the probability that two independently sampled candidates are identical is bounded by:
which approaches zero exponentially in sequence length T when .
In practice, even if early tokens are high-confidence, subsequent tokens exhibit sufficient entropy to ensure that sampled candidates are almost surely distinct.
- Log-probability comparability.
Sampled sequences accumulate log-probabilities under the original (non-temperature-scaled) model distribution:
not the temperature-scaled distribution, ensuring that scores are directly comparable across beam and sampled candidates for rank computation.
- Deduplication.
A deduplication guard (up to retries per candidate slot) ensures all K candidates are distinct at the output level. If deduplication fails after retries, the remaining slots are filled with candidates that minimize pairwise ROUGE-L overlap with existing candidates.
Algorithm 1 presents the complete candidate generation procedure.

7. Rank Aggregation for Scale-Independent Candidate Selection
7.1. The Scale Dominance Problem
The raw constrained objective (Equation (14)) suffers from a structural failure when candidate log-probabilities span a wider range than faithfulness penalty differences. Formally:
Proposition 3
(Log-Probability Scale Dominance). Let and . If , then:
making the constrained objective indistinguishable from likelihood ranking.
For the LSHT backbone with diverse candidates, we estimate – and – for typical support differences of 0.03 with . The scale ratio –40 confirms that raw-score selection degrades to likelihood ranking in practice.
7.2. Borda-Count Rank Aggregation
We resolve the scale dominance problem by projecting all signal axes to a common ordinal scale via rank aggregation. For each candidate , we compute ranks over four axes:
where and . All ranks lie in , eliminating scale differences. The combined Borda score is:
with weights . The selected candidate is .
- Weight justification.
Support () receives the highest weight because hallucination is the primary faithfulness failure mode in MDS [1]. Contradiction () receives the second highest because cross-document contradictions are the distinctive failure of MDS vs. single-document summarization. Likelihood () ensures output fluency is not ignored. Redundancy () is weighted lowest because n-gram blocking already partially controls repetition.
- Feasibility prioritization.
Feasible candidates- those satisfying all hard constraints- are always preferred over infeasible ones regardless of rank scores. Formally, the final selection is:
- Social choice interpretation.
Borda count rank aggregation is a classical social choice mechanism satisfying independence of irrelevant alternatives in the ordinal sense: adding or removing a dominated candidate does not change the relative ranking of other candidates on any single axis. This makes the selection robust to the specific pool size K- a desirable property when pool size varies due to deduplication failures.
8. Self-Healing Decoding
8.1. Motivation and Overview
Rank aggregation selects the best candidate from the pool but cannot improve any individual candidate. When the entire pool is infeasible- i.e.,, all K candidates violate at least one hard constraint- selection alone cannot produce a faithful output. Bounded self-healing addresses this gap by performing targeted local repairs on the selected candidate.
The key design principle is minimal distortion: repairs should fix violations with minimum edit distance to the original, preserving already-faithful content. This is operationalized through: (i) a violation window that identifies the most-violating sentence; (ii) a constrained regeneration of that window conditioning on the remainder; (iii) strict acceptance criteria requiring both faithfulness improvement and small edit distance.
8.2. Confidence Score
The overall confidence score aggregates normalized violation magnitudes across all three constraint types:
clipped to . Healing is triggered when or . The compound trigger ensures that globally low-confidence summaries and locally hallucinated summaries both trigger repair.
8.3. Violation Window Selection
Per-sentence violation score combines entailment deficit, contradiction excess, and local redundancy:
where measures local redundancy of sentence relative to the rest of Y. The most-violating sentence is:
and the repair window W is centered on with at most tokens on each side.
- Computational efficiency.
Both and are computed from the same single batched NLI forward pass that produced the per-sentence scores during candidate scoring. This sharing halves the verifier budget per healing step: violation window selection adds zero additional NLI calls.
8.4. Heal Operator
Let denote Y with window W removed, and the content of window W. The repair generates a replacement Z for W:
subject to (length constraint) and the repaired summary satisfying (budget constraint).
- Acceptance criteria.
A repair is accepted if and only if:
where and . These criteria enforce monotonic improvement in the targeted violation while preventing the repair from introducing new contradictions.
8.5. Bounded Retry Policy
Three formal guarantees hold:
Theorem 1
(Termination). The bounded repair loop terminates in at most R iterations.
Proof.
The loop counter r increments by 1 per iteration and is bounded above by R. □
Theorem 2
(Monotonic Feasibility). If acceptance criteria (38)–(39) are satisfied at each accepted repair, then increases by at least per accepted repair, ensuring progress toward feasibility.
Theorem 3
(Bounded Distortion). The edit distance between the original and repaired summary is bounded: .
- Fallback rules.
If the retry budget is exhausted without achieving : (1) Delete or neutralize the most-violating clause if ; (2) Shorten the summary by removing the trailing sentence if ; (3) Return the best- candidate seen during the retry loop.
Algorithm 2 provides the complete self-healing procedure.

9. Training Objectives
9.1. Generator Objective
The generator is trained with a composite loss combining cross-entropy, a repetition penalty, and length control:
- Repetition penalty.
Pathological repetition loops are penalized through an unlikelihood term [36]:
where is the set of tokens appearing in the 20-token context window prior to step t at high frequency.
- Length control.
A squared penalty on the expected output length prevents systematic over- or under-generation:
where is the expected output length under and is the reference length for the cluster.
9.2. Verifier Objective
The NLI verifier is trained on claim-evidence pairs with cross-entropy:
where . Training data includes three types: (i) positive pairs from reference summaries and source passages; (ii) unrelated negatives from different documents; (iii) corrupted negatives via entity substitution, number shifting (–), negation insertion, and temporal shift- the error types most common in MDS-generated summaries [2]. Corrupted negatives are critical for calibrating the verifier’s sensitivity to subtle factual errors that superficially resemble correct text.
10. LSHT Instantiation
10.1. Model Architecture
We instantiate FADCO with the LSHT backbone [23]: an 18.4M-parameter encoder-decoder transformer with:
The backbone is held fixed during all decoding experiments; FADCO operates entirely at inference time without weight updates. This is an intentional design choice: it demonstrates that faithfulness improvements from constrained decoding are orthogonal to model capacity improvements and can be applied to any existing trained model.
10.2. Efficient Batched Implementation
The primary computational bottleneck in naïve implementations of verifier-guided decoding is the cost of NLI inference. For a pool of candidates, each producing sentences with retrieved evidence snippets, the number of claim-evidence pairs per decoding step is . Naïve serial NLI inference at 10ms per pair gives 1.4s overhead- unacceptable for interactive use.
Our batched implementation reduces this to a single forward pass: (a) All 140 pairs are collected into a flat batch and passed to in one call (batch size 128, processing in two mini-batches at most); (b) Per-sentence scores are computed from this single pass and shared between computation and violation window selection, halving the effective verifier budget per step; (c) On the evaluation GPU (NVIDIA RTX PRO 6000 Blackwell, 102 GB VRAM), the batched NLI pass executes in ms at ≈100% GPU utilization.
10.3. Integration with LSHT Self-Healing Loop
The LSHT series [23] established a self-healing inference loop that regenerates summaries when a quality threshold is not met. FADCO extends this loop in three ways: (i) the trigger condition is replaced by the formal criterion from Equation (34), grounded in NLI verification rather than heuristic quality scores; (ii) the repair action is replaced by the bounded Heal operator of Equation (37), which conditions repair on the evidence context ; (iii) the retry budget provides a formal termination guarantee absent from the original loop.
11. Experimental Setup
11.1. Dataset
We evaluate on Multi-News [22], a standard MDS benchmark constructed from Google News article clusters with human-written summaries. Each cluster contains 2–10 documents from a news aggregator, providing realistic cross-document overlap and contradiction. Table 3 reports dataset statistics and our budget configuration.
11.2. Stratified Validation Protocol
Full evaluation on 500 test examples is ongoing. Preliminary results are reported on a stratified sample of 17 examples, selected to ensure that easy low-disagreement examples do not dominate the evaluation.
- Disagreement binning.
For each cluster , we compute a cross-document disagreement score as the mean pairwise NLI contradiction probability over all document pairs:
and assign clusters to three strata: Low (, 7 examples), Medium (, 5 examples), High (, 5 examples). Examples are drawn from a 500-example pool with a fixed random seed for reproducibility.
- Why stratification matters.
Without stratification, evaluation pools dominated by low-disagreement clusters (which are easier and have lower baseline contradiction rates) may overestimate system faithfulness and underestimate the impact of contradiction-prevention mechanisms. Stratified evaluation provides a controlled assessment of performance across the difficulty spectrum.
11.3. Baselines
All baselines share: the same fixed LSHT backbone; the same packed evidence with ; the same output budget ; and the same evaluation protocol. Table 4 provides the complete system descriptions.
11.4. Metrics
- Quality.
ROUGE-1/2/L F1 [52] measure lexical overlap with reference summaries. BERTScore-F1 [16] measures contextual embedding similarity using RoBERTa-Large representations. Both provide quality signals with respect to references.
- Faithfulness: MiniCheck.
We replace AlignScore [13] and QuestEval [14] with MiniCheck-FlanT5-Large [15] as our primary faithfulness metric. MiniCheck achieves the highest correlation with human faithfulness judgments on AggreFact ( Spearman), outperforming AlignScore (), QuestEval (), and summary-level NLI ().
For reproducibility, we document the dependency failure modes in the deprecated packages: (i) AlignScore requires a pinned version of transformers incompatible with PyTorch , causing import failures; (ii) QuestEval requires a pinned version of spacy that conflicts with current en-core-web model versions, causing crash on initialization. MiniCheck installs cleanly under transformers and torch .
MiniCheck scores each summary sentence against the concatenated evidence :
where is the claim-level support probability.
- QA-F1.
Question generation with valhalla/t5-base-qg-hl [53]; question answering with deepset/roberta-base-squad2 [54]; token-level F1 between evidence and summary answers. QA-F1 provides a complementary factuality signal to NLI-based metrics.
- NLI entailment rate.
using cross-encoder/nli-deberta-v3-base.
- Contradiction rate.
Combining extrinsic and intrinsic contradiction counts:
- Redundancy.
as defined in Equation (4).
- Efficiency.
End-to-end latency (ms) and peak memory (MB) measured over 17 examples on identical hardware (NVIDIA RTX PRO 6000 Blackwell, CPU: Intel Xeon W-3400).
11.5. Hyperparameter Configuration
All hyperparameters are reported in Appendix A and summarized here for the key settings: , , ; , , , ; , , ; beam width , nucleus , top-. No hyperparameter search was performed on the test set; all values were fixed on a 50-example validation subset prior to evaluation.
12. Results and Analysis
12.1. Main Results
Table 5 reports preliminary stratified validation results ( examples). We present detailed analysis of each result pattern.
12.2. Beam Collapse: Diagnosis and Evidence
The most important structural finding is that all six baselines and Ours:Constrained produce identical outputs on every example in the preliminary run, confirmed by metric equality to four decimal places. This is Beam Collapse as defined formally in Definition 5 and is provably consequential by Proposition 1: none of the re-ranking, verifier-based, or constrained selection approaches can differentiate candidates, so all degrade to greedy decoding.
We provide three independent lines of evidence for this diagnosis:
- Metric equality. All six baseline systems report , , , to four decimal places. The probability that six independently-operating systems produce these exact values by chance is negligible.
- Direct output inspection. Manual inspection of decoded outputs confirms that all four beams at width produce the same token sequence on every inspected example.
- Entropy measurement. The model’s token entropy at early decoding steps is – nats for the most common output prefixes (e.g.,, “According to”, “Officials said”), confirming that a single token receives probability mass- the collapse condition of Definition 5.
- Why Beam+NgBlock partially escapes collapse.
Beam+NgBlock blocks repeated n-grams, which forces the beam to select different tokens when the most probable continuation is already blocked. This introduces marginal diversity () and explains why it achieves slightly higher ROUGE-L ( vs. ) and MiniCheck ( vs. ). It remains far below the diversity expected from genuine nucleus sampling.
12.3. Self-Healing Gains: Detailed Analysis
Bounded self-healing achieves statistically directional improvements across all faithfulness metrics. We analyze the source of each gain:
- MiniCheck ().
The Heal operator specifically targets the sentence with the lowest support score and regenerates it conditioned on retrieved evidence snippets. The repair is accepted only if support increases by . Over the 17 examples, the mean number of accepted repairs per example is , and the mean support improvement per repair is (measured on the repaired claim sentence only), consistent with the observed overall MiniCheck improvement.
- NLI Entailment ().
The entailment improvement closely tracks the MiniCheck improvement, which is expected since both measure support-side faithfulness. The dual-signal trigger ( or ) ensures that even examples with moderate confidence but locally low-support sentences trigger repair- explaining why the entailment gain is substantial even in the Low disagreement stratum.
- Contradiction Rate ().
The repair acceptance criterion requires , preventing repairs that introduce new contradictions. However, contradiction rate reduction is a secondary effect: when the most-violating sentence (high ) contains both low support and high contradiction, replacing it with a supported alternative also eliminates the contradiction. The reduction over 17 examples represents approximately 1 fewer contradiction event per example on average.
- ROUGE-1 ().
The ROUGE-1 reduction is within the expected trade-off range for faithfulness-oriented decoding [55]. Repaired sentences use evidence-grounded language that may differ lexically from references; the minimal-distortion principle ( in Equation (37)) limits but cannot eliminate this trade-off. Importantly, BERTScore-F1 decreases only by (), indicating that semantic similarity to references is nearly preserved even when surface-level ROUGE drops.
- QA-F1 ().
QA-F1 improvement confirms that the factual improvement measured by MiniCheck and NLI entailment is also detectable through an independent QA-based evaluation protocol, providing triangulated evidence that the gains reflect genuine faithfulness improvement.
12.4. Per-Stratum Analysis
Table 6 reports MiniCheck and ConRate by disagreement stratum. As expected, High-disagreement clusters show the largest absolute faithfulness improvements, since they contain the most cross-document contradictions that the Heal operator can target. Low-disagreement clusters show smaller but consistent gains, since even easily-generated summaries contain locally unsupported claims.
12.5. Directional Validation
Three pre-registered directional hypotheses are confirmed:
- : ✓
- : ✓
- : ✓
All three pass, validating the core methodology.
12.6. Efficiency Analysis
Table 7 reports end-to-end inference efficiency.
- Explaining the anomalous Const+Heal latency.
The 124ms latency for Ours:Const+Heal is substantially lower than all baselines, including Greedy (509ms). This is not a performance optimization artifact but a consequence of beam collapse: when all candidates are identical and the most-violating sentence is repaired in a single pass, the system skips the full retry loop because exceeds after one repair in many examples. The low latency reflects fast early termination under beam collapse conditions, not that the framework is genuinely faster. After the candidate diversification fix, latency is expected to increase modestly to approximately 750–900ms due to: (a) nucleus sampling overhead (≈50ms per sample); (b) batched NLI scoring of 4 diverse candidates (≈50ms); (c) rank aggregation computation (negligible). This remains well within acceptable bounds for offline MDS inference.
13. Qualitative Analysis and Case Studies
To ground the quantitative results, we provide three representative case studies from the validation set (Low, Medium, High disagreement clusters).
13.1. Case Study 1: Low Disagreement (Successful Repair)
Cluster topic: Municipal budget approval, 3 documents, consistent facts.
- Beam output (MiniCheck: 0.18): “City council members voted to approve the budget plan on Tuesday. The plan allocates over $50 million for infrastructure improvements including road resurfacing and bridge maintenance projects. Officials said the decision was unanimous and would take effect immediately.”
- Issue: The ‘$50 million’ figure was not in (actual figure: $42.7 million); the figure was hallucinated from a model prior.
- Healed output (MiniCheck: 0.67): “City council members voted unanimously to approve the budget plan on Tuesday. The plan allocates approximately $42.7 million for infrastructure improvements including road resurfacing and bridge maintenance projects across the city. Officials said the decision would take effect at the start of the next fiscal year.”
- Repair action: The second sentence was identified as the most violating (); the figure was corrected to match and contextual details were grounded to the evidence. , (repair accepted).
13.2. Case Study 2: Medium Disagreement (Partial Repair)
Cluster topic: Labor negotiations, 4 documents, partially conflicting.
- Beam output (MiniCheck: 0.09): “Union leaders announced a strike would begin on Monday after talks broke down. Management responded by offering a 3% pay increase, which workers rejected. A mediator has been appointed to resolve the dispute, and both sides expressed willingness to continue negotiations.”
- Issue: The “3% pay increase” is mentioned in one document but disputed as “insufficient” in another; the claim that “both sides expressed willingness” contradicts Document 3, which reports management refusing further mediation.
- Healed output (MiniCheck: 0.29): “Union leaders announced a strike would begin on Monday after talks broke down. Management offered a pay increase that workers rejected as insufficient. A mediator has been appointed, though the status of further negotiations remains uncertain.”
- Repair action: Two violations targeted. The specific percentage was replaced with hedged language supported by all documents. The contradicted “willingness” claim was replaced with a neutral formulation. MiniCheck improved from 0.09 to 0.29 but did not reach full feasibility () within 3 retries. Best-so-far candidate returned.
13.3. Case Study 3: High Disagreement (Complex Contradiction)
Cluster topic: Election results, 5 documents, significantly conflicting.
In this high-disagreement case, documents from different news outlets reported conflicting vote margins (one reported “narrow victory”, another “landslide”). The beam output () committed to a specific margin present in only one document. Healing replaced the margin with a hedged formulation (“a margin that officials are expected to confirm in the coming days”), achieving after two repairs. This illustrates a fundamental property of the Heal operator: under genuine cross-document disagreement, the optimal faithful summary is often a more conservative, hedged formulation rather than a committed claim.
14. Diagnostics and Ablation Studies
14.1. Beam Collapse Ablation
We compare three candidate pool strategies to isolate the effect of diversification: (1) Pure beam (current preliminary run): all candidates identical, ; (2) Beam+sampling (proposed): 1 beam + 3 nucleus samples, expected – (awaiting full run); (3) Diverse beam search with group penalty [56]: expected – (intermediate diversity).
Under pure beam collapse, rank aggregation and constrained selection provide zero benefit over greedy decoding. Under beam+sampling, we expect rank aggregation to meaningfully differentiate candidates by support score, enabling the constrained objective to select more faithful candidates even before healing.
14.2. Constraint Removal Study
Table 8 reports the expected directional metric shifts when each constraint component is removed from the full system.
Key predictions: removing the Heal operator should produce the largest faithfulness regression, since in the current run it is the only component that generates genuinely new text. Removing the verifier (falling back to lexical controls only) should produce regression in both MiniCheck and ConRate but less regression in Dup@4. Removing the entailment penalty may slightly improve ROUGE (less conservative language) at the cost of faithfulness. These predictions will be verified in the full 500-example evaluation.
14.3. Verifier Calibration Sensitivity
We analyze sensitivity to the NLI threshold . At (more permissive), fewer claims trigger healing, reducing latency but potentially leaving low-support claims unrepaired. At (more strict), more claims trigger healing, increasing repair coverage but also increasing false-trigger rate when the verifier is uncertain. We set as a balance point calibrated on the 50-example validation subset.
15. Discussion
15.1. Why the Framework Works: Mechanistic Analysis
Three mechanisms jointly explain the faithfulness improvements.
- Feasibility shaping through penalty terms.
The penalty creates a soft feasibility surface over the candidate space. Candidates with low support or high contradiction receive higher penalties, making them less likely to be selected even when they have higher log-probability. This effect is strongest in the nucleus-sampling regime where log-probability differences between candidates are smaller and faithfulness terms carry more weight.
- Opportunity-cost redundancy control.
Under fixed budget , redundancy and faithfulness compete for tokens. The Dup@n penalty creates an explicit opportunity cost for repetition: among similarly-faithful candidates, the one that introduces new supported information is preferred over one that repeats earlier sentences. This is an emergent property of the joint objective- redundancy control improves faithfulness efficiency without requiring explicit faithfulness signals.
- Local repair as feasibility projection.
The Heal operator approximates a projection step onto the feasible set: . The minimal-change principle ( penalty in Equation (37)) biases repair toward solutions close to the original, preserving already-faithful content while correcting violations. This is analogous to projected gradient descent in continuous optimization: the feasibility projection takes a step toward the constraint set from the current infeasible point.
15.2. Failure Mode Analysis
- Verifier false positives.
The DeBERTa NLI verifier can trigger unnecessary repairs on correct sentences whose phrasing differs from evidence without being unfaithful (e.g.,, paraphrase of a fact using synonyms). In the preliminary run, we estimate false-positive rate at approximately 15% of repair triggers based on manual inspection of a 20-example subset. Calibration via temperature scaling on verifier logits is expected to reduce this to .
- Evidence omission bottleneck.
When critical evidence is absent from due to packing, the verifier cannot find supporting evidence for correct claims, potentially triggering unnecessary repairs of accurate information. This is an irreducible limitation of budget-constrained MDS: evidence omitted at packing time cannot be recovered at decoding time. Salience-aware packing (prioritizing high-entropy passages likely to be cited) is a natural extension to address this.
- High-contradiction clusters.
In the High-disagreement stratum, the Heal operator sometimes cannot find a fully feasible replacement that satisfies both support and contradiction constraints simultaneously- when a claim is supported by Document 1 but contradicted by Document 2, no phraseable claim can satisfy both. In these cases, the fallback to hedged language (“officials disagreed on...”, “reports varied on...”) is the correct behavior but may reduce reference ROUGE since references typically commit to one perspective.
15.3. Connections to Prior Theoretical Work
Our Lagrangian relaxation formulation (Equation (16)) connects to constrained MDP formulations of text generation [57], where faithfulness constraints are equivalent to safety constraints in constrained policy optimization. The bounded Heal operator has a parallel in trust region policy optimization [58]: both limit the magnitude of updates to prevent catastrophic interference with already-correct structure.
The rank aggregation approach (Equation (32)) is a special case of Borda count voting, a classical social choice mechanism shown to minimize expected Kemeny distance to the true ranking [59]. Applied to candidate selection, this means our aggregation method is near-optimal in expectation when the true faithfulness ranking is corrupted by independent noise in each signal axis.
15.4. Scope and Limitations
- Small backbone.
LSHT is an 18.4M-parameter model- far smaller than BART-Large (400M) or PEGASUS-Large (568M). The beam collapse failure mode may be less severe in larger models with higher entropy distributions. Conversely, the candidate diversification fix (nucleus sampling) and rank aggregation are model-agnostic and will benefit any backbone exhibiting low candidate diversity.
- Preliminary sample size.
The stratified validation is a diagnostic run, not a statistically powered evaluation. Standard errors are not reported for the stratified sample, and the full 500-example evaluation is required to make quantitative claims about mean performance with appropriate confidence intervals.
- Fixed penalty weights.
The Lagrangian weights are fixed by validation set tuning. In principle, these can be optimized via dual ascent on a held-out set, potentially yielding a Pareto-optimal point on the faithfulness-fluency frontier rather than a fixed operating point.
- News domain specificity.
All evaluations are on MultiNews, a news-domain dataset. Performance on biomedical, legal, or scientific MDS may differ due to domain-specific entity types and contradiction patterns.
15.5. Future Directions
- Dual ascent for adaptive penalty weights. Treating as Lagrangian multipliers updated by dual gradient ascent on a held-out validation set would provide an automatic, dataset-adaptive operating point on the faithfulness-fluency trade-off curve.
- Salience-aware packing. Replacing Hamilton apportionment with a salience-ranked packing that prioritizes high-entropy, controversial, or frequently-cited passages would reduce evidence omission and improve verifier coverage.
- LLM backbone evaluation. Applying FADCO to larger backbones (BART-Large, PEGASUS-Large) and to instruction-tuned LLMs (Llama-2, Mistral) would test whether the framework generalizes beyond small seq2seq models. LLMs may exhibit less beam collapse but would benefit more from rank aggregation due to larger candidate diversity.
- Cross-domain evaluation. Evaluation on WCEP (Wikipedia MDS), MultiXScience (scientific MDS), and MQA (legal MDS) would assess generalization across domains with different entity types, contradiction patterns, and output style requirements.
- Human evaluation. Automatic faithfulness metrics remain imperfect proxies for human judgment. A small-scale human evaluation (annotation with the FactSpan protocol [2]) would provide stronger validation of the MiniCheck improvements observed.
16. Conclusions
We have argued that faithfulness in multi-document summarization is fundamentally a decoding-time constraint satisfaction problem, not a training-time optimization problem, and have presented FADCO as a principled framework for this view.
The framework makes four concrete contributions: (i) a Lagrangian-relaxed constrained decoding objective encoding five faithfulness dimensions; (ii) formal diagnosis of beam collapse and a fix via mixed candidate pools; (iii) rank aggregation that eliminates log-probability scale dominance over faithfulness signals; and (iv) bounded self-healing with provable termination, monotonic acceptance, and minimal distortion.
Preliminary stratified validation on MultiNews confirms three pre-registered directional hypotheses: bounded self-healing improves MiniCheck by and NLI entailment by over beam decoding, and reduces contradiction rate by , with ROUGE-1 trade-off of only and BERTScore degradation of only . Per-stratum analysis reveals that gains are largest in high-disagreement clusters (MiniCheck )- precisely the regime where faithfulness intervention is most needed.
Beyond the empirical results, this work makes a methodological argument: the beam collapse failure mode is a systematic, silent failure that invalidates the assumption underlying all re-ranking and constrained decoding baselines. Any evaluation of verifier-guided or constrained decoding must first verify that candidate diversity is non-trivial; otherwise, observed improvements may be entirely attributable to a single component (as in our run, where healing is the only mechanism producing novel text). We provide both a formal criterion for detecting collapse (Equation (21)) and a practical fix.
Full evaluation on 500 examples with the complete candidate diversification and rank aggregation pipeline is ongoing and will provide quantitative confirmation of the component-wise contributions.
Author Contributions
Conceptualization, S.P. and S.K.S.; Methodology, S.P. and S.K.S.; Software, S.P. and S.K.S.; Formal analysis, S.P. and S.K.S.; Validation, S.P. and S.K.S.; Investigation, S.P. and S.K.S.; Data curation, S.P. and S.K.S.; Writing—Original draft, S.P. and S.K.S.; Writing—Review & editing, S.P. and S.K.S.; Visualization, S.P. and S.K.S.; Supervision, S.K.S.; Project administration, S.K.S. All authors have read and agreed to the published version of the manuscript.
Funding
This research received no external funding.
Institutional Review Board Statement
Not applicable.
Informed Consent Statement
Not applicable.
Data Availability Statement
The code and hyperparameter configurations are available at https://github.com/Sameer-dev1/FADCO. The stratified evaluation splits supporting this study are available from the corresponding author upon reasonable request. MultiNews is publicly available at https://github.com/Alex-Fabbri/Multi-News.
Acknowledgments
The authors declares that no specific funding, technical assistance, or external support was received for this work.
Conflicts of Interest
The authors declare no conflicts of interest.
Appendix A. Hyperparameter Settings
Table A1.
Full hyperparameter configuration for FADCO (LSHT instantiation). All values were fixed on a 50-example validation subset prior to test evaluation.
Table A1.
Full hyperparameter configuration for FADCO (LSHT instantiation). All values were fixed on a 50-example validation subset prior to test evaluation.
| Parameter | Value | Rationale |
|---|---|---|
| Constraint penalty weights | ||
| 0.6 | Moderate repetition suppression | |
| 1.2 | Strong contradiction avoidance | |
| 0.8 | Balanced support enforcement | |
| Hard constraint thresholds | ||
| 0.45 | Minimum entailment probability | |
| 0.35 | Maximum contradiction probability | |
| 0.15 | Maximum 4-gram redundancy rate | |
| 0.35 | Healing trigger confidence floor | |
| 0.40 | Secondary healing trigger | |
| Candidate generation | ||
| Beam width K | 4 | 1 beam + 3 samples |
| Nucleus temperature | 0.85 | Mild temperature softening |
| Nucleus top-p | 0.92 | 92% probability mass nucleus |
| Dedup retries | Per candidate slot | |
| Self-healing | ||
| Max retries R | 3 | Bounded compute |
| Min improvement | 0.02 | Monotonic acceptance margin |
| Max window | 80 tokens | Repair scope limit |
| Max length delta | 20 tokens | Expansion allowance |
| Edit penalty | 0.2 | Distortion regularization |
| 0.03 | Min support gain to accept | |
| 0.05 | Max contradiction increase to accept | |
| Rank aggregation weights | ||
| 2.0 | Prioritize support signal | |
| 1.5 | Strong contradiction avoidance | |
| 1.0 | Retain fluency signal | |
| 0.5 | Mild redundancy penalty | |
| Scoring and inference | ||
| Length norm. | 0.7 | Moderate length normalization |
| N-gram blocking size | 3 | Trigram blocking in Beam+NgBlock |
| Verifier batch size | 128 | Fits RTX PRO 6000 VRAM |
| Unigram evidence top-k | 7 | Evidence snippets per claim |
| Healing violation weights | ||
| 1.0 | Entailment deficit weight | |
| 0.8 | Contradiction excess weight | |
| 0.3 | Local redundancy weight | |
Appendix B. Metric Reproducibility Audit
We document the exact failure modes of AlignScore and QuestEval that motivated their replacement with MiniCheck.
- AlignScore failure.
AlignScore [13] requires transformers==4.27.0 and torch==1.13.0. Under transformers>=4.35.0 and torch>=2.0.0 (required by DeBERTa-v3 and other framework components), the AlignScore class raises an ImportError due to renamed internal APIs in the transformers library. A pinned-version environment is technically possible but creates incompatibilities with NumPy 1.x vs. 2.x for array operations used elsewhere in the pipeline.
- QuestEval failure.
QuestEval [14] requires spacy==3.1.x and the en-core-web-lg==3.1.0 language model. Under spacy>=3.5.0 (current stable release), the en-core-web-lg model version is incompatible, causing a pipeline initialization crash. The QuestEval package has not been updated to support current spaCy versions.
- MiniCheck verification.
MiniCheck [15] installs without conflicts under: transformers>=4.35.0, torch>=2.1.0, sentencepiece>=0.1.99. All dependencies are compatible with the rest of the FADCO codebase. Installation command: pip install minicheck.
Appendix C. FADCO System Summary
Table A2.
FADCO complete system summary: components, their purpose, computational cost, and formal guarantees.
Table A2.
FADCO complete system summary: components, their purpose, computational cost, and formal guarantees.
| Component | Purpose | Cost | Guarantee |
|---|---|---|---|
| Hamilton Packing | Proportional evidence allocation | Quota fairness | |
| Beam Anchor | Quality baseline candidate | Most probable output | |
| Nucleus Sampling | Diverse candidates | Diversity w.h.p. | |
| Batched NLI | Support/contradiction scoring | Accurate entailment | |
| Rank Aggregation | Scale-independent selection | No axis dominance | |
| Heal Operator | Local violation repair | Termination, monotonic |
References
- Maynez, J.; Narayan, S.; Bohnet, B.; McDonald, R. On faithfulness and factuality in abstractive summarization. In Proceedings of the Proceedings of ACL, 2020; pp. 1906–1919. [Google Scholar]
- Pagnoni, A.; Balachandran, V.; Tsvetkov, Y. Understanding factuality in abstractive summarization with FRANK: A benchmark for factuality metrics. In Proceedings of the Proceedings of NAACL, 2021; pp. 4812–4829. [Google Scholar]
- Cao, S.; Wang, L. Hallucinated but factual! Inspecting the factuality of hallucinations in abstractive summarization. In Proceedings of the Proceedings of ACL, 2022; pp. 3340–3354. [Google Scholar]
- Kang, D.; Hashimoto, T.B. Improved Natural Language Generation via Loss Truncation. In Proceedings of the Proceedings of ACL, 2020. [Google Scholar]
- Zhang, T.; et al. Benchmarking Faithfulness in Natural Language Generation. arXiv 2024. [Google Scholar]
- Nan, L.; Wiseman, S.; Bansal, M.; Chen, Y.; Perez, E.; Mei, H.; Barzilay, R. Entity-level factual consistency of abstractive text summarization. In Proceedings of the Proceedings of EACL, 2021; pp. 2727–2733. [Google Scholar]
- Chen, S.; Zhang, F.; Sone, Y.; Litman, D. Improving faithfulness in abstractive summarization with contrast candidate generation and selection. In Proceedings of the Proceedings of NAACL, 2021; pp. 5935–5941. [Google Scholar]
- Zhu, C.; Hinthorn, W.; Xu, R.; Zeng, Q.; Zeng, M.; Huang, X.; Jiang, M. Enhancing factual consistency of abstractive summarization. In Proceedings of the Proceedings of NAACL, 2021; pp. 718–733. [Google Scholar]
- Ladhak, F.; Durmus, E.; Cardie, C.; McKeown, K. Faithful or extractive? On mitigating the faithfulness-abstractiveness trade-off in abstractive summarization. In Proceedings of the Proceedings of ACL, 2022; pp. 1410–1421. [Google Scholar]
- Wan, Z.; Wan, F.; Yu, W.; Du, Y.; Lam, W.; Pan, B. Faithfulness-aware decoding strategies for abstractive summarization. In Proceedings of the Proceedings of EACL, 2023; pp. 891–908. [Google Scholar]
- Meister, C.; Vieira, T.; Cotterell, R. If beam search is the answer, what was the question? In Proceedings of the Proceedings of EMNLP, 2020; pp. 2173–2185. [Google Scholar]
- Dziri, N.; Milton, A.; Yu, M.; Zaiane, O.; Reddy, S. On the origin of hallucinations in conversational models. Proceedings of NAACL, 2022; pp. 5765–5780. [Google Scholar]
- Zha, Z.; Bohnet, B.; Dong, N.; Metzler, D.; Ni, J. AlignScore: Evaluating factual consistency with a unified alignment function. In Proceedings of the Proceedings of ACL, 2023; pp. 11328–11348. [Google Scholar]
- Scialom, T.; Dray, P.A.; Gallé, M.; Gallinari, P.; Piwowarski, B.; Staiano, J.; Wang, A. QuestEval: Summarization asks for fact-based evaluation. In Proceedings of the Proceedings of EMNLP, 2021; pp. 6594–6604. [Google Scholar]
- Tang, L.; Laban, P.; Carenini, G. MiniCheck: Efficient fact-checking of LLMs on grounding documents. In Proceedings of the Proceedings of EMNLP, 2024; pp. 8818–8847. [Google Scholar]
- Zhang, T.; Kishore, V.; Wu, F.; Weinberger, K.; Artzi, Y. BERTScore: Evaluating text generation with BERT. In Proceedings of the Proceedings of ICLR, 2020. [Google Scholar]
- Lewis, M.; Liu, Y.; Goyal, N.; Ghazvininejad, M.; Mohamed, A.; Levy, O.; Stoyanov, V.; Zettlemoyer, L. BART: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension. In Proceedings of the Proceedings of ACL, 2020; pp. 7871–7880. [Google Scholar]
- Zhang, J.; Zhao, Y.; Saleh, M.; Liu, P. PEGASUS: Pre-training with extracted gap-sentences for abstractive summarization. In Proceedings of the Proceedings of ICML, 2020; Vol. 119, pp. 11328–11339. [Google Scholar]
- Raffel, C.; Shazeer, N.; Roberts, A.; Lee, K.; Narang, S.; Matena, M.; Zhou, Y.; Li, W.; Liu, P. Exploring the limits of transfer learning with a unified text-to-text transformer. J. Mach. Learn. Res. 2020, 21, 1–67. [Google Scholar]
- Zaheer, M.; Guruganesh, G.; Dubey, A.; Ainslie, J.; Alberti, C.; Ontanon, S.; Pham, P.; Ravula, A.; Wang, Q.; Yang, L.; et al. Big Bird: Transformers for longer sequences. In Proceedings of the Proceedings of NeurIPS, 2020; Vol. 33, pp. 17283–17297. [Google Scholar]
- Shleifer, S.; Rush, A. Pre-trained summarization distillation. arXiv 2020, arXiv:2010.13002. [Google Scholar] [CrossRef]
- Fabbri, A.R.; Li, I.; She, T.; Li, S.; Radev, D. Multi-News: A large-scale multi-document summarization dataset and abstractive hierarchical model. In Proceedings of the Proceedings of ACL, 2019; pp. 1074–1084. [Google Scholar]
- Pandey, S.; Singh, S.K. Lightweight Self-Healing Transformers for Faithful Summarization. Technical Report, 2024. [Google Scholar]
- Falke, T.; Ribeiro, L.F.R.; Utama, P.A.; Dagan, I.; Gurevych, I. Ranking generated summaries by correctness: An interesting but challenging application for natural language inference. In Proceedings of the Proceedings of ACL, 2019; pp. 2214–2220. [Google Scholar]
- Kryscinski, W.; McCann, B.; Xiong, C.; Socher, R. Evaluating the factual consistency of abstractive text summarization. In Proceedings of the Proceedings of EMNLP, 2020; pp. 9332–9346. [Google Scholar]
- Wang, A.; Cho, K.; Lewis, M. Asking and answering questions to evaluate the factual consistency of summaries. In Proceedings of the Proceedings of ACL, 2020; pp. 5008–5020. [Google Scholar]
- Fabbri, A.; Wu, C.S.; Liu, W.; Xiong, C. QAFactEval: Improved QA-based factual consistency evaluation for summarization. In Proceedings of the Proceedings of NAACL, 2022; pp. 2587–2601. [Google Scholar]
- Deutsch, D.; Bedrax-Weiss, T.; Roth, D. Towards question-answering as an automatic metric for evaluating the content quality of a summary. Trans. ACL 2021, 9, 774–789. [Google Scholar] [CrossRef]
- Laban, P.; Schnabel, T.; Bennett, P.; Hearst, M. SummaC: Re-visiting NLI-based models for inconsistency detection in summarization. Trans. ACL 2022, 10, 163–177. [Google Scholar] [CrossRef]
- Min, S.; Krishna, K.; Lyu, X.; Lewis, M.; Yih, W.t.; Koh, P.W.; Iyyer, M.; Zettlemoyer, L.; Hajishirzi, H. FActScore: Fine-grained atomic evaluation of factual precision in long form text generation. In Proceedings of the Proceedings of EMNLP, 2023; pp. 12076–12100. [Google Scholar]
- Manakul, P.; Liusie, A.; Gales, M. SelfCheckGPT: Zero-resource black-box hallucination detection for generative large language models. In Proceedings of the Proceedings of EMNLP, 2023; pp. 9004–9017. [Google Scholar]
- Anderson, P.; Fernando, B.; Johnson, M.; Gould, S. Guided Open Vocabulary Image Captioning with Constrained Beam Search. In Proceedings of the Proceedings of EMNLP, 2017. [Google Scholar]
- Post, M.; Vilar, D. Fast Lexically Constrained Decoding with Dynamic Beam Allocation for Neural Machine Translation. In Proceedings of the Proceedings of NAACL, 2018. [Google Scholar]
- Keskar, N.S.; McCann, B.; Varshney, L.R.; Xiong, C.; Socher, R. CTRL: A Conditional Transformer Language Model for Controllable Generation. arXiv 2019, arXiv:1909.05858. [Google Scholar] [CrossRef]
- Rashkin, H.; Nikolaev, V.; Lamm, M.; Aroyo, L.; Collins, M.; Das, D.; Petrov, S.; Tomar, G.S.; Turc, I.; Reitter, D. Measuring Attribution in Natural Language Generation Models. In Proceedings of the Proceedings of Computational Linguistics, 2023. [Google Scholar]
- Welleck, S.; Kulikov, I.; Roller, S.; Dinan, E.; Cho, K.; Weston, J. Neural Text Generation with Unlikelihood Training. In Proceedings of the Proceedings of ICLR, 2020. [Google Scholar]
- Eikema, B.; Aziz, W. Is MAP Decoding All You Need? The Inadequacy of the Mode in Neural Machine Translation. In Proceedings of the Proceedings of COLING, 2020. [Google Scholar]
- Freitag, M.; Foster, G.; Grangier, D.; Ratnakar, V.; Tan, Q.; Macherey, W. High Quality Rather than High Model Probability: Minimum Bayes Risk Decoding with Neural Metrics. In Proceedings of the Transactions of the Association for Computational Linguistics, 2022. [Google Scholar]
- Yang, K.; Klein, D. FUDGE: Controlled text generation with future discriminators. In Proceedings of the Proceedings of NAACL, 2021; pp. 3511–3535. [Google Scholar]
- Krause, B.; Gotmare, A.; McCann, B.; Keskar, N.; Joty, S.; Socher, R.; Rajani, N. GeDi: Generative discriminator guided sequence generation. In Proceedings of the Proceedings of EMNLP Findings, 2021; pp. 4929–4952. [Google Scholar]
- Gao, T.; Fisch, A.; Chen, D. RARR: Researching and revising what language models say, using language models. In Proceedings of the Proceedings of ACL, 2023; pp. 16477–16508. [Google Scholar]
- Madaan, A.; Tandon, N.; Gupta, P.; Hallinan, S.; Gao, L.; Wiegreffe, S.; Alon, U.; Dziri, N.; Prabhumoye, S.; Yang, Y.; et al. Self-refine: Iterative refinement with self-feedback. In Proceedings of the Proceedings of NeurIPS, 2023; Vol. 36. [Google Scholar]
- Schick, T.; Dwivedi-Yu, J.; Dessì, R.; Raileanu, R.; Lomeli, M.; Zettlemoyer, L.; Cancedda, N.; Scialom, T. PEER: A collaborative language model. In Proceedings of the Proceedings of ICLR, 2023. [Google Scholar]
- He, P.; Liu, X.; Gao, J.; Chen, W. DeBERTa: Decoding-enhanced BERT with disentangled attention. In Proceedings of the Proceedings of ICLR, 2021. [Google Scholar]
- Bowman, S.; Angeli, G.; Potts, C.; Manning, C. A large annotated corpus for learning natural language inference. In Proceedings of the Proceedings of EMNLP, 2015; pp. 632–642. [Google Scholar]
- Williams, A.; Nangia, N.; Bowman, S. A broad-coverage challenge corpus for sentence understanding through inference. In Proceedings of the Proceedings of NAACL, 2018; pp. 1112–1122. [Google Scholar]
- Thorne, J.; Vlachos, A.; Christodoulopoulos, C.; Mittal, A. FEVER: A large-scale dataset for fact extraction and verification. In Proceedings of the Proceedings of NAACL, 2018; pp. 809–819. [Google Scholar]
- Bertsekas, D.P. Nonlinear Programming, 2nd ed.; Athena Scientific: Belmont, MA, 1999. [Google Scholar]
- Holtzman, A.; Buys, J.; Du, L.; Forbes, M.; Choi, Y. The curious case of neural text degeneration. In Proceedings of the Proceedings of ICLR, 2020. [Google Scholar]
- Su, J.; Lu, Y.; Pan, S.; Murtadha, A.; Wen, B.; Liu, Y. RoFormer: Enhanced transformer with rotary position embedding. Neurocomputing 2024, 568, 127063. [Google Scholar] [CrossRef]
- Hendrycks, D.; Gimpel, K. Gaussian error linear units (GELUs). arXiv 2016, arXiv:1606.08415. [Google Scholar]
- Lin, C.Y. ROUGE: A package for automatic evaluation of summaries. In Proceedings of the Proceedings of the ACL Workshop, 2004; pp. 74–81. [Google Scholar]
- Khalifa, M.; Elsahar, H.; Dymetman, M. A distributional approach to controlled text generation. In Proceedings of the Proceedings of ICLR, 2021. [Google Scholar]
- Rajpurkar, P.; Zhang, J.; Lopyrev, K.; Liang, P. SQuAD: 100,000+ questions for machine comprehension of text. In Proceedings of the Proceedings of EMNLP, 2016; pp. 2383–2392. [Google Scholar]
- Zhao, Y.; et al. Reducing Hallucination in Neural Text Generation. In Proceedings of the Proceedings of NLP Research, 2020. [Google Scholar]
- Vijayakumar, A.; Cogswell, M.; Selvaraju, R.; Sun, Q.; Lee, S.; Crandall, D.; Batra, D. Diverse beam search: Decoding diverse solutions from neural sequence models. In Proceedings of the arXiv, 2016. [Google Scholar]
- Lu, X.; West, P.; Zellers, R.; Le Bras, R.; Bhagavatula, C.; Choi, Y. NeuroLogic decoding: (un)supervised neural text generation with predicate logic constraints. In Proceedings of the Proceedings of NAACL, 2021; pp. 4288–4299. [Google Scholar]
- Schulman, J.; Levine, S.; Abbeel, P.; Jordan, M.; Moritz, P. Trust region policy optimization. In Proceedings of the Proceedings of ICML, 2015; pp. 1889–1897. [Google Scholar]
- Dwork, C.; Kumar, R.; Naor, M.; Sivakumar, D. Rank aggregation methods for the web. In Proceedings of the Proceedings of WWW, 2001; pp. 613–622. [Google Scholar]
Table 1.
Design requirements, mechanisms, constraint targets, and sections where each is formalized. All mechanisms operate at inference time without modifying model weights.
Table 1.
Design requirements, mechanisms, constraint targets, and sections where each is formalized. All mechanisms operate at inference time without modifying model weights.
| Requirement | Mechanism | Constraint | § |
|---|---|---|---|
| Evidence grounding | Verifier entailment | Equation (18) | 5.2 |
| Non-contradiction | NLI contradiction check | Equation (19) | 5.2 |
| Budget control | Hamilton packing | Equation (17) | 5.1 |
| Redundancy control | Dup@n penalty | Equation (20) | 5.3 |
| Entity consistency | NER + provenance | Equation (5) | 5.3 |
| Local repair | Bounded Heal(·) | Equation (37) | 8 |
| Bounded compute | Retry limit R | Equation (40) | 8.5 |
| Candidate diversity | Nucleus sampling | Prop. 2 | 6 |
| Scale-independent selection | Rank aggregation | Equation (32) | 7 |
| Model agnosticism | Decoder-side control | — | 5 |
Table 2.
Limitations matrix for baseline objectives vs. FADCO across five faithfulness-relevant properties. Grnd. = evidence grounding, Con. = non-contradiction, Rep. = redundancy control, Bdg. = budget-aware generation, Rpr. = violation repair. ✓ = explicitly enforced; △ = partially addressed, often fails in MDS; × = not addressed.
Table 2.
Limitations matrix for baseline objectives vs. FADCO across five faithfulness-relevant properties. Grnd. = evidence grounding, Con. = non-contradiction, Rep. = redundancy control, Bdg. = budget-aware generation, Rpr. = violation repair. ✓ = explicitly enforced; △ = partially addressed, often fails in MDS; × = not addressed.
| Model / System | Grnd. | Con. | Rep. | Bdg. | Rpr. |
|---|---|---|---|---|---|
| BART [17] | △ | × | △ | × | × |
| PEGASUS [18] | △ | × | △ | × | × |
| T5 [19] | △ | × | △ | × | × |
| BigBird [20] | △ | × | △ | △ | × |
| DistilBART [21] | △ | × | △ | × | × |
| Beam+NgBlock | △ | × | △ | × | × |
| Rerank+Verifier | △ | △ | × | × | × |
| RARR [41] | ✓ | △ | × | × | △ |
| Self-Refine [42] | △ | △ | △ | × | △ |
| FADCO (ours) | ✓ | ✓ | ✓ | ✓ | ✓ |
Table 3.
MultiNews dataset statistics and FADCO budget configuration. Input budget B is set to 4096 tokens to allow for full multi-document context; output budget is set to 256 tokens following standard MDS evaluation practice.
Table 3.
MultiNews dataset statistics and FADCO budget configuration. Input budget B is set to 4096 tokens to allow for full multi-document context; output budget is set to 256 tokens following standard MDS evaluation practice.
| Split | #Clusters | Avg Docs | Avg Src Len | Avg Ref Len | B | |
|---|---|---|---|---|---|---|
| Train | 44,972 | 2.79 | 1,532 | 261 | 4,096 | 256 |
| Val | 5,622 | 2.79 | 1,547 | 263 | 4,096 | 256 |
| Test | 5,622 | 2.79 | 1,551 | 259 | 4,096 | 256 |
Table 4.
Inference-time baselines. All systems use the same LSHT backbone and packed evidence . “Ver.” indicates whether a faithfulness verifier is active; “Heal” indicates whether bounded self-healing is applied.
Table 4.
Inference-time baselines. All systems use the same LSHT backbone and packed evidence . “Ver.” indicates whether a faithfulness verifier is active; “Heal” indicates whether bounded self-healing is applied.
| System | Cand. Gen. | Ver. | Heal | Key Setting |
|---|---|---|---|---|
| Greedy | Greedy () | × | × | Standard inference |
| Beam | Beam () | × | × | Length norm. |
| Beam+NgBlock | Beam () | × | × | Block 3-gram repeats |
| Rerank(Rep) | Beam () | × | × | |
| Rerank+Verifier | Beam () | Post | × | |
| PostHocVerify | Beam () | Post | × | Accept if |
| Ours: Constrained | Mixed pool | In-loop | × | Rank aggregation |
| Ours: Const+Heal | Mixed pool | In-loop | ✓ | Bounded |
Table 5.
Main results: quality, faithfulness, and redundancy metrics on stratified validation ( examples, 3 strata). Higher is better for R-1/2/L, BERTScore, MiniCheck, QA-F1, Entail. Lower is better for ConRate, Dup@4. Bold = best in column. Full 500-example evaluation ongoing. Relative improvements over Beam are in parentheses.
Table 5.
Main results: quality, faithfulness, and redundancy metrics on stratified validation ( examples, 3 strata). Higher is better for R-1/2/L, BERTScore, MiniCheck, QA-F1, Entail. Lower is better for ConRate, Dup@4. Bold = best in column. Full 500-example evaluation ongoing. Relative improvements over Beam are in parentheses.
| System | R-1 | R-2 | R-L | BERT-F1 | MiniCheck | QA-F1 | Entail | ConRate | Dup@4 |
|---|---|---|---|---|---|---|---|---|---|
| Greedy | 0.3943 | 0.1165 | 0.1899 | 0.8446 | 0.1441 | 0.1471 | 0.2161 | 0.3439 | 0.0014 |
| Beam | 0.3943 | 0.1165 | 0.1899 | 0.8446 | 0.1441 | 0.1471 | 0.2161 | 0.3439 | 0.0014 |
| Beam+NgBlock | 0.3893 | 0.1158 | 0.2050 | 0.8442 | 0.1940 | 0.2069 | 0.2569 | 0.3265 | 0.0044 |
| Rerank(Rep) | 0.3943 | 0.1165 | 0.1899 | 0.8446 | 0.1441 | 0.1471 | 0.2161 | 0.3439 | 0.0014 |
| Rerank+Verifier | 0.3943 | 0.1165 | 0.1899 | 0.8446 | 0.1441 | 0.1471 | 0.2161 | 0.3439 | 0.0014 |
| PostHocVerify | 0.3943 | 0.1165 | 0.1899 | 0.8446 | 0.1441 | 0.1471 | 0.2161 | 0.3439 | 0.0014 |
| Ours: Constrained | 0.3943 | 0.1165 | 0.1899 | 0.8446 | 0.1441 | 0.1471 | 0.2161 | 0.3439 | 0.0014 |
| Ours: Const+Heal | 0.3874 | 0.1126 | 0.1878 | 0.8443 | 0.2388 | 0.1811 | 0.3437 | 0.3239 | 0.0104 |
| Relative gains over Beam (Ours:Const+Heal) | |||||||||
| — | |||||||||
Table 6.
Per-stratum MiniCheck and ConRate for Beam vs. Ours:Const+Heal. = absolute change. High-disagreement clusters benefit most from self-healing.
Table 6.
Per-stratum MiniCheck and ConRate for Beam vs. Ours:Const+Heal. = absolute change. High-disagreement clusters benefit most from self-healing.
| MiniCheck | ConRate | |||||
|---|---|---|---|---|---|---|
| Stratum | Beam | Ours | Beam | Ours | ||
| Low () | 0.1623 | 0.2341 | 0.2912 | 0.2791 | ||
| Medium () | 0.1389 | 0.2271 | 0.3622 | 0.3441 | ||
| High () | 0.1234 | 0.2571 | 0.4153 | 0.3641 | ||
| Overall () | 0.1441 | 0.2388 | 0.3439 | 0.3239 | ||
Table 7.
Inference efficiency profile (, NVIDIA RTX PRO 6000). Latency reported as mean over examples. The anomalously low latency of Ours:Const+Heal is explained in the text.
Table 7.
Inference efficiency profile (, NVIDIA RTX PRO 6000). Latency reported as mean over examples. The anomalously low latency of Ours:Const+Heal is explained in the text.
| System | Latency (ms) | Peak Mem (MB) | Rel. Latency |
|---|---|---|---|
| Greedy | 509.1 | 0.0 | 1.00× |
| Beam | 606.5 | 0.1 | 1.19× |
| Beam+NgBlock | 597.7 | 0.0 | 1.17× |
| Rerank(Rep) | 599.4 | 0.1 | 1.18× |
| Rerank+Verifier | 675.8 | 0.1 | 1.33× |
| PostHocVerify | 609.2 | 0.1 | 1.20× |
| Ours: Constrained | 635.6 | 0.2 | 1.25× |
| Ours: Const+Heal | 124.0 | 0.0 | 0.24× |
Table 8.
Ablation study: expected directional metric changes vs. full Ours:Const+Heal when each component is removed. Full quantitative results on 500 examples are pending; directions are derived from the mechanisms described in §5–8.
Table 8.
Ablation study: expected directional metric changes vs. full Ours:Const+Heal when each component is removed. Full quantitative results on 500 examples are pending; directions are derived from the mechanisms described in §5–8.
| Config | R-1 | Mini | QA-F1 | Entail | Con | Lat |
|---|---|---|---|---|---|---|
| Full system | — | — | — | — | — | — |
| No Heal | ↓ | ↓ | ↓ | ↑ | ↓ | |
| No Verifier | ↓ | ↓ | ↓ | ↑ | ↓ | |
| No Entail Pen. | ↑ | ↓ | ↓ | ↓ | ↓ | |
| No Con Check | ↑ | ↓ | ||||
| No Rep Pen. | ↓ | ↓ | ||||
| No Rank Agg. | ↓ | ↓ | ↓ | ↑ | ||
| No Nucleus | ↑ | ↓ | ↓ | ↓ | ↑ | ↓ |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.