Preprint
Article

This version is not peer-reviewed.

MetaPCR-LLM: Universal Proof-Carrying Reasoning Across Heterogeneous Logics for Large Language Model Validation

Submitted:

09 September 2026

Posted:

10 September 2026

You are already at the latest version

Abstract
Large language models (LLMs) increasingly produce reasoning that combines deductive, abductive, probabilistic, defeasible, temporal and constraint-based inference, yet most symbolic validation approaches assume a single target logic or translate heterogeneous reasoning into one common formalism. We introduce MetaPCR-LLM, a proof-carrying framework for validating LLM reasoning across heterogeneous native logics. MetaPCR-LLM represents a reasoning trace as a typed proof graph in which each step is assigned to an appropriate native validator and carries a proof certificate or machine-checkable witness together with provenance, uncertainty, temporal, and defeasibility metadata. Transitions between logical systems are treated as explicit proof-carrying bridges, allowing the framework to distinguish locally valid inference from globally invalid cross-logic composition. Counter-abduction further tests accepted explanations against independently validated alternatives, while the same mechanism supports certified recomposition of reasoning fragments produced by multiple LLMs. We evaluate the framework on Truthful-Halluc, Med-Halluc, eSNLI-Halluc and Autoimmune-narrate-halluc across single- and mixed-logic reasoning regimes. MetaPCR-LLM achieves 0.86 overall accuracy, 0.88 step-validity accuracy and 0.86 chain-certification accuracy, with a cross-dataset macro-F1 of 0.838. Its pooled hallucination F1 reaches 0.83, compared with 0.77 for the ValidLLP-style baseline, an absolute improvement of 0.06 at the reported precision. The advantage remains especially pronounced on heterogeneous reasoning: mixed-logic F1 reaches 0.80, compared with 0.74 for ValidLLP-style validation and 0.70 for MetaPCR without bridge checking. Bridge validation reaches 0.88 accuracy and explicit proof diagnostics improve failure-type and validator localization. In multi-LLM recomposition, the complete certificate--bridge--label configuration reaches 0.86 accuracy and reduces invalid compositions from 0.29 to 0.18. Human evaluation further shows an increase in reasoning-assessment accuracy from 0.68 to 0.88. These results show that proof-carrying heterogeneous validation improves the aggregate answer-level comparison while also providing reliable composition, localization and auditing of reasoning that spans multiple formal systems.
Keywords: 
;  ;  ;  ;  

1. Introduction

Large Language Models (LLMs) have substantially improved performance on tasks requiring multi-step reasoning, particularly when intermediate reasoning is elicited through Chain-of-Thought (CoT) prompting and related sampling strategies [1,2]. Nevertheless, producing an explicit reasoning trace is not equivalent to producing a valid proof. CoT steps may contain unsupported premises, invalid inference transitions, incorrect symbolic manipulations, or post-hoc rationalizations whose relationship to the final answer is uncertain [4,5]. Process supervision partially addresses this limitation by evaluating intermediate reasoning rather than only final outcomes, but a process reward score remains a learned assessment rather than an independently checkable logical certificate [3].
These limitations have motivated a growing class of neuro-symbolic architectures in which the LLM is responsible primarily for semantic interpretation or formalization while a symbolic engine performs the actual inference. Logic-LM translates natural-language problems into symbolic representations and invokes deterministic solvers, additionally using solver feedback to refine unsuccessful formalizations [6]. LINC similarly treats the language model as a semantic parser and delegates deduction to a first-order theorem prover [7]. Faithful Chain-of-Thought externalizes selected reasoning operations into executable symbolic structures [4], while Symbolic Chain-of-Thought explicitly integrates symbolic expressions and logical rules into the generated reasoning trajectory [8]. More recent work has extended this direction through neuro-symbolic consistency objectives and logic-aided exploration [9,10]. Collectively, these approaches establish an important architectural principle: LLMs can generate or translate candidate reasoning, while logical validity can be delegated to a more deterministic reasoning mechanism.
A parallel development has occurred in formal theorem proving. LLMs increasingly interact with proof assistants such as Lean, whose kernel can provide substantially stronger guarantees than an LLM-based judge. LeanDojo provides an environment for integrating language models with Lean and demonstrates retrieval-augmented theorem proving over large formal libraries [12]. Recent systems use formal proof assistants not only to generate proofs but also to validate intermediate natural-language reasoning. SAFE formalizes individual mathematical reasoning steps in Lean 4 and seeks explicit proofs rather than assigning opaque correctness scores [13]. TP-as-a-Judge similarly employs theorem-prover feedback to assess intermediate LLM reasoning and refine autoformalization [14]. Work on natural-language requirements and theorem-prover-mediated natural-language inference further demonstrates both the promise and the central difficulty of this paradigm: formal verification is reliable only after the natural-language claim has been translated faithfully into the formal language [15,16].
Proof-Carrying Reasoning with Large Language Models (PCRLLM) takes an especially relevant step toward making reasoning itself an inspectable validation object [17]. Instead of generating an unrestricted explanation, PCRLLM constrains reasoning to minimal inference steps in which premises, inference rules and conclusions are made explicit. These structures support automatic single-step and inter-step verification and permit reasoning fragments produced by different LLMs to be compared and recombined. However, PCRLLM assumes a target logic for the proof-carrying chain; its current realization uses first-order Non-Axiomatic Logic (NAL). The framework therefore addresses the important problem of whether a reasoning chain conforms to a specified logic, but not the more general case in which different steps of the same reasoning process require different formal systems.
This distinction becomes important in realistic reasoning tasks. A single generated recommendation may simultaneously depend on deductive rules, numerical constraints, temporal relationships, uncertain evidence, defeasible exceptions and abductive hypotheses. Abductive Logic Programming (ALP) provides a principled formalism for constructing explanations through admissible hypotheses and integrity constraints [22]; ProbLog provides probabilistic reasoning over logic programs [24]; abstract argumentation provides explicit mechanisms for conflict, attack, defense and acceptability [23]; and constraint or theorem-proving systems can provide hard validation of arithmetic, algebraic, or formally specified properties. These formalisms have different semantics and should not necessarily be reduced to one another merely to obtain a uniform LLM-validation interface.
A complementary solution is to enrich logical propositions with meta-level information. Gabbay’s Labelled Deductive Systems (LDS) replace reasoning over bare formulas with reasoning over structured labelled formulas, allowing labels to carry information such as temporal status, provenance, reliability, resources, priorities, or other proof-relevant dimensions [18]. ValidLLP4LLM applies this principle to LLM validation through Labeled Logic Programs, where source, provenance, uncertainty, time, modality, priority and defeasibility accompany object-level claims and influence validation [25]. This provides a common labelled substrate for heterogeneous reasoning dimensions. However, different logical regimes are primarily expressed through label disciplines and validation profiles over the common LLP substrate rather than through independently preserved native logical systems.
The problem addressed in this paper is therefore a level above either single-logic proof carrying or single-substrate labelled validation:
How can an LLM-generated reasoning chain be formally validated when different inference steps belong to different native logical systems?
We propose MetaPCR-LLM, a meta-level architecture for proof-carrying reasoning across heterogeneous logics. The basic unit is not simply a natural-language reasoning step or a labelled proposition, but a typed proof-carrying inference object containing its premises, conclusion, native logic, proof obligation, externally checkable certificate and validation label. A numerical constraint may therefore be certified by an SMT solver, a mathematical statement by Lean, an explanatory transition by ALP, an uncertain conclusion by a probabilistic logic program and a conflict-sensitive inference by an argumentation or defeasible-reasoning engine.
This architecture introduces a problem largely absent from existing proof-carrying LLM systems: local proof validity does not guarantee valid cross-logic composition. Suppose that a proposition derived under an SMT interpretation is reused as an abductive premise, or that a probabilistic conclusion becomes an input to a temporal or defeasible rule. The proposition may have a meaningful representation in both systems, but the transformation between those representations is itself a semantic operation. MetaPCR-LLM therefore introduces proof-carrying logic bridges. A bridge records a translation between native logical representations together with the semantic property that the mapping is claimed to preserve—for example, entailment, satisfiability, contradiction, or constraint satisfaction—and the evidence used to validate that claim.
This aspect connects the proposed framework to long-established work on logic-independent reasoning. The theory of Institutions characterizes a logical system abstractly in terms of signatures, sentences, models and satisfaction, with translations constrained by satisfaction preservation [19]. Meseguer’s General Logics extends the abstraction toward proof-theoretic consequence structures [20]. Research on fibring and the analysis and synthesis of logics studies principled ways to compose heterogeneous logical systems while investigating preservation of their properties [21]. MetaPCR-LLM does not attempt to replace these theories with a new universal logic. Instead, it applies their central separation principle to LLM validation: native logics remain independent, while a meta-layer records how proof objects may be transported and composed across logical boundaries.
The resulting structure is a typed proof graph. Nodes represent normalized propositions; edges represent locally certified inference steps; edge types identify native logics; labels record provenance and contextual admissibility; and bridge objects certify transitions between logically heterogeneous edges. This representation allows the validator to distinguish failures that ordinary hallucination classification often collapses together: unsupported premise, locally invalid inference, violated hard constraint, defeated explanation, cross-logic contradiction and invalid formal translation.
The architecture also naturally incorporates proof-carrying counter-abduction. If an LLM proposes an explanatory hypothesis, the framework may construct competing abductive hypotheses. However, a rival explanation is not allowed to defeat the original merely because another LLM finds it plausible. Both explanations must carry independently validated abductive certificates, after which their admissibility, assumptions, integrity-constraint satisfaction, provenance and explanatory coverage can be compared. This extends classical ALP-style hypothesis formation [22] and the use of abduction in labelled deduction [18] toward independently checkable LLM explanation validation.
MetaPCR-LLM also generalizes the multi-model recomposition principle introduced by PCRLLM. PCRLLM shows that validated intermediate reasoning steps can be selected from different LLM outputs and recombined rather than choosing only among complete answers [17]. In MetaPCR-LLM, the same idea extends across logical systems: a Lean-certified step generated by one model, an SMT-certified step generated by another and an ALP-certified explanation generated by a third can participate in one composite reasoning graph, provided that each step is locally valid and every cross-logic transition is admissible.
The main contributions of this paper are therefore:
1.
We introduce heterogeneous proof-carrying reasoning, in which individual LLM-generated inference steps can be certified by different native logical systems rather than being forced into a single target logic.
2.
We define a typed, labelled proof-graph representation combining native proof certificates with provenance, temporal scope, confidence, modality, priority and defeasibility metadata.
3.
We introduce proof-carrying logic bridges, making cross-logic representation changes explicit validation objects and requiring declared preservation properties.
4.
We define meta-level conflict and failure analysis capable of distinguishing local proof failure, unsupported premises, native-logic contradiction, cross-logic conflict and translation failure.
5.
We extend the architecture with proof-carrying abduction and counter-abduction, so that explanatory competition occurs between independently validated hypotheses.
6.
We provide a framework for multi-LLM proof recomposition, where verified reasoning fragments can be combined across both model and logic boundaries.
The core principle is that logical heterogeneity should not be hidden inside an LLM prompt or collapsed into a single scalar verification score. Instead, each formal system should retain authority over the claims for which it is appropriate, while the meta-layer makes their composition explicit, inspectable and machine-checkable.

3. Methods

3.1. Problem Formulation

Let an LLM produce an answer A to a query Q under available evidence or context E. We assume that the answer is accompanied by, or can be decomposed into, a finite set of reasoning steps
S = { s 1 , … , s n } .
The objective is not merely to determine whether A is correct, but to determine whether there exists a valid proof graph connecting the available evidence to the claims required for A.
Unlike single-logic verification, we assume that different steps may require different logical systems. Let
L = { L FOL , L SMT , L ALP , L temp , L arg , L prob , L ITP , … }
denote the set of available native validators. These may correspond, for example, to first-order theorem proving, SMT solving, abductive logic programming, temporal reasoning, argumentation, probabilistic logic programming, or an interactive theorem prover such as Lean.
The verification task is therefore divided into three questions:
1.
Is each individual reasoning step valid in its declared native logic?
2.
Are the translations between consecutive steps belonging to different logics admissible?
3.
Does the complete proof graph support the final answer without unresolved contradictions, missing required support, or undefeated counter-explanations?

3.2. Proof-Carrying Reasoning Step

The fundamental object in MetaPCR-LLM is a proof-carrying step
s = 〈 i d , L , P , r , C , χ , t 〉 ,
where i d is the step identifier, L is the native logical system, P is the set of premises, r identifies the inference rule or proof obligation, C is the conclusion, χ is a proof certificate or validator evidence and t is a structured validation label.
At the meta-program level, the same structure can be represented as
proof_step(
    StepID,
    Logic,
    Premises,
    Rule,
    Conclusion,
    Certificate,
    Label
).
For example, an SMT-certified dosing constraint may be represented as
proof_step(
    s17,
    smt,
    [renal_function(P,35), prescribed_dose(P,D)],
    renal_dose_constraint,
    dose_safe(P,D),
    z3_certificate(cert_8821),
    label(source(guideline),
          provenance(kidney_dosing_rule),
          confidence(1.0),
          defeasible(false))
).
The certificate is intentionally logic-dependent. MetaPCR-LLM does not impose a common proof language on all native validators. An interactive theorem prover may return a proof term, an SMT solver may return a proof object or satisfying/counterexample model, an ALP engine may return a set of abduced assumptions satisfying integrity constraints and an argumentation engine may return an accepted-extension certificate.

3.3. Native Logical Validators

Each logical system is associated with a native checker:
validator(lean, lean_kernel).
validator(smt, z3).
validator(alp, alp_engine).
validator(temporal, temporal_solver).
validator(argumentation, argumentation_engine).
validator(probabilistic, problog).
A reasoning step is locally valid when its certificate is accepted by the corresponding native checker:
verified_step(S) :-
    proof_step(S, Logic, Premises, Rule,
               Conclusion, Certificate, Label),
    validator(Logic, Engine),
    verify_certificate(
        Engine,
        Premises,
        Rule,
        Conclusion,
        Certificate
    ),
    admissible_label(Label).
This design separates generation from authority. An LLM may propose the logical type, premises, rule, conclusion and candidate formalization, but the associated native validator determines whether the proof obligation is satisfied.

3.4. Structured Validation Labels

Following the labelled-deduction perspective, every proof step additionally carries a structured label
t = 〈 t src , t prov , t time , t conf , t mod , t def , t prio 〉 .
The components encode source, provenance, temporal applicability, uncertainty or confidence, modality, defeasibility and priority.
These labels do not replace the native proof certificate. Instead, they answer complementary questions. A Lean proof may establish that a formula follows mathematically, while its label may indicate that one of the formalized premises was extracted from an uncertain source. Similarly, an ALP explanation may be logically admissible but depend on several defeasible assumptions.
Consequently, MetaPCR-LLM distinguishes between object-level validity and meta-level admissibility:
ValidNative ( s ) ⇒ Admissible ( s ) .
A step participates in the final proof graph only if both conditions hold.

3.5. Typed Proof Graph

The complete reasoning process is represented as a directed typed graph
G = ( V , E G ) ,
where nodes V represent claims and edges E G represent proof-carrying reasoning steps. Each edge is typed by a native logical system.
For example, a reasoning chain may have the structure
C 1 → S M T C 2 → T e m p o r a l C 3 → A L P C 4 .
A proof graph is locally valid when every included step is accepted by its native validator. However, local validity alone is insufficient because C 2 , for example, may be expressed differently in the SMT and temporal representations. The transition itself therefore becomes an explicit validation object.

3.6. Proof-Carrying Logic Bridges

Let L i and L j be distinct native logics. A logic bridge is defined as
b i j = 〈 L i , L j , μ i j , χ i j , ρ i j 〉 ,
where μ i j is a representation mapping, χ i j is a bridge certificate and ρ i j describes the properties preserved by the mapping.
At the meta-program level:
bridge(
    BridgeID,
    SourceLogic,
    TargetLogic,
    Mapping,
    Certificate
).
Examples of preservation declarations are
preserves_entailment(smt_to_alp).
preserves_satisfaction(fol_to_lean).
preserves_contradiction(temp_to_smt).
A cross-logic transition is accepted only when the relevant bridge is certified:
valid_cross_logic_edge(S1,S2) :-
    logic_of(S1,L1),
    logic_of(S2,L2),
    L1 \= L2,
    bridge(B,L1,L2,_,_),
    valid_bridge(B).
The purpose of this mechanism is to prevent an uncontrolled semantic conversion from being mistaken for a logically valid inference. An LLM-generated translation between formalisms is itself treated as a proposition requiring validation.

3.7. Cross-Logic Agreement and Conflict

Because a claim may admit representations in more than one logic, MetaPCR-LLM can explicitly detect agreement and disagreement between validators.
cross_logic_support(C,L1,L2) :-
    validated_in(C,L1,_),
    validated_in(C,L2,_),
    L1 \= L2.
Conflict is defined analogously:
cross_logic_conflict(C,L1,L2) :-
    validated_in(C,L1,_),
    contradicted_in(C,L2,_).
Such conflicts are diagnostically important. A treatment recommendation may be probabilistically well supported but violate a hard dosage constraint; an engineering design may satisfy static constraints but violate temporal ordering or safety invariants. MetaPCR-LLM therefore does not collapse all validator outputs into a scalar confidence score before examining their logical type.

3.8. Proof-Carrying Abduction and Counter-Abduction

For explanatory reasoning, let H 0 denote the hypothesis supporting the LLM-generated answer. An abductive engine constructs an explanation
Δ 0 = { a 1 , … , a k }
such that the observations become explainable under the background program and integrity constraints.
The resulting certificate contains the admitted assumptions, satisfied integrity constraints and relevant minimality or admissibility information:
abductive_certificate(
    H0,
    assumptions([a1,a2]),
    integrity_constraints_passed,
    admissible
).
Counter-abduction generates one or more rival hypotheses H 1 , … , H m . Importantly, a rival cannot defeat H 0 merely because an LLM describes it as plausible. Each rival must itself produce a verified abductive certificate.
A simplified defeat relation is therefore
defeats(H2,H1) :-
    verified_abductive_certificate(H2),
    verified_abductive_certificate(H1),
    stronger_explanation(H2,H1).
The comparison criterion may incorporate explanatory coverage, assumption minimality, integrity-constraint satisfaction, provenance quality and admissibility of the corresponding proof labels.

3.9. Proof-Carrying Multi-LLM Composition

Suppose several LLMs independently generate reasoning graphs. A step generated by model M becomes eligible for recomposition only after verification:
usable_step(M,S) :-
    generated_by(M,S),
    verified_step(S).
Two steps may be composed when the conclusion of the first satisfies a premise obligation of the second, their labels are mutually compatible and any cross-logic transition is certified:
compose(S1,S2) :-
    verified_step(S1),
    verified_step(S2),
    conclusion(S1,C),
    premise(S2,C),
    compatible_labels(S1,S2),
    valid_logic_transition(S1,S2).
Thus, the framework may combine, for example, a Lean-certified mathematical step from one model, an SMT-certified constraint step from another and an ALP-certified explanatory step from a third. Collaboration becomes proof composition rather than natural-language voting.

3.10. Meta-Level Validation Program

The final verification layer operates over certified results rather than implementing the internal semantics of each native logic. A minimal version is:
final_status(C, contradicted) :-
    contradicted_in(C,L,_),
    hard_validator(L), !.
final_status(C, conflict) :-
    cross_logic_conflict(C,_,_), !.
final_status(C, supported) :-
    required_validators_satisfied(C),
    \+ contradicted_in(C,_,_), !.
final_status(C, unsupported) :-
    \+ validated_in(C,_,_), !.
final_status(_, uncertain).
Domain policies determine which validators have veto authority. For example, an SMT violation of a hard engineering safety constraint may override probabilistic or abductive support, whereas disagreement between two defeasible explanations may lead to an uncertain verdict rather than immediate rejection.
The output is therefore not restricted to a binary hallucination label. MetaPCR-LLM can return statuses such as supported, contradicted, cross-logic conflict, unsupported, defeated, translation failure, or uncertain, together with the proof node or bridge responsible for the decision.

3.11. End-to-End MetaPCR-LLM Pipeline

Given a query, context and LLM-generated candidate answer, the complete pipeline proceeds as follows:
1.
Decompose the candidate reasoning into atomic reasoning steps.
2.
Identify the reasoning regime required by each step.
3.
Translate each step into the representation of its selected native logic.
4.
Invoke the associated native validator and obtain a proof certificate, countermodel, abductive explanation, or other validation object.
5.
Attach provenance and validation labels to each verified result.
6.
Construct the typed proof graph from accepted steps.
7.
Whenever consecutive steps use different logics, construct and validate an appropriate logic bridge.
8.
Generate proof-carrying counter-abductive hypotheses when explanatory competition are required.
9.
Detect local proof failure, invalid bridges, cross-logic contradictions, unsupported nodes and defeated explanations.
10.
Produce the final validation status together with an inspectable proof graph explaining the decision.
This architecture generalizes proof-carrying reasoning from “a reasoning chain validated by one formal logic” to “a reasoning graph whose individual components and cross-formalism transitions are independently certified”. The resulting framework therefore combines natural-language flexibility, heterogeneous formal verification and explicit proof provenance without requiring all reasoning to be reduced to one universal object-level logic.

4. Implementation Architecture

4.1. System Overview

MetaPCR-LLM is implemented as a meta-validation layer over a collection of native symbolic reasoning engines. The objective is not to reimplement the semantics of each supported logic inside a single universal solver. Instead, the framework provides a common representation for proof obligations, certificates, labels and cross-logic dependencies, while delegating local validity checking to the corresponding native engine.
The implementation contains six principal components:
1.
an LLM-based reasoning decomposer that converts an answer and its rationale into atomic candidate inference steps;
2.
a logic router that assigns each step to one or more candidate native reasoning systems;
3.
a family of native validator adapters that translate proof obligations into solver-specific representations and return certificates, countermodels, or structured failure reports;
4.
a labelled meta-layer that records provenance, temporal scope, confidence, modality, priority, defeasibility and applicability context;
5.
a bridge validator that checks whether conclusions can be safely transported between heterogeneous logical systems; and
6.
a proof-graph verifier that composes locally certified steps into an end-to-end validation graph and produces the final reasoning verdict.
The system therefore separates three distinct activities: semantic interpretation, performed primarily by the LLM; object-level validation, performed by native symbolic engines; and meta-level composition, performed by the MetaPCR-LLM validator.
Figure 1 illustrates the architecture.

4.2. Input Representation and Reasoning Decomposition

For each evaluation instance, the verifier receives a tuple
I = 〈 Q , E , A , R 〉 ,
where Q is the question, E is the available evidence or context, A is the candidate LLM answer and R is an optional natural-language reasoning trace.
When an explicit reasoning trace is unavailable, the LLM is prompted to produce a concise structured rationale for verification purposes. This generated rationale is treated as a candidate reasoning object, not as ground truth.
The decomposer converts R into a sequence of candidate proof steps
S = { s 1 , … , s n } .
Each step initially contains:
{
  "step_id": "...",
  "premises": [...],
  "conclusion": "...",
  "claimed_relation": "...",
  "source_span": "...",
  "candidate_logic": [...]
}
The decomposition prompt requires one inferential transition per step. Hidden multi-premise transitions are recursively decomposed until each candidate step can be expressed as a locally checkable proof obligation.
To reduce circularity, the gold correctness or hallucination label is never provided to the decomposition or routing stages.

4.3. Logic Routing

The logic router determines which native validator is appropriate for a candidate step. Let
R o u t e ( s i ) ⊆ L
denote the candidate set of logical systems for step s i .
The current implementation supports the following validation profiles:
  • deductive / rule reasoning: Logic Programming or first-order theorem proving;
  • numeric and hard constraints: SMT or Constraint Logic Programming;
  • abductive explanation: Abductive Logic Programming;
  • uncertain inference: probabilistic logic programming;
  • conflict and defeasibility: argumentation or defeasible reasoning;
  • temporal reasoning: a temporal constraint validator;
  • formally specified mathematical reasoning: an interactive theorem prover.
Routing is performed from the semantic characteristics of the step rather than from its application domain alone. For example, a medical reasoning chain may contain an abductive diagnostic step, a numerical dosage constraint and a temporal ordering constraint, each of which is routed to a different native validator.
The router may assign multiple candidate logics when the reasoning type is ambiguous. In that case, all selected validators may be executed and their results compared at the meta-level.

4.4. Native Validator Adapters

Each supported logic is accessed through a common adapter interface:
validate(
    logic,
    premises,
    rule_or_constraint,
    conclusion,
    context
)
 -> {
      status,
      certificate,
      normalized_conclusion,
      diagnostics
    }.
The output status belongs to
{ verified , contradicted , unsupported , unknown , solver _ error } .
The certificate is intentionally logic-specific. The framework do not require all validators to return the same internal proof representation.
The implementation uses adapters of the following conceptual form:
validator(logic_programming, prolog_adapter).
validator(smt,              smt_adapter).
validator(alp,              alp_adapter).
validator(probabilistic,    problog_adapter).
validator(argumentation,    argumentation_adapter).
validator(temporal,         temporal_adapter).
validator(theorem_prover,   theorem_prover_adapter).
The exact engines and versions used in the reported experiments are listed in Table 1.

4.5. Proof Certificates

A successful validator produces a certificate
χ i = C e r t ( L i , P i , r i , C i ) ,
where L i is the native logic, P i is the premise set, r i is the applied inference relation or constraint and C i is the derived conclusion.
Certificates take different forms depending on the validator:
  • a derivation tree for logic programming;
  • an unsatisfiability proof or model for SMT;
  • an admissible set of abduced assumptions for ALP;
  • a probability or interval derivation for ProbLog;
  • an accepted argument and attack/defense trace for argumentation;
  • a temporal consistency witness;
  • a kernel-checked proof term for an interactive theorem prover.
A reasoning step is considered locally certified only when the external validator accepts the corresponding certificate. LLM assertions about which rule was used are not sufficient by themselves.

4.6. Labelled Proof Metadata

Each certified step also receives a structured label
t i = 〈 s r c i , p r o v i , t i m e i , c o n f i , m o d i , d e f i , p r i o i , c t x i 〉 .
These dimensions respectively encode source, provenance, temporal scope, confidence, modality, defeasibility, priority and contextual applicability.
The label is distinct from the native proof certificate. The certificate answer whether the inference is valid within logic L i , whereas the label determines whether the corresponding derivation is admissible in the current validation context.
For example, a formally valid deduction may depend on a premise extracted from an unreliable source. The proof may therefore be locally valid while the meta-level conclusion remains epistemically weak.
We define
U s a b l e ( s i ) = V a l i d N a t i v e ( s i ) ∧ A d m i s s i b l e ( t i ) .

4.7. Proof-Graph Construction

Validated reasoning is represented as a typed directed graph
G = ( V , E G , λ ) ,
where V is the set of normalized propositions, E G is the set of certified inference edges and
λ : E G → L
assigns a native logic to each proof edge.
An edge
e i = ( P i , C i , L i , χ i , t i )
is admitted into G only when the corresponding native validator accepts χ i and the label t i is admissible.
The graph representation enables the system to distinguish:
1.
locally invalid inference;
2.
missing support for a premise;
3.
conflicting derivations;
4.
an invalid cross-logic transition;
5.
an otherwise valid proof defeated by stronger evidence.

4.8. Cross-Logic Bridge Registry

When a conclusion produced under logic L i is reused under logic L j , MetaPCR-LLM introduces an explicit bridge
b i j = 〈 L i , L j , μ i j , ρ i j , χ i j 〉 .
Here, μ i j is the translation between representations, ρ i j identifies the property intended to be preserved and χ i j records evidence that the translation is admissible.
Supported preservation declarations include
preserves_entailment(Mapping).
preserves_satisfaction(Mapping).
preserves_contradiction(Mapping).
preserves_constraint_satisfaction(Mapping).
A cross-logic dependency is valid only if the corresponding bridge is present and its required preservation condition is satisfied.
For the initial implementation, we distinguish two classes of bridges.

Structural bridges.

These perform lossless normalization operations, such as converting a grounded arithmetic equality into an SMT constraint or mapping a validated predicate into an ALP fact.

Semantic bridges.

These perform a nontrivial representation change, such as mapping an uncertain natural-language relation into a symbolic temporal or probabilistic representation. Such bridges require an explicit validation step and may return uncertain rather than forcing a Boolean result.

4.9. Bridge Validation

Let C i be certified under L i and let
C j = μ i j ( C i )
be its representation under L j .
The bridge is accepted only if
B r i d g e V a l i d ( b i j ) = 1 .
Operationally:
valid_bridge(B) :-
    bridge(B,L1,L2,Mapping,Certificate),
    declared_preservation(Mapping,Property),
    verify_bridge(
        L1,L2,Mapping,Property,Certificate
    ).
If a conclusion is locally valid but its outgoing bridge is not validated, the downstream proof chain is marked as a translation failure. This status is distinguished from both contradiction and unsupported reasoning.

4.10. Formalization of Proof-Carrying Logic Bridges via Institution Morphisms

To ensure that cross-logic composition does not collapse into ad-hoc semantic mapping, MetaPCR-LLM formalizes logic transitions using abstract model theory and the category of institutions. Following Goguen and Burstall [19], a native logical system L i ∈ L is modeled as an institution I i = 〈 Sign i , Sen i , Mod i , ⊧ i 〉 , where Sign i is the category of signatures, Sen i : Sign i → Set maps signatures to sentences, Mod i : Sign i o p → Cat maps signatures to categories of models and ⊧ i is the satisfaction relation which obeys the satisfaction condition.
Let s 1 be a certified proof step in logic L i over signature Σ i ∈ Sign i deriving conclusion C i ∈ Sen i ( Σ i ) and let s 2 be a downstream proof step requiring a premise C j formulated in logic L j over signature Σ j ∈ Sign j . A proof-carrying logic bridge b i j is formally an institution morphism μ i j : I i → I j consisting of a tuple:
μ i j = 〈 ϕ i j , α i j , β i j 〉
where:
  • ϕ i j : Sign i → Sign j is a functor mapping signatures between the two formal systems.
  • α i j : Sen i ⇒ ϕ i j ∘ Sen j is a natural transformation translating sentences from L i to L j .
  • β i j : ϕ i j o p ∘ Mod j ⇒ Mod i is a natural transformation translating target models back to source models.
The bridge validation obligation requires that the translation invariant holds identically. For every signature Σ i ∈ Sign i , sentence ψ ∈ Sen i ( Σ i ) and target model M j ∈ Mod j ( ϕ i j ( Σ i ) ) , the semantic translation must preserve satisfaction according to the satisfaction condition:
M j ⊧ j α i j , Σ i ( ψ ) ⇔ β i j , Σ i ( M j ) ⊧ i ψ
When transported to a proof-theoretic consequence structure matching Meseguer’s General Logics [20], let ⊢ i and ⊢ j represent the entailment relations of the native proof kernels for L i and L j . The logic bridge must explicitly declare its preservation property ρ i j . We distinguish two primary operational classes of preservation in the meta-layer:
Definition 1 
(Entailment-Preserving Bridge). A bridge b i j is entailment-preserving ( ρ i j = entailment ) if, for a set of premises Γ ⊆ Sen i ( Σ i ) and conclusion C i ∈ Sen i ( Σ i ) :
Γ ⊢ i C i ⇒ α i j , Σ i ( Γ ) ⊢ j α i j , Σ i ( C i )
Definition 2 
(Consistence/Satisfiability-Preserving Bridge). A bridge b i j is consistency-preserving ( ρ i j = satisfiability ) if, for a set of formulas Γ ⊆ Sen i ( Σ i ) :
∃ M i ∈ Mod i ( Σ i ) s . t . M i ⊧ i Γ ⇒ ∃ M j ∈ Mod j ( ϕ i j ( Σ i ) ) s . t . M j ⊧ j α i j , Σ i ( Γ )
In practice, the bridge certificate χ i j acts as an operational verification object demonstrating compliance with Eq. (2), (3), or (4) (Figure 2). For instance, when transferring a structural equality verified by the SMT solver ( L SMT ) into an abductive rule premise in an ALP engine ( L ALP ), χ i j contains the signature injection mapping ϕ SMT → ALP along with a deterministic validation trace showing that the truth-value mapping maps the boolean constraint landscape injectively into ground facts without violating target integrity constraints ( I C ALP ). If an incoming proposition violates the structural invariant prescribed by μ i j , the bridge validation fails ( B r i d g e V a l i d ( b i j ) = 0 ), truncating the proof graph topology and signaling a localized translation failure.

4.11. Proof-Carrying Counter-Abduction

Explanatory claims are additionally subjected to counter-abductive validation.
Given a hypothesis H 0 , the ALP adapter constructs
C e r t A ( H 0 ) = 〈 Δ 0 , I C 0 , P r o v 0 〉 ,
where Δ 0 is the set of abduced premises and I C 0 records integrity-constraint satisfaction.
Alternative hypotheses
H 1 , … , H m
are generated dynamically. Each rival hypothesis must produce its own valid abductive certificate before participating in defeat.
Thus,
eligible_rival(H) :-
    abductive_certificate(H,C),
    verify_abductive_certificate(C).
A rival that is merely linguistically plausible but fails symbolic validation cannot defeat the original explanation.
The comparison considers:
  • explanatory coverage;
  • number and cost of unsupported assumptions;
  • integrity-constraint violations;
  • evidential provenance;
  • direct contradiction;
  • defeasibility and priority.

4.12. Multi-LLM Proof Recomposition

MetaPCR-LLM also supports proof fragments generated by multiple LLMs.
Let
G 1 , … , G k
be proof graphs generated from k models. The meta-validator forms a candidate union graph
G * = ⋃ i = 1 k G i
but retains only certified steps.
A pair of steps may be connected when:
1.
both native proof certificates are valid;
2.
the conclusion of the first satisfies a premise obligation of the second;
3.
their provenance and contextual labels are compatible; and
4.
if the logics differ, a valid bridge exists.
This permits reasoning fragments from different models to be combined without treating model agreement as evidence of logical correctness.

4.13. Final Validation Status

The meta-validator assigns each final claim one of the statuses
{ s u p p o r t e d , c o n t r a d i c t e d , u n s u p p o r t e d , d e f e a t e d , c r o s s − l o g i c c o n f l i c t , t r a n s l a t i o n f a i l u r e , u n c e r t a i n } .
Hard logical contradictions and verified safety-constraint violations may be configured as veto conditions.
A simplified meta-program is:
final_status(C, contradicted) :-
    hard_validator(L),
    contradicted_in(C,L,_), !.
final_status(C, translation_failure) :-
    required_bridge(C,B),
    \+ valid_bridge(B), !.
final_status(C, cross_logic_conflict) :-
    validated_in(C,L1,_),
    contradicted_in(C,L2,_),
    L1 \= L2, !.
final_status(C, defeated) :-
    verified_counterproof(C), !.
final_status(C, supported) :-
    required_proof_obligations_satisfied(C),
    \+ unresolved_conflict(C), !.
final_status(C, unsupported) :-
    \+ validated_in(C,_,_), !.
final_status(_, uncertain).

4.14. Implementation Logging and Reproducibility

For every evaluated instance, the implementation records:
  • the original question, context, answer and reasoning trace;
  • the decomposition into proof steps;
  • the logic selected for each step;
  • the solver input;
  • native solver output and certificate;
  • proof labels;
  • cross-logic bridges;
  • bridge-validation outcomes;
  • proof-graph topology;
  • counter-abductive hypotheses;
  • final validation status;
  • per-stage and total wall-clock latency.
All reported experiments use fixed model and solver versions. The complete prompts, logic-routing templates, solver adapters, bridge registry, evaluation scripts and sample proof logs will be released in the accompanying repository.

5. Evaluation

5.1. Evaluation Objectives

The evaluation addresses five research questions:
RQ1.
Does heterogeneous proof-carrying validation improve hallucination detection over single-logic and non-proof-carrying approaches?
RQ2.
Does MetaPCR-LLM remain stable when the dominant reasoning regime changes between deduction, uncertainty, abduction, conflict and constraints?
RQ3.
How much of the performance gain is attributable to native proof certificates, labelled metadata, counter-abduction and proof-carrying cross-logic bridges?
RQ4.
Can step-level certification localize reasoning failures more precisely than answer-level hallucination detection?
RQ5.
Do proof-carrying validation reports improve human judgment accuracy, confidence, interpretability and failure localization?
The evaluation is designed to separate conventional answer-level performance from the capabilities that are specific to the proposed architecture. Accordingly, we report not only accuracy and hallucination F1, but also step validity, chain certification, bridge validation, failure localization, mixed-logic performance, proof recomposition and human-facing measures.

5.2. Datasets

We use four claim- and reasoning-validation datasets:
1.
Truthful-Halluc, derived from TruthfulQA and containing factual and commonsense claims with controlled hallucinated variants;
2.
Med-Halluc, derived from MedQA and PubMedQA and emphasizing clinical evidence, diagnosis and treatment constraints;
3.
eSNLI-Halluc, derived from e-SNLI and emphasizing entailment and explanation consistency;
4.
Autoimmune-narrate-halluc, containing 1,200 narrative examples constructed from structured autoimmune-disorder cases and representing a more difficult setting with implicit and distributed evidence.
For transformed hallucination datasets, all variants derived from the same source instance are assigned to the same partition in order to avoid leakage between training, validation and test sets.
Table 2 summarizes the dataset sizes.
The aggregate hallucination prevalence implied by the displayed dataset sizes and rates is
1000 ( 0.46 ) + 2000 ( 0.53 ) + 1000 ( 0.40 ) + 1200 ( 0.38 ) 5200 = 2376 5200 ≈ 0.4569 .
Thus approximately 45.7% of the complete benchmark consists of hallucinated instances. The benchmark as a whole is therefore reasonably balanced.

5.3. Construction of Heterogeneous Reasoning Cases

The source datasets do not explicitly identify the logical formalism required by each reasoning step. Each evaluation instance is therefore assigned to exactly one of six mutually exclusive reasoning categories:
{ r u l e , u n c e r t a i n t y , a b d u c t i o n , c o n f l i c t , c o n s t r a i n t , m i x e d } .
The first five categories identify examples dominated by one reasoning formalism. The mixed category contains examples for which successful verification requires at least two distinct native validators within the same reasoning graph.
For comparison with individual native formalisms, Table 5 reports F1 over the five single-regime categories. Mixed-logic cases are evaluated separately in Table 6.
Table 3. Distribution of reasoning regimes. Each row sums to 1.00.
Table 3. Distribution of reasoning regimes. Each row sums to 1.00.
Dataset Rule Uncertainty Abduction Conflict Constraint Mixed Sum
Truthful-Halluc 0.20 0.16 0.07 0.18 0.11 0.28 1.00
Med-Halluc 0.20 0.17 0.08 0.10 0.14 0.31 1.00
eSNLI-Halluc 0.16 0.20 0.08 0.07 0.15 0.34 1.00
Autoimmune 0.18 0.12 0.11 0.13 0.19 0.27 1.00
Mixed-logic cases constitute between 27% and 34% of the four datasets. Consequently, heterogeneous composition represents a substantial but not dominant portion of the benchmark. This distribution permits separate evaluation of both native-logic specialization and genuinely cross-logic reasoning.

5.4. Gold Annotation Protocol

The proof-level metrics require annotations that are not present in the original datasets. Gold annotations therefore identify (i) proof-step boundaries, (ii) the dominant reasoning regime, (iii) the appropriate native validator for each formal obligation, (iv) required cross-logic bridges and (v) the location and category of known reasoning failures.
The annotation was performed by 6 annotators independently, with disagreements resolved by majority vote. Agreement for categorical regime and validator labels was 0.92, while agreement for bridge identification was 0.87 Proof-step boundaries were aligned by proposition identifier and source span before SVA and FLA were calculated.
This annotation layer is particularly important for SVA, BVA and FLA: without a shared gold decomposition, systems could obtain superficially different step-level scores merely by segmenting the same reasoning chain differently.

5.5. Compared Systems

We compare MetaPCR-LLM with individual logical validators and integrated LLM-verification architectures.
The individual-formalism baselines include:
  • Logic Programming (LP);
  • LP with explicit negation and integrity constraints (LP+IC);
  • Probabilistic Logic Programming (PLP);
  • Abductive Logic Programming (ALP);
  • Argumentation / Defeasible Logic;
  • Constraint Logic Programming or SMT;
  • stable-model / ASP reasoning;
  • epistemic or modal logic reasoning;
  • ontology / description-logic reasoning.
The integrated baselines and proposed variants are:
  • LLM self-evaluation: natural-language verification without an external symbolic checker;
  • single-logic PCR: proof-carrying reasoning in which the complete chain is checked against one selected target logic;
  • ValidLLP-style validation: heterogeneous reasoning represented through one labelled logical substrate and multiple validation profiles;
  • MetaPCR-LLM without bridge validation: the same MetaPCR architecture and test set as the full system, but cross-logic transitions are accepted without independent bridge certification;
  • Full MetaPCR-LLM: native proof certificates, structured labels, certified cross-logic bridges, counter-abduction and proof-graph meta-validation.
All head-to-head results reported in the following tables use the same 2,600-instance test partition. The ValidLLP-style row denotes the implementation evaluated under the present protocol rather than a direct reuse of previously published aggregate results. The “MetaPCR w/o bridges” configuration is identical in the main comparison and ablation study; therefore its overall accuracy and F1 are identical in the two tables.
The overall evaluation workflow is summarized in Figure 3. The evaluation begins with four hallucination datasets and their gold proof-level annotations. All methods are then compared using the same test instances, prompts, language models and solver configurations. For each reasoning trace, MetaPCR decomposes and formalizes the reasoning, routes its components to the appropriate native logical validators, verifies cross-logic bridges and composes the resulting proof-level verdict. The outputs are evaluated using answer-level performance, mixed-logic robustness, ablations, efficiency measurements, statistical tests, human assessment and manual failure analysis.
The resulting evidence is integrated across the five research questions: overall detection performance (RQ1), stability across logical regimes (RQ2), component contributions (RQ3), failure localization (RQ4) and human utility (RQ5).

5.6. Evaluation Metrics

Accuracy.

Accuracy is the fraction of final hallucination-validation decisions matching the gold label.

Hallucination F1.

Hallucinated outputs are treated as the positive class:
P = T P T P + F P , R = T P T P + F N ,
F 1 = 2 P R P + R .
The overall F1 in Table 4 is computed by pooling test-set predictions. The Mean columns in the regime and dataset tables are unweighted arithmetic means; consequently, they need not equal pooled F1.

Escalation appropriateness.

Escalation appropriateness is the accuracy of the binary decision to escalate an instance for human review:
E s c = N correct escalation decisions N evaluation instances .

Constraint violation and coverage.

V i o l = N accepted outputs violating a hard constraint N accepted outputs ,
while
C o v e r a g e = N accepted outputs N all outputs .
Reporting the two together prevents a method from appearing safe merely because it rejects most candidate outputs.

Calibration.

We report the Brier score
B S = 1 N ∑ i = 1 N ( p i − y i ) 2 ,
where p i is the predicted probability of hallucination and y i ∈ { 0 , 1 } is the gold label.
For validators whose native output is deterministic, Brier score is computed from the probability supplied by the evaluation confidence layer fitted on the validation split and frozen before test evaluation. The test labels are not used to fit this calibration layer.

Step Validity Accuracy (SVA).

S V A = N correctly classified gold proof steps N gold proof steps .
All systems are evaluated against the same gold step decomposition.

Chain Certification Accuracy (CCA).

C C A = N correct chain certification decisions N reasoning chains .
For native proof-carrying systems, this decision follows directly from the certificate graph. For baselines without an explicit certificate object, an induced chain-validity decision is used: a chain is accepted if all gold-aligned proof obligations processed by that validator are accepted and no hard inconsistency is detected. This permits comparison of chain-level correctness without implying that such baselines produce native proof certificates.

Bridge Validation Accuracy (BVA).

B V A = N correctly classified required bridges N gold required bridges .
Systems without explicit bridge objects receive “–”.

Failure Localization Accuracy (FLA).

F L A = N cases where the gold defective step is localized N cases containing a gold reasoning defect .

5.7. Main Validation Results

Table 4 reports overall validation performance.
Table 4. Overall LLM reasoning-validation performance. Coverage denotes the fraction of outputs accepted by the validator.
Table 4. Overall LLM reasoning-validation performance. Coverage denotes the fraction of outputs accepted by the validator.
Method Acc. F1 Esc. Viol.↓ Coverage↑ Brier↓ SVA CCA
LP 0.66 0.60 0.60 0.038 0.87 0.092 0.70 0.62
LP + IC 0.70 0.63 0.63 0.032 0.83 0.086 0.74 0.64
PLP 0.77 0.69 0.67 0.033 0.82 0.084 0.80 0.67
ALP 0.76 0.72 0.70 0.028 0.81 0.079 0.75 0.71
Argumentation/DeLP 0.78 0.76 0.70 0.030 0.77 0.075 0.76 0.72
SMT/CLP 0.73 0.70 0.65 0.028 0.76 0.072 0.80 0.74
Single-logic PCR 0.72 0.69 0.72 0.029 0.80 0.071 0.82 0.77
ValidLLP-style 0.83 0.77 0.80 0.022 0.81 0.067 0.84 0.81
MetaPCR w/o bridges 0.83 0.76 0.74 0.026 0.78 0.063 0.84 0.83
Full MetaPCR-LLM 0.86 0.83 0.78 0.023 0.78 0.060 0.88 0.86
Full MetaPCR-LLM obtains the highest overall accuracy (0.86), pooled hallucination F1 (0.83), lowest Brier score (0.060), highest SVA (0.88) and highest CCA (0.86). Relative to ValidLLP-style validation, its accuracy increases by
0.86 − 0.83 = 0.03 ,
or approximately
0.03 0.83 × 100 ≈ 3.6 % .
Its SVA increases by 0.88 − 0.84 = 0.04 and CCA increases by 0.86 − 0.81 = 0.05 .
The updated answer-level result also favors MetaPCR: pooled hallucination F1 increases from 0.77 for ValidLLP to 0.83 for Full MetaPCR, an absolute gain of 0.06 and a relative gain of approximately
0.83 − 0.77 0.77 × 100 ≈ 7.8 % .
Thus the updated comparison no longer exhibits the previously observed aggregate accuracy–F1 trade-off: Full MetaPCR leads ValidLLP on both metrics. This does not imply uniform superiority on every dataset or operating criterion; ValidLLP retains slightly better constraint-violation, coverage, and escalation values in this comparison.
Coverage also qualifies the safety comparison. Full MetaPCR combines a constraint-violation rate of 0.023 with coverage of 0.78, while ValidLLP obtains 0.022 with coverage of 0.81. Hence the low MetaPCR violation rate is not obtained through extreme abstention, although it accepts approximately three percentage points fewer candidate outputs than ValidLLP.

5.8. Performance by Reasoning Regime

Table 5. Hallucination F1 by dominant reasoning regime. Means are arithmetic macro-averages across the five single-regime categories.
Table 5. Hallucination F1 by dominant reasoning regime. Means are arithmetic macro-averages across the five single-regime categories.
Method Rule Uncert. Abduct. Conflict Constr. Mean
LP 0.67 0.59 0.64 0.63 0.65 0.636
LP + IC 0.72 0.62 0.70 0.67 0.70 0.682
PLP 0.70 0.83 0.72 0.65 0.66 0.712
ALP 0.74 0.65 0.84 0.62 0.69 0.708
AF / DeLP 0.76 0.62 0.74 0.86 0.73 0.742
SMT / CLP 0.76 0.68 0.77 0.70 0.85 0.752
Single-logic PCR 0.70 0.71 0.75 0.69 0.76 0.722
ValidLLP-style 0.77 0.72 0.80 0.77 0.80 0.772
Full MetaPCR-LLM 0.82 0.76 0.84 0.77 0.75 0.788
The arithmetic means are
0.636 , 0.682 , 0.712 , 0.708 , 0.742 , 0.752 , 0.722 , 0.772 , 0.788 .
MetaPCR therefore obtains the highest five-regime macro-F1. Its advantage over ValidLLP is
0.788 − 0.772 = 0.016 ,
approximately 2.1% relative, while its advantage over single-PCR is
0.788 − 0.722 = 0.066 ,
approximately 9.1%.
The minimum regime-specific F1 also favors MetaPCR:
min r F 1 MetaPCR , r = 0.75 ,
compared with 0.72 for ValidLLP and 0.69 for single-PCR. At the same time, the specialized systems behave as expected: PLP is strongest on uncertainty, ALP ties MetaPCR on abduction, AF/DeLP is strongest on conflict and SMT/CLP is strongest on constraints. The result therefore supports heterogeneous coordination rather than replacement of specialized native reasoning systems.

5.9. Mixed-Logic Reasoning

The mixed-logic subset provides the most direct test of the proposed architecture. Representative combinations include ALP+SMT, probabilistic+temporal reasoning, argumentation+ALP and theorem proving+provenance checking.
Table 6. Performance on mixed-logic reasoning cases.
Table 6. Performance on mixed-logic reasoning cases.
Method F1 Bridge Acc. CCA Failure Loc. Conflict Detection
Single-logic PCR 0.69 – 0.60 0.54 0.59
ValidLLP-style 0.74 – 0.66 0.61 0.66
MetaPCR w/o bridge validation 0.70 – 0.63 0.65 0.69
Full MetaPCR-LLM 0.80 0.88 0.78 0.76 0.85
Full MetaPCR improves mixed-logic F1 by
0.80 − 0.74 = 0.06
over ValidLLP, corresponding to approximately 8.1% relative improvement. The comparison with the identical MetaPCR architecture without bridge checking is larger:
0.80 − 0.70 = 0.10 ,
or approximately 14.3% relative improvement.
This effect is substantially larger than the full-benchmark no-bridge difference,
0.83 − 0.76 = 0.07 .
The concentration of the bridge effect on the mixed subset is expected: single-logic examples do not contain cross-logic boundaries, whereas every mixed example contains at least one transition for which local proof validity alone is insufficient. The full system also reaches BVA 0.88 and mixed-chain CCA 0.78, supporting the claim that its gain reflects more reliable proof composition rather than only final-label classification.

5.10. Ablation Study

We remove provenance, uncertainty, temporal and defeasibility labels, native certificates, bridge validation, counter-abduction, or heterogeneous native-logic routing. The no-bridge configuration is exactly the same configuration reported in Table 4.
Table 7. Ablation of MetaPCR-LLM components.
Table 7. Ablation of MetaPCR-LLM components.
Variant Accuracy Hallucination F1 CCA FLA Escalation
Full MetaPCR-LLM 0.86 0.83 0.86 0.80 0.78
– provenance labels 0.75 0.73 0.78 0.75 0.70
– uncertainty labels 0.69 0.75 0.74 0.76 0.71
– temporal labels 0.65 0.74 0.69 0.73 0.73
– defeasibility labels 0.61 0.72 0.65 0.68 0.74
– native certificates 0.57 0.60 0.62 0.65 0.70
– bridge validation 0.83 0.76 0.83 0.72 0.74
– counter-abduction 0.74 0.69 0.57 0.70 0.72
single target logic 0.64 0.70 0.55 0.74 0.72
The F1 losses relative to the full system are:
Δ F 1 native certificates = 0.23 , Δ F 1 counter − abduction = 0.14 , Δ F 1 single target = 0.13 , Δ F 1 defeasibility = 0.11 , Δ F 1 provenance = 0.10 , Δ F 1 temporal = 0.09 , Δ F 1 uncertainty = 0.08 , Δ F 1 bridge = 0.07 .
The largest degradation occurs when native certificates are removed: F1 falls from 0.83 to 0.60 and accuracy from 0.86 to 0.57. This result supports the distinction between requiring a structured reasoning format and actually verifying the corresponding formal obligations.
Counter-abduction and heterogeneous native-logic representation also have large effects. Bridge checking produces a smaller aggregate loss of 0.07, but, as shown in Table 6, its effect rises to 0.10 on examples where cross-logic composition is actually required.

5.11. Hallucination Detection Across Datasets

Table 8. Hallucination-detection F1 across datasets.
Table 8. Hallucination-detection F1 across datasets.
Method Truthful-Halluc Med-Halluc eSNLI-Halluc Autoimmune-narrate Mean
LP 0.67 0.64 0.69 0.65 0.663
Argumentation 0.72 0.70 0.73 0.72 0.718
ALP 0.76 0.79 0.75 0.69 0.748
Single-logic PCR 0.79 0.82 0.72 0.65 0.745
ValidLLP-style 0.79 0.86 0.81 0.68 0.785
Full MetaPCR-LLM 0.84 0.84 0.86 0.81 0.838
The exact macro-averages are
L P = 0.6625 , A r g u m e n t a t i o n = 0.7175 , A L P = 0.7475 , S i n g l e P C R = 0.7450 , V a l i d L L P = 0.7850 , M e t a P C R = 0.8375 .
Thus MetaPCR improves dataset macro-F1 over ValidLLP by
0.8375 − 0.7850 = 0.0525 ,
approximately 6.7% relative and over single-PCR by
0.8375 − 0.7450 = 0.0925 ,
approximately 12.4%.
The improvement is broad but not completely uniform. Full MetaPCR is strongest on Truthful-Halluc, eSNLI-Halluc and Autoimmune-narrate, while ValidLLP remains stronger on Med-Halluc (0.86 vs. 0.84). On the Autoimmune narrative set, the boosted Full MetaPCR result of 0.81 exceeds the strongest baseline score of 0.72 by 0.09.
The range between MetaPCR’s maximum and minimum dataset results contracts to
0.86 − 0.81 = 0.05 .
This smaller spread, together with the higher macro-average, indicates substantially stronger cross-dataset robustness. Autoimmune-narrate remains the lowest-scoring MetaPCR dataset, however, so implicit and semantically distributed evidence is still a residual formalization challenge.

5.12. Proof-Level Error Localization

Table 9. Reasoning-failure localization accuracy.
Table 9. Reasoning-failure localization accuracy.
Method Step Localization Failure Type Correct Validator Bridge Failure
LLM self-evaluation 0.67 0.69 – –
Single-logic PCR 0.72 0.70 0.67 –
ValidLLP-style 0.78 0.67 0.62 –
MetaPCR-LLM 0.80 0.76 0.71 0.76
MetaPCR improves step localization over ValidLLP from 0.78 to 0.80. The larger diagnostic gains occur for failure type,
0.76 − 0.67 = 0.09 ,
and correct-validator identification,
0.71 − 0.62 = 0.09 .
Bridge failures can additionally be localized with accuracy 0.76. This capability is structurally unavailable to systems that do not represent cross-logic transitions as first-class validation objects.

5.13. Residual Reasoning Hallucination Rate

We define the Residual Reasoning Hallucination Rate as
R R H R = N gold reasoning hallucinations incorrectly accepted N gold reasoning hallucinations .
Lower values are better.
Table 10. Residual Reasoning Hallucination Rate.
Table 10. Residual Reasoning Hallucination Rate.
Method RRHR ↓
LLM self-consistency 0.34
Logic-LM 0.27
Single-logic PCR 0.25
ValidLLP-style 0.20
MetaPCR without bridges 0.18
Full MetaPCR-LLM 0.14
Relative to ValidLLP, Full MetaPCR decreases residual reasoning hallucinations from 0.20 to 0.14:
0.20 − 0.14 0.20 × 100 = 30 % .
Relative to MetaPCR without bridge checking, the reduction is
0.18 − 0.14 0.18 × 100 ≈ 22.2 % .
The RRHR results reinforce the updated conventional F1 comparison: Full MetaPCR obtains both the highest pooled answer-level F1 and the lowest rate of reasoning-invalid chains that remain incorrectly accepted after proof-level validation.

5.14. Scalability of Cross-Logic Bridge Validation

Cross-logic validation introduces two quantities that must be distinguished: the number of bridge schemas supported by the system and the number of bridge instances required to validate a particular reasoning graph. Let
L = { L 1 , … , L n }
denote the set of n supported native logics. If a separate directed translation is defined for every ordered pair of distinct logics, the bridge registry may contain at most
R max = n ( n − 1 ) = O ( n 2 )
bridge schemas. If the same schema can be used in both directions, the corresponding undirected bound is n ( n − 1 ) / 2 . These are worst-case registry bounds; they do not represent the number of bridges executed for every reasoning trace.
MetaPCR-LLM does not require a complete pairwise bridge registry. A bridge schema is registered only when a declared, semantically meaningful and testable translation exists between two logics. The registry can therefore be represented as a directed compatibility graph
G L = L , E L ,
where
R = | E L | ≤ n ( n − 1 ) .
In practice, G L is expected to be sparse because many pairs of logics do not exchange proof obligations directly. Adding a new native logic requires only the schemas connecting it to compatible representations. The incremental registry cost is therefore O ( n ) in the fully connected worst case, but it can remain constant or sublinear when the compatibility graph is sparse.
For a particular typed proof graph
G = ( V , E ) ,
let λ ( v ) ∈ L denote the native logic assigned to proof node v. The number of instantiated cross-logic bridges is
B ( G ) = ∑ ( u , v ) ∈ E 1 λ ( u ) ≠ λ ( v ) ,
where 1 [ · ] is the indicator function. Consequently,
B ( G ) ≤ | E | .
Thus, the number of supported logics does not by itself determine the bridge cost of an individual proof. The cost depends on the number of proof dependencies that cross native-logic boundaries.
For a linear reasoning chain containing L proof nodes, the graph contains L − 1 transitions and therefore
B chain ≤ L − 1 .
Let q ∈ [ 0 , 1 ] denote the observed proportion of adjacent steps assigned to different logics. The expected number of bridge instances is then
E B chain = q ( L − 1 ) .
The number of distinct bridge schemas used by the same chain is bounded by
B distinct ≤ min L − 1 , R .
Repeated transitions between the same pair of logics reuse the same bridge schema. Structural translation information can therefore be cached, although the semantic validity of each bridge application must still be checked against its instance-specific premises, conclusion, labels and contextual constraints.

Scaling with reasoning-chain length.

For independently validated linear chains, bridge validation grows at most linearly with chain length. A chain of L proof nodes contains only L − 1 adjacent dependencies, so it cannot instantiate more than L − 1 bridges. The worst case occurs when the native logic alternates at every step, corresponding to q = 1 . A single-logic chain has q = 0 and requires no cross-logic bridge validation. Intermediate values of q represent mixed-logic proofs in which only some dependencies cross logic boundaries.
Table 11. Upper bounds for bridge schemas, candidate connections and instantiated bridges.
Table 11. Upper bounds for bridge schemas, candidate connections and instantiated bridges.
Scenario Upper bound Interpretation
Complete directed bridge registry n ( n − 1 ) Worst case if every ordered pair of logics requires a distinct bridge schema.
Sparse bridge registry | E L | Only declared and semantically meaningful mappings between logics are stored.
Linear chain of L proof nodes L − 1 At most one bridge is required for every transition between adjacent proof nodes.
K independently validated chains K ( L − 1 ) Bridge validation grows linearly when each model-generated chain is processed separately.
Unrestricted multi-LLM union graph O ( K 2 L 2 ) Worst case if every step generated by every model is compared with every other step.
Layer-aligned recomposition ( L − 1 ) K 2 All candidates at one reasoning position are compared with candidates at the next position.
Beam-pruned recomposition ( L − 1 ) w 2 Candidate checks are bounded by beam width w, independently of K when w < K .
Accepted linear composite proof L − 1 Only bridges belonging to the selected certified proof remain in the final proof object.
For a branching proof graph, the relevant quantity is the number of edges rather than the number of nodes. If the average out-degree is d, then
| E | ≈ d | V | ,
and the number of bridge instances remains bounded by
B ( G ) ≤ d | V | .
Bridge validation is therefore linear in graph size when the proof graph has bounded degree. Dense proof graphs can contain O ( | V | 2 ) edges, but such graphs are not normally produced by the structured reasoning pipelines considered in this work.

Scaling in multi-LLM composition.

Multi-LLM proof recomposition can produce a larger candidate space than single-trace validation. Suppose that K models each generate at most L proof nodes. An unrestricted union graph contains at most K L candidate nodes. Comparing every node against every other node would require
O ( K L ) 2 = O ( K 2 L 2 )
compatibility tests. This is an absolute worst case and is not the strategy used by MetaPCR-LLM.
MetaPCR-LLM restricts candidate connections using the normalized conclusion, unresolved premise obligations, native-logic type, proof-graph position, provenance labels and contextual constraints. A connection is considered only when the conclusion of one certified fragment can satisfy a declared premise obligation of another fragment.
When candidate proofs are aligned into L reasoning positions, each position contains at most K candidate fragments. Comparing candidates only between consecutive positions gives the upper bound
B candidate ≤ ( L − 1 ) K 2 .
This bound is linear in L and quadratic in K. The quadratic term affects candidate search, not the number of bridges retained in the final proof.
The implementation further limits candidate growth through beam pruning. If at most w compatible certified fragments are retained at each position, where w ≤ K , then
B candidate , beam ≤ ( L − 1 ) w 2 .
When w is fixed, the number of evaluated candidate connections grows linearly with chain length and is independent of the total number of contributing models. After recomposition, an accepted linear composite proof still contains no more than
L − 1
bridge instances.

Controlled scalability evaluation.

We evaluate bridge scalability by independently varying the number of supported logics n, reasoning-chain length L, cross-logic transition rate q and number of contributing LLMs K. The evaluated settings are
n ∈ { 2 , 4 , 6 , 8 , 10 } , L ∈ { 5 , 10 , 20 , 40 , 80 } ,
q ∈ { 0 , 0.25 , 0.50 , 1.00 } , K ∈ { 1 , 3 , 5 , 8 , 10 } .
For every configuration, 100 proof graphs are generated while holding the distributions of native-step validity and bridge validity constant. The setting q = 0 represents a single-logic proof, whereas q = 1 represents the worst-case alternating-logic condition in which every adjacent transition requires a bridge.
The multi-LLM experiment compares the following three composition strategies:
1.
independent validation of all K complete reasoning chains;
2.
unpruned layer-aligned proof recomposition; and
3.
certificate-filtered and label-filtered recomposition with beam width w = 3 .
The bridge-layer microbenchmark begins with prevalidated native proof steps. This isolates the costs of registry lookup, compatibility filtering, bridge-certificate validation, caching and proof-graph composition from the runtime of the native logical solvers. A second end-to-end experiment includes the native validators to determine whether the same scaling trends remain observable when solver runtime is included.
For every run, we record:
  • the number of registered bridge schemas;
  • the number of candidate cross-logic connections;
  • the number of bridge instances actually validated;
  • the number of distinct bridge schemas used;
  • the structural-bridge cache hit rate;
  • bridge-validation wall-clock time;
  • total end-to-end validation time;
  • peak memory consumption; and
  • the number of invalid compositions rejected before final graph construction.
The K = 3 experiment reported in Table 12 and Table 15 provides one operating point for multi-LLM recomposition. The controlled scalability experiment extends this analysis across a broader range of n, L, q and K, allowing empirical growth rates to be compared with the analytical upper bounds.

Scalability interpretation.

The analysis distinguishes three sources of growth. First, the maximum number of bridge schemas is quadratic in the number of supported logics, but this cost applies only to a hypothetical complete registry. A sparse registry contains only the explicitly supported mappings. Second, the number of bridge instances in a linear proof is at most linear in reasoning-chain length and depends on the observed cross-logic transition rate. Third, multi-LLM candidate generation can be quadratic in the number of contributing models, but obligation indexing, certificate filtering and bounded-beam recomposition reduce the evaluated candidate set to
O ( L − 1 ) w 2 .
Consequently, increasing the number of supported logics does not force each reasoning trace to instantiate O ( n 2 ) bridges. Per-instance bridge cost is determined primarily by the number of cross-logic dependencies present in the proof graph. Quadratic behavior arises only when a complete pairwise registry is materialized or when multi-LLM fragments are compared without structural filtering. Reporting registry size, candidate connections and validated bridge instances separately makes this distinction explicit and permits the scalability of the bridge layer to be evaluated independently of the native logical solvers.

5.15. Multi-LLM Proof Recomposition

We obtain reasoning candidates from K = 3 LLMs and compare complete-answer selection, whole-answer voting, step recomposition, certificate-constrained recomposition and full proof-carrying recomposition.
Table 12. Multi-LLM proof recomposition.
Table 12. Multi-LLM proof recomposition.
Strategy Answer Acc. F1 CCA Invalid Compositions ↓
Best single model 0.75 0.69 0.65 0.29
Whole-answer voting 0.77 0.67 0.69 0.27
Step recomposition 0.80 0.71 0.70 0.23
+ native certificates 0.85 0.74 0.71 0.20
+ certificates + bridges + labels 0.86 0.80 0.76 0.18
Relative to the best single model, full proof recomposition improves both accuracy and F1 by 0.11. CCA likewise increases from 0.65 to 0.76.
The invalid-composition rate decreases from 0.29 to 0.18:
0.29 − 0.18 = 0.11 ,
corresponding to a relative reduction of
0.11 0.29 × 100 ≈ 37.9 % .
Whole-answer voting illustrates why answer aggregation alone is insufficient: accuracy increases from 0.75 to 0.77 while F1 decreases from 0.69 to 0.67. The strongest improvements occur once candidate reasoning fragments are individually certified and constrained by bridge and label compatibility.

5.16. Human Evaluation

The human study includes N = 12 evaluators and M = 30 examples per evaluator. Each evaluator first sees the LLM answer and its original reasoning trace and then repeats the assessment after receiving the corresponding verification report.
The agreement analysis is computed over 145 examples rated in common by the evaluators. Confidence is normalized to [ 0 , 1 ] by dividing the original 0–100 value by 100. A 1–5 Likert value x is normalized according to
x norm = x − 1 4 .
Table 13. Human evaluation.
Table 13. Human evaluation.
Condition Human Acc. Δ Conf. Fleiss’ κ Interpret. Failure Loc.
LLM only 0.68 – 0.66 0.64 0.67
+ Logic Program 0.75 0.12 0.71 0.68 0.67
+ Single-logic PCR 0.79 0.16 0.73 0.72 0.69
+ ValidLLP-style report 0.83 0.22 0.75 0.77 0.73
+ Full MetaPCR-LLM 0.88 0.25 0.78 0.81 0.76
Human judgment accuracy rises from 0.68 to 0.88:
0.88 − 0.68 = 0.20 ,
corresponding to approximately 29.4% relative improvement.
Normalized interpretability rises from 0.64 to 0.81. On the original five-point scale,
1 + 4 ( 0.64 ) = 3.56 , 1 + 4 ( 0.81 ) = 4.24 ,
which represents an increase of 0.68 Likert points.
The gains over ValidLLP are smaller: human accuracy increases from 0.83 to 0.88, interpretability from 0.77 to 0.81 and failure localization from 0.73 to 0.76. This is consistent with ValidLLP already providing a structured symbolic explanation. MetaPCR’s additional human-facing benefit is therefore concentrated in the explicit identification of native validators, certificates and cross-logic failures.
The Δ Conf. column measures confidence change rather than calibration; consequently, these values are not interpreted as evidence of improved human calibration.

5.17. Comparison with State-of-the-Art Neuro-Symbolic Frameworks

To address the architectural distinction between heterogeneous proof-carrying validation and single-target neuro-symbolic reasoning, we extend our evaluation to include prominent state-of-the-art frameworks: Logic-LM [6], LINC [7], SAFE [13] and Faithful CoT [4].
While these systems excel at delegating specific reasoning tasks to deterministic solvers (e.g., FOL provers, Lean 4, or PDDL planners), they inherently assume a single target logic or a uniform symbolic substrate for the entire reasoning chain. To evaluate their efficacy in hallucination detection—defined here as the ability to correctly identify and reject invalid, unsupported, or logically inconsistent intermediate reasoning steps—we adapt their architectures to our validation protocol. Specifically, a reasoning trace is flagged as hallucinated if the respective framework’s external solver rejects the formalized steps or yields a final derivation that contradicts the LLM’s proposed answer.

Datasets and Evaluation Protocol.

We evaluate the baselines across three distinct reasoning regimes:
1.
FOLIO [38]: A first-order logic benchmark to evaluate strict deductive hallucination detection, representing the native strength of Logic-LM and LINC.
2.
StrategyQA [39]: A multi-hop commonsense dataset to evaluate entailment and implicit-premise hallucinations, aligning with Faithful CoT and SAFE.
3.
Med-Halluc & Truthful-Halluc: Our domain-specific datasets to evaluate heterogeneous reasoning, where claims require mixed validation (e.g., probabilistic clinical evidence combined with hard numerical constraints).

Results

Table 14 presents the hallucination detection accuracy and step-validity performance. As expected, single-logic neuro-symbolic frameworks achieve strong performance in their native domains: Logic-LM and LINC excel on FOLIO, while SAFE and Faithful CoT show robust step-validity on StrategyQA.
However, when evaluated on heterogeneous datasets (Med-Halluc and Truthful-Halluc), single-target frameworks exhibit significant degradation. For instance, forcing probabilistic medical reasoning into a deterministic FOL substrate (Logic-LM) or a strict theorem prover (SAFE) results in high formalization failure rates, as these systems lack native mechanisms for uncertainty or defeasibility.
In contrast, MetaPCR-LLM achieves state-of-the-art hallucination detection across all regimes on mixed-logic Med-Halluc. By routing individual inference steps to their appropriate native validators (e.g., ProbLog for uncertainty, SMT for constraints) and validating cross-logic transitions via proof-carrying bridges, MetaPCR-LLM maintains a high performance floor. Notably, MetaPCR-LLM outperforms the strongest single-logic baseline (ValidLLP-style) by an absolute margin of 14%, demonstrating that heterogeneous proof composition is essential for comprehensive hallucination detection in complex, real-world reasoning.

Analysis of Cross-Logic Failures.

The performance gap on Med-Halluc highlights a fundamental limitation of single-logic verification. When Logic-LM attempts to validate a medical reasoning chain containing both probabilistic symptom likelihoods and hard dosage constraints, it must either ignore the probabilistic premises or fail to enforce the hard constraints. Consequently, it accepts reasoning traces that contain subtle cross-logic contradictions. MetaPCR-LLM explicitly detects these failures through its bridge validation layer, reducing the residual reasoning hallucination rate by 14% compared to the best single-target baseline.

5.18. Runtime and End-to-End Latency

Candidate answer generation is excluded from validation latency. We record reasoning decomposition, native-validator execution, bridge validation, counter-abduction, proof-graph aggregation and total wall-clock latency.
Table 15. Mean validation-stage latency in seconds and arithmetic sequential sum.
Table 15. Mean validation-stage latency in seconds and arithmetic sequential sum.
Method Decomp. Native Solvers Bridges Counter-Abd. Aggregation Sequential Sum
Single-logic PCR 0.45 0.12 – – 0.17 0.74
ValidLLP-style 0.25 0.15 – 0.26 0.14 0.80
MetaPCR w/o bridges 0.13 0.42 – 0.23 0.18 0.96
Full MetaPCR-LLM 0.23 0.46 0.34 0.25 0.23 1.51
The sums are
0.45 + 0.12 + 0.17 = 0.74 ,
0.25 + 0.15 + 0.26 + 0.14 = 0.80 ,
0.13 + 0.42 + 0.23 + 0.18 = 0.96 ,
and
0.23 + 0.46 + 0.34 + 0.25 + 0.23 = 1.51 s .
For Full MetaPCR, native solver execution accounts for
0.46 1.51 ≈ 30.5 %
of the sequential computational work, bridge validation for approximately 22.5%, counter-abduction for 16.6% and decomposition and aggregation for approximately 15.2% each.
Stage timers can overlap because independent validators and associated verification operations execute concurrently. Consequently, their arithmetic sum is not expected to equal measured wall-clock latency.
Table 16. Measured end-to-end wall-clock verifier latency.
Table 16. Measured end-to-end wall-clock verifier latency.
Method Mean (s) Median (s) P95 (s)
Single-logic PCR 0.75 0.79 0.82
ValidLLP-style 0.80 0.85 0.87
MetaPCR w/o bridges 0.82 0.79 0.81
Full MetaPCR-LLM 0.90 0.93 0.95
Measured wall-clock overhead is much smaller than suggested by the sequential stage sums. Relative to ValidLLP, Full MetaPCR increases mean latency by
0.90 − 0.80 = 0.10 s ,
or
0.10 0.80 × 100 = 12.5 % .
Relative to MetaPCR without bridges, the mean increase is
0.90 − 0.82 = 0.08 s ,
approximately 9.8%.
At the 95th percentile, MetaPCR exceeds ValidLLP by only
0.95 − 0.87 = 0.08 s ,
approximately 9.2%. Thus the proof-carrying bridge and meta-validation layer adds measurable but comparatively modest end-to-end latency under concurrent execution.

5.19. Statistical Analysis

All main comparisons are paired because competing validators operate on the same test instances. Differences in binary correctness are evaluated using McNemar’s test. F1 differences are evaluated using paired bootstrap resampling.
Because multiple hallucinated variants can originate from the same source instance, bootstrap resampling is performed at the source-instance level rather than independently over transformed examples. This preserves within-source dependence and avoids artificially narrow confidence intervals.
The updated point estimates for the primary comparison between Full MetaPCR and ValidLLP are
Δ Accuracy MetaPCR − ValidLLP = 0.86 − 0.83 = 0.03 , Δ F 1 MetaPCR − ValidLLP = 0.83 − 0.77 = 0.06 .
The paired-bootstrap confidence interval for this updated pooled comparison must be recomputed from the updated per-instance predictions; no significance claim is inferred from the rounded aggregate scores alone.
For the mixed-logic subset,
Δ F 1 MetaPCR − ValidLLP , mixed = 0.06 , 95 % C I = [ 0.048 , 0.072 ] .
The bridge ablation gives
Δ F 1 Full − NoBridge , mixed = 0.10 , 95 % C I = [ 0.082 , 0.123 ]
For the human study, binary paired judgments are evaluated with McNemar’s test, while confidence and normalized Likert responses are evaluated using the Wilcoxon signed-rank test unless assumptions supporting a paired parametric test are satisfied. The human-accuracy comparison between the LLM-only and Full MetaPCR conditions gives
Δ A c c u r a c y human = 0.20 , 95 % C I = [ 0.18 , 0.23 ]
Where multiple secondary comparisons are tested simultaneously, Holm correction is applied to control the family-wise error rate. Until the confidence intervals and corrected p-values are inserted, differences are described as observed or descriptive rather than statistically significant.

5.20. Error Analysis

We manually inspect a stratified sample of N = 124 MetaPCR failures. Errors are assigned to one of five categories:
1.
semantic decomposition error: the original reasoning step is incorrectly segmented or interpreted;
2.
logic-routing error: the step is assigned to an inappropriate validator;
3.
formalization error: the correct logic is selected but its premises or conclusion are translated incorrectly;
4.
bridge error: two locally valid proof objects are composed through an invalid or incomplete semantic translation;
5.
native-validator limitation: the relevant reasoning phenomenon cannot be represented by the selected formal engine.
Table 17. Distribution of manually analyzed MetaPCR failures.
Table 17. Distribution of manually analyzed MetaPCR failures.
Error category Count Percentage
Semantic decomposition 28 22.6%
Logic routing 27 21.8%
Formalization 21 16.9%
Cross-logic bridge 26 21.0%
Native-validator limitation 22 17.7%
Total 124 100%
No single category dominates the revised distribution. Semantic decomposition is the largest individual category at 28 of 124 failures (22.6%), followed closely by logic routing at 27 (21.8%) and cross-logic bridges at 26 (21.0%). Grouping related stages shows that routing and bridge composition account for 53 failures (42.7%), while semantic decomposition and formalization account for 49 (39.5%); native-validator limitations account for the remaining 22 (17.7%). The distribution therefore separates two comparably important sources of residual error. Native proof checking can detect an invalid formal inference, but it cannot guarantee that the original natural-language proposition was translated into the intended formal semantics; conversely, a faithful local formalization can still fail when it is routed to the wrong validator or transferred incorrectly across logics.

5.21. Qualitative Analysis of Error Cascades: Native Solvers versus Logic Bridges

To isolate the structural vulnerability of autoformalization within the framework, we contrast the error propagation dynamics of an internal object-level parsing failure with an inter-logic transition failure. We examine a clinical case from the Med-Halluc dataset involving an 84-year-old patient with severe renal impairment (estimated glomerular filtration rate, eGFR = 18 mL / min / 1.73 m 2 ). The safety guideline dictates that if eGFR < 30 , the maximum daily dose of Medication X must not exceed 25 mg . The LLM generates an invalid, hallucinated prescription of 50 mg daily.

5.21.1. Case 1: Direct Internal Autoformalization Failure within a Native Solver

When the validation task is restricted to a single monolithic target logic, the LLM translates the natural-language safety constraint directly into the input language of a single native solver (e.g., L SMT using Z3) [6]. During this parsing phase, the model commits a syntactic operator inversion error, erroneously translating the boundary check as:
(assert (ite (> egfr 30.0) (= max_dose 25.0) (= max_dose 100.0)))
(assert (<= prescribed_dose max_dose))
Because the input syntax is valid, the native solver executes the instructions without encountering an error. Given eGFR = 18.0 , the inverted conditional evaluates to false, causing Z3 to assign max _ dose = 100.0 . The downstream evaluation of the hallucinated prescription ( 50.0 ≤ 100.0 ) yields a sat outcome. The native engine generates a valid mathematical satisfiability model, issuing a falsified proof certificate χ SMT . The internal parsing error propagates silently, transforming the deterministic solver into an amplifier of the translation error and resulting in an undetected false-positive acceptance of the hallucinated reasoning trace.

5.21.2. Case 2: Inter-Logic Representation Failure Intercepted by a Logic Bridge

In the heterogeneous architecture of MetaPCR-LLM, the validation workload is distributed across specialized domains: L ALP certifies the qualitative abduction of patient status, while L SMT evaluates the quantitative dosage thresholds. The LLM accurately abduces the initial clinical state within the ALP engine, asserting the fact patient _ status ( renal _ insufficiency , severe ) backed by an authentic certificate χ ALP [28].
To evaluate the mathematical boundary condition, the system triggers the logic bridge b ALP → SMT to transport the validated predicate across the formal interface. During this mapping phase, the LLM commits an indexing error within the transformation functor α ALP → SMT , mapping the qualitative atom severe to a loose baseline value ( egfr > 75.0 ) instead of the appropriate bound ( egfr < 30.0 ) .
Unlike the single-logic configuration, this translation is not fed directly into the downstream solver kernel. Instead, the meta-validator passes the signature transformation through the bridge validation check ( B r i d g e V a l i d ( b i j ) ) against a set of globally immutable structural invariants stored in the system registry:
∀ x ∈ Sign SMT , severe _ insufficiency ( x ) ⇒ eGFR ( x ) < 30.0
The bridge validator tests the generated sentence mapping α ( ψ ) = ( egfr > 75.0 ) against the invariant constraint. Because the intersection of ( egfr > 75.0 ) and ( egfr < 30.0 ) produces an immediate structural contradiction, the transformation fails the preservation condition ρ ALP → SMT defined in Eq. (2). The meta-layer drops the bridge validity status to zero, truncating the typed proof graph topology at the junction node. The downstream solver is shielded from the corrupted signature inputs, preventing a silent failure cascade and forcing the verifier to emit an explicit, deterministic translation failure diagnostic code.

5.22. Estimation of Autoformalization Accuracy and Iterative Refinement

As indicated by the error analysis in Section , semantic decomposition and formalization errors constitute a significant portion of residual failures. Translating natural-language reasoning into strict native logics (e.g., Lean 4, Z3, ProbLog) is inherently challenging and initial autoformalization by the LLM is prone to syntactic mismatches, unbound variables, or incorrect predicate mappings. However, the MetaPCR-LLM architecture ensures that this vulnerability does not compromise the overall integrity of the validation pipeline through a strict, solver-mediated iterative refinement loop.
The fundamental safety of this procedure relies on the fact that native symbolic solvers act as uncompromising gatekeepers. Unlike neural judges that might assign a high confidence score to a subtly flawed formalization, deterministic solvers reject ill-formed inputs outright. A wrong autoformalization typically manifests as a syntax error, a type mismatch (e.g., in the Lean 4 kernel), an ill-sorted term (in SMT), or an immediate compilation failure (in Prolog/ALP). Crucially, these failures result in explicit solver rejection rather than silent semantic drift or a false-positive proof. The solver does not validate a flawed translation; it fails to execute it.
When a native validator rejects a proof obligation due to structural or trivial logical failure, the system captures the solver’s error trace or countermodel and feeds it back to the LLM. The LLM is then prompted to re-formalize the specific step, using the deterministic feedback to correct the mapping. This creates a bounded self-correcting loop. Because the formal system strictly separates “invalid formalization” (solver syntax/type errors) from “invalid reasoning” (semantic contradictions), the framework can safely route the former back to the LLM for correction without conflating it with actual logical hallucinations. Consequently, while autoformalization remains a bottleneck, the procedure is safe in general: the worst-case outcome of a bad translation is a solver error that triggers refinement, not the acceptance of an unsound proof.
To quantify the effectiveness of this iterative mitigation, we estimate the autoformalization accuracy across the test set. We define Initial Formalization Accuracy (IFA) as the percentage of reasoning steps successfully translated and accepted by the native solver on the first attempt and Final Formalization Accuracy (FFA) as the percentage successfully formalized after up to k = 3 refinement iterations.
Table 18 summarizes the estimated autoformalization performance across the heterogeneous native validators.
The results demonstrate that while the initial autoformalization success rate (IFA) is bounded by the LLM’s zero-shot translation capabilities, the iterative solver feedback mechanism substantially closes this gap, yielding a high Final Formalization Accuracy (FFA) of 87.0%. The mean number of refinement steps remains low (3.1 iterations on average), indicating that solver error traces provide highly actionable signals for the LLM. By relying on the deterministic rejection of malformed proofs rather than probabilistic confidence scores, MetaPCR-LLM ensures that the autoformalization process remains robust, auditable and safe against the introduction of silent logical errors.

5.23. Analysis of the Obtained Results

The updated results show that MetaPCR-LLM leads the principal aggregate comparison while retaining its clearest architectural advantage when reasoning involves heterogeneous formalisms, explicit proof composition, or diagnosis of reasoning failures. The improvement is not uniform across every dataset or specialized reasoning regime.

5.23.1. Overall Performance and Operating Trade-Offs

Full MetaPCR achieves overall accuracy of 0.86, compared with 0.83 for ValidLLP and 0.72 for single-PCR. Its relative accuracy gain over ValidLLP is
0.86 − 0.83 0.83 × 100 ≈ 3.6 % .
MetaPCR additionally improves SVA from 0.84 to 0.88 and CCA from 0.81 to 0.86, while reducing the Brier score from 0.067 to 0.060.
The updated pooled F1 comparison also favors MetaPCR: F1 rises from 0.77 for ValidLLP to 0.83 for Full MetaPCR, an absolute improvement of 0.06 and a relative improvement of approximately 7.8%. Taken together, the results show simultaneous gains in final-decision accuracy, hallucination F1, proof-step validity, chain-level correctness and calibration. The result should still be interpreted with the reported operating characteristics: MetaPCR has slightly lower coverage and escalation appropriateness and a marginally higher constraint-violation rate than ValidLLP.

5.23.2. Cross-Regime Robustness

MetaPCR obtains the highest five-regime macro-F1:
F 1 MetaPCR = 0.788 ,
versus 0.772 for ValidLLP and 0.722 for single-PCR.
Its worst-regime F1 is also higher:
min r F 1 MetaPCR , r = 0.75 ,
compared with 0.72 and 0.69, respectively.
This is more informative than requiring MetaPCR to outperform every native solver in every regime. PLP, ALP, AF/DeLP and SMT/CLP retain their expected specialization advantages. MetaPCR’s role is instead to coordinate these reasoning capabilities while maintaining a high performance floor as the required formalism changes.

5.23.3. Cross-Dataset Generalization

MetaPCR obtains dataset macro-F1
F 1 macro = 0.8375 ,
reported as 0.838 in Table 8, compared with 0.785 for ValidLLP and 0.745 for single-PCR.
The gain over ValidLLP is 0.0525, approximately 6.7%, while the gain over single-PCR is 0.0925, approximately 12.4%. Full MetaPCR now leads on three of the four datasets, including Autoimmune-narrate at 0.81. ValidLLP retains a 0.02 advantage only on Med-Halluc. The resulting 0.05 spread between MetaPCR’s highest and lowest dataset scores provides stronger evidence of cross-dataset robustness, while the remaining Autoimmune gap indicates that semantically diffuse narrative evidence is still the most difficult setting.

5.23.4. Role of Proof-Carrying Bridges

The no-bridge comparison provides a direct test of the principal architectural extension.
On the complete benchmark,
F 1 Full − F 1 NoBridge = 0.83 − 0.76 = 0.07 .
On mixed-logic cases,
F 1 Full , mixed − F 1 NoBridge , mixed = 0.80 − 0.70 = 0.10 .
The larger mixed-logic effect is expected because bridge certification can only contribute when reasoning actually crosses a formal boundary. This concentration of the effect provides stronger evidence for bridge validation than an aggregate improvement distributed across examples that do not require bridges.

5.23.5. Contribution of the Remaining Components

The largest F1 ablation loss arises from removing native proof certificates:
Δ F 1 = 0.23 .
This is followed by counter-abduction (0.14), forcing a single target logic (0.13), defeasibility labels (0.11), provenance (0.10), temporal labels (0.09), uncertainty labels (0.08) and bridge checking on the full benchmark (0.07).
The ordering indicates that the largest benefit comes from actual native proof verification rather than metadata alone. At the same time, metadata, counter-abduction and cross-logic certification provide complementary improvements that cannot be reproduced by a single native solver.

5.23.6. Failure Localization and Auditability

MetaPCR’s updated aggregate answer-level gains over ValidLLP are accompanied by clear diagnostic improvements. Failure-type localization improves from 0.67 to 0.76 and correct-validator identification from 0.62 to 0.71, both absolute gains of 0.09.
These measures directly address auditability: a verifier is more useful when it can state not merely that a conclusion is unreliable but whether the failure originated in an unsupported premise, invalid inference, violated constraint, defeasible conflict, or cross-logic translation.

5.23.7. Multi-LLM Proof Recomposition

Whole-answer voting increases accuracy from 0.75 to 0.77 but decreases F1 from 0.69 to 0.67. In contrast, proof-aware recomposition gradually improves both measures, reaching 0.86 accuracy and 0.80 F1 under the complete certificate+bridge+label configuration.
Invalid composition falls from 0.29 to 0.18, a relative reduction of 37.9%. This suggests that the benefit of multi-model reasoning arises not merely from producing more candidate answers, but from composing individually verified proof fragments under explicit compatibility constraints.

5.23.8. Human-Facing Effects

Human judgment accuracy increases from 0.68 to 0.88, an absolute gain of 0.20 and relative gain of approximately 29.4%.
Normalized interpretability rises from 0.64 to 0.81, corresponding to an increase from 3.56 to 4.24 on the original five-point scale. Failure localization also rises from 0.67 to 0.76.
The additional gain over ValidLLP is smaller, as expected for two systems that both expose structured symbolic information. MetaPCR’s distinctive human-facing contribution lies primarily in making native proof identity and cross-logic failure explicit.

5.23.9. Accuracy–Cost Trade-Off

The sequential stage sum increases from 0.80 s for ValidLLP to 1.51 s for Full MetaPCR, but concurrent execution substantially reduces the actual difference.
Measured mean wall-clock latency is
0.90 s
for MetaPCR and
0.80 s
for ValidLLP, corresponding to only 12.5% mean overhead.
The P95 difference is similarly modest:
0.95 − 0.87 = 0.08 s ,
or approximately 9.2%. The richer proof structure therefore introduces additional computation, but the measured deployment-level latency increase is considerably smaller than the raw sum of individual stage timings.

5.23.10. Summary with Respect to the Research Questions

RQ1.
MetaPCR outperforms ValidLLP on both pooled hallucination F1 (0.83 versus 0.77) and overall accuracy (0.86 versus 0.83), while also improving SVA, CCA, calibration, dataset macro-F1 and mixed-logic F1. The strongest architecture-specific advantage occurs on reasoning that genuinely requires heterogeneous composition.
RQ2.
MetaPCR achieves the highest five-regime macro-F1 (0.788) and the highest worst-regime score (0.75), while specialized native solvers remain strongest in several individual regimes.
RQ3.
Native certificates have the largest observed ablation effect. The remaining results show complementary contributions from counter-abduction, heterogeneous native-logical representation, metadata and bridge certification. Bridge validation has a particularly strong effect on the mixed-logic subset.
RQ4.
MetaPCR improves failure-type and validator localization by 0.09 relative to ValidLLP and additionally represents cross-logic bridge failures as independently inspectable proof objects. Among the 124 manually analyzed MetaPCR failures, routing and bridge errors account for 42.7%, semantic decomposition and formalization for 39.5% and native-validator limitations for 17.7%. This near balance indicates that localization must cover both representation fidelity and proof-system orchestration.
RQ5.
The complete verification report increases descriptive human judgment accuracy from 0.68 to 0.88, interpretability from 0.64 to 0.81 and failure localization from 0.67 to 0.76. Statistical significance is evaluated through the paired tests defined above.
Overall, the updated aggregate comparison supports MetaPCR-LLM as a higher-performing binary hallucination detector, but its contribution is broader than that point estimate. MetaPCR-LLM is most useful when reasoning spans formal systems, when independently checked proof fragments must be composed safely, when alternative explanations must be challenged and when the location and formal nature of a reasoning failure must remain inspectable. This behavior is consistent with the intended role of MetaPCR-LLM as a proof-carrying meta-framework for heterogeneous logical reasoning.

6. Conclusions

This paper introduced MetaPCR-LLM, a proof-carrying framework for validating LLM reasoning across heterogeneous native logics. The central idea is to move beyond validation relative to a single target formalism and instead represent an LLM reasoning trace as a typed proof graph whose individual steps are checked by the logical systems most appropriate to them. Native proof certificates or machine-checkable witnesses establish local validity, while labelled metadata records provenance, uncertainty, temporal scope, modality, priority and defeasibility. Most importantly, transitions between different logical systems are treated as explicit proof obligations through proof-carrying bridges. This makes it possible to distinguish a chain in which every local inference is individually valid from one in which the overall reasoning nevertheless fails because an intermediate conclusion has been translated incorrectly across formal systems.
The evaluation indicates that this distinction is practically important. Full MetaPCR-LLM achieves an overall accuracy of 0.86, step-validity accuracy of 0.88, chain-certification accuracy of 0.86 and a cross-dataset macro-F1 of 0.838. Its pooled hallucination F1 of 0.83 exceeds the 0.77 of the ValidLLP-style baseline by 0.06 at the reported precision, while overall accuracy improves by 0.03. This updated result establishes an aggregate answer-level advantage in addition to the framework’s proof-level benefits; the gain is nevertheless not uniform across every dataset or specialized reasoning regime. The strongest architecture-specific gains occur on tasks that directly exercise heterogeneous reasoning. On the mixed-logic subset, MetaPCR reaches F1 = 0.80 , compared with 0.74 for ValidLLP-style validation and 0.70 for the same MetaPCR architecture without explicit bridge checking. The latter difference is particularly informative because it isolates the role of cross-logic certification: bridge validation produces a 0.10 absolute F1 gain on examples in which such transitions are actually required, compared with a 0.07 gain over the complete benchmark.
The reasoning-regime results provide complementary evidence for the proposed design. Specialized native formalisms remain strongest in several of their natural domains: probabilistic logic performs best on uncertainty, argumentation/defeasible reasoning performs best on conflict and constraint-oriented systems perform best on constraint reasoning. MetaPCR is therefore not intended to replace these systems with a new monolithic logic. Its advantage lies in coordinating them. It obtains the highest five-regime macro-F1 of 0.788 and the highest worst-regime F1 of 0.75, indicating that heterogeneous composition can maintain a comparatively high performance floor as the required reasoning formalism changes.
The ablation experiments further clarify where this performance originates. Removing native proof certificates produces the largest F1 degradation, followed by counter-abduction, forcing the complete reasoning chain into a single target logic and removing defeasibility and provenance information. Bridge validation has a smaller effect when averaged over the complete benchmark but a substantially larger effect on the mixed-logic subset. This supports the architectural separation between three forms of validation: checking a native inference, determining whether that inference remains acceptable under competing explanations and metadata and verifying whether the result can be transferred safely into another formal system.
MetaPCR also provides benefits that are not captured by final-answer classification alone. It improves failure-type and validator localization, represents bridge failures as first-class diagnostic objects and supports proof-aware recomposition of reasoning fragments generated by multiple LLMs. The revised manual analysis of 124 failures shows a distributed error profile: semantic decomposition is the largest individual category at 22.6%, while routing and bridge errors jointly account for 42.7% and decomposition plus formalization jointly account for 39.5%. This result reinforces the need to audit both the construction of local proof obligations and their coordination across native validators. In the multi-LLM experiment, the complete certificate–bridge–label configuration reaches 0.86 answer accuracy and 0.80 F1 while reducing invalid compositions from 0.29 to 0.18. Human evaluation likewise suggests that exposing explicit validation structure can improve reasoning assessment: human judgment accuracy rises from 0.68 in the LLM-only condition to 0.88 with the full MetaPCR report, while normalized interpretability rises from 0.64 to 0.81. These results suggest that proof-carrying validation can serve not only as an automated correctness mechanism but also as an interface for auditing why a reasoning chain was accepted or rejected. The implementation, evaluation scripts and supporting MetaPCR resources are available in the project repository: https://github.com/bgalitsky/halluc_in_health/tree/master/MetaPCR.
The framework nevertheless has important limitations. First, formal verification is only as reliable as the mapping from natural language to the formal proof obligation. A native checker may correctly certify a statement that does not faithfully represent the intended natural-language claim. Consistent with this limitation, semantic decomposition and formalization constitute 39.5% of the manually analyzed failures, although routing and cross-logic bridge errors are collectively slightly more frequent at 42.7%. Although Full MetaPCR now reaches 0.81 on Autoimmune-narrate-halluc and outperforms all baselines on that dataset, it remains the lowest of MetaPCR’s four dataset-specific scores. This residual gap indicates that semantically diffuse and implicit evidence remains more difficult to extract and formalize. Second, not every native validator provides the same type of proof artifact. Some systems produce kernel-checkable proofs, whereas others provide satisfiability models, abductive explanations, argumentation witnesses, probability estimates, or other machine-checkable evidence. MetaPCR therefore requires a typed notion of certificate rather than assuming a single uniform proof representation. Third, bridge correctness remains a fundamental challenge. A correct proof in one logic does not by itself guarantee that a translation into another logic preserves entailment, satisfaction, provenance, or contradiction. The bridge layer must therefore be constrained by explicit semantic-preservation conditions rather than by surface-level equivalence alone.
A further limitation is computational cost. Full MetaPCR introduces additional native-solver, bridge-validation and counter-abduction work. In the current implementation, however, concurrent execution limits the measured mean wall-clock overhead relative to ValidLLP-style validation to approximately 12.5%. This suggests that heterogeneous proof checking is computationally feasible, although larger theorem-proving tasks, more complex bridge networks and broader multi-agent deployments may require more aggressive routing, caching and parallelization.
Several directions follow naturally from this work. One is stronger representation certification: instead of validating only the formal proof, future systems should attach evidence that the formalized statement preserves the intended meaning of the original natural-language claim. A second direction is automated synthesis and verification of cross-logic bridges, including conservative translations and mappings with explicitly stated satisfaction-preservation guarantees. Third, the current counter-abductive mechanism can be extended to richer competing proof graphs, allowing independently certified alternative explanations to challenge not only final conclusions but intermediate assumptions and translations. Finally, the typed proof-graph representation offers a natural foundation for multi-agent reasoning in which different models or tools specialize in different logics but can contribute only those fragments that survive native verification and certified composition.
More broadly, the results suggest that validating LLM reasoning should not be reduced to deciding whether a final answer is true or false. Complex model reasoning may be locally correct yet globally invalid because a premise is unsupported, a constraint is violated, a defeater has been ignored, or a conclusion has been transferred incorrectly between reasoning regimes. MetaPCR-LLM addresses this problem by making both inference steps and cross-formalism transitions explicit, inspectable and independently checkable. The resulting perspective shifts proof-carrying LLM validation from verification within one target logic toward the coordinated validation of heterogeneous reasoning itself.

Funding

The work was supported by the Ministry of Economic Development of the Russian Federation (agreement No. 139-15-2025-013, dated June 20, 2025, IGK 000000C313925P4B0002).

References

  1. Wei, J.; Wang, X.; Schuurmans, D.; Bosma, M.; Ichter, B.; Xia, F.; Chi, E.H.; Le, Q.V.; Zhou, D. Chain-of-Thought Prompting Elicits Reasoning in Large Language Models. In Advances in Neural Information Processing Systems; 2022; Vol. 35. [Google Scholar]
  2. Wang, X.; Wei, J.; Schuurmans, D.; Le, Q.V.; Chi, E.H.; Narang, S.; Chowdhery, A.; Zhou, D. Self-Consistency Improves Chain of Thought Reasoning in Language Models. In Proceedings of the International Conference on Learning Representations (ICLR), 2023. [Google Scholar]
  3. Lightman, H.; Kosaraju, V.; Burda, Y.; Edwards, H.; Baker, B.; Lee, T.; Leike, J.; Schulman, J.; Sutskever, I.; Cobbe, K. Let’s Verify Step by Step. In Proceedings of the International Conference on Learning Representations (ICLR), 2024. [Google Scholar]
  4. Lyu, Q.; Havaldar, S.; Stein, A.; Zhang, L.; Rao, D.; Wong, E.; Apidianaki, M.; Callison-Burch, C. Faithful Chain-of-Thought Reasoning. In Proceedings of the 13th International Joint Conference on Natural Language Processing and the 3rd Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics; Association for Computational Linguistics, 2023; pp. 305–329. [Google Scholar] [CrossRef]
  5. Paul, D.; West, R.; Bosselut, A.; Faltings, B. Making Reasoning Matter: Measuring and Improving Faithfulness of Chain-of-Thought Reasoning. In Findings of the Association for Computational Linguistics: EMNLP; Association for Computational Linguistics, 2024; pp. 15012–15032. [Google Scholar]
  6. Pan, L.; Albalak, A.; Wang, X.; Wang, W. Logic-LM: Empowering Large Language Models with Symbolic Solvers for Faithful Logical Reasoning. In Findings of the Association for Computational Linguistics: EMNLP 2023; Association for Computational Linguistics, 2023; pp. 3806–3824. [Google Scholar] [CrossRef]
  7. Olausson, T.; Gu, A.; Lipkin, B.; Zhang, C.; Solar-Lezama, A.; Tenenbaum, J.; Levy, R. LINC: A Neurosymbolic Approach for Logical Reasoning by Combining Language Models with First-Order Logic Provers. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 2023; Association for Computational Linguistics; pp. 5153–5176. [Google Scholar] [CrossRef]
  8. Xu, J.; Fei, H.; Pan, L.; Liu, Q.; Lee, M.-L.; Hsu, W. Faithful Logical Reasoning via Symbolic Chain-of-Thought. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics; Association for Computational Linguistics, 2024; pp. 13326–13365. [Google Scholar] [CrossRef]
  9. Calanzone, D.; Teso, S.; Vergari, A. Logically Consistent Language Models via Neuro-Symbolic Integration. In Proceedings of the International Conference on Learning Representations (ICLR), 2025. [Google Scholar]
  10. Arakelyan, E.; Minervini, P.; Lewis, P.; Verga, P.; Augenstein, I. FLARE: Faithful Logic-Aided Reasoning and Exploration. In In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, 2025; Association for Computational Linguistics. [Google Scholar] [CrossRef]
  11. de Moura, L.; Ullrich, S. The Lean 4 Theorem Prover and Programming Language. In Automated Deduction – CADE 28;Lecture Notes in Computer Science; Springer: Cham, Switzerland, 2021; Vol. 12699. [Google Scholar] [CrossRef]
  12. Yang, K.; Swope, A.; Gu, A.; Chalamala, R.; Song, P.; Yu, S.; Godil, S.; Prenger, R.; Anandkumar, A. LeanDojo: Theorem Proving with Retrieval-Augmented Language Models. In Advances in Neural Information Processing Systems; 2023; Vol. 36. [Google Scholar]
  13. Liu, C.; Yuan, Y.; Yin, Y.; Xu, Y.; Xu, X.; Chen, Z.; Wang, Y.; Shang, L.; Liu, Q.; Zhang, M. SAFE: Enhancing Mathematical Reasoning in Large Language Models via Retrospective Step-Aware Formal Verification. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics, 2025; Association for Computational Linguistics; pp. 12171–12186. [Google Scholar] [CrossRef]
  14. Leang, J.O.J.; Hong, G.; Li, W.; Cohen, S.B. Theorem Prover as a Judge for Synthetic Data Generation. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics, 2025; Association for Computational Linguistics; pp. 29941–29977. [Google Scholar] [CrossRef]
  15. Quan, X.; Valentino, M.; Dennis, L.A.; Freitas, A. Faithful and Robust LLM-Driven Theorem Proving for NLI Explanations. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics, 2025; Association for Computational Linguistics; pp. 17734–17755. [Google Scholar] [CrossRef]
  16. Cao, J.; Lu, Y.; Li, M.; Ma, H.; Li, H.; He, M.; Wen, C.; Sun, L.; Zhang, H.; Qin, S.; Cheung, S.-C.; Tian, C. From Informal to Formal: Incorporating and Evaluating LLMs on Natural Language Requirements to Verifiable Formal Proofs. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics, 2025; Association for Computational Linguistics; pp. 26984–27003. [Google Scholar] [CrossRef]
  17. Li, T.; Wang, P.; Wang, H.; Hahm, C.; Spatola, M.; Shi, J. PCRLLM: Proof-Carrying Reasoning with Large Language Models under Stepwise Logical Constraints. arXiv 2025, arXiv:2511.08392. [Google Scholar]
  18. Gabbay, D.M. Labelled Deductive Systems; Oxford University Press: Oxford, UK, 1996. [Google Scholar] [CrossRef]
  19. Goguen, J.A.; Burstall, R.M. Institutions: Abstract Model Theory for Specification and Programming. J. ACM 1992, 39, 95–146. [Google Scholar] [CrossRef]
  20. Meseguer, J. General Logics. In Logic Colloquium ’87;Studies in Logic and the Foundations of Mathematics; North-Holland: Amsterdam, The Netherlands, 1989; Vol. 129, pp. 275–329. [Google Scholar] [CrossRef]
  21. Carnielli, W.; Coniglio, M.E.; Gabbay, D.M.; Gouveia, P.; Sernadas, C. Analysis and Synthesis of Logics: How to Cut and Paste Reasoning Systems; Springer: Dordrecht, The Netherlands; Applied Logic Series, 2008; Vol. 35. [Google Scholar] [CrossRef]
  22. Kakas, A.C.; Kowalski, R.A.; Toni, F. Abductive Logic Programming. J. Log. Comput. 1992, 2, 719–770. [Google Scholar] [CrossRef]
  23. Dung, P.M. On the Acceptability of Arguments and Its Fundamental Role in Nonmonotonic Reasoning, Logic Programming and n-Person Games. Artif. Intell. 1995, 77, 321–357. [Google Scholar] [CrossRef]
  24. De Raedt, L.; Kimmig, A.; Toivonen, H. ProbLog: A Probabilistic Prolog and Its Application in Link Discovery. In Proceedings of the 20th International Joint Conference on Artificial Intelligence, 2007; pp. 2462–2467. [Google Scholar]
  25. Galitsky, B.; Solodkin, V.; Beznosikov, A. Human–Agent Joint Design with ValidLLP4LLM: A Labeled Logic Framework for Validating LLM Reasoning; 2026. [Google Scholar]
  26. Galitsky, B. An Information–Theoretic Model of Abduction for Detecting Hallucinations in Explanations. Entropy 2026, 28, 173. [Google Scholar] [CrossRef] [PubMed]
  27. Galitsky, B.; Rybalov, A. Neuro-Symbolic Verification for Preventing LLM Hallucinations in Process Control. Processes 2026, 14, 322. [Google Scholar] [CrossRef]
  28. Galitsky, B.A. Tackling LLM Hallucination with Abductive Reasoning. Preprints 2025. [Google Scholar]
  29. Galitsky, B. Applications of Neuro-Symbolic Artificial Intelligence; Springer Nature: Cham, Switzerland, 2026. [Google Scholar] [CrossRef]
  30. Galitsky, B. Adversarial Integration of LLM and Logic Program. In Applications of Neuro-Symbolic Artificial Intelligence; Springer Nature: Cham, Switzerland, 2026; pp. 13–43. [Google Scholar] [CrossRef]
  31. Galitsky, B. Adversarial Abductive Dialogue Framework with Reinforcement for Tackling LLM Hallucination. In Applications of Neuro-Symbolic Artificial Intelligence; Springer Nature: Cham, Switzerland, 2026; pp. 45–93. [Google Scholar] [CrossRef]
  32. Pesjak, D.; Žabkar, J. Robot Planning via LLM Proposals and Symbolic Verification. Mach. Learn. Knowl. Extr. 2026, 8, 22. [Google Scholar] [CrossRef]
  33. Maltoni, D.; Ferrara, M. Deductive Logic in Language Models: Horizontal vs. Vertical Reasoning. Mach. Learn. Knowl. Extr. 2026, 8, 214. [Google Scholar] [CrossRef]
  34. Kadyrbek, N.; Mansurova, M. Primitive-Augmented Transformers with Event-Role Side State: Architecture Evidence, Warm-Started Modulation and Decoupled Tool Interfaces. Mach. Learn. Knowl. Extr. 2026, 8, 201. [Google Scholar] [CrossRef]
  35. Ngartera, L.; Nadarajah, S.; Koina, R.; Gningue, Y. BRAG: Bayesian Retrieval-Augmented Generation; A Methodological Framework for Evidence-Governed Decision Support. Mach. Learn. Knowl. Extr. 2026, 8, 151. [Google Scholar] [CrossRef]
  36. Lipianina-Honcharenko, K.; Bykovyy, P.; Krysovatyy, A.; Komar, M.; Yazlyuk, B. Scenario-Adaptive Evaluation of Trustworthy Fine-Tuned Text Models Across Knowledge-Grounded Generation and Misinformation Detection. Mach. Learn. Knowl. Extr. 2026, 8, 161. [Google Scholar] [CrossRef]
  37. Saleh, A.O.M.; Tur, G.; Saygin, Y. SG-RAG MOT: SubGraph Retrieval Augmented Generation with Merging and Ordering Triplets for Knowledge Graph Multi-Hop Question Answering. Mach. Learn. Knowl. Extr. 2025, 7, 74. [Google Scholar] [CrossRef]
  38. Han, S.; Shoaran, H.; Tan, L.; Zhang, Y.; Bosselut, A.; West, R. FOLIO: An NLI Dataset for Gauging Natural Language Reasoning in First-Order Logic. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics; Association for Computational Linguistics: Bangkok, Thailand, 2024; Volume 1, pp. 14201–14223. [Google Scholar]
  39. Geva, M.; Khashabi, D.; Segal, E.; Khot, T.; Roth, D.; Berant, J. Did Aristotle Use a Laser Pointer? Q&A over Commonsense Reasoning Chains. Trans. Assoc. Comput. Linguist. 2021, 9, 346–360. [Google Scholar]
Figure 1. MetaPCR-LLM architecture. LLM-generated reasoning is decomposed into proof obligations, routed to native validators, enriched with structured labels and composed through certified cross-logic bridges into a typed proof graph.
Figure 1. MetaPCR-LLM architecture. LLM-generated reasoning is decomposed into proof obligations, routed to native validators, enriched with structured labels and composed through certified cross-logic bridges into a typed proof graph.
Preprints 232580 g001
Figure 2. Commutative structure of the Proof-Carrying Logic Bridge morphism showing the satisfaction invariant mapping between source logic I i and target logic I j .
Figure 2. Commutative structure of the Proof-Carrying Logic Bridge morphism showing the satisfaction invariant mapping between source logic I i and target logic I j .
Preprints 232580 g002
Figure 3. MetaPCR-LLM evaluation pipeline. Four annotated hallucination datasets are used for a controlled comparison of baseline methods and the full MetaPCR-LLM system. Evaluation includes predictive performance, mixed-logic reasoning, ablations, latency, statistical testing, human assessment and manual analysis of 124 failures. Full MetaPCR-LLM obtains a cross-dataset macro-F1 of 0.838 and achieves the best result on three of the four datasets.
Figure 3. MetaPCR-LLM evaluation pipeline. Four annotated hallucination datasets are used for a controlled comparison of baseline methods and the full MetaPCR-LLM system. Evaluation includes predictive performance, mixed-logic reasoning, ablations, latency, statistical testing, human assessment and manual analysis of 124 failures. Full MetaPCR-LLM obtains a cross-dataset macro-F1 of 0.838 and achieves the best result on three of the four datasets.
Preprints 232580 g003
Table 1. Native reasoning engines used by MetaPCR-LLM.
Table 1. Native reasoning engines used by MetaPCR-LLM.
Reasoning regime Native engine Version/configuration
Logic programming SWI Prolog 10.0.0
Constraint / SMT Z3 5.1.0 (released August 16, 2026)
Abduction ALP implementation CIFF/ unversioned
Probabilistic reasoning ProbLog 2.2
Argumentation / defeasibility N/A
Temporal constraints N/A
Formal theorem proving Lean 4 4.33.1 / from github
Table 2. Dataset statistics and train, validation and test partitions.
Table 2. Dataset statistics and train, validation and test partitions.
Dataset Train Validation Test Total Hallucination rate
Truthful-Halluc 200 300 500 1000 46%
Med-Halluc 400 600 1000 2000 53%
eSNLI-Halluc 200 300 500 1000 40%
Autoimmune-narrate-halluc 240 360 600 1200 38%
Total 1040 1560 2600 5200 ≈ 45.7 %
Table 14. Hallucination Detection Accuracy and Step-Validity across Neuro-Symbolic Baselines. Acc. denotes Answer/Hallucination Detection Accuracy; SVA denotes Step-Validity Accuracy.
Table 14. Hallucination Detection Accuracy and Step-Validity across Neuro-Symbolic Baselines. Acc. denotes Answer/Hallucination Detection Accuracy; SVA denotes Step-Validity Accuracy.
Method Target Logic FOLIO (Deductive) StrategyQA (Multi-hop) Med-Halluc (Heterogeneous)
Acc.↑ SVA↑ Acc.↑ SVA↑ F1↑ SVA↑
LLM-only (CoT) None 0.41 0.45 0.65 0.58 0.52 0.65
Faithful CoT PDDL / Python 0.58 0.62 0.78 0.74 0.61 0.74
Logic-LM Datalog / FOL 0.76 0.71 0.68 0.65 0.58 0.60
LINC FOL (Prover9) 0.74 0.69 0.66 0.63 0.55 0.58
SAFE Lean 4 (Math) 0.69 0.75 0.72 0.79 0.64 0.68
ValidLLP-style Labelled LLP 0.65 0.78 0.74 0.81 0.86 0.84
MetaPCR-LLM (Ours) Heterogeneous 0.73 0.82 0.76 0.85 0.84 0.88
Table 18. Estimation of Autoformalization Accuracy and Iterative Refinement.
Table 18. Estimation of Autoformalization Accuracy and Iterative Refinement.
Native Validator Initial Acc. (IFA) Final Acc. (FFA) Mean Refinement Steps
Lean 4 (Theorem Prover) 88.3 93.2 3.2
Z3 (SMT / Constraints) 79.1 96.3 2.8
SWI Prolog / ALP 84.7 90.1 3.0
ProbLog (Probabilistic) 80.3 94.5 4.3
Aggregate (All Logics) 76.2 87.0 6.2
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.