Preprint
Article

This version is not peer-reviewed.

MAVERICK: Breaking the FPR–FNR Seesaw in LLM Output Verification Through Strategy-Differentiated Multi-Agent Consensus

  † These authors contributed equally: Siping Dong, Qiang Yao and Zixuan Dong.

Submitted:

26 August 2026

Posted:

26 August 2026

You are already at the latest version

Abstract
The FPR–FNR seesaw—a mathematical constraint on any single-pipeline verifier, where lowering the false-negative rate necessarily raises false positives—is inherent to single-agent judgment, not merely an engineering limitation. We prove that, under conditional error independence, strategy-differentiated multi-agent consensus transcends this trade-off: when agents with deliberately different evidence sources, thresholds and judgment criteria must all agree before an output is automatically released, the system-level FPR falls as p^k—ambiguous cases are escalated to human review rather than silently misclassified. We instantiate this in(MAVERICK), a three-agent architecture, and validate it on LLM reference hallucination (N=5,094): 93.4% unambiguous verdicts,0% FPR on 1,280 fabricated references, and zero misses, consistent with p³ ≤ 2.4×10⁻⁵. Cross-domain pilots on legal citation and clinical trial registry verification (N=400) transfer with zero silent errors among automated verdicts. The architecture is domain-agnostic: wherever LLM outputs can be independently verified, consensus among differentiated verifiers offers a mathematically grounded alternative to model scaling.
Keywords: 
;  ;  ;  ;  ;  ;  ;  ;  ;  
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.