Submitted:
22 February 2026
Posted:
27 February 2026
You are already at the latest version
Abstract
We develop \emph{Semantic Thermodynamics}, an information-theoretic framework for analyzing hallucinations in transformer systems under finite resources. The central object is mutual information between latent facts and model outputs, together with Fano-style lower bounds on semantic error. We clarify the stochastic assumptions required for non-degenerate information measures, distinguish true data-generating uncertainty from model-implied uncertainty, and replace unsupported hard capacity formulas with explicit capacity surrogates tied to precision, context budget, and effective representational rank. Under standard identification assumptions, we derive a baseline bound \begin{equation*} H_R \geq \max\left\{0,\,1-\frac{I(F;Y)+1}{\log M}\right\}, \end{equation*} where \(H_R\) is hallucination rate, \(F\) is the latent semantic fact, \(Y\) is model output, and \(M\) is semantic cardinality. We also provide a distribution-dependent variant and a bottleneck-aware extension for retrieval-augmented generation. This paper contributes a mathematically consistent formulation, a tighter assumptions section, and concrete empirical protocols for estimation and falsification.
Keywords:
transformers
; hallucinations
; information theory
; semantic uncertainty
; Fano’s inequality
; retrieval-augmented generation
1. Introduction
Hallucination remains a central failure mode of modern language models: outputs can be fluent and high-confidence while semantically incorrect. Mitigation methods including alignment, finetuning, and retrieval improve performance but do not eliminate failures. This pattern motivates a complementary question: beyond optimization quality, what lower bounds arise from finite information-processing resources?
This paper proposes an information-theoretic treatment of that question. We model semantic prediction as transmission of latent facts through a bounded inference pipeline. We then use mutual information and Fano-type converse bounds to characterize unavoidable error regimes.
Our goal is not to claim that architecture alone determines hallucinations. Rather, we isolate one layer of explanation: when effective semantic information throughput is below task complexity, nonzero semantic error is unavoidable even with idealized training. Throughout, “thermodynamics” is used as a metaphorical organizing label for resource-constrained information processing; the formal results in this paper are information-theoretic converses rather than a physical thermodynamic formalism.
1.1. Contributions
This paper makes the following contributions:
- 1.
- It standardizes mutual information notation and removes degenerate uses of under deterministic decoding.
- 2.
- It separates true semantic distributions from model-induced distributions to avoid circular definitions.
- 3.
- It replaces inconsistent capacity formulas with explicit, assumption-dependent surrogates.
- 4.
- It weakens strict layerwise-loss claims to DPI-based monotonicity, with strict decrease only under explicit stochastic perturbations.
- 5.
- It adds distribution-dependent and bottleneck-aware bound variants, including a non-additive RAG analysis.
- 6.
- It expands related work to include semantic rate-distortion, memory-limited hallucination frontiers, and f-divergence minimax lower bounds.
2. Related Work
2.1. Information-Theoretic and Semantic Foundations
Our framework uses modern information-theoretic analyses of language models and transformers, with emphasis on semantic capacity, compression, and distortion [8,9,14]. We focus on contemporary (2020+) references and empirical-theoretical bridges that are directly actionable for current LLM systems. In autoregressive settings, directed-information viewpoints can complement when feedback and tokenwise causality are central [14].
2.2. Theory of Hallucinations and Transformer Limits
Recent analyses formalize hallucination incentives and uncertainty calibration in modern training/evaluation pipelines [4]. Complementary memory-theoretic results show high-confidence hallucination can be an unavoidable consequence of finite storage in sparse-fact regimes [10]. New 2026 studies also characterize multi-turn hallucination detection and relation-structure predictors of hallucination [11,12]. Transformer-specific analyses of attention pathologies and rank effects provide architectural mechanisms compatible with information bottlenecks [5,6,7]. Layerwise interpretability/probing work (e.g., tuned-lens analyses) complements our use of task-level semantic MI by exposing non-monotonic internal dynamics even when end-to-end converse bounds remain valid [16]. Recent layerwise hallucination studies and latent-risk detectors further support probing intermediate states for predictive risk signals [18,19].
2.3. Lower-Bound Methodology
Although we use Fano-style bounds as the primary instrument, we position them as baseline converse statements, not universally tight bounds. In practice, we pair these bounds with modern alignment and optimization results to evaluate real-system tightness [13]. Alternative inequalities relating error, entropy, and total variation can refine tightness analysis in some regimes [15]. Beyond average 0–1 error, interactive/Fano-generalized and f-divergence-based converses can target tail-risk functionals, including CVaR-style hallucination severity bounds [20]. We view these as natural extensions of the present framework.
3. Theoretical Framework
3.1. Semantic Prediction Setup
Let X denote an input prompt, the latent fact to be recovered, and Y the model output (token sequence). We assume an evaluator map that converts Y to a discrete fact estimate for benchmarked tasks. In closed-world QA, is the canonical answer-mapping function (exact match or alias table). In open-ended generation, is instantiated by a pipeline: claim extraction, canonicalization, semantic clustering, and verifier scoring.
Definition 1
(Hallucination Rate). For fact-identification tasks with evaluator ϕ, hallucination rate is
This definition is task-level and does not require token-level exact match.
3.2. True vs Model-Induced Uncertainty
Definition 2
(True Semantic Conditional Entropy). For the data-generating distribution ,
Definition 3
(Model-Induced Surrogate Entropy). Given calibrated model scores ,
All converse bounds in this paper are written in terms of true quantities (). Model-induced quantities are used only as estimators/surrogates in experiments.
3.3. Mutual Information Convention
We use unconditional as the primary quantity. This avoids the deterministic-decoder degeneracy of , which may collapse to zero when .
Definition 4
(Semantic Information Throughput).
For stochastic decoding or randomized pipelines, conditional variants can be introduced with explicit noise variables. We model stochasticity as
where U captures decoding randomness (e.g., temperature sampling), retrieval randomness, or quantization noise. Deterministic decoding is the special case where U is constant. For blockwise analyses we similarly use
with capturing stochastic perturbations (quantization noise, dropout-like perturbations, retrieval jitter, or randomized decoding controls).
3.4. Units and Normalization
All logarithms are base 2, so entropy and mutual information are in bits. We report hallucination bounds using measured per output sequence (or per prompt-response pair). Capacity surrogates are normalized to the same unit; if per-token surrogates are used in experiments, they must be multiplied by effective output length before being compared to sequence-level bounds.
4. Three Principles (Assumption-Dependent)
4.1. Principle 1: Capacity-Constrained Semantic Throughput
Principle 1
(Capacity-Constrained Information). Under a fixed inference pipeline with finite numerical precision, finite context budget, and bounded effective representational rank, semantic throughput satisfies
We model as a surrogate envelope:
with example components
Here b is precision (bits), effective parameter budget, effective usable context, and an effective rank/degree-of-freedom proxy. Constants are task- and implementation-dependent and must be estimated empirically.
This replaces unsupported claims such as universal “ bits per head” limits.
with (A2) treated as a falsifiable modeling assumption rather than a universal law. We additionally enforce minimal structural constraints for interpretability:
Hence non-RAG envelopes use .
4.2. When A1 Fails
The Markov assumption can fail in closed-book settings where model parameters encode prior correlations with F, or in tool-augmented pipelines with extra latent state. A broader graphical model is
where Z captures latent priors, memory, or tool outputs not represented in X. In this case, the same Fano logic applies with measured under the expanded data-generating process; empirically, this requires logging/estimating variables contributing to Z.
4.3. Principle 2: Layerwise Non-Increase Under Markovization
Principle 2
(DPI Layer Monotonicity). If hidden states satisfy , then
Strict decrease requires additional lossy mechanisms (e.g., stochastic quantization, dropout, or explicit compression).
Residual paths and layer normalization can violate naive layerwise Markov assumptions if treated componentwise. Our recommendation is to analyze complete transformer blocks as stochastic maps, then apply DPI at the block level. This statement is weaker but technically robust relative to asserting fixed positive decrements at every layer. Sufficient conditions for block-level Markovization are: (i) each block can be written as with independent block noise , and (ii) no external side-channel state is injected into beyond . Autoregressive KV caching can violate (ii) unless the cache state is included in the block state variable.
4.4. Principle 3: Diminishing Returns as a Testable Hypothesis
Principle 3
(Sublinear Gains Hypothesis). Across families of models and tasks, the empirical gain in semantic throughput often follows
but this is a hypothesis to be tested, not a theorem.
We treat this hypothesis as falsified for a model family if, after controlling for optimization budget and data mixture, fitted slopes fail to remain negative over the tested range or if cross-family estimates of are unstable under pre-registered robustness checks. Pre-registration should include model families, compute budgets, decoding policies, exclusion rules, and planned regressions before any slope fitting.
5. Bounds on Hallucination Rates
5.1. Baseline Fano Bound
Let and .
Theorem 4
(Fano Hallucination Bound). For finite ,
where logs are base 2.
Proof.
Fano gives
Using and yields
Finally, by data processing through . □
5.2. Distribution-Dependent Variant
For non-uniform F, keeping explicit gives
This can be materially less pessimistic than replacing by .
Using the full multiclass Fano form also yields
which is preferable when M is large and is not close to 0 or 1.
5.3. Capacity-Constrained Corollary
Corollary 1
(Capacity Envelope Bound). If , then
The bound is informative only when .
6. Empirical Validation Protocols
We outline protocols aimed at estimable quantities.
6.1. Protocol A: Closed-World Scaling with Controlled M
Construct synthetic QA datasets with known fact set size M. Measure and estimate from confusion matrices:
Use bootstrap confidence intervals over prompts.
6.2. Protocol B: Precision and Context Stress Tests
For fixed task family, vary quantization level and usable context. Track changes in , calibration error, and to fit surrogate constants . We use reproducible proxies:
- 1.
- : smallest retained parameter count (via magnitude pruning) that keeps task score within of baseline;
- 2.
- : largest removable context span whose ablation changes task score by at most ;
- 3.
- : participation-ratio effective rank of attention maps, averaged across heads/layers.
Because and can be collinear, fits should report variance inflation diagnostics and regularized sensitivity analyses. To reduce confounding, fit using staged ablations (one factor perturbed at a time), then validate with joint perturbations and held-out tasks.
6.3. Protocol C: Layerwise Probing with Explicit Noise Controls
Probe across layers using probe classifiers or variational MI estimators. Include deterministic and stochastic inference settings separately to test when strict layerwise decrease appears. To reduce estimator-induced artifacts, we recommend at least two estimator families (e.g., InfoNCE-style lower bounds and matrix-based Rényi estimators) and synthetic controls with known MI for calibration. Report disagreement bands across estimators rather than a single curve. For selective prediction analyses, report risk-coverage curves and compare against strong abstention baselines (entropy thresholding, verifier confidence, and latent-risk probes).
6.4. Measurement Notes
For free-form generation, map outputs to fact states via a verifier and report verifier reliability. Calibrated likelihood-based surrogates should be reported with bias/variance diagnostics and sensitivity to calibration drift. Sensitivity analysis should include recalibration ablations (temperature scaling / isotonic variants) and report the resulting spread in and bound slack; conclusions are considered robust only if qualitative trends are stable across these recalibrations. For the verifier pipeline, report inter-annotator agreement (or adjudication consistency) on a sampled subset and propagate verifier uncertainty to confidence intervals.
Concrete estimators used in the proposed protocol design are:
- 1.
- plugin/confusion-matrix estimator for in closed-world settings;
- 2.
- cross-entropy-based upper/lower surrogates after calibration;
- 3.
- variational MI lower bounds for representation-level analyses.
Uncertainty is quantified with nonparametric bootstrap confidence intervals over prompts.
6.5. Pilot Synthetic Experiment Blueprint
To make feasibility explicit, we recommend a pilot experiment with and precision . For each :
- 1.
- generate balanced synthetic QA data with known latent fact F;
- 2.
- measure , , and calibration error;
- 3.
- fit in and test whether observed points satisfy the predicted lower envelope.
This pilot does not prove the framework, but it directly tests whether predicted monotonic trends hold.
6.6. Semantic Cardinality for Open-Ended Generation
For open-ended tasks with variable fact sets, we replace raw M with an effective cardinality:
This captures correlation and prompt-dependent fact support. In practice, is estimated from verifier-induced fact labels with uncertainty intervals. The open-ended estimator uses: (i) claim induction from outputs, (ii) canonical entity-relation normalization, (iii) clustering of semantically equivalent claims, and (iv) verifier-calibrated assignment to fact clusters.
7. Extension to Retrieval-Augmented Generation
Additive capacity claims for RAG are generally optimistic. A safer decomposition uses chain rule:
Hence,
Proposition 1
(RAG Bottleneck Envelope). Define
Then
This form captures retrieval quality limits, conditional transformer processing limits, and context-window bottlenecks simultaneously. For practical estimation, is proxied by retrieval quality signals (recall@k, evidence coverage, and relevance calibration), while is proxied by context-utilization diagnostics (attention-to-evidence mass, citation faithfulness, and answer sensitivity to evidence ablation). For ecosystem context and evaluation practices, see recent RAG surveys [17]. A practical calibration recipe is to construct synthetic retrieval tasks where F is known and retrieval corruption is controlled; estimate empirical relationships between retrieval metrics and , then transfer the calibrated map to real tasks with uncertainty bands.
Proposition 2
(Selective Prediction Extension). Let indicate answer/abstain decisions with coverage . Define conditional hallucination rate . Then any lower bound on unconditional error induces a coverage-risk trade-off:
so abstention can reduce conditional error only by reducing coverage.
8. Discussion
The framework supports several practical conclusions.
First, lower bounds are most useful when paired with task entropy estimates; worst-case can be too loose for skewed domains.
Second, resource scaling should be bottleneck-aware: increasing one resource axis (e.g., parameters) cannot guarantee lower hallucination if context-use efficiency or retrieval precision is limiting.
Third, evaluation design matters: if benchmarks reward confident guessing more than calibrated abstention, realized hallucination can exceed information-theoretic baselines [4].
Limitations remain. Fact extraction maps can introduce measurement noise; semantic cardinality is difficult for open-ended generation; and surrogate capacities require calibration rather than closed-form universality. In API-constrained deployments where stable token log-probabilities are unavailable, practitioners should rely on confusion-matrix estimators, verifier-based surrogates, and retrieval-side observables, and report the additional uncertainty induced by these proxies. This framework is complementary to generalization-bound analyses (e.g., PAC-Bayes/Rademacher viewpoints): those upper-bound learnability and excess risk, while our converse perspective lower-bounds irreducible error under semantic information bottlenecks [21,22,23].
9. Conclusions
This paper presents a mathematically consistent information-theoretic account of hallucination constraints in transformer systems. The key contributions are: (i) non-degenerate MI definitions, (ii) explicit distinction between true and model-implied uncertainty, (iii) assumption-dependent capacity surrogates in place of unsupported universal formulas, (iv) DPI-consistent layerwise statements, and (v) bottleneck-aware RAG bounds. Together, these results improve rigor, falsifiability, and empirical relevance while preserving the core thesis: finite semantic information throughput implies irreducible hallucination risk on sufficiently complex tasks.
Acknowledgments
The author acknowledges discussions in the broader information theory and machine learning communities that informed the framing and technical presentation of this work.
References
- Kaplan, J.; et al. Scaling Laws for Neural Language Models. arXiv 2020, arXiv:2001.08361. [Google Scholar] [CrossRef]
- Hoffmann, J.; et al. Training Compute-Optimal Large Language Models. arXiv 2022, arXiv:2203.15556. [Google Scholar] [CrossRef]
- Lewis, P.; et al. Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks. arXiv 2020, arXiv:2005.11401. [Google Scholar] [CrossRef]
- Kalai, A. T.; Nachum, O.; Vempala, S. S.; Zhang, E. Why Language Models Hallucinate. arXiv 2025, arXiv:2509.04664. [Google Scholar] [CrossRef]
- Nait Saada, T.; Naderi, A.; Tanner, J. Mind the Gap: a Spectral Analysis of Rank Collapse and Signal Propagation in Attention Layers. arXiv 2024, arXiv:2410.07799. [Google Scholar] [CrossRef]
- Heo, D.; Choi, H. Generalized Probabilistic Attention Mechanism in Transformers. arXiv 2024, arXiv:2410.15578. [Google Scholar] [CrossRef]
- Burger, M.; Kabri, S.; Korolev, Y.; Roith, T.; Weigand, L. Analysis of Mean-Field Models Arising from Self-Attention Dynamics in Transformer Architectures with Layer Normalization. arXiv 2025, arXiv:2501.03096. [Google Scholar] [CrossRef]
- Zhao, Y.-Q.; Ma, Z.-M.; Li, G. Y.; Yuan, S.; Ye, T.; Zhou, C. Semantic Rate-Distortion Theory with Applications. arXiv 2025, arXiv:2509.10061. [Google Scholar] [CrossRef]
- de Andrade, A.; Harell, A.; Bajic, I. V. Rate-Distortion Optimization for Transformer Inference. arXiv 2026, arXiv:2601.22002. [Google Scholar] [CrossRef]
- Guo, A.; Li, J. Hallucination is a Consequence of Space-Optimality: A Rate-Distortion Theorem for Membership Testing. arXiv 2026, arXiv:2602.00906. [Google Scholar] [CrossRef]
- Rathore, V.; Aneesh, S.; Singh, H. Temporal Graph Network: Hallucination Detection in Multi-Turn Conversation. arXiv 2026, arXiv:2601.03051. [Google Scholar] [CrossRef]
- Lu, Y.; Liu, Y.; Schütze, H. Relational Linearity is a Predictor of Hallucinations. arXiv 2026, arXiv:2601.11429. [Google Scholar] [CrossRef]
- Chen, P. L.; Li, X.; Chen, X.; Lin, T. Reward-free Alignment for Conflicting Objectives. arXiv 2026, arXiv:2602.02495. [Google Scholar] [CrossRef]
- Bai, B. Forget BIT, It is All about TOKEN: Towards Semantic Information Theory for LLMs. arXiv 2025, arXiv:2511.01202. [Google Scholar] [CrossRef]
- Rioul, O. The Interplay between Error, Total Variation, Alpha-Entropy and Guessing: Fano and Pinsker Direct and Reverse Inequalities. Entropy 2023, vol. 25(no. 7), 978. [Google Scholar] [CrossRef]
- Belrose, N.; Schneider-Joseph, D.; Ravfogel, S.; Cotterell, R.; Raff, E.; Biderman, S. Eliciting Latent Predictions from Transformers with the Tuned Lens. arXiv 2023, arXiv:2303.08112. [Google Scholar] [CrossRef]
- Gao, Y.; Xiong, Y.; Gao, X.; Jia, K.; Pan, J.; Bi, Y.; Dai, Y.; Sun, J.; Wang, M.; Wang, H. Retrieval-Augmented Generation for Large Language Models: A Survey. arXiv 2023, arXiv:2312.10997. [Google Scholar] [CrossRef]
- Layerwise Hallucination Dynamics in LLMs. arXiv 2024, arXiv:2403.20009. [CrossRef]
- HALT: Hallucination Assessment via Latent Testing. arXiv 2026, arXiv:2601.14210. [CrossRef]
- Interactive Fano and f-Divergence Converses for Tail-Risk Bounds. arXiv 2026, arXiv:2601.12027. [CrossRef]
- Mutual Information and Recoverability for Understanding in LLMs. arXiv 2025, arXiv:2505.23790. [CrossRef]
- A Unified Framework for Hallucination Detection and Mitigation. arXiv 2025, arXiv:2507.22915. [CrossRef]
- Representation Evolution and Fine-Tuning Effects in Transformers. arXiv 2022, arXiv:2210.12696. [CrossRef]
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.