Preprint
Article

This version is not peer-reviewed.

Semantic Thermodynamics of Transformer Architectures: A Framework for Understanding Hallucination Constraints

Submitted:

22 February 2026

Posted:

27 February 2026

You are already at the latest version

Abstract
We develop \emph{Semantic Thermodynamics}, an information-theoretic framework for analyzing hallucinations in transformer systems under finite resources. The central object is mutual information between latent facts and model outputs, together with Fano-style lower bounds on semantic error. We clarify the stochastic assumptions required for non-degenerate information measures, distinguish true data-generating uncertainty from model-implied uncertainty, and replace unsupported hard capacity formulas with explicit capacity surrogates tied to precision, context budget, and effective representational rank. Under standard identification assumptions, we derive a baseline bound \begin{equation*} H_R \geq \max\left\{0,\,1-\frac{I(F;Y)+1}{\log M}\right\}, \end{equation*} where \(H_R\) is hallucination rate, \(F\) is the latent semantic fact, \(Y\) is model output, and \(M\) is semantic cardinality. We also provide a distribution-dependent variant and a bottleneck-aware extension for retrieval-augmented generation. This paper contributes a mathematically consistent formulation, a tighter assumptions section, and concrete empirical protocols for estimation and falsification.
Keywords: 
;  ;  ;  ;  ;  

1. Introduction

Hallucination remains a central failure mode of modern language models: outputs can be fluent and high-confidence while semantically incorrect. Mitigation methods including alignment, finetuning, and retrieval improve performance but do not eliminate failures. This pattern motivates a complementary question: beyond optimization quality, what lower bounds arise from finite information-processing resources?
This paper proposes an information-theoretic treatment of that question. We model semantic prediction as transmission of latent facts through a bounded inference pipeline. We then use mutual information and Fano-type converse bounds to characterize unavoidable error regimes.
Our goal is not to claim that architecture alone determines hallucinations. Rather, we isolate one layer of explanation: when effective semantic information throughput is below task complexity, nonzero semantic error is unavoidable even with idealized training. Throughout, “thermodynamics” is used as a metaphorical organizing label for resource-constrained information processing; the formal results in this paper are information-theoretic converses rather than a physical thermodynamic formalism.

1.1. Contributions

This paper makes the following contributions:
1.
It standardizes mutual information notation and removes degenerate uses of I ( F ; Y ∣ X ) under deterministic decoding.
2.
It separates true semantic distributions from model-induced distributions to avoid circular definitions.
3.
It replaces inconsistent capacity formulas with explicit, assumption-dependent surrogates.
4.
It weakens strict layerwise-loss claims to DPI-based monotonicity, with strict decrease only under explicit stochastic perturbations.
5.
It adds distribution-dependent and bottleneck-aware bound variants, including a non-additive RAG analysis.
6.
It expands related work to include semantic rate-distortion, memory-limited hallucination frontiers, and f-divergence minimax lower bounds.

3. Theoretical Framework

3.1. Semantic Prediction Setup

Let X denote an input prompt, F ∈ F the latent fact to be recovered, and Y the model output (token sequence). We assume an evaluator map ϕ that converts Y to a discrete fact estimate F ^ = ϕ ( Y ) for benchmarked tasks. In closed-world QA, ϕ is the canonical answer-mapping function (exact match or alias table). In open-ended generation, ϕ is instantiated by a pipeline: claim extraction, canonicalization, semantic clustering, and verifier scoring.
Definition 1
(Hallucination Rate). For fact-identification tasks with evaluator ϕ, hallucination rate is
H R : = Pr [ F ^ ≠ F ] .
This definition is task-level and does not require token-level exact match.

3.2. True vs Model-Induced Uncertainty

Definition 2
(True Semantic Conditional Entropy). For the data-generating distribution p ★ ( f ∣ x ) ,
H ★ ( F ∣ X = x ) : = − ∑ f p ★ ( f ∣ x ) log p ★ ( f ∣ x ) .
Definition 3
(Model-Induced Surrogate Entropy). Given calibrated model scores q θ ( f ∣ x ) ,
H ˜ θ ( F ∣ X = x ) : = − ∑ f q θ ( f ∣ x ) log q θ ( f ∣ x ) .
All converse bounds in this paper are written in terms of true quantities ( p ★ ). Model-induced quantities are used only as estimators/surrogates in experiments.

3.3. Mutual Information Convention

We use unconditional I ( F ; Y ) as the primary quantity. This avoids the deterministic-decoder degeneracy of I ( F ; Y ∣ X ) , which may collapse to zero when Y = g ( X ) .
Definition 4
(Semantic Information Throughput). 
I S : = I ( F ; Y ) .
For stochastic decoding or randomized pipelines, conditional variants can be introduced with explicit noise variables. We model stochasticity as
Y = g θ ( X , U ) , U ∼ p ( u ) ,
where U captures decoding randomness (e.g., temperature sampling), retrieval randomness, or quantization noise. Deterministic decoding is the special case where U is constant. For blockwise analyses we similarly use
H ( ℓ + 1 ) = T ℓ ( H ( ℓ ) , ξ ℓ ) ,
with ξ ℓ capturing stochastic perturbations (quantization noise, dropout-like perturbations, retrieval jitter, or randomized decoding controls).

3.4. Units and Normalization

All logarithms are base 2, so entropy and mutual information are in bits. We report hallucination bounds using I ( F ; Y ) measured per output sequence (or per prompt-response pair). Capacity surrogates are normalized to the same unit; if per-token surrogates are used in experiments, they must be multiplied by effective output length before being compared to sequence-level bounds.

4. Three Principles (Assumption-Dependent)

4.1. Principle 1: Capacity-Constrained Semantic Throughput

Principle 1
(Capacity-Constrained Information). Under a fixed inference pipeline with finite numerical precision, finite context budget, and bounded effective representational rank, semantic throughput satisfies
I ( F ; Y ) ≤ C surr .
We model C surr as a surrogate envelope:
C surr : = min { C param , C ctx , C arith } ,
with example components
C param = α b P eff ,
C ctx = β b L eff ,
C arith = γ b r attn .
Here b is precision (bits), P eff effective parameter budget, L eff effective usable context, and r attn an effective rank/degree-of-freedom proxy. Constants ( α , β , γ ) are task- and implementation-dependent and must be estimated empirically.
This replaces unsupported claims such as universal “ log   d bits per head” limits.
( Assumption A 1 ) F → X → Y , ( A 2 ) I ( F ; Y ) ≤ C surr ,
with (A2) treated as a falsifiable modeling assumption rather than a universal law. We additionally enforce minimal structural constraints for interpretability:
C surr ≥ 0 , ∂ C surr ∂ b , ∂ C surr ∂ P eff , ∂ C surr ∂ L eff , ∂ C surr ∂ r attn ≥ 0 , I ( F ; Y ) ≤ H ( F ) .
Hence non-RAG envelopes use min { C surr , H ( F ) } .

4.2. When A1 Fails

The Markov assumption F → X → Y can fail in closed-book settings where model parameters encode prior correlations with F, or in tool-augmented pipelines with extra latent state. A broader graphical model is
F → ( X , Z ) → Y ,
where Z captures latent priors, memory, or tool outputs not represented in X. In this case, the same Fano logic applies with I ( F ; Y ) measured under the expanded data-generating process; empirically, this requires logging/estimating variables contributing to Z.

4.3. Principle 2: Layerwise Non-Increase Under Markovization

Principle 2
(DPI Layer Monotonicity). If hidden states satisfy F → H ( ℓ ) → H ( ℓ + 1 ) , then
I ( F ; H ( ℓ + 1 ) ) ≤ I ( F ; H ( ℓ ) ) .
Strict decrease requires additional lossy mechanisms (e.g., stochastic quantization, dropout, or explicit compression).
Residual paths and layer normalization can violate naive layerwise Markov assumptions if treated componentwise. Our recommendation is to analyze complete transformer blocks as stochastic maps, then apply DPI at the block level. This statement is weaker but technically robust relative to asserting fixed positive decrements Δ ℓ > 0 at every layer. Sufficient conditions for block-level Markovization are: (i) each block can be written as H ( ℓ + 1 ) = T ℓ ( H ( ℓ ) , ξ ℓ ) with independent block noise ξ ℓ , and (ii) no external side-channel state is injected into T ℓ beyond H ( ℓ ) . Autoregressive KV caching can violate (ii) unless the cache state is included in the block state variable.

4.4. Principle 3: Diminishing Returns as a Testable Hypothesis

Principle 3
(Sublinear Gains Hypothesis). Across families of models and tasks, the empirical gain in semantic throughput often follows
∂ I ( F ; Y ) ∂ C surr ∝ C surr − δ , δ > 0 ,
but this is a hypothesis to be tested, not a theorem.
We treat this hypothesis as falsified for a model family if, after controlling for optimization budget and data mixture, fitted slopes fail to remain negative over the tested C surr range or if cross-family estimates of δ are unstable under pre-registered robustness checks. Pre-registration should include model families, compute budgets, decoding policies, exclusion rules, and planned regressions before any slope fitting.

5. Bounds on Hallucination Rates

5.1. Baseline Fano Bound

Let M = | F | and F ^ = ϕ ( Y ) .
Theorem 4
(Fano Hallucination Bound). For finite M ≥ 2 ,
H R ≥ 1 − I ( F ; F ^ ) + 1 log M ≥ 1 − I ( F ; Y ) + 1 log M ,
where logs are base 2.
Proof. 
Fano gives
H ( F ∣ F ^ ) ≤ h 2 ( H R ) + H R log ( M − 1 ) ≤ 1 + H R log M .
Using H ( F ∣ F ^ ) = H ( F ) − I ( F ; F ^ ) and H ( F ) ≤ log M yields
H R ≥ 1 − I ( F ; F ^ ) + 1 log M .
Finally, I ( F ; F ^ ) ≤ I ( F ; Y ) by data processing through ϕ . □

5.2. Distribution-Dependent Variant

For non-uniform F, keeping H ( F ) explicit gives
H R ≥ H ( F ) − I ( F ; Y ) − 1 log ( M − 1 ) .
This can be materially less pessimistic than replacing H ( F ) by log M .
Using the full multiclass Fano form also yields
H ( F ∣ F ^ ) ≤ h 2 ( H R ) + H R log ( M − 1 ) ,
which is preferable when M is large and H R is not close to 0 or 1.

5.3. Capacity-Constrained Corollary

Corollary 1
(Capacity Envelope Bound). If I ( F ; Y ) ≤ min { C surr , H ( F ) } , then
H R ≥ max 0 , 1 − min { C surr , H ( F ) } + 1 log M .
The bound is informative only when min { C surr , H ( F ) } < log M − 1 .
Equivalent informativity condition : I ( F ; Y ) < log M − 1 .

6. Empirical Validation Protocols

We outline protocols aimed at estimable quantities.

6.1. Protocol A: Closed-World Scaling with Controlled M

Construct synthetic QA datasets with known fact set size M. Measure H R and estimate I ( F ; F ^ ) from confusion matrices:
I ^ ( F ; F ^ ) = H ^ ( F ) − H ^ ( F ∣ F ^ ) .
Use bootstrap confidence intervals over prompts.

6.2. Protocol B: Precision and Context Stress Tests

For fixed task family, vary quantization level and usable context. Track changes in H R , calibration error, and I ^ ( F ; F ^ ) to fit surrogate constants ( α , β , γ ) . We use reproducible proxies:
1.
P eff ( ϵ ) : smallest retained parameter count (via magnitude pruning) that keeps task score within ϵ of baseline;
2.
L eff ( ϵ ) : largest removable context span whose ablation changes task score by at most ϵ ;
3.
r attn : participation-ratio effective rank of attention maps, averaged across heads/layers.
Because L eff and r attn can be collinear, fits should report variance inflation diagnostics and regularized sensitivity analyses. To reduce confounding, fit α , β , γ using staged ablations (one factor perturbed at a time), then validate with joint perturbations and held-out tasks.

6.3. Protocol C: Layerwise Probing with Explicit Noise Controls

Probe I ( F ; H ( ℓ ) ) across layers using probe classifiers or variational MI estimators. Include deterministic and stochastic inference settings separately to test when strict layerwise decrease appears. To reduce estimator-induced artifacts, we recommend at least two estimator families (e.g., InfoNCE-style lower bounds and matrix-based Rényi estimators) and synthetic controls with known MI for calibration. Report disagreement bands across estimators rather than a single curve. For selective prediction analyses, report risk-coverage curves and compare against strong abstention baselines (entropy thresholding, verifier confidence, and latent-risk probes).

6.4. Measurement Notes

For free-form generation, map outputs to fact states via a verifier and report verifier reliability. Calibrated likelihood-based surrogates should be reported with bias/variance diagnostics and sensitivity to calibration drift. Sensitivity analysis should include recalibration ablations (temperature scaling / isotonic variants) and report the resulting spread in I ^ ( F ; F ^ ) and bound slack; conclusions are considered robust only if qualitative trends are stable across these recalibrations. For the verifier pipeline, report inter-annotator agreement (or adjudication consistency) on a sampled subset and propagate verifier uncertainty to I ^ ( F ; F ^ ) confidence intervals.
Concrete estimators used in the proposed protocol design are:
1.
plugin/confusion-matrix estimator for I ^ ( F ; F ^ ) in closed-world settings;
2.
cross-entropy-based upper/lower surrogates after calibration;
3.
variational MI lower bounds for representation-level analyses.
Uncertainty is quantified with nonparametric bootstrap confidence intervals over prompts.

6.5. Pilot Synthetic Experiment Blueprint

To make feasibility explicit, we recommend a pilot experiment with M ∈ { 16 , 32 , 64 , 128 } and precision b ∈ { 16 , 8 , 4 } . For each ( M , b ) :
1.
generate balanced synthetic QA data with known latent fact F;
2.
measure H R , I ^ ( F ; F ^ ) , and calibration error;
3.
fit ( α , β , γ ) in C surr and test whether observed points satisfy the predicted lower envelope.
This pilot does not prove the framework, but it directly tests whether predicted monotonic trends hold.

6.6. Semantic Cardinality for Open-Ended Generation

For open-ended tasks with variable fact sets, we replace raw M with an effective cardinality:
M eff ( x ) : = exp H ( F ∣ X = x ) , M ¯ eff : = exp E x [ H ( F ∣ X = x ) ] .
This captures correlation and prompt-dependent fact support. In practice, H ( F ∣ X = x ) is estimated from verifier-induced fact labels with uncertainty intervals. The open-ended estimator uses: (i) claim induction from outputs, (ii) canonical entity-relation normalization, (iii) clustering of semantically equivalent claims, and (iv) verifier-calibrated assignment to fact clusters.

7. Extension to Retrieval-Augmented Generation

Additive capacity claims for RAG are generally optimistic. A safer decomposition uses chain rule:
I ( F ; Y , R ) = I ( F ; R ) + I ( F ; Y ∣ R ) .
Hence,
I ( F ; Y , R ) ≤ min { H ( F ) , I ( F ; R ) + C tr ∣ R , C ctx RAG } .
Proposition 1
(RAG Bottleneck Envelope). Define
C eff RAG : = min { H ( F ) , I ( F ; R ) + C tr ∣ R , C ctx RAG } .
Then
H R ≥ max 0 , 1 − C eff RAG + 1 log M .
This form captures retrieval quality limits, conditional transformer processing limits, and context-window bottlenecks simultaneously. For practical estimation, I ( F ; R ) is proxied by retrieval quality signals (recall@k, evidence coverage, and relevance calibration), while C tr ∣ R is proxied by context-utilization diagnostics (attention-to-evidence mass, citation faithfulness, and answer sensitivity to evidence ablation). For ecosystem context and evaluation practices, see recent RAG surveys [17]. A practical calibration recipe is to construct synthetic retrieval tasks where F is known and retrieval corruption is controlled; estimate empirical relationships between retrieval metrics and I ( F ; R ) , then transfer the calibrated map to real tasks with uncertainty bands.
Proposition 2
(Selective Prediction Extension). Let q ( x ) ∈ { 0 , 1 } indicate answer/abstain decisions with coverage κ = E [ q ( X ) ] . Define conditional hallucination rate H R ( κ ) = Pr [ F ^ ≠ F ∣ q ( X ) = 1 ] . Then any lower bound on unconditional error induces a coverage-risk trade-off:
Pr [ F ^ ≠ F ] = κ H R ( κ ) ≥ max 0 , 1 − I ( F ; Y ) + 1 log M ,
so abstention can reduce conditional error only by reducing coverage.

8. Discussion

The framework supports several practical conclusions.
First, lower bounds are most useful when paired with task entropy estimates; worst-case log M can be too loose for skewed domains.
Second, resource scaling should be bottleneck-aware: increasing one resource axis (e.g., parameters) cannot guarantee lower hallucination if context-use efficiency or retrieval precision is limiting.
Third, evaluation design matters: if benchmarks reward confident guessing more than calibrated abstention, realized hallucination can exceed information-theoretic baselines [4].
Limitations remain. Fact extraction maps can introduce measurement noise; semantic cardinality is difficult for open-ended generation; and surrogate capacities require calibration rather than closed-form universality. In API-constrained deployments where stable token log-probabilities are unavailable, practitioners should rely on confusion-matrix estimators, verifier-based surrogates, and retrieval-side observables, and report the additional uncertainty induced by these proxies. This framework is complementary to generalization-bound analyses (e.g., PAC-Bayes/Rademacher viewpoints): those upper-bound learnability and excess risk, while our converse perspective lower-bounds irreducible error under semantic information bottlenecks [21,22,23].

9. Conclusions

This paper presents a mathematically consistent information-theoretic account of hallucination constraints in transformer systems. The key contributions are: (i) non-degenerate MI definitions, (ii) explicit distinction between true and model-implied uncertainty, (iii) assumption-dependent capacity surrogates in place of unsupported universal formulas, (iv) DPI-consistent layerwise statements, and (v) bottleneck-aware RAG bounds. Together, these results improve rigor, falsifiability, and empirical relevance while preserving the core thesis: finite semantic information throughput implies irreducible hallucination risk on sufficiently complex tasks.

Acknowledgments

The author acknowledges discussions in the broader information theory and machine learning communities that informed the framing and technical presentation of this work.

References

  1. Kaplan, J.; et al. Scaling Laws for Neural Language Models. arXiv 2020, arXiv:2001.08361. [Google Scholar] [CrossRef]
  2. Hoffmann, J.; et al. Training Compute-Optimal Large Language Models. arXiv 2022, arXiv:2203.15556. [Google Scholar] [CrossRef]
  3. Lewis, P.; et al. Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks. arXiv 2020, arXiv:2005.11401. [Google Scholar] [CrossRef]
  4. Kalai, A. T.; Nachum, O.; Vempala, S. S.; Zhang, E. Why Language Models Hallucinate. arXiv 2025, arXiv:2509.04664. [Google Scholar] [CrossRef]
  5. Nait Saada, T.; Naderi, A.; Tanner, J. Mind the Gap: a Spectral Analysis of Rank Collapse and Signal Propagation in Attention Layers. arXiv 2024, arXiv:2410.07799. [Google Scholar] [CrossRef]
  6. Heo, D.; Choi, H. Generalized Probabilistic Attention Mechanism in Transformers. arXiv 2024, arXiv:2410.15578. [Google Scholar] [CrossRef]
  7. Burger, M.; Kabri, S.; Korolev, Y.; Roith, T.; Weigand, L. Analysis of Mean-Field Models Arising from Self-Attention Dynamics in Transformer Architectures with Layer Normalization. arXiv 2025, arXiv:2501.03096. [Google Scholar] [CrossRef]
  8. Zhao, Y.-Q.; Ma, Z.-M.; Li, G. Y.; Yuan, S.; Ye, T.; Zhou, C. Semantic Rate-Distortion Theory with Applications. arXiv 2025, arXiv:2509.10061. [Google Scholar] [CrossRef]
  9. de Andrade, A.; Harell, A.; Bajic, I. V. Rate-Distortion Optimization for Transformer Inference. arXiv 2026, arXiv:2601.22002. [Google Scholar] [CrossRef]
  10. Guo, A.; Li, J. Hallucination is a Consequence of Space-Optimality: A Rate-Distortion Theorem for Membership Testing. arXiv 2026, arXiv:2602.00906. [Google Scholar] [CrossRef]
  11. Rathore, V.; Aneesh, S.; Singh, H. Temporal Graph Network: Hallucination Detection in Multi-Turn Conversation. arXiv 2026, arXiv:2601.03051. [Google Scholar] [CrossRef]
  12. Lu, Y.; Liu, Y.; Schütze, H. Relational Linearity is a Predictor of Hallucinations. arXiv 2026, arXiv:2601.11429. [Google Scholar] [CrossRef]
  13. Chen, P. L.; Li, X.; Chen, X.; Lin, T. Reward-free Alignment for Conflicting Objectives. arXiv 2026, arXiv:2602.02495. [Google Scholar] [CrossRef]
  14. Bai, B. Forget BIT, It is All about TOKEN: Towards Semantic Information Theory for LLMs. arXiv 2025, arXiv:2511.01202. [Google Scholar] [CrossRef]
  15. Rioul, O. The Interplay between Error, Total Variation, Alpha-Entropy and Guessing: Fano and Pinsker Direct and Reverse Inequalities. Entropy 2023, vol. 25(no. 7), 978. [Google Scholar] [CrossRef]
  16. Belrose, N.; Schneider-Joseph, D.; Ravfogel, S.; Cotterell, R.; Raff, E.; Biderman, S. Eliciting Latent Predictions from Transformers with the Tuned Lens. arXiv 2023, arXiv:2303.08112. [Google Scholar] [CrossRef]
  17. Gao, Y.; Xiong, Y.; Gao, X.; Jia, K.; Pan, J.; Bi, Y.; Dai, Y.; Sun, J.; Wang, M.; Wang, H. Retrieval-Augmented Generation for Large Language Models: A Survey. arXiv 2023, arXiv:2312.10997. [Google Scholar] [CrossRef]
  18. Layerwise Hallucination Dynamics in LLMs. arXiv 2024, arXiv:2403.20009. [CrossRef]
  19. HALT: Hallucination Assessment via Latent Testing. arXiv 2026, arXiv:2601.14210. [CrossRef]
  20. Interactive Fano and f-Divergence Converses for Tail-Risk Bounds. arXiv 2026, arXiv:2601.12027. [CrossRef]
  21. Mutual Information and Recoverability for Understanding in LLMs. arXiv 2025, arXiv:2505.23790. [CrossRef]
  22. A Unified Framework for Hallucination Detection and Mitigation. arXiv 2025, arXiv:2507.22915. [CrossRef]
  23. Representation Evolution and Fine-Tuning Effects in Transformers. arXiv 2022, arXiv:2210.12696. [CrossRef]
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.