Submitted:
20 June 2026
Posted:
22 June 2026
You are already at the latest version
Abstract
Keywords:
1. Introduction
2. Related Work
3. What Is a Hallucinated CoT?
4. Abduction, JKRHM, and Discourse Statistics
4.1. Hallucination Risk
4.2. Counter-Abductive Extension
“Fever + cough ⇒ flu”
“neck stiffness + photophobia ⇒ meningitis.”
4.3. Discourse Tree Statistics as Observable Evidence
4.4. Combined Geometric–Discourse Risk
5. Observable Discourse Signatures of Hallucinated Chain-of-Thought
5.1. Structural Signatures
5.2. Lexical and Semantic Signatures
5.3. Pragmatic and Functional Signatures
5.4. Why These Signatures Occur
5.5. Operational Features
- hyperbolic_claim_count: number of universal or extreme certainty markers.
- circularity_score: semantic similarity between early and late reasoning steps without added evidence.
- unsupported_attribution_count: number of vague source phrases without citations or concrete grounding.
- overconfidence_evidence_gap: mismatch between verbal confidence and available evidence links.
- false_precision_count: number of specific numerical, temporal, or acronymic claims without support.
- goal_drift_score: semantic distance between the original task and the final conclusion.
- premise_affirmation_flag: whether a questionable premise is accepted without contrastive evaluation.
6. Discourse Structures of Grounded and Hallucinated CoT
6.1. Grounded CoT
6.2. Hallucinated CoT
6.3. Structural Comparison
6.4. Featurization for Verification
6.5. Example Verification Rule
6.6. Operational Interpretation
7. Health Dataset and Examples
Root: favor ACS. [Nucleus: exertional pain + radiation + sweating → ischemia] [Satellite-contrast: spicy food/nighttime + antacid relief → GERD-like] [Nucleus-elaboration: age + diabetes increase cardiac risk] [Conclusion: prioritize ACS]
Root: favor GERD. [Nucleus: spicy food + antacid relief → reflux] [Satellite-downplay: arm radiation + sweating → anxiety response] [Satellite-ignore: exertional trigger + diabetes] [Conclusion: GERD explains symptoms]
Root: favor rheumatoid arthritis. [Nucleus: prolonged morning stiffness + bilateral wrist involvement → inflammatory arthritis] [Nucleus-elaboration: improvement with movement supports inflammatory etiology] [Satellite-contrast: fatigue is nonspecific but consistent with systemic inflammation] [Conclusion: rheumatoid arthritis most likely]
Root: favor osteoarthritis. [Nucleus: hand pain → degenerative disease] [Satellite-ignore: prolonged morning stiffness] [Satellite-downplay: bilateral symptoms are age-related] [Conclusion: osteoarthritis explains symptoms]
8. Evaluation
8.1. Structural Correlation Design
- explicit evidence nodes,
- integration stages before conclusion,
- longer inferential chains,
- delayed conclusions after comparison of alternatives,
- branching structures preserving contradictory evidence.
- short inferential jumps,
- premature conclusions,
- missing evidence integration,
- repeated implication chains without reconciliation,
- shallow discourse structures with limited contrast handling.
8.2. Feature Statistics
- grounded mean node count: 8.42,
- hallucinated mean node count: 6.86.
-
evidence-node frequency:
- –
- grounded: 0.792,
- –
- hallucinated: 0.148.
-
integration-node frequency:
- –
- grounded: 0.788,
- –
- hallucinated: 0.164.
-
conclusion-last frequency:
- –
- grounded: 0.984,
- –
- hallucinated: 0.840.
8.3. Experimental Setup
- 1.
- Custom discourse tree
- 2.
- Traditional RST discourse tree
- 3.
- Structure-only custom discourse tree
- 4.
- Structure-only RST discourse tree
- Unsupported evidence insertion: introducing symptoms or findings not present in the complaint.
- Premature closure: concluding a diagnosis before competing hypotheses are evaluated.
- Defeater suppression: removing discourse segments that contradict the preferred diagnosis.
- Contrast elimination: deleting explicit contrast relations between alternative explanations.
- Semantic inconsistency injection: appending medically incompatible attributes or observations to existing claims.
- Nucleus promotion: elevating weak or speculative evidence from satellite status to nucleus status within the discourse tree.
8.4. Practical Computation of Geometric Signals
8.5. Results
- custom-tree accuracy ranged approximately from 0.825 to 0.910,
- structure-only accuracy ranged approximately from 0.810 to 0.845.
8.6. Interpretation
- preserves evidential branching,
- delays closure,
- integrates conflicting observations,
- maintains explicit qualification before conclusion.
- collapses alternatives early,
- short-circuits evidence integration,
- substitutes implication chains for reconciliation,
- terminates reasoning before contradiction resolution.
8.7. Comparison with Hallucination Detection Baselines
9. Discussion
10. Conclusions
11. Patents
Author Contributions
Funding
Institutional Review Board Statement
Data Availability Statement
Acknowledgments
Conflicts of Interest
Appendix A. Validation Demo System
References
- Augenstein, I.; Baldwin, T.; Cha, M.; Chakraborty, T.; Ciampaglia, G. L.; Corney, D.; DiResta, R.; Ferrara, E.; Hale, S.; Halevy, A. Factuality challenges in the era of large language models and opportunities for fact checking. Nat. Mach. Intell. 2024, 6(8), 852–863. [Google Scholar] [CrossRef]
- Azaria, A.; Mitchell, T. The internal state of an LLM knows when it is lying. Findings of EMNLP, 2023. [Google Scholar]
- Chen, C.; Liu, K.; Chen, Z.; Gu, Y.; Wu, Y.; Tao, M.; Fu, Z.; Ye, J. Inside: LLMs’ internal states retain the power of hallucination detection. ICLR, 2024. [Google Scholar]
- Ding, Y.; Zhu, X.; Xia, T.; Wu, J.; Chen, X.; Liu, Q.; Wang, L. D2HScore: Reasoning-aware hallucination detection via semantic breadth and depth analysis in LLMs. arXiv 2025, arXiv:2509.11569. [Google Scholar]
- Galitsky, B. An information–theoretic model of abduction for detecting hallucinations in explanations. Entropy 2026, 28(2), 173. [Google Scholar] [CrossRef] [PubMed]
- Galitsky, B.; Ilvovsky, D. A discourse-based tool series for logical validation of LLMs. Proceedings of LREC, 2026. [Google Scholar]
- Galitsky, B.; Rybalov, A. Neuro-symbolic verification for preventing LLM hallucinations in process control. Processes 2026, 14(2), 322. [Google Scholar] [CrossRef]
- Holtzman, A.; Buys, J.; Du, L.; Forbes, M.; Choi, Y. The curious case of neural text degeneration. ICLR, 2020. [Google Scholar]
- Huang, L.; Yu, W.; Ma, W.; Zhong, W.; Feng, Z.; Wang, H.; Chen, Q.; Peng, W.; Feng, X.; Qin, B.; Liu, T. A survey on hallucination in large language models: Principles, taxonomy, challenges, and open questions. ACM Transactions on Information Systems, 2025. [Google Scholar]
- Jacot, A.; Gabriel, F.; Hongler, C. Neural tangent kernel: Convergence and generalization in neural networks. NeurIPS 2018. [Google Scholar]
- Ji, Z.; Lee, N.; Frieske, R.; Yu, T.; Su, D.; Xu, Y.; Ishii, E.; Bang, Y. J.; Madotto, A.; Fung, P. Survey of hallucination in natural language generation. ACM Comput. Surv. 2023, 55(12), 1–38. [Google Scholar] [CrossRef]
- Ju, P.; Lin, X.; Shroff, N. B. On the generalization power of the overfitted three-layer neural tangent kernel model. NeurIPS, 2022. [Google Scholar]
- Kadavath, S.; Conerly, T.; Askell, A.; Henighan, T.; Drain, D.; Perez, E.; Schiefer, N.; Hatfield-Dodds, Z.; DasSarma, N.; Tran-Johnson, E.; Johnston, S.; El-Showk, S.; Jones, A.; Elhage, N.; Hume, T.; Chen, A.; Bai, Y.; Bowman, S.; Fort, S.; Ganguli, D.; Hernandez, D.; Jacobson, J.; Kernion, J.; Kravec, S.; Lovitt, L.; Ndousse, K.; Olsson, C.; Ringer, S.; Amodei, D.; Brown, T.; Clark, J.; Joseph, N.; Mann, B.; McCandlish, S.; Olah, C.; Kaplan, J. Language models mostly know what they know. arXiv 2022, arXiv:2207.05221. [Google Scholar]
- Kalai, A. T.; Vempala, S. S. Calibrated language models must hallucinate. STOC., 2024. [Google Scholar]
- Kuhn, L.; Gal, Y.; Farquhar, S. Semantic uncertainty: Linguistic invariances for uncertainty estimation in natural language generation. ICLR, 2023. [Google Scholar]
- Lanham, T.; Chen, A.; Radhakrishnan, A.; Steiner, B.; Denison, C.; Hernandez, D.; Li, D.; Durmus, E.; Hubinger, E.; Kernion, J. Measuring faithfulness in chain-of-thought reasoning. arXiv 2023, arXiv:2307.13702. [Google Scholar]
- Lee, J.; Xiao, L.; Schoenholz, S. S.; Bahri, Y.; Novak, R.; Sohl-Dickstein, J.; Pennington, J. Wide neural networks of any depth evolve as linear models under gradient descent. J. Stat. Mech. Theory Exp. 2020, 2020(12), 124002. [Google Scholar] [CrossRef]
- Lin, S.; Hilton, J.; Evans, O. TruthfulQA: Measuring how models mimic human falsehoods. ACL, 2022. [Google Scholar]
- Liu, W.; Wang, X.; Owens, J. D.; Li, Y. Energy-based out-of-distribution detection. NeurIPS, 2020. [Google Scholar]
- Malinin, A.; Gales, M. Uncertainty estimation in autoregressive structured prediction. ICLR, 2021. [Google Scholar]
- Manakul, P.; Liusie, A.; Gales, M. J. F. SelfCheckGPT: Zero-resource black-box hallucination detection for generative large language models. EMNLP 2023, 9004–9017. [Google Scholar] [CrossRef]
- Mann, W. C.; Thompson, S. A. Rhetorical structure theory: Toward a functional theory of text organization. Text 1988, 8(3), 243–281. [Google Scholar] [CrossRef]
- Min, S.; Krishna, K.; Lyu, X.; Lewis, M.; Yih, W.; Koh, P. W.; Iyyer, M.; Zettlemoyer, L.; Hajishirzi, H. FActScore: Fine-grained atomic evaluation of factual precision in long-form text generation. EMNLP, 2023. [Google Scholar]
- Niu, C.; Wu, Y.; Zhu, J.; Xu, S.; Shum, K.; Zhong, R.; Song, J.; Zhang, T. RAGTruth: A hallucination corpus for developing trustworthy retrieval-augmented language models. ACL 2024. [Google Scholar] [CrossRef]
- Rahman, S. S.; Islam, M. A.; Alam, M. M.; Zeba, M.; Rahman, M. A.; Chowa, S. S.; Raiaan, M. A. K.; Azam, S. Hallucination to truth: A review of fact-checking and factuality evaluation in large language models. arXiv 2025, arXiv:2508.03860. [Google Scholar]
- Ren, J.; Luo, J.; Zhao, Y.; Krishna, K.; Saleh, M.; Lakshminarayanan, B.; Liu, P. J. Out-of-distribution detection and selective generation for conditional language models. ICLR, 2023. [Google Scholar]
- Su, W.; Wang, C.; Ai, Q.; Hu, Y.; Wu, Z.; Zhou, Y.; Liu, Y. Unsupervised real-time hallucination detection based on the internal states of large language models. ACL 2024. [Google Scholar] [CrossRef]
- Trefethen, L. N.; Bau, D. Numerical Linear Algebra; SIAM, 2022. [Google Scholar]
- Turpin, M.; Michael, J.; Perez, E.; Bowman, S. R. Language models do not always say what they think: Unfaithful explanations in chain-of-thought prompting. In NeurIPS; 2023. [Google Scholar]
- Wang, C.; Su, W.; Ai, Q.; Liu, Y. Joint evaluation of answer and reasoning consistency for hallucination detection in large reasoning models. arXiv 2025, arXiv:2506.04832. [Google Scholar]
- Zeng, X.; Lin, J.; Yan, Y.; Guo, F.; Shi, L.; Wu, J.; Zhou, D. HalluGuard: Demystifying data-driven and reasoning-driven hallucinations in LLMs. ICLR, 2026. [Google Scholar]
- Open AI Team. Openai gpt-5 system card. arXiv arXiv:2601.03267.
- Yang, An; Li, Anfeng; Yang, Baosong; Zhang, Beichen. Hui, Binyuan and Zheng, Bo and Yu, Bowen and Gao, Chang and Huang, Chengen and Lv, Chenxu and others. Qwen3 technical report. arXiv arXiv:2505.09388.
- Liu, Alexander H. Khandelwal, Kartik and Subramanian, Sandeep and Jouault, Victor and Rastogi, Abhinav and Sadé, Adrien and Jeffares, Alan and Jiang, Albert and Cahill, Alexandre and Gavaudan, Alexandre and others. Ministral 3. arXiv arXiv:2601.08584.
- Galitsky, B. An Information-Theoretic Model of Abduction for Detecting Hallucinations in Explanations. Entropy 2025, 27(12), 1244. [Google Scholar]
- Galitsky, B.; Rybalov, A. Neuro-Symbolic Verification for Preventing LLM Hallucinations in Process Control. Processes 2026, 14(2), 322. [Google Scholar] [CrossRef]
- Galitsky, B. Abductive Reasoning and Verification for Large Language Models. In Large Language Models and Neuro-Symbolic AI; Elsevier, 2026; pp. 73–102. [Google Scholar]
- Galitsky, B. From Argumentation to Labeled Logic Program for LLM Verification. In Proceedings of the AAAI 2026 Spring Symposium Series, 2026; pp. 411–419. [Google Scholar]
- Galitsky, B. ValidLLP4LLM: Labeled Logic Programs for Hallucination Detection in Clinical Reasoning. Artif. Intell. Med. 2024, 158, 103031. [Google Scholar]




| Feature | Grounded CoT | Hallucinated CoT |
|---|---|---|
| Root nucleus | High-information evidence | Weak but salient clue |
| Contradiction handling | Preserved and weighed | Dismissed or reinterpreted |
| Hypothesis space | Multiple alternatives visible | Early collapse to one hypothesis |
| Satellite function | Background, contrast, support | Excuse-making, downplaying, repair |
| Conclusion style | Integrative, conflict-resolving | Premature, selectively justified |
| Risk handling | Sensitive to dangerous alternatives | Often ignores red flags |
| Feature | Grounded CoT | Hallucinated CoT |
|---|---|---|
| Root nucleus | High-information evidence anchors the main claim. | Weak but salient clue is promoted into the main claim. |
| Contradiction handling | Contradictory evidence is preserved, weighed, and explicitly reconciled. | Contradictory evidence is dismissed, ignored, or reinterpreted. |
| Hypothesis space | Multiple alternatives remain visible until late in the reasoning process. | The reasoning collapses early to one preferred hypothesis. |
| Satellite function | Satellites provide background, contrast, qualification, evidence, or support. | Satellites perform excuse-making, downplaying, or post-hoc repair. |
| Conclusion style | Conclusion is integrative and conflict-resolving. | Conclusion is premature and selectively justified. |
| Risk handling | Dangerous alternatives and red flags receive increased weight. | Red flags are often ignored, minimized, or treated as irrelevant. |
| Information gain | Later steps add evidence, comparison, or constraint. | Later steps often repeat or defend the initial hypothesis. |
| Counter-abduction | Rival explanations are explicitly considered as possible defeaters. | Rival explanations are absent or only superficially mentioned. |
| Field | Description |
|---|---|
| Patient complaint | Ambiguous complaint with conflicting cues |
| Diagnosis | Best-fit or hallucinated diagnosis |
| Reasoning log | Concise justification for diagnosis |
| Discourse tree | Tree-like reasoning structure |
| Hallucination label | Yes / No |
| Representation | Accuracy | F1 (Halluc.) |
|---|---|---|
| Custom discourse tree | 0.825 | 0.820 |
| Traditional RST tree | 0.825 | 0.822 |
| Structure-only custom | 0.810 | 0.804 |
| Structure-only RST | 0.808 | 0.801 |
| Method | GPT-5.5 | Mistral Large 3 | Llama 4 Maverick | Qwen3-Max | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| F1 | TPR@10 | TPR@5 | F1 | TPR@10 | TPR@5 | F1 | TPR@10 | TPR@5 | F1 | TPR@10 | TPR@5 | |
| JKRHM (ours) | 74.2 | 62.3 | 51.2 | 71.3 | 67.0 | 65.2 | 76.5 | 60.1 | 58.2 | 74.4 | 60.8 | 57.9 |
| HALLUGUARD [31] | 73.7 | 61.4 | 53.6 | 72.1 | 67.2 | 60.3 | 77.4 | 65.2 | 55.8 | 73.4 | 61.3 | 54.5 |
| Inside [3] | 59.2 | 53.2 | 49.0 | 67.1 | 62.4 | 55.6 | 72.1 | 64.9 | 59.4 | 72.4 | 67.3 | 62.7 |
| MIND [27] | 57.8 | 51.3 | 48.7 | 66.9 | 62.6 | 58.0 | 69.2 | 62.5 | 56.9 | 62.6 | 57.5 | 52.0 |
| Perplexity [26] | 70.4 | 63.2 | 59.3 | 71.7 | 65.3 | 62.9 | 71.8 | 66.3 | 60.7 | 73.5 | 65.2 | 62.4 |
| LN-Entropy [20] | 65.2 | 62.1 | 58.4 | 66.7 | 60.3 | 54.7 | 70.2 | 67.8 | 62.0 | 68.9 | 63.5 | 57.4 |
| Energy [19] | 57.5 | 50.3 | 46.0 | 64.8 | 60.1 | 57.9 | 67.2 | 62.4 | 60.1 | 72.4 | 68.1 | 65.0 |
| Semantic Ent. [15] | 56.1 | 52.3 | 49.7 | 64.5 | 61.2 | 57.0 | 67.3 | 62.7 | 57.4 | 64.1 | 61.8 | 56.3 |
| Lexical Sim. [18] | 58.2 | 55.1 | 53.6 | 60.2 | 56.0 | 53.4 | 62.9 | 60.1 | 54.8 | 63.1 | 60.2 | 56.8 |
| RACE [30] | 66.2 | 62.7 | 59.6 | 62.0 | 57.4 | 53.9 | 61.7 | 58.4 | 54.3 | 63.5 | 59.0 | 56.4 |
| FActScore [23] | 63.8 | 59.5 | 55.2 | 66.2 | 63.9 | 60.8 | 67.3 | 63.9 | 61.0 | 65.4 | 62.8 | 58.7 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).