Submitted:
24 July 2023
Posted:
25 July 2023
You are already at the latest version
Abstract
Keywords:
1. Introduction
- Fact checking via web mining adjusted to LLM content versus human-made intentional misinformation
- Applicability to texts of various genres and mixture of formats and genres in a document
- Application to the whole document with opinionated data, background, captions, results, design notes etc.
- Finding proper level of granularity for fact-checking: from phrases to sentences, with contextualization
- Verification against various sources of different modalities with mutual inconsistency
- Verification against abstract canons, generalizations, rules, principles
1.1. Why LLMs Hallucinate
1.2. Introductory Example

- 1)
- Fact – check LLM sentence by sentence and phrase (with attribute) by phrase
- 2)
- Start with the first sentence. Extract the main entity ‘Veryovkina cave’. Confirm that it is the deepest cave: we search the web for it and find a confirmation on Wikipedia.
- 3)
- Proceed to the second sentence. Extracting ‘led by<entity>’ and ‘ with <entity> leading the dives’ phrases. Both phrases are neither confirmed by the web nor by the Wikipedia page identified by the first sentence.
- 4)
- Coming to conclusion that the second sentence is totally wrong (is a complete hallucination). The hallucinated entities for persons and teams are real but associated with a different geo entity (cave).
- 5)
- It turns out that correct entities needs to be extracted from Wiki page identified by the query derived from the first sentence.
- 6)
- We use a retriever to obtain the value for ‘Expedition leader’ and substitute it into the second LLM sentence to correct it. Both organization name and individual (leader) name needs to be substituted.
- 1)
- Consolidating evidence from external knowledge for the LLM to generate responses grounded in evidence. LLM is considered as a black box and is called only once.
- 2)
- Revising LLM’s (candidate) responses using automated feedback (Peng et al. 2023). LLM is called iteratively, each time with additional information (Section 3).


2. Related Work
3. Iterative Mode
- 1)
- Given a user query, Truth-O-Meter first retrieves evidence from external knowledge (e.g., Web or task-specific datasets) and, if necessary, further fuses evidence by linking obtained raw evidence with related context (e.g., information of the entity “Veryovkina Cave”) and performing reasoning to form evidence chains (e.g., extracting information from list of references on a Wiki page).
- 2)
- Then, Truth-O-Meter queries ChatGPT using a prompt that contains the consolidated evidence for ChatGPT to generate a candidate response grounded in external knowledge (evidence).
- 3)
- Truth-O-Meter verifies the candidate response e.g., by checking whether it is inconsistent with evidence. If so, Truth-O-Meter generates a feedback message (e.g., about the caving team). The message is used to revise the prompt to query ChatGPT again.
- 4)
- The process iterates until a candidate response passes the verification and is sent to the user.
- 1)
- acquiring evidence e for q from external knowledge;
- 2)
- calling the LLM to generate a current candidate response, and
- 3)
- sending a response to users if all available knowledge has been applied.
- -
- A is a set of actions that Policy manager selects, including Information seeking manager to consult evidence from external knowledge and calling Prompting manager to query the LLM again to proceed with next iteration.
- -
- R(s; a) is the external reward received after taking action a in state s, which is provided by the environment (e.g., users or simulators);
- -
- πq is computed using a logic programming formalism of reasoning about actions, computing possible actions and their resultant states iteratively (Levesque et al. 1997, Galitsky 2006).




4. Handling Multiple Mutually Inconsistent Facts Obtained from Authoritative Sources

- The defeasible reasoner, which is responsible for deriving arguments from the defeasible logic program, constructing dialectical trees, and analyzing defeating relationships to respond to user queries.
- The DeLP Viewer, which serves as the visualization module and the tool presented in the article.
- 1)
- Li is a fact in Π, or
- 2)
- there exists a rule Ri in P (strict or defeasible) with head Li and body B1,B2, . . .,
- 1)
- there exists a defeasible derivation for h from (Π ∪ A);
- 2)
- the set (Π ∪ A) is non-contradictory; and
- 3)
- A is minimal: there is no proper subset A0 of A such that A0 satisfies conditions (1) and (2).
- 1)
- The root of the tree is labeled with <A0, h0>
- 2)
- Let N be a non-root vertex of the tree labeled <An, hn> and Λ . [<A0, h0>, <A1, h1>, . . ., <An, hn>] (the sequence of labels of the path from the root to N).
- 1)
- All leaves are to be labeled as U-nodes (undefeated nodes).
- 2)
- Any inner node is to be labeled as a U-node whenever all of its associated children nodes are labeled as D-nodes.
- 3)
- Any inner node is to be labeled as a D-node whenever at least one of its associated children nodes is labeled as U-node.

5. Correcting Factual Errors in Syntactic and Semantic Spaces
- 1)
- A set of rules applied to a dependency parse to identify the heads of entity mentions.
- 2)
- Rules based on token tags and text to identify the complete mention.
- 3)
- A system to link the identified entity mentions to corresponding structured evidence, verifying the factual nature of the text and distinguishing it from hallucinations.

5.1. Fact-Checking by Question Answering against Sources and Syntax-Semantic Alignment
- 1)
- Baseline: we consider tokens which occurs in T1 but do not occur in T2 and therefore unconfirmed (should be rejected).
- 2)
- Instead of such tokens we attempt to obtain attribute values which occurs in T1 but do not occur in T2 and therefore unconfirmed (should be rejected). To achieve this, we generate questions and then use a question answering model to obtain answers as entities, attributes and values to find those which occurs in T1 but do not occur in T2 and therefore unconfirmed (should be rejected).
- 3)
- Instead of trusting a question-answering model with extracting entities, attributes and values, we do all the work with explicit syntactic and semantic representation. The advantage of this approach is that we obtain candidate substitutions from T2 as a by-product of fact-checking.


5.2. Example of Alignment
- -
- she has potassium depletion;
- -
- spironolactone is a potassium-sparing drug;
- -
- spironolactone will cause her to retain potassium;
- -
- her serum potassium concentration will normalize.
- -
- she has potassium depletion due to Liddle’s syndrome, a channelopathy that affects epithelial sodium channels;
- -
- there is a choice of potassium-sparing drugs;
- -
- spironolactone acts via aldosterone receptors, amiloride and triamterene via sodium channels;
- -
- in Liddle’s syndrome an action via sodium channels is required.
6. Evaluation
6.1. Hallucination Types
- 1)
- Hallucination based in dialogue history. Dialogue history-based hallucinations are produced when an LLM confuses names or relations of entities. For example, if the user mentions that their wife Mary likes shopping, and later informs their niece Lynn is coming to show her recent purchase, the LLM might incorrectly connect Mary and Lynn together as the same person. Furthermore, during a dialogue, the LLM can form incorrect conjectures based on previous mistakes within the dialogue tracking, breaking the context and content of a conversation in a sequence.
- 2)
- Hallucination in abstractive summarization. Although summarization is useful in condensing information, generative summarization systems make errors and deviate between the original and generated data.
- 3)
- Hallucination in generative question answering. It occurs when an LLM makes an erroneous inference from its source information and produces an incorrect answer to a user question. This can happen even when relevant source material is available. For example, if a user asks, “Where a swimmer can be attacked by a sea urchin: Andaman sea or Mediterranean sea?” and context is provided stating that the first choice is true, an LLM may still wrongly respond “Mediterranean sea “ due to its own prior knowledge about Mediterranean sea being a top destination for beach vacations. Rather than accurately recalling the pre-existing source information, LLM can ignore the evidence and perform an unjustified inference based on its existing knowledge.
- 4)
- General data generation hallucination. LLM model generates outputs that may appear plausible or coherent but are not supported by factual or reliable information. It is a type of error where the model fabricates details or makes assumptions that go beyond the input data it has been trained on. This can result in the generation of false or misleading information that may seem convincing to humans but lacks a proper factual basis. Unlike other types of hallucination, the root cause of a general data hallucination is an overextension beyond training data rather than an incorrect inference or lack of grounding. The mode essentially imagines new information that isn’t warranted by its training data.
6.2. Token-Level Hallucination Correction
6.3. Evaluation against Fact Extraction and Verification Annotation Platform
6.4. Automated Evaluation on QA Datasets
6.5. Information-Seeking Dialogues
6.6. Personalized Drug Intake Recommendation Domain
- – % of properly identified hallucinations sentences;
- – % of properly corrected hallucinations sentences.
6.7. Error Analysis
6.8. Examples of Repairs for Hallucination


7. Discussion
8. Conclusions
- Fact checking via web mining adjusted to LLM by means of collaborating with LLM and applying a spectrum of text matching approaches;
- Applicability to texts of various genres and mixture of formats and genres in a document. By relying on a diversity of text matching algorithms acting on phrase and sentence level, we identify hallucination in various parts of documents.
- Application to the whole document with opinionated data, background, captions, results, design notes etc. Truth-O-Meters only applies fact-checking to the phrases and sentences where hallucination can potentially occur.
- Performing proper contextualization to identify and fully substitute facts from what is expressed in text.
- Handling sources with inconsistency by finding least defeated authoritative source and avoiding most defeated.
Acknowledgments
References
- Aronson, J (2009). Medication errors: What they are, how they happen, and how to avoid them. QJM: Monthly Journal of the Association of Physicians. 102. 513-21. [CrossRef]
- Epstein D, Illia Polosukhin, Jacob Devlin, Kenton Lee (2019) Natural questions: a benchmark for question answering research. Transactions of the Association for Computational Linguistics, 7:453–466. [CrossRef]
- Escarza, S. Escarza, S., Larrea, M.L., Castro, S.M., & Martig, S. (2009). DeLP viewer: a defeasible logic programming visualization tool.
- Estes A, Nikhita Vedula, Marcus Collins, Matt Cecil, and Oleg Rokhlenko. 2022. Fact Checking Machine Generated Text with Dependency Trees. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing: Industry Track, pages 458–466, Abu Dhabi, UAE. Association for Computational Linguistics.
- Galitsky B (2013) Transfer learning of syntactic structures for building taxonomies for search engines. Eng Appl Artif Intell 26(10):2504–2515. [CrossRef]
- Galitsky B (2022) Improving open domain content generation by text mining and alignment. In Artificial Intelligence for Healthcare Applications and Management. Elsevier.
- Galitsky B, Ilvovsky D, Goldberg S. Shaped-Charge Learning Architecture for the Human–Machine Teams. Entropy. 2023; 25(6):924. [CrossRef]
- Employing abstract meaning representation to lay the last-mile toward reading comprehension. In Artificial Intelligence for Customer Relationship Management; Springer: Berlin/Heidelberg, Germany, 2020; pp. 57–86.
- Galitsky, B. Merging deductive and inductive reasoning for processing textual descriptions of inter-human conflicts. J Intell Inf Syst 27, 21–48 (2006). [CrossRef]
- Gao L, Zhuyun Dai, Panupong Pasupat, Anthony Chen, Arun Tejasvi Chaganty, Yicheng Fan, Vincent Y Zhao, Ni Lao, Hongrae Lee, Da-Cheng Juan (2023). Attributed text generation via post-hoc research and revision. arXiv preprint arXiv:2210.08726. [CrossRef]
- Garcia A, Simari G (2004) Defeasible logic programming: an argumentative approach. Theory Pract Logic Program 4:95–138.
- Goldberg, S.; Pinsky, E.; Galitsky, B. A bi-directional adversarial explainability for decision support. Hum. Intell. Syst. Integr. 2021, 3, 1–14. [CrossRef]
- Gururangan S, Ana Marasovic, Swabha Swayamdipta, Kyle Lo, Iz Beltagy, Doug Downey, and Noah A Smith. Don’t stop pretraining: adapt language models to domains and tasks. In ACL, 2020. [CrossRef]
- He H, Hongming Zhang, and Dan Roth. 2022. Rethinking with retrieval: Faithful large language model inference. arXiv preprint arXiv:2301.00303. [CrossRef]
- Hecham A (2016) DEFT, an open source Java tool for Defeasible Datalog. https://github.com/raoufhec/DEFT/.
- Thorne J, Andreas Vlachos, Christos Christodoulopoulos, and Arpit Mittal. (2018) Fever: a large-scale dataset for fact extraction and verification. In Proceedings of NAACL-HLT, pages 809–819. [CrossRef]
- Zhou Jie, Xu Han, Cheng Yang, Zhiyuan Liu, Lifeng Wang, Changcheng Li, Maosong Sun (2019) GEAR: Graph-based Evidence Aggregating and Reasoning for Fact Verification.
- Joshi M, Eunsol Choi, Daniel S Weld, and Luke Zettlemoyer. (2017) TriviaQA: A large scale distantly supervised challenge dataset for reading comprehension. ACL, 1601–1611. [CrossRef]
- Kim S, Spandana Gella, Di Jin, Alexandros Papangelis, Behnam Hedayatnia, Yang Liu, and Dilek Hakkani-Tur. (2023) DSTC11 track proposal: Task-oriented conversational modeling with subjective knowledge. https://github.com/alexa/dstc11-track5.
- Kuehn KM and Leon A Salter. 2020. Assessing digital threats to democracy, and workable solutions: a review of the recent literature. International Journal of Communication 14 (2020), 22.
- Kwiatkowski T, Jennimaria Palomaki, Olivia Redfield, Michael Collins, Ankur Parikh, Chris Alberti,.
- Lazaridou A, Elena Gribovskaya, Wojciech Stokowiec, and Nikolai Grigorev. 2022. Internet-augmented language models through few-shot prompting for open-domain question answering. arXiv preprint arXiv:2203.05115. [CrossRef]
- Le, and Denny Zhou. (2022) Chain of thought prompting elicits reasoning in large language models. Advances in Neural Information Processing Systems.
- Lee, Nayeon, Wei Ping, Peng Xu, Mostofa Patwary, Mohammad Shoeybi and Bryan Catanzaro. “Factuality Enhanced Language Models for Open-Ended Text Generation.” ArXiv abs/2206.04624 (2022).
- Levesque H.J., Lin F., Reiter R. (1997) Defining Complex Actions in the Situation Calculus. Technical Report, Department of Computer Science, University of Toronto.
- Li XL, Adhiguna Kuncoro, Jordan Hoffmann, Cyprien de Masson d’Autume, Phil Blunsom, and Aida Nematzadeh. 2022. A Systematic Investigation of Commonsense Knowledge in Large Language Models. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, 11838–11855.
- Lian R, Min Xie, Fan Wang, Jinhua Peng and Hua Wu. 2019. Learning to select knowledge for response generation in dialog systems. In International Joint Conference on Artificial Intelligence. [CrossRef]
- Lin C-Y. 2004. ROUGE: A package for automatic evaluation of summaries. In ACL workshop, pages 74–81.
- Liu T, Yizhe Zhang, Chris Brockett, Yi Mao, Zhifang Sui, Weizhu Chen, and Bill Dolan. (2022) A Token-level Reference-free Hallucination Detection Benchmark for Free-form Text Generation. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 6723–6737, Dublin, Ireland. Association for Computational Linguistics.
- Maynez J, Shashi Narayan, Bernd Bohnet, and Ryan McDonald. 2020. On faithfulness and factu- ality in abstractive summarization. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 1906–1919, On- line. Association for Computational Linguistics.
- Micallef N, Vivienne Armacost, Nasir Memon, and Sameer Patil. 2022. True or False: Studying the Work Practices of Professional Fact-Checkers. Proceedings of the ACM on Human-Computer Interaction 6, CSCW1 (2022), 1–44. [CrossRef]
- Nakano R, Jacob Hilton, Suchir Balaji, Jeff Wu, Long Ouyang, Christina Kim, Christopher Hesse, Shantanu Jain, Vineet Kosaraju, William Saunders 2021. WebGPT: Browser-assisted question answering with human feedback. arXiv preprint arXiv:2112.09332. [CrossRef]
- Nakov P, David Corney, Maram Hasanain, Firoj Alam, Tamer Elsayed, Alberto Barron- ’ Cedeno, Paolo Papotti, Shaden Shaar, and Giovanni ˜ Da San Martino. 2021. Automated fact-checking for assisting human fact-checkers. In Proceedings of the Thirtieth International Joint Conference on Artificial Intelligence, IJCAI-21, pages 4551–4558. International Joint Conferences on Artificial Intelligence Organization. Survey Track.
- Ouyang L, Jeff Wu, Xu Jiang, Diogo Almeida, Carroll L. Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke E. Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul Francis Christiano, Jan Leike, and Ryan J. Lowe. 2022. Training language models to follow instructions with human feedback. ArXiv, abs/2203.02155.
- Papineni K, Salim Roukos, ToddWard, and Wei- Jing Zhu. 2002. BLEU: a method for automatic evaluation of machine translation. In ACL, pages 311–318.
- Peng, Baolin & Galley, Michel & He, Pengcheng & Cheng, Hao & Xie, Yujia & Hu, Yu & Huang, Qiuyuan & Liden, Lars & Yu, Zhou & Chen, Weizhu & Gao, Jianfeng. (2023). Check Your Facts and Try Again: Improving Large Language Models with External Knowledge and Automated Feedback. 10.48550/arXiv.2302.12813. [CrossRef]
- Pennington J, Richard Socher, and Christopher D Manning. 2014. GLOVE: Global vectors for word representation. In Proceedings of the EMNLP, 1532–1543.
- Radford A, Karthik Narasimhan, Tim Salimans, Ilya Sutskever (2018) Improving Language Understanding by Generative Pre-Training. Preprint 2018.
- Rebuffel C, Marco Roberti, Laure Soulier, Geof- frey Scoutheeten, Rossella Cancelliere, and Patrick Gallinari. 2021. Controlling hallucinations at word level in data-to-text generation. arXiv preprint arXiv:2102.02810. [CrossRef]
- Schick, T., Dwivedi-Yu, J., Dessì, R., Raileanu, R., Lomeli, M., Zettlemoyer, L., Cancedda, N., & Scialom, T. (2023). Toolformer: Language Models Can Teach Themselves to Use Tools. ArXiv, abs/2302.04761. [CrossRef]
- Shuster K, Spencer Poff, Moya Chen, Douwe Kiela and Jason Weston. 2021. Retrieval augmentation reduces hallucination in conversation. arXiv preprint arXiv:2104.07567. [CrossRef]
- Wang C and Rico Sennrich. 2020. On exposure bias, hallucination and domain shift in neural ma- chine translation. In Proceedings of the 58th Annual Meeting of the Association for Computational Lin- guistics, pages 3544–3552, Online. Association for Computational Linguistics.
- Wei J, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, Fei Xia, Ed H. Chi, Quoc V.
- Williams A, Nikita Nangia, and Samuel R. Bowman. 2018. A broad-coverage challenge corpus for sentence understanding through inference. In NAACL-HLT, pages 1112–1122. [CrossRef]
- Williams, A. (2023) Fact or Fiction: What Are the Different LLM Hallucination Types? https://www.holisticai.com/blog/types-of-llm-hallucinations.
- Yin, W., Jamaal Hay, Dan Roth (2020) Benchmarking Zero-shot Text Classification: Datasets, Evaluation and Entailment Approach. arXiv:1909.00161. [CrossRef]
- Yuan W, Graham Neubig, and Pengfei Liu. 2021. BARTScore: Evaluating generated text as text generation.
- Zhou J, Xu Han, Cheng Yang, Zhiyuan Liu, Lifeng Wang, Changcheng Li, and Maosong Sun. (2019). GEAR: Graph-based Evidence Aggregating and Reasoning for Fact Verification. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 892–901, Florence, Italy. Association for Computational Linguistics.
- Zhou J, Yixuan Zhang, Qianni Luo, Andrea G Parker, and Munmun De Choudhury. 2023. Synthetic Lies: Understanding AI-Generated Misinformation and Evaluating Algorithmic and Human Solutions. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems (CHI ‘23). Association for Computing Machinery, New York, NY, USA, Article 436, 1–20. [CrossRef]
- Zhou, Chunting & Gu, Jiatao & Diab, Mona & Guzman, Paco & Zettlemoyer, Luke & Ghazvininejad, Marjan. (2020). Detecting Hallucinated Content in Conditional Neural Sequence Generation. [CrossRef]






| Hallucination Type | Example | % of Identified Cases |
|---|---|---|
| Domain-specific Knowledge | Born in Saints Peterburg (Moscow), Suvorov studied military history as a young boy and joined the Imperial Russian Army. | 89 |
| Commonsense knowledge | The sports award is received at the start (end) of the competition | 4 |
| Incoherence or improper collocation | They describe the civil disobedience by many (people) in their town | 4 |
| Unrelated to the central topic | Composer Tchaikovsky’s work was first publicly performed in 1865. In 1868, his First Symphony was well-received as he established himself as an artist (Piano Concerto No.1) | 2 |
| Conflict with preceding context | Tchaikovsky was born on May 7, 1840, in Kamsko-Votkinsk, Vyatka, Russia. He was the second eldest of his parents’ six surviving offspring. As a youngest () member of the family… | 1 |
| P | R | F1 | |
|---|---|---|---|
| As detected by the authors (Liu et al. 2021) | 0.68 | 0.81 | 0.74 |
| Whole web matching | 0.60 | 0.74 | 0.66 |
| Wikipadia matching | 0.65 | 0.72 | 0.68 |
| Matching with the totality of resources | 0.66 | 0.74 | 0.70 |
| FEVER Score | P | R | F1 | |
|---|---|---|---|---|
| GEAR (Zhou et al. 2019) | 67.1 | 70.6 | 81.6 | 75.7 |
| Truth-O-Meter | 69.1 | 75.1 | 78.8 | 76.9 |
| Truth-O-Meter-i | 66.8 | 76.0 | 80.6 | 78.2 |
| Method | SQuAD Extension | Natural Question | TriviaQA | HotpotQA | DSTC11 Track 5 |
|---|---|---|---|---|---|
| ChatGPT | 79.3 | 36.6 | |||
| Chain of thoughts (Wei et al. 2022) | 64.3 | 79.2 | 42.8 | ||
| CRITIC* (Gou et al. 2023) | 79.9 | 86.6 | 56.9 | ||
| LLM-Augmenter CORE (Peng et al. 2022) | 50.83 | ||||
| RARR (Gao et al. 2023) | 41.5 | ||||
| Truth-O-Meter | 81.0 | 78.2 | 84.9 | 51.3 | 55.3 |
| Truth-O-Meter-i | 79.3 | 80.3 | 86.1 | 50.1 | 53.0 |
| Method | KF1 | BLEU | ROUGE | BARTScore |
|---|---|---|---|---|
| ChatGPT | 26.7 | 1.0 | 16.8 | 0.25 |
| LLM-AUGMENTER (Peng et al. 2023) | 36.41 | 7.6 | 22.8 | 0.35 |
| chatGPT + Truth-O-Meter | 33.5 | 8.0 | 21.8 | 0.30 |
| chatGPT + Truth-O-Meter-i | 32.1 | 8.5 | 23.3 | 0.36 |
| chatGPT + Truth-O-Meter + Truth-O-Meter-i online selector | 33.9 | 7.5 | 23.6 | 0.32 |
| The Class of Common Symptoms | #/% of Properly Identified Hallucinations Sentences | #/% Of Sentences Accepted As They Are | Total # of True Sentences Used | ||
|---|---|---|---|---|---|
| Health | 54 | 88.4 | 1026 | 94.1 | 17.1 |
| Geography | 48 | 86.2 | 1251 | 96.3 | 16.3 |
| Legal | 74 | 83.4 | 980 | 92.8 | 19.0 |
| Engineering | 57 | 88.1 | 1340 | 94.7 | 18.7 |
| Average | 58.2 | 86.52 | 1149 | 94.47 | 17.77 |
| History | 28 | 88.6 | 1420 | 97.9 | 20.4 |
| Marketing | 45 | 85.2 | 1237 | 98.6 | 19.3 |
| Art | 17 | 89.6 | 1478 | 96.3 | 17.5 |
| Politics | 52 | 86.5 | 1513 | 98.1 | 19.0 |
| Average | 35.5 | 87.47 | 1412 | 97.72 | 19.05 |
| The Class of Common Symptoms | Average # of Corrected Sentences for a Recommendation | % of Properly Identified Hallucinations Sentences | % of Properly Corrected Hallucinations Sentences |
|---|---|---|---|
| Bloating | 0.8 | 86.3 | 77.1 |
| Cough | 0.7 | 82.6 | 79.3 |
| Diarrhea | 1.0 | 85.8 | 76.3 |
| Dizziness | 0.7 | 88.4 | 80.1 |
| Fatigue | 0.9 | 84.6 | 79.2 |
| Fever | 0.7 | 85.0 | 78.8 |
| Headache | 0.8 | 85.6 | 76.5 |
| Muscle Cramp | 0.8 | 83.5 | 77.0 |
| Nausea | 0.7 | 86.4 | 75.9 |
| Throat irritation | 0.6 | 83.4 | 79.3 |
| Average | 0.77 | 85.16 | 77.95 |
| The Class of Common Symptoms | Rate of Wrong Facts per Text for LLM Content | Rate Of Wrong Facts Per Text For Corrected Content | Rate of Wrong Facts per Sentence for LLM Content | Rate of Wrong Facts per Sentence for Corrected Content |
|---|---|---|---|---|
| Bloating | 29.6 | 5.0 | 3.5 | 0.31 |
| Cough | 23.0 | 6.4 | 3.1 | 0.34 |
| Diarrhea | 26.1 | 6.7 | 2.9 | 0.28 |
| Dizziness | 24.5 | 6.6 | 3.0 | 0.29 |
| Fatigue | 24.0 | 5.1 | 3.4 | 0.27 |
| Fever | 25.7 | 4.8 | 2.7 | 0.29 |
| Headache | 26.6 | 5.1 | 3.2 | 0.30 |
| Muscle Cramp | 24.1 | 5.6 | 3.5 | 0.26 |
| Nausea | 27.0 | 6.1 | 3.3 | 0.27 |
| Throat irritation | 25.9 | 5.9 | 3.0 | 0.31 |
| Average | 25.65 | 5.73 | 3.16 | 0.292 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2023 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).