Preprint
Review

This version is not peer-reviewed.

Hallucination in Large Language Models: Epistemic Status, Detection, and Validation Criteria in Higher Education

Submitted:

22 September 2026

Posted:

23 September 2026

You are already at the latest version

Abstract
False statements produced by a language model arrive with the same confident form as true ones, with no mark to guide verification, and this indistinction bears fully on the university classroom. This article reviews the peer-reviewed literature under a declared protocol that retained thirty-two works and one continuously updated resource, drawn from technical research on detection, epistemology, cognitive psychology, critical thinking, and educational research. The argument starts from an epistemic and functional characterization of the system, to which we attribute the lack of access of its own to the output and of a criterion of truth of its own, whence it follows that every correction comes from a reference external to the generative process, which we call external anchoring. On that basis, the classes of divergence are ordered according to whether they are measured against facts about the world or against the material supplied by the user, the detection procedures are reviewed together with their declared limits, and four criteria applicable from a public interface are formulated, grouped into two families and crossed with the typology in a table. The closing sections examine two effects on the reader, cognitive satisfaction and the externalization of epistemic vigilance, and the institutional conditions that their correction demands.
Keywords: 
;  ;  ;  ;  ;  ;  ;  

1. Introduction

Consulting conversational systems has rapidly become part of the ordinary academic work of students and teachers, and this incorporation is running ahead of the training needed to evaluate what such systems deliver, a situation that UNESCO’s guidance on generative artificial intelligence in education and research acknowledged when it warned that the errors of these tools tend to go unnoticed unless the user has solid knowledge of the subject [1]. The warning bears fully on those who consult in order to learn, since in that position the background beliefs of the domain are still being formed. What distinguishes these errors from those carried by any documentary source is their presentation, given that they arrive with the same assertive surface as true statements and without any signal that directs attention toward what is doubtful. Hence the decision to verify rests entirely with the reader, without the form of the text ever prompting it, and the problem reaches squarely into university education, where the evaluation of sources is part of what is taught.
The phenomenon received the name hallucination through an analogy that specialized discussion has been questioning for several years; Smith et al. [2] objected that calling inaccurate output by this name implies accepting that the system perceives, and proposed confabulation for its closeness to the clinical picture in which someone produces a false account without intent to deceive, while Hicks et al. [3] argued that the activity of these models is more accurately described as bullshit in Frankfurt’s sense, given their indifference to the truth of what they produce. Šekrst [4], for her part, warned that much of the vocabulary used to describe the phenomenon carries unnecessary presuppositions. We retain the term hallucination because of its consolidated circulation in the technical literature, where even the works that adopt confabulation reserve it for a subset of cases and place it within the established vocabulary [5], and we use it in a limited way, as the label for a class of outputs and without attributing to the system any perception, intention, or mental state.
The operational definition with which we work understands by hallucination an output that diverges from facts about the world, or from the governing instruction and the material supplied by the user, presented with high apparent confidence, a property we take in a formal sense, verifiable in the text itself through the absence of markers of doubt or hedging and through a predominantly assertive syntax, so that the statement arrives with no signal of its own uncertainty. The two axes announced in that definition, factuality and faithfulness, organize the entire article and stand in an asymmetric relation to each other, given that faithfulness presupposes a governing instruction or a supplied document and is therefore the derived case, while factuality operates even when the user poses an open query and receives an answer with no declared source. The object is further restricted to the textual outputs of language models, so that the divergences specific to multimodal generation remain outside the scope of this review.
The discussion about the mind, consciousness, and understanding of these systems remains open and admits diverse approaches without a conclusive answer [6], and its detailed examination is not necessary for the argument we advance, which is why it lies outside the scope of this work. The work rests instead on a thesis of an epistemic and functional order, according to which the system has neither access of its own to the output it produces nor a criterion of truth of its own with which to evaluate it, deficits that we assert at the level of functioning and justification, so that the thesis holds whatever position one takes on that discussion. From this formulation it follows that every procedure capable of flagging an error imports an external anchor, whereby correction is relocated without ever being internalized, and that the evaluation of the output falls entirely on whoever receives it.
The available literature on the phenomenon developed along two fronts that rarely meet; on the technical side, systematic reviews ordered the classes of divergence and gathered the procedures by which each is detected, together with their performance and their declared limits [7,8,9], while on the educational side empirical research examined the effects of using these tools on the performance and critical thinking of those who use them, with findings ranging from an association between frequent use and lower critical-thinking performance [10] to the persistence of incorrect information in texts written with the system’s assistance [11]. What remains little explored is the stretch that joins both fronts, namely the derivation of criteria executable by the user from what technical research documents and from what the epistemic status of the output warrants expecting, and it is in that stretch that the present work is situated.
Against this backdrop, technical research developed a broad repertoire of detection procedures, most of which presuppose conditions of model access and computation beyond the reach of someone querying from a public interface. The distance between that repertoire and the one available to students and teachers delimits the problem this article addresses, whose objective is to review what the peer-reviewed literature documents about the classes of hallucination and the procedures by which each is detected, in order to formulate on that basis validation criteria executable with the means of a public interface, together with the formative consequences and institutional conditions that their exercise requires. To this end the work takes the form of a review, with the protocol declared in the section following this introduction, and it addresses students and teachers alike, who occupy the same user position before the generated output.
What the work adds to this discussion falls on two levels, and on the conceptual level lies the thesis of external anchoring, by which we hold that the correction introduced by any detection procedure comes from outside the generative process, together with two categories that describe the effect of the output on whoever receives it, cognitive satisfaction, the name for the closure of inquiry in the face of an answer of finished form, and the externalization of epistemic vigilance, the name for transferring to the system the functions of cross-checking what is asserted and correcting it, both marked as our own and distinguished in the body from the neighboring notions with which they might be confused. To the operational level belong the ordering of the criteria into two families and their crossing with the classes of hallucination in Table 1, which offers an instrument for direct classroom use, together with the epistemic function that query design, which we call the prompt as a new maieutics, receives in these pages relative to the one it already had in the first author’s earlier work.
The path begins with the declaration of the review protocol by which the corpus was constituted and continues with the examination of the epistemic status of the generated output, where the thesis just stated is argued with the aid of contemporary epistemology. The next section sets out the typology of the phenomenon together with the documented detection procedures, each with its declared limit, and treats the available figures as situated magnitudes; on that material the validation criteria addressed to the user are then formulated, along with the table that crosses them with the classes of hallucination, while the last section of the body examines the formative consequences and institutional conditions that accompany those criteria, before the conclusions. The article thus crosses technical evidence on detection with conceptual and formative work, so that what computational research documents may find translation into teaching and assessment decisions in higher education.

2. Materials and Methods

The article was built as a narrative review with a declared protocol, without primary data collection, so that what is offered here derives from the triangulation of documentary sources and of previously published secondary data. The unit of analysis was the published work, considered on two levels, that of the literature documenting and measuring the false outputs of language models and that of the epistemic, logical, and formative frameworks through which those outputs can be interpreted. The gathering of the corpus followed thematic cores defined by the function each work fulfills within the argument, a criterion that distributes a single discipline across several cores and requires counting each work only once, in the core where it supports the main claim for which it was retained. The cores so defined cover the technical definition of hallucination with its detection procedures, epistemology, cognitive psychology and metacognition, critical thinking with its logical basis, and the use of generative systems in higher education. The retained corpus comprises thirty-two published works and one continuously updated resource, with the greatest weight in the epistemological core, with nine works, and in the technical core, with eight, followed by cognitive psychology and educational research, with six each, and by critical thinking, with three.
The search was conducted on Scopus, Web of Science, ACM Digital Library, IEEE Xplore, SpringerLink, and PhilPapers, without language restriction, and entry into the corpus was conditional on editorial verification and on the availability of a verifiable identifier, checked in each case against the publication of record; preprints were excluded. On that basis a few bounded admissions operated, which we declare. Peer-reviewed conference proceedings were admitted when no journal publication documented the claim at the same level of detail, a situation that arose with the sources that established the separation between faithfulness to the source and factuality and with those that described verification procedures without external resources. Opinion and commentary pieces published in peer-reviewed journals were retained for their terminological or conceptual contribution, without drawing any empirical data from them. Normative documents disseminated through official channels of international organizations were admitted with that status, and dynamic resources are cited with their access date and with the figure current on that date.
The currency of the literature weighs unevenly across cores, which is why we set a window from 2021 to 2026 for works on hallucination and detection procedures, where the succession of architectures and measurement protocols quickly ages any result, while entry into the cores of epistemology, logic, and cognitive psychology was governed by theoretical relevance, given that their contributions retain value regardless of year of publication. Hence the corpus retains works such as that of Searle [12] on the distinction between syntax and semantics or that of Sperber et al. [13] on epistemic vigilance. The window further admits a bounded exception within the technical core itself, that of the work of Maynez et al. [14], which separated faithfulness to the source from factuality with respect to the world without any later publication documenting that finding in the same detail. Figures taken from updatable resources are recorded with the date on which they were consulted, together with the version of the evaluation method that produced them, so that a reader can situate the figure and repeat the consultation.
The treatment of the retained works consisted of a comparison among the available taxonomies, aimed at reducing the variety of labels to the axes along which the phenomenon is actually measured, and of the extraction, for each documented detection procedure, of the performance it reports together with the limit declared by its own authors. The contributions of the epistemological and formative cores were incorporated through controlled transposition, that is, by taking from each framework the function it fulfills in its literature of origin and declaring the point at which its scope is extended to the problem examined in these pages. Every empirical datum supporting a claim in the article comes from peer-reviewed literature.
Limits that must be declared follow from the design, beginning with the absence of classroom observation and of any detection experiment of our own, so that the work neither estimates the prevalence of false outputs in a given institutional context nor measures the ability of those who read them to recognize them. The figures taken from the literature function as situated magnitudes, dependent on the architecture evaluated, on the prompt with which it is queried, and on the measurement method employed, which is why they are neither aggregated with one another nor projected onto populations other than those in which they were obtained. The scope is further restricted to the textual hallucinations of language models, and multimodal outputs remain outside the corpus and the claims of the article.

3. Epistemic Status of the Generated Output

Asking about the hallucinations of a language model first requires specifying what kind of object the model delivers. Public discussion tends to shift quickly toward the question of the machine’s mind, and that shift carries a high cost, because it ties the evaluation of any output to a controversy that remains open and shows no sign of closing soon. This work takes a different path and advances what we call here an epistemic-functional thesis, according to which the system has neither access of its own to the output it produces nor a criterion of truth of its own with which to evaluate it. Both deficits are asserted at the level of functioning and justification, so that the thesis holds whatever position is held on mind, consciousness, or understanding.
The discussion on intelligence and consciousness in artificial systems today admits multiple approaches and no absolute answer [6], and its detailed treatment would exceed the object of this review and, more importantly, would be unnecessary for the argument that follows. Indeed, whoever attributes understanding to a language model and whoever denies it can agree on something more modest and verifiable, namely that the system does not consult a record of truth of its own before issuing a statement and that neither can it indicate, from within, which of its statements is false; this is the common ground on which we build the rest of the section.

3.1. Agency Without Intelligence

The starting point was offered by Floridi [15], who proposed decoupling agency from intelligence and describing generative models as a form of agency capable of carrying out tasks successfully without that success implying understanding of what is carried out, a formulation the author later extended by inscribing artificial agency within a thesis of multiple realizability [16]. What interests us in this decoupling is its epistemic yield, because it allows us to name precisely the criterion of success that governs the production of the output, a criterion of a statistical and linguistic order, the probable continuation of a sequence given a context; the correspondence between what is stated and the state of affairs to which the statement refers lies entirely outside that criterion. In this sense, a text can be successful by the first measure and false by the second, without anything in the process registering the divergence.
The first half of the thesis deserves careful development, since to say that the system lacks access of its own to its output is not equivalent to saying that it lacks access to its own internal states in a technical sense, because there are procedures that inspect activations and probability distributions. What is missing is the relation in which a statement presents itself to the one who issues it as something that this same issuer holds and can answer for. The output is produced without there being a place from which it is assumed, and this lack explains why the same system can state with identical disposition a proposition and its negation when the context of generation changes.
To this consideration we can add the Chinese room argument of Searle [12], taken up here in a restricted use; from his argument we keep the distinction between the syntactic and the semantic level, without adjudicating anything about the understanding or the mental states of the system. What is relevant to our object is that an output can fully satisfy the rules of syntactic formation, be semantically coherent as well, with terms that combine compatibly and internal references that hold throughout the text, and at the same time be, at the epistemic level, partially or totally false. Formal correctness and internal coherence leave the question of truth value untouched, and in that gap lies the phenomenon we call hallucination here.
Šekrst [4] worked on this same intersection of hallucination, epistemology, and cognition, and warned that much of the vocabulary used to describe the phenomenon carries unnecessary presuppositions, an observation we share and that motivates the limited use of the term in this article. Her proposal consists in applying epistemological frameworks such as reliabilism to assess the reliability of outputs, on the assumption that artificial systems are comparable to human cognition in their exposure to errors of judgment. Here we distance ourselves from that assumption, because the comparison with human error presupposes a subject who makes mistakes and who, under ordinary conditions, has resources for noticing those mistakes; the absence of this second condition is what gives the problem its own status.

3.2. The Subject of Competence

Virtue epistemology offers a finer vocabulary for naming that absence; in this sense, Sosa [17] distinguished animal knowledge, understood as apt belief, that is, belief whose correctness manifests a competence of the subject according to the structure of accuracy, adroitness, and aptness, from reflective knowledge, which further requires that this apt belief be aptly noted by the agent itself. Hila [18], in turn, transferred the distinction to the domain of language models and holds that they achieve a reliabilist, externalist kind of justification without instantiating the internalist standards that produce knowledge, so that they reliably transmit information whose reflective basis was established beforehand by human beings.
Hila’s reading is productive and yet concedes more than necessary, since placing the models at the level of animal knowledge assumes that there is a competence and, with it, a subject whose competence it is, because in Sosa’s scheme aptness consists in the correctness of the result manifesting a skill attributable to whoever bears it. In a language model, statistical regularity belongs to the trained system and its correctness is assessed from outside, with no bearer to whom the skill could be ascribed as its own. It should be added that aptness, in that framework, is predicated of beliefs, and the output of a model lacks the propositional attitude that would make it a belief. For these reasons the analogy with animal knowledge retains expository value and loses force when taken literally.
The foregoing does not require taking sides in the discussion about the mind, because the argument proceeds at the level of attribution and justification; it suffices to observe that the lack of a subject of competence leaves no place for what could house a criterion of truth of its own. In virtue epistemology that criterion rests on the agent’s capacity to note its own belief as apt; where there is no agent to whom the skill can be ascribed, neither is reflective noting possible, and the evaluation of the output shifts entirely to whoever receives it.
A remark should be added on the way the output reaches the user, a text that presents itself with the appearance of testimony, with the structure of a statement that someone holds and for which someone answers, and that at the same time lacks anyone who holds it in that sense. The disposition the reader adopts toward human testimony, with its ordinary mechanisms of attributing responsibility and evaluating the source, is thus activated before an object that does not meet the conditions to receive it.

3.3. Apparent Confidence and a Criterion of Truth of Its Own

The definition of hallucination with which we work in this article incorporates the high apparent confidence of the output, a property we understand in a formal sense, as a feature verifiable in the text itself, given by the absence of markers of doubt or hedging and by a predominantly assertive syntax, so that the statement presents itself without any signal of its own uncertainty. So understood, apparent confidence is established by inspecting the output and does not require attributing to the system any state of subjective certainty. The peer-reviewed literature on verbalized uncertainty supports this characterization by documenting that deployed models are reluctant to mark uncertainty even when they answer incorrectly, and that when prompted to state their confidence they lean toward overconfidence [19,20]; the available frequency figures, for their part, support the persistence and magnitude of the phenomenon, which belongs to a different order of evidence.
Taken together, both properties explain much of the practical problem the user faces, since a false statement accompanied by marks of hesitation would by itself invite verification, whereas the same statement, formulated with the assertive surface that also characterizes the model’s true statements, offers no differential signal. Such is the case with a nonexistent bibliographic reference that arrives with plausible authorship, a relevant journal, volume, pages, and a well-formed digital identifier, elements that make it indistinguishable from a real reference until someone tries to retrieve it. Against this backdrop, the reader is left without the cue that in ordinary communication directs attention toward what is doubtful, and the available experiments show that users rely on generated answers regardless of whether they are marked by expressions of certainty [19], so that verification is no longer triggered by the form of the message and comes to depend on a deliberate decision by whoever carries it out.
A clarification is in order regarding the way this deficit operates; when the user sets a governing instruction, a document to summarize, or a set of data to interpret, the output can be evaluated by its fit with that instruction, and that fit constitutes a derived case, because it depends on there being an instruction involved. Divergence from facts about the world, by contrast, operates even when no governing instruction has been given, and that is why the epistemic problem arises in its purest form where the user poses an open query and receives an answer with no declared source.
It remains to clarify in what sense the output is not correctable from within, since the strong formulation, according to which the system lacks any detection mechanism, is easily refuted by the documented procedures that identify hallucinations with appreciable degrees of success. The formulation we advance here is more restricted and, for that very reason, more resilient, namely that the system lacks a criterion of truth of its own. Every detection procedure, even an automatic one, imports an external anchor, whether a source corpus against which to check the output, the comparison between successive samples from the same model, another model acting as evaluator, or a reference set by human beings; in every case the criterion of correctness comes from outside the generative process, so that these mechanisms relocate correction without internalizing it.
This distinction has consequences that we note now, given that a procedure evaluating the stability of the output across successive samples detects instability, and a falsehood the model produces stably passes through that filter unnoticed. A procedure that checks the output against a source corpus evaluates faithfulness to that corpus, and its scope ends where the corpus ends. What we call here external anchoring names precisely this dependence, each technique inherits the scope and the limits of the reference it imports, and none substitutes for the criterion of truth absent in the system.
A foreseeable objection should be addressed, since technical literature calls certain detection procedures reference-free, because they dispense with an external knowledge base, and one of them even presents itself as a zero-resource procedure [8,21]. That label describes the input the procedure dispenses with and leaves the decisive point untouched, because the criterion of correctness still comes from the contrast between successive samples and from thresholds set by human beings, none of which belongs to the generative act. The same holds for procedures that inspect the internal states of the model, where what is read are signals of uncertainty or inconsistency, and their translation into a verdict about truth requires a label established from outside.
The epistemic status of the generated output is thus delimited as the product of agency without intelligence, formally correct and epistemically indifferent to the truth value of what it states, presented without any signal of its own uncertainty and not correctable from within for lack of a criterion of truth of its own. From these features it follows that validation falls on the reader, and that its exercise requires explicit criteria, because the form of the output does not provide them.

4. Typology and Documented Detection

With the epistemic status of the output delimited, it is time to name in its technical terminology what we described at the functional level in the previous section, since the specialized literature has a precise vocabulary for distinguishing the classes of divergence and the procedures by which each is detected. That vocabulary, however, does not form a unified system and carries the history of the problems it was solving, so that its careless transfer to the discussion of language models produces nestings that none of the sources authorizes. We should therefore reconstruct the genealogy of the terms before fixing them, so as to establish what each category measures and against what reference it measures it. What follows adopts one pair of axes as the governing one and places the other schemes in their corresponding position.

4.1. Factuality and Faithfulness

The first consolidated scheme came from natural language generation, where Ji et al. [7] took up a distinction already established in research on automatic summarization, translation, and data-to-text generation, and separated intrinsic hallucination, which contradicts the source content, from extrinsic hallucination, which the source allows neither to verify nor to refute. It is worth noting that both categories are defined entirely with respect to a given source, given that the framework presupposes tasks in which the system receives an input document and produces a transformation of it. In this sense, the scheme is rigorous within its reference class and becomes inapplicable where there is no input document, that is, in the open query that constitutes the most widespread use of conversational systems.
An earlier precedent had already posed the problem that this scheme could not solve; indeed, Maynez et al. [14] evaluated with human annotators the summaries produced by several neural systems and separated faithfulness to the source document from factuality with respect to the world, upon finding that part of the content unsupported by the document was nonetheless true in light of general knowledge. It follows from this finding that the two properties can come apart, since a text faithful to a mistaken source will be false, while a text unfaithful to a correct source may be right by another route. This dissociation makes it productive to treat them as intersecting axes and advises against merging them into a single scale of correctness.
Huang et al. [8] took up this dissociation and proposed a redefined taxonomy for language models, organized into two primary types, of which factuality hallucination marks the discrepancy between what is generated and verifiable real-world facts, and is subdivided according to whether the model contradicts an established fact or fabricates content that admits no check against any fact; faithfulness hallucination marks the divergence from the user’s input, whether the instruction or the supplied context, and further incorporates the internal inconsistency of the generated text itself. The proposal presents itself explicitly as a redefinition of the earlier vocabulary and not as a layer added on top of it, a point worth retaining, because the way both schemes coexist in a single text depends on it.
From this genealogy follows a decision that governs the rest of the article, to adopt factuality and faithfulness as governing axes and to keep the distinction between intrinsic and extrinsic hallucination in its proper place, within the faithfulness axis and not parallel to it. Treating both pairs as two levels of a single hierarchy, with factuality and faithfulness subordinated to the intrinsic and the extrinsic, produces an inconsistency that is hard to overcome, because it would require defining the discrepancy with facts about the world within a scheme that only knows how to measure against an input document. Zhang et al. [9] further offered a third ordering, with categories that distinguish conflict with the input, conflict with the previous context, conflict with world knowledge, and its correspondence with the axes of Huang et al. is close enough not to multiply the schemes in the exposition.
The two axes do not divide the epistemic problem of interest here equally, given that faithfulness presupposes a governing instruction or a supplied document, while factuality operates even when the user poses an open query and receives an answer with no declared source. The asymmetry noted in the previous section thus receives its technical name, the derived case corresponds to the faithfulness axis and the general case to the factuality axis. Added to this is a practical corollary of considerable scope, since faithfulness admits verification against an object the user has at hand, whereas factuality requires leaving the exchange for a reference that users must obtain on their own.

4.2. Detection Procedures

Technical literature has produced a broad repertoire of detection procedures, whose reading in light of the previous section calls for constant caution, since each procedure measures against a given reference and its scope ends where that reference ends. Ordering them by the axis they actually evaluate, rather than by their computational sophistication, makes it possible to see what is covered and what remains beyond the reach of each family, and avoids the impression that high detection performance amounts to a guarantee about the truth value of the output.
On the faithfulness side operate the procedures that check the output against the supplied document, with textual inference metrics that estimate whether each claim in the generated text follows from the source content; Ji et al. [7] documented the origin of this family in summarization research and Huang et al. [8] gathered its later development in the domain of language models. Their performance is appreciable and their limit is declared from the design, since such a procedure does not rule on the truth of the claim, only on its support in the document, so that a summary faithful to an erroneous source receives the highest score without anything in the procedure registering the problem.
On the factuality side operate the procedures that check the output against an external knowledge base or against retrieval results, and there the limit shifts toward the coverage and currency of that base, in addition to the difficulty of converting a natural-language claim into a verifiable query. When the claim involves a poorly documented fact, an attribution of authorship, or a bibliographic reference, the procedure frequently returns absence of evidence, a result that is not equivalent to established falsehood and that shifts the decision back to the person querying.
A third family dispenses with documentary reference and examines the stability of the output across successive samples from the same model, a procedure that Manakul et al. [21] developed under the label zero-resource and whose anchoring was situated in the previous section. What we must specify here is what it measures, namely instability in generation, a magnitude the literature takes as an indicator of hallucination by empirical correlation and not by definition. Zhang et al. [9] placed alongside it the uncertainty estimation procedures, those that read the probabilities assigned to generated tokens and those that ask the model about its own confidence expressed in natural language; in all these cases the signal obtained is one of stability or confidence, and its translation into a verdict about truth requires a threshold established from outside.
Added to this landscape is the use of one model as evaluator of another, a procedure that Huang et al. [8] recorded for its low cost and broad coverage, and whose limit is that the evaluator inherits the deficiencies of the evaluated. A restriction that the technical literature rarely thematizes concerns the conditions of application of these procedures, since most of them presuppose programmatic access to the model, the ability to generate multiple samples, or the availability of a comparison corpus, conditions that do not accompany students or teachers querying a conversational system from a public interface. The repertoire available to research and the repertoire available to the ordinary user do not coincide, and we hold here that this distance, which follows from what has just been described and not from a conjecture about user behavior, defines much of the formative problem.

4.3. Figures and Their Scope

Frequency figures function in this work as situated magnitudes, dependent on the architecture evaluated, the prompt used, and the measurement method, and reading them requires declaring in each case what was measured. The resource most cited in public discussion is the Vectara Hallucination Leaderboard, a continuously updated repository that evaluates the factual consistency of summaries produced over a proprietary corpus of more than seven thousand seven hundred articles, with a prompt instructing the model to summarize using only the information in the passage provided. Accessed on September 20, 2026, with data updated as of May 11, 2026, and evaluation by means of the HHEM-2.3 model, the resource reports hallucination rates ranging from 1.8% for the best-placed model to 24.2% for the worst-placed [22].
Two clarifications accompany this figure inseparably, and the first concerns the axis, since the measurement compares the summary with the source document and belongs entirely to the faithfulness axis, so that it warrants no claim about how often a model diverges from facts about the world in an open query; the resource’s own maintainers state that they do not evaluate general factual accuracy and explain that decision by the impossibility of determining what data each model saw during training. The second concerns the procedure, since the metric rewards extractiveness, and a system that copied the source document would obtain a zero rate without this saying anything about the quality of the summary, a limit the resource itself explicitly acknowledges.
Let us add an observation that the resource’s own listing allows one to verify by simple inspection, namely that the ranking does not track the scale or recency of the models, since small and older systems appear in the top positions while several of the most recent and largest models occupy intermediate or low places. This dispersion reinforces the treatment of the figure as a situated magnitude and advises against reading it as an indicator of general capability, all the more so since the resource itself warns against its use as an isolated metric.
A second figure concerns the link between apparent confidence and error, and comes from peer-reviewed literature; Zhou et al. [19] explicitly prompted several deployed models to state their confidence when answering questions, and found among the answers issued with confidence markers an average error rate of 47%. The condition of the procedure is inseparable from the figure, because it corresponds to answers obtained under a prompt that asks for a statement of confidence and does not describe the hallucination frequency of a model in ordinary use; with that condition stated, the figure supports what the argument asks of it, namely that the confidence marker does not discriminate between correct and incorrect answers.
From this section as a whole we retain an ordering and a caveat, the first of which distributes the classes of hallucination along two axes, one measuring against facts about the world and the other against the instruction or the supplied document, and distributes the detection procedures according to the axis each evaluates, with the limit each family declares. The caveat concerns the figures, which report on a given task, corpus, and prompt, and which lose their meaning when extracted from those conditions to state a general property of the systems.

5. Validation Criteria for the User

In the previous section we noted an asymmetry that now organizes the work, since the repertoire of procedures developed by research presupposes conditions of access and computation beyond the reach of someone querying a conversational system from a public interface. The question of validation therefore calls for criteria that can be formulated in this second register, executable with the means students and teachers have at hand, rather than domestic versions of computational techniques. A note on nomenclature is also in order, since the names factuality and faithfulness are reserved for the classes of hallucination established in the preceding section, while the criteria are named after the operation the user performs, so that no criterion is defined by the type it must detect. So ordered, the criteria are grouped into two families, consistency and verification, which this section crosses at the end with the typology in an intersection table.

5.1. Formal and Informal Logic as a Basis

Formulating the criteria first requires establishing where they are formulated from, a question this work resolves by continuity with an already published position, according to which the basis of critical thinking lies in formal and informal logic [6,23,24], with critique understood in its most precise sense, that of analysis, as fixed by the Kantian use of the term, and with conceptual clarity operating as a task prior to any evaluation. Clarification is needed on the scope of the former, since the formal logic we refer to is symbolic or mathematical logic, whose evolution fed the development of computer logic, and not the syllogistic that continues to be taught as the classical route. Informal logic, for its part, contributes to the examination of the fallacies and biases of ordinary language, a front that no computational metric covers and that is accessible through attentive reading, without any instrument.
This basis imposes a limit that we declare at the outset, given that formal logic evaluates the validity of inference and not the truth of premises, so that impeccable reasoning from a false premise can deliver a false conclusion with all the outward marks of correctness. In Section 3 we stated the consequence for the case at hand, namely that a falsehood the model produces stably passes unnoticed through any consistency filter, and semantic coherence, an intralinguistic property by which terms combine without contradiction and internal references hold, is entirely compatible with the falsity of what is asserted. Hence the consistency family functions in what follows as a low-cost preliminary filter, capable of discarding defective outputs with one reading and without leaving the text, while the weight of validation falls on the verification family, the only one that introduces a reference independent of the model that produced the output.
Conceptual clarity precedes both families and therefore deserves separate mention, since the refinement of the terms at stake determines what is checked and against what reference it is checked. A student who asks about the state of a debate without first having fixed the meaning of the term that names it receives an answer whose correctness is undecidable, given that the claim may be true under one sense of the term and false under another, so that the disagreement lies in meaning rather than in fact. Fixing that meaning when formulating the query, with the precision required by the definition of a term within an argumentative fabric, thus constitutes the first validation operation, prior in the order of execution to the consistency review and to any check against sources, since without it none of the later criteria knows what to measure against.

5.2. Consistency Criteria

The consistency family brings together two criteria that operate without leaving the material the model itself delivers, and the first of them, internal coherence, examines a single output. Its minimal form consists in registering contradictions between claims within the same text, a task that the length of the answers makes less trivial than it seems, given that the initial claim and the one contradicting it often appear several paragraphs of development apart. Its demanding form, the one that takes up the basis just established, examines the structure of the inference when the output presents an argument, checking that the conclusion actually follows from the stated premises. This examination incorporates the front of informal logic, attentive to the deceptions of ordinary language that the generated text reproduces naturally, among which we find hasty generalization from an isolated case and the presentation as established consensus of what is one position among several.
The second criterion, cross-sample stability, transfers to the user’s register what the previous section described as a zero-resource procedure, and its execution consists in posing the same query in independent sessions, with the previous conversation closed, and then comparing the answers obtained. Divergences in concrete data, the reference that changes year or journal from one session to the next and the figure that shifts without warning, mark the place where one should stop. We set the limit in Section 3 and repeat it here because it governs the use of the criterion, what the comparison measures is instability in generation, a magnitude the literature takes as an indicator by empirical correlation, and agreement between successive samples confirms nothing and systematic error repeats itself identically without triggering any signal.
Both criteria share a virtue and a deficiency worth retaining before moving to the next family, since they are carried out with the sole resources of the text and the user’s time, without access to external databases or tools, and are therefore immediately applicable in the classroom; their scope, however, ends at the surface of language, where contradiction and instability become visible, while the relation of what is said to the world remains outside their jurisdiction. Both import a criterion foreign to the generative process, in the sense we established in Section 3, and what keeps them within this family is that the material on which they operate remains the one the model itself delivers. An internally coherent text, stable across sessions and formally impeccable, fully satisfies this family and may be false in each of its claims, which is why validation requires further review.

5.3. Verification Criteria

With the second family, a reference the model did not produce comes into play, and its first criterion, checking against an external reference, corresponds to the factuality axis and operates where the user posed an open query without supplying any document. The operation consists in independently retrieving the source that documents the claim, and its clearest case is offered by the bibliographic reference, which can be checked through the digital identifier, the journal’s catalog, or the institutional repository, with a result that admits no degrees, the work either exists with those data or it does not. Outside that case the criterion becomes considerably more costly, requiring the user to locate a relevant source and to compare it with the claim, a task that returns the user to the documentary work the query sought to shorten.
This criterion carries two limits inherited from the automatic procedures operating on the same axis, and the first concerns the negative result, since the impossibility of finding documentary support for a claim is not equivalent to establishing its falsity, so that the user is left with a well-founded suspicion rather than a verdict. The second concerns cost; external verification consumes time that grows with the number of checkable claims in the output, and an extensive answer may contain dozens of them; hence the criterion is applied hierarchically, attending first to the claims that support the conclusion and to those whose falsity would carry consequences, as one proceeds with any documentary source of unknown reliability.
When the user has delivered a document or set a governing instruction, checking against the supplied source, the second criterion of this family is applied to an object already at hand, and the cost of the operation drops considerably. Its execution requires separating in the output what is supported by the document from what the model added, a distinction that reproduces in the user’s register what the typology named intrinsic divergence and extrinsic divergence, and that leaves the added material ready to pass to the previous criterion. The faithfulness check stops, it should be stressed, at the relation between the output and the document, so that a summary entirely faithful to a mistaken source satisfies this criterion and nonetheless calls for the external check when what is at stake is truth value.
The asymmetry between the two criteria suggests a prior operation that the user fully controls; the condition that makes verification cheap, having a document and a governing instruction, is not given in advance and is produced through query design. The prompt understood as a new maieutics, a category of our own formulated when recovering the Socratic method for interaction with language systems [6], receives here an epistemic function that that work did not attribute to it and that we advance as a contribution of this article, namely that the deliberate formulation of the query, with the source attached and the scope delimited, turns the general case into a derived case and creates the object against which the output can be checked. Whoever asks without delimiting receives an answer that only admits external verification; whoever delimits the query and supplies the material obtains, in the same act, the reference that validation needs.

5.4. Intersection of Types and Criteria

The elements gathered so far admit a joint presentation, shown in Table 1, which crosses the typology of the previous section with the documented procedures and with the criteria just formulated, so that each class of hallucination is accompanied by what research can do with it and by what the user can do without the means of research. A horizontal reading of Table 1 shows the distance between the two registers, which in several rows forces the user into a slower and less conclusive operation than the corresponding automatic procedure; a vertical reading of its second column recalls, for its part, that no procedure rules on truth and that all of them inherit the scope of the reference against which they measure.
Table 1. Intersection of hallucination types, detection procedures, and validation criteria.
Table 1. Intersection of hallucination types, detection procedures, and validation criteria.
Type of hallucination Documented detection procedure and declared limit Applicable validation criterion User operation and capacity exercised
Factuality by contradiction of an established fact Automatic check against an external knowledge base or against retrieval results [8]; the scope ends at the coverage and currency of that base Checking against an external reference (verification) Independently retrieve the source that documents the fact and compare the claim with it; conceptual clarity about what is being checked
Factuality by fabrication without a checkable referent, as in the case of the citation, the attribution of authorship, or the unrecorded datum The same procedures return absence of evidence, a result that does not establish falsehood Checking against an external reference (verification) Attempt the actual retrieval of the cited object by identifier, catalog, or repository; distinction between absence of evidence and established falsehood
Intrinsic faithfulness, in which the output contradicts the supplied document or the governing instruction Textual inference metrics over the source document [7,14]; they rule on support, not on truth Checking against the supplied source (verification) Compare claim by claim against the delivered material; analytical reading and design of the governing instruction
Extrinsic faithfulness, in which the output adds content that the supplied document does not allow to be verified The same procedures flag the claim as unsupported, without resolving its truth value Checking against the supplied source and, for the added content, checking against an external reference Separate what is supported from what is added and submit the added content to external checking; prioritization of what is to be verified
Faithfulness by internal inconsistency of the generated text Zero-resource procedures over successive samples and uncertainty estimation [9,21]; they measure instability and declared confidence, not truth Internal coherence and cross-sample stability (consistency) Read the output in search of contradiction and fallacy, and repeat the query in an independent session to compare; symbolic logic and informal logic
Authors’ elaboration based on [7,8,9,14,21]. The procedures listed in the second column presuppose programmatic access to the model, generation of multiple samples, or availability of a comparison corpus, conditions beyond the reach of a user querying from a public interface. None of the consistency criteria detects a falsehood that the model produces stably.
This section establishes an order of operations and a caveat about its distribution, given that validation begins with fixing the meaning of terms, continues with a consistency review that discards defective outputs at low cost, and culminates in checking against a reference the model did not produce, the only moment at which the relation between what is said and the world actually comes into play. The caveat concerns the distribution of effort, because the initial operations are cheap and limited while the one that decides on truth is costly whenever users must obtain the reference on their own, so that a use of these criteria that stops at coherence and stability leaves untouched the problem that motivated the article; deliberate query design mitigates that cost by bringing the reference into the exchange, without eliminating it, since the supplied document calls for validation in its own right.

6. Formative Consequences and Institutional Conditions

In Section 4 we held that the distance between the detection repertoire of research and that of the ordinary user defines much of the formative problem, and in the previous section that distance took the form of an order of operations whose decisive moment, checking against a reference the model did not produce, is also the costliest. It remains to examine whoever receives the output, together with the conditions that an educational institution can offer so that the burden of validation does not fall entirely on the individual; the consequences are formulated at the conceptual level, and empirical evidence enters with its population and design declared, without any claim to estimate how often what it describes occurs in classrooms.

6.1. Cognitive Satisfaction

Apparent confidence, which in Section 3 we established as a formal property of the output, has a correlate on the reader’s side, and we call that correlate cognitive satisfaction, a category with which we designate the closure of inquiry that occurs when an answer presents itself as finished and the user takes it as sufficient without having mediated in its construction or its verification. The category describes the structure of that acceptance and refrains from estimating how many students or teachers experience it. The relief that accompanies this closure rests, however, on something the user does not see, given that the end of the answer coincides with the end of a sequence governed by a statistical and linguistic criterion foreign to the justification of what is said, in keeping with the absence of a criterion of truth of its own that we attribute to the system.
The mechanism of this closure admits a reading from the theory of epistemic vigilance; Sperber et al. [13] argued that understanding an utterance requires an attitude of tentative trust that ends in acceptance unless vigilance finds reasons to doubt. The generated output offers few occasions for such reasons to arise from its form, since the false statement shares with the true one the assertive surface and internal coherence. In this sense, cognitive satisfaction names the moment at which the absence of visible reasons to doubt is taken as sufficient reason to accept, a shift that goes unnoticed because the understanding of the text proceeded without a hitch.
Among the neighboring literature, the closest notion in the study of these tools is cognitive offloading, which Risko and Gilbert [25] defined as the use of physical action to alter the information-processing requirements of a task so as to reduce cognitive demand; offloading transfers a task to an external resource, whereas cognitive satisfaction falls on inquiry, so that someone who entrusts a search to the system can keep open the question of the correctness of what was obtained and someone who merely reads an answer can close it. Kruglanski and Webster [26], for their part, described the need for cognitive closure as a desire for definite knowledge, present as a stable disposition and as a state elicited by the situation; cognitive satisfaction does not presuppose that motivation, since its condition of appearance lies in the finished form of the output.
Simon [27], in turn, attributed to organisms with limited information and computational capacity the choice of the alternative that meets their aspiration level, the behavior he named satisficing, and in cognitive satisfaction that threshold is met by the appearance of completeness, without the content having been measured against any criterion. The notion closest in meaning is the feeling of rightness that Thompson et al. [28] studied as a metacognitive experience that accompanies an intuitive answer and predicts how much it is reconsidered; that feeling, however, accompanies an answer the reasoner produced, whereas cognitive satisfaction falls on a text generated by a system for which no one answers, so that the signal that closes the review comes from an object whose production the user did not carry out.
The evidence we gather does not measure this category, although Dell’Acqua et al. [29] documented the separation between apparent quality and correctness on which it depends. In their preregistered experiment with 758 consultants who worked with GPT-4 on tasks of their trade, those who used the tool on the task designed to exceed the capability of the system were 19 percentage points less likely to reach a correct answer than those who worked without it, and their recommendations received, whether correct or not, higher ratings of coherence and persuasiveness from evaluators who did not know the solution. These are qualified professionals operating in their own domain, and from the figure we retain that expertise did not prevent the acceptance of erroneous outputs, which is why the category reaches teachers as much as students.

6.2. Externalization of Epistemic Vigilance

Sperber et al. [13] distinguished a vigilance directed at the source, which assesses whether the informant is competent and benevolent, and another directed at the content, which examines its coherence with the background beliefs of the addressee; in the ordinary case both presuppose a communicator who, in asserting, claims sufficient epistemic authority to expect trust, and in Section 3 we noted that the generated output has the appearance of testimony without anyone who holds it. We hold that a consequence for vigilance follows from that lack, given that the criteria of competence and benevolence find no bearer to which they can be ascribed, in correspondence with the absence of a subject of competence that we established against Hila [18], so that vigilance directed at the source loses its object and the weight shifts toward the content and toward the receiver.
It could be objected that users evaluate the system as a whole, in the way Sperber et al. [13] described the tentative trust placed in a search engine on the basis of its observed success, but that general impression does not discriminate between tasks of similar difficulty, and Dell’Acqua et al. [29] showed that the assistance of the system improved some tasks and degraded others within the same workflow, with a boundary that was not evident to the participants. Shao [30] arrived at a convergent observation from misinformation research, noting that the available interventions were designed for identifiable human sources and that hallucinations, having neither author nor agenda, might be questioned less often, and Bielik and Krell [31] extended Sperber’s framework toward the evaluation of the receiver and of the receiver’s biases. What we add to these developments is the reason for the shift, which stems from the epistemic status of the output and therefore affects every user.
With the object of vigilance over the source lost, the formative consequence of greatest scope falls on vigilance over content, which operates on background beliefs that the student does not yet possess in the domain being learned [13]; UNESCO’s guidance makes the very possibility of noticing the error depend on that prior knowledge [1]. The situation corresponds to factuality in an open query, the general case of the typology, where only checking against an external reference is available and where validation is most costly, so that whoever consults in order to learn occupies the position in which verifying is at once most necessary and most onerous, a position that teachers share outside their specialty.
To the offloading of a task may then be added the externalization of epistemic vigilance, which consists in transferring to the system the functions of cross-checking what it asserts and correcting it, in the sense that the literature on technology-mediated offloading has already given to externalization [32]. We hold that what is distinctive about our notion lies in the recipient, since someone who entrusts a calculation to a calculator transfers it to an instrument with a correctness criterion fixed in advance, whereas someone who asks the model to confirm its answer transfers vigilance to a system without a criterion of truth of its own or a subject of competence, so that the delegation lacks a recipient capable of receiving it and hands back to the system what in Section 3 we attributed to whoever receives the output. Vigilance is thus non-delegable to the system that produced the output, while dependence on other sources, among them the references used in verification, remains subject to the ordinary vigilance that Sperber et al. [13] made compatible with trust in what others communicate.
Two other studies document neighboring phenomena whose scope should be specified, the first being that of Gerlich [10], who found among 666 participants in the United Kingdom a negative correlation between frequent use of AI tools and critical-thinking performance, a relation his model presented as mediated by cognitive offloading and which the author himself qualified because of its reliance on self-reported measures; his notion of critical thinking, the ability to analyze, evaluate, and synthesize information, is broader than the logical basis we adopt [6], so that the finding is compatible with our approach without measuring it. In an experiment with university students, Urban et al. [11] found that writing with the assistance of ChatGPT did not reduce the incorrect information taken from an article generated by the system, and that greater overestimation of one’s own performance was moderately associated with incorporating it.

6.3. Institutional Conditions

The non-delegable character of vigilance with respect to the system does not authorize placing it entirely on the individual either, given that Parasuraman and Manzey [33], in reviewing studies on automation bias, concluded that it affects novices and experts alike and cannot be prevented by training or instructions, a finding obtained with decision-support systems predating language models and which we take as a transposed indication. Sperber et al. [13] further noted that vigilance takes institutional forms, such as the certification of competence or peer review, capable of filtering better than the sum of spontaneous vigilances; an institution that merely exhorts people to verify shifts the burden from itself onto the isolated user.
One condition concerns the basis of the criteria, symbolic formal logic and informal logic with its examination of fallacies and biases, whose teaching should be extended to all programs, since exposure to generated output does not distinguish between disciplines; a critical-thinking course built on these two pillars and designed to be taught across all the programs of a university offers a precedent [6]. Training includes teachers, exposed as users to the same cognitive satisfaction, and in UNESCO’s guidance Miao and Holmes [1] recommended developing their capacity for an appropriate use of these tools and institutionally evaluating the long-term effects of these tools on critical thinking.
The cost of checking against an external reference also admits institutional intervention, insofar as access to databases and repositories turns into an ordinary operation what the isolated user must obtain on their own, without eliminating the distance described in Section 4 but reducing the part that depends on means. Alongside that access operates training in query design, in the sense of the prompt as a new maieutics that we specified in Section 5, oriented toward supplying the source and delimiting the scope, a proposal that converges with the structured and verifiable prompt construction that Shao [30] called for in user education. This training differs from generic instruction in prompt writing, since in the experiment by Dell’Acqua et al. [29] the group that received such instruction registered the largest drop in correctness on the task that exceeded the capability of the system.
Assessment offers an additional point of support through tasks that require the verification trail, that is, the record of the claims checked and of the reference against which each was checked, an assignment that by its design renders insufficient the closure we described as cognitive satisfaction and that makes concrete the human responsibility for the accuracy of generated content and the revision of the design of written tasks raised by Miao and Holmes [1]. The requirement also applies to the materials teachers prepare with these tools, whose validation is part of their responsibility.
The conditions described share an assumption, namely that the institution has the means to assume part of the cost of keeping vigilance active on the reader’s side, and that distribution is today part of what it means to educate in the use of these tools.

7. Conclusions

The path we have followed rests on a thesis of an epistemic and functional order, according to which the system has neither access of its own to the output it produces nor a criterion of truth of its own with which to evaluate it, deficits that we assert at the level of functioning and justification, so that the thesis holds whatever position one takes on mind, consciousness, or understanding. From there it follows that every procedure capable of flagging an error imports an external anchor, whether a source corpus, the comparison between successive samples from the same model, another evaluating system, or a reference set by human beings, whereby correction is relocated without ever being internalized. Hallucination thus appears as a possibility inscribed in the way these systems operate, and the formative problem consists in having criteria to deal with it while that possibility remains open.
On that basis, the typology ordered the phenomenon along two axes, factuality, which operates in the open query and constitutes the general case, and faithfulness, which presupposes a governing instruction or a supplied document and is therefore the derived case, with the distinction between the intrinsic and the extrinsic housed within this second axis. The detection procedures documented by research are uneven in scope and each carries a declared limit, while the available figures behave as situated magnitudes, dependent on the architecture evaluated, the prompt, and the measurement method, and none of them warrants a general error rate. In Section 4 we further noted that the detection repertoire within reach of research and the one available to the ordinary user do not coincide, and that distance underlies much of what this work asks of university education.
The response we offered to that distance consisted of four criteria grouped into two families, separated by the origin of the material on which they operate, so that consistency, with internal coherence and cross-sample stability, works on what the model itself delivers, while verification, with checking against an external reference and against the supplied source, introduces material the model did not produce. The order of execution places conceptual clarity first, which fixes the meaning of the terms used in the query, and assigns to formal logic in its symbolic version a bounded yield, since it evaluates the validity of an inference and leaves the truth of its premises untouched; consistency accordingly functions as a low-cost preliminary filter and the weight of validation falls on verification, whose cost lies in the need to obtain the reference. Table 1, which closes Section 5, crosses the classes of hallucination with the documented procedures and with the operation corresponding to the user, and makes visible that a stable falsehood passes unimpeded through the consistency criteria.
On the side of whoever receives the output, the apparent confidence we had established as a formal property of the text finds its correlate in cognitive satisfaction, a category with which we name the closure of inquiry that ensues when an answer presents itself as finished and the user takes it as sufficient without having mediated in its construction or its verification; alongside it operates the externalization of epistemic vigilance, that is, the transfer to the system of the functions of cross-checking what it asserts and correcting it. Vigilance directed at the source loses its object, given that competence and benevolence find no bearer to which they can be ascribed, and the weight shifts toward the content and toward the receiver, with the added difficulty that the student does not yet possess the background beliefs of the domain being learned, a position that teachers share outside their specialty. Vigilance is thus non-delegable to the system that produced the output, without this restriction authorizing its placement entirely on the individual, and from this double restriction follow the institutional conditions we proposed, ranging from the logical basis extended to all programs to an assessment that requires the verification trail, with access to references and training in query design between the two.
Among what these pages add to the discussion is, at the conceptual level, the thesis of external anchoring, by which we hold that no detection mechanism internalizes the criterion of truth the system lacks, and the pair of categories with which we describe the effect of the output on whoever receives it, cognitive satisfaction as the name for the closure that finished form favors, and the externalization of epistemic vigilance in the specification we gave it, where what is distinctive lies in the recipient of the delegation, a system without a criterion of truth of its own or a subject of competence, incapable of receiving what is entrusted to it. At the operational level, the ordering of the criteria into two families and their crossing with the classes of hallucination offer an instrument for direct classroom use, and query design, which we call the prompt as a new maieutics, receives an additional epistemic function, that of turning the general case into a derived case and thereby creating the object against which to check the output.
The scope of what has been argued admits limits that we note in closing, since the work reviews literature without observing classrooms or measuring the prevalence of the phenomenon in a given population, and the categories we proposed describe a structure of acceptance without estimating how often it occurs. Nor does the peer-reviewed literature yet offer a measure of the ability of students and teachers to identify a false output in their own field, a gap that points to a line of research with immediate consequences for the design of teaching, to which is added the empirical testing of the effect that an assessment organized around the verification trail would have on the validation operations students carry out.
The implication of greatest scope for higher education institutions follows from the distribution of the cost of validation, insofar as a university that merely exhorts people to verify transfers to the isolated user a burden it can partly assume, and assuming it amounts to treating validation as a matter of curriculum and assessment design rather than as an individual virtue of the student or the teacher. The trust placed in these tools is defensible when it rests on criteria the user applies, and ceases to be so when it rests on the finished form of the answer, and the formative task consists in sustaining by institutional means the difference between the one situation and the other, while the use of these systems as an assistant retains its place in learning as long as vigilance over what they assert remains active.

Author Contributions

Conceptualization, W.O.C.M. and R.D.M.R.; methodology, W.O.C.M.; investigation, W.O.C.M. and R.D.M.R.; data curation, R.D.M.R.; writing—original draft preparation, W.O.C.M.; writing—review and editing, W.O.C.M. and R.D.M.R.; supervision, W.O.C.M. All authors have read and agreed to the published version of the manuscript.

Funding

This research was funded internally by Universidad Politécnica Salesiana (Ecuador) through the research project “Humanidades e inteligencia artificial: potencialidades y limitaciones a partir de una perspectiva interdisciplinaria” [Humanities and Artificial Intelligence: Potentialities and Limitations from an Interdisciplinary Perspective].

Institutional Review Board Statement

Not applicable. The study did not involve humans or animals.

Data Availability Statement

No new data were created or analyzed in this study. All sources examined are publicly available publications, identified in the reference list; the figures taken from the cited dynamic resource correspond to the access date stated in the text.

Acknowledgments

During the preparation of this manuscript, the authors used Claude (Anthropic; Claude Opus 5) for the purposes of editorial support and for verifying bibliographic records.

Conflicts of Interest

The authors declare no conflicts of interest.

References

  1. Miao, F.; Holmes, W. Guidance for Generative AI in Education and Research; UNESCO: Paris, France, 2023. [Google Scholar] [CrossRef]
  2. Smith, A.L.; Greaves, F.; Panch, T. Hallucination or confabulation? Neuroanatomy as metaphor in large language models. PLoS Digit. Health 2023, 2, e0000388. [Google Scholar] [CrossRef] [PubMed]
  3. Hicks, M.T.; Humphries, J.; Slater, J. ChatGPT is bullshit. Ethics Inf. Technol. 2024, 26, 38, Correction: Ethics Inf. Technol. 2024, 26, 46. https://doi.org/10.1007/s10676-024-09785-3. [Google Scholar] [CrossRef]
  4. Šekrst, K. Chinese Chat Room: AI hallucinations, epistemology and cognition. Stud. Log. Gramm. Rhetor. 2024, 69, 365–381. [Google Scholar] [CrossRef]
  5. Farquhar, S.; Kossen, J.; Kuhn, L.; Gal, Y. Detecting hallucinations in large language models using semantic entropy. Nature 2024, 630, 625–630. [Google Scholar] [CrossRef] [PubMed]
  6. Cárdenas-Marín, W.O. Nuevos horizontes de la filosofía ante el auge de la(s) inteligencia(s) artificial(es). In Bioética. Inteligencia artificial, cambio climático y muerte asistida; Bolaños Vivas, R.F., Ed.; Abya-Yala: Quito, Ecuador, 2025; pp. 13–34, (In Spanish). [Google Scholar] [CrossRef]
  7. Ji, Z.; Lee, N.; Frieske, R.; Yu, T.; Su, D.; Xu, Y.; Ishii, E.; Bang, Y.J.; Madotto, A.; Fung, P. Survey of hallucination in natural language generation. ACM Comput. Surv. 2023, 55, 248. [Google Scholar] [CrossRef]
  8. Huang, L.; Yu, W.; Ma, W.; Zhong, W.; Feng, Z.; Wang, H.; Chen, Q.; Peng, W.; Feng, X.; Qin, B.; Liu, T. A survey on hallucination in large language models: Principles, taxonomy, challenges, and open questions. ACM Trans. Inf. Syst. 2025, 43, 42. [Google Scholar] [CrossRef]
  9. Zhang, Y.; Li, Y.; Cui, L.; Cai, D.; Liu, L.; Fu, T.; Huang, X.; Zhao, E.; Zhang, Y.; Chen, Y.; Wang, L.; Luu, A.T.; Bi, W.; Shi, F.; Shi, S. Siren’s song in the AI ocean: A survey on hallucination in large language models. Comput. Linguist. 2025, 51, 1373–1418. [Google Scholar] [CrossRef]
  10. Gerlich, M. AI tools in society: Impacts on cognitive offloading and the future of critical thinking. Societies 2025, 15, 6, Correction: Societies 2025, 15, 252. [Google Scholar] [CrossRef]
  11. Urban, M.; Brom, C.; Lukavský, J.; Děchtěrenko, F.; Hein, V.; Svacha, F.; Kmoníčková, P.; Urban, K. ‘ChatGPT can make mistakes. Check important info.’ Epistemic beliefs and metacognitive accuracy in students’ integration of ChatGPT content into academic writing. Br. J. Educ. Technol. 2025, 56, 1897–1918. [Google Scholar] [CrossRef]
  12. Searle, J.R. Minds, brains, and programs. Behav. Brain Sci. 1980, 3, 417–424. [Google Scholar] [CrossRef]
  13. Sperber, D.; Clément, F.; Heintz, C.; Mascaro, O.; Mercier, H.; Origgi, G.; Wilson, D. Epistemic vigilance. Mind Lang. 2010, 25, 359–393. [Google Scholar] [CrossRef]
  14. Maynez, J.; Narayan, S.; Bohnet, B.; McDonald, R. On faithfulness and factuality in abstractive summarization. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, Online, 5–10 July 2020; pp. 1906–1919. [Google Scholar] [CrossRef]
  15. Floridi, L. AI as agency without intelligence: On ChatGPT, large language models, and other generative models. Philos. Technol. 2023, 36, 15. [Google Scholar] [CrossRef]
  16. Floridi, L. AI as agency without intelligence: On artificial intelligence as a new form of artificial agency and the multiple realisability of agency thesis. Philos. Technol. 2025, 38, 30. [Google Scholar] [CrossRef]
  17. Sosa, E. A Virtue Epistemology: Apt Belief and Reflective Knowledge; Volume I; Clarendon Press: Oxford, UK, 2007; pp. 22–43. [Google Scholar]
  18. Hila, A. The epistemological consequences of large language models: Rethinking collective intelligence and institutional knowledge. AI Soc. 2026, 41, 79–97. [Google Scholar] [CrossRef]
  19. Zhou, K.; Hwang, J.D.; Ren, X.; Sap, M. Relying on the unreliable: The impact of language models’ reluctance to express uncertainty. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Bangkok, Thailand, 11–16 August 2024; pp. 3623–3643. [Google Scholar] [CrossRef]
  20. Ulmer, D.; Lorson, A.; Titov, I.; Hardmeier, C. Anthropomimetic uncertainty: What verbalized uncertainty in language models is missing. Trans. Assoc. Comput. Linguist. 2026, 14, 1505–1540. [Google Scholar] [CrossRef]
  21. Manakul, P.; Liusie, A.; Gales, M.J.F. SelfCheckGPT: Zero-resource black-box hallucination detection for generative large language models. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, Singapore, 6–10 December 2023; pp. 9004–9017. [Google Scholar] [CrossRef]
  22. Vectara Hallucination Leaderboard. data updated as of 11 May 2026, evaluation by means of HHEM-2.3. Available online: https://github.com/vectara/hallucination-leaderboard (accessed on 20 September 2026).
  23. Pinker, S. Racionalidad. Qué es, por qué escasea y cómo promoverla; Paidós: Barcelona, Spain, 2021. (In Spanish) [Google Scholar]
  24. Talbot, M. Critical Reasoning. Romp Through the Foothills of Logic for Complete Beginners; Metafore, 2017. [Google Scholar]
  25. Risko, E.F.; Gilbert, S.J. Cognitive offloading. Trends Cogn. Sci. 2016, 20, 676–688. [Google Scholar] [CrossRef] [PubMed]
  26. Kruglanski, A.W.; Webster, D.M. Motivated closing of the mind: ‘Seizing’ and ‘freezing’. Psychol. Rev. 1996, 103, 263–283. [Google Scholar] [CrossRef] [PubMed]
  27. Simon, H.A. Rational choice and the structure of the environment. Psychol. Rev. 1956, 63, 129–138. [Google Scholar] [CrossRef] [PubMed]
  28. Thompson, V.A.; Prowse Turner, J.A.; Pennycook, G. Intuition, reason, and metacognition. Cogn. Psychol. 2011, 63, 107–140. [Google Scholar] [CrossRef] [PubMed]
  29. Dell’Acqua, F.; McFowland, E., III; Mollick, E.; Lifshitz, H.; Kellogg, K.C.; Rajendran, S.; Krayer, L.; Candelon, F.; Lakhani, K.R. Navigating the jagged technological frontier: Field experimental evidence of the effects of artificial intelligence on knowledge worker productivity and quality. Organ. Sci. 2026, 37, 403–423. [Google Scholar] [CrossRef]
  30. Shao, A. New sources of inaccuracy? A conceptual framework for studying AI hallucinations. Harv. Kennedy Sch. Misinformation Rev. 2025, 6. [Google Scholar] [CrossRef]
  31. Bielik, T.; Krell, M. Developing and evaluating the extended epistemic vigilance framework. J. Res. Sci. Teach. 2025, 62, 869–895. [Google Scholar] [CrossRef]
  32. Skulmowski, A. The cognitive architecture of digital externalization. Educ. Psychol. Rev. 2023, 35, 101. [Google Scholar] [CrossRef]
  33. Parasuraman, R.; Manzey, D.H. Complacency and bias in human use of automation: An attentional integration. Hum. Factors 2010, 52, 381–410. [Google Scholar] [CrossRef] [PubMed]
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.