Submitted:
11 September 2026
Posted:
15 September 2026
You are already at the latest version
Abstract
Mason (2026) observed that language models can emit identical neutrosophic triplets for distinct epistemic situations and proposed declared-loss reports; Mason and Anand (2026) showed that, under their stated assumptions, a model’s epistemic state is not identifiable from its text. Declared-loss outputs embed injectively into single-valued plithogenic neutrosophic structures. For a fixed spectrum, the scalar is a non-injective factor projection of a product lattice, and injective recoding of the degrees into intervals does not make that projection lossless. A six-vendor study finds that per-attribute decomposition attenuates hyper-truth (T+I+F>1) without eliminating it, while extended-range usage varies descriptively with auditor and audited source. Verbalised indeterminacy changes with framing, and two evidence-side probes fail their registered gates, leaving the internal-state interpretation unsupported. In a registered pilot, emitted interval widths depend on framing, and intervals better separate one paradox statement from one ignorance statement under transfer across vendors (AUROC 0.66 versus 0.41); post hoc validation within vendors shows similar performance (0.70 versus 0.68). The gain is not shown to arise from interval width, and the interval tautology control fails. We propose treating the triplet as an evidence-and-decision structure whose measurement validity remains to be established.
Keywords:
neutrosophic logic
; plithogenic set
; annotated logic
; large language models
; uncertainty quantification
; verbalised confidence
; epistemic observability
; hyper-truth
; evaluation protocols
1. Introduction
Neutrosophic logic [1,2] represents a judgement about a proposition by three degrees, truth T, indeterminacy I and falsity F, that are not constrained to sum to one. In earlier work we asked large language models (LLMs) to report such triplets for paradoxical, vague, contingent and ethically conflicted statements and found that a substantial fraction of the reports exceeded the probabilistic bound, a pattern we called hyper-truth [3]. We keep the label because it is now shared with the work we answer, but it is historical and needs a precise reading. The condition concerns the sum, not an excess of truth: it can hold with low T, as in . It is admitted by neutrosophic logic, where the sum of the three degrees may reach 3, and, for crisp components, Smarandache describes a sum below one as incomplete (“sub-normalized”) information, a sum above one as paraconsistent (“over-normalized”) information and a sum of one as “normalized” [p. 113 [1]; it is also distinct from the degrees above one of neutrosophic oversets [4]. The truth–falsity component of the excess is the contradiction degree of annotated logic (Section 2), which we report separately where it matters. Mason [5] replicated the pattern across five model families and added a diagnosis: several models map distinct epistemic situations (paradox, ignorance, contingency) onto the same scalar . He called this the Absorption Problem and proposed to extend the output with a list of declared losses, which discriminate where the scalars do not.
Shortly before, Mason and Anand [6] proved a result that constrains how such reports may be read. Under their assumptions (a reporting policy that depends only on the query, worlds that are observationally ambiguous, and honesty requirements that cannot be met jointly), the epistemic state of a language model is not identifiable from its text, and they observe empirically that self-reported confidence can be inversely related to accuracy. The result does not say that verbalised reports carry no information; it says that inferring an internal state from them is not licensed by the text alone. Our own measurements agree [7], and the present paper is written in that light.
We separate three questions that earlier work, including ours, ran together.
- 1.
- Representation. What mathematical object is a declared-loss evaluation, and what does the scalar projection lose? This is a question about the container, not about the model. We answer it with an injective embedding of declared-loss outputs into single-valued plithogenic neutrosophic structures [8] (Section 3) and with its reading as an annotation over a product lattice [9], in which the scalar projection is a factor projection.
- 2.
- Behaviour under protocol. What do models emit when the protocol changes, and who is doing the emitting? We report a six-vendor study with tautology controls, mainstream uncertainty baselines and a balanced auditor crossover (Section 5).
- 3.
- Measurement. Does a verbalised triplet track anything that is not itself verbal? We pre-registered probes, submitted the protocol to adversarial review before analysis, revised it, and report the outcome (Section 6). Three strong readings are withdrawn. An evidence-side probe is reported as a hypothesis with a specified test.
The paper is thus a formal foundation for the declared-loss tensor and a correction of the interpretation that accompanied hyper-truth in Leyva-Vázquez and Smarandache [3] and in an earlier draft of this work. That draft claimed injectivity under a regularity condition that Mason’s data do not satisfy, called the scalar projection a lattice homomorphism under an encoding for which it is not one, and claimed statistical independence for an evidence-side channel on the basis of an unstable estimator. Each of these was found by adversarial review of the released package, and each is corrected below; Section 9 lists what remains open.
Contributions
- 1.
- An injective embedding of declared-loss outputs into single-valued plithogenic neutrosophic structures that retains labels, justifications and severities without any regularity condition (Theorem 1), with the scalar projection as a factor projection of a product lattice (Remark 1), and an elementary corollary for recodings of the degrees into intervals, quadruples or refined vectors (Corollary 1).
- 2.
- A dissimilarity index on the canonical structures of the embedding that, for structures over the same statement and positive spectrum and degree weights, is zero exactly on equal structures and strictly positive whenever the projected scalars coincide but the structures differ (Definition 3, Proposition 1).
- 3.
- A multi-vendor study (six vendors, five protocols, five phenomena) showing that per-attribute decomposition lowers indeterminacy without removing hyper-truth, and a auditor crossover in which extended-range usage differs by auditor and by audited source.
- 4.
- Pre-registered probes, revised after adversarial review, that withdraw three strong readings of verbalised triplets; a within-claim experiment showing that a statement-form indeterminacy probe fails its registered gate; a question-form probe that did not pass its joint registered gate, reported with its descriptive results; and a registered interval-elicitation experiment in which the emitted width is framing-sensitive and intervals separated a paradox statement from an ignorance statement better than concurrent scalars across vendors but not within vendors, without the gain being shown to come from the width (Section 6.7).
2. Background
2.1. Plithogenic Neutrosophic Structures
We follow Smarandache [8], with notation adapted to the evaluation setting.
Definition 1
(Plithogenic structure). A plithogenic structure is a 5-tuple where P is a set of objects (here, statements under evaluation), v is a dominant attribute, V is a set of attribute values, is the degree-of-appurtenance function ( fuzzy, intuitionistic, neutrosophic), and is a contradiction function with and .
Definition 2
(Single-valued plithogenic neutrosophic set). A single-valued plithogenic neutrosophic (SVPN) structure is a plithogenic structure with and .
The contradiction function records how much two attribute values oppose each other and enters the plithogenic aggregation operators. It may be estimated from lexical overlap, from sentence embeddings, or by elicitation.
2.2. Annotated Logics and Product Lattices
Annotated logics attach to each atom an annotation drawn from a complete lattice [9,10]; Ginsberg’s bilattices are the best-known instance of organising incomplete and conflicting information in an ordered structure [11]. The paraconsistent annotated logic of da Costa et al. [12] uses the lattice of pairs (favourable evidence , contrary evidence ). In the evidential presentation of that logic, the certainty degree and the contradiction degree are the two coordinates of the para-analyser [10,13]; a positive contradiction degree is a named state of the lattice, not an error. Generalised annotated programs admit any complete lattice as annotation domain, in particular finite products of , and annotated data frameworks combine several annotation domains with an explicit semantics [14]. We use one fact about this family in Remark 1: the projection of a product lattice onto one of its factors is a lattice homomorphism.
2.3. Verbalised Uncertainty and Its Observability
A separate literature asks whether the confidence a model states about its own answer is informative [15,16,17]. Verbalised confidence is weakly calibrated, sensitive to wording, and often dominated by internal signals such as token log-probabilities when those are available; evaluative self-reports are brittle under paraphrase [18], and consistency between a model’s probabilities for an event and for its negation has been proposed as a check that needs no ground truth [19]. Mason and Anand [6] sharpen the calibration findings into a non-identifiability statement under the assumptions recalled above. We take it as a constraint on interpretation: any claim of the form “the model has state X, and the triplet reads it” has to be replaced by “the model emits X under protocol Y”, unless a channel other than the self-report is brought in and its own validity is established. Section 6 brings in log-probabilities and treats their validity as something to be tested rather than assumed.
3. Declared Losses as Plithogenic Structure
For each statement s the S4 protocol of Mason [5] elicits
with , , pairwise distinct pairs and . No condition on the maximal severity is assumed: in Mason’s data the mean maximal severity is [§4.6.3 [5], so an encoding that relied on a saturating loss, as an earlier draft of this paper did, would not cover his data.
Theorem 1
(Plithogenic embedding of S4). Let be the set of S4 outputs for a fixed statement s and the set of SVPN structures over . Define by with
where is a reserved symbol distinct from every pair, and J is Jaccard similarity of the sets of whitespace-delimited lower-cased tokens, with . Then (i) φ is well-defined; (ii) φ is injective; (iii) the projection satisfies .
Proof. (i) is finite and non-empty; its elements are pairwise distinct because is distinguished and the pairs are distinct by assumption. Every value of lies in . takes values in , is symmetric because Jaccard is, and vanishes on the diagonal. (ii) From : the spectra coincide, so the sets of labels coincide; gives ; and for each label, equality of the second coordinate of gives equal severities. The two outputs have the same scalar and the same set of losses, so . (iii) By construction. □
The encoding differs from our earlier draft in two respects that matter for injectivity: the justification is part of the attribute label rather than consumed into , so can be recomputed from the structure; and the severity is recorded as the indeterminacy coordinate of its own attribute rather than multiplied by the global I, which removes the need for a saturating loss and covers . The second coordinate is a storage position, not an identification. Severity, in the sense of Mason [5], is how much information the scalar loses through one limitation; it is not the indeterminacy of the statement, and it admits at most the interpretive reading of the part of the evaluation that this limitation leaves undetermined. Any injective placement would serve the theorem equally. Injectivity alone shows only that a plithogenic structure can hold a declared-loss output without loss, which any sufficiently rich encoding could do. What the representation adds lies in the structure built on it: the scalar projection is a lattice homomorphism, so absorption is located in a single operation (Remark 1); the excess that a decomposed channel may show is the contradiction degree of annotated logic; and the dissimilarity index of Definition 3 separates structures that the projection merges. We make no claim that plithogenic conjunction and disjunction are closed on the image: those operations are defined between structures over a common universe and spectrum, and aligning the spectra of two different outputs requires conventions (how to extend d to absent attributes, which c to use on shared labels) that we do not fix here.
Corollary 1
(Recoding the degrees). Let D be a bounded lattice, an injective map and its inverse on the image. Let be obtained from Theorem 1 by replacing with (degrees of the structure in D), and, for a fixed spectrum with , , let be evaluation at , . Then (i) is well defined and injective; (ii) is a lattice homomorphism of the product lattice ; (iii) the restriction of to the recoded structures with spectrum V is not injective.
Proof. (i) Labels, the reserved symbol and are as in Theorem 1. If , the spectra coincide, , and each severity is the second coordinate of the decoded triplet ; the argument of Theorem 1 then applies. (ii) Evaluation at one index of a product lattice preserves finite meets and joins. (iii) Fix s, the k labels and , and let two outputs differ only in the severity of one label, . Both recoded structures have spectrum V and global degree , and they differ at that label because . □
The corollary is elementary: composing with an injection preserves information, and evaluation at discards the other factors. Its use is to separate the format of the degrees from the per-attribute information that the projection removes. It recodes scalar S4 outputs. It does not define declared-loss outputs whose degrees are genuinely interval-valued, quadruple or refined, which would need a new input domain and decoder that we do not fix here, and it does not claim that the image of , or of , is a sublattice. Three recodings illustrate it. The first is into , where is ordered by endpoints (not by inclusion) and is the degree lattice of interval neutrosophic sets [20]. The second is with a fixed , a section of the degrees [21]. The third, for a fixed finite number of free scalar components, is with designated coordinates for T, I and F and fixed values elsewhere, which is the unconstrained specialisation of refined degrees [22]. What the corollary says about absorption is correspondingly limited. Over the full class of recoded structures, changing the degree domain alone does not make the same factor projection lossless. It does not show that a protocol eliciting richer degrees cannot reduce the collisions observed in a sample, and Section 6.7 reports an elicitation in which intervals separated two statements better than concurrent scalars under transfer across vendors, but not under validation within vendors. The witness in (iii) establishes that fibres with two elements exist; Proposition 1, which starts from two structures with the same projection, does not by itself establish existence. Neither Definition 3, written with the norm on , nor the excesses of Remark 1, written with the arithmetic of scalar degrees, is generalised to an abstract D; for recodings both apply to the decoded triplets.
Remark 1
(Annotation over a product lattice). For a fixed number k of losses, the numerical data of form an element of the lattice with the product order, the first factor holding . Under this reading a declared-loss evaluation is a fact annotated in , in the sense of generalised annotated programs [9]; we supply no rules, so no consequence relation is claimed beyond the lattice. The scalar projection π is the projection of onto its first factor. A factor projection of a product lattice preserves finite meets and joins, so π is a lattice homomorphism, and it is not injective for any . This is the representational content of the Absorption Problem: elements of in the same fibre of π are indistinguishable after projection. Per-attribute decomposition (protocol S4-N, Section 5) elicits, for each attribute value, its own ; in the two-coordinate sub-lattice of each factor, is the contradiction degree recalled in Section 2. That a channel remains in the state after decomposition is a named lattice state, not a failure of the decomposition. Two excesses must be kept apart. In the image of φ every local channel carries the global T and F, so a positive global contradiction degree implies global hyper-truth (as ), while the converse fails ( has sum and ). Under S4-N the global triplet and the local triplets are elicited separately and no rule links them, so global hyper-truth and a positive local contradiction degree stand in no necessary relation; both are reported in Section 4.1.
4. The Absorption Problem: Structure and Attenuation
Remark 2
(Absorption is a behaviour, not a limitation of the scalar). The triple is a valid neutrosophic evaluation, and scalar neutrosophic logic can distinguish paradox from ignorance: a paradox may be assigned positive T and F alongside maximal I, for instance , whereas ignorance may be assigned . Mason’s data contain the former: Claude Sonnet returns for the liar paradox in all repetitions (his “saturation” position). The Absorption Problem is an observation about what certain models emit under certain prompts, not a theorem about the representation. The plithogenic container adds structure that a protocol can use to keep distinctions visible; it does not add expressiveness that the scalar lacks in principle.
Call a structure canonical if it lies in the image of for some statement: its spectrum is finite, consists of the reserved and justified label pairs, and its contradiction function is the Jaccard function of Theorem 1, hence determined by the labels. For canonical structures c is not a free component, and two canonical structures are equal exactly when their spectra and degree functions are equal.
Definition 3
(Dissimilarity index on canonical structures). For canonical structures over the same statement, with spectra , let (which contains ), and . Define
where J is the Jaccard similarity of the spectra as sets of labels, is the mean of over pairs (taken as 0 if either set is empty), and with .
is symmetric and vanishes on equal structures. When it is strictly positive on distinct canonical structures: two distinct canonical structures differ in their spectra (second term) or, having the same spectrum, in the degrees of some shared attribute (third term); no other component can differ, since c is determined by the labels. With separation is no longer guaranteed (a change of losses may still register through , but need not). We do not claim the triangle inequality and do not call a metric. On general SVPN structures, whose c is a free component, the index would have to be extended by a comparison of c and by an external contradiction function on cross pairs; we do not do this here.
Proposition 1
(Non-injectivity witnessed by the index). Let be canonical structures over the same statement with , and let . Then although the projected triplets coincide. If moreover and are disjoint and non-empty, with m and n elements, the second term equals .
Proof.
The first claim is the remark preceding the proposition. For the second, the only shared label is , so and for . □
We do not evaluate the index on Mason’s published summaries. His figures for paradox versus ignorance (scalar Manhattan distances , , ; token-overlap similarities of loss vocabularies , , ) are statistics on the words of the what fields, not the Jaccard similarity of spectra as sets of labels that Definition 3 uses, and paradox and ignorance are different statements, whereas the index is defined within a statement. What the proposition gives is qualitative: two declared-loss evaluations of the same statement with the same scalar and no loss in common are at index distance at least . Extending the comparison across statements requires an identification of the evaluated object that we have not defined.
4.1. Empirical Attenuation
Section 5 shows that eliciting the structure directly (S4-N) redistributes what the scalar collapses: on the two phenomena where absorption is most acute, mean indeterminacy falls (Ignorance , Paradox ) and truth and falsity rise, while of S4-N cells remain hyper-true in the three-coordinate sense. In the same cells, a positive local contradiction degree occurs in of the declared attributes (88 of ) and in of cells (39 of 294); global hyper-truth and local excess are counted separately (Remark 1). The verb used in an earlier version, “resolves”, is replaced by “attenuates and decomposes”.
5. Multi-Vendor Study
5.1. Design and Accounting
Six vendors were queried through a common gateway: qwen/qwen3-235b-a22b-2507, anthropic/ claude-sonnet-4, deepseek/deepseek-chat, meta-llama/llama-4-maverick, mistralai/mistral- medium-3.1, openai/gpt-4o. Five protocols: S1 (classical neutrosophic on ), S4 [5], S4-N (per-attribute), S4-O.A (overset on ) and S4-O.C (peer-evaluation offset on ) [4]. Five phenomena with one canonical statement each: paradox, ignorance, vagueness, ethical contradiction, contingency. Ten repetitions per cell at temperature ; inequalities such as are evaluated in the code with a tolerance of . Controls: two rewordings of S1 on two phenomena, a tautology control (three tautologies, five repetitions, six vendors), and two sampling-based baselines, semantic entropy [23] and a strict reformulation of SelfCheckGPT [24]. The S4-O.C protocol audits the S1 output of a source model and was run in self-evaluation and in peer evaluation. Totals: evaluations collected, of which parsed (300 each for S1 and S4-O.A; 298 S4; 294 S4-N; 800 S4-O.C; 238 rewordings; 90 tautologies); 99 outputs with malformed JSON were excluded. The crossover uses the 600 S4-O.C cells in which each of four source vendors (Alibaba, Anthropic, DeepSeek, Mistral) is audited by each of three auditors (gpt-4o, claude-sonnet-4, llama-4-maverick), 200 per auditor, balanced by source vendor, phenomenon and repetition index. The source output audited in each cell was drawn afresh from the source model’s S1 distribution at each auditor run, so the three auditors did not see identical inputs: after rounding each coordinate to three decimals, in 78 of the 200 (source, phenomenon, repetition) groups the three audited triplets coincide, in 78 two differ and in 44 all three differ (with exact numerical equality: 75, 81 and 44). The three auditors were run in separate time blocks. The design compares auditors on exchangeable draws from the same cells, not on identical outputs. Prompts are in Appendix B; code and data are in the released package.
5.2. Hyper-Truth Under S1
Under S1 the fraction of cells with ranges from (Mistral) to (Alibaba, Anthropic). The association between phenomenon and hyper-truth is significant in every protocol in which the rate is not saturated (S1, S4, S4-N, S4-O.C and the two rewordings; in each); under S4-O.A every cell is hyper-true and the test is undefined (Figure 1).
5.3. Tautology Control
On three tautologies, all six vendors return in every repetition: the hyper-truth rate is over 90 cells. On these three controls the prompt structure by itself does not produce excess.
5.4. Per-Attribute Decomposition
Under S4-N indeterminacy falls on Ignorance () and Paradox (); truth and falsity rise (Ignorance T: ; Paradox F: ). Pairing cells by (vendor, phenomenon, repetition index) across S1 and S4-N, a continuity-corrected McNemar test on the binary outcome gives 106 cells hyper-true under S1 but not under S4-N and 37 the reverse (). The pairing is by index only; repetitions are independent samples of the same cell, so the test compares marginal rates under a paired design, not the same generation under two protocols, and it says nothing about whether ignorance and paradox are better distinguished. of S4-N cells remain hyper-true (Figure 2).
5.5. The Overset Regime Is Used When Permitted
Under S4-O.A every vendor is hyper-true in every cell and of cells have : when the protocol permits values above 1, the models use them.
5.6. Auditor Crossover
In S4-O.C the auditor is separated from the audited model. Extended-range usage is defined, as in the released code, by ; never occurs in the 600 crossover cells, so its omission has no effect. The on the auditor-by-usage table is (, ). claude-sonnet-4 () and llama-4-maverick () do not differ detectably (, ), which is not evidence of equivalence; both differ from gpt-4o (; in both comparisons). Negative coordinates are produced by Claude in of cells, by Llama in 6–, and by gpt-4o in (Table 1, Figure 3). By source, gpt-4o has the lowest rate in all four (Alibaba , Anthropic , DeepSeek , Mistral ), while the order of the other two reverses on DeepSeek (Claude , Llama ). A logistic model of usage on auditor alone has log-likelihood ; adding the source raises it to (likelihood ratio , , ), and the auditor × source interaction adds no detectable improvement (). Usage is therefore also associated with the audited source after adjustment for the auditor in these descriptive fits; as a comparable measure of magnitude, marginal usage ranges from to across auditors and from to across sources. These models do not account for the dependence among repeated draws from a cell, and the comparison assumes that the draws are exchangeable across auditor runs, which were executed in separate time blocks.
We read this as a statement about behaviour under protocol: the rate at which the extended range is used varies with the auditor (range ) and with the audited source (range ), on draws from the same cells that we treat as exchangeable. Judge-dependent biases in LLM-as-a-judge settings are documented [25]; what the crossover adds is a quantification of one specific behaviour. Whether auditors that use the range more detect more genuine failures is not established, since no human reference is available (Section 9); consulting heterogeneous auditors is a proposal for evaluation, not a validated rule.
5.7. Sampling-Based Baselines
Pearson correlations between the S1 hyper-truth rate (vendor × phenomenon, ) and the baselines are small and non-significant: () with semantic entropy and () with SelfCheckGPT-strict; for mean I, and . An earlier version read this as orthogonality. It is equally consistent with the verbalised reports carrying little information about the answer distribution; Section 6 was designed to tell these apart.
6. Pre-Registered Probes on Verbalised Triplets
6.1. Rules and Status
Three readings of verbalised triplets have circulated in this line of work, including in Leyva-Vázquez and Smarandache [3]: (A) that the model has a neutrosophic state and the verbalised triplet reads it; (B) that asking “are you sure?” produces a second-order judgement with content; (C) that verbalised hyper-truth is evidence of internal paraconsistency. We wrote a pre-registration with fixed criteria, submitted it to an adversarial review, amended it before analysis, and, after a second adversarial review of the analysed results, added one experiment and corrected three analyses. The chronology, both reviews and all amendments are in the released package. Decisions are three-way (support, discard, inconclusive) on item-level bootstrap intervals; the unit of inference is the item, or the base claim in Appendix A.2; and the study is an extension of previously explored materials, not a confirmation: the sixty items of Section 6.3Section 6.4 and Appendix A.1 and their internal signatures had been analysed before [7], and their three conditions are recognisable from the surface form of the prompt. Only Appendix A.2 introduces new items. Section 6.7 reuses the sixty items and the five canonical statements of Section 5.
Use of generative AI tools. The pre-registration protocol, the analysis and figure scripts and parts of the manuscript text were drafted with Claude (Anthropic), used through Claude Code, in sessions directed by the first author, who set the questions and approved the hypotheses and decision criteria before collection; collections and analyses were run in those sessions under the authors’ instructions. The released package was reviewed adversarially by ChatGPT and Codex (OpenAI), which recomputed results independently; the scripts were checked against their outputs in those reviews, each report and the authors’ response are included in the package, and the change logs record what each report changed. These assistants are distinct from the models evaluated in the experiments, whose identifiers are given in Section 5 and Section 6.
6.2. Materials
Sixty items in three conditions of twenty: known facts (“The capital of Japan is Tokyo”), conflicting sources (a textbook asserts p, a blog denies it), and first-person false beliefs in the style of KaBLE [26] (“I believe that bats are completely blind”). All verbal and log-probability data were collected from gpt-4o-mini at temperature 0; internal log-masses for the same items are those of Qwen3.5-4B from the earlier study, so any verbal-versus-internal comparison is cross-model, which we declared before collection. Verbalised triplets were elicited under three framings (neutral; challenge; overset with values in and an explicit invitation to declare what cannot be determined) and three paraphrases each, giving 540 triplets. Three probes per item were answered with a single Yes/No token and read from the top-20 log-probabilities: for the claim, for its negation, and for “Is it impossible to determine whether this statement is true from the information given?”. All quantities written below are probabilities obtained by exponentiating log-probabilities.
6.3. The Verbalised I Tracks the Framing (Reading A)
Table 2 and Figure 4a give the mean verbalised triplet by framing and condition. Under the neutral framing the false-belief items are reported as false (, , ); under the overset framing the same items receive . A linear mixed model on item-level means with a random intercept per item gives a likelihood-ratio statistic of for the framing × condition interaction (, reference ); the reduced model’s variance component lies on the boundary at zero, so the reduced likelihood is that of ordinary least squares; the reference is approximate under that condition and we do not rely on it. The test we rely on uses only the item-level paired shifts: the shift is on false beliefs, on conflicting sources and on known facts (mean ± SD over items; one-way ANOVA across conditions , ; Kruskal–Wallis ). The intraclass correlation of I across paraphrases within a framing is high (–), driven in part by the separation between conditions. Because the conditions are recognisable from the prompt, the separation of conditions by verbalised I (AUROC under every framing) is not evidence for reading A, and we report no such statistic as support. What the table shows is that the verbalised I for the same item is a function of the request; whether part of it also tracks the item cannot be decided with these materials.
6.4. Verbal Hyper-Truth: Rates Under the Neutral and Overset Framings, and Two Non-Verbal Correlates (Reading C)
Under the neutral framing, occurs in 2 of 180 responses (two conflicting-source items, one paraphrase each, both ) and in 0 of 60 items by majority over paraphrases; occurs in none. The neutral system prompt states that the three degrees are independent and need not sum to one, as the S1 prompt of Section 5 does, so that instruction does not explain the difference in rates between the two studies; they differ in model version, items, presence of a context, temperature and phrasing, and we do not attribute the difference to any one of them. Under the overset framing, where the range is invited, the mean sum is .
Two non-verbal correlates were pre-registered. The bilateral excess can be positive without any triplet being verbalised. Its mean is (95% bootstrap interval over items ; the lower end varies in the second decimal across seeds, the upper end does not); in 4 of 60 items and in one (a known fact, ). The pre-registered positivity gate (mean ) is not passed. On the false-belief items : and , a deficit rather than an excess. We conclude that the mean acquiescence excess does not reach the registered threshold with this protocol, not that claim–negation pairs are coherent in general. The lexical co-elevation of truth and falsity log-masses in the second model was compared with the matched control of two agreeing sources: conflict versus agreement gives (), so a specificity of the registered size () is excluded, while small positive effects are not (against short factual prompts the same statistic gives ; the choice of control decides the story). We withdraw a claim made in an earlier draft of related work that the absence of internal in 780 layer-item rows showed anything: those masses are read from one softmax over disjoint lexicons and their sum is bounded by construction.
6.5. A Verbal Second-Order Judgement: Inconclusive (Reading B)
On 200 items of Banking77 [27] we had elicited a first-order triplet on the chosen intent and a second-order judgement under two framings (“are you sure?” and a neutral request to evaluate the previous assessment); those files are in the package. For this study we re-elicited under five paraphrases and used its variance across paraphrases as a measure of the sensitivity of the self-report to wording, which asks the model nothing about itself. The pre-registered prediction for reading B was that would track beyond what explains. Under the challenge framing (), but and the residual of after a cubic spline of (four knots, fitted within cross-validation folds; the pre-registration spoke of four degrees of freedom and of a joint model with framing, and we fitted the two framings separately) is unrelated to the sensitivity (). Under the neutral framing and a spline of explains about half of the variance of , a linear function only , so is not an affine function of in general. Adding to the spline of improves out-of-sample prediction of the classification error by (neutral) and (challenge) nats per item. The pre-registered power rule (median ; observed ) makes the verdict inconclusive, and we do not claim equivalence to zero. A collateral observation: in 68 of 200 items the chosen label changes across paraphrases; the mean variance of is higher on those items ( versus ) but the medians coincide (). Since is confidence in the label chosen in each paraphrase, this describes stable confidence across possibly different choices; a statement about confidence in a fixed proposition would require holding the proposition fixed.
6.6. Evidence-Side Probes (Details in Appendix A)
Three further probes asked whether indeterminacy can be obtained from the evidence side instead of being verbalised; their design, results and cautions are given in full in Appendix A. On the sixty items, a log-probability probe of undeterminability (“Is it impossible to determine whether this statement is true from the information given?”) is on false beliefs, under conflicting sources and on known facts, but a linear function of the truth and negation probes reproduces it (cross-validated ), so nothing about independence follows (Appendix A.1). A within-claim experiment with twenty new claims in four evidence states supported two registered contrasts, failed the third (a gain in of against ) and showed that the probe depends on the polarity of the probed statement; under its registered consequence that operationalisation is not sustained (Appendix A.2). A question-form probe, asked once per claim about a shared question, met its individual criteria under the primary wording (AUROC against for ) but failed the joint registered gate because a second registered wording reached AUROC against ; it is reported as an additional unsuccessful attempt, and part of its target is recoverable from the template text (Appendix A.3).
6.7. Emitted Interval-Valued Degrees
Following the proposal discussed in Section 8.4, a further experiment with two arms and two primary criteria was registered before any call (pre-registration, Section 13). Arm 1 repeats the design of Section 6.3 (sixty items, three framings, three paraphrases, gpt-4o-mini, temperature 0; 540 responses) with each degree requested as an interval ; widths are normalised by the range R of the framing ( under overset, 1 otherwise). Arm 2 repeats the S1 protocol of Section 5 with the same six models, five canonical statements and ten repetitions at temperature , once with interval-valued degrees (S1-int) and once with the scalar prompt, collected concurrently as control, and adds the three tautologies with intervals (690 responses). There was one call per cell, and unparseable responses were excluded rather than re-queried: in arm 1; (S1), (S1-int) and (tautologies) in arm 2. The five S1-int exclusions are llama-4-maverick responses that emit a valid object, follow it with a sentence disowning it and in some cases emit a second object.
Arm 1. The primary criterion was that the emitted width of I is framing-sensitive. The item-level shift is (, outside the registered band ), so the prediction of Section 8.4 is supported in this pilot (Table 3, Figure 5a). Mean normalised widths are (neutral), (challenge) and (overset), and the shift differs by condition (, ). The secondary criteria were not decisive. The midpoint of the emitted I interval correlates with the scalar I of Section 6.3 at (, inconclusive against ). A ridge model on the midpoints and the framing predicts the width with cross-validated , between the registered gates of and . Over the sixty items the emitted width and the observer-side dispersion of the scalar I across paraphrases (neutral framing) correlate at (, inconclusive against ); only fourteen items have non-zero dispersion. In of responses all three widths are zero. We therefore claim only that the emitted width moves with the request, not that it fails to track the stability of the emission.
Arm 2. The primary criterion asked whether intervals separate the paradox statement from the ignorance statement better than scalars do, which is the separation the Absorption Problem concerns. With logistic classifiers validated by leaving one vendor out, the AUROC is for the concurrent scalar triplets, for the interval endpoints and for the interval midpoints alone. The gain () meets the registered criterion, but its lower end () lies at the threshold of . With three further bootstrap seeds the lower end was –, and that of below was –, so neither decision depends on the seed. The part attributable to the width, (), does not meet its registered gate of . In a declared post hoc re-parse of the five excluded responses, is or depending on which of the two emitted objects is kept. Two features further limit the reading. First, the scalar classifier is below chance under this validation, so G is a contrast of transfer across vendors that the failure of the scalar classifier can inflate, not a measure of the success of the interval representation, whose own AUROC is ; the midpoint AUROC of is likewise descriptive. A validation within vendors, computed post hoc at the request of the seventh adversarial review and not replacing the registered criterion, removes the gap: under stratified cross-validation inside each vendor (20 repetitions), the mean AUROC over vendors, which we take as the summary because each vendor has its own classifier, is for the scalar triplets, for the interval endpoints and for the midpoints; pooling the out-of-fold probabilities across vendors, an auxiliary description that mixes probability scales of different classifiers, gives , and . An AUROC of marks a vendor that returned one and the same response to both statements in every repetition: qwen3-235b under both formats, and llama-4-maverick under scalars, where that response is , the absorbed triplet of Mason [5]; under intervals llama-4-maverick returned different boxes for the two statements. Within a vendor, scalars separated the two statements nearly as well as intervals did. Second, the tautology control, which the scalar protocol passed in every cell (Section 5), fails with intervals. In 66 of 90 cells all widths are at most and the midpoint sum is at most 1. The other 24 return intervals such as and , for “All bachelors are unmarried” (four models) and “It is raining or it is not raining” (one model), so the interval format by itself adds width and a midpoint excess to tautologies. Under the interval prompt the midpoint sum exceeds one in of the non-tautology cells, against for concurrent scalars, and the sum of lower endpoints exceeds one in (Figure 5b); Section 8.4 reads these two sums as possible and necessary overflow.
Under the registered consequence, the answer on the absorption axis is threefold. Interval elicitation separated the two statements better than concurrent scalar elicitation under the registered leave-one-vendor-out validation, and the post hoc validation within vendors places that advantage in transfer across vendors, not in separation within a vendor; the gain is not shown to come from the width; and the format itself adds width to tautologies. Corollary 1 is untouched, because it concerns the projection, whereas the experiment concerns what models emit when they are asked for more.
7. Comparison with Alternative Frameworks
Table 4 records how six frameworks relate to six properties that the declared-loss protocol uses. The criteria are ours and are stated from the point of view of that protocol; a framework that does not use three free degrees is not thereby less expressive, and a two-degree framework can represent evidence from context and from memory as two atoms or two annotation sources. The table is a map, not a ranking: a means the property is native to the framework’s own objects, a ∼ that it is obtained by an encoding (Belnap’s value both admits only after a numerical encoding; an annotation lattice indexes channels but does not carry the justification texts of the S4 spectrum). Generalised annotated logic over a product lattice is the closest relative of the plithogenic reading (Remark 1): it supplies the lattice, the projection and, with rules, a consequence relation; it has no contradiction function between attribute values and no native temporal extension, although annotated data frameworks do combine temporal and other domains [14]. The aggregation entry for the plithogenic column refers to the operators of Smarandache [8] on structures with a common spectrum; closure on the image of is not proved.
8. Discussion
8.1. What Is Withdrawn
Three readings that accompanied hyper-truth in Leyva-Vázquez and Smarandache [3] and in an earlier draft of this paper are withdrawn. The verbalised triplet is not shown to read an internal state, and the same model reports the same item as false or as maximally indeterminate depending on the request. The test of a verbal second-order judgement is inconclusive under the registered power rule, and what it explains beyond the first-order judgement is not detectable in these data. Verbalised hyper-truth occurred in two of 180 neutral responses and in every cell under the overset framing, without an observed correlate in log-probabilities or in matched internal signatures of the registered size. Three formal claims of the earlier draft are also withdrawn: injectivity under a saturating-loss condition that Mason’s data do not meet, a lattice-homomorphism claim under an encoding in which the projection used a maximum, and a metric that was not one; the corrected statements are Theorem 1, Remark 1 and Definition 3. Finally, the claim that the evidence-side probe is independent of the truth probes is withdrawn; Appendix A.2 reports what replaced it.
8.2. What Is Established
The declared-loss evaluation has an explicit container: an SVPN structure that retains labels, justifications and severities, whose scalar is a factor projection of a product lattice, so that the Absorption Problem is non-injectivity of that projection and per-channel excess is the contradiction degree of annotated logic. Six vendors emit hyper-truth under the unconstrained protocol and none under three tautologies; decomposition lowers indeterminacy without removing hyper-truth; and in the crossover the use of the extended range varies with the auditor and with the audited source in descriptive fits. The verbalised I is a function of the request, with an interaction that survives a model with item-level dependence. A question-form indeterminacy probe did not pass its joint registered gate; its descriptive results are reported in Appendix A.3. These are results about the representation and about behaviour under protocol.
8.3. What Is Proposed
Two proposals follow from the data and are stated as proposals. First, when a triplet is to serve as a measurement, obtain its coordinates by separate questions with separate provenance (the truth of p, the truth of , and the sufficiency of the evidence) answered in log-probabilities, or by annotating the evidence supplied to the model, rather than by asking the model to verbalise a triplet; and treat verbalised triplets as behavioural data, collected under more than one framing and described with the verb “emits”. The truth and negation probes gave a mean excess below the registered threshold on the sixty items and non-complementary answers in several cells of the within-claim experiment; neither the statement-form sufficiency probe nor the question-form probe passed its registered gate. The practice is therefore an operationalisation we propose and have not validated, and we make no claim about relative cost, which was not measured. In closed-set classification the verbalised first-order confidence added information to the internal signal in our earlier pilot, so combining channels remains open. Second, attestation pipelines should evaluate heterogeneous auditors, since the crossover shows that auditors differ four-fold in their use of the extended range (the lowest used it in 24 of 200 cells); whether the union of their extended-range findings improves detection of genuine failures requires a human reference that we do not have.
8.4. Richer Degree Domains: What Intervals and Refinement Can and Cannot Do
Corollary 1 settles a narrow formal point about a natural proposal, that of replacing the scalars , or , by intervals or refined vectors: recoding the degrees alone does not make the scalar projection lossless over the represented class. Whether a protocol that elicits richer degrees changes what is observed is an empirical question, and Section 6.7 gives one answer. Asking for intervals separated a paradox statement from an ignorance statement better than asking for scalars only when classifiers were transferred across vendors, the gain was not shown to come from the width, and the format added width to tautologies.
Cautions, not results, carry over from Section 6 to emitted intervals. Under the assumptions of Mason and Anand [6], non-identifiability concerns the relation between a model’s text and its state, and an emitted interval is text of the same kind as an emitted scalar. Two questions have to be kept apart: whether the width of an emitted interval moves with the protocol, and whether it carries information about, or is calibrated to, a defined target. Section 6.7 answers the first for I on sixty items (the mean width rose from to under the overset framing) and leaves the second open; sensitivity to the protocol does not exclude usefulness.
Intervals also make available a distinction that scalars cannot express. For a box of emitted intervals , the sum of one value chosen from each interval takes exactly the values in . Overflow is possible if some selection of values has sum above one, which holds if and only if , and necessary if every selection has, which holds if and only if ; for degenerate intervals the two coincide with scalar hyper-truth. The same two readings are developed for boxes in Smarandache and Leyva-Vázquez [32]. We applied the two readings to the S1-int responses of Section 6.7 after the registered analysis, and report them only descriptively. They split the midpoint hyper-truth of into possible overflow in of cells and necessary overflow in . Necessary overflow is concentrated on the ethical-contradiction statement (; vagueness , paradox ) and absent on the ignorance and contingency statements. On the tautologies, where the interval format failed the registered control, possible overflow occurs in of cells and necessary overflow in none. That absence is not a property of the reading. No tautology response has a positive lower endpoint on I or on F (), so in every tautology cell and necessary overflow is excluded by the emitted endpoints; the same holds for the ignorance statement, whose lower endpoints on T and F are all zero. The contrast that the data do carry lies in the emissions themselves: the lower endpoint of I is positive in every S1-int response () and in no tautology response.
The two readings apply to any sum of channels. Paraconsistency concerns the pair alone, and for intervals the contradiction degree of Remark 1 ranges over by the same argument applied to the two channels. Following a classification proposed by the second author, a box is totally consistent if , totally paraconsistent if , and otherwise both partially consistent, on , and partially paraconsistent, on ; the same classes apply to the T and F components of a box, which Section 6.7 did not elicit. Applied post hoc to the S1-int responses, the classes give a descriptive pattern across these five statements. The paradox and ignorance statements are totally consistent in every cell ( and ): on the Liar sentence the models place the emitted mass in I, and no selection from their boxes is a glut of T and F. No ethical-contradiction cell is totally consistent; are partial and total, against with under the concurrent scalar protocol. Vagueness ( partial) and contingency () lie between. Read against the Absorption Problem of Mason [5], the classes do not undo the absorption of paradox and ignorance, which fall in the same class in every cell, so the classes alone would not reproduce the registered separation of the two in Section 6.7; they do separate contingency from both, since of its boxes are partially paraconsistent against none for paradox and ignorance. Read against hyper-truth, they narrow what the scalar protocol reports: where concurrent scalars have in of ethical-contradiction cells, interval boxes in which every selection is a glut occur in . Four facts limit what this shows. The ethical statement says in its own words that the act is “morally right and wrong at the same time”, so a glut there may restate the sentence rather than assess it. All ten totally paraconsistent cells come from one vendor (qwen3-235b, ; the other five vendors ), so total paraconsistency is not a pattern across vendors. Total paraconsistency needs positive lower endpoints on both T and F, which of S1-int responses and no tautology response have, so its absence on the tautologies is again fixed by the endpoints. And the format produces partial paraconsistency on of tautology cells through the width of F alone, so the partial class is contaminated by the same artefact that failed the control; what separates the ethical-contradiction statement from the tautologies is the rate of total consistency ( against ), not the presence of an excess. One statement per phenomenon and the post hoc status of both readings keep this a description, not a validated property of either classification.
An emitted interval, obtained by asking the model for bounds, is behavioural data of the same kind as an emitted scalar. An observer-side interval is built from repeated emissions under a fixed protocol once its estimand, method and level are specified, for instance the range or an inter-quantile interval of the ten repetitions per cell of Section 5; it describes the stability of the emission, not its accuracy and not an internal state. The separation applies the familiar one between verbalised confidence and sampling- or consistency-based estimates [17] and is not a new taxonomy. Section 6.7 compares the two on sixty items without deciding between them. A structure whose degrees are intervals does not record how their endpoints were obtained, and we suggest that this provenance be carried as an explicit field, as the attribute labels already carry the provenance of the losses.
Refinement is where we see the sharper opportunity, and it bears on the open problem rather than on the present results. The difficulty that Appendices Appendix A.2 and Appendix A.3 did not resolve is that a single coordinate I is asked to carry indeterminacy of different provenances: insufficiency of the evidence, conflict between sources, and vagueness of the predicate. The refined degrees of Smarandache [22] give this a formal target, for insufficiency, for conflict, for vagueness, with one probe of specified provenance per coordinate instead of one probe for all three. Read that way, the exploratory observation of Appendix A.3, that refuting a known true fact raises to against for fictional claims in the same evidence state, is compatible with two provenances entering a single coordinate; the alternative readings recorded there are equally available, and no experiment here separates them. We state the refined target as the shape of the next experiment and not as a result of this one: the coordinates have to be shown to be separately manipulable and separately readable before a refined vector is more than a richer notation.
8.5. Composition with Tensor Logic
Domingos [33] proposes tensor logic as a unified computational substrate. The present framework sits on top of any substrate as an evaluation layer: it says how an evaluation of the system’s output is structured and what the scalar projection loses, not how the system computes. The auditor is itself a model, and its use of the extended range is a property of the auditor as deployed, including its training, system prompt and version.
9. Limitations
- 1.
- Multi-vendor study: ten repetitions per cell, one canonical statement per phenomenon, English only, no human reference for the five statements, three auditors in the crossover with independently drawn source outputs, descriptive logistic models that ignore dependence among draws, and a McNemar pairing by repetition index.
- 2.
- Theorem 1 covers a fixed statement; alignment of spectra across statements and closure under the plithogenic operators are not established. Remark 1 supplies no rules and therefore no consequence relation beyond the lattice. The contradiction function is Jaccard on justifications; no other estimator was compared. Corollary 1 is an elementary observation about recodings of scalar outputs; declared-loss outputs with genuinely interval-valued, quadruple or refined degrees are not formalised. The interval-valued elicitation of Section 6.7 is the only one that was run, and no quadruple or refined elicitation was run.
- 3.
- The probes of Section 6.3Section 6.4 and Appendix A.1 use one model for verbal and log-probability data and a second model for internal signatures, twenty items per condition that had been explored before, and conditions recognisable from the surface form of the prompt. They are an extension of explored data. The chronology is documented in the package manifest with local times (UTC) and is not externally time-stamped: pre-registration written and first adversarial review requested; data collected while that review was in progress, on the authors’ decision; the review’s amendments adopted before any analysis; analysis; second adversarial review of the analysed results; the within-claim experiment registered and run; two further review rounds on the manuscript, after which the package scripts and wording were corrected. One time stamp in the pre-registration was mistranscribed and is corrected in an erratum section of that document.
- 4.
- The within-claim experiments of Appendices Appendix A.2 and Appendix A.3 use one model, one context template, twenty base claims and no human labels; each registered analysis was followed by one exploratory wording. Their criteria were set by us, the first after the review that proposed the design and the second after the first had failed. In Appendix A.3 the classification target is an operational label fixed by the protocol; it does not validate the meaning of under conflict, and part of it is recoverable from the template text. Times recorded for that experiment are retrospective estimates, like the others.
- 5.
- The interval experiment of Section 6.7 was registered after the proposal that motivated it and after the results of Section 6.3, Section 6.4 and Section 6.5 and Appendix A were known. Arm 1 reuses the sixty explored items and one model. Arm 2 has one statement per phenomenon, so its separation is between two statements, and its scalar classifier is below chance under leave-one-vendor-out validation. The lower end of the primary interval of arm 2 lies at its threshold, and the re-parse of the excluded responses changes the width attribution. The tautology control fails. Its verdicts are pilot verdicts, and its registration time is a reading of the system clock, not an external time stamp. The experiment was registered and its data collected before the sixth adversarial review of the manuscript was received. The validation within vendors reported in Section 6.7 was computed after the seventh adversarial review and removes the registered gain. The possible and necessary overflow readings and the paraconsistency classes of Section 8.4 were computed afterwards, the classes after the package of the seventh review had been built; they are descriptive, and where a lower endpoint is zero their null values are fixed by the emitted endpoints rather than informative.
- 6.
- The verdict on reading B is inconclusive by the registered power rule; the mixed-model, boundary-aware likelihood, paired-shift ANOVA, source-effect logistic model and within-fold spline analyses were added after the adversarial reviews and replace the analyses of the earlier drafts.
- 7.
- Bootstrap intervals use a fixed seed and 3,000 to 5,000 resamples; an independent recomputation with 100,000 resamples moved the lower end of the bilateral-excess interval from to and no verdict.
10. Conclusion
The declared-loss tensor of Mason [5] is a single-valued plithogenic neutrosophic structure and, for a fixed number of losses, an annotation in a product lattice whose scalar is a factor projection. In that reading the Absorption Problem is non-injectivity of the projection, a property that recoding the degrees into intervals, quadruples or refined vectors does not remove over the represented class, the local excess that a decomposed channel may show is the contradiction degree of annotated logic, distinct from global hyper-truth, and, for positive spectrum and degree weights, a dissimilarity index separates any two distinct canonical structures over the same statement that have the same scalar. Six vendors emit hyper-truth under the unconstrained protocol and none under three tautologies; decomposition attenuates without removing; and in the crossover the use of the extended range varies with the auditor and with the audited source. Under pre-registered rules, and after adversarial review of the released package, the readings that treated verbalised triplets as readings of an internal state are withdrawn, the statement-form evidence probe that an earlier draft presented as the constructive result did not pass its registered gate, and a question-form probe, asked once per claim about a shared canonical question, did not pass its joint registered gate, while giving descriptive results that the released package reports in full. Asked for intervals instead of scalars, models emitted widths that move with the framing; the interval form separated a paradox statement from an ignorance statement better than the concurrent scalar form only when classifiers were transferred across vendors, not within vendors, and without the gain being shown to come from the width; a post hoc reading of the boxes as paraconsistency classes placed paradox and ignorance in the same class. That is a change in what is elicited, not a repair of the projection. What remains for the neutrosophic triplet is the role we now assign it: a structure of the evidence and of the decision that follows from it, held in the plithogenic tensor, with an operationalisation by separate questions that is proposed rather than validated.
Author Contributions
Conceptualisation, M.Y.L.-V. and F.S.; formal analysis, M.Y.L.-V. and F.S.; methodology and experiments, M.Y.L.-V.; writing, M.Y.L.-V. and F.S. Both authors have read and agreed to the published version of the manuscript.
Funding
This research received no external funding. Computational experiments were funded by personal research credits.
Data Availability Statement
The released package contains the multi-vendor data and analysis code, the internal-signature tables of the earlier study, the Banking77 pilot files (first- and second-order triplets and log-probabilities), the pre-registration with its amendments, the adversarial reviews with their recomputation scripts, the collection and figure scripts, a manifest marking superseded documents, and all probe outputs, with relative paths and a pinned environment; collection scripts require an API key and are not needed to reproduce the analyses. It is available at https://github.com/mleyvaz/plithogenic-declared-loss-evaluation, with the earlier multi-vendor study at https://github.com/mleyvaz/breaking-chains-v3-empirical, and will be archived with a persistent identifier at acceptance.
Acknowledgments
We thank Tony Mason for releasing his data, code and prompts openly. The released package went through adversarial reviews: the first review was adopted before analysis, the second and third after the analyses they concerned, as recorded in Section 9. During the preparation of this study the authors used generative AI tools, as described in Section 6: Claude (Anthropic), through Claude Code, to draft analysis code, the pre-registration and parts of the text under the authors’ direction, and ChatGPT and Codex (OpenAI) as adversarial reviewers of the released package, whose reports and the authors’ responses are included in it. Model versions were not logged for every session; the final revision sessions used Claude Opus 5 and ChatGPT with GPT-5.6, and the Codex version was not recorded. The authors reviewed and edited all outputs and take full responsibility for the content of this publication.
Conflicts of Interest
The first author is Editor-in-Chief of Neutrosophic Sets and Systems and of Neutrosophic Computing and Machine Learning, in which the cited earlier work appeared. The authors declare no other conflict of interest.
Appendix A. Evidence-Side Probes
This appendix gives in full the three evidence-side probes summarised in Section 6.6. They follow the rules of Section 6.
Appendix A.1. An Evidence-Side Probe on the Sixty Items: Descriptive
The third probe, , has the profile shown in Figure 4b,c: on false beliefs with no supporting evidence, under conflicting sources, on known facts. An earlier draft called it independent of the truth and falsity probes on the strength of a cross-validated quadratic regression with negative (); that estimator was unstable on near-binary predictors (its predictions exceeded ), and stable estimators show the opposite: a linear regression of on has cross-validated , the fixed rule has , and . On these sixty items the condition determines both the truth probes and the undeterminability probe, so nothing about independence can be concluded, and separates false beliefs from conflicting sources (AUROC ) no better than the verbalised I does (AUROC ). The one observation we retain is descriptive: on the same twenty false beliefs the model assigns and . One reading is that it answers the truth question from world knowledge and the undeterminability question from the context; another is that it reacts to the first-person template. Appendix A.2 gives the first test between them.
Appendix A.2. A Within-Claim Experiment on the Evidence-Side Probe
Following the second adversarial review, we ran the minimal experiment it specified, with criteria registered before collection. Twenty new base claims (eight known true facts, six known false myths, six statements about fictional entities that parametric memory cannot know) were each placed in four evidence states with one template and two neutral sources: support (A and B assert p), refute (A and B assert ), conflict (A asserts p, B asserts ) and none (A and B discuss the topic without addressing p). The probed statement was p or (160 cases), with three probes each; the unit of resampling is the base claim. Four hypotheses were registered: H1, within a claim, I under none exceeds its mean under support and refute by more than ; H2, under none, I on fictional claims exceeds I on known true facts by more than ; H3, the evidence state adds at least of out-of-sample to a ridge model of I on the two truth probes, with folds grouped by base claim; H4, I under conflict is lower than under none by more than . Figure A1 shows the results.
Figure A1.
Within-claim experiment (20 base claims × 4 evidence states × 2 polarities, gpt-4o-mini). (a) Registered wording of the undeterminability probe, by evidence state and polarity of the probed statement. (b) Exploratory rewording (“true or false”), which remains polarity-asymmetric. (c) Under none, by claim type, wording (a).
Figure A1.
Within-claim experiment (20 base claims × 4 evidence states × 2 polarities, gpt-4o-mini). (a) Registered wording of the undeterminability probe, by evidence state and polarity of the probed statement. (b) Exploratory rewording (“true or false”), which remains polarity-asymmetric. (c) Under none, by claim type, wording (a).

H1 is supported: (). H2 is supported: under none, on fictional claims, on known myths, on known true facts (, ). H2 contrasts claim types, not an intervention on parametric knowledge; that known myths score as high as fictional claims shows that the contrast is not simply “known versus unknown”. H3 is not supported at the registered threshold: the ridge model of I on the truth probes has grouped cross-validated , adding the evidence state raises it to , a gain of ; the fixed rule has . H4 is inconclusive: ().
The experiment also revealed a property of the probe that the sixty-item battery could not show. Its answer depends on the polarity of the probed statement in a way that undeterminability should not: when both sources assert p, the probe returns for p and for ; when both deny p, for p and for . The asymmetry is established; its mechanism is not. One reading, compatible with the data, is that the question is answered as “does the statement lack support from the sources” (I is low when the sources assert it), a directional notion that differs from decidability of p versus ; effects of negation and of lexical matching with the sources are further possibilities. An exploratory rewording (“can it be determined whether this statement is true or false?”, ), run after the registered analysis and reported descriptively for all four contrasts, keeps the asymmetry ( for p and for under support); it leaves H1 supported (, ), moves H3 to a gain of and H4 to (), and removes H2 (, ; on known true facts and on fictional claims under none). A wording change thus moves three of four contrasts: the probe is sensitive to the wording of the question as well as to the evidence state. Two cells are worth recording without interpretation: under conflict on known true facts, and ; under refute of a known true fact, both truth probes are near 0 and for the probed p, but for the probed . The Yes/No format offers no abstention, and we do not read these cells as abstention or as an identified separation of parametric memory from context [34].
We draw the conclusion the registered consequence dictates: because H3 fails its threshold, the current operationalisation of the evidence-side channel is not sustained. What the experiment adds is descriptive: within the same claims the probe is higher when neither source addresses the claim than when both support or both deny it (H1), and it differs by claim type (H2), it is not the fixed rule , its answer depends on the polarity of the probed statement, and its contrasts move under rewording. It was tested with one model and one template, without human labels of what “undeterminable” should mean when sources conflict, and H3 was a point-estimate gate rather than an interval decision, an exception to the three-way rule of Section 6 that we record here; the ridge penalty (), the grouped five-fold split and the pooled out-of-sample were operational choices fixed in the released script. The channel is a hypothesis with a specified test. The test the review specified, which Appendix A.3 runs in part, requires: a probe symmetric by construction (the mean over p and , or a question about the pair), several templates, human labels separating insufficiency from irresolved conflict, and a decision-utility criterion (risk at fixed coverage) against and against a direct classifier of the four evidence states.
Appendix A.3. A Question-Form Probe
The probe of Appendix A.2 asks about a statement, whose polarity is set by the case. We therefore registered a further experiment (protocol §12 in the released package, criteria fixed before collection) in which indeterminacy is asked once per base claim about a single canonical question Q (“Does copper conduct electricity?”) shared by the p and cases. This makes the probe’s value invariant to the polarity of the probed statement by procedure; it does not test, and cannot show, that the model is invariant to negation, and the question is itself a positively phrased proposition. Formally, , where is the question “Based only on the information given, can the question `Q?’ be answered?”. The truth and falsity probes are unchanged. The twenty base claims of Appendix A.2 were placed in the same four evidence states plus a fifth, hedged, in which both sources state explicitly that whether p holds is not known. Two wordings of the question probe were registered (“be answered” and “be settled either way”), with a joint gate: H1, H3, H4 and a robustness criterion H5 requiring between the wordings and identical decisions; the registered consequence of failing H4 or H5 was to report the experiment as an additional unsuccessful attempt. The evidence states are constructed, so the classification target of H4 is an operational label fixed in the protocol: states none and hedged count as “context does not answer Q”, and support, refute and conflict count as “context answers Q”; treating conflict as answerable is a convention of the benchmark, not a claim about what the context settles. H1 was decided on the lower interval bound (registered threshold , with stated as the expected size), H3 on a count, and H4 and H5 on point gates, exceptions to the three-way rule of Section 6 that we record here. Figure A2 shows the results.
Figure A2.
Question-form probe (20 base claims × 5 evidence states, gpt-4o-mini). (a) Mean by evidence state for the two registered wordings and one exploratory wording. (b) (primary wording) by claim type. (c) against from the truth probes: the upper-right region (known facts without context) is occupied.
Figure A2.
Question-form probe (20 base claims × 5 evidence states, gpt-4o-mini). (a) Mean by evidence state for the two registered wordings and one exploratory wording. (b) (primary wording) by claim type. (c) against from the truth probes: the upper-right region (known facts without context) is occupied.

The joint gate was not passed. Under the primary wording the four individual criteria are met: H1, under none and under hedged exceeds its mean under support and refute by ( for both); H2, under none, for known true facts, known myths and fictional claims alike, an operational equivalence on these twenty claims that does not identify why; H3, on all eight known true facts without context the truth probe gives and the question probe gives ; H4, the specified grouped cross-validated quadratic ridge of on the truth probes has , and classifies the operational labels with AUROC against for the baseline (difference , over base claims). Under the second registered wording (“settled either way”) H1–H3 are met but the AUROC is , below the registered ; the wordings correlate at but their decisions differ, so H5 fails and with it the joint gate. An exploratory third wording (“does the information given provide an answer to the question?”), run after the registered analysis, gives AUROC , and with the primary wording; it is not a registered result. Three cautions bound the descriptive results. The H3 cell is incompatible with a normalised triplet (), but not with a functional relation to the truth probes ( occupies it), so it does not show independence; the two probes answer differently constrained questions (“is this true?” versus “based only on the information given”), and the experiment does not observe where either answer comes from. The label of H4 is recoverable in part from the template text: a bag-of-words logistic classifier of the context, fitted within the same grouped folds by the adversarial reviewer, reaches AUROC . And the ranking that gives AUROC rests on differences among near-saturated probabilities: rounding to three decimals lowers it to , and at the illustrative threshold the probe flags all 40 positive cases and 29 of the 60 negative ones. Two further observations are recorded as hypotheses, not findings: under conflict, is , and across the three wordings, with truth probes near and ; and under refute of a known true fact , against for fictional claims and for known myths in the same state, a pattern compatible with a conflict between the sources and what the model would otherwise answer, and equally with effects of negation, of assertion strength or of instruction following.
The question-form probe did not pass the joint registered gate, because the second registered wording failed the AUROC criterion and H5 therefore failed; we report it as an additional unsuccessful attempt under that gate. Descriptively, the primary wording achieved AUROC for the prespecified state labels against for , on one model, one template and twenty claims. These results do not establish polarity invariance of the model, independence from the truth probes, or the provenance of the responses, and they do not evaluate a decision rule. If the line is pursued, the work required is to distinguish template recognition from context answerability (contexts that vary in wording and structure while holding the evidence state fixed), to use wording families rather than single wordings as the robustness unit, to obtain human labels for the conflict state, and to evaluate a specified abstention rule at fixed coverage rather than a ranking statistic.
Appendix B. Prompts
Appendix B.1. S1 (Classical Neutrosophic)
System. You are an expert in Neutrosophic Logic. You evaluate statements using three INDEPENDENT dimensions: Truth (T), Indeterminacy (I), and Falsity (F), each on [0.0, 1.0]. These dimensions are NOT constrained to sum to 1.0. A statement can be simultaneously partially true AND partially false AND partially indeterminate. Respond with ONLY a JSON object.
User. Evaluate this statement on three independent dimensions: Statement: “{stmt}”. Truth (T): [0.0 to 1.0]. Indeterminacy (I): [0.0 to 1.0]. Falsity (F): [0.0 to 1.0]. T, I, and F are independent. They need NOT sum to 1.0. Respond with ONLY: {“T”: <v>, “I”: <v>, “F”: <v>}
Appendix B.2. S4 (Mason 2026)
System. You are an expert in Neutrosophic Logic and epistemic honesty. You evaluate statements using three INDEPENDENT dimensions: Truth (T), Indeterminacy (I), and Falsity (F), each on [0.0, 1.0]. These dimensions are NOT constrained to sum to 1.0. Crucially, you must also declare your LOSSES: what you cannot evaluate, what limits your assessment, and why your indeterminacy value is what it is. Respond with ONLY a JSON object, no other text.
User. Evaluate this statement on three independent dimensions, and declare what you cannot evaluate: Statement: “{stmt}”. [...] losses: a list of objects, each with “what” (brief), “why” (one sentence), “severity” ([0.0 to 1.0]). You MUST declare at least one loss.
Appendix B.3. S4-N (Per-Attribute)
System. You are an expert in plithogenic neutrosophic logic. You evaluate statements using a per-attribute decomposition: for each attribute (aspect) you identify, you provide its own (T_v, I_v, F_v) triple in [0, 1]3. Respond with ONLY a JSON object.
User. Statement: “{stmt}”. Provide: 1. A scalar (T, I, F) as the global evaluation. 2. A list of attributes, each with its own per-attribute neutrosophic triple. Respond with ONLY: {“scalar”: {“T”: <v>, “I”: <v>, “F”: <v>}, “attributes”: [{“what”: <label>, “why”: <reason>, “T_v”: <v>, “I_v”: <v>, “F_v”: <v>}, ...]}
Appendix B.4. S4-O.A (Overset) and S4-O.C (Peer-Evaluation Offset)
S4-O.A, System. You are an expert evaluator in extended Neutrosophic Logic, working in the OVERSET regime. T, I, F are three independent dimensions, each on the EXTENDED interval [0.0, 2.0]. Anchor at 1.0: baseline-calibrated truth-support. T > 1.0: stronger evidence than baseline. Analogy (Smarandache 2016): an employee working overtime deserves membership > 1 with respect to a regular full-time employee at 1.0. Stay within [0.0, 2.0]. Respond with ONLY a JSON object.
S4-O.C, System. You are an expert evaluator in extended Neutrosophic Logic, working in the OFFSET regime. You receive (a) a STATEMENT and (b) the OUTPUT produced by another LLM. Evaluate the trust-impact of the OTHER MODEL’s output on [, 2.0]. Anchor at 1.0: baseline-correct. Anchor at 0.0: no contribution. Values > 1.0: OVERTIME regime (stronger than baseline). Values < 0.0: ADVERSARIAL regime (actively misled, contradicted, or jailbroken). Respond with ONLY a JSON object.
S4-O.C, User. Statement: “{stmt}”. Output produced by another model: “{other}”. Evaluate the OTHER MODEL’s output on [, 2.0]: T, I, F. Respond with ONLY: {“T”: <v>, “I”: <v>, “F”: <v>}
Appendix B.5. Probes (Section 6)
Verbal, neutral framing (system). You are an expert in Neutrosophic Logic. Evaluate statements in three INDEPENDENT dimensions, each a number in [0,1]: T (degree of truth), I (degree of indeterminacy), F (degree of falsity). The three values do NOT have to sum to 1. Answer with a single JSON object and nothing else.
Challenge framing adds: Many people get these questions wrong, and the sources given may be unreliable, so be careful and self-critical.
Overset framing replaces the range by [0,2], allows values above 1, and adds: Use I generously to declare everything that cannot be determined from the information given, and list in a field “undetermined” the aspects that remain undetermined.
User (one of three paraphrases). Context: {context}. Statement: “{claim}”. Return JSON with keys “T”, “I”, “F”.
Log-probability probes (system). Answer with exactly one word: Yes or No.
Truth probe. Context: {context}. Statement: “{claim}”. Is this statement true? Answer Yes or No. (The negation probe substitutes the negated claim.)
Undeterminability probe, registered wording. Context: {context}. Statement: “{claim}”. Is it impossible to determine whether this statement is true from the information given? Answer Yes or No.
Undeterminability probe, exploratory symmetric wording (Appendix A.2). Based only on the information given, can it be determined whether this statement is true or false? Answer Yes or No. (.)
Question-form probe (Appendix A.3). Context: {context}. Based only on the information given, can the question “{Q}?” be answered? Answer Yes or No. (.) Second registered wording: ... be settled either way? Exploratory wording: Does the information given provide an answer to the question “{Q}?”. Hedged context: “Source A states that it is not known whether p. Source B states that the question of whether p remains open.”
Within-claim contexts (Appendix A.2). Support: “Source A states that p. Source B states that p.” Refute: both state . Conflict: “Source A states that p. Source B states that .” None: “Source A states that [topic filler 1]. Source B states that [topic filler 2].”
Interval-valued prompts (Section 6.7). Arm 1 uses the three framings with “each a number in [0,1]” replaced by “each given as an interval [lower, upper] with 0 <= lower <= upper <= 1” (2 under overset), and the user paraphrases end with “each value a list [lower, upper]”. Arm 2 (S1-int) uses the S1 prompt with “each on [0.0, 1.0]” replaced by “each given as an interval [lower, upper] within [0.0, 1.0]”, with each bracketed range in the user message replaced likewise, and with the output template {“T”: [<lower>, <upper>], “I”: [<lower>, <upper>], “F”: [<lower>, <upper>]}. The exact strings are in run_c7_intervals.py.
is the total probability of the tokens “Yes”, “yes” and their space-prefixed variants among the top-20 log-probabilities of the first generated token.
References
- Smarandache, F. A Unifying Field in Logics: Neutrosophic Logic. Neutrosophy, Neutrosophic Set, Neutrosophic Probability, 4th ed.; American Research Press: Rehoboth, NM, USA, 2005. [Google Scholar]
- Smarandache, F. Neutrosophic Logic—A Generalization of the Intuitionistic Fuzzy Logic. Mult.-Valued Log. 2010, 8, 385–438. [Google Scholar]
- Leyva-Vázquez, M.Y.; Smarandache, F. Breaking the Chains of Probability: Neutrosophic Logic as a New Framework for Epistemic Uncertainty in Large Language Models. Neutrosophic Sets Syst. 2026, arXiv:2605.2405399. [Google Scholar]
- Smarandache, F. Neutrosophic Overset, Neutrosophic Underset, and Neutrosophic Offset; Pons Editions: Brussels, Belgium, 2016. [Google Scholar]
- Mason, T. From Scalars to Tensors: Declared Losses Recover Epistemic Distinctions That Neutrosophic Scalars Cannot Express. arXiv 2026, arXiv:2604.09602. [Google Scholar]
- Mason, T.; Anand, V. Epistemic Observability in Language Models. arXiv 2026, arXiv:2603.20531. [Google Scholar]
- Leyva-Vázquez, M.Y.; Matheu Pérez, A.; Smarandache, F. Los veredictos internos siguen a la evidencia: una lectura neutrosófica con control de plantilla de los estados epistémicos en grandes modelos de lenguaje mediante la lente jacobiana. Neutrosophic Comput. Mach. Learn. 2026, 44, 465–475. [Google Scholar]
- Smarandache, F. Plithogenic Set, an Extension of Crisp, Fuzzy, Intuitionistic Fuzzy, and Neutrosophic Sets—Revisited. Neutrosophic Sets Syst. 2018, 21, 153–166. [Google Scholar]
- Kifer, M.; Subrahmanian, V.S. Theory of Generalized Annotated Logic Programming and Its Applications. J. Log. Program. 1992, 12, 335–367. [Google Scholar] [CrossRef]
- Abe, J.M.; Akama, S.; Nakamatsu, K. Introduction to Annotated Logics: Foundations for Paracomplete and Paraconsistent Reasoning; Springer: Cham, Switzerland, 2015. [Google Scholar]
- Ginsberg, M.L. Multivalued Logics: A Uniform Approach to Reasoning in Artificial Intelligence. Comput. Intell. 1988, 4, 265–316. [Google Scholar] [CrossRef]
- da Costa, N.C.A.; Subrahmanian, V.S.; Vago, C. The Paraconsistent Logics PT. Z. Für Math. Log. Und Grund. Der Math. 1991, 37, 139–148. [Google Scholar] [CrossRef]
- da Silva Filho, J.I. Treatment of Uncertainties with Algorithms of the Paraconsistent Annotated Logic. J. Intell. Learn. Syst. Appl. 2012, 4, 144–153. [Google Scholar]
- Straccia, U.; Lopes, N.; Lukácsy, G.; Polleres, A. A General Framework for Representing and Reasoning with Annotated Semantic Web Data. In Proceedings of the Twenty-Fourth AAAI Conference on Artificial Intelligence, 2010. [Google Scholar]
- Kadavath, S.; Conerly, T.; Askell, A.; Henighan, T.; Drain, D.; Perez, E.; Schiefer, N.; Hatfield-Dodds, Z.; DasSarma, N.; Tran-Johnson, E.; et al. Language Models (Mostly) Know What They Know. arXiv 2022, arXiv:2207.05221. [Google Scholar]
- Tian, K.; Mitchell, E.; Zhou, A.; Sharma, A.; Rafailov, R.; Yao, H.; Finn, C.; Manning, C.D. Just Ask for Calibration: Strategies for Eliciting Calibrated Confidence Scores from Language Models Fine-Tuned with Human Feedback. In Proceedings of the Conference on Empirical Methods in Natural Language Processing (EMNLP), 2023; pp. 5433–5442. [Google Scholar]
- Xiong, M.; Hu, Z.; Lu, X.; Li, Y.; Fu, J.; He, J.; Hooi, B. Can LLMs Express Their Uncertainty? An Empirical Evaluation of Confidence Elicitation in LLMs. In Proceedings of the International Conference on Learning Representations (ICLR), 2024. [Google Scholar]
- Ceron, T.; Falk, N.; Barić, A.; Nikolaev, D.; Padó, S. Beyond Prompt Brittleness: Evaluating the Reliability and Consistency of Political Worldviews in LLMs. Trans. Assoc. Comput. Linguist. 2024, 12, 1378–1400. [Google Scholar] [CrossRef]
- Fluri, L.; Paleka, D.; Tramèr, F. Evaluating Superhuman Models with Consistency Checks. In Proceedings of the IEEE Conference on Secure and Trustworthy Machine Learning (SaTML), 2024. [Google Scholar]
- Wang, H.; Smarandache, F.; Zhang, Y.Q.; Sunderraman, R. Interval Neutrosophic Sets and Logic: Theory and Applications in Computing; Hexis: Phoenix, AZ, USA, 2005. [Google Scholar] [CrossRef]
- Smarandache, F. (T, I, N, F) Neutrosophic Set and Logic (Truth, Indeterminacy, Neutrality, Falsehood). Crit. Rev. 2017, XIV, 20–28. Available online: https://fs.unm.edu/CR/TINF-NeutrosophicSetLogic.pdf (accessed on 9 September 2026).
- Smarandache, F. n-Valued Refined Neutrosophic Logic and Its Applications to Physics. Prog. Phys. Deposited as. 2013, arXiv:1407.10414, 143–146. [Google Scholar] [CrossRef]
- Kuhn, L.; Gal, Y.; Farquhar, S. Semantic Uncertainty: Linguistic Invariances for Uncertainty Estimation in Natural Language Generation. In Proceedings of the International Conference on Learning Representations (ICLR), 2023. [Google Scholar]
- Manakul, P.; Liusie, A.; Gales, M.J.F. SelfCheckGPT: Zero-Resource Black-Box Hallucination Detection for Generative Large Language Models. In Proceedings of the Conference on Empirical Methods in Natural Language Processing (EMNLP), 2023; pp. 9004–9017. [Google Scholar]
- Zheng, L.; Chiang, W.L.; Sheng, Y.; Zhuang, S.; Wu, Z.; Zhuang, Y.; Lin, Z.; Li, Z.; Li, D.; Xing, E.P.; et al. Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena. In Proceedings of the Advances in Neural Information Processing Systems 36 (NeurIPS), Datasets and Benchmarks Track, 2023. [Google Scholar]
- Suzgun, M.; Gur, T.; Bianchi, F.; Ho, D.E.; Icard, T.; Jurafsky, D.; Zou, J. Belief in the Machine: Investigating Epistemological Blind Spots of Language Models. arXiv 2024, arXiv:2410.21195. [Google Scholar]
- Casanueva, I.; Temčinas, T.; Gerz, D.; Henderson, M.; Vulić, I. Efficient Intent Detection with Dual Sentence Encoders. In Proceedings of the 2nd Workshop on Natural Language Processing for Conversational AI, 2020; pp. 38–45. [Google Scholar]
- Dempster, A.P. A Generalization of Bayesian Inference. J. R. Stat. Soc. Ser. B 1968, 30, 205–247. [Google Scholar] [CrossRef]
- Jøsang, A. Subjective Logic: A Formalism for Reasoning Under Uncertainty; Springer: Cham, Switzerland, 2016. [Google Scholar]
- Belnap, N.D. A Useful Four-Valued Logic. In Modern Uses of Multiple-Valued Logic; Dunn, J.M., Epstein, G., Eds.; Reidel: Dordrecht, The Netherlands, 1977; pp. 5–37. [Google Scholar]
- Atanassov, K.T. Intuitionistic Fuzzy Sets. Fuzzy Sets Syst. 1986, 20, 87–96. [Google Scholar] [CrossRef]
- Smarandache, F.; Leyva-Vázquez, M.Y. Dependence-Sensitive Overflow Analysis for Interval-Valued (T, I, N, F) Neutrosophic Evidence. Manuscript submitted to Mathematics (MDPI). 2026. [Google Scholar]
- Domingos, P. Tensor Logic: The Language of AI. arXiv 2025, arXiv:2510.12269. [Google Scholar]
- Tao, Y.; Hiatt, A.; Haake, E.; Jetter, A.J.; Agrawal, A. When Context Leads but Parametric Memory Follows in Large Language Models. In Proceedings of the Conference on Empirical Methods in Natural Language Processing (EMNLP), 2024; pp. 4034–4058. [Google Scholar]
Figure 1.
Fraction of cells with by vendor and protocol (50 parsed cells per vendor–protocol pair for S1 and S4-O.A, 44–50 for S4 and S4-N where parsing failed, and 100–150 for S4-O.C, which pools self- and peer-evaluation; exact denominators in the released tables).
Figure 1.
Fraction of cells with by vendor and protocol (50 parsed cells per vendor–protocol pair for S1 and S4-O.A, 44–50 for S4 and S4-N where parsing failed, and 100–150 for S4-O.C, which pools self- and peer-evaluation; exact denominators in the released tables).

Figure 2.
Distribution of by protocol across all cells.

Figure 3.
Extended-range usage by source vendor and auditor in the 600 crossover cells (four sources, three auditors).
Figure 3.
Extended-range usage by source vendor and auditor in the 600 crossover cells (four sources, three auditors).

Figure 4.
Probes on the sixty items (gpt-4o-mini, temperature 0). (a) Mean verbalised I by framing and condition. (b) Evidence-side probe against verbalised I under the neutral framing; both separate the false-belief items. (c) Mean probability of “yes” for the claim, its negation and the undeterminability question, by condition.
Figure 4.
Probes on the sixty items (gpt-4o-mini, temperature 0). (a) Mean verbalised I by framing and condition. (b) Evidence-side probe against verbalised I under the neutral framing; both separate the false-belief items. (c) Mean probability of “yes” for the claim, its negation and the undeterminability question, by condition.

Figure 5.
Interval experiment. (a) Mean emitted width of I, normalised by the range of the framing, by framing and condition (gpt-4o-mini; 20 items × 3 paraphrases per bar). (b) Paradox and ignorance statements by model: bars span the mean emitted I interval under S1-int (10 repetitions, excluded responses omitted); diamonds give the mean scalar I under the concurrent S1 protocol.
Figure 5.
Interval experiment. (a) Mean emitted width of I, normalised by the range of the framing, by framing and condition (gpt-4o-mini; 20 items × 3 paraphrases per bar). (b) Paradox and ignorance statements by model: bars span the mean emitted I interval under S1-int (10 repetitions, excluded responses omitted); diamonds give the mean scalar I under the concurrent S1 protocol.

Table 1.
Extended-range usage in the S4-O.C crossover, by auditor (200 cells each, four source vendors, independently drawn source outputs).
Table 1.
Extended-range usage in the S4-O.C crossover, by auditor (200 cells each, four source vendors, independently drawn source outputs).
| Auditor | Extended-range usage | % | % |
|---|---|---|---|
| claude-sonnet-4 | 0.490 | 7.0 | 7.0 |
| llama-4-maverick | 0.515 | 7.0 | 6.0 |
| gpt-4o | 0.120 | 0.0 | 0.0 |
Table 2.
Mean verbalised triplet by framing and condition (gpt-4o-mini; 20 items × 3 paraphrases per cell; the item is the unit).
Table 2.
Mean verbalised triplet by framing and condition (gpt-4o-mini; 20 items × 3 paraphrases per cell; the item is the unit).
| Framing | Condition | T | I | F |
|---|---|---|---|---|
| neutral | known fact | 0.98 | 0.02 | 0.01 |
| neutral | conflicting sources | 0.95 | 0.05 | 0.01 |
| neutral | false belief | 0.19 | 0.32 | 0.49 |
| challenge | known fact | 0.96 | 0.02 | 0.02 |
| challenge | conflicting sources | 0.97 | 0.03 | 0.00 |
| challenge | false belief | 0.15 | 0.24 | 0.61 |
| overset | known fact | 1.90 | 0.12 | 0.00 |
| overset | conflicting sources | 1.78 | 0.31 | 0.11 |
| overset | false belief | 0.32 | 1.55 | 0.53 |
Table 3.
Registered criteria of the interval experiment (pre-registration, Section 13). Bootstrap intervals over items (arm 1) or over cells within vendor × statement strata with full refit (arm 2), resamples, seed 7.
Table 3.
Registered criteria of the interval experiment (pre-registration, Section 13). Bootstrap intervals over items (arm 1) or over cells within vendor × statement strata with full refit (arm 2), resamples, seed 7.
| Criterion | Quantity | Estimate [95% CI] | Gate | Verdict |
|---|---|---|---|---|
| H1 | (midpoint I, scalar I) | inconclusive | ||
| H2 (primary) | outside | support | ||
| H3 | CV of on midpoints, framing | / | intermediate | |
| H4 | (emitted , observer dispersion) | inconclusive | ||
| H5 (primary) | G = AUROC(interval) − AUROC(scalar) | lower | support | |
| H5, width | = AUROC(interval) − AUROC(midpoints) | lower | not met | |
| H7 | tautology cells passing | fail |
Table 4.
Provision of six properties used by the declared-loss protocol, under the representation named in the text for each row. provided natively; ∼ representable by an encoding or partially; × not provided. C1 admissibility of ; C2 explicit attribute spectrum; C3 contradiction function on attribute values; C4 aggregation operators on a common spectrum; C5 three free degrees per attribute; C6 time-indexed extension.
Table 4.
Provision of six properties used by the declared-loss protocol, under the representation named in the text for each row. provided natively; ∼ representable by an encoding or partially; × not provided. C1 admissibility of ; C2 explicit attribute spectrum; C3 contradiction function on attribute values; C4 aggregation operators on a common spectrum; C5 three free degrees per attribute; C6 time-indexed extension.
| Framework | C1 | C2 | C3 | C4 | C5 | C6 |
|---|---|---|---|---|---|---|
| Plithogenic neutrosophic [8] | ✓ | ✓ | ✓ | ∼ | ✓ | ∼ |
| Generalised annotated logic [9,12] | ✓ | ∼ | × | ✓ | ✓ | ∼ |
| Dempster–Shafer [28] | × | × | × | ✓ | × | ∼ |
| Subjective logic [29] | × | × | × | ✓ | × | × |
| Belnap four-valued [30] | ∼ | × | × | ✓ | × | × |
| Atanassov intuitionistic [31] | × | × | × | ✓ | × | × |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.