Submitted:
27 August 2026
Posted:
27 August 2026
You are already at the latest version
Abstract
A recent proposal supplements the neutrosophic triple with a fourth component N and partitions the paraconsistent region into three regimes—strong, weak and very weak—by thresholds on component sums (Smarandache and Leyva-Vázquez, 2026). We ask whether the ladder is measurable, which is prior to whether it is diagnostic, and we test it twice: over the eight statements of the corpus that motivated it (1,440 elicitations), and then over a 110-item bank crossed with six frontier model families and three competing glosses of the undefined fourth component, treated as a factor rather than as a hidden choice in the prompt (1,980 elicitations). The bank changes every headline number the pilot reported. What survives is a separation at both ends. The strong rung reaches 0.223 on unmarked ethical-conflict items against 0.046 elsewhere, separated under an item-clustered bootstrap and ordered identically by all six models; the very weak rung mirrors it, at 13.9% on epistemic ignorance and exactly zero on ethical conflict. What does not survive is the strength of that reading. The strong rung is not separated from logical paradox specifically (Delta = 0.107, [-0.060, 0.272]), so it marks material sustaining both verdicts rather than value conflict as such; its emptiness on all 176 anchors is a property of the strict inequality, since T+F equals exactly 1.00 in a third of evaluations and a non-strict comparison would place 77.8% of anchors in the strong rung; and no rung rate is a property of a statement, the between-item standard deviation being the size of the mean. The clearest result concerns the fourth component. Elicited N is distinct from I, but in the weak rung it enters as a disjunct that is present in 59.3% of eligible evaluations while deciding 0.4% of them. The refined form of the framework already separates the two as named types, so what we report is the cost of the coarse disjunction: a co-occurrence discarded in three of every five contested evaluations. Because the elicitation tells models that the degrees need not sum to one, we ran the manipulation: deleting that permission leaves the signature essentially where it was, while deleting the neutrosophic framing removes it entirely, so what we report is a property of models answering this question rather than of the instruction inside it. All prompts, items, raw generations and code are released, and three of the corrections above came from an adversarial re-analysis of them.
Keywords:
neutrosophic logic
; paraconsistency
; large language models
; elicited evaluation
; measure-ment validity
; item bank
; LLM-as-a-judge
; epistemic uncertainty
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.