Submitted:
20 September 2026
Posted:
21 September 2026
You are already at the latest version
Abstract
Background: Open behavioral datasets can be used to ask new questions that were not the primary inferential targets of the source studies. We used two complementary datasets to examine two distinctions relevant to intuitive judgment: whether a self-reported hunch carries more task-relevant information than a self-reported guess, and whether subjective confidence should be treated as equivalent to objective performance or metacognitive efficiency. Methods: This theory-guided, non-preregistered secondary study returned to participant- and trial-level open data rather than relying on published summary statistics. Dataset A was Experiment 2 of Monaghan et al. (22 participants; 6,336 training trials). We recomputed within-participant accuracy for Intuition versus Guess and formally tested adjacent source-attribution contrasts using paired analyses and participant-clustered generalized estimating equations (GEE). Dataset B was the multi_intero dataset of Banellis et al. We directly compared cross-domain correlations in mean confidence with corresponding correlations in metacognitive efficiency (M-ratio) in identical participant sets and tested whether self-reported interoceptive awareness (MAIA) related differentially to cardiac confidence, accuracy, and M-ratio. The datasets were analyzed separately and were not pooled. Results: In Dataset A, Intuition-attributed decisions were more accurate than Guess-attributed decisions (mean difference = 0.146, 95% CI 0.086–0.205; t(21) = 5.12, p < 0.001; dz = 1.09); the adjusted trial-level association was OR = 1.56 (95% CI 1.25–1.96). Exploratory adjacent contrasts were also positive for Recollection versus Intuition and Rule knowledge versus Recollection. In Dataset B, cross-domain confidence correlations exceeded corresponding M-ratio correlations (delta r = 0.448–0.616; all Holm-adjusted p < 0.001). MAIA was associated with cardiac confidence (r = 0.262) but not cardiac accuracy (r = 0.019) or cleaned cardiac M-ratio (r = −0.001), and the correlation differences were significant. Conclusions: These analyses do not validate a unified construct or mechanism of intuition. They provide complementary constraints on its measurement: a reported hunch was not equivalent to a guess in Dataset A, while confidence was not equivalent to objective performance or metacognitive efficiency in Dataset B. Because intuition, confidence, performance, and metacognitive efficiency were not measured jointly in the same task, stronger claims about their within-person architecture require prospective within-task testing.
Keywords:
intuition
; metacognition
; confidence
; implicit learning
; interoception
; meta-d-prime
; source attribution
; open data
1. Introduction
Intuition is commonly used to describe judgments that seem informative before a person can fully articulate why the judgment should be trusted. This description raises at least two separable empirical questions. First, does a self-reported hunch contain task-relevant information beyond a self-reported guess? Second, should the subjective strength of a judgment, usually indexed by confidence, be treated as equivalent to objective performance or metacognitive efficiency? These questions are often discussed together in theories of intuition, yet they need not be answered by a single dataset or a single measurement scale.
Metacognition research provides a framework for separating some of these quantities. Mean confidence reflects metacognitive bias, or the general tendency to use higher or lower confidence ratings. Metacognitive sensitivity concerns how well confidence discriminates correct from incorrect decisions, whereas metacognitive efficiency expresses second-order sensitivity relative to first-order task sensitivity. Signal-detection measures such as meta-d-prime and M-ratio were developed because raw confidence and simple performance-confidence correlations can be influenced by response criteria and first-order performance [1,2,3,4]. Recent large-scale evaluation further shows that currently available metacognitive measures differ in psychometric behavior and can remain dependent on task performance [5].
A complementary literature addresses access to the basis of a judgment. Dienes and Scott distinguished structural knowledge from judgment knowledge and showed that a person can possess a sense that a judgment is correct while lacking reportable access to the structure supporting it [6]. Subsequent work using subjective source-attribution measures has separated guesses, intuitions, recollections, and rule-based responses [7,8]. In Monaghan et al.'s cross-situational learning paradigm, participants classified the basis of each training decision as Guess, Intuition, Recollection, or Rule knowledge [9]. The source study predicted and described an accuracy ordering across these categories and also compared broader implicit and explicit groupings [9]. We therefore do not present the categories or their descriptive ordering as new. Instead, we return to the archived trial-level data to ask a new inferential question: how large and robust are the direct within-participant differences between adjacent reported decision-basis categories when they are estimated under a common analysis plan, with multiplicity control and trial-level adjustment?
Interoception provides a separate opportunity to ask a different question relevant to intuition research: whether subjective confidence behaves like objective performance or model-based metacognitive efficiency. Garfinkel and colleagues argued that interoceptive accuracy, subjective sensibility, and awareness should not be collapsed into a single dimension [10]. More recent psychophysical tasks estimate cardiac and respiratory sensitivity while collecting trial-level confidence [11,12]. Banellis and colleagues reported weak correspondence across several objective and metacognitive dimensions while confidence showed stronger domain-general covariance [13]. The public multi_intero dataset therefore permits direct, same-participant tests of whether cross-domain confidence correlations exceed corresponding M-ratio correlations, and whether self-reported bodily awareness aligns more closely with confidence than with objective accuracy or metacognitive efficiency.
The present study does not combine these datasets as if they measured one latent construct. Rather, it asks two new, complementary questions of existing open data. Dataset A addresses the informational value associated with a reported hunch relative to a reported guess. Dataset B addresses whether confidence can be treated as interchangeable with objective performance or metacognitive efficiency. The analyses were formulated after the source datasets and publications were available and were not preregistered; they are therefore theory-guided and hypothesis-generating. H1 (primary) tested whether Intuition-attributed decisions were more accurate than Guess-attributed decisions in Dataset A. An exploratory extension formally tested adjacent contrasts across the four source-attribution categories, and H4 characterized change in explicit source attribution across learning blocks. In Dataset B, H2 tested whether cross-domain correlations in mean confidence exceeded corresponding correlations in M-ratio in identical participant sets, and H3 tested whether MAIA related more strongly to cardiac confidence than to cardiac accuracy or M-ratio. The contribution lies in deriving new inferential contrasts from the raw data rather than in treating previously published observations as novel.
2. Methods
2.1. Study Design, Transparency, and Data Provenance
We performed theory-guided secondary analyses of two publicly available behavioral datasets. No new participants were recruited and no new observations were generated. The central methodological principle was to ask new inferential questions of the original participant- and trial-level data rather than to reproduce the source articles' conclusions from published summary statistics. Variables and observations were reorganized only as required to define the prespecified reanalysis contrasts, and all effect estimates reported here were recalculated from the archived data. Because the source papers and source results were available before this reanalysis was conceived, the analyses are not presented as confirmatory replications.
The two datasets were analyzed separately. We did not pool participants, effect sizes, or variables across them because the paradigms, estimands, and awareness measures are not commensurate and the key constructs were not measured jointly in the same people or trials. Their relationship in the present study is therefore one of theoretical complementarity rather than statistical integration. We also performed a provenance audit of the earlier Cardioception repository associated with the Heart Rate Discrimination Task [11]. After normalizing identifier delimiters, all 206 unique participant identifiers represented in its session-1 psychophysics file were nested within the VMP1 cohort of the later multi_intero repository. We therefore did not count Cardioception as an independent replication sample.
2.2. Dataset A: Source Attribution During Cross-Situational Learning
We reanalyzed Experiment 2 from Monaghan, Schoetensack, and Rebuschat [9]. The archived public training file contained 22 university students and 6,336 trials (288 trials per participant across 12 blocks). In the source study, participants were native English speakers; 18 were female, and mean age was 20.23 years (SD 3.01) [9]. After each first-order decision, participants reported its basis using four categories. Guess was defined as a true guess comparable to a coin flip; Intuition as feeling that the decision was correct while being unable to explain why and following a hunch; Recollection as conscious memory for previously encountered material; and Rule knowledge as use of a conscious verbalizable rule [9].
We preserved the source labels and did not reclassify observations on the basis of performance. This avoids defining low-awareness observations post hoc using an outcome-related threshold, a practice that can create regression-to-the-mean artifacts [15]. The primary outcome was each participant's mean accuracy on intuition-attributed trials minus mean accuracy on guess-attributed trials.
The archived training file yielded at least one Guess trial for 22 participants, Intuition for 22, Recollection for 21, and Rule knowledge for 19. These raw-file denominators were used throughout the present reanalysis. Some category-specific degrees of freedom reported for descriptive one-sample tests in the source publication do not match these archived counts, although several published category means and t statistics are reproducible from the public file. To avoid propagating denominator ambiguity, all inferential statistics reported here were recalculated directly from the archived trial-level data.
2.3. Dataset B: Multimodal Interoception and Metacognition
We reanalyzed the public multi_intero dataset reported by Banellis et al. [13]. The source study included a Heart Rate Discrimination Task (HRDT) with cardiac-interoceptive and matched auditory-exteroceptive conditions and a Respiratory Resistance Sensitivity Task (RRST). Following source-study quality control, HRDT data were available for 513 participants and RRST data for 267; 241 participants contributed usable data across all three modalities [13]. The original study reported local ethics approval and informed consent.
Mean confidence was treated as a measure of metacognitive bias. Metacognitive efficiency was represented by M-ratio (meta-d-prime/d-prime) as supplied in the repository [2,3,13]. To match the repository processing logic used in the existing reanalysis package, negative M-ratio values were set to missing and values more than three median absolute deviations from the modality-specific median were removed. This yielded 417 cardiac, 485 auditory, and 257 respiratory cleaned M-ratio values. The MAIA analysis used the repository variable maia_full_mean, a participant-level summary derived from the 32-item MAIA instrument [14]. MAIA was treated as a self-report measure and was not reinterpreted as objective sensory accuracy.
2.4. Statistical Analysis
H1 (primary). For each Monaghan participant, accuracy was averaged separately for Intuition and Guess trials. We used a paired t test, reported the paired mean difference with a 95% t-based confidence interval and paired-samples Cohen dz, and calculated a participant-bootstrap confidence interval using 20,000 resamples. As a robustness analysis retaining all 6,336 trials, we fitted a binomial generalized estimating equation (GEE) with participant as the clustering variable and an exchangeable working correlation structure. The model included source-attribution category, centered training block, target action, and target picture and used bias-reduced sandwich standard errors. The GEE estimates associations between reported decision basis and accuracy; it is not interpreted as a causal effect of choosing an attribution label.
Exploratory adjacent-category analysis. We directly compared Intuition versus Guess, Recollection versus Intuition, and Rule knowledge versus Recollection using paired participant-level contrasts restricted to participants who contributed both categories. The three paired P values formed one family and were Holm-adjusted. The same three adjacent contrasts were then estimated as linear combinations from the categorical GEE described above and were Holm-adjusted as a separate robustness family. Because Rule knowledge was sparse for several participants, we repeated the paired contrasts after requiring at least 5, 10, and 20 trials in both categories of each contrast. As an additional sensitivity analysis only, the four source labels were coded 0-3 in their source-study order and entered as a single predictor in the adjusted GEE. This ordered model was not treated as primary because it imposes equal spacing between subjective categories that has not been psychometrically established.
H2. For each modality pair (cardiac-respiratory, cardiac-auditory, respiratory-auditory), analysis was restricted to identical participants with both confidence and cleaned M-ratio measures in both modalities. We calculated the Pearson correlation between mean-confidence measures and the Pearson correlation between M-ratio measures. The inferential quantity was their difference (delta r = r_confidence - r_M-ratio), estimated with 20,000 participant bootstrap resamples. The three H2 contrasts formed one Holm-adjusted family.
H3. We calculated Pearson correlations between MAIA full-scale mean and cardiac mean confidence, objective HRDT accuracy, and cleaned cardiac M-ratio, with Fisher-transformed 95% confidence intervals. We then formally compared the MAIA-confidence correlation with the MAIA-accuracy correlation and separately with the MAIA-M-ratio correlation in identical participants using participant-bootstrap delta-r intervals. The two H3 contrasts formed one Holm-adjusted family. As sensitivity analyses, we fitted HC3 heteroskedasticity-consistent linear models for standardized cardiac confidence and cleaned M-ratio, including objective accuracy and source cohort; the confidence model additionally tested a MAIA-by-cohort interaction.
H4 (exploratory). We modeled whether a trial was assigned an explicit source category (Recollection or Rule knowledge) rather than a non-explicit category (Guess or Intuition) as a function of training block using participant-clustered binomial GEE. This analysis characterizes change in reported decision basis over exposure and does not identify a discrete unconscious-to-conscious transition.
All tests were two-sided. Bootstrap analyses used a fixed seed (20260914). No prospective power calculation was performed because sample sizes were fixed by the archived datasets; emphasis was placed on effect estimates and confidence intervals. Analyses were performed in Python using pandas, NumPy, SciPy, and statsmodels.
Table 1.
Data sources, analytic roles, and denominator architecture.
| Dataset | Source sample used here | Operational construct | Role in present study |
| Monaghan et al., Experiment 2 | N=22; 6,336 training trials; 12 blocks | Trial-level source report (Guess, Intuition, Recollection, Rule knowledge) and objective accuracy | Dataset A: primary Intuition-vs-Guess reanalysis; exploratory adjacent source-attribution contrasts and temporal source-report analysis |
| Banellis et al. multi_intero - HRDT | Post-QC N=513 | Cardiac and auditory confidence; HRDT accuracy; cardiac/auditory M-ratio | Dataset B: confidence-versus-M-ratio and MAIA association contrasts; does not measure intuition directly |
| Banellis et al. multi_intero - RRST | Post-QC N=267; all three modalities N=241 | Respiratory confidence and M-ratio | Dataset B: cross-domain confidence-versus-M-ratio contrasts; does not measure intuition directly |
| Cardioception repository | Not counted as independent evidence | Earlier HRDT dataset; session-1 identifiers nested within VMP1 | Provenance/overlap audit only |
Note. Denominators vary by measure because source tasks, availability, and M-ratio quality-control procedures yield different analytic samples. The datasets were not pooled, and no cross-dataset coefficient is interpreted as a within-person association.
3. Results
3.1. H1: Intuition-Attributed Decisions Versus Guesses
All 22 participants contributed both Guess and Intuition trials. Participant-level accuracy averaged 0.528 (bootstrap 95% CI 0.481-0.569) for Guess and 0.673 (0.604-0.739) for Intuition. The paired difference was 0.146 accuracy units (14.6 percentage points; 95% CI 0.086-0.205), t(21)=5.12, p<0.001, dz=1.09. The participant-bootstrap interval was 0.090-0.198.
The trial-level robustness model was consistent with the participant-level result. After adjustment for training block, target action, and target picture, Intuition-attributed trials had higher odds of a correct response than Guess-attributed trials (OR=1.56, 95% CI 1.25-1.96, p<0.001; 6,336 trials clustered within 22 participants).
3.2. Exploratory Ordered Pattern Across Reported Decision-Basis Categories
The four raw-file category means reproduced the source-study descriptive ordering: Guess 0.528 (N=22), Intuition 0.673 (N=22), Recollection 0.794 (N=21), and Rule knowledge 0.928 (N=19). Because this ordering was predicted and described in the source publication [9], it is not presented here as a novel discovery. The present contribution is the direct formal comparison of adjacent categories using a common reanalysis framework.
All three paired adjacent contrasts were positive. Intuition exceeded Guess by 0.146 (95% CI 0.086-0.205; dz=1.09; Holm-adjusted p<0.001), Recollection exceeded Intuition by 0.135 (95% CI 0.068-0.202; dz=0.92; Holm-adjusted p<0.001), and Rule knowledge exceeded Recollection by 0.147 (95% CI 0.029-0.264; dz=0.60; Holm-adjusted p=0.017). Participant-bootstrap intervals were 0.090-0.198, 0.077-0.198, and 0.031-0.243, respectively.
The adjusted trial-level model yielded the same ordering. The adjacent odds ratios were 1.56 (95% CI 1.25-1.96) for Intuition versus Guess, 1.85 (1.44-2.38) for Recollection versus Intuition, and 6.42 (3.53-11.68) for Rule knowledge versus Recollection; all three remained significant after Holm correction. Sparse-category sensitivity analyses requiring at least 5, 10, or 20 observations in both categories preserved positive and statistically supported adjacent differences. For example, with a minimum of 10 trials per category, the differences were 0.147, 0.143, and 0.180, respectively, with all Holm-adjusted p values <0.001.
In the additional ordered-code sensitivity model, each one-step increase in the prespecified source ordering was associated with higher adjusted odds of a correct response (OR per step=1.88, 95% CI 1.71-2.07, p<0.001). Because this model assumes equal spacing between subjective categories, it is reported only as a compact sensitivity summary rather than as evidence for a psychometric continuum or graded consciousness.
Figure 1.
Decision accuracy by reported decision basis in Monaghan et al. Experiment 2. Points show participant-level mean accuracy and error bars show 95% participant-bootstrap confidence intervals (20,000 resamples). The dashed horizontal line marks chance accuracy (0.50). Category labels preserve the source-study instructions.
Figure 1.
Decision accuracy by reported decision basis in Monaghan et al. Experiment 2. Points show participant-level mean accuracy and error bars show 95% participant-bootstrap confidence intervals (20,000 resamples). The dashed horizontal line marks chance accuracy (0.50). Category labels preserve the source-study instructions.

Table 2.
Adjacent source-attribution contrasts in Monaghan Experiment 2.
| Contrast | N | Mean difference | 95% CI | dz | Holm-adjusted p | Adjusted GEE OR (95% CI) |
| Intuition - Guess | 22 | 0.146 | 0.086 to 0.205 | 1.09 | <0.001 | 1.56 (1.25-1.96) |
| Recollection - Intuition | 21 | 0.135 | 0.068 to 0.202 | 0.92 | <0.001 | 1.85 (1.44-2.38) |
| Rule knowledge - Recollection | 19 | 0.147 | 0.029 to 0.264 | 0.60 | 0.017 | 6.42 (3.53-11.68) |
Note. Participant-level contrasts are paired and use only participants contributing both adjacent categories. GEE contrasts are linear combinations from a single categorical model adjusted for centered block, target action, and target picture with participant clustering and bias-reduced sandwich standard errors. Participant-level and GEE contrast families were Holm-adjusted separately.
3.3. H2: Cross-Domain Confidence Versus Metacognitive Efficiency
The source publication had already reported that mean confidence showed more cross-domain correspondence than M-ratio [13]. Our reanalysis tested whether those correlation strengths differed when calculated in identical participants. In the 173 participants contributing all measures for the cardiac-respiratory comparison, confidence correlated across modalities at r=0.486 whereas M-ratio correlated at r=0.038; delta r=0.448 (95% bootstrap CI 0.254-0.640). For cardiac-auditory data (N=398), the corresponding correlations were r=0.584 and r=0.088; delta r=0.496 (0.376-0.612). For respiratory-auditory data (N=225), they were r=0.642 and r=0.025; delta r=0.616 (0.468-0.767). All three delta-r tests remained p<0.001 after Holm correction.
These results support a difference in correlation magnitude rather than merely a contrast between statistically significant and non-significant individual correlations.
Figure 2.
Formal within-sample contrasts of cross-domain correlation strength. Each point is delta r = correlation in mean confidence minus correlation in cleaned M-ratio for participants with all four measures required by that modality-pair comparison. Error bars are 95% participant-bootstrap confidence intervals (20,000 resamples).
Figure 2.
Formal within-sample contrasts of cross-domain correlation strength. Each point is delta r = correlation in mean confidence minus correlation in cleaned M-ratio for participants with all four measures required by that modality-pair comparison. Error bars are 95% participant-bootstrap confidence intervals (20,000 resamples).

3.4. H3: Self-Reported Interoceptive Awareness Aligns More with Confidence Than with Objective or Efficiency Measures
MAIA full-scale mean was available for 551 participants and for 501 participants with cardiac mean-confidence and accuracy measures. MAIA correlated with cardiac mean confidence (r=0.262, 95% CI 0.178-0.342, p<0.001) but not with objective cardiac accuracy (r=0.019, 95% CI -0.069 to 0.106, p=0.672). The same-participant difference between these correlations was delta r=0.243 (95% bootstrap CI 0.130-0.352, Holm-adjusted p<0.001).
Among the 407 participants with MAIA and cleaned cardiac M-ratio, the MAIA-M-ratio association was essentially zero (r=-0.001, 95% CI -0.098 to 0.096, p=0.988). The MAIA-confidence correlation was stronger by delta r=0.286 (95% bootstrap CI 0.154-0.418, Holm-adjusted p<0.001). In an HC3 model adjusting standardized cardiac confidence for objective accuracy and cohort, standardized MAIA remained associated with confidence (beta=0.261, 95% CI 0.170-0.352, p<0.001), and there was no evidence of a MAIA-by-cohort interaction (p=0.679). Conversely, the standardized MAIA coefficient in the cleaned M-ratio sensitivity model was 0.000 (95% CI -0.095 to 0.096, p=0.998).
3.5. H4: Exploratory Change in Reported Decision Basis Across Learning
Reported decision basis shifted across the 12 Monaghan training blocks. In block 1, 43.4% of trials were labeled Guess and 38.1% Intuition, whereas by block 12 the corresponding proportions were 10.2% and 17.6%. Recollection and Rule knowledge together increased from 18.6% to 72.1%. A participant-clustered GEE estimated a 23% increase in the odds of an explicit source attribution for each successive block (OR per block=1.23, 95% CI 1.15-1.31, p<0.001). Because block simultaneously indexes accumulated exposure and learning, this temporal association does not identify the cognitive mechanism producing the shift.
Table 3.
Main inferential results outside the adjacent-category analysis.
| Analysis | N | Effect estimate | 95% CI | Adjusted p | Interpretive boundary |
| H1 primary: Intuition - Guess accuracy | 22 | Mean difference 0.146; dz=1.09 | 0.086 to 0.205 | <0.001 | Reported decision basis differs in associated accuracy; not proof of unconscious knowledge |
| H1 robustness: trial-level GEE | 6,336 trials / 22 clusters | OR 1.56 | 1.25 to 1.96 | <0.001 | Adjusted within-task association |
| H2: Cardiac-respiratory delta r | 173 | 0.448 | 0.254 to 0.640 | <0.001* | Confidence correlation minus M-ratio correlation |
| H2: Cardiac-auditory delta r | 398 | 0.496 | 0.376 to 0.612 | <0.001* | Same-participant correlation contrast |
| H2: Respiratory-auditory delta r | 225 | 0.616 | 0.468 to 0.767 | <0.001* | Same-participant correlation contrast |
| H3: MAIA-confidence minus MAIA-accuracy delta r | 501 | 0.243 | 0.130 to 0.352 | <0.001* | Self-report aligns more with confidence than objective accuracy |
| H3: MAIA-confidence minus MAIA-M-ratio delta r | 407 | 0.286 | 0.154 to 0.418 | <0.001* | Self-report aligns more with confidence than model-based efficiency |
| H4: explicit attribution per block | 6,336 trials / 22 clusters | OR 1.23 per block | 1.15 to 1.31 | <0.001 | Temporal association; mechanism not identified |
Note. *Holm-adjusted within the H2 or H3 contrast family. CI = confidence interval; GEE = generalized estimating equation; MAIA = Multidimensional Assessment of Interoceptive Awareness; M-ratio = meta-d-prime/d-prime.
4. Discussion
The present secondary analyses support two deliberately narrow conclusions rather than one unified dissociation claim. In Dataset A, decisions labeled as Intuition were more accurate than decisions labeled as true Guess, and exploratory formal contrasts further quantified adjacent differences across the source-attribution categories. In Dataset B, mean confidence generalized across sensory domains more strongly than M-ratio, while self-reported interoceptive awareness aligned with confidence but not with objective cardiac accuracy or cleaned cardiac M-ratio. These results constrain different parts of the measurement problem: a reported hunch is not behaviorally equivalent to guessing, and confidence is not equivalent to objective performance or metacognitive efficiency. Because these constructs were not measured jointly in the same task, the analyses do not demonstrate a single within-person architecture of intuition, and they do not show that the relevant information was processed unconsciously.
The present study should be understood as a secondary reanalysis that asks new questions of existing raw data, not as a claim that the source studies failed to analyze their own primary aims. Monaghan et al. introduced the source-attribution categories and reported the descriptive accuracy ordering in their learning paradigm [9]. Banellis et al. characterized multiple dimensions of interoception and reported stronger domain-generality for confidence than for several objective or metacognitive measures [13]. Those observations provide the empirical setting for the present work but are not claimed as novel findings.
The added value of the current reanalysis lies in the inferential contrasts constructed from the archived participant- and trial-level observations. In Dataset A, the Intuition-versus-Guess difference is estimated directly within participant with an effect size and bootstrap interval and is then re-estimated across all 6,336 trials after adjustment for block and stimulus identities. The three adjacent source-attribution contrasts are tested under a common framework with multiplicity control and sparse-category sensitivity analyses. In Dataset B, confidence-versus-M-ratio differences are tested as direct correlation contrasts in identical participant sets rather than inferred from whether separate correlations reach significance. The provenance audit further prevents a nested cardiac dataset from being misinterpreted as independent replication evidence.
Dataset A provides the most direct observation concerning intuition. Under the source instructions, Guess denoted a decision equivalent to a coin flip, whereas Intuition denoted a hunch that felt correct despite an inability to explain why [9]. The 14.6-percentage-point within-participant accuracy advantage therefore shows that these two reports were associated with different amounts of task-relevant information in this paradigm. Put simply, a reported hunch was not behaviorally equivalent to a reported guess. This result is compatible with the distinction between judgment knowledge and reportable structural knowledge [6] and with work showing that intuitions can emerge before fully explicit knowledge [16]. It does not identify the source of the information supporting the hunch, and it does not establish that the relevant information was unconscious. Partial, weak, difficult-to-verbalize, or strategically inaccessible knowledge remain viable explanations.
The exploratory adjacent contrasts provide a more detailed description of Dataset A. Accuracy increased across the source-study ordering from Guess to Intuition to Recollection to Rule knowledge. Because this ordering was already predicted and described in the source publication, our contribution is the formal quantification of adjacent differences and their robustness, not discovery of the ordering itself. The pattern should be described as an ordered association with reported decision basis, not as a measured continuum of consciousness. The category labels are not known to have equal psychological distances, which is why the 0-3 ordered GEE was retained only as a sensitivity analysis. Likewise, the large Rule-versus-Recollection odds ratio should not be interpreted as a universal magnitude of explicit-access gain; Rule trials became common later in learning, when overall task performance was also higher, and residual time-varying confounding cannot be eliminated by block adjustment alone.
Dataset B addresses a different question. It does not contain the Guess/Intuition/Recollection/Rule source-attribution measure and therefore does not test intuition directly. Instead, it tests whether confidence can be treated as interchangeable with objective performance or metacognitive efficiency. The direct delta-r analyses show that cross-domain confidence correlations were materially stronger than corresponding M-ratio correlations in the same participants. MAIA showed a parallel separation: self-reported interoceptive awareness was associated with how confident participants tended to be, but not with objective HRDT accuracy or cleaned cardiac M-ratio. This does not imply that MAIA or confidence is inferior as a measure; it indicates that self-report, confidence disposition, objective performance, and model-based efficiency capture non-equivalent aspects of behavior and experience [5,10,14].
Taken together, the datasets provide convergent constraints rather than a unified validation of intuition. Dataset A shows that a reported hunch can carry more task-relevant information than a reported guess. Dataset B shows that confidence should not automatically be used as a proxy for objective performance or metacognitive efficiency. These findings motivate, but do not themselves prove, a cautious working description of intuitive judgment as a decision that may be informed by task-relevant information while the basis of that information is not fully available for explicit report. The phrase 'not fully available' is intentionally weaker than 'unconscious' and allows graded, incomplete, unstable, or difficult-to-verbalize access. Because the relevant constructs were not measured jointly, the present data cannot determine the within-person relationship between intuition and confidence.
Strong alternative interpretations remain. Source-attribution categories are subjective reports and can be influenced by instructions, response criteria, memory for the decision process, and demand characteristics. Newell and Shanks argued that claims about unconscious influence often exceed what awareness measures can establish [18], and Shanks demonstrated how post-hoc selection on low-awareness observations can produce misleading dissociations [15]. The present analyses avoid selecting trials because an awareness measure happened to be non-significant, but they cannot establish what information was consciously represented at the moment of choice or where the information supporting an intuitive response originated. Likewise, M-ratio is a model-based measure of metacognitive efficiency rather than a direct readout of consciousness, and recent psychometric work shows that metacognitive measures retain task dependencies [5].
No neural variable was analyzed, and no dataset used here measured intuition, confidence, objective performance, metacognitive efficiency, and source accessibility within the same task. The present study therefore cannot adjudicate neural mechanisms, determine whether intuition and confidence have a linear or nonlinear relationship, or identify the origin of the information expressed in an intuitive judgment. A decisive next study would measure first-order discrimination, trial-level confidence, and trial-level source accessibility prospectively within the same preregistered task, with enough observations in each source category to estimate metacognitive efficiency. Such a design could directly test whether informative hunches occupy a distinct region of the performance-confidence-accessibility space and could subsequently be extended with time-resolved neural measurements.
4.1. Strengths
The study has several practical strengths. It returns to participant- and trial-level open data and recalculates the inferential quantities required by research questions that were not the primary targets of the source publications. The primary Dataset A comparison is within participant and is supported by a clustered trial-level model. Exploratory adjacent contrasts were subjected to multiplicity control and sparse-trial sensitivity analyses. Dataset B correlation differences were tested directly in identical participant sets. Finally, explicit separation of the two datasets prevents complementary evidence from being overstated as a joint within-person dissociation, while the data-provenance audit reduces the risk of pseudo-replication across overlapping repositories.
4.2. Limitations
First, this is a non-preregistered secondary analysis developed after the source studies were published. The inferential contrasts are new, but they were specified with knowledge of the datasets and literature; even small P values should therefore be interpreted as hypothesis-generating rather than as confirmatory evidence for a prospectively specified theory.
Second, Monaghan Experiment 2 is small at the participant level (N=22), despite its dense repeated-measures structure. The Intuition-versus-Guess effect is large within this task, but its generalizability to other populations, tasks, or definitions of intuition is unknown. The four source categories are also unevenly represented across participants and learning blocks, especially for Rule knowledge early in training.
Third, the source-attribution categories are subjective. Above-chance or higher accuracy on Intuition trials indicates that the label carries information relative to Guess within this task, but it does not prove an absence of conscious structural knowledge. Similarly, an ordered association among the four labels does not establish an interval scale of accessibility or consciousness.
Fourth, the two datasets answer different questions and do not provide a common within-participant test. Dataset A measures reported decision basis and objective accuracy but not the confidence/M-ratio architecture used in Dataset B. Dataset B measures confidence, objective performance, M-ratio, and self-reported interoceptive awareness but not the source-attribution categories used to define Intuition in Dataset A. Accordingly, the present study supports complementary measurement constraints, not a unified latent model of intuition.
Fifth, individual M-ratio estimates and self-report scales have measurement limitations. M-ratio can be unstable or dependent on first-order performance and confidence behavior, while MAIA indexes subjective interoceptive dimensions rather than objective sensory accuracy [5,14]. Cleaning rules and sensitivity models reduce but do not remove these limitations.
Sixth, the archived Monaghan file yields category denominators that differ from some degrees of freedom reported in the source article's descriptive one-sample tests. We therefore rely on transparent raw-file denominators and fully recomputed statistics. This discrepancy does not alter the primary within-participant Intuition-versus-Guess result or the trial-level GEE, but it reinforces the need to make the reanalysis code and denominator logic publicly auditable.
5. Conclusion
This study asked new inferential questions of two existing open datasets by returning to their participant- and trial-level observations and recalculating the relevant contrasts. In Dataset A, a reported hunch was more accurate than a reported guess, showing that the two subjective decision-basis reports were not behaviorally equivalent. In Dataset B, confidence generalized across domains more strongly than M-ratio, and self-reported interoceptive awareness tracked confidence more than objective cardiac performance or metacognitive efficiency. These findings should not be combined into a claim that intuition, confidence, performance, and metacognitive efficiency have been jointly dissociated, because they were not measured together. Instead, they provide complementary constraints for future theories: intuition should not be equated with guessing, and confidence should not be assumed to index either objective performance or metacognitive efficiency. The data do not identify where intuitive information comes from, demonstrate unconscious processing, establish a neural mechanism, or define intuition universally. A preregistered within-task study measuring performance, confidence, metacognitive efficiency, and source accessibility together is required for direct validation of the proposed architecture.
Ethics
This manuscript reports secondary analyses of publicly available, deidentified datasets and involved no new participant recruitment or contact. The source studies report their original ethics approval and informed-consent procedures [9,13]. Any additional institutional determination required by the target journal or the authors' local regulations should be confirmed before submission.
Data and Code Availability
The Monaghan et al. source data and code are publicly available through OSF project 2xzye (https://osf.io/2xzye/). The Banellis et al. source data and analysis repository are publicly available at https://github.com/embodied-computation-group/multi_intero. The complete reanalysis script, fixed random seed, software-version record, and machine-readable results used for the present manuscript should be deposited in a persistent public repository before submission.
Supplementary Materials
The following supporting information can be downloaded at the website of this paper posted on Preprints.org.
References
- Fleming, S.M.; Lau, H.C. How to measure metacognition. Front Hum. Neurosci. 2014, 8, 443. [Google Scholar] [CrossRef]
- Maniscalco, B.; Lau, H. A signal detection theoretic approach for estimating metacognitive sensitivity from confidence ratings. Conscious Cogn. 2012, 21(1), 422–430. [Google Scholar] [CrossRef]
- Fleming, S.M. HMeta-d: hierarchical Bayesian estimation of metacognitive efficiency from confidence ratings. Neurosci. Conscious 2017, 2017(1), nix007. [Google Scholar] [CrossRef]
- Rahnev, D.; Fleming, S.M. How experimental procedures influence estimates of metacognitive ability. Neurosci. Conscious 2019, 2019(1), niz009. [Google Scholar] [CrossRef]
- Rahnev, D. A comprehensive assessment of current methods for measuring metacognition. Nat. Commun. 2025, 16, 701. [Google Scholar] [CrossRef]
- Dienes, Z.; Scott, R. Measuring unconscious knowledge: distinguishing structural knowledge and judgment knowledge. Psychol. Res. 2005, 69(5-6), 338–351. [Google Scholar] [CrossRef]
- Mealor, A.D.; Dienes, Z. The speed of metacognition: taking time to get to know one's structural knowledge. Conscious Cogn. 2013, 22(1), 123–136. [Google Scholar] [CrossRef]
- Rebuschat, P. Measuring implicit and explicit knowledge in second language research. Lang. Learn. 2013, 63(3), 595–626. [Google Scholar] [CrossRef]
- Monaghan, P.; Schoetensack, C.; Rebuschat, P. A single paradigm for implicit and statistical learning. Top. Cogn. Sci. 2019, 11(3), 536–554. [Google Scholar] [CrossRef]
- Garfinkel, S.N.; Seth, A.K.; Barrett, A.B.; Suzuki, K.; Critchley, H.D. Knowing your own heart: distinguishing interoceptive accuracy from interoceptive awareness. Biol. Psychol. 2015, 104, 65–74. [Google Scholar] [CrossRef]
- Legrand, N.; Nikolova, N.; Correa, C.; et al. The heart rate discrimination task: a psychophysical method to estimate the accuracy and precision of interoceptive beliefs. Biol. Psychol. 2022, 168, 108239. [Google Scholar] [CrossRef]
- Nikolova, N.; Harrison, O.; Toohey, S.; et al. The respiratory resistance sensitivity task: an automated method for quantifying respiratory interoception and metacognition. Biol. Psychol. 2022, 170, 108325. [Google Scholar] [CrossRef]
- Banellis, L.; Nikolova, N.; Ehmsen, J.F.; et al. Interoceptive ability is uncorrelated across respiratory and cardiac axes in a large scale psychophysical study. Commun. Psychol. 2026, 4, 43. [Google Scholar] [CrossRef]
- Mehling, W.E.; Price, C.; Daubenmier, J.J.; Acree, M.; Bartmess, E.; Stewart, A. The Multidimensional Assessment of Interoceptive Awareness (MAIA). PLoS ONE 2012, 7(11), e48230. [Google Scholar] [CrossRef]
- Shanks, D.R. Regressive research: the pitfalls of post hoc data selection in the study of unconscious mental processes. Psychon. Bull. Rev. 2017, 24(3), 752–775. [Google Scholar] [CrossRef]
- Weinberger, A.B.; Green, A.E. Dynamic development of intuitions and explicit knowledge during implicit learning. Cognition 2022, 222, 105008. [Google Scholar] [CrossRef]
- Jurchis, R.; Preda, A.; Costea, A.; Skora, L. Non-conscious knowledge of complex regularities supports instrumental conditioning: a registered report. Cortex 2026, 203, 187–210. [Google Scholar] [CrossRef]
- Newell, B.R.; Shanks, D.R. Unconscious influences on decision making: a critical review. Behav. Brain Sci. 2014, 37(1), 1–19. [Google Scholar] [CrossRef]
- Dienes, Z. Subjective measures of unconscious knowledge. Prog. Brain Res. 2008, 168, 49–64. [Google Scholar] [CrossRef]
- Rebuschat, P.; Williams, J.N. Implicit and explicit knowledge in second language acquisition. Appl. Psycholinguist. 2012, 33(4), 829–856. [Google Scholar] [CrossRef]
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the author. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.