Preprint
Article

This version is not peer-reviewed.

Eye-Tracking Integration in Computerized Social-Cognition Tasks: A Feasibility and Signal-Characterization Pilot Study in Healthy Adults

Submitted:

27 August 2026

Posted:

31 August 2026

You are already at the latest version

Abstract
Background: Social-cognition assessment commonly relies on endpoint behavioral scores, which provide limited information about how socially relevant visual information is sampled during task performance. Eye-tracking may complement these measures, but the interpretation of gaze-derived metrics depends strongly on task structure, stimulus format, and spatial resolution. Objective: This pilot study examined the feasibility and task-related informativeness of integrating eye-tracking into two computerized social-cognition tasks with different perceptual demands: a full-face facial emotion-recognition task and the Reading the Mind in the Eyes Test. Methods: Nineteen healthy adults completed the Spanish Test de Reconocimiento Emocional en Caras (TREC) and the Reading the Mind in the Eyes Test (RMET) while gaze behavior was recorded with a screen-based Tobii Pro Nano eye tracker. Fixation count, cumulative fixation duration, and reaction time were extracted. For the TREC, gaze allocation was additionally characterized across predefined facial areas of interest, including the eyes, nose, mouth, and facial hemifields. Analyses were hypothesis-generating; p-values were interpreted as descriptive indices rather than confirmatory evidence. Results: Eye-tracking could be implemented across both computerized tasks and provided task-dependent gaze information. In the full-face TREC, visual sampling was mainly distributed across the eyes and nose, with lower allocation to the mouth, suggesting that full-face stimuli may be useful for spatially resolved gaze characterization. In contrast, the RMET elicited higher fixation count, longer cumulative fixation duration, and longer response time, but its restricted eye-region format mainly allowed global inspection metrics rather than facial-region comparisons. Performance-related gaze patterns were interpreted in light of the restricted variance and near-ceiling distribution of TREC scores. Conclusions: This pilot study is consistent with the feasibility of embedding eye-tracking into computerized social-cognition assessment and illustrates how the interpretability of gaze-derived measures depends on task format. Full-face paradigms may be more informative for area-of-interest analyses of facial sampling, whereas eye-region paradigms mainly inform inspection time and response dynamics under constrained visual input. These findings should be considered exploratory and do not establish stable gaze profiles, mechanisms, or clinical markers. Instead, they help define a methodological basis for larger, preregistered studies testing whether task-sensitive gaze-derived measures can contribute to the characterization of social-cognitive processing.
Keywords: 
;  ;  ;  ;  ;  ;  ;  ;  ;  
Subject: 
Social Sciences  -   Psychology

1. Introduction

Social cognition refers to the neurocognitive capacity to extract meaning from other agents by integrating perceptual, attentional, affective, and inferential processes. It comprises partially dissociable domains—including social perception, facial emotion recognition, affective processing, theory of mind, empathy, attributional style, and social knowledge—that allow individuals to detect socially relevant cues, evaluate their significance, and adjust behavior accordingly(Adolphs, 2009; Arioli et al., 2018; Bölte, 2025). Faces and eyes are privileged social signals through which individuals infer emotion, intention, trust, threat, affiliation, and mental states(Haxby et al., 2000; Pitcher, 2025). Although human social cognition is expanded by language, symbolic representation, cultural learning, and explicit mental-state reasoning, it builds on conserved social-attentional mechanisms shared across social species, especially sensitivity to faces, gaze direction, emotional displays, and approach–avoidance cues(Arioli et al., 2018; Frith & Frith, 2007; Rosati, 2026; Shafiei et al., 2026). Neurobiologically, these functions depend on distributed perceptual, salience-related, limbic, and associative cortical systems that support facial encoding, affective valuation, contextual integration, and social prediction(Adolphs, 2009; Cavallo et al., 2026). Across development, these systems are progressively calibrated through experience(Bohn et al., 2026), and cross-cultural evidence indicates that face-scanning strategies may vary in the relative use of eye, mouth, and central facial information(Caldara, 2017; Hessels et al., 2025). Thus, social-cognitive performance is not only a matter of giving a correct or incorrect response; the same behavioral outcome may arise from different configurations of perceptual sensitivity, attentional allocation, affective salience, and inferential processing(Alkan, 2025; Arioli et al., 2018). This makes social cognition clinically important and methodologically complex: to understand performance, it is necessary to examine not only what response is given, but how social information is sampled before that response emerges(Alkan, 2025; De Biase et al., 2026; Hooge et al., 2026).
Despite this complexity, social-cognition assessment remains largely anchored in endpoint behavioral metrics, especially accuracy scores from facial emotion-recognition tasks and the Reading the Mind in the Eyes Test(Hafner et al., 2026; Higgins et al., 2025). These instruments have been essential for documenting impairments across schizophrenia-spectrum disorders, autism, mood disorders, neurodegenerative diseases, and other conditions affecting social functioning(Bölte, 2025; Cavallo et al., 2026; Corbera et al., 2025; Dodich et al., 2025). However, accuracy scores provide limited mechanistic information. A correct or incorrect response compresses several processing stages—perceptual encoding, attentional selection, affective evaluation, evidence accumulation, and inferential decision-making—into a single outcome(Arioli et al., 2018). As a result, equivalent performance may conceal different strategies, and similar impairment may reflect disruptions at different levels of the social-cognitive system(Alkan, 2025; Bölte, 2025). This limitation is especially relevant in heterogeneous or compensatory contexts. Some individuals may achieve preserved performance through slower, broader, or more effortful exploration; others may fail because they do not sample the most informative facial regions, misattribute affective salience, or struggle to integrate perceptual evidence into social meaning(Yang et al., 2024). Task structure further complicates interpretation: full-face paradigms, eye-region tasks, static images, dynamic stimuli, and different response formats impose distinct perceptual and inferential constraints(Hafner et al., 2026; Higgins et al., 2025). Therefore, performance cannot be treated as a transparent measure of social cognition without considering how the task organizes information sampling. This has motivated growing interest in process-level measures that can characterize how socially meaningful evidence is acquired before an explicit response is made(Alkan, 2025).
Eye-tracking may help address this gap because it captures visual sampling during task execution(Rahal & Fiedler, 2019). Fixation location, fixation duration, scan patterns, and response timing can provide indirect information about how attention is allocated across socially informative facial features and whether performance appears to rely on focused extraction of relevant cues, diffuse exploration, or compensatory effort(Alkan, 2025). Its value, however, depends on task design. Full-face emotion-recognition tasks provide a spatially rich field in which gaze can be distributed across the eyes, nose, mouth, and facial laterality; eye-region tasks, by contrast, restrict visual exploration while increasing the inferential demand placed on limited information(Hafner et al., 2026; Haxby et al., 2000). The present pilot study examines the feasibility and preliminary informativeness of integrating eye-tracking into computerized social-cognition assessment in healthy adults, using a full-face facial emotion-recognition task and the Reading the Mind in the Eyes Test. Rather than establishing diagnostic thresholds or normative standards, it adopts a signal-characterization approach: determining whether gaze-derived metrics provide preliminary, task-dependent patterns under controlled conditions and whether task format shapes the granularity of information obtained. Characterizing such preliminary patterns in non-clinical participants may represent a useful step before testing whether clinical deviations reflect altered perceptual encoding, attentional prioritization, affective salience, processing efficiency, or compensatory visual strategies(Corbera et al., 2025; Dodich et al., 2025).

2. Methods

2.1. Participants

The pilot sample was recruited by the IDIVAL Mental Illness Research Department’s neuropsychology lab at the Marqués de Valdecilla Research Institute (IDIVAL), Santander, Spain, as part of the Spanish National Research Project PI18/00212, “Influence of attachment style on social cognition and cognitive biases in individuals with first-episode psychosis, chronic schizophrenia, and healthy controls.” The final sample comprised 19 healthy adults (11 males and 8 females), all of whom received detailed information about the study and provided written informed consent. The study protocol was approved by CEIC-Cantabria (internal code: 2019.024) and conducted in accordance with the principles outlined in the Declaration of Helsinki (1975, revised in 2013). Inclusion criteria were:
(i)
Absence of any mental disorder and related treatment during the previous year, assessed using the Composite International Diagnostic Interview (CIDI).
(ii)
Age between 18 and 65 years.
(iii)
Absence of traumatic brain injury, dementia, or intellectual disability (IQ < 70).
(iv)
Absence of substance abuse.

2.2. Social Cognition Tasks

Participants completed two computerized social-cognition tasks: the Spanish version of the Test de Reconocimiento Emocional en Caras (TREC) and the Spanish adaptation of the Reading the Mind in the Eyes Test (RMET). The two tasks were selected because they impose different perceptual constraints on social-information sampling: the TREC presents full-face emotional stimuli, allowing gaze allocation to be examined across multiple facial regions, whereas the RMET restricts visual input to the eye region and therefore places greater inferential demand on limited perceptual information.
(i)
Facial Emotion Recognition Task (TREC)(Baron-Cohen et al., 1997; Huerta-Ramos et al., 2021). The Spanish version of the TREC was administered as a computerized full-face facial emotion-recognition task. The task consists of 20 items, each corresponding to a photograph of the face of a Caucasian woman. Half of the items depict emotional expressions of negative valence and half depict emotional expressions of positive valence. For each item, participants were required to identify the emotion expressed by the face by selecting one of two response options. In the present study, the TREC was used to assess facial emotion recognition under spatially rich visual conditions, allowing eye-tracking analyses of region-specific gaze allocation across the eyes, nose, mouth, and facial hemifields.
(ii)
Reading the Mind in the Eyes Test (RMET)(Baron-Cohen et al., 2001; Fernández-Abascal et al., 2013). The Spanish adaptation of the RMET was also administered in computerized format. The task consists of 36 black-and-white images originally taken from magazines, showing only the eye and eyebrow region of different Caucasian individuals, with an equal number of male and female faces. Each item depicts a complex mental or emotional state. Participants were required to select, from four response alternatives, the word that best described what the person in the image was thinking or feeling. In contrast to the TREC, the RMET restricts visual information to the eye region, preventing equivalent comparisons across full-face areas of interest and allowing assessment of global visual inspection during mental-state attribution under constrained perceptual input.

2.3. Eye-Tracking Measures and Data-Quality Considerations

Gaze behavior was recorded during both computerized social-cognition tasks using a screen-based Tobii Pro Nano eye tracker operating at a sampling frequency of 60 Hz. The eye tracker was integrated into custom software that synchronized stimulus presentation, response registration, gaze-data acquisition, and real-time supervision of the acquisition stream. Stimuli were presented on a 24-inch display, and participants were seated at an approximate viewing distance of 60–65 cm from the screen. Testing was conducted under stable, soft ambient lighting conditions to support tracking stability and reduce avoidable signal variability.
Before administration of the computerized social-cognition tasks, each participant completed a standard 5-point calibration and validation sequence using the manufacturer’s native SDK interface. During calibration, participants were instructed to follow a calibration target presented at five distinct screen locations. Immediately after calibration, the validation layout was visually checked by the supervisor through the acquisition interface. When visible geometric displacement, unstable gaze mapping, or signal loss was detected, the calibration procedure was repeated before proceeding with task administration. Thus, calibration quality was addressed procedurally through real-time visual validation and recalibration when needed.
During task administration, the supervisor interface allowed real-time monitoring of gaze-data quality and acquisition stability. The system monitored data-quality information during acquisition, including indicators related to gaze accuracy, sample precision, tracking ratio, and signal loss. However, quantitative participant-level calibration accuracy logs were not permanently exported as standalone data files. Therefore, calibration quality was available for procedural supervision during data collection, but not for retrospective participant-level numerical reporting.
Raw gaze samples were processed using a velocity-threshold identification filter (I-VT). A discrete fixation was registered when gaze remained spatially stable according to the velocity-threshold procedure for a minimum duration of 150 ms. The following eye-tracking-derived variables were extracted:
(i)
Fixation count: defined as the number of discrete fixations registered during a task or within a given area of interest, indexing how often visual attention was allocated to that stimulus region.
(ii)
Cumulative fixation duration: defined as the total time, in milliseconds, spent fixating on a task or area of interest, indexing the amount of visual inspection devoted to that stimulus region.
(iii)
Reaction time: defined as the time elapsed between stimulus onset and participant response, indexing response speed during social-cognitive judgment.
For the TREC, non-overlapping geometric areas of interest (AOIs) were manually delineated around the eyes, nose, and mouth using anatomical facial landmarks. These AOIs were used to quantify region-specific visual sampling across facial features. In addition, each facial stimulus was divided along the vertical facial midline into symmetrical left and right hemifacial regions to examine gaze laterality. For each AOI and hemifacial region, fixation count and cumulative fixation duration were extracted. Exploratory comparisons examined gaze allocation across facial regions, emotional valence, sex, age, and task-performance groups.
For the RMET, gaze measures were computed at the task level because the stimuli were restricted to the eye and eyebrow region. This stimulus format prevented equivalent comparisons across full-face AOIs such as eyes, nose, and mouth. Therefore, RMET eye-tracking analyses focused on global visual inspection indices, including fixation count, cumulative fixation duration, and reaction time, during mental-state attribution under constrained perceptual input.
The computerized assessment was implemented within a multi-screen supervisor–participant architecture in which task presentation, response registration, gaze-data acquisition, and real-time monitoring were synchronized during administration (Figure 1). The participant-facing display presented the computerized social-cognition tasks and response options, whereas the supervisor interface controlled task administration and enabled monitoring of acquisition streams and task progression. In the present pilot study, the analytic focus was restricted to behavioral and eye-tracking-derived measures.

2.4. Statistical Analysis

Statistical analyses were performed using SPSS version 23.0. Descriptive statistics were used to summarize sociodemographic, behavioral, and eye-tracking variables. Continuous variables are reported as means and standard deviations, and categorical variables as frequencies and percentages. Within-task comparisons across facial areas of interest and emotional valence were conducted using paired-samples t-tests when the same participants contributed data to both conditions. Between-group comparisons were conducted using independent-samples t-tests for sex, age group, and task-performance group comparisons. Pearson correlation coefficients were calculated to examine associations between behavioral performance, age, education, and gaze-derived measures. Effect sizes were estimated using Cohen’s d. Given the pilot nature of the study, the small sample size, and the number of exploratory contrasts, all analyses were interpreted as hypothesis-generating. No confirmatory inference was made from individual p-values, which were treated as nominal descriptive indices rather than evidence of statistically robust effects. Interpretation prioritized effect sizes, directionality, and consistency with task structure.

3. Results

The sample included 19 healthy adults, of whom 11 were men and 8 were women. Descriptive characteristics are shown in Table 1. Mean age was 35.89 years (SD = 11.35), and mean education was 13.16 years (SD = 2.32). Mean TREC score was 18.00 out of 20 (SD = 1.49), and mean RMET score was 25.10 (SD = 2.20). No sex-related differences were observed in age, years of education, TREC performance, or RMET performance (all p > .05). Most participants were right-handed (89.47%). As specified in the statistical analysis plan, p-values are reported descriptively. TREC scores showed restricted variance and near-ceiling performance; therefore, performance-group results are presented separately as ancillary descriptive analyses.
During the TREC, fixation count varied across facial areas of interest (Table 2; Figure 2B–C). In the total sample, participants showed a mean of 89.35 fixations on the eyes (SD = 51.40), 83.26 on the nose (SD = 32.96), and 21.43 on the mouth (SD = 16.84). Post hoc descriptive contrasts indicated more fixations on the eyes than on the mouth (t = −17.58, p < .001) and on the nose than on the mouth (t = −16.00, p < .001). The corresponding Cohen’s d values were 0.14 for eyes versus nose, 1.78 for eyes versus mouth, and 2.36 for nose versus mouth.
Cumulative fixation duration showed a comparable distribution across TREC facial areas of interest. In the total sample, cumulative fixation duration was 210.63 ms for the eyes (SD = 205.29), 173.16 ms for the nose (SD = 115.83), and 49.60 ms for the mouth (SD = 50.35). Post hoc descriptive contrasts indicated greater cumulative fixation duration on the eyes than on the mouth (t = −13.94, p < .001) and on the nose than on the mouth (t = −10.70, p < .001). The corresponding Cohen’s d values were 0.22 for eyes versus nose, 1.08 for eyes versus mouth, and 1.38 for nose versus mouth.
TREC performance-group analyses are shown in Table 3 and Figure 2D. Participants scoring above the TREC mean showed more fixations on the eyes than participants scoring at or below the TREC mean (t = −2.15, p = .05, d = 0.97). They also showed greater cumulative fixation duration on the eyes (t = −2.33, p = .03, d = 1.00).
TREC gaze laterality also differed by performance group. Participants scoring above the TREC mean showed more fixations on the left hemiface (M = 154.47, SD = 64.63) than participants scoring at or below the TREC mean (M = 33.68, SD = 31.37; t = −5.42, p < .001). Participants scoring at or below the TREC mean showed more fixations on the right hemiface (M = 137.63, SD = 39.02) than participants scoring above the TREC mean (M = 70.84, SD = 56.14; t = 3.07, p = .007). TREC accuracy was positively correlated with left-side facial fixations (r = .64, p = .003).
No sex-related differences were observed in the spatial pattern of TREC gaze allocation. Women showed shorter TREC response times than men (3.33 vs. 4.05 seconds; t = −1.96, p = .06, d = 0.88). In age-group analyses, participants older than 35 years showed more fixations on the eyes (t = −2.30, p = .03, d = 1.03) and greater cumulative fixation duration on the eyes (t = −2.12, p = .05, d = 0.91). Age was positively correlated with fixation count on the eyes (r = .63, p = .003) and cumulative fixation duration on the eyes (r = .49, p = .03). Negative expressions elicited more fixations than positive expressions (t = −2.65, p = .02, d = 0.56). Cumulative fixation duration did not differ by emotional valence.
In the RMET, participants showed a mean fixation count of 467.16 (SD = 171.99) and a mean cumulative fixation duration of 2215.87 ms (SD = 1346.71; Table 4). RMET gaze metrics did not differ according to sex, age, education, performance level, or emotional valence. Participants with higher educational level showed higher RMET scores (t = −2.10, p = .05, d = 0.99).
At the task level, RMET values were higher than TREC values for fixation count, cumulative fixation duration, and reaction time (Figure 2A). Mean response time was 7.50 seconds for the RMET and 3.75 seconds for the TREC.

4. Discussion

This pilot study examined whether eye-tracking can provide task-dependent information about visual sampling during computerized social-cognition assessment. Within its exploratory design, the results suggest that gaze-derived metrics are interpretable only in relation to task structure, stimulus layout, task demands, and AOI definition. From this perspective, the results suggest that gaze-derived metrics may yield coherent patterns of visual sampling, but that their meaning depends critically on task structure, stimulus layout, task demands, and area-of-interest definition(Rahal & Fiedler, 2019).
The most consistent descriptive pattern was observed in the full-face facial emotion-recognition task, where gaze was preferentially allocated to the eyes and nose, with markedly less sampling of the mouth. This pattern should not be interpreted as simple dominance of the eye region, but rather as structured allocation of attention across facial regions with different potential informational value. Classical models of face perception propose that facial processing depends on distributed systems encoding both invariant and changeable facial information(Haxby et al., 2000). Neuropsychological evidence also indicates that the eye region can be especially informative for emotion recognition(Adolphs et al., 2005). More recent work suggests that superior face-recognition performance may depend not only on the amount of information sampled, but also on the computational value of the sampled facial information(Dunn et al., 2025). Within this framework, the present pattern may reflect a strategy in which participants combine eye-region sampling with more central facial integration, although this remains a hypothesis because facial diagnosticity, emotional category, and stimulus salience were not experimentally manipulated.
Ancillary exploratory analyses suggested a possible association between TREC performance and eye-region sampling, broadly consistent with the idea that more accurate performance may be supported by attention to diagnostically informative facial features. However, because TREC performance showed restricted variance and near-ceiling scores, these findings should not be interpreted as robust evidence of gaze–performance coupling. Previous eye-tracking work has reported positive associations between expression-recognition performance and attention to the eyes(Hall et al., 2010). However, the present findings must be interpreted cautiously. The performance groups were defined using a sample-derived cut-off, the sample was small, and the analyses were exploratory. It therefore remains unclear whether greater eye-region sampling reflects a more efficient perceptual strategy, increased task engagement, slower evidence accumulation, or compensatory processing. Disentangling these possibilities will require larger samples, trial-level modelling, preregistered contrasts, and continuous performance measures, because fixation-derived metrics require task-specific interpretation rather than generic interpretation as direct markers of attention or cognition(Hooge et al., 2026; Rahal & Fiedler, 2019).
The exploratory laterality findings raise an additional, highly provisional hypothesis regarding the spatial organization of visual sampling. Participants with higher performance showed greater fixation on the left hemiface, whereas those with lower performance showed greater fixation on the right hemiface. This pattern is compatible with previous evidence suggesting lateralized biases in face processing(Parente & Tommasi, 2008). It is also broadly compatible with models proposing right-hemisphere involvement in processing changeable and socioemotional facial cues(Haxby et al., 2000). However, the present data do not allow discrimination between neurocognitive, stimulus-related, or methodological explanations. Laterality effects may depend on stimulus composition, emotional category, segmentation procedures, or display characteristics, and should therefore be considered exploratory signals rather than evidence of a stable mechanism.
The RMET showed a different descriptive profile. Compared with the TREC, it elicited higher fixation count, longer cumulative fixation duration, and longer response times. This pattern may reflect more prolonged inspection under conditions of restricted perceptual input, consistent with the structure of the RMET as an eye-region task requiring mental-state attribution from limited visual information(Baron-Cohen et al., 2001). However, because the RMET stimulus is restricted to the eye region, it does not allow spatially resolved analysis of gaze allocation across facial features. The higher number and duration of fixations should therefore not be interpreted straightforwardly as better, worse, or more socially informative scanning. This distinction is important because the RMET is widely used as a measure of theory of mind, but recent psychometric work has questioned whether it indexes a unitary construct, suggesting instead that performance may reflect heterogeneous perceptual, lexical, semantic, and inferential components(Hafner et al., 2026; Higgins et al., 2025). The present findings are consistent with the broader methodological point that the same eye-tracking metric can have different meanings across tasks, depending on the perceptual information made available by the paradigm(Hooge et al., 2026; Rahal & Fiedler, 2019).
The absence of clear sex differences in gaze allocation should also be interpreted cautiously. The present study was not powered to detect sex effects, and the male and female subsamples were small. Previous work has reported sex-related differences in facial-expression recognition and face scanning, including greater attention to the eyes in women in some samples(Hall et al., 2010). Therefore, the present null pattern should not be interpreted as evidence that sex plays no role in social-cognitive visual sampling. Rather, it indicates that this pilot sample did not provide reliable support for sex-related differences in gaze allocation. Similarly, the observed associations with age and emotional valence should be considered exploratory. Age-related differences in face scanning and emotion recognition have been documented previously(Noh & Isaacowitz, 2013). Emotional salience and context can also shape how faces are inspected and interpreted(Noh & Isaacowitz, 2013). However, these effects were not the primary focus of the present study and were not corrected for multiple comparisons.
From a methodological perspective, the study illustrates how eye-tracking can be integrated into computerized social-cognition tasks and how its interpretability depends on task format. Eye-tracking has been described as a fine-grained process-tracing method for cognitive and affective mechanisms(Rahal & Fiedler, 2019). However, area-of-interest analyses require careful definition, reporting, and interpretation because AOI-based results depend on spatial accuracy, stimulus layout, and how gaze samples are assigned to regions of interest(Hooge et al., 2026). In this study, the TREC provided spatially resolved information about gaze allocation across facial features, whereas the RMET provided global indices of inspection under perceptual constraint. These differences do not indicate that one task is superior to the other. Rather, they suggest that full-face emotion-recognition tasks may be more informative for studying spatial allocation across facial features, whereas eye-region tasks are better suited to examining inspection time, response latency, and inference under restricted perceptual input(Baron-Cohen et al., 2001; Hafner et al., 2026).
The study may also inform the design of future clinical research, although no clinical claims can be drawn from the present data. Social-cognitive impairments are well established across schizophrenia-spectrum disorders(Green et al., 2015). Recent clinical frameworks continue to emphasize the need for more precise and process-oriented assessment of social cognition(Green et al., 2015). Eye-tracking may contribute to this effort by providing information about how social stimuli are sampled, rather than only whether they are correctly interpreted(Rahal & Fiedler, 2019). However, the present study does not demonstrate diagnostic validity, predictive value, or biomarker potential. Instead, it provides a methodological basis for future studies examining whether altered social cognition reflects differences in perceptual sampling, attentional prioritization, affective salience, or compensatory visual strategies.

4.1. Strengths & Limitations

A strength of this pilot study is its process-level design: eye-tracking was embedded in computerized social-cognition tasks to complement accuracy scores with information on how visual information was sampled during task performance. The use of two tasks with different perceptual constraints is also informative. The full-face TREC allowed preliminary spatially resolved characterization of gaze allocation across facial regions and hemifields, whereas the eye-region RMET provided global inspection metrics under restricted visual input. This contrast clarifies that gaze-derived measures are not task-independent markers, but metrics whose meaning depends on stimulus structure and inferential demand.
The study also has limitations that define its inferential scope. The sample was small, as expected in a pilot study, so findings should not be interpreted as normative patterns. Analyses were exploratory, uncorrected for multiple comparisons, and partly based on sample-derived cut-offs; future studies should use larger samples, preregistered hypotheses, continuous models, and trial-level analyses. The healthy adult sample precludes clinical or diagnostic conclusions, and static stimuli limit ecological generalization. Although calibration was performed and visually validated before task administration, quantitative participant-level calibration accuracy logs were not permanently exported as standalone data files. Therefore, residual spatial error cannot be fully quantified retrospectively, and AOI-level results should be interpreted as exploratory spatial-sampling estimates rather than definitive high-precision gaze-localization measures. Future studies should retain participant-level calibration accuracy, precision, tracking ratio, and signal-loss metrics to strengthen the interpretability and reproducibility of AOI-based conclusions. These limitations define the appropriate interpretive scope of the study: a feasibility and signal-characterization step to identify task-dependent gaze metrics for later validation.

4.2. Further Directions

Future studies should test the robustness of these task-dependent gaze patterns in larger, preregistered samples using continuous, trial-level models that jointly examine fixation count, cumulative fixation duration, reaction time, accuracy, emotional valence, and task format. This would help determine whether the observed patterns reflect stable visual-sampling strategies or sample-specific exploratory signals. Methodologically, future work should refine the interpretation of gaze metrics by improving calibration, standardizing area-of-interest procedures, and testing the reliability of each eye-tracking index. Direct comparisons between full-face, eye-region, dynamic, and context-rich stimuli are needed to clarify which measures are task-specific and which generalize across social-cognitive paradigms. Clinically, this approach should be extended to populations with heterogeneous social-cognitive difficulties, including schizophrenia-spectrum disorders, autism, mood disorders, and neurocognitive disorders. Eye-tracking should be further tested as a process-level complement to behavioral accuracy, with the aim of identifying whether altered performance reflects atypical facial sampling, lateralized exploration, emotional-valence sensitivity, slowed inference, or compensatory visual strategies.

5. Conclusions

This pilot study indicates that eye-tracking can be integrated into computerized social-cognition assessment and can provide task-dependent information about visual sampling. Full-face emotion recognition provided spatially resolved indices of gaze allocation across facial regions, whereas eye-region mental-state inference mainly yielded global inspection metrics under constrained perceptual input. These findings should be considered exploratory and do not establish stable gaze profiles, mechanisms, diagnostic utility, or clinical markers. Instead, they help define a methodological basis for larger, preregistered studies testing whether task-sensitive gaze-derived measures can contribute to the characterization of social-cognitive processing.

Author Contributions

Conceptualization, R.A.-A., S.O. and A.D.-P.; methodology, M.S.-R. and L.R.-C.; software, L.R.-C.; investigation, R.A.-A., M.S.-R., E.S.-S. and A.D.-P.; data curation, M.S.-R. and L.R.-C.; formal analysis, M.S.-R. and L.R.-C.; visualization, L.R.-C. and A.D.-P.; interpretation of results, R.A.-A. and A.D.-P.; writing—original draft preparation, R.A.-A. and A.D.-P.; writing—review and editing, R.A.-A., M.S.-R., E.S.-S., L.R.-C., S.O. and A.D.-P.; supervision, R.A.-A. and A.D.-P.; project administration, R.A.-A.; funding acquisition, S.O. and R.A.-A. All authors read and approved the submitted version of the manuscript.

Funding

This project was funded by Ambar Telecomunicaciones SL within the CogniBio project, supported through the 2019 call of the línea de subvenciones INNOVA of the Consejería de Innovación, Industria, Turismo y Comercio del Gobierno de Cantabria, and by a Miguel Servet contract awarded to Dr Rosa Ayesa-Arriola by the Instituto de Salud Carlos III (CP18/00003). No pharmaceutical company provided financial support for this study.

Institutional Review Board Statement

This study was performed in accordance with the ethical principles of the Declaration of Helsinki. The study protocol was reviewed and approved by the Comité de Ética de Investigación Clínica de Cantabria (CEIC-Cantabria; approval code 2019.024; approval date 15 February 2019).:

Data Availability Statement

The de-identified data supporting the findings of this study are not publicly available because of participant-privacy and data-protection considerations. They may be made available by the corresponding author upon reasonable request, subject to applicable ethical, institutional, and data-protection requirements.

Acknowledgments

The authors would like to thank the Photonic Engineering Group at the University of Cantabria for its technological support, Ambar Telecomunicaciones SL for its contribution to the CogniBio project, and all volunteers who participated in this pilot study. The authors also acknowledge the IDIVAL Mental Illness Research Department for its support in participant recruitment, study coordination, and research infrastructure.

Conflicts of Interest

The authors have no relevant financial or non-financial interests to disclose.

References

  1. Adolphs, R. The social brain: Neural basis of social knowledge. Annu. Rev. Psychol. 2009, 60, 693–716. [Google Scholar] [CrossRef] [PubMed]
  2. Adolphs, R.; Gosselin, F.; Buchanan, T. W.; Tranel, D.; Schyns, P.; Damasio, A. R. A mechanism for impaired fear recognition after amygdala damage. Nature 2005, 433(7021), 68–72. [Google Scholar] [CrossRef] [PubMed]
  3. Alkan, N. Recognition and Misclassification Patterns of Basic Emotional Facial Expressions: An Eye-Tracking Study in Young Healthy Adults. J. Eye Mov. Res. 2025, 18(5), 53. [Google Scholar] [CrossRef] [PubMed]
  4. Arioli, M.; Crespi, C.; Canessa, N. Social Cognition through the Lens of Cognitive and Clinical Neuroscience. BioMed Res. Int. 2018, 4283427. [Google Scholar] [CrossRef] [PubMed]
  5. Baron-Cohen, S.; Wheelwright, S.; Hill, J.; Raste, Y.; Plumb, I. The “Reading the Mind in the Eyes” Test revised version: A study with normal adults, and adults with Asperger syndrome or high-functioning autism. J. Child Psychol. Psychiatry Allied Discip. 2001, 42(2), 241–251. [Google Scholar]
  6. Baron-Cohen, S.; Wheelwright, S.; Jolliffe, T. Is there a “language of the eyes”? Evidence from normal adults, and adults with autism or Asperger syndrome. Vis. Cogn. 1997, 4(3), 311–331. [Google Scholar] [CrossRef]
  7. Bohn, M.; Prein, J. C.; Ayikoru, A.; Bednarski, F. M.; Dzabatou, A.; Frank, M. C.; Henderson, A. M. E.; Isabella, J.; Kalbitz, J.; Kanngiesser, P.; Keşşafoğlu, D.; Köymen, B.; Manrique-Hernandez, M. V.; Magazi, S.; Mújica-Manrique, L.; Ohlendorf, J.; Olaoba, D.; Pieters, W. R.; Pope-Caldwell, S.; …; Haun, D. B. M. A universal of human social cognition: Children from 17 communities process gaze in similar ways. Child Dev. 2026, 97(1), 219–232. [Google Scholar] [CrossRef] [PubMed]
  8. Bölte, S. Social cognition in autism and ADHD. Neurosci. Biobehav. Rev. 2025, 169, 106022. [Google Scholar] [CrossRef] [PubMed]
  9. Caldara, R. Culture Reveals a Flexible System for Face Processing. Curr. Dir. Psychol. Sci. 2017, 26(3), 249–255. [Google Scholar] [CrossRef]
  10. Cavallo, N. D.; Giacobbe, C.; Baiano, C.; Maietta, P.; Trojano, L.; Moretta, P.; Marcuccio, L.; Esposito, F.; Santangelo, G. Neural correlates of social cognition in stroke and traumatic brain injury: A systematic review. In Cognitive, Affective, & Behavioral Neuroscience; 2026. [Google Scholar] [CrossRef] [PubMed]
  11. Corbera, S.; Kurtz, M. M.; Achim, A. M.; Agostoni, G.; Amado, I.; Assaf, M.; Barlati, S.; Bechi, M.; Cavallaro, R.; Ikezawa, S.; Okano, H.; Okubo, R.; Penadés, R.; Uchino, T.; Vita, A.; Yamada, Y.; Bell, M. D. International perspective on social cognition in schizophrenia: Current stage and the next steps. Eur. Psychiatry 2025, 68(1), e9. [Google Scholar] [CrossRef] [PubMed]
  12. De Biase, R.; Ponari, M.; Sagliano, L. I have my eye on the emotion: A review of eye-tracking studies exploring processing of emotional facial expressions. Psychol. Res. 2026, 90(3), 80. [Google Scholar] [CrossRef] [PubMed]
  13. Dodich, A.; Panzavolta, A.; Funghi, G.; Meli, C.; Festari, C.; Chatzikostopoulos, T.; Chicherio, C.; Clarens, F.; de Oliveira, F. F.; Filardi, M.; Ibanez, A.; Invernizzi, L.; Lebouvier, T.; Logroscino, G.; MacPherson, S. E.; Manca, R.; Marra, C.; Matias-Guiu, J. A.; Montembeault, M.; …; Cerami, C. International consensus for the assessment of social cognition in neurocognitive disorders: Framework definition and clinical recommendations of the SIGNATURE initiative. Alzheimer’s Res. Ther. 2025, 18, 6. [Google Scholar] [CrossRef] [PubMed]
  14. Dunn, J. D.; Varela, V.; Popovic, B.; Summersby, S.; Miellet, S.; White, D. Super-recognizers sample visual information of superior computational value for facial recognition. Proceedings. Biol. Sci. 2025, 292(2058), 20252005. [Google Scholar] [CrossRef] [PubMed]
  15. Fernández-Abascal, E. G.; Cabello, R.; Fernández-Berrocal, P.; Baron-Cohen, S. Test-retest reliability of the ‘Reading the Mind in the Eyes’ test: A one-year follow-up study. Mol. Autism 2013, 4(1), 33. [Google Scholar] [CrossRef] [PubMed]
  16. Frith, C. D.; Frith, U. Social cognition in humans. Curr. Biol. 2007, 17(16), R724–32. [Google Scholar] [CrossRef] [PubMed]
  17. Green, M. F.; Horan, W. P.; Lee, J. Social cognition in schizophrenia. Nat. Rev. Neurosci. 2015, 16(10), 620–631. [Google Scholar] [CrossRef] [PubMed]
  18. Hafner, R. M.; Johnson, B. N.; Tone, E. B.; Kivity, Y.; Levy, K. N.; Bedwell, J. S. A multi-site psychometric evaluation of the Reading the Mind in the Eyes Test – Revised. Personal. Individ. Differ. 2026, 250, 113526. [Google Scholar] [CrossRef]
  19. Hall, J. K.; Hutton, S. B.; Morgan, M. J. Sex differences in scanning faces: Does attention to the eyes explain female superiority in facial expression recognition? Cogn. Emot. 2010, 24(4), 629–637. [Google Scholar] [CrossRef]
  20. Haxby, J. V.; Hoffman, E. A.; Gobbini, M. I. The distributed human neural system for face perception. Trends Cogn. Sci. 2000, 4(6), 223–233. [Google Scholar] [CrossRef] [PubMed]
  21. Hessels, R. S.; Iwabuchi, T.; Niehorster, D. C.; Funawatari, R.; Benjamins, J. S.; Kawakami, S.; Nyström, M.; Suda, M.; Hooge, I. T. C.; Sumiya, M.; Heijnen, J. I. P.; Teunisse, M. K.; Senju, A. Gaze behavior in face-to-face interaction: A cross-cultural investigation between Japan and The Netherlands. Cognition 2025, 263, 106174. [Google Scholar] [CrossRef] [PubMed]
  22. Higgins, W. C.; Kaplan, D. M.; Deschrijver, E.; Ross, R. M. Why most research based on the Reading the Mind in the Eyes Test is unsubstantiated and uninterpretable: A response to Murphy and Hall (2024). Clin. Psychol. Rev. 2025, 115(102530), 1–5. [Google Scholar] [CrossRef] [PubMed]
  23. Hooge, I. T. C.; Nyström, M.; Niehorster, D. C.; Andersson, R.; Foulsham, T.; Nuthmann, A.; Hessels, R. S. The fundamentals of eye tracking part 6: Working with areas of interest. Behav. Res. Methods 2026, 58(3), 65. [Google Scholar] [CrossRef] [PubMed]
  24. Huerta-Ramos, E.; Ferrer-Quintero, M.; Gómez-Benito, J.; González-Higueras, F.; Cuadras, D.; Del Rey-Mejías, A. L.; Usall, J.; Ochoa, S. Translation and validation of the Baron-Cohen Face Test in the Spanish population. Actas Esp. De Psiquiatr. 2021, 49(3), 106–113. [Google Scholar]
  25. Noh, S. R.; Isaacowitz, D. M. Emotional Faces in Context: Age Differences in Recognition Accuracy and Scanning Patterns. In Emotion; Washington, D.C., 2013; Volume 13, 2, pp. 238–249. [Google Scholar] [CrossRef] [PubMed]
  26. Parente, R.; Tommasi, L. A bias for the female face in the right hemisphere. Laterality 2008, 13(4), 374–386. [Google Scholar] [CrossRef] [PubMed]
  27. Pitcher, D. Neuropsychological evidence of a third visual pathway specialized for social perception. Nat. Commun. 2025, 16(1), 5774. [Google Scholar] [CrossRef] [PubMed]
  28. Rahal, R.-M.; Fiedler, S. Understanding cognitive and affective mechanisms in social psychology through eye-tracking. J. Exp. Soc. Psychol. 2019, 85, 103842. [Google Scholar] [CrossRef]
  29. Rosati, A. G. Cognitive adaptations in humans and other great apes; 2026; pp. 384–396. [Google Scholar] [CrossRef]
  30. Shafiei, M.; Reik, M.; Görner, M.; Taubert, N.; Giese, M.; Thier, P. Rhesus monkeys use both eye and head gaze to reallocate covert spatial attention facilitating visual perception. Cogn. Affect. Behav. Neurosci. 2026. [Google Scholar] [CrossRef] [PubMed]
  31. Yang, Q.; Fu, Y.; Yang, Q.; Yin, D.; Zhao, Y.; Wang, H.; Zhang, H.; Sun, Y.; Xie, X.; Du, J. Eye movement characteristics of emotional face recognizing task in patients with mild to moderate depression. Front. Neurosci. 2024, 18. [Google Scholar] [CrossRef] [PubMed]
Figure 1. Proposed multi-screen architecture for computerized social-cognition assessment with real-time biometric monitoring. Schematic representation of the experimental setup. The participant-facing display presented the computerized social-cognition tasks and response options, while the supervisor interface-controlled task administration and enabled real-time monitoring of acquisition streams and task progression. The system was designed to synchronize stimulus presentation, response registration, gaze-data recording, and concurrent signal visualization, allowing subsequent extraction of behavioral and eye-tracking-derived measures, including fixation count, cumulative fixation duration, reaction time, and area-of-interest metrics. In the present pilot study, analyses focused on gaze-derived and behavioral variables. The diagram is schematic and does not display participant-identifiable data.
Figure 1. Proposed multi-screen architecture for computerized social-cognition assessment with real-time biometric monitoring. Schematic representation of the experimental setup. The participant-facing display presented the computerized social-cognition tasks and response options, while the supervisor interface-controlled task administration and enabled real-time monitoring of acquisition streams and task progression. The system was designed to synchronize stimulus presentation, response registration, gaze-data recording, and concurrent signal visualization, allowing subsequent extraction of behavioral and eye-tracking-derived measures, including fixation count, cumulative fixation duration, reaction time, and area-of-interest metrics. In the present pilot study, analyses focused on gaze-derived and behavioral variables. The diagram is schematic and does not display participant-identifiable data.
Preprints 230505 g001
Figure 2. Looking dynamics across social-cognition tasks. Panel A shows a task-level descriptive comparison between the full-face facial emotion-recognition task (TREC) and the eye-region mental-state attribution task (RMET), including mean fixation count, cumulative fixation duration, and reaction time. Panel B shows TREC gaze allocation by facial area of interest, with fixation count and cumulative fixation duration for the eyes, nose, and mouth in the total sample; data are presented as mean ± SD, and brackets indicate post hoc contrasts reported for the total sample. Panel C presents an effect-size heatmap for TREC facial-region contrasts, showing Cohen’s d values for fixation count and cumulative fixation duration; positive values indicate higher values for the first facial region named in each contrast. Panel D shows TREC gaze laterality by performance group, with fixation count on the left and right hemiface in participants scoring above the sample mean versus those scoring at or below the sample mean; data are presented as mean ± SD, and asterisks denote nominal exploratory p-value thresholds and should not be interpreted as confirmatory evidence. **p < .01, ***p < .001. TREC, Test de Reconocimiento Emocional en Caras; RMET, Reading the Mind in the Eyes Test; AOI, area of interest.
Figure 2. Looking dynamics across social-cognition tasks. Panel A shows a task-level descriptive comparison between the full-face facial emotion-recognition task (TREC) and the eye-region mental-state attribution task (RMET), including mean fixation count, cumulative fixation duration, and reaction time. Panel B shows TREC gaze allocation by facial area of interest, with fixation count and cumulative fixation duration for the eyes, nose, and mouth in the total sample; data are presented as mean ± SD, and brackets indicate post hoc contrasts reported for the total sample. Panel C presents an effect-size heatmap for TREC facial-region contrasts, showing Cohen’s d values for fixation count and cumulative fixation duration; positive values indicate higher values for the first facial region named in each contrast. Panel D shows TREC gaze laterality by performance group, with fixation count on the left and right hemiface in participants scoring above the sample mean versus those scoring at or below the sample mean; data are presented as mean ± SD, and asterisks denote nominal exploratory p-value thresholds and should not be interpreted as confirmatory evidence. **p < .01, ***p < .001. TREC, Test de Reconocimiento Emocional en Caras; RMET, Reading the Mind in the Eyes Test; AOI, area of interest.
Preprints 230505 g002
Table 1. Sample characteristics by sex.
Table 1. Sample characteristics by sex.
Total (N = 19) Men (n = 11) Women (n = 8)
Mean SD Mean SD Mean SD t p
Age (years) 35.89 11.35 37.10 13.95 34.25 6.92 -0.58 .57
Education (years) 13.16 2.32 13.00 2.57 13.38 2.07 0.35 .73
TREC score 18.00 1.49 17.82 1.66 18.25 1.28 0.61 .55
RMET score 25.10 2.20 24.55 1.92 25.88 2.47 1.32 .20
Right-handedness (%) 89.47 90.91 87.50
Table 2. TREC fixation count and cumulative fixation duration by facial area of interest.
Table 2. TREC fixation count and cumulative fixation duration by facial area of interest.
Outcome Sample Eyes (E) Nose (N) Mouth (M) Cohen’s d
Mean SD Mean SD Mean SD Eyes vs Nose Eyes vs Mouth Nose vs Mouth post hoc
Fixation count Total sample 89.35 51.40 83.26 32.96 21.43 16.84 0.14 1.78 2.36 E > M: t = -17.58, p < .001; N > M: t = -16.00, p < .001
Men 101.33 51.89 87.74 35.87 21.33 20.81 0.30 2.02 2.26
Women 72.88 49.11 77.11 29.68 21.57 10.51 -0.10 1.44 2.49
Fixation duration (ms) Total sample 210.63 205.29 173.16 115.83 49.60 50.35 0.22 1.08 1.38 E > M: t = -13.94, p < .001; N > M: t = -10.70, p < .001
Men 221.44 122.86 186.47 90.80 54.87 58.56 0.32 1.73 1.72
Women 195.75 293.88 154.86 148.53 42.35 38.91 0.18 0.73 1.04
Table 3. TREC gaze laterality by performance group.
Table 3. TREC gaze laterality by performance group.
TREC > 18 TREC ≤ 18 Between-group
Mean SD Mean SD t p
Left hemiface fixations 154.47 64.63 33.68 31.37 -5.42 < .001
Right hemiface fixations 70.84 56.14 137.63 39.02 3.07 .007
Table 4. RMET task-level gaze metrics by sex.
Table 4. RMET task-level gaze metrics by sex.
Total sample Men Women
Mean SD Mean SD Mean SD p Cohen’s d
Number of fixations 467.16 171.99 508.98 165.55 409.65 174.30 > .05 0.58
Fixation duration (ms) 2215.87 1346.71 2432.10 1276.39 1918.55 1470.27 > .05 0.37
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.