Submitted:
19 August 2026
Posted:
21 August 2026
You are already at the latest version
Abstract
Large language models (LLMs) are increasingly used for qualitative coding, yet it remains unclear how closely their coding outputs align with frameworks developed by human experts. Existing studies tend to report overall similarity but rarely distinguish thematic granularity from dimension-level semantic alignment or systematically identify which constructs are stably recovered. This study therefore compares human expert coding with three zero-shot LLM-generated qualitative analyses using a structured framework-level mapping method. The method maps human-coded dimensions to LLM-generated qualitative frameworks and classifies each relationship as direct, partial or merged, or not recovered. To quantify alignment, three evaluation metrics are used: Direct Coverage (DC), which measures strict boundary recovery; Broad Coverage (BC), which includes partial or merged mappings; and Frequency-Weighted Mapping (FWM), which weights each dimension by its human coding frequency to reflect construct prominence. The method is demonstrated using a human-derived emotional-intelligence framework developed from 34 Chinese behavioral-event interviews and three zero-shot frameworks generated from the same corpus by DeepSeek V4.0 PRO, Qwen 3.6, and GPT 5.6 under an identical prompt. Two senior experts jointly evaluated 36 dimension-level relationships. The three LLM outputs achieved direct coverage ranging from 75.0% to 91.7%, broad coverage from 83.3% to 100.0%, and frequency-weighted mapping from 80.3% to 94.9%. Eight of the 12 human-derived dimensions were directly recovered across all three LLM frameworks, whereas guidance and motivation, big-picture awareness, care and support orientation, and teamwork showed variable alignment. The framework-level comparison distinguishes thematic organization from construct-boundary recovery and supports the use of LLMs as complementary analytical assistants while retaining human responsibility for contextual interpretation and construct definition.
Keywords:
large language models
; qualitative data analysis
; framework mapping
; zero-shot prompting
; emotional intelligence
; human–AI comparison
1. Introduction
Large language models (LLMs) are increasingly used to generate codes, organize themes, and summarize interview corpora. Their speed and linguistic fluency make them attractive for qualitative analysis, but these capabilities do not establish that model-generated categories reproduce human interpretation, preserve contextual meaning, or maintain comparable construct boundaries. The central methodological question is therefore comparative: when human researchers and LLMs analyze the same corpus, where do their frameworks converge, and where do they differ in construct boundaries and thematic granularity?
Emotional intelligence (EI) provides a useful test case because ability, competency, emotional–social, and trait traditions share a recognizable core but define the boundaries of the construct differently [1,2,3,4,5,6]. These differences are especially visible in whether resilience, adaptation, communication, teamwork, care, and influence are treated as parts of EI [7,8]. Framework reconstruction must therefore identify not only familiar emotional content but also the boundaries among related behavioral dimensions.
Qualitative interview methods are well suited to domain-specific construct development when analytic procedures and links between interpretation and evidence are made explicit [9]. Thematic analysis identifies patterned meanings across a corpus [10], while grounded theory and the constant comparative method support the inductive development of categories and their relationships [11,12,13]. These approaches rely on sustained engagement with the data, researcher judgment, reflexivity, and transparent evidence use; without such safeguards, sampling, analytic, or confirmation biases may be reproduced [14]. LLMs can accelerate code generation, theme organization, and comparison across large text collections, but speed and fluency do not establish interpretive validity.
Existing human–LLM comparisons generally report partial convergence while retaining a central role for human review, but their evaluation methods and reporting practices remain heterogeneous [15,16,17,18,19,20,21,22,23,24,25]. This leaves unresolved the extent to which different LLM outputs recover the construct boundaries established through human qualitative analysis.
A particularly close methodological precedent is Shanwetter Levit and Saban’s comparison of investigator-led analysis, ChatGPT-4o, and Gemini Advance Pro 1.5 across 33 oncology interviews [28]. Their study applied the same zero-shot prompt to both models and compared convergent and divergent themes. However, two questions remain insufficiently resolved for framework reconstruction: whether different LLM outputs generated under a shared zero-shot prompt recover the conceptual boundaries established by human analysis of the same corpus, and which human-derived dimensions recur or vary across outputs.
To address these gaps, the present study compares a human-led grounded-theory framework with three LLM analyses generated from the same 34 Chinese behavioral-event interviews concerning EI. The same zero-shot thematic-analysis prompt was used for all three outputs. RQ1 asks: To what extent do the three LLM outputs align with the dimensions of the human-led, domain-specific EI framework? RQ2 asks: Which human-derived dimensions are shared across the three outputs, and which vary across outputs? The study applies a framework-level mapping procedure that examines both how the LLMs organize themes and whether those themes correspond in meaning to individual human-defined dimensions. Direct coverage, broad coverage, and frequency-weighted mapping provide complementary descriptive summaries of strict, inclusive, and prominence-sensitive framework alignment. Using the human framework as the comparison reference, the procedure identifies where zero-shot LLM analyses converge with, compress, fragment, or omit human-defined constructs.
Figure 1 summarizes the study design, from interview data collection through parallel human and LLM analysis to framework-level comparison and the two research questions.
2. Conceptual Background and Related Research
2.1. Emotional Intelligence as a Multi-Model Construct
The modern EI literature developed from the proposition that emotion can function as information within thought and action. Salovey and Mayer initially defined EI in terms of monitoring one’s own and others’ feelings, discriminating among them, and using this information to guide thinking and behavior [1]. The subsequent ability model organized EI into four branches: perceiving emotion, using emotion to facilitate thought, understanding emotional meanings and transitions, and managing emotion in oneself and others [2,3]. In this tradition, EI is a form of mental ability directed toward emotion-laden information rather than a general label for socially desirable behavior.
Competency approaches broaden the focus from maximal emotional ability to observable patterns of effective behavior. Boyatzis organized emotional and social competencies around self-awareness, self-management, social awareness, and relationship management [4]. The relationship-management domain explicitly includes influence, coaching, conflict management, inspirational leadership, and teamwork. This behavioral emphasis is particularly relevant to interview research because participants often describe EI through what a person does in difficult interactions rather than through an abstract account of emotional reasoning.
The emotional–social model similarly combines intrapersonal and interpersonal functioning with adaptability and stress management [5]. It therefore provides a conceptual basis for treating flexibility and psychological resilience as part of an EI framework rather than as unrelated outcomes. Trait EI takes a different position. It concerns emotion-related self-perceptions, including emotionality, self-control, sociability, and well-being, and is situated within personality structure [6]. The present human framework was derived from behavioral-event accounts rather than a trait self-report instrument, so similarity to a trait label does not imply that the underlying construct was operationalized in the same way.
These traditions share a recognizable core but differ at the construct boundary. Self-awareness, perception of others, empathy, and emotion regulation appear across several traditions. Resilience, adaptation, communication, influence, collective responsibility, and teamwork depend more strongly on whether EI is defined narrowly as an emotional ability or more broadly as a set of emotional and social competencies [7,8]. Table 1 summarizes these differences and their conceptual relevance to the human reference framework used in this study.
2.2. LLMs as Qualitative Analysts
Using an LLM as a qualitative analyst can refer to several distinct operations: segmenting text, proposing codes, grouping codes into categories, abstracting themes, selecting quotations, and producing an integrated account. These operations should not be collapsed into a single claim of “qualitative performance.” In empirical social research, AI-supported qualitative content-analysis systems have been evaluated not only for categorization accuracy but also for how algorithmic recommendations affect researchers’ reflection and decision processes [29]. De Paoli demonstrated that an LLM can produce an inductive thematic analysis with plausible thematic structure, while also emphasizing the limits of treating generated output as equivalent to human interpretation [15]. More generally, fluent text generation does not guarantee logical reasoning or reliable compliance with strict analytic instructions [26]. Evaluation must therefore distinguish a plausible thematic structure from faithful recovery of the construct boundaries established through human analysis.
Direct human–LLM comparisons show substantial but incomplete convergence. Mathis et al. found that open-source LLMs could generate themes with substantial similarity to a conventional human analysis of healthcare interviews, but still required validation and careful methodological control [16]. Wachinger et al. likewise observed overlap between ChatGPT and a human researcher while documenting prompt sensitivity, imperfections, and differences in analytic nuance [17]. Zeng et al. compared ChatGPT and DeepSeek with human coding of 15 nursing interviews and reported strong coding consistency (Cohen’s ) and substantial time savings. They also identified differences in the interpretation of implicit emotional cues and analytic scope, reinforcing the need for human validation of model-generated themes [30]. In the closest precedent to the present study, Shanwetter Levit and Saban compared investigator-led analysis, ChatGPT-4o, and Gemini Advance Pro 1.5 across 33 oncology interviews. The investigator emphasized psychosocial and emotional depth, whereas the LLMs more readily identified structural, temporal, and logistical patterns; the authors consequently proposed a complementary human–AI model rather than substitution [28].
Recent work has also moved from one-off prompting toward structured coding workflows. Qiao et al. used LLMs for inductive coding of semistructured maternal-health interviews and reported an 81% reduction in coding time, while retaining human evaluation of generated codes [18]. Than et al. tested multiple prompt structures for qualitative coding and showed that generative LLM workflows can approximate hand-coded outputs under some conditions, while raising methodological and epistemological questions about what is being reproduced [19]. A systematic review and controlled experiment further showed that structured prompting can improve human-rated usefulness, positioning prompt design as a cognitive interface rather than a neutral formatting choice [31]. Grounded-theory and protocol studies likewise demonstrate that instructions shape how models organize categories and relationships [20,32,33]. Because prompt design is consequential, a shared zero-shot prompt provides a useful descriptive setting for examining how different LLM outputs organize the same corpus without introducing a second prompt workflow into the comparison.
These empirical possibilities do not remove the epistemological distinction between computational coding and reflexive qualitative inquiry. Jowsey et al. argue that generative AI cannot perform reflexive thematic analysis because reflexivity depends on a situated, accountable, and interpretive human researcher [27]. A more defensible position is to treat LLMs as bounded analytic assistants that can widen comparison, challenge a codebook, or accelerate organization, while human researchers remain responsible for theoretical positioning, contextual judgment, negative-case analysis, and the final account [34]. Hybrid workflows can scale code application while retaining human interpretive leadership and iterative codebook refinement [35]. Related research illustrates how computationally identified topics can be combined with researcher review and interpretation in qualitative analysis [36]. This distinction is especially important when an LLM output is compared with a human-derived framework: alignment can indicate conceptual recoverability, but it does not establish equivalence of analytic process.
Reporting practices remain uneven. Kempny et al.’s scoping review found LLM use across the qualitative workflow, most commonly for coding assistance and theme identification. Reported human–LLM agreement varied from 36% to 99%, and 97% of included studies incorporated human verification [25]. Cultural interpretation adds another layer: a comparative Japanese study found that apparent thematic performance could coexist with challenges in culturally situated interpretation [37]. In a low-resource Kuwaiti dialect, comparisons of 11 LLMs against manually labeled data similarly demonstrated substantial variation across models and zero-shot and few-shot prompt conditions [38]. These findings support comparisons that distinguish output granularity from dimension-level semantic alignment.
3. Materials and Methods
3.1. Study Design and Source Materials
This study used a retrospective human–LLM document comparison based on 34 Chinese behavioral-event interview transcripts containing approximately 140,000 Chinese characters. The original purposive sample included adult professionals and senior trainees across multiple role levels and functions within the same structured, hierarchical organizational setting. Two senior experts jointly reviewed the transcripts and removed fields unrelated to the emotional-intelligence analysis before coding. The same two experts conducted the human analysis and jointly carried out the LLM dimension mapping described below. The resulting corpus was processed through one human-led qualitative analysis and three archived zero-shot LLM outputs. The human-led analysis served as the reference framework for comparing thematic structure, thematic granularity, and alignment with human-derived dimensions.
The corpus contained 34 interviews. The present document-level comparison does not make population-level claims about the interview participants. Detailed role, tenure, and demographic characteristics are not reported because they are not required for the analytic comparison and could increase re-identification risk within a bounded organizational setting.
3.2. Interview Design, Sampling, and Data Collection
The original researchers purposively selected participants across role levels and occupational functions to obtain varied behavioral-event accounts. They used behavioral-event interviewing, a form of in-depth interviewing organized around the Situation–Task–Action–Result (STAR) sequence. Questions 3 and 5 of the interview guide were the principal STAR questions. They asked participants to recount one or two events that best represented high EI and one or two events that represented low EI, and then prompted for the situation, assigned task, concrete behavior, and outcome. The remaining questions addressed prior learning and work experience, the functions and components of EI, abilities associated with high or low EI, and possible approaches to EI development. Together, the interviews elicited accounts of EI in collaborative, coordination, practical, educational, and administrative situations. The interview guide is provided in Appendix A.1.
Interview appointments were normally arranged approximately one week in advance and reconfirmed two to three days before the scheduled time. Depending on circumstances, interviews were conducted face to face or by telephone and typically lasted 30–45 min. Recording was used only after consulting the participant; otherwise, the interviewer took written notes. The source report states that participation was voluntary, that confidentiality was promised, and that interviews were managed to avoid unnecessary disclosure of sensitive information. All coding and dimension mapping were conducted on the original Chinese text. The dimension labels used in this article were selected to preserve the meaning of the source categories; for example, “care and support orientation” reflects the source category’s emphasis on relational care and emotional support.
3.3. Human Grounded-Theory Analysis
The two senior experts conducted the human analysis in NVivo 11 Plus. Before formal coding, they jointly reviewed a word-frequency search using a minimum word length of three Chinese characters and a stop-word list as an initial familiarization step. The theoretical framework itself was developed through open, axial, and selective coding rather than from word frequency.
For framework construction, the first 25 interviews, approximately three quarters of the corpus, were examined sentence by sentence while the coding frame remained open. Relevant expressions were converted into initial concepts and grouped into categories. Axial coding then linked related concepts and categories through a paradigm involving causes, phenomena, contexts, influencing conditions, and behavioral strategies. During selective coding, the experts compared and integrated the main categories into a higher-order framework through joint discussion, constant comparison, and reference to existing theoretical knowledge.
The remaining nine interviews were coded separately from the construction set using the same procedures and standards. The experts assessed theoretical saturation by determining whether the holdout interviews introduced new categories, category types, or important relationships that required revision of the framework.
3.4. Zero-Shot LLM Outputs and Prompt
The three zero-shot analyses were generated through API calls to DeepSeek V4.0 PRO, Qwen 3.6, and GPT 5.6. For readability, DeepSeek, Qwen, and GPT are used as short forms in figures and tables.
Each model received all 34 transcripts and the same zero-shot prompt. The prompt instructed the LLM to identify EI-related first-order and second-order themes, define each theme, and provide two verbatim quotations with participant identifiers and line numbers. It explicitly prohibited the use of an existing EI model and prohibited rewriting, supplementing, or inventing quotations. The prompt is provided in Appendix A.2.
3.5. Framework-Level Human–LLM Mapping Procedure
The framework-level mapping procedure used the 12 human-derived dimensions as reference units for the comparison. For each of the three LLM outputs, two senior experts jointly compared every human dimension with the LLM’s second-order themes using the theme names, definitions, and supporting quotations. The closest relationship was classified as direct, partial or merged, or not mapped. A direct mapping required substantially similar central meaning and scope. A partial mapping indicated that the human dimension appeared only within a broader, narrower, or adjacent LLM theme. For descriptive scoring, direct mappings were assigned 1, partial or merged mappings 0.5, and non-mappings 0. This procedure produced a 12-by-3 matrix, or 36 dimension-level comparisons. The classifications describe semantic overlap among the final outputs; they do not imply that an LLM reproduced the human analytic process. Interpretive differences were discussed until the two experts reached consensus; independent pre-consensus ratings were not retained, so no inter-rater statistic is reported.
Three descriptive summaries were used to present alignment at the output level. Let denote the number of human-derived reference dimensions, let denote the consensus mapping score for output m and dimension i, and let denote the human coding frequency of dimension i. The indicator function equals 1 when its condition is true and 0 otherwise. Direct coverage (DC) measures strict recovery of dimension boundaries:
Broad coverage (BC) measures whether a dimension is represented either directly or within a partial or merged theme:
Frequency-weighted mapping (FWM) gives greater descriptive weight to dimensions that appeared more frequently in the human coding. A direct mapping contributes the full dimension frequency, a partial or merged mapping contributes half of that frequency, and a non-mapping contributes zero:
All three metrics range from 0 to 100%. DC is the strictest summary, BC is inclusive of partial semantic representation, and FWM is prominence-sensitive. They are descriptive measures of alignment with the human reference framework rather than estimates of population-level accuracy or inferential test statistics. When dimension-level convergence across the three outputs is described, the same mapping values are averaged across outputs.
4. Results
4.1. The Human Reference Framework
The preliminary word-frequency review highlighted expressions corresponding to “interpersonal relationships” and “dealing with people,” indicating the prominence of relational interaction in the corpus. Open coding produced 356 coded excerpts and 39 categories. Axial coding consolidated these categories into 12 EI dimensions, and selective coding organized the dimensions into three higher-order categories. When the remaining nine interviews were coded as the holdout set, they introduced alternative wording for existing ideas but no new category, category type, or important relationship. The two experts therefore judged the framework theoretically saturated.
The most frequently coded dimensions were empathy (46), organizational coordination (44), flexibility (36), awareness of others (35), and guidance and motivation (33). Care and support orientation and teamwork were least frequent (18 each). The frequencies of the 12 dimensions sum to 356 and are used as descriptive weights in the mapping metric. Table 2 presents the complete human reference framework derived from the 34 behavioral-event interviews.
The three higher-order categories represent a progression from intrapersonal functioning, through interpersonal perception and situational adaptation, to the enactment of emotional understanding in collective activity. They should not be read as psychometric factors or as stages that every interaction must follow. Rather, they organize the behavioral meanings developed in the human analysis and clarify distinctions among dimensions that share related content.
4.1.1. Self-Awareness and Regulation
The first higher-order category combined self-awareness, psychological resilience, emotion regulation, and big-picture awareness. The source material treated self-awareness as knowledge of one’s traits, capabilities, emotional state, and current role. One participant linked EI to clarity about one’s orientation, standpoint, knowledge structure, personal identity, and current developmental stage. Emotion regulation was defined behaviorally rather than as emotional suppression alone. Another participant described how emotional fluctuation can make communication impatient and cause a person to transmit a message different from the one intended. This connection between emotional state and communicative effect explains why regulation was positioned as a context-specific capability.
Psychological resilience integrated setback tolerance, persistence, self-discipline, recovery, and self-motivation. One participant used the metaphor of a rechargeable power bank, explaining that frustration had to be absorbed and converted into energy for continued self-motivation. Big-picture awareness extended self-regulation toward shared responsibility. Another participant contrasted personal preference with the consequences for a collaborative group, arguing that refusing an assigned task without considering its negative effect on others represented low EI. The category therefore joined internal monitoring with behavior governed by responsibility and shared priorities.
The four dimensions describe different temporal and behavioral aspects of self-related EI. Self-awareness is diagnostic: it concerns recognizing one’s emotional state, capabilities, role, and limitations. Emotion regulation concerns the management of an immediate response, especially when irritation or pressure could distort communication. Psychological resilience concerns recovery and sustained action after difficulty rather than control of a single emotional episode. Big-picture awareness extends self-management beyond the individual by connecting emotional restraint and initiative with responsibility for shared goals.
4.1.2. Environmental Perception and Adaptation
The second higher-order category contained empathy, awareness of others, flexibility, and care and support orientation. Empathy referred to deliberate perspective-taking. One participant described an appropriate problem-solving attitude as remaining calm and “putting oneself in the other person’s position.” Awareness of others concerned the perception of emotional cues and circumstances. Another participant explained that facial expression, gaze, and subtle body movement can indicate whether another person is friendly, resistant, willing to communicate, or feels offended. The distinction between empathy and awareness of others is important: the former concerns adopting another perspective, whereas the latter concerns detecting the other’s state before deciding how to respond.
Flexibility connected emotional understanding to situational adaptation. One participant stated that a person cannot remain within a single personal viewpoint but must consider “who [they are] interacting with, what the environment is, and what kind of situation it is.” Care and support orientation captured sincere care for peers and collaborators. Another participant recalled that treating others with genuine feeling encouraged reciprocal trust, willingness to work together, and positive feedback for task performance. This dimension was therefore not generic helpfulness; it represented emotionally attentive relationship building in the source setting.
The original analysis implies a sequence among these dimensions without treating them as interchangeable. Awareness of others concerns detecting emotional and situational cues. Empathy concerns interpreting a situation from the other person’s perspective. Care and support orientation concerns translating that understanding into sustained relational behavior, while flexibility concerns adjusting one’s own communication and action to the person and setting.
4.1.3. Communication, Coordination, and Influence
The third higher-order category included guidance and motivation, organizational coordination, teamwork, and communication. Organizational coordination reflected a bridge function frequently described by participants. One participant described “a connecting or coordinating role” between receiving task-related information, carrying out one’s own work, allocating tasks, and managing collaborative relationships. Communication referred both to accuracy and acceptability. Another participant observed that when a person cannot express an intended meaning clearly, or communicates it too rigidly, the interaction may create rather than resolve conflict.
Teamwork emphasized the use of collective knowledge and coordinated effort. One participant associated high EI with uniting colleagues around a shared objective across routine and temporary tasks. Guidance and motivation described a more directional form of influence. Another participant distinguished adapting oneself from trying to force the other person to change, then described the higher-level capacity to use language, expression, and action to guide another person’s emotion toward a constructive state. Together, these four dimensions show that the source framework treated EI not only as recognition and regulation but also as the capacity to organize collective action.
The four dimensions also operate at different levels of collective behavior. Communication concerns whether meaning is expressed accurately and in a form the other person can accept. Organizational coordination concerns arranging interdependent roles, relationships, and resources. Teamwork concerns mobilizing collective knowledge and effort toward a shared task. Guidance and motivation concerns influencing trust, emotion, cohesion, and willingness to act.
4.2. Thematic Structure across LLM Outputs
The three zero-shot LLM outputs contained six to ten first-order themes and 13 to 22 second-order themes. Figure 2 provides a structural overview. Panel A compares the numbers of higher-order and dimension-level categories in the human framework with the first- and second-order themes in the LLM outputs. Panel B summarizes direct coverage, broad coverage, and frequency-weighted mapping across the 12 human dimensions. DS-Z, QW-Z, and GPT-Z denote the zero-shot outputs from DeepSeek V4.0 PRO, Qwen 3.6, and GPT 5.6, respectively.
Table 3 gives the numerical results. The columns labeled “1st” and “2nd” report the numbers of first- and second-order themes. “Direct,” “Partial,” and “Missing” report the numbers of human dimensions assigned to each mapping category. Direct coverage counts only direct mappings, whereas broad coverage also includes partial or merged mappings; “Weighted” denotes frequency-weighted mapping. Direct coverage ranged from 75.0% for Qwen 3.6 to 91.7% for DeepSeek V4.0 PRO. Broad coverage ranged from 83.3% for Qwen 3.6 to 100.0% for GPT 5.6. The frequency-weighted mapping scores ranged from 80.3% to 94.9%.
The FWM numerators are the summed human coding frequencies of directly mapped dimensions plus one half of the frequencies of partially mapped dimensions. DeepSeek therefore contributes 338 weighted frequency units; Qwen contributes ; and GPT contributes . This calculation preserves the distinction between a framework that mentions a dimension within a broader theme and one that reconstructs it as a direct dimension.
4.3. Dimension-Level Mapping across LLM Outputs
The output-level summaries do not indicate which human dimensions were consistently represented or varied across outputs. Figure 3 therefore displays the complete dimension-level mapping across the three LLM outputs. Each cell shows the strongest semantic relationship identified between a human dimension and the second-order themes of an LLM output: D denotes a direct mapping, P denotes a partial or merged mapping, and N denotes no clear mapping.
Self-awareness, psychological resilience, emotion regulation, empathy, awareness of others, flexibility, organizational coordination, and communication received direct mappings in all three outputs. Big-picture awareness and care and support orientation received direct mappings in two outputs and were not mapped in the third. Teamwork was represented in all three outputs but received one direct and two partial mappings. Guidance and motivation varied most across outputs: it received one direct mapping, one partial mapping, and one non-mapping, producing the lowest mean consensus-mapping score of 50.0%.
4.4. Dimension Prominence and Cross-Output Recovery
The dimension-level matrix shows which dimensions varied, but it does not show whether a dimension’s prominence in the human analysis was related to how consistently the outputs recovered it. To examine this relationship, Figure 4 plots each human-derived dimension by its human coding frequency and its mean consensus mapping averaged across the three outputs. Because each output contributes a direct (1), partial or merged (0.5), or non-mapping (0) score, the three-output mean can take only three values: 50.0%, 66.7%, or 100.0%.
Most of the more frequently coded dimensions clustered at full recovery. Empathy (46), organizational coordination (44), flexibility (36), awareness of others (35), emotion regulation (30), and communication (28) were mapped directly in all three outputs, as were the less frequent self-awareness and psychological resilience (20 each). The dimensions that fell below full recovery were mostly lower-frequency boundary constructs: big-picture awareness (28), care and support orientation (18), and teamwork (18) each reached 66.7%. The salient exception was guidance and motivation (33), the fifth most frequently coded dimension, which recorded the lowest mean consensus mapping of 50.0%. In this corpus, higher human coding frequency was descriptively associated with more consistent recovery for most dimensions, but the pattern was not universal and no inferential relationship was tested. Guidance and motivation combined high prominence with the greatest cross-output divergence, reinforcing that it lies on the contested boundary between EI and context-specific influence.
4.5. Similarity across Zero-Shot Outputs
The three mapping summaries compare each output with the human reference framework but do not describe how similar the outputs were to one another. Figure 5 reports the pairwise Jaccard similarity of the sets of directly recovered human dimensions across the three zero-shot outputs.
The GPT and DeepSeek outputs were most similar (0.92), the GPT and Qwen outputs were intermediate (0.83), and the DeepSeek and Qwen outputs were least similar (0.75). The lower DeepSeek–Qwen overlap reflects the two human dimensions that Qwen did not directly recover rather than conflicting direct mappings of dimensions shared by the two outputs. Even the lowest pairwise value indicates that the three outputs agreed on most of the directly recovered dimensions, which is consistent with the stable reconstructable core identified in the dimension-level analysis.
5. Discussion
This study examined how three LLM outputs generated under the same zero-shot prompt reconstructed a human-derived emotional intelligence framework from the same interview corpus. The outputs were compared with a 12-dimension human framework derived from 34 behavioral-event interviews. Two senior experts jointly evaluated thematic structure and completed 36 dimension-level mapping decisions. The comparison focused on what each output reconstructed and how its thematic granularity related to coverage of the human-derived framework; it was not designed to rank general model capability.
5.1. Framework Reconstruction: Stable Core and Output-Dependent Boundaries
The three LLM outputs mapped to substantial portions of the human-led EI framework without being supplied the framework itself. Self-awareness, psychological resilience, emotion regulation, empathy, awareness of others, flexibility, organizational coordination, and communication received direct mappings in every output. These dimensions were expressed through recurring and behaviorally explicit contrasts in the source material, including recognizing one’s own state, noticing another person’s state, regulating emotion rather than reacting impulsively, adapting to the person and situation, coordinating interdependent work, communicating effectively, and recovering from setbacks. Within this corpus, they formed a stable reconstructable core under the shared zero-shot prompt.
This stable core also corresponds to areas of overlap among major EI traditions. Self-awareness, awareness of others, empathy, and emotion regulation are central to ability, competency, and emotional–social accounts, while psychological resilience and flexibility are consistent with emotional–social emphases on stress management and adaptability. This combination of behavioral prominence and conceptual familiarity may help explain their recovery across all three outputs, and Figure 4 shows that most of the frequently coded dimensions were fully recovered while the divergent dimensions were generally lower-frequency boundary constructs. The high agreement among outputs on directly recovered dimensions (Figure 5) points to the same stable core. The pattern is compatible with prior qualitative comparisons in which LLMs reproduced broad human themes while differing in nuance and organization [15,16,17]. It nevertheless demonstrates construct recovery, not equivalence of analytic process: the human framework resulted from line-by-line coding, constant comparison, expert consultation, and saturation analysis, whereas the LLMs generated complete frameworks from the supplied corpus and prompt.
Guidance and motivation varied most across outputs: it was directly mapped in the DeepSeek output, absent from the Qwen output, and partially represented in the GPT output through a broader theme concerning trust, support, and cooperation. Big-picture awareness and care and support orientation were each directly mapped in two outputs and absent from one, while teamwork was represented in all three but merged into broader cooperation themes in two. These dimensions lie closer to the contested boundary between EI and context-specific social or organizational competence. They extend EI toward collective responsibility, sustained relational concern, teamwork, cohesion, and interpersonal influence, making their separation more dependent on the analyst’s conceptual framing.
The contrast between stable and output-dependent dimensions is therefore substantive rather than numerical alone. Emotion regulation, awareness of others, flexibility, communication, and organizational coordination were commonly expressed through visible actions and interactional cues. Guidance and motivation, care and support orientation, and big-picture awareness required finer distinctions among influence, concern, collective responsibility, and shared goals. The human reference framework thus identified both where the LLM outputs converged and where a broad theme preserved general meaning while merging or omitting a more specific construct boundary.
5.2. Thematic Granularity and Framework Coverage
The three outputs differed substantially in thematic granularity. DeepSeek produced 22 second-order themes, GPT produced 18, and Qwen produced 13. These counts did not translate directly into framework coverage. DeepSeek had the highest direct coverage (91.7%), whereas GPT achieved the highest broad coverage (100.0%) despite producing fewer themes. Qwen generated the smallest theme inventory and had the lowest direct and broad coverage, but still mapped at least partially to ten of the 12 human dimensions.
The comparison therefore distinguishes output size from semantic coverage. A larger theme inventory may separate concepts that another output combines, but additional themes do not guarantee that every human-defined dimension will be recovered. Conversely, a more compact output may cover a dimension only by merging it into a broader construct. Theme count is thus a measure of organizational granularity rather than a standalone measure of qualitative quality.
5.3. Roles of Human Researchers and LLMs
This study compares the products of human and LLM analysis; it does not ask whether LLMs possess or simulate emotional intelligence. Human researchers defined the research problem, designed the interviews, interpreted behavioral events in context, compared positive and negative cases, and decided how the constructs should be defined. The two senior experts also developed the human framework and mapped the 36 human–LLM relationships. The LLMs worked from the transcripts and the shared prompt. They identified, organized, and summarized patterns, but they did not determine the research purpose or make the final theoretical decisions.
The human framework served as the comparison reference and was itself shaped by interview evidence, existing EI theory, expert consultation, and contextual knowledge. Agreement indicates that an output recovered a human-derived dimension from the corpus; omission or merging identifies a boundary that was less clearly preserved; and model-specific themes provide alternative abstractions for human review. The comparison therefore supports interpretive evaluation rather than a simple correct–incorrect judgment.
This distinction leads to a clear division of labor. LLMs can provide parallel analyses, identify possible gaps in a human codebook, highlight themes that may have been divided too narrowly or combined too broadly, and speed up the organization of large text collections. Human researchers remain responsible for defining constructs, interpreting context, verifying evidence, examining negative cases, protecting the data, and deciding the final theoretical account. This division of labor is consistent with recent human-centered LLM workflows in qualitative research [18,19,20,30].
5.4. Methodological Implications for Human–LLM Qualitative Data Analysis
From a data-analysis perspective, the findings support a two-layer approach to human–LLM qualitative comparison. Output structure should first be separated from framework alignment because theme counts describe granularity but not construct recovery. Semantic alignment should then be assessed at the dimension level using theme names, definitions, and supporting context rather than labels alone. These layers answer different questions and should not be collapsed into a single performance score.
Prior human–LLM qualitative studies have commonly summarized similarity through Jaccard coefficients, cosine similarity, expert acceptance rates, or agreement statistics [16,18]. Such measures are useful for overall similarity but do not necessarily distinguish an exact reconstruction of a human-defined dimension from its incorporation into a broader LLM theme. The three summaries used here retain the distinctions recorded in the mapping matrix rather than collapsing them into one undifferentiated similarity estimate.
The two-expert consensus procedure also demonstrates how human judgment can be made explicit within this workflow. The experts applied a stated mapping rubric and resolved interpretive differences through discussion. This approach does not convert qualitative judgment into an objective benchmark, but it makes the comparison more transparent. The methodological contribution is therefore not a ranking of models; it is a framework-level mapping procedure for examining what different zero-shot LLM outputs reconstruct and how they organize that reconstruction relative to a human-derived framework.
5.5. Limitations and Future Directions
First, each model was represented by one zero-shot output, so the present comparison does not show whether the same framework structure or dimension mappings would recur across repeated generations. Future studies should generate repeated outputs for each model–prompt condition and compare the stability of theme inventories and dimension-level mappings across runs.
Second, the human framework and the human–LLM mapping matrix remain expert interpretations, even though two senior experts reviewed the corpus, conducted the analysis, and reached consensus throughout the study. Because the same experts developed the reference framework and mapped the LLM outputs, their dual role may have favored the human framework’s construct boundaries. Future research should invite independent expert teams to reproduce the mapping process and should report agreement before consensus using a statistic appropriate to the coding scale and design [39]. This would distinguish agreement produced by independent interpretation from agreement reached through discussion.
Third, the human and LLM analyses differed in theoretical constraint and corpus use. The human analysis used grounded-theory procedures, a construction–saturation split, prior EI knowledge, and expert consultation, whereas each LLM received the full corpus under a zero-shot prompt that prohibited imposing an existing EI model. The results therefore describe recovery of the human framework rather than equivalence of analytic process. Future controlled studies should use standardized corpus partitions and compare theory-informed and theory-naive instructions directly.
Fourth, the comparison used one context-specific Chinese interview corpus and one human reference framework. The dimensions that appeared stable here may not be equally recoverable in other languages, populations, or substantive domains. Future research should repeat the same zero-shot framework-comparison procedure across multiple interview corpora and use independent expert teams to examine cross-domain generalizability.
6. Conclusions
This study compared a human-derived 12-dimension EI framework with three LLM qualitative reconstructions generated from the same interview corpus under a shared zero-shot prompt. Two senior experts jointly conducted the framework comparison. For RQ1, all three outputs recovered substantial portions of the human framework, with direct coverage ranging from 75.0% to 91.7% and broad coverage ranging from 83.3% to 100.0%. For RQ2, eight dimensions were directly mapped in all three archived outputs, while guidance and motivation varied most across outputs; big-picture awareness, care and support orientation, and teamwork also differed in mapping strength. The shared dimensions largely reflected constructs common to major EI traditions, whereas the variable dimensions marked context-dependent extensions into collective responsibility, relational care, teamwork, and interpersonal influence.
The study demonstrates a framework-level mapping procedure that evaluates thematic granularity and dimension alignment separately. The central methodological conclusion is that the number of generated themes is not equivalent to recovery of human-defined construct boundaries. Repeated generations, independent expert replication, and cross-domain corpora are needed to establish stability and generalizability. Until then, LLMs are best used as comparative assistants, while human researchers retain responsibility for construct boundaries, contextual interpretation, and the final theoretical account.
Author Contributions
Conceptualization, Yuhan Xie; Methodology, Yuhan Xie; Software, Yuhan Xie and Shifeng Lei; Validation, Yuhan Xie and Shifeng Lei; Formal analysis, Yuhan Xie; Investigation, Yuhan Xie, Shifeng Lei and Guang Yang; Resources, Yuhan Xie and Shifeng Lei; Data curation, Yuhan Xie and Shifeng Lei; Writing – original draft, Yuhan Xie and Shifeng Lei; Writing – review & editing, Yuhan Xie, Shifeng Lei, Guang Yang and Feng Yao; Visualization, Guang Yang; Supervision, Guang Yang and Feng Yao; Project administration, Guang Yang and Feng Yao; Funding acquisition, Guang Yang and Feng Yao. All authors have read and agreed to the published version of the manuscript.
Funding
This research received no external funding.
Data Availability Statement
The original contributions presented in this study are included in the supplementary material. Further inquiries can be directed to the corresponding authors.
Conflicts of Interest
The authors declare no conflicts of interest.
Appendix A. Interview Guide and Zero-Shot Prompt
Appendix A.1. Interview Guide
The interview guide is as follows:
- Subject to the requirements of confidentiality, please briefly introduce your basic background.
- What role do you think emotional intelligence plays in your work and daily life, and what aspects does it include?
- Please describe one or two events in your work or daily life that you believe best demonstrate high emotional intelligence. Where possible, describe in detail the situation, task, behavior, and outcome.
- Which personal abilities, qualities, or behaviors do you think demonstrate high emotional intelligence?
- Please describe one or two events in your work or daily life that you believe best demonstrate low emotional intelligence. Where possible, describe in detail the situation, task, behavior, and outcome.
- Which personal abilities, qualities, or behaviors do you think demonstrate low emotional intelligence?
- In your view, which areas should be addressed when developing emotional intelligence?
Appendix A.2. Zero-Shot Prompt
Please analyze the following 34 verbatim interview transcripts concerning emotional intelligence. Identify first-order and second-order themes related to emotional intelligence, and provide a concise definition for each theme. For each theme, provide two verbatim quotations from the interview transcripts, preceded by the relevant participant identifier and corresponding line number. Induce the themes only from the current interview texts; do not use, infer, or impose any existing emotional-intelligence model or dimensions. Do not rewrite, supplement, or fabricate quotations. If textual evidence is insufficient, state this explicitly.
References
- Salovey, P.; Mayer, J.D. Emotional intelligence. Imagin. Cogn. Personal. 1990, 9, 185–211. [Google Scholar] [CrossRef]
- Mayer, J.D.; Salovey, P.; Caruso, D.R. Emotional intelligence: Theory, findings, and implications. Psychol. Inq. 2004, 15, 197–215. [Google Scholar] [CrossRef] [PubMed]
- Mayer, J.D.; Roberts, R.D.; Barsade, S.G. Human abilities: Emotional intelligence. Annu. Rev. Psychol. 2008, 59, 507–536. [Google Scholar] [CrossRef] [PubMed]
- Boyatzis, R.E. Competencies as a behavioral approach to emotional intelligence. J. Manag. Dev. 2009, 28, 749–770. [Google Scholar] [CrossRef]
- Bar-On, R. The Bar-On model of emotional-social intelligence (ESI). Psicothema 2006, 18 (Suppl.), 13–25. Available online: https://pubmed.ncbi.nlm.nih.gov/17295953/. [PubMed]
- Petrides, K.V.; Pita, R.; Kokkinaki, F. The location of trait emotional intelligence in personality factor space. Br. J. Psychol. 2007, 98, 273–289. [Google Scholar] [CrossRef] [PubMed]
- Ashkanasy, N.M.; Daus, C.S. Rumors of the death of emotional intelligence in organizational behavior are vastly exaggerated. J. Organ. Behav. 2005, 26, 441–452. [Google Scholar] [CrossRef]
- Cherniss, C. Emotional intelligence: Toward clarification of a concept. Ind. Organ. Psychol. 2010, 3, 110–126. [Google Scholar] [CrossRef]
- Lim, W.M. What is qualitative research? An overview and guidelines. Australas. Mark. J. 2025, 33(2), 199–229. [Google Scholar] [CrossRef]
- Braun, V.; Clarke, V. Using thematic analysis in psychology. Qual. Res. Psychol. 2006, 3, 77–101. [Google Scholar] [CrossRef]
- Chapman, A.L.; Hadfield, M.; Chapman, C.J. Qualitative research in healthcare: An introduction to grounded theory using thematic analysis. J. R. Coll. Physicians Edinb. 2015, 45, 201–205. [Google Scholar] [CrossRef] [PubMed]
- Glaser, B.G. The constant comparative method of qualitative analysis. Soc. Probl. 1965, 12, 436–445. [Google Scholar] [CrossRef]
- Glaser, B.G.; Strauss, A.L. The Discovery of Grounded Theory: Strategies for Qualitative Research; Aldine: Chicago, IL, USA, 1967. [Google Scholar]
- Daly, J.; Lumley, J. Bias in qualitative research designs. Aust. N. Z. J. Public Health 2002, 26, 299–300. [Google Scholar] [CrossRef] [PubMed]
- De Paoli, S. Performing an inductive thematic analysis of semi-structured interviews with a large language model: An exploration and provocation on the limits of the approach. Soc. Sci. Comput. Rev. 2024, 42, 997–1019. [Google Scholar] [CrossRef]
- Mathis, W.S.; Zhao, S.; Pratt, N.; Weleff, J.; De Paoli, S. Inductive thematic analysis of healthcare qualitative interviews using open-source large language models: How does it compare to traditional methods? Comput. Methods Programs Biomed. 2024, 255, 108356. [Google Scholar] [CrossRef] [PubMed]
- Wachinger, J.; Bärnighausen, K.; Schäfer, L.N.; Scott, K.; McMahon, S.A. Prompts, pearls, imperfections: Comparing ChatGPT and a human researcher in qualitative data analysis. Qual. Health Res. 2025, 35(9), 951–966. [Google Scholar] [CrossRef] [PubMed]
- Qiao, S.; Fang, X.; Wang, J.; Zhang, R.; Li, X.; Kang, Y. Generative AI for thematic analysis in a maternal health study: Coding semistructured interviews using large language models. Appl. Psychol. Health Well-Being 2025, 17, e70038. [Google Scholar] [CrossRef] [PubMed]
- Than, N.; Fan, L.; Law, T.; Nelson, L.K.; McCall, L. Updating “The Future of Coding”: Qualitative coding with generative large language models. Sociol. Methods Res. 2025, 54, 849–888. [Google Scholar] [CrossRef]
- Yue, Y.; Liu, D.; Lv, Y.; Hao, J.; Cui, P. A practical guide and assessment on using ChatGPT to conduct grounded theory: Tutorial. J. Med. Internet Res. 2025, 27, e70122. [Google Scholar] [CrossRef] [PubMed]
- Li, K.D.; Fernandez, A.M.; Schwartz, R.; Rios, N.; Carlisle, M.N.; Amend, G.M.; Patel, H.V.; Breyer, B.N. Comparing GPT-4 and human researchers in health care data analysis: Qualitative description study. J. Med. Internet Res. 2024, 26, e56500. [Google Scholar] [CrossRef] [PubMed]
- Bijker, R.; Merkouris, S.S.; Dowling, N.A.; Rodda, S.N. ChatGPT for automated qualitative research: Content analysis. J. Med. Internet Res. 2024, 26, e59050. [Google Scholar] [CrossRef] [PubMed]
- Barros, C.F.; Azevedo, B.B.; Graciano Neto, V.V.; Kassab, M.; Kalinowski, M.; do Nascimento, H.A.D.; Bandeira, M.C.G.S.P. Large language model for qualitative research: A systematic mapping study. In In Proceedings of the 2025 IEEE/ACM International Workshop on Methodological Issues with Empirical Studies in Software Engineering (WSESE), 2025; pp. 48–55. [Google Scholar] [CrossRef]
- Lee, J.; Park, S.; Shin, J.; Cho, B. Analyzing evaluation methods for large language models in the medical field: A scoping review. BMC Med. Inform. Decis. Mak. 2024, 24, 366. [Google Scholar] [CrossRef] [PubMed]
- Kempny, C.; Frings, J.; Rust, P.; Meister, S.; Fehring, L. The use and methodological reporting of large language models in qualitative research: A scoping review. BMC Med. Res. Methodol. 2026, 26, 137. [Google Scholar] [CrossRef] [PubMed]
- Han, Z.; Battaglia, F.; Mansuria, K.; Heyman, Y.; Terlecky, S.R. Beyond text generation: Assessing large language models’ ability to reason logically and follow strict rules. AI 2025, 6, 12. [Google Scholar] [CrossRef]
- Jowsey, T.; Braun, V.; Clarke, V.; Lupton, D.; Fine, M. We reject the use of generative artificial intelligence for reflexive qualitative research. Qual. Inq. 2025. [Google Scholar] [CrossRef]
- Shanwetter Levit, N.; Saban, M. When investigator meets large language models: A qualitative analysis of cancer patient decision-making journeys. npj Digit. Med. 2025, 8, 336. [Google Scholar] [CrossRef] [PubMed]
- Reis, T.; Dumberger, L.; Bruchhaus, S.; Krause, T.; Schreyer, V.; Bornschlegl, M.X.; Hemmje, M.L. AI-based user empowerment for empirical social research. Big Data Cogn. Comput. 2024, 8, 11. [Google Scholar] [CrossRef]
- Zeng, Y.; Chen, Y.; Yu, J.; Yang, J.; Fu, Y.; Chen, J.; Guan, Q.; He, H.-G. Advancing qualitative analysis in nursing research: Comparing text analysis rigor and efficiency of large language models with human-coded analysis for interview data. Int. J. Nurs. Stud. 2026, 182, 105584. [Google Scholar] [CrossRef] [PubMed]
- Cantini, A.; De Mauro, A. Better prompts, better usefulness: A systematic review and experimental evaluation of structured prompting techniques in large language models. Big Data Cogn. Comput. 2026, 10, 224. [Google Scholar] [CrossRef]
- Zhou, Y.; Yuan, Y.; Huang, K.; Hu, X. Can ChatGPT perform a grounded theory approach to do risk analysis? An empirical study. J. Manag. Inf. Syst. 2024, 41, 982–1015. [Google Scholar] [CrossRef]
- Goyanes, M.; Lopezosa, C.; Jordá, B. Thematic analysis of interview data with ChatGPT: Designing and testing a reliable research protocol for qualitative research. Qual. Quant. 2025, 59, 5491–5510. [Google Scholar] [CrossRef]
- Hitch, D. Artificial intelligence augmented qualitative analysis: The way of the future? Qual. Health Res. 2024, 34, 595–606. [Google Scholar] [CrossRef] [PubMed]
- Dunivin, Z.O. Scaling hermeneutics: A guide to qualitative coding with LLMs for reflexive content analysis. EPJ Data Sci. 2025, 14, 28. [Google Scholar] [CrossRef]
- Romero, J.D.; Feijoo-Garcia, M.A.; Nanda, G.; Newell, B.; Magana, A.J. Evaluating the performance of topic modeling techniques with human validation to support qualitative analysis. Big Data Cogn. Comput. 2024, 8, 132. [Google Scholar] [CrossRef]
- Sakaguchi, K.; Sakama, R.; Watari, T. Evaluating ChatGPT in qualitative thematic analysis with human researchers in the Japanese clinical context and its cultural interpretation challenges: Comparative qualitative study. J. Med. Internet Res. 2025, 27, e71521. [Google Scholar] [CrossRef] [PubMed]
- Alostad, H. Large language models as Kuwaiti annotators. Big Data Cogn. Comput. 2025, 9, 33. [Google Scholar] [CrossRef]
- McHugh, M.L. Interrater reliability: The kappa statistic. Biochem. Medica 2012, 22, 276–282. [Google Scholar] [CrossRef]
Figure 1.
Overall study framework from data collection to human–LLM comparison.

Figure 2.
Thematic structure and framework alignment across three LLM outputs.

Figure 3.
Dimension-level mapping across three LLM outputs.

Figure 4.
Human coding frequency versus mean consensus mapping across the three zero-shot outputs. Each point is one human-derived dimension, positioned by its coding frequency and its mean mapping score (50.0%, 66.7%, or 100.0%) and colored by higher-order category.
Figure 4.
Human coding frequency versus mean consensus mapping across the three zero-shot outputs. Each point is one human-derived dimension, positioned by its coding frequency and its mean mapping score (50.0%, 66.7%, or 100.0%) and colored by higher-order category.

Figure 5.
Pairwise Jaccard similarity of the directly recovered human dimensions across the three zero-shot outputs.
Figure 5.
Pairwise Jaccard similarity of the directly recovered human dimensions across the three zero-shot outputs.

Table 1.
Major emotional-intelligence traditions.
| Tradition | Core view of EI | Representative components | Relevance to the human framework | Key sources |
|---|---|---|---|---|
| Ability model | Ability to process emotion-related information | Perceiving, using, understanding, and managing emotion | Closely related to self-awareness, awareness of others, empathy, and emotion regulation | [1,2,3] |
| Competency model | Observable emotional and social competencies in effective behavior | Self-awareness, self-management, social awareness, and relationship management | Extends toward communication, coordination, teamwork, and guidance and motivation | [4] |
| Emotional–social model | Integration of emotional functioning, social functioning, adaptation, and stress management | Intrapersonal and interpersonal functioning, adaptability, and stress management | Provides conceptual links to psychological resilience, flexibility, and care and support orientation | [5] |
| Trait model | Emotion-related self-perceptions within personality | Emotionality, self-control, sociability, and well-being | Overlaps with several labels, but differs from the behavioral-event basis of the present framework | [6] |
Table 2.
Human emotional-intelligence reference framework.
| Higher-order category | Dimension | Working definition | Frequency |
|---|---|---|---|
| Self-awareness and regulation | Self-awareness | Understanding one’s traits, abilities, current role, learning needs, and emotional state. | 20 |
| Psychological resilience | Facing setbacks constructively, recovering through self-motivation, and sustaining effort under difficult conditions. | 20 | |
| Emotion regulation | Controlling impulses, avoiding uncontrolled anger, maintaining rationality, and stabilizing emotion during interaction. | 30 | |
| Big-picture awareness | Demonstrating responsibility, initiative, attention to shared goals, and the ability to balance collective and personal interests. | 28 | |
| Environmental perception and adaptation | Empathy | Taking another person’s perspective, considering their feelings, and understanding the intention behind their behavior. | 46 |
| Awareness of others | Detecting emotional change through expression, gaze, tone, and behavior while understanding personal circumstances and preferences. | 35 | |
| Flexibility | Judging one’s role and setting, adapting communication and action to changing conditions, and avoiding rigid thinking. | 36 | |
| Care and support orientation | Caring for and supporting peers and collaborators and sustaining constructive relationships through sincere emotional concern. | 18 | |
| Communication, coordination, and influence | Guidance and motivation | Building moral credibility and trust, creating cohesion, influencing emotion, and stimulating others’ initiative. | 33 |
| Organizational coordination | Coordinating people and resources across interdependent roles, resolving conflicts, and maintaining working relationships. | 44 | |
| Teamwork | Mobilizing team members, drawing on collective knowledge, allocating work, and cooperating to complete shared tasks. | 18 | |
| Communication | Expressing intentions accurately, speaking logically and tactfully, and using language acceptable to the listener and setting. | 28 |
Table 3.
Output-level human–LLM mapping results.
| Output | 1st | 2nd | Direct | Partial | Missing | Direct (%) | Broad (%) | Weighted (%) |
|---|---|---|---|---|---|---|---|---|
| DeepSeek zero-shot | 10 | 22 | 11 | 0 | 1 | 91.7 | 91.7 | 94.9 |
| Qwen zero-shot | 6 | 13 | 9 | 1 | 2 | 75.0 | 83.3 | 80.3 |
| GPT zero-shot | 6 | 18 | 10 | 2 | 0 | 83.3 | 100.0 | 92.8 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.