Submitted:
20 September 2026
Posted:
21 September 2026
You are already at the latest version
Abstract
This developmental paper investigates how simulation-based learning mechanisms translate across diverse management courses. Preliminary data from a summer 2026 pilot (n=29, single leadership course) suggested that consequence-driven simulations paired with personalized feedback can expose gaps between students’ conceptual knowledge and enacted judgment. Building on these findings, we propose a developmental process in which consequential simulation experiences create productive struggle, struggle makes performance-self-perception discrepancies visible, and credible feedback paired with structured reflection supports mental-model revision and adaptive judgment. The current study scales this design across six courses in Fall 2026 and seven courses in Spring 2027, examining whether this process generalizes beyond leadership education to functional areas including human resources, employment law, organizational behavior, ethics, organizational development, and training and development. We further examine whether task characteristics such as ambiguity, rule boundedness, and relational complexity shape when simulated struggle becomes productive. Data collection is underway in Fall 2026; we present the research design, theoretical framework, and preliminary summer 2026 findings that inform the expanded implementation. This work examines how AI-enabled pedagogies may close theory-practice gaps while surfacing challenges around feedback credibility, equity of access, and the boundaries of simulation fidelity across management sub-disciplines.
Keywords:
simulation-based learning
; management education
; AI-enabled pedagogies
; productive struggle
; consequence-driven simulations
; personalized feedback
; mental-model revision
; adaptive judgment
; theory-practice gaps
; simulation fidelity
; leadership education
; organizational behavior
Introduction: The Translation Problem Across the Management Curriculum
Management education faces a persistent challenge: students demonstrate conceptual mastery yet struggle to enact that knowledge under pressure (Day & Dragoni, 2015). This knowing-doing gap appears across sub-disciplines - students can define employment law concepts but freeze when conducting termination conversations; they understand motivation theory yet default to directive approaches when managing underperformance. Traditional pedagogies privilege declarative knowledge, creating what we term "curriculum-specific brittleness": learning remains siloed within course boundaries rather than integrating into adaptive professional judgment.
AI-driven immersive and gamified simulations represent an emerging experiential technology that may address this brittleness by creating consequence-driven learning environments where decisions trigger immediate, emotionally realistic stakeholder responses (Lateef, 2010; Salas et al., 2009). Our summer 2026 pilot examined this mechanism in undergraduate leadership education, finding that six sequential simulations paired with structured reflection produced three outcomes: characteristic dip-then-climb learning trajectories consistent with productive failure frameworks (Kapur, 2016), personalized career profiles that functioned as "mirrors with teeth" surfacing gaps between self-perception and behavior, and changes in how students approached uncertain management problems, including greater ambiguity tolerance and more adaptive judgment.
That pilot raised a critical question: Do these mechanisms generalize beyond leadership-specific contexts to the broader management curriculum? Leadership simulations privilege interpersonal dynamics - stakeholder emotions, trust-building, and communication style. Do simulations addressing employment law compliance, HR process design, or ethical decision frameworks produce similar developmental arcs? Or do task and simulation characteristics such as ambiguity, rule boundedness, relational complexity, emotional salience, temporal dynamics, consequence immediacy, and the number of defensible solutions shape how students experience productive struggle, interpret algorithmic feedback, and transfer learning across contexts?
This research proposal reports on an ongoing multi-course implementation scaling the original design from one leadership class (summer 2026) to thirteen course sections across six disciplines (fall 2026 and spring 2027). We examine whether simulation-based learning mechanisms replicate, adapt, or fail when exported beyond their origin context.
Theoretical Framework: From Productive Struggle to Adaptive Judgment
Consequential Simulation and Productive Struggle
Kapur’s (2008, 2016) productive failure framework challenges the assumption that learning environments should minimize errors. Instead, learners may benefit from encountering complex problems before receiving expert solutions because struggle can activate prior knowledge, expose the limits of existing approaches, and create receptivity to new frameworks (Kapur & Bielaczyc, 2012). In the present study, we distinguish productive struggle from a simple decline in performance. A lower simulation score indicates that an approach may not have worked, but failure becomes developmentally meaningful when the experience reveals that a learner’s existing knowledge, assumptions, or behavioral strategies are insufficient for the situation.
AI-enabled simulations create this opportunity by requiring students to enact management knowledge in consequence-rich situations rather than simply recognize or describe correct concepts. Students must diagnose problems, respond to stakeholders, and make decisions under uncertainty while observing the consequences of those decisions. In the summer pilot, early performance declines often occurred when students attempted to adjust after initial feedback. We treat these dip-then-climb patterns as possible indicators of productive struggle rather than as evidence of productive failure. The central theoretical question is whether struggle exposes limitations in students’ existing mental models and, with appropriate feedback and reflection, leads them to revise how they approach subsequent situations.
Task Characteristics as Boundary Conditions
Transfer research distinguishes near transfer (applying learning to similar contexts) from far transfer (applying learning to dissimilar contexts). Rather than treating course membership itself as the theoretical moderator, we focus on characteristics of the simulation task that vary across management domains. These include ambiguity, rule boundedness, relational complexity, emotional salience, temporal dynamics, consequence immediacy, and the number of defensible solutions. These characteristics may influence the strength of earlier stages of the developmental process and the credibility of feedback.
Task structure: Employment law scenarios often involve applying rules and procedures, whereas ethical dilemmas may involve competing values with multiple defensible responses. Greater ambiguity may make existing assumptions more visible because students cannot rely on a single known rule, while more structured tasks may make diagnostic feedback clearer.
Temporal dynamics and consequence immediacy: Training design decisions (e.g., needs assessment, learning objectives) may represent consequences that unfold over longer periods, whereas conflict de-escalation decisions can generate immediate stakeholder responses. The timing and immediacy of consequences may affect how easily students connect their decisions to outcomes and recognize that their existing approach was insufficient.
Relational complexity and emotional salience: Some simulations center on processes or policies, while others require students to respond to specific stakeholders with visible emotional states. Relational complexity and emotional salience may make performance-self-perception discrepancies especially noticeable when students believe they are empathetic, collaborative, or effective communicators but their enacted behavior produces different outcomes.
These dimensions create theoretically meaningful boundary conditions for the developmental process. Fall 2026 courses span this design space, allowing us to examine whether the same underlying learning process operates across domains and whether task characteristics strengthen or weaken the earlier stages of that process.
Feedback, Reflection, and Mental-Model Revision
The summer pilot suggests that one important developmental moment occurs when students recognize a discrepancy between how they believe they approach management situations and how they actually perform in them. The SKIVE career profile (Skills, Knowledge, Identity, Values, Epistemology) functioned as an external mirror that made behavioral patterns visible. Students used profile language to articulate gaps between perceived strengths and enacted behavior (e.g., strong analysis but weaker listening or patience). Self-discrepancy theory suggests that inconsistencies involving individuals’ self-perceptions can create psychologically meaningful discrepancy signals that motivate self-regulation (Higgins, 1987). Extending this logic to the simulation context, we refer to performance-self-perception discrepancy as the perceived inconsistency between how learners believe they will perform and evidence of how they actually behave in the simulation. Recognition of this discrepancy may prompt learners to reconsider their assumptions about their capabilities and identify areas for development.
Feedback and reflection play different roles in transforming this discrepancy into learning. Diagnostic feedback helps students identify what happened, where their approach was ineffective, and why; structured reflection helps them consider what assumptions drove their response and what they should do differently. Together, these processes can support mental-model revision - a change in how students understand, diagnose, or approach a management problem. However, this process depends on perceived feedback credibility. When students view AI-generated feedback as accurate, transparent, and fair, they are more likely to accept and use that feedback to make sense of performance discrepancies and revise their mental models. When credibility is low, the same discrepancy may produce defensiveness or rejection. Thus, perceived feedback credibility is positioned as a moderator within the developmental process rather than simply as an outcome of the simulation experience.
Figure 1 in the Appendix summarizes the proposed theoretical process. We argue that AI-enabled simulations create consequential experiences that can generate productive struggle when students’ existing knowledge, assumptions, or strategies prove insufficient. This struggle can make discrepancies between students’ self-perceptions and enacted performance visible. Diagnostic feedback helps students understand what occurred, while structured reflection helps them examine why they responded as they did and what underlying assumptions or strategies may need to change. When feedback is perceived as credible, this process is more likely to result in mental-model revision and, ultimately, more adaptive judgment and transfer to subsequent situations. Task and simulation characteristics including ambiguity, rule boundedness, relational complexity, emotional salience, temporal dynamics, consequence immediacy, and the number of defensible solutions serve as boundary conditions that may strengthen or weaken earlier stages of this developmental process (see Figure 1).
Research Questions and Hypotheses
RQ1: Do consequential simulation experiences produce similar patterns of productive struggle across management domains, and how do task characteristics shape those trajectories?
H1: Task characteristics that increase ambiguity, relational complexity, emotional salience, consequence immediacy, and the availability of multiple defensible solutions will be associated with greater productive struggle and stronger recognition of performance-self-perception discrepancies, whereas greater rule boundedness may constrain these effects.
RQ2: How do diagnostic feedback and structured reflection help students translate recognized performance-self-perception discrepancies into mental-model revision?
H2: Perceived feedback credibility will strengthen the relationship between discrepancy recognition and mental-model revision, such that students who perceive AI-generated feedback as more accurate, transparent, and fair will demonstrate greater mental-model revision.
RQ3: Does mental-model revision lead to more adaptive judgment across subsequent simulations and novel management situations?
H3: Students who demonstrate greater mental-model revision will show stronger subsequent performance improvement and greater evidence of adaptive judgment and transfer.
Exploratory Question: Which task and simulation characteristics enable productive struggle and strengthen the developmental process across diverse management contexts?
Method
Sample and Context
Data are being collected across thirteen course sections spanning fall 2026 (six courses, n ≈ 150) and spring 2027 (seven courses, n ≈ 180). Courses include Introduction to HR, Employment Law, Organizational Behavior, Ethical Decision-Making in Organizations, Organizational Development and Change Management, and Training and Development. All courses are undergraduate; student populations include sophomores through seniors across primarily various business and social science majors.
Simulation Design
Each course includes six AI-driven simulations spaced evenly across the term. Simulations present branching workplace scenarios requiring diagnosis, stakeholder management, and decision-making under uncertainty. AI-controlled characters respond to student communication based on emotional state modeling and natural language processing. Key features include adaptive dialogue, visible emotional states, Otto the AI coach prompting perspective-taking, SKIVE career profiles tracking performance dimensions, and immediate post-simulation feedback. Because cross-course comparisons depend on score meaning, the expanded study will also document how performance dimensions are scored and assess the extent to which scores are calibrated and comparable across simulations and course contexts.
Simulations are customized to course content while maintaining structural consistency:
- Employment Law: Scenarios require applying discrimination, harassment, and wrongful termination statutes to ambiguous fact patterns (e.g., performance termination of protected-class employee, ADA accommodation requests).
- Organizational Behavior: Scenarios require applying motivation, group dynamics, and change theories to team performance and conflict situations.
- Ethics: Scenarios present values dilemmas (e.g., transparency versus loyalty, fairness versus efficiency) with no clear right answer.
- HR & Training: Scenarios require designing selection systems, appraisal processes, or training programs based on strategic needs analysis.
- OD/Change: Scenarios require diagnosing organizational problems and designing interventions under resistance and political complexity.
Data Collection
Students complete a structured reflection at term’s end, responding to prompts about performance trajectories, career profile interpretation, learning mechanisms, theory-practice connections, and concrete developmental gains. Reflection length ranges 450-1,600 words (median 850 words). Reflections are graded for depth of analysis but not performance quality, encouraging candid self-assessment.
Additional data include:
- Platform-generated performance scores across simulations
- Pre-post self-report surveys measuring leadership self-efficacy, comfort with ambiguity, listening orientation, adaptive style endorsement, and feedback receptivity
- Mid-term profile interpretation exercises where students analyze SKIVE dimensions and set developmental goals
Analysis Plan
Qualitative analysis will use a hybrid deductive-inductive coding approach applied to student reflections. Theory-driven codes will capture productive struggle, discrepancy recognition, feedback credibility, reflection, mental-model revision, and adaptive judgment, while inductive coding will identify unanticipated learning mechanisms and barriers. Codes will be compared across task and course contexts to distinguish common processes from domain-specific patterns. Because reflections are graded for depth, findings from reflections will be interpreted alongside platform performance and survey data rather than treated as stand-alone evidence of learning.
Quantitative analysis will examine:
- Performance trajectories: Model repeated simulation scores over time to examine individual learning trajectories, including early struggle and subsequent recovery, rather than relying only on dip-then-climb categories.
- Discrepancy and mental-model revision: Examine whether students who recognize larger gaps between self-perception and enacted performance show greater evidence of revised assumptions, strategies, and subsequent adaptive judgment.
- Feedback credibility: Test whether perceived credibility of AI-generated feedback strengthens the relationship between discrepancy recognition and mental-model revision, and whether credibility varies with task or simulation characteristics.
- Prediction models: Does simulation performance predict final grades? Do trajectory characteristics (early dip magnitude, recovery slope) predict learning outcomes differently across courses?
Summer 2026 Pilot Findings: Foundation for Multi-Course Scaling
The summer 2026 leadership course pilot (n=29) established baseline findings that inform the expanded implementation. Three primary outcomes emerged from qualitative analysis of student reflections paired with self-report trajectory data.
The Productive Struggle Trajectory
Forty-one percent of students reported dip-then-climb patterns - initial scores of approximately 58, a dip to 55 in Simulation 2, then steady climb through 64, 70, and 74 to approximately 79 by Simulation 6. Students repeatedly described overcorrecting after initial feedback: "The feedback basically showed that I probe well, but I commit late" (Student 23). The dip proved pedagogically valuable; several students explicitly named it as their most instructive moment: "The dip in my third simulation taught me more than the improvement in my sixth" (Student 6).
The Mirror with Teeth
Twenty-four students (83%) referenced SKIVE profile data as revelatory: "This term’s simulations were less about learning new leadership theory and more about seeing, with real numbers, exactly where my leadership instincts are solid and where they are still aspirational" (Student 1). The profile surfaced a consistent archetype: analytical strength paired with relational development needs. Sixteen students cited analytical/diagnostic ability as a strength and only three as a growth area, while active listening showed the mirror pattern - five strength, fourteen growth area. Patience showed similar skew (three strength, twelve growth area).
However, trust in algorithmic assessment proved fragile. Two to three students questioned score validity, and this skepticism correlated with minimal reported growth. The threshold between productive discomfort and defensive rejection appeared narrow.
Changes in Thinking and Adaptive Judgment
Comparing students’ descriptions of their Simulation 1 mindset against final reflections revealed measurable belief changes:
- The belief that leadership means having answers fell from 21 students to 4
- Comfort with ambiguity rose from 6 to 19
- Listening before solving rose from 7 to 22
- Viewing one’s style as adaptive rather than fixed rose from 5 to 18
These patterns suggest that simulations may influence not only specific skills but also how students approach uncertain management problems. Some changes, such as greater comfort with ambiguity, may reflect epistemological beliefs about certainty and knowledge; others, such as listening before solving or adapting one’s style, are better understood as changes in behavioral strategy or adaptive judgment. The expanded study therefore treats adaptive judgment as the broader developmental outcome while examining epistemological change as one possible component.
Design Lessons
Three design elements emerged as critical: (1) consequence loops that make decisions visible ("I got to see the effect my decisions made"), (2) emotional realism of AI stakeholders ("Learning to reflectively listen to teammates feeling ’anxious’ or ’defensive’"), and (3) Otto’s Socratic prompting ("Sometimes I wanted Otto to just give me the answer, but instead the conversation made me think from another angle").
The most common barrier was the locked profile dimension mismatch - Identity, Values, and Epistemology required ten simulations while only six were assigned. Students’ appetite for exactly those dimensions tripled in the second half of the sequence, suggesting deeper self-understanding unlocks only after accumulated performance history.
Unanswered Questions Driving Multi-Course Design
The pilot succeeded in a leadership context optimized for interpersonal, ambiguous, relational learning. Three questions remain:
- Will the dip-then-climb trajectory replicate when simulations assess rule application (Employment Law), process design (HR/Training), or theory application (OB)?
- Will students trust algorithmic assessment of technical competencies (legal compliance, needs analysis accuracy) differently than relational competencies (empathy, listening)?
- Will changes in adaptive judgment (including comfort with ambiguity and flexible thinking) transfer across domains, or are they leadership-specific?
The fall/spring multi-course implementation directly tests these questions.
Next Steps: Fall 2026 and Spring 2027 Data Collection
Fall 2026 Implementation (Currently Underway)
Six courses are currently running with simulations integrated:
- Introduction to HR (n≈25)
- Employment Law (n≈30)
- Organizational Behavior (n≈30)
- Ethical Decision-Making in Organizations (n≈25)
- Organizational Development and Change Management (n≈20)
- Training and Development (n≈20)
Each follows the six-simulation structure with end-of-term structured reflections. Pre-post surveys are administered in weeks 1 and 15. Mid-term profile interpretation exercises occur after Simulation 3. Fall data collection will conclude in December 2026, enabling preliminary cross-course analysis before spring implementation.
Spring 2027 Refinements
Spring 2027 will add one section and incorporate design refinements based on fall findings:
- All seven spring courses will include mid-term profile interpretation workshops explicitly teaching students to read SKIVE dimensions as developmental observations rather than performance verdicts, addressing the trust-building challenge identified in summer.
- Employment Law and Training simulations will be revised to increase ambiguity where fall data suggests current scenarios function more as practice problems than developmental dilemmas. For example, employment law scenarios will add fact patterns where multiple legal interpretations are defensible, testing whether productive struggle becomes more salient when right answers are less clear.
- One OB section will serve as a comparison condition removing AI coach (Otto) dialogues to isolate the contribution of perspective-taking prompts versus self-directed reflection.
Where feasible, spring courses will include a novel end-of-term transfer scenario that differs from previously practiced simulations. This will provide a more direct test of whether students can apply revised mental models and adaptive judgment beyond the situations in which they learned them.
Timeline and Analysis Milestones
- December 2026: Complete fall data collection (reflections, surveys, platform performance data)
- January 2027: Preliminary fall analysis focusing on trajectory replication and trust patterns across courses
- March 2027: WAM presentation of preliminary findings and solicitation of developmental feedback
- May 2027: Complete spring data collection
- Summer 2027: Full cross-course comparative analysis
- Fall 2027: Manuscript development integrating both semesters
Anticipated Contributions
This research will contribute to three conversations:
- Productive failure theory: Clarifying when simulated struggle becomes productive by specifying a process through which consequential experience exposes limitations in existing mental models and credible feedback plus reflection supports revision and adaptive judgment.
- Simulation-based learning design: Identifying task characteristics and design features that strengthen or weaken specific stages of the learning process across management sub-disciplines.
- AI in management education: Explaining how perceived feedback credibility shapes whether students use or reject AI-generated developmental feedback, while identifying which AI-supported design features warrant further comparison with simpler or non-AI experiential approaches.
Discussion: Theoretical and Practical Implications
Generalizability as Empirical Question
The summer pilot established proof-of-concept for leadership education. The multi-course expansion treats generalizability as an empirical question rather than an assumption. The central issue is not simply whether one course produces a larger performance dip than another, but whether the same developmental process - simulation experience, productive struggle, discrepancy recognition, diagnostic feedback, reflection, mental-model revision, and adaptive judgment/transfer - operates across different task contexts. If the process varies systematically with ambiguity, rule boundedness, relational complexity, emotional salience, temporal dynamics, consequence immediacy, or the number of defensible solutions, these task characteristics may represent meaningful boundary conditions for simulation-based learning.
Trust and Algorithmic Assessment
The fragility of student trust in algorithmic feedback represents a central moderator in the proposed process. If students dismiss profile data as inaccurate, opaque, or unfair, performance-self-perception discrepancies may lead to defensiveness rather than revision. When feedback is perceived as credible, however, discrepancy can become a developmental signal that prompts reflection and mental-model change. Making assessment criteria transparent and contestable may therefore support both learning and critical AI literacy - students learning to evaluate algorithmic claims rather than accept or reject them wholesale.
Curriculum Integration Challenges
Scaling from one course to thirteen surfaces implementation barriers invisible in pilot contexts. Course scheduling constraints, faculty technology adoption curves, and the need for simulation content aligned to specific learning objectives all moderate feasibility. The spring comparison condition (Otto versus no-Otto) represents one approach to isolating value-added: if outcomes don’t differ, simpler designs may be preferable for widespread adoption.
Conclusions
This research proposal extends AI-driven simulation research from a single leadership course to a multi-course, multi-discipline investigation of how consequential simulation experiences become developmental. The proposed process model argues that simulation alone is not sufficient: productive struggle must expose limitations in learners’ existing approaches, students must recognize meaningful discrepancies between self-perception and enacted performance, and credible feedback paired with structured reflection must support mental-model revision. The expected outcome is adaptive judgment - the ability to modify reasoning and behavior as contextual demands change.
As AI-enabled pedagogies proliferate, understanding the conditions under which they support learning becomes essential. The current study examines whether this developmental process generalizes across management tasks characterized by different levels of ambiguity, rule boundedness, relational complexity, emotional salience, temporal dynamics, consequence immediacy, and numbers of defensible solutions. It also highlights perceived feedback credibility as a critical moderator determining whether AI-generated assessment becomes a developmental mirror or is rejected by learners. Ongoing developmental feedback will inform spring data collection, subsequent analysis, and the translation of findings into actionable design principles for management educators navigating AI-enhanced curricula.
References
- Day, D. V.; Dragoni, L. Leadership development: An outcome-oriented review based on time and levels of analyses. Annual Review of Organizational Psychology and Organizational Behavior 2015, 2, 133–156. [Google Scholar] [CrossRef]
- Higgins, E. T. Self-discrepancy: a theory relating self and affect. Psychological review 1987, 94(3), 319. [Google Scholar] [CrossRef]
- Kapur, M. Productive failure. Cognition and Instruction 2008, 26(3), 379–424. [Google Scholar] [CrossRef]
- Kapur, M. Examining productive failure, productive success, unproductive failure, and unproductive success in learning. Educational Psychologist 2016, 51(2), 289–299. [Google Scholar] [CrossRef]
- Kapur, M.; Bielaczyc, K. Designing for productive failure. Journal of the Learning Sciences 2012, 21(1), 45–83. [Google Scholar] [CrossRef]
- Lateef, F. Simulation-based learning: Just like the real thing. Journal of Emergencies, Trauma, and Shock 2010, 3(4), 348–352. [Google Scholar] [CrossRef] [PubMed]
- Narciss, S. Feedback strategies for interactive learning tasks. In Handbook of research on educational communications and technology, 3rd ed.; Spector, J. M., Merrill, M. D., van Merriënboer, J., Driscoll, M. P., Eds.; Routledge, 2008; pp. 125–144. [Google Scholar]
- Salas, E.; Wildman, J. L.; Piccolo, R. F. Using simulation-based training to enhance management education. Academy of Management Learning & Education 2009, 8(4), 559–573. [Google Scholar] [CrossRef]
Figure 1.
theoretical process model: how AI-enabled simulations produce developmental leaning.

Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.