Preprint
Hypothesis

This version is not peer-reviewed.

Enhancing Engagement in Learners with Special Educational Needs through physical and immersive Musical Toys: A Mixed-Methods Comparative Study Protocol

Submitted:

26 August 2026

Posted:

31 August 2026

You are already at the latest version

Abstract
Musical play gives learners with Special Educational Needs (SEN) an immediate way to act on, perceive, and reorganize a learning environment through sound. This study protocol proposes a randomized, counterbalanced, within-participant mixed-methods comparison of two matched forms of musical interaction: physical/tangible musical toys and virtual/immersive musical environments controlled through touch, gesture, or movement. Engagement is treated as a multidimensional process encompassing behavioral participation and persistence, emotional involvement, and cognitive-agentic investment in exploration, choice, and task completion. A common event ontology will align systematic observation of tangible interaction with time-stamped digital logs, while educator ratings and structured qualitative observations will contextualize individual response patterns. The protocol emphasizes functional equivalence between conditions, accessibility, sensory and motor fit, learner agency, and standardized adult mediation rather than assuming the superiority of either medium. Mixed-effects models will estimate within-learner condition effects and order effects, and integrated qualitative analysis will examine why particular affordances support or constrain engagement for different learners. The resulting framework is intended to inform the design and evaluation of accessible serious games, digital musical instruments, and hybrid physical-digital learning experiences for inclusive education.
Keywords: 
;  ;  ;  ;  ;  ;  ;  
Subject: 
Arts and Humanities  -   Music

1. Introduction

A learner reaches for a drum, a resonant object, or a luminous digital surface and waits for the environment to answer. In that short interval, participation is already more than attention: action has been initiated, a sensory consequence is anticipated, emotion accompanies the result, and a decision is made about whether to continue, vary the action, invite another person, or withdraw. For learners with Special Educational Needs (SEN), such moments are especially informative because engagement cannot be reduced to compliance, stillness, or time spent in front of a device. It is a dynamic relation among the learner, the task, the material or digital medium, the adult who mediates access, and the possibilities for making something happen in ways that remain intelligible and personally meaningful. Recent engagement research accordingly distinguishes behavioral, emotional, cognitive, and agentic forms of involvement, emphasizing that each dimension may serve a different function in learning [1]. Evidence from SEN populations also indicates that engagement is shaped by relational and participation conditions rather than by learner characteristics alone [2].
Music offers an unusually dense setting in which these dimensions can be observed together. Producing a sound requires an action; sustaining a musical exchange requires persistence and temporal coordination; variation requires discrimination and decision-making; shared music invites mutual attention and turn-taking; and the affective qualities of sound can render success, surprise, anticipation, and frustration immediately perceptible. Within inclusive education, musical activity has therefore been approached not merely as curricular content but as a participation environment in which technological mediation can widen the range of feasible actions and expressive forms [3]. A recent systematic review of music lessons for pupils with SEN nevertheless shows that pedagogical support remains unevenly organized and that research still provides limited guidance on how support practices can be matched to heterogeneous learner needs [4]. This methodological deficit matters because a technology may be technically accessible while still failing to produce educationally meaningful participation. The educational value of the interaction is generated by the fit among learner, affordance, task, and mediation, a view consistent with pedagogical accounts that treat inclusion as adaptive regulation rather than as the simple addition of assistive resources [5].
The present protocol focuses on a specific comparison within this broader problem: physical or tangible musical toys versus virtual or immersive musical toys. The word toy is used here in an educational sense to denote an interactive object or environment whose musical rules can be explored through play, not to imply triviality or an absence of instructional intention. Tangible systems include manipulable sound-producing objects and physical interfaces through which learners can strike, press, squeeze, move, rotate, or combine elements. Virtual and immersive systems include digital environments in which musical consequences are generated through touch, gesture, bodily movement, or spatial interaction. Both forms can be playful, technically sophisticated, socially mediated, and instructionally purposeful; what differs is the architecture through which action becomes feedback. The study therefore avoids treating physical and digital media as competing stages of technological progress. Instead, it asks whether different action-feedback architectures yield distinct profiles of engagement and whether those profiles vary in relation to individual access needs.

1.1. Engagement as a Multidimensional and Relational Outcome

Engagement is frequently operationalized through visible participation, yet observable activity alone is an incomplete indicator of educational involvement. A learner can remain physically active while repeating actions without exploration, can appear behaviorally quiet while listening with sustained cognitive investment, or can complete a task under dense prompting while exercising little agency. A multidimensional model allows these possibilities to be distinguished. Behavioral engagement refers in this protocol to active participation, persistence, initiation, turn-taking, and task completion. Emotional engagement concerns observable enjoyment, approach, frustration, avoidance, and willingness to remain in or return to the activity. Cognitive-agentic engagement concerns purposeful exploration, decision-making, strategy variation, error recovery, self-initiated choice, and attempts to shape the activity rather than only respond to it [1]. The agentic component is retained within the cognitive domain for parsimony in the primary three-domain analysis, while its indicators are also reported separately in secondary analyses.
For learners with SEN, the measurement problem is intensified by heterogeneity in communication, motor control, sensory processing, attention regulation, and familiarity with conventional classroom response formats. Measures that privilege speech, rapid manual responses, or a single normative expression of positive affect can convert access differences into apparent engagement differences. The protocol therefore defines engagement through multiple observable channels and interprets each indicator in relation to the learner's documented access profile. This relational orientation is also consistent with evidence that teacher-student relationships and opportunities for participation are associated with engagement and disengagement among learners with SEN [2]. Consequently, educator mediation is treated as a standardized component of the experimental situation and as an interpretive variable, rather than as background noise to be ignored.

1.2. Musical Interaction, Embodiment, and Multisensory Access

Musical interaction is particularly suited to an embodied analysis because sound unfolds as a consequence of action while simultaneously reorganizing subsequent action. Rhythm invites temporal prediction; changes in intensity or timbre make the effects of force and movement perceptible; repeated patterns allow anticipation; and shared pulse can support interpersonal coordination. Recent empirical work with children who have special needs has associated embodied musical engagement with attention control, emotion regulation, persistence, and cooperative behavior [6], while studies in autism have linked active rhythm engagement to social interaction and expressive communication pathways [7]. Randomized and feasibility studies likewise indicate that structured music-based activity can support engagement, initiation, attention to communication, and language-related outcomes for some autistic learners [8,9], and a recent meta-analysis reports positive effects of music therapy on communication and social skills, while also documenting substantial heterogeneity across interventions [10]. These findings do not establish that any musical technology will produce the same effects; rather, they justify closer examination of the interaction mechanisms through which music becomes accessible and motivating.
The multisensory character of musical play is educationally relevant because auditory information can be coupled with tactile, visual, proprioceptive, and spatial feedback. This coupling can create redundant or complementary channels, but it can also increase sensory load. Accessible design therefore requires more than maximizing stimulation. It requires controllable intensity, predictable contingencies, clear mappings between action and consequence, opportunities for withdrawal, and the possibility of adapting the interface to motor and sensory preferences. Recent reviews of natural and digital musical instruments for autistic children similarly emphasize both the expressive potential of technology and the need to account for individual sensory and interaction profiles [11]. The proposed comparison treats sensory controllability and action-feedback intelligibility as measurable affordance properties rather than as assumed benefits of one medium.

1.3. Tangible Musical Toys: Materiality, Direct Manipulation, and Shared Visibility

Physical musical toys provide information through material properties that are inseparable from the act of manipulation: weight, resistance, texture, vibration, spatial position, and the visible trajectory of the body. When action and sound are tightly coupled, a learner can test a hypothesis by varying force, timing, location, or movement and immediately hear or feel the consequence. These qualities can lower abstraction demands and can make participation visible to peers and adults, which is especially relevant in joint musical activity. Tangible instruments can also support distributed attention because the object, the learner's movement, and the partner's response occupy a shared physical space. Yet materiality is not automatically inclusive. Conventional instruments often presuppose particular ranges of motion, grip, timing precision, bilateral coordination, or strength, and even adapted objects may create barriers if their physical affordances do not match the learner's access profile.
Work on Accessible Digital Musical Instruments (ADMIs) demonstrates why the physical-digital distinction should not be treated as binary. Many ADMIs are tangible interfaces whose physical surfaces are digitally mapped, allowing the form of the gesture and the resulting sound to be decoupled and reconfigured. A social-ecological design framework developed in SEN school settings shows that accessibility emerges from the interaction among instrument design, learners, educators, institutional routines, and musical goals [12]. The I-Ork project similarly situates accessible instrument design within music technology, human-computer interaction, music education, therapy, and community music, underscoring the value of adaptable mappings and collaborative participation [13]. Participatory design research with autistic children further indicates that collaborative accessible instruments can be developed around movement and shared musical action when teacher knowledge and learner experience inform design choices [14]. These studies motivate the protocol's emphasis on functional affordances rather than on hardware category alone.

1.4. Virtual and Immersive Musical Toys: Programmability, Adaptive Mapping, and Traceability

Virtual and immersive musical environments relocate part of the interaction logic from the material object to a programmable mapping layer. A touch, gesture, or body movement can be translated into pitch, timbre, spatialized sound, visual animation, or haptic response; the sensitivity of the mapping can be modified without changing the learner's physical repertoire; task complexity can be progressively altered; and feedback can be synchronized across modalities. This programmability can support personalization and can preserve a stable task structure while changing the access route. Deployment work with an elastic interactive display for autistic children, for example, showed sustained musical motivation and a wide range of sound and gesture variation over repeated sessions [15]. Embodied systems such as OSMoSIS have mapped full-body movement to sound to support playful interaction and imaginative activity [16,17], while multisensory environments such as MusicTraces combine bodily movement, sound, and visual traces to support collaborative activity with autistic people and individuals with complex needs [18].
Immersive interaction also creates a computational opportunity that physical objects do not automatically provide: the system can record every registered event with a timestamp and task state. Such logs can reveal initiation latency, action frequency, temporal spacing, choice sequences, navigation paths, response to prompts, and patterns of exploration that are difficult to reconstruct from global ratings. Haptic co-design with Deaf and Hard-of-Hearing children shows how non-auditory musical channels can be treated as expressive media in their own right rather than as secondary compensations [19], and recent interactive systems such as uCue illustrate how musical preference, choice, and formative listening can be represented within child-centered digital interfaces [20]. At the same time, computational traceability can create an asymmetry in comparative research if the digital condition is measured with richer instrumentation than the physical condition. The present protocol addresses this problem by translating both conditions into a common event ontology, with video-coded tangible events represented using the same action classes as digital logs.

1.5. Inclusive Multimedia Design: Agency, Accessibility, and Educator Mediation

Inclusive multimedia design should not equate access with successful input registration. An interface may detect a gesture reliably while leaving the learner little meaningful control over musical structure; conversely, a slower or less precise action can represent strong engagement if it reflects deliberate choice, experimentation, or social negotiation. Disability-oriented analyses of musical interfaces have shown that design assumptions can either reproduce normative expectations about performance or enable alternative forms of musical authorship [21]. Research on game-based learning for learners with disabilities likewise emphasizes the joint role of technology, rules, community, division of labor, and target outcomes, arguing against interpretations that isolate the digital artifact from its activity system [22]. Reviews of artificial intelligence and virtual reality in inclusive education further indicate that personalization and immersive access must be evaluated alongside usability, accessibility, teacher readiness, and the possibility of new forms of exclusion [23,24].
For this reason, the protocol treats agency as a design and measurement principle. Agency is operationalized through self-initiated actions, meaningful choice, successful modification of the environment, voluntary continuation, and the capacity to influence the pace or direction of the activity. Educator behavior is standardized through a prompt hierarchy so that differences in adult assistance do not masquerade as differences between media. At the same time, educator observations are preserved qualitatively because educators possess contextual knowledge about communication modes, fatigue, sensory regulation, and atypical expressions of preference that may not be inferable from logs alone. This dual treatment of mediation—as both a controlled component and a source of interpretive evidence—reflects inclusive pedagogical perspectives in which teaching action is adaptive, situated, and oriented toward making participation possible through proportionate regulation rather than through uniform demands [25,26]. Virtual environments can be especially useful when they allow perspective, movement, and response conditions to be manipulated systematically, as shown in recent work using a virtual classroom videogame to study perspective taking in children [27].
Figure 1 summarizes the conceptual model used in the study. The two media conditions are first matched at the level of musical goal, choice structure, duration, and adult prompting. Their different affordances are then expected to influence proximal sensorimotor, relational, and access processes, which in turn may be expressed through behavioral, emotional, and cognitive-agentic engagement. Learner access profile and educator mediation are treated as moderators across the interaction cycle.

1.6. Research Gap, Objectives, and Research Questions

Existing research provides strong reasons to study musical technologies in inclusive education, but the evidence base remains difficult to compare across systems. Studies often examine one instrument, one platform, one diagnostic group, or one therapeutic goal, and they use engagement as a broad descriptor rather than as a multidimensional outcome. Moreover, physical and digital musical tools are rarely compared under equivalent task rules within the same learners. This makes it difficult to distinguish a medium effect from novelty, task difficulty, adult support, musical content, or differences in measurement density. The present study is designed to isolate the interaction medium more carefully by holding the musical task architecture constant while varying the physical versus virtual/immersive mode of action and feedback.
The primary objective is to estimate within-learner differences in behavioral, emotional, and cognitive-agentic engagement across matched tangible and immersive musical activities. A second objective is to identify whether individual access profiles and specific affordance properties help explain heterogeneous responses. A third objective is methodological: to test a common event representation that combines systematic observation and digital logs without treating one data source as inherently more valid than the other. This objective aligns the educational question with multimedia research on instrumented interaction and multimodal learning analytics [28,29,30].
The study addresses five research questions: (RQ1) How does behavioral engagement differ between tangible and immersive musical play when musical goals, choice structure, duration, and educator prompting are matched? (RQ2) How does emotional engagement differ between conditions, particularly in enjoyment, frustration/avoidance, and willingness to continue? (RQ3) How does cognitive-agentic engagement differ in purposeful exploration, decision-making, strategy variation, error recovery, and self-initiated choice? (RQ4) To what extent do sensory, motor, communication, and prior-experience profiles moderate within-learner condition effects? (RQ5) How do educator observations and event-level interaction traces explain cases in which quantitative engagement indicators converge or diverge? No directional superiority hypothesis is imposed. The principal expectation is heterogeneity: different affordance configurations are expected to fit different learners and different engagement functions.
Table 1. Theoretical comparison of the two musical interaction conditions. 
Table 1. Theoretical comparison of the two musical interaction conditions. 
Analytical dimension Physical/tangible musical toys Virtual/immersive musical environments Comparative implication
Action-feedback coupling Direct material manipulation; acoustic and tactile consequence may be physically co-located. Programmable mapping of touch, gesture, or movement to sound and visual/haptic feedback. Latency and mapping transparency must be documented in both conditions.
Embodiment Weight, resistance, texture, vibration, reach, and spatial manipulation. Body movement can be amplified, filtered, or remapped through sensing and software. Motor demand should be matched to the learner rather than standardized only by device.
Agency Choice may be constrained by physical layout and available objects. Choice sets, sensitivity, and response rules can be dynamically personalized. Agency is measured through effective control, not merely number of options.
Social visibility Actions and objects are directly visible in a shared space. Actions may be distributed across screen, sensor field, avatar, or projected space. Turn-taking and shared attention are coded explicitly.
Measurement Requires structured observation and, where available, sensor instrumentation. Produces native time-stamped logs in addition to observation. Both are translated into a common event ontology.
Accessibility risk Mechanical effort, reach, grip, or conventional performance assumptions. Visual complexity, tracking failure, latency, sensory load, interface abstraction. Barriers are logged as interaction events rather than attributed to learner deficit.

2. Materials and Methods

2.1. Study Design and Comparative Logic

The study is designed as a randomized, counterbalanced, within-participant crossover protocol with an embedded qualitative strand. Each learner experiences both the tangible condition (A) and the immersive condition (B). Participants are assigned to sequence AB or BA using computer-generated block randomization with balanced allocation. The within-participant structure is preferred because SEN populations are heterogeneous and between-group comparisons can be dominated by stable differences in communication, motor access, sensory profile, prior musical experience, or support needs. Each learner therefore serves as their own primary comparator, while sequence and period effects are modeled explicitly.
The quantitative strand estimates condition effects on pre-specified engagement outcomes. The qualitative strand examines interaction episodes, educator observations, access barriers, and learner-specific responses that help interpret the quantitative pattern. Integration occurs at the participant-by-condition level through joint displays rather than by treating qualitative notes as anecdotal supplements. The design is comparative but not evaluative in the sense of selecting a universal winner. Its target is conditional fit: which medium-affordance configuration is associated with which form of engagement, for which learner, under which access conditions.
Figure 2. Planned randomized counterbalanced crossover sequence. AB and BA orders distribute novelty and period effects across conditions; the same outcome framework is applied after each condition.
Figure 2. Planned randomized counterbalanced crossover sequence. AB and BA orders distribute novelty and period effects across conditions; the same outcome framework is applied after each condition.
Preprints 230287 g002

2.2. Setting, Participants, and Eligibility

The intended sample comprises school-aged learners with documented SEN enrolled in inclusive or specialized educational provision. Recruitment is planned through cooperating educational settings. Because the protocol is intended to examine access heterogeneity rather than a single diagnostic category, eligibility is based on educational need and capacity to participate in at least one supported cause-and-effect musical interaction, not on diagnosis alone. Inclusion criteria are: (a) current enrollment in a participating educational setting; (b) formal documentation of SEN or equivalent educational support eligibility; (c) ability to indicate assent, preference, continuation, or withdrawal through speech, augmentative communication, gesture, gaze, body movement, or an individually recognized response; and (d) availability of a parent/legal guardian and an educator who can provide contextual information. Exclusion criteria are limited to temporary conditions that make participation unsafe on the scheduled day, uncorrected sensory conditions that preclude access to both planned forms of feedback, or clinical restrictions that contraindicate the movement or sensory exposure required by the protocol.
The sample will be characterized descriptively by age, sex, educational setting, SEN category where ethically and educationally relevant, communication mode, mobility/access method, sensory accommodations, prior musical experience, prior exposure to digital or immersive systems, and level of routine educational support. Diagnostic labels will not be used as explanatory shortcuts. Wherever subgroup analyses are feasible, they will be framed around functional access variables, because two learners sharing a diagnostic label may require very different mappings, sensory intensities, or prompting strategies.

2.3. Sample-Size Planning and Sensitivity

The primary inferential contrast is the within-participant condition effect. In the absence of pilot variance components for the final instruments, the recruitment target is anchored to a conservative paired-condition approximation and will be confirmed by simulation after pilot calibration of the mixed-effects model. For a two-sided alpha of 0.05, a standardized within-participant effect of d_z = 0.45 requires 54 complete participant pairs for 90% power. Allowing approximately 15% incomplete data due to absence, fatigue, withdrawal, unusable video, or technical failure yields a recruitment target of 64 participants. This target is a planning minimum rather than a claim about the final realized sample. Figure 4 displays the sensitivity of the complete-pair requirement across plausible standardized effects.
The final power analysis will use the planned mixed-effects model and empirically estimated within-participant correlation from a small protocol calibration sample that will not be used to test the primary hypotheses unless specified in a preregistered amendment. Recruitment will continue until the preregistered complete-case target or the planned recruitment ceiling is reached. Any deviation will be reported with reasons rather than justified retrospectively through observed power.
Figure 3. Sample-size sensitivity guide based on a two-sided paired-condition approximation. The final power statement will be updated using simulation from the mixed-effects model after protocol calibration; the figure is intended as a transparent planning reference rather than as a substitute for model-based power analysis.
Figure 3. Sample-size sensitivity guide based on a two-sided paired-condition approximation. The final power statement will be updated using simulation from the mixed-effects model after protocol calibration; the figure is intended as a transparent planning reference rather than as a substitute for model-based power analysis.
Preprints 230287 g003

2.4. Experimental Conditions and Functional Matching

Both conditions will implement the same four musical functions: (1) cause-and-effect sound triggering; (2) controlled variation of at least one musical parameter such as pitch, timbre, intensity, or tempo; (3) construction or continuation of a short rhythmic or sonic sequence; and (4) a shared or turn-based musical exchange with an educator or peer when appropriate. The tangible condition will use physical musical toys and manipulable sound-producing objects. The immersive condition will use an interactive digital environment in which equivalent musical functions are accessed through touch, gesture, or body movement. Specific hardware and software will be documented before data collection, including input technology, output devices, sampling or synthesis engine, display modality, audio chain, tracking frequency, measured action-to-feedback latency, and the version of all software used.
Functional equivalence is prioritized over visual similarity. A physical shaker and a gesture-controlled digital object need not look alike, but they must instantiate comparable opportunities to initiate sound, vary a parameter, repeat an action, stop, and make a choice. Before participant recruitment, a technical matching audit will verify that both conditions offer the same number of task phases, comparable choice cardinality, equivalent nominal duration, the same musical material where feasible, the same adult prompt hierarchy, and a maximum acceptable feedback latency. When a learner requires an access adaptation, the goal is not to force identical motor gestures across conditions but to preserve the same educational function with the least restrictive access route.
Table 2. Condition-matching matrix to be completed and frozen before preregistration. 
Table 2. Condition-matching matrix to be completed and frozen before preregistration. 
Feature Tangible condition Immersive condition Matching rule
Session duration 30 min planned 30 min planned Equal nominal duration; learner-led early stop permitted.
Familiarization 5 min 5 min Same explanation and prompt hierarchy.
Free exploration 8 min 8 min Same number of available musical functions.
Goal-guided tasks 12 min 12 min Same musical goal, choice cardinality, and success criterion.
Shared/turn-based play 5 min 5 min Same partner role and turn structure.
Feedback latency Measured at device level Measured at system level Report median and 95th percentile; keep below preregistered tolerance.
Adult prompting Five-level prompt hierarchy Five-level prompt hierarchy Prompt level time-stamped/coded in both conditions.
Data capture Video + coded events (+ sensors if available) Native logs + video + coded events Common event ontology and synchronized clock.

2.5. Session Procedure

Each participant will first complete a baseline access-profile session with an educator and researcher. This session will document preferred communication, motor access, relevant sensory accommodations, previous experience with musical or interactive technologies, and any individualized stop signals. A brief familiarization with both media will occur before randomization so that the main contrast is not simply first contact with one device. No performance criterion is required during familiarization; the purpose is to establish a minimal understanding that an action can produce a musical consequence.
Each experimental session is planned for approximately 30 min and follows the same temporal structure: 5 min of supported familiarization/reorientation, 8 min of learner-led free exploration, 12 min of goal-guided musical tasks, and 5 min of shared or turn-based play. A learner may pause or stop at any time. The educator follows a five-level prompt hierarchy: (0) no prompt, (1) environmental cue, (2) gestural/model prompt, (3) verbal or symbolic prompt, and (4) individualized physical or access support already used in the learner's educational plan. Prompts are delivered only after a preregistered response interval unless safety or distress requires immediate intervention. The interval between the two condition blocks will be at least one school day. Where feasible, sessions will occur at similar times of day and in the same room to limit context-related variation.

2.6. Engagement Outcomes and Operational Definitions

Engagement is operationalized through a multi-source measurement matrix. No single indicator is treated as the construct itself. The behavioral domain combines active interaction time, self-initiated actions, persistence after unsuccessful attempts, task completion, turn-taking, and prompt dependence. The emotional domain combines observable positive affect, approach behavior, frustration/avoidance, recovery after difficulty, and willingness to continue or repeat the activity. The cognitive-agentic domain combines purposeful variation, choice differentiation, rule discovery, strategy change, error recovery, independent task sequencing, and attempts to modify the activity. Because affect can be expressed idiosyncratically, educator interpretation is recorded separately and does not overwrite the primary observational code.
The primary outcome will be the proportion of scorable session time classified as actively engaged according to the behavioral coding protocol. Secondary outcomes include domain composite scores and event-level indicators. Composites will be created only if reliability and dimensionality are acceptable in the calibration data; otherwise, pre-specified indicators will be analyzed separately with multiplicity control. Learner refusal, withdrawal, self-regulation breaks, and sensory avoidance are coded as meaningful events rather than automatically categorized as negative engagement, because a voluntary stop can itself represent agency.
Table 3. Planned engagement measurement matrix. 
Table 3. Planned engagement measurement matrix. 
Domain Indicator Operationalization Primary data source
Behavioral Active interaction time Proportion of scorable time in task-relevant manipulation, exploration, listening-with-orienting, or shared musical action. Video coding; digital task state/logs.
Behavioral Initiation and persistence Latency to first self-initiated action; number and duration of sustained action sequences; return after unsuccessful action. Common event ontology.
Behavioral Prompt dependence Highest and mean prompt level; proportion of actions following prompts versus self-initiation. Observer codes + facilitator event marker.
Emotional Approach/enjoyment Observable approach, positive affect, voluntary continuation, request to repeat. Video coding + educator rating.
Emotional Frustration/avoidance Observable distress, withdrawal, repeated rejection, dysregulation; recovery time coded separately. Video coding + educator contextual note.
Cognitive-agentic Purposeful exploration Systematic variation of action or parameter rather than undifferentiated repetition. Video/log sequence analysis.
Cognitive-agentic Choice and control Distinct choices, self-initiated changes, successful environment modification, voluntary stopping. Common event ontology.
Cognitive-agentic Strategy and error recovery Change in action after mismatch/failure; successful completion after strategy revision. Video/log event sequence.
Relational secondary Turn-taking/shared attention Reciprocal alternation, partner-oriented action, coordinated start/stop or shared sonic goal. Video coding; facilitator event marker.

2.7. Common Event Ontology and Multimodal Data Architecture

A central methodological component is a common event ontology that represents interactions from both conditions in the same analytic language. Every coded or logged event will contain, at minimum: participant pseudonym, condition, session, timestamp, task state, action class, feedback class, prompt level, success/mismatch flag where applicable, and annotation confidence. Action classes will include initiate, repeat, vary, choose, stop, request/help, partner-oriented action, and recovery/change-strategy. Feedback classes will include auditory, visual, haptic/tactile, combined, delayed/failed, and educator-mediated feedback. The ontology will be frozen before confirmatory analysis and published with the analysis code.
Digital events will be exported directly from the immersive system using a synchronized clock. Tangible events will be annotated from video with event-marking software and, where feasible, supplemented by simple sensors. Synchronization checks will be performed at the beginning and end of each session using a visible/audible marker. Derived features will include event rates, inter-event intervals, initiation latency, exploration breadth, sequence depth, prompt-adjusted autonomy, and choice diversity. Exploration entropy may be estimated for participants with sufficient event counts, but it will remain a secondary descriptive feature because high entropy can reflect productive exploration or disorganized responding depending on task context. Multimodal learning analytics research has repeatedly cautioned that data fusion is informative only when modalities are theoretically aligned and temporally synchronized [28,29,30]. The protocol therefore prioritizes interpretable event features over opaque high-dimensional prediction.
Figure 4. Planned synchronized data architecture. Tangible observations and immersive digital logs are transformed into a shared event representation before feature derivation, statistical modeling, and qualitative integration.
Figure 4. Planned synchronized data architecture. Tangible observations and immersive digital logs are transformed into a shared event representation before feature derivation, statistical modeling, and qualitative integration.
Preprints 230287 g004

2.8. Observation Coding, Training, and Reliability

A structured coding manual will define each engagement event, boundary rule, ambiguous case, and condition-independent example. Coders will complete training on recordings not included in the confirmatory dataset, followed by calibration until acceptable agreement is reached. At least 25% of sessions, stratified by condition and participant support level, will be independently double-coded. Continuous and proportion outcomes will be evaluated using intraclass correlation coefficients with an a priori target of at least 0.75 for confirmatory use; ordinal codes will use weighted kappa, and event occurrence will additionally be examined through tolerance-window matching when timestamp agreement matters. If reliability falls below the threshold, the code definition will be revised and affected material recoded before unblinding the primary condition comparison.
Complete blinding to condition is impossible because the physical form of the interface is visible on video. Bias will therefore be reduced through other means: coders will not be informed of the study's expected direction, condition labels will be replaced with neutral codes during annotation, the primary analysis script will be written before decoding the condition key where feasible, and automated log-derived features will be computed from pre-specified functions.

2.9. Educator Measures and Qualitative Strand

Immediately after each session, the educator will complete a brief structured rating form addressing perceived engagement, sensory fit, motor access, communication opportunity, frustration, autonomy, and degree of assistance required. Ratings will use behaviorally anchored response categories. Educators will also record a short structured observation describing one episode in which the system supported access and one episode in which it constrained access. This symmetrical prompt is intended to reduce technology-positive reporting bias.
A purposive subset of cases will be selected for deeper qualitative analysis after the quantitative dataset has been quality-checked but before final interpretation. Selection will maximize explanatory value by including: learners with consistently high engagement in both conditions, learners with consistently low engagement, learners showing a strong tangible preference, learners showing a strong immersive preference, and learners whose quantitative and educator-rated outcomes diverge. Qualitative coding will use a deductive-inductive framework organized around agency, sensory fit, motor fit, predictability, social visibility, adult mediation, and task meaning, while allowing new categories to emerge. At least two researchers will review the framework and negative cases; disagreements will be resolved through analytic discussion and documented memoing rather than reduced to a single agreement coefficient.

2.10. Statistical Analysis Plan

Analyses will be conducted in R (version to be frozen at preregistration) and, where appropriate, reproduced in an open analysis notebook. The primary outcome—proportion of active engagement time—will be modeled using a mixed-effects framework appropriate to its distribution. If values are sufficiently interior to (0,1), beta mixed regression will be preferred; if a substantial mass occurs at 0 or 1, a zero/one-inflated or binomial-time formulation will be considered according to the preregistered decision rule. The principal fixed effect is condition (tangible versus immersive). Period, randomized order, session block, and measured task duration will be included as design covariates. Participant will be modeled as a random intercept; educational site/class will be added as a random effect if the number of clusters permits stable estimation. The primary estimand is the average within-learner condition contrast with a 95% confidence interval.
Count outcomes such as self-initiated actions or strategy changes will use Poisson or negative-binomial mixed models based on dispersion diagnostics. Binary outcomes will use logistic mixed models. Approximately continuous ratings or standardized composites will use linear mixed models with residual diagnostics and robust alternatives if assumptions are materially violated. Moderation analyses will test pre-specified interactions between condition and functional access variables (sensory accommodation needs, motor access method, communication mode, prior digital familiarity, and prior music experience) only where cell support is adequate. These interactions are exploratory unless the final preregistration specifies otherwise.
Multiplicity will be controlled hierarchically. The primary outcome will be tested at two-sided alpha = 0.05. Secondary indicators within behavioral, emotional, and cognitive-agentic families will be adjusted using a false-discovery-rate procedure, while effect sizes and confidence intervals will be emphasized over dichotomous significance labels. Period-by-condition and order-by-condition terms will be inspected for carryover. If substantial carryover is detected, a sensitivity analysis restricted to the first condition period will be reported, acknowledging the loss of within-participant efficiency. All model specifications, exclusions, transformations, and convergence remedies will be documented.
Table 4. Pre-specified analysis family and estimands. 
Table 4. Pre-specified analysis family and estimands. 
Outcome type Example Planned model Primary interpretation
Proportion Active engagement time Beta mixed model or preregistered boundary alternative Adjusted within-learner condition difference / marginal contrast.
Count Self-initiated actions; strategy changes Poisson or negative-binomial mixed model Condition rate ratio with 95% CI.
Latency/time Time to first initiation; recovery time Log-transformed LMM or survival/mixed model if censoring is substantial Condition difference or time ratio.
Ordinal Educator ratings; prompt level Cumulative-link mixed model where feasible Shift in probability of higher/lower category.
Binary Task completion; voluntary repeat request Logistic mixed model Condition odds ratio and marginal probabilities.
Sequence feature Choice diversity; exploration entropy LMM/GLMM after minimum-event threshold Secondary descriptive/associational contrast.
Moderator Functional access variable × condition Interaction in corresponding mixed model Heterogeneity of within-learner condition effect.

2.11. Missing Data, Protocol Deviations, and Robustness

Every incomplete session will be assigned a reason code: learner withdrawal, absence, fatigue/health, equipment failure, recording failure, facilitator deviation, or other documented cause. Mixed-effects models use available repeated observations under a missing-at-random assumption, but this assumption will not be treated as self-validating. The distribution of missingness by condition and order will be reported, and sensitivity analyses will examine whether results change after excluding sessions affected by major technical or procedural deviations. Single-value imputation will not be used for primary outcomes.
Novelty is a particular concern when comparing immersive technology with familiar physical objects. Familiarization, counterbalancing, repeated exposure where scheduling permits, and explicit modeling of session block are therefore built into the design. A secondary analysis will estimate whether the condition contrast changes from first to later exposure. If the interaction attenuates sharply, the study will interpret the initial difference as partly novelty-related rather than as an enduring media effect.

2.12. Mixed-Methods Integration

Quantitative and qualitative evidence will be integrated using participant-level joint displays. For each learner, the display will juxtapose condition-specific quantitative indicators, event-sequence features, educator ratings, access adaptations, and coded qualitative affordance episodes. Integration will classify evidence as convergent, complementary, divergent, or silent. A digital system, for example, may produce more logged actions but lower educator-rated autonomy if many actions follow prompts; conversely, a tangible condition may produce fewer actions but longer sustained sequences and more partner-oriented engagement. These are not measurement contradictions to be averaged away; they are different descriptions of how engagement is organized.
The final interpretation will therefore distinguish intensity from quality, activity from agency, and successful task completion from independent participation. This structure is intended to produce design knowledge that can inform multimedia systems for education rather than a single rank ordering of media.

2.13. Ethics, Accessibility, and Data Governance

This manuscript describes a prospective study protocol. At the time of submission, no participants have been recruited and no empirical data have been collected. The study involves minors and potentially vulnerable participants and will not begin until the competent institutional ethics body has approved the final protocol. Written informed consent will be obtained from parents or legal guardians, and learner assent will be sought using communication methods accessible to each participant. Assent is treated as ongoing: withdrawal behavior, an established stop signal, or sustained refusal will terminate or pause the session even if guardian consent has been provided. Participation or withdrawal will have no effect on educational provision.
Sensory safety will be addressed through pre-session profiling, adjustable sound level and visual intensity, removal of unnecessary flashing effects, accessible seating/positioning, planned breaks, and immediate access to the learner's usual regulation supports. Audio output will remain within safe listening levels. No interface will require a participant to tolerate a stimulus that they reject merely to complete the experimental sequence. Physical assistance, if part of the learner's routine support, will be documented as a prompt event rather than concealed.
All participant data will be pseudonymized. Raw video will be encrypted and access-restricted because it contains identifiable data from minors. Public data release will prioritize de-identified derived event tables, codebooks, analysis scripts, and aggregated outcomes. Raw video will not be openly released. Digital logs will exclude names, email addresses, precise location, and unnecessary device identifiers. The data-management plan will specify retention, access roles, deletion procedures, and the lawful basis for processing under applicable European data-protection requirements.

2.14. Preregistration, Reproducibility, and Protocol Amendments

Before confirmatory data collection, the authors will preregister the hypotheses/research questions, sample-size rule, randomization procedure, condition-matching audit, primary and secondary outcomes, exclusion rules, coding manual, power simulation, and statistical models on an open repository such as the Open Science Framework. Analysis code for the common event ontology and derived features will be version-controlled. Any protocol amendment after preregistration will be timestamped, justified, and distinguished from the original plan in the final report.
The multimedia system description will include enough technical information for reproduction: hardware and sensor specifications, software version, input mapping, output modality, task-state logic, event-log schema, latency measurement procedure, and accessibility parameters. If proprietary components prevent code release, the manuscript will document the functional behavior and interfaces needed for independent replication.

3. Planned Analytical Outputs and Falsifiable Interpretive Logic

Because this manuscript is a study protocol, no empirical results are reported. The study is designed, however, so that several possible outcome patterns carry distinct interpretations. A consistent condition advantage across behavioral, emotional, and cognitive-agentic domains would support a broad media-affordance effect, provided that order and novelty effects are small. A domain-specific advantage—for example, higher exploratory diversity in the immersive condition but longer sustained interaction in the tangible condition—would support the view that media organize engagement differently rather than globally. Strong condition-by-access-profile interactions would indicate that personalization and functional fit explain more variance than medium category alone. Finally, weak or negligible condition effects accompanied by substantial between-learner heterogeneity would argue against universal claims about physical or immersive superiority and would redirect design attention toward individualized mappings and educator mediation.
Qualitative integration is specifically intended to test interpretations that cannot be inferred safely from event volume. More actions may indicate active exploration, but they may also arise from rapid unsuccessful attempts. Longer duration may indicate persistence, but it may also reflect difficulty leaving a task. Fewer prompts may reflect autonomy, but they may also occur because the educator recognizes a learner's need for processing time. The joint-display analysis will therefore search for explanatory cases that challenge the most convenient quantitative narrative.

4. Discussion

4.1. Scientific Contribution to Inclusive Multimedia Research

The protocol contributes to inclusive multimedia research by making the interaction architecture, rather than the device label, the unit of comparison. Physical and immersive musical toys are defined through affordances—how action is sensed, mapped, fed back, shared, and adapted—so that the comparison remains meaningful even as specific commercial systems change. This is particularly important in music technology, where tangible interfaces may be digitally mediated and virtual systems may depend on highly embodied movement. Recent accessible-instrument research has already moved toward social-ecological and participatory models of design [12,13,14,19,20,21]. The proposed study extends that orientation into comparative measurement by asking whether matched affordances produce distinct engagement profiles within the same learners.
A second contribution is computational. The common event ontology gives the physical condition a trace representation comparable to native digital logs. This reduces a recurring methodological asymmetry in which digital systems appear analytically richer simply because they record more events. The ontology also makes it possible to connect educational observation with multimedia analytics while retaining interpretable constructs. Rather than using multimodal data because more channels are technologically available, the protocol fuses only channels that correspond to explicit engagement or affordance hypotheses [28,29,30].
A third contribution concerns inclusive pedagogy. The design recognizes that adaptation is not a methodological contaminant to be eliminated. If a learner uses eye gaze in one condition and a large-area touch surface in another, the motor actions differ, but the educational function—making a meaningful musical choice—can remain equivalent. This principle is compatible with inclusive approaches that seek proportional, context-sensitive regulation of learning demands and with perspectives on music, technology, and the learner's broader life project [25,26,31]. The comparison therefore standardizes purpose and opportunity while allowing access routes to vary.

4.2. Methodological Strengths

The within-participant randomized crossover structure reduces confounding by stable learner differences; counterbalancing distributes first-exposure effects; a common task architecture limits differences in musical goal and adult prompting; the common event ontology prevents the immersive condition from benefiting solely from native logging; and multi-source outcomes reduce reliance on a single expression of engagement. Reliability procedures, explicit missing-data codes, preregistration, and public analysis code are intended to make analytic decisions inspectable.
The protocol also separates access from performance. A tracking failure, unreachable object, sensory overload, or misunderstood mapping is recorded as an interaction-system event before it is interpreted as disengagement. This distinction is especially important when studying learners with disabilities because system failure can otherwise be misclassified as learner failure. Participatory and disability-oriented instrument research has shown the importance of taking such design assumptions seriously [14,19,21].

4.3. Anticipated Limitations and Boundary Conditions

Several limitations are anticipated. First, functional matching cannot make physical and immersive systems identical, and some affordances are inherently medium-specific. The protocol addresses this by reporting those differences explicitly rather than implying perfect equivalence. Second, novelty may favor the immersive condition for some learners and increase anxiety for others; repeated exposure and order modeling can reduce but not remove this effect. Third, an educationally heterogeneous SEN sample improves ecological relevance but limits diagnostic subgroup inference. Functional access moderators are therefore prioritized, although statistical power for interactions may remain limited.
Fourth, engagement is context-sensitive. A condition that supports exploration in a 30-min study session may not support sustained curriculum participation over a semester. The study should therefore be regarded as evidence about proximal interaction and engagement, not as proof of long-term learning effects. Fifth, educator ratings provide valuable contextual knowledge but can be influenced by expectations about technology or familiarity with the learner. The protocol mitigates this through behaviorally anchored ratings, symmetrical prompts about benefits and constraints, and comparison with independently coded events. Finally, the planned event ontology is intentionally interpretable and may omit subtle qualities of musical expression. Future studies may extend it with acoustic, movement, physiological, or gaze features when those modalities answer a clear question and can be collected without disproportionate burden.

4.4. Implications for System Design and Educational Practice

If the study reveals heterogeneous engagement profiles rather than a single media advantage, the practical implication will be to design musical systems as configurable repertoires of action-feedback relations. Educators and designers would then select among tangible resistance, large-movement control, touch precision, haptic feedback, visual support, simplified choice sets, or adaptive mappings according to the engagement function being targeted. A learner who benefits from material predictability may require a different configuration from a learner whose motor range is better amplified by motion tracking. Hybrid environments may become especially valuable when they preserve the physical clarity of tangible action while using software to personalize sound mappings and collect interpretable traces.
For educators, the measurement framework offers a way to discuss engagement without reducing it to attention or compliance. A system that increases independent choice but produces fewer total actions may still represent an educational gain; a system that produces excitement but requires continuous prompting may not support autonomy. The proposed behavioral, emotional, and cognitive-agentic matrix encourages decisions based on a profile of participation and on the learner's capacity to influence the activity.

5. Conclusions

This study protocol proposes a controlled but ecologically informed comparison of physical and immersive musical toys for learners with SEN. Its central premise is that the educational question is not whether digital interaction is better than material interaction, but how different action-feedback architectures organize participation, emotion, exploration, and agency for different learners. By combining a randomized within-participant design, functional task matching, systematic observation, common event coding, native digital logs, educator interpretation, and mixed-effects analysis, the protocol is designed to produce evidence that is both pedagogically meaningful and computationally traceable.
The expected value of the study lies in the resolution of its comparison. Instead of treating engagement as a single score or technology as a uniform intervention, the design examines multiple forms of engagement and the conditions under which they emerge. Such evidence can support more accountable development of accessible digital musical instruments, serious games, immersive environments, and hybrid physical-digital experiences in which learner agency and educational purpose remain visible within the multimedia system itself.

Supplementary Materials

The final study report is expected to provide the following supplementary materials: coding manual; common event ontology; condition-matching audit; educator rating form; randomization code; statistical analysis scripts; and, where licensing permits, technical configuration files for the immersive musical environment.

Author Contributions

Conceptualization, A.D.P. and M.D.T.; methodology, A.D.P. and M.D.T.; validation, A.D.P. and M.D.T.; formal analysis plan, A.D.P. and M.D.T.; investigation protocol, A.D.P. and M.D.T.; resources, A.D.P. and M.D.T.; writing-original draft preparation, A.D.P. and M.D.T.; writing-review and editing, A.D.P. and M.D.T.; visualization, A.D.P. and M.D.T.; supervision, A.D.P. and M.D.T.; project administration, A.D.P. All authors have read and agreed to the submitted version of the manuscript.

Funding

The preparation of this study protocol received no specific grant from any funding agency in the public, commercial, or not-for-profit sectors and no dedicated financial support.

Institutional Review Board Statement

This manuscript describes a prospective study protocol. At the time of submission, no participants have been recruited and no empirical data have been collected. The study will be submitted to the competent Institutional Ethics Committee for review and approval before participant recruitment and data collection begin. No study procedure involving human participants will be initiated until formal ethical approval has been obtained. The study will be conducted in accordance with the Declaration of Helsinki and applicable institutional and European ethical and data-protection requirements.

Data Availability Statement

No empirical data are reported at the protocol stage. After study completion, de-identified derived event data, the codebook, and analysis scripts will be deposited in a public repository to the extent permitted by ethics and data-protection requirements. Raw identifiable video data from minors will not be released openly and will be managed under controlled institutional access.

Conflicts of Interest

The authors declare that they have no known competing financial interests or personal relationships that could have appeared to influence the work reported in this paper.

References

  1. Reeve, J. Specialized Purpose of Each Type of Student Engagement. Educ. Psychol. Rev. 2025, 37. [Google Scholar] [CrossRef]
  2. Pérez-Salas, C.P.; Parra, V.; Sáez-Delgado, F.; Olivares, H. Influence of Teacher-Student Relationships and Special Educational Needs on Student Engagement and Disengagement: A Correlational Study. Front. Psychol. 2021, 12, 708157. [Google Scholar] [CrossRef]
  3. Di Paolo, A.; Todino, M.D. Inclusive Music Education in the Digital Age: The Role of Technology and Edugames in Supporting Students with Special Educational Needs. Encyclopedia 2025, 5, 102. [Google Scholar] [CrossRef]
  4. Mommo, O.; Sutela, K.; Mononen, R. Inclusion and Pedagogical Support for Students with Special Educational Needs in Music Lessons: A Systematic Review. Res. Stud. Music Educ. 2025, 47, 403–423. [Google Scholar] [CrossRef]
  5. Di Paolo, A.; Zollo, I. Special Education, Music and Simplexity: Some Theoretical Reflections. Form. Insegn. 2022, 20, 238–251. [Google Scholar] [CrossRef]
  6. Lu, Y.; Viladot, L.; Chen, X.; Fan, Y.; Yuan, J. The Role of Embodied Musical Engagement in Enhancing Non-Cognitive Skills and Rehabilitation Outcomes in Children with Special Needs. Front. Psychol. 2026, 17, 1798176. [Google Scholar] [CrossRef]
  7. Fram, N.R.; Liu, T.; Lense, M.D. Social Interaction Links Active Musical Rhythm Engagement and Expressive Communication in Autistic Toddlers. Autism Res. 2024, 17, 338–354. [Google Scholar] [CrossRef]
  8. Yum, Y.N.; Poon, K.; Lau, W.K.-W.; Ho, F.C.; Sin, K.F.K.; Chung, K.M.; Lee, H.Y.; Liang, D. Music Therapy Improves Engagement and Initiation for Autistic Children with Mild Intellectual Disabilities: A Randomized Controlled Study. Autism Res. 2024, 17, 2702–2722. [Google Scholar] [CrossRef]
  9. Williams, T.I.; Loucas, T.; Sin, J.; Jeremic, M.; Meyer, S.; Boseley, S.; Fincham-Majumdar, S.; Aslett, G.; Renshaw, R.; Liu, F. Using Music to Assist Language Learning in Autistic Children with Minimal Verbal Language: The MAP Feasibility RCT. Autism 2024, 28, 2515–2533. [Google Scholar] [CrossRef]
  10. Shi, Z.; Wang, S.; Chen, M.; Hu, A.; Long, Q.; Lee, Y. The Effect of Music Therapy on Language Communication and Social Skills in Children with Autism Spectrum Disorder: A Systematic Review and Meta-Analysis. Front. Psychol. 2024, 15, 1336421. [Google Scholar] [CrossRef]
  11. Pavlou, E.-S.; Garmpis, A. Enhancing Social Skills in Children with Autism Spectrum Disorder Through Natural Musical Instruments and Innovative Digital Musical Instruments: A Literature Review. Societies 2025, 15, 53. [Google Scholar] [CrossRef]
  12. Förster, A.; Schnell, N. Designing Accessible Digital Musical Instruments for Special Educational Needs Schools-A Social-Ecological Design Framework. Int. J. Child-Comput. Interact. 2024, 41, 100666. [Google Scholar] [CrossRef]
  13. Mandanici, M.; Bergamino, G.; Valente, S. The I-Ork Project: A Technological Approach to Inclusive Music Making and Therapy. Front. Educ. 2025, 10, 1552302. [Google Scholar] [CrossRef]
  14. Iványi, B.A.; Tjemsland, T.B.; May, L.; Robidoux, M.; Serafin, S. Participatory Design of a Collaborative Accessible Digital Musical Interface with Children with Autism Spectrum Condition. In Proceedings of the International Conference on New Interfaces for Musical Expression, Utrecht, The Netherlands, 2024; pp. 43–51. [Google Scholar] [CrossRef]
  15. Monarca, I.; Tentori, M.; Cibrian, F.L. Understanding the Musical Interaction of Children with Autism Spectrum Disorder Using Elastic Display. Pers. Ubiquitous Comput. 2023, 27, 1843–1860. [Google Scholar] [CrossRef]
  16. Ragone, G. Designing Embodied Musical Interaction for Children with Autism. In Proceedings of the 22nd International ACM SIGACCESS Conference on Computers and Accessibility (ASSETS 2020), Virtual Event, Greece, 2020; pp. 1–4. [Google Scholar] [CrossRef]
  17. Ragone, G.; Good, J.; Howland, K. OSMoSIS: Interactive Sound Generation System for Children with Autism. In Proceedings of the 2020 ACM Interaction Design and Children Conference: Extended Abstracts, London, UK, 2020; pp. 151–156. [Google Scholar] [CrossRef]
  18. Bauer, V.; Padovano, T.; Gianotti, M.; Caslini, G.; Garzotto, F. MusicTraces: A Collaborative Music and Paint Activity for Autistic People. Extended Abstracts of the 2024 CHI Conference on Human Factors in Computing Systems, Honolulu, HI, USA, 2024; Article 247, pp. 1–7. [Google Scholar] [CrossRef]
  19. May, L.; Malik, R.; Thomas, A. Co-Designing Haptic Instruments With Deaf and Hard-of-Hearing Children. In Proceedings of the International Conference on New Interfaces for Musical Expression, Utrecht, The Netherlands, 2024; pp. 52–61. [Google Scholar] [CrossRef]
  20. Karwankar, A.; Ruggiero, E.; Lipkin, Z.; Iyer, M.K.; Brugel, S.; Khatiwada, P.; Stevens, D.; Mauriello, M.L. uCue: An Interactive Musical Interface to Enhance Formative Listening Experiences for Children with ASD. In Proceedings of the 2025 ACM Interaction Design and Children Conference (IDC 2025), Reykjavik, Iceland, 2025; pp. 340–357. [Google Scholar] [CrossRef]
  21. Duarte, E.G.; Cossette, I.; Wanderley, M.M. Analysis of Accessible Digital Musical Instruments through the Lens of Disability Models: A Case Study with Instruments Targeting d/Deaf People. Front. Comput. Sci. 2023, 5, 1158476. [Google Scholar] [CrossRef]
  22. Tlili, A.; Denden, M.; Duan, A.; Padilla-Zea, N.; Huang, R.; Sun, T.; Burgos, D. Game-Based Learning for Learners With Disabilities-What Is Next? A Systematic Literature Review From the Activity Theory Perspective. Front. Psychol. 2022, 12, 814691. [Google Scholar] [CrossRef]
  23. Chalkiadakis, A.; Seremetaki, A.; Kanellou, A.; Kallishi, M.; Morfopoulou, A.; Moraitaki, M.; Mastrokoukou, S. Impact of Artificial Intelligence and Virtual Reality on Educational Inclusion: A Systematic Review of Technologies Supporting Students with Disabilities. Educ. Sci. 2024, 14, 1223. [Google Scholar] [CrossRef]
  24. Peretti, S.; Pino, M.C.; Caruso, F.; Di Mascio, T. Evaluating the Potential of Immersive Virtual Reality-Based Serious Games Interventions for Autism: A Pocket Guide Evaluation Framework. Educ. Sci. 2024, 14, 377. [Google Scholar] [CrossRef]
  25. Aiello, P.; Pace, E.M.; Sibilio, M. A Simplex Approach in Italian Teacher Education Programmes to Promote Inclusive Practices. Int. J. Incl. Educ. 2023, 27, 1163–1176. [Google Scholar] [CrossRef]
  26. Di Tore, S.; Aiello, P.; Sibilio, M.; Berthoz, A. Simplex Didactics: Promoting Transversal Learning through the Training of Perspective Taking. J. E-Learn. Knowl. Soc. 2020, 16, 34–49. [Google Scholar] [CrossRef]
  27. Beatini, V.; Cohen, D.; Di Tore, S.; Pellerin, H.; Aiello, P.; Sibilio, M.; Berthoz, A. Measuring Perspective Taking with the “Virtual Class” Videogame: A Child Development Study. Comput. Hum. Behav. 2024, 151, 108012. [Google Scholar] [CrossRef]
  28. Emerson, A.; Cloude, E.B.; Azevedo, R.; Lester, J. Multimodal Learning Analytics for Game-Based Learning. Br. J. Educ. Technol. 2020, 51, 1505–1526. [Google Scholar] [CrossRef]
  29. Mu, S.; Cui, M.; Huang, X. Multimodal Data Fusion in Learning Analytics: A Systematic Review. Sensors 2020, 20, 6856. [Google Scholar] [CrossRef]
  30. Caskurlu, S.; Ocak, C.; Dai, C.-P. The Scope of Multimodal Learning Analytics in K-8: A Systematic Review. J. Learn. Anal. 2025, 12, 224–236. [Google Scholar] [CrossRef]
  31. Di Paolo, A.; Rescigno, A.; Todino, M.D. Semplessità e Progetto di Vita tra Musica e Tecnologie Didattiche: Traiettorie Inclusive. Educ. Sci. Soc. 2025, 16, 349–374. [Google Scholar] [CrossRef]
Figure 1. Affordance-to-engagement framework guiding the matched comparison of tangible and immersive musical play. The model does not assume that either condition is globally superior; it specifies candidate processes through which different interaction architectures may support or constrain engagement.
Figure 1. Affordance-to-engagement framework guiding the matched comparison of tangible and immersive musical play. The model does not assume that either condition is globally superior; it specifies candidate processes through which different interaction architectures may support or constrain engagement.
Preprints 230287 g001
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.