Preprint
Article

This version is not peer-reviewed.

Beyond the Cursor: What Real-Time Score Following Can and Cannot Tell Us About Piano Learning

Submitted:

20 August 2026

Posted:

21 August 2026

You are already at the latest version

Abstract
Real-time score following is usually shown to users as a cursor moving through the score, but its educational value depends on what happens after position is estimated, not on the estimate itself. This structured applied review draws on music-alignment, instrumental-learning, practice-analysis, and self-regulation research to separate four functions: navigation, recovery, repetition, and practice support. Correct localization is not the same as understanding. A system relocalizing after a jump does not mean a learner recovered musically. Detected repetition does not establish deliberate practice, and coverage metrics do not establish learning. Working from these boundaries, the article proposes the Score-Following Educational Action Matrix (SF-EAM), an unvalidated conceptual framework linking alignment state, evidence provenance, permitted educational action, ambiguity, verification, and inference boundary. The framework calls for explicit unlocated, ambiguous, and relocalizing states, and for suspending judgment when position evidence is thin. Score following, on this view, is infrastructure for pedagogical interaction — not a substitute for pedagogical assessment.
Keywords: 
;  ;  ;  ;  
Subject: 
Arts and Humanities  -   Music
Introduction
A moving cursor is the most visible face of score following. When it tracks correctly, the interface looks like it knows where the performer is — and that apparent achievement gets pressed into service for page turning, synchronized visualization, accompaniment, error detection, replay, and practice review. The trouble starts when these functions are treated as interchangeable. Knowing a likely score position is not the same as knowing why a student stopped, whether a repeated passage reflects a good strategy, whether a wrong note was understood, or whether any learning happened at all.
The distinction matters historically because educational score following is not new. The Piano Tutor built at Carnegie Mellon used score following to detect student errors within an interactive system for beginning pianists (Dannenberg et al., 1990). A later review of instrument-pedagogy technologies places the Piano Tutor lineage among the earliest projects to combine polyphonic score following, page turning, performance analysis, feedback, and instructional design (Acquilino & Scavone, 2022). Tekin et al. (2005) described a system meant to accompany piano students practicing at home while handling mistakes and jumps, though their first reported experiments were monophonic. Han et al. (2013) went further, proposing a real-time system that identified a beginner’s current location and note errors from smartphone audio. None of this supports a claim that educational score following — or score-linked feedback for beginners, or practice-oriented handling of mistakes and jumps — is itself a new idea.
The technical literature has also moved well past strictly linear performances. Automatic page turning already showed a navigation-oriented use of real-time machine listening (Arzt et al., 2008). Arzt and Widmer (2010) built “any-time” tracking meant to tolerate omissions, forward and backward jumps, unexpected repetitions, and mid-piece restarts. The Piano Music Companion later combined live piano identification and position tracking with rehearsal, visualization, and automatic page-turning (Arzt et al., 2014). Offline synchronization work handled structural differences such as inserted or omitted repeats, evaluating JumpDTW on Beethoven piano sonatas (Fremerey et al., 2010). Real-time followers have also been built to tolerate performance errors and arbitrary repeats or skips in practice and rehearsal (Nakamura et al., 2016). Offline work has aligned realistic practice sessions where performers jump between passages and repeat material (Jiang et al., 2019), and more recent work uses score alignment to reconstruct unscripted practice and enable score-driven browsing of recorded attempts (Raphael, 2025). Taken together, this body of work leaves little room to claim novelty for navigation, discontinuity handling, repetition tracking, or practice-session reconstruction.
Current benchmarking reinforces why application and alignment need to be kept separate. Park et al. (2025) systematically compared real-time audio score-following approaches on solo-piano datasets and called for a more unified benchmarking environment — a contribution about representation, algorithms, and evaluation, not educational outcomes. Emerging work such as CODA addresses real-time recovery from notation-level discontinuities in image-based score following (Yang et al., 2026), but this is preprint-stage technical evidence, not pedagogical validation.
A second, separate literature complicates how practice traces get read educationally. Duke et al. (2009) found that retention among advanced pianists was not significantly related to practice time, total practice trials, or complete practice trials — the quality and organization of practice mattered more than sheer quantity. Nielsen (2001) described instrumental practice as cyclic self-regulation involving goals, planning, task strategies, monitoring, and evaluation. Miksza and Brenner (2023) used automated offline score following to document measures played during violin practice, but read those traces alongside diaries, behavioral observation, and stimulated recall rather than treating position data as a full account of self-regulation on its own.
These two literatures meet at a useful boundary. Score following can make a practice trajectory inspectable, but the trajectory carries no automatic pedagogical meaning. A return to measure 32 might be an intentional drill, a restart after failure, an interpretive comparison, a teacher-requested loop, or simply an accidental relocation by the tracking system. A cursor that jumps successfully after a pause can demonstrate system recovery while telling us nothing about the student’s cognitive or musical recovery.
This article asks: what educational actions can reasonably be built on real-time score following, and what has to stay outside the inference? Through a structured applied review, it distinguishes four actions — navigation, recovery, repetition, and practice support — and proposes the Score-Following Educational Action Matrix (SF-EAM). SF-EAM is author-developed and unvalidated. It does not measure learning, diagnose musicianship, or prescribe an optimal practice strategy; its job is to connect alignment evidence to a bounded educational action while preserving uncertainty and the need for teacher or learner verification.
Method
This study uses a structured applied literature-review design rather than a systematic review or meta-analysis. The topic spans music information retrieval, human-computer interaction, instrumental music education, and practice research, and an integrative approach fits when heterogeneous sources answer related but non-identical questions — provided technical performance, educational outcomes, and authorial synthesis stay distinguishable (Whittemore & Knafl, 2005).
Searches were conducted and updated on 18 August 2026. Discovery used targeted scholarly web searches followed by verification against primary publisher, conference, institutional, author-hosted, or openly accessible full-text sources wherever available. Twenty-five search families combined score following or audio-to-score alignment with piano, music education, beginner feedback, page turning, practice, rehearsal, errors, repeats, skips, discontinuities, relocalization, self-regulated practice, performance feedback, digital scaffolding, and uncertainty-aware educational action. Additional searches targeted historical score-following surveys, piano-practice systems, “any-time” tracking, rehearsal companions, mixed or null feedback effects, and exact combinations of educational action, verification, provenance, and prohibited inference — testing how novel the proposed synthesis actually is.
Sources were sorted into five evidence roles. Technical score-following evidence covered real-time or offline alignment, piano-specific representations, discontinuities, and recovery. Educational prior art included tutoring, practice-at-home accompaniment, beginner evaluation, page turning, rehearsal companions, and practice-support systems that explicitly used score following (Dannenberg et al., 1990; Tekin et al., 2005; Acquilino & Scavone, 2022). Practice-process evidence covered repetition, self-regulation, strategy, and retention; the broader mapping of music-practice research by How et al. (2022) reinforced the need to treat practice as a heterogeneous research domain rather than a single countable behavior. Human-outcome evidence was kept only when a study actually measured learners or performers. Emerging technical work available only as a preprint or other non-equivalent publication form was held separate and never treated as peer-reviewed educational validation.
Not every source needed to be real-time to qualify. Offline studies were retained when they directly established prior art for non-linear practice reconstruction, repetition detection, or score-driven review, since these functions constrain what can be claimed as novel for a real-time educational framework. Consumer marketing pages, generic music-app descriptions, internal product testing, and studies where score following was not a substantive component were excluded from the evidentiary synthesis.
The analysis applied a claim-evidence boundary to each candidate educational claim, classifying it as direct external evidence, technical prior art, emerging context, authorial synthesis, proposed framework, adjacent evidence, or blocked claim. Guiding questions included: What does the follower actually observe? What alignment state does that support? Which educational action becomes possible? What ambiguity remains? What additional verification is needed? Which learner-level or pedagogical conclusion would go beyond the evidence? A separate search audit records 25 search families, a source-integrity audit records the access basis and permitted use of every retained reference, and a claim-evidence-contradiction ledger records 37 central claims or inference boundaries.
A single reviewer conducted this review, with OpenAI ChatGPT (GPT-5.6 Sol) used for search-term expansion, evidence organization, drafting support, language revision, and document preparation. AI output was not treated as a scholarly source. No new participant data were generated or analyzed, and no proprietary platform user data, telemetry, internal quality-assurance results, or product performance tests were used as evidence of educational effectiveness. The search was deliberately broad and adversarial but remains search-bounded — an equivalent framework may exist under terminology this review did not recover.
Score Following Is Position Infrastructure, Not Pedagogy
Score following estimates a relationship between an unfolding performance and a reference score. Orio et al. (2003) described score following as synchronizing a computer with a performer playing a known score, reviewing a research history already about two decades old at that point. In a real-time system, observations arrive incrementally and the follower has to update its estimate while the performance continues. Dannenberg and Raphael (2006) describe score alignment as infrastructure for interactive applications such as accompaniment, while Henkel et al. (2019) formulate image-based score following as a multimodal sequential decision problem in which an agent learns to move through sheet music in response to audio. These are computationally different formulations, but both make the same point: the immediate technical output is score position or progression, not a direct measure of learning.
Educational systems build additional layers on top of that estimate. The Piano Tutor analyzed performance relative to a score and used an expert system plus multimedia instruction to respond to novice mistakes (Dannenberg et al., 1990). Han et al. (2013) aligned smartphone audio with a reference score to report current position and incorrect notes for beginner musicians. These examples point to a crucial architectural distinction: score following makes score-referenced interpretation possible, while the educational meaning of an event comes from a separate decision rule, model, or teaching interaction layered on top.
That boundary matters because an alignment error can turn into a pedagogical error. If a system is a measure behind, a correct note can get labeled wrong. If repeated material admits two plausible positions, a confident cursor can hide real uncertainty. If a pianist’s pedal and legato prolong spectral energy past the notated duration, audio features can suggest a delayed score position even when the performance is musically fine. Li and Duan (2016) identified this sustained effect as a specific source of delay error in piano score following and built preprocessing to reduce it.
The historical literature also warns against mistaking a successful demonstration for robust everyday use. Puckette and Lippe (1992) discussed score-following behavior in practical musical settings; later systems explicitly targeted non-linear deviations (Arzt & Widmer, 2010) and piano rehearsal (Arzt et al., 2014). Contemporary benchmarking strengthens the evidence base without erasing this application dependence — Park et al. (2025) compare real-time piano-following approaches across representations and methods, which is precisely why alignment performance needs its own evaluation before an educational layer can treat a position estimate as ground truth.
SF-EAM responds to these documented failure modes by introducing explicit alignment states that must be established before any educational action is permitted. An unlocated state means the available evidence cannot support choosing a position; a located state denotes a stable estimate adequate for the stated application; an ambiguous state holds two or more plausible positions open; a relocalizing state marks a suspected break in continuity while the system searches for a new position; and a recovered state records that a new stable position has been established after the discontinuity. This five-state taxonomy is proposed rather than validated — it has not been tested as a psychological, pedagogical, or psychometric classification. A recovered technical state, in particular, says nothing about the learner’s understanding or recovery.
Table 1. Proposed alignment states before educational action.
Table 1. Proposed alignment states before educational action.
State Operational meaning Permitted system behavior Verification / next action Prohibited inference
Unlocated Evidence is insufficient to select a score position Do not attach correctness or practice meaning to a location Ask performer/teacher to continue, select a point, or provide more evidence Student is lost; student made an error
Located A position estimate is stable enough for the intended application Cursor/page/segment can be synchronized with stated provenance Continue monitoring for contradiction Student understands the passage; execution is correct
Ambiguous Two or more positions remain plausible Show uncertainty or withhold a unique cursor Wait for disambiguating evidence or allow manual confirmation Choose the most likely location and score the student
Relocalizing A discontinuity, pause, skip, or contradiction has invalidated continuity Suspend position-dependent error labels Search broadly and preserve prior trace Student failed to recover
Recovered A new stable position is established after discontinuity Resume permitted score-linked actions from new position Mark recovery event and verify stability Learner recovery, resilience, or understanding
Note. The five alignment states are an authorial SF-EAM taxonomy derived from documented alignment problems and are proposed and unvalidated.
Navigation: Where Are We?
Navigation is the educationally relevant use that sits closest to what the system actually measures. Automatic page turning demonstrates the principle plainly: a real-time follower can control a digital or physical score page during performance (Arzt et al., 2008). “Any-time” tracking extends navigation to performances containing omissions, jumps, repetitions, and restarts (Arzt & Widmer, 2010), while the Piano Music Companion applies live piano position tracking to visualization, rehearsal, and page turning (Arzt et al., 2014). Image-based following maps audio directly onto a location in sheet music (Henkel et al., 2019), and Shan and Tsai (2021) show, in an offline piano context, how alignment can drive a score-following video with a line display and cursor.
Within SF-EAM, proposed navigation affordances include maintaining the current page, scrolling to the active system, highlighting the estimated measure, linking replay to notation, or helping teacher and learner refer to the same score location. These stay close to established score-position and page-navigation capabilities (Arzt et al., 2008; Arzt & Widmer, 2010; Arzt et al., 2014; Henkel et al., 2019) — they are application examples, not demonstrated learning effects.
Navigation still inherits documented alignment and mapping problems. Repeated or structurally similar material, written or performed repeats, omissions, and jumps can create competing alignment hypotheses (Arzt & Widmer, 2010; Fremerey et al., 2010; Shan & Tsai, 2021). In image- or page-based systems, musical position and visual score location are also distinct representational problems (Henkel et al., 2019; Shan & Tsai, 2021). SF-EAM treats reliable score-to-display mapping as a separate evidentiary requirement rather than assuming a correct musical position guarantees notehead-level visual placement.
The reviewed evidence establishes technical and application prior art for synchronized navigation, but it does not show that an accurate cursor or automatic page turn, on its own, improves learning, attention, comprehension, or cognitive load. The warranted claim is narrower: score following can automate or support score navigation when the musical alignment — and any separate score-to-display mapping — are reliable enough. Educational benefits beyond navigation need direct human evidence.
Recovery: What Happens After Continuity Breaks?
Practice routinely contains discontinuities: stops, restarts, local repetitions, backward or forward jumps, starts from the middle of a piece. Tekin et al. (2005) explicitly framed mistakes and jumps as requirements for a score-following system meant for piano students practicing at home. Arzt and Widmer (2010) targeted omissions, jumps, unexpected repetitions, and mid-piece restarts in real-time tracking. Nakamura et al. (2016) later addressed real-time alignment under performance errors and arbitrary repeats and skips, tying that behavior explicitly to practice and rehearsal; their reported experiments were monophonic, with polyphonic extension left as future work.
The same problem shows up in piano-specific and score-synchronization work. Fremerey et al. (2010) handled structural differences such as inserted or omitted repeats in offline score-performance synchronization, evaluating the approach on Beethoven piano sonatas. Jiang et al. (2019) aligned realistic practice where musicians move non-linearly through a score, identifying do-overs and jumps. Shan and Tsai (2021) treat unknown repeats and jumps as a central bottleneck when aligning real piano recordings to raw sheet music. In 2026, Yang, Chen and Han proposed CODA, an emerging real-time image-based approach designed to recover from notation-level discontinuities; since it is a preprint at the time of this review, it is used only as current technical context, not as established educational evidence.
The word “recovery” needs particular care here. A system can recover its score position. A learner can recover a performance after a memory lapse, an interpretive disruption, or a technical error. These are different events. Successful relocalization demonstrates something about the follower — it re-established a usable position estimate — and nothing about whether the student diagnosed the mistake, held musical continuity, used an effective strategy, or built resilience.
This distinction motivates an authorial fail-closed policy within SF-EAM: position-dependent error feedback is suspended while the tracker is relocalizing, because the reference location itself is unstable. Once a new location is backed by stable evidence, navigation or annotation can resume and the discontinuity gets recorded. This is a proposed epistemic policy, not an experimentally demonstrated superiority claim — its purpose is to stop an uncertain alignment hypothesis from turning into a confident learner-level verdict.
SF-EAM also proposes an explicit failure policy. When alignment stays ambiguous, the system can wait for further musical evidence or invite manual confirmation through a rehearsal mark, measure selection, or equivalent control. These are conservative design choices, not empirically validated pedagogical interventions. The evidentiary claim is simple: when the location estimate is not adequately supported, downstream score-dependent judgments should not be presented as if the location were known.
Repetition: Detecting Return Without Judging Its Value
Score alignment can make returns to previously aligned material observable. Fremerey et al. (2010) established structural repeat/jump handling in score-performance synchronization; Jiang et al. (2019) explicitly identify do-overs and jumps in realistic practice; Raphael (2025) uses score-aligned practice excerpts for score-driven browsing and visualizing how often regions are visited. The technical capacity to segment or count revisits is prior art, not a new capability of SF-EAM.
From the SF-EAM perspective, interpreting repetition pedagogically requires evidence beyond location history. Practice research describes goals, strategies, monitoring, evaluation, and qualitative differences among practice behaviors that don’t reduce to a trial count (Nielsen, 2001; Duke et al., 2009; Miksza et al., 2018; How et al., 2022). A score-alignment trace can document that material was revisited, but position alone doesn’t say whether the return served fingering, voicing, rhythm, pedaling, memorization, correction, exploration, or something else entirely.
Direct piano-practice research adds an important constraint. Duke et al. (2009) observed 17 graduate and advanced-undergraduate piano majors and found no significant relationship between next-day retention rankings and practice time, total practice trials, or complete trials. Relationships were stronger for characteristics tied to correctness and practice strategy. A score-following system should resist the temptation to turn repetition count into a productivity score.
SF-EAM treats repetition displays as descriptive evidence: an interface may report how many times a passage was traversed, whether revisits were consecutive or distributed, approximate tempo across attempts, or the order sections were revisited in. That data can be presented for teacher-student interpretation, but the trace alone doesn’t establish whether repetition was strategic, unnecessary, corrective, exploratory, successful, or deliberate — those readings need additional evidence such as audio, performance features, stated goals, cross-attempt comparison, or teacher/learner reflection.
As an authorial inference boundary, SF-EAM does not equate the most frequently visited passage with the learner’s greatest weakness. High visit frequency is compatible with several purposes, and practice research does not support reducing quality to amount alone (Duke et al., 2009; Nielsen, 2001; How et al., 2022). Visit frequency stays location-history evidence — it does not automatically become a diagnosis.
Practice: From Alignment Trace to Reflection
Practice support is the richest and most inference-heavy action considered here. Realistic practice is never a single uninterrupted performance. Jiang et al. (2019) designed offline alignment specifically for non-linear practice, where musicians jump between passages to drill technical challenges or pursue other goals. Raphael (2025) extends this into a score-centered representation of unscripted practice, partitioning recordings into aligned excerpts and supporting score-driven navigation through previous attempts. For real-time non-linear behavior, peer-reviewed prior art already includes “any-time” tracking with jumps, repetitions, and restarts (Arzt & Widmer, 2010), the Piano Music Companion in rehearsal settings (Arzt et al., 2014), and error/repeat/skip-aware alignment motivated by practice and rehearsal (Nakamura et al., 2016).
These systems establish substantial prior art, so this article does not claim novelty for practice-session reconstruction, score-based browsing, visit-frequency visualization, replay by clicking notation, or online tracking of free movement through a score. The question here is different: how should such functions be interpreted educationally?
Research on self-regulated music practice suggests why a location trace alone falls short. Nielsen (2001) describes advanced instrumental practice as involving specific goals, strategic planning, self-instruction, task strategies, selective monitoring, and self-evaluation. Miksza et al. (2018) similarly treat self-regulated practice as a multidimensional process that can be examined and supported pedagogically. Position data may show what material was addressed and when, but it doesn’t directly reveal the student’s goal, attention, judgment, or reason for changing strategy.
Miksza and Brenner (2023) offer a particularly useful methodological example. Their study of advanced violinists used automated offline score following to document measures played, but combined it with questionnaires, practice diaries, behavioral observations, and stimulated recall — the score-following data were one evidence stream within a broader account of self-regulation, not a stand-in for the learner’s mental process.
Human-outcome studies using score-following-enabled applications demand the same caution. Ou et al. (2025) reported a four-month quasi-experiment with 40 violin majors in which the experimental group used Violy, an application that explicitly incorporated score-following technology alongside Note-by-Note real-time pitch evaluation, visual score comparison, demonstrations, accompaniment, audition reports, practice records, and other functions. The intervention was associated with changes in performance and some learner measures, but the design does not isolate score following as the causal component, and it does not establish an effect in piano instruction specifically.
Adjacent feedback research further argues against treating “technology” as a uniformly effective intervention. Nusseck et al. (2025) studied 25 advanced piano students using short Disklavier playback-feedback interventions; students in the intervention groups rated their second performances more favorably on some musical parameters, while ratings by professional pianists showed no significant effects. Rom and Woody (2026), by contrast, found greater short-term performance-accuracy gains for 60 high-school string players under digital scaffold conditions than under control practice, with the largest gains in the condition combining an aural model, playback, and visual evaluative feedback. Neither study isolates score following. Together with the bundled Violy study, they illustrate why outcomes from broader digital-feedback systems should inform the educational context without being attributed automatically to score-following accuracy.
As an authorial SF-EAM proposition, an aligned practice trace can be exposed for reflection without being converted into a learning score. Candidate functions include reconstructing a timeline of excerpts, showing bounded coverage and revisit descriptors, linking notation to recorded attempts, and permitting teacher or learner annotations. These are proposed affordances grounded in established alignment and browsing capabilities; their effects on reflection, practice quality, or learning remain open empirical questions.
Score-Following Educational Action Matrix (SF-EAM)
The Score-Following Educational Action Matrix (SF-EAM) is proposed as an author-developed, unvalidated framework for connecting score-following evidence to educational action. It is not an alignment algorithm, assessment scale, diagnostic instrument, causal model, or predictor of learning. The framework assumes an educational system should first establish an alignment state and only then activate actions whose evidentiary requirements are met.
SF-EAM organizes four actions. Navigation asks where the current performance is located and uses position to synchronize display or access. Recovery asks whether the system can establish a new location after continuity has broken. Repetition asks whether material has been revisited and how those revisits are organized. Practice support asks how a sequence of aligned events can be represented for review, comparison, annotation, or teacher-student discussion. The sequence isn’t developmental — a beginner may use practice review, and an advanced pianist may need simple navigation.
Each action carries its own inference boundary. Navigation cannot establish comprehension. Recovery cannot establish learner recovery. Repetition cannot establish practice quality or intention. Practice analytics cannot establish mastery, self-regulation, or learning gain without additional evidence. This is not a limitation to hide — it’s a design requirement that keeps the system’s claims auditable.
The framework also records provenance, because score-following systems operate on different representations with different observables and failure modes. Prior literature includes symbolic-score alignment, audio-to-score methods, and direct image-based following (Orio et al., 2003; Dannenberg & Raphael, 2006; Henkel et al., 2019; Park et al., 2025). In piano, sustained sound and reverberation can delay audio-based alignment (Li & Duan, 2016), while repeated structures and discontinuities complicate localization (Shan & Tsai, 2021). SF-EAM proposes reporting enough representation provenance to keep an algorithmic estimate distinguishable from direct observation.
Table 2. Score-Following Educational Action Matrix (SF-EAM): proposed and unvalidated.
Table 2. Score-Following Educational Action Matrix (SF-EAM): proposed and unvalidated.
Educational action Operational question Minimum evidence Warranted output Main failure / ambiguity Verification Inference boundary
Navigation Where is the performance now? Located state; reliable score/display mapping Cursor, page, system, measure highlight; score-linked replay Repeated material; wrong visual mapping; latency Confirm against performance/score when uncertainty appears Understanding, correctness, learning
Recovery Can the follower establish a new position after continuity breaks? Discontinuity detected; relocalizing → stable new location Suspend score-dependent labels; resume from verified new position; mark discontinuity False jump; long ambiguity; pedal/reverb; polyphonic mismatch Allow more evidence or manual location confirmation Learner recovery, resilience, error diagnosis
Repetition Was material revisited, and how? Stable alignment path across returns Visit count, segment loops, attempt order, timing descriptors Intentional vs accidental restart; repeated-score ambiguity Pair trace with audio, goal, annotation or teacher/student explanation Practice quality, weakness, deliberateness
Practice support How can aligned session history support review? Trace with provenance, uncertainty, and segment history Coverage map, excerpt browser, linked replay, annotations, bounded comparisons Missing goals; misalignment; metric overinterpretation; bundled feedback Combine trace with listening, goals, reflection, teacher judgment Mastery, SRL, attention, motivation, learning gain
Note. SF-EAM is a conceptual framework and has not been validated as an assessment instrument, diagnostic model, causal model, or predictor of learning.
Design Principles for Piano Instruction
The principles in this section are authorial design propositions drawn from the preceding evidence synthesis. They are not a validated instructional protocol and should not be read as effects demonstrated by the cited technical literature.
The first design proposition: expose uncertainty rather than animate it away. When the follower is unlocated or ambiguous, the interface can pause score-dependent feedback or explicitly show that it is locating rather than displaying a unique position without sufficient evidence. This operationalizes the distinction between a supported position estimate and an unresolved hypothesis; its usability and effect on trust still need empirical testing.
Second, SF-EAM labels technical recovery and learner recovery separately. A system message such as “position recovered at measure 48” can describe the tracker, when its location evidence supports that statement. “You recovered well” is a performance judgment that needs independent musical or behavioral criteria. The distinction is conceptual and evidence-based; its value as interface wording hasn’t been experimentally evaluated yet.
Third, SF-EAM proposes gating score-dependent error feedback by localization. Historical systems establish that score-referenced error detection and feedback are feasible educational designs (Dannenberg et al., 1990; Han et al., 2013; Acquilino & Scavone, 2022). SF-EAM’s additional rule is authorial: when the reference location is ambiguous or being re-established, the system withholds a score-dependent verdict rather than convert uncertain alignment into false precision.
Fourth, repetition is represented as history, not rank. Visit maps and loop counts can describe patterns of work, but Duke et al. (2009) show why practice amount alone is an inadequate proxy for next-day retention quality in their advanced-pianist task. SF-EAM therefore favors comparing attempts, listening, and goal annotation over an automated productivity score based only on repetition quantity.
Fifth, practice-support systems should tolerate non-linear learner action. This is grounded technically in systems designed for jumps, restarts, repeats, and rehearsal (Arzt & Widmer, 2010; Arzt et al., 2014; Nakamura et al., 2016) and in offline reconstructions of realistic practice (Jiang et al., 2019; Raphael, 2025). Whether such tolerance improves learner agency or practice quality isn’t established by those technical studies.
Sixth, the reference representation should be treated as provenance, not invisible truth. A symbolic score, MIDI file, scanned PDF, or image-derived representation can encode different information and introduce different mapping errors. If visual cursor geometry was estimated separately from musical alignment, the interface shouldn’t imply notehead-level precision unless that mapping is actually supported.
Seventh, SF-EAM proposes that practice summaries retain access to the underlying attempt. A heat map or visit count can serve as a navigation index, but a pedagogical interpretation should stay traceable to the audio/performance evidence and, where relevant, to the learner’s or teacher’s stated goal — an auditability principle, not a demonstrated learning effect.
Finally, the framework does not prescribe real-time feedback as universally superior. The reviewed score-following literature establishes technical real-time capabilities, while adjacent human studies show digital-feedback effects depend on intervention composition and outcome measure: Nusseck et al. (2025) found positive changes in some student self-ratings but no significant effects in professional external ratings, whereas Rom and Woody (2026) found short-term accuracy gains from bundled digital scaffolds in string players. Neither answers whether real-time score-following feedback beats delayed review in piano, so timing stays an open empirical design question.
When Score Following Can Mislead
SF-EAM treats confident mislocalization as a higher epistemic risk than temporary uncertainty, because downstream labels depend on the selected reference position. If a system silently relocates to a repeated passage, later score-dependent judgments can get attached to the wrong location. This is a framework-level risk analysis, not a demonstrated claim about user harm — it motivates explicit uncertainty and relocalization states.
Piano acoustics create another risk. Li and Duan (2016) show that sustain pedal and reverberation extend note energy past notated durations and can create delay errors for score followers. A pedagogical system that reads every alignment lag as poor timing would be confusing an acoustic property of the instrument for a learner’s fault.
Discontinuities create a different kind of ambiguity. Shan and Tsai (2021) identify repeats and jumps as a bottleneck in real-world piano alignment, while Nakamura et al. (2016) explicitly model arbitrary repeats and skips. A follower lacking such mechanisms may look stable only because the performer happened to behave linearly. Educational use should be tested against the non-linear behavior real practice actually contains.
Benchmark accuracy can mislead too, when carried directly into pedagogy. Park et al. (2025) provide a systematic real-time piano benchmark, but a benchmark metric answers whether an algorithm tracks a reference under defined datasets and evaluation tolerances — not whether a displayed cursor is useful to a learner, whether feedback lands at an instructionally appropriate moment, or whether teachers agree with the resulting interpretation.
Within SF-EAM, practice analytics count as high-inference outputs, since location history doesn’t by itself encode goals, attention, technical choice, or self-evaluation. Self-regulation studies rely on additional behavioral or reflective evidence to get at such constructs (Nielsen, 2001; Miksza et al., 2018; Miksza & Brenner, 2023). A coverage map can establish that a passage was visited, and a sequence of short repetitions can establish repeated returns — but whether those events reflect deliberate refinement, unresolved difficulty, or something else is a separate pedagogical judgment.
Bundled educational technologies can also obscure attribution. Ou et al. (2025) offer valuable evidence that a score-following-enabled practice application can be part of a broader learning environment associated with human outcomes, but the app bundled multiple forms of feedback and instruction — a causal claim that score following itself produced the gains would go beyond what the design supports.
Interface incentives may shape what students attend to during practice, though the direction and size of that influence remain empirical questions. A design that rewards full-run completion, labels frequently visited regions as problems, or compresses practice into a single score may nudge users toward readings the underlying trace doesn’t actually warrant. SF-EAM treats such representations as hypotheses about practice support, not established pedagogical effects.
Applications for Piano Teaching
The applications below are constructed pedagogical scenarios derived from SF-EAM and the reviewed technical affordances. They illustrate bounded uses of the framework — they are not intervention effects reported by the cited studies.
For a beginning student, the lowest-risk application is orientation. During a slow reading task, the interface can hold the current measure while letting the student stop or change tempo. If location becomes ambiguous, the cursor can pause rather than penalize notes. The teacher can use the same position cue to point to the exact place in the score without treating the cue as an assessment.
For passage work, the system can record local loops. A teacher might assign measures 18–22 and ask the student to compare three attempts. Score following can locate the attempts and link them to audio; the teacher and student then decide whether rhythm, fingering, balance, articulation, pedaling, or another goal changed. The educational value comes from the comparison and discussion, not from the loop count itself.
For recovery practice, a teacher can intentionally ask the student to start from several landmarks or continue after a simulated interruption. The follower tests its own ability to relocate and provides score access once stable — but any evaluation of the student’s recovery strategy has to use musical criteria independent of the mere fact that the software regained position.
For lesson review, an aligned practice or lesson recording can become a score-indexed archive. A student can click a measure and hear the most recent attempt, compare earlier and later versions, or attach a note describing what was being practiced — extending the cursor from a transient display into a navigation structure for reflection.
For teacher education, SF-EAM works as an evidence-literacy exercise. Preservice or in-service teachers can classify statements such as “the follower relocalized after the repeat,” “the student recovered from the mistake,” “measure 24 was repeated six times,” and “measure 24 is the student’s main weakness.” The first and third can be direct system descriptions under appropriate alignment evidence; the second and fourth require pedagogical interpretation.
For research, score following can serve as instrumentation rather than intervention. Miksza and Brenner (2023) used offline following to document measures played while studying self-regulated practice. Similar approaches can quantify location history in piano practice while still relying on diaries, interviews, teacher judgments, or audio analysis for constructs that position alone can’t represent.
The National Association for Music Education (NAfME, 2022) emphasizes pathways for practitioners to understand and apply research and effective practices. SF-EAM translates that priority into an evidence-literacy routine: use alignment to organize access, then reserve musical and learning judgments for evidence that can actually support them.
Limitations and Research Agenda
This review is limited by its structured applied design. It is not a systematic review, and the search cannot prove that no equivalent framework exists anywhere. Technical score-following literature is extensive and uses shifting terminology across accompaniment, alignment, tracking, synchronization, page turning, rehearsal, tutoring, and practice analysis. Within the expanded search and claim-level audit conducted here, no prior source combines, in one educational reporting matrix, explicit alignment states, evidence provenance, permitted downstream action, verification requirement, and prohibited pedagogical inference across navigation, system recovery, repetition, and practice support. This is a search-bounded observation, not proof of global priority.
The evidence base is heterogeneous. Some studies evaluate alignment algorithms, some describe educational systems, some investigate music practice without score following at all, and a smaller set includes human learners using score-following-enabled or adjacent digital-feedback applications. These sources shouldn’t be pooled into one effect estimate. Ou et al. (2025) provide direct human evidence for a multi-feature violin application that explicitly includes score following, Nusseck et al. (2025) report mixed self-rating versus external-rating findings in advanced piano feedback, and Rom and Woody (2026) report positive short-term effects of bundled digital scaffolds in high-school strings. None isolates the causal effect of real-time score following in piano — these sources define different links in an evidence chain, and, just as importantly, show where that chain breaks.
SF-EAM itself is unvalidated. Its action categories and alignment states haven’t been tested for inter-rater reliability, teacher usability, student comprehension, cognitive load, or learning effect. A first empirical study should ask piano teachers whether the distinctions among navigation, system recovery, repetition history, practice support, and pedagogical inference match how they actually make decisions.
A second research program should compare interface policies under controlled alignment uncertainty. The same follower could present a continuously moving cursor, an uncertainty-aware cursor, or a fail-closed locating state, letting researchers examine trust calibration, teacher error detection, student reliance, and willingness to override the system — testing whether epistemic transparency is usable, not just principled.
A third program should study recovery as two separate variables. Technical recovery can be measured as time or musical distance required for the follower to re-establish position after a discontinuity. Learner recovery can be evaluated with musical and behavioral criteria such as continuity, restart strategy, error correction, or successful resumption. Correlating the two would show when system behavior supports the human recovery process and when it gets in the way.
A fourth program should investigate repetition semantics. Position traces can identify returns, but studies could combine those traces with student-stated goals, think-aloud or stimulated recall, audio-performance features, and teacher ratings to distinguish corrective repetition, exploratory repetition, memorization, interpretive refinement, and unproductive restarting.
Finally, longitudinal piano studies are needed before any learning claims can be made. A score-following interface might improve access to evidence without improving practice, or it might help students reflect more effectively only after teacher modeling. Outcomes should include transfer, retention, adaptive practice choices, and the ability to interpret one’s own performance — not just follower accuracy or interface engagement.
Conclusions
Real-time score following can do far more than move a cursor, but the cursor is still a useful reminder of what the technology fundamentally provides: an estimate of where an unfolding performance corresponds to a score. Educational value begins once that estimate supports an action. Navigation uses position directly. Recovery requires the system to admit continuity has broken and establish a new position. Repetition uses the path history to identify returns. Practice support turns aligned events into material for review and reflection.
The literature also makes the limits clear. Score following has decades of technical history, and educational or practice-oriented uses substantially predate this review (Orio et al., 2003; Tekin et al., 2005; Acquilino & Scavone, 2022). Real-time systems have supported page turning, beginner feedback, jumps, repeats, restarts, rehearsal, and piano tracking (Arzt et al., 2008; Arzt & Widmer, 2010; Arzt et al., 2014; Han et al., 2013; Nakamura et al., 2016). Offline systems already reconstruct structural deviations and non-linear practice and support score-driven review (Fremerey et al., 2010; Jiang et al., 2019; Raphael, 2025). Piano-specific work shows that sustained sound and discontinuities complicate localization, while practice research shows that repetition quantity alone can’t account for retention or learning.
SF-EAM responds by placing an explicit evidence boundary between alignment and pedagogy. A position estimate may justify moving a page. A recovered position may justify resuming score-linked interaction. A repeated segment may justify displaying a revisit. A practice trace may justify a review interface. None of these observations, on its own, establishes understanding, resilience, deliberate strategy, mastery, self-regulation, or learning. Beyond the cursor, the central design question is therefore not just whether the system can follow the pianist — it’s whether the system can state honestly what following actually allows it to know.

References

  1. Acquilino, A.; Scavone, G. Current state and future directions of technologies for music instrument pedagogy. Frontiers in Psychology 2022, 13, 835609. [Google Scholar] [CrossRef] [PubMed]
  2. Arzt, A.; Böck, S.; Flossmann, S.; Frostel, H.; Gasser, M.; Liem, C. C. S.; Widmer, G. The piano music companion. In Proceedings of the European Conference on Artificial Intelligence; IOS Press, 2014; p. 1221--1222. [Google Scholar] [CrossRef]
  3. Arzt, A.; Widmer, G. Towards effective “any-time” music tracking. In Proceedings of the Starting AI Researchers’ Symposium; IOS Press, 2010; p. 24--36. [Google Scholar]
  4. Arzt, A.; Widmer, G.; Dixon, S. Automatic page turning for musicians via real-time machine listening. In Proceedings of the European Conference on Artificial Intelligence; IOS Press, 2008; p. 241--245. [Google Scholar] [CrossRef]
  5. Dannenberg, R. B.; Raphael, C. Music score alignment and computer accompaniment. Communications of the ACM 2006, 49(8), 38--43. [Google Scholar] [CrossRef]
  6. Dannenberg, R. B.; Sanchez, M.; Joseph, A.; Capell, P.; Joseph, R.; Saul, R. A computer-based multi-media tutor for beginning piano students. Interface - Journal of New Music Research 1990, 19(2--3), 155--173. [Google Scholar] [CrossRef]
  7. Duke, R. A.; Simmons, A. L.; Cash, C. D. It’s not how much; it’s how: characteristics of practice behavior and retention of performance skills. Journal of Research in Music Education 2009, 56(4), 310--321. [Google Scholar] [CrossRef]
  8. Fremerey, C.; Müller, M.; Clausen, M. Handling repeats and jumps in score-performance synchronization. In Proceedings of the International Society for Music Information Retrieval Conference; 2010; p. 243--248. [Google Scholar]
  9. Han, Y.; Kwon, S.; Lee, K.; Lee, K. A musical performance evaluation system for beginner musician based on real-time score following. In Proceedings of the International Conference on New Interfaces for Musical Expression; KAIST, 2013; p. 120--121. [Google Scholar] [CrossRef]
  10. Henkel, F.; Balke, S.; Dorfer, M.; Widmer, G. Score following as a multi-modal reinforcement learning problem. Transactions of the International Society for Music Information Retrieval 2019, 2(1), 67--81. [Google Scholar] [CrossRef]
  11. How, E. R.; Tan, L.; Miksza, P. A PRISMA review of research on music practice. Musicae Scientiae 2022, 26(3), 675--697. [Google Scholar] [CrossRef]
  12. Jiang, Y.; Ryan, F.; Cartledge, D.; Raphael, C. Offline score alignment for realistic music practice. In Proceedings of the Sound and Music Computing Conference; 2019; p. 387--393. [Google Scholar] [CrossRef]
  13. Li, B.; Duan, Z. An approach to score following for piano performances with the sustained effect. IEEE/ACM Transactions on Audio, Speech, and Language Processing 2016, 24(12), 2425--2438. [Google Scholar] [CrossRef]
  14. Miksza, P.; Blackwell, J.; Roseth, N. E. Self-regulated music practice: microanalysis as a data collection technique and inspiration for pedagogical intervention. Journal of Research in Music Education 2018, 66(3), 295--319. [Google Scholar] [CrossRef]
  15. Miksza, P.; Brenner, B. A descriptive study of intra-individual change in advanced violinists’ music practice. Update: Applications of Research in Music Education 2023, 41(3), 22--36. [Google Scholar] [CrossRef]
  16. Nakamura, T.; Nakamura, E.; Sagayama, S. Real-time audio-to-score alignment of music performances containing errors and arbitrary repeats and skips. IEEE/ACM Transactions on Audio, Speech, and Language Processing 2016, 24(2), 329--339. [Google Scholar] [CrossRef]
  17. National Association for Music Education. 2022 strategic plan . 2022. Available online: https://nafme.org/wp-content/uploads/2023/03/NAfME-2022-Strategic-Plan.pdf.
  18. Nielsen, S. Self-regulating learning strategies in instrumental music practice. Music Education Research 2001, 3(2), 155--167. [Google Scholar] [CrossRef]
  19. Nusseck, M.; Wild, F.; Sischka, C.; Spahn, C. Effects of audio feedback interventions with the Disklavier on the performance of piano students. Frontiers in Psychology 2025, 16, 1568021. [Google Scholar] [CrossRef] [PubMed]
  20. Orio, N.; Lemouton, S.; Schwarz, D. Score following: state of the art and new developments. In Proceedings of the International Conference on New Interfaces for Musical Expression; 2003; p. 36--41. [Google Scholar] [CrossRef]
  21. Ou, J.; Nogueira, J.; Qin, C. Exploring the impact of AI-assisted practice applications on music learners’ performance, self-efficacy, and self-regulated learning. Frontiers in Psychology 2025, 16, 1675762. [Google Scholar] [CrossRef] [PubMed]
  22. Park, J.; Cancino-Chacón, C. E.; Chiruthapudi, S.; Nam, J. A systematic evaluation of real-time audio score following for piano performance. In Proceedings of the International Society for Music Information Retrieval Conference; ISMIR, 2025; p. 105--113. [Google Scholar] [CrossRef]
  23. Puckette, M.; Lippe, C. Score following in practice. In Proceedings of the International Computer Music Conference; International Computer Music Association, 1992; p. 182--185. [Google Scholar]
  24. Raphael, C. Face the music: summarizing unscripted music practice from audio. In Proceedings of the International Conference on Computer Supported Education; 2025; p. 685--691. [Google Scholar] [CrossRef]
  25. Rom, B.; Woody, R. H. The effect of digital scaffolds on performance gains made in practice by high school string players. Journal of Research in Music Education 2026. [Google Scholar] [CrossRef]
  26. Shan, M.; Tsai, T. J. Automatic generation of piano score following videos. Transactions of the International Society for Music Information Retrieval 2021, 4(1), 29--41. [Google Scholar] [CrossRef]
  27. Tekin, M. E.; Anagnostopoulou, C.; Tomita, Y. Towards an intelligent score following system: handling of mistakes and jumps encountered during piano practicing. In Computer Music Modeling And Retrieval (Lecture Notes in Computer Science; Springer, 2005; Vol. 3310, p. 211--219. [Google Scholar] [CrossRef]
  28. Whittemore, R.; Knafl, K. The integrative review: updated methodology. Journal of Advanced Nursing 2005, 52(5), 546--553. [Google Scholar] [CrossRef] [PubMed]
  29. Yang, Y.; Chen, R.; Han, J. CODA: Cascaded Online Discontinuity-Aware Alignment for Real-Time Image-Based Score Following [Preprint]. arXiv 2026. https://arxiv.org/abs/2607.21899. [Google Scholar]
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.