Submitted:
21 August 2026
Posted:
24 August 2026
You are already at the latest version
Abstract
A one-to-one music lesson normally unfolds in a shared room, where teacher and student can move, reposition, and inspect the instrument from any angle. Synchronous online teaching compresses all of that into a fixed rectangle. This structured applied literature review asks what visual information teachers actually need during instrumental instruction, and how camera framing might be organized around the teaching task rather than around whatever equipment happens to be on hand. Evidence drawn from distance music teaching, gesture research, teacher observation, visual attention studies, and educational technology shows that bodily modeling, hand and arm movement, posture, facial communication, score reference, and contact with the instrument all carry pedagogically relevant information. But camera position by itself does not reliably determine what a teacher notices or how that teacher evaluates instruction — seeing more is not the same as noticing more. The review proposes the Pedagogical Visual Access Matrix (PVAM), an unvalidated framework linking pedagogical task, visual target, candidate view, information loss, and complementary view. PVAM is meant to make camera-design decisions explicit, while keeping a firm line between visual access, teacher noticing, pedagogical interpretation, and what actually happens to learning.
Keywords:
instrumental music education
; online music teaching
; camera framing
; visual attention
; gesture
; multiview
; piano pedagogy
Introduction
Instrumental teaching runs on more than sound and speech. A teacher listens for intonation, rhythm, articulation, tone, and phrasing, while also tracking posture, breathing, hand position, arm movement, embouchure, bow path, pedal use, gaze, and the student's relationship with the instrument and the score. Teachers use their own bodies too — modeling actions, pointing, gesturing, demonstrating, coordinating verbal explanation with movement. Research on one-to-one music teaching has long treated this physical, nonverbal layer as part of the pedagogical exchange itself, not as background noise (Bremmer & Nijs, 2020; Kurkul, 2007; Simones, 2019; Simones, Rodger, et al., 2015; Simones, Schroeder, et al., 2015).
Synchronous online instruction changes the terms under which all of this happens. A face-to-face lesson unfolds in a shared three-dimensional space: teacher and student can reposition themselves, shift attention, look at the instrument from a different angle, and move between score, body, and instrument without touching a single button. Videoconferencing collapses that space into one or more camera frames. Whatever falls outside the frame — or is too small to read, or blocked, or pushed off to a different view — simply isn't there at that moment.
Early distance-music research already found both the promise and the limits of this shift. Dammers (2009), writing in Update: Applications of Research in Music Education, found that internet-based trumpet lessons could work at a basic level, but video delay and limited visual control got in the way. Orman and Whitaker (2010) then compared face-to-face and synchronous distance lessons across more than 28,800 video frames — teacher modeling, focus of attention, eye contact, student performance time, and other nonverbal behaviors. Riley et al. (2016) later showed that lower-latency audiovisual systems could support kinds of distance music-making that ordinary videoconferencing simply couldn't handle. Taken together, these studies make one point clearly: remote music teaching is not face-to-face teaching piped through a neutral channel. The medium itself changes what can be heard, seen, coordinated, and acted on.
Technology has since widened the visual channels available to teachers. King et al. (2019) used an audiovisual mixer that gave three camera angles inside digitally delivered instrumental lessons. Piano-specific practitioner literature covers similar ground: Casarotti (2021) wrote about visual communication through virtual webcams and multiple angles, and Hamond (2021) described combining OBS Studio, two camera angles, and other visual resources in online higher-education piano teaching. More recent accounts describe teachers stitching together videoconferencing, screen-based resources, multiple camera angles, and score display all at once (Stephens-Himonides & Young, 2025). None of this is new — multiview teaching and task-sensitive visual communication are established practice. What remains unresolved is narrower: making explicit which visual target a given task actually requires, what a chosen view necessarily hides, when a second view earns its place on screen, and what can't be inferred from visibility alone.
That last point matters because seeing more isn't the same as noticing more, and noticing isn't the same as interpreting correctly. Studies of music-teacher observation show that visual focus and camera placement don't automatically shape what observers conclude. Duke and Prickett (1987) manipulated the visual focus of applied-music instruction and tracked how observers evaluated it. Madsen and Cassidy (2005) found that observers kept commenting more on the teacher even when the video focused on students. Buonviri and Paney (2022), more recently, found no significant effect of camera placement on preservice teachers' reflection comments. Hicken and Duke (2023), using eye-tracking, found the opposite kind of result — differences in attention allocation tied to teaching experience and expertise. Read together, these findings argue against treating a camera angle as an intervention with a predictable pedagogical effect.
So the more useful question isn't "which camera is best" or "how many cameras does a lesson need." It's "what visual information does this particular pedagogical task actually require?" This structured applied literature review works through that question, pulling together evidence from synchronous instrumental teaching, piano pedagogy, embodied and gestural teaching, teacher observation, visual attention, and educational technology. Specifically: across empirical and conceptual literature, what visual information is tied to observation, demonstration, gesture, and feedback in instrumental instruction — and what does that suggest for camera-framing and multiview decisions in synchronous teaching? The review builds toward the Pedagogical Visual Access Matrix (PVAM), an author-proposed and unvalidated framework mapping pedagogical tasks to visual targets, candidate views, likely information loss, and complementary views. The goal is to make visual-design decisions explicit while holding a hard boundary between visual access, teacher noticing, pedagogical interpretation, and learner outcomes.
Method
This is a structured applied literature review, not an attempt to estimate an intervention effect or claim exhaustive coverage of online music education research. The aim was narrower and more practical: synthesizing evidence directly relevant to the visual information available to teachers during synchronous instrumental instruction. The design follows Update's expectations for standalone literature reviews — a stated rationale, explicit search methods, inclusion and exclusion criteria, and an applied synthesis rather than a study-by-study catalog.
Searches ran, and were updated, in August 2026. Four families of title, abstract, and keyword terms guided searches across OpenAlex-oriented scholarly discovery tools and publisher/index platforms: (a) instrumental OR piano teaching combined with online, synchronous, videoconference, or distance and camera, video, or visual; (b) instrumental OR piano teaching combined with gesture, body, modeling, demonstration, or nonverbal communication; (c) music teacher combined with visual attention, observation, noticing, or eye tracking; and (d) camera placement, camera angle, framing, multiview, or multiple camera angles combined with music teaching or instrumental instruction. Sensitivity searches added terms like visual communication, task-specific camera angle, overhead view, side view, score/keyboard/body, and multimedia piano studio. Recent reviews of synchronous instrumental teaching and online piano education served as coverage checks and starting points for backward citation chaining (Løkke Jakobsen et al., 2025; Turan, 2026). Reference lists from the most directly relevant primary studies were also mined for earlier work on applied-music observation, gesture, and nonverbal communication.
A study made the cut if it filled at least one of four roles. Direct evidence: empirical research on synchronous or distance instrumental teaching where video, physical behavior, visual interaction, or camera configuration mattered. Mechanism evidence: empirical or conceptual work on bodily demonstration, gesture, nonverbal communication, visual attention, or observation in instrumental and music-teacher settings. Design-context evidence: reviews or teacher-centered technology studies that help mark the boundary of what technology can and can't be assumed to do pedagogically. Practice-oriented prior art — professional articles, conference experience reports — was kept when it documented concrete visual or multiview practices relevant to assessing novelty; these sources were never treated as evidence of causal educational effectiveness. Studies focused purely on asynchronous tutorial videos, general educational videoconferencing with no music-specific mechanism, or engineering performance work with no teaching connection, were left out.
The synthesis is organized around visual target, not around technology. During analysis, evidence was grouped by five recurring questions: What aspect of the learner or instrument matters pedagogically here? What view could make that target visible? What might get lost or occluded in that view? When does a second, complementary view earn its place? And — the question that shaped the whole review — what can't be inferred from visibility alone? The literature supports claims about what visual information can be made available and how teachers may distribute attention across it. It does not support a general causal claim that adding cameras improves learning, and this review doesn't make that claim either.
This is not a systematic review or a meta-analysis, and it doesn't pretend to be. It's a single-reviewer, task-centered synthesis with broad but search-bounded coverage. OpenAI ChatGPT (GPT-5.6 Sol) supported search-term expansion, evidence organization, drafting, language revision, and document preparation; none of its output was treated as a scholarly source. The PVAM itself is an authorial synthesis built from the reviewed evidence — it has not been validated as a scale, an observational instrument, a causal model, or a predictor of teaching effectiveness.
Why Visual Access Matters in Instrumental Teaching
The visual frame matters in the first place because instrumental teaching is embodied. Bremmer and Nijs (2020) describe bodily engagement in instrumental and vocal pedagogy through physical modeling, action demonstration, pedagogical gesture, and touch. That analysis matters here because these practices depend on perceptible relationships among body, movement, instrument, and where the learner's attention is. When a teacher demonstrates a bow stroke, a breathing action, a hand shape, or a movement pathway, part of the instructional content travels through motion itself.
Piano pedagogy offers unusually detailed evidence on this point. Simones, Schroeder, et al. (2015) studied one-to-one piano lessons and identified physical gestures woven into teacher-student communication. Simones, Rodger, et al. (2015) went further, reporting that teachers' gestural behavior shifted with didactic intention and student proficiency — gesture isn't a fixed decorative layer sitting on top of verbal instruction. Simones (2019) then proposed the Teacher Behaviour and Gesture framework for studying hand gestures across instrumental and vocal contexts. None of this proves every gesture must stay visible at all times. It does establish that bodily action carries pedagogical meaning, and that meaning can disappear when framing leaves it out.
Nonverbal communication reaches past the hands, too. Kurkul (2007) examined nonverbal communication in one-to-one music performance instruction, reinforcing that teacher-student interaction runs on more than spoken propositions. Daugvilaite (2021), studying students, parents, and instrumental teachers after a shift to online lessons, found that participants connected the teacher's physical absence to real losses — in nonverbal communication, gestures, scaffolding, and tactile approaches. De Bruin (2021) documented something similar: experienced instrumental teachers adapting their interaction and their "ways of showing and telling" over sustained online teaching. Together these studies place visual access inside a relational, communicative system — not just a technical imaging problem.
The design implication is that the instructional target can shift fast. A teacher might start by reading a student's facial response, then check whole-body posture, then zero in on the hand, then point to something in the score, then demonstrate the same passage. A single fixed frame gives continuity, but it can't guarantee equally useful access to every one of those targets. A very wide frame keeps context but loses detail; a close view gains detail but can strip away the bodily relationships that give that detail its meaning. So the pedagogical problem here isn't image quality in the abstract — it's the tradeoff between contextual breadth and task-relevant resolution.
What Changes in Synchronous Online Instrumental Teaching?
Distance instruction keeps exposing what happens when a studio lesson gets translated into a mediated audiovisual environment. Dammers (2009) found basic instructional functionality in Skype-based trumpet lessons, but flagged limited visual control as one real constraint. Orman and Whitaker (2010) compared face-to-face and videoconferenced private lessons and found differences in teacher modeling, student performance time, eye contact, and off-task behavior — though many focus-of-attention differences turned out small. That mixed pattern is itself informative: mediation changed some behaviors without uniformly degrading or improving the lesson.
King et al. (2019) went beyond a single webcam, pairing Skype with an audiovisual mixer capable of three camera angles. Their work matters here because it shows multiview instrumental teaching already has a real history. Piano-pedagogy sources back this up. Casarotti (2021) gave professional guidance on visual communication using virtual-webcam tools and varied perspectives; Hamond (2021) reported using OBS Studio to combine two camera angles with other visual resources in synchronous higher-education piano teaching, while also flagging the need for systematic research into whether such resources actually help learning — a boundary that still matters here. So this review's contribution can't rest on the idea of multiple cameras, or on the general notion of matching a view to instructional content; both are already established. The narrower question is how visual access gets reported and planned in a way that makes target, occlusion, complementarity, and the limits of inference explicit.
In piano teaching specifically, Comeau et al. (2019) analyzed verbal and physical behaviors across on-site and distance lessons given by the same teacher. Their design surfaces something important: behaviors that are trivially simple to produce in a shared room can require explicit repositioning, verbalization, or mediation once distance enters the picture. Daugvilaite (2021) found something parallel — online piano students could grow more independent while still reporting real consequences tied to the absence of physical and nonverbal interaction. None of this supports either technological pessimism or technological triumphalism. Online teaching can support real learning; it just changes the conditions under which teacher and student coordinate visually.
Low-latency systems make a related point from a different angle. Riley et al. (2016) found that participants rated LOLA as more effective than PolyCom and Skype for certain synchronous performance situations. That result is about latency and audiovisual interaction, not camera framing per se, but it points to a broader design principle: a technical feature earns pedagogical relevance when it addresses a real demand of the musical task. Camera design should follow the same logic. More visual channels aren't valuable on their own — they're resources whose worth depends entirely on the task they're serving.
From Camera Placement to Pedagogical Visual Access
"Camera placement" sounds like a hardware decision — front, side, overhead, close, wide. The literature reviewed here suggests that's too coarse a unit of analysis. A camera position only affects pedagogy indirectly. Its immediate effect is to change what information sits inside the frame; everything after that depends on attention, expertise, interpretation, and action.
Observation research makes this visible. Duke and Prickett (1987) manipulated what observers could see while they evaluated one-to-one violin instruction. Madsen and Cassidy (2005) compared teacher-focused and student-focused videotaped music classes and found that participants — across experience levels — kept commenting more on the teacher regardless of where the video's focus sat; the focus of observation made no measurable difference. Buonviri and Paney (2022) compared head-mounted and tripod-mounted recordings of preservice peer teaching and, again, found no significant difference in comment frequency tied to camera placement. None of these studies concerns synchronous instrumental lessons narrowly, but all of them bear directly on any claim that changing the view will automatically redirect pedagogical attention.
Hicken and Duke (2023) supply the complementary piece. Using wearable eye tracking, they studied gaze behavior in music teachers with different experience and expertise levels as they watched brief performance videos, and found real differences in attention allocation and visual scan patterns. That doesn't mean expertise guarantees a correct diagnosis. It means the same visual scene gets sampled differently by different observers. Visual access is necessary for noticing something — it just isn't sufficient to determine what gets noticed.
Put these findings together and a four-stage boundary emerges: visual availability, teacher noticing, pedagogical interpretation, pedagogical action. A camera shapes the first stage directly, by determining what's visible. It can influence — but not determine — the second, by making some regions prominent and others harder to see. The third stage runs on professional knowledge, musical context, learner history, and how reliable the observable cue actually is. The fourth depends on instructional judgment. Learning outcomes sit beyond all four stages, and nothing about the camera setup alone lets you infer them.
This boundary also explains why multiview design should stay selective. Display several simultaneous views at equal prominence, and the system may add available information while also adding competition for attention. The music literature reviewed here offers no general threshold for when that competition turns harmful — so a claim that more simultaneous views are simply better wouldn't be supported by current sources. A more defensible principle is complementarity: bring in an additional view when it supplies information the primary view is materially missing.
Pedagogical Visual Targets
The synthesis surfaced a recurring set of visual targets that can organize camera decisions without locking the framework to one instrument or one platform. These targets overlap — they aren't independent factors, and shouldn't be read as a psychometric scale.
Face and relational communication. Face, gaze direction, and upper-body behavior support interpersonal coordination and let a teacher gauge whether an explanation is landing, confusing, effortful, or losing the student. Daugvilaite (2021) and de Bruin (2021) both put communication and relational connection at the center of online instrumental teaching. A front-facing or moderately wide view earns its place when the immediate task is conversation, explanation, affective attunement, or turn-taking. The tradeoff is detail: a frame built for face and torso may render finger action or pedal behavior too small to actually inspect.
Whole-body posture and coordination. Instrumental actions sit inside larger bodily organizations. Bremmer and Nijs (2020) emphasize physical modeling and action demonstration; technology reviews note visual systems built for movement and posture-related feedback (Acquilino & Scavone, 2022). Many tasks call for relationships rather than isolated body parts — head to torso, shoulder to arm, seating to instrument, breathing movement to playing action. A wide frontal or lateral view preserves those relationships, at the cost of fine-grained visibility.
Hands, arms, and contact with the instrument. Piano gesture studies show teachers communicating through hand and arm movement, adapting gesture to instructional intention (Simones, Rodger, et al., 2015; Simones, Schroeder, et al., 2015). Close or overhead views can make finger patterns, hand shape, key contact, bow path, stick trajectory, or valve interaction easier to read. But a close view can also make it impossible to tell whether that local action is coordinated with the rest of the body. "Hands visible" and "technique visible," in other words, aren't synonyms.
Instrument-specific interfaces. Some visual targets belong to the instrument itself — the full keyboard, a bow-string contact point, a percussion striking zone, a guitarist's left-hand position, a wind player's embouchure. These justify task-specific framing because they occupy different spatial regions and need different resolution. The view that shows a pianist's entire keyboard may render the face poorly; the view built for embouchure may cut off everything below the shoulders. The PVAM treats candidate views as instrument- and task-dependent rather than universal, for exactly this reason.
Pedal, feet, and lower-body actions. Pedal use on piano and other foot-controlled instruments is a clean case of visual occlusion — a conventional upper-body webcam can omit the action entirely. The camera literature reviewed here doesn't establish that a dedicated pedal view improves learning, so this framework doesn't claim that either. It simply classifies pedal action as information that stays unavailable unless the lower-body region is deliberately included, or supplied through a second view.
Score and shared reference. Instrumental teaching often means directing attention to a specific spot in the notation. Orman and Whitaker (2010) noted pointing to specific locations in the music as one of the face-to-face behaviors their analysis captured; remote environments often need alternative ways to establish that same shared reference. A score view, screen share, or document camera increases shared access to notation, but it competes for screen space with body and instrument views. The task-centered question is whether the current instructional action depends more on reading a location, executing a physical action, or communicating relationally.
Teacher demonstration. Demonstration flips the usual direction of observation — the student's need to see the teacher becomes primary. Bremmer and Nijs (2020) and Simones (2019) both support treating bodily demonstration and gesture as meaningful channels in their own right. A multiview environment can therefore be symmetrical in capability while asymmetrical in use: during student performance the teacher may need a full-body or hand view of the learner, while during teacher modeling, the learner may need a close view of the teacher's relevant movement. Designing views around pedagogical roles is a more precise approach than assigning one fixed "teacher camera" and one fixed "student camera."
The Pedagogical Visual Access Matrix
The Pedagogical Visual Access Matrix (PVAM) is proposed here as a reporting and planning framework built from the synthesis above. It is not a validated scale, an assessment instrument, a causal model, or a claim that task-specific camera selection is itself a new idea — multiple-camera teaching and task-sensitive visual communication are established prior practice (Casarotti, 2021; Hamond, 2021; King et al., 2019). What's new, if anything, is combining visual target, candidate view, information loss or occlusion, complementary view, and an inference boundary into one task-centered structure.
PVAM has six fields: pedagogical task, visual target, required visibility, candidate primary view, information potentially lost or occluded, and complementary view. A seventh field — an inference boundary — can be added wherever a teacher or designer might otherwise read too much into what the image shows. Seeing a student's wrist angle, for instance, doesn't by itself establish the cause of a technical problem; seeing a facial expression doesn't establish motivation. The matrix is built to stay descriptive before it turns interpretive.
The underlying logic is "primary plus complementary." One view should normally carry the main instructional target. A complementary view earns its place only when a second target can't be adequately represented in the primary frame and matters to the immediate task. This keeps multiview design from turning into a contest to cram in as many feeds as possible, and it gives a clear rationale for switching views: a switch is pedagogically motivated when the instructional target itself changes, not for its own sake.
The matrix also supports preparation before a lesson even starts. A teacher can identify the recurring tasks in a repertoire or technique unit and ask whether the default framing gives adequate access to each one. A piano teacher working on seating and arm organization might prioritize a side or wide view; when the task shifts to fingering or hand redistribution, an overhead or closer view becomes more useful. These are applications of the framework — not evidence that any particular configuration improves outcomes.
Table 1.
Pedagogical Visual Access Matrix (PVAM): Proposed, Unvalidated Framework.
| Task | Target | Need | View | Loss | Alt | Limit |
|---|---|---|---|---|---|---|
| Relational communication / explanation | Face, gaze, upper torso | Facial and conversational cues visible at useful size | Front / medium-wide | Hands, lower body, pedals may be small or absent | Close hand or lower-body view when task changes | Visibility does not establish engagement or motivation |
| Whole-body posture / coordination | Head, torso, shoulders, arms, instrument relationship | Body relationships visible simultaneously | Wide front or side | Fine finger or contact detail reduced | Close hand / instrument view | Posture image alone does not establish cause or discomfort |
| Hand / finger technique | Hands, fingers, contact points | Fine movement and instrument contact visible | Overhead or close oblique | Face and whole-body context reduced | Wide or side context view | Local movement does not equal complete technique diagnosis |
| Register / instrument navigation | Full relevant instrument span plus both hands | Spatial travel and redistribution visible | Overhead / wide instrument view | Facial and lower-body information reduced | Front or side view | Visible route does not reveal musical intention by itself |
| Embouchure / bow / striking interface | Instrument-specific contact region | Relevant interface visible at close resolution | Close task-specific view | Global posture often excluded | Wide / side view | Close detail can hide upstream body organization |
| Pedal / foot action | Foot, pedal, lower leg | Press, release, depth/coordination visible | Low side / dedicated pedal view | Hands, face, score absent | Hand or whole-body view | Pedal motion alone does not establish acoustic or musical adequacy |
| Score reference / reading | Shared notation and current location | Notation legible to both participants | Screen share / document view | Body and instrument may shrink or disappear | Body / instrument view | Score location does not reveal physical execution |
| Teacher modeling | Teacher's task-relevant movement | Demonstrated action visible at useful scale | View matched to modeled action | Other teacher/student cues may disappear | Return to relational/default view | Demonstration visibility does not prove learner uptake |
When More Views May Not Be Better
Multiview systems look attractive because they seem to restore information a single webcam loses. The literature counsels caution, though. Camera placement didn't redirect preservice teachers' reflective focus in Buonviri and Paney (2022); visual focus didn't determine ratings in Madsen and Cassidy (2005). The expertise-related differences Hicken and Duke (2023) found suggest that adding visual information doesn't standardize what observers actually attend to.
Several practical risks follow from this. Simultaneous views compete for limited display space, so each one can end up too small to show the feature it's meant to show. Frequent switching can break continuity if learner or teacher has to keep reorienting to a different spatial layout. Technical management eats into instructional attention — especially when a teacher has to control cameras, audio, screen sharing, and notation while also listening and teaching at the same time; King et al. (2019) documented exactly this kind of strain even in a deliberately enhanced audiovisual setup. And a close-up can invite overconfidence: detail becomes visible while the causal context that would explain it stays out of frame.
For these reasons, PVAM doesn't prescribe a fixed number of cameras. One camera can be enough for a task if it captures the relevant target at usable resolution. Two views may be necessary when a task depends on spatially separated targets — face and pedal, say, or whole-body posture alongside hand detail. Three or more views might be justified in some settings, but nothing in the current literature supports a general rule that each additional view adds pedagogical benefit.
This is also where teacher-centered technology design earns its keep. Michałko et al. (2022) found that involving instrumental teachers directly in technology design reduces the mismatch between what a platform does and what users actually need. Acquilino and Scavone (2022) survey a wide range of technologies for instrument pedagogy while flagging how varied the functions and learning contexts really are. A task-centered visual framework fits this orientation: technology should answer a pedagogical requirement that can be named, not manufacture a requirement out of whatever features a platform happens to ship with.
Applications for Music Teaching
The first application here is diagnostic preparation — not diagnosis of the student, but preparation before the lesson even begins. A teacher can name, in advance, what kind of information the planned work will actually require. If the lesson centers on posture, breathing, or whole-arm coordination, a frame that shows only face and hands is under-specified. If the lesson centers on precise fingering, a distant full-body shot is equally under-specified. PVAM turns these mismatches into explicit design decisions instead of accidents of default settings.
Second, teachers can settle on a stable default view and save changes for moments when the pedagogical target genuinely shifts. Stability matters because students shouldn't have to decode a constantly changing interface on top of processing musical feedback. A default relational or whole-body view supports conversation and overall coordination; a close-up gets introduced for a specific demonstration or technical check, then dismissed once its job is done.
Third, teachers can say out loud why they're changing the view. "I'm switching to the side view so I can see the relationship between your shoulder, elbow, and wrist" or "show me the full keyboard for this passage — I want to watch how you move between registers" makes the observational logic transparent to the learner. It also signals something to the student: camera positioning isn't surveillance, it's a way of making a specific musical or physical problem visible.
Fourth, teachers should keep what's visible separate from what's inferred. A frame might show a collapsed hand shape, a lifted shoulder, a late pedal release — none of which automatically reveals cause, discomfort, understanding, or the right intervention. The image is one source of evidence among several; a teacher can use it, then ask questions, listen to the resulting sound, request a repetition, or switch the view before drawing any conclusion.
Fifth, institutions and teacher-education programs can teach camera literacy as part of online pedagogy — and that means more than showing someone how to plug in a webcam. Teachers can practice identifying visual targets, choosing framing, recognizing occlusion, and reflecting on what a camera simply can't show. Observation studies like Duke and Prickett (1987), Madsen and Cassidy (2005), Buonviri and Paney (2022), and Hicken and Duke (2023) are useful training material here precisely because they show that professional seeing involves both access and attention — two different things.
Finally, platform designers could make pedagogical roles legible in the interface itself. Labels like "front," "side," and "overhead" describe geometry; labels like "posture," "hands," or "pedal" describe intended information. A system could support either approach, but the second one nudges users to think in terms of task rather than hardware. That's a design proposition drawn from the literature reviewed here — not an empirically validated claim of superiority.
The National Association for Music Education (NAfME, 2022) emphasizes practitioner pathways for understanding and applying research and effective practice. PVAM serves that same professional-learning goal by turning camera setup into an explicit evidence question rather than a device preference.
Limitations and Research Agenda
This review has real limits. It's a structured applied review conducted by a single reviewer — not a systematic review or meta-analysis. The search was deliberately broad across adjacent literatures but focused on one narrow applied problem, and equivalent concepts may well exist under terminology the search strings didn't catch. Task-specific camera recommendations and multiview visual communication clearly predate PVAM in professional, conference, and empirical literature (Casarotti, 2021; Hamond, 2021; King et al., 2019). Within the sources this search identified, no research-synthesis framework explicitly combined pedagogical task, visual target, candidate primary view, information loss or occlusion, complementary view, and an inference boundary — but that's a search-bounded observation, not a claim of priority, and it's certainly not proof that no equivalent framework exists elsewhere.
The evidence base is also uneven by design. Direct synchronous instrumental studies, piano case studies, conceptual work on embodiment, teacher-observation experiments, eye-tracking research, and technology reviews answer genuinely different questions. They're combined here to illuminate mechanisms and design constraints — not pooled as if they measured the same outcome. Evidence from classroom observation and preservice reflection, in particular, informs the boundary between framing and attention without directly telling us what a studio teacher notices during a live lesson.
PVAM itself remains unvalidated. Future work should start with descriptive validity: do experienced instrumental teachers actually agree on the visual targets required for common tasks? Do those targets shift by instrument, repertoire, learner level, or teaching philosophy? A second step could test perceptual sufficiency — can teachers reliably identify specified observable features across different framings? Only after that should research ask whether task-matched visual configurations change feedback quality, lesson efficiency, student understanding, practice behavior, or learning outcomes.
Research should also look directly at switching and information competition. A multiview layout might improve access to spatially separated targets while simultaneously raising visual search demands. Eye tracking could test whether teachers actually use complementary views as intended, or just ignore them. Studies could compare simultaneous multiview displays against teacher-controlled switching, student-controlled switching, or automatic task-linked views — measuring both technical performance and pedagogical behavior, rather than assuming a technically available feed gets pedagogically used.
Finally, future studies should keep camera evidence distinct from other evidence channels. Instrumental diagnosis usually draws on sound, score, verbal report, prior knowledge of the student, and repeated performance all at once. Visual information is one part of a multimodal teaching ecology — PVAM is meant to sharpen that one part, not elevate it above the rest.
Conclusion
Synchronous instrumental teaching runs into a simple but consequential problem: a teacher can't respond to visual information that isn't there, and making information available is no guarantee it gets noticed, interpreted correctly, or turned into effective instruction. Research on distance lessons, gesture, embodiment, teacher observation, and visual attention all point the same direction — camera framing is an informational decision, not a technology preference.
This review's central recommendation follows from that: start with the pedagogical action — observing posture, inspecting hand movement, monitoring embouchure, establishing a score location, watching pedal use, communicating face to face, demonstrating a movement — and then choose the view that makes that specific target visible at usable resolution. Additional views earn their place only when they restore information the primary view necessarily leaves out.
The Pedagogical Visual Access Matrix offers a vocabulary for making those decisions explicit. It doesn't claim any particular camera angle improves learning, that more cameras are automatically better, or that visibility can substitute for professional judgment. Its narrower contribution is separating four steps that digital teaching often collapses into one: what can be seen, what gets noticed, what gets interpreted, and what gets done about it. That separation gives synchronous instrumental teaching a more deliberate practical foundation — and gives future research something testable to work with.
Ethics Statement
This literature review did not involve human participants, animals, or new primary data; ethics approval was not required.
Data Availability Statement
No new datasets were generated or analyzed for this review. The evidence base consists of the publicly cited scholarly and technical sources listed in the References.
Conflicts of Interest
The author declares no conflicts of interest.
References
- Acquilino, A.; Scavone, G. Current state and future directions of technologies for music instrument pedagogy. Frontiers in Psychology 2022, 13, 835609. [Google Scholar] [CrossRef]
- Bremmer, M.; Nijs, L. The role of the body in instrumental and vocal music pedagogy: A dynamical systems theory perspective on the music teacher's bodily engagement in teaching and learning. Frontiers in Education 2020, 5, 79. [Google Scholar] [CrossRef]
- Buonviri, N. O.; Paney, A. S. Effects of camera placement on undergraduates' peer teaching reflection. Journal of Music Teacher Education 2022, 31(3), 37–48. [Google Scholar] [CrossRef]
- Casarotti, J. P. Effective visual communication in piano lessons. American Music Teacher 2021, 71(1), 14–17. Available online: https://www.jstor.org/stable/e27143454.
- Comeau, G.; Lu, Y.; Swirp, M. On-site and distance piano teaching: An analysis of verbal and physical behaviours in a teacher, student and parent. Journal of Music, Technology & Education 2019, 12(1), 49–77. [Google Scholar] [CrossRef]
- Dammers, R. J. Utilizing internet-based videoconferencing for instrumental music lessons. Update: Applications of Research in Music Education 2009, 28(1), 17–24. [Google Scholar] [CrossRef]
- Daugvilaite, D. Exploring perceptions and experiences of students, parents and teachers on their online instrumental lessons. Music Education Research 2021, 23(2), 179–193. [Google Scholar] [CrossRef]
- de Bruin, L. R. Instrumental music educators in a COVID landscape: A reassertion of relationality and connection in teaching practice. Frontiers in Psychology 2021, 11, 624717. [Google Scholar] [CrossRef]
- Duke, R. A.; Prickett, C. A. The effect of differentially focused observation on evaluation of instruction. Journal of Research in Music Education 1987, 35(1), 27–37. [Google Scholar] [CrossRef]
- Hamond, L. F. Práticas pedagógicas no ensino superior de piano online: OBS Studio, VMPK, Reaper e Synthesia [Conference paper]. XXV Congresso Nacional da Associação Brasileira de Educação Musical; 2021. Available online: https://abemeducacaomusical.com.br/anais_congresso/v4/papers/704/public/704-4299-1-PB.pdf.
- Hicken, L. K.; Duke, R. A. Differences in attention allocation in relation to music teacher experience and expertise. Journal of Research in Music Education 2023, 70(4), 369–384. [Google Scholar] [CrossRef]
- King, A.; Prior, H.; Waddington-Jones, C. Exploring teachers' and pupils' behaviour in online and face-to-face instrumental lessons. Music Education Research 2019, 21(2), 197–209. [Google Scholar] [CrossRef]
- Kurkul, W. W. Nonverbal communication in one-to-one music performance instruction. Psychology of Music 2007, 35(2), 327–362. [Google Scholar] [CrossRef]
- Løkke Jakobsen, M.; Hebert, D. G.; Ørngreen, R. Synchronous online instrumental music teaching in cross-cultural learning contexts. International Journal of Music Education 2025, 43(2), 288–309. [Google Scholar] [CrossRef]
- Madsen, K.; Cassidy, J. W. The effect of focus of attention and teaching experience on perceptions of teaching effectiveness and student learning. Journal of Research in Music Education 2005, 53(3), 222–233. [Google Scholar] [CrossRef]
- Michałko, A.; Campo, A.; Nijs, L.; Leman, M.; Van Dyck, E. Toward a meaningful technology for instrumental music education: Teachers' voice. Frontiers in Education 2022, 7, 1027042. [Google Scholar] [CrossRef]
- National Association for Music Education. 2022 strategic plan. 2022. Available online: https://nafme.org/wp-content/uploads/2023/03/NAfME-2022-Strategic-Plan.pdf.
- Orman, E. K.; Whitaker, J. A. Time usage during face-to-face and synchronous distance music lessons. American Journal of Distance Education 2010, 24(2), 92–103. [Google Scholar] [CrossRef]
- Riley, H.; MacLeod, R. B.; Libera, M. Low latency audio video: Potentials for collaborative music making through distance learning. Update: Applications of Research in Music Education 2016, 34(3), 15–23. [Google Scholar] [CrossRef]
- Simones, L. L. A framework for studying teachers' hand gestures in instrumental and vocal music contexts. Musicae Scientiae 2019, 23(2), 231–249. [Google Scholar] [CrossRef]
- Simones, L. L.; Rodger, M.; Schroeder, F. Communicating musical knowledge through gesture: Piano teachers' gestural behaviours across different levels of student proficiency. Psychology of Music 2015, 43(5), 723–735. [Google Scholar] [CrossRef]
- Simones, L.; Schroeder, F.; Rodger, M. Categorizations of physical gesture in piano teaching: A preliminary enquiry. Psychology of Music 2015, 43(1), 103–121. [Google Scholar] [CrossRef]
- Stephens-Himonides, C.; Young, M. Adding to the knowledge of the TPACK framework: A case study of female identity in performance, education, and technology. Frontiers in Education 2025, 10, 1522739. [Google Scholar] [CrossRef]
- Turan, H. S. A thematic review of online piano education research (2019–2024). Music Education Research Advance online publication. 2026. [Google Scholar] [CrossRef]
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.