Submitted:
29 August 2026
Posted:
31 August 2026
You are already at the latest version
Abstract
Accurate control of stimuli timing is essential in psychophysical experiments, yet the stimulus onset asynchrony (SOA) specified within a virtual reality (VR) application may differ from that physically delivered to the participant. This study presents a wearable wireless sensor framework for measuring physical audiovisual SOAs on a trial-by-trial basis during VR experiments. Visual onset is detected using a photodiode-based synchronization marker, whereas auditory onset is detected directly from the analog audio signal using a portable wireless sensing module. The framework is characterized across different rendering pipelines, audio buffer configurations, spatial audio conditions, and wired and wireless headset connections, and subsequently validated during psychophysical experiments under both stationary and walking conditions. Before compensation, systematic audiovisual offsets ranged from 118 to 185 ms across configurations. Condition-specific temporal compensation reduced mean SOA errors to approximately zero, although residual trial-to-trial variability remained, with standard deviations of 11–32 ms. During psychophysical validation, the framework successfully measured physical audiovisual onset in 99.94\% of trials. These results demonstrate that trial-by-trial physical monitoring can complement conventional pre-experimental calibration, enabling behavioral responses to be associated with the audiovisual timing actually delivered on each trial and supporting more accurate and reproducible audiovisual psychophysics in dynamic VR.
Keywords:
wearable sensors
; virtual reality
; audiovisual timing
; timing validation
; temporal accuracy
1. Introduction
Immersive virtual reality (VR) has become an increasingly employed experimental platform for investigating human perception and behavior across a wide range of disciplines, including audio and visual perception, multisensory integration, human–computer interaction, rehabilitation, locomotion research, and clinical engineering [1]. Compared with traditional monitor-based paradigms, VR combines precise experimental control with greater ecological validity by allowing participants to naturally interact with virtual environments through head movements, body motion, and locomotion [2]. This combination of controlled sensory stimulation and interaction has made VR particularly attractive for investigating perceptual and cognitive processes under conditions that more closely resemble real-world behavior [3].
Among the many applications of immersive VR, a growing body of research is focusing on the temporal aspects of multisensory stimulation. Many experimental paradigms rely on the precise manipulation of the temporal relationship between visual and auditory events to investigate how the nervous system integrates information from different sensory modalities. Simultaneity Judgment task (SJ), Temporal Order Judgment task (TOJ) [4], the Sound-Induced Flash Illusion (SIFI) [5], the Bounce Perception Illusion [6], Temporal Bisection task [7], and other audiovisual temporal perception paradigms all use temporal delays between sensory stimuli as their primary independent variable. Traditionally, these audiovisual psychophysical experiments have been conducted using conventional display-based setups controlled by software environments specifically developed for vision science and psychophysics, such as Psychtoolbox [8] and PsychoPy [9]. These platforms have been extensively characterized with respect to stimulus timing and provide researchers with a high degree of confidence that the temporal relationship specified within the experimental software closely matches the physical timing of stimulus [10]. As VR becomes increasingly adopted for psychophysical research [11], however, this assumption can no longer be taken for granted. While modern game engines such as Unity and Unreal Engine allow experimenters to specify stimulus timing with millisecond precision at the software level, this alone does not guarantee that visual and auditory stimuli are presented with the requested temporal offset in the physical world [12,13,14].
The distinction between nominal and physical delays is particularly important in VR systems, where visual and auditory signals are generated through independent hardware and software pipelines. The nominal delay corresponds to the temporal delay specified within the experimental software, whereas the physical delay is the actual delay separating the onset of the sensory events perceived by the participant. For visual stimuli, the final presentation time may be influenced by multiple stages of the rendering and display pipeline (rendering latency, frame scheduling, display refresh rate, pixel response characteristics, reprojection techniques), as well as by headset-specific processing. Similarly, the onset of auditory stimuli may depend on the audio pipeline such as audio buffer size, operating system scheduling, audio engine implementation, spatial audio rendering, digital-to-analog conversion, wired or wireless transmission, and the characteristics of the playback device. As a consequence, the physical delay may systematically differ from the programmed value and may also vary across repeated trials because of temporal jitter introduced by the underlying hardware and software architecture.
Such discrepancies are not merely technical considerations but have direct implications for audiovisual psychophysics. Simultaneity judgment and temporal order judgment paradigms are widely used to characterize temporal aspects of multisensory integration, often through the estimation of psychometric functions from which measures such as the point of subjective simultaneity (PSS) and the temporal binding window (TBW) are derived [15]. These measures are commonly interpreted as reflecting the temporal characteristics of multisensory integration and are frequently compared across experimental conditions [4], participant populations [16,17], and laboratories [18]. However, if the physical stimulus onset asynchrony (SOA) differs from the nominal value used for data analysis, the resulting psychometric functions may be systematically shifted or distorted, potentially biasing estimates of perceptual parameters. Accurate validation of the physical SOA is therefore essential not only for obtaining reliable psychophysical measurements but also for improving the comparability and reproducibility of audiovisual experiments conducted across different VR platforms and hardware configurations [19].
To this end, several approaches have been proposed to characterize temporal performance in VR systems. Initial work focused on the measurement of motion-to-photon latency, defined as the delay between a user’s movement and the corresponding visual update displayed by the head-mounted display (HMD) [20]. A variety of measurement systems based on photodiodes, inertial sensors, microcontrollers, and oscilloscopes have been proposed to quantify this latency and benchmark the performance of VR rendering and display pipelines [21,22,23]. Although these studies have played a fundamental role in improving the temporal performance of immersive systems, they primarily address the visual channel and do not directly investigate the temporal relationship between auditory and visual stimuli, which represents the critical variable in audiovisual psychophysical experiments.
More recently, greater attention has been devoted to the empirical validation of audiovisual timing. The first systematic investigations into audiovisual timing in VR focused on quantifying the temporal accuracy of stimulus presentation. Tachibana et al. [14] developed a systematic methodology to evaluate the temporal accuracy of visual, auditory, and audiovisual stimulus presentation in VR using HTC Vive and Oculus Rift head-mounted displays. Their experimental setup combined a photodiode to detect visual stimulus onset with a microphone to detect sound onset while presenting sequences of 1000 visual, auditory, or audiovisual stimuli of different durations generated through Python. Their measurements revealed a visual latency of approximately 18 ms, whereas auditory latencies ranged from 37 ms on the HTC Vive Pro to 58 ms on the Oculus Rift. Consequently, audiovisual stimuli programmed to be simultaneous exhibited physical SOAs of approximately 20 ms and 40 ms, respectively, highlighting the importance of experimental validation before conducting audiovisual psychophysical experiments. This issue became particularly relevant with the widespread adoption of Unity and Unreal Engine for VR experiments, making further research necessary to characterise the temporal performance of these platforms. Therefore, Fucci et al. [13] proposed a low-cost hardware platform for measuring audio and visual latencies in VR based on a microcontroller, photodiode, microphone, and oscilloscope. Their measurements revealed substantial differences in both visual and auditory end-to-end latency across VR platforms, rendering engines, and V-sync configurations, with visual latency ranging from approximately 29 to 68 ms and audio latency from approximately 145 to 190 ms. Although they demonstrated that both visual and auditory end-to-end latencies vary substantially across VR platforms and software configurations, the visual and auditory channels were characterized independently, without directly quantifying the resulting physical audiovisual SOA. A further step in this direction was taken by Eckhoff et al. [12] who performed a comprehensive evaluation of audiovisual timing accuracy and precision across multiple HMDs, rendering engines, and software configurations. To directly measure the SOA, the authors developed a dummy head equipped with integrated microphones positioned at the ears and photodiodes placed over the eyes, enabling simultaneous acquisition of auditory and visual stimulus onsets. Their measurements demonstrated that both the accuracy and precision of the physical SOA strongly depended on the combination of HMD, rendering engine, and audio output, with systematic differences observed even within the same headset when different software configurations were used. Performance ranged from approximately 22 ms for the Varjo XR-3 using Unreal Engine and external headphones to over 150 ms for the Microsoft HoloLens 2 using Unreal Engine and the integrated speaker. Furthermore, they showed that experimentally calibrating the stimulus presentation within the Unity application substantially reduced the audiovisual timing bias while also decreasing temporal jitter, highlighting that accurate SOA control requires empirical calibration rather than relying solely on nominal software timing.
Collectively, these studies have substantially advanced the characterization of temporal performance in VR systems and have established the importance of physically validating audiovisual timing. However, they have been designed as calibration systems intended to characterize the temporal performance of a VR setup under static laboratory conditions. These approaches are valuable for hardware benchmarking but generally provide limited support for verifying the physical SOA of each trial during the actual execution of an experiment. Indeed, these systems rely on wired instrumentation and microphone-based acoustic measurements that are impractical during tasks involving locomotion or active interaction with the virtual environment. As VR experiments increasingly incorporate natural behaviors such as walking, reaching, and embodied interaction, there is a growing need for sensing solutions that can operate directly during experimental sessions without constraining participant movement.
To address this need, this study presents a wearable wireless sensor framework for the trial-by-trial measurement of the physical audiovisual SOA during VR experiments. Rather than serving solely as a calibration tool for characterizing the temporal performance of a VR system before data collection, the proposed framework is designed to operate throughout the experiment, enabling continuous verification of stimulus timing under the same conditions experienced by the participant. The system combines dedicated sensing of visual and auditory stimulus onsets with embedded wireless acquisition and automated signal processing to estimate the physical SOA of individual trials. This enables both the characterization of timing performance, through the evaluation of systematic bias, temporal jitter, repeatability, and detection reliability, and the use of physically measured SOAs in the subsequent analysis of psychophysical data. Finally, the framework is validated under both static and dynamic VR conditions, demonstrating its applicability to experimental paradigms involving participant movement and providing a practical tool for improving the accuracy and reproducibility of audiovisual psychophysical research in virtual reality.
1.1. System Architecture
The proposed measurement framework consists of two main components: a visual sensing module and an audio sensing module, which communicate with the same acquisition and synchronization unit responsible for measuring the timing of the events recorded by the two modules. As illustrated in Figure 1, the framework operates alongside the VR application, which generates the audiovisual stimuli, exchanges synchronization events with the acquisition unit, and records the acquired timestamps. The visual sensing module detects the physical onset of a dedicated synchronization marker displayed on the computer monitor simultaneously with the visual stimulus presented in the VR environment. In parallel, the audio sensing module identifies the physical onset of auditory stimuli directly from the analog audio output. The acquisition unit records the events generated by the visual and audio sensing modules, determines their physical timestamps, and simultaneously acquires the synchronization events sent by the VR application when visual and auditory stimuli are generated. The timestamps associated with both the VR triggered and physical events are then transmitted to the VR application, where they are automatically linked to the corresponding experimental trial and stored for subsequent offline analysis.
1.2. Acquisition and Synchronization Unit
The acquisition and synchronization unit consists of an Arduino Nano ESP32 which serves as the central acquisition and synchronization unit of the proposed framework. All nominal events generated by the VR application and all physical events detected by the sensing modules are timestamped using the Arduino internal clock, providing a common temporal reference. To improve measurement robustness and to ensure a one-to-one correspondence between each programmed stimulus and its physical measurement, the acquisition process is event-driven. Rather than continuously monitoring the sensor outputs, the visual and audio sensing modules are enabled only after the VR application triggers the generation of the corresponding stimulus. Whenever the VR application generates a visual or auditory stimulus, it immediately transmitts a synchronization event to the Arduino Nano ESP32 through the serial port. Receipt of this event enabled the corresponding sensing module, which remains active only until the physical onset of the stimulus is detected. The resulting timestamp is then recorded and the sensing module returns to its idle state while awaiting the next synchronization event.
For every audiovisual stimulus presentation, the acquisition unit therefore records four events: the trigger visual event generated by the VR application, the corresponding physical visual onset detected by the visual sensing module, the trigger auditory event generated by the VR application, and the corresponding physical auditory onset detected by the audio sensing module. The recorded timestamps are subsequently used to estimate the physical audiovisual SOA, compare the physical timing with the nominal SOA for quantifying systematic timing bias, temporal jitter, and repeatability. This synchronization strategy ensures that each measured onset was uniquely associated with its corresponding experimental event while minimizing false detections and reducing unnecessary sensor activity between successive stimulus presentations.
1.3. Visual Sensing Module
The visual sensing module is designed to detect the physical onset of each visual stimulus by monitoring a dedicated synchronization marker displayed simultaneously with the visual stimulus generated by the VR application, allowing to measure the visual onset independently of the position and appearance of the experimental scene.
The synchronization marker consists of a black square patch (2 × 2 cm) displayed at a fixed location on the computer monitor. Whenever a visual stimulus is presented, the marker changes from black to white and remains visible for 100 ms before returning to its initial state. Since the marker is positioned at the peripheral region of the head-mounted display’s field of view, it doesn’t interfere with the experimental task while providing a reliable reference for visual onset detection.
The synchronization marker is monitored using a PD333-3C/H0/L2 photodiode (Everlight Electronics Co., Ltd.), featuring an active area of 19.6 , a spectral sensitivity of 400–1100 nm, and a response time of 45 ns. The photodiode is mounted on a small cardboard support and is fixed directly over the synchronization marker using adhesive tape. The cardboard shields the sensing element reducing the influence of ambient illumination and improving measurement robustness.
Following the software synchronization trigger described in Section 1.2, the photodiode signal is acquired through one analog input of the Arduino Nano ESP32 using its onboard analog-to-digital converter (ADC, 0–3.3 V input range) and sampled using the analogRead() function. Visual onset is identified when the acquired signal exceeds a predefined detection threshold. The threshold is determined empirically by measuring the sensor output while the synchronization marker is displayed in its black and white states. The black marker consistently produced values between approximately 0 and 10 ADC counts, whereas the white marker generated values around 100–110 ADC counts. Based on these measurements, a detection threshold of 90 ADC counts is selected to reliably discriminate the onset of the synchronization marker while minimizing false detections due to measurement noise. Once the threshold crossing is detected, the Arduino Nano ESP32 immediately records the corresponding physical timestamp using its internal clock. The visual sensing module is then disabled until the next software synchronization trigger is received.
1.4. Wireless Audio Sensing Module
The wireless audio sensing module is designed to detect the physical onset of auditory stimuli directly from the analog audio signal delivered to the participant’s headphones. Unlike the visual sensing module, which relies on an external synchronization marker, the proposed audio module monitors the same analog signal presented to the participant, ensuring that the detected onset corresponds to the actual auditory stimulus delivered during the experiment.
The module consists of an analog signal conditioning circuit connected to a XIAO ESP32-S3 Plus microcontroller, powered by a rechargeable lithium battery and enclosed in a compact portable housing. The conditioning circuit receives the left and right audio channels independently from the auxiliary (AUX) audio output through a dedicated input connector. A second AUX connector routes the same audio signal to the participant’s headphones, allowing the module to monitor the audio stream without modifying the signal delivered to the listener. This configuration can be connected either to the computer AUX output during static laboratory experiments or directly to the headset AUX output during dynamic VR experiments, enabling the same sensing module to be used in both experimental conditions.
The analog conditioning circuit performs AC coupling, signal buffering, and level shifting to adapt the audio signal to the input range of the XIAO ESP32-S3 Plus analog-to-digital converter. The left and right audio channels are processed independently before being sampled by the microcontroller using the analogRead() function.
Auditory onset detection is based on threshold crossing. Whenever either audio channels exceeds the threshold, the XIAO ESP32-S3 Plus immediately transmits an Audio Detected message to the Arduino Nano ESP32 using the ESP-NOW wireless communication protocol. The average transmission delay between the two microcontrollers is experimentally characterized through repeated bidirectional ("ping-pong") measurements (see Section 1.7.1).
Following the synchronisation signal generated by the VR application (see Section 1.2), the Arduino Nano ESP32 initiates the capture of audio events and waits for the message from the XIAO ESP32-S3 Plus to record the corresponding physical audio timestamp using its internal clock.
1.5. VR Setup
Two experimental setups are implemented to validate the proposed sensing framework under both stationary and locomotion-based VR conditions. Both applications are developed in Unity using the OpenXR framework and executed on Windows 11. The first application is implemented in Unity 2021.3.21f1 using the Universal Render Pipeline (URP), whereas the second application is implemented in Unity 2022.3.62f3 using the High Definition Render Pipeline (HDRP). In both applications, vertical synchronization (VSync) is disabled and the Unity fixed timestep is left at its default value.
Both experimental setups employ the Meta Quest 3 head-mounted display, operating at a refresh rate of 72 Hz. The HMD is connected to the host computer either via cable or wirelessly using Meta Horizon Link. External headphones are used to deliver the auditory stimuli and to provide the analog audio signal monitored by the wireless sensing module.
Audiovisual stimuli are generated within Unity using independent rendering and audio pipelines. Visual events are presented by triggering a Visual Effect Graph (VFX Graph), whereas auditory stimuli are reproduced using the AudioSource.PlayScheduled() function to achieve sample-accurate audio playback. The desired nominal SOA is implemented by delaying the onset of one sensory modality relative to the other through coroutine-based timing using the WaitForSeconds() function.
The initial characterization of the framework employs predetermined nominal SOAs, whereas the psychophysical experiments generate SOAs online using an adaptive procedure. Additional hardware specifications for the two experimental setups are summarized in Table 1.
1.6. Stimuli
1.6.1. Visual Stimuli
To demonstrate the applicability of the proposed framework across different experimental paradigms, two visual stimuli are employed. In the first validation experiment, the visual stimulus consists of a white flash of 45 ms while in the second experiment, the visual event corresponds to the explosion of a virtual balloon with a duration of 45 ms. Since visual onset is determined exclusively by the dedicated synchronization marker described in Section 1.3, the framework remains independent of the position, appearance, and temporal evolution of the experimental stimulus.
1.6.2. Auditory Stimuli
Consistent with the visual modality, the first and second experiments employ different auditory stimuli. In the first experiment, the auditory stimulus consists of a 45 ms broadband white-noise burst. In the second experiment, the auditory stimulus corresponds to the sound of a balloon pop, also lasting 45 ms. Both stimuli are spatialized in real time using the 3DTune-In Toolkit with a generic KEMAR head-related transfer function (HRTF) and include the reverberation characteristics of the room.
1.6.3. Stimulus Onset Asynchrony
For each audiovisual presentation, the VR application generates two software synchronization triggers corresponding to the visual and auditory stimuli. These triggers are transmitted to the Arduino Nano ESP32 via the serial interface and timestamped using the Arduino internal clock, yielding the software trigger timestamps and .
The physical onset of the visual and auditory stimuli is independently detected by the visual and audio sensing modules, producing the corresponding physical timestamps and , respectively. Since the synchronization marker is displayed on the external monitor while the participant observes the visual stimulus through the HMD, the measured visual timestamp is corrected by a constant monitor-to-HMD latency, experimentally determined prior to the experiments.
The nominal stimulus onset asynchrony, which is the programmed SOA, is defined as
where positive values indicate that the auditory stimulus follows the visual stimulus and negative values indicate that it precedes it.
The physical stimulus onset asynchrony is computed as
Finally, the SOA error for each trial is defined as
which represents the discrepancy between the programmed audiovisual timing and the physically measured timing.
For the initial characterization of the sensing framework, nominal SOAs range from -500 ms to +500 ms in 100 ms increments, with 10 repetitions acquired for each SOA.
1.7. Framework Validation Protocol
The proposed sensing framework is validated through a three-stage protocol consisting of system calibration, timing characterization, and experimental validation. This protocol is designed to first identify systematic timing biases introduced by the VR system, then to correct these systematic temporal deviations, and finally demonstrate its applicability during real psychophysical experiments.
1.7.1. System Calibration
Two preliminary calibration procedures are performed before evaluating the framework.
The first calibration quantifies the latency between the synchronization marker displayed on the external monitor and the corresponding visual stimulus presented inside the head-mounted display. Two identical photodiodes are employed simultaneously: one is positioned over the synchronization marker on the monitor, while the second is mounted inside the HMD and detects a full-screen flash presented concurrently with the marker. Fifty repeated measurements are acquired to characterize the latency between the synchronization marker displayed on the external monitor and the corresponding visual stimulus presented inside the HMD. The resulting monitor-to-HMD delay was subsequently used to correct the visual timestamps during timing analysis.
The second calibration evaluates the communication latency of the wireless ESP-NOW protocol between the XIAO ESP32-S3 Plus and the Arduino Nano ESP32. A bidirectional “ping-pong” communication protocol is employed in which 1000 round-trip exchanges are performed. The one-way transmission delay is estimated as half of the measured round-trip time, allowing the average communication latency and its variability to be characterized.
1.7.2. Timing Characterization
Following system calibration, the timing performance of the proposed framework is characterized under the two experimental configurations described in Section 1.5. The first setup is evaluated under stationary conditions with the HMD connected to the host computer through a wired connection. The second setup is evaluated under both stationary and locomotion-based conditions with the HMD connected wirelessly.
For each experimental configuration, timing characterization is performed under different audio rendering conditions by enabling or disabling real-time spatial audio processing through the 3DTune-In Toolkit. In addition, the Unity audio buffer configuration is systematically varied using the available latency settings (Best Latency, Good Latency, and Best Performance) to evaluate its effect on audiovisual synchronization.
For each tested condition, the systematic temporal offset of the VR system is estimated from the measured SOA errors. Specifically, the mean timing error is first computed separately for positive and negative values of the eSOA:
The systematic temporal offset is then estimated as
The estimated offset is subsequently compensated directly within the Unity application. Depending on the experimental configuration, the compensation is implemented either by delaying the visual stimulus or by advancing the auditory stimulus to minimize the systematic timing error. Following implementation of the compensation, the complete timing characterization is repeated under the same experimental conditions. The mean and standard deviation of the resulting values are then computed for each configuration to evaluate the effectiveness of the proposed correction strategy.
1.7.3. Psychophysical Validation
The proposed sensing framework is finally validated during real psychophysical experiments to demonstrate its applicability in audiovisual perception studies without modifying the experimental protocol. Participants perform an audiovisual Simultaneity Judgment Task (SJT), in which 120 audiovisual stimulus pairs are presented using an adaptive SOA selection procedure. After each stimulus pair, participants are asked to report whether the auditory and visual stimuli were perceived as synchronous or asynchronous.
Based on the timing characterization described in the previous section, the systematic temporal offset estimated for each experimental setup is compensated before data collection using the corresponding offset. All psychophysical experiments are performed using spatial audio rendering through the 3DTune-In Toolkit and the Unity audio buffer configuration set to Good Latency. Although the Best Latency setting provides the highest temporal performance during preliminary characterization, the Good Latency setting is selected for the psychophysical experiments because repeated preliminary tests revealed occasional audio playback instability with the Best Latency configuration, including audible distortion and discontinuous playback. These instabilities are particularly observed for simultaneous audiovisual presentations (SOA = 0 ms). Good Latency is therefore selected as a compromise between temporal performance and reliable audio reproduction during experimental sessions.
The first experimental setup described in Section 1.5 is evaluated under stationary conditions with the Meta Quest 3 connected to the host computer through a wired connection. The second experimental setup is evaluated using the Meta Quest 3 connected wirelessly and includes both stationary and locomotion-based conditions. Five participants are tested for each experimental condition.
Throughout all experiments, the proposed sensing framework continuously records the physical onset of every audiovisual stimulus pair, enabling trial-by-trial verification of the effective SOA delivered during psychophysical testing.
2. Results
2.1. System Calibration
The preliminary calibration procedures show that both the monitor-to-HMD latency and the wireless communication delay are small and reproducible. The synchronization marker displayed on the external monitor follows the corresponding visual onset in the HMD by 5 ms [5 ± 0.1 ms], and this delay is subsequently taken into account when calculating the error between the nominal and the physical SOA.
The ESP-NOW communication latency is characterized across two independent runs. The estimated one-way latency is 1.502 ± 0.563 ms in the first run and 1.512 ± 0.556 ms in the second, with minimum latencies of 1.154 and 1.167 ms and maximum latencies of 7.612 and 6.067 ms, respectively. These measurements indicate a consistent average transmission delay of approximately 1.5 ms across repeated runs, although occasional longer transmission times are observed. Given its small magnitude relative to the temporal variability of the audiovisual presentation system, the average wireless transmission delay is considered negligible and is therefore not compensated in the subsequent SOA measurements.
2.2. Timing Characterization
Before temporal compensation, a systematic discrepancy between nominal and physically measured SOAs is observed across all tested configurations (Figure 2; Table 2). The magnitude of this discrepancy varies according to the rendering pipeline and Unity audio buffer setting. Across all tested configurations, the Best Latency setting consistently produces the smallest systematic offsets, whereas Best Performance produces the largest. Similarly, offsets are consistently smaller with the URP than with the HDRP rendering pipeline. In contrast, no clear systematic differences are observed between spatialized and non-spatialized audio conditions or between wired and wireless HMD connections. Overall, the estimated systematic offset ranges from 118 ms in the configuration combining URP, spatialized audio, wireless HMD connection, and the Best Latency setting, to 185 ms with HDRP, spatialized audio, wired connection, and the Best Performance setting.
Following condition-specific compensation, the systematic component of the SOA error is substantially reduced, with mean residual errors approaching zero across all configurations (Figure 2; Table 2). Trial-to-trial variability nevertheless remains and differs across system configurations, with standard deviations ranging from 11 to 32 ms. The Best Latency setting generally yields the lowest residual variability, with the exception of the URP configuration with spatialized audio and wireless connection, whereas Best Performance consistently produces the highest variability. Residual variability is also generally lower with HDRP than with URP, except for the spatialized-audio condition with a wired HMD connection.
Considering the 72 Hz HMD refresh rate, a single display frame corresponds to approximately 13.9 ms. Residual variability is therefore generally on the order of one to two display-frame intervals, with some configurations showing standard deviations below a single frame interval and the largest observed variability corresponding to approximately 2.3 frame intervals. Thus, while condition-specific compensation effectively removed the systematic timing bias, trial-to-trial variability in the physical SOA persists across individual stimulus presentations.
2.3. Psychophysical Validation
During the psychophysical experiments, the framework successfully measures the physical audiovisual SOA on a trial-by-trial basis under all three experimental conditions, including during participant locomotion (Figure 3). Mean SOA errors remain close to zero across conditions, averaging ms in the static cabled condition, ms in the static wireless condition, and ms during walking with a wireless HMD connection.
Residual trial-to-trial variability differs across experimental configurations. The static cabled condition shows the lowest variability, with a mean within-participant standard deviation of 9.20 ms (range: 8.39-10.38 ms). Variability is higher in the static wireless condition, averaging 17.47 ms (range: 14.22-27.28 ms), whereas the walking wireless condition shows an intermediate variability of 13.50 ms (range: 11.87-15.79 ms). Thus, the use of a wireless HMD connection is associated with greater SOA variability than the cabled configuration, while locomotion did not produce a further increase in temporal variability relative to the static wireless condition.
Across the 1,800 experimental trials, the framework successfully detected both physical stimulus onsets in 1,799 trials, corresponding to an overall detection rate of 99.94%. The only missing measurement was an auditory onset during the walking condition. Because the corresponding software trigger was correctly recorded, the missing physical onset can be reconstructed using the mean trigger-to-physical auditory latency derived from the remaining valid trials, allowing the trial to be retained in the dataset.
3. Discussion
The present study confirms that the temporal relationship specified between auditory and visual stimuli within a VR application does not necessarily correspond to the audiovisual SOA physically delivered to the observer. Before temporal compensation, substantial systematic discrepancies are observed across all tested configurations, with the magnitude of the offset varying as a function of the rendering pipeline and audio buffer configuration. In particular, the Best Latency setting consistently yields smaller offsets than Good Latency and Best Performance, while URP generally produces smaller offsets than HDRP rendering pipeline. Conversely, no equally consistent pattern emerges as a function of audio spatialization or wired versus wireless HMD connection. This configuration-dependent variability is consistent with previous VR timing characterizations, which have shown that audiovisual latency can vary substantially across HMDs, rendering environments, and audio configurations [12,13,14]. Taken together, these findings reinforce the idea that audiovisual timing should be considered a property of the complete experimental setup rather than depending solely on any individual hardware or software component. Importantly, the systematic audiovisual offset observed across configurations can be effectively compensated by adjusting the relative timing of the auditory and visual stimuli within the experimental software, either by delaying one modality or advancing the other according to the direction and magnitude of the measured offset. Following condition-specific compensation, the mean SOA error approaches zero across all tested configurations, consistent with previous evidence showing that hardware-based characterization can be used to identify and compensate for systematic audiovisual timing biases in VR [12,24]. However, correcting the systematic offset does not eliminate trial-to-trial temporal variability. In the present study, residual variability ranges from approximately 11 to 32 ms during system characterization and remains evident during the psychophysical experiments, despite mean SOA errors close to zero. This distinction reflects two complementary aspects of temporal performance: accuracy, which can be improved by identifying and compensating for systematic timing offsets, and precision, which depends on the variability of stimulus timing across repeated presentations. The residual variability observed after compensation therefore indicates that an accurately calibrated system does not necessarily reproduce the intended SOA identically on every trial. The persistence of residual variability after temporal compensation highlights the potential advantage of extending timing validation beyond pre-experimental calibration. While previous approaches have demonstrated that physical audiovisual timing can be characterized and corrected before data collection [24], the present framework enables the physical SOA to be measured continuously on a trial-by-trial basis during the experiment itself. This provides an additional level of temporal information, as each behavioral response can be associated with the audiovisual timing that was physically delivered on that specific trial and can be used directly in psychophysical analyses, when fitting psychometric functions and estimating temporal-perception measures.
Importantly, the wearable and wireless implementation extends this approach to dynamic VR paradigms in which participants are free to move. The framework maintains reliable physical onset detection during locomotion, with an overall detection reliability of 99.94%. Moreover, the parallel recording of software triggers provides temporal redundancy: when a physical onset is occasionally missed, the corresponding trigger remains available and can be used, together with the trigger-to-physical latency, to estimate the missing timestamp, allowing the trial to be retained for the analysis. Together, these results demonstrate the feasibility of trial-by-trial audiovisual timing validation not only in stationary VR experiments, but also in more dynamic paradigms involving natural body movement and locomotion.
3.1. Practical Considerations for Audiovisual Timing in VR
Based on the present findings and previous timing characterizations, several practical considerations can be identified for researchers aiming to achieve accurate and reliable audiovisual timing in VR experiments. These considerations concern both the initial validation of the experimental setup and the strategies used to minimize and monitor timing errors during data collection.
First, audiovisual timing should be empirically validated for each specific experimental setup rather than inferred from values reported for other systems. The timing values observed in the present study should therefore not be interpreted as reference values that can be directly transferred to other experimental configurations but each experimental setup should be independently verified.
Second, physical stimulus events should, whenever possible, be measured and timestamped using a dedicated acquisition device operating independently of the timing of the host computer. Conventional operating systems and experimental software do not operate under strict real-time constraints, and their workload and scheduling may therefore introduce unpredictable delays between software commands and physical stimulus presentation. In the present framework, physical events are referenced to the internal clock of the Arduino Nano ESP32 before the corresponding timestamps are transferred to the host computer. Consequently, variability in the subsequent USB communication affects when the information becomes available to the experimental application but not the timestamp already assigned to the event. This principle is consistent with the hardware-based timing approach advocated by Eckhoff et al. [12]. When additional communication stages are required, their contribution should also be independently characterized. In the present wireless audio module, for example, ESP-NOW communication introduces an additional delay between audio detection and timestamping by the central microcontroller and this delay was experimentally characterized.
Third, systematic audiovisual offsets should be quantified and, when present, compensated before experimental data collection. Once the direction and magnitude of the offset have been established, compensation can be implemented at the software level by delaying one sensory modality or advancing the other accordingly. Importantly, physical timing should be measured again after compensation to verify that the correction has effectively removed the systematic component of the error.
Fourth, temporal variability should be minimized in addition to correcting systematic offsets. On the auditory side, the present characterization indicates that smaller audio buffers generally provide better temporal performance, with the Best Latency configuration showing the lowest residual variability in most conditions. However, minimizing buffer size introduces an important practical trade-off. During preliminary testing, the Best Latency setting occasionally resulted in unstable audio reproduction, including discontinuities and audible distortion, particularly for simultaneous audiovisual presentations. The smallest available buffer should therefore not necessarily be selected by default; rather, researchers should identify the smallest buffer that provides both satisfactory temporal precision and stable, artifact-free audio reproduction on their specific system [12]. On the visual side, stable rendering performance should similarly be prioritized. Maintaining a stable frame rate and avoiding dropped or reprojected frames is particularly important in timing-sensitive paradigms, as visual stimulus onset is ultimately constrained by the display refresh cycle and rendering pipeline [25].
Finally, when the experimental question is particularly sensitive to small variations in audiovisual timing, trial-by-trial physical monitoring can complement pre-experimental calibration. Calibration can effectively reduce systematic timing bias, but the residual variability observed in the present study demonstrates that the physical SOA still varies across individual stimulus presentations. The appropriate level of timing validation should therefore depend on the temporal sensitivity of the experimental question: for some paradigms, characterization and compensation of the average offset may be sufficient, whereas experiments requiring fine-grained estimates of audiovisual timing may benefit from continuous trial-by-trial measurement.
3.2. Limitations
Some limitations should be considered when interpreting the present findings. First, the timing characterization was performed using a specific combination of hardware and software, including a single HMD model, host computer, VR development environment, and spatial audio implementation. Although different rendering pipelines, audio buffer configurations, and HMD connection modalities were examined, the quantitative timing values reported here cannot be assumed to generalize to other experimental systems. This limitation further emphasizes the importance of characterizing each experimental setup independently.
Second, during experimental operation, visual onset within the HMD is not measured directly on every trial. Instead, the framework detects the physical onset of a synchronization marker displayed on the external monitor and estimates the corresponding HMD onset using a separately characterized monitor-to-HMD temporal difference. This approach enables unobtrusive visual monitoring during dynamic experiments but assumes that the temporal relationship between the external marker and the HMD display remains sufficiently stable throughout data collection.
Similarly, auditory onset is detected by the wearable module before being communicated via ESP-NOW to the central microcontroller, where the event is timestamped. The resulting timestamp therefore includes the wireless transmission latency, which was small and reproducible on average in the present implementation but represents an additional source of measurement uncertainty.
Finally, the psychophysical validation involved five participants per experimental condition. While this sample is sufficient to demonstrate the feasibility of the framework under stationary and locomotion-based experimental conditions, the limited amount of data restricts the strength of conclusions regarding its long-term measurement reliability. In particular, although a high physical-onset detection rate is observed in the present experiments, a larger number of experimental sessions and stimulus presentations would be required to provide a more robust characterization of detection failures and their frequency under different operating conditions.
3.3. Future Applications and Extensions
Beyond audiovisual SOA validation, the modular architecture of the proposed framework could be extended to characterize additional aspects of stimulus and response timing. First, although the present study focuses on the temporal relationship between auditory and visual events, the same sensing modules could be used independently to characterize unimodal timing, including auditory–auditory or visual–visual stimulus onsets. Such measurements would complement the audiovisual characterization performed here and extend the framework to experimental paradigms in which the temporal relationship between successive events within the same sensory modality is critical.
The framework could also be adapted to measure stimulus offset in addition to onset, allowing the physical duration of auditory and visual stimuli to be characterized. This would enable the validation of both the accuracy and precision of stimulus duration and the identification of the minimum durations that can be reliably reproduced by a given experimental system, extending previous VR timing characterizations such as that of Tachibana et al [14]. Such an extension would be particularly relevant for paradigms in which stimulus duration represents the experimental variable of interest, including temporal bisection, duration discrimination, and temporal reproduction tasks.
Finally, the same architecture could be extended to the measurement of behavioral response timing. Accurate reaction-time estimation requires not only precise stimulus-onset measurement but also reliable timestamping of the participant’s response, as timing errors introduced by VR systems can affect both components [25]. A physical response device, such as a push button connected directly to the central microcontroller, could be timestamped using the same internal clock employed for the sensory events. Reaction time could therefore be computed from stimulus and response timestamps expressed within a common temporal reference, while the response event could subsequently be transmitted to the VR application via the serial connection. This extension would allow the framework to support experimental paradigms requiring accurate measurement of both sensory timing and behavioral responses.
4. Conclusions
This study presents a wearable wireless framework for the trial-by-trial validation of physical audiovisual stimulus timing in immersive VR. The results demonstrate that substantial discrepancies can occur between nominal and physically delivered SOAs and that their magnitude depends on the specific hardware and software configuration. Although condition-specific calibration can effectively compensate for systematic temporal offsets, residual trial-to-trial variability remains and cannot be captured by average calibration alone. By continuously measuring the physical onset of auditory and visual stimuli during experimental sessions, the proposed framework enables each behavioral response to be associated with the SOA actually delivered on that trial. Importantly, the system maintained reliable onset detection under both stationary and locomotion-based conditions, demonstrating its applicability to dynamic VR paradigms in which conventional wired measurement approaches may be impractical. Overall, the framework provides a practical approach for complementing pre-experimental calibration with continuous timing verification, potentially improving the accuracy, transparency, and reproducibility of audiovisual psychophysical experiments conducted in immersive and interactive virtual environments.
Author Contributions
Conceptualization, E.A.B.V. and A.C.; methodology, E.A.B.V., A.C. and G.C.; software, E.A.B.V. and G.C.; validation, E.A.B.V. and G.C.; formal analysis, E.A.B.V. and A.C.; investigation, E.A.B.V.; resources, A.C. and G.C.; data curation, E.A.B.V.; writing—original draft preparation, E.A.B.V.; writing—review and editing, E.A.B.V., A.C. and G.C.; visualization, E.A.B.V.; supervision, A.C.; project administration, E.A.B.V. and A.C. All authors have read and agreed to the published version of the manuscript.
Funding
This research received no external funding
Institutional Review Board Statement
The study was conducted in accordance with the Declaration of Helsinki and approved by the Ethics Committee for Research of the University of Genoa (CERA; approval no. 2023.69, 21 September 2023).
Informed Consent Statement
Informed consent was obtained from all subjects involved in the study.
Data Availability Statement
The data supporting the findings of this study are available from the corresponding author upon reasonable request.
Acknowledgments
During the preparation of this manuscript, the authors used ChatGPT (OpenAI) for English language revision and for the generation of graphical elements used in Figure 1. The authors reviewed and edited the output and take full responsibility for the content of this publication.
Conflicts of Interest
The authors declare no conflicts of interest.
References
- Martin, D.; Malpica, S.; Gutierrez, D.; Masia, B.; Serrano, A. Multimodality in VR: A Survey. ACM Comput. Surv. (CSUR) 2022, 54, 216:1–216:36. [Google Scholar] [CrossRef]
- Parsons, T.D. Virtual Reality for Enhanced Ecological Validity and Experimental Control in the Clinical, Affective and Social Neurosciences. Front. Hum. Neurosci. 2015, 9, 660. [Google Scholar] [CrossRef]
- Loomis, J.M.; Blascovich, J.J.; Beall, A.C. Immersive virtual environment technology as a basic research tool in psychology. Behav. Res. Methods Instrum. Comput. 1999, 31, 557–564. [Google Scholar] [CrossRef]
- Stevenson, R.A.; Wallace, M.T. Multisensory temporal integration: task and stimulus dependencies. Exp. Brain Res. 2013, 227, 249–261. [Google Scholar] [CrossRef]
- Shams, L.; Ma, W.J.; Beierholm, U. Sound-induced flash illusion as an optimal percept. Neuroreport 2005, 16, 1923–1927. [Google Scholar] [CrossRef]
- Grassi, M.; Casco, C. Audiovisual bounce-inducing effect: When sound congruence affects grouping in vision. Atten. Percept. Psychophys. 2010, 72, 378–386. [Google Scholar] [CrossRef]
- Gori, M.; Sandini, G.; Burr, D. Development of Visuo-Auditory Integration in Space and Time. Front. Integr. Neurosci. 2012, 6. [Google Scholar] [CrossRef]
- Brainard, D.H. The Psychophysics Toolbox. Spat. Vis. 1997, 10, 433–436. [Google Scholar] [CrossRef]
- Peirce, J.W. PsychoPy—Psychophysics software in Python. J. Neurosci. Methods 2007, 162, 8–13. [Google Scholar] [CrossRef]
- Bridges, D.; Pitiot, A.; MacAskill, M.R.; Peirce, J.W. The timing mega-study: comparing a range of experiment generators, both lab-based and online. PeerJ 2020, 8, e9414. [Google Scholar] [CrossRef]
- Bohil, C.J.; Alicea, B.; Biocca, F.A. Virtual reality in neuroscience research and therapy. Nat. Rev. Neurosci. 2011, 12, 752–762. [Google Scholar] [CrossRef]
- Eckhoff, D.; Schnupp, J.; Cassinelli, A. Temporal precision and accuracy of audio-visual stimuli in mixed reality systems. PLoS ONE 2024, 19, e0295817. [Google Scholar] [CrossRef]
- Fucci, V.; Liu, J.; You, Y.; Cuijpers, R. Measuring Audio-Visual Latencies in Virtual Reality Systems; 2024; pp. 137–149. [Google Scholar] [CrossRef]
- Tachibana, R.; Matsumiya, K. Accuracy and precision of visual and auditory stimulus presentation in virtual reality in Python 2 and 3 environments for human behavior research. Behav. Res. Methods 2022, 54, 729–751. [Google Scholar] [CrossRef]
- Vatakis, A.; Balci, F.; Di Luca, M.; Correa, A. Timing and Time Perception: Procedures, Measures, & Applications; BRILL, 2018. [Google Scholar] [CrossRef]
- Stevenson, R.A.; Zemtsov, R.K.; Wallace, M.T. Individual differences in the multisensory temporal binding window predict susceptibility to audiovisual illusions. J. Exp. Psychol. Hum. Percept. Perform. 2012, 38, 1517–1529. [Google Scholar] [CrossRef]
- Stevenson, R.A.; Segers, M.; Ferber, S.; Barense, M.D.; Camarata, S.; Wallace, M.T. Keeping time in the brain: Autism spectrum disorder and audiovisual temporal processing. Autism Res. 2016, 9, 720–738. Available online: https://onlinelibrary.wiley.com/doi/pdf/10.1002/aur.1566. [CrossRef]
- Parise, C.V.; Parise, E.; Parise, A. Perceiving audiovisual synchrony: A quantitative synthesis of simultaneity and temporal order judgments from 185 studies. Neurosci. Biobehav. Rev. 2026, 181, 106449. [Google Scholar] [CrossRef]
- Plant, R.R. A reminder on millisecond timing accuracy and potential replication failure in computer-based psychology experiments: An open letter. Behav. Res. Methods 2016, 48, 408–411. [Google Scholar] [CrossRef]
- Raaen, K.; Kjellmo, I. Measuring Latency in Virtual Reality Systems. In Proceedings of the Entertainment Computing - ICEC 2015; Chorianopoulos, K., Divitini, M., Baalsrud Hauge, J., Jaccheri, L., Malaka, R., Eds.; Cham, 2015; pp. 457–462. [Google Scholar] [CrossRef]
- Warburton, M.; Mon-Williams, M.; Mushtaq, F.; Morehead, J.R. Measuring motion-to-photon latency for sensorimotor experiments with virtual reality systems. Behav. Res. Methods 2023, 55, 3658–3678. [Google Scholar] [CrossRef]
- Stauffert, J.P.; Niebling, F.; Latoschik, M.E. Simultaneous Run-Time Measurement of Motion-to-Photon Latency and Latency Jitter. In Proceedings of the 2020 IEEE Conference on Virtual Reality and 3D User Interfaces (VR), 2020; pp. 636–644, ISSN 2642-5254. [Google Scholar] [CrossRef]
- Choi, S.W.; Lee, S.; Seo, M.W.; Kang, S.J. Time Sequential Motion-to-Photon Latency Measurement System for Virtual Reality Head-Mounted Displays. Electronics 2018, 7, 171. [Google Scholar] [CrossRef]
- Clouston, G.; Davidson, M.; Alais, D. Perception of audio-visual synchrony is modulated by walking speed and step-cycle phase, 2024. Pages: 2024.07.21.604456 Section: New Results. [CrossRef]
- Wiesing, M.; Fink, G.R.; Weidner, R. Accuracy and precision of stimulus timing and reaction times with Unreal Engine and SteamVR. PLoS ONE 2020, 15, e0231152. [Google Scholar] [CrossRef]
Figure 1.
Overview of the proposed framework for trial-by-trial audiovisual timing validation in virtual reality. The VR experimental setup generates the auditory and visual stimuli together with their corresponding software triggers. Visual onset is monitored through a synchronization marker displayed on the external monitor and detected by a photodiode. Auditory onset is detected from the analog audio signal by the wireless audio sensing module, which includes a signal-conditioning circuit and a XIAO ESP32-S3 microcontroller; detected events are transmitted to the acquisition unit via ESP-NOW. The Arduino Nano ESP32 serves as the central acquisition and synchronization unit, recording software triggers ( and ) and physical auditory and visual onset timestamps ( and ) on a common clock and storing the resulting trial-by-trial timing data.
Figure 1.
Overview of the proposed framework for trial-by-trial audiovisual timing validation in virtual reality. The VR experimental setup generates the auditory and visual stimuli together with their corresponding software triggers. Visual onset is monitored through a synchronization marker displayed on the external monitor and detected by a photodiode. Auditory onset is detected from the analog audio signal by the wireless audio sensing module, which includes a signal-conditioning circuit and a XIAO ESP32-S3 microcontroller; detected events are transmitted to the acquisition unit via ESP-NOW. The Arduino Nano ESP32 serves as the central acquisition and synchronization unit, recording software triggers ( and ) and physical auditory and visual onset timestamps ( and ) on a common clock and storing the resulting trial-by-trial timing data.

Figure 2.
Distribution of audiovisual SOA errors () measured during timing characterization before (Biased) and after (Corrected) condition-specific temporal compensation. Results are shown for the Universal Render Pipeline (URP) and High Definition Render Pipeline (HDRP) under three audio and HMD configurations: spatialized audio with a cabled HMD connection, spatialized audio with a wireless HMD connection, and standard stereo audio with a cabled HMD connection. For each configuration, distributions are reported for the three Unity audio buffer settings: Best Latency (BL), Best Performance (BP), and Good Latency (GL). Individual points represent repeated measurements and the associated distributions illustrate their variability. Black markers indicate group medians. In the Biased condition, separate black markers indicate the medians of the distributions of SOA errors below and above zero; in the Corrected condition, a single black marker indicates the mean of the entire distribution.
Figure 2.
Distribution of audiovisual SOA errors () measured during timing characterization before (Biased) and after (Corrected) condition-specific temporal compensation. Results are shown for the Universal Render Pipeline (URP) and High Definition Render Pipeline (HDRP) under three audio and HMD configurations: spatialized audio with a cabled HMD connection, spatialized audio with a wireless HMD connection, and standard stereo audio with a cabled HMD connection. For each configuration, distributions are reported for the three Unity audio buffer settings: Best Latency (BL), Best Performance (BP), and Good Latency (GL). Individual points represent repeated measurements and the associated distributions illustrate their variability. Black markers indicate group medians. In the Biased condition, separate black markers indicate the medians of the distributions of SOA errors below and above zero; in the Corrected condition, a single black marker indicates the mean of the entire distribution.

Figure 3.
Trial-by-trial audiovisual SOA errors () measured during the psychophysical validation experiments. Results are shown for five participants (S1–S5) under three experimental conditions: static with a cabled HMD connection, static with a wireless HMD connection, and walking with a wireless HMD connection. Individual points represent the SOA error measured for each experimental trial, and the associated distributions illustrate within-participant variability. Black markers indicate group means.
Figure 3.
Trial-by-trial audiovisual SOA errors () measured during the psychophysical validation experiments. Results are shown for five participants (S1–S5) under three experimental conditions: static with a cabled HMD connection, static with a wireless HMD connection, and walking with a wireless HMD connection. Individual points represent the SOA error measured for each experimental trial, and the associated distributions illustrate within-participant variability. Black markers indicate group means.

Table 1.
Hardware and software specifications of the two VR experimental setups.
| Component | First setup | Second setup |
|---|---|---|
| Unity version | 2021.3.21f1 | 2022.3.62f3 |
| Rendering Pipeline | URP | HDRP |
| XR framework | OpenXR | OpenXR |
| Operating system | Windows 11 | Windows 11 |
| Head-mounted display | Meta Quest 3 | Meta Quest 3 |
| Refresh rate | 72 Hz | 72 Hz |
| Connection | Cabled with Meta Horizon Link | Wifi with Meta Horizon Link |
| Audio output | External headphones | External headphones |
| CPU | Intel Core Ultra 7 265F | Intel Core Ultra 7 265F |
| GPU | NVIDIA GeForce RTX 5060 Ti | NVIDIA GeForce RTX 5060 Ti |
| RAM | 32 GB | 32 GB |
Table 2.
Timing characterization under the different experimental configurations before and after offset compensation. The estimated offset is computed from the mean timing errors measured for positive () and negative () timing errors. After compensation, the residual mean timing error and its standard deviation are reported. BL = Best Latency, GL = Good Latency, BP = Best Performance.
Table 2.
Timing characterization under the different experimental configurations before and after offset compensation. The estimated offset is computed from the mean timing errors measured for positive () and negative () timing errors. After compensation, the residual mean timing error and its standard deviation are reported. BL = Best Latency, GL = Good Latency, BP = Best Performance.
| Biased | Offset-Compensated | |||||
|---|---|---|---|---|---|---|
| Pipeline | Configuration | Offset | Mean() | SD() | ||
| Audio spatialized and cabled | ||||||
| URP | BL | 0.127 | -0.118 | 0.122 | -0.002 | 0.011 |
| BP | 0.183 | -0.155 | 0.169 | 0.001 | 0.021 | |
| GL | 0.148 | -0.134 | 0.141 | 0.000 | 0.013 | |
| HDRP | BL | 0.151 | -0.134 | 0.143 | 0.002 | 0.014 |
| BP | 0.201 | -0.170 | 0.185 | 0.001 | 0.022 | |
| GL | 0.169 | -0.149 | 0.159 | -0.001 | 0.016 | |
| Audio spatialized and Wifi linked | ||||||
| URP | BL | 0.120 | -0.116 | 0.118 | 0.002 | 0.032 |
| BP | 0.170 | -0.144 | 0.157 | 0.000 | 0.026 | |
| GL | 0.138 | -0.128 | 0.133 | 0.003 | 0.029 | |
| HDRP | BL | 0.157 | -0.138 | 0.148 | 0.002 | 0.019 |
| BP | 0.196 | -0.173 | 0.184 | 0.001 | 0.020 | |
| GL | 0.167 | -0.143 | 0.155 | 0.002 | 0.017 | |
| Audio stereo standard and cabled | ||||||
| URP | BL | 0.133 | -0.123 | 0.128 | 0.000 | 0.013 |
| BP | 0.183 | -0.154 | 0.168 | 0.001 | 0.017 | |
| GL | 0.148 | -0.134 | 0.141 | 0.000 | 0.015 | |
| HDRP | BL | 0.143 | -0.130 | 0.137 | -0.003 | 0.011 |
| BP | 0.189 | -0.171 | 0.180 | 0.000 | 0.015 | |
| GL | 0.167 | -0.147 | 0.157 | 0.000 | 0.012 | |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.