Preprint
Article

This version is not peer-reviewed.

Talker-Learning Transfers Across Modalities

Submitted:

21 July 2026

Posted:

21 July 2026

You are already at the latest version

Abstract
It has been shown that the brain allows transfer of sensory experience across modalities. For instance, our laboratory has shown that one hour of experience lipreading from a talker results in better identification of heard speech (presented in noise) from that same talker than of heard speech from a different talker. Conversely, experience identifying heard speech of a talker facilitates better lipreading from that same talker compared to a different talker. These findings can be interpreted through the supramodal learning hypothesis which suggests that both speech and talker learning involve the extraction of supramodal characteristics—such as articulatory actions—available across modalities. A recent study in our laboratory has shown that training perceivers to identify visual point-light talkers facilitates learning of those same talkers from auditory sinewave speech, compared to a different group of talkers. The current study examined whether this talker-learning transfer might work in the opposite direction. Our results indicated that talkers learned through auditory sinewave speech were better recognized when presented in visual point-light speech. These results are consistent with the previous study showing that talker-specific articulatory information can be shared across modalities.
Keywords: 
;  ;  
Subject: 
Social Sciences  -   Psychology

1. Introduction

There is evidence that speech signals convey information about both the talker and the message being conveyed. The speech literature has shown that talkers can be identified through talker specific vocal qualities (e.g., fundamental frequency of phonation and breathiness) (see Carrell, 1984; Sheffert et al., 2002). However, talkers can also be identified through their articulatory style or idiolect (Fellowes et al., 1997; Remez et al., 1997; Sheffert et al., 2002). For instance, Remez and colleagues (1997) have shown that talkers can still be identified even when the speech signals are reduced to articulatory information only. In their study, talkers were identified through sinewave speech— signals are composed of time-varying sinusoidal patterns similar to the natural resonances of naturally produced utterances.
Notably, this talker-specific articulatory information is also available for lipreading. As shown by Rosenblum and colleagues (2007), talkers can be identified through the visual modality after isolating visible articulatory information using a point-light method. To create point-light stimuli, fluorescent dots are placed on the lips, teeth, jaw, and face of talkers while they are filmed against a black background. The final stimuli remove the typical facial features and retain articulatory information (Rosenblum et al., 1996; Rosenblum & Saldana, 1996). This articulatory information is conveyed through movement since still point-images cannot be identified as faces. However, dynamic point-light faces (despite removing facial characteristics) provide information about talkers through their unique articulatory style. This allows familiar talkers to be identified at better than chance levels (Rosenblum et al., 2007) and also facilitates matching of point-light faces (Rosenblum et al., 2002).
Adding visual speech to the auditory signal has been shown to facilitate talker learning. There is evidence that training to recognize talkers audiovisually results in improved auditory-only talker identification (see Sheffert & Olson, 2004; von Kriegstein et al., 2008; Zadoorian & Rosenblum, 2023). An explanation for these results is the connection between the face and voice processing brain areas (e.g., Black et al., 2011; Schall & von Kriegstein, 2014; von Kriegstein & Giraud, 2006). Importantly, this cross-modal transfer of talker learning is not due to associative experience, as presenting written names along with voices does not facilitate talker learning (see von Kriegstein & Giraud, 2006; von Kriegstein et al., 2008). Rather, these findings suggest that speech perception relies on shared representations that are accessible across sensory modalities.
Consistent with this view, theories of multisensory speech perception propose that speech signals are integrated through common articulatory information available across modalities. The supramodal learning hypothesis proposes that perceivers extract this shared articulatory information during speech learning, allowing experience in one modality (e.g., auditory speech) to facilitate perception in another modality (e.g., visual speech) (Fowler, 2004; Rosenblum, 2008). For instance, there is evidence that learning to lipread from a talker for about an hour resulted in better performance of hearing the same talker’s speech in noise compared to hearing the speech from a different talker (Rosenblum et al., 2007). Furthermore, the same cross-modal training also works in the opposite direction. Learning to perceive speech from a talker in noise improves subsequent lipreading performance for that same talker compared to lipreading performance for a different talker (Sanchez et al., 2013). These results suggest that talker-specific articulatory information embedded in speech can also be shared across modalities.
Importantly, this supramodal representation may not be limited to phonetic information but may also include talker-specific articulatory characteristics for talker recognition. Lachs and Pisoni (2004) have demonstrated that talkers can be matched using sinewave and point-light talker stimuli. Similarly, in a more recent study, Simmons and colleagues (2021) found evidence for cross-modal talker learning. After learning to identify talkers in the visual modality (using the point-light technique), participants were better able to identify the same talkers in the auditory modality (through sinewave speech). Taken together these results suggest that learning to identify a talker in one modality can be shared across modalities.

1.1. Purpose of the Current Study

As stated above, a recent study has shown that talker learning can be shared across modalities. Simmons and colleagues (2021) initially trained participants to identify point-light faces and were then tested identifying sinewave voices. Half of the participants were trained on the same set of talkers across the point-light and sinewave blocks, and the other half were trained on a different set of talkers across the modalities. As indicated by their results, learning to identify talkers during point-light training facilitated subsequent learning of the same talkers during sinewave training compared to participants who were trained with different sets of talkers. These results suggest that learning to identify talkers in the visual modality can transfer to the auditory modality.
The goal of the current study was to see whether this cross-modal talker transfer can occur in the opposite direction. In Experiment 1, we demonstrated that, in an online setting, learning to identify talkers through sinewave speech facilitated subsequent identification of the same talkers when they were presented as point-light faces. Experiment 2 replicated and extended these findings by validating the effect in an in-person setting.

2. Materials and Methods

Experiment 1: Online Study of Cross-Modal Talker Identity Using the Same Sentence

2.1. Method

Due to COVID related protocol, the experiment was administered online through Pavlovia.org (an online platform used for experiments created through PsychoPy; Peirce et al., 2019). Each participant was monitored by an experimenter over a Zoom connection. All participants were asked to share their screen and were instructed to use earphones. The experiment took about an hour and a half to complete, and participants received course credit for participating.

2.2. Participants

A total of forty-eight undergraduates (Mean age = 19.55, S = 1.73) from the University of California, Riverside were included in the study. Twenty-four (19 females) participants were randomly assigned to the same-talker group and twenty-four (16 females) participants were randomly assigned to the different-talker group. These sample sizes are consistent with past studies (e.g., Simmons et al., 2021). Forty-seven of these participants were native English speakers. All participants reported having normal hearing and normal (or corrected) vision.

2.3. Materials

Stimuli were recordings of 3 male and 2 female native English talkers. Talkers were video recorded using a SONY DRC-TRV11 (Tokyo, Japan) camcorder uttering sentences from the Bamford-Kowal-Bench sentence list (Bench et al., 1979).
Dynamic Point-Light Speech: Point-light videos were created by placing 15 small fluorescent dots (each 0.12 inch in diameter) on the cheeks, forehead, and jawline of the talkers’ face. Additionally, another 15 dots were placed on the talkers’ mouth, including their teeth, tongue, and lips. Talkers were filmed articulating speech while placing their face inside a black cardboard box with four plastic masks attached to it. Importantly, both the black cardboard and the masks were covered with numerous dots to prevent participants from recognizing faces by memorizing the dot patterns. This technique results in displays consisting of just white moving dots (hiding facial characteristics) against a black background while maintaining critical articulatory information (see Figure 1a). To further prevent participants from memorizing the location of the dots, each talker was recorded nine times using a quasi-random configuration to change the place of the dots. For more details about this method see Rosenblum et al., (2002) and Simmons et al., (2021).
Sinewave Speech: Audio from the point-light video recordings were extracted using the Final Cut Pro software and were normalized (89 total; eight audios for talker F1 and nine for all other talkers). Praat software (Boersma, 2001) was then used to create the sinewave replicas of the stimuli by extracting the center frequencies of the first three formants (see Figure 1b). This technique removes natural voice quality by maintaining time-varying Spectro-temporal information (e.g., Remez et al., 1997; Sheffert et al., 2002). For more details see Simmons et al., (2021).
The stimuli differed only during the sinewave training phase. Participants assigned to the different-talker condition were trained with speech produced by five talkers and were subsequently tested with a different set of five talkers during the point-light phase. In contrast, participants assigned to the same-talker condition were trained with a different set of five talkers and were tested with the same five talkers during the point-light phase. Thus, the point-light talkers were identical across conditions, whereas the sinewave talkers differed. Assignment to the same- and different-talker conditions was counterbalanced across participants.

2.4. Procedure

Participants initially received instructions for the sinewave training phase, and they were not aware that they would be presented with point-light videos later. After completing the sinewave training phase, they were given a brief break and were then presented with instructions for the point-light test. Participants assigned to the same-talker group were never told they were presented with the same talkers during the point-light test phase. Similarly, those assigned to the different-talker group were not told they were presented with different talkers. For this reason, the names assigned to the talkers were different in two modality phases of the experiment. The experiment consisted of four total phases: familiarization for sinewave talkers, sinewave training, familiarization for point-light talkers, and point-light test. During all these phases, the talkers’ utterances of the sentence “The football game is over.” were used.
Sinewave Familiarization Phase: During this phase, participants were informed that they would learn to identify talkers based on hearing their voices. They were told that the voices would not sound normal but instead the voices were described as sounding like whistles. On a given trial, participants were first presented with the talker’s name as text on the screen (e.g., “This is Liz”) then heard the talker uttering the sentence “The football game is over” followed by another text presentation of the talker’s name (e.g., “That was Liz”). During this phase, there were two repetitions of each of the talker’s utterance (using two different utterances) for a total of 20 randomized trials. Participants did not make any responses during this phase.
Sinewave Training Phase: During the sinewave training phase, participants were presented with eight repetitions of four sinewave utterances (taken from different recording from those presented during the familiarization phase) of the five talkers. This created a total of 160 training trials. On a given trial, a single utterance was presented, and participants were asked to press 1 to 5 keys on their keyboard, with each talker assigned a number corresponding to the numbers on the keyboard. Participants were able to see the talkers’ names and corresponding numbers on the bottom of their screen until they made a response. After each response, participants were given feedback (i.e., correct vs. incorrect) followed by the correct name of the talker (e.g., “That was Liz”).
Point-Light Familiarization Phase: After the sinewave training, participants were told they would be identifying talkers by watching their point-light videos. Participants were informed that the faces presented in the videos would appear different from natural faces, consisting of white dots moving against a black background. On a given trial, participants were first presented with the talker’s name on the screen (e.g., “This is Joe”), followed by the talker’s silent articulation of the point-light video. After viewing each point-light face, the talker’s name was presented on the screen again (e.g., “That was Joe”). Similar to the sinewave phase, there were two repetitions of each of the talkers’ utterances (using two different utterances) for a total of 20 randomized trials. Participants did not make any responses during this phase.
Point-light Test Phase: This phase was similar to that of the sinewave training phase. Participants were presented with four point-light utterances (different than the ones presented during the familiarization phase) of the five talkers. Each utterance was repeated eight times, totaling fully randomized 160 trials. Participants assigned to the same-talker group were presented with different utterances from those used in the sinewave training phase. On a given trial, participants were presented with a talker’s point-light video and were then instructed to identify the talker by pressing 1 to 5 keys on their keyboard. Each talker was assigned a number corresponding to the numbers on their keyboard. The names of the talkers appeared on the screen until participants responded, after which they received feedback indicating whether their response was correct or incorrect, along with the correct name of the talker.
Participants completed the experiment on a laptop or desktop computer using Chrome, Microsoft Edge, or Firefox. The average screen size used was 14.36 ″ (ranging from 11″ to 27″). About 81% of our participants were using headphones (e.g., Air pods, over-ear headphones) during the sinewave training phase. Each participant adjusted the sound to a comfortable listening level.
Experiment 2: In-Person Validation of Experiment 1 Results
The results of Experiment 1 indicated that learning to identify talkers through sinewave speech facilitates learning of the same talkers when presented in point-light speech. These results suggest that talker-specific characteristics are available in both auditory and visual modalities and can also be shared across modalities. To validate these findings, we replicated Experiment 1 in person.

2.5. Participants

A total of fifty-one undergraduates (Mean age = 19.16, S = 1.41) were recruited from the University of California, Riverside. Twenty-four participants (12 females) were randomly assigned to the same-talker group and twenty-seven (15 females) were randomly assigned to the different-talker group. We made every effort to ensure that the sample size for this experiment was consistent with that of Experiment 1. All participants were Native English speakers with self-reported normal hearing and normal-to-corrected vision and received course credit as compensation.

2.6. Materials and Procedure

The same materials as Experiment 1 were used. The training procedures for both sinewave and point-light phases were the same as Experiment 1. Since the study was conducted in-person, the experiment was administered through PsychoPy (Pierce, 2007), using a Mac OS system. Participants listened to the sinewave stimuli using the Sony MDR 7506 headphones.

3. Results

Experiment 1: Online Study of Cross-Modal Talker Identity Using the Same Sentence
During the sinewave training, participants trained with the same-talker group (M = 0.72, S = 0.12) identified the talkers at better than chance levels (20%), t(23) = 29.90, p < 0.001, Cohen’s d = 0.12. Similarly, those trained with the different-talker group (M = 0.65, S = 0.14) identified the talkers at better than chance (20%), t(23) = 23.15, p < 0.001, Cohen’s d = 0.14. Each talker was also identified at better than chance (10%) for both groups (corrected α = 0.005) (see Figure 2).
During the point-light training, participants identified the talkers at better than chance (20%) for both the same-talker (M = 0.55, S = 0.16); t(23) = 16.60, p < 0.001, Cohen’s d = 0.16 and different-talker groups (M = 0.46, S = 0.16); t(23) = 14.25, p < 0.001, Cohen’s d = 0.16. All talkers were identified at the above chance (20%) for both groups (corrected α = 0.001) (see Figure 3). Taken together, these results suggest that participants were able to identify talkers from sinewave and point-light utterances.
Effects of same vs. different talker training: Talker identification during sinewave training was not found to depend on whether participants were trained with the same or different talker groups, t(8) = 1.45, p = 0.09, Cohen’s d = 0.07. However, talkers were better identified during the point-light test phase when participants had been trained with those same speakers during the sinewave training phase, t(4) = 7.75, p < 0.001, Cohen’s d = 0.03. These results show that learning to identify the talkers auditorily through sinewave speech facilitated better identification of those same talkers presented visually in point-light speech.
Experiment 2: In-Person Validation of Experiment 1 Results
During the sinewave training, participants trained with the same-talker group (M = 0.72, S = 0.09) identified the talkers at better than chance (20%), t(23)= 41.22, p < 0.001, Cohen’s d = 0.09. Similarly, participants trained with the different-talker group (M = 0.69, S = 0.07) also identified the talkers at better than chance (20%), t(26)= 38.59, p < 0.001, Cohen’s d = 0.09. All talkers were identified at better than chance (10%) for both the same-talker and different-talker groups (corrected α = 0.005) (see Figure 4).
Furthermore, participants trained with the same-talker group during the point-light test phase (M = 0.56, S = 0.13) identified the talkers at better than chance (20%), t(23)= 21.07, p < 0.001, Cohen’s d = 0.13. Also, those trained with the different-talker group (M = 0.51, S = 0.13) identified the talkers at better than chance (20%), t(26)= 19.70, p < 0.001, Cohen’s d = 0.13. Talkers included during the point-light test phase, were also identified at above chance (20%) for both groups (corrected α = 0.001) (see Figure 5).
Effects of same vs. different talker training: Talker During sinewave training, participants trained with the same-talker group performed similarly to those trained with the different-talker group, t(8)= 0.60, p = 0.28, Cohen’s d = 0.08. However, during the point-light test phase, talkers were better identified when they had been trained in sinewave, t(4)= -3.22, p = 0.02, Cohen’s d = 0.04. Consistent with the findings from Experiment 1, auditory training with sinewave speech enhanced participants’ ability to identify the same talkers when they were later presented visually using point-light speech.

4. Discussion

This study examined how articulatory information contributes to learning talker identities across sensory modalities. Previous work by Simmons and colleagues (2021) demonstrated that talker-identity learning can transfer from point-light speech to sinewave speech. The present study investigated whether this cross-modal transfer is bidirectional by examining whether learning to identify talkers through sinewave speech facilitates subsequent identification of those same talkers when they are presented as point-light faces. Both sinewave and point-light speech preserve information about a talker’s articulatory patterns (e.g., Lachs & Pisoni, 2004; Remez et al., 1997; Rosenblum et al., 2007).
In Experiment 1, participants were trained to identify a set of five talkers auditorily using sinewave speech. They were then tested on their ability to identify either the same talkers or a different set of talkers when presented visually as point-light faces. Participants trained with the same talkers during the sinewave phase identified those talkers more accurately during the point-light phase than participants trained with different talkers, demonstrating that talker-specific learning transferred from the auditory to the visual modality. To determine whether this effect generalized beyond the online testing environment, Experiment 2 replicated the same procedure in an in-person setting. The results replicated those of Experiment 1, indicating that learning talker-specific articulatory information auditorily facilitates subsequent visual identification of the same talkers.
Together, these findings provide evidence that talker-identity learning transfers across sensory modalities. These cross-modal talker-learning findings are consistent with previous research (see Simmons et al., 2021) and align with the supramodal learning hypothesis (Fowler, 2004; Rosenblum, 2008), which proposes that perceivers extract shared articulatory information across modalities. Thus, learning to recognize a talker through one sensory modality may facilitate recognition in another modality because both modalities provide access to common talker-specific articulatory characteristics. This shared articulatory information may reflect the unique production patterns that characterize individual talkers (see Fellowes et al., 1997; Remez et al., 1997; Sheffert et al., 2002). Consistent with this possibility, talker-specific phonetic cues have been shown to support speech perception across sensory modalities (Rosenblum et al., 2007; Sanchez et al., 2013).
The findings of the current study have practical implications. As shown by previous research, those with moderate hearing loss or cochlear implant(s) have difficulty identifying talkers (Cullington & Zeng, 2011; Vongphoe & Zeng, 2005). Therefore, the multisensory training might help those individuals to better recognize talkers, which may ease their everyday interactions. The findings can also ultimately help researchers to design telecommunications systems in order to help those with hearing and language disorders.

Funding

This research received no external funding.

Institutional Review Board Statement

The study was conducted in accordance with the Declaration of Helsinki, and approved by the Institutional Review Board of University of California, Riverside (protocol code HS 05-035 approved on July 2019).

Data Availability Statement

The data presented in this study are available on request from the corresponding author.

Conflicts of Interest

The authors declare no conflict of interest.

References

  1. Bench, J.; Kowal, Å.; Bamford, J. The BKB (Bamford-Kowal-Bench) sentence lists for partially-hearing children. Br. J. Audiol. 1979, 13(3), 108–112. [Google Scholar] [CrossRef] [PubMed]
  2. Blank, H.; Anwander, A.; von Kriegstein, K. Direct structural connections between voice-and face-recognition areas. J. Neurosci. 2011, 31(36), 12906–12915. [Google Scholar] [CrossRef] [PubMed]
  3. Boersma, Paul. Praat, a system for doing phonetics by computer. Glot Int. 2001, 5:9/10, 341–345. [Google Scholar]
  4. Carrell, T. D. CONTRIBUTIONS OF FUNDAMENTAL FREQUENCY, FORMANT SPACING, AND GLOTTAL WAVEFORM TO TALKER IDENTIFICATION (SPEECH, PERCEPTION, SPEAKER); Indiana University, 1984. [Google Scholar]
  5. Cullington, H. E.; Zeng, F. G. Comparison of bimodal and bilateral cochlear implant users on speech recognition with competing talker, music perception, affective prosody discrimination and talker identification. Ear Hear. 2011, 32(1), 16. [Google Scholar] [CrossRef] [PubMed]
  6. Fellowes, J. M.; Remez, R. E.; Rubin, P. E. Perceiving the sex and identity of a talker without natural vocal timbre. Percept. Psychophys. 1997, 59(6), 839–849. [Google Scholar] [CrossRef] [PubMed]
  7. Fowler, C. A. Speech as a supramodal or amodal phenomenon. 2004. [Google Scholar] [CrossRef] [PubMed]
  8. Lachs, L.; Pisoni, D. B. Crossmodal source identification in speech perception. Ecol. Psychol. 2004, 16(3), 159–187. [Google Scholar] [CrossRef] [PubMed]
  9. Peirce, J. W. PsychoPy—psychophysics software in Python. J. Neurosci. Methods 2007, 162(1-2), 8–13. [Google Scholar] [CrossRef] [PubMed]
  10. Peirce, J. W.; Gray, J. R.; Simpson, S.; MacAskill, M. R.; Höchenberger, R.; Sogo, H.; Kastman, E.; Lindeløv, J. PsychoPy2: experiments in behavior made easy. Behav. Res. Methods. 2019. [Google Scholar] [CrossRef] [PubMed]
  11. Remez, R. E.; Fellowes, J. M.; Rubin, P. E. Talker identification based on phonetic information. J. Exp. Psychol. Hum. Percept. Perform. 1997, 23(3), 651. [Google Scholar] [CrossRef] [PubMed]
  12. Rosenblum, L. D. Speech perception as a multimodal phenomenon. Curr. Dir. Psychol. Sci. 2008, 17(6), 405–409. [Google Scholar] [CrossRef] [PubMed]
  13. Rosenblum, L.D.; Johnson, J.A.; Saldafia, H.M. Visual kinematic information for embellishing speech in noise. J. Speech Hear. Res. 39 1996, 1159–11. [Google Scholar] [CrossRef]
  14. Rosenblum, L. D.; Miller, R. M.; Sanchez, K. Lip-read me now, hear me later: Cross-modal transfer of speaker familiarity effects. Psychol. Sci. 2007, 18(5), 392–396. [Google Scholar] [CrossRef] [PubMed]
  15. Rosenblum, L. D.; Niehus, R. P.; Smith, N. M. Look who’s talking: recognizing friends from visible articulation. Perception 2007, 36(1), 157–159. [Google Scholar] [CrossRef] [PubMed]
  16. Rosenblum, L. D.; Saldaña, H. M. An audiovisual test of kinematic primitives for visual speech perception. J. Exp. Psychol. Hum. Percept. Perform. 1996, 22(2), 318. [Google Scholar] [CrossRef] [PubMed]
  17. Rosenblum, L. D.; Yakel, D. A.; Baseer, N.; Panchal, A.; Nodarse, B. C.; Niehus, R. P. Visual speech information for face recognition. Percept. Psychophys. 2002, 64(2), 220–229. [Google Scholar] [CrossRef] [PubMed]
  18. Sanchez, K.; Dias, J. W.; Rosenblum, L. D. Experience with a talker can transfer across modalities to facilitate lipreading. Atten. Percept. Psychophys. 2013, 75(7), 1359–1365. [Google Scholar] [CrossRef] [PubMed]
  19. Schall, S.; von Kriegstein, K. Functional connectivity between face-movement and speech-intelligibility areas during auditory-only speech perception. PLoS ONE 2014, 9(1), e86325. [Google Scholar] [CrossRef] [PubMed]
  20. Sheffert, S. M.; Olson, E. Audiovisual speech facilitates voice learning. Percept. Psychophys. 2004, 66(2), 352–362. [Google Scholar] [CrossRef] [PubMed]
  21. Sheffert, S. M.; Pisoni, D. B.; Fellowes, J. M.; Remez, R. E. Learning to recognize talkers from natural, sinewave, and reversed speech samples. J. Exp. Psychol. Hum. Percept. Perform. 2002, 28(6), 1447. [Google Scholar] [CrossRef] [PubMed]
  22. Simmons, D.; Dorsi, J.; Dias, J. W.; Rosenblum, L. D. Cross-modal transfer of talker-identity learning. Atten. Percept. Psychophys. 2021, 83(1), 415–434. [Google Scholar] [PubMed]
  23. Vongphoe, M.; Zeng, F. G. Speaker recognition with temporal cues in acoustic and electric hearing. J. Acoust. Soc. Am. 2005, 118(2), 1055–1061. [Google Scholar] [CrossRef] [PubMed]
  24. von Kriegstein, K.; Dogan, Ö.; Grüter, M.; Giraud, A. L.; Kell, C. A.; Grüter, T.; Kiebel, S. J. Simulation of talking faces in the human brain improves auditory speech recognition. Proc. Natl. Acad. Sci. 2008, 105(18), 6747–6752. [Google Scholar] [CrossRef] [PubMed]
  25. Von Kriegstein, K.; Giraud, A. L. Implicit multisensory associations influence voice recognition. PLoS Biol. 2006, 4(10), e326. [Google Scholar] [CrossRef] [PubMed]
  26. Zadoorian, S.; Rosenblum, L. D. The benefit of bimodal training in voice learning. Brain Sci. 2023, 13(9), 1260. [Google Scholar] [CrossRef] [PubMed]
Figure 1. Examples of stimuli of a talker’s speech utterance of the sentence “The football game is over.” Figure b is adapted from Simmons and colleagues (2021).
Figure 1. Examples of stimuli of a talker’s speech utterance of the sentence “The football game is over.” Figure b is adapted from Simmons and colleagues (2021).
Preprints 224226 g001
Figure 2. This figure represents the talker identification accuracy for sinewave training conducted online. Talkers F1-M5 (shown in dark gray) were part of the same-talker group, while talkers F2-M6 (shown in horizontal gray stripes) were included in the different-talker training group. In this experiment, the same sentence was used during both the sinewave and point-light phases.
Figure 2. This figure represents the talker identification accuracy for sinewave training conducted online. Talkers F1-M5 (shown in dark gray) were part of the same-talker group, while talkers F2-M6 (shown in horizontal gray stripes) were included in the different-talker training group. In this experiment, the same sentence was used during both the sinewave and point-light phases.
Preprints 224226 g002
Figure 3. This figure represents the talker identification accuracy during the point-light test phase for both the same-talker and different-talker groups. In this online experiment, the same sentence was used during both the sinewave and point-light phases.
Figure 3. This figure represents the talker identification accuracy during the point-light test phase for both the same-talker and different-talker groups. In this online experiment, the same sentence was used during both the sinewave and point-light phases.
Preprints 224226 g003
Figure 4. This figure represents the talker identification accuracy for sinewave training conducted in-person. Talkers F1-M5 (shown in dark gray) were part of the same-talker group, while talkers F2-M6 (shown in horizontal gray stripes) were included in the different-talker training group. In this experiment, the same sentence was used during both the sinewave and point-light phases.
Figure 4. This figure represents the talker identification accuracy for sinewave training conducted in-person. Talkers F1-M5 (shown in dark gray) were part of the same-talker group, while talkers F2-M6 (shown in horizontal gray stripes) were included in the different-talker training group. In this experiment, the same sentence was used during both the sinewave and point-light phases.
Preprints 224226 g004
Figure 5. This figure represents the talker identification accuracy during the point-light test phase for both the same-talker and different-talker groups. In this in-person experiment, the same sentence was used during both the sinewave and point-light phases.
Figure 5. This figure represents the talker identification accuracy during the point-light test phase for both the same-talker and different-talker groups. In this in-person experiment, the same sentence was used during both the sinewave and point-light phases.
Preprints 224226 g005
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.
Prerpints.org logo

Preprints.org is a free preprint server supported by MDPI in Basel, Switzerland.

Subscribe

© 2026 MDPI (Basel, Switzerland) unless otherwise stated

Accessibility

Disclaimer

Terms of Use

Privacy Policy

Privacy Settings