Preprint
Article

This version is not peer-reviewed.

Condition-Aware Personalized Modeling of Cross-Register Vocal Stability for Singing Assessment

Submitted:

03 September 2026

Posted:

03 September 2026

You are already at the latest version

Abstract
Vocal stability is central to skilled singing, yet objective assessment remains difficult because adaptations to pitch, vowel, register, and comfort can resemble technical instability. We present a condition-aware framework for personalized assessment in a known-singer setting. Condition-Aware Residual Estimation (CARE) estimates each singer’s expected acoustic representation under specific vocal conditions and predicts teacher-rated quality from deviations from this baseline. A soprano-pretrained SFT branch provides three-level score probabilities, integrated with CARE through strict cross-fitting. Using 295 recordings from 12 singers and 79 independent recording groups, CARE achieved a mean absolute error of 0.837 and Pearson’s r = 0.715. SFT achieved 57.63% accuracy and a macro-F1 of 0.582, while fusion reached 65.76% accuracy and a macro-F1 of 0.652. We also introduce an exploratory Vocal Stability and Smoothness Index (VSSI) combining cross-register residual change, within-group variability, and bootstrap uncertainty. By separating condition-related adaptation from atypical deviation, the framework could support teacher-guided problem localization, individualized practice, and longitudinal monitoring. These educational uses and generalization to unseen singers, teachers, and devices require prospective validation.
Keywords: 
;  ;  ;  ;  ;  
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.