Submitted:
07 September 2026
Posted:
09 September 2026
You are already at the latest version
Abstract
Purpose: Respiratory-event rates quantify event frequency but do not describe how respiratory-related acoustic activity is expressed across an overnight recording. We investigated whether a score-defined digital acoustic representation captures information not reducible to annotated respiratory-event rate.Methods: We analyzed 32 overnight home respiratory polygraphy recordings from the APSAA dataset, including audio, respiratory annotations, and oxygen saturation. Twenty-eight recordings yielded at least one qualifying acoustic episode (8,512 episodes). Five acoustic features—episode density, relative intensity, duration, long-event occurrence, and temporal distribution—formed a composite score defining three acoustic phenotypes using fixed heuristic thresholds. Temporal compaction quantified short inter-episode intervals and was adjusted for episode quantity.Results: Scores were computationally reproduced with identical phenotype assignment in 28/28 recordings (mean absolute difference 0.0011; maximum 0.0046). Two recordings with comparable annotation-derived respiratory-event rates (30.45 vs 33.38 events/h; 217 vs 232 annotated events) and mean SpO₂ (91.75% vs 91.61%) contained 382 vs 5 acoustic episodes and scores of 0.685 vs 0.145. Load-adjusted temporal compaction differed among phenotypes (H=8.192, p=0.0166, ε²=0.248). The effect persisted after removing the historical 120-s interval exclusion (H=8.108, p=0.0174, ε²=0.244); residual association with episode count remained. The low- versus intermediate-score contrast was robust, whereas the high-score contrast was less stable.Conclusions: Overnight respiratory audio can provide reproducible digital acoustic phenotypes describing the amount and temporal organization of detected acoustic activity beyond annotated event frequency. These are signal representations, not clinical phenotypes or sleep stages. Repeated-night studies are required to establish within-person stability, transitions, and clinical meaning.

Keywords:
digital phenotyping
; respiratory acoustics
; overnight audio
; temporal organization
; sleep-disordered breathing
; reproducibility
1. Introduction
Obstructive sleep apnea (OSA), a major form of sleep-disordered breathing (SDB), affects a substantial proportion of the adult population worldwide [1]. Clinical characterization relies heavily on the apnea–hypopnea index (AHI), which summarizes the frequency of scored respiratory events per hour of sleep [2]. Although clinically established, event frequency does not retain the duration or the associated physiological burden of individual events [2,3], nor does it preserve their temporal order or distribution within the night. Recordings with similar event rates may contain activity concentrated in short periods, dispersed across extended intervals, or distributed unevenly over time. These temporal patterns cannot be reconstructed from a single hourly rate.
This limitation is relevant because SDB is physiologically heterogeneous [4,5]. Similar AHI values may arise from different combinations of upper-airway collapsibility, ventilatory control, arousal threshold, and neuromuscular responsiveness [6,7,8]. An acoustic recording cannot be assumed to identify these mechanisms directly. It can, however, preserve another observable layer of the overnight recording: how respiratory-related acoustic activity is expressed and organized over time.
Respiratory acoustics provide a non-invasive source of information about overnight breathing. Previous studies have used respiratory sounds to detect apnea-related events, estimate OSA severity, predict oxygen desaturation, or classify recordings using spectral, temporal, and machine-learning features [9,10,11,12,13]. Even where temporal and nonstationary characteristics have been incorporated, they have served as predictors of an externally defined event label or severity class [12]. A complementary question is whether the acoustic recording itself can be represented in a way that preserves multidimensional differences in how respiratory-related acoustic activity is expressed across the night.
We use the term digital acoustic phenotype for such a representation. The concept of the digital phenotype was introduced to describe characteristics of an individual captured through interaction with digital technology [14], and has since been developed as the moment-by-moment quantification of state from data collected by personal devices in everyday settings [15]. Our use is deliberately narrower. Here, digital acoustic phenotype refers to a score-defined computational description of observable acoustic characteristics within a single overnight recording. It does not denote a clinical OSA phenotype, an inferred sleep stage, or the longitudinal in-situ digital phenotype derived from repeated personal-device measurements. The representation combines properties of detected acoustic episodes including their density, relative intensity, duration, occurrence of prolonged episodes, and distribution across the recording. The resulting score and its fixed heuristic thresholds provide a deterministic description of the recording rather than a data-driven discovery of latent biological classes.
A useful digital acoustic phenotype should capture more than a relabeling of respiratory-event frequency. This distinction can be examined in two ways. First, recordings with comparable independently annotated respiratory-event rates may nevertheless differ in their acoustic representation. Second, a property of temporal organization that does not enter the score itself can be examined after accounting for the number of detected acoustic episodes. Here, temporal compaction describes the tendency of consecutive acoustic episodes to occur after short intervening intervals and provides an additional characterization of how detected activity is distributed through time.
Using overnight audio and respiratory polygraphy from the APSAA dataset [16,17], we examined whether a score-defined digital acoustic phenotype could be reproducibly reconstructed from a frozen acoustic-event representation and whether this representation contained information not reducible to annotated respiratory-event frequency alone. We reconstructed the acoustic score and phenotype assignments, compared acoustic expression with annotation-derived respiratory-event rate, and examined load-adjusted temporal compaction across the score-defined phenotypes. Because APSAA provides one overnight recording per participant, the present study evaluates recording-level digital acoustic phenotyping rather than within-person reproducibility. Phenotype stability, transitions, and longitudinal meaning remain questions for repeated-night studies, particularly given the substantial night-to-night variability of sleep-disordered breathing [18].
2. Methods
2.1. Study Design and Data Source
This exploratory study was a secondary analysis of version 2.0 of the Audio-Polygraphy Dataset for Sleep Apnea Analysis (APSAA), a dataset containing one synchronized overnight audio and respiratory polygraphy recording from each of 32 adults, available upon reasonable request through its Zenodo repository [16,17]. The recordings were obtained from 30 men and two women between September 2021 and April 2022 at the Sleep Unit of Dr. Sagaz Hospital in Jaén, Spain [17]. Individuals younger than 18 years and patients receiving continuous positive airway pressure therapy were not included in the source dataset [17].
Respiratory polygraphy was performed using an Embletta Gold system and included nasal airflow, oronasal thermal airflow, thoracic and abdominal respiratory effort, peripheral oxygen saturation (SpO₂), pulse rate, and a snore-pressure signal. Respiratory events were initially identified using Embla RemLogic 3.1 and subsequently reviewed manually under medical supervision according to American Academy of Sleep Medicine criteria [17]. The dataset does not contain electroencephalography, electrooculography, or electromyography; therefore, sleep stages, electroencephalographically confirmed sleep duration, and cortical arousals were unavailable.
Audio was originally recorded using a ZOOM H4n Pro recorder positioned approximately 30 cm above the participant's head at a sampling rate of 44,100 Hz. During preparation of the public dataset, the source investigators converted the audio to a single channel and downsampled it to 4,000 Hz. The specific upstream resampling algorithm was not reported. The source investigators also removed the first and last 60 min of each recording to reduce wake-related content, remove instructions or other staff speech, and protect participant anonymity [17]. Accordingly, all recording-duration measures in the present study refer to the duration of the distributed files rather than the complete acquisition period or electroencephalographically confirmed sleep time.
The audio and polygraphy signals were acquired using separate devices and were not synchronized by the recording equipment. The APSAA release provides files temporally aligned by the source investigators [16,17]. The present analysis used these synchronized files without modifying their timestamps. The separately available automatic synchronization algorithm described by the dataset authors was not executed, and audio–polygraphy alignment was not independently re-estimated. Analyses involving the temporal relationship between acoustic and polygraphic signals therefore depend on the alignment supplied with the dataset.
The source study received approval from the Provincial Research Ethics Committee of Jaén, Spain, and all participants provided written informed consent [16,17]. The present study involved no additional participant recruitment, intervention, or collection of identifiable information and used only the de-identified data distributed through the APSAA repository.
2.2. Acoustic Signal Processing and Episode Representation
The analysis used the monaural 4,000-Hz WAV files distributed with APSAA. Audio was analyzed at its nominal sampling rate of 4,000 Hz without further resampling. Recording duration was defined as the complete duration of the distributed audio file, calculated as the number of audio samples divided by the sampling rate. This duration was used as the recording-level time denominator throughout the acoustic analysis.
A root-mean-square (RMS) amplitude envelope was calculated using a frame length of 512 samples and a hop length of 128 samples, corresponding to a 128-ms analysis window and a 32-ms frame step. RMS calculation was performed without frame centering. The resulting RMS sequence was smoothed using a 20-frame uniform moving-average filter, corresponding to approximately 0.64 s.
For each recording, an individual acoustic reference was defined as the 75th percentile of the smoothed RMS envelope, and the detection threshold was set at 50% of this reference. A candidate reduced-amplitude acoustic episode began when the smoothed envelope fell below the threshold and ended when it subsequently returned to or exceeded the threshold. Consecutive below-threshold frames were treated as a single candidate interval, and intervals lasting at least 5 s were retained as atomic acoustic episodes.
A below-threshold interval remaining open at the end of the RMS sequence was closed at the temporal boundary immediately following the final available RMS frame and was retained if its resulting duration satisfied the same 5-s criterion. This envelope-derived endpoint may precede the end of the raw audio file by the residual duration not represented by the RMS frame sequence. No additional exclusion of episodes near the beginning or end of a recording was applied.
For each atomic episode, onset, offset, duration, and RMS amplitude relative to the individual acoustic reference were derived. The segmentation parameters were inherited from the original exploratory analysis and were not optimized against respiratory-polygraphy annotations. Detected intervals were therefore treated as reduced-amplitude respiratory acoustic episodes rather than acoustically identified apneas or hypopneas.
A merged representation was subsequently constructed from the retained atomic episodes. The 5-s minimum-duration criterion was applied before merging. Consecutive retained episodes separated by less than 3 s were combined into a single acoustic complex extending from the onset of the first constituent episode to the offset of the last. No border exclusion was applied before or after merging. The duration of a merged complex therefore included intervening above-threshold gaps shorter than 3 s; these intervals were not interpreted as continuous acoustic suppression, apnea duration, or hypoxic time.
The merged representation included the relative-amplitude attribute used in the frozen phenotype reconstruction. Because the exact aggregation rule by which this attribute was derived from constituent atomic episodes was not independently re-established in the present audit, the phenotype analysis used the amplitude values preserved in the frozen merged representation rather than redefining this operation.
The merged representation was used to construct the score-defined digital acoustic phenotype described in Section 2.3. Across the 32 APSAA recordings, 28 yielded at least one qualifying acoustic complex, producing a frozen representation of 8,512 merged complexes used for phenotype reconstruction and the analyses reported here. Four recordings yielded no qualifying complexes and therefore could not contribute to phenotype construction or analyses requiring detected acoustic episodes.
For episode-rate calculations and division of each recording into temporal segments, recording duration was defined from the complete distributed audio file rather than from the endpoint of the RMS envelope or electroencephalographically confirmed sleep time.
2.3. Digital acoustic Phenotype Construction
A recording-level acoustic score was reconstructed from the frozen merged-complex representation described in Section 2.2. The score originated from the original exploratory analysis and was retained without refitting its component definitions, weights, or category thresholds. It was a deterministic composite of five features describing the amount, amplitude, duration, and temporal distribution of acoustic activity within a recording. The score was not derived by clustering, supervised classification, or optimization against clinical outcomes.
For each of the 28 recordings containing at least one merged acoustic complex, five normalized components were calculated. Intermediate feature values were rounded according to the original implementation before normalization: complexes per hour to two decimal places, mean relative amplitude to three decimal places, mean complex duration to one decimal place, and late long-event fraction to three decimal places. Late temporal bias was not rounded before entering the composite calculation. The final composite score was rounded to three decimal places.
Density (D) represented the number of merged complexes per hour of distributed recording. After rounding the event rate rr to two decimal places, density was calculated as D=min(r/30,1)D = \min(r/30, 1).
Relative amplitude (A) was based on the mean relative-amplitude value across merged complexes, as defined in Section 2.2. After rounding the recording-level mean relative amplitude aa to three decimal places, the component was calculated as A=min(a/0.8,1)A = \min(a/0.8, 1).
Duration (U) represented the mean duration of merged complexes. After rounding mean duration uu to one decimal place, the component was calculated as U=min(u/60,1)U = \min(u/60, 1).
Two additional components described the temporal distribution of complexes within the recording. Each merged complex was assigned to the first, middle, or final third of the complete distributed recording according to its temporal midpoint. Late long-event fraction (L) was the proportion of complexes assigned to the final third whose duration exceeded 20 s, rounded to three decimal places. When no complexes occurred in the final third, LL was set to zero.
Late temporal bias (T) compared the number of complexes in the final third (n_late) with the number in the first third (n_early). It was calculated as:T = min[n_late / (n_early + 1), 2] / 2.The additive constant prevented division by zero, and capping the ratio at 2 before division by 2 constrained the normalized component to the interval from 0 to 1. T was not rounded before entering the composite score.
The composite acoustic score was calculated as S=0.30D+0.20A+0.15U+0.15L+0.20TS = 0.30D + 0.20A + 0.15U + 0.15L + 0.20T and rounded to three decimal places. Thus, density contributed 30% of the score, relative amplitude 20%, mean duration 15%, late long-event fraction 15%, and late temporal bias 20%. These components were not assumed to represent statistically independent or physiologically equivalent dimensions; they were retained as the fixed components of the original exploratory score.
The rounded composite score was converted into three score-defined digital acoustic phenotypes using the original fixed thresholds: High-score for S≥0.60S \geq 0.60, Intermediate-score for 0.42≤S<0.600.42 \leq S < 0.60, and Low-score for S<0.42S < 0.42. The four recordings without qualifying merged acoustic complexes did not receive a score or phenotype assignment. These categories are deterministic discretizations of a continuous acoustic score rather than data-driven clusters or independently discovered latent classes.
The term digital acoustic phenotype is therefore used here to denote a score-defined computational representation of acoustic expression within a single overnight recording. It does not denote a clinical obstructive sleep apnea phenotype, an inferred sleep stage, or a validated physiological subtype.
Score values and phenotype assignments were reconstructed without refitting the original component weights or category thresholds. Reproduction was evaluated against the frozen score representation and original category assignments. Subsequent comparisons with annotation-derived respiratory measures and temporal compaction were treated as evaluations external to the score rather than as inputs to phenotype construction.
2.4. Annotation-Derived Respiratory Measures
Respiratory-polygraphy annotations distributed with APSAA were analyzed separately from the acoustic-event representation. These annotations were not used to detect acoustic complexes, construct the acoustic score, define phenotype thresholds, or assign recordings to digital acoustic phenotypes. They were used only as recording-level external respiratory measures for comparison with the acoustic representation.
For each recording, annotated respiratory events were counted when their annotation labels corresponded to Obstructive Apnea, Central Apnea, Hypopnea, or Mixed Apnea. Other annotation types, including desaturation, snoring, and annotations indicating absence of respiratory effort, were not included in this count. An annotation-derived respiratory event rate was calculated as the number of included respiratory events divided by the complete duration of the corresponding distributed audio recording, expressed as events per hour.
This denominator differs from that used in the original exploratory calculation, in which event count was divided by the timestamp of the final annotation of any type. The complete distributed-recording duration was used in the present analysis because it provides a consistent recording-level denominator independent of annotation timing. Consequently, annotation-derived respiratory event rates reported here may differ slightly from values produced by the original exploratory calculation.
The annotation-derived respiratory event rate was not interpreted as the apnea-hypopnea index (AHI). APSAA respiratory polygraphy does not include electroencephalography, electro-oculography, or electromyography and therefore does not provide electroencephalographically confirmed total sleep time. The measure consequently represents the frequency of selected annotated respiratory events per hour of distributed recording rather than events per hour of sleep.
Peripheral oxygen saturation (SpO₂) was obtained from the polygraphy data distributed with APSAA. For recording-level SpO₂ summaries, samples outside the physiologically bounded interval from 50% to 100% were excluded. This filtering rule was specified in the submitted Methods and was retained because it reproduces the reported recording-level SpO₂ values. The recovered exploratory function did not itself implement this filter; therefore, its exact historical implementation provenance could not be independently established.
After filtering, mean SpO₂ was calculated from the remaining samples for each recording. The excluded samples represented a small fraction of the available SpO₂ series in the audited anchor recordings and included values as low as 0%, consistent with signal dropout or invalid sensor values rather than interpretable oxygen saturation measurements.
Annotation-derived respiratory event rate and mean SpO₂ were retained as external recording-level descriptors. Neither measure contributed to calculation of the acoustic score or phenotype assignment.
2.5. Temporal Compaction
Temporal compaction was examined as a recording-level property of the temporal organization of detected acoustic complexes. This analysis was performed separately from construction of the acoustic score and did not contribute to phenotype assignment.
For each recording, merged acoustic complexes were ordered by onset time. The inter-complex gap was defined as the interval between the onset of the current complex and the offset of the immediately preceding complex. Because constituent episodes separated by less than 3 s were merged, gaps between consecutive merged complexes were ≥3 s by construction.
In the primary analysis, gaps greater than 0 s and less than 120 s were considered eligible. Temporal compaction was defined as the proportion of these eligible inter-complex gaps that were shorter than 10 s:
Temporal compaction = number of eligible gaps <10 s / number of eligible gaps
Higher values therefore indicate that a greater proportion of consecutive detected acoustic complexes occurred after short intervening intervals, whereas lower values indicate a more temporally dispersed sequence within the eligible interval range.
All 28 recordings contributing to the temporal-compaction analysis had at least one eligible inter-complex gap. The number of eligible gaps varied substantially across recordings, from 1 to 771. Temporal compaction was calculated directly as the recording-level proportion of eligible gaps shorter than 10 s. Each recording therefore contributed one compaction value to subsequent participant-level analyses regardless of the number of eligible gaps underlying that proportion. The sparsest recording, 20220130C, contained one eligible gap, which was not shorter than 10 s, yielding a compaction value of 0.000.
The 120-s upper bound was retained from the exploratory inter-event analysis as an operational eligibility rule rather than as a physiological definition of recovery. Its effect was strongly dependent on acoustic-complex quantity: the fraction of positive inter-complex gaps excluded by the 120-s cutoff was strongly inversely associated with the number of detected complexes (Spearman ρ = −0.925), ranging from 75% in recording 20220130C to 0.6% in recording 20211104C. Because this rule disproportionately removed long intervals from sparse recordings, the primary analysis was complemented by a sensitivity analysis including all positive inter-complex gaps.
Temporal compaction was interpreted only as a property of the temporal arrangement of the detected acoustic-complex sequence. It was not interpreted as a measure of physiological recovery time, respiratory recovery capacity, resilience, or clinical severity.
2.6. Adjustment for Acoustic-Complex Quantity
Because the observed temporal-compaction proportion can vary with the number of detected acoustic complexes, compaction was adjusted at the recording level for acoustic-complex quantity. For each recording, the observed compaction value was defined as the number of short eligible gaps (<10 s) divided by the total number of eligible gaps (<120 s).
Expected temporal compaction was estimated using a leave-one-participant-out (LOPO) binomial generalized linear model. In each iteration, one recording was excluded and the model was fitted to the remaining recordings, with the recording-level short-gap fraction as the response, log(1 + n_complexes) as the predictor, and the number of eligible gaps as the frequency weight. The fitted model was then used to estimate the expected temporal-compaction proportion for the excluded recording. This weighting allowed recordings with different numbers of eligible gaps to contribute according to the number of binomial observations underlying their recording-level proportion.
For each recording, load-adjusted temporal compaction was expressed as the residual:
residual = observed temporal compaction − LOPO-expected temporal compaction
Positive residuals indicate a greater proportion of short gaps than expected for the number of detected acoustic complexes, whereas negative residuals indicate a lower proportion than expected for that acoustic-complex quantity. The LOPO procedure was deterministic and involved no random sampling or random seed.
As a sensitivity analysis, temporal compaction was additionally adjusted for acoustic-complex burden rather than acoustic-complex count. Acoustic-complex burden was defined as the proportion of distributed recording duration occupied by merged acoustic complexes. Before transformation, burden proportions were clipped to the interval [10⁻⁶, 1 − 10⁻⁶] and then transformed using the logit function. The LOPO binomial model otherwise retained the same specification as the count-adjusted model, including the recording-level short-gap proportion as the response and the number of eligible gaps as the frequency weight, with logit-transformed acoustic-complex burden replacing log(1 + n_complexes) as the predictor. For each held-out recording, the burden-adjusted residual was calculated as the observed temporal-compaction proportion minus its LOPO-predicted proportion.
The binomial models assumed binomial conditional variance. No quasi-binomial dispersion parameter or beta-binomial model was estimated. Consequently, potential overdispersion and the substantial variation in the number of eligible gaps contributing to individual recording-level proportions were not explicitly modeled beyond the binomial frequency weights. These features were considered when interpreting the adjusted residuals and in participant-level sensitivity analyses.
The adjustment was intended to reduce the dependence of temporal compaction on acoustic-event load and was not interpreted as establishing statistical independence from event quantity or burden.
2.7. Sensitivity and Influence Analyses
Sensitivity to the 120-s eligibility cutoff was examined by recalculating temporal compaction using all positive inter-complex gaps, without an upper-duration restriction, while retaining the <10-s definition of a short gap. The corresponding recording-level compaction proportions were adjusted for acoustic-complex quantity using the same leave-one-participant-out procedure described in Section 2.6. This analysis evaluated the extent to which the primary findings depended on the operational exclusion of gaps ≥120 s.
Sensitivity to the definition of acoustic load was evaluated using the alternative acoustic-complex-burden adjustment described in Section 2.6. This analysis was used to determine whether the group comparisons were materially altered when acoustic load was represented by the proportion of recording duration occupied by acoustic complexes rather than by acoustic-complex count.
Because the Low-score phenotype contained only four recordings, the influence of individual Low-score recordings was examined by repeating the relevant group comparisons after omitting each Low-score recording in turn. Each omission left three Low-score recordings. With a three-recording group, the group median is the value of the middle-ranked observation; consequently, across the four leave-one-out omissions, the Low-score median could assume only two distinct values, and the four omissions yielded only two distinct group-median configurations rather than four independent replications. These analyses were therefore interpreted as influence and directional-consistency checks, not as independent evidence of replication or robustness.
Finally, the association between temporal compaction and the late temporal-bias component (T) of the acoustic score was examined as an audit of potential construct overlap. Temporal compaction was not an input to the acoustic score, but both quantities describe aspects of temporal distribution within the recording. Their recording-level association was therefore examined to determine whether temporal compaction substantially recapitulated the temporal-bias component of the score. This analysis was treated as a construct-dependence audit rather than as an independent validation test.
2.8. Statistical Analysis
Continuous recording-level variables were summarized using medians and interquartile ranges unless otherwise specified. Comparisons across the three score-defined acoustic phenotypes (High-score, Intermediate-score, and Low-score) were performed using the Kruskal–Wallis test. Effect size for the three-group comparison was quantified using epsilon-squared (ε²). Pairwise phenotype comparisons were performed using two-sided Mann–Whitney U tests, with Holm adjustment applied jointly to the three pairwise comparisons within each analysis.
Associations between continuous recording-level variables were assessed using Spearman rank correlation. This included analyses used to characterize dependence on acoustic-complex quantity and the construct-dependence audit between temporal compaction and the late temporal-bias component (T) of the acoustic score.
The primary phenotype comparison used temporal-compaction residuals obtained from the count-adjusted LOPO binomial model described in Section 2.6. Sensitivity analyses using all positive inter-complex gaps and acoustic-complex-burden adjustment were analyzed using the same recording-level group-comparison framework. Leave-one-Low-score-out analyses were treated as influence analyses rather than as independent hypothesis tests or replications. The omnibus and pairwise statistics were recalculated after each omission to characterize the sensitivity of the observed contrasts to individual Low-score recordings; pairwise p values within each omission were Holm-adjusted regardless of whether the corresponding Kruskal–Wallis p value crossed the 0.05 threshold. Results across omissions were not combined or interpreted as independent replications.
All analyses were conducted at the recording level, with one overnight recording per participant. Individual inter-complex gaps were not treated as independent observations in phenotype-level inferential tests. Statistical tests were two-sided, and p < 0.05 was considered statistically significant. Because the study was exploratory and involved a small sample, p values were interpreted together with effect sizes, direction and consistency of effects, and sensitivity analyses rather than as standalone evidence.
No a priori sample-size or power calculation was performed. This was a secondary analysis of an existing dataset, and all APSAA recordings satisfying the analysis requirements were included.
Signal processing and statistical analyses were performed in Python using NumPy, SciPy, pandas, librosa, and statsmodels.
3. Results
3.1. Processing Yield and Score-Defined Acoustic Phenotypes
All 32 APSAA recordings were processed using the fixed acoustic-segmentation procedure. Twenty-eight recordings (87.5%) yielded at least one qualifying merged acoustic complex and were therefore eligible for score construction and phenotype assignment. Four recordings yielded no qualifying complexes and were not assigned an acoustic score or score-defined phenotype.
Across the 28 recordings with qualifying acoustic complexes, the canonical merged representation contained 8,512 acoustic complexes. Application of the deterministic acoustic-score rule classified 12 recordings as High-score, 12 as Intermediate-score, and four as Low-score.
The reconstructed score distribution showed a narrow separation around the 0.60 High-score threshold. The lowest reconstructed High-score was 0.602, whereas the highest reconstructed Intermediate-score was 0.598, corresponding to a gap of 0.004. This separation was of the same order as the discrepancy between the historical and reconstructed scores: the largest absolute reconstruction discrepancy was 0.00462. The High-score and Intermediate-score categories should therefore be interpreted as heuristic discretizations of a continuous score rather than as empirically separated classes.
The internal structure of the score reinforced this interpretation. The density component reached its normalization cap of 1.0 in 21 of the 28 scored recordings (75.0%), so the most heavily weighted component of the composite score was constant across three-quarters of the sample and contributed no between-recording discrimination in those recordings. Effective differentiation among recordings therefore rested on the remaining four components.
Zero acoustic yield did not indicate absence of independently annotated respiratory activity. Annotation-derived respiratory event rates in the four zero-yield recordings ranged from 15.74 to 71.07 events per distributed recording hour, against a median of 34.90 events/h among the 28 recordings that did yield complexes. Their sample-based T90 values ranged from 1.48% to 55.48%, with two below and two above the median of 17.23% observed in the yielding recordings, and all four recordings had valid oximetry in more than 99.7% of samples. Recording 20220123A, which contained the second-highest annotated respiratory event count in the dataset (537 events, 71.07 events/h) with a mean SpO₂ of 86.63% and a T90 of 55.48%, yielded no qualifying acoustic complexes. Two recordings bound this observation in the opposite direction: 20220125B, with the highest annotated count in the dataset (636 events), yielded 781 complexes, and 20220124A, with the highest T90 (78.00%), yielded 238. Detection failure under the fixed segmentation rule was therefore neither equivalent to absence of annotated respiratory events nor monotonically related to respiratory or oximetric burden, and the four zero-yield recordings represent failures of the acoustic rule rather than recordings without respiratory-event burden.
Table 1.
Recording-level acoustic features and score-defined digital acoustic phenotypes (n = 28).
| Recording | Phenotype | Acousticscore | Complexes(n) | Complexesper hour | Burden(%) | Annotation-derivedrate (events/h) | Temporalcompaction | Load-adjustedresidual |
| 20211001B | High | 0.684 | 382 | 53.60 | 30.78 | 30.45 | 0.341 | +0.097 |
| 20220206C | High | 0.666 | 288 | 40.14 | 15.13 | 29.96 | 0.243 | −0.007 |
| 20220124A | High | 0.666 | 238 | 32.51 | 11.67 | 73.08 | 0.150 | −0.106 |
| 20211025C | High | 0.649 | 424 | 58.98 | 32.11 | 56.62 | 0.209 | −0.041 |
| 20220121A | High | 0.638 | 220 | 28.92 | 7.86 | 62.71 | 0.099 | −0.162 |
| 20211220D | High | 0.626 | 274 | 35.81 | 12.52 | 25.61 | 0.202 | −0.051 |
| 20211019A | High | 0.618 | 415 | 53.05 | 35.77 | 43.21 | 0.285 | +0.038 |
| 20211104C | High | 0.611 | 638 | 80.69 | 58.94 | 71.33 | 0.175 | −0.088 |
| 20210928B | High | 0.610 | 264 | 33.50 | 13.60 | 15.74 | 0.282 | +0.033 |
| 20211022B | High | 0.607 | 316 | 41.31 | 27.07 | 15.95 | 0.376 | +0.132 |
| 20210926B | High | 0.606 | 208 | 30.85 | 11.58 | 19.28 | 0.271 | +0.020 |
| 20220105D | High | 0.602 | 466 | 58.92 | 44.92 | 54.24 | 0.264 | +0.018 |
| 20211011D | Intermediate | 0.598 | 469 | 57.68 | 28.09 | 51.04 | 0.184 | −0.068 |
| 20211212D | Intermediate | 0.585 | 494 | 61.37 | 26.13 | 42.36 | 0.183 | −0.070 |
| 20211107B | Intermediate | 0.572 | 379 | 53.30 | 29.05 | 36.43 | 0.255 | +0.007 |
| 20211107C | Intermediate | 0.563 | 309 | 50.57 | 15.28 | 43.21 | 0.236 | −0.015 |
| 20211111C | Intermediate | 0.554 | 326 | 49.80 | 23.18 | 46.59 | 0.405 | +0.163 |
| 20211026B | Intermediate | 0.535 | 339 | 47.73 | 23.00 | 21.12 | 0.323 | +0.077 |
| 20220125B | Intermediate | 0.532 | 781 | 98.49 | 43.26 | 80.21 | 0.249 | +0.006 |
| 20211004D | Intermediate | 0.503 | 309 | 41.16 | 20.72 | 5.06 | 0.358 | +0.113 |
| 20211005C | Intermediate | 0.491 | 221 | 31.11 | 11.50 | 25.06 | 0.194 | −0.060 |
| 20210919B | Intermediate | 0.477 | 194 | 26.15 | 7.25 | 4.31 | 0.176 | −0.080 |
| 20211022A | Intermediate | 0.474 | 242 | 31.61 | 10.60 | 18.28 | 0.270 | +0.020 |
| 20220211B | Intermediate | 0.432 | 129 | 17.62 | 6.61 | 18.98 | 0.367 | +0.122 |
| 20210923A | Low | 0.410 | 60 | 8.04 | 1.69 | 37.77 | 0.129 | −0.138 |
| 20220116B | Low | 0.345 | 78 | 11.54 | 4.80 | 66.46 | 0.186 | −0.076 |
| 20211109B | Low | 0.268 | 44 | 6.20 | 1.87 | 16.48 | 0.053 | −0.234 |
| 20220130C | Low | 0.145 | 5 | 0.72 | 0.16 | 33.38 | 0.000 | −0.273 |
Recordings are ordered by reconstructed acoustic score. Phenotype assignment used the fixed thresholds of the deterministic score: High, S ≥ 0.60; Intermediate, 0.42 ≤ S < 0.60; Low, S < 0.42. These ranges are discretizations of a continuous score and do not denote clinical phenotypes, sleep stages, or validated physiological subtypes; the two recordings adjacent to the 0.60 threshold differ by 0.004.
Acoustic complexes are merged reduced-amplitude acoustic episodes as defined in the Methods. Complexes per hour and burden use the complete distributed recording duration as denominator. Temporal compaction is the proportion of eligible inter-complex gaps (>0 s and <120 s) shorter than 10 s. The load-adjusted residual is observed compaction minus the leave-one-participant-out expected value from a binomial model on log(1 + complex count); negative values indicate fewer short gaps than expected for that acoustic-complex quantity.
The annotation-derived respiratory event rate is the number of manually reviewed obstructive, central and mixed apneas plus hypopneas divided by the distributed recording duration in hours. It is not an apnea–hypopnea index, because the denominator is recording time rather than electroencephalographically confirmed total sleep time.
Recording 20220130C contributed a single eligible inter-complex interval, so its compaction value of 0.000 derives from one observation. Four further recordings (20210919A, 20211214D, 20220116A, 20220123A) yielded no qualifying complexes and are not listed.
3.2. Internal Structure of the Reconstructed Acoustic Score
The acoustic score was a weighted composite of five prespecified components representing acoustic-complex density (D), relative amplitude (A), duration (U), prolonged late complexes (L), and late temporal bias (T). Inspection of the reconstructed components showed that the density component reached its upper bound in 22 of the 28 scored recordings (78.6%), reducing its ability to differentiate recordings across much of the analytic sample.
The distinction between the High-score and Intermediate-score ranges occurred primarily within components already incorporated into the score rather than through an independently measured characteristic. In particular, the late temporal-bias component (T) showed the clearest separation between these two score ranges. Because T contributes directly to the composite score, this difference describes the internal structure of the scoring rule and should not be interpreted as independent validation of the High-score and Intermediate-score categories.
Consistent with their narrow separation around the 0.60 threshold, the High-score and Intermediate-score recordings were therefore treated as adjacent heuristic ranges of a continuous acoustic score rather than as empirically established discrete classes.
Figure 1.
Distribution and internal structure of the reconstructed acoustic score (n = 28). (A) Recordings ordered by reconstructed acoustic score. Dashed and dotted lines mark the fixed thresholds of 0.42 and 0.60 that define the three score-defined ranges. The distribution is continuous across both thresholds; the two recordings adjacent to the 0.60 boundary are separated by 0.004 (20211011D, S = 0.598; 20220105D, S = 0.602), which is smaller than the maximum absolute difference between reconstructed and originally reported scores across the 28 recordings. (B) The five normalized components of the deterministic score, by score-defined range. Points are individual recordings; horizontal bars are group medians. The density component reaches its cap of 1.0 in 21 of 28 recordings, and High-score and Intermediate-score recordings are indistinguishable in this component as well as in relative amplitude and duration. The two ranges separate visibly only in late temporal bias. Because the score assigns range membership, differences between ranges in these components are constitutive rather than independent evidence of group structure.
Figure 1.
Distribution and internal structure of the reconstructed acoustic score (n = 28). (A) Recordings ordered by reconstructed acoustic score. Dashed and dotted lines mark the fixed thresholds of 0.42 and 0.60 that define the three score-defined ranges. The distribution is continuous across both thresholds; the two recordings adjacent to the 0.60 boundary are separated by 0.004 (20211011D, S = 0.598; 20220105D, S = 0.602), which is smaller than the maximum absolute difference between reconstructed and originally reported scores across the 28 recordings. (B) The five normalized components of the deterministic score, by score-defined range. Points are individual recordings; horizontal bars are group medians. The density component reaches its cap of 1.0 in 21 of 28 recordings, and High-score and Intermediate-score recordings are indistinguishable in this component as well as in relative amplitude and duration. The two ranges separate visibly only in late temporal bias. Because the score assigns range membership, differences between ranges in these components are constitutive rather than independent evidence of group structure.

3.3. Relationship Between Acoustic Representation and Annotation-Derived Respiratory-Event Rate
Acoustic yield was moderately associated with the annotation-derived respiratory-event rate at the recording level (Spearman ρ = 0.435, p = 0.021 for acoustic complexes per hour). The composite acoustic score showed a weaker and non-significant association (ρ = 0.283, p = 0.145), consistent with four of its five components describing properties other than the amount of detected acoustic activity. The acoustic representation therefore carried information partly, but not wholly, shared with the annotated respiratory-event rate.
This partial association was not preserved by the score-defined discretization. Annotation-derived event rates did not differ across the three score-defined phenotypes (Kruskal–Wallis H = 0.815, p = 0.665, ε² = 0.000). Median rates were 36.83 events/h (Q1–Q3, 24.03–58.14) in the High-score group, 30.74 (18.81–44.05) in the Intermediate-score group, and 35.58 (29.16–44.94) in the Low-score group, with widely overlapping interquartile ranges and no monotonic ordering across ranges.
The dissociation was most evident at the level of individual recordings. Recording 20211001B had an annotation-derived respiratory-event rate of 30.45 events/h, derived from 217 annotated respiratory events, and yielded 382 qualifying acoustic complexes (53.60 complexes/h) with an acoustic score of 0.685. Recording 20220130C had a slightly higher annotated rate of 33.38 events/h, derived from a larger number of annotated events (n = 232), yet yielded only five qualifying acoustic complexes (0.72 complexes/h) with an acoustic score of 0.145. The approximately 74-fold difference in acoustic-complex rate between these two recordings therefore occurred alongside comparable, and in fact marginally higher, annotated respiratory-event burden in the recording with the lower acoustic yield.
Divergence between the two orderings was not confined to this pair. Recording 20220124A, with the second-highest annotation-derived rate in the cohort (73.08 events/h), was assigned to the High-score range, whereas recording 20211004D, with the lowest annotated rate (5.06 events/h), was assigned to the Intermediate-score range; the two adjacent score ranges therefore spanned a fourteen-fold difference in annotated event rate.
Accordingly, the acoustic score and its score-defined phenotypes should be interpreted as descriptors of the detected acoustic representation rather than as alternative estimates of respiratory-event frequency. Low acoustic yield in particular cannot be read as low annotated respiratory-event burden. The present analysis does not establish independence between acoustic organization and annotated respiratory-event burden — a moderate recording-level association was observed — but shows that the two orderings were not equivalent and that the score-defined grouping did not recover the annotated ordering.
Figure 2.
Acoustic yield and annotation-derived respiratory-event rate across the 28 recordings yielding qualifying acoustic complexes. Each point is one recording, coloured by score-defined phenotype. The annotation-derived rate is the number of manually reviewed obstructive, central and mixed apneas plus hypopneas divided by the distributed recording duration; it is not an apnea–hypopnea index. Acoustic complexes per hour were moderately associated with the annotated rate (Spearman ρ = 0.435, p = 0.021), whereas the composite acoustic score showed a weaker, non-significant association (ρ = 0.283, p = 0.145) and the annotated rate did not differ across score-defined phenotypes (Kruskal–Wallis H = 0.815, p = 0.665, ε² = 0.000). Four recordings discussed in the text are labelled with their annotated rate and acoustic-complex rate: the anchor pair 20211001B and 20220130C, which have comparable annotated rates but differ approximately 74-fold in acoustic yield, and 20220124A and 20211004D, which occupy adjacent score ranges while spanning a fourteen-fold difference in annotated rate.
Figure 2.
Acoustic yield and annotation-derived respiratory-event rate across the 28 recordings yielding qualifying acoustic complexes. Each point is one recording, coloured by score-defined phenotype. The annotation-derived rate is the number of manually reviewed obstructive, central and mixed apneas plus hypopneas divided by the distributed recording duration; it is not an apnea–hypopnea index. Acoustic complexes per hour were moderately associated with the annotated rate (Spearman ρ = 0.435, p = 0.021), whereas the composite acoustic score showed a weaker, non-significant association (ρ = 0.283, p = 0.145) and the annotated rate did not differ across score-defined phenotypes (Kruskal–Wallis H = 0.815, p = 0.665, ε² = 0.000). Four recordings discussed in the text are labelled with their annotated rate and acoustic-complex rate: the anchor pair 20211001B and 20220130C, which have comparable annotated rates but differ approximately 74-fold in acoustic yield, and 20220124A and 20211004D, which occupy adjacent score ranges while spanning a fourteen-fold difference in annotated rate.

3.4. Temporal Compaction Across Score-Defined Acoustic Phenotypes
Temporal compaction differed across the three score-defined acoustic phenotypes after adjustment for acoustic-complex quantity using the leave-one-participant-out (LOPO) binomial model (Kruskal–Wallis H = 8.192, p = 0.0166, ε² = 0.248). Median raw compaction was 0.254 (Q1–Q3, 0.195–0.283) in the High-score group, 0.252 (0.192–0.332) in the Intermediate-score group, and 0.091 (0.039–0.143) in the Low-score group. After quantity adjustment, median residuals were +0.005 (−0.060 to +0.035), +0.006 (−0.062 to +0.086), and −0.186 (−0.244 to −0.123), respectively.
The magnitude of this difference was particularly evident in the Low-score recordings. Their median observed compaction was 0.091, meaning that approximately 9% of eligible inter-complex intervals were shorter than 10 s, whereas the LOPO model predicted a median compaction of 0.270 for these recordings based on acoustic-complex quantity. This expected value was close to the median raw compaction actually observed in the High-score (0.254) and Intermediate-score (0.252) groups. Thus, despite their lower acoustic-complex counts, the Low-score recordings were predicted to exhibit substantially more short-interval compaction than was observed.
Pairwise comparisons confirmed this pattern. Count-adjusted compaction was lower in the Low-score group than in the Intermediate-score group (U = 47 of 48 possible rank-order units; rank-biserial correlation = 0.958 for Intermediate-score relative to Low-score; median pairwise difference = 0.200; Holm-adjusted p = 0.0066) and the High-score group (U = 44 of 48; rank-biserial correlation = 0.833 for High-score relative to Low-score; median pairwise difference = 0.172; Holm-adjusted p = 0.0264). These large rank-biserial values reflect near-complete rank separation in comparisons involving only four Low-score recordings and should therefore be interpreted together with the markedly unequal group sizes. In contrast, High-score and Intermediate-score recordings showed little separation (U = 64 of 144; rank-biserial correlation = −0.111; median pairwise difference = −0.026; Holm-adjusted p = 0.665).
The primary analysis therefore did not show a graded High–Intermediate–Low ordering of temporal compaction. Rather, the group-level effect was concentrated in the Low-score recordings, while High-score and Intermediate-score recordings had nearly identical median count-adjusted residuals. The robustness of the Low-score contrasts to alternative compaction definitions, burden adjustment, and the influence of individual Low-score recordings is examined in Section 3.5.
Table 2.
Acoustic and temporal features by score-defined digital acoustic phenotype (n = 28).
| Feature | High-score (n = 12) | Intermediate-score (n = 12) | Low-score (n = 4) |
| Acoustic representation | |||
| Acoustic score ᵃ | 0.631 | 0.525 | 0.292 |
| Acoustic complexes, n | 302.0 | 317.5 | 52.0 |
| Complexes per recording hour ᵃ | 45.6 | 47.2 | 6.6 |
| Acoustic-complex burden, % | 21.10 | 21.86 | 1.78 |
| Complex duration, s | 10.88 | 10.99 | 9.64 |
| Relative amplitude ratio ᵃ | 0.329 | 0.336 | 0.440 |
| Temporal organization | |||
| Inter-complex gap, s | 19.46 | 21.53 | 67.79 |
| Temporal compaction | 0.254 (0.195–0.283) | 0.252 (0.192–0.332) | 0.091 (0.039–0.143) |
| Load-adjusted compaction residual | +0.005 (−0.060 to +0.035) | +0.006 (−0.062 to +0.086) | −0.186 (−0.244 to −0.123) |
| External respiratory measures | |||
| Annotation-derived respiratory event rate, /h ᵃ | 41.5 | 32.7 | 38.5 |
| Recording | |||
| Recording duration, h | 7.63 | 7.37 | 7.02 |
Values are medians (interquartile range shown in parentheses where reported). Recordings were assigned to score-defined ranges using the fixed thresholds of the deterministic acoustic score: High-score, S ≥ 0.60; Intermediate-score, 0.42 ≤ S < 0.60; Low-score, S < 0.42. These ranges are discretizations of a continuous score and do not denote clinical phenotypes, sleep stages, or validated physiological subtypes.
Acoustic complexes are merged reduced-amplitude acoustic episodes as defined in the Methods. Complexes per recording hour and acoustic-complex burden use the complete distributed recording duration as denominator. Relative amplitude is the mean within-complex smoothed RMS amplitude divided by the participant-specific 75th-percentile reference; because a complex is a below-threshold interval, a higher ratio indicates a shallower amplitude reduction. Temporal compaction is the proportion of eligible inter-complex gaps (>0 s and <120 s) shorter than 10 s. The load-adjusted residual is observed compaction minus the leave-one-participant-out expected value from a binomial model on log(1 + complex count); negative values indicate fewer short gaps than expected for that acoustic-complex quantity.
The annotation-derived respiratory event rate is the number of manually reviewed obstructive, central and mixed apneas plus hypopneas divided by the distributed recording duration in hours. It is not an apnea–hypopnea index, because the denominator is recording time rather than electroencephalographically confirmed total sleep time, and it should not be read as an independently measured clinical severity value.
Because acoustic-complex density contributes 0.30 of the composite score, differences between score-defined ranges in complex count, rate, or burden are constitutive of the classification and are not independent validation of it.
ᵃ Reported as the group mean; requires recomputation as median (interquartile range) for consistency with the Methods.
Four recordings (20210919A, 20211214D, 20220116A, 20220123A) yielded no qualifying acoustic complexes under the fixed segmentation rule and were not assigned a score or a score-defined range.
Figure 3.
Temporal compaction across score-defined acoustic phenotypes. (A) Raw temporal compaction for the 28 recordings yielding qualifying acoustic complexes. Temporal compaction was defined as the proportion of eligible inter-complex intervals — positive and shorter than 120 s — that were shorter than 10 s. Points represent individual recordings; horizontal markers indicate medians and vertical bars the interquartile range. Median raw compaction was 0.254 (interquartile range 0.195–0.283) in the High-score group, 0.252 (0.192–0.332) in the Intermediate-score group, and 0.091 (0.039–0.143) in the Low-score group. The Low-score recording with a value of 0.000 (20220130C) contributed a single eligible inter-complex interval. (B) Temporal-compaction residuals after leave-one-participant-out (LOPO) adjustment for acoustic-complex quantity. The dashed line at zero represents the compaction expected from acoustic-complex quantity alone. Median residuals were +0.005 (−0.060 to +0.035), +0.006 (−0.062 to +0.086), and −0.186 (−0.244 to −0.123), respectively (Kruskal–Wallis H = 8.192, p = 0.0166, ε² = 0.248). Low-score recordings differed from Intermediate-score recordings (rank-biserial correlation 0.958; median pairwise difference 0.200; Holm-adjusted p = 0.0066) and from High-score recordings (rank-biserial correlation 0.833; median pairwise difference 0.172; Holm-adjusted p = 0.0264), whereas High-score and Intermediate-score recordings did not differ (rank-biserial correlation −0.111; Holm-adjusted p = 0.665) and are shown without a bracket. The large rank-biserial values for comparisons involving the Low-score group reflect near-complete rank separation between groups of markedly unequal size. For Low-score recordings, median observed compaction was 0.091 against a median LOPO-expected value of 0.270, close to the observed medians of the other two groups. Score-defined ranges are discretizations of a continuous acoustic score and do not denote clinical phenotypes, sleep stages, or validated physiological subtypes.
Figure 3.
Temporal compaction across score-defined acoustic phenotypes. (A) Raw temporal compaction for the 28 recordings yielding qualifying acoustic complexes. Temporal compaction was defined as the proportion of eligible inter-complex intervals — positive and shorter than 120 s — that were shorter than 10 s. Points represent individual recordings; horizontal markers indicate medians and vertical bars the interquartile range. Median raw compaction was 0.254 (interquartile range 0.195–0.283) in the High-score group, 0.252 (0.192–0.332) in the Intermediate-score group, and 0.091 (0.039–0.143) in the Low-score group. The Low-score recording with a value of 0.000 (20220130C) contributed a single eligible inter-complex interval. (B) Temporal-compaction residuals after leave-one-participant-out (LOPO) adjustment for acoustic-complex quantity. The dashed line at zero represents the compaction expected from acoustic-complex quantity alone. Median residuals were +0.005 (−0.060 to +0.035), +0.006 (−0.062 to +0.086), and −0.186 (−0.244 to −0.123), respectively (Kruskal–Wallis H = 8.192, p = 0.0166, ε² = 0.248). Low-score recordings differed from Intermediate-score recordings (rank-biserial correlation 0.958; median pairwise difference 0.200; Holm-adjusted p = 0.0066) and from High-score recordings (rank-biserial correlation 0.833; median pairwise difference 0.172; Holm-adjusted p = 0.0264), whereas High-score and Intermediate-score recordings did not differ (rank-biserial correlation −0.111; Holm-adjusted p = 0.665) and are shown without a bracket. The large rank-biserial values for comparisons involving the Low-score group reflect near-complete rank separation between groups of markedly unequal size. For Low-score recordings, median observed compaction was 0.091 against a median LOPO-expected value of 0.270, close to the observed medians of the other two groups. Score-defined ranges are discretizations of a continuous acoustic score and do not denote clinical phenotypes, sleep stages, or validated physiological subtypes.

3.5. Sensitivity, Influence, and Construct-Dependence Analyses
The primary compaction finding was first tested without the 120-s upper eligibility cutoff, retaining all positive inter-complex intervals while preserving the <10-s definition of a short interval and the same LOPO quantity adjustment. The group difference remained detectable (Kruskal–Wallis H = 8.108, p = 0.0174, ε² = 0.244). Pairwise comparisons again showed lower adjusted compaction in Low-score than Intermediate-score recordings (Holm-adjusted p = 0.0033) and Low-score than High-score recordings (Holm-adjusted p = 0.0396), whereas High-score and Intermediate-score recordings remained similar (Holm-adjusted p = 0.795). Thus, the primary group pattern was not dependent on excluding inter-complex intervals ≥120 s.
Adjustment for acoustic-complex burden rather than complex count produced a similar, although attenuated, group effect (Kruskal–Wallis H = 7.096, p = 0.0288, ε² = 0.204). The attenuation is expected, because acoustic-complex burden incorporates both the number and the duration of detected complexes and therefore removes more between-recording variance than complex count alone. Median burden-adjusted residuals were +0.005 (Q1–Q3, −0.058 to +0.036) in the High-score group, −0.001 (−0.055 to +0.088) in the Intermediate-score group, and −0.147 (−0.195 to −0.092) in the Low-score group; the two higher score ranges therefore had essentially null median residuals under this adjustment as well. Pairwise comparisons again showed lower residuals in Low-score recordings relative to Intermediate-score recordings (rank-biserial correlation = 0.875; median pairwise difference = 0.147; Holm-adjusted p = 0.0231) and High-score recordings (rank-biserial correlation = 0.792; median pairwise difference = 0.152; Holm-adjusted p = 0.0396), while High-score and Intermediate-score recordings did not differ (rank-biserial correlation = −0.111; Holm-adjusted p = 0.665). The Low-score compaction pattern was therefore not specific to adjustment by complex count alone.
Because the Low-score group contained only four recordings, we next repeated the primary count-adjusted analysis four times, each time excluding one Low-score recording. The Low-score median residual remained negative in all four omissions. The Low–Intermediate contrast remained significant after Holm correction in all four analyses (4/4; adjusted p range, 0.0132–0.0264). In contrast, the Low–High comparison remained significant in only one of four omissions: removal of 20220116B yielded Holm-adjusted p = 0.0176, whereas the other three omissions produced adjusted p values of 0.0615, 0.0615, and 0.0967. The omnibus group test remained significant in three of four omissions; its largest p value was 0.0508, after exclusion of 20220130C, the recording with raw compaction of 0.000 and only one eligible inter-complex interval. These leave-one-out results should not be interpreted as four independent replications: each omission retains three of the four Low-score recordings, so the analyses share most of their data. The procedure therefore quantifies influence and directional consistency rather than independent reproducibility. Within this constraint, the Low–Intermediate contrast was substantially more stable than the Low–High contrast.
Finally, we examined whether temporal compaction merely reproduced the late temporal-bias component (T), which contributes directly to the acoustic score. At the recording level, compaction and T were not associated under the primary <120-s eligibility rule (Spearman ρ = −0.096, p = 0.627; Pearson r = 0.020) or when all positive inter-complex intervals were retained (Spearman ρ = −0.035, p = 0.861; Pearson r = 0.030). The concordance of near-zero rank-based and linear correlations under both interval definitions is compatible with an absence of association between the two quantities, although a sample of 28 recordings cannot exclude a weak association. This construct-dependence analysis does not establish compaction as an independent physiological variable; it shows only that the observed compaction pattern is not numerically reducible to the late temporal-bias component of the deterministic score.
Table 3.
Sensitivity, influence, and construct-dependence analyses of temporal compaction (n = 28).
| A. Sensitivity of the group comparison to interval eligibility and load definition | ||||||||||||
| Analysis | H | p | ε² | High vs Int.p (Holm) | High vs Lowp (Holm) | Int. vs Lowp (Holm) | ||||||
| Count-adjusted, gaps <120 s (primary) | 8.192 | 0.0166 | 0.248 | 0.665 | 0.0264 | 0.0066 | ||||||
| Count-adjusted, all positive gaps | 8.108 | 0.0174 | 0.244 | 0.795 | 0.0396 | 0.0033 | ||||||
| Burden-adjusted, gaps <120 s | 7.096 | 0.0288 | 0.204 | 0.665 | 0.0396 | 0.0231 | ||||||
| B. Leave-one-Low-score-out analyses (count-adjusted, gaps <120 s) | ||||||||||||
| Recordingomitted | n | H | p | ε² | Low-score medianresidual | High vs Lowp (Holm) | Int. vs Lowp (Holm) | |||||
| 20210923A | 27 | 6.721 | 0.0347 | 0.197 | −0.250 | 0.0615 | 0.0132 | |||||
| 20211109B | 27 | 6.721 | 0.0347 | 0.197 | −0.166 | 0.0615 | 0.0132 | |||||
| 20220116B | 27 | 7.441 | 0.0242 | 0.227 | −0.246 | 0.0176 | 0.0132 | |||||
| 20220130C | 27 | 5.959 | 0.0508 | 0.165 | −0.140 | 0.0967 | 0.0264 | |||||
| C. Association between temporal compaction and the late temporal-bias component (T) | ||||||||||||
| Eligible-interval definition | Spearman ρ | p | Pearson r | p | ||||||||
| Gaps >0 s and <120 s (primary) | −0.096 | 0.627 | 0.020 | 0.920 | ||||||||
| All positive gaps | −0.035 | 0.861 | 0.030 | 0.878 | ||||||||
Int., Intermediate-score. Group comparisons used the Kruskal–Wallis test with epsilon-squared (ε²) as effect size; pairwise comparisons used two-sided Mann–Whitney U tests with Holm adjustment across the three pairwise contrasts within each analysis.
Panel A. Temporal compaction is the proportion of eligible inter-complex gaps shorter than 10 s. The primary analysis restricts eligibility to gaps longer than 0 s and shorter than 120 s; the second row removes the upper bound. Count adjustment uses a leave-one-participant-out binomial model on log(1 + complex count); burden adjustment substitutes the logit of acoustic-complex burden as the predictor. Attenuation of ε² under burden adjustment is expected, because acoustic-complex burden incorporates both the number and the duration of detected complexes.
Panel B. The primary count-adjusted analysis was repeated four times, each time excluding one of the four Low-score recordings; the leave-one-participant-out adjustment was refitted within each omission. Because each omission retains three of the four Low-score recordings, the four analyses share most of their data and are not independent replications; the panel documents influence and directional consistency. Recording 20220130C had a raw compaction of 0.000 derived from a single eligible inter-complex interval.
Panel C. The late temporal-bias component (T) is one of the five components of the deterministic acoustic score; temporal compaction is not an input to that score. Both rank-based and linear correlations are reported to address monotonic and linear dependence separately. A sample of 28 recordings cannot exclude a weak association.
4. Discussion
4.1. From a Single-Night Acoustic Representation to a Candidate Digital Acoustic Phenotype
The present study examined whether overnight respiratory acoustics could be represented in a form that preserves organization beyond conventional event frequency. A deterministic five-component score was reconstructed without refitting in all 28 recordings that yielded qualifying acoustic complexes, producing score-defined Low-, Intermediate-, and High-score acoustic ranges. These ranges did not reproduce the ordering of annotation-derived respiratory-event rate, and recordings with similar annotated rates could occupy markedly different positions in the acoustic representation. Temporal compaction added a second level of distinction: Low-score recordings exhibited fewer short inter-complex intervals than expected from acoustic-complex quantity, although this effect was substantially more stable relative to the Intermediate-score than to the High-score group.
We use the term digital acoustic phenotype here in a deliberately restricted sense. It denotes a score-defined computational representation of acoustic expression within a single overnight recording. It does not denote a validated clinical phenotype of obstructive sleep apnea, an inferred sleep stage, a physiological endotype, or the longitudinal in-situ digital phenotype described in the broader digital-phenotyping literature, in which the term denotes moment-by-moment quantification of individual state from personal devices in everyday settings [14,15]. The present dataset contains one recording per participant and therefore cannot establish within-person stability, transitions, or state sensitivity.
This distinction is important because the potential value of the representation lies not in replacing the apnea–hypopnea index or other clinical event summaries, but in preserving acoustic structure that those summaries discard [2,3]. Whether the same representation can be acquired repeatedly through personal or consumer microphones, remain stable when the individual is stable, and change meaningfully when physiological state changes is a subsequent longitudinal validation question rather than a result of the present study.
4.2. What the Score-Defined Acoustic Phenotype Captures and Does not Capture
The reconstructed score is a compact description of how the detector represents respiratory acoustic activity across a recording. It combines acoustic-complex density, relative amplitude, duration, prolonged complexes occurring late in the recording, and late temporal bias. None of these components individually establishes a physiological state, and their weighted combination was specified deterministically rather than learned from an external clinical outcome. The three labels therefore mark operational ranges along a continuous score, not three discrete biological classes.
Two features of the reconstruction make this explicit. The score distribution was continuous across both thresholds: the highest Intermediate-score value was 0.598 and the lowest High-score value 0.602, a separation of the same order as the largest discrepancy between the historical and reconstructed scores. And the density component, which carries the greatest weight, reached its normalization cap in 21 of the 28 scored recordings, so the most heavily weighted term was constant across three-quarters of the sample. The High- and Intermediate-score ranges also showed nearly identical median compaction residuals and no detectable difference in the primary comparison. The threshold at S = 0.60 is therefore useful for reproducing the prespecified framework, but should not be read as a discontinuity in the underlying data.
The four recordings with no qualifying acoustic complexes define the outer limit of the representation. Zero acoustic yield did not imply absence of annotated respiratory abnormality. Recording 20220123A contained 537 annotated respiratory events, the second-highest count in the dataset, with an annotation-derived rate of 71.07 events/h and a T90 of 55.48%, yet yielded no qualifying complexes under the fixed detector. The converse also held, and in both dimensions of burden: the recording with the highest annotated event count and the recording with the highest oximetric burden each yielded several hundred acoustic complexes. Detection failure therefore did not track respiratory or oximetric severity in either direction. A zero-yield recording is a detector outcome, not a negative respiratory finding, and the acoustic representation describes only the signal structure that passes the specified amplitude, duration, and merging rules.
4.3. Acoustic Representation Beyond Annotated Event Frequency
The acoustic and annotation-derived measurements interrogate different properties of the same recordings. Respiratory-event annotations enumerate externally classified obstructive, central, mixed, and hypopneic events; the present detector identifies periods satisfying a fixed acoustic amplitude-duration rule and merges nearby detections into complexes. One-to-one correspondence between the two is neither required nor observed.
The two orderings were related but not equivalent. Acoustic yield showed a moderate recording-level association with the annotated rate, yet the composite score was more weakly associated, the score-defined ranges did not differ in annotated event frequency, and divergence occurred at both extremes of the annotated distribution. The anchor pair illustrates the strongest case: recordings 20211001B and 20220130C had annotation-derived rates of 30.45 and 33.38 events/h yet yielded 53.60 and 0.72 acoustic complexes/h, with the lower-yield recording in fact containing slightly more annotated respiratory events than the higher-yield one.
This matters because most previous acoustic work has ultimately been evaluated by its capacity to reproduce apnea labels, AHI, or AHI-derived severity categories [9,10,11,12,13]. The present analysis asks a different question. Event rate compresses an entire night into a frequency; the acoustic representation retains other dimensions of how detected activity occurs and is distributed. This is consistent with the recognized physiological heterogeneity of sleep-disordered breathing, in which similar AHI values can arise from different combinations of upper-airway collapsibility, ventilatory control, arousal threshold, and neuromuscular responsiveness [4,5,6,7,8]. The present data identify none of these mechanisms and the acoustic ranges should not be given mechanistic labels; they support only the narrower proposition that respiratory recordings contain organizational information discarded by frequency summaries.
Because the source study used respiratory polygraphy without electroencephalography, electro-oculography, or electromyography [16,17], the annotation-derived rate used here is not an apnea–hypopnea index based on measured sleep time. The present analyses therefore compare two representations available within this dataset rather than testing an acoustic replacement for polysomnographic diagnosis.
4.4. Temporal Compaction Beyond Acoustic Quantity
Temporal compaction asks a narrower question than event frequency: once qualifying complexes have been detected, how tightly are they organized in time? Because this quantity could depend mechanically on how many complexes were available to generate intervals, the primary analysis adjusted for acoustic-complex count using a leave-one-participant-out model.
The resulting pattern was not a three-level gradient. High- and Intermediate-score recordings were essentially indistinguishable after adjustment, whereas Low-score recordings were displaced toward lower compaction. Median observed compaction in the Low-score group was 0.091, against a model-predicted median of 0.270 — close to the values actually observed in the two higher ranges. The Low-score pattern therefore cannot be attributed simply to having fewer detected complexes.
Several analyses delimit the finding. Removing the 120-s upper interval cutoff retained the overall pattern. Replacing complex count with acoustic-complex burden attenuated but did not eliminate the group effect, the attenuation being expected because burden incorporates both the number and the duration of complexes. Compaction showed neither monotonic nor linear association with the late temporal-bias component of the score under either interval definition, so it is not a numerical restatement of that component.
The influence analysis places a boundary on the result. The Low–Intermediate comparison survived all four leave-one-Low-score-out analyses, whereas the Low–High comparison survived only one, and the omnibus test reached p = 0.0508 after removal of 20220130C. Because these analyses share most of the same Low-score observations, they are influence tests rather than independent replications. The pairwise effect sizes involving the Low-score group were large, reflecting near-complete rank separation between groups of markedly unequal size, and describe separation in this dataset rather than a precise population estimate. The appropriate interpretation is asymmetric: the data support reduced quantity-adjusted compaction in Low-score relative to Intermediate-score recordings more strongly than a stable distinction between Low- and High-score recordings.
4.5. From Recording-Level Representation to Longitudinal Digital Sensing
The audio analyzed here was acquired with a dedicated recorder positioned above the participant's head during supervised home polygraphy [16,17], not with a personal or consumer device. Nothing in the present study establishes that the same representation can be obtained from smartphone microphones, and the study compared no devices, placements, compression formats, or domestic acoustic environments. Previous work has shown that smartphone-recorded breathing sounds can support automated analysis of sleep-disordered breathing [9,10,11], which makes transfer technically plausible; but plausibility is not demonstration, and the representation described here is derived from acoustic information alone, so its transfer to repeated low-burden microphone acquisition remains to be tested prospectively.
Technical reproducibility would need to be established before any longitudinal claim. The relevant questions are whether complex detection and compaction estimates remain stable across repeated acquisitions and robust to differences in hardware, microphone placement, background noise, and audio processing. Only after that would repeated-night recordings be able to address what a single-night design cannot: whether the acoustic representation is a stable within-person characteristic, a variable recording-level state, or both. Night-to-night variability in sleep-disordered breathing is itself substantial [18], so a meaningful within-person change would first have to be distinguished from ordinary variation between nights.
Establishing clinical meaning would be a further step. An individualized deviation from a personal baseline would not by itself indicate deterioration; linking such deviations to independent physiological measurements, symptoms, treatment response, or patient-centered outcomes requires prospective designs beyond what is available here.
4.6. Limitations
Several limitations determine how far these findings generalize. The cohort comprised 32 participants with one overnight recording each, and 28 recordings produced qualifying acoustic complexes. The score-defined ranges were highly unbalanced at the lower end, with only four Low-score recordings; this small group drives the principal inferential distinction and makes its estimates sensitive to individual observations. The marked sex imbalance of the source dataset further restricts generalizability. Repeated-night data are unavailable, so neither within-person stability of the score nor of temporal compaction can be established.
The acoustic detector was inherited rather than optimized against respiratory-event annotations. Its fixed threshold, minimum-duration criterion, and merging rule determine which portions of the signal become qualifying complexes, and four recordings produced zero yield despite documented respiratory-event burden. The 120-s eligibility cutoff and the 10-s short-interval threshold are operational definitions rather than established physiological boundaries; the all-positive-interval sensitivity analysis reduces dependence on the former, but independent datasets are required to determine whether the observed structure generalizes to other recording conditions and detector implementations.
The score has additional structural limitations. Its ranges arise from deterministic thresholds applied to a continuous composite, one component was saturated in most scored recordings, and several components contribute directly to distinctions subsequently described between ranges. Internal differences among score components therefore cannot serve as independent validation of the ranges. Likewise, the absence of detectable association between compaction and late temporal bias addresses numerical redundancy but does not establish a distinct biological mechanism.
Finally, the present reconstruction does not support claims about recovery dynamics, physiological resilience, sleep-stage identity, or downstream oxygen-desaturation timing. What the data support is narrower: a fixed respiratory-acoustic detector generates a reproducible recording-level representation; that representation is not ordered simply by annotation-derived respiratory-event frequency; and a small subset of Low-score recordings shows reduced short-interval temporal compaction that persists under alternative interval and quantity-adjustment definitions, with stability constrained by the small number of Low-score recordings. These findings position temporal acoustic organization as a measurable property worth prospective investigation rather than an established physiological phenotype.
5. Conclusions
This study shows that overnight respiratory audio can be transformed into a reproducible score-defined digital acoustic phenotype. The representation describes how detected acoustic respiratory activity is expressed across a recording through measurable properties of the signal: complex density, relative amplitude, duration, occurrence of prolonged late complexes, and temporal distribution. It was reconstructed deterministically, without refitting, in all recordings that yielded qualifying acoustic complexes.
These phenotypes are score-defined digital representations. They are not sleep stages, clinical phenotypes of obstructive sleep apnea, physiological endotypes, or diagnostic categories, and the ranges they define are discretizations of a continuous score rather than empirically separated classes.
The representation is related to annotation-derived respiratory-event frequency but does not reproduce its ordering. Recordings with comparable annotated rates occupied markedly different positions in the acoustic representation, and quantity-adjusted temporal compaction was lower in the Low-score range, with this contrast more consistently supported relative to the Intermediate-score range than to the High-score range. The framework therefore captures both how much acoustic activity is detected and how that activity is distributed in time.
The technological implication is that overnight audio need not be analysed solely as an attempt to reproduce the apnea–hypopnea index or to classify individual respiratory events. Because the representation derives from acoustic information alone, its transfer to repeated low-burden microphone acquisition is technically plausible, but the present audio was obtained with a dedicated recorder during supervised polygraphy and no device comparison was performed. Establishing technical reproducibility across acquisition conditions is a prerequisite for any longitudinal application.
The present study establishes the recording-level representation, not the clinical meaning of its longitudinal change. Whether within-person deviations from an individual's usual acoustic pattern correspond to relevant changes in sleep, respiratory physiology, treatment response, or health state remains to be determined prospectively.
Funding
No funding was received for this research.
Data Availability Statement
The APSAA dataset (Zenodo DOI: 10.5281/zenodo.14096541) is available upon reasonable request through its repository, subject to the conditions established by the dataset authors. Analysis code and derived feature datasets supporting the present study are available upon reasonable request to the corresponding author.
Ethical Approval
This study analyzed a de-identified dataset (APSAA, Zenodo DOI: 10.5281/zenodo.14096541) obtained under the conditions established by the dataset authors, permitting academic use with appropriate citation. No procedures were performed on human participants by the author. For this type of study formal consent is not required.
AI Assistance Disclosure
The author used artificial intelligence tools to assist with computational workflow development, statistical scripting support, and figure preparation. All scientific decisions, data analysis, interpretation, manuscript writing, and final approval were performed and critically reviewed by the author, who retains full responsibility for the content of this work.
Conflicts of Interest
The author certifies no affiliations with or involvement in any organization or entity with any financial interest (such as honoraria, educational grants, participation in speakers’ bureaus, membership, employment, consultancies, stock ownership, or other equity interest, and expert testimony) or non-financial interest (such as personal or professional relationships, affiliations, knowledge or beliefs) in the subject matter or materials discussed in this manuscript. A patent application covering methods described in this work is currently pending (BR 10 2026 006549 8).
References
- Benjafield, A.V.; Ayas, N.T.; Eastwood, P.R.; et al. Estimation of the global prevalence and burden of obstructive sleep apnoea: a literature-based analysis. Lancet Respir. Med. 2019, 7(8), 687–698. [Google Scholar] [CrossRef] [PubMed]
- Mohammadieh, A.M.; Sutherland, K.; Cistulli, P.A. Moving beyond the AHI. J. Clin. Sleep Med. 2021, 17(11), 2335–2336. [Google Scholar] [CrossRef] [PubMed]
- Azarbarzin, A.; Sands, S.A.; Stone, K.L.; et al. The hypoxic burden of sleep apnoea predicts cardiovascular disease-related mortality: the Osteoporotic Fractures in Men Study and the Sleep Heart Health Study. Eur. Heart J. 2019, 40(14), 1149–1157. [Google Scholar] [CrossRef] [PubMed]
- Zinchuk, A.V.; Gentry, M.J.; Concato, J.; Yaggi, H.K. Phenotypes in obstructive sleep apnea: a definition, examples and evolution of approaches. Sleep Med. Rev. 2017, 35, 113–123. [Google Scholar] [CrossRef] [PubMed]
- Malhotra, A.; Mesarwi, O.; Pepin, J.L.; Owens, R.L. Endotypes and phenotypes in obstructive sleep apnea. Curr. Opin. Pulm. Med. 2020, 26(6), 609–614. [Google Scholar] [CrossRef] [PubMed]
- Eckert, D.J.; White, D.P.; Jordan, A.S.; Malhotra, A.; Wellman, A. Defining phenotypic causes of obstructive sleep apnea: identification of novel therapeutic targets. Am. J. Respir. Crit. Care Med. 2013, 188(8), 996–1004. [Google Scholar] [CrossRef] [PubMed]
- Wellman, A.; Edwards, B.A.; Sands, S.A.; et al. A simplified method for determining phenotypic traits in patients with obstructive sleep apnea. J. Appl. Physiol. 2013, 114(7), 911–922. [Google Scholar] [CrossRef] [PubMed]
- Tolbert, T.M.; Schmickl, C.N.; Gell, L.K.; et al. Research priorities for translating endophenotyping of adult obstructive sleep apnea to the clinic: an official American Thoracic Society research statement. Am. J. Respir. Crit. Care Med. 2025, 211(9), 1562–1583. [Google Scholar] [CrossRef] [PubMed]
- Cho, S.W.; Jung, S.J.; Shin, J.H.; Won, T.B.; Rhee, C.S.; Kim, J.W. Evaluating prediction models of sleep apnea from smartphone-recorded sleep breathing sounds. JAMA Otolaryngol. Head Neck Surg. 2022, 148(6), 515–521. [Google Scholar] [CrossRef] [PubMed]
- Le, V.L.; Kim, D.; Cho, E.; et al. Real-time detection of sleep apnea based on breathing sounds and prediction reinforcement using home noises: algorithm development and validation. J. Med. Internet Res. 2023, 25, e44818. [Google Scholar] [CrossRef] [PubMed]
- Han, S.C.; Kim, D.; Rhee, C.S.; et al. In-home smartphone-based prediction of obstructive sleep apnea in conjunction with level 2 home polysomnography. JAMA Otolaryngol. Head Neck Surg. 2024, 150(1), 22–29. [Google Scholar] [CrossRef] [PubMed]
- Kim, J.; Kim, T.; Lee, D.; Kim, J.W.; Lee, K. Exploiting temporal and nonstationary features in breathing sound analysis for multiple obstructive sleep apnea severity classification. BioMed Eng. Online 2017, 16(1), 6. [Google Scholar] [CrossRef] [PubMed]
- Kim, J.W.; Shin, J.; Lee, K.; Won, T.B.; Rhee, C.S.; Cho, S.W. Prediction of oxygen desaturation by using sound data from a noncontact device: a proof-of-concept study. Laryngoscope 2022, 132(4), 901–905. [Google Scholar] [CrossRef] [PubMed]
- Jain, S.H.; Powers, B.W.; Hawkins, J.B.; Brownstein, J.S. The digital phenotype. Nat. Biotechnol. 2015, 33(5), 462–463. [Google Scholar] [CrossRef] [PubMed]
- Insel, T.R. Digital phenotyping: technology for a new science of behavior. JAMA 2017, 318(13), 1215–1216. [Google Scholar] [CrossRef] [PubMed]
- González Martínez, F.D.; de la Torre Cruz, J.; Carabias Orti, J.J.; et al. Audio-Polygraphy Dataset for Sleep Apnea Analysis (APSAA). Version 2.0. Zenodo 2024. [Google Scholar] [CrossRef]
- Gonzalez-Martinez, F.; De La Torre-Cruz, J.; Carabias-Orti, J.; et al. Polygraph and audio synchronization applied to apnea event analysis based on non-negative matrix factorization. EURASIP J. Audio Speech Music Process 2025, 2025, 24. [Google Scholar] [CrossRef]
- Stöberl, A.S.; Schwarz, E.I.; Haile, S.R.; Turnbull, C.D.; Rossi, V.A.; Stradling, J.R.; Kohler, M. Night-to-night variability of obstructive sleep apnea. J. Sleep Res. 2017, 26(6), 782–788. [Google Scholar] [CrossRef] [PubMed]
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.