Submitted:
29 September 2026
Posted:
30 September 2026
You are already at the latest version
Abstract
Training deep neural networks on limited music datasets remains prone to overfitting. While spectrogram-level augmentations such as time masking, frequency masking, and Mixup are widely adopted, the temporal segment extraction strategy—the procedure for sampling fixed-length clips from full-length tracks—has received comparatively little attention. We systematically investigate Random Segment Extraction (RSE), a technique that independently samples a random temporal window at each training epoch. Unlike fixed-start cropping or exhaustive pre-cutting, RSE exposes the model to diverse temporal regions across iterations without modifying the underlying audio signal. We analyze its efficacy through three complementary mechanisms: data diversity amplification, temporal positional invariance, and implicit regularization. The method admits a lightweight implementation requiring only minimal modification to the data loader. Experiments on GTZAN (999 usable tracks, 10 genres) using ResNet-18/34/50/101 under five-fold stratified cross-validation demonstrate that RSE yields consistent accuracy gains of 8.5–12.1 percentage points over the standard fixed-segment baseline (average +9.8 pp). A systematic duration sweep (T = 4–28 s) reveals stable performance across a sevenfold range of input lengths, with peak accuracy reaching 86.6% for ResNet-18 (T = 24 s) and 85.9% for ResNet-50 (T = 28 s). Furthermore, RSE outperforms the best exhaustive pre-cutting variant (T = 6 s, 83.1%) by 2.3 pp while using only one-fifth as many training samples per fold. These results establish epoch-level temporal randomization as a low-cost, signal-preserving augmentation strategy that improves classification performance without altering the input representation.
Keywords:
music genre classification
; GTZAN
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.