Submitted:
18 September 2026
Posted:
21 September 2026
You are already at the latest version
Abstract
General-purpose audio foundation models (FMs) — wav2vec 2.0, HuBERT, WavLM, AST, CLAP — have transformed speech and environmental-sound processing, but their value for underwater acoustic target recognition (UATR) is unknown: UATR operates on radiated vessel noise far outside the pretraining distribution of any of these models, and our companion work has shown that even models trained on underwater audio transfer catastrophically across water areas [2]. We benchmark six public FMs (wav2vec2-base, wav2vec2-large, HuBERT-base, WavLM-base+, AST, CLAP-HTSAT) as frozen feature extractors on all four UATR corpora of our series (Oceanship, QiandaoEar22, VTUAD, ShipsEar; 258,893 segments), under two label views (per-corpus full-label and a shared two-class view) and two regimes: within-corpus linear probing and cross-corpus zero-shot probe transfer (72 ordered pairs). Findings: (1) Within each corpus, linear probes on frozen FM embeddings beat the majority-class baseline on macro-F1 on every model–corpus pair of the full-label view (on raw accuracy, four of six models sit below the corrected VTUAD majority rate — an imbalance artifact we flag); the accuracy-strongest full-label FM on QiandaoEar22 and ShipsEar (CLAP) matches trained log-mel ResNet-18 probes within corpora in accuracy but trails them in macro-F1 on three of four corpora, and trails fine-tuned ResNet-18 (e.g., QiandaoEar22 macro-F1 0.258 vs 0.292 for the fine-tuned reference of [4], §4.1), with environmental-audio pretraining (AST, CLAP) beating speech pretraining in 6 of 8 full-label cells. (2) Model scale does not help: wav2vec2-large never beats wav2vec2-base on both metrics of any corpus — macro-F1 degrades in 3 of 4 corpora (accuracy in 2 of 4). (3) Cross-corpus transfer is where FMs pay off: under the series protocol and against all available trained references, the best frozen-FM probe — selected per direction on the primary metric adopted for this benchmark, macro-F1 (chosen because of the documented class imbalance, §4.1–4.2), as a descriptive oracle upper envelope — beats the trained ResNet-18′s zero-shot macro-F1 on 11 of 12 anchored directions (point estimates against the tested anchors under the fixed probe pipeline; in 8 directions the paired-bootstrap 95% CI excludes zero for all three anchor seeds). On Oceanship→QiandaoEar22, a linear probe on frozen WavLM embeddings reaches macro-F1 0.736, numerically exceeding the available ResNet-18 target-oracle reference (0.551) of our trained-model UDA benchmark [1] under the series protocol; a 5,000-run permutation control (the true score sits at the 96.3rd–97.4th percentile of random readouts) and comparability analyses indicate that this exceptional direction should be interpreted cautiously rather than as a controlled superiority claim (§4.3, §4.5). When the FM is instead chosen by source-validation performance — the deployable selection rule — the transferred probe still beats the trained source_only baseline in 9 of 12 directions under the same fixed probe pipeline. (4) All FM results are deterministic (frozen pipeline, seed 42); we therefore report exact values without seed variance and attach paired-bootstrap intervals to all twelve cross-corpus comparisons, with a recording-cluster-level bootstrap as a dependence check (§4.3, §4.5). Frozen audio-FM embeddings are not a replacement for in-domain training; within the literature we surveyed, no earlier study has reported a four-corpus, full 12-direction frozen-embedding cross-water benchmark under a shared label-mapping protocol, and within the companion series frozen FM probes are the first representation family shown to transfer across water areas without any adaptation.

Keywords:
underwater acoustic target recognition
; audio foundation models
; linear probing
; domain shift
; transfer learning
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.