Preprint
Article

This version is not peer-reviewed.

Self-Supervised Pretraining for Underwater Acoustic Target Recognition: A Three-Corpus SimCLR Pilot with a Sobering Verdict

Submitted:

11 September 2026

Posted:

14 September 2026

You are already at the latest version

Abstract
Labels are the bottleneck in underwater acoustic target recognition (UATR): hydrophones collect unlabeled audio essentially for free, while verified vessel labels remain scarce. Self-supervised learning (SSL) promises to convert the former into reusable representations, but its value for ship-radiated noise has been quantified only on single, leakage-prone corpora. We report a controlled three-corpus pilot: a ResNet-18 pretrained with SimCLR on 176,481 unlabeled log-mel segments pooled from the train splits of Oceanship, QiandaoEar22, and VTUAD, then evaluated on each corpus under linear probing and full fine-tuning against ImageNet and random initialization (15 corpus–initialization–protocol cells × three seeds: 45 runs). The result is predominantly negative at this budget. Under linear probing the SSL representation is significantly worse than ImageNet features on two of three corpora (VTUAD: −0.083 accuracy, p = 0.017; QiandaoEar22: −0.053, p = 0.012). Fine-tuning largely erases the differences within 30 epochs (paired accuracy gaps ≤ 0.039, mostly non-significant). One unestablished hint survives: QiandaoEar22 macro-F1 +0.0372 (p = 0.310). A diagnostic val–test gap on VTUAD, plus a recording-identity probe that reads recording ID near-perfectly from the SSL embeddings, shows the contrastive objective encodes recording identity, not vessel type. A normalization control rules out the main alternative explanation. We analyze why in-corpus contrastive pretraining under-delivers and derive requirements for the follow-on large-scale study. One limitation bounds all of this: pretraining was a single run (n = 1), so encoder-level variance is unmeasured, and all results are preliminary observations on one specific encoder, not general conclusions about SimCLR on this task.
Keywords: 
;  ;  ;  ;  

1. Introduction

Passive sonar archives grow faster than anyone can label them. Cabled observatories and autonomous recorders accumulate years of hydrophone audio. Yet the public labeled corpora used to train underwater acoustic target recognition (UATR) classifiers (ShipsEar [1], DeepShip [2], Oceanship [3], QiandaoEar22 [4], VTUAD [5]) remain small, site-specific, and taxonomically incompatible. Our companion cross-dataset benchmark [6] showed that supervised classifiers trained on one corpus collapse when transferred to another (0–4% accuracy within a shared label space). Follow-up studies quantified what domain adaptation with [7] and without [8] source data can recover. All of these results point to the same root cause: supervised features overfit the acquisition conditions of their training site. The natural remedy is self-supervised learning (SSL) on the abundant unlabeled audio itself. SSL has transformed representation learning in vision [9,10], speech [11], and general audio [12].
Evidence that SSL helps UATR specifically is, however, thinner than the enthusiasm suggests. Masked-spectrogram reconstruction with a Swin Transformer reaches 80.22% on DeepShip [13]. Contrastive learning on unlabeled single-hydrophone archives produces embeddings that transfer across benchmarks [14]. Linear probing of pretrained audio models enables low-cost ship-noise recognition [15]. All three lines evaluate within-corpus, and two of them [13,15] use splits in which adjacent segments of one recording can appear on both sides of the train/test boundary. In that regime recording-identity cues inflate scores. Whether SSL learns vessel-type structure, as opposed to recording or channel structure, is precisely the question that our earlier cross-dataset results [6] make urgent. It cannot be answered by within-corpus accuracy alone.
This paper is a deliberately small, controlled pilot that asks the question directly. We pretrain a ResNet-18 [16] with SimCLR [9] on the pooled, de-labeled train splits of three acoustically and taxonomically distinct corpora (176,481 segments). We then measure the resulting representation two ways on each corpus: (i) linear probing of the frozen backbone, which isolates the linear separability of the learned features, and (ii) full fine-tuning, which measures the value of the representation as an initialization. Both are compared against the two standard alternatives: ImageNet initialization [17], the de-facto default in UATR including all our previous work [6,7,8], and random initialization. The comparison runs over 15 corpus–initialization–protocol cells and three seeds (45 runs) in total. Because pretraining draws on the same corpora as the downstream tasks, the setup is deliberately favorable to SSL. If in-distribution contrastive pretraining cannot beat ImageNet features here, claims about foundation-model-style pretraining for UATR need qualification, not amplification.
The verdict is sobering. Under linear probing, the SSL representation is significantly worse than ImageNet features on VTUAD (−0.083 accuracy, −0.041 macro-F1) and QiandaoEar22 (−0.053 accuracy), and only marginally better on Oceanship. Under full fine-tuning, all three initializations converge to statistically indistinguishable performance on two corpora, and SSL is significantly worse than ImageNet on the third. A diagnostic val–test gap on VTUAD (validation accuracy 0.46 vs. test 0.27 for the SSL probe) points to the mechanism. Together with a recording-identity probe that reaches 0.997 train accuracy on the SSL embeddings (vs. 0.962 ImageNet, 0.874 random; Section 4.3), it demonstrates that the contrastive objective has learned recording/session identity instead of vessel-type structure. This is the same confound Hummel et al. [15] identified in pretrained audio embeddings. We report these results in full, analyze the mechanisms, and extract concrete design requirements for the large-scale follow-on study (multi-site ONC and SanctSound pretraining) that this pilot was intended to de-risk. Our contributions are:
(1) To our knowledge, the first controlled three-corpus evaluation of in-corpus SSL pretraining for UATR, with a leakage-safe protocol: recording/clip-level frozen splits [6], segment-level evaluation, two downstream protocols, three initializations, three seeds, and paired statistics, all on an identical front-end and backbone so that initialization is the only varied factor.
(2) A quantified negative result. At this pilot budget, in-corpus SimCLR pretraining (176 k segments, 40 epochs) does not produce representations that beat ImageNet features: linear probing is significantly worse on two of three corpora, and fine-tuning largely washes out initialization effects at 30 epochs. This bounds what “just pretrain on more hydrophone data of the same kind” can buy.
(3) A mechanism diagnosis. The contrastive objective under our augmentation family (time/frequency roll, gain jitter, SpecAugment masks) shows the signature of recording-identity coding: a 0.18–0.20 val–test accuracy gap on VTUAD, stable across seeds and absent for ImageNet features, under clip-disjoint splits that make recording identity non-predictive at test time. A recording-identity linear probe demonstrates the coding directly (0.997 train accuracy against a 0.015 chance level). We connect this to the recording-ID dominance documented by [15] and to the channel-dominated feature geometry of [6].
(4) Design requirements for scale-up. The pilot isolates four requirements for the follow-on large-scale pretraining study: cross-site pretraining corpora (not in-corpus), channel-varying augmentations, longer schedules with larger batches, and recording-identity control evaluations. The follow-on study must satisfy them to have a chance of succeeding where this pilot did not.

3. Materials and Methods

3.1. Datasets and Splits

We use the three corpora and frozen recording-level splits of [6]: Oceanship (open coastal ocean, ONC cabled-observatory hydrophones, AIS vessel-type labels, 15 classes), QiandaoEar22 (confined freshwater lake, self-contained recorder, 8 target-presence view labels), and VTUAD (rebuilt from ONC Strait of Georgia archives with an automatic AIS labeling pipeline, 12 classes, recording/clip-level frozen splits of 369/80/76 recordings so that no recording crosses partitions; MMSI metadata would permit hull-level isolation, but the frozen splits do not enforce it) [5,6,19]1. All audio is preprocessed identically: 16 kHz, 3 s segments with 1.5 s hop, single-channel 128-bin log-mel spectrograms (80 dB floor). Table 1 gives segment counts and the majority-class chance level of each test split. All three test sets are class-imbalanced, so macro-F1 is the primary metric throughout, with accuracy reported alongside.
The pretraining corpus is the union of the three train splits with labels discarded: 176,481 segments, of which VTUAD contributes 93,592 (53%). The pooled distribution is therefore dominated by the ONC Strait of Georgia recording conditions, a composition bias that resonates with VTUAD also being the worst probing case (Section 4.3). The resonance runs deeper than shared difficulty. The corpus that dominates pretraining is also the corpus where the diagnostic validation–test gap is largest (0.18–0.20, Section 4.3), i.e., where the encoder’s recording-identity coding is most strongly expressed. An encoder over-exposed to one site’s recording conditions has the most recording-specific structure to harvest there. The composition bias thus supports the recording-identity interpretation of Section 4.3 instead of offering an alternative to it. This makes the pilot in-corpus (transductive at the corpus level): the SSL model sees the downstream distribution, only without labels. The choice is deliberate. It is the most favorable setting for SSL short of adding external data, so a negative result here is maximally informative.

3.2. SimCLR Pretraining

The encoder is ResNet-18 with the first convolution replaced by a single-channel 7×7 kernel (identical geometry to our supervised baselines [6,7]), producing 512-d embeddings. A two-layer projection head (512→512→128, ReLU) is appended for the contrastive loss and discarded afterward. Two augmented views of each segment are generated by: uniform time roll (±16 frames), frequency roll (±4 bins), gain jitter (×U(0.8, 1.25)), and SpecAugment masking [20] (2 frequency masks ≤ 24 bins, 2 time masks ≤ 32 frames). Note what this family does not vary: recording device, noise floor, propagation conditions, or session. This exclusion is deliberate. The augmentation family matches the one used by our supervised pipeline [6,7,8], keeping the pilot directly comparable to the supervised installments. Section 5 discusses its consequences. Training uses the NT-Xent loss at temperature τ = 0.2, AdamW [21] (lr 3×10⁻⁴, weight decay 10⁻⁴, cosine schedule), batch size 256, mixed precision, 40 epochs, seed 42. This is roughly an order of magnitude fewer update steps and a 16× smaller batch than the default 1,000-epoch SimCLR reference configuration [9]. Against the shorter 100-epoch SimCLR recipe the update-step count is comparable; the order-of-magnitude gap is specifically relative to the 1,000-epoch reference. We return to this budget gap in Section 5. Inputs are standardized with pooled-corpus statistics. The final training loss is 1.752, far below the chance level ln(2·256−1) ≈ 6.24, confirming that the optimization learned non-trivial view consistency (no collapse). The full run takes ~72 min on an RTX 5070 Ti Laptop GPU.

3.3. Downstream Protocols

Linear probe. The backbone is frozen; embeddings for all splits are extracted in a single forward pass, and a linear head (512→C) is trained on the cached tensors with AdamW (lr 10⁻³, weight decay 10⁻⁴), batch 64, ≤100 epochs, early stopping (patience 15) on validation accuracy. Probing is run for ImageNet and SSL initializations (a random-frozen probe is uninformative and omitted).
Full fine-tuning. All parameters are trained with cross-entropy, AdamW (lr 3×10⁻⁴, weight decay 10⁻⁴, cosine), batch 64, ≤30 epochs, patience 8 on validation accuracy, with SpecAugment (2 frequency masks ≤ 16 bins, 2 time masks ≤ 24 frames) on training inputs (the recipe of our supervised baselines [7]). Fine-tuning is run for random, ImageNet, and SSL initializations. ImageNet initialization uses ResNet18_Weights.IMAGENET1K_V1 with conv1 collapsed by summing over input channels, exactly as in [6,7]. All runs standardize inputs with per-dataset train statistics. Everything downstream (splits, front-end, optimizer, selection rule of best validation accuracy, evaluation of held-out test accuracy and macro-F1) is identical across initializations. Seeds 42/43/44 give n = 3 paired measurements per cell. The design matrix is 3 corpora × 3 inits × 3 seeds (fine-tune) + 3 corpora × 2 inits × 3 seeds (probe) = 45 configurations.

3.4. Statistics

We report mean ± std over seeds with the population std (ddof = 0), matching the convention of [7,8]. Paired comparisons (SSL minus reference on the same seed) use a two-sided paired t-test at df = 2 with the ddof = 1 sample std, as in [8]. With n = 3 seeds, statistical power is limited and only large effects are detectable. We treat p < 0.05 as pilot-level significant and always discuss conclusions together with effect sizes and per-seed patterns.

4. Results

4.1. Linear Probing: SSL Features Lose to ImageNet Features on Two of Three Corpora

Table 2 gives the linear-probing results (per-seed values in Table A2). The SSL representation is the weaker one almost everywhere. On VTUAD the gap is large and significant in both metrics (−0.0829 accuracy, p = 0.017; −0.0410 macro-F1, p = 0.008, consistent across all three seeds). On QiandaoEar22 the accuracy gap is significant (−0.0533, p = 0.012) while the macro-F1 gap is small and not significant (−0.0144, p = 0.414). The SSL probe loses mainly on the majority noise class. Only on Oceanship does the SSL probe edge ahead (+0.0059 accuracy, p = 0.188; +0.0211 macro-F1, p = 0.056, same-sign on all three seeds but short of significance at n = 3).

4.2. Fine-Tuning: Initialization Barely Matters, and Where It Does, SSL Hurts

Table 3 gives the fine-tuning results (per-seed values in Table A1). The headline is how flat the table is. On Oceanship and QiandaoEar22, random, ImageNet, and SSL initialization converge to statistically indistinguishable accuracy and macro-F1 (all |Δ| ≤ 0.038 on these two corpora, all p ≥ 0.095; the largest absolute gap in the full table is 0.039). Even the comparison practitioners assume is settled, ImageNet vs. random, shows no reliable advantage at this 30-epoch budget. Random initialization matches or slightly exceeds ImageNet on Oceanship (0.3622 vs. 0.3556 accuracy, p = 0.276) and on QiandaoEar22 macro-F1 (0.3227 vs. 0.2919, p = 0.408), with a non-significant VTUAD accuracy gap in ImageNet’s favor (0.3411 vs. 0.3564, p = 0.227). Where initialization does matter, it matters against SSL. On VTUAD, SSL fine-tuning is significantly worse than ImageNet in accuracy (−0.0388, p = 0.034), though the macro-F1 gap is not significant (−0.0212, p = 0.178). VTUAD supplies 53% of the pretraining corpus, yet its SSL fine-tuning is the only cell in Table 3 where SSL is significantly worse. This is consistent with the recording-identity diagnosis of Section 4.3. The encoder harvested the most recording-specific structure on the corpus it saw most, and fine-tuning needs more time to overwrite that structure. The single encouraging cell for SSL is QiandaoEar22 macro-F1, where SSL fine-tuning posts the best mean (0.3291 vs. 0.2919 ImageNet, +0.0372, p = 0.310, same sign on two of three seeds). We treat it as a hint, short of a finding.
Two absolute-level observations frame the table. First, no configuration on VTUAD exceeds its majority-class chance accuracy (best: ImageNet 0.3564 vs. chance 0.3567). The classifiers trade accuracy on the dominant tug class for minority-class coverage, which is exactly why macro-F1 is the primary metric. Second, the macro-F1 levels (0.12–0.33) quantify how hard these leakage-safe splits are relative to the within-corpus numbers customary in the literature, consistent with [6].

4.3. Diagnosis: The SSL Representation Encodes Recording Identity, Not Vessel Type

The VTUAD probe rows expose the mechanism. For the SSL probe, validation accuracy at the selected checkpoint is 0.456–0.462 across seeds, while test accuracy is 0.258–0.277, a 0.18–0.20 gap that is stable across seeds. For the ImageNet probe on the same splits the gap is only 0.033–0.037 (val 0.376–0.389, test 0.339–0.356). The splits are recording-level: class-stratified at the recording level with a fixed seed, so that no recording crosses partitions [6], and identical across methods. The gap is therefore a property of the representation, not of split luck or head-training randomness. The linear head is trained on embeddings of the very train split the encoder was contrastively fitted on. It finds, and is selected for, linear structure that simply does not persist into the held-out clips. One explicit alternative explanation must be disposed of first. The SSL encoder was pretrained on inputs standardized with pooled-corpus statistics (mean −51.7443, std 14.1745), whereas the downstream probes standardize with per-dataset train statistics (Section 3.3). A normalization mismatch could in principle handicap the SSL probe on VTUAD. We re-ran the VTUAD SSL probe with the pooled-corpus statistics instead. Test accuracy moves only from 0.2663 ± 0.0080 to 0.2759 ± 0.0185, an improvement within seed noise, as the overlapping standard deviations indicate. The gap to the ImageNet probe narrows from −0.0829 to −0.0733 (≈12%), and macro-F1 (0.1353 ± 0.0034) remains well below the ImageNet probe’s 0.1637 ± 0.0069. Normalization mismatch therefore cannot explain the VTUAD probing deficit; the control, if anything, hardens the diagnosis. The reading we favor is instance/recording-identity coding, and a recording-identity probe demonstrates it directly. A linear head trained to predict the recording ID (361 classes, majority chance level 0.015; these 361 classes are the mappable recording IDs of the 369 train-split recordings, after dropping 8 background-class recordings with cross-group filename collisions)2 from the train-split embeddings reaches 0.9969 ± 0.0001 on SSL embeddings. That is stably above ImageNet features (0.9618 ± 0.0016) and random features (0.8739 ± 0.0422). A competing reading of the val–test gap is generic fragility: SSL features might simply tolerate any split perturbation worse, with nothing specific to recordings. The probe comparison excludes it. Generic fragility would not make SSL embeddings more recording-separable than the ImageNet control, yet they are (0.997 vs. 0.962). That ImageNet features also reach 0.962 is expected and carries no diagnostic weight. Adjacent segments of one recording share hydrophone, noise floor, and propagation conditions, so any reasonable representation is highly recording-separable. The diagnostic weight is carried by the two ways SSL goes beyond that baseline: the probe is near-perfect on SSL embeddings (0.997 vs. 0.962), and only the SSL probe pairs this with a 0.18–0.20 validation–test accuracy gap on the vessel-type task (vs. 0.033–0.037 for ImageNet). Read together, the contrastive objective has optimized recording-identity separability at the expense of vessel-type semantics. The better the encoder tells recordings apart, the more the linear head can harvest recording-specific structure that does not generalize across the recording-disjoint split boundary. Two caveats bound the claim: the measurement is train-fit (the frozen splits are recording-disjoint, so no generalization-to-unseen-recordings reading is available), and the task is intrinsically easy in this high-dimensional, many-segments-per-recording regime (even random features are linearly fittable to 0.81–0.91). The defensible statement is that recording identity is near-perfectly linearly separable in the SSL embeddings and stably more so than in both controls. NT-Xent rewards telling individual training segments apart. The part of that code that a linear probe can harvest (recording sessions, channel conditions, hull-specific noise) is exactly the part that clip-disjoint evaluation punishes. This is the in-corpus analogue of Hummel et al.’s record-ID finding [15], where pretrained audio embeddings were shown to be geometrically dominated by recording identity. It is also the geometric echo of [6]’s result that acquisition channel dominates the supervised feature space. The mechanism is straightforward in hindsight. Our augmentation family (Section 3.2) perturbs time alignment, gain, and local time–frequency content, but never the channel or the recording context. The cheapest invariance the NT-Xent objective can purchase is therefore precisely “same recording, same channel”, a feature set that cross-water deployment punishes instead of rewarding [6]. QiandaoEar22 shows a milder version of the same pattern (SSL probe validation 0.475–0.482 vs. test 0.442–0.467, with the accuracy loss concentrated on the majority noise class). Oceanship, where the SSL probe is mildly better, is the corpus whose test split shares the most recording conditions with train.

5. Discussion

What the pilot establishes. In-corpus SimCLR pretraining, at 176 k segments and 40 epochs, provides no free upgrade over ImageNet initialization for UATR. As a feature extractor it is significantly worse on two of three corpora. As an initialization it is statistically indistinguishable from random on two corpora and significantly worse than ImageNet on the third. The flatness of Table 3 also carries an independent practical message. Within this pilot’s 30-epoch fine-tuning budget, we observe no reliable advantage of ImageNet initialization over random. Random matches ImageNet on Oceanship accuracy (0.3622 vs. 0.3556, p = 0.276) and exceeds it on QiandaoEar22 macro-F1 (0.3227 vs. 0.2919, p = 0.408). At this pilot’s budget we detect no significant advantage of ImageNet initialization. This negative result needs verification at a larger seed count before it can be generalized. For SSL itself the bar to clear is then no longer “beat a strong transfer prior” but “beat training from scratch”, and in-corpus contrastive pretraining does not clear it.
Why, and why this was worth measuring directly. The pilot rules out the cheap hypothesis that volume of same-site unlabeled data alone yields useful invariances. The failure mode we observe is augmentation-induced: with channel-invariant augmentations, contrastive learning optimizes exactly the invariance (recording identity) that cross-corpus deployment breaks [6,15]. The omission is specific. The family varies time alignment, frequency alignment, gain, and local time–frequency masking, but never the channel. Hydrophone response and reverberation are identical between the two views of every segment. That gap may be exactly why SSL bought nothing in this pilot. Adding channel simulation (hydrophone-response and reverberation variation) to the augmentation family is accordingly the first change the follow-on study should make, ahead of any increase in scale. The two remedies operate at different levels. Channel-simulation augmentation answers “what to learn”, because the direct cause of this failure is an augmentation family that never perturbs the channel. Scaling up (longer schedules, larger batches) answers “how well to learn”, the SimCLR budget shortfall. They target different levels, so channel simulation ranks ahead of scale-up. This suggests four concrete design requirements for the large-scale follow-on study, each now backed by a measurement instead of intuition. (i) Pretrain across sites and hydrophones (multi-station ONC archives [22], SanctSound [23]), so that recording identity is no stable feature to latch onto. (ii) Add channel-varying augmentations (simulated reverb/noise-floor variation, band-pass perturbations) so view consistency cannot be achieved by channel matching. (iii) Scale the schedule and batch; the vision literature’s gains from longer training and larger negative pools [9] have not been given a chance at 40 epochs/batch 256. (iv) Include recording-identity probes and label-shuffling controls [15] as standard evaluation, so that “the representation learned something” cannot be confused with “the representation learned recordings.” Masked-reconstruction objectives (MAE-style [10,12,13]) are the natural comparison arm, since reconstruction losses do not reward instance discrimination and may be less prone to identity coding.
Honest caveats. This is a pilot, and its statistical resolution is deliberately coarse. With n = 3 seeds only large effects are detectable, and several same-sign trends (Oceanship probe +0.021 F1; QiandaoEar22 fine-tune +0.037 F1) may be real but are not established. (The n = 1 pretraining limitation is flagged separately below, because it bounds everything reported.) We varied one SSL method at one scale on one backbone. VICReg-style non-contrastive objectives [14,18], larger encoders, and external pretraining corpora are all untested here. The QiandaoEar22 fine-tune macro-F1 hint and the Oceanship probe hint point where a scaled study might first look for a win. Finally, our pretraining corpus unions three sites but is dominated by VTUAD volume (93,592 of 176,481 segments, 53%), and the pretrain normalization uses pooled statistics. The Section 4.3 control shows the normalization choice explains only ≈12% of the VTUAD probing deficit. Corpus-composition balancing (e.g., equalizing the three corpora’s volumes) is another untested variable. The observed pattern (the largest validation–test gap on exactly the corpus that dominates pretraining) is precisely what the recording-identity account predicts, so a composition-balanced replication is also the sharpest test of that account. The follow-on study should revisit both choices systematically.
A limitation we want to flag prominently: pretraining itself is n = 1. The three seeds vary only the downstream probe and fine-tuning; all 45 configurations share a single SSL encoder from a single pretraining run. Encoder-level (pretraining-run) variance is therefore entirely unmeasured. A different pretraining initialization, or even a different data order, could yield an encoder whose downstream behavior differs from the one reported here. Our conclusions are strictly conditioned on this one pretrained model. Follow-up work should run multiple independent pretraining seeds, in addition to downstream seeds, to quantify encoder-level variance.

6. Conclusions

We conducted a controlled three-corpus pilot of self-supervised pretraining for underwater acoustic target recognition. The pilot runs SimCLR on 176,481 unlabeled segments, evaluated by linear probing and full fine-tuning against ImageNet and random initialization over 15 cells and three seeds (45 runs). The result is predominantly negative and, we argue, valuable. In-corpus contrastive pretraining produces representations that significantly underperform ImageNet features under linear probing on two of three corpora (VTUAD −0.083 accuracy, p = 0.017; QiandaoEar22 −0.053, p = 0.012). They also confer no measurable advantage as a fine-tuning initialization. A diagnostic val–test gap is consistent with the learned features encoding recording identity instead of vessel type. The pilot converts “pretrain on more hydrophone audio” from an assumption into a set of testable design requirements: cross-site corpora, channel-varying augmentations, longer schedules, and identity-control evaluations. These requirements define the follow-on large-scale study on multi-station ONC and SanctSound archives.

Funding

This research was funded by the National Natural Science Foundation of China under Grant 12501435.

Author Contributions (CRediT)

Conceptualization, H.Y. and W.W.; methodology, H.Y.; software, H.Y. and Y.C.; validation, H.Y. and X.W.; formal analysis, H.Y. and G.C.; investigation, H.Y.; resources, W.W.; data curation, H.Y. and T.L.; writing—original draft preparation, H.Y.; writing—review and editing, W.W. and G.C.; visualization, H.Y.; supervision, W.W.; project administration, W.W.; funding acquisition, W.W. In addition, T.L. performed engineering verification of the pretraining and evaluation pipeline. All authors have read and agreed to the published version of the manuscript.

Data availability

All per-run artifacts (histories, checkpoints, summary JSONs), the aggregation script (aggregate_paper5_ssl.py), and the authoritative aggregated numbers (paper5_ssl_final_results.md) accompany this manuscript; split definitions are those of [6].

Conflicts of Interest

Hao Yuan, Yu Chen, Tian Li, and Xinyu Wu are also affiliated with CSSC-LINCOM Electronics (Wuhan) Co., Ltd. (affiliation 2). The authors declare no other competing interests.

Appendix A. Per-Seed Results

All 45 per-seed runs; val_acc is the selection criterion (best validation accuracy) at the selected epoch.
Table A1. Fine-tuning, per seed.
Table A1. Fine-tuning, per seed.
Dataset Init Seed Test acc Test macro-F1 Val acc Best epoch
Oceanship imagenet 42 0.3674 0.1545 0.3544 5
Oceanship imagenet 43 0.3503 0.1796 0.3542 7
Oceanship imagenet 44 0.3492 0.1340 0.3535 4
Oceanship random 42 0.3652 0.1696 0.3578 7
Oceanship random 43 0.3602 0.1502 0.3579 6
Oceanship random 44 0.3612 0.1553 0.3559 4
Oceanship ssl 42 0.3503 0.1606 0.3521 5
Oceanship ssl 43 0.3511 0.1641 0.3441 6
Oceanship ssl 44 0.3571 0.1482 0.3565 4
QiandaoEar22 imagenet 42 0.5587 0.2759 0.5118 8
QiandaoEar22 imagenet 43 0.5557 0.3102 0.5063 10
QiandaoEar22 imagenet 44 0.5391 0.2895 0.5092 12
QiandaoEar22 random 42 0.5400 0.2734 0.5028 8
QiandaoEar22 random 43 0.4954 0.3151 0.4937 4
QiandaoEar22 random 44 0.5564 0.3795 0.4939 10
QiandaoEar22 ssl 42 0.5391 0.3598 0.5088 18
QiandaoEar22 ssl 43 0.5296 0.2985 0.5194 10
QiandaoEar22 ssl 44 0.5449 0.3290 0.5018 9
VTUAD imagenet 42 0.3510 0.2329 0.4555 5
VTUAD imagenet 43 0.3637 0.2574 0.4857 3
VTUAD imagenet 44 0.3546 0.2310 0.4646 10
VTUAD random 42 0.3186 0.2343 0.4418 3
VTUAD random 43 0.3613 0.1816 0.4560 13
VTUAD random 44 0.3434 0.2334 0.4591 8
VTUAD ssl 42 0.3019 0.2054 0.4591 2
VTUAD ssl 43 0.3210 0.2222 0.4855 12
VTUAD ssl 44 0.3301 0.2301 0.4870 20
Table A2. Linear probing, per seed.
Table A2. Linear probing, per seed.
Dataset Init Seed Test acc Test macro-F1 Val acc Best epoch
Oceanship imagenet 42 0.3138 0.1037 0.3156 25
Oceanship imagenet 43 0.3170 0.1068 0.3153 13
Oceanship imagenet 44 0.3169 0.0816 0.3187 17
Oceanship ssl 42 0.3255 0.1256 0.3202 25
Oceanship ssl 43 0.3210 0.1185 0.3210 11
Oceanship ssl 44 0.3188 0.1113 0.3168 6
QiandaoEar22 imagenet 42 0.5204 0.2444 0.4825 2
QiandaoEar22 imagenet 43 0.5046 0.2625 0.4816 4
QiandaoEar22 imagenet 44 0.4977 0.2496 0.4840 4
QiandaoEar22 ssl 42 0.4665 0.2293 0.4746 2
QiandaoEar22 ssl 43 0.4415 0.2241 0.4816 4
QiandaoEar22 ssl 44 0.4547 0.2599 0.4783 14
VTUAD imagenet 42 0.3388 0.1729 0.3759 21
VTUAD imagenet 43 0.3531 0.1622 0.3877 6
VTUAD imagenet 44 0.3556 0.1561 0.3888 13
VTUAD ssl 42 0.2771 0.1250 0.4616 2
VTUAD ssl 43 0.2637 0.1225 0.4557 9
VTUAD ssl 44 0.2580 0.1206 0.4619 10

References

  1. Santos-Domínguez, D.; Torres-Guijarro, S.; Cardenal-López, A.; Pena-Gimenez, A. ShipsEar: An Underwater Vessel Noise Database. Appl. Acoust. 2016, 113, 64–69. [Google Scholar] [CrossRef]
  2. Irfan, M.; Jiangbin, Z.; Ali, S.; Iqbal, M.; Masood, Z.; Hamid, U. DeepShip: An Underwater Acoustic Benchmark Dataset and a Separable Convolution Based Autoencoder for Classification. Expert Syst. With Appl. 2021, 183, 115270. [Google Scholar] [CrossRef]
  3. Li, Z.; Xiang, S.; Yu, T.; Gao, J.; Ruan, J.; Hu, Y.; Liu, T.; Fu, Y. Oceanship: A Large-Scale Dataset for Underwater Audio Target Recognition. In Proceedings of the International Conference on Intelligent Computing (ICIC); Springer: Singapore, 2024; pp. 475–486. [Google Scholar]
  4. Du, X.; Hong, F. QiandaoEar22: A High-Quality Noise Dataset for Identifying Specific Ship from Multiple Underwater Acoustic Targets Using Ship-Radiated Noise. EURASIP J. Adv. Signal Process. 2024, 2024, 96. [Google Scholar] [CrossRef]
  5. Domingos, L.C.F.; Santos, P.E.; Skelton, P.S.M.; Brinkworth, R.S.A.; Sammut, K. An Investigation of Preprocessing Filters and Deep Learning Methods for Vessel Type Classification With Underwater Acoustic Data. IEEE Access 2022, 10, 117582–117596. [Google Scholar] [CrossRef]
  6. Yuan, H.; Wang, W.; Zeng, L.; Chen, G.; Li, T.; Hou, X. When Underwater Acoustic Recognition Fails Across Datasets: A Cross-Dataset Benchmark Revealing Zero-Transfer and Label Shift. Preprints (companion paper). 2026, 202609.0919.v1. [Google Scholar] [CrossRef]
  7. Yuan, H.; Wang, W. Unsupervised Domain Adaptation for Cross-Water Underwater Acoustic Target Recognition: A Four-Corpus Benchmark on Oceanship, QiandaoEar22, an AIS-Auto-Labeled VTUAD Reconstruction, and ShipsEar. In Preprints; (companion paper); 2026. [Google Scholar]
  8. Yuan, H.; Wang, W. Source-Free Domain Adaptation for Underwater Acoustic Target Recognition: When Batch Normalization Statistics Alone Suffice. In Preprints; (companion paper); 2026. [Google Scholar]
  9. Chen, T.; Kornblith, S.; Norouzi, M.; Hinton, G. A Simple Framework for Contrastive Learning of Visual Representations. Proc. 37th Int. Conf. Mach. Learn. (ICML) 2020, PMLR 119, 1597–1607. [Google Scholar]
  10. He, K.; Chen, X.; Xie, S.; Li, Y.; Dollár, P.; Girshick, R. Masked Autoencoders Are Scalable Vision Learners. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2022; pp. 16000–16009. [Google Scholar]
  11. Baevski, A.; Zhou, Y.; Mohamed, A.; Auli, M. wav2vec 2.0: A Framework for Self-Supervised Learning of Speech Representations. Adv. Neural Inf. Process. Syst. 33 (NeurIPS) 2020, 12449–12460. [Google Scholar]
  12. Huang, P.-Y.; Xu, H.; Li, J.; Baevski, A.; Auli, M.; Galuba, W.; Metze, F.; Feichtenhofer, C. Masked Autoencoders That Listen. Adv. Neural Inf. Process. Syst. 35 (NeurIPS) 2022, 28708–28720. [Google Scholar] [CrossRef]
  13. Xu, K.; Xu, Q.; You, K.; Zhu, B.; Feng, M.; Feng, D.; Liu, B. Self-Supervised Learning-Based Underwater Acoustical Signal Classification via Mask Modeling. J. Acoust. Soc. Am. 2023, 154, 5–15. [Google Scholar] [CrossRef] [PubMed]
  14. Hummel, H.I.; Gansekoele, A.; Bhulai, S.; van der Mei, R. The Computation of Generalized Embeddings for Underwater Acoustic Target Recognition Using Contrastive Learning. Appl. Acoust. 2026, 243, 111103. [Google Scholar] [CrossRef]
  15. Hummel, H.I.; Bhulai, S.; van der Mei, R.; Ghani, B. Linear Probing Enables Ship-Radiated Noise Recognition with Pretrained Audio Embeddings. Ecol. Inform. 2026, 95, 103709. [Google Scholar] [CrossRef]
  16. He, K.; Zhang, X.; Ren, S.; Sun, J. Deep Residual Learning for Image Recognition. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2016; pp. 770–778. [Google Scholar]
  17. Deng, J.; Dong, W.; Socher, R.; Li, L.-J.; Li, K.; Fei-Fei, L. ImageNet: A Large-Scale Hierarchical Image Database. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2009; pp. 248–255. [Google Scholar]
  18. Bardes, A.; Ponce, J.; LeCun, Y. VICReg: Variance-Invariance-Covariance Regularization for Self-Supervised Learning. In Proceedings of the International Conference on Learning Representations (ICLR), 2022. [Google Scholar]
  19. Yuan, H.; Wang, W. VTUAD: An AIS-Auto-Labeled Vessel-Type Underwater Acoustic Dataset Reconstructed from Ocean Networks Canada Archives. Preprints 2026, 202609.0617.v1, (companion paper); dataset: https://doi.org/10.5281/zenodo.22274549. [Google Scholar] [CrossRef]
  20. Park, D.S.; Chan, W.; Zhang, Y.; Chiu, C.-C.; Zoph, B.; Cubuk, E.D.; Le, Q.V. SpecAugment: A Simple Data Augmentation Method for Automatic Speech Recognition. Proceedings of Interspeech, 2019; pp. 2613–2617. [Google Scholar]
  21. Loshchilov, I.; Hutter, F. Decoupled Weight Decay Regularization. In Proceedings of the International Conference on Learning Representations (ICLR), 2019. [Google Scholar]
  22. Ocean Networks Canada data archive. Available online: https://data.oceannetworks.ca (accessed on 2026).
  23. NOAA/NPS SanctSound project. Available online: https://sanctsound.noaa.gov (accessed on 2026).
1
VTUAD as used throughout this series is the reconstructed corpus of our companion dataset descriptor (Yuan, H.; Wang, W. VTUAD: An AIS-Auto-Labeled Vessel-Type Underwater Acoustic Dataset Reconstructed from Ocean Networks Canada Archives, companion paper, Preprints 2026, 202609.0617.v1, https://doi.org/10.20944/preprints202609.0617.v1; dataset: https://doi.org/10.5281/zenodo.22274549), rebuilt from the ONC Strait of Georgia archives with an automatic AIS labeling pipeline. Reference [5] cites the original VTUAD release of Domingos et al.; the reconstruction differs from it in archive coverage, label source, and split definitions, and it is the reconstruction that all numbers in this paper are measured on.
2
The frozen VTUAD train split contains 369 recordings, but the probe’s manifest-based ID mapping drops 8 background-class recordings whose segments are double-registered in the feature manifest under both background/* and vessel/* group names (1,511 segments with conflicting group assignments, excluded as ambiguous), leaving 361 mappable recording IDs and 92,081 of the 93,592 train segments. The 8 dropped recordings are all background-class; no vessel-class recording is affected. The full account of this filename collision is given in the companion dataset descriptor [19] (arXiv:2609.XXXXX; dataset: https://doi.org/10.5281/zenodo.22274549).
Table 1. Corpora, frozen splits (segment counts), and test-split majority-class chance level. Splits are recording/clip-disjoint [6]; the same train splits, de-labeled, constitute the pretraining corpus.
Table 1. Corpora, frozen splits (segment counts), and test-split majority-class chance level. Splits are recording/clip-disjoint [6]; the same train splits, de-labeled, constitute the pretraining corpus.
Dataset Classes Train Val Test Test majority chance
Oceanship 15 56,998 12,242 12,206 0.2260 (Tug)
QiandaoEar22 8 25,891 5,426 4,324 0.2761 (noise)
VTUAD 12 93,592 20,479 18,202 0.3567 (tug)
Table 2. Linear probing, test set, mean ± std over seeds 42/43/44 (ddof = 0). Δ columns: SSL minus ImageNet, paired per seed; p from the paired t-test (df = 2). Significant pilot-level gaps in bold.
Table 2. Linear probing, test set, mean ± std over seeds 42/43/44 (ddof = 0). Δ columns: SSL minus ImageNet, paired per seed; p from the paired t-test (df = 2). Significant pilot-level gaps in bold.
Dataset ImageNet acc SSL acc Δacc (p) ImageNet F1 SSL F1 ΔF1 (p)
Oceanship 0.3159 ± 0.0015 0.3218 ± 0.0028 +0.0059 (0.188) 0.0974 ± 0.0112 0.1185 ± 0.0058 +0.0211 (0.056)
QiandaoEar22 0.5076 ± 0.0095 0.4542 ± 0.0102 −0.0533 (0.012) 0.2522 ± 0.0076 0.2378 ± 0.0158 −0.0144 (0.414)
VTUAD 0.3492 ± 0.0074 0.2663 ± 0.0080 −0.0829 (0.017) 0.1637 ± 0.0069 0.1227 ± 0.0018 −0.0410 (0.008)
Table 3. Full fine-tuning, test set, mean ± std over seeds 42/43/44 (ddof = 0). Δ: SSL minus the column’s reference, paired per seed (df = 2 t-test). Significant pilot-level gaps in bold. Majority-class chance levels are 0.2260 / 0.2761 / 0.3567 (Table 1).
Table 3. Full fine-tuning, test set, mean ± std over seeds 42/43/44 (ddof = 0). Δ: SSL minus the column’s reference, paired per seed (df = 2 t-test). Significant pilot-level gaps in bold. Majority-class chance levels are 0.2260 / 0.2761 / 0.3567 (Table 1).
Dataset Metric Random ImageNet SSL ΔSSL−IN (p) ΔSSL−RND (p)
Oceanship acc 0.3622 ± 0.0022 0.3556 ± 0.0083 0.3528 ± 0.0030 −0.0028 (0.743) −0.0094 (0.095)
F1 0.1584 ± 0.0082 0.1560 ± 0.0186 0.1576 ± 0.0068 +0.0016 (0.873) −0.0007 (0.930)
QiandaoEar22 acc 0.5306 ± 0.0258 0.5512 ± 0.0086 0.5379 ± 0.0063 −0.0133 (0.305) +0.0073 (0.651)
F1 0.3227 ± 0.0436 0.2919 ± 0.0141 0.3291 ± 0.0250 +0.0372 (0.310) +0.0064 (0.890)
VTUAD acc 0.3411 ± 0.0175 0.3564 ± 0.0053 0.3177 ± 0.0118 −0.0388 (0.034) −0.0234 (0.110)
F1 0.2164 ± 0.0246 0.2404 ± 0.0120 0.2192 ± 0.0103 −0.0212 (0.178) +0.0028 (0.903)
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.