Preprint
Article

This version is not peer-reviewed.

Stabilized Multi-Objective Optimization for Camouflage Attack in Remote Sensing

Submitted:

11 September 2026

Posted:

15 September 2026

You are already at the latest version

Abstract
We investigate the vulnerability of multi-task maritime surveillance systems where object detection and semantic segmentation are deployed concurrently. Existing adversarial patches often fail in the satellite domain due to extreme scale variations and gradient interference between disparate architectures. To address this challenge, we propose Stabilized Multi-Objective Optimization (SMO), a framework that jointly targets Mask R-CNN and U-Net with a single adversarial patch. Our approach introduces a logarithmic-inversion loss to provide stable, bounded gradients, with various strategies to dynamically balance seven distinct loss heads. Unlike traditional patch-based attacks where orientation, geometry and scale invariance are achieved by Expectation over Transformation (EoT), we utilize a U-Net generator to synthesize hull-aligned textures, ensuring this invariance implicitly across diverse settings. Experiments on the Airbus Ship Detection, DOTA, and VisDrone datasets demonstrate that SMO reduces both AP and mIoU significantly, exposing a critical security gap in integrated maritime monitoring frameworks.
Keywords: 
;  ;  ;  ;  

1. Introduction

Modern remote sensing systems integrate satellite and aerial imagery with deep learning for applications such as maritime surveillance, border security, environmental monitoring, and disaster response [1]. In these systems, heterogeneous perception pipelines may combine region-proposal-based object detectors, such as Mask R-CNN [2], with dense semantic segmentation networks, such as U-Net [3]. Although these architectures provide complementary representations of the same scene, their increasing use also raises concerns regarding the robustness of remote sensing perception systems to adversarial perturbations.
Deep neural networks are known to be vulnerable to carefully optimized perturbations that can substantially alter their predictions. In remote sensing, adversarial patch generation is particularly challenging because targets may appear at different scales, orientations, geometries, and Ground Sample Distances (GSDs). Existing attacks are typically optimized for a single prediction task or architecture, and their effectiveness can degrade when transferred between fundamentally different models, such as object detectors and semantic segmenters. In addition, jointly attacking multiple prediction objectives introduces competing loss terms with substantially different magnitudes and gradient directions, making optimization difficult.
To address these challenges, we propose Stabilized Multi-Objective Optimization (SMO), a geometry-aware adversarial camouflage framework that jointly targets object detection and semantic segmentation models. SMO employs a U-Net-based generator to synthesize a spatially coherent adversarial texture and applies it through a target-aligned mask so that the perturbation follows the geometry of the object. To improve optimization across heterogeneous objectives, we introduce a logarithmic-inversion adversarial loss and incorporate a feature-inconsistency objective. We further investigate several strategies for balancing the resulting loss terms, including uniform weighting, static heuristics, cyclic optimization, adaptive stochastic optimization, and GradNorm-based gradient normalization.
The proposed framework is evaluated on the Airbus Ship Detection, DOTA, and VisDrone datasets. Experiments show that SMO substantially degrades both detection and segmentation performance and outperforms the evaluated adversarial patch baselines on the primary Airbus benchmark. Cross-dataset and cross-architecture experiments further demonstrate transferability to YOLO-based detectors and several convolutional detection and segmentation architectures, while also revealing reduced effectiveness against some architecturally dissimilar models. Additional experiments analyze patch coverage, patch geometry, multi-patch configurations, optimization strategies, and the contribution of individual SMO components.
The main contributions of this paper are summarized as follows:
  • We formulate adversarial camouflage against heterogeneous remote sensing perception systems as a unified multi-objective optimization problem that jointly targets object detection and semantic segmentation.
  • We introduce a logarithmic-inversion adversarial objective together with feature inconsistency to improve optimization across multiple heterogeneous detector and segmentation losses.
  • We develop a U-Net-based geometry-aware perturbation mechanism that produces target-aligned adversarial textures and accommodates variations in object scale, orientation, and geometry without relying on a fixed-size rectangular patch.
  • We systematically investigate multiple fixed and adaptive loss-balancing strategies and analyze their influence on adversarial optimization.
  • We validate the proposed framework on Airbus, DOTA, and VisDrone through baseline comparisons, cross-model transferability experiments, patch-configuration studies, and component-level ablations.
Figure 1. Comparison of single-task and joint-model adversarial suppression. A segmentation-optimized patch ( 1 st row) transfers poorly to the region-proposal-based detector, while a detection-optimized patch ( 2 nd row) does not consistently suppress the dense-prediction segmenter. In contrast, the proposed Stabilized Multi-Objective Optimization (SMO) framework ( 3 rd row) jointly optimizes the adversarial texture against both architectures, resulting in more balanced degradation across detection and segmentation.
Figure 1. Comparison of single-task and joint-model adversarial suppression. A segmentation-optimized patch ( 1 st row) transfers poorly to the region-proposal-based detector, while a detection-optimized patch ( 2 nd row) does not consistently suppress the dense-prediction segmenter. In contrast, the proposed Stabilized Multi-Objective Optimization (SMO) framework ( 3 rd row) jointly optimizes the adversarial texture against both architectures, resulting in more balanced degradation across detection and segmentation.
Preprints 232964 g001

3. Methodology

3.1. Threat Model

We consider a white-box, digital targeted attack against systems designed to distinguish ship from background. Specifically, we assume the attacker can train against a surrogate model, with full access to the model architecture and representative training data, to design a single, fixed patch for a camouflage attack. This white-box setting serves as a conservative, worst-case evaluation. It is realistic in remote sensing contexts where architectures, pre-trained weights, and training recipes are often shared or can be approximated from public datasets. For practicality, the attacker is constrained to modifying only the ship surface area (e.g., replacing image pixel values of the ship with the trained patch) and cannot alter sensor hardware or adapt the patch at inference time. We adopt this threat model to: (i) provide a reproducible upper bound on system vulnerability, and (ii) evaluate the transferability of the learned patch to black-box victims, demonstrating risk even when exact model parameters remain unavailable.

3.2. System Overview

The proposed SMO framework aims to synthesize a geometry-aware adversarial texture that concurrently suppresses region-proposal and dense-prediction pipelines. As illustrated in Figure 2, the pipeline comprises three core functional modules: (1) a generative model G that maps a latent noise manifold z to a universal adversarial patch p, (2) a geometry-aware poisoning function P that integrates the patch onto the ship hull x to form the adversarial input x adv , and (3) a dual-paradigm victim suite consisting of a detector F (Mask R-CNN) and a segmentor U (U-Net). These interactions are formalized as follows:
p = G ( z ) , x adv = P ( x , p )
( y d , m s ) = ( F ( x ) , U ( x ) ) , ( y d adv , m s adv ) = ( F ( x adv ) , U ( x adv ) )
where y d denotes the set of bounding boxes, classification scores, and instance masks, while m s represents the pixel-wise binary segmentation mask.
Figure 2. Schematic of the proposed SMO framework. A generator G synthesizes adversarial textures p onto ships via a geometry-aware function P from noise input z. The resulting image x adv simultaneously targets a detector F (Mask R-CNN) and segmenter U (U-Net). SMO employs a logarithmic-inversion to balance seven loss heads—including RPN, ROI, and feature inconsistency—ensuring stable gradient fusion and comprehensive multi-task evasion.
Figure 2. Schematic of the proposed SMO framework. A generator G synthesizes adversarial textures p onto ships via a geometry-aware function P from noise input z. The resulting image x adv simultaneously targets a detector F (Mask R-CNN) and segmenter U (U-Net). SMO employs a logarithmic-inversion to balance seven loss heads—including RPN, ROI, and feature inconsistency—ensuring stable gradient fusion and comprehensive multi-task evasion.
Preprints 232964 g002

3.3. Generator Model

To mitigate high-frequency noise and improve cross-sensor transferability, we optimize over a latent manifold, utilizing a U-Net architecture to impose a structural prior that ensures spatial coherency. At each iteration t, the generator transforms a trainable latent tensor z t into a spatially-consistent patch: p t = G ( z t ) . The optimization simultaneously updates the weights of G and the coefficients of z. To prevent overfitting and ensure generalization, we employ an early-stopping mechanism governed by a dual-objective criterion:
∥ z t − z t − 1 ∥ 2 < ϵ or Ψ ( x adv ) < τ
where ϵ is a latent stability threshold and Ψ ( · ) represents a performance metric (e.g., AP for detection or IoU for segmentation). Upon convergence, the generator is frozen, yielding a universal patch p for camouflage attack during inference.

3.4. Geometry-Aware Poisoning Function

The adversarial texture synthesized by G is a global signal p ∈ R 3 × H × W , where H and W are the sizes of the input image. We define a geometry-aware poisoning function P to fit ship’s geometry. The adversarial image x adv is obtained through a convex combination of the clean image and the patch signal:
x adv = P ( x , p ) = x ⊙ ( 1 − M hull ) + p ⊙ M hull
where M hull denotes a binary elliptical mask strictly localized within the ship’s hull boundaries, and ⊙ represents the Hadamard product.

3.5. Standard Object Detection and Segmentation Objectives

To achieve comprehensive camouflage attack, the framework must suppress the standard supervisory signals of both the regional-proposal detector F and the dense segmentor U. We briefly list these objectives as follows:
Detector Objectives ( L F ): The Mask R-CNN detector [2] utilizes a two-stage optimization process. The first stage, the Region Proposal Network (RPN), minimizes the objectness score L rpn - obj and the anchor box regression loss L rpn - reg . The second stage, the ROIHead, performs region-wise classification and refines bounding box regression by minimizing L roi - cls and L roi - reg , respectively. Additionally, an instance-level mask head minimizes the binary cross-entropy L mask within each proposal. Together, these constitute the set of detector loss heads:
L F = { L rpn - obj , L rpn - reg , L roi - cls , L roi - reg , L mask }
Segmentor Objective ( L seg ): The U-Net segmentor performs pixel-level semantic classification by minimizing a hybrid Dice and Cross-Entropy objective:
L seg = DiceCE ( m d , m )
These six task-specific losses define the standard performance manifold. In the following sections, we describe how our logarithmic-inversion scheme transforms these losses into a stabilized adversarial training objective.

3.6. Stabilized Adversarial Optimization

Logarithmic Inversion for Gradient Stability.

To perform camouflage attack, the goal of training is to compromise those standard objectives listed above. A primary challenge in multi-objective adversarial training is the instability arising from the disparate magnitudes of the loss signals. A naive sign-reversal (i.e., L adv = − L ) often yields a volatile optimization landscape, characterized by gradient explosion when L is large or vanishing gradients when L approaches zero. To enforce a stable, bounded gradient profile, we propose a logarithmic-inversion scheme. For each standard task loss L i ∈ { L F , L seg } , we define its adversarial counterpart as:
L i adv = log 1 + λ L i + α
where λ > 0 modulates the inversion strength and α > 0 serves as a numerical stabilizer. Since L i adv is a monotonically decreasing function of L i , minimizing the adversarial objective inherently drives the generator to maximize the model’s prediction error. The superiority of this formulation lies in its derivative with respect to the task loss:
d L i adv d L i = − λ ( L i + α ) ( L i + α + λ )
This design ensures that when the attack is failing (i.e., L i → 0 ), the gradient remains bounded away from zero: | d L i adv / d L i | ≈ λ / ( α ( α + λ ) ) . This prevents gradient vanishing and provides a persistent update signal to the generator. Conversely, as the attack succeeds (i.e., L i → ∞ ), the gradient naturally decays at a rate of O ( L i − 2 ) , providing a smooth saturation that prevents the optimizer from over-focusing on already-evaded loss heads.

Feature Inconsistency and Joint Objective.

Beyond surface-level predictions, we aim to disrupt the latent representations within the detection backbone. We introduce a feature-inconsistency loss, formulated as the squared L 2 distance between the RPN feature embeddings of the clean and adversarial inputs:
L feat = ∥ RPN ( x ) − RPN ( x adv ) ∥ 2 2 , L feat adv = log 1 + λ L feat + α
By maximizing the divergence in the feature manifold, we encourage large perceptual shifts that further degrade the robustness of the region proposal mechanism. The final total adversarial loss is constructed as a weighted composition of the seven stabilized adversarial terms:
L adv = ∑ i ∈ S w i L i adv + w feat L feat adv
where S represents the set of detector losses and the segmentation loss. This joint objective ensures that the synthesized patch p is optimized to simultaneously bypass ship detection of both the detector F and the segmentor U.

3.7. Adversarial Training Dynamics and Multi-Objective Strategies

The adversarial objective L adv is composed of seven distinct loss terms, creating a complex optimization landscape where individual objectives may compete or vary significantly in scale. We investigate six strategic configurations designed to stabilize the training of the adversarial patch and ensure robust convergence:
  • Setting A (Primitive Optimization): Directly optimizes p from noise z while bypassing G, establishing an unconstrained baseline for adversarial efficacy.
  • Setting B (Uniform Weighting): Optimizes both G and z with an equal weight for each of the components in L adv , assuming equal contribution among them.
  • Setting C (Static Heuristics): This approach utilizes fixed weights determined through extensive trial and error, thereby leveraging empirical intuition to mitigate the negative influence of several dominant loss signals.
  • Setting D (Cyclic Partitioning): Optimizes a single loss term per iteration, with an aim to decouple the objectives to prevent gradient interference.
  • Setting E (Adaptive Stochasticity): Utilizes a probabilistic sampling strategy where the optimization focus is biased toward high-impact terms, allowing the model to dynamically address the most significant deficiencies.
  • Setting F (Dynamic Gradient Normalization): Implements GradNorm [33] to automatically calibrate weights in real-time. By balancing gradient magnitudes across the seven terms, this strategy ensures that no single objective overwhelms the learning process, maintaining a stable and consistent descent.
A comprehensive comparative analysis of these strategies is detailed in Section 5.5, while the finalized training procedure is formalized in Algorithm 1.
Algorithm 1 Proposed Adversarial Patch Training
 1:
Input: Target detectors { F , U } (Mask R-CNN, U-Net), patch generator G, dataset D train , batch size b.
 2:
Output: Optimized adversarial patch p.
 3:
Initialize: Latent noise tensor z ∼ N ( 0 , I ) .
 4:
while training objective not converged do
 5:
    Sample a mini-batch X = { x i } i = 1 b from D train .
 6:
    Synthesize Patch:  p = G ( z ) .
 7:
    Adversarial Projection: For each x i ∈ X , generate x i adv = P ( x i , p ) .
 8:
    Loss Computation: Evaluate multi-objective adversarial loss L adv by processing { X , X adv } through F and U.
 9:
    Optimization Step: Update weights of G and latent vector z via the selected strategy (e.g., GradNorm) as detailed in Section 3.7.
10:
end while

3.8. Baseline Competitors and Comparative Objectives

We benchmark our approach against four general-purpose adversarial patch attacks: 1) Universal Patch (Brown et al. [8]), designed for classification and adapted here for segmentation, 2) DPatch (Liu et al. [11]), which directly optimizes pixels to minimize detector confidence, 3) ShapeShifter (Chen et al. [34]), which introduces multi-objective optimization to attack both classification and localization, and 4) Scale-Adaptive Attack (Lu et al. [23]), which conditions optimization on an empirical scale distribution. We also compare against several remote sensing and aerial camouflage methods  [21,24,25,35]. While all baselines rely on the Expectation-over-Transformation (EoT) framework [8] to explicitly model spatial variance via stochastic sampling, our method bypasses this requirement. By leveraging a U-Net generator to produce hull-aligned textures, the model implicitly internalizes geometric and scale invariance during synthesis. This enables robust attacks across diverse orientations and resolutions without the computational overhead of explicit spatial augmentations. Please refer to Section A.1 of the Appendix for further details.
Core Methodological Distinctions: Our proposed SMO framework diverges from the aforementioned baselines in three fundamental aspects:
  • Joint optimization: Existing attacks focus on isolated tasks either for classification or detection, our approach executes a joint optimization across detector and segmentation, and prevents the "objective interference".
  • Generative Patch Parameterization: Baseline methods optimize raw pixel values p directly. We parameterize the adversarial manifold via a generator G, ensuring the synthesis textures are more resilient to satellite surveillance.
  • Implicit Geometric Invariance: Existing methods rely on explicit EoT sampling to handle various variations. Our method utilizes hull-aligned extraction to conform naturally to the ship’s variance of orientation and scale.

4. Experimental Setup and Evaluation Metrics

4.1. Maritime and Aerial Benchmarks.

We primarily evaluate on the Airbus ship dection dataset [36], filtering for the 42,556 images containing valid ship instances (40,428 training/2,128 testing). The dataset’s example images and multi-instance distribution are shown in Figure 3 and Figure 4. To validate generalizability and ensure a rigorous comparison against standard remote sensing camouflage attacks [21,35], we extend our evaluation to the DOTA [37] and VisDrone [38] benchmarks. This allows us to confirm the transferability of our SMO optimization method over different object types and dataset. Table 1 summarizes the information about all the three dataset.

4.2. Victim Model Architectures and Implementation Details.

We target a Mask R-CNN [2] with a ResNet-101 FPN backbone as the primary victim model. To evaluate patch transferability, we further benchmark against ResNet-50 and ResNeXt backbones, as well as YOLOv2 and YOLOv3 [14] paradigms. All optimization pipelines are built on the Detectron2 framework [39], with our generator G and SMO logic integrated into a PyTorch-based pipeline. This setup facilitates concurrent gradient backpropagation from heterogeneous models, ensuring a rigorous evaluation across distinct architectural designs. More details about dataset and hyperparameters used during the training can be found in Appendix sections A, B.

4.3. Performance Metrics

To evaluate adversarial impact across detection and segmentation, we employ a suite of standard metrics based on Intersection-over-Union (IoU). For object detection, we classify predictions as True Positives (TP) or False Positives (FP) at an IoU threshold of 0.50, while semantic segmentation metrics are computed at the pixel level. Based on these counts, we report Precision, Recall, and the F1-score to capture the balance between localization and classification. Our primary measure of adversarial efficacy is Average Precision (AP), representing the area under the precision-recall curve. Specifically, we report AP 50 and the COCO-standard mAP, averaged across IoU thresholds from 0.50 to 0.95, to assess performance across varying levels of spatial rigor.

5. Results and Discussions

5.1. Results on Airbus Dataset

The comparative performance on the Airbus Ship Detection dataset is summarized in Table 2, where we delineate efficacy across object scales using AP s (small), AP m (medium), and AP l (large). SMO demonstrates a decisive advantage, consistently degrading detection performance across the entire scale spectrum. A critical observation is the scale-sensitivity of traditional baselines. For instance, while DPatch [11] exhibits a significant reduction in AP s , this is primarily a consequence of its fixed 32 × 32 spatial footprint, which physically occludes a majority of small ship instances. However, its effectiveness diminishes precipitously for larger vessels (minimal reduction in AP l ), as the fixed-size patch occupies an insignificant fraction of the object’s feature map. In contrast, our approach utilizes a generative, geometry-aware extraction mechanism that adaptively scales the adversarial texture relative to the ship’s hull.
This ensures a balanced degradation across all categories without requiring total occlusion. SMO maintains adversarial capability on large targets, underscoring the importance of our "hull-aligned" synthesis. Visual comparisons of the patches generated by each baseline are provided in Figure 5, illustrating the transition from rigid, square perturbations to our fluid, geometry-conforming textures.

5.2. Impact of Patch Coverage Percentage

To quantify the sensitivity of the victim models to adversarial surface area, we evaluate performance as a function of hull coverage ( 0 % − 100 % ). As detailed in Table 3, we observe a non-linear degradation in all metrics, with a critical ‘collapse threshold’ occurring beyond 50 % coverage. At full occlusion ( 100 % ), precision and recall plummet to 2.3 % and 4.4 % , respectively. These results validate that while the patch is most potent when conforming to the majority of the ship’s hull, the SMO framework begins to compromise feature extraction logic even at partial coverage, demonstrating high efficiency in adversarial settings.

5.3. Results on DOTA and VisDrone Datasets

To assess the generalizability of the proposed SMO framework beyond satellite-based maritime contexts, we benchmark against state-of-the-art camouflage attacks [21,35] on the DOTA [37] and VisDrone [38] datasets. In these experiments, we train the adversarial patch by jointly optimizing through Mask R-CNN and U-Net and utilize YOLOv2 and YOLOv3 as victim detectors to evaluate the attack’s efficacy against one-stage, anchor-based architectures. As summarized in Table 4, our method consistently induces a more profound degradation in AP 50 than the baseline competitors across all configurations.
A notable result occurs on the DOTA benchmark using the YOLOv2 detector: while the strongest existing baseline achieves a 57.1% reduction in AP, our SMO framework collapses the detection performance from an 88.0% baseline to a mere 6.0%, representing a 93.1% relative decrease. This significant performance gap suggests that while traditional camouflage attacks successfully incorporate physical-realism priors (e.g., TV and NPS losses), they lack the gradient stability and multi-objective alignment necessary to fully compromise the detector’s feature extraction logic. The consistent superiority of our approach across disparate YOLO versions further validates the transferability of our synthesized textures, confirming that our generative strategy successfully captures universal adversarial vulnerabilities in the remote sensing domain.

5.4. Results on Transferability Experiments

To evaluate the generalizability of our SMO-synthesized textures, we deployed patches trained on the Mask R-CNN (ResNet-101) victim against a diverse suite of architectures. As shown in Table 5, the adversarial potency remains exceptionally high across structurally related convolutional frameworks, with the source model experiencing a catastrophic collapse in F 1 and mIoU (dropping from 70.2/86.4 to 16.3/28.5). We observe significant cross-task transfer to U-Net and FPN backbones, confirming that our multi-objective optimization captures fundamental geometric vulnerabilities within the CNN manifold. Conversely, a ’robustness gap’ emerges in Vision Transformer (ViT) and MobileNet paradigms. This suggests that the global self-attention mechanisms of ViTs are inherently more resilient to the locally-optimized, hull-aligned perturbations intended for convolutional kernels—a finding that highlights architectural diversity as a critical defense for maritime surveillance.
Figure 6. Patches generated by different training strategies.
Figure 6. Patches generated by different training strategies.
Preprints 232964 g006

5.5. Comparison of Patch Optimization Strategies

We evaluated the training strategies described in Section 3.7 with a 40% coverage. Table 6 shows that the fixed-weights strategy performs best. The GradNorm and stochastic variants also show promise, though they are less aggressive.
GradNorm helped identify suitable coefficients for each loss, and the results indicate that U-Net–based optimization outperforms direct optimization (Adam) in both control and effectiveness. These findings underscore the importance of structured, multi-loss balancing when optimizing adversarial patches.

5.6. Effect of Patch Shape

We evaluate the influence of patch shape on adversarial efficacy using a Mask R-CNN and U-Net pipeline (Table 7). The oval (elliptical) patch proved most effective, reducing segmentation F1 to 30.6 and detection F1 to 16.6. This suggests that elongated contours more effectively disrupt the spatial context and boundary features prioritized by these architectures. In contrast, circular and rectangular geometries were less disruptive, with the rectangular patch preserving the highest performance across both tasks. These results indicate that patch geometry is a critical factor in adversarial robustness for shape-sensitive models.

5.7. Large Patch vs. Small Patches

Covering 60% of a ship’s surface with a single monolithic patch may be impractical due to complex superstructures. We therefore compare a single-patch strategy against multi-patch configurations (comprising two or three disjoint patches) while maintaining a fixed total coverage budget of 60% to ensure a fair evaluation. Figure 7 shows that distributed layouts achieve slightly higher adversarial effectiveness (Table 8), suggesting that multi-patch designs not only accommodate physical structural constraints but may also disrupt feature extraction more effectively, making them a preferred choice for practical camouflage applications.

5.8. Ablation Study

The proposed method consists of three components: the U-Net generator G to generate the adversarial patch, the log-adversarial loss { L i adv , i ∈ Standard loss }, and the feature inconsistence loss L feat adv . We performed an ablation study to investigate the contribution of each component. In experiments, when the log-adversarial loss was omitted, the traditional negative adversarial loss was employed. Table 9 demonstrates that including each component improves the proposed model for camouflage attack, and the most significant component is the U-Net generator model, indicating that utilizing a U-Net to generate the adversarial patch is much more effective than optimizing the patch directly.

6. Conclusions

This work introduces a geometry-aware and resolution-robust adversarial attack framework for maritime surveillance systems. The proposed method generates hull-aligned adversarial textures using a stabilized multi-objective optimization (SMO) strategy, enabling a single attack to simultaneously disrupt detection, localization, and segmentation tasks. To address conflicts between different network objectives, the framework employs a logarithmic-inversion loss and dynamic gradient balancing, improving attack stability across varying image scales and resolutions. Experiments on multiple remote sensing benchmarks show that the approach significantly outperforms conventional attacks targeting only detection or segmentation models. The results reveal major security vulnerabilities in modern multi-task maritime monitoring systems and highlight the need for stronger geometry-aware defense strategies.

Author Contributions

Conceptualization, O.R.R., K.A.I., J.L., and H.W.; methodology, O.R.R., H.W., and J.L.; software, O.R.R., and K.A.I.; validation, O.R.R. and K.A.I.; formal analysis, O.R.R.; investigation, O.R.R. and L.Z.; resources, O.R.R, J.L., C.X., and H.W.; data curation, O.R.R.; writing—original draft preparation, O.R.R and J.L.; writing—review and editing, O.R.R, K.A.I., L.Z., R.N., C.X., H.W., and J.L.; visualization, O.R.R.; supervision, R.N., C.X., H.W. and J.L.; project administration, J.L.; funding acquisition, J.L. All authors have read and agreed to the published version of the manuscript.

Appendix A.

Appendix A.1. Baseline Competitors and Comparative Objectives

We benchmark our approach against general-purpose adversarial patch attacks:
Brown et al. [8] (classification-based attack): Brown et al. generate a universal adversarial patch under the Expectation over Transformation (EoT) framework T . For our ship-detection setting, we adapt the original image-classification objective to binary ship/background segmentation:
min p E T ∼ T L cls f ( T ( P ( x , p ) ) ) , y
where p is the adversarial patch, P ( x , p ) overlays p onto image x, T denotes random physical transformations, f ( · ) is the victim model, and y is the target ship label.
DPatch [11] (object-detection attack): DPatch suppresses detector objectness/confidence scores by directly optimizing patch pixels under EoT:
min p E T ∼ T L obj f ( T ( P ( x , p ) ) ) , y
where L obj minimizes objectness or detection confidence for true ship instances.
ShapeShifter [34] (detection and localization attack): ShapeShifter jointly attacks detector classification and localization objectives:
min p E T ∼ T [ L cls f ( T ( P ( x , p ) ) ) , y cls + λ loc L loc f ( T ( P ( x , p ) ) ) , y loc ]
where L loc is the bounding-box regression loss and λ loc balances classification and localization objectives.
Scale-Adaptive [23] (multi-scale detection attack): The scale-adaptive method improves robustness across object sizes by sampling scales s ∼ S during optimization:
min p E s ∼ S E T ∼ T L task f ( T s ( P s ( x , p ) ) ) , y
where S is the empirical object-scale distribution, and T s , P s denote scale-conditioned transformations and patch placement operations.
Camouflage Attacks for Aerial Imagery [21,24,25,35]: These methods extend detector-oriented patch attacks by incorporating physical-realism constraints such as smoothness and printability. All methods optimize adversarial patches against a pretrained YOLO detector [40]:
min p E s ∼ S E T ∼ T [ L task f ( T s ( P s ( x , p ) ) ) , y + λ TV L TV ( p ) + λ NPS L NPS ( p ) ]
where L task includes detector classification and box-regression losses, L TV enforces spatial smoothness, L NPS improves printability, and λ TV , λ NPS control their relative importance.

Appendix A.2. Dataset Statistics and Examples

We evaluated our method on three datasets covering ships, airplanes, and vehicles. Figure A1, Figure A2 and Figure A3 show representative examples with bounding-box annotations.
Figure A1. Examples from the Ship dataset with detected ships.
Figure A1. Examples from the Ship dataset with detected ships.
Preprints 232964 g0a1
Figure A2. Examples from the DOTA dataset with detected planes.
Figure A2. Examples from the DOTA dataset with detected planes.
Preprints 232964 g0a2
Figure A3. Examples from VisDrone with detected cars.
Figure A3. Examples from VisDrone with detected cars.
Preprints 232964 g0a3

Appendix A.3. Hyperparameters for Model Pretraining and Adversarial Patch Training

Table A1 summarizes the hyperparameters used to train all pre-trained models. Most networks are trained using a fixed learning rate of 1 e − 4 . For deeper or higher-capacity models, we adopt the WarmUpMultiStepLR schedule, which starts at 1 e − 6 and increases by a factor of 10 after the first 1000 iterations until reaching 1 e − 4 . This stabilizes early optimization for large architectures.
In Section 3.7, we described different adversarial patch training strategies, and Section 5.5 presented their performance. Table A2 lists the coefficients for the different strategies.
Table A1. Hyperparameters used for model pretraining. Larger models used the WarmUpMultiStepLR schedule.
Table A1. Hyperparameters used for model pretraining. Larger models used the WarmUpMultiStepLR schedule.
Model Structure Backbone # Parameters Learning Rate
UNet ResNet34 24M 1 e − 4
FPN ResNet34 23M 1 e − 4
Mask R-CNN FPN ResNet101 101M WarmupMultiStepLR
Faster R-CNN FPN RetinaNet50 38M WarmupMultiStepLR
DeepLabV3+ MobileNet 4M 1 e − 4
Faster R-CNN C4 ResNet50 32M WarmupMultiStepLR
R-CNN FPN ResNetXt 142M WarmupMultiStepLR
PAN EfficientNet 6M 1 e − 4
Mask-RCNN MViTV2-T 43M WarmupMultiStepLR
Mask R-CNN ViTDet-B 110M WarmupMultiStepLR
Cascade Mask R-CNN Swin-B 139M WarmupMultiStepLR
Table A2. Coefficients of the Loss Components for Adversarial Patch Training. Values in GradNorm represent the converged coefficients.
Table A2. Coefficients of the Loss Components for Adversarial Patch Training. Values in GradNorm represent the converged coefficients.
Training Strategy α rpn - obj α rpn - reg α feat α roi - cls α roi - reg α mask α seg
Uniform Weighting 1.0 1.0 1.0 1.0 1.0 1.0 1.0
Static Heuristics 1.0 1e-3 1e-3 1.0 1e-3 1e-2 1e-2
GradNorm 1.0 1e-8 1e-8 1e-8 1e-8 1e-8 1e-8

Appendix A.4. Additional Results on the Transferability of Adversarial Patches

We investigated the transferability of adversarial patches optimized with different models. In this experiment, we obtained three adversarial patches by optimizing them with Faster R-CNN, U-Net, and jointly with both models, respectively. Note that Faster R-CNN detects objects through region proposals, whereas U-Net performs detection via segmentation. Using these three adversarial patches, we attacked both the Faster R-CNN and U-Net models. Table A3 and Table A4 list the performance of the camouflage attacks. We observe that the jointly trained patch yields a more balanced degradation in detection performance for both the Faster R-CNN and U-Net models. In contrast, patches trained solely on Faster R-CNN show limited transferability to the segmentation model, while U-Net–trained patches heavily degrade segmentation but do not reliably affect detection.
Table A3. Performances of camouflage attacks on Faster R-CNN using patches optimized with different models, where “Joint Optimization” denotes that the patch was optimized using both Faster R-CNN and U-Net.
Table A3. Performances of camouflage attacks on Faster R-CNN using patches optimized with different models, where “Joint Optimization” denotes that the patch was optimized using both Faster R-CNN and U-Net.
Applied Patch Precision Recall F1 mIoU
No patch 76.7 96.6 85.5 87.8
Patch with Faster R-CNN 31.7 39.0 34.9 39.3
Patch with U-Net 78.5 95.8 86.3 82.0
Joint Optimization 25.9 36.2 30.2 38.4
Table A4. Performances of Camouflage Attacks on U-Net Using Patches Optimized with Different Models.
Table A4. Performances of Camouflage Attacks on U-Net Using Patches Optimized with Different Models.
Applied Patch Precision Recall F1 mIoU
No patch 86.9 91.9 88.2 80.9
Patch with Faster R-CNN 82.2 39.3 49.7 35.0
Patch with U-Net 42.1 1.8 8.9 1.82
Joint Optimization 74.1 19.8 32.8 17.8

Appendix A.5. Additional Camouflage Attack Examples

Figure A5 presents additional ship-detection results by Mask R-CNN and U-Net under camouflage attack using the adversarial patch jointly optimized by the proposed model, illustrating successful, partially successful, and failed adversarial cases. Figure A4 shows examples from the DOTA dataset where YOLOv2 is subjected to camouflage attacks, with the adversarial patch optimized by the proposed method.

Appendix A.6. Discussions

Results in appendix strengthen the core findings of the paper. The dataset examples further demonstrate that the method generalizes across object types and imaging conditions, reinforcing the practical relevance of the proposed adversarial design.
Figure A4. Detection results on DOTA images. The first row shows YOLOv2 predictions, and the second row shows Mask R-CNN predictions, for both clean and adversarial inputs.
Figure A4. Detection results on DOTA images. The first row shows YOLOv2 predictions, and the second row shows Mask R-CNN predictions, for both clean and adversarial inputs.
Preprints 232964 g0a4
Figure A5. Camouflage attack examples for Mask R-CNN (detection) and U-Net (segmentation) on clean and adversarial ship images.
Figure A5. Camouflage attack examples for Mask R-CNN (detection) and U-Net (segmentation) on clean and adversarial ship images.
Preprints 232964 g0a5

References

  1. ESRI. A Look Back: How a GIS Team Guided Response and Recovery After 9/11. https://www.esri.com/about/newsroom/blog/how-maps-guided-9-11-response-and-recovery, 2021. Accessed: 12 Nov. 2025.
  2. He, K.; Gkioxari, G.; Dollár, P.; Girshick, R. Mask r-cnn. In Proceedings of the Proceedings of the IEEE international conference on computer vision, 2017, pp. 2961–2969.
  3. Ronneberger, O.; Fischer, P.; Brox, T. U-net: Convolutional networks for biomedical image segmentation. In Proceedings of the Medical image computing and computer-assisted intervention–MICCAI 2015: 18th international conference, Munich, Germany, October 5-9, 2015, proceedings, part III 18. Springer, 2015, pp. 234–241.
  4. Biggio, B.; Corona, I.; Maiorca, D.; Nelson, B.; Šrndić, N.; Laskov, P.; Giacinto, G.; Roli, F. Evasion attacks against machine learning at test time. In Proceedings of the Joint European conference on machine learning and knowledge discovery in databases. Springer, 2013, pp. 387–402.
  5. Szegedy, C.; Zaremba, W.; Sutskever, I.; Bruna, J.; Erhan, D.; Goodfellow, I.; Fergus, R. Intriguing properties of neural networks. arXiv preprint arXiv:1312.6199 2013.
  6. Kurakin, A.; Goodfellow, I.; Bengio, S. Adversarial machine learning at scale. arXiv preprint arXiv:1611.01236 2016.
  7. Nesti, F.; Rossolini, G.; Nair, S.; Biondi, A.; Buttazzo, G. Evaluating the robustness of semantic segmentation for autonomous driving against real-world adversarial patch attacks. In Proceedings of the Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, 2022, pp. 2280–2289.
  8. Brown, T.B.; Mané, D.; Roy, A.; Abadi, M.; Gilmer, J. Adversarial patch. arXiv preprint arXiv:1712.09665 2017.
  9. Thys, S.; Van Ranst, W.; Goedemé, T. Fooling automated surveillance cameras: Adversarial patches to attack person detection. In Proceedings of the Proceedings of the IEEE/CVF conference on computer vision and pattern recognition workshops, 2019, pp. 0–0.
  10. Song, D.; Eykholt, K.; Evtimov, I.; Fernandes, E.; Li, B.; Rahmati, A.; Tramer, F.; Prakash, A.; Kohno, T. Physical adversarial examples for object detectors. In Proceedings of the 12th USENIX workshop on offensive technologies (WOOT 18), 2018.
  11. Liu, X.; Yang, H.; Liu, Z.; Song, L.; Li, H.; Chen, Y. Dpatch: An adversarial patch attack on object detectors. arXiv preprint arXiv:1806.02299 2018.
  12. Saha, A.; Subramanya, A.; Patil, K.; Pirsiavash, H. Role of spatial context in adversarial robustness for object detection. In Proceedings of the Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops, 2020, pp. 784–785.
  13. Ren, S.; He, K.; Girshick, R.; Sun, J. Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks. In Proceedings of the Advances in Neural Information Processing Systems; Cortes, C.; Lawrence, N.; Lee, D.; Sugiyama, M.; Garnett, R., Eds. Curran Associates, Inc., 2015, Vol. 28.
  14. Redmon, J.; Farhadi, A. Yolov3: An incremental improvement. arXiv preprint arXiv:1804.02767 2018.
  15. Zhang, Y.; Foroosh, H.; David, P.; Gong, B. CAMOU: Learning physical vehicle camouflages to adversarially attack detectors in the wild. In Proceedings of the International Conference on Learning Representations, 2018.
  16. Wang, D.; Jiang, T.; Sun, J.; Zhou, W.; Gong, Z.; Zhang, X.; Yao, W.; Chen, X. Fca: Learning a 3d full-coverage vehicle camouflage for multi-view physical adversarial attack. In Proceedings of the Proceedings of the AAAI conference on artificial intelligence, 2022, Vol. 36, pp. 2414–2422.
  17. Wen, H.; Chang, S.; Zhou, L. Light projection-based physical-world vanishing attack against car detection. In Proceedings of the ICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2023, pp. 1–5.
  18. Wei, X.; Yu, J.; Huang, Y. Infrared adversarial patches with learnable shapes and locations in the physical world. International Journal of Computer Vision 2024, 132, 1928–1944.
  19. Hu, Z.; Chu, W.; Zhu, X.; Zhang, H.; Zhang, B.; Hu, X. Physically realizable natural-looking clothing textures evade person detectors via 3d modeling. In Proceedings of the Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023, pp. 16975–16984.
  20. Jing, P.; Tang, Q.; Du, Y.; Xue, L.; Luo, X.; Wang, T.; Nie, S.; Wu, S. Too good to be safe: Tricking lane detection in autonomous driving with crafted perturbations. In Proceedings of the 30th USENIX Security Symposium (USENIX Security 21), 2021, pp. 3237–3254.
  21. Zhang, Y.; Zhang, Y.; Qi, J.; Bin, K.; Wen, H.; Tong, X.; Zhong, P. Adversarial Patch Attack on Multi-Scale Object Detection for UAV Remote Sensing Images. Remote Sensing 2022, 14. [CrossRef]
  22. Deng, B.; Zhang, D.; Dong, F.; Zhang, J.; Shafiq, M.; Gu, Z. Rust-Style Patch: A Physical and Naturalistic Camouflage Attacks on Object Detector for Remote Sensing Images. Remote Sensing 2023, 15. [CrossRef]
  23. Lu, M.; Li, Q.; Chen, L.; Li, H. Scale-adaptive adversarial patch attack for remote sensing image aircraft detection. Remote Sensing 2021, 13, 4078.
  24. Pan, Y.; Wang, H. ShipCamou: Adversarial camouflage against optical remote sensing image ship detector. In Proceedings of the 1st Aerospace Frontiers Conference (AFC 2024), 2024.
  25. Liu, C.; Ding, P.; Zheng, Z.; Wang, H.; Zhu, B.; Xu, T.; Han, Z.; Wang, J. Adversarial Patch Attack for Ship Detection via Localized Augmentation. arXiv preprint arXiv:2508.21472 2025.
  26. Goodfellow, I.; Pouget-Abadie, J.; Mirza, M.; Xu, B.; Warde-Farley, D.; Ozair, S.; Courville, A.; Bengio, Y. Generative adversarial nets. In Proceedings of the Advances in Neural Information Processing Systems (NeurIPS), 2014, Vol. 27.
  27. Poursaeed, O.; Katsman, I.; Gao, B.; Belongie, S. Generative adversarial perturbations. In Proceedings of the Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2018, pp. 4422–4431.
  28. Xiao, C.; Li, B.; Zhu, J.Y.; He, W.; Nie, M.; Song, D. Generating adversarial examples with adversarial networks. In Proceedings of the Proceedings of the Twenty-Seventh International Joint Conference on Artificial Intelligence (IJCAI), 2018, pp. 3905–3911.
  29. Qiu, H.; Xiao, C.; Yang, L.; Li, B.; Guo, J.; Li, M. SemanticAdv: Generating adversarial examples via attribute-conditioned image editing. In Proceedings of the European Conference on Computer Vision (ECCV). Springer, 2020, pp. 19–37.
  30. Athalye, A.; Engstrom, L.; Ilyas, A.; Kwok, K. Synthesizing robust adversarial examples. In Proceedings of the International conference on machine learning. PMLR, 2018, pp. 284–293.
  31. Yu, T.; Kumar, S.; Gupta, A.; Levine, S.; Hausman, K. Gradient surgery for multi-task learning. In Proceedings of the Advances in Neural Information Processing Systems (NeurIPS), 2020, Vol. 33, pp. 5824–5836.
  32. Kendall, A.; Gal, Y.; Cipolla, R. Multi-Task Learning Using Uncertainty to Weigh Losses for Scene Geometry and Semantics. In Proceedings of the Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2018, pp. 7482–7491.
  33. Chen, Z.; Badrinarayanan, V.; Lee, C.Y.; Rabinovich, A. GradNorm: Gradient normalization for adaptive loss balancing in deep multitask networks. In Proceedings of the International Conference on Machine Learning (ICML). PMLR, 2018, pp. 794–803.
  34. Chen, S.T.; Cornelius, C.; Martin, J.; Chau, D.H. Shapeshifter: Robust physical adversarial attack on faster r-cnn object detector. In Proceedings of the Machine Learning and Knowledge Discovery in Databases: European Conference, ECML PKDD 2018, Dublin, Ireland, September 10–14, 2018, Proceedings, Part I 18. Springer, 2019, pp. 52–68.
  35. Den Hollander, R.; Adhikari, A.; Tolios, I.; van Bekkum, M.; Bal, A.; Hendriks, S.; Kruithof, M.; Gross, D.; Jansen, N.; Perez, G.; et al. Adversarial patch camouflage against aerial detection. In Proceedings of the Artificial intelligence and machine learning in defense applications II. SPIE, 2020, Vol. 11543, pp. 77–86.
  36. Competition, K. Airbus Ship Detection Challenge, 2024.
  37. Xia, G.S.; Bai, X.; Ding, J.; Zhu, Z.; Belongie, S.; Luo, J.; Datcu, M.; Pelillo, M.; Zhang, L. DOTA: A large-scale dataset for object detection in aerial images. In Proceedings of the Proceedings of the IEEE conference on computer vision and pattern recognition, 2018, pp. 3974–3983.
  38. Zhu, P.; Wen, L.; Du, D.; Bian, X.; Fan, H.; Hu, Q.; Ling, H. Detection and tracking meet drones challenge. IEEE Transactions on Pattern Analysis and Machine Intelligence 2021, 44, 7380–7399.
  39. Microsoft. Welcome to detectron2’s documentation!, 2024.
  40. Redmon, J.; Farhadi, A. YOLO9000: Better, faster, stronger. In Proceedings of the Proceedings of the IEEE conference on computer vision and pattern recognition, 2017, pp. 7263–7271.
Figure 3. Example of the Airbus dataset
Figure 3. Example of the Airbus dataset
Preprints 232964 g003
Figure 4. ship count distribution per image in Airbus dataset.
Figure 4. ship count distribution per image in Airbus dataset.
Preprints 232964 g004
Figure 5. Adversarial patches optimized by different methods.
Figure 5. Adversarial patches optimized by different methods.
Preprints 232964 g005
Figure 7. A ship with a single large patch, two patches, and three patches with a fixed coverage of 60% during camouflage attack.
Figure 7. A ship with a single large patch, two patches, and three patches with a fixed coverage of 60% during camouflage attack.
Preprints 232964 g007
Table 1. Dataset statistics. Although the native resolutions vary, all inputs are resized to the corresponding dimensions listed below.
Table 1. Dataset statistics. Although the native resolutions vary, all inputs are resized to the corresponding dimensions listed below.
Dataset #Train data #Train objects #Val data #Val objects Input size Object type
Airbus 40,428 70,964 2,128 7,915 768×768×3 Ship
DOTA 2,186 16,426 673 5,298 1024×1024×3 Plane
VisDrone 6,133 144,867 515 14,064 640×640×3 Car
Table 2. Comparison with baselines on Airbus. APs, APm, and AP1 denote AP for small, medium, and large ships.
Table 2. Comparison with baselines on Airbus. APs, APm, and AP1 denote AP for small, medium, and large ships.
Method AP AP50 AP 75 AP s AP m AP l mIoU
No Patch 75.9 96.2 87.9 63.8 79.6 83.7 87.8
Brown [8] 42.9 80.9 38.6 18.3 44.4 71.8 73.5
ShapeShifter [34] 43.0 80.9 38.8 18.6 44.7 72.0 73.6
Dpatch [11] 14.9 18.7 16.6 0.0 8.1 71.3 49.4
ScaleAdaptive [23] 20.5 33.5 21.6 17.0 23.9 20.3 50.3
ShipCamou [24] 16.2 31.5 14.3 14.6 29.2 0.9 50.2
Chun Liu [25] 17.7 33.1 17.4 14.9 33.7 0.2 51.5
SMO 3.0 11.3 1.5 5.3 3.0 0.0 40.7
Table 3. Performance with increasing patch coverage on ships for the detection-based algorithm under camouflage attack.
Table 3. Performance with increasing patch coverage on ships for the detection-based algorithm under camouflage attack.
Coverage (%) 0 25 50 75 100
Precision 55.0 29.8 11.1 4.7 2.3
Recall 96.4 78.6 30.8 10.6 4.4
F1 Score 70.0 43.2 16.3 6.5 3.0
mIoU 86.4 71.0 32.0 16.3 10.3
Table 4. Transferability results on DOTA and VisDrone. Adversarial patches are generated using SMO with a Mask R-CNN and U-Net source pipeline and evaluated on YOLO-based victim detectors. ‘Pre’ and ‘Post’ denote AP 50 before and after the attack. ‘Reduction’ reports the relative drop in AP 50 .
Table 4. Transferability results on DOTA and VisDrone. Adversarial patches are generated using SMO with a Mask R-CNN and U-Net source pipeline and evaluated on YOLO-based victim detectors. ‘Pre’ and ‘Post’ denote AP 50 before and after the attack. ‘Reduction’ reports the relative drop in AP 50 .
Model Patch Method Dataset AP 50 (Pre) AP 50 (Post) Reduction (%)
YOLOv2 Den et al. [35] DOTA (Planes) 88.0 37.7 57.1
YOLOv2 SMO DOTA (Planes) 88.0 6.0 93.1
YOLOv3 Zhan et al. [21] VisDrone (Vehicles) 99.0 59.7 39.6
YOLOv3 SMO VisDrone (Vehicles) 99.0 15.5 84.3
Table 5. Transferability of a camouflage adversarial patch optimized on Mask R-CNN with a ResNet101 backbone. The patch is evaluated on different architectures and backbones across instance segmentation, object detection, and semantic segmentation tasks. ‘Pre’ and ‘Post’ denote performance before and after the attack, respectively.
Table 5. Transferability of a camouflage adversarial patch optimized on Mask R-CNN with a ResNet101 backbone. The patch is evaluated on different architectures and backbones across instance segmentation, object detection, and semantic segmentation tasks. ‘Pre’ and ‘Post’ denote performance before and after the attack, respectively.
Task Model Backbone Pre Post
F1 mIoU F1 mIoU
Instance Seg. Mask R-CNN ResNet101 70.2 86.4 16.3 28.5
Mask R-CNN MViT2-T 85.5 87.8 42.3 54.6
Mask R-CNN ViTDet-B 94.0 84.9 89.7 80.2
Detection R-CNN FPN ResNeXt 68.4 85.7 31.4 66.8
Faster R-CNN RetinaNet50 85.9 83.0 30.9 22.0
Faster R-CNN C4 ResNet50 48.7 84.0 32.5 62.2
Segmentation U-Net ResNet34 88.2 80.9 30.6 16.4
DeepLabV3+ MobileNet 84.5 75.2 77.6 66.0
PAN EfficientNet 81.6 70.7 77.7 63.3
FPN ResNet50 79.4 68.5 53.6 39.7
Table 6. Impact of patch optimization strategies at 40% coverage per object. Patch updates in the last five rows are guided by a U-Net architecture using different multi-loss balancing methods.
Table 6. Impact of patch optimization strategies at 40% coverage per object. Patch updates in the last five rows are guided by a U-Net architecture using different multi-loss balancing methods.
Strategy Precision Recall AP mIoU
Pre-attack 55.2 96.4 73.3 86.4
Random Patch 56.0 97.9 64.3 84.1
Primitive Optimization 16.4 38.3 3.7 34.0
Uniform Weighting 11.5 29.5 2.3 29.5
Cyclic Partitioning 11.1 30.8 2.7 32.0
Adaptive Stochasticity 49.7 82.3 48.5 69.5
GradNorm 24.1 56.1 11.5 54.3
Static Heuristics 8.1 25.3 1.1 28.5
Table 7. Impact of different patch shapes on performance after attacks.
Table 7. Impact of different patch shapes on performance after attacks.
Shape Segmentation Detection
F1 mIoU F1 mIoU
Pre-attack 88.2 80.9 70.2 86.4
Circle 54.5 32.1 17.2 40.6
Rectangle 50.7 35.3 40.2 50.2
Oval 30.6 16.4 16.3 28.5
Table 8. Effect of splitting a single large patch into multiple smaller patches on post-attack detection performance. Total patched area is fixed at 60% of the ship.
Table 8. Effect of splitting a single large patch into multiple smaller patches on post-attack detection performance. Total patched area is fixed at 60% of the ship.
Patches AP AP 50 AP 75 AP s AP m AP l mIoU
One (Large) 13.1 27.9 11.4 13.8 14.8 1.4 45.0
Two (Small) 7.9 22.5 4.3 7.6 10.4 0.2 36.0
Three (Small) 14.4 26.2 13.9 7.8 17.7 3.3 31.0
Table 9. Ablation study on G, L i adv , and L feat adv in the proposed SMO method.
Table 9. Ablation study on G, L i adv , and L feat adv in the proposed SMO method.
G L i adv L feat adv AP AP 50 AP 75 AP s AP m AP l
✗ ✓ ✓ 26.3 43.2 27.6 26.5 29.8 7.2
✓ ✗ ✓ 17.2 36.4 15.1 2.5 23.7 5.6
✓ ✓ ✗ 13.1 27.9 11.4 13.8 14.8 1.4
✓ ✓ ✓ 8.5 16.5 7.7 10.3 10.3 0.5
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.