Submitted:
01 October 2024
Posted:
01 October 2024
You are already at the latest version
Abstract

Keywords:
1. Introduction
- Defensive distillation uses knowledge distillation to train a student model using soft labels obtained from a teacher model, reducing model’s sensitivity to noise. Initially seen as strong, it was later shown to be insufficient against subtle and powerful attacks such as the C&W attack [13].
- Randomized smoothing is the application of differential privacy to adversarial robustness. By adding random noise to inputs and averaging predictions, it allows to reduce the model's sensitivity to small perturbations. Despite being considered as a reliable defense, it offers limited robustness against specific norm-bounded attacks, and efforts are ongoing to address this issue [37].
- The proposed MP method leverages auxiliary tasks related to the main task within a multi-task learning model to generate perturbations, maintaining the primary directionality needed to improve the adversarial accuracy of the main task. At the same time, directional diversity is injected through a random weighted summation of the main task perturbation and those from the auxiliary tasks.
- In line with our previous work [43], the proposed method is applicable even when no auxiliary task is given for multi-task learning and provides better robustness in addition to improving the generalization performance of the multi-task model.
- We conduct the experiments on five benchmark datasets, demonstrating that the proposed method can improve adversarial robustness (adversarial accuracy) and recognition performance on original data (clean accuracy) in several attack methods and datasets. We also analyzed the relationship between the characteristics of data distribution and the diversity of perturbation directions, as well as their impact on the model's performance.
2. Materials and Methods
2.1. Generating Mixed Perturbation
2.2. Adversarial Training of Victim models
2.2.1. Adversarial Training using Mixed Perturbation
2.2.2. Adversarial Training using Multi-task Attack
3. Results
3.1. Datasets and Victim Models
3.2. Attack Methods
3.3. Metrics
3.4. Experimental Results and Analysis
3.4.1. Experimental Results with Mixed Perturbation
3.4.2. Analysis on the Effect of Mixed Perturbation
3.4.3. Comparison of Three Perturbation Generation Methods
4. Discussion
5. Conclusions
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Conflicts of Interest
References
- Sarker, I.H. Deep learning: A comprehensive overview on techniques, taxonomy, applications and research directions. SN Computer Science, 2021, 2(6), 420.
- Gambín, Á.F.; Yazidi, A.; Vasilakos, A.; Haugerud, H.; Djenouri, Y. Deepfakes: Current and future trends. Artificial Intelligence Review, 2024, 57(3), 64.
- Marchal, N.; Xu, R.; Elasmar, R.; Gabriel, I.; Goldberg, B.; Isaac, W. Generative AI misuse: A taxonomy of tactics and insights from real-world data. arXiv 2024, arXiv:2406.13843. [Google Scholar]
- Anderljung, M.; Hazell, J. Protecting society from AI misuse: When are restrictions on capabilities warranted? arXiv 2023, arXiv:2303.09377. [Google Scholar]
- Madiega, T. Artificial intelligence act. European Parliament: European Parliamentary Research Service, 2021.
- Szegedy, C.; Zaremba, W.; Sutskever, I.; Bruna, J.; Erhan, D.; Goodfellow, I.; Fergus, R. Intriguing properties of neural networks. In 2nd International Conference on Learning Representations, 2014, January.
- Nguyen, A.; Yosinski, J.; Clune, J. Deep neural networks are easily fooled: High confidence predictions for unrecognizable images. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition; 2015; pp. 427–436. [Google Scholar]
- Hendrycks, D.; Carlini, N.; Schulman, J.; Steinhardt, J. Unsolved problems in ML safety. arXiv 2021, arXiv:2109.13916. [Google Scholar]
- Oprea, A.; Vassilev, A. Adversarial machine learning: A taxonomy and terminology of attacks and mitigations (No. NIST Artificial Intelligence (AI) 100-2 E2023 (Withdrawn)). National Institute of Standards and Technology, 2023.
- Goodfellow, I.J.; Shlens, J.; Szegedy, C. Explaining and harnessing adversarial examples. arXiv 2014, arXiv:1412.6572, 2014. [Google Scholar]
- Kurakin, A.; Goodfellow, I.J.; Bengio, S. Adversarial examples in the physical world. In Artificial Intelligence Safety and Security, 2018, pp. 99-112. Chapman and Hall/CRC.
- Madry, A.; Makelov, A.; Schmidt, L.; Tsipras, D.; Vladu, A. Towards deep learning models resistant to adversarial attacks. In International Conference on Learning Representations, 2018, February.
- Carlini, N.; Wagner, D. Towards evaluating the robustness of neural networks. In Proceedings of the IEEE Symposium on Security and Privacy (SP); 2017; pp. 39–57. [Google Scholar]
- Finlayson, S.G.; Bowers, J.D.; Ito, J.; Zittrain, J.L.; Beam, A.L.; Kohane, I.S. Adversarial attacks on medical machine learning. Science, 2019; 363, 1287–1289. [Google Scholar]
- Angelos, F.; Panagiotis, T.; Rowan, M.; Nicholas, R.; Sergey, L.; Yarin, G. Can autonomous vehicles identify, recover from, and adapt to distribution shifts? In Proceedings of the IEEE International Conference on Machine Learning, 2020, pp. 3145-3153, November.
- Papernot, N.; McDaniel, P.; Goodfellow, I.; Jha, S.; Celik, Z.B.; Swami, A. Practical black-box attacks against machine learning. In Proceedings of the ACM Asia Conference on Computer and Communications Security (ASIACCS), 2017, pp. 506-519, April.
- Yue, K.; Jin, R.; Wong, C.W.; Baron, D.; Dai, H. Gradient obfuscation gives a false sense of security in federated learning. In 32nd USENIX Security Symposium (USENIX Security 23), 2023, pp. 6381-6398.
- Papernot, N.; McDaniel, P.; Wu, X.; Jha, S.; Swami, A. Distillation as a defense to adversarial perturbations against deep neural networks. In 2016 IEEE Symposium on Security and Privacy (SP), 2016, pp. 582-597. IEEE.
- Kuang, H.; Liu, H.; Wu, Y.; Satoh, S.I.; Ji, R. Improving adversarial robustness via information bottleneck distillation. Advances in Neural Information Processing Systems, 2024, 36.
- Nesti, F.; Biondi, A.; Buttazzo, G. Detecting adversarial examples by input transformations, defense perturbations, and voting. IEEE Transactions on Neural Networks and Learning Systems 2021, 34, 1329–1341. [Google Scholar] [CrossRef] [PubMed]
- Prakash, N. Moran, S. Garber, A. DiLillo, and J. Storer, "Deflecting adversarial attacks with pixel deflection," in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2018, pp. 8571-8580.
- Chen, J.; Raghuram, J.; Choi, J.; Wu, X.; Liang, Y.; Jha, S. Revisiting adversarial robustness of classifiers with a reject option. In Proceedings of the AAAI Workshop on Adversarial Machine Learning Beyond, 2021, December.
- Crecchi, F.; Melis, M.; Sotgiu, A.; Bacciu, D.; Biggio, B. FADER: Fast adversarial example rejection. Neurocomputing 2022, 470, 257–268. [Google Scholar] [CrossRef]
- Aldahdooh, A.; Hamidouche, W.; Fezza, S.A.; Déforges, O. Adversarial example detection for DNN models: A review and experimental comparison. Artificial Intelligence, 2022, Review, pp. 1-60.
- Lecuyer, M.; Atlidakis, V.; Geambasu, R.; Hsu, D.; Jana, S. Certified robustness to adversarial examples with differential privacy. In Proceedings of the IEEE Symposium on Security and Privacy (SP), 2019, pp. 656-672, May.
- Wang, H.; Zhang, A.; Zheng, S.; Shi, X.; Li, M.; Wang, Z. Removing batch normalization boosts adversarial training. In Proceedings of the International Conference on Machine Learning (ICML), 2022, pp. 23433-23445, June.
- Zhang, H.; Yu, Y.; Jiao, J.; Xing, E.; El Ghaoui, L.; Jordan, M. Theoretically principled trade-off between robustness and accuracy. In International Conference on Machine Learning, 2019, pp. 7472-7482, May. PMLR.
- Wong, E.; Rice, L.; Kolter, J.Z. Fast is better than free: Revisiting adversarial training. arXiv 2020, arXiv:2001.03994. [Google Scholar]
- Athalye, A.; Carlini, N.; Wagner, D. Obfuscated gradients give a false sense of security: Circumventing defenses to adversarial examples. In Proceedings of the International Conference on Machine Learning (ICML), 2018, pp. 274-283, July.
- Tramer, F.; Carlini, N.; Brendel, W.; Madry, A. On adaptive attacks to adversarial example defenses. In Advances in Neural Information Processing Systems (NeurIPS), 2020, 33, pp. 1633-1645.
- Liu, Z.; Liu, Q.; Liu, T.; Xu, N.; Lin, X.; Wang, Y.; Wen, W. Feature distillation: DNN-oriented JPEG compression against adversarial examples. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2019, pp. 860-868, June.
- Xu, W. Feature squeezing: Detecting adversarial examples in deep neural networks. arXiv 2017, arXiv:1704.01155. [Google Scholar]
- Nesti, F.; Biondi, A.; Buttazzo, G. Detecting adversarial examples by input transformations, defense perturbations, and voting. IEEE Transactions on Neural Networks and Learning Systems, 2021, 34(3), 1329-1341.
- Chen, Y.; Zhang, M.; Li, J.; Kuang, X.; Zhang, X.; Zhang, H. Dynamic and diverse transformations for defending against adversarial examples. In 2022 IEEE International Conference on Trust, Security and Privacy in Computing and Communications (TrustCom), 2022, pp. 976-983, December.
- Klingner, M.; Kumar, V.R.; Yogamani, S.; Bär, A.; Fingscheidt, T. Detecting adversarial perturbations in multi-task perception. In 2022 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), 2022, pp. 13050-13057, October.
- Carlini, N.; Wagner, D. Adversarial examples are not easily detected: Bypassing ten detection methods. In Proceedings of the 10th ACM Workshop on Artificial Intelligence and Security (AISec), 2017, pp. 3-14, November.
- Pfrommer, S.; Anderson, B.G.; Sojoudi, S. Projected randomized smoothing for certified adversarial robustness. arXiv 2023, arXiv:2309.13794. [Google Scholar]
- Tsipras, D.; Santurkar, S.; Engstrom, L.; Turner, A.; Madry, A. Robustness may be at odds with accuracy. In Proceedings of the International Conference on Learning Representations (ICLR); 2019; pp. 1–24. [Google Scholar]
- H. Sicong et al.; "Interpreting adversarial examples in deep learning: A review," ACM Computing Surveys, 2023.
- Mao, C.; et al. Multitask learning strengthens adversarial robustness. In Proceedings of the European Conference on Computer Vision (ECCV); 2020; pp. 158–174.
- Huang, T.; Menkovski, V.; Pei, Y.; Wang, Y.; Pechenizkiy, M. Direction-aggregated attack for transferable adversarial examples. ACM Journal of Emerging Technologies in Computing Systems (JETC), 2022, 18(3), 1-22.
- Li, Z.; Yin, B.; Yao, T.; Guo, J.; Ding, S.; Chen, S.; Liu, C. Sibling-attack: Rethinking transferable adversarial attacks against face recognition. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition; 2023; pp. 24626–24637. [Google Scholar]
- Hyun, C.; Park, H. Multi-task learning with self-defined tasks for adversarial robustness of deep networks. IEEE Access 2024, 12, 83248–83259. [Google Scholar] [CrossRef]
- Hassani, H.; Javanmard, A. The curse of overparametrization in adversarial training: Precise analysis of robust generalization for random features regression. The Annals of Statistics, 2024, 52(2), 441-465.
- Wu, B.; Chen, J.; Cai, D.; He, X.; Gu, Q. Do wider neural networks really help adversarial robustness? Advances in Neural Information Processing Systems 2021, 34, 7054–7067. [Google Scholar]
- Lee, S.W.; Lee, R.; Seo, M.S.; Park, J.C.; Noh, H.C.; Ju, J.G.; ... & Choi, D.G. Multi-task learning with task-specific feature filtering in low-data condition. Electronics, 2021, 10(21), 2691.
- LeCun, Y.; Bottou, L.; Bengio, Y.; Haffner, P. Gradient-based learning applied to document recognition. Proceedings of the IEEE, 1998, 86(11), 2278-2324.
- Krizhevsky, A.; Hinton, G. Learning multiple layers of features from tiny images. 2009.
- Netzer, Y.; Wang, T.; Coates, A.; Bissacco, A.; Wu, B.; Ng, A.Y. Reading digits in natural images with unsupervised feature learning. In NIPS Workshop on Deep Learning and Unsupervised Feature Learning, 2011, vol. 2011, no. 5, pp. 7, December.
- Stallkamp, J.; Schlipsing, M.; Salmen, J.; Igel, C. Man vs. computer: Benchmarking machine learning algorithms for traffic sign recognition. Neural Networks, 2012, 32, 323–332. [Google Scholar] [CrossRef] [PubMed]
- Le, Y.; Yang, X. Tiny ImageNet Visual Recognition Challenge. CS 231N, 2015, vol. 7, no. 7, pp. 3.
- Ma, X.; Niu, Y.; Gu, L.; Wang, Y.; Zhao, Y.; Bailey, J.; Lu, F. Understanding adversarial attacks on deep learning based medical image analysis systems. Pattern Recognition, 2021, 110, 107332. [Google Scholar] [CrossRef]
- Xiong, P.; Tegegn, M.; Sarin, J.S.; Pal, S.; Rubin, J. It Is All About Data: A Survey on the Effects of Data on Adversarial Robustness. ACM Computing Surveys, 2024, 56(7), 1-41.
- Ghamizi, S.; Cordy, M.; Papadakis, M.; Le Traon, Y. Adversarial robustness in multi-task learning: Promises and illusions. In Proceedings of the AAAI Conference on Artificial Intelligence, 2022, 36(1), pp. 697-705, June.
- Deng, J.; Guo, J.; Xue, N.; Zafeiriou, S. Arcface: Additive angular margin loss for deep face recognition. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2019, (pp. 4690-4699).







| Type of Perturbation | STA [12] | MTA [40] | MP |
|---|---|---|---|
| Illustration | ![]() |
![]() |
![]() |
| Perturbation formula |
|
| Dataset | Train | Test | Class | Channel | Size | Auxiliary task 1 | Auxiliary task 2 | Network |
|---|---|---|---|---|---|---|---|---|
| MNIST | 60,000 | 10,000 | 10 | Gray | 28x28 | odd, even | prime, composite | LeNet |
| SVHN | 73,257 | 26,032 | 10 | RGB | 32x32 | odd, even | prime, composite | ResNet-18 |
| GTSRB | 39,209 | 12,630 | 43 | RGB | 112x112 | circle, polygon | character, symbol | AlexNet |
| CIFAR10 | 50,000 | 10,000 | 10 | RGB | 32x32 | animal, vehicle | sky, ground, water | WRN-32 |
| Tiny ImageNet | 100,000 | 10,000 | 200 | RGB | 64x64 | natural, artefact | animal, machine, others | ResNet-18 |
| Dataset | Inter-class variance |
|---|---|
| MNIST | 10.3211 |
| SVHN | 14.4643 |
| GTSRB | 67.2574 |
| CIFAR10 | 19.0552 |
| Tiny ImageNet | 19.2206 |
| Dataset | Task Combination |
Clean samples | PGD targeted | PGD untargeted | CW untargeted | ||||
|---|---|---|---|---|---|---|---|---|---|
| STA trained |
MP trained |
STA trained |
MP trained |
STA trained |
MP trained |
STA trained |
MP trained |
||
| MNIST | Main+Aux1 | 99.09 | 99.05 | 98.03 | 98.10 | 94.46 | 94.35 | 82.96 | 83.84 |
| Main+Aux2 | 98.95 | 99.08 | 98.05 | 97.96 | 94.82 | 94.24 | 87.16 | 81.90 | |
| Main+Aux1+Aux2 | 98.99 | 98.97 | 98.13 | 97.85 | 95.13 | 93.16 | 86.84 | 73.04 | |
| SVHN | Main+Aux1 | 90.42 | 90.31 | 69.99 | 70.68 | 54.26 | 52.32 | 56.14 | 53.64 |
| Main+Aux2 | 89.56 | 88.53 | 68.72 | 69.22 | 53.20 | 51.48 | 54.23 | 50.68 | |
| Main+Aux1+Aux2 | 90.05 | 91.05 | 70.82 | 70.36 | 54.18 | 53.33 | 54.34 | 53.98 | |
| GTSRB | Main+Aux1 | 90.05 | 92.90 | 76.28 | 75.23 | 58.72 | 63.48 | 57.44 | 54.03 |
| Main+Aux2 | 88.87 | 92.01 | 75.12 | 73.90 | 59.20 | 62.07 | 56.65 | 54.69 | |
| Main+Aux1+Aux2 | 89.83 | 94.81 | 76.18 | 73.64 | 59.94 | 68.04 | 56.48 | 51.53 | |
| CIFAR10 | Main+Aux1 | 84.62 | 83.34 | 68.63 | 70.44 | 45.23 | 48.12 | 39.78 | 41.56 |
| Main+Aux2 | 84.79 | 84.18 | 67.83 | 69.85 | 45.77 | 47.29 | 40.68 | 42.18 | |
| Main+Aux1+Aux2 | 84.94 | 82.57 | 67.85 | 69.77 | 46.68 | 47.33 | 42.13 | 40.68 | |
| Tiny ImageNet |
Main+Aux1 | 17.71 | 22.36 | 18.21 | 21.75 | 6.16 | 5.80 | 7.83 | 8.89 |
| Main+Aux2 | 17.88 | 21.43 | 18.04 | 20.33 | 5.51 | 5.15 | 7.57 | 8.10 | |
| Main+Aux1+Aux2 | 18.39 | 22.52 | 18.31 | 21.20 | 6.03 | 5.29 | 8.16 | 8.36 | |
| Dataset | ||||
|---|---|---|---|---|
| MNIST | 8.88 | 3.67 | 4.04 | 0.37 |
| SVHN | 13.63 | 1.93 | 2.23 | 0.30 |
| GTSRB | 3.72 | 4.39 | 6.45 | 2.06 |
| CIFAR10 | 5.15 | 3.01 | 4.08 | 1.07 |
| Tiny ImageNet | 5.07 | 2.88 | 3.81 | 0.93 |
| Dataset | Task Combination |
PGD targeted | PGD untargeted | ||||
|---|---|---|---|---|---|---|---|
| STA trained |
MTA trained |
MP trained |
STA trained |
MTA trained |
MP trained |
||
| MNIST | Main+Aux1 | 98.03 | 97.85 | 98.10 | 94.46 | 93.77 | 94.35 |
| Main+Aux2 | 98.05 | 97.83 | 97.96 | 94.82 | 93.90 | 94.24 | |
| Main+Aux1+Aux2 | 98.13 | 97.91 | 97.85 | 95.13 | 93.55 | 93.16 | |
| SVHN | Main+Aux1 | 69.99 | 70.33 | 70.68 | 54.26 | 51.78 | 52.32 |
| Main+Aux2 | 68.72 | 69.31 | 69.22 | 53.20 | 51.16 | 51.48 | |
| Main+Aux1+Aux2 | 70.82 | 71.20 | 70.36 | 54.18 | 51.50 | 53.33 | |
| GTSRB | Main+Aux1 | 76.28 | 76.00 | 75.23 | 58.72 | 59.84 | 63.48 |
| Main+Aux2 | 75.12 | 75.27 | 73.90 | 59.20 | 59.91 | 62.07 | |
| Main+Aux1+Aux2 | 76.18 | 74.60 | 73.64 | 59.94 | 59.44 | 68.04 | |
| CIFAR10 | Main+Aux1 | 67.83 | 68.69 | 70.44 | 45.23 | 45.82 | 48.12 |
| Main+Aux2 | 67.85 | 68.89 | 69.85 | 45.77 | 45.67 | 47.29 | |
| Main+Aux1+Aux2 | 68.64 | 68.94 | 69.77 | 46.68 | 46.05 | 47.33 | |
| Tiny ImageNet |
Main+Aux1 | 18.21 | 23.64 | 21.75 | 6.16 | 6.55 | 5.80 |
| Main+Aux2 | 18.04 | 22.21 | 20.33 | 5.51 | 7.45 | 5.15 | |
| Main+Aux1+Aux2 | 18.30 | 18.25 | 21.20 | 6.03 | 6.30 | 5.29 | |
| Dataset | Model | Main task | One Aux task | Two Aux tasks |
|---|---|---|---|---|
| MNIST | LeNet | 47.89 | 74.55 | 76.12 |
| SVHN | ResNet18 | 1352.04 | 2054.92 | 2248.94 |
| GTSRB | AlexNet | 502.66 | 533.75 | 659.65 |
| CIFAR10 | WRN-32 | 5032.58 | 7818.13 | 11393.31 |
| Tiny ImageNet | ResNet18 | 1290.05 | 2148.45 | 2795.09 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2024 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).


