Submitted:
30 June 2023
Posted:
30 June 2023
You are already at the latest version
Abstract
Keywords:
1. Introduction
2. Related Work
2.1. Deep Learning-Based Fusion Methods
2.2. Transformer
3. Proposed Method
3.1. Framework Overview
3.2. Infrared Feature Extraction Module
3.3. Visible Feature Extraction Module

3.4. Merge Module
3.5. Loss Function
4. Experiments
4.1. Experimental Configuration and Experimental Details
4.2. Comparison Methods and Evaluation Indicators
4.3. Ablation Experiments
4.3.1. Qualitative Comparisons
4.3.2. Quantitative Comparisons
4.4. Comparative Experiments
4.4.1. Qualitative Comparisons
4.4.2. Quantitative Comparisons
4.5. Generalization Experiments
4.5.1. Subjective Results
4.5.2. Objective Results
4.6. Detecting Performance
4.6.1. Subjective Results
4.6.2. Objective Results
| Methods | mAP@0.5 | mAP@0.9 | ||||
|---|---|---|---|---|---|---|
| Person | Car | Average | Person | Car | Average | |
| IR | 0.6307 | 0.3023 | 0.4665 | 0.2562 | 0.3013 | 0.2788 |
| VIS | 0.4953 | 0.7240 | 0.6096 | 0.1901 | 0.4358 | 0.3129 |
| ADF | 0.6935 | 0.7208 | 0.7072 | 0.2456 | 0.4505 | 0.3480 |
| IVFusion | 0.7288 | 0.7040 | 0.7164 | 0.1768 | 0.3733 | 0.2750 |
| GF | 0.6562 | 0.7300 | 0.6931 | 0.2415 | 0.4603 | 0.3509 |
| DenseFuse | 0.6915 | 0.7353 | 0.7134 | 0.2413 | 0.4425 | 0.3419 |
| DDcGAN | 0.4010 | 0.6968 | 0.5489 | 0.1072 | 0.3550 | 0.2316 |
| IFCNN | 0.7038 | 0.7305 | 0.7172 | 0.2541 | 0.4108 | 0.3324 |
| PMGI | 0.6990 | 0.6788 | 0.6889 | 0.2238 | 0.3448 | 0.2843 |
| GAN-FM | 0.7450 | 0.7548 | 0.7499 | 0.2409 | 0.4178 | 0.3293 |
| YDTR | 0.7149 | 0.5708 | 0.6428 | 0.2348 | 0.4972 | 0.3660 |
| Our | 0.7241 | 0.7388 | 0.7223 | 0.2458 | 0.4865 | 0.3661 |
5. Conclusions
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Conflicts of Interest
References
- Yang, Z.; Yu, W.; Liang, P.; Guo, H.; Xia, L.; Zhang, F.; Ma, Y.; Ma, J. Deep Transfer Learning for Military Object Recognition under Small Training Set Condition. Neural Comput & Applic 2019, 31, 6469–6478. [Google Scholar] [CrossRef]
- Ma, J.; Ma, Y.; Li, C. Infrared and Visible Image Fusion Methods and Applications: A Survey. Inf. Fusion 2019, 45, 153–178. [Google Scholar] [CrossRef]
- Schnelle, S.R.; Chan, A.L. Enhanced Target Tracking through Infrared-Visible Image Fusion. In Proceedings of the 14th International Conference on Information Fusion; 2011; pp. 1–8. [Google Scholar]
- Ma, W.; Wang, K.; Li, J.; Yang, S.X.; Li, J.; Song, L.; Li, Q. Infrared and Visible Image Fusion Technology and Application: A Review. Sensors 2023, 23, 599. [Google Scholar] [CrossRef]
- Zhang, L.; Yang, X.; Wan, Z.; Cao, D.; Lin, Y. A Real-Time FPGA Implementation of Infrared and Visible Image Fusion Using Guided Filter and Saliency Detection. Sensors 2022, 22, 8487. [Google Scholar] [CrossRef]
- Jia, W.; Song, Z.; Li, Z. Multi-Scale Fusion of Stretched Infrared and Visible Images. Sensors 2022, 22, 6660. [Google Scholar] [CrossRef]
- Liu, Y.; Wu, Z.; Han, X.; Sun, Q.; Zhao, J.; Liu, J. Infrared and Visible Image Fusion Based on Visual Saliency Map and Image Contrast Enhancement. Sensors 2022, 22, 6390. [Google Scholar] [CrossRef]
- Huang, Z.; Yang, B.; Liu, C. RDCa-Net: Residual Dense Channel Attention Symmetric Network for Infrared and Visible Image Fusion. Infrared Phys. Technol. 2023, 130, 104589. [Google Scholar] [CrossRef]
- Wang, H.; Wang, J.; Xu, H.; Sun, Y.; Yu, Z. DRSNFuse: Deep Residual Shrinkage Network for Infrared and Visible Image Fusion. Sensors 2022, 22, 5149. [Google Scholar] [CrossRef]
- Zheng, X.; Yang, Q.; Si, P.; Wu, Q. A Multi-Stage Visible and Infrared Image Fusion Network Based on Attention Mechanism. Sensors 2022, 22, 3651. [Google Scholar] [CrossRef]
- Liu, Y.; Wang, Z. Simultaneous Image Fusion and Denoising with Adaptive Sparse Representation. IET Image Processing 2015, 9, 347–357. [Google Scholar] [CrossRef]
- Yang, B.; Li, S. Visual Attention Guided Image Fusion with Sparse Representation. Optik 2014, 125, 4881–4888. [Google Scholar] [CrossRef]
- Bulanon, D.M.; Burks, T.F.; Alchanatis, V. Image Fusion of Visible and Thermal Images for Fruit Detection. Biosyst. Eng. 2009, 103, 12–22. [Google Scholar] [CrossRef]
- Yu, X.; Ren, J.; Chen, Q.; Sui, X. A False Color Image Fusion Method Based on Multi-Resolution Color Transfer in Normalization YCbCr Space. Optik 2014, 125, 6010–6016. [Google Scholar] [CrossRef]
- Cvejic, N.; Bull, D.; Canagarajah, N. Region-Based Multimodal Image Fusion Using ICA Bases. IEEE Sens. J. 2007, 7, 743–751. [Google Scholar] [CrossRef]
- Mitianoudis, N.; Stathaki, T. Pixel-Based and Region-Based Image Fusion Schemes Using ICA Bases. Inf. Fusion 2007, 8, 131–142. [Google Scholar] [CrossRef]
- Yin, M.; Duan, P.; Liu, W.; Liang, X. A Novel Infrared and Visible Image Fusion Algorithm Based on Shift-Invariant Dual-Tree Complex Shearlet Transform and Sparse Representation. Neurocomputing 2017, 226, 182–191. [Google Scholar] [CrossRef]
- Deng, J.; Xuan, X.; Wang, W.; Li, Z.; Yao, H.; Wang, Z. A Review of Research on Object Detection Based on Deep Learning. J. Phys.: Conf. Ser. 2020, 1684, 012028. [Google Scholar] [CrossRef]
- Tian, C.; Fei, L.; Zheng, W.; Xu, Y.; Zuo, W.; Lin, C.-W. Deep Learning on Image Denoising: An Overview. Neural Netw. 2020, 131, 251–275. [Google Scholar] [CrossRef]
- Liu, Y.; Chen, X.; Peng, H.; Wang, Z. Multi-Focus Image Fusion with a Deep Convolutional Neural Network. Inf. Fusion 2017, 36, 191–207. [Google Scholar] [CrossRef]
- Liu, Y.; Chen, X.; Cheng, J.; Peng, H.; Wang, Z. Infrared and Visible Image Fusion with Convolutional Neural Networks. Int. J. Wavelets Multiresolut Inf. Process. 2018, 16, 1850018. [Google Scholar] [CrossRef]
- Li, H.; Wu, X.-J. DenseFuse: A Fusion Approach to Infrared and Visible Images. IEEE Trans. on Image Process. 2019, 28, 2614–2623. [Google Scholar] [CrossRef] [PubMed]
- Ma, J.; Yu, W.; Liang, P.; Li, C.; Jiang, J. FusionGAN: A Generative Adversarial Network for Infrared and Visible Image Fusion. Inf. Fusion 2019, 48, 11–26. [Google Scholar] [CrossRef]
- Ma, J.; Xu, H.; Jiang, J.; Mei, X.; Zhang, X.-P. DDcGAN: A Dual-Discriminator Conditional Generative Adversarial Network for Multi-Resolution Image Fusion. IEEE Trans. on Image Process. 2020, 29, 4980–4995. [Google Scholar] [CrossRef] [PubMed]
- Zhang, H.; Xu, H.; Xiao, Y.; Guo, X.; Ma, J. Rethinking the Image Fusion: A Fast Unified Image Fusion Network Based on Proportional Maintenance of Gradient and Intensity. In Proceedings of the AAAI Conference on Artificial Intelligence; 34; pp. 12797–12804. [Google Scholar] [CrossRef]
- Xu, H.; Ma, J.; Jiang, J.; Guo, X.; Ling, H. U2Fusion: A Unified Unsupervised Image Fusion Network. IEEE Trans. Pattern Anal. Mach. Intell. 2022, 44, 502–518. [Google Scholar] [CrossRef] [PubMed]
- Tang, W.; He, F.; Liu, Y. YDTR: Infrared and Visible Image Fusion via Y-Shape Dynamic Transformer. IEEE Trans. Multimedia 2022, 1–16. [Google Scholar] [CrossRef]
- Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A.N.; Kaiser, Ł.; Polosukhin, I. Attention Is All You Need. In Proceedings of the Advances in Neural Information Processing Systems; Guyon, I., Luxburg, U.V., Bengio, S., Wallach, H., Fergus, R., Vishwanathan, S., Garnett, R., Eds.; Curran Associates, Inc., 2017; Vol. 30.
- Devlin, J.; Chang, M.-W.; Lee, K.; Toutanova, K. BERT: Pre-Training of Deep Bidirectional Transformers for Language Understanding 2019.
- Yang, Z.; Dai, Z.; Yang, Y.; Carbonell, J.; Salakhutdinov, R.R.; Le, Q.V. XLNet: Generalized Autoregressive Pretraining for Language Understanding. . In Proceedings of the Advances in Neural Information Processing Systems; Wallach, H., Larochelle, H., Beygelzimer, A., Alché-Buc, F. d’, Fox, E., Garnett, R., Eds.; Curran Associates, Inc., 2019; Vol. 32.
- Dosovitskiy, A.; Beyer, L.; Kolesnikov, A.; Weissenborn, D.; Zhai, X.; Unterthiner, T.; Dehghani, M.; Minderer, M.; Heigold, G.; Gelly, S.; et al. An Image Is Worth 16x16 Words: Transformers for Image Recognition at Scale 2021. 2021. [Google Scholar]
- Chen, J.; Lu, Y.; Yu, Q.; Luo, X.; Adeli, E.; Wang, Y.; Lu, L.; Yuille, A.L.; Zhou, Y. TransUNet: Transformers Make Strong Encoders for Medical Image Segmentation 2021. 2021. [Google Scholar]
- Wang, W.; Xie, E.; Li, X.; Fan, D.-P.; Song, K.; Liang, D.; Lu, T.; Luo, P.; Shao, L. Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without Convolutions 2021. 2021. [Google Scholar]
- Zheng, S.; Lu, J.; Zhao, H.; Zhu, X.; Luo, Z.; Wang, Y.; Fu, Y.; Feng, J.; Xiang, T.; Torr, P.H.S.; et al. Rethinking Semantic Segmentation from a Sequence-to-Sequence Perspective with Transformers. In Proceedings of the 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR); IEEE: Nashville, TN, USA, June, 2021; pp. 6877–6886. [Google Scholar]
- Yan, B.; Peng, H.; Fu, J.; Wang, D.; Lu, H. Learning Spatio-Temporal Transformer for Visual Tracking. In Proceedings of the 2021 IEEE/CVF International Conference on Computer Vision (ICCV); IEEE: Montreal, QC, Canada, October 2021; pp. 10428–10437. [Google Scholar]
- Ren, P.; Li, C.; Wang, G.; Xiao, Y.; Du, Q.; Liang, X.; Chang, X. Beyond Fixation: Dynamic Window Visual Transformer. In Proceedings of the 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR); IEEE: New Orleans, LA, USA, June 2022; pp. 11977–11987. [Google Scholar]
- Vs, V.; Jose Valanarasu, J.M.; Oza, P.; Patel, V.M. Image Fusion Transformer. In Proceedings of the 2022 IEEE International Conference on Image Processing (ICIP); IEEE: Bordeaux, France, 16 October 2022; pp. 3566–3570. [Google Scholar]
- Zhao, H.; Nie, R. DNDT: Infrared and Visible Image Fusion Via DenseNet and Dual-Transformer. In Proceedings of the 2021 International Conference on Information Technology and Biomedical Engineering (ICITBE); IEEE: Nanchang, China, December 2021; pp. 71–75. [Google Scholar]
- Fu, Y.; Xu, T.; Wu, X.; Kittler, J. PPT Fusion: Pyramid Patch Transformerfor a Case Study in Image Fusion 2022.
- Rao, D.; Xu, T.; Wu, X.-J. TGFuse: An Infrared and Visible Image Fusion Approach Based on Transformer and Generative Adversarial Network. IEEE Trans. on Image Process. 2023, 1–1. [Google Scholar] [CrossRef]
- Zhang, Y.; Tian, Y.; Kong, Y.; Zhong, B.; Fu, Y. Residual Dense Network for Image Super-Resolution 2018.
- Ma, J.; Zhou, Y. Infrared and Visible Image Fusion via Gradientlet Filter. Computer Vision and Image Understanding 2020, 197–198. [Google Scholar] [CrossRef]
- Bavirisetti, D.P.; Dhuli, R. Fusion of Infrared and Visible Sensor Images Based on Anisotropic Diffusion and Karhunen-Loeve Transform. IEEE Sens. J. 2016, 16, 203–209. [Google Scholar] [CrossRef]
- Li, G.; Lin, Y.; Qu, X. An Infrared and Visible Image Fusion Method Based on Multi-Scale Transformation and Norm Optimization. Inf. Fusion 2021, 71, 109–129. [Google Scholar] [CrossRef]
- Zhang, H.; Yuan, J.; Tian, X.; Ma, J. GAN-FM: Infrared and Visible Image Fusion Using GAN With Full-Scale Skip Connection and Dual Markovian Discriminators. IEEE Trans. Comput. Imaging 2021, 7, 1134–1147. [Google Scholar] [CrossRef]
- Zhang, Y.; Liu, Y.; Sun, P.; Yan, H.; Zhao, X.; Zhang, L. IFCNN: A General Image Fusion Framework Based on Convolutional Neural Network. Inf. Fusion 2020, 54, 99–118. [Google Scholar] [CrossRef]
- Wang, Z.; Bovik, A.C.; Sheikh, H.R.; Simoncelli, E.P. Image Quality Assessment: From Error Visibility to Structural Similarity. IEEE Trans. on Image Process. 2004, 13, 600–612. [Google Scholar] [CrossRef]
- Aslantas, V.; Kurban, R. A Comparison of Criterion Functions for Fusion of Multi-Focus Noisy Images. Opt. Commun. 2009, 282, 3231–3242. [Google Scholar] [CrossRef]
- Mukaka, M.M. A Guide to Appropriate Use of Correlation Coefficient in Medical Research. Malawi Med. J. 2012, 24, 69–71. [Google Scholar]
- Rajkumar, S.; Mouli, P.V.S.S.R.C. Infrared and Visible Image Fusion Using Entropy and Neuro-Fuzzy Concepts. In ICT and Critical Infrastructure: Proceedings of the 48th Annual Convention of Computer Society of India- Vol I; Satapathy, S.C., Avadhani, P.S., Udgata, S.K., Lakshminarayana, S., Satapathy, S.C., Avadhani, P.S., Udgata, S.K., Lakshminarayana, S., Eds.; Advances in Intelligent Systems and Computing; Springer International Publishing: Cham, 2014; pp. 24893–100. ISBN 978-3-319-03106-4. [Google Scholar]
- Aslantas, V.; Bendes, E. A New Image Quality Metric for Image Fusion: The Sum of the Correlations of Differences. AEU - International Journal of Electronics and Communications 2015, 69, 1890–1896. [Google Scholar] [CrossRef]
- Chen, Y.; Blum, R.S. A New Automated Quality Assessment Algorithm for Image Fusion. Image and Vis. Comput. 2009, 27, 1421–1432. [Google Scholar] [CrossRef]
- Ge, Z.; Liu, S.; Wang, F.; Li, Z.; Sun, J. YOLOX: Exceeding YOLO Series in 2021 2021.
















| Datasets | Methods | Quality Metrics | |||||
|---|---|---|---|---|---|---|---|
| SSIM | MSE | CC | PSNR | SCD | QCB | ||
| RoadScene | D-Trans | 0.7141 | 125.1179 | 0.7754 | 27.1705 | 1.2882 | 0.4773 |
| D-RDB | 0.7317 | 72.9455 | 0.7762 | 29.5688 | 1.2598 | 0.5086 | |
| O-Trans | 0.7306 | 68.2725 | 0.7837 | 29.8866 | 1.2907 | 0.5027 | |
| O-RDB | 0.7192 | 66.7998 | 0.7732 | 30.0231 | 1.3115 | 0.5112 | |
| E-FEM | 0.7330 | 82.1417 | 0.7815 | 29.0312 | 1.2557 | 0.5021 | |
| Our | 0.7277 | 46.6912 | 0.7990 | 31.5830 | 1.3218 | 0.5469 | |
| TNO | D-Trans | 0.7125 | 117.3409 | 0.5523 | 27.4721 | 1.5780 | 0.4768 |
| D-RDB | 0.7289 | 83.4061 | 0.5416 | 29.1615 | 1.4354 | 0.4805 | |
| O-Trans | 0.7539 | 83.2827 | 0.5279 | 29.1897 | 1.4338 | 0.4892 | |
| O-RDB | 0.7605 | 83.8169 | 0.5553 | 29.1480 | 1.5119 | 0.4945 | |
| E-FEM | 0.7624 | 90.1372 | 0.5351 | 28.7422 | 1.4452 | 0.4844 | |
| Our | 0.7587 | 80.6446 | 0.5452 | 29.3554 | 1.5180 | 0.5000 | |
| Methods | Quality Metrics | |||||
|---|---|---|---|---|---|---|
| SSIM | MSE | CC | PSNR | SCD | QCB | |
| ADF | 0.6909 | 93.2764 | 0.7810 | 28.4483 | 1.0776 | 0.5285 |
| IVFusion | 0.4642 | 60.1766 | 0.6855 | 30.3363 | 0.9305 | 0.4535 |
| GF | 0.7190 | 90.6036 | 0.7769 | 28.5830 | 1.2868 | 0.5465 |
| DenseFuse | 0.7453 | 93.8907 | 0.7851 | 28.4170 | 1.0819 | 0.5440 |
| DDcGAN | 0.5589 | 48.4804 | 0.7410 | 31.5353 | 1.1833 | 0.4594 |
| IFCNN | 0.7045 | 99.6146 | 0.7694 | 28.1568 | 1.1557 | 0.4973 |
| PMGI | 0.6777 | 26.9935 | 0.7120 | 34.2283 | 0.9673 | 0.5852 |
| GAN-FM | 0.6590 | 52.2970 | 0.7680 | 30.9994 | 1.3848 | 0.5327 |
| YDTR | 0.7231 | 131.9041 | 0.7771 | 26.9622 | 1.1619 | 0.5236 |
| Our | 0.7277 | 46.6912 | 0.7990 | 31.5830 | 1.3218 | 0.5469 |
| Methods | Quality Metrics | |||||
|---|---|---|---|---|---|---|
| SSIM | MSE | CC | PSNR | SCD | QCB | |
| ADF | 0.7085 | 101.6540 | 0.5407 | 28.1317 | 1.4522 | 0.4888 |
| IVFusion | 0.5224 | 133.1557 | 0.4016 | 27.4738 | 1.0104 | 0.4795 |
| GF | 0.7529 | 99.9707 | 0.5368 | 28.2123 | 1.5735 | 0.4891 |
| DenseFuse | 0.7547 | 104.4317 | 0.5068 | 28.0037 | 1.4427 | 0.4774 |
| DDcGAN | 0.5785 | 107.4098 | 0.5120 | 28.2019 | 1.4089 | 0.4497 |
| IFCNN | 0.7230 | 108.8827 | 0.5246 | 27.7912 | 1.5154 | 0.4804 |
| PMGI | 0.7067 | 87.2343 | 0.5363 | 29.7962 | 1.4668 | 0.4823 |
| GAN-FM | 0.6763 | 100.7944 | 0.5049 | 28.3592 | 1.5058 | 0.4540 |
| YDTR | 0.7443 | 131.2808 | 0.5213 | 26.9846 | 1.4979 | 0.4478 |
| Our | 0.7587 | 80.6446 | 0.5452 | 29.3554 | 1.5180 | 0.5000 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2023 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).