Submitted:
15 February 2025
Posted:
17 February 2025
You are already at the latest version
Abstract
Keywords:
1. Introduction
- An optimized ConvLSTM-based encoder-decoder structure that accelerates network integration and enhances the learning of the spatial-temporal features.
- A newly proposed supervised attention module that mitigates the high computational overhead of self-attention by guiding the model to focus on features most relevant to blur information when transferring features between different network scales.
- The introduction of a multi-loss function based on the fast Fourier transform (FFT), enabling the model to learn deblurring features in the frequency domain.
- The development of a new dataset collected in diverse environments, which outperforms existing datasets in several aspects, reducing the challenges posed by discrepancies between simulated data and real-world images.
2. Related Works
2.1. Encoder-Decoder Structure
2.2. Multi-Scale Network
2.3. Self-Attention
2.4. Residual Block
3. Deep Supervised Attention Network (DSANet)
3.1. Supervised Attenuation Module
3.2. Multi-Loss Function
3.3. A New Dataset
4. Experimental Results

5. Conclusions
Acknowledgments
Conflicts of Interest
References
- Kim, T.H.; Ahn, B.; Lee, K.M. Dynamic scene deblurring. In Proceedings of the IEEE International Conference on Computer Vision, Sydney, Australia, 2013; pp. 3160–3167. [CrossRef]
- Sun, J.; Cao, W.; Xu, Z.; Ponce, J. Learning a convolutional neural network for non-uniform motion blur removal. In Proceedings of the IEEE Computer Vision and Pattern Recognition, Boston, USA, 2015; pp. 769–777. [CrossRef]
- Schuler, C.J.; Hirsch, M.; Harmeling, S.; Scholkopf, B. Learning to deblur. IEEE Transactions on Pattern Analysis and Machine Intelligence 2016, 38, 1439–1451. [CrossRef]
- Nah, S.; Kim, T.H.; Lee, K.M. Deep multi-scale convolutional neural network for dynamic scene deblurring. In Proceedings of the IEEE Computer Vision and Pattern Recognition, Honolulu, USA, 2017; pp. 257–265. [CrossRef]
- Ren, W.; Zhang, J.; Pan, J.; Liu, S.; Ren, J.S.; Du, J.; Cao, X.; Yan, M.H. Deblurring dynamic scenes via spatially varying recurrent neural networks. IEEE Transactions on Pattern Analysis and Machine Intelligence 2022, 44, 3974–3987. [CrossRef]
- Tao, X.; Gao, H.; Shen, X.; Wang, J.; Jia, J. Scale-recurrent network for deep image deblurring. In Proceedings of the IEEE Computer Vision and Pattern Recognition, Salt Lake City, USA, 2018; pp. 8174–8182. [CrossRef]
- Kupyn, O.; Martyniuk, T.; Wu, J.; Wang, Z. DeblurGAN-v2: Deblurring (orders-of-magnitude) faster and better. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Seoul, Korea, 2019; pp. 8877–8886. [CrossRef]
- Kuldeep, P.; Rajagopalan, A.N. Region-adaptive dense network for efficient motion deblurring. In Proceedings of the AAAI Conference on Artificial Intelligence, 2020, 34, 11882–11889. [CrossRef]
- Badrinarayanan, V.; Kendall, A.; Cipolla, R. SegNet: A deep convolutional encoder-decoder architecture for image segmentation. IEEE Transactions on Pattern Analysis and Machine Intelligence 2017, 39, 2481–2495. [CrossRef]
- Ronneberger, O.; Fischer, P.; Brox, T. U-Net: Convolutional networks for biomedical image segmentation. In Proceedings of the Medical Image Computing and Computer-Assisted Intervention, Munich, Germany, 2015; pp. 234–241.
- Noh, H.; Hong, S.; Han, B. Learning deconvolution network for semantic segmentation. In Proceedings of the IEEE International Conference on Computer Vision, Santiago, Chile, 2015; pp. 1520–1528. [CrossRef]
- Long, J.; Shelhamer, E.; Darrell, T. Fully convolutional networks for semantic segmentation. In Proceedings of the IEEE Computer Vision and Pattern Recognition, Boston, USA, 2015; pp. 3431–3440. [CrossRef]
- Eigen, D.; Fergus, R. Predicting depth, surface normals and semantic labels with a common multi-scale convolutional architecture. In Proceedings of the IEEE International Conference on Computer Vision, Santiago, Chile, 2015; pp. 2650–2658. [CrossRef]
- Mathieu, M.; Couprie, C.; LeCun, Y. Deep multiscale video prediction beyond mean square error. arXiv 2015, arXiv:1511.05440.
- Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A.N.; Kaiser, L.; Polosukhin, I. Attention is all you need. In Proceedings of the NIPS, Long Beach, USA, 2017; pp. 5998–6008. arXiv:1706.03762.
- Wang, F.; Jiang, M.; Qian, C.; Yang, S.; Li, C.; Zhang, H.; Wang, X.; Tang, X. Residual attention network for image classification. In Proceedings of the IEEE Computer Vision and Pattern Recognition, Honolulu, USA, 2017; pp. 6450–6458. [CrossRef]
- Woo, S.; Park, J.; Lee, J.; Kweon, I.S. CBAM: Convolutional block attention module. In Proceedings of the European Conference on Computer Vision, Munich, Germany, 2018; pp. 3–19.
- Hu, J.; Shen, L.; Sun, G. Squeeze-and-excitation networks. In Proceedings of the IEEE Computer Vision and Pattern Recognition, Salt Lake City, USA, 2018; pp. 7132–7141. [CrossRef]
- Tsai, F.J.; Peng, Y.T.; Tsai, C.C.; Lin, Y.Y. BANet: A blur-aware attention network for dynamic scene deblurring. IEEE Transactions on Image Processing 2022, 31, 6789–6799. [CrossRef]
- He, K.; Zhang, X.; Ren, S.; Sun, J. Deep residual learning for image recognition. In Proceedings of the IEEE Computer Vision and Pattern Recognition, Las Vegas, USA, 2016; pp. 770–778. [CrossRef]
- Nair, V.; Hinton, G.E. Rectified linear units improve restricted Boltzmann machines. In Proceedings of the International Conference on Machine Learning, Haifa, Israel, 2010; pp. 807–814.
- Wang, X.; Yu, K.; Wu, S.; Gu, J.; Liu, Y.; Dong, C.; Qiao, Y.; Change, C.R. ESRGAN: Enhanced super-resolution generative adversarial networks. In Proceedings of the European Conference on Computer Vision, 2018; pp. 1–16.
- Yi, P.; Wang, Z.; Jiang, K.; Jiang, J.; Ma, J. Progressive fusion video super-resolution network via exploiting non-local spatio-temporal correlations. In Proceedings of the IEEE International Conference on Computer Vision, Seoul, Korea, 2019; pp. 3106–3115. [CrossRef]
- Hochreiter, S.; Schmidhuber, J. Long short-term memory. Neural Computation 1997, 9, 1735–1780. [CrossRef]
- Ledig, C.; Theis, L.; Twitter, W.S. Photo-realistic single image super-resolution using a generative adversarial network. In Proceedings of the IEEE Computer Vision and Pattern Recognition, Honolulu, USA, 2017; pp. 105–114. [CrossRef]
- Johnson, J.; Alahi, A.; Li, F. Perceptual losses for real-time style transfer and super-resolution. In Proceedings of the European Conference on Computer Vision, Amsterdam, The Netherlands, 2016; pp. 694–711.
- Jianbo, J.; Cao, Y.; Song, Y.; Lau, R. Look deeper into depth: Monocular depth estimation with semantic booster and attention driven loss. In Proceedings of the European Conference on Computer Vision, Munich, Germany, 2018; pp. 53–69.
- Ignatov, A.; Kobyshev, N.; Timofte, R.; Vanhoey, K. DSLR-quality photos on mobile devices with deep convolutional networks. In Proceedings of the IEEE International Conference on Computer Vision, Venice, Italy, 2017; pp. 3297–3305. [CrossRef]
- Su, S.; Delbracio, M.; Wang, J.; Sapiro, G.; Heidrich, W.; Wang, O. Deep video deblurring for hand-held cameras. In Proceedings of the IEEE Computer Vision and Pattern Recognition, Honolulu, USA, 2017; pp. 237–246. [CrossRef]
- Zhou, S.; Zhang, J.; Zuo, W.; Xie, H.; Pan, J.; Ren, J.S. DAVANet: Stereo deblurring with view aggregation. In Proceedings of the IEEE Computer Vision and Pattern Recognition, Long Beach, USA, 2019; pp. 10988–10997. [CrossRef]
- Lu, B.; Chen, J.C.; Chellappa, R. Unsupervised domain-specific deblurring via disentangled representations. In Proceedings of the IEEE Computer Vision and Pattern Recognition, Long Beach, USA, 2019; pp. 10217–10226. [CrossRef]
- Nimisha, T.M.; Sunil, K.; Rajagopalan, A.N. Unsupervised class-specific deblurring. In Proceedings of the European Conference on Computer Vision, Munich, Germany, 2018; pp. 353–369.
- Lai, W.S.; Huang, J.B.; Hu, Z.; Ahuja, N.; Yang, M.H. A comparative study for single image blind deblurring. In Proceedings of the IEEE Computer Vision and Pattern Recognition, Las Vegas, USA, 2016; pp. 1701–1709. [CrossRef]
- Rim, J.; Lee, H.; Won, J.; Cho, S. Real-world blur dataset for learning and benchmarking deblurring algorithms. In Proceedings of the European Conference on Computer Vision, 2020; pp. 184–201.
- Zhang, K. Deblurring by realistic blurring. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2020; pp. 2734–2743. [CrossRef]
- Goodfellow, I.; Pouget-Abadie, J.; Mirza, M.; Xu, B.; Warde-Farley, D.; Ozair, S.; Courville, A.; Bengio, Y. Generative adversarial nets. In Proceedings of the NIPS, 2014.
- He, K.; Zhang, X.; Ren, S.; Sun, J. Delving deep into rectifiers: Surpassing human-level performance on ImageNet classification. In Proceedings of the IEEE International Conference on Computer Vision, Santiago, Chile, 2015; pp. 1026–1034. [CrossRef]
- Kingma, D.P.; Ba, J. Adam: A method for stochastic optimization. arXiv 2014, arXiv:1412.6980.
- Zhang, H.; Dai, Y.; Li, H.; Koniusz, P. Deep stacked hierarchical multi-patch network for image deblurring. In Proceedings of the IEEE Computer Vision and Pattern Recognition, Long Beach, USA, 2019; pp. 5978–5986.
- Shi, X.; Chen, Z.; Wang, H.; Yeung, D.Y.; Wong, W.K.; Woo, W.C. Convolutional LSTM network: A machine learning approach for precipitation nowcasting. In Proceedings of the International Conference on Neural Information Processing Systems, 2015; pp. 802–810.








| Model | PSNR | SSIM | Time (ms) | Size (MB) |
| MSCNN | 29.23 | 0.9162 | 4300 | 303.6 |
| RNNDeblur | 29.19 | 0.9306 | 1400 | 37.1 |
| SRN | 30.60 | 0.9323 | 1600 | 33.6 |
| DMPHN | 31.25 | 0.9483 | 424 | 86.8 |
| DSANet1 | 31.38 | 0.9485 | 254 | 32.2 |
| DSANet2 | 31.55 | 0.9490 | 254 | 32.2 |
| DSANet++ | 31.65 | 0.9492 | 254 | 32.2 |
| k = 1 | k = 2 | k = 3 | |
| PSNR | 30.5 | 31.4 | 31.65 |
| SSIM | 0.9402 | 0.9485 | 0.9492 |
| Time (ms) | 921 | 534 | 253 |
| Model | PSNR | SSIM | Time (ms) | Size (MB) |
| 1Ed Model | 27.56 | 0.9255 | 80 | 5.3 |
| 2Ed Model | 29.4 | 0.9358 | 145 | 12.4 |
| 3Ed Model | 30.24 | 0.9401 | 280 | 31.5 |
| DSANet | 31.65 | 0.9492 | 254 | 32.2 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2025 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).