Submitted:
06 December 2024
Posted:
06 December 2024
You are already at the latest version
Abstract
Keywords:
1. Introduction
2. Related Work
2.1. Monocular Depth Estimation
2.2. Radar-Camera Depth Estimation
3. Materials and Methods
3.1. Overall Structure
3.2. Generate Semi-Dense Depth Estimation
3.2.1. Loss Function
3.3. Lower-Layer Bidirectional Feature Fusion
3.4. Higher-Layer Attention Mechanism and Feature Fusion
3.4.1. Loss Function
3.5. Parallel Attention Mechanism
3.5.1. Channel Attention
3.5.2. Position Attention
4. Experiments
4.1. Datasets and Experimental Environment
4.2. Training Details and Evaluation Metrics
4.3. Comparison and Analysis of Results
4.3.1. Results Analysis
4.3.2. Regional Result Analysis
5. Conclusions
Author Contributions
Funding
Data Availability Statement
Conflicts of Interest
References
- Jiang, Y.; Wu, Y.; Zhang, J.; Wei, J.; Peng, B.; Qiu, C.W. Dilemma in optical identification of single-layer multiferroics. Nature 2023, 619, E40–E43. [Google Scholar] [CrossRef] [PubMed]
- Jiang, Y.; He, A.; Zhao, R.; Chen, Y.; Liu, G.; Lu, H.; Zhang, J.; Zhang, Q.; Wang, Z.; Zhao, C.; others. Coexistence of photoelectric conversion and storage in van der Waals heterojunctions. Physical Review Letters 2021, 127, 217401. [Google Scholar] [CrossRef] [PubMed]
- Yao, S.; Guan, R.; Huang, X.; Li, Z.; Sha, X.; Yue, Y.; Lim, E.G.; Seo, H.; Man, K.L.; Zhu, X.; others. Radar-camera fusion for object detection and semantic segmentation in autonomous driving: A comprehensive review. IEEE Transactions on Intelligent Vehicles 2023. [Google Scholar] [CrossRef]
- Zhang, J.; Zhang, J.; Qi, Y.; Gong, S.; Xu, H.; Liu, Z.; Zhang, R.; Sadi, M.A.; Sychev, D.; Zhao, R.; others. Room-temperature ferroelectric, piezoelectric and resistive switching behaviors of single-element Te nanowires. Nature Communications 2024, 15, 7648. [Google Scholar] [CrossRef] [PubMed]
- Masoumian, A.; Rashwan, H.A.; Cristiano, J.; Asif, M.S.; Puig, D. Monocular depth estimation using deep learning: A review. Sensors 2022, 22, 5353. [Google Scholar] [CrossRef] [PubMed]
- Jiang, Y.; He, A.; Luo, K.; Zhang, J.; Liu, G.; Zhao, R.; Zhang, Q.; Wang, Z.; Zhao, C.; Wang, L.; others. Giant bipolar unidirectional photomagnetoresistance. Proceedings of the National Academy of Sciences 2022, 119, e2115939119. [Google Scholar] [CrossRef] [PubMed]
- Tran, D.M.; Ahlgren, N.; Depcik, C.; He, H. Adaptive active fusion of camera and single-point lidar for depth estimation. IEEE Transactions on Instrumentation and Measurement 2023, 72, 1–9. [Google Scholar] [CrossRef]
- Shao, S.; Pei, Z.; Chen, W.; Liu, Q.; Yue, H.; Li, Z. Sparse pseudo-lidar depth assisted monocular depth estimation. IEEE Transactions on Intelligent Vehicles 2023. [Google Scholar] [CrossRef]
- Jiang, Y.; Ma, X.; Wang, L.; Zhang, J.; Wang, Z.; Zhao, R.; Liu, G.; Li, Y.; Zhang, C.; Ma, C.; others. Observation of Electric Hysteresis, Polarization Oscillation, and Pyroelectricity in Nonferroelectric p-n Heterojunctions. Physical Review Letters 2023, 130, 196801. [Google Scholar] [CrossRef] [PubMed]
- Wei, Z.; Zhang, F.; Chang, S.; Liu, Y.; Wu, H.; Feng, Z. Mmwave radar and vision fusion for object detection in autonomous driving: A review. Sensors 2022, 22, 2542. [Google Scholar] [CrossRef] [PubMed]
- Lin, J.T.; Dai, D.; Van Gool, L. Depth estimation from monocular images and sparse radar data. 2020 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS). IEEE, 2020, pp. 10233–10240.
- Long, Y.; Morris, D.; Liu, X.; Castro, M.; Chakravarty, P.; Narayanan, P. Radar-camera pixel depth association for depth completion. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2021, pp. 12507–12516.
- Tang, J.; Tian, F.P.; Feng, W.; Li, J.; Tan, P. Learning guided convolutional network for depth completion. IEEE Transactions on Image Processing 2020, 30, 1116–1129. [Google Scholar] [CrossRef] [PubMed]
- Yan, Z.; Wang, K.; Li, X.; Zhang, Z.; Li, J.; Yang, J. RigNet: Repetitive image guided network for depth completion. European Conference on Computer Vision. Springer, 2022, pp. 214–230.
- Eigen, D.; Puhrsch, C.; Fergus, R. Depth map prediction from a single image using a multi-scale deep network. Advances in neural information processing systems 2014, 27. [Google Scholar]
- Eigen, D.; Fergus, R. Predicting depth, surface normals and semantic labels with a common multi-scale convolutional architecture. Proceedings of the IEEE international conference on computer vision, 2015, pp. 2650–2658.
- Laina, I.; Rupprecht, C.; Belagiannis, V.; Tombari, F.; Navab, N. Deeper depth prediction with fully convolutional residual networks. 2016 Fourth international conference on 3D vision (3DV). IEEE, 2016, pp. 239–248.
- He, K.; Zhang, X.; Ren, S.; Sun, J. Deep residual learning for image recognition. Proceedings of the IEEE conference on computer vision and pattern recognition, 2016, pp. 770–778.
- Cao, Y.; Wu, Z.; Shen, C. Estimating depth from monocular images as classification using deep fully convolutional residual networks. IEEE Transactions on Circuits and Systems for Video Technology 2017, 28, 3174–3182. [Google Scholar] [CrossRef]
- Fu, H.; Gong, M.; Wang, C.; Batmanghelich, K.; Tao, D. Deep ordinal regression network for monocular depth estimation. Proceedings of the IEEE conference on computer vision and pattern recognition, 2018, pp. 2002–2011.
- Piccinelli, L.; Sakaridis, C.; Yu, F. idisc: Internal discretization for monocular depth estimation. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023, pp. 21477–21487.
- Guizilini, V.; Ambrus, R.; Pillai, S.; Raventos, A.; Gaidon, A. 3d packing for self-supervised monocular depth estimation. Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2020, pp. 2485–2494.
- Johnston, A.; Carneiro, G. Self-supervised monocular trained depth estimation using self-attention and discrete disparity volume. Proceedings of the ieee/cvf conference on computer vision and pattern recognition, 2020, pp. 4756–4765.
- Godard, C.; Mac Aodha, O.; Firman, M.; Brostow, G.J. Digging into self-supervised monocular depth estimation. Proceedings of the IEEE/CVF international conference on computer vision, 2019, pp. 3828–3838.
- Lyu, X.; Liu, L.; Wang, M.; Kong, X.; Liu, L.; Liu, Y.; Chen, X.; Yuan, Y. Hr-depth: High resolution self-supervised monocular depth estimation. Proceedings of the AAAI conference on artificial intelligence, 2021, Vol. 35, pp. 2294–2301.
- Jung, H.; Park, E.; Yoo, S. Fine-grained semantics-aware representation enhancement for self-supervised monocular depth estimation. Proceedings of the IEEE/CVF International Conference on Computer Vision, 2021, pp. 12642–12652.
- Feng, C.; Wang, Y.; Lai, Y.; Liu, Q.; Cao, Y. Unsupervised monocular depth learning using self-teaching and contrast-enhanced SSIM loss. Journal of Electronic Imaging 2024, 33, 013019–013019. [Google Scholar] [CrossRef]
- Lo, C.C.; Vandewalle, P. Depth estimation from monocular images and sparse radar using deep ordinal regression network. 2021 IEEE International Conference on Image Processing (ICIP). IEEE, 2021, pp. 3343–3347.
- Gasperini, S.; Koch, P.; Dallabetta, V.; Navab, N.; Busam, B.; Tombari, F. R4Dyn: Exploring radar for self-supervised monocular depth estimation of dynamic scenes. 2021 International Conference on 3D Vision (3DV). IEEE, 2021, pp. 751–760.
- Zhang, X.; Zhu, J.; Wang, D.; Wang, Y.; Liang, T.; Wang, H.; Yin, Y. A gradual self distillation network with adaptive channel attention for facial expression recognition. Applied Soft Computing 2024, 161, 111762. [Google Scholar] [CrossRef]
- Bi, M.; Zhang, Q.; Zuo, M.; Xu, Z.; Jin, Q. Bi-directional long short-term memory model with semantic positional attention for the question answering system. Transactions on Asian and Low-Resource Language Information Processing 2021, 20, 1–13. [Google Scholar] [CrossRef]
- Caesar, H.; Bankiti, V.; Lang, A.H.; Vora, S.; Liong, V.E.; Xu, Q.; Krishnan, A.; Pan, Y.; Baldan, G.; Beijbom, O. nuscenes: A multimodal dataset for autonomous driving. Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2020, pp. 11621–11631.
- Singh, A.D.; Ba, Y.; Sarker, A.; Zhang, H.; Kadambi, A.; Soatto, S.; Srivastava, M.; Wong, A. Depth estimation from camera image and mmwave radar point cloud. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023, pp. 9275–9285.
- Ma, F.; Karaman, S. Sparse-to-dense: Depth prediction from sparse depth samples and a single image. 2018 IEEE international conference on robotics and automation (ICRA). IEEE, 2018, pp. 4796–4803.
- Wang, T.H.; Wang, F.E.; Lin, J.T.; Tsai, Y.H.; Chiu, W.C.; Sun, M. Plug-and-play: Improve depth estimation via sparse data propagation. arXiv preprint, arXiv:1812.08350 2018.












| Experimental Platform | Environment Configuration |
|---|---|
| Operating systems | Ubuntu18.04 |
| Programming Languages | Python 3.8 |
| CPU | Intel(R) Xeon(R) Platinum 8352V |
| GPU | NVIDIA RTX A5000 |
| CUDA | 11.3 |
| Eval Distance | Method | Radar frames | Images | MAE↓ | RMSE↓ |
|---|---|---|---|---|---|
| 5*50m | RC-PDA [12] | 5 | 3 | 2225.0 | 4156.5 |
| RC-PDA with HG | 5 | 3 | 2315.7 | 4321.6 | |
| DORN [28] | 5(x3) | 1 | 1926.6 | 4124.8 | |
| Singh [33] | 1 | 1 | 1727.7 | 3746.8 | |
| Ours | 1 | 1 | 1646.5 | 3589.3 | |
| 5*70m | RC-PDA [12] | 5 | 3 | 3326.1 | 6700.6 |
| RC-PDA with HG | 5 | 3 | 3485.6 | 7002.9 | |
| DORN [28] | 5(x3) | 1 | 2380.6 | 5252.7 | |
| Singh [33] | 1 | 1 | 2073.2 | 4825.0 | |
| Ours | 1 | 1 | 1942.6 | 4574.1 | |
| 10*80m | RC-PDA [12] | 5 | 3 | 3713.6 | 7692.8 |
| RC-PDA with HG | 5 | 3 | 3884.3 | 8008.6 | |
| DORN [28] | 5(x3) | 1 | 2467.7 | 5554.3 | |
| Lin [11] | 3 | 1 | 2371.0 | 5623.0 | |
| R4Dyn [29] | 4 | 1 | N/A | 6434.0 | |
| Sparse-to-dense [34] | 3 | 1 | 2374.0 | 5628.0 | |
| PnP [35] | 3 | 1 | 2496.0 | 5578.0 | |
| Singh [33] | 1 | 1 | 2179.3 | 4898.7 | |
| Ours | 1 | 1 | 2052.9 | 4658.7 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2024 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).