Submitted:
20 June 2025
Posted:
20 June 2025
You are already at the latest version
Abstract

Keywords:
1. Introduction
- Integration of a CBAM module within the YOLOv8 backbone to enhance spatial and channel-wise focus on structural features of doorways.
- Implementation of a CGCAFusion module for multi-scale contextual refinement to improve accuracy in segmenting semantically similar regions.
- Development of a lightweight dual head structure for simultaneous door segmentation and monocular depth estimation module (DEM).
- Implementation of an Alignment Estimation Module (AEM) to correct door misalignment and provide real-time guidance to users.
2. Materials and Methods
- 1)
- Convolutional Block Attention Module (CBAM): This sub module enhances the feature representation in the YOLOv8 backbone by applying channel and spatial attention, allowing the model to focus more effectively on salient regions associated with doorways, as detailed in Section 2.2.
- 2)
- Content-Guided Convolutional Attention Fusion Module (CGCAFusion): This sub module is responsible for dynamically building contextual relationships based on input feature content using a content-guided convolutional attention mechanism, improving structural segmentation performance in cluttered indoor scenes, as detailed in Section 2.3.
- 3)
- Depth Estimation Module (DEM): The objective of this sub module is to predict relative depth information from RGB inputs, eliminating the need for external depth sensors while providing spatial context to support alignment estimation (Section 2.4).
- 4)
- Alignment Estimation Module (AEM): The final sub module is designed to compute the lateral offset of the detected doorway with respect to the image center to suggest directional guidance (e.g., move left, right, or stay centered) for wheelchair navigation, as described in Section 2.5.
2.1. Dataset Description
2.2. Convolutional Block Attention Module (CBAM)
- 1)
- Channel Attention Module (CAM): This module infers channel-wise importance using both global average pooling (GAP) and global max pooling (GMP), followed by a shared multi-layer perceptron (MLP) [39] (Equations (1) and (2)).where , denotes the sigmoid function, and ⊗ denotes element-wise multiplication.
- 2)
2.3. Content Guided Convolutional Attention Fusion Module (CGCAFusion)
- CGA (Content-Guided Attention) [44]: A coarse-to-fine spatial attention mechanism that refines each feature channel by learning spatial saliency within channels.
- CAFM (Convolutional Attention Fusion Module) [45]: A simplified transformer-inspired structure that extracts global features using self-attention while preserving local features via depthwise convolutions.
2.4. Depth Estimation Module (DEM)
2.5. Alignment Estimation Module (AEM)
3. Results and Discussion
3.1. Structural Attention Enhancement through CBAM
3.2. Improved Precision and Semantic Fusion using CGCAFusion
3.3. Performance of Unsupervised DEM
3.4. Accurate Alignment Estimation for Intelligent Guidance
4. Conclusions
Author Contributions
Funding
Data Availability Statement
Conflicts of Interest
Abbreviations
| YOLO | You Only Look Once |
| CBAM | Convolutional Block Attention Module |
| CGCAFusion | Content Guided Convolutional Attention Fusion Module |
| DEM | Depth Estimation Module |
| AEM | Alignment Estimation Module |
| CAM | Channel Attention Module |
| GAP | Global Average Pooling |
| GMP | Global Max Pooling |
| MLP | Multi-layer Perceptron |
| SAM | Spatial Attention Module |
| CGA | Content Guided Attention |
| CAFM | Convolutional Attention Fusion Module |
| mAP | Mean Average Precision |
| MAE | Mean Absolute Error |
References
- Dickinson, L. Autonomy and motivation a literature review. System 1995, 23, 165–174. [Google Scholar] [CrossRef]
- Atkinson, J. Autonomy and mental health. In Ethical issues in mental health; Springer, 1991; pp. 103–126.
- Mayo, N.E.; Mate, K.K.V. Quantifying Mobility in Quality of Life. In Quantifying Quality of Life: Incorporating Daily Life into Medicine; Wac, K., Wulfovich, S., Eds.; Springer International Publishing: Cham, 2022; pp. 119–136. [Google Scholar] [CrossRef]
- Meijering, L. Towards meaningful mobility: a research agenda for movement within and between places in later life. Ageing and Society 2021, 41, 711–723. [Google Scholar] [CrossRef]
- Stevens, G.A.; White, R.A.; Flaxman, S.R.; Price, H.; Jonas, J.B.; Keeffe, J.; Leasher, J.; Naidoo, K.; Pesudovs, K.; Resnikoff, S.; et al. Global prevalence of vision impairment and blindness: magnitude and temporal trends, 1990–2010. Ophthalmology 2013, 120, 2377–2384. [Google Scholar] [CrossRef] [PubMed]
- Fricke, T.R.; Jong, M.; Naidoo, K.S.; Sankaridurg, P.; Naduvilath, T.J.; Ho, S.M.; Wong, T.Y.; Resnikoff, S. Global prevalence of visual impairment associated with myopic macular degeneration and temporal trends from 2000 through 2050: systematic review, meta-analysis and modelling. British Journal of Ophthalmology 2018, 102, 855–862. [Google Scholar] [CrossRef]
- Sliwa, K.; of Disease Study 2013 Collaborators, G.B.; et al. Global, regional, and national incidence, prevalence, and years lived with disability for 301 acute and chronic diseases and injuries in 188 countries, 1990–2013: a systematic analysis for the Global Burden of Disease Study 2013. The Lancet 2015, 743–800. [Google Scholar]
- Park, S.J.; Ahn, S.; Park, K.H. Burden of visual impairment and chronic diseases. JAMA ophthalmology 2016, 134, 778–784. [Google Scholar] [CrossRef] [PubMed]
- Maresova, P.; Javanmardi, E.; Barakovic, S.; Barakovic Husic, J.; Tomsone, S.; Krejcar, O.; Kuca, K. Consequences of chronic diseases and other limitations associated with old age–a scoping review. BMC public health 2019, 19, 1–17. [Google Scholar] [CrossRef]
- Mahadevaswamy, U.; Rohith, M.; Arivazhagan, R. Development of a Semi-Automatic Wheelchair System for Improved Mobility and User Independence. In Proceedings of the 2024 3rd International Conference on Automation, Computing and Renewable Systems (ICACRS). IEEE; 2024; pp. 36–43. [Google Scholar]
- Karim, S.; Que, B.; Que, J.; Reyes, L.; Lim, L.G.; Bandala, A.A.; Vicerra, R.R.P.; Dadios, E.P. Design, fabrication, and testing of a semi-autonomous wheelchair. In Proceedings of the 2017IEEE 9th International Conference on Humanoid, Nanotechnology, Information Technology, Communication and Control, Environment and Management (HNICEM). IEEE. 2017; pp. 1–7. [Google Scholar]
- Hasan, S.; Faisal, F.; Sabrin, S.; Tong, Z.; Hasan, M.; Debnath, D.; Hossain, M.S.; Siddique, A.H.; Alam, J. A Simplified Approach to Develop Low Cost Semi-Automated Prototype of a Wheelchair. InUniversity of Science and Technology Annual (USTA) 2020. [Google Scholar]
- Kim, E.Y. Wheelchair navigation system for disabled and elderly people. Sensors 2016, 16, 1806. [Google Scholar] [CrossRef]
- Sanders, D.; Tewkesbury, G.; Stott, I.J.; Robinson, D. Simple expert systems to improve an ultrasonic sensor-system for a tele-operated mobile-robot. Sensor Review 2011, 31, 246–260. [Google Scholar] [CrossRef]
- Zheng, T.; Duan, Z.; Wang, J.; Lu, G.; Li, S.; Yu, Z. Research on distance transform and neural network lidar information sampling classification-based semantic segmentation of 2d indoor room maps. Sensors 2021, 21, 1365. [Google Scholar] [CrossRef] [PubMed]
- Gallo, V.; Shallari, I.; Carratù, M.; Laino, V.; Liguori, C. Design and Characterization of a Powered Wheelchair Autonomous Guidance System. Sensors 2024, 24, 1581. [Google Scholar] [CrossRef] [PubMed]
- Perra, C.; Kumar, A.; Losito, M.; Pirino, P.; Moradpour, M.; Gatto, G. Monitoring Indoor People Presence in Buildings Using Low-Cost Infrared Sensor Array in Doorways. Sensors 2021, 21. [Google Scholar] [CrossRef]
- Sahoo, S.; Choudhury, B. Voice-activated wheelchair: An affordable solution for individuals with physical disabilities. Management Science Letters 2023, 13, 175–192. [Google Scholar] [CrossRef]
- Sahoo, S.K.; Choudhury, B.B. Autonomous navigation and obstacle avoidance in smart robotic wheelchairs. Journal of Decision Analytics and Intelligent Computing 2024, 4, 47–66. [Google Scholar] [CrossRef]
- Ess, A.; Schindler, K.; Leibe, B.; Van Gool, L. Object detection and tracking for autonomous navigation in dynamic environments. The International Journal of Robotics Research 2010, 29, 1707–1725. [Google Scholar] [CrossRef]
- Qiu, Z.; Lu, Y.; Qiu, Z. Review of ultrasonic ranging methods and their current challenges. Micromachines 2022, 13, 520. [Google Scholar] [CrossRef]
- Lecrosnier, L.; Khemmar, R.; Ragot, N.; Decoux, B.; Rossi, R.; Kefi, N.; Ertaud, J.Y. Deep learning-based object detection, localisation and tracking for smart wheelchair healthcare mobility. International journal of environmental research and public health 2021, 18, 91. [Google Scholar] [CrossRef]
- Ju, M.; Luo, H.; Wang, Z.; Hui, B.; Chang, Z. The application of improved YOLO V3 in multi-scale target detection. Applied Sciences 2019, 9, 3775. [Google Scholar] [CrossRef]
- Bewley, A.; Ge, Z.; Ott, L.; Ramos, F.; Upcroft, B. Simple online and realtime tracking. In Proceedings of the 2016 IEEE international conference on image processing (ICIP). Ieee; 2016; pp. 3464–3468. [Google Scholar]
- Zhang, T.; Li, J.; Jiang, Y.; Zeng, M.; Pang, M. Position detection of doors and windows based on dspp-yolo. Applied Sciences 2022, 12, 10770. [Google Scholar] [CrossRef]
- Iandola, F.; Moskewicz, M.; Karayev, S.; Girshick, R.; Darrell, T.; Keutzer, K. Densenet: Implementing efficient convnet descriptor pyramids. arXiv preprint arXiv:1404.1869, arXiv:1404.1869 2014.
- He, K.; Zhang, X.; Ren, S.; Sun, J. Spatial pyramid pooling in deep convolutional networks for visual recognition. IEEE transactions on pattern analysis and machine intelligence 2015, 37, 1904–1916. [Google Scholar] [CrossRef] [PubMed]
- Mochurad, L.; Hladun, Y. Neural network-based algorithm for door handle recognition using RGBD cameras. Scientific Reports 2024, 14, 15759. [Google Scholar] [CrossRef]
- Sandler, M.; Howard, A.; Zhu, M.; Zhmoginov, A.; Chen, L.C. Mobilenetv2: Inverted residuals and linear bottlenecks. In Proceedings of the Proceedings of the IEEE conference on computer vision and pattern recognition, 2018, pp. 3–19.
- Hussain, M. YOLOv5, YOLOv8 and YOLOv10: The Go-To Detectors for Real-time Vision, 2024, [arXiv:cs.CV/2407.02988].
- Wei, L.; Tong, Y. Enhanced-YOLOv8: A new small target detection model. Digital Signal Processing 2024, 153, 104611. [Google Scholar] [CrossRef]
- Sharma, P.; Tyagi, R.; Dubey, P. Bridging the Perception Gap A YOLO V8 Powered Object Detection System for Enhanced Mobility of Visually Impaired Individuals. In Proceedings of the 2024 First International Conference on Technological Innovations and Advance Computing (TIACOMP); 2024; pp. 107–117. [Google Scholar] [CrossRef]
- Choi, E.; Dinh, T.A.; Choi, M. Enhancing Driving Safety of Personal Mobility Vehicles Using On-Board Technologies. Applied Sciences 2025, 15. [Google Scholar] [CrossRef]
- Zou, Z.; Chen, K.; Shi, Z.; Guo, Y.; Ye, J. Object detection in 20 years: A survey. Proceedings of the IEEE 2023, 111, 257–276. [Google Scholar] [CrossRef]
- Tennekoon, S.; Wedasingha, N.; Welhenge, A.; Abhayasinghe, N.; Murray Am, I. Advancing Object Detection: A Narrative Review of Evolving Techniques and Their Navigation Applications. IEEE Access 2025, 13, 50534–50555. [Google Scholar] [CrossRef]
- Woo, S.; Park, J.; Lee, J.Y.; Kweon, I.S. Cbam: Convolutional block attention module. In Proceedings of the Proceedings of the European conference on computer vision (ECCV), 2018, pp.
- Zhang, Z.; Zou, Y.; Tan, Y.; Zhou, C. YOLOv8-seg-CP: a lightweight instance segmentation algorithm for chip pad based on improved YOLOv8-seg model. Scientific Reports 2024, 14, 27716. [Google Scholar] [CrossRef]
- Ramôa, J.; Lopes, V.; Alexandre, L.; Mogo, S. Real-time 2D–3D door detection and state classification on a low-power device. SN Applied Sciences 2021, 3. [Google Scholar] [CrossRef]
- Kruse, R.; Mostaghim, S.; Borgelt, C.; Braune, C.; Steinbrecher, M. Multi-layer perceptrons. In Computational intelligence: a methodological introduction; Springer, 2022; pp. 53–124.
- Hu, J.; Shen, L.; Sun, G. Squeeze-and-excitation networks. In Proceedings of the Proceedings of the IEEE conference on computer vision and pattern recognition, 2018, pp. 7132–7141.
- Hou, Q.; Zhou, D.; Feng, J. Coordinate attention for efficient mobile network design. In Proceedings of the Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2021, pp.
- Zhu, L.; Wang, X.; Ke, Z.; Zhang, W.; Lau, R.W. Biformer: Vision transformer with bi-level routing attention. In Proceedings of the Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2023, pp. 10323–10333.
- Xu, W.; Wan, Y. ELA: Efficient local attention for deep convolutional neural networks. arXiv preprint arXiv:2403.01123, arXiv:2403.01123 2024.
- Chen, Z.; He, Z.; Lu, Z.M. DEA-Net: Single image dehazing based on detail-enhanced convolution and content-guided attention. IEEE Transactions on Image Processing 2024, 33, 1002–1015. [Google Scholar] [CrossRef]
- Hu, S.; Gao, F.; Zhou, X.; Dong, J.; Du, Q. Hybrid convolutional and attention network for hyperspectral image denoising. IEEE Geoscience and Remote Sensing Letters 2024. [Google Scholar] [CrossRef]
- Jocher, G. YOLOv5 by Ultralytics, 2020. [CrossRef]
- Wang, C.Y.; Bochkovskiy, A.; Liao, H.Y.M. YOLOv7: Trainable bag-of-freebies sets new state-of-the-art for real-time object detectors. In Proceedings of the Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2023, pp. 7464–7475.
- Jocher, G.; Qiu, J.; Chaurasia, A. Ultralytics YOLO, 2023.
- He, K.; Gkioxari, G.; Dollár, P.; Girshick, R. Mask r-cnn. In Proceedings of the Proceedings of the IEEE international conference on computer vision, 2017; pp. 2961–2969.













| Parameters | Value |
|---|---|
| Epochs | 150 |
| Ir0 | 0.002 |
| Irf | 0.002 |
| Momentum | 0.9 |
| Batchsize | 16 |
| Cache | False |
| Input image size | 640 x 640 |
| Optimizer | AdamW |
| Model | mAP50 (Box) | mAP50 (Mask) | Params (M) | FPS | Model size (M) | Inference time (ms) |
|---|---|---|---|---|---|---|
| Mask R-CNN | 0.825 | 0.814 | 45.96 | 105.7 | 346.52 | 9.33 |
| YOLOv5n-seg | 0.783 | 0.609 | 2.53 | 239 | 5.14 | 4.62 |
| YOLOv5s-seg | 0.808 | 0.621 | 7.74 | 252.1 | 15.6 | 4.26 |
| YOLOv7-seg | 0.953 | 0.911 | 37.98 | 147.23 | 78.1 | 6.9 |
| YOLOv8n-seg | 0.896 | 0.814 | 3.26 | 1490 | 6.45 | 0.64 |
| YOLOv8s-seg | 0.872 | 0.808 | 11.79 | 720 | 22.73 | 1.39 |
| YOLOv8x-seg | 0.873 | 0.696 | 71.75 | 120 | 137.26 | 8.62 |
| Proposed Model | 0.958 | 0.924 | 2.96 | 1560 | 3.6 | 0.42 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2025 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).