Submitted:
06 June 2026
Posted:
08 June 2026
You are already at the latest version
Abstract
Keywords:
1. Introduction
2. Related Work
2.1. Assistive Vision Systems for the Visually Impaired
2.2. Monocular Depth Estimation
2.3. Retrieval-Augmented Generation in Computer Vision
2.4. Agentic AI Systems
3. Materials and Methods
3.1. System Architecture Overview
3.2. Object Detection Layer (Unchanged from v1)
3.3. Depth Awareness Module
3.4. Temporal Memory Layer (ChromaDB RAG)
3.5. Temporal Threat Score (TTS): The Urgency Equation
| TTS Range | Urgency Level | Voice Response Strategy |
| TTS < 0.40 | LOW | Describe obstacle, no alert |
| 0.40 ≤ TTS < 0.62 | MEDIUM | Describe + caution warning |
| TTS ≥ 0.62 | HIGH | Strong warning + recommend action |




3.6. Agentic Decision Layer
3.7. Latency Optimizations
4. Experimental Evaluation
4.1. Experimental Setup
4.2. Ablation Study
4.3. Memory Recall Evaluation
| Original Scene Description | Query Paraphrase | Cosine Distance | Retrieved |
| a wooden chair blocking the hallway | chair placed in the middle of the passageway | 0.3458 | Y |
| a glass door at the end of the corridor | transparent door at the hallway entrance | 0.3160 | Y |
| a staircase going up on the left side | steps leading upward to the left | 0.2355 | Y |
| a person standing near the window | someone standing by the window | 0.0410 | Y |
| a desk with a laptop computer on it | table with a computer on top | 0.3250 | Y |
| a bicycle leaning against the wall | bike propped up against the wall | 0.2422 | Y |
| boxes stacked in the middle of the room | cardboard boxes piled up in the center of the room | 0.1454 | Y |
| an open door blocking the walkway | door swung open into the path | 0.3235 | Y |
| a shopping cart parked in the aisle | shopping cart left in the aisle | 0.1452 | Y |
| a low coffee table in front of the sofa | small table positioned in front of the couch | 0.2431 | Y |
| water puddle on the floor near the entrance | wet patch on the floor by the door | 0.3579 | Y |
| a backpack left on the ground | backpack left sitting on the ground | 0.1032 | Y |
| a trash bin near the doorway | trash bin placed near the doorway | 0.0567 | Y |
| a dog resting on the carpet | pet dog lying on the rug | 0.2388 | Y |
| a large potted plant blocking the hallway | large plant pot placed in the corridor | 0.2163 | Y |
4.4. Depth Estimation Accuracy
| True Distance (m) | Raw Predicted (m) | Calibrated (m) | Raw |Error| (m) | Cal. |Error| (m) | Rel. Error (raw) | Detected By | Fallback |
| 0.5 | 0.6 | 0.82 | 0.100 | 0.320 | 20.0% | — | Yes |
| 1.0 | 1.2 | 1.34 | 0.200 | 0.340 | 20.0% | Door handle | No |
| 1.5 | 1.4 | 1.51 | 0.100 | 0.010 | 6.7% | — | Yes |
| 2.0 | 1.9 | 1.95 | 0.100 | 0.050 | 5.0% | — | Yes |
| 3.0 | 2.2 | 2.21 | 0.800 | 0.790 | 26.7% | Door handle | No |
| 4.0 | 4.1 | 3.85 | 0.100 | 0.150 | 2.5% | Door handle | No |
| 5.0 | 5.8 | 5.32 | 0.800 | 0.320 | 16.0% | Door handle | No |
| MAE (raw) | 0.314 m | ||||||
| MAE (calibrated) | 0.285 m | ||||||
| RMSE (raw) | 0.441 m | ||||||
| RMSE (calibrated) | 0.374 m |
4.5. Component Latency Analysis
| Component | CPU Mean (N=371) | GPU Mean (N=20) | GPU Speedup |
| Parallel inference (YOLO + BLIP + Depth) | 5,523 ms | 236 ms | 23× |
| ChromaDB memory query | 270 ms | 103 ms | 2.6× |
| Anthropic agent call (Claude Haiku) | 1,417 ms | 1,010 ms | 1.4× |
| Sum of component means | 7,210 ms | 1,349 ms | 5.3× |
5. Discussion
5.1. Innovations and Key Contributions
5.2. Limitations
5.3. Ethical Considerations and Privacy
5.4. Comparative Positioning
6. Conclusions
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Acknowledgments
Conflicts of Interest
References
- Wahwah, S.; Chen, Y. My Eye AI: A Hybrid Cloud-Mobile Object Detection System for the Visually Impaired Using YOLOv11, OWL-ViT, and BLIP. Also Present. At. 2025 IEEE 16th Int. Symp. Auton. Decentralized Syst. (ISADS) J. Artif. Intell. Technol. 2025. [Google Scholar] [CrossRef]
- Lewis, P.; Perez, E.; Piktus, A.; Petroni, F.; Karpukhin, V.; Goyal, N.; Küttler, H.; Lewis, M.; Yih, W.; Rocktäschel, T.; Riedel, S.; Kiela, D. Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks. in Proceedings of NeurIPS 2020, 2020. [Google Scholar]
- Yang, L.; Kang, B.; Huang, Z.; Xu, X.; Feng, J.; Zhao, H. Depth Anything: Unleashing the Power of Large-Scale Unlabeled Data. in Proceedings of CVPR 2024, 2024; pp. 10371–10381. [Google Scholar]
- Yang, L.; Kang, B.; Huang, Z.; Zhao, Z.; Xu, X.; Feng, J.; Zhao, H. “Depth Anything V2,” in Proceedings of NeurIPS. arXiv 2024, arXiv:2406.09414. [Google Scholar]
- Li, J.; Li, D.; Xiong, C.; Hoi, S. BLIP: Bootstrapping Language-Image Pre-Training for Unified Vision-Language Understanding and Generation. in Proceedings of ICML 2022, 2022. [Google Scholar]
- Khanam, R.; Hussain, M. YOLOv11: An Overview of the Key Architectural Enhancements. arXiv 2024, arXiv:2410.17725. [Google Scholar] [CrossRef]
- Minderer, M.; Gritsenko, A.; Stone, A.; Neumann, M.; Weissenborn, D.; Dosovitskiy, A.; Mahendran, A.; Arnab, A.; Dehghani, M.; Shen, Z.; Wang, X.; Zhai, X.; Kipf, T.; Houlsby, N. Simple Open-Vocabulary Object Detection with Vision Transformers. in Proceedings of ECCV 2022, 2022. [Google Scholar]
- Algamdi, S. A. A behaviour-adaptive AI assistant enhancing accessibility and usability for blind users through real-time interaction personalization. Sci. Rep. 2026. [Google Scholar] [CrossRef] [PubMed]
- Huang, Y.; Xu, J.; Pei, B.; He, Y.; Chen, G.; Yang, L.; Chen, X.; Wang, Y.; Nie, Z.; Liu, J.; Fan, G.; Lin, D.; Fang, F.; Li, K.; Yuan, C.; Wang, Y.; Qiao, Y.; Wang, L. Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model. Shanghai Artif. Intell. Lab.>, arXiv 2024. [Google Scholar]
- Zheng, X.; Weng, Z.; Lyu, Y.; Jiang, L.; Xue, H.; Ren, B.; Paudel, D.; Sebe, N.; Van Gool, L.; Hu, X. Retrieval Augmented Generation and Understanding in Vision: A Survey and New Outlook. arXiv 2025. [Google Scholar] [CrossRef]
- Yu, X.; Saniie, J. Visual Impairment Spatial Awareness System for Indoor Navigation and Daily Activities. J. Imaging 2025, vol. 11(no. 9). [Google Scholar] [CrossRef] [PubMed]
- Baig, M. S. A.; Gillani, S. A.; Shah, S. M.; Aljawarneh, M.; Khan, A. A.; Siddiqui, M. H. AI-based Wearable Vision Assistance System for the Visually Impaired: Integrating Real-Time Object Recognition and Contextual Understanding Using Large Vision-Language Models. In Hamdard University / Applied Science Private University; 2024. [Google Scholar]
- Arsalwad, et al. YOLOInsight: Artificial Intelligence-Powered Assistive Device for Visually Impaired Using Internet of Things and Real-Time Object Detection. Cureus J. Comput. Sci. 2024. [Google Scholar] [CrossRef]
- Udayakumar, D.; Gopalakrishnan, S.; Raghuram, A.; Kartha, A.; Krishnan, A. K.; Ramamirtham, R.; Muthangi, R.; Raju, R. Artificial intelligence-powered smart vision glasses for the visually impaired. Indian J. Ophthalmol. 2025, vol. 73 suppl. 3, S490–S497. [Google Scholar] [CrossRef] [PubMed]
- Google, “Lookout — Assisted Vision,” Google Play Store. 2024. Available online: https://play.google.com/store/apps/details?id=com.google.android.apps.accessibility.reveal.
- Microsoft, “Seeing AI,” Microsoft Garage. March 2024. Available online: https://www.microsoft.com/en-us/garage/wall-of-fame/seeing-ai/.
- Birkl, 17 R.; Wofk, D.; Müller, M. MiDaS v3.1 — A Model Zoo for Robust Monocular Relative Depth Estimation. In Intel Labs; 2023. [Google Scholar]
- Zeng, N.; Hou, H.; Yu, F. R.; Shi, S.; He, Y. T. SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding. In Guangdong Laboratory of AI and Digital Economy / Shenzhen University; June 2025. [Google Scholar]
- Kumar, D. R.; Thakkar, H. K.; Merugu, S.; Gunjan, V. K.; Gupta, S. K. Object Detection System for Visually Impaired Persons Using Smartphone. In SpringerLink; 2022. [Google Scholar] [CrossRef]
- Shimakawa, M.; Matsushita, K.; Taguchi, I.; Okuma, C.; Kiyota, K. Smartphone Apps of Obstacle Detection for Visually Impaired and its Evaluation. In ACM; 2019. [Google Scholar] [CrossRef]
- Sridevi, N. T. Navigation Assistance with Real-Time Object Detection for Visually Impaired — A Mobile App. ReadyTensor, 2024. [Google Scholar]
- Richards, M. Software Architecture Patterns, 2nd ed.; O’Reilly Online Learning, 2023. [Google Scholar]
- Google Open Images Dataset V7. Available online: https://storage.googleapis.com/openimages/web/index.html.
- GBD 2019 Blindness and Vision Impairment Collaborators, Causes of blindness and vision impairment in 2020 and trends over 30 years, and prevalence of avoidable blindness in relation to VISION 2020: the Right to Sight: an analysis for the Global Burden of Disease Study. Lancet Glob. Health 2021, vol. 9(no. 2), e144–e160. [CrossRef]
- Envision AI, “Envision Glasses,” letsenvision.com. 2025. Available online: https://www.letsenvision.com/glasses.
- Bhattacharya, S.; Islam, R.; Hossain, M. S. SeeSay: Leveraging Large Language Models and Retrieval-Augmented Generation for Assistive Scene Understanding and Navigation for the Visually Impaired. arXiv 2024, arXiv:2410.03771. [Google Scholar]
- Shaikh, S. “Seeing AI App Launches on Android — Including new and updated features and new languages,” Microsoft Accessibility Blog. December 2023. Available online: https://blogs.microsoft.com/accessibility/seeing-ai-app-launches-on-android-including-new-and-updated-features-and-new-languages/.
- Be My Eyes, Introducing Be My AI (formerly Virtual Volunteer) for People who are Blind or Have Low Vision, Powered by OpenAI’s GPT-4. bemyeyes.com, March 2023. Available online: https://www.bemyeyes.com/news/introducing-be-my-ai-formerly-virtual-volunteer-for-people-who-are-blind-or-have-low-vision-powered-by-openais-gpt-4/.
- OrCam Technologies. OrCam MyEye 3 Pro — The most advanced wearable solution for visual impairment. orcam.com. 2024. Available online: https://www.orcam.com/en-us/orcam-myeye-3-pro.
- Shaikh, S. “What’s new with Seeing AI,” Microsoft Accessibility Blog, August 2019. Available online: https://blogs.microsoft.com/accessibility/seeing-ai-2/.
| Configuration | mAP@0.5 | p95 Latency (CPU) | p95 Latency (GPU) | Action Accuracy | Memory Recall |
| Baseline (v1) — YOLO+BLIP only | 0.578* | 7,186 ms | 288 ms | N/A | N/A |
| + Depth only | 0.578* | 13,651 ms | 289 ms | 72.5% | N/A |
| + Memory only | 0.578* | 9,190 ms | 380 ms† | 45.0% | 100.0% |
| Full v2 (all three) | 0.578* | 10,007 ms | 3,637 ms | 100.0% | 100.0% |
| Dimension | Seeing AI | Envision (App/Glasses) | Be My AI | OrCam MyEye 3 Pro | My Eye AI v2 |
| OCR / Text reading | ● | ● | ● | ● | ● |
| Object / scene description | ● | ● | ● | ◐ | ● |
| Face recognition | ● | ● | ○ | ● | ◐ |
| Live human assistance | ○ | ● (Ally / Aira) | ● (volunteers) | ○ | ○ |
| Monocular metric depth estimation | ○ | ○ | ○ | ○ | ● |
| Persistent temporal scene memory (RAG) | ○ | ○ | ○ | ○ | ● |
| Agentic LLM reasoning with composite urgency score | ○ | ◐ (Ask Envision / ally Q&A) | ◐ (Be My AI conversational Q&A) | ◐ (Just Ask Q&A) | ● |
| Offline / on-device operation | ◐ (limited) | ○ | ○ | ● (core functions) | ○ |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.