Submitted:
17 September 2024
Posted:
17 September 2024
You are already at the latest version
Abstract
Keywords:
1. Introduction and Related Work
2. Methodology
- Auxiliary: These entities are detected to determine if a pavement detection failure occurred due to occlusion rather than incorrect detection. Prompts are "car", "vehicle", "pole", "tree".
- Sidepaths: This query is primarily aimed at detecting sidewalks, but it may also retrieve other sidepaths, such as paved shoulders. Prompts are "sidewalk", "sidepath", "sideway", "sidetrack", and "lateral".
- Roads: Focused on detecting motorised pathways. Prompts are "road" and "street".
- Pathways: These are intended to detect any kind of traversable way. Prompts are "way", "path", "pathway", "pavement", and "track".
- Surface Pavement Types: Prompts directly target the property, with additional ones included for broader testing. Prompts are "sett", "grass", "cobblestone", "earth", "soil", "dirt", "sand", "concrete", "paving stones", "chip seal", "gravel", "compacted", "asphalt", "concrete plates", and "ground".
- Abstract Concepts: Words that do not have a unique or physical representation are used to test the model's responses. Prompts include anything, nothing, something, void, and thing.
3. Results and Discussion
3.1. Free Prompting and Random Imagery Evaluation
3.2. Evaluation of Open-Vocabulary Algorithms for Standardized Pavement Classification
3.3. Combine and Fine-Tune Models
4. Conclusions
Supplementary Materials
Author Contributions
Funding
Data Availability Statement
Acknowledgments
Conflicts of Interest
References
- Hamim, O.F.; Kancharla, S.R.; Ukkusuri, S. Mapping Sidewalks on a Neighborhood Scale from Street View Images. Environment and Planning B: Urban Analytics and City Science 2023. https://doi.org/10.1177/23998083231200445. [CrossRef]
- Serna, A.; Marcotegui, B. Urban Accessibility Diagnosis from Mobile Laser Scanning Data. ISPRS Journal of Photogrammetry and Remote Sensing 2013, 84, 23–32. https://doi.org/10.1016/j.isprsjprs.2013.07.001. [CrossRef]
- Vestena, K. de M.; Camboim, S.; Santos, D.R. dos OSM Sidewalkreator: A QGIS Plugin for an Automated Drawing of Sidewalk Networks for OpenStreetMap. European Journal of Geography 2023. https://doi.org/10.48088/ejg.k.ves.14.4.066.084. [CrossRef]
- Wood, J. Sidewalk City: Remapping Public Spaces in Ho Chi Minh City. Geographical Review 2016, 108, 486–488. https://doi.org/10.1111/gere.12239. [CrossRef]
- Zhou, Z.; Lin, Y.; Li, Y. Large Language Model Empowered Participatory Urban Planning 2024.
- Nadkarni, P.M.; Ohno-Machado, L.; Chapman, W.W. Natural Language Processing: An Introduc-Tion. Journal of the American Medical Informatics Association 2011, 18, 544–551. https://doi.org/10.1136/amiajnl-2011-000464. [CrossRef]
- Wang, X.; Ji, L.; Yan, K.; Sun, Y.; Song, R. Expanding the Horizons: Exploring Further Steps in Open-Vocabulary Segmentation. In Pattern recognition and computer vision; Liu, Q., Wang, H., Ma, Z., Zheng, W., Zha, H., Chen, X., Wang, L., Ji, R., Eds.; Springer Nature Singapore, 2024; pp. 407–419.
- Eichstaedt, J.C.; Kern, M.L.; Yaden, D.B.; Schwartz, H.A.; Giorgi, S.; Park, G.; Hagan, C.A.; Tobolsky, V.A.; Smith, L.K.; Buffone, A.; et al. Closed- and Open-Vocabulary Approaches to Text Analysis: A Review, Quantitative Comparison, and Recommendations. Psychological Methods 2021, 26, 398–427. https://doi.org/10.1037/met0000349. [CrossRef]
- Zhu, C.; Chen, L. A survey on open-vocabulary detection and segmentation: Past, present, and future 2023.
- Zareian, A.; Dela Rosa, K.; Hu, D.H.; Chang, S. Open-Vocabulary Object Detection Using Captions. CoRR 2020. https://doi.org/2011.10678.
- Yang, G.; Ye, Z.; Zhang, R.; Huang, K. A Comprehensive Survey of Zero-Shot Image Classification: Methods, Implementation, and Fair Evaluation. Applied Computing and Intelligence 2022, 2, 1–31.
- Lampert, C.H.; Nickisch, H.; Harmeling, S. Attribute-Based Classification for Zero-Shot Visual Object Cat-Egorization. IEEE Transactions on Pattern Analysis and Machine Intelligence 2014, 36, 453–465. https://doi.org/10.1109/TPAMI.2013.140. [CrossRef]
- Rohrbach, M.; Stark, M.; Schiele, B. Evaluating Knowledge Transfer and Zero-Shot Learning in a Large-Scale Setting. In CVPR 2011; IEEE, 2011; pp. 1641–1648.
- Zhang, D.; Han, J.; Cheng, G.; Yang, M.-H. Weakly Supervised Object Localization and Detection: A Survey. IEEE Trans. Pattern Anal. Mach. Intell. 2021, 1–1. https://doi.org/10.1109/TPAMI.2021.3074313. [CrossRef]
- Vo, H.V.; Siméoni, O.; Gidaris, S.; Bursuc, A.; Pérez, P.; Ponce, J. Active Learning Strategies for Weakly-Supervised Object Detection 2022.
- Blasiis, M.D.; Benedetto, A.; Fiani, M. Mobile Laser Scanning Data for the Evaluation of Pavement Surface Distress. Remote Sensing 2020, 12, 942. https://doi.org/10.3390/rs12060942. [CrossRef]
- Praticò, F.; Vaiana, R. A Study on the Relationship between Mean Texture Depth and Mean Profile Depth of Asphalt Pavements. Construction and Building Materials 2015, 101, 72–79. https://doi.org/10.1016/j.conbuildmat.2015.10.021. [CrossRef]
- Fidalgo, C.D.; Santos, I.M.; Nogueira, C. de A.; Portugal, M.C.S.; Martins, L.M.T. Urban Sidewalks, Dysfunction and Chaos on the Projected Floor. The Search for Accessible Pavements and Sustainable Mobility. In Proceedings of the Proceedings of the 7th International Congress on Scientific Knowledge; 2021.
- Vaitkus, A.; Andriejauskas, T.; Šernas, O.; Čygas, D.; Laurinavičius, A. Definition of concrete and composite precast concrete pavements texture. Transport 2019. https://doi.org/10.3846/transport.2019.10411. [CrossRef]
- Zeng, Z.; Boehm, J. Exploration of an Open Vocabulary Model on Semantic Segmentation for Street Scene Imagery. ISPRS International Journal of Geo-Information 2024, 13, 153. https://doi.org/10.3390/ijgi13050153. [CrossRef]
- Cordts, M.; Omran, M.; Ramos, S.; Rehfeld, T.; Enzweiler, M.; Benenson, R.; Franke, U.; Roth, S.; Schiele, B. The Cityscapes Dataset for Semantic Urban Scene Understanding. In Proceedings of the Proceedings of the IEEE conference on com-puter vision and pattern recognition; 2016; pp. 3213–3223.
- Geiger, A.; Lenz, P.; Stiller, C.; Urtasun, R. Vision Meets Robotics: The KITTI Dataset. The International Journal of Robotics Research 2013, 32, 1231–1237.
- Ros, G.; Sellart, L.; Materzynska, J.; Vazquez, D.; Lopez, A.M. The SYNTHIA Dataset: A Large Collection of Synthetic Images for Semantic Segmentation of Urban Scenes. In Proceedings of the Proceedings of the IEEE conference on computer vision and pattern recognition; 2016; pp. 3234–3243.
- Yu, H.; Yang, Z.; Tan, L.; Wang, Y.; Sun, W.; Sun, M.; Tang, Y. Methods and Datasets on Semantic Seg-Mentation: A Review. Neurocomputing 2018, 304, 82–103. https://doi.org/10.1016/j.neucom.2018.03.037. [CrossRef]
- Hao, S.; Zhou, Y.; Guo, Y. A brief survey on semantic segmentation with deep learning. Neurocomputing 2020, 406, 302–321. https://doi.org/10.1016/j.neucom.2019.11.118. [CrossRef]
- Mo, Y.; Wu, Y.; Yang, X.; Liu, F.; Liao, Y. Review the State-of-the-Art Technologies of Semantic Segmenta-Tion Based on Deep Learning. Neurocomputing 2022, 493, 626–646. https://doi.org/10.1016/j.neucom.2022.01.005. [CrossRef]
- Zou, J.; Guo, W.; Wang, F. A Study on Pavement Classification and Recognition Based on VGGNet-16 Transfer Learning. Electronics 2023, 12, 3370. https://doi.org/10.3390/electronics12153370. [CrossRef]
- Zhang, C.; Nateghinia, E.; Miranda-Moreno, L.F.; Sun, L. Pavement Distress Detection Using Convolu-Tional Neural Network (CNN): A Case Study in Montreal, Canada. International Journal of Transportation Science and Technology 2022, 11, 298–309. https://doi.org/10.1016/j.ijtst.2021.04.008. [CrossRef]
- Riid, A.; Lõuk, R.; Pihlak, R.; Tepljakov, A.; Vassiljeva, K. Pavement Distress Detection with Deep Learning Using the Orthoframes Acquired by a Mobile Mapping System. Applied Sciences 2019, 9, 4829. https://doi.org/10.3390/app9224829. [CrossRef]
- Mesquita, R.; Ren, T.I.; Mello, C.; Silva, M. Street Pavement Classification Based on Navigation through Street View Imagery. AI & SOCIETY 2022. https://doi.org/10.1007/s00146-022-01520-0. [CrossRef]
- Hosseini, M.; Miranda, F.; Lin, J.; Silva, C.T. CitySurfaces: City-scale semantic segmentation of sidewalk materials. CoRR 2022, 2201.
- Radford, A.; Kim, J.W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; Mishkin, P.; Clark, J.; et al. Learning Transferable Visual Models from Natural Language Supervision 2021.
- Liu, S.; Zeng, Z.; Ren, T.; Li, F.; Zhang, H.; Yang, J.; Li, C.; Yang, J.; Su, H.; Zhu, J.; et al. Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection 2023.
- Grinberger, A.Y.; Minghini, M.; Juhász, L.; Yeboah, G.; Mooney, P. OSM Science—The Academic Study of the OpenStreetMap Project, Data, Contributors, Community, and Applications. IJGI 2022, 11, 230. https://doi.org/10.3390/ijgi11040230. [CrossRef]
- Zeng, Y.; Huang, Y.; Zhang, J.; Jie, Z.; Chai, Z.; Wang, L. Investigating Compositional Challenges in Vision-Language Models for Visual Grounding. In Proceedings of the Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR); June 2024; pp. 14141–14151.
- Rajabi, N.; Kosecka, J. Q-GroundCAM: Quantifying Grounding in Vision Language Models via GradCAM 2024.
- Wang, S.; Kim, D.; Taalimi, A.; Sun, C.; Kuo, W. Learning Visual Grounding from Generative Vision and Language Model 2024.
- Quarteroni, S.; Dinarelli, M.; Riccardi, G. Ontology-Based Grounding of Spoken Language Understanding. In Proceedings of the 2009 IEEE Workshop on Automatic Speech Recognition & Understanding; IEEE: Moreno, Italy, December 2009; pp. 438–443.
- Baldazzi, T.; Bellomarini, L.; Ceri, S.; Colombo, A.; Gentili, A.; Sallinger, E. Fine-Tuning Large Enterprise Language Models via Ontological Reasoning 2023.
- Jullien, M.; Valentino, M.; Freitas, A. Do Transformers Encode a Foundational Ontology? Probing Abstract Classes in Natural Language 2022.
- FRC CSC RAS / Moscow, Russia; Larionov, D.; RUDN University / Moscow, Russia; Shelmanov, A.; Skoltech / Moscow, Russia; FRC CSC RAS / Moscow, Russia; Chistova, E.; FRC CSC RAS / Moscow, Russia; RUDN University / Moscow, Russia; Smirnov, I.; et al. Semantic Role Labeling with Pretrained Language Models for Known and Unknown Predicates. In Proceedings of the Proceedings - Natural Language Processing in a Deep Learning World; Incoma Ltd., Shoumen, Bulgaria, October 22 2019; pp. 619–628.
- Smith, M.K.; Welty, C.; McGuinness, D.L. OWL Web Ontology Language Guide 2004.
- Kirillov, A.; Mintun, E.; Ravi, N.; Mao, H.; Rolland, C.; Gustafson, L.; Xiao, T.; Whitehead, S.; Berg, A.C.; Lo, W.Y.; et al. Segment anything 2023.
- Meta Mapillary. GitHub XXXX.
- Vestena, K. GitHub - kauevestena/deep_pavements_dataset. GitHub XXXX.
- Fan, Q.; Tao, X.; Ke, L.; Ye, M.; Zhang, Y.; Wan, P.; Wang, Z.; Tai, Y.-W.; Tang, C.-K. Stable Segment Anything Model 2023.
- Hetang, C.; Xue, H.; Le, C.; Yue, T.; Wang, W.; He, Y. Segment Anything Model for Road Network Graph Extraction 2024.
- Son, J.; Jung, H. Teacher–Student Model Using Grounding DINO and You Only Look Once for Multi-Sensor-Based Object Detection. Applied Sciences 2024, 14, 2232. https://doi.org/10.3390/app14062232. [CrossRef]
- Dong, X.; Bao, J.; Zhang, T.; Chen, D.; Gu, S.; Zhang, W.; Yuan, L.; Chen, D.; Wen, F.; Yu, N. CLIP Itself Is a Strong Fine-Tuner: Achieving 85.7% and 88.0% Top-1 Accuracy with ViT-B and ViT-L on ImageNet 2022.
- Nguyen, T.; Ilharco, G.; Wortsman, M.; Oh, S.; Schmidt, L. Quality Not Quantity: On the Interaction between Dataset Design and Robustness of Clip. Advances in Neural Information Processing Systems 2022, 35, 21455–21469.
- Fang, A.; Ilharco, G.; Wortsman, M.; Wan, Y.; Shankar, V.; Dave, A.; Schmidt, L. Data Determines Distributional Robustness in Contrastive Language Image Pre-Training (Clip). In Proceedings of the International Conference on Machine Learning; PMLR, 2022; pp. 6216–6234.
- Tu, W.; Deng, W.; Gedeon, T. A Closer Look at the Robustness of Contrastive Language-Image Pre-Training (Clip). Advances in Neural Information Processing Systems 2024, 36.
- Mumuni, F.; Mumuni, A. Segment Anything Model for Automated Image Data Annotation: Empirical Studies Using Text Prompts from Grounding DINO 2024.
- Kaue-Vestena/Clip-Vit-Base-Patch32-Finetuned-Surface-Materials. Hugging Face. Title of Thesis. Level of Thesis, Degree-Granting University, Location of University, Date of Completion, 2024.
- Eimer, T.; Lindauer, M.; Raileanu, R. Hyperparameters in Reinforcement Learning and How To Tune Them 2023.
- Tong, Q.; Liang, G.; Bi, J. Calibrating the Adaptive Learning Rate to Improve Convergence of ADAM. Neurocomputing 2022, 481, 333–356. https://doi.org/10.1016/j.neucom.2022.01.014. [CrossRef]
- Reddi, S.J.; Kale, S.; Kumar, S. On the Convergence of Adam and Beyond 2019.
- Yu, J.; Wang, Z.; Vasudevan, V.; Yeung, L.; Seyedhosseini, M.; Wu, Y. CoCa: Contrastive Cap-Tioners Are Image-Text Foundation Models 2022.
- Code, P.W. Papers with Code - ImageNet Benchmark (Image Classification. GitHub 2024.








| Package | Model Name | Base Algorithm | Main Image Dataset | Model Size (GB) | Image Parameters (Millions) | Text Parameters (Millions) | Image GFLOPS | Text GFLOPS | Package | Model Name |
| CLIP | RN101 | CNN | OpenAI | 0.28 | 56.26 | 63.43 | 19.54 | 5.96 | CLIP | RN101 |
| CLIP | RN50x64 | CNN | OpenAI | 1.26 | 420.38 | 202.88 | 529.11 | 23.55 | CLIP | RN50x64 |
| CLIP | ViT-B/32 | VT | OpenAI | 0.34 | 87.85 | 63.43 | 8.82 | 5.96 | CLIP | ViT-B/32 |
| CLIP | ViT-L/14 | VT | OpenAI | 0.89 | 303.97 | 123.65 | 162.03 | 13.3 | CLIP | ViT-L/14 |
| Open CLIP | ViT-H-14-378-quickgelu | VT | dfn5b | 3.95 | 632.68 | 354.03 | 1006.96 | 47.09 | Open CLIP | ViT-H-14-378-quickgelu |
| Open CLIP | EVA02-E-14-plus | VT | laion2b_s9b_b144k | 10.1 | 4350.56 | 694.33 | 2264.33 | 97.86 | Open CLIP | EVA02-E-14-plus |
| row | Prompt | 1 | 2 | 3 | 4 | 5 | 6 | 7 | 8 | hit % |
| I | car | car | truck | car | car | car panel* | car | car | car | 100 |
| II | poles | street lamp | power pole | building pole | street lamp | street lamp | guard rail ** | street lamp | street lamp | 87.5 |
| III | vehicle | car | truck | car | truck | fence** | car | car | truck | 87.5 |
| IV | tree | tree | tree | tree | tree | tree | vegetation | tree | tree | 100 |
| row | Prompt | 1 | 2 | 3 | 4 | 5 | 6 | 7 | 8 | hit % |
| I | sidewalk | sidewalk | sidewalk | road | sidewalk | sidewalk | a sidepath | kerbs | sidewalk | 75 |
| II | sidepath | road ** | road ** | kerb | road ** | road | road ** | road | road | 0 |
| III | sideway | car panel | car | truck | car | car | car panel | road | car | 0 |
| IV | sidetrack | car | car panel | building | car | car panel ** | truck container | car | car | 0 |
| V | lateral | road ** | road | road and sidewalk | car ** | car | vegetation | road | car ** | 0 |
| VI | walkway | road and sidepath | road ** | road ** | road ** | partial sidewalk | all pavements | road | road ** | 12.5 |
| row | Prompt | 1 | 2 | 3 | 4 | 5 | 6 | 7 | 8 | hit % |
| I | street | road | road | road | road | road | road | road | road | 100 |
| II | road | road | road | road | road | road | road | road | road | 100 |
| row | Prompt | 1 | 2 | 3 | 4 | 5 | 6 | 7 | 8 | hit % |
| I | track | car | road | road | road | road | road | road | road | 87.5 |
| II | pavement | road | road | a path | road | road -- | road | walkway | road | 100 |
| III | way | bus | road markings | road | road | bus | road | sidewalk | road | 62.5 |
| IV | path | road | road | road | road | road | road | road | road | 100 |
| V | pathway | road | road | sidewalk | road | road | road | road | road | 100 |
| row | Prompt | 1 | 2 | 3 | 4 | 5 | 6 | 7 | 8 | hit % |
| I | sett | car ** | car ** | car ** | car ** | asphalt road ** | tree ** | car ** | car ** | 0 |
| II | grass | grass sidepath | grass sidewalk | grass sidewalk | grass in a garden | grass sidepath | plants ** | grass stripe | grass sidepath | 87.5 |
| III | cobblestone | grass sidepath | a pathway with stones | bush row | dirt sidepath | sidepath with vegetation | asphalt road | asphalt road | asphalt road | 0 |
| IV | earth | car ** | gravel sidewalk ** | asphalt road ** | car panel ** | asphalt road ** | a tire ** | asphalt road ** | car ** | 0 |
| V | soil | an unpaved sidepath | asphalt road ** | asphalt road | soy crop | car | soil sidepath | grass sidewalk | grass sidepath | 12.5 |
| VI | dirt | asphalt road ** | asphalt road ** | car ** | unpaved sidewalk ** | asphalt road ** | asphalt road and car panel ** | a grass slope ** | asphalt road ** | 0 |
| VII | sand | car ** | asphalt road ** | asphalt road ** | sand sidewalk | asphalt road ** | asphalt road ** | asphalt road (partial) ** | car | 12.5 |
| VIII | concrete | asphalt road ** | compacted pathway ** | asphalt road ** | asphalt road ** | concrete sidewalk | guard rail ** | asphalt road ** | asphalt road ** | 12.5 |
| IX | paving-stones | asphalt road | soil sidepath ** | concrete sidewalk ** | asphalt road | asphalt road ** | asphalt road | cobblestone sidewalk stripe | asphalt road | 0 |
| X | chipseal | car | truck | car | car | car | car | car | car | 0 |
| XI | gravel | asphalt road ** | asphalt pathway ** | asphalt pathway ** | sidepath with vegetation | car | asphalt road ** | car | wall | 0 |
| XII | compacted | truck | car | car | car | car | car | car | car | 0 |
| XIII | asphalt | asphalt road | asphalt road | asphalt road | asphalt road | asphalt pathway | asphalt road | asphalt road | asphalt road | 100 |
| XIV | Concrete plates | truck | nothing | nothing | truck | truck | truck | truck | a stairway | 25 |
| XV | ground | asphalt road ** | asphalt road ** | asphalt road and ground sidewalk | asphalt road and asphalt sidepath | asphalt road ** | asphalt road ** | asphalt road | asphalt road ** | 0 |
| row | Prompt | 1 | 2 | 3 | 4 | 5 | 6 | 7 | 8 | hit % |
| I | anything | car | landscape | road | person | road and sky | car panel | vegetation and sky | road and car panel | N/A |
| II | nothing | asphalt road and noise | asphalt road and noise | car | asphalt road and noise | noisy asphalt road and sky | sky and noise | truck | road, wall, and noise | N/A |
| III | something | car hood | car panel | traffic signal | car | car | nothing | truck | truck | N/A |
| IV | anything | road, sidewalk, and vegetation | noisy asphalt road and noisy sky | car | tree | landscape and noisy road | nothing | car | building | N/A |
| V | void | asphalt road | car | car | nothing | car | car | car | car | N/A |
| Hyperparameter | Original | Best Result |
| Batch size | 256 | 320 |
| Learning Rate | 5.00E-05 | 5.00E-06 |
| Weight Decay | 0.2 | 0.5 |
| AMSGrad Method | Deactivated | Activated |
| Betas | (0.9,0.98) | (0.9,0.98) |
| Epsilon | 1.00E-06 | 1.00E-06 |
| training epochs | 100 | 100 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2024 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).