Submitted:
06 November 2025
Posted:
07 November 2025
You are already at the latest version
Abstract
Keywords:
1. Introduction

- We propose RS-ZeroSeg, a novel end-to-end architecture specifically designed for Open-Vocabulary Remote Sensing Image Semantic Segmentation, which effectively integrates general VLM capabilities with specialized remote sensing knowledge.
- We introduce three key architectural innovations: the Dual-Stream Feature Extractor (DSFE), the Multi-Scale Contextual Alignment Module (MS-CAM), and the Category-Adaptive Refinement Head (CARH), each tailored to address unique challenges in remote sensing imagery.
- We achieve state-of-the-art performance on multiple challenging remote sensing benchmarks (FLAIR, FAST, ISPRS Potsdam, FloodNet), demonstrating RS-ZeroSeg’s superior generalization ability and robustness compared to existing methods.
- We provide comprehensive ablation studies and in-depth analyses of various training strategies, including the impact of different training datasets and partial freezing policies for VLM and remote sensing backbones, offering valuable insights for future research in this domain.
2. Related Work
2.1. Open-Vocabulary Semantic Segmentation in Remote Sensing
2.2. Vision-Language Models for Image Segmentation
3. Method
3.1. Overall Architecture of RS-ZeroSeg

3.2. Dual-Stream Feature Extractor (DSFE)
3.3. Multi-Scale Contextual Alignment Module (MS-CAM)
3.4. Category-Adaptive Refinement Head (CARH)
3.5. Training Objective
3.6. Training Strategy and Implementation Details
4. Experiments
4.1. Experimental Setup
4.2. Comparison with State-of-the-Art Methods
4.3. Ablation Studies
4.4. Impact of Different Training Data Sources
4.5. Analysis of Partial Freezing Strategies
4.6. Generalization to Novel Categories
4.7. Efficiency Analysis
4.8. Impact of VLM Backbone Choice
4.9. Sensitivity to Alignment Loss Weight ()
4.10. Human Evaluation
5. Conclusions
References
- Shin, R.; Lin, C.; Thomson, S.; Chen, C.; Roy, S.; Platanios, E.A.; Pauls, A.; Klein, D.; Eisner, J.; Van Durme, B. Constrained Language Models Yield Few-Shot Semantic Parsers. In Proceedings of the Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics; 2021; pp. 7699–7715. [Google Scholar] [CrossRef]
- Zhou, Y.; Li, X.; Wang, Q.; Shen, J. Visual In-Context Learning for Large Vision-Language Models. In Proceedings of the Findings of the Association for Computational Linguistics, ACL 2024, Association for Computational Linguistics. Bangkok, Thailand and virtual meeting, August 11-16, 2024; pp. 15890–15902. [Google Scholar]
- Long, Q.; Wu, Y.; Wang, W.; Pan, S.J. Does in-context learning really learn? rethinking how large language models respond and solve tasks via in-context learning. arXiv preprint, 2024; arXiv:2404.07546. [Google Scholar]
- Wang, X.; Ruder, S.; Neubig, G. Multi-view Subword Regularization. In Proceedings of the Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. Association for Computational Linguistics; 2021; pp. 473–482. [Google Scholar] [CrossRef]
- Shan, X.; Wu, D.; Zhu, G.; Shao, Y.; Sang, N.; Gao, C. Open-Vocabulary Semantic Segmentation with Image Embedding Balancing. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2024, Seattle, WA, USA, June 16-22, 2024; IEEE, 2024; pp. 28412–28421. [Google Scholar] [CrossRef]
- Cho, S.; Shin, H.; Hong, S.; Arnab, A.; Seo, P.H.; Kim, S. CAT-Seg: Cost Aggregation for Open-Vocabulary Semantic Segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2024, Seattle, WA, USA, June 16-22, 2024. IEEE, 2024, pp. 4113–4123. [CrossRef]
- Csordás, R.; Irie, K.; Schmidhuber, J. The Devil is in the Detail: Simple Tricks Improve Systematic Generalization of Transformers. In Proceedings of the Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, 2021, pp. 619–634. [CrossRef]
- Chen, Z.; Huang, H.; Liu, B.; Shi, X.; Jin, H. Semantic and Syntactic Enhanced Aspect Sentiment Triplet Extraction. In Proceedings of the Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021. Association for Computational Linguistics, 2021, pp. 1474–1483. [CrossRef]
- Wang, D.; Ding, N.; Li, P.; Zheng, H. CLINE: Contrastive Learning with Semantic Negative Examples for Natural Language Understanding. In Proceedings of the Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). Association for Computational Linguistics, 2021, pp. 2332–2342. [CrossRef]
- Montariol, S.; Martinc, M.; Pivovarova, L. Scalable and Interpretable Semantic Change Detection. In Proceedings of the Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. Association for Computational Linguistics, 2021, pp. 4642–4652. [CrossRef]
- Tian, Y.; Xu, S.; Cao, Y.; Wang, Z.; Wei, Z. An Empirical Comparison of Machine Learning and Deep Learning Models for Automated Fake News Detection. Mathematics 2025, 13. [Google Scholar] [CrossRef]
- Long, Q.; Wang, M.; Li, L. Generative Imagination Elevates Machine Translation. In Proceedings of the Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, 2021, pp. 5738–5748.
- Hu, Y.; Lee, C.H.; Xie, T.; Yu, T.; Smith, N.A.; Ostendorf, M. In-Context Learning for Few-Shot Dialogue State Tracking. In Proceedings of the Findings of the Association for Computational Linguistics: EMNLP 2022. Association for Computational Linguistics, 2022, pp. 2627–2643. [CrossRef]
- Zhou, Y.; Shen, J.; Cheng, Y. Weak to strong generalization for large language models with multi-capabilities. In Proceedings of the The Thirteenth International Conference on Learning Representations, 2025.
- Huang, X.; Lin, Z.; Sun, F.; Zhang, W.; Tong, K.; Liu, Y. Enhancing Document-Level Question Answering via Multi-Hop Retrieval-Augmented Generation with LLaMA 3. arXiv preprint, 2025; arXiv:2506.16037. [Google Scholar]
- Huang, X.; Wang, Z.; Liu, X.; Tian, Y.; Leng, Q. Towards Interpretable and Consistent Multi-Step Mathematical Reasoning in Large Language Models. Available at SSRN 5680042 2025. [Google Scholar]
- Long, Q.; Deng, Y.; Gan, L.; Wang, W.; Pan, S.J. Backdoor attacks on dense retrieval via public and unintentional triggers. In Proceedings of the Second Conference on Language Modeling; 2025. [Google Scholar]
- Koto, F.; Lau, J.H.; Baldwin, T. IndoBERTweet: A Pretrained Language Model for Indonesian Twitter with Effective Domain-Specific Vocabulary Initialization. In Proceedings of the Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, 2021, pp. 10660–10668. [CrossRef]
- Xu, J.; Zhou, H.; Gan, C.; Zheng, Z.; Li, L. Vocabulary Learning via Optimal Transport for Neural Machine Translation. In Proceedings of the Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). Association for Computational Linguistics, 2021, pp. 7361–7373. [CrossRef]
- Che, W.; Feng, Y.; Qin, L.; Liu, T. N-LTP: An Open-source Neural Language Technology Platform for Chinese. In Proceedings of the Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing: System Demonstrations. Association for Computational Linguistics, 2021, pp. 42–49. [CrossRef]
- Hardalov, M.; Arora, A.; Nakov, P.; Augenstein, I. Cross-Domain Label-Adaptive Stance Detection. In Proceedings of the Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, 2021, pp. 9011–9028. [CrossRef]
- Wang, P.; Zhu, Z.; Liang, D. Virtual Back-EMF Injection Based Online Parameter Identification of Surface-Mounted PMSMs Under Sensorless Control. IEEE Transactions on Industrial Electronics 2024. [Google Scholar] [CrossRef]
- Lin, Z.; Lan, J.; Anagnostopoulos, C.; Tian, Z.; Flynn, D. Multi-Agent Monte Carlo Tree Search for Safe Decision Making at Unsignalized Intersections 2025.
- Wang, P.; Zhu, Z.Q.; Feng, Z. Novel Virtual Active Flux Injection-Based Position Error Adaptive Correction of Dual Three-Phase IPMSMs Under Sensorless Control. IEEE Transactions on Transportation Electrification 2025. [Google Scholar] [CrossRef]
- Gu, J.; Stefani, E.; Wu, Q.; Thomason, J.; Wang, X. Vision-and-Language Navigation: A Survey of Tasks, Methods, and Future Directions. In Proceedings of the Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, 2022, pp. 7606–7623. [CrossRef]
- Huang, P.Y.; Patrick, M.; Hu, J.; Neubig, G.; Metze, F.; Hauptmann, A. Multilingual Multimodal Pre-training for Zero-Shot Cross-Lingual Transfer of Vision-Language Models. In Proceedings of the Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. Association for Computational Linguistics, 2021, pp. 2443–2459. [CrossRef]
- Sun, S.; Chen, Y.C.; Li, L.; Wang, S.; Fang, Y.; Liu, J. LightningDOT: Pre-training Visual-Semantic Embeddings for Real-Time Image-Text Retrieval. In Proceedings of the Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. Association for Computational Linguistics, 2021, pp. 982–997. [CrossRef]
- Ross, C.; Katz, B.; Barbu, A. Measuring Social Biases in Grounded Vision and Language Embeddings. In Proceedings of the Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. Association for Computational Linguistics, 2021, pp. 998–1008. [CrossRef]
- Zhou, Y.; Song, L.; Shen, J. Improving Medical Large Vision-Language Models with Abnormal-Aware Feedback. arXiv preprint, 2025; arXiv:2501.01377. [Google Scholar]
- Xu, Y.; Xu, Y.; Lv, T.; Cui, L.; Wei, F.; Wang, G.; Lu, Y.; Florencio, D.; Zhang, C.; Che, W.; et al. LayoutLMv2: Multi-modal Pre-training for Visually-rich Document Understanding. In Proceedings of the Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). Association for Computational Linguistics, 2021, pp. 2579–2591. [CrossRef]
- Tian, Z.; Lin, Z.; Zhao, D.; Zhao, W.; Flynn, D.; Ansari, S.; Wei, C. Evaluating scenario-based decision-making for interactive autonomous driving using rational criteria: A survey. arXiv preprint, 2025; arXiv:2501.01886. [Google Scholar]
- Lin, Z.; Lan, J.; Anagnostopoulos, C.; Tian, Z.; Flynn, D. Safety-Critical Multi-Agent MCTS for Mixed Traffic Coordination at Unsignalized Intersections. IEEE Transactions on Intelligent Transportation Systems, 2025; 1–15. [Google Scholar] [CrossRef]
- Wang, P.; Zhu, Z.; Liang, D. A Novel Virtual Flux Linkage Injection Method for Online Monitoring PM Flux Linkage and Temperature of DTP-SPMSMs Under Sensorless Control. IEEE Transactions on Industrial Electronics 2025. [Google Scholar] [CrossRef]
- Vu, T.; Lester, B.; Constant, N.; Al-Rfou’, R.; Cer, D. SPoT: Better Frozen Model Adaptation through Soft Prompt Transfer. In Proceedings of the Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, 2022, pp. 5039–5059. [CrossRef]
- Agrawal, M.; Hegselmann, S.; Lang, H.; Kim, Y.; Sontag, D. Large language models are few-shot clinical information extractors. In Proceedings of the Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, 2022, pp. 1998–2022. [CrossRef]


| Method | FLAIR | FAST | Potsdam | FloodNet | Avg |
|---|---|---|---|---|---|
| EBSeg [5] | 21.26 | 18.53 | 5.68 | 35.26 | 20.18 |
| CAT-SEG [6] | 19.99 | 13.90 | 38.79 | 37.89 | 27.64 |
| SCAN [7] | 18.49 | 8.56 | 5.60 | 39.23 | 17.97 |
| SED [8] | 14.65 | 12.63 | 28.64 | 22.57 | 19.62 |
| GSNet [4] | 20.00 | 16.61 | 45.75 | 42.63 | 31.25 |
| RS-ZeroSeg (Ours) | 20.55 | 17.02 | 46.10 | 43.08 | 31.69 |
| Variant | FLAIR | FAST | Potsdam | FloodNet | Avg |
|---|---|---|---|---|---|
| w/o MS-CAM | 18.12 | 14.50 | 39.20 | 40.50 | 28.08 |
| w/o CARH | 18.70 | 15.80 | 41.30 | 41.80 | 29.40 |
| Base DSFE | 17.90 | 13.20 | 35.60 | 39.00 | 26.43 |
| Ours | 20.55 | 17.02 | 46.10 | 43.08 | 31.69 |
| CLIP | RSIB | FLAIR | FAST | Potsdam | FloodNet | Avg |
|---|---|---|---|---|---|---|
| Freeze | Full | 11.80 | 12.90 | 22.50 | 17.90 | 16.28 |
| Freeze | Attention | 11.90 | 13.80 | 28.40 | 29.10 | 20.80 |
| Freeze | Freeze | 16.20 | 12.80 | 29.50 | 30.60 | 22.28 |
| Full | Full | 9.80 | 4.50 | 21.30 | 27.00 | 15.65 |
| Full | Attention | 15.90 | 14.60 | 39.00 | 36.10 | 26.40 |
| Full | Freeze | 18.20 | 10.70 | 25.50 | 33.10 | 21.88 |
| Attention | Full | 19.60 | 14.50 | 37.40 | 37.20 | 27.18 |
| Attention | Attention | 20.10 | 15.20 | 39.40 | 38.10 | 28.20 |
| Attention | Freeze | 20.55 | 17.02 | 46.10 | 43.08 | 31.69 |
| Category Set | FLAIR | FAST | Potsdam | FloodNet |
|---|---|---|---|---|
| Base Categories | 25.10 | 20.80 | 50.20 | 45.15 |
| Novel Categories | 15.80 | 13.20 | 41.00 | 40.50 |
| VLM Backbone | FLAIR | FAST | Potsdam | FloodNet | Avg |
|---|---|---|---|---|---|
| CLIP-ViT-B/32 | 18.90 | 16.10 | 42.50 | 39.80 | 29.33 |
| CLIP-ViT-B/16 | 20.00 | 16.50 | 44.80 | 42.10 | 30.85 |
| CLIP-ViT-L/14 | 20.55 | 17.02 | 46.10 | 43.08 | 31.69 |
| FLAIR | FAST | Potsdam | FloodNet | Avg | |
|---|---|---|---|---|---|
| 0.0 | 19.80 | 16.30 | 43.20 | 41.50 | 30.20 |
| 0.1 | 20.10 | 16.70 | 44.50 | 42.20 | 30.88 |
| 0.2 | 20.55 | 17.02 | 46.10 | 43.08 | 31.69 |
| 0.3 | 20.30 | 16.80 | 45.70 | 42.70 | 31.38 |
| 0.4 | 19.90 | 16.50 | 44.90 | 42.00 | 30.83 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2025 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).