Submitted:
23 June 2026
Posted:
25 June 2026
You are already at the latest version
Abstract
Keywords:
1. Introduction


2. Preliminaries
2.1. Vision-Language Models: A Brief Overview
2.1.1. Network Architectures
2.1.2. Training Strategy
2.2. Low-Level Vision: Formulation
3. Direct VLM Adaptation for Low-Level Vision
3.1. Visual Encoder Adaptation: Handling Details
| Dataset | Year | Related Method | Short Description | Type |
|---|---|---|---|---|
| Paired training samples [75] | 2024 | RestoreAgent [75] | A 23K-pair instruction dataset labeled with the optimal restoration pipeline |
Real |
| CleanBench [74] | 2025 | JarvisIR [74] | 150K synthetic pairs and 80K real degraded images accompanied by self-constructed instruction–response pairs |
Real and Syn |
| SlowAgent [32] | 2025 | HybridAgent [32] | A total of 70K pairs composed of degraded images, predicted degradation types, and tool-calling instructions. |
Syn |
| FeedbackAgent [32] | 2025 | HybridAgent [32] | 66K pairs used to determine whether an image is clean and whether further restoration is required. |
Syn |
| Low-level vision VQA dataset [77] | 2024 | UniProcessor [77] | Covering 30 degradation types and over 70K image patches, with diverse QA instructions generated via templates |
Syn |
| InstructIR Prompt Set [99] | 2024 | InstructIR [99] | An instruction set with 10K+ distinct prompts covering 7 restoration/ enhancement tasks |
Syn |
| PromptFix Dataset [88] | 2024 | PromptFix [88] | A 1,013,320-sample instruction-following dataset spanning 7 restoration/ editing tasks (e.g., dehazing, low-light, deblurring) |
Real and Syn |
| PixTalk Dataset [195] | 2025 | PixTalk [195] | A 106K RAW↔sRGB paired dataset with language prompts for photorealistic image processing and editing |
Real |
| Tri-IR [196] | 2025 | InstructRestore [196] | A 536,945-triplet dataset consisting of a high-quality image, a target-region description, and a region mask for region instruction restoration |
Real and Syn |
| Pico-Banana-400K [197] | 2025 | Pico-Banana-400K [197] | A 400K instruction-based image editing dataset built on real photos, paired with edited results for instruction-following enhancement |
Real and Syn |

3.2. Language Branch Adaptation: Bridging Modalities
| Generation approach | Method | Explanation | Utilized VLMs |
|---|---|---|---|
| Direct Generation | RAP-SR [54] | Automatically generates image descriptions using Florence-2. | Florence-2 [199] |
| RamIR [200] | Generates degradation () and clean () embeddings via BLIP. | BLIP [21] | |
| CyclicPrompt [102] | Encodes BLIP descriptions into , combining learnable and visual . | BLIP [21] | |
| LDR [51] | Describes weather types and occluded regions via language instructions. | LISA [201] | |
| Diff-Dehazer [202] | Derives positive prompts by removing “haze” from BLIP-2 descriptions. | BLIP-2 [203] | |
| DA-CLIP [29] | Uses triplets of HQ images, descriptive texts, and degradation labels. | BLIP [21] | |
| Progressive generation | PBI [48] | CogVLM describes masked objects; Mistral-7B rewrites them as addition instructions. | CogVLM [204], Mistral-7B [205] |
| MTGFusion [49] | Encodes six types of GPT-4 text prompts via BLIP into vectors. | GPT-4 [206], BLIP [21] | |
| Chen’s Method [50] | Fuses coarse-grained (BLIP-2) and fine-grained (MiniGPT-4) descriptions via templates. | BLIP-2 [203], MiniGPT-4 [207] | |
| Multi level generation | XPSR [89] | Generates high-level semantic and low-level detail descriptions via LLaVA. | LLaVA [22] |
| PASD [93] | Encodes object/location words (ResNet/YOLO) and scene contents (BLIP). | BLIP [21] | |
| MTGFusion [49] | Generates separate source, shared, and unique descriptions. | GPT-4 [206], BLIP [21] | |
| FILM [95] | Combines scene (BLIP-2), object (GRIT), and mask (SAM) prompts for ChatGPT generation. | BLIP-2 [203], ChatGPT [208] | |
| Region level generation | DreamPaint [209] | Generates descriptive text based on Kosmos-2 object detection boxes. | LLaVA-1.6-Vicuna-7B [187] |
| Anywhere [110] | Generates text descriptions for foreground regions extracted by RMBG-1.4. | Gemini-Pro-Vision [210] |
3.3. Output Head Adaptation: From Tokens to Pixels
3.4. Parameter-Efficient Fine-Tuning (PEFT) in Restoration
4. VLM as Auxiliary for Low-Level Vision
4.1. VLM as Semantic Provider
4.2. VLM as Degradation Interpreter
4.3. VLM as Quality Evaluator
4.4. VLM as Intelligent Controller

5. Extended Applications
5.1. Medical Image Processing
5.2. Remote Sensing Data Processing
5.3. Spatiotemporal and Geometric Modality Processing
| Tasks | Datasets | Size | Source | Type | Description |
|---|---|---|---|---|---|
| SR | DIV2K [230] | 800/100/100 | NTIRE 2017 | Syn | A high-resolution image benchmark containing diverse natural scenes with synthetically generated degradations. |
| RealSR [231] | 595 | ICCV 2019 | Real | A real-capture dataset obtained by varying the camera focal length under controlled settings. | |
| DRealSR [232] | 2,507 | ECCV 2020 | Real | A real-world dataset acquired with focal-length adjustments across indoor and outdoor environments. | |
| LLIE | LOLv1 [233] | 485/15 | BMVC 2018 | Real | A real-world paired low-/normal-light dataset collected under diverse illumination conditions. |
| LIME [234] | 10 | TIP 2016 | Real | A real-world low-light image set without paired ground truth, used for no-reference low-light enhancement evaluation. | |
| LOLv2 [235] | 1,589/200 | TIP 2021 | Real&Syn | A combined dataset including real pairs with multi-stage imaging and synthetic pairs using luminance transformations. | |
| Dehaze | RESIDE [236] | 13,000/990 | TIP 2019 | Real&Syn | A dehazing benchmark with synthetic and real subsets, covering diverse indoor/outdoor conditions. |
| NH-Haze [237] | 55 | CVPRW 2020 | Real | A real-captured outdoor dataset using controlled haze machines, providing paired hazy/clear images. | |
| Haze-4K [238] | 4,000 | MM 2021 | Syn | A synthetic benchmark including paired images with their transmission maps and atmospheric light annotations. | |
| Inpainting | CelebA [239] | 200,000 | ICCV 2015 | Real | A large-scale face benchmark with over 10,000 identities, used in semantic inpainting and human-centric restoration. |
| CelebA-HQ [240] | 30,000 | ArXiv 2017 | Real | A high-quality subset of CelebA produced through super-resolution, suitable for high-fidelity face restoration tasks. | |
| MSCOCO [241] | 328,124 | ECCV 2014 | Real | A general-purpose benchmark offering diverse scenes across 91 object categories with segmentation annotations. | |
| Derain | Rain100H [242] | 1,800/100 | CVPR 2017 | Syn | A synthetic benchmark simulating heavy rain streaks with multiple directional patterns for supervised deraining. |
| RainDrop [243] | 861/239 | CVPR 2018 | Real | A real-world dataset by capturing paired images through dual glass surfaces, modeling adherent raindrop. | |
| GT-RAIN [244] | 26,124/2,100 | ECCV 2022 | Real | A real paired deraining dataset with ground-truth clean/rainy pairs captured under controlled non-rain variations. | |
| Deblur | GoPro [245] | 2,103/1,111 | CVPR 2017 | Syn | A motion-blur benchmark with high-frame-rate video capture, approximating camera shake via temporal integration. |
| RealBlur [246] | 3,758/980 | ECCV 2020 | Real | A motion blur dataset acquired using paired short-/long-exposure imaging in both RAW and JPEG formats. | |
| HIDE [247] | 8,422 | ICCV 2019 | Syn | A human-centered dynamic-scene dataset with both near-field and far-field motion. | |
| CT denoising | 2016 NIH-AAPM-Mayo [248] | 5,936 | Med. Phys. 2017 | Real | A dataset with full-dose CT scans from 10 patients, widely used in dose-reduction and denoising studies. |
| Mayo-2016 [249] | 4,800/1,136 | Med. Phys. 2021 | Real | A dataset comprising abdominal CT images with normal-dose and simulated quarter-dose ones from 10 subjects. | |
| Mayo-2020 [249] | 2,400/580 | Med. Phys. 2021 | Real | A dataset with abdominal CT images from 100 patients reconstructed at of the routine radiation dose. |
| Methods | RealSR [231] | DRealSR [232] | Params | ||||||
|---|---|---|---|---|---|---|---|---|---|
| PSNR↑ | SSIM↑ | LPIPS↓ | MUSIQ↑ | PSNR↑ | SSIM↑ | LPIPS↓ | MUSIQ↑ | [M] | |
| LR Image | 27.80 | 0.7928 | 0.3356 | 34.40 | 28.02 | 0.8374 | 0.4410 | 20.54 | N/A |
| CoSeR [252] | 21.24 | 0.6109 | 0.2438 | 70.29 | 19.95 | 0.5350 | 0.2702 | 70.18 | 720+1936 |
| XPSR [89] | 24.19 | 0.6870 | 0.3517 | 70.23 | 26.62 | 0.7220 | 0.3864 | 67.84 | 868+1066 |
| PASD [93] | 25.93 | 0.7105 | 0.2806 | 65.60 | 29.09 | 0.7937 | 0.2893 | 34.63 | 361+1314 |
| MegaSR [94] | 23.49 | 0.6903 | 0.3072 | 70.03 | 25.74 | 0.7367 | 0.3258 | 64.14 | — |
| PURE [215] | 22.83 | 0.6079 | 0.3821 | 72.37 | 24.73 | 0.6459 | 0.3452 | 73.28 | 7000 |
| Methods | CBSD68 [253] | Urban100 [254] | Params | ||||
|---|---|---|---|---|---|---|---|
| [M] | |||||||
| Noised Image | 24.84/0.592 | 20.54/0.418 | 15.00/0.219 | 24.95/0.629 | 20.69/0.480 | 15.18/0.292 | N/A |
| InstructIR-3D [99] | 34.15/0.933 | 31.52/0.890 | 28.30/0.804 | 34.12/0.945 | 31.80/0.917 | 28.63/0.861 | 16+17 |
| TextPromptIR [255] | 34.17/0.936 | 31.52/0.893 | 28.26/0.805 | 34.76/0.951 | 32.54/0.929 | 29.47/0.882 | 124+110 |
| CLIPDenoising [256] | 33.97/0.930 | 31.02/0.878 | 26.69/0.731 | 33.15/0.930 | 30.72/0.893 | 26.27/0.769 | 11+9 |
| VLU-Net [53] | 34.13/0.935 | 31.48/0.892 | 28.23/0.804 | 34.92/0.952 | 32.71/0.930 | 29.61/0.883 | 35+88 |
| DFPIR [257] | 34.32/0.934 | 31.71/0.897 | 28.49/0.814 | 34.79/0.952 | 32.57/0.930 | 29.53/0.883 | 31+63 |
| Methods | Snow100K-L [258] | Outdoor-Rain [259] | RainDrop [243] | Params | |||
|---|---|---|---|---|---|---|---|
| PSNR↑ | SSIM↑ | PSNR↑ | SSIM↑ | PSNR↑ | SSIM↑ | [M] | |
| Degraded Image | 18.67 | 0.6468 | 12.89 | 0.5293 | 23.81 | 0.8363 | N/A |
| MPerceiver [260] | 31.11 | 0.9180 | 31.25 | 0.9246 | 33.62 | 0.9300 | — |
| ADSM [70] | 30.81 | 0.9058 | 31.28 | 0.9227 | 30.71 | 0.9188 | — |
| VLCIR [45] | 32.28 | 0.9296 | 32.62 | 0.9447 | 33.39 | 0.9489 | — |
| CyclicPrompt [102] | 32.16 | 0.9265 | 32.81 | 0.9371 | 32.57 | 0.9454 | 30+486 |
| M2Restore [113] | 31.20 | 0.9110 | 32.01 | 0.9550 | 31.73 | 0.9430 | — |
| Methods | LOLv1 [233] | LOLv2-Real [235] | LIME [234] | Params | ||||
|---|---|---|---|---|---|---|---|---|
| PSNR↑ | SSIM↑ | LPIPS↓ | PSNR↑ | SSIM↑ | LPIPS↓ | NIQE↓ | [M] | |
| Low-Light Image | 7.77 | 0.191 | 0.417 | 9.72 | 0.196 | 0.394 | 4.351 | N/A |
| NeRCo [261] | 22.95 | 0.785 | 0.311 | 25.17 | 0.785 | 0.338 | 3.803 | 46+102 |
| CFWD [66] | 29.19 | 0.872 | 0.197 | 29.86 | 0.891 | 0.193 | 3.568 | 22+151 |
| CLIP-LLA [262] | 22.66 | 0.882 | 0.140 | 20.23 | 0.844 | 0.164 | — | 2+151 |
| HVI-CIDNet+ [263] | 28.85 | 0.894 | 0.058 | 24.31 | 0.873 | 0.107 | 3.760 | 69+246 |
| GPP-LLIE [119] | 27.51 | 0.872 | 0.081 | 29.23 | — | 0.055 | 4.24 | 76+55 |
| Methods | Rain100H [242] | Rain100L [242] | Test100 [264] | Params | Time | |||
|---|---|---|---|---|---|---|---|---|
| PSNR↑ | SSIM↑ | PSNR↑ | SSIM↑ | PSNR↑ | SSIM↑ | [M] | [s] | |
| Rainy Image | 12.13 | 0.344 | 25.52 | 0.814 | 21.11 | 0.640 | N/A | N/A |
| DA-CLIP [29] | 27.07 | 0.817 | 34.50 | 0.948 | 27.95 | 0.857 | 49+125 | 10.898 |
| TVI-Derain [47] | 30.51 | 0.893 | 37.72 | 0.973 | 31.17 | 0.912 | 20+278 | — |
| TexDepth-Derain [265] | 33.97 | 0.949 | 41.97 | 0.991 | — | — | 26+1469 | — |
| PTG-RM-cRestormer [52] | 31.77 | 0.913 | 39.27 | 0.985 | 32.30 | 0.934 | 27+86 | 0.142 |
| AWRaCLe [266] | 27.20 | 0.840 | 35.70 | 0.966 | 27.58 | 0.860 | 35+151 | 0.221 |
| Methods | MagicBrush [267] | AnyEdit [268] | Params | Time | ||||
|---|---|---|---|---|---|---|---|---|
| L1↓ | DINO↑ | CVS↑ | L1↓ | DINO↑ | CVS↑ | [B] | [s] | |
| ML-MGIE [76] | 0.083 | 84.54 | 90.70 | 0.211 | 60.31 | 78.85 | 0.86+7.43 | 5.11 |
| InstructPix2Pix [269] | 0.101 | 71.46 | 85.22 | 0.167 | 65.93 | 77.90 | 0.86+0.12 | 7.76 |
| ICEdit [270] | 0.088 | 84.37 | 90.17 | 0.163 | 66.66 | 80.90 | 12.00+4.82 | 9.95 |
| UltraEdit-SD3 [271] | 0.042 | 87.90 | 89.70 | 0.152 | 70.50 | 80.43 | 2.00+5.52 | 3.09 |
| AnySD [268] | 0.066 | 88.10 | 87.60 | 0.148 | 72.45 | 81.68 | 0.86+0.12 | 3.16 |
| EditMGT [272] | 0.109 | 76.96 | 86.22 | 0.147 | 68.80 | 82.74 | 1.01+3.70 | 12.75 |

6. Experiments
6.1. Datasets
6.2. Evaluation metrics


6.3. Experimental Results
7. Future Directions
7.1. Advancing VLM Architectures for Low-Level Vision
7.2. Expanding VLM Capabilities and Applications
7.3. Optimizing VLM Performance and Deployment
7.4. Addressing Inherent Challenges and New Paradigms
8. Conclusions
Funding
Conflicts of Interest
References
- Yang, W.; Zhou, F.; Zhu, R.; Fukui, K.; Wang, G.; Xue, J.H. Deep learning for image super-resolution. Neurocomputing 2020, 398, 291–292. [Google Scholar] [CrossRef]
- Saharia, C.; Ho, J. Image super-resolution via iterative refinement. IEEE Trans. Pattern Anal. Mach. Intell. 2022, 713–726. [Google Scholar]
- Biyouki, S.A.; Hwangbo, H. A comprehensive survey on deep neural image deblurring. arXiv 2023, arXiv:2310.04719. [Google Scholar]
- He, C.; Zhang, R.; Chen, Z.; Yang, B.; Fang, C.; Lin, Y.; Xiao, F.; Farsiu, S. UnfoldLDM: Deep Unfolding-based Blind Image Restoration with Latent Diffusion Priors. arXiv 2025, arXiv:2511.18152. [Google Scholar]
- Liu, Y.; Zhao, G.; Gong, B.; Li, Y.; Raj, R.; Goel, N.; Kesav, S.; Gottimukkala, S.; Wang, Z.; Ren, W.; et al. Improved techniques for learning to dehaze and beyond: A collective study. arXiv 2018, arXiv:1807.00202. [Google Scholar]
- Fang, C.; He, C.; Xiao, F.; Zhang, Y.; Tang, L.; Zhang, Y.; Li, K.; Li, X. Real-world Image Dehazing with Coherence-based Label Generator and Cooperative Unfolding Network. NeurIPS 2024. [Google Scholar]
- Xia, B.; Zhang, Y.; Wang, S.; Wang, Y.; Wu, X.; Tian, Y.; Yang, W.; Van Gool, L. DiffIR: Efficient diffusion model for image restoration. In Proceedings of the ICCV, 2023; pp. 13095–13105. [Google Scholar]
- He, C.; Li, K.; Xu, G.; Zhang, Y.; Hu, R.; Guo, Z.; Li, X. Degradation-resistant unfolding network for heterogeneous image fusion. In Proceedings of the ICCV, 2023; pp. 12611–12621. [Google Scholar]
- Xu, G.; He, C.; Wang, H.; Zhu, H.; Ding, W. DM-Fusion: Deep Model-Driven Network for Heterogeneous Image Fusion. IEEE Trans. Neural Netw. Learn. Syst. 2023. [Google Scholar]
- He, C.; Fang, C.; Zhang, Y.; Ye, T.; Li, K.; Tang, L.; Guo, Z.; Li, X.; Farsiu, S. Reti-Diff: Illumination degradation image restoration with Retinex-based latent diffusion model. ICLR, 2025. [Google Scholar]
- He, C.; Zhang, R.; Xiao, F.; Fang, C.; Tang, L.; Zhang, Y.; Farsiu, S. UnfoldIR: Rethinking Deep Unfolding Network in Illumination Degradation Image Restoration. CVPR, 2026. [Google Scholar]
- Kaur, A.; et al. A complete review on image denoising techniques for medical images. NPL 2023, 55, 7807–7850. [Google Scholar] [CrossRef]
- Han, L.; Zhao, Y.; Lv, H.; Zhang, Y.; Liu, H.; Bi, G. Remote sensing image denoising based on deep and shallow feature fusion and attention mechanism. Remote Sens. 2022, 14, 1243. [Google Scholar] [CrossRef]
- He, C.; Li, K.; Zhang, Y.; Zhang, Y.; Guo, Z.; Li, X. Strategic Preys Make Acute Predators: Enhancing Camouflaged Object Detectors by Generating Camouflaged Objects. ICLR, 2024. [Google Scholar]
- Dong, C.; Loy, C.C.; He, K.; Tang, X. Image super-resolution using deep convolutional networks. IEEE Trans. Pattern Anal. Mach. Intell. 2015, 38, 295–307. [Google Scholar] [CrossRef]
- Zhang, K.; Zuo, W.; Chen, Y.; Meng, D.; Zhang, L. Beyond a Gaussian denoiser: Residual learning of deep CNN for image denoising. IEEE Trans. Image Process. 2017, 26, 3142–3155. [Google Scholar] [CrossRef] [PubMed]
- Zamir, S.W.; Arora, A.; Khan, S.; Hayat, M.; Khan, F.S.; Yang, M.H. Restormer: Efficient transformer for high-resolution image restoration. In Proceedings of the CVPR, 2022; pp. 5728–5739. [Google Scholar]
- Zhu, K.; Gu, J.; You, Z.; Qiao, Y.; Dong, C. An Intelligent Agentic System for Complex Image Restoration Problems. arXiv 2024, arXiv:2410.17809. [Google Scholar]
- Zhou, Y.; Cao, J.; Zhang, Z.; Wen, F.; Jiang, Y.; Jia, J.; Liu, X.; Min, X.; Zhai, G. Q-Agent: Quality-Driven Chain-of-Thought Image Restoration Agent through Robust Multimodal Large Language Model. arXiv 2025, arXiv:2504.07148. [Google Scholar]
- Radford, A.; Kim, J.W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; et al. Learning transferable visual models from natural language supervision. In Proceedings of the ICML, 2021; pp. 8748–8763. [Google Scholar]
- Li, J.; Li, D.; Xiong, C.; Hoi, S. BLIP: Bootstrapping language-image pre-training for unified vision-language understanding and generation. In Proceedings of the ICML, 2022; pp. 12888–12900. [Google Scholar]
- Liu, H.; Li, C.; Li, Y.; Lee, Y.J. Improved baselines with visual instruction tuning. In Proceedings of the CVPR, 2024; pp. 26296–26306. [Google Scholar]
- Jia, C.; Yang, Y.; Xia, Y.; Chen, Y.T.; Parekh, Z.; Pham, H.; Le, Q.; Sung, Y.H.; Li, Z.; Duerig, T. Scaling up visual and vision-language representation learning with noisy text supervision. In Proceedings of the ICML. PMLR, 2021; pp. 4904–4916. [Google Scholar]
- Yi, X.; Xu, H.; Zhang, H.; Tang, L.; Ma, J. Text-IF: Leveraging semantic text guidance for degradation-aware and interactive image fusion. In Proceedings of the CVPR, 2024; pp. 27026–27035. [Google Scholar]
- Yu, Y.; Du, D.; Zhang, L.; Luo, T. Unbiased multi-modality guidance for image inpainting. In Proceedings of the ECCV, 2022; pp. 668–684. [Google Scholar]
- Zhang, X.; Zhang, H.; Wang, G.; Zhang, Q.; Zhang, L. ClearAIR: A Human-Visual-Perception-Inspired All-in-One Image Restoration. In Proceedings of the AAAI, 2026. [Google Scholar]
- Xu, J.; Wu, M.; Hu, X.; Fu, C.W.; Dou, Q.; Heng, P.A. Towards Real-World Adverse Weather Image Restoration: Enhancing Clearness and Semantics with Vision-Language Models. In Proceedings of the ECCV, 2024; pp. 147–164. [Google Scholar]
- Morawski, I.; He, K.; Dangi, S.; Hsu, W.H. Leveraging Content and Context Cues for Low-Light Image Enhancement. IEEE Trans. Multimedia 2025. [Google Scholar]
- Luo, Z.; Gustafsson, F.K.; Zhao, Z.; Sjölund, J.; Schön, T.B. Controlling vision-language models for multi-task image restoration. arXiv 2023, arXiv:2310.01018. [Google Scholar]
- Lei, X.; Zhang, W.; Luo, B.; Liang, H.; Cao, W.; Lin, Q. DACESR: Degradation-Aware Conditional Embedding for Real-World Image Super-Resolution. arXiv 2026, arXiv:2602.23890. [Google Scholar]
- Jiang, X.; Li, G.; Chen, B.; Zhang, J. Multi-Agent Image Restoration. arXiv 2025, arXiv:2503.09403. [Google Scholar]
- Li, B.; Li, X.; Lu, Y.; Chen, Z. Hybrid agents for image restoration. arXiv 2025, arXiv:2503.10120. [Google Scholar]
- Tang, J.; Xiao, H.; Li, X.; Wang, W.; Gong, Z. ChatCAD: An MLLM-Guided Framework for Zero-shot CAD Drawing Restoration. In Proceedings of the ICASSP, 2025; pp. 1–5. [Google Scholar]
- Wu, H.; Zhang, Z.; Zhang, E.; Chen, C.; Liao, L.; Wang, A.; Xu, K.; Li, C.; Hou, J.; Zhai, G.; et al. Q-Instruct: Improving low-level visual abilities for multi-modality foundation models. In Proceedings of the CVPR, 2024; pp. 25490–25500. [Google Scholar]
- Ai, Y.; Huang, H.; He, R. LoRA-IR: Taming Low-Rank Experts for Efficient All-in-One Image Restoration. arXiv 2024, arXiv:2410.15385. [Google Scholar]
- Chen, B.; Chen, K.; Liu, L.; Shi, Z.; Zou, Z. Leveraging language-aligned visual knowledge for remote sensing image spectral super-resolution. In Proceedings of the WHISPERS, 2024; pp. 1–5. [Google Scholar]
- Wu, Z.; Chen, Y.; Yokoya, N.; He, W. MP-HSIR: A Multi-Prompt Framework for Universal Hyperspectral Image Restoration. arXiv 2025, arXiv:2503.09131. [Google Scholar]
- Lee, C.M.; Cheng, C.H.; Lin, Y.F.; Cheng, Y.C.; Liao, W.T.; Hsu, C.C.; Yang, F.E.; Wang, Y.C.F. PromptHSI: Universal hyperspectral image restoration framework for composite degradation. arXiv 2024, arXiv:2411.15922. [Google Scholar]
- Eslami, S.; de Melo, G.; Meinel, C. Does CLIP benefit visual question answering in the medical domain as much as it does in the general domain? arXiv 2021, arXiv:2112.13906. [Google Scholar]
- Eslami, S.; Meinel, C.; De Melo, G. PubmedCLIP: How much does CLIP benefit visual question answering in the medical domain? In Proceedings of the EACL, 2023; pp. 1181–1193. [Google Scholar]
- Boecking, B.; Usuyama, N.; Bannur, S.; Castro, D.C.; Schwaighofer, A.; Hyland, S.; Wetscherek, M.; Naumann, T.; Nori, A.; Alvarez-Valle, J.; et al. Making the most of text semantics to improve biomedical vision–language processing. In Proceedings of the ECCV, 2022; pp. 1–21. [Google Scholar]
- Zhang, Z.; Shen, H.; Zhao, T.; Guan, Z.; Chen, B.; Wang, Y.; Jia, X.; Cai, Y.; Shang, Y.; Yin, J. ImageRAG: Enhancing Ultra High Resolution Remote Sensing Imagery Analysis with ImageRAG. arXiv 2024, arXiv:2411.07688. [Google Scholar]
- Ranzinger, M.; Heinrich, G.; Molchanov, P.; Kautz, J.; Catanzaro, B.; Tao, A. FeatSharp: Your Vision Model Features, Sharper. arXiv 2025, arXiv:2502.16025. [Google Scholar]
- Luo, Z.; Gustafsson, F.K.; Zhao, Z.; Sjölund, J.; Schön, T.B. Photo-realistic image restoration in the wild with controlled vision-language models. In Proceedings of the CVPR, 2024; pp. 6641–6651. [Google Scholar]
- Shao, M.; Liu, W.; Li, Q.; Meng, L. Controlling vision-language model for enhancing image restoration. IVC 2025, 105538. [Google Scholar] [CrossRef]
- Yu, L.; Lezama, J.; Gundavarapu, N.B.; Versari, L.; Sohn, K.; Minnen, D.; Cheng, Y.; Birodkar, V.; Gupta, A.; Gu, X.; et al. Language Model Beats Diffusion–Tokenizer is Key to Visual Generation. arXiv 2023, arXiv:2310.05737. [Google Scholar]
- Yang, Q.; Li, P.; Jin, J.; Jin, G.; Song, T.; Fan, S.; Hou, H. Textual-visual interaction for enhanced single image deraining using adapter-tuned VLMs. Vis. Comput. 2026, 42, 193. [Google Scholar] [CrossRef]
- Wasserman, N.; Rotstein, N.; Ganz, R.; Kimmel, R. Paint by inpaint: Learning to add image objects by removing them first. arXiv 2024, arXiv:2404.18212. [Google Scholar]
- Wang, Z.; Zhao, L.; Zhang, J.; Song, R.; Song, H.; Meng, J.; Wang, S. Multi-Text Guidance Is Important: Multi-Modality Image Fusion via Large Generative Vision-Language Model. Int. J. Comput. Vis. 2025, 1–23. [Google Scholar] [CrossRef]
- Chen, X.; Bai, C.; Wu, Z.; Wu, X.; Zou, Q.; Xia, Y.; Wang, S. Coarse-to-fine text injecting for realistic image super-resolution. Neurocomputing 2025, 129591. [Google Scholar] [CrossRef]
- Yang, H.; Pan, L.; Yang, Y.; Liang, W. Language-driven all-in-one adverse weather removal. In Proceedings of the CVPR, 2024; pp. 24902–24912. [Google Scholar]
- Xu, X.; Kong, S.; Hu, T.; Liu, Z.; Bao, H. Boosting image restoration via priors from pre-trained models. In Proceedings of the CVPR, 2024; pp. 2900–2909. [Google Scholar]
- Zeng, H.; Wang, X.; Chen, Y.; Su, J.; Liu, J. Vision-Language Gradient Descent-driven All-in-One Deep Unfolding Networks. In Proceedings of the CVPR, 2025; pp. 7524–7533. [Google Scholar]
- Wang, J.; Fan, Q.; Chen, J.; Gu, H.; Huang, F.; Ren, W. RAP-SR: RestorAtion Prior Enhancement in Diffusion Models for Realistic Image Super-Resolution. In Proceedings of the AAAI; 2025; Vol. 39, pp. 7727–7735. [Google Scholar] [CrossRef]
- Wolters, P.; Bastani, F.; Kembhavi, A. Zooming out on zooming in: Advancing super-resolution for remote sensing. arXiv 2023, arXiv:2311.18082. [Google Scholar]
- Yin, J.; He, Y.; Zhang, M.; Zeng, P.; Wang, T.; Lu, S. PromptLnet: Region-adaptive aesthetic enhancement via prompt guidance in low-light enhancement net. arXiv 2025, arXiv:2503.08276. [Google Scholar]
- Yin, S.; Fu, C.; Zhao, S.; Li, K.; Sun, X. A survey on multimodal large language models. NSR 2024, 11, nwae403. [Google Scholar] [CrossRef] [PubMed]
- Zhang, J.; Huang, J.; Jin, S.; Lu, S. Vision-language models for vision tasks: A survey. IEEE Trans. Pattern Anal. Mach. Intell. 2024. [Google Scholar]
- Ghosh, A.; Acharya, A.; Saha, S.; Jain, V.; Chadha, A. Exploring the frontier of vision-language models: A survey of current methodologies and future directions. arXiv 2024, arXiv:2404.07214. [Google Scholar]
- Danish, S.; Sadeghi-Niaraki, A.; Khan, S.U.; Dang, L.M.; Tightiz, L.; Moon, H. A comprehensive survey of vision-language models: Pretrained models, fine-tuning, prompt engineering, adapters, and benchmark datasets. Inf. Fusion 2025, 103623. [Google Scholar]
- Jiang, J.; Zuo, Z.; Wu, G.; Jiang, K.; Liu, X. A survey on all-in-one image restoration: Taxonomy, evaluation and future trends. IEEE Trans. Pattern Anal. Mach. Intell. 2025. [Google Scholar]
- Liu, M.; Shu, H.; Cui, Y.; Zhou, X.; Cao, H.; Ren, W.; Shi, B.; Knoll, A.C. Language-Driven Image Restoration and Semantic-Aware Quality Assessment: A Survey. 2026. [Google Scholar] [CrossRef] [PubMed]
- Zhang, J.; Cheng, S.; Sun, Q.; Liu, J.; Luyang, W.; Feng, C.; Fang, C.; et al. Ultra High-Resolution Image Inpainting with Patch-Based Content Consistency Adapter. In Proceedings of the ICCV; 2025; pp. 16991–17000. [Google Scholar] [CrossRef]
- Liu, Z.; Zhu, L.; Shi, B.; Zhang, Z.; Lou, Y.; Yang, S.; Xi, H.; Cao, S.; Gu, Y.; Li, D.; et al. NVILA: Efficient frontier visual language models. In Proceedings of the CVPR, 2025; pp. 4122–4134. [Google Scholar]
- Wei, Y.; Zhang, Y.; Li, K.; Wang, F.; Tang, S.; Zhang, Z. Leveraging vision-language prompts for real-world image restoration and enhancement. Comput. Vis. Image Underst. 2025, 250, 104222. [Google Scholar]
- Xue, M.; He, J.; Wang, W.; Zhou, M. Low-light image enhancement via CLIP-Fourier guided wavelet diffusion. arXiv 2024, arXiv:2401.03788. [Google Scholar]
- Liang, Z.; Li, C.; Zhou, S.; Feng, R.; Loy, C.C. Iterative Prompt Learning for Unsupervised Backlit Image Enhancement. In Proceedings of the ICCV, 2023; pp. 8060–8069. [Google Scholar]
- Chang, Z.; Weng, S.; Zhang, P.; Li, Y.; Li, S.; Shi, B. L-CAD: Language-based Colorization with Any-level Descriptions using Diffusion Priors. In Proceedings of the NeurIPS; 2023; Vol. 36, pp. 77174–77186. [Google Scholar] [CrossRef]
- Morawski, I.; He, K.; Dangi, S.; Hsu, W.H. Unsupervised Image Prior via Prompt Learning and CLIP Semantic Guidance for Low-Light Image Enhancement. In Proceedings of the CVPR, 2024; pp. 5971–5981. [Google Scholar]
- Wen, Y.; Gao, T.; Li, Z.; Zhang, J.; Zhang, K.; Chen, T. All-in-one Weather-degraded Image Restoration via Adaptive Degradation-aware Self-prompting Model. IEEE Trans. Multimedia 2025. [Google Scholar]
- Shao, M.; Liu, Y.; Cheng, Y.; Wan, Y.; Wang, C. Adaptive Fuzzy Degradation Perception Based on CLIP Prior for All-in-one Image Restoration. IEEE Trans. Fuzzy Syst. 2024. [Google Scholar]
- Li, X.; Liu, J.; Chen, Z.; Zou, Y.; Ma, L.; Fan, X.; Liu, R. Contourlet residual for prompt learning enhanced infrared image super-resolution. In Proceedings of the ECCV; Springer, 2024; pp. 270–288. [Google Scholar]
- Zhang, X.; Ma, J.; Wang, G.; Zhang, Q.; Zhang, H.; Zhang, L. Perceive-IR: Learning to perceive degradation better for all-in-one image restoration. IEEE Trans. Image Process. 2025. [Google Scholar]
- Lin, Y.; Lin, Z.; Chen, H.; Pan, P.; Li, C.; Chen, S.; Wen, K.; Jin, Y.; Li, W.; Ding, X. JarvisIR: Elevating autonomous driving perception with intelligent image restoration. In Proceedings of the CVPR, 2025; pp. 22369–22380. [Google Scholar]
- Chen, H.; Li, W.; Gu, J.; Ren, J.; Chen, S.; Ye, T.; Pei, R.; Zhou, K.; et al. RestoreAgent: Autonomous image restoration agent via multimodal large language models. arXiv 2024, arXiv:2407.18035. [Google Scholar]
- Fu, T.J.; Hu, W.; Du, X.; Wang, W.Y.; Yang, Y.; Gan, Z. Guiding instruction-based image editing via multimodal large language models. arXiv 2023, arXiv:2309.17102. [Google Scholar]
- Duan, H.; Min, X.; Wu, S.; Shen, W.; Zhai, G. UniProcessor: a text-induced unified low-level image processor. In Proceedings of the ECCV. Springer, 2024; pp. 180–199. [Google Scholar]
- Wu, Y.; Zhang, Z.; Chen, J.; Tang, H.; Li, D.; Fang, Y.; Zhu, L.; Xie, E.; et al. VILA-U: a unified foundation model integrating visual understanding and generation. arXiv 2024, arXiv:2409.04429. [Google Scholar]
- Qu, L.; Zhang, H.; Liu, Y.; Wang, X.; Jiang, Y.; Gao, Y.; Ye, H.; Du, D.K.; et al. Tokenflow: Unified image tokenizer for multimodal understanding and generation. In Proceedings of the CVPR, 2025; pp. 2545–2555. [Google Scholar]
- Chen, Z.; Wang, C.; Chen, X.; Xu, H.; Huang, R.; Zhou, J.; Han, J.; Xu, H.; Liang, X. SemHiTok: A unified image tokenizer via semantic-guided hierarchical codebook for multimodal understanding and generation. arXiv 2025, arXiv:2503.06764. [Google Scholar]
- Zhu, L.; Wei, F.; Lu, Y. Beyond Text: Frozen Large Language Models in Visual Signal Comprehension. In Proceedings of the CVPR, 2024; pp. 27047–27057. [Google Scholar]
- Eteke, C.; Griessel, A.; Kellerer, W.; Steinbach, E. BIR-Adapter: A Low-Complexity Diffusion Model Adapter for Blind Image Restoration. arXiv 2025, arXiv:2509.06904. [Google Scholar]
- Hu, Q.; Fan, L.; Luo, Y.; Yu, Y.; Guo, X.; Fan, Q. Text-Aware Real-World Image Super-Resolution via Diffusion Model with Joint Segmentation Decoders. arXiv 2025, arXiv:2506.04641. [Google Scholar]
- Fang, Y.; Chen, Y.; Yin, S.; Hu, Q.; Yao, J.; Zhang, Y.; Zhang, X.; Wang, Y. One-Step Diffusion Transformer for Controllable Real-World Image Super-Resolution. In Proceedings of the CVPR, 2026. [Google Scholar]
- Zheng, B.; Gu, J.; Li, S. LM4LV: A frozen large language model for low-level vision tasks. arXiv 2024, arXiv:2405.15734. [Google Scholar]
- Pang, Z.; Xie, Z.; Man, Y.; Wang, Y.X. Frozen transformers in language models are effective visual encoder layers. arXiv 2023, arXiv:2310.12973. [Google Scholar]
- Chen, X.; Liu, Y.; Pu, Y.; Zhang, W.; Zhou, J.; Qiao, Y.; Dong, C. Learning a low-level vision generalist via visual task prompt. In Proceedings of the ACM MM, 2024; pp. 2671–2680. [Google Scholar]
- Zeng, Z.; Hua, H.; Fu, J.; Luo, J.; et al. PromptFix: You Prompt and We Fix the Photo. NeurIPS 2024, 37, 40000–40031. [Google Scholar] [CrossRef]
- Qu, Y.; Yuan, K.; Zhao, K.; Xie, Q.; Hao, J.; Sun, M.; Zhou, C. XPSR: Cross-modal priors for diffusion-based image super-resolution. In Proceedings of the ECCV. Springer, 2024; pp. 285–303. [Google Scholar]
- Zhang, Y.; Zhang, H.; Chai, X.; Cheng, Z.; Xie, R.; Song, L.; Zhang, W. Diff-Restorer: Unleashing visual prompts for diffusion-based universal image restoration. arXiv 2024, arXiv:2407.03636. [Google Scholar]
- Kong, D.; Li, F.; Wang, Z.; Xu, J.; Pei, R.; Li, W.; Ren, W. Dual Prompting Image Restoration with Diffusion Transformers. In Proceedings of the CVPR, 2025. [Google Scholar]
- Jiang, Y.; Zhang, Z.; Xue, T.; Gu, J. AutoDIR: Automatic all-in-one image restoration with latent diffusion. In Proceedings of the ECCV, 2024; pp. 340–359. [Google Scholar]
- Yang, T.; Wu, R.; Ren, P.; Xie, X.; Zhang, L. Pixel-aware stable diffusion for realistic image super-resolution and personalized stylization. In Proceedings of the ECCV, 2024; pp. 74–91. [Google Scholar]
- Li, X.; Wu, J.; Huang, X.; Chen, C.; Guan, W.; Hua, X.S.; Nie, L. MegaSR: Mining Customized Semantics and Expressive Guidance for Image Super-Resolution. arXiv 2025, arXiv:2503.08096. [Google Scholar]
- Zhao, Z.; Deng, L.; Bai, H.; Cui, Y.; Zhang, Z.; Zhang, Y.; Qin, H.; Chen, D.; Zhang, J.; Wang, P.; et al. Image fusion via vision-language model. arXiv 2024, arXiv:2402.02235. [Google Scholar]
- Fei, S.; Ye, T.; Wang, L.; Zhu, L. LucidFlux: Caption-Free Universal Image Restoration via a Large-Scale Diffusion Transformer. In Proceedings of the ICLR, 2026. [Google Scholar]
- Jiang, L.; Liu, X.; Tong, X.; Li, Z.; Liu, J.; Tang, J.; Wu, G. Disentangled Textual Priors for Diffusion-based Image Super-Resolution. In Proceedings of the CVPR, 2026. [Google Scholar]
- Sung, M.; Ham, S.; Kim, K.; Yoon, Y.; Yun, S.; Kim, I.M.; Kang, J.M. GLYPH-SR: Can We Achieve Both High-Quality Image Super-Resolution and High-Fidelity Text Recovery via VLM-guided Latent Diffusion Model? arXiv 2025, arXiv:2510.26339. [Google Scholar]
- Conde, M.V.; Geigle, G.; Timofte, R. InstructIR: High-quality image restoration following human instructions. In Proceedings of the ECCV, 2024; pp. 1–21. [Google Scholar]
- Qi, C.; Tu, Z.; Ye, K.; Delbracio, M.; Milanfar, P.; Chen, Q.; Talebi, H. SPIRE: Semantic prompt-driven image restoration. In Proceedings of the ECCV, 2024; pp. 446–464. [Google Scholar]
- Chen, Z.; Zhang, Y.; Gu, J.; Yuan, X.; Kong, L.; Chen, G.; Yang, X. Image super-resolution with text prompt diffusion. arXiv 2023, arXiv:2311.14282. [Google Scholar]
- Liao, R.; Li, F.; Wei, Y.; Shi, Z.; Zhang, L.; Bai, H.; Wang, M. Prompt to Restore, Restore to Prompt: Cyclic Prompting for Universal Adverse Weather Removal. arXiv 2025, arXiv:2503.09013. [Google Scholar]
- Zhang, C.; Yang, W.; Li, X.; Han, H. MMGInpainting: Multi-modality guided image inpainting based on diffusion models. IEEE Trans. Multimedia 2024. [Google Scholar]
- Li, Y.; Bian, Y.; Ju, X.; Zhang, Z.; Shan, Y.; Zou, Y.; Xu, Q. BrushEdit: All-in-one image inpainting and editing. arXiv 2024, arXiv:2412.10316. [Google Scholar]
- Chiu, M.T.; Zhou, Y.; Zhang, L.; Lin, Z.; Barnes, C.; Amirghodsi, S.; Shechtman, E.; Shi, H. Brush2Prompt: Contextual Prompt Generator for Object Inpainting. In Proceedings of the CVPR, 2024; pp. 12636–12645. [Google Scholar]
- Xie, H.; Du, K.; Yan, Q.; Lu, S.; Han, J.; Chen, H.; Hu, H.; Hu, J. EAM: Enhancing Anything with Diffusion Transformers for Blind Super-Resolution. arXiv 2025, arXiv:2505.05209. [Google Scholar]
- Wang, R.; Li, W.; Liu, X.; Li, C.; Zhang, Z.; Min, X.; Zhai, G. HazeCLIP: Towards language guided real-world image dehazing. In Proceedings of the ICASSP. IEEE, 2025; pp. 1–5. [Google Scholar]
- Zuo, Y.; Zheng, Q.; Wu, M.; Jiang, X.; Li, R.; Wang, J.; Zhang, Y.; Mai, G.; Wang, L.V.; Zou, J.; et al. 4KAgent: Agentic Any Image to 4K Super-Resolution. arXiv 2025, arXiv:2507.07105. [Google Scholar]
- Yu, F.; Gu, J.; Li, Z.; Hu, J.; Kong, X.; Wang, X.; He, J.; et al. Scaling up to excellence: Practicing model scaling for photo-realistic image restoration in the wild. In Proceedings of the CVPR, 2024; pp. 25669–25680. [Google Scholar]
- Xie, T.; Ma, R.; Wang, Q.; Ye, X.; Liu, F.; Tai, Y.; Zhang, Z.; Wang, L.; Yi, Z. Anywhere: A Multi-Agent Framework for User-Guided, Reliable, and Diverse Foreground-Conditioned Image Generation. In Proceedings of the AAAI; 2025; Vol. 39, pp. 7410–7418. [Google Scholar] [CrossRef]
- Zhang, R.; Yang, Z.; Pan, L. DehazeMamba: Large multi-modal model guided single image dehazing via mamba. Vis. Intell. 2025, 3, 11. [Google Scholar]
- Tan, Z.; Wu, Y.; Liu, Q.; Chu, Q.; Lu, L.; Ye, J.; Yu, N. Exploring the application of large-scale pre-trained models on adverse weather removal. IEEE Trans. Image Process. 2024. [Google Scholar]
- Wang, Y.; Li, Y.; Zheng, Z.; Zhang, X.P.; Wei, M. M2Restore: Mixture-of-Experts-based Mamba-CNN Fusion Framework for All-in-One Image Restoration. arXiv 2025, arXiv:2506.07814. [Google Scholar]
- Zhang, R.; Yang, H.; Yang, Y.; Fu, Y.; Pan, L. LMHaze: Intensity-aware Image Dehazing with a Large-scale Multi-intensity Real Haze Dataset. In Proceedings of the MM Asia, 2024; pp. 1–1. [Google Scholar]
- Huang, X.; Zhang, Q.; Hu, J.F.; Zheng, W.S. CLIP-RestoreX: Restore Image Structure and Perception in Exposure Correction. In Proceedings of the AAAI; 2025; Vol. 39, pp. 3760–3768. [Google Scholar] [CrossRef]
- Cai, J.; Yang, K.; Ding, J.; Fu, L.; Ouyang, L.; Li, J.; Shen, J.; Meng, Z. Degradation-Aware Image Enhancement via Vision-Language Classification. arXiv 2025, arXiv:2506.05450. [Google Scholar]
- Lan, Y.; Cui, Z.; Luo, X.; Liu, C.; Wang, N.; Zhang, M.; Su, Y.; Liu, D. When Schrödinger Bridge Meets Real-World Image Dehazing with Unpaired Training. In Proceedings of the ICCV, 2025; pp. 8756–8765. [Google Scholar]
- Kim, J.H.; Cho, P.H.; Kim, C.; Min, J.; Lee, J.; Park, J.; Choi, Y.; Kim, S. UniT: Unified Diffusion Transformer for High-fidelity Text-Aware Image Restoration. In Proceedings of the ICLR, 2026. [Google Scholar]
- Zhou, H.; Dong, W.; Liu, X.; Zhang, Y.; Zhai, G.; Chen, J. Low-light image enhancement via generative perceptual priors. In Proceedings of the AAAI; 2025; Vol. 39, pp. 10752–10760. [Google Scholar] [CrossRef]
- Zhang, Y.; Li, H.; Zhang, S.; Wang, R.; He, B.; Dou, H.; Yan, J.; Zhang, Y.; Wu, F. LLMCO4MR: LLMs-Aided Neural Combinatorial Optimization for Ancient Manuscript Restoration from Fragments with Case Studies on Dunhuang. In Proceedings of the ECCV. Springer; 2024; pp. 253–269. [Google Scholar]
- Liu, Y.; Chen, X.; Ma, X.; Wang, X.; Zhou, J.; Qiao, Y.; Dong, C. Unifying Image Processing as Visual Prompting Question Answering. In Proceedings of the ICML, 2024; pp. 30873–30891. [Google Scholar]
- Zeng, Y.; Fu, J.; Amirpour, H.; Wang, H.; Yue, G.; Liu, H.; Chen, Y.; Zhou, W. CLIP-DQA: Blindly Evaluating Dehazed Images from Global and Local Perspectives Using CLIP. arXiv 2025, arXiv:2502.01707. [Google Scholar]
- Cheng, C.; Xu, T.; Wu, X.J.; Zhou, T.; Li, H.; Tang, Z.; Kittler, J. EvaNet: Towards More Efficient and Consistent Infrared and Visible Image Fusion Assessment. IEEE Trans. Pattern Anal. Mach. Intell. 2026. [Google Scholar]
- Ma, Y.; Xia, F.; Lin, L.; Guan, X. LEGO: LLM-enhanced genetic optimization for underwater robot image restoration. Pattern Recognit. 2025, 111782. [Google Scholar]
- Cai, Z.; Zhang, J.; Yuan, X.; Jiang, P.T.; Chen, W.; Tang, B.; Yao, L.; Wang, Q.; Chen, J.; Li, B. Q-Ponder: A Unified Training Pipeline for Reasoning-based Visual Quality Assessment. arXiv 2025, arXiv:2506.05384. [Google Scholar]
- Lai, J.; Chen, S.; Lin, Y.; Ye, T.; Liu, Y.; et al. SnowMaster: Comprehensive Real-world Image Desnowing via MLLM with Multi-Model Feedback Optimization. In Proceedings of the CVPR, 2025; pp. 4302–4312. [Google Scholar]
- Ai, Y.; Zhou, X.; Huang, H.; Han, X.; Chen, Z.; You, Q.; Yang, H. DreamClear: High-Capacity Real-World Image Restoration with Privacy-Safe Dataset Curation. NeurIPS 2024, 37, 55443–55469. [Google Scholar] [CrossRef]
- Deng, J.; Wu, X.; Yang, Y.; Zhu, C.; Wang, S.; Wu, Z. Acquire and then Adapt: Squeezing out Text-to-Image Model for Image Restoration. arXiv 2025, arXiv:2504.15159. [Google Scholar]
- Wang, J.; Chan, K.C.; Loy, C.C. Exploring CLIP for assessing the look and feel of images. In Proceedings of the AAAI; 2023; Vol. 37, pp. 2555–2563. [Google Scholar] [CrossRef]
- Yang, S.; Wu, T.; Shi, S.; Lao, S.; Gong, Y.; Cao, M.; Wang, J.; Yang, Y. MANIQA: Multi-dimension attention network for no-reference image quality assessment. In Proceedings of the CVPR, 2022; pp. 1191–1200. [Google Scholar]
- Li, Z.; Jin, J.; Cai, S.; Lin, W. R4-CGQA: Retrieval-based Vision Language Models for Computer Graphics Image Quality Assessment. arXiv 2026, arXiv:2603.10578. [Google Scholar]
- Zhang, W.; Zhai, G.; Wei, Y.; Yang, X.; Ma, K. Blind Image Quality Assessment via Vision-Language Correspondence: A Multitask Learning Perspective. In Proceedings of the CVPR, 2023; pp. 14071–14081. [Google Scholar]
- Wen, W.; Zhi, T.; Fan, K.; Li, Y.; Peng, X.; Zhang, Y.; Liao, Y.; Li, J.; Zhang, L. Self-Evolving Vision-Language Models for Image Quality Assessment via Voting and Ranking. In Proceedings of the ICLR, 2026. [Google Scholar]
- Zhu, H.; Tian, Y.; Ding, K.; Chen, B.; Chen, B.; Wang, S.; Lin, W. AgenticIQA: An Agentic Framework for Adaptive and Interpretable Image Quality Assessment. arXiv 2025, arXiv:2509.26006. [Google Scholar]
- Lu, Z.; Xia, Q.; Wang, W.; Wang, F. CLIP-aware domain-adaptive super-resolution. arXiv 2025, arXiv:2505.12391. [Google Scholar]
- Gaintseva, T.; Benning, M.; Slabaugh, G. RAVE: Residual vector embedding for CLIP-guided backlit image enhancement. In Proceedings of the ECCV, 2024; pp. 412–428. [Google Scholar]
- Ogino, Y.; Toizumi, T.; Ito, A. CURVE: Clip-Utilized Reinforcement Learning for Visual Image Enhancement via Simple Image Processing. In Proceedings of the ICIP. IEEE, 2025; pp. 427–432. [Google Scholar]
- Qiao, J.; Cai, M.; Li, W.; Liu, Y.; Huang, X.; He, G.; Xie, J.; Hu, J.; Chen, X.; Lin, S. RealSR-R1: Reinforcement Learning for Real-World Image Super-Resolution with Vision-Language Chain-of-Thought. arXiv 2025, arXiv:2506.16796. [Google Scholar]
- Wu, B.; Liu, Y.; Zhang, C.; Zhao, Y.; Wang, W. LRPO: Enhancing Blind Face Restoration through Online Reinforcement Learning. arXiv 2025, arXiv:2509.23339. [Google Scholar]
- Xu, X.; Chu, R.; Wang, J.; Zhou, K.; Shu, W.; Yang, H.; Lim, S.N.; Chen, H.; Lin, L. Enhancing Diffusion-based Restoration Models via Difficulty-Adaptive Reinforcement Learning with IQA Reward. arXiv 2025, arXiv:2511.01645. [Google Scholar]
- Cho, U.; Kim, N. A-IDE: Agent-Integrated Denoising Experts. arXiv 2025, arXiv:2503.16780. [Google Scholar]
- Cui, X.; Li, Z.; Li, P.; Hu, Y.; Shi, H.; Cao, C.; He, Z. ChatEdit: Towards multi-turn interactive facial image editing via dialogue. In Proceedings of the EMNLP, 2023; pp. 14567–14583. [Google Scholar]
- Liu, Z.; Yu, Y.; Ouyang, H.; Wang, Q.; Cheng, K.L.; Wang, W.; Liu, Z.; Chen, Q.; Shen, Y. MagicQuill: An intelligent interactive image editing system. In Proceedings of the CVPR, 2025; pp. 13072–13082. [Google Scholar]
- Thawkar, O.; Shaker, A.; Mullappilly, S.S.; Cholakkal, H.; Anwer, R.M.; Khan, S.; Laaksonen, J.; Khan, F.S. XrayGPT: Chest radiographs summarization using medical vision-language models. arXiv 2023, arXiv:2306.07971. [Google Scholar]
- Nath, V.; Li, W.; Yang, D.; Myronenko, A.; Zheng, M.; Lu, Y.; Liu, Z.; Yin, H.; Tang, Y.; Guo, P.; et al. VILA-M3: Enhancing vision-language models with medical expert knowledge. arXiv 2024, arXiv:2411.12915. [Google Scholar]
- Alkhaldi, A.; Alnajim, R.; Alabdullatef, L.; Alyahya, R.; Chen, J.; Zhu, D.; Alsinan, A.; Elhoseiny, M. MiniGPT-Med: Large language model as a general interface for radiology diagnosis. arXiv 2024, arXiv:2407.04106. [Google Scholar]
- Wu, C.; Zhang, X.; Zhang, Y.; Wang, Y.; Xie, W. MedKLIP: Medical knowledge enhanced language-image pre-training for X-Ray diagnosis. In Proceedings of the ICCV, 2023; pp. 21372–21383. [Google Scholar]
- Moor, M.; Huang, Q.; Wu, S.; Yasunaga, M.; Dalmia, Y.; Leskovec, J.; Zakka, C.; Reis, E.P.; Rajpurkar, P. Med-Flamingo: a multimodal medical few-shot learner. In Proceedings of the ML4H, 2023; pp. 353–367. [Google Scholar]
- Li, C.; Wong, C.; Zhang, S.; Usuyama, N.; Liu, H.; Yang, J.; Naumann, T.; Poon, H.; Gao, J. LLaVA-Med: Training a large language-and-vision assistant for biomedicine in one day. NeurIPS 2023, 36, 28541–28564. [Google Scholar]
- Bannur, S.; Hyland, S.; Liu, Q.; Perez-Garcia, F.; Ilse, M.; Castro, D.C.; Boecking, B.; Sharma, H.; Bouzid, K.; Thieme, A.; et al. Learning to exploit temporal structure for biomedical vision-language processing. In Proceedings of the CVPR, 2023; pp. 15016–15027. [Google Scholar]
- Chen, J.; Gui, C.; Ouyang, R.; Gao, A.; Chen, S.; Chen, G.H.; Wang, X.; Zhang, R.; Cai, Z.; Ji, K.; et al. HuatuoGPT-Vision, towards injecting medical visual knowledge into multimodal LLMs at scale. arXiv 2024, arXiv:2406.19280. [Google Scholar]
- Lin, T.; Zhang, W.; Li, S.; Yuan, Y.; Yu, B.; Li, H.; He, W.; Jiang, H.; Li, M.; Song, X.; et al. HealthGPT: A Medical Large Vision-Language Model for Unifying Comprehension and Generation via Heterogeneous Knowledge Adaptation. arXiv 2025, arXiv:2502.09838. [Google Scholar]
- Monajatipoor, M.; Rouhsedaghat, M.; Li, L.H.; Jay Kuo, C.C.; Chien, A.; Chang, K.W. BERTHop: An effective vision-and-language model for chest X-Ray disease diagnosis. In Proceedings of the MICCAI, 2022; pp. 725–734. [Google Scholar]
- Dong, A.; Xu, J.; Wang, L.; Lv, G.; Zhao, G.; Cheng, J. TDMF: Text-Guided Denoising and Interactive Medical Image Fusion. In Proceedings of the ICASSP. IEEE, 2025; pp. 1–5. [Google Scholar]
- Chen, Z.; Chen, T.; Wang, C.; Gao, Q.; Niu, C.; Wang, G.; Shan, H. Low-Dose CT denoising with language-engaged dual-space alignment. In Proceedings of the BIBM, 2024; pp. 3088–3091. [Google Scholar]
- Chen, Z.; Chen, T.; Wang, C.; Gao, Q.; Xie, H.; Niu, C.; Wang, G.; Shan, H. LangMamba: A Language-driven Mamba Framework for Low-Dose CT Denoising with Vision-language Models. IEEE Trans. Radiat. Plasma Med. Sci. 2025. [Google Scholar]
- Zhang, X.; Cai, A.; Wang, S.; Wang, L.; Zheng, Z.; Li, L.; Yan, B. Dual-Domain CLIP-Assisted Residual Optimization Perception Model for Metal Artifact Reduction. arXiv 2024, arXiv:2408.14342. [Google Scholar]
- Wu, B.; Hao, S.; Wang, W. Semantic-aware Guidance for Blind Super-resolution of Remote Sensing Images. IEEE Geosci. Remote Sens. Lett., 2025. [Google Scholar]
- Chen, B.; Chen, K.; Yang, M.; Zou, Z.; Shi, Z. SeG-SR: Integrating Semantic Knowledge into Remote Sensing Image Super-Resolution via Vision-Language Model. arXiv 2025, arXiv:2505.23010. [Google Scholar]
- Jian, L.; Liu, J.; Wu, S.; Chen, L. CLIPPan: Adapting CLIP as A Supervisor for Unsupervised Pansharpening. arXiv 2025, arXiv:2511.10896. [Google Scholar]
- Zhang, M.; Li, L.; Gao, F.; Zhang, Q.; Guo, J. Multimodal Prior Learning with Double Constraint Alignment for Snapshot Spectral Compressive Imaging. In Proceedings of the IJCAI; 2025; pp. 2359–2367. [Google Scholar] [CrossRef]
- Wang, X.; Zheng, J.; Hu, Y.; Zhu, H.; Yu, Q.; Zhou, Z. From 2D CAD Drawings to 3D Parametric Models: A Vision-Language Approach. arXiv 2024, arXiv:2412.11892. [Google Scholar]
- Mallis, D.; Karadeniz, A.S.; Cavada, S.; Rukhovich, D.; Foteinopoulou, N.; Cherenkova, K.; Kacem, A.; Aouada, D. CAD-Assistant: Tool-Augmented VLLMs as Generic CAD Task Solvers. arXiv 2025, arXiv:2412.13810. [Google Scholar]
- Xu, Y.; Song, Z.; Lu, J. Universal Video Face Restoration Method Based on Vision-Language Model. In Proceedings of the ACML, 2025. [Google Scholar]
- Liu, J.; Zhang, J.; Yang, S.; Xiang, J.; Wang, X.; Zhao, J.; Yang, Z.; Zhao, J. Towards General-Purpose Video Reconstruction through Synergy of Grid-Splicing Diffusion and Large Language Models. IEEE Trans. Circuits Syst. Video Technol. 2025. [Google Scholar]
- Ren, J.; Chen, H.; Ye, T.; Wu, H.; Zhu, L. Triplane-smoothed video dehazing with CLIP-enhanced generalization. Int. J. Comput. Vis. 2025, 133, 475–488. [Google Scholar]
- Liu, J.; Liu, Y.; Zhang, Y.; Meng, Z.; Tai, Y.W.; Tang, C.K. VP-LLM: Text-Driven 3D Volume Completion with Large Language Models through Patchification. arXiv 2024, arXiv:2406.05543. [Google Scholar]
- Wang, M.; Pi, H.; Li, R.; Qin, Y.; Tang, Z.; Li, K. VLScene: Vision-Language Guidance Distillation for Camera-Based 3D Semantic Scene Completion. arXiv 2025, arXiv:2503.06219. [Google Scholar]
- Hu, X.; Shi, C.; Yang, C.; Chen, M.; Ding, J.; Wei, T.; Wei, C.; Yu, Z.; Tan, M. SRSplat: Feed-Forward Super-Resolution Gaussian Splatting from Sparse Multi-View Images. arXiv 2025, arXiv:2511.12040. [Google Scholar]
- Wu, J.; Bian, J.W.; Li, X.; Wang, G.; Reid, I.; Torr, P.; Prisacariu, V.A. GaussCtrl: Multi-View Consistent Text-Driven 3D Gaussian Splatting Editing. arXiv 2024, arXiv:2403.08733. [Google Scholar]
- Haque, A.; Tancik, M.; Efros, A.A.; Holynski, A.; Kanazawa, A. Instruct-NeRF2NeRF: Editing 3D Scenes with Instructions. arXiv 2023, arXiv:2303.12789. [Google Scholar]
- Zhang, Z.; Qin, C.; Guo, C.; Zhang, Y.; Xue, C.; Cheng, M.M.; Li, C. RAM++: Robust Representation Learning via Adaptive Mask for All-in-One Image Restoration. arXiv 2025, arXiv:2509.12039. [Google Scholar]
- Yang, Y.; Zhang, C.; Yang, Z.; Gao, Y.; Qin, Y.; Li, K.; Sun, X.; Yang, J.; Gu, Y. RESTORE: Towards Feature Shift for Vision-Language Prompt Learning. arXiv 2024, arXiv:2403.06136. [Google Scholar]
- Simonyan, K.; Zisserman, A. Very deep convolutional networks for large-scale image recognition. arXiv 2014, arXiv:1409.1556. [Google Scholar]
- He, K.; Zhang, X.; Ren, S.; Sun, J. Deep residual learning for image recognition. In Proceedings of the CVPR, 2016; pp. 770–778. [Google Scholar]
- Tan, M.; Le, Q. EfficientNet: Rethinking model scaling for convolutional neural networks. In Proceedings of the ICML, 2019; pp. 6105–6114. [Google Scholar]
- Dosovitskiy, A. An image is worth 16x16 words: Transformers for image recognition at scale. arXiv 2020, arXiv:2010.11929. [Google Scholar]
- Yao, L.; Huang, R.; Hou, L.; Lu, G.; Niu, M.; Xu, H.; Liang, X.; Li, Z.; Jiang, X.; Xu, C. FILIP: Fine-grained interactive language-image pre-training. arXiv 2021, arXiv:2111.07783. [Google Scholar]
- Mu, N.; Kirillov, A.; Wagner, D.; Xie, S. SLIP: Self-supervision meets language-image pre-training. In Proceedings of the ECCV, 2022; pp. 529–544. [Google Scholar]
- Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A.N.; Kaiser, Ł.; Polosukhin, I. Attention is all you need. NeurIPS 2017, 30. [Google Scholar]
- Devlin, J.; Chang, M.W.; Lee, K.; Toutanova, K. BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the NAACL-HLT, 2019; pp. 4171–4186. [Google Scholar]
- Radford, A.; Wu, J.; Child, R.; Luan, D.; Amodei, D.; Sutskever, I.; et al. Language models are unsupervised multitask learners. OpenAI Blog 2019, 1, 9. [Google Scholar]
- Pi, R.; Gao, J.; Diao, S.; Pan, R.; Dong, H.; Zhang, J.; Yao, L.; Han, J.; Xu, H.; Kong, L.; et al. DetGPT: Detect what you need via reasoning. arXiv 2023, arXiv:2305.14167. [Google Scholar]
- Bai, J.; Bai, S.; Yang, S.; Wang, S.; Tan, S.; Wang, P.; Lin, J.; Zhou, C.; Zhou, J. Qwen-VL: A frontier large vision-language model with versatile abilities. arXiv 2023, arXiv:2308.129661, 3. [Google Scholar]
- Ye, Q.; Xu, H.; Xu, G.; Ye, J.; Yan, M.; Zhou, Y.; Wang, J.; Hu, A.; Shi, P.; Shi, Y.; et al. mPLUG-Owl: Modularization empowers large language models with multimodality. arXiv 2023, arXiv:2304.14178. [Google Scholar]
- Wang, W.; Chen, Z.; Chen, X.; Wu, J.; Zhu, X.; Zeng, G.; Luo, P.; Lu, T.; et al. VisionLLM: Large language model is also an open-ended decoder for vision-centric tasks. NeurIPS 2023, 36, 61501–61513. [Google Scholar] [CrossRef]
- Liu, H.; Li, C.; Wu, Q.; Lee, Y.J. Visual instruction tuning. NeurIPS 2023, 36, 34892–34916. [Google Scholar] [CrossRef]
- Zhang, X.; Wu, C.; Zhao, Z.; Lin, W.; Zhang, Y.; Wang, Y.; Xie, W. PMC-VQA: Visual instruction tuning for medical visual question answering. arXiv 2023, arXiv:2305.10415. [Google Scholar]
- Ziegler, D.M.; Stiennon, N.; Wu, J.; Brown, T.B.; Radford, A.; Amodei, D.; Christiano, P.; Irving, G. Fine-tuning language models from human preferences. arXiv 2019, arXiv:1909.08593. [Google Scholar]
- Sun, Z.; Shen, S.; Cao, S.; Liu, H.; Li, C.; Shen, Y.; Gan, C.; Gui, L.; Wang, Y.X.; Yang, Y.; et al. Aligning large multimodal models with factually augmented RLHF. In Proceedings of the ACL, 2024; pp. 13088–13110. [Google Scholar]
- Rafailov, R.; Sharma, A.; Mitchell, E.; Manning, C.D.; Ermon, S.; Finn, C. Direct preference optimization: Your language model is secretly a reward model. NeurIPS 2023, 36, 53728–53741. [Google Scholar] [CrossRef]
- Shao, Z.; Wang, P.; Zhu, Q.; Xu, R.; Song, J.; Bi, X.; Zhang, H.; Zhang, M.; et al. DeepSeekMath: Pushing the limits of mathematical reasoning in open language models. arXiv 2024, arXiv:2402.03300. [Google Scholar]
- Fan, L.; Zhang, F.; Fan, H.; Zhang, C. Brief review of image denoising techniques. Vis. Comput. Ind. Biomed. Art. 2019, 7. [Google Scholar] [PubMed]
- Yang, W.; Tan, R.T.; Wang, S.; Fang, Y.; Liu, J. Single image deraining: From model-based to data-driven and beyond. IEEE Trans. Pattern Anal. Mach. Intell. 2020, 43, 4059–4077. [Google Scholar]
- Conde, M.V.; Lu, Z.; Timofte, R. PixTalk: Controlling Photorealistic Image Processing and Editing with Language. In Proceedings of the ICCV; 2025; pp. 19269–19279. [Google Scholar] [CrossRef]
- Liu, S.; Ma, J.; Sun, L.; Kong, X.; Zhang, L. InstructRestore: Region-Customized Image Restoration with Human Instructions. arXiv 2025, arXiv:2503.24357. [Google Scholar]
- Qian, Y.; Bocek-Rivele, E.; Song, L.; Tong, J.; Yang, Y.; Lu, J.; Hu, W.; Gan, Z. Pico-Banana-400K: A Large-Scale Dataset for Text-Guided Image Editing. arXiv 2025, arXiv:2510.19808. [Google Scholar]
- Zhang, L.; Rao, A.; Agrawala, M. Adding conditional control to text-to-image diffusion models. In Proceedings of the ICCV, 2023; pp. 3836–3847. [Google Scholar]
- Xiao, B.; Wu, H.; Xu, W.; Dai, X.; Hu, H.; Lu, Y.; Zeng, M.; Liu, C.; Yuan, L. Florence-2: Advancing a unified representation for a variety of vision tasks. In Proceedings of the CVPR, 2024; pp. 4818–4829. [Google Scholar]
- Tang, A.; Wu, Y.; Zhang, Y. RamIR: Reasoning and action prompting with Mamba for all-in-one image restoration. Appl. Intell. 2025, 55, 258. [Google Scholar] [CrossRef]
- Lai, X.; Tian, Z.; Chen, Y.; Li, Y.; Yuan, Y.; Liu, S.; Jia, J. LISA: Reasoning segmentation via large language model. In Proceedings of the CVPR, 2024; pp. 9579–9589. [Google Scholar]
- Lan, Y.; Cui, Z.; Liu, C.; Peng, J.; Wang, N.; Luo, X.; Liu, D. Exploiting Diffusion Prior for Real-World Image Dehazing with Unpaired Training. In Proceedings of the AAAI; 2025; Vol. 39, pp. 4455–4463. [Google Scholar] [CrossRef]
- Li, J.; Li, D.; Savarese, S.; Hoi, S. BLIP-2: Bootstrapping language-image pre-training with frozen image encoders and large language models. In Proceedings of the ICML, 2023; pp. 19730–19742. [Google Scholar]
- Wang, W.; Lv, Q.; Yu, W.; Hong, W.; Qi, J.; Wang, Y.; Ji, J.; Yang, Z.; Zhao, L.; XiXuan, S.; et al. CogVLM: Visual expert for pretrained language models. NeurIPS 2024, 37, 121475–121499. [Google Scholar] [CrossRef]
- Jiang, A.Q.; Sablayrolles, A.; Mensch, A.; Bamford, C.; Chaplot, D.S.; de las Casas, D.; Bressand, F.; Lengyel, G.; Lample, G.; Saulnier, L.; et al. Mistral 7B. arXiv 2023, arXiv:2310.06825. [Google Scholar]
- Achiam, J.; Adler, S.; Agarwal, S.; Ahmad, L.; Akkaya, I.; Aleman, F.L.; Almeida, D.; Altenschmidt, J.; Altman, S.; Anadkat, S.; et al. GPT-4 technical report. arXiv 2023, arXiv:2303.08774. [Google Scholar]
- Zhu, D.; Chen, J.; Shen, X.; Li, X.; Elhoseiny, M. MiniGPT-4: Enhancing vision-language understanding with advanced large language models. arXiv 2023, arXiv:2304.10592. [Google Scholar]
- Brown, T.; Mann, B.; Ryder, N.; Subbiah, M.; Kaplan, J.D.; Dhariwal, P.; Neelakantan, A.; Shyam, P.; Sastry, G.; Askell, A.; et al. Language models are few-shot learners. NeurIPS 2020, 33, 1877–1901. [Google Scholar]
- Fanelli, N.; Vessio, G.; Castellano, G. I Dream My Painting: Connecting MLLMs and Diffusion Models via Prompt Generation for Text-Guided Multi-Mask Inpainting. In Proceedings of the WACV; 2025; pp. 6073–6082. [Google Scholar] [CrossRef]
- Team, G.; Anil, R.; Borgeaud, S.; Alayrac, J.B.; Yu, J.; Soricut, R.; Schalkwyk, J.; Dai, A.M.; Hauth, A.; Millican, K.; et al. Gemini: a family of highly capable multimodal models. arXiv 2023, arXiv:2312.11805. [Google Scholar]
- Van Den Oord, A.; Vinyals, O.; et al. Neural discrete representation learning. NeurIPS 2017, 30. [Google Scholar]
- Houlsby, N.; Giurgiu, A.; Jastrzebski, S.; Morrone, B.; De Laroussilhe, Q.; Gesmundo, A.; Attariyan, M.; Gelly, S. Parameter-efficient transfer learning for NLP. In Proceedings of the ICML, 2019; pp. 2790–2799. [Google Scholar]
- Hu, E.J.; Shen, Y.; Wallis, P.; Allen-Zhu, Z.; Li, Y.; Wang, S.; Wang, L.; Chen, W.; et al. LoRA: Low-rank adaptation of large language models. ICLR, 2022. [Google Scholar]
- Fan, G.; Zhou, S.; Hua, Z.; Li, J.; Zhou, J. LLaVA-based semantic feature modulation diffusion model for underwater image enhancement. Inform. Fus. 2025, 103566. [Google Scholar]
- Wei, H.; Liu, S.; Yuan, C.; Zhang, L. Perceive, Understand and Restore: Real-World Image Super-Resolution with Autoregressive Multimodal Generative Models. arXiv 2025, arXiv:2503.11073. [Google Scholar]
- Wang, C.; An, W.; Jiang, K.; Liu, X.; Jiang, J. LLV-FSR: Exploiting Large Language-Vision Prior for Face Super-resolution. arXiv 2024, arXiv:2411.09293. [Google Scholar]
- Sun, H.; Li, W.; Liu, J.; Zhou, K.; Chen, Y.; Guo, Y.; Li, Y.; Pei, R.; Peng, L.; Yang, Y. Beyond Pixels: Text Enhances Generalization in Real-World Image Restoration. arXiv 2024, arXiv:2412.00878. [Google Scholar]
- Kirillov, A.; Mintun, E.; Ravi, N.; Mao, H.; Rolland, C.; Gustafson, L.; Xiao, T.; Whitehead, S.; Berg, A.C.; Lo, W.Y.; et al. Segment anything. In Proceedings of the ICCV, 2023; pp. 4015–4026. [Google Scholar]
- Wu, H.; Zhang, Z.; Zhang, W.; Chen, C.; Liao, L.; Li, C.; Gao, Y.; Wang, A.; Zhang, E.; Sun, W.; et al. Q-Align: Teaching LMMs for visual scoring via discrete text-defined levels. arXiv 2023, arXiv:2312.17090. [Google Scholar]
- Johnson, J.; Alahi, A.; et al. Perceptual losses for real-time style transfer and super-resolution. In Proceedings of the ECCV, 2016; pp. 694–711. [Google Scholar]
- He, X.; Li, L.; Wang, Y.; Zheng, H.; Cao, K.; Yan, K.; Li, R.; Xie, C.; Zhang, J.; Zhou, M. Training-Free Large Model Priors for Multiple-in-One Image Restoration. arXiv 2024, arXiv:2407.13181. [Google Scholar]
- Wang, Z.; Wu, Z.; Agarwal, D.; Sun, J. MedCLIP: Contrastive learning from unpaired medical images and text. In Proceedings of the EMNLP; 2022; Vol. 2022, p. 3876. [Google Scholar] [CrossRef]
- Sun, Q.; Fang, Y.; Wu, L.; Wang, X.; Cao, Y. Eva-CLIP: Improved training techniques for CLIP at scale. arXiv 2023, arXiv:2303.15389. [Google Scholar]
- Kumari, S.; Singh, P. Data efficient deep learning for medical image analysis: A survey. arXiv 2023, arXiv:2310.06557. [Google Scholar]
- Zhang, S.; Xu, Y.; Usuyama, N.; Bagga, J.; Tinn, R.; Preston, S.; Rao, R.; Wei, M.; Valluri, N.; Wong, C.; et al. Large-scale domain-specific pretraining for biomedical vision-language processing. arXiv 2023, arXiv:2303.00915. [Google Scholar]
- He, Y.; Guo, P.; Tang, Y.; Myronenko, A.; Nath, V.; Xu, Z.; Yang, D.; Zhao, C.; Simon, B.; Belue, M.; et al. Vista3D: Versatile imaging segmentation and annotation model for 3D computed tomography. arXiv 2024, arXiv:2406.05285. [Google Scholar]
- Yang, Z.; Chen, Y.; Wang, Z.; Shan, H.; Chen, Y.; Zhang, Y. Patient-level anatomy meets scanning-level physics: Personalized federated Low-Dose CT denoising empowered by large language model. arXiv 2025, arXiv:2503.00908. [Google Scholar]
- Kim, K.; Na, Y.; Ye, S.J.; Lee, J.; Ahn, S.S.; Park, J.E.; Kim, H. Controllable text-to-image synthesis for multi-modality MR images. In Proceedings of the WACV, 2024; pp. 7936–7945. [Google Scholar]
- Wang, X.; Yu, K.; Wu, S.; Gu, J.; Liu, Y.; Dong, C.; Qiao, Y.; Change Loy, C. ESRGAN: Enhanced super-resolution generative adversarial networks. In Proceedings of the ECCVW, 2018; pp. 0–0. [Google Scholar]
- Agustsson, E.; Timofte, R. NTIRE 2017 Challenge on Single Image Super-Resolution: Dataset and Study. In Proceedings of the CVPRW, July 2017. [Google Scholar]
- Cai, J.; Zeng, H.; Yong, H.; Cao, Z.; Zhang, L. Toward real-world single image super-resolution: A new benchmark and a new model. In Proceedings of the ICCV, 2019; pp. 3086–3095. [Google Scholar]
- Wei, P.; Xie, Z.; Lu, H.; Zhan, Z.; Ye, Q.; Zuo, W.; Lin, L. Component divide-and-conquer for real-world image super-resolution. In Proceedings of the ECCV, 2020; pp. 101–117. [Google Scholar]
- Wei, C.; Wang, W.; Yang, W.; Liu, J. Deep Retinex decomposition for low-light enhancement. arXiv 2018, arXiv:1808.04560. [Google Scholar]
- Guo, X.; Li, Y.; Ling, H. LIME: Low-light image enhancement via illumination map estimation. IEEE Trans. Image Process. 2016, 26, 982–993. [Google Scholar] [CrossRef]
- Yang, W.; Wang, W.; Huang, H.; Wang, S.; Liu, J. Sparse gradient regularized deep Retinex network for robust low-light image enhancement. IEEE Trans. Image Process. 2021, 30, 2072–2086. [Google Scholar] [CrossRef] [PubMed]
- Li, B.; Ren, W.; Fu, D. Benchmarking single-image dehazing and beyond. IEEE Trans. Image Process. 2018, 492–505. [Google Scholar] [CrossRef]
- Ancuti, C.; Ancuti, C. An image dehazing benchmark with non-homogeneous hazy images. In Proceedings of the CVPRW, 2020; pp. 444–445. [Google Scholar]
- Liu, Y.; Zhu, L. From synthetic to real: Image dehazing collaborating with real data. In Proceedings of the ACM MM, 2021; pp. 50–58. [Google Scholar]
- Liu, Z.; Luo, P.; Wang, X.; Tang, X. Deep learning face attributes in the wild. In Proceedings of the ICCV, 2015; pp. 3730–3738. [Google Scholar]
- Karras, T.; Aila, T.; Laine, S.; Lehtinen, J. Progressive growing of GANs for improved quality, stability, and variation. arXiv 2017, arXiv:1710.10196. [Google Scholar]
- Lin, T.Y.; Maire, M.; Belongie, S.; Hays, J.; Perona, P.; Ramanan, D.; Dollár, P.; Zitnick, C.L. Microsoft coco: Common objects in context. In Proceedings of the ECCV, 2014; pp. 740–755. [Google Scholar]
- Yang, W.; Tan, R.T.; Feng, J.; Liu, J.; Guo, Z.; Yan, S. Deep joint rain detection and removal from a single image. In Proceedings of the CVPR, 2017; pp. 1357–1366. [Google Scholar]
- Qian, R.; Tan, R.T.; Yang, W.; Su, J.; Liu, J. Attentive generative adversarial network for raindrop removal from a single image. In Proceedings of the CVPR, 2018; pp. 2482–2491. [Google Scholar]
- Ba, Y.; Zhang, H.; Yang, E.; Suzuki, A.; Pfahnl, A.; Chandrappa, C.C.; de Melo, C.M.; You, S.; Soatto, S.; Wong, A.; et al. Not Just Streaks: Towards Ground Truth for Single Image Deraining. In Proceedings of the ECCV, 2022; pp. 723–740. [Google Scholar]
- Nah, S.; Hyun Kim, T.; Mu Lee, K. Deep multi-scale convolutional neural network for dynamic scene deblurring. In Proceedings of the CVPR, 2017; pp. 3883–3891. [Google Scholar]
- Rim, J.; Lee, H.; Won, J.; Cho, S. Real-world blur dataset for learning and benchmarking deblurring algorithms. In Proceedings of the ECCV, 2020; pp. 184–201. [Google Scholar]
- Shen, Z.; Wang, W.; Lu, X.; Shen, J.; Ling, H.; Xu, T.; Shao, L. Human-aware motion deblurring. In Proceedings of the ICCV, 2019; pp. 5572–5581. [Google Scholar]
- McCollough, C.H.; Bartley, A.C.; Carter, R.E.; Chen, B.; Drees, T.A.; Edwards, P.; Holmes, D.R., III; Huang, A.E.; Khan, F.; Leng, S.; et al. Low-Dose CT for the detection and classification of metastatic liver lesions: results of the 2016 low dose CT grand challenge. Med. Phys. 2017, 44, e339–e352. [Google Scholar] [CrossRef] [PubMed]
- Moen, T.R.; Chen, B.; Holmes, D.R., III; Duan, X.; Yu, Z.; Yu, L.; Leng, S.; Fletcher, J.G.; McCollough, C.H. Low-Dose CT image and projection dataset. Med. Phys. 2021, 48, 902–911. [Google Scholar] [PubMed]
- Zhu, H.; Wu, W.; Zhu, W.; Jiang, L.; Tang, S.; Zhang, L.; Liu, Z.; Loy, C.C. CelebV-HQ: A large-scale video facial attributes dataset. In Proceedings of the ECCV, 2022; pp. 650–667. [Google Scholar]
- Dai, W.; Li, J.; Li, D.; Tiong, A.; Zhao, J.; Wang, W.; Li, B.; Fung, P.N.; Hoi, S. InstructBLIP: Towards general-purpose vision-language models with instruction tuning. NeurIPS 2023, 36, 49250–49267. [Google Scholar] [CrossRef]
- Sun, H.; Li, W.; Liu, J.; Chen, H.; Pei, R.; Zou, X.; Yan, Y.; Yang, Y. CoSeR: Bridging image and language for cognitive super-resolution. In Proceedings of the CVPR, 2024; pp. 25868–25878. [Google Scholar]
- Martin, D.; Fowlkes, C.; Tal, D.; Malik, J. A database of human segmented natural images and its application to evaluating segmentation algorithms and measuring ecological statistics. In Proceedings of the ICCV; IEEE, 2001; Vol. 2, pp. 416–423. [Google Scholar]
- Huang, J.B.; Singh, A.; Ahuja, N. Single image super-resolution from transformed self-exemplars. In Proceedings of the CVPR, 2015; pp. 5197–5206. [Google Scholar]
- Yan, Q.; Jiang, A.; Chen, K.; Peng, L.; Yi, Q.; Zhang, C. Textual prompt guided image restoration. arXiv 2023, arXiv:2312.06162. [Google Scholar]
- Cheng, J.; Liang, D.; Tan, S. Transfer CLIP for generalizable image denoising. In Proceedings of the CVPR, 2024; pp. 25974–25984. [Google Scholar]
- Tian, X.; Liao, X.; Liu, X.; Li, M.; Ren, C. Degradation-Aware Feature Perturbation for All-in-One Image Restoration. In Proceedings of the CVPR, 2025; pp. 28165–28175. [Google Scholar]
- Liu, Y.F.; Jaw, D.W.; Huang, S.C.; Hwang, J.N. DesnowNet: Context-aware deep network for snow removal. IEEE Trans. Image Process. 2018, 27, 3064–3073. [Google Scholar] [CrossRef]
- Li, R.; Cheong, L.F.; Tan, R.T. Heavy rain image restoration: Integrating physics model and conditional adversarial learning. In Proceedings of the CVPR, 2019; pp. 1633–1642. [Google Scholar]
- Ai, Y.; Huang, H.; Zhou, X.; Wang, J.; He, R. Multimodal prompt perceiver: Empower adaptiveness generalizability and fidelity for all-in-one image restoration. In Proceedings of the CVPR, 2024; pp. 25432–25444. [Google Scholar]
- Yang, S.; Ding, M.; Wu, Y.; Li, Z.; Zhang, J. Implicit neural representation for cooperative low-light image enhancement. In Proceedings of the ICCV, 2023; pp. 12918–12927. [Google Scholar]
- Song, S. Noise-Resilient Low-Light Image Enhancement with CLIP Guidance and Pixel-Reordering Subsampling. Electronics 2025, 14, 4839. [Google Scholar] [CrossRef]
- Yan, Q.; Shi, K.; Feng, Y.; Hu, T.; Wu, P.; Pang, G.; Zhang, Y. HVI-CIDNet+: Beyond Extreme Darkness for Low-Light Image Enhancement. arXiv 2025, arXiv:2507.06814. [Google Scholar]
- Zhang, H.; Sindagi, V.; Patel, V.M. Image de-raining using a conditional generative adversarial network. IEEE Trans. Circuits Syst. Video Technol. 2019, 30, 3943–3956. [Google Scholar] [CrossRef]
- Wei, X.; Ye, X.; Mei, X.; Wang, J.; Ma, H. A single image deraining algorithm guided by text generation based on depth information conditions. Appl. Soft Comput. 2025, 113506. [Google Scholar]
- Rajagopalan, S.; Patel, V.M. AWRaCLe: All-weather image restoration using visual in-context learning. In Proceedings of the AAAI; 2025; Vol. 39, pp. 6675–6683. [Google Scholar] [CrossRef]
- Zhang, K.; Mo, L.; Chen, W.; Sun, H.; Su, Y. MagicBrush: A Manually Annotated Dataset for Instruction-Guided Image Editing. arXiv 2024, arXiv:2306.10012. [Google Scholar]
- Yu, Q.; Chow, W.; Yue, Z.; Pan, K.; Wu, Y.; Wan, X.; Li, J.; Tang, S.; Zhang, H.; Zhuang, Y. AnyEdit: Mastering Unified High-Quality Image Editing for Any Idea. arXiv 2025, arXiv:2411.15738. [Google Scholar]
- Brooks, T.; Holynski, A.; Efros, A.A. InstructPix2Pix: Learning to follow image editing instructions. In Proceedings of the CVPR, 2023; pp. 18392–18402. [Google Scholar]
- Zhang, Z.; Xie, J.; Lu, Y.; Yang, Z.; Yang, Y. In-Context Edit: Enabling Instructional Image Editing with In-Context Generation in Large Scale Diffusion Transformer. arXiv 2025, arXiv:2504.20690. [Google Scholar]
- Zhao, H.; Ma, X.; Chen, L.; Si, S.; Wu, R.; An, K.; Yu, P.; Zhang, M.; Li, Q.; Chang, B. UltraEdit: Instruction-based Fine-Grained Image Editing at Scale. arXiv 2024, arXiv:2407.05282. [Google Scholar]
- Chow, W.; Li, L.; Kong, L.; Li, Z.; Xu, Q.; Song, H.; Ye, T.; Wang, X.; Bai, J.; Xu, S.; et al. EditMGT: Unleashing Potentials of Masked Generative Transformers in Image Editing. arXiv 2026, arXiv:2512.11715. [Google Scholar]
- Cai, Z.; Yeh, C.F.; Xu, H.; Liu, Z.; Meyer, G.; Lei, X.; Zhao, C.; Li, S.W.; Chandra, V.; Shi, Y. DepthLM: Metric Depth From Vision Language Models. arXiv 2025, arXiv:2509.25413. [Google Scholar]
- Zhang, J.; Zhou, S.; Liu, B.; Kadambi, A.; Fan, Z. SpatialStack: Layered Geometry-Language Fusion for 3D VLM Spatial Reasoning. arXiv 2026, arXiv:2603.27437. [Google Scholar]
- Schuhmann, C.; Beaumont, R.; Vencu, R.; Gordon, C.; Wightman, R.; Cherti, M.; Coombes, T.; Katta, A.; Mullis, C.; Wortsman, M.; et al. LAION-5B: An open large-scale dataset for training next generation image-text models. NeurIPS 2022, 35, 25278–25294. [Google Scholar] [CrossRef]
- Sharma, P.; Ding, N.; Goodman, S.; Soricut, R. Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning. In Proceedings of the ACL, 2018; pp. 2556–2565. [Google Scholar]
- Ordonez, V.; Kulkarni, G.; Berg, T. Im2Text: Describing images using 1 million captioned photographs. NeurIPS 2011, 24. [Google Scholar]
- Chen, X.; Fang, H.; Lin, T.Y.; Vedantam, R.; Gupta, S.; Dollár, P.; Zitnick, C.L. Microsoft COCO captions: Data collection and evaluation server. arXiv 2015, arXiv:1504.00325. [Google Scholar]
- Yuan, L.; Chen, D.; Chen, Y.L.; Codella, N.; Dai, X.; Gao, J.; Hu, H.; Huang, X.; Li, B.; Li, C.; et al. Florence: A new foundation model for computer vision. arXiv 2021, arXiv:2111.11432. [Google Scholar]
- Alayrac, J.B.; Donahue, J.; Luc, P.; Miech, A.; Barr, I.; Hasson, Y.; Lenc, K.; Mensch, A.; Millican, K.; Reynolds, M.; et al. Flamingo: a visual language model for few-shot learning. NeurIPS 2022, 35, 23716–23736. [Google Scholar] [CrossRef]
- Chen, X.; Wang, X.; Changpinyo, S.; Piergiovanni, A.; Padlewski, P.; Salz, D.; Goodman, S.; Grycner, A.; Mustafa, B.; Beyer, L.; et al. PaLI: A jointly-scaled multilingual language-image model. arXiv 2022, arXiv:2209.06794. [Google Scholar]
- Zhang, R.; Isola, P.; Efros, A.A.; Shechtman, E.; Wang, O. The unreasonable effectiveness of deep features as a perceptual metric. In Proceedings of the CVPR, 2018; pp. 586–595. [Google Scholar]
- Heusel, M.; Ramsauer, H.; Unterthiner, T.; Nessler, B.; Hochreiter, S. GANs trained by a two time-scale update rule converge to a local Nash equilibrium. NeurIPS 2017, 30. [Google Scholar]
- Ke, J.; Wang, Q.; Wang, Y.; Milanfar, P.; Yang, F. MUSIQ: Multi-scale image quality transformer. In Proceedings of the ICCV, 2021; pp. 5148–5157. [Google Scholar]
- Yang, C.; Dong, R.; Lam, K.M. Vision-Language Model Guided Image Restoration. arXiv 2025, arXiv:2512.17292. [Google Scholar]
- Xiao, F.; Feng, J.; Hu, P.; Zhang, D.; Xu, L.; Qin, G.; Li, L.; He, C.; Farsiu, S. Qualiteacher: Quality-conditioned pseudo-labeling for real-world image restoration. arXiv 2026, arXiv:2603.08030. [Google Scholar]
- Xiao, F.; Hu, P.; Xu, L.; Guo, X.; Qin, G.; Shen, Y.; Fang, C.; Zhang, R.; He, C.; Farsiu, S. Beyond Ground-Truth: Leveraging Image Quality Priors for Real-World Image Restoration. CVPR, 2026. [Google Scholar]
- Gu, A.; Dao, T. Mamba: Linear-time sequence modeling with selective state spaces. arXiv 2023, arXiv:2312.00752. [Google Scholar]
- Ramesh, A.; Dhariwal, P.; Nichol, A.; Chu, C.; Chen, M. Hierarchical text-conditional image generation with CLIP latents. arXiv 2022, arXiv:2204.061251, 3. [Google Scholar]
- Hurst, A.; Lerer, A.; Goucher, A.P.; Perelman, A.; Ramesh, A.; Clark, A.; Ostrow, A.; Welihinda, A.; Hayes, A.; Radford, A.; et al. GPT-4o system card. arXiv 2024, arXiv:2410.21276. [Google Scholar]
- Pu, Y.; Zhuo, L.; Zhu, K.; Xie, L.; Zhang, W.; Chen, X.; Gao, P.; Qiao, Y.; Dong, C.; Liu, Y. Lumina-OmniLV: A unified multimodal framework for general low-level vision. arXiv 2025, arXiv:2504.04903. [Google Scholar]
- Blau, Y.; Michaeli, T. The perception-distortion tradeoff. In Proceedings of the CVPR, 2018; pp. 6228–6237. [Google Scholar]
- He, C.; Zhang, R.; Zhang, D.; Xiao, F.; Fan, D.P.; Farsiu, S. Nested Unfolding Network for Real-World Concealed Object Segmentation. arXiv 2025, arXiv:2511.18164. [Google Scholar]
- Zhu, H.; Luo, M.D.; Wang, R.; Zheng, A.H.; He, R. Deep audio-visual learning: A survey. Int. J. Autom. Comput. 2021, 18, 351–376. [Google Scholar] [CrossRef]
- Duan, J.; Yu, S.; Tan, H.L.; Zhu, H.; Tan, C. A survey of Embodied AI: From simulators to research tasks. IEEE Trans. Emerg. Top. Comput. Intell. 2022, 6, 230–244. [Google Scholar] [CrossRef]
- Wang, T.; Mao, X.; Zhu, C.; Xu, R.; Lyu, R.; et al. EmbodiedScan: A holistic multi-modal 3D perception suite towards Embodied AI. In Proceedings of the CVPR, 2024; pp. 19757–19767. [Google Scholar]
- Zhang, C.; Gong, D.; He, J.; Zhu, Y.; Sun, J.; Zhang, Y. UIR-LoRA: Achieving Universal Image Restoration through Multiple Low-Rank Adaptation. arXiv 2024, arXiv:2409.20197. [Google Scholar]
- Ren, P.; Xiao, Y.; Chang, X.; Huang, P.Y.; Li, Z.; Chen, X.; Wang, X. A comprehensive survey of neural architecture search: Challenges and solutions. CSUR 2021, 54, 1–34. [Google Scholar] [CrossRef]
- Wang, P.; Luo, X.; Xie, Y.; Qu, Y. Data-free Distillation with Degradation-prompt Diffusion for Multi-weather Image Restoration. arXiv 2024, arXiv:2409.03455. [Google Scholar]
- He, C.; Shen, Y.; Fang, C.; Xiao, F.; Tang, L. Diffusion Models in Low-Level Vision: A Survey. IEEE Trans. Pattern Anal. Mach. Intell. 2024. [Google Scholar]
- Chen, D.; Wang, D.; Darrell, T.; Ebrahimi, S. Contrastive test-time adaptation. In Proceedings of the CVPR, 2022; pp. 295–305. [Google Scholar]
- Tang, J.; Chen, J.; Wei, W.; Xu, X.; Liu, R.; Wu, X.; Xie, Q.; Wu, J.; Zhang, L.; Chen, Q. Robust-R1: Degradation-Aware Reasoning for Robust Visual Understanding. arXiv 2025, arXiv:2512.17532. [Google Scholar]
- Wei, Y.; Zhang, Z.; Ren, J.; Xu, X.; Hong, R.; Yang, Y.; Yan, S.; Wang, M. Clarity ChatGPT: An interactive and adaptive processing system for image restoration and enhancement. arXiv 2023, arXiv:2311.11695. [Google Scholar]
- Orjuela, D.Y.G.; Scappatura, L.; Di Gennaro, V.; Izzo, R.A.; Bardaro, G.; Matteucci, M. Improving Robustness of Vision-Language-Action Models by Restoring Corrupted Visual Inputs. arXiv 2026, arXiv:2602.01158. [Google Scholar]
- Poria, S.; Majumder, N.; Hung, C.Y.; Bagherzadeh, A.A.; Li, C.; Kwok, K.; Wang, Z.; Tan, C.; Wu, J.; Hsu, D. 10 open challenges steering the future of vision-language-action models. In Proceedings of the AAAI, 2026; pp. 39771–39779. [Google Scholar]










| Model | Year | Technical points | Function | Paper | Open |
|---|---|---|---|---|---|
| LLaVA-Med | 2023 | Self-supervised instruction tuning; Two-stage curriculum learning. | Medical VQA, Image captioning | [149] | Y |
| Med-Flamingo | 2023 | Few-shot learning; OpenFlamingo pre-training. | Rationale generation, Case solving | [148] | Y |
| BioViL-T | 2023 | CNN-Transformer hybrid; Temporal image-report pairing. | Disease progression modeling | [150] | Y |
| XrayGPT | 2023 | Frozen MedCLIP [222] and LLM; Alignment layer tuning. | Radiograph summarization & QA | [144] | Y |
| MedKLIP | 2023 | Entity-level alignment; Structured triplet supervision. | Zero-shot diagnosis, Grounding | [147] | Y |
| VILA-M3 | 2024 | Expert-guided tuning; Integration of expert models. | Medical VQA, Classification | [145] | Y |
| MiniGPT-Med | 2024 | Frozen EVA [223] projection to LLaMA2; Task token guidance. | Radiology diagnosis, Report generation | [146] | Y |
| HuatuoGPT | 2024 | VLM-assisted data reformatting; Large-scale alignment. | Cross-modality reasoning | [151] | Y |
| HealthGPT | 2025 | Heterogeneous LoRA; Task-aware feature selection. | Modality conversion, Super-resolution | [152] | Y |
| Failure mode | Likely cause | Partial mitigation |
|---|---|---|
| Text hallucination | Over-association with semantic priors; weak text grounding | JarvisIR [74] human-feedback alignment stage |
| Noise introduction | Low-quality-caption noise leaking into the prior | VLMIR [285] caption-alignment cosine-similarity loss |
| Limited generative recovery | Insufficient quality-conditioned generative capacity | Quality-conditioned supervision: QualiTeacher [286], IQPIR [287] |
| Zero-shot failure | Poor transfer to unseen degradations | EvoQuality [133] (voting + GRPO); DA-CLIP [29] |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).