首页 >  , Vol. , Issue () : -

摘要

全文摘要次数: 1280 全文下载次数: 2161
引用本文:

DOI:

10.11834/jrs.20254427

收稿日期:

2024-09-25

修改日期:

2025-04-18

PDF Free   EndNote   BibTeX
用于零样本分类的遥感视觉语言模型:综述
檀晓萌1,2, 席博博3,4, 薛长斌1, 李云松5, 徐海涛1
1.中国科学院国家空间科学中心 复杂航天系统电子信息技术重点实验室;2.中国科学院大学;3.中国科学院国家空间科学中心;4.西安电子科技大学;5.西安电子科技大学 通信工程学院
摘要:

视觉语言模型在零-小样本分类、图文检索、图像字幕、视觉问答和视觉定位等多种多模态任务上取得了显著的进展。然而,大多数方法依赖于通用数据集的预训练,导致在如遥感、医学等特殊领域的泛化性能不佳。最近提出了许多遥感视觉语言模型,这些新模型通过构建大规模的遥感图像文本对数据来微调通用视觉语言模型,以实现具备地理感知能力的遥感域专用视觉语言模型。本文以零样本分类任务为主线,总结并分析了遥感视觉语言模型的最新发展,主要是利用遥感数据对通用视觉语言模型微调来构建遥感域专用模型,整体驱动方式依赖于高质量标注的遥感数据和高性能的算力资源。此外,当前模型的发展较为分散多样,这使得遥感视觉语言模型的统一基准评价难以建立。为解决上述问题,未来或许可结合模型架构设计、微调技术等进行算力优化,同时依据多样任务特点逐步完善评价体系。

Remote Sensing Visual Language Models for Zero-Shot Classification: A Survey
Abstract:

Visual language models have made remarkable progress in various multimodal tasks such as zero-shot classification, image-text retrieval, image captioning, visual question answering, and visual localization. However, most methods rely on pre-training on general datasets, resulting in poor generalization performance in special domains such as remote sensing and medicine. Many remote sensing visual language models have been proposed recently, which aim to achieve remote sensing domain-specific visual language models with geo-awareness by fine-tuning general visual language models by constructing large-scale remote sensing image-text pair data. In this paper, we summarize and analyze the latest developments of remote sensing visual language models with zero-shot classification tasks as the main line, mainly using remote sensing data to fine-tune the general visual language model to build a remote sensing domain-specific model. The overall approach to model development heavily depends on high-quality annotated remote sensing data and high-performance computing resources. However, the current landscape of model development remains fragmented and highly diverse, posing significant challenges in establishing a unified benchmark for evaluating remote sensing visual-language models. To address these issues, future research could focus on integrating advancements in model architecture design and fine-tuning techniques to optimize computational efficiency. Additionally, a systematic evaluation framework should be progressively refined to account for the specific characteristics of different remote sensing tasks.

本文暂时没有被引用!

欢迎关注学报微信

遥感学报交流群