首页 >  , Vol. , Issue () : -

摘要

全文摘要次数: 507 全文下载次数: 766
引用本文:

DOI:

10.11834/jrs.20255437

收稿日期:

2025-10-15

修改日期:

2025-12-25

PDF Free   EndNote   BibTeX
遥感跨模态图文检索:关键技术与挑战
王懿婧, 唐旭, 韩硕, 杜瑞琦
西安电子科技大学 人工智能学院 西安
摘要:

遥感跨模态图文检索作为连接自然语言与遥感影像的桥梁,旨在构建高效的双向语义关联,是遥感数据智能化分析的关键技术。本文全面概述了遥感跨模态图文检索领域的技术演进与研究现状。首先,细致分析了主流基准数据集在规模、场景类别及文本标注质量方面的特点,并介绍了通用的评价指标体系,为后续方法研究奠定基础。其次,回顾了文本特征表示从传统统计到深度学习,以及遥感影像特征表示从手工特征到深度神经网络的技术突破。再次,以是否采用跨模态预训练模型为划分,深入剖析了基于非跨模态预训练和基于跨模态预训练方法的原理及模型特点,并通过在三个主流数据集上的实验对比,揭示了跨模态预训练方法的性能优势及其不同微调策略的数据适配规律。最后,本文总结了当前研究面临的细粒度语义对齐、多源数据融合与跨域泛化、以及时序动态匹配机制缺失等核心挑战,并展望了未来在细粒度特征增强、多源异构数据协同建模和时序感知对齐机制和开发等方面的研究方向,以期推动遥感跨模态图文检索技术在实际应用中的深入发展。

Remote sensing cross-modal image-text retrieval: key technologies and challenges
Abstract:

With the rapid development of remote sensing technology and the continuous expansion of application scenarios, efficient retrieval and intelligent utilization of massive remote sensing data have become core demands in the field of remote sensing information processing. Remote sensing cross-modal image-text retrieval (RS-CMITR), as a key technology connecting visual perception and language understanding, establishes semantic associations between natural language descriptions and remote sensing image content to achieve efficient bidirectional interaction between text-to-image retrieval and image-to-text retrieval. This technology breaks down modality barriers and enables precise semantic localization of remote sensing data, providing intelligent solutions for disaster detection, environmental monitoring, and urban planning. This paper aims to systematically review the technological evolution, core methodologies, and major challenges in the RS-CMITR field, providing comprehensive reference for in-depth research in this domain. This review systematically analyzes RS-CMITR technology from datasets, feature representation, model architectures, to learning paradigms. First, we introduce three mainstream benchmark datasets—UCM-Captions, RSICD, and RSITMD—analyzing their characteristics in terms of scale, scene diversity, and text annotation quality, and explain the evaluation metrics including Recall@K and mean Recall. Second, we review the evolution of text feature representation methods from traditional statistical approaches to deep learning methods, as well as the development of remote sensing image feature representation from hand-crafted features to deep neural networks. Third, using cross-modal pre-training as the classification criterion, we categorize existing RS-CMITR methods into two major classes: methods based on non-cross-modal pre-training (including dual-encoder architecture and fusion encoder architecture) and methods based on cross-modal pre-training (including full fine-tuning, prompt learning, and adapter learning). We provide in-depth analysis of the technical principles, model characteristics, advantages, and conduct comprehensive performance comparisons of representative algorithms across the benchmark datasets. Experimental comparisons reveal several important findings regarding RS-CMITR methods. Cross-modal pre-training methods consistently demonstrate superior retrieval performance compared to non-cross-modal pre-training methods across all benchmark datasets. Different fine-tuning strategies exhibit distinct data adaptability patterns: full fine-tuning and adapter learning excel on large-scale datasets, while prompt learning shows advantages on small-scale datasets, highlighting the effectiveness of parameter-efficient fine-tuning. Dataset quality, particularly text diversity, significantly influences model performance. The review demonstrates that RS-CMITR has evolved from traditional feature engineering to deep learning-driven intelligent retrieval paradigms, with cross-modal pre-training combined with parameter-efficient fine-tuning emerging as the mainstream technical approach. Despite significant progress in RS-CMITR technology, three core challenges remain: (1) Fine-grained semantic alignment is difficult, as existing methods struggle to capture subtle differences between similar land cover types and establish precise image-text correspondences; (2) Multi-source data fusion and cross-domain generalization capabilities are insufficient, with models showing significant performance degradation in cross-domain and cross-sensor tasks; (3) Temporal dynamic matching mechanisms are absent, as current research focuses on static images and cannot effectively model temporal changes in land cover. Future research should focus on enhancing fine-grained feature representation, collaborative modeling of multi-source heterogeneous data, and construction of temporal-aware dynamic alignment mechanisms to advance RS-CMITR technology from theory to practical applications.

本文暂时没有被引用!

欢迎关注学报微信

遥感学报交流群