首页 >  , Vol. , Issue () : -

摘要

全文摘要次数: 257 全文下载次数: 350
引用本文:

DOI:

10.11834/jrs.20264292

收稿日期:

2024-07-15

修改日期:

2026-03-01

PDF Free   EndNote   BibTeX
基于Cross-Transformer差异特征融合的高分辨率光学遥感影像变化检测
彭代锋, 周顶蔚, 管海燕
南京信息工程大学
摘要:

针对卷积网络在变化检测中面临的感受野有限、跨层次特征交互不足及长距离上下文依赖建模困难等问题,本文提出基于交叉Transformer差分特征融合的变化检测网络(Cross-Transformer-based Difference Feature Fusion Network for Change Detection, CTDFFNet)。首先,利用ResNet18网络构建多尺度局部特征表达,并结合Cross-Transformer模块学习差异特征表示以实现全局上下文感知和多层次特征交互,增强模型对复杂场景变化的敏感度;接着,解码部分利用密集连接结构实现特征的多路径传播和信息的充分利用,同时抑制梯度消失;最后,结合自适应通道增强模块(Adaptive Channel Enhancement Module, ACEM)动态调整通道权重以增强关键信息特征表达,并利用Sigmoid层生成最后变化图。为验证本文方法有效性,采用LEVIR-CD、CDD和WHU-CD等三个变化检测数据集进行实验。定量结果表明,与对比方法相比,CTDFFNet在三个数据集均取得最优精度指标,其F1分数分别达到91.52%、95.83%和93.02%。此外,目视结果表明,CTDFFNet能有效抵抗复杂背景、阴影和光照变化等因素干扰,显著增强网络对不同尺度目标的检测识别效果。

High-resolution optical remote sensing image change detection based on Cross-Transformer difference feature fusion
Abstract:

Recently, deep learning-based change detection (CD) models, particularly Convolutional Neural Networks (CNNs), have emerged as the de facto mainstream approach. However, such methods suffer from several inherent architectural limitations. On the one hand, their inherently constrained receptive fields hinder the acquisition of global contextual information, which is critical for modeling long-range dependencies and interpreting complex geographical scenes. On the other hand, the limited interaction across multi-level features results in suboptimal integration of high-level semantic information and low-level spatial details. To address these challenges, this paper proposes a novel CD network termed Cross-Transformer-based Difference Feature Fusion Network (CTDFFNet), which systematically integrates global context perception, multi-level feature interaction, and adaptive feature refinement into a unified framework. The proposed CTDFFNet adopts an encoder-decoder architecture, leveraging a ResNet18 backbone pre-trained on ImageNet in the encoding stage to extract multi-scale local feature representations from bi-temporal remote sensing images. To mitigate the limited receptive field of standard convolutions, a Cross-Transformer module is introduced to learn discriminative difference features. By leveraging a hierarchical cross-attention mechanism, it captures long-range spatial dependencies while facilitating cross-level feature interaction, thereby improving the sensitivity to changes under complex conditions such as varying illumination, shadow effects, or seasonal variations. In the decoding stage, a dense connection mechanism is employed, where each layer receives inputs from all preceding layers. This design promotes feature reuse and ensures comprehensive utilization of hierarchical information, thereby mitigating the gradient vanishing problem and enabling more stable training. Furthermore, an Adaptive Channel Enhancement Module (ACEM) is integrated to dynamically recalibrate channel-wise feature responses by learning attention weights, thereby emphasizing change-salient features while suppressing irrelevant background noise. Finally, a Sigmoid layer is applied to generate the final change map. To comprehensively evaluate the effectiveness and generalization capability of CTDFFNet, extensive experiments were conducted on three publicly available CD datasets: LEVIR-CD, CDD, and WHU-CD. Quantitative experimental results demonstrate that CTDFFNet consistently outperforms state-of-the-art methods across all three datasets, achieving the highest F1 scores of 91.52%, 95.88%, and 93.02% on LEVIR-CD, CDD, and WHU-CD, respectively. These significant improvements confirm the effectiveness of integrating Transformer mechanisms into CD frameworks. Visual results reveal that by effectively suppressing the interference from complex backgrounds and environmental variations, CTDFFNet generates change maps with notably fewer false alarms and missed detections, while maintaining high precision in delineating change objects of varying scales. To conclude, the proposed CTDFFNet effectively addresses the inherent limitations in CNN-based CD frameworks through three key innovations: (1) The Cross-Transformer module captures long-range dependencies and facilitates cross-level feature interaction, enhancing sensitivity to complex scene changes. (2) A densely connected decoding stage facilitates multi-scale feature utilization while alleviating gradient vanishing for stable training. (3) The ACEM further suppresses background interference by dynamically emphasizing change-salient features. Extensive experiments across multiple datasets demonstrate that CTDFFNet consistently outperforms existing methods, exhibiting strong generalization and robustness under challenging conditions. These findings underscore its potential for various remote sensing CD applications. Future work will explore extensions to multi-modal CD scenarios and investigate label-efficient learning strategies to reduce dependence on large-scale annotated data.

本文暂时没有被引用!

欢迎关注学报微信

遥感学报交流群