下载中心
优秀审稿专家
优秀论文
相关链接
首页 > , Vol. , Issue () : -
摘要

针对遥感影像多视图立体任务中存在的特征匹配精度低、预测深度图存在噪声和边缘重建不完整等问题,提出一种基于扩散约束的多视图立体网络(DiffusionMVS)。首先,在特征金字塔网络的基础上,设计基于特征增强的多尺度特征提取模块MFE-FPN(Multi-scale Feature Enhancement Feature Pyramid Network)来增强网络学习多视图遥感影像特征的能力;其次,提出自适应特征聚合模块AFA(Adaptive Feature Aggregation)来动态整合不同层次的特征以捕获目标边缘的深度细节特征;最后,设计基于扩散约束的代价体优化模型DCM(Diffusion Constrained Module),通过优化存在噪声点的深度值分布来消除预测深度图存在的噪声干扰,并结合边缘引导的Transformer网络优化深度图边缘重建效果。实验结果显示,在WHU-TLC和LuoJia-MVS数据集测试中,与基准模型相比,DiffusionMVS网络的MAE指标分别提升了28.11%和3.37%,展示了较好的重建性能和泛化能力。
Large-scale 3D scene reconstruction based on remote sensing images provides critical support for smart city development, map navigation, virtual reality, and digital twin systems. Existing 3D reconstruction algorithms predominantly rely on feature matching techniques and demonstrate satisfactory performance in small-scale or structurally simple scenes. Due to the intricate terrain features and noise interference in complex or large-scale environments, there are significant challenges such as suboptimal reconstruction accuracy and incomplete modeling, which hinder the effectiveness of these methods. To address the issues of low feature matching accuracy, high noise in predicted depth maps, and incomplete edge reconstruction in multi-view stereo for remote sensing images, this paper proposes a diffusion-constrained multi-view stereo network comprising a Multi-scale Feature Enhancement Feature Pyramid Network (MFE-FPN), an Adaptive Feature Aggregation module (AFA), and a Diffusion Constrained Module (DCM). The proposed method consists of several steps. First, the network takes N multi-view remote sensing images as input, with the first image serving as the reference and the remaining N-1 as source images. It adopts a three-stage coarse-to-fine strategy to progressively predict depth maps. The network utilizes the MFE-FPN module to extract multi-scale features from the input images, thereby generating hierarchical feature representations. Secondly, the top-level features from the FPN are mapped through an edge-aware network to compute edge-aware features, which are subsequently fused with the multi-scale features. Next, an Adaptive Feature Aggregation module is designed to aggregate the multi-scale features, forming a matching cost volume. Subsequently, a Diffusion Constraint Module is introduced to integrate cost volume features with edge-aware features. Following this, an Edge-guided Transformer (EGT) is employed to enhance the representation of edge details during the denoising stage. Finally, the cost volume features are regularized and regressed to depth, resulting in the final reconstructed depth map. Additionally, an edge-aware loss function is constructed during training to better preserve edge information in the predicted depth maps. Experimental results on the WHU-TLC and LuoJia-MVS datasets show that the MAE (Mean Absolute Error) metric of the DiffusionMVS network is improved by 28.11% and 3.37%, respectively, compared to other methods, demonstrating superior reconstruction performance. However, in terms of inference time, the proposed method does not achieve the best performance due to the relatively low operational efficiency of the diffusion constraint module. Nevertheless, it achieves a better balance between accuracy and efficiency, making it more suitable for remote sensing stereo reconstruction tasks. On the self constructed dataset of oil and gas stations, the results verify the model’s capability for more detailed geometric feature reconstruction, which benefits from its excellent performance in edge preservation and generalization in unseen scenarios. Moreover, ablation experiment results confirm that the proposed MFE FPN, AFA, and DCM modules can effectively enhance the accuracy of depth map reconstruction.
