Mosaicking multiple remote sensing images is an effective approach to rapidly expand the coverage of remote sensing imagery, and an indispensable process in large-area mapping using remote sensing imagery. Differences in sensor properties, imaging conditions, and processing algorithms cause radiometric variations, requiring consistency correction during mosaicking. However, most existing methods employ linear correction models based on statistical information between images, which are often inadequate for addressing nonlinear discrepancies caused by the above differences. This study proposes a nonlinear correction method for achieving radiometric consistency in multiple remote sensing images, based on global optimization. A quadratic local constraint model, combined with a global mean-variance preservation constraint, ensures smooth transitions and prevents color shifts. Validation with multitemporal high-resolution images shows the method outperforms Wallis filtering, Histogram matching, linear correction, and trust-region reflective in handling nonlinear radiance differences. The mosaic products exhibit natural visual quality and smooth color transitions, achieving color balance and avoiding excessive radiometric correction. The method reduces the mean of absolute mean differences and the mean of absolute standard deviation differences by 30.5% and 44.5% versus uncorrected images, and by 17.5% and 6.1% versus linear correction. The color distance drops 81.2%, while the gradient loss remains low (0.64), preserving spectral fidelity and structural details. In terms of global radiometric stability, the method, maintaining the overall mean value above 93.4% of all the uncorrected images, an improvement of 10.3% compared to the linear method, achieving effective radiometric balance and consistency, and effectively enhancing the mapping quality of large-area remote sensing imagery.
Hyperspectral remote sensing satellites are an important tool in Earth observation technology, with hyperspectral remote sensing images (HSIs) providing rich spectral information of the Earth's surface. However, due to various degradation factors in the imaging chain, HSIs are highly susceptible to stripe noise interference, which significantly degrades image quality. More troublingly, the stripe noise is easily mixed with the edges of the ground objects and background, making it difficult to effectively separate them. In addition, existing methods have not sufficiently integrated frequency-domain and spectral information for stripe noise removal, leading to an inability to completely remove the noise. Moreover, the issue of preserving the original image structure and spectral information during the denoising process has been overlooked. In response to these challenges, this article proposes the Kolmogorov-Arnold dynamic gating network with dual-domain multiscale feature aggregation for hyperspectral stripe noise removal ($\text{KAN-D}{2}\text{Net}$). $\text{KAN-D}{2}\text{Net}$ models stripe noise in the frequency domain to capture its anisotropic characteristics, integrates multiscale spatial features to enhance expression capability, and uses a learnable graph structure to focus on key channels, improving spectral recovery. In addition, it employs a KAN-based dynamic gating mechanism with a learnable activation function to adapt to complex data and boost nonlinear feature expression. Comparison with state-of-the-art methods shows that $\text{KAN-D}{2}\text{Net}$ effectively suppresses stripe noise and preserves original information, outperforming several existing techniques in both simulated and real hyperspectral data experiments.
Drone small-object detection remains challenging due to extreme scale variation, dense spatial distribution, and frequent occlusion in aerial images. Small objects often occupy very limited pixel regions and exhibit ambiguous structural boundaries, making their discriminative features easily overwhelmed by background clutter. These challenges lead to unstable multiscale representation and degraded localization precision in existing detectors. To address these challenges, combining Mamba and YOLO, we propose the HEdge-MamYOLO, which consists of three components. First, we design a frequency-Mamba collaborative high-frequency enhancement module (FM-CHFEM) to extract high-frequency components of enhanced visible edges from occluded objects in the frequency domain and leverage Mamba global selective scanning combined with spatial similarity to unoccluded objects of the same type to achieve feature enhancement and recognition of occluded objects. Second, we propose the dynamic scale feature fusion module (DSFFM), by covering large-scale objects and adaptively matches small-scale object features, achieving balanced capture of objects at different scales under FM-CHFEM-based feature extraction. Finally, we design the lightweight low-level feature fusion head (LLFFH) by integrating high-resolution low-level features with DSFFM's scale-adaptive features, which sparsely processes feature maps via partial convolution (PConv). This approach preserves details of small objects and reduces computational overhead. Extensive experiments on the VisDrone2019 and UAVDT aerial image datasets, with mAP50 metrics reaching 52.5% and 33.4%, respectively, which is better than other state-of-the-art (SOTA) models, demonstrate the effectiveness of our proposed method.
In satellite remote sensing, disparity estimation is a core technology for three-dimensional reconstruction. However, satellite stereo image disparity estimation faces two primary challenges: 1) The limitations in the quality and scale of available datasets. 2) The persistent difficulty for current methods to simultaneously achieve both fine-grained small-scale and robust large-scale disparity estimation. To address these issues, this study builds a satellite stereo image disparity estimation dataset (WHU-SSIDE) using GaoFen-7 stereo images and open-source LiDAR data through a meticulously designed technical pipeline. The dataset comprises 3,737 data pairs, each containing left-view and right-view images, high-quality disparity ground truth, and corresponding metadata. The dataset construction pipeline incorporates a multi-level geometry correction mechanism that ensures sub-pixel accuracy (MAE < 1 pixel) for disparity ground truth. Additionally, a key advantage of this dataset lies in its presence of high-rise building areas, which introduces a number of large-disparity regions. As a result, the dataset establishes a highly reliable and challenging benchmark for satellite stereo disparity estimation. Furthermore, we propose an end-to-end disparity estimation network called metadata-informed multi-range geometric encoding network (Meta-MRGE). This network is a hybrid architecture that integrates multi-range geometry-aware cost volume with an iterative optimization mechanism. Moreover, Meta-MRGE embeds satellite imaging parameters to guide the feature extraction and matching process. These modules enable the network to more effectively capture features and achieve joint optimization for multi-range disparity estimation. As a result, it improved the accuracy of both large-scale and small-scale disparity estimation in various remote sensing scenarios. Experimental results demonstrate that Meta-MRGE outperforms state-of-the-art methods on both the WHU-SSIDE and WHU-Stereo datasets. On WHU-SSIDE, it achieves the lowest end-point error (EPE) of 1.2485 and the lowest root mean square error (RMSE) of 4.0683. On WHU-Stereo, it surpasses the second-best method by 2.50% in D1 and 0.1725 in EPE. Notably, the method shows significant improvements in large-disparity regions. Ablation studies validate the effectiveness of the individual technical components within Meta-MRGE. Furthermore, generalization and metadata sensitivity tests further confirmed the robustness of Meta-MRGE to input data. This study provides a foundational solution for satellite stereo disparity estimation, particularly in challenging regions. The dataset and code are publicly available at https://github.com/zhanggb1997/WHU-SSIDE.
As the cost of small satellites decreases annually and their performance improves, hundreds of high-resolution satellites are now in orbit. This has increased the demand for satellite stereo processing toward detailed observations of the Earth. However, satellites are limited by their orbit, making it difficult to observe all surfaces of a target. To address this, multiple satellites can be used for collaborative observation. However, imaging quality and timing vary across different satellites, making multi-source and multi-temporal stereo data processing a key research focus. To advance this area, this study developed a high-precision digital surface model (DSM) and satellite images to create a multi-source and multi-temporal satellite stereo matching dataset, verifying the accuracy of the dataset. The dataset demonstrated row and disparity accuracies better than one pixel for the epipolar images. Consequently, our dataset supports algorithm testing and model training, enhancing multi-satellite collaborative observations.
This article proposes an unsupervised knowledge distillation (KD) framework for satellite multiview stereo (MVS) reconstruction under label-free settings. A teacher-student paradigm is adopted, where a teacher MVS network is first trained using self-supervised multiview geometric constraints, including photometric consistency, feature similarity, smoothness, and structural similarity. The teacher produces multiview height predictions, and geometry-aware uncertainty is inferred from cross-view consistency during multiview fusion. Based on the fused height estimates and uncertainty cues, reliable pseudo-labels are constructed in the form of expectation-based height maps, pixel-wise uncertainty maps, and visibility masks. These pseudo-labels supervise a lightweight student network through probability-volume distillation, combined with uncertainty-weighted height regression and cross-view geometric consistency constraints. Extensive experiments on representative satellite benchmarks demonstrate that the proposed framework achieves competitive accuracy among unsupervised methods while significantly reducing computational cost and eliminating the need for LiDAR ground truth or pretrained models, providing a scalable solution for large-scale satellite 3-D reconstruction.
Inconsistent radiometric response characteristics of nighttime sensor detectors can interfere with weak nighttime surface radiation, resulting in radiometric inconsistencies in captured images. This degrades image quality and compromises information extraction from such image products. Relative radiometric calibration involves calibrating response differences among imaging detectors to eliminate radiometric inconsistencies, which is particularly critical for nighttime sensors. However, existing on-orbit calibration methods require onboard calibration devices, reliable attitude control, and dedicated calibration imaging tasks. The limited hardware capabilities and attitude control capacity of micro nighttime satellites pose significant challenges for radiometric calibration. This study proposes a deep learning-based on-orbit relative radiometric calibration method for nighttime sensors to address these limitations, representing an effort to apply deep learning techniques to relative radiometric calibration of remote sensing satellites. A Multi-Gain Radiometric Response Transfer Network (MGRT-Net) was developed to characterize the radiometric response relationships of nighttime sensors across gain settings. A daytime low-gain radiometric calibration reference was established using a uniform calibration method. The calibration reference is transferred to nighttime images to calibrate the radiometric parameters of the nighttime sensor in high-gain mode. Using Luojia1-01 nighttime images, MGRT-Net accurately simulated radiometric responses across different gain settings and effectively transferred daytime low-gain radiometric calibration references to nighttime images. After MGRT-Net-based transfer correction, radiometric response errors among sensor detectors in nighttime high-gain images were effectively eliminated. Under low- and high-brightness, the streaking metrics (SMs) exceeded 0.027% (76.5% improvement) and 0.028% (75.6% improvement), respectively, outperforming existing methods. Compared with analytical calibration methods, the proposed method requires no prior model or additional satellite requirements and demonstrates strong adaptability. This study provides a reference for micro-remote sensing satellites constrained by manufacturing costs and validates the applicability of deep learning methods for relative radiometric calibration.
Relative radiometric calibration is essential for improving image quality by adjusting the radiometric response model of each detector and eliminating detector-level systematic errors arising from variations in radiometric characteristics. Planar array sensors in remote sensing satellites now employ millions of imaging detectors - a hundredfold increase over linear push-broom sensors - creating unprecedented calibration challenges. Current on-orbit calibration methods rely on on-board equipment (on-board calibration) or large-area homogeneous surface data (uniform calibration), both requiring dedicated imaging tasks and consuming valuable orbit resources. The latter is often preferred for satellites lacking onboard calibration devices; however, as satellite resolution improves, the scarcity of sufficiently uniform calibration sites further constrain these approaches. This study develops an efficient, high-accuracy on-orbit relative radiometric calibration method for planar array sensors that overcomes these limitations. Our approach corrects geometric distortions and aligns sequential images using geometric consistency constraints. Calibration reference is derived through probabilistic statistical analysis, and detector-level calibration coefficients are refined via multi-scale error correction. We validated the method using the Jilin-1 02B satellite, a commercial optical remote sensing platform with a plane array sensor. The method achieved relative radiometric correction accuracies of 1.14% to 2.35%, across three spectral bands, with red and green band performance exceeding alternative calibration methods and blue band accuracy comparable to the uniform method. Results demonstrate the calibration coefficients ensure accurate correction for features of varying brightness across the sensor's full dynamic range. Critically, the method enables efficient, high-accuracy on-orbit calibration using any standard observation data without dedicated imaging tasks, eliminating orbit resource constraints and providing a scalable solution for current and next-generation remote sensing satellites equipped with planar array sensors.
With the rapid development of the low-altitude economy,the construction of low-altitude transportation infrastructure has progressed from conceptual exploration to scaled practice.The digital air-route network not only guides the construction of facility and air-internet networks,but also provides the service network with followable routes,making it a priority task in building low-altitude transportation infrastructure.However,the existing methods for constructing digital air-route networks insufficiently consider risk quantification,lack structured topology,and omit essential route attributes.In addition,they have not clarified the required types and geometric accuracies of geographic and constraint elements in low-altitude environments.Therefore,it is necessary to further improve relevant methodologies to better guide the construction and application of digital air-route networks for the low-altitude economy.To address these issues,this study begins with the interaction mechanism between unmanned aerial vehicles(UAVs)and their geographic constraint environments.It identifies the categories of geographic and constraint elements required for digital air-route network construction and specifies the geometric accuracy requirements for geographic elements.The feasibility and adequacy of spaceborne remote sensing techniques for acquiring these elements are analyzed.Based on these findings,a construction method for digital air-route networks is proposed,integrating geographic and constraint information while jointly optimizing topological structure and risk.Field experiments are conducted in Anyang to verify the feasibility of this method,including validation of the spaceborne geographic information base,meteorological constraints,the digital air-route network itself,and the communication and positioning quality along the routes.Results show that spaceborne remote sensing data achieve a DSM vertical accuracy better than 2m,a building white model accuracy of 3.83m,an overall obstacle recognition accuracy of 80.77%,and a land cover classification accuracy of 79.5%.These results collectively meet the meter level geometric and surface-attribute resolution requirements for digital air-route network construction.Compared with manually designed routes,the air-route network generated with this method reduces route length by 7.6%,cruise time by 12.6%,and the proportion of high-risk segments by 7.6%,while increasing the nonlinearity coefficient by 8.2%.Compared with pilot-planned ad-hoc routes,route length decreases by 4.2%,cruise time by 3.4%,and the nonlinearity coefficient improves by 18.5%.Overall,the proposed method effectively improves airspace utilization,reduces flight risk,and enhances flight efficiency,fulfilling the operational requirements for UAVs to fly,fly safely,and fly efficiently in large-scale low-altitude operations.
This article presents an innovative geometric calibration mode for a linear array camera used in the Jilin-1 satellite. Changing the camera from observing the ground to observing space and treating stars as control points allow for the accurate determination of the geometric parameters of the camera. In this article, several important calibration procedures and strategies are presented to gradually eliminate the influence of astronomical and attitude errors. Self-verification of the calibration image results in a star-based calibration accuracy is better than 0.3 pixels. Further verification of the star image that leads to an interior orientation accuracy is better than 0.6 pixels. The accuracy of geometric positioning without control is verified by ground images of different regions to be 30 m. The interior orientation accuracy of ground images processed by star-based calibration parameters is better than 1.5 m. The ground image corrected using the star-based calibration parameters is well overlaid on the reference image. The accuracies obtained with star-based and ground-based calibrations are comparable. However, ground calibration fields require high maintenance costs and may suffer from harsh imaging conditions. Contrarily, star-based calibration is not affected by the weather and climate conditions on the ground, which can be photographed from any location along the satellite orbit. Therefore, star-based calibration can meet the requirements for high-frequency geometric calibration of subsequent satellite constellations construction.
With the rapid development of hyperspectral image classification (HSIC) technology, its applications in geological exploration and environmental monitoring have become increasingly prominent. Recently, Mamba has garnered significant attention owing to its outstanding performance in long-range sequence modeling and linear computational complexity. However, Mamba still exhibits significant limitations in HSIC: first, it does not fully consider the hierarchical spatial-contextual representation and nonlinear spectral interactions in hyperspectral images; second, its sequential processing approach leads to the loss of spatial structural information and feature redundancy. In response, this study proposes a structure-enhanced spatial-spectral dynamic gating Mamba (SEDGM) that leverages the collaborative design of spatial and spectral gating Mamba mechanisms to extract and exploit key regional features of hyperspectral data. The spatial branch employs hierarchical gating Mamba (HGM) to capture multidirectional pixel sequences and extract the hierarchical spatial features and their intrinsic relationships. In contrast, the spectral branch utilizes a random shuffled gating Mamba to disrupt the fixed order of traditional spectral sequences and capture higher order spectral couplings, effectively characterizing the cooperative variation patterns of spectral features. Both branches employ a dynamic gating mechanism that weights features based on sequence centrality, dynamically activating feature sequences. Additionally, shape-specific offset-aware attention (OAA) is incorporated into each branch to enhance the structured features that were lacking in the Mamba sequences. Finally, a spectral-oriented feature review module (SOFRM) is incorporated to achieve dynamic feature fusion and optimized refinement. Experiments were conducted on four large-scale benchmark hyperspectral imaging (HSI) datasets, with SEDGM achieving significant improvements in classification performance, validating the effectiveness of this approach in HSIC tasks.
In the process of mosaicking regional synthetic aperture radar (SAR) intensity images, multiple images with significant brightness anomalies can cause a considerable number of pixels to exceed the grayscale quantization range. Applying traditional color harmonization methods increases this issue, causing a loss of brightness information. We propose a multi-objective gray consistency correction method designed explicitly for mosaicking regional SAR intensity images with brightness anomalies to address this. We constructed a two-objective optimization model to ensure regional image gray consistency and mitigate brightness information loss. The truncation values of brightness anomaly images were selected as decision variables, maximizing the overall gray consistency of overlapping image pairs and minimizing the number of pixels with grayscale values that were out of bounds as the objective functions. To synchronously solve the truncation values of brightness anomaly images and linear stretch parameters of all images, a hybrid framework that combines the non-dominated sorting genetic algorithm II (NSGA-II) with the quadratic programming (QP) algorithm was proposed. Two large-area experimental results show that the proposed method achieves a balanced optimization between gray consistency and brightness information loss for regional SAR intensity image mosaicking. Compared with the traditional method, our method reduces brightness information loss by 99.552–99.647% and 99.973–99.969%, respectively, while maintaining better peak signal-to-noise ratio performance.
Accurate disparity estimation of high-resolution satellite remote sensing stereo images serves as a crucial method for generating precise digital surface models. However, the complex intractable regions in satellite images (textureless regions, repeated texture regions, occlusion regions) pose serious challenges for accurate disparity estimation. To enhance the matching accuracy within intractable regions, a dual branch multi-scale stereo matching network for high-resolution satellite stereo images is proposed. First, a dual branch feature extraction module is designed which can perform efficient downsampling. This module can enhance the scene awareness capability of the model, enabling it to extract multi-scale feature maps and construct multi-scale cost volumes. Then, the cost aggregation process is executed in a coarse-to-fine manner. The method employs a simple hourglass structure and leverages low-scale information to guide the aggregation of high-scale cost volumes. Next, a disparity-channel attention mechanism is proposed for the cost aggregation process to obtain more representative feature information. Finally, a simple disparity refinement module is designed by utilizing both intensity and gradient information of the left image to improve the local details of the disparity map. Experiments are performed separately on the GaoFen-7 and US3D datasets. The experimental results indicate that the proposed method is conducive to improving the matching accuracy within intractable regions of satellite images. The structure of the proposed network is simple, which can effectively reduce the network parameters and realize the lightweight of the model.
While systematic radiometric errors can be mitigated through on-orbit calibration, correcting low-frequency, time-varying errors often induced by complex imaging dynamics remains a greater challenge. This study introduces a correction method tailored for optical remote sensing platforms to overcome such errors in JiTian-03 imagery, particularly those arising from high-agility curve imaging. The process begins with relative radiometric calibration to correct high-frequency inconsistencies among sensor detectors, thereby suppressing striping and banding noise in the imagery. Next, the image is decomposed into low-frequency background and high-frequency detail components. A column-wise moment-matching algorithm is then applied to the low-frequency background, aligning each column’s first- and second-order central moments with global statistics to derive correction coefficients. This approach effectively compensates for low-frequency radiometric errors while preserving high-frequency texture details in the image. Experimental results confirm that, following correction of both high- and low-frequency radiometric errors, JT-03 imagery exhibits a substantial reduction in striping, banding, and brightness nonuniformity. The correction yields a uniform radiometric response across the entire field of view while preserving ground texture fidelity and image detail. As a result, the radiometric uniformity of JT-03 image products is significantly enhanced. Quantitatively, post high-frequency calibration, the striping index across all JT-03 bands improved to within 0.1% and the standard deviation of column means to 4.7%. Subsequent low-frequency correction reduced the standard deviation of column means by an average of 73.89%, achieving a final value of better than 1.3%.
Hyperspectral image (HSI) classification is an essential technique in hyperspectral remote sensing applications, and Transformers have shown great potential in this field. However, factors such as the distribution of land covers, spectral variability, and high dimensionality often cause many Transformers to lose information related to inter-class heterogeneity and boundary details. To address these issues, we propose a dual-branch Transformer network that combines sparsification and high-low frequency interaction, aiming to address challenges such as diverse surface cover and spectral feature mining with lower computational cost. The sparse Transformer branch is used to extract multi-scale local-global dependency features from HSIs, preserving finegrained and wide-range contextual information. The high-low frequency interaction Transformer branch is combined to reveal the spectral response differences of objects indifferent spectral bands. Deep interaction between the two parallel branches is achieved by perceiving channel correlations and spatial positions. Extensive experiments on four public HSI datasets demonstrate the effectiveness and competitiveness of the proposed method, and the code is made publicly available at https://github.com/youngboy03/MSFI-CNet.
Deep learning-based disparity estimation methods have demonstrated significant potential in optical satellite stereo image applications. However, learning-based methods remain susceptible to domain shifts caused by spatiotemporal variations and stereo-sensor heterogeneity. To address these challenges, we propose a hierarchical domain adaptation disparity estimation (HDADE) framework for optical satellite stereo images. HDADE was structured with a four-stage technique pipeline to improve the training data quality and diversity, explicitly align the spectral and stereo distribution, implicitly enhance the robustness of feature extraction and matching, and directly facilitate feature alignment with the target domain. This hierarchical framework systematically mitigates disparity estimation accuracy degradation in cross-domain scenarios. Cross-spatiotemporal and cross-payload generalization experiments were conducted based on the WHU_Stereo and US3D datasets. The experimental results show that HDADE significantly outperformed other advanced methods and possessed plug-and-play versatility. Notably, greater domain shift scene transfer experiments indicated that, with limited annotation data, HDADE has the potential for large-scale automatic applications.
Building height significantly influences urban development and evolution. Previous studies on building height estimation using digital surface models (DSMs) have predominantly addressed simple, single-environmental scenarios, often yielding unsatisfactory results across diverse environments. This study introduces a novel method for estimating building height by integrating scene classification with spatial geometric relationships. Initially, raw data are processed to derive the various data types required for this approach. Environmental scene classification, based on vegetation and shadows analysis, is then performed. Subsequently, the building height is estimated either directly from the DSM or through road height prediction. The proposed method is validated using a scene image from Wuhan, Hubei Province, China. The results demonstrate that the estimated building height maintains high accuracy in complex environments with significant vegetation and shadow coverage, achieving a mean absolute error of 1.84 m. Furthermore, the proposed method outperforms existing DSM-based techniques. This approach is adaptable for high-precision building height estimation across various environments and holds substantial application potential, facilitating further research in urban-related scenarios.
Reconstructing missing information due to cloud occlusion is an effective means of enhancing the utilization of low-, medium-, and high-resolution optical remote sensing images. However, singletemporal-based methods have limitations regarding the demand for cloud-free reference data and the applicability of specific datadriven models to real-world scenarios. It is more unable to realize mutitemporal reconstruction. To address this, we propose the Dual-Decoupling Inter-correction Multitemporal Reconstruction network (DDIM-RecNet), a unified framework designed for single- and multitemporal cloud occlusion reconstruction of low-, medium-, and high-resolution images. DDIM-RecNet innovatively decouples remote sensing images into ground object and imaging environment components using dedicated inter-correction modules, coupled with targeted loss functions. Additionally, an imaging environment enhancement module ensures spatial consistency between reconstructed and original regions. Compared with classical models, such as U-Net, RFR-Net, STGAN, PSTCR, BSN, GLDF-RecNet, and IDF-CR, DDIM-RecNet achieved excellent visual reconstruction results and the best quantitative evaluation indicators under Gaofen-1 (2 m), Sentinel-2 (10 m), Landsat-8 (30 m) single/multitemporal images. Taking Gaofen-1 (2 m) as an example, compared with the suboptimal model, the clarity of the DDIM-RecNet model in the three bands was improved by 0.44, 0.70, and 0.85 respectively under singletemporal reconstruction; the clarity of DDIM-RecNet was improved by 0.55, 0.43, and 0.35 respectively under mutitemporal cloud occlusion.
Accurate height estimation from satellite imagery is essential for digital surface model (DSM) generation in remote sensing applications. Although recent deep learning-based multiview stereo (MVS) methods have advanced height prediction, they often struggle with edge degradation, particularly around building outlines and terrain discontinuities. To address this issue, we propose SP-MVS, a structure-preserving MVS network that explicitly enhances edge-aware height estimation. The network integrates a triple-branch structural encoder equipped with a U-Net backbone and the convolutional block attention module to capture multiscale spatial and texture features. A Mamba-based cost volume regularization module is introduced to model long-range cross-view dependencies with low computational overhead. Furthermore, we design a multistage edge-texture consistency loss to guide the network in preserving sharp structural boundaries throughout the estimation process. SP-MVS directly produces high-quality height maps and serves as a critical component in precise DSM generation. Extensive experiments on the WHU-TLC and US3D datasets demonstrate that SP-MVS achieves superior accuracy and sharper boundary delineation compared to state-of-the-art methods.