
Building height is an essential variable for describing the urban vertical landscape and evaluating urban resilience to natural disasters and climate change adaptation. Building-scale height maps are key to quantifying urban structures across multiple relevant scales (e.g., building, block, and landscape scales). Although some studies have filled in missing building height data using various methods, the scarcity of temporally consistent and spatially balanced building height samples continues to limit mapping accuracy. This is particularly true in the Global South, where urban vertical data remain scarce. In this study, we developed an easy-to-use mapping framework to produce building-scale heights for global major cities in 2020 by combining active and passive remotely sensed data. Specifically, we proposed a novel approach for generating building-scale height samples through integrating Global Ecosystem Dynamics Investigation (GEDI) data and building footprints. Combining spatial-spectral characteristics effectively mitigated the impact of positional offsets in GEDI footprints during sample construction. Twelve cities at different stages of development were selected as the study area. Validation results indicated good agreement between derived samples and reference data (r = 0.87; RMSE = 5.78 m). Random forest models calibrated with local samples, neighboring high-value samples, and variables derived from building footprints as well as Sentinel-1 and Sentinel-2 data produced accurate building height maps (r = 0.81; RMSE = 11.48 m). Our framework outperforms existing products in accuracy stability, alleviating the global scarcity of building-scale height samples. This contributes to building-scale height data availability and reliability.
For Mars landing site selection, it is of great engineering significance to identify landing ellipses that not only satisfy the geometric constraints of entry, descent, and landing (EDL) but also offer high scientific value. In this study, we develop an optimization framework that transforms landing ellipse identification from predefined discrete candidate evaluation into a suitability-map-driven spatial optimization problem. Within this framework, a Particle Swarm Optimization-based Search (PSOS) strategy is introduced as a representative optimization strategy to identify high-value landing ellipses in a continuous suitability map. The performance of PSOS is evaluated against four discrete search strategies: Threshold-based Search (TS), Moving Ellipse Search (MES), Global Random Search (GRS), and High-value Local Search (HLS). The results show that the PSOS can effectively identify high-suitability landing ellipses (suitability > 0.77) with solution consistency comparable to MES (Jaccard index ≈ 0.6), while exhibiting lower variability across independent runs than stochastic search methods such as GRS and HLS. Furthermore, the optimization effort of PSOS can be adjusted through algorithm parameters, enabling a flexible trade-off between solution quality and computational cost when exact global optimality is not required. The proposed framework provides a complementary approach to exhaustive search for automated landing ellipse identification and supports iterative mission analysis for future Mars landing site selection.
Over the last six decades, Mars exploration missions (both successful and unsuccessful) have inevitably altered the Martian surface and environment. Consequently, accurately detecting non-salient anthropogenic changes resulting from these exploration activities plays an essential role in environmental evaluation. However, precise identification of changed areas and their semantic information on Mars is challenging due to interference from natural changes and complex terrain. To address this, we propose an efficient contrast-guided hierarchical unsupervised method, i.e., CGHU-CD, for sample-free Martian anthropogenic change detection (MACD) using high-resolution orbital images. The proposed approach first performs a global patch-level detection, employing feature description and adaptive thresholding, to maximize the removal of unchanged regions, yielding coarse-grained change candidate pairs. A local pixel-level detection method then exploits the contrastive features of anthropogenic changes, including brightness and texture differences relative to the surrounding Martian background, to progressively enhance change information and refine the detection results. Furthermore, an automatic classification method incorporates inherent object brittleness and size information to distinguish different change types. To evaluate the effectiveness of the designed approach, we constructed a MACD dataset collected from High-Resolution Imaging Science Experiment orbiter images, which include four Mars exploration missions. Extensive experiments comparing our CGHU-CD with seven advanced techniques demonstrate its superior capability in detecting non-salient anthropogenic changes of diverse sizes, shapes, and appearances across varied Mars scenarios. It exceeded these existing unsupervised methods by 2.08% to 78.43% in the F1-score measure. The CGHU-CD also achieved superior classification performance to existing techniques, with an increase of 2.77% to 23.28% in the Kappa value. At last, the sample-free CGHU-CD method exhibited superior potential for detecting anthropogenic changes at the Zhurong landing site across diverse sensors and spatial resolutions.
Earth observation increasingly combines satellite and ground systems that record different manifestations of the same phenomenon. Spatial clustering and temporal persistence inflate apparent agreement, creating uncertainty over where joint products are statistically supported. Marine lightning is a stringent case because offshore energy, shipping, and coastal infrastructure depend on geostationary optical imagers and coastal low-frequency networks. We develop a four-state factorial null framework that evaluates FY-4A Lightning Mapping Imager–ground-network correspondence under observed, spatially randomized, temporally displaced, and jointly randomized states across four Chinese marginal seas and 215 coastal VLF/LF stations. Spatial and temporal storm backgrounds interact non-additively, contributing 11.5–32.8 percentage points of correspondence after block-bootstrap propagation of spatial autocorrelation. Combining this interaction — rather than any single bias-corrected match rate — with distance support and multi-year stability yields joint summaries in the Bohai and Yellow seas, local calibration in the East China Sea, and parallel optical and radio evidence in the northern South China Sea. These classes remain stable across 20 matching criteria and six summers. Independent datasets over six non-Chinese coasts reproduce the factorial structure, while calendar-date randomization reduces the interaction towards zero. Controlled event loss and location error contract fusion support, and timing displacement adds 1.6–2.7 percentage points of apparent matches. The framework defines a transferable decision sequence for satellite–ground observing pairs with different observation entities, common support, countable correspondence, and clustered persistence. This study demonstrates the sequence for marine lightning; matching windows, support limits, coefficients, and application classes remain locally estimated.
As a virtual component of the geospatial digital twin (GDT), the virtual geographic environment (VGE) has significant advantages in enhancing human geographic cognition and analyzing and solving practical engineering problems. However, existing virtual geographic scene modeling methods still have significant gaps compared to GDT, making it difficult to balance the ability to represent dynamic objects as required by GDT while preserving the geographic characteristics of VGE, such as spatiotemporal reference and multiscale representation. Against this backdrop, this paper takes highway tunnel scenes as an example. It proposes a modeling framework that leverages the synergy between perception data and scene knowledge to construct a virtual geographic scene for GDT. First, we constructed a multi-domain integrated knowledge graph of highway tunnel scenes to enhance the standardization and completeness of the modeling process. Second, we established a dynamic object detection network based on an improved YOLOv8 model, endowing the virtual geographic scene with the capability to represent dynamic objects. Third, we proposed a tunnel scene twin modeling method that integrates a knowledge graph with a perception network, enabling effective fusion of the static base and dynamic objects. Finally, we selected a tunnel in China as the study area, developed a prototype system, and conducted experimental analyses. The results demonstrate that the proposed method successfully constructs a GDT-oriented virtual geographic scene for highway tunnels. For the city-level model, over 93.6 % of points have an error of less than 0.1 m; for the component-level model, over 95.4 % of points have an error of less than 0.01 m; the model’s semantic accuracy reaches 95 %; the average modeling time for dynamic targets is 11.8 ms; and the average rendering efficiency of the virtual scene is 133.2 FPS. Furthermore, compared to existing scene modeling methods, the constructed scene integrates advantages such as geospatial and temporal references, multiscale representation, and dynamic object representation. Research results provide an example reference for intelligent management of urban traffic, and provide new ideas for virtual geographical scene modeling.
The use of satellite observations to detect and quantify methane emissions from localized sources is constantly increasing. The comparison between results generated by different sensors aboard satellites can be challenging due to differences in spatial and spectral resolution, retrieval approaches, and acquisition timing. In this study, we evaluate the capability of the PRISMA hyperspectral satellite to detect methane plumes and reproduce plume structures, benchmarking its methane detection and estimates against corresponding data provided by the GHGSat constellation. Geolocation refinement is applied to PRISMA imagery, followed by methane retrievals using the MAG1C algorithm. Results show that PRISMA detects almost 60% of the emission sources identified by GHGSat and reproduces the main spatial trends of methane enhancements along the plume. The analysis along the plume centerline highlights consistent patterns of higher methane concentrations near the source, followed by gradual dilution along the plume tail. A multi-sensor comparison, including EnMAP, further illustrates the temporal evolution of a plume by showing the observed methane flux trend over a two hour interval. Overall, the correlation between PRISMA and GHGSat flux estimates showed a Pearson coefficient of 0.95, although PRISMA systematically seems to reveal lower methane enhancements with respect to GHGSat. The results obtained in this work confirm that hyperspectral satellite observations can provide a consistent representation of methane plume dynamics across different sensors, looking forward robust cross-sensor analysis of emission sources.
The task of referring remote sensing image segmentation (RRSIS) aims to predict segmentation masks of target objects based on text descriptions. Effective alignment between textual and visual modalities is crucial for achieving high performance. However, most existing RRSIS methods overlook the inherent domain gap between the two modalities and naively fuse visual and text features. Moreover, they lack guidance from domain-specific priors during alignment, often leading to inaccurate segmentation results. In this paper, we propose a novel method for RRSIS, namely TIANet. This method combines implicit and explicit alignment approaches to achieve fine-grained image-text alignment. Furthermore, to compensate for the lack of domain priors during alignment, we leverage the Remote Sensing Vision Foundation Model (RS-VFM) to guide both alignment strategies. On one hand, the Prior-Guided Visual-Text Feature Alignment Module (PVTAM) leverages domain priors to fuse textual and visual features in a stage-wise manner, thereby achieving cross-modal implicit alignment at the feature level. On the other hand, the Dynamic Bidirectional Alignment and Prior Distillation (DBAPD) loss is adopted to implement explicit cross-modal alignment during training. Furthermore, to enhance the perception ability of small targets, the Kernel-Adaptive Small Target Enhancement Module (KASTEM) is used in the decoder stage. Extensive experiments on the RefSegRS and RRSIS-D datasets show that our TIANet achieves excellent performance compared with other advanced methods, with a 1.22% mIoU improvement over existing state-of-the-art approaches on the RefSegRS dataset.
Spatially continuous three-dimensional tropospheric delay fields are essential for Earth observation and atmospheric sensing. The ERA5 reanalysis offers continuous coverage but is delayed by several days, whereas Global Navigation Satellite System (GNSS) observations are accurate and near-real-time but pointwise. This study develops Tropoformer-S, an iTransformer-based framework that reconstructs the regional delay field by integrating the two sources. The framework forecasts the ERA5-derived level-wise Zenith Tropospheric Delay (ZTD) at 15 pressure levels from 1000 to 550 hPa for 1 to 120 h, where the variate-token attention explicitly models the spatial correlation among grid points. The forecast field is converted into station ZTD through a physically guided retrieval chain and refined by GNSS-informed corrections along two complementary paths, harmonic self-correction for stations with local GNSS records and spatial reconstruction of the bias field for stations without. Over mainland China, the forecasts attain a level-mean RMSE of 2.84 cm against the ERA5-derived ZTD and outperform five alternative deep-learning architectures. The overall RMSE is 3.20 cm in 2022 and 3.23 cm in 2023 against GNSS, on average 18.64 % and 33.10 % below GPT3 and HGPT2, and the radiosonde evaluation shows a near-zero bias. The harmonic correction path lowers the RMSE to 2.79 and 2.98 cm and reduces the network bias to near zero, while the spatial correction path reaches 2.10 to 2.66 cm at withheld stations under nine network densities and largely removes the lead dependence. The framework thus fills the ERA5 latency gap with continuous tropospheric delay information for geodetic and atmospheric applications.
Lakes across the Tibetan Plateau exhibit pronounced responses to climatic variability and glacier fluctuations. However, the compound effects of glacier and climate forcing on lake dynamics remain insufficiently quantified, owing to limited observational continuity and reliance on single-factor qualitative or regression-based methods that are not well suited to representing multivariate nonlinear relationships. In this study, a Glacier–Climate–Lake Coupled Probabilistic Framework (GCL-CPF) was developed by integrating remote sensing–based reconstruction of glacier–lake dynamics with glacier process modeling and Vine Copula theory to characterize multivariate dependence. The framework quantifies lake expansion probability under bivariate coupling among glacier melt runoff (GMR), temperature (Temp), and precipitation (Prec). Results show lake expansion is not a monotonic response to coupled forcing but exhibits probability maxima under specific meltwater conditions, indicating a regime-dependent rather than continuously increasing response. Meltwater–precipitation combinations are associated with higher probabilities of lake expansion when favorable joint conditions occur. Temperature conditions correspond to either lower or higher lake expansion probabilities depending on meltwater–precipitation conditions, indicating a context-dependent relationship. In June, lakes generally exhibit higher probabilities of contraction characterized by limited runoff and precipitation, while occasional expansion corresponds to conditions with higher temperature and precipitation. In July and August, increased precipitation and rising temperature-induced glaciers melt correspond to higher probabilities of rapid lake expansion, while extreme heat conditions in some years are associated with pronounced lake shrinkage. This study captures compound relationships and dependence structures among multiple drivers. The GCL-CPF framework provides a transferable basis for quantifying lake dynamics in glacierized regions and predicting hydrological responses under climate change.
Precise crop mapping via remote sensing is critical for the rational utilization of cultivated land resources. Remote sensing techniques and deep learning-based semantic segmentation methods provide effective approaches to large-scale crop mapping. However, due to variations among remote sensing imaging systems, significant distributional shifts exist across different remote sensing datasets, which limit the generalizability of deep learning models. To address the issue, we propose a semantic segmentation network named PLGCA-SAM for crop mapping in remote sensing imagery. PLGCA-SAM integrates the contextual perception capability of PLGCA with the powerful generalization of SAM. Specifically, we introduce a task-adaptive encoding (TAE) module, which includes a SAM adapter to align SAM’s generic visual features with the PLGCA feature space, and a task-guided mechanism that refines the fused representation using a task-aware loss. This design enables more effective feature integration and improves adaptability across diverse remote sensing datasets. Experiments conducted on GF-2 and Sentinel-2 datasets demonstrate that PLGCA-SAM achieves 2.03% and 1.50% improvements in mIoU, respectively, over SOTA methods. Our code will be obtained by: https://github.com/Hanhlab/PLGCA-SAM.git.
Moisture-transport anomalies are closely associated with meteorological drought, but the distinct and compound roles of source-region supply and pathway variability remain insufficiently quantified. Using CMFD meteorological data and FNL-driven FLEXPART simulations for 2005–2024, this study developed a multi-index conditional-probability framework for the middle and lower reaches of the Yangtze River (MLYR). Three standardized 10-day indices represented source-region precipitation contribution (SPC), moisture pathway intensity (MPI), and moisture pathway frequency (MPF). Gaussian copulas were used to estimate conditional drought probabilities under individual and paired moisture anomalies. Climatologically, Pathways 2 and 3 accounted for 45.4% of total pathway transport, while the Northwest Pacific and southern China–Indochina Peninsula together contributed 35.9% of source moisture. However, negative anomalies in the South China Sea and Pathway 2 were associated with higher individual drought probabilities than those in the Northwest Pacific and Pathway 3, despite their smaller climatological shares. Across 210 pair-matched configurations, compound conditional drought probability exceeded the stronger constituent single-anomaly probability in 181 cases (86.2%), with a median increase of 7.55 percentage points. After calendar-year-block permutation and BH-FDR correction, ΔP was significantly greater than zero in 164 pairs, and all nine top-ranked combinations exceeded 50% conditional probability. These results show that mean contribution magnitude was not a reliable indicator of anomaly-conditioned drought association and that elevated drought probability is frequently, but not universally, associated with concurrent weakening across moisture-transport components. The framework provides a process-oriented basis for diagnosing drought-relevant transport configurations and can be adapted to other basins after region-specific recalibration and independent validation.
Urban flood (UF) poses major challenges to sustainable development and public safety. Geographic information-based studies commonly apply machine learning (ML) with UF samples and multi-domain factors to assess urban flood susceptibility (UFS) and identify priority regions for UF management. However, most existing approaches rely solely on the latest UFS map, overlooking the dynamic nature of UF and thereby reducing the reliability of the identified priority areas. To address this limitation, we propose an ML-based framework for spatiotemporal UFS estimation to delineate more accurate and practical UF priority regions. The framework integrates: (1) Natural Language Processing to extract flood locations from social media (Weibo), and (2) a Spatiotemporal Susceptibility Extraction (SSE) method that combines the spatial and temporal characteristics of UFS. Guangzhou, a megacity in China with frequent UF events and extensive Weibo usage, is employed as the case study. We found: (1) incorporating UF locations extracted from social media effectively complements official records, as evidenced by both the increased information content of the positive class across multiple flood-conditioning factors and the improved classifier performance and robustness; and (2) compared with priority areas directly delineated from the latest-year UFS map, the SSE-derived UF priority areas show higher accuracy when evaluated against ground-truth UF data, further supported by high-resolution remote sensing imagery, while also exhibiting more spatially concentrated and continuous patterns. The proposed framework enables rapid and reliable identification of UF priority areas and provides practical decision support for UF prevention and management in rapidly urbanizing megacities under resource constraints.
Traditional regional landslide hazard assessments primarily focus on static susceptibility, often overlooking critical post-failure kinematic characteristics and their destructive potential. To address this gap, this study proposes a landslide intensity assessment framework integrating both landslide size and mobility characteristics, using the Ms6.8 Luding earthquake (2022) as a representative case. We developed intensity indices to reflect destructive potential across three dimensions: cumulative effects of all landslides within a slope unit, average landslide intensity, and maximum single-event intensity. An interpretable log-Gaussian Generalized Additive Model was employed for regional modeling, while the synergistic effects of potential controlling factors were systematically quantified. Results demonstrate that the proposed indices effectively capture multiple dimensions of destructive potential of regional landslide intensity. Under spatial cross-validation, the models showed satisfactory predictive performance and robustness, with predictions closely matching observations. Notably, the total effect model achieved Pearson’s correlation coefficient (R) of 0.809, Nash-Sutcliffe efficiency coefficient (NSE) of 0.648. Particularly, this study is the first to explicitly quantify nonlinear interactions among key controls (e.g., slope, seismic motion, fault), revealing their important contributions to landslide intensity assessment. By expanding the paradigm from occurrence probability to destructive intensity modeling, this study further complements and refines the existing landslide hazard assessment framework. The proposed framework provides a robust, interpretable tool for identifying high-risk zones, facilitating more effective disaster mitigation and emergency planning in seismically active regions.
Hyperspectral and multispectral image fusion (HMIF) aims to reconstruct high spatial resolution hyperspectral images by jointly leveraging the complementary spectral information from hyperspectral imagery and the spatial details from high spatial resolution multispectral imagery. Existing deep learning-based methods typically improve fusion performance by enlarging the receptive field to capture richer spatial–spectral contextual information, which inevitably increase computational complexity and may introduce redundant spatial–spectral interactions. To address these issues, this study proposes a new multiscale Spatial-Conditioned Spectral Routing Network (SCSRNet). Specifically, a Dynamic Fusion Gate (DFG) module is first developed to adaptively regulate the contributions of hyperspectral and multispectral features based on local spatial characteristics. Subsequently, a Large-Kernel Spatial–Spectral Residual Block (LK-SSRB) is designed, in which a Multi-Scale Spatial Perception Module (MSPM) captures spatial contextual information with large receptive fields, while a Spatial-Conditioned Spectral Routing Module (SCSRM) is innovatively designed to dynamically establish spectral dependencies guided by spatial priors. Additionally, a Residual Scaling Connection (RSC) is introduced to improve optimization stability and feature propagation efficiency. Extensive experiments conducted on three public datasets (i.e. Pavia Center, Chikusei, and MDAS) and a self-collected UAV-based Daxing dataset have demonstrated that the proposed method consistently outperforms seven state-of-the-art approaches in both quantitative metrics and visual quality. The performance gains of the proposed SCSRNet observed on the MDAS dataset achieves improvements of 6.0%, 9.8%, 0.7%, 9.7%, and 7.2% over the second-best comparative method. Furthermore, SCSRNet achieves superior reconstruction performance with relatively low computational complexity, underscoring its effectiveness and practicality for HMIF.
LST-based surface thermal anomaly extraction is an important task in remote sensing data analysis and is widely used in disaster monitoring and environmental assessment. This study proposes an isolated-neighborhood-based surface thermal anomaly extraction framework, termed SRAIN, for extracting spatially localized LST-based surface thermal anomalies. Taking coal-fired power plants in China from 2013 to 2024 as the primary study object and using Landsat 8/9 land surface temperature (LST) data as input, each sample is represented by a 12-month LST feature vector to characterize intra-annual monthly thermal-state variations. Drawing upon Tobler’s First and Second Laws of Geography, SRAIN assumes that spatial anomalies exhibit an “isolated” characteristic. Initial neighborhoods are constructed using Voronoi diagrams and dynamically merged using Jaccard and Silhouette coefficients to form integrated neighborhoods, thereby constituting isolated neighborhoods composed of both initial and integrated neighborhoods. In the anomaly extraction stage, ten anomaly detection algorithms are integrated, a confidence–intensity hybrid anomaly score (G-score) is proposed, and adaptive threshold optimization is performed using a beam-search-inspired strategy. In the anomaly pattern analysis stage, detected anomalies are classified into three modes: Concentrated (63.39%), Stretched (33.04%), and Diffuse (3.57%). Under the proxy-label-based relative evaluation framework, SRAIN improves proxy-label-based relative performance and spatial autocorrelation for LST-based surface thermal anomalies in the coal-fired power plant scenario. Compared with the best single algorithm, the fused G-score reduces the proxy-label-based misclassification rate by 12.9% and increases the spatial autocorrelation coefficient by 10.8%. In addition, the earthquake and fire cases provided in the Supplementary Materials preliminarily illustrate the potential transferability of the framework to other LST-based surface thermal anomaly scenarios, although systematic cross-scenario validation remains necessary in future studies. Overall, SRAIN provides a transferable framework for complex LST-based surface thermal anomaly extraction and offers a new perspective for environmental monitoring of anthropogenic heat sources such as coal power plants.
Optimal segmentation scale is a core parameter in object-based image analysis (OBIA). However, its variation with spatial resolution and land-cover type remains insufficiently characterized, constraining cross-resolution scale transfer in multi-source remote sensing mapping workflows. Using a coastal wetland landscape near the Yellow River estuary in Dongying, China, as the study area, this study systematically characterized optimal segmentation scales for seven land-cover types—tidal flat, bare land, buildings, water bodies, roads, cropland, and vegetation—across six spatial-resolution levels from 0.02 to 30 m. The optimal scale of each land-cover type was extracted using the Ratio of Mean Difference to Neighbors (absolute) to Standard Deviation (RMAS) method. Local variance (LV) quantified spatial heterogeneity, and perimeter-area fractal dimension (PAFRAC) provided an auxiliary descriptor of boundary complexity. Empirical regression models relating optimal scale to spatial resolution and LV were developed, and their utility was evaluated through classification and search-efficiency analyses.The results show that: (1) at the same spatial resolution, land-cover types with lower LV, namely tidal flat and water bodies, were consistently associated with larger optimal segmentation scales, whereas land-cover types with higher LV, namely buildings and roads, were associated with smaller optimal segmentation scales, indicating a stable negative cross-sectional correspondence. The PAFRAC results provided auxiliary support for this correspondence from the perspective of boundary complexity. (2) The functional form of the resolution–scale response systematically diverged by land-cover type: tidal flat and water bodies followed exponential-function models (R2 = 0.985–0.994), whereas bare land, buildings, vegetation, roads, and cropland followed power-function models (R2 = 0.931–0.996). This difference in function type was also reproduced in the LV–scale model, showing a consistent empirical pattern across the two model formulations. (3) In the 0.02 m ultra-high-resolution unmanned aerial vehicle (UAV) imagery, scale reversal was observed for buildings and vegetation: their optimal segmentation scales were lower than the corresponding values at 0.5 m and 2 m, deviating from the monotonic trend over the 0.5–30 m range. This phenomenon may reflect a shift in the dominant source of spatial heterogeneity from macroscopic patch morphology to micro-scale internal texture; however, because of the limited coverage of the UAV data, this interpretation requires further verification. (4) Across the five full-area satellite-image resolutions, overall accuracy (OA) varied non-monotonically with spatial resolution. The 10 m imagery achieved the highest OA (OA = 97.85%, Kappa = 0.97). Tidal flat and water bodies showed stable accuracy across the five satellite-image resolutions, whereas the accuracies of buildings and roads decreased markedly at spatial resolutions coarser than 10 m. (5) In an in-sample model evaluation, direct use of the model-predicted scales yielded lower performance than the RMAS-derived optimal scales, whereas prediction-centered local searches covered all 35 spatial-resolution–land-cover-type combinations and reduced candidate scales and estimated segmentation-only processing time by 41.8% and 38.7%, respectively.The empirical equations provide a quantitative basis for cross-resolution scale-search initialization in the present study area under the present data and processing conditions. Their applicability to structurally similar coastal wetlands requires independent multi-site validation.
Rapid detection of damaged buildings is critical for effective emergency response following sudden disasters. Although single-temporal methods enable rapid detection without requiring paired pre-event imagery, they often lack disaster-specific awareness, leading to limited accuracy and weak generalization. To address these critical limitations, we propose a disaster perception network (DPNet) tailored for single-temporal high-resolution remote sensing imagery, enabling more precise detection of damaged buildings in various disaster scenarios. DPNet integrates disaster perception module (DPM) to adaptively fuse disaster semantics with building damage features and employs the dynamic semantic-guided multi-task loss (DSML) to enforce cross-task semantic consistency, effectively reducing both false positives and missed detections. Extensive experiments conducted on a large-scale global dataset, encompassing 24 disaster events and 7 distinct disaster types, indicate that DPNet outperforms existing state-of-the-art models by over 7 % in mean intersection over union. Moreover, it significantly surpasses all comparative models on 7 additional unseen disaster events, further validating its generalization capability. These results demonstrate that DPNet enables robust building damage detection, providing a highly reliable and efficient technical solution for global disaster emergency response.
As the spatial resolution of remote sensing imagery continues to improve, Earth observation scenes exhibit greater diversity in spatial and spectral characteristics, together with increasingly complex surface structures across multiple scales. In such scenarios, global land-cover patterns and local high-frequency details are often strongly coupled, posing significant challenges for accurate scene interpretation and boundary delineation. From a feature representation perspective, spatial-domain features are effective for modeling global context and high-level semantics, whereas frequency-domain representations facilitate the separation of structural components at different scales and enhance fine boundary details, but have limited capability in explicitly capturing semantic relationships. Consequently, most existing semantic segmentation methods operate in a single domain, either spatial or frequency, which restricts their ability to jointly model semantic consistency and diverse spatial structures. To address this issue, we propose a semantic segmentation network termed Spatial–Frequency Collaborative Modeling Network (SFMNet), which jointly learns complementary representations in both spatial and frequency domains. By leveraging spatial-domain semantic priors to guide frequency-domain structural enhancement, SFMNet enables coordinated modeling of land-cover semantics and multi-scale surface structures. In addition, cross-scale interaction between low- and high-frequency components and adaptive multi-level feature fusion are introduced to improve structural consistency across hierarchical feature representations. Extensive experiments on the ISPRS Potsdam, Vaihingen, and LoveDA datasets demonstrate that SFMNet consistently outperforms existing methods in both overall segmentation accuracy and boundary delineation quality. Specifically, SFMNet achieves mean Intersection-over-Union scores of 81.30%, 70.23%, and 55.07%, and Boundary mean Intersection-over-Union scores of 65.42%, 56.76%, and 43.74% on the three datasets, respectively.
Geostationary satellite-based PM2.5 mapping increasingly relies on data-driven models using top-of-atmosphere (TOA) reflectance. However, most existing approaches train a single estimator on temporally aggregated samples, implicitly assuming a stationary relationship between TOA radiative signals and near-surface PM2.5. This assumption overlooks the pronounced diurnal variability captured by geostationary satellites and limits the representation of hour-dependent aerosol–radiation interactions. To address this limitation, we developed a physically constrained multi-task learning framework, termed the Multi-Temporal Decoupled Attention Network (MTDAN), for hourly PM2.5 estimation from geostationary satellite TOA reflectance. The framework integrates three components: (1) a physically guided surface–atmosphere decoupling module that suppresses surface-background interference and enhances aerosol-sensitive information; (2) a multi-task attention architecture that jointly learns shared and hour-specific relationships between PM2.5 and explanatory variables; and (3) a temporal-constrained loss function that exploits diurnal continuity to improve the consistency of sequential predictions. Experiments over eastern China using year-long Himawari-8 observations and ground-based PM2.5 measurements demonstrated the effectiveness of the proposed framework. MTDAN achieved an overall site-based validation R2 of 0.89 with an RMSE of 9.27 μg m⁻3, outperforming representative machine-learning and deep-learning models by reducing estimation errors by 21.7 %–27.8 %. The model successfully reproduced hourly pollution evolution, captured localized hotspots, and reconstructed season-dependent diurnal PM2.5 dynamics. Ablation analyses further confirmed the complementary contributions of physical decoupling, adaptive feature weighting, and temporal-consistency constraints. These results demonstrate that explicitly modeling hour-dependent aerosol–radiation relationships can substantially improve geostationary satellite-based PM2.5 estimation. The proposed framework provides an interpretable and robust solution for high-frequency air-quality monitoring and offers a general strategy for integrating physical constraints and temporal context into Earth observation applications.
Bacterial leaf blight (BLB) remains a major biotic constraint to global rice production, requiring faster and scalable detection methods than conventional field scouting to minimise yield losses. Early detection of BLB is essential for timely intervention and yield protection, yet very few studies have examined the use of remote sensing, particularly during the early growth phase of the rice crop. This study evaluated the potential of canopy-level hyperspectral data for identifying BLB in early vegetative phase. Field trials were conducted at the International Rice Research Institute, where canopy hyperspectral reflectance was collected from healthy and inoculated plots at the stem elongation and booting stages of the rice crop. Biophysical and biochemical parameters were also collected during the stem elongation stage. Four spectral transformation techniques, including first and second derivatives (FD, SD), continuum removal (CR), and continuous wavelet transform (CWT), were applied to canopy spectral data to identify BLB-sensitive features. To reduce the high dimensionality of hyperspectral features, the top 5 % most separable features were extracted, and redundancies were removed using hierarchical clustering combined with variance inflation factor filtering. Support vector machines (SVM), random forests (RF), and XGBoost classifiers were trained, and their performance on BLB detection was evaluated. Statistical tests revealed that transformed spectra provided greater class separability than original reflectance data, particularly in identifying bands within the near-infrared (NIR, 710 nm to 1333 nm) and shortwave infrared (SWIR, 1451 nm to 2344 nm) regions. The SVM model using 20 selected features achieved the highest F1-score on an independent test set (73 %), outperforming RF and XGBoost (67 % and 68 %, respectively), and yielded a recall/sensitivity of 98 % for BLB identification. Notably, the first-derivative feature at 1076 nm was the most informative predictor for BLB detection, as determined by the separability test and absolute SVM coefficient. Most of the identified sensitive wavelengths correspond to changes in chlorophyll, leaf area index (LAI), and fresh biomass. Using a controlled experiment, this proof-of-concept study demonstrates the potential of canopy hyperspectral data, combined with spectral transformations and machine learning algorithms, to support early BLB detection and provides a methodological basis for future validation in complex regional-scale environments.