Urban resilience is critical in the context of global disruptions such as the COVID-19 pandemic, yet recovery processes are rarely uniform across urban economies. How resilience varies across economic sectors and geographical contexts, and the extent to which it is structured by business essentiality, remain poorly understood at the global scale. Here, we use the COVID-19 lockdowns as a natural experiment and analyze monthly satellite-derived nighttime light trajectories for 105 city–sector combinations across 48 global cities between 2018 and 2023. We identify four archetypal sectoral trajectories capturing distinct resilience responses and two dominant dimensions associated with their variation: economic essentiality and sectoral structure. These trajectories are strongly structured by essentiality, with resilient trajectories concentrated in essential sectors, whereas Chronic Decline is disproportionately associated with non-essential sectors, particularly in Europe. A pronounced geographic divergence is also evident: cities in Latin America and Asia more frequently exhibit Resilient or Full Recovery trajectories, whereas European cities show widespread and persistent declines. The observed resilience trajectories vary systematically across cities and sectors, reflecting geographically contingent, sector-specific patterns associated with essentiality. Our findings suggest that city-level resilience assessments may overlook persistent sectoral differences and highlight the value of sector-sensitive assessments of urban resilience. Urban resilience differed across sectors and regions, with essential sectors more often resilient and European sectors more often in decline, according to a study analyzing satellite nighttime light trajectories for 105 city-sector combinations across 48 cities.
Few-shot learning has increasingly been explored in remote sensing scene classification as an effective approach to learning from limited examples. However, existing methods suffer from several limitations, including the need for well-annotated auxiliary datasets, limited generalisation and a tendency to overfit on training data. In this paper, a novel semi-supervised self-organised prototype tree-based method (S3OPT) is proposed for few-shot remote sensing scene classification. S3OPT progressively constructs a hierarchical prototype tree from image embeddings in a top-down, discriminatory manner across multiple levels of granularity, capturing interclass similarities and intra-class variations to enable automated class separation. Using pretrained convolutional neural networks for feature extraction and exploiting pseudo-labelling, S3OPT reduces the need for extensive manual labelling and enhances generalisation through self-training from unlabelled samples. Thanks to the prototype-based nature, S3OPT offers high transparency and its reasoning is based on the mutual similarity between images, ensuring explainability in internal reasoning and decision-making. Extensive experiments on four widely used benchmark datasets demonstrate the great classification accuracy of the proposed S3OPT under standard few-shot learning protocols. It achieved results superior to or on par with state-of-the-art methods for few-shot remote sensing scene classification, delivering up to a 15% accuracy increase without the requirement for computationally expensive training and/or fine-tuning.
Flood type information is critical for effective flood risk management because dominant flood types are associated with distinct hydrodynamic behaviour, contamination pathways, and recovery trajectories. However, most operational flood mapping products provide only binary inundation extent, offering limited information for interpreting the dominant flood type and its likely impact characteristics. Existing flood type classification approaches rely predominantly on hydrometeorological observations and modelling, which are often unavailable in data-scarce regions and can be unstable in mechanism-complex environments such as estuarine deltas, urban river corridors, and coastal cities. To address these limitations, this study proposes a multi-CNN framework that integrates flood type classification directly into satellite-based flood mapping. The framework first uses a U-Net model for flood extent segmentation and then applies a CNN-based classifier for scene-level flood type identification by combining satellite imagery with auxiliary topographic and hydrological-context features, including DEM information and water connectivity ratios. Several CNN architectures were compared for flood type classification, with Inception-ResNet selected based on its performance–complexity trade-offs. To support model training and evaluation, this study introduces the World Flood Type (WFT) dataset, a new multi-event flood dataset containing 464 flood scenes from 120 flood events across 48 countries. Results show effective performance in both inundation segmentation and flood type classification. The U-Net model achieved 83.8% overall accuracy on an independent test dataset, while the final Inception-ResNet-based classifier achieved 93.0% scene-level overall accuracy and a scene-level macro-F1 score of 0.73 under dominant-mechanism conditions. These findings demonstrate the feasibility of extending conventional flood extent products with scene-level attribution of dominant flood types using remotely sensed imagery and ancillary geospatial data, thereby providing more decision-relevant information for emergency response and post-disaster assessment.
Accurate crowd detection (CD) is critical for public safety and historical pattern analysis, yet existing methods relying on ground and aerial imagery suffer from limited spatio-temporal coverage. The development of very-fine-resolution (VFR) satellite sensor imagery (e.g., similar to 0.3 m spatial resolution) provides unprecedented opportunities for large-scale crowd activity analysis, but it has never been considered for this task. To address this gap, we proposed CrowdSat-Net, a novel point-based convolutional neural network, which features two innovative components: Dual-Context Progressive Attention Network (DCPAN) to improve feature representation of individuals by aggregating scene context and local individual characteristics, and High-Frequency Guided Deformable Upsampler (HFGDU) that recovers high-frequency information during upsampling through frequency-domain guided deformable convolutions. To validate the effectiveness of CrowdSat-Net, we developed CrowdSat, the first VFR satellite imagery dataset designed specifically for CD tasks, comprising over 120 k manually labeled individuals from multi-source satellite platforms (Beijing-3 N, Jilin-1 Gaofen-04A and Google Earth) across China. In the experiments, CrowdSat-Net was compared with eight state-of-the-art point-based CD methods (originally designed for ground or aerial imagery and satellite-based animal detection) using CrowdSat and achieved the largest F1-score of 66.12 % and Precision of 73.23 %, surpassing the second-best method by 0.80 % and 6.83 %, respectively. Moreover, extensive ablation experiments validated the importance of the DCPAN and HFGDU modules. Furthermore, cross-regional evaluation further demonstrated the spatial generalizability of CrowdSat-Net. This research advances CD capability by providing both a newly developed network architecture for CD and a pioneering benchmark dataset to facilitate future CD development. The source code is available at https://github.com/Tong-777777/CrowdSat-Net.
Accurate, reliable, and up-to-date information on wildlife populations is crucial for biodiversity conservation in the face of unprecedented biodiversity loss worldwide. However, monitoring wildlife populations at large scales remains challenging. Advances in satellite remote sensing, particularly very-high-resolution satellite data, offer new opportunities for monitoring wildlife from space, and new machine learning techniques present great potential for detecting wildlife with remarkable speed and accuracy. Here, we introduce a deep learning pipeline for automatically detecting and counting large migratory ungulate herds (wildebeest and zebra) at the individual level in the Serengeti-Mara ecosystem from submeter-resolution satellite imagery. We apply the pipeline to implement the first-ever population census of large-size ungulates in the Serengeti-Mara ecosystem through a satellite survey and generate the total count of the whole population. The model shows robust performance across diverse landscapes with an overall F1-score of 84.75% (Precision: 87.85%, Recall: 81.86%) on an independent test dataset containing 11,594 animals and achieves good transferability spatially and temporally. This research showcases the capability of satellite remote sensing and deep learning techniques to accurately locate and count very large populations of terrestrial mammals in open landscapes. It provides a new perspective on monitoring wildlife populations and animal migration, which will facilitate the understanding of animal behavior and ecology as well as improve the conservation of the whole ecosystem in the face of rapid environmental changes.
Mapping forest above-ground biomass (AGB) is crucial for monitoring forest ecosystems and assessing the success of conservation initiatives such as the REDD + carbon projects. Traditional field-based approaches to measuring AGB, however, face significant challenges, due to high financial costs and logistical constraints. Remote sensing, including both active and passive sensors, presents a promising and cost-effective alternative, yet its practical utility and accuracy for capturing forest AGB in diverse and complex ecosystems remains largely unexplored. This research used an extensive national forest inventory (NFI) dataset to evaluate the ability to map the AGB of the Miombo woodlands in Zambia across four agro-ecological zones using both multi-seasonal SAR (Sentinel-1A) and optical (Landsat-8 OLI) imagery. A multi-level experiment was designed to (i) compare the accuracy of AGB estimation using SAR and optical data when used independently, and in combination, using a Random Forest regression model, (ii) assess the effect of seasonality on the accuracy of AGB estimation when using SAR and optical datasets, and (iii) evaluate the effect of variation in climatic and environmental conditions on AGB estimation. Experimental results show that multi-seasonal images (across the rainy, hot and dry seasons) outperformed single-season and annual images. Combining SAR backscatter in the hot season, optical bands in the dry season, and vegetation indices in the hot season produced the most accurate AGB model (R = 0.69, MAE = 14.01 Mg ha(-1) and RMSE = 18.23 Mg ha(-1)). The models performed distinctly across different agro-ecological zones (R = 0.44 - 0.79), suggesting that fitting local models could be beneficial. These results based on the extensive NFI of Zambia demonstrate that seasonal effects and fitting local models can lead to more accurate AGB estimation within the Miombo woodlands, which is of significance for ongoing REDD + carbon projects in Zambia and other African countries.
Change detection(CD) is important for Earth observation, emergency response and time-series understanding. Recently, data availability in various modalities has increased rapidly, and multimodal change detection (MCD) is gaining prominence. Given the scarcity of datasets and labels for MCD, unsupervised approaches are more practical for MCD. However, previous methods typically either merely reduce the gap between multimodal data through transformation or feed the original multimodal data directly into the discriminant network for difference extraction. The former faces challenges in extracting precise difference features. The latter contains the pronounced intrinsic distinction between the original multimodal data; direct extraction and comparison of features usually introduce significant noise, thereby compromising the quality of the resultant difference image. In this article, we proposed the MaCon framework to synergistically distill the common and discrepancy representations. The MaCon framework unifies mask reconstruction (MR) and contrastive learning (CL) self-supervised paradigms, where the MR serves the purpose of transformation while CL focuses on discrimination. Moreover, we presented an optimal sampling strategy in the CL architecture, enabling the CL subnetwork to extract more distinguishable discrepancy representations. Furthermore, we developed an effective silent attention mechanism that not only enhances contrast in output representations but stabilizes the training. Experimental results on both multimodal and monomodal datasets demonstrate that the MaCon framework effectively distills the intrinsic common representations between varied modalities and manifests state-of-the-art performance across both multimodal and monomodal CD. Such findings imply that the MaCon possesses the potential to serve as a unified framework in the CD and relevant fields. Source code will be publicly available once the article is accepted.
Spatio-Temporal Image Fusion (STIF) methods usually require sets of images with matching spatial and spectral resolutions captured by different sensors. To facilitate the application of STIF methods, we propose and compare two different standardization approaches. The first method is based on traditional upscaling of the fine-resolution images. The second method is a sharpening approach called Anomaly Based Satellite Image Standardization (ABSIS) that blends the overall features found in the fine-resolution image series with the distinctive attributes of a specific coarse-resolution image to produce images that more closely resemble the outcome of aggregating the fine-resolution images. Both methods produce a significant increase in accuracy of the Unpaired Spatio Temporal Fusion of Image Patches (USTFIP) STIF method, with the sharpening approach increasing the spectral and spatial accuracies of the fused images by up to 49.46% and 78.40%, respectively.
Greater accessibility to China’s vast archive of satellite Earth observations could enhance scientific progress, disaster preparedness, and international cooperation.
Tropical evergreen forests, among the most diverse and complex forest ecosystems, host an extraordinary variety of plants and animal species, underscoring the importance of monitoring their spatio-temporal dynamics. Despite substantial advances in tracking tropical forest disturbances with Landsat images, the understanding of post-disturbances recovery, particularly small-scale disturbances, remains limited. Here, by enhancing the subtle signal caused by small-scale disturbances, we propose a method (i.e., framework) to simultaneously monitor tropical evergreen forest disturbances and post-disturbance recovery using annual maximal Landsat canopy openings. A baseline evergreen forest cover (EFC) map was automatically generated by integrating median normalized difference fraction index (NDFI) and maximal self-referenced NDFI (rNDFI) images. Utilizing the baseline map and long-term spatio-temporally filtered annual maximal rNDFI images, three additional thematic maps were produced: first forest disturbance year (FDY), interannual forest disturbance frequency (FDF), and post-disturbance forest recovery (PFR). For the study areas of Brazil Mato Grosso, the Democratic Republic of the Congo (DRC) Kasai and Indonesia Kalimantan Tengah, the results demonstrated that the proposed framework not only detected more forest disturbance events, particularly enhancing the small-scale disturbances, but also predicted forest disturbances more accurately than the Global Forest Change (GFC) forest loss and Joint Research Centre (JRC) Tropical Moist Forest (TMF) deforestation and degradation products. The generated interannual FDF and PFR maps exhibit high accuracies, with the R-2, mean absolute error (MAE), and root mean square error (RMSE) in the ranges 0.82-0.95, 0.47-1.15, and 1.33-2.88, respectively, which are more accurate than the benchmark results extracted from the JRC TMF product. Our results are more sensitive to the interannual canopy opening and recovery of small-scale disturbances due to smallholder clearing and selective logging than JRC TMF. This research offers an effective framework for addressing the existing knowledge gap on postchange dynamics after forest disturbances and the net carbon change of tropical evergreen forests.
Semantic change detection (SCD) involves the simultaneous extraction of changed regions and their corresponding semantic classifications (pre- and post-change) in remote sensing images (RSIs). Despite recent advancements in vision foundation models (VFMs), the fast-segment anything model has demonstrated insufficient performance in SCD. In this article, we propose a novel VFMs architecture for SCD, designated as VFM-ReSCD. This architecture integrates a side adapter (SA) into the VFM-ReSCD to fine-tune the fast segment anything model (FastSAM) network, enabling zero-shot transfer to novel image distributions and tasks. This enhancement facilitates the extraction of spatial features from very high-resolution (VHR) RSIs. Moreover, we introduce a recurrent neural network (RNN) to model semantic correlation and capture feature changes. We evaluated the proposed methodology on two benchmark datasets. Extensive experiments show that our method achieves state-of-the-art (SOTA) performances over existing approaches and outperforms other CNN-based methods on two RSI datasets.
Land surface temperature (LST) data are crucial for global climate change research. While remote sensing data serve as a key source for LST, single-source sensor data often lack spatiotemporal continuity due to long satellite revisit intervals and cloud cover. Spatiotemporal fusion, which combines the strengths of multiple sources, can increase the available information. However, most current spatiotemporal fusion methods are designed for local-scale applications. This research proposes the Global Spatiotemporal Fusion Model (GLOSTFM) to generate global LST products. GLOSTFM, built on image pyramid principles, addresses the computational and complexity challenges of global-scale spatiotemporal fusion. Moreover, the model utilizes data from the novel Fengyun-3D satellite, which has a daily revisit capability and provides LST products separately derived from its thermal infrared (MERSI, 1 km) and microwave (MWRI, 25 km) sensors. By leveraging the cloud-penetrating capabilities of the microwave data to compensate for missing information, GLOSTFM increases the available information and reduces observational uncertainties. The results showcase high processing efficiency and enhanced spatiotemporal continuity, with an average RMSE of 2.874 K and an excellent R-2 of 0.980. The utility of the GLOSTFM model for monitoring urban heat island effects in Beijing was explored to illustrate one application among a broad range of potential applications of the proposed GLOSTFM that require global data on LST across the Earth's surface.
Fractional vegetation cover (FVC) is a critical component of ecosystems, global climate change, and the carbon cycle. Several FVC products have been released, the most widely used of which are the GLASS FVC products (including the GLASS-Moderate Resolution Imaging Spectroradiometer (MODIS) and GLASS-AVHRR FVC products). Specifically, the GLASS-MODIS FVC product covers the period from 2000 to present with a 500-m spatial resolution, whereas the GLASS-AVHRR FVC product is available from 1982 to present with a coarser spatial resolution of 5-km. For local monitoring of patterns of change in vegetation, however, there is a great need for fine spatial resolution (e.g., 500-m in this article) and long-term time-series FVC datasets. To this end, we proposed to reconstruct a 500-m, 8-day historical MODIS FVC dataset (1982-2000) by making full use of the advantages of the existing GLASS-MODIS FVC (a fine spatial resolution of 500-m) and GLASS-AVHRR FVC (long-term coverage from 1982 to the present) products covering China in this article. The known GLASS-AVHRR FVC product was first used to fit the relationship between the FVC data after 2000 and before 2000, based on a random forest (RF) model. The trained relationship was migrated to the GLASS-MODIS FVC product, that is, predicting the MODIS FVC before 2000 based on the input of MODIS FVC after 2000. The validation using 64 scenes of Landsat FVC reference data revealed that the predicted historical MODIS FVC dataset has a reliable accuracy with a correlation coefficient (CC) value of 0.84, a root-mean-square error (RMSE) of 0.14, a Bias of 0.04, and an unbiased RMSE (ubRMSE) of 0.12. Moreover, an accuracy evaluation in seven different regions in 1999 suggested that the historical MODIS FVC is closer to the Landsat FVC than the GEOV2 FVC product. Overall, the 500-m, 8-day MODIS FVC dataset (1982-2000) in China can provide important historical data for long-term, local monitoring of vegetation, which has great potential in supporting studies in a range of application areas, including ecology, hydrology, and climatology. This dataset is available at https://doi.org/10.6084/m9.figshare.24616446.v1
Tropical dry forests, such as the Miombo woodlands, play crucial roles both as an effective global carbon sink, and as the source of livelihood for a vast number of local communities. However, mapping Miombo woodlands accurately into definable classes is a great challenge due to their sparse and heterogeneous nature and their alteration due to anthropogenic impacts. Nevertheless, such mapping is important to underpin management and conservation efforts. We explored the potential of Sentinel (S-1) and Sentinel-2 (S-2) seasonal and multi-seasonal images for two tasks: (i) mapping Land Use Land Cover (LULC) such as to identify the Miombo woodlands and (ii) mapping three specific forest classes (reference, degraded and regrowth forests) within the Miombo woodlands of Zambia. The Random Forest (RF) algorithm within Google Earth Engine (GEE) was selected for the LULC classification while a U-Net convolutional neural network (CNN) was applied to classify the different types of forest. Models were trained, validated, and tested using ground validation data. The RF model achieved an overall accuracy of 93 % for LULC classification, with the forest class F1-scores ranging from 93 % to 96 % across different seasons. The U-Net CNN effectively delineated the Miombo woodlands into reference, degraded and regrowth forests, with respective F1-scores of 85 %, 73 % and 72 %. Combining multi-seasonal S-1 and S-2 images and their derivatives yielded the greatest accuracy for LULC mapping, while combining the year-round S-1 and S-2 bands produced the highest F1-scores for the forest type classification. The hierarchical approach employed was, thus, demonstrated to be effective, providing more nuanced functional forest information. The approach holds great promise for mapping and monitoring programs aiming to manage and conserve the Miombo woodlands sustainably.
Solar-induced chlorophyll fluorescence (SIF) is a crucial variable toward timely and effective monitoring of vegetation productivity, as well as physiological and biochemical parameters, across extensive areas. Among these advances, the Tropospheric Monitoring Instrument (TROPOMI) SIF has significantly increased the spatiotemporal resolution and data coverage compared with previous sensors. However, TROPOMI SIF data suffer from nonuniform sampling, swath gaps, and cloud contamination, resulting in numerous instances of missing data. In this article, we proposed a physical and spatial information-aided gap filling (PSGF) method, which effectively addresses the missing data problem, generating a Spatially Seamless, 0.05 degrees, daily SIF (S2-SIF) dataset globally at a spatial resolution of 0.05 degrees from 2018 to 2021. Through missing data simulation experiments conducted in six regions worldwide, we demonstrated consistency between the reference SIF and filled SIF, with a correlation coefficient (CC) of 0.659. Furthermore, validation using in situ data from 35 SIF and gross primary productivity (GPP) ground sites yielded a CC of approximately 0.70 for the SIF sites and CC values above 0.60 between the ground GPP and filled SIF. In addition, consistency was observed between the filled SIF datasets and two other SIF products across 11 vegetation types, confirming the reliability of the filled SIF data and the efficacy of the PSGF method. The produced filled SIF data are made publicly available and should greatly increase the applicability of the daily SIF data for a wide range of applications, including quantifying the photosynthesis of vegetation and accurately estimating GPP globally.