Long time series impervious surface mapping (ISM) is important for understanding urban expansion, environmental impacts, and urban planning. There are some historical global ISM products, such as GAIA and NUACI datasets, whereas they may not meet the user’s diverse application needs in the aspects of mapping timeliness, temporal resolution, and spatial resolution. Therefore, this study proposes an automatic, rapid, and continuous impervious surface mapping and updating framework based on historical land cover datasets without using other labeled data, to improve the updating speed and spatio-temporal resolution of impervious surface maps. The main process is divided into three steps: (1) Multi-temporal samples for classification were obtained by using GAIA dataset, FROM-GLC dataset and the unsupervised continuous change detection (CCD) algorithm; (2) Quarterly long time series ISM results (ISMs) were obtained by using multi-temporal samples and quarterly features; (3) The final results were obtained by using the post-processing operations in the obtained quarterly long time series ISMs to improve the mapping accuracy. The proposed framework is applied to eight cities around the world, and the total overall accuracy (OA) and Kappa of long time series ISMs with post-processing in the eight cities are 92.64% and 0.8525, respectively, improving the OA and Kappa of those without post-processing by 1.41% and 0.0281, respectively, and those of GAIA dataset by 4.57% and 0.0914, respectively, which proved the effectiveness of the proposed method. This study also analyzed the spatial patterns of impervious surface expansion in eight cities and identified different spatial patterns of expansion that existed among the cities, while capturing the abrupt change in the spatial patterns of expansion in Rosario and Novosibirsk after the second quarter of 2021. The proposed framework achieved rapid mapping and updating of impervious surface without any labeled samples, and has the potential to map the global impervious surface continuously.
Remote sensing image captioning is becoming an advanced technique for earth understanding in domains such as urban development, transportation management, and environmental monitoring. However, most existing remote sensing captioning methods still produce visually accurate yet task-irrelevant descriptions, limiting their applicability in further downstream tasks. In this paper, we introduce a global multi-task image captioning dataset (EarthReport) to promote the applicable remote sensing image understanding, which features reasoning tasks such as construction, traffic, ecology, land use statistics, and general summaries, encompassing 10,950 images, land-cover masks, and 1.06 million captions from 14 countries across six continents. To our knowledge, EarthReport is currently the largest available multi-task dataset for remote sensing image captioning. Furthermore, we propose a remote sensing task-conditioned captioning framework (RSCoCap) to generate specific captions, in which a task controller is designed to integrate dynamic task prompts with object-guided visual features and to leverage a fine-tuned large language model for task-oriented caption generation. Experiments on EarthReport show that RSCoCap exceeds previous methods by more than 1 similar to 2%in Bleu and Cider, indicating that our work provides a strong benchmark for applicable Earth Vision's comprehensive understanding. Applying our framework to 20 representative global cities, we identified four patterns of city development stage types. These patterns show spatial variation within cities, influenced by distance from the city center.
Hyperspectral video data provide complementary contextual cues across spectral, spatial, and temporal dimensions for modeling object dynamics under challenging conditions. Many existing hyperspectral video object tracking (HVOT) approaches organize spatial-spectral and temporal modeling in successive stages, leaving room for closer interaction among video-level contextual cues. To address this, we propose HucrTrack, a unified contextual reasoning framework for HVOT trained by parameter-efficient fine-tuning (PEFT). HucrTrack forms synchronized hyperspectral and false-color representations from each hyperspectral cube and enhances spatial-spectral features through a weight-shared dual-representation backbone with unified contextual cue modeling. To effectively leverage contextual dynamics, we design a unified contextual reasoning module (UCRM) composed of three key components: memory dynamics unit (MDU), contextual injection unit (CIU), and selective retrieval unit (SRU). Specifically, MDU maintains a frame-wise dynamic memory via Mamba's hidden states; CIU hierarchically integrates this memory into the spectral-spatial backbone features; and SRU selectively retrieves relevant contextual information to reinforce the tracking representation. In contrast to representative stepwise designs, HucrTrack enables concurrent, unified reasoning over all three dimensions within one recurrent process. Extensive experiments on ten benchmarks demonstrate that HucrTrack compares favorably with existing trackers in both robustness and generalization.
In recent years, optical remote sensing imagery has played an increasingly vital role in Earth observation, but cloud contamination exists as an inevitable degradation. Combining synthetic aperture radar (SAR) and optical data with machine learning offers a promising solution for reconstructing clear-sky satellite imagery. Nevertheless, several challenges persist, including insufficient attention to large cloud cover, difficulties in restoring temporal changes, and limited practicality of deep models. To address these issues, this paper introduces a novel deep learning-based cloud removal framework, termed Began+, which integrates bi-temporal SAR-optical data to deal with cloudy images with high cover ratios. The Began+ framework comprises two primary components: a deep network and a flexible post-processing step, combining the strengths of data-driven models for restoring change information and traditional gap-filling algorithms for mitigating radiance discrepancies. First, a bi-output enhanced generative adversarial network, abbreviated as Began, is designed for image synthesis, featuring an enhanced channel-wise fusion block (ECFB) and a multi-scale depth-wise convolution residual block (MDRB). By applying the dual-tasking optimization and co-learning strategy, the Began model identifies potential change areas from bi-temporal SAR and pre-temporal optical inputs, guiding the synthesis of target optical images. Second, a range of cloud masking and gap-filling techniques can be optionally employed to effectively reduce radiometric discrepancies between the synthesized images and the cloudy data, ultimately yielding high-quality, clear-sky imagery. To meet the big data requirements of deep learning, we constructed two globally distributed cloud removal datasets, named BiS1L8-CR and BiS1S2-CR. Supported by these datasets, extensive experiments demonstrated that the Began+ framework effectively captures bi-temporal change features, reconstructing precise surface information in both Landsat-8 and Sentinel-2 satellite images under large cloud cover. Compared to the latest solutions and algorithms, our proposed Began+ framework exhibits significant advantages from both qualitative and quantitative perspectives in both simulated and real experiments. Furthermore, without strict constraints on input timing, the Began+ framework enables accurate reconstruction of large-scale dual-sensor imagery under high-ratio cloud cover, effectively restoring changing surfaces and improving the quality of unsupervised vegetation extraction.
Precise quantification of gross primary productivity (GPP) at fine resolution is crucial for regional carbon cycle analysis. However, existing GPP products frequently fail to capture fine-scale variability due to an inherent trade-off between satellite spatiotemporal resolution and the coarse-scale design of most existing GPP models. To address this issue, a machine learning-based downscaling and correction framework was designed to generate high-accuracy monthly GPP at 30-m resolution (DS-FC-GPP) through fusing multiple remote sensing model-based products along with eddy covariance (EC) data. Moderate resolution imaging spectroradiometer (MODIS) and global land surface satellite (GLASS) GPP products were simultaneously downscaled from 500 m to 30 m resolution using a fine-resolution vegetation index time series. Subsequently, a unified fusion and correction framework was established based on EC data. The analysis revealed that DS-FC-GPP exhibited enhanced spatial patterns with finer details and maintained consistency with the original products when aggregated back to coarse resolution. In the spatial validation, the leave-one-site-out validation achieved a coefficient of determination (R-2) of 0.828 with a mean absolute error (MAE) of 39.497 gC.m(-2)& centerdot;month(-1), while the leave-one-region-out scheme yielded a comparable performance with an R-2 of 0.812 and an MAE of 42.506 gC.m(-2)& centerdot;month(-1). The consistently strong performance across both validation strategies demonstrates the robustness of the proposed framework. Moreover, DS-FC-GPP better captured intra-annual peaks and valleys, maintaining stable accuracy throughout the year. The transfer experiment over China indicated that the DS-FC-GPP framework exhibits robust generalization and transferability. The present work highlights the potential of multisource data fusion and downscaling to promote fine-resolution GPP estimation, offering valuable insights into terrestrial carbon cycle processes.
China's New-Type Urbanization Plan since 2015-the world's largest urbanization endeavor-reshapes the nation's socioeconomic landscape but lacks high-precision, fine-scale progress monitoring. Urban construction sites (UCSs)-barometers of urban spatial expansion and renewal-offer a detailed observational window. A sub-meter-resolution deep-learning framework for nationwide UCS mapping is proposed. Using a Segment Anything Model-enabled weakly supervised method for pixel-level UCS annotation, a spectral-texture dual-branch segmentation network with 94.3% overall accuracy identifies 541 177 UCSs (including 10-m² micro-sites) across 372 cities. K-means clustering partitions cities into four typologies, uncovering a dual-track parallel pattern (incremental expansion + stock optimization) vs the classical 'growth-decline-renewal' trajectory. Spatial analysis shows that UCS construction correlates with annual PM₂.₅ concentrations; green dust-proof net coverage (<10%) fails to curb pollution. The framework serves as a 'microscope' for evaluating new-type urbanization and supports sustainable planning.
Binary semantic segmentation in remote sensing (RS) imagery faces persistent challenges due to complex object appearances, ambiguous boundaries, and high similarity between foreground and background, all of which introduce significant uncertainty into the prediction process. Existing approaches often treat uncertainty as either a global attribute or a pixel-level estimate, overlooking the critical role of spatial and contextual interactions. To address these limitations, we propose the Progressive Uncertainty-Guided Segmentation Network (PUGNet), a unified framework that explicitly models uncertainty in a context-aware manner. PUGNet decomposes uncertainty into three distinct components: foreground uncertainty, background uncertainty, and contextual uncertainty. This tripartite modeling enables more precise handling of local ambiguities and global inconsistencies. We adopt a coarse-to-fine decoding strategy that progressively refines features through two specialized modules. The Dynamic Uncertainty-Aware Module enhances regions of high foreground and background uncertainty using Gaussian-based modeling and contrastive learning. The Entropy-Driven Refinement Module quantifies contextual uncertainty via entropy and facilitates adaptive refinement through multi-scale context aggregation. Extensive experiments on ten public benchmark datasets, covering both single-temporal (e.g., building and cropland extraction) and bi-temporal (e.g., building change detection) binary segmentation tasks, demonstrate that PUGNet consistently achieves superior segmentation accuracy and uncertainty reduction, establishing a new state of the art in RS binary segmentation. The full implementation of the proposed framework and all experimental results can be accessed at https://github.com/Henryjiepanli/PU_RS.
Multi-temporal hyperspectral imagery (HSI) has become a powerful tool for change detection (CD) owing to its rich spectral signatures and detailed spatial information. Nevertheless, the application of paired HSIs is constrained by the scarcity of annotated training data. While unsupervised domain adaptation (UDA) offers a potential solution by transferring change detection knowledge from source to target domains, two critical limitations persist: (1) the labor-intensive process of acquiring and annotating source-domain paired samples, and (2) the suboptimal transfer performance caused by substantial cross-domain distribution discrepancies. To address these challenges, we present a Temporal Self-Construction Cross-Domain learning (TSCCD) framework for UDA-based HSI-CD. Our TSCCD framework introduces an innovative temporal self-construction mechanism that synthesizes bi-temporal source-domain data from existing HSI classification datasets while simultaneously performing initial data-level alignment. Furthermore, we develop a reweighted amplitude maximum mean discrepancy (MMD) metric to enhance feature-level domain adaptation. The proposed architecture incorporates an attention-based Kolmogorov-Arnold network (KAN) with high-frequency feature augmentation within an encoder-decoder structure to effectively capture change characteristics. Comprehensive experiments conducted on three benchmark HSI datasets demonstrate that TSCCD achieves superior performance compared to current state-of-the-art methods in HSI change detection tasks. Codes are available at https://github.com/Zhoutya/TSCCD.
While the COVID-19 pandemic has affected global public health and economies, evidence on its impact on urban expansion remains limited. Here we show, using high-dynamic global urban expansion mapping, that global urban areas expand by 6,722.9 km² during 2020–2022, 27% less than expected in the absence of the pandemic. The global urban expansion rate decreases by 29%, implying a 9.8-month delay relative to a counterfactual no-pandemic scenario. Average urban expansion area in upper-middle-income countries is more than five times that in lower-middle-income countries, while high-income countries also experience substantial reductions in expansion rates. Inequality in urban expansion intensity increases, with 90.6% of the net increase arising from differences among countries within the same income group. Additional country-month analyses suggest that strict government containment policies are more consistently associated with the shortfall in urban expansion than COVID-19 case rates alone. These findings demonstrate the potential of monthly satellite mapping for assessing short-term disruptions to urbanization. Consider this: The COVID-19 pandemic slowed global urban expansion to 6722.9 square kilometers in 2020–2022, 27 percent below expected levels and causing a 9.8-month delay, with widening inequality driven mainly within income groups according to monthly satellite mapping and country-month analysis.
Change detection using multitemporal remote sensing imagery was a critical tool for monitoring land surface dynamics. However, conventional methods often suffered from limited spatial-temporal feature integration, seasonal noise interference, and poor interpretability. To address these limitations, we proposed a temporal-consistent spectral-spatial fusion network (TC-SSFN) that jointly modeled local spatial-spectral structures and long-term temporal dynamics through four key modules: a spatial feature extraction (SFE) module for hierarchical spatial encoding, a temporal series modeling (TSM) module employing patch-based self-attention, a quarter-based seasonal encoding (QSE) module for suppressing interannual seasonal noise, and a deep supervision (DS) strategy that injected label-guided constraints into intermediate layers to enhance feature discriminability and interpretability. Experimental results across five diverse scenes in China demonstrated that TC-SSFN achieved state-of-the-art performance, with an average F1 score of 93.33% and IOU of 87.69%, consistently outperforming baseline methods. Ablation and clustering analyses further validated the complementary role of each module and revealed that supervision-induced feature structuring improved semantic separability. The quarter-based temporal segmentation strategy grouped multitemporal images by season and was combined with a differencing strategy. This combined approach effectively balanced sensitivity to both gradual and abrupt changes while preserving more meaningful change patterns. Overall, TC-SSFN provided a robust, interpretable, and generalizable framework for land cover change detection in complex, seasonally dynamic environments, offering practical value for large-scale monitoring and long-term remote sensing applications. The codes of our proposed TC-SSFN model are available at https://github.com/ChenyinD/TC-SSFN, and the datasets used in this study can be accessed at https://doi.org/10.5281/zenodo.18222086
Change detection (CD) holds significant value in fields such as environmental monitoring, disaster assessment, and resource surveys. Hyperspectral images (HSIs), with their rich spectral information, offer unique advantages for CD. The rise of deep learning has substantially propelled the flourishing development of HSICD. Model architectures have undergone continuous evolution, while learning paradigms have gradually expanded from supervised learning, which relies heavily on labeled data to encompass semisupervised and unsupervised approaches. Nevertheless, significant imaging differences, high redundancy and coarse spatial resolution, and the scarcity of annotations continue to pose challenges to the generalizability and robustness of the existing models. To systematically analyze intelligent CD technologies and challenges, this article comprehensively reviews deep learning-based intelligent HSICD methods, focusing on the comparative analysis of learning strategies under different supervision information. Furthermore, it consolidates widely used datasets and performance benchmarks within the field to provide reliable reference points for method evaluation while also exploring key future development directions. This review aims to offer theoretical underpinnings and directional guidance for further research in HSICD, thereby advancing the role of remote sensing (RS) change interpretation techniques.
Fine-grained ship recognition in remote sensing imagery is essential for maritime applications. However, its development is hindered by two challenges: 1) the limited granularity of existing ship detection datasets, and 2) the disturbance of complex maritime conditions as well as the arbitrary ship orientations and distributions. To address the first issue, we annotated a large-scale fine-grained ship instance detection dataset (LAFI), comprising 48,717 ship instances worldwide with 49 categories. To tackle the challenges of marine disturbance and diverse ship status, we proposed a controllable generative knowledge-driven ship detection framework (COSD). It employs a controllable diffusion model guided by ship-marine textual prompt to generate millions of synthetic images that not only preserve ship structures but also cover diverse sea and weather conditions for robust pretraining. The pretraining stage then utilizes masked reconstruction to learn component-level cues under occlusion, clutter, fog, and illumination changes. Furthermore, a heterogeneous feature alignment decoder is designed to align multi-modal metrics of orientation and distribution features in the latent space, allowing for accurate representation of diverse ship status. Extensive experiments on two benchmark datasets showed that our method respectively increased 0.011 and 0.030 mean average precision (mAP@50) over SOTA methods, particularly in scenarios involving small, densely packed and arbitrary oriented ships.
Ultra-high-resolution Synthetic Aperture Radar (SAR) and optical image registration is a fundamental prerequisite for multi-sensor remote sensing data fusion across various remote sensing applications. Ultra-high-resolution imagery reveals objects and details that are imperceptible at lower resolution. However, the intrinsic heterogeneity between SAR and optical imaging mechanisms emphasizes different visual attributes, thereby introducing new registration challenges manifested as more pronounced detail mismatches. To tackle the above problems, a global ultra-high-resolution SAR and optical image registration benchmark dataset (GUSO) is constructed. Covering more than 319 cities across 78 countries on six continents, the GUSO dataset provides spatial resolution between 0.16 m and 0.98 m and contains 593335 image pairs. It stands as the largest and highest-resolution benchmark to date for SAR and optical image registration, encompassing diverse scenes such as urban, rural, plain, hill, water, and disaster. Based on GUSO, the frequency-guided hierarchical registration method (FHReg) is proposed in this paper. To mitigate the disturbance of excessive details between ultra-high-resolution SAR and optical images, FHReg employs a wavelet-embedded feature extraction module to project the data into the wavelet domain, thereby achieving a unified details representation of different modality within a shared feature space. Furthermore, to progressively cope with large displacements and modality-induced ambiguities, the bidirectional coarse matching module establishes cross-modal correlations to generate forward and backward coarse matches, which are then refined by the localized fine matching module within neighborhood regions, ultimately producing high-confidence correct matches for transformation parameter estimation. Experimental results demonstrate that FHReg achieves robust registration performance across five common scenes in the GUSO dataset, achieving an RMSE of 2.79 pixels. Extending to unseen disaster scene in the GUSO dataset, FHReg demonstrates significant advantages in inference accuracy, achieving an RMSE of 3.08 pixels and a runtime of 0.04 s per patch, outperforming 12 state-of-the-art methods. Finally, when applied to real-world wide-swath disaster imagery, including both the Japan earthquake (26.48km2) and the United States wildfire (248.69km2), FHReg both accomplishes registration within 7 s and achieves an RMSE of better than 3 pixels, highlighting its precise registration capability and robust cross-domain generalization. The code and dataset used in this study are publicly available at https://github.com/vision-heng/GUSO.
Rapid post-disaster response requires timely, building-level information on whether structures remain intact, are damaged, or are destroyed. Post-event optical imagery, however, may be unavailable because of cloud, smoke, or darkness. The Bright Challenge evaluated all-weather building damage mapping from a submeter-resolution pre-event optical image and a post-event SAR image. Participants were required to detect and delineate each building and assign exactly one of three mutually exclusive damage labels. The challenge extended the globally distributed \textsc{Bright} dataset with instance-level annotations for about 291,000 buildings across 16 disaster events spanning seven disaster types. The final phase was evaluated exclusively on two 2025 events absent from training: a wildfire event in California and a hurricane in Jamaica. A total of 157 participants made 1,289 submissions, and 46 teams entered the final phase. The two winning solutions achieved test mAPs of 0.182 and 0.181, approximately 8.7 times the public baseline of 0.021, but remained far below the best in-domain holdout score of 0.513. Across teams ranked in both phases, performance declined sharply and the rank order changed substantially. The two leading solutions independently favored modality-specific encoding, staged or late optical--SAR fusion, and an optical-dominant separation of building localization from damage recognition. The winning method additionally used scene-aware threshold adjustment and pseudo-label adaptation. These results identify cross-event generalization and stable severity discrimination as the principal remaining challenges. All data, annotations, baseline code, and winning solutions are publicly available at https://github.com/ChenHongruixuan/BRIGHT.
This article presents the scientific outcomes of the 2024 Data Fusion Contest (DFC24) organized by the Image Analysis and Data Fusion Technical Committee of the IEEE Geoscience and Remote Sensing Society, the Space for Climate Observatory, the Centre national d'etudes spatiales, the National Aeronautics and Space Administration, and the Centre Europeen de Recherche et de Formation Avancee et Calcul Scientifique. The contest aims to advance image analysis and data fusion algorithms that generate reliable flood maps from multimodal Earth observation imagery. The DFC24 provides a large-scale, multimodal flood mapping benchmarking dataset and comprises two challenging competition tracks on the flood mapping task, one based on synthetic aperture radar imagery, and another using passive-optical imagery. Additional features, such as a digital terrain model and land-use and water occurrence, are also provided to the participants. This article presents the methods and results obtained by the first and second-ranked teams of each track. During the development phase, 1935 people registered for the contest, while at the end, 46 for Track 1 and 52 for Track 2 teams competed during the test phase in the two tracks, respectively. The data of this contest are openly available to the community for further research, development, and refinement of geospatial artificial intelligence, data fusion, and flood mapping methods.
Unmanned Aerial Vehicle (UAV) multispectral video object tracking is critical for real-world applications. While multispectral imaging offers complementary spectral cues beyond the visible range, tracking in aerial scenarios remains challenging due to data scarcity, suboptimal spectral-spatial modulation, and discrete sequential temporal modeling. To this end, we propose CASS, a context-aware memory framework with spectral-spatial modulation, which integrates spectral, spatial, and temporal cues for UAV multispectral tracking. CASS introduces two lightweight modules: (i) the efficient spectral-spatial modulation (ESSM) module, which modulates spatial representations through spectral-guided fusion, and (ii) the context state space reasoning (CSSR) module, which leverages evolving state space memory to retain long-term temporal cues and mitigate error propagation during cross-frame reasoning. By integrating these components in a parameter-efficient fine-tuning fashion, CASS achieves both efficient modulation and context-aware tracking. Evaluations on UAV multispectral benchmark (MUST) and ground-based hyperspectral benchmarks (NIR, RedNIR, VIS, MSSOT, MSVT) demonstrate CASS’s superior performance for both general hyperspectral tracking and specific UAV perception.
Due to the substantial domain gaps in Remote Sensing (RS) images that are characterized by variabilities such as location, wavelength, and sensor type, Remote Sensing Domain Generalization (RSDG) has emerged as a critical and valuable research frontier, focusing on developing models that generalize effectively across diverse scenarios. However, research in this area remains underexplored: (1) Current cross-domain methods primarily focus on Domain Adaptation (DA), which adapts models to predefined domains rather than to unseen ones; (2) Few studies target the RSDG issue, especially for semantic segmentation tasks. Existing related models are developed for specific unknown domains, struggling with issues of underfitting on other unseen scenarios; (3) Existing RS foundation models tend to prioritize in-domain performance over cross-domain generalization. To this end, we introduce the first vision foundation model for RSDG semantic segmentation, CrossEarth. CrossEarth demonstrates strong cross-domain generalization through a specially designed data-level Earth-Style Injection pipeline and a model-level Multi-Task Training pipeline. In addition, for the semantic segmentation task, we have curated an RSDG benchmark comprising 32 semantic segmentation scenarios across various regions, spectral bands, platforms, and climates, providing comprehensive evaluations of the generalizability of future RSDG models. Extensive experiments on this collection demonstrate the superiority of CrossEarth over existing state-of-the-art methods.
Remote sensing images acquired by unmanned aerial vehicles (UAVs) and satellites are often degraded by adverse weather, illumination variation, and imaging artifacts, which may co-occur and jointly induce global distribution shifts and local structural corruption. Although All-in-One image restoration offers an appealing unified alternative to task-specific pipelines, existing methods still suffer from weak or implicit degradation cues and parameter redundancy caused by full-rank multi-expert designs with overlapping restoration behaviors. We propose CoRE-UIR (Common and Residual Experts for Universal Image Restoration), a prior-guided global–local framework centered on the Common-and-Residual Expert Block (CoRE). CoRE explicitly decomposes restoration capacity into a common dense expert for degradation-invariant restoration and low-rank residual experts for degradation-specific compensation, enabling adaptive specialization without redundant expert replication. Built on this design, Degradation Prior Embedding (DPE) adapts frozen CLIP features into an explicit restoration-oriented prior, while Global Feature Modulation (GFM) aligns global feature statistics before local residual compensation. We also construct MDVD-108K (Multi-Degradation VisDrone), a large-scale UAV restoration dataset covering both single and compound degradations, together with a real-world test set. Extensive experiments on multiple datasets show that CoRE-UIR improves the overall average PSNR by 1.05 dB while running 11.83× faster and reducing peak memory by 85.3% relative to the strongest baseline, BaryIR, thereby maintaining a favorable quality-efficiency trade-off.Evaluations on downstream tasks and unseen degradation also validate the generalizability of CoRE-UIR. The code and dataset will be released at https://github.com/zzaiyan/CoRE-UIR.
Accurate building footprint databases are fundamental for sustainable urbanization yet face persistent updating challenges due to the rapid pace of urban change. Traditional methods rely on bi-temporal image comparison for change detection, which requires a large number of new labels to retrain the model, which is costly. We propose a passive updating paradigm that eliminates the reliance on historical imagery and leverages a lightweight adaptive strategy applied to Segment Anything Model (SAM) to minimize labeling costs. Furthermore, we propose a Cross Modal Temporal Fusion (CMTF) module that combines features from historical building footprints with those from recent imagery, alleviating the burden of small-sample training. The training process utilizes a semi-supervised approach, enabling the model to learn from both labeled and unlabeled regions, with labeled regions comprising only 0.4% of the building samples. Besides, we propose the RIO dataset, a sub-meter bi-temporal building footprint update dataset for studying building changes in rapidly developing areas. In addition, this work is validated on a range of cities worldwide, including Christchurch (post-earthquake reconstruction) and Beijing-Shanghai (megacity expansion). This work advances urban building renewal by overcoming the reliance on paired historical imagery for change detection and the need for large amounts of up-to-date labels. This approach offers a scalable solution for monitoring SDG 11 (Sustainable Cities and Communities), enabling less developed countries to use free and open product data to track urban expansion patterns with only a few labels.