Urban environmental indicators (UEI), such as building morphology, influence residential electricity consumption (REC) by directly reflecting baseline energy demand and indirectly modulating demand through the urban heat island (UHI) effect. However, existing research frequently overlooks the cascading causal relationships among UEI, UHI, and REC, leading to biased assessments of their mechanisms. To address this gap, this study employs a causal inference framework based on double machine learning (DML) to quantify the energy-heat response intensity and decompose the double pathways through which UEI influences REC. The results reveal multi-scale spatial heterogeneity in energy-heat response intensity (city-scale: 254.9-2531.9 MWh/ degrees C; plot-scale: 300-6100 MWh/ degrees C), with stronger responses observed in economically developed regions. Decomposition analysis reveals that direct pathways of the UEI on REC account for 72-86% of the total effect, while the indirect pathways mediated by UHI contribute 14-28%. By separating these cascading pathways, the study demonstrates that traditional linear regression, which neglects thermal mediation, underestimates the effect of UEI on REC by approximately 12% relative to DML. Scenario simulations further show that expanding green spaces can indirectly reduce REC by 2890 MWh through UHI mitigation. These findings offer actionable insights for implementing spatially differentiated strategies to optimize urban energy systems under increasing heat stress.
Urban renewal prediction is critical for sustainable development, yet existing methods often overlook complex spatial dependencies and lack explainability. This study proposes an explainable hierarchical graph network for urban renewal (URHGN) prediction. It employs a hierarchical graph to model building and community level spatial interactions with GNNExplainer to predict renewal potential and quantify driving factors. Applied to Beijing, URHGN achieves an F1-score of 0.855 f 0.012, outperforming traditional machine learning (0.700-0.775) and single-layer graph methods (0.789-0.810). Explainability analysis reveals that building-level features like floor (importance: 0.485 f 0.030) are primary drivers while community-level context like house price (0.343 f 0.003) provides essential supplements. Spatial relationships prove more influential than node features, with contribution scores of 0.271 f 0.019 and 0.124 f 0.017 respectively. The model identifies 15,990 buildings with very high renewal potential (scores > 0.8), advancing explainable GeoAI (XGeoAI) methodologies for evidence-based urban planning and establishing a foundation for future dynamic models incorporating temporal changes. The code is available at: https://github.com/kkxiaoqin/URHGN.
The transition towards climate-neutral mobility remains challenging for cities worldwide, particularly in balancing travel efficiency with emission reduction goals. This study develops a policy-oriented framework to optimize travel mode splits, demonstrating how behavioral adaptations can contribute to sustainable urban mobility without extensive infrastructure investments. Using Beijing as a case study, we formulate a multi-objective optimization model to identify optimal travel mode distributions between origin-destination pairs, considering both travel efficiency and carbon emissions. Our results reveal significant potential through strategic mode shifts: the optimization could reduce average travel time by up to 6.3 min while cutting carbon emissions by up to 384.1 tCO2 per trip. The effectiveness varies across urban contexts, with optimization potential ranging from 11 % to 43 %, suggesting targeted policy interventions. Different scenarios, prioritizing either emission reduction or efficiency improvement, help identify high-priority areas for implementing mode shift strategies. Sensitivity analyses demonstrate the framework’s robustness across various contexts, including scenarios of vehicle electrification. These findings provide evidence-based support for policymakers to design targeted interventions that effectively influence travel behavior towards climate-neutral mobility.
As core nodes of economic activity and consumption, urban commercial districts generate visitor flow dynamics that directly impact the operational efficiency of commercial facilities and the effectiveness of urban spatial planning. Origin-Destination (OD) flow prediction has been widely employed to accurately uncover urban spatial mobility patterns from residential areas to commercial districts, serving as a pivotal tool for elucidating the alignment between commercial attractiveness and residential demand. However, existing models often overemphasize geospatial features while neglecting the intrinsic socio-economic attributes of both commercial districts and residential communities. Consequently, these approaches fail to adequately characterize the supply-demand alignment, as they limit their scope primarily to spatial proximity and overlook the consumption similarity inherent in economic activities. To address this gap, this study proposes the Multi-Feature Spatial Representation Learning Network (MFSRNet), integrating commercial attractiveness, residential purchasing power, and demographic composition. Furthermore, unlike traditional functional similarity graphs that rely on static land-use attributes, we incorporate a Dual Graph Convolutional Network specifically designed to capture the latent economic homophily. This network concurrently learns spatial dependencies based on geographic proximity and consumption similarity derived from purchasing power and housing prices, thereby enabling a sophisticated simulation of individual commercial mobility decisions. Extensive experiments conducted on real-world data from Beijing demonstrate that MFSRNet consistently outperforms baseline models by at least 13.83%. Overall, this study presents a competitive model with enhanced performance and strong generalizability for commercial OD flow prediction. It offers data-driven decision-making support to enterprises for market positioning, business optimization, and precision marketing, while aiding urban planning authorities in optimizing the allocation of commercial resources so as to promote balanced regional economic development.
Semantic segmentation plays a crucial role in numerous remote sensing (RS) applications. Despite the success of multimodal RS segmentation models, integrating large-scale visual priors from foundation models with multimodal information remains challenging due to modality incompatibility and increased computational costs. To address this, we propose MmSAM, an efficient fine-tuning framework that applies Segment Anything Model 2 (SAM2) to multimodal RS semantic segmentation. Unlike traditional feature fusion paradigms in multimodal segmentation, we do not treat additional modalities as equal inputs to the main modality but as prompts for it. We employ the Mixture-of-Experts (MoE) mechanism to construct hard- and soft-MoE as multimodal prompters, sparsifying the model architecture while extracting multimodal features, effectively controlling the computational load. Additionally, we introduce several fine-tuning methods to enhance the performance of the SAM2 image encoder and perform end-to-end modification to better adapt the model to downstream tasks. Experimental results on two public multimodal RS datasets demonstrate that MmSAM significantly outperforms the single-modal SAM2 baseline by similar to 2.5% and similar to 2.2% in mean intersection over union (mIoU), respectively. Furthermore, MmSAM achieves state-of-the-art performance with lower computational cost, making it highly suitable for consumer-level deployments. The code will be available at: https://github.com/W-qp/MmSAM.
As a key component of meteorological disasters, precipitation underpins decision-making in economic and social sectors dependent on weather data. Mitigating its socio-economic impacts requires long-lead-time, high-resolution forecasting, with spatio-temporal feature learning being central to this task. Recently, spatio-temporal graph neural networks (GNNs) have emerged as proliferating models for capturing meteorological dynamic patterns. However, existing precipitation forecasting methods face two major challenges: 1) GNNs rely on pairwise relationships, limiting their ability to capture dynamic high-order spatio-temporal interactions critical to climate evolution and 2) precipitation's nonstationary, highly fluctuating nature discerns long-term temporal patterns. To address these issues, this article proposes a spatio-temporal dynamic hypergraph neural network (STDHGNN) for large-scale, long-term online precipitation forecasting. It develops an adaptive graph and hypergraph generation method to simulate precipitation aggregation and multidirectional diffusion (driven by atmospheric circulation, urban agglomerations, or topographic lifting, etc.), and a dynamic replay mechanism to adapt time and frequency domains for online learning of long-term precipitation's key frequencies and periodicity. Extensive experiments on 26-year datasets covering five nations show STDHGNN outperforms state-of-the-art for up to 30-day lead times. This work provides new insights for re-evaluating precipitation's spatio-temporal patterns and innovates high-order interaction integration in deep learning (DL).
Abstract The rapid decline of Arctic sea ice requires accurate prediction of its concentration (SIC) and thickness (SIT). We introduce IceCT, a deep learning model using a Mixture‐of‐Experts (MoE) framework to generate monthly SIC and SIT forecasts at a 12.5 km resolution up to 6 months ahead. Its architecture features two specialized experts—one for SIC, one for SIT—enhanced with Multi‐Head Self‐Attention (MHSA) to capture spatial patterns. A novel Spatial Gate adaptively weights the experts' contributions across the grid. Evaluations show IceCT outperforms other deep learning baselines, achieving an overall Mean Absolute Error (MAE) of 4.56% for SIC and 0.140 m for SIT, and achieves enhanced predictive skill in forecasting sea ice during the crucial summer melting period. Additionally, the model mitigates the Arctic spring predictability barrier by learning to rely more on SIT information for forecasts initialized in spring.
Urban vitality is crucial for measuring urban sustainable development. Despite substantial progress in unraveling the spatial distribution of urban vitality, research on its temporal forecasting is less investigated. A significant challenge in urban vitality forecasting arises from the proliferation of spatial nodes, which leads to a reduction in the distinguishability among nodes. This phenomenon reduces the ability of transformer attention to detect node-specific differences, thereby decreasing the forecasting performance. To address this, we propose a spatial heterogeneity-aware deep learning framework for urban vitality forecasting, namely Sphereformer. Take advantage of practice from the geographical sciences, an angular-based attention module was proposed to deal with homogeneous noise on large-scale spatial areas, thereby improving the differentiation representation of the model. Additionally, we use Point-of-Interests (POI) as a covariate to further increase differentiation, thereby better achieving spatial heterogeneity modeling. Experiments demonstrate our model outperforms state-of-the-art models on a large-scale real-world dataset of urban vitality. This research provides a new point for integrating geographical statistics with deep learning for urban vitality forecasting.
Against rapid urbanization, ensuring carbon emission equity is critical for inclusive sustainable development. Using Beijing's smart card data, we develop a bottom-up accounting model to quantify travel carbon emissions in public transportation at the Traffic Analysis Zone scale, assessing equity through three dimensions: vertical (Palma ratio), horizontal (Gini coefficient), and spatial (Carbon Burden Index). Results show that the stratification and efficiency differences in public transportation services drive emission characteristics. Subway travel time is approximately twice that of buses, and the travel distance is 2.5-3.4 times longer, resulting in 80.9% higher per capita emissions. Weekdays, weekends, and holidays reflect distinct carbon emission patterns driven by 'commute-driven,' 'mixed demand,' and 'leisure-oriented' travel behaviors, respectively. During weekday peak periods, low-income groups bear nearly half of the travel emission burden, with per capita emissions 69% higher than those of high-income groups. Marginalized groups bear hidden environmental costs for urban economic efficiency. Commuting demand and jobs-housing spatial mismatches exacerbate equity disparities. Substituting short trips with shared bikes reduces emissions by 2%. The study reveals the socioeconomic and spatial differentiation mechanisms of urban public transportation travel carbon emissions. These findings provide a scientific basis for formulating low-carbon transportation policies that balance equity and efficiency.
Climate change is intensifying urban heat risks, making green infrastructure crucial for heat-stress mitigation. Yet most assessments focus on greenspace coverage rather than realized cooling services, and rarely examine the divergence between the two, particularly in rapidly urbanizing countries. Using remote sensing data and spatial modeling across 29 Chinese cities, we quantified greenspace coverage, cooling efficiency (CE) and cooling capacity at 1-km resolution. Realized cooling capacity did not simply mirror greenspace coverage; spatially heterogeneous CE reshaped how green resources were translated into cooling services. This coverage–capacity divergence was unevenly associated with social groups: higher-price neighborhoods showed greater cooling capacity, areas with larger elderly populations showed weaker alignment in half of the cities, and areas with more children generally experienced more favorable conditions. Allocation simulations further revealed trade-offs between aggregate cooling and distributional outcomes. These findings show that equal greenspace coverage does not guarantee equitable cooling outcomes, shifting assessment from distributional inequality toward service inequity.
Dockless bike-sharing is widely promoted for carbon reduction benefits by replacing other travel modes. However, studies evaluating bike-sharing’s environmental impact often overlook travel distance’s influence on mode choice and lifecycle emissions, leading to systematic overestimation. This study incorporates these overlooked factors using Monte Carlo simulation and life cycle assessment in Shenzhen, China. Results reveal that neglecting distance and embodied emissions overestimates carbon reductions by 44.9%, with overestimation reaching 60% under vehicle electrification scenarios. While bike-sharing achieves meaningful carbon reductions (163.19 tons daily in Shenzhen, 95% CI: [161.69, 164.68]), the substantial overestimation in existing studies undermines evidence-based policymaking. Our findings demonstrate that short-distance trips - the primary use case for bike-sharing - show the most significant overestimation. These results urgently call for policymakers to reassess bike-sharing environmental benefits using more comprehensive methodologies, particularly as transportation electrification reduces the relative environmental advantages of bike-sharing systems.
Abstract Urban greenspace inequality has become a major sustainability challenge, yet the roles of different greenspace types in mitigating exposure inequality remain poorly understood, even as cropland is increasingly being incorporated into urban landscape thinking. Here, we developed a population‐weighted greenspace exposure framework that integrates greenspace composition, seasonal vegetation condition, and population distribution to assess landscape greenspace (LGS) and cropland exposure across 320 Chinese cities from 2010 to 2020. Monthly NDVI, land‐cover data, and population grids were combined to quantify human greenspace exposure, and the Gini index was used to measure exposure inequality. Our results show that cropland is an important but long‐overlooked component of urban greenspace exposure. Although average LGS exposure remained higher nationally, cropland served a larger share of the population in nearly half of the sampled cities, with mean exposure reaching about 82% of that of LGS. Under dynamic urban boundaries, cropland showed a significant compensatory effect in 2010, 2015, and 2020, reducing LGS exposure inequality by about 30% on average. Annual association analyses further indicate that cropland exposure was more strongly linked to lower inequality than LGS exposure; by 2020, its mitigating contribution was more than four times larger. However, the effects of both greenspace types weakened over the decade. Fixed‐boundary analysis shows that this compensation reflected both a structural spatial pattern and the compositional effects of boundary expansion. These findings indicate that cropland compensation depends mainly on spatial allocation, highlighting the need to combine targeted LGS provision with selective cropland retention and multifunctional integration.
Hyperspectral image (HSI) classification models are highly sensitive to distribution shifts caused by real-world degradations such as noise, blur, compression, and atmospheric effects. To address this challenge, we propose HyperTTA (Test-Time Adaptable Transformer for Hyperspectral Degradation), a unified framework that enhances model robustness under diverse degradation conditions. First, we construct a multi-degradation hyperspectral benchmark that systematically simulates nine representative degradations, enabling comprehensive evaluation of robust classification. Based on this benchmark, we develop a Spectral-Spatial Transformer Classifier (SSTC) with a multi-level receptive field mechanism and label smoothing regularization to capture multi-scale spatial context and improve generalization. Furthermore, we introduce a lightweight test-time adaptation strategy, the Confidence-aware Entropy-minimized LayerNorm Adapter (CELA), which dynamically updates only the affine parameters of LayerNorm layers by minimizing prediction entropy on high-confidence unlabeled target samples. This strategy ensures reliable adaptation without access to source data or target labels. Experiments on two benchmark datasets demonstrate that HyperTTA outperforms state-of-the-art baselines across a wide range of degradation scenarios. The code will be made publicly available at https://github.com/halfcoder1/HyperTTA.
As the spatial resolution of remote sensing imagery continues to improve, Earth observation scenes exhibit greater diversity in spatial and spectral characteristics, together with increasingly complex surface structures across multiple scales. In such scenarios, global land-cover patterns and local high-frequency details are often strongly coupled, posing significant challenges for accurate scene interpretation and boundary delineation. From a feature representation perspective, spatial-domain features are effective for modeling global context and high-level semantics, whereas frequency-domain representations facilitate the separation of structural components at different scales and enhance fine boundary details, but have limited capability in explicitly capturing semantic relationships. Consequently, most existing semantic segmentation methods operate in a single domain, either spatial or frequency, which restricts their ability to jointly model semantic consistency and diverse spatial structures. To address this issue, we propose a semantic segmentation network termed Spatial–Frequency Collaborative Modeling Network (SFMNet), which jointly learns complementary representations in both spatial and frequency domains. By leveraging spatial-domain semantic priors to guide frequency-domain structural enhancement, SFMNet enables coordinated modeling of land-cover semantics and multi-scale surface structures. In addition, cross-scale interaction between low- and high-frequency components and adaptive multi-level feature fusion are introduced to improve structural consistency across hierarchical feature representations. Extensive experiments on the ISPRS Potsdam, Vaihingen, and LoveDA datasets demonstrate that SFMNet consistently outperforms existing methods in both overall segmentation accuracy and boundary delineation quality. Specifically, SFMNet achieves mean Intersection-over-Union scores of 81.30%, 70.23%, and 55.07%, and Boundary mean Intersection-over-Union scores of 65.42%, 56.76%, and 43.74% on the three datasets, respectively.
Recent advances in deep learning and multimodal data fusion technologies have significantly enhanced hyperspectral image (HSI) classification performance. Nevertheless, classification accuracy of hyperspectral data continues to degrade substantially under diverse degradation scenarios, such as noise interference, spectral distortion, or reduced resolution. To robustly address this challenge, this paper proposes a novel cross-modal guided classification framework that integrates active remote sensing data (e.g., LiDAR) to improve classification resilience under degraded conditions. Specifically, we introduce a Cross-Modal Feature Pyramid Guidance (CMFPG) module, which effectively utilizes cross-modal information across multiple levels and scales to guide hyperspectral feature extraction and fusion, thereby enhancing modeling stability in degraded environments. Additionally, we develop the HyperGroupMix module, which enhances cross-domain adaptability through grouping spectral bands, extracting statistical features, and transferring features across samples. Experimental results conducted under complex degradation conditions demonstrate that our proposed method exhibits stable high-level classification accuracy and robustness in overall performance. The code is accessible at: https://github.com/miliwww/CMGF
Real-world aerial image super-resolution (SR) remains particularly challenging because degradations in remote-sensing imagery involve random combinations of anisotropic blur, signal-dependent noise, and unknown downsampling kernels. Most existing SR methods either rely on simplified degradation assumptions or lack semantic perception of degradation, resulting in limited generalization to real-world conditions. To address these gaps, we propose a novel diffusion-based SR framework that integrates Multi-modal Large Language Models (MLLMs) and self-supervised contrastive learning for extracting degradation-insensitive representation. Specifically, we introduce a contrastive learning strategy into a ControlNet module, where the HR and LR counterparts of the same image are regarded as positive pairs, while representations from different images serve as negative pairs, enabling the network to learn degradation-insensitive structural features. To further enhance semantic awareness of degradation, an MLLM-generated change caption is incorporated into the diffusion process as textual guidance, allowing the model to explicitly perceive and reconstruct different degradation types. Moreover, a classifier-free guidance (CFG) distillation strategy compresses the original dual-branch diffusion model into a single lightweight network, substantially improving inference efficiency while maintaining high reconstruction fidelity. Extensive experiments conducted on various datasets have showcased the superior performance of our proposed model compared to existing state-of-the-art methods. Furthermore, our distillation algorithm achieves a twofold reduction in inference time compared to its non-distilled counterpart, making it more feasible for real-time and resource-limited applications.
Spatial processes are fundamental to understanding complex patterns and dynamics across diverse systems but remain challenging to model due to the competing demands of predictive accuracy and interpretability. Traditional spatial regression techniques such as geographically weighted regression (GWR) provide interpretable, location-specific coefficients yet systematically underfit complex, high-dimensional, and nonlinear spatial relationships prevalent in modern geospatial data sets. In contrast, machine learning methods achieve higher prediction accuracy but sacrifice explicit quantification of spatially varying relationships and compatibility with spatial model evaluation criteria. To address these limitations, we propose GWRBoost, an ensemble learning framework that integrates gradient boosting with GWR models to simultaneously improve model performance and enhance the interpretability of coefficient estimates across space. GWRBoost uses a stage-wise residual passing mechanism to leverage spatial dependence and optimize model performance, globally preserving intrinsically interpretable, location-specific linear coefficients that support coefficient-based local explanations. We demonstrate the effectiveness of GWRBoost through synthetic and empirical case studies including regional, national, and global scales in the domains of computational social science, ecology, and public health. Our results show that GWRBoost consistently outperforms traditional global regression models and leads spatially varying methods, delivering superior predictive accuracy and more precise estimation of spatially heterogeneous relationships. Importantly, GWRBoost retains compatibility with model selection metrics, enabling robust comparative evaluation. GWRBoost provides a scalable, interpretable, and generalizable approach for spatial process modeling, bridging the gap between accuracy and explainability, and offering significant value to a broad range of spatial disciplines.
While urban mitigation strategies have historically focused on operational energy, the embodied carbon emissions of the built environment remain a critical yet under-measured challenge. This study develops a multi-scale, high-resolution framework for quantifying and analyzing embodied carbon emissions from the urban built environment at both inter-city and intra-city scales, using 35 Chinese cities as a case study. Inter-city analysis indicates pronounced heterogeneity in total embodied carbon emissions, ranging from 45 Mt in Lhasa to 2103 Mt in Beijing. Intra-city analysis using 500 m × 500 m grids shows pronounced spatial heterogeneity, with high-carbon zones concentrated in central urban functional areas. Land-use assessment indicates that residential land accounts for the largest share of embodied carbon (43.5%), followed by transportation (18.7%) and industrial land (15.5%), while commercial and public administration lands contribute 10.2% and 12.1%, respectively. Furthermore, bivariate spatial analysis shows that the relationship between built environment carbon emissions and GDP, population, and green space coverage is significantly heterogeneous across cities. These findings provide a granular evidence base for precision carbon management and offer a scalable tool for integrating embodied carbon into urban environmental impact assessment and planning.
Local leisure events satisfy personalized needs and boost local tourism vitality. However, research on leisure event preference recommendations remains inadequate. Event-Based Social Networks (EBSNs) combine online social and offline interactions, but users face information overload, requiring efficient recommendation systems. Current systems face two key challenges: cold-start problems for new users/events and data sparsity. Research shows social relationships help mitigate cold-start issues, while users' interests and social connections change over time, with recent behaviors being more predictive than long-term ones-a fact often overlooked. To address these issues, we propose ERDGAT, a dynamic graph attention network model for event recommendations in EBSNs. The model extracts event features, mines user preferences from historical events, models social relationships using graph attention networks, and captures recent preference features through temporal social networks with long short-term memory networks. Experiments on the Douban Events dataset demonstrate ERDGAT significantly outperforms baseline methods in recommendation accuracy and cold-start mitigation, improving NDCG@10 by 26.5%.