Urban region representation learning has emerged as a fundamental approach for diverse urban analytics tasks, where each neighborhood is encoded as a dense embedding vector for effective downstream applications. However, existing approaches suffer from insufficient multi-modal alignment and inadequate spatial relationship modeling, limiting their representation quality and generalizability. To address these challenges, we propose UrbanMMCL, a novel self-supervised framework that integrates multi-modal multi-view contrastive pre-training with unified fine-tuning for comprehensive urban representation learning. UrbanMMCL employs a dual-stage architecture. First, cross-modal contrastive learning aligns diverse data modalities including remote sensing imagery, street view imagery, location encodings, and Vision-Language Model (VLM)-generated textual descriptions. Second, multi-view adaptive graph contrastive learning captures complex spatial relationships across human mobility, functional similarity, and geographic distance perspectives. The framework then fine-tunes all parameters with the learned representations for effective adaptation to downstream tasks. Comprehensive experiments demonstrate that UrbanMMCL consistently outperforms state-of-the-art methods across pollutant emission prediction, population density estimation, and land use classification with minimal fine-tuning requirements, thereby advancing foundation model development for diverse Geo-AI applications.
Human mobility is closely associated with regional economies and demographics. Cellular automata (CA)-based models are widely used for urban simulation and provide a powerful framework to support sustainable development of urban agglomerations under increasingly dynamic conditions. However, the explicit integration of human mobility into CA-based simulations of urban agglomerations remains limited. Here, we proposed a human mobility-enhanced cellular automata (HME-CA) model to simulate the dynamic evolution of economies and demographics in urban agglomerations. The human mobility intensity and mobility-based neighborhoods were generated from large-scale mobile phone positioning data to capture the tele-coupling effect as flow-based connectivity. The contribution of human mobility to urban simulation was estimated by comparing with four random forest-based CA models. Taking the Pearl River Delta (PRD) region as a case, these results showed that the HME-CA model outperformed baseline approaches, yielding lower mean absolute percentage errors (16.00% for demographic and 24.00% for economic simulations) and root mean square error (1.02 for demographic and 2.96 for economic simulations). Moreover, human mobility-based neighborhoods can enhance economic and demographic simulations, particularly in less developed areas. The findings highlight the pivotal role of human mobility in shaping urban simulations and provide valuable insights to support informed decision-making for urban agglomerations.
Rising urban temperatures are increasing heat-related risks for active transportation users, particularly cyclists. Yet the hyperlocal heat burden experienced by bike-share users remains poorly understood. This study integrates 1-m resolution hourly Universal Thermal Climate Index (UTCI) maps with 1,189,667 Citi Bike trips across 767 stations in New York City during August 2018. Using generalized linear models, K-means clustering, and non-parametric tests, we examine how user attributes, trip characteristics, and temporal patterns are associated with accumulated heat stress, defined as route-mean UTCI multiplied by trip duration. Short-term Customers and weekend riders accumulate higher heat stress than Subscribers and weekday riders, primarily because of longer trip durations, although route-level thermal intensity also contributes. Gender- and age-related differences are also observed, but these patterns reflect cumulative thermal burden rather than thermal intensity or physiological vulnerability alone. Clustering identifies three behaviorally distinct trip profiles: High-Burden Long-Duration Riders, Low-Burden Routine Riders, and Standard Weekday Riders. These findings show that cyclist heat burden is shaped by both urban microclimates and travel behavior, offering evidence for prioritizing shade, cooling infrastructure, and heat-risk communication in bike-share systems.
Urban soundscapes shape residents' health and well-being while reflecting the geographic and functional structure of cities. However, existing approaches often underperform in small-sample settings, insufficiently integrate heterogeneous urban datasets, and rely on opaque "black-box" models, which limits their planning utility. To fill these gaps, an interpretable semi-supervised framework is proposed for fine-grained urban soundscape mapping and mechanism analysis that fuses multi-source geospatial data. First, a multi-dimensional indicator system is constructed by combining field sound measurements with multi-source geospatial data. Second, a heterogeneous co-training regressor (HetCoReg) is developed for sound pressure level (SPL) and dataaugmented Random Forest (RF) for sound source classification. Finally, we incorporate Shapley Additive Explanations(SHAP) to quantify global and local feature effects, and conduct spatial analysis. Using Shekou, Shenzhen as a case study, the HetCoReg model achieved R2 = 0.87 for SPL, while the augmented RF model reached precision = 0.93 for source classification. Results show that traffic and human sounds dominate, natural sounds cluster in parks, and mechanical sounds concentrate near industrial zones. Road density and green space are identified as key drivers. Overall, the framework enhances prediction under data-scarce conditions and provides actionable insights for soundscape-informed planning and public health protection.
Deformation monitoring of long-span bridges is essential for evaluating their structural health and safety. Traditional methods are labor-intensive and low-frequency. Image-based approaches offer precise deformation results but deteriorate with distance; thus, they cannot meet the requirements of long-span bridges. To fill these gaps, this study proposes a real-time deformation monitoring approach by integrating cameras with inertial sensors (RTDCI). Cameras capture the relative deformation of targets using computer vision. Multiple cameras are serially integrated to extend the measurement range. Inertial sensors compensate for camera motion errors alongside structural deformation. A multi-camera deformation measurement model is developed to process both images and inertial data. Laboratory and field experiments were conducted. Results demonstrate that RTDCI achieves an error of 0.32 mm with a 30 Hz frequency and 34 frames per second (FPS) in a vibrating laboratory environment. Field experiments exhibit millimeter-level accuracy when compared with hydrostatic leveling and ground-based interferometric radar. It offers an alternative deformation monitoring approach for long-span infrastructures and can be extended to complex construction environments.
Origin-Destination (OD) flow matrices are critical for urban mobility analysis, supporting traffic forecasting, infrastructure planning, and policy design. Existing methods face two key limitations: (1) reliance on costly auxiliary features (e.g., Points of Interest, socioeconomic statistics) with limited spatial coverage, and (2) fragility to spatial topology changes, where reordering urban regions disrupts the structural coherence of generated flows. We propose Sat2Flow, a structure-aware diffusion framework that generates structurally coherent OD flows using only satellite imagery. Our approach employs a multi-kernel encoder to capture diverse regional interactions and a permutation-aware diffusion process that maintains consistency across regional orderings. Through joint contrastive training linking satellite features with OD patterns and equivariant diffusion training enforcing structural invariance, Sat2Flow ensures topological robustness under arbitrary regional reindexing. Experiments on real-world datasets show that Sat2Flow outperforms physics-based and data-driven baselines in accuracy while preserving flow distributions and spatial structures under index permutations. Sat2Flow offers a globally scalable solution for OD flow generation in data-scarce environments, eliminating region-specific auxiliary data dependencies while maintaining structural robustness for reliable mobility modeling.
Traffic forecasting is a critical task in intelligent transportation systems, requiring accurate modeling of spatiotemporal dependencies among traffic sensors. Traditional deep-learning methods face two key challenges: (1) decoupled spatial-temporal pipelines that process and fuse spatial and temporal dimensions separately fail to capture their intricate interdependencies; and (2) state-of-the-art (SOTA) models relying on Transformer architectures often struggle to balance computational efficiency with representational capacity. To address these limitations, we propose ST-Camba, a novel decoupled-free spatiotemporal graph fusion state space model that unifies spatial and temporal dimensions within a single framework. ST-Camba is the first to integrate a spatial dimension axis into state space equations, enabling effective coupled spatiotemporal modeling through graph convolutions while inheriting the linear complexity advantage of Mamba series models. Additionally, we design an Adaptive Spatial Structure (ASS) Injector and a Lerp-based Gated Unit (LGU) to facilitate adaptive spatial structure capture and control information flow in spatiotemporal modeling. Extensive experiments on flow and speed prediction tasks across standard datasets demonstrate ST-Camba's superiority. Specifically, on the PEMS07 dataset, our model achieves a 1.8% reduction in MAE compared to other baselines, while reducing computational costs by up to 14.5%. This work underscores the necessity of coupled spatiotemporal modeling and provides a theoretical foundation for scalable solutions in urban traffic systems.
Equitable access to homeless shelters is a critical urban and transportation challenge, particularly for vulnerable populations who often rely on non-motorized travel. This study introduces a Gaussian Two-Step Floating Catchment Area (G2SFCA) method to evaluate nighttime shelter accessibility across the City of Los Angeles. We then utilize Ordinary Least Squares (OLS) and Geographically Weighted Regression (GWR) to investigate the spatially varying relationships between accessibility and neighborhood-level socioeconomic factors. Results reveal distinct spatial inequality, with “service deserts” concentrated in the city’s affluent western and northern areas, while shelter services are clustered in historically disadvantaged southern communities. The GWR results confirm that the influence of key predictors, such as housing costs, public assistance enrollment, and demographic composition, is not uniform, but spatially heterogeneous, with some relationships even inverting across the city. These findings suggest that shelter access is shaped by a complex set of local socioeconomic conditions, highlighting the limits of global models. These findings offer empirical support for developing place-based, spatially sensitive policies to address service gaps for vulnerable urban populations.
Mountainous areas are crucial repositories of natural resources and biodiversity hotspots but remain ecologically vulnerable. With the acceleration of urbanization, impervious surfaces have progressively expanded from flat terrain to hillside terrain, forming unique hillside urbanization landscapes. Although hillside urbanization alleviates the pressure of cultivated land loss in plains and enhances population carrying capacity, its vertical warming effect (temperature increases along hillside areas) has not been systematically investigated. Taking the Guangdong-Hong Kong-Macao Greater Bay Area (GBA) as a case study, this study integrates multi-source remote sensing data to precisely extract high-resolution impervious surfaces and land surface temperature (LST). We developed a methodological framework to quantify the vertical warming effect of hillside urbanization and analyzed its spatiotemporal characteristics from 2000 to 2020. Our results showed that 21.40% of impervious surface expansion areas occurred on hillside areas (slopes above 5 degrees) in the GBA, primarily occupying croplands and forests. The mean vertical warming effect in hillside areas increased from 1.42 degrees C to 1.64 degrees C, with the area affected by warming expanding by approximately 206%. In addition, the warming effect showed significant correlations with the landscape patterns of the size, proportion, and aggregation of impervious surface patches. Our findings provide quantitative evidence to support the scientific implementation of hillside development policies and offer theoretical insights for regional thermal environment mitigation strategies.
Urban freight transport is becoming increasingly complex under rapid urbanization. Understanding spatio-temporal distribution of truck dwelling behavior is essential to improve logistics efficiency and urban sustainability. Yet, the freight activity patterns are not well captured by static representation methods. This study introduces a contrastive graph representation learning framework to characterize the spatio-temporal and behavioral features of truck dwelling locations using large-scale GPS trajectory data. We construct a time-varying directed truck flow network. Building on edge convolution, we develop a modified Edge Convolution Network (ECN) with attention-based aggregation to learn spatial dependencies and temporal dynamics. These embeddings are used to delineate logistics activity zones through unsupervised clustering. The framework is tested and evaluated with a massive truck trajectory dataset and compared with other baseline methods. The results prove that the modified ECN model achieves better performance than baselines and reveals interpretable spatio-temporal activity patterns providing valuable insights in logistics amelioration.
Urban sewer pipelines are prone to diverse faults, such as cracks, erosion, and root intrusion. Effective and efficient inspection methods are essential for large-scale urban sewer pipe networks. This paper presented a collaborative inspection approach to inspect urban sewer pipes, which integrates robotic pipe capsules (RPCs) with lightweight deep learning and spatial optimization. A bi-level network is built to represent diverse movements of workers and the RPCs and their collaboration. A specialized lightweight deep neural network is designed to identify faults with images captured by PRC in real time. The worker and RPC routes are spatially optimized with hybrid meta-heuristics. An experiment in Shenzhen, China, demonstrated that it achieves a balanced accuracy of 83.43% with 7.64 frames per second, which outperforms baseline methods. The presented method provides an alternative approach for large-scale urban sewer pipe networks.
Accurately predicting fine-grained urban mobility is essential for optimizing transportation, accessibility, and urban management. However, existing approaches often depend on dynamic data such as trajectories or signaling records, which are sparsely available across cities, thereby limiting their applicability and generalizability to new urban contexts. To address these limitations, this study proposes a Large Model Enhanced Multimodal Representations (LMEMR) framework to learn hourly grid-level mobility dynamics solely from static geospatial data-including remote sensing imagery, building data, street view imagery, and points of interest-which are widely accessible. Large vision-language models are employed to generate natural-language descriptions of each modality, enriching the data with human-understandable semantics. A dual-level contrastive learning strategy aligns raw and textual features both within and across modalities, mitigating semantic gaps and enhancing multimodal consistency. Spatial dependencies are modeled through a graph attention network, and temporal dynamics are captured via a transformer encoder to produce 24-hour mobility sequences. Results from Shenzhen demonstrate that LMEMR outperforms the baseline CLIP model, achieving an R-2 of 0.856and an 18.07% reduction in MAE. Ablation experiments confirm the effectiveness of semantic enhancement, spatial graph reasoning, and cross-modal fusion. Overall, this research reveals the potential of static multimodal data for dynamic mobility inference, offering a scalable, interpretable, and privacy-friendly solution for smart city planning and management.
Hospital bypass behaviour, where patients forgo nearby hospitals for distant providers, challenges the efficiency of healthcare systems. However, its socioeconomic status (SES) drivers and equity implications remain poorly understood. Here we analyse 96 million mobile phone users across 11 Chinese cities integrated with behavioural models quantifying the trade-off between spatial proximity and hospital quality. Results show 80.63% of patients bypass their nearest hospital with 211.84% increased travel distance. Low-SES patients exhibit paradoxical bypass patterns with 14.01% lower bypass rate but stronger bypass intensity (skip 8.02% more hospitals) than high-SES patients. Behavioural modelling reveals that low-SES patients have 5.9-fold stronger willingness to travel for hospital quality. Their intense travel burden represents a compensatory mechanism to overcome spatial quality deficits, and this preference-driven spatial sorting unintentionally reduces 42.15% experienced social segregation within hospitals. Our findings highlight the need to address the structural spatial mismatch in medical quality rather than relying on geographic proximity or access restrictions only. Using data from 96 million mobile phone users across 11 cities in China, this analysis sheds light on the trade-off between spatial proximity and hospital quality in patterns of healthcare utilization, showing that 80.63% of patients bypass their nearest hospital with 211.84% increased travel distance.
Drive-by vehicle-borne mobile sensing with third-party vehicles has the advantages of high precision, low cost, appropriate coverage, and high timeliness, when compared to satellite-based, Unmanned Aerial Vehicle (UAV)-borne, or ground-based station monitoring. However, the non-prescriptive (or even unpredictable) behaviors of third-party vehicles can lead to imbalanced sampling. To obtain a good mobile sensing scheme with a maximum spatial coverage and a minimal spatial sampling heterogeneity, this study selected the best hybrid bus-taxi fleet to install sensors by proposing cooperative mobile sensing optimization models. As the traveling behavior patterns of taxis and buses mined from a huge amount of historical data were fully given consideration into optimization models, the sensors installed on the proposed mobile sensing taxi-bus fleet could automatically collect data, achieving the largest urban spatial range without operational intervention. Experimental results demonstrated the benefits of our solution in terms of global sensing ratios, sensing heterogeneity, cost savings, and geographical sampling equality. The proposed models can be used to monitor a variety of urban environmental objects, including air pollutants, noise, road roughness, and urban 3D scenes.
As extreme heat events intensify due to climate change and urbanization, cities face increasing challenges in mitigating outdoor heat stress. While traditional physical models such as SOLWEIG and ENVI-met provide detailed assessments of human-perceived heat exposure, their computational demands limit scalability for city-wide planning. In this study, we propose GSM-UTCI, a multimodal deep learning framework designed to predict both average and hourly Universal Thermal Climate Index (UTCI) at 1-meter hyperlocal resolution across daytime hours. The model fuses surface morphology (nDSM), high-resolution land cover data, and hourly meteorological conditions using a feature-wise linear modulation (FiLM) architecture that dynamically conditions spatial features on atmospheric context. Trained on SOLWEIG-derived UTCI maps, GSM-UTCI effectively emulates the physical model, achieving an R-2 of 0.9151 and a mean absolute error (MAE) of 0.41 degrees C for daytime average UTCI prediction. Across 11 daytime hours (8 a.m.-6 p.m.), it maintains robust hourly performance with an average R-2 of 0.8044 and an average MAE of 0.64 degrees C. Additionally, GSM-UTCI reduces inference time from days to under five minutes for an entire city, enabling rapid city-wide simulations. To demonstrate its planning relevance, we apply GSM-UTCI to simulate systematic landscape transformation scenarios in Philadelphia. Results demonstrate clear and spatially heterogeneous cooling effects. Notably, converting impervious surfaces to tree canopy leads to the largest reduction in heat exposure, lowering the number of residents experiencing strong heat stress (UTCI > 32 degrees C) by over 374,000, with an average UTCI decrease of 4.18 degrees C in affected areas. Tract-level analysis further reveals strong alignment between thermal reduction potential and land cover proportions. These findings highlight the value of the GSM-UTCI framework as a scalable decision-support tool for climate adaptation in Philadelphia, while its extension to other cities remains a key direction for future validation.
Street view imagery (SVI), with its rich visual information, is increasingly recognized as a valuable data source for urban research. Particularly, by leveraging computer vision techniques, SVI can be used to calculate various urban form indices (e.g., Green View Index, GVI), providing a new approach for large-scale quantitative assessments of urban environments. However, SVI data collected at the same location in different seasons can yield varying urban form indices due to phenological changes, even when the urban form remains constant. Numerous studies overlook this kind of seasonal bias. To address this gap, we propose a systematic analytical framework for quantifying and evaluating seasonal bias in SVI, drawing on more than 262,000 images from 40 cities worldwide. This framework encompasses three aspects: seasonal bias within urban areas, seasonal bias across cities on a global scale, and the impact of seasonal bias in practical applications. The results reveal that (1) seasonal bias is evident, with an average mean absolute percentage error (MAPE) of 54 % for GVI across all sampled cities, and it is particularly pronounced in areas with significant seasonal bias; (2) seasonal bias is strongly correlated with geographic location, with greater bias observed in cities with lower average rainfall and temperatures; and (3) in practical applications, ignoring seasonal bias may result in analytical errors (e.g., an ARI of 0.35 in clustering). By identifying and quantifying seasonal bias in SVI, this study contributes to improving the accuracy of urban environmental assessments based on street view data and provides new theoretical support for the broader application of such data on a global scale.
Traditional studies of urban functions often rely on static classifications, failing to capture the inherently dynamic nature of urban environments. This paper introduces the Spatio-temporal Graph for Dynamic Urban Functions (STG4DUF), a novel framework that combines multimodal data fusion and self-supervised learning to uncover dynamic urban functionalities without ground truth labels. The framework features a dual-branch encoder and dynamic graph architecture that integrates diverse urban data sources: street view imagery, building vector data, Points of Interest (POI), and hourly mobile phone-based human trajectory data. Through a self-supervised learning approach combining dynamic graph neural networks with Spatio-Temporal Fuzzy C-Means (STFCM), STG4DUF extracts parcel-level functional patterns and their temporal dynamics. Using Shenzhen as a case study, we validate the framework through static proxy tasks and demonstrate its effectiveness in capturing multi-scale urban dynamics. Our analysis, based on pyramid functional-semantic interpretation, uncovers intricate functional topics related to human activity, livability, social services, and industrial development, along with their temporal transitions and mixing patterns. These insights provide valuable guidance for evidence-based smart city planning and policy-making.
Urban functional zone (UFZ) classification is essential for understanding city dynamics, supporting urban planning, and enabling effective resource allocation. Traditional approaches rely heavily on remote sensing imagery, which often lacks the contextual information necessary for distinguishing between zones with similar visual features but different functions. This study proposes a novel multi-modal bi-branch deep learning model, named BibDL, which integrates remote sensing imagery with Point-of-Interest (POI) data for UFZ classification. The BibDL model leverages the complementary strengths of these data sources: remote sensing provides spatial and structural information, while POI data offers insights into human activities and land use patterns. Experimental results demonstrate that the BibDL model significantly outperforms a baseline model trained only on imagery, achieving higher F1 scores of 0.975 and Kappa coefficients of 0.953 across multiple UFZ categories. In particular, the BibDL model shows improved performance in challenging zones such as Commerce, Public, and Academia, which are often misclassified when using imagery alone. An ablation study highlights the substantial accuracy gains achieved by incorporating POI data, underscoring the value of a multi-modal approach for UFZ classification. The findings suggest that combining remote sensing imagery with contextual POI density image offers a powerful framework for more precise, context-aware UFZ classification, with implications for urban planning, smart city development, and sustainable resource management.