With the increasing availability of multi-source spatial datasets, ensuring data consistency and accuracy has become a critical challenge in spatial analysis. Traditional global indicators such as Mean Error (ME), Mean Absolute Error (MAE), Mean Relative Error (MRE), Root Mean Square Error (RMSE), Correlation Coefficient (CC), and Univariate Linear Regression (ULR) have been widely used to evaluate overall discrepancies between pair-wise datasets. However, these global metrics fail to reveal spatial details of discrepancies between these datasets. To address this limitation, we established a spatially explicit and scale aware validation framework by embedding conventional global indicators within a geographically weighted (GW) modeling context by introducing GW ME, GW MAE, GW MRE, GW RMSE, GW CC, and univariate GWR to provide spatially detailed assessments of discrepancies. The study also emphasizes the crucial role of bandwidth selection in determining the scale of analysis, with large bandwidths smoothing local variations and small bandwidths preserving fine-scale details. Using GPWv4 data, the Gridded Population of China Dataset (GPCD) and the Seventh National Population Census (NPC7) of China uniformly collected in 2020 as a case study, we validate the effectiveness of this approach at the county level. Results show that GW indicators successfully reveal spatial heterogeneities in data discrepancies, which are overlooked by global measures. It is noteworthy that the local validation framework is always straightforward to apply to other types of spatial data sets. Overall, this study provides a systematic and scalable method for multi-source spatial data validation to enhance quality control, improve integration strategies, and facilitate more informed decision-making in spatial analysis and geographic studies.
Geographically weighted mean (GWM) is a fundamental geographically weighted descriptive statistic for characterizing spatial heterogeneity of a geographic variable, yet its application has long been constrained by the absence of a principled bandwidth selection approach. In practice, bandwidth choice is often guided by rules of thumb or cross-validation strategies that lack clear interpretation and theoretical grounding. To address this limitation, we here proposed an analysis of variance (ANOVA) based approach of bandwidth optimization for GWM. By conceptualizing bandwidth-defined local neighborhoods as spatial groupings and explicitly quantifying between- and within- neighborhood variability through the F-statistic, the proposed approach identifies the bandwidth at which spatial heterogeneity is most effectively expressed. It demonstrated that bandwidth should not be viewed merely as a smoothing parameter, but rather as a scale-controlling mechanism that determines which spatial structures are emphasized or suppressed in GWM. Evaluations using both simulated datasets with diverse spatial patterns and an empirical case study, the proposed method yields stable and interpretable bandwidth choices, while avoiding the pitfalls of overfitting associated with overly small bandwidths and excessive smoothing induced by overly large ones. Overall, the proposed approach enhances the interpretability and methodological transparency of GWM and offers a generalizable choice to its scale selection, with potential applicability to a wide range of spatial data exploration and modeling tasks.
City-level urban renewal is increasingly deployed to achieve sustainable development through the profound spatial restructuring of urban resources. However, it remains unclear whether the highly heterogeneous outcomes of these interventions drive spatially uneven development. In this study, we develop a multi-scale analytical framework to examine spatial disparities in the sustainability performance within the context of Wuhan's 2015-2019 ”Three Old” renewal program. Incorporating scale-consistent sustainability indicators at the city, street, and site scales, the framework systematically identifies where disparities arise and traces the mechanisms through which they are produced. At the city scale, renewal investment demonstrates a pronounced sectoral bias. Resources are disproportionately channeled into growth-enhancing infrastructures, such as transportation and public security, leaving basic social services like education and healthcare largely deprioritized. At the street scale, disadvantaged neighborhoods receive basic ecological improvements, but renewal gains primarily cluster within established zones at the expense of spatial equity. At the site scale, regional spatial contexts prevail over the influence of renewal types on spatial patterns. High quality locations trigger synergistic gains to mitigate site-specific deficits. Conversely, deprived environments constrain renewal efforts regardless of the project type. Overall, the proposed multi-scale approach elucidates how city-level resource allocation is translated into differentiated local outcomes. It reveals localized disparities hidden within aggregated evaluation results, and provides diagnostic insight into place-specific mechanisms governing intervention effectiveness. Thereby, the framework supports the design of more spatially balanced and sustainable urban renewal policies.
Abstract. Accurate aviation emission inventories are essential for environmental impact assessment and mitigation, yet existing estimates predominantly rely on aircraft performance models and standardized assumptions, introducing significant uncertainties. Here, we present a high-resolution, four-dimensional aviation emission inventory for China's civil aviation sector in 2023, built directly from Quick Access Recorder (QAR) data covering 4.52 million flights (~96.8 % of national commercial movements). By coupling second-level fuel-flow measurements with observed flight-state and meteorological parameters via the Boeing Fuel Flow Method 2 and a dynamic thermodynamic correction, this dataset yields temporally and spatially explicit emission estimates for CO₂, NOₓ, SO₂, PM, CO, and HC across all flight phases. In 2023, China's civil aviation consumed 28.4 Mt of fuel, emitting 89.3 Mt CO₂, 634.6 kt NOₓ, 28.4 kt SO₂, 8.46 kt PM, 44.3 kt CO, and 6.35 kt HC. Monte Carlo analysis demonstrates that direct physical observations substantially constrain inventory uncertainty, reducing coefficients of variation (CV) for major pollutants to 2 %–10 % and shifting the dominant error source from activity-level assumptions to emission index parameterizations. High temporal resolution reveals pronounced phase-specific variations: taxiing accounts for 23.7 % of HC and 19.4 % of CO emissions due to low-thrust incomplete combustion, while the brief high-thrust acceleration segment preceding climb contributes 18.0 % of total PM. Spatially, emissions are highly concentrated, with the top 10 % of grid cells accounting for 74.2 % of national CO₂ emissions, exhibiting significant clustering (Global Moran's I = 0.3389, p < 0.05) along major corridors. Providing second-level and high resolution, this dataset serves as an observation-based benchmark for refining existing inventories and assessing near-airport air quality. The dataset is publicly available at https://doi.org/10.5281/zenodo.20828076 (Lu et al., 2026).
Cotton is one of the most economically valuable fiber crops worldwide. Understanding the factors influencing its yield under climate change is crucial for managing climate risks, ensuring a stable textile supply, and promoting sustainable development globally and regionally. This study systematically analyzed the annual yield trends in 82 cotton-producing countries from 1992 to 2021. It investigated the driving factors of cotton yield using eight key indicators, including resource inputs, management practices, and climatic variables. Results show that global cotton yields increased by 27 t km-2 annually, with higher growth in high-income and upper-middle-income countries, the latter achieving yields twice the global average. In contrast, lower-middle-income countries experienced slower growth, while low-income countries faced slight declines. Arid regions maintained the highest yields, with significant improvements across tropical, temperate, and arid zones. Using a varying coefficient spatiotemporal regression (geographically and temporally weighted regression (GTWR)), the study further revealed significant regional variations in the driving factors of cotton yield. Globally, cotton-planted areas and labor inputs negatively impacted yields, as larger areas can lead to insufficient management and high labor density indicates a lack of mechanization. In contrast, fertilizer application and sunlight improved yields, while rising temperatures suppressed them. Regional differences were also observed: mechanization enhanced yields in high-income countries, increased land allocation to cotton improved yields in tropical regions, and fertilizer use drove yield growth in low-income regions. This study highlights the spatial and temporal heterogeneity of global cotton production and offers scientific insights to inform region-specific agricultural management policies and sustainable development strategies.
As the spatial resolution of remote sensing imagery continues to improve, Earth observation scenes exhibit greater diversity in spatial and spectral characteristics, together with increasingly complex surface structures across multiple scales. In such scenarios, global land-cover patterns and local high-frequency details are often strongly coupled, posing significant challenges for accurate scene interpretation and boundary delineation. From a feature representation perspective, spatial-domain features are effective for modeling global context and high-level semantics, whereas frequency-domain representations facilitate the separation of structural components at different scales and enhance fine boundary details, but have limited capability in explicitly capturing semantic relationships. Consequently, most existing semantic segmentation methods operate in a single domain, either spatial or frequency, which restricts their ability to jointly model semantic consistency and diverse spatial structures. To address this issue, we propose a semantic segmentation network termed Spatial–Frequency Collaborative Modeling Network (SFMNet), which jointly learns complementary representations in both spatial and frequency domains. By leveraging spatial-domain semantic priors to guide frequency-domain structural enhancement, SFMNet enables coordinated modeling of land-cover semantics and multi-scale surface structures. In addition, cross-scale interaction between low- and high-frequency components and adaptive multi-level feature fusion are introduced to improve structural consistency across hierarchical feature representations. Extensive experiments on the ISPRS Potsdam, Vaihingen, and LoveDA datasets demonstrate that SFMNet consistently outperforms existing methods in both overall segmentation accuracy and boundary delineation quality. Specifically, SFMNet achieves mean Intersection-over-Union scores of 81.30%, 70.23%, and 55.07%, and Boundary mean Intersection-over-Union scores of 65.42%, 56.76%, and 43.74% on the three datasets, respectively.
Quantifying the unequal supply and demand of ecosystem services (ESs) is a prerequisite for hierarchical ecological governance decisions. However, previous studies have largely overlooked the scale effect of spatially adjacent units and the role of spatial compactness in shaping inequality. To address these research gaps, this study conducted a survey in six counties within the Danjiangkou Basin in China. By adopting a moving window-based local Gini coefficient method, we quantified the inequality in the supply and demand of ESs in this region, and introduced a refined coefficient of variation to measure spatial compactness, analyzing the impact of urbanization on this inequality. The results indicate that the inequality in the supply and demand of ESs in this region is gradually intensifying. However, from a local perspective, the inequality exhibits significant spatial heterogeneity, decreasing gradually from urban centers to suburbs and rural areas, while maintaining strong spatial continuity. Furthermore, we found that urbanization is the primary factor exacerbating this inequality, while compact urban development can mitigate it. The findings of this study can provide practical guidance for cross-county ecological coordination, ecological restoration, and sustainable urban development.
Spatial heterogeneity and correlation are two primary geographical effects of spatial data. Geographically weighted regression (GWR) and its extensions were proposed to quantitively analyze the heterogeneous features in data relationships. An integrative distance metric is usually adopted to calculate proximity-based weights for model calibration for these techniques. However, it could be defective when dealing with higher dimensional data, eg spatio-temporal data (3-D), and geographical flow data (4-D). This study proposes a new local model, namely geographical and temporal density regression (GTDR), to deal with objects of flexible dimensions by reconsidering the spatial weights and experimental investigation of GWR. We use a Nelder-Mead algorithm to optimize each kernel function's bandwidth for every dimension. To validate its performance, we conduct three sets of simulation experiments with 2-D, 3-D, and 4-D data, respectively, and compare them to conventional techniques. Results indicate the apparent advantages of GTDR in treating each dimension individually instead of calculating an integrative distance in traditional ways, such as spatio-temporal or flow distances. All in all, the GTDR technique shows a promising ability in fitting data with higher and diverse dimensions, and exploring heterogeneities in temporal, spatial, spatio-temporal or more complex structural data relationships.
Leaf chlorophyll content (LCC) is vital for photosynthesis and ecosystem functioning; it influences carbon, water, and energy exchanges while serving as an indicator of photosynthetic activity and nitrogen levels in precision agriculture. Hyperspectral data enable precise LCC monitoring by extracting spectral indices through optimal band combination (OBC) and predicting LCC with machine learning. However, OBC faces dimensionality issues, and machine learning models often overlook geographical influences, potentially reducing prediction accuracy. This study hypothesizes that developing spectral indices from important wavelengths and integrating geospatial data into machine learning models can address these issues and increase prediction accuracy. To test this hypothesis, a framework was developed that first uses elastic net (EN) and the successive projection algorithm (SPA) for wavelength selection, followed by spectral index creation with OBC and ranking with random forest (RF). Support vector regression (SVR), random forest regression (RFR), and geographically weighted least squares support vector regression (GWLS-SVR) were then used to assess the prediction accuracy. Finally, the optimal variables and regression model were identified. The results revealed that the EN- and SPA-based indices had stronger correlations and importance than defined indices. The double-difference index (DDn) and the antireflectance index (ARI) are the most robust three-dimensional and two-dimensional spectral indices, respectively. GWLS-SVR requires fewer indices (1-4) to achieve optimal results, with EN-DDn (2R519-R775-R936)-GWLSSVR performing best (R2 = 0.95, RMSE = 0.61, PBIAS = -0.02). This research presents a robust framework with strong adaptability for estimating LCC in a specific study area and region, demonstrating substantial potential for the precise estimation of agroforestry vegetation parameters.
Assessing regional economic development is key for advancing towards the Sustainable Development Goals and ensuring sustainable societal progress. Traditional evaluation methods focus on basic economic metrics like population and capital, which may not fully reflect the complexities of economic activities. Nighttime light (NTL) has been validated as an alternative indicator for regional economic development, yet limitations persist in its evaluation. This study integrates OpenStreetMap (OSM) data and NTL data, providing a novel data integration approach for evaluating economic development. The study uses mainland of China as a case, applying ordinary least squares (OLS) and geographically weighted regression (GWR) to evaluate OSM and NTL data across provincial, municipal, and county levels. It compares OSM, NTL, and their combined use, providing key empirical insights for enhancing data fusion models. The study results reveal: (1) NTL data is more accurate for provincial-level economic activity, while OSM data excels at the county level. (2) GWR demonstrates superior capability over OLS in revealing the spatial dynamics of economic development across scales. (3) Through the integration of both datasets, it is observed that, compared to single-data modeling, the performance is enhanced at the city scale and county scale. The study demonstrates that combining OSM and NTL data effectively assesses economic development in both developed and underdeveloped areas at provincial, municipal, and county levels. The study offers a straightforward and efficient approach to data integration. The findings offer new research perspectives and scientific support for sustainable regional economic growth, particularly valuable in data-scarce, underdeveloped areas.
Existing qualitative direction-relation matrix models employ rigid classification schemes, limiting their ability to differentiate directional relationships between multiple targets within the same directional tile. This paper proposes two quantitative matrix models for qualitative direction-relation with differing levels of precision. Based on directional tile partitioning derived from qualitative direction-relation models, the new models achieve quantitative expression of qualitative directionality through two distinct descriptive parameters: order and coordinate. The order matrix utilizes angular and displacement measurements as sequential variables, capturing the directional sequence characteristics within the same directional tile. The coordinate matrix employs direction-relation coordinates as matrix elements, integrating directional and distance relationships to identify the distribution of targets at varying distances along the same line of sight. These two novel models operate at distinct scales and achieve soft classification of directional relationships, substantially enhancing descriptive precision. Furthermore, they serve as foundational quantitative frameworks for the qualitative direction-relation models, establishing a bridge between quantitative and qualitative models. Experimental assessment confirms that the new models substantially improve directional relationship precision through their quantitative elements while supporting various application domains.
Spatial heterogeneity or nonstationarity in spatial data and relationships has elicited increasing attention in the field of spatial statistics.To explore this fundamental phenomenon,researchers have extensively developed place-or location-specific methods and local statistical techniques that assume data relationships to be spatially variant.In line with the principle of spatial dependence depicted by the first law of geography,the Geographically Weighted(GW)regression technique was proposed to incorporate spatial weights into location-wise regression model calibrations to highlight spatial heterogeneities in data relationships by outputting spatially varying coefficient estimates.With this distance-decaying schema for calculating spatial weights,a series of GW models has been proposed for fine-scaled spatial analysis in descriptive,explanatory,interpretive,and predictive scenarios,including GW descriptive statistics,basic GW regression and extensions,GW discriminant analysis,GW principal component analysis,GW machine learning,and GW artificial neural network.These GW models form a continually evolving technical framework for identifying spatially nonstationary features or patterns in various disciplines or fields,including geography,social science,biology,public health,and environment science. In this study,we systematically sorted the theoretical and technical frameworks of GW models.First,we summarized the essence and rules for applying the family of GW models,including catering for spatially heterogeneous or nonstationary features and relationships in geographic variables and outputting location-dependent metrics or estimates by calculating the spatial weight matrix and the distance-decaying principle of spatial dependence presented by Tobler's first law of geography.With regard to the common and fundamental parts of GW models,we conducted hypothesis tests on spatial heterogeneity or nonstationarity,provided a general definition of distance metrics in geography,calculated spatial weights,and performed bandwidth optimization. With regard to descriptive,explanatory,interpretive,and predictive scenarios,the potential usages of each GW model were discussed from four analysis perspectives.We recommend the use of univariate GW descriptive statistics,such as GW average,GW quantile,GW standard deviation,and GW skewness,to help users grasp the spatially heterogeneous distribution of a geographic variable.For exploratory data analysis with multivariate spatial data,the GW correlation coefficient and GW principal component analysis are recommended.GW regression and its rich extensions,especially multiscale GW regression,are powerful tools in interpretive analysis and have been widely applied.When data relationships are studied comprehensively,accurate predictions are usually obtained in data analytics.The usages of GW regression and geographically and temporally weighted regression in predictions are straightforward,and the prediction accuracy is further improved when artificial intelligence technologies,such as GW machine learning,GW artificial neural network,and geographically neural network weighted regression,are incorporated. The increasing popularity of GW models has resulted in the development of several software packages,standalone programs,and toolkits,including the R package GWmodel and GWmodelS,which are new,free,user-friendly,high-performance standalone software that incorporate spatial data management and mapping tools and GW model functions.However,further improvement is needed before GW models can become all-around,quantitative,analytical frameworks for spatial heterogeneity because of drawbacks in theoretical foundation,technical completeness,complementarity,and evolution to spatiotemporal dimensions.
Rapid urbanization poses a serious threat to regional ecological security (ES) and has led to significant ecological losses. To assess and mitigate ecological security (ES) losses caused by urban expansion, this study simulated urban expansion patterns in Northwest China by 2050 under the Shared Socioeconomic Pathways and Representative Concentration Pathways (SSP-RCP) scenarios. A coupled ERI (Ecological Risk Index)-EH (Ecological Health)-ESs (Ecosystem Services) model was employed to quantitatively analyze the regional ES development levels under different scenarios. From the perspectives of landscape pattern dynamics and ecological degradation, the study evaluated the impacts of urban expansion on ES across various development modes. Scenario comparisons were then used to implement ES zoning control strategies. The results reveal that while the northwest and southeast regions show relatively high ES levels, a gradual decline is observed towards central areas. The Economic Priority Development (EPD) model enhances urban land cohesion and mitigates landscape trade-offs on ES. Among the three scenarios, the ranking of ES loss areas is: EPD > Business as Usual (BAU) > Ecological Priority Protection (EPP), with the EPD scenario causing 2.524 times more ecological loss than the EPP scenario. Using provincial, municipal, county, and ecological function zones, the optimal urban development plan is selected based on the highest ES potential. This study provides strategic guidance for spatial planning by offering a practical and scalable framework to reduce ecological degradation while promoting balanced and sustainable regional development.
The increasing popularity of geographically weighted (GW) techniques has resulted in the development of several software packages. Their ongoing updates and enhancements always require extraordinary efforts. In this study, we introduced a fundamental C++ library, namely libgwmodel. With its updates and maintenance in the future, all the associated products could be uniformly maintained and upgraded. Moreover, its development will greatly facilitate the further achievement of more GW models or tools, even from different teams, so as to form a comprehensive foundation for further development and extensions of GW models.
Flight trajectory data from Quick Access Recorders (QAR) is critical for ensuring flight safety and conducting performance analysis. However, the inherent uncertainties and noise present in this data necessitate the use of filtering techniques. The conventional Kalman filter, widely applied in civil aviation, exhibits limitations when addressing varying noise levels across different aircraft types and flight phases. This study addresses these challenges through a multistep approach. First, QAR data from Daocheng Yading Airport underwent preprocessing, including data cleaning, resampling, key field selection, and target trajectory extraction. Next, the Kalman filter’s adaptive capabilities were enhanced and applied specifically to three-dimensional trajectory data. Finally, a comparative analysis was conducted with the segmented noise matrix adjustment method. The results demonstrate that the adaptive Kalman filter effectively preserves essential data characteristics while streamlining the filtering process.
The emergence of the Digital Twin of Earth (DTE) signifies a pivotal advancement in the field of the Digital Earth, which includes observations, simulations, and predictions regarding the state of the Earth system and its temporal evolution. While the Geospatial Data Cubes (GDCs) hold promise in probing the DTE across time, existing GDC approaches are not well-suited for real-time data ingestion and processing. This paper proposes the design and implementation of a real-time ingestion and processing approach in Geospatial Data Cube (RTGDC), achieved by ingesting real-time observation streams into data cubes. The methodology employs a publish/subscribe model for efficient observation ingestion. To optimize observation processing within the data cube, a distributed streaming computing framework is utilized in the implementation of the RTGDC. This approach significantly enhances the real-time capabilities of GDCs, providing an efficient solution that brings the realization of the DTE concept within closer reach. In RTGDC, two cases involving local and global scales were implemented, and their real-time performance was evaluated based on latency, throughput, and parallel efficiency.