
Recent research in the field of geographic information science shows that reproducibility and replicability of publications have substantial room for improvement and proposes various actions to improve the situation. However, the impact of these actions remains unclear. This study investigates the combined effect of novel author guidelines and workflow review process, which award badges for successful reproductions, on the potential reproducibility of articles published in the AGILE conference series proceedings over the past decade. While replicating the approach of previous studies, this work expands the scope of prior reproducibility assessments and systematically compares the findings for the AGILE conference with those of the GIScience conference series proceedings, which has not undergone similar changes to guidelines or procedures. Results indicate that the reproducibility guidelines and the review process measurably improved the potential reproducibility of AGILE publications. The comparison with GIScience papers further suggests that clear and enforced guidance is a key driver for change. Our findings demonstrate the value of institutional policies and community norms in fostering reproducible research in the GIScience field and identify pathways for its ongoing improvement.
The proliferation of location-sensing technologies, including smartphones and standalone GPS devices, has transformed human mobility studies. A central methodological concern in these studies is positioning accuracy, as errors in positioning can bias estimates of mobility patterns and environmental exposure. While the recent rollout of next-generation 5G technology promises sub-meter accuracy, most evidence comes from simulations or controlled trials, leaving everyday mobility performance largely untested, particularly in comparison with 4G smartphones and standalone GPS devices. This study systematically evaluates the positioning errors of 5G New Radio (NR) smartphones compared with 4G smartphones and standalone GPS devices in everyday mobility scenarios. Data were collected from 27 participants across walking, biking, and driving routes on the Western University campus, encompassing five environmental contexts: open space, between buildings, under tree canopy, in building, and underground. Positioning errors were assessed using three complementary approaches: point-based, path-based, and area based analyses. Results demonstrate that 5G smartphones consistently outperform 4G devices and standalone GPS in accuracy, particularly in building and underground, achieving lower median errors, higher spatial path fidelity, and improved indoor localization metrics. This advancement broadens the applicability of 5G smartphones across diverse research in health geography and urban planning.
This paper demonstrates an approach to the application of generalized additive models (GAMs) with space-time smooths to model coefficient processes that vary over space and time. The approach is to create and evaluate multiple GAMs, each with the predictor variables specified in different ways. It emphasizes the need to determine the nature of the space-time dependencies present in the data relationships rather than to assume them, based on the perceived data generating process, especially if this is unknown. The approach is explored using simulated coefficient data with known space-time dependencies. The GAMs are compared with multiscale geographically and temporally weighted regression (MGTWR) models and are shown to have marginally weaker predictive performance and to be marginally better at coefficient recovery. The inferential costs of misspecifying the target-to-predictor variable relationships in the GAMs is quantified both for individual variable main effects and interacting misspecifications. The approach is then applied to an empirical case study of NDVI (as a proxy for forest productivity) informed by precipitation and temperature in the Chaco dry rainforest of South America. The best GAM is determined and its space-time varying coefficient estimates are investigated. The methods and results are discussed and several areas of further work and enhancements to the stgam R package used to undertake this analysis are identified.
The generation of uniform, gridded data from spatially discontinuous station values and the assessment of their accuracy are essential for water resources assessment and management, and for climate change studies, especially in semi-arid environments. Spatial and temporal grids have been generated in recent years as a basis for several studies. This work has two objectives: a) to evaluate the accuracy of the grids available for Spain by comparing their monthly values with stations not used in their estimation or prediction, and b) to verify the improvement in accuracy using Multi-Model Ensembles based on machine learning. A dual ensemble approach is presented: (i) multiple individual Random Forest (RF) ensembles per weather station, using only the information from the station, and (ii) spatially distributed grid prediction using a single ensemble model that incorporates all the information from the nearest stations and their distances using Random Forest Spatial Interpolation -RFSI-). Both models were used to generate monthly data grids of maximum, minimum and mean temperature, and total precipitation, with high spatial resolution (5 km). Seven datasets: Iberia01, STEAD, AEMET, SIMPA, EOBSv27 and STEAD, were used as predictors. Accuracy was estimated using the root mean square error, the percentage bias and the Nash-Sutcliffe efficiency index obtained using block cross-validation buffering (LOOBUF-CV), robust to spatial autocorrelation. The significance of the differences was assessed using ANOVA with heteroscedasticity correction in the residuals. Preliminary results indicate that multi-model ensembles using RF outperform individual grids. Among other reasons, ensembles aggregate the different representations of meteorological processes included in each grid and reduce the uncertainty associated with each individual grids.
This paper addresses the problem of constructing an ideal projection of an ellipsoid of revolution onto a sphere based on the Airy criterion. General equations required to solve the problem are derived. In particular, the Euler-Urmaev system is obtained, allowing a clear illustration of Gauss's theorem that a distortion-free projection between these surfaces cannot exist. The Euler-Ostrogradsky system is also derived to find the projection that minimizes distortion according to the Airy criterion. Natural boundary conditions for the ideal projection are analyzed. It is shown that on the boundary of the mapping region, Tissot's indicatrices are aligned either along the normals or tangents to the boundary, and one of the extremal linear scale factors is equal to unity. Since the value of the Airy criterion depends not only on the projection's mapping functions but also on the radius of the sphere, an additional integral condition is introduced alongside the Euler-Urmaev system, the Euler-Ostrogradsky system, and the natural boundary conditions. According to this condition, the integral of the area distortion over the entire mapping region of the ideal projection must be equal to zero. Two specific cases are examined in detail: projection of the entire ellipsoid and of a region bounded by a parallel. For comparison, conformal projections optimized according to the Airy criterion were also constructed for the same mapping regions. The resulting ideal projections can be used in geodesy for solving direct and inverse geodetic problems, and in cartography for constructing double projections.
Checkpoint data are generated by movement past fixed "checkpoints," such as smart-card readers. The heterogeneity of checkpoint data sources, data structures, data models, and granularities means existing research on checkpoint-based movement analytics relies on bespoke or ad-hoc analytics frameworks. This work addresses this gap by developing a consistent framework for analyzing network-constrained checkpoint movement data based on the "cordon network." The cordon network is a simple, graph-based computational structure that captures the underlying spatial structure and heterogeneous granularity of movement through checkpoints. This paper explores the design, development, and testing of an analytics toolkit founded on the cordon network. The approach and its ability to handle heterogeneous checkpoint data added within transportation networks is validated using three diverse transportation case studies. While this study focuses on transportation-related checkpoint data, the discussion and conclusions outline key factors for extending the framework to other domains, providing guidelines for checkpoint movement analysis across different contexts.
This paper introduces and evaluates a novel method for privacy-preserving distance computations. The method is based on randomized geometric surface calculations and replaces coordinates with contextual variables representing information about the coordinates or the distances between coordinates. The method is presented with an accompanying step-by-step workflow. Its applicability is demonstrated with real-world spatial data sets from Germany and the Netherlands that contain information about hospital and school locations. Open data was used to enable reproducibility. The method's utility is evaluated in detail using correlations, the relative root mean squared error (RRMSE), a Monte Carlo simulation, and the Wasserstein distance. The results show that the method yields high correlations, provides reasonably accurate results as an RRMSE of about 20% is achieved, converges fast, and preserves the spatial distribution of the true coordinates.
Cycling practice is quickly increasing around the world, giving rise to the development of devoted infrastructure to protect its users and offer them a more enjoyable ride. Mobility infrastructure is represented in geographical databases, but these databases are often centered on car and pedestrian mobility. This causes some data quality problems like the lack of completeness or freshness. Volunteered geographical information (VGI) is affected by this kind of problem with a variable extent relying on the contributors' wish and skills. Research on VGI evolution for a network mainly focuses on the main usage of a road section, ignoring secondary information related to other road users of a specific section. This paper contains two contributions. To model the evolution, we define a multiplex graph where each layer represents a snapshot. It is implemented with an infrastructure class based on how cyclists perceive an infrastructure. We also present two complementary VGI road network evolution methods with a usage-centric approach on cycling. These approaches are adaptable for any usage of the network and are based on the multiplex graph. The first approach is based on the road sections, analyzing the evolution of each section individually. The second approach is based on randomly generated starting/ending points. These methods are illustrated in the Centre-Val de Loire region with OpenStreetMap.
This paper gives a new possible realization of the oblique Lambert Azimuthal Equal-Area map projection for the ellipsoid of revolution. Unlike the realization available in previous literature, the authalic sphere used for the derivation has very low distortion at the neighbourhood of a freely chosen standard parallel. For this reason, the distortions caused by this authalic sphere can be neglected. It is shown that this realization gives a better approximation of the azimuthal equal-area mapping of the sphere in terms of angular distortions. Interesting side results of the study include a numerically stable inverse formulation for the azimuthal equal-area map of the sphere and mathematical connections between the Gaussian conformal sphere and the low-distortion authalic sphere.
The integration of large language models (LLMs) with geographic information science (GIScience) represents a new frontier in interdisciplinary research that combines advanced natural language processing with sophisticated spatial data analysis. This paper explores the synergistic potential of combining the natural language understanding and generation capabilities of LLMs with the expertise of GIScience in handling complex geospatial data. By exploring the specific contributions that LLMs can offer to GIScience, such as improving data processing, analysis, and visualization, and the mutual benefits that GIScience can offer to LLMs in terms of spatial reasoning and conceptual frameworks, we outline a comprehensive framework and a research agenda for this integration. Furthermore, we address the societal and ethical implications of this convergence, highlighting the challenges of bias, misinformation, and environmental impact. Through this exploration, we aim to set the stage for innovative applications in urban planning, environmental analysis, and beyond, while emphasizing the need for responsible use of AI.
Indoor wayfinders often rely on verbal route directions, particularly in situations where other navigational aids may be unavailable or less effective. Ensuring the clarity and validity of these instructions is particularly important for navigation in complex indoor environments, such as airports and malls. However, current methods lack a reliable, systematic approach to computationally ensuring the a-priori validity of route instructions, failing to provide certainty to agents that they will be able to follow instructions successfully. Here we show a novel computational model for validating indoor route instructions, applicable to a wide range of indoor environments and turn-based grammars. Using a synthetic dataset of indoor floorplans with varying complexities, we demonstrate the model's capability to validate route instructions systematically. We systematize the requirements for route instruction validation in the framework which assesses instructions based on understandability, executability, path-following, and destination guidance. Our findings highlight the effectiveness of nuanced grammars, such as 8-sector grammar, for complex layouts and confirm the applicability of simpler grammars, like 4-sector grammar, for right-angle constrained environments. Importantly, we identify a transition point where the benefits of increased grammatical complexity on the descriptions of the turns are no longer productively supporting a reduction in turn ambiguity in the environments. This research shifts the field from subjective, time-consuming human evaluations to a computational approach, enhancing the reliability of indoor navigation systems.
Virginia's seventeenth- and eighteenth-century land patents survive primarily as narrative metes-and-bounds descriptions, limiting spatial analysis. This study systematically evaluates current-generation large language models (LLMs) in converting these prose abstracts into research-grade latitude/longitude coordinates. A digitized corpus of 5,471 Virginia patent abstracts (1695–1732) is released, with 43 rigorously verified test cases for benchmarking. Six OpenAI models across three architectures—o-series, GPT-4-class, and GPT-3.5—were tested under two paradigms: direct-to-coordinate and tool-augmented chain-of-thought invoking external geocoding APIs. Results were compared against a professional GIS workflow, Stanford NER geoparser, Mordecai-3 neural geoparser, and a county-centroid heuristic. The top single-call model, o3-2025-04-16, achieved a mean error of 23 km (median 14 km), a 67% improvement over professional GIS methods and 70% better than Stanford NER. A five-call ensemble further reduced errors to 19 km (median 12 km) at minimal additional cost (~USD 0.20 per grant). Paired Wilcoxon tests confirm ensemble superiority (W=629, p=0.03 vs. single-shot). A patentee-name redaction ablation slightly increased error (~9%), showing reliance on metes-and-bounds reasoning rather than memorization. The cost-effective gpt-4o-2024-08-06 model maintained a 28 km mean error at USD 1.09 per 1,000 grants, establishing a strong cost-accuracy benchmark. External geocoding tools offer no measurable benefit for this task. These findings demonstrate that LLMs can georeference early-modern records as accurately and significantly faster and cheaper than traditional GIS workflows, enabling scalable spatial analysis of colonial archives.
Recent workshops brought together several developers, educators and users of software packages extending popular languages for spatial data handling, with a primary focus on R, Python and Julia. Common challenges discussed included handling of spatial or spatio-temporal support, geodetic coordinates, in-memory vector data formats, data cubes, inter-package dependencies, packaging upstream libraries, differences in habits or conventions between the GIS and physical modeling communities, and statistical models. The following set of recommendations have been formulated: (i) considering software problems across data science language silos helps to understand and standardise analysis approaches, also outside the domain of formal standardisation bodies; (ii) whether attribute variables have block or point support, and whether they are spatially intensive or extensive has consequences for permitted operations, and hence for software implementing those; (iii) handling geometries on the sphere rather than on the flat plane requires modifications to the logic of simple features, (iv) managing communities and fostering diversity is a necessary, on-going effort, and (v) tools for cross-language development need more attention and support.
Open data initiatives and infrastructures play an essential role in favoring better data access, participation, and transparency in government operations and decision-making. Open Geographical Data Infrastructures (OGDIs) allow citizens to access and scrutinize government and public data, thereby enhancing accountability and evidence-based decision-making. This encourages citizen engagement and participation in public affairs and offers researchers, non-governmental organizations, civil society, and business sectors novel opportunities to analyze and disseminate large amounts of geographical data and to address social, urban, and environmental challenges. In Latin America, while recent open government agendas have shown an inclination towards transparency, citizen participation, and collaboration, only a limited number of OGDIs allow unrestricted use and re-use of their data. Given the region's cultural, social, and economic disparities, there is a contrasting digital divide that significantly impacts how OGDIs are being developed. Therefore, this paper analyses recent progress in developing OGDIs in Latin America, technological gaps, and open geographical data initiatives. The main results denote an early development of OGDIs in the region. Nevertheless, this opens the door for the timely involvement of citizens and non-government sectors to share needs, experiences, knowledge, and expertise, as well as to address a transboundary research agenda. Challenges are discussed from multiple perspectives: data, methodological, governmental and readiness, and potential impact. This analysis is aimed at researchers, policymakers, and practitioners interested in the specific challenges and progress of OGDIs in Latin America, while also contributing to the global conversation on best practices and lessons learned in implementing OGDIs across different contexts.
A city is an intricate system where interactions between transport, land use, the environment, and the population occur at various scales. This complexity makes it challenging to predict and govern these interactions. However, big data on human activity patterns allows researchers to discover dynamic, temporary patterns in the activity landscape and understand the choreographies of people's behavior to enhance urban areas' vitality through planning. In this article, we hypothesized that a higher diversity of urban spatio-functional and socio-economic features indicates higher urban vitality in Tallinn, Estonia. We explored multi-sourced indexes to interpret this formation of urban vitality using complex agent variables of location, cluster, diversity, and similar actors generating self-organizing patterns of urban life. We used functional and morphological components and socio-economic data identified as traditional, `slow' vitality measures (SM), and mobile phone location data as dynamic metrics (DM), respectively. We analyzed them in a geographic information system (GIS) environment to measure the types of spatial configurations, temporal variation of vital places, and their correlation. The results indicate a positive correlation (r=0.5116) between the slow metrics and the high mobile phone activity. These correlations demonstrate that cell phone data provides a detailed and accurate view of people's daily rhythms and choreographies. The diversity indicators offer a new method to interpret urban vitality in cities and make planning decisions that support its emergence.
The color landscape is an essential aspect of each village, representing both natural scenery and human history. However, previous research has not provided a thorough and quantitative assessment of the color spatial pattern of regions and their surroundings. In this study, color patches were extracted from 3D real scene models, and color landscape indices were used to quantify the color landscape pattern. A questionnaire was utilized to establish the association between the color landscape indices and the Scenic Beauty Estimation (SBE) scores, which was then used to predict the SBE without the need for another questionnaire. The results showed that: 1) the color landscape indices extracted using 3D real scene models can reveal the scenic beauty of villages, with different villages presenting various color landscape patterns; 2) the SBE scores obtained through the questionnaire have a strong correlation with various color landscape indices, such as COHESION, LPI, SPILT, Y-MPS, and GE-MPS; 3) the SBE model based on color landscape indices was developed using stepwise linear regression, with an R2 value of 0.822 and an average error of 0.248, which can predict SBE in various places without the use of a questionnaire. This study introduces a new perspective and approach for estimating scenic beauty, which will help with rural planning and beautiful countryside development.
This research is aimed to solve the tweet/user geolocation prediction task and provide a flexible methodology for the geotagging of textual big data. The suggested approach implements neural networks for natural language processing (NLP) to estimate the location as coordinate pairs (longitude, latitude) and two-dimensional Gaussian Mixture Models (GMMs). The scope of proposed models has been finetuned on a Twitter dataset using pretrained Bidirectional Encoder Representations from Transformers (BERT) as base models. Performance metrics show a median error of fewer than 30 km on a worldwide-level, and fewer than 15 km on the US-level datasets for the models trained and evaluated on text features of tweets' content and metadata context. Our source code and data are available at https://github.com/K4TEL/geo-twitter.git
This paper reviews trends in GeoAI research and discusses cutting-edge advances in GeoAI and its roles in accelerating environmental and social sciences. It addresses ongoing attempts to improve the predictability of GeoAI models and recent research aimed at increasing model explainability and reproducibility to ensure trustworthy geospatial findings. The paper also provides reflections on the importance of defining the "science" of GeoAI in terms of its fundamental principles, theories, and methods to ensure scientific rigor, social responsibility, and lasting impacts.
In recent decades, the analysis of different geographic scales for studying the spatial patterning of crime has profoundly deepened our theoretical grasp of crime dynamics. However, a similar investigation is lacking when it comes to the patterning of offender residences, despite there being clear theoretical and empirical reasons for doing so, among them, the close relationship between where offenders live and where their corresponding crimes are committed. This paper delves into the concentration and variance of offender residences across different levels of spatial aggregation. The data used contains the locations of residence for known offenders in Birmingham between the years 2006 and 2016. Resident locations are aggregated to Output Areas (OA), nested within Lower Super Output Areas (LSOA), further nested within Middle Super Output Areas (MSOA). Descriptive and model-based statistics are deployed to quantify concentration and variation at each spatial scale. Results suggest that most variance (~48%) in offender residence concentrations is attributable to the largest spatial scale (MSOA level). Output Areas capture approximately 38\% of the variance. Findings open up discussions on the role of urban development in determining the appropriateness of spatial scale.
Street networks are ubiquitous components of cities, guiding their development and enabling movement from place to place; street networks are also the critical components of many urban analytical methods. However, their graph representation is often designed primarily for transportation purposes. This representation is less suitable for other use cases where transportation networks need to be simplified as a mandatory pre-processing step, e.g., in the case of morphological analysis, visual navigation, or drone flight routing. While the urgent demand for automated pre-processing methods comes from various fields, it is still an unsolved challenge. In this article, we tackle this challenge by proposing a cheap computational heuristic for the identification of "face artifacts", i.e., geometries that are enclosed by transportation edges but do not represent urban blocks. The heuristic is based on combining the frequency distributions of shape compactness metrics and area measurements of street network face polygons. We test our method on 131 globally sampled large cities and show that it successfully identifies face artifacts in 89\% of analyzed cities. Our heuristic of detecting artifacts caused by data being collected for another purpose is the first step towards an automated street network simplification workflow. Moreover, the proposed face artifact index uncovers differences in structural rules guiding the development of cities in different world regions.