Large-scale pre-trained machine learning models have reshaped our understanding of artificial intelligence across numerous domains, including our own field of geography. As with any new technology, trust has taken on an important role in this discussion. In this chapter, we examine the multifaceted concept of trust in foundation models, particularly within a geographic context. As reliance on these models increases and they become relied upon for critical decision-making, trust, while essential, has become a fractured concept. Here we categorize trust into three types: epistemic trust in the training data, operational trust in the model's functionality, and interpersonal trust in the model developers. Each type of trust brings with it unique implications for geographic applications. Topics such as cultural context, data heterogeneity, and spatial relationships are fundamental to the spatial sciences and play an important role in developing trust. The chapter continues with a discussion of the challenges posed by different forms of biases, the importance of transparency and explainability, and ethical responsibilities in model development. Finally, the novel perspective of geographic information scientists is emphasized with a call for further transparency, bias mitigation, and regionally-informed policies. Simply put, this chapter aims to provide a conceptual starting point for researchers, practitioners, and policy-makers to better understand trust in (generative) GeoAI.
Abstract. Unused energy is released into the environment from a number of sources, such as energy-intensive industries, whereas it could be used in district heating (DH) systems to make a significant contribution to decarbonising the heating sector. Despite its essential role, surplus heat (SH) is currently under-utilised and large amounts are wasted. Furthermore, there is a lack of detailed data in the literature on the sources that could potentially be captured. This paper examines the historical distribution of DH systems and their relationship with SH in Denmark, a country with very advanced DH. It also considers the potential for medium and low-temperature sources to be included alongside high-temperature industrial SH, which has been the dominant resource to date. Specifically, it utilises visual and statistical interpretation of the geospatial distribution of the utilised amounts of SH at municipal level. The analysis shows through classified choropleth mapping the dynamic temporal and spatial changes in the field, where the potentials depend on the availability of sources and their corresponding activities, but also shows that many municipalities with high and increasing DH shares could utilise SH from various sources that are not currently used such as in the food and beverages and retail and trade sector. Further potential from manufacturing and industrial activities could also be included, especially in the municipalities of the capital region, while significant low-temperature SH across the country would require a boost from heat pumps or upgrading of DH systems.
Open educational resources (OER) are teaching, learning, or research resources freely available for use and reuse. Despite their potential, OER uptake in existing education systems remains low, primarily due to challenges in locating suitable resources. This study addresses this challenge by proposing and implementing a workflow applying the FAIR (Findable, Accessible, Interoperable, and Reusable) principles to OER. We demonstrated this framework within the Earth System Sciences as an application domain. We constructed a knowledge graph of approximately 500 FAIR OER, each annotated with structured metadata using the Schema.org vocabulary and made accessible through a SPARQL endpoint. To bridge the gap between making resources queryable and enabling their practical reuse, we employed a transformer-based language model (Sentence-BERT). The model was fine-tuned using few-shot learning on a domain-specific dataset of course-description pairs. This specialized model was then used to map the OER collection against over 200 university courses across five academic programs at a German university, based on semantic similarity between OER descriptions and university course descriptions. Expert evaluation of the model’s recommendations demonstrated 74% accuracy in identifying reusable OER for university courses. Notably, even with limited training data, fine-tuning the Sentence-BERT model significantly improved performance, resulting in a 16% reduction in mean squared error compared to the base model. This study provides both a generalizable methodology and a practical demonstration of how FAIR principles can streamline OER discovery, potentially accelerating OER uptake in higher education.
Spatially aggregated data on socio-demographic groups often fail to capture the population's spatial heterogeneity in cities. This poses challenges for urban planning, particularly when addressing the needs of groups such as migrants or families with children. Moreover, the commonly provided aggregated units, such as census tracts, vary in size and across data sources. Existing literature on disaggregation typically handles individual subgroups separately, ignoring their interrelations in the downscaling process. This article explores the potentials of multi-output regression models for simultaneous spatial downscaling of multiple groups and conducts a detailed spatial error analysis using individualized neighborhoods. We experiment with self-training gradient-boosting trees and fully convolutional neural networks, assessing the quality of results against ground truth data at the target resolution. We show that the evaluation of the disaggregated results at this detailed resolution requires unconventional methods. The methodology proves convenient and achieves high-accuracy results using input datasets of building features.
Education on spatial data infrastructures is an important building block of study programs focusing on spatial data. The focus is on equipping students with essential skills for future contributions to spatial data infrastructure development and application. Despite a wide range of available teaching materials, their reuse is hindered by a lack of harmonization and integration into learning modules. This paper reports on the activities in the "Open Educational Resources for Spatial Data Infrastructures" (OER4SDI) project, which addresses this gap by developing easily reusable teaching materials for spatial data infrastructures. OER4SDI strives to create high-quality, findable, accessible, interoperable, and easily reusable (FAIR) educational resources, fostering efficient knowledge exchange in the field of spatial data infrastructures. The project puts a special focus on the modularity of open educational resources to enable reuse in a wide variety of settings. We describe the workflows and technical setup we use to enable collaborative development of open educational resources and their continuous revision, which is vital in this rapidly evolving field. Three example open educational resources for spatial data infrastructures are given, dealing with the "AAA Data Catalog", "OGC Open API Features", and "Knowledge Graphs". We share the lessons learned from the project to foster the development of more high-quality open educational resources in our community.
Population growth in urban centres and the intensification of segregation phenomena associated with international mobility require improved urban planning and decision-making. More effective planning in turn requires better analysis and geospatial modelling of residential locations, along with a deeper understanding of the factors that drive the spatial distribution of various migrant groups. This study examines the factors that influence the distribution of migrants at the local level and evaluates their importance using machine learning, specifically the variable importance measures produced by the random forest algorithm. It is conducted on high spatial resolution (100×100 grid cells) register data in Amsterdam and Copenhagen, using demographic, housing and neighbourhood attributes for 2018. The results distinguish the ethnic and demographic composition of a location as an important factor in the residential distribution of migrants in both cities. We also examine whether certain migrant groups pay higher prices in the most attractive areas, using spatial statistics and mapping for 2008 and 2018. We find evidence of segregation in both cities, with Western migrants having higher purchasing power than non-Western migrants in both years. The method sheds light on the determinants of migrant distribution in destination cities and advances our understanding of the application of geospatial artificial intelligence to urban dynamics and population movements.
Human migration is associated with a dynamic process that stimulates the generation of large data streams that hold valuable information and is possible to comprehend through different visualisation methods. In this research, we present an analytical approach to understanding human migration patterns in urban areas of Amsterdam and Copenhagen. Unlike other methods, we propose to map migration data within grid cells by retaining data characteristics and plotting statistical calculation analysis for better visualisation and analytics.
Over the past decades, the world has experienced increasing heatwave intensity, frequency, and duration. This trend is projected to increase into the future with climate change. At the same time, the global population is also projected to increase, largely in the world’s cities. This urban growth is associated with increased heat in the urban core, compared to surrounding areas, exposing residents to both higher temperatures and more intense heatwaves than their rural counterparts. Regional studies suggest that Asia and Africa will be significantly affected. How many people may be exposed to levels of extreme heat events in the future remains unclear. Identifying the range in number of potentially exposed populations and where the vulnerable are located can help planners prioritize adaption efforts. We project the ranges of population exposed to heatwaves at varying levels to 2,100 for three future periods of time (2010–2039, 2040–2069, 2070–2099) using the Shared Socio-Economic Pathways (SSPs) and the Representative Concentration Pathways (RCPs). We hypothesize that the largest populations that will be exposed to very warm heatwaves are located in Asia and Africa. Our projections represent the warmest heatwaves for 15 days during these three periods. By the 2070–2099 period, the exposure levels to extreme heatwaves (>42°) exceed 3.5 billion, under the sustainability scenario (RCP2.6-SSP1). The number of those exposed in cities climbs with greater projected climate change. The largest shares of the exposed populations are located in Southern Asia and tropical countries Western and Central Africa. While this research demonstrates the importance of this type of climate change event, urban decision-makers are only recently developing policies to address heat. There is an urgent need for further research in this area.
Data normalization for removing the influence of population density in Population Geography is a common procedure that may come with an unperceived risk. In this regard, data are constrained to a constant sum and they are therefore not independent observations, a fundamental requirement for applying standard multivariate statistical tools. Compositional Data (CoDa) techniques were developed to solve the issues that the standard statistical tools have with close data (i.e., spurious correlations, predictions outside the range, and sub-compositional incoherence) but they are still not commonly used in the field. Hence, we present in this article a case study where we analyse at parish level the spatial distribution of Danes, Western migrants and non-Western migrants in the Capital region of Denmark. By applying CoDa techniques, we have been able to identify the spatial population segregation in the area and we have recognized some patterns that can be used for interpreting housing prices variations. Our exercise is a basic example of the potential of CoDa techniques, which generate more robust and reliable results than standard statistical procedures, but it can be generalized to other population datasets with more complex structures.
Accurate and consistent estimations on the present and future population distribution, at fine spatial resolution, are fundamental to support a variety of activities. However, the sampling regime, sample size, and methods used to collect census data are heterogeneous across temporal periods and/or geographic regions. Moreover, the data is usually only made available in aggregated form, to ensure privacy. In an attempt to address these issues, several previous initiatives have addressed the use of spatial disaggregation methods to produce high-resolution gridded datasets describing the human population distribution, although these projects have usually not addressed specific population subgroups. This paper describes a spatial disaggregation method based on self-training regression models, innovating over previous studies in the simultaneous prediction of disaggregated counts for multiple inter-related variables, by leveraging multioutput models based on gradient tree boosting. We report on experiments for two case studies, using high-resolution data (i.e., counts for different subgroups available at a resolution of 100 meters) for the municipality of Amsterdam and the region of Greater Copenhagen. Results show that the proposed approach can capture spatial heterogeneity and the dependency on local factors, outperforming alternatives (e.g., seminal disaggregation algorithms, or approaches leveraging individual regression models for each variable) in terms of averaged error metrics, and also upon visual inspection of spatial variation in the resulting maps.
ABSTRACT Stone walls in the landscape of Denmark are protected not only for their cultural and historical significance but also for their vital role in supporting local biodiversity. Many stone wall structures have either disappeared, suffered substantial damage, or had segments removed. Additionally, as it stands today, the registry of these structures, managed by each municipality, is outdated and incomplete. Leveraging recent developments in Machine Learning and Convolutional Neural Networks (CNNs), we analyze the publicly available terrain data (40 cm resolution) derived from the Danish LiDAR data, using a U-Net-like CNN model to assess the stone walls dataset and provide for an update of the registry. While the Digital Terrain Model (DTM) alone provided good results, better results were obtained when adding Height Above Terrain (HAT) and an additional DTM layer with a Sobel filter applied. Using a pixel-wise evaluation, there was an overall agreement of 93% between ground truth and prediction of stone walls in a validation area and 88% overall agreement for the whole predicted area. Good generalizability was found when externally validating the model on new data, showing positive results for both the existing stone walls and predicting new potential ones upon visualisation. The method performed best in open areas, however positive results were also seen in forested areas, although denser areas and urban areas presented as challenging. Given the lack of a reference dataset or other studies on this specific matter, the evaluation of our study was heavily based on the stone walls registry itself complemented by visual inspection of the predictions and on the ground in the Danish municipality of Ærø. Automating the process of identifying and updating the stone walls registry in Denmark is of great relevance to the local governments. We suggest the development of a Decision Support System to allow municipalities access to the results of this method.
Abstract. The new concept of Open Spatial Data Infrastructures (Open SDIs) has emerged from an increased interest in open data initiatives together with national and international directives, such as the EU Open Data Directive (Directive (EU) 2019/1024), and the large investment of European public authorities in developing SDIs for sharing spatial data within public authorities. Open SDIs have the potential to boost reaching SDIs’ general aims and goals of facilitating the exchange and sharing of spatial data to support planning and decision-making by including public participation and increased openness in all aspects of SDIs, including Open SDI Education. The open SPatial data Infrastructure eDucation nEtwoRk (SPIDER) project aims to address Open SDI Education by particular emphasis on studying Active Learning and Teaching (ALT) methods for SDI education. This article provides a theoretical basis of ALT for SDI methodologies. We show in which way ALT practices were already implemented in SDI education at the Partner universities before the COVID-19 pandemic. We also describe how the pandemic functioned as a catalyst for implementing ALT practices to an online environment, and how students evaluated these practices. The outcomes of our research can serve as an inspiration for SDI education in other countries.
Identifying centralities in cities helps determine how public space is perceived and utilized in everyday life. Sustainable mobility, social sustainability and spatial justice can be examined by investigating centralities in the urban form. In this study, we investigate configurational centralities in metropolitan Copenhagen created by the road network based on space syntax analysis and active centralities of land-use patterns with a geographical approach. The purpose of the research is to present a reproducible methodology for determining the active and configurational centralities. Using this methodology, we explore the meaning of the centralities in terms of pedestrian and cyclist accessibility, as well as the role of the configurational centralities in shaping land-use patterns. The results serve as input to an analysis of their relation through Kernel Density Correlation and spatial correlation. The results of correlations indicate that areas close to the city centre and around the Finger Plan – Copenhagen’s strategic development plan – tend to be more central and favourable for pedestrians and cyclists. On the contrary, central areas far from the city centre, especially in Northern Copenhagen, and areas between the axes of the Finger Plan are more car-oriented since centralities are dispersed and located around highways or road segments designed for cars. The workflow presented in this paper is provided as a set of open-source R scripts that draw largely on data from OpenStreetMap, thus enabling replications of the study for other cities.
Recent research has shown promising results for estimating structural area, volume, and population from Sentinel 1 and 2 data at a 10 by 10‐m spatial resolution. These studies were, however, conducted in homogeneous countries in Northern Europe. This study presents a deep learning methodology for population estimation in areas geographically distinct from Northern Europe. The two case study areas are Ghana and Egypt's Mediterranean coast, with supplementary ground truth data collected from Uganda, Kenya, Tanzania, Palestine, and Israel. This study aims to answer the question: How can we use Deep Learning to map structural area and type to derive population estimates for Ghana and Egypt based on Sentinel data? At 10 by 10‐m resolution, the accuracy of the presented area predictions is similar to the Google Open Buildings dataset. An intercomparison of the presented population predictions is made with global state‐of‐the‐art spatial population estimates, and the results are promising, with the proposed methodology showing comparable or better results than the state‐of‐the‐art for the study areas.
Several global and regional efforts have been undertaken to map human-made settlements and their characteristics, including building material, area, volume, and population. However, given the unprecedented amount of Earth observation data and processing power available, there is a timely need for developing novel approaches for mapping these characteristics at higher spatial and temporal resolution. Such information is key to effectively answering questions related to population growth, pollution, disaster management, risk assessments, spatial planning, and even generating business cases in peri-urban and rural areas. While such data is available from mapping agencies or commercial companies in some countries, there are many countries where this is not the case. The main objective of this study is to propose an Inception-ResNet inspired deep learning approach to estimate the characteristics and location of human-made structures, including estimates of population, based on Earth observation data from the Copernicus Programme. The study investigates the effects on prediction accuracy using data from different orbital directions and interferometric coherence from Sentinel 1 data and different band combinations of Sentinel 2 data as model input variables. The model is trained and evaluated on a nationwide Danish case study, where the national mapping agency provides high-quality open data on humanmade structures, which serves as the ground truth data for the study. Our findings reveal that it is possible to design models that, on average, perform within 2.6% total absolute percentage error for area predictions, 7.7% for volume and 17% for population at 10 by 10 m scale using only Copernicus data and deep learning models. The models achieved 98.68% binary accuracy for extracting structural area when all test sites were merged. Combining Sentinel 1 and 2 input variables yielded the best results, while adding interferometric coherence did not significantly improve accuracy. Furthermore, including data from both orbital directions of the Sentinel 1 constellation significantly improved model performance.
Financial Inclusion is, in many ways, a spatial planning issue: Where do financial institutions provide services, how far do customers travel to access mobile money, which services are available where and how is agent cash-flow handled? Utilising geodata can contribute significantly to measuring financial access and thus assist in improving Financial Inclusion by expanding the reach of services and locating areas of economic exclusion. This study presents a new Spatial Decision Support System, with a frontend embedded directly within a spreadsheet interface, that enables measuring and planning financial access through geospatial analysis and Earth Observation derived products. The purpose is to complement existing Financial Inclusion measures, which rely significantly on large-scale representative household surveys to quantify financial access and opportunity to proxy quality of inclusion. The Decision Support System relies on Earth Observation and Public Participatory GIS, which enables a decoupling from the census cycle and global reach. Our findings indicate that a geospatial approach to measuring and making decisions regarding the location of financial access points can positively affect both tracking and delivering Financial Inclusion and reducing the urban–rural service cliff-edge. Our proposed geospatial methodology is useful for decision-makers in two ways: a) It allows the measurement of the large-scale geospatial reach of financial services – useful for decision-makers, planners, politicians, national statistical offices, and NGOs in charge of tracking progress towards the Sustainable Development Goals. b) It helps with planning and optimising services for local financial entities such as mobile money agents, brick and mortar bank branches, and less formal saving mechanisms such as saving clubs. The Spatial Decision Support System is currently used by several Financial Service Providers in Ghana and undergoing implementation for one in North-western Tanzania.
This research describes the change in temperatures across approximately 270 tropical cities from 1960 to 2020 with a focus on urban warming. It associates urban growth indicators with temperature variations in tropical climate zones (tropical rainforest, tropical monsoon, and tropical wet-dry savanna). Our findings demonstrate that over time while temperatures have increased across the tropics, urban residents have experienced higher temperatures (minimum and maximum) than those living outside of cities. Moreover, in certain tropical zones, over the study period, temperatures have risen faster in urban areas than the background (non-urban) temperatures. The results also suggest that with continuing climate change and urban growth, temperatures will continue to rise at higher than background levels in tropical cities unless mitigation measures are implemented. Several fundamental characteristics of urban growth including population size, population density, infrastructure and urban land use patterns are factors associated with variations in temperatures. We find evidence that dense urban forms (compact residential and industrial developments) are associated with higher temperatures and population density is a better predictor of variation in temperatures than either urban population size or infrastructure in most tropic climate zones. Infrastructure, however, is a better predictor of temperature increases in wet-dry savanna tropical climates than population density. There are a number of potential mitigation measures available to urban managers to address heat. We focus on ecological services, but whether these services can address the projected increasing heat levels is unclear. More local research is necessary to untangle the various contributions to increasing heat in cities and evaluate whether these applications can be effective to cool tropical cities as temperature continue to rise. Our methods include combining several different datasets to identify differences in daily, seasonal, and annual maximum and minimum temperatures.
Urbanization and climate change are among the most important global trends affecting human well-being during the twenty-first century. One region expected to undergo enormous urbanization and be significantly affected by climate change is Africa. Studies already find increases in temperature and high temperature events for the region. How many people will be exposed to heat events in the future remains unclear. This paper attempts to provide a first estimate of the number of African urban residents exposed to very warm 15-day heat events (>42 degrees C). Using the Shared Socio-economic Pathways and Representative Concentration Pathways framework we estimate the numbers of exposed, sensitive (those younger than 5 and older than 64 years), and those in low-income nations, with gross national products of $4000 ($2005, purchasing power parity), from 2010 to 2100. We examine heat events both with and without urban heat island estimates. Our results suggest that at the low end of the range, under pathways defined as sustainable (SSP 1) and low relative levels of climate change (RCP 2.6) without including the urban heat island effect there will be large populations (>300 million) exposed to very warm heat wave by 2100. Alternatively, by 2100, the high end exposure level is approximately 2.0 billion for SSP 4 under RCP 4.5 where the urban heat island effect is included.
Abstract. Cities expand rapidly with international migration significantly contributing to urban growth and urban population change. However, cities miss out on a great opportunity of reclaiming valuable knowledge on future population distribution due to the lack of established tools and methodologies to project where it is more likely for people of specific socio-demographic groups to set up home. The present work suggests that spatially explicit projections can play a significant role as a tool for urban planning and for managing diversity creatively, especially when a combination of social, demographic and topographic data is utilized. Machine learning techniques have demonstrated capabilities to capture relationships among this plethora of urban features to estimate future population distribution. We present a flexible, ML-based methodology for high-resolution gridded population projections by demographic characteristics, and specifically by region of origin, for the capital region of Copenhagen, Denmark, by combining various socio-demographic and topographic input layers.