While question-answering (QA) has gained attention in recent years, an often overlooked condition for receiving a valid answer is the quality of the question. The ability to formulate sufficiently clear and specific questions is crucial, in particular, for indirect, geo-analytical question-answering systems that need to generate valid geoprocessing workflows. We present a new method for formulating geo-analytical questions and specifying their semantics by integrating a controlled natural language for expressing geo-analytic intentions with a semantic model of geo-analytical concept transformations. This method is implemented in a customized Google Blockly interface, where users can construct questions by selecting and connecting blocks representing different conceptual components of geo-analytical questions. We conducted a user study to compare the Blockly-based interface with a free-form interface. Results show that the Blockly-based interface effectively constrains question formulation, leading to more standardized, complete, and interpretable questions regardless of users' GIS expertise levels. However, the Blockly-based interface requires more learning effort and receives a medium-to-low usability rating. Our study shows the potential for template-based geo-analytical QA to foster a particular level of quality as input to QA systems. Future research should combine this controlled language with generative models to enable higher quality, adaptability, and scalability of question answering.
While recent WHO systematic reviews have comprehensively assessed the direct health effects of radiofrequency electromagnetic field (RF-EMF) exposure, its potential indirect impacts on human health via ecosystem disruption remain unstudied. Therefore, we propose a Planetary Health Impact Assessment (PHIA) approach, which incorporates both direct and ecologically mediated pathways. Developing the underlying framework requires a method for organizing and visualizing complex, interdisciplinary knowledge. This study explores an approach for constructing a PHIA framework in the form of knowledge graphs (KGs). Using RF-EMF exposure from mobile telecommunication technologies as a case study, we developed an expert-based KG in collaboration with 12 specialists. We further evaluated the potential of an artificial intelligence (AI)-based tool, incorporating Natural Language Processing (NLP) and Deep Learning, to extract relevant information from scientific literature and generate KGs to explore ways to enhance the expert-based approach. Experts developed and visualized jointly the hypothesized pathways linking RF-EMF exposure to direct health effects on organisms and indirect effects on human health through ecological consequences. The AI tool quickly processed large volumes of literature and visualized it into KGs with varied structures but required extensive expert validation due to limitations in precision and context sensitivity. The expert-based KG can serve as organizer of the available knowledge and as a first step in PHIA development. While AI tools offer potential for exploratory analysis, they currently require substantial human oversight and cannot replace expert judgment. The resulting KGs also identified possible gaps in the scientific literature.
The integration of large language models (LLMs) with geographic information science (GIScience) represents a new frontier in interdisciplinary research that combines advanced natural language processing with sophisticated spatial data analysis. This paper explores the synergistic potential of combining the natural language understanding and generation capabilities of LLMs with the expertise of GIScience in handling complex geospatial data. By exploring the specific contributions that LLMs can offer to GIScience, such as improving data processing, analysis, and visualization, and the mutual benefits that GIScience can offer to LLMs in terms of spatial reasoning and conceptual frameworks, we outline a comprehensive framework and a research agenda for this integration. Furthermore, we address the societal and ethical implications of this convergence, highlighting the challenges of bias, misinformation, and environmental impact. Through this exploration, we aim to set the stage for innovative applications in urban planning, environmental analysis, and beyond, while emphasizing the need for responsible use of AI.
Urban planning can help tackle environmental health issues. We demonstrate that agent-based simulation can discover unintended as well as intended environmental health effects and social inequalities of intervention scenarios. We developed, calibrated and validated UrbHealth-ABM, an empirically grounded agent-based model of Amsterdam The Netherlands, integrating data and models of individual mobility choices, traffic, air pollution, physical activity and personal exposure. We used the 2019 parking price increase as a natural experiment, confirming the models’ accuracy in predicting traffic reduction. Projections for the planned 2030 no emission zone show a significant reduction of nitrogen dioxide exposure for everyone and an increase in transport-related physical activity, especially for less affluent outer-city residents. However, disproportionate increases in their travel times raise equity concerns. Several 15-minutes city scenarios reveal that although driving may increase to distant destinations, nitrogen dioxide exposure decreases overall. Moreover, transport-related physical activity might decrease in the Amsterdam context due to shorter active travel distances, but travel time savings could be used for mitigation strategies.
Geographers have a shared intuition about geographic information, but this intuition is not well understood. In this work, we study the extent to which geo-analysts of varying skill levels possess the ability to distinguish between cartographic representations of two kinds of quantities, extensive quantities which can be summed, and intensive quantities which cannot. Data were collected from two sample groups in two online experiments. Participants were first tasked to select the odd-one-out from three maps with the same visualization style, and second to choose whether a given data attribute best fits a choropleth or a proportional symbol map visualization. The results show GIS expertise and cartographic skill predict answer accuracy and extensive quantities are distinguished correctly significantly less than intensive quantities. Evidence of the Use of Extensive and Intensive Quantity Concepts in Experienced Readers of Cartographic Maps Les g & eacute;ographes ont une intuition partag & eacute;e & agrave; propos de l'information g & eacute;ographique, mais cette intuition n'est pas tr & egrave;s bien comprise. Dans ce travail nous & eacute;tudions dans quelle mesure des g & eacute;o-analystes de diff & eacute;rents niveaux de comp & eacute;tences poss & egrave;dent la capacit & eacute; de diff & eacute;rentier des repr & eacute;sentations cartographiques de deux types de quantit & eacute;, les grandeurs extensives qui peuvent & ecirc;tre additionn & eacute;es et les grandeurs intensives qui ne peuvent pas l'& ecirc;tre. Les donn & eacute;es ont & eacute;t & eacute; collect & eacute;es aupr & egrave;s de deux groupes d'& eacute;chantillons lors de deux exp & eacute;riences en ligne. Il a & eacute;t & eacute; demand & eacute; aux participants en premier de s & eacute;lectionner la solution qui fait exception parmi trois cartes ayant le m & ecirc;me style de visualisation puis de dire pour un attribut propos & eacute; si la meilleure repr & eacute;sentation serait une visualisation choropl & egrave;the ou une visualisation avec des symboles proportionnels. Les r & eacute;sultats montrent que l'expertise SIG et les comp & eacute;tences cartographiques permettent de pr & eacute;dire la pr & eacute;cision de la r & eacute;ponse et que les grandeurs extensives sont significativement moins bien distingu & eacute;es que les grandeurs intensives.
Scenario microsimulations like agent-based models can account for feedbacks and spatio-temporal and social heterogeneity when projecting future intervention impacts. Addressing air pollution exposure requires traffic scenario models (i.e. of car-free zones). Traditional air pollution models do not meet all requirements for traffic scenario microsimulation: isolating traffic emission, integrating relevant dispersion moderators, while computationally efficient, interoperable and valid. We propose a hybrid model of land use regression- based baseline concentrations and on-road emissions in conjunction with cellular automata-based off-road dispersion. The model efficiently assesses air pollution, while accounting for meteorological and morphological dispersion processes. We calibrate using genetic algorithms and externally validate the model based on mobile measurements and fixed-site routine monitoring data of NO2 concentrations across Amsterdam. Our model achieves an external validation R2 of 0.60 and 0.48 s computation time in a 50 m x 50 m raster. Further, we successfully projected the NO2 reduction of the first Covid-19 lockdown traffic scenario (R2 0.57).
We argue that in order to justify a modeling approach for a particular purpose, we need to better understand the experimental structure that is supposed to be represented by a given model application. For this purpose, we introduce a logic for specifying causal as well as spatio-temporal experiments, based on which we reinterpret Sinton's structure of spatial information from a pragmatic, experimental viewpoint. We illustrate the use of this logic based on a landuse modeling example, showing to what extent remote sensing and simulation approaches can be justified by decomposing the example into experiments required for answering its main question.
Exposure is a central concept of the health and behavioural sciences needed to study the influence of the environment on the health and behaviour of people within a spatial context. While an increasing number of studies measure different forms of exposure, including the influence of air quality, noise, and crime, the influence of land cover on physical activity, or of the urban environment on food intake, we lack a common conceptual model of environmental exposure that captures its main structure across all this variety. Against the background of such a model, it becomes possible not only to systematically compare different methodological approaches but also to better link and align the content of the vast amount of scientific publications on this topic in a systematic way. For example, an important methodical distinction is between studies that model exposure as an exclusive outcome of some activity versus ones where the environment acts as a direct independent cause (active vs. passive exposure). Here, we propose an information ontology design pattern that can be used to define exposure and to model its variants. It is built around causal relations between concepts including persons, activities, concentrations, exposures, environments and health risks. We formally define environmental stressors and variants of exposure using Description Logic (DL), which allows automatic inference from the RDF-encoded content of a paper. Furthermore, concepts can be linked with data models and modelling methods used in a study. To test the pattern, we translated competency questions into SPARQL queries and ran them over RDF-encoded content. Results show how study characteristics can be classified and summarized in a manner that reflects important methodical differences.
Human behavior may be one of the most challenging phenomena to model and validate. This paper proposes a method for automatically extracting and compiling evidence on human behavior determinants into a knowledge graph. The method (1) extracts associations of behavior determinants and choice options in relation to study groups and moderators from published studies using Natural Language Processing and Deep Learning, (2) synthesizes the extracted evidence into a knowledge graph, and (3) sub-selects the model components and relationships that are relevant and robust. The method can be used to either (4a) construct a structurally valid simulation model before proceeding with calibration or (4b) to validate the structure of existing simulation models. To demonstrate the feasibility of the method, we discuss an example implementation with mode of transport as behavior choice. We find that including non-frequently studied significant behavior determinants drastically improves the model's explanatory power in comparison to only including frequently studied variables. The paper serves as a proof-of-concept which can be reused, extended or adapted for various purposes.
The concept of validity is a cornerstone of science. Given this central role, it is somewhat surprising to find that validity remains a rather obscure concept. Unfortunately, the term is often reduced to a matter of ground truth data, seemingly because we fail to come to grips with it. In this paper, instead, we take a purpose-based approach to the validity of spatio-temporal models. We argue that a model application is valid only if the model delivers an answer to a particular spatio-temporal question specifying some experiment including spatio-temporal controls and measures. Such questions constitute the information purposes of models, forming an intermediate layer in a pragmatic knowledge pyramid with corresponding levels of validity. We introduce a corresponding question-based grammar that allows us to formally distinguish among contemporary inference, prediction, retrodiction, projection, and retrojection models. We apply the grammar to corresponding examples and discuss the possibilities for validating such models as a means to a given end.
Abstract. There is an increasing trend of applying AIbased automated methods to geoscience problems. An important example is a geographic question answering (geoQA) focused on answer generation via GIS workflows rather than retrieval of a factual answer. However, a representative question corpus is necessary for developing, testing, and validating such generative geoQA systems. We compare five manually constructed geographical question corpora, GeoAnQu, Giki, GeoCLEF, GeoQuestions201, and Geoquery, by applying a conceptual transformation parser. The parser infers geo-analytical concepts and their transformations from a geographical question, akin to an abstract GIS workflow. Transformations thus represent the complexity of geo-analytical operations necessary to answer a question. By estimating the variety of concepts and the number of transformations for each corpus, the five corpora can be compared on the level of geo-analytical complexity, which cannot be done with purely NLP-based methods. Results indicate that the questions in GeoAnQu, which were compiled from GIS literature, require a higher number as well as more diverse geo-analytical operations than questions from the four other corpora. Furthermore, constructing a corpus with a sufficient representation (including GIS) may require an approach targeting a uniquely qualified group of users as a source. In contrast, sampling questions from large-scale online repositories like Google, Microsoft, and Yahoo may not provide the quality necessary for testing generative geoQA systems.
Current artificial intelligence (AI) approaches to handle geographic information (GI) reveal a fatal blindness for the information practices of exactly those sciences whose methodological agendas are taken over with earth-shattering speed. At the same time, there is an apparent inability to remove the human from the loop, despite repeated efforts. Even though there is no question that deep learning has a large potential, for example, for automating classification methods in remote sensing or geocoding of text, current approaches to GeoAI frequently fail to deal with the pragmatic basis of spatial information, including the various practices of data generation, conceptualization and use according to some purpose. We argue that this failure is a direct consequence of a predominance of structuralist ideas about information. Structuralism is inherently blind for purposes of any spatial representation, and therefore fails to account for the intelligence required to deal with geographic information. A pragmatic turn in GeoAI is required to overcome this problem.
The recent success of large language models and AI chatbots such as ChatGPT in various knowledge domains has a severe impact on teaching and learning Geography and GIScience. The underlying revolution is often compared to the introduction of pocket calculators, suggesting analogous adaptations that prioritize higher-level skills over other learning content. However, using ChatGPT can be fraudulent because it threatens the validity of assessments. The success of such a strategy therefore rests on the assumption that lower-level learning goals are substitutable by AI, and supervision and assessments can be refocused on higher-level goals. Based on a preliminary survey on ChatGPT's quality in answering questions in Geography and GIScience, we demonstrate that this assumption might be fairly naive, and effective control in assessments and supervision is required.
The term homeomerosity refers to when a whole and its parts are the same kind of thing. For instance, a computer and its processor can both be classified as machines. Homeomerosity is a prerequisite for meaningful addition and subtraction. For example, adding the area sizes of two independent regions gives another area size, but adding an area size and a number of hours yields a number with a peculiar unit. In earlier work, homeomerosity has been formalized with respect to mereological parthood, but not in concurrence with a notion of class subsumption. Both are essential to homeomerosity, as a part can only be observed to be of the same kind as the whole if they are observed to be of some kinds in the first place. In this work, we use formal concept analysis to organize conceptual representations of parts and wholes in a shared contextual model. In our doing so, we show wholes and parts can be represented by sub-concepts of a concept with respect to which they are homeomerous.
Since the European information economy faces insufficient access to and joint utilization of data, data ecosystems increasingly emerge as economical solutions in B2B environments. Contrarily, in B2C ambits, concepts for sharing and monetizing personal data have not yet prevailed, impeding growth and innovation. Their major pitfall is European data protection law that merely ascribes human data subjects a need for data privacy while widely neglecting their economic participatory claims to data. The study reports on a design science research (DSR) approach addressing this gap and proposes an abstract reference system architecture for an ecosystem centered on humans with personal data. In this DSR approach, multiple methods are embedded to iteratively build and evaluate the artifact, i.e., structured literature reviews, design recovery, prototyping, and expert interviews. Managerial contributions embody novel design knowledge about the conceptual development of human-centric B2C data ecosystems, considering their legal, ethical, economic, and technical constraints.
BACKGROUND AND AIM: With ever more people living in cities worldwide, it becomes increasingly important to understand the positive and negative impacts of the urban habitat on livability, health behaviors and health outcomes. However, implementing interventions that tackle the exposome in complex urban systems can be costly and have long-term, sometimes unforeseen and indirect, impacts. Hence, it is crucial to not only assess the health impact of interventions, but also its cost-effectiveness and the social distributional impacts of possible urban exposome interventions before implementing them. METHOD: Spatial agent-based modeling can capture complex behavior-environment interactions, exposure dynamics, and social outcomes in a spatial context. We present our work on agent-based modeling of transport interventions in the context of the city of Amsterdam. The goal of our model is to capture the health impacts of multiple hypothetical transport intervention scenarios, such as car-free zones, replacement of parking places with active transport infrastructure, and separation of active and motorized transport routes. Our agent-based modeling approach entails the integration of a behavioral model of people's mobility choices and dynamic physical models of environmental stressors (e.g. air pollution). RESULTS: Together these sub-models result in an exposure interaction that approximates personal behavioral and environmental exposure for different population groups within an urban environment, e.g. based on demographics, the neighbourhoods they live in or on their social economic circumstances. Consequently, the accumulated health impacts for different interventions and population groups are assessed using exposure-response functions. CONCLUSIONS: We present our model architecture, the strength and limitations of the method, and our findings on effects of transport interventions.
The ever-growing amounts of data offer companies many opportunities for data-driven-value generation which, in turn, can be multiplied by leveraging data across company boundaries in evolving data ecosystems. However, while such systems increasingly emerge in B2B environments enabling systematic sharing and utilization of “industrial data”, comparable concepts in B2C ambits have not yet prevailed. Despite the rising importance of personal data in the information economy, B2C data ecosystems represent a widely unexplored research area. To remedy this gap, the study generates design principles for human centric B2C data ecosystems to aid in their development. For this purpose, a qualitative interview study with experts of interdisciplinary domains and a structured literature review are conducted both embedded into a methodology for generating design principles. On this basis, derived design principles help to understand peculiarities of data ecosystems in B2C ambits and provide solutions to overcome their obstacles identified in the empirical investigation.
Edzer J. Pebesma合作论文数Institute for Geoinformatics, University of Muenster2