Georeferencing text documents has typically relied on either gazetteer-based methods to assign geographic coordinates to place names or on language modelling approaches that associate textual terms with geographic locations. However, many location descriptions specify positions relatively with spatial relationships, making geocoding based solely on place names or geo-indicative words inaccurate. This issue frequently arises in biological specimen collection records, where locations are often described through narratives rather than coordinates if they pre-date GPS. Accurate georeferencing is vital for biodiversity studies, yet the process remains labour-intensive, leading to a demand for automated georeferencing solutions. This paper explores the potential of Large Language Models (LLMs) to georeference complex locality descriptions automatically, focusing on the biodiversity collections domain. We first identified effective prompting patterns, then fine-tuned an LLM using Quantized Low-Rank Adaptation (QLoRA) on biodiversity datasets from multiple regions and languages. Our approach outperforms existing baselines with an average, across datasets, of 65% of records within a 10 km radius, for a fixed amount of training data. The best results (New York state) were 85% within 10 km and 67% within 1 km. The selected LLM performs well for lengthy, complex descriptions, highlighting its potential for georeferencing intricate locality descriptions.
Accurately georeferencing textual locality descriptions in natural history records (e.g. "9m [miles] E. of Lake Manapouri, junction of Manapouri Road \& State Highway 94") remains a major bottleneck for mobilizing biodiversity data. Manual georeferencing is labour-intensive and difficult to scale, while existing automated methods rely primarily on linguistic cues and neglect the spatial reasoning central to human map-based georeferencing. This paper introduces a fully automated multimodal pipeline for georeferencing complex location descriptions that integrates Foundational Models including Large Language Models (LLMs) and Vision-Language Models (VLMs), and geospatial knowledge bases into a unified engineering pipeline. The system comprises five modular components that perform spatial entity extraction, toponym refinement, map generation, and visual–textual grounding respectively. The proposed GeoVLM employs parameter-efficient fine-tuning with discrete grid-cell prediction to achieve efficient, interpretable spatial reasoning over cartographic inputs. Experiments conducted on a real-world dataset of herbarium records show that the multimodal approach achieves up to three-fold improvements in georeferencing accuracy compared to existing LLM and rule-based baselines, with over 90\% of records located within 10 km of the ground truth. The results demonstrate that multimodal learning can bridge linguistic and cartographic representations, providing a scalable and generalizable solution for automated geospatial intelligence. The results highlight VLMs’ capacity to comprehend cartographic information and establish a foundation for scalable, globally generalizable georeferencing of unstructured textual data across domains.
Effective disease surveillance in wild fish populations is essential for food security and biodiversity conservation, but data acquisition can be limited by ad hoc reporting and resource-intensive laboratory diagnostics. We developed and evaluated a computer vision pipeline to detect saprolegniasis-like infections, a devastating disease in salmonids that manifests as visible signs. Compiling a dataset of 4526 images (494 infected, 4032 healthy) from citizen science platforms and stakeholders, we used data augmentation to address the significant class imbalance. We then fine-tuned and compared four pre-trained convolutional neural network architectures (EfficientNetV2S, EfficientNetV2B0, ResNet50, and MobileNetV3S), chosen to represent a range of standard and efficient models, to classify healthy versus infected fish across datasets of varying host taxonomic specificity. The EfficientNetV2S model achieved the highest performance on a Salmo spp. specific dataset, with a mean recall (proportion of infected fish images correctly identified) of 0.898 (+/- 0.043) and precision (proportion of correctly identified infected fish among all fish identified as infected) of 0.858 (+/- 0.067). Performance varied with host taxonomic scope, with models achieving lower metrics on broader host taxa datasets. Despite challenges including variable image quality, water surface reflections, and inherent class imbalance, these results show computer vision can support large-scale disease surveillance in wild fish populations. Computer vision-based surveillance could enable earlier outbreak detection and targeted diagnostics, improving freshwater ecosystem health management. While successful implementation hinges on acquiring sufficient high-quality imagery, this study highlights the potential of applying tailored Artificial Intelligence tools for monitoring visually detectable diseases across diverse wildlife species.
Biological specimen records collected by natural history institutions constitute a rich source of temporal geographic knowledge, capturing biodiversity information about regional landscapes as they were recorded at different times. Using digitised data from the Allan Herbarium (New Zealand), this study identifies place names in these specimen locality descriptions that are absent from current gazetteers; we refer to these as non-gazetteer place names (NGPs). These place names are typically historical, vernacular, or colloquial and were used as landmarks to describe a specimen's location at the time of collection. We then investigate the problem of georeferencing the NGPs using only the limited information available in the specimen records. To resolve this, we leverage repeated occurrences of the same place name across specimen records with different specimen locations and spatial relation terms, extracting and inverting these relations to derive constraints on NGP locations. This approach is instantiated within deterministic, probabilistic, and LLM-based methods, enabling a comparative analysis of their strengths and limitations for text-based spatial inference. On a pseudo-NGP benchmark, probabilistic inference achieves the highest accuracy (median error 1.43 km; A@1 km 36
The development, integration, and maintenance of geospatial databases rely heavily on efficient and accurate matching procedures of Geospatial Entity Resolution (ER). While resolution of points-of-interest (POIs) has been widely addressed, resolution of entities with diverse geometries has been largely overlooked. This is partly due to the lack of a uniform technique for embedding heterogeneous geometries seamlessly into a neural network framework. Existing neural approaches simplify complex geometries to a single point, resulting in significant loss of spatial information. To address this limitation, we propose Omni, a geospatial ER model featuring an omni-geometry encoder. This encoder is capable of embedding point, line, polyline, polygon, and multi-polygon geometries, enabling the model to capture the complex geospatial intricacies of the places being compared. Furthermore, Omni leverages transformer-based pre-trained language models over individual textual attributes of place records in an Attribute Affinity mechanism. The model is rigorously tested on existing point-only datasets and a new diverse-geometry geospatial ER dataset. Omni produces up to 12% (F1) improvement over existing methods. Furthermore, we test the potential of Large Language Models (LLMs) to conduct geospatial ER, experimenting with prompting strategies and learning scenarios, comparing the results of pre-trained language model-based methods with LLMs. Results indicate that LLMs show competitive results.
Millions of biological sample records collected in the last few centuries archived in natural history collections are un-georeferenced. Georeferencing complex locality descriptions associated with these collection samples is a highly labour-intensive task collection agencies struggle with. None of the existing automated methods exploit maps that are an essential tool for georeferencing complex relations. We present preliminary experiments and results of a novel method that exploits multi-modal capabilities of recent Large Multi-Modal Models (LMM). This method enables the model to visually contextualize spatial relations it reads in the locality description. We use a grid-based approach to adapt these auto-regressive models for this task in a zero-shot setting. Our experiments conducted on a small manually annotated dataset show impressive results for our approach (similar to 1 km Average distance error) compared to uni-modal georeferencing with Large Language Models and existing georeferencing tools. The paper also discusses the findings of the experiments in light of an LMM's ability to comprehend fine-grained maps. Motivated by these results, a practical framework is proposed to integrate this method into a georeferencing workflow.
Large-scale disasters can often result in catastrophic consequences on people and infrastructure. Situation awareness about such disaster impacts generated by authoritative data from in-situ sensors, remote sensing imagery, and/or geographic data is often limited due to atmospheric opacity, satellite revisits, and time limitations. This often results in geo-temporal information gaps. In contrast, impact-related social media posts can act as "geo-sensors" during a disaster, where people describe specific impacts and locations. However, not all locations mentioned in disaster-related social media posts relate to an impact. Only the impacted locations are critical for directing resources effectively. e.g., "The death toll from a fire which ripped through the Greek coastal town of #Mati stood at 80, with dozens of people unaccounted for as forensic experts tried to identify victims who were burned alive #Greecefires #AthensFires #Athens #Greece." contains impacted location "Mati" and non-impacted locations "Greece" and "Athens". This research uses Large Language Models (LLMs) to identify all locations, impacts and impacted locations mentioned in disaster-related social media posts. In the process, LLMs are fine-tuned to identify only impacts and impacted locations (as distinct from other, non-impacted locations), including locations mentioned in informal expressions, abbreviations, and short forms. Our fine-tuned model demonstrates efficacy, achieving an F1-score of 0.69 for impact and 0.74 for impacted location extraction, substantially outperforming the pre-trained baseline. These robust results confirm the potential of fine-tuned language models to offer a scalable solution for timely decision-making in resource allocation, situational awareness, and post-disaster recovery planning for responders.
Gazetteers typically store data on place names, place types, and the associated coordinates. They play an essential role in disambiguating place names in online geographical information retrieval systems for navigation and mapping, detecting and disambiguating place names in text, and providing coordinates. Currently, there are many gazetteers in use derived from many sources, with no commonly accepted standard for encoding the data. Most gazetteers are also very limited in the extent to which they represent the multiple facets of the named places yet they have potential to assist user search for locations with specific physical, commercial, social, or cultural characteristics. With a focus on understanding digital gazetteer technologies and advancing their future effectiveness for information retrieval, we provide a review of data sources, components, software and data management technologies, data quality and volunteered data, and methods for matching sources that refer to the same real-world places. We highlight the need for future work on richer representation of named places, the temporal evolution of place identity and location, and the development of more effective methods for data integration.
Recent approaches to geo-referencing X posts have focused on the use of language modelling techniques that learn geographic region-specific language and use this to infer geographic coordinates from text. These approaches rely on large amounts of labelled data to build accurate predictive models. However, obtaining significant volumes of geo-referenced data from Twitter, recently renamed X, can be difficult. Further, existing language modelling approaches can require the division of a given area into a grid or set of clusters, which can be dataset-specific and challenging for location prediction at a fine-grained level. Regression-based approaches in combination with deep learning address some of these challenges as they can assign coordinates directly without the need for clustering or grid-based methods. However, such approaches have received only limited attention for the geo-referencing task. In this paper, we adapt state-of-the-art neural network models for the regression task, focusing on geo-referencing wildlife Tweets where there is a limited amount of data. We experiment with different transfer learning techniques for improving the performance of the regression models, and we also compare our approach to recently developed Large Language Models and prompting techniques. We show that using a location names extraction method in combination with regression-based disambiguation, and purely regression when names are absent, leads to significant improvements in locational accuracy over using only regression.
Museum collection records are a source of historic data for species occurrence, but little attention is paid to the associated descriptions of habitat at the sample locations. We propose that artificial intelligence methods have potential to use these descriptions for reconstructing past habitat, to address ecological and evolutionary questions.
Topological data analysis (TDA) is an emerging field of research, which considers the application of topology to data analysis. Recently, these methods have been successfully applied to research problems in the field of geographical information science (GIS) and there is much potential for future applications. In this article, we provide an introduction to the fundamentals of TDA for GIS researchers and practitioners and highlight specific benefits that TDA methods provide relative to some conventional methods. We focus on the method of persistent homology, which is the most commonly used TDA method. We describe how persistent homology can be applied to data types commonly encountered in the GIScience domain, namely sets of points, networks and sequences of images. We also describe the application of persistent homology to two specific GIS problems, which are the point pattern analysis of UK city pubs and the analysis of UK rainfall radar imagery. In each case we stress the specific benefits of TDA methods that include, for example, generating an output signature in a form that can be subject to subsequent analyses; identification of void regions in point patterns; and providing a relatively simple method to track objects in spatio-temporal images.
Vast numbers of biological specimens (e.g. flora, fauna, soils) are stored in collections globally. Many of these have only a natural-language location description, such as â200ft above and south of main highway, 1.1 miles west of Porters Passâ, and numerical coordinates are unknown. The BioWhere project is pioneering methods to automatically determine the geographic coordinates (georeferences) of complex location descriptions. Particular challenges are posed by the variable accuracy of recent and historical data that might be used to train models to predict geographic coordinates from the natural-language descriptions; by the presence of historical place names in the descriptions that are not stored in existing gazetteers; and by the vague and context-sensitive nature (e.g. above, on, south of) of the descriptions. We are addressing these challenges by extending the latest transformer-based deep learning models to parse locality descriptions, and to build models for specific spatial terms that incorporate geographic context and data quality to more accurately predict georeferences. We also describe a gazetteer that contains enriched cultural content to support georeferencing of historical records, and to serve as a store of New Zealand MÄori cultural knowledge for future generations.
The locations of objects are often described in natural language relative to some other object using vague and context-sensitive spatial relation terms (e.g
Spatial prepositions have been studied in some detail from multiple disciplinary perspectives. However, neither the semantic similarity of these prepositions, nor the relationships between the multiple senses of different spatial prepositions, are well understood. In an empirical study of 24 spatial prepositions, we identify the degree and nature of semantic similarity and extract senses for three semantically similar groups of prepositions using t-SNE, DBSCAN clustering, and Venn diagrams. We validate the work by manual annotation with another data set. We find nuances in meaning among proximity and adjacency prepositions, such as the use of close to instead of near for pairs of lines, and the importance of proximity over contact for the next to preposition, in contrast to other adjacency prepositions.
We propose a novel model of street network connectivity which uses a method from the field of applied topology called “persistent homology”. The output from this model is a pair of density functions which model the relative strength and frequency of connected components and cycles in the network. In this context, strength is a function of street type, such as motorway or residential, with more significant street types providing greater connectivity. The pair of density functions output from the model are easily interpreted and provide novel insights into the connectivity properties of different street networks. We demonstrate the usefulness of this model through an analysis of UK and US city street networks. This analysis identifies tangible similarities and differences in the connectivity of different cities plus ways in which the connectivity of individual cities might be improved.
. Geographical context is required of many information retrieval tasks in which the target of the search may be documents, images or records which are referenced to geographical space only by means of place names. Often there may be an imprecise match between the query name and the names associated with candidate sources of information. There is a need therefore for geographical information retrieval facilities that can rank the relevance of candidate information with respect to geographical closeness as well as semantic closeness with respect to the topic of interest. Here we present an ontology of place that combines limited coordinate data with qualitative spatial relationships between places. This parsimonious model of place is intended to support information retrieval tasks that may be global in scope. The ontology has been implemented with a semantic modelling system linking non-spatial conceptual hierarchies with the place ontology. An hierarchical distance measure is combined with Euclidean distance between place centroids to create a hybrid spatial distance measure. This can be combined with thematic distance, based on classification semantics, to create an integrated semantic closeness measure that can be used for a relevance ranking of retrieved objects.
Spatial language incorporates descriptions of locations, routes, and landscapes, and is used by humans daily. Research has addressed a wide range of aspects of spatial language, including its form; the ways in which it is selected and applied; and cognitive, geometric, and functional factors affecting its use. Furthermore, much work has been done on the automation of spatial language extraction, analysis, interpretation, and generation. To introduce the Special Issue on this broad topic, this paper reviews spatial language research framed by an extension to the well-known semantic triangle, the "spatial semantic pyramid," which represents both human spatial language and relevant computational research. By introducing it, we hope to stimulate discussion about gaps and future directions in this important research field.
Despite the potential of social media for environmental monitoring, concerns remain about the quality and reliability of the information automatically extracted. Notably there are many observations of wildlife on Twitter, but their automated detection is a challenge due to the frequent use of wildlife related words in messages that have no connection with wildlife observation. We investigate whether and what type of supervised machine learning methods can be used to create a fully automated text classification model to identify genuine wildlife observations on Twitter, irrespective of species type or whether Tweets are geo-tagged. We perform experiments with various techniques for building feature vectors that serve as input to the classifiers, and consider how they affect classification performance. We compare three classification approaches and perform an analysis of the types of features that are indicative for genuine wildlife observations on Twitter. In particular, we compare some classical machine learning algorithms, widely used in ecology studies, with state-of-the-art neural network models. Results showed that the neural network-based model Bidirectional Encoder Representations from Transformers (BERT) outperformed the classical methods. Notably this was the case for a relatively small training corpus, consisting of less than 3000 instances. This reflects that fact that the BERT classifier uses a transfer learning approach that benefits from prior learning on a very much larger collection of generic text. BERT performed particularly well even for Tweets that employed specialised language relating to wildlife observations. The analysis of possible indicative features for wildlife Tweets revealed interesting trends in the usage of hashtags that are unrelated to official citizen science campaigns. The findings from this study facilitate more accurate identification of wildlife-related data on social media which can in turn be used for enriching citizen science data collections.
Spatial relations in natural language are frequently expressed through prepositions. Thus, in the locative expressions “New York in the United States” and “the house on the river” the prepositions “in” and “on,” respectively, serve to communicate the relationships in space between the subject and object of the preposition. Automatic detection of the use of prepositions in a spatial and in particular a geo‐spatial sense that refers to geographic context is of interest in supporting automated methods for determining the actual geographic location referred to by locative expressions. This work focuses on disambiguation of prepositions in natural language, with the goal of distinguishing whether a preposition is used in a specifically geo‐spatial sense. We conduct machine learning experiments that demonstrate the clear benefit for geo‐spatial sense detection of using transformer model deep learning methods when compared with a variety of methods, that include Naive Bayes, support vector machine, and random forest classifiers with handcrafted linguistic features, and a bag of words approach with a meta‐classifier that adds geo‐spatial features. The best performance was obtained with the Bidirectional Encoder Representation from Transformer‐based XLNet transformer model, with a best precision of 0.96 and an F1 score of 0.94 when evaluated on a corpus of natural language expressions that were annotated for this task. We also conducted experiments to detect generic spatial sense, in which the best F1 score, of 0.95, was again obtained with XLNet.