
Large-scale pretrained foundation models have transformed the machine learning paradigm. However, their applicability to Geospatial Tabular Data (GTD), particularly for regression tasks, remains poorly understood. This study systematically investigated TabPFN, a recent tabular foundation model, across synthetic data-generating processes and real-world geospatial regression tasks. We showed that while TabPFN generally outperforms state-of-the-art baselines, its performance degrades on large datasets and under strong, localised spatial dependence. To address these limitations, we proposed Geospatial Sparse Attention (GSA), a spatially informed inference strategy that injects geospatial inductive bias into the off-the-shelf TabPFN framework. The resulting model, TabPFN-GSA, prunes redundant attention calculations to better balance local and global spatial effects while improving scalability to large datasets. Empirical results showed that TabPFN-GSA delivers more accurate and robust predictions, particularly for large-scale GTD. Theoretically, this work advances our understanding of the strengths and limits of tabular foundation models in spatial contexts. Methodologically, it offers TabPFN-GSA, a principled, spatially explicit bridge between classical spatial modelling and modern foundation models.
Methods such as multi-factor geographically weighted machine learning (MFGWML), geographically neural network weighted regression (GNNWR), and its enhanced version-geographically convolutional neural network weighted regression (GCNNWR)-have improved spatial non-stationarity modeling. However, they are over-reliant on spatial proximity weighting and underutilize spatial neighborhood information. Therefore, this study proposes a novel Geographically Conditioned Multiscale Convolutional Neural Network (GCMCNN) that downscales MODIS land surface temperature (LST) data to 100 m resolution by integrating two- and three-dimensional surface features. GCMCNN enhances the capture of spatial heterogeneity through a geographically aware (geo-aware) mechanism conditioned by location and attributes, while improving the characterization of spatial correlation by developing a multiscale convolutional neural network (CNN). GCMCNN model outperformed four benchmark models (random forest, MFGWML, GNNWR, and GCNNWR) in LST downscaling, reducing RMSE by 15.4-26.16% and MAE by 15.05-23.59%, and increasing R2 by 39.81-148.09%. GCMCNN improved the model performance of CNN through its geo-aware mechanism and multi-level spatial feature extraction and fusion, reducing RMSE and MAE by 6.6% and 8%, respectively, and increasing R2 by 7.5%. Furthermore, spatial block cross-validation better reflects the model's true generalization ability than random cross-validation. This study highlights the importance of accurately modeling spatial heterogeneity and correlation in LST downscaling.
Geography and neuroscience share a core interest in understanding human behaviour in spatial environments. Yet, interdisciplinary collaboration is often limited by differences in methodology and epistemological assumptions. To help bridge this divide, we introduce a translational link between cognitive models of spatial processing in neuroscience and representations of geographic space. Building on long-standing theories that the brain predicts possible futures, the predictive map hypothesis suggests that locations in space are encoded according to their association with possible locations in the future. Here, we adapt a formal instantiation of this idea, the successor representation (SR), to urban space resulting in the geographic successor representation (gSR). We show that this cognitive model of geographic representation produces compelling unique features of urban space while remaining closely aligned with brain mechanisms of spatial processing. We outline several promising directions for extending this work, and propose that the gSR and its variants may provide spatial representations capable of supporting deeper integration between geographic and neuroscientific research.
Sketch maps are a widely utilized method for assessing how participants encode and externalize knowledge of large-scale environments. Even though many such spaces include an important vertical component, participants tend to omit, or distort vertical information when drawing 2D sketch maps. Recent advancement in Virtual Reality sketching interfaces open the question whether sketch maps could be realised in 3D. We recruited participants across two experiments - one in a layered indoor environment, one in a volumetric urban environment - and assessed their 2D and 3D sketch maps by coding the occurrence and correctness of qualitative spatial relations across all three dimensions. Results show that when tasked with drawing vertically-complex spaces on traditional 2D pen-and-paper medium, participants omit a lot of (mostly vertical) information that they do in fact store in their cognitive model of that space and can externalise in a Virtual Reality-based 3D sketch map. This indicates that 2D sketch maps may be a bottleneck for expressing valid parts of participants' spatial knowledge of vertically-complex spaces, while 3D sketch maps offer a promising alternative.
Geographical study areas (GSAs) anchor empirical research to specific locations and are essential for geographically aware knowledge organization, retrieval and spatial meta-analysis. However, GSA information is rarely stored in structured form in bibliographic databases and instead appears as unstructured text in article titles and abstracts, hindering large-scale spatial analyses of scientific knowledge production. This study proposes an LLM-assisted unified framework to systematically extract, disambiguate and classify multidimensional GSA information from large-scale article metadata. The proposed method follows an 'Expert-Teacher-Student' framework. First, a dual-dimensional GSA taxonomy integrating spatial scale and spatial attributes was constructed through expert-LLM collaboration. Second, a retrieval-augmented annotation pipeline generated high-quality supervision data by combining LLM ensemble reasoning with external geospatial knowledge verification. Third, a lightweight unified model was developed via parameter-efficient fine-tuning to jointly perform GSA extraction and classification, reducing annotation costs and mitigating error propagation. Experiments demonstrate strong performance with high computational efficiency. Applying the framework to 163,781 geography-related articles (2010-2024) reveals significant research attention-population mismatch, epistemic biases and scale disparities in global knowledge production. The proposed framework advances geographically aware literature mining and provides a scalable foundation for spatial bibliometrics and GIScience.
Spatial cross-sectional data encapsulate rich information on spatial processes, forming a critical foundation for examining causation between variables. Detecting and quantifying such causation is essential for understanding complex natural and human phenomena. Measuring causal strengths from spatial cross sectional data, however, remains challenging, as existing methods often suffer from high false positive rates when quantifying causation. To address this gap, we propose a Geographical Cross Mapping Cardinality (GCMC) model that quantifies causal strength based on the intersectional cardinality of neighborhoods in reconstructed state space, and incorporates the DeLong placement method to evaluate the statistical significance of causal strength estimates. We validate GCMC using a simulated three variable causal benchmark and three representative spatial cross sectional datasets with known causal structures, and further assess its sensitivity to observational noise. Results demonstrate that GCMC effectively captures causation across weak, moderate, and strong coupling regimes while maintaining a low false positive rate and robust performance under noise. As a new extension of empirical dynamic modeling for spatial cross sectional data, GCMC complements existing methods and enables more reliable spatial causal inference.
Causal inference in geographical sciences faces the challenge of isolating treatment effects from high-dimensional observational data, complicated by spatial non-stationarity and persistent confounding. Double machine learning (DML) offers a powerful solution for high-dimensional debiasing through orthogonalization and cross-fitting, but traditional variants overlook spatial heterogeneity by treating space as a simple covariate. To address this, we introduce DML-Geo, an ensemble extension of DML for estimating spatially varying causal effects. Retaining the orthogonalization procedure of DML at its first stage, DML-Geo augments the second stage with three complementary estimators, namely a linear regression model for covariate-driven effects, a generalized additive model (GAM) for spatially smoothed additive effects, and geographically weighted regression (GWR) for localized patterns. Robustness is further enhanced by an adaptive weighting scheme based on inter-model correlations to aggregate outputs from these variants, complemented by a bootstrap procedure for significance testing. Extensive simulations confirm DML-Geo's superior precision and stability relative to its component models and competing baselines. In real-world applications to housing prices and mental health outcomes, DML-Geo uncovers interpretable spatial causal effect patterns, offering place-specific insights to support policy decisions. DML-Geo provides a flexible toolkit for geospatial causal inference that does not require causal graphs or strong structural assumptions.