Media texts convey emotions and stances that can shape the evolution of public opinion, which calls for comparable quantitative sentiment analysis. However, most existing approaches assign a single sentiment score to an entire article, making it difficult to distinguish functional differences across sentences and to clarify who expresses the sentiment and what the evaluation targets, thereby limiting interpretability and cross-source comparability. To address this issue, we propose a fine-grained sentiment quantification method for media texts that jointly considers sentence types and opinion holder-target structure. The method obtains sentence-level sentiment scores and simultaneously extracts sentence types, opinion holders, and opinion targets, enabling article-level structured quantification and comparison under a unified evaluation setting. In our implementation, a large language model (LLM) is primarily used for semantic parsing and structured extraction. Experiments demonstrate that the proposed method delivers stable performance on the sentiment score regression task (R2 = 0.899, MAE = 0.088, MSE = 0.027; relative to the strongest fine-tuned pretrained language model baseline in our comparison, RoBERTa with R2 = 0.871, this corresponds to a 2.8-percentage-point gain in R2 and an 8.3% reduction in MAE), and effectively supports opinion holder-target identification (holder weighted average F1 = 0.812; target loose F1 = 0.691 in a supplementary evaluation). Building on these outputs, the method can further reveal the spatial distribution of sentiment bias in global media coverage, highlighting relative sentiment patterns in cross-national narratives.
Extracting spatiotemporal disaster knowledge from massive, heterogeneous social media data is crucial for urban flood management but it remains technically challenging. This study proposes a unified framework that integrates Chinese Multimodal Large Language Models with agents to automate cross-modal extraction, using water depth as a representative case. Systematic evaluations against deep learning baselines demonstrate that knowledge-guided prompting enhances geographic extraction precision by 10%-25%. Specifically, DeepSeek R1 and Doubao-1.5-thinking-vision-pro excel in textual and visual tasks, respectively, and they are jointly adopted to maximize overall performance. The framework is applied to urban flood events in China (July-August 2021) to demonstrate practical utility. The framework constructs a structured database containing 96,826 spatiotemporal water depth records by processing 1.53 million texts and 240,000 images. The results successfully capture the evolution of flood events, verifying the robustness of the proposed approach. This study establishes a validated, data-driven approach using agent-based Multimodal Large Language Models, providing essential knowledge infrastructure for post-event analysis and urban flood risk assessment.
High-quality multimodal datasets are essential for developing vision-language models, yet publicly available figure-text resources in specialized scientific domains remain limited. To address this gap, we present a large-scale figure-text pair dataset constructed from map-related scientific literature (FTPD-ML). The dataset was derived from 96,859 scientific publications in cartography, geography, remote sensing, and related disciplines, and provides 75,702 publicly redistributable figure-text pairs under Creative Commons Attribution (CC BY) licenses with multiple levels of textual representations, including original figure captions, standardized caption variants, and contextual paragraph descriptions, together with bibliographic metadata. To ensure compliance with copyright and licensing requirements, records from non-redistributable publications are represented by metadata only. The dataset supports a range of multimodal research tasks, including image-text retrieval, image captioning, scientific document understanding, and cartography-oriented vision-language studies.
China's poverty alleviation and elimination campaign (PAEC, 2015-2021) aimed to eliminate absolute poverty and reduce social inequality, making improvements in the accessibility and equity of basic educational facilities a key component. However, existing research on the accessibility of educational facilities in China has predominantly focused on developed urban areas or specific regions, lacking nationwide spatiotemporal assessments. To address this gap, this study systematically evaluates the accessibility and equity of basic educational facilities (kindergartens, primary, and secondary schools) across mainland of China during the PAEC. Travel time cost, derived using multi-source geospatial datasets and the nearest neighbor method, was adopted as the primary indicator for assessing educational accessibility and equity. The results revealed pronounced disparities in accessibility between Eastern and Western China. Overall, both the accessibility and equity of basic educational facilities improved substantially during the campaign, with the average travel time per person reducing by approximately 50 %. Notably, the rate of improvement in impoverished regions was nearly double that observed in non-impoverished areas. Although widespread improvements in educational equity have occurred, the urban-rural disparity persists as a primary barrier to achieving comprehensive educational fairness. This study offers empirical evidence and methodological innovations for optimizing educational resource allocation and provides high-resolution temporal data to support the monitoring and evaluation of progress toward the Sustainable Development Goals (SDGs) in education.
The Himalayas host the planet's highest-elevation mountain forests, serving as natural sentinels of climate change. Despite evidence of accelerated shifting alpine treelines, changes in the entire Himalayan forest system remain understudied. In this study, we derive tree cover from long-term Landsat observations using machine learning algorithms and find a significant overall increase in tree cover by 2.5 percentage points from 1990 to 2020. At the scale of the entire Himalayas, the fastest increase occurred at higher elevations (1,000 to 3,500 m above sea level (a.s.l.)) at an annual rate exceeding 0.15% yr−1, whereas within the monsoonal Himalayas, even higher rates occurred between 1,000 and 2,500 m a.s.l. (>0.3% yr−1). We analyze forest change over the three decades in relation to climate and land cover datasets and validate the results with Google Earth Pro imagery. Forest gains in the monsoonal and westerly Himalayas were mainly associated with rising temperatures and changing precipitation patterns, respectively. Forest losses were dominated by the expansion of cropland and built-up areas, especially in the lower elevations. These findings highlight the value of long-term satellite records for detecting elevation-dependent forest changes and supporting sustainable forest management in the Himalayas.
Traditional deep learning methods and econometric model have played a crucial role in the field of data mining, particularly in the prediction of socioeconomic outcomes. However, socioeconomic information is unable to be directly extracted from remote sensing data. So, in this paper, we propose a method to leverage transfer learning to predict socioeconomic indicators (outcomes) through satellite imagery. Specifically, we use road network types as a proxy for socioeconomic factors, which is more effectively and stably than using nightlight. We have extracted eleven distinct road topological features to generate reasonable road network types. Given the unique characteristics of road networks, we have constructed and fine-tuned a hybrid pre-trained model that combines ResNet50 and Vision Transformer architectures for the transfer learning task. Through extensive experiments conducted across multiple regions, we demonstrated that our approach outperforms state-of-the-art methods in this field. This work highlights the potential of leveraging road network types as a proxy for socioeconomic information and the effectiveness of our transfer learning-based framework in extracting valuable insights from satellite imagery to support socioeconomic policy decisions. The code had released in https://github.com/xiachan254/PredSocecOut .
PM2.5 pollution remains a critical environmental and public health challenge in China despite post-2015 improvements. However, our understanding of the spatially heterogeneous and nonlinear associations of its driving factors remains limited. To fill this gap, we employ an explainable geospatial artificial intelligence (GeoAI) framework that integrates the geographical random forest (GRF) model and the Shapley additive explanations (SHAP) approach to examine the associations between 16 determinants and PM2.5 concentrations. Based on a nationwide and multi-year analysis across 288 cities selected from all 336 Chinese cities between 2015 and 2022, the results show that GRF achieves at least a 0.04 higher R-2 than baseline models. Our analysis reveals three findings. First, population density is the most influential factor in 52.39% of cities; combined with temperature, road density, and gas supply, these four dominate over 95% of cities. Second, drivers exhibit significant spatially varying and nonlinear associations. For instance, population density correlates positively with PM2.5 in the North China Plain but negatively in sparsely populated areas; and the association of temperature follows an inverted U-shaped pattern. Third, these spatial and nonlinear associations undergo temporal changes. These findings offer insights for future environmental management strategies to mitigate the negative impacts of various drivers.
High-quality knowledge graphs for ore-forming systems and mineral exploration are essential for "knowledgedriven" prospecting, but they remain difficult to construct from geological exploration texts. Geological prose is long, specialised and mechanism-rich, with multi-scale spatiotemporal coupling and ambiguous terminology, which challenges conventional information extraction methods. Deep-learning models such as BERT-BiLSTMCRF require large labelled datasets and often yield redundant entities, misconnected relations and geologically inconsistent results. We propose OntoGRC (Ontology-guided Generate-Reflect-Correct prompting), a zeroshot ore-forming and mineral exploration knowledge extraction framework that couples a domain ontology with a three-stage prompting workflow. An ore-forming systems and mineral exploration knowledge ontology spanning the metallogenic-prospecting chain is first constructed and formally represented in OWL 2 DL, defining concepts, attributes and semantic relations (e.g., metallogenic dynamics, tectono-lithologic coupling) for use as a structured constraint. OntoGRC then uses this ontology as a semantic anchor to decompose extraction into three stages: ontology-constrained generation of candidate triples, reflective completion using the original text and preliminary results, and ontology-constrained correction that verifies entity types, relation categories, coreference and domain-range consistency while preserving original textual expressions and relation strength. Experiments on two Chinese geological exploration text datasets-including a manually annotated benchmark derived from a Taxkorgan-Yecheng Fe-Pb-Zn assessment report-show that OntoGRC substantially outperforms conventional sequence-labelling and neural relation-extraction baselines. On the benchmark, it achieves F1-scores of 0.886 for entity recognition and 0.880 for relation extraction. Ablation and iteration analyses demonstrate that ontology guidance is essential, and that most performance gains are obtained within three to four generate-reflect-correct cycles. Case studies further illustrate fine-grained entity typing, cross-sentence subject completion and multi-dimensional relation structuring. The results indicate that ontology-guided, LLM-executed and rule-corrected framework can extract structured, traceable and ontology-consistent knowledge from geological texts, supporting mineral exploration knowledge graph construction and subsequent knowledgedriven prospectivity analysis.
Textual spatial correlation quantifies the degree of association between knowledge described in text and specific geographic spaces. For example, the statement ‘the nutria is an invasive species’ is valid only in certain regions. In geographic information science, measuring textual spatial correlation is not only a fundamental prerequisite for spatiotemporal computing, but also a cornerstone for high-precision geo-artificial intelligence applications. To address the limitations of annotation dependency and low computational efficiency in the existing methods, this study proposed a Spatial Correlation Index and corresponding calculation method. This method fully leverages implicit spatial knowledge from a large-scale knowledge base and through a semantic-matching mechanism, enabling the efficient calculation of textual spatial correlation. The effectiveness of the method was evaluated based on the spatial correlation calculations of textual data and entities in a knowledge graph. Results demonstrate that the proposed method achieves performance comparable to few-shot prompted GPT-4.1 and DeepSeek-V3, while outperforming their zero-shot counterparts, with average F1-score improvements of 18.65% and 27.27%, respectively. Meanwhile, the proposed method reduces the average calculation time by 77.37% and 77.74%, respectively. This study achieved a breakthrough in efficiently measuring textual spatial correlations, thereby providing an essential technical foundation for spatial-related intelligent computing.
Vector-data similarity (VD_SM) quantifies the similarity between objects’ spatial and attribute characteristics and underpins retrieval, recommendation, and infringement detection. Fundamental differences in the representation of spatial and attribute features challenge coupled VD_SM measurement, and existing methods rarely integrate both aspects. This study proposes a coupled VD_SM computed through a dual-embedding network optimized with triplet loss. Triplet-ResNet101 learns spatial embeddings from vector images, and Triplet-AttTabNet learns attribute embeddings from attribute tables. Triplet loss projects both embeddings into a shared metric space in which similar features are positioned closer than dissimilar ones. The spatial and attribute embeddings are concatenated, and VD_SM is computed as the Euclidean distance between the concatenated vectors. Experimental results show that (1) Triplet-ResNet101 and Triplet-AttTabNet achieve triplet ranking accuracies of 99.09% and 98.32%, respectively, both surpassing baseline methods; and (2) VD_SM attains an average retrieval accuracy of 92.74%, while a questionnaire yields 90% satisfaction. These results demonstrated that the triplet-loss-based dual-embedding framework provides an effective approach for jointly measuring spatial and attribute similarity in vector data.
Geographic computation is an important process in geographic information systems to detect, predict, and simulate geographic entities, events, and phenomena, which is performed through a series of geographic models over geographic data. However, selecting and sequencing appropriate models is challenging for users with limited knowledge. To automate the process of linking models into workflows, a knowledge graph-based approach is proposed. In this approach, the first part is to construct a knowledge graph that integrates knowledge from geographic models and domain experts. Then, an algorithm is designed to assist the constructed knowledge graph in automating model linking. This paper takes the geomorphological classification of the Hengduan Mountains in China as a case study, which geomorphological classification maps are generated by performing querying and computing through the geomorphological classification knowledge graph. Experimental results demonstrate that the proposed knowledge graph-based approach links the models into workflows automatically and generates reliable classification results.
The Hindu Kush Himalaya (HKH), known as the "Water Tower of Asia", faces mounting challenges from climate change, accelerated glacial retreat, intensified land-use changes, and transboundary water management complexities, leading to significant hydrological transformations that threaten regional water sustainability. To address the urgent need for precise water monitoring, we produced the comprehensive 10-meter resolution surface water dataset (HKH-SWD10m) for the HKH region spanning from 2016 to 2022. The dataset is produced by developing a Vision Transformer-based deep learning network optimized for rapid, automated water extraction. The resulting network achieved exceptional performance metrics on our test dataset, with an Overall Accuracy (OA) of 0.9981, an Intersection over Union (IoU) of 0.9734, and a Kappa of 0.9855. Extensive validation using 15,000 stratified random sampling points demonstrated high accuracy with an OA of 0.9787, a Producer's Accuracy (PA) of 0.9638, a User's Accuracy (UA) of 0.9856, and a Kappa of 0.9476. Comparative analysis with the 30-meter resolution Global Surface Water (GSW) dataset revealed that HKH-SWD10m is generally consistent with the GSW product but captures more small water bodies while providing superior boundary delineation precision. Based on the HKH-SWD10m dataset, we analyzed changes of surface water in the HKH over the years 2016-2022. Our interannual analysis (2016-2022) not only corroborates previous hydrological findings but reveals novel sub-regional divergence in surface water trends when analyzed through national and basin-level frameworks, suggesting localized climate impacts. This dataset advances hydrological monitoring capabilities by offering unprecedented spatial-temporal resolution for the HKH, serving as a critical resource for water security assessments, ecosystem management, and climate adaptation strategies. The HKH-SWD10m dataset is publicly available through Zenodo (https://doi.org/10.5281/zenodo.15067176) and National Earth System Science Data Center of China (https://doi.org/10.12041/geodata.551748804486886.ver1.db).
Geographic knowledge graphs (KGs) mainly describe static facts and have difficulty representing changes, greatly limiting their application in geographic information retrieval and geographic spatio-temporal processes. By analyzing the spatio-temporal features and evolution of geographic elements, this paper measures the degree of correlation between spatio-temporal information and tuples and further accurately characterizes the semantic knowledge of tuples to support accurate computation and inference of KGs. This paper proposes a novel knowledge-guided quantitative measure framework for spatio-temporal correlation by considering rules and dependency syntax from natural language texts. Firstly, the natural language processing (NLP) stage preprocess the texts and extracts the candidate tuples by dependency syntactic analysis and rule matching. Secondly, we model the spatio-temporal correlation measures by considering semantic (entity types and tuple predicate) and syntactic features (dependency distance and dependency path). Finally, we establish a specific threshold value with the extracted candidates and performing multiple levels of categorization to form the final spatio-temporal correlation strength (strong, moderate, and weak). The experimental results with a large dataset indicate that the proposed method achieves an F-score of over 0.73, which is better than those of the existing methods. The proposed spatio-temporal correlation framework has more advantages in representing geographic evolutionary knowledge, revealing the evolution mechanism of geographic elements and the evolutionary reasons.
Geographic knowledge graphs (GKG), central to GeoAI, represent the culmination of knowledge engineering in the era of geographic big data. Knowledge Graph Embedding (KGE) transforms entities and relationships within a knowledge graph into a low-dimensional vector space, effectively capturing their semantic and structural properties. Geographic object knowledge encompasses both intrinsic features and spatiotemporal characteristics, with temporal features indicating the existence or state changes of objects. However, neglecting temporal aspects-such as order, continuity, granularity, and periodicity-during vector calculations can distort the embedding space, reducing the effectiveness of time-sensitive geographic queries, link prediction, and recommendations. This study introduced a temporal feature encoder and designed a fusion mechanism that integrated geographic objects and temporal features. Grounded in logical query tasks, this approach aims to enhance the temporal expressiveness by refining temporal embeddings, thereby improving query accuracy for time-sensitive tasks. A comparative analysis was conducted to evaluate the effects of different baseline models, temporal encoders, and temporal feature weights on the performance of geographic queries.
Volunteered Geographic Information (VGI) images are a vital source of visual information and image features in the GIS field. The advancement of artificial intelligence, particularly large language models and generative AI, now enables the generation of seemingly lifelike images from textual prompts (Artificial Intelligence Generated Content - AIGC). This raises a pertinent question: can AIGC image features serve as a viable alternative to VGI features in downstream GIS tasks, especially where VGI is scarce or difficult to obtain? This paper conducts an exploratory study comparing VGI and AIGC image features as inputs for a geographic recommendation model, specifically investigating the impact of image elements, colors, and spatial structures. The results indicate that current AIGC images, generated from general textual descriptions, cannot fully substitute for VGI images. This is primarily due to AIGC's challenges in accurately replicating the specific elements, colors, and spatial relationships inherent in real-world VGI. However, the study suggests AIGC images hold significant potential. When provided with more specific information about elements and particularly their colors, AIGC's performance approached, and in some color-focused tests, even slightly surpassed that of VGI images. This implies AIGC's main deficiency is its current understanding of real-world object characteristics and their visual representation, notably the basic knowledge of elements and their associated colors. We propose these shortcomings could be addressed by integrating geographic knowledge bases in future AIGC development. These findings aim to guide AIGC's application in GIS by identifying current limitations and areas for focused improvement.
Geographic knowledge graph (GKG) embedding (referred to as knowledge representation learning) enables the mapping of geographic entities and geographic relationships into a continuous vector space, thereby better capturing the semantic and structural information among entities in geographic space. The representation learning of GKGs requires generating corresponding negative samples based on positive samples. Negative sampling is an essential component of GKG embedding models. However, traditional negative sample generation algorithms suffer from high error rates and poor adaptability to GKGs, leading to potential replacements of entities that may not be logically reasonable in terms of geographic relationships, thus reducing the model's performance and generalization ability. To address the aforementioned challenge, this paper proposes a novel negative sampling approach for GKGs embedding, predicated on the amalgamation of entity semantic congruence and clustering techniques. Primarily, the proposed methodology employs the FastText model to encode entities within the KG into vector representations, thereby aiming to procure precise semantic embeddings of said entities. Subsequently, for improved delineation of semantically akin entities, the K-means algorithm is invoked to partition the entities into distinct clusters. Within the purview of negative sampling, entities within the same cluster as the substituted entity are randomly selected as negative instances, thus augmenting the fidelity of negative sample generation. Finally, experiments and analyses were conducted on the open-source datasets WN18, FB15k, and the constructed geographic dataset GKG12. The experimental results show that the proposed method can effectively improve the accuracy of GKG representation learning compared to the existing representation models.
Geographic knowledge graph (GeoKG) alignment is important for the integration and knowledge discovery of multisource geographic information and the generation of large-scale and high-quality knowledge graphs (KGs). However, the existing models/technologies face many challenges when dealing with large-scale multisource complex GeoKG alignment tasks, including the inconsistency of attribute and relationship values caused by domain differences, the inability to perceive relationships and entities, and missing geographic domain training data. To address these issues, we propose a GeoKG alignment model based on depth relationships and neighborhood awareness (named DRNA-GCNE). The DRNA-GCNE model adopts a graph neural network as the infrastructure and uses the graph attention technique to evaluate and weight the entity's relationship attributes dynamically, thus enhancing the ability to perceive structural and semantic information in the GeoKG; concurrently, the relationship information and the multihop neighbor characteristics of the entity are effectively integrated, and the representation of the entity is further enriched. Finally, the training technique of normalized loss mining for multiple negative samples is shown. This approach increases the model's capacity for generalization. The DRNA-GCNE model, as evaluated on two public datasets and our GeoEA2024 Chinese dataset, significantly outperforms current GeoKG entity alignment methods across key metrics.
High-quality Geographic Knowledge Graphs (GeoKGs) are highly anticipated for their potential to provide reliable semantic support in geographical knowledge reasoning, training Geographic Large Language Models (Geo-LLMs), enabling geographical recommendation, and facilitating various geospatial knowledge-driven tasks. However, there is a lack of a standardized quality assessment methodology and clearly defined evaluative indicators in the field of GeoKGs research. This research uses the Preferred Reporting Items for Systematic Reviews and Meta-Analyses (PRISMA) methodology to conduct a systematic review of literature and standards in the field of GeoKG in an effort to fill the gap. First, using the lifecycle theory as a guide, we outline and propose five groups including twenty assessment criteria and their accompanying calculation techniques for evaluating GeoKG quality. Then, expanding on this foundation, we present a streamlined evaluation scheme for GeoKGs that relies on just seven key measures, discussing their applicability, utility, and weight scheme in greater detail. After applying the GeoKG quality framework, we stated three key tasks emerge as priorities: the creation of specialized assessment tools, the formation of worldwide standards, and the building of large-scale, high-quality GeoKGs. We believe this thorough and systematic GeoKG quality assessment technique will help construct high-quality GeoKGs and promote GeoKGs as an engine for geo-intelligence applications including Geospatial Artificial Intelligence (GeoAI) systems, Sustainable Development Goals (SDGs) analyzers, and Virtual Geographic Environments (VGEs) models.
In recent years, geographic knowledge graphs (GeoKGs) have shown great promise in representing spatio-temporal and event-driven knowledge. However, existing knowledge graph embedding approaches mainly focus on structural patterns and often overlook the dynamic evolution of entities in both time and space, which limits their effectiveness in downstream reasoning tasks. To address this, we propose a spatio-temporal evolutionary knowledge embedding approach (ST-EKA) that enhances entity representations by modeling their evolution through type-aware encoding, temporal and spatial decay mechanisms, and context aggregation. ST-EKA integrates four core components, including an entity encoder constrained by relational type consistency, a temporal encoder capable of handling both time points and intervals through unified sampling and feedforward encoding, a multi-scale spatial encoder that combines geometric coordinates with semantic attributes, and an evolutionary knowledge encoder that employs attention-based spatio-temporal weighting to capture contextual dynamics. We evaluate ST-EKA on three representative GeoKG datasets—GDELT, ICEWS, and HAD. The results demonstrate that ST-EKA achieves an average improvement of 6.5774% in AUC and 5.0992% in APR on representation learning tasks. In question answering tasks, it yields a maximum average increase of 1.7907% in AUC and 0.5843% in APR. Notably, it exhibits superior performance in chain queries and complex spatio-temporal reasoning, validating its strong robustness, good interpretability, and practical application value.
Efficient maintenance of the Earth system science knowledge graph (ESSKG) is essential for structuring and evolving scientific knowledge in the context of AI-driven research. However, the rapid emergence of new terminology and reliance on manual curation constrain the timeliness and scalability of updates. A multidimensional feature deep alignment mechanism (MFDA) and an end-to-end update workflow (EKUM) are introduced to address this challenge. MFDA integrates structural, relational, and contextual features through two-stage fusion with cross feature gating, dynamic temperature contrastive alignment, and a sparsified high order neighborhood module. EKUM operationalizes MFDA in a pipeline for corpus preparation, term recognition, alignment, and graph update. Experiments show that MFDA surpasses character based, representation learning, and large language model (LLM) enhanced baselines, achieving precision 0.943 and right to left mean reciprocal rank (MRR) 0.677, exceeding the best-performing baseline by over 11%. Deployed within EKUM, 2889 new terms from 1478 peer-reviewed articles (2012-2024) are integrated, expanding the ESSKG to 6352 entities and 10 747 relations with refined hierarchies and cross layer links. Longitudinal evaluation from 2013 to 2022 indicates stable precision and slower error growth than baselines, mitigating cascading errors. Layer and regional growth patterns reflect differences in data coverage, observing infrastructure, terminology maturity, and cataloging practices, underscoring a scalable, timely, and interpretable pathway for ESSKG evolution.