Spoken language offers a natural, hands-free interface for specifying an arbitrary target in dense remote-sensing imagery, yet existing referring remote-sensing image segmentation benchmarks accept only written expressions. To bridge this gap, we introduce , a spoken-query benchmark derived from RISBench that adds accent- and voice-diverse speech while preserving the original image, mask, and data splits. Its hard evaluation sets combine rotor, wind, and mixed interference with three signal-to-noise levels. We also propose , an efficient bilateral network that combines a boundary-preserving visual path with token-preserving speech encoding, kernel linear cross-modal attention, and a resolution refinement head. The design conditions visual features at two scales without materializing a dense speech–visual affinity matrix, then restores fine boundaries using high-resolution visual features. On the clean test split, with Swin-Base achieves 62.09% mean intersection over union (mIoU) and 68.22% overall intersection over union (oIoU), outperforming the strongest audio-adapted remote-sensing baseline by 5.38 and 2.08 percentage points, respectively. It retains the best hard-set mIoU at 54.09%. To the best of our knowledge, this is the first benchmark and model study of full-sentence spoken-query referring segmentation for remote-sensing imagery. The code will be made publicly available.
The foundation model has recently attracted significant attention due to its exceptional generalisability and outstanding adaptability. However, when it comes to data-driven wind farm wake modelling, due to the high cost of data generation and the complexity of flow characteristics, dimension reduction technology is still the mainstream pre-processing procedure to alleviate the significant challenge inherent in the task itself. Existing methods are still far from achieving a foundation model with high adaptability and excellent scalability. To fill the research, we propose the FlowFormer framework to serve as a foundation model. Specifically, we design a Transformer-based framework, FlowFormer, for flow field prediction by directly taking the LES-simulated data without introducing any dimensionality reduction operations. Moreover, a semi-supervised training strategy is designed to address the problem of over-fitting caused by dimension complexity. The overall mean absolute error of the developed FlowFormer is 6.660% compared to the freestream wind speed for multi-step iterative prediction. The results for a utility-scale farm consisting of 81 turbines demonstrate the high adaptability and excellent scalability of the proposed FlowFormer. Significantly, a qualitative experiment demonstrates that FlowFormer can handle changing-yaw conditions using only fixed-yaw training data, highlighting its excellent flexibility. The demo is available at https://github.com/warwick-icse/FlowFormer.
The rapid advancement of wireless communication technologies has improved the connectivity and operational capabilities of robotic vehicles, enabling their use in applications such as urban monitoring, infrastructure inspection, and disaster response. However, these systems still face key challenges, including limited computational resources, unstable connectivity, and inefficient energy use. Existing coverage exploration methods often neglect real-world wireless communication constraints, relying on idealized conditions or lacking integration of physical and communication factors, which limits their practical effectiveness. This paper proposes a framework that leverages wireless channel modeling for viewpoint and trajectory planning. First, the framework employs ray tracing-based propagation modeling and wireless channel mapping to achieve precise wireless link-level modeling. Next, it integrates a multi-objective viewpoint optimization strategy to balance spatial coverage and communication reliability. Finally, it incorporates a communication-driven trajectory planning method to adapt to realistic wireless conditions. The experimental results demonstrate the effectiveness of the proposed method in achieving reliable coverage exploration and robust wireless communication. This work enhances the resilience and efficiency of robotic operations in urban environments by enabling robust path planning, navigation, and target tracking through wireless communication, contributing to broader sustainable development goals.
Public map service platforms (PMSPs) aggregate spatial data, offering geographic information services across various domains such as health, environment, ocean, and agriculture. Spatial interactions between users and PMSPs give rise to virtual trajectories. Accurate multi-domain virtual trajectory classification is crucial for establishing user profiles and enabling personalized recommendations. However, virtual trajectories derived from access logs exhibit temporal noise, characterized by localized layer sequence fluctuations due to misordered trajectory points, which degrades data quality. We propose the local dynamic sorting algorithm to ensure layer sequence smoothness through local trajectory point reordering. Furthermore, current research often neglects temporal features, specifically zooming and panning operational sequences during transitions, focusing primarily on spatial and semantic features of browsing targets, thereby limiting classification accuracy. We present a simplified representation of virtual trajectories as sequences of , where browsing targets facilitate the extraction of spatial and semantic features, while operations capture temporal features. We develop an operation representation model, a color-template method, and a POI spatial co-occurrence model to extract these features, subsequently integrated into a temporal classification model. Evaluation using real-world data from Tianditu demonstrates a 15.02% improvement in virtual trajectory classification accuracy compared to the benchmark. This study contributes to the precise delineation of PMSP users, fostering the development of personalized, intelligent geographic information services.
The increasing demand for spatiotemporal data and modeling tasks in geosciences has made geospatial code generation technology a critical factor in enhancing productivity. Although large language models (LLMs) have demonstrated potential in code generation tasks, they often encounter issues such as refusal to code or hallucination in geospatial code generation due to a lack of domain-specific knowledge and code corpora. To address these challenges, this paper presents and open-sources the GeoCode-PT and GeoCode-SFT corpora, along with the GeoCode-Eval evaluation dataset. Additionally, by leveraging QLoRA and LoRA for pretraining and fine-tuning, we introduce GeoCode-GPT-7B, the first LLM focused on geospatial code generation, fine-tuned from Code Llama-7B. Furthermore, we establish a comprehensive geospatial code evaluation framework, incorporating option matching, expert validation, and prompt engineering scoring for LLMs, and systematically evaluate GeoCode-GPT-7B using the GeoCode-Eval dataset. Experimental results reveal that GeoCode-GPT significantly outperforms existing models across multiple tasks. For multiple-choice tasks, its accuracy improves by 9.1% to 32.1%. In code summarization, it achieves superior scores in completeness, accuracy, and readability, with gains ranging from 1.7 to 25.4 points. For code generation, its performance in accuracy, readability, and executability surpasses benchmarks by 1.2 to 25.1 points. Grounded in the fine-tuning paradigm, this study introduces and validates an approach to enhance LLMs in geospatial code generation and associated tasks. These findings extend the application boundaries of such models in geospatial domains and offer a robust foundation for exploring their latent potential.
Deep convolutional neural network has strong feature extraction and fitting capabilities and perform well in hyperspectral image classification tasks. However, due to its huge parameters, complex structure and high energy consumption, it is difficult to be used in mobile edge computing. Spiking neural network (SNN) has the characteristics of event-driven and low energy consumption and has developed rapidly in image classification. But it usually requires more time steps to achieve optimal accuracy. This paper designs a faster residual multi-branch SNN (FRM-SNN) based on leaky integrate-and-fire neurons for HSI classification. The network uses the residual multi-branch module (RMM) as the basic unit for feature extraction. The RMM is composed of spiking mixed convolution and spiking point convolution, which can effectively extract spatial spectral features. Secondly, to address the problem of non-differentiability of Dirac function spiking propagation, a simple and efficient arcsine approximate derivative was designed for gradient proxy, and the classification performance, testing time, and training time of various approximate derivative algorithms were analyzed and evaluated under the same network architecture. Experimental results on six public HSI data sets show that compared with advanced SNN-based HSI classification algorithms, the time step, training time and testing time required for FRM-SNN to achieve optimal accuracy are shortened by approximately 84%, 63% and 70%. This study has important practical significance for promoting the engineering application of HSI classification algorithms in unmanned autonomous devices such as spaceborne and airborne systems.
Urban functional structures and daily rhythms significantly impact population mobility. Detecting and quantifying thematic activity changes in the short-term aggregated inflow and outflow of urban population mobility (referred to as black holes and volcanoes, respectively) contribute to the economy and public services. Current research is focused on the changes in intensity of single-region population activity, but it overlooks the daily rhythms of the aggregated flows and latent thematic activity changes in urban populations. Here, we propose an Aggregated Inflow-Outflow Thematic Detection (AIOTD) method. It can detect thematic activity changes in population inflows and outflows from a spatiotemporal aggregation perspective by leveraging traffic flow theory and semantic models. Considering the stationary of flow sequences within the same time periods and the spatial continuity of flows, we designed a spatial aggregation method based on the relative ratios of population flows in spatiotemporal units. This method enables a quantitative depiction of the spatiotemporal evolution of black holes and volcanoes. Furthermore, due to the spatial proximity of land features and category imbalances, we utilized the SMOTETomek-Place2vec model to construct a spatial context information dataset at the grid level, enhancing the accuracy of capturing thematic activities. Results demonstrate that our method outperforms existing approaches in capturing the number of spatiotemporal units for black hole and volcano clusters at both the 500 m grid and 1000 m grid scales, in terms of both semantics and spatiotemporal dimensions. It reveals the spatiotemporal complementarity and thematic cross-symmetry of urban population mobility between black holes and volcanoes. By applying this method to daytime and nighttime economies and public services, we quantified changes in the fine-grained functional service vitality and the distribution of functional facilities in commuting zones. These findings offer guidance for urban services and insights into the behavioral preferences of residents.
Flood disasters rank as the most prevalent natural calamities of the twenty-first century, incurring extensive human and economic losses globally. As a crucial source for disaster monitoring, social media data exhibits high variability and ambiguity, with current research lacking targeted multidimensional semantic analysis, resulting in coarse granularity and limited accuracy. To address this problem, this study proposes a framework and method synthesizing multiple semantic features to extract fine-grained disaster information. Static embeddings representing stable semantics and dynamic embeddings representing changing semantics are fused to extract toponyms, with the depth-first search used to generate addresses through the toponym tree. Guiding prompts incorporating domain-specific knowledge of disasters are designed for the large language model, with an iterative feedback process refining location-based disaster information. Finally, the reliability of social media-sourced information is assessed by comparing extracted flooded locations with actual monitoring data. The case study on the Zhengzhou ‘7·20’ flood event demonstrates the effectiveness of our semantic fusion approach, achieving an F1 score of 0.9384 for address extraction, with accuracies of 0.8485 and 0.8788 for waterlogging depth and trapped individuals, respectively. This research offers a practical framework for nuanced perception and timely rescue operations in urban disaster management.
As a novel and challenging task, referring segmentation combines computer vision and natural language processing to localise and segment objects based on textual descriptions. While Referring Image Segmentation (RIS) has been extensively studied in natural images, little attention has been given to aerial imagery, particularly from Unmanned Aerial Vehicles (UAVs). The unique challenges of UAV imagery, including complex spatial scales, occlusions, and varying object orientations, render existing RIS approaches ineffective. A key limitation has been the lack of UAV-specific datasets, as manually annotating pixel-level masks and generating textual descriptions is labour-intensive and time-consuming. To address this gap, we design an automatic labelling pipeline that leverages pre-existing UAV segmentation datasets and the Multimodal Large Language Models (MLLM) for generating textual descriptions. Furthermore, we propose Aerial Referring Transformer (AeroReformer), a novel framework for UAV Referring Image Segmentation (UAV-RIS), featuring a Vision-Language Cross-Attention Module (VLCAM) for effective cross-modal understanding and a Rotation-Aware Multi-Scale Fusion (RAMSF) decoder to enhance segmentation accuracy in aerial scenes. Extensive experiments on two newly developed datasets demonstrate the superiority of AeroReformer over existing methods, establishing a new benchmark for UAV-RIS. The datasets and code are publicly available at https://github.com/lironui/AeroReformer.
While leveraging large language models (LLMs) for intelligent geospatial modeling has garnered significant attention, the limited domain-specific knowledge of LLMs often leads to inefficient or unreliable geo-analysis model generation. Crowdsourced geoprocessing scripts encapsulate extensive expert knowledge for different geospatial modeling tasks, where code snippets are strategically combined into functional steps to build application-specific modeling processes. However, extracting these modeling processes from heterogeneous geoprocessing scripts and integrating them for reuse remains challenging due to the complexity of code interdependencies, the heterogeneity of scripting approaches, and the need for domain-specific customization. To address this, we propose S-GMKG, a knowledge graph that systematically extracts and integrates modeling processes from scripts as structured semantic units. Two strategies are introduced: a skeleton-based extraction method and a knowledge-enhanced chain of thought (CoT) approach, which facilitate automated modeling process extraction for S-GMKG via prompt engineering. Furthermore, a self-canonicalization and knowledge augmentation process is proposed to refine the S-GMKG. Consequently, S-GMKG serves as a robust external knowledge source to provide interpretable, graph-based modeling solutions and synergizes with LLMs for geospatial tasks. We implemented the S-GMKG using 4820 geoprocessing scripts and evaluated it across various LLMs. Results indicate that most scripts in the S-GMKG can be represented as modeling processes with 3-7 functional steps, with the proposed strategies achieving 3.2%-14.5% higher recall rates in relationship identification for these functional steps. Case studies in two distinct scenarios demonstrate the practicality of S-GMKG, particularly in collaborating with LLMs to generate code for geospatial modeling.
Fine-scale population estimation (FPE) is crucial for urban management. After training, the bottom-up FPE models can be applied independently of census data. However, given the lack of real fine-scale population data, the existing bottom-up methods typically apply models trained on coarse-grained census data to FPE directly, causing estimation bias induced by scale effects. Traffic analysis zones (TAZs) balance geo-semantics and fine granularity, but their potential as population analysis units has not been fully exploited. Hence, we developed a weakly-supervised TAZ-scale bottom-up population estimation method (WSTP). Specifically, to mitigate scale effects, the weakly-supervised training procedure involves fine-scale feature input, model prediction, spatial aggregation, and coarse-scale supervision is designed to ensure that WSTP consistently focuses on FPE throughout the training and prediction phases. To enable weakly-supervised training using census data, we treated TAZs as graph nodes and designed a Spatial Aggregation Layer to aggregate TAZ-scale population predictions into communities. Given the diverse distribution patterns across age groups, we decomposed the FPE task by age groups. The experiments showed that WSTP significantly outperformed the baselines with R2 values of 0.821 and 0.785 at the community and TAZ scales, respectively, indicating that WSTP can mitigate scale effects and produce high-resolution, accurate population data.
Differences in the spatial scale and component elements of urban scenes affect the analysis of spatiotemporal changes in population distribution (PD) patterns. Fixed scales and geometric forms constrained previous analyses of PD and spatiotemporal changes, which neglected the realistic characteristics of urban scenarios. These limitations hinder the ability to capture the diversity and spatiotemporal representation of PD patterns across multiple spatial scales. This study developed a multi-spatial scale population analysis unit (PAU) construction method considering the scenario heterogeneity and long/short-term patterns of spatiotemporal population changes (PC). Initially, we decompose the multi-scale changes in population time series patterns and their relationships with scene features along the temporal dimension to analyze the primary and secondary factors. Starting from the temporal characteristics of the PD and PC, we established a validation and correction method for the factors. Finally, combined with a multi-feature clustering method, a multi-scale PAU construction method driven by scene feature factors is proposed. Experiments were conducted at different spatial scales in both routine and emergent scenarios. The results indicated that this method can help to obtain more homogeneous analysis regions and enhance the stability and phase pattern representations of changes in PDs, thereby improving the interpretability of analytical results.
When performing surveying and mapping missions using drones, motion blur is an unavoidable issue induced by several factors such as vibration, turbulence and wind during operation. Such blurring can significantly degrade the image quality, adversely affecting the accuracy and reliability of downstream applications. In this paper, to effectively eliminate the motion blur in the UAV-captured images, we propose the NAFormer based on the well-established Nonlinear Activation Free Network (NAFNet), which introduces Transformer-based blocks to further enhance its motion-deblurring ability to UAV images. The experiments based on the UAVid dataset demonstrate the effectiveness of the proposed framework.
Super-resolution, which aims to reconstruct high-resolution (HR) images from low-resolution (LR) images, has drawn considerable attention and has been intensively studied in computer vision and remote sensing communities. Super-resolution technology is especially beneficial for unmanned aerial vehicles (UAVs), as the number and resolution of images captured by UAVs are highly limited by physical constraints such as flight altitude and load capacity. In the wake of the successful application of deep learning methods in the super-resolution task, in recent years, a series of super-resolution algorithms have been developed. In this article, for the super-resolution of UAV images, a novel network based on the state-of-the-art Swin Transformer is proposed with better efficiency and competitive accuracy. Meanwhile, as one of the essential applications of the UAV is land cover and land use monitoring, simple image quality assessments such as the peak-signal-to-noise ratio (PSNR) and the structural similarity index measure (SSIM) are not enough to comprehensively measure the performance of an algorithm. Therefore, we further investigate the effectiveness of super-resolution methods using the accuracy of semantic segmentation. The code is available at https://github.com/lironui/GeoSR.
Spatial interaction centrality reflects the relative importance of population mobility within a location in urban population mobility. Population mobility networks visually represent urban population mobility, with mobility features and network topology contributing to the quantification of spatial interaction centrality of locations (i.e., geographical nodes). However, existing centrality measures rarely consider mobility features and network topology simultaneously. Centrality quantification also ignores the differences in distance effects between long- and short-distance trips. These factors have led to the inaccurate quantification of centrality. We propose an algorithm called k-dis-weight-shell that quantifies the spatial interaction centrality of geographical nodes at different spatiotemporal scales. Considering the different effects of distance on long- and short-distance trips, we use a spatial continuous wavelet transformation to estimate the radiation radius of geographical nodes. Then, by combining network topology with mobility features (mobility distance and flow), the algorithm transforms them into a ranked order of spatial interaction centrality. Tested in Wuhan and Chengdu, our algorithm outperforms six existing benchmarks. For cases in urban planning and epidemic management, results show that k-dis-weight-shell effectively distinguishes similarities and differences between the distribution of population mobility's spatial interaction centrality and the urban center hierarchy at a coarse spatiotemporal scale. Additionally, it reveals a double wave phenomenon of spatiotemporal correlation between population mobility and COVID-19 transmission before and after lockdown at a fine spatiotemporal scale.
Public Map Service Platforms (PMSPs) provide embedded map services in domains such as forests and rivers. Users from different domains (Domain Users) prefer specific spatial features, and extracting the Browsing Interests of Domain Users (BIDUs) can help elucidate users’ access intentions and provide suitable recommendations. Previous research has found that access frequency of spatial features is an indicator of users’ browsing interests; however, high-frequency spatial features are sparsely distributed, resulting in inaccurate extraction of browsing interests. Our objective is to model the spatial co-occurrence of spatial features and employ BIDUs extraction to address this limitation. First, to extract spatial features in tiles, we proposed a k-nearest neighbor method for Point-of-Interest (POI) extraction and a template-based method for Land Uses/Land Covers extraction. Then, we developed the word2vec model to construct a POI semantic space to quantify spatial co-occurrence and employed multi-domain user classification to verify its effectiveness. Finally, a combined word2vec and singular value decomposition model is proposed to perform topic extraction as a representation of BIDUs. Compared with the baseline models, the proposed model integrates spatial co-occurrence from massive POIs to achieve high-accuracy BIDU extraction. Our findings can help construct domain user profiles and support the development of intelligent PMSPs.
Information on the population distribution at the building scale can help governments make supplemental decisions to address complex urban management issues. However, the discontinuity and strong spatial heterogeneity of research units at the building scale make it challenging to fuse multi-source geographic data, which causes significant errors in population estimation. To address this problem, this study proposes a method for population estimation at the building scale based on Dual-Environment Feature Fusion (DEFF). The dual environments of buildings were constructed by splitting the physical boundaries and extracting features suitable for the dual-environment scale from multi-source geographic data to describe the complex environmental features of buildings. Meanwhile, Data Quality Weighting based Technique for Order of Preference by Similarity to Ideal Solution (DQW-TOPSIS) method was proposed to assign appropriate weights to the features of the external environment for better feature fusion. Finally, a regression model was established using dual-environment features for building-scale population estimation. The experimental areas chosen for this study were Jianghan and Wuchang Districts, both located in Wuhan City, China. The estimated results of the DEFF were compared with those of the ablation experiments, as well as three publicly accessible population datasets, specifically LandScan, WorldPop, and GHS-POP, at the community scale. The evaluation results showed that DEFF had an R2 of approximately 0.8, Mean Absolute Error (MAE) of approximately 1200, Root Mean Square Error (RMSE) of approximately 1700, and both Mean Absolute Percentage Error (MAPE) and Symmetric Mean Absolute Percentage Error (SMAPE) of approximately 26%, indicating an improved performance and verifying the validity of the proposed method for fine-scale population estimation.
Public Map Service Platforms (PMSPs) aggregate and disseminate the earth observation data. Leveraging spatiotemporal preference patterns derived from browsing targets within complex virtual trajectories on PMSPs aids in constructing user-profiles and comprehending their intentions. However, complex virtual trajectories, characterized by numerous trajectory points and overlapping pyramidal spatial structures, introduce inefficiencies and inaccuracies during browsing target extraction. To mitigate this, we propose an Optimized Spatial Structure Segmentation (OSSS) method that divides complex virtual trajectories into sub-trajectories with simplified spatial structures, enhancing the efficiency of browsing target extraction. Spatiotemporal reconstruction of these sub-trajectories establishes sequences of browsing targets, revealing patterns of interest transitions. Moreover, recognizing the spatial uncertainty inherent in complex virtual trajectories, which results in imprecise matching between browsing targets and spatial features, we introduce a spatial co-occurrence semantic modeling approach. This involves constructing a POI semantic space and introducing a semantic similarity matching method to reduce spatial uncertainty and refine the accuracy of mining spatiotemporal preference patterns. We evaluate these methods using real-world data from Tianditu. Results demonstrate that the OSSS improves extraction efficiency by 3.45 times and accuracy by 18.84%. Additionally, the semantic similarity approach combined with spatial co-occurrence effectively mitigates spatial uncertainty. This research contributes to advancing the intelligence of PMSPs.
With the escalating global climate changes and rapid urbanization, urban flood disasters have become more frequent and severe, hindering progress toward global sustainability goals. The rapid growth of information technology provides abundant data for disaster monitoring, but transitioning from static to dynamic and structured to heterogeneous data poses challenges for organized flood disaster information analysis. This study conducts an in-depth analysis of flood disaster components and spatiotemporal characteristics, exploring qualitative information expression in diverse contexts and quantitative dynamic analysis. Using dimensions like spatiotemporal, semantic, and hierarchical correlations, we introduce a multi-level, multi-dimensional “circumstance-event” model for flood disaster events. Leveraging proximity and semantic relationships, we establish a dynamic semantic association network among geographical units, supporting disaster circumstance division using semantic correlation graphs. Introducing bidirectional relationship discrimination, we enhance the traditional group evolution discovery method to determine the evolution of circumstances in consecutive time windows. Taking the “7·20” heavy rain in Zhengzhou, China as a case study, we observe significant improvement in circumstance division based on semantic correlation graphs, with an average modularity exceeding 0.745. Our bidirectional community evolution method reveals the evolution sequences of circumstances throughout the flood disaster period in Zhengzhou. This approach effectively organizes the entire flood disaster event lifecycle and dynamically analyzes event evolution, offering vital support for comprehensive disaster management.