Accurate characterization of individual tree attributes and spatial distributions is pivotal for scientifically guided forest thinning decisions, yet traditional methods for identifying target trees are subjective and labor-intensive. To address these challenges within the context of typical Chinese fir (Cunninghamia lanceolata) plantations in southern China, we propose a Geographic Object-based Integrated Multi-module Image Analysis (GEOIMIA) framework, incorporating three interconnected modules. First, an object-based classification module applies heuristic sequential optimization to achieve high object-level accuracy (97% on the validation set) for four categories-Chinese fir, broadleaf trees, other tree species, and others-across entire images, thereby enabling precise global semantic masks and robust species delineation. Second, we developed a dedicated detection network, AMC-Net, tailored to high-resolution UAV remote sensing data under low illumination conditions, enabling precise crown detection for Chinese fir with AP50 = 96.3% and AP50-95 = 58.3%. Model compression experiments further reduced AMC-Net's parameters to 1.1 million (only 45.8% of the original) while maintaining comparable detection performance (APweighted = 62.1%). Third, a decision module supported by Maskable Proximal Policy Optimization (Maskable PPO) dynamically optimizes thinning priorities, effectively balancing canopy closure (CC), stem density (SD), and stand stability (SS) in real-time. Adhering to the Regulations for forest tending (GB/T15781-2015), this module implements selective thinning strategies that remove denser, smaller, or weaker trees and retain those with higher ecological and silvicultural value. Consequently, GEOIMIA offers an economical, efficient, and scalable solution for automated thinning operations in mixed-species forests.
Accurately estimating gravitational direction from crowdsourced driving record data is challenging due to prevalent low frame rates and unreliable inertial measurement unit (IMU) data. These limitations restrict the applicability of conventional sequential processing methods, while existing vision-based approaches often lack robustness in complex urban environments. To address these issues, we propose a Robust Vanishing Point-based method for Gravitational Direction Estimation (RVP-GDE), a novel single-image framework for gravity direction estimation. The core of our approach is a robust vanishing point estimator based on an adapted weighted Density-Based Spatial Clustering of Applications with Noise (DBSCAN) clustering algorithm. This method is particularly suitable for this task as it does not require pre-specification of the number of clusters, and is explicitly designed to identify and exclude noise points. We further enhance the algorithm by incorporating segment length reliability through a weighting scheme. Experimental results on a dedicated dataset demonstrate that the mean absolute error between estimated and ground-truth angles is 0.78 degrees for pitch and 0.72 degrees for roll, outperforming several baseline methods and achieving accuracy comparable to consumer-grade IMUs. This work provides an effective, sensor-independent solution for camera orientation estimation, with significant potential for large-scale crowdsourced data processing applications such as traffic sign localization for HD map updating.
Vision-Language-Action (VLA) models have become a central paradigm for embodied intelligence. However, most existing approaches are built on large-scale Transformers, resulting in substantial inference latency and energy consumption that limit their practical deployment in low-power, real-time scenarios. We propose SpikeVLA, an end-to-end spiking VLA framework for embodied navigation with energy-efficient inference, consisting of three key components. (i) a spiking vision encoder, Spike-V, that replaces dense continuous computation with event-driven spiking representations to reduce the energy cost of visual representation learning, (ii) a multimodal spiking large language model, Spike-L, that reformulates cross-modal reasoning with spiking dynamics and token-level event-driven sparsity to further lower inference overhead, and (iii) a spiking action policy network, Spike-A, that uses Laplacian-kernel population coding and end-to-end reinforcement learning to produce stable, robust continuous control under low-energy constraints. Experiments on multimodal interaction and robotic control tasks show that SpikeVLA significantly reduces energy consumption and computational overhead while maintaining competitive performance, highlighting its potential for low-power, real-time embodied intelligence.
ABSTRACT Advances in understanding the interactions between the built environment (BE) and traffic spatiotemporal congestion patterns (TCPs) have significantly contributed to sustainable transportation development. However, existing BE‐TCP studies remain fragmented in terms of measurement, built environment indicators, spatial scale and methodology selection, which restrict cross‐study comparability and policy generalization. To address these gaps, this review systematically analyses 214 papers collected from Web of Science, Scopus, and Google Scholar published between 1997 and 2025. This review makes three main contributions. First, it synthesizes the BE measurement from the conventional 5Ds framework (density, diversity, design, distance to transit, and destination accessibility) to an expanded 7Ds by incorporating two additional dimensions: demand management and demographics. Second, it identifies four hierarchical geospatial scales of BE‐TCP relationships—disaggregated, neighborhood, aggregated, and regional scales, and highlights the scale‐dependent nature of BE‐TCP effects. Third, it summarizes the methodological evolution from classical linear to spatial nonlinear, from direct to mediation, from singular to synergic, from static to dynamic, and from correlation to causality approaches. The review reveals that BE‐TCP relationships are not uniform but are characterized by strong spatiotemporal heterogeneity, multi‐scale effects, and regional dependence. Based on these advances, three critical challenges and four opportunities are proposed. Overall, this review highlights the need to move BE‐TCP research from fragmented empirical associations toward spatiotemporal, multiscale, and context‐sensitive explanations that can better support transferable and locally adaptive policy‐making.
3D object detection, a pivotal task in autonomous driving systems, confronts the challenge of performance degradation under occlusion and adverse weather conditions. Fusing multi-modal point clouds from LiDAR and 4D radar is an effective way to solve this problem. However, current LiDAR and 4D radar fusion methods fuse the features of heterogeneous point clouds at the bird’s-eye-view (BEV) level and ignore geometric inconsistencies, which makes them sensitive to adverse conditions. To address the issues above, we propose a LiDAR and 4D radar fusion model (FusionBev) for accurate and robust 3D object detection. In this study, we focus on how to make the network fully fuse LiDAR and 4D radar data at the voxel level and ensure geometric consistency. Further, we propose a cross-fusion module (CF) to aggregate the features of LiDAR and 4D radar voxels. After voxel encodings by CF module, we design a redundant down-sampling strategy (RD) to learn the multi-scale features. Finally, a geometry-consistent module (GC) is designed to solve the problem of geometric offset between sensors. We conduct extensive experiments across multiple public datasets to evaluate the effectiveness and robustness of our model. Notably, FusionBev achieves 89.2 % mAP on the VoD dataset and 64.9 % mAP on the K-Radar dataset. Compared to the recent LiDAR-4D radar fusion method (L4DR), we achieve more than twice the inference speed (27.6 FPS> 13.1 FPS) with less than half the GPU memory (2.81 GB < 6.31 GB).
Building function is a description of building usage. The accessibility of its information is essential for urban research, including urban morphology, urban environment, and human activity patterns. Existing building function classification methodologies face two major bottlenecks: (1) poor model interpretability and (2) inadequate multimodal feature fusion. Although large models with strong interpretability and efficient multimodal data fusion capabilities offer promising potential for addressing the bottlenecks, they remain limited in processing multimodal spatial datasets. Their performance in building function classification is therefore also unknown. To the best of our knowledge, there is a lack of multimodal building function classification datasets, which results in the challenge of effectively performing their performance evaluation. Meanwhile, prevailing building function categorization schemes remain coarse, which hinders their ability to support finer-grained urban research in the future. To bridge the gap, we constructed a novel multimodal and fine-grained dataset - BuildingSense - for building function classification, comprising over 34 000 buildings, 60 000 annotated images, 71 654 POIs, and 3400 building description texts in 26 distinct categories. Based on BuildingSense, we evaluated the performance of four state-of-the-art large models from the perspective of classification outcomes and reasoning processes. The results demonstrate that large models can effectively comprehend multimodal spatial data, challenging the conventional concept. Based on that, three directions for future research can be key: (1) build a categorized inference example database, (2) develop cost-effective classification models, and (3) quantify the confidence of model outputs. Our findings not only provide insights into the development of subsequent large model-based classification methods but also contribute to the advancement of multimodal fusion-based classification methods. The dataset and code of this paper can be accessed through 10.6084/m9.figshare.30645776.v2 .
Mapping in urban environments is often disrupted by moving objects, which contaminate LiDAR observations. To address this issue, we propose a complete LiDAR-inertial odometry framework designed to ensure map consistency and cleanness for mobile laser mapping systems. Our approach features three key innovations: (1) We design a visibility-based moving-object removal module that utilizes motion cues and object-level clustering to consistently detect and eliminate points from moving entities. (2) We propose a novel two-pass strategy that couples moving-object removal with LiDAR-inertial odometry. The first pass provides priors for moving-object filtering, while the second pass refines the trajectory using only points from static objects. Crucially, this iterative update builds a clean map from the start while maintaining real-time performance. (3) In addition to real-time mapping, the framework is scalable for professional large-scale mobile mapping. We add a global pose graph optimization module with long-range revisit and short-range co-visibility constraints to further optimize trajectories and improve map consistency. Experiments on large-scale public datasets and challenging private sequences demonstrate that the framework achieves state-of-the-art performance in moving-object removal (achieving the best F1 score in 7 out of 8 sequences) and outperforms state-of-the-art general-purpose and dynamic-aware LIO methods (achieving the lowest trajectory error in all 8 sequences). Furthermore, our experiments confirm that superior trajectories systematically enhance moving-object removal, validating the synergistic effect between map cleanness and consistency. The code will be available at https://github.com/Shidabot/CC-MLS.
Accurate rail wear measurement under dynamic conditions remains challenging, particularly for single-line structured light systems that provide only partial profile observations. The core difficulty lies in distinguishing pose misalignment from wear-induced geometric deformation under noise-contaminated and incomplete measurements, which leads to an ill-posed registration problem. This study formulates dynamic rail wear measurement using a single-line structured light sensor as a partial profile registration problem in which pose misalignment and wear induced deformation must be separated. Instead of using weighted registration as a generic fitting tool, the proposed method introduces a wear aware constraint allocation strategy that reduces the influence of rail head and gauge corner regions, where wear is likely to occur, and assigns greater importance to geometrically stable regions. This design prevents worn geometry from biasing rigid alignment and preserves wear related deviations for subsequent quantification. Field experiments were conducted on an operational metro line at speeds of 20, 40, and 50 km/h to evaluate dynamic profile acquisition and preprocessing performance. Quantitative wear validation was performed in the 50 km/h section, where the proposed method achieved an RMSE of 0.066 mm for total wear against manual reference measurements. Compared with conventional rigid registration approaches, the proposed method improves registration stability under the tested partial-profile conditions and reduces alignment bias caused by wear.
The rapid detection and geolocation of road events are critical for urban traffic management and public safety. With the growing variety of traffic information sources, efficiently processing textual data from news report and social media and delivering GIS/LBS services remains challenging due to delayed information retrieval, imprecise geographic localization, and insufficient integration between textual and spatial information in existing methods. To address these limitations, this study proposes RoadEventGPT, a framework for the automatic extraction and geospatial localization of road events. By integrating large language models (LLMs) with geospatial tools via prompt engineering and few-shot learning, RoadEventGPT enables complex geospatial reasoning and tool invocation, forming an end-to-end pipeline that converts unstructured text into structured, georeferenced road event data. The framework decomposes the task into text preprocessing, event extraction, and geolocation. Experimental results demonstrate that RoadEventGPT significantly outperforms traditional rule-based methods, with high F1-scores for geographic entity recognition and spatial scene classification. It also exhibits strong robustness and adaptability across diverse road-related spatial scenarios, enabling effective handling of different types of spatial contexts in road event analysis.
The complexity of urban environments influences pedestrians’ walkability, which is especially significant for people living in mega cities. While many studies identify influential factors, how these factors shape pedestrian wayfinding through complex and spatially varied mechanisms remains underexplored. This study addresses this gap by using a novel pedestrian navigation dataset as a proxy to quantify the perceived complexity of walking environments. By integrating multi-scale urban features—four at the macro-level and 14 at the micro-level derived from Street View Imagery—we systematically uncover the key correlates of navigation demand and their underlying effects. The results reveal that a combination of factors such as the number of Points of Interest, transportation accessibility, proportion of people in view, and intersection count are positively associated with pedestrians’ navigation behavior. More importantly, we demonstrate that their relationship is profoundly non-linear and exhibits strong spatial heterogeneity. These results are further validated through population normalization, sensitivity tests, and temporal comparisons between weekdays and weekends. Such analyses confirm the robust and independent association between environmental complexity and navigation behavior. By operationalizing these complex interrelationships, our work advances the theoretical framework for urban environmental complexity. The findings provide crucial evidence for moving beyond a "one-size-fits-all" approach, offering targeted, context-aware insights to foster truly human-centered urban planning and design.
Accurate and dynamic Intersection Turning Control (ITC) information is a fundamental component for intelligent traffic management. However, widely used open-source maps, such as OpenStreetMap, often suffer from incomplete attribute data and update lags. While crowdsourced trajectory data offers a promising solution, existing extraction methods typically rely on absolute frequency thresholds, lacking a systematic quantification of uncertainty. To address this, we propose an integrated framework for ITC extraction and confidence evaluation using mobile navigation trajectories. The framework first abstracts intersection geometry and refines trajectory segments to ensure spatial alignment. Building on this, a Balanced Random Forest model is utilized to robustly classify movement modes. Crucially, to quantify the reliability of these extracted rules, we introduce a Bayesian hierarchical confidence evaluation model that infers the posterior probability of movement legitimacy from the observed data. This probabilistic approach explicitly quantifies the reliability of each turning rule, effectively distinguishing between rare legal maneuvers and anomalous violations. Experimental results from a case study in Shanghai demonstrate that the proposed framework achieves an overall extraction accuracy of 93.75%, with an optimal confidence threshold identified between 0.60 and 0.65. The study validates that incorporating confidence evaluation significantly enhances the semantic correctness of map updates, providing a robust solution for maintaining real-time, lane-level urban road networks.
Due to the underground or enclosed nature of subway systems, satellite navigation signals are often difficult to receive. Without reliable signals, mobile measurement systems are prone to significant positioning errors. To address this challenge, this study proposes an improved method that leverages control points, dynamic calibration, and advanced data fusion. A unified coordinate system is established for all sensors, ensuring that measurements from different devices, such as inertial navigation systems and laser scanners, can be integrated seamlessly. System errors are identified and corrected through dynamic calibration to maintain long-term accuracy. Object detection technologies are used to identify subway control points, enabling trajectory correction through coordinate-assisted updates. The calibrated parameters are then applied to fuse inertial navigation data with laser scanning data, effectively reducing the accumulation of errors. This approach generates high-precision three-dimensional point clouds with absolute coordinates, suitable for applications such as subway inspection and tunnel monitoring in signal-denied environments. Results show that the method achieves deviations within 1 cm for horizontal measurements and 5 mm for vertical measurements. This significant improvement enhances the precision and reliability of laser scanning data in underground applications. The study also analyzes the spatial distribution of control points on point cloud quality, providing a practical solution for improving mobile measurement systems where satellite signals are unavailable.
Tunnels are essential components of modern transportation infrastructure, where their structural integrity is crucial for ensuring the safety and efficiency of train operations as well as public security. This study introduces an innovative submillimeter imaging inspection method for tunnels, leveraging line scan cameras to simultaneously address the challenges of achieving high-resolution imaging and optimized data processing. A comprehensive tunnel imaging inspection system, integrating both hardware and software components with multiple synchronized line scan cameras to achieve high-speed, wide-area imaging. To address image distortions, a robust correction framework was developed, incorporating lens shadow correction, white balance adjustment, directional alignment, tilt correction, and orthorectification. Furthermore, an innovative multi-view image stitching method was introduced, in which pixel energy in overlapping regions was defined based on structural and color differences to guide seam selection and blending. Dynamic programming was employed to determine globally optimal stitching lines, effectively mitigating ghosting and double images caused by uneven lighting and parallax while maintaining high processing efficiency. The system is compatible with both monochrome and color line scan cameras, offering versatility to adapt to diverse project requirements. Extensive field tests in metro and high-speed rail tunnels demonstrated that the system can achieve a spatial resolution of 0.2 mm/pixel, which qualifies as submillimeter-level resolution. This high-resolution imaging capability enables the precise detection of small cracks, spalling, and other surface defects, providing an efficient, automated, and accurate alternative to traditional inspection methods. This advancement significantly improves tunnel maintenance operations.
Real-time traffic event information is essential for various applications, including travel service improvement, vehicle map updating, and road management decision optimization. With the rapid advancement of Internet, text published from network platforms has become a crucial data source for urban road traffic events due to its strong real-time performance and wide space-time coverage and low acquisition cost. Due to the complexity of massive, multi-source web text and the diversity of spatial scenes in traffic events, current methods are insufficient for accurately and comprehensively extracting and geographizing traffic events in a multi-dimensional, fine-grained manner, resulting in this information cannot be fully and efficiently utilized. Therefore, in this study, we proposed a “data preparation - event extraction - event geographization” framework focused on traffic events, integrating geospatial information to achieve efficient text extraction and spatial representation. First, the text data is preprocessed, with road-related information extracted and summarized to prepare for subsequent tasks. Next, a step-wise method for automated extraction is introduced. Trigger words and rules of spatial relationship are set to identify spatial elements within the text, then dictionaries of proper and general names are applied to further recognize candidate entities. Finally, we adopt a method for entity disambiguation by introducing spatial constraints such as direction. Based on spatial scenes, entities representing different elements are organized to perform spatial computing, realizing the multi-dimensional geographization of events. A case study in Shanghai demonstrated the effectiveness of the proposed method, showing that it improves the completeness and accuracy of traffic event extraction while enhancing the diversity and accuracy of geographization.
Extrinsic calibration between light detection and ranging (LiDAR) and inertial measurement unit (IMU) is fundamental to autonomous platforms, as the quality of motion estimation and sensor fusion relies heavily on precise intersensor alignment. Traditional methods are typically conducted in controlled laboratory settings, which is impractical for field deployment. On-the-fly calibration, while highly convenient, is often hindered by intense motion and the lack of omnidirectional excitation on ground platforms such as vehicle-mounted systems. To overcome these limitations, in this article, we propose a real-time LiDAR-IMU extrinsic calibration framework. The method first constructs a nonuniform motion model to precisely estimate LiDAR motion and then applies a decoupled calibration strategy to independently solve and optimize rotation and translation parameters. Furthermore, to address degraded observability under planar motion, we introduce an excitation analysis module that compensates for insufficient excitation and enhances calibration robustness. We validate effectiveness and stability on public and self-collected datasets, reporting mean +/- standard deviation (SD) for all six extrinsic parameters (repeatability) and RMSE to a reference (accuracy). On fully excited sequences, our method achieves the smallest SDs on 15/18 IMU-axis pairs, improves calibration accuracy by an average of 6.92%, and reduces processing time by 29.33% compared with state-of-the-art baselines. Under weak excitation, it attains the smallest SDs in 11/12 parameters and improves calibration results by 25.10% with comparable runtime. Overall, the approach enables accurate and robust LiDAR-IMU integration, improving perception quality and system reliability.
Conventional end-to-end autonomous driving methods often rely on explicit global scene representations, which typically consist of 3D object detection, online mapping, and motion prediction. In contrast, human drivers selectively attend to task-relevant regions and implicitly reason over the broader traffic context. Motivated by this observation, we introduce a lightweight end-to-end autonomous driving framework, InsightDrive. Unlike approaches that directly embed large language models (LLMs), InsightDrive introduces an Insight scene representation that jointly models attention-centric explicit scene representation and reasoning-centric implicit scene representation, so that scene understanding aligns more closely with human cognitive patterns for trajectory planning. To this end, we employ Chain-of-Thought (CoT) instructions to model human driving cognition and design a task-level Mixture-of-Experts (MoE) adapter that injects this knowledge into the autonomous driving model at negligible parameter cost. We further condition the planner on both explicit and implicit scene representations and employ a diffusion-based generative policy, which produces robust trajectory predictions and decisions. The overall framework establishes a knowledge distillation pipeline that transfers human driving knowledge to LLMs and subsequently to onboard models. Extensive experiments on the nuScenes and Navsim benchmarks demonstrate that InsightDrive achieves significant improvements over conventional scene representation approaches.
Traffic signs provide important traffic information for automatic driving, and accurate and complete traffic sign data of HD (High Definition) map provides important data support for intelligent transportation, automatic driving and other emerging service industries. Driving record data fills the data gap of crowd-source updating in HD maps, and the crowd-source updating method of road traffic facilities in HD maps using massive driving record data has become a new research hotspot. In this paper, an incremental HD map traffic sign crowd-source update method is proposed based on the driving record data. The traffic sign detection results are matched with the existing traffic signs in the HD map for traffic sign change detection, and the added results are optimized and fused for position, and the new sign positions are optimized using the unchanged signs to obtain the optimized new traffic sign positions. The experiments in Shanghai show that the matching method can meet the matching requirements of crowd-source updating; the accuracy of the traffic sign positions after position optimization and crowd-source fusion is obviously improved, with an average plane error of 3.69 m and a standard deviation of error of 3.29 m, which can provide data support for crowd-source updating of the HD map.
Traffic congestion is significantly affected by the built environment. Existing studies predominantly examine this through correlation analysis, overlooking causal mechanisms. This omission leads to unreliable feature selection in policy models and hinders evidence-based interventions. To address this, this study proposes a three-stage causal framework that rigorously assesses built environment impacts. The first stage identifies statistically significant correlations using multivariable least squares regression. The second stage applies five causal inference models - Granger causality, structural equation model, causal forest, causal impact, and convergent cross mapping - to uncover causality. The third stage assesses how the identified causal factors shape congestion patterns in perpetually congested roadways (PCRs). Applied to New York City (NYC), the United States, the results reveal 19 correlated and 11 causal impacts. Our key findings include: (1) Transit accessibility is the most robust causal factor, while built environment diversity exhibits time-dependent variability; (2) traffic light design demonstrates bidirectional causality with congestion; (3) PCRs exhibit four distinct spatiotemporal patterns, with bridge-related congestion having the most consistent impact. These results yielded policy recommendations for NYC transportation planning: (i) improve the first-and-last-mile connectivity through micro-mobility; (ii) deploy artificial intelligence-driven adaptive traffic signals; (iii) expand the capacity of critical bridge corridors near PCRs.