Frequency-Modulated Continuous-Wave (FMCW) LiDAR is capable of acquiring dense point clouds along with additional Doppler measurements. For discrete-time-based FMCW LiDAR odometry methods, motion distortion correction for both 3D and Doppler measurements should be performed before scan matching. However, this correction is inherently less accurate than continuous-time methods, and the residual errors are challenging to model. In this paper, we introduce CT-FLO, the first FMCW LiDAR odometry method that employs a linear continuous-time trajectory representation. Our method employs a sliding-window-based factor graph that incorporates point-to-plane factors, Doppler velocity factors, and a marginalization factor. We evaluate the proposed method on real-world datasets and demonstrate that it achieves excellent performance compared to state-of-the-art methods. The average root mean square of translational RPE, calculated using 100-meter sub-trajectories, was 1.7 m on geometrically challenging sequences and 0.19 m on feature-rich urban sequences. Additionally, an ablation study was conducted to investigate the contribution of the Doppler factor. We found that in a feature-rich environment with a longer time interval (e.g., 100 ms) for the control points, the Doppler factor can have a negative impact on the overall accuracy. Furthermore, we performed comparative experiments that determined the optimal time interval on a vehicle platform to be 40 ms for the test dataset. These results indicate that the optimal time interval is closely correlated with the geometric structure of the environment and the vehicle's motion state.
In autonomous driving, 3D object detection is essential for accurate perception and reliable decision-making. However, object motion and ego-motion often induce cross-frame spatiotemporal inconsistencies in BEV-based detectors, leading to temporal BEV feature misalignment and degraded spatiotemporal consistency. To address these challenges, we propose Co-Fusion4D, a unified framework that explicitly preserves cross-frame spatiotemporal consistency and suppresses temporal feature drift. Co-Fusion4D adopts a current-frame-centric strategy, treating the current frame as the primary source of information while selectively incorporating historical frames after spatiotemporal filtering and alignment. This dominant-complementary mechanism effectively mitigates cumulative alignment errors, suppresses noisy feature propagation, and exploits reliable temporal cues for a more consistent BEV representation. In addition, Co-Fusion4D integrates a Dual Attention Fusion (DAF) module to further enhance spatiotemporal feature interaction. DAF jointly leverages intra-frame spatial attention and inter-frame temporal attention to adaptively align and fuse multi-frame features, emphasizing motion-consistent regions while suppressing spurious correlations. By departing from conventional uniform fusion paradigms, this design substantially improves the temporal stability and discriminative capability of BEV representations. Extensive experiments on the nuScenes benchmark demonstrate that Co-Fusion4D achieves state-of-the-art performance, with 74.9
Abstract: Localization and mapping capabilities are essential for intelligent robotic systems, and LiDAR has proven to be an efficient sensor for these tasks. However, a single LiDAR suffers from a limited field of view (FoV) and system fragility, while multi‑LiDAR configurations introduce challenges of spatio‑temporal inconsistency and computational overhead. In this paper, we propose an intuitive multi‑LiDAR‑inertial odometry system that explicitly accounts for both temporal and spatial inconsistency, especially among heterogeneous LiDARs. To improve data synchronization, a novel data rearrangement strategy is introduced to prevent sensor failure. A dedicated LiDAR measurement model is formulated that incorporates anisotropic noise propagated through extrinsic parameters, thereby addressing point‑wise spatial discrepancy. For temporal uncertainty, motion priors derived from time deviations of individual points are integrated. Ultimately, a stream‑based extended Kalman filter (EKF) framework is employed for measurement‑level state update, processing each LiDAR and IMU measurement directly. Extensive experiments were conducted on benchmark datasets as well as self‑collected data involving different platforms and heterogeneous LiDARs, comparing our system against other state‑of‑the‑art methods. A carefully designed field test in a dense forest further demonstrated practical applicability. The diameter at breast height (DBH) and relative positions of marked trees were recorded and compared to evaluate the internal consistency of the reconstructed point cloud maps. Overall, our system achieves the best performance, with average error improvements exceeding 50%, demonstrating great promise for complex, real‑world environments.
Seasonal measurement of canopy parameters plays a critical role in assessing ecological and environmental health. Mountainous forests represent a primary reservoir of terrestrial carbon stocks. However, generating accurate 3D maps of these environments presents significant technical challenges. In this paper, a novel quadric-constrained LiDAR-inertial system is proposed for UGV-UAV collaborative mapping. With IMU preintegration, each laser scan is first undistorted and then segmented to derive ground features, edge features and planar features. Efficient dual-optimization based on point-to-line and point-to-plane correspondences is employed for optimal pose estimation of UGV. Leveraging this initial pose, local surfaces are extracted from the UAV map and modeled as quadric surfaces. These surfaces serve as geometric priors within the factor graph, where quadric factors actively constrain vertical drift during optimization. In the backend, a flexible factor graph is maintained for global consistency, incorporating IMU factors, odometry factors as well as quadric factors. A large-scale field test is conducted in a mountainous forest with over 50 m elevation difference. Compared to other state-of-theart frameworks, our system obtains the least 0.23% error ratio in elevation direction and exhibits the best 47.48% in the overall mapping performance, showing a promising potential for integrated understory and canopy measurements in complex forest environments. (c) 2026 COSPAR. Published by Elsevier B.V. All rights are reserved, including those for text and data mining, AI training, and similar technologies.
In this paper, we propose a LiDAR-based aerial exploration framework that enables coverage path (CP) guidance directly on point cloud maps. Existing CP-guided methods mitigate frequent revisitations by incorporating unknown-zone centroids as topological nodes, yet this operation fundamentally relies on occupancy grid representations that explicitly encode free and unknown states. Such representations are incompatible with point cloud-driven frameworks that eschew volumetric grids. To bridge this gap, we introduce a spatial inference mechanism that distinguishes explored from unknown zones using only geometric evidence, enabling unknown zone node creation without auxiliary occupancy structures. Building on this enabler, we further design a hierarchical planning architecture that integrates global CP sequencing with local path refinement, which not only guides the UAV toward unknown zones but also ensures smooth trajectory execution. Extensive simulation and real-world experiments validate that our framework effectively eliminates revisitation redundancy and reduces exploration time against state-of-the-art baselines.
In this paper, we propose RAEM, a robust autonomous exploration framework for quadruped robots operating in multi-floor environments. Most existing ground-robot exploration approaches rely on planar traversability representations, which cannot adequately represent the overlapping structures and cross-floor connectivity of multi-floor buildings. Although tomography-based representations provide effective traversability modeling for multi-floor navigation, maintaining a global tomography map incurs substantial computational overhead for online exploration with frequent replanning. Moreover, sparse and fragmented LiDAR observations in stairwells can degrade local traversability estimation, leading to irregular viewpoint placement and temporary topological disconnections. To address these challenges, RAEM adopts a hybrid local-global traversability representation, in which a local tomography map and an explicitly categorized local 3D grid map are used for online terrain analysis and connectivity evaluation, while an elevation-aware global topological graph is incrementally constructed from these local spatial representations for efficient cross-floor exploration planning. We further introduce a staircase center alignment strategy to reduce abrupt yaw variations during climbing and a dual path searching mechanism to recover guidance paths when the global topology is locally disconnected. Extensive simulation and real-world experiments demonstrate robust and computationally stable autonomous exploration across multi-floor structures, including continuous exploration of a five-floor stairwell.
3D object detection in driving scenarios is particularly challenging due to factors such as sensor noise, occlusions, and the inherent sparsity of LiDAR point clouds, which can lead to the loss or incompleteness of key features, in turn affecting perception performance. To address these challenges, we propose Co-Fix3D, an advanced detection framework that integrates Local and Global Enhancement (LGE) modules to refine Bird's Eye View (BEV) features. The LGE module employs Discrete Wavelet Transform (DWT) to refine local features at a fine scale, which helps capture frequency details and subtle variations in the environment, and incorporates an attention mechanism to enhance global feature representations across the entire scene. Moreover, we adopt multi-head LGE modules that each concentrate on targets with varying levels of detection difficulty, further improving our overall perception performance. On the nuScenes dataset, Co-Fix3D achieves a new SOTA performance with 69.4% mAP and 73.5% NDS compared to other competing methods, while on the multimodal benchmark, it achieves 72.3% mAP and 74.7% NDS, respectively.
Accurate localization and detailed environmental perception are critical for planetary exploration, particularly in unstructured and complex terrains. To support such efforts, we present the WHU-PA3D dataset, a planetary analog dataset collected in the Mars-like Qaidam Basin. This dataset includes eight handheld LiDAR point cloud sequences and seven UAV image sequences, documenting diverse terrains such as yardangs, water-eroded landscapes, and canyons. It integrates multimodal data, including RGB imagery, LiDAR, and IMU measurements, offering a realistic testing platform for SLAM and image classification algorithms. Benchmarking experiments reveal the strengths and limitations of existing algorithms in handling complex and unstructured environments. The WHU-PA3D dataset provides a useful reference for planetary surface localization, geological analysis, and mission planning.
Multi-robot exploration in non-exposed spaces, such as underground tunnels, disaster-stricken areas, and planetary subsurface environments (e.g., lunar lava tubes, Martian caves), presents significant challenges due to limited perception, complex terrain, and the absence of Global Navigation Satellite System signals. Overcoming these challenges requires a combination of advanced perception techniques, intelligent path planning algorithms, and multi-robot coordination strategies. This paper systematically categorizes state-of-the-art methods into three key areas: environmental perception, path planning, and multi-agent coordination. We review the progression from traditional exploration methods to artificial intelligence-based techniques, including deep reinforcement learning and graph neural networks, which have demonstrated improved multi-robot performance. Moreover, we examine sensor fusion techniques, uncertainty modeling, and adaptive decision-making frameworks. Furthermore, we analyze major challenges, including limited communication, real-time adaptability, and decentralized coordination in multi-robot systems. Addressing these challenges requires more robust task allocation mechanisms, enhanced map representation strategies, and efficient learning-based control frameworks. Finally, we discuss future research directions, emphasizing scalable multi-robot intelligence, enhanced perception models, and resilient collaboration frameworks to improve exploration efficiency in extreme environments.
Phase unwrapping poses a critical challenge in 3D reconstruction, particularly due to the presence of noise and discontinuities that compromise the accuracy of phase extraction. Existing convolutional neural network (CNN)-based methods have struggled to effectively integrate traditional approaches while fully utilizing both multi-frequency, i.e., temporal, and spatial information of the phase. In this paper, we propose a novel phase unwrapping method, STPhaseNet, which incorporates both temporal and spatial phase information into the CNN framework. Specifically, we introduce a temporal feature fusion module and a local attention mechanism to extract and integrate features from different frequency phases. To further leverage spatial phase information, we develop a spatial information extraction module that enlarges the local receptive field of the convolution and assigns weights based on the phase information of horizontal and vertical coordinates. Additionally, we design a globally optimized gradient residual loss function to exploit spatial constraints more effectively. To address the lack of real-world training data, we apply a Random Matrix Enlargement (RME) method to generate high-quality dual-frequency wrapped phase data along with corresponding absolute phases for training purposes. Extensive experiments demonstrate that STPhaseNet outperforms existing methods, achieving superior performance in phase unwrapping tasks.
In transmission lines, regular inspections are crucial for maintaining their safe operation. Automatic and accurate detection of power transmission facility components (power components) in inspection imagery is an effective way to monitor the status of electrical assets within the Right of Ways (RoWs). However, the multitude of smallscale objects (e.g. grading rings, vibration dampers) in inspection imagery poses enormous challenges. To address these challenges, we propose a coarse-to-fine object detector named RF-DET. It adopts a Refocus Framework to refine the detection accuracy of small objects within the Regions of Interest of the Power Components (P-RoIs) generated through explicit context. On the basis above, an Implicit Context Aggregation Attention Module (ICAM) is proposed. ICAM utilizes a multi-branch structure to capture and aggregate multidirectional positional and global information, enabling in-depth mining of the implicit context among small objects. To verify the performance of this detector, a benchmark dataset named DOPI-UAV is constructed, comprising 4,438 UAV oblique images and 54,591 instances, encompassing six common categories of power components and one category of defect. Experimental results show that RF-DET achieves mAP of 62.7%, 55.7%, 84.6%, and 52.8% on the DOPI-UAV, Tower, CPLID, and InsD datasets, respectively. Compared to the state-ofthe-art method, such as YOLOv9, RF-DET attains significant performance improvements, with increases of 5.2% in mAP and 6.4% in mAP50, respectively. Especially, the APS shows an improvement of 8.3%. The datasets and codes are available at https://github.com/DCSI2022/RF-DET.
Plane instance segmentation from RGB-D data is critical for BIM-related tasks. However, existing deep-learning methods rely on only RGB bands, overlooking depth information. To address this, PlaneSAM, a Segment-Anything-Model-based network, is proposed. It fully integrates RGB-D bands using a dual-complexity backbone: a simple branch primarily for the D band and a high-capacity branch mainly for RGB bands. This structure facilitates effective D-band learning with limited data, preserves EfficientSAM's RGB feature representations, and enables task-specific fine-tuning. To improve adaptability to RGB-D domains, a self-supervised pretraining strategy is introduced. EfficientSAM's loss is also optimized for large-plane segmentation. Additionally, plane detection is performed using Faster R-CNN, enabling fully automatic segmentation. State-of-the-art performance is achieved on multiple datasets, with <10% additional overhead compared to EfficientSAM. The proposed dual-complexity backbone shows strong potential for transferring RGB-based foundation models to RGB+X domains in other scenarios, while the pretraining strategy is promising for other data-scarce tasks.
Recent advancements in Bird's Eye View (BEV)-based 3D object detection have highlighted its potential to enhance scene understanding in autonomous driving applications. However, existing BEV-based methods utilizing point clouds for 3D object detection face significant challenges due to inherent sparsity and noise, which often compromise the accuracy of BEV representations. Furthermore, in multimodal 3D object detection, the lack of depth information in images can lead to distortions in the image BEV features generated through view transformations, further leading to inaccuracies in the fused BEV representation. To overcome these limitations, we introduce BEVFix, an innovative end-to-end 3D object detection method designed to refine BEV representations. BEVFix starts by generating a mask based on the point cloud distribution to identify specific regions requiring repair. This is followed by our WaveRefiner, which employs Discrete Wavelet Transform (DWT) for multi-frequency decomposition and utilizes a Feed-Forward Network (FFN) to isolate noise while selectively retaining critical features. These components work synergistically to reduce noise and enhance BEV representations. Experiments on the nuScenes and Waymo datasets demonstrate that BEVFix significantly improves performance, achieving state-of-the-art results. The source code will be publicly available at https://github.com/WenxuanLi-whu/Co-Fix3d.
Multi-focus image fusion is a highly challenging task in robotic vision. Although significant progress has been made in deep learning-based Multi-focus image fusion, existing methods still have some limitations in generating satisfactory results, especially for continuous outdoor scenes. There are two main problems: i) for outdoor scenes, the defocus diffusion effect can lead to blurred areas in near focus and far focus images; ii) for continuous depth scenes, there is no boundary between near and far focus images. To deal with the above problems, this paper proposes an end-to-end multi-focus image fusion method - XFusion, which takes an X-shaped multi-scale convolutional module as the basic network and the transformer module as the fusion unit. The fusion unit utilizes cross-focus feature statistics to effectively de-blur the defocused area to a certain extent. Furthermore, we constructed a new dataset for outdoor scene multi-focus image fusion. Experiments show that the proposed XFusion achieves excellent performance compared to other competitors on our dataset and several other benchmark datasets. The source code and dataset can be found on https://github.com/Shouxi-Zhao/XFusion .
Non-exposed spaces,such as indoor environments,underground utility tunnels,and natural caves,are partially enclosed areas that have gained increasing attention as urbanization progresses.This growing demand for exploration has significant implications.Investigating these spaces and acquiring spatial information within them can support the development of new infrastructure and facilitate digital transformation.However,exploring non-exposed spaces presents several challenges due to their complex structures,signal isolation from external sources,and degraded conditions.In such spaces,external positioning signals,such as GNSS,are often unavailable,complicating the reliance on these signals for localization.Additionally,many non-exposed spaces suffer from degraded environmental conditions,which further hinder self-localization.The intricate internal structures of these spaces also pose safety risks for human entry.The advancement of unmanned systems technology presents a promising solution to these challenges.To address these issues,we designed an autonomous aerial unmanned system equipped with panoramic LiDAR,providing a wide field of view,and integrated with modules for localization,mapping,planning,and control,enabling autonomous flight in uncharted spaces.The system consists of five main components:sensor input,localization and mapping,planning,control,and the unmanned system itself.Sensor data from LiDAR and IMU are utilized for state estimation and real-time mapping.The system generates an occupancy grid map for trajectory planning,followed by optimization.Commands are then sent to the flight controller,which integrates manual and planned inputs to maintain stable flight.The system's pose is monitored through Mavros,ensuring autonomous flight by controlling the motors via ESCs.Additionally,we propose a method for autonomous exploration of non-exposed spaces that involves point cloud mapping based on manually or autonomously assigned target points on a pre-established map.To validate the proposed autonomous aerial unmanned system and exploration method,we conducted experimental verifications in both simulated and real-world scenarios.We first selected a typical indoor scenario,"Indoorl,"within the XTDrone simulation environment under GNSS-denied conditions.The simulation utilized the Iris drone,supported by PX4 firmware with PX4 software-in-the-loop communication,and an Intel RealSense D455 simulation module for capturing visible and depth images.The simulation results demonstrated a detection efficiency of 23.94 m3/s using the proposed exploration method.Subsequently,real-world experiments were conducted in a section of an underground parking lot at Wuhan University's Xinghu Experimental Building.The autonomous aerial unmanned system demonstrated stable flight,achieving a detection efficiency of 53.94 m3/s in complex environments,such as corridors,pipelines,and rooms.The experimental results confirm the feasibility of the proposed method.Real-world experiments achieved a exploration efficiency exceeding 50%m3/s,validating the system's capability to efficiently explore non-exposed spaces.This demonstrates the significant potential of the system for future applications.Further research will focus on viewpoint generation based on specific targets to enhance the system's intelligence,enabling intelligent exploration of non-exposed spaces.This will improve the system's autonomy and adaptability in complex environments.
Multirobot Simultaneous localization and mapping (SLAM) has become a hot spot all over the world while there still exist some key issues to be addressed for collaborative mapping, especially for heterogenous platforms. This paper presents a self-designed LiDAR-inertial-mapping system on lightweight manned helmet and unmanned ground vehicle (UGV) platforms. A stable-triangular-descriptor-based place recognition method is proposed to detect overlap between varied field of view (FoV), followed by a Delaunay verification to restrain the global similarity. Based on the conjugated plane primitives across the platforms, bundle adjustment is derived to enhance the map consistency. All frames within the sliding window will be mutually constrained and optimized through the collective pose graph. Experiments on place recognition, pose estimation and point cloud mapping are conducted in different kinds of environments. Compared to other famous works (e.g. FastLIO2, LIOSAM, PointLIO), our system can attain the best results in almost all cases. The final improvements against them of mapping precision reach 15.71% with same LiDAR and 14.49% for heterogenous LiDARs on UGV. As for the helmet, this disparity varies from 0.6% for dataset with small intersection to 22.33% at most.
With the development of urbanization, underground spaces have become an important part of human life. Accurately surveying and describing the spatial information of underground spaces is of significant importance. However, the complex environment of underground spaces, often characterized by darkness, narrowness, lack of structure, and GNSS-denied conditions, presents tremendous challenges for intelligent information acquisition and analysis in such environments. To address these challenges, we propose ALM-LED, an autonomous LiDAR mapping framework designed for underground environments, which integrates a cost-effective, Luojia explorer anti-collision drone system featuring a lightweight LiDAR sensor and a carbon fiber frame. This framework consists of two main modules: localization and mapping, and planning and control. Localization and mapping integrates LiDAR point cloud data, IMU data, as well as flight control barometer and magnetometer sensor data, enabling robust localization and high-precision mapping in GNSS-denied underground environments. Planning and control constructs a triple constrained cost function for flight trajectory optimization based on smoothness, dynamic feasibility, and collision penalty terms, providing autonomous flight paths for the anti-collision drone system and combining MAVROS, achieving robust control. To validate the proposed system and methods, we conducted experiments in one simulation scenario and two real-world underground scenarios. The experiments demonstrate that ALM-LED achieves average mapping efficiencies exceeding 100 m3/s in simulated environments and 50 m3/s in real-world scenarios when applied to underground spaces. The flight trajectory estimated by the localization and mapping subsystem is nearly identical to the target trajectory. Point cloud maps with volumes of 879 m3, 26313 m3, 22240m3 and 115m3 were generated in four real-world scenarios, with point cloud map accuracies reaching 0.034m, 0.31m, 0.088m and 0.053m, respectively. The experimental results indicate that ALM-LED can achieve efficient and accurate information acquisition in underground spaces, demonstrating high application potential. To support the research community, the key source code for this work is publicly available at the following repository: https://github.com/DCSI2022/ALM-LED.
Since the 1960s,remote sensing science and technology has emerged as a competitive high-tech field,with major countries striving to advance their capabilities.It has become a fundamental tool for human research in the earth system science and the comprehensive application of aerospace information across multiple domains.Recently,two significant developments warrant attention:First,in 2022,the Ministry of Education of China officially recognized Remote Sensing Science and Technology as a first-level interdisciplinary discipline within the graduate education framework,thereby strengthening foundational research in remote sensing and broadening its application areas.Second,the rise of artificial intelligence technologies,particularly deep learning,has ushered in a new paradigm for data-driven analysis and application of remote sensing data.While remote sensing fundamentally belongs to the domain of electromagnetic radiation physics,the associated physical models have been indispensable for the development of quantitative remote sensing.Nevertheless,the data-driven deep learning paradigm has introduced transformative ideas and methodologies to the field.Moving forward,the synergy between physical models and artificial intelligence will undoubtedly shape the future trajectory of remote sensing research and applications.In this context,a deeper exploration of the core concepts and fundamental issues in remote sensing science is crucial for achieving significant technological breakthroughs and scientific discoveries within this discipline. This article begins by examining the physical origins of remote sensing science,focusing on the interaction between ground objects and electromagnetic waves,which produces spectral radiation images under specific conditions.It explores the characteristics of various remote sensing methods across the electromagnetic spectrum,including solar reflected radiation in the visible to shortwave infrared remote sensing,daylight-induced chlorophyll fluorescence(SIF)remote sensing,laser remote sensing,both medium and longwave infrared remote sensing,and microwave remote sensing.The fundamental theoretical issues in remote sensing science are categorized into three primary characteristics:radiative,spectral,and temporal characteristics,along with five major effects:scale,atmospheric,angular,adjacent,and transfer effects.The former pertains to the intrinsic physical and chemical properties of ground objects within the electromagnetic spectrum,while the latter relates to factors such as imaging scale,atmospheric conditions,observation angle,and background environment.This discussion includes the expression and variation patterns of remote sensing features of land objects formed under diverse observation modes and conditions. The radiative characteristics reflect overall difference in term of radiation across different electromagnetic bands for various land covers,closely tied to geophysical and chemical properties.The spectral characteristics of land cover manifest as variations in the intensity of reflected and emitted signals with wavelength,highlighting significant differences in absorption,reflection,and emission behaviors among different materials,known as spectral characteristics.Temporal characteristics pertain to the systematic changes in spectral reflection or emission over time,aiding in remote sensing identification or feature inversion of land cover.The scale effect refers to the changes in remote sensing observation characteristics due to variations in pixel area size,influenced by spatial resolution or point scanning density(e.g.,laser scanning spot density).The atmospheric effect describes how electromagnetic waves are impacted by the absorption,scattering,and emission from atmospheric particles during remote sensing imaging,leading to radiation distortion in image data.The angular effect highlights the directional nature of the interaction between land cover and electromagnetic waves,resulting in significant anisotropic characteristics and variations in radiation values based on the angles of incident radiation,remote sensing observation,and electromagnetic wave wavelength.The adjacent effect refers to the influence of spatial structure heterogeneity among land features,which can create cross-radiation contributions from non-target pixels to target pixels,dependent on spatial distribution and remote sensing observation mode.Finally,the transfer effect encompasses the changes in imaging quality after the electromagnetic signal of the ground objects entering the remote sensing system,including the processes such as photoelectric conversion,signal transmission,and digital recording. The review and discussion presented in this article on the fundamental issues of remote sensing science aim to deepen theoretical research in the field,particularly in the context of artificial intelligence.This exploration is intended to foster innovative methods in remote sensing technology and applications,promote the collaborative evolution of AI for Science and Science for AI in remote sensing,and encourage profound cross-disciplinary integration between remote sensing and other fields.
LiDAR-Inertial simultaneous localization and mapping (LI-SLAM) plays a crucial role in various applications such as robot localization and low-cost 3D mapping. However, factors including inaccurate motion distortion estimation and pose graph constraints, and frequent LiDAR feature degeneracy present significant challenges for existing LI-SLAM methods. To address these issues, we propose DALI-SLAM, an accurate and robust LI-SLAM that consists of degeneracy-aware LiDAR-inertial odometry (DA-LIO) with a dual spline-based motion distortion correction (DS-MDC) module, and multi-constraint pose graph optimization (MC-PGO). Considering the cumulative errors of micro-electromechanical systems (MEMS) inertial measurement unit (IMU) integration, two continuous-time trajectories in the sliding window are fitted to update the discrete IMU poses for accurate motion distortion correction. In the LiDAR-inertial fusion stage, LiDAR feature degeneracy is detected by analyzing the Jacobian matrix and a remapping strategy is introduced into the updating of error state Kalman Filter (ESKF) to mitigate the influence of degeneracy. Furthermore, in the back-end optimization stage, three types of submap constraints are accurately built with dedicated strategy through a robust variant of the iterative closest point (ICP) method. The proposed method is comprehensively validated using data collected from a helmet-based laser scanning system (HLS) in representative indoor and outdoor environments. Experiment results demonstrate that the proposed method outperforms the SOTA methods on the test data. Specifically, the proposed DS-MDC module reduces trajectory root mean square errors (RMSEs) by 7.9 %, 5.8 %, and 3.1 %, while the degeneracy-aware update strategy achieves additional reductions of 43.3 %, 17.7%, and 4.9 %, respectively, across three typical sequences compared to existing methods, thereby effectively improving trajectory accuracy. Furthermore, the results of DA-LIO demonstrate a maximum RMSE within 1 kilometer of approximately 1 meter in outdoor environments, achieving superior performance compared to the SOTA method FAST-LIO2. After performing MC-PGO, the RMSEs of the trajectories are reduced by 25.2 %, 9.2 %, and 52.4 %, respectively, across three typical sequences, demonstrating better performance compared to the SOTA method HBA.
Transformers have been seldom employed in point cloud roof plane instance segmentation, which is the focus of this study, and existing superpoint Transformers suffer from limited performance due to the use of low-quality superpoints. To address this challenge, we establish two criteria that high-quality superpoints for Transformers should satisfy and introduce a corresponding two-stage superpoint generation process. The superpoints generated by our method not only have accurate boundaries, but also exhibit consistent geometric sizes and shapes, both of which greatly benefit the feature learning of superpoint Transformers. To compensate for the limitations of deep learning features when the training set size is limited, we incorporate multidimensional handcrafted features into the model. Additionally, we design a decoder that combines a Kolmogorov-Arnold Network with a Transformer module to improve instance prediction and mask extraction. Finally, our network's predictions are refined using traditional algorithm-based postprocessing. For evaluation, we annotated a real-world dataset and corrected annotation errors in the existing RoofN3D dataset. Experimental results show that our method achieves state-of-the-art performance on our dataset, as well as both the original and reannotated RoofN3D datasets. Moreover, our model is not sensitive to plane boundary annotations during training, significantly reducing the annotation burden. Through comprehensive experiments, we also identified key factors influencing roof plane segmentation performance: in addition to roof types, variations in point cloud density, density uniformity, and 3D point precision have a considerable impact. These findings underscore the importance of incorporating data augmentation strategies that account for point cloud quality to enhance model robustness under diverse and challenging conditions.