Multi-modality sensor fusion has emerged as a prevailing trend in 3D object detection tasks.However,existing research predominantly emphasizes the efficient fusion of data from diverse sensors,overlooking the potential severe con-sequences of calibration failures.In this paper,we present an innovative analysis and prediction of scenarios that could lead to fusion algorithm failures,along with introducing remedial measures to enhance model robustness.Specifically,leveraging our predicted outcomes,we proactively generate similar hazardous scenarios during the model training phase to facilitate generalization capabilities.Subsequently,we introduce a query mechanism during data fusion to identify the ap-propriate fusion target in the event of miscalibration.Evaluation on the nuScenes dataset demonstrates that our approach can mitigate model instability by up to 90%,and our framework can be seamlessly adapted to other fusion algorithms.
Roadside perception achieves comprehensive traffic scene understanding with multiple infrastructure-mounted sensors, and LiDAR-camera fusion plays a vital role. A critical prerequisite for effective fusion is accurate spatiotemporal calibration, including intrinsic calibration, extrinsic calibration and time synchronization. To ensure precise spatiotemporal calibration of different sensors, calibration targets are commonly used to provide distinct reference points. However, existing target-based calibration methods typically require manual detection steps, constraining the acquisition of temporal information for time synchronization. Additionally, in roadside scenarios, target placement is restricted to the drivable road surface, limiting the intrinsic calibration performance. Moreover, most existing methods solely rely on a single 2D reprojection error for optimization, resulting in suboptimal calibration performance. To alleviate these issues, we propose a joint intrinsic and extrinsic spatiotemporal calibration system using a novel target consisting of a high-reflectivity circular marker and four AprilTags. With this target, we first design an automatic target detection method based on physical and geometric rules to track target trajectories, eliminating the need for manual intervention. Then, we propose a joint optimization method to estimate the intrinsic and extrinsic parameters using trajectories from different sensors. After that, we represent the trajectories as GMMs to jointly refine the extrinsic parameters both in 3D space and on the 2D plane. Finally, time synchronization is achieved based on the spatially aligned trajectories. In simulator evaluations, the proposed method achieves maximum mean translation, rotation, and time synchronization errors of 1.5 cm, 0.13°, and 16 ms, respectively, outperforming existing methods. Real-world experiments further validate the effectiveness of our proposed method.
Fall detection is essential for safeguarding the health of elderly persons, enabling timely alerts to family members or the community. Millimeter-wave (mmWave) radar offers an effective solution, as it is privacy-preserving, non-invasive, and highly sensitive to motion. However, most existing approaches rely on multi-input, multi-output mmWave radar to generate 4D point clouds or range-angle heatmaps, significantly raising device costs. In this paper, we propose GR-Fall, a fall detection system with integrated gait recognition designed for indoor environments using single-input, single-output mmWave radar. To achieve high performance in various environments, we develop a data augmentation algorithm for target heatmaps and a cross-attention-based heatmap fusion framework for efficient fall detection. Furthermore, we introduce an innovative fall alarm mechanism based on joint fall-gait detection. This mechanism activates alerts when a person is detected having difficulty moving after a fall, thus minimizing unnecessary alarms and reducing strain on community resources. To evaluate GR-Fall, we recruit 33 volunteers and collect 5,799 instances across four different environments. Experimental results show that GR-Fall achieves 98.1% precision and 98.7% recall in new environments and with new participants, outperforming other state-of-the-art heatmap-based methods.
Extrinsic calibration is a fundamental step in sensor fusion systems. However, existing methods often lack generalization capabilities when facing diverse hardware configurations, sensor poses, and environmental conditions, hindering their large-scale deployment. To address this limitation, we propose a general extrinsic calibration method, CalibWorkflow. Our core innovation lies in positioning multimodal large language models (MLLMs) as "visual guides" for the calibration process, leveraging their powerful vision-language understanding capabilities to guide parameter search and refinement. This reliance on visual scene understanding, rather than specific geometric features or sensor characteristics, enables the method to generalize effectively across diverse hardware and environmental conditions. Specifically, CalibWorkflow employs a three-stage calibration pipeline: initial parameter search, coarse optimization, and fine optimization. First, it utilizes the MLLM to assess the visual consistency between the projected point cloud and the image, rapidly determining an initial range for the extrinsic parameters. Next, the MLLM serves as a differential evaluator, giving simple "better" or "worse" feedback on parameter changes to guide the search through the parameter space. Finally, the method refines the calibration by matching edge features and performing non-linear optimization. Extensive experiments are conducted across six diverse scenarios and four heterogeneous sensor combinations. CalibWorkflow achieves state-of-the-art sub-degree and centimeter-level accuracy on four datasets and demonstrates highly competitive performance on others. These results thoroughly validate the generalization and robustness when facing various scenarios. Codes will be available.
The fusion of LiDARs and cameras has been increasingly adopted in autonomous driving for perception tasks. The performance of such fusion-based algorithms largely depends on the accuracy of sensor calibration, which is challenging due to the difficulty of identifying common features across different data modalities. Previously, many calibration methods involved specific targets and/or manual intervention, which has proven to be cumbersome and costly. Learning-based online calibration methods have been proposed, but their performance is barely satisfactory in most cases. These methods usually suffer from issues such as sparse feature maps, unreliable cross-modality association, inaccurate calibration parameter regression, etc. In this paper, to address these issues, we propose CalibFormer, an end-to-end network for automatic LiDAR-camera calibration. We aggregate multiple layers of camera and LiDAR image features to achieve high-resolution representations. A multi-head correlation module is utilized to identify correlations between features more accurately. Lastly, we employ transformer architectures to estimate accurate calibration parameters from the correlation information. Our method achieved a mean translation error of $0.8751 \mathrm{cm}$ and a mean rotation error of $0.0562 ^{\circ}$ on the KITTI dataset, surpassing existing state-of-the-art methods and demonstrating strong robustness, accuracy, and generalization capabilities.
Numerous roadside perception datasets have been introduced to propel advancements in autonomous driving and intelligent transportation systems research and development. However, it has been observed that the majority of their concentrates is on urban arterial roads, inadvertently overlooking residential areas such as parks and campuses that exhibit entirely distinct characteristics. In light of this gap, we propose CORP, which stands as the first public benchmark dataset tailored for multi-modal roadside perception tasks under campus scenarios. Collected in a university campus, CORP consists of over 205k images plus 102k point clouds captured from 18 cameras and 9 LiDAR sensors. These sensors with different configurations are mounted on roadside utility poles to provide diverse viewpoints within the campus region. The annotations of CORP encompass multi-dimensional information beyond 2D and 3D bounding boxes, providing extra support for 3D seamless tracking and instance segmentation with unique IDs and pixel masks for identifying targets, to enhance the understanding of objects and their behaviors distributed across the campus premises. Unlike other roadside datasets about urban traffic, CORP extends the spectrum to highlight the challenges for multi-modal perception in campuses and other residential areas.
Fusing the data of millimeter-wave Radar sensors and high-definition cameras has emerged as a viable approach to achieving precise 3D object detection for roadside traffic surveillance. For roadside perception systems, earlier studies have pointed out that it is better to perform the fusion on the 2D image plane than on the BEV plane (which is popular for on-car perception systems), especially when the perception range is large (e.g., > 150 m). Image-plane fusion requires critical transformations, like perspective projection from the Radar's BEV to the camera's 2D plane and reverse IPM. However, real-world issues like uneven terrain and sensor movement degrade these transformations' precision, impacting fusion effectiveness. To alleviate these issues, we propose a geometry-based Radar-camera fusion method on the ground, namely FARFusion V2. Specifically, we extend the ground-plane assumption in FARFusion [20] to support arbitrary shapes by formulating the ground height as an implicit representation based on geometric transformations. By incorporating the ground information, we can enhance Radar data with target height measurements. Consequently, we can thus project the enhanced Radar data onto the 2D plane to obtain more accurate depth information, thereby assisting the IPM process. A real-time parameterized transformation parameters estimation module is further introduced to refine the view transformation processes. Moreover, considering various measurement noises across these two sensors, we introduce an uncertainty-based depth fusion strategy into the 2D fusion process to maximize the probability of obtaining the optimal depth value. Extensive experiments are conducted on our collected roadside OWL benchmark, demonstrating the excellent localization capacity of FARFusion V2 in far-range scenarios. Our method achieves an average location accuracy of 0.771m when we extend the detection range up to 500m.
We consider a quadrature-based finite difference discretization of one-dimensional scalar linear nonlocal conservation laws. The range of nonlocal interactions is allowed to vary in the spatial domain. We are particularly concerned with the convergence of the discrete approximation both in the nonlocal setting and in the local limit as the horizon parameter approaches zero. We present the first complete proof of the convergence of numerical discretization to both the nonlocal regime and the local limit of all feasible kernels, which in particular, establishes the asymptotically compatibility of the numerical scheme. We also present numerical results to demonstrate the effect of the variable horizon on the wave propagation described by the nonlocal model.
The recent trend of fusing complementary data from LiDARs and cameras for more accurate perception has made the extrinsic calibration between the two sensors critically important. Indeed, to align the sensors spatially for proper data fusion, the calibration process usually involves estimating the extrinsic parameters between them. Traditional LiDAR–camera calibration methods often depend on explicit targets or human intervention, which can be prohibitively expensive and cumbersome. Recognizing these weaknesses, recent methods usually adopt the autonomic targetless calibration approach, which can be conducted at a much lower cost. This paper presents a thorough review of these automatic targetless LiDAR–camera calibration methods. Specifically, based on how the potential cues in the environment are retrieved and utilized in the calibration process, we divide the methods into four categories: information theory based, feature based, ego-motion based, and learning based methods. For each category, we provide an in-depth overview with insights we have gathered, hoping to serve as a potential guidance for researchers in the related fields.
We study the propagation of singularities in solutions of linear convection equations with spatially heterogeneous nonlocal interactions. A spatially varying nonlocal horizon parameter is adopted in the model, which measures the range of nonlocal interactions. Via heterogeneous localization, this can lead to the seamless coupling of the local and nonlocal models. We are interested in understanding the impact on singularity propagation due to the heterogeneities of the nonlocal horizon and the local and nonlocal transition. We first analytically derive equations to characterize the propagation of different types of singularities for various forms of nonlocal horizon parameters in the nonlocal regime. We then use asymptotically compatible schemes to discretize the equations and carry out numerical simulations to illustrate the propagation patterns in different scenarios.