Multi-sensor fusion has significant potential in perception tasks for both indoor and outdoor environments. Especially under challenging conditions such as adverse weather and low-light environments, the combined use of millimeter-wave radar and RGB-D sensors has shown distinct advantages. However, existing multi-sensor datasets in the fields of autonomous driving and robotics often lack high-quality millimeter-wave radar data. To address this gap, we present a new multi-sensor dataset:RadarRGBD. This dataset includes RGB-D data, millimeter-wave radar point clouds, and raw radar matrices, covering various indoor and outdoor scenes, as well as low-light environments. Compared to existing datasets, RadarRGBD employs higher-resolution millimeter-wave radar and provides raw data, offering a new research foundation for the fusion of millimeter-wave radar and visual sensors. Furthermore, to tackle the noise and gaps in depth maps captured by Kinect V2 due to occlusions and mismatches, we fine-tune an open-source relative depth estimation framework, incorporating the absolute depth information from the dataset for depth supervision. We also introduce pseudo-relative depth scale information to further optimize the global depth scale estimation. Experimental results demonstrate that the proposed method effectively fills in missing regions in sensor data. Our dataset and related documentation will be publicly available at: https://github.com/song4399/RadarRGBD.
According to the technical requirements of the large field of view and wide spectral band in polarization spectral imaging detection, a wide spectral band and large field of view polarization imaging spectrometer based on Polarimetric-spectral intensity modulation (PSIM) was designed. For the front telescope group, according to the achromatic analysis of existing domestic glass materials, the achromatic glass from visible to short-wave infrared is selected. By controlling the light angle of the PSIM module in the mirror group, the incident angle demand on the PSIM module in the large field of view is realized. Based on the results of the analysis, optical design software is used to optimize the design. The design results show that the front telescopic system can achieve high-quality imaging with a wavelength range of 400 similar to 1 700 nm, a field angle of 72 degrees, a focal length of 20 mm, and an F-number of 4. The transfer function of the detector at the cut-off frequency in the full spectrum is better than 0.4, and the maximum incidence angle on the PSIM module is +/- 4.99 degrees, effectively ensuring the consistency of polarization modulation in each field of view. The post-spectral spectroscopic system uses a convex grating based on the Offner structure. The optimization results show that the point array of each band is less than one pixel and the MTF of the central wavelength at the Nyquist frequency of the detector reaches 0.6, and all indicators meet the design requirements. This paper has important practical significance for the engineering of polarization spectral imaging instruments based on PSIM wide spectrum and also has certain guiding significance for the achromatic design of wide-spectrum optical systems.
As the application of deep learning in general optical images becomes more and more widespread, the field of remote sensing images also begins to pay attention to the application of deep learning methods. Although deep learning detection algorithms have achieved better results than traditional detection algorithms, the detection results for poorly imaged SAR images still need improvement, and processing poorly imaged noisy SAR images is still a big challenge for existing algorithms. To address the problem of low precision and recall of existing algorithms for noisy SAR image detection, we propose a convolutional neural network detection algorithm based on min-pooling. First, we design a feature processing layer with min-pooling as the main structure to suppress the noise and then use a feature fusion layer to compensate for the missing information caused by pooling. To avoid problems such as redundancy in computation caused by the anchor-base algorithm, we choose the anchor-free algorithm as the main structure of ship detection. Finally, the model is evaluated using ordinary SAR image datasets and noisy SAR image datasets. Experiment results show that our proposed method has a better detection effect for noisy SAR images than other object detection models.
In order to meet the requirements of high-precision polarization detection, the influence of the polarization effect of the telescope group needs to be considered in the design of the channel-type polarization spectrometer, and the corresponding analysis and optimization should be carried out. First, the influencing factors of the polarization effect of the telescope group are analyzed, and the Mueller matrix model of the telescope group considering the polarization effect of the film system is established by using coordinate transformation and Mueller matrix multiplication method, which is brought into the polarization demodulation model of the channel polarimeter. Next, by simultaneously controlling the transmittance and phase retardation of S light and P light, a corresponding film system with low polarization effect is designed. Finally, the polarization effect simulation of the telescope group coated with different film systems is carried out by using the method of polarization ray tracing. The simulation results show that at the wavelengths of 580 nm and 750 nm, the polarization detection accuracy of the film system with low polarization effect and the film system with high polarization effect does not change significantly. At the wavelength of 420 nm, the polarization detection accuracy of the edge field of view of the low polarization effect film system is 3. 22% higher than that of the common film system. The low polarization effect film system effectively reduces the polarization effect of the telescope group of the channel-type polarization spectrometer and improves the polarization detection accuracy of the instrument.
A channeled spectropolarimeter is a powerful tool for the simultaneous measurement of the intensity, spectral, and polarization information of a target. However, the fore-optics introduce additional polarization information, which leads to inaccurate reconstruction of the Stokes parameters. In this study, we propose a simple method for polarimetric calibration and Stokes parameters reconstruction for a fieldable channeled spectropolarimeter. The polarization effects of the fore-optics and phase factors of the high-order retarders at varying view angles are considered and calibrated independently using a single reference beam. Moreover, the misalignment of the retarders is also considered. Simulation results demonstrate that the polarization effects of fore-optics can be precisely determined, enhancing the measurement accuracy of the Stokes parameters by approximately an order of magnitude. The effectiveness of the proposed method is also verified experimentally.
In this paper, we study the problem of unsupervised object detection from 3D point clouds in self-driving scenes. We present a simple yet effective method that exploits (i) point clustering in near-range areas where the point clouds are dense, (ii) temporal consistency to filter out noisy unsupervised detections, (iii) translation equivariance of CNNs to extend the auto-labels to long range, and (iv) self-supervision for improving on its own. Our approach, OYSTER (Object Discovery via Spatio-Temporal Refinement), does not impose constraints on data collection (such as repeated traversals of the same location), is able to detect objects in a zero-shot manner without supervised finetuning (even in sparse, distant regions), and continues to self-improve given more rounds of iterative self-training. To better measure model performance in self-driving scenarios, we propose a new planning-centric perception metric based on distance-to-collision. We demonstrate that our unsupervised object detector significantly outperforms unsupervised baselines on PandaSet and Argoverse 2 Sensor dataset, showing promise that self-supervision combined with object priors can enable object discovery in the wild. For more information, visit the project website: https://waabi.ai/research/oyster.
With the continuously thriving popularity around the world, fitness activity analytic has become an emerging research topic in computer vision. While a variety of new tasks and algorithms have been proposed recently, there are growing hunger for data resources involved in high-quality data, fine-grained labels, and diverse environments. In this paper, we present FLAG3D, a large-scale 3D fitness activity dataset with language instruction containing 180K sequences of 60 categories. FLAG3D features the following three aspects: 1) accurate and dense 3D human pose captured from advanced MoCap system to handle the complex activity and large movement, 2) detailed and professional language instruction to describe how to perform a specific activity, 3) versatile video resources from a high-tech MoCap system, rendering software, and cost-effective smartphones in natural environments. Extensive experiments and in-depth analysis show that FLAG3D contributes great research value for various challenges, such as cross-domain human action recognition, dynamic human mesh recovery, and language-guided human action generation. Our dataset and source code are publicly available at https://andytang15.github.io/FLAG3D.
Unmanned aerial vehicle (UAV)-based multispectral remote sensing effectively monitors agro-ecosystem functioning and predicts crop yield. However, the timing of the remote sensing field campaigns can profoundly impact the accuracy of yield predictions. Little is known on the effects of phenological phases on skills of high-frequency sensing observations used to predict maize yield. It is also unclear how much improvement can be gained using multi-temporal compared to mono-temporal data. We used a systematic scheme to address those gaps employing UAV multispectral observations at nine development stages of maize (from second-leaf to maturity). Next, the spectral and texture indices calculated from the mono-temporal and multi-temporal UAV images were fed into the Random Forest model for yield prediction. Our results indicated that multi-temporal UAV data could remarkably enhance the yield prediction accuracy compared with mono-temporal UAV data (R2 increased by 8.1% and RMSE decreased by 27.4%). For single temporal UAV observation, the fourteenth-leaf stage was the earliest suitable time and the milking stage was the optimal observing time to estimate grain yield. For multi-temporal UAV data, the combination of tasseling, silking, milking, and dough stages exhibited the highest yield prediction accuracy (R2 = 0.93, RMSE = 0.77 t·ha−1). Furthermore, we found that the Normalized Difference Red Edge Index (NDRE), Green Normalized Difference Vegetation Index (GNDVI), and dissimilarity of the near-infrared image at milking stage were the most promising feature variables for maize yield prediction.
A channeled spectropolarimeter can simultaneously obtain intensity, spectral, and polarization information. In the traditional model, the retarders must be oriented at specific angles. However, misalignments of the retarders are inevitable during assembly, and the status of the retarders is sensitive to environmental perturbations, which affects the performance of the channeled spectropolarimeter. In this study, a general channeled spectropolarimeter model was derived, in which the retarder orientations can be arbitrary and unknown. Meanwhile, the system is unaffected by environmental perturbation because it can self-calibrate to avoid fluctuations in the retarder orientations and phase retardations. The effectiveness and robustness of the model were verified through simulations and experiments.
In this paper, we propose an end-to-end self-driving network featuring a sparse attention module that learns to automatically attend to important regions of the input. The attention module specifically targets motion planning, whereas prior literature only applied attention in perception tasks. Learning an attention mask directly targeted for motion planning significantly improves the planner safety by performing more focused computation. Furthermore, visualizing the attention improves interpretability of end-to-end self-driving.
Modern self-driving systems heavily rely on deep learning. As a consequence, their performance is influenced significantly by the quality and richness of the training data. Data collection platforms can generate many hours of raw data on a daily basis, however, it is not feasible to label everything. Therefore, it is critical to have a mechanism to identify "what to label". Active learning approaches identify examples to label, but their interestingness is tied to a fixed model performing a particular task. These assumptions are not valid in self-driving, where we must solve a diverse set of tasks (i.e., perception, motion forecasting, and planning) and models frequently evolve over time. In this paper, we introduce a novel approach to dataset selection that exploits a diverse set of criteria that quantize interestingness of traffic scenes. Our experiments on a wide range of tasks and models demonstrate that the proposed curation pipeline is able to select datasets that lead to better generalization and improved performance.
Unmanned aerial vehicle (UAV) system is an emerging remote sensing tool for profiling crop phenotypic characteristics, as it distinctly captures crop real-time information on field scales. For optimizing UAV agro-monitoring schemes, this study investigated the performance of single-source and multi-source UAV data on maize phenotyping (leaf area index, above-ground biomass, crop height, leaf chlorophyll concentration, and plant moisture content). Four UAV systems [i.e., hyperspectral, thermal, RGB, and Light Detection and Ranging (LiDAR)] were used to conduct flight missions above two long-term experimental fields involving multi-level treatments of fertilization and irrigation. For reducing the effects of algorithm characteristics on maize parameter estimation and ensuring the reliability of estimates, multi-variable linear regression, backpropagation neural network, random forest, and support vector machine were used for modeling. Highly correlated UAV variables were filtered, and optimal UAV inputs were determined using a recursive feature elimination procedure. Major conclusions are (1) for single-source UAV data, LiDAR and RGB texture were suitable for leaf area index, above-ground biomass, and crop height estimation; hyperspectral outperformed on leaf chlorophyll concentration estimation; thermal worked for plant moisture content estimation; (2) model performance was slightly boosted via the fusion of multi-source UAV datasets regarding leaf area index, above-ground biomass, and crop height estimation, while single-source thermal and hyperspectral data outperformed multi-source data for the estimation of plant moisture and leaf chlorophyll concentration, respectively; (3) the optimal UAV scheme for leaf area index, above-ground biomass, and crop height estimation was LiDAR + RGB + hyperspectral, while considering practical agro-applications, optical Structure from Motion + customer-defined multispectral system was recommended owing to its cost-effectiveness. This study contributes to the optimization of UAV agro-monitoring schemes designed for field-scale crop phenotyping and further extends the applications of UAV technologies in precision agriculture.
In the past few years we have seen great advances in object perception (particularly in 4D space-time dimensions) thanks to deep learning methods. However, they typically rely on large amounts of high-quality labels to achieve good performance, which often require time-consuming and expensive work by human annotators. To address this we propose an automatic annotation pipeline that generates accurate object trajectories in 3D space (i.e., 4D labels) from LiDAR point clouds. The key idea is to decompose the 4D object label into two parts: the object size in 3D that's fixed through time for rigid objects, and the motion path describing the evolution of the object's pose through time. Instead of generating a series of labels in one shot, we adopt an iterative refinement process where online generated object detections are tracked through time as the initialization. Given the cheap but noisy input, our model produces higher quality 4D labels by re-estimating the object size and smoothing the motion path, where the improvement is achieved by exploiting aggregated observations and motion cues over the entire trajectory. We validate the proposed method on a large-scale driving dataset and show a 25% reduction of human annotation efforts. We also showcase the benefits of our approach in the annotator-in-the-loop setting.
3D object detection is a key component of many robotic applications such as self-driving vehicles. While many approaches rely on expensive 3D sensors such as LiDAR to produce accurate 3D estimates, methods that exploit stereo cameras have recently shown promising results at a lower cost. Existing approaches tackle this problem in two steps: first depth estimation from stereo images is performed to produce a pseudo LiDAR point cloud, which is then used as input to a 3D object detector. However, this approach is suboptimal due to the representation mismatch, as the two tasks are optimized in two different metric spaces. In this paper we propose a model that unifies these two tasks and performs them in the same metric space. Specifically, we directly construct a pseudo LiDAR feature volume (PLUME) in 3D space, which is then used to solve both depth estimation and object detection tasks. Our approach achieves state-of-the-art performance with much faster inference times when compared to existing methods on the challenging KITTI benchmark [1].
In this paper, we tackle the problem of detecting objects in 3D and forecasting their future motion in the context of self-driving. Towards this goal, we design a novel approach that explicitly takes into account the interactions between actors. To capture their spatial-temporal dependencies, we propose a recurrent neural network with a novel Transformer [1] architecture, which we call the Interaction Transformer. Importantly, our model can be trained end-to-end, and runs in real-time. We validate our approach on two challenging real-world datasets: ATG4D [2] and nuScenes [3]. We show that our approach can outperform the state-of-the-art on both datasets. In particular, we significantly improve the social compliance between the estimated future trajectories, resulting in far fewer collisions between the predicted actors.
The process of radiometric calibration would be coupled with the polarization properties of an optical system for spectropolarimetry, which would have significant influences on reconstructed Stokes parameters. In this paper, we propose a novel polarization radiometric calibration model that decouples the radiometric calibration coefficient and polarization properties of an optical system. The alignment errors of the polarization module and the variation of the retardations at different fields of view are considered and calibrated independently. According to these calibration results, the input Stokes parameters at different fields of view can be reconstructed accurately through the proposed model. Simulations are performed for the presented calibration and reconstruction methods, which indicate that the measurement accuracy of polarization information is improved compared with the traditional undecoupled calibration method.
We present a novel method for testing the safety of self-driving vehicles in simulation. We propose an alternative to sensor simulation, as sensor simulation is expensive and has large domain gaps. Instead, we directly simulate the outputs of the self-driving vehicle's perception and prediction system, enabling realistic motion planning testing. Specifically, we use paired data in the form of ground truth labels and real perception and prediction outputs to train a model that predicts what the online system will produce. Importantly, the inputs to our system consists of high definition maps, bounding boxes, and trajectories, which can be easily sketched by a test engineer in a matter of minutes. This makes our approach a much more scalable solution. Quantitative results on two large-scale datasets demonstrate that we can realistically test motion planning using our simulations.
Sensor simulation is a key component for testing the performance of self-driving vehicles and for data augmentation to better train perception systems. Typical approaches rely on artists to create both 3D assets and their animations to generate a new scenario. This, however, does not scale. In contrast, we propose to recover the shape and motion of pedestrians from sensor readings captured in the wild by a self-driving car driving around. Towards this goal, we formulate the problem as energy minimization in a deep structured model that exploits human shape priors, reprojection consistency with 2D poses extracted from images, and a ray-caster that encourages the reconstructed mesh to agree with the LiDAR readings. Importantly, we do not require any ground-truth 3D scans or 3D pose annotations. We then incorporate the reconstructed pedestrian assets bank in a realistic LiDAR simulation system by performing motion retargeting, and show that the simulated LiDAR data can be used to significantly reduce the amount of annotated real-world data required for visual perception tasks.
Modern autonomous driving systems rely heavily on deep learning models to process point cloud sensory data; meanwhile, deep models have been shown to be susceptible to adversarial attacks with visually imperceptible perturbations. Despite the fact that this poses a security concern for the self-driving industry, there has been very little exploration in terms of 3D perception, as most adversarial attacks have only been applied to 2D flat images. In this paper, we address this issue and present a method to generate universal 3D adversarial objects to fool LiDAR detectors. In particular, we demonstrate that placing an adversarial object on the rooftop of any target vehicle to hide the vehicle entirely from LiDAR detectors with a success rate of 80%. We report attack results on a suite of detectors using various input representation of point clouds. We also conduct a pilot study on adversarial defense using data augmentation. This is one step closer towards safer self-driving under unseen conditions from limited training data.
In this paper we show that High-Definition (HD) maps provide strong priors that can boost the performance and robustness of modern 3D object detectors. Towards this goal, we design a single stage detector that extracts geometric and semantic features from the HD maps. As maps might not be available everywhere, we also propose a map prediction module that estimates the map on the fly from raw LiDAR data. We conduct extensive experiments on KITTI as well as a large-scale 3D detection benchmark containing 1 million frames, and show that the proposed map-aware detector consistently outperforms the state-of-the-art in both mapped and un-mapped scenarios. Importantly the whole framework runs at 20 frames per second.