Radar and LiDAR have been widely used in autonomous driving as LiDAR provides rich structure information, and radar demonstrates high robustness under adverse weather. Recent studies highlight the effectiveness of fusing radar and LiDAR point clouds. However, challenges remain due to the modality misalignment and information loss during feature extractions. To address these issues, we propose a 4D radar-LiDAR framework to mutually enhance their representations. Initially, the indicative features from radar are utilized to guide both radar and LiDAR geometric feature learning. Subsequently, to mitigate their sparsity gap, the shape information from LiDAR is used to enrich radar BEV features. Extensive experiments on the View-of-Delft (VoD) dataset demonstrate our approach’s superiority over existing methods, achieving the highest mAP of 71.76% across the entire area and 86.36% within the driving corridor. Especially for cars, we improve the AP by 4.17% and 4.20% due to the strong indicative features and symmetric shapes.
In recent years, approaches based on radar object detection have made significant progress in autonomous driving systems due to their robustness under adverse weather compared to LiDAR. However, the sparsity of radar point clouds poses challenges in achieving precise object detection, highlighting the importance of effective and comprehensive feature extraction technologies. To address this challenge, this paper introduces a comprehensive feature extraction method for radar point clouds. This study first enhances the capability of detection networks by using a plug-and-play module, GeoSPA. It leverages the Lalonde features to explore local geometric patterns. Additionally, a distributed multi-view attention mechanism, DEMVA, is designed to integrate the shared information across the entire dataset with the global information of each individual frame. By employing the two modules, we present our method, MUFASA, which enhances object detection performance through improved feature extraction. The approach is evaluated on the VoD and TJ4DRaDSet datasets to demonstrate its effectiveness. In particular, we achieve state-of-the-art results among radar-based methods on the VoD dataset with the mAP of 50.24
Similarity estimation between periodic biological signals is crucial for vital sign-sensing applications. In this paper, we propose a novel approach for the accurate alignment and similarity estimation of such signals, named wave Dynamic Time Warping (wave DTW), and its variant, wave derivative DTW. Proposed methods take advantage of the inherent structure of periodic signals and align signal segments rather than entire sequences. Wave DTW employs a two-dimensional feature vector, derived from the signal's amplitude envelope and phase, to perform segment point-by-point alignment using DTW. Validation of the proposed algorithms based on Beth Israel Deaconess Medical Center Photoplethysmogram and Respiration Dataset (BIDMC PPG and Respiration Dataset) demonstrates that wave DTW and wave derivative DTW significantly outperform state-of-the-art algorithms DTW and derivative DTW in terms of path misalignment and path warpings and provide a more intuitive signal alignment.
Meta-few-shot learning algorithms, such as Model-Agnostic Meta-Learning (MAML) and Almost No Inner Loop (ANIL), enable machines to learn complex tasks quickly with limited data and based on previous experience. By maintaining the inner loop head of the neural network, ANIL leads to simpler computations and reduces the complexity of MAML. Despite its benefits, ANIL suffers from issues like accuracy variance, slow initial learning, and overfitting, hardening its adaptation and generalization. This work proposes “Look-Ahead ANIL” (LaANIL), an enhancement to ANIL for better learning. LaANIL reorganizes ANIL’s internal architecture, integrating parallel computing techniques (to process multiple training examples simultaneously across computing units) and incorporating Nesterov momentum (which accelerates convergence by adjusting the learning rate based on past gradient information and extracting informative features for look-ahead gradient computation). These additional features make our model more state-of-the-art capable and better edge-compatible and thus improve few-short learning by enabling models to quickly adapt to new information and tasks. LaANIL’s effectiveness is validated on established meta-few-shot learning datasets, including FC100, CIFAR-FS, Mini-ImageNet, CUBirds-200-2011, and Tiered-ImageNet. The proposed model achieved an increased validation accuracy by 7 ± 0.7% and a variance reduction by 44 ± 4% in two-way two-shot classification as well as increased validation by 5 ± 0.4% and a variance reduction by 18 ± 2% in five-way five-shot classification on the FC100 dataset and similarly performed well on other datasets.
The nameedge intelligence, also known asEdge AI, is a recent term used in the past few years to refer to the confluence of machine learning, or broadly speaking artificial intelligence, with edge computing. In this article, we revise the concepts regarding edge intelligence, such as cloud, edge, and fog computing, the motivation to use edge intelligence, and compare current approaches and analyze application scenarios. To provide a complete review of this technology, previous frameworks and platforms for edge computing have been discussed in this work to provide the general view of the basis for Edge AI. Similarly, the emerging techniques to deploy deep learning models at the network edge, as well as specialized platforms and frameworks to do so, are review in this article. These devices, techniques, and frameworks are analyzed based on relevant criteria at the network edge, such as latency, energy consumption, and accuracy of the models, to determine the current state of the art as well as current limitations of the proposed technologies. Because of this, it is possible to understand the current possibilities to efficiently deploy state-of-the-art deep learning models at the network edge based on technologies such as artificial intelligence accelerators, tensor processing units, and techniques that include federated learning and gossip training. Finally, the challenges of Edge AI are discussed in the work, as well as the future directions that can be extracted from the evolution of the edge computing and Internet of Things approaches.
Hand gesture and motion sensing offer an intuitive and natural form of human-machine interface. Air-writing systems allow users to draw alpha-numerical or linguistic characters in the virtual board in air through hand gestures. Traditionally, radar-based air-writing systems have been based on a network of radars, at least three, to localize the hand target through trilateration algorithm followed by tracking to extract the drawn trajectory, which is then followed by recognition of the drawn character by either Long-Short Term Memory (LSTM) utilizing the sensed trajectory or Deep Convolutional Neural Network (DCNN) utilizing a reconstructed 2D image from the trajectory. However, the practical deployments of such systems are limited since the detection of the finger or hand target by all three radars cannot be guaranteed leading to failure of the trilateration algorithm. Further placement of three or more radars for the air-writing solution is neither always physically plausible nor cost-effective. Furthermore, these solutions do not exploit the full potentials of deep neural networks, which are generally capable of learning features implicitly. In this paper, we propose an air-writing system based on a network of sparse radars, i.e. strictly less than three, using 1D DCNN-LSTM-1D transposed DCNN architecture to reconstruct and classify the drawn character utilizing only the range information from each radar. The paper employs real data using one and two 60 GHz milli-meter wave radar sensors to demonstrate the success of the proposed air-writing solution.
The increasing significance of technology in daily lives led to the need for the development of convenient methods of human-computer interaction (HCI). Given that the existing HCI approaches exhibit various limitations, hand gesture recognition-based HCI may serve as a more intuitive mode of human-machine interaction in many situations. In addition, the system has to be deployable on low-power devices for applicability in broadly defined Internet of Things (IoT) and smart home solutions. Recent advances exhibit the potential of deep learning models for gesture classification, whereas they are still limited to high-performance hardware. Embedded neural network accelerators are constrained in terms of available memory, central processing unit (CPU) clock speed, graphics processing unit (GPU) performance, and a number of supported operations. The aforementioned problems are addressed in this paper by namely two approaches - simplifying the signal processing pipeline to avoid recurrent structures and efficient topological design. This paper employs an intuitive scheme allowing for the generation of the data in the compressed form from the sequence of range-Doppler images (RDI). Thus, it allows for the design of a neural classifier avoiding the usage of recurrent layers. The proposed framework has been optimized for Intel® Neural Compute Stick 2 (Intel® NCS 2), at the same time achieving promising classification accuracy of 97.57%. To confirm the robustness of the proposed algorithm, five independent persons have been involved in the algorithm testing process.
Traditional cloud-based infrastructures are not enough for the current demands of Internet of Things (IoT) applications. Two major issues are the limitations in terms of latency and network bandwidth. In recent years, the concepts of fog computing and edge computing were proposed to alleviate these limitations by moving data processing capabilities closer to the network edge. Considering IoT growth and development forecasts, we believe the full potential of IoT can, in many cases, only be unlocked by combining cloud, fog and edge computing. This paper discusses four possible approaches for distributing workload among these levels. We also highlight developments and possibilities as well as consider challenges for implementation in the areas of hardware, machine learning, security, privacy and communication.