Robust object 6D pose tracking is critical for robotic systems operating in dynamic and occluded scenes. Per-frame estimators are accurate but computationally expensive, while current trackers struggle with fast motion and complete occlusion due to their reliance on continuous visibility. To address these challenges, we present RRTrack, an efficient, recoverable object 6D pose tracker that enables robust tracking through fast motion and target disappearance--reappearance. RRTrack introduces a 2D--6D closed-loop tracking strategy that integrates memory-based video object segmentation (VOS) with 6D pose refinement. The 2D branch maintains target localization, and the 6D branch verifies geometric consistency before memory updates. In addition, a DINOv2-based dual-bank template matching module is developed to recover lost targets by jointly exploiting offline synthetic templates and online observation anchors while maintaining real-time efficiency. We also introduce a synthetic RGB-D benchmark comprising three robotic scenarios with fast motion and full occlusion. Experimental results on the synthetic benchmark demonstrate that RRTrack improves equal-subset mean ADD-S AR by 66.3\% and ADD-S AUC by 65.7\% over FoundationPose while achieving 55.2 FPS. Real-world experiments further validate the robustness of RRTrack under noisy sensing conditions. Project page: https://github.com/7kevin24/RRTrack
The internal state of the battery seriously affects its lifespan, safety, and reliability. CT scanning can obtain internal images of the battery without damaging it. This study focused on the discrimination of main internal defects in batteries: electrode crack, electrode deformation, inclusion and burr, among which inclusions are further divided into bubbles and metal particles. We first fused features extracted by 0 degrees and 90 degrees Gabor filtering kernels to enhance the discrimination between defects and electrode backgrounds. For the localization and type discrimination of battery defects, an improved YOLOv8 method, incorporating a Convolutional Block Attention Module in the Backbone layer and a lightweight Group Shuffle Convolution module in the Neck layer, is designed for defect localization and type discrimination. We finally achieved an average detection accuracy of 99.2%, reducing the parameter count by 6k on the basis of the original YOLOv8 model. The detection speed for a single image is 9.5ms.
Bin-picking is a practical and challenging robotic manipulation task, where accurate 6D pose estimation plays a pivotal role. The workpieces in bin-picking are typically textureless and randomly stacked in a bin, which poses a significant challenge to 6D pose estimation. Existing solutions are typically learning-based methods, which require object-specific training. Their efficiency of practical deployment for novel workpieces is highly limited by data collection and model retraining. Zero-shot 6D pose estimation is a potential approach to address the issue of deployment efficiency. Nevertheless, existing zero-shot 6D pose estimation methods are designed to leverage feature matching to establish point-to-point correspondences for pose estimation, which is less effective for workpieces with textureless appearances and ambiguous local regions. In this paper, we propose ZeroBP, a zero-shot pose estimation framework designed specifically for the bin-picking task. ZeroBP learns Position-Aware Correspondence (PAC) between the scene instance and its CAD model, leveraging both local features and global positions to resolve the mismatch issue caused by ambiguous regions with similar shapes and appearances. Extensive experiments on the ROBI dataset demonstrate that ZeroBP outperforms state-of-the-art zero-shot pose estimation methods, achieving an improvement of 9.1% in average recall of correct poses.
Direct pose estimation networks aim to directly regress the 6D poses of target objects in the scene image using a neural network. These direct methods offer efficiency and an optimal optimization target, presenting significant potential for practical applications. However, due to the complex and implicit mappings between input features and target pose parameters, direct methods are challenging to train and prone to overfitting on mappings seen during training, resulting in limited effectiveness and generalization capability on unseen mappings. Existing methods focus primarily on improvements of the network architecture and training strategies, with less attention given to mappings. In this work, we propose a geometric constraints learning approach, which enables networks to explicitly capture and utilize the geometric mappings between inputs and optimization targets for pose estimation. Specifically, we introduce a residual pose transformation formula that preserves pose transformation constraints within both the 2D image plane and the 3D space while decoupling the absolute pose distribution, thereby addressing the pose distribution gap issue. We further design a Geo6D mechanism based on the formula, which enables the network to explicitly utilize geometric constraints for pose estimation by reconstructing the inputs and outputs. We select two different methods as our baseline and extensive experiments show that Geo6D enhances the performance and reduces the dependence on extensive training data, remaining effective even with only 10% of the typical data volume.
Visual detection of micro aerial vehicles (MAVs) has received increasing research attention in recent years due to its importance in many applications. However, the existing approaches based on either appearance or motion features of MAVs still face challenges when the background is complex, the MAV target is small, or the computation resource is limited. In this paper, we propose a global-local MAV detector that can fuse both motion and appearance features for MAV detection under challenging conditions. This detector first searches MAV targets using a global detector and then switches to a local detector which works in an adaptive search region to enhance accuracy and efficiency. Additionally, a detector switcher is applied to coordinate the global and local detectors. A new dataset is created to train and verify the effectiveness of the proposed detector. This dataset contains more challenging scenarios that can occur in practice. Extensive experiments on three challenging datasets show that the proposed detector outperforms the state-of-the-art ones in terms of detection accuracy and computational efficiency. In particular, this detector can run with near real-time frame rate on NVIDIA Jetson NX Xavier, which demonstrates the usefulness of our approach for real-world applications. The dataset is available at https://github.com/WestlakeIntelligentRobotics/GLAD. In addition, A video summarizing this work is available at https://youtu.be/Tv473mAzHbU.
Visual detection of micro aerial vehicles (MAVs) is an important problem in many tasks such as vision-based swarming of MAVs. This article studies vision-based 6-D pose estimation to detect a 3-D bounding box of a target MAV, and then, estimate its 3-D position and 3-D attitude. The 3-D attitude information is critical to better estimate the target's velocity since the attitude and motion are dynamically coupled. In this article, we propose a novel 6-D pose estimation method, whose novelties are threefold. First, we propose a novel centroid point-guided keypoint localization network that outperforms the state-of-the-art methods in terms of both accuracy and efficiency. Second, while there are no publicly available real-world datasets for 6-D pose estimation for MAVs up to now, we propose a high-quality dataset based on an automatic dataset collection method. Third, since the dataset is collected in an indoor environment but detection tasks are usually in outdoor environments, we propose a self-training-based unsupervised domain adaption method to transfer the method from indoor to outdoor. Finally, we show that the estimated 6-D pose especially the 3-D attitude can significantly help improve the target's velocity estimation.
Uni6D is the first 6D pose estimation approach to employ a unified backbone network to extract features from both RGB and depth images. We discover that the principal reasons of Uni6D performance limitations are Instance-Outside and Instance-Inside noise. Uni6D's simple pipeline design inherently introduces Instance-Outside noise from background pixels in the receptive field, while ignoring Instance-Inside noise in the input depth data. In this paper, we propose a two-step denoising approach for dealing with the aforementioned noise in Uni6D. To reduce noise from non-instance regions, an instance segmentation network is utilized in the first step to crop and mask the instance. A lightweight depth denoising module is proposed in the second step to calibrate the depth feature before feeding it into the pose regression network. Extensive experiments show that our Uni6Dv2 reliably and robustly eliminates noise, outperforming Uni6D without sacrificing too much inference efficiency. It also reduces the need for annotated real data that requires costly labeling.
Visual anomaly detection plays a crucial role in not only manufacturing inspection to find defects of products during manufacturing processes, but also maintenance inspection to keep equipment in optimum working condition particularly outdoors. Due to the scarcity of the defective samples, unsupervised anomaly detection has attracted great attention in recent years. However, existing datasets for unsupervised anomaly detection are biased towards manufacturing inspection, not considering maintenance inspection which is usually conducted under outdoor uncontrolled environment such as varying camera viewpoints, messy background and degradation of object surface after long-term working. We focus on outdoor maintenance inspection and contribute a comprehensive Maintenance Inspection Anomaly Detection (MIAD) dataset which contains more than 100K high-resolution color images in various outdoor industrial scenarios. This dataset is generated by a 3D graphics software and covers both surface and logical anomalies with pixel-precise ground truth. Extensive evaluations of representative algorithms for unsupervised anomaly detection are conducted, and we expect MIAD and corresponding experimental results can inspire research community in outdoor unsupervised anomaly detection tasks. Worthwhile and related future work can be spawned from our new dataset.
Malicious use of micro aerial vehicles (MAVs) has become a serious threat to public safety and personal privacy in recent years. Motivated by this problem, we propose a systematic approach to monitor the intrusion of malicious MAVs based on a novel type of panoramic stereo camera networks. Each sensing node of such a network consists of 16 lenses that can form a 360-degree panoramic vision system. The 16 lenses further form 8 pairs of stereo cameras that can directly localize aerial targets. The effective range for a sensing node localizing a MAV like DJI M300 could reach 80 meters, which is much farther than existing commercial stereo cameras. In terms of algorithms, we propose i) a novel visual MAV detection algorithm based primarily on motion features of MAVs, ii) an efficient stereo localization algorithm based on sparse feature points, and iii) robust multi-target tracking and trajectory fusion algorithm to fuse the observations of different sensing nodes. The effectiveness, robustness, and accuracy of the proposed algorithms together with the overall system have been verified by extensive experimental tests. To the best of our knowledge, this is the first systematic approach to detect, localize, and track unknown MAVs in the literature. Our approach provides a scalable solution to securely cover large areas of interest against malicious MAV intrusion. Note to Practitioners—Micro aerial vehicles (MAVs) have been widely used in many domains nowadays. However, they have also brought many safety problems. To monitor the intrusion of malicious MAVs, this paper proposes a novel type of panoramic stereo camera networks that can detect, localize, and track multiple MAVs simultaneously. Such a network consists of a number of sensing nodes and a central node. Each sensing node is able to detect, localize, and track multiple MAV targets. The role of the central node is to fuse the observations from multiple sensing nodes to generate more accurate trajectories of the MAV targets and in the meantime secure a large area in a coordinated way. This paper presents the details of the prototype of the system and the key algorithms therein.
Electrocardiogram(ECG) is commonly utilized in clinical diagnosis and health monitoring. However, ECG acquisition is cumbersome because ECG electrodes must be attached to the skin tightly. Ballistocardiogram(BCG), originating from the heartbeat, can be captured without attaching the BCG sensors to the skin, making BCG collecting convenient and non-feeling. However, BCG diagnostic experience is limited compared with ECG. In this paper, we propose a model, called Rec-AUNet, to reconstruct ECG from BCG so that we can combine the convenient measurement of BCG and diagnostic experience on ECG. Rec-AUNet utilizes an attentive UNet-based neural network with encoding paths to better capture the temporal features and with attentive paths to preserve the spatial features. To better evaluate the coherence of the global waveform and the fidelity of the local physiological features between the synthetic ECG and the original ECG, we deliberately design the person identification task as the semantic metric. The synthetic ECG from our proposed model achieves an accuracy of 93.75% and 94.05% in person identification on the Kansas-Dataset and selfcollected data, respectively, outperforming existing algorithms.
Numerous 6D pose estimation methods have been proposed that employ end-to-end regression to directly estimate the target pose parameters. Since the visible features of objects are implicitly influenced by their poses, the network allows inferring the pose by analyzing the differences in features in the visible region. However, due to the unpredictable and unrestricted range of pose variations, the implicitly learned visible feature-pose constraints are insufficiently covered by the training samples, making the network vulnerable to unseen object poses. To tackle these challenges, we proposed a novel geometric constraints learning approach called Geo6D for direct regression 6D pose estimation methods. It introduces a pose transformation formula expressed in relative offset representation, which is leveraged as geometric constraints to reconstruct the input and output targets of the network. These reconstructed data enable the network to estimate the pose based on explicit geometric constraints and relative offset representation mitigates the issue of the pose distribution gap. Extensive experimental results show that when equipped with Geo6D, the direct 6D methods achieve state-of-the-art performance on multiple datasets and demonstrate significant effectiveness, even with only 10% amount of data.
This article introduces the solutions of the “MicroalgaeDetector” team for the IEEE UV 2022 Vision Meets Algae Object Detection Challenge. This challenge focus on developing computer vision detection algorithm to automatically detect marine microalgae from microscopy images. Automatic localization and identification of microalgae are anticipated to be accomplished concurrently during image analysis, which will simplify downstream cell analysis and lay the groundwork for algae identification using image data in conjunction with biomorphological traits. In this competition, we observe that the training dataset has a serious class imbalance problem, and some classes are in a state of few samples, which greatly limits the performance of both single stage detectors and multi-stage detectors. There are also issues with tiny objects in high-resolution images and serious bounding box annotation inconsistencies. To address the aforementioned competition challenges of few samples, unbalanced categories, noisy annotations and small objects in this competition, we propose a robust and high-performance algae detection method (RAD), which can precisely localize and identify marine microalgae in microscopy images. In the proposed RAD, we develop a class-specific copy-paste strategy to achieve instance-level re-sampling, which resolves the problem of the data imbalance. We also introduce several training/inference strategies and a bag of tricks that brings more or less performance boost. In order to increase robustness, we also train multiple expert models to ensemble them. Our RAD wins the competition after achieving 58.192% mAP in the test dataset.
In RGB-D based 6D pose estimation, direct regression approaches can directly predict the 3D rotation and translation from RGB-D data, allowing for quick deployment and efficient inference. However, directly regressing the absolute translation of the pose suffers from diverse object translation distribution between the training and testing datasets, which is usually caused by the diversity of pose distribution of objects in 3D physical space. To this end, we generalize the pin-hole camera projection model to a residual-based projection model and propose the projective residual regression (Res6D) mechanism. Given a reference point for each object in an RGB-D image, Res6D not only reduces the distribution gap and shrinks the regression target to a small range by regressing the residual between the target and the reference point, but also aligns its output residual and its input to follow the projection equation between the 2D plane and 3D space. By plugging Res6D into the latest direct regression methods, we achieve state-of-the-art overall results on datasets including Occlusion LineMOD (ADD(S): 79.7%), LineMOD (ADD(S): 99.5%), and YCB-Video datasets (AUC of ADD(S): 95.4%).
Deep learning-based methods have recently shown great promise in the defect detection task. However, current methods rely on large-scale annotated data and are unable to adapt a trained deep learning model to new samples that were not observed during training. To address this issue, we propose a new siamese defect-aware attention network (SDANet) with a template comparison detection strategy that improves the defect detection technique for matching new samples without rapidly collecting new data and retraining the model. In SDANet, the siamese feature pyramid network is used to extract multi-scale features from input and template images, the defect-aware attention module is proposed to obtain inconsistency between input and template features and use it to enhance abnormality in input image features, and the self-calibration module is developed to calibrate the alignment error between the input and template features. SDANet can be used as a plug-in module to enable most existing mainstream detection algorithms to detect defects using not only the features of defects, but also the inconsistency between features of the inspected image and the template image. Extensive experiments on two publicly available industrial defect detection benchmarks highlight the effectiveness of our method. SDANet can be seamlessly integrated into mainstream detection methods and improve the mAP of mainstream detection algorithms on unseen samples by 12% on average which outperforms current state-of-the-art method by 7.7%. It can also improve the performance in seen samples by 4.3% on average. SDANet can be used in general defect detection applications of industrial manufacturing.
Unsupervised anomaly detection and localization, as of one the most practical and challenging problems in computer vision, has received great attention in recent years. From the time the MVTec AD dataset was proposed to the present, new research methods that are constantly being proposed push its precision to saturation. It is the time to conduct a comprehensive comparison of existing methods to inspire further research. This paper extensively compares 13 papers in terms of the performance in unsupervised anomaly detection and localization tasks, and adds a comparison of inference efficiency previously ignored by the community. Meanwhile, analysis of the MVTec AD dataset are also given, especially the label ambiguity that affects the model fails to achieve full marks. Moreover, considering the proposal of the new MVTec 3D-AD dataset, this paper also conducts experiments using the existing state-of-the-art 2D methods on this new dataset, and reports the corresponding results with analysis.
As RGB-D sensors become more affordable, using RGB-D images to obtain high-accuracy 6D pose estimation results becomes a better option. State-of-the-art approaches typically use different backbones to extract features for RGB and depth images. They use a 2D CNN for RGB images and a per-pixel point cloud network for depth data, as well as a fusion network for feature fusion. We find that the essential reason for using two independent backbones is the "projection breakdown" problem. In the depth image plane, the projected 3D structure of the physical world is preserved by the 1D depth value and its built-in 2D pixel coordinate (UV). Any spatial transformation that modifies UV, such as resize, flip, crop, or pooling operations in the CNN pipeline, breaks the binding between the pixel value and UV coordinate. As a consequence, the 3D structure is no longer preserved by a modified depth image or feature. To address this issue, we propose a simple yet effective method denoted as Uni6D that explicitly takes the extra UV data along with RGB-D images as input. Our method has a Unified CNN framework for 6D pose estimation with a single CNN backbone. In particular, the architecture of our method is based on Mask R-CNN with two extra heads, one named RT head for directly predicting 6D pose and the other named abc head for guiding the network to map the visible points to their coordinates in the 3D model as an auxiliary module. This end-to-end approach balances simplicity and accuracy, achieving comparable accuracy with state of the arts and 7.2x faster inference speed on the YCB-Video dataset.
The essence of unsupervised anomaly detection is to learn the compact distribution of normal samples and detect outliers as anomalies in testing. Meanwhile, the anomalies in real-world are usually subtle and fine-grained in a high-resolution image especially for industrial applications. Towards this end, we propose a novel framework for unsupervised anomaly detection and localization. Our method aims at learning dense and compact distribution from normal images with a coarse-to-fine alignment process. The coarse alignment stage standardizes the pixel-wise position of objects in both image and feature levels. The fine alignment stage then densely maximizes the similarity of features among all corresponding locations in a batch. To facilitate the learning with only normal images, we propose a new pretext task called non-contrastive learning for the fine alignment stage. Non-contrastive learning extracts robust and discriminating normal image representations without making assumptions on abnormal samples, and it thus empowers our model to generalize to various anomalous scenarios. Extensive experiments on two typical industrial datasets of MVTec AD and BenTech AD demonstrate that our framework is effective in detecting various real-world defects and achieves a new state-of-the-art in industrial unsupervised anomaly detection.
Unsupervised anomaly detection and localization is crucial to the practical application when collecting and labeling sufficient anomaly data is infeasible. Most existing representation-based approaches extract normal image features with a deep convolutional neural network and characterize the corresponding distribution through non-parametric distribution estimation methods. The anomaly score is calculated by measuring the distance between the feature of the test image and the estimated distribution. However, current methods can not effectively map image features to a tractable base distribution and ignore the relationship between local and global features which are important to identify anomalies. To this end, we propose FastFlow implemented with 2D normalizing flows and use it as the probability distribution estimator. Our FastFlow can be used as a plug-in module with arbitrary deep feature extractors such as ResNet and vision transformer for unsupervised anomaly detection and localization. In training phase, FastFlow learns to transform the input visual feature into a tractable distribution and obtains the likelihood to recognize anomalies in inference phase. Extensive experimental results on the MVTec AD dataset show that FastFlow surpasses previous state-of-the-art methods in terms of accuracy and inference efficiency with various backbone networks. Our approach achieves 99.4% AUC in anomaly detection with high inference efficiency.
On account of a large scale of dataset need to be annotated to train the deep learning based modern object detection model, zero-shot object detection has become an important research field which aims to simultaneously localize and recognize unseen objects that are not observed during training. In order to improve the performance of zero-shot object detection, recent state of the art methods tend to make complicated modifications to the modern object detectors in terms of the model structure, loss function and training process. They always take the simple modification as a baseline, and think it is worse than more complicated methods. In contrast, we find that simple modification can achieve better performance. Considering that the redundant modification may increase the risk of over-fitting in seen classes and reduce generalization performance on unseen classes, we propose a visual language based succinct zero-shot object detection framework, which only replaces the classification branch in the modern object detector with a lightweight visuallanguage network. Since zero-shot object detection is a classic multi-modal learning protocol which consists of a visual feature space and a language space, our visual-language network learns the visual language alignment from the image and language data of seen classes and transfers this alignment to detect unseen objects. Following the Occam's razor principle that "Entities should not be multiplied unnecessarily", extensive experimental results show that our succinct framework can suppress all existing zero-shot object detection methods on several benchmarks and gets the new state-of-the-art.
Using computed tomography (CT) images to assist pseudomyxoma peritonei (PMP) diagnosis is noninvasive and fast compared with puncturing detection. However, it is time and energy-demanding to detect and annotate lesions on CT scans for radiologists. Thus, the automatic segmentation of PMP lesions is of great potential to reduce the burden on radiologists and improve PMP diagnostic efficiency. This paper proposed a Dual External Contextual Attention Network (DECANet) to segment PMP lesions automatically. Our network is derived from ResUNet, and we design a module named dual external contextual attention to extract high-level features to improve PMP lesion segmentation accuracy. The PMP segmentation dataset are collected from 38 patients and annotated by experienced radiologists. The proposed network achieves good performance with a dice coefficient of 88.68% and a mean Intersection over Union (mIoU) of 79.40%, outperforming other networks including UNet, AttentionUNet, and ResUNet.