With the growing demand for electricity, live-line working for distribution network has gained increasing attention worldwide to ensure uninterrupted power supply. Live-line working robots have significantly enhanced the efficiency of maintance and ensured the safety of operators, leading to their growing application. In this paper, we propose a dual-arm collaborative live-line working robot system for the distribution network, with design considerations for both hardware and software. On the hardware level, the robot is primarily composed of two UR5e robot arms and is equipped with various visual sensors to get information about work environment. We use the Touch haptic device of 3D Systems to teleoperate the UR5e robot arm. On the software level, we propose a master-slave heterogeneous teleoperation method, where the position and orientation of UR5e end-effector are respectively controlled through incremental mapping and one-to-one mapping. This system not only enhances operators' safety but also improves the sense of immersion and operational accuracy, demonstrating significant practical value and broad application prospects.
Existing deepfake detection methods heavily rely on static labeled datasets. However, with the proliferation of generative models, real-world scenarios are flooded with massive amounts of unlabeled fake face data from unknown sources. This presents a critical dilemma: detectors relying solely on existing data face generalization failure, while manual labeling for this new stream is infeasible due to the high realism of fakes. A more fundamental challenge is that, unlike typical unsupervised learning tasks where categories are clearly defined, real and fake faces share the same semantics, which leads to a decline in the performance of traditional unsupervised strategies. Therefore, there is an urgent need for a new paradigm designed specifically for this scenario to effectively utilize these unlabeled data. Accordingly, this paper proposes a dual-path guided network (DPGNet) to address two key challenges: (1) bridging the domain differences between faces generated by different generative models; and (2) utilizing unlabeled image samples. The method comprises two core modules: text-guided cross-domain alignment, which uses learnable cues to unify visual and textual embeddings into a domain-invariant feature space; and curriculum-driven pseudo-label generation, which dynamically utilizes unlabeled samples. Extensive experiments on multiple mainstream datasets show that DPGNet significantly outperforms existing techniques,, highlighting its effectiveness in addressing the challenges posed by the deepfakes using unlabeled data.
Achieving robust locomotion on complex terrains remains a challenge due to high dimensional control and environmental uncertainties. This paper introduces a teacher prior framework based on the teacher student paradigm, integrating imitation and auxiliary task learning to improve learning efficiency and generalization. Unlike traditional paradigms that strongly rely on encoder-based state embeddings, our framework decouples the network design, simplifying the policy network and deployment. A high performance teacher policy is first trained using privileged information to acquire generalizable motion skills. The teacher's motion distribution is transferred to the student policy, which relies only on noisy proprioceptive data, via a generative adversarial mechanism to mitigate performance degradation caused by distributional shifts. Additionally, auxiliary task learning enhances the student policy's feature representation, speeding up convergence and improving adaptability to varying terrains. The framework is validated on a humanoid robot, showing a great improvement in locomotion stability on dynamic terrains and significant reductions in development costs. This work provides a practical solution for deploying robust locomotion strategies in humanoid robots.
LiDAR-based loop closure detection is a crucial part of realizing robust SLAM algorithms for intelligent vehicles with LiDAR sensors. Existing methods often reduce the keypoint dimension to encode the global descriptor, which sacrifices the freedom of loop detection and correction. Based on the 6-DOF rigid transformation property of spatial triangles, we propose an algorithm for extracting and describing 3D keypoints from high-resolution spinning LiDAR intensity images to encode triangle descriptors, termed intensity triangle descriptor (ITD) . In comparison to the direct extraction of keypoints from the point cloud, the use of image-derived feature points provides additional photometric texture information and better handles uneven spatial density of the point cloud, which is advantageous in unstructured and geometrically degraded scenes. To enhance the stability of keypoints, the spatial positions of multi-frame image feature points are registered to a keyframe by an odometer for voxel downsampling and non-maximum suppression, with the objective of reducing unstable feature points. For high discrimination, the neighbor image patches of each vertex (keypoint) are aggregated to estimate a Gaussian mixture model (GMM) as the keypoint signature. An efficient two-stage loop closure detection method is then proposed for ITD, consisting of candidate retrieval based on triangle side lengths and vertex GMMs, followed by geometric verification of matched descriptor pairs. The effectiveness of the proposed method is evaluated on the STheReO, FusionPortable, and our self-collected datasets.
Hybrid robots, which combine crawling and flying capabilities, have gained significant attention in recent years due to their potential in powerline inspection. This paper provides insights into the design of hybrid robots, offering a comprehensive analysis of the system’s design requirements and its solutions. Firstly, we present a detailed overview of the hybrid robot’s system architecture, encompassing the design aspects of its mechanical, hardware, and software components. Secondly, we analyze the dynamics of the robot in crawling mode, deducing the conditions necessary for maintaining balance on powerlines and traversing inclined powerlines. Thirdly, we investigated the kinematics and dynamics of the robot in flight mode, examining how the structural design of the hybrid robot specifically impacts its system dynamics and control responsiveness. Finally, we verify our analyses through a series of simulations and real-world tests, confirming the effectiveness and feasibility of the design strategies presented.
Pose estimation of train couplers is a crucial task for the train uncoupling robot. Existing pose estimation methods are commonly employed for extracting convex-shaped objects; however, their applicability is limited when dealing with the slender tubular geometry characteristic of train couplers. To address this issue, we propose a novel method for estimating the pose of train couplers directly from RGB images. Our method utilizes a robust backbone composed of pure convolution modules to effectively capture essential geometric features from monocular images. We also incorporate dynamic snake convolution into the network to specifically focus on the slender tubular structural feature of the train coupler. Finally, a novel adaptive weighted pose loss function is proposed to further improve the precision of the train coupler pose. Experimental results on our custom coupler dataset and the Linemod dataset demonstrate significant improvements achieved by our proposed method.
This letter presents the first trajectory planning method for hybrid robot to perform powerline inspection involving obstacle navigation and landing. We develop a geometric model that incorporates constraints for landing the hybrid robot on a powerline, obstacle avoidance, and objectives that maximize the visibility of the powerline during flight. The trajectory generation is achieved via solving a multiple shooting nonlinear programming problem with respect to system dynamics and geometric constraints. The formulation of the problem accommodates both powerline-to-powerline and air-to-powerline trajectory planning scenarios. It runs onboard and is capable of generating trajectories within 50 ms, regardless of whether the hybrid robot's initial state is positioned on the powerline or hovering above it. Through simulation experiments, we illustrate the impact of our proposed geometric model on trajectory planning. Furthermore, real-world experimental results validate the efficacy of the proposed planning method. Compared with the existing feedback-control-based work, the landing and obstacle navigation time are significantly reduced.
Due to the uneven depth scale in the linear disparity space, calculating stereo disparity and subsequently converting it into depth, although widely used, leads to nonlinear error amplification. Moreover, the limited availability of depth annotations in stereo datasets has impeded the progress of end-to-end depth estimation techniques. This paper introduces a semi-supervised method for depth estimation using a multitask network. The multi-task network comprises two branches for disparity estimation and depth estimation. During training, it leverages a pre-trained stereo disparity network to provide dense depth self-supervision, expediting the training of the depth branch. This network offers an effective solution to the scarcity of stereo depth datasets and the sparsity of depth information in point cloud annotations. The efficacy of the algorithm is validated on the KITTI 3D object dataset using sparse point clouds as depth annotations, showcasing remarkable depth estimation capabilities. Additionally, the paper transforms obtained depth maps into pseudo-lidar for 3D object detection, achieving promising results on the KITTI dataset.
Operating in narrow spaces is an important challenge in the development of robots. Redundant manipulators are one way to solve this problem, but their mechanism design and control method still have much room for improvement. In this paper, we propose a coiled cable-conduit-driven hyper-redundant manipulator (C-CDHRM) with great slenderness and flexibility. In terms of mechanism design, it considers both compactness and operability. By imitating the structure and behavior of a constricting snake, it can be uncoiled sequentially from a coiled storage state, led by the head. In terms of control methods, we propose a multi-layer control system that can make remote operations more accurate and reliable. On the one hand, guiding, segmenting, and following the path overcome the planning ambiguity caused by redundancy. On the other hand, conduit transmission modeling and cable length correction overcome the nonlinear mapping of cable-driven joints and were verified in experiments. Through tests, the mobile integrated system composed of C-CDHRM has an excellent performance in operation precision and accuracy, ensuring safety and accessibility in narrow spaces. Finally, in field experiments, the inspection and cleaning of various types of electrical equipment have been successfully completed, showing excellent application prospects.
In recent years, various types of inspection robots have been developed to automate powerline inspection. The hybrid robot combines the advantages of climbing and flying robots and has a promising prospect in powerline inspection. But landing a hybrid robot on the target powerline among multiple ones is challenging. Flights require robust detection of powerlines and stable tracking of the target powerline. We propose a complete solution for the autonomous landing of a hybrid robot on a powerline. First, a special feature extraction operator and the corresponding density-based feature recognition algorithm are designed to detect multiscale powerlines. Second, a binocular vision-based depth estimation method for the landing point in the powerline is described. Third, two spatio-temporal dictionaries are established to track the target one in multiple powerlines. Meanwhile, landing strategies and control methods are presented to achieve a stable landing task. Finally, a hybrid robot is designed to validate the proposed method. The experiment results demonstrate the accuracy of the powerline detection and depth estimation algorithm, as well as the effectiveness of the robot in tracking and landing tasks.
The detection of pests plays a crucial role in intelligent early warning systems of injurious insects and diseases in precision agriculture. However, pests strong concealment and mobility pose significant challenges to their timely detection. In this paper, we propose a novel approach called Multi-scale Dense YOLO (MD-YOLO) for detecting three typical small target lepidopteran pests on sticky insect boards. In MD-YOLO, we design three key components: the image feature extraction part, the feature fusion network, and the prediction module. To enhance the utilization of feature maps and mitigate information loss, we incorporate DenseNet blocks and an adaptive attention module (AAM) into the feature extraction part. The AAM helps capture relevant image details and improves the model’s ability to exploit feature representations effectively. For effective feature integration, our feature fusion network incorporates both a feature extraction path and a feature aggregation path. This enables the deep network to leverage spatial location information from the shallower network, thereby enhancing the detection accuracy. Experimental results demonstrate the effectiveness of MD-YOLO, with detection results achieving an mAP@.5 value of 86.2%, an F1 score of 79.1%, and an IoU value of 88.1%. We conduct extensive experiments to compare MD-YOLO with state-of-the-art models, and the results showcase its superiority. Furthermore, we design an Internet of Things (IoT) system that demonstrates MD-YOLO’s performance in real-world field scenes, highlighting its practical applicability.
In this letter, we propose RI-LIO, a new reflectivity image assisted tightly-coupled LiDAR-inertial odometry (LIO) framework that introduces additional reflectivity texture information to efficiently reduce the drift of geometric-only methods. To achieve this, we construct an iterated extended Kalman filter framework by blending the point-to-plane geometric measurement and the reflectivity image measurement. Specifically, the geometric measurement is defined as the distance from the raw point of a new scan to its nearest neighbor plane in the global incremental kd-tree map. The searched nearest neighbor point is used to render a sparse reflectivity image after LiDAR motion distortion information is given by its corresponding raw point. Then, the reflectivity measurement is built to align the sparse reflectivity image with the dense reflectivity image of the current scan by minimizing the photometric errors directly. In addition, based on the mechanism of high-resolution LiDAR, a corrected spherical projection model is proposed to project spatial points into the image frame. Finally, extensive experiments are conducted in structured, unstructured and challenging open field scenarios. The results demonstrate that the proposed method outperforms existing geometric-only methods in terms of robustness and accuracy, especially in the rotation direction.
Scoliosis is a common disease of the spine and requires regular monitoring due to its progressive properties. A preferred indicator to assess scoliosis is by the Cobb angle, which is currently measured either manually by the relevant medical staff or semi-automatically, aided by a computer. These methods are not only labor-intensive but also vary in precision by the inter-observer and intra-observer. Therefore, a reliable and convenient method is urgently needed. With the development of computer vision and deep learning, it is possible to automatically calculate the Cobb angles by processing X-ray or CT/MR/US images. In this paper, the research progress of Cobb angle measurement in recent years is reviewed from the perspectives of computer vision and deep learning. By comparing the measurement effects of typical methods, their advantages and disadvantages are analyzed. Finally, the key issues and their development trends are also discussed.
In order to solve the vibration problem of low-damping flexible mechanism in motion, a cascade vibration suppression (CVS) trajectory planning method combining the advantages of the optimal double S-curve trajectory planning method and the input shaping method is proposed. The CVS method can effectively suppress the vibration in the period of uniform motion. The correctness of this method is proved by theoretical analysis and digital simulation. Finally, the application results of the proposed method in the FAST feed support system show that it can effectively improve the motion accuracy of the parallel cable mechanism, especially when the motion speed is constant, improving the position tracking accuracy by over 20%.
To achieve large-scale and high bandwidth environmental monitoring, this paper proposes a hybrid network structure based on the ZigBee network and Mesh network, which consists of the backbone network and nodes, branch network and nodes, and monitoring center. Both the backbone network node and branch network node with sensor information collection and wireless communication functions are designed. And a monitoring center is developed to display environmental information in real time. With the hybrid-mode network, the environment monitoring experiments were carried out to verify the feasibility of the system. The results show that compared to traditional environmental monitoring systems, the proposed system has the advantage of strong networking ability, large bandwidth, long transmission distance, and low cost.
A hybrid transmission line inspection robot is a combination of a multirotor and a power line landing mechanism, which can perform both airborne and on-wire inspections. An improved design of a hybrid automatic inspection system is demonstrated in this paper, planned to land on a shield line. Installed with a lightweight landing mechanism consists of a belt system drive linear motion mechanism and a gear rack drive wire clamper, the robot can perform both airborne and on-wire inspection to powerlines. The finite element analysis on key components was conducted. In the field, the robot was found to have an acceptable flight performance.
The failure of an insulator may compromise the safety of the entire power transmission system. Therefore, insulator defect detection is vital for the safe operation of power systems. However, insulator defects in an insulator image may have varying sizes, and several currently available methods do not have satisfactory detection accuracy for small defects. To address this issue, we propose an improved detection network for small insulator defects with a batch normalization convolutional block attention module (BN-CBAM) and a feature fusion module. The BN-CBAM is designed to better exploit channel information and enhance the effect of different channels on the feature map. In addition, we propose a feature fusion module that fuses multi-scale features from different layers to improve small object detection performance. Moreover, to address the scarcity of aerial images, a data augmentation method based on the fusion of the target segment and background is introduced. Experiments demonstrate that the proposed method achieves better small insulator defect detection performance than other state-of-the-art approaches. In addition, data augmentation methods enrich sample diversity and enhance the generalizability of the network.
The application of traditional 3D reconstruction methods such as structure-from-motion and simultaneous localization and mapping are typically limited by illumination conditions, surface textures, and wide baseline viewpoints in the field of robotics. To solve this problem, many researchers have applied learning-based methods with convolutional neural network architectures. However, simply utilizing convolutional neural networks without taking other measures into account is computationally intensive, and the results are not satisfying. In this study, to obtain the most informative images for reconstruction, we introduce a residual block to a 2D encoder for improved feature extraction, and propose an attentive latent unit that makes it possible to select the most informative image being fed into the network rather than choosing one at random. The recurrent visual attentive network is injected into the auto-encoder network using reinforcement learning. The recurrent visual attentive network pays more attention to useful images, and the agent will quickly predict the 3D volume. This model is evaluated based on both single- and multi-view reconstructions. The experiment results show that the recurrent visual attentive network increases prediction performance in a way that is superior to other alternative methods, and our model has desirable capacity for generalization.
Image-based segmentation of overhead power lines is critical for power line inspection. Real-time segmentation helps the inspection robot avoid obstacles or land on the wire during the inspection task. It is challenging for several studies to achieve real-time overhead power line segmentation with high accuracy. In addition, cluttered background brings great difficulties to overhead power lines segmentation. To address these issues, an efficient parallel branch network for real-time overhead power line segmentation is proposed. Our framework combines a context branch that generates useful global information with a spatial branch that preserves high-resolution segmentation details. The asymmetric factorized depth-wise bottleneck (AFDB) module is designed in the context branch to achieve more efficient short-range feature extraction and provide a large receptive field. Furthermore, the subnetwork-level skip connections in the classifier are proposed to fuse long-range features and lead to high accuracy. Experiments demonstrate that our framework achieves more than 90% segmentation accuracy.
随着无人机应用领域的不断扩大,无人机对周围环境感知的需求逐渐加大,激光雷达技术的应用成为无人机研究领域的重要发展趋势.为解决无人机在飞行过程中周围环境复杂导致现场定位信号较差的问题,基于三维激光雷达,结合大疆无人机的Onboard SDK技术,设计机载激光点云采集系统以及激光点云处理算法.为了降低激光采集系统的操作难度,设计避障算法,在飞行既定路线的同时结合激光点云进行避障飞行.利用激光点云SLAM算法,实现了无人机对周围环境的三维建模,以达到无人机对周围环境感知的目的,并通过实际测试验证了所提方法的有效性.