Real-time semantic segmentation is a prerequisite for autonomous vehicles. Despite significant advancements, attaining an optimal balance between segmentation accuracy and efficiency remains a challenge for existing algorithms. To this end, we propose an orthogonal dual-path network, terms ODPNet, for real-time semantic segmentation of urban road scenes. First, the network is decoupled into two orthogonal paths, namely the Semantic Path and the Refinement Path, to encode high-level semantics and low-level details, respectively. Subsequently, we introduce an Atrous Decoupled Pyramid Pooling Module (ADPPM) to overcome the receptive field rigidity of traditional square pooling. By synergizing asymmetric pooling with atrous convolutions, the ADPPM effectively aggregates multi-scale contextual information and simultaneously enhances the modeling of long-range dependencies along orthogonal spatial dimensions. Furthermore, we design a Multi-View Aggregation Module (MVAM) to bridge the semantic-spatial gap through a trans-dimensional attention mechanism, thereby enhancing feature discriminability. Additionally, we leverage boundary supervision via the Canny operator to explicitly guide the preservation of fine-grained structural details to better perform orthogonalization of the dual-path. Extensive experiments on the Cityscapes and CamVid datasets demonstrate the effectiveness of ODPNet, showcasing a superior trade-off between inference speed and segmentation accuracy. Specifically, on the Cityscapes dataset, ODPNet-Lite achieves 78.9% mIoU at 162.1 FPS, while ODPNet-Base and ODPNet-Deep achieve 79.9% mIoU at 64.2 FPS and 80.4% mIoU at 48.2 FPS, respectively.
With the growing demand for rotating machinery health monitoring, single-sensor fault diagnosis shows clear limitations in noise immunity, information representation and adaptability to complex operating conditions. Taking rolling bearings as a typical example, this paper reviews multi-source data fusion methods for equipment fault diagnosis and prediction from the perspectives of data-level, feature-level, decision-level and multi-level fusion. Existing studies show that data-level fusion preserves raw information, feature-level fusion balances representation and efficiency, decision-level fusion improves flexibility and robustness, and multi-level fusion enhances diagnosis through cross-layer collaboration.
Semantic segmentation is a key technology for autonomous vehicles to understand the surrounding scenes. Multi-branch network architectures have demonstrated their efficiency and effectiveness in real-time semantic segmentation tasks. Although PIDNet achieves a balance between performance and efficiency, it is inadequate in fine-grained segmentation and multi-scale feature integration. In this paper, we propose a real-time semantic segmentation network, called ARFNet, which employs a three-branch structure. To enhance the perception and fusion of multi-scale features, we introduce a Hierarchical Dense Atrous Pyramid Module (HDAPM). Additionally, we propose a novel Trans-Dimensional Interaction Module (TDIM) to systematically enhance feature representation through cross-dimensional attention mechanisms. The effectiveness of our method is demonstrated by extensive experiments on the Cityscapes and CamVid datasets, showing that it achieves a promising trade-off between inference speed and segmentation accuracy. Specifically, ARFNet achieves 79.8% mIoU at 94.8 FPS on the Cityscapes dataset and 80.9% mIoU at 154.1 FPS on the CamVid dataset.
This paper presents a review-oriented comparative analysis of deep reinforcement learning (DRL) for intelligent control of dual-arm robots. Instead of focusing on a single control algorithm, it organizes recent studies into an algorithm-task-metric framework and extracts quantitative evidence from representative applications including cooperative grasping, assembly, transportation, obstacle-aware planning, contact-rich control, and sim-to-real transfer. PPO, MAPPO, MADDPG, and SAC are compared in terms of success rate, convergence behavior, trajectory smoothness, force regulation, safety constraints, and transferability. Key design factors and future trends, including reward design, multimodal perception, safe reinforcement learning, sample efficiency, and real-robot deployment, are summarized to provide practical guidance for dual-arm intelligent cooperative control.
At present, the development prospect of robots is broad, and the types of robots are increasing day by day. Facing the scenario of multiple robots providing services, the problem of optimal service scheduling becomes more and more prominent. Optimal service scheduling for the cloud robot, this paper based on SOA architecture of cloud, the service type, the latitude and longitude of the ratings and service as the index, USES the circular area search algorithm(CAS),to determine the optimal service decisions, when certain cloud robot is scheduling, cloud service robots to applicants need planning between indoor and outdoor path, The outdoor path planning of cloud robot is carried out by amap navigation, and the indoor path planning is carried out by extended A * algorithm to realize the indoor and outdoor service scheduling of cloud robot. The experimental results show that the optimal service decision and indoor and outdoor scheduling task under CAS algorithm are realized by integrating the longitude and latitude, service score and service type of cloud robot when the service applicant applies for the cloud robot, which provides reference for the motion control of cloud robot.
Real-time semantic segmentation has broad prospects in computer vision. Existing state-of-the-art approaches generally employ bilateral networks to encode spatial and contextual information. Nevertheless, the real-time methods exhibit unsatisfactory performance. To address this problem, in this work we propose a novel shared trunk and dual-branch network named STDBNet for real-time semantic segmentation. In particular, we devise a Shared Trunk Module, which efficiently diminishes superfluous channels and parameters. Subsequently, we present a Split Dual-Branch Module, consisting of Detail Branch and Semantic Branch, with the former capturing detailed information and the latter capturing contextual information correspondingly. Furthermore, at the Semantic Branch, an Efficient Pyramid Pooling Module is designed towards the end to further expand the receptive fields and harvest multi-scale contextual information. Finally, to efficiently merge the features from the two branches and then get the final results, we introduce an Attention-optimized Feature Fusion Module, which utilizes redesigned spatial attention mechanism for feature augmentation. Extensive experiments conducted on the Cityscapes dataset indicate that the devised STDBNet obtains 77.6% mIoU with 91.5 FPS on one GeForce 2080Ti GPU, which outperforms BiSeNet with a favorable trade-off between accuracy and efficiency.
This paper presents a cost-effective and efficient closed-loop calibration method to enhance the absolute positioning accuracy of industrial robots. The method utilizes an economical and highly accurate set of simple measurement devices, including a device mounted on the robot's end effector and a spherical constraint device positioned within the robot's workspace. A novel local area measurement method is employed for the spherical constraint device, effectively converting position errors into the errors in the sphere center position of the constraint device. Furthermore, the error model considers both kinematic parameter deviations of the robot itself as well as those of the measurement device, along with non-kinematic errors caused by link self-weight. The closed-loop calibration model relies on position constraints, enabling identification and compensation of all error parameters using Levenberg Marquardt method after substituting measured data. The effectiveness of this proposed method is demonstrated through simulation and experimentation. In simulations, average position error reduces from 8.249 mm to 0.8384 mm, while absolute mean difference in sphere center positions obtained using our measurement equipment decreases from 1.206 mm to 0.1027 mm. Significant reductions in both position error and difference in sphere center position are indicated after calibration. This proves that the absolute mean value of the sphere center position difference can be used as the basis for judging the size of the robot’s end position error. Experimental results show that absolute mean difference in sphere center position decreases from 0.4319 mm to 0.02869 mm, leading to substantial improvement in overall positioning accuracy.
With the advent of the 5G era, the development of IoT technology has been accelerated. Due to the continuous increase in the amount of data waiting to be processed from the edge, edge nodes may struggle to handle such a vast amount of data. Therefore, the technology of cloud-edge collaboration has emerged, and how to achieve cloud-edge collaborative task scheduling has become a current research hotspot. This article provides a detailed exposition of the relevant work on task scheduling in the cloud-edge environment, and outlines the common optimization objectives in the cloud-edge collaboration scenario. The methods used to solve task scheduling problems are classified and summarized, including heuristic, heuristic algorithm based on linear programming, and meta-heuristic algorithms. The advantages and disadvantages of each algorithm are analyzed. Finally, the development trends of large-scale task scheduling in the cloud-edge environment are discussed, providing valuable insights for achieving real-time performance, efficiency, and energy conservation in the Internet of Things.
The animal husbandry industry is undergoing a transition towards intelligent breeding with advancements in artificial intelligence technology. Ensuring the health and safety of livestock and poultry requires automated assessment of their condition, with object detection being the primary focus. However, the existing object detection methods demonstrate subpar accuracy when applied to caged chicken detection. Moreover, their deployment on embedded devices is hindered by the significant size of the network models. To address these issues, an Improved -You Only Look Once version 5 (YOLOv5s) network detection model is proposed, which is based on YOLOv5s object detection algorithms. The cross -stage partial structure of the neck is modified to enhance the residual structure. Additionally, a convolutional block attention module is incorporated to extract crucial features. The distance intersection over union non -maximum suppression algorithm is employed to improve the identification of overlapping chickens. Furthermore, depthwise separable convolution is implemented to reduce network complexity. Experimental results on the custom dataset demonstrate that the Improved-YOLOv5s model achieves a mean average precision of 98.28%, surpassing the original model by 5.33%. It also exhibits strong performance across all other evaluation metrics while significantly reducing network parameters. In comparison to SSD, Faster-RCNN, YOLOv3, YOLOv4, YOLOv5s, and YOLOX-m, the Improved-YOLOv5s model demonstrates notable advantages in detection accuracy, network complexity, and detection speed.
To improve the insufficient adaptability of single degree-of-freedom (DOF) closed-chain legs to unknown surroundings, this paper studies a novel adjustable closed-chain leg mechanism. Based on the traditional six-bar Stephenson-III leg mechanism, we add one degree-of-freedom to realize the adjustment of frame rod position, by which different shapes of foot trajectory curves are obtained. The height of the foot trajectory can be adjusted longitudinally to enhance the obstacle-surmounting ability, and the length of the walking stride can be changed in the transverse direction to achieve the landing point placement of the swing leg. The structure optimization, gait analysis and obstacle-surmounting simulation in different scenarios are carried out. The test results of walking ability and obstacle surmounting performance in specific surrounding show that the designed closed-chain leg mechanism can increase the maximum obstacle-crossing height and walking stability, which proves the promising characteristics of the proposed leg mechanism in the application of multi-legged robots.
Ultra-Wide Band (UWB) positioning system has become the main research object in indoor positioning field due to its superior positioning performance. Accuracy at the scale of centimeters is achievable in an ideal positioning scenario. Nonetheless, the intricate and changeable indoor surroundings may cause UWB signals to propagate beyond the line-of-sight (LOS) and weaken in signal strength, resulting in a considerable reduction in the precision of positioning. In this paper, methods for identifying non-line-of-sight (NLOS) errors, alleviating NLOS errors, and integrating multi-sensor information are analyzed and studied to provide reference for improving the NLOS positioning accuracy of UWB.
With the continuous development of Machine Learning, Reinforcement Learning has achieved excellent results in many fields. An improved Deep Deterministic Policy Gradient (DDPG) algorithm is proposed in this paper to optimize Reinforcement Learning control performance, because of the issue of sparse rewards during training in continuous action spaces such as robotic arms, leading to slow convergence and low success rates. The algorithm merges the DDPG algorithm with the Hindsight Experience Relay (HER) algorithm, enabling the agent to learn from the failed experience when learning. Moreover, a double experience replay buffer is introduced, comprising a primary buffer and a secondary buffer containing high-reward experiences, which enhances the sampling probability of high-quality samples. At long last, a PyBullet simulation environment was utilized to execute a simulation experiment. The robotic arm using the improved DDPG algorithm achieved a success rate of nearly 100% in about 220 epochs of training, and the algorithm convergence speed was greatly improved. The trained robotic arm can reach any target point in the workspace. The efficacy of the improved DDPG algorithm is demonstrated by the outcomes, providing a valuable reference for further exploration into the intelligent control of robotic arms for object grasping.
In order to improve the performance of lane detection algorithms under complex scenes like obstacles, we proposed a multi-lane detection method based on dual attention mechanism. Firstly, we designed a lane segmentation network based on a spatial and channel attention mechanism. With this, we obtained a binary image which shows lane pixels and the background region. Then, we introduced HNet which can output a perspective transformation matrix and transform the image to a bird's eye view. Next, we did curve fitting and transformed the result back to the original image. Finally, we defined the region between the two-lane lines near the middle of the image as the ego lane. Our algorithm achieves a 96.63% accuracy with real-time performance of 134 FPS on the Tusimple dataset. In addition, it obtains 77.32% of precision on the CULane dataset. The experiments show that our proposed lane detection algorithm can detect multi-lane lines under different scenarios including obstacles. Our proposed algorithm shows more excellent performance compared with the other traditional lane line detection algorithms.
Under the complex environment, in the process of sensing and understanding the surrounding scene with an unmanned system, the traditional single-source panorama provides limited information and is prone to interference from lighting, it is difficult to capture hidden objects, and the user's three-dimensional (3D) experience is poor. To solve this problem, we propose a real-time panoramic map modeling method based on multisource image fusion and 3D rendering with intensity, hue, saturation transform and wavelet technology, open graphics library. Five general image sharpness evaluation methods are used to analyze the single-source and multisource close-range and distant images during the day and night. The results show that in the daytime, the sharpness of the multisource image is significantly improved compared with the single-source visible light image, especially the energy gradient evaluation can reach 3.4 times. Similarly, compared with the single-source infrared image, the multisource image at night also improve significantly, the energy gradient evaluation can reach 1.4 times. This conclusion is verified by the subsequent experiments of several indoor scenes, the 3D rendering processing technology of panoramic map is realized, which improves user experience of 3D stereoscopic vision, and will provide multisource image fusion and 3D rendering technical support for unmanned real-time panoramic map modeling. (c) 2023 SPIE and IS&T
To improve the performance of image semantic segmentation on accuracy and efficiency for practical applications, in this study, we propose a real-time semantic segmentation algorithm based on improved BiSeNet. First, the redundancy of certain channels and parameters of BiSeNet is eliminated by sharing the heads of dual branches, and the affluent shallow features are effectively extracted at the same time. Subsequently, the shared layers are divided into dual branches, namely, the detail branch and the semantic branch, which are used to extract detailed spatial information and contextual semantic information, respectively. Furthermore, both the channel attention mechanism and spatial attention mechanism are introduced into the tail of the semantic branch to enhance the feature representation; thus the BiSeNet is optimized by using dual attention mechanisms to extract contextual semantic features more effectively. Finally, the features of the detail branch and semantic branch are fused and up-sampled to the resolution of the input image to obtain semantic segmentation. Our proposed algorithm achieves 77. 2% mIoU on accuracy with real-time performance of 95. 3 FPS on Cityscapes dataset and 73. 8% mIoU on accuracy with realtime performance of 179. 1 FPS on CamVid dataset. The experiments demonstrate that our proposed semantic segmentation algorithm achieves a good trade-off between accuracy and efficiency. Furthermore, the performance of semantic segmentation is significantly improved compared with BiSeNet and other existing algorithms.
Aiming at the problems in the path planning of the manipulator by artificial potential field (APF) method, a method combining the APF adaptive variable step size in the joint space and goal-biased rapidly-exploring random tree (RRT) is proposed. The APF obstacle avoidance planning is performed in the joint space to reduce the number of inverse kinematics and the sudden change of joint angles. The collision and target unreachability problems in the path planning are solved by improving the repulsive and gravitational potential field functions. The Cauchy probability distribution is used to change the joint angle step size through the distance between the end point and the obstacle. By adjusting the bias of the RRT algorithm, suitable temporary target points are generated to solve the local minima problem of the APF. The obstacle avoidance simulation of the manipulator is carried out in the presence of local minima of the APF. Adaptive variable-step path planning can generate smooth trajectories and improve the search efficiency. The goal-biased RRT selects the temporary target point and the overall path length becomes smaller. The picking manipulator can effectively meet the requirements of obstacle avoidance picking tasks under the improved algorithm.
In the traditional direct visual odometry, it is difficult to satisfy the photometric invariant assumption due to the influence of illumination changes in the real environment, which will lead to errors and drift. This paper proposes an improved direct visual odometry system, which combines luminosity and depth information. The algorithm proposed in this paper uses Kinect 2 to collect RGB images with the corresponding depth information, and selects points with large changes of gray gradient to construct a luminosity error function and uses the corresponding depth information to construct a depth error function. The two error functions are merged into one function and converted into the least squares function of the pose of camera, the Levenberg-Marquardt algorithm is used to solve the camera pose. Finally, the Graph optimization theory and the g2o library are used to optimize the initial pose. Experiments show that the algorithm can reduce the error to a certain extent and reduce the drift caused by illumination changes.
为解决室内三维地图场景中移动机器人的路径规划与避障问题,将机器人实际体积纳入考虑范畴.定义机器人为正方体包围盒,讨论了机器人体积与周围障碍物的关系,提出了安全区域的概念,并对A?路径规划算法进行拓展.为解决三维点云节点数量大难以进行处理的问题,将数据格式转换成了八叉树结构,并在八叉树上创建了最优路径搜索方案,以提升数据处理效果.通过室内三维地图场景中移动机器人的最优路径搜索,实验结果表明上述方法能够在三维场景里有效解决机器人的路径规划与避障问题.
A multi-channel (9*X) data acquisition system based on FPGA and MCU is developed, which can be used for signal field strength detection of base station equipment such as RFID. The system consists of data acquisition array, central controller, embedded firmware, and upper computer. Using the modular design method, the data acquisition array can be designed by changing the number of channels. The single-channel data acquisition module has a sampling resolution of 16 bits, a sampling rate of 200 kSPS/ch and dynamic range 70dB. The master-slave architecture is used to realize multi-channel synchronous acquisition of data. When the central controller communicates with the upper computer, it controls the operation of the single-channel data acquisition module and obtains the sampling data in a parallel manner. The embedded firmware is written in Verilog and C language to realize data acquisition, buffering, and sending and receiving of uplink and downlink signaling. The upper computer has the functions of system control and communication, establishes a data link with the central controller through the Ethernet cable, and obtains multi-channel sampling data through signaling. The system verifies the performance of the multi-channel data acquisition system through Analog-to-digital converter, single-channel data acquisition module and system performance test.
In order to realize the compliant grinding and polishing of complex surfaces, the traditional impedance control cannot adjust the system impedance, so the force tracking deviation is large, the adaptive variable impedance control can adapt to the environmental changes to adjust the system impedance and realize the stable tracking of the contact force, but the overshoot is too large. In this paper, the constant force control of compliant grinding and polishing processing is studied on the designed macro-mini robotic system, and an adaptive variable impedance control algorithm with pre-PD adjustment is proposed. The control law is used to update the damping term in the impedance parameters of the system, and realize the constant force control process in which the grinding and polishing head is approximately perpendicular to the machined surface, and the stability and convergence of the force control algorithm are proved. Through simulation and experiment, the four algorithms of impedance control, adaptive impedance control, adaptive variable impedance control and the proposed PD-adaptive variable impedance control are compared and analyzed in the force tracking performance of plat surface, sloped surface and curved surface, and then in the grinding and polishing experiments to verify the grinding and polishing force tracking error, the force tracking errors of the other three control algorithms are controlled within 2.08 N, 1.30 N, and 1.34 N, and the overshoot amounts to 19.2%, 56.3%, and 11.1%. The proposed algorithm can control the force tracking error within 0.78 N, the overshoot is less than 4.5%, which verifies that the proposed force control algorithm is more superior in force control performance and is suitable for force control scenarios of actual grinding and polishing.