
Driven by breakthroughs in next-generation artificial intelligence,embodied intelligence is rapidly advanc-ing into industrial manufacturing.In flexible manufacturing,industrial embodied intelligence faces three core challenges:accurate process modeling and monitoring under limited perception,dynamic balancing between flexible adaptation and high-precision control,and the integration of general-purpose skills with specialized industrial operations.Accordingly,this survey reviews existing work from three viewpoints:Industrial Eye,Industrial Hand,and Industrial Brain.At the perception level(Industrial Eye),multimodal data fusion and real-time modeling in complex dynamic settings are examined.At the con-trol level(Industrial Hand),flexible,adaptive,and precise manipulation for complex manufacturing processes is analyzed.At the decision level(Industrial Brain),intelligent optimization methods for process planning and line scheduling are sum-marized.By considering multi-level collaboration and interdisciplinary integration,this work reveals the key technological pathways of embodied intelligence for closed-loop optimization of perception-decision-execution in manufacturing systems.A three-stage evolution model for the development of embodied intelligence in flexible manufacturing scenarios,comprising cognition enhancement,skill transition,and system evolution,is proposed,and future development trends are examined,to offer both a theoretical framework and practical guidance for the interdisciplinary advancement of industrial embodied intelligence in the context of flexible manufacturing.
3D Multi-Object Tracking (MOT) is an important part of the unmanned vehicle perception module. Most methods optimize object detection and data association independently. These methods make the network structure complicated and limit the improvement of MOT accuracy. we proposed a 3D MOT framework based on simultaneous optimization of object detection and scene flow estimation. In the framework, a detection-guidance scene flow module is proposed to relieve the problem of incorrect inter-frame assocation. For more accurate scene flow label especially in the case of motion with rotation, a box-transformation-based scene flow ground truth calculation method is proposed. Experimental results on the KITTI MOT dataset show competitive results over the state-of-the-arts and the robustness under extreme motion with rotation.
The objective of this work is to expand upon previous works, considering socially acceptable behaviours within robot navigation and interaction, and allow a robot to closely approach static and dynamic individuals or groups. The space models developed in this dissertation are adaptive, that is, capable of changing over time to accommodate the changing circumstances often existent within a social environment. The space model's parameters' adaptation occurs with the end goal of enabling a close interaction between humans and robots and is thus capable of taking into account not only the arrangement of the groups, but also the basic characteristics of the robot itself. This work also further develops a preexisting approach pose estimation algorithm in order to better guarantee the safety and comfort of the humans involved in the interaction, by taking into account basic human sensibilities. The algorithms are integrated into ROS's navigation system through the use of the $costmap2d$ and the $move\_base$ packages. The space model adaptation is tested via comparative evaluation against previous algorithms through the use of datasets. The entire navigation system is then evaluated through both simulations (static and dynamic) and real life situations (static). These experiments demonstrate that the developed space model and approach pose estimation algorithms are capable of enabling a robot to closely approach individual humans and groups, while maintaining considerations for their comfort and sensibilities.
Multispectral imagery is frequently incorporated into agricultural tasks, providing valuable support for applications such as image segmentation, crop monitoring, field robotics, and yield estimation. From an image segmentation perspective, multispectral cameras can provide rich spectral information, helping with noise reduction and feature extraction. As such, this paper concentrates on the use of fusion approaches to enhance the segmentation process in agricultural applications. More specifically, in this work, we compare different fusion approaches by combining RGB and NDVI as inputs for crop row detection, which can be useful in autonomous robots operating in the field. The inputs are used individually as well as combined at different times of the process (early and late fusion) to perform classical and DL-based semantic segmentation. In this study, two agriculture-related datasets are subjected to analysis using both deep learning (DL)-based and classical segmentation methodologies. The experiments reveal that classical segmentation methods, utilizing techniques such as edge detection and thresholding, can effectively compete with DL-based algorithms, particularly in tasks requiring precise foreground-background separation. This suggests that traditional methods retain their efficacy in certain specialized applications within the agricultural domain. Moreover, among the fusion strategies examined, late fusion emerges as the most robust approach, demonstrating superiority in adaptability and effectiveness across varying segmentation scenarios. The dataset and code is available at https://github.com/Cybonic/MISAgriculture.git.
Model-based systems engineering (MBSE) is a methodology that exploits system representation during the entire system life-cycle. The use of formal models has gained momentum in robotics engineering over the past few years. Models play a crucial role in robot design; they serve as the basis for achieving holistic properties, such as functional reliability or adaptive resilience, and facilitate the automated production of modules. We propose the use of formal conceptualizations beyond the engineering phase, providing accurate models that can be leveraged at runtime. This paper explores the use of Category Theory, a mathematical framework for describing abstractions, as a formal language to produce such robot models. To showcase its practical application, we present a concrete example based on the Marathon 2 experiment. Here, we illustrate the potential of formalizing systems -- including their recovery mechanisms -- which allows engineers to design more trustworthy autonomous robots. This, in turn, enhances their dependability and performance.
针对五指机械手抓取成功率低和抓取任务相对简单的问题,基于区域姿态解算方法设计并实现了一种满足不同复杂任务的五指抓取系统.首先,设计了一种气动五指软爪,该软爪使用多种不同刚度的材料制成,由1根主气管和5根支气管驱动,控制复杂度较低,机械性能和抓取性能良好,适用不同的抓取策略.进一步结合软爪的特点提出了基于区域姿态解算的抓取策略.通过预测人手抓取物体时在物体上的接触区域,求解接触区域与软爪指尖在空间上的姿态解,计算软爪的抓取姿态和关节弯曲角度,该策略能够生成高鲁棒性的抓取姿态.然后,设计了包含大量物体的接触数据集.对数据集中的物体标注人手在抓取操作中指尖的接触区域,并尽可能地去除场景信息,提高了数据集的通用性,可用作基准数据集测试算法性能.最后,设计了一系列实验来验证软爪和抓取策略在复杂场景下的抓取性能,实验结果表明了所设计的五指软爪抓取系统在复杂场景下的有效性和可靠性.
捕蝇草的叶片运动具备动作迅速、可逆、控制方便、结构简单等优点.本文从捕蝇草的快速运动中获得灵感,从微观和宏观层面上研究捕蝇草的基本结构和运动学特性.微观层面上,使用植物切片法和组织透明技术观察分析叶片的脉管系统结构、细胞尺寸及横向、纵向上细胞的排布规律.宏观层面上,利用高速摄像机和Kinovea图像分析软件分析叶片边缘的水平位移、速度、加速度、张开角度等的变化规律,采用非接触全场应变测量系统获得捕蝇草叶片闭合过程中外表面的应变情况;基于多孔介质模型,研究水分传输引起的细胞膨胀变形运动的规律,以及在流体作用下仿生柔性叶片的弯曲变形规律;最后使用3D打印技术制备仿生柔性叶片驱动器,进行实验验证.研究表明,本文提出的仿生柔性叶片驱动器通过流固耦合仿真技术能够在2 MPa的液压下在1 s内完成弯曲变形,最大弯曲角度为32.27°.同时,制备的仿生柔性叶片驱动器样机的实验结果与仿真结果基本吻合,验证了流固耦合仿真技术的可行性与正确性,完成了对捕蝇草叶片快速闭合运动的复现,闭合过程中实现了叶片的快速弯曲.
针对基于地图的移动机器人导航框架部署在动态复杂环境时出现的问题,提出一种基于时序-双延迟深度确定性策略梯度(TS-TD3)的无地图导航方法.首先,将动态场景(具有环境部分可观测性)的导航任务定义为部分可观测马尔可夫决策过程(POMDP).其次,引入经过长短期记忆组件处理的历史信息作为模型的输入,为策略网络的确定性策略梯度引入历史信息基准,以处理隐藏在环境观测集合中的状态信息,将关注导航动作时序关联性的评价标准引入评价网络.再次,通过专家经验网络在训练前期指导策略网络的输出,以规范导航动作.最后,建立演员-评论家框架的深度强化学习(DRL)端到端模型,根据传感器感知结果直接输出控制动作.与主流DRL方法进行对比实验,在仿真实验中,该方法运动轨迹自然、稳定、具有连续性,能处理多动态障碍物交汇情况,整体导航效果表现最优;在真实动态环境的测试中,模型未作调整直接部署在未知环境中,模型的导航效果和泛化性得到验证.
本综述涵盖了深度学习技术应用到SLAM(同步定位与地图创建)领域的最新研究成果,重点介绍和总结了深度学习在前端跟踪、后端优化、语义建图和不确定性估计中的研究成果,展望了深度学习下视觉SLAM的发展趋势,为后继者了解与应用深度学习技术、研究移动机器人自主定位和建图问题的可行性方案提供助力.
针对人机共融场景中移动机器人与人类关注区域交叠的情景,提出了一种考虑人类视线区域约束的社交导航方法.首先,基于感知到的注视方向和注视目标,确定视线状态并构建视域交互模型.然后,借鉴电磁场和极限环理论,引入包括视线静电力、视线磁偏转力和视域交互势场力在内的扩展社会力对机器人的行为和人类视线之间的关系进行建模.进而驱动机器人根据任务类型以符合社交规则的方式主动避让或接近人的关注区域,在保证人类心理舒适度的同时提升了服务机器人的社交性.仿真实验部分的定性定量分析及相应的问卷评估表明了在机器人社交导航中引入人类视线区域约束的必要性,也进一步验证了所设计导航方法的有效性和适用性.
针对多样性目标在非结构化环境中的抓取位姿难以估计的问题,提出一种基于上下文聚合策略的轻量级编/解码抓取位姿检测网络.首先,以编/解码网络架构为基础,利用深度可分离卷积层与混洗单元构建目标特征深度分离-融合提取块,减少编码网络参数量,增强网络对抓取区域特征的提取能力;其次,利用双线性插值法和深度可分离卷积层建立深度分离-重构块,在恢复高层特征丢失信息的同时,有效减少解码网络的参数量;最后,针对可抓取区域像素点与目标物体全貌之间的非一致性问题,基于交叉熵辅助损失和自注意力机制,提出一种抓取区域上下文聚合策略,引导网络增强可抓取目标区域特征的表征能力,抑制非抓取像素点的冗余特征.实验结果表明,所提网络在Cornell数据集的图像拆分与对象拆分子集上抓取检测准确率分别可达97.8%与93.8%,单张图像检测速度可达64.93张/秒;在Jacquard数据集上抓取检测准确率可达95.1%,单张图像检测速度可达60.6张/秒.与对比网络相比,所提网络不仅计算量与参数量较小,而且抓取检测的准确率与速度均有明显提升,在真实场景下对9种物体的抓取检测验证中,抓取成功率达到93.3%.
水下六足机器人具有丰富的步态样式和冗余的肢体结构,凭借其离散式的地面支撑和对水下障碍、礁坪等复杂特殊地形的极强适应性,具有广泛的应用前景.本文通过充分的文献调研和总结,对目前水下六足机器人平台发展现状进行了综述.针对水下六足机器人的海底爬行能力,分别从稳定性判据、路径规划与自适应行走方法3项关键技术进行了分析和总结;在稳定性判据方面,分别针对水下六足机器人静态稳定性判据与动态稳定性判据进行阐述;在路径规划方面,在目前典型陆地六足机器人路径规划方法的基础上,结合水下六足机器人独有的运动特性进行阐述;针对水下六足机器人自适应行走方面,分别阐述传统自适应调整方法与目前基于深度强化学习的方法进行介绍;最后对这些关键技术的未来发展趋势进行了展望.
深海环境的长周期监测已成为人类分析和认识海洋生态系统和海洋环境变换过程的必要手段,而深海长期驻留自主水下机器人(LRAUV)系统则是实现深海环境长周期监测的有效装备,成为近年来深海领域的研究热点.本文通过梳理、总结前人的研究,首先对LRAUV系统的含义及其工作模式进行了介绍;对LRAUV系统的典型应用领域进行了分析与展望;然后,对国内外LRAUV系统进行了综述,包括其配置、功能与技术参数等内容,并总结了其发展趋势;最后对LRAUV系统长期生存、基站支撑、AUV自主对接、能源补充及数据传输、探测作业等关键技术的研究现状、难点问题及未来发展方向进行了综述与分析.
针对复杂环境和不确定的动态模型问题,提出了一种自适应神经网络控制算法,并且基于李雅普诺夫理论保证了闭环系统的稳定性.相较而言,基于模型的控制策略需要精准地知道微机器人的动力学以及周围环境参数,而本文提出的基于径向基函数神经网络的状态反馈控制策略可以有效地根据状态与期望轨迹在线估计系统模型的不确定项.最后,在所开发的磁场驱动的微型机器人系统上进行了 2项实验并进行了对比,以验证所提出的控制器的有效性.结果表明,曲线与直线的轨迹跟踪均方根误差分别达到6.2204像素与6.4279像素,明显优于传统的PID(比例-积分-微分)算法.
为提高类人机器人面部情感迁移的时空一致性并降低机械运动约束的影响,提出一种基于Trans-former 架构和B样条平滑约束的机器人面部情感迁移网络RFEFormer.该网络由面部形变编码子网和驱动序列生成子网组成.在面部形变编码子网中,为表征帧内不同层次、不同粒度的空间信息,基于域内形变注意力和域间协作注意力双重机制构建帧内空间注意力模块并嵌入到Transformer编码器中;在驱动序列生成子网中,利用Transformer解码器实现面部时空序列和历史电机驱动序列的交叉注意以及未来电机驱动序列的多步预测,并引入三次B样条平滑约束实现预测序列的规整.实验结果表明:RFEFormer网络的电机驱动偏差、面部形变逼真度和电机运动平滑度分别为3.21%、89.48%和90.63%,且实时面部情感迁移帧率大于25帧/秒.与相关方法相比,RFEFormer网络在满足实时性的同时提升了逼真度、平滑度等时序指标性能,而人类感官对这些指标更为敏感、也更为关注.
Aiming at the actual clinical navigation needs of minimally invasive spine surgery, an augmented reality-based navigation system for minimally invasive spinal surgery is designed and developed. The scan and reconstruction of the patient’s spine lesions are realized by CT(computed tomography) before surgery, and the appropriate spine model is imported into the Unity-3 D platform to add control scripts for it. To achieve the synchronous display of medical images of surgical instrument and lesion model, the calibration ball is adopted to calibrate surgical instruments, the Polaris Vega optical tracker is used to track surgical instruments in real time, the coordinate conversion relationship between surgical instrument and Polaris Vega optical tracker is established, and the posture information of the identified surgical instrument is sent to HoloLens device in real time. During surgery, it provides doctors with three-dimensional images to visualize spinal lesions, helping doctors to locate lesions and navigate surgical instruments. The experimental test on the spine model show that the system navigation error is less than 2.8 mm, which can meet the requirements of clinical application in spine surgery.
The serial elastic joint provides a compliant design solution for lower limb rehabilitation exoskeleton, but there are still many challenges in the design of its structure and control. Aiming at the human-robot interaction(HRI)requirements of rehabilitation exoskeletons, a series elastic joint integrated with the novel elastic element is proposed, which can realize linear/nonlinear stiffness switching subject to external loads, while achieving structural compliance. Secondly, the rigid-flexible coupling dynamic model of the modular joint is established, and a step-by-step decoupling identification of the physical system is carried out. Under the framework of model predictive control, a control algorithm is designed for nonlinear and non-convex optimization problems through an iterative linearization method. Finally, several groups of trajectory tracking experiments and disturbance experiments are conducted, and the control bandwidth of the system is discussed. Experimental results show that the proposed control method can track different reference trajectories, effectively reduce control energy consumption and suppress external disturbances under the condition of control constraints.
面向经肛内镜微创手术(TEM),设计了一种基于万向轴关节和串联杆结构的柔性机器人.基于柔性体变形的恒曲率模型建立了万向轴关节的正运动学,并提出了一种万向轴关节的正运动学优化方法,建立了柔性机器人驱动空间、关节空间以及工作空间之间的映射关系.设计并实现了柔性机器人的主从控制方法.采用光学定位系统对柔性机器人末端定位误差进行检测,实验结果表明,柔性机器人末端的绝对定位误差小于6.00 mm,重复定位误差小于3.00 mm.
Currently, underwater robots are used in the fields of exploration, monitoring, search & rescue, and maintenance in ocean development. As a multi-joint and highly flexible underwater robot, the underwater snake robots are superior to traditional underwater robots in terms of motion efficiency, function execution, environment adaptation, and autonomous learning. For the motion efficiency problem of underwater snake robots with a rear thruster, an efficient motion mode combining lateral undulation motion and thruster propulsion is proposed. Firstly, the dynamic model of underwater snake robot is established. Then, the motion efficiency of the snake robot is evaluated by the transportation economic metric method. Based on this evaluation method, the NSGA-II(non-dominated sorting genetic algorithm II) is used to optimize four kinematic parameters, including three lateral undulation gait parameters(the amplitude of joint motion, the frequency of joint motion,the phase shift between the joints) and the thruster force. The optimization results of the three motion modes show that: in the low-speed segment, the lateral undulation mode is of the highest efficiency; in the medium speed segment, the hybrid motion mode combining lateral undulation motion and thruster propulsion, is of the highest efficiency, and in this mode, not only the speed is much faster than that of the lateral undulation mode, but also the overall motion efficiency is higher than that in the thruster mode; in the high-speed segment, the thruster mode is of the highest efficiency and the fastest forward motion speed.Finally, the effectiveness of the proposed motion mode is verified by pool experiments. It maximizes the movement ability of the underwater snake robots with thrusters, improves the movement efficiency of the robot, and elongates the battery life.
现有的多AGV(自动导引车)系统处理死锁的方案往往约束过强,压缩了潜在的性能优化空间。本文提出一种高度灵活的死锁避免算法,通过分析系统状态图中的宏环结构并结合银行家算法来实现状态图的链状结构判断,在确保算法高效性(最坏情形时间复杂度为O((|V|+|E|)|A|),其中V、E、A分别代表节点、边、AGV)的同时,实现了灵活的死锁避免。通过离散事件系统仿真及实际系统应用验证了算法的有效性,结果表明,在典型路线图上,该算法相较于经典的银行家算法及其变种,容许覆盖率提升高于16%,在使用相同任务分配、路径规划算法的情况下,任务平均完成时间降低了15%,具有更高的灵活性,有效提升了系统性能优化的潜力。