International Conference on Automation, Control and Robotics Engineering(2023)
School of Instrument Science &; Engineering
被引用0|浏览16
摘要
The use of monocular cameras for ego-motion estimation in autonomous driving is a fundamental technique with significant potential for development. Recently, unsupervised CNNs based schemes have rapidly advanced due to their label-free advantages. However, traditional CNNs lack the ability to capture global dependencies, which are essential for sequential tasks like pose, depth, and optical flow estimation. In this paper, according to the characteristics of these three-independent tasks, we designed different network frameworks fully based on the transformer separately to improve performance by taking advantage of the global dependencies of different tasks. After unsupervised training, the model can independently predict the pose, depth, and optical flow results for the vehicle, solely based on a monocular camera. The experimental results conducted on the KITTI demonstrate that the proposed method obtain satisfactory results in all three tasks.