International Conference on Remote Sensing, Mapping, and Image Processing (RSMIP 2024)(2024)
被引用0|浏览2
摘要
Tracking targets accurately and robustly in visually complex environments poses a formidable challenge. To address this, capturing a resilient appearance representation is essential while augmenting the model's ability to generalize and cope with diverse challenges such as object deformation, variations in illumination, changes in scale, and motion blur. This paper presents a method for sturdy tracking in intricate scenarios, employing the efficient convolution operator (ECO) tracker. Our approach incorporates the 2 key concepts: a) extracting profound features via the Conformer network by increasing the number of underlying channels, and b) flexibly adjusting the fusion weight for shallow and profound features based on factors such as the peak-to-sidelobe ratio and the joint score of adjacent frame trajectory smoothness. This approach enhances the model's adaptability and generalization in intricate environments by capitalizing on the complementary aspects of deeper-layer and shallow-layer features. Experimental outcomes validate the algorithm's efficacy in tackling varied challenges related to target tracking in complex environments, ensuring robust tracking with consistently high accuracy.