The exploration and resolution of persistent noise incursions within the tracking sequences, especially the occlusion, illumination variations, and fast motion, have garnered substantial attention for their functional properties in enhancing the accuracy and robustness of visual object trackers. However, existing visual object trackers, equipped with template updating mechanisms or calibration strategy, heavily rely on time-consuming historical data to achieve optimal tracking performance, impeding their real-time tracking capabilities. To address these challenges, this paper introduces a long-short term dual level memory augmented transformer structure aided visual object predictor (MeAP). The key contributions of MeAP can be summarized as follows: 1) the formulation of a noise model for specific invasion events based on incursion effects and corresponding template strategies serving as the foundation for more efficient memory utilization; 2) The memory exploration scheme based online tracking mask-based feature extraction strategy and the transformer architecture is introduced to mitigate the impact of noise invasion during memory vector construction; 3) the memory utilization scheme based target basic feature and dual feature target mask predictor is provided to implement the scene-edge feature for mask-based feature extraction method and jointly predict the accurate location of the tracking target.. Extensive experiments conducted on OTB100, NFS, VOT2021, and AVisT benchmarks demonstrate that MeAP, with its introduced modules, achieves comparable tracking performances against other state-of-the-art (SOTA) trackers, and operates at an average speed of 31 frames per second (FPS) across 4 benchmarks.