XR-1: Towards Versatile Vision-Language-Action Models Via Learning Unified Vision-Motion Representations | AMiner