Layered motion estimation (LME) in X-ray fluoroscopy is a challenging, ill-posed and non-convex problem due to transparency effects and the way the image is defined. Minimizing an energy formulation of layered motion estimation is computationally expensive. For clinical usability of this approach, we propose to use primal-dual optimization parallelized using a graphical processing unit (GPU) to reduce the overall run-time of this algorithm. Experimentally this method is able to substantially reduce target registration error by 70% on manually annotated landmarks on five distinct image sequences compared to the static baseline, similar to prior work on this domain. However, the overall runtime of our method on a conventional GPU is less than 3.3 seconds compared to several minutes for the state of the art. Considering typical framerates of X-ray fluoroscopy devices, this runtime makes the application of layered motion estimation feasible for many clinical workflows.