Underwater pipelines play an important role in offshore oil and gas transmission, and underwater cables are vital for global interconnection communication. However, due to the harsh underwater environment, underwater pipelines and cables are easily damaged. In order to complete the inspection of underwater pipelines and cables, researchers have proposed many inspection platforms, detection and tracking methods to improve inspection efficiency and automation level. In this article, the state-of-the-art techniques for underwater pipeline and cable external inspection are comprehensively reviewed. First, underwater pipeline and cable inspection platforms are classified and introduced, and their advantages and disadvantages are discussed. Then, underwater pipeline and cable detection methods, including acoustic method, passive optical method, active optical method, and electromagnetic method, are introduced, and their performance comparison is provided. Besides, underwater pipeline and cable tracking methods are analyzed. Finally, we discuss the key issues and future directions of underwater pipeline and cable inspection. This review will help researchers better understand the current research progress in underwater pipeline and cable inspection platforms, detection and tracking methods, which provide guidance for further research and deployment.
A significant challenge in exoskeleton robotics is the need to dynamically adapt control profiles to individual motion preferences, thereby ensuring both efficient and comfortable assistance. Currently, since user experience can serve as a comprehensive metric for evaluating the effectiveness of assistance, user preference-based optimization methods have been widely studied for parameter tuning. However, the existing methods rely heavily on extensive human-robot online interactions and suffer from slow optimization speed, which not only induces user fatigue but also compromises optimization effectiveness. Therefore, this paper aims to explore an efficient preference-based optimization framework for personalized exoskeleton assistance that can learn optimal parameters with minimal interaction. We propose a preference-based Bayesian optimization (PbBO) approach that can improve sample efficiency by leveraging knowledge about the sampling distribution of candidate sets. For optimizing six control parameters, PbBO can converge to user-preferred parameters with 90.7
Offline reinforcement learning (RL) aims to optimize a policy using collected data without online interactions. Model-based approaches are particularly appealing for addressing offline RL challenges because of their capability to mitigate the limitations of data coverage through data generation using models. Nonetheless, a prevalent issue in offline RL is the overestimation caused by distribution shift. This study proposes a novel model-based offline RL algorithm named Conservative Reward for model-based Offline Policy optimization (CROP). CROP introduces a streamlined objective that concurrently minimizes estimation error and the rewards of random actions, thereby yielding a robustly conservative reward estimator. Theoretical analysis shows that the designed conservative reward mechanism leads to a conservative policy evaluation and mitigates distribution shift. Experiments showcase that with the simple modification to reward estimation, CROP can conservatively estimate the reward and achieve competitive performance with existing methods. The source code will be available after acceptance.
Underwater three-dimensional (3-D) technology is of great significance in underwater structure inspection, underwater terrain mapping, and underwater archaeology. Underwater structured light systems (SLS) are widely used in underwater 3-D reconstruction because of good measurement accuracy and strong robustness. However, the existing underwater SLS have the drawback of low measurement efficiency and are not suitable for the underwater motion objects 3-D reconstruction. To overcome these issues, an event-based underwater SLS for high-speed 3-D reconstruction is proposed in this article. First, an underwater SLS, including an event camera and a scanning laser, is developed, and a robust laser stripe extraction algorithm is proposed to suppress noise data in the event stream. Second, an underwater self-scanning structured light measurement model considering both camera refraction and laser plane refraction is established, and an efficient model parameter calibration method is presented, which does not require complex calibration processes in water. Finally, underwater 3-D reconstruction experiments are conducted, including underwater 3-D reconstruction in different turbidity water and underwater 3D reconstruction of three motion objects. The reconstruction accuracy is less than 4 mm, and the measurement efficiency can reach 500 Hz. To the best of our knowledge, this article first achieves dense 3-D reconstruction of underwater motion object.
Adaptive torque prediction in dynamic exoskeleton scenarios requires expensive motion capture systems, which are infeasible in complex outdoor environments. Trajectory prediction has emerged as one of the effective approaches to address such an issue. However, the core challenges of exoskeleton trajectory prediction are twofold: establishing the mapping from multi-modal features to trajectory information; constructing the mapping from trajectory to torque. For the former, most existing methods perform only single-step prediction and neglect inter-subject trajectory variability, thereby limiting the trajectory optimization space and prediction generalization. To address this, this paper proposes a fast flow matching method that enables accurate trajectory prediction and better generalization for real-time performance, where trajectory generation errors and encoded observations are used to guide the training direction. For the second challenge, due to the high dynamics of the human-robot system and the strong coupling between perception and control, simple control methods struggle to achieve efficient assistance based on the predicted trajectory. This paper utilizes model predictive control and designs a novel optimization objective to optimize torque, ensuring the exoskeleton achieves comfortable and robust assistance. By integrating the above two components, the unified policy, denoted as ExoTraj, is developed to enable adaptive assistance in complex outdoor scenarios without high data acquisition cost. Experimental results show that compared to traditional methods, ExoTraj reduces cross-subject prediction error by 14.0% during the online phase and maintains robustness against external noise. Relative to the zero torque condition, ExoTraj decreases metabolic rate by 11.5-24.4%, heart rate by 1.7-19.5%, and peak muscle activation levels by 10.9-41.3%, respectively.
High-precision mapping across two-dimensional and three-dimensional vascular images enables 3D surgical navigation for cerebrovascular interventions, which assists surgeons in improving procedural accuracy and avoiding damage to critical blood vessels and neural structures. However, cross-dimensional mapping of vascular images faces challenges, including significant cumulative errors and poor robustness. In this paper, we propose a Vascular-anatomy and Projective-prior based Mapping (VPMap) framework, which leverages the anatomical features and projective priors of vessels to achieve robust and precise mapping of 2D/3D vascular structures. Firstly, a Vascular-Anatomy-aware Segmentation (VASeg) model is designed by leveraging the local morphological features and global structural information to reduce the fractures and outliers of segmented vessels. Secondly, a Projective-Prior-enhanced Hierarchical Registration (PPHReg) method is proposed, in which the vessel skeletons are hierarchically matched and weighted with projective priors to achieve robust and accurate pose calibration. Then, comprehensive experiments are conducted to evaluate the performance of the proposed framework. VASeg achieves the best segmentation results in an in-house and two public vessel datasets, while PPHReg shows robust and accurate registration performance in an in-house 2D/3D intracranial artery registration dataset. Moreover, the effectiveness of VPMap is further validated through joint experiments on the 2D/3D image-pairs collected from real surgeries. Finally, based on the millimeter-level fusion accuracy of VPMap, a surgical navigation system is developed, which achieves precise 3D guidance for cerebrovascular interventions and demonstrates the potential to improve surgical safety. The code is available at https://github.com/bytewandering/VPMap.
The increasing incidence of structural heart disease has emerged as a significant global health challenge. Recently, ultrasound-guided interventional techniques have demonstrated the potential to replace traditional treatment methods due to their advantages, including the absence of radiation risks and the ability to provide real-time monitoring during surgeries. This study addresses challenges in instrument identification caused by factors such as low imaging resolution, variable instrument morphology in different sections, and the tendency for instruments to disappear or become obstructed. To tackle these issues, we propose a deep learning-based method for Video Instance Segmentation (VIS). The proposed method consists of three key components: 1) A frame-level detector extracts features from each image frame and generates queries for object instances. 2) A model architecture integrates the spatial feature extraction capabilities of convolutional neural networks with the temporal modeling capabilities of transformers, utilizing a windowed attention mechanism to capture inter-frame dependencies. 3) A multi-video segment joint memory learning mechanism is introduced, which stores historical query features in a shared memory bank. This enhances tracking robustness, particularly when instruments disappear or deform. Experimental results show that the model achieves an Average Precision (AP) of 40.640, significantly surpassing the performance of other mainstream VIS models. This method effectively improves the accuracy of instrument recognition and tracking stability in ultrasound images, thereby enhancing the safety and success rate of interventional surgeries. It also contributes to the promotion of this advanced surgical technique.
Objective: Detection of mild cognitive impairment (MCI), a precursor to dementia, is critical for timely intervention. Functional near-infrared spectroscopy (fNIRS) offers non-invasive, cost-effective, and motion-tolerant brain activity monitoring, but existing machine learning approaches for MCI classification using fNIRS face two limitations: 1) underutilization of complementary information between resting-state and task-state data, and 2) high feature dimensionality relative to small sample sizes, limiting model robustness and generalizability. We propose a spatio-temporal feature engineering framework addressing these gaps.Methods: Resting-state fNIRS signals are processed via independent component analysis to derive subject-specific spatial filters, which are then clustered into a universal population-level filter set. This filter set isolates spatial features from task-state signals. Then, temporal feature selection combines variance-based and advanced methods to further reduce dimensionality by identifying discriminative task-evoked time points relevant to MCI detection. The framework integrates fNIRS spatial filtering (resting-state) and temporal selection (task-state) critical for MCI detection.Results: Validated on 104 participants, this framework achieved a single-run best of 90.91% accuracy for cognitively normal vs. MCI classification, with 91.07% feature dimensionality reduction, suggest the potential for generalizable MCI detection and efficient model retraining for expanding clinical data. Feature analysis reveals (1) universal spatial filters linked to MCI biomarkers and (2) temporal weights highlighting critical decision time points during cognitive tasks.Conclusion: By resolving the integration gap between resting-state neurovascular patterns with task-evoked hemodynamic dynamics while reducing dimensionality, the framework achieves higher accuracy and interpretability, advancing fNIRS-based MCI detection.
Dexterous manipulation, which refers to the ability of a robotic hand or multi-fingered end-effector to skillfully control, reorient, and manipulate objects through precise, coordinated finger movements and adaptive force modulation, enables complex interactions similar to human hand dexterity. With recent advances in robotics and machine learning, there is a growing demand for these systems to operate in complex and unstructured environments. Traditional model-based approaches struggle to generalize across tasks and object variations due to the high dimensionality and complex contact dynamics of dexterous manipulation. Although model-free methods such as reinforcement learning (RL) show promise, they require extensive training, large-scale interaction data, and carefully designed rewards for stability and effectiveness. Imitation learning (IL) offers an alternative by allowing robots to acquire dexterous manipulation skills directly from expert demonstrations, capturing fine-grained coordination and contact dynamics while bypassing the need for explicit modeling and large-scale trial-and-error. This survey provides an overview of dexterous manipulation methods based on imitation learning, details recent advances, and addresses key challenges in the field. Additionally, it explores potential research directions to enhance IL-driven dexterous manipulation. Our goal is to offer researchers and practitioners a comprehensive introduction to this rapidly evolving domain.
Accurate segmentation of multi-branched blood vessels from Digital Subtraction Angiography (DSA) images is essential to improve efficiency and safety of vascular interventional procedures. However, the high-speed flow of contrast agents may lead to incomplete visualization of multiple blood vessel branches and unclear boundary contours, resulting in the so-called partial labeling issue. This significantly undermines the network's ability to extract and understand the features of multi-branched vascular structures with uncertain region, thereby severely impairing the accuracy of the recognition results. In this paper, we introduce a novel pseudo-label guided multi-task learning strategy, capable of effectively learning feature representation completion under partial label supervision. Specifically, a pretext task branch that generates boundary pseudo-label signals extracts absence structural information and transfers it to the target task for multi-branch vascular segmentation. To achieve more precise semantic-supplementing between tasks, an affinity-based criss-cross feature propagation (CCFP) module is designed to dynamically fill semantic and structural voids caused by missing categories. Furthermore, to mitigate performance degradation caused by unreliable pseudo-labels, a unique loss function is proposed to constrain redundant information at both the pixel and structural levels. We validate our approach through the creation of an in-house DSA dataset composed of six sub-datasets, each containing different vascular branches. Extensive experimental results demonstrate that our method not only addresses the challenge of partial labeling but also strikes a balance between pixel-wise accuracy and the preservation of structural integrity, offering potential value in the field of clinical applications. Note to Practitioners-Abdominal multi-branch vascular segmentation from DSA images plays a crucial role in enhancing computer-assisted interventional surgery. However, due to the high labor costs and specialized expertise required, collecting large-scale DSA vascular dataset with multiple annotations at pixel level is challenging. Practically, the collected datasets are typically annotated for the segmentation task of only a single type of vessel, while unrelated categories are marked as background, leading to the problem of partial labeling. Importantly, in clinical practice to improve autonomy in interventional surgery, a unified segmentation model for multi-branch vessels is critical. Considering all these practical needs to medical applications and benefit healthcare, we proposed a pseudo-label guided multi-task learning framework for abdominal multi-branch vascular segmentation under the supervision of these partially annotated images. Extensive experiments show the superiority of our method on multiple testing categories. Besides, our approach can be adapted into semi-automatic annotation software, utilizing partially annotated datasets to generate complete annotations, and requiring only minimal fine-tuning to build large-scale, fully annotated datasets. This significantly reduces both the labor and time costs associated with manual labeling. Consequently, our method holds great potential for broad application across various medical image datasets.
Multimodal image registration is a crucial prerequisite for the automation and intelligence of interventional surgical medical robots. In endovascular aneurysm repair, due to limitations in imaging principles and hemodynamic effects, single-frame DSA images often fail to provide a complete representation of the vascular structure. This is particularly true for blood vessels that run parallel to the X-ray beam, as they are difficult to visualize in the DSA images. To address this issue, this study proposes an abdominal aortic vessel registration network, HDCAR, based on preoperative CTA 3D vascular models and intraoperative DSA images, aiming to enhance vascular completeness and spatial consistency in intraoperative imaging. The HDCAR network integrates multiple optimization modules to improve registration accuracy and robustness. First, the K-Sample module is employed to filter DSA images, enhancing the uniformity of intra-vascular structures and improving contrast between vessels and surrounding tissues. Second, depth information is incorporated to strengthen cross-dimensional spatial feature fusion, thereby optimizing the alignment between preoperative 3D models and intraoperative 2D images. Additionally, the network utilizes a dual-rectangular-window-based cross-attention mechanism and the RankC module to enhance both global contextual relationships and local feature representations. The ASPP module is further employed to extract multi-scale feature information, improving the model’s ability to capture vascular structures. Finally, a two-stage hybrid loss function is applied to optimize network parameters, ensuring precise and stable image registration. Experimental results demonstrate that the HDCAR network achieves high-precision vascular registration across multi-modal images, significantly improving the completeness and accuracy of intraoperative vascular imaging. This provides more precise imaging support for endovascular aneurysm repair procedures and holds great potential for clinical applications.
The fusion of electroencephalography (EEG) and electromyography (EMG) holds significant potential for clinical motor intention decoding. However, existing methods are constrained by the dual challenges of data scarcity and high physiological heterogeneity across patient populations. To address these limitations, this article proposes the brain-muscle-based motor planning to execution network (BM-MP2E), a framework leveraging brain-muscle complementarity to ensure reliable decoding across the full impairment spectrum. We introduce a physiologically informed architecture that explicitly mirrors the neuromuscular transmission pathway. In particular, movement-related cortical potentials (MRCPs) for motor planning, $\alpha \beta $ bands for activation, and EMG for execution are extracted to constrain the model's solution space. Furthermore, a subject-adaptive fusion module is designed to dynamically modulate modality contributions based on individual functional status. Experimental validation on 21 patients (ranging from disorders of consciousness (DoCs) to stroke) in a unimanual five-class task demonstrates that BM-MP2E achieves an average accuracy of 60.81%, yielding a 6.94% improvement over single-modality EMG. The average learned weights were 0.265 (MRCP), 0.323 (alpha,beta), and 0.412 (EMG). Quantitative analysis reveals a strong positive correlation between adaptive EMG weights and motor function scores (r = 0.884 and p < 0.0001). This confirms a compensatory measurement logic: the system automatically prioritizes stable cortical features in severe cases while progressively leveraging high-fidelity peripheral signals as function recovers. These findings validate BM-MP2E as an adaptive clinical solution capable of facilitating motor assistance across the entire recovery spectrum.
Medical image segmentation takes an important position in various clinical applications. 2.5D-based segmentation models bridge the computational efficiency of 2D-based models with the spatial perception capabilities of 3D-based models. However, existing 2.5D-based models primarily adopt a single encoder to extract features of target and neighborhood slices, failing to effectively fuse inter-slice information, resulting in suboptimal segmentation performance. In this study, a novel momentum encoder-based inter-slice fusion transformer (MOSformer) is proposed to overcome this issue by leveraging inter-slice information from multi-scale feature maps extracted by different encoders. Specifically, dual encoders are employed to enhance feature distinguishability among different slices. One of the encoders is moving-averaged to maintain consistent slice representations. Moreover, an inter-slice fusion transformer (IF-Trans) module is developed to fuse inter-slice multi-scale features. MOSformer is evaluated on three benchmark datasets (Synapse, ACDC, and AMOS), achieving a new state-of-the-art with 85.63%, 92.19%, and 85.43% DSC, respectively. These results demonstrate MOSformer’s competitiveness in medical image segmentation.
Accurate vessel segmentation in X-ray angiograms is crucial for numerous clinical applications. However, the scarcity of annotated data presents a significant challenge, which has driven the adoption of self-supervised learning (SSL) methods such as masked image modeling (MIM) to leverage large-scale unlabeled data for learning transferable representations. Unfortunately, conventional MIM often fails to capture vascular anatomy because of the severe class imbalance between vessel and background pixels, leading to weak vascular representations. To address this, we introduce Vascular anatomy-aware Masked Image Modeling (VasoMIM), a novel MIM framework tailored for X-ray angiograms that explicitly integrates anatomical knowledge into the pre-training process. Specifically, it comprises two complementary components: anatomy-guided masking strategy and anatomical consistency loss. The former preferentially masks vessel-containing patches to focus the model on reconstructing vessel-relevant regions. The latter enforces consistency in vascular semantics between the original and reconstructed images, thereby improving the discriminability of vascular representations. Empirically, VasoMIM achieves state-of-the-art performance across three datasets. These findings highlight its potential to facilitate X-ray angiogram analysis.
Offline preference-based reinforcement learning (PbRL) offers an effective approach to addressing the challenges of designing rewards and mitigating the high costs associated with online interaction. However, since labeling preference needs real-time human feedback, acquiring sufficient preference labels is challenging. To solve this, this article proposes an offline PbRL with a high sample efficiency (LEASE ) algorithm, where a learned transition model is leveraged to generate unlabeled preference data. Considering the pretrained reward model may generate incorrect labels for unlabeled data, we design an uncertainty-aware mechanism to ensure the performance of the reward model, where only high-confidence and low-variance data are selected. Moreover, the generalization bound of the reward model is provided to analyze the factors influencing reward accuracy, and the policy learned by LEASE has a theoretical improvement guarantee. The above developed theory is based on a state-action pair, which can be easily combined with other offline algorithms. The experimental results show that LEASE can achieve comparable performance to the baseline under fewer preference data without online interaction.
In interventional surgery, using dynamic contrast imaging to evaluate the boundary morphology of abdominal aortic aneurysms is widely regarded as the gold standard. However, due to the high blood flow velocity within the abdominal aorta, single-frame images lack comprehensive vascular information, leading to uncertainty in the boundaries of multi-branch vessels. This study proposes a multi-frame fusion segmentation network for abdominal aortic vessels based on DSA image sequences, termed L-MFFUnet. The proposed method consists of a two-stage network, including a rapid multi-branch extraction network and a dynamic sequence fusion module. By integrating temporal frame encoding, MFFUnet captures and incorporates temporal information from frame sequences. The dual-rectangular attention module enhances the network’s sensitivity to vascular features. The FreqPass module improves intra-class feature uniformity and inter-class distinctiveness. Additionally, the fusion of inter-frame optical flow information ensures the temporal continuity of vascular structures. Meanwhile, the LAF module, which combines LSTM and attention mechanisms, dynamically captures subtle pulsations and achieves precise fusion of different vascular branches. Extensive experimental results demonstrate that the proposed method exhibits excellent vascular segmentation performance in high dynamic contrast imaging environments, highlighting its significant potential for clinical applications.
Maneuverability is a critical factor in the clinical advantages of magnetic guidewires (magwires) for minimally invasive vascular interventions. However, the absence of standardized evaluation metrics often results in inconsistent or subjective assessments. This study proposes the Manipulation Degree, a novel metric that quantitatively evaluates maneuver-ability of magwires in an objective and standardized manner. The Manipulation Degree is defined by two key factors: the workspace area, which is the total surface area reachable by the magwires tip, and instability. It exhibits a positive correlation with workspace area and a negative correlation with instability, offering a concise evaluation framework. Experimental validation using three representative magwires demonstrates that increasing the number of magnetic segments and decreasing the magwire’s diameter increase the Manipulation Degree. Furthermore, magwires exhibiting a high Manipulation Degree in less constrained scenarios maintain this advantage in highly constrained scenarios. A strong linear correlation (R2 ≥ 0.99) is also observed between the Manipulation Degree and both delivery time and translation force, highlighting its predictive value for clinical performance. These findings establish the Manipulation Degree as a robust and practical metric for assessing the maneuverability of magwires, offering insights for the design optimization of magwires.
Analyzing the characteristic waves in electrocardiograms is a crucial method for diagnosing heart diseases. Current methods for detecting characteristic waves primarily focus on feature matching based on the morphology of these waves or subsequences of their morphology. However, individual differences and diseases can cause specific changes in the characteristic waves, making it challenging for feature matching schemes to effectively detect these waves under various disease conditions. Although deep learning has shown great potential in efficiently extracting nonlinear features of characteristic waves in the presence of diseases, its reliability remains a concern for healthcare professionals.To address these challenges, we propose an adaptive waveform correction filter to correct and enhance the ECG waveform. By reducing the specificity between patients and diseases, the data distribution of the characteristic waves is optimized, enabling bidirectional matching between the waves and their detection methods. This, in turn, improves the stability of detection. The filtered results are then fed into an Interpretable Temporal Modeling Neural Network (IT-Net) that we propose in this paper. IT-Net can detect characteristic waves in the corrected and enhanced standard 12-lead ECGs, and its temporal modeling approach more effectively captures the sequential dependencies among the characteristic waves. The temporal dependencies identified by IT-Net ensure its reliability in waveform segmentation.Experiments conducted on a public dataset demonstrate that the mean and variance of detection errors for the P wave, QRS complex, and T wave are as follows: Pon is-2.5 +/- 13.9, Pend is 0.48 +/- 11.3, QRSon is 0.04 +/- 7.2, QRSend is 0.2 +/- 7.0, Ton is-3.6 +/- 16.7, and Tend is 1.2 +/- 14.3, proving the stability of the proposed filter and IT-Net. Additionally, we validated the performance of IT-Net on a clinical dataset, and the results show that the method can effectively detect characteristic waves under disease conditions.
Predicting knee joint trajectories from surface electromyography (sEMG) signals holds a significant application value in various fields such as rehabilitation engineering and prosthetics control. However, existing prediction methods often struggle to achieve satisfactory performance due to limited dataset sizes and poor cross-subject generalization capabilities. In this paper, we propose an effective framework that integrates motion decoupling with a conditional diffusion model to address these challenges. Our approach decomposes knee joint angles into shared motion patterns across subjects and individual-specific amplitude parameters, enabling dual-task collaborative modeling that considers both commonalities and individual differences. Furthermore, the conditional diffusion model is employed to generate high-quality synthetic sEMG samples, effectively expanding the available data resources. Experiments conducted on data from 11 subjects demonstrate that our approach achieves a Root Mean Square Error ( RMSE ) of 3.59 ± 0.88 ^∘ , outperforming the non-decoupled model (4.61 ± 1.58 ^∘ ), the model without diffusion (4.85 ± 1.62 ^∘ ), the Bidirectional Long Short-Term Memory (Bi-LSTM) (6.75 ± 1.33 ^∘ ) and the traditional LSTM baseline (6.88 ± 1.59 ^∘ ).
Automatic vessel segmentation plays a pivotal role in the development of next-generation interventional navigation systems for surgical robotics. However, current approaches still suffer from suboptimal segmentation performance under challenging intraoperative conditions, such as low-signal-to-noise ratio (SNR), small or slender vessels, and strong interference. In this study, a novel SPatial-frequency learning and graph-based channel InteRactiOn Network (SPIRONet) is proposed to address the above issues. To overcome low-SNR vessel appearance and small or slender branches, dual spatial and frequency encoders are utilized, where the frequency encoder captures global vessel continuity that is less affected by local noise fluctuations, while the spatial encoder preserves fine vessel details. A cross-attention fusion module is further introduced to adaptively integrate this complementary spatial and frequency information. Moreover, to suppress interference from non-target vessels and vessel-like structures, a graph-based channel interaction module is designed to model channel-wise correlations, enhancing consistent vessel-related responses while suppressing task-irrelevant activations. Extensive experimental results on five challenging datasets demonstrate that the proposed method achieves competitive and consistently strong performance compared with existing methods. For example, SPIRONet achieves IoU improvements of +0.87%, +0.52%, +0.23%, +1.39%, and +2.22% over the strongest competing methods on CADSA, CAXF, DCA1, XCAD, and ARCADE, respectively. Moreover, SPIRONet achieves an inference speed of 21 FPS with a 512 & times; 512 input size, meeting the real-time requirements of interventional scenarios (6-12 FPS). These promising results indicate SPIRONet's potential for integration into interventional navigation systems. Code is available at https://github.com/Dxhuang-CASIA/SPIRONet.