To address force signal distortion and neural injury risks caused by respiratory-induced spinal displacement coupled with system control delays (>200 ms) in robot-assisted laminectomy surgery, this study proposes a Swin Transformer-based Prediction Network (STP-Net). STP-Net replaces the standard Transformer with a shifted window multi-head self-attention mechanism (alternating W-MSA and SW-MSA), reducing computational complexity from O(n(2)) to O(n).The encoder-decoder architecture incorporates hierarchical downsampling and future-masking to extract global features from long sequences. Experimental results demonstrate that for 150-frame inputs predicting 75 frames (approximate to 300 ms), STP-Net achieves a prediction error (MSE =0.145 mm, MAE =0.074 mm) - 76% and 66% lower than LSTM and Informer, respectively. With a per-frame inference latency of 8.4 ms, the total closed-loop delay is < 25 ms. In the closed-loop respiratory compensation experiment, the motion trajectories of the active and passive arms showed an exceptional correlation (r = 0.996), while end-effector contact force fluctuations were reduced to merely 1.6 N. This approach achieves high-accuracy and low-delay respiratory compensation in spinal surgery, providing a critical technical foundation for enhancing the safety of autonomous surgical robots.
Spinal surgical robots help improve the safety of spinal surgery and are expected to address the issue of uneven development of medical resources in various regions. The localization of spinal anatomical landmarks in CT images is an important basis for achieving autonomous planning of surgical robot paths. Manual selection of spinal anatomical landmarks often involves a significant degree of subjectivity, and doctors with insufficient experience often find it difficult to ensure the accuracy of landmark selection. This paper proposes a two-stage method for localizing spinal segment anatomical landmarks in CT images under arbitrary views. First, based on 3D MFS-Net, precise localization of the spatial regions of each spinal segment in CT images is achieved. On this basis, 3D MSH-Net is used to precisely determine the anatomical landmarks of each spinal segment. The average 3D IOU of 3D MFS-Net for spinal segment spatial region localization can reach 0.91, and the average positioning error of 3D MSH-Net for anatomical landmarks is only 0.537mm. Both 3D MFS-Net and 3D MSH-Net use lightweight design, which can perform rapid reasoning at speeds of 38.41FPS and 10.28FPS, respectively.
Motion artifacts present in magnetic resonance imaging (MRI) can seriously interfere with clinical diagnosis. Removing motion artifacts is a straightforward solution and has been extensively studied. However, paired data are still heavily relied on in recent works and the perturbations in k-space (frequency domain) are not well considered, which limits their applications in the clinical field. To address these issues, we propose a novel unsupervised purification method which leverages pixel-frequency information of noisy MRI images to guide a pre-trained diffusion model to recover clean MRI images. Specifically, considering that motion artifacts are mainly concentrated in high-frequency components in k-space, we utilize the low-frequency components as the guide to ensure correct tissue textures. Additionally, given that high-frequency and pixel information are helpful for recovering shape and detail textures, we design alternate complementary masks to simultaneously destroy the artifact structure and exploit useful information. Quantitative experiments are performed on datasets from different tissues and show that our method achieves superior performance on several metrics. Qualitative evaluations with radiologists also show that our method provides better clinical feedback.
Laminectomy represents an effective surgical procedure for the treatment of lumbar spinal stenosis. Due to the intricate anatomical structure of the lumbar spine, meticulous surgical path planning is essential to ensure the safety of the procedure and enhance the likelihood of successful outcomes. This study aims to implement multi-objective optimization techniques in the context of laminectomy, with a particular emphasis on identifying the optimal reference cutting path for the lamina. In clinical practice, the cutting path is typically characterized as a relatively straight line, akin to making an incision through the lamina with a sharp, rigid plane. Consequently, the optimal reference cutting path can be established by determining the ideal reference cutting plane. In our methodology, the cutting contour, defined as the intersection of the cutting plane with the lamina, is treated as a variable. Key features of the lamina are extracted and classified into three objective functions: the average thickness of the lamina, the derivative of the entry point of the cutting path, and the degree of overlap between adjacent vertebrae. We then apply multi-objective optimization algorithms and utilize the weighted sum method to solve the multi-objective problems in the laminectomy task. The experimental findings are validated by an enhanced laminectomy plane evaluation system, demonstrating that the automatically generated cutting planes achieve a high level of excellence (94%), thereby satisfying the surgical requirements validated by professional surgeons.
BACKGROUND:Robotic-assisted unilateral biportal endoscopic surgery (UBE) is a more accurate and safer technique than traditional open surgical operations. The penetration recognition of ultrasonic drilling remains one of the challenging techniques of robotic-assisted UBE surgery. METHODS:We propose a force and VAE-MLP-based method for real-time penetration recognition. During the ultrasonic drilling procedure, the force signals are collected and denoised via Kalman filtering first. The pre-processed data are then used to extract hidden features and perform classification by Variational Autoencoder (VAE) and Multilayer Perceptron (MLP), respectively, ultimately achieving real-time penetration recognition. RESULTS:Our method achieves superior accuracy (99.32% vs. 95.90%) and faster inference speed (17 vs. 33 ms) compared to the classic time-series classification algorithm. Robotic ex vivo bone experiments further validated its efficacy. CONCLUSION:The force and VAE-MLP framework enables fast and accurate penetration detection, which offers a reliable and efficient solution for minimizing nerve damage in UBE surgery.
Small object detection in uncrewed aerial vehicle (UAV) images is one of the critical aspects for its widespread application. However, due to limited feature extraction for small objects and complex backgrounds, there remain significant issues of missed detections and false alarms. This article proposes a real-time small object detection network for UAV images based on cross spatial frequency domain and position relation (CSFPR-RTDETR). First, we propose a cross-spatial-frequency domain hybrid (CSFH) feature extraction network, which incorporates frequency-domain processing based on the CSP network to effectively capture global contextual features and enhance the distinction between small objects and backgrounds. Second, we propose a position relation decoder that incorporates the two novel geometric priors: IoU and relative angle. Through rational characterization of spatial correlations, this design significantly strengthens the spatial perception capability of the model, thereby improving the detection performance for densely distributed small objects. Finally, we design an efficient small-object high-frequency hybrid encoder, integrating the P2 detection head and proposing a mixed high-frequency enhancement fusion module (MHE-Fusion) to extract fine-grained high-frequency features of small objects, further boosting detection performance. The experimental results demonstrate that CSFPR-RTDETR achieves superior performance on the VisDrone, AI-TOD, and HIT-UAV datasets, with mAP50 metrics reaching 42.3%, 55.4%, and 83.1%, respectively, which is better than other SOTA models. Compared to RT-DETR, CSFPR-RTDETR reduces the parameters of the network by 29.1% while significantly enhancing detection performance: the mAP50 metrics reach notable improvements of 4.6%, 4.4%, and 1.5% on the three datasets, respectively. The source code is available at https://github.com/HuLei-JXNU/CSFPR-RTDETR
The calibration of ultrasound probes is essential for three-dimensional ultrasound reconstruction and navigation. However, the existing calibration methods are often cumbersome and inadequate in accuracy. In this paper, a hybrid mathematical model, Dimensionality Reduction and Homography Transformation (DRHT), is proposed. The model characterizes the relationship between the image plane of ultrasound and projected calibration lines and homography transformation. The homography transformation, which can be estimated using the singular value decomposition method, reduces the dimensionality of the calibration data and could significantly accelerate the computation of image points in ultrasonic three-dimensional reconstruction. Experiments comparing the DRHT method with the PLUS library demonstrated that DRHT outperformed the PLUS algorithm in terms of accuracy (0.89 mm vs. 0.92 mm) and efficiency (268 ms vs. 761 ms). Furthermore, high-precision calibration can be achieved with only four images, which greatly simplifies the calibration process and enhances the feasibility of the clinical application of this model.
To mitigate the limited texture fidelity and perceptual realism of Real-ESRGAN on infrared imagery, we propose an enhanced discriminator: GFCSDiscriminator, which is equipped with a Gated Fine-Grained Channel-Spatial Module (GFCS-Module). The module contains a Fine-Grained Channel-Spatial Attention Block (FCSA Block) that jointly models global-local channel dependencies and spatial context, dynamically reallocating feature weights to emphasize minute textures and edges. An Attention Gate is further introduced to strengthen semantic coherence between high- and low-level features. Experiments on an infrared dataset show that the proposed discriminator surpasses the original Real-ESRGAN in Natural Image Quality Evaluator (NIQE) and Learned Perceptual Image Patch Similarity (LPIPS), delivering more natural texture restoration.
Infrared images have less available information compared to visible images, and the applying of high-frequency details and edge information can directly influence the quality of super-resolution (SR) reconstruction of infrared images. However, most existing SR methods have a single activation mode for high-frequency features and over-dependently increase the network depth to improve performance. To address these problems, we design a variable GELU (VGELU), which introduces a learnable parameter a based on GELU to suppress low-frequency features and noise by adaptively changing the slope of GELU in high-frequency feature extraction. In addition, we propose an attention-enhanced CATS-RCF (ACR) network in the strong edge feature extraction module (SEFEM), which introduces coordinate attention based on CATS-RCF (CR) to enhance the edge weights of infrared low-resolution (LR) images and improve the effect of edge extraction. To fully fuse high-frequency features and edge information, we further design an edge feature fusion block (EFFB), which effectively fuses edge information from different dimensions. Our edge-enhanced and variable activation network (EVAN) is constructed by applying the proposed VGELU, SEFEM with EFFB. The comprehensive experiments demonstrate the superiority of our EVAN over other comparison methods.
The accuracy of 3D-printed guides alignment during robot-assisted spinal surgery is significantly influenced by the size of the contact area for registration. An undersized contact area may result in unstable registration, whereas an oversized contact area can increase the dissection area of soft tissues. This paper introduces a novel medium contact (MC) guide design procedure based on a four-point alignment method, which is specifically tailored for posterior laminectomy procedures. A comparative finite element analysis against the widely utilized full contact (FC) guides demonstrated that the MC guide, which occupies only 36.28% of the area of the FC guide, shows lower stress concentration and displacement when subjected to external forces according to the ABAQUS simulation environment. To verify the clinical efficacy of the MC guide, an in vivo animal experiment was conducted. The MC guide, post-printing, was employed for alignment, and successfully complete the registration of a robot-assisted spinal decompression surgery. The MC guide proposed herein not only ensures registration stability but also minimizes patient injury and surgeon’s workload, and offers considerable reference value in clinical applications.
OBJECTIVE This study aimed to introduce a novel artificial intelligence (AI)-based robotic system for autonomous planning of spinal posterior decompression and verify its accuracy through a cadaveric model. METHODS Seventeen vertebrae from 3 cadavers were included in the study. Three thoracic vertebrae (T9-11) and 3 lumbar vertebrae (L3-5) were selected from each cadaver. After obtaining CT data, the robotic system independently planned the laminectomy path based on AI algorithms before the surgical procedure and automatically performed the decompression during the procedure. A postoperative CT scan was performed, and the deviation of each cutting plane from the preoperative plan was quantitatively analyzed to evaluate the accuracy and safety of the cuts. The duration of laminectomy was also recorded. RESULTS A total of 285 cuts were made on thoracic and lumbar vertebrae. The average duration for unilateral longitudinal cutting was 16.38 +/- 4.76 minutes, while for transverse cutting it was 4.44 +/- 1.52 minutes. In terms of accuracy assessment, 3 levels were divided based on the distance between the actual cutting plane and the preplanned plane: 77 (84%) were grade A, 15 (16%) were grade B, and none were grade C. Regarding safety assessment, 74 (80%) were designated safe (grade A), with 18 (20%) classified as uncertain (grade B). CONCLUSIONS The results confirm the accuracy and preliminary safety of the robotic system for autonomous planning and cutting of spinal decompression.
The YOLOx-s network does not sufficiently meet the accuracy demand of equipment detection in the autonomous inspection of distribution lines by Unmanned Aerial Vehicle (UAV) due to the complex background of distribution lines, variable morphology of equipment, and large differences in equipment sizes.Therefore, aiming at the difficult detection of power equipment in UAV inspection images, we propose a multi-equipment detection method for inspection of distribution lines based on the YOLOx-s.Based on the YOLOx-s network, we make the following improvements: 1) The Receptive Field Block (RFB) module is added after the shallow feature layer of the backbone network to expand the receptive field of the network.2) The Coordinate Attention (CA) module is added to obtain the spatial direction information of the targets and improve the accuracy of target localization.3) After the first fusion of features in the Path Aggregation Network (PANet), the Adaptively Spatial Feature Fusion (ASFF) module is added to achieve efficient re-fusion of multi-scale deep and shallow feature maps by assigning adaptive weight parameters to features at different scales.4) The loss function Binary Cross Entropy (BCE) Loss in YOLOx-s is replaced by Focal Loss to alleviate the difficulty of network convergence caused by the imbalance between positive and negative samples of small-sized targets.The experiments take a private dataset consisting of four types of power equipment: Transformers, Isolators, Drop Fuses, and Lightning Arrestors.On average, the mean Average Precision (mAP) of the proposed method can reach 93.64%, an increase of 3.27%.The experimental results show that the proposed method can better identify multiple types of power equipment of different scales at the same time, which helps to improve the intelligence of UAV autonomous inspection in distribution lines.
Image registration is an important prerequisite for image fusion, and its accuracy will affect the quality of fused images. Visible images are enhanced with detailed textures, but their quality is heavily influenced by lighting conditions; infrared images are not affected by illumination and distance, which makes the two complement-ary images. However, the disparate imaging mechanism leads to a great difference between the two images, so it is difficult to register them. This paper introduces MTIVRNet, a network for registering infrared and visible images using modal transformation as its foundation. CycleGAN is a framework used to train two sets of images, such as infrared and visible images, by transforming one modality to another. In this case, it is used to convert visible images into pseudo-infrared images. The aim is to bridge the spectral difference and make the images appear similar in both modalities. Then, the OD-Dense module proposed in this paper is used for feature extraction, which is mainly composed of Omni-dimensional Dynamic Convolution (ODConv). By utilizing ODConv, the model benefits from enhanced feature extraction capabilities without significantly increasing the number of parameters. The experimental results show that the evaluation metric PCK of MTIVRN et on the infrared and visible images dataset FLIR is 9.48% higher than the basic network when α=0.1, and superior to other comparison methods in vision.
Percutaneous transforaminal endoscopic discectomy (PTED) is a decompression surgery on patients with lumbar disc herniation and spinal canal stenosis in a minimally invasive environment, which greatly shorten the rehabilitation cycle of patients. However, the puncture of traditional PTED is performed under non-direct vision, which relies on the surgeon’s clinical experience heavily and can easily cause collateral damage. In this paper, we propose a robotic positioning method in PTED based on X-ray image and DLT algorithm. The end-effector with three rings (EETR) for PTED positioning has been specially designed. During the operation, the 3D-2D transformation matrix is calculated by direct linear transform (DLT) algorithm. Finally, the motion parameters are calculated to control the robot. Based on our method, only one X-ray image is needed to complete the positioning, which also works well in a large range of deflection angle between EETR and target puncture channel. Thus, it greatly simplifies the surgical process and reduces the radiation exposure time. By conducting a series of comparative experiments on the positioning of model bones, the translation error is less than l.lmm and rotation error is less than 0.8°, when the deflection angle is less than 30°. And the translation error is 1.61mm and rotation error is 1.98°, when the angle reached 60°. It meets the requirements for PTED positioning, and has improved the positioning accuracy.
For robot-assisted pelvic fracture reduction, at least two bone needles need to be inserted into the ilium of the affected pelvis, and the robot clamping device is connected with the bone needles. The biomechanical properties of the pelvic musculoskeletal tissues are different with the different Spatial Position and Orientation (SPO) of the bone needles. In order to determine the optimal SPO of bone needle pairs, the constraints between the bone needles and the pelvis are analyzed, and the SPO vectors of 150 groups bone needles are obtained by the KNN-hierarchical clustering method; a batch modeling method of bone needles with different SPO is proposed. 150 finite element models of damaged pelvic musculoskeletal tissue with different SPO of bone needles are established and simulated. The stress and strain distribution homogenization of musculoskeletal tissue with bone needles as evaluation index, the simulation results of 150 models are evaluated. Results show that, the anterior superior iliac spine and the anterior inferior iliac spine are suitable regions to place bone needles in the pelvis, and the optimal distribution of the needle combination is found in this region. The overall stress and strain distribution of the damaged pelvic musculoskeletal tissue under the large reduction force is the best.
High-precision image segmentation of the spine in computed tomography (CT) images is important for the diagnosis of spinal diseases and surgical path planning. Manual segmentation is often tedious and time consuming. Thus, an automatic segmentation algorithm is expected to solve this problem. However, because different areas are scanned, the number of spines in the original CT image and the coverage area are often different, making it extremely difficult to directly conduct a fully autonomous spine segmentation. In this study, we propose a two-stage automatic spine segmentation method based on 3D Swin Transformer. In the first stage, the 3D Swin-YoloX algorithm is used to achieve an accurate positioning of each spine segment in the CT images. In the second stage, 3D Swin-UNet is used to achieve a high-precision segmentation of the spine. Using an open dataset, the average Dice of our approach can reach 0.942 and the average Hausdorff distance can reach 6.24, indicating a higher accuracy in comparison with other published methods. Our proposed method can effectively eliminate any adverse effects of the different scanning areas on a spinal image segmentation and has a high application value.
Robotic end effector accurately operated in small and confined surgical area in laminectomy can improve the efficiency of robotic surgery. This paper proposed a two-DOFs planar grinding end-effector for robotic laminectomy, which can move along any curve within the planar operation area. The end-effector adopts an arc RCM mechanism, which kinematic analysis gives a working space of an annular sector with a minimum radius of 137 mm, a maximum radius of 162.3 mm, a central angle of 1.65 rad. The structure rationality of the end-effector is verified through static analysis. The prototype was developed and controlled based on TwinCAT-PC architecture. Experiments of positioning accuracy give a repeated positioning accuracy of 0.1 mm. Experiments of ex vivo bone grinding gives an operation accuracy of 1.5 mm in a simulated clinical environment. The proposed prototype can be used for robotic laminectomy in combination with the navigation robotic arm.
Objective: This study aims to use artificial intelligence to realize the automatic planning of laminectomy, and verify the method. Methods: We propose a two-stage approach for automatic laminectomy cutting plane planning. The first stage was the identification of key points. 7 key points were manually marked on each CT image. The Spatial Pyramid Upsampling Network (SPU-Net) algorithm developed by us was used to accurately locate the 7 key points. In the second stage, based on the identification of key points, a personalized coordinate system was generated for each vertebra. Finally, the transverse and longitudinal cutting planes of laminectomy were generated under the coordinate system. The overall effect of planning was evaluated. Results: In the first stage, the average localization error of the SPU-Net algorithm for the seven key points was 0.65mm. In the second stage, a total of 320 transverse cutting planes and 640 longitudinal cutting planes were planned by the algorithm. Among them, the number of horizontal plane planning effects of grade A, B, and C were 318(99.38%), 1(0.31%), and 1(0.31%), respectively. The longitudinal planning effects of grade A, B, and C were 622(97.18%), 1(0.16%), and 17(2.66%), respectively. Conclusions: In this study, we propose a method for automatic surgical path planning of laminectomy based on the localization of key points in CT images. The results showed that the method achieved satisfactory results. More studies are needed to confirm the reliability of this approach in the future.
目的 测量不同脊柱组织的电阻抗,基于支持向量机建立电阻抗数据的组织分类算法并验证算法的准确性,寻找不同组织电阻抗分类阈值.方法 取离体脊柱组织,应用电化学分析仪采集10~100 kHz频率范围内皮质骨、松质骨、脊髓、肌肉、髓核的电阻抗.将两只猪采集的数据集分别作为训练集和测试集,应用主成分分析降维至二维数据,训练和验证基于支持向量机(SVM)建立的分类算法,应用集成学习的方法计算不同组织分类的电阻抗阈值.结果 5种组织在10~100 kHz的测量频率内,电阻抗值差异有统计学意义(P<0.001).应用主成分分析降维的数据集建立的支持向量机分类算法识别不同组织的准确率为100%.应用集成学习建立的多个分类器计算出了不同组织的电阻抗分类阈值.结论 基于支持向量机可以实现脊柱术区组织电阻抗的准确识别,有望应用于临床协助医生提升组织识别准确率.