Accurate and automated Cobb angle measurement is essential for the diagnosis and management of Adolescent Idiopathic Scoliosis (AIS). However, existing deep learning approaches often ignore the intrinsic anatomical dependencies of the spine, which results in physiologically implausible estimations. We propose a two-stage anatomical perception and reasoning framework based on keypoint detection for Cobb angle measurement. The anatomical perception stage is built upon a Local-to-Global Aggregation Backbone (LGAB), which combines convolutional neural networks and Transformers to jointly capture local anatomical details and global spinal context. The resulting feature representations are further refined by a Semantic Guidance Module (SGM) to facilitate cross-scale feature fusion and improve vertebral keypoint localization. In the second stage anatomical reasoning, an Anatomical Reasoning Network (ARN) models all detected vertebral center keypoints as a structured graph. By embedding anatomical priors and applying dynamic relational reasoning, the ARN refines the spatial configuration of keypoints, which ensures anatomical coherence and corrects subtle localization errors. The optimized keypoints are then used to calculate the Cobb angle. The proposed framework achieves a Symmetric Mean Absolute Percentage Error (SMAPE) of 6.61 Error_Center ) of 24.58 pixels on the public AASCE 2019 Challenge dataset. Furthermore, the framework achieves a comparable CMAE of 2.11° on an independent clinical dataset. These results demonstrate that the proposed method provides a reliable foundation for precise and automated Cobb angle measurement.
The accurate measurement of the Cobb angle in spinal radiographs is of great significance for the diagnosis and treatment decision of adolescent idiopathic scoliosis. To avoid variability in Cobb angle measurement, this paper develops a U-Net-based model: Multi-scale Wavelet Convolution-Enhanced Cobb angle Estimation framework (MWCE-CE), which segments vertebrae for Cobb angle measurement. The core of this framework is a segmentation network named WA-SEUNet, which is specifically designed to improve vertebrae segmentation accuracy. First, a Wavelet residual feature extraction module is used to enhance the model's ability to recognize specific vertebrae textures and resist noise at the encoder stage. Second, a dual-stream channel attention module is used to integrate the local to global information at the decoder stage. Therefore, the network can accurately extract key features leading to the accurate segmentation of vertebrae, from which the Cobb angle is calculated. We conduct extensive experiments on the publicly available AASCE 2019 challenge dataset. The average Dice similarity coefficient is 0.9432, and the accuracy reaches 0.9833 in the segmentation task. The circular mean absolute error is 2.24 degrees in the Cobb angle measurement task. Compared with the state-of-the-art methods, the results demonstrated the superiority of the proposed method for scoliosis evaluation.
Objective: The biomechanical properties of the lumbar spine is crucial for assisting the diagnosis, treatment, and prevention of spinal diseases. Traditional biomechanical analysis methods, especially the finite element analysis, require extensive computational resources, precise material property definitions, and complex meshing processes to accurately model the biomechanical behavior of the lumbar spine. While deep learning is introduced to enhance efficiency and accuracy, challenges like data dependency and lack of physical consistency remain. Methods: We propose a novel framework that consists of a 3D generative adversarial network for data augmentation together with a dual-channel vision transformer to extract geometric and physical information. We also introduce a physics-guided mechanism into training phase, ensuring model consistency with mechanical principles. Results: The proposed method achieved an Intersection over Union of 0.8332 and a Mean Squared Error of 0.0002. The five vertebrae of the lumbar spine are processed in 87 milliseconds, which is approximately 3000 times faster than traditional finite element methods. Conclusion: Our framework demonstrates high accuracy and substantial computational efficiency, offering a reliable alternative to conventional biomechanical modeling. Significance: This enables real-time lumbar spine analysis for diagnosis, surgical planning, and personalized treatment.
Three-dimensional (3D) reconstruction of the spine from X-ray images is of great significance for the diagnosis of diseases such as scoliosis. However, traditional methods are time-consuming and laborious. Most existing deep learning methods focus on 3D reconstruction from biplanar or multi X-ray images. Compared with these reconstructions, the reconstruction from a single image is more difficult, but the cost and radiation are lower. In this study, the spinal reconstruction network (SR-Net) with encoder-decoder architecture is developed to reconstruct the 3D spinal model from a single X-ray image. First, the wavelet feature extraction module is proposed to address the interference of noise and artifacts in the X-ray image. Second, we design a dimension transformation module to convert the two-dimensional features extracted by the encoder into 3D features. Finally, the feature fusion module is developed for skip connections between the encoder and decoder. The experimental results on two datasets demonstrate that our method outperforms other state-of-the-art methods. In addition, experimental results on the real X-ray images show that the SR-Net has the potential for clinical applications.
BACKGROUND:Understanding spinal biomechanics is essential for exploring the functions of the spine and the pathogenesis of related diseases. Traditional numerical methods for biomechanical analysis are computationally expensive, while Physics-Informed Neural Network (PINN) struggles with complex solid geometries. This study develops an Enhanced Physics-Informed Neural Network that integrates Geometric Features (EPINN-GF) to address these limitations in predicting stress distribution for the complex spinal geometries. STUDY OBJECTIVES:The primary objective of this study is to improve both the accuracy and efficiency of predicting stress distributions in the complex spinal geometries using the proposed EPINN‑GF framework. The secondary objective is to facilitate the clinical translation to support the diagnosis, surgical planning, and personalized treatment for the spinal disorders, particularly Adolescent Idiopathic Scoliosis (AIS). METHODS:This study proposes EPINN-GF, which innovatively integrates geometric features and dual loss constraints. Specifically, the eigenvectors of the Laplacian operator, derived from spinal geometry coordinates, are utilized to capture both global and local structural characteristics of the spine. These eigenvectors, along with their corresponding normalized coordinates, serve as inputs to EPINN-GF. Furthermore, the equilibrium equations, which describe the balance between internal and external forces acting on the material, are embedded in the network's loss function. RESULTS:Experimental results demonstrate the superiority of EPINN-GF in spinal stress prediction, with a mean squared error (MSE) of 227.59 MPa compared to 398.41 MPa, 271.43 MPa, and 343.61 MPa for Collocation-based PINN (C-PINN), Graph Neural Network (GNN), and Deep Energy Method (DEM), respectively. Despite the training time of 1917.71 s for the EPINN-GF, which is slightly longer than those of 1713.03 s and 1883.68 s for C-PINN and GNN, respectively, it achieves higher stress prediction accuracy, making it a promising tool for spinal disease diagnosis and modeling complex physical systems. CONCLUSIONS:EPINN-GF accurately predicts stress in the geometrically complex spinal regions and multi-physical environments, offering the potential clinical value for the spinal disorders by enabling personalized treatment planning, optimizing surgical strategies, and supporting early AIS progression prediction.
BACKGROUND:To investigate the clinical implications of serum carcinoembryonic antigen (CEA), neuron specific enolase (NSE), and pro-gastrin-releasing peptide (ProGRP) in small cell esophageal carcinoma (SCEC). METHODS:Receiver operating characteristic (ROC) curves were used to determine the area under the curve (AUC), sensitivity, and specificity of serum markers for differentiating SCEC from patients with esophageal squamous carcinoma (ESCC), esophageal adenocarcinoma (EAC) and healthy subjects. RESULTS:The combination of ProGRP and NSE demonstrated significant diagnostic efficacy in distinguishing SCEC from ESCC, EAC, and healthy individuals. After treatment, serum levels of ProGRP, NSE, and CEA in SCEC patients with disease control decreased significantly compared to pre-treatment levels. Conversely, the serum levels of ProGRP and NSE in patients with disease progression after treatment were significantly increased compared to those before treatment. Compared with those in the follow-up phase, the levels of ProGRP, NSE, and CEA significantly increased after tumor progression in SCEC patients. This study further exhibited that serum levels of ProGRP above 45.15 pg/mL and NSE above 14.70 ng/mL during follow-up after first-line treatment in SCEC patients were associated with poor progression-free survival (PFS). Moreover, significant associations were observed between baseline serum concentrations of ProGRP, NSE, CEA and both PFS and cancer-specific survival (CSS) in SCEC patients. CONCLUSIONS:ProGRP and NSE have significant value in the diagnosis and differential diagnosis of SCEC, and are also associated with the clinical stage and prognosis of SCEC patients, as well as are effective indicators for evaluating therapeutic efficacy and disease recurrence.
Biomechanical analysis studies the mechanical structure, function and movement of biological systems using mechanical methods. Recent research shows that more researchers tend to use Physics Informed Neural Networks (PINN) instead of finite elements for biomechanical analysis because it greatly shortens the analysis time. However, most existing approaches have two shortcomings: (1) Random initialization of small samples leads to the local optimal solution; (2) Single attribute leads to stress deviation. To tackle the shortcomings mentioned above, we proposed a biomechanical analysis method to obtain the global optimal solution in small sample CT data sets and automatically distinguish bone components. Specifically, to solve the problem that small sample CT data easily produces local optimal solutions and leads to poor generalization ability, we try to purposefully reconstruct the initialization weight matrix and design an initialization weight method based on spinal CT image features to alleviate this problem and thus improve the generalization ability of the model. To distinguish different bone compositions, a bone composition differentiation method based on PINN was designed, which can automatically distinguish different bone compositions. At the same time, leveraging the acquired global optimal solution and the adaptive matching parameters of the bone components, we employ multiple sets of physical rules to constrain feature mapping through numerical regression. This approach enhances the accuracy of biomechanical analysis. Extensive experiments on three representative thoracic and lumbar spine datasets achieved an accuracy of 91.09%, demonstrating the effectiveness of the method. The training and testing time for each sample does not exceed 8ms, demonstrating the technique’s applicability.
The 3D spinal model plays a crucial role in the assessment and treatment decision of adolescent idiopathic scoliosis. The complex 3D shape of the spine cannot be fully captured by a single radiograph. A 3D spine reconstruction framework is developed in this study. First, a dual-training strategy for Generative Adversarial Networks (GANs) is proposed, which generates high-quality 3D spinal structures. Second, an adaptive scale-agnostic attention mechanism is integrated to establish cross-layer feature correlations and dynamically allocate weights. This mechanism ensures the preservation of the crucial information across all scales throughout the feature extraction process. The proposed method has been validated on 49 cases of scoliosis. Experiments show that surface overlap and volume Dice coefficient are 0.92 and 0.94, respectively. Compared with the state-of-the-art methods, the proposed method reduces the average surface distance by 0.16 mm. The results demonstrate its effectiveness in reconstructing the 3D spine from a single radiograph.
Adolescent idiopathic scoliosis (AIS) is a three-dimensional spine deformity governed of the spine. A child’s Risser stage of skeletal maturity must be carefully considered for AIS evaluation and treatment. However, there are intra-observer and inter-observer inaccuracies in the Risser stage manual assessment. A multi-task learning approach is proposed to address the low precision issue of manual assessment. With our developed multi-task learning approach, the iliac area is extracted and forwarded to the improved Swin Transformer for Risser stage assessment. The spatial and channel reconstruction convolutional Swin block is adapted to each stage of the Swin Transformer to achieve better performance. The Risser stage assessment based on iliac region extraction had an overall accuracy of 81.53 https://github.com/xyz911015/Risser-stage-assessment
Purpose: Adolescent idiopathic scoliosis (AIS) is a three-dimensional spine deformity governed by lateral curvature and axial vertebral rotation (AVR). Estimating AVR is important for treatment decisions and the prediction of AIS progression. However, manual AVR measurements have intra-observer and inter-observer errors, and therefore this paper proposes an automatic AVR measurement method to address the low precision issue of manual measurement. Method: We develop an improved feature extraction module for vertebral landmark detection and pedicle segmentation. The improved coordinate convolution layer combined Polarized Self-Attention mechanism is applied in the feature extraction module to help the High-resolution Network to extract coordinate information. Based on vertebral landmark detection and pedicle segmentation, we propose an automatic AVR measurement algorithm to estimate the AVR. Results: The mean radial error (MRE) of vertebral landmark detection is 2.70 mm pixels, and the dice coefficient of pedicle segmentation is 72.45%. Compared with the original model, the MRE decreases by 2.13 mm, and the dice coefficient improves by 4.61%. For the AVR measurement results, the testing set contains 37 spine radiographs, including 481 vertebrae and 962 AVR measurements (481 vertebrae by two observers). The average measurement error is lower than 5 degrees, which is within the standard of clinical measurement error. Conclusion: The results demonstrate that the proposed network performs well in vertebral landmark detection and pedicle segmentation, and the proposed AVR measurement is adopted for clinic diagnosis of AIS. Significance: Our method achieves automatic AVR measurement, reducing the error introduced by manual measurement and improving the efficiency of orthopedists.
Skin cancer is a significant public health issue, and computer-aided diagnosis technology can effectively alleviate this burden. Accurate identification of skin lesion types is crucial when employing computer-aided diagnosis. This study proposes a multi-level attention cascaded fusion model based on Swin-T and ConvNeXt. It employed hierarchical Swin-T and ConvNeXt to extract global and local features, respectively, and introduced residual channel attention and spatial attention modules for further feature extraction. Multi-level attention mechanisms were utilized to process multi-scale global and local features. To address the problem of shallow features being lost due to their distance from the classifier, a hierarchical inverted residual fusion module was proposed to dynamically adjust the extracted feature information. Balanced sampling strategies and focal loss were employed to tackle the issue of imbalanced categories of skin lesions. Experimental testing on the ISIC2018 and ISIC2019 datasets yielded accuracy, precision, recall, and F1-Score of 96.01%, 93.67%, 92.65%, and 93.11%, respectively, and 92.79%, 91.52%, 88.90%, and 90.15%, respectively. Compared to Swin-T, the proposed method achieved an accuracy improvement of 3.60% and 1.66%, and compared to ConvNeXt, it achieved an accuracy improvement of 2.87% and 3.45%. The experiments demonstrate that the proposed method accurately classifies skin lesion images, providing a new solution for skin cancer diagnosis.
Background: Microwave ablation (MWA) is a minimally invasive alternative for the treatment of unresectable liver tumors. To verify the effectiveness and safety of MWA, it is critical to measure the temperature variation and assess the regions of the microwave-induced thermal lesions. Purpose: Recent studies have indicated that the locations of optimally matched Gabor atoms (LOMGA) from ultrasound radiofrequency (RF) echo signals allow accurate and stable scatterer spacing estimation. Herein, a harmonic-based LOMGA method is proposed to estimate the scatterer spacing for improving the assessment of microwave-induced thermal lesions. Methods: The mean scatterer spacing (MSS) is estimated via the LOMGA method incorporating the selection of concise atoms from separated second-harmonic RF echo signals with the pulse-inversion algorithm for thermal lesion evaluation. In vitro experiments, 10 fresh porcine liver samples were ablated at different time nodes during the ablation period, and 200 sets of second-harmonic and fundamental RF echo signals were randomly selected from the regions of interest in the coagulated liver samples for MSS estimation. The means and standard deviations of the MSSs, as well as the linear regression for the mean MSSs, were calculated from fundamental and second-harmonic signals for comparison and evaluation, the receiver operating characteristic (ROC) curves for the 200 sets of fundamental-based and harmonic-based MSS estimates from the 10 liver samples at five pairs of adjacent time nodes were calculated, and one-way analysis of variance (ANOVA) tests were performed for the five pairs of adjacent time nodes. The fundamental and harmonic-based p-values in the ANOVA tests and the areas under the ROC curves (AUCs) were calculated to statistically analyze the differences in the MSSs between adjacent time nodes. Results: The harmonic-based increments in the intensity variation and coherent components were larger than the fundamental-based increments with the increasing ablation time. The harmonic-based MSSs from the 10 liver samples at five pairs of adjacent time nodes were found to be highly statistically significant (p < 0.01). Thus, the harmonic-based MSSs had greater variations. Compared with the fundamental-based results, for the five preset ST values, the average increment in the harmonic-based mean slopes was 69.22% and the average decrement in the mean standard deviations was 11.67% for the linear-fitting MSS results, and the results were statistically significant (p < 0.05). Conclusion: Harmonic-based MSSs are more sensitive and robust to variations in coagulated tissues, which is advantageous for the assessment of microwave-induced thermal lesions.
Biomechanics are crucial for diagnosing and analyzing Adolescent Idiopathic Scoliosis (AIS) in CT images. Recent research suggests that spine biomechanics follow equilibrium equations, but the exact mathematical model and parameters are not fully understood. In this study, we use equilibrium equations and state space models to develop a computational biomechanics framework. Specifically, we create a model combining biomechanical analysis based on equilibrium equations with state-space image segmentation. First, we propose PM-UNet, a segmentation method using a state-space model to capture bone structure in CT images. Then, Physical Information Neural Networks (PINN) with equilibrium equations, using control matrices a and ss, calculate biomechanical properties. We validate the model by comparing results with FEM calculations. The experimental results show that the PM-UNet and PINN framework effectively extracts mechanical features and calculates biomechanical properties, providing a foundation for advanced analysis and improving surgical planning and brace design for AIS.
新型冠状病毒肺炎在全球范围迅速蔓延,为快速准确地对其诊断,进而阻断疫情传播链,提出一种基于深度学习的分类网络DLDA-A-DenseNet.首先将深层密集聚合结构与DenseNet-201结合,对不同阶段的特征信息聚合,以加强对病灶的识别及定位能力;其次提出高效多尺度长程注意力以细化聚合的特征;此外针对CT图像数据集类别不均衡问题,使用均衡抽样训练策略消除偏向性.在中国胸部CT图像调查研究会提供的数据集上测试,所提方法较原始DenseNet-201在准确率、召回率、精确率、F1分数和Kappa系数提高了 2.24%、3.09%、2.09%、2.60%和3.48%;并在COVID-CISet图像数据集上测试,取得99.50%的最优准确率.结果表明,对比其他方法,提出的新冠肺炎CT图像分类方法充分提取了 CT切片的病灶特征,具有更高的精度和良好的泛化性.
Essential matrix (E-matrix) estimation is a crucial aspect of pose estimation. In this study, we developed an end-to-end method (E-net) to estimate the E-matrix without correspondences. A pair of the corresponding images was placed in the twin transformer architecture to simultaneously extract the features. We developed a feature matching module for matching the extracted features based on their commonalities. To avoid excessive network parameters, matched features with their weights obtained by multilayer perceptron were transmitted to the flatten layer, where the Max-Pooling was used to eliminate their useless portions. We further constructed three self-defined layers to ensure that E-matrix is rank-2 with 5 degrees of freedom using reserved helpful features. Besides, we presented two self-defined loss functions (Loss1 and Loss2) to train the E-net and improve the estimated E-matrix's accuracy. E-net's performance was evaluated on the KITTI and TUM SLAM datasets using two self-defined metrics, M1 (mean value of matching error) and M2 (mean squared value of matching error). The E-net achieved M1 0.107 and M2 0.091 on the KITTI dataset and M1 0.253 and M2 0.144 on the TUM SLAM dataset. The results demonstrated that the E-net trained with self-defined loss functions outperforms other algorithms when compared to the 5-point algorithm of M1 10.411 and M2 8.332.
PURPOSE:The purpose of this study was to develop and evaluate a deep learning network for three-dimensional reconstruction of the spine from biplanar radiographs. METHODS:The proposed approach focused on extracting similar features and multiscale features of bone tissue in biplanar radiographs. Bone tissue features were reconstructed for feature representation across dimensions to generate three-dimensional volumes. The number of feature mappings was gradually reduced in the reconstruction to transform the high-dimensional features into the three-dimensional image domain. We produced and made eight public datasets to train and test the proposed network. Two evaluation metrics were proposed and combined with four classical evaluation metrics to measure the performance of the method. RESULTS:In comparative experiments, the reconstruction results of this method achieved a Hausdorff distance of 1.85 mm, a surface overlap of 0.2 mm, a volume overlap of 0.9664, and an offset distance of only 0.21 mm from the vertebral body centroid. The results of this study indicate that the proposed method is reliable.
Considering the existing spinal computer tomography(CT)and magnetic resonance(MR)image segmentation models have limitations in segmentation performance, this paper proposed a spinal segmentation method MAU-Net based on U-shaped network.Firstly, this paper introduced coordinate attention module into the encoder of U-shaped network, which enabled the network accurately capture the spatial position information and embedded it into the channel attention.Secondly, this paper proposed dual-branch channel cross fusion module based on Transformer, it could replace the skip connection for multi-scale feature fusion.Finally, this paper proposed a feature fusion attention module to better fuse the semantic differences between Transformer and Convolution network.On scoliosis CT dataset, Dice reached 0.929 6,IoU reached 0.859 7.On the public MR dataset SpineSagT2Wdataset3,compared with FCN,Dice improved by 14.46%.Experimental results show that this method can effectively reduce the false segmentation area of vertebrae.
The disease analysis of the lumbar spine often requires a large number of three-dimensional (3D) models. Currently, there is a lack of 3D model of the lumbar spine for research, especially for the diseases such as scoliosis where it is difficult to collect sufficient data in a short period of time. To solve this problem, we develop an end-to-end network based on 3D variational autoencoder for randomly generating 3D lumbar spine model. In this network, the dual path encoder structure is used to fit two individual variables, i.e., mean and variance. Spatial coordinate attention modules are added to the encoder to improve the learning ability of the network to the 3D spatial structure of the lumbar spine. To enhance the power of the network to reconstruct the lumbar spine, a regularization loss is added to constrain the distribution loss. Additionally, Gaussian noise layers are added to the decoder to improve the authenticity and diversity of generated model. The experiments were conducted on the data of the entire lumbar spine and the individual lumbar vertebra, respectively. The results showed that the voxel intersection over union was 0.588 and 0.684, the voxel Dice coefficient was 0.739 and 0.811, the average surface distance was 0.807 and 1.189, and the Hausdorff distance was 2.615 and 3.710, for the entire lumbar spine and individual lumbar vertebra, respectively. The developed approach is comparable to the most commonly used model generation method of statistical shape model (SSM) in both visual and objective indicators, while the developed approach does not require the landmarks that is needed in the SSM method. Therefore, this fully automatic method can be easily used for population-based modeling of the lumbar spine which has the potential to be a powerful clinical tool.
Background Ultrasound image segmentation is challenging due to the low signal-to-noise ratio and poor quality of ultrasound images. With deep learning advancements, convolutional neural networks (CNNs) have been widely used for ultrasound image segmentation. However, due to the intrinsic locality of convolutional operations and the varying shapes of segmentation objects, segmentation methods based on CNNs still face challenges with accuracy and generalization. In addition, Transformer is a network architecture with self-attention mechanisms that performs well in the field of computer vision. Based on the characteristics of Transformer and CNNs, we propose a hybrid architecture based on Transformer and U-Net with joint loss for ultrasound image segmentation, referred to as TU-Net. Methods TU-Net is based on the encoder-decoder architecture and includes encoder, parallel attention mechanism and decoder modules. The encoder module is responsible for reducing dimensions and capturing different levels of feature information from ultrasound images; the parallel attention mechanism is responsible for capturing global and multiscale local feature information; and the decoder module is responsible for gradually recovering dimensions and delineating the boundaries of the segmentation target. Additionally, we adopt joint loss to optimize learning and improve segmentation accuracy. We use experiments on datasets of two types of ultrasound images to verify the proposed architecture. We use the Dice scores, precision, recall, Hausdorff distance (HD) and average symmetric surface distance (ASD) as evaluation metrics for segmentation performance. Results For the brachia plexus and fetal head ultrasound image datasets, TU-Net achieves mean Dice scores of 79.59% and 97.94%; precisions of 81.25% and 98.18%; recalls of 80.19% and 97.72%; HDs (mm) of 12.44 and 6.93; and ASDs (mm) of 4.29 and 2.97, respectively. Compared with those of the other six segmentation algorithms, the mean values of TU-Net increased by approximately 3.41%, 2.62%, 3.74%, 36.40% and 31.96% for the Dice score, precision, recall, HD and ASD, respectively.