Cardiovascular-kidney-metabolic (CKM) syndrome represents a composite disease state driven by glucose and lipid metabolic dysregulation. The Cholesterol, High-density lipoprotein, and Glucose (CHG) index reflects the composite burden of these metabolic factors. However, its impact on the onset, stage-wise progression, and prognosis of CKM syndrome remains unclear. This study utilized data from two large-scale prospective cohorts. In the UK Biobank (UKB) cohort, Fine-Gray proportional subdistribution hazards models were employed to investigate associations between CHG levels and the risk of incident cardiovascular disease (CVD), chronic kidney disease (CKD), and type 2 diabetes mellitus (T2DM), as well as the risk of progression from CKM Stage 0–1 to 2–3 and Stage 1–3 to 4. In the Beijing Anzhen Hospital cohort, Cox regression models assessed the impact of CHG on major adverse cardiovascular and cerebrovascular events (MACCE) in patients with CKM Stage 4 (established coronary artery disease). A causal forest algorithm was used to identify high-value subpopulations, and restricted cubic splines (RCS) characterized dose-response relationships. Sensitivity analyses were conducted to verify result robustness. The study included 370,916 participants free of CKM diseases from UKB (median follow-up: 16.5 years) and 8,494 patients with CKM Stage 4 from Beijing Anzhen Hospital (median follow-up: 645 days). In the general population, each 1-SD increase in the CHG index was significantly associated with an increased risk of incident T2DM (HR: 1.47; 95
This study aimed to develop and validate a visualization-based predictive model for evaluating the risk of pedicle screw loosening (PSL) following posterior lumbar interbody fusion (PLIF) in patients with lumbar degenerative conditions. A total of 466 consecutive patients undergoing primary PLIF with pedicle screw fixation for lumbar degenerative disease were retrospectively enrolled. Patients with prior spinal surgery, deformity, fracture, tumor, ankylosing spondylitis, metabolic bone disease, fixation involving more than four levels, screw redirection or malposition, or non-PSL-related reoperation within 12 months were excluded. Patients treated in 2016–2021 formed the derivation cohort, and those treated in 2022–2023 formed the temporal internal validation cohort. Multivariable logistic regression was used to identify independent predictors and develop a PSL risk-scoring model. Model robustness was assessed using 3,000-iteration bootstrap resampling, and a 5-fold cross-validation framework was used to generate receiver operating characteristic curves and calculate the area under the curve (AUC). Calibration and decision curve analyses were performed to evaluate model calibration and clinical utility. In the derivation cohort, the PSL incidence within 1 year was 22.74
OBJECTIVE This study aimed to develop a deep learning (DL) model for the detection of cervical spinal cord compression on cervical radiographs and compare its performance with spine surgeons. METHODS The authors conducted a retrospective study on consecutive hospitalized patients who underwent cervical spine radiography and MRI at their center. Data from 600 patients were randomly divided into the training (n = 480), validation (n = 60), and internal test (n = 60) sets. Additionally, patients from another center were included as an external test set (n = 60). MR images were used as the gold standard for determining the presence of cervical segmental compression. The model was trained on cervical radiographs, where a segmentation-based localization algorithm was first developed to identify cervical segments, followed by a binary classification to diagnose spinal cord compression. Furthermore, the gradient-weighted class activation mapping (Grad-CAM) was used to visualize the area with high feature densities extracted by the model. Model performance was evaluated based on accuracy, sensitivity, specificity, F1 score, and area under the receiver operating characteristic curve (AUC), and compared to the diagnoses of two spine surgeons. RESULTS In the internal test set, the model achieved 94.67% accuracy and an AUC of 0.9911, significantly outperforming the two spine surgeons (69.09% and 71.18% accuracy, p < 0.05). In the external test set, the model achieved 93.33% accuracy with an AUC of 0.9868. Compared to the reference standard, the kappa coefficients for the model, reader 1, and reader 2 were 0.893, 0.378, and 0.422, respectively. Grad-CAM showed high feature density in the intervertebral discs, intervertebral foramina, and facet joints. CONCLUSIONS The DL model developed in this study achieved binary classification of cervical spinal cord compression and localized the affected segments on cervical radiographs, demonstrating superior diagnostic performance to that of spine surgeons. This DL model ensures high detection rates for cervical spinal cord compression and holds promise for clinical diagnosis, particularly in resource-limited or remote settings.
Background:Outpatients presenting with low back pain (LBP) often require efficient preconsultation triage and early differential diagnostic support. Large language models may assist these text-based tasks, but their performance under different clinical information conditions remains unclear. Objective:This study aimed to compare the performance of ChatGPT (5.1; OpenAI) and DeepSeek (V3.2; DeepSeek AI) in musculoskeletal disorders (MSDs) triage and the differential diagnosis of outpatients with LBP using real-world outpatient records under 2 simulated information conditions. Methods:This retrospective comparative study was conducted at a tertiary academic teaching hospital in Beijing. A total of 160 cases were included using a balanced design across 8 diagnostic categories (20 per category); 6 MSDs and 2 non-MSDs. Evaluation was performed in 2 phases: Phase 1 (chief complaint) and Phase 2 (structured questionnaire with 7 domains or 33 items), both executed in a zero-shot setting using standardized prompts. Outcomes included (1) triage accuracy, (2) preliminary diagnosis accuracy, and (3) differential diagnosis agreement. In Phase 2, 3 senior orthopedic evaluators additionally rated model rationales across 5 domains using a 5-point Likert scale. Results:For triage accuracy across all 160 cases, DeepSeek V3.2 improved from 84.4% to 90.6% (risk difference [RD] 6.2%, 95% CI -0.7% to 13.3%), and ChatGPT 5.1 improved from 75.6% to 93.1% (RD 17.5%, 95% CI 10.2%-24.9%). For preliminary diagnosis accuracy across the 120 musculoskeletal cases, DeepSeek V3.2 improved from 48.3% to 76.7% (RD 28.3%, 95% CI 16.8%-38.8%), whereas ChatGPT 5.1 improved from 35.0% to 87.5% (RD 52.5%, 95% CI 42.8%-60.6%). The mean number of correct differential diagnoses increased from 1.27 (SD 0.71) to 2.02 (SD 0.74) for DeepSeek V3.2 and from 1.34 (SD 0.70) to 2.03 (SD 0.77) for ChatGPT 5.1. In Phase 2, rationale ratings were generally good for both models, with ChatGPT 5.1 scoring higher in understanding and reasoning. Recognition of multiple myeloma (MM) remained limited, improving only from 45% to 55% (DeepSeek V3.2) and 55% to 60% (ChatGPT 5.1). Structured input reduced safety-risk errors in both models, but residual errors remained, especially for MM and metastatic spinal tumor. Conclusions:Both ChatGPT 5.1 and DeepSeek V3.2 demonstrated potential in text-based triage and differential diagnosis of MSDs for LBP, with structured clinical information generally improving performance, particularly for preliminary diagnosis accuracy and differential diagnosis agreement. However, their suboptimal sensitivity for red-flag conditions such as MM highlights significant safety concerns, indicating that they should not be used as stand-alone triage tools without clinician oversight. ChatGPT 5.1 showed stronger reasoning with structured inputs based on rationale ratings, whereas DeepSeek V3.2 showed better performance under chief-complaint-only input, with significantly higher Phase 1 preliminary diagnostic accuracy and numerically higher Phase 1 triage accuracy. These findings underscore the need for further model refinement, rigorous prospective validation, and integration with clinician oversight before clinical implementation.
Adolescent idiopathic scoliosis (AIS) is a complex three-dimensional spinal deformity, frequently requiring fusion surgery. An optimal fusion surgical strategy can not only achieve effective correction but also reduce the incidence of postoperative complications. Recently, several researchers have refined and expanded AIS fusion surgical strategies based on the Lenke classification system, which is the current international standard for AIS. Therefore, this study aims to review the advances in fusion level selection and surgical approaches for AIS based on this classification. Databases such as PubMed, Embase, Web of Science, Scopus, Cochrane Database, China National Knowledge Infrastructure, Wanfang Database, and China Biomedical Literature Database were queried for articles using the keywords “adolescent idiopathic scoliosis”, “fusion surgery”, “Lenke classification system”, “Lenke 1”, “Lenke 2”, “Lenke 3”, “Lenke 4”, “Lenke 5” and “Lenke 6”. Over the past decade, fusion surgical guidelines based on the Lenke classification have been refined, with new strategies emerging. We summarize the latest AIS fusion surgical strategies with recent research results. However, the fusion strategy based on the Lenke classification system has undergone no revolutionary changes. The selection of surgical designs for certain subtypes remains controversial. The fusion surgical strategy based on the Lenke classification system remains the standard for AIS surgical treatment. With the advancement of surgical technologies, further optimization of surgical strategies and the development of three-dimensional classification systems are potential future directions.
This study aimed to evaluate the predictive value of CT-based S1 Hounsfield Unit (HU) measurements for postoperative pedicle screw loosening (PSL) in patients with lumbar degenerative diseases who underwent posterior lumbar interbody fusion (PLIF). Consecutive patients who underwent PLIF at our institution between January 2016 and June 2024 were retrospectively analyzed. The L1 and S1 HU values were obtained using CT, whereas L1–L4 and S1 vertebral bone quality (VBQ) scores were obtained using MRI. Multivariate logistic regression analysis was performed to identify independent predictors of PSL. The area under the receiver operating characteristic curve (AUC) was used to assess the predictive performance of bone quality parameters. Optimal cutoff values were determined using the Youden index. A total of 285 patients were included. The PSL rate was 21.40
ABSTRACT Objective Kümmell’s disease (KD) represents a delayed form of osteoporotic vertebral collapse and shares clinical features with osteoporotic vertebral compression fractures (OVCF). Percutaneous kyphoplasty (PKP) is commonly performed for both conditions, yet comparative evidence and predictors of 1‐year outcomes in KD remain limited. This study aimed to evaluate the efficacy and safety of PKP for the treatment of stage I and II KD and to identify factors associated with the outcomes at 1‐year follow‐up. Methods We included 387 inpatients with KD or OVCF who underwent PKP from January 2016 to December 2022. All patients were assigned to the KD group (n = 107) and the OVCF group (n = 280). The difference of demographic data (age, gender, surgical segment, osteoporosis severity and disease duration), clinical efficacy (visual analog scale and Oswestry disability index of pre‐operation, 3‐day post‐operation, 3‐month post‐operation, and 1‐year post‐operation), complications (bone cement leakage during surgery and postoperative refractures), and radiographic parameters (anterior vertebral height and kyphotic angle of pre‐operation, post‐operation, and 1‐year follow‐up) was analyzed. Intergroup comparisons of continuous variables were performed using the Student's t‐test. Repeated measures ANOVA with Bonferroni post hoc correction was used to evaluate intra‐group differences of visual analog scale (VAS), Oswestry disability index (ODI), anterior vertebral height (AVH) and kyphotic angle (KA) across different time points. Multivariate logistic regression analysis was employed to identify the independent factors influencing VAS and ODI scores during the follow‐up period. Results The disease duration of the KD group was much longer than that of the OVCF group. Significant improvements in VAS and ODI were observed at three‐day, three‐month, and one‐year after PKP. Multivariate regression analysis identified the blocky cement distribution pattern, higher preoperative VAS, and higher preoperative ODI as independent risk factors for suboptimal recovery during follow‐up. Besides, the KD group had lower AVH and larger KA than the OVCF group preoperatively. Both groups showed significant improvements in AVH and KA after PKP. However, the KD group had a higher rate of type II bone cement leakage (BCL) and more severe cemented vertebral collapse at the final follow‐up. The mean bone cement volume was significantly greater in the KD group. Refracture rates were similar between the two groups during follow‐up. Conclusions PKP can effectively alleviate back pain, improve functional impairment, and correct local deformity in KD patients, with the risks of BCL and vertebral collapse or refractures during follow‐up. It is suitable for KD treatment without nerve injury symptoms. Appropriate measures should be taken to reduce the risk of complications.
BackgroundThis study meticulously outlines the evolution of the burden of malignant neoplasm of bone and articular cartilage (MNBAC) among different age and sex groups in China from 1990 to 2021, analyzes the global impact of the disease, and predicts the trend of disease burden up to 2035.MethodsLeveraging public data from the Global Burden of Disease (GBD) database spanning 1990 to 2021, this study thoroughly analyzed the characteristics of the burden of MNBAC in China and globally, including its incidence, prevalence, mortality, and disability-adjusted life years (DALYs). The joinpoint analysis method was employed to calculate the average annual percentage change (AAPC) and its 95% uncertainty interval, revealing the trend of MNBAC's impact. Furthermore, Bayesian age-period-cohort (BAPC) model was used to forecast changes in disease burden leading up to 2035.ResultsBetween 1990 and 2021, the age-standardized incidence rate (ASIR) of MNBAC in China increased from 0.65 to 1.42 per 100 000 people, and the global ASIR rose from 0.97 to 1.11 per 100 000. The AAPC for China's ASIR, age-standardized prevalence rate (ASPR), age-standardized mortality rate (ASMR), and age-standardized DALYs rate (ASDR) were 2.59%, 2.71%, 1.52%, and 1.26%, and the AAPC of ASIR and ASPR of the global burden of MNBAC were 0.44% and 0.51%, respectively. The effect of age and sex on the burden of MNBAC showed significant differences. Forecasting analyses suggest that from 2022 to 2035, the burden of MNBAC in China and globally will show a declining trend.ConclusionsFrom 1990 to 2021, the disease burden of MNBAC in China has been rising among the population, particularly pronounced among older men. Although forecasts indicate a gradual reduction in the future burden of MNBAC, given China's large population base and the increasing trend of population aging, MNBAC will pose a public health challenge in China.
Background:Vertebral compression fractures (VCFs) impose a substantial clinical and health care burden, and their management relies on timely access to evidence-based guidelines. Large language models (LLMs) may help clinicians rapidly obtain guideline-related information, but their performance on VCF guidelines remains unclear. Objective:This study aimed to evaluate the performance of LLMs, including DeepSeek-R1 and ChatGPT-5, in generating responses consistent with VCF clinical guidelines. Methods:Using the 2024 North American Spine Society VCF clinical guidelines as the reference standard, 34 open-ended and 87 closed-ended questions were submitted to DeepSeek-R1 and ChatGPT-5. Four senior spine surgeons independently rated responses to both closed-ended and open-ended questions using a 5-point Likert scale for accuracy, consistency, self-awareness, and fabrication/falsification. For open-ended questions, comprehensiveness, clarity, and trust and confidence were additionally assessed. Subgroup analyses were performed by question type, recommendation grade, and VCF subtype, with direct comparisons between models. Results:A total of 726 responses were generated for 121 questions. For closed-ended questions, ChatGPT-5 and DeepSeek-R1 showed comparable performance in accuracy (P=.11), self-awareness (P=.10), and fabrication/falsification (P=.10). DeepSeek-R1 demonstrated better consistency than ChatGPT-5 for both closed-ended and open-ended questions (P<.001 and P=.001, respectively). For open-ended questions, the models differed significantly in comprehensiveness (P=.03) and trust and confidence (P=.02), but not in accuracy (P=.42), self-awareness (P=.22), fabrication/falsification (P=.64), or clarity (P=.48). Closed-ended questions generally outperformed open-ended questions. Responses to grade A-C recommendations outperformed grade I recommendations in accuracy, consistency, and fabrication/falsification (all P≤.001) but scored lower in self-awareness (P<.001). No significant differences were observed across VCF subtypes. Conclusions:Under a standardized clinician-oriented prompting condition, ChatGPT-5 and DeepSeek-R1 showed generally high but variable scores across evaluation dimensions, with important deficiencies remaining, particularly in interventional and surgical treatment recommendations and in questions linked to recommendation grade I. Because these findings were obtained in a controlled prompting setting, caution is warranted when extrapolating them to other query styles, clinical scenarios, or LLMs.
Exploring large language models (LLMs) performance in the specific medical domain can help understand their generalizability in real-world application. We assessed the predictive and decision-support value of two state-of-the-art LLMs in predicting bone cement leakage (BCL) and new vertebral fractures (NVF) after percutaneous kyphoplasty (PKP) and to compare them with those of traditional machine learning (TML) and spine surgeon. This study utilized combined retrospective and prospective data at a single tertiary hospital. Two LLMs (GPT-5 and DeepSeek R1) with zero- and few-shot strategy, five TML models, and two spine surgeons with/without exposure to LLM responses, were asked to predict complications based on demographic, perioperative baseline, and radiographic data. We also tested LLMs' ability to predict complication subtype. For BCL prediction, both LLMs demonstrated acceptable performance (F1-score, 0.857-0.871; MCC, 0.164-0.332) under zero-shot conditions, comparable to TML models (F1-score, 0.758-0.867; MCC, 0.265-0.416), and slightly superior to surgeons alone (F1-score, 0.675-0.684; MCC, 0.074-0.185). Few-shot prompting enhanced specificity but yielded uncertain overall gains. For NVF prediction, the zero-shot LLM performance was poor (F1-score, 0.309; MCC, 0.044) but improved with few-shot learning. The RBF-SVM model showed the best performance for NVF prediction (F1-score, 0.536; MCC, 0.414). LLM explanations enhanced surgeon performance in BCL prediction but not in NVF. LLMs showed poor prediction of complication subtypes. The findings suggest that current LLMs hold diverse predictive performances for different complications after PKP, they are still immature for real clinical applicability and need further improvement.
Objective:To develop and externally validate a dual-mechanism deep learning (DL) model that integrates vertebral segmentation and lesion detection for automated evaluation of lumbar degeneration and structured report generation on plain radiographs. Methods:In this retrospective study, 5,964 patients who underwent standing anteroposterior and lateral lumbar radiographs at a single institution and 600 patients from a public dataset (BUU-Spine) were included. Vertebral corners from T11-L5 (and S1 on lateral views) and 7 degenerative findings (scoliosis, straightened/preserved lordosis, spondylolisthesis, disc space narrowing, osteophytes, vertebral compression, and abdominal aortic calcification) were annotated by 3 spine surgeons. Two independently trained, parallel networks were developed, including a ResNet-based segmentation network and a YOLOv8-based detection network. A rule-based integration strategy reconciled both outputs and generated structured diagnostic reports. Segmentation accuracy, quantitative measurement agreement, diagnostic performance, and clinical acceptability of reports were evaluated. Results:Intra- and interobserver landmark distances within 3 mm reached 96% and >95%, respectively. On the internal test set, the percentage of correct keypoints within 3 mm was 95.7%-98.6%, with intraclass correlation coefficients of 0.84-0.89 and Pearson correlation coefficient (r) of 0.90-0.94 for key radiographic parameters. The segmentation- and detection-based models achieved precision of 92.2%-96.9% and 91.7%-95.5%, and recall of 91.6%-94.8% and 93.3%-95.2%, respectively. Under the dual-positive condition, the integrated model yielded the highest precision (93.8%-97.3%), whereas the any-positive condition achieved the highest recall (94.1%-97.6%). Of 596 automatically generated structured reports, 557 (93.4%) were deemed clinically acceptable. Conclusion:The proposed dual-mechanism DL framework enables accurate, multilesion assessment of lumbar degeneration and generation of clinically acceptable structured reports from plain radiographs, supporting workflow optimization in lumbar spine imaging.
The optimal surgical strategy for geriatric patients with mild degenerative lumbar scoliosis (DLS) and lumbar spinal stenosis (LSS) remains controversial. Although percutaneous transforaminal endoscopic decompression (PTED) offers the advantages of minimal invasiveness, some older patients present with degenerative and stenotic features that may favor posterior lumbar interbody fusion (PLIF). This single-center retrospective study included geriatric patients (≥ 70 years) with DLS and LSS who underwent surgery between January 2016 and March 2023. Patients were divided into two groups based on the surgical approach: the PTED and PLIF groups. Clinical outcomes were assessed using the Oswestry Disability Index (ODI) and visual analog scale (VAS) preoperatively and at 3 months, 1 year, and 2 years after surgery. Patient satisfaction was assessed using the modified MacNab criteria. Radiological evaluations included measurements of the Cobb angle, lumbar lordosis, segmental coronal angle, and disc height. Additionally, complications and radiographic adjacent segment degeneration were recorded. A total of 68 geriatric patients (42 in the PTED group and 26 in the PLIF group) were included in this study. The baseline demographic and radiological characteristics were comparable between the PLIF and PTED groups. Similarly, both surgical techniques significantly improved postoperative clinical outcomes. However, compared to the PTED group, the PLIF group showed significantly lower VAS scores for back pain at 3 months, 1 year, and 2 years, along with lower leg pain VAS and ODI scores at 2 years (P < 0.05 for all). The satisfaction rates were comparable between the PLIF and PTED groups (88.46
OBJECTIVE:This study aimed to compare the efficacy of endplate Hounsfield unit (HU) values and endplate bone quality (EBQ) scores in predicting cage subsidence (CS) after posterior lumbar interbody fusion (PLIF) in older patients and identify the most discriminative bone mineral density (BMD) assessment indicator. METHODS:This retrospective analysis included consecutive patients who underwent PLIF at the authors' institution between January 2016 and February 2024. Clinical data were collected for all patients. Propensity scores were used to match patients with and without CS, and the matched cohort was subjected to conditional logistic regression to investigate the association between radiographic factors and CS. L1 and endplate HU values were derived from CT scans, whereas vertebral bone quality (VBQ) and EBQ scores were derived from MR images. Receiver operating characteristic curve analysis was conducted to assess the predictive value of endplate HU values and EBQ scores for CS and further compare their predictive value with that of L1 HU values and VBQ scores. RESULTS:This study included 130 matched patients. The CS group demonstrated lower L1 (p < 0.001) and endplate HU (p < 0.001) values and higher VBQ (p = 0.002) and EBQ (p < 0.001) scores compared with the non-CS group. The conditional logistic regression analysis identified L1 HU value (OR 0.99, 95% CI 0.97-0.99; p = 0.036), endplate HU value (OR 0.99, 95% CI 0.98-0.99; p = 0.044), VBQ score (OR 2.40, 95% CI 1.34-4.32; p = 0.038), and EBQ score (OR 4.46, 95% CI 2.16-9.18; p = 0.003) as independent predictors of CS, demonstrating areas under the curve of 0.722, 0.815, 0.648, and 0.782, respectively. The optimal cutoff for the endplate HU value in predicting CS was 262.11 (sensitivity 83.08%, specificity 73.85%). CONCLUSIONS:Endplate HU values indicated a relatively higher predictive performance for CS compared with EBQ scores and served as the most discriminative BMD indicator in patients who underwent PLIF. Measuring the endplate HU value preoperatively helps surgeons select a more appropriate surgical plan and is expected to improve patient outcomes.
Atherosclerosis is a leading cause of cardiovascular diseases, necessitating the identification of novel therapeutic targets. By utilizing Mendelian randomization, the aim is to identify potential drug targets for the treatment of atherosclerosis. Using results from genomewide association studies and protein quantitative trait loci data, we genetically identified causal effects of levels of several circulating proteins on atherosclerosis risk, validated using multiple statistical approaches, and identified promising drug targets. These findings demonstrated promising targets worthy of further clinical study, especially lipoprotein (a) and proprotein convertase subtilisin/kexin type 9.
This study examines the causal relationship between spermidine levels and coronary artery disease (CHD) risk using a bidirectional Mendelian Randomization (MR) approach. We employed genetic variants as instrumental variables to assess the influence of genetically predicted spermidine levels on CHD risk and vice versa. Data for the MR analysis were sourced from the UK Biobank and genome-wide association study datasets, focusing on single nucleotide polymorphisms (SNPs) associated with spermidine levels and CHD. The study also utilized liquid chromatography-tandem mass spectrometry (LC-MS/MS) for accurate quantification of spermidine in plasma samples. Our analysis identified a significant association between lower genetically predicted spermidine levels and increased CHD risk. The LC-MS/MS results supported the accurate measurement of spermidine, highlighting its feasibility as a clinical biomarker. The findings suggest that reduced spermidine levels may be a significant risk factor for CHD. This study supports the potential of spermidine as a biomarker for CHD risk assessment and its development as a therapeutic target. The integration of genetic and biochemical methodologies enhances our understanding of the role of spermidine in cardiovascular health and its utility in managing CHD risk.
Background:Osteoporosis is a prevalent skeletal disorder characterized by decreased bone mass and increased fracture risk; however, it frequently remains underdiagnosed due to limited health care resources and its asymptomatic progression. Deep learning (DL) provides a promising solution for automated screening using computed tomography (CT) scans, enabling earlier detection and improved management. Objective:This systematic review and meta-analysis aimed to investigate the diagnostic performance of DL models in diagnosing osteoporosis based on CT scans. Methods:This study was conducted under the PRISMA (Preferred Reporting Items for Systematic Reviews and Meta-Analyses) guidelines using articles extracted from PubMed, Scopus, Web of Science (Core), and Embase (Ovid). Studies involving adult participants who underwent CT and in which DL was applied for osteoporosis diagnosis were included. The QUADAS-2 (Quality Assessment of Diagnostic Accuracy Studies-2) tool was used to estimate the risk of bias in each study. The confusion matrices from the included studies were extracted to summarize the diagnostic performance of DL models for osteoporosis. Within a bivariate random-effects framework, sensitivity and specificity were jointly synthesized to yield the summary estimates. Heterogeneity was quantified with Higgins I² statistics. Subgroup analyses were performed to explore potential sources of heterogeneity among the included studies. Results:This review included 24 studies, encompassing CT images from 29,808 participants. All studies used conventional CT scans and used DL-based architectures. Fifteen, 6, and 3 studies were assessed as having a low, uncertain, and high risk of bias, respectively. The meta-analysis included 20 studies. The pooled sensitivity and specificity were 0.88 (95% CI 0.85-0.91; I2=83.69%) and 0.94 (95% CI 0.91-0.96; I2=95.07%) for osteoporosis diagnosis; 0.81 (95% CI 0.76-0.85; I2=82.38%) and 0.92 (95% CI 0.90-0.94; I2=79.05%) for osteopenia identification; and 0.95 (95% CI 0.92-0.97; I2=98.28%) and 0.93 (95% CI 0.91-0.95; I2=94.93%) for normal case identification. The area under the curve of the DL models for identifying osteoporosis, osteopenia, and normal cases was 0.96 (95% CI 0.93-0.97), 0.94 (95% CI 0.92-0.96), and 0.98 (95% CI 0.96-0.99), respectively. Subgroup analyses revealed that models based on DenseNet variants (P<.01), multislice input (P<.01), 3D architecture (P<.01), and CT as the reference standard (P<.01) demonstrated superior diagnostic performance. Conclusions:This study indicated that CT-based DL models achieve promising diagnostic performance for osteoporosis. However, substantial heterogeneity among the included studies, limited external validation, and incomplete end-to-end pipelines constrain the generalizability of the proposed models. Further research is warranted to support their clinical translation and standardized application.
STUDY DESIGN:Retrospective study. OBJECTIVE:To develop a deep learning (DL) model to predict bone cement leakage (BCL) subtypes during percutaneous kyphoplasty (PKP) using preoperative computed tomography (CT) as well as employing multicenter data to evaluate the effectiveness and generalizability of the model. BACKGROUND:DL excels at automatically extracting features from medical images. However, there is a lack of models that can predict BCL subtypes based on preoperative images. MATERIALS AND METHODS:This study included an internal data set for DL model training, validation, and testing as well as an external data set for additional model testing. Our model integrated a segment localization module based on vertebral segmentation through three-dimensional (3D) U-Net with a classification module based on 3D ResNet-50. Vertebral level mismatch rates were calculated, and confusion matrixes were used to compare the performance of the DL model with that of spine surgeons in predicting BCL subtypes. Furthermore, the simple Cohen kappa coefficient was used to assess the reliability of spine surgeons and the DL model against the reference standard. RESULTS:A total of 901 patients containing 997 eligible segments were included in the internal data set. The model demonstrated a vertebral segment identification accuracy of 96.9%. It also showed high area under the curve (AUC) values of 0.734 to 0.831 and sensitivities of 0.649 to 0.900 for BCL prediction in the internal data set. Similar favorable AUC values of 0.709 to 0.818 and sensitivities of 0.706 to 0.857 were observed in the external data set, indicating the stability and generalizability of the model. Moreover, the model outperformed nonexpert spine surgeons in predicting BCL subtypes, except for type II. CONCLUSION:The model achieved satisfactory accuracy, reliability, generalizability, and interpretability in predicting BCL subtypes, outperforming nonexpert spine surgeons. This study offers valuable insights for assessing osteoporotic vertebral compression fractures, thereby aiding preoperative surgical decision-making. LEVEL OF EVIDENCE:Level III.
BackgroundThe best treatment yielding clinical benefits was still equivocal and controversial for the treatment of cervical radicular pain (CRP). This study aimed to propose a novel combination strategy of percutaneous cervical nucleoplasty (PCN) and ultrasound-guided pulsed radiofrequency (PRF) of cervical nerve root for CRP, and to compare its therapeutic effects with PRF alone.Methods120 CRP patients who satisfied the inclusion requirements between January 2016 and March 2019 were retrospectively analyzed and split into PCN + PRF and PRF groups. The propensity score matching (PSM) technique was used to correct the imbalanced confounding variables between the groups. Then, clinical outcomes including the visual analog scale (VAS) score, Neck Disability Index (NDI) score, clinical assessment scale for cervical spondylosis (CASCS), modified MacNab criteria, radiological parameters, and complications were evaluated.ResultsIn all, 120 patients were used to calculate the propensity score, producing 26 matched pairs that were monitored for a minimum of a year. When compared to the preoperative data, both groups' neck pain VAS scores, arm pain VAS scores, NDI scores, and CASCS scores saw a significant improvement during the follow-up period (p < 0.001). However, patients in the PRF group noted higher neck pain VAS scores, arm pain VAS scores, NDI scores, and CASCS scores than those in the PRF + PCN group at the final follow-up (p < 0.05). The decrease in surgical level disc height was more pronounced in the PRF + PCN group at the final follow-up (P < 0.05). The ROM was reduced in the PRF group but increased in the PRF + PCN group at the final follow-up (P < 0.01). Based on the modified MacNab criteria, the PRF and PCN + PRF groups had excellent and good rates of 76.92% and 84.62%, respectively, with no statistically significant difference (P > 0.05).ConclusionWe present and describe a novel strategy for the combined treatment of CRP in chronic cervical radicular pain using ultrasound-guided percutaneous disc radiofrequency ablation PCN and spinal nerve root pulse radiofrequency PRF, which is both effective and safe throughout the treatment process, reducing pain and improving function.
PurposeTo assess the clinical and radiological outcomes of lumbar endoscopic decompression for the treatment of lumbar spinal stenosis (LSS) with concurrent degenerative lumbar scoliosis (DLS).MethodsThis study retrospectively reviewed 97 patients with LSS and DLS who underwent lumbar endoscopic decompression between 2016 and 2021. The average follow-up duration was 52.9 months. Another 97 LSS patients without DLS were selected as the control group. The pre- and postoperative visual analog score (VAS) and the Oswestry disability index (ODI) were recorded and analyzed to compare clinical outcomes. Radiological findings, such as coronal balance and intervertebral disc height, have also been reported.ResultsBoth groups' mean VAS scores for back pain, leg pain, and ODI were significantly improved two weeks after surgery and at the final follow-up (p < 0.001). There was no significant difference in the prevalence of surgical complications or patient satisfaction rates. However, patients in the DLS group reported more severe back pain at the final follow-up than those in the LSS group (p = 0.039). Radiological follow-up revealed no significant deterioration in coronal imbalance or loss of disc height in either group.ConclusionLumbar endoscopic decompression can be a safe and effective surgical technique for treating LSS with DLS, particularly in elderly patients with poor general conditions.