BACKGROUND:Obesity is a major risk factor for OSA, and visceral adiposity may mediate its cardiometabolic consequences. However, studies evaluating the impact of CPAP on visceral obesity have yielded conflicting results. RESEARCH QUESTION:What is the relationship between OSA and visceral adipose tissue (VAT) volume and metabolic activity, and how does short-term CPAP therapy modify these measures? STUDY DESIGN AND METHODS:Adults with newly diagnosed, moderate to severe OSA underwent combined [18F]-fluoro-2-deoxy-D-glucose (FDG) PET imaging and MRI before and after short-term CPAP therapy. Using a novel deep learning approach to segment abdominal adipose compartments, we quantified adipose volumes and metabolic activity using mean standardized uptake values (SUVmean). The primary outcome was VAT SUVmean. Secondary outcomes included VAT and subcutaneous adipose tissue (SAT) volume, and VAT to SAT volume ratio. Associations between OSA severity and adipose metrics and CPAP effects were assessed using multivariable regression and linear mixed-effects models. RESULTS:Among 134 participants, OSA severity was not associated with increased VAT SUVmean in adjusted analyses (P = .70) and CPAP did not alter VAT metabolic activity significantly after 3 months (P = .66). A modest reduction in VAT to SAT volume ratio (P = .03) was observed after CPAP therapy, despite no changes in weight or total abdominal adipose tissue volume. Although CPAP did not affect VAT volume (P = .09) overall, we observed a significant reduction in VAT volume in patients with obesity (-2.08%; 95% CI, -3.9 to -0.24; P = .03). INTERPRETATION:In this study, short-term CPAP therapy was not associated with a reduction in VAT metabolic activity, though a modest shift in fat distribution from the visceral to subcutaneous compartment was observed. Notably, patients with obesity had a significant reduction in VAT volume, highlighting potential heterogeneity of treatment effects. Our findings underscore the need for longer-term studies to evaluate how CPAP therapy and, importantly, adjunct weight loss pharmacotherapies may alter fat distribution and visceral adiposity, potentially modifying OSA-related cardiometabolic risk.
Rationale: Obstructive sleep apnea (OSA) is associated with an increased risk of cardiovascular disease (CVD), but the effectiveness of continuous positive airway pressure (CPAP) therapy in improving cardiovascular outcomes has been inconsistent. Visceral obesity has emerged as a potential mediator in this relationship and is a known risk factor for CVD. This study aimed to explore the relationship between OSA severity and abdominal obesity metrics, including visceral and subcutaneous adipose tissue (VAT and SAT) volumes and VAT metabolic activity, and to assess changes in these metrics following CPAP intervention, using [18F]-Fluoro-2-deoxy-D-glucose (FDG) positron emission tomography (PET) combined with magnetic resonance imaging (MRI). Methods: We retrospectively analyzed PET/MRI scans from 115 adults with newly diagnosed OSA, both before and after three months of CPAP therapy. OSA severity was determined using portable sleep testing, defined by the respiratory disturbance index (pRDI). A deep learning model, using a transfer learning approach, segmented regions of interest (ROIs) within the subcutaneous and visceral adipose tissue compartments on MRI. Adipose volume and SUVmean values were calculated for the ROIs at each time point. Log-transformed linear regression and linear mixed-effects models were used to evaluate associations between OSA severity, adipose tissue metrics, and the effects of CPAP, adjusted for age, sex, hypertension, diabetes, diet, and body mass index (BMI). Results: Participants had an average age of 47.0 years (standard deviation [SD] 11.97), were predominantly male (84.0%), and had an average BMI of 31.85 kg/m² (SD 5.08). The mean pRDI was 32.53 events/hour (SD 19.21). OSA severity was not significantly associated with baseline VAT SUVmean (0.02%, confidence interval [CI] [-0.21-0.25], p=0.87). CPAP therapy did not significantly alter VAT SUVmean (-2.17%, CI [-5.02-0.77], p=0.15), VAT volume (-0.98%, CI [-2.41-0.47], p=0.19), or weight (0.30%, CI [-0.83-1.43], p=0.61). However, there was a significant reduction in the VAT/SAT volume ratio (-1.70%, CI [-3.28-0.09], p=0.04). Conclusion: This study found no change in visceral adipose tissue metabolic activity after three months of CPAP. However, despite no significant changes in weight following CPAP, a notable reduction in the VAT/SAT volume ratio suggests that CPAP may differentially impact abdominal fat distribution. Future research should explore how alternative therapies for OSA, such as GLP-1 receptor agonists, may modulate abdominal fat distribution and metabolic activity in patients with obesity-related OSA.
Background and purpose:Predicting hepatocellular carcinoma (HCC) response to Stereotactic Body Radiation Therapy (SBRT) can be challenging. Here, we assessed the value of a radiomics-based machine learning (ML) approach for predicting HCC response to SBRT, using pre-treatment and early post-treatment magnetic resonance imaging (MRI). Materials and Methods:This retrospective single-center study included 87 patients (M 67, mean age 65.3 ± 9.1y) with HCC treated with SBRT who underwent gadoxetate MRI both pre- and early post-treatment (around 9.5 weeks). Tumor radiomics features were extracted on pre- and post-SBRT MRIs on pre- and post-contrast T1-weighted imaging (T1WI) [pre-contrast, arterial phase (AP), portal venous phase (PVP), transitional phase and hepatobiliary phase]. Long term response was assessed using modified RECIST criteria. Different ML models were developed based on 1st and 2nd order radiomics features to predict long-term objective response (partial and complete response) versus no response (stable and progressive disease). The cohort was randomly divided into training/validation (70 %) and testing 30 %. Results:A total of 87 tumors were assessed (mean size 2.7 ± 1.6 cm). Objective long-term response was observed in 43 (49.4 %) patients. The best predictive outcomes were achieved using models combining pre- and early post-treatment radiomics, with top performing model combining pre-treatment T1WI-pre-contrast, pre-treatment T1WI-AP and post-treatment T1WI-PVP, achieving an AUC of 0.85 [95 % CI: 0.67---1], sensitivity of 0.7 and specificity of 1. Conclusions:Our initial findings show promising results for ML radiomics in predicting long-term response of HCC to SBRT, which may have implications for management decisions.
DeepSeek is a newly introduced large language model (LLM) designed for enhanced reasoning, but its medical-domain capabilities have not yet been evaluated. Here we assessed the capabilities of three LLMs- DeepSeek-R1, ChatGPT-o1 and Llama 3.1-405B-in performing four different medical tasks: answering questions from the United States Medical Licensing Examination (USMLE), interpreting and reasoning on the basis of text-based diagnostic and management cases, providing tumor classification according to RECIST 1.1 criteria and providing summaries of diagnostic imaging reports across multiple modalities. In the USMLE test, the performance of DeepSeek-R1 (accuracy 0.92) was slightly inferior to that of ChatGPT-o1 (accuracy 0.95; P = 0.04) but better than that of Llama 3.1-405B (accuracy 0.83; P < 10-3). For text-based case challenges, DeepSeek-R1 performed similarly to ChatGPT-o1 (accuracy of 0.57 versus 0.55; P = 0.76 and 0.74 versus 0.76; P = 0.06, using New England Journal of Medicine and Médicilline databases, respectively). For RECIST classifications, DeepSeek-R1 also performed similarly to ChatGPT-o1 (0.74 versus 0.81; P = 0.10). Diagnostic reasoning steps provided by DeepSeek were deemed more accurate than those provided by ChatGPT and Llama 3.1-405B (average Likert score of 3.61, 3.22 and 3.13, respectively, P = 0.005 and P < 10-3). However, summarized imaging reports provided by DeepSeek-R1 exhibited lower global quality than those provided by ChatGPT-o1 (5-point Likert score: 4.5 versus 4.8; P < 10-3). This study highlights the potential of DeepSeek-R1 LLM for medical applications but also underlines areas needing improvements.
Hepatocellular carcinoma (HCC) surveillance primarily relies on ultrasound (U/S), which often exhibits decreased sensitivity in high-risk populations, such as individuals with cirrhosis or obesity. Abbreviated magnetic resonance imaging (AMRI) offers a potential alternative by employing targeted MRI sequences to enhance HCC detection. AMRI encompasses three primary strategies: non-contrast, dynamic contrast-enhanced, and hepatobiliary phase imaging, showing potential for overcoming U/S limitations in these populations. This study investigates the application of deep learning (DL) techniques to automate HCC tumor detection and segmentation within dynamic contrast-enhanced (Dyn-AMRI) protocols. Specifically, we leverage the capabilities of Vision Transformers (ViTs) to analyze complex image data and extract relevant features. Additionally, a novel heuristic is introduced to enhance the segmentation performance of the MedNeXt architecture. Our aim is to develop a robust DL pipeline for accurate HCC detection and segmentation on Dyn-AMRI, ultimately improving diagnostic outcomes.
The objective of this study is to describe the prevalence of inflammatory cardiopulmonary findings in a prospective cohort of long coronavirus disease (LC) patients. Methods: Subjects with a history of coronavirus disease 2019 infection, persistent cardiopulmonary symptoms 9-12 mo after initial infection, and a clinical assessment compatible with LC underwent cardiopulmonary 18F-FDG PET/MRI, dual-energy CT (DECT) of the lungs, and plasma protein analysis (subgroup). A control group that included subjects with a history of acute severe acute respiratory syndrome coronavirus 2 infection but without cardiopulmonary symptoms at recruitment was also characterized. Results: Ninety-eight patients (median age, 48.5 y; 47% men) were enrolled. The most common LC symptom was shortness of breath (80%), and 27% of participants were hospitalized. Of the subjects, 90% presented abnormalities in DECT, with 67% and 59% of participants demonstrating pulmonary infiltrates and abnormal perfusion, respectively. PET/MRI was abnormal for 57% of subjects: 24% showed cardiac involvement suggestive of myocarditis, 22% presented uptake reminiscent of pericarditis, 11% showed periannular uptake, and 30% showed vascular uptake (aortic or pulmonary). There was no myocardial, pericardial, periannular, or pulmonary uptake on the PET/MRI scans of the control group (n = 9). Analysis of plasma protein concentrations showed significant differences between the LC and the control groups. Lastly, the plasma protein profile was significantly different among LC patients with abnormal and normal PET/MRI. Conclusion: In LC subjects evaluated up to a year after coronavirus disease 2019 infection, our results indicate a high prevalence of abnormalities on PET/MRI and DECT, as well as significant differences in the peripheral biomarker profile, which might warrant further monitoring to exclude the development of complications such as pulmonary hypertension and valvular disease.
Background: Accurate quantification of visceral (VAT) and subcutaneous adipose tissue (SAT) is critical for understanding the cardiometabolic consequences of obstructive sleep apnea (OSA) and other chronic diseases. This study validates a customization framework using pre-trained networks for the development of automated VAT/SAT segmentation models using hybrid positron emission tomography (PET)/magnetic resonance imaging (MRI) data from OSA patients. While the widespread adoption of deep learning models continues to accelerate the automation of repetitive tasks, establishing a customization framework is essential for developing models tailored to specific research questions. Methods: A UNet-ResNet50 model, pre-trained on RadImageNet, was iteratively trained on 59, 157, and 328 annotated scans within a closed-loop system on the Discovery Viewer platform. Model performance was evaluated against manual expert annotations in 10 independent test cases (with 80-100 MR slices per scan) using Dice similarity coefficients, segmentation time, intraclass correlation coefficients (ICC) for volumetric and metabolic agreement (VAT/SAT volume and standardized uptake values [SUVmean]), and Bland-Altman analysis to evaluate the bias. Results: The proposed deep learning pipeline substantially improved segmentation efficiency. Average annotation time per scan was 121.8 min (manual segmentation), 31.8 min (AI-assisted segmentation), and only 1.2 min (fully automated AI segmentation). Segmentation performance, assessed on 10 independent scans, demonstrated high Dice similarity coefficients for masks (0.98 for VAT and SAT), though lower for contours/boundary delineation (0.43 and 0.54). Agreement between AI-derived and manual volumetric and metabolic VAT/SAT measures was excellent, with all ICCs exceeding 0.98 for the best model and with minimal bias. Conclusions: This scalable and accurate pipeline enables efficient abdominal fat quantification using hybrid PET/MRI for simultaneous volumetric and metabolic fat analysis. Our framework streamlines research workflows and supports clinical studies in obesity, OSA, and cardiometabolic diseases through multi-modal imaging integration and AI-based segmentation. This facilitates the quantification of depot-specific adipose metrics that may strongly influence clinical outcomes.
Rationale: Central abdominal obesity, particularly visceral adiposity may be a key player in mediating obstructive sleep apnea (OSA)-related cardiovascular disease (CVD) risk. Accurately measuring changes in visceral (VAT) and subcutaneous adipose tissue (SAT) volumes and metabolic activity could be crucial for evaluating the effectiveness of OSA therapies such as continuous positive airway pressure (CPAP) and novel weight-loss drugs. Manual analysis of abdominal adipose tissue on MRI can be time-intensive. We developed a dynamic training approach leveraging pre-trained AI models for abdominal fat segmentation in patients with OSA who underwent 18F-FDG positron emission tomography (PET) / magnetic resonance imaging (MRI), before and after CPAP. Methods: We utilized the AI Discovery Viewer (DV) platform, a web application for developing and deploying Medical AI models. In total, 328 abdominal PET/MRI scans from OSA patients were annotated within DV, with contours delineated for external (EXT) and internal (INT) SAT, as well as exclusionary (EXC) regions (i.e. kidneys, bone marrow). Initial training was conducted on a RadImageNet (RIN) UNet-ResNet50 model with 40 manually annotated cases, allowing for rapid model learning and facilitating AI-assisted annotation. This closed-loop system within DV enabled continuous fine-tuning of models with new annotations (Figure 1). Three versions of the models were assessed against manual segmentations in Osirix/Horos for segmentation speed, contour accuracy (Dice score), and VAT/SAT volumes and SUV. Performance was analyzed using the Wilcoxon Signed-Rank test. Results: The models achieved an average processing time of 1.14±0.19 minutes per scan, while expert-corrected AI segmentations took 31.8±17 minutes in DV, versus 134.5±27 minutes (Horos) and 96.8±7.8 minutes (Osirix) manually. Fat mask Dice scores showed high reliability, exceeding 0.98 for INT and EXT and reaching 0.83 for EXC. Contour Dice scores improved with each model iteration: INT from 0.39 to 0.45, EXT from 0.52 to 0.55, and EXC from 0.34 to 0.38. For VAT/SAT metrics, no significant differences were found between AI and manual annotations, except for VAT SUV mean (p=0.039), although the mean difference of 0.01 was not clinically significant. Conclusion: In summary, the abdominal adipose volumes and metabolic activity values derived using our AI models demonstrate a reasonable correlation to manual segmentation values. This novel approach has accelerated the time-for-annotation process by a factor of four and promises continued improvements with further model refinement. It demonstrates promise for abdominal fat quantification measures in OSA, to explore how therapies such as GLP-1 receptor agonists may modulate abdominal fat distribution and metabolic activity.
To evaluate radiomics features’ reproducibility using inter-package/inter-observer measurement analysis in renal masses (RMs) based on MRI and to employ machine learning (ML) models for RM characterization. 32 Patients (23M/9F; age 61.8 ± 10.6 years) with RMs (25 renal cell carcinomas (RCC)/7 benign masses; mean size, 3.43 ± 1.73 cm) undergoing resection were prospectively recruited. All patients underwent 1.5 T MRI with T2-weighted (T2-WI), diffusion-weighted (DWI)/apparent diffusion coefficient (ADC), and pre-/post-contrast-enhanced T1-weighted imaging (T1-WI). RMs were manually segmented using volume of interest (VOI) on T2-WI, DWI/ADC, and T1-WI pre-/post-contrast imaging (1-min, 3-min post-injection) by two independent observers using two radiomics software packages for inter-package and inter-observer assessments of shape/histogram/texture features common to both packages (104 features; n = 26 patients). Intra-class correlation coefficients (ICCs) were calculated to assess inter-observer and inter-package reproducibility of radiomics measurements [good (ICC ≥ 0.8)/moderate (ICC = 0.5–0.8)/poor (ICC < 0.5)]. ML models were employed using reproducible features (between observers and packages, ICC > 0.8) to distinguish RCC from benign RM. Inter-package comparisons demonstrated that radiomics features from T1-WI-post-contrast had the highest proportion of good/moderate ICCs (54.8–58.6
In this retrospective study, we annotated 44 structures on two datasets: an internal dataset of 1,518 MRI sequences from 843 patients at the Mount Sinai Health System, and an external dataset of 397 MRI sequences from 263 patients for benchmarking. The internal dataset trained the nnU-Net model MRAnnotator, which demonstrated strong generalizability on the external dataset. MRAnnotator outperformed existing models such as TotalSegmentator MRI and MRSegmentator on both datasets, achieving an overall average Dice score of 0.878 on the internal dataset and 0.875 on the external set. Model weights are available on GitHub, and the external test set can be shared upon request.
Virtual reality (VR) has emerged as a technology with huge improvements in the last few years, offering immersive experiences and paving the way for novel applications across various fields. In the medical field, VR is advancing at a notable rate, allowing for new innovations ranging from pain treatment to rehabilitation and showing its potential to reconfigure the brain to modify how our mind perceives our body and environment. The immersive nature of VR can be reflected in people's physiological and psychological responses, making stress management a key focus to obtain potentially transformative effects as a regulatory stress treatment. This pilot study focuses on a comprehensive assessment, both physiologically and psychologically, of the impact of stressful versus relaxing VR experiences. Our exploratory study examines the effectiveness of this emerging technology in manipulating stress levels in a group of 20 healthy volunteers. In a randomized cross-over design, each subject was exposed to both experiences in two independent sessions, during which their physiological signals were continuously measured. Additionally, blood draws and psychological tests were obtained before and after each VR experience. The results show physiological changes consistent with the paradigm of the experience and supported by self-reported psychological scores. Laboratory findings reveal statistical differences in cortisol levels when comparing changes in the Stress versus Relax experience. Additionally, statistically significant changes in white blood cell counts are observed when comparing pre- versus post-Stress VR. These results provide a first attempt to explore how VR scenarios can modulate individual stress levels and the immune response. Future exploratory avenues may include the implementation of VR-based treatments for stress modulation aimed at mitigating the detrimental effects of stress on mental health, cardiovascular disease, and the immune system.
Patients recovered from COVID-19 may develop long-COVID symptoms in the lung. For this patient population (post-COVID patients), they may benefit from longitudinal, radiation-free lung MRI exams for monitoring lung lesion development and progression. The purpose of this study was to investigate the performance of a spiral ultrashort echo time MRI sequence (Spiral-VIBE-UTE) in a cohort of post-COVID patients in comparison with CT and to compare image quality obtained using different spiral MRI acquisition protocols. Lung MRI was performed in 36 post-COVID patients with different acquisition protocols, including different spiral sampling reordering schemes (line in partition or partition in line) and different breath-hold positions (inspiration or expiration). Three experienced chest radiologists independently scored all the MR images for different pulmonary structures. Lung MR images from spiral acquisition protocol that received the highest image quality scores were also compared against corresponding CT images in 27 patients for evaluating diagnostic image quality and lesion identification. Spiral-VIBE-UTE MRI acquired with the line in partition reordering scheme in an inspiratory breath-holding position achieved the highest image quality scores (score range = 2.17-3.69) compared to others (score range = 1.7-3.29). Compared to corresponding chest CT images, three readers found that 81.5% (22 out of 27), 81.5% (22 out of 27) and 37% (10 out of 27) of the MR images were useful, respectively. Meanwhile, they all agreed that MRI could identify significant lesions in the lungs. The Spiral-VIBE-UTE sequence allows for fast imaging of the lung in a single breath hold. It could be a valuable tool for lung imaging without radiation and could provide great value for managing different lung diseases including assessment of post-COVID lesions.
Purpose: To assess the accuracy of a machine learning (ML) approach based on magnetic resonance (MR) imaging radiomic quantification obtained before treatment and early after treatment for prediction of early hepatocellular carcinoma (HCC) response to yttrium-90 transarterial radioembolization (TARE).Materials and Methods: In this retrospective single-center study of 76 patients with HCC, baseline and early (1-2 months) post-TARE MR images were collected. Semiautomated tumor segmentation facilitated extraction of shape, first-order histogram, and custom signal intensity-based radiomic features, which were then trained (n = 46) using a ML XGBoost model and validated on a separate cohort (n = 30) not used in training to predict treatment response assessed at 4-6 months (based on modified Response and Evaluation Criteria in Solid Tumors criteria). Performance of this ML radiomic model was compared with those of models comprising clinical parameters and standard imaging characteristics using area under the receiver operating curve (AUROC) analysis for prediction of complete response (CR).Results: Seventy-six tumors with a mean (+/- SD) diameter of 2.6 cm +/- 1.6 were included. Sixty, 12, 1, and 3 patients were classified as having CR, partial response, stable disease, and progressive disease, respectively, at 4-6 months posttreat-ment on the basis of MR images. In the validation cohort, the radiomic model showed good performance (AUROC, 0.89) for prediction of CR, compared with models comprising clinical and standard imaging criteria (AUROC, 0.58 and 0.59, respectively). Baseline imaging features appeared to be more heavily weighted in the radiomic model.Conclusions: The use of ML modeling of radiomic data combining baseline and early follow-up MR imaging could predict HCC response to TARE. These models need to be investigated further in an independent cohort.
Background: Patellofemoral anatomy has not been well characterized. Applying deep learning to automatically measure knee anatomy can provide a better understanding of anatomy, which can be a key factor in improving outcomes. Methods: 483 total patients with knee CT imaging (April 2017–May 2022) from 6 centers were selected from a cohort scheduled for knee arthroplasty and a cohort with healthy knee anatomy. A total of 7 patellofemoral landmarks were annotated on 14,652 images and approved by a senior musculoskeletal radiologist. A two-stage deep learning model was trained to predict landmark coordinates using a modified ResNet50 architecture initialized with self-supervised learning pretrained weights on RadImageNet. Landmark predictions were evaluated with mean absolute error, and derived patellofemoral measurements were analyzed with Bland–Altman plots. Statistical significance of measurements was assessed by paired t-tests. Results: Mean absolute error between predicted and ground truth landmark coordinates was 0.20/0.26 cm in the healthy/arthroplasty cohort. Four knee parameters were calculated, including transepicondylar axis length, transepicondylar-posterior femur axis angle, trochlear medial asymmetry, and sulcus angle. There were no statistically significant parameter differences (p > 0.05) between predicted and ground truth measurements in both cohorts, except for the healthy cohort sulcus angle. Conclusion: Our model accurately identifies key trochlear landmarks with ~0.20–0.26 cm accuracy and produces human-comparable measurements on both healthy and pathological knees. This work represents the first deep learning regression model for automated patellofemoral annotation trained on both physiologic and pathologic CT imaging at this scale. This novel model can enhance our ability to analyze the anatomy of the patellofemoral compartment at scale.
The rapid rise of artificial intelligence (AI) in medicine in the last few years highlights the importance of developing bigger and better systems for data and model sharing. However, the presence of Protected Health Information (PHI) in medical data poses a challenge when it comes to sharing. One potential solution to mitigate the risk of PHI breaches is to exclusively share pre-trained models developed using private datasets. Despite the availability of these pre-trained networks, there remains a need for an adaptable environment to test and fine-tune specific models tailored for clinical tasks. This environment should be open for peer testing, feedback, and continuous model refinement, allowing dynamic model updates that are especially important in the medical field, where diseases and scanning techniques evolve rapidly. In this context, the Discovery Viewer (DV) platform was developed in-house at the Biomedical Engineering and Imaging Institute at Mount Sinai (BMEII) to facilitate the creation and distribution of cutting-edge medical AI models that remain accessible after their development. The all-in-one platform offers a unique environment for non-AI experts to learn, develop, and share their own deep learning (DL) concepts. This paper presents various use cases of the platform, with its primary goal being to demonstrate how DV holds the potential to empower individuals without expertise in AI to create high-performing DL models. We tasked three non-AI experts to develop different musculoskeletal AI projects that encompassed segmentation, regression, and classification tasks. In each project, 80% of the samples were provided with a subset of these samples annotated to aid the volunteers in understanding the expected annotation task. Subsequently, they were responsible for annotating the remaining samples and training their models through the platform's "Training Module". The resulting models were then tested on the separate 20% hold-off dataset to assess their performance. The classification model achieved an accuracy of 0.94, a sensitivity of 0.92, and a specificity of 1. The regression model yielded a mean absolute error of 14.27 pixels. And the segmentation model attained a Dice Score of 0.93, with a sensitivity of 0.9 and a specificity of 0.99. This initiative seeks to broaden the community of medical AI model developers and democratize the access of this technology to all stakeholders. The ultimate goal is to facilitate the transition of medical AI models from research to clinical settings.
To assess the role of pretreatment multiparametric (mp)MRI-based radiomic features in predicting pathologic complete response (pCR) of locally advanced rectal cancer (LARC) to neoadjuvant chemoradiation therapy (nCRT). This was a retrospective dual-center study including 98 patients (M/F 77/21, mean age 60 years) with LARC who underwent pretreatment mpMRI followed by nCRT and total mesorectal excision or watch and wait. Fifty-eight patients from institution 1 constituted the training set and 40 from institution 2 the validation set. Manual segmentation using volumes of interest was performed on T1WI pre-/post-contrast, T2WI and diffusion-weighted imaging (DWI) sequences. Demographic information and serum carcinoembryonic antigen (CEA) levels were collected. Shape, 1st and 2nd order radiomic features were extracted and entered in models based on principal component analysis used to predict pCR. The best model was obtained using a k-fold cross-validation method on the training set, and AUC, sensitivity and specificity for prediction of pCR were calculated on the validation set. Stage distribution was T3 (n = 79) or T4 (n = 19). Overall, 16 (16.3