Background Since the introduction of TotalSegmentator CT, there has been demand for a similar robust automated MRI segmentation tool that can be applied across all MRI sequences and anatomic structures. Purpose To develop and evaluate an automated MRI segmentation model for robust segmentation of major anatomic structures independent of MRI sequence. Materials and Methods In this retrospective study, an nnU-Net model (TotalSegmentator MRI) was trained on MRI and CT scans to segment 80 anatomic structures relevant for use cases such as organ volumetry, disease characterization, surgical planning, and opportunistic screening. Images were randomly sampled from routine clinical studies to represent real-world examples. Dice scores were calculated between the predicted segmentations and expert radiologist segmentations to evaluate model performance on an internal test set and two external test sets and against two publicly available models and TotalSegmentator CT. The Wilcoxon signed rank test was used to compare model performance. The proposed model was applied to a separate internal dataset containing abdominal MRI scans to investigate age-dependent volume changes. Results A total of 1143 scans (616 MRI, 527 CT; median patient age, 61 years [IQR, 50-72 years]) were split into a training set (n = 1088; CT and MRI) and an internal test set (n = 55; MRI only). The two external test sets (AMOS and CHAOS) contained 20 MRI scans each, and the aging-study dataset contained 8672 abdominal MRI scans (median patient age, 59 years [IQR, 45-70 years]). The proposed model had a Dice score of 0.839 for the 80 anatomic structures in the internal test set and outperformed two other models (Dice score of 0.862 vs 0.759 for 40 anatomic structures and 0.838 vs 0.560 for 13 anatomic structures; P < .001 for both). On the TotalSegmentator CT test set (89 CT scans), the performance of the proposed model almost matched that of TotalSegmentator CT (Dice score, 0.966 vs 0.970; P < .001). The aging study demonstrated a strong correlation between age and organ volume (eg, age and liver volume: ρ = -0.096; P < .0001). Conclusion The proposed open-source, easy-to-use model allows for automatic, robust segmentation of 80 structures, extending the capabilities of TotalSegmentator to MRI scans from any MRI sequence. The ready-to-use online tool is available at https://totalsegmentator.com; the model, at https://github.com/wasserth/TotalSegmentator; and the dataset, at http://zenodo.org/records/14710732. © RSNA, 2025 Supplemental material is available for this article. See also the editorial by Kitamura in this issue.
The aim of this study was to develop an open-source nnU-Net-based AI model for combined detection and segmentation of unruptured intracranial aneurysms (UICA) in 3D TOF-MRI and compare models trained on datasets with aneurysm-like differential diagnoses. This retrospective study (2020–2023) included 385 anonymized 3D TOF-MRI images from 345 patients (mean age 59 years, 60% female) at multiple centers plus 113 subjects from the ADAM challenge. Images featured untreated or possible UICA and differential diagnoses. Four distinct training datasets were created, and the nnU-Net framework was used for model development. Performance was assessed on a separate test set using sensitivity and false positive (FP)/case rate for detection and DICE score and NSD (normalized surface distance, 0.5 mm threshold) for segmentation. Segmentation performance on the test set was also compared to a second human reader. The four models achieved overall sensitivity between 82 and 85% and an FP/case rate of 0.20 to 0.31, with no significant differences ( p = 0.90 and p = 0.16) between them. The primary model showed 85% sensitivity and 0.23 FP/case rate, outperforming the ADAM-challenge winner (61%) and a nnU-Net trained on ADAM data (51%) in sensitivity ( p < 0.05). Mean DICE (0.73) and NSD (0.84 for 0.5 mm threshold) for correctly detected UICA did not significantly differ from human reader performance. Our open-source, nnU-Net-based AI model (available at https://zenodo.org/records/13386859 ) demonstrates high sensitivity, low FP rates, and consistent segmentation accuracy for UICA detection and segmentation in 3D TOF-MRI, suggesting its potential to improve clinical diagnosis and monitoring of UICA.
AIM:Spine fractures are a frequent and relevant diagnosis, but systematic documentation is time-consuming and sometimes overlooked. A deep learning pipeline for opportunistic fracture detection in computed tomography (CT) spine images of varying field-of-views is introduced. MATERIALS AND METHODS:This retrospective study builds on 452 CTs of the lumbar/thoracolumbar spine. Patients were included based on the evidence of ≥1 vertebral body fracture and excluded in case of history of spinal surgery or pathologic fractures. The collective was split into training/validation (405) and test (47) sets. An open-source spine dataset was used to train a preliminary segmentation model, which was applied on the training set. The resulting segmentation was post-processed to remove posterior vertebral structures and if needed, manually refined by a radiologist. Using the refined version as new training data, a final segmentation nnU-net was trained. Sagittal slices from each vertebra were labelled individually with regard to fracture evidence. Slices without fracture were used as negative class. Twenty seven thousand nineteen slices (20,396 negative, 6,623 positive) trained a classification algorithm using resnet18. Two senior readers independently assessed fractures in the test set to obtain a consensual ground truth. The segmentation-classification pipeline was applied to the test set and compared with the ground truth. RESULTS:The segmentation model correctly segmented 330/339 (97%) vertebrae. Considering every segmented vertebra, the classifier detected fractures with 88% sensitivity, 95% specificity, and 93% accuracy. CONCLUSION:A deep learning pipeline was built and shown to accurately detect fractures on CT images. The final models as well as our code material are available at https://github.com/usb-radiology/VertebraeFx.
Background Report writing skills are a core competency to be acquired during residency, yet objective tools for tracking performance are lacking. Purpose To investigate whether the Jaccard index, derived from report comparison, can objectively illustrate learning curves in report writing performance throughout radiology residency. Materials and Methods Retrospective data from 246 984 radiology reports written from September 2017 to November 2022 in a tertiary care radiology department were included. Reports were scored using the Jaccard similarity coefficient (ie, a quantitative expression of the amount of edits performed; range, 0-1) of residents' draft (unsupervised initial attempt at a complete report) or preliminary reports (following joint readout with attending physicians) and faculty-reviewed final reports. Weighted mean Jaccard similarity was compared between years of experience using Welch analysis of variance with post hoc testing overall, per imaging division, and per modality. Relationships with years and quarters of resident experience were assessed using Spearman correlation. Results This study included 53 residents (mean report count, 4660 ± 3546; 1-5 years of experience). Mean Jaccard similarity of preliminary reports increased by 6% from 1st-year to 5th-year residents (0.86 ± 0.22 to 0.92 ± 0.15; P < .001). Spearman correlation demonstrated a strong relationship between residents' experience and higher report similarity when aggregated for years (rs = 0.99 [95% CI: 0.85, 1.00]; P < .001) or quarters of experience (rs = 0.90 [95% CI: 0.73, 0.96]; P < .001). For residents' draft reports, Jaccard similarity increased by 14% over the course of the 5-year residency program (0.68 ± 0.27 to 0.82 ± 0.23; P < .001). Subgroup analysis confirmed similar trends for all imaging divisions and modalities (eg, in musculoskeletal imaging, from 0.77 ± 0.31 to 0.91 ± 0.16 [P < .001]; rs = 0.98 [95% CI: 0.72, 1.00] [P < .001]). Conclusion Residents' report writing performance increases with experience. Trends can be quantified with the Jaccard index, with a 6% improvement from 1st- to 5th-year residents, indicating its effectiveness as a tool for evaluating training progress and guiding education over the course of residency. © RSNA, 2024 Supplemental material is available for this article. See also the editorial by Bruno in this issue.
Abstract Background Outcome prediction after catheter ablation for atrial fibrillation (AF) based on electronic health records (EHR) using machine learning (ML) is not yet established. Purpose The aim of the study was to assess the value of ML methods to predict the risk of AF recurrence after catheter ablation based on easily accessible EHR. Methods We analyzed 1362 patients from our prospective registry. Only patients undergoing a first AF ablation were studied. Follow-up was performed at 3, 6 and 12 months after the ablation including 24-hour and 7-day Holter electrocardiogram (ECG). Four different analyses with increasing complexity out of a set of overall 22 features using seven simple and three ensemble method ML algorithms were performed: model 1 (including age, sex, height, weight, and established anamnestic risk factors: AF-type, coronary artery disease, myocardial infarction, stroke, heart failure, hypertension, current smoker, diabetes, renal insufficiency), model 2: model 1 plus additional AF-history features (duration of AF in months, previous typical atrial flutter, number of previous electrical cardioversions, number of failed antiarrhythmic drugs), model 3: model 2 plus simple echocardiographic parameters (left ventricular ejection fraction, LA diameter), model 4: model 3 plus biomarkers (CRP, creatinine, NTproBNP). To evaluate the performance of the prediction, we report C-statistics, accuracy, and recall (sensitivity). The C-statistics of five risk scores for AF recurrence (APPLE, DR-FLASH, BASE-AF2, ATLAS, CHA2DS2-Vasc) and a logistic regression with a forward selection were calculated for comparison. Results Of the 1362 patients, 996 (73%) were male, mean age was 62±10 years and 60% presented with paroxysmal AF. Recurrence of AF during 1-year follow-up was observed in 458 patients (34%). The results for the four models using the eleven different ML algorithms are summarized in the Figure. In detail, the C-statistics were modest with a maximal value of 0.602 for the random forest (RF) classifier for the most complex model 4. The performance for the risk scores was maximal for the APPLE score (C-statistics 0.572) and comparable for the forward-selection logistic regression including five features with a C-statistics of 0.607 (95% CI 0.573-0.641). Conclusion In conclusion, outcome prediction of AF recurrence after catheter ablation with easily accessible EHR data remains challenging and cannot be improved simply by applying machine learning methods. Whether a more comprehensive feature selection, such as more complex ECG parameters instead of the selection of easily available features from the EHR, a combination with deep neural network-based feature extraction, or unsupervised ML might improve the result in a clinically relevant way needs further investigation.Results of the eleven ML algorithms
Introduction: Pulmonary transit time (PTT) is the time it takes blood to pass from the right ventricle to the left ventricle via the pulmonary circulation, making it a potentially useful marker for heart failure. We assessed the association of PTT with diastolic dysfunction (DD) and mitral valve regurgitation (MVR). Methods: We evaluated routine stress perfusion cardiovascular magnetic resonance (CMR) scans in 83 patients including assessment of PTT with simultaneously available echocardiographic assessment. Relevant DD and MVR were defined as exceeding Grade I (impaired relaxation and mild regurgitation). PTT was determined from CMR rest perfusion scans. Normalized PTT (nPTT), adjusted for heart rate, was calculated using Bazett's formula. Results: Higher PTT and nPTT values were associated with higher grade DD and MVR. The diagnostic accuracy for the prediction of DD as quantified by the area under the ROC curve (AUC) was 0.73 (CI 0.61-0.85; p = 0.001) for PTT and 0.81 (CI 0.71-0.89; p < 0.001) for nPTT. For MVR, the diagnostic performance amounted to an AUC of 0.80 (CI 0.68-0.92; p < 0.001) for PTT and 0.78 (CI 0.65-0.90; p < 0.001) for nPTT. PTT values < 8 s rule out the presence of DD and MVR with a probability of 70% (negative predictive value 78%). Conclusion: CMR-derived PTT is a readily obtainable hemodynamic parameter. It is elevated in patients with DD and moderate to severe MVR. Low PTT values make the presence of DD and MVR-as assessed by echocardiography-unlikely.
Purpose:To present a deep learning segmentation model that can automatically and robustly segment all major anatomic structures on body CT images.Materials and Methods:In this retrospective study, 1204 CT examinations (from 2012, 2016, and 2020) were used to segment 104 anatomic structures (27 organs, 59 bones, 10 muscles, and eight vessels) relevant for use cases such as organ volumetry, disease characterization, and surgical or radiation therapy planning. The CT images were randomly sampled from routine clinical studies and thus represent a real-world dataset (different ages, abnormalities, scanners, body parts, sequences, and sites). The authors trained an nnU-Net segmentation algorithm on this dataset and calculated Dice similarity coefficients to evaluate the model's performance. The trained algorithm was applied to a second dataset of 4004 whole-body CT examinations to investigate age-dependent volume and attenuation changes.Results:The proposed model showed a high Dice score (0.943) on the test set, which included a wide range of clinical data with major abnormalities. The model significantly outperformed another publicly available segmentation model on a separate dataset (Dice score, 0.932 vs 0.871; P < .001). The aging study demonstrated significant correlations between age and volume and mean attenuation for a variety of organ groups (eg, age and aortic volume [rs = 0.64; P < .001]; age and mean attenuation of the autochthonous dorsal musculature [rs = -0.74; P < .001]).Conclusion:The developed model enables robust and accurate segmentation of 104 anatomic structures. The annotated dataset (https://doi.org/10.5281/zenodo.6802613) and toolkit (https://www.github.com/wasserth/TotalSegmentator) are publicly available.Keywords: CT, Segmentation, Neural Networks Supplemental material is available for this article. © RSNA, 2023See also commentary by Sebro and Mongan in this issue.
Purpose/objective: Reliable detection of thoracic aortic dilatation (TAD) is mandatory in clinical routine. For ECG -gated CT angiography, automated deep learning (DL) algorithms are established for diameter measurements according to current guidelines. For non-ECG gated CT (contrast enhanced (CE) and non-CE), however, only a few reports are available. In these reports, classification as TAD is frequently unreliable with variable result quality depending on anatomic location with the aortic root presenting with the worst results. Therefore, this study aimed to explore the impact of re-training on a previously evaluated DL tool for aortic measurements in a cohort of non-ECG gated exams. Methods & materials: A cohort of 995 patients (68 +/- 12 years) with CE (n = 392) and non-CE (n = 603) chest CT exams was selected which were classified as TAD by the initial DL tool. The re-trained version featured improved robustness of centerline fitting and cross-sectional plane placement. All cases were processed by the re-trained DL tool version. DL results were evaluated by a radiologist regarding plane placement and diameter measurements. Measure-ments were classified as correctly measured diameters at each location whereas false measurements consisted of over-/under-estimation of diameters. Results: We evaluated 8948 measurements in 995 exams. The re-trained version performed 8539/8948 (95.5%) of diameter measurements correctly. 3765/8948 (42.1%) of measurements were correct in both versions, initial and re-trained DL tool (best: distal arch 655/995 (66%), worst: Aortic sinus (AS) 221/995 (22%)). In contrast, 4456/8948 (49.8%) measurements were correctly measured only by the re-trained version, in particular at the aortic root (AS: 564/995 (57%), sinotubular junction: 697/995 (70%)). In addition, the re-trained version performed 318 (3.6%) measurements which were not available previously. A total of 228 (2.5%) cases showed false measurements because of tilted planes and 181 (2.0%) over-/under-segmentations with a focus at AS (n = 137 (14%) and n = 73 (7%), respectively). Conclusion: Re-training of the DL tool improved diameter assessment, resulting in a total of 95.5% correct measurements. Our data suggests that the re-trained DL tool can be applied even in non-ECG-gated chest CT including both, CE and non-CE exams.
Aims Pulmonary transit time (PTT) is the time blood takes to pass from the right ventricle to the left ventricle via pulmonary circulation. We aimed to quantify PTT in routine cardiovascular magnetic resonance imaging perfusion sequences. PTT may help in the diagnostic assessment and characterization of patients with unclear dyspnoea or heart failure (HF). Methods and results We evaluated routine stress perfusion cardiovascular magnetic resonance scans in 352 patients, including an assessment of PTT. Eighty-six of these patients also had simultaneous quantification of N-terminal pro-brain natriuretic peptide (NTproBNP). NT-proBNP is an established blood biomarker for quantifying ventricular filling pressure in patients with presumed HF. Manually assessed PTT demonstrated low inter-rater variability with a correlation between raters >0.98. PTT was obtained automatically and correctly in 266 patients using artificial intelligence. The median PTT of 182 patients with both left and right ventricular ejection fraction >50% amounted to 6.8 s (Pulmonary transit time: 5.9-7.9 s). PTT was significantly higher in patients with reduced left ventricular ejection fraction (<40%; P < 0.001) and right ventricular ejection fraction (<40%; P < 0.0001). The area under the receiver operating characteristics curve (AUC) of PTT for exclusion of HF (NT-proBNP <125 ng/L) was 0.73 (P < 0.001) with a specificity of 77% and sensitivity of 70%. The AUC of PTT for the inclusion of HF (NT-proBNP >600 ng/L) was 0.70 (P < 0.001) with a specificity of 78% and sensitivity of 61%. Conclusion PTT as an easily, even automatically obtainable and robust non-invasive biomarker of haemodynamics might help in the evaluation of patients with dyspnoea and HF.
Rationale and Objectives: To assess the effects of a change from free text reporting to structured reporting on resident reports, the proofreading workload and report turnaround times in the neuroradiology daily routine.Materials and Methods: Our neuroradiology section introduced structured reporting templates in July 2019. Reports dictated by resi-dents during dayshifts from January 2019 to March 2020 were retrospectively assessed using quantitative parameters from report com-parison. Through automatic analysis of text-string differences between report states (i.e. draft, preliminary and final report), Jaccard similarities and edit distances of reports following read-out sessions as well as after report sign-off were calculated. Furthermore, turn-around times until preliminary and final report availability to clinicians were investigated. Parameters were visualized as trending line graphs and statistically compared between reporting standards.Results: Three thousand five hundred thirty-eight reports were included into analysis. Mean Jaccard similarity of resident drafts and staff-reviewed final reports increased from 0.53 +/- 0.37 to 0.79 +/- 0.22 after the introduction of structured reporting (p < .001). Both mean overall edits on draft reports by residents following read-out sessions (0.30 +/- 0.45 vs. 0.09 +/- 0.29; p < .001) and by staff radiologists during report sign-off (0.17 +/- 0.28 vs. 0.12 +/- 0.23, p < .001) decreased. With structured reporting, mean turnaround time until preliminary report availability to clinicians decreased by 20.7 minutes (246.9 +/- 207.0 vs. 226.2 +/- 224.9; p < .001). Similarly, final reports were available 35.0 minutes faster on average (558.05 +/- 15.1 vs. 523.0 +/- 497.3; p = .002).Conclusion: Structured reporting is beneficial in the neuroradiology daily routine, as resident drafts require fewer edits in the report review process. This reduction in proofreading workload is likely responsible for lower report turnaround times.
Purpose Thoracic aortic (TA) dilatation (TAD) is a risk factor for acute aortic syndrome and must therefore be reported in every CT report. However, the complex anatomy of the thoracic aorta impedes TAD detection. We investigated the performance of a deep learning (DL) prototype as a secondary reading tool built to measure TA diameters in a large-scale cohort. Material and methods Consecutive contrast-enhanced (CE) and non-CE chest CT exams with “normal” TA diameters according to their radiology reports were included. The DL-prototype (AIRad, Siemens Healthineers, Germany) measured the TA at nine locations according to AHA guidelines. Dilatation was defined as >45 mm at aortic sinus, sinotubular junction (STJ), ascending aorta (AA) and proximal arch and >40 mm from mid arch to abdominal aorta. A cardiovascular radiologist reviewed all cases with TAD according to AIRad. Multivariable logistic regression (MLR) was used to identify factors (demographics and scan parameters) associated with TAD classification by AIRad. Results 18,243 CT scans (45.7% female) were successfully analyzed by AIRad. Mean age was 62.3 ± 15.9 years and 12,092 (66.3%) were CE scans. AIRad confirmed normal diameters in 17,239 exams (94.5%) and reported TAD in 1,004/18,243 exams (5.5%). Review confirmed TAD classification in 452/1,004 exams (45.0%, 2.5% total), 552 cases were false-positive but identification was easily possible using visual outputs by AIRad. MLR revealed that the following factors were significantly associated with correct TAD classification by AIRad: TAD reported at AA [odds ratio (OR): 1.12, p < 0.001] and STJ (OR: 1.09, p = 0.002), TAD found at >1 location (OR: 1.42, p = 0.008), in CE exams (OR: 2.1–3.1, p < 0.05), men (OR: 2.4, p = 0.003) and patients presenting with higher BMI (OR: 1.05, p = 0.01). Overall, 17,691/18,243 (97.0%) exams were correctly classified. Conclusions AIRad correctly assessed the presence or absence of TAD in 17,691 exams (97%), including 452 cases with previously missed TAD independent from contrast protocol. These findings suggest its usefulness as a secondary reading tool by improving report quality and efficiency.
Objectives An end-to-end method is introduced to build a combined segmentation-classification pipeline using deep learning for opportunistic fracture detection in CT spine images of varying field-of-views. Materials and Methods This retrospective study builds on 452 CTs of the lumbar/thoracolumbar spine. Patients were included based on the evidence of ≥ 1 vertebral body fracture and excluded in case of history of spinal surgery or pathologic fractures. The collective was split into training/validation (405) and test (47) sets. An open-source pre-segmented spine dataset was used to train a preliminary segmentation model, which was applied on the training set. The resulting segmentation was post-processed to remove posterior vertebral structures and if needed manually refined by a radiologist. Using the refined version as new training data, a final segmentation nnU-net was trained. Sagittal slices from each vertebra were labelled individually with regard to fracture evidence. Slices without signs of fracture were used as negative class. 27,019 slices (20,396 negative, 6,623 positive) trained a classification algorithm using resnet18. Two senior readers independently assessed fractures in the test set to obtain a consensual ground truth. The segmentation-classification pipeline was applied to the test set and compared to the ground truth. Results The segmentation model correctly segmented 330/339 (97%) vertebrae. Considering every segmented vertebra, the classifier detected fractures with 88% sensitivity, 95% specificity and 93% accuracy. Conclusion Our two-step method can help to detect spine fractures on images of varying field-of-views, with an accuracy comparable to that of a radiologist in-training. The final models as well as our code material are available at . ### Competing Interest Statement The authors have declared no competing interest. ### Funding Statement This study did not receive any funding ### Author Declarations I confirm all relevant ethical guidelines have been followed, and any necessary IRB and/or ethics committee approvals have been obtained. Yes The details of the IRB/oversight body that provided approval or exemption for the research described are given below: Ethics committee/IRB of the University Hospital Basel gave ethical approval for this work. I confirm that all necessary patient/participant consent has been obtained and the appropriate institutional forms have been archived, and that any patient/participant/sample identifiers included were not known to anyone (e.g., hospital staff, patients or participants themselves) outside the research group so cannot be used to identify individuals. Yes I understand that all clinical trials and any other prospective interventional studies must be registered with an ICMJE-approved registry, such as ClinicalTrials.gov. I confirm that any such study reported in the manuscript has been registered and the trial registration ID is provided (note: if posting a prospective study registered retrospectively, please provide a statement in the trial ID field explaining why the study was not registered in advance). Yes I have followed all appropriate research reporting guidelines and uploaded the relevant EQUATOR Network research reporting checklist(s) and other pertinent material as supplementary files, if applicable. Yes All code data in the present study are available online at .
Background Workloads in radiology departments have constantly increased over the past decades. The resulting radiologist fatigue is considered a rising problem that affects diagnostic accuracy. Purpose To investigate whether data mining of quantitative parameters from the report proofreading process can reveal daytime and shift-dependent trends in report similarity as a surrogate marker for resident fatigue. Materials and Methods Data from 117 402 radiology reports written by residents between September 2017 and March 2020 were extracted from a report comparison tool and retrospectively analyzed. Through calculation of the Jaccard similarity coefficient between residents' preliminary and staff-reviewed final reports, the amount of edits performed by staff radiologists during proofreading was quantified on a scale of 0 to 1 (1: perfect similarity, no edits). Following aggregation per weekday and shift, data were statistically analyzed by using simple linear regression or one-way analysis of variance (significance level, P < .05) to determine relationships between report similarity and time of day and/or weekday reports were dictated. Results Decreasing report similarity with increasing work hours was observed for day shifts (r = -0.93 [95% CI: -0.73, -0.98]; P < .001) and weekend shifts (r = -0.72 [95% CI: -0.31, -0.91]; P = .004). For day shifts, negative linear correlation was strongest on Fridays (r = -0.95 [95% CI: -0.80, -0.99]; P < .001), with a 16% lower mean report similarity at the end of shifts (0.85 ± 0.24 at 8 am vs 0.69 ± 0.32 at 5 pm). Furthermore, mean similarity of reports dictated on Fridays (0.79 ± 0.35) was lower than that on all other weekdays (range, 0.84 ± 0.30 to 0.86 ± 0.27; P < .001). For late shifts, report similarity showed a negative correlation with the course of workweeks, showing a continuous decrease from Monday to Friday (r = -0.98 [95% CI: -0.70, -0.99]; P = .007). Temporary increases in report similarity were observed after lunch breaks (day and weekend shifts) and with the arrival of a rested resident during overlapping on-call shifts. Conclusion Decreases in report similarity over the course of workdays and workweeks suggest aggravating effects of fatigue on residents' report writing performances. Periodic breaks within shifts potentially foster recovery. © RSNA, 2021.
BACKGROUND:Manually performed diameter measurements on ECG-gated CT-angiography (CTA) represent the gold standard for diagnosis of thoracic aortic dilatation. However, they are time-consuming and show high inter-reader variability. Therefore, we aimed to evaluate the accuracy of measurements of a deep learning-(DL)-algorithm in comparison to those of radiologists and evaluated measurement times (MT).METHODS:We retrospectively analyzed 405 ECG-gated CTA exams of 371 consecutive patients with suspected aortic dilatation between May 2010 and June 2019. The DL-algorithm prototype detected aortic landmarks (deep reinforcement learning) and segmented the lumen of the thoracic aorta (multi-layer convolutional neural network). It performed measurements according to AHA-guidelines and created visual outputs. Manual measurements were performed by radiologists using centerline technique. Human performance variability (HPV), MT and DL-performance were analyzed in a research setting using a linear mixed model based on 21 randomly selected, repeatedly measured cases. DL-algorithm results were then evaluated in a clinical setting using matched differences. If the differences were within 5 mm for all locations, the cases was regarded as coherent; if there was a discrepancy >5 mm at least at one location (incl. missing values), the case was completely reviewed.RESULTS:HPV ranged up to ±3.4 mm in repeated measurements under research conditions. In the clinical setting, 2,778/3,192 (87.0%) of DL-algorithm's measurements were coherent. Mean differences of paired measurements between DL-algorithm and radiologists at aortic sinus and ascending aorta were -0.45±5.52 and -0.02±3.36 mm. Detailed analysis revealed that measurements at the aortic root were over-/underestimated due to a tilted measurement plane. In total, calculated time saved by DL-algorithm was 3:10 minutes/case.CONCLUSIONS:The DL-algorithm provided coherent results to radiologists at almost 90% of measurement locations, while the majority of discrepent cases were located at the aortic root. In summary, the DL-algorithm assisted radiologists in performing AHA-compliant measurements by saving 50% of time per case.
Artificial intelligence can assist in cardiac image interpretation. Here, we achieved a substantial reduction in time required to read a cardiovascular magnetic resonance (CMR) study to estimate left atrial volume without compromising accuracy or reliability. Rather than deploying a fully automatic black-box, we propose to incorporate the automated LA volumetry into a human-centric interactive image-analysis process. Atri-U, an automated data analysis pipeline for long-axis cardiac cine images, computes the atrial volume by: (i) detecting the end-systolic frame, (ii) outlining the endocardial borders of the LA, (iii) localizing the mitral annular hinge points and constructing the longitudinal atrial diameters, equivalent to the usual workup done by clinicians. In every step human interaction is possible, such that the results provided by the algorithm can be accepted, corrected, or re-done from scratch. Atri-U was trained and evaluated retrospectively on a sample of 300 patients and then applied to a consecutive clinical sample of 150 patients with various heart conditions. The agreement of the indexed LA volume between Atri-U and two experts was similar to the inter-rater agreement between clinicians (average overestimation of 0.8 mL/m2 with upper and lower limits of agreement of − 7.5 and 5.8 mL/m2, respectively). An expert cardiologist blinded to the origin of the annotations rated the outputs produced by Atri-U as acceptable in 97% of cases for step (i), 94% for step (ii) and 95% for step (iii), which was slightly lower than the acceptance rate of the outputs produced by a human expert radiologist in the same cases (92%, 100% and 100%, respectively). The assistance of Atri-U lead to an expected reduction in reading time of 66%—from 105 to 34 s, in our in-house clinical setting. Our proposal enables automated calculation of the maximum LA volume approaching human accuracy and precision. The optional user interaction is possible at each processing step. As such, the assisted process sped up the routine CMR workflow by providing accurate, precise, and validated measurement results.
OBJECTIVES:To evaluate the performance of a deep convolutional neural network (DCNN) in detecting and classifying distal radius fractures, metal, and cast on radiographs using labels based on radiology reports. The secondary aim was to evaluate the effect of the training set size on the algorithm's performance.METHODS:A total of 15,775 frontal and lateral radiographs, corresponding radiology reports, and a ResNet18 DCNN were used. Fracture detection and classification models were developed per view and merged. Incrementally sized subsets served to evaluate effects of the training set size. Two musculoskeletal radiologists set the standard of reference on radiographs (test set A). A subset (B) was rated by three radiology residents. For a per-study-based comparison with the radiology residents, the results of the best models were merged. Statistics used were ROC and AUC, Youden's J statistic (J), and Spearman's correlation coefficient (ρ).RESULTS:The models' AUC/J on (A) for metal and cast were 0.99/0.98 and 1.0/1.0. The models' and residents' AUC/J on (B) were similar on fracture (0.98/0.91; 0.98/0.92) and multiple fragments (0.85/0.58; 0.91/0.70). Training set size and AUC correlated on metal (ρ = 0.740), cast (ρ = 0.722), fracture (frontal ρ = 0.947, lateral ρ = 0.946), multiple fragments (frontal ρ = 0.856), and fragment displacement (frontal ρ = 0.595).CONCLUSIONS:The models trained on a DCNN with report-based labels to detect distal radius fractures on radiographs are suitable to aid as a secondary reading tool; models for fracture classification are not ready for clinical use. Bigger training sets lead to better models in all categories except joint affection.KEY POINTS:• Detection of metal and cast on radiographs is excellent using AI and labels extracted from radiology reports. • Automatic detection of distal radius fractures on radiographs is feasible and the performance approximates radiology residents. • Automatic classification of the type of distal radius fracture varies in accuracy and is inferior for joint involvement and fragment displacement.
The use of artificial intelligence (AI) is a powerful tool for image analysis that is increasingly being evaluated by radiology professionals. However, due to the fact that these methods have been developed for the analysis of nonmedical image data and data structure in radiology departments is not “AI ready”, implementing AI in radiology is not straightforward. The purpose of this review is to guide the reader through the pipeline of an AI project for automated image analysis in radiology and thereby encourage its implementation in radiology departments. At the same time, this review aims to enable readers to critically appraise articles on AI-based software in radiology.
Objective To assess the diagnostic performance of a deep learning-based algorithm for automated detection of acute and chronic rib fractures on whole-body trauma CT. Materials and Methods We retrospectively identified all whole-body trauma CT scans referred from the emergency department of our hospital from January to December 2018 (n = 511). Scans were categorized as positive (n = 159) or negative (n = 352) for rib fractures according to the clinically approved written CT reports, which served as the index test. The bone kernel series (1.5-mm slice thickness) served as an input for a detection prototype algorithm trained to detect both acute and chronic rib fractures based on a deep convolutional neural network. It had previously been trained on an independent sample from eight other institutions (n = 11455). Results All CTs except one were successfully processed (510/511). The algorithm achieved a sensitivity of 87.4% and specificity of 91.5% on a per-examination level [per CT scan: rib fracture(s): yes/no]. There were 0.16 false-positives per examination (= 81/510). On a per-finding level, there were 587 true-positive findings (sensitivity: 65.7%) and 307 false-negatives. Furthermore, 97 true rib fractures were detected that were not mentioned in the written CT reports. A major factor associated with correct detection was displacement. Conclusion We found good performance of a deep learning-based prototype algorithm detecting rib fractures on trauma CT on a per-examination level at a low rate of false-positives per case. A potential area for clinical application is its use as a screening tool to avoid false-negative radiology reports.
Purpose: To design and evaluate a self-trainable natural language processing (NLP)-based procedure to classify unstructured radiology reports. The method enabling the generation of curated datasets is exemplified on CT pulmonary angiogram (CTPA) reports. Method: We extracted the impressions of CTPA reports created at our institution from 2016 to 2018 (n = 4397; language: German). The status (pulmonary embolism: yes/no) was manually labelled for all exams. Data from 2016/2017 (n = 2801) served as a ground truth to train three NLP architectures that only require a subset of reference datasets for training to be operative. The three architectures were as follows: a convolutional neural network (CNN), a support vector machine (SVM) and a random forest (RF) classifier. Impressions of 2018 (n = 1377) were kept aside and used for general performance measurements. Furthermore, we investigated the dependence of classification performance on the amount of training data with multiple simulations. Results: The classification performance of all three models was excellent (accuracies: 97 %-99 %; F1 scores 0.88-0.97; AUCs: 0.993-0.997). Highest accuracy was reached by the CNN with 99.1 % (95 % CI 98.5-99.6 %). Training with 470 labelled impressions was sufficient to reach an accuracy of > 93 % with all three NLP architectures. Conclusion: Our NLP-based approaches allow for an automated and highly accurate retrospective classification of CTPA reports with manageable effort solely using unstructured impression sections. We demonstrated that this approach is useful for the classification of radiology reports not written in English. Moreover, excellent classification performance is achieved at relatively small training set sizes.
Objectives To investigate the most common errors in residents’ preliminary reports, if structured reporting impacts error types and frequencies, and to identify possible implications for resident education and patient safety. Material and methods Changes in report content were tracked by a report comparison tool on a word level and extracted for 78,625 radiology reports dictated from September 2017 to December 2018 in our department. Following data aggregation according to word stems and stratification by subspecialty (e.g., neuroradiology) and imaging modality, frequencies of additions/deletions were analyzed for findings and impression report section separately and compared between subgroups. Results Overall modifications per report averaged 4.1 words, with demonstrably higher amounts of changes for cross-sectional imaging (CT: 6.4; MRI: 6.7) than non-cross-sectional imaging (radiographs: 0.2; ultrasound: 2.8). The four most frequently changed words (right, left, one, and none) remained almost similar among all subgroups (range: 0.072–0.117 per report; once every 9–14 reports). Albeit representing only 0.02% of analyzed words, they accounted for up to 9.7% of all observed changes. Subspecialties solely using structured reporting had substantially lower change ratios in the findings report section (mean: 0.2 per report) compared with prose-style reporting subspecialties (mean: 2.0). Relative frequencies of the most changed words remained unchanged. Conclusion Residents’ most common reporting errors in all subspecialties and modalities are laterality discriminator confusions (left/right) and unnoticed descriptor misregistration by speech recognition (one/none). Structured reporting reduces overall error rates, but does not affect occurrence of the most common errors. Increased error awareness and measures improving report correctness and ensuring patient safety are required. Key Points • The two most common reporting errors in residents’ preliminary reports are laterality discriminator confusions (left/right) and unnoticed descriptor misregistration by speech recognition (one/none). • Structured reporting reduces the overall the error frequency in the findings report section by a factor of 10 (structured reporting: mean 0.2 per report; prose-style reporting: 2.0) but does not affect the occurrence of the two major errors. • Staff radiologist review behavior noticeably differs between radiology subspecialties.