Artificial intelligence (AI) can automatically delineate lesions on computed tomography (CT) and generate radiology report content, yet progress is limited by the scarcity of publicly available CT datasets with lesion-level annotations. To bridge this gap, we introduce CT-Bench, a first-of-its-kind benchmark dataset comprising two components: a Lesion Image and Metadata Set containing 20,335 lesions from 7,795 CT studies with bounding boxes, descriptions, and size information, and a multitask visual question answering benchmark with 2,850 QA pairs covering lesion localization, description, size estimation, and attribute categorization. Hard negative examples are included to reflect real-world diagnostic challenges. We evaluate multiple state-of-the-art multimodal models, including vision-language and medical CLIP variants, by comparing their performance to radiologist assessments, demonstrating the value of CT-Bench as a comprehensive benchmark for lesion analysis. Moreover, fine-tuning models on the Lesion Image and Metadata Set yields significant performance gains across both components, underscoring the clinical utility of CT-Bench.
Multimodal large language models have significantly improved in assessing RSNA Case of the Day Questions over the past year, with OpenAI’s latest models outperforming those from Google and Meta; notably, OpenAI o1 had an accuracy comparable to expert radiologists.
Objective To report data from the first three years of operation of the RSNA-ACR 3D Printing Registry. Methods Data from June 2020 to June 2023 was extracted, including demographics, indications, workflow and user assessments. Clinical indications were stratified by 12 organ systems. Imaging modalities, printing technologies and number of parts per case were assessed. Effort data was analyzed, dividing staff into provider and non-provider categories. The opinions of clinical users were evaluated through a Likert-scale questionnaire, and estimates of procedure time saved were collected. Results A total of 20 sites and 2,637 cases were included, consisting of 1,863 anatomic models and 774 anatomic guides. Mean patient age for models and guides was 42.4 ± 24.5 years and 56.3 ± 18.5 years respectively. Cardiac models were the most common type of models (27.2%), and neurologic guides were the most common type of guides (42.4%). Material jetting, vat photopolymerization and material extrusion were the most common printing technologies used overall (85.6% of all cases). On average, providers spent 92.4 minutes and non-providers spent 335.0 minutes per case. Providers spent most time on consultation (33.6 minutes), while non-providers focused most on segmentation (148.0 minutes). Confidence in treatment plans increased after using 3D printing (p<.001). Estimated procedure time savings for 155 cases was 40.5 ± 26.1 minutes. Conclusion 3D printing is performed in healthcare facilities for many clinical indications. The registry provides insight into the technologies and workflows used to create anatomic models and guides, and the data shows clinical benefits from 3D printing.
Imaging utilization has increased dramatically in recent years, and at least some of these studies are not appropriate for the clinical scenario. The development of large language models (LLMs) may address this issue by providing a more accessible reference resource for ordering providers, but their relative performance is currently understudied. Evaluate and compare the relative appropriateness and usefulness of imaging recommendations generated by eight publicly available models in response to neuroradiology clinical scenarios. Twenty-four common neuroradiology clinical scenarios were selected which often yield suboptimal imaging utilization. Questions were crafted to assess the ability of LLMs to provide accurate and actionable advice. The LLMs were assessed in August 2023 using natural-language 1-2 sentence queries requesting advice about optimal image ordering given certain clinical parameters. Eight of the most well-known LLMs were chosen for evaluation: ChatGPT, GPT4, Bard (Versions 1 and 2), Bing Chat, Llama 2, Perplexity, and Claude. The models were graded by three fellowship-trained neuroradiologists on whether their advice was "optimal" or "not optimal" according to the ACR Appropriateness Criteria or the New Orleans Head CT Criteria. The raters also ranked the models based on the appropriateness, helpfulness, concision, and source-citations in their response. The models varied in their ability to deliver an "optimal" recommendation based on these scenarios as follows: ChatGPT (20/24), GPT4 (23/24), Bard 1 (13/24), Bard 2 (14/24), Bing Chat (14/24), Llama (5/24), Perplexity (19/24), and Claude (19/24). The median ranks of the LLMs were as follows: ChatGPT (3), GPT4 (1.5), Bard 1 (4.5), Bard 2 (5), Bing Chat (6), Llama (7.5), Perplexity (4), and Claude (3). Characteristic errors are described and discussed. GPT-4, ChatGPT, and Claude generally outperformed Bard, Bing Chat, and Llama 2. This study evaluates the performance of a greater variety of publicly available LLMs in settings that more closely mimic real-world use cases as well as discussing the practical challenges of doing so. This is the first study to evaluate and compare a wide range of publicly available LLMs to determine appropriateness of their neuroradiology imaging recommendations.
Background GPT-4V (GPT-4 with vision, ChatGPT; OpenAI) has shown impressive performance in several medical assessments. However, few studies have assessed its performance in interpreting radiologic images. Purpose To assess and compare the accuracy of GPT-4V in assessing radiologic cases with both images and textual context to that of radiologists and residents, to assess if GPT-4V assistance improves human accuracy, and to assess and compare the accuracy of GPT-4V with that of image-only or text-only inputs. Materials and Methods Seventy-two Case of the Day questions at the RSNA 2023 Annual Meeting were curated in this observer study. Answers from GPT-4V were obtained between November 26 and December 10, 2023, with the following inputs for each question: image only, text only, and both text and images. Five radiologists and three residents also answered the questions in an "open book" setting. For the artificial intelligence (AI)-assisted portion, the radiologists and residents were provided with the outputs of GPT-4V. The accuracy of radiologists and residents, both with and without AI assistance, was analyzed using a mixed-effects linear model. The accuracies of GPT-4V with different input combinations were compared by using the McNemar test. P < .05 was considered to indicate a significant difference. Results The accuracy of GPT-4V was 43% (31 of 72; 95% CI: 32, 55). Radiologists and residents did not significantly outperform GPT-4V in either imaging-dependent (59% and 56% vs 39%; P = .31 and .52, respectively) or imaging-independent (76% and 63% vs 70%; both P = .99) cases. With access to GPT-4V responses, there was no evidence of improvement in the average accuracy of the readers. The accuracy obtained by GPT-4V with text-only and image-only inputs was 50% (35 of 70; 95% CI: 39, 61) and 38% (26 of 69; 95% CI: 27, 49), respectively. Conclusion The radiologists and residents did not significantly outperform GPT-4V. Assistance from GPT-4V did not help human raters. GPT-4V relied on the textual context for its outputs. © RSNA, 2024 Supplemental material is available for this article. See also the editorial by Katz in this issue.
From basic research to the bedside, precise terminology is key to advancing medicine and ensuring optimal and appropriate patient care. However, the wide spectrum of diseases and their manifestations superimposed on medical team-specific and discipline-specific communication patterns often impairs shared understanding and the shared use of common medical terminology. Common terms are currently used in medicine to ensure interoperability and facilitate integration of biomedical information for clinical practice and emerging scientific and educational applications alike, from database integration to supporting basic clinical operations such as billing. Such common terminologies can be provided in ontologies, which are formalized representations of knowledge in a particular domain. Ontologies unambiguously specify common concepts and describe the relationships between those concepts by using a form that is mathematically precise and accessible to humans and machines alike. RadLex® is a key RSNA initiative that provides a shared domain model, or ontology, of radiology to facilitate integration of information in radiology education, clinical care, and research. As the contributions of the computational components of common radiologic workflows continue to increase with the ongoing development of big data, artificial intelligence, and novel image analysis and visualization tools, the use of common terminologies is becoming increasingly important for supporting seamless computational resource integration across medicine. This article introduces ontologies, outlines the fundamental semantic web technologies used to create and apply RadLex, and presents examples of RadLex applications in everyday radiology and research. It concludes with a discussion of emerging applications of RadLex, including artificial intelligence applications. © RSNA, 2023 Quiz questions for this article are available in the supplemental material.
MRI is an essential diagnostic imaging modality for many knee conditions; however, it is not indicated in the setting of advanced knee arthritis. Inappropriate MRI imaging adds to health care costs and may delay definitive management for many patients. The primary purpose of this study was to ascertain the frequency of inappropriate MRI scans performed at one Veterans' Administration Medical Center (VAMC). We performed a retrospective chart review of all knee MRIs ordered over a 6-month period. Inappropriate MRI was defined as MRI performed prior to radiographs (XRs), or in the presence of XRs demonstrating severe osteoarthritis, without leading to a nonarthroplasty procedure of the knee. Of the 304 cases reviewed, 36.8% (112) of the MRIs were deemed inappropriate, 33 were ordered by orthopedists, and 79 were ordered by other health care providers. Of the 33 ordered by orthopedists, 25 were ordered by retired/nonsurgical orthopedists. Obtaining an MRI delayed care by an average of 29.2 days. Of the 252 cases that had XR prior to MRI, none included all four views in the standard knee XR series and only four had weightbearing images. Over a third of knee MRIs performed at this VAMC were inappropriate and delayed care. Additionally, no XRs in our study contained all the necessary views to properly assess knee arthritis. These concerning findings signify a potential opportunity for education in diagnostic strategies, to better patient care and resource utilization in the VAMC.
The purpose of this study was to evaluate the feasibility of translation of RadLex lexicon from English to German performed by Google Translate, using the RadLex ontology as ground truth. The same comparison was also performed for German to English translations. We determined the concordance rate of the Google Translate-rendered translations (for both English to German and German to English translations) to the official German RadLex (translations provided by the German Radiological Society) and English RadLex terms via character-by-character concordance analysis (string matching). Specific term characteristics of term character count and word count were compared between concordant and discordant translations using t-tests. Google Translate-rendered translations originally considered incongruent (2482 English terms and 2500 German terms) were then reviewed by German and English-speaking radiologists to further evaluate clinical utility. Overall success rates of both methods were calculated by adding the percentage of terms marked correct by string comparison to the percentage marked correct during manual review extrapolated to the terms that had been initially marked incorrect during string analysis. 64,632 English and 47,425 German RadLex terms were analyzed. 3507 (5.4%) of the Google Translate-rendered English to German translations were concordant with the official German RadLex terms when evaluated via character-by-character concordance. 3288 (6.9%) of the Google Translate-rendered German to English translations matched the corresponding English RadLex terms. Human review of a random sample of non-concordant machine translations revealed that 95.5% of such English to German translations were understandable, whereas 43.9% of such German to English translations were understandable. Combining both string matching and human review resulted in an overall Google Translate success rate of 95.7% for English to German translations and 47.8% for German to English translations. For certain radiologic text translation tasks, Google Translate may be a useful tool for translating multi-language radiology reports into a common language for natural language processing and subsequent labeling of datasets for machine learning. Indeed, string matching analysis alone is an incomplete method for evaluating machine translation. However, when human review of automated translation is also incorporated, measured performance improves. Additional evaluation using longer text samples and full imaging reports is needed. An apparent discordance between English to German versus German to English translation suggests that the direction of translation affects accuracy.
OBJECTIVE:There is a paucity of utility and cost data regarding the launch of 3D printing in a hospital. The objective of this project is to benchmark utility and costs for radiology-based in-hospital 3D printing of anatomic models in a single, adult academic hospital.METHODS:All consecutive patients for whom 3D printed anatomic models were requested during the first year of operation were included. All 3D printing activities were documented by the 3D printing faculty and referring specialists. For patients who underwent a procedure informed by 3D printing, clinical utility was determined by the specialist who requested the model. A new metric for utility termed Anatomic Model Utility Points with range 0 (lowest utility) to 500 (highest utility) was derived from the specialist answers to Likert statements. Costs expressed in United States dollars were tallied from all 3D printing human resources and overhead. Total costs, focused costs, and outsourced costs were estimated. The specialist estimated the procedure room time saved from the 3D printed model. The time saved was converted to dollars using hospital procedure room costs.RESULTS:The 78 patients referred for 3D printed anatomic models included 11 clinical indications. For the 68 patients who had a procedure, the anatomic model utility points had an overall mean (SD) of 312 (57) per patient (range, 200-450 points). The total operation cost was $213,450. The total cost, focused costs, and outsourced costs were $2,737, $2,180, and $2,467 per model, respectively. Estimated procedure time saved had a mean (SD) of 29.9 (12.1) min (range, 0-60 min). The hospital procedure room cost per minute was $97 (theoretical $2,900 per patient saved with model).DISCUSSION:Utility and cost benchmarks for anatomic models 3D printed in a hospital can inform health care budgets. Realizing pecuniary benefit from the procedure time saved requires future research.
The existing fellowship imaging informatics curriculum, established in 2004, has not undergone formal revision since its inception and inaccurately reflects present-day radiology infrastructure. It insufficiently equips trainees for today’s informatics challenges as current practices require an understanding of advanced informatics processes and more complex system integration. We sought to address this issue by surveying imaging informatics fellowship program directors across the country to determine the components and cutline for essential topics in a standardized imaging informatics curriculum, the consensus on essential versus supplementary knowledge, and the factors individual programs may use to determine if a newly developed topic is an essential topic. We further identified typical program structural elements and sought fellowship director consensus on offering official graduate trainee certification to imaging informatics fellows. Here, we aim to provide an imaging informatics fellowship director consensus on topics considered essential while still providing a framework for informatics fellowship programs to customize their individual curricula.
In the American healthcare system, proper documentation is required in order to ensure accurate medical billing and reimbursement using the current procedural terminology (CPT) code system. Today, limited reimbursement for three-dimensional (3D) printed anatomic models and guides may be obtained using the established Category III CPT codes for 3D printing. To capture information about the number and types of 3D printed anatomic models and guides that are being created in hospitals, the Radiological Society of North America and the American College of Radiology established a 3D printing registry to collect data which will be used to support future efforts relating to reimbursement for clinical 3D printing. This chapter will discuss proper documentation for 3D printed anatomic models and guides, CPT codes for 3D printing, and current efforts to demonstrate widespread use of 3D printing in medicine.
Data collection and labeling is one of the main challenges in employing machine learning algorithms in a variety of real-world applications with limited data. While active learning methods attempt to tackle this issue by labeling only the data samples that give high information, they generally suffer from large computational costs and are impractical in settings where data can be collected in parallel. Batch active learning methods attempt to overcome this computational burden by querying batches of samples at a time. To avoid redundancy between samples, previous works rely on some ad hoc combination of sample quality and diversity. In this paper, we present a new principled batch active learning method using Determinantal Point Processes, a repulsive point process that enables generating diverse batches of samples. We develop tractable algorithms to approximate the mode of a DPP distribution, and provide theoretical guarantees on the degree of approximation. We further demonstrate that an iterative greedy method for DPP maximization, which has lower computational costs but worse theoretical guarantees, still gives competitive results for batch active learning. Our experiments show the value of our methods on several datasets against state-of-the-art baselines.
We compare the continuous and discrete truncated Wigner approximations - TWA and dTWA, respectively - of various spin models' dynamics to exact analytical and numerical solutions. We account for all components of spin-spin correlations on equal footing, facilitated by a recently introduced geometric correlation matrix visualization (CMV) technique. We find that the two approximations capture different aspects of the dynamics of correlations: dTWA captures periodic partial revivals of spin-spin correlations missed in the TWA; but at least for some conditions, TWA better captures the directions in which spins have correlations. We also find that at modestly short times, the dominant error in both approximations is to substantially suppress spin correlations along one direction. In addition to assessing the accuracy of approximations, our comparisons illustrate the utility of CMVs as an intuitive visual tool to understand the full complexity of spin correlations.
To demonstrate the 3D printed appearance of glenoid morphologies relevant to shoulder replacement surgery and to evaluate the benefits of printed models of the glenoid with regard to surgical planning. A retrospective review of patients referred for shoulder CT was performed, leading to a cohort of nine patients without arthroplasty hardware and exhibiting glenoid changes relevant to shoulder arthroplasty planning. Thin slice CT images were used to create both humerus-subtracted volume renderings of the glenoid, as well as 3D surface models of the glenoid, and 11 printed models were created. Volume renderings, surface models, and printed models were reviewed by a musculoskeletal radiologist for accuracy. Four fellowship-trained orthopaedic surgeons specializing in shoulder surgery reviewed each case individually as follows: First, the source CT images were reviewed, and a score for the clarity of the bony morphologies relevant to shoulder arthroplasty surgery was given. The volume rendering was reviewed, and the clarity was again scored. Finally, the printed model was reviewed, and the clarity again scored. Each printed model was also scored for morphologic complexity, expected usefulness of the printed model, and physical properties of the model. Mann–Whitney–Wilcoxon signed rank tests of the clarity scores were calculated, and the Spearman’s ρ correlation coefficient between complexity and usefulness scores was computed. Printed models demonstrated a range of glenoid bony changes including osteophytes, glenoid bone loss, retroversion, and biconcavity. Surgeons rated the glenoid morphology as more clear after review of humerus-subtracted volume rendering, compared with review of the source CT images (p = 0.00903). Clarity was also better with 3D printed models compared to CT (p = 0.00903) and better with 3D printed models compared to humerus-subtracted volume rendering (p = 0. 00879). The expected usefulness of printed models demonstrated a positive correlation with morphologic complexity, with Spearman’s ρ 0.73 (p = 0.0108). 3D printing of the glenoid based on pre-operative CT provides a physical representation of patient anatomy. Printed models enabled shoulder surgeons to appreciate glenoid bony morphology more clearly compared to review of CT images or humerus-subtracted volume renderings. These models were more useful as glenoid complexity increased.
Objective This paper describes the unified LOINC/RSNA Radiology Playbook and the process by which it was produced. Methods The Regenstrief Institute and the Radiological Society of North America (RSNA) developed a unification plan consisting of six objectives 1) develop a unified model for radiology procedure names that represents the attributes with an extensible set of values, 2) transform existing LOINC procedure codes into the unified model representation, 3) create a mapping between all the attribute values used in the unified model as coded in LOINC (ie, LOINC Parts) and their equivalent concepts in RadLex, 4) create a mapping between the existing procedure codes in the RadLex Core Playbook and the corresponding codes in LOINC, 5) develop a single integrated governance process for managing the unified terminology, and 6) publicly distribute the terminology artifacts. Results We developed a unified model and instantiated it in a new LOINC release artifact that contains the LOINC codes and display name (ie LONG_COMMON_NAME) for each procedure, mappings between LOINC and the RSNA Playbook at the procedure code level, and connections between procedure terms and their attribute values that are expressed as LOINC Parts and RadLex IDs. We transformed all the existing LOINC content into the new model and publicly distributed it in standard releases. The organizations have also developed a joint governance process for ongoing maintenance of the terminology. Conclusions The LOINC/RSNA Radiology Playbook provides a universal terminology standard for radiology orders and results.
Standard clinical terms, codes, and ontologies promote clarity and interoperability. Within radiology, there is a variety of relevant content resources, tools and technologies. These provide the basis for fundamental imaging workflows such as reporting and billing, and also facilitate a range of applications in quality improvement and research. This article reviews the key characteristics of lexicons, coding systems, and ontologies. A number of standards are described, including International Classification of Diseases-10-Clinical Modification (ICD-10-CM), Current Procedural Terminology (CPT), Systematized Nomenclature of Medicine—Clinical Terms (SNOMED CT), Logical Observation Identifiers Names and Codes (LOINC), and RadLex. Tools for accessing this material are reviewed, such as the National Center for Biomedical Ontology BioPortal system. Web services are discussed as a mechanism for semantic application development. Several example systems, workflows, and research applications using semantic technology are also surveyed.
Medical three-dimensional (3D) printing has expanded dramatically over the past three decades with growth in both facility adoption and the variety of medical applications. Consideration for each step required to create accurate 3D printed models from medical imaging data impacts patient care and management. In this paper, a writing group representing the Radiological Society of North America Special Interest Group on 3D Printing (SIG) provides recommendations that have been vetted and voted on by the SIG active membership. This body of work includes appropriate clinical use of anatomic models 3D printed for diagnostic use in the care of patients with specific medical conditions. The recommendations provide guidance for approaches and tools in medical 3D printing, from image acquisition, segmentation of the desired anatomy intended for 3D printing, creation of a 3D-printable model, and post-processing of 3D printed anatomic models for patient care.
We compare the continuous and discrete truncated Wigner approximations - TWA and dTWA, respectively - of various spin modelsu0027 dynamics to exact analytical and numerical solutions. We account for all components of spin-spin correlations on equal footing, facilitated by a recently introduced geometric correlation matrix visualization (CMV) technique. We find that the two approximations capture different aspects of the dynamics of correlations: dTWA captures periodic partial revivals of spin-spin correlations missed in the TWA; but at least for some conditions, TWA better captures the directions in which spins have correlations. We also find that at modestly short times, the dominant error in both approximations is to substantially suppress spin correlations along one direction. In addition to assessing the accuracy of approximations, our comparisons illustrate the utility of CMVs as an intuitive visual tool to understand the full complexity of spin correlations.