PURPOSE:To evaluate the safety and efficacy of ultrasound-guided percutaneous thrombin injection for the treatment of upper extremity pseudoaneurysms. METHODS:An institutional database containing 8,316,467 radiology reports was searched for suitable cases over a 241-month period. Fourteen female and 10 male patients, average age of 69.7 years (range 29-93) underwent a total of 26 procedures for the management of upper extremity pseudoaneurysms, involving the radial (n = 9), brachial (n = 9) or other upper extremity arteries (n = 6). Baseline demographic and pseudoaneurysm characteristics were documented, together with primary and secondary success, failures, and complications. All procedures were performed with real-time ultrasound guidance. RESULTS:The mean pseudoaneurysm volume was 9.93 cm3 (range 0.06-111.62 cm3). Twelve cases were related to central line placement or arterial access. Primary success was obtained in 50% (n = 12) after a single ultrasound-guided thrombin injection, and secondary success was achieved in an additional six (for a total success of 75%). Success was highest for the treatment of brachial artery pseudoaneurysms (87.5%), and in those who were diagnosed within 7 days of the inciting event, findings that were statistically significant (p-value 0.046 and 0.002, respectively). CONCLUSIONS:Ultrasound-guided percutaneous thrombin injection is safe and effective for managing upper extremity pseudoaneurysms.
Rib fractures commonly result from traumatic injury and often require hospitalization for pain control and supportive pulmonary care. Although the use of mobile health technology to share patient-generated health data has increased, it remains limited in patients with traumatic injuries. We sought to assess the feasibility of mobile health tracking in patients with rib fractures by using a smartphone app to monitor postdischarge recovery. We encountered patient, institutional, and process-related obstacles that limited app use. The success of future work requires the acknowledgment of these limitations and the use of an implementation science framework to effectively integrate technological tools for personalized trauma care.
Since 2016 the Healthcare Information and Management Systems Society (HIMSS) and the Society for Imaging Informatics in Medicine (SIIM) have collaborated to generate a series of white papers summarizing important topics of interest to both communities-the HIMSS-SIIM Enterprise Imaging Community (HSEIC) White Papers.As of December 2023, these papers reached a significant milestone with over 100,000 accesses.To celebrate this accomplishment and the renaming of SIIM's Journal of Digital Imaging (JDI) to the Journal of Imaging Informatics in Medicine (JIIM), we invited the authors of these white papers (Table 1) to provide an update on what the impact has been and what the future may still hold for these important topics.
Routine concordance evaluation between pathology and imaging findings was introduced for CT-guided biopsies. To analyze malignancy rate in concordant, discordant, and indeterminate non-malignant results of CT-guided lung biopsies. Concordance between pathology results and imaging findings of consecutive patients undergoing CT-guided lung biopsy between 7/1/2016 and 9/30/2021 was assessed during routine meetings by procedural radiologists. Concordant was defined as pathology consistent with imaging findings; discordant was used when pathology could not explain imaging findings; indeterminate when pathology could explain imaging findings but there was concern for malignancy. Recommendations for discordant and indeterminate were provided. All the malignant results were concordant. Pathology of repeated biopsy, surgical sample, or follow-up was considered reference standard. Consecutive 828 CT-guided lung biopsies were performed on 795 patients (median age 70 years, IQR 61–77), 423/828 (51 • A routine radiology-pathology concordance evaluation of CT-guided lung biopsies classified 224 non-malignant results as concordant, discordant, or indeterminate. • The percentage of malignancy on follow-up was significantly different in concordant (2 • Time to definitive diagnosis was significantly shorter with repeat biopsy (33 days), compared to imaging follow-up (114 days), p = 0.01.
Self-supervised representation learning on image-text data facilitates crucial medical applications, such as image classification, visual grounding, and cross-modal retrieval. One common approach involves contrasting semantically similar (positive) and dissimilar (negative) pairs of data points. Drawing negative samples uniformly from the training data set introduces false negatives, i.e., samples that are treated as dissimilar but belong to the same class. In healthcare data, the underlying class distribution is nonuniform, implying that false negatives occur at a highly variable rate. To improve the quality of learned representations, we develop a novel approach that corrects for false negatives. Our method can be viewed as a variant of debiased contrastive learning that uses estimated sample-specific class probabilities. We provide theoretical analysis of the objective function and demonstrate the proposed approach on both image and paired image-text data sets. Our experiments illustrate empirical advantages of sample-specific debiasing.
Vision-language pretraining has been shown to produce high-quality visual encoders which transfer efficiently to downstream computer vision tasks. While generative language models have gained widespread attention, image captioning has thus far been mostly overlooked as a form of cross-modal pretraining in favor of contrastive learning, especially in medical image analysis. In this paper, we experiment with bidirectional captioning of radiology reports as a form of pretraining and compare the quality and utility of learned embeddings with those from contrastive pretraining methods. We optimize a CNN encoder, transformer decoder architecture named RadTex for the radiology domain. Results show that not only does captioning pretraining yield visual encoders that are competitive with contrastive pretraining (CheXpert competition multi-label AUC of 89.4%), but also that our transformer decoder is capable of generating clinically relevant reports (captioning macro-F1 score of 0.349 using CheXpert labeler) and responding to prompts with targeted, interactive outputs.
To report outcomes of post-surgical benign biliary strictures treated with percutaneous transhepatic biliary drainage (PTBD) and balloon dilation. Patients with a history of surgery and who developed a benign biliary stricture were included in this retrospective, IRB-approved study (n = 44). Patients having biliary leaks, or with complete occlusion of the bile duct due to surgical clips were excluded (n = 28). Drain size, type (internal-external or external), balloon dilation, and type and size of balloon used was evaluated for all biliary procedures. These characteristics were compared between patients needing (n = 8) and not needing (n = 36) surgical revision. Continuous variables are reported as median (IQR) and compared with Wilcoxon rank sum test, categorical variables as numbers with percentages and compared using chi-square test. Patients were 59 (42-67) years old with 67% being female. Hepaticojejunostomy was the most common surgery (30%) followed by Roux-en-Y gastric bypass (16%), pancreaticoduodenectomy (14%), transplant (14%), Billroth II (9%), cholecystectomy (7%) and others (9%). Duct to enteric anastomosis was present in 52%, duct to duct in 11% and no anastomosis in 36%. Prior non-PTBD interventions included ERCP dilation and stent placement in 9 (20%) and failed ERCP in 10 (23%) patients. The first PTBD was done 28 (6-98) months after surgery, with a drain size of ≥10F in 20 (45%) and < 10F in 24 (55%) patients. The drain was internal-external in 41 (93%) patients. Patients underwent a total of 4 (3-8) procedures before drain removal, with drain upsizing performed in 23 (52%) patients, leading to a largest drain size of ≤8F in 10 (23%), 10F in 17 (39%), 12F in 14 (32%) and ≥14F in 3 (6%) patients. Balloon dilation was performed in 40 (91%) patients, with the largest balloon size being ≤8 mm in 13 (30%), 10 mm in 20 (45%), 12 mm in 6 (14%), and 14 mm in 1 (2%) patient. A cutting balloon was used in 19 (43%) patients. Surgical revision was needed in 8 (18%) patients, with a repeat PTBD performed in 5 (11%) patients. The rate of surgical revision was lower in patients who received a largest drain of ≥10F (12%) compared with those who received a largest drain of <10F (40%, P = 0.042). The rate of surgical revision was 7.4% in patients who were dilated with a balloon of ≥10 mm compared with 30.7% in those who were dilated with a balloon of <10 mm (P = 0.053). The median drain-free survival was 41 (11-57) months for the entire cohort. Postsurgical benign biliary strictures treated with PTBD with a drain of ≥10F and dilated with a balloon of ≥10 mm are less likely to need surgical revision.
Purpose To assess the accuracy, completeness, and readability of patient educational material produced by a machine learning model and compare the output to that provided by a societal website. Materials and Methods Content from the Society of Interventional Radiology Patient Center website was retrieved, categorized, and organized into discrete questions. These questions were entered into the ChatGPT platform, and the output was analyzed for word and sentence counts, readability using multiple validated scales, factual correctness, and suitability for patient education using the Patient Education Materials Assessment Tool for Printable Materials (PEMAT-P) instrument. Results A total of 21,154 words were analyzed, including 7,917 words from the website and 13,377 words representing the total output of the ChatGPT platform across 22 text passages. Compared to the societal website, output from the ChatGPT platform was longer and more difficult to read on 4 of 5 readability scales. The ChatGPT output was incorrect for 12 (11.5%) of 104 questions. When reviewed using the PEMAT-P tool, the ChatGPT content scored lower than the website material. Content from both the website and ChatGPT were significantly above the recommended fifth or sixth grade level for patient education, with a mean Flesch-Kincaid grade level of 11.1 (±1.3) for the website and 11.9 (±1.6) for the ChatGPT content. Conclusions The ChatGPT platform may produce incomplete or inaccurate patient educational content, and providers should be familiar with the limitations of the system in its current form. Opportunities may exist to fine-tune existing large language models, which could be optimized for the delivery of patient educational content.
Image-text multimodal representation learning aligns data across modalities and enables important medical applications, e.g., image classification, visual grounding, and cross-modal retrieval. In this work, we establish a connection between multimodal representation learning and multiple instance learning. Based on this connection, we propose a generic framework for constructing permutation-invariant score functions with many existing multimodal representation learning approaches as special cases. Furthermore, we use the framework to derive a novel contrastive learning approach and demonstrate that our method achieves state-of-the-art results in several downstream tasks.
Accurately assessing pulmonary edema severity is critical for making treatment decisions in congestive heart failure patients. However, the current scale for quantifying pulmonary edema based on chest radiographs does not have well-characterized severity levels, with substantial inter-radiologist disagreement. In this study, we investigate whether comparisons documented in radiology reports can accurately characterize pulmonary edema progression. We propose a rules-based natural language processing approach to assess the change in a patient's pulmonary edema status (better, worse, no change) by performing pairwise comparisons of consecutive radiology reports, using regular expressions and heuristics derived from clinical knowledge. Evaluated against ground-truth labels from radiology experts, our labeler extracts comparisons describing the progression of pulmonary edema with 0.875 precision and 0.891 recall. We also demonstrate the potential utility of comparison labels in providing additional fine-grained information over noisier labels produced by models that directly estimate severity level.
Importance:Following up on recommendations from radiologic findings is important for patient care, but frequently there are failures to carry out these recommendations. The lack of reliable systems to characterize and track completion of actionable radiology report recommendations poses an important patient safety challenge.Objectives:To characterize actionable radiology recommendations and, using this taxonomy, track and understand rates of loop closure for radiology recommendations in a primary care setting.Design, Setting, and Participants:Radiology reports in a primary care clinic at a large academic center were redesigned to include actionable recommendations in a separate dedicated field. Manual review of all reports generated from imaging tests ordered between January 1 and December 31, 2018, by primary care physicians that contained actionable recommendations was performed. For this quality improvement study, a taxonomy system that conceptualized recommendations was developed based on 3 domains: (1) what is recommended (eg, repeat a test or perform a different test, specialty referral), (2) specified time frame in which to perform the recommended action, and (3) contingency language qualifying the recommendation. Using this framework, a 2-stage process was used to review patients' records to classify recommendations and determine loop closure rates and factors associated with failure to complete recommended actions. Data analysis was conducted from April to July 2021.Main Outcomes and Measures:Radiology recommendations, time frames, and contingencies. Rates of carrying out vs not closing the loop on these recommendations in the recommended time frame were assessed.Results:A total of 598 radiology reports were identified with structured recommendations: 462 for additional or future radiologic studies and 196 for nonradiologic actions (119 specialty referrals, 47 invasive procedures, and 43 other actions). The overall rate of completed actions (loop closure) within the recommended time frame was 87.4%, with 31 open loop cases rated by quality expert reviewers to pose substantial clinical risks. Factors associated with successful loop closure included (1) absence of accompanying contingency language, (2) shorter recommended time frames, and (3) evidence of direct radiologist communication with the ordering primary care physicians. A clinically significant lack of loop closure was found in approximately 5% of cases.Conclusions and Relevance:The findings of this study suggest that creating structured radiology reports featuring a dedicated recommendations field permits the development of taxonomy to classify such recommendations and determine whether they were carried out. The lack of loop closure suggests the need for more reliable systems.
Despite technological advances in the analysis of digital images for medical consultations, many health information systems lack the ability to correlate textual descriptions of image findings linked to the actual images. Images and reports often reside in separate silos in the medical record throughout the process of image viewing, report authoring, and report consumption. Forward-thinking centers and early adopters have created interactive reports with multimedia elements and embedded hyperlinks in reports that connect the narrative text with the related source images and measurements. Most of these solutions rely on proprietary single-vendor systems for viewing and reporting in the absence of any encompassing industry standards to facilitate interoperability with the electronic health record (EHR) and other systems. International standards have enabled the digitization of image acquisition, storage, viewing, and structured reporting. These provide the foundation to discuss enhanced reporting. Lessons learned in the digital transformation of radiology and pathology can serve as a basis for interactive multimedia reporting (IMR) across image-centric medical specialties. This paper describes the standard-based infrastructure and communications to fulfill recently defined clinical requirements through a consensus from an international workgroup of multidisciplinary medical specialists, informaticists, and industry participants. These efforts have led toward the development of an Integrating the Healthcare Enterprise (IHE) profile that will serve as a foundation for interoperable interactive multimedia reporting.
Rationale and objectives: In response to COVID-19, our institution implemented three virtual readout systems: a commercial HIPAA compliant web-based video conferencing platform used for screen-sharing (Starleaf), an interactive control sharing system integrated into PACS allowing simultaneous multi-user mouse control over images (Collaborate), and the telephone. Our aim was to assess overall satisfaction with and perceived effectiveness of these virtual readout methods to optimize best practices for the future. Materials and methods: An IRB-exempt survey was electronically distributed to 64 trainees and 76 attendings at one tertiary-care institution via Survey Monkey. Questions focused on overall satisfaction, perceived effectiveness, technical difficulties, and continued future use of the three virtual readout strategies. Answers were collected with Likert scales, tick boxes, and open-ended questions. Results: 32/64 trainees (50%) and 32/76 attendings (42%) completed the survey. Trainees and attendings were more satisfied with screen sharing (Starleaf) and perceived it more effective than control sharing (Collaborate) or the telephone (p < 0.0001). Respondents experienced more technical difficulties with control sharing versus screen sharing (p = 0.0004) with a negative correlation between level of technical difficulties and satisfaction with screen sharing (r =-0.50, p < 0.0001) and control sharing (r =-0.38, p = 0.0006). Trainees and faculty supported a combination of in-person and virtual readouts in the future (p < 0.0001). Conclusion: Platforms mirroring in-person readouts, such as Starleaf, are preferred by both trainees and attendings over non-screen sharing platforms such as the telephone. However, technical stability determines satisfaction between similar platforms. Both trainees and attendings support incorporation of virtual readout methods in combination with traditional in-person readouts in the post-COVID-19 era.
Automated analysis of chest radiography using deep learning has tremendous potential to enhance the clinical diagnosis of diseases in patients. However, deep learning models typically require large amounts of annotated data to achieve high performance – often an obstacle to medical domain adaptation. In this paper, we build a data-efficient learning framework that utilizes radiology reports to improve medical image classification performance with limited labeled data (fewer than 1000 examples). Specifically, we examine image-captioning pretraining to learn high-quality medical image representations that train on fewer examples. Following joint pretraining of a convolutional encoder and transformer decoder, we transfer the learned encoder to various classification tasks. Averaged over 9 pathologies, we find that our model achieves higher classification performance than ImageNet-supervised and in-domain supervised pretraining when labeled training data is limited.
We propose and demonstrate a representation learning approach by maximizing the mutual information between local features of images and text. The goal of this approach is to learn useful image representations by taking advantage of the rich information contained in the free text that describes the findings in the image. Our method trains image and text encoders by encouraging the resulting representations to exhibit high local mutual information. We make use of recent advances in mutual information estimation with neural network discriminators. We argue that the sum of local mutual information is typically a lower bound on the global mutual information. Our experimental results in the downstream image classification tasks demonstrate the advantages of using local features for image-text representation learning.
Adoption of machine learning models in healthcare requires end users’ trust in the system. Models that provide additional supportive evidence for their predictions promise to facilitate adoption. We define consistent evidence to be both compatible and sufficient with respect to model predictions. We propose measures of model inconsistency and regularizers that promote more consistent evidence. We demonstrate our ideas in the context of edema severity grading from chest radiographs. We demonstrate empirically that consistent models provide competitive performance while supporting interpretation.
Artificial intelligence (AI) models for decision support have been developed for clinical settings such as radiology, but little work evaluates the potential impact of such systems. In this study, physicians received chest X-rays and diagnostic advice, some of which was inaccurate, and were asked to evaluate advice quality and make diagnoses. All advice was generated by human experts, but some was labeled as coming from an AI system. As a group, radiologists rated advice as lower quality when it appeared to come from an AI system; physicians with less task-expertise did not. Diagnostic accuracy was significantly worse when participants received inaccurate advice, regardless of the purported source. This work raises important considerations for how advice, AI and non-AI, should be deployed in clinical environments.
William M. Wells III合作论文数Surgical Planning Laboratory, Department of Radiology, Brigham and Women's Hospital, Harvard Medical School;Division of Health Sciences and Technology, Massachusetts Institute of Technology3