Social and behavioral determinants of health (SDOH) play a significant role in shaping health outcomes, and extracting these determinants from clinical notes is a first step to help healthcare providers systematically identify opportunities to provide appropriate care and address disparities. Progress on using NLP methods for this task has been hindered by the lack of high-quality publicly available labeled data, largely due to the privacy and regulatory constraints on the use of real patients' information. This paper introduces a new dataset, SDOH-NLI, that is based on publicly available notes and which we release publicly. We formulate SDOH extraction as a natural language inference (NLI) task, and provide binary textual entailment labels obtained from human raters for a cross product of a set of social history snippets as premises and SDOH factors as hypotheses. Our dataset differs from standard NLI benchmarks in that our premises and hypotheses are obtained independently. We evaluate both "off-the-shelf" entailment models as well as models fine-tuned on our data, and highlight the ways in which our dataset appears more challenging than commonly used NLI datasets.
Machine learning systems show significant promise for forecasting patient adverse events via risk scores. However, these risk scores implicitly encode assumptions about future interventions that the patient is likely to receive, based on the intervention policy present in the training data. Without this important context, predictions from such systems are less interpretable for clinicians. We propose a joint model of intervention policy and adverse event risk as a means to explicitly communicate the model's assumptions about future interventions. We develop such an intervention policy model on MIMIC-III, a real world de-identified ICU dataset, and discuss some use cases that highlight the utility of this approach. We show how combining typical risk scores, such as the likelihood of mortality, with future intervention probability scores leads to more interpretable clinical predictions.
Chronic kidney disease (CKD) represents a slowly progressive disorder that can eventually require renal replacement therapy (RRT) including dialysis or renal transplantation. Early identification of patients who will require RRT (as much as 1 year in advance) improves patient outcomes, for example by allowing higher-quality vascular access for dialysis. Therefore, early recognition of the need for RRT by care teams is key to successfully managing the disease. Unfortunately, there is currently no commonly used predictive tool for RRT initiation. In this work, we present a machine learning model that dynamically identifies CKD patients at risk of requiring RRT up to one year in advance using only claims data. To evaluate the model, we studied approximately 3 million Medicare beneficiaries for which we made over 8 million predictions. We showed that the model can identify at risk patients with over 90% sensitivity and specificity. Although additional work is required before this approach is ready for clinical use, this study provides a basis for a screening tool to identify patients at risk within a time window that enables early proactive interventions intended to improve RRT outcomes.
Harmonization of local source concepts to standard clinical terminologies is a prerequisite for multi-center data aggregation and sharing. Challenges in automating the mapping process stem from the idiosyncratic source encoding schemes adopted by different health systems and the lack of large publicly available training data. In this study, we aim to develop a scalable and generalizable machine learning tool to facilitate standardizing laboratory observations to the Logical Observation Identifiers Names and Codes (LOINC). Specifically, we leverage the contextual embedding from pre-trained T5 models and propose a two-stage fine-tuning strategy based on contrastive learning to enable learning in a few-shot setting without manual feature engineering. Our method utilizes unlabeled general LOINC ontology and data augmentation to achieve high accuracy on retrieving the most relevant LOINC targets when limited amount of labeled data are available. We further show that our model generalizes well to unseen targets. Taken together, our approach shows great potential to reduce manual effort in LOINC standardization and can be easily extended to mapping other terminologies.
Objectives Few machine learning (ML) models are successfully deployed in clinical practice. One of the common pitfalls across the field is inappropriate problem formulation: designing ML to fit the data rather than to address a real-world clinical pain point. Methods We introduce a practical toolkit for user-centred design consisting of four questions covering: (1) solvable pain points, (2) the unique value of ML (eg, automation and augmentation), (3) the actionability pathway and (4) the model’s reward function. This toolkit was implemented in a series of six participatory design workshops with care managers in an academic medical centre. Results Pain points amenable to ML solutions included outpatient risk stratification and risk factor identification. The endpoint definitions, triggering frequency and evaluation metrics of the proposed risk scoring model were directly influenced by care manager workflows and real-world constraints. Conclusions Integrating user-centred design early in the ML life cycle is key for configuring models in a clinically actionable way. This toolkit can guide problem selection and influence choices about the technical setup of the ML problem.
While it has been well known in the ML community that deep learning models suffer from instability, the consequences for healthcare deployments are under characterised. We study the stability of different model architectures trained on electronic health records, using a set of outpatient prediction tasks as a case study. We show that repeated training runs of the same deep learning model on the same training data can result in significantly different outcomes at a patient level even though global performance metrics remain stable. We propose two stability metrics for measuring the effect of randomness of model training, as well as mitigation strategies for improving model stability.
Physicians write clinical notes with abbreviations and shorthand that are difficult to decipher. Abbreviations can be clinical jargon (writing "HIT" for "heparin induced thrombocytopenia"), ambiguous terms that require expertise to disambiguate (using "MS" for "multiple sclerosis" or "mental status"), or domain-specific vernacular ("cb" for "complicated by"). Here we train machine learning models on public web data to decode such text by replacing abbreviations with their meanings. We report a single translation model that simultaneously detects and expands thousands of abbreviations in real clinical notes with accuracies ranging from 92.1%-97.1% on multiple external test datasets. The model equals or exceeds the performance of board-certified physicians (97.6% vs 88.7% total accuracy). Our results demonstrate a general method to contextually decipher abbreviations and shorthand that is built without any privacy-compromising data.
We develop a no-reference binocular image quality assessment model that operates on static stereoscopic images. The model deploys 2D and 3D features extracted from stereopairs to assess the perceptual quality they present when viewed stereoscopically. Both symmetric- and asymmetric-distorted stereopairs are handled by accounting for binocular rivalry using a classic linear rivalry model. The NSS features are used to train a support vector machine model to predict the quality of a tested stereopair. The model is tested on the LIVE 3D Image Quality Database, which includes both symmetric- and asymmetric-distorted stereoscopic 3D images. The experimental results show that our proposed model significantly outperforms the conventional 2D full-reference QA algorithms applied to stereopairs, as well as the 3D full-reference IQA algorithms on asymmetrically distorted stereopairs.
We develop a framework for assessing the quality of stereoscopic images that have been afflicted by possibly asymmetric distortions. An intermediate image is generated which when viewed stereoscopically is designed to have a perceived quality close to that of the cyclopean image. We hypothesize that performing stereoscopic QA on the intermediate image yields higher correlations with human subjective judgments. The experimental results confirm the hypothesis and show that the proposed framework significantly outperforms conventional 2D QA metrics when predicting the quality of stereoscopically viewed images that may have been asymmetrically distorted.
We describe two studies that were aimed towards increasing our understanding of how the visibility of distortions on stereoscopically viewed 3D images is affected by scene content and distortion types. By assuming that subjects' performance would be highly correlated with the visibility of local distorted patches, we analyzed subjects' performance in locating distortion patches when viewing stereoscopic 3D images. Subjects' performances are measured by whether they successfully locate a local distorted patch, the times they spent to finish the task, and subjective quality ratings given by subjects. The visual data used in this work are co-registered stereo images with co-registered "ground truth" range (depth) data. Varied statistical analysis methods were used to discuss the significance of our observations. Three observations are drawn from our analyses. First, blur, JPEG, and JP2K distortions in stereo 3D images may be suppressed if one of the left or right views is undistorted. Second, contrast masking does not occur, or is reduced, while viewing white noise distorted stereo 3D images. Third, there is no depth/disparity masking effect when viewing stereo 3D images, but there may be (conversely) depth-related facilitation effects for blur, JPEG, and JP2K distorted stereo 3D images.
We develop a framework for assessing the quality of stereoscopic images that have been afflicted by possibly asymmetric distortions. An intermediate image is generated which when viewed stereoscopically is designed to have a perceived quality close to that of the cyclopean image. We hypothesize that performing stereoscopic QA on the intermediate image yields higher correlations with human subjective judgments. The experimental results confirm the hypothesis and show that the proposed framework significantly outperforms conventional 2D QA metrics when predicting the quality of stereoscopically viewed images that may have been asymmetrically distorted.
We describe a study that aims towards enhancing our understanding of the perception of H.264/AVC compressed stereoscopic 3D videos, in particular spatial video quality, depth quality, visual comfort and overall 3D video quality. The results of this study indicate that the human subjects have diverse opinions on depth quality scores but a high agreement on spatial video quality. Their agreement on overall 3D video quality is intermediate relative to that on spatial video quality and depth quality. Based on our analysis, we propose to use separate quality assessment models: spatial video quality models and depth quality models.
We develop an algorithm that predicts the best presentation of a stereo 3D image in the sense of viewers' preference. The algorithm operates in three steps. First, the 3D image is classified as either a “foreground dominant” or “background dominant” image. Next, for “foreground dominant” images, a model of the stereoacuity function is used to optimize the perceptual 3D resolution; for “background dominant” images, the nearest surface is placed in the 3D plane of the display screen. A human study was conducted to assess the algorithm and showed that the proposed model produced 3D images which had the best 3D quality scores among several candidate algorithms.
The increasing number of demanding consumer video applications, as exemplified by cell phone and other low-cost digital cameras, has boosted interest in no-reference objective image and video quality assessment (QA) algorithms. In this paper, we focus on no-reference image and video blur assessment. We consider natural scenes statistics models combined with multi-resolution decomposition methods to extract reliable features for QA. The algorithm is composed of three steps. First, a probabilistic support vector machine (SVM) is applied as a rough image quality evaluator. Then the detail image is used to refine the blur measurements. Finally, the blur information is pooled to predict the blur quality of images. The algorithm is tested on the LIVE Image Quality Database and the Real Blur Image Database; the results show that the algorithm has high correlation with human judgments when assessing blur distortion of images.
The development of real-time image and video quality assessment algorithms is an important direction on which little research has focused. Towards this end, we present a design of real-time implementable full-reference image/video quality algorithms that are based on the Structural SIMilarity (SSIM) index and multi-scale SSIM (MS-SSIM) index. The proposed algorithms, which modify SSIM/MS-SSIM to achieve speed of execution, were tested on the LIVE Image Quality Database and LIVE Video Quality Database. The experimental results show that the performance of the new, fast algorithms is commensurate with that of SSIM and MS-SSIM, but with much lower computational complexity. Indeed, the proposed Fast MS-SSIM algorithm is 10 times faster (lower complexity) than the MS-SSIM algorithm, while the proposed Fast SSIM is 2.68 times faster than SSIM without parallel computing optimization.