The goal of this radiomic analysis is to quantify the sensitivity of radiomic features on computed tomography (CT) image pre-processing parameters and use machine learning (ML) techniques to identify the radiomic features that are highly predictive of shoulder arthroplasty outcomes. An ML framework auto-segmented 3D masks of the deltoid muscle and scapula bone from pre-operative CT images of 1949 primary anatomic total shoulder arthroplasty (aTSA)/reverse total shoulder arthroplasty (rTSA) patients. Radiomic features were extracted after various image pre-processing protocols and assessed for reproducibility. The radiomic features deemed robust to image pre-processing were used to train ML predictive outcomes models. Feature importance data were rank-ordered to identify the radiomic features that were highly predictive of pain, motion, and function before and after aTSA/rTSA. A sensitivity analysis identified 37 deltoid muscle and 38 scapular bone radiomic features that were robust, reproducible, and unique across image pre-processing parameters. The most predictive deltoid muscle radiomic measurements were normalized volume, elongation, flatness, fat percentage, sphericity, and max 2D diameter column. The most predictive scapular bone radiomic measurements were flatness, sphericity, elongation, max 2D diameter column, and max 2D diameter slice. Radiomic data of the deltoid and scapula were highly predictive of pain, motion, and function before and after aTSA and rTSA. Radiomic data were more predictive than patient comorbidities, diagnosis, and implant type/size data, but less predictive than pre-operative active range of motion measurements and patient reported outcome measures, 3D measurements from planning software, or patient demographic data. Future work is required to clinically validate these radiomic features before they can be deployed in clinical decision support tools.Level of Evidence Level III, Retrospective Comparative Outcome Study.
Introduction: We developed a computed tomography (CT)-based tool designed for automated segmentation of deltoid muscles, enabling quantification of radiomic features and muscle fatty infiltration. Prior to use in a clinical setting, this machine learning (ML)-based segmentation algorithm requires rigorous validation. The aim of this study is to conduct shoulder expert validation of a novel deltoid ML auto-segmentation and quantification tool. Materials and Methods: A SwinUnetR-based ML model trained on labeled CT scans is validated by three expert shoulder surgeons for 32 unique patients. The validation evaluates the quality of the auto-segmented deltoid images. Specifically, each of the three surgeons reviewed the auto-segmented masks relative to CT images, rated masks for clinical acceptance, and performed a correction on the ML-generated deltoid mask if the ML mask did not completely contain the full deltoid muscle, or if the ML mask included any tissue other than the deltoid. Non-inferiority of the ML model was assessed by comparing ML-generated to surgeon-corrected deltoid masks versus the inter-surgeon variation in metrics, such as volume and fatty infiltration. Results: The results of our expert shoulder surgeon validation demonstrates that 97% of ML-generated deltoid masks were clinically acceptable. Only two of the ML-generated deltoid masks required major corrections and only one was deemed clinically unacceptable. These corrections had little impact on the deltoid measurements, as the median error in the volume and fatty infiltration measurements was <1% between the ML-generated deltoid masks and the surgeon-corrected deltoid masks. The non-inferiority analysis demonstrates no significant difference between the ML-generated to surgeon-corrected masks relative to inter-surgeon variations. Conclusions: Shoulder expert validation of this CT image analysis tool demonstrates clinically acceptable performance for deltoid auto-segmentation, with no significant differences observed between deltoid image-based measurements derived from the ML generated masks and those corrected by surgeons. These findings suggest that this CT image analysis tool has potential to reliably quantify deltoid muscle size, shape, and quality. Incorporating these CT image-based measurements into the pre-operative planning process may facilitate more personalized treatment decision making, and help orthopedic surgeons make more evidence-based clinical decisions.
Background: The goal of this study is to analyze a registry of preoperative computed tomography (CT) images of anatomic total shoulder arthroplasty (aTSA) and reverse total shoulder arthroplasty (rTSA) patients, quantify the radiomics of the deltoid muscle and scapular bone, and identify the radiomic features that are most predictive of pain, motion, and function before and after aTSA/rTSA. Methods: Preoperative CT images and clinical data from 4,009 primary shoulder arthroplasty patients were retrospectively analyzed. Next, three-dimensional masks of the deltoid (n = 2,597) and scapula (n = 3,358) were auto-segmented from CT images, and radiomic features were extracted using Py-Radiomics. These radiomics features were then used to train machine-learning regression models to predict pain, motion, and function before and after aTSA/rTSA. Finally, a clustering analysis was performed using the most predictive radiomic features to identify unique deltoid and scapula morphological groups/classes relevant to clinical outcomes before and after aTSA/rTSA. Results: Incorporating radiomic features into the machine-learning models improved the accuracy of 70.5% of deltoid model outcome predictions and 67.3% of scapular model outcome predictions. Analysis of feature importance data demonstrated that the most predictive radiomic features were numerical representations of deltoid and scapula shape and size. Notably, most shape-based radiomic features were more predictive of aTSA/rTSA outcomes than any patient demographic data (except age), comorbidity data, implant data, or diagnosis data. Finally, a radiomic-based clustering analysis identified several deltoid muscle and scapula bone morphologies associated with differences in clinical outcomes before and after aTSA/rTSA. Conclusion: This analysis of >4,000 preoperative CT scans identified numerous radiomic features of the deltoid and scapula that were highly predictive of pain, motion, and function before and after aTSA/rTSA. These most predictive radiomic features were aggregated into unique morphological clusters of deltoids and scapula that were associated with differences in clinical outcomes before and after aTSA/rTSA. Shape-based radiomic features were more predictive than first-order and second-order radiomic features, suggesting that these more interpretable measurements are more clinically relevant and could be more readily incorporated into future radiomic-based clinical decision support tools. Future work is required to further validate these radiomic findings and refine the proposed clustering analysis.
Background/Objectives: Artificial intelligence (AI) is increasingly integrated into everyday life, including the complex and highly regulated healthcare sector. Given healthcare’s essential role in safeguarding human life and well-being, AI deployment requires careful oversight to ensure safety, effectiveness, and ethical compliance. This paper aims to examine the current regulatory landscapes governing AI in healthcare, particularly in the European Union (EU) and the United States (USA), and to propose practical tools to support the responsible development and implementation of AI systems. Methods: The study reviews key regulatory frameworks, ethical guidelines, and expert recommendations from international bodies, professional associations, and governmental institutions in the EU and USA. Based on this analysis, the paper develops structured questionnaires tailored for AI developers and implementers to help operationalize regulatory and ethical expectations. Results: The proposed questionnaires address critical gaps in existing frameworks by providing actionable, lifecycle-oriented tools that span AI development, deployment, and clinical use. These instruments support compliance and ethical integrity while promoting transparency and accountability. Conclusions: The structured questionnaires can serve as practical tools for health technology assessments, public procurement, accreditation processes, and training initiatives. By aligning AI system design with regulatory and ethical standards, they contribute to building trustworthy, safe, and innovative AI applications in healthcare.
Heart disease is one of the leading causes of death worldwide, highlighting the urgent need for early detection and effective preventative treatment. Recent years have seen the development of prediction models to support clinical judgment in the medical field, making machine learning a powerful tool. Decision Tree, Naïve Bayes, K-Nearest Neighbors (KNN), and Support Vector Machine (SVM) are four standard machine learning techniques for heart disease prediction that are examined in this study. The models are trained and assessed on a publicly available dataset that comprises several health-related variables, including blood pressure, cholesterol, age, kind of chest pain, and blood sugar levels. Key performance metrics such as F1-score, recall, accuracy, and precision are utilized to assess performance. In the evaluation of prediction accuracy, the Decision Tree classifier performed better than the other algorithms, with Naïve Bayes ranking second. As per the results, there is a great deal of promise for practical applications in medical diagnostics using these models, especially Decision Trees. This study demonstrates how machine learning techniques can be reliable, data-driven tools to assist healthcare professionals in identifying individuals who are at risk of heart disease. In the end, this will enhance patient care and optimize medical resources.
Background: Machine learning (ML)-based clinical decision support tools (CDSTs) make personalized predictions for different treatments; by comparing predictions of multiple treatments, these tools can be used to optimize decision making for a particular patient. However, CDST prediction accuracy varies for different patients and also for different treatment options. If these differences are sufficiently large and consistent for a particular subcohort of patients, then that bias may result in those patients not receiving a particular treatment. Such level of bias would deem the CDST "unfair." The purpose of this study is to evaluate the "fairness" of ML CDST-based clinical outcomes predictions after anatomic (aTSA) and reverse total shoulder arthroplasty (rTSA) for patients of different demographic attributes. Methods: Clinical data from 8280 shoulder arthroplasty patients with 19,249 postoperative visits was used to evaluate the prediction fairness and accuracy associated with the following patient demographic attributes: ethnicity, sex, and age at the time of surgery. Performance of clinical outcome and range of motion regression predictions were quantified by the mean absolute error (MAE) and performance of minimal clinically important difference (MCID) and substantial clinical benefit classification predictions were quantified by accuracy, sensitivity, and the F1 score. Fairness of classification predictions leveraged the "four-fifths" legal guideline from the US Equal Employment Opportunity Commission and fairness of regression predictions leveraged established MCID thresholds associated with each outcome measure. Results: For both aTSA and rTSA clinical outcome predictions, only minor differences in MAE were observed between patients of different ethnicity, sex, and age. Evaluation of prediction fairness demonstrated that 0 of 486 MCID (0%) and only 3 of 486 substantial clinical benefit (0.6%) classification predictions were outside the 20% fairness boundary and only 14 of 972 (1.4%) regression predictions were outside of the MCID fairness boundary. Hispanic and Black patients were more likely to have ML predictions out of fairness tolerance for aTSA and rTSA. Additionally, patients < 60 years old were more likely to have ML predictions out of fairness tolerance for rTSA. No disparate predictions were identified for sex and no disparate regression predictions were observed for forward elevation, internal rotation score, American Shoulder and Elbow Surgeons Standardized Shoulder Assessment Form score, or global shoulder function. Conclusion: The ML algorithms analyzed in this study accurately predict clinical outcomes after aTSA and rTSA for patients of different ethnicity, sex, and age, where only 1.4% of regression predictions and only 0.3% of classification predictions were out of fairness tolerance using the proposed fairness evaluation method and acceptance criteria. Future work is required to externally validate these ML algorithms to ensure they are equally accurate for all legally protected patient groups. Level of evidence: Basic Science Study; Validation of Computer Modeling (c) 2023 Journal of Shoulder and Elbow Surgery Board of Trustees. All rights reserved.
Background: Despite the importance of the deltoid to shoulder biomechanics, very few studies have quantified the three-dimensional shape, size, or quality of the deltoid muscle, and no studies have correlated these measurements to clinical outcomes after anatomic (aTSA) and/or reverse (rTSA) total shoulder arthroplasty in any statistically/scientifically relevant manner. Methods: Preoperative computer tomography (CT) images from 1057 patients (585 female, 469 male; 799 primary rTSA and 258 primary aTSA) of a single platform shoulder arthroplasty prosthesis (Equinoxe; Exactech, Inc., Gainesville, FL) were analyzed in this study. A machine learning (ML) framework was used to segment the deltoid muscle for 1057 patients and quantify 15 different muscle characteristics, including volumetric (size, shape, etc.) and intensity-based Hounsfield (HU) measurements. These deltoid measurements were correlated to postoperative clinical outcomes and utilized as inputs to train/test ML algorithms used to predict postoperative outcomes at multiple postoperative timepoints (1 year, 2–3 years, and 3–5 years) for aTSA and rTSA. Results: Numerous deltoid muscle measurements were demonstrated to significantly vary with age, gender, prosthesis type, and CT image kernel; notably, normalized deltoid volume and deltoid fatty infiltration were demonstrated to be relevant to preoperative and postoperative clinical outcomes after aTSA and rTSA. Incorporating deltoid image data into the ML models improved clinical outcome prediction accuracy relative to ML algorithms without image data, particularly for the prediction of abduction and forward elevation after aTSA and rTSA. Analyzing ML feature importance facilitated rank-ordering of the deltoid image measurements relevant to aTSA and rTSA clinical outcomes. Specifically, we identified that deltoid shape flatness, normalized deltoid volume, deltoid voxel skewness, and deltoid shape sphericity were the most predictive image-based features used to predict clinical outcomes after aTSA and rTSA. Many of these deltoid measurements were found to be more predictive of aTSA and rTSA postoperative outcomes than patient demographic data, comorbidity data, and diagnosis data. Conclusions: While future work is required to further refine the ML models, which include additional shoulder muscles, like the rotator cuff, our results show promise that the developed ML framework can be used to evolve traditional CT-based preoperative planning software into an evidence-based ML clinical decision support tool.
Background: Improvement in internal rotation (IR) after anatomic (aTSA) and reverse (rTSA) total shoulder arthroplasty is difficult to predict, with rTSA patients experiencing greater variability and more limited IR improvements than aTSA patients. The purpose of this study is to quantify and compare the IR score for aTSA and rTSA patients and create supervised machine learning that predicts IR after aTSA and rTSA at multiple postoperative time points. Methods: Clinical data from 2270 aTSA and 4198 rTSA patients were analyzed using 3 supervised machine learning techniques to create predictive models for internal rotation as measured by the IR score at 6 postoperative time points. Predictions were performed using the full input feature set and 2 minimal input feature sets. The mean absolute error (MAE) quantified the difference between actual and predicted IR scores for each model at each time point. The predictive accuracy of the XGBoost algorithm was also quantified by its ability to distinguish which patients would achieve clinical improvement greater than the minimal clinically important difference (MCID) and substantial clinical benefit (SCB) patient satisfaction thresholds for IR score at 2-3 years after surgery. Results: rTSA patients had significantly lower mean IR scores and significantly less mean IR score improvement than aTSA patients at each postoperative time point. Both aTSA and rTSA patients experienced significant improvements in their ability to perform activities of daily living (ADLs); however, aTSA patients were significantly more likely to perform these ADLs. Using a minimal feature set of preoperative inputs, our machine learning algorithms had equivalent accuracy when predicting IR score for both aTSA (0.92-1.18 MAE) and rTSA (1.03-1.25 MAE) from 3 months to >5 years after surgery. Furthermore, these predictive algorithms identified with 90% accuracy for aTSA and 85% accuracy for rTSA which patients will achieve MCID IR score improvement and predicted with 85% accuracy for aTSA patients and 77% accuracy for rTSA which patients will achieve SCB IR score improvement at 2-3 years after surgery. Discussion: Our machine learning study demonstrates that active internal rotation can be accurately predicted after aTSA and rTSA at multiple postoperative time points using a minimal feature set of preoperative inputs. These predictive algorithms accurately identified which patients will, and will not, achieve clinical improvement in IR score that exceeds the MCID and SCB patient satisfaction thresholds. (C) 2021 Journal of Shoulder and Elbow Surgery Board of Trustees. All rights reserved.
We use machine learning to create predictive models from preoperative data to predict the Shoulder Arthroplasty Smart (SAS) score, the American Shoulder and Elbow Surgeons (ASES) score, and the Constant score at multiple postoperative time points and compare the accuracy of each algorithm for anatomic total shoulder arthroplasty (aTSA) and reverse total shoulder arthroplasty (rTSA). Clinical data from 2270 patients who underwent aTSA and 4198 patients who underwent rTSA were analyzed using 3 supervised machine learning techniques to create predictive models for the SAS, ASES, and Constant scores at 6 different postoperative time points using a full input feature set and the 2 different minimal feature sets. Mean absolute errors (MAEs) quantified the difference between actual and predicted outcome scores for each model at each postoperative time point. The performance of each model was also quantified by its ability to predict improvement greater than the minimal clinically important difference (MCID) and the substantial clinical benefit (SCB) patient satisfaction thresholds for each outcome measure at 2-3 years after surgery. All 3 machine learning techniques were more accurate at predicting aTSA and rTSA outcomes using the SAS score (aTSA: ±7.41 MAE; rTSA: ±7.79 MAE), followed by the Constant score (aTSA: ±8.32 MAE; rTSA: ±8.30 MAE) and finally the ASES score (aTSA: ±10.86 MAE; rTSA: ±10.60 MAE). These prediction accuracy trends were maintained across the 3 different model input categories for each of the SAS, ASES, and Constant models at each postoperative time point. For patients who underwent aTSA, the XGBoost predictive models achieved 94%-97% accuracy in MCID with an area under the receiver operating curve (AUROC) between 0.90-0.97 and 89%-94% accuracy in SCB with an AUROC between 0.89-0.92 for the 3 clinical scores using the full feature set of inputs. For patients who underwent rTSA, the XGBoost predictive models achieved 95%-99% accuracy in MCID with an AUROC between 0.88-0.96 and 88%-92% accuracy in SCB with an AUROC between 0.81-0.89 for the 3 clinical scores using the full feature set of inputs. Our study demonstrated that the SAS score predictions are more accurate than the ASES and Constant predictions for multiple supervised machine learning techniques, despite requiring fewer input data for the SAS model. In addition, we predicted which patients will and will not achieve clinical improvement that exceeds the MCID and SCB thresholds for each score; this highly accurate predictive capability effectively risk-stratifies patients for a variety of outcome measures using only preoperative data. Level III; Retrospective Comparative Study
Pressure Injuries are localized damages to the skin caused by sustained pressure. It is a common yet preventable disease affecting millions of patients. While there are multiple scales to determine if a patient has pressure injury, these methods suffer from high inter-rater subjectivity. To address this problem we create predictive models for pressure injury using Centers for Medicare Medicaid Services claims data. The models show relatively good predictive performance, we also explore aspects of the model where they will be deployed in a real world clinical settings.
The issue of bias and fairness in healthcare has been around for centuries. With the integration of AI in healthcare the potential to discriminate and perpetuate unfair and biased practices in healthcare increases many folds. The tutorial focuses on the challenges, requirements and opportunities in the area of fairness in healthcare AI and the various nuances associated with it. The problem healthcare as a multi-faceted systems level problem that necessitates careful consideration of different notions of fairness in healthcare to corresponding concepts in machine learning is elucidated via different real world examples.
Background: A machine learning analysis was conducted on 5774 shoulder arthroplasty patients to create predictive models for multiple clinical outcome measures after anatomic total shoulder arthroplasty (aTSA) and reverse total shoulder arthroplasty (rTSA). The goal of this study was to compare the accuracy associated with a full feature set predictive model (ie, full model, comprising 291 parameters) and a minimal feature set model (ie, abbreviated model, comprising 19 input parameters) to predict clinical outcomes to assess the efficacy of using a minimal feature set of inputs as a shoulder arthroplasty clinical decision-support tool. Methods: Clinical data from 2153 primary aTSA patients and 3621 primary rTSA patients were analyzed using the XGBoost machine learning technique to create and test predictive models for multiple outcome measures at different postoperative time points via the full and abbreviated models. Mean absolute errors (MAEs) quantified the difference between actual and predicted outcomes, and each model also predicted whether a patient would experience clinical improvement greater than the patient satisfaction anchor-based thresholds of the minimal clinically important difference and substantial clinical benefit for each outcome measure at 2-3 years after surgery. Results: Across all postoperative time points analyzed, the full and abbreviated models had similar MAEs for the American Shoulder and Elbow Surgeons score (+/- 11.7 with full model vs. +/- 12.0 with abbreviated model), Constant score (+8.9 vs. +/- 9.8), Global Shoulder Function score (+/- 1.4 vs. +/- 1.5), visual analog scale pain score (+/- 1.3 vs. +/- 1.4), active abduction (+/- 20.4 degrees - vs. +/- 21.8 degrees), forward elevation (+/- 17.6 degrees vs. +/- 19.2 degrees), and external rotation (+/- 12.2 degrees vs. +/- 12.6 degrees). Marginal improvements in MAEs were observed for each outcome measure prediction when the abbreviated model was supplemented with data on implant size and/or type and measurements of native glenoid anatomy. The full and abbreviated models each effectively risk stratified patients using only preoperative data by accurately identifying patients with improvement greater than the minimal clinically important difference and substantial clinical benefit thresholds. Discussion: Our study showed that the full and abbreviated machine learning models achieved similar accuracy in predicting clinical outcomes after aTSA and rTSA at multiple postoperative time points. These promising results demonstrate an efficient utilization of machine learning algorithms to predict clinical outcomes. Our findings using a minimal feature set of only 19 preoperative inputs suggest that this tool may be easily used during a surgical consultation to improve decision making related to shoulder arthroplasty. (C) 2020 Journal of Shoulder and Elbow Surgery Board of Trustees. All rights reserved.
Datasets from Electronic Health Records (EHRs) are increasingly large and complex, creating challenges in their use for predictive modeling. The two major challenges are large-scale and high-dimensionality. One of the common way to address the large-scale challenge is through use of data phenotypes: clinically relevant characteristic groupings that can be expressed as logical queries (e.g., “senior patients with diabetes”). With the increasing use of machine learning across the continuum of care, phenotypes play an important role in modeling for population management, clinical trials, observational and interventional research, and quality measures. Yet, phenotype interpretation can often be difficult and require post-hoc clarifications from experienced clinicians. For example, detailed analysis may be needed to find that all patients in a a phenotype are diabetic seniors with complications from previous surgery. Moreover, the high-dimensionality problem is often addressed either separately or simultaneously with phenotyping by dimension reduction methods that may further hinder interpretability. In this paper, we introduce the notion of interpretable data phenotypes generated by an unsupervised learning technique. Methods are designed to disambiguate relative feature memberships, thus facilitating general clinical validation, and alleviating the problem of high-dimensionality. The empirical study applies the proposed unsupervised interpretable phenotyping method to a real world healthcare dataset (MIMIC), then uses hospital length of stay as a reference prediction task. The results demonstrate that the proposed method produces phenotypes with improved interpretability and without diminishing the quality of prediction results.
BACKGROUND:We propose a new clinical assessment tool constructed using machine learning, called the Shoulder Arthroplasty Smart (SAS) score to quantify outcomes following total shoulder arthroplasty (TSA).METHODS:Clinical data from 3667 TSA patients with 8104 postoperative follow-up reports were used to quantify the psychometric properties of validity, responsiveness, and clinical interpretability for the proposed SAS score and each of the Simple Shoulder Test (SST), Constant, American Shoulder and Elbow Surgeons Standardized Shoulder Assessment Form (ASES), University of California Los Angeles (UCLA), and Shoulder Pain and Disability Index (SPADI) scores.RESULTS:Convergent construct validity was demonstrated, with all 6 outcome measures being moderately to highly correlated preoperatively and highly correlated postoperatively when quantifying TSA outcomes. The SAS score was most correlated with the UCLA score and least correlated with the SST. No clinical outcome score exhibited significant floor effects preoperatively or postoperatively or significant ceiling effects preoperatively; however, significant ceiling effects occurred postoperatively for each of the SST (44.3%), UCLA (13.9%), ASES (18.7%), and SPADI (19.3%) measures. Ceiling effects were more pronounced for anatomic than reverse TSA, and generally, men, younger patients, and whites who received TSA were more likely to experience a ceiling effect than TSA patients who were female, older, and of non-white race or ethnicity. The SAS score had the least number of patients with floor and ceiling effects and also exhibited no response bias in any patient characteristic analyzed in this study. Regarding clinical interpretability, patient satisfaction anchor-based thresholds for minimal clinically importance difference and substantial clinical benefit were quantified for all 6 outcome measures; the SAS score thresholds were most similar in magnitude to the Constant score. Regarding responsiveness, all 6 outcome measures detected a large effect, with the UCLA exhibiting the most responsiveness and the SST exhibiting the least. Finally, each of the SAS, ASES, Constant, and SPADI scores had similarly large standardized response mean and effect size responsiveness.DISCUSSION:The 6-question SAS score is an efficient TSA-specific outcome measure with equivalent or better validity, responsiveness, and clinical interpretability as 5 other historical assessment tools. The SAS score has an appropriate response range without floor or ceiling effects and without bias in any target patient characteristic, unlike the age, gender, or race/ethnicity bias observed in the ceiling scores with the other outcome measures. Because of these substantial benefits, we recommend the use of the new SAS score for quantifying TSA outcomes.
Prediction of diabetes and its various complications has been studied in a number of settings, but a comprehensive overview of problem setting for diabetes prediction and care management has not been addressed in the literature. In this document we seek to remedy this omission in literature with an encompassing overview of diabetes complication prediction as well as situating this problem in the context of real world healthcare management. We illustrate various problems encountered in real world clinical scenarios via our own experience with building and deploying such models. In this manuscript we illustrate a Machine Learning (ML) framework for addressing the problem of predicting Type 2 Diabetes Mellitus (T2DM) together with a solution for risk stratification, intervention and management. These ML models align with how physicians think about disease management and mitigation, which comprises these four steps: Identify, Stratify, Engage, Measure.
With the increased adoption of AI in healthcare, there is a growing recognition and demand to regulate AI in healthcare to avoid potential harm and unfair bias against vulnerable populations. Around a hundred governmental bodies and commissions as well as leaders in the tech sector have proposed principles to create responsible AI systems. However, most of these proposals are short on specifics which has led to charges of ethics washing. In this tutorial we offer a guide to help navigate through complex governmental regulations and explain the various constituent practical elements of a responsible AI system in healthcare in the light of proposed regulations. Additionally, we breakdown and emphasize that the recommendations from regulatory bodies like FDA or the EU are necessary but not sufficient elements of creating a responsible AI system. We elucidate how regulations and guidelines often focus on epistemic concerns to the detriment of practical concerns e.g., requirement for fairness without explicating what fairness constitutes for a use case. FDA's Software as a medical device document and EU's GDPR among other AI governance documents talk about the need for implementing sufficiently good machine learning practices. In this tutorial we elucidate what that would mean from a practical perspective for real world use cases in healthcare throughout the machine learning cycle i.e., Data Management, Data Specification, Feature Engineering, Model Evaluation, Model Specification, Model Explainability, Model Fairness, Reproducibility, checks for data leakage and model leakage. We note that conceptualizing responsible AI as a process rather than an end goal accords well with how AI systems are used in practice. We also discuss how a domain centric stakeholder perspective translates into balancing requirements for multiple competing optimization criteria.
An important psychometric parameter of validity that is rarely assessed is predictive value. In this study we utilize machine learning to analyze the predictive value of 3 commonly used clinical measures to assess 2-year outcomes after total shoulder arthroplasty (TSA). XGBoost was used to analyze data from 2790 TSA patients and create predictive algorithms for the American Shoulder and Elbow Surgeons (ASES), Constant, and the University of California Los Angeles (UCLA) scores and also quantify the most meaningful predictive features utilized by these measures and for all questions comprising each measure to rank and compare their value to predict 2-year outcomes after TSA. Our results demonstrate that the ASES, Constant, and UCLA measures rarely considered the most-predictive features relevant to 2-year TSA outcomes and that each outcome measure was composed of questions with different distributions of predictive value. Specifically, the questions composing the UCLA score were of greater predictive value than the Constant questions, and the questions composing the Constant score were of greater predictive value than the ASES questions. We also found the preoperative Shoulder Pain and Disability Index (SPADI) score to be of greater predictive value than the preoperative ASES, Constant, and UCLA scores. Finally, we identified the types of preoperative input questions that were most-predictive (subjective self-assessments of pain and objective measurements of active range of motion and strength) and also those that were least-predictive of 2-year TSA outcomes (subjective task-specific activities of daily living questions). Machine learning can quantify the predictive value of the ASES, Constant, and UCLA scores after TSA. Future work should utilize this and related techniques to construct a more efficient and effective clinical outcome measure that incorporates subjective and objective input questions to better account for the preoperative factors that influence postoperative outcomes after TSA. Level III; Retrospective Comparative Study
Fairness in AI and machine learning systems has become a fundamental problem in the accountability of AI systems. While the need for accountability of AI models is near ubiquitous, healthcare in particular is a challenging field where accountability of such systems takes upon additional importance, as decisions in healthcare can have life altering consequences. In this paper we present preliminary results on fairness in the context of classification parity in healthcare. We also present some exploratory methods to improve fairness and choosing appropriate classification algorithms in the context of healthcare.
The issue of bias and fairness in healthcare has been around for centuries. With the integration of AI in healthcare the potential to discriminate and perpetuate unfair and biased practices in healthcare increases many folds The tutorial focuses on the challenges, requirements and opportunities in the area of fairness in healthcare AI and the various nuances associated with it. The problem healthcare as a multi-faceted systems level problem that necessitates careful of different notions of fairness in healthcare to corresponding concepts in machine learning is elucidated via different real world examples.
Over the past several years, across the globe, there has been an increase in people seeking care in emergency departments (EDs). ED resources, including nurse staffing, are strained by such increases in patient volume. Accurate forecasting of incoming patient volume in emergency departments (ED) is crucial for efficient utilization and allocation of ED resources. Working with a suburban ED in the Pacific Northwest, we developed a tool powered by machine learning models, to forecast ED arrivals and ED patient volume to assist end-users, such as ED nurses, in resource allocation. In this paper, we discuss the results from our predictive models, the challenges, and the learnings from users' experiences with the tool in active clinical deployment in a real world setting.
Joseph A. Konstan合作论文数Department of Computer Science and Engineering, College of Science and Engineering, University of Minnesota4
Lyndon Kennedy合作论文数Yahoo! Research3
John Riedl合作论文数Department of Computer Science and Engineering, College of Science and Engineering, University of Minnesota1