The growing adoption of diagnostic and prognostic algorithms in healthcare has led to concerns about the perpetuation of algorithmic bias against disadvantaged groups of individuals. Deep learning methods to detect and mitigate bias have revolved around modifying models, optimization strategies, and threshold calibration with varying levels of success and tradeoffs. However, there have been limited substantive efforts to address bias at the level of the data used to generate algorithms in healthcare datasets. We create a simple metric (AEquity) that utilizes a learning curve approximation to distinguish and mitigate bias via guided dataset collection or relabeling. We demonstrate this metric in two well-known examples: chest X-rays and healthcare cost utilization, and detect novel biases in the National Health and Nutrition Examination Survey. We demonstrate that utilizing AEquity to guide data-centric collection for each diagnostic finding in the chest radiograph dataset decreased bias by between 29% and 96.5% when measured by differences in area-under-the-curve. When we examined Black patients on Medicaid, at the intersection of race and socioeconomic status, we found that AEquity-based interventions reduced bias across a number of different fairness metrics including overall false negative rate by 33.3% (Bias Reduction Absolute = 1.88 x 10-1; 95% CI (1.4x10-1, 2.5x10-1); Bias Reduction (%) 33.3% (95% CI, 26.6-40.0)), Precision Bias by 7.50x10-2; 95% CI (7.48x10-2, 7.51x10-2); Bias Reduction (%) 94.6% (95% CI, 94.5-94.7%); False Discovery Rate by 94.5% (Absolute Bias Reduction = 3.50x10-2; 95% CI: (3.49x10-2, 3.50x10-2). Similarly, AEquity-guided data collection demonstrates bias reduction of up to 80% on mortality prediction with the National Health and Nutrition Examination Survey (Bias Reduction Absolute = 0.08; 95% CI (0.07, 0.09)). Additionally, we benchmark against balanced empirical risk minimization and calibration and we show that AEquity-guided data collection outperforms both standard approaches. Moreover, we demonstrate that AEquity works on fully connected networks, convolutional neural networks such as ResNet-50, transformer architectures such as on VIT-B-16, an 86 million parameter Vision Transformer, and nonparametric methods such as LightGBM In short, we demonstrate AEquity is a robust tool by applying it to different datasets and algorithms, intersectional analyses and measuring its effectiveness with respect to a range of traditional fairness metrics.
BACKGROUND:Acute kidney injury (AKI) is common in SARS-CoV-2 infection and COVID-19, often leading to long-term kidney dysfunction. However, the transcriptomic features of AKI severity and its long-term effects are underexplored. METHODS:We performed bulk RNA sequencing on peripheral blood mononuclear cells (PBMCs) from hospitalized SARS-CoV-2 patients and complemented these findings with proteomic data from the same cohort. We compared the functional enrichment findings with historical sepsis-AKI data and subsequently examined the association between molecular signatures and long-term kidney function changes. RESULTS:In 283 patients, 57 had mild AKI (stage 1) and 49 had severe AKI (stage 2 or 3). Following adjustments for age, sex, severity of infection, and pre-existing chronic kidney disease (CKD), we identified 6,432 differentially expressed genes (DEGs) in the severe AKI vs. control comparison, 840 in the mild AKI vs. control, and 1,213 in the severe vs. mild AKI comparison (FDR<0.05). Common pathways included unfolded protein response, cellular response to stress via eIF2, and IFN-g-mediated inflammatory response. Severe AKI was linked to pathways involved in mitochondrial dysfunction and endoplasmic reticulum stress. Proteomic analysis confirmed 40 established AKI and inflammation biomarkers, while gene-set enrichment of transcription regulators revealed additional biomarkers for severe AKI. Comparison with PBMC transcriptomics from sepsis-related AKI showed significant functional overlap (30%). Analysis of post-discharge eGFR data in 115 patients identified 177 DEGs for severe vs. control, 106 for mild vs. control, and 46 for severe vs. mild AKI. Key associations included kidney function decline related to carbohydrate and mitochondrial metabolism, inflammatory-response, and cardiovascular regulation. CONCLUSIONS:We demonstrate that severe AKI in SARS-CoV-2 infection is linked to mitochondrial dysfunction and ER stress. The functional overlap with sepsis-AKI suggests potential broader therapeutic applicability. Long-term kidney dysfunction is influenced by disruptions in cellular energy metabolism and immune response.
Importance Increased intracranial pressure (ICP) is associated with adverse neurological outcomes, but needs invasive monitoring. Objective Development and validation of an AI approach for detecting increased ICP (aICP) using only non-invasive extracranial physiological waveform data. Design Retrospective diagnostic study of AI-assisted detection of increased ICP. We developed an AI model using exclusively extracranial waveforms, externally validated it and assessed associations with clinical outcomes. Setting MIMIC-III Waveform Database (2000-2013), a database derived from patients admitted to an ICU in an academic Boston hospital, was used for development of the aICP model, and to report association with neurologic outcomes. Data from Mount Sinai Hospital (2020-2022) in New York City was used for external validation. Participants Patients were included if they were older than 18 years, and were monitored with electrocardiograms, arterial blood pressure, respiratory impedance plethysmography and pulse oximetry. Patients who additionally had intracranial pressure monitoring were used for development (N=157) and external validation (N=56). Patients without intracranial monitors were used for association with outcomes (N=1694). Exposures Extracranial waveforms including electrocardiogram, arterial blood pressure, plethysmography and SpO 2 . Main Outcomes and Measures Intracranial pressure > 15 mmHg. Measures were Area under receiver operating characteristic curves (AUROCs), sensitivity, specificity, and accuracy at threshold of 0.5. We calculated odds ratios and p-values for phenotype association. Results The AUROC was 0.91 (95% CI, 0.90-0.91) on testing and 0.80 (95% CI, 0.80-0.80) on external validation. aICP had accuracy, sensitivity, and specificity of 73.8% (95% CI, 72.0%-75.6%), 99.5% (95% CI 99.3%-99.6%), and 76.9% (95% CI, 74.0-79.8%) on external validation. A ten-percentile increment was associated with stroke (OR=2.12; 95% CI, 1.27-3.13), brain malignancy (OR=1.68; 95% CI, 1.09-2.60), subdural hemorrhage (OR=1.66; 95% CI, 1.07-2.57), intracerebral hemorrhage (OR=1.18; 95% CI, 1.07-1.32), and procedures like percutaneous brain biopsy (OR=1.58; 95% CI, 1.15-2.18) and craniotomy (OR = 1.43; 95% CI, 1.12-1.84; P < 0.05 for all). Conclusions and Relevance aICP provides accurate, non-invasive estimation of increased ICP, and is associated with neurological outcomes and neurosurgical procedures in patients without intracranial monitoring.
Increased intracranial pressure (ICP) >= 15 mmHg is associated with adverse neurological outcomes, but needs invasive intracranial monitoring. Using the publicly available MIMIC-III Waveform Database (2000-2013) from Boston, we developed an artificial intelligence-derived biomarker for elevated ICP (aICP) for adult patients. aICP uses routinely collected extracranial waveform data as input, reducing the need for invasive monitoring. We externally validated aICP with an independent dataset from the Mount Sinai Hospital (2020-2022) in New York City. The AUROC, accuracy, sensitivity, and specificity on the external validation dataset were 0.80 (95% CI, 0.80-0.80), 73.8% (95% CI, 72.0-75.6%), 73.5% (95% CI 72.5-74.5%), and 73.0% (95% CI, 72.0-74.0%), respectively. We also present an exploratory analysis showing aICP predictions are associated with clinical phenotypes. A ten-percentile increment was associated with brain malignancy (OR = 1.68; 95% CI, 1.09-2.60), intracerebral hemorrhage (OR = 1.18; 95% CI, 1.07-1.32), and craniotomy (OR = 1.43; 95% CI, 1.12-1.84; P < 0.05 for all).
Univariate high-frequency time series are dominant data sources for many medical, economic and environmental applications. In many of these domains, the time series are tied to real-time changes in state. In the intensive care unit, for example, changes and intracranial pressure waveforms can indicate whether a patient is developing decreased blood perfusion to the brain during a stroke, for example. However, most representation learning to resolve states is conducted in an offline, batch-dependent manner. In high frequency time-series, high intra-state and inter-sample variability makes offline, batch-dependent learning a relatively difficult task. Hence, we propose Spatial Resolved Temporal Networks (SpaRTeN), a novel composite deep learning model for online, unsupervised representation learning through a spatially constrained latent space. SpaRTeN maps waveforms to states, and learns time-dependent representations of each state. Our key contribution is that we generate clinically relevant representations of each state for intracranial pressure waveforms.
Background Acute kidney injury (AKI) is a known complication of COVID-19 and is associated with an increased risk of in-hospital mortality. Unbiased proteomics using biological specimens can lead to improved risk stratification and discover pathophysiological mechanisms. Methods Using measurements of ~4000 plasma proteins in two cohorts of patients hospitalized with COVID-19, we discovered and validated markers of COVID-associated AKI (stage 2 or 3) and long-term kidney dysfunction. In the discovery cohort (N= 437), we identified 413 higher plasma abundances of protein targets and 40 lower plasma abundances of protein targets associated with COVID-AKI (adjusted p <0.05). Of these, 62 proteins were validated in an external cohort (p <0.05, N =261). Results We demonstrate that COVID-AKI is associated with increased markers of tubular injury ( NGAL ) and myocardial injury. Using estimated glomerular filtration (eGFR) measurements taken after discharge, we also find that 25 of the 62 AKI-associated proteins are significantly associated with decreased post-discharge eGFR (adjusted p <0.05). Proteins most strongly associated with decreased post-discharge eGFR included desmocollin-2 , trefoil factor 3 , transmembrane emp24 domain-containing protein 10 , and cystatin-C indicating tubular dysfunction and injury. Conclusions Using clinical and proteomic data, our results suggest that while both acute and long-term COVID-associated kidney dysfunction are associated with markers of tubular dysfunction, AKI is driven by a largely multifactorial process involving hemodynamic instability and myocardial damage.
The human skeletal form underlies our ability to walk on two legs, but unlike standing height, the genetic basis of limb lengths and skeletal proportions is less well understood. Here we applied a deep learning model to 31,221 whole body dual-energy X-ray absorptiometry (DXA) images from the UK Biobank (UKB) to extract 23 different image-derived phenotypes (IDPs) that include all long bone lengths as well as hip and shoulder width, which we analyzed while controlling for height. All skeletal proportions are highly heritable (∼40-50%), and genome-wide association studies (GWAS) of these traits identified 179 independent loci, of which 102 loci were not associated with height. These loci are enriched in genes regulating skeletal development as well as associated with rare human skeletal diseases and abnormal mouse skeletal phenotypes. Genetic correlation and genomic structural equation modeling indicated that limb proportions exhibited strong genetic sharing but were genetically independent of width and torso proportions. Phenotypic and polygenic risk score analyses identified specific associations between osteoarthritis (OA) of the hip and knee, the leading causes of adult disability in the United States, and skeletal proportions of the corresponding regions. We also found genomic evidence of evolutionary change in arm-to-leg and hip-width proportions in humans consistent with striking anatomical changes in these skeletal proportions in the hominin fossil record. In contrast to cardiovascular, auto-immune, metabolic, and other categories of traits, loci associated with these skeletal proportions are significantly enriched in human accelerated regions (HARs), and regulatory elements of genes differentially expressed through development between humans and the great apes. Taken together, our work validates the use of deep learning models on DXA images to identify novel and specific genetic variants affecting the human skeletal form and ties a major evolutionary facet of human anatomical change to pathogenesis.
AbstractBackgroundBroad adoption of artificial intelligence (AI) algorithms in healthcare has led to perpetuation of bias found in datasets used for algorithm training. Methods to mitigate bias involve approaches after training leading to tradeoffs between sensitivity and specificity. There have been limited efforts to address bias at the level of the data for algorithm generation.MethodsWe generate a data-centric, but algorithm-agnostic approach to evaluate dataset bias by investigating how the relationships between different groups are learned at different sample sizes. We name this method AEquity and define a metric AEq. We then apply a systematic analysis of AEq values across subpopulations to identify and mitigate manifestations of racial bias.FindingsWe demonstrate that AEquity helps mitigate different biases in three different chest radiograph datasets, a healthcare costs dataset, and when using tabularized electronic health record data for mortality prediction. In the healthcare costs dataset, we show that AEquity is a more sensitive metric of label bias than model performance. AEquity can be utilized for label selection when standard fairness metrics fail. In the chest radiographs dataset, we show that AEquity can help optimize dataset selection to mitigate bias, as measured by nine different fairness metrics across nine of the most frequent diagnoses and four different protected categories (race, sex, insurance status, age) and the intersections of race and sex. We benchmark against approaches currently used after algorithm training including recalibration and balanced empirical risk minimization. Finally, we utilize AEquity to characterize and mitigate a previously unreported bias in mortality prediction with the widely used National Health and Nutrition Examination Survey (NHANES) dataset, showing that AEquity outperforms currently used approaches, and is effective at both small and large sample sizes.InterpretationAEquity can identify and mitigate bias in known biased datasets through different strategies and an unreported bias in a widely used dataset.SummaryAEquity, a machine learning approach can identify and mitigate bias the level of datasets used to train algorithms. We demonstrate it can mitigate known cases of bias better than existing methods, and detect and mitigate bias that was previously unreported.EVIDENCE IN CONTEXTEvidence before this studyMethods to mitigate algorithmic bias typically involve adjustments made after training, leading to a tradeoff between sensitivity and specificity. There have been limited efforts to mitigate bias at the level of the data.Added value of this studyThis study introduces a machine learning based method, AEquity, which analyzes the learnability of data from subpopulations at different sample sizes, which can then be used to intervene on the larger dataset to mitigate bias. The study demonstrates the detection and mitigation of bias in two scenarios where bias had been previously reported. It also demonstrates the detection and mitigation of bias the widely used National Health and Nutrition Examination Survey (NHANES) dataset, which was previously unknown.Implications of all available evidenceAEquity is a complementary approach that can be used early in the algorithm lifecycle to characterize and mitigate bias and thus prevent perpetuation of algorithmic disparities.
BACKGROUND:Modern machine learning and deep learning algorithms require large amounts of data; however, data sharing between multiple healthcare institutions is limited by privacy and security concerns. SUMMARY:Federated learning provides a functional alternative to the single-institution approach while avoiding the pitfalls of data sharing. In cross-silo federated learning, the data do not leave a site. The raw data are stored at the site of collection. Models are created at the site of collection and are updated locally to achieve a learning objective. We demonstrate a use case with COVID-19-associated AKI. We showed that federated models outperformed their local counterparts, even when evaluated on local data in the test dataset, and performance was like those being used for pooled data. Increases in performance at a given hospital were inversely proportional to dataset size at a given hospital, which suggests that hospitals with smaller datasets have significant room for growth with federated learning approaches. KEY MESSAGES:This short article provides an overview of federated learning, gives a use case for COVID-19-associated acute kidney injury, and finally details the issues along with some potential solutions.
Background AKI is a heterogeneous syndrome. Current subphenotyping approaches have only used limited laboratory data to understand a much more complex condition. Methods We focused on patients with AKI from the Assessment, Serial Evaluation, and Subsequent Sequelae in AKI (ASSESS-AKI). We used hierarchical clustering with Ward linkage on biomarkers of inflammation, injury, and repair/health. We then evaluated clinical differences between subphenotypes and examined their associations with cardiorenal events and death using Cox proportional hazard models. Results We included 748 patients with AKI: 543 (73%) of them had AKI stage 1, 112 (15%) had AKI stage 2, and 93 (12%) had AKI stage 3. The mean age (+/- SD) was 64 (13) years; 508 (68%) were men; and the median follow- up was 4.7 (Q1: 2.9, Q3: 5.7) years. Patients with AKI subphenotype 1 (N=181) had the highest kidney injury molecule (KIM-1) and troponin T levels. Subphenotype 2 (N=250) had the highest levels of uromodulin. AKI subphenotype 3 (N=159) comprised patients with markedly high pro-brain natriuretic peptide and plasma tumor necrosis factor receptor-1 and -2 and low concentrations of KIM-1 and neutrophil gelatinase- associated lipocalin. Finally, patients with subphenotype 4 (N=158) predominantly had sepsisAKI and the highest levels of vascular/kidney inflammation (YKL-40, MCP-1) and injury (neutrophil gelatinase- associated lipocalin, KIM-1). AKI subphenotypes 3 and 4 were independently associated with a higher risk of death compared with subphenotype 2 and had adjusted hazard ratios of 2.9 (95% confidence interval, 1.8 to 4.6) and 1.6 ( 95% confidence interval, 1.01 to 2.6, P = 0.04), respectively. Subphenotype 3 was also independently associated with a three-fold risk of CKD and cardiovascular events. Conclusions We discovered four AKI subphenotypes with differing clinical features and biomarker profiles that are associated with longitudinal clinical outcomes.
The human skeletal form underlies bipedalism, but the genetic basis of skeletal proportions (SPs) is not well characterized. We applied deep-learning models to 31,221 x-rays from the UK Biobank to extract a comprehensive set of SPs, which were associated with 145 independent loci genome-wide. Structural equation modeling suggested that limb proportions exhibited strong genetic sharing but were independent of width and torso proportions. Polygenic score analysis identified specific associations between osteoarthritis and hip and knee SPs. In contrast to other traits, SP loci were enriched in human accelerated regions and in regulatory elements of genes that are differentially expressed between humans and great apes. Combined, our work identifies specific genetic variants that affect the skeletal form and ties a major evolutionary facet of human anatomical change to pathogenesis.
PURPOSE: Facial aging is a multifactorial process involving both soft tissues and bony structures1. Factors including volume loss, gravity, muscle laxity, and cellular damage contribute to decreased skin and soft tissue elasticity, resulting in increased mobility of facial soft tissues1-2. Variations in soft tissue integrity have been attributed to age, sun exposure, and smoking. While facial aging has been extensively examined histologically, the present study sought to leverage clinical MRI to quantify facial soft tissue movement (STM) and correlate with environmental and demographic factors. MATERIALS & METHODS: Sixty-eight patients underwent high resolution MRI scans, which included two identical scans at the beginning and end of imaging separated by approximately 45 minutes. MRIcron was used to label 49 reproducible bony and soft tissue facial landmarks on all scans. An avatar scan was used to co-register and scale all scans. For each patient, early and late scans were coregistered, and mathematical voxel-wise absolute differences were used to render composite maps to highlight the change in soft tissue configuration over the 45 minute gap. Lines between close neighbor landmarks offered corresponding paths by which movement could be compared between patients. Linear regression was used to correlate average absolute differences with age, sex, smoking status, and sun exposure. RESULTS: Age was positively correlated with the greatest STM compared to sun exposure, sex, and smoking status. In the upper face, age was correlated with STM in the forehead (glabella to superior orbits, p=0.026), bony orbits (p-range=0.001-0.023), and orbital soft tissue (orbits to medial/lateral canthi, p-range=0.001-0.014). Age was associated with bilateral midface STM in the infraorbital region (between malar eminence, bilateral canthi and inferior orbit, p-range=0.001-0.019) and zygomatic region (malar eminence to auditory canal, p-range=0.001-0.032). In the lower face, age correlated with STM around the mouth and philtrum (lips, oral commissures, columella and nares, p-range=0.028-0.001), and between the mandible and mentum (p-range=0.001-0.031). Sun exposure was only associated with STM in the oral region (lips, columella, and nares, p-range=0.001-0.049) and infraorbital/nasal region (nares to medial canthus, p=0.001). Sex was only associated with STM around the mandibular angle (p=0.004). Smoking was not found to be associated with significant bilateral STM. CONCLUSIONS: To our knowledge this is the first study to examine facial STM using clinical in vivo MRI. This methodology identified facial regions most susceptible to changes from aging and various environmental factors. These results provide further understanding of the natural facial aging process and may be helpful in identifying rejuvenative targets in the future. References: 1. Ilankovan V. Anatomy of ageing face. Br J Oral Maxillofac Surg. 2014 Mar;52(3):195-202. doi: 10.1016/j.bjoms.2013.11.013. Epub 2013 Dec 23. PMID: 24370442. 2. Freytag L, Alfertshofer MG, Frank K, et al. Understanding Facial Aging Through Facial Biomechanics: A Clinically Applicable Guide for Improved Outcomes. Facial Plast Surg Clin North Am. 2022 May;30(2):125-133. doi: 10.1016/j.fsc.2022.01.001. PMID: 35501049.
The adoption of diagnosis and prognostic algorithms in healthcare has led to concerns about the perpetuation of bias against disadvantaged groups of individuals. Deep learning methods to detect and mitigate bias have revolved around modifying models, optimization strategies, and threshold calibration with varying levels of success. Here, we generate a data-centric, model-agnostic, task-agnostic approach to evaluate dataset bias by investigating the relationship between how easily different groups are learned at small sample sizes (AEquity). We then apply a systematic analysis of AEq values across subpopulations to identify and mitigate manifestations of racial bias in two known cases in healthcare - Chest X-rays diagnosis with deep convolutional neural networks and healthcare utilization prediction with multivariate logistic regression. AEq is a novel and broadly applicable metric that can be applied to advance equity by diagnosing and remediating bias in healthcare datasets.
Sample size estimation is a crucial step in experimental design but is understudied in the context of deep learning. Currently, estimating the quantity of labeled data needed to train a classifier to a desired performance, is largely based on prior experience with similar models and problems or on untested heuristics. In many supervised machine learning applications, data labeling can be expensive and time-consuming and would benefit from a more rigorous means of estimating labeling requirements. Here, we study the problem of estimating the minimum sample size of labeled training data necessary for training computer vision models as an exemplar for other deep learning problems. We consider the problem of identifying the minimal number of labeled data points to achieve a generalizable representation of the data, a minimum converging sample (MCS). We use autoencoder loss to estimate the MCS for fully connected neural network classifiers. At sample sizes smaller than the MCS estimate, fully connected networks fail to distinguish classes, and at sample sizes above the MCS estimate, generalizability strongly correlates with the loss function of the autoencoder. We provide an easily accessible, code-free, and dataset-agnostic tool to estimate sample sizes for fully connected networks. Taken together, our findings suggest that MCS and convergence estimation are promising methods to guide sample size estimates for data collection and labeling prior to training deep learning models in computer vision.
Journal of the American Society of Nephrology 33(11S):p 409-410, November 2022. | DOI: 10.1681/ASN.20223311S1409d
The fundamental challenge in machine learning is ensuring that trained models generalize well to unseen data. We developed a general technique for ameliorating the effect of dataset shift using generative adversarial networks (GANs) on a dataset of 149,298 handwritten digits and dataset of 868,549 chest radiographs obtained from four academic medical centers. Efficacy was assessed by comparing area under the curve (AUC) pre- and post-adaptation. On the digit recognition task, the baseline CNN achieved an average internal test AUC of 99.87% (95% CI, 99.87-99.87%), which decreased to an average external test AUC of 91.85% (95% CI, 91.82-91.88%), with an average salvage of 35% from baseline upon adaptation. On the lung pathology classification task, the baseline CNN achieved an average internal test AUC of 78.07% (95% CI, 77.97-78.17%) and an average external test AUC of 71.43% (95% CI, 71.32-71.60%), with a salvage of 25% from baseline upon adaptation. Adversarial domain adaptation leads to improved model performance on radiographic data derived from multiple out-of-sample healthcare populations. This work can be applied to other medical imaging domains to help shape the deployment toolkit of machine learning in medicine.
Clinical EHR data is naturally heterogeneous, where it contains abundant sub-phenotype. Such diversity creates challenges for outcome prediction using a machine learning model since it leads to high intra-class variance. To address this issue, we propose a supervised pre-training model with a unique embedded k-nearest-neighbor positive sampling strategy. We demonstrate the enhanced performance value of this framework theoretically and show that it yields highly competitive experimental results in predicting patient mortality in real-world COVID-19 EHR data with a total of over 7,000 patients admitted to a large, urban health system. Our method achieves a better AUROC prediction score of 0.872, which outperforms the alternative pre-training models and traditional machine learning methods. Additionally, our method performs much better when the training data size is small (345 training instances).
Rationale & Objective: The association between cannabis use and chronic kidney disease (CKD) is controversial. We aimed to assess association of CKD with cannabis use in a large cohort study and then assess causality using Mendelian randomization with a genome-wide association study (GWAS).Study Design: Retrospective cohort study and genome-wide association study.Setting & Participants: The retrospective study was conducted on the All of Us cohort (N=223,35 4). Genetic instruments for cannabis use disorder were identified from 3 GWAS: the Psychiatric Genomics Consortium Substance Use Disorders, iPSYCH, and deCODE (N=38 4,032). Association between genetic instruments and CKD was investigated in the CKDGen GWAS (N > 1.2 million).Exposure: Cannabis consumption.Outcomes: CKD outcomes included: cyst atin-C and creatinine-based kidney function, proteinuria, and blood urea nitrogen.Analytical Approach: We conducted association analyses to test for frequency of cannabis use and CKD. To evaluate causality, we performed a 2 -sample Mendelian randomization.Results: In the retrospective study, compared to former users, less than monthly (OR, 1.01; 95% CI, 0.87-1.18; P = 0.87) and monthly cannabis users (OR, 1.15; 95% CI, 0.8 6-1.52; P = 0.33) did not have higher CKD odds. Conversely, weekly (OR, 1.28; 95% CI, 1.01-1.6 0; P = 0.04) and daily use (OR, 1.25; 95% CI, 1.04-1.5 0; P = 0.02) was significantly associated with CKD, adjusted for multiple confounders. In Mendelian randomization, genetic liability to cannabis use disorder was not associated with increased odds for CKD (OR, 1.00; 95% CI, 0.9 9-1.01; P = 0.96). These results were robust across different Mendelian randomi-zation techniques and multiple kidney traits.Limitations: Likely underreporting of cannabis use. In Mendelian randomization, genetic instruments were identified in the GWAS that included in-dividuals primarily of European ancestry.Conclusions: Despite the epidemiological associ-ation between cannabis use and CKD, there was no evidence of a causal effect, indicating con-founding in observational studies.
Purpose of review Risk stratification for chronic kidney is becoming increasingly important as a clinical tool for both treatment and prevention measures. The goal of this review is to identify how machine learning tools contribute and facilitate risk stratification in the clinical setting. Recent findings The two key machine learning paradigms to predictively stratify kidney disease risk are genomics-based and electronic health record based approaches. These methods can provide both quantitative information such as relative risk and qualitative information such as characterizing risk by subphenotype. Summary The four key methods to stratify chronic kidney disease risk are genomics, multiomics, supervised and unsupervised machine learning methods. Polygenic risk scores utilize whole genome sequencing data to generate an individual's relative risk compared with the population. Multiomic methods integrate information from multiple biomarkers to generate trajectories and prognostic different outcomes. Supervised machine learning methods can directly utilize the growing compendia of electronic health records such as laboratory results and notes to generate direct risk predictions, while unsupervised machine learning methods can cluster individuals with chronic kidney disease into subphenotypes with differing approaches to care.
Food intake behavior is regulated by a network of appetite-inducing and appetite-suppressing neuronal populations throughout the brain. The parasubthalamic nucleus (PSTN), a relatively unexplored population of neurons in the posterior hypothalamus, has been hypothesized to regulate appetite due to its connectivity with other anorexigenic neuronal populations and because these neurons express Fos, a marker of neuronal activation, following a meal. However, the individual cell types that make up the PSTN are not well characterized, nor are their functional roles in food intake behavior. Here, we identify and distinguish between two discrete PSTN subpopulations, those that express tachykinin-1 (PSTNTac1 neurons) and those that express corticotropin-releasing hormone (PSTNCRH neurons), and use a panel of genetically encoded tools in mice to show that PSTNTac1 neurons play an important role in appetite suppression. Both subpopulations increase activity following a meal and in response to administration of the anorexigenic hormones amylin, cholecystokinin (CCK), and peptide YY (PYY). Interestingly, chemogenetic inhibition of PSTNTac1, but not PSTNCRH neurons, reduces the appetite-suppressing effects of these hormones. Consistently, optogenetic and chemogenetic stimulation of PSTNTac1 neurons, but not PSTNCRH neurons, reduces food intake in hungry mice. PSTNTac1 and PSTNCRH neurons project to distinct downstream brain regions, and stimulation of PSTNTac1 projections to individual anorexigenic populations reduces food consumption. Taken together, these results reveal the functional properties and projection patterns of distinct PSTN cell types and demonstrate an anorexigenic role for PSTNTac1 neurons in the hormonal and central regulation of appetite.