Atypical visual perception is often described in autism spectrum disorder (ASD); however, few studies have characterized ocular conditions in ASD using basic vision metrics such as those collected in routine eye exams. The current study uses electronic health record (EHR) codes to establish ocular phenotypes across individuals with and without neurodevelopmental diagnoses, including ASD. Using a population health approach, we assessed ocular conditions (identified based on medical codes from the EHR) in N = 7518 pediatric patients across 4 groups: n = 1196 with ASD, n = 156 with Intellectual Disability (ID), n = 347 with Language Disorder (LD), and n = 5819 matched controls (MC). We grouped and summarized ocular conditions across 5 ocular classes, including: (1) Visual impairment; (2) Refractive error, Accommodative & Vergence disorders; (3) Eye movements, Strabismus & Oculomotor Disorders; (4) Retinal disorders & Ocular disease; (5) Photosensitivity & Atypical Pupil response. We find an increased rate of ocular conditions in diagnostic groups compared to matched controls across classes 1 and 3. This study highlights the use of EHR data to curate ocular condition metrics collected in clinical care. The characterization of ocular anomalies across categories using EHR data offers a scalable method to improve our understanding of vision phenotypes that may be present in children with ASD and other neurodevelopmental differences.
Amidst the opioid crisis, understanding the genetic basis of opioid use disorder (OUD) is crucial for identifying biological mechanisms and intervention points. However, genome-wide association studies (GWASs) have been hampered by inadequate sample sizes and often the use of control populations not assessed for prior opioid exposure. Because opioid exposure is a prerequisite for the development of OUD, consideration of exposure history in controls is important. Electronic health record data (EHR) paired with genomic information allow a broader sampling of patients with OUD and exposed controls. We leveraged data across two healthcare systems to evaluate the impact of using controls not screened for opioid exposure ('generic') versus minimally opioid-exposed control ('exposed'). First, at the phenotypic level, we conducted phenome-wide association studies (PheWAS) to compare the medical comorbidity profiles of OUD cases when using generic versus exposed controls. While PheWAS results for OUD-related comorbidities were more pronounced when using the generic group, 83% of the disease associations were overlapping and of similar effect sizes. Second, at the genetic level, we conducted GWAS (cases vs. generic; cases vs. exposed) and assessed differences in genetic correlations and degrees of phenotypic misclassification. Genetic results were concordant across control groups based on heritability (generic: 0.16 ± 0.07 vs. 0.10 ± 0.07), associations with the coding OPRM1 variant rs1799971 (pgeneric = 8.83E-03 vs. pexposed = 1.83E-02) and genetic correlations with prior OUD GWAS (rg-generic = 0.83 ± 0.26 vs. rg-exposed = 0.78 ± 0.27). Although GWASs were limited by sample size (Ngeneric = 6269, Nexposed = 6365), compared to an independent OUD GWAS (N = 425 944), the dilution value for the two GWAS was not different from 1, suggesting no major impact of phenotypic misclassification. This study represents the first effort to enhance OUD genetic research through optimization of control definitions using EHR data. Generic controls ascertained within the US health systems, where exposure to prescription opioids is high, offer a practical alternative for genetic studies of OUD.
Introduction: Recreational and medical cannabis use (CU) information is often available within the electronic health record (EHR) in a format that is impractical for health care provider use. Transformation of free-text EHR documentation in notes to discrete elements is possible using natural language processing (NLP) and has the potential to characterize CU efficiently. The objective of this study was to develop an NLP algorithm to identify CU documentation within unstructured EHR clinical notes. Methods: We identified EHR notes with cannabis-related terminologies through a keyword search among all Geisinger patients with at least one encounter between 1/1/2013 and 6/30/2022. We trained four NLP models to classify CU documentation within notes into six categories based on time, context, and reliability, as identified through manual annotation. We compared the demographic characteristics of patients with a positive CU classification using the best-performing model to those of the studied sample. Results: Of the over 1.7 million eligible patients, 150,726 (8.6%) were flagged as cannabis users. Bio-ClinicalBERT, a transformer-based NLP model, achieved close to human performance in classifying CU (weighted precision=91.4, recall=93.3, F-score=92.4). An unadjusted analysis showed that cannabis users had higher body mass index (BMI) and were at least nine-fold more likely to use tobacco, alcohol, or illicit substances. Conclusion: Our study evaluated the prevalence of CU documentation across the entire corpus of EHR notes data based on available data without population segmentation over a 9.5-year period. The NLP methodologies used achieved performance close to that of human annotation and laid the foundation for identifying and classifying CU within unstructured data sources, with future applications in research and patient care.
The human brain folds in utero, primarily during late gestation. Shortly after birth, cortical folding patterns are established and remain stable thereafter, making them promising early neurodevelopmental markers. Yet it is unclear whether representations given by current neuroimaging foundation models capture cortical folding variability. Here, we introduce Champollion, a self-supervised learning framework that learns interpretable local representations of cortical folding from structural MRI. Optimized on representative folding-related tasks, Champollion accurately captures known folding patterns across cortical regions and external datasets. In a comprehensive benchmark, it consistently outperforms neuroimaging and general-purpose foundation models. Furthermore, Champollion reveals richer genetic associations than conventional morphometric descriptors and identifies localized folding signatures associated with incomplete hippocampal inversion, prematurity, and maternal smoking. These results establish cortical folding as a rich and largely untapped source of neurodevelopmental information and illustrate how pre-processing and architectural inductive biases can recover biologically meaningful signals overlooked by current generalist foundation models.
Sensory processing differences, particularly within the visual domain, are common in neurodevelopmental conditions, including autism. Studies examining hierarchical processing of figures containing global (i.e., gist) and local (i.e., detail) elements are inconsistent but converge on a common theme in relation to autism: slowed global processing and a locally-oriented default. We examined behavioral and pupillary responses in adults with varying levels of autistic traits during a free-viewing hierarchical processing task. Results showed that participants were both more likely and faster to report global elements, but contrary to our hypothesis, differences in level of autistic traits were unrelated to spontaneous reporting of global vs. local elements. When examining phase-based analysis of pupillary responses, participants high on autistic traits showed more early and less later constriction within trials. Further, trajectory-based pupillary analysis revealed two trajectories, one characterized by constriction and the other dilation, and results showed that the dilation group disproportionately included low traits individuals. Findings suggest that although high and low traits groups showed similar behavioral responses, visual strategies used may differ, as indicated by pupillometry. This study advances our understanding of the relationship between autistic traits and visual processing, laying groundwork for further investigations into neurodivergent visual processing mechanisms.
Introduction: We are in the midst of an opioid epidemic. In the USA, more than a third of the country knows someone who has died from an opioid overdose. Prescription opioids (e.g., oxycodone, hydrocodone, and fentanyl) are commonly used and misused, and it has been estimated that approximately 8–12% of individuals who misuse opioids will subsequently develop an opioid use disorder (OUD). While emphasis has been placed on understanding OUD and the associated adverse effects, there remains a critical gap in systematically characterizing the multifactorial pathways (e.g., behavioral, clinical, genetic, and socio-demographic characteristics) that contribute to the transition from initial use to misuse to OUD. Methods: To address this gap, we introduce the Prescription Opioid Medication Survey (POMS), an online 120-item assessment that compiles multiple validated and standardized instruments. POMS is intended for individuals with any lifetime prescription opioid use. POMS captures various aspects of prescription opioid use including data on opioid use patterns, subjective effects (e.g., euphoria, nausea), problematic use, withdrawal, OUD, overdose, treatment history, and remission. It also addresses comorbid risk factors such as surgical history, chronic pain, other substance use disorders (SUD; e.g., nicotine, alcohol, cannabis, stimulants), other addictive behaviors (i.e., gambling, sexual behaviors, and gaming), and family history of SUD and other addictive behaviors. Mental health assessments, including screening for depression and anxiety, self-reports of eight psychiatric disorders (anxiety, depression, bipolar, schizophrenia, attention-deficit/hyperactivity disorder, post-traumatic stress disorder, obsessive-compulsive disorder, eating disorders), and related mental health conditions (e.g., loneliness, suicide, trauma) are included, along with data on personality traits (e.g., risk-taking, delay discounting, wisdom) and socio-demographic factors. POMS is intended to be administered in clinical settings and large population-based cohorts, facilitating data collection that can enable discoveries to inform better prevention and intervention strategies for OUD. Conclusion: POMS offers a comprehensive tool for systematically capturing the multifactorial risk factors associated with opioid misuse and OUD, providing insights that can inform prevention and intervention strategies.
Background Every year in the United States, over 80,000 individuals die from opioid overdoses, and over 5.7 million are affected by opioid use disorder (OUD). Electronic health records (EHRs) linked with biobanks can increase sample sizes in genome-wide association studies (GWAS). The standard approach for case identification within EHRs is to use the International Classification of Diseases (ICD) codes. However, we have demonstrated previously in both BioVU (biorepository of Vanderbilt University Medical Center) and Geisinger Health System (GHS) that ICD codes alone are insufficient for case identification and that patients prescribed long-term opioids frequently demonstrate problematic opioid use based on clinical note review. These findings suggest potential overlap in the phenotype of individuals prescribed long-term opioids and formally diagnosed OUD patients, but the utility for combining these groups for genetic studies remains to be evaluated. Methods Here we performed two OUD GWAS and varied the case definitions within two large scale EHR-linked biobanks, BioVU and GHS (including 217,774 individuals with high genetic similarity to the 1000 Genomes Project [1KG] European reference population [1KG-EU-clustered]). We based the two case definitions on our previously published work, including one that included individuals with at least one OUD ICD code and a second definition that included patients with more than 10 opioid prescriptions in 12 months. Two GWAS were then performed using these case definitions: 1) OUD ICD code group and/ or individuals prescribed long-term opioids (N case=26,582; N control=191,192) and 2) OUD ICD code group only (N case=7,958; N control=209,816). GWAS summary statistics were meta-analyzed, and SNP-based heritability and genetic correlation (rg) with comorbid traits. Results The GWAS meta-analysis that included both OUD ICD code group and/ or long-term opioid use as cases identified three genome-wide significant loci: rs1799971 (P=2.91e-9), mapped to OPRM1, a missense variant that influences the affinity of the mu-opioid receptor; rs9989877 (P=1.71e-12), mapped to NRXN1-DT, which was previously implicated in smoking initiation and pain intensity; and rs9289747 (P=1.78e-8), which was not reported in prior GWAS. GWAS including only OUD ICD cases did not identify any genome-wide significant associations. GWAS using both case groups identified similar genetic correlations with opioid-related traits, including OUD (0.65±0.07 vs ICD-only 0.80±0.08) and general addiction (0.62±0.05 vs 0.73±0.06). However, GWAS using both case groups showed a higher genetic correlation with chronic pain (0.65±0.04 vs 0.42±0.04), suggesting that including individuals prescribed long-term opioids as cases may also be enriching for pain-related signals. Discussion We conclude that in hospital-based biobanks where OUD patients and individuals prescribed long-term opioids exhibit similar comorbidity profiles, defining cases based on long-term opioid exposure alongside those identified by OUD ICD codes can enhance GWAS discovery power. However, this approach may also capture pain-related genetic signals, underscoring the need for future analyses to disentangle the genetic contributions of OUD and chronic pain. Additionally, we note that the comorbidity profiles of individuals prescribed long-term opioids may vary across hospitals. Therefore, further validation will be conducted at additional PsycheMERGE sites and All of Us.
The brain surface is composed of humps called gyri, separated by grooves called sulci. Although the main folds are common to all individuals, their shape varies, making them unique to each individual. Cortical folding may contain biomarkers that have yet to be deciphered. While conventional geometric approaches fail to fully characterize the high inter-individual variability, recent efforts in large-scale MRI data collection allow us to leverage the statistical power of deep neural networks. Here, we introduce Champollion V0, a self-supervised learning (SSL) algorithm to sort sulcal variability based on 21,070 subjects from the UKBioBank dataset. We revisit from scratch an existing model and optimize its ability to retrieve hand-labeled patterns defined by the neuroscientific community. Under linear evaluation on the latent space, Champollion V0 significantly improves the detection of three different kinds of folding patterns: the presence of a parallel sulcus (AUC increases from 73 R^2 increases on each of the six main geometric features), respectively in the cingulate, the orbital and the central region. These hand-labeled patterns were found to be correlated to neurodevelopmental pathologies. Champollion V0 could enable the automatic labeling of larger datasets for future studies. The code can be found on Github .
Medical Marijuana (MMJ) is available in Pennsylvania and participation in the state-regulated program requires a patient to register and receive a certification by an approved physician. There is currently no integration of MMJ certification data in Pennsylvania into health records that would allow for physicians to rapidly identify patients that are using MMJ, as there are with other scheduled drugs. This absence of a formal data sharing structure necessitates tools that aid in consistent documentation practices to enable comprehensive patient care. Customized smart data elements (SDE) were made available to clinicians at an integrated health system, Geisinger, following MMJ legalization in Pennsylvania. The purpose of this project was to examine and contextualize the use of MMJ SDEs in the Geisinger population. We accomplished this goal by developing a systematic chart review protocol, with the goal of creating a tool that resulted in consistent human data extraction. We developed a chart review protocol for extracting MMJ-related information. The protocol was developed between August to December of 2022 and focused on a patient group that received one of several MMJ SDE between 1/25/2019 and 5/26/2022. Characteristics were first identified on a small pilot sample of patients (N=5), which were then iteratively reviewed to optimize for consistency. Following the pilot, two reviewers were assigned 200 patient charts, selected randomly from the larger cohort, with a third reviewer examining a subsample to determine reliability. We then summarized the clinician-level and patient-level features from n=156 charts with a table-format SDE that best captured MMJ information. We found the chart review protocol was feasible for those with minimal medical background to complete, with high inter-rater reliability (Kappa = 0.966 (p <0.001), 95% CI (0.954 - 0.978)). MMJ certification was largely documented by nurses and medical assistants (87.2%) and typically within primary care settings (68.6%). The SDE has 6 pre-set field prompts, including certifying provider, authorized dispensary, certifying conditions, dosage, product, and active ingredient. We found preset fields were overall well-recorded (76.6% across all fields). Individual fields were more heterogeneous in terms of completion, with dispensary specified in 87.8% of documentation, certifying provider specified in 61.5% of documentation, and product dose specified in only 30.8% of documentation. This method of chart review yields high quality data extraction that can serve as a model for other health record inquiries. Our evaluation showed relatively high completeness of SDE fields, primarily by clinical staff responsible for rooming patients. Improving adoption and fidelity of SDE data collection may present a valuable data source for future research on patient MMJ use and treatment efficacy and outcomes. N/A
BackgroundInformation regarding opioid use disorder (OUD) status and severity is important for patient care. Clinical notes provide valuable information for detecting and characterizing problematic opioid use, necessitating development of natural language processing (NLP) tools, which in turn requires reliably labeled OUD-relevant text and understanding of documentation patterns. ObjectiveTo inform automated NLP methods, we aimed to develop and evaluate an annotation schema for characterizing OUD and its severity, and to document patterns of OUD-relevant information within clinical notes of heterogeneous patient cohorts. MethodsWe developed an annotation schema to characterize OUD severity based on criteria from the Diagnostic and Statistical Manual of Mental Disorders, 5th edition. In total, 2 annotators reviewed clinical notes from key encounters of 100 adult patients with varied evidence of OUD, including patients with and those without chronic pain, with and without medication treatment for OUD, and a control group. We completed annotations at the sentence level. We calculated severity scores based on annotation of note text with 18 classes aligned with criteria for OUD severity and determined positive predictive values for OUD severity. ResultsThe annotation schema contained 27 classes. We annotated 1436 sentences from 82 patients; notes of 18 patients (11 of whom were controls) contained no relevant information. Interannotator agreement was above 70% for 11 of 15 batches of reviewed notes. Severity scores for control group patients were all 0. Among noncontrol patients, the mean severity score was 5.1 (SD 3.2), indicating moderate OUD, and the positive predictive value for detecting moderate or severe OUD was 0.71. Progress notes and notes from emergency department and outpatient settings contained the most and greatest diversity of information. Substance misuse and psychiatric classes were most prevalent and highly correlated across note types with high co-occurrence across patients. ConclusionsImplementation of the annotation schema demonstrated strong potential for inferring OUD severity based on key information in a small set of clinical notes and highlighting where such information is documented. These advancements will facilitate NLP tool development to improve OUD prevention, diagnosis, and treatment.
Background: Medical marijuana (MMJ) is available in Pennsylvania, and participation in the state-regulated program requires patient registration and receiving certification by an approved physician. Currently, no integration of MMJ certification data with health records exists in Pennsylvania that would allow clinicians to rapidly identify patients using MMJ, as exists with other scheduled drugs. This absence of a formal data sharing structure necessitates tools aiding in consistent documentation practices to enable comprehensive patient care. Customized smart data elements (SDEs) were made available to clinicians at an integrated health system, Geisinger, following MMJ legalization in Pennsylvania. Objective: The purpose of this project was to examine and contextualize the use of MMJ SDEs in the Geisinger population. We accomplished this goal by developing a systematic protocol for review of medical records and creating a tool that resulted in consistent human data extraction. Methods: We developed a protocol for reviewing medical records for extracting MMJ-related information. The protocol was developed between August and December of 2022 and focused on a patient group that received one of several MMJ SDEs between January 25, 2019, and May 26, 2022. Characteristics were first identified on a pilot sample (n=5), which were then iteratively reviewed to optimize for consistency. Following the pilot, 2 reviewers were assigned 200 randomly selected patients' medical records, with a third reviewer examining a subsample (n=30) to determine reliability. We then summarized the clinician- and patient-level features from 156 medical records with a table-format SDE that best captured MMJ information. Results: We found the review protocol for medical records was feasible for those with minimal medical background to complete, with high interrater reliability (kappa=0.966; P< .001; odds ratio 0.97, 95% CI 0.954-0.978). MMJ certification was largely documented by nurses and medical assistants (n=138, 88.5%) and typically within primary care settings (n=107, 68.6%). The SDE has 6 preset field prompts with heterogeneous documentation completion rates, including certifying conditions (n=146, 93.6%), product (n=145, 92.9%), authorized dispensary (n=137, 87.8%), active ingredient (n=130, 83.3%), certifying provider (n=96, 61.5%), and dosage (n=48, 30.8%). We found preset fields were overall well-recorded (mean 76.6%, SD 23.7% across all fields). Primary diagnostic codes recorded at documentation encounters varied, with the most frequent being routine examinations and testing (n=34, 21.8%), musculoskeletal or nervous conditions, and signs and symptoms not classified elsewhere (n=21, 13.5%). Conclusions: This method of reviewing medical records yields high-quality data extraction that can serve as a model for other health record inquiries. Our evaluation showed relatively high completeness of SDE fields, primarily by clinical staff responsible for rooming patients, with an overview of conditions under which MMJ is documented. Improving the adoption and fidelity of SDE data collection may present a valuable data source for future research on patient MMJ use, treatment efficacy, and outcomes.
Opioid misuse, addiction, and associated overdose deaths remain global public health crises. Despite the tremendous need for pharmacological treatments, current options are limited in number, use, and effectiveness. Fundamental leaps forward in our understanding of the biology driving opioid addiction are needed to guide development of more effective medication-assisted therapies. This Review focuses on the omics-identified biological features associated with opioid addiction. Recent GWAS have begun to identify robust genetic associations, including variants in OPRM1, , FURIN, , and the gene cluster SCAI / PPP6C / RABEPK . An increasing number of omics studies of postmortem human brain tissue examining biological features (e.g., histone modification and gene expression) across different brain regions have identified broad gene dysregulation associated with overdose death among opioid misusers. Drawn together by meta-analysis and multi-omic systems biology, and informed by model organism studies, key biological pathways enriched for opioid addiction-associated genes are emerging, which include specific receptors (e.g., GABAB receptors, GPCR, and Trk) linked to signaling pathways (e.g., Trk, ERK/MAPK, orexin) that are associated with synaptic plasticity and neuronal signaling. Studies leveraging the agnostic discovery power of omics and placing it within the context of functional neurobiology will propel us toward much-needed, field-changing breakthroughs, including identification of actionable targets for drug development to treat this devastating brain disease.
Background Participant recruitment in rural and hard-to-reach (HTR) populations can present unique challenges. These challenges are further exacerbated by the need for low-cost recruiting, which often leads to use of web-based recruitment methods (eg, email, social media). Despite these challenges, recruitment strategy statistics that support effective enrollment strategies for underserved and HTR populations are underreported. This study highlights how a recruitment strategy that uses email in combination with follow-up, mostly phone calls and email reminders, produced a higher-than-expected enrollment rate that includes a diversity of participants from rural, Appalachian populations in older age brackets and reports recruitment and demographic statistics within a subset of HTR populations. Objective This study aims to provide evidence that a recruitment strategy that uses a combination of email, telephonic, and follow-up recruitment strategies increases recruitment rates in various HTR populations, specifically in rural, older, and Appalachian populations. Methods We evaluated the overall enrollment rate of 1 recruitment arm of a larger study that aims to understand the relationship between genetics and substance use disorders. We evaluated the enrolled population’s characteristics to determine recruitment success of a combined email and follow-up recruitment strategy, and the enrollment rate of HTR populations. These characteristics included (1) enrollment rate before versus after follow-up; (2) zip code and county of enrollee to determine rural or urban and Appalachian status; (3) age to verify recruitment in all eligible age brackets; and (4) sex distribution among age brackets and rural or urban status. Results The email and follow-up arm of the study had a 17.4% enrollment rate. Of the enrolled participants, 76.3% (4602/6030) lived in rural counties and 23.7% (1428/6030) lived in urban counties in Pennsylvania. In addition, of patients enrolled, 98.7% (5956/6030) were from Appalachian counties and 1.3% (76/6030) were from non-Appalachian counties. Patients from rural Appalachia made up 76.2% (4603/6030) of the total rural population. Enrolled patients represented all eligible age brackets from ages 20 to 75 years, with the 60-70 years age bracket having the most enrollees. Females made up 72.5% (4371/6030) of the enrolled population and males made up 27.5% (1659/6030) of the population. Conclusions Results indicate that a web-based recruitment method with participant follow-up, such as a phone call and email follow-up, increases enrollment numbers more than web-based methods alone for rural, Appalachian, and older populations. Adding a humanizing component, such as a live person phone call, may be a key element needed to establish trust and encourage patients from underserved and rural areas to enroll in studies via web-based recruitment methods. Supporting statistics on this recruitment strategy should help researchers identify whether this strategy may be useful in future studies and HTR populations.
BACKGROUND:Opioid addiction is a worldwide public health crisis. In the United States, for example, opioids cause more drug overdose deaths than any other substance. However, opioid addiction treatments have limited efficacy, meaning that additional treatments are needed. METHODS:To help address this problem, we used network-based machine learning techniques to integrate results from genome-wide association studies of opioid use disorder and problematic prescription opioid misuse with transcriptomic, proteomic, and epigenetic data from the dorsolateral prefrontal cortex of people who died of opioid overdose and control individuals. RESULTS:We identified 211 highly interrelated genes identified by genome-wide association studies or dysregulation in the dorsolateral prefrontal cortex of people who died of opioid overdose that implicated the Akt, BDNF (brain-derived neurotrophic factor), and ERK (extracellular signal-regulated kinase) pathways, identifying 414 drugs targeting 48 of these opioid addiction-associated genes. Some of the identified drugs are approved to treat other substance use disorders or depression. CONCLUSIONS:Our synthesis of multiomics using a systems biology approach revealed key gene targets that could contribute to drug repurposing, genetics-informed addiction treatment, and future discovery.
Less common orbitofrontal cortex (OFC) sucogyral patterns are observed at higher rates among those witth psychopathology. Previous work has assumed demographic characteristics have no influence on OFC sulcogyral patterns. However, the influence of sociodemographic and health-related characteristics on OFC patterns within a neurotypical population has not been formally evaluated. We used structural brain magnetic resonance imaging (MRI) from a cohort from the Human Connectome Project (HCP) with existing OFC sulcogyral characterizations (n = 238); none of the participants had psychiatric diagnoses. We evaluated distributions of participant demographics (i.e., age), socioeconomic factors (i.e., employment), and health history-related factors (i.e., smoking history) by OFC sulcogyral pattern within each hemisphere. We then used logistic regression to estimate the odds of OFC sulcogyral pattern by participant characteristics. Distributions of study sample characteristics did not vary substantially by OFC sulcogyral pattern type within either hemisphere. Findings from logistic regression analyses suggest no association between OFC sulcogyral pattern and any of the demographic or socioeconomic characteristics. Two health history-related characteristics, body mass index (BMI) and smoking history, were associated with increased odds of having specific OFC pattern types. For example, individuals with obesity had 2.65 increased odds (95% CI: 1.17, 6.65) of having OFC sulcogyral pattern Type II, III, or IV, compared with Type I in the left hemisphere with normal BMIs. We did not observe substantial influence of demographic or socioeconomic characteristics on OFC sulcogyral patterns. These results confirm assumptions made in previous work that demographic and socioeconomic characteristics do not seem to impact OFC patterns. We do show some evidence for an influence of health history-related factors (obesity and smoking history); future work should evaluate whether these and other phenotypic risk factors interact to modify the relationship between psychiatric diagnoses and OFC sulcogyral patterns.
Genome-wide association analyses (GWAS) of substance use disorders (SUDs) have historically been severely underpowered, due to the difficulty of assembling adequately large, and diverse sample sizes; but traditional ascertainment strategies are often costly and slow. Large health systems with biorepositories and electronic health records (EHR) are emerging as complementary “Big Data” tools for genetic studies of SUDs. EHR contain a variety of data types, including structured data from billing, laboratory test results, and unstructured data from physician notes, yielding extensive longitudinal data that are the byproduct of routine clinical care. Although EHR were not initially designed for genetic research, this wealth of data provides opportunities for SUD genetics, with the potential to support analyses across the phenotypic spectrum. The speakers in this session will present compelling phenotypic and genomic insights from the Substance Use Disorder Workgroup within the PsycheMERGE Consortium. Sandra Sanchez-Roige will demonstrate the value of “phecodes”, or instances of specific diagnostic codes on two or more occasions, as powerful tools for genetic studies of tobacco use disorders (TUD). She will describe a multi-ancestry GWAS meta-analysis of TUD in 898,680 individuals (739,895 European, 114,420 African American, 44,365 Latin American) based on EHR across five biobanks. Analyses of EHR have utility beyond GWAS, they facilitate prediction models that take advantage of the longitudinal structure of EHR for prediction of SUDs and treatment response. Maria Niarchou will present the development and validation of an algorithm for identifying cases and controls of Alcohol Use Disorder (AUD) in Vanderbilt's EHR data and biobank. This approach combines billing codes, structured data, and natural language processing techniques to enhance the accuracy of phenotypic classification of AUD. Brandon Coombes will next evaluate prediction of SUD EHR diagnosis using polygenic data from the largest addiction GWAS compiled by the Psychiatric Genomics Consortium. Using EHR data from the Mayo Clinic Biobank (N=46,000) along with PsycheMERGE sites, this talk will investigate how polygenic scores for addiction, measured via traditional ascertainment, associate with SUDs identified using diagnosis codes from the EHR, and whether SUD-specific polygenic scores improve prediction beyond that predicted by addiction genetic risk. Rachel Kember will describe multi-ancestry phenome-wide association analyses of genetic liability for problematic alcohol use in four PsycheMERGE sites (AFR N=27,494; EUR N=131,500). While heavy alcohol use may potentiate the course of mental and medical illnesses, there is accumulating evidence suggesting common genetic architectures across AUD and other conditions. Her presentation will further demonstrate that genetic liability for AUD is associated with other SUDs, psychiatric, and medical disorders, even in the absence of an AUD diagnosis. Identifying shared pathways of risk has high translational potential, allowing tailoring of treatments for multiple medical conditions. Lea K Davis will serve as discussant. These efforts exemplify how multi-disciplinary teams with complementary expertise can collaborate to enhance discovery for substance use disorders.
Background: Sleep disturbances, gastrointestinal problems, and atypical heart rate are commonly observed in patients with autism spectrum disorder (ASD) and may relate to underlying function of the autonomic nervous system (ANS). The overall objective of the current study was to quantitatively characterize features of ANS function using symptom scales and available electronic health record (EHR) data in a clinically and genetically characterized pediatric cohort. Methods: We assessed features of ANS function via chart review of patient records adapted from items drawn from a clinical research questionnaire of autonomic symptoms. This procedure coded for the presence and/or absence of targeted symptoms and was completed in 3 groups of patients, including patients with a clinical neurodevelopmental diagnosis and identified genetic etiology (NPD, n = 244), those with an ASD diagnosis with no known genetic cause (ASD, n = 159), and age and sex matched controls (MC, n = 213). Symptoms were assessed across four main categories: (1) Mood, Behavior, and Emotion; (2) Secretomotor, Sensory Integration; (3) Urinary, Gastrointestinal, and Digestion; and (4) Circulation, Thermoregulation, Circadian function, and Sleep/Wake cycles. Results: Chart review scores indicate an increased rate of autonomic symptoms across all four sections in our NPD group as compared to scores with ASD and/or MC. Additionally, we note several significant relationships between individual differences in autonomic symptoms and quantitative ASD traits. Conclusion: These results highlight EHR review as a potentially useful method for quantifying variance in symptoms adapted from a questionnaire or survey. Further, using this method indicates that autonomic features are more prevalent in children with genetic disorders conferring risk for ASD and other neurodevelopmental diagnoses.
Background: We used structured and unstructured electronic health record (EHR) data to develop and validate an approach to identify moderate/severe opioid use disorder (OUD) that includes individuals without prescription opioid use or chronic pain, an underrepresented population. Methods: Using electronic diagnosis grouper text from EHRs of similar to 1 million patients (2012-2020), we created indicators of OUD-with "tiers" indicating OUD likelihood-combined with OUD medication (MOUD) orders. We developed six sub-algorithms with varying criteria (multiple vs single MOUD orders, multiple vs single tier 1 indicators, tier 2 indicators, tier 3 and 4 indicators). Positive predictive values (PPVs) were calculated based on chart review to determine OUD status and severity. We compared demographic and clinical characteristics of cases identified by the sub-algorithms. Results: In total, 14,852 patients met criteria for one of the sub-algorithms. Five sub-algorithms had PPVs >= 0.90 for any severity OUD; four had PPVs >= 0.90 for moderate/severe OUD. Demographic and clinical characteristics differed substantially between groups. Of identified OUD cases, 31.3% had no past opioid analgesic orders, 79.7% lacked evidence of chronic prescription opioid use, and 43.5% lacked a chronic pain diagnosis. Discussion: Incorporating unstructured data with MOUD orders yielded an approach that adequately identified moderate/severe OUD, identified unique demographic and clinical sub-groups, and included individuals without prescription opioid use or chronic pain, whose OUD may stem from illicit opioids. Findings show that incorporating unstructured data strengthens EHR algorithms for identifying OUD and suggests approaches limited to populations with prescription opioid use or chronic pain exclude many individuals with OUD.
Introduction: Opioid use disorders (OUDs) constitute a major public health issue, and we urgently need alternative methods for characterizing risk for OUD. Electronic health records (EHRs) are useful tools for understanding complex medical phenotypes but have been underutilized for OUD because of challenges related to underdiagnosis, binary diagnostic frameworks, and minimally characterized reference groups. As a first step in addressing these challenges, a new paradigm is warranted that characterizes risk for opioid prescription misuse on a continuous scale of severity, i.e., as a continuum. Methods: Across sites within the PsycheMERGE network, we extracted prescription opioid data and diagnoses that co-occur with OUD (including psychiatric and substance use disorders, pain-related diagnoses, HIV, and hepatitis C) for over 2.6 million patients across three health registries (Vanderbilt University Medical Center, Mass General Brigham, Geisinger) between 2005 and 2018. We defined three groups based on levels of opioid exposure: no prescriptions, minimal exposure, and chronic exposure and then compared the comorbidity profiles of these groups to the full registries and to those with OUD diagnostic codes. Results: Our results confirm that EHR data reflects known higher prevalence of substance use disorders, psychiatric disorders, medical, and pain diagnoses in patients with OUD diagnoses and chronic opioid use. Comorbidity profiles that distinguish opioid exposure are strikingly consistent across large health systems, indicating the phenotypes described in this new quantitative framework are robust to health systems differences. Conclusion: This work indicates that EHR prescription opioid data can serve as a platform to characterize complex risk markers for OUD using existing data.