Adverse drug events (ADEs) cost society lives and an estimated $30 billion per year in the USA alone. Their prevalence has led to the public losing trust in the safety of drugs, especially generics (e.g., Eban, 2019). These concerns have motivated the wide study of methods for general ADE discovery, but discovering ADEs in generic drugs challenges causal discovery methods with a scenario of multiple treatments over time, a scenario which presents new problems and opportunities for machine learning. In response, this research develops methods for causal discovery based on analyzing controlled before-after studies with differential prediction and temporal inverse probability weighting. These methods are easy to realize by employing off-the-shelf machine learning classifiers. Experiments on both synthetic and real electronic health records demonstrate the ability of the methods to control for confounding, discover generic-specific ADEs in synthetic data, and hypothesize brand-generic differences in real-world data that agree with known ones. These are the abilities that causal discovery methods need for helping establish the facts of generic drug safety.
Objective Diverticular disease (DD) is one of the most prevalent conditions encountered by gastroenterologists, affecting ~50% of Americans before the age of 60. Our aim was to identify genetic risk variants and clinical phenotypes associated with DD, leveraging multiple electronic health record (EHR) data sources of 91,166 multi-ancestry participants with a Natural Language Processing (NLP) technique. Materials and methods We developed a NLP-enriched phenotyping algorithm that incorporated colonoscopy or abdominal imaging reports to identify patients with diverticulosis and diverticulitis from multicenter EHRs. We performed genome-wide association studies (GWAS) of DD in European, African and multi-ancestry participants, followed by phenome-wide association studies (PheWAS) of the risk variants to identify their potential comorbid/pleiotropic effects in clinical phenotypes. Results Our developed algorithm showed a significant improvement in patient classification performance for DD analysis (algorithm PPVs ≥ 0.94), with up to a 3.5 fold increase in terms of the number of identified patients than the traditional method. Ancestry-stratified analyses of diverticulosis and diverticulitis of the identified subjects replicated the well-established associations between ARHGAP15 loci with DD, showing overall intensified GWAS signals in diverticulitis patients compared to diverticulosis patients. Our PheWAS analyses identified significant associations between the DD GWAS variants and circulatory system, genitourinary, and neoplastic EHR phenotypes. Discussion As the first multi-ancestry GWAS-PheWAS study, we showcased that heterogenous EHR data can be mapped through an integrative analytical pipeline and reveal significant genotype-phenotype associations with clinical interpretation. Conclusion A systematic framework to process unstructured EHR data with NLP could advance a deep and scalable phenotyping for better patient identification and facilitate etiological investigation of a disease with multilayered data.
OBJECTIVE:Natural language processing (NLP) systems convert unstructured text into analyzable data. Here, we describe the performance measures of NLP to capture granular details on nodules from thyroid ultrasound (US) reports and reveal critical issues with reporting language.METHODS:We iteratively developed NLP tools using clinical Text Analysis and Knowledge Extraction System (cTAKES) and thyroid US reports from 2007 to 2013. We incorporated nine nodule features for NLP extraction. Next, we evaluated the precision, recall, and accuracy of our NLP tools using a separate set of US reports from an academic medical center (A) and a regional health care system (B) during the same period. Two physicians manually annotated each test-set report. A third physician then adjudicated discrepancies. The adjudicated "gold standard" was then used to evaluate NLP performance on the test-set.RESULTS:A total of 243 thyroid US reports contained 6,405 data elements. Inter-annotator agreement for all elements was 91.3%. Compared with the gold standard, overall recall of the NLP tool was 90%. NLP recall for thyroid lobe or isthmus characteristics was: laterality 96% and size 95%. NLP accuracy for nodule characteristics was: laterality 92%, size 92%, calcifications 76%, vascularity 65%, echogenicity 62%, contents 76%, and borders 40%. NLP recall for presence or absence of lymphadenopathy was 61%. Reporting style accounted for 18% errors. For example, the word "heterogeneous" interchangeably referred to nodule contents or echogenicity. While nodule dimensions and laterality were often described, US reports only described contents, echogenicity, vascularity, calcifications, borders, and lymphadenopathy, 46, 41, 17, 15, 9, and 41% of the time, respectively. Most nodule characteristics were equally likely to be described at hospital A compared with hospital B.CONCLUSIONS:NLP can automate extraction of critical information from thyroid US reports. However, ambiguous and incomplete reporting language hinders performance of NLP systems regardless of institutional setting. Standardized or synoptic thyroid US reports could improve NLP performance.
Introduction Currently, one of the commonly used methods for disseminating electronic health record (EHR)-based phenotype algorithms is providing a narrative description of the algorithm logic, often accompanied by flowcharts. A challenge with this mode of dissemination is the potential for under-specification in the algorithm definition, which leads to ambiguity and vagueness. Methods This study examines incidents of under-specification that occurred during the implementation of 34 narrative phenotyping algorithms in the electronic Medical Record and Genomics (eMERGE) network. We reviewed the online communication history between algorithm developers and implementers within the Phenotype Knowledge Base (PheKB) platform, where questions could be raised and answered regarding the intended implementation of a phenotype algorithm. Results We developed a taxonomy of under-specification categories via an iterative review process between two groups of annotators. Under-specifications that lead to ambiguity and vagueness were consistently found across narrative phenotype algorithms developed by all involved eMERGE sites. Discussion and conclusion Our findings highlight that under-specification is an impediment to the accuracy and efficiency of the implementation of current narrative phenotyping algorithms, and we propose approaches for mitigating these issues and improved methods for disseminating EHR phenotyping algorithms.
Chronic Kidney Disease (CKD) represents a slowly progressive disorder that is typically silent until late stages, but early intervention can significantly delay its progression. We designed a portable and scalable electronic CKD phenotype to facilitate early disease recognition and empower large-scale observational and genetic studies of kidney traits. The algorithm uses a combination of rule-based and machine-learning methods to automatically place patients on the staging grid of albuminuria by glomerular filtration rate (“A-by-G” grid). We manually validated the algorithm by 451 chart reviews across three medical systems, demonstrating overall positive predictive value of 95% for CKD cases and 97% for healthy controls. Independent case-control validation using 2350 patient records demonstrated diagnostic specificity of 97% and sensitivity of 87%. Application of the phenotype to 1.3 million patients demonstrated that over 80% of CKD cases are undetected using ICD codes alone. We also demonstrated several large-scale applications of the phenotype, including identifying stage-specific kidney disease comorbidities, in silico estimation of kidney trait heritability in thousands of pedigrees reconstructed from medical records, and biobank-based multicenter genome-wide and phenome-wide association studies.
Assumptions are made about the genetic model of single nucleotide polymorphisms (SNPs) when choosing a traditional genetic encoding: additive, dominant, and recessive. Furthermore, SNPs across the genome are unlikely to demonstrate identical genetic models. However, running SNP-SNP interaction analyses with every combination of encodings raises the multiple testing burden. Here, we present a novel and flexible encoding for genetic interactions, the elastic data-driven genetic encoding (EDGE), in which SNPs are assigned a heterozygous value based on the genetic model they demonstrate in a dataset prior to interaction testing. We assessed the power of EDGE to detect genetic interactions using 29 combinations of simulated genetic models and found it outperformed the traditional encoding methods across 10%, 30%, and 50% minor allele frequencies (MAFs). Further, EDGE maintained a low false-positive rate, while additive and dominant encodings demonstrated inflation. We evaluated EDGE and the traditional encodings with genetic data from the Electronic Medical Records and Genomics (eMERGE) Network for five phenotypes: age-related macular degeneration (AMD), age-related cataract, glaucoma, type 2 diabetes (T2D), and resistant hypertension. A multi-encoding genome-wide association study (GWAS) for each phenotype was performed using the traditional encodings, and the top results of the multi-encoding GWAS were considered for SNP-SNP interaction using the traditional encodings and EDGE. EDGE identified a novel SNP-SNP interaction for age-related cataract that no other method identified: rs7787286 (MAF: 0.041; intergenic region of chromosome 7)–rs4695885 (MAF: 0.34; intergenic region of chromosome 4) with a Bonferroni LRT p of 0.018. A SNP-SNP interaction was found in data from the UK Biobank within 25 kb of these SNPs using the recessive encoding: rs60374751 (MAF: 0.030) and rs6843594 (MAF: 0.34) (Bonferroni LRT p: 0.026). We recommend using EDGE to flexibly detect interactions between SNPs exhibiting diverse action.
Accurate estimation of healthcare costs is crucial for healthcare systems to plan and effectively negotiate with insurance companies regarding the coverage of patient-care costs. Greater accuracy in estimating healthcare costs would provide mutual benefit for both health systems and the insurers that support these systems by better aligning payment models with patient-care costs. This study presents the results of a generalizable machine learning approach to predicting medical events built from 40 years of data from >860,000 patients pertaining to >6,700 prescription medications, courtesy of Marshfield Clinic in Wisconsin. It was found that models built using this approach performed well when compared to similar studies predicting physician prescriptions of individual medications. In addition to providing a comprehensive predictive model for all drugs in a large healthcare system, the approach taken in this research benefits from potential applicability to a wide variety of other medical events.
Atopic dermatitis (AD) is a common chronic inflammatory skin disease that affects up to 30% of children and is highly heritable. To date, 31 genetic loci have been identified through genome-wide association studies (GWASs), mapping to genes involved in immune regulation and skin barrier deficiencies.1Paternoster L. Standl M. Waage J. Baurecht H. Hotze M. Strachan D.P. et al.Multi-ancestry genome-wide association study of 21,000 cases and 95,000 controls identifies new risk loci for atopic dermatitis.Nat Genet. 2015; 47: 1449-1456Crossref PubMed Scopus (362) Google Scholar However, despite the higher prevalence of AD in African Americans (AAs; 15.9% vs 9.7% in European ancestry),2Shaw T.E. Currie G.P. Koudelka C.W. Simpson E.L. Eczema prevalence in the United States: data from the 2003 National Survey of Children's Health.J Invest Dermatol. 2011; 131: 67-73Abstract Full Text Full Text PDF PubMed Scopus (542) Google Scholar this population has been underrepresented in GWASs and specific genetic risk factors remain elusive. In this study, we report an AD GWAS in the largest AA sample studied to date. For case and control selection, we developed, validated, and implemented an AD phenotyping algorithm based on electronic health record data mining, in collaboration with the electronic Medical Records and Genomics (eMERGE) network. Cases were defined as subjects 2 months or older with either (1) 2 or more AD diagnoses (International Classification of Diseases, Ninth Revision [ICD-9] code 691.8), on separate calendar days, and 1 or more prescriptions of AD-related medications (see Table E1 in this article's Online Repository at www.jacionline.org); or (2) 3 or more AD diagnoses, on separate calendar days. Subjects with scabies, Wiscott-Aldrich syndrome, allergic purpura, ichtyosis congenita, and chromosomal anomalies were excluded. Controls were subjects with no ICD-9 code for AD, no ICD-9 codes for any skin condition, allergy, asthma, nutritional deficiencies, or skin cancer (see Table E2 in this article's Online Repository at www.jacionline.org), and no relevant medications for AD (see Table E1). The algorithm was validated internally at the Center for Applied Genomics (CAG) and also externally at the Marshfield Clinic, with positive predictive values over 92% for cases and controls at both sites (for details, see this article's Online Repository at www.jacionline.org). Implementation of the algorithm at CAG accrued 12,345 AD cases and controls with imputed genotyping data (see Table E3 in this article's Online Repository at www.jacionline.org), including 5,843 of AA ancestry (2,420 cases and 3,423 controls; see Table E4 in this article's Online Repository at www.jacionline.org), which were selected for the analysis. GWASs on 11,743,754 variants in the AA sample did not result in any variants surpassing genome-wide significance; however, 150 variants at 49 loci showed nominal P values of less than 10−5 that were followed-up in an independent AA sample from the SAGE II study. Meta-analysis of the 2 AA samples resulted in a genome-wide significant association at rs3811419 (C>T) (Allele C: β (SE) = 0.573 (0.098); PMETA = 5.64 × 10−9; MAF(T) = 3.7%; PCAG = 1.48 × 10−7; PSAGEII = .011) 1.5 kb upstream of RORC (Fig 1; see Table E5 in this article's Online Repository at www.jacionline.org). This association also replicated in the 6502 AD cases (N = 510) and controls (N = 5998) of European ancestry that were identified using the same phenotyping algorithm (β (SE) = 0.405 (0.149); P = 6.61 × 10−3; MAF = 11.6%). We explored the potential cis-expression quantitative trait locus (eQTL) effects of rs3811419 (C) and 7 variants in linkage disequilibrium (LD) (r2 > 0.8) with it, in skin and blood, using HaploReg v4,3Ward L.D. Kellis M. HaploReg: a resource for exploring chromatin states, conservation, and regulatory motif alterations within sets of genetically linked variants.Nucleic Acids Res. 2012; 40: D930-D934Crossref PubMed Scopus (1655) Google Scholar the National Center for Biotechnology Information (NCBI) Genotype-Tissue Expression version 6 data,4Lonsdale J. Thomas J. Salvatore M. Phillips R. Lo E. Shad S. et al.The Genotype-Tissue Expression (GTEx) project.Nat Genet. 2013; 45: 580-585Crossref PubMed Scopus (4307) Google Scholar and studies reporting eQTL effects in those tissues (see this article's Online Repository at www.jacionline.org). Seven of the 8 single nucleotide polymorphisms (SNPs) investigated showed eQTL effects on the expression of THEM4 in blood (top effect for rs3811418 (G); P value = 6.22 × 10−14) and 5 showed effects in skin (top effect for rs11582525 (G); P value = 1.50 × 10−5) (see Table E6 in this article's Online Repository at www.jacionline.org). Ferreira et al5Ferreira M.A. Vonk J.M. Baurecht H. Marenholz I. Tian C. Hoffman J.D. et al.Shared genetic origin of asthma, hay fever and eczema elucidates allergic disease biology.Nat Genet. 2017; 49: 1752-1757Crossref PubMed Scopus (274) Google Scholar recently reported association at this locus in a large meta-analysis of allergic disease in individuals of European ancestry. Their top variant, rs11204896, was independent of rs3811419 and did not replicate in our AA sample. Only 9 of the 136 variants reported by Ferreira et al showed nominal significance in the AA analysis (P value < .05). Of those, 7 had the same direction of effect (see Table E7 in this article's Online Repository at www.jacionline.org). It is noteworthy, however, that the authors report that rs11204896-C is also an eQTL for THEM4,5Ferreira M.A. Vonk J.M. Baurecht H. Marenholz I. Tian C. Hoffman J.D. et al.Shared genetic origin of asthma, hay fever and eczema elucidates allergic disease biology.Nat Genet. 2017; 49: 1752-1757Crossref PubMed Scopus (274) Google Scholar supporting a potential role for THEM4 in atopy. Follow-up of associated variants might be needed to exclude the effect being due to LD rather than a causal role of THEM4. eQTL effects were also found for C2CD4D (top effect for rs72692781 (T) = −0.21; P value = 3.2 × 10−5) in skin (Table E6). Interestingly, all variants queried also had weak eQTL effects over FLG-AS1 (FLG antisense RNA 1) in skin (rs3811419 (T), effect = −0.23; P value = 3.60 × 10−3). FLG-AS1 is a long noncoding RNA (lncRNA), located 500 Kb from rs3811419, that overlaps some of the coding FLG and FLG2 sequence on the antisense strand. Although the exact function of FLG-AS1 has not been characterized to date, lncRNAs are increasingly recognized as key players in the regulation of gene expression in developmental and disease processes (reviewed in Beermann et al6Beermann J. Piccoli M.T. Viereck J. Thum T. Non-coding RNAs in development and disease: background, mechanisms, and therapeutic approaches.Physiol Rev. 2016; 96: 1297-1325Crossref PubMed Scopus (1116) Google Scholar). One of the most well-recognized pathophysiologic differences in AD of AAs compared with Europeans and Asians is the absence of genetically determined FLG deficiency. Although null FLG mutations are commonly found in European and Asian populations, they are very rare in individuals of AA ancestry and not considered a risk factor.7Margolis D.J. Gupta J. Apter A.J. Ganguly T. Hoffstad O. Papadopoulos M. et al.Filaggrin-2 variation is associated with more persistent atopic dermatitis in African American subjects.J Allergy Clin Immunol. 2014; 133: 784-789Abstract Full Text Full Text PDF PubMed Scopus (118) Google Scholar The rs3811419-C risk allele would be associated with an increase in the expression of FLG-AS1, which could be translated in a decrease in FLG expression. This eQTL effect over FLG-AS could constitute an alternate mechanism leading to FLG-related skin barrier deficiency in AA that should be further explored. To further characterize the functional consequences of the associated locus, we analyzed the methylation status of blood-derived DNA from 374 AA CAG subjects, generated on the Infinium HumanMethylation450 BeadChip (details in this article's Online Repository at www.jacionline.org). Methylation data were analyzed as the response variable in a linear regression, with the genotype of rs3811419 and the 7 variants in LD with it (r2 > 0.8) as the predictors, similar to an expression QTL analysis. Sex, age, and 10 principal components were included as covariates. The top association was for rs72692781 (T), a variant in LD with rs3811419 (r2 = 0.97), and methylation of probe cg22228337 (β (SE) = −0.187 (0.026); P = 6.70 × 10−12) located 3.1 kb from the variant (chr1:151,802,807). Other associations passing the significance threshold of 10−6 were found for the top SNP rs3811419 (C) and methylation probe cg08477332 (chr1: 153590243; β (SE) = 1.133 (0.225); P = 8.05 × 10−7); and for rs7540799 (T) and probes cg00975746 (chr1: 161409970; β (SE) = 0.619 (0.129); P = 2.30 × 10−6) and cg00304520 (chr1: 161582581; β (SE) = 0.984 (0.208); P = 3.28 × 10−6). Based on data published by the BIOS consortium database, probe cg22228337 is an eQTM for THEM5 (https://genenetwork.nl/biosqtlbrowser/), a gene almost exclusively expressed in skin.4Lonsdale J. Thomas J. Salvatore M. Phillips R. Lo E. Shad S. et al.The Genotype-Tissue Expression (GTEx) project.Nat Genet. 2013; 45: 580-585Crossref PubMed Scopus (4307) Google Scholar Information on the other probes was not available. The strengths and limitations of the use of an EHR-based algorithm for case-control selection for genetic studies has already been described by our group.8Almoguera B. Vazquez L. Mentch F. Connolly J. Pacheco J.A. Sundaresan A.S. et al.Identification of four novel loci in Asthma in European American and African American populations.Am J Respir Crit Care Med. 2017; 195: 456-463Crossref PubMed Scopus (59) Google Scholar However, we could not validate the performance of the algorithm in terms of confirmation of known AD and allergy genetic loci reported in adults (see Tables E7 and E8 in this article's Online Repository at www.jacionline.org)1Paternoster L. Standl M. Waage J. Baurecht H. Hotze M. Strachan D.P. et al.Multi-ancestry genome-wide association study of 21,000 cases and 95,000 controls identifies new risk loci for atopic dermatitis.Nat Genet. 2015; 47: 1449-1456Crossref PubMed Scopus (362) Google Scholar: only 11q13.1 (OVOL1) replicated in the AA population. Interestingly, OVOL1 was also one of the top significant genes showing transethnic effects after the MANTRA meta-analysis performed by Paternoster et al.1Paternoster L. Standl M. Waage J. Baurecht H. Hotze M. Strachan D.P. et al.Multi-ancestry genome-wide association study of 21,000 cases and 95,000 controls identifies new risk loci for atopic dermatitis.Nat Genet. 2015; 47: 1449-1456Crossref PubMed Scopus (362) Google Scholar These results suggest that AD genetic factors described to date cannot be extrapolated to AA populations and specific genetic factors need to be uncovered. In conclusion, we have identified the first genome-wide significant association for AD in AAs and replicated the association in individuals of European ancestry. The top SNP at the locus, rs3811419, is an eQTL for THEM4, a gene previously associated with allergy5Ferreira M.A. Vonk J.M. Baurecht H. Marenholz I. Tian C. Hoffman J.D. et al.Shared genetic origin of asthma, hay fever and eczema elucidates allergic disease biology.Nat Genet. 2017; 49: 1752-1757Crossref PubMed Scopus (274) Google Scholar and FLG-AS1, a long noncoding RNA that overlaps the fillagrin gene. Although we provide replication of the association in individuals of European ancestry, replication in an independent AA sample and elucidation of the exact biological role of these variants in AD is needed. We thank all patients and control subjects for their participation in the study and the eMERGE phenotyping group and researchers. Download .docx (.04 MB) Help with docx files Online Repository text Download .docx (.01 MB) Help with docx files Online References 1-23 Download .docx (.04 MB) Help with docx files Tables E1-E8 Download .docx (.09 MB) Help with docx files Fig E1 Download .docx (.04 MB) Help with docx files Fig E2
Electronic health records (EHR) are valuable to define phenotype selection algorithms used to identify cohorts ofpatients for sequencing or genome wide association studies (GWAS). To date, the electronic medical records and genomics (eMERGE) network institutions have developed and applied such algorithms to identify cohorts with associated DNA samples used to discover new genetic associations. For complex diseases, there are benefits to stratifying cohorts using comorbidities in order to identify their genetic determinants. The objective of this study was to: (a) characterize comorbidities in a range of phenotype-selected cohorts using the Johns Hopkins Adjusted Clinical Groups® (ACG®) System, (b) assess the frequency of important comorbidities in three commonly studied GWAS phenotypes, and (c) compare the comorbidity characterization of cases and controls. Our analysis demonstrates a framework to characterize comorbidities using the ACG system and identified differences in mean chronic condition count among GWAS cases and controls. Thus, we believe there is great potential to use the ACG system to characterize comorbidities among genetic cohorts selected based on EHR phenotypes.
Uterine fibroids affect up to 77% of women by menopause and account for up to $34 billion in healthcare costs each year. Although fibroid risk is heritable, genetic risk for fibroids is not well understood. We conducted a two-stage case-control meta-analysis of genetic variants in European and African ancestry women with and without fibroids classified by a previously published algorithm requiring pelvic imaging or confirmed diagnosis. Women from seven electronic Medical Records and Genomics (eMERGE) network sites (3,704 imaging-confirmed cases and 5,591 imaging-confirmed controls) and women of African and European ancestry from UK Biobank (UKB, 5,772 cases and 61,457 controls) were included in the discovery genome-wide association study (GWAS) meta-analysis. Variants showing evidence of association in Stage I GWAS (P < 1 × 10-5) were targeted in an independent replication sample of African and European ancestry individuals from the UKB (Stage II) (12,358 cases and 138,477 controls). Logistic regression models were fit with genetic markers imputed to a 1000 Genomes reference and adjusted for principal components for each race- and site-specific dataset, followed by fixed-effects meta-analysis. Final analysis with 21,804 cases and 205,525 controls identified 326 genome-wide significant variants in 11 loci, with three novel loci at chromosome 1q24 (sentinel-SNP rs14361789; P = 4.7 × 10-8), chromosome 16q12.1 (sentinel-SNP rs4785384; P = 1.5 × 10-9) and chromosome 20q13.1 (sentinel-SNP rs6094982; P = 2.6 × 10-8). Our statistically significant findings further support previously reported loci including SNPs near WT1, TNRC6B, SYNE1, BET1L, and CDC42/WNT4. We report evidence of ancestry-specific findings for sentinel-SNP rs10917151 in the CDC42/WNT4 locus (P = 1.76 × 10-24). Ancestry-specific effect-estimates for rs10917151 were in opposite directions (P-Het-between-groups = 0.04) for predominantly African (OR = 0.84) and predominantly European women (OR = 1.16). Genetically-predicted gene expression of several genes including LUZP1 in vagina (P = 4.6 × 10-8), OBFC1 in esophageal mucosa (P = 8.7 × 10-8), NUDT13 in multiple tissues including subcutaneous adipose tissue (P = 3.3 × 10-6), and HEATR3 in skeletal muscle tissue (P = 5.8 × 10-6) were associated with fibroids. The finding for HEATR3 was supported by SNP-based summary Mendelian randomization analysis. Our study suggests that fibroid risk variants act through regulatory mechanisms affecting gene expression and are comprised of alleles that are both ancestry-specific and shared across continental ancestries.
In the version of this article originally published, one of the two authors with the name Wei Zhao was omitted from the author list and the affiliations for both authors were assigned to the single Wei Zhao in the author list. In addition, the ORCID for Wei Zhao (Department of Biostatistics and Epidemiology, Perelman School of Medicine, University of Pennsylvania, Philadelphia, PA, USA) was incorrectly assigned to author Wei Zhou. The errors have been corrected in the HTML and PDF versions of the article.
Benign prostatic hyperplasia (BPH) results in a significant public health burden due to the morbidity caused by the disease and many of the available remedies. As much as 70% of men over 70 will develop BPH. Few studies have been conducted to discover the genetic determinants of BPH risk. Understanding the biological basis for this condition may provide necessary insight for development of novel pharmaceutical therapies or risk prediction. We have evaluated SNP-based heritability of BPH in two cohorts and conducted a genome-wide association study (GWAS) of BPH risk using 2,656 cases and 7,763 controls identified from the Electronic Medical Records and Genomics (eMERGE) network. SNP-based heritability estimates suggest that roughly 60% of the phenotypic variation in BPH is accounted for by genetic factors. We used logistic regression to model BPH risk as a function of principal components of ancestry, age, and imputed genotype data, with meta-analysis performed using METAL. The top result was on chromosome 22 in SYN3 at rs2710383 (p-value = 4.6 × 10−7; Odds Ratio = 0.69, 95% confidence interval = 0.55–0.83). Other suggestive signals were near genes GLGC, UNCA13, SORCS1 and between BTBD3 and SPTLC3. We also evaluated genetically-predicted gene expression in prostate tissue. The most significant result was with increasing predicted expression of ETV4 (chr17; p-value = 0.0015). Overexpression of this gene has been associated with poor prognosis in prostate cancer. In conclusion, although there were no genome-wide significant variants identified for BPH susceptibility, we present evidence supporting the heritability of this phenotype, have identified suggestive signals, and evaluated the association between BPH and genetically-predicted gene expression in prostate.
With the rapid expansion of applied 3D computational vision, shape descriptors have become increasingly important for a wide variety of applications and objects from molecules to planets. Appropriate shape descriptors are critical for accurate (and efficient) shape retrieval and 3D model classification. Several spectral-based shape descriptors have been introduced by solving various physical equations over a 3D surface model. In this paper, for the first time, we incorporate a specific group of techniques in statistics and machine learning, known as manifold learning, to develop a global shape descriptor in the computer graphics domain. The proposed descriptor utilizes the Laplacian Eigenmap technique in which the Laplacian eigenvalue problem is discretized using an exponential weighting scheme. As a result, our descriptor eliminates the limitations tied to the existing spectral descriptors, namely dependency on triangular mesh representation and high intra-class quality of 3D models. We also present a straightforward normalization method to obtain a scale-invariant descriptor. The extensive experiments performed in this study show that the present contribution provides a highly discriminative and robust shape descriptor under the presence of a high level of noise, random scale variations, and low sampling rate, in addition to the known isometric-invariance property of the Laplace-Beltrami operator. The proposed method significantly outperforms state-of-the-art algorithms on several non-rigid shape retrieval benchmarks.
Accurate diagnosis of lung nodules is essential for detection and assessment of lung cancer. The present contribution proposes a descriptive model for diagnostic classification of lung nodules by jointly using deep and spectral features from the 3D surface structure of nodules. To the best of our knowledge, this is the first work that utilizes a point cloud (PC)-based deep network for extracting nodule shape features. The PC-based deep network takes into account the 3D context of a nodule; meanwhile, it is extensively less computationally intensive. The spectral features prevent over-fitting, a common problem of deep networks trained by relatively small dataset in the medical imaging domain, and compensates for missing information of mesh connections. Experimental results reveal that our descriptive model demonstrates high sensitivity (87.23%) as well as high specificity (89.80%) with a total accuracy of 88.54% for reliable and accurate prediction of lung nodule malignancy.
Epidemiological studies identifying biological markers of disease state are valuable, but can be time-consuming, expensive, and require extensive intuition and expertise. Furthermore, not all hypothesized markers will be borne out in a study, suggesting that higher quality initial hypotheses are crucial. In this work, we propose a high-throughput pipeline to produce a ranked list of high-quality hypothesized marker laboratory tests for diagnoses. Our pipeline generates a large number of candidate lab-diagnosis hypotheses derived from machine learning models, filters and ranks them according to their potential novelty using text mining, and corroborate final hypotheses with logistic regression analysis. We test our approach on a large electronic health record dataset and the PubMed corpus, and find several promising candidate hypotheses.
Summary Doctors do not know whether treatment of high parathyroid hormone levels is linked to better outcomes in their patients with kidney disease. In this study, lower parathyroid hormone levels at baseline were linked to lower risk of fracture, vascular events, and death in people with kidney disease. Purpose Chronic kidney disease (CKD) affects ~ 20% of older adults, and secondary hyperparathyroidism (HPT) is a common condition in these patients. To what degree HPT predicts fractures, vascular events, and mortality in pre-dialysis CKD patients is debated. In stage 3 and 4 CKD patients, we assessed relationships between baseline serum PTH levels and subsequent 10-year probabilities of clinical fractures, vascular events, and death. Methods We used Marshfield Clinic Health System electronic health records to analyze data from adult CKD patients receiving care between 1985 and 2013, and whose PTH was measured using a second-generation assay. Covariates included PTH, age, gender, tobacco use, vascular disease, diabetes, hypertension, hyperlipidemia, obesity, GFR, and use of osteoporosis medications. Results Five thousand one hundred eight subjects had a mean age of 68 ± 17 years, 48% were men, and mean follow-up was 23 ± 10 years. Fractures, vascular events, and death occurred in 18%, 71%, and 56% of the cohort, respectively. In univariate and multivariate models, PTH was an independent predictor of fracture, vascular events, and death. The hazards of fracture, vascular events and death were minimized at a baseline PTH of 0, 69, and 58 pg/mL, respectively. Conclusions We found that among individuals with stage 3 and 4 CKD, PTH was an independent predictor of fractures, vascular events, and death. Additional epidemiologic studies are needed to confirm these findings. If a target PTH range can be confirmed, then randomized placebo-controlled trials will be needed to confirm that treating HPT reduces the risk of fracture, vascular events, and death.
To better understand the real-world effects of pharmacogenomic (PGx) alerts, this study aimed to characterize alert design within the eMERGE Network, and to establish a method for sharing PGx alert response data for aggregate analysis. Seven eMERGE sites submitted design details and established an alert logging data dictionary. Six sites participated in a pilot study, sharing alert response data from their electronic health record systems. PGx alert design varied, with some consensus around the use of active, post-test alerts to convey Clinical Pharmacogenetics Implementation Consortium recommendations. Sites successfully shared response data, with wide variation in acceptance and follow rates. Results reflect the lack of standardization in PGx alert design. Standards and/or larger studies will be necessary to fully understand PGx impact. This study demonstrated a method for sharing PGx alert response data and established that variation in system design is a significant barrier for multi-site analyses.
PURPOSE As a preliminary evaluation of the outcomes of implementing pharmacogenetic testing within a large rural healthcare system, patients who received pre-emptive pharmacogenetic testing and warfarin dosing were monitored until June 2017. SUMMARY Over a 20-month period, 749 patients were genotyped for VKORC1 and CYP2C9 as part of the electronic Medical Records and Genomics Pharmacogenetics (eMERGE PGx) study. Of these, 27 were prescribed warfarin and received an alert for pharmacogenetic testing pertinent to warfarin; 20 patients achieved their target international normalized ratio (INR) of 2.0-3.0, and 65% of these patients achieved target dosing within the recommended pharmacogenetic alert dose (± 0.5 mg/day). Of these, 10 patients had never been on warfarin prior to the alert and were further evaluated with regard to time to first stable target INR, bleeds and thromboembolic events, hospitalizations, and mortality. There was a general trend of faster time to first stable target INR when the patient was initiated at a warfarin dose within the alert recommendation versus a dose outside of the alert recommendation with a mean (± SD) of 34 (± 28) days versus 129 (± 117) days, respectively. No trends regarding bleeds, thromboembolic events, hospitalization, or mortality were identified with respect to the pharmacogenetic alert. The pharmacogenetic alert provided pharmacogenetic dosing information to prescribing clinicians and appeared to deploy appropriately with the correct recommendation based upon patient genotype. CONCLUSION Implementing pharmacogenetic testing as a standard of care service in anticoagulation monitoring programs may improve dosage regimens for patients on anticoagulation therapy.
Background: Implementing clinical phenotypes across a network is labor intensive and potentially error prone. Use of a common data model may facilitate the process. Methods: Electronic Medical Records and Genomics (eMERGE) sites implemented the Observational Health Data Sciences and Informatics (OHDSI) Observational Medical Outcomes Partnership (OMOP) Common Data Model across their electronic health record (EHR)-linked DNA biobanks. Two previously implemented eMERGE phenotypes were converted to OMOP and implemented across the network. Results: It was feasible to implement the common data model across sites, with laboratory data producing the greatest challenge due to local encoding. Sites were then able to execute the OMOP phenotype in less than one day, as opposed to weeks of effort to manually implement an eMERGE phenotype in their bespoke research EHR databases. Of the sites that could compare the current OMOP phenotype implementation with the original eMERGE phenotype implementation, specific agreement ranged from 100% to 43%, with disagreements due to the original phenotype, the OMOP phenotype, changes in data, and issues in the databases. Using the OMOP query as a standard comparison revealed differences in the original implementations despite starting from the same definitions, code lists, flowcharts, and pseudocode. Conclusion: Using a common data model can dramatically speed phenotype implementation at the cost of having to populate that data model, though this will produce a net benefit as the number of phenotype implementations increases. Inconsistencies among the implementations of the original queries point to a potential benefit of using a common data model so that actual phenotype code and logic can be shared, mitigating human error in reinterpretation of a narrative phenotype definition.