Background: Renal glucosuria is a rare inheritable trait caused by loss-of-function variants in the gene that encodes SGLT2 (i.e., SLC5A2). The genetics of renal glucosuria is poorly understood and even less is known on how loss-of-function variants in SLC5A2 may affect response to SGLT2 inhibitors, a new class of medication gaining popularity to treat diabetes by artificially inducing glucosuria. Methods: We used two biobanks that link genomic with electronic health record data to study the genetics of renal glucosuria. This included 245,394 participants enrolled in the All of Us (AoU) Research Program and 11,011 enrolled in Marshfield Clinic’s Personalized Research Project (PMRP). Association studies in AoU and PMRP identified 10 variants that reached an experiment-wise Bonferroni threshold in either cohort, nine were novel. PMRP was further used as a recruitment source for a prospective SGLT2 pharmacogenetic trial. During a glucose tolerance test, the trial measured urine glucose concentrations in 15 SLC5A2 variant-positive individuals and 15 matched wild types with and without an SGLT2 inhibitor. Results: This trial demonstrated that carriers of SLC5A2 risk variants may be more sensitive to SGLT2 inhibitors compared to wild types (P=0.075). Based on population data, 2% of an ethnically diverse population carry rare variants in SLC5A2 and are at risk for renal glucosuria. Conclusions: As a result, 2% of individuals being treated with SGLT2 inhibitors may respond differently to this new class of medication compared to the general population suggesting a larger investigation into SLC5A2 variants and SGLT2 inhibitors is needed.
Abstract Background: Multiple myeloma (MM) is the second most common hematologic malignancy in the United States, with over 30,000 new cases diagnosed annually. The development of novel therapies and immunotherapies has improved the 5-year relative survival rate from 54% in 2020 to 62% in 2025. However, relapsed and refractory MM (RRMM) continues to be a major barrier to further survival gains. Personalized treatment strategies that integrate multi-omics profilings with drug sensitivity testing offer promise for improving RRMM outcomes. Yet, implementation of such approaches remains limited, particularly in community hospitals serving underrepresented and rural populations. A key obstacle is the lack of efficient ex vivo drug screening platforms for identifying effective therapies tailored to individual patients. Existing ex vivo MM culture systems are complex, costly, and labor-intensive. They often rely on co-culture with mixed cell types using expensive devices, supplemented with cytokines and antibodies over 5–7 days, introducing bias and limiting scalability. Thus, we established a simplified ex vivo protocol for drug sensitivity testing that can be feasibly implemented in community hospital settings. Methods: Our two versions of simplified ex vivo protocol were published with detailed instructions (PMID: 40654055 and 38977990). Briefly, CD138⁺ plasma cells were enriched from patient bone marrow samples and cultured for 18 hours in RPMI-1640 supplemented with 20% heat-inactivated autologous serum, maintaining approximately 80% cell viability. The protocol does not require high-end instrumentation or PhD-level expertise. This streamlined approach enables drug sensitivity testing in a 96-well plate format using primary patient cells. Aims: This study aims to adapt the existing 1-day ex vivo drug sensitivity testing protocol from a 96-well plate format to 384-well and/or 1,536-well plate formats. Additionally, the study seeks to apply the protocol to assess the efficacy of current treatment regimens in primary multiple myeloma samples. Results: We successfully adapted the protocol for use in a 384-well plate format by seeding 50 µL of CD138⁺ cells at 0.3 × 10⁶/mL per well. Ongoing efforts are focused on enhancing cell viability by fine tuning of these conditions and integrating robotic process automation using cost-effective BioTek instrumentation. Using this protocol, we also evaluated residual MM cells from a patient undergoing treatment. < 1% cell viability was detected after 18 hours of culture, and a follow-up bone marrow biopsy confirmed a reduction of MM cells to undetectable level. We are currently collecting additional residual samples to assess whether this approach can reliably measure treatment effectiveness. Implications for Clinical Practice: Adapting the protocol to 384- and 1,536-well plate formats enables drug sensitivity testing of over 300 FDA-approved anti-tumor agents at multiple concentrations, considering that the myeloma cell number in patient samples is limited. Consistent results across multiple specimens may support the protocol's use for evaluating therapeutic effectiveness alongside minimal residual disease (MRD) assays. Significance & Conclusion: We are establishing a practical, accessible protocol for ex vivo drug sensitivity testing and treatment evaluation that can be implemented in community hospital settings, particularly in rural and under-represented areas. This approach has the potential to improve clinical decision-making and advance personalized myeloma care in underserved populations.
Background: Although inhaled corticosteroids (ICS) are the first-line therapy for patients with persistent asthma, many patients continue to have exacerbations. We developed machine learning models to predict the ICS response in patients with asthma. Methods: The subjects included asthma patients of European ancestry (n = 1371; 448 children; 916 adults). A genome-wide association study was performed to identify the SNPs associated with ICS response. Using the SNPs identified, two machine learning models were developed to predict ICS response: (1) least absolute shrinkage and selection operator (LASSO) regression and (2) random forest. Results: The LASSO regression model achieved an AUC of 0.71 (95% CI 0.67–0.76; sensitivity: 0.57; specificity: 0.75) in an independent test cohort, and the random forest model achieved an AUC of 0.74 (95% CI 0.70–0.78; sensitivity: 0.70; specificity: 0.68). The genes contributing to the prediction of ICS response included those associated with ICS responses in asthma (TPSAB1, FBXL16), asthma symptoms and severity (ABCA7, CNN2, PTRN3, and BSG/CD147), airway remodeling (ELANE, FSTL3), mucin production (GAL3ST), leukotriene synthesis (GPX4), allergic asthma (ZFPM1, SBNO2), and others. Conclusions: An accurate risk prediction of ICS response can be obtained using machine learning methods, with the potential to inform personalized treatment decisions. Further studies are needed to examine if the integration of richer phenotype data could improve risk prediction.
Objective: To determine if host genetics may be a risk factor for severe blastomycosis. Design: A cohort of patients who had contracted blastomycosis underwent targeted SNP (single nucleotide polymorphism) genotyping. The genetics of these patients were compared to a set of age and gender-matched controls and between patients with severe versus mild to moderate blastomycosis. Participants: Patients with a diagnosis of blastomycosis prior to 2017 were contacted for enrollment in this study. A phone hotline was also set up to allow interested participants from outside the Marshfield Clinic Health System to request enrollment. Methods: SNP frequency was assessed for significant differences between the patient cohort and controls and between patients with severe versus mild to moderate blastomycosis. We also tested the effect of Blastomyces species identified in clinical isolates on disease symptoms and severity. Results: No significant differences were found in SNP frequency between cases and controls or between those with severe or mild to moderate blastomycosis. We did detect significant differences in symptom frequency and disease severity by Blastomyces species. Conclusions: Our study did not identify any genetic risk factors for blastomycosis. Instead, the species of Blastomyces causing the infection had a significant effect on disease severity.
Globally, half a billion people are employed in animal agriculture and are directly exposed to the associated microorganisms. However, the extent to which such exposures affect resident human microbiomes is unclear. Here we conducted a longitudinal profiling of the nasal and faecal microbiomes of 66 dairy farmers and 166 dairy cows over a year-long period. We compare farmer microbiomes to those of 60 age-, sex- and ZIP code-matched people with no occupational exposures to farm animals (non-farmers). We show that farming is associated with microbiomes containing livestock-associated microbes; this is most apparent in the nasal bacterial community, with farmers harbouring a richer and more diverse nasal community than non-farmers. Similarly, in the gut microbial communities, we identify more shared microbial lineages between cows and farmers from the same farms. Additionally, we find that shared microbes are associated with antibiotic resistance genes. Overall, our study demonstrates the interconnectedness of human and animal microbiomes. Longitudinal profiling of the nasal and faecal microbiomes of 66 dairy farmers and 166 dairy cows over a year-long period shows that microbes acquired from cow microbiomes introduce clinically relevant antimicrobial resistance genes to farmer guts.
Multiple sclerosis (MS) is a complex autoimmune disease in which both the roles of genetic susceptibility and environmental/microbial factors have been investigated. More than 200 genetic susceptibility variants have been identified along with the dysbiosis of gut microbiota, both independently have been shown to be associated with MS. We hypothesize that MS patients harboring genetic susceptibility variants along with gut microbiome dysbiosis are at a greater risk of exhibiting the disease. We investigated the genetic risk score for MS in conjunction with gut microbiota in the same cohort of 117 relapsing remitting MS (RRMS) and 26 healthy controls. DNA samples were genotyped using Illumina’s Infinium Immuno array-24 v2 chip followed by calculating genetic risk score and the microbiota was determined by sequencing the V4 hypervariable region of the 16S rRNA gene. We identified two clusters of MS patients, Cluster A and B, both having a higher genetic risk score than the control group. However, the MS cases in cluster B not only had a higher genetic risk score but also showed a distinct gut microbiome than that of cluster A. Interestingly, cluster A which included both healthy control and MS cases had similar gut microbiome composition. This could be due to (i) the non-active state of the disease in that group of MS patients at the time of fecal sample collection and/or (ii) the restoration of the gut microbiome post disease modifying therapy to treat the MS. Our study showed that there seems to be an association between genetic risk score and gut microbiome dysbiosis in triggering the disease in a small cohort of MS patients. The MS Cluster A who have a higher genetic risk score but microbiome profile similar to that of healthy controls could be due to the remitting phase of the disease or due to the effect of disease modifying therapies.
It is well known that common variants in specific genes influence drug metabolism and response, but it is currently unknown what fraction of patients are given prescriptions over a lifetime that could be contraindicated by their pharmacogenomic profiles. To determine the clinical utility of pharmacogenomics over a lifetime in a general patient population, we sequenced the genomes of 300 deceased Marshfield Clinic patients linked to lifelong medical records. Genetic variants in 33 pharmacogenes were evaluated for their lifetime impact on drug prescribing using extensive electronic health records. Results show that 93% of the 300 deceased patients carried clinically relevant variants. Nearly 80% were prescribed approximately three medications on average that may have been impacted by these variants. Longitudinal data suggested that the optimal age for pharmacogenomic testing was prior to age 50, but the optimal age is greatly influenced by the stability of the population in the healthcare system. This study emphasizes the broad clinical impact of pharmacogenomic testing over a lifetime and demonstrates the potential application of genomic medicine in a general patient population for the advancement of precision medicine.
Background: Despite improved 5-year survival in multiple myeloma (MM), relapsed and/or refractory multiple myeloma (RRMM) remains a big challenge. Forkhead box transcription factor FOXM1 is a key regulator of metabolism and cell cycle progression in RRMM. FOXM1 is highly expressed in OPM2 and Delta47 cells compared to 9 other myeloma cell lines. Inhibiting FOXM1 function by deleting the FOXM1 gene or using the small-molecule FOXM1 inhibitor, NB73, suppresses OPM2 and Delta47 cells in vitro and in vivo. We hypothesized that FOXM1-targeted inhibition of RRMM may be deepened by combining NB73 with established or candidate myeloma drugs. Materials and methods: We used 3 tools to develop FOXM1-targeted combinatorial therapies by: (1) combining NB73 with 6 widely used myeloma drugs; (2) repurposing screens of cancer drugs in FOXM1-knockout (FOXM1 KO) vs FOXM1-proficient (FOXM1 WT) OPM2 cells; and (3) combining NB73 with compounds recommended by CMap analysisof FOXM1 KO vs FOXM1 WTcells. RNA seq analysis was used to elucidate pathways underlying drug synergy. Results: We adopted ZIP drug synergy scoring assay to examine the interactions between NB73 and the 6 myeloma drugs in MM cells. ZIP score takes the merits of both Loewe score that appraises drugs targeting the same pathway and Bliss score that appraises drugs targeting different pathways. Venetoclax, a BCL2 inhibitor treating MM with t(11;14) translocation, synergized with NB73 in killing OPM2 and Delta47 cells. In contrast, taking advantage of FOXM1 WT and FOXM1 KO MM cells, we conducted repurposing screens of FDA-approved anti-cancer drug library (NCI AOD X) in these cells. We identified Dasatinib and Panobinostat that killed FOXM1 KO OPM2 cells much more than FOXM1 WT cells. We also identified 3 drugs from the repurposing screens in FOXM1 WT and FOXM1 KO Delta47 cells. Thirdly, CMap, a large-scale compendium of functional perturbations in cultured human cells coupled to a gene expression read-out, implicated a synergy between NB73 and Thapsigargin that was verified by ZIP drug synergy scoring assay. Collectively, we have identified a total of 7 drugs (6 FDA-approved) synergizing with NB73 in killing OPM2 and/or Delta47 cells in which FOXM1 is highly expressed. Since Venetoclax is a current myeloma drug, we have focused on how inhibiting FOXM1 sensitizes MM cells to Venetoclax. Firstly, we showed the synergy between NB73 and Venetoclax in inducing apoptosis in OPM2 and Delta47 cells. We used the experimental condition in apoptosis assay to treat OPM2 cells for RNA sequencing study. Principal Component Analysis of RNA seq data indicated that the clusters of NB73 and the combo were well separated from each other and from the clusters of DMSO and Venetoclax. Gene Set Enrichment Analysis of these 4 groups presented many changes in tumor-associated pathways. Pathways regulated by a single drug ( e.g. MYC, p53, IFN-α) were further aggravated by the combo. Pathways ( e.g. Apoptosis) were not enriched by NB73 or Venetoclax, but enriched by the combo. Hyperactivation of MYC pathway is prevalent in MM. The combo significantly repressed MYC, PLK1, CCNA2 and CDC20 in MYC pathway through several mechanisms. For example, both NB73 and Venetoclax decreased MYC RNA, and the combo further lowered MYC, showing a “sum-up” effect. In contrast, the single drug lightly increased CDC20, but the combo significantly decreased it, presenting a “ de novo” effect. Since these 4 genes are important tumor promoters and inhibiting any one markedly represses tumor growth, we are addressing how the combo inhibits them. Ongoing experiments: We transplanted OPM2 cells into NSG mice and established the MM xenograft for in vivo test. NB73 is being given via sub-Q and Venetoclax is being given via oral gavage (n=5/group, 4 groups). The tumor burden will be monitored by live imaging every 2 weeks because OPM2 cells express Luciferase and the survival curve will be drawn. We are recruiting MM patients for bone marrow cells in order to develop a simplified protocol of culturing MM cells ex vivo that will be used to test new drug combinations targeting FOXM1 or other molecular vulnerabilities in MM. Conclusion: We identified FOXM1-targeted therapies consisting of novel FOXM1 inhibitor NB73 and FDA-approved drugs in vitro. We are addressing how the combo represses MYC pathway to sensitize MM cells to BCL2 inhibitor Venetoclax.
Background Multiple sclerosis (MS) is a complex autoimmune disease in which both the roles of genetic susceptibility and environmental/microbial factors have been investigated. More than 200 genetic susceptibility variants have been identified along with the dysbiosis of gut microbiota, both independently have been shown to be associated with MS. We hypothesize that MS patients harboring genetic susceptibility variants along with gut microbiome dysbiosis are at a greater risk of exhibiting the disease. We investigated the polygenic risk score for MS in conjunction with gut microbiota in the same cohort of 117 relapsing remitting MS (RRMS) and 26 healthy controls. DNA samples were genotyped using Illumina’s Infinium Immuno array-24 v2 chip followed by calculating polygenic risk score and the microbiota was determined by sequencing the V4 hypervariable region of the 16S rRNA gene. Results We identified two clusters of MS patients, Cluster A and B both having a higher polygenic risk score than the control group. The Cluster B with the higher polygenic risk score had a distinct gut microbiota, different than the Cluster A. MS group whose microbiome was similar to that of the control group despite a higher genetic risk score than the control group. This could be due to i) the non-active state of the disease in that group of MS patients at the time of fecal sample collection and/or ii) the restoration of the gut microbiome post disease modifying therapy to treat the MS. Conclusion Our study showed that there seems to be association between polygenic risk score and gut microbiome dysbiosis in triggering the disease in a small cohort of MS patients. The MS Cluster A who have a higher polygenic risk score but microbiome profile similar to that of healthy controls could be due to the remitting phase of the disease or due to the effect of disease modifying therapies.
The Electronic Medical Records and Genomics (eMERGE) Network, established in 2007, is a consortium of academic and integrated health systems conducting discovery and implementation research in translational genomics. Here, we outline the history of the network, highlight major impacts and lessons learned, and present the tools and resources developed for large-scale genomic analyses and translation into a clinical setting. The network developed methods to extract phenotypes from the electronic medical record to perform genome-wide and phenome-wide association studies. Recruited cohorts were clinically sequenced off a custom panel for targeted sequencing of variants and monogenic disease risks and returned to participants to investigate the impact of return of genomic results. After generating a 105,000 participant-imputed genome-wide association study (GWAS) dataset for discovery, the network enrolled and sequenced 24,998 participants. Integration of these results into the medical record and the effects of results on participants provided key lessons to the field. These learned lessons inform genetic research in diverse populations and provide insights into the clinical impact of return and implementation of genomic medicine using the electronic medical record. The lessons produced by the eMERGE Network can be utilized by other consortia as translational genomic medicine research evolves.
Andrea R. Waksmunski, Michelle Grunin, Tyler G. Kinzy, Robert P. Igo Jr, Jonathan L. Haines, and Jessica N. Cooke Bailey; for the International Age-Related Macular Degeneration Genomics Consortium Department of Genetics and Genome Sciences, Case Western Reserve University, Cleveland, Ohio, United States Cleveland Institute for Computational Biology, Case Western Reserve University, Cleveland, Ohio, United States Department of Population and Quantitative Health Sciences, Case Western Reserve University, Cleveland, Ohio, United States
Objective:We describe a stratified sampling design that combines electronic health records (EHRs) and United States Census (USC) data to construct the sampling frame and an algorithm to enrich the sample with individuals belonging to rarer strata. Materials and Methods:This design was developed for a multi-site survey that sought to examine patient concerns about and barriers to participating in research studies, especially among under-studied populations (eg, minorities, low educational attainment). We defined sampling strata by cross-tabulating several socio-demographic variables obtained from EHR and augmented with census-block-level USC data. We oversampled rarer and historically underrepresented subpopulations. Results:The sampling strategy, which included USC-supplemented EHR data, led to a far more diverse sample than would have been expected under random sampling (eg, 3-, 8-, 7-, and 12-fold increase in African Americans, Asians, Hispanics and those with less than a high school degree, respectively). We observed that our EHR data tended to misclassify minority races more often than majority races, and that non-majority races, Latino ethnicity, younger adult age, lower education, and urban/suburban living were each associated with lower response rates to the mailed surveys. Discussion:We observed substantial enrichment from rarer subpopulations. The magnitude of the enrichment depends on the accuracy of the variables that define the sampling strata and the overall response rate. Conclusion:EHR and USC data may be used to define sampling strata that in turn may be used to enrich the final study sample. This design may be of particular interest for studies of rarer and understudied populations.
BACKGROUND:We characterised the phenotypic consequence of genetic variation at the PCSK9 locus and compared findings with recent trials of pharmacological inhibitors of PCSK9. METHODS:Published and individual participant level data (300,000+ participants) were combined to construct a weighted PCSK9 gene-centric score (GS). Seventeen randomized placebo controlled PCSK9 inhibitor trials were included, providing data on 79,578 participants. Results were scaled to a one mmol/L lower LDL-C concentration. RESULTS:The PCSK9 GS (comprising 4 SNPs) associations with plasma lipid and apolipoprotein levels were consistent in direction with treatment effects. The GS odds ratio (OR) for myocardial infarction (MI) was 0.53 (95% CI 0.42; 0.68), compared to a PCSK9 inhibitor effect of 0.90 (95% CI 0.86; 0.93). For ischemic stroke ORs were 0.84 (95% CI 0.57; 1.22) for the GS, compared to 0.85 (95% CI 0.78; 0.93) in the drug trials. ORs with type 2 diabetes mellitus (T2DM) were 1.29 (95% CI 1.11; 1.50) for the GS, as compared to 1.00 (95% CI 0.96; 1.04) for incident T2DM in PCSK9 inhibitor trials. No genetic associations were observed for cancer, heart failure, atrial fibrillation, chronic obstructive pulmonary disease, or Alzheimer's disease - outcomes for which large-scale trial data were unavailable. CONCLUSIONS:Genetic variation at the PCSK9 locus recapitulates the effects of therapeutic inhibition of PCSK9 on major blood lipid fractions and MI. While indicating an increased risk of T2DM, no other possible safety concerns were shown; although precision was moderate.
Standard analyses applied to genome-wide association data are well designed to detect additive effects of moderate strength. However, the power for standard genome-wide association study (GWAS) analyses to identify effects from recessive diplotypes is not typically high. We proposed and conducted a gene-based compound heterozygosity test to reveal additional genes underlying complex diseases. With this approach applied to iron overload, a strong association signal was identified between the fibroblast growth factor-encoding gene, FGF6, and hemochromatosis in the central Wisconsin population. Functional validation showed that fibroblast growth factor 6 protein (FGF-6) regulates iron homeostasis and induces transcriptional regulation of hepcidin. Moreover, specific identified FGF6 variants differentially impact iron metabolism. In addition, FGF6 downregulation correlated with iron-metabolism dysfunction in systemic sclerosis and cancer cells. Using the recessive diplotype approach revealed a novel susceptibility hemochromatosis gene and has extended our understanding of the mechanisms involved in iron metabolism.
Background The eMERGE III Network was tasked with harmonizing genetic testing protocols linking multiple sites and investigators.Methods DNA capture panels targeting 109 genes and 1551 variants were constructed by two clinical sequencing centers for analysis of 25,000 participant DNA samples collected at 11 sites where samples were linked to patients with electronic health records. Each step from sample collection, data generation, interpretation, reporting, delivery and storage, were developed and validated in CAP/CLIA settings and harmonized across sequencing centers.Results A compliant and secure network was built and enabled ongoing review and reconciliation of clinical interpretations while maintaining communication and data sharing between investigators. Mechanisms for sustained propagation and growth of the network were established. An interim data freeze representing 15,574 sequenced subjects, informed the assay performance for a range of variant types, the rate of return of results for different phenotypes and the frequency of secondary findings. Practical obstacles for implementation and scaling of clinical and research findings were identified and addressed. The eMERGE protocols and tools established are now available for widespread dissemination.Conclusions This study established processes for different sequencing sites to harmonize the technical and interpretive aspects of sequencing tests, a critical achievement towards global standardization of genomic testing. The network established experience in the return of results and the rate of secondary findings across diverse biobank populations. Furthermore, the eMERGE network has accomplished integration of structured genomic results into multiple electronic health record systems, setting the stage for clinical decision support to enable genomic medicine.
Translational multidisciplinary research is important for the Center for Devices and Radiological Health's efforts for utilizing real‐world data (RWD) to enhance predictive evaluation of medical device performance in patient subpopulations. As part of our efforts for developing new RWD‐based evidentiary approaches, including in silico discovery of device‐related risk predictors and biomarkers, this study aims to characterize the sex/race‐related trends in hip replacement outcomes and identify corresponding candidate single nucleotide polymorphisms (SNPs). Adverse outcomes were assessed by deriving RWD from a retrospective analysis of hip replacement hospital discharge data from the National Inpatient Sample (NIS). Candidate SNPs were explored using pre‐existing data from the Personalized Medicine Research Project (PMRP). High‐Performance Integrated Virtual Environment was used for analyzing and visualizing putative associations between SNPs and adverse outcomes. Ingenuity Pathway Analysis (IPA) was used for exploring plausibility of the sex‐related candidate SNPs and characterizing gene networks associated with the variants of interest. The NIS‐based epidemiologic evidence showed that periprosthetic osteolysis (PO) was most prevalent among white men. The PMRP‐based genetic evidence associated the PO‐related male predominance with rs7121 (odds ratio = 4.89; 95% confidence interval = 1.41−17.05) and other candidate SNPs. SNP‐based IPA analysis of the expected gene expression alterations and corresponding signaling pathways suggested possible role of sex‐related metabolic factors in development of PO, which was substantiated by ad hoc epidemiologic analysis identifying the sex‐related differences in metabolic comorbidities in men vs. women with hip replacement‐related PO. Thus, our in silico study illustrates RWD‐based evidentiary approaches that may facilitate cost/time‐efficient discovery of biomarkers for informing use of medical products.
BACKGROUND:Proteomic approaches allow measurement of thousands of proteins in a single specimen, which can accelerate biomarker discovery. However, applying these technologies to massive biobanks is not currently feasible because of the practical barriers and costs of implementing such assays at scale. To overcome these challenges, we used a "virtual proteomic" approach, linking genetically predicted protein levels to clinical diagnoses in >40 000 individuals. METHODS:We used genome-wide association data from the Framingham Heart Study (n=759) to construct genetic predictors for 1129 plasma protein levels. We validated the genetic predictors for 268 proteins and used them to compute predicted protein levels in 41 288 genotyped individuals in the Electronic Medical Records and Genomics (eMERGE) cohort. We tested associations for each predicted protein with 1128 clinical phenotypes. Lead associations were validated with directly measured protein levels and either low-density lipoprotein cholesterol or subclinical atherosclerosis in the MDCS (Malmö Diet and Cancer Study; n=651). RESULTS:In the virtual proteomic analysis in eMERGE, 55 proteins were associated with 89 distinct diagnoses at a false discovery rate q<0.1. Among these, 13 associations involved lipid (n=7) or atherosclerosis (n=6) phenotypes. We tested each association for validation in MDCS using directly measured protein levels. At Bonferroni-adjusted significance thresholds, levels of apolipoprotein E isoforms were associated with hyperlipidemia, and circulating C-type lectin domain family 1 member B and platelet-derived growth factor receptor-β predicted subclinical atherosclerosis. Odds ratios for carotid atherosclerosis were 1.31 (95% CI, 1.08-1.58; P=0.006) per 1-SD increment in C-type lectin domain family 1 member B and 0.79 (0.66-0.94; P=0.008) per 1-SD increment in platelet-derived growth factor receptor-β. CONCLUSIONS:We demonstrate a biomarker discovery paradigm to identify candidate biomarkers of cardiovascular and other diseases.
Individuals participating in biobanks and other large research projects are increasingly asked to provide broad consent for open-ended research use and widespread sharing of their biosamples and data. We assessed willingness to participate in a biobank using different consent and data sharing models, hypothesizing that willingness would be higher under more restrictive scenarios. Perceived benefits, concerns, and information needs were also assessed. In this experimental survey, individuals from 11 US healthcare systems in the Electronic Medical Records and Genomics (eMERGE) Network were randomly allocated to one of three hypothetical scenarios: tiered consent and controlled data sharing; broad consent and controlled data sharing; or broad consent and open data sharing. Of 82,328 eligible individuals, exactly 13,000 (15.8%) completed the survey. Overall, 66% (95% CI: 63%-69%) of population-weighted respondents stated they would be willing to participate in a biobank; willingness and attitudes did not differ between respondents in the three scenarios. Willingness to participate was associated with self-identified white race, higher educational attainment, lower religiosity, perceiving more research benefits, fewer concerns, and fewer information needs. Most (86%, CI: 84%-87%) participants would want to know what would happen if a researcher misused their health information; fewer (51%, CI: 47%-55%) would worry about their privacy. The concern that the use of broad consent and open data sharing could adversely affect participant recruitment is not supported by these findings. Addressing potential participants' concerns and information needs and building trust and relationships with communities may increase acceptance of broad consent and wide data sharing in biobank research.
Motivation Pedigree analysis is a longstanding and powerful approach to gain insight into the underlying genetic factors in human health, but identifying, recruiting and genotyping families can be difficult, time consuming and costly. Development of high throughput methods to identify families and foster downstream analyses are necessary. Results This paper describes simple methods that allowed us to identify 173 368 family pedigrees with high probability using basic demographic data available in most electronic health records (EHRs). We further developed and validate a novel statistical method that uses EHR data to identify families more likely to have a major genetic component to their diseases risk. Lastly, we showed that incorporating EHR-linked family data into genetic association testing may provide added power for genetic mapping without additional recruitment or genotyping. The totality of these results suggests that EHR-linked families can enable classical genetic analyses in a high-throughput manner. Availability and implementation Pseudocode is provided as supplementary information. Contact HEBBRING.SCOTT@marshfieldresearch.org. Supplementary information Supplementary data are available at Bioinformatics online.