Electronic health records, biobanks, and wearable biosensors contain multiple high-dimensional clinical data (HDCD) modalities (e.g., ECG, Photoplethysmography (PPG), and MRI) for each individual. Access to multimodal HDCD provides a unique opportunity for genetic studies of complex traits because different modalities relevant to a single physiological system (e.g., circulatory system) encode complementary and overlapping information. We propose a novel multimodal deep learning method, M-REGLE, for discovering genetic associations from a joint representation of multiple complementary HDCD modalities. We showcase the effectiveness of this model by applying it to several cardiovascular modalities. M-REGLE jointly learns a lower representation (i.e., latent factors) of multimodal HDCD using a convolutional variational autoencoder, performs genome wide association studies (GWAS) on each latent factor, then combines the results to study the genetics of the underlying system. To validate the advantages of M-REGLE and multimodal learning, we apply it to common cardiovascular modalities (PPG and ECG), and compare its results to unimodal learning methods in which representations are learned from each data modality separately, but the downstream genetic analyses are performed on the combined unimodal representations. M-REGLE identifies 19.3% more loci on the 12-lead ECG dataset, 13.0% more loci on the ECG lead I + PPG dataset, and its genetic risk score significantly outperforms the unimodal risk score at predicting cardiac phenotypes, such as atrial fibrillation (Afib), in multiple biobanks.
AbstractBackgroundDespite the growing interest in the use of human genomic data for drug target identification and validation, the extent to which the spectrum of human disease has been addressed by genome-wide association studies (GWAS), or by drug development, and the degree to which these efforts overlap remain unclear.MethodsIn this study we harmonize and integrate different data sources to create a sample space of all the human drug targets and diseases and identify points of convergence or divergence of GWAS and drug development efforts.ResultsWe show that only 612 of 11,158 diseases listed in Human Disease Ontology have an approved drug treatment in at least one region of the world. Of the 1414 diseases that are the subject of preclinical or clinical phase drug development, only 666 have been investigated in GWAS. Conversely, of the 1914 human diseases that have been the subject of GWAS, 1121 have yet to be investigated in drug development.ConclusionsWe produce target-disease indication lists to help the pharmaceutical industry to prioritize future drug development efforts based on genetic evidence, academia to prioritize future GWAS for diseases without effective treatments, and both sectors to harness genetic evidence to expand the indications for licensed drugs or to identify repurposing opportunities for clinical candidates that failed in their originally intended indication.
Metabolomic age models have been proposed for the study of biological aging, however, they have not been widely validated. We aimed to assess the performance of newly developed and existing nuclear magnetic resonance spectroscopy (NMR) metabolomic age models for prediction of chronological age (CA), mortality, and age-related disease. Ninety-eight metabolic variables were measured in blood from nine UK and Finnish cohort studies (N ≈31,000 individuals, age range 24-86 years). We used nonlinear and penalized regression to model CA and time to all-cause mortality. We examined associations of four new and two previously published metabolomic age models, with aging risk factors and phenotypes. Within the UK Biobank (N ≈102,000), we tested prediction of CA, incident disease (cardiovascular disease (CVD), type-2 diabetes mellitus, cancer, dementia, and chronic obstructive pulmonary disease), and all-cause mortality. Seven-fold cross-validated Pearson's r between metabolomic age models and CA ranged between 0.47 and 0.65 in the training cohort set (mean absolute error: 8-9 years). Metabolomic age models, adjusted for CA, were associated with C-reactive protein, and inversely associated with glomerular filtration rate. Positively associated risk factors included obesity, diabetes, smoking, and physical inactivity. In UK Biobank, correlations of metabolomic age with CA were modest (r = 0.29-0.33), yet all metabolomic model scores predicted mortality (hazard ratios of 1.01 to 1.06/metabolomic age year) and CVD, after adjustment for CA. While metabolomic age models were only moderately associated with CA in an independent population, they provided additional prediction of morbidity and mortality over CA itself, suggesting their wider applicability.
Background Ketone bodies (KBs) are an alternative energy supply for brain functions when glucose is limited. The most abundant ketone metabolite, 3-β-hydroxybutyrate (BOHBUT), has been suggested to prevent or delay cognitive impairment, but the evidence remains unclear. We triangulated observational and Mendelian randomization (MR) studies to investigate the association and causation between KBs and cognitive function. Methods In observational analyses of 5506 participants aged ≥ 45 years from the Whitehall II study, we used multiple linear regression to investigate the associations between categorized KBs and cognitive function scores. Two-sample MR was carried out using summary statistics from an in-house KBs meta-analysis between the University College London-London School of Hygiene and Tropical Medicine-Edinburgh-Bristol (UCLEB) Consortium and Kettunen et al. ( N = 45,031), and publicly available summary statistics of cognitive performance and Alzheimer’s disease (AD) from the Social Science Genetic Association Consortium ( N = 257,841), and the International Genomics of Alzheimer’s Project ( N = 54,162), respectively. Both strong ( P < 5 × 10 −8 ) and suggestive ( P < 1 × 10 −5 ) sets of instrumental variables for BOHBUT were applied. Finally, we performed cis -MR on OXCT1 , a well-known gene for KB catabolism. Results BOHBUT was positively associated with general cognitive function ( β = 0.26, P = 9.74 × 10 −3 ). In MR analyses, we observed a protective effect of BOHBUT on cognitive performance (inverse variance weighted: β IVW = 7.89 × 10 −2 , P IVW = 1.03 × 10 −2 ; weighted median: β W-Median = 8.65 × 10 −2 , P W-Median = 9.60 × 10 −3 ) and a protective effect on AD ( β IVW = − 0.31, odds ratio: OR = 0.74, P IVW = 3.06 × 10 −2 ). Cis -MR showed little evidence of therapeutic modulation of OXCT1 on cognitive impairment. Conclusions Triangulation of evidence suggests that BOHBUT has a beneficial effect on cognitive performance. Our findings raise the hypothesis that increased BOHBUT may improve general cognitive functions, delaying cognitive impairment and reducing the risk of AD.
Increased blood lipid levels are heritable risk factors of cardiovascular disease with varied prevalence worldwide owing to different dietary patterns and medication use1. Despite advances in prevention and treatment, in particular through reducing low-density lipoprotein cholesterol levels2, heart disease remains the leading cause of death worldwide3. Genome-wideassociation studies (GWAS) of blood lipid levels have led to important biological and clinical insights, as well as new drug targets, for cardiovascular disease. However, most previous GWAS4–23 have been conducted in European ancestry populations and may have missed genetic variants that contribute to lipid-level variation in other ancestry groups. These include differences in allele frequencies, effect sizes and linkage-disequilibrium patterns24. Here we conduct a multi-ancestry, genome-wide genetic discovery meta-analysis of lipid levels in approximately 1.65 million individuals, including 350,000 of non-European ancestries. We quantify the gain in studying non-European ancestries and provide evidence to support the expansion of recruitment of additional ancestries, even with relatively small sample sizes. We find that increasing diversity rather than studying additional individuals of European ancestry results in substantial improvements in fine-mapping functional variants and portability of polygenic prediction (evaluated in approximately 295,000 individuals from 7 ancestry groupings). Modest gains in the number of discovered loci and ancestry-specific variants were also achieved. As GWAS expand emphasis beyond the identification of genes and fundamental biology towards the use of genetic variants for preventive and precision medicine25, we anticipate that increased diversity of participants will lead to more accurate and equitable26 application of polygenic scores in clinical practice. A genome-wide association meta-analysis study of blood lipid levels in roughly 1.6 million individuals demonstrates the gain of power attained when diverse ancestries are included to improve fine-mapping and polygenic score generation, with gains in locus discovery related to sample size.
ABSTRACTCommon SNPs are predicted to collectively explain 40-50% of phenotypic variation in human height, but identifying the specific variants and associated regions requires huge sample sizes. Here we show, using GWAS data from 5.4 million individuals of diverse ancestries, that 12,111 independent SNPs that are significantly associated with height account for nearly all of the common SNP-based heritability. These SNPs are clustered within 7,209 non-overlapping genomic segments with a median size of ~90 kb, covering ~21% of the genome. The density of independent associations varies across the genome and the regions of elevated density are enriched for biologically relevant genes. In out-of-sample estimation and prediction, the 12,111 SNPs account for 40% of phenotypic variance in European ancestry populations but only ~10%-20% in other ancestries. Effect sizes, associated regions, and gene prioritization are similar across ancestries, indicating that reduced prediction accuracy is likely explained by linkage disequilibrium and allele frequency differences within associated regions. Finally, we show that the relevant biological pathways are detectable with smaller sample sizes than needed to implicate causal genes and variants. Overall, this study, the largest GWAS to date, provides an unprecedented saturated map of specific genomic regions containing the vast majority of common height-associated variants.
Background and Aims : The causal relevance of the cholesterol and triglyceride (TG) content of lipoproteins other than low-density lipoproteins (LDL-C) in coronary heart disease (CHD) is uncertain.Results: Univariable MR indicated a causal association with CHD of the TG content of six lipoprotein subfractions and the cholesterol content of 10 lipoprotein subfractions. In MVMR analysis, the TG content of four subfractions displayed associations with CHD independently of the cholesterol content in the same subfraction, while the cholesterol content of 10 subfractions displayed a causal association with CHD independent of the TG content of the corresponding subfractions. The cholesterol but not the triglyceride content in subfractions referred to as triglyceride-rich lipoproteins (TRL) displayed the largest association with CHD (MVMR odds ratio [OR] 2.73 to 14.31 per 1 SD increase in the subfraction), though the large effect estimates may reflect model instability.Conclusions: The cholesterol and TG content of certain lipoprotein subfractions other LDL may contribute to risk of CHD and may be relevant risk factors to target in drug development. Background and Aims : The causal relevance of the cholesterol and triglyceride (TG) content of lipoproteins other than low-density lipoproteins (LDL-C) in coronary heart disease (CHD) is uncertain. Results: Univariable MR indicated a causal association with CHD of the TG content of six lipoprotein subfractions and the cholesterol content of 10 lipoprotein subfractions. In MVMR analysis, the TG content of four subfractions displayed associations with CHD independently of the cholesterol content in the same subfraction, while the cholesterol content of 10 subfractions displayed a causal association with CHD independent of the TG content of the corresponding subfractions. The cholesterol but not the triglyceride content in subfractions referred to as triglyceride-rich lipoproteins (TRL) displayed the largest association with CHD (MVMR odds ratio [OR] 2.73 to 14.31 per 1 SD increase in the subfraction), though the large effect estimates may reflect model instability. Conclusions: The cholesterol and TG content of certain lipoprotein subfractions other LDL may contribute to risk of CHD and may be relevant risk factors to target in drug development.
A major challenge of genome-wide association studies (GWASs) is to translate phenotypic associations into biological insights. Here, we integrate a large GWAS on blood lipids involving 1.6 million individuals from five ancestries with a wide array of functional genomic datasets to discover regulatory mechanisms underlying lipid associations. We first prioritize lipid-associated genes with expression quantitative trait locus (eQTL) colocalizations and then add chromatin interaction data to narrow the search for functional genes. Polygenic enrichment analysis across 697 annotations from a host of tissues and cell types confirms the central role of the liver in lipid levels and highlights the selective enrichment of adipose-specific chromatin marks in high-density lipoprotein cholesterol and triglycerides. Overlapping transcription factor (TF) binding sites with lipid-associated loci identifies TFs relevant in lipid biology. In addition, we present an integrative framework to prioritize causal variants at GWAS loci, producing a comprehensive list of candidate causal genes and variants with multiple layers of functional evidence. We highlight two of the prioritized genes, CREBRF and RRBP1, which show convergent evidence across functional datasets supporting their roles in lipid biology.
Interleukin 6 (IL-6) is a multifunctional cytokine with both pro- and anti-inflammatory properties with a heritability estimate of up to 61%. The circulating levels of IL-6 in blood have been associated with an increased risk of complex disease pathogenesis. We conducted a two-staged, discovery and replication meta genome-wide association study (GWAS) of circulating serum IL-6 levels comprising up to 67 428 (ndiscovery = 52 654 and nreplication = 14 774) individuals of European ancestry. The inverse variance fixed effects based discovery meta-analysis, followed by replication led to the identification of two independent loci, IL1F10/IL1RN rs6734238 on chromosome (Chr) 2q14, (Pcombined = 1.8 × 10-11), HLA-DRB1/DRB5 rs660895 on Chr6p21 (Pcombined = 1.5 × 10-10) in the combined meta-analyses of all samples. We also replicated the IL6R rs4537545 locus on Chr1q21 (Pcombined = 1.2 × 10-122). Our study identifies novel loci for circulating IL-6 levels uncovering new immunological and inflammatory pathways that may influence IL-6 pathobiology.
Glycemic traits are used to diagnose and monitor type 2 diabetes and cardiometabolic health. To date, most genetic studies of glycemic traits have focused on individuals of European ancestry. Here we aggregated genome-wide association studies comprising up to 281,416 individuals without diabetes (30% non-European ancestry) for whom fasting glucose, 2-h glucose after an oral glucose challenge, glycated hemoglobin and fasting insulin data were available. Trans-ancestry and single-ancestry meta-analyses identified 242 loci (99 novel; P < 5 × 10−8), 80% of which had no significant evidence of between-ancestry heterogeneity. Analyses restricted to individuals of European ancestry with equivalent sample size would have led to 24 fewer new loci. Compared with single-ancestry analyses, equivalent-sized trans-ancestry fine-mapping reduced the number of estimated variants in 99% credible sets by a median of 37.5%. Genomic-feature, gene-expression and gene-set analyses revealed distinct biological signatures for each trait, highlighting different underlying biological pathways. Our results increase our understanding of diabetes pathophysiology by using trans-ancestry studies for improved power and resolution. A trans-ancestry meta-analysis of GWAS of glycemic traits in up to 281,416 individuals identifies 99 novel loci, of which one quarter was found due to the multi-ancestry approach, which also improves fine-mapping of credible variant sets.
Abstract Background Low socio-economic position (SEP) is a risk factor for multiple health outcomes, but its molecular imprints in the body remain unclear. Methods We examined SEP as a determinant of serum nuclear magnetic resonance metabolic profiles in ∼30 000 adults and 4000 children across 10 UK and Finnish cohort studies. Results In risk-factor-adjusted analysis of 233 metabolic measures, low educational attainment was associated with 37 measures including higher levels of triglycerides in small high-density lipoproteins (HDL) and lower levels of docosahexaenoic acid (DHA), omega-3 fatty acids, apolipoprotein A1, large and very large HDL particles (including levels of their respective lipid constituents) and cholesterol measures across different density lipoproteins. Among adults whose father worked in manual occupations, associations with apolipoprotein A1, large and very large HDL particles and HDL-2 cholesterol remained after adjustment for SEP in later life. Among manual workers, levels of glutamine were higher compared with non-manual workers. All three indicators of low SEP were associated with lower DHA, omega-3 fatty acids and HDL diameter. At all ages, children of manual workers had lower levels of DHA as a proportion of total fatty acids. Conclusions Our work indicates that social and economic factors have a measurable impact on human physiology. Lower SEP was independently associated with a generally unfavourable metabolic profile, consistent across ages and cohorts. The metabolites we found to be associated with SEP, including DHA, are known to predict cardiovascular disease and cognitive decline in later life and may contribute to health inequalities.
Drug target Mendelian randomization (MR) studies use DNA sequence variants in or near a gene encoding a drug target, that alter the target’s expression or function, as a tool to anticipate the effect of drug action on the same target. Here we apply MR to prioritize drug targets for their causal relevance for coronary heart disease (CHD). The targets are further prioritized using independent replication, co-localization, protein expression profiles and data from the British National Formulary and clinicaltrials.gov. Out of the 341 drug targets identified through their association with blood lipids (HDL-C, LDL-C and triglycerides), we robustly prioritize 30 targets that might elicit beneficial effects in the prevention or treatment of CHD, including NPC1L1 and PCSK9, the targets of drugs used in CHD prevention. We discuss how this approach can be generalized to other targets, disease biomarkers and endpoints to help prioritize and validate targets during the drug development process.
Data and data-driven technologies are playing an increasingly influential role in health care, helping to detect disease earlier, move care closer to home, encourage health-promoting behaviours, and improve the efficiency of service delivery. Although data-driven technologies have potential for good, they can also exacerbate existing health inequalities, which are deep-rooted and have been laid bare during the COVID-19 pandemic. In this Comment, we examine how structural inequalities, biases, and racism in society are easily encoded in datasets and in the application of data science, and how this practice can reinforce existing social injustices and health inequalities. Approaching the problem from the perspective of data scientists, we follow the stages in an analytical pipeline to consider how and where things can go wrong. We then outline the essential role of data scientists in tackling racism and discrimination. Structural racism—defined as the macrolevel systems, ideologies, and processes that interact with one another to produce cumulative and chronic adverse outcomes for people from ethnic minority groups—is deeply entrenched in our society. Despite being a relatively new field, data science has undoubtedly been shaped by these social forces. Indeed, the closely related discipline of statistics played a pivotal role in the development and justification of race science (ie, the claim that there is an evolutionary basis for inequalities in social outcomes between racial groups), which has been used to justify slavery, discrimination, and racist ideologies.1Saini A Superior: the return of race science. Beacon Press, Boston, MA2019Google Scholar Today, structural racism influences the data science workforce and the hierarchies within it, the datasets collected and who is represented within them, and the research questions pursued and prioritised. These factors mean that data science might not equitably benefit people from backgrounds that are underrepresented in the workforce and in the datasets. A heterogeneous digital workforce might be less prone to groupthink, can better understand and interpret real-world problems, such as the unmet needs of a wider range of stakeholders, and can help mitigate inherent bias stemming from technological and digital processes. Yet recent reports show that, across the USA, UK, and EU, there has been little progress in achieving better representation within the sector.2HarnhamGlobal data and analytics diversity report.https://www.harnham.com/harnham-data-analytics-diversity-report-2021Date accessed: January 25, 2021Google Scholar Along the data science analysis pipeline, there are common points where insights can be both affected by and result in racism, including in design, input, analysis, and application. For instance, the way in which analytical problems are framed and selected is influenced by many factors, including the availability of funding, and the interests and backgrounds of those planning and conducting the analyses. When we combine the historical context, which excluded ethnic minority individuals from scientific institutions and failed to recognise their contributions,3Ileka KM McCluney CL Robinson RAS White coats, Black scientists.https://hbr.org/2020/09/white-coats-black-scientistsDate: Sept 23, 2020Date accessed: November 3, 2020Google Scholar with the continuing underrepresentation of ethnic minority groups in technology and data science,2HarnhamGlobal data and analytics diversity report.https://www.harnham.com/harnham-data-analytics-diversity-report-2021Date accessed: January 25, 2021Google Scholar and the evidence that Black scientists are less likely to receive research and innovation funding than their White counterparts,4Hoppe TA Litovitz A Willis KA et al.Topic choice contributes to the lower rate of NIH awards to African-American/black scientists.Sci Adv. 2019; 5eaaw7238Crossref PubMed Scopus (141) Google Scholar it is unsurprising that a White and Western lens is pervasive in health data science. During the input stage, the data sources used for health research reflect the willingness and ability of individuals to provide data, as well as the priorities of those collecting and investing in the data. Therefore, data reflect inequalities and injustices in society. There are a variety of reasons for the lower participation rates of ethnic minority groups in research, ranging from mistrust and fear of the medical establishment, and stigma related to research participation, to exclusion by design.5George S Duran N Norris K A systematic review of barriers and facilitators to minority research participation among African Americans, Latinos, Asian Americans, and Pacific Islanders.Am J Public Health. 2014; 104: e16-e31Crossref PubMed Scopus (517) Google Scholar Where research uses routinely collected data such as electronic health records, recording of data on ethnicity is often poor or patchy. There are also examples of racial bias in treatment being encoded and fed into algorithms that determine who needs extra care, thereby placing Black people at an even greater disadvantage.6Obermeyer Z Powers B Vogeli C Mullainathan S Dissecting racial bias in an algorithm used to manage the health of populations.Science. 2019; 366: 447-453Crossref PubMed Scopus (467) Google Scholar Analytical decisions, such as how variables are defined, can also perpetuate racism and inequalities. Darshali Vyas and colleagues give examples from nine clinical specialties of race-adjusted algorithms that "risk baking inequity into the system", by interpreting racial inequalities in the underlying data as immutable biological facts rather than as reflecting the societal effects of racism.7Vyas DA Eisenstein LG Jones DS Hidden in plain sight: reconsidering the use of race correction in clinical algorithms.N Engl J Med. 2020; 383: 874-882Crossref PubMed Scopus (178) Google Scholar The authors go on to distinguish between the use of race in descriptive statistics, for which it plays a crucial role in epidemiological analyses, and in prediction tools or prescriptive clinical guidelines.7Vyas DA Eisenstein LG Jones DS Hidden in plain sight: reconsidering the use of race correction in clinical algorithms.N Engl J Med. 2020; 383: 874-882Crossref PubMed Scopus (178) Google Scholar Other key decisions made during analysis, such as the choice of model performance metrics, might mask a weak true-positive rate or might not sufficiently capture how the models fare across different groups. Despite the promise of advanced analytical approaches, many models have had little clinical utility compared with their performance in research settings. This discrepancy is partly due to study design, logistical implementation challenges, human factors, and data shift.8Kelly CJ Karthikesalingam A Suleyman M Corrado G King D Key challenges for delivering clinical impact with artificial intelligence.BMC Med. 2019; 17: 195Crossref PubMed Scopus (187) Google Scholar Without the mechanisms in place to monitor and understand poor model translation and performance in live settings, the real-world impacts that could particularly harm ethnic minorities might be missed.6Obermeyer Z Powers B Vogeli C Mullainathan S Dissecting racial bias in an algorithm used to manage the health of populations.Science. 2019; 366: 447-453Crossref PubMed Scopus (467) Google Scholar There are a number of steps that all of us in the health data science community can take to combat structural racism and its effects. First, we need to educate ourselves on the ways in which data science perpetuates racism and embed this understanding in future data scientists. Examples of this approach might include unpicking how race is conceptualised in the field, introducing modules on ethnic and other inequalities in data science teaching curricula, and supporting research and debate on the relationship between data science and health inequalities.3Ileka KM McCluney CL Robinson RAS White coats, Black scientists.https://hbr.org/2020/09/white-coats-black-scientistsDate: Sept 23, 2020Date accessed: November 3, 2020Google Scholar Second, we can seek out diverse and representative perspectives from patients and the general public, and integrate these perspectives into our research governance, ethics, and analysis plans. This approach includes engagement to ensure that new technology meets the needs of underserved communities, through partnering, for example, with community-based organisations to agree on the ethical use of datasets and on the definition of ethnic categories used. Many useful resources exist for researchers wanting to involve the public in the way they identify, prioritise, design, conduct, and disseminate their research (eg, the National Institute for Health Research's INVOLVE national advisory group). These practices could help build greater trust in the use of ethnicity data, because people might be more inclined to trust systems in which they feel represented and their best interests are demonstrably acknowledged and addressed. Placing greater emphasis on intersectional analysis in data science would also provide a more comprehensive view of the interplay between different social determinants of health and oppression for some groups of people (eg, being a Black person, a woman, and someone from a low-income household). Third, we can make the collection and reporting of disaggregated ethnicity data routine. When data on ethnicity is recorded, stratifying analyses by ethnicity can ensure trends across the wider population are not masking that of subgroups. Without these data and ethnicity-disaggregated analyses, people from disadvantaged groups will not be able to effectively lobby for change or hold leaders to account and services will not be designed with the needs of these groups in mind. The COVID-19 pandemic has further illustrated the value of data-driven approaches to addressing racial disparities in health outcomes. For example, Brigham Health's intersectional approach to data analysis highlighted early on the need for improved translation services and outreach to non-English speaking communities. Finally, we must take organisational action to address the low diversity in health data science. This approach might include reviewing and updating hiring processes; ensuring representation on executive leadership teams, boards, and expert panels; developing leadership pathways to support emerging leaders from historically underrepresented backgrounds; creating inclusive working environments that are a safe space to share ideas and concerns; and actively listening to and learning from the experiences of data scientists from ethnic minority groups (eg, Black in AI, the Shuri Network, One HealthTech).9Choo E Seven things organisations should be doing to combat racism.Lancet. 2020; 396: 157Summary Full Text Full Text PDF PubMed Scopus (3) Google Scholar As a practical first step, researchers and patients involved in analyses could adopt the tools and practices available to improve generalisability, documentation quality, transparency, and reproducibility for ethical and race-sensitive data-driven insights.10Morley J Floridi L Kinsey L Elhalal A From what to how: an initial review of publicly available AI ethics tools, methods and research to translate principles into practices.Sci Eng Ethics. 2020; 26: 2141-2168Crossref PubMed Scopus (47) Google Scholar Individual actions could include diversifying research and newsfeeds, joining communities of practice from non-Western origins (eg, Data Science Africa), following approaches promoted by the EU's Responsible Research Innovation framework, and engaging in discussions on the topic through events such as the Conference on Fairness, Accountability, and Transparency by the Association for Computing Machinery. As data stewards in a world that is increasingly data-driven, data scientists have a responsibility to tackle the different forms of racism that manifest themselves in our sector. Inaction perpetuates existing inequalities and racism; as practitioners, we all need to take more action to address racism and ensure that the benefits from the use of health data are shared equitably. MM is a Company Director of OneHealthTech and a non-executive director of the Eastern Academic Health Science Network. All other authors declare no competing interests. Ethnic bias in data linkageIn The Lancet Digital Health, Hannah Knight and colleagues1 highlight stages in the data science pipeline that are affected by and lead to racism. Data linkage is a further stage in which ethnic bias can be encoded into datasets. Ethnic bias occurs when linkage error (false or missed matches) is more likely to occur for particular ethnic groups. The problem of ethnic bias in health data linkage is well described in the literature2 and is concerning because health data are widely used for monitoring, service planning, research, evaluation, and policy. Full-Text PDF Open Access
Background Nuclear magnetic resonance (NMR) spectroscopy allows triglycerides to be subclassified into 14 different classes based on particle size and lipid content. We recently showed that these subfractions have differential associations with cardiovascular disease events. Here we report the distributions and define reference interval ranges for 14 triglyceride-containing lipoprotein subfraction metabolites. Methods Lipoprotein subfractions using the Nightingale NMR platform were measured in 9073 participants from four cohort studies contributing to the UCL-Edinburgh-Bristol consortium. The distribution of each metabolite was assessed, and reference interval ranges were calculated for a disease-free population, by sex and age group (<55, 55-65, >65 years), and in a subgroup population of participants with cardiovascular disease or type 2 diabetes. We also determined the distribution across body mass index and smoking status. Results The largest reference interval range was observed in the medium very-low density lipoprotein subclass (2.5th 97.5th percentile; 0.08 to 0.68 mmol/L). The reference intervals were comparable among male and female participants, with the exception of triglyceride in high-density lipoprotein. Triglyceride subfraction concentrations in very-low density lipoprotein, intermediate-density lipoprotein, low-density lipoprotein and high-density lipoprotein subclasses increased with increasing age and increasing body mass index. Triglyceride subfraction concentrations were significantly higher in ever smokers compared to never smokers, among those with clinical chemistry measured total triglyceride greater than 1.7 mmol/L, and in those with cardiovascular disease, and type 2 diabetes as compared to disease-free subjects. Conclusion This is the first study to establish reference interval ranges for 14 triglyceride-containing lipoprotein subfractions in samples from the general population measured using the nuclear magnetic resonance platform. The utility of nuclear magnetic resonance lipid measures may lead to greater insights for the role of triglyceride in cardiovascular disease, emphasizing the importance of appropriate reference interval ranges for future clinical decision making.
Aims Elevated low-density lipoprotein cholesterol (LDL-C) is a risk factor for cardiovascular disease; however, there is uncertainty about the role of total triglycerides and the individual triglyceride-containing lipoprotein sub-fractions. We measured 14 triglyceride-containing lipoprotein sub-fractions using nuclear magnetic resonance and examined associations with coronary heart disease and stroke. Methods Triglyceride-containing sub-fraction measures were available in 11,560 participants from the three UK cohorts free of coronary heart disease and stroke at baseline. Multivariable logistic regression was used to estimate the association of each sub-fraction with coronary heart disease and stroke expressed as the odds ratio per standard deviation increment in the corresponding measure. Results The 14 triglyceride-containing sub-fractions were positively correlated with one another and with total triglycerides, and inversely correlated with high-density lipoprotein cholesterol (HDL-C). Thirteen sub-fractions were positively associated with coronary heart disease (odds ratio in the range 1.12 to 1.22), with the effect estimates for coronary heart disease being comparable in subgroup analysis of participants with and without type 2 diabetes, and were attenuated after adjustment for HDL-C and LDL-C. There was no evidence for a clear association of any triglyceride lipoprotein sub-fraction with stroke. Conclusions Triglyceride sub-fractions are associated with increased risk of coronary heart disease but not stroke, with attenuation of effects on adjustment for HDL-C and LDL-C.
Background Interleukin 6 concentration is associated with myocardial injury, heart failure, and mortality after myocardial infarction. In the Norwegian tocilizumab non–ST‐segment–elevation myocardial infarction trial, the first randomized trial of interleukin 6 blockade in myocardial infarction, concentration of both C‐reactive protein and troponin T were reduced in the active treatment arm. In this follow‐up study, an aptamer‐based proteomic approach was employed to discover additional plasma proteins modulated by tocilizumab treatment to gain novel insights into the effects of this therapeutic approach. Methods and Results Plasma from percutaneous coronary intervention–treated patients, 24 in the active intervention and 24 in the placebo‐control arm, drawn 48 hours postrandomization were randomly selected for analysis with the SOMAscan assay. Employing slow off‐rate aptamers, the relative abundance of 1074 circulating proteins was measured. Proteins identified as being significantly different between groups were subsequently measured by enzyme immunoassay in the whole trial cohort (117 patients) at all time points (days 1–3 [7 time points] and 3 and 6 months). Five proteins identified by the SOMAscan assay, and subsequently confirmed by enzyme immunoassay, were significantly altered by tocilizumab administration. The acute‐phase proteins lipopolysaccharide‐binding protein, hepcidin, and insulin‐like growth factor‐binding protein 4 were all reduced during the hospitalization phase, as was the monocyte chemoattractant C‐C motif chemokine ligand 23. Proteinase 3, released primarily from neutrophils, was significantly elevated. Conclusions Employing the SOMAscan aptamer‐based proteomics platform, 5 proteins were newly identified that are modulated by interleukin 6 antagonism and may mediate the therapeutic effects of tocilizumab in non–ST‐segment–elevation myocardial infarction.
Sleep is an essential human function but its regulation is poorly understood. Using accelerometer data from 85,670 UK Biobank participants, we perform a genome-wide association study of 8 derived sleep traits representing sleep quality, quantity and timing, and validate our findings in 5,819 individuals. We identify 47 genetic associations at P < 5 × 10 −8 , of which 20 reach a stricter threshold of P < 8 × 10 −10 . These include 26 novel associations with measures of sleep quality and 10 with nocturnal sleep duration. The majority of identified variants associate with a single sleep trait, except for variants previously associated with restless legs syndrome. For sleep duration we identify a missense variant (p.Tyr727Cys) in PDE11A as the likely causal variant. As a group, sleep quality loci are enriched for serotonin processing genes. Although accelerometer-derived measures of sleep are imperfect and may be affected by restless legs syndrome, these findings provide new biological insights into sleep compared to previous efforts based on self-report sleep measures.
Liver dysfunction and type 2 diabetes (T2D) are consistently associated. However, it is currently unknown whether liver dysfunction contributes to, results from, or is merely correlated with T2D due to confounding. We used Mendelian randomization to investigate the presence and direction of any causal relation between liver function and T2D risk including up to 64,094 T2D case and 607,012 control subjects. Several biomarkers were used as proxies of liver function (i.e., alanine aminotransferase [ALT], aspartate aminotransferase [AST], alkaline phosphatase [ALP], and γ-glutamyl transferase [GGT]). Genetic variants strongly associated with each liver function marker were used to investigate the effect of liver function on T2D risk. In addition, genetic variants strongly associated with T2D risk and with fasting insulin were used to investigate the effect of predisposition to T2D and insulin resistance, respectively, on liver function. Genetically predicted higher circulating ALT and AST were related to increased risk of T2D. There was a modest negative association of genetically predicted ALP with T2D risk and no evidence of association between GGT and T2D risk. Genetic predisposition to higher fasting insulin, but not to T2D, was related to increased circulating ALT. Since circulating ALT and AST are markers of nonalcoholic fatty liver disease (NAFLD), these findings provide some support for insulin resistance resulting in NAFLD, which in turn increases T2D risk.
Carotid artery intima media thickness (cIMT) and carotid plaque are measures of subclinical atherosclerosis associated with ischemic stroke and coronary heart disease (CHD). Here, we undertake meta-analyses of genome-wide association studies (GWAS) in 71,128 individuals for cIMT, and 48,434 individuals for carotid plaque traits. We identify eight novel susceptibility loci for cIMT, one independent association at the previously-identified PINX1 locus, and one novel locus for carotid plaque. Colocalization analysis with nearby vascular expression quantitative loci (cis-eQTLs) derived from arterial wall and metabolic tissues obtained from patients with CHD identifies candidate genes at two potentially additional loci, ADAMTS9 and LOXL4 . LD score regression reveals significant genetic correlations between cIMT and plaque traits, and both cIMT and plaque with CHD, any stroke subtype and ischemic stroke. Our study provides insights into genes and tissue-specific regulatory mechanisms linking atherosclerosis both to its functional genomic origins and its clinical consequences in humans.