Most existing genotype-by-environment interaction (G×E) methods assume a known causal direction as an assumption that often does not hold and can lead to biased estimates and spurious findings. To address this, we introduce the Genetic Causality Inference Model (GCIM), a novel approach designed to infer causal directions in G×E studies. GCIM integrates polygenic risk scores (PRS) for both the exposure and the outcome to strengthen causal inference and reduce spurious interaction signals. We evaluated GCIM using simulated data across varying genetic and residual correlation settings and compared its performance to existing PRS-by-environment (PRS×E) models under both null and alternative G×E scenarios. GCIM was also applied to real-world UK Biobank data in both causal directions. GCIM consistently outperformed existing methods by accurately identifying the absence of G×E variance and avoiding false positives, even in the presence of strong phenotypic heteroscedasticity due to residual heterogeneity. Other methods often generated spurious associations, especially under reverse causality. Applying GCIM to UK Biobank data, we investigated 11 circulating biomarkers (including liver enzymes, lipids, and inflammatory markers) and three anthropometric traits (BMI, body fat, and waist-to-hip ratio [WHR]). GCIM identified that bilirubin modulates genetic effects on BMI and WHR, while body fat modulates genetic effects on C-reactive protein, with associations remaining significant after multiple testing corrections. Overall, GCIM provides a more reliable framework for GxE analysis, particularly under challenging conditions such as residual heterogeneity and uncertain causal direction. However, further development is needed to improve its statistical power.
Mapping the pleiotropic effect of genetic variation on biological processes and complex phenotypes is fundamental to extracting translational insight from genome-wide association studies (GWAS). Here we present The Human Genotype-Phenotype Map (GPMap), a repository of colocalizing genetic associations across 15,997 complex traits and 2.7 million molecular measurements, leveraging common and rare variants and cis- and trans-acting effects across disaggregated tissue types and single cell datasets to trace the complex pathways through which they act. We identify over 49.3 million colocalizing trait pairs, which aggregate into 97,393 colocalization groups, representing distinct pleiotropic variants based on shared genetic signals, with 55.8% of genome-wide significant disease-associated loci colocalizing with at least one molecular trait. This insight facilitates clustering of complex health and disease phenotypes based on genetic architecture, and the dissection of polygenic traits reflecting the composite impact of many underlying processes. We show that leveraging pleiotropic information can enhance the selection of genetic instruments for causal inference approaches and improves prediction of drug trial success. This open-source resource is available at https://gpmap.opengwas.io, with functionality for user GWAS upload. ### Competing Interest Statement TRG and GH receive funding from Biogen for research not presented here. TRG receives funding from Biogen, GSK, Roche and Novartis for research not presented here. LP was part of an Innovative Medicines Initiative-European funded consortia (biomap-imi.eu) with multiple industry partners. JJ holds a permanent position at Genomics England for work unrelated to this project. AH consults for Altis Medicines (Cambridge, US) on unrelated work. ### Funding Statement This study/research is funded by the National Institute for Health and Care Research (NIHR) Bristol Biomedical Research Centre (BRC). The views expressed are those of the author(s) and not necessarily those of the NIHR or the Department of Health and Social Care. The research was carried out in, and supported by, the MRC Integrative Epidemiology Unit (MC\_UU\_00032/1 and MC\_UU\_00032/3). ### Author Declarations I confirm all relevant ethical guidelines have been followed, and any necessary IRB and/or ethics committee approvals have been obtained. Yes The details of the IRB/oversight body that provided approval or exemption for the research described are given below: The set of datasets are from these sources: https://opengwas.io/ https://yanglab.westlake.edu.cn/data/brainmeta https://www.nature.com/articles/s41467-018-04558-1 https://doi.org/10.1038/s41586-023-06592-6 https://opengwas.io/ https://www.nature.com/articles/s41586-021-04103-z https://app.genebass.org/ https://azphewas.com/ https://www.gtexportal.org/home/downloads/adult-gtex/overview https://onek1k.org/ http://www.godmc.org.uk/ https://www.eqtlgen.org/ I confirm that all necessary patient/participant consent has been obtained and the appropriate institutional forms have been archived, and that any patient/participant/sample identifiers included were not known to anyone (e.g., hospital staff, patients or participants themselves) outside the research group so cannot be used to identify individuals. Yes I understand that all clinical trials and any other prospective interventional studies must be registered with an ICMJE-approved registry, such as ClinicalTrials.gov. I confirm that any such study reported in the manuscript has been registered and the trial registration ID is provided (note: if posting a prospective study registered retrospectively, please provide a statement in the trial ID field explaining why the study was not registered in advance). Yes I have followed all appropriate research reporting guidelines, such as any relevant EQUATOR Network research reporting checklist(s) and other pertinent material, if applicable. Yes We have provided multiple ways for users to use and interact with the data. For an interactive exploration of the data, please visit https://gpmap.opengwas.io/. To upload a GWAS that will be run against all existing data in the database, either use the website (https://gpmap.opengwas.io?showUpload=true) or use the R package gpmapr (https://github.com/MRCIEU/gpmapr). Vignettes have been created to illustrate how to use the R package (https://mrcieu.r-universe.dev/gpmapr). We have also published the database of results, along with all fine-mapped summary statistics in a cloud storage bucket. As the cost to hosting this amount of data is high, we have set this to a 'requester pays' model, where the user pays for the cost to download. The bucket is available at https://console.cloud.google.com/storage/browser/genotype-phenotype-map. The code for this project is available over 3 different publicly available GitHub repositories. The data processing pipeline is available at https://github.com/MRCIEU/genotype-phenotype-map, the website and API are available at https://github.com/MRCIEU/genotype-phenotype-api, and the R package is available at https://github.com/MRCIEU/gpmapr. As the data is large and available via both the R package and website, we have not included the results as supplementary material. The data from the original studies is available via the R package, source_url is provided for each colocalization result.
Abstract Confounding is a central challenge in observational studies. Here, we propose a framework for identifying confounders of two non-causally related traits by employing cross-trait pleiotropy analysis to detect genetic loci that affect both traits and multi-trait colocalisation to identify molecular phenotypes mediating these effects. We apply this approach to the analysis of C-reactive protein (CRP) - a non-specific marker of inflammation - and 10 inflammation-related cancers. In UK Biobank, higher pre-diagnostic CRP levels are associated with increased risk of multiple cancers, but bidirectional Mendelian randomization provides little evidence for a causal relationship. Cross-trait genetic analyses identify 92 loci with shared CRP-cancer effects including those with established roles in cancer and 50 novel loci such as RSPO3 (breast cancer) and GCKR (colorectal cancer). Integration with proteomic and single-cell transcriptomic data identified putative molecular mediators at 24 loci including plasma TLR1 levels in breast cancer and CD4 + T cell IRF5 expression in kidney cancer. Notably, 15 candidate effector genes encode targets of approved or investigational medications, including IL6 , PDE4D , and CASP8 , indicating potential opportunities for their repurposing for cancer prevention. The proposed approach provides a generalisable framework for leveraging non-causal phenotypic relationships to yield insights into disease mechanisms and therapeutic targets for disease prevention.
Background Higher adiposity in early-life has consistently been associated with a reduced risk of breast cancer in later life, with Mendelian randomization (MR) studies supporting a causal effect. However, concerns have been raised that selection bias, particularly collider stratification due to selective participation or survival, may induce spurious protective MR estimates. Methods We triangulated across empirical analyses and simulations to evaluate whether selection-induced bias could plausibly explain the inverse effect estimate of early-life adiposity on breast cancer risk. First, we analysed proxy-genotype Mendelian randomization (MR) analyses of breast cancer in relatives, in which participant genotype is used as a proxy for relatives’ genotype, and conducted family-based simulations to assess whether attenuation in relative-based estimates could arise without selection bias. Second, we performed multivariable MR analyses of parental survival to evaluate survival-related selection mechanisms. Third, we conducted extensive simulations to quantify the magnitude of bias introduced by selection under a range of plausible and extreme scenarios, including interaction-driven selection. Results The weaker proxy-genotype MR estimates of breast cancer in relatives, compared with MR estimates for an individual's own breast cancer, were reproduced in family-based simulations without selection bias, indicating that this pattern does not provide evidence for selection bias. Multivariable MR analyses of parental survival indicated that survival differences are primarily driven by mid-to-late adulthood, not early-life, adiposity, providing little support for survival-related selection acting through early-life adiposity. In simulation analyses, additive selection produced minimal bias, while interaction-driven selection generated increasing distortion; however, even under extreme scenarios, the magnitude of bias was insufficient to replicate the observed protective effect estimate. In simulations where selection depended on mid-to-late adulthood rather than early-life adiposity, bias was expressed primarily in mid-to-late adulthood MR estimates, with little distortion of early-life MR estimates. Across all simulated scenarios, the combined pattern of empirical univariable and multivariable MR findings was not reproduced by selection alone. Conclusions Although selection bias can influence MR estimates, our findings suggest that plausible selection mechanisms are unlikely to substantively explain the observed inverse effect estimate of early-life adiposity on breast cancer risk. These results support a causal interpretation of the strong protective effect estimate of early-life adiposity on breast cancer risk and highlight the value of triangulating evidence across complementary approaches when evaluating bias in lifecourse MR. Plain Language Summary Previous studies have found that having a larger body size in early life is linked to a lower risk of developing breast cancer later in life. Mendelian randomization studies, which use genetic variation to investigate possible causal effects, have supported this finding. However, it has been suggested that the result could be explained by selection bias arising from who survives or participates in the studies analysed. We used several complementary approaches to evaluate this possibility. Together, the findings provide little evidence that the selection mechanisms examined explain the protective effect estimate of early-life adiposity on breast cancer. This work also provides a practical framework for investigating selection bias in lifecourse Mendelian randomization studies.
Abstract Whether and how genetic regulation of DNA methylation (DNAm) change across the lifespan remains unclear. Here, we map age-dependent methylation quantitative trait loci (longitudinal mQTLs) using linear mixed models applied to repeated blood DNAm measures from birth, childhood, adolescence and adulthood in the Avon Longitudinal Study of Parents and Children. We identify 2,210 longitudinal mQTLs (2,393 SNP-CpG pairs; 7.3% trans ) and observe consistent genotype-by-age effects in two independent cohorts of diverse ancestries (Pearson’s r = 0.85 in the Generation R Study; r = 0.56 in the Drakenstein Child Health Study). Longitudinal mQTLs show increasing effects with age at half of loci and associations with multiple phenotypes. CpGs with longitudinal mQTLs are more heritable and enriched in regulatory elements and pathways related to multicellular organism development and cell adhesion. These results chart dynamic genetic influences on the human methylome and provide a novel perspective on epigenetic regulation.
Parental smoking has been linked to several adverse offspring cardiometabolic outcomes; however, evidence is conflicting regarding the causal and long-term nature of these associations. We investigated the effects of maternal and paternal smoking, capturing exposure before, during, and after pregnancy, on eleven offspring cardiometabolic risk factors related to body composition, blood pressure, glucose, and lipid levels in adulthood. We applied a multi-method intergenerational Mendelian randomization (MR) framework, combining two-sample MR (outcome GWAS, n = up to 564,160) and one-sample MR analyses (n = up to 17,484 genotyped mother-father-offspring trios with offspring cardiometabolic risk factors) from the HUNT cohort, Norway, and the UK Biobank and ALSPAC cohorts in the United Kingdom. Smoking behaviour was instrumented using genome-wide significant variants for smoking initiation and heaviness from large genome-wide association studies (2019 and 2022), with additional analyses of the CHRNA5 variant rs16969968. Using the two-sample MR approach, we found an average change in adult offspring waist-hip ratio (WHR) per one standard deviation (SD) increase in maternal cigarettes smoked per day of 0.25 SD (95
Abstract Altered affect and cognitive dysfunction are burdensome features of many neuropsychiatric conditions that are highly comorbid, remain poorly understood, and have few efficacious treatments. Exploring their genetic architecture and causal relationships may provide insight into their aetiology and comorbidity. Compared to related but distinct traits (depression, wellbeing, neuroticism), findings from genome-wide association studies (GWAS) of positive and negative affect may be informative due to these phenotypes being less heterogenous (compared to broader phenotypes of wellbeing and depression), whilst still retaining important clinical relevance as potential targets for indirect intervention (compared to neuroticism which may be less modifiable). Using data from the Lifelines Cohort Study, we conducted the first GWAS of positive and negative affect using a validated measure (N = 57,946), and four cognitive domains: working memory, reaction time, learning and memory, and executive function (N ≥ 35,729). We then assessed genetic overlap and potential causal relationships using genetic correlation and bidirectional Mendelian randomization (MR) analyses, incorporating large GWAS on related—albeit distinct—phenotypes (depression, anxiety, wellbeing, general cognitive ability [GCA]). We identified one SNP that reached genome-wide significance ( p < 5 × 10 –8 ) for reaction time, and many independent SNPs with suggestive associations for other phenotypes (N = 11–20). For most phenotypes, exploratory gene mapping indicated that SNPs with suggestive associations have higher gene expression in brain tissue compared to other tissues; however, this only met Bonferroni-corrected p -value threshold criteria for positive affect and visual learning and memory. Genetic correlations between negative and positive affect suggest that they are dissociable constructs ( r g = − 0.18). GCA has higher genetic overlap with negative affect than with positive affect ( r g = − 0.19 vs − 0.06), which could indicate that negative affect and GCA have a higher shared neural basis and/or they exhibit causal relationships. Supporting the latter, MR analyses indicated that higher GCA may reduce negative affect, depression, and anxiety, and increase wellbeing, with little impact on positive affect. Conversely, MR analyses indicated that higher risk of depression and lower wellbeing may causally reduce GCA. Together, these findings suggest that interventions that indirectly influence GCA may be valid targets to prevent negative affect, while interventions that indirectly influence depression/wellbeing may be valid targets for GCA.
Iron is essential for both humans and pathogens, yet its genetic regulation remains understudied in African populations. Here, we report genome-wide association studies of six iron-related biomarkers in 3928 children from five sites across Africa, with replication in 2868 African American adults and investigate associations with severe malaria and bacteremia. We identify previously unreported loci at genome-wide significance, for transferrin at GTF3C5, and for hepcidin at CHCHD7/SDR16C5. Variants tagging the DUP4 haplotype, encoding the Dantu blood group (rs552439837) are associated with soluble transferrin receptor levels. Variants at GTF3C5 (rs2905094) and DUP4 confer protection against severe malaria and bacteremia. The CHCHD7/SDR16C5 variant (rs73596248) increases hepcidin levels and is associated with reduced risk of Klebsiella pneumoniae and Staphylococcus aureus bacteremia. Polygenic risk scores derived from European data show limited transferability to African populations. In this work, we demonstrate new genetic insights into iron regulation and highlight iron's role in host-pathogen interactions.
Genome-wide association studies (GWAS) are conventionally conducted in cohorts spanning a wide age-range. These studies typically assume that genetic associations are constant across different ages. Some traits, however, may have age-varying genetic associations. This has implications for the interpretation of genetic effects derived in downstream applications, such as Mendelian randomization (MR) analyses. In this study we conducted a series of age-stratified GWAS on individuals aged 40-69 years in the UK Biobank, for body-mass index (BMI) and three blood pressure traits (systolic, diastolic and pulsatile pressure (PP)) in 2-year age strata (N up to 26,330). We used a meta-regression approach to systematically identify single nucleotide polymorphisms (SNPs) with evidence for age interaction effects among trait-associated GWAS signals and additional loci genome-wide. Within an MR framework, we examine the relationship between BMI and blood pressure traits on cardiovascular and cardiometabolic outcomes (type-2 diabetes (T2D), stroke, peripheral artery disease (PAD), heart failure, coronary heart disease and atrial fibrillation). Next, we describe the effect of the SNP*Age interaction on those relationships in a modified inverse-variance weighted (ivw) analysis. We identified differential enrichment of age-interaction effects, which was trait dependent. For example, 10.3% of BMI discovery SNPs had evidence for an age-interaction in our data compared to 44.7% for PP (at P < 0.05). Our downstream MR and modified ivw analyses highlight the influence of age on the genetically predicted relationship between PP and adverse cardiovascular outcomes. For example, our results indicated that an increased rate of change in genetically predicted PP across the age period is associated with higher susceptibility to PAD (interaction odds ratio = 2.71; P = 1.82x10-13; 95%-CI: 2.08-3.53). The data generated in this project provides a valuable resource for further exploration of mechanisms relevant to the genetic architecture of complex traits and all summary data has been made accessible to the research community.
The explosion of Mendelian Randomization (MR) submissions of dubious quality to journals globally is well recognized. Contributing to this deluge of publications may be the poor understanding among practitioners, reviewers, and editors of gene-environment equivalence, the fundamental principle of MR.
The theorised risk that confounded rare variant associations will emerge from population based genetic studies has not been investigated empirically. Here, we use 306,991 sequenced exomes from the UK Biobank to demonstrate that recent demography is poorly captured by common and rare variant principal components, and accounting for haplotype sharing does not eliminate false-positive rare variant associations with non-heritable spatially structured traits. Through re-analysis of 155 phenotypes in siblings, we show a trend of higher effect estimates bias for non-uniformly distributed traits, suggesting population stratification is most pervasive in these settings. Despite its spatial structure, bias of rare variant associations with height appeared most strongly influenced by assortative mating. We explore the risk of elevated false discovery rates for recent variants private to extended families sharing polygenic liability to extreme phenotypes, as well as through local linkage with common causal variants. Overall, we consider the complex confounding mechanisms that can impact rare variant studies and demonstrate family-based approaches can enable important sensitivity analyses.
Previous evidence suggests that higher prepubertal adiposity protects against breast cancer risk. Whether this protection extends into early adulthood remains uncertain. We conducted genome-wide association studies on body mass index (BMI) in nulliparous women from menarche to <40 years across five cohorts, with additional analyses in three subintervals of this life stage. Results were meta-analyzed, and two-sample univariable and multivariable Mendelian randomization was applied within a lifecourse framework to assess the effect of BMI on breast cancer risk. Between menarche and <40 years, we observed heterogeneity in genetic effects. Genome-wide correlations further suggest that BMI during this early adult period may be partly influenced by distinct genetic factors compared with adiposity at other life stages. Higher genetically proxied BMI between menarche and 40 years reduced breast cancer risk. This protective effect attenuated after adjusting for prepubertal adiposity. These findings refine our understanding of adiposity's role in breast cancer and highlight earlier life stages as critical windows for risk modulation.
Background/Objectives: Childhood appetitive traits are heritable behavioural phenotypes hypothesized to link genetic susceptibility to obesity risk. Yet their genetic architecture and role in mediating polygenic adiposity risk remain poorly understood. Methods: We conducted the largest survey of childhood eating behaviour to date, allowing us to perform genome-wide association studies of six appetitive domains derived from 18 items of the parent-reported Children's Eating Behaviour Questionnaire in up to 31,018 eight-year-old children from the Norwegian Mother, Father and Child Cohort Study (MoBa). A trio-based design enabled decomposition of direct and indirect genetic effects on appetite and BMI. Results: We identified ten independent genome-wide significant loci for childhood eating behaviour, primarily across Food Responsiveness, Satiety Responsiveness, and Food Fussiness, eight of which lie at established childhood or adult BMI loci. Food Responsiveness and Satiety Responsiveness showed both phenotypic and genetic correlations with BMI trajectories from early childhood through adolescence. Statistical mediation analyses indicated that 22.1% and 10.4% of the aggregated genetic association with BMI at age 8 could be decomposed through these traits, respectively. Locus-specific patterns further suggested mechanistic pathways, with the FTO locus acting predominantly via Food Responsiveness, and the ADCY3 locus via Satiety Responsiveness. Trio analyses demonstrated that both BMI and eating behaviour associations were predominantly explained by children's inherited alleles, with minimal contribution from indirect effect from parental adiposity, although parental genetic liability influenced reporting of Satiety Responsiveness. Conclusions: Childhood appetitive traits capture a substantial proportion of genetic susceptibility to adiposity through distinct eating behaviour pathways (under standard mediation assumptions). These effects are primarily driven by the child's own genotype rather than indirect parental influences, positioning appetite as a plausible, biologically grounded target for early obesity prevention.
Abstract A greater genetic susceptibility has been proposed as an explanation of the greater rates of cardiovascular and metabolic disease in South Asian relative to European populations. We first demonstrate that after accounting for technical artefacts the genetic effects for related traits are largely consistent between ancestral groups, which downplays the role of GxG or GxE interactions driving differential prevalence. If higher genetic susceptibility in South Asians is due to selective pressures acting through adiposity-related traits in the evolutionary past, signatures of selection should be evident at loci associated with cardiometabolic disease and other causally related traits (e.g. fat distribution). We tested for enrichment of several selection statistics (F ST, XP-EHH and XP-nSL) at loci associated with a range of traits related to cardiometabolic disease, in comparison to a null distribution of linkage disequilibrium (LD) score and minor allele frequency (MAF) matched SNPs. Loci associated with a subset of these traits (Type 2 diabetes mellitus, trunk fat percentage, body fat percentage and trunk fat mass) exhibited enrichment for F ST , consistent with a moderate adaptive explanation for their cross-population differentiation. In contrast, none of the studied traits were enriched for haplotype-based statistics, indicative that cross population genetic divergence is unlikely to have been driven by recent selective sweeps but has rather likely arisen from either ancient selection or recent polygenic selection acting on standing variation.
Mendelian randomization (MR) is a popular statistical technique that uses genetic variants to explore causal relationships in observational epidemiology. Summary-level MR, the most common form, relies on published GWAS summary statistics to estimate causal effects between exposures and outcomes. However, empirical analyses tend to ignore issues relating to Winner's Curse of instrument effects, weak instrument bias and sample overlap. Our simulations and empirical analyses using the UK Biobank indicate that such mechanisms can induce substantial bias in routine MR approaches. We propose MR Simulated Sample Splitting (MR-SimSS), a novel method that corrects this bias requiring no additional data beyond GWAS summary statistics for the exposure and outcome of interest. It operates by simulating statistically independent sets of summary statistics, analogous to what would be produced by splitting the individual-level data into independent subsets, which can then be plugged into existing two-sample MR methods. With sufficient instrument variants, MR-SimSS is robust to a range of sample overlap scenarios, providing a practical and modular solution to Winner's Curse and weak instrument bias.
BACKGROUND:Epidemiological studies suggest associations of asthma with the psychosis spectrum (psychotic experiences, bipolar disorder, schizophrenia), but the mechanisms underlying these associations remain unclear. METHODS:We examined the relationship between asthma and psychosis-related outcomes using observational cohort, polygenic score, and Mendelian randomization (MR) analyses, and assessed the possibility of shared genetic underpinnings between these traits using genetic colocalisation. RESULTS:Results from a UK population-based prospective birth cohort suggest that asthma at age 7 and polygenic risk for asthma are associated with psychotic experiences in early adulthood. Results from two-sample MR analyses do not support causal relationships of genetic liability to asthma with bipolar disorder or schizophrenia. Instead, genetic correlation and colocalization analyses point to the presence of shared genetic etiology between the conditions. We identified 8 genomic regions with potentially shared causal genes between asthma and bipolar disorder or asthma and schizophrenia. Using genetically predicted mRNA expression in condition-relevant tissues (brain and lung), we identified 16 genes shared between conditions, which include FADS1, SLC4A10, BDH2, and CISD2. CONCLUSIONS:Our results suggest that population-level associations of asthma with psychosis spectrum conditions could be due to shared molecular mechanisms involving fatty acid metabolism, ion channel activity, and iron homeostasis.
Personality traits describe stable differences in how people think, feel and behave, and how they interact with and experience their social and physical environments1,2. Many questions remain unanswered about associations between DNA and personality traits, such as their robustness, their generalizability and the biological and social pathways through which they act. Here we meta-analyse data across 46 cohorts comprising 611,037 to 1.14 million participants with European-like and African-like genomes for genome-wide association studies (GWAS) of the Big Five personality traits (extraversion, agreeableness, conscientiousness, neuroticism and openness to experience), and data from up to 50,725 participants for within-family GWAS. We identify 1,260 lead genetic variants associated with personality, including 824 novel variants3. Common genetic variants explain a moderate 4.8-9.3% of the variance in measures of each trait, and 9.3-13.3% among instruments with typical measurement reliability. Genetic associations with personality are highly consistent but not identical across geography, reporter (self versus close other), age group and measurement instrument, and we find minimal spousal assortment for personality in recent history. In contrast to many other social and behavioural traits4,5, within-family GWAS and polygenic index analyses indicate that genetic associations with personality are minimally confounded by the shared family environment. Polygenic prediction, genetic correlation and Mendelian randomization analyses indicate that personality traits have widespread, potentially causal associations with consequential behaviours and life outcomes. Overall, we find that the genetic architecture of personality is robustly generalizable, minimally confounded and widely relevant to human experience.
Mendelian randomization has evolved from a niche methodology to a widely adopted research approach. In this Perspective, we briefly present a bibliometric analysis of the Mendelian randomization literature to inform a discussion of how Mendelian randomization studies are conducted and how they do not fully realize the potential of the data and techniques available to empirically examine the reliability of assumptions. We propose that future progress will depend on integrating empirical evidence from molecular, cellular, animal and quasi-experimental studies to assess its assumptions and causal claims. We also highlight how the shifting landscape of genetic and genomic data presents new challenges and opportunities for the Mendelian randomization framework, providing a deeper understanding of causal mechanisms.