
Although individual traits related to metabolic dysfunction-associated steatotic liver disease (MASLD) have been investigated through large-scale genome-wide association studies (GWASs), the shared genetic susceptibility across these traits remains unclear. We therefore conducted a multivariate GWAS of key MASLD-related traits to elucidate their common genetic architecture. We applied genomic structural equation modeling to model a latent genetic factor (MASLD-F) underlying genetically correlated MASLD-related traits, leveraging their GWAS-derived genetic correlations. We then performed functional annotations, including fine-mapping, transcriptome-wide association study, and cell- and tissue-type-specific enrichment analyses, and conducted Mendelian randomization analyses to identify modifiable risk factors. Our multivariate MASLD-F GWAS identified 50 independent variants across 48 genomic loci. Transcriptomic imputation identified several MASLD-F-associated genes, including ARNTL, NPC1, BTBD10, VDAC2, TSKU, SFMBT1, and ABHD17C. We observed significant enrichment of MASLD-F-related genetic signals predominantly in brain tissues, pancreatic islets, and the adrenal gland. Additionally, six modifiable risk factors and four modifiable protective factors for MASLD-F were identified. These findings reveal a complex shared genetic architecture underlying MASLD components, thereby expanding our understanding of disease pathogenesis and providing novel insights for precision medicine and public health interventions.
Genome-wide association studies (GWASs) have been extensively adopted to depict the underlying genetic architecture of complex traits. Recent studies show that knockoff-based methods can identify variants with unique, potentially causal effects on phenotypes. However, their statistical validity and effectiveness in studies with related individuals, such as the UK Biobank, remain unexplored. In this paper, we extensively evaluate a simple and effective analytical strategy that integrates GhostKnockoffs and state-of-the-art marginal association tests. We show that this approach is robust to arbitrary relatedness structure as long as the input Z-scores are derived from valid generalized linear mixed models. This robustness also extends GhostKnockoffs to other GWASs settings, including meta-analysis of studies with sample overlap when the input score test Z-scores are properly calibrated, and association test statistics beyond score tests in independent sample settings. We demonstrate the method's validity and practical utility using simulation studies and a meta-analysis of nine European ancestral genome-wide association studies and whole exome/genome sequencing studies for the Alzheimer's disease.
Mendelian randomisation (MR) is an approach to causal inference that uses genetic variants to infer whether or not a causal effect exists, unbiased by unobserved confounding. MR estimation usually considers the effect of a single exposure on an outcome; it has recently been extended to explore potential effects of multiple exposures using multivariable MR (MVMR). Existing MVMR models are restricted to a few exposure traits in a single estimation, particularly if those traits are highly correlated. However, for many relationships of interest there are many highly correlated exposures which may have a causal effect on the outcome. MVMR Bayesian model averaging (MVMR-BMA) provides a hypothesis-free exposure selection approach with many correlated exposures. Although potentially very powerful, BMA approaches to estimation are not commonly applied in epidemiological studies. Here we describe the application of MVMR-BMA to the selection of maternal metabolites that are causal for offspring birthweight. We describe the inputs and outputs of the model in detail and discuss the appropriate sensitivity analyses, illustrating these with our application. Through this, we hope to provide a guide to help other researchers, who are potentially unfamiliar with the terminology of Bayesian analysis, but would like to apply the method to their data.
There is a need for genetic analytical methods that integrate multi-individual identity-by-descent (IBD) tools with phenotypic enrichment testing to discover novel shared haplotypes contributing to disease traits. Existing tools are designed to identify IBD sharing and leave interpretation and phenotype association tests to further analyses. Here we present Distant Relatedness for Identification and Variant Evaluation (DRIVE) v3, a python command-line interface tool that identifies networks of participants who share an identical haplotype at a given genomic location. Given phenotypic data, DRIVE additionally estimates significant enrichment of dichotomous traits within networks. DRIVE is designed for efficient use across large-scale genetic data resources, featuring a versatile application programming interface and a backend structure designed for flexible integration into existing analytical pipelines. In this work, we describe the implementation of DRIVE v3 and illustrate two applications of the tool to an autosomal dominant condition and to an autosomal recessive condition, cardiomyopathy and cystic fibrosis, respectively. These applications highlight the substantial performance improvements between v1 and v3 and demonstrate practically how the newer features of DRIVE such as the enrichment test can be used in the interpretation of the identified networks.
Parkinson's disease (PD) is a complex neurodegenerative disorder with a significant genetic component. While genome-wide association studies (GWAS) have been instrumental in identifying genetic variants associated with PD, the reliance on large sample sizes and population-level analyses may overlook variants with lower minor allele frequencies or individual-specific relevance. Individualized Bayesian Inference (IBI) offers a promising method to complement GWAS by identifying and prioritizing candidate genetic markers at both the individual and patients-like-me subgroup levels. This study evaluates the application of IBI to PD genetics, using GWAS as a baseline for comparison. We analyzed genetic data from the Fox Insight online study, including 8840 individuals (8585 PD cases and 255 controls). IBI prioritized variants that were not detected or were ranked substantially lower by GWAS, including variants within or near genes with prior PD association. The top 200 IBI SNPs showed stronger predictive performance in ANN models (AUC = 0.79) than the top 200 GWAS SNPs (AUC = 0.72), providing complementary support for the utility of IBI-based prioritization in this cohort. Notably, IBI highlighted variants with lower minor allele frequencies that GWAS did not detect. This study demonstrates the utility of IBI as a complementary tool for prioritizing PD-related candidate variants and genes for further investigation.
The Cancer dependency maps (DepMap) identify genetic dependencies in cancer cells using large-scale loss-of-function screens, providing a foundation for cancer-specific treatment strategies. However, discrepancies exist between cancer cell line models (CCLs) and patient-derived tumor models, particularly in translating findings to clinical settings. To bridge this gap, computational approaches such as artificial intelligence-based domain adaptation can assist in aligning laboratory and patient-derived molecular data, thereby improving the translation of preclinical findings into personalized treatment strategies. We developed a deep unsupervised domain adaptation (UDA) algorithm to align features between source and target domains. It was trained on labeled CCLs data from the source domain and unseen, unlabeled CCL data from the target domain. The trained model was applied to predict the dependency map of breast cancer (BC) patients in The Cancer Genome Atlas (TCGA). To validate its performance, we used the predicted BC dependency map to classify ER + /HER2 + BC subtype statuses and identify synthetic lethality (SL) gene pairs for drug discovery. Our model demonstrated high accuracy in predicting cancer dependency maps for patient-derived tumors. The generated maps showed excellent performance in predicting ER + /HER2+ subtype statuses, with an area under the curve of the receiver operating characteristic (AUC-ROC) of 0.99. Notably, our analysis also identified two potential synthetic lethality gene pairs: PBRM1-NF2 and PBRM1-CTNND2, which can be potentially used for developing precision therapies for ER + /HER2+ breast cancer. Domain adaptation is a promising approach for transferring biological knowledge between different cancer models and improving patient-specific treatment strategies.
Electronic health records (EHRs) are valuable sources of data but are susceptible to biases from missing data and sample selection, often due to clinically informative visiting processes and non-probability sampling. This research explores whether genetic data, typically measured on nearly all participants in EHR-linked biobanks, can be used to mitigate these biases. Simulations were performed under conditions of missing completely at random (MCAR), missing at random (MAR), and missing not at random (MNAR) within random and biased sampling frameworks. We evaluated PRS-informed imputation, PRS-uninformed imputation, and complete case analysis across these scenarios in terms of bias, coverage, and root mean square error (RMSE) of the regression coefficient estimates. A real-world example using data from the Michigan Genomics Initiative (MGI, n = 68,063) compared the effectiveness of these methods against national benchmark estimates. PRS-informed imputation generally reduced bias and RMSE and improved coverage, particularly under MAR conditions in random samples. In analyses of biased samples of n = 10,000 with MAR exposure-only missingness, weighted, PRS-informed imputation analyses showed substantially lower percent bias (0.6%) and closer to nominal coverage (89.1%) compared to weighted, complete case analyses (9.4%; 74.3%). The MGI estimates showed that PRS-informed approaches aligned more closely with national benchmarks than unweighted complete case analysis. Leveraging genetic data with sample weighting can help reduce bias in outcome-exposure association estimates derived from biobank data. When available, researchers should consider including PRS for imputation and survey methods for sample weighting when estimating outcome-exposure association coefficients in a target population of interest, recognizing that benefits may vary by outcome and data structure.
Intrahepatic cholestasis of pregnancy (ICP) is a pregnancy-specific liver disorder characterized by elevated total bile acid (TBA) levels, leading to adverse maternal and fetal outcomes. While genetic factors contribute to ICP and bile acid metabolism, the underlying mechanisms remain incompletely understood, particularly in East Asian populations. We conducted a genome-wide association study (GWAS) in 13,357 pregnant women from a Chinese cohort to investigate genetic determinants of TBA levels and ICP. Meta-analysis was performed by combining our data with the Shenzhen cohort. Post-GWAS analyses included pathway enrichment, integrative analysis with liver single-cell RNA sequencing (scRNA-seq) data, and Mendelian randomization (MR). We identified genome-wide significant associations at CYP7A1 (rs4738680, p = 1.08 × 10-10) and SLC39A9 (rs17107007, p = 3.46 × 10-28) for TBA, and at SLC39A9 (rs17107007, p = 3.34 × 10-16) for ICP. Pathway analysis highlighted bile acid synthesis and metabolism pathways for TBA and immune-related pathways for ICP. Integrative scRNA-seq analysis revealed enrichment of TBA-associated genes in hepatocytes and ICP-associated genes in neutrophils. MR analysis suggested a potential causal effect of estrone on TBA levels. Our findings provide novel insights into the genetic architecture of TBA and ICP, emphasizing the roles of bile acid metabolism in hepatocytes and immune dysregulation in ICP pathogenesis. The identified genetic loci and pathways may inform future research on risk prediction and targeted therapies for ICP.
Whole genome sequence (WGS) data provides opportunities for comprehensive evaluation of variants that may influence complex traits. However, prioritizing the large number of variants, particularly those in non-coding regions, is a challenge. Here we present an approach that uses pedigree-based haplotyping to identify the risk haplotype and resulting set of prioritized variants in a region of interest (ROI) defined by identity-by-descent (IBD) sharing among familial cases. The approach is applicable for use in both a full range of pedigree sizes and for the full allele frequency spectrum of variants without the need for a large reference sample. By determining haplotype sharing among individuals with WGS data, we demonstrate the ability to accurately identify a risk haplotype and a strongly reduced list of potential risk alleles for a trait of interest along with the cases who carry the risk haplotype. This is important in the context of complex traits where the disease may be etiologically heterogeneous even within a single pedigree. Application to both simulated and real Alzheimer's disease family data shows that the approach leads to accurate risk-haplotype identification with marked reduction in the number of potential trait-associated variants. Simulation also shows that the approach provides accurate risk haplotypes in ROIs.
Using principles of Mendelian genetics, probability theory, and mutation-specific knowledge, Mendelian risk prediction models identify those at high risk of carrying a heritable cancer susceptibility variant and assess future risk of cancer. Our previously-validated Fam3PRO model is a generalizable and computationally efficient Mendelian risk prediction framework that incorporates an arbitrary number of gene-cancer associations. In practice, from a model training perspective, there may be uncertainty in estimating the population-level model parameters necessary for rare gene-cancer associations. From a clinical perspective, it may be infeasible to obtain a detailed patient family history for many cancers. Motivated by the context of pre-screening for germline testing of a broad hereditary cancer gene panel, we propose a Mendelian model that aggregates information across genes and cancers, reducing patient burden and bypassing the need for robust parameter estimation for rare genes and syndromes. We evaluated this aggregate model through simulations and applied it to two independent clinical cohorts. We show that when the clinical goal is to assess patient risk of carrying a pathogenic variant for any cancer susceptibility gene, the aggregate model can give results comparable to a Mendelian model that considers many genes and cancers individually, while greatly simplifying model assumptions and user input.
In the last decade, genome-wide association studies (GWAS) have identified tens of thousands of common variants associated with a wide array of complex traits and diseases. Integration of GWAS with molecular data has informed the development of statistical tools for causal gene discovery. In this paper, we give an overview of commonly used causal inference methods and discuss the strengths and limitations of colocalization, Mendelian randomization (MR) and network-based approaches. Colocalization is often used to assess whether the genetic association signals for two traits arise from the same causal variant, thereby strengthening inferred causal associations. MR was developed to tackle issues of confounding and reverse causality, providing a rigorous approach to causal inference and demonstrating improved false discovery rates. Unlike MR, network-based analyses employ a discovery approach and model complex relationships between multiple variables. All causal inference methods are, to varying degrees, susceptible to spurious associations due to genetic confounding, pleiotropy and linkage disequilibrium. Here, we discuss the latest developments in the field of causal gene inference and limitations of these methods. We give an overview of interplay between different approaches as well as practical applications with reference to published examples in context of heart disease.
Accurate prediction of disease risk and other complex traits across different populations is essential for clinical and research purposes. However, genetic differences among ancestries, such as allelic frequencies and genetic architecture, can affect the performance of polygenic risk score (PGS) methods in cross-ancestry prediction. To address this issue, we conducted a formal test of seven polygenic prediction methods applicable across ancestries for five traits (BMI, standing height, LDL-, HDL- and total-cholesterol) from the UK Biobank dataset. We demonstrate that, GBLUP and PRS-CSx outperformed other methods for highly polygenic traits like height and BMI. In contrast, PRSice and PolyPred performed best for less polygenic traits like cholesterol, with PRS-CSx being comparable with larger sample sizes. We also observed that utilizing concordant SNPs, which have the same effect direction across diverse ancestries, can improve the accuracy of cross-ancestry PGS models. Furthermore, we found that the transferability of PGS across ancestries varied depending on the trait. Understanding the strengths and limitations of different methods and approaches is important for future methodological development and improvement, enabling better interpretation and application of PGS results in clinical and research settings.
A longstanding aim of developmental psychology and epidemiology is to understand the causal effects of parental phenotypes on offspring outcomes. Traditional approaches often fail to account for confounding and reverse causation. We evaluate the use of Mendelian randomisation with non-inherited variants (MR-NIV) to address these limitations. MR-NIV leverages non-inherited genetic variants to instrument the parental phenotype independent of the offspring's genotype. We used Directed Acyclic Graphs and simulations to validate MR-NIV and explore robustness to assortative mating. In contrast to an alternative MR method which adjusts the parental genotype for offspring genotype, MR-NIV can be robust to assortative mating when used without trio data. In settings without trio data, MR-NIV outperformed the adjustment method. The adjustment method outperformed MR-NIV in settings with trio data. Applying MR-NIV to the Avon Longitudinal Study of Parents and Children, we assessed the causal effect of parental smoking on offspring smoking initiation at age 16. Results were consistent with observational studies, suggesting a meaningful increase in the risk of offspring smoking due to parental smoking. However, larger sample sizes will be necessary to provide a precise answer. MR-NIV offers a promising extension of Mendelian randomisation for studying the developmental environment.
Summary-data-based multivariable Mendelian randomization (MVMR) methods, such as MVMR-Egger, MVMR-IVW, MVMR median-based, and MVMR-PRESSO, are used to assess the causal effects of multiple risk factors on disease. However, accounting for variances in the summary statistics of risk factors remains a challenge. We propose a linear mixed model with measurement error correction (LMM-MEC) that accounts for the variance in summary statistics for both disease outcomes and risk factors. First, under the NOME assumption, we apply a linear mixed model to account for variance in disease summary statistics by treating it as fixed- or random-effects, depending on whether there is heterogeneity in the effect sizes of the genetic variants on the disease outcome. Next, we relax the NOME assumption and further take the estimation error (or variance) in the summary statistics of risk factors into consideration by measurement models through a regression calibration approach. In a simulation study, using independent genetic variants as instrumental variables (IV), our method showed comparable performance to existing MVMR methods under conditions of no pleiotropy or with balanced pleiotropy on the disease outcome, and it achieved slightly improved coverage rates and power under directional pleiotropy. When genetic variants are in low to moderate linkage disequilibrium (LD) (0 < ρ $\rho $ 2 ≤ 0.3), our method showed comparable performance to MVMR-Egger, although both methods showed reduced coverage rates and power compared to situations where genetic variants as IVs are in LD. In the application study, we examined causal associations between correlated cholesterol biomarkers and longevity. By including 739 genetic variants selected based on p values < 5 × 10-5 from GWAS and allowing for low LD ( ρ $\rho $ 2 ≤ 0.1), our method identified that large LDL-c levels were causally associated with a lower likelihood of achieving longevity.
Genome-wide association studies (GWAS) have been instrumental in identifying genetic variants associated with complex traits and diseases, including Alzheimer's disease (AD). However, traditional GWAS approaches often focus on European populations, which may lead to loss of power and limit the generalizability of findings across diverse ancestries. On the other hand, LS-Imputation, a nonparametric trait imputation method, leverages GWAS summary statistics and genotype data to impute missing traits, which can then be used for GWAS and other downstream analyses. Although LS-Imputation has been applied successfully to European populations, its performance in non-European populations would be hindered by smaller sample sizes, leading to reduced imputation accuracy. To address these limitations, we propose two novel variants of LS-Imputation-LS-Imputation-Combined and LS-Imputation-Transfer-designed to integrate multi-ancestry GWAS data and enhance imputation performance. LS-Imputation-Combined optimally combines GWAS summary statistics from multiple ancestries, while LS-Imputation-Transfer sequentially refines imputed trait values across ancestries using stochastic gradient descent. We evaluate these methods using data from the UK Biobank and the Alzheimer's Disease Sequencing Project (ADSP), first applying them to high-density lipoprotein (HDL) cholesterol levels as a proof-of-concept before focusing on imputing AD status in Black individuals for genetic association analysis. Our results demonstrate that integrating multi-ancestry GWAS data improves trait imputation accuracy, with LS-Imputation-Transfer achieving the highest performance.
Environmental contexts may increase or decrease the heritability of a phenotype. Or equivalently, people with some genotypes may be more (or less) sensitive to environmental differences. Such genetic sensitivity to environmental contexts is called gene-environment interaction (G × E). While G × E has been robustly detected in twin and model organism studies for numerous phenotypes, there is a lingering perception that existing genome-wide G × E methods struggle to identify substantively significant and replicable interactions. We propose a novel method for examining G × E heritability using genetic marginal effects from genome-wide G × E analyses and Linkage Disequilibrium Score Regression (LDSC). We demonstrate the effectiveness of our method for body mass index (BMI) using biological sex (binary) and age (continuous) as moderators. Using the same procedures for both binary and continuous moderators, we detect robust evidence for G × E. Our results are consistent with findings from twin G × E studies of BMI and are more sensitive to environmental moderation than other LDSC-based methods. We conclude that BMI heritability is substantially more sensitive to variation in sex and age than is currently appreciated. Extending this method to other phenotype-moderator combinations has the potential to reveal G × E across numerous outcomes and moderators.
Complex disorders such as depression and alcohol use involve numerous genetic variants, and implicated loci continue to grow with sample size. This proliferation hampers interpretability, as the mechanisms by which so many variants jointly contribute to pathophysiology remain unclear. In contrast, classical Mendelian diseases arise from a single causal locus and are easier to interpret. We thus introduce Mendelianization-an algorithm distinct from Mendelian randomization-that learns weighted combinations of outcomes so that each aggregated phenotype concentrates association at one locus. We prove that this locus is causal under four structural assumptions natural to genetic data. The method handles partial sample overlap, provides calibrated hypothesis tests, maps coefficients to interpretable scales, and quantifies the degree of Mendelianism using summary z -statistics alone. In experiments, Mendelianization enhances statistical power to detect Mendelian symptom profiles even in heterogeneous disorders like major depression, generalized anxiety, and alcohol use disorder. An R implementation is available at github.com/ericstrobl/Mendelianization.
Mediation analysis is a pivotal tool for elucidating the indirect effect of an environmental factor or treatment on disease through potentially high-dimensional omics data, such as gene expression profiles. However, traditional mediation analysis methods tailored for binary outcomes often rely on the rare disease assumption in logistic regression and provide inadequate measures of total mediation effect when multiple mediators have effects in different directions. In this paper, we develop a MEdiation analysis framework in LOgistic regression for high-Dimensional mediators and a binarY outcome (MELODY). It leverages a second-moment-based measure analogous to the R 2 ${R}^{2}$ for linear models to quantify the total mediation effect. We also develop a variable selection procedure for high-dimensional data to reduce bias introduced by non-mediators. Our comprehensive simulations demonstrate the superior performance of MELODY in scenarios with non-rare disease binary outcomes and high-dimensional mediators. We apply MELODY to the Framingham Heart Study of over 5000 individuals to analyze the mediation effects of metabolomics and transcriptomics data on the pathways from sex to incident coronary heart disease.
Mendelian Randomization (MR) is a human genetics method for inferring causal relationships between risk factors and diseases. A common focus of MR studies has been on the causal inference of a single risk factor on a single disease. This has led to the successful discovery of numerous causal risk factors for disease. However, it remains unclear how much each causal risk factor contributes to disease collectively. Here, we introduce the concept of "causality explained", that provides an estimate of the causal variance explained by a phenome-wide set of risk factors on complex diseases to assess how much causality can be potentially explained. The model is based on principal component regression which is a multivariate linear regression based on principal component analysis. In complement, we propose the "polyfactorial index" to assess the trajectory of causality explained as risk factors are sequentially added into the model, to characterize the causal architecture for a complex disease. We demonstrate that our model correctly assesses the causality explained and causal architecture in simulations across a wide range of parameters. To build our model, we used a phenome-wide set of 222 traits from the UK Biobank compared to a set of 5 known risk factors for coronary artery disease. We observed that the phenome-wide set explains almot 45% of causality compared to 28.73% for the set of known risk factors. In addition, we tested our approach on 13 complex diseases and showed that the phenome-wide set can explain between 27% for anorexia to 80% for schizophrenia, with increasing trajectories of causality explained. We propose the "omnicausal model", which posits that a large number of risk factors explains a very small portion individually to disease but collectively explain most of the causal variance. We distinguished core and peripheral causal factors that explain respectively a larger and a smaller part of causal variance. This approach provides insights into the relative importance of individual risk factors as well as the collective impact of multiple causal risk factors on disease, providing insights into the causal architecture of disease.