It is increasingly recognized that participation bias can pose problems for genetic studies. Recently, to overcome the challenge that genetic information of nonparticipants is unavailable, it is shown that by comparing the IBD (identity by descent) shared and not-shared segments between participating relative pairs, one can estimate the genetic component underlying participation. That, however, does not directly address how to adjust estimates of heritability and genetic correlation for phenotypes correlated with participation. Here, we demonstrate a way to do so by adopting a statistical framework that separates the genetic and nongenetic correlations between participation and these phenotypes. Crucially, our method avoids making the assumption that the effect of the genetic component underlying participation is manifested entirely through these other phenotypes. Applying the method to 12 UK Biobank phenotypes, we found eight that have significant genetic correlations with participation, including body mass index, educational attainment, and smoking status. For most of these phenotypes, without adjustments, estimates of heritability and the absolute value of genetic correlation would have underestimation biases.
Purpose: To examine the effects of serum growth hormone (GH) and insulin-like growth factor-1 (IGF-1) on choroidal structures with different blood glucose levels in patients with diabetes mellitus (DM) with acromegaly without diabetic retinopathy. Methods: Eighty-eight eyes of 44 patients with acromegaly were divided into a nondiabetic group (23 patients, 46 eyes) and a diabetic group (21 patients, 42 eyes). Forty-four age- and sex-matched healthy controls and 21 patients with type 2 DM without diabetic retinopathy were also included. Linear regression models with a simple slope analysis were used to identify the correlation and interaction between endocrine parameters and choroidal thickness (ChT), total choroidal area (TCA), luminal area (LA), stromal area (SA), and choroidal vascular index (CVI). Results: Our study revealed significant increases in the ChT, LA, SA, and TCA in patients with acromegaly compared with healthy controls, with no difference in the CVI. Comparatively, patients with DM with acromegaly had greater ChT than matched patients with type 2 DM, with no significant differences in other choroidal parameters. The enhancement of SA, LA and TCA caused by an acromegalic status disappeared in patients with diabetic status, whereas ChT and CVI were not affected by the interaction. In the diabetic acromegaly, higher IGF-1 (P = 0.006) and GH levels (P = 0.049), longer DM duration (P = 0.007), lower blood glucose (P = 0.001), and the interaction between GH and blood glucose were associated independently with thicker ChT. Higher GH levels (P = 0.016, 0.004 and 0.007), longer DM duration (P = 0.022, 0.013 and 0.013), lower blood glucose (P = 0.034, 0.011 and 0.01), and the interaction of IGF-1 and blood glucose were associated independently with larger SA, LA, and TCA. As blood glucose levels increased, the positive correlation between serum GH level and ChT diminished, and became insignificant when blood glucose was more than 7.35 mM/L. The associations between serum IGF-1 levels and LA, SA, and TCA became increasingly negative, with LA, becoming signif- icantly and negatively associated to the GH levels only when blood glucose levels were more than 8.59 mM/L. Conclusions: Acromegaly-related choroidal enhancements diminish in the presence of DM. In diabetic acromegaly, blood glucose levels are linked negatively with changes in choroidal metrics and their association with GH and IGF-1. Translational Relevance: We revealed the potential beneficial impacts of IGF-1 and GH on structural measures of the choroid in patients with DM at relatively well-controlled blood glucose level, which could provide a potential treatment target for diabetic retinopathy.
TWAS have shown great promise in extending GWAS loci to a functional understanding of disease mechanisms. In an effort to fully unleash the TWAS and GWAS information, we propose MTWAS, a statistical framework that partitions and aggregates cross-tissue and tissue-specific genetic effects in identifying gene-trait associations. We introduce a non-parametric imputation strategy to augment the inaccessible tissues, accommodating complex interactions and non-linear expression data structures across various tissues. We further classify eQTLs into cross-tissue eQTLs and tissue-specific eQTLs via a stepwise procedure based on the extended Bayesian information criterion, which is consistent under high-dimensional settings. We show that MTWAS significantly improves the prediction accuracy across all 47 tissues of the GTEx dataset, compared with other single-tissue and multi-tissue methods, such as PrediXcan, TIGAR, and UTMOST. Applying MTWAS to the DICE and OneK1K datasets with bulk and single-cell RNA sequencing data on immune cell types showcases consistent improvements in prediction accuracy. MTWAS also identifies more predictable genes, and the improvement can be replicated with independent studies. We apply MTWAS to 84 UK Biobank GWAS studies, which provides insights into disease etiology.
Abstract Mendelian randomization (MR) is a statistical approach to inferring the causal relationships from genome-wide association studies (GWAS) by using genetic variants as instrumental variables (IVs). As IVs, the selected genetic variants should be solely associated with the exposure of interest, and have no associations with confounders and the studied outcome except through the exposure. Sometimes the selected genetic variants have effects on the outcome through other pathways, a phenomenon known as horizontal pleiotropy, which makes the selected variants violate IV assumptions. Two different approaches have been proposed to address this issue: one is to improve robustness to pleiotropic effects, and the other one is to generalize univariate exposure to multivariate cases so that we can incorporate possible pathways into MR analysis. Compared to pleiotropy-robust methods, multivariable Mendelian randomization (MVMR) can uncover the exposures having direct effects on the outcome. However, measuring all possible pathways from genetic variants to the outcome in MVMR analysis is difficult. Although MVMR methods with robustness to unmeasured pleiotropy have been proposed recently, they are statistically inefficient as a result of ignoring correlations among exposures and distributions of effect sizes. Given the limitations from both directions, we propose a novel method named MVMR-PRESS that can infer causal relationships between multivariate exposures and the outcome with robustness to horizontal pleiotropic effects from unconsidered pathways. MVMR-PRESS estimates causal effects by using a Penalized Regression on Summary Statistics from GWAS considering both correlations among exposures and distributions of effect sizes. One merit of MVMR-PRESS is that samples in GWAS of different exposures can have overlaps, which allows us to include GWAS summary statistics from the same cohort or consortium. Simulation experiments showed that our method achieved the smallest bias and highest power compared to existing pleiotropy-robust MR and MVMR methods while the type 1 error rate was well-controlled. Applying MVMR-PRESS to publicly available GWAS summary statistics demonstrated that body mass index (BMI), height (HT), and low-density lipoprotein cholesterol (LDL) have significant causal effects on coronary artery disease (CAD), and BMI and high-density lipoprotein cholesterol (HDL) were causally related to type 2 diabetes (T2D). On the contrary, the causal estimates of triglycerides (TG), HDL, and total cholesterol (TC) on CAD, and the estimates of HT, LDL, TG, and TC on T2D were non-significant. In addition, we found no evidence suggesting BMI, HT, and lipid levels have causal effects on inflammatory bowel disease (IBD) and schizophrenia (SCZ).
Heritability is a fundamental concept in genetic studies, measuring the genetic contribution to complex traits and bringing insights about disease mechanisms. The advance of high-throughput technologies has provided many resources for heritability estimation. Linkage disequilibrium (LD) score regression (LDSC) estimates both heritability and confounding biases, such as cryptic relatedness and population stratification, among single-nucleotide polymorphisms (SNPs) by using only summary statistics released from genome-wide association studies. However, only partial information in the LD matrix is utilized in LDSC, leading to loss in precision. In this study, we propose LD eigenvalue regression (LDER), an extension of LDSC, by making full use of the LD information. Compared to state-of-the-art heritability estimating methods, LDER provides more accurate estimates of SNP heritability and better distinguishes the inflation caused by polygenicity and confounding effects. We demonstrate the advantages of LDER both theoretically and with extensive simulations. We applied LDER to 814 complex traits from UK Biobank, and LDER identified 363 significantly heritable phenotypes, among which 97 were not identified by LDSC.
Accurate estimate of relatedness is important for genetic data analyses, such as heritability estimation and association mapping based on data collected from genome-wide association studies. Inaccurate relatedness estimates may lead to biased heritability estimations and spurious associations. Individual-level genotype data are often used to estimate kinship coefficient between individuals. The commonly used sample correlation-based genomic relationship matrix (scGRM) method estimates kinship coefficient by calculating the average sample correlation coefficient among all single nucleotide polymorphisms (SNPs), where the observed allele frequencies are used to calculate both the expectations and variances of genotypes. Although this method is widely used, a substantial proportion of estimated kinship coefficients are negative, which are difficult to interpret. In this paper, through mathematical derivation, we show that there indeed exists bias in the estimated kinship coefficient using the scGRM method when the observed allele frequencies are regarded as true frequencies. This leads to negative bias for the average estimate of kinship among all individuals, which explains the estimated negative kinship coefficients. Based on this observation, we propose an unbiased estimation method, UKin, which can reduce kinship estimation bias. We justify our improved method with rigorous mathematical proof. We have conducted simulations as well as two real data analyses to compare UKin with scGRM and three other kinship estimating methods: rGRM, tsGRM, and KING. Our results demonstrate that both bias and root mean square error in kinship coefficient estimation could be reduced by using UKin. We further investigated the performance of UKin, KING, and three GRM-based methods in calculating the SNP-based heritability, and show that UKin can improve estimation accuracy for heritability regardless of the scale of SNP panel.
Recent advances in single-cell technologies have enabled the characterization of epigenomic heterogeneity at the cellular level. Computational methods for automatic cell type annotation are urgently needed given the exponential growth in the number of cells. In particular, annotation of single-cell chromatin accessibility sequencing (scCAS) data, which can capture the chromatin regulatory landscape that governs transcription in each cell type, has not been fully investigated. Here we propose EpiAnno, a probabilistic generative model integrated with a Bayesian neural network, to annotate scCAS data automatically in a supervised manner. We systematically validate the superior performance of EpiAnno for both intra- and inter-dataset annotation on various datasets. We further demonstrate the advantages of EpiAnno for interpretable embedding and biological implications via expression enrichment analysis, partitioned heritability analysis, enhancer identification, cis-coaccessibility analysis and pathway enrichment analysis. In addition, we show that EpiAnno has the potential to reveal cell type-specific motifs and facilitate scCAS data simulation. The investigation of single-cell epigenomics with technologies such as single-cell chromatin accessibility sequencing (scCAS) presents an opportunity to expand the understanding of gene regulation at the cellular level. The authors develop a probabilistic generative model to better characterize cell heterogeneity and accurately annotate the cell type of scCAS data.
Openness-weighted association study (OWAS) is a method that leverages the in silico prediction of chromatin accessibility to prioritize genome-wide association studies (GWAS) signals, and can provide novel insights into the roles of non-coding variants in complex diseases. A prerequisite to apply OWAS is to choose a trait-related cell type beforehand. However, for most complex traits, the trait-relevant cell types remain elusive. In addition, many complex traits involve multiple related cell types. To address these issues, we develop OWAS-joint, an efficient framework that aggregates predicted chromatin accessibility across multiple cell types, to prioritize disease-associated genomic segments. In simulation studies, we demonstrate that OWAS-joint achieves a greater statistical power compared to OWAS. Moreover, the heritability explained by OWAS-joint segments is higher than or comparable to OWAS segments. OWAS-joint segments also have high replication rates in independent replication cohorts. Applying the method to six complex human traits, we demonstrate the advantages of OWAS-joint over a single-cell-type OWAS approach. We highlight that OWAS-joint enhances the biological interpretation of disease mechanisms, especially for non-coding regions.
Abstract Accurate estimate of relatedness is important for genetic data analyses, such asheritability estimation and association mapping based on data collected fromgenome-wide association studies. Inaccurate relatedness estimates may lead tobiased heritability estimations and spurious associations. Individual-level genotypedata are often used to estimate kinship coefficient between individuals. Thecommonly used sample correlation-based genomic relationship matrix (scGRM)method estimates kinship coefficient by calculating the average samplecorrelation coefficient among all single nucleotide polymorphisms (SNPs), wherethe observed allele frequencies are used to calculate both the expectations andvariances of genotypes. Although this method is widely used, a substantialproportion of estimated kinship coefficients are negative, which are difficult tointerpret. In this paper, through mathematical derivation, we show that thereindeed exists bias in the estimated kinship coefficient using the scGRM methodwhen the observed allele frequencies are regarded as true frequencies. This leadsto negative bias for the average estimate of kinship among all individuals, whichexplains the estimated negative kinship coefficients. Based on this observation,we propose an unbiased estimation method, UKin, which can reduce kinshipestimation bias. We justify our improved method with rigorous mathematicalproof. We have conducted simulations as well as two real data analyses tocompare UKin with scGRM and another widely used kinship estimating method:KING. Our results demonstrate that both bias and root mean square error inkinship coefficient estimation could be reduced by using UKin. We furtherinvestigated the performance of UKin, scGRM and KING in calculating theSNP-based heritability, and show that UKin can improve estimation accuracy forheritability regardless of the scale of SNP panel.
MOTIVATION:Polygenic risk score (PRS) has been widely exploited for genetic risk prediction due to its accuracy and conceptual simplicity. We introduce a unified Bayesian regression framework, NeuPred, for PRS construction, which accommodates varying genetic architectures and improves overall prediction accuracy for complex diseases by allowing for a wide class of prior choices. To take full advantage of the framework, we propose a summary-statistics-based cross-validation strategy to automatically select suitable chromosome-level priors, which demonstrates a striking variability of the prior preference of each chromosome, for the same complex disease, and further significantly improves the prediction accuracy. RESULTS:Simulation studies and real data applications with seven disease datasets from the Wellcome Trust Case Control Consortium cohort and eight groups of large-scale genome-wide association studies demonstrate that NeuPred achieves substantial and consistent improvements in terms of predictive r2 over existing methods. In addition, NeuPred has similar or advantageous computational efficiency compared with the state-of-the-art Bayesian methods. AVAILABILITY AND IMPLEMENTATION:The R package implementing NeuPred is available at https://github.com/shuangsong0110/NeuPred. SUPPLEMENTARY INFORMATION:Supplementary data are available at Bioinformatics online.
We study the law of the iterated logarithm (LIL) for the maximum likelihood estimation of the parameters (as a convex optimization problem) in the generalized linear models with independent or weakly dependent (ρ-mixing) responses under mild conditions. The LIL is useful to derive the asymptotic bounds for the discrepancy between the empirical process of the log-likelihood function and the true log-likelihood. The strong consistency of some penalized likelihood-based model selection criteria can be shown as an application of the LIL. Under some regularity conditions, the model selection criterion will be helpful to select the simplest correct model almost surely when the penalty term increases with the model dimension, and the penalty term has an order higher than O(log log n) but lower than O(n): Simulation studies are implemented to verify the selection consistency of Bayesian information criterion.
Accurate estimate of relatedness is important for genetic data analyses, such as association mapping and heritability estimation based on data collected from genome-wide association studies. Inaccurate relatedness estimates may lead to spurious associations and biased heritability estimations. Individual-level genotype data are often used to estimate kinship coefficient between individuals. The commonly used sample correlation-based genomic relationship matrix (scGRM) method estimates kinship coefficient by calculating the average sample correlation coefficient among all single nucleotide polymorphisms (SNPs), where the observed allele frequencies are used to calculate both the expectations and variances of genotypes. Although this method is widely used, a substantial proportion of estimated kinship coefficients are negative, which are difficult to interpret. In this paper, through mathematical derivation, we show that there indeed exists bias in the estimated kinship coefficient using the scGRM method when the observed allele frequencies are regarded as true frequencies. This leads to negative bias for the average estimate of kinship among all individuals, which explains the estimated negative kinship coefficients. Based on this observation, we propose an unbiased estimation method, UKin, which can reduce the bias. We justify our improved method with rigorous mathematical proof. We have conducted simulations as well as two real data analyses to demonstrate that both bias and root mean square error in kinship coefficient estimation can be reduced by using UKin. Further simulations indicate that the power in association mapping can also be improved by using our unbiased kinship estimates to adjust for cryptic relatedness. Author summary Inference of relatedness plays an important role in genetic data analysis. Many methods have been proposed to estimate kinship coefficients, including the commonly used genomic relationship matrix method. However, a substantial proportion of the kinship coefficients estimated by this method are negative, which is difficult to interpret. In this paper, through mathematical derivation, we show that there indeed exists a negative bias in this approach. To correct for this bias, we propose a new kinship coefficient estimation method, UKin, which is unbiased without requiring extra genetic information nor added computational complexity. The better performance of UKin in reducing bias and root mean squared error is demonstrated through theory, simulations and analysis of data from the young-onset breast cancer and familial intracranial aneurysm studies.
High-dimensional correlated binary data arise in many areas, such as observed genetic variations in biomedical research. Data simulation can help researchers evaluate efficiency and explore properties of different computational and statistical methods. Also, some statistical methods, such as Monte Carlo methods, rely on data simulation. Lunn and Davies proposed linear time complexity methods to generate correlated binary variables with three common correlation structures. However, it is infeasible to specify unequal probabilities in their methods. In this article, we introduce several computationally efficient algorithms that generate high-dimensional binary data with specified correlation structures and unequal probabilities. Our algorithms have linear time complexity with respect to the dimension for three commonly studied correlation structures, namely exchangeable, decaying-product and K-dependent correlation structures. In addition, we extend our algorithms to generate binary data of specified nonnegative correlation matrices satisfying the validity condition with quadratic time complexity. We provide an R package, CorBin, to implement our simulation methods. Compared to the existing packages for binary data generation, the time cost to generate a 100-dimensional binary vector with the common correlation structures and general correlation matrices can be reduced up to 105 folds and 103 folds, respectively, and the efficiency can be further improved with the increase of dimensions. The R package CorBin is available on CRAN at https://cran.r-project.org/.
Mobility restrictions have been a heated topic during the global pandemic of coronavirus disease 2019 (COVID-19). However, multiple recent findings have verified its importance in blocking virus spread. Evidence on the association between mobility, cases imported from abroad and local medical resource supplies is limited. To reveal the association, this study quantified the importance of inter- and intra-country mobility in containing virus spread and avoiding hospitalizations during early stages of COVID-19 outbreaks in India, Japan, and China. We calculated the time-varying reproductive number (Rt) and duration from illness onset to diagnosis confirmation (Doc), to represent conditions of virus spread and hospital bed shortages, respectively. Results showed that inter-country mobility fluctuation could explain 80%, 35%, and 12% of the variance in imported cases and could prevent 20 million, 5 million, and 40 million imported cases in India, Japan and China, respectively. The critical time for screening and monitoring of imported cases is 2 weeks at minimum and 4 weeks at maximum, according to the time when the Pearson’s Rs between Rt and imported cases reaches a peak (>0.8). We also found that if local transmission is initiated, a 1% increase in intra-country mobility would result in 1430 (±501), 109 (±181), and 10 (±1) additional bed shortages, as estimated using the Doc in India, Japan, and China, respectively. Our findings provide vital reference for governments to tailor their pre-vaccination policies regarding mobility, especially during future epidemic waves of COVID-19 or similar severe epidemic outbreaks.
Recently polygenetic risk score (PRS) has been successfully used in the risk prediction of complex human diseases. Many studies incorporated internal information, such as effect size distribution, or external information, such as linkage disequilibrium, functional annotation, and pleiotropy among multiple diseases, to optimize the performance of PRS. To leverage on multiomics datasets, we developed a novel flexible transcriptional risk score (TRS), in which messenger RNA expression levels were imputed and weighted for risk prediction. In simulation studies, we demonstrated that single‐tissue TRS has greater prediction power than LDpred, especially when there is a large effect of gene expression on the phenotype. Multitissue TRS improves prediction accuracy when there are multiple tissues with independent contributions to disease risk. We applied our method to complex traits, including Crohn's disease, type 2 diabetes, and so on. The single‐tissue TRS method outperformed LDpred and AnnoPred across the tested traits. The performance of multitissue TRS is trait‐dependent. Moreover, our method can easily incorporate information from epigenomic and proteomic data upon the availability of reference datasets.
AbstractMotivationIdentification and interpretation of non-coding variations that affect disease risk remain a paramount challenge in genome-wide association studies (GWAS) of complex diseases. Experimental efforts have provided comprehensive annotations of functional elements in the human genome. On the other hand, advances in computational biology, especially machine learning approaches, have facilitated accurate predictions of cell-type-specific functional annotations. Integrating functional annotations with GWAS signals has advanced the understanding of disease mechanisms. In previous studies, functional annotations were treated as static of a genomic region, ignoring potential functional differences imposed by different genotypes across individuals.ResultsWe develop a computational approach, Openness Weighted Association Studies (OWAS), to leverage and aggregate predictions of chromosome accessibility in personal genomes for prioritizing GWAS signals. The approach relies on an analytical expression we derived for identifying disease associated genomic segments whose effects in the etiology of complex diseases are evaluated. In extensive simulations and real data analysis, OWAS identifies genes/segments that explain more heritability than existing methods, and has a better replication rate in independent cohorts than GWAS. Moreover, the identified genes/segments show tissue-specific patterns and are enriched in disease relevant pathways. We use rheumatic arthritis and asthma as examples to demonstrate how OWAS can be exploited to provide novel insights on complex diseases.Availability and implementationThe R package OWAS that implements our method is available at https://github.com/shuangsong0110/OWAS.Supplementary informationSupplementary data are available at Bioinformatics online.
Genetic risk prediction is an important problem in human genetics, and accurate prediction can facilitate disease prevention and treatment. Calculating polygenic risk score (PRS) has become widely used due to its simplicity and effectiveness, where only summary statistics from genome-wide association studies are needed in the standard method. Recently, several methods have been proposed to improve standard PRS by utilizing external information, such as linkage disequilibrium and functional annotations. In this paper, we introduce EB-PRS, a novel method that leverages information for effect sizes across all the markers to improve prediction accuracy. Compared to most existing genetic risk prediction methods, our method does not need to tune parameters nor external information. Real data applications on six diseases, including asthma, breast cancer, celiac disease, Crohn's disease, Parkinson's disease and type 2 diabetes show that EB-PRS achieved 307.1%, 42.8%, 25.5%, 3.1%, 74.3% and 49.6% relative improvements in terms of predictive r2 over standard PRS method with optimally tuned parameters. Besides, compared to LDpred that makes use of LD information, EB-PRS also achieved 37.9%, 33.6%, 8.6%, 36.2%, 40.6% and 10.8% relative improvements. We note that our method is not the first method leveraging effect size distributions. Here we first justify our method by presenting theoretical optimal property over existing methods in this class of methods, and substantiate our theoretical result with extensive simulation results. The R-package EBPRS that implements our method is available on CRAN.
This paper proposes a novel underactuated hand, COSA-FBA hand, which performs coupled and self-adaptive grasping. The COSA-FBA hand consists of 5 fingers and 14 degrees of freedom. The COSA-FBA finger consists of an actuator embedded into its proximal phalange, a five-gear mechanism and a spring, which allows it to finish multiple grasping postures such as enveloping and pinching, depending on the shape, size, and dimension of the object. The COSA-FBA hand has advantages of concise structure, low cost, high efficiency and space-saving of the palm. The coupled and self-adaptive hand has a wide arrange of applications.
Traditional underactuated fingers with the coupled and self-adaptive (COSA) grasping mode cannot parallel grip objects; fingers with the parallel and self-adaptive (PASA) grasping mode cannot hook objects. This paper reviewed all grasping modes of underactuated fingers and proposes a novel switching concept to combine the COSA mode and the PASA mode. This paper designs a novel underactuated robotic finger with parallel-coupled switchable (PCSS) mode, called PCSS finger. The PCSS finger consists of a pulley-belt speed-up transmission, an open-loop tendon transmission, two springs, a mechanical limit and a switching mechanism. The PCSS finger can conduct either the PASA grasp or the COSA grasp, being able to grasp a larger range of objects than traditional COSA fingers and PASA fingers.