There are primarily two computational approaches to alternative splicing (AS) detection using short reads: splice junction-based and exon-based approaches. Despite their shared goal of addressing the same biological problem, these approaches have not been reconciled before. We devised a novel graph structure and algorithm aimed at mapping between the exonic parts and splicing events detected by the two different methods. Through simulations, we demonstrated disparities in sensitivity and specificity between splice junction-based and exon-based methods. When applied to empirical data, there were large discrepancies in the results, suggesting that the methods are complementary. With the discrepancies localized to individual events and exonic parts, we were able to gain insights into the strengths and weaknesses inherent in each approach. Finally, we integrated the results to generate a comprehensive list of both common and unique AS events detected by both methodologies.
Spatial barcoding-based transcriptomic (ST) data require deconvolution for cellular-level downstream analysis. Here we present SDePER, a hybrid machine learning and regression method to deconvolve ST data using reference single-cell RNA sequencing (scRNA-seq) data. SDePER tackles platform effects between ST and scRNA-seq data, ensuring a linear relationship between them while addressing sparsity and spatial correlations in cell types across capture spots. SDePER estimates cell-type proportions, enabling enhanced resolution tissue mapping by imputing cell-type compositions and gene expressions at unmeasured locations. Applications to simulated data and four real datasets showed SDePER’s superior accuracy and robustness over existing methods.
Background:The development and progression of Alzheimer's disease (AD) is a complex process that can change over time, during which genetic influences on phenotypes may also fluctuate. Incorporating longitudinal phenotypes in genome wide association studies (GWAS) could help unmask genetic loci with time-varying effects. In this study, we incorporated a varying coefficient test in a longitudinal GWAS model to identify single nucleotide polymorphisms (SNPs) that may have time- or age-dependent effects in AD. Methods:Genotype data from 1,877 participants in the Alzheimer's Neuroimaging Data Initiative (ADNI) were imputed using the Haplotype Reference Consortium (HRC) panel, resulting in 9,573,130 SNPs. Subjects' longitudinal impairment status at each visit was considered as a binary and clinical phenotype. Participants' composite standardized uptake value ratio (SUVR) derived from each longitudinal amyloid PET scan was considered as a continuous and biological phenotype. The retrospective varying coefficient mixed model association test (RVMMAT) was used in longitudinal GWAS to detect time-varying genetic effects on the impairment status and SUVR measures. Post-hoc analyses were performed on genome-wide significant SNPs, including 1) pathway analyses; 2) age-stratified genotypic comparisons and regression analyses; and 3) replication analyses using data from the National Alzheimer's Coordinating Center (NACC). Results:Our model identified 244 genome-wide significant SNPs that revealed time-varying genetic effects on the clinical impairment status in AD; among which, 12 SNPs on chromosome 19 were successfully replicated using data from NACC. Post-hoc age-stratified analyses indicated that for most of these 244 SNPs, the maximum genotypic effect on impairment status occurred between 70 to 80 years old, and then declined with age. Our model further identified 73 genome-wide significant SNPs associated with the temporal variation of amyloid accumulation. For these SNPs, an increasing genotypic effect on PET-SUVR was observed as participants' age increased. Functional pathway analyses on significant SNPs for both phenotypes highlighted the involvement and disruption of immune responses- and neuroinflammation-related pathways in AD. Conclusion:We demonstrate that longitudinal GWAS models with time-varying coefficients can boost the statistical power in AD-GWAS. In addition, our analyses uncovered potential time-varying genetic variants on repeated measurements of clinical and biological phenotypes in AD.
Many genetic studies contain rich information on longitudinal phenotypes that require powerful analytical tools for optimal analysis. Genetic analysis of longitudinal data that incorporates temporal variation is important for understanding the genetic architecture and biological variation of complex diseases. Most of the existing methods assume that the contribution of genetic variants is constant over time and fail to capture the dynamic pattern of disease progression. However, the relative influence of genetic variants on complex traits fluctuates over time. In this study, we propose a retrospective varying coefficient mixed model association test, RVMMAT, to detect time-varying genetic effect on longitudinal binary traits. We model dynamic genetic effect using smoothing splines, estimate model parameters by maximizing a double penalized quasi-likelihood function, design a joint test using a Cauchy combination method, and evaluate statistical significance via a retrospective approach to achieve robustness to model misspecification. Through simulations, we illustrated that the retrospective varying-coefficient test was robust to model misspecification under different ascertainment schemes and gained power over the association methods assuming constant genetic effect. We applied RVMMAT to a genome-wide association analysis of longitudinal measure of hypertension in the Multi-Ethnic Study of Atherosclerosis. Pathway analysis identified two important pathways related to G-protein signaling and DNA damage. Our results demonstrated that RVMMAT could detect biologically relevant loci and pathways in a genome scan and provided insight into the genetic architecture of hypertension.
As human complex diseases are influenced by the interplay of genes and environment, detecting gene-environment interactions (G×E) can shed light on biological mechanisms of diseases and play an important role in disease risk prediction. Development of powerful quantitative tools to incorporate G×E in complex diseases has potential to facilitate the accurate curation and analysis of large genetic epidemiological studies. However, most of existing methods that interrogate G×E focus on the interaction effects of an environmental factor and genetic variants, exclusively for common or rare variants. In this study, we proposed two tests, MAGEIT_RAN and MAGEIT_FIX, to detect interaction effects of an environmental factor and a set of genetic markers containing both rare and common variants, based on the MinQue for Summary statistics. The genetic main effects in MAGEIT_RAN and MAGEIT_FIX are modeled as random or fixed, respectively. Through simulation studies, we illustrated that both tests had type I error under control and MAGEIT_RAN was overall the most powerful test. We applied MAGEIT to a genome-wide analysis of gene-alcohol interactions on hypertension in the Multi-Ethnic Study of Atherosclerosis. We detected two genes, CCNDBP1 and EPB42, that interact with alcohol usage to influence blood pressure. Pathway analysis identified sixteen significant pathways related to signal transduction and development that were associated with hypertension, and several of them were reported to have an interactive effect with alcohol intake. Our results demonstrated that MAGEIT can detect biologically relevant genes that interact with environmental factors to influence complex traits.
During the early phase of the COVID-19 pandemic, infected patients presented with symptoms similar to bacterial pneumonias and were treated with antibiotics before confirmation of a bacterial or fungal co-infection. We reasoned that wastewater surveillance could reveal potential relationships between reduced antimicrobial stewardship, specifically misprescribing antibiotics to treat viral infections, and the occurrence of antimicrobial resistance (AMR) in an urban community. Here, we analyzed microbial communities and AMR profiles in sewage samples from a wastewater treatment plant (WWTP) and a community shelter in Las Vegas, Nevada during a COVID-19 surge in December 2020. Using a respiratory pathogen and AMR enrichment next-generation sequencing panel, we identified four major phyla in the wastewater, including Actinobacteria, Firmicutes, Bacteroidetes and Proteobacteria. Consistent with antibiotics that were reportedly used to treat COVID-19 infections (e.g., fluoroquinolones and beta-lactams), we also measured a significant spike in corresponding AMR genes in the wastewater samples. AMR genes associated with colistin resistance (mcr genes) were also identified exclusively at the WWTP, suggesting that multidrug resistant bacterial infections were being treated during this time. We next compared the Las Vegas sewage data to local 2018–2019 antibiograms, which are antimicrobial susceptibility profile reports about common clinical pathogens. Similar to the discovery of higher levels of beta-lactamase resistance genes in sewage during 2020, beta-lactam antibiotics accounted for 51 ± 3 % of reported antibiotics used in antimicrobial susceptibility tests of 2018–2019 clinical isolates. Our data highlight how wastewater-based epidemiology (WBE) can be leveraged to complement more traditional surveillance efforts by providing community-level data to help identify current and emerging AMR threats.
Neoantigens are derived from tumor-specific somatic mutations. Neoantigen-based synthesized peptides have been under clinical investigation to boost cancer immunotherapy efficacy. The promising results prompt us to further elucidate the effect of neoantigen expression on patient survival in breast cancer. We applied Kaplan–Meier survival and multivariable Cox regression models to evaluate the effect of neoantigen expression and its interaction with T-cell activation on overall survival in a cohort of 729 breast cancer patients. Pearson’s chi-squared tests were used to assess the relationships between neoantigen expression and clinical pathological variables. Spearman correlation analysis was conducted to identify correlations between neoantigen expression, mutation load, and DNA repair gene expression. ERCC1, XPA, and XPC were negatively associated with neoantigen expression, while BLM, BRCA2, MSH2, XRCC2, RAD51, CHEK1, and CHEK2 were positively associated with neoantigen expression. Based on the multivariable Cox proportional hazard model, patients with a high level of neoantigen expression and activated T-cell status showed improved overall survival. Similarly, in the T-cell exhaustion and progesterone receptor (PR) positive subgroups, patients with a high level of neoantigen expression showed prolonged survival. In contrast, there was no significant difference in the T-cell activation and PR negative subgroups. In conclusion, neoantigens may serve as immunogenic agents for immunotherapy in breast cancer.
A number of deep neural network (DNN)-based models have been applied to help classify and detect severe cardiovascular diseases using 12-lead electrocardiogram (ECG) signals. These models, however, suffer from poor performance in detecting one or two specific cardiac abnormalities, like ST-segment abnormalities, with their accuracy lingering at only around 60%, which in turn limits their applicability in clinical practice. In this paper, we show that the convolution layers of these DNN models can cause the diminishment of the key features of ST-segment abnormalities, making it, compared to cardiac arrhythmia and normal sinus rhythm (NSR), hard to be classified out from the ECG. Correspondingly, we, in this paper, propose a novel DNN-based model that is able to achieve high accuracy of detecting 9 classes of rhythms. In specific, the moving averages of the ECG signals are used as the second input to the proposed model so that it can help fully mine the features relevant to the ST abnormalities. In addition, as opposed to the convolution layers, our model utilizes the inception-residual layers to preserve features from the shallow layers for reuse in the deep layers. On top of these network architecture improvements, we further introduce a customized differential layer so that all the relevant features can be preserved and/or amplified for classification purposes. Trained with CPSC 2018 dataset, our proposed model is able to accurately classify the 9 rhythm classes, with an F1 score as high as 89.7%, up by merely 4.2% from the best result reported in the literature. As far as the ST segment abnormalities are concerned, the proposed model achieves an 89.6% F1 score for STD and an 80.8% F1 score for STE, which is 7.8% and 13.1% respectively higher than that of the best result. The proposed method is thus poised to become a viable solution for cardiovascular health monitoring with the increasing availability of portable and home-based ECG devices.
The effect of immunosuppression blockade therapies depends on the infiltration of effector T cells and other immune cells in tumor. However, it is unclear how molecular pathways regulate the infiltration of immune cells, as well as how interactions between tumor-infiltrating immune cells and T cell activation affect breast cancer patient survival. CIBERSORT was used to estimate the relative abundance of 22 immune cell types. The association between mRNAs and immune cell abundance were assessed by Spearman correlation analysis. Enriched pathways were identified using MetaCore pathway analysis. The interactions between the T cell activation status and the abundance of tumor-infiltrating immune cells were evaluated using Kaplan-Meier survival and multivariate Cox regression models in a publicly available dataset of 1081 breast cancer patients. The role of tumor-infiltrating B cells in antitumor immunity, immune response of T cell subsets, and breakdown of CD4+ T cell peripheral tolerance were positively associated with M1 macrophage and CD8+ T cell but negatively associated with M2 macrophage. Abundant plasma cell was associated with prolonged survival (HR = 0.46, 95% CI: 0.32-0.67), and abundant M2 macrophage was associated with shortened survival (HR = 1.78, 95% CI: 1.23-2.60). There exists a significant interaction between the T cell activation status and the resting DC abundance level (p = 0.025). Molecular pathways associated with tumor-infiltrating immune cells provide future directions for developing cancer immunotherapies to control immune cell infiltration, and further influence T cell activation and patient survival in breast cancer.
Effector CD8+ T cell activation and its cytotoxic function are positively correlated with improved survival in breast cancer. tRNA-derived fragments (tRFs) have recently been found to be involved in gene regulation in cancer progression. However, it is unclear how interactions between expression of tRFs and T cell activation affect breast cancer patient survival. We used Kaplan–Meier survival and multivariate Cox regression models to evaluate the effect of interactions between expression of tRFs and T cell activation on survival in 1081 breast cancer patients. Spearman correlation analysis and weighted gene co-expression network analysis were conducted to identify genes and pathways that were associated with tRFs. tRFdb-5024a, 5P_tRNA-Leu-CAA-4-1, and ts-49 were positively associated with overall survival, while ts-34 and ts-58 were negatively associated with overall survival. Significant interactions were detected between T cell activation and ts-34 and ts-49. In the T cell exhaustion group, patients with a low level of ts-34 or a high level of ts-49 showed improved survival. In contrast, there was no significant difference in the activation group. Breast cancer related pathways were identified for the five tRFs. In conclusion, the identified five tRFs associated with overall survival may serve as therapeutic targets and improve immunotherapy in breast cancer.
Clostridioides difficile infections (CDI) are the leading cause of nosocomial antibiotic-associated diarrhea. C. difficile produces dormant spores that serve as infectious agents. Bile salts in the gastrointestinal tract signal spores to germinate into toxin-producing cells. As spore germination is required for CDI onset, anti-germination compounds may serve as prophylactics. CamSA, a synthetic bile salt, was previously shown to inhibit C. difficile spore germination in vitro and in vivo. Unexpectedly, a single dose of CamSA was sufficient to offer multi-day protection from CDI in mice without any observable toxicity. To study this intriguing protection pattern, we examined the pharmacokinetic parameters of CamSA. CamSA was stable to the gut of antibiotic-treated mice but was extensively degraded by the microbiota of non-antibiotic-treated animals. Our data also suggest that CamSA's systemic absorption is minimal since it is retained primarily in the intestinal lumen and liver. CamSA shows weak interactions with CYP3A4, a P450 hepatic isozyme involved in drug metabolism and bile salt modification. Like other bile salts, CamSA seems to undergo enterohepatic circulation. We hypothesize that the cycling of CamSA between the liver and intestines serves as a slow-release mechanism that allows CamSA to be retained in the gastrointestinal tract for days. This model explains how a single CamSA dose can prevent murine CDI even though spores are present in the animal's intestine for up to four days post-challenge.
Longitudinal phenotypes have been increasingly available in genome-wide association studies (GWAS) and electronic health record-based studies for identification of genetic variants that influence complex traits over time. For longitudinal binary data, there remain significant challenges in gene mapping, including misspecification of the model for phenotype distribution due to ascertainment. Here, we propose L-BRAT (Longitudinal Binary-trait Retrospective Association Test), a retrospective, generalized estimating equation-based method for genetic association analysis of longitudinal binary outcomes. We also develop RGMMAT, a retrospective, generalized linear mixed model-based association test. Both tests are retrospective score approaches in which genotypes are treated as random conditional on phenotype and covariates. They allow both static and time-varying covariates to be included in the analysis. Through simulations, we illustrated that retrospective association tests are robust to ascertainment and other types of phenotype model misspecification, and gain power over previous association methods. We applied L-BRAT and RGMMAT to a genome-wide association analysis of repeated measures of cocaine use in a longitudinal cohort. Pathway analysis implicated association with opioid signaling and axonal guidance signaling pathways. Lastly, we replicated important pathways in an independent cocaine dependence case-control GWAS. Our results illustrate that L-BRAT is able to detect important loci and pathways in a genome scan and to provide insights into genetic architecture of cocaine use.
Aligned DNA sequences have been widely used for quantitatively analyzing and interpreting evolutionary processes. By comparing the information between intraspecific polymorphism with interspecific divergence in two sibling species, Poisson random field (PRF) theory offers a statistical framework with which various genetic parameters such as natural selection intensity, mutation rate and speciation time can be effectively estimated. A recently developed time-inhomogeneous PRF model has reinforced the original method by removing the assumption of stationary site frequency, but it keeps the assumption that the two sibling species share same effective population size with their ancestral species. This paper explores a relaxation of this biologically unrealistic assumption by hypothesizing that each of the two descendant species experienced a sudden change in population size at the times of the divergence from their most recent common ancestor. The newly developed PRF model with non-constant population size is applied to a set of 91 genes in African population of Drosophila to make statistical inference of the various genetic parameters under a hierarchical Bayesian framework and carried out with a multi-layer Markov chain Monte Carlo sampling scheme. In order to meet the intensive computational demand, a R program is integrated with C++ code and a parallel executing technique is designed to run the program with multiple CPU cores.
Genome-wide association studies (GWAS) have identified over 100 loci associated with schizophrenia. Most of these studies test genetic variants for association one at a time. In this study, we performed GWAS of the molecular genetics of schizophrenia (MGS) dataset with 5334 subjects using multivariate Bayesian variable selection (BVS) method Posterior Inference via Model Averaging and Subset Selection (piMASS) and compared our results with the previous univariate analysis of the MGS dataset. We showed that piMASS can improve the power of detecting schizophrenia-associated SNPs, potentially leading to new discoveries from existing data without increasing the sample size. We tested SNPs in groups to allow for local additive effects and used permutation test to determine statistical significance in order to compare our results with univariate method. The previous univariate analysis of the MGS dataset revealed no genome-wide significant loci. Using the same dataset, we identified a single region that exceeded the genome-wide significance. The result was replicated using an independent Swedish Schizophrenia Case–Control Study (SSCCS) dataset. Based on the SZGR 2.0 database we found 63 SNPs from the best performing regions that are mapped to 27 genes known to be associated with schizophrenia. Overall, we demonstrated that piMASS could discover association signals that otherwise would need a much larger sample size. Our study has important implication that reanalyzing published datasets with BVS methods like piMASS might have more power to discover new risk variants for many diseases without new sample collection, ascertainment, and genotyping.
We have developed a Poisson random field model for estimating the distribution of selective effects of newly arisen nonsynonymous mutations that could be observed as polymorphism or divergence in samples of two related species under the assumption that the two species populations are not at mutation-selection-drift equilibrium. The model is applied to 91Drosophila genes by comparing levels of polymorphism in an African population of D. melanogaster with divergence to a reference strain of D. simulans. Based on the difference of gene expression level between testes and ovaries, the 91 genes were classified as 33 male-biased, 28 female-biased, and 30 sex-unbiased genes. Under a Bayesian framework, Markov chain Monte Carlo simulations are implemented to the model in which the distribution of selective effects is assumed to be Gaussian with a mean that may differ from one gene to the other to sample key parameters. Based on our estimates, the majority of newly-arisen nonsynonymous mutations that could contribute to polymorphism or divergence in Drosophila species are mildly deleterious with a mean scaled selection coefficient of -2.81, while almost 86% of the fixed differences between species are driven by positive selection. There are only 16.6% of the nonsynonymous mutations observed in sex-unbiased genes that are under positive selection in comparison to 30% of male-biased and 46% of female-biased genes that are beneficial. We also estimated that D. melanogaster and D. simulans may have diverged 1.72 million years ago.
The Friedman test is often used for a randomized complete block design when the normality assumption is not satisfied or the data are ordinal. The Friedman test can be viewed as an extension of the sign test for multiple measurements within each subject or block. We propose a modified Friedman test based on the Wilcoxon sign rank approach. Coincidentally, Tukey proposed a test statistic similar to our proposed test, but blocks are ranked by the minimum difference within each block. In the proposed test, we use the variance of block to rank the blocks, with the least variance being ranked the smallest. In both Tukey test and the modified Friedman test, linear ranks are used for blocks and treatments. The Tukey test belongs to the family of weighted-ranking test from Quade (1979).The modified Friedman test, the Friedman test and the Tukey test are compared under various conditions and the results indicate that the proposed test is generally more powerful than the Friedman test and the Tukey test when the number of groups is small.
Cancer incidence disparities exist among specific Asian American populations. However, the existing reports exclude data from large metropoles like Chicago, Houston and New York. Moreover, incidence rates by subgroup have been underestimated due to the exclusion of Asians with unknown subgroup. Cancer incidence data for 2009 to 2011 for eight states accounting for 68% of the Asian American population were analyzed. Race for cases with unknown subgroup was imputed using stratified proportion models by sex, age, cancer site and geographic regions. Age-standardized incidence rates were calculated for 17 cancer sites for the six largest Asian subgroups. Our analysis comprised 90,709 Asian and 1,327,727 non-Hispanic white cancer cases. Asian Americans had significantly lower overall cancer incidence rates than non-Hispanic whites (336.5 per 100,000 and 541.9 for men, 299.6 and 449.3 for women, respectively). Among specific Asian subgroups, Filipino men (377.4) and Japanese women (342.7) had the highest overall incidence rates while South Asian men (297.7) and Korean women (275.9) had the lowest. In comparison to non-Hispanic whites and other Asian subgroups, significantly higher risks were observed for colorectal cancer among Japanese, stomach cancer among Koreans, nasopharyngeal cancer among Chinese, thyroid cancer among Filipinos, and liver cancer among Vietnamese. South Asians had remarkably low lung cancer risk. Overall, Asian Americans have a lower cancer risk than non-Hispanic whites, except for nasopharyngeal, liver and stomach cancers. The unique portrayal of cancer incidence patterns among specific Asian subgroups in this study provides a new baseline for future cancer surveillance research and health policy.
Sensitivity and specificity are often used to assess the performance of a diagnostic test with binary outcomes. Wald-type test statistics have been proposed for testing sensitivity and specificity individually. In the presence of a gold standard, simultaneous comparison between two diagnostic tests for noninferiority of sensitivity and specificity based on an asymptotic approach has been studied by Chen et al. (2003). However, the asymptotic approach may suffer from unsatisfactory type I error control as observed from many studies, especially in small to medium sample settings. In this paper, we compare three unconditional approaches for simultaneously testing sensitivity and specificity. They are approaches based on estimation, maximization, and a combination of estimation and maximization. Although the estimation approach does not guarantee type I error, it has satisfactory performance with regard to type I error control. The other two unconditional approaches are exact. The approach based on estimation and maximization is generally more powerful than the approach based on maximization.
The resistance gene Pi-ta has been effectively used to control rice blast disease, but some populations of cultivated and wild rice have evolved resistance. Insights into the evolutionary processes that led to this resistance during crop domestication may be inferred from the population history of domesticated and wild rice strains. In this study, we applied a recently developed statistical method, time-dependent Poisson random field model, to examine the evolution of the Pi-ta gene in cultivated and weedy rice. Our study suggests that the Pi-ta gene may have more recently introgressed into cultivated rice, indica and japonica, and U.S. weedy rice from the wild species, O. rufipogon. In addition, the Pi-ta gene is under positive selection in japonica, tropical japonica, U.S. cultivars and U.S. weedy rice. We also found that sequences of two domains of the Pi-ta gene, the nucleotide binding site and leucine-rich repeat domain, are highly conserved among all rice accessions examined. Our results provide a valuable analytical tool for understanding the evolution of disease resistance genes in crop plants.