Acute kidney injury (AKI) and chronic kidney disease (CKD) are two interconnected clinical conditions, both defined by degree of functional impairment, but with heterogeneous clinical trajectories. Using new transcriptomic technologies, recent studies have described the cellular diversity in the healthy and injured kidney at the single cell level. Here, we used single nucleus transcriptomics to investigate the molecular diversity and commonalities in kidney biopsies from over 150 participants with AKI and CKD enrolled within the Kidney Precision Medicine Project (KPMP) and did so at the patient participant level. Using an unsupervised approach, we identified two multi-cellular programs associated with clinical and histopathological features of acute injury and chronic damage, respectively. We found that these programs are expressed across patients with AKI and CKD, supporting shared, rather than distinct, underlying molecular mechanisms. These programs capture tissue-level compositional changes towards adaptive and failed-repair states in tubular epithelial cells, as well as intra-cellular molecular changes characteristic of stress in all cell types. We identified subunits of the NFkB and AP-1 complexes, as well as members of the STAT family, as putative upstream regulators of the acute and chronic programs. We were able to map these continuous molecular measures of acute injury and chronic damage to urine and plasma protein profiles obtained at time of biopsy. These non-invasive protein signatures were predictive of renal outcomes in an independent cohort of 44 thousand participants from the UK biobank. In summary, unbiased identification of cellular programs in kidney disease biopsies defined molecular programs of injury cutting across conventional disease categorization and established a non-invasive molecular link to long term patient outcomes.
MOTIVATION:Convergence analysis can characterize genetic elements underlying morphological adaptations. However, its performance on regulatory elements is limited due to their modular composition of transcription factor motifs, which have rapid turnover and experience different evolutionary pressures. RESULTS:We introduce phyloConverge, a phylogenetic method that performs scalable, fine-grained local convergence analysis of genomic elements at flexible length scales. Using a benchmarking case of convergent subterranean mammal adaptation, phyloConverge identifies rate-accelerated conserved noncoding elements (CNEs) with high specificity and statistical robustness relative to competing methods. From CNE-level scoring, we detect the convergent regression of entire CNE units and highlight the contrast that subterranean-associated coding region regression is highly specific to ocular functions, whereas regulatory element regression is enriched for accompanying neuronal phenotypes and other developmental processes. From transcription factor motif-level scoring, we dissect elements into subregions with uneven convergence signals and demonstrate the modular adaptation of CNEs with high functional specificity. Finally, we demonstrate phyloConverge's scalability to perform high-resolution convergence analysis genome-wide. AVAILABILITY AND IMPLEMENTATION:phyloConverge is available at https://github.com/ECSaputra/phyloConverge.
High-throughput gene expression profiling measures individual gene expression across conditions. However, genes are regulated in complex networks, not as individual entities, limiting the interpretability of gene expression data. Machine learning models that incorporate prior biological knowledge are a powerful tool to extract meaningful biology from gene expression data. Pathway-level information extractor (PLIER) is an unsupervised machine learning method that defines biological pathways by leveraging the vast amount of published transcriptomic data. PLIER converts gene expression data into known pathway gene sets, termed latent variables (LVs), to substantially reduce data dimensionality and improve interpretability. In the current study, we trained the first mouse PLIER model on 190,111 mouse brain RNA-sequencing samples, the greatest amount of training data ever used by PLIER. We then validated the mousiPLIER approach in a study of microglia and astrocyte gene expression across mouse brain aging. mousiPLIER identified biological pathways that are significantly associated with aging, including one latent variable (LV41) corresponding to striatal signal. To gain further insight into the genes contained in LV41, we performed k-means clustering on the training data to identify studies that respond strongly to LV41. We found that the variable was relevant to striatum and aging across the scientific literature. Finally, we built a Web server (http:/mousiplier.greenelab.com/) for users to easily explore the learned latent variables. Taken together, this study defines mousiPLIER as a method to uncover meaningful biological processes in mouse brain transcriptomic studies.
Regular exercise has many physical and brain health benefits, yet the molecular mechanisms mediating exercise effects across tissues remain poorly understood. Here we analyzed 400 high-quality DNA methylation, ATAC-seq, and RNA-seq datasets from eight tissues from control and endurance exercise-trained (EET) rats. Integration of baseline datasets mapped the gene location dependence of epigenetic control features and identified differing regulatory landscapes in each tissue. The transcriptional responses to 8 weeks of EET showed little overlap across tissues and predominantly comprised tissue-type enriched genes. We identified sex differences in the transcriptomic and epigenomic changes induced by EET. However, the sex-biased gene responses were linked to shared signaling pathways. We found that many G protein-coupled receptor-encoding genes are regulated by EET, suggesting a role for these receptors in mediating the molecular adaptations to training across tissues. Our findings provide new insights into the mechanisms underlying EET-induced health benefits across organs.
DNA methylation comprises a cumulative record of lifetime exposures superimposed on genetically determined markers. Little is known about methylation dynamics in humans following an acute perturbation, such as infection. We characterized the temporal trajectory of blood epigenetic remodeling in 133 participants in a prospective study of young adults before, during, and after asymptomatic and mildly symptomatic SARS‐CoV‐2 infection. The differential methylation caused by asymptomatic or mildly symptomatic infections was indistinguishable. While differential gene expression largely returned to baseline levels after the virus became undetectable, some differentially methylated sites persisted for months of follow‐up, with a pattern resembling autoimmune or inflammatory disease. We leveraged these responses to construct methylation‐based machine learning models that distinguished samples from pre‐, during‐, and postinfection time periods, and quantitatively predicted the time since infection. The clinical trajectory in the young adults and in a diverse cohort with more severe outcomes was predicted by the similarity of methylation before or early after SARS‐CoV‐2 infection to the model‐defined postinfection state. Unlike the phenomenon of trained immunity, the postacute SARS‐CoV‐2 epigenetic landscape we identify is antiprotective.
High throughput gene expression profiling is a powerful approach to generate hypotheses on the underlying causes of biological function and disease. Yet this approach is limited by its ability to infer underlying biological pathways and burden of testing tens of thousands of individual genes. Machine learning models that incorporate prior biological knowledge are necessary to extract meaningful pathways and generate rational hypothesis from the vast amount of gene expression data generated to date. We adopted an unsupervised machine learning method, Pathway-level information extractor (PLIER), to train the first mouse PLIER model on 190,111 mouse brain RNA-sequencing samples, the greatest amount of training data ever used by PLIER. mousiPLER converted gene expression data into a latent variables that align to known pathway or cell maker gene sets, substantially reducing data dimensionality and improving interpretability. To determine the utility of mousiPLIER, we applied it to a mouse brain aging study of microglia and astrocyte transcriptomic profiling. We found a specific set of latent variables that are significantly associated with aging, including one latent variable (LV41) corresponding to striatal signal. We next performed k-means clustering on the training data to identify studies that respond strongly to LV41, finding that the variable is relevant to striatum and aging across the scientific literature. Finally, we built a web server (http://mousiplier.greenelab.com/) for users to easily explore the learned latent variables. Taken together this study provides proof of concept that mousiPLIER can uncover meaningful biological processes in mouse transcriptomic studies. Significance statement Analysis of RNA-sequencing data commonly generates differential expression of individual genes across conditions. However, genes are regulated in complex networks, not as individual entities. Machine learning models that incorporate prior biological information are a powerful tool to analyze human gene expression. However, such models are lacking for mouse despite the vast number of mouse RNA-seq datasets. We trained a mouse pathway-level information extractor model (mousiPLIER). The model reduced the data dimensionality from over 10,000 genes to 196 latent variables that map to prior pathway and cell marker gene sets. We demonstrated the utility of mousiPLIER by applying it to mouse brain aging data and developed a web server to facilitate the use of the model by the scientific community.
The identification of a COVID-19 host response signature in blood can increase the understanding of SARS-CoV-2 pathogenesis and improve diagnostic tools. Applying a multi-objective optimization framework to both massive public and new multi-omics data, we identified a COVID-19 signature regulated at both transcriptional and epigenetic levels. We validated the signature's robustness in multiple independent COVID-19 cohorts. Using public data from 8,630 subjects and 53 conditions, we demonstrated no cross-reactivity with other viral and bacterial infections, COVID-19 comorbidities, or confounders. In contrast, previously reported COVID-19 signatures were associated with significant cross-reactivity. The signature's interpretation, based on cell-type deconvolution and single-cell data analysis, revealed prominent yet complementary roles for plasmablasts and memory T cells. Although the signal from plasmablasts mediated COVID-19 detection, the signal from memory T cells controlled against cross-reactivity with other viral infections. This framework identified a robust, interpretable COVID-19 signature and is broadly applicable in other disease contexts. A record of this paper's transparent peer review process is included in the supplemental information.
Early SARS-CoV-2 diagnosis is a key nonpharmacological strategy to contain the current pandemic. Nucleic acid amplification tests (NAATs), the reference standard for SARS-CoV-2 diagnosis, are poorly sensitive during the first 4 days after infection, with false negative rates estimated in the range 67–100%.1 Here, we implemented a new assay that shows increased sensitivity to SARS-CoV-2 infection during the early window of NAAT false negativity. Host response assays (HRAs) are emerging as a new paradigm for infection diagnosis,2 recently implemented to discriminate viral from bacterial infections,3, 4 and to detect early respiratory viral illnesses.5 Unlike NAATs that target viral genetic material, HRAs target transcriptional alterations in the host blood. These alterations may become detectable by RT-PCR as early as 12 hours after viral challenge.6 Given the potentially higher sensitivity early in infection, we set out to implement the first HRA for SARS-CoV-2 diagnosis. Our work leveraged the COVID-19 Health Action Response for Marines (CHARM), a prospective study that identified incident SARS-CoV-2 infection among US Marine recruits from May 12 to November 5 2020.7, 8 The cohort included 3249 predominantly young, male participants. Participants were typically tested by an FDA-approved NAAT for SARS-CoV-2 three times during an initial 2-week quarantine, and then biweekly for 6 weeks during basic training (Figure 1A). Most infected participants were asymptomatic at the first positive NAAT and none required hospitalization. During basic training, 45.1% of participants showed a SARS-CoV-2 NAAT positive result at one or more time points. The high infection rate, along with the longitudinal design, made the CHARM study highly instrumental for benchmarking a new SARS-CoV-2 diagnostic assay. The strategy to develop a SARS-CoV-2 HRA followed four main steps (Figure 1B-E): (1) bioinformatics-driven identification of a SARS-CoV-2 host response signature; (2) technical implementation; (3) cross-sectional benchmark, by comparing HRA and NAAT results from different participants at randomly selected time points; (4) longitudinal benchmark, by comparing HRA and NAAT repeated measures over time for the same participants. The first challenge we faced was to identify a host transcriptional response specific for SARS-CoV-2 infection. We aimed to find a compact set of 40–50 genes whose expression in the blood would indicate SARS-CoV-2-infection, but not related infections such as influenza. To address this problem, we curated a compendium of public blood transcriptomes from 15 COVID-19 studies and from 112 studies on a wide variety of viral and bacterial infections (Figure 1B, Supporting information). Furthermore, the compendium included transcriptomes on COVID-19 comorbidities (e.g., obesity, hypertension) and risk factors (e.g., age, sex) that might act as potential confounders. Applying a combination of meta-analysis and optimization techniques to the data compendium, we identified 41 genes that together provided robust SARS-CoV-2 detection (receiver operator curve (ROC) area under the curve (AUC) 0.7–0.9), and low cross-reactivity with other infections and confounding factors (ROC AUC ≤0.5). Next, we implemented a HRA with three main components: whole blood collection through a PAXgene Blood RNA Tube (BD Biosciences, San Jose, CA, USA); measurement of the expression levels of the 41 transcripts on an integrated fluidic circuit; sample interpretation through a machine learning algorithm (Figure 1C, Supporting information). The algorithm was based on a regularized logistic regression classifier, taking as input the combined expression levels of the 41 transcripts measured in a blood sample, and returning as output the sample interpretation in one of the following classes: SARS-CoV-2 positive; SARS-CoV-2 negative; inclusive, in case of highly uncertain interpretation. The algorithm was developed using a training set of 245 SARS-CoV-2 positive and 296 SARS-CoV-2 negative samples from the CHARM study. To control for viral cross-reactivity, the training set included 63 blood samples from subjects in a vaccine trial after H3N2 influenza virus challenge.9 During algorithm training, the influenza samples were treated as SARS-CoV-2 negative. We performed extensive tests to ensure that the machine learning-generated interpretation calls were highly reproducible across sample technical replicates. We first assessed the HRA performance in a cross-sectional way (Figure 1D, Supporting information). We extracted samples from the SARS-CoV-2 positive (n = 93) and negative (n = 93) groups at random time points, disregarding the participants’ testing history. All of these samples were from participants not contributing to the training data, to avoid leakage from the training to the benchmark data. Using a NAAT-based comparator as the reference standard, HRA had a positive percent agreement (PPA) of 96.6% (95% confidence interval (CI), 90.7–98.9%), an negative percent agreement (NPA) of 97.7% (95% CI, 92.2–99.4%) (Table 1). To assess cross-reactivity, we used 33 additional influenza samples from subjects in the influenza vaccine trial cohort used for training. Two samples produced inconclusive HRA results, and the cross-reactivity rate was 4/31 = 12.9% (95% CI, 4.2–30.7%) (Table 1). Overall, the cross-sectional benchmark demonstrated a high concordance between HRA and NAAT results. We then performed a longitudinal benchmark by comparing HRA and NAAT repeated measures for the same participants over time (Figure 1E). The goal of this assessment was to explore whether HRA could anticipate SARS-CoV-2 diagnosis compared to NAAT. Due to the absence of a reference standard for SARS-CoV-2 diagnosis prior to NAAT positivity, we performed a validation study.10 We reasoned that some study participants were infected before their first positive NAAT result, but undetected due to low NAAT sensitivity early in infection. First, we defined groups of samples with higher and lower risk for NAAT early false negativity, based on phylogenetic and epidemiological evidence (Supporting information). Second, we compared HRA results in the two groups (Table 1; Figure 2). In the higher-risk group, HRA was positive before NAAT in 10 of 15 participants (66.6%). In the lower-risk group, HRA was positive in 0 of 8 participants (0%). The results support an earlier SARS-CoV-2 diagnosis using HRA as compared to NAAT (Fisher exact test, p = 0.0027). Limitations of our study include an unknown generalizability beyond young, healthy, male participants; some cross-reactivity with influenza and possibly with other infections such as other coronaviruses; lack of knowledge of when SARS-CoV-2 exposure occurred, or of when NAAT would first turn positive with more frequent testing. Since the beginning of the COVID-19 pandemic, several diagnostic technologies have been proposed including surface-enhanced Raman spectroscopy and field-effect transistor-based biosensors. Compared to these and other technologies, the main advantage of HRAs is the potentially higher sensitivity early in infection. This benefit should be assessed relative to the additional cost associated with blood draws. Although a cost-benefit analysis was beyond the scope of our work, we envisage scenarios where using an HRA may be cost-effective. These scenarios include, for example, hospitals and nursing homes where the need to ensure virus-free environments is of critical importance. In conclusion, our work provides the first implementation of a SARS-CoV-2 HRA and initial evidence that monitoring the host response can anticipate NAAT infection diagnosis. We thank Daniel Chawla, BS, and Steven Kleinstein, PhD, for the signature meta-analysis, Aliza Rubenstein, PhD; and Elena Zaslavsky, PhD, for signature identification, and David King PhD; Charles Park PhD; Mandi Wong; Harry Lin; Vishal Thakore; and Apoorva Sooranahalli Meng for implementation of host response assay in a microfluidic system. AC, JG, YG, VDN, NR, and SCS are coinventors of the Host Response Assay. JG and NR are employees of Fluidigm. AGL has nothing to declare. Defense Advanced Research Projects Agency (contract number N6600119C4022) has provided the funding for this work. Role of the Funder/Sponsor: The funder had no role in any aspect of the study. AGL is a military service member. This work was prepared as part of his official duties. Title 17, US Code §105 provides that copyright protection under this title is not available for any work of the US Government. Title 17, US code §101 defines a US Government work as a work prepared by a military service member or employee of the US Government as part of that person's official duties. The views expressed in the article are those of the authors and do not necessarily express the official policy and position of the US Navy, the Department of Defense, the US government or the institutions affiliated with the authors. Please note: The publisher is not responsible for the content or functionality of any supporting information supplied by the authors. Any queries (other than missing content) should be directed to the corresponding author for the article.
Male sex is a major risk factor for SARS-CoV-2 infection severity. To understand the basis for this sex differ-ence, we studied SARS-CoV-2 infection in a young adult cohort of United States Marine recruits. Among 2,641 male and 244 female unvaccinated and seronegative recruits studied longitudinally, SARS-CoV-2 infec-tions occurred in 1,033 males and 137 females. We identified sex differences in symptoms, viral load, blood transcriptome, RNA splicing, and proteomic signatures. Females had higher pre-infection expression of anti-viral interferon-stimulated gene (ISG) programs. Causal mediation analysis implicated ISG differences in number of symptoms, levels of ISGs, and differential splicing of CD45 lymphocyte phosphatase during infec-tion. Our results indicate that the antiviral innate immunity set point causally contributes to sex differences in response to SARS-CoV-2 infection. A record of this paper's transparent peer review process is included in the supplemental information.
Motivation Single-cell RNA-seq analysis has emerged as a powerful tool for understanding inter-cellular heterogeneity. Due to the inherent noise of the data, computational techniques often rely on dimensionality reduction (DR) as both a pre-processing step and an analysis tool. Ideally, dimensionality reduction should preserve the biological information while discarding the noise. However if the dimensionality reduction is to be used directly to gain biological insight it must also be interpretable – that is the individual dimensions of the reduction should correspond to specific biological variables such as cell-type identity or pathway activity. Maximizing biological interpretability necessitates making assumption about the data structures and the choice of the model is critical. Results We present a new probabilistic single-cell factor analysis model, N on-negative I ndependent F actor A nalysis (NIFA), that incorporates different interpretability inducing assumptions into a single modeling framework. The key advantage of our NIFA model is that it simultaneously models uni- and multi-modal latent factors, and thus isolates discrete cell-type identity and continuous pathway activity into separate components. We apply our approach to a range of datasets where cell-type identity is known, and we show that NIFA-derived factors outperform results from ICA, PCA, NMF and scCoGAPS (an NMF method designed for single-cell data) in terms of disentangling biological sources of variation. Studying an immunotherapy dataset in detail, we show that NIFA is able to reproduce and refine previous findings in a single analysis framework and enables the discovery of new clinically relevant cell states. Availability NFIA is a R package which is freely available at GitHub ( https://github.com/wgmao/NIFA ). Contact mchikina@pitt.edu Supplementary information Supplementary data are available at Bioinformatics online.
Physiological and morphological adaptations to extreme environments arise from the molecular evolution of protein-coding regions and regulatory elements (REs) that regulate gene expression. Comparative genomics methods can characterize genetic elements that underlie the organism-level adaptations, but convergence analyses of REs are often limited by their evolutionary properties. A RE can be modularly composed of multiple transcription factor binding sites (TFBS) that may each experience different evolutionary pressures. The modular composition and rapid turnover of TFBS also enables a compensatory mechanism among nearby TFBS that allows for weaker sequence conservation/divergence than intuitively expected. Here, we introduce phyloConverge , a comparative genomics method that can perform fast, fine-grained local convergence analysis of genetic elements. phyloConverge calibrates for local shifts in evolutionary rates using a combination of maximum likelihood-based estimation of nucleotide substitution rates and phylogenetic permutation tests. Using the classical convergence case of mammalian adaptation to subterranean environments, we validate that phyloConverge identifies rate-accelerated conserved non-coding elements (CNEs) that are strongly correlated with ocular tissues, with improved specificity compared to competing methods. We use phyloConverge to perform TFBS-scale and nucleotide-scale scoring to dissect each CNE into subregions with uneven convergence signals and demonstrate its utility for understanding the modularity and pleiotropy of REs. Subterranean-accelerated regions are also enriched for molecular pathways and TFBS motifs associated with neuronal phenotypes, suggesting that subterranean eye degeneration may coincide with a remodeling of the nervous system. phyloConverge offers a rapid and accurate approach for understanding the evolution and modularity of regulatory elements underlying phenotypic adaptation.
Recent single-cell studies of cancer in both mice and humans have identified the emergence of a myofibroblast population specifically marked by the highly restricted leucine-rich-repeat-containing protein 15 (LRRC15) 1 – 3 . However, the molecular signals that underlie the development of LRRC15 + cancer-associated fibroblasts (CAFs) and their direct impact on anti-tumour immunity are uncharacterized. Here in mouse models of pancreatic cancer, we provide in vivo genetic evidence that TGFβ receptor type 2 signalling in healthy dermatopontin + universal fibroblasts is essential for the development of cancer-associated LRRC15 + myofibroblasts. This axis also predominantly drives fibroblast lineage diversity in human cancers. Using newly developed Lrrc15– diphtheria toxin receptor knock-in mice to selectively deplete LRRC15 + CAFs, we show that depletion of this population markedly reduces the total tumour fibroblast content. Moreover, the CAF composition is recalibrated towards universal fibroblasts. This relieves direct suppression of tumour-infiltrating CD8 + T cells to enhance their effector function and augments tumour regression in response to anti-PDL1 immune checkpoint blockade. Collectively, these findings demonstrate that TGFβ-dependent LRRC15 + CAFs dictate the tumour-fibroblast setpoint to promote tumour growth. These cells also directly suppress CD8 + T cell function and limit responsiveness to checkpoint blockade. Development of treatments that restore the homeostatic fibroblast setpoint by reducing the population of pro-disease LRRC15 + myofibroblasts may improve patient survival and response to immunotherapy.
We investigated serological responses following a SARS-CoV-2 outbreak in spring 2020 on a US Marine recruit training base. 147 participants that were isolated during an outbreak of respiratory illness were enrolled in this study, with visits approximately 6 and 10 weeks post-outbreak (PO). This cohort is comprised of young healthy adults, ages 18-26, with a high rate of asymptomatic infection or mild symptoms, and therefore differs from previously reported longitudinal studies on humoral responses to SARS-CoV-2, which often focus on more diverse age populations and worse clinical presentation. 80.9% (119/147) of the participants presented with circulating IgG antibodies against SARS-CoV-2 spike (S) receptor-binding domain (RBD) at 6 weeks PO, of whom 97.3% (111/114) remained positive, with significantly decreased levels, at 10 weeks PO. Neutralizing activity was detected in all sera from SARS-CoV-2 IgG positive participants tested (n=38) at 6 and 10 weeks PO, without significant loss between time points. IgG and IgA antibodies against SARS-CoV-2 RBD, S1, S2, and the nucleocapsid (N) protein, as well neutralization activity, were generally comparable between those participants that had asymptomatic infection or mild disease. A multiplex assay including S proteins from SARS-CoV-2 and related zoonotic and human endemic betacoronaviruses revealed a positive correlation for polyclonal cross-reactivity to S after SARS-CoV-2 infection. Overall, young adults that experienced asymptomatic or mild SARS-CoV-2 infection developed comparable humoral responses, with no decrease in neutralizing activity at least up to 10 weeks after infection.
MOTIVATION:RNA-seq technology provides unprecedented power in the assessment of the transcription abundance and can be used to perform a variety of downstream tasks such as inference of gene-correlation network and eQTL discovery. However, raw gene expression values have to be normalized for nuisance biological variation and technical covariates, and different normalization strategies can lead to dramatically different results in the downstream study.RESULTS:We describe a generalization of singular value decomposition-based reconstruction for which the common techniques of whitening, rank-k approximation and removing the top k principal components are special cases. Our simple three-parameter transformation, DataRemix, can be tuned to reweigh the contribution of hidden factors and reveal otherwise hidden biological signals. In particular, we demonstrate that the method can effectively prioritize biological signals over noise without leveraging external dataset-specific knowledge, and can outperform normalization methods that make explicit use of known technical factors. We also show that DataRemix can be efficiently optimized via Thompson sampling approach, which makes it feasible for computationally expensive objectives such as eQTL analysis. Finally, we apply our method to the Religious Orders Study and Memory and Aging Project dataset, and we report what to our knowledge is the first replicable trans-eQTL effect in human brain.AVAILABILITYAND IMPLEMENTATION:DataRemix is an R package which is freely available at GitHub (https://github.com/wgmao/DataRemix).SUPPLEMENTARY INFORMATION:Supplementary data are available at Bioinformatics online.
The explosive growth of next-generation sequencing data enhances our ability to understand biological process at an unprecedented resolution. Meanwhile organizing and utilizing this tremendous amount of data becomes a big challenge. High-throughput technology provides us a snapshot of all underlying biological activities, but this kind of extremely high-dimensional data is hard to interpret. Due to the curse of dimensionality, the measurement is sparse and far from enough to shape the actual manifold in the high-dimensional space. On the other hand, the measurements may contain structured noise such as technical or nuisance biological variation which can interfere downstream interpretation. Generative modeling is a powerful tool to make sense of the data and generate compact representations summarizing the embedded biological information. This thesis introduces three generative models that help amplifying biological signals buried in the noisy bulk and single-cell RNA-seq data. In Chapter 2, we propose a semi-supervised deconvolution framework called PLIER which can identify regulations in cell-type proportions and specific pathways that control gene expression. PLIER has inspired the development of MultiPLIER and has been used to infer context-specific genotype effects in the brain. In Chapter 3, we construct a supervised transformation named DataRemix to normalize bulk gene expression profiles in order to maximize the biological findings with respect to a variety of downstream tasks. By reweighing the contribution of hidden factors, we are able to reveal the hidden biological signals without any external dataset-specific knowledge. We apply DataRemix to the ROSMAP dataset and report the first replicable trans-eQTL effect in human brain. In Chapter 4, we focus on scRNA-seq and introduce NIFA which is an unsupervised decomposition framework that combines the desired properties of PCA, ICA and NMF. It simultaneously models uni- and multi-modal factors isolating discrete cell-type identity and continuous pathway-level variations into separate components. The work presented in Chapter 2 has been published as a journal article. The work in Chapter 3 and Chapter 4 are under submission and they are available as preprints on bioRxiv.
MOTIVATION:When different lineages of organisms independently adapt to similar environments, selection often acts repeatedly upon the same genes, leading to signatures of convergent evolutionary rate shifts at these genes. With the increasing availability of genome sequences for organisms displaying a variety of convergent traits, the ability to identify genes with such convergent rate signatures would enable new insights into the molecular basis of these traits.RESULTS:Here we present the R package RERconverge, which tests for association between relative evolutionary rates of genes and the evolution of traits across a phylogeny. RERconverge can perform associations with binary and continuous traits, and it contains tools for visualization and enrichment analyses of association results.AVAILABILITY AND IMPLEMENTATION:RERconverge source code, documentation and a detailed usage walk-through are freely available at https://github.com/nclark-lab/RERconverge. Datasets for mammals, Drosophila and yeast are available at https://bit.ly/2J2QBnj.SUPPLEMENTARY INFORMATION:Supplementary data are available at Bioinformatics online.
A major challenge in gene expression analysis is to accurately infer relevant biological insights, such as variation in cell-type proportion or pathway activity, from global gene expression studies. We present pathway-level information extractor (PLIER) ( https://github.com/wgmao/PLIER and http://gobie.csb.pitt.edu/PLIER ), a broadly applicable solution for this problem that outperforms available cell proportion inference algorithms and can automatically identify specific pathways that regulate gene expression. Our method improves interstudy replicability and reveals biological insights when applied to trans-eQTL (expression quantitative trait loci) identification.
Vehicular mobile crowd sensing is a fast-emerging paradigm to collect data about the environment by mounting sensors on vehicles such as taxis. An important problem in vehicular crowd sensing is to design payment mechanisms to incentivize drivers (agents) to collect data, with the overall goal of obtaining the maximum amount of data (across multiple vehicles) for a given budget. Past works on this problem consider a setting where each agent operates in isolation---an assumption which is frequently violated in practice. In this paper, we design an incentive mechanism to incentivize agents who can engage in arbitrary collusions. We then show that in a "homogeneous" setting, our mechanism is optimal, and can do as well as any mechanism which knows the agents' preferences a priori. Moreover, if the agents are non-colluding, then our mechanism automatically does as well as any other non-colluding mechanism. We also show that our proposed mechanism has strong (and asymptotically optimal) guarantees for a more general "heterogeneous" setting. Experiments based on synthesized data and real-world data reveal gains of over 30\% attained by our mechanism compared to past literature.