Sample multiplexing reduces cost and batch effects in population based large-scale single-cell genomics studies but requires accurate and scalable computational demultiplexing. Existing genotype-based methods, such as demuxlet, provide high accuracy but can be computationally slow and memory intensive as the number of cells, donors, and informative variants increases. Here, we introduce fastdemux, a scalable genotype-based demultiplexing framework based on a diagonal linear discriminant analysis (DLDA) model that substantially improves computational efficiency while maintaining accurate donor assignment. Using a pooled single-cell RNA-seq dataset from unrelated donors, we benchmarked fastdemux against demuxlet, vireo, and demuxalot. fastdemux achieved comparable or improved demultiplexing accuracy while reducing runtime and peak memory usage by orders of magnitude relative to alternative methods. Performance remained robust across varying sequencing depths and genotype SNP filtering thresholds. In addition, the DLDA framework naturally extends to doublet and higher-order multiplet detection. We also show that fastdemux works well with scATAC-seq data where genetic variants are more sparsely covered. Together, these results establish fastdemux as an efficient and scalable solution for genetic demultiplexing of pooled single-cell datasets.
Background:Inflammatory Bowel Disease (IBD) is characterized by chronic intestinal inflammation and is associated with both altered gut microbiome composition and host genetic risk. Both host genetic variants and the gut microbiome can affect host gene expression in the colon; however, it remains unclear whether interactions between the two (genotype × microbiome, GxM) shape intestinal gene regulation in humans and their contribution to IBD risk. Methods:We analyzed publicly available data for 86 individuals (64 patients with IBD and 22 controls) in the Inflammatory Bowel Disease Multi'omics Database consisting of host genotype, host gene expression, and mucosal gut microbiome (16S rRNA) data from rectal and ileum biopsies. We performed expression Quantitative Trait Locus (eQTL) mapping and then used computational fine-mapping to identify likely causal variants. We tested whether microbial taxa modify genetic effects on host gene expression. We then integrated GxM eQTLs with IBD, Crohn's Disease (CD) and Ulcerative Colitis (UC) Genome-Wide Association Study results by leveraging Transcriptome-Wide Association Studies and colocalization methods. Results:We found 3,777 and 3,694 host genes with eQTLs in the rectum and in the ileum, respectively (FDR = 10%). Using the fine-mapped eQTLs, we found 36 GxM interactions for 31 host genes with 22 microbial taxa in the rectum and 30 GxM interactions in the ileum for 15 host genes and 20 taxa (FDR = 10%). Taxa with GxM interactions clustered into two distinct groups with opposing effects on host gene regulation and reflected distinct functions of microbes in the gut. i.e, butyrate producers versus sulfate reducers. Integration with IBD GWAS revealed that 23 variants with GxM regulated the expression of host genes putatively causal for IBD, CD or UC (FDR = 10%), thus identifying microbes that can either amplify or buffer genetic risk. Conclusions:Our results show evidence of genetic effects on host gene expression that are modulated by microbiome composition, and provide insight into how IBD risk could be reduced by targeting specific microbial taxa contingent on host genotype.
Loneliness has been linked to increased risk of cardiovascular disease, which disproportionately affects African American adults. Dysregulation of the hypothalamic-pituitary-adrenal (HPA) axis, as measured by long-term accumulation of cortisol in hair, may be one pathway through which loneliness increases cardiovascular disease risk. However, the relationship between loneliness and hair cortisol levels among African American adults has not yet been explored. Further, both loneliness and cortisol activity differ across age and sex. To better understand the association between loneliness and HPA axis activity among middle-aged and older African American adults, the present study examined the degree to which age and sex interacted with loneliness to predict hair cortisol concentrations. Data were obtained from 340 African American adults (Mage = 66.06, SD = 5.46, range = 55-75; 87.1 % female), who provided hair samples and reported their loneliness level as a part of The Heart of Detroit Study. Results showed that sex significantly moderated the association between loneliness and hair cortisol. Loneliness was positively associated with hair cortisol concentrations in male, but not female, participants. These findings suggest that sex-specific associations may exist between loneliness and hair cortisol.
Colorectal cancer (CRC) is associated with changes in the microbial communities in the tumor microenvironment. Although metabolic reprogramming is an important feature of host cells in CRC, little is known about metabolic changes in the tumor-associated microbiota and how these microbial metabolic alterations can contribute to disease. Here, we investigated metabolic host-microbiome interactions in CRC using complementary computational and experimental approaches. Using patient-specific in silico metabolic models across three independent datasets, we discovered that Fusobacterium, a cancer-promoting taxon, consistently grows faster in tumor-associated versus normal tissue-associated microbiomes. This finding prompted us to investigate whether host metabolic changes drive these microbial growth advantages. By integrating our metabolic predictions with host transcriptomics data, we identified correlations between tumor gene expression and the growth of CRC-associated taxa (including Porphyromonadaceae, Blautia, and Streptococcus), as well as associations between host genes and microbial metabolism of dietary components (including choline, amino acids, and starch). To test whether these correlations reflect causal relationships, we simulated spent medium experiments in silico, demonstrating that Blautia preferentially grows on metabolites produced by tumor versus normal host cells. We further validated the direct impact of microbes on host metabolism using an in vitro system, where colon cancer cells exposed to human microbiomes showed gene expression changes in response to specific taxa including Bilophila, Anaerotruncus, and Escherichia. Together, these findings reveal a metabolic dialogue between host and microbiome in CRC, where tumor metabolic reprogramming creates a favorable environment for pathogenic microbes, which in turn may reinforce tumorigenic processes through metabolic crosstalk.
BACKGROUND:Aging is associated with increasing systemic inflammation (inflammaging) alongside declining immune function (immunosenescence); stress is hypothesized to exacerbate these effects. However, whether stress and age interact to effect inflammatory markers among African Americans is unknown. We hypothesized that stress would amplify age-related changes in inflammation and immune function, such that greater stress would result in greater circulating and lower stimulated markers of inflammation at greater ages. METHODS:Data are from The Heart of Detroit Study (n = 522 African Americans aged 55-75 years). Analyses included 312 participants who provided usable blood samples and had C-reactive protein (CRP) < 10 mg/L. Markers of basal inflammation [interleukin (IL)-6, IL-8, IL-10, tumor necrosis factor (TNF)-α, macrophage migration inhibitory factor (MIF), CRP] and ex vivo lipopolysaccharide-stimulated inflammation (IL-1β, IL-6, IL-8, IL-10, TNF-α) were assessed. Main effects of perceived stress and age on inflammatory markers were tested with linear regression; an interaction term was then added to assess moderation. Significant interactions were probed with simple slopes analyses. RESULTS:Age and stress interacted to affect the basal cytokine composite score, IL-8, and MIF, but not the stimulated composite score or CRP. Among participants with average to high stress, greater age was associated with higher basal cytokine composite scores; with high stress, greater age was also associated with higher IL-8. With low stress, greater age was associated with lower MIF. CONCLUSION:Results suggest that moderate to high perceived stress exacerbates inflammaging. No such effect on stimulated inflammation was observed. Intriguingly, low stress may protect against inflammaging.
cis-regulatory elements (CREs) control gene transcription dynamics across cell types and in response to the environment. In asthma, multiple immune cell types play an important role in the inflammatory process. Genetic variants in CREs can also affect gene expression response dynamics and contribute to asthma risk. However, the regulatory mechanisms underlying control of transcriptional dynamics across different environmental contexts and cell types at single-cell resolution remain to be elucidated. To resolve this question, we performed single-cell ATAC-seq (scATAC-seq) in peripheral blood mononuclear cells (PBMCs) from 16 children with asthma. PBMCs were activated with phytohemagglutinin (PHA) or lipopolysaccharide (LPS) and treated with dexamethasone (DEX), an anti-inflammatory glucocorticoid. We analyzed changes in chromatin accessibility, measured transcription factor motif activity, and identified treatment- and cell-type-specific transcription factors that drive changes in both gene expression mean and variability. We observed a strong positive linear dependence between motif response and their target gene expression changes but a negative relationship with changes in target gene expression variability. This result suggests that an increase of transcription factor binding tightens the variability of gene expression around the mean. We then annotated genetic variants in chromatin accessibility peaks and response motifs, followed by computational fine-mapping of expression quantitative trait loci (eQTL) from a pediatric asthma cohort. We found that eQTLs were 5-fold enriched in peaks with response motifs and refined the credible set for 410 asthma risk genes, with 191 having the causal variant in response motifs. In conclusion, scATAC-seq enhances the understanding of molecular mechanisms for asthma risk variants mediated by gene expression.
Social factors influence health outcomes and life expectancy. Individuals living in poverty often have adverse health outcomes related to chronic inflammation that affect the cardiovascular, renal, and pulmonary systems. Negative psychosocial experiences are associated with transcriptional changes in genes associated with complex traits. However, the underlying molecular mechanisms by which poverty increases the risk of disease and health disparities are still not fully understood. To bridge the gap in our understanding of the link between living in poverty and adverse health outcomes, we performed RNA-sequencing of blood immune cells from 204 participants of the Healthy Aging in Neighborhoods of Diversity across the Life Span (HANDLS) study in Baltimore, Maryland. We identified 138 genes differentially expressed in association with poverty. Genes differentially expressed were enriched in wound healing and coagulation processes. Of the genes differentially expressed in individuals living in poverty, EEF1DP7 and VIL1 are also associated with hypertension in transcriptome-wide association studies. Our results suggest that living in poverty influences inflammation and the risk for cardiovascular disease through gene expression changes in immune cells.
We present multi-integration of transcriptome-wide association studies and colocalization (Multi-INTACT), an algorithm that models multiple gene products (e.g. encoded RNA transcript and protein levels) to implicate causal genes and relevant gene products. In simulations, Multi-INTACT achieves higher power than existing methods, maintains calibrated false discovery rates, and detects the true causal gene product(s). We apply Multi-INTACT to GWAS on 1,408 metabolites, integrating the GTEx expression and UK Biobank protein QTL datasets. Multi-INTACT infers 52% to 109% more metabolite causal genes than protein-alone or expression-alone analyses and indicates both gene products are relevant for most gene nominations.
Psychological stress is linked to elevated markers of chronic inflammation, whereas social support is associated with lower levels; yet, the molecular mechanisms mediating these effects are poorly understood. We investigated gene regulatory variation in peripheral blood mononuclear cells (PBMCs) from 165 self-reported African American adults (aged 50-89 years) using single-cell RNA sequencing (scRNA-seq) and single-cell chromatin accessibility (scATAC-seq). Self-reported psychological stress and social support were associated with differential expression of 1,956 and 1,296 genes, respectively (10% FDR), primarily in CD4+ T cells and monocytes. Interferon signaling genes showed high expression in individuals with high psychological stress and low expression in those with high social support; this pattern mirrored gene expression in individuals with elevated circulating inflammatory markers (IFN-γ, TNF-α, IL-6). Genome-wide transcription factor (TF) motif analysis identified stress- and social support-associated changes in motif activity for 70 and 116 TFs, respectively, with 87 motifs enriched near differentially expressed genes. In CD4+ T cells, high psychological stress corresponded to increased IRF and STAT TF motif activity (interferon pathway), while social support was associated with reduced activity and expression in these pathways. We used an immune challenge paradigm (i.e., LPS stimulation), which confirmed the biological pathways of these gene regulatory effects. Our results demonstrate that psychological stress and social support modulate immune gene regulation at the single-cell level, revealing mechanistic links between psychosocial factors and inflammation, and suggesting that social support may promote immunological health.
Gut microbiomes of urban communities are compositionally different from their rural counterparts, and are associated with immune dysregulation and gastrointestinal disease. However, it is unknown whether these compositional differences impact host physiology, and through what mechanisms. Here, we used human colonic epithelial cells to directly compare host transcriptional changes induced by gut microbiomes from urban versus rural communities. We co-cultured host cells with live, stool-derived gut microbiomes from Rwanda, Ghana, Nigeria, Malaysia, and the United States, and quantified transcriptional responses using RNA-seq. We found that urban microbiomes affected innate immune pathways, including TNF signaling and bacterial antigen recognition. We also found that high-diversity microbiomes elicited a stronger host transcriptional response, while low-diversity microbiomes triggered epithelial restructuring and glycolysis. Finally, specific taxa driving these effects, including Bifidobacterium adolescentis and Bacteroides dorei, correlated with lifestyle factors such as diet. These findings demonstrate that urbanization-associated microbiome changes directly influence host epithelial gene expression.
Psychosocial experiences affect children's development and health. Urban environments in particular expose children to chronic stress, which impacts the HPA axis and leads to an altered glucocorticoid response. Because glucocorticoids have a major anti-inflammatory action, negative psychosocial experiences may be associated with immunological disorders, including asthma. However, the molecular mechanisms underlying the effect of negative psychosocial experiences on asthma and their interaction with genetic risk remain unknown.
Loss of function mutations in the checkpoint kinase gene CHEK2 are associated with increased risk of breast and other cancers. Most of the 3,188 unique amino acid changes that can result from non-synonymous single nucleotide variants (SNVs) of CHEK2, however, have not been tested for their impact on the function of the CHEK2-enocded protein (CHK2). One successful approach to testing the function of variants has been to test for their ability to complement mutations in the yeast ortholog of CHEK2, RAD53. This approach has been used to provide functional information on over 100 CHEK2 SNVs and the results align with functional assays in human cells and known pathogenicity. Here we tested all but two of the 4,887 possible SNVs in the CHEK2 open reading frame for their ability to complement RAD53 mutants using a high throughput technique of deep mutational scanning (DMS). Among the non-synonymous changes, 770 were damaging to protein function while 2,417 were tolerated. The results correlate well with previous structure and function data and provide a first or additional functional assay for all the variants of uncertain significance identified in clinical databases. Combined, this approach can be used to help predict the pathogenicity of CHEK2 variants of uncertain significance that are found in susceptibility screening and could be applied to other cancer risk genes.
Estimating and testing for differences in molecular phenotypes (e.g., gene expression, chromatin accessibility, transcription factor binding) across conditions is an important part of understanding the molecular basis of gene regulation. These phenotypes are commonly measured using high-throughput high-resolution count data that reflect how the phenotypes vary along the genome. Multiple methods have been proposed to help exploit these highresolution measurements for differential expression analysis. However, they ignore the count nature of the data, instead using normal distributions that work well only for data with large sample sizes or high counts. Here we develop count-based methods to address this problem. We model the data for each sample using an inhomogeneous Poisson process with spatially structured underlying intensity function and then, building on multiscale models for the Poisson process, estimate and test for differences in the underlying intensity function across samples (or groups of samples). Using both simulation and real ATAC-seq data, we show that our method outperforms previous normal-based methods, especially in situations with small sample sizes or low counts.
Genotype x environment interactions (GxE) have long been recognized as a key mechanism underlying human phenotypic variation. Technological developments over the past 15 years have dramatically expanded our appreciation of the role of GxE in both gene regulation and complex traits. The richness and complexity of these datasets also required parallel efforts to develop robust and sensitive statistical and computational approaches. Although our understanding of the genetic architecture of molecular and complex traits has been maturing, a large proportion of complex trait heritability remains unexplained. Furthermore, there are increasing efforts to characterize the effect of environmental exposure on human health. We therefore review GxE in human gene regulation and complex traits, advocating for a comprehensive approach that jointly considers genetic and environmental factors in human health and disease. Genotype x environment interactions are a key mechanism underlying human phenotypic variation and contribute to our understanding of the genetic architecture of human traits, with possible applications in personalized medicine.
Integrative genetic analysis of molecular and complex trait data, including colocalization analysis and transcriptome-wide association studies (TWAS), has shown promise in linking GWAS findings to putative causal genes (PCGs) underlying complex diseases. However, existing methods have notable limitations: TWAS tend to produce an excess of false-positive PCGs, while colocalization analysis often lacks sufficient statistical power, resulting in many false negatives. This paper introduces a probabilistic fine-mapping method, INTERFACE, which is designed to identify putative causal genes while accounting for direct variant-to-trait effects within genomic regions harboring multiple gene candidates. INTERFACE leverages interpretable, data-informed priors that incorporate both colocalization and TWAS evidence, enhancing the sensitivity and specificity of PCG inference and setting it apart from existing methods. Additionally, INTERFACE implements analytical measures to improve the accuracy of gene-to-trait effect estimation. We apply INTERFACE to METSIM plasma metabolite GWASs and UK Biobank pQTL data to identify causal genes regulating blood metabolite levels and demonstrate the unique biological insights INTERFACE provides. ### Competing Interest Statement The authors have declared no competing interest.
RNA sequencing (RNA-seq) is a high-throughput sequencing technique used to analyze and quantify the entire population or subsets of cellular RNA. This chapter presents experimental protocols and data analysis methods for polyadenylated RNA-seq, with a focus on research questions related to quantification of gene and isoform expression, comparison between groups, and studies of the genetic regulation of gene and isoform expression. A comprehensive discussion of study design considerations introduces important concepts and presents real-life situations to show researchers how to obtain well-balanced and calibrated datasets. These best practices apply beyond RNA-seq and also provide a solid foundation for other high-throughput sequencing studies.
Genetic variants in gene regulatory sequences can modify gene expression and mediate the molecular response to environmental stimuli. In addition, genotype–environment interactions (GxE) contribute to complex traits such as cardiovascular disease. Caffeine is the most widely consumed stimulant and is known to produce a vascular response. To investigate GxE for caffeine, we treated vascular endothelial cells with caffeine and used a massively parallel reporter assay to measure allelic effects on gene regulation for over 43,000 genetic variants. We identified 665 variants with allelic effects on gene regulation and 6 variants that regulate the gene expression response to caffeine (GxE, false discovery rate [FDR] < 5%). When overlapping our GxE results with expression quantitative trait loci colocalized with coronary artery disease and hypertension, we dissected their regulatory mechanisms and showed a modulatory role for caffeine. Our results demonstrate that massively parallel reporter assay is a powerful approach to identify and molecularly characterize GxE in the specific context of caffeine consumption.
BACKGROUND:Cardiovascular disease disproportionately affects African Americans. Psychosocial factors, including the experience of and emotional reactivity to racism and interpersonal stressors, contribute to the etiology and progression of cardiovascular disease through effects on health behaviors, stress-responsive neuroendocrine axes, and immune processes. The full pathway and complexities of these associations remain underexamined in African Americans. The Heart of Detroit Study aims to identify and model the biopsychosocial pathways that influence cardiovascular disease risk in a sample of urban middle-aged and older African American adults.METHODS:The proposed sample will be composed of 500 African American adults between the ages of 55 and 75 from the Detroit urban area. This longitudinal study will consist of two waves of data collection, two years apart. Biomarkers of stress, inflammation, and cardiovascular surrogate endpoints (i.e., heart rate variability and blood pressure) will be collected at each wave. Ecological momentary assessments will characterize momentary and daily experiences of stress, affect, and health behaviors during the first wave. A proposed subsample of 60 individuals will also complete an in-depth qualitative interview to contextualize quantitative results. The central hypothesis of this project is that interpersonal stressors predict poor cardiovascular outcomes, cumulative physiological stress, poor sleep, and inflammation by altering daily affect, daily health behaviors, and daily physiological stress.DISCUSSION:This study will provide insight into the biopsychosocial pathways through which experiences of stress and discrimination increase cardiovascular disease risk over micro and macro time scales among urban African American adults. Its discoveries will guide the design of future contextualized, time-sensitive, and culturally tailored behavioral interventions to reduce racial disparities in cardiovascular disease risk.
Integrative genetic association methods have shown great promise in post-GWAS (genome-wide association study) analyses, in which one of the most challenging tasks is identifying putative causal genes and uncovering molecular mechanisms of complex traits. Recent studies suggest that prevailing computational approaches, including transcriptome-wide association studies (TWASs) and colocalization analysis, are individually imperfect, but their joint usage can yield robust and powerful inference results. This paper presents INTACT, a computational framework to integrate probabilistic evidence from these distinct types of analyses and implicate putative causal genes. This procedure is flexible and can work with a wide range of existing integrative analysis approaches. It has the unique ability to quantify the uncertainty of implicated genes, enabling rigorous control of false-positive discoveries. Taking advantage of this highly desirable feature, we further propose an efficient algorithm, INTACT-GSE, for gene set enrichment analysis based on the integrated probabilistic evidence. We examine the proposed computational methods and illustrate their improved performance over the existing approaches through simulation studies. We apply the proposed methods to analyze the multi-tissue eQTL data from the GTEx project and eight large-scale complex-and molecular-trait GWAS datasets from multiple consortia and the UK Biobank. Overall, we find that the proposed methods markedly improve the existing putative gene implication methods and are particularly advantageous in evaluating and iden-tifying key gene sets and biological pathways underlying complex traits.
Puberty is an important developmental period marked by hormonal, metabolic and immune changes. Puberty also marks a shift in sex differences in susceptibility to asthma. Yet, little is known about the gene expression changes in immune cells that occur during pubertal development. Here we assess pubertal development and leukocyte gene expression in a longitudinal cohort of 251 children with asthma. We identify substantial gene expression changes associated with age and pubertal development. Gene expression changes between pre- and post-menarcheal females suggest a shift from predominantly innate to adaptive immunity. We show that genetic effects on gene expression change dynamically during pubertal development. Gene expression changes during puberty are correlated with gene expression changes associated with asthma and may explain sex differences in prevalence. Our results show that molecular data used to study the genetics of early onset diseases should consider pubertal development as an important factor that modifies the transcriptome.