BackgroundLung adenocarcinoma shows distinct differences between males and females in incidence, prognosis, and treatment response, suggesting unique molecular mechanisms that remain underexplored. This study aims to identify sex-specific molecular signatures and therapeutic targets in lung adenocarcinoma using multi-omics approaches to inform personalized treatment strategies.MethodsWe conducted an integrative analysis of transcriptomic and proteomic data from the Clinical Proteomic Tumor Analysis Consortium (CPTAC) and The Cancer Genome Atlas (TCGA) datasets, comparing male and female lung adenocarcinoma profiles. Transcription factor activity was assessed using TIGER on gene expression data, while kinase activity was evaluated with PTM-SEA on proteomic data. These results were combined to build a kinase-transcription factor signaling network. Potential sex-specific drugs were identified using the PRISM drug screening database.ResultsThe analysis revealed significant sex-based differences in transcription factor and kinase activity. Notably, NR3C1, AR, and AURKA exhibited sex-biased expression and activity. The constructed signaling network highlighted druggable pathways linked to cancer-related processes, with distinct profiles in males and females. PRISM screening identified glucocorticoid receptor agonists and aurora kinase inhibitors as promising sex-specific therapeutic candidates.ConclusionsOur findings underscore the importance of considering sex differences in lung adenocarcinoma molecular profiles. The integration of transcriptomic and proteomic data reveals sex-specific pathways and potential therapies, paving the way for personalized treatment approaches tailored to male and female patients.
RATIONALE Occupational exposures including Vapors, Gas, Dust or Fumes (VGDF) and additional Social Determinants of Health (SDOH) may contribute to adverse respiratory outcomes, although molecular causes are understudied. DNA methylation is influenced by multiple environmental exposures and may reveal novel insights into complex diseases. Studying the interaction between VGDF, SDOH and DNA methylation (DNAm) on lung function may reveal molecular mechanisms associated with COPD. METHODS COPDGene is a large-scale multicenter longitudinal cohort. The IlluminaEPIC array was used to assay leukocyte DNA methylation (DNAm) for 5,433 COPDGene blood samples from the five-year visit. VGDF were categorized by self-report of occupational exposures. Poverty was defined as self-reported annual income less than $15,000. Robust linear regression was performed across all samples to address the association between site-specific CpG methylation and lung function (FEV1), including the evaluation of interactions between VGDF and poverty and DNAm. All models were adjusted for age, age2, sex, height, height2, pack-years of smoking, cell proportions, genetic ancestry and the smoking-associated CpG site near AHRR (cg05575921). P-values were set at 10-3 for interactions. Gene set enrichment analysis was performed using GO and KEGG. Epigenetic age acceleration was calculated using the difference between the Horvath pan-tissue clock and chronologic age in years. RESULTS The prevalence of dusty jobs (yes, N=1405) and fumes (yes, N=1402) is 26%, with 786 subjects reporting exposures to both dust and fumes; 243 with VGDF were classified with income in the poverty range. We identified 4,101 CpGs demonstrating a significant interaction between methylation and fumes exposure, and 23,635 significant marks for the interaction term between dust exposure and DNAm in association with FEV1. There were 1,207 overlapping associations between CpGs for dust and fumes for FEV1. In a three-way comparison that included 2,201 associations for the interaction between poverty and DNAm, we identified 29 differentially methylated genes, including F2RL3, MYLK and ITPK1. Pathway enrichment included inflammatory/immune processes associated with VGDF and poverty. We observed epigenetic age acceleration of 1.8 years associated with occupational dust and 2.3 years associated with occupational fumes; epigenetic age acceleration increased to 2.7 years (95% CI: 2.0, 3.4) when VGDF exposures were associated with low income. CONCLUSIONS Lung function may be impacted by VGDF and poverty through epigenetic mechanisms. Analyses of VGDF and lung function highlight molecular associations related to both inflammation and accelerated aging, with further perturbation in the context of social disparities, thus revealing compounding effects on lung health.
Rationale: In COPD, women tend to be diagnosed earlier, experience more severe symptoms with less smoking exposure, and suffer from worse overall symptoms. One of the X chromosomes in each cell of females undergoes inactivation through DNA methylation to balance gene dosage with males, but some genes can escape X chromosome inactivation (XCI). These escape genes vary among individuals and tissues and have been linked to certain diseases. Our study aims to examine the variability of XCI patterns in women and their impact on COPD. Methods: We analyzed peripheral blood DNA methylation data in 2,995 women from Phase 1 and 2,626 from Phase 2 of the COPDGene longitudinal cohort, and blood gene expression for the XIST gene from Phase 2. XCI status was defined by gene promoter methylation levels. To evaluate COPD differences in women, we used logistic regression to test for associations between methylation escape status with COPD affection status, lung function (FEV1, FEV1/FVC) and emphysema (LAA-950), accounting for age, smoking history, blood cell counts, income, and education. Results: Methylation levels of X chromosome probes showed intra-class correlations between Phases 1 and 2 significantly higher than autosomal probes, suggestive of stable inactivation patterns of one X chromosome allele through methylation. We identified 66 (7%) X chromosome genes that appear to escape inactivation in blood. These escape genes exhibited greater variability over the 5-year time frame, and were enriched for GO terms related to gene expression regulation and immune processes. The escape status of seven genes (ALG13, ARSD, FIRRE, IQSEC2, LHFPL1, MIR223, MORF4L2-AS1) was associated with COPD affection status, worse lung function, and increased emphysema (p<0.05). XCI escape status was associated with income and education, highlighting how social determinants of health may influence epigenetic drivers of COPD in women. The negative impact of escape status on lung function was more pronounced in women with lower income and education levels. Finally, we found that the expression of XIST, a key regulator of XCI, is correlated with methylation of multiple escape genes (median Pearson coefficient of 0.2, p<0.05) suggesting a mechanistic association between XIST expression and escape status. Conclusion: Our results show the association of higher rates of XCI escape of multiple genes with COPD onset and progression. The association of X chromosome patterns and COPD outcomes varies by education and income, suggesting that social determinants of health might influence COPD disparities among women through biological processes related to the X chromosome.
Chronic obstructive pulmonary disease (COPD) often develops at an earlier age in women than in men, with worse respiratory symptoms despite lower smoking exposure. However, most preventive and therapeutic strategies ignore biological sex differences in COPD. Our goal was to better understand sex-specific gene regulatory processes in lung tissue and the molecular basis for sex differences in COPD onset and severity. We analyzed lung tissue gene expression and DNA methylation data from 747 individuals in the Lung Tissue Research Consortium and 85 individuals in an independent dataset. We identified sex differences in COPD-associated gene regulation using gene regulatory networks. We used linear regression to test for sex-biased associations of methylation with lung function, emphysema, smoking, and age. Analyzing gene regulatory networks in the control group, we identified that genes involved in the extracellular matrix (ECM) have higher transcriptional factor targeting in female subjects than in male subjects. However, this pattern is reversed in COPD, with men showing stronger regulatory targeting of ECM-related genes than women. Smoking exposure, age, lung function, and emphysema were all associated with sex-specific differential methylation of ECM-related genes. We identified sex-based gene regulatory patterns of ECM-related genes associated with lung function and emphysema. Multiple factors, including epigenetics, smoking, aging, and cell heterogeneity, influence sex-specific gene regulation in COPD. Our findings underscore the importance of considering sex as a key factor in disease susceptibility and severity.
BACKGROUND:Technological advances in sequencing and computation have allowed deep exploration of the molecular basis of diseases. Biological networks have proven to be a valuable framework for analyzing omics data and modeling regulatory interactions between genes and proteins. Large collaborative projects, such as The Cancer Genome Atlas (TCGA), have provided a rich resource for building and validating new computational methods, resulting in a plethora of open-source software for downloading, preprocessing, and analyzing those data. However, for an end-to-end analysis of regulatory networks, a coherent and reusable workflow is essential to integrate all relevant packages into a robust pipeline. FINDINGS:We developed tcga-data-nf, a Nextflow workflow that allows users to reproducibly infer regulatory networks from the thousands of samples in TCGA using a single command. The workflow can be divided into 3 main steps: multiomic data, such as RNA sequencing and methylation, are (i) downloaded, (ii) preprocessed, and (iii) analyzed to infer regulatory network models with the Network Zoo. The workflow is powered by the NetworkDataCompanion R package, a standalone collection of functions for managing, mapping, and filtering TCGA data. Here, we demonstrate how the pipeline can be used to investigate the differences between colon cancer subtypes attributed to epigenetic mechanisms. Lastly, we provide a database of pregenerated networks for the 10 most common cancer types that can be readily accessed by the public. CONCLUSIONS:tcga-data-nf is a complete, yet flexible and extensible, framework that enables the reproducible inference and analysis of cancer regulatory networks, bridging a gap in the current universe of software tools for analyzing TCGA data.
The rising incidence of lung cancer among individuals without a history of smoking highlights the need to explore biological mechanisms underlying this phenomenon. Our study aims to identify gene regulatory mechanisms that drive lung cancer risk among never-smokers by analyzing how gene regulatory networks differ between individuals with and without lung cancer, depending on their smoking history. We used RNA-Seq data from the Lung Tissue Research Consortium (LTRC) collected via TOPMed, comprising non-cancerous lung tissue samples from 344 individuals with non-small cell carcinoma and 329 lung tissue samples from individuals without cancer. Across all, 18% reported no history of smoking. We ran differential gene expression analysis by voom, adjusting for age, sex, COPD status, and examining statistical interactions between smoking (never/former) and cancer status. We analyzed sample-specific transcription factor-gene regulatory networks generated by PANDA-LIONESS. We used linear regression to evaluate the association between gene targeting score (measured by gene indegree) and the interaction between smoking and cancer status. We ran gene set enrichment analysis with genes ranked by the corresponding smoking by cancer interaction coefficients. Differential expression analysis of lung tissues from individuals with and without cancer showed that genes overexpressed in cancer are enriched for canonical cancer pathways, such as p53, MAPK, and WNT signaling pathways (FDR<0.05). We found a significant interaction between cancer and smoking status, indicating these cancer-related pathways had higher enrichment in former-smokers compared to never smokers. Gene regulatory network analysis indicated that these differential expression patterns are possibly driven by increased transcriptional targeting of these cancer-related pathways. Among never-smokers, individuals with cancer showed higher targeting of metabolic pathways (e.g., fructose and mannose, glutathione, phenylalanine, histidine, arginine, and proline metabolism) compared to lung tissue from individuals without cancer (FDR<0.05), highlighting a possible role of metabolic pathways in tumor development, uniquely among never-smokers. A negative correlation between metabolic pathway targeting scores and age was observed exclusively in individuals with cancer without a history of smoking (Mean Pearson R=-0.24). Our findings reveal transcriptional targeting differences in lung cancer by smoking history. Among individuals without a history of smoking, the increased targeting of metabolic pathways in cancer, which is pronounced in younger compared to older individuals, may contribute to the higher risk of early-onset lung cancer among never-smokers, offering new avenues for understanding and addressing the disease in populations without a history of smoking. Camila Lopes-Ramos, Enakshi Saha, Jeong Yun, Craig Hersh, Edwin Silverman, Dawn DeMeo, John Quackenbush, Kimberly Glass. Lung cancer gene regulatory networks reflect smoking history and age-related metabolic pathway alterations [abstract]. In: Proceedings of the American Association for Cancer Research Annual Meeting 2025; Part 1 (Regular Abstracts); 2025 Apr 25-30; Chicago, IL. Philadelphia (PA): AACR; Cancer Res 2025;85(8_Suppl_1):Abstract nr 7478.
Rationale: Dietary polyunsaturated fatty acids (PUFAs) found in plant and seed oils and fatty fish have immunomodulatory properties that may influence lung function. Smoking impacts DNA methylation, inflammation, and lung function and may mediate the effects of PUFAs on the lungs. We analyzed omega-3 and omega-6 PUFA biomarker associations with lung function, and investigated whether these associations are mediated through DNA methylation of the AHRR gene, an epigenetic marker of smoking dose and duration. Methods: We evaluated spirometry, AHRR DNA methylation and PUFA biomarker data from 3854 participants in the Genetic Epidemiology of COPD (COPDGene) Study. Infinium MethylationEPIC BeadChip DNAm data was assayed through the NIH NHLBI TOPMed program using peripheral blood leukocyte DNA. Plasma PUFAs were quantified by gas chromatography-mass spectrometry (GC-MS). PUFA associations with FEV1 were estimated using robust linear regression accounting for key lung function covariates. Mediation analyses estimating the average causal effect of PUFAs mediated through AHRR (cg05575921) methylation were performed with the R package ‘mediation’. Results: Higher blood levels of omega-3 PUFAs alpha-linolenic acid (ALA) and eicosapentaenoic acid (EPA) were positively associated with FEV1 while omega-6 PUFA arachidonic acid (AA) was negatively associated with FEV1. Specifically, one standard deviation (SD) higher ALA and EPA were associated with 33.0 (95% confidence interval [CI] 11.2 – 54.8) and 42.4 (95% CI 22.9 – 61.8) ml higher FEV1, while one SD higher AA was associated with 84.3 (95% CI 61.6-107.0) ml lower FEV1. Associations of docosahexaenoic acid (DHA), docosapentaenoic acid (DPA), and linoleic acid (LA) with FEV1 were not statistically significant. Mediation analyses revealed significant (p<2E-16) causal mediation through AHRR methylation for EPA and ALA, with the proportion mediated estimated at 0.27 and 0.19 respectively, while there was no statistical evidence (p=0.18) of mediation for AA. In these models, increases in AHRR methylation attributed to one SD higher blood levels of ALA and EPA were associated with 11.6 (95% CI 7.5 – 16.4) and 9.0 (95% CI 5.14 – 13.9) ml higher FEV1. Analyses stratified by smoking status suggest these effects may be realized primarily in former smokers. Conclusions: Omega-3 PUFAs may impact lung function through influences on smoking-associated DNA methylation levels, while the impacts of omega-6 PUFAs may be independent of the epigenetic effects of smoking. These findings shed light on molecular mechanisms underlying intersections of PUFA nutritional status, smoking, and lung function, and highlight a future role for precision nutrition for lung health.
Lung adenocarcinoma (LUAD) exhibits differences between the sexes in incidence, prognosis, and therapy, suggesting underexplored molecular mechanisms. We conducted an integrative multi-omics analysis using the Clinical Proteomic Tumor Analysis Consortium (CPTAC) and The Cancer Genome Atlas (TCGA) datasets to contrast transcriptomes and proteomes between sexes. We used TIGER to analyze TCGA-LUAD expression data and found sex-biased activity of transcription factors (TFs); we used PTM-SEA with CPTAC-LUAD proteomics data and found sex-biased kinase activity. We combined these to construct a kinase-TF signaling network and discovered druggable pathways linked to cancer-related processes. We also found significant sex biases in clinically relevant TFs and kinases, including NR3C1, AR, and AURKA. Using the PRISM drug screening database, we identified potential sex-specific drugs, such as glucocorticoid receptor agonists and aurora kinase inhibitors. Our findings emphasize the importance of considering sex and using multi-omics network methods to discover personalized cancer therapies.
Aging is the primary risk factor for many cancer types, including lung adenocarcinoma (LUAD). To understand how aging-related alterations in the regulation of key cellular processes might affect LUAD risk and survival, we built individual-specific gene regulatory networks integrating gene expression, transcription factor protein-protein interaction, and sequence motif data, using PANDA/LIONESS algorithms, for non-cancerous lung samples from GTEx project and LUAD samples from TCGA. In healthy lung, pathways involved in cell proliferation and immune response were increasingly targeted with age; these aging-associated alterations were accelerated by smoking and resembled oncogenic shifts observed in LUAD. Aging-associated genes showed greater aging-biased targeting patterns in individuals with LUAD compared to healthier counterparts, a pattern suggestive of age acceleration. Using drug repurposing tool CLUEreg, we found small molecule drugs that may potentially alter the accelerating aging profiles we found. We defined a network-informed aging signature that was associated with survival in LUAD.
Lung adenocarcinoma (LUAD) has been observed to have significant sex differences in incidence, prognosis, and response to therapy. However, the molecular mechanisms responsible for these disparities have not been investigated extensively. Sample-specific gene regulatory network methods were used to analyze RNA sequencing data from non-cancerous human lung samples from The Genotype Tissue Expression Project (GTEx) and lung adenocarcinoma primary tumor samples from The Cancer Genome Atlas (TCGA); results were validated on independent data. We observe that genes associated with key biological pathways including cell proliferation, immune response and drug metabolism are differentially regulated between males and females in both healthy lung tissue, as well as in tumor, and that these regulatory differences are further perturbed by tobacco smoking. We also uncovered significant sex bias in transcription factor targeting patterns of clinically actionable oncogenes and tumor suppressor genes, including AKT2 and KRAS. Using differentially regulated genes between healthy and tumor samples in conjunction with a drug repurposing tool, we identified several small-molecule drugs that might have sex-biased efficacy as cancer therapeutics and further validated this observation using an independent cell line database. These findings underscore the importance of including sex as a biological variable and considering gene regulatory processes in developing strategies for disease prevention and management.
Gene regulatory networks (GRNs) are effective tools for inferring complex interactions between molecules that regulate biological processes and hence can provide insights into drivers of biological systems. Inferring coexpression networks is a critical element of GRN inference, as the correlation between expression patterns may indicate that genes are coregulated by common factors. However, methods that estimate coexpression networks generally derive an aggregate network representing the mean regulatory properties of the population and so fail to fully capture population heterogeneity. Bayesian optimized networks obtained by assimilating omic data (BONOBO) is a scalable Bayesian model for deriving individual sample-specific coexpression matrices that recognizes variations in molecular interactions across individuals. For each sample, BONOBO assumes a Gaussian distribution on the log-transformed centered gene expression and a conjugate prior distribution on the sample-specific coexpression matrix constructed from all other samples in the data. Combining the sample-specific gene coexpression with the prior distribution, BONOBO yields a closed-form solution for the posterior distribution of the sample-specific coexpression matrices, thus allowing the analysis of large data sets. We demonstrate BONOBO's utility in several contexts, including analyzing gene regulation in yeast transcription factor knockout studies, the prognostic significance of miRNA-mRNA interaction in human breast cancer subtypes, and sex differences in gene regulation within human thyroid tissue. We find that BONOBO outperforms other methods that have been used for sample-specific coexpression network inference and provides insight into individual differences in the drivers of biological processes.
There is increasing recognition that the sex chromosomes, X and Y, play an important role in health and disease that goes beyond the determination of biological sex. Loss of the Y chromosome (LOY) in blood, which occurs naturally in aging men, has been found to be a driver of cardiac fibrosis and heart failure mortality. LOY also occurs in most solid tumors in males and is often associated with worse survival, suggesting that LOY may give tumor cells a growth or survival advantage. We analyzed LOY in lung adenocarcinoma (LUAD) using both bulk and single-cell expression data and found evidence suggesting that LOY affects the tumor immune environment by altering cancer/testis antigen expression and consequently facilitating tumor immune evasion. Analyzing immunotherapy data, we show that LOY and changes in expression of particular cancer/testis antigens are associated with response to pembrolizumab treatment and outcome, providing a new and powerful biomarker for predicting immunotherapy response in LUAD tumors in males.
Correlation networks can provide important insights into biological systems by uncovering intricate interactions between genes and their molecular regulators. However, methods for estimating co-expression networks generally derive an aggregate population-specific network that represents the mean regulatory properties of the entire population and hence falls short in capturing heterogeneity across individuals. While numerous methods have been proposed to estimate sample-specific co-expression networks, they fail to estimate positive semidefinite correlation networks and, hence, are subject to misinterpretation. To fill this gap in co-expression network inference, we introduce BONOBO (Bayesian Optimized Networks Obtained By assimilating Omics data), a scalable Bayesian model for deriving individual sample-specific co-expression networks by acknowledging heterogeneity in molecular interactions across individuals. For each sample, BONOBO imposes a Gaussian distribution on the log-transformed, centered gene expression and a conjugate Inverse Wishart prior distribution on the sample-specific co-expression matrix constructed from assimilating all other samples in the data. BONOBO yields a closed-form solution for the posterior distribution of the sample-specific co-expression matrices by combining the sample-specific gene expression with the prior distribution. We demonstrate the advantages of BONOBO using several simulated and real datasets. BONOBO is computationally scalable and available as open-source software through the Network Zoo package (from netZooPy v0.10.0; netzoo.github.io). A preprint associated with this abstract can be found on bioRxiv (doi: 10.1101/2023.11.16.567119v1).
Abstract Sex differences in lung adenocarcinoma (LUAD) are evident in incidence rates, prognostic outcomes, and therapy responses, yet the underlying molecular mechanisms driving these disparities remain underexplored. In this study, we conducted a comprehensive proteogenomic analysis encompassing 38 females and 73 males with LUAD from the Clinical Proteomic Tumor Analysis Consortium (CPTAC) dataset. Employing Transcription Inference using Gene Expression and Regulatory data (TIGER), we inferred sex-differentially activated transcription factors (TFs) from The Cancer Genome Atlas (TCGA) LUAD gene expression data and identified sex-differentially activated kinases using CPTAC protein phosphorylation data. We further constructed a comprehensive kinase-TF signaling network by integrating these sex-differentially activated kinases with TFs, identifying all paths shorter than 3 in the protein interaction networks to highlight druggable pathways. Our analyses revealed that many proteins exhibit not only sex-biased abundance but also sex-biased phosphorylation and acetylation. Furthermore, these sex-biased proteins were associated with critical biological pathways including cell proliferation, immune response, and metabolism. Using kinase-TF signaling networks, we found substantial sex bias in the activities of clinically actionable TFs and kinases, including the glucocorticoid receptor (NR3C1), AR, AURKA, CDK6, and MAPK14. Leveraging the PRISM cancer cell line screening database, we identified several small-molecule drugs, such as glucocorticoid receptor agonists and aurora kinase inhibitors, potentially exhibiting sex-specific efficacy as LUAD therapeutics. Our findings showed that the activity of some clinically relevant TFs and kinases differ by sex in LUAD, underscoring the need to consider sex as a biological variable and the utility of multi-omics integrative protein signaling networks in advancing our understanding of cancer biology and the development of sex-aware therapeutics. Citation Format: Chen Chen, Enakshi Saha, Dawn L. DeMeo, John Quackenbush, Camila M. Lopes-Ramos. Unveiling sex differences in lung adenocarcinoma through multi-omics integrative protein signaling networks [abstract]. In: Proceedings of the American Association for Cancer Research Annual Meeting 2024; Part 1 (Regular Abstracts); 2024 Apr 5-10; San Diego, CA. Philadelphia (PA): AACR; Cancer Res 2024;84(6_Suppl):Abstract nr 3490.
Abstract Background The association between genetic variants on the X chromosome to risk of COPD has not been fully explored. We hypothesize that the X chromosome harbors variants important in determining risk of COPD related phenotypes and may drive sex differences in COPD manifestations. Methods Using X chromosome data from three COPD-enriched cohorts of adult smokers, we performed X chromosome specific quality control, imputation, and testing for association with COPD case–control status, lung function, and quantitative emphysema. Analyses were performed among all subjects, then stratified by sex, and subsequently combined in meta-analyses. Results Among 10,193 subjects of non-Hispanic white or European ancestry, a variant near TMSB4X, rs5979771, reached genome-wide significance for association with lung function measured by FEV1/FVC ( $$\beta$$ β 0.020, SE 0.004, p 4.97 × 10–08), with suggestive evidence of association with FEV1 ( $$\beta$$ β 0.092, SE 0.018, p 3.40 × 10–07). Sex-stratified analyses revealed X chromosome variants that were differentially trending in one sex, with significantly different effect sizes or directions. Conclusions This investigation identified loci influencing lung function, COPD, and emphysema in a comprehensive genetic association meta-analysis of X chromosome genetic markers from multiple COPD-related datasets. Sex differences play an important role in the pathobiology of complex lung disease, including X chromosome variants that demonstrate differential effects by sex and variants that may be relevant through escape from X chromosome inactivation. Comprehensive interrogation of the X chromosome to better understand genetic control of COPD and lung function is important to further understanding of disease pathology. Trial registration Genetic Epidemiology of COPD Study (COPDGene) is registered at ClinicalTrials.gov, NCT00608764 (Active since January 28, 2008). Evaluation of COPD Longitudinally to Identify Predictive Surrogate Endpoints Study (ECLIPSE), GlaxoSmithKline study code SCO104960, is registered at ClinicalTrials.gov, NCT00292552 (Active since February 16, 2006). Genetics of COPD in Norway Study (GenKOLS) holds GlaxoSmithKline study code RES11080, Genetics of Chronic Obstructive Lung Disease.
Additional file 4: Table S3. Replication Examination Of Associations From Other Studies In This XWAS Meta-analysis
Differential network topology contrasting the short-term with the long-term survival network. A) Differential network modules significantly enriched for overrepresentation of Gene Ontology terms. Modules are shown for both discovery datasets (D1 and D2). Two modules were significant in both dataset (indicated with M1 and M2). The color in the heatmap indicates the enrichment of genes in the module as observed/expected value (Obs/Exp), with purple representing enriched modules in discovery dataset 1 and green enriched modules in discovery dataset 2. B) "Beeswarm" plot visualizing the distribution of average log differential modularity scores from ALPACA, for transcription factors (TF) and genes (Gene). Higher scores mean higher differential modularity. TFs and genes with significant differential modularity between the two survival groups are labeled.
Additional file 3: Table S2. XWAS Meta-analysis Top ACE2 Xp22.2 Variants
We collected studies in glioblastoma that included survival thresholds through a PubMed search (search term: glioblastoma short-term prognostic signature). As can be seen from this table, a wide range of survival thresholds is used, with thresholds for short-term survival ranging between 7-24 months, and thresholds for long-term survival between 13-36 months after initial diagnosis. The threshold of 1.7yrs, which we chose for our comparison in the discovery datasets, corresponds to 20 months, which lies right in between these two ranges.
To ensure that our results were stable across multiple survival thresholds, we repeated the comparative network analysis in the two discovery datasets across multiple survival thresholds, which we ranged from zero to the maximum follow-up, with intervals of 0.1 years, keeping comparisons that had group sizes of at least ten patients. This figure reports the -log10FDR and GSEA enrichment scores (ES) for the different thresholds, for A) discovery dataset 1 and B) discovery dataset 2. Significant results (FDR<0.05) are highlighted with red circles.