There are significant sex differences in cancer incidence, yet the underlying regulatory mechanisms in normal tissues remain poorly understood. We studied 8,279 gene regulatory networks across 29 non-cancerous tissues and compared network centrality by sex. Cancer genes were differentially targeted by transcription factors in males and females, with an overrepresentation on the X chromosome, particularly among X-inactivation escapees, and key signaling pathways such as WNT, NOTCH, and p53. We observed higher targeting of cancer-related pathways in females for tissues that have higher tumor incidence in females (breast, lung, and thyroid) and higher targeting in males for tissues with increased tumor incidence in males (stomach, colon, and liver), a pattern replicated in independent lung data. Sex-biased transcription factors were enriched for sex hormone response elements. These findings suggest that sex-biased transcriptional programs in normal tissues contribute to sex differences in cancer incidence and should be considered in cancer prevention strategies.
Gene regulatory networks (GRNs) provide a mechanistic framework for understanding how transcription factors coordinate gene expression to establish cellular identity and phenotype. Methods that integrate gene expression with motif-derived regulatory priors and other sources of biological information have substantially advanced gene regulatory network inference by reconstructing condition-specific regulatory architecture. These approaches estimate the evidence supporting regulatory interactions and have proven remarkably successful in a wide range of biological applications. A complementary view of regulatory networks, however, seeks to estimate the effect of those interactions on gene expression itself, providing a framework in which regulatory edges can be interpreted as activating or inhibitory influences on transcription. We developed GIRAFFE, a biologically informed matrix factorization framework that jointly estimates transcription factor activities and gene regulatory networks by integrating gene expression, motif-based regulatory priors, and transcription factor protein-protein interactions. GIRAFFE estimates signed partial regulatory effects whose magnitude and sign can be interpreted as the strength and direction of transcriptional regulation. Building directly on the biological framework established by methods such as PANDA, GIRAFFE provides a complementary representation of gene regulatory networks that emphasizes mechanistic interpretation while remaining scalable, flexible, and computationally efficient. Across synthetic benchmarks, six human tissues, yeast transcription factor perturbation experiments, and liver hepatocellular carcinoma, GIRAFFE accurately reconstructs regulatory interactions while distinguishing activating from inhibitory regulation with high accuracy. The inferred networks recover known features of tissue-specific regulation, correctly classify regulatory effects in transcription factor perturbation experiments, and identify biologically coherent changes in regulatory programs associated with liver cancer. Together, these results demonstrate that estimating the direction of transcriptional regulation provides a complementary perspective on gene regulatory networks that facilitates biological interpretation and hypothesis generation.
Background Glioblastoma or GBM (IDH wild-type) is an aggressive brain tumor that is notoriously resistant to treatment, with an average survival time of 17 months. While the overall outcome is poor for both males and females, sex differences in GBM incidence and outcome suggest sex-specific biological mechanisms underlie tumorigenesis. In contrast, low-grade glioma (LGG) is a less aggressive brain tumor that tends to have a better prognosis and a longer survival time.Methods To understand mechanisms contributing to treatment resistance in GBM in both males and females, we inferred gene regulatory networks (GRNs) for males and females with LGG and GBM using RNA-seq data from The Cancer Genome Atlas (TCGA). We analyzed these to identify both sex-specific and sex-stratified gene regulation in GBM. We then validated these results on a separate cohort, the Repository of Molecular BRAin Neoplasia DaTa (REMBRANDT).Results We found sex-specific differential targeting of several pathways, including hypoxia and related pathways (carbohydrate metabolism, innate immune processes, and extracellular matrix pathways) known to be dysregulated in hypoxic conditions, in GBM when compared against LGG. After further evaluating the co-regulation of sex-specific pathways in GBM, we found that females exhibited a greater degree of co-regulation between hypoxia and hypoxia-associated transcriptional programs with the aforementioned downstream pathways than did males.Conclusions Our results suggest that dysregulation of hypoxia-related pathways in GBM plays a female-specific role in resistance to treatment and overall outcomes.
Abstract Background: There is currently no ovarian cancer biomarker appropriate for screening which is partially due to using retrospective clinical samples obtained at the time of diagnosis for biomarker discovery. Thus, we sought to discover novel plasma proteomic biomarkers for ovarian cancer early detection using prospectively collected blood samples. Method: We evaluated 10,778 plasma proteins measured using the SomaScan v5.0 assay in blood drawn at least three years prior to ovarian cancer diagnosis and matched controls in the Prostate, Lung, Colorectal and Ovarian Cancer Screening Trial (PLCO; n=98, training dataset) and the Nurses’ Health Studies (NHS; n=99, replication dataset). We used a conditional logistic regression to identify individual proteins associated with ovarian cancer in the two datasets separately. We also compared plasma proteins for early-stage and late-stage ovarian cancer in blood collected at diagnosis in the PreOperative Pelvic Mass Study to age-matched population-based controls (PreOp; n=134). Then we used Elastic Net to develop a proteomic-based score to discriminate ovarian cancer cases from controls in PLCO, compared to a model with CA125 alone, and validated the proteomic-based score performance in NHS by calculating the area under the receiver operating characteristic curve (AUC) and 95% confidence interval (CI). Results: Plasma proteins associated with ovarian cancer in blood samples collected prospectively differed from those associated with blood samples collected at diagnosis of early-stage disease compared to controls. There were 99 proteins associated with ovarian cancer diagnosed at least 3 years from blood collection (p<0.05) in PLCO, where 2 proteins, RCN3 and OBP2B, replicated in NHS (p<0.05). Of these 99 proteins, majority were not associated with ovarian cancer in PreOp and only three proteins overlapped (i.e., SERPINF2, ASAH2, BAGE3). In PLCO, adding a proteomic-based score comprised of 56 proteins to a model with CA125 alone significantly (p=0.02) improved discriminating ovarian cancer cases from controls with an AUC (95%CI) from 0.65(0.50-0.80) to 0.86(0.67,1.00). In NHS, proteomic-based score resulted in an AUC of 0.60(0.49-0.72) with marginal significance. Conclusion: Our results revealed plasma proteomic profiles differ between prospectively collected blood samples at least 3 years prior to diagnosis and blood samples collected at time of diagnosis regardless of stage. We developed a proteomics-based score that improved upon CA-125 alone, although application to an independent cohort did not demonstrate a strong improvement. However, differences between studies (e.g., menopausal status and hormone therapy use) may explain this variation. Citation Format: Nan Lin, Ngo Long, Allison F. Vitonis, Tara Eicher, SHELLEY TWOROGER, Simon T. Dillon, Towia A. Libermann, Daniel W. Cramer, John Quackenbush, Kathryn L. Terry, Naoko Sasamoto. Development and validation of a plasma proteomics signature for earlier diagnosis of ovarian cancer using prospectively collected blood samples [abstract]. In: Proceedings of the American Association for Cancer Research Annual Meeting 2026; Part 1 (Regular Abstracts); 2026 Apr 17-22; San Diego, CA. Philadelphia (PA): AACR; Cancer Res 2026;86(7 Suppl):Abstract nr 2310.
A growing body of literature has investigated the relationship between age and biomolecular changes, leading to conclusions that aging occurs in discrete molecular "waves." Data summary tools such as LOESS and sliding window analyses like DE-SWAN are common approaches that have gained acceptance in recent years. We demonstrate via simple simulations that these tools can identify non-linear patterns of aging where they do not exist. Specifically, we show that (i) clustering of molecular trajectories using LOESS can lead to artifactual characteristic patterns of molecular aging, (ii) "waves" of aging identified using the combination of LOESS and DE-SWAN in real data are not robust to changes in the underlying age distribution and are not supported by valid permutation testing, and (iii) DE-SWAN alone can generate pronounced "waves" of nonlinear molecular aging in linear data due to differences in statistical power along the age continuum. Our results specifically challenge the statistical support for discrete aging crests inferred in the literature, but do not rule out nonlinear molecular aging or age-associated transitions that may be detectable using other cohorts and statistical models.
Motivation:Gene regulatory networks undergo dynamic restructuring during development and disease. Identifying when and how these networks change is crucial for understanding developmental and disease transitions, yet existing change-point detection methods often ignore network structure or lack interpretable community assignments. Results:We present PARROT (Phase-Altering Regulatory Rewiring Over Time), a framework for detecting change-points in dynamic networks using Stochastic Block Models. PARROT jointly estimates change-point locations and community structure across four network classes: unipartite and bipartite with either Gaussian or Bernoulli edge models. Simulations demonstrate improved performance and community recovery compared to other methods. Applications to human cardiac differentiation and mouse lung development data successfully recovered known phase boundaries. PARROT identifies both which genes are reassigned across modules and how the connections change between states. Availability:PARROT is available as an R package at https://github.com/cchen22/PARROT. Contact:chenchen9945@gmail.com. Supplementary information:Supplementary data are available at Bioinformatics online.
The Lancet Commission on the Future of Care and Clinical Research in Autism proposed the construct of "profound autism" as a recognizable subtype of autism. Supporters argue that this classification is necessary to ensure that autistic persons with severe impairment receive appropriate research attention and policy support, whereas critics contend that the construct lacks scientific validity and may reflect social or political considerations more than biological distinction. To inform this debate, we evaluate whether the proposed "profound autism" category represents a distinct genetic phenotype using multiple molecular data types collected in a large cohort. Across genomic, transcriptomic, and regulatory analyses, we find no evidence supporting "profound autism" as a biologically distinct phenotypic group. Instead, differences emerge primarily in inferred gene regulatory networks distinguishing nonspeaking from speaking autistic children, suggesting potential regulatory mechanisms contributing to speech ability. These findings suggest that future research into severe impairment may be more productive if focused on specific traits-such as speech impairment-rather than attempting to define a distinct biological subtype within the multidimensional phenomenon of autism.
Techniques for evaluating gene regulatory network (GRN) inference methods typically focus on recovering small ground-truth networks or on benchmarking against simulated data. However, both approaches have important limitations and fail to capture the biological variability present in real datasets. FERRET is a framework for benchmarking single-cell GRN inference methods based on a simple biological assumption: independent estimates of the regulatory network from the same cellular state should resemble one another more closely than estimates from distinct cellular states. Rather than relying on incomplete or simulated ground truth, FERRET quantifies network robustness using two complementary metrics: Robustness Area Under the Curve ( RAUC ), an AUC-like measure of within-cell-type network similarity relative to between-cell-type similarity, and Monotonicity , which assesses the consistency of network similarity across edge-weight cutoffs. FERRET also supports biological validation through pathway enrichment analysis. We validate FERRET using experimentally derived ChIP-seq networks from B lymphocytes and fibroblasts as positive controls and randomly generated networks as negative controls, showing that biologically related networks receive high robustness scores whereas randomly generated networks receive scores consistent with chance. Finally, we apply FERRET to multiple GRN inference methods on real single-cell RNA-sequencing datasets to identify methods that produce the most robust, biologically informative regulatory networks. GRAPHICAL ABSTRACT:
In analyzing gene regulatory network models, a common question is how members of a particular set of genes are connected. For example, one might want to explore network relationship between a set of differentially expressed genes, a gene set previously reported in the literature, or elements of one or more pathways. BLOBFISH uses a breadth-first search algorithm adapted to bipartite graphs to identify a compact subnetwork connecting the members of a pre-specified set of genes, providing a regulatory context that can shed light on specific mechanisms involved in a phenotype and its development. We demonstrate the use of BLOBFISH to extract gene regulatory subnetworks reflecting tissue specificity using publicly available data from the Genotype Tissue Expression (GTEx) project.
We previously described three subtypes of Waldenstrom's Macroglobulinemia (WM) using a multi-omic data set derived from 249 untreated patients with WM who had detectable MYD88 mutations. The B-cell like (BCL) and Plasma cell like (PCL) subtypes have distinct cell of origin but evolve from a shared transcriptional subtype (Early WM) found at early stages of disease development (Hunter et al, Manuscript in Review. Research Square, preprint 2025). We also described the WM evolutionary score (EScore) that corresponds to the progression from smoldering WM to symptomatic disease correlating with time to first therapy and WM bone marrow involvement. The transcription factor (TF) network differences that underlie these findings remain poorly understood. In this study, we examined changes in implied gene regulatory networks between WM and healthy donor memory B-cells (HDMB; CD19+CD27+), with EScore, and between WM subtypes. Findings were further analyzed for expression, alternative splicing, and TF motif enrichment in differentially regulated ATAC peaks. Sample specific gene regulatory networks were modeled using PANDA (Glass et al. Plos One, 2013) followed by LIONESS (Kuijer et al. iScience, 2019). Together, these calculate an interaction score for each TF/gene pair for per sample. These scores were summed up by TF and by gene to model to net regulatory connectivity known as the target score (TS). LIONESS and TS data were analyzed using limma in R and top networks were derived from the results. Top TF and gene hits were analyzed further for changes in expression and isoform composition based on Gencode augmented by novel isoforms discovered in our IsoSeq long read analysis of HDMB and WM samples (Hunter et al, ASH 2024). TF motif enrichment from differentially regulated ATACSeq peaks were calculated using MEME Suite's SEA. Network analysis of EScore TF enrichment revealed AHR and TFAP4 as top drivers. AHR is a regulator of xenobiotic metabolism which was already known to be negatively associated with EScore from our previous Hallmark gene set enrichment analysis of the expression data. Expression of AHR was negatively correlated with EScore (r = -0.73; p<0.0001) with the median expression falling by a factor of 12 from the earliest to latest EScore levels. AHR has been reported to inhibit B-cell class switching and plasma cell differentiation while upregulating the pro-inflammatory PTGS2 (cyclooxygenase-2) and anti-apoptotic BCL2L1 (BCL-XL) genes. The correlation with PTGS2 (r=0.641; p<0.0001) and BCL2L1 (r=0.266; p<0.0001) was observed in this study. Notably BCL2 was negatively correlated (r=-0.552; p<0.0001) consistent with a change in BCL2 family dependency during progression to symptomatic disease. Unlike AHR, TFAP4 was both transcriptionally upregulated and showed increased TS with EScore, consistent with its reported role as an oncogene. Most of the AP-1 transcription factor family also appeared in the top EScore network results including JUNB, FOS, FOSL1, FOSL2, ATF3, and MAF. Moreover, JUNB, FOS, FOSL2, and FOSB were all top enriched motifs in differentially closed chromatin between early and late EScore (p<0.0001 for all). Transcriptional levels of FOS and FOSL2 decreased >15x with EScore (p<0.0001 for both) while FOSB and JUNB levels were largely steady despite notable TS enrichment. Isoform analysis further revealed that while FOSB expression in WM is over 2x higher than HDMB expression, HDMB expression is primarily intron retained (ENST00000587358) while WM primarily expressed the full-length transcript (ENST00000353609). Comparing HDMB to WM Subtypes revealed strong differences in TF and gene TS with corresponding isoform dysregulation including PRDM1, PRDM4, IL15, RCAN3, CRIP2, and CD19. Despite increased plasmacytic differentiation, CD19 was expressed at only slightly lower levels than HDMB. However, HDMB primarily expressed a novel isoform skipping exon 2 while WM expressed the canonical transcript (p <0.0001). Given the important role of CD19 in BCR and TLR signaling, our findings indicate a critical dependency for CD19 in WM. Our extensive regulatory network analysisutilizing a large multi-omic data set identified several novel therapeutic targets including TFAP4 and CD19, as well as the AHR, TFAP4, and AP-1 transcription factors as key drivers of inflammatory signaling, BCL2 family dysregulation and malignant cell growth for symptomatic disease progression in WM.
BackgroundLung adenocarcinoma shows distinct differences between males and females in incidence, prognosis, and treatment response, suggesting unique molecular mechanisms that remain underexplored. This study aims to identify sex-specific molecular signatures and therapeutic targets in lung adenocarcinoma using multi-omics approaches to inform personalized treatment strategies.MethodsWe conducted an integrative analysis of transcriptomic and proteomic data from the Clinical Proteomic Tumor Analysis Consortium (CPTAC) and The Cancer Genome Atlas (TCGA) datasets, comparing male and female lung adenocarcinoma profiles. Transcription factor activity was assessed using TIGER on gene expression data, while kinase activity was evaluated with PTM-SEA on proteomic data. These results were combined to build a kinase-transcription factor signaling network. Potential sex-specific drugs were identified using the PRISM drug screening database.ResultsThe analysis revealed significant sex-based differences in transcription factor and kinase activity. Notably, NR3C1, AR, and AURKA exhibited sex-biased expression and activity. The constructed signaling network highlighted druggable pathways linked to cancer-related processes, with distinct profiles in males and females. PRISM screening identified glucocorticoid receptor agonists and aurora kinase inhibitors as promising sex-specific therapeutic candidates.ConclusionsOur findings underscore the importance of considering sex differences in lung adenocarcinoma molecular profiles. The integration of transcriptomic and proteomic data reveals sex-specific pathways and potential therapies, paving the way for personalized treatment approaches tailored to male and female patients.
Objective:Most, if not all, thyroid disorders are more prevalent in females than males. However, sex differences in thyroid disease risk varies with age, e.g. although the risk of anaplastic thyroid carcinoma (ATC), the most aggressive form of thyroid cancer, increases with age for both sexes, age of diagnosis for most thyroid cancers is lower for females than males. In contrast, the risk of Hashimoto's thyroiditis (HT), an autoimmune condition, is higher in ages 30-50 than older age groups, and females have a higher risk than males at any given age. These age- and sex-dependent variations suggest that thyroid aging is a sex-biased process, where gene regulatory patterns evolve with age differentially between sexes. Yet the underlying molecular mechanisms remain poorly understood. Methods:To characterize sex-specific aging-related changes in gene regulation in healthy thyroid and disease, we constructed individual-specific gene regulatory networks using a two-step approach. First, we estimated gene-gene co-expression networks for each individual using BONOBO. Second, we integrated these networks with sex-specific transcription factor (TF) motif data and protein-protein interaction priors using PANDA to infer individual-level TF-driven gene regulatory networks. Results:We found that within normal thyroid, regulatory patterns of cancer-related genes and biological pathways involved in cell proliferation, immune response, and metabolic processes vary by age in a nonlinear manner. However, the direction and rate of age-related changes differ between sexes. In females we detect two inflection points around ages 40 and 60, when the gene regulatory patterns in healthy thyroid show significant change for most pathways. This drastic change in gene regulatory networks is mostly driven by estrogen and androgen receptor TFs. To understand how aging-related changes in gene regulation drive risk of thyroid disorders, we investigated two diseases that have known sex- and age-specific differences: HT and ATC. We observed that in ATC, disease-relevant immune and metabolic pathways change with age in the same direction as they change in disease. In contrast, in HT, disease-related pathway targeting patterns are in the opposite direction of those in aging. Moreover, in age groups where HT is most diagnosed, TF-targeting patterns of disease-associated immune, metabolic and cell proliferation pathways in females were closer towards the disease state than in males, emphasizing the influence of sex-biased regulatory patterns in increasing thyroid disease risk in females. Conclusion:In thyroid tissue, genes related to immune response and metabolic processes, are regulated by TFs in an age- and sex-biased manner. These age- and sex-specific gene regulatory variations may contribute to the variation in risk of thyroid conditions with age and an overall higher risk of diseases in females compared to males, thus emphasizing the need for tailored screening and prevention strategies.
Rationale: In COPD, women tend to be diagnosed earlier, experience more severe symptoms with less smoking exposure, and suffer from worse overall symptoms. One of the X chromosomes in each cell of females undergoes inactivation through DNA methylation to balance gene dosage with males, but some genes can escape X chromosome inactivation (XCI). These escape genes vary among individuals and tissues and have been linked to certain diseases. Our study aims to examine the variability of XCI patterns in women and their impact on COPD. Methods: We analyzed peripheral blood DNA methylation data in 2,995 women from Phase 1 and 2,626 from Phase 2 of the COPDGene longitudinal cohort, and blood gene expression for the XIST gene from Phase 2. XCI status was defined by gene promoter methylation levels. To evaluate COPD differences in women, we used logistic regression to test for associations between methylation escape status with COPD affection status, lung function (FEV1, FEV1/FVC) and emphysema (LAA-950), accounting for age, smoking history, blood cell counts, income, and education. Results: Methylation levels of X chromosome probes showed intra-class correlations between Phases 1 and 2 significantly higher than autosomal probes, suggestive of stable inactivation patterns of one X chromosome allele through methylation. We identified 66 (7%) X chromosome genes that appear to escape inactivation in blood. These escape genes exhibited greater variability over the 5-year time frame, and were enriched for GO terms related to gene expression regulation and immune processes. The escape status of seven genes (ALG13, ARSD, FIRRE, IQSEC2, LHFPL1, MIR223, MORF4L2-AS1) was associated with COPD affection status, worse lung function, and increased emphysema (p<0.05). XCI escape status was associated with income and education, highlighting how social determinants of health may influence epigenetic drivers of COPD in women. The negative impact of escape status on lung function was more pronounced in women with lower income and education levels. Finally, we found that the expression of XIST, a key regulator of XCI, is correlated with methylation of multiple escape genes (median Pearson coefficient of 0.2, p<0.05) suggesting a mechanistic association between XIST expression and escape status. Conclusion: Our results show the association of higher rates of XCI escape of multiple genes with COPD onset and progression. The association of X chromosome patterns and COPD outcomes varies by education and income, suggesting that social determinants of health might influence COPD disparities among women through biological processes related to the X chromosome.
Chronic obstructive pulmonary disease (COPD) and idiopathic pulmonary fibrosis (IPF) are phenotypically divergent disorders arising from similar exposures (including cigarette smoke). Differences in DNA methylation may drive the exposed lung towards COPD vs. IPF. To characterize differential methylation in COPD and IPF lung tissue relative to controls, we conducted epigenome-wide association studies of COPD and IPF in lung tissue from the Lung Tissue Research Consortium (N=1029), adjusting for age, sex, smoke exposure, ancestry, estimated cell type composition, and plate. "Switch probes" were defined as CpGs differentially methylated in COPD vs. control and IPF vs. control in opposite directions. Gaussian graphical models were used to mine network properties of switch probes. Differential methylation of genes related to COPD/IPF in the literature was assessed. Switch probe methylation was compared with previously reported gene expression to identify multi-omic switches. We found 13,313 CpGs were associated with COPD and 43,359 with IPF (3,163 overlapping). We identified 1,091 switch CpGs enriched for endocytosis, glycosphingolipid biosynthesis, and pathways in cancer. 24 genes exhibited multi-omic switch behavior, many related to lipid metabolism ( ACSL1 ; FASN ; LPCAT1 ; MED27 ; NCOR2 ). LPCAT1 is of particular interest due to its role in maintaining phosphatidylcholine, the majority component of surfactant. Further related to surfactant, we observed strong divergent methylation and expression of ATP11A , which facilitates endocytosis of surfactant lipids. CONCLUSIONS Our findings suggest multi-omic switch-like regulation may underlie differential COPD/IPF etiology. Future investigation of LPCAT1 and ATP11A could provide new mechanistic understanding and therapeutic avenues.
Large-scale, open-access data sets such as the Genotype Tissue Expression Project (GTEx) and The Cancer Genome Atlas (TCGA) include multi-omic data on large numbers of samples along with extensive clinical and phenotypic information. These datasets provide a unique opportunity to discover correlations among clinical and genomic data features that can lead to testable hypotheses and new discoveries. SEAHORSE (http://seahorse.networkmedicine.org/) is a web-based database and search tool for exploratory data analysis in which we have pre-computed statistical associations between available data elements. An easy-to-use user interface allows users to explore significant associations using tabulated summary statistics, data visualizations, and functional enrichment analyses (using RNA-seq data) for identified sets of genes. We describe the motivation and construction of SEAHORSE and demonstrate its utility by documenting several surprising association patterns observed across multiple tissues in GTEx and multiple different cancer types in TCGA.
Chronic obstructive pulmonary disease (COPD) often develops at an earlier age in women than in men, with worse respiratory symptoms despite lower smoking exposure. However, most preventive and therapeutic strategies ignore biological sex differences in COPD. Our goal was to better understand sex-specific gene regulatory processes in lung tissue and the molecular basis for sex differences in COPD onset and severity. We analyzed lung tissue gene expression and DNA methylation data from 747 individuals in the Lung Tissue Research Consortium and 85 individuals in an independent dataset. We identified sex differences in COPD-associated gene regulation using gene regulatory networks. We used linear regression to test for sex-biased associations of methylation with lung function, emphysema, smoking, and age. Analyzing gene regulatory networks in the control group, we identified that genes involved in the extracellular matrix (ECM) have higher transcriptional factor targeting in female subjects than in male subjects. However, this pattern is reversed in COPD, with men showing stronger regulatory targeting of ECM-related genes than women. Smoking exposure, age, lung function, and emphysema were all associated with sex-specific differential methylation of ECM-related genes. We identified sex-based gene regulatory patterns of ECM-related genes associated with lung function and emphysema. Multiple factors, including epigenetics, smoking, aging, and cell heterogeneity, influence sex-specific gene regulation in COPD. Our findings underscore the importance of considering sex as a key factor in disease susceptibility and severity.
BACKGROUND:Technological advances in sequencing and computation have allowed deep exploration of the molecular basis of diseases. Biological networks have proven to be a valuable framework for analyzing omics data and modeling regulatory interactions between genes and proteins. Large collaborative projects, such as The Cancer Genome Atlas (TCGA), have provided a rich resource for building and validating new computational methods, resulting in a plethora of open-source software for downloading, preprocessing, and analyzing those data. However, for an end-to-end analysis of regulatory networks, a coherent and reusable workflow is essential to integrate all relevant packages into a robust pipeline. FINDINGS:We developed tcga-data-nf, a Nextflow workflow that allows users to reproducibly infer regulatory networks from the thousands of samples in TCGA using a single command. The workflow can be divided into 3 main steps: multiomic data, such as RNA sequencing and methylation, are (i) downloaded, (ii) preprocessed, and (iii) analyzed to infer regulatory network models with the Network Zoo. The workflow is powered by the NetworkDataCompanion R package, a standalone collection of functions for managing, mapping, and filtering TCGA data. Here, we demonstrate how the pipeline can be used to investigate the differences between colon cancer subtypes attributed to epigenetic mechanisms. Lastly, we provide a database of pregenerated networks for the 10 most common cancer types that can be readily accessed by the public. CONCLUSIONS:tcga-data-nf is a complete, yet flexible and extensible, framework that enables the reproducible inference and analysis of cancer regulatory networks, bridging a gap in the current universe of software tools for analyzing TCGA data.
The rising incidence of lung cancer among individuals without a history of smoking highlights the need to explore biological mechanisms underlying this phenomenon. Our study aims to identify gene regulatory mechanisms that drive lung cancer risk among never-smokers by analyzing how gene regulatory networks differ between individuals with and without lung cancer, depending on their smoking history. We used RNA-Seq data from the Lung Tissue Research Consortium (LTRC) collected via TOPMed, comprising non-cancerous lung tissue samples from 344 individuals with non-small cell carcinoma and 329 lung tissue samples from individuals without cancer. Across all, 18% reported no history of smoking. We ran differential gene expression analysis by voom, adjusting for age, sex, COPD status, and examining statistical interactions between smoking (never/former) and cancer status. We analyzed sample-specific transcription factor-gene regulatory networks generated by PANDA-LIONESS. We used linear regression to evaluate the association between gene targeting score (measured by gene indegree) and the interaction between smoking and cancer status. We ran gene set enrichment analysis with genes ranked by the corresponding smoking by cancer interaction coefficients. Differential expression analysis of lung tissues from individuals with and without cancer showed that genes overexpressed in cancer are enriched for canonical cancer pathways, such as p53, MAPK, and WNT signaling pathways (FDR<0.05). We found a significant interaction between cancer and smoking status, indicating these cancer-related pathways had higher enrichment in former-smokers compared to never smokers. Gene regulatory network analysis indicated that these differential expression patterns are possibly driven by increased transcriptional targeting of these cancer-related pathways. Among never-smokers, individuals with cancer showed higher targeting of metabolic pathways (e.g., fructose and mannose, glutathione, phenylalanine, histidine, arginine, and proline metabolism) compared to lung tissue from individuals without cancer (FDR<0.05), highlighting a possible role of metabolic pathways in tumor development, uniquely among never-smokers. A negative correlation between metabolic pathway targeting scores and age was observed exclusively in individuals with cancer without a history of smoking (Mean Pearson R=-0.24). Our findings reveal transcriptional targeting differences in lung cancer by smoking history. Among individuals without a history of smoking, the increased targeting of metabolic pathways in cancer, which is pronounced in younger compared to older individuals, may contribute to the higher risk of early-onset lung cancer among never-smokers, offering new avenues for understanding and addressing the disease in populations without a history of smoking. Camila Lopes-Ramos, Enakshi Saha, Jeong Yun, Craig Hersh, Edwin Silverman, Dawn DeMeo, John Quackenbush, Kimberly Glass. Lung cancer gene regulatory networks reflect smoking history and age-related metabolic pathway alterations [abstract]. In: Proceedings of the American Association for Cancer Research Annual Meeting 2025; Part 1 (Regular Abstracts); 2025 Apr 25-30; Chicago, IL. Philadelphia (PA): AACR; Cancer Res 2025;85(8_Suppl_1):Abstract nr 7478.
Lung adenocarcinoma (LUAD) exhibits differences between the sexes in incidence, prognosis, and therapy, suggesting underexplored molecular mechanisms. We conducted an integrative multi-omics analysis using the Clinical Proteomic Tumor Analysis Consortium (CPTAC) and The Cancer Genome Atlas (TCGA) datasets to contrast transcriptomes and proteomes between sexes. We used TIGER to analyze TCGA-LUAD expression data and found sex-biased activity of transcription factors (TFs); we used PTM-SEA with CPTAC-LUAD proteomics data and found sex-biased kinase activity. We combined these to construct a kinase-TF signaling network and discovered druggable pathways linked to cancer-related processes. We also found significant sex biases in clinically relevant TFs and kinases, including NR3C1, AR, and AURKA. Using the PRISM drug screening database, we identified potential sex-specific drugs, such as glucocorticoid receptor agonists and aurora kinase inhibitors. Our findings emphasize the importance of considering sex and using multi-omics network methods to discover personalized cancer therapies.
Aging is the primary risk factor for many cancer types, including lung adenocarcinoma (LUAD). To understand how aging-related alterations in the regulation of key cellular processes might affect LUAD risk and survival, we built individual-specific gene regulatory networks integrating gene expression, transcription factor protein-protein interaction, and sequence motif data, using PANDA/LIONESS algorithms, for non-cancerous lung samples from GTEx project and LUAD samples from TCGA. In healthy lung, pathways involved in cell proliferation and immune response were increasingly targeted with age; these aging-associated alterations were accelerated by smoking and resembled oncogenic shifts observed in LUAD. Aging-associated genes showed greater aging-biased targeting patterns in individuals with LUAD compared to healthier counterparts, a pattern suggestive of age acceleration. Using drug repurposing tool CLUEreg, we found small molecule drugs that may potentially alter the accelerating aging profiles we found. We defined a network-informed aging signature that was associated with survival in LUAD.