Breast cancer remains a significant global health challenge due to its complexity, which arises from multiple genetic and epigenetic mutations that originate in normal breast tissue. Traditional machine learning models often fall short in addressing the intricate gene interactions that complicate drug design and treatment strategies. In contrast, our study introduces GEMDiff, a novel computational workflow leveraging a diffusion model to bridge the gene expression states between normal and tumor conditions. GEMDiff augments RNAseq data and simulates perturbation transformations between normal and tumor gene states, enhancing biomarker identification. GEMDiff can handle large-scale gene expression data without succumbing to the scalability and stability issues that plague other generative models. By avoiding the need for task-specific hyper-parameter tuning and specific loss functions, GEMDiff can be generalized across various tasks, making it a robust tool for gene expression analysis. The model's ability to augment RNA-seq data and simulate gene perturbations provides a valuable tool for researchers. This capability can be used to generate synthetic data for training other machine learning models, thereby addressing the issue of limited biological data and enhancing the performance of predictive models. The effectiveness of GEMDiff is demonstrated through a case study using breast mRNA gene expression data, identifying 307 core genes involved in the transition from a breast tumor to a normal gene expression state. GEMDiff is open source and available at https://github.com/xai990/GEMDiff.git under the MIT license.
Genes involved in centrosome function, microtubule dynamics, and mitotic regulation are critical for normal cell division. In cancer, dysregulation of these processes can lead to chromosomal instability and tumor progression. The genes CEP72, HAUS4, TUBGCP4, HAUS2, PLK1, and OFD1 play essential roles in cell cycle regulation, mitosis, and microtubule dynamics, particularly in spindle assembly and centrosome function. These genes are involved in interconnected cellular processes related to mitotic progression, and their altered expression may be linked to cancer progression and survival outcomes. This study analyzes the expression profiles for these six genes across thirteen cancer types, aiming to identify associations between dysregulation of these genes and cancer progression, as well as their impact on survival outcomes. We analyzed the expression of these six genes and the tumor suppressor genes APC and PTEN using RNA sequencing data from TCGA cancerous tissues and normal control tissues from GTEx for the thirteen cancer types. Differential gene expression was assessed using the DESeq2 R package, which models RNA sequencing data using a negative binomial distribution. Our analysis identified significant changes in gene expression across cancers, including bladder (BLCA), colon (COAD, READ), esophagus (ESCA), kidney (KICH, KIRC, KIRP), liver (LIHC), prostate (PRAD), stomach (STAD), and thyroid (THCA). Elevated CEP72 expression was associated with lower progression-free survival (PFS). We observed consistent downregulation of CEP72, HAUS2, and PLK1 across several cancers, with significantly adjusted p-values (padj < 1E-5). OFD1, HAUS4, and PTEN showed tissue-specific expression patterns, with upregulation in cancers such as prostate (padj = 1.1E-11) and thyroid (padj = 8.5E-5). We further identified cancer-specific expression profiles for these six genes involved in cell division. These findings suggest complex regulatory mechanisms influencing tumor progression and metastasis. Further analysis of these mitotic and centrosomal proteins revealed significant cancer-specific correlations between gene expression levels of CEP72, HAUS4, TUBGCP4, HAUS2, PLK1, OFD1 and patient outcomes, including PFS, disease-specific survival, and overall survival. This study underscores the interconnectedness of centrosomal and mitotic genes, suggesting that downregulation of key proteins like CEP72, HAUS2, and PLK1 disrupts essential cell migration, proliferation, and adhesion pathways. Conversely, upregulation of HAUS4 may enhance tumor suppressor functions and immune modulation. Analyzing these genes within their broader biological networks provides valuable insights into their roles in cancer progression, offering potential as cancer-specific biomarkers or therapeutic targets. Christopher L. Farrell, Amy Turner, Alex Feltus, Victoria Cipollino. Centrosomal and mitotic gene expression profiles across cancer types: Implications for tumor progression and survival outcomes. [abstract]. In: Proceedings of the American Association for Cancer Research Annual Meeting 2025; Part 1 (Regular Abstracts); 2025 Apr 25-30; Chicago, IL. Philadelphia (PA): AACR; Cancer Res 2025;85(8_Suppl_1):Abstract nr 3343.
Homo sapiens and Neanderthals underwent hybridization during the Middle/Upper Paleolithic age, culminating in retention of small amounts of Neanderthal-derived DNA in the modern human genome. In the current study, we address the potential roles Neanderthal single nucleotide polymorphisms (SNP) may be playing in autism susceptibility in samples of black non-Hispanic, white Hispanic, and white non-Hispanic people using data from the Simons Foundation Powering Autism Research (SPARK), Genotype-Tissue Expression (GTEx), and 1000 Genomes (1000G) databases. We have discovered that rare variants are significantly enriched in autistic probands compared to race-matched controls. In addition, we have identified 25 rare and common SNPs that are significantly enriched in autism on different ethnic backgrounds, some of which show significant clinical associations. We have also identified other SNPs that share more specific genotype-phenotype correlations but which are not necessarily enriched in autism and yet may nevertheless play roles in comorbid phenotype expression (e.g., intellectual disability, epilepsy, and language regression). These results strongly suggest Neanderthal-derived DNA is playing a significant role in autism susceptibility across major populations in the United States.
Legumes establish a symbiotic relationship with nitrogen-fixing rhizobia by developing nodules. Nodules are modified lateral roots that undergo changes in their cellular development in response to bacteria, but the transcriptional reprogramming that occurs in these root cells remains largely uncharacterized. Here, we describe the cell-type-specific transcriptome response of Medicago truncatula roots to rhizobia during early nodule development in the wild-type genotype Jemalong A17, complemented with a hypernodulating mutant (sunn-4) to expand the cell population responding to infection and subsequent biological inferences. The analysis identifies epidermal root hair and stele sub-cell types associated with a symbiotic response to infection and regulation of nodule proliferation. Trajectory inference shows cortex-derived cell lineages differentiating to form the nodule primordia and, posteriorly, its meristem, while modulating the regulation of phytohormone-related genes. Gene regulatory analysis of the cell transcriptomes identifies new regulators of nodulation, including STYLISH 4, for which the function is validated.
Admixture refers to the mixing of genetic ancestry from different populations. Admixture is important for genomic medicine because it can affect how an individual responds to certain medications, how they metabolize drugs, and susceptibility to certain diseases. For example, some genetic variants associated with drug metabolism and response may be more common in certain populations, and individuals with admixed ancestry may have a different frequency of these variants than individuals from the ancestral populations. Understanding the patterns of admixture in a population can also help researchers identify new genetic variants associated with diseases or traits and develop more personalized and targeted treatments. In this study, we compared and classified the known and self-reported genetic backgrounds from 1000 Genomes Project and admixed samples from GTEx projects using supervised, unsupervised and statistical classification methodologies. We developed a novel tool called Admix-AI that uses a one-dimensional convolutional neural network to understand and classify admixed genetic backgrounds using 213 DNA-marker based genetic background labels. Admix-AI can be used to discover admixed proportions in samples and ultimately aid personalized genomic medicine by identifying specific biomarker systems. We compared Admix-AI to the existing admixture categorization software and found our tool to be computationally faster with 2× speedup and streamlined usage. Admix-AI is available as open-source code under GPL version 3.0 license at https://github.com/rpauly/Admix-AI .
Summary Large-scale and whole-cell modeling has multiple challenges, including scalable model building and module communication bottlenecks (e.g. between metabolism, gene expression, signaling, etc). We previously developed an open-source, scalable format for a large-scale mechanistic model of proliferation and death signaling dynamics, but communication bottlenecks between gene expression and protein biochemistry modules remained. Here, we developed two solutions to communication bottlenecks that speed up simulation by ~4-fold for hybrid stochastic-deterministic simulations and by over 100-fold for fully deterministic simulations. Availability and Implementation Source code is freely available at https://github.com/birtwistlelab/SPARCED/releases/tag/v1.1.0 implemented in python, and supported on Linux, Windows, and MacOS (via Docker). Contact Marc Birtwistle mbirtwi@clemson.edu Supplementary information N/A
AbstractNodule number regulation in legumes is controlled by a feedback loop that integrates nutrient and rhizobia symbiont status signals to regulate nodule development. Signals from the roots are perceived by shoot receptors, including a CLV1-like receptor-like kinase known as SUNN in the annual medicMedicago truncatula. In the absence of functional SUNN, the autoregulation feedback loop is disrupted, resulting in hypernodulation. To elucidate early autoregulation mechanisms disrupted inSUNNmutants, we searched for genes with altered expression in the loss-of-functionsunn-4mutant and included therdn1-2autoregulation mutant for comparison. We identified constitutively altered expression of small groups of genes insunn-4roots, including higher levels of transcription factorNF-YA2, and insunn-4shoots. All genes with verified roles in nodulation that were induced in wild type roots during the establishment of nodules were also induced insunn-4, including, surprisingly, autoregulation genesTML2andTML1. Among all genes with a differential response to rhizobia in wild type roots, only an isoflavone-7-O-methyltransferase gene (Medtr7g014510) was found to be unresponsive insunn-4. In shoot tissues of wild type, eight rhizobia-responsive genes were identified, including a MYB family transcription factor gene (Medtr3111880) which remained at a baseline level insunn-4; three genes were found to be induced by rhizobia in shoots ofsunn-4but not wild type. We also cataloged the temporal induction profiles of many small secreted peptide (MtSSP) genes in nodulating root tissues, encompassing members of twenty-four peptide families, including the CLE and IRON MAN families. The discovery that expression ofTMLgenes in roots, a key factor in inhibiting nodulation in response to autoregulation signals, is also triggered insunn-4in the section of roots analyzed suggests that the mechanism of TML regulation inM. truncatulamay be more complex than published models.
Nodule number regulation in legumes is controlled by a feedback loop that integrates nutrient and rhizobia symbiont status signals to regulate nodule development. Signals from the roots are perceived by shoot receptors, including a CLV1-like receptor-like kinase known as SUNN in Medicago truncatula. In the absence of functional SUNN, the autoregulation feedback loop is disrupted, resulting in hypernodulation. To elucidate early autoregulation mechanisms disrupted in SUNN mutants, we searched for genes with altered expression in the loss-of-function sunn-4 mutant and included the rdn1-2 autoregulation mutant for comparison. We identified constitutively altered expression of small groups of genes in sunn-4 roots and in sunn-4 shoots. All genes with verified roles in nodulation that were induced in wild-type roots during the establishment of nodules were also induced in sunn-4, including autoregulation genes TML2 and TML1. Only an isoflavone-7-O-methyltransferase gene was induced in response to rhizobia in wild-type roots but not induced in sunn-4. In shoot tissues of wild-type, eight rhizobia-responsive genes were identified, including a MYB family transcription factor gene that remained at a baseline level in sunn-4; three genes were induced by rhizobia in shoots of sunn-4 but not wild-type. We cataloged the temporal induction profiles of many small secreted peptide (MtSSP) genes in nodulating root tissues, encompassing members of twenty-four peptide families, including the CLE and IRON MAN families. The discovery that expression of TML2 in roots, a key factor in inhibiting nodulation in response to autoregulation signals, is also triggered in sunn-4 in the section of roots analyzed, suggests that the mechanism of TML regulation of nodulation in M. truncatula may be more complex than published models.
We report a public resource for examining the spatiotemporal RNA expression of 54,893 M. truncatula genes during the first 72 hours of response to rhizobial inoculation. Using a methodology that allows synchronous inoculation and growth of over 100 plants in a single media container, we harvested the same segment of each root responding to rhizobia in the initial inoculation over a time course, collected individual tissues from these segments with laser capture microdissection, and created and sequenced RNA libraries generated from these tissues. We demonstrate the utility of the resource by examining the expression patterns of a set of genes induced very early in nodule signaling, as well as two gene families (CLE peptides and nodule specific PLAT-domain proteins) and show that despite similar whole root expression patterns, there are tissue differences in expression between the genes. Using a rhizobial response data set generated from transcriptomics on intact root segments, we also examined differential temporal expression patterns and determined that, after nodule tissue, the epidermis and cortical cells contained the most temporally patterned genes. We circumscribed gene lists for each time and tissue examined and developed an expression pattern visualization tool. Finally, we explored transcriptomic differences between the inner cortical cells that become nodules and those that do not, confirming that the expression of ACC synthases distinguishes inner cortical cells that become nodules and provide and describe potential downstream genes involved in early nodule cell division.
Domain science applications in fields such as Genomics and High-Energy Particle Physics use geographically distributed data federations for publishing and accessing datasets. Data is typically replicated among data federation nodes to improve efficiency and fault tolerance. While replication strategies are well documented in distributed database instances (e.g., Apache Cassandra), replication among distributed data storage nodes can be ad-hoc. Replication over wide area networks can also require global coordination (or global shared state) which is not ideal when multiple organizations are involved. In this paper, we introduce GNSGA, which stands for Greedy Non-dominated Sorting Genetic Algorithm II. It is an optimization algorithm that combines greedy and non-dominated sorting genetic algorithms to solve multi-objective optimization problems. The “greedy” aspect of the algorithm refers to the use of a greedy strategy in the selection of nodes, while the “Non-dominated Sorting Genetic Algorithm II (NSGA-II)” is a fast non-dominated multi-objective optimization algorithm with an elite retention strategy. Replication decisions in GNSGA are based on the local properties and resource availability of the data storage nodes. By incorporating Greedy and NSGA-II algorithms, GNSGA optimizes multiple conflicting objectives to satisfy replica placement constraints such as cost, time, and storage capacity. We compared GNSGA with popular replica placement strategies, such as closest node replication, shortest transfer time, and a Particle Swarm Optimization (PSO)-based replication algorithm. We performed simulations and an actual deployment on the NSF's FABRIC testbed for evaluation. The results demonstrate that GNSGA consistently selects nodes to reduce replication time by 5.8%-15.4% while satisfying replication constraints (i.e., cost, time, and storage). We also show that GNSGA is beneficial for replicating large files over wide area networks.
Abstract The remarkable flexibility and adaptability of generative adversarial networks (GANs) have led to the proliferation of its models in bioinformatics research. Proteomic and transcriptomic profiles have been shown to be promising methods for discovering and identifying disease biomarkers. However, those analyses were performed by trained human examiners making the process tedious, time consuming, and hard to standardize. With the development of GANs, it is now possible to reduce computational costs and human time for bioinformatics analysis to produce effective biomarkers. Moreover, GANs help address the lack of phenotypic state transitional gene expression data as well as avoid protected human data constraints by generating RNA sequencing (RNA‐seq) data from random vectors. The purpose of this review is to summarize the use of GAN approaches and techniques to augment RNA‐seq expression data and identify clinically useful biomarkers. We compare different studies that use different types of GAN models to examine the biomarkers. Also, we identify research gaps and challenges that apply GANs to bio‐informatics. Finally, we propose potential directions for future research.
We report a public resource for examining the spatiotemporal RNA expression of 54,893 Medicago truncatula genes during the first 72 h of response to rhizobial inoculation. Using a methodology that allows synchronous inoculation and growth of more than 100 plants in a single media container, we harvested the same segment of each root responding to rhizobia in the initial inoculation over a time course, collected individual tissues from these segments with laser capture microdissection, and created and sequenced RNA libraries generated from these tissues. We demonstrate the utility of the resource by examining the expression patterns of a set of genes induced very early in nodule signaling, as well as two gene families (CLE peptides and nodule specific PLAT-domain proteins) and show that despite similar whole-root expression patterns, there are tissue differences in expression between the genes. Using a rhizobial response dataset generated from transcriptomics on intact root segments, we also examined differential temporal expression patterns and determined that, after nodule tissue, the epidermis and cortical cells contained the most temporally patterned genes. We circumscribed gene lists for each time and tissue examined and developed an expression pattern visualization tool. Finally, we explored transcriptomic differences between the inner cortical cells that become nodules and those that do not, confirming that the expression of 1-aminocyclopropane-1-carboxylate synthases distinguishes inner cortical cells that become nodules and provide and describe potential downstream genes involved in early nodule cell division. [Formula: see text] Copyright © 2023 The Author(s). This is an open access article distributed under the CC BY-NC-ND 4.0 International license .
Background Lung cancer is the leading cause of cancer death in both men and women. The most common lung cancer subtype is non-small cell lung carcinoma (NSCLC) comprising about 85% of all cases. NSCLC can be further divided into three subtypes: adenocarcinoma (LUAD), squamous cell carcinoma (LUSC), and large cell lung carcinoma. Specific genetic mutations and epigenetic aberrations play an important role in the developmental transition to a specific tumor subtype. The elucidation of normal lung versus lung tumor gene expression patterns and regulatory targets yields biomarker systems that discriminate lung phenotypes (i.e., biomarkers) and provide a foundation for the discovery of normal and aberrant gene regulatory mechanisms. Results We built condition-specific gene co-expression networks (csGCNs) for normal lung, LUAD, and LUSC conditions. Then, we integrated normal lung tissue-specific gene regulatory networks (tsGRNs) to elucidate control-target biomarker systems for normal and cancerous lung tissue. We characterized co-expressed gene edges, possibly under common regulatory control, for relevance in lung cancer. Conclusions Our approach demonstrates the ability to elucidate csGCN:tsGRN merged biomarker systems based on gene expression correlation and regulation. The biomarker systems we describe can be used to classify and further describe lung specimens. Our approach is generalizable and can be used to discover and interpret complex gene expression patterns for any condition or species.
In response to colonization by rhizobia bacteria, legumes are able to form nitrogen-fixing nodules in their roots, allowing the plants to grow efficiently in nitrogen-depleted environments. Legumes utilize a complex, long-distance signaling pathway to regulate nodulation that involves signals in both roots and shoots. We measured the transcriptional response to treatment with rhizobia in both the shoots and roots of Medicago truncatula over a 72-h time course. To detect temporal shifts in gene expression, we developed GeneShift, a novel computational statistics and machine learning workflow that addresses the time series replicate the averaging issue for detecting gene expression pattern shifts under different conditions. We identified both known and novel genes that are regulated dynamically in both tissues during early nodulation including leginsulin, defensins, root transporters, nodulin-related, and circadian clock genes. We validated over 70% of the expression patterns that GeneShift discovered using an independent M. truncatula RNA-Seq study. GeneShift facilitated the discovery of condition-specific temporally differentially expressed genes in the symbiotic nodulation biological system. In principle, GeneShift should work for time-series gene expression profiling studies from other systems.
BACKGROUND:Quantification of gene expression from RNA-seq data is a prerequisite for transcriptome analysis such as differential gene expression analysis and gene co-expression network construction. Individual RNA-seq experiments are larger and combining multiple experiments from sequence repositories can result in datasets with thousands of samples. Processing hundreds to thousands of RNA-seq data can result in challenges related to data management, access to sufficient computational resources, navigation of high-performance computing (HPC) systems, installation of required software dependencies, and reproducibility. Processing of larger and deeper RNA-seq experiments will become more common as sequencing technology matures.RESULTS:GEMmaker, is a nf-core compliant, Nextflow workflow, that quantifies gene expression from small to massive RNA-seq datasets. GEMmaker ensures results are highly reproducible through the use of versioned containerized software that can be executed on a single workstation, institutional compute cluster, Kubernetes platform or the cloud. GEMmaker supports popular alignment and quantification tools providing results in raw and normalized formats. GEMmaker is unique in that it can scale to process thousands of local or remote stored samples without exceeding available data storage.CONCLUSIONS:Workflows that quantify gene expression are not new, and many already address issues of portability, reusability, and scale in terms of access to CPUs. GEMmaker provides these benefits and adds the ability to scale despite low data storage infrastructure. This allows users to process hundreds to thousands of RNA-seq samples even when data storage resources are limited. GEMmaker is freely available and fully documented with step-by-step setup and execution instructions.
Today's big data science communities manage their data publication and replication at the application layer. These communities utilize myriad mechanisms to publish, discover, and retrieve datasets - the result is an ecosystem of either centralized, or otherwise a collection of ad-hoc data repositories. Publishing datasets to centralized repositories can be process-intensive, and those repositories do not accept all datasets. The ad-hoc repositories are difficult to find and utilize due to differences in data names, metadata standards, and access methods. To address the problem of scientific data publication and storage, we have designed Hydra, a secure, distributed, and decentralized data repository made of a loose federation of storage servers (nodes) provided by user communities. Hydra runs over Named Data Networking (NDN) and utilizes the State Vector Sync (SVS) protocol that lets individual nodes maintain a "global view" of the system. Hydra provides a scalable and resilient data retrieval service, with data distribution scalability achieved via NDN's built-in data anycast and in-network caching and resiliency against individual server failures through automated failure detection and maintaining a specific degree of replication. Hydra utilizes "Favor", a locally calculated numerical value to decide which nodes will replicate a file. Finally, Hydra utilizes data-centric security for data publication and node authentication. Hydra uses a Network Operation Center (NOC) to bootstrap trust in Hydra nodes and data publishers. The NOC distributes user and node certificates and performs the proof-of-possession challenges. This technical report serves as the reference for Hydra. It outlines the design decisions, the rationale behind them, the functional modules, and the protocol specifications.
AbstractThe mechanisms that coordinate cellular gene expression are highly complex and intricately interconnected. Thus, it is necessary to move beyond a fully reductionist approach to understanding genetic information flow and begin focusing on the networked connections between genes that organize cellular function. Continued advancements in computational hardware, coupled with the development of gene correlation network algorithms, provide the capacity to study networked interactions between genes rather than their isolated functions. For example, gene coexpression networks are used to construct gene relationship networks using linear metrics such as Spearman or Pearson correlation. Recently, there have been tools designed to deepen these analyses by differentiating between intrinsic vs extrinsic noise within gene expression values, identifying different modules based on tissue phenotype, and capturing potential nonlinear relationships. In this report, we introduce an algorithm with a novel application of image-based segmentation modalities utilizing blob detection techniques applied for detecting bigenic edges in a gene expression matrix. We applied this algorithm called EdgeCrafting to a bulk RNA-sequencing gene expression matrix comprised of a healthy kidney and cancerous kidney data. We then compared EdgeCrafting against 4 other RNA expression analysis techniques: Weighted Gene Correlation Network Analysis, Knowledge Independent Network Construction, NetExtractor, and Differential gene expression analysis.
Autism Spectrum Disorder (ASD) is a complex neurodevelopmental disorder characterized by challenges in social communication as well as repetitive or restrictive behaviors. Many genetic associations with ASD have been identified, but most associations occur in a fraction of the ASD population. Here, we searched for eQTL-associated DNA variants with significantly different allele distributions between ASD-affected and control. Thirty significant DNA variants associated with 174 tissue-specific eQTLs from ASD individuals in the SPARK project were identified. Several significant variants fell within brain-specific regulatory regions or had been associated with a significant change in gene expression in the brain. These eQTLs are a new class of biomarkers that could control the myriad of brain and non-brain phenotypic traits seen in ASD-affected individuals.
Background Thyroid cancer (THCA) is the most common endocrine malignancy and incidence is increasing. There is an urgent need to better understand the molecular differences between THCA tumors at different pathologic stages so appropriate diagnostic, prognostic, and treatment strategies can be applied. Transcriptome State Perturbation Generator (TSPG) is a tool created to identify the changes in gene expression necessary to transform the transcriptional state of a source sample to mimic that of a target. Methods We used TSPG to perturb the bulk RNA expression data from various THCA tumor samples at progressive stages towards the transcriptional pattern of normal thyroid tissue. The perturbations produced were analyzed to determine if there are consistently up- or down-regulated genes or functions in certain stages of tumors. Results Some genes of particular interest were investigated further in previous research. SLC6A15 was found to be down-regulated in all stage 1-3 samples. This gene has previously been identified as a tumor suppressor. The up-regulation of PLA2G12B in all samples was notable because the protein encoded by this gene belongs to the PLA2 superfamily, which is involved in metabolism, a major function of the thyroid gland. REN was up-regulated in all stage 3 and 4 samples. The enzyme renin encoded by this gene, has a role in the renin-angiotensin system; this system regulates angiogenesis and may have a role in cancer development and progression. This is supported by the consistent up-regulation of REN only in later stage tumor samples. Functional enrichment analysis showed that olfactory receptor activities and similar terms were enriched for the up-regulated genes which supports previous research concluding that abundance and stimulation of olfactory receptors is linked to cancer. Conclusions TSPG can be a useful tool in exploring large gene expression datasets and extracting the meaningful differences between distinct classes of data. We identified genes that were characteristically perturbed in certain sample types, including only late-stage THCA tumors. Additionally, we provided evidence for potential transcriptional signatures of each stage of thyroid cancer. These are potentially relevant targets for future investigation into THCA tumorigenesis.