Muscle-derived stem/progenitor cells (MDSPCs) are an adult stem cell population with demonstrated regenerative and rejuvenative potential distinct from other muscle progenitor cells. However, their molecular identity and developmental status remain poorly defined. Using single-cell transcriptomics and proteomics, we comprehensively profiled murine MDSPCs across age groups. We show that MDSPCs exist along a transcriptional continuum of maturation-ranging from metabolically active, proliferative early-stage cells to late-stage, lineage-committed myogenic populations. While lacking canonical pluripotency markers, early-stage MDSPCs express gene programs associated with embryonic progenitor identity, suggesting a non-canonical, multipotent-like state. These features distinguish them from both satellite cells and committed myoblasts. Aging reshapes this continuum by reducing stemness-associated signatures while enhancing differentiation programs and oxidative stress. Our identification of distinct MDSPC states provide critical insights into mechanisms that underly tissue regeneration and aging. These findings offer a blueprint for development of future regenerative therapies to combat age-related functional decline.
The NIH Common Fund Data Ecosystem (CFDE) integrates data resources from 18 NIH Common Fund programs for discovery and integrative analysis. These programs generate valuable but heterogeneous datasets that can be difficult to discover, access, and reuse. CFDE aims to provide a collaborative, community-built infrastructure that links and enriches Common Fund programs. We describe the evolution, structure, and core technologies of CFDE, including practical approaches that support submission, integration, visualization, and public release of multimodal data. Training programs and workforce initiatives lower barriers to adoption. CFDE has devised solutions to critical issues facing cross-program initiatives, including data scale and heterogeneity, dataset integration, and long-term sustainability. We demonstrate the utility of linking Common Fund resources through integrative tools and cross-dataset queries to yield insights that would otherwise be infeasible. Collectively, CFDE shows that a standards-driven, federated approach enhances and unifies cross-disciplinary resources, fostering collaboration and data-driven discovery.
Triple-negative breast cancer (TNBC) remains the most aggressive breast cancer subtype, with limited treatment options and variable response to immune checkpoint inhibitors. While tumor-infiltrating lymphocytes have been extensively studied, the integration of system-level peripheral immune dynamics with mechanistic immune regulation underlying therapeutic response and resistance remain poorly defined. Here, we integrate systems-level immune state modeling with pathway-level mechanistic inference to analyze single-cell RNA sequencing of peripheral blood mononuclear cells from advanced TNBC patients treated with paclitaxel alone (chemotherapy) or in combination with anti-PD-L1 antibody atezolizumab (combination). This framework leverages treatment arm, longitudinal sampling, and clinical response to resolve coordinated immune programs across lymphoid and myeloid compartments. Using this approach, we identified distinct treatment- and response-specific immune states in pre- and post-treatment. Chemotherapy responders displayed pre-treatment adaptive immune priming, whereas combination therapy responders exhibited pre-existing effector T cell activity coupled with tumor tissue PD-L1 expression. In contrast, chemotherapy non-responders developed persistent post-treatment immune dysregulation in regulatory and terminal effector programs, while combination therapy non-responders demonstrated maladaptive remodeling of adaptive and innate lymphoid compartments, including dysfunctional NK and metabolically reprogrammed myeloid populations. Across both regimens, pathways involving protein translation, metabolic adaptation, and stress signaling emerged as critical modulators of response. These findings suggest that coordinated adaptive-innate immune dynamics underlie therapeutic efficacy, whereas systemic immune exhaustion and myeloid immunoregulation lead to resistance. Projection of these peripheral immune programs onto independent I-SPY2 showed concordant associations with tumor immune phenotypes and pathological complete response, supporting generalizability of the identified systemic immune states. Our study demonstrates the utility of an integrative systems-level approach for linking peripheral immune state organization with mechanistic insights, informing immune response and resistance in TNBC.
Biomarkers are essential tools for disease detection, risk assessment, therapeutic monitoring, and precision medicine. However, biomarker data are dispersed across heterogeneous resources, inconsistently reported in the literature, and rarely standardized for computational use. This fragmentation limits reproducibility, cross-study integration, and the discovery of novel biomarker and disease relationships. We developed BiomarkerKB, a knowledgebase designed to harmonize and integrate biomarker information under a standardized data model. The model follows the FDA-NIH BEST biomarker definition and captures both core fields (biomarker entity, condition, exposure agent) and contextual metadata (specimen, biomarker role, evidence, provenance). Biomarker data and related annotations were either curated from publications or collected from public resources (e.g., OpenTargets, GWAS Catalog, ClinVar, CIViC, OncoMX) and also contributed by Common Fund Data Coordinating Centers and Early Detection Research Network (EDRN). Standardization was achieved using ontologies and reference resources such as Disease Ontology, UBERON, UniProtKB, and HUGO Gene Nomenclature Committee (HGNC) gene symbols. BiomarkerKB data were ingested into a Neo4j-based knowledge graph and integrated with the Common Fund Data Ecosystem (CFDE) Knowledge Graph. The initial release of BiomarkerKB contains over 200,000 biomarker-disease associations spanning genes, proteins, metabolites, glycans, and chemical elements. The knowledge graph comprises more than 300,000 nodes and 1.2 million edges, enabling structured exploration of biomarker relationships within CFDE data as demonstrated through knowledge graph query-based use cases presented in this study. A publicly accessible web portal (https://biomarkerkb.org) provides keyword search, filtering, data downloads, and access to graph visualization to support both researchers and computational analyses. BiomarkerKB addresses a critical gap in biomarker informatics by providing a unified, FAIR, and computationally enabled framework for biomarker knowledge access. ### Competing Interest Statement The authors have declared no competing interest. NIH Common Fund, https://ror.org/001d55x84, U24OD038423, OT2OD032092
Recent widespread adoption of cerebral organoid protocols has led to many new studies assessing the human-specific features of neural diseases. However, not all organoid studies employ proper quality control, which limits the physiological relevance of their findings. Here, we discuss the stages of in vivo neocortex formation and how those stages are recapitulated in organoid protocols. We then present the first guide for real-time operator removal of maldeveloped organoids in shaking culture. Finally, we show preliminary work on an organoid imaging and mesofluidic control platform for automated quality control of organoid development. Taken together, this approach for assessing the morphological features of organoids will improve the rigor and reproducibility of organoid studies, increase effect sizes of physiologically relevant disease etiology, and pave the way for cortical organoid GMP in high-throughput. Clinical Relevance:This establishes high-throughput, visual brain organoid quality control for disease studies and preclinical testing.
As lipidomics approaches its 25th anniversary, we explore how lipid research has matured over the years while highlighting emerging innovations that are expanding our ability to study these diverse, life-critical biomolecules. In particular, we showcase the community-driven, open-access databases, software, and educational resources made freely available through the ELIXIR Core Data Resource LIPID MAPS for the benefit of both established and new researchers.
The NIH Common Fund Data Ecosystem (CFDE) program was established to facilitate data accessibility and interoperability across multiple Common Fund (CF) programs, promote collaborations and accelerate discoveries by combining diverse data types from different CF programs. The CFDE Data Resource Center (DRC) was tasked with developing two web-based portals: an Information Portal to serve information about the CFDE, and a Data Portal to host harmonized metadata and processed data contributed by participating CF Data Coordination Centers (DCCs) and other sources. To achieve these goals, the CFDE DRC developed the CFDE Workbench, a web-based platform that hosts processed data, metadata, tools, use cases, and analyses developed by the CFDE. The Cross-Cut Metadata Model (C2M2) and several other processed data are hosted by the CFDE Workbench, including set libraries (XMTs), Knowledge Graph (KG) assertions, and attribute tables. These processed data formats make information derived from CF programs more findable, accessible, interoperable, and reusable (FAIR), and artificial intelligence (AI)-ready for cross-DCC knowledge discovery. Besides serving data, metadata, and code assets, the CFDE Workbench has also developed several tools that utilize these resources to enable cross-CF-program knowledge discovery use cases. Overall, the CFDE Workbench is a platform that consolidates efforts toward making CF resources harmonized, FAIR, and AI-ready. The CFDE Workbench website is available from https://cfde.cloud.
Background: Breast cancer cell heterogeneity and cellular fate are governed by a variety of molecular mechanisms. The LINCS (Library of Integrated Network-based Cellular Signatures) consortium performed multi-omics experiments on normal breast epithelial cells, MCF10a, to deduce temporal mechanisms of regulation and cell state signatures contributing to pro-oncogenic phenotypes. Methods: Normal breast epithelial cells were treated with oncogenic ligands such as EGF, HGF, and OSM, and multi-modal measurements including Reverse phase protein Assay (RPPA), RNA-seq, ATAC-seq, and Cyclic immune-fluorescence (Cyclic-IF) were performed. Result: In our integrated analyses of the data to elucidate mechanisms, contextual functional networks were constructed by integrating protein signaling, transcription factor activity, and gene expression. Phenotypic changes in response to the ligands consisted of cell cycle modifications leading to oncogenic events such as loss of apoptosis and induction of EMT (Epithelial to Mesenchymal Transition). Activation of mTOR was observed with all ligands, which led to the activation of E2F1. Downstream transcriptomic regulation of E2F1 led to an increase in both oncogenic signaling and EMT. Additionally, under OSM treatment, activation of STAT3 facilitated the enhancement of EMT via transcriptomic regulation of JUN and FOS. These findings were further validated using the chromatin changes seen in ATAC-seq and protein localization as seen in Cyclic-IF assay. Conclusion: This analysis provides valuable insights into the mechanisms of transcriptional regulation during oncogenesis in normal breast cells treated with growth factors and can aid in the discovery of novel drug targets and treatments
Abstract Quadruple negative breast cancer (QNBC) is an aggressive subtype of triple negative breast cancer (TNBC) that lacks androgen receptor (AR) expression. QNBC is highly proliferative and is present in more than half of all TNBC cases. QNBCs lack all traditional breast cancer targets (ER, PR, HER2, AR) and chemotherapy is highly cytotoxic. Thus, there is an unmet need for actionable and more cytocompatible targets for QNBC. Kinesin family member C1 (KIFC1) is a highly cancer-cell specific microtubule-binding protein that impedes cancer cells with excess centrosomes (hallmark of cancer) from undergoing apoptosis. Preclinical in vitro and in vivo evidence suggests that inhibition of KIFC1 eradicates cancer cells, while sparing normal, healthy cells. Here, we investigate KIFC1 as a potential actionable biomarker for QNBC in vitro and in vivo TNBC models as well as via GeoMx spatial transcriptomic analysis. Knockdown of AR in TNBC cells led to upregulation of KIFC1, β-catenin/TCF4 expression, and TCF4-mediated transcription as well as increased cell proliferation and reduced apoptosis. Conversely, upregulation of AR signaling in TNBC cells produced the opposite effect. QNBC cells were more sensitive to KIFC1 inhibition than AR-positive TNBC cells. We validated in multiple independent publicly available breast cancer patient cohorts that AR gene expression negatively correlates with KIFC1, β-catenin, TCF4, and TCF4-target gene expression. In a large TNBC tissue set (n=250) from Louisiana State University Health Sciences Center, KIFC1 expression is upregulated in QNBC relative to AR-positive TNBC samples. Spatial transcriptomic analysis of this tissue set revealed that KIFC1, β-catenin, and TCF4 expression are upregulated in the epithelial or stromal compartments of AR-low relative to AR-high expressing TNBC tumors. Our findings suggest that KIFC1 may be upregulated in QNBCs via increased β-catenin/TCF4-mediated signaling, and that inhibition of KIFC1 may suppress QNBC cell proliferation. Collectively, our work provides preclinical evidence that KIFC1 may serve as a potential actionable QNBC biomarker that could improve clinical management of an aggressive subpopulation of TNBC patients. Citation Format: Benecia Jackson, Zahra Mesrizadeh, Robert Lou, Yate-Ching Yuan, Daniel Schmolze, Rania Bakkar, Jerneja Tomsic, Cristal Resto, Nancy Sanchez, Padmashree Rida, Shankar Subramaniam, Lucio Miele, Victoria Seewaldt, Nikita Jinna. Loss of androgen receptor expression in triple negative breast cancer upregulates kinesin family member C1 (KIFC1) via increased β-catenin/TCF4 signaling [abstract]. In: Proceedings of the American Association for Cancer Research Annual Meeting 2026; Part 1 (Regular Abstracts); 2026 Apr 17-22; San Diego, CA. Philadelphia (PA): AACR; Cancer Res 2026;86(7 Suppl):Abstract nr 5082.
Searching and learning from aggregated public metabolomics data spanning thousands of studies remained largely inaccessible. Here we present StructureMASST, a web-based application enabling scalable, structure-centric searches across public metabolomics repositories using molecule names or chemical representations. It queries a precomputed knowledgebase of 2.19 billion spectral matches and 420 million metadata links, supports modification-tolerant and mass-shift searches, and maps chemical structures across taxonomy, biological context and environmental conditions to accelerate discovery.
The biological heterogeneity of triple-negative breast cancer (TNBC) poses significant challenges for diagnosis, prognosis, and treatment. While prior TNBC subtype classifications exist, they are not widely used clinically. Here, we aimed to subtype TNBC based on transcriptomic profiles using cell type and state heterogeneity in tumor tissue from 250 pre-treatment women (127 African-American and 123 European-American). We identified three major subtypes and three distinct groups exhibiting unique cell-type composition and mechanisms: Subtype-1 immune signaling/T-cell response; Subtype-2 pro-fibrotic and immune desert; Subtype-3 fatty acid and nuclear receptor signaling. Subtype-1 showed potential responsiveness to immunotherapy, while Subtypes-2 and 3 suggested alternative therapeutic targets. In Subtype-3, which contained a patient group with high ESR1, (but not high ERα protein expression) we identified putative mutations in the gene that are unique to these patients. This framework provides a path toward personalized TNBC treatment and is accessible through a user-friendly RShiny application for clinical use.
We performed a systems vaccinology analysis to investigate immune responses in humans to an H5N1 influenza vaccine, with and without the AS03 adjuvant, to identify factors influencing antibody response magnitude and durability. Our findings revealed a platelet and adhesion-related blood transcriptional signature on day 7 that predicted the longevity of the antibody response, suggesting a potential role for platelets in modulating antibody response durability. As platelets originate from megakaryocytes, we explored the effect of thrombopoietin (TPO)-mediated megakaryocyte activation on antibody response longevity. We found that TPO administration enhanced the durability of vaccine-induced antibody responses. TPO-activated megakaryocytes also promoted survival of human bone-marrow plasma cells through integrin β1/β2-mediated cell–cell interactions, along with survival factors APRIL and the MIF–CD74 axis. Using machine learning, we developed a classifier based on this platelet-associated signature, which predicted antibody response longevity across six vaccines from seven independent trials, highlighting a conserved mechanism for vaccine durability. Pulendran and colleagues define a molecular signature that can be used to predict the durability of antibody responses to vaccination and reveal important insights into the mechanisms by which vaccines induce durable immunity.
Early-onset Alzheimer’s disease (EOAD) is a complex disease that occurs at an early age at onset (AAO) before 65 years, constituting 5-6% of all AD cases and remains poorly understood. Patient-derived induced pluripotent stem cells (iPSCs) have been used to model different forms of EOAD that display heterogeneous disease mechanisms. We examined iPSC-derived neurons from both familial EOAD harboring mutations in PSEN1 A79V , PSEN2 N141I , and APP V717I and non-familial EOAD patients at an early AAO. RNA-seq for familial and non-familial EOAD patients as well as ATAC-seq for familial EOAD patients were carried out to characterize the gene expression and chromatin accessibility changes, respectively. Differential expression and enrichment analysis, TF activity identification, and co-expression module detection were performed for familial EOAD RNA-seq. Clustering and surrogate neuron marker classification were performed for non-familial EOAD RNA-seq. Differential peak analysis, TF motif footprinting and peak functional enrichment were performed for familial EOAD ATAC-seq. Our approach allowed us to identify the correlation between gene expression and chromatin accessibility associated with key disease familial EOAD endotypes. We identified limitations with our non-familial EOAD neuron model to study sporadic AD, providing evidence that these neurons present variation of differentiation across patient clones, patient variability and an immature culture state. Common endotypes were identified across three familial EOAD mutations such as dedifferentiation of a mature neuron to a less differentiated quasi-neuron state and repression of mitochondrial function and metabolism. Integrative analysis allowed us to ascertain the master transcriptional regulators associated with these endotypes, including REST, ASCL1, and ZIC family members (activation), as well as NRF1 (repression). Our non-familial EOAD study showed a modest difference in expression profiling and a limited number of differentially expressed genes (DEGs) between diseased and control subjects. iPSC-derived neurons demonstrated that familial EOAD mutations share common regulatory changes within endotypes with varying severity, leading to reversion to a less-differentiated neuron state. Extending the usage of these neurons to non-familial EOAD may not serve as ideal to study sporadic AD. Overall, we have demonstrated that human neuron modeling can be applied to different forms of EOAD to understand the disease etiology better.
Despite being information rich, the vast majority of untargeted mass spectrometry data are underutilized; most analytes are not used for downstream interpretation or reanalysis after publication. The inability to dive into these rich raw mass spectrometry datasets is due to the limited flexibility and scalability of existing software tools. Here we introduce a new language, the Mass Spectrometry Query Language (MassQL), and an accompanying software ecosystem that addresses these issues by enabling the community to directly query mass spectrometry data with an expressive set of user-defined mass spectrometry patterns. Illustrated by real-world examples, MassQL provides a data-driven definition of chemical diversity by enabling the reanalysis of all public untargeted metabolomics data, empowering scientists across many disciplines to make new discoveries. MassQL has been widely implemented in multiple open-source and commercial mass spectrometry analysis tools, which enhances the ability, interoperability and reproducibility of mining of mass spectrometry data for the research community.
Recent applications of foundation models in biology have focused on pretraining using large-scale single-cell datasets comprising millions of cells, across diverse patho-physiological states. These models are then fine-tuned for downstream tasks such as cell-type classification. In this study, we evaluated the performance of three widely-used foundation models in biology—scGPT, SCMAMBA-2, and Geneformer—and a statistical baseline (Seurat v5) on cell-type classification under Gaussian noise perturbation. We used two curated datasets, referred to as Myeloid (13k cells) and hPancreas (15k cells). Surprisingly, we found that the baseline performance of the foundation models was inferior to that of the statistical model, even without any added perturbation. Although we note that model size can affect performance, Geneformer’s accuracy outperformed scGPT and SCMAMBA-2 by 5% on average across all datasets despite having 40% fewer trainable parameters. Nonetheless, the statistical baseline still outperformed Geneformer by 9% in accuracy. Based on these findings, we hypothesized that the conventional training paradigm used by foundation models for single-cell tasks consistently underperforms statistical models due to a lack of essential biological context. To investigate this, we evaluated whether performance degradation stems from early data embedding steps—such as binning or gene normalization during tokenization, and sampling bias. First, to better understand tokenization-related artifacts, we introduced controlled Gaussian noise to gene expression values before tokenization, amplifying downstream distortions introduced by the tokenization process (all models were trained for identical step durations with identical hyperparameters). On the Myeloid dataset, following the introduction of Gaussian noise perturbation to 20% of cells, both scGPT and SCMAMBA-2 saw an 11% decrease in accuracy while Geneformer saw an 8% decrease in accuracy. This difference in performance may be due to the difference in encoding methods used by both models. scGPT and SCMAMBA-2 use a bin-based tokenization strategy, in contrast to Geneformer’s rank-value encoding, which normalizes gene expression values using a predetermined encoding constant. Although binning captures general trends in count data, it fails to preserve relative expression at the gene level, resulting in significant information loss. Applying the three models to the Myeloid dataset also revealed that scGPT’s prediction distribution is biased toward overrepresented cell types in the training data, while underrepresenting rarer classes. CD14 cells are overpredicted by 16% (among most abundant cell types) by scGPT. Geneformer, however, maintains a more stable prediction distribution with a 6% (CD14) increase and outperforms scGPT and SCMAMBA-2 by 26% in macro F1 score (unweighted metric). Based on our findings we assert that the gap in contextual encoding in bin-based tokenization is what contributes to the less-nuanced learning. Recent research that integrated cellular-ontology during training showed improved performance to both scGPT and Geneformer. Our results underscore a fundamental issue, that foundation models lack critical biological context that would allow for them to make the nuanced inferences required for complex biological analyses. The compression of single-cell data from raw counts to embedding vectors can span several orders of magnitude, and lead to significant loss of information. As a result, methods must adapt to prioritize contextual integration during tokenization to ensure sufficient information for the model. ### Competing Interest Statement The authors have declared no competing interest.
Objective: High-throughput biological data, with its vast complexity and higher dimensions, continues to require innovative analytic methodologies for meaningful exploration. Most methods for reducing data dimensions overlook the shape and topology of data, even though these are vital components of the data structure and complexity. This study leverages topological data analysis (TDA) and shows, using breast cancer (BC) gene expression data as an illustrative example, the power of including the shape of data. Results: In addition to delineating the known subtypes of BC, TDA identifies a new subtype within luminal B cancer along with the features that define the subtype. The final outcome is shown via three-dimensional (3D) scatter plots which demonstrate how the underlying patterns that we identified through TDA map to 3D space. Conclusions: The new subtype, obtained unsupervised and validated by prior knowledge, demonstrates the power of embedding the topology and shape of data in the analyses.
Heterogeneity of breast cancer poses several challenges for detection and treatment. With next-generation sequencing, we can now map the transcriptional profile of each patient's breast tissue, which has the potential for identifying and characterizing cancer subtypes. However, the large dimensionality of this transcriptomic data and the heterogeneity between the molecular profiles of breast cancers poses a barrier to identifying minimal markers and mechanistic consequences. In this study, we develop an autoencoder to identify a reduced set of gene markers that characterize the four major breast cancer subtypes with the accuracy of 82.38%. The reduced feature space created by our model captures the functional characteristics of each breast cancer subtype highlighting mechanisms that are unique to each subtype as well as those that are shared. Our high prediction accuracy shows that our markers can be valuable for breast cancer subtype detection and have the potential to provide insights into mechanisms associated with each subtype.
Monge's disease, or Chronic Mountain Sickness (CMS), is a chronic high-altitude disorder characterized by hypoxia-induced excessive erythrocytosis (EE), elevating the risk of stroke and myocardial infarction. Using RNA-seq and ATAC-seq, we profiled iPSC-derived erythroid cells from CMS and non-CMS subjects under normoxia and hypoxia to identify statistically significant, disease-associated transcriptional and chromatin accessibility changes. RNA-seq revealed induction of inflammatory, stress, and erythropoiesis programs in CMS even under normoxia, including robust activation of JAK/STAT signaling, upregulation of heme metabolism and VEGF, and accelerated erythrocyte lineage commitment alongside repression of Notch and WNT/β-catenin. Hypoxia amplified this dysregulated state, and critically, activated NFκB-driven inflammatory signaling together with canonical HIF targets. ATAC-seq revealed pronounced hypoxia-induced changes, with increased accessibility within inflammatory and erythrocyte lineage genes occurring concomitantly with decreased accessibility within pluripotency and ectodermal lineage genes. Pharmacological NFκB inhibition in CMS cells significantly reduced EE ( p -value <0.0001), whereas NFκB activation in non-CMS cells was sufficient to drive EE ( p -value <0.01), confirming the causal role inferred by our multiomics analyses. Collectively, our multiomics and functional experiments substantiate a coordinated chromatin-transcription paradigm favoring an inflammatory axis that, through hypoxia-driven NFκB activation, accelerates stress-induced erythroid commitment and underlies EE in CMS.