We sought to identify universal organizing principles behind phenotypic variation within cell types. Pareto optimality describes how trade-offs between optimal solutions account for variation, predicting that the boundary points of a data distribution reflect specialized functions. We hypothesized that transcriptomic variation was explained by Pareto optimality across all cell types. We then used the Tabula Sapiens Atlas of single-cell RNA sequencing across cell types and tissues in the human body to test this hypothesis and found that most cell types adhere to this theory. This enabled us to use this principled method to characterize the functions performed by each cell type. These phenotypes are derived from an unbiased approach and do not incorporate ideas from existing biological models or theories, and yet in many cases they recapitulate our understanding of the functions of major cell types. Ultimately, we conclude that multiobjective optimization broadly shapes the observed phenotypic variation within cell types. This finding enables us to write explicit representations of the low-dimensional manifolds on which transcriptomes of single cells reside. This can inform the design of the next generation of virtual cell language models, which aim to statistically learn low-dimensional transcriptomic manifolds.
A small number of people living with HIV (PLWH) develop broadly neutralizing antibodies (bNAbs) targeting multiple HIV strains. Although several viral and immune factors contribute to bNAb development, the genetic and environmental factors driving this response remain largely unknown. We performed combined cell-free DNA (cfDNA) and cell-free RNA (cfRNA) sequencing in 42 plasma samples from a longitudinal cohort of 14 PLWH (7 who develop bNAbs and 7 matched controls). This approach enabled us to non-invasively monitor the host transcriptome, viral genetic variation, and microbiome composition during HIV infection, and to identify molecular correlates of bNAb development. We find that development of bNAbs is associated with a transcriptomic signature of early immune activation characterized by elevated levels of MHC class I antigen presentation genes. This signature is independent of viral load or CD4 count and declines over time. In addition to host features, we recovered sufficient viral reads to reconstruct HIV consensus sequences, supporting the utility of cfRNA for viral genotyping. Finally, we also identified an enrichment of several microbial taxa in bNAb producers and increased levels of GB virus C (GBV-C), a non-pathogenic lymphotropic virus. Our findings suggest a distinct early immune activation profile in PLWH who develop bNAbs. More broadly, we show that combined cfDNA/cfRNA sequencing can reveal relationships between a protective immunogenic response to HIV infection, the host immune system, and microbiome, highlighting its potential for biomarker discovery in future vaccine and therapeutic studies.
How dendrites of different neurons segregate into discrete spatial domains during neural circuit assembly is poorly understood. Here, using the Drosophila olfactory system, we found that heterophilic interactions between two cell-surface proteins Teneurin-m (Ten-m) and Capricious (Caps) drive dendrite segregation. Ten-m and Caps are expressed in largely inverse patterns across projection neuron (PN) types when PNs are establishing their dendritic territories. Loss of Ten-m in Ten-m+ PNs causes their dendrites to invade Caps+ territories, whereas loss of Caps in Caps+ PNs causes dendrite invasion into Ten-m+ territories. Structure-guided mutations that abolish Ten-m-Caps binding disrupt dendrite segregation, whereas the same mutation on Ten-m preserves its homophilic attraction in a synaptic partner matching assay. These results support a model in which mutual repulsions between two inversely expressed cell-surface proteins drive dendrite segregation into discrete glomerular territories.
Cellular morphology is tightly linked to function, but how subcellular transcript localization contributes remains unclear. Using microglia, the brain's resident macrophages, as a model, we combined multiplexed error-robust fluorescence in situ hybridization with immunohistochemistry to map how morphology and subcellular mRNA localization interact with function in young and aged mouse brains. We show that mRNA spatial organization varies across microglial states and defines distinct localization patterns within their processes, revealing morphological heterogeneity within transcriptomically defined populations. Notably, we found a subpopulation of disease-associated-like microglia with a ramified morphology (that is, displaying numerous processes), challenging the conventional assumption between morphology and microglial states. Finally, we found that aging may reshape mRNA distributions and their co-localization networks, shifting microglial programs from intracellular signaling and regulation of phagocytosis toward migration and catabolic regulation. Our findings highlight the role of subcellular transcript organization in shaping microglial morphology and function, offering new avenues for studying and modulating microglial states in health, disease and aging.
Inflammatory responses occur within the complex spatial context of tissues and organs, and many questions remain about how tissue structure and cellular communication shape their spatiotemporal dynamics. Here, we use a multiplexed RNA in situ hybridization approach, together with analytical tools, to study inflammatory gene expression in the larval zebrafish tailfin in response to a bath of lipopolysaccharide. We use this model system to address whether spatial structure emerges in the tissue response even absent the spatial variation introduced by a pathogen. We find that epithelial cells in the tailfin express several proinflammatory genes, and that across these genes, the uniform stimulus triggers a spatially nonuniform response. We use a graph-based spectral decomposition method to analyze its structure, and find that it is consistent with a diffusion-consumption model of secondary signaling. Overall, long-wavelength modes dominate the signal, creating zones of activation which account for a majority of the variation in gene expression. Our results show that epithelial cells are important producers of proinflammatory effector molecules in this system, and that tissue induces spatial correlations even absent a structured input.
While coding regions in the genome have a direct interpretation in terms of protein products, significant fractions are non-coding and yet control essential biological functions. Unlike the genetic code, there is no "lookup table" that identifies where regulatory proteins, known as transcription factors (TFs), bind. Here, we extract these binding sites by distilling sequences of nucleotide letters into collective coordinates (hyperletters) representing the binding sites that are active under specific environmental conditions. Going beyond local information footprints between individual bases and expression levels, our information blueprint algorithm compresses the global information by optimising filters that simultaneously scan an entire promoter sequence. Inspired by renormalisation-group techniques, we identify TF binding sites as coarse-grained variables combining groups of correlated mutations with the highest collective impact on gene expression. We validate our approach on experimental data for E. coli and discover novel regulatory elements illustrating its deployment at scale across growth conditions.
Engulfment by macrophages is critical for waste clearance in the vertebrate brain. Understanding clearance mechanisms may open new therapeutic possibilities to counter brain aging and neurodegenerative diseases. However, few in vivo models exist to study engulfment in the brain and characterize this process during aging and across species. Here we present a genetic model for secretion of a fluorescent protein by neurons in the brain of the African turquoise killifish, the shortest-lived vertebrate that can be bred in captivity. We use this model to identify a population of brain macrophages in the killifish responsible for engulfment of material from the brain extracellular space. Intriguingly, many of these cells bear similarities to mammalian border-associated and monocyte-derived macrophages, rare subsets of macrophages in mouse and human brains noted for their engulfment capabilities. We also find that in our model, killifish brain macrophages decline in engulfment capacity with age. This work highlights how vertebrate brain macrophages, particularly those at brain border regions, can play a critical role in clearance and provides an opportunity to test interventions that can boost engulfment by these macrophages to promote brain resilience in old age and disease.
Developing a universal representation space for cells that encompasses the tremendous molecular diversity of cell types across species would be transformative for cell biology. Recent work using single-cell transcriptomic approaches to create molecular definitions of cell types in the form of cell atlases has provided the necessary data for such an endeavour1–3. Here we present the universal cell embedding (UCE) foundation model. UCE was trained on a large corpus of cell data using self-supervision, creating a unified biological latent space that can represent cells across diverse tissues and species. This latent space captures important biological variation despite the presence of experimental noise. UCE’s universality means that new cells can be embedded with no data labelling, model training or fine-tuning. We used UCE to create the Integrated Mega-scale Atlas, embedding 36 million cells, with more than 1,000 uniquely named cell types, from hundreds of experiments, dozens of tissues and eight species. We gain insights into the organization of cell types and tissues within the space. UCE’s embedding space exhibits emergent behaviour, identifying biology that it was never trained for, such as identifying developmental lineages and embedding data from species that were not included in the training set. Overall, by enabling a universal representation for every cell state and type, UCE is a valuable tool for analysis, annotation and hypothesis generation over single-cell data. The universal cell embedding foundation model learns to capture the organization and variation of cells by training on 36 million cells from hundreds of experiments, dozens of tissues and eight species.
Generative AI (Gen-AI) has shown a remarkable impact in several biological research areas, from protein folding and de novo design to pathogenic mutation prediction. However, it remains unclear whether these molecular-level successes can translate to cellular and multicellular insights relevant to fields ranging from immunology to cancer and neurodegeneration. This arises from the intricate nature of the molecular mechanisms that determine cellular and organismal behavior, the lack of sufficient training data, and the multicellular nature of most pathophysiologic phenotypes. Novel Gen-AI frameworks are likely needed to integrate prior biological knowledge, such as molecular interaction networks, as well as guiding principles focusing the community’s attention on solving biologically and translationally relevant problems. Drawing inspiration from Hilbert’s list of 23 mathematical problems that have focused the mathematical community’s attention for more than a century, we propose fifteen grand AI challenges to focus the biomedical community’s attention on critically relevant questions, most of which still lack effective predictive methodologies.
Mouse lemurs (Microcebus spp.) are an emerging primate model organism, but their genetics, cellular and molecular biology remain largely unexplored. In an accompanying paper1, we performed large-scale single-cell RNA sequencing of 27 organs from mouse lemurs. We identified more than 750 molecular cell types, characterized their transcriptomic profiles and provided insight into primate evolution of cell types. Here we use the generated atlas to characterize mouse lemur genes, physiology, disease and mutations. We uncover thousands of previously unidentified lemur genes and hundreds of thousands of new splice junctions including over 85,000 primate splice junctions missing in mice. We systematically explore the lemur immune system by comparing global expression profiles of key immune genes in health and disease, and by mapping immune cell development, trafficking and activation. We characterize primate-specific and lemur-specific physiology and disease, including molecular features of the immune program, lemur adipocytes and metastatic endometrial cancer that resembles the human malignancy. We present expression patterns of more than 400 primate genes missing in mice, many with similar expression patterns to humans and some implicated in human disease. Finally, we provide an experimental framework for reverse genetic analysis by identifying naturally occurring nonsense mutations in three primate immune genes missing in mice and by analysing their transcriptional phenotypes. This work establishes a foundation for molecular and genetic analyses of mouse lemurs and prioritizes primate genes, isoforms, physiology and disease for future study.
ABSTRACT Despite the increasing prevalence of neurodegenerative diseases, the molecular characterization of the brain remains challenging due to limited access to the tissue. Cerebrospinal fluid (CSF) contains a significant proportion of molecular contents originating from the brain, and characterizing these molecules has served as a surrogate to evaluate molecular dysregulation in the brain. Here we performed cell-free messenger RNA (cf-mRNA) RNA-sequencing on 52 human CSF samples, and further compared their transcriptomic profiles to matched plasma samples. In addition, we evaluated the molecular dysregulation of cf-mRNA in CSF between individuals with Alzheimer’s disease (AD) and non-cognitively impaired (NCI) controls. The molecular content of CSF cf-mRNA was distinct from plasma cf-mRNA, with a substantially higher number of brain-associated genes identified in CSF. We identified a large set of dysregulated gene transcripts in the CSF cf-mRNA population of individuals with AD, and these gene transcripts were used to establish a diagnostic classifier to discriminate AD from NCI subjects. Notably, the gene transcripts were enriched in biological processes closely associated with AD, such as brain development and synaptic signaling. We also discovered a subset of gene transcripts within AD subjects that exhibit a strong correlation between CSF and plasma cf-mRNA. This study not only reveals the novel cf-mRNA content of CSF but also highlights the potential of CSF cf-mRNA profiling as a tool to garner pathophysiological insights into AD.
Type 1 diabetes (T1D) is characterized by the autoimmune destruction of most insulin-producing β-cells, along with dysregulated glucagon secretion from pancreatic α-cells. We conducted an integrated analysis that combines electrophysiological and transcriptomic profiling, along with machine learning, of islet cells from T1D donors to investigate the mechanisms underlying their dysfunction. Surviving β-cells exhibit altered electrophysiological properties and transcriptomic signatures indicative of increased antigen presentation, metabolic reprogramming, and impaired protein translation. In α-cells, we observed hyper-responsiveness and increased exocytosis, which are associated with upregulated immune signaling, disrupted transcription factor localization and lysosome homeostasis, as well as dysregulation of mTORC1 complex signaling. Notably, key genetic risk signals for T1D were enriched in transcripts related to α-cell dysfunction, including MHC class I which were closely linked with α-cell dysfunction. Our data provide novel insights into the molecular underpinnings of islet cell dysfunction in T1D, highlighting pathways that may be leveraged to preserve residual β-cell function and modulate α-cell activity. These findings underscore the complex interplay between immune signaling, metabolic stress, and cellular identity in shaping islet cell phenotypes in T1D. ### Competing Interest Statement The authors have declared no competing interest.
High-resolution video capillaroscopy shows promise as a non-invasive method to directly observe blood cells as they flow through and interact with the microvasculature. When imaging microvasculature in the nailfold, the inherent shaking of the fingers requires physical stabilization to reduce image noise. However, any external force applied to the fingers will affect the blood flow rate. To address this, we designed an inflatable "finger-lock" for nailfold capillaroscopy, stabilizing the finger against a coverslip with constant pressure. Testing 72 participants, we demonstrated that increasing "finger-lock" pressure improves video stability and decreases the velocity of capillary blood cells. Hence, capillary blood cell velocity measurements require monitoring and controlling external pressure. Our work introduces a method to perturb local capillary blood flow and measure the microvascular response, enabling further studies investigating how person-specific factors (e.g., age and disease) impact vascular health.
Breakthroughs in biomedical research are transforming healthcare through smarter diagnostics, advanced drug delivery, regenerative therapies, and innovative clinical tools. This progress is driven by a shared vision to connect discovery with real-world impact through interdisciplinary collaboration, innovation, and purpose-driven research for a healthier future. On March 5–7, 2025, the Terasaki Institute for Biomedical Innovation (TIBI) hosted its third annual Terasaki Innovation Summit in Los Angeles, California, bringing together leading investigators, entrepreneurs, and translational experts to discuss advances in bioengineering. The summit featured keynote presentations, panel discussions, and sessions on commercialization and intellectual property. Excellence in translational science was celebrated by honoring four leaders with the Paul Terasaki Innovation Award, the Hisako Terasaki Young Investigator Award, and the Keith Terasaki Mid-Career Innovation Award. This Voices article captures key reflections from speakers and contributors on the state and future of biomedical innovation—and what it means to turn visionary science into transformative solutions.
All cells respond to changes in both their internal milieu and the environment around them through the regulation of their genes. Despite decades of effort, there remain huge gaps in our knowledge of both the function of many genes (the so-called y-ome) and how they adapt to changing environments via regulation. Here we describe a joint experimental and theoretical dissection of the regulation of a broad array of over 100 biologically interesting genes in E. coli across 39 diverse environments, permitting us to discover the binding sites and transcription factors that mediate regulatory control. Using a combination of mutagenesis, massively parallel reporter assays, mass spectrometry and tools from information theory and statistical physics, we go from complete ignorance of a promoter's environment-dependent regulatory architecture to predictive models of its behavior. As a proof of principle of the biological insights to be gained from such a study, we chose a combination of genes from the y-ome, toxin-antitoxin pairs, and genes hypothesized to be part of regulatory modules; in all cases, we discovered a host of new insights into their underlying regulatory landscape and resulting biological function.
The Tabula Sapiens is a reference human cell atlas containing single cell transcriptomic data from more than two dozen organs and tissues. Here we report Tabula Sapiens 2.0 which includes data from nine new donors, doubles the number of cells in Tabula Sapiens, and adds four new tissues. This new data includes four donors with multiple organs contributed, thus providing a unique data set in which genetic background, age, and epigenetic effects are controlled for. We analyzed the combined Tabula Sapiens data for expression of transcription factors, thereby providing putative cell type specificity for nearly every human transcription factor and as well as new insights into their regulatory roles. We analyzed the molecular phenotypes of senescent cells across the entire data set, providing new insight into both the universal attributes of senescence as well as those aspects of human senescence that are specific to particular organs and cell-types. Similarly, we analyzed sex-specific gene expression across all of the identified cell types and discovered which cell types and genes have the most distinct sex based gene expression profiles. Finally, to enable accessible analysis of the voluminous medical records of Tabula Sapiens donors, we created a web application powered by a large language model that allows users to ask general questions about the health history of the donors.
Mouse lemurs are the smallest and fastest reproducing primates, as well as one of the most abundant, and they are emerging as a model organism for primate biology, behaviour, health and conservation. Although much has been learnt about their ecology and phylogeny in Madagascar and their physiology, little is known about their cellular and molecular biology. Here we used droplet-based and plate-based single-cell RNA sequencing to create Tabula Microcebus, a transcriptomic atlas of 226,000 cells from 27 mouse lemur organs opportunistically obtained from four donors clinically and histologically characterized. Using computational cell clustering, integration and expert cell annotation, we define and biologically organize more than 750 lemur molecular cell types and their full gene expression profiles. This includes cognates of most classical human cell types, including stem and progenitor cells, and differentiating cells along the developmental trajectories of spermatogenesis, haematopoiesis and other adult tissues. We also describe dozens of previously unidentified or sparsely characterized cell types. We globally compare expression profiles to define the molecular relationships of cell types across the body, and explore primate cell and gene expression evolution by comparing lemur transcriptomes to those of human, mouse and macaque. This reveals cell-type-specific patterns of primate specialization and many cell types and genes for which the mouse lemur provides a better human model than mouse 1 . The atlas provides a cellular and molecular foundation for studying this model primate and establishes a general approach for characterizing other emerging model organisms.
B cells generate pathogen-specific antibodies and play an essential role in providing adaptive protection against infection. Antibody genes are modified in evolutionary processes acting on the B cell populations within an individual. These populations proliferate, differentiate, and migrate to long-term niches in the body. However, the dynamics of these processes in the human immune system are primarily inferred from mouse studies. We addressed this gap by sequencing the antibody repertoire and transcriptomes from single B cells in four immune-rich tissues from six individuals. We find that B cells descended from the same pre-B cell ("lineages") often colocalize within the same tissue, with the bone marrow harboring the largest excess of lineages without representation in other tissues. Within lineages, cells with different levels of somatic hypermutation are uniformly distributed among tissues and functional states. This suggests that the relative probabilities of localization and differentiation outcomes change negligibly during affinity maturation, and quantitatively agrees with a simple dynamical model of B cell differentiation. While lineages strongly colocalize, we find individual B cells nevertheless appear to make independent differentiation decisions. Proliferative antibody-secreting cells, however, deviate from these global patterns. These cells are often clonally expanded, their clones appear universally distributed among all sampled organs, and form lineages with an excess of cells of the same type. Collectively, our findings show the limits of peripheral blood monitoring of the immune repertoire, and provide a probabilistic model of the dynamics of antibody memory formation in humans.
Aging induces region-specific functional decline across the brain. The cerebellum, critical for motor coordination and cognitive function, undergoes significant structural and functional changes with age. The molecular mechanisms driving cerebellar aging, particularly the role of cerebellar glia, including microglia, remain poorly understood. Here, we used single-nuclei RNA sequencing (snRNA-seq), microglial bulk RNA-seq, and multiplexed error-robust fluorescence in situ hybridization (MERFISH) to characterize transcriptional changes associated with cellular aging in the mouse cerebellum. We discovered that microglia exhibited the most pronounced age-related changes of all cell types and that their transcriptional signatures pointed to enhanced neuroprotective immune activation and reduced lipid-droplet accumulation compared to hippocampal microglia. Furthermore, cerebellar microglia in aged mice, compared to young mice, were found in closer proximity to granule cells. This relationship was characterized using the newly defined neuron-associated microglia score, which captures proximity-dependent transcriptional changes and suggests a novel microglial responsiveness. These findings underscore the unique adaptations of the cerebellum during aging and its potential resilience to Alzheimers disease (AD) related pathology, providing crucial insight into region-specific mechanisms that may shape disease susceptibility. ### Competing Interest Statement The authors have declared no competing interest.