
Heteroplasmy is the mixture of mutant and wild-type mitochondrial DNA (mtDNA) within each of our cells. Heteroplasmy levels in cells, tissues, and organisms change over time, thus contributing to mitochondrial disease, aging, and evolution. Germline and pedigree studies first revealed heteroplasmy shifts between generations and have long offered a window into the dynamics of mtDNA inheritance through single oocytes. Single-cell technologies are now uncovering similar mechanisms that operate in somatic tissues throughout life. Stochastic processes (relaxed replication and vegetative segregation, enhanced through genetic bottlenecks) generate cell-to-cell variation, while selection mechanisms such as intercellular competition, mitophagy, and preferential replication allow or drive directional shifts. Single-cell sequencing, mtDNA imaging, and genetic screening, combined with mtDNA-editing technology and heteroplasmic model systems, have transformed our ability to dissect these processes, revealing heteroplasmy dynamics at molecular resolution. These approaches are uncovering quantifiable principles governing heteroplasmy across cell types and life stages, transforming our understanding from descriptive observations to predictive mechanistic models and novel therapeutic avenues.
Recent research has significantly advanced our understanding of how variation in the genome shapes human biology, yet the sex chromosomes remain among its least explored regions. Technical and conceptual challenges have historically limited their inclusion in genomic studies, despite their influence on gene regulation, development, and disease. Here, we review key aspects of sex chromosome biology, beginning with their evolutionary origins as ordinary autosomes. We highlight how variation in sex chromosome copy number, including typical differences between males and females as well as aneuploidies, provides insight into the roles of the X and Y chromosomes across the human life span, from early embryonic events, such as X chromosome inactivation, to later processes, including reproduction and aging. Finally, we outline emerging innovations that are enabling more comprehensive integration of the sex chromosomes into genomic research, laying the foundation for a more inclusive and mechanistic understanding of their contributions to human diversity.
Below the scale of long-range compartments and topologically associating domains (TADs) lie a diverse set of local chromatin organization features. Recent maps and single-molecule assays reveal modular sub-TAD units, nucleosome clutches, micro- and nanodomains, packing domains, microcompartments, and stripes arising from an interplay of loop extrusion, epigenetic affinity and condensates, and polymerase motion. We highlight how cohesin regulation and its crosstalk with transcription shape local topology, how microcompartments and affinity-driven hubs guide enhancer-promoter communication, and how nucleosome positioning and spacing set the energetic landscape on which these forces act. We also outline how sub-TAD architecture is related to dynamics and single-molecule heterogeneity, describe additional looping mechanisms and several case studies of mesoscale structure regulation, and discuss perspectives on how technological advances can build a more mechanistic understanding of mesoscale chromatin organization. Overall, we argue that the submegabase structure of chromatin, though complex, is an essential length scale for understanding transcriptional regulation.
Noncoding variants occur within noncoding genes as well as within the regulatory nontranslated regions of protein-coding genes. It is important to be aware that these variants have been increasingly implicated in developmental disease through a variety of mechanisms. However, they remain difficult to interpret clinically due to their unclear effect on transcript or protein abundance compared with coding variants. Here, we review methods to identify pathogenic noncoding variants in rare disease, which can present challenges due to the inaccessibility of disease-relevant tissue for many conditions. We explore experimental approaches such as high-throughput functional assays, omic data integration, and long-read sequencing. We also review computational methods for annotating and filtering variants, as well as machine learning methods for predicting variant effect and pathogenicity. We discuss the recent discovery of several developmental syndromes caused by noncoding variants and propose an integrated approach to identifying pathogenic noncoding variants within this patient cohort.
TDP-43 is an RNA-binding protein that regulates multiple aspects of RNA processing, and its mislocalization from the nucleus to the cytoplasm is a defining feature of amyotrophic lateral sclerosis (ALS). While both loss- and gain-of-function mechanisms contribute to disease, the discovery of cryptic splicing has shed light on the downstream consequences of TDP-43 nuclear clearance for neuronal health. Here, we highlight how loss of nuclear TDP-43 can drive a cascade of events that lead to the impairment of cellular proteostasis and result in a positive feedback loop that perpetuates neuronal dysfunction. This sustains the appearance of cryptic splicing events in genes that are involved in key pathways for the maintenance of axonal homeostasis and synaptic transmission. In contrast to their detrimental effects on neuronal health, cryptic splicing mechanisms may be harnessed to develop novel therapeutic strategies, unprecedentedly expanding the availability of therapeutic avenues for TDP-43 proteinopathies.
Numerous human diseases are caused by changes in gene expression levels. In addition, changing the expression levels of specific genes can lead to therapeutic benefits for several diseases. Nuclease-deficient gene-editing proteins fused to transcriptional modulators that target gene regulatory elements have emerged as powerful, programmable, and customizable systems to modify gene expression for therapeutic benefits. Several of these systems have already been used in the clinic, and many more are under development. Here, we review these emerging technologies and assess their therapeutic potential, their delivery, and related challenges in the clinic.
The advent of next-generation sequencing has expanded our understanding of the genotypic, pathobiological, and phenotypic spectrum of human disease, helping to inform more personalized patient care. Current clinical guidelines are prompting the generation of large volumes of clinical diagnostic genome sequencing data, but we remain unable to interpret variants found in the noncoding 98.5% of sequencing data. In this review, we discuss the known and emerging mechanisms by which noncoding variants cause human disease. Through the lens of immunity, we integrate insights from population genetics, evolutionary and functional genomics, and in silico strategies to propose a framework for identifying and characterizing potential disease-relevant noncoding variants with regulatory impact on gene expression. By tackling the assessment of this vast black box of missing genetic contributions to disease, we hope to improve diagnostic yields and clinical management for more unsolved patients.
Organoids have reshaped biomedical research by providing stem cell-derived model systems that capture key aspects of tissue organization, homeostasis, and disease. Their physiological relevance and adaptability have made them indispensable tools for studying development, regeneration, and tumor biology under controlled experimental conditions. This is particularly powerful in the human setting, where organoids offer an experimentally accessible alternative to in vivo studies that are ethically and practically unfeasible. Central to their success is the ability to apply diverse perturbation strategies, ranging from targeted genetic edits and pharmacological interventions to microenvironmental and biomechanical manipulations, that reveal the molecular logic of cellular and tissue function. In this review, we discuss the current landscape of organoid perturbation studies, highlighting methodological advances, representative applications, and what these efforts have taught us about cellular behavior in complex systems. By outlining methodological innovations and conceptual insights, we aim to establish a framework for using organoids not only as descriptive models but as predictive systems for probing and engineering human tissue behavior.
Richard Gibbs interviews James (Jim) Lupski about his training in New York and work in Houston to elucidate the role of complex genomic rearrangements in human genetic diseases. The challenges and excitement of developing human personalized genomics and the advantages of clinical translation of genome methods for both patients and researchers are discussed.
Artificial intelligence (AI) technologies have recently undergone transformative growth in capabilities. In human genetics, AI is rapidly advancing our ability to reveal the effects of genetic variation. This review explores recent progress and remaining challenges across the diverse applications of AI in genotype-to-phenotype mapping, from predicting the functional and clinical consequences of mutations, to identifying causal genes, to estimating disease risk. Particular emphasis is placed on the growing utility of general-purpose foundation models trained on massive genomic data, including DNA and protein language models, alongside areas where narrower machine-learning approaches still dominate. The review concludes with key considerations for future progress and impact.
Arthrogryposis multiplex congenita (AMC) is characterized by congenital joint contractures in two or more body areas resulting from reduced or absent fetal movements. AMC exhibits marked phenotypic and genetic heterogeneity, as it is a symptom rather than a disease and may be part of a large number of unrelated conditions. Despite advances in genomic approaches and the increasing number of newly identified genes, disease-gene identification was achieved in fewer than half of cases in several reports, including in a French cohort of 367 AMC patients. The most frequent cause of AMC in these reports was a primary involvement of skeletal muscle. In the French cohort, the most frequent mode of inheritance was autosomal recessive (68.3%); in autosomal dominant or X-linked form, a high proportion of de novo variants (24%) was observed, indicating that this mechanism plays a prominent part in this developmental condition. Accurate genetic diagnosis is critical for more tailored management of AMC and possibly other organ involvement.
Gene regulatory networks (GRNs) define the regulatory relationships among molecules such as transcription factors, chromatin remodelers, and target genes. GRNs play a critical role in diverse biological processes, including development, disease manifestation, and evolution. However, fully characterizing these networks across multiple cell types and states remains a significant challenge. Recent advances in single-cell omics have dramatically enhanced our ability to measure biological systems at unprecedented resolution. These technologies have opened new avenues for computational methods to infer GRNs, offering deeper insights into cell type-specific mechanisms, causality, and dynamic regulatory processes. This review summarizes the current state of GRN inference from single cell omic datasets, with a particular focus on dynamics and perturbations, and outlines key open challenges that must be addressed to advance the field.
In this article, I recount and reflect on the US Department of Energy's contributions to sequence the human genome as part of the Human Genome Project.
Understanding the drivers of human brain specialization, and how specialized properties are codified during development and evolution, seems to be within reach for the first time. Improved cell-based experimental models of the human brain have empowered the field to address some of the most fundamental questions about our brains, including mechanisms of neurodevelopment, the etiology of neurological disease, and the underpinnings of human-to-human variation in brain function and response. The emergence of scalable in vitro systems has enabled investigation of interindividual variation within large human cohorts in both normal development and disease processes, which is fundamental to developing effective and personalized treatments. This review explores recent advancements in organoid technology, highlighting future directions that employ interdisciplinary approaches to enhance the physiological relevance of these models. This work promises to bring us ever closer to understanding not only what makes a brain human but also how each of our brains is human in unique ways.
Newborn screening for phenylketonuria began in the United States in the early 1960s, and it expanded one disease at a time until the development of tandem mass spectrometry. This technology allowed for screening many conditions simultaneously, but its uneven adoption led to wide disparities. A collaboration between the American College of Medical Genetics and Genomics and the US Health Resources and Services Administration resulted in a recommended uniform screening panel. Newborn sequencing (NBSeq) identifies many monogenic disorders, although to date it cannot identify all cases identified by tandem mass spectrometry. NBSeq has the potential to reduce diagnostic odysseys and increase health equity, but it could also exacerbate disparities and cause psychosocial and clinical harms due to overdiagnosis, oversurveillance, and/or overtreatment. By expanding beyond previously established public health screening principles, NBSeq also challenges the mandatory nature of current screening. In this review, we examine the promise and perils of NBSeq.
2025 marks the twenty-fifth anniversary of the completion of a working draft of the 3-Gb human genome sequence and its availability in public databases to promote research into human health and disease for the benefit of all. The sequence was produced by the International Human Genome Sequencing Consortium, which comprised sequencing centers from six countries who together undertook the largest collaborative biological project to date. Under the leadership of Sir John Sulston, the United Kingdom played a significant role in the project through the Sanger Centre (now the Wellcome Sanger Institute), which was founded in 1992 with support from the Wellcome Trust, a charitable foundation funding medical research. The Sanger Centre contributed approximately one-third of the final human genome sequence generated by the Human Genome Project and, along with the European Bioinformatics Institute, developed Ensembl, one of the major databases providing free access to genomic data and annotation for biomedical research. As a result of a chance meeting, I came to work at the Sanger Centre (and later the Wellcome Sanger Institute) from 1992 to 2007, initially as a scientific administrator and later as the Human Genome Project manager and head of sequencing. Over that period, the Sanger Centre became one of the largest genome sequencing centers in the world and began its transition to become a world-leading center in genomics research to advance biology and health.
Collectively, various tandem and interspersed repetitive sequences make up approximately half the human genome, yet we have only begun to understand the potential functions of "junk" DNA. Here, we provide a brief overview of various types of repeats, but a full treatment of the repeat genome (repeatome) is beyond the scope of any review. Hence, we focus primarily on less established functions of a few major repeat classes, including pericentromeric satellites and abundant degenerate interspersed repeats, short interspersed nuclear elements (Alu), and long interspersed nuclear elements (L1). A theme developed throughout is how sequence organization in the human karyotype provides insights into potential functions within nuclear structure. For example, millions of small tandem major satellite repeats can form bodies that sequester nuclear factors, or the segmental organization of interspersed repeats may underpin the nuclear compartmentalization of heterochromatin and euchromatin. Decoding the vast repeatome is an exciting frontier being enabled by recent technological advancements. However, identifying the extent of meaningful information in repeats will likely require concepts that go well beyond impacts for individual genes, to new ways to identify and interpret broad patterns of genome-wide organization and nucleus-wide regulation.
Molecular profiling of DNA and RNA from pediatric cancers by next-generation sequencing has been demonstrated to improve diagnosis and prognosis and to identify somatic alterations indicating vulnerability to targeted therapies. Hence, much like in the treatment of adult cancers, molecular profiling is now routinely utilized in clinical workflows for pediatric cancers as a companion to conventional pathology diagnosis. Many variants of unknown significance identified through DNA profiling are being characterized by saturation genome editing, enabled by CRISPR editing technology and clever functional assays. Newer technologies and analytics are revealing additional structural complexity around cancer drivers and gene fusions in pediatric cancer DNA. Similarly, computational methods such as rare variant association studies and polygenic risk scoring are being used to identify novel cancer susceptibility. Together, these advances are expanding our understanding of pediatric cancer's complexity and fueling the development of emerging methods such as liquid biopsy-based monitoring.
Over the last decade, a set of very short (3-51 nt) and highly conserved microexons have been found to crucially influence a set of diverse protein functions and interactions. Advancements in RNA sequencing and analysis pipelines have revealed an enrichment for the alternative splicing of microexons in a subset of tissues and cell types, especially across the central nervous system. Microexons are thought to fine-tune important developmental processes such as synaptogenesis by preserving the protein's reading frame upon inclusion. Dysregulation of microexon splicing has been linked to several neurological conditions, including autism spectrum disorder and schizophrenia, as well as metabolic disorders like diabetes and various cancer types. This review discusses the expanding body of literature on the molecular and organismal consequences of microexon inclusion, emphasizing their evolutionary conservation, tissue specificity, and functional diversity. It also explores the potential for therapeutic interventions, including pharmacological modulation, on microexon splicing and splicing regulators like SRRM3 and SRRM4, offering perspectives on targeting diseases related to microexon misregulation. More research is needed to better understand similarities and differences between microexon functions across tissues, pathologies, and species.
Aneuploidy, characterized by the gain or loss of chromosomes, plays a critical role in both cancer and congenital aneuploidy syndromes. For any aneuploidy, we can distinguish between its general effects and its chromosome-specific effects. General effects refer to the common cellular stresses induced by aneuploidy, such as impaired proliferation, proteotoxic stress, and altered metabolism, which occur regardless of the specific chromosome involved and profoundly impact cellular and organismal functions. These generalized stresses often hinder cell fitness but can also, under certain conditions, contribute to cancer progression. In contrast, chromosome-specific effects arise from the altered dosage of particular genes on the gained or lost chromosome. These effects vary depending on the chromosome involved and can provide specific fitness effects in cancer cells or distinct developmental phenotypes in congenital aneuploidies like Down syndrome. Understanding the interplay between these two levels of effects is crucial for deciphering the outcomes of aneuploidy. This review synthesizes current knowledge and discusses future directions for unraveling the hallmarks of aneuploidy.