
Genome editing technologies have advanced from nuclease-based reagents that generate programmed DNA double-strand breaks, which can cause deleterious effects, to next-generation reagents that perform controlled DNA modification through double-strand break-independent mechanisms, such as base editing and prime editing. Although these approaches enable precise small-scale sequence changes, methods for programmable insertion of large DNA cargos have been limited. The ability to write entire genes or large regions into the genome could transform the treatment of genetically heterogeneous disorders, for which numerous pathogenic variants underlie a common disease and mutation-specific editing strategies are impractical. Recent advances in computational genome mining have accelerated the discovery of naturally occurring enzymes with novel biochemical and functional properties, including recombinases and transposases capable of large-scale modifications. Moreover, directed evolution, rational engineering and expanded homologue discovery are enabling the repurposing and optimization of these systems for genome engineering. Here we review recent technology development efforts that harness diverse enzymes for kilobase-scale genome engineering, with a particular focus on CRISPR-associated transposase systems. The programmable insertion of large DNA sequences could overcome fundamental limitations of current genome-editing technologies. In this Review, Lampe, King and Sternberg describe recent advances in recombinases, retrotransposons and CRISPR-associated transposases, and discuss how these systems are enabling mutation-agnostic genome modification.
Despite sequencing advances, genomic resources remain uneven and rarely comparable by design. Turning today’s genomes into analysis-ready resources requires species-dense sampling anchored in natural history collections. Large-scale genome initiatives are transforming biodiversity genomics, but genomes remain taxonomically sparse and difficult to compare across species. Kapli and colleagues argue that natural history collections should play a central role in building species-dense genomic resources, creating the foundation needed for comparative and predictive biodiversity genomics.
Recent studies have shown that telomere attrition functions as a powerful and broadly effective tumour-suppressor mechanism that halts emerging cancers when telomeric DNA becomes depleted. This Review highlights new insights into how telomere attrition guards against cancer, including its role in limiting cell proliferation, the importance of telomere length at birth for lifelong cancer prevention, the mechanisms governing telomere length regulation and how telomerase activation enables malignant cells to bypass this barrier. In contrast to the popular perception that telomere shortening is harmful, the findings demonstrate that normal telomere attrition limits the risk of cancer and suggest that counteracting telomere attrition in healthy individuals (for example, by activating telomerase) could enable tumour outgrowth. Thus, such interventions are best reserved for patients with diseases caused by excessively short telomeres. The new data also further underscore the promise of telomerase inhibition as a broadly effective cancer therapy.
Comparative genetics aims to resolve the genetic and molecular architecture of complex traits and the evolutionary constraints governing biology. Recent improvements in genome assembly, coupled with population-scale multi-omics, have generated high-resolution maps of genetic and functional variation across species. These advances have been particularly transformative for farmed animals, which have a rich reservoir of genetic diversity and provide an opportunity for systematic, full-lifespan, pan-tissue functional annotation not always possible for humans. Given the physiological similarities and shared selective environments between humans and farmed animals, a unified comparative framework that integrates human and animal omics data will enable a bidirectional flow of information and bridge the gap between statistical association and causal mechanism. This framework will be essential to accelerate the parallel development of precision agriculture and biomedicine.
In human population studies, proteogenomics integrates genomic and proteomic data to uncover how genetic variation shapes protein expression and function. A central focus of proteogenomics is the identification of protein quantitative trait loci (pQTL), which link genetic variants to protein abundance and offer critical insights into the molecular basis of disease. Recent years have seen pQTL studies scale rapidly, driven by advances in high-throughput proteomic platforms, the expansion of large biobank datasets and increasingly powerful cross-cohort meta-analyses, thereby greatly extending the depth, breadth and translational potential of proteogenomics. This Review highlights recent progress in proteogenomics with a focus on pQTL mapping and its downstream integration with complementary molecular data to provide insights into human disease. We also address current challenges and propose future directions to harness proteogenomics for the development of personalized therapies and improvement of health outcomes.
In this Journal Club, Pamela Yeh recalls a landmark study by Bollenbach et al. that used genetic perturbations in Escherichia coli to uncover the mechanism underlying a classic suppressive drug interaction.
The widespread adoption of electronic health records (EHRs), which capture patient-specific longitudinal information on diagnoses, laboratory tests, clinical procedures, and outcomes, has created unprecedented opportunities to study diseases at scale. Integrating EHR data with genomic information offers novel ways to understand disease heterogeneity, identify biomarkers and therapeutic targets, and predict disease risk to improve clinical decision-making at scale. Recent methodological advances in machine learning (ML) and artificial intelligence (AI) can handle data with high dimensionality, high levels of noise, and irregular temporality better than traditional statistical approaches. Progress in EHR-linked biobank development, data standardization pipelines, and architectures for modelling and implementation have accelerated the advancement of the field and warrant an assessment of current capabilities and limitations. This Review highlights AI and ML frameworks for integrating genomic, multi-omics, and EHR data, and discusses how these approaches are reshaping genomics research as well as clinical practice.
Christian Landry reflects on a 2002 study by Brem, Yvert and colleagues, which combined yeast genetics and genome-wide expression profiling to reveal the genetic architecture underlying gene expression and transformed our understanding of regulatory variation.
The ageing process is often portrayed as a steady linear decline, as reflected by the methodological frameworks used to study it. Although linear models successfully capture relevant age-related mechanisms, they may fail to identify key transition states occurring during the lifespan. Indeed, ageing entails a series of transitions - from early development, through adolescence, to older age - that are each characterized by distinct biological remodelling events. The study of these underlying patterns, through an appropriate framework, is essential because interventions may only be effective during specific windows when the system transitions to a new state. This Perspective highlights the evidence for non-linearity in ageing, overviews current analytical approaches to capture non-linear processes and then discusses important considerations for ageing research going forward.
In this Tools of the Trade article, Khalid Al-Zahrani introduces CRISPR-KOALA, a new CRISPR-based platform that enables simultaneous loss- and gain-of-function screening in vivo, helping to pinpoint cancer driver genes hidden within large chromosome-scale copy number alterations.
Emerging computational approaches that measure the functional impact of HLA variation offer an opportunity to represent aspects of antigen-presentation diversity as quantitative genomic measures rather than as solely as categorical genetic features. The era of population genomics now provides the scale needed to benchmark these approaches and determine whether such measures can serve as proxies for immune diversity. Computational metrics such as HLA evolutionary divergence offer new ways to quantify functional variation in antigen presentation. In this Comment, the authors discuss how population-scale genomics can benchmark these approaches and assess their potential as quantitative measures of human immune diversity.
Genetic variation influences human physiology across biological scales from molecules to cells, tissues, organs and the whole organism. Unravelling how variants and their genetic effects propagate across these levels, through molecular interactions, cellular programmes and tissue architectures, to shape phenotypes remains a central challenge in human genetics. Resolving this challenge requires deciphering the genetic architecture of each biological layer and developing systems-level analyses that aim to integrate across scales. Network-based and computational approaches, including artificial intelligence, offer opportunities to move beyond statistical associations towards a context-aware, mechanistic understanding of the genetics underlying human traits and disease, although integration across layers remains limited. Here we review recent advances in mapping genetic effects across biological scales, from intracellular networks that capture molecular interactions, through single-cell and spatial omics approaches that define cellular and tissue contexts, to population-scale imaging genomics that links genetic variation to organ-level and organismal phenotypes. We discuss emerging strategies and remaining challenges for integrating these layers into mechanistic models of genotype-phenotype relationships.
In this Journal Club, Tatsuya Nobori recalls a landmark paper by Melotto et al., who showed that stomata are active components of plant innate immunity rather than passive entry points for bacterial pathogens.
Reference genomes serve as a coordinate system and are central to almost all analyses in genomics. However, linear reference genomes are based on a single individual or a small number of individuals and do not represent genetic diversity. Recent advances in de novo genome assembly, powered by long-read sequencing technologies, now enable the sequence reconstruction of many genomes to reference quality. These pangenomes integrate sequences from multiple individuals into graph-based or multi-haplotype representations, capturing genetic variation beyond a single linear reference. The widespread adoption of such pangenome references, which encode a diverse set of haplotypes, thus removes biases and enables the discovery of variants relative to all included haplotype backgrounds. The emergence of corresponding computational tools for analysing structural variants and complex genetic loci opens up opportunities in genome-wide association studies and rare-disease genetics. Here we review these opportunities, as well as challenges concerning pangenomes that need to be addressed by the research community.
Cohesin is a protein complex that shapes 3D genome organization through two distinct mechanisms. First, cohesin tethers replicated chromatids from DNA replication until mitosis. This process, known as sister chromatid cohesion, ensures accurate chromosome segregation and enables high-fidelity DNA repair through homologous recombination between the sister chromatids. Second, cohesin organizes the genome during interphase by dynamically extruding chromatin loops, structures that have key roles in gene regulation. Recent work has shown that, in addition to the well-established repair functions of sister chromatid cohesion, cohesin-mediated chromatin looping is closely linked to the repair of DNA double-strand breaks - one of the most toxic DNA lesions. In this Review, we discuss the central roles of cohesin in maintaining genome stability, with emphasis on the cellular response to DNA double-strand breaks. We review how dynamic loop structures facilitate signalling of repair events and promote long-range chromatin motions that underpin the repair process. Overall, its dual mode of action - cohesion and loop extrusion - positions cohesin as a central regulator of chromatin architecture and genome maintenance.
Imaging-derived phenotypes (IDPs) developed from medical imaging data, such as magnetic resonance imaging, computed tomography and X-ray scans, are traits that provide quantitative information on anatomical and functional properties of organs and tissues. IDPs are powerful tools for identifying biomarkers and studying disease mechanisms. When coupled with genetic data, IDPs can be analysed as heritable phenotypes using modern gene mapping methods to uncover genotype-phenotype relationships. The field of imaging genomics is rapidly maturing, with the emergence of high-quality imaging datasets collected in biobank-scale cohorts and sophisticated computational methods for extracting IDPs from imaging data, including tools that leverage machine learning. Here we review common imaging modalities and analytical approaches for developing IDPs, discuss biological insights gleaned from the large-scale genetic analysis of imaging traits and highlight emerging areas and remaining challenges that must be overcome to realize the full potential of IDPs for genetic analysis.
In this Tools of the Trade article, Yujie Chen describes CHARM, a single-cell four-omics sequencing method that profiles genome conformation, chromatin accessibility, histone modifications and gene expression.
Many plants sense and respond to temperature signals over long seasonal timescales to robustly time the transition to reproductive development - a process that will have to cope with the rapidly changing climate. Decades of genetic and molecular analyses have identified many genes influencing flowering in response to environmental signals. Among these, the regulation of one gene, FLOWERING LOCUS C (FLC), is central to requiring and responding to the prolonged cold of winter, aligning flowering with spring. Recent advances in our mechanistic understanding of FLC regulation suggest that, in contrast to a classic temperature signalling cascade, different aspects of natural fluctuating temperatures are monitored in parallel over various timescales and integrated through chromatin memory mechanisms, providing a robust way to accommodate climate variation. Dissecting the complexity and the source of adaptive variation in these flowering mechanisms helps in understanding the consequences of climate change with far-reaching implications for crop improvement.
In this Tools of the Trade article, Dmitry Kretov and Daniel Cifuentes describe RBPscan (RNA-binding protein specificity and contextual analysis via nucleotide editing), a method that measures the strength of in vivo interactions between RNA binding proteins and their RNA targets.