Centromeres play essential roles in chromosome segregation and genome stability, yet they remain among the least characterized regions of the human genome. Despite advances in long-read sequencing and complete genome assembly, the extreme repetitiveness and structural complexity of these regions still challenge population-scale analysis, obscuring their mutational dynamics. The Human Pangenome Reference Consortium has now accurately assembled over 6,000 centromeres, providing an opportunity to catalog global centromere variation. However, centromeric regions have been systematically excluded from pangenome alignments due to the technical challenge of aligning their highly repetitive tandem arrays and extreme structural variability. Here we introduce Centrolign, a graph-based multiple sequence alignment tool that combines a uniqueness-driven objective function with partial-order partial-order alignment to accurately align alpha satellite higher-order repeats. By prioritizing rare matches within tandem arrays and leveraging extended centromere-spanning haplotypes formed by suppressed recombination, Centrolign produces progressive multiple sequence alignments that preserve ancestral repeat organization. Applied across human centromeres, these alignments reveal the phylogenetic structure of similar satellite array haplotypes and enable precise estimation of variation rates, structural variant frequencies, and spatial patterns of mutation within satellite arrays. Integrating Centrolign graphs with repeat annotation tools and pangenome mapping algorithms allows accurate variant calling and genotyping from long reads without prior assembly. Moreover, we show that centromere haplotypes can be accurately subtyped with k-mers alone. Together, these advances establish a robust framework for incorporating centromeres into broader pangenomes, and population genomics in general, advancing our understanding of human genome evolution and diversity.
A pangenome reference overcomes the inherent limitation of any individual reference genome by integrating the variation present in a population. We present the Human Pangenome Reference Consortium's (HPRC) Release 2 (HPRC2), an openly available, second phase pangenome that is an approximately fivefold expansion in genome number over HPRC Release 1 (HPRC1) and measurable improvement in genome completeness, contiguity, and accuracy. Selecting samples with a principled algorithm prioritising common variant coverage, HPRC2 contributes 460 haplotypes that together capture over 99% of common variation observed in the All of Us Research Program v8 cohort. Combining high-coverage long and ultra-long reads with modern assemblers and polishers, we produce thousands of telomere-to-telomere (T2T) chromosomes, and relative to HPRC1 halve the number of structurally unreliable regions as well as individual base errors per haplotype. We complement the assemblies with whole genome multiple alignments and gene annotations, and derive formal pangenome coordinate systems for addressing off-reference variation, demonstrating that individual human genomes contain more than one hundred thousand variants not succinctly described with respect to existing reference genomes. We also present the first matched long-read backed pantranscriptome and panepigenome at this scale, provide continuous local-ancestry estimates spanning every genome, and outline a host of new tools and applications that leverage the pangenome resource for improved genomics analysis.
Balanced mating type polymorphisms offer a distinct window into the forces shaping sexual reproduction strategies. Multiple hermaphroditic genera in Juglandaceae, including walnuts (Juglans) and hickories (Carya), show a 1:1 genetic dimorphism for male versus female flowering order (heterodichogamy). We map two distinct Mendelian inheritance mechanisms to ancient (>37 million years old) genus-wide structural DNA polymorphisms. The dominant haplotype for female-first flowering in Juglans contains tandem repeats of the 3 untranslated region of a gene putatively involved in trehalose-6-phosphate metabolism and is associated with increased cis gene expression in developing male flowers, possibly mediated by small RNAs. The Carya locus contains similar to 20 syntenic genes and shows molecular signatures of sex chromosome-like evolution. Inheritance mechanisms for heterodichogamy are deeply conserved, yet may occasionally turn over, as in sex determination.
The maintenance of stable mating type polymorphisms is a classic example of balancing selection, underlying the nearly ubiquitous 50/50 sex ratio in species with separate sexes. One lesser known but intriguing example of a balanced mating polymorphism in angiosperms is heterodichogamy - polymorphism for opposing directions of dichogamy (temporal separation of male and female function in hermaphrodites) within a flowering season. This mating system is common throughout Juglandaceae, the family that includes globally important and iconic nut and timber crops - walnuts (Juglans), as well as pecan and other hickories (Carya). In both genera, heterodichogamy is controlled by a single dominant allele. We fine-map the locus in each genus, and find two ancient (>50 Mya) structural variants involving different genes that both segregate as genus-wide trans-species polymorphisms. The Juglans locus maps to a ca. 20 kb structural variant adjacent to a probable trehalose phosphate phosphatase (TPPD-1), homologs of which regulate floral development in model systems. TPPD-1 is differentially expressed between morphs in developing male flowers, with increased allele-specific expression of the dominant haplotype copy. Across species, the dominant haplotype contains a tandem array of duplicated sequence motifs, part of which is an inverted copy of the TPPD-1 3' UTR. These repeats generate various distinct small RNAs matching sequences within the 3' UTR and further downstream. In contrast to the single-gene Juglans locus, the Carya heterodichogamy locus maps to a ca. 200-450 kb cluster of tightly linked polymorphisms across 20 genes, some of which have known roles in flowering and are differentially expressed between morphs in developing flowers. The dominant haplotype in pecan, which is nearly always heterozygous and appears to rarely recombine, shows markedly reduced genetic diversity and is over twice as long as its recessive counterpart due to accumulation of various types of transposable elements. We did not detect either genetic system in other heterodichogamous genera within Juglandaceae, suggesting that additional genetic systems for heterodichogamy may yet remain undiscovered.
Sugar pine, Pinus lambertiana Douglas, is a keystone species of montane forests from Baja California to southern Oregon. Like other North American white pines, populations of sugar pine have been greatly reduced by the disease white pine blister rust (WPBR) caused by a fungal pathogen, Cronartium ribicola, that was introduced into North America early in the twentieth century. Major gene resistance to WPBR segregating in natural populations has been documented in sugar pine. Indeed, the dominant resistance gene in this species, Cr1, was genetically mapped, although not precisely. Genomic single nucleotide polymorphisms (SNPs) placed in a large scaffold were reported to be associated with the allele for this major gene resistance (Cr1(R)). Forest restoration efforts often include sugar pine seed derived from the rare resistant individuals (typically Cr1(R)/Cr1(r)) identified through an expensive 2-year phenotypic testing program. To validate and geographically characterize the variation in this association and investigate its potential to expedite genetic improvement in forest restoration, we developed a simple PCR-based, diploid genotyping of DNA from needle tissue. By applying this to range-wide samples of susceptible and resistant (Cr1(R)) trees, we show that the SNPs exhibit a strong, though not complete, association with Cr1(R). Paralleling earlier studies of the geographic distribution of Cr1(R) and the inferred demographic history of sugar pine, the resistance-associated SNPs are marginally more common in southern populations, as is the frequency of Cr1(R). Although the strength of the association of the SNPs with Cr1(R) and thus, their predictive value, also varies with geography, the potential value of this new tool in quickly and efficiently identifying candidate WPBR-resistant seed trees is clear.
Existing human genome assemblies have almost entirely excluded repetitive sequences within and near centromeres, limiting our understanding of their organization, evolution, and functions, which include facilitating proper chromosome segregation. Now, a complete, telomere-to-telomere human genome assembly (T2T-CHM13) has enabled us to comprehensively characterize pericentromeric and centromeric repeats, which constitute 6.2% of the genome (189.9 megabases). Detailed maps of these regions revealed multimegabase structural rearrangements, including in active centromeric repeat arrays. Analysis of centromere-associated sequences uncovered a strong relationship between the position of the centromere and the evolution of the surrounding DNA through layered repeat expansions. Furthermore, comparisons of chromosome X centromeres across a diverse panel of individuals illuminated high degrees of structural, epigenetic, and sequence variation in these complex and rapidly evolving regions.
Understanding the genomic and environmental basis of cold adaptation is key to understand how plants survive and adapt to different environmental conditions across their natural range. Univariate and multivariate genome-wide association (GWAS) and genotype-environment association (GEA) analyses were used to test associations among genome-wide SNPs obtained from whole-genome resequencing, measures of growth, phenology, emergence, cold hardiness, and range-wide environmental variation in coastal Douglas-fir (Pseudotsuga menziesii). Results suggest a complex genomic architecture of cold adaptation, in which traits are either highly polygenic or controlled by both large and small effect genes. Newly discovered associations for cold adaptation in Douglas-fir included 130 genes involved in many important biological functions such as primary and secondary metabolism, growth and reproductive development, transcription regulation, stress and signaling, and DNA processes. These genes were related to growth, phenology and cold hardiness and strongly depend on variation in environmental variables such degree days below 0c, precipitation, elevation and distance from the coast. This study is a step forward in our understanding of the complex interconnection between environment and genomics and their role in cold-associated trait variation in boreal tree species, providing a baseline for the species’ predictions under climate change.
ABSTRACTJuglans (walnuts), the most speciose genus in the walnut family (Juglandaceae) represents most of the family’s commercially valuable fruit and wood-producing trees and includes several species used as rootstock in agriculture for their resistance to various abiotic and biotic stressors. We present the full structural and functional genome annotations of six Juglans species and one outgroup within Juglandaceae (Juglans regia, J. cathayensis, J. hindsii, J. microcarpa, J. nigra, J. sigillata and Pterocarya stenoptera) produced using BRAKER2 semi-unsupervised gene prediction pipeline and additional in-house developed tools. For each annotation, gene predictors were trained using 19 tissue-specific J. regia transcriptomes aligned to the genomes. Additional functional evidence and filters were applied to multiexonic and monoexonic putative genes to yield between 27,000 and 44,000 high-confidence gene models per species. Comparison of gene models to the BUSCO embryophyta dataset suggested that, on average, genome annotation completeness was 89.6%. We utilized these high quality annotations to assess gene family evolution within Juglans and among Juglans and selected Eurosid species, which revealed significant contractions in several gene families in J. hindsii including disease resistance-related Wall-associated Kinase (WAK) and Catharanthus roseus Receptor-like Kinase (CrRLK1L) and others involved in abiotic stress response. Finally, we confirmed an ancient whole genome duplication that took place in a common ancestor of Juglandaceae using site substitution comparative analysis.SIGNIFICANCEHigh-quality full genome annotations for six species of walnut (Juglans) and a wingnut (Pterocarya) outgroup were constructed using semi-unsupervised gene prediction followed by gene model filtering and functional characterization. These annotations represent the most comprehensive set for any hardwood genus to date. Comparative analyses based on the gene models uncovered rapid evolution in multiple gene families related to disease-response and a whole genome duplication in a Juglandaceae common ancestor.
The genomic architecture and molecular mechanisms controlling variation in quantitative disease resistance loci are not well understood in plant species and have been barely studied in long-generation trees. Quantitative trait loci mapping and genome-wide association studies were combined to test a large single nucleotide polymorphism (SNP) set for association with quantitative and qualitative white pine blister rust resistance in sugar pine. In the absence of a chromosome-scale reference genome, a high-density consensus linkage map was generated to obtain locations for associated SNPs. Newly discovered associations for white pine blister rust quantitative disease resistance included 453 SNPs involved in wide biological functions, including genes associated with disease resistance and others involved in morphological and developmental processes. In addition, NBS-LRR pathogen recognition genes were found to be involved in quantitative disease resistance, suggesting these newly reported genes are qualitative genes with partial resistance, they are the result of defeated qualitative resistance due to avirulent races, or they have epistatic effects on qualitative disease resistance genes. This study is a step forward in our understanding of the complex genomic architecture of quantitative disease resistance in long-generation trees, and constitutes the first step towards marker-assisted disease resistance breeding in white pine species.
Screening panel. The 288 Pyrus accessions screened with the Axiom™ 700 K Pear Genotyping Array. For each sample, the Table shows the accession’s Plant Introduction (PI) number, the inventory lot identifier, the assigned taxon and common plant name, the origin, the group to which the species belongs (as in Challice and Westwood [60]), the source of the sample, if it failed or passed quality check and the reason for failure. (XLSX 28 kb)
Both a source of diversity and the development of genomic tools, such as reference genomes and molecular markers, are equally important to enable faster progress in plant breeding. Pear (Pyrus spp.) lags far behind other fruit and nut crops in terms of employment of available genetic resources for new cultivar development. To address this gap, we designed a high-density, high-efficiency and robust single nucleotide polymorphism (SNP) array for pear, with the main objectives of conducting genetic diversity and genome-wide association studies. By applying a two-step design process, which consisted of the construction of a first ‘draft’ array for the screening of a small subset of samples, we were able to identify the most robust and informative SNPs to include in the Applied Biosystems™ Axiom™ Pear 70 K Genotyping Array, currently the densest SNP array for pear. Preliminary evaluation of this 70 K array in 1416 diverse pear accessions from the USDA National Clonal Germplasm Repository (NCGR) in Corvallis, OR identified 66,616 SNPs (93% of all the tiled SNPs) as high quality and polymorphic (PolyHighResolution). We further used the Axiom Pear 70 K Genotyping Array to construct high-density linkage maps in a bi-parental population, and to make a direct comparison with available genotyping-by-sequencing (GBS) data, which suggested that the SNP array is a more robust method of screening for SNPs than restriction enzyme reduced representation sequence-based genotyping. The Axiom Pear 70 K Genotyping Array, with its high efficiency in a widely diverse panel of Pyrus species and cultivars, represents a valuable resource for a multitude of molecular studies in pear. The characterization of the USDA-NCGR collection with this array will provide important information for pear geneticists and breeders, as well as for the optimization of conservation strategies for Pyrus.
Statistics about the parental genetic maps of the F1 population P16.009. The number of markers, the length in cM, the average distance between markers (in cM) and the length of the largest gap (in cM) are reported for each Linkage Group (LG) and for the two maps. (XLSX 11 kb)
Despite critical roles in chromosome segregation and disease, the repetitive structure and vast size of centromeres and their surrounding heterochromatic regions impede studies of genomic variation. Here we report the identification of large-scale haplotypes (cenhaps) in humans that span the centromere-proximal regions of all metacentric chromosomes, including the arrays of highly repeated α-satellites on which centromeres form. Cenhaps reveal deep diversity, including entire introgressed Neanderthal centromeres and equally ancient lineages among Africans. These centromere-spanning haplotypes contain variants, including large differences in α-satellite DNA content, which may influence the fidelity and bias of chromosome transmission. The discovery of cenhaps creates new opportunities to investigate their contribution to phenotypic variation, especially in meiosis and mitosis, as well as to more incisively model the unexpectedly rich evolution of these challenging genomic regions.
Genomic analysis in Juglans (walnuts) is expected to transform the breeding and agricultural production of both nuts and lumber. To that end, we report here the determination of reference sequences for six additional relatives of Juglans regia: Juglans sigillata (also from section Dioscaryon), Juglans nigra, Juglans microcarpa, Juglans hindsii (from section Rhysocaryon), Juglans cathayensis (from section Cardiocaryon), and the closely related Pterocarya stenoptera While these are 'draft' genomes, ranging in size between 640Mbp and 990Mbp, their contiguities and accuracies can support powerful annotations of genomic variation that are often the foundation of new avenues of research and breeding. We annotated nucleotide divergence and synteny by creating complete pairwise alignments of each reference genome to the remaining six. In addition, we have re-sequenced a sample of accessions from four Juglans species (including regia). The variation discovered in these surveys comprises a critical resource for experimentation and breeding, as well as a solid complementary annotation. To demonstrate the potential of these resources the structural and sequence variation in and around the polyphenol oxidase loci, PPO1 and PPO2 were investigated. As reported for other seed crops variation in this gene is implicated in the domestication of walnuts. The apparently Juglandaceae specific PPO1 duplicate shows accelerated divergence and an excess of amino acid replacement on the lineage leading to accessions of the domesticated nut crop species, Juglans regia and sigillata.
SummaryOver the last 20 years, global production of Persian walnut (Juglans regia L.) has grown enormously, likely reflecting increased consumption due to its numerous benefits to human health. However, advances in genome‐wide association (GWA) studies and genomic selection (GS) for agronomically important traits in walnut remain limited due to the lack of powerful genomic tools. Here, we present the development and validation of a high‐density 700K single nucleotide polymorphism (SNP) array in Persian walnut. Over 609K high‐quality SNPs have been thoroughly selected from a set of 9.6 m genome‐wide variants, previously identified from the high‐depth re‐sequencing of 27 founders of the Walnut Improvement Program (WIP) of University of California, Davis. To validate the effectiveness of the array, we genotyped a collection of 1284 walnut trees, including 1167 progeny of 48 WIP families and 26 walnut cultivars. More than half of the SNPs (55.7%) fell in the highest quality class of ‘Poly High Resolution’ (PHR) polymorphisms, which were used to assess the WIP pedigree integrity. We identified 151 new parent‐offspring relationships, all confirmed with the Mendelian inheritance test. In addition, we explored the genetic variability among cultivars of different origin, revealing how the varieties from Europe and California were differentiated from Asian accessions. Both the reconstruction of the WIP pedigree and population structure analysis confirmed the effectiveness of the Applied Biosystems™ Axiom™ J. regia 700K SNP array, which initiates a novel genomic and advanced phase in walnut genetics and breeding.
Dissecting the genetic and genomic architecture of complex traits is essential to understand the forces maintaining the variation in phenotypic traits of ecological and economical importance. Whole-genome resequencing data were used to generate high-resolution polymorphic single nucleotide polymorphism (SNP) markers and genotype individuals from common gardens across the loblolly pine (Pinus taeda) natural range. Genome-wide associations were tested with a large phenotypic dataset comprising 409 variables including morphological traits (height, diameter, carbon isotope discrimination, pitch canker resistance), and molecular traits such as metabolites and expression of xylem development genes. Our study identified 2335 new SNP × trait associations for the species, with many SNPs located in physical clusters in the genome of the species; and the genomic location of hotspots for metabolic × genotype associations. We found a highly polygenic basis of quantitative inheritance, with significant differences in number, effects size, genomic location and frequency of alleles contributing to variation in phenotypes in the different traits. While mutation-selection balance might be shaping the genetic variation in metabolic traits, balancing selection is more likely to shape the variation in expression of xylem development genes. Our work contributes to the study of complex traits in nonmodel plant species by identifying associations at a whole-genome level.
A reference genome sequence for Pseudotsuga menziesii var. menziesii (Mirb.) Franco (Coastal Douglas-fir) is reported, thus providing a reference sequence for a third genus of the family Pinaceae. The contiguity and quality of the genome assembly far exceeds that of other conifer reference genome sequences (contig N50 = 44,136 bp and scaffold N50 = 340,704 bp). Incremental improvements in sequencing and assembly technologies are in part responsible for the higher quality reference genome, but it may also be due to a slightly lower exact repeat content in Douglas-fir vs. pine and spruce. Comparative genome annotation with angiosperm species reveals gene-family expansion and contraction in Douglas-fir and other conifers which may account for some of the major morphological and physiological differences between the two major plant groups. Notable differences in the size of the NDH-complex gene family and genes underlying the functional basis of shade tolerance/intolerance were observed. This reference genome sequence not only provides an important resource for Douglas-fir breeders and geneticists but also sheds additional light on the evolutionary processes that have led to the divergence of modern angiosperms from the more ancient gymnosperms.
The 22-gigabase genome of loblolly pine (Pinus taeda) is one of the largest ever sequenced. The draft assembly published in 2014 was built entirely from short Illumina reads, with lengths ranging from 100 to 250 base pairs (bp). The assembly was quite fragmented, containing over 11 million contigs whose weighted average (N50) size was 8206 bp. To improve this result, we generated approximately 12-fold coverage in long reads using the Single Molecule Real Time sequencing technology developed at Pacific Biosciences. We assembled the long and short reads together using the MaSuRCA mega-reads assembly algorithm, which produced a substantially better assembly, P. taeda version 2.0. The new assembly has an N50 contig size of 25 361, more than three times as large as achieved in the original assembly, and an N50 scaffold size of 107 821, 61% larger than the previous assembly.
We investigate the utility and scalability of new read cloud technologies to improve the draft genome assemblies of the colossal, and largely repetitive, genomes of conifers. Synthetic long read technologies have existed in various forms as a means of reducing complexity and resolving repeats since the outset of genome assembly. Recently, technologies that combine subhaploid pools of high molecular weight DNA with barcoding on a massive scale have brought new efficiencies to sample preparation and data generation. When combined with inexpensive light shotgun sequencing, the resulting data can be used to scaffold large genomes. The protocol is efficient enough to consider routinely for even the largest genomes. Conifers represent the largest reference genome projects executed to date. The largest of these is that of the conifer Pinus lambertiana (sugar pine), with a genome size of 31 billion bp. In this paper, we report on the molecular and computational protocols for scaffolding the P. lambertiana genome using the library technology from 10× Genomics. At 247,000 bp, the NG50 of the existing reference sequence is the highest scaffold contiguity among the currently published conifer assemblies; this new assembly’s NG50 is 1.94 million bp, an eightfold increase.