Abstract Kiwifruit is an economically and nutritionally important fruit crop with extremely high contents of vitamin C. However, the previously released versions of kiwifruit genomes all have a mass of unanchored or missing regions. Here, we report a highly continuous and completely gap-free reference genome of Actinidia chinensis cv. ‘Hongyang’, named Hongyang v4.0, which is the first to achieve two de novo haploid-resolved haplotypes, HY4P and HY4A. HY4P and HY4A have a total length of 606.1 and 599.6 Mb, respectively, with almost the entire telomeres and centromeres assembled in each haplotype. In comparison with Hongyang v3.0, the integrity and contiguity of Hongyang v4.0 is markedly improved by filling all unclosed gaps and correcting some misoriented regions, resulting in ~38.6–39.5 Mb extra sequences, which might affect 4263 and 4244 protein-coding genes in HY4P and HY4A, respectively. Furthermore, our gap-free genome assembly provides the first clue for inspecting the structure and function of centromeres. Globally, centromeric regions are characterized by higher-order repeats that mainly consist of a 153-bp conserved centromere-specific monomer (Ach-CEN153) with different copy numbers among chromosomes. Functional enrichment analysis of the genes located within centromeric regions demonstrates that chromosome centromeres may not only play physical roles for linking a pair of sister chromatids, but also have genetic features for participation in the regulation of cell division. The availability of the telomere-to-telomere and gap-free Hongyang v4.0 reference genome lays a solid foundation not only for illustrating genome structure and functional genomics studies but also for facilitating kiwifruit breeding and improvement.
In this chapter, we introduce the main components of the Legume Information System ( https://legumeinfo.org ) and several associated resources. Additionally, we provide an example of their use by exploring a biological question: is there a common molecular basis, across legume species, that underlies the photoperiod-mediated transition from vegetative to reproductive development, that is, days to flowering? The Legume Information System (LIS) holds genetic and genomic data for a large number of crop and model legumes and provides a set of online bioinformatic tools designed to help biologists address questions and tasks related to legume biology. Such tasks include identifying the molecular basis of agronomic traits; identifying orthologs/syntelogs for known genes; determining gene expression patterns; accessing genomic datasets; identifying markers for breeding work; and identifying genetic similarities and differences among selected accessions. LIS integrates with other legume-focused informatics resources such as SoyBase ( https://soybase.org ), PeanutBase ( https://peanutbase.org ), and projects of the Legume Federation ( https://legumefederation.org ).
SoyBase, the USDA-ARS soybean genetic database, is a comprehensive repository for professionally curated genetics, genomics and related data resources for soybean. SoyBase contains the most current genetic, physical and genomic sequence maps integrated with qualitative and quantitative traits. The quantitative trait loci (QTL) represent more than 18 years of QTL mapping of more than 90 unique traits. SoyBase also contains the well-annotated ‘Williams 82’ genomic sequence and associated data mining tools. The genetic and sequence views of the soybean chromosomes and the extensive data on traits and phenotypes are extensively interlinked. This allows entry to the database using almost any kind of available information, such as genetic map symbols, soybean gene names or phenotypic traits. SoyBase is the repository for controlled vocabularies for soybean growth, development and trait terms, which are also linked to the more general plant ontologies. SoyBase can be accessed at http://soybase.org.
We report reference-quality genome assemblies and annotations for two accessions of soybean (Glycine max) and for one accession of Glycine soja, the closest wild relative of G. max. The G. max assemblies provided are for widely used US cultivars: the northern line Williams 82 (Wm82) and the southern line Lee. The Wm82 assembly improves the prior published assembly, and the Lee and G. soja assemblies are new for these accessions. Comparisons among the three accessions show generally high structural conservation, but nucleotide difference of 1.7 single-nucleotide polymorphisms (snps) per kb between Wm82 and Lee, and 4.7 snps per kb between these lines and G. soja. snp distributions and comparisons with genotypes of the Lee and Wm82 parents highlight patterns of introgression and haplotype structure. Comparisons against the US germplasm collection show placement of the sequenced accessions relative to global soybean diversity. Analysis of a pan-gene collection shows generally high conservation, with variation occurring primarily in genomically clustered gene families. We found approximately 40-42 inversions per chromosome between either Lee or Wm82v4 and G. soja, and approximately 32 inversions per chromosome between Wm82 and Lee. We also investigated five domestication loci. For each locus, we found two different alleles with functional differences between G. soja and the two domesticated accessions. The genome assemblies for multiple cultivated accessions and for the closest wild ancestor of soybean provides a valuable set of resources for identifying causal variants that underlie traits for the domestication and improvement of soybean, serving as a basis for future research and crop improvement efforts for this important crop species.
Reconstruction of an ancestral genome for a set of plant species has been a challenging task because of complex histories that may include whole-genome duplications, segmental duplications, independent gene duplications or losses, diploidization and rearrangement events. Here, we describe the reconstruction a hypothetical ancestral genome for the papilionoid legumes (the largest subfamily within the third largest family in flowering plants), and evaluate the results relative to phylogenetic and chromosomal count data for this group of legumes, spanning 294 diverse papilionoid genera. To reconstruct the ancestral genomes for nine legume species with sequenced genomes, we used a maximum likelihood approach combined with a novel method for identifying informative markers for this purpose. Analyzing genomes from four species within the Phaseoleae, two in Dalbergieae, two in the 'inverted repeat loss' clade, and one in the Robinieae, we infer a common ancestral genome with nine chromosomes. The reconstructed genome structural histories are consistent with chromosomal and phylogenetic histories, but we also infer that a common ancestor with nine chromosomes was probably intermediate to an earlier state of 14 chromosomes following a whole-genome duplication that pre-dated the radiation of the papilionoid legumes, evidence for which is found in early-diverging papilionoid lineages.
Like many other crops, the cultivated peanut ( Arachis hypogaea L.) is of hybrid origin and has a polyploid genome that contains essentially complete sets of chromosomes from two ancestral species. Here we report the genome sequence of peanut and show that after its polyploid origin, the genome has evolved through mobile-element activity, deletions and by the flow of genetic information between corresponding ancestral chromosomes (that is, homeologous recombination). Uniformity of patterns of homeologous recombination at the ends of chromosomes favors a single origin for cultivated peanut and its wild counterpart A. monticola . However, through much of the genome, homeologous recombination has created diversity. Using new polyploid hybrids made from the ancestral species, we show how this can generate phenotypic changes such as spontaneous changes in the color of the flowers. We suggest that diversity generated by these genetic mechanisms helped to favor the domestication of the polyploid A. hypogaea over other diploid Arachis species cultivated by humans.
The factors behind genome size evolution have been of great interest, considering that eukaryotic genomes vary in size by more than three orders of magnitude. Using a model of two wild peanut relatives, Arachis duranensis and Arachis ipaensis, in which one genome experienced large rearrangements, we find that the main determinant in genome size reduction is a set of inversions that occurred in A. duranensis, and subsequent net sequence removal in the inverted regions. We observe a general pattern in which sequence is lost more rapidly at newly distal (telomeric) regions than it is gained at newly proximal (pericentromeric) regions - resulting in net sequence loss in the inverted regions. The major driver of this process is recombination, determined by the chromosomal location. Any type of genomic rearrangement that exposes proximal regions to higher recombination rates can cause genome size reduction by this mechanism. In comparisons between A. duranensis and A. ipaensis, we find that the inversions all occurred in A. duranensis. Sequence loss in those regions was primarily due to removal of transposable elements. Illegitimate recombination is likely the major mechanism responsible for the sequence removal, rather than unequal intrastrand recombination. We also measure the relative rate of genome size reduction in these two Arachis diploids. We also test our model in other plant species and find that it applies in all cases examined, suggesting our model is widely applicable.
Recombination (R) rate and linkage disequilibrium (LD) analyses are the basis for plant breeding. These vary by breeding system, by generation of inbreeding or outcrossing and by region in the chromosome. Common bean (Phaseolus vulgaris L.) is a favored food legume with a small sequenced genome (514 Mb) and n = 11 chromosomes. The goal of this study was to describe R and LD in the common bean genome using a 768-marker array of single nucleotide polymorphisms (SNP) based on Trans-legume Orthologous Group (TOG) genes along with an advanced-generation Recombinant Inbred Line reference mapping population (BAT93 x Jalo EEP558) and an internationally available diversity panel. A whole genome genetic map was created that covered all eleven linkage groups (LG). The LGs were linked to the physical map by sequence data of the TOGs compared to each chromosome sequence of common bean. The genetic map length in total was smaller than for previous maps reflecting the precision of allele calling and mapping with SNP technology as well as the use of gene-based markers. A total of 91.4% of TOG markers had singleton hits with annotated Pv genes and all mapped outside of regions of resistance gene clusters. LD levels were found to be stronger within the Mesoamerican genepool and decay more rapidly within the Andean genepool. The recombination rate across the genome was 2.13 cM / Mb but R was found to be highly repressed around centromeres and frequent outside peri-centromeric regions. These results have important implications for association and genetic mapping or crop improvement in common bean.
David Bertioli and colleagues report the genomes of Arachis duranensis and Arachis ipaensis, the diploid ancestors of cultivated peanut, Arachis hypogaea. Their analyses are a first step in understanding the evolution of the peanut's tetraploid genome. Cultivated peanut (Arachis hypogaea) is an allotetraploid with closely related subgenomes of a total size of ∼2.7 Gb. This makes the assembly of chromosomal pseudomolecules very challenging. As a foundation to understanding the genome of cultivated peanut, we report the genome sequences of its diploid ancestors (Arachis duranensis and Arachis ipaensis). We show that these genomes are similar to cultivated peanut's A and B subgenomes and use them to identify candidate disease resistance genes, to guide tetraploid transcript assemblies and to detect genetic exchange between cultivated peanut's subgenomes. On the basis of remarkably high DNA identity of the A. ipaensis genome and the B subgenome of cultivated peanut and biogeographic evidence, we conclude that A. ipaensis may be a direct descendant of the same population that contributed the B subgenome to cultivated peanut.
Summary Lupins are important grain legume crops that form a critical part of sustainable farming systems, reducing fertilizer use and providing disease breaks. It has a basal phylogenetic position relative to other crop and model legumes and a high speciation rate. Narrow‐leafed lupin (NLL; Lupinus angustifolius L.) is gaining popularity as a health food, which is high in protein and dietary fibre but low in starch and gluten‐free. We report the draft genome assembly (609 Mb) of NLL cultivar Tanjil, which has captured >98% of the gene content, sequences of additional lines and a dense genetic map. Lupins are unique among legumes and differ from most other land plants in that they do not form mycorrhizal associations. Remarkably, we find that NLL has lost all mycorrhiza‐specific genes, but has retained genes commonly required for mycorrhization and nodulation. In addition, the genome also provided candidate genes for key disease resistance and domestication traits. We also find evidence of a whole‐genome triplication at around 25 million years ago in the genistoid lineage leading to Lupinus. Our results will support detailed studies of legume evolution and accelerate lupin breeding programmes.
Legume Information System (LIS), at http://legumeinfo.org, is a genomic data portal (GDP) for the legume family. LIS provides access to genetic and genomic information for major crop and model legumes. With more than two-dozen domesticated legume species, there are numerous specialists working on particular species, and also numerous GDPs for these species. LIS has been redesigned in the last three years both to better integrate data sets across the crop and model legumes, and to better accommodate specialized GDPs that serve particular legume species. To integrate data sets, LIS provides genome and map viewers, holds synteny mappings among all sequenced legume species and provides a set of gene families to allow traversal among orthologous and paralogous sequences across the legumes. To better accommodate other specialized GDPs, LIS uses open-source GMOD components where possible, and advocates use of common data templates, formats, schemas and interfaces so that data collected by one legume research community are accessible across all legume GDPs, through similar interfaces and using common APIs. This federated model for the legumes is managed as part of the ‘Legume Federation’ project (accessible via http://legumefederation.org), which can be thought of as an umbrella project encompassing LIS and other legume GDPs.
UNLABELLEDGeneContent is a software system to infer the genome phylogeny based on an additive genome distance that can be estimated from the extended gene content data, which contains the genome-wide information (absence of a gene family, presence as single copy or presence as duplicates) across multiple species. GeneContent can also be used to explore the genome-wide evolutionary pattern of gene loss and proliferation.AVAILABILITYDistribution packages of GeneContent for both Microsoft Windows and Linux operating systems are available at http://xgu.zool.iastate.eduCONTACTxgu@iastate.edu.