Here, we describe a collection of genomic database portals, SoyBase (https://soybase.org), Legume Information System (https://legumeinfo.org), and PeanutBase (https://peanutbase.org), that support breeding and research work in the legume plant family. The legume family includes important crops such as soybean, peanut, common bean, lentils, chickpeas, as well as approximately 20,000 other species that are important in all terrestrial ecosystems. Beyond the value of the portals for species in this large clade (as well as for plant biology more generally), the database and site architecture of these portals will be of interest to developers of similar genomic sites, as the data management and software solutions are generic and should be applicable to a wide variety of organisms. The architecture for these sites has been designed for rapid, modular, flexible development well suited to genomic data and to rapid incorporation of new data. Website content is handled with a static site generator (Jekyll). Interactive applications are developed using javascript encapsulated as Web Components that access back-end data via APIs for stability and flexibility. This architecture allows for both code portability and for customization to serve the unique needs of each research community.
Apios americana and Apios priceana are tuber-forming legumes native to North America with ecological value and agricultural potential. The lack of genomic resources has limited comparative studies and crop improvement for these species. Here, we report high-quality, haplotype-resolved, chromosome-level genome assemblies for both the species. The assemblies were generated from high-fidelity long-read data, and the primary assemblies were scaffolded using Omni-C chromatin conformation maps. The genome sizes of the primary haplotype assemblies were 1.53 Gb for A. americana and 1.85 Gb for A. priceana, each represented by 11 pseudochromosomes. The BUSCO completeness scores ranged from 98.6% to 99.0%. Approximately 26,000 predicted genes (30,000–33,000 predicted mRNA transcripts) were identified per haplotype in A. americana and A. priceana, respectively. Repeat annotation revealed that over 80% of both genomes consist of interspersed repetitive elements, with the most abundant being long terminal repeat (LTR) retrotransposons. These genomic resources will support trait mapping and structural variation analyses in Apios, and more broadly, comparative genomics within the legume family.
Root nodule symbiosis (RNS) is found in approximately 16–18 widely-separated lineages within the “nitrogen-fixing nodulation clade (NFNC)”. Although modeling of trait gain and loss across approximately 13,000 species within the rosid group indicates multiple gains and losses, there is no consensus about whether RNS had a single or multiple origins; and our understanding is fragmentary regarding the molecular mechanisms underlying those changes. Evolution of a new organ and functions involves many thousands of genes; but the evolutionary histories for many of these genes may be uninformative regarding RNS evolution. A portion of the genes, however, are likely to be derived from prior gene duplications and to have acquired new functions or to have come under new regulatory patterns. Whole genome duplications (WGDs) could conceivably enable the necessary neo- or sub-functionalization for new roles in the nodule. All species that exhibit RNS share a history of several ancient WGDs; but the last such common WGD for these species was the “gamma” paleohexaploidy that occurred early in the core eudicot lineage, ~120 Mya. This presents a puzzle: If legume RNS within the NFNC only arose in the Late Cretaceous, several tens of millions of years after the gamma event, what explains the long, seemingly quiescent interval and the many eudicot lineages without RNS? This study focuses on a collection of gene superfamilies with additional independent WGDs that appear to have occurred in the interim period, after the gamma triplication and prior to the evolution of RNS, identifying several that are both essential for RNS and that show evidence of critical roles of both ancient WGDs and more recent local duplications.
Mungbean (Vigna radiata (L.) R. Wilczek) is a vital source of digestible proteins and is well-suited for the plant-based protein industry. In this study, we analyzed pod morphological traits in the Iowa Mungbean Diversity (IMD) panel of 372 genotypes (2022-2023) using image-analysis-based phenotyping on 2,418 pod images. Pod morphological traits were extracted using deep learning image analysis, achieving excellent agreement with manual measurements (r > 0.96 for pod length (PL) and seed-per-pod (SPP)). Four complementary genome-wide association studies models identified 65 significant SNPs (-log10(P) ≥ 5.56) associated with pod curvature, length, width, and SPP traits. A significant SNP (5_35265704) on chromosome 4 was linked to pod dimensional traits, length, width, and curvature. A candidate gene, Virad04G0076900, located 15.6 kb from this SNP, is part of the GH3 gene family and has an Arabidopsis ortholog (AT4G27260) known for influencing organ elongation, pod, and seed development. Another SNP, 5_210437 on chromosome 6, has been found to be significantly associated with both PL and SPP. A candidate gene, Virad06G0002400 (36.5 kb from this SNP), encodes a potassium transporter and shares homology with the Arabidopsis gene HAK5 (AT4G13420), known to influence pod growth. Image-based measurements achieved genomic prediction accuracies ranging from 0.61 to 0.85 across various traits, demonstrating comparable accuracy to manual methods for linear traits and up to 22% improvement for complex shape traits. These results highlight the potential of deep learning-assisted phenomics integrated with genomic tools to accelerate selection for improved pod architecture in mungbean breeding programs across the Midwestern United States and globally.
The legume family originated ca. 60-65 million years ago and soon diversified into at least six lineages (now extant subfamilies). The signal of whole genome duplications (WGD) is apparent in species sampled from all six subfamilies. The early diversification has posed difficulties for resolving the legume backbone structure and the timing of WGDs, especially in Caesalpinioideae where the diversification and WGD signals coincide. In this study, we report the genome sequences and annotations for Cercis canadensis (Cercidoideae) and Chamaecrista fasciculata (Caesalpinioideae) to help resolve the timings of WGDs relative to subfamily origins and the ancestral legume karyotype. Analyses of genome assemblies from four subfamilies within Fabaceae show that the last common ancestor of all legumes likely had seven chromosomes, with a genome structure similar to the extant Cercis genome. The retained karyotype structure, the lack of a WGD in the last 100+ Mya (Cercis and the lineage leading to it following the eudicot γ whole-genome triplication), and the unusually slow rates of nucleotide substitution and structural evolution in the Cercis genome underscore its utility as a genomic proxy for the last common ancestor of all legume species. Our analysis supports an allopolyploid origin of Caesalpinioideae, with progenitors from lineages along the backbone of the legume phylogeny. Rapid diversification and the inferred allopolyploid origin of Caesalpinioideae provide a partial explanation for the difficulty in resolving the backbone of the legume phylogeny and early Caesalpinioideae diversification.
Improving seed size and weight is a major breeding goal in mungbean (Vigna radiata (L.) R. Wilczek). Improved genomic resources and precision phenotyping may enable more efficient selection for seed trait improvement. In this study, we integrated a deep learning-based segment anything model phenotyping pipeline with genome-wide association studies (GWAS), comparative mapping, and genomic prediction (GP) to dissect the genetic architecture of seed size, shape, and weight traits in the Iowa mungbean diversity panel. The zero-shot segmentation approach reliably captured seed size traits, which exhibited high heritability (H2 > 0.90) and strong correlation with seed weight (r > 0.91). A multi-model GWAS identified 82 unique single nucleotide polymorphisms (SNPs) across all seven traits, of which 13 were major pleiotropic SNPs governing multiple seed dimensions, including high-confidence regions on chromosomes 1, 4, and 6 that explained over 20% of the phenotypic variance. Within these SNP regions, comparative mapping highlighted candidate genes including an ABC transporter (Virad01G0084400) and two colocated candidates, NPGR1 (Virad06G0255600) and a RING-type E3 ubiquitin-protein ligase (Virad06G0255800), presented as hypothesis-generating candidates for seed size regulation. GP using genomic best linear unbiased prediction (gBLUP) produced moderate to high accuracies for seed size and weight traits (r = 0.76-0.84). Incorporating significant GWAS SNPs (gBLUP + SNPs) yielded slight improvements, suggesting the standard gBLUP model is sufficiently robust for selection. Collectively, this study provides the most comprehensive genomic dissection of seed size and weight traits in mungbean to date, providing candidate loci and genomic prediction models that can accelerate genetic improvement for seed yield and quality.
Contradictory lines of evidence have made it difficult to resolve the phylogenetic history of the legume diversification era; this is true for the backbone topology, and for the number and timing of whole genome duplications (WGDs). By analyzing the transcriptomic data for 473 gene families in 76 species covering all six accepted legume subfamilies, we assessed the phylogenetic relationships of the legume backbone and uncovered evidence of independent whole genome duplications in each of the six legume subfamilies. Three subfamilies - Cercidoideae, Dialioideae, and Caesalpinioideae - bear evidence of an allopolyploid duplication pattern suggestive of ancient hybridization. In Cercidoideae and Dialioideae, the hybridization appears to be within-subfamily, with the genera Cercis and Poeppigia apparently unduplicated descendants of one of the parental lineages. In Caesalpinioideae, the hybridization appears to involve a member of the Papilionoideae lineage, and some other lineage, potentially extinct. Several independent lines of evidence converged on a single backbone hypothesis and the above hypotheses of reticulate evolution: phylogenies calculated from both superalignments and from multi-tree coalescent-based analyses; concordance factor analysis of the set of gene family alignments and topologies; and direct inference of reticulation events via maximum pseudo-likelihood implemented by PhyloNet.
This study, the first in a three-part series, lays the foundation for understanding the origin of the peanut crop (Arachis hypogaea). Its subsequent evolution is explored in the two papers that follow. The evidence that A. hypogaea originated from a single hybridization event between Arachis duranensis and Arachis ipaënsis less than 10 000 years ago was already very strong. Here, we extend this evidence using more than 1600 single-nucleotide polymorphisms to make an almost exhaustive comparison of wild Arachis section germplasm conserved ex situ with the A and B subgenomes of divergent, sequenced cultivated peanuts. The wild relatives of peanut are highly selfing and their geocarpy means they plant their own seeds, allowing them to persist as discrete populations for millennia. This unusual biology creates a rare opportunity for genetic archaeology: ancestral lineages can be identified with exceptional precision. Our results reaffirm a single origin for the cultigen, identifying A. duranensis from Río Seco and A. ipaënsis K 30076 as the closest known relatives of the A and B subgenomes of peanut. As a genomic resource, we generated a chromosome-scale assembly of the Río Seco A. duranensis K 30065 and confirmed that it is more closely related to the A subgenome of peanut than the current reference genome (V14167). Even if somewhat closer wild accessions were found through new field collections, they would still belong to the same ancestral lineage. With this level of evidence, the origin of peanut is now known in greater detail than that of any other ancient polyploid crop.
Cultivar 'Williams 82' has served as the reference genome for the soybean research community since 2008, but is known to have areas of genomic heterogeneity among different sub-lines. This work provides an updated assembly (version Wm82.a6) derived from a specific sub-line known as 'Wm82-ISU-01' (seeds available under USDA accession PI 704477). The genome was assembled using Pacific BioSciences HiFi reads and integrated into chromosomes using HiC. The 20 soybean chromosomes assembled into a genome of 1.01Gb, consisting of 36 contigs. The genome annotation identified 48,387 gene models, named in accordance with previous assembly versions Wm82.a2 and Wm82.a4. Comparisons of Wm82.a6 with other near-gapless assemblies of 'Williams 82' reveal large regions of genomic heterogeneity, including regions of differential introgression from the genotype 'Kingwa' within approximately 30 Mb and 25 Mb segments on chromosomes 03 and 07, respectively. Additionally, our analysis revealed a previously unknown large (~20 Mb) heterogeneous region in the pericentromeric region of chromosome 12, where Wm82.a6 matches the 'Williams' haplotype while the other two near-gapless assemblies do not match the haplotype of either parent of 'Williams 82'. In addition to the Wm82.a6 assembly, we also assembled the genome of soybean line 'Fiskeby III', a rich resource for abiotic stress resistance genes. A genome comparison of Wm82.a6 with 'Fiskeby III' revealed the nucleotide and structural polymorphisms between the two genomes within a QTL region for iron deficiency chlorosis resistance. The Wm82.a6 and 'Fiskeby III' genomes described here will enhance comparative and functional genomics capacities and applications in the soybean community.
SUMMARY:Identification of allelic or corresponding genes (pan-genes) within a species or genus is important for discovery of biologically significant genetic conservation and variation. Similarly, identification of orthologs (gene families) across wider evolutionary distances is important for understanding the genetic basis for similar or differing traits. Especially in plants, several complications make identification of pan-genes and gene families challenging, including whole-genome duplications, evolutionary rate differences among lineages, and varying qualities of assemblies and annotations. Here, we document and distribute a set of workflows that we have used to address these problems. RESULTS:Pandagma is a set of configurable workflows for identifying and comparing pan-gene sets and gene families for annotation sets from eukaryotic genomes, using a combination of homology, synteny, and expected rates of synonymous change in coding sequence. AVAILABILITY AND IMPLEMENTATION:The Pandagma workflows, example configurations, implementation details, and scripts for retrieving public datasets, are available at https://github.com/legumeinfo/pandagma.
### Competing Interest Statement The authors have declared no competing interest.
Abstract This strategic plan summarizes the major accomplishments achieved in the last quinquennial by the soybean [Glycine max (L.) Merr.] genetics and genomics research community and outlines key priorities for the next 5 years (2024–2028). This work is the result of deliberations among over 50 soybean researchers during a 2‐day workshop in St Louis, MO, USA, at the end of 2022. The plan is divided into seven traditional areas/disciplines: Breeding, Biotic Interactions, Physiology and Abiotic Stress, Functional Genomics, Biotechnology, Genomic Resources and Datasets, and Computational Resources. One additional section was added, Training the Next Generation of Soybean Researchers, when it was identified as a pressing issue during the workshop. This installment of the soybean genomics strategic plan provides a snapshot of recent progress while looking at future goals that will improve resources and enable innovation among the community of basic and applied soybean researchers. We hope that this work will inform our community and increase support for soybean research.
Background Mung bean ( Vigna radiata (L.) Wilczek), is an important pulse crop in the global south. Early flowering and maturation are advantageous traits for adaptation to northern and southern latitudes. This study investigates the genetic basis of the Days-to-Flowering trait (DTF) in mung bean, combining genome-wide association studies (GWAS) in mung bean and comparisons with orthologous genes involved with control of DTF responses in soybean ( Glycine max (L) Merr) and Arabidopsis ( Arabidopsis thaliana ). Results The most significant associations for DTF were on mung bean chromosomes 1, 2, and 4. Only the SNPs on chromosomes 1 and 4 were heavily investigated using downstream analysis. The chromosome 1 DTF association is tightly linked with a cluster of locally duplicated FERONIA ( FER ) receptor-like protein kinase genes, and the SNP occurs within one of the FERONIA genes. In Arabidopsis, an orthologous FERONIA gene ( AT3G51550 ), has been reported to regulate the expression of the FLOWERING LOCUS C ( FLC ). For the chromosome 4 DTF locus, the strongest candidates are Vradi04g00002773 and Vradi04g00002778 , orthologous to the Arabidopsis PhyA and PIF3 genes, encoding phytochrome A (a photoreceptor protein sensitive to red to far-red light) and phytochrome-interacting factor 3, respectively. The soybean PhyA orthologs include the classical loci E3 and E4 (genes GmPhyA3, Glyma.19G224200, and GmPhyA2, Glyma.20G090000 ). The mung bean PhyA ortholog has been previously reported as a candidate for DTF in studies conducted in South Korea. Conclusion The top two identified SNPs accounted for a significant proportion (~ 65%) of the phenotypic variability in mung bean DTF by the six significant SNPs (39.61%), with a broad-sense heritability of 0.93. The strong associations of DTF with genes that have orthologs with analogous functions in soybean and Arabidopsis provide strong circumstantial evidence that these genes are causal for this trait. The three reported loci and candidate genes provide useful targets for marker-assisted breeding in mung beans.
Introduction Virginia-type peanut, Arachis hypogaea subsp. hypogaea, is the second largest market class of peanut cultivated in the United States. It is mainly used for large-seeded, in-shell products. Historically, Virginia-type peanut cultivars were developed through long-term recurrent phenotypic selection and wild species introgression projects. Contemporary genomic technologies represent a unique opportunity to revolutionize the traditional breeding pipeline. While there are genomic tools available for wild and cultivated peanuts, none are tailored specifically to applied Virginia-type cultivar development programs. Methods and respective results Here, the first Virginia-type peanut reference genome, “Bailey II”, was assembled. It has improved contiguity and reduced instances of manual curation in chromosome arms. Whole-genome sequencing and marker discovery was conducted on 66 peanut lines which resulted in 1.15 million markers. The high marker resolution achieved allowed 34 unique wild species introgression blocks to be cataloged in the A. hypogaea genome, some of which are known to confer resistance to one or more pathogens. To enable marker-assisted selection of the blocks, 111 PCR Allele Competitive Extension assays were designed. Forty thousand high quality markers were selected from the full set that are suitable for mid-density genotyping for genomic selection. Genomic data from representative advanced Virginia-type peanut lines suggests this is an appropriate base population for genomic selection. Discussion The findings and tools produced in this research will allow for rapid genetic gain in the Virginia-type peanut population. Genomics-assisted breeding will allow swift response to changing biotic and abiotic threats, and ultimately the development of superior cultivars for public use and consumption.
Motivation Genotyping by sequencing is a powerful tool for investigating genetic variation in plants, but many economically important plants are allopolyploids, where homoeologous similarity obscures the subgenomic origin of reads and confounds allelic and homoeologous SNPs. Recent polyploid genotyping methods use allelic frequencies, rate of heterozygosity, parental cross or other information to resolve read assignment, but good subgenomic references offer the most direct information. The typical strategy aligns reads to the joint reference, performs diploid genotyping within each subgenome, and filters the results, but persistent read misassignment results in an excess of false heterozygous calls. Results We introduce the Comprehensive Allopolyploid Genotyper (CAPG), which formulates an explicit likelihood to weight read alignments against both subgenomic references and genotype individual allopolyploids from whole genome resequencing (WGS) data. We demonstrate CAPG in allotetraploids, where it performs better than GATK’s HaplotypeCaller applied to reads aligned to the combined subgenomic references. Availability Code and tutorials are available at https://github.com/Kkulkarni1/CAPG.git .
The fatty acid composition of seed oil is a major determinant of the flavor, shelf-life, and nutritional quality of peanuts. Major QTLs controlling high oil content, high oleic content, and low linoleic content have been characterized in several seed oil crop species. Here we employ genome-wide association approaches on a recently genotyped collection of 787 plant introduction accessions in the USDA peanut core collection, plus selected improved cultivars, to discover markers associated with the natural variation in fatty acid composition, and to explain the genetic control of fatty acid composition in seed oils. Overall, 251 single nucleotide polymorphisms (SNPs) had significant trait associations with the measured fatty acid components. Twelve SNPs were associated with two or three different traits. Of these loci with apparent pleiotropic effects, 10 were associated with both oleic (C18:1) and linoleic acid (C18:2) content at different positions in the genome. In all 10 cases, the favorable allele had an opposite effect - increasing and lowering the concentration, respectively, of oleic and linoleic acid. The other traits with pleiotropic variant control were palmitic (C16:0), behenic (C22:0), lignoceric (C24:0), gadoleic (C20:1), total saturated, and total unsaturated fatty acid content. One hundred (100) of the significantly associated SNPs were located within 1000 kbp of 55 genes with fatty acid biosynthesis functional annotations. These genes encoded, among others: ACCase carboxyl transferase subunits, and several fatty acid synthase II enzymes. With the exception of gadoleic (C20:1) and lignoceric (C24:0) acid content, which occur at relatively low abundance in cultivated peanut, all traits had significant SNP interactions exceeding a stringent Bonferroni threshold (α = 1%). We detected 7,682 pairwise SNP interactions affecting the relative abundance of fatty acid components in the seed oil. Of these, 627 SNP pairs had at least one SNP within 1000 kbp of a gene with fatty acid biosynthesis functional annotation. We evaluated 168 candidate genes underlying these SNP interactions. Functional enrichment and protein-to-protein interactions supported significant interactions (p- value < 1.0E-16) among the genes evaluated. These results show the complex nature of the biology and genes underlying the variation in seed oil fatty acid composition and contribute to an improved genotype-to-phenotype map for fatty acid variation in peanut seed oil. Key phrases SNP Genotyping, Genome-wide Association Study (GWAS), GWAS of interacting SNPs (GWASi), Pleiotropy, Seed fatty acid composition, Oleic-Linoleic acid ratio.
Accounting for field variation patterns plays a crucial role in interpreting phenotype data and, thus, in plant breeding. Several spatial models have been developed to account for field variation. Spatial analyses show that spatial models can successfully increase the quality of phenotype measurements and subsequent selection accuracy for continuous data types such as grain yield and plant height. The phenotypic data for stress traits are usually recorded in ordinal data scores but are traditionally treated as numerical values with normal distribution, such as iron deficiency chlorosis (IDC). The effectiveness of spatial adjustment for ordinal data has not been systematically compared. The research objective described here is to evaluate methods for spatial adjustment of ordinal data, using soybean IDC as an example. Comparisons of adjustment effectiveness for spatial autocorrelation were conducted among eight different models. The models were divided into three groups: Group I, moving average grid adjustment; group II, geospatial autoregressive regression (SAR) models; and Group III, tensor product penalized P-splines. Results from the model comparison show that the effectiveness of the models depends on the severity of field variation, the irregularity of the variation pattern, and the model used. The geospatial SAR models outperform the other models for ordinal IDC data. Prediction accuracy for the lines planted in the IDC high-pressure area is 11.9% higher than those planted in low-IDC-pressure regions. The relative efficiency of the mixed SAR model is 175%, relative to the baseline ordinary least squares model. Even though the geospatial SAR model is the best among all the compared models, the efficiency is not as good for ordinal data types as for numeric data.
Polyploidy and life-strategy transitions between annuality and perenniality often occur in flowering plants. However, the evolutionary propensities of polyploids and the genetic bases of such transitions remain elusive. We assembled chromosome-level genomes of representative perennial species across the genus Glycine including five diploids and a young allopolyploid, and constructed a Glycine super-pangenome framework by integrating 26 annual soybean genomes. These perennial diploids exhibit greater genome stability and possess fewer centromere repeats than the annuals. Biased subgenomic fractionation occurred in the allopolyploid, primarily by accumulation of small deletions in gene clusters through illegitimate recombination, which was associated with pre-existing local subgenomic differentiation. Two genes annotated to modulate vegetative–reproductive phase transition and lateral shoot outgrowth were postulated as candidates underlying the perenniality–annuality transition. Our study provides insights into polyploid genome evolution and lays a foundation for unleashing genetic potential from the perennial gene pool for soybean improvement. Assemblies of six representative perennial Glycine genomes and a comparison with annual soybean genomes reveal evolutionary patterns, differentiation and adaptation of annual and perennial genomes and mechanisms driving subgenome fractionation.
Mung bean [Vigna radiata (L.) Wilczek] is a drought-tolerant, short-duration crop, and a rich source of protein and other valuable minerals, vitamins, and antioxidants. The main objectives of this research were (1) to study the root traits related with the phenotypic and genetic diversity of 375 mung bean genotypes of the Iowa (IA) diversity panel and (2) to conduct genome-wide association studies of root-related traits using the Automated Root Image Analysis (ARIA) software. We collected over 9,000 digital images at three-time points (days 12, 15, and 18 after germination). A broad sense heritability for days 15 (0.22-0.73) and 18 (0.23-0.87) was higher than that for day 12 (0.24-0.51). We also reported root ideotype classification, i.e., PI425425 (India), PI425045 (Philippines), PI425551 (Korea), PI264686 (Philippines), and PI425085 (Sri Lanka) that emerged as the top five in the topsoil foraging category, while PI425594 (unknown origin), PI425599 (Thailand), PI425610 (Afghanistan), PI425485 (India), and AVMU0201 (Taiwan) were top five in the drought-tolerant and nutrient uptake "steep, cheap, and deep" ideotype. We identified promising genotypes that can help diversify the gene pool of mung bean breeding stocks and will be useful for further field testing. Using association studies, we identified markers showing significant associations with the lateral root angle (LRA) on chromosomes 2, 6, 7, and 11, length distribution (LED) on chromosome 8, and total root length-growth rate (TRL_GR), volume (VOL), and total dry weight (TDW) on chromosomes 3 and 5. We discussed genes that are potential candidates from these regions. We reported beta-galactosidase 3 associated with the LRA, which has previously been implicated in the adventitious root development via transcriptomic studies in mung bean. Results from this work on the phenotypic characterization, root-based ideotype categories, and significant molecular markers associated with important traits will be useful for the marker-assisted selection and mung bean improvement through breeding.
In this chapter, we introduce the main components of the Legume Information System ( https://legumeinfo.org ) and several associated resources. Additionally, we provide an example of their use by exploring a biological question: is there a common molecular basis, across legume species, that underlies the photoperiod-mediated transition from vegetative to reproductive development, that is, days to flowering? The Legume Information System (LIS) holds genetic and genomic data for a large number of crop and model legumes and provides a set of online bioinformatic tools designed to help biologists address questions and tasks related to legume biology. Such tasks include identifying the molecular basis of agronomic traits; identifying orthologs/syntelogs for known genes; determining gene expression patterns; accessing genomic datasets; identifying markers for breeding work; and identifying genetic similarities and differences among selected accessions. LIS integrates with other legume-focused informatics resources such as SoyBase ( https://soybase.org ), PeanutBase ( https://peanutbase.org ), and projects of the Legume Federation ( https://legumefederation.org ).