The vast amount of genomic data available for various plant species has enabled a greater understanding of genomic regions associated with phenotypic variation, mainly through approaches such as Genome-Wide Association Studies (GWAS). However, relying solely on GWAS results limits our understanding of trait variation, as it focuses on individual SNPs without considering the linked variants that may also influence the trait. As a follow-up to GWAS, haplotype analysis focusing on local regions of the genome can provide detailed insights into the genetic variants associated with a trait. crosshap, a local haplotyping R-package, identifies and visualises haplotype structures, allowing users to explore allelic diversity and associated phenotypic variation in targeted genomic regions. This chapter demonstrates local haplotyping analysis using crosshap to understand the genetic control of flowering time in domesticated and wild soybean, informing breeding strategies and crop improvement efforts; however, crosshap can be applied to any species where suitable data is available.
Genome-wide association studies (GWAS) are a valuable approach to identify single-nucleotide polymorphisms (SNPs) associated with a phenotype of interest. There are now a variety of R-packages and command-line tools available to perform GWAS. Here, we provide an example downloading and filtering SNP data, followed by GWAS analysis using the R-package rMVP.
Orphan crops are important sources of nutrition in developing regions and many are tolerant to biotic and abiotic stressors; however, modern crop improvement technologies have not been widely applied to orphan crops due to the lack of resources available. There are orphan crop representatives across major crop types and the conservation of genes between these related species can be used in crop improvement. Machine learning (ML) has emerged as a promising tool for crop improvement. Transferring knowledge from major crops to orphan crops and using machine learning to improve accuracy and efficiency can be used to improve orphan crops.
Pyrus pyrifolia, commonly known as sand pear, is a key economic fruit tree in temperate regions that possesses highly diverse germplasm resources for pear quality improvement. However, research on the relationship between resistance and fruit quality traits in the breeding of fruit species like pear is limited. Pan-transcriptomes effectively capture genetic information from coding regions and reflect variations in gene expression between individuals. Here, we constructed a pan-transcriptome based on 506 samples from different tissues of sand pear, and explored the intrinsic relationships among phenotypes and the selection for disease resistance during improvement based on expression presence/absence variations (ePAVs). The pan-transcriptome in this study contains 156,744 transcripts, among which the novel transcripts showed significant enrichment in the defense response. Interestingly, disease resistance genes are highly expressed in landraces of pear but have been selected against during the improvement of this perennial tree species. We found that the genetically diverse landraces can be divided into two subgroups and inferred that they have undergone different dispersal processes. Through co-expression network analysis, we confirmed that the formation of stone cells in pears, the synthesis of fruit anthocyanins, and the ability to resist stress are interrelated. They are jointly regulated by several modules, and the expression of regulatory genes has significant correlations with these three processes. Moreover, we identified candidate genes such as HKL1 that may affect sugar content and are missing from the reference genome. This study provides insights into the associations between complex fruit traits, while providing a database resource for pear disease resistance and fruit quality breeding.
Multi-omics assisted prediction of disease resistance mechanisms using machine learning has the potential to accelerate the breeding of resistant legume varieties. Grain legumes, such as soybean (Glycine max (L.) Merr.), chickpea (Cicer arietinum L.), and lentil (Lens culinaris Medik.) play an important role in combating micronutrient malnutrition in the growing human population. However, plant diseases significantly reduce grain yield, causing 10–40
Lupin crops provide nutritious seeds as an excellent source of dietary protein. However, extensive genomic resources are needed for crop improvement, focusing on key traits such as nutritional value and climate resiliency, to ensure global food security based on sustainable and healthy diets for all. Such resources can be derived either from related lupin species or crop wild relatives, which represent a large and untapped source of genetic variation for crop improvement. Here, we report genome assemblies of the cross-compatible species Lupinus cosentinii (Mediterranean) and its pan-Saharan wild relative L. digitatus, which are well adapted to drought-prone environments and partially domesticated. We show that both species are tetraploids, and their repetitive DNA content differs considerably from that of the main lupin crops L. angustifolius and L. albus. We present the complex evolutionary process within the rough-seeded lupins as a species-based model involving polyploidization and rediploidization. Our data also provide the foundation for a systematic analysis of genomic diversity among lupin species to promote their exploitation for crop improvement and sustainable agriculture.
This study identifies and classifies resistance gene analogues (RGAs) in the genomes of Brassica nigra, Sinapis arvensis and Sinapis alba using the RGAugury pipeline. RGAs were categorised into four main classes: receptor-like kinases (RLKs), receptor-like proteins (RLPs), nucleotide-binding leucine-rich repeat (NLR) proteins and transmembrane-coiled-coil (TM-CC) genes. A total of 4499 candidate RGAs were detected, with species-specific proportions. RLKs were the most abundant across all genomes, followed by TM-CCs and RLPs. The sub-classification of RLKs and RLPs identified LRR-RLKs, LRR-RLPs, LysM-RLKs, and LysM-RLPs. Atypical NLRs were more frequent than typical ones in all species. Atypical NLRs were more frequent than typical ones in all species. We explored the relationship between chromosome size and RGA count using regression analysis. In B. nigra and S. arvensis, larger chromosomes generally harboured more RGAs, while S. alba displayed the opposite trend. Exceptions were observed in all species, where some larger chromosomes contained fewer RGAs in B. nigra and S. arvensis, or more RGAs in S. alba. The distribution and density of RGAs across chromosomes were examined. RGA distribution was skewed towards chromosomal ends, with patterns differing across RGA types. Sequence hierarchical pairwise similarity analysis revealed distinct gene clusters, suggesting evolutionary relationships. The study also identified homologous genes among RGAs and non-RGAs in each species, providing insights into disease resistance mechanisms. Finally, RLKs and RLPs were co-localised with reported disease resistance loci in Brassica, indicating significant associations. Phylogenetic analysis of cloned RGAs and QTL-mapped RLKs and RLPs identified distinct clusters, enhancing our understanding of their evolutionary trajectories. These findings provide a comprehensive view of RGA diversity and genomics in these Brassicaceae species, providing valuable insights for future research in plant disease resistance and crop improvement.
Brassica carinata is considered an orphan crop, yet it is vital for understanding the evolution of the triangle of U Brassica species. The availability of a genome reference for this species has allowed for the interrogation of the genomic and genetic underpinnings of important traits, including disease resistance. In this study, we report a comprehensive analysis of resistance gene analogs (RGAs) in the first genome assembly for B. carinata (zd-1). A total of 2570 RGAs were predicted, of which 2020 were transmembrane leucine-rich repeats and 550 were nucleotide-binding site leucine-rich repeats. Gene duplication events affected 65.2% of the RGAs, which were classified as either intergenomic or intragenomic duplications. The contrasting patterns of these gene duplication events between the two subgenomes (B and C) support previous findings indicating the presence of subgenome dominance in this species, a characteristic that is shared with the other allopolyploid Brassicas. Comparative analysis with its diploid progenitors, B. nigra and B.oleracea, revealed conservation of genomic features among these species, while phylogenetic analysis suggests that B. carinata RGAs have undergone extensive expansion. This study is the first to analyze the complete set of RGAs within the B. carinata genome, presenting a comprehensive view of the disease resistance landscape in this species.
Machine learning use in plant phenotyping has grown exponentially. These algorithms empowered the use of image data to measure plant traits rapidly and to predict the effect of genetic and environmental conditions on plant phenotype. However, the lack of interpretability in machine learning models has limited their usefulness in gaining insights into the underlying biological processes that drive plant phenotypes. Explainable AI (XAI) emerges to help understand the 'why' behind machine learning model predictions and allow researchers to investigate the most influential features that lead to prediction, classification or segmentation results. Understanding the mechanisms behind model prediction is also central to sanity-checking models, increasing model reliability and identifying dataset biases that may limit the model's applicability across different conditions. This review introduces the concept of XAI and presents current algorithms, emphasizing their suitability for different data types or machine learning algorithms. The use of XAI to leverage trait information is highlighted, showcasing how recent studies employed model explanations to recognize the features that impact plant phenotype. Overall, this review presents a framework for using XAI to gain insights into intricate biological processes driving plant phenotypes, underscoring the significance of transparency and interpretability in machine learning.
The soybean [Glycine max (L.) Merr.] pangenome has been studied and shown to be an invaluable resource for investigating structural variations (SVs), from which different genomic markers were successfully developed and employed for genome-wide association studies (GWAS). Among the SVs markers, gene presence-and-absence variations (PAVs) have been developed in soybean, but have not been widely utilized for association analyses. Here, we reported GWAS and haplotype analysis of seed protein and oil content for two diverse panels, comprised over 500 soybean accessions evaluated in multiple field environments using three marker datasets, whole genome sequence (WGS)-single-nucleotide polymorphisms (SNPs), 50 K-SNPs, and PAVs. The analyses identified new quantitative trait loci (QTL) for protein and oil content, along with the validation of previously reported QTL for these traits. This includes a well-studied QTL on chromosome (Chr.) 20 and another one on Chr. 05 for protein and/or oil. Importantly, this study is the first to report a new genomic locus for both protein and oil mapped to Chr. 08. Gene ontology annotations and expression profiles suggested candidate genes. Further analyses using haplotype-based markers led to the identification of multiple haplotype blocks encompassing candidate genes. Among these, Glyma.05G243400 on Chr. 05 and Glyma.08G109900 and Glyma.08G110000 on Chr. 08 were identified as promising targets. These genes can be incorporated into soybean breeding programs to enhance the selection of desirable protein and oil phenotypes through a haplotype-based breeding approach.
Global food security depends heavily on a few staple crops, while orphan crops, despite being less studied, offer the potential benefits of environmental adaptation and enhanced nutritional traits, especially in a changing climate. Major crops have benefited from genomics-based breeding, initially using single genomes and later pangenomes. Recent advances in DNA sequencing have enabled pangenome construction for several orphan crops, offering a more comprehensive understanding of genetic diversity. Orphan crop research has now entered the pangenomics era and applying these pangenomes with advanced selection methods and genome editing technologies can transform these neglected species into crops of broader agricultural significance.
Polyploidy, also known as whole-genome duplication (WGD), is a significant evolutionary force in green plants, especially angiosperms. The dynamic nature of polyploid genomes generates genetic diversity and drives the evolution of novel traits and adaptations. Pangenomics is emerging as a major frontier in plant genome research, with a rapidly growing number of pangenomes for individual species and associated analyses providing novel agronomic and evolutionary insights. Polyploid genome analysis can be confounded by intraspecific variation when relying on a single reference genome assembly. The use of pangenomes that better represent the genomic diversity of a species helps overcome this limitation. However, a major gap remains between the number of pangenomic studies in polyploid compared to diploid species, despite the widespread prevalence of WGD, limiting the potential of the pangenome framework for characterizing and understanding polyploid genomes. Furthermore, most polyploid pangenome studies have focused on domesticated crop species, and natural populations have rarely been examined. In addition to applications in crop improvement, pangenomes can provide insights into the ecological and evolutionary impact of polyploidy. Here, we summarize recent pangenome studies in polyploid plants and highlight promising topics for future research. We hope this article will encourage the growth of pangenomic studies in polyploid systems, particularly in natural populations.
Crop diseases pose a major threat to global food security, causing substantial yield losses and economic damage each year. Plant disease epidemiology studies the dynamics of plant-pathogen interactions and their impact on disease outcomes, considering environmental influences at a population level. While recent advances in artificial intelligence (AI) and machine learning (ML) have introduced innovative tools for disease prediction and management, most applications have focused on plant disease detection, classification and severity quantification using imaging technologies and sensor-based data. However, their use in plant disease epidemiology, particularly in understanding host-pathogen interactions and the ecology and evolution of the pathosystems remains limited due to the complexity of multi-scale interactions. In this review, we first propose an updated plant disease epidemiology ‘disease pyramid’ model, incorporating ecological and evolutionary components into the traditional ‘disease triangle’ model. Following this, we discuss current ML applications in plant disease epidemiology, while highlighting both challenges and opportunities. We offer insights into potential input datasets that could significantly enhance the predictability and accuracy of ML models, while also outlining future directions for this rapidly evolving field. The aim of this review is to draw the reader's attention to the knowledge gap in the application of ML in plant disease epidemiology and showcase the vast potential for expanding the scope of more in-depth and comprehensive research in this field in the future.
Predicting phenotypes from a combination of genetic and environmental factors is a grand challenge of modern biology. Slight improvements in this area have the potential to save lives, improve food and fuel security, permit better care of the planet, and create other positive outcomes. In 2022 and 2023 the first open-to-the-public Genomes to Fields (G2F) initiative Genotype by Environment (GxE) prediction competition was held using a large dataset including genomic variation, phenotype and weather measurements and field management notes, gathered by the project over nine years. The competition attracted registrants from around the world with representation from academic, government, industry, and non-profit institutions as well as unaffiliated. These participants came from diverse disciplines include plant science, animal science, breeding, statistics, computational biology and others. Some participants had no formal genetics or plant-related training, and some were just beginning their graduate education. The teams applied varied methods and strategies, providing a wealth of modeling knowledge based on a common dataset. The winners strategy involved two models combining machine learning and traditional breeding tools: one model emphasized environment using features extracted by Random Forest, Ridge Regression and Least-squares, and one focused on genetics. Other high-performing teams methods included quantitative genetics, classical machine learning/deep learning, mechanistic models, and model ensembles. The dataset factors used, such as genetics; weather; and management data, were also diverse, demonstrating that no single model or strategy is far superior to all others within the context of this competition.
Crop disease detection is important due to its significant impact on agricultural productivity and global food security. Traditional disease detection methods often rely on labour-intensive field surveys and manual inspection, which are time-consuming and prone to human error. In recent years, the advent of imaging technologies coupled with machine learning (ML) algorithms has offered a promising solution to this problem, enabling rapid and accurate identification of crop diseases. Previous studies have demonstrated the potential of image-based techniques in detecting various crop diseases, showcasing their ability to capture subtle visual cues indicative of pathogen infection or physiological stress. However, the field is rapidly evolving, with advancements in sensor technology, data analytics and artificial intelligence (AI) algorithms continually expanding the capabilities of these systems. This review paper consolidates the existing literature on image-based crop disease detection using ML, providing a comprehensive overview of cutting-edge techniques and methodologies. Synthesizing findings from diverse studies offers insights into the effectiveness of different imaging platforms, contextual data integration and the applicability of ML algorithms across various crop types and environmental conditions. The importance of this review lies in its ability to bridge the gap between research and practice, offering valuable guidance to researchers and agricultural practitioners.
The tomato (Solanum lycopersicum L.), a principal fruit crop, exhibits significant genetic diversity shaped by domestication and breeding. Analysis of the gene-based super-pangenome, a catalogue of all genes across diverse genome-sequenced tomatoes, has not yet been fully explored. Here, we present a comprehensive analysis of the gene-based super-pangenome across 61 genetically diverse tomato varieties, revealing 59 066 orthologous groups, thereby providing a detailed genetic framework for understanding the evolution of tomatoes. Our phylogenetic analysis recalibrates the position of S. galapagense, challenging existing paradigms of tomato evolution. Identification of genes linked to key agronomic traits such as fruit size, ripening and stress tolerance, along with their presence/absence variation among accessions, offers a rich source of genetic markers for breeding programs. The study also highlights the impact of whole-genome triplication (WGT) and tandem gene duplication (TD) events on gene family expansion, particularly in distant wild relatives. The analysis of the LRR-RLK gene family, important for plant development and defence, reveals substantial sequence diversity and conservation. Rapidly evolving genes and those under positive selection, such as HAI3, CYP711A1/MAX1, WRKY9 and CNGC15, are implicated in stress tolerance and defence mechanisms. The identification of these genes, along with specific pathogenesis-related genes in distant wild relatives, suggests potential strategies to improve fruit shelf life, fruit set and stress tolerance in elite tomato cultivar breeding. Additionally, we have developed the tomatoPangenome platform, integrating genomic and pangenomic data, gene families and tools, to support sustainable production of high-quality, climate-resilient tomatoes and advance selective breeding for future food security.
Brassica species, which include economically important Brassica crops grown around the globe, are important as popular vegetables, forage, and oilseed crops, supplying food for humans and animals. Despite their importance, these crops face increasing challenges from biotic and abiotic stresses, exacerbated by climate change and the evolving threat of crop pathogens. Enhancing crop resilience against these stresses has become a key priority to ensure stable crop production. Recent advancements in genomic studies on Brassica crops and their pathogens have facilitated the deployment of CRISPR/Cas systems in breeding major Brassica crops. This review highlights recent progress in CRISPR/Cas-based gene editing technologies to improve resistance to pathogens and enhance tolerance to drought, salinity, and extreme temperatures. It also summarises the molecular mechanisms underlying crop responses to these stresses. Furthermore, the review discusses the workflow for employing the CRISPR/Cas system to boost stress tolerance and resistance, outlines the associated challenges, and explores prospects based on gene editing research in Brassica species.
There was an error in the original publication [...]
Abstract The surge in high‐throughput technologies has empowered the acquisition of vast genomic datasets, prompting the search for genetic markers and biomarkers relevant to complex traits. However, grappling with the inherent complexities of high dimensionality and sparsity within these datasets poses formidable hurdles. The immense number of features and their potential redundancy demand efficient strategies for extracting pertinent information and identifying significant markers. Feature selection is important in large genomic data as it helps in enhancing interpretability and computational efficiency. This study focuses on addressing these challenges through a comprehensive investigation into genomic feature selection methodologies, employing a rich soybean (Glycine max L. Merr.) dataset comprising 966 lines with over 5.5 million single nucleotide polymorphisms. Emphasizing the “small n large p” dilemma prevalent in contemporary genomic studies, we compared the efficacy of traditional genome‐wide association studies (GWAS) with two prominent machine learning tools, random forest and extreme gradient boosting, in pinpointing predictive features. Utilizing the expansive soybean dataset, we assessed the performance of these methodologies in selecting features that optimize predictive modeling for various phenotypes. By constructing predictive models based on the selected features, we ascertain the comparative prediction accuracies, thereby illuminating the strengths and limitations of these feature selection methodologies in the realm of genomic data analysis.