
Studying the phenotypic diversity of landraces and older advanced cultivars, along with their distribution, is important for breeding and germplasm conservation. To this sense, we analyzed the phenotypic diversity and spreading of 108 durum wheat accessions, including 56 landraces (LANs), 15 Moroccan cultivars (MCs), and 37 North American cultivars (NACs), using thirteen phenotypic descriptors. The collection studied was phenotypically very diverse (Shannon-Weaver index, H'mean=0.66). Spike color (H'mean=0.95), glume hairiness and spike waxiness (H'mean=0.89), peduncle attitude (H'mean=0.82), awn color (H' mean=0.75), and spike density and grain color (H'mean= 0.73) were the most diverse traits. The results also showed that NACs (H'mean=0.78) and the LANs (H'mean=0.69) were more diverse than MCs (H'mean=0.53). Multiple correspondence analyses explained 19.31% of the total variability and demonstrated a clear distinction between genotypes of Morocco and North America and between the LANs and MCs. These results of phenotypic diversity and distribution of studied genotypes can make a major leap forward in national breeding programs and germplasm collection expeditions.
Inositol phosphate kinases (IPKs) play vital roles in the synthesis and regulation of cellular levels of inositol polyphosphates, which perform vital functions in eukaryotic cells as second messengers. Due to the vital biological roles of the inositol phosphates, the kinases involved in their synthesis have gained great attention. The IPK gene family has not been extensively studied in bread wheat (Triticum aestivum). In the present study, we reported a genome-wide identification, phylogenetic analysis and expression patterns of the IPK gene family in wheat. Gene structure, genome distribution, motif conservation, and gene ontology enrichment analysis were carried out systematically. A total of 24 inositol phosphate kinase (IPK) genes were identified in the wheat genome that belonged to 8 homoeologous groups. The TaIPK genes were distributed on chromosomes 1, 2, 3, 4, 5 and 7 in all the three sub-genomes A, B, and D. Based on phylogenetic analysis, these genes were classified into four subfamilies. The subfamilies were defined based on conserved domains, motifs, chromosome locations, and gene structures. Eight pairs of paralogous IPKs were identified based on the phylogenetic relationships among wheat IPKs. Gene ontology enrichment analysis revealed the significant role of IPK genes in different biological and molecular processes in addition to their role in inositol phosphate signaling. At different developmental stages, all the 24 TaIPKs exhibited diverse expression patterns in different tissues showing diversity in their biological functions. The current study identified and explored the properties of 24 IPK genes in the wheat genome at diverse levels, thus providing a strong foundation for researchers to understand the family of these kinases in wheat.
Pinus roxburghii (P. roxburghii) is an important ecological and industrial conifer of the Himalayas. Plant boom is incredibly impacted by way of environmental stresses, in particular high ranges of nitrogen (N) and ozone (O3). This work aims to explore how these factors affect the total protein content and chlorophyll levels of P. roxburghii. Three consortiums were inoculated with two-year-old P. roxburghii seedlings. The stresses of 100 kg N h−1 and 100 ppb O3 were applied for 1 month to study their effect on chlorophyll level and total protein content. To evaluate their potential mitigating effects, the fungal consortium presented promising outcomes for the chosen plant species. The elevated metabolic activities and photosynthesis rate were determined by improved total protein content and high chlorophyll level (p
Plant phenomics is the study of plant growth and performance using multidisciplinary approaches, including high-speed computing, computer vision, sensing, and robotics. Numerous phenotyping platforms, ranging from laboratories to growth chambers, greenhouses, and finally the field, have been established through substantial efforts devoted by the plant phenotyping community. These platforms offer opportunities to measure plant traits at different scales and automation levels, with reasonable throughput and accuracy. Different sensor technologies and data analysis pipelines have been successfully implemented in plant phenomics. However, the available phenotyping platforms are not without limitations, preventing the broader application of plant phenomics in crop breeding, vegetation monitoring, and precision agriculture. The present review summarizes recent advances in various phenotyping platforms in plant phenomics, highlights their limitations, and proposes research directions for future studies in this area.
Object detection and image segmentation are powerful computer vision techniques that have been widely used in plant phenotyping tasks, such as the identification of crop diseases/insects and the measurement/counting of organs (leaves, stems, fruits, etc.). This chapter summarizes recent object detection and image segmentation techniques and public datasets developed for plant phenotyping, discusses their current challenges, particularly focusing on model generalization capability, and offers a perspective on future approaches to make these techniques more helpful to the plant science community.
Genome-wide association studies (GWAS) associate genomic polymorphisms with phenotypes. As the cost of genotyping using next-generation sequencing (NGS) has fallen, GWAS has become a popular technique, taking advantage of population diversity to rapidly identify genomic loci associated with phenotypic traits. This chapter provides a brief introduction to the principles and practicalities of using GWAS, primarily on its use in diploid crop species, and concentrating on GWAS employing mixed linear models (MLM) that accommodate the effects of population structure and genetic relatedness among the study population. Core processes that comprise a GWAS study are reviewed, and some recommendations concerning software and best practices for achieving meaningful results are discussed.
Extensive transcriptome studies of over 40 plant species have shown that very broad regions of plant genomes contain noncoding RNAs (ncRNAs). ncRNAs are known to regulate gene expression, and are involved in DNA methylation, DNA structural modification, histone modification, RNA degradation, RNA masking, and translational regulation. The biological roles of ncRNAs impact plant growth, response to biotic and abiotic stresses, and regulate inherited properties of gene expression. Therefore, it is important to elucidate the ncRNA regulatory mechanisms in plants, including their expression and the interaction of their targets such as RNAs, DNAs, and proteins. In this review, recent research pertaining to the biogenesis and functions of ncRNAs is summarized and discussed.
Nowadays high-throughput sequencing and computational techniques are indispensable tools to interpret the regulations between transcription factors (TFs) and their downstream target genes. Also, plenty of analysis tools and online resources have been developed to assist the construction of gene regulatory networks (GRNs). In this chapter, we illustrate the wet-lab and dry-lab approaches used in the construction of GRNs, online databases for plant TFs and cis-elements, and advanced analysis of GRNs. At the end of this chapter, we discuss the potential aspects of studies of TFs and cis-elements. Overall, this chapter summarizes the basic framework and public resources for the construction of GRNs in plants.
In the past two decades, technological advancements in DNA sequencing and mass spectrometry-based protein sequencing have made large-scale genome, transcriptome, and proteome analyses accessible to individual researchers. Consequently, many omics databases have been constructed for various organisms, which has facilitated a diversity of studies on genetics, functional genomics, molecular biology, and other research areas. In this chapter, we introduce key plant omics databases together with some useful online analysis tools that enable users without any bioinformatics expertise to access omics data. In the early part of this chapter, we have focused on databases for the model plant Arabidopsis, in concurrence with a summary of the related milestone studies for genome, epigenome, transcriptome, and proteome. Subsequently, we introduce omics databases for major crops, including rice, wheat, maize, soybean, tomato, and pepper, which have contributed not only to functional genomics, but also to agronomic studies. We also introduce omics databases for bryophytes, a basal lineage of land plants, which have contributed to plant evolutionary studies. Moreover, we introduce plant omics portals that facilitate the performance of integrated omics data analysis at a species-wide level. Thus, this chapter provides a useful guide for researchers to obtain and exploit ample online resources to promote their own studies.
In this post-genomic era, we now have easy access to the genetic information of entire living organisms, and that information has been essential for biological research. The prosperity of genomics resulted from the progress of DNA sequence technologies, the development of computational analysis environments, and the establishment of biological resources. Plant genomics is one of the research fields that has strongly benefited from these technical advances. This chapter presents the evolution and transition of DNA sequence technologies and gives concrete examples of the proliferation in plant genomic research.
Genome editing is an innovative technology that has brought about a revolution in the field of genetic engineering. In particular, the emergence of clustered regularly interspaced short palindromic repeat (CRISPR)-CRISPR-associated protein 9 (Cas9) has created a breakthrough in genetic engineering by enabling easy, efficient, and precise genome manipulation. In this chapter, we introduce recent advances in plant genome editing technologies that include precise genome editing by gene targeting, base editing, and a new strategy called prime editing. The exploration of new CRISPR-Cas tools has been continuously reported, resulting in successful applications of engineered Cas9 proteins or newly identified CRISPR-Cas systems such as Cas9-NG, CRISPR-CasF, and CRISPR-Cas type I-D systems. The applications of CRISPR-Cas are already not limited to genome editing but have also expanded to other purposes, such as RNA editing by CRISPR-Cas13, transcriptional control, and epigenetic modification by CRISPR-dCas9 fused with effector proteins. The toolbox will be updated further in the near future, bringing new approaches to achieve precise genome editing and genome manipulation in plants. These developments will strongly facilitate the functional analysis of plant genes and contribute to plant breeding.
The advent of proteomic techniques has made it possible to identify a broad spectrum of proteins in living systems. Its capability is especially useful for crop breeding as it gives clues, not only about nutritional value, but also about yield, and how these factors are affected by adverse conditions. This chapter describes recent progress in plant proteomics and highlights the achievements made in understanding the proteins of major soybean under flooding stress. The main emphasis will be on crop responses to abiotic and biotic stresses. Rigorous genetic testing of the role of possibly important proteins can be conducted. Currently, a massive amount of research output in DNA, mRNA, and protein levels is available, and it is suggested that now the proteome is a key data layer to be applied in practical crop breeding.
Recent progress of deep learning (DL) frameworks has allowed various high-quality classifications and regressions to be developed. Their performances have often exceeded human standards, especially in image diagnoses, so we may be able to reproduce artificial professional eyes for various objectives. Furthermore, recent development of explainable DL technology (or explainable AI (X-AI)), which can visualize the relevance of DL predictions, would allow biological interpretations of the predictions, potentially providing insights that only DL models may be able to recognize. Nevertheless, the application of DL frameworks has still progressed less in plant science than in animal, medical, and social sciences. In this chapter, mainly focused on classification and regression analyses, we introduce the current applications and potential prospects of DL technologies with respect to plant images and genetic sequences. Here, we take as examples taxonomic classification, disease/stress diagnosis, non-invasive prediction, and implementation with automated sorting systems based on plant images, as well as prediction of gene functions, such as expression patterns and protein folding structures, based on DNA or amino acid sequences. For beginners in the use of DL techniques, tips and precautions regarding practical application of DL frameworks are also provided.
In recent years, there have been growing opportunities to apply artificial intelligence to experimental research data. Artificial intelligence, often abbreviated as AI, is a field of computer science in the development of software algorithms that mimics human intellectual behaviors. Machine learning is a subset of artificial intelligence to facilitate data-driven numerical learning strategy, and deep learning is, again, a subset of machine learning. There are multiple algorithms in deep learning, and the selection of deep learning algorithms is a key problem for applying them to plant omics. The optimality of the algorithm depends on task design and input modality, while the plant omics can be broken down into multiple omics layers of genomics, transcriptomics, proteomics, metabolomics, phenomics, and so on. In this chapter, major deep learning approaches will be introduced with their applications in life science, particularly in plant omics.
The expression profile of a given gene represents the changes in expression levels and it can be established from a variety of experiments including different tissues, organs, developmental stages, and treatments. By using the wealth of publicly available RNA-seq data, we can obtain an overview of genome-wide gene expression profiles under a wide range of experimental conditions. Such gene expression profiles are particularly useful to mine a gene set having a similar expression profile. The similarity indicates that these genes may be involved in the same biological process and/or be regulated by the same transcription factor. For large-scale expression data, the gene expression network (GEN) analysis allows evaluation of the significance of the similarity in expression profiles among genes in a genome-wide manner. In order to efficiently predict the biological functions and regulatory mechanisms of expression for individual genes in GENs, we can assign various kinds of functional descriptions in GENs. This chapter introduces the types of GENs and the methods used to construct them, in addition to the database built from large-scale analysis of public datasets and the integration of existing annotation of the literature (e.g., knowledge-based) to GENs.
Epigenetic regulation is highly conserved in eukaryotes. In general, events mediated by the genome - such as the regulation of gene activity and establishment and arrangement of chromatin structure - are coordinated by epigenetic and chromatin-remodeling factors. The concept of epigenetics was coined by Waddington (1957). In the context of embryogenesis, cells are affected by particular circumstances during development. This idea involves the possibility of "epi" factors that determine cell fate ("epi" is a Greek prefix meaning "external action or effect"). Currently, the definition of epigenetics remains interpretive in various contexts; however, the concept is based on "heritable phenotype changes that do not involve alterations in DNA sequence" and involves "nonheritable chromatin-level regulation that does not involve alterations in DNA sequence." In this chapter, the basic concepts of genome-wide regulation of histone modifications and DNA methylation - known as epigenetic and chromatin regulatory factors - are the primary focus, with an emphasis on understanding the fundamental mechanisms of epigenetic regulation.
Experimental resources of various plant species are provided from the core facilities of the National BioResource Project (NBRP) of Japan. Among them, Arabidopsis thaliana is a well-known model plant widely used in the international research community. Arabidopsis resources are maintained and distributed through collaboration with centers in the USA and UK. The project also preserves and distributes biological materials of crop species that are indispensable for human beings, such as rice, wheat, tomato, and legumes. The aim of the NBRP is to promote life science research by providing resources of the highest quality level to ensure the reproducibility of research. The NBRP contributes to the promotion of plant science, including omics studies.
Plant cells contain a variety of organelles, whose functions support many biological processes. One of the characteristics of plant organelles is that they can dynamically change their functions in response to developmental stages, environmental changes, and external stimuli. To elucidate the molecular mechanisms of organelle function, omics analyses have been conducted at various levels, including genomics, transcriptomics, and proteomics. This chapter outlines current omics approaches for membrane-bound organelles. In recent years, imaging techniques have become indispensable in life science research, and in contributing to the visualization of organelle dynamics. This chapter also introduces useful databases on plant organelles.
Under the new paradigms of integrative and network biology, comparison of expression changes among a small subset of candidate genes across phenotypic variants is hardly informative or conclusive in the context of cellular response regulation and the underlying genetic mechanisms. Integration of global changes in gene expression under multiple conditions with existing genomic databases that have been curated systematically are key for the efficient extraction of robust and biologically meaningful patterns and signatures that are reflective of cellular states. Despite the increasing availability of a wide array of computational tools, the resolution of RNA-seq-based transcriptome profiling is as good as the experimental design that determines the window of information revealed relative to the hypothesis being tested, and this intricacy is often underestimated. In this chapter, we discuss the important aspects of data analytics and the basic principles that must be taken into consideration to better bridge the design of the wet-lab experiments with the requirements of a robust dry-lab knowledge dissection and integration. We also highlight the unique assumptions and requirements between transcriptome experiments conducted using plant genetic models with comprehensive and annotated genomes for reference-guided assembly and extraction of biological knowledge, in comparison with the non-model plant species, which rely on a de novo assembly of transcriptome datasets followed by homology-based comparison with closely related species with reference genome.
Metabolomics aims to analyze the so-called metabolome, which is composed of all the low-molecular-weight compounds present in biological systems. Approximately 1 million metabolites are estimated to exist in the plant kingdom. The remarkable development of a range of analytical platforms enables the separation, detection, characterization, and (semi-)quantification of primary and secondary metabolites in plants. Furthermore, spatial metabolomics techniques such as mass spectrometry and magnetic resonance imaging have shown great progress in visualizing the localization of metabolites in plant tissues. Based on the great contribution of the metabolomics community, guidelines for metabolite identification confidence and the collection of metabolomics metadata are proposed. Here, the analytical targets and techniques in plant metabolomics are summarized. Furthermore, the progress made in generating "reliable" metabolomics data for big data biology is introduced. In conclusion, metabolomics data analyses will boost the progress of big data biology, though the validation of data reliability is important. When researchers upload metabolomics data, minimum reporting standards as proposed by the Metabolomics Standards Initiative should be provided along with metabolite identification confidence. More data should be shared worldwide to foster the development of computational approaches for outcome interpretation. Finally, the integration of other omics with metabolomics data and the synthesis on a system level are crucial steps toward linking genotype-metabotype-phenotype relationships and implementing "metabolic editing" in plants.