Crop evolutionary history and domestication processes are key issues for better conservation and effective use of crop genetic diversity. Black and white fonio (Digitaria iburua and D. exilis, respectively) are two small indigenous grain cereals grown in West Africa. The relationship between these two cultivated crops and wild Digitaria species is still unclear. Here, we analyse whole genome sequences of 265 accessions comprising these two cultivated species and their close wild relatives. We show that white and black fonio were the result of two independent domestications without gene flow. We infer a cultivation expansion that began at the outset of the CE era, coinciding with the earliest discovered archaeological fonio remains in Nigeria. Fonio population sizes declined a few centuries ago, probably due to a combination of several factors, including major social and agricultural changes, intensification of the slave trade and the introduction of new, less labour-intensive crops. The key knowledge and genomic resources outlined here will help to promote and conserve these neglected climate-resilient crops and thereby provide an opportunity to tailor agriculture to the changing world.
To what extent overdominance might contribute to the maintenance of genetic diversity within genomes is still an ongoing research question. Pseudo-overdominance created by the complementation of deleterious alleles in heterozygotes has recently become a subject of particular interest. Simulations and theory suggest that pseudo-overdominance may occur in low recombining regions. Here, we conduct a comprehensive investigation of large low-recombining (LLR) regions in cultivated populations of pearl millet, an outcrossing diploid African cereal. We examine seven large regions ranging from 5 to 88 Mb and six of them are pericentromeric. These LLR regions exhibit an excess of heterozygotes, a distinctive hallmark of overdominance. They display a tendency toward a higher diversity and a larger ratio of non-synonymous and deleterious variants. We conduct a more in-depth study of the largest 88 Mb region, identified on chromosome 3. Interestingly, haplotypes of this region have been introgressed from wild relatives. Using long read sequencing, we confirm their strong divergence and the presence of inversions across one of them. One of the haplotypes seems to be highly deleterious in the homozygous state. A total of 17% of the cultivated pearl millet genome exhibit a local population structure suggestive of overdominance or possibly pseudo-overdominance. Our empirical results contribute to the accumulation of knowledge, which will enhance our understanding of the potential role of overdominance or pseudo-overdominance in maintaining genetic diversity, particularly in low recombining regions.
Backgrounds and Aims In palms, many dioecious species have emerged from at least eight independent events; the mechanisms of sex determination remain poorly understood. Here, we identify and compare the sex chromosomes of Kerriodoxa elegans with those of the well-studied date palm (Phoenix dactylifera), which evolved dioecy independently from a monoclinous common ancestor.Methods We developed target sequence capture kits and inferred sex-linked genes using a probabilistic approach in both species.Key Results We find a striking similarity between the sex-linked regions of K. elegans and P. dactylifera, with the majority of sex-linked genes being common between the two species. However, we confirm that these regions evolved independently, much later than the split between the lineages.Conclusions This case of convergent evolution seems to be unique in plants so far, and raises questions on the mechanisms of sex determination. This could be explained by the presence of genes related to floral sex development and sex determination in this region, which have been recruited during the evolution of sex chromosomes, even though the genes involved may differ between the two species.
During herbivory, chewing insects deposit complex oral secretions (OS) onto the plant wound. Understanding how plants respond to the different cues of herbivory remains an active area of research. In this study, we used an herbivory-mimick experiment to investigate the early transcriptional response of rice plants leaves to wounding, OS, and OS microbiota from Spodoptera frugiperda larvae. Wounding induced a massive early response associated to hormones such as jasmonates. This response switched drastically upon OS treatment indicating the activation of OS specific pathways. When comparing native and dysbiotic OS treatments, we observed few gene regulation. This suggests that in addition to wounding the early response in rice is mainly driven by the insect compounds of the OS rather than microbial. However, microbiota affected genes encoding key phytohormone synthesis enzymes, suggesting an additional modulation of plant response by OS microbiota.
Seedling root traits impact plant establishment under challenging environments. Pearl millet is one of the most heat and drought tolerant cereal crops that provides a vital food source across the sub-Saharan Sahel region. Pearl millet’s early root system features a single fast-growing primary root which we hypothesize is an adaptation to the Sahelian climate. Using crop modeling, we demonstrate that early drought stress is an important constraint in agrosystems in the Sahel where pearl millet was domesticated. Furthermore, we show that increased pearl millet primary root growth is correlated with increased early water stress tolerance in field conditions. Genetics including genome-wide association study and quantitative trait loci (QTL) approaches identify genomic regions controlling this key root trait. Combining gene expression data, re-sequencing and re-annotation of one of these genomic regions identified a glutaredoxin-encoding gene PgGRXC9 as the candidate stress resilience root growth regulator. Functional characterization of its closest Arabidopsis homolog AtROXY19 revealed a novel role for this glutaredoxin (GRX) gene clade in regulating cell elongation. In summary, our study suggests a conserved function for GRX genes in conferring root cell elongation and enhancing resilience of pearl millet to its Sahelian environment.
AbstractPaspalum notatum Flüggé is an economically important subtropical fodder grass that is widely used in the Americas. Here, we report a new chromosome-scale genome assembly and annotation of a diploid biotype collected in the center of origin of the species. Using Oxford Nanopore long reads, we generated a 557.81 Mb genome assembly (N50 = 56.1 Mb) with high gene completeness (BUSCO = 98.73%). Genome annotation identified 320 Mb (57.86%) of repetitive elements and 45,074 gene models, of which 36,079 have a high level of confidence. Further characterisation included the identification of 59 miRNA precursors together with their putative targets. The present work provides a comprehensive genomic resource for P. notatum improvement and a reference frame for functional and evolutionary research within the genus.
Pearl millet (Pennisetum glaucum (L.)) R. Br. syn. Cenchrus americanus (L.) Morrone) is an important crop in South Asia and sub-Saharan Africa which contributes to ensuring food security. Its genome has an estimated size of 1.76 Gb and displays a high level of repetitiveness above 80%. A first assembly was previously obtained for the Tift 23D2B1-P1-P5 cultivar genotype using short-read sequencing technologies. This assembly is, however, incomplete and fragmented with around 200 Mb unplaced on chromosomes. We report here an improved quality assembly of the pearl millet Tift 23D2B1-P1-P5 cultivar genotype obtained with an approach combining Oxford Nanopore long reads and Bionano Genomics optical maps. This strategy allowed us to add around 200 Mb at the chromosome-level assembly. Moreover, we strongly improved continuity in the order of the contigs and scaffolds within the chromosomes, particularly in the centromeric regions. Notably, we added more than 100 Mb around the centromeric region on chromosome 7. This new assembly also displayed a higher gene completeness with a complete BUSCO score of 98.4% using the Poales database. This more complete and higher quality assembly of the Tift 23D2B1-P1-P5 genotype now available to the community will help in the development of research on the role of structural variants and more broadly in genomics studies and the breeding of pearl millet.
Background Developing high yielding varieties is a major challenge for breeders tackling the challenges of climate change in agriculture. The panicle (inflorescence) architecture of rice is one of the key components of yield potential and displays high inter- and intra-specific variability. The genus Oryza features two different crop species: Asian rice ( Oryza sativa L.) and the African rice ( O. glaberrima Steud). One of the main morphological differences between the two independently domesticated species is the structure (or complexity) of the panicle, with O. sativa displaying a highly branched panicle, which in turn produces a larger number of grains than that of O. glaberrima . The genetic interactions that govern the diversity of panicle complexity within and between the two species are still poorly understood. Results To identify genetic factors linked to panicle architecture diversity in the two species, we used a set of 60 Chromosome Segment Substitution Lines (CSSLs) issued from third generation backcross (BC 3 DH) and carrying genomic segments from O. glaberrima cv. MG12 in the genetic background of O. sativa Tropical Japonica cv. Caiapó. Phenotypic data were collected for rachis and primary branch length, primary, secondary and tertiary branch number and spikelet number. A total of 15 QTLs were localized on chromosomes 1, 2, 3, 7, 11 and 12 and QTLs associated with enhanced secondary and tertiary branch numbers were detected in two CSSLs. Furthermore, BC 4 F 3:5 lines carrying different combinations of substituted segments were produced to decipher the effects of the identified QTL regions on variations in panicle architecture. A detailed analysis of phenotypes versus genotypes was carried out between the two parental genomes within these regions in order to understand how O. glaberrima introgression events may lead to alterations in panicle traits. Conclusion Our analysis led to the detection of genomic variations between O. sativa cv. Caiapó and O. glaberrima cv. MG12 in regions associated with enhanced panicle traits in specific CSSLs. These regions contain a number of key genes that regulate panicle development in O. sativa and their interspecific genomic variations may explain the phenotypic effects observed.
Using long reads provides higher contiguity and better genome assemblies. However, producing such high quality sequences from raw reads requires to chain a growing set of tools, and determining the best workflow is a complex task. To tackle this challenge, we developed CulebrONT, an open-source, scalable, modular and traceable Snakemake pipeline for assembling long reads data. CulebrONT enables to perform tests on multiple samples and multiple long reads assemblers in parallel, and can optionally perform, downstream circularization and polishing. It further provides a range of assembly quality metrics summarized in a final user-friendly report. CulebrONT alleviates the difficulties of assembly pipelines development, and allow users to identify the best assembly options.
Abstract Discovered in the 1960s, Meloidogyne graminicola is a root‐knot nematode species considered as a major threat to rice production. Yet, its origin, genomic structure, and intraspecific diversity are poorly understood. So far, such studies have been limited by the unavailability of a sufficiently complete and well‐assembled genome. In this study, using a combination of Oxford Nanopore Technologies and Illumina sequencing data, we generated a highly contiguous reference genome (283 scaffolds with an N50 length of 294 kb, totaling 41.5 Mb). The completeness scores of our assembly are among the highest currently published for Meloidogyne genomes. We predicted 10,284 protein‐coding genes spanning 75.5% of the genome. Among them, 67 are identified as possibly originating from horizontal gene transfers (mostly from bacteria), which supposedly contribute to nematode infection, nutrient processing, and plant defense manipulation. Besides, we detected 575 canonical transposable elements (TEs) belonging to seven orders and spanning 2.61% of the genome. These TEs might promote genomic plasticity putatively related to the evolution of M. graminicola parasitism. This high‐quality genome assembly constitutes a major improvement regarding previously available versions and represents a valuable molecular resource for future phylogenomic studies of Meloidogyne species. In particular, this will foster comparative genomic studies to trace back the evolutionary history of M. graminicola and its closest relatives.
South Green is a bioinformatics platform dedicated to the genetics and genomics of tropical and Mediterranean plants of agronomic interest and their related pathogens. It federates a network of bioinformaticians belonging to different units and institutes of Montpellier (Alliance Bioversity CIAT, CIRAD, INRAE and IRD) with a multidisciplinary expertise ranging from the data integration, bioinformatics software development, sequencing data analyses and high-performance computing. Exchange and collaborative developments are fostered through regular hands-on sessions on synergistic themes (information systems, pangenomic methods, graphical visualizations and workflows managers). The South Green web portal (www.southgreen.fr) gathers all the information systems and tools developed and supported by the platform. Indeed, South Green ensures the development of original information systems such as GreenPhyl, SNiPlay, Gigwa, AgroLD or the Genome Hubs, and offers sequencing data analysis pipelines through two workflow managers: Galaxy and TOGGLe. Overall, South Green's mission is to promote these original tools as well as their interoperability. A significant part of activities also comprises hands-on trainings that are regularly offered in the local community as well as with partners in the Africa and Asia on the following topics: Galaxy, NGS analyses, R, Perl, Linux, HPC administration (southgreenplatform.github.io/trainings/). Besides, South Green provides access to computing facilities for both users and developers engaged in this scientific area. South Green is part of the network of platforms of the French Institute of Bioinformatics (IFB).
The African bonytongue (Heterotis niloticus) is an excellent candidate for fish farming because it has outstanding biological characteristics and zootechnical performances. However, the absence of sexual dimorphism does not favor its reproduction in captivity or the understanding of its reproductive behavior. Moreover, no molecular data related to its reproduction is yet available. This study therefore focuses on the structural identification of the different molecular actors of vitellogenesis expressed in the pituitary gland, the liver and the ovary of H. niloticus. A transcriptomic approach based on de novo RNA sequencing of the pituitary gland, ovary and liver of females in vitellogenesis led to the creation of three transcriptomes. In silico analysis of these transcriptomes identified the sequences of pituitary hormones such as prolactin (PRL), luteinizing hormone (LH) and follicle-stimulating hormone (FSH) and their ovarian receptors (PRLR, FSHR, LHR). In the liver and ovary, estrogen receptors (ER) beta and gamma, liver vitellogenins (VtgB and VtgC) and their ovarian receptors (VLDLR) were identified. Finally, the partial transcript of an ovarian Vtg weakly expressed compared to hepatic Vtg was identified based on structural criteria. Moreover, a proteomic approach carried out from mucus revealed the presence of one Vtg exclusively in females in vitellogenesis. In this teleost fish that does not exhibit sexual dimorphism, mucus Vtg could be used as a sexing biomarker based on a non-invasive technique compatible with the implementation of experimental protocols in vivo.
ABSTRACTThe advent of NGS has intensified the need for robust pipelines to perform high-performance automated analyses. The required softwares depend on the sequencing method used to produce raw data (e.g. Whole genome sequencing, Genotyping By Sequencing, RNASeq) as well as the kind of analyses to carry on (GWAS, population structure, differential expression). These tools have to be generic and scalable, and should meet the biologists needs.Here, we present the new version of TOGGLe (Toolbox for Generic NGS Analyses), a simple and highly flexible framework to easily and quickly generate pipelines for large-scale second- and third-generation sequencing analyses, including multi-sample and multi-threading support. TOGGLe is a workflow manager designed to be as effortless as possible to use for biologists, so the focus can remain on the analyses. Pipelines are easily customizable and supported analyses are reproducible and shareable. TOGGLe is designed as a generic, adaptable and fast evolutive solution, and has been tested and used in large-scale projects on various organisms. It is freely available at http://toggle.southgreen.fr/, under the GNU GPLv3/CeCill-C licenses) and can be deployed onto HPC clusters as well as on local machines.
South Green (www.southgreen.fr) est une plateforme de bio-informatique dediee `a la genetique et la genomique des plantes tropicales et mediterraneennes d'interet agronomique et de leurs pathogenes. Elle federe un reseau de bio informaticiens appartenant a differentes unites et instituts de Montpellier (Bioversity International, CIRAD, INRA et IRD) soit environ une vingtaine de personnes en interaction avec les equipes de recherches, avec une expertise multidisciplinaire allant de l'integration de donnees et de connaissance au developpement de logiciels en bio-informatique, a l'analyse de donnees de sequencage (detection de polymorphismes et variants structuraux, pangenomique, metagenomique, analyse differentielle de donnees RNAseq) et le calcul haute performance. Le plateforme South Green a pour objectifs de : Promouvoir des outils originaux issus de la recherche methodologique. Promouvoir l'interoperabilite des applications developpees au sein du reseau. Centraliser l'ensemble des logiciels et systemes d'information developpees au sein d'un portail Web unique (http://www.southgreen.fr) Promouvoir les echanges et les developpements collaboratifs Proposer des formations en bio-informatique, bio analyse de donnees et `a l'utilisation de clusters de calcul Promouvoir la demarche qualite au sein du reseau. Proposer un support pour le calcul a haute performance La plateforme assure le developpement de systemes d'informations et d'outils innovants, necessaires aux projets scientifiques, realises au sein de la plateforme et en lien avec l'analyse des donnees produites par les technologies de sequencage a haut debit (annotation des genomes et de transcriptomes, phylogenie, genotypage) tels que GreenPhylDB, SNiPlay, Gigwa ou AgroLD. Elle propose egalement des pipelines d'analyses de donnees de sequencage au travers de deux gestionnaires de workflows : Galaxy et TOGGLe. Enfin, impliquee dans plusieurs projets de sequencage international, elle possede une forte expertise en developpement de ”genome hub” qu'elle a deploye sur de nombreuses plantes (bananier, cafeier, manioc, cacaoyer) au niveau duquel sont disponibles de nombreuses applications utiles pour l'etude de ces genomes. La plateforme assure aussi des formations specialisees en bio-informatique au niveau national et international (analyse des donnees de sequencage haut debit, Galaxy) et informatique (logiciel R, Perl, Linux). Les ressources sont disponibles sur le site https://southgreenplatform.github.io/trainings/. South Green s'inscrit dans le reseau des plateformes de l'Institut Francais de Bio-informatique (IFB) et fait partie du reseau Renabi (Reseau national des plates-formes bio-informatiques).
Dear biologist, have you ever dreamed of using the whole power of those numerous NGS tools that your bioinformatician colleagues uses through this awful list of command lines ? Dear bioinformatician, have you ever wished for a really quick way of designing a new NGS pipeline without having to retype again dozens of code lines to readapt your scripts or starting from scratch ? So, be happy ! TOGGLE is for you ! With TOGGLE (TOolbox for Generic nGs anaLysEs), you can create your own pipeline through an easy and user-friendly approach. Indeed, TOGGLE integrate a large set of NGS softwares and utilities to easily design pipelines able to handle hundreds of samples. The pipelines can start from Fastq (plain or compressed), SAM, BAM or VCF (plain or compressed) files, with parallel (by sample) and global analyses (by multi samples). Moreover, TOGGLE offers an easy way to configure your pipeline with a single configuration file: – Organizing the different steps of workflow, – Setting the parameters for the different softwares, – Managing storage space through compressing/deleting intermediate data, – Determining the way the jobs are managed (serial or parallel jobs through scheduler SGE, SLURM and LSF). TOGGLE can work on your laptop, on a single machine server as well as on a HPC system, as a local instance or in a Docker machine. The only limit will be your available space on the storage system, not the amount of samples to be treated or the number of steps. TOGGLE was used on different organisms, from a single sample to more than one hundred at a time, in RNAseq, DNAreseq/SNP discovery and GBS analyses. List of bioinformatics tools included: – Cleaning and Quality checking : FastQC, CutAdapt, fastxTrimmer – Assembly : TransAbyss, Trinity, TGI-CL – Mapping : BWA (aln/sampe and mem), TopHat – SAM/BAM management : picardTools, SAMtools, GATK – SNP calling/cleaning/annotation : SAMtools, GATK, VarScan, snpEff – ReadCount : HTseq-count – Structural Variations : BreakDancer, Pindel Web site: https://github.com/SouthGreenPlatform/TOGGLE (Texte integral)
BACKGROUND:The explosion of NGS (Next Generation Sequencing) sequence data requires a huge effort in Bioinformatics methods and analyses. The creation of dedicated, robust and reliable pipelines able to handle dozens of samples from raw FASTQ data to relevant biological data is a time-consuming task in all projects relying on NGS. To address this, we created a generic and modular toolbox for developing such pipelines.RESULTS:TOGGLE (TOolbox for Generic nGs anaLysEs) is a suite of tools able to design pipelines that manage large sets of NGS softwares and utilities. Moreover, TOGGLE offers an easy way to manipulate the various options of the different softwares through the pipelines in using a single basic configuration file, which can be changed for each assay without having to change the code itself. We also describe one implementation of TOGGLE in a complete analysis pipeline designed for SNP discovery for large sets of genomic data, ready to use in different environments (from a single machine to HPC clusters).CONCLUSION:TOGGLE speeds up the creation of robust pipelines with reliable log tracking and data flow, for a large range of analyses. Moreover, it enables Biologists to concentrate on the biological relevance of results, and change the experimental conditions easily. The whole code and test data are available at https://github.com/SouthGreenPlatform/TOGGLE .
We present here the first curated collection of wild and cultivated African rice species. For that, we designed specific SNPs and were able to structure these very low diverse species.