Advances in high-fidelity long-read (HiFi-LR) sequencing technologies have opened new opportunities to explore the microbial genomic diversity of complex environments, such as soils. While short-read (SR) sequencing has enabled broad insights at the gene level, the limited read length constrains the reconstruction of complete metagenome-assembled genomes (MAGs). HiFi-LR, in contrast, improves assembly continuity and completeness, supporting higher-resolution taxonomic and functional annotation. However, the cost and relatively low throughput of HiFi-LR sequencing can limit genome recovery, particularly at the binning stage, where coverage depth is critical. In this study, we assess the benefit of combining HiFi-LR and SR sequencing for genome-resolved characterization of a soil microbiome. We generated metagenomic data for a tunnel-cultivated soil sample using high-coverage Illumina SRs as well as a combination of two HiFi-LR sequencing platforms (PacBio Sequel II and PacBio Revio). We found that assemblies generated from pooled HiFi-LRs data alone exhibited higher completeness than those generated from ultra-deep SR data. Incorporating SR-derived coverage information into the binning of HiFi-LR contigs increased MAG recovery by 24
BACKGROUND:There is considerable interest in the high-throughput discovery and genotyping of single nucleotide polymorphisms (SNPs) to accelerate genetic mapping and enable association studies. This study provides an assessment of EST-derived and resequencing-derived SNP quality in maritime pine (Pinus pinaster Ait.), a conifer characterized by a huge genome size ( approximately 23.8 Gb/C).METHODOLOGY/PRINCIPAL FINDINGS:A 384-SNPs GoldenGate genotyping array was built from i/ 184 SNPs originally detected in a set of 40 re-sequenced candidate genes (in vitro SNPs), chosen on the basis of functionality scores, presence of neighboring polymorphisms, minor allele frequencies and linkage disequilibrium and ii/ 200 SNPs screened from ESTs (in silico SNPs) selected based on the number of ESTs used for SNP detection, the SNP minor allele frequency and the quality of SNP flanking sequences. The global success rate of the assay was 66.9%, and a conversion rate (considering only polymorphic SNPs) of 51% was achieved. In vitro SNPs showed significantly higher genotyping-success and conversion rates than in silico SNPs (+11.5% and +18.5%, respectively). The reproducibility was 100%, and the genotyping error rate very low (0.54%, dropping down to 0.06% when removing four SNPs showing elevated error rates).CONCLUSIONS/SIGNIFICANCE:This study demonstrates that ESTs provide a resource for SNP identification in non-model species, which do not require any additional bench work and little bio-informatics analysis. However, the time and cost benefits of in silico SNPs are counterbalanced by a lower conversion rate than in vitro SNPs. This drawback is acceptable for population-based experiments, but could be dramatic in experiments involving samples from narrow genetic backgrounds. In addition, we showed that both the visual inspection of genotyping clusters and the estimation of a per SNP error rate should help identify markers that are not suitable to the GoldenGate technology in species characterized by a large and complex genome.
De novo OTU picking from large metabarcoding read datasets is at the same time a current and a complex task, and several methods coexist to perform it. We present here the outcome of a collective project developed within Working Group on « Data Analysis and Storage » in DNAqua.net. Our aim has been to organize a thorough comparison of OTU composition according to some selected methods called by the wrappers, in a diversity of situations. This has been done by disposing of a set of different datasets, and a set of different methods, applying each method on each dataset, and comparing the results. We have deliberately chosen to work with cleaned datasets only, and not to include cleaning in the process. We have worked with a set of about 60 different datasets, some environmental, some as mock communities, produced by six teams, in different countries (D, F, I, T, UK), each with specific markers for different organisms. All datasets have been cleaned beforehand by the team proposing it. We have installed four different tools for building OTUs by unsupervised clustering : Swarm (Mahé et al. 2015), Vsearch (Rognes et al. 2016) with the same receipe for all datasets, usearch (Edgar 2010) with a unique command, the same for all datasets, and yapotu, which computes pairwise Smith-Waterman distances between all reads of a given dataset, and then clusters them with graph based techniques. Yapotu approach is expected to be the most accurate one, as there are no heuristics in the calculations. We have harmonized common input/output format for the four methods, to make comparisons. Here is a summary of the indicators selected for comparing results. We have first computed basic indicators per sample and method, like the number of OTU, the number of singletons, the number of OTUs with ten reads or more (after dereplication), and the fraction of reads that have been allocated to an OTU. The four methods displayed a great variety of counts, with highest number of OTUs and singletons for Swarm, then slighltly equivalent figures (but a smaller number of singletons) for yapotu, and significantly smaller counts for Vsearch and Usearch. However, the counts for the number of OTUs with 10 reads or more are much more convergent between the four methods. We have then compared rank-size curves, which have been computed for all pairs (sample by method). Here again, yapotu and swarm results are very similar, whereas Vsearch and Usearch sometimes are close to the former pattern, sometimes very different (I attach a figure?) We then have computed 10 different diversity indices, like OTU richness, Shannon, Chao, eveness. Here again, results provided by Swarm and Yapotu are very similar, with very strong correlations between indices over all samples by method, whereas the correlations with Vsearch and Usearch is very poor. Finally, we have computed all contingency tables (in a sparse format) between all pairs of methods (hence, 6 pairs) for all samples, which accurately describe whether OTUs composition are similar or dissimilar between methods. We have observed that swarm OTUs are systematically nested within yapotu OTUs, and most often, there is a one to one correspondence between a Swarm and a Yapotu OTU. As a conclusion, we show that - Swarm and yapotu yield very similar results including for fine details, the only diffrence being a larger number of singletons provided by Swarm ; - This shows that Swarm OTUs are very close to OTUs built by single linkage Clustering on Smith-Waterman pairwise distances, and consolidates these approaches. - Very often, both Vsearch and Usearch diverge from those convergent results, but not always, and it is not easy to understand when and why. Some further investigations are needed therefore. All datasets will be publicly available for further benchmarking of a wider set of methods and datasets.
Des effets indésirables receveurs (EIR), peuvent apparaître après transfusion de concentrés plaquettaires (CP) chez des patients. Le temps de stockage des CP peut influencer leurs apparitions. Nous avons comparé le transcriptome plaquettaire de CP stockés pendant 2 jours (j2) et pendant 4 jours (j4), transfusés et ayant induit un EIR. Le but étant de mettre en évidence les gènes potentiellement impliqués dans l'apparition d'EIR. Nous avons comparé 2 EIR/témoins (TCP) pour J2 et 2 EIR/TCP pour j4. Les données transcriptomiques ont été générées par séquençage via Ion Proton™. Les fichiers bruts ont été pré-traités avec le logiciel Partek Flow®. L'analyse différentielle des gènes et leur interprétation ont été réalisées avec TranSCApp (application interne), utilisant les bases de données Gene Ontology, KEGG pathaways et STRING db. À j2, 373 gènes DE ont été identifiés, 168 gènes sous-exprimés et 205 sur-exprimés, qui semblent impliquer une sous-régulation de la voie de signalisation ribosomale et une sur-régulation de l'activation cellulaire : plaquettaire et immunitaire. À j4, 79 gènes DE ont été identifiés, 40 sous-exprimés et 39 sur-exprimés, impliqués notamment dans la machinerie transcriptionnelle et traductionnelle, ainsi que la voie de signalisation ribosomale principalement enrichie. À j2, les plaquettes semblent maintenir leurs fonctions essentielles (hémostase et réponse immunitaire), contrairement à j4. Ceci pourrait s'expliquer par la dégradation des ARNm, due en partie à la dérégulation des protéines ribosomales (RPL et RPS). Ces observations renforcent l'idée des lésions de stockage.
Saccharomyces cerevisiae is the main actor of wine fermentation but at present, still little is known about the factors impacting its distribution in the vineyards. In this study, 23 vineyards and 7 cellars were sampled over 2 consecutive years in the Bordeaux and Bergerac regions. The impact of geography and farming system and the relation between grape and vat populations were evaluated using a collection of 1374 S. cerevisiae merlot grape isolates and 289 vat isolates analyzed at 17 microsatellites loci. A very high genetic diversity of S. cerevisiae strains was obtained from grape samples, higher in conventional farming system than in organic one. The geographic appellation and the wine estate significantly impact the S. cerevisiae population structure, whereas the type of farming system has a weak global effect. When comparing cellar and vineyard populations, we evidenced the tight connection between the two compartments, based on the high proportion of grape isolates (25%) related to the commercial starters used in the cellar and on the estimation of bidirectional geneflows between the vineyard and the cellar compartments.
Winemakers are increasingly keen to limit the use of commercial yeasts in order to reduce oenological inputs. The preparation of an indigenous winery-made fermentation starter from grapes called ‘pied de cuve’ (PdC) is becoming popular, especially in organic farming systems. However, the implementation of the PdC method is still empirical and knowledge is lacking regarding the impact of PdC on S. cerevisiae diversity during alcoholic fermentation. In this study, the impact of PdC on S. cerevisiae genetic diversity and wine composition was evaluated at an industrial scale. Despite very low initial population level of S. cerevisiae before inoculation, the use of PdC was as efficient as Active Dry Yeast in terms of fermentation kinetics and chemical analyses on the resulting wines, except for one modality. At mid-fermentation, the diversity of S. cerevisiae strains was different depending on the PdC used, and was also different from that in the spontaneous fermentation with, in some cases, clonal expansion. Our results provide evidence that the use of PdC could secure the fermentation process more efficiently than spontaneous fermentation.
Application of high-throughput sequencing technologies to microsatellite genotyping (SSRseq) has been shown to remove many of the limitations of electrophoresis-based methods and to refine inference of population genetic diversity and structure. We present here a streamlined SSRseq development workflow that includes microsatellite development, multiplexed marker amplification and sequencing, and automated bioinformatics data analysis. We illustrate its application to five groups of species across phyla (fungi, plant, insect and fish) with different levels of genomic resource availability. We found that relying on previously developed microsatellite assay is not optimal and leads to a resulting low number of reliable locus being genotyped. In contrast, de novo ad hoc primer designs gives highly multiplexed microsatellite assays that can be sequenced to produce high quality genotypes for 20–40 loci. We highlight critical upfront development factors to consider for effective SSRseq setup in a wide range of situations. Sequence analysis accounting for all linked polymorphisms along the sequence quickly generates a powerful multi-allelic haplotype-based genotypic dataset, calling to new theoretical and analytical frameworks to extract more information from multi-nucleotide polymorphism marker systems.
Les plaquettes jouent un rôle essentiel dans l'hémostase et la thrombose. Malgré la mise en œuvre de la leucoréduction, les effets indésirables receveurs (EIR) n'ont pas disparu. Le but de ce travail est de caractériser l'ensemble des modifications du transcriptome plaquettaire dans les mélanges de concentrés plaquettaires (MCP) impliqués dans un EIR. Une analyse transcriptomique des plaquettes a été réalisée dans 6 MCP associés à l'apparition d'un EIR versus 6 témoins appariés. Le séquençage a été réalisé avec la technologie Ion-Proton®, l'alignement des reads avec le logiciel STAR sur le génome de référence hg19, la quantification des gènes à l'aide de l'algorithme EM du logiciel Partek® Flow. Les packages R (DESeq2, Edge R, Limma voom) ont été utilisés pour les analyses statistiques. Les protéines codant pour les gènes différentiellement exprimées (DE) qui impactent sur les processus biologiques (BP) et sur les voies de signalisation ont été étudiées via une analyse d'interaction gène-gène (String db) et une analyse fonctionnelle (PANTHER). Soixante dix neuf gènes significativement DE (p > 0,05) ont été identifiés. Un réseau constitué de 15 gènes fortement connectés a été mis en évidence, avec SPTBN1 comme gène hub. L'analyse fonctionnelle révèle, entre autre, un enrichissement de la voie de transduction du signal impliquant RHOC et ARHGAP30 (retrouvés en interaction dans String db). Des modifications profondes du transcriptome des MCP ont été observées. Ces observations renforcent l'idée que l'apparition d'un EIR puisse être associée à la modification de la structure plaquettaire et plus précisément par remodelage de son cytosquelette.
Over the past decade, a new strategy was developed to bypass the difficulties to genetically engineer some microbial species by transferring (or "cloning") their genome into another organism that is amenable to efficient genetic modifications and therefore acts as a living workbench. As such, the yeast Saccharomyces cerevisiae has been used to clone and engineer genomes from viruses, bacteria, and algae. The cloning step requires the insertion of yeast genetic elements in the genome of interest, in order to drive its replication and maintenance as an artificial chromosome in the host cell. Current methods used to introduce these genetic elements are still unsatisfactory, due either to their random nature (transposon) or the requirement for unique restriction sites at specific positions (TAR cloning). Here we describe the CReasPy-cloning, a new method that combines both the ability of Cas9 to cleave DNA at a user-specified locus and the yeast's highly efficient homologous recombination to simultaneously clone and engineer a bacterial chromosome in yeast. Using the 0.816 Mbp genome of Mycoplasma pneumoniae as a proof of concept, we demonstrate that our method can be used to introduce the yeast genetic element at any location in the bacterial chromosome while simultaneously deleting various genes or group of genes. We also show that CReasPy-cloning can be used to edit up to three independent genomic loci at the same time with an efficiency high enough to warrant the screening of a small (<50) number of clones, allowing for significantly shortened genome engineering cycle times.
The archaeological site of Santa Ana-La Florida (SALF), located in the Ecuadorian upper Amazon, is in the region of Theobroma spp. greatest genetic diversity, thus making it ideal to investigate the origins of domestication of this enigmatic tree. We present research showing that the residents of SALF were involved in the domestication of cacao, traditionally thought to have been first domesticated in Mesoamerica and/or Central America. We used three independent lines of evidence—starch grains, theobromine residues and ancient DNA—dating from approximately 5,300 years ago, to establish the earliest evidence of T. cacao use in the Americas, the first unequivocal archaeological example of its pre-Columbian use in South America and reveal the upper Amazon region as the oldest centre of cacao domestication yet identified. We suggest that new paleoethnobotanical research will expand our knowledge of this process, including the timing, locations, and uses of cacao by Indigenous South Americans.
Brettanomyces bruxellensis is a unicellular fungus of increasing industrial and scientific interest over the past 15 years. Previous studies revealed high genotypic diversity amongst B. bruxellensis strains as well as strain-dependent phenotypic characteristics. Genomic assemblies revealed that some strains harbour triploid genomes and based upon prior genotyping it was inferred that a triploid population was widely dispersed across Australian wine regions. We performed an intraspecific diversity genotypic survey of 1488 B. bruxellensis isolates from 29 countries, 5 continents and 9 different fermentation niches. Using microsatellite analysis in combination with different statistical approaches, we demonstrate that the studied population is structured according to ploidy level, substrate of isolation and geographical origin of the strains, underlying the relative importance of each factor. We found that geographical origin has a different contribution to the population structure according to the substrate of origin, suggesting an anthropic influence on the spatial biodiversity of this microorganism of industrial interest. The observed clustering was correlated to variable stress response, as strains from different groups displayed variation in tolerance to the wine preservative sulfur dioxide (SO2). The potential contribution of the triploid state for adaptation to industrial fermentations and dissemination of the species B. bruxellensis is discussed.
Cacao (Theobroma cacao L.) is an important economic crop, yet studies of its domestication history and early uses are limited. Traditionally, cacao is thought to have been first domesticated in Mesoamerica. However, genomic research shows that T. cacao's greatest diversity is in the upper Amazon region of northwest South America, pointing to this region as its centre of origin. Here, we report cacao use identified by three independent lines of archaeological evidence-cacao starch grains, absorbed theobromine residues and ancient DNA-dating from approximately 5,300 years ago recovered from the Santa Ana-La Florida (SALF) site in southeast Ecuador. To our knowledge, these findings constitute the earliest evidence of T. cacao use in the Americas and the first unequivocal archaeological example of its pre-Columbian use in South America. They also reveal the upper Amazon region as the oldest centre of cacao domestication yet identified.
Starmerella bacillaris is an ascomycetous yeast ubiquitously present in grapes and fermenting grape musts. In this report, we present the draft genome sequence of the S. bacillaris type strain CBS 9494, isolated from sweet botrytized wines, which will contribute to the study of this genetically heterogeneous wine yeast species.
[This corrects the article DOI: 10.1128/MRA.00872-18.].
We have designed a new efficient dimensionality reduction algorithm in order to investigate new ways of accurately characterizing the biodiversity, namely from a geometric point of view, scaling with large environmental sets produced by NGS (∼ 10^5 sequences). The approach is based on Multidimensional Scaling (MDS) that allows for mapping items on a set of n points into a low dimensional euclidean space given the set of pairwise distances. We compute all pairwise distances between reads in a given sample, run MDS on the distance matrix, and analyze the projection on first axis, by visualization tools. We have circumvented the quadratic complexity of computing pairwise distances by implementing it on a hyperparallel computer (Turing, a Blue Gene Q), and the cubic complexity of the spectral decomposition by implementing a dense random projection based algorithm. We have applied this data analysis scheme on a set of 10^5 reads, which are amplicons of a diatom environmental sample from Lake Geneva. Analyzing the shape of the point cloud paves the way for a geometric analysis of biodiversity, and for accurately building OTUs (Operational Taxonomic Units), when the data set is too large for implementing unsupervised, hierarchical, high-dimensional clustering.
MicroRNAs are key factors in the regulation of gene expression and their deregulation has been directly linked to various pathologies such as cancer. The use of small molecules to tackle the overexpression of oncogenic miRNAs has proved its efficacy and holds the promise for therapeutic applications. Here we describe the screening of a 640-compound library and the identification of polyamine derivatives interfering with in vitro Dicer-mediated processing of the oncogenic miR-372 precursor (pre-miR-372). The most active inhibitor is a spermine-amidine conjugate that binds to the pre-miR-372 with a KD of 0.15 µM, and inhibits its in vitro processing with a IC50 of 1.06 µM. The inhibition of miR-372 biogenesis was confirmed in gastric cancer cells overexpressing miR-372 and a specific inhibition of proliferation through de-repression of the tumor suppressor LATS2 protein, a miR-372 target, was observed. This compound modifies the expression of a small set of miRNAs and its selective biological activity has been confirmed in patient-derived ex vivo cultures of gastric carcinoma. Polyamine derivatives are promising starting materials for future studies about the inhibition of oncogenic miRNAs and, to the best of our knowledge, this is the first report about the application of functionalized polyamines as miRNAs interfering agents.
Oaks are an important part of our natural and cultural heritage. Not only are they ubiquitous in our most common landscapes1 but they have also supplied human societies with invaluable services, including food and shelter, since prehistoric times2. With 450 species spread throughout Asia, Europe and America3, oaks constitute a critical global renewable resource. The longevity of oaks (several hundred years) probably underlies their emblematic cultural and historical importance. Such long-lived sessile organisms must persist in the face of a wide range of abiotic and biotic threats over their lifespans. We investigated the genomic features associated with such a long lifespan by sequencing, assembling and annotating the oak genome. We then used the growing number of whole-genome sequences for plants (including tree and herbaceous species) to investigate the parallel evolution of genomic characteristics potentially underpinning tree longevity. A further consequence of the long lifespan of trees is their accumulation of somatic mutations during mitotic divisions of stem cells present in the shoot apical meristems. Empirical4 and modelling5 approaches have shown that intra-organismal genetic heterogeneity can be selected for6 and provides direct fitness benefits in the arms race with short-lived pests and pathogens through a patchwork of intra-organismal phenotypes7. However, there is no clear proof that large-statured trees consist of a genetic mosaic of clonally distinct cell lineages within and between branches. Through this case study of oak, we demonstrate the accumulation and transmission of somatic mutations and the expansion of disease-resistance gene families in trees.
BACKGROUND:Atypical Myeloproliferative Neoplasms (aMPN) share characteristics of MPN and Myelodysplastic Syndromes. Although abnormalities in cytokine signaling are common in MPN, the pathophysiology of atypical MPN still remains elusive. Since deregulation of microRNAs is involved in the biology of various cancers, we studied the miRNome of aMPN patients.METHODS:MiRNome and mutations in epigenetic regulator genes ASXL1, TET2, DNMT3A, EZH2 and IDH1/2 were explored in aMPN patients. Epigenetic regulation of miR-10a and HOXB4 expression was investigated by treating hematopoietic cell lines with 5-aza-2'deoxycytidine, valproic acid and retinoic acid. Functional effects of miR-10a overexpression on cell proliferation, differentiation and self-renewal were studied by transducing CD34+ cells with lentiviral vectors encoding the pri-miR-10a precursor.RESULTS:MiR-10a was identified as the most significantly up-regulated microRNA in aMPN. MiR-10a expression correlated with that of HOXB4, sitting in the same genomic locus. The transcription of these two genes was increased by DNA demethylation and histone acetylation, both necessary for optimal expression induction by retinoic acid. Moreover, miR-10a and HOXB4 overexpression seemed associated with DNMT3A mutation in hematological malignancies. However, overexpression of miR-10a had no effect on proliferation, differentiation or self-renewal of normal hematopoietic progenitors.CONCLUSIONS:MiR-10a and HOXB4 are overexpressed in aMPN. This overexpression seems to be the result of abnormalities in epigenetic regulation mechanisms. Our data suggest that miR-10a could represent a simple marker of transcription at this genomic locus including HOXB4, widely recognized as involved in stem cell expansion.
Understanding the mechanisms behind the typicity of regional wines inevitably brings attention to microorganisms associated with their production. Oenococcus oeni is the main bacterial species involved in wine and cider making. It develops after the yeast-driven alcoholic fermentation and performs the malolactic fermentation, which improves the taste and aromatic complexity of most wines. Here, we have evaluated the diversity and specificity of O. oeni strains in six regions. A total of 235 wines and ciders were collected during spontaneous malolactic fermentations and used to isolate 3,212 bacterial colonies. They were typed by multilocus variable analysis, which disclosed a total of 514 O. oeni strains. Their phylogenetic relationships were evaluated by a second typing method based on single nucleotide polymorphism (SNP) analysis. Taken together, the results indicate that each region holds a high diversity of strains that constitute a unique population. However, strains present in each region belong to diverse phylogenetic groups, and the same groups can be detected in different regions, indicating that strains are not genetically adapted to regions. In contrast, greater strain identity was seen for cider, white wine, or red wine of Burgundy, suggesting that genetic adaptation to these products occurred. IMPORTANCE:This study reports the isolation, genotyping, and geographic distribution analysis of the largest collection of O. oeni strains performed to date. It reveals that there is very high diversity of strains in each region, the majority of them being detected in a single region. The study also reports the development of an SNP genotyping method that is useful for analyzing the distribution of O. oeni phylogroups. The results show that strains are not genetically adapted to regions but to specific types of wines. They reveal new phylogroups of strains, particularly two phylogroups associated with white wines and red wines of Burgundy. Taken together, the results shed light on the diversity and specificity of wild strains of O. oeni, which is crucial for understanding their real contribution to the unique properties of wines.