In the absence of a DNA template, the ab initio production of long double-stranded DNA molecules of predefined sequences is particularly challenging. The DNA synthesis step remains a bottleneck for many applications such as functional assessment of ancestral genes, analysis of alternative splicing or DNA-based data storage. In this report we propose a fully in vitro protocol to generate very long double-stranded DNA molecules starting from commercially available short DNA blocks in less than 3 days using Golden Gate assembly. This innovative application allowed us to streamline the process to produce a 24 kb-long DNA molecule storing part of the Declaration of the Rights of Man and of the Citizen of 1789 . The DNA molecule produced can be readily cloned into a suitable host/vector system for amplification and selection.
In absence of DNA template, the ab initio production of long double-stranded DNA molecules of predefined sequences is particularly challenging. The DNA synthesis step remains a bottleneck for many applications such as functional assessment of ancestral genes, analysis of alternative splicing or DNA-based data storage. We propose in this report a fully in vitro protocol to generate very long double-stranded DNA molecule starting from commercially available short DNA blocks in less than 3 days. This innovative application of Golden Gate assembly allowed us to streamline the assembly process to produce a 24 kb long DNA molecule storing part of the Universal Declaration of Human rights and citizens. The DNA molecule produced can be readily cloned into suitable host/vector system for amplification and selection.
Background Streptococcus thermophilus is a Gram-positive bacterium widely used as starter in the dairy industry as well as in many traditional fermented products. In addition to its technological importance, it has also gained interest in recent years as beneficial bacterium due to human health-promoting functionalities. The objective of this study was to inventory the main health-promoting properties of S. thermophilus and to study their intra-species diversity at the genomic and genetic level within a collection of representative strains. Results In this study various health-related functions were analyzed at the genome level from 79 genome sequences of strains isolated over a long time period from diverse products and different geographic locations. While some functions are widely conserved among isolates (e.g., degradation of lactose, folate production) suggesting their central physiological and ecological role for the species, others including the tagatose-6-phosphate pathway involved in the catabolism of galactose, and the production of bioactive peptides and gamma-aminobutyric acid are strain-specific. Most of these strain-specific health-promoting properties seems to have been acquired via horizontal gene transfer events. The genetic basis for the phenotypic diversity between strains for some health related traits have also been investigated. For instance, substitutions in the galK promoter region correlate with the ability of some strains to catabolize galactose via the Leloir pathway. Finally, the low occurrence in S. thermophilus genomes of genes coding for biogenic amine production and antibiotic resistance is also a contributing factor to its safety status. Conclusions The natural intra-species diversity of S. thermophilus , therefore, represents an interesting source for innovation in the field of fermented products enriched for healthy components that can be exploited to improve human health. A better knowledge of the health-promoting properties and their genomic and genetic diversity within the species may facilitate the selection and application of strains for specific biotechnological and human health-promoting purpose. Moreover, by pointing out that a substantial part of its functional potential still defies us, our work opens the way to uncover additional health-related functions through the intra-species diversity exploration of S. thermophilus by comparative genomics approaches.
This study aimed to provide efficient recognition of bacterial strains on personal computers from MinION (Nanopore) long read data. Thanks to the fall in sequencing costs, the identification of bacteria can now proceed by whole genome sequencing. MinION is a fast, but highly error -prone sequencing device and it is a challenge to successfully identify the strain content of unknown simple or complex microbial samples. It is heavily constrained by memory management and fast access to the read and genome fragments. Our strategy involves three steps: indexing of known genomic sequences for a given or several bacterial species; a request process to assign a read to a strain by matching it to the closest reference genomes; and a final step looking for a minimum set of strains that best explains the observed reads. We have applied our method, called ORI, on 77 strains of Streptococcus thermophilus. We worked on several genomic distances and obtained a detailed classification of the strains, together with a criterion that allows merging of what we termed 'sibling' strains, only separated by a few mutations. Overall, isolated strains can be safely recognized from MinION data. For mixtures of several non-sibling strains, results depend on strain abundance.
HAL is a multi-disciplinary open access archive for the deposit and dissemination of scientific research documents, whether they are published or not. The documents may come from teaching and research institutions in France or abroad, or from public or private research centers. L’archive ouverte pluridisciplinaire HAL, est destinée au dépôt et à la diffusion de documents scientifiques de niveau recherche, publiés ou non, émanant des établissements d’enseignement et de recherche français ou étrangers, des laboratoires publics ou privés. Bacterial strains identification using Oxford Nanopore sequencing Grégoire Siekaniec, Emeline Roux, Eric Guédon, Jacques Nicolas
Inferring genome-scale metabolic networks in emerging model organisms is challenged by incomplete biochemical knowledge and partial conservation of biochemical pathways during evolution. Therefore, specific bioinformatic tools are necessary to infer biochemical reactions and metabolic structures that can be checked experimentally. Using an integrative approach combining genomic and metabolomic data in the red algal model Chondrus crispus, we show that, even metabolic pathways considered as conserved, like sterols or mycosporine-like amino acid synthesis pathways, undergo substantial turnover. This phenomenon, here formally defined as "metabolic pathway drift," is consistent with findings from other areas of evolutionary biology, indicating that a given phenotype can be conserved even if the underlying molecular mechanisms are changing. We present a proof of concept with a methodological approach to formalize the logical reasoning necessary to infer reactions and molecular structures, abstracting molecular transformations based on previous biochemical knowledge.
Long-read sequencing currently provides sequences of several thousand base pairs. It is therefore possible to obtain complete transcripts, offering an unprecedented vision of the cellular transcriptome. However the literature lacks tools for de novo clustering of such data, in particular for Oxford Nanopore Technologies reads, because of the inherent high error rate compared to short reads. Our goal is to process reads from whole transcriptome sequencing data accurately and without a reference genome in order to reliably group reads coming from the same gene. This de novo approach is therefore particularly suitable for non-model species, but can also serve as a useful pre-processing step to improve read mapping. Our contribution both proposes a new algorithm adapted to clustering of reads by gene and a practical and free access tool that allows to scale the complete processing of eukaryotic transcriptomes. We sequenced a mouse RNA sample using the MinION device. This dataset is used to compare our solution to other algorithms used in the context of biological clustering. We demonstrate that it is the best approach for transcriptomics long reads. When a reference is available to enable mapping, we show that it stands as an alternative method that predicts complementary clusters.
Because of the increasing size and complexity of available graph structures in experimental sciences like molecular biology, techniques of graph visualization tend to reach their limit. To assist experimental scientists into the understanding of the underlying phenomena, most visualization methods are based on the organization of edges and nodes in clusters. Among recent ones, Power Graph Analysis is a lossless compression of the graph based on the search of cliques and bicliques, improving the readability of the overall structure. Royer et al. introduced a heuristic approach providing approximate solutions to this NP-complete problem. Later, Bourneuf et al. formalized the heuristic using Formal Concept Analysis. This paper proposes to extend this work by a formalization of the graph compression search space. It shows that (1) the heuristic cannot always achieve an optimal compression, and (2) the concept lattice associated to a graph enables a more complete exploration of the search space. Our conclusion is that the search for graph compression can be usefully associated with the search for patterns in the concept lattice and that, conversely, confusing sets of objects and attributes brings new interesting problems for FCA.
In order to predict the behavior of a biological system, one common approach is to perform a simulation on a dynamic model. Boolean networks allow to analyze the qualitative aspects of the model by identifying its steady states and attractors. Each of them, when possible, is associated with a phenotype which conveys a biological interpretation. Phenotypes are characterized by their signatures, provided by domain experts. The number of steady states tends to increase with the network size and the number of simulation conditions, which makes the biological interpretation difficult. As a first step, we explore the use of Formal Concept Analysis as a symbolic bi-clustering technics to classify and sort the steady states of a Boolean network according to biological signatures based on the hierarchy of the roles the network components play in the phenotypes. FCA generates a lattice structure describing the dependencies between proteins in the signature and steady-states of the Boolean network. We use this lattice (i) to enrich the biological signatures according to the dependencies carried by the network dynamics, (ii) to identify variants to the phenotypes and (iii) to characterize hybrid phenotypes. We applied our approach on a T helper lymphocyte (Th) differentiation network with a set of signatures corresponding to the sub-types of Th. Our method generated the same classification as a manual analysis performed by experts in the field, and was also able to work under extended simulation conditions. This led to the identification and prediction of a new hybrid sub-type later confirmed by the literature.
Lately, long read sequencing technologies, referred to as Third Generation Sequencing, (TGS, Pacific Bioscience [1] and Nanopore [2]) have brought the opportunity to sequence full-length RNA molecules. In doing so they relax the constraint of transcript reconstruction prior to study complete RNA transcripts. By avoiding limitations of previous technologies, [3, 4] and giving access to the trancripts structure, they might contribute to complement and improve transcriptomes studies. This is particularly crucial for non model species where assembly was required. Many biological questions (finding gene signatures for a trait, finding expressed variants...)[5, 6] are classically addressed using transcriptome sequencing. However, this gain in length is at the cost of a computationally challenging error rate (up to 15%) that disqualifies previous short-reads methods. In this work we propose to support the analysis of RNA long read sequencing with a clustering method that works at the gene level. It enables to group transcripts that emerged from a same gene. From the clusters, the expression of each gene is obtained and related transcripts are identied, even when no reference is available.
This work addresses the problem of grouping by genes long reads expressed in a whole transcriptome sequencing data set. Long read sequencing produces several thousands base-pair long sequences, although showing high error rate in comparison to short reads. Long reads can cover full-length RNA transcripts and thus are of high interest to complete references. However, the literature is lacking tools to cluster such data de novo, in particular for Oxford Nanopore Technologies reads. As a consequence, we propose a novel algorithm based on community detection and its implementation. Since solution is meant to be reference-free (de novo), it is especially well-tailored for non model species. We demonstrate it performs well on a real mouse data set. When a reference is available, we show that it stands as an alternative to mapping. In addition, we show that quick assessment of gene's expression is a straightforward use case of our solution.
This work addresses the problem of assigning a set of long reads issued from a de novo transcriptomics study to clusters by genes they originate from. The different transcripts of a gene give long reads sharing similar sequences and our work makes use of this fact to retrieve the right cluster of reads for each gene from the graph of similarity between reads. We propose a method based on the use of the clustering coefficient (CC) and the search of a minimal cut in the graph with a greedy procedure favoring nodes with a high degree and high CC. Our approach compares favorably to state of the art methods. We provide results on the mouse brain transcriptome which show that the approach achieves a high precision level and a good level of recall despite not using any reference genome.
Background The emergence of functions in biological systems is a long-standing issue that can now be addressed at the cell level with the emergence of high throughput technologies for genome sequencing and phenotyping. The reconstruction of complete metabolic networks for various organisms is a key outcome of the analysis of these data, giving access to a global view of cell functioning. The analysis of metabolic networks may be carried out by simply considering the architecture of the reaction network or by taking into account the stoichiometry of reactions. In both approaches, this analysis is generally centered on the outcome of the network and considers all metabolic compounds to be equivalent in this respect. As in the case of genes and reactions, about which the concept of essentiality has been developed, it seems, however, that some metabolites play crucial roles in system responses, due to the cell structure or the internal wiring of the metabolic network. Results We propose a classification of metabolic compounds according to their capacity to influence the activation of targeted functions (generally the growth phenotype) in a cell. We generalize the concept of essentiality to metabolites and introduce the concept of the phenotypic essential metabolite (PEM) which influences the growth phenotype according to sustainability, producibility or optimal-efficiency criteria. We have developed and made available a tool, Conquests , which implements a method combining graph-based and flux-based analysis, two approaches that are usually considered separately. The identification of PEMs is made effective by using a logical programming approach. Conclusion The exhaustive study of phenotypic essential metabolites in six genome-scale metabolic models suggests that the combination and the comparison of graph, stoichiometry and optimal flux-based criteria allows some features of the metabolic network functionality to be deciphered by focusing on a small number of compounds. By considering the best combination of both graph-based and flux-based techniques, the Conquests python package advocates for a broader use of these compounds both to facilitate network curation and to promote a precise understanding of metabolic phenotype.
Phenotypic plasticity is an adaptive process widespread in Insects. This involves the capacity of embryos to sense external local environmental cues during their development with alternative routes providing a final phenotype adapted to those new cues. In our lab, we describe and analyse the gene expression regulation rules during the establishment of the phenotypic plasticity of the reproductive mode in aphids due to the change of the photoperiod in the fall. The development of sexual and asexual embryos that differ mainly by their production of haploid meiotic gametes or diploid non-recombinant gametes is thus compared. We developed an integrative genomics approach that aims to define all the functional DNA elements in the formation of alternative morphs. We thus combine genomics (nucleic acid sequencing), bioinformatics and mathematic modelling to construct gene networks describing and modelling the molecular processes acting in the phenotypic plasticity of the reproductive mode. We annotated functional DNA elements (open chromatin, long non-coding RNAs, microRNAs, mRNAs), predicted their putative interactions and modelled their network functioning. We identified two major biological processes as important signatures of these networks: “oogenesis” and “central nervous system”. We analysed more deeply some of the candidates to fill the gap between prediction, hypotheses and biological observations.
Molecular biology produces and accumulates huge amounts of data that are generally integrated within graphs of molecules linked by various interactions. Exploring potentially interesting substructures (clusters, motifs) within such graphs requires proper abstraction and visualization methods. Most layout techniques (edge and nodes spatial organization) prove insufficient in this case. Royer et al. introduced in 2008 Power graph analysis, a dedicated program using classes of nodes with similar properties and classes of edges linking node classes to achieve a lossless graph compression. The contributions of this paper are twofold. First, we formulate and study this issue in the framework of Formal Concept Analysis. This leads to a generalized view of the initial problem offering new variants and solving approaches. Second, we state the FCA modeling problem in a logical setting, Answer Set programming, which provides a great flexibility for the specification of concept search spaces.
Plant genomes contain a particularly high proportion of repeated structures of various types. This chapter proposes a guided tour of available software that can help biologists to look for these repeats and check some hypothetical models intended to characterize their structures. Since transposable elements are a major source of repeats in plants, many methods have been used or developed for this large class of sequences. They are representative of the range of tools available for other classes of repeats and we have provided a whole section on this topic as well as a selection of the main existing software. In order to better understand how they work and how repeats may be efficiently found in genomes, it is necessary to look at the technical issues involved in the large-scale search of these structures. Indeed, it may be hard to keep up with the profusion of proposals in this dynamic field and the rest of the chapter is devoted to the foundations of the search for repeats and more complex patterns. The second section introduces the key concepts that are useful for understanding the current state of the art in playing with words, applied to genomic sequences. This can be seen as the first stage of a very general approach called linguistic analysis that is interested in the analysis of natural or artificial texts. Words, the lexical level, correspond to simple repeated entities in texts or strings. In fact, biologists need to represent more complex entities where a repeat family is built on more abstract structures, including direct or inverted small repeats, motifs, composition constraints as well as ordering and distance constraints between these elementary blocks. In terms of linguistics, this corresponds to the syntactic level of a language. The last section introduces concepts and practical tools that can be used to reach this syntactic level in biological sequence analysis.
The Molecular Distance Geometry Problem (MDGP) is the problem of finding the possible conformations of a molecule by exploiting available information about distances between some atom pairs. When particular assumptions are satisfied, the MDGP can be discretized, so that the search domain of the problem becomes a tree. This tree can be explored by using an interval Branch & Prune (iBP) algorithm. In this context, the order given to the atoms of the molecules plays an important role. In fact, the discretization assumptions are strongly dependent on the atomic ordering, which can also impact the computational cost of the iBP algorithm. In this work, we propose a new partial discretization order for protein backbones. This new atomic order optimizes a set of objectives, that aim at improving the iBP performances. The optimization of the objectives is performed by Answer Set Programming (ASP), a declarative programming language that allows to express our problem by a set of logical constraints. The comparison with previously proposed orders for protein backbones shows that this new discretization order makes iBP perform more efficiently.
Targeting the deleterious effects of Transforming Growth Factor TGF-β without affecting its physiological role is the common goal of therapeutic strategies aiming at curing fibrosis, the final outcome of all chronic liver disease. The pleiotropic effects of TGF-β are linked to the complex nature of its activation and signaling net- works which understanding requires modeling approaches. Our group recently developed a model of TGF-beta signal propagation based on guarded transitions (ref, Andrieux et al, 2014). In this initial work, we explored the combinatorial complexity of cell signaling, developing a discrete formalism based on guarded transitions. We imported the whole database Pathway Interaction Database into a single unified model of signal transduction. We detected 16,000 chains of reactions linking TGF-β to at least one of 159 target genes in the nucleus. The size and complexity of this model place it beyond current understanding. Its analysis requires automated tools for identifying general patterns. Currently, we focus on designing one reasoning method based on Semantic Web technologies for the analysis of signaling pathways. Our method aims at leveraging external domain knowledge represented in biomedical ontologies and linked databases to rank these candidates. We consider a signaling pathway as a set of proteins involved in the respons of a cell to an external stimulus and influencing at least one gene. The underlying reasoning methods are based on graph topological analysis, formal concepts analysis (FCA) and semantic similarity and particularity measures. First, we determine the formal concepts, maximal bi-cliques, between proteins sets and genes. Then, to determine the biological relevance of theses gene clusters, we calculate a similarity score for each cluster based on Wang semantic similarity. Using such approaches, we identify groups of genes sharing signaling networks.