Abstract Understanding genome regulation is limited by the complexity of molecular interactions in living cells. Cell-free systems provide a simplified platform for studying gene expression, but low mRNA levels have prevented RNA-seq. To address this, we develop an active learning workflow combining Bayesian optimization with automated high-throughput experimentation to systematically explore over 1.6 million buffer compositions, experimentally testing 653. We identify a “mRNA-optimized” buffer (20-fold increase in mRNA yield) and a “trade-off” buffer (13-fold increase while maintaining protein production). Using direct RNA-seq, we profile the T7 phage transcriptome in cell-free systems and compare it with a purified T7-RNAP transcription system and phage-infected bacteria. This comparative analysis reveals distinct regulatory layers: the T7-RNAP system captures promoter-strength hierarchies but lacks RNA degradation, whereas cell-free systems provide an accurate estimation of in vivo expression and reveal mRNA maturation sites. This work establishes cell-free transcriptomics as a controlled approach to study genome regulation.
The complete genome sequence of the strain GD386 of Mycobacterium avium subspecies hominissuis isolated from a patient with lung infection in France was determined. The genome was sequenced using the PacBio technology, yielding a genome size of 5,562,671 nucleotides with no identified plasmids.
Milk oligosaccharides are bioactive components that regulate the composition of the neonatal microbiota and exert immunomodulatory functions. Their beneficial effects depend on their structure. Numerous studies have shown intra- and inter-species variation in the structural composition and concentration of these compounds in mammalian milk, yet the biological significance of such variation remains poorly understood. Automated natural language processing methods are promising tools for extracting and gathering structured data from unstructured texts to get insight into the biological significance of milk oligosaccharide variation across mammals. These methods require training and evaluation on manually annotated text corpora. While annotated corpora exist for chemical substances, none are specifically designed for training natural language processing models to extract information on milk oligosaccharides. To this end, we propose MilkOligoCorpus, a new gold standard for milk oligosaccharide composition in mammalian species. MilkOligoCorpus' annotation scheme is a rich entity/relation model designed to describe the diversity pattern of milk oligosaccharides according to female factor variability and to help better understand the structure-related function of milk oligosaccharides. MilkOligoCorpus consists of abstracts (15) and extracts (15) from 20 full text articles indexed by PubMed annotated with entities related to individuals, samples, oligosaccharides and oligosaccharide quantification linked by binary and n-ary relationships. To address data interoperability across disparate publications and databases, four terminological resources were also developed to assign unique identifiers to the entities, supported by external ontologies. This paper presents the creation of the MilkOligoCorpus and its associated schema, along with the development of annotation guidelines and terminological resources. We also present experimental results obtained by baseline information extraction models on the corpus.
Cheeses are fermented dairy products consumed worldwide. Their global diversity results from various local variables, including technological practices, as well as the metabolic activity of diverse microorganisms. In Europe, this typicity is exemplified by Protected Designation of Origin (PDO) cheeses, for which genetic diversity remains largely unexplored. Combining culturomics (n = 373 bacterial genomes) and metagenomic (n = 146 metagenomes), we performed a national-scale survey of the microbial diversity encompassing 44 French PDO cheeses. Taxonomic (bacteria, fungi and viruses) and functional profiling reveal a high diversity in the cheese rind, mainly driven by the cheese technology. We also reconstructed 1,119 bacterial metagenome-assembled genomes (MAGs) encompassing seven phyla, including Actinomycetota, Bacillota, Pseudomonadota and Bacteroidota. Using GTDB as a reference, we identified 221 MAGs encompassing 46 genera, as well as 44 bacterial isolate genomes encompassing eight genera, which represent potentially 81 new species (based on <95% ANI). These species were particularly numerous among the genera Halomonas, Psychrobacter and Brachybacterium. Similar results were observed when compared with the cFMD database. We combined our genomic and metagenomic datasets into a catalog of 26.2 million protein clusters, with 50% of these clusters remaining unassigned to a known function and taxonomy. We illustrated the potential of this resource by searching for methionine gamma-lyase (MGL), an enzyme playing a significant role in cheese flavor. This protein was predominantly found in Pseudoalteromonas, a potentially new MGL-producing genus, Serratia, Pseudomonas, Proteus and Hafnia, and its prevalence varied with cheese technology. Our study provides a substantial genomic resource for food microbiologists and cheesemakers to further explore the biotechnological potential of PDO cheese biodiversity. ### Competing Interest Statement The authors have declared no competing interest. France Génomique, ANR-10-INBS-0908 Centre national interprofessionnel de l'économie laitière, https://ror.org/00w571a59 Genoscope, https://ror.org/028pnqf58 Commissariat à l'Énergie Atomique et aux Énergies Alternatives, https://ror.org/00jjx8s55 Institut National de Recherche pour l'Agriculture, l'Alimentation et l'Environnement, https://ror.org/003vg9w96
Understanding the rate and nature of spontaneous mutations is crucial for understanding and modulating the pace and trajectory of evolution. Yet, both remain poorly characterized in double-stranded DNA (dsDNA) phages, despite their relevance for phage-based therapies. Here, we address this gap for the dsDNA phage lambda, using four complementary approaches: mutation accumulation assay combined with whole-genome sequencing, mutation visualization assay, duplex sequencing and fluctuation assay. We find that the mutation rate of wild-type phage lambda is 4.9 ± 1.8 x10-9 per base per replication, approximately 15 times lower than previously estimated and about 20 times higher than that of its host, Escherichia coli ( E. coli ). Inactivation of Mismatch Repair (MMR), a major conserved cellular system for mutation avoidance, increases the lambda mutation rate by only 2-to 10-fold, in contrast to the approximately 150-fold increase in E. coli , however lambda does not exhibit the characteristic mutational bias associated with MMR deficiency. Interestingly replication of the lambda genome generates an error spectrum distinct from that of E. coli , characterized by a marked increase in transversions, poorly repaired by MMR. Together, these results reveal that lambda exhibits a replication error profile that is less amenable to repair, likely contributing to its elevated mutation rate. ### Competing Interest Statement The authors have declared no competing interest. Agence Nationale de la Recherche, https://ror.org/00rbzpz17, ANR-20-CE12-0008
The study of microbial metabolic interactions within food microbiomes represents a key scientific approach for improving the quality and health benefits of food. In such studies, methods based on gene expression levels (metatranscriptomic) analysis are promising. However specific tools are required to overcome the challenges posed by food microbiomes, in particular the high variability of microbiomes between samples and the difficulty of automatically inferring the annotation of metabolic functions across taxa. To adress this gap, we present the Food Microbiome Metabolic Modules (F3M) tool suite, which comprises (1) a curated database containing about 1,985 functional genes representing key fermentative metabolic reactions in food microbiomes, (2) a F3M Builder for generating F3M-annotated gene catalogs and mapping of metatranscriptomic reads, and then (3) an F3M R package to parse and aggregate gene expression data by taxonomic and functional categories for downstream analysis. The F3M taxonomy is organized according to the Genome Taxonomy Database (GTDB) nomenclature, whereas the F3M functional repertoire is structured hierarchically into 183 metabolic modules, which enable multi-scale analysis of inter-organism metabolic interactions and meaningful fermentative outputs (e.g., primary alcohols, acetate). Notably, a dedicated 'redox' module captures oxido-reduction mechanisms and NADH-dependent pathways central to fermentation, while an 'uptake' module complements the metabolic pathways to trace potential metabolite exchanges across taxa. Together, the F3M suite provides a robust framework for uncovering functional dynamics within food microbiomes. The F3M tool suite is available as open-source.
The rapid growth of microbiome research has led to the development of numerous bioinformatics tools and databases, but information about them remains fragmented across disparate, often outdated cataloging efforts, hindering resource discovery and utilization. To address this critical gap, the ELIXIR Microbiome Community proposes the development of MiCoReCa (Microbiome Community Resource Catalogue), a comprehensive, dynamic, open-access catalogue of microbiome-related bioinformatics resources (tools, workflows, training, standards, and databases). Leveraging our community's expertise, this initiative will utilize standardized ontologies like EDAM and cross-reference established platforms like bio.tools and WorkflowHub to create a centralized, findable inventory. A key feature is the community-driven process for identifying and curating missing ontological terms and metadata, ensuring MiCoReCa's accuracy and relevance in collaboration with partner platforms. Furthermore, the catalogue will integrate links to training materials from TeSS to support appropriate tool usage, and connect with OpenEBench for benchmarking capabilities. This project will not only provide a vital resource for the microbiome field, enhancing research efficiency and reproducibility, but will also establish a sustainable, adaptable infrastructure potentially applicable to other ELIXIR Communities. This effort represents a significant contribution by the ELIXIR Microbiome Community to streamline microbiome bioinformatics.
Despite advances in transcriptomics, understanding of genome regulation remains limited by the complex interactions within living cells. To address this, we performed cell-free transcriptomics by developing a platform using an active learning workflow to explore over 1,000,000 buffer conditions. This enabled us to identify a buffer that increased mRNA yield by 20-fold, enabling cell-free transcriptomics. By employing increasingly complex conditions, our approach untangles the regulatory layers controlling genome expression. ### Competing Interest Statement The authors have declared no competing interest. ANR program, ANR-24-CE44-4467, ANR-11-IDEX-0003, ANR-22-PEBB-0008 UE HORIZON BIOS program, 101070281
The rapid growth of microbiome research has led to the development of numerous bioinformatics tools and databases, but information about them remains fragmented across disparate, often outdated cataloging efforts, hindering resource discovery and utilization. To address this critical gap, the ELIXIR Microbiome Community proposes the development of MiCoReCa (Microbiome Community Resource Catalogue), a comprehensive, dynamic, open-access catalogue of microbiome-related bioinformatics resources (tools, workflows, training, standards, and databases). Leveraging our community's expertise, this initiative will utilize standardized ontologies like EDAM and cross-reference established platforms like bio.tools and WorkflowHub to create a centralized, findable inventory. A key feature is the community-driven process for identifying and curating missing ontological terms and metadata, ensuring MiCoReCa's accuracy and relevance in collaboration with partner platforms. Furthermore, the catalogue will integrate links to training materials from TeSS to support appropriate tool usage, and connect with OpenEBench for benchmarking capabilities. This project will not only provide a vital resource for the microbiome field, enhancing research efficiency and reproducibility, but will also establish a sustainable, adaptable infrastructure potentially applicable to other ELIXIR Communities. This effort represents a significant contribution by the ELIXIR Microbiome Community to streamline microbiome bioinformatics.
Streptococcus suis is a bacterial pathogen responsible for infections in pigs and in wild fauna that can also lead to severe infections in humans. Increasing antimicrobial resistance (AMR) has been described for this zoonotic pathogen worldwide. Since most of these AMR genes are carried by mobile genetic elements (MGEs), they can largely disseminate by horizontal gene transfer. Taking advantage of the large set of genomes available for this species, an exhaustive search of integrative and conjugative elements (ICEs) and integrative and mobilizable elements (IMEs) was undertaken in a representative set of 400 selected high-quality genomes of S. suis. We examined how these elements vary across phylogenetic clades and ecotypes and their association with AMR genes and defence systems (DSs), including restriction-modification (RM), CRISPR and also less studied DSs. This investigation identified 569 ICEs, belonging to the 7 families previously described in streptococci, inserted in 12 distinct specific integration sites. Additionally, 1,035 IMEs characterized by 11 distinct relaxase families and integrated in 10 specific chromosomal sites were detected in the 400 genomes of S. suis. New associations between ICE/IME and AMR genes were discovered. A huge diversity of putative DSs was observed including 2,035 RM systems, 124 CRISPR systems and systems belonging to 20 other categories, most of them described as efficient against phages and plasmids. Furthermore, most of the spacers associated with CRISPR systems target these MGEs rather than integrative elements. In addition, many integrative elements appear to carry an orphan methylase that could help them escape RM systems. Altogether, this points out that ICEs and IMEs are spared by DSs and play a major role in AMR dissemination in S. suis. In addition, most of the strains have the full set of genes required for competence, i.e. for the acquisition of extracellular DNA by natural transformation. This suggests a high risk of AMR dissemination in S. suis.
Soil microbial communities respond quickly to natural and/or anthropic-induced changes in environmental conditions. Metagenomics allows studying taxa that are often overlooked in microbiota studies, such as protists or viruses. Here, we employed metagenomics to characterise microbial successions after wheat straw input in a 4-month in-situ field study. We compared microbial successions patterns with those obtained by high throughput amplicon sequencing on the same soil samples to validate metagenomics as a tool for the fine analysis of microbial population dynamics in situ. Taxonomic patterns were concordant between the two methodologies but metagenomics allowed studying all the microbial groups simultaneously. Notably, our results evidenced that each domain displayed a specific dynamic pattern after wheat straw amendment. For instance, viral sequences multiplied in the early phase of straw decomposition, in parallel to copiotrophic bacteria, suggesting a “kill-the-winner” pattern that, to our knowledge, had not been observed before in soil. Altogether, our results highlighted that both inter and intra-domain trophic interactions were impacted by wheat amendment and these patterns depended on the land use history. Our study highlights that top-down regulation by microbial predators or viruses might play a key role in soil microbiota dynamics and structure.
An exhaustive analysis was performed on more than 2000 microbiotas from French Protected Designation of Origin (PDO) cheeses, covering most cheese families produced throughout the world. Thanks to a complete and accurate set of associated metadata, we have carried out a deep analysis of the ecological drivers of microbial communities in milk and "terroir" cheeses. We show that bacterial and fungal microbiota from milk differed significantly across dairy species while sharing a core microbiome consisting of four microbial species. By contrast, no microbial species were detected in all ripened cheese samples. Our network analysis suggested that the cheese microbiota was organized into independent network modules. These network modules comprised mainly species with an overall relative abundance lower than 1%, showing that the most abundant species were not those with the most interactions. Species assemblages differed depending on human drivers, dairy species, and geographical area, thus demonstrating the contribution of regional know-how to shaping the cheese microbiota. Finally, an extensive analysis at the milk-to-cheese batch level showed that a high proportion of cheese taxa were derived from milk under the influence of the dairy species and protected designation of origin.
There is a growing interest in milk oligosaccharides (MOs) because of their numerous benefits for newborns’ and long-term health. A large number of MO structures have been identified in mammalian milk. Mostly described in human milk, the oligosaccharide richness, although less broad, has also been reported for a wide range of mammalian species. The structure of MOs is particularly difficult to report as it results from the combination of 5 monosaccharides linked by various glycosidic bonds forming structurally diverse and complex matrices of linear and branched oligosaccharides. Exploring the literature and extracting relevant information on MO diversity within or across species appears promising to elucidate structure-function role of MOs. Currently, given the complexity of these molecules, the main issues in exploring literature to extract relevant information on MO diversity within or across species relate to the heterogeneity in the way authors refer to these molecules. Herein, we provide a thesaurus (MilkOligoThesaurus) including the names and synonyms of MOs collected from key selected articles on mammalian milk analyses. MilkOligoThesaurus gathers the names of the MOs with a complete description of their monosaccharide composition and structures. When available, each unique MO molecule is linked to its ID from the NCBI PubChem and ChEBI databases. MilkOligoThesaurus is provided in a tabular format. It gathers 245 unique oligosaccharide structures described by 22 features (columns) including the name of the molecule, its abbreviation, the chemical database IDs if available, the monosaccharide composition, chemical information (molecular formula, monoisotopic mass), synonyms, its formula in condensed form, and in abbreviated condensed form, the abbreviated systematic name, the systematic name, the isomer group, and scientific article sources. MilkOligoThesaurus is also provided in the SKOS (Simple Knowledge Organization System) format. This thesaurus is a valuable resource gathering MO naming variations that are not found elsewhere for (i) Text and Data Mining to enable automatic annotation and rapid extraction of milk oligosaccharide data from scientific papers; (ii) biology researchers aiming to search for or decipher the structure of milk oligosaccharides based on any of their names, abbreviations or monosaccharide compositions and linkages.
Next generation sequencing offers several ways to study microbial communities. For agri-food sciences, identifying species in diverse food ecosystems is key for both food sustainability and food security. The aim of this study was to compare metabarcoding pipelines and markers to determine fungal diversity in food ecosystems, from Illumina short reads. We built mock communities combining the most representative fungal species in fermented meat, cheese, wine and bread. Four barcodes (ITS1, ITS2, D1/D2 and RPB2) were tested for each mock and on real fermented products. We created a database, including all mock species sequences for each barcode to compensate for the lack of curated data in available databases. Four bioinformatics tools (DADA2, QIIME, FROGS and a combination of DADA2 and FROGS) were compared. Our results clearly showed that the combined DADA2 and FROGS tool gave the most accurate results. Most mock community species were not identified by the RPB2 barcode due to unsuccessful barcode amplification. When comparing the three rDNA markers, ITS markers performed better than D1/D2, as they are better represented in public databases and have better specificity to distinguish species. Between ITS1 and ITS2, differences in the best marker were observed according to the studied ecosystem. While ITS2 is best suited to characterize cheese, wine and fermented meat communities, ITS1 performs better for sourdough bread communities. Our results also emphasized the need for a dedicated database and enriched fungal-specific public databases with novel barcode sequences for 118 major species in food ecosystems.
Background:The use of omics data for monitoring the microbial flow of fresh meat products along a production line and the development of spoilage prediction tools from these data is a promising but challenging task. In this context, we produced a large multivariate dataset (over 600 samples) obtained on the production lines of two similar types of fresh meat products (poultry and raw pork sausages). We describe a full analysis of this dataset in order to decipher how the spoilage microbial ecology of these two similar products may be shaped differently depending on production parameter characteristics.Methods:Our strategy involved a holistic approach to integrate unsupervised and supervised statistical methods on multivariate data (OTU-based microbial diversity; metabolomic data of volatile organic compounds; sensory measurements; growth parameters), and a specific selection of potential uncontrolled (initial microbiota composition) or controlled (packaging type; lactate concentration) drivers.Results:Our results demonstrate that the initial microbiota, which is shown to be very different between poultry and pork sausages, has a major impact on the spoilage scenarios and on the effect that a downstream parameter such as packaging type has on the overall evolution of the microbial community. Depending on the process, we also show that specific actions on the pork meat (such as deboning and defatting) elicit specific food spoilers such as Dellaglioa algida, which becomes dominant during storage. Finally, ecological network reconstruction allowed us to map six different metabolic pathways involved in the production of volatile organic compounds involved in spoilage. We were able connect them to the different bacterial actors and to the influence of packaging type in an overall view. For instance, our results demonstrate a new role of Vibrionaceae in isopropanol production, and of Latilactobacillus fuchuensis and Lactococcus piscium in methanethiol/disylphide production. We also highlight a possible commensal behavior between Leuconostoc carnosum and Latilactobacillus curvatus around 2,3-butanediol metabolism.Conclusion:We conclude that our holistic approach combined with large-scale multi-omic data was a powerful strategy to prioritize the role of production parameters, already known in the literature, that shape the evolution and/or the implementation of different meat spoilage scenarios.