The metagenome analysis of complex environments with thousands of datasets, such as those in the Sequence Read Archive, requires substantial computational resources for it to be completed within a reasonable time frame. Efficient use of infrastructure is essential, and analyses must be fully reproducible with publicly available workflows to ensure transparency. Here, we introduce the Metagenomics-Toolkit, a scalable, data-agnostic workflow that automates the analysis of short and long metagenomic reads from Illumina and Oxford Nanopore Technology devices, respectively. The Metagenomics-Toolkit provides standard features such as quality control, assembly, binning, and annotation, along with unique capabilities including plasmid identification, recovery of unassembled microbial community members, and discovery of microbial interdependencies through dereplication, co-occurrence, and genome-scale metabolic modeling. Additionally, the Metagenomics-Toolkit includes a machine learning-optimized assembly step that adjusts peak RAM usage to match actual requirements, reducing the need for high-memory hardware. It can be executed on user workstations and includes optimizations for efficient cloud-based cluster execution. We compare the Metagenomics-Toolkit with five widely used metagenomics workflows and demonstrate its capabilities on 757 sewage metagenome datasets to investigate a possible sewage core microbiome. The Metagenomics-Toolkit is open source and available at https://github.com/metagenomics/metagenomics-tk.
In recent years, modern life sciences research underwent a rapid development driven mainly by the technical improvements in analytical areas leading to miniaturization, parallelization, and high throughput processing of biological samples. This has led to the generation of huge amounts of experimental data. To meet these rising demands, the German Network for Bioinformatics Infrastructure (de.NBI) was established in 2015 as a national bioinformatics consortium aiming to provide high quality bioinformatics services, comprehensive training, powerful computing capacities (de.NBI Cloud) as well as connections to the European Life Science Infrastructure ELIXIR, with the goal to assist researchers in exploring and exploiting data more effectively. Since its foundation, de.NBI Cloud has formed the scientific and collaborative backbone for new major German initiatives like NFDI or EOSC-Life in the European sector of computational biosciences. Above all, the cooperation with various NFDI consortia such as NFDI4Biodiversity, DataPLANT, GHGA, FAIRagro or NFDI4Microbiota showcases the power, range and flexibility of the de.NBI Cloud, especially for the national life science community. In conclusion, the de.NBI Cloud provides the ability to unlock the full potential of research data and enables easier collaboration across different ecosystems and research areas, which in turn enables scientists to innovate and scale-up their data-driven research, not only in the life and computational biosciences, but across the different science domains addressed by the NFDI.
Granting access to the full functionality of OpenStack’s Horizon Dashboard can overwhelm inexperienced users when accessing cloud infrastructures. Therefore, we developed SimpleVM as one of two project types available to users within the federated de.NBI Cloud. Based on OpenStack and interacting with its API, SimpleVM provides easy Life Science AAI-granted access to VMs for basic data-processing use-cases, but can also serve sophisticated, cluster-based computing and workshop scenarios.
Research on biogas-producing microbial communities aims at elucidation of correlations and dependencies between the anaerobic digestion (AD) process and the corresponding microbiome composition in order to optimize the performance of the process and the biogas output. Previously, Lachnospiraceae species were frequently detected in mesophilic to moderately thermophilic biogas reactors. To analyze adaptive genome features of a representative Lachnospiraceae strain, Anaeropeptidivorans aminofermentans M3/9T was isolated from a mesophilic laboratory-scale biogas plant and its genome was sequenced and analyzed in detail. Strain M3/9T possesses a number of genes encoding enzymes for degradation of proteins, oligo- and dipeptides. Moreover, genes encoding enzymes participating in fermentation of amino acids released from peptide hydrolysis were also identified. Based on further findings obtained from metabolic pathway reconstruction, M3/9T was predicted to participate in acidogenesis within the AD process. To understand the genomic diversity between the biogas isolate M3/9T and closely related Anaerotignum type strains, genome sequence comparisons were performed. M3/9T harbors 1,693 strain-specific genes among others encoding different peptidases, a phosphotransferase system (PTS) for sugar uptake, but also proteins involved in extracellular solute binding and import, sporulation and flagellar biosynthesis. In order to determine the occurrence of M3/9T in other environments, large-scale fragment recruitments with the M3/9T genome as a template and publicly available metagenomes representing different environments was performed. The strain was detected in the intestine of mammals, being most abundant in goat feces, occasionally used as a substrate for biogas production.
Genomic surveillance of the SARS-CoV-2 pandemic is crucial and mainly achieved by amplicon sequencing protocols. Overlapping tiled-amplicons are generated to establish contiguous SARS-CoV-2 genome sequences, which enable the precise resolution of infection chains and outbreaks. We investigated a SARS-CoV-2 outbreak in a local hospital and used nanopore sequencing with a modified ARTIC protocol employing 1200 bp long amplicons. We detected a long deletion of 168 nucleotides in the ORF8 gene in 76 samples from the hospital outbreak. This deletion is difficult to identify with the classical amplicon sequencing procedures since it removes two amplicon primer-binding sites. We analyzed public SARS-CoV-2 sequences and sequencing read data from ENA and identified the same deletion in over 100 genomes belonging to different lineages of SARS-CoV-2, pointing to a mutation hotspot or to positive selection. In almost all cases, the deletion was not represented in the virus genome sequence after consensus building. Additionally, further database searches point to other deletions in the ORF8 coding region that have never been reported by the standard data analysis pipelines. These findings and the fact that ORF8 is especially prone to deletions, make a clear case for the urgent necessity of public availability of the raw data for this and other large deletions that might change the physiology of the virus towards endemism.
ELIXIR Germany (ELIXIR-DE) is the German National Node of ELIXIR Europe. By signing the ELIXIR Collaboration Agreement on 10 February 2020, the national node ELIXIR Germany is fully established. ELIXIR Germany is coordinated from Bielefeld University and is funded by the German government. The Administration Office of ELIXIR Germany is the central support entity for the German ELIXIR Node and for the Central Coordination Unit (CCU). ELIXIR Germany actively participates in the research activities of the ELIXIR platform, i. e. Data, Training, Compute, Interoperability, and Tools Platform. A remarkable contribution is highlighted by the involvement of the two databases, SILVA and BRENDA, in ELIXIR Data Platform as ELIXIR Core Data Resources. Within the Compute Platform, the Cloud of the German Network for Bioinformatics Infrastructure (de.NBI) has built connections to the European cloud infrastructures through its involvement in the EOSC-Life project and the adaptation of ELIXIR´s Authentication and Authorization Infrastructure (AAI) system. The members of ELIXIR Germany also make a decisive contribution in the current situation in COVID-19 research. The corona pandemic is a major challenge for our society and therefore requires special attention by the established scientific structures. It is important to advance molecular biological research of coronaviruses, especially of SARS-CoV-2, and to investigate both medical and epidemiological aspects of the infection process in COVID-19. These research areas generate large amounts of data requiring intensive bioinformatics analysis. Here, ELIXIR Germany provides a wide collection of analysis programs and the compute capacities of the network's own de.NBI cloud. For example, the GALAXY project provides workflows to analyze genomics, evolution and cheminformatics data related to COVID-19 and SARS-CoV-2. Both, analysis programs and compute capacities can be used free of charge by researchers in the life sciences. Overall, ELIXIR-DE channels efforts into the three pillars of ELIXIR´s COVID-19 response: 1. Connecting national COVID-19 data platforms to create federated European COVID-19 Data Spaces; 2. Fostering good data management to make COVID-19 data open, FAIR and reusable over the long term; 3. Providing open tools, workflows and computational resources to drive reproducible and collaborative science. Moreover, ELIXIR-DE reports all on-going and emerging community and national responses to the COVID-19 outbreak to ELIXIR Europe.
The academic de.NBI Cloud offers compute resources for life science research in Germany. At the beginning of 2017, de.NBI Cloud started to implement a federated cloud consisting of five compute centers, with the aim of acting as one resource to their users. A federated cloud introduces multiple challenges, such as a central access and project management point, a unified account across all cloud sites and an interchangeable project setup across the federation. In order to implement the federation concept, de.NBI Cloud integrated with the ELIXIR authentication and authorization infrastructure system (ELIXIR AAI) and in particular Perun, the identity and access management system of ELIXIR. The integration solves the mentioned challenges and represents a backbone, connecting five compute centers which are based on OpenStack and a web portal for accessing the federation.This article explains the steps taken and software components implemented for setting up a federated cloud based on the collaboration between de.NBI Cloud and ELIXIR AAI. Furthermore, the setup and components that are described are generic and can therefore be used for other upcoming or existing federated OpenStack clouds in Europe.
The explosive growth in taxonomic metagenome profiling methods over the past years has created a need for systematic comparisons using relevant performance criteria. The Open-community Profiling Assessment tooL (OPAL) implements commonly used performance metrics, including those of the first challenge of the initiative for the Critical Assessment of Metagenome Interpretation (CAMI), together with convenient visualizations. In addition, we perform in-depth performance comparisons with seven profilers on datasets of CAMI and the Human Microbiome Project. OPAL is freely available at https://github.com/CAMI-challenge/OPAL .
The German Network for Bioinformatics Infrastructure (de.NBI) is a national and academic infrastructure funded by the German Federal Ministry of Education and Research (BMBF). The de.NBI provides (i) service, (ii) training, and (iii) cloud computing to users in life sciences research and biomedicine in Germany and Europe and (iv) fosters the cooperation of the German bioinformatics community with international network structures. The de.NBI members also run the German node (ELIXIR-DE) within the European ELIXIR infrastructure. The de.NBI / ELIXIR-DE training platform, also known as special interest group 3 (SIG 3) ‘Training & Education’, coordinates the bioinformatics training of de.NBI and the German ELIXIR node. The network provides a high-quality, coherent, timely, and impactful training program across its eight service centers. Life scientists learn how to handle and analyze biological big data more effectively by applying tools, standards and compute services provided by de.NBI. Since 2015, more than 300 training courses were carried out with about 6,000 participants and these courses received recommendation rates of almost 90% (status as of July 2020). In addition to face-to-face training courses, online training was introduced on the de.NBI website in 2016 and guidelines for the preparation of e-learning material were established in 2018. In 2016, ELIXIR-DE joined the ELIXIR training platform. Here, the de.NBI / ELIXIR-DE training platform collaborates with ELIXIR in training activities, advertising training courses via TeSS and discussions on the exchange of data for training events essential for quality assessment on both the technical and administrative levels. The de.NBI training program trained thousands of scientists from Germany and beyond in many different areas of bioinformatics.
The academic de.NBI Cloud offers compute resources for life science research in Germany. At the beginning of 2017, de.NBI Cloud started to implement a federated cloud consisting of five compute centers, with the aim of acting as one resource to their users. A federated cloud introduces multiple challenges, such as a central access and project management point, a unified account across all cloud sites and an interchangeable project setup across the federation. In order to implement the federation concept, de.NBI Cloud integrated with the ELIXIR authentication and authorization infrastructure system (ELIXIR AAI) and in particular Perun, the identity and access management system of ELIXIR. The integration solves the mentioned challenges and represents a backbone, connecting five compute centers which are based on OpenStack and a web portal for accessing the federation.This article explains the steps taken and software components implemented for setting up a federated cloud based on the collaboration between de.NBI Cloud and ELIXIR AAI. Furthermore, the setup and components that are described are generic and can therefore be used for other upcoming or existing federated OpenStack clouds in Europe.
A common Authentication and Authorisation Infrastructure (AAI) that would allow single sign-on to services has been identified as a key enabler for European bioinformatics. ELIXIR AAI is an ELIXIR service portfolio for authenticating researchers to ELIXIR services and assisting these services on user privileges during research usage. It relieves the scientific service providers from managing the user identities and authorisation themselves, enables the researcher to have a single set of credentials to all ELIXIR services and supports meeting the requirements imposed by the data protection laws. ELIXIR AAI was launched in late 2016 and is part of the ELIXIR Compute platform portfolio. By the end of 2017 the number of users reached 1000, while the number of relying scientific services was 36. This paper presents the requirements and design of the ELIXIR AAI and the policies related to its use, and how it can be used for serving some example services, such as document management, social media, data discovery, human data access, cloud compute and training services.
Reconstructing the genomes of microbial community members is key to the interpretation of shotgun metagenome samples. Genome binning programs deconvolute reads or assembled contigs of such samples into individual bins. However, assessing their quality is difficult due to the lack of evaluation software and standardized metrics. Here, we present Assessment of Metagenome BinnERs (AMBER), an evaluation package for the comparative assessment of genome reconstructions from metagenome benchmark datasets. It calculates the performance metrics and comparative visualizations used in the first benchmarking challenge of the initiative for the Critical Assessment of Metagenome Interpretation (CAMI). As an application, we show the outputs of AMBER for 11 binning programs on two CAMI benchmark datasets. AMBER is implemented in Python and available under the Apache 2.0 license on GitHub.
In metagenome analysis, computational methods for assembly, taxonomic profiling and binning are key components facilitating downstream biological data interpretation. However, a lack of consensus about benchmarking datasets and evaluation metrics complicates proper performance assessment. The Critical Assessment of Metagenome Interpretation (CAMI) challenge has engaged the global developer community to benchmark their programs on datasets of unprecedented complexity and realism. Benchmark metagenomes were generated from ~700 newly sequenced microorganisms and ~600 novel viruses and plasmids, including genomes with varying degrees of relatedness to each other and to publicly available ones and representing common experimental setups. Across all datasets, assembly and genome binning programs performed well for species represented by individual genomes, while performance was substantially affected by the presence of related strains. Taxonomic profiling and binning programs were proficient at high taxonomic ranks, with a notable performance decrease below the family level. Parameter settings substantially impacted performances, underscoring the importance of program reproducibility. While highlighting current challenges in computational metagenomics, the CAMI results provide a roadmap for software selection to answer specific research questions.
Background: The production of biogas takes place under anaerobic conditions and involves microbial decomposition of organic matter. Most of the participating microbes are still unknown and non-cultivable. Accordingly, shotgun metagenome sequencing currently is the method of choice to obtain insights into community composition and the genetic repertoire.Findings: Here, we report on the deeply sequenced metagenome and metatranscriptome of a complex biogas-producing microbial community from an agricultural production-scale biogas plant. We assembled the metagenome and, as an example application, show that we reconstructed most genes involved in the methane metabolism, a key pathway involving methanogenesis performed by methanogenic Archaea. This result indicates that there is sufficient sequencing coverage for most downstream analyses.Conclusions: Sequenced at least one order of magnitude deeper than previous studies, our metagenome data will enable new insights into community composition and the genetic potential of important community members. Moreover, mapping of transcripts to reconstructed genome sequences will enable the identification of active metabolic pathways in target organisms.
A standard for creating interchangeable bioinformatics software containers Software has proliferated in bioinformatics and so have the problems associated with it: missing or unobtainable code, difficult to install dependencies, unreproducible workflows, all with terrible user experiences. We have created bioboxes with the aim to make accessing and using bioinformatics software more simple and more easy. http://bioboxes.org/
We introduce the open-source community "bioboxes" which has the aim of simplifying bioinformatics tools through adopting common interfaces. The Docker project makes installing, running and reproducing the output of an application easier because all the dependencies can be provided by creating a "container" of the application. The bioboxes project furthers this concept by creating an interface standard so that containerised software of the same biobox type can be seamlessly interchanged. Feedback and contributions of new biobox interfaces are managed using the Github issue system. Bioboxes provides documentation and software to help developers follow this standard. The bioboxes.org website provides instructions and guide on how to build a biobox. We also provide a validator tool that tests whether a container follows a biobox interface and thereby helps the developer create their own bioboxes. Furthermore we provide bioboxes in Github and our public Docker Hub repository for download and use. Researchers with a whole selection of bioboxes at their fingertips can comprehensively evaluate many similar tools that have the same interface and improve the quality of their research. A system of well-defined bioboxes is an important step towards the making provenance easier and research more shareable and reproducible.
Software is now both central and essential to modern biology, yet lack of availability, difficult installations, and complex user interfaces make software hard to obtain and use. Containerisation, as exemplified by the Docker platform, has the potential to solve the problems associated with sharing software. We propose bioboxes: containers with standardised interfaces to make bioinformatics software interchangeable.
Alexander Sczyrba合作论文数Universität Bielefeld
Technische Fakultät
AG Praktische Informatik11