The HMMER web server, available at https://www.ebi.ac.uk/Tools/hmmer, provides online access to tools from the HMMER software suite (http://hmmer.org/) for protein analysis using profile hidden Markov models. Users can perform sequence similarity searches against a range of regularly updated protein sequence databases or annotate protein sequences with domains and families using profile HMM libraries from protein family databases. Since the 2018 update, the continued exponential growth of sequence databases has necessitated substantial infrastructural improvements to maintain search performance speed and service reliability. To achieve this, the web interface has been completely reengineered using modern web technologies (JavaScript and React), providing users with an enhanced experience, including session-based search history and streamlined results visualization. The web application programming interface has been rewritten to better support programmatic access with updated endpoints and JSON-based responses. The infrastructure has been redesigned to efficiently handle searches against much larger databases through horizontal scaling and asynchronous job processing. Target database offerings have been updated to reflect current usage patterns and data availability. The HMMER web server is free and open to all users, and there is no login requirement.
Multi-omics datasets are an increasingly prevalent and necessary resource for achieving scientificadvances in microbial ecosystem research. However, they present twin challenges to researchinfrastructures: firstly the utility of multi-omics datasets relies entirely on interoperability ofomics layers, i.e. on formalised data linking. Secondly, microbiome derived data typically leadto computationally expensive analyses, and so rely on the availability of high performancecomputing (HPC) or cloud infrastructures. These challenges can be better met by combining the resources of multiple groups, services and infrastructures. In this BioHackathon Europe 2024 project, we envisioned a "federated microbiome analysis service" and worked on three tracks of development towards it: mapping metagenomics metadata standards to Schema.org and Bioschemas terms, rendering Nextflow workflow executions as RO-Crates, and tooling for creating, viewing and interlinking human-readable RO-Crate previews.
SUMMARY:Holo-omics is an emerging research area that integrates multi-omic datasets from the host organism and its microbiome to study their interactions. Recently, curated and openly accessible holo-omic databases have been developed. The HoloFood database, for instance, provides nearly 10 000 holo-omic profiles for salmon and chicken under controlled treatments. However, bridging the gap between holo-omic data resources and algorithmic frameworks remains a challenge. Combining the latest advances in statistical programming with curated holo-omic data sets can facilitate the design of open and reproducible research workflows in the emerging field of holo-omics. AVAILABILITY AND IMPLEMENTATION:HoloFoodR R/Bioconductor package and the source code are available under the open-source Artistic License 2.0 at the package homepage https://doi.org/10.18129/B9.bioc.HoloFoodR.
The HoloFood project used a hologenomic approach to understand the impact of host-microbiota interactions on salmon and chicken production by analysing multiomic data, phenotypic characteristics, and associated metadata in response to novel feeds. The project's raw data, derived analyses, and metadata are deposited in public, open archives (BioSamples, European Nucleotide Archive, MetaboLights, and MGnify), so making use of these diverse data types may require access to multiple resources. This is especially complex where analysis pipelines produce derived outputs such as functional profiles or genome catalogues. The HoloFood Data Portal is a web resource that simplifies access to the project datasets. For example, users can conveniently access multiomic datasets derived from the same individual or retrieve host phenotypic data with a linked gut microbiome sample. Project-specific metagenome-assembled genome and viral catalogues are also provided, linking to broader datasets in MGnify. The portal stores only data necessary to provide these relationships, with possible linking to the underlying repositories. The portal showcases a model approach for how future multiomics datasets can be made available. Database URL: https://www.holofooddata.org.
The availability of public metaproteomics, metagenomics and metatranscriptomics data in public resources such as MGnify (for metagenomics/metatranscriptomics) and the PRIDE database (for metaproteomics), continues to increase. When these omics techniques are applied to the same samples, their integration offers new opportunities to understand the structure (metagenome) and functional expression (metatranscriptome and metaproteome) of the microbiome. Here, we describe a pilot study aimed at integrating public multi-meta-omics datasets from studies based on human gut and marine hatchery samples. Reference search databases (search DBs) were built using assembled metagenomic (and metatranscriptomic, where available) sequence data followed by de novo gene calling, using both data from the same sampling event and from independent samples. The resulting protein sets were evaluated for their utility in metaproteomics analysis. In agreement with previous studies, the highest number of peptide identifications was generally obtained when using search DBs created from the same samples. Data integration of the multi-omics results was performed in MGnify. For that purpose, the MGnify website was extended to enable the visualisation of the resulting peptide/protein information from three reanalysed metaproteomics datasets. A workflow (https://github.com/PRIDE-reanalysis/MetaPUF) has been developed allowing researchers to perform equivalent data integration, using paired multi-omics datasets. This is the first time that a data integration approach for multi-omics datasets has been implemented from public data available in the world-leading MGnify and PRIDE resources.
Resolving the microbiome of the Atlantic salmon Salmo salar gut is challenged by a low microbial diversity often dominated by one or two species of bacteria, and high levels of host contamination in sequencing data. Nevertheless, existing metabarcoding and metagenomic studies consistently resolve a putative beneficial Mycoplasma species as the most abundant organism in gut samples. The remaining microbiome is heavily influenced by factors such as developmental stage and water salinity. We profiled the salmon gut microbiome across 540 salmon samples in differing conditions with a view to capture the genomic diversity that can be resolved from the salmon gut. The salmon were exposed to 3 different nutritional additives: seaweed, blue mussel protein and silaged blue mussel protein, including both pre-smolts (30-60 g salmon reared in freshwater) as well as post-smolts (300–600 g salmon reared in saltwater). Using genome-resolved metagenomics, we generated a catalogue of 11 species-level bacterial MAGs from 188 input metagenome assembled genomes, with 5 species not found in other catalogues. This highlights that our understanding of salmon gut microbial diversity is still incomplete. A prevalent bacterial genome annotated as Mycoplasmoidaceae is present in adult fish, and a comparison of functions revealed significant sub-species variation. Juvenile fish have a different microbial diversity, dominated by a species of Pseudomonas aeruginosa. We also present the first viral catalogue for salmon including prophage sequences which can be linked to the bacterial MAGs.
Natural products biosynthesised by microbes are an important component of the pharmacopeia with a vast array of biomedical and industrial applications, in addition to their key role in mediating many ecological interactions. One approach for the discovery of these metabolites is the identification of biosynthetic gene clusters (BGCs), genomic units which encode the molecular machinery required for producing the natural product. Genome mining has revolutionised the discovery of BGCs, yet metagenomic assemblies represent a largely untapped source of natural products. The imbalanced distribution of BGC classes in existing databases restricts the generalisation of detection patterns and limits the ability of mining methods to recognise a broader spectrum of BGCs. This problem is further intensified in metagenomic datasets, where BGC genes may be incomplete. This work presents SanntiS, a new machine learning-based tool for identifying BGCs. SanntiS achieved high precision and recall in both genomic and metagenomic datasets, effectively capturing a broad range of BGCs. Application of SanntiS to MGnify metagenomic assemblies led to a resource containing 1.9 million BGC predictions with associated contextual data from diverse biomes and demonstrates a significant fraction of novelty compared to equivalent isolate genomes datasets. Subsequent experimental validation of a novel antimicrobial peptide detected solely by SanntiS, further demonstrates the potential of this approach for uncovering novel bioactive compounds.
The MGnify platform (https://www.ebi.ac.uk/metagenomics) facilitates the assembly, analysis and archiving of microbiome-derived nucleic acid sequences. The platform provides access to taxonomic assignments and functional annotations for nearly half a million analyses covering metabarcoding, metatranscriptomic, and metagenomic datasets, which are derived from a wide range of different environments. Over the past 3 years, MGnify has not only grown in terms of the number of datasets contained but also increased the breadth of analyses provided, such as the analysis of long-read sequences. The MGnify protein database now exceeds 2.4 billion non-redundant sequences predicted from metagenomic assemblies. This collection is now organised into a relational database making it possible to understand the genomic context of the protein through navigation back to the source assembly and sample metadata, marking a major improvement. To extend beyond the functional annotations already provided in MGnify, we have applied deep learning-based annotation methods. The technology underlying MGnify's Application Programming Interface (API) and website has been upgraded, and we have enabled the ability to perform downstream analysis of the MGnify data through the introduction of a coupled Jupyter Lab environment.
MGnify is EMBL-EBI's metagenomics resource. MGnify's recently launched Notebook Server provides an online Jupyter Lab environment for users to explore programmatic access to MGnify's datasets using Python or R. Here, we report several developments to the Notebook Server completed during the BioHackathon Europe 2022. The developments range from establishing an instance of the notebooks on the Galaxy platform, to adding new notebooks and Jupyter UI extensions enabling more users to perform downstream analysis tasks on MGnify's extensive metagenomics datasets.
An increasingly common output arising from the analysis of shotgun metagenomic datasets is the generation of metagenome-assembled genomes (MAGs), with tens of thousands of MAGs now described in the literature. However, the discovery and comparison of these MAG collections is hampered by the lack of uniformity in their generation, annotation and storage. To address this, we have developed MGnify Genomes, a growing collection of biome-specific non-redundant microbial genome catalogues generated using MAGs and publicly available isolate genomes. Genomes within a biome-specific catalogue are organised into species clusters. For species that contain multiple conspecific genomes, the highest quality genome is selected as the representative, always prioritising an isolate genome over a MAG. The species representative sequences and annotations can be visualised on the MGnify website and the full catalogue and associated analysis outputs can be downloaded from MGnify servers. A suite of online search tools is provided allowing users to compare their own sequences, ranging from a gene to sets of genomes, against the catalogues. Seven biomes are available currently, comprising over 300,000 genomes that represent 11,048 non-redundant species, and include 36 taxonomic classes not currently represented by cultured genomes. MGnify Genomes is available at https://www.ebi.ac.uk/metagenomics/browse/genomes/.
Metagenomics is a culture-independent method for studying the microbes inhabiting a particular environment. Comparing the composition of samples (functionally/taxonomically), either from a longitudinal study or cross-sectional studies, can provide clues into how the microbiota has adapted to the environment. However, a recurring challenge, especially when comparing results between independent studies, is that key metadata about the sample and molecular methods used to extract and sequence the genetic material are often missing from sequence records, making it difficult to account for confounding factors. Nevertheless, these missing metadata may be found in the narrative of publications describing the research. Here, we describe a machine learning framework that automatically extracts essential metadata for a wide range of metagenomics studies from the literature contained in Europe PMC. This framework has enabled the extraction of metadata from 114,099 publications in Europe PMC, including 19,900 publications describing metagenomics studies in European Nucleotide Archive (ENA) and MGnify. Using this framework, a new metagenomics annotations pipeline was developed and integrated into Europe PMC to regularly enrich up-to-date ENA and MGnify metagenomics studies with metadata extracted from research articles. These metadata are now available for researchers to explore and retrieve in the MGnify and Europe PMC websites, as well as Europe PMC annotations API.
We present the results of a study investigating the sizes and morphologies of redshift 4 < z < 8 galaxies in the CANDELS (Cosmic Assembly Near-infrared Deep Extragalactic Legacy Survey) GOODS-S (Great Observatories Origins Deep Survey southern field), HUDF (Hubble Ultra-Deep Field) and HUDF parallel fields. Based on non-parametric measurements and incorporating a careful treatment of measurement biases, we quantify the typical size of galaxies at each redshift as the peak of the lognormal size distribution, rather than the arithmetic mean size. Parametrizing the evolution of galaxy half-light radius as r(50) alpha (1 + z)(n), we find n = -0.20 +/- 0.26 at bright UV-luminosities (0.3L(*(z=3)) < L < L-*) and n = -0.47 +/- 0.62 at faint luminosities (0.12L(*) < L < 0.3L(*)). Furthermore, simulations based on artificially redshifting our z similar to 4 galaxy sample show that we cannot reject the null hypothesis of no size evolution. We show that this result is caused by a combination of the size-dependent completeness of high-redshift galaxy samples and the underestimation of the sizes of the largest galaxies at a given epoch. To explore the evolution of galaxy morphology we first compare asymmetry measurements to those from a large sample of simulated single Sersic profiles, in order to robustly categorize galaxies as either 'smooth' or 'disturbed'. Comparing the disturbed fraction amongst bright (M-1500 <= -20) galaxies at each redshift to that obtained by artificially redshifting our z similar to 4 galaxy sample, while carefully matching the size and UV-luminosity distributions, we find no clear evidence for evolution in galaxy morphology over the redshift interval 4 < z < 8. Therefore, based on our results, a bright (M-1500 <= -20) galaxy at z similar to 6 is no more likely to be measured as 'disturbed' than a comparable galaxy at z similar to 4, given the current observational constraints.
We present the results of a new search for bright star-forming galaxies at z ~ 7 within the UltraVISTA DR2 and UKIDSS UDS DR10 data, which together provide 1.65 sq deg of near-infrared imaging with overlapping optical and Spitzer data. Using a full photo-z analysis to identify high-z galaxies and reject contaminants, we have selected a sample of 34 luminous (-22.7 < M_UV < -21.2) galaxies with 6.5 < z < 7.5. Crucially, the deeper imaging provided by UltraVISTA DR2 confirms all of the robust objects previously uncovered by Bowler et al. (2012), validating our selection technique. Our sample includes the most massive galaxies known at z ~ 7, with M_* ~ 10^{10} M_sun, and the majority are resolved, consistent with larger sizes (r_{1/2} ~ 1 - 1.5 kpc) than displayed by less massive galaxies. From our final sample, we determine the form of the bright end of the rest-frame UV galaxy luminosity function (LF) at z ~ 7, providing strong evidence that the bright end of the z = 7 LF does not decline as steeply as predicted by the Schechter function fitted to fainter data. We consider carefully, and exclude the possibility that this is due to either gravitational lensing, or significant contamination of our galaxy sample by AGN. Rather, our results favour a double power-law form for the galaxy LF at high z, or, more interestingly, a LF which simply follows the form of the dark-matter halo mass function at bright magnitudes. This suggests that the physical mechanism which inhibits star-formation activity in massive galaxies (i.e. AGN feedback or some other form of `mass quenching') has yet to impact on the observable galaxy LF at z ~ 7, a conclusion supported by the estimated masses of our brightest galaxies which have only just reached a mass comparable to the critical `quenching mass' of M_* = 10 ^{10.2} M_sun derived from studies of the mass function of star-forming galaxies at lower z.
We present the results of a study investigating the sizes and morphologies of redshift 4 < z < 8 galaxies in the CANDELS GOODS-S, HUDF and HUDF parallel fields. Based on non-parametric measurements and incorporating a careful treatment of measurement biases, we quantify the typical size of galaxies at each redshift as the peak of the log-normal size distribution, rather than the arithmetic mean size. Parameterizing the evolution of galaxy half-light radius as $r_{50} \propto (1+z)^n$, we find $n = -0.20 \pm 0.26$ at bright UV-luminosities ($0.3L_{*(z=3)} < L < L_*$) and $n = -0.47 \pm 0.62$ at faint luminosities ($0.12L_* < L < 0.3L_*$). Furthermore, simulations based on artificially redshifting our z~4 galaxy sample show that we cannot reject the null hypothesis of no size evolution. We show that this result is caused by a combination of the size-dependent completeness of high-redshift galaxy samples and the underestimation of the sizes of the largest galaxies at a given epoch. To explore the evolution of galaxy morphology we first compare asymmetry measurements to those from a large sample of simulated single Sersic profiles, in order to robustly categorise galaxies as either `smooth' or `disturbed'. Comparing the disturbed fraction amongst bright ($M_{UV} \leq -20$) galaxies at each redshift to that obtained by artificially redshifting our z~4 galaxy sample, while carefully matching the size and UV-luminosity distributions, we find no clear evidence for evolution in galaxy morphology over the redshift interval 4 < z < 8. Therefore, based on our results, a bright ($M_{UV} \leq -20$) galaxy at z~6 is no more likely to be measured as `disturbed' than a comparable galaxy at z~4, given the current observational constraints.
We present the results of a study investigating the rest-frame ultra-violet (UV) spectral slopes of redshift z 5 Lyman-break galaxies (LBGs). By combining deep Hubble Space Telescope imaging of the CANDELS and HUDF fields with ground-based imaging from the UKIDSS Ultra Deep Survey (UDS), we have produced a large sample of z 5 LBGs spanning an unprecedented factor of >100 in UV luminosity. Based on this sample we find a clear colour-magnitude relation (CMR) at z 5, such that the rest-frame UV slopes (beta) of brighter galaxies are notably redder than their fainter counterparts. We determine that the z 5 CMR is well described by a linear relationship of the form: d beta = (-0.12 +/- 0.02) d Muv, with no clear evidence for a change in CMR slope at faint magnitudes (i.e. Muv > -18.9). Using the results of detailed simulations we are able, for the first time, to infer the intrinsic (i.e. free from noise) variation of galaxy colours around the CMR at z 5. We find significant (12 sigma) evidence for intrinsic colour variation in the sample as a whole. Our results also demonstrate that the width of the intrinsic UV slope distribution of z 5 galaxies increases from Delta(beta)=0.1 at Muv=-18 to Delta(beta)=0.4 at Muv=-21. We suggest that the increasing width of the intrinsic galaxy colour distribution and the CMR itself are both plausibly explained by a luminosity independent lower limit of beta=-2.1, combined with an increase in the fraction of red galaxies in brighter UV-luminosity bins.
We present the results of a study investigating the sizes and morphologies of redshift 4
We use the new ultra-deep, near-infrared imaging of the Hubble Ultra-Deep Field (HUDF) provided by our UDF12 Hubble Space Telescope (HST) Wide Field Camera 3/IR campaign to explore the rest-frame ultraviolet (UV) properties of galaxies at redshifts z > 6.5. We present the first unbiased measurement of the average UV power-law index, , (f(lambda) alpha lambda(beta)) for faint galaxies at z similar or equal to 7, the first meaningful measurements of at z similar or equal to 8, and tentative estimates for a new sample of galaxies at z similar or equal to 9. Utilizing galaxy selection in the new F140W (J(140)) imaging to minimize colour bias, and applying both colour and power-law estimators of beta, we find = -2.1 +/- 0.2 at z similar or equal to 7 for galaxies with M-UV similar or equal to -18. This means that the faintest galaxies uncovered at this epoch have, on average, UV colours no more extreme than those displayed by the bluest star-forming galaxies at low redshift. At z similar or equal to 8 we find a similar value, = -1.9 +/- 0.3. At z similar or equal to 9, we find = -1.8 +/- 0.6, essentially unchanged from z similar or equal to 6 to 7 (albeit highly uncertain). Finally, we show that there is as yet no evidence for a significant intrinsic scatter in beta within our new, robust z similar or equal to 7 galaxy sample. Our results are most easily explained by a population of steadily star-forming galaxies with either similar or equal to solar metallicity and zero dust, or moderately sub-solar (similar or equal to 10-20 per cent) metallicity with modest dust obscuration (A(V) similar or equal to 0.1-0.2). This latter interpretation is consistent with the predictions of a state-of-the-art galaxy-formation simulation, which also suggests that a significant population of very-low metallicity, dust-free galaxies with beta similar or equal to -2.5 may not emerge until M-UV > -16, a regime likely to remain inaccessible until the James Webb Space Telescope.
We present a new determination of the ultraviolet (UV) galaxy luminosity function (LF) at redshift z similar or equal to 7 and 8, and a first estimate at z similar or equal to 9. An accurate determination of the form and evolution of the galaxy LF during this era is of key importance for improving our knowledge of the earliest phases of galaxy evolution and the process of cosmic reionization. Our analysis exploits to the full the new, deepest Wide Field Camera 3/infrared imaging from our Hubble Space Telescope (HST) Ultra-Deep Field 2012 (UDF12) campaign, with dynamic range provided by including a new and consistent analysis of all appropriate, shallower/wider area HST survey data. Our new measurement of the evolving LF at z similar or equal to 7 to 8 is based on a final catalogue of similar or equal to 600 galaxies, and involves a step-wise maximum-likelihood determination based on the photometric redshift probability distribution for each object; this approach makes full use of the 11-band imaging now available in the Hubble Ultra-Deep Field (HUDF), including the new UDF12 F140W data, and the latest Spitzer IRAC imaging. The final result is a determination of the z similar or equal to 7 LF extending down to UV absolute magnitudes M-1500 = -16.75 (AB mag) and the z similar or equal to 8 LF down to M-1500 = -17.00. Fitting a Schechter function, we find M-1500(*) = -19.90(-0.28)(+0.23), log phi* = -2.96(-0.23)(+0.18) and a faint-end slope alpha = -1.90(-0.15)(+0.14) at z similar or equal to 7, and M-1500* = -20.12(-0.48)(+0.37), log phi* = -3.35(-0.47)(+0.28) and alpha = -2.02(+0.23)(+0.22) at z similar or equal to 8. These results strengthen previous suggestions that the evolution at z > 7 appears more akin to 'density evolution' than the apparent 'luminosity evolution' seen at z similar or equal to 5 - 7. We also provide the first meaningful information on the LF at z similar or equal to 9, explore alternative extrapolations to higher redshifts, and consider the implications for the early evolution of UV luminosity density. Finally, we provide catalogues (including derived z(phot), M-1500 and photometry) for the most robust z similar to 6.5-11.9 galaxies used in this analysis. We briefly discuss our results in the context of earlier work and the results derived from an independent analysis of the UDF12 data based on colour-colour selection.