Missing data is a long-standing issue in phylogenetic inference, which often results in high levels of taxonomic instability, obscuring otherwise well-supported relationships. Multiple approaches have been developed to deal with the negative effects of ineffective overlap on tree resolution, often by identifying taxa for removal. Here, we repurpose a heuristic method developed to identify unstable taxa in morphological data matrices, concatabominations, and combine it with a novel gene-tree jackknifing on matrix representation of trees to identify candidates for targeted sequencing. Using a multilocus caecilian data set, we illustrate the method's capacity to identify candidate taxa and loci for additional sequencing, compare the results with those of the mathematics-based gene sampling sufficiency approach, and explore the terrace space associated with the multilocus data set. We show that our approach yields tractable numbers of loci/taxa for targeted sequencing that successfully mitigate topological instability due to ineffective overlap, even when modest amounts of data are added.
Sepsis caused by multidrug-resistant (MDR) bacteria, particularly Escherichia coli, represents a significant global health threat due to high morbidity, mortality, and limited treatment options. E. coli, a major causative agent of bloodstream infections, has evolved highly virulent and MDR strains, which contribute to the increasing burden of antimicrobial resistance (AMR), complicating clinical management and reducing the efficacy of conventional antibiotic therapies. This study highlights the negative implications of MDR E. coli sepsis through the genomic and phenotypic characterization of E. coli 266631E isolated from a sepsis patient. Antimicrobial susceptibility testing and minimum inhibitory concentration analysis revealed resistance to multiple antibiotics, including amoxicillin, cefotaxime, ciprofloxacin, gentamicin, and tobramycin. Whole genome sequencing identified a broad array of AMR genes encoding resistance to various antibiotic classes, such as macrolides, fluoroquinolones, aminoglycosides, carbapenems, and cephalosporins. Notably, the CTX-M-15 gene, a key extended-spectrum β-lactamase determinant, was found in both the bacterial chromosome and an IncF-type plasmid, emphasizing the potential for horizontal gene transfer and rapid dissemination of resistance. Confirming the taxonomy of the isolate through querying its 16S rRNA sequence and genome in recognised bacterial taxonomic databases presented a challenge. The isolate showed genetic similarity to E. coli, E. fergusonii, and Shigella species despite their phenotypic differences and variations in their pathogenic traits. However, a simple phenotypic laboratory procedure, based on the biochemical and cultural differences among these bacteria in Coliform ChromoSelect Agar, confirmed the isolate as E. coli. This study highlights the importance of applying phenotypic approaches to support genomic identification of clinically significant bacteria, as this will help guide proper diagnosis and treatment of life-threatening infections like sepsis. ### Competing Interest Statement The authors have declared no competing interest.
The prevalence and abundance of Acidobacteriota raise concerns about their ecological function and metabolic activity in the environment. Studies have reported the potential of some members of Acidobacteriota to interact with plants and play a significant role in biogeochemical cycles. However, their role in this context has not been extensively studied. Here, we performed a comprehensive genomic analysis of 758 metagenome-assembled genomes (MAGs). Our analysis revealed a high frequency of plant growth-promoting traits (PGPTs) genes in the Acidobacteriaceae, Bryobacteraceae, Koribacteraceae, and Pyrinomonadaceae families. The colonization and competitive exclusion classes of PGPTs were found to be present in numerous Acidobacteriota members. In addition, these PGPTs also include genes involved in nitrogen fixation, phosphorus solubilization, exopolysaccharide production, siderophore production, and plant growth hormone production. Expression of such genes was found to be transcriptionally active in different environments. In addition, we identified numerous carbohydrate-active enzymes and peptidases involved in plant polymer degradation. By applying an in-depth insight into the diversity of the phylum, we expand the understanding of the role played by Acidobacteriota in the cycling of carbon, nitrogen, sulfur, and trace metals. Together, this study underscores the distinct potential ecological roles of each of these taxonomic groups, providing valuable insights for future research.
Background: Robust methods to track pathogens support public health surveillance. Both wastewater (WW) and individual whole genome sequencing (WGS) are used to assess viral variant diversity and spread. However, their relative performance and the information provided by each approach have not been sufficiently quantified. Therefore, we conducted a comparative evaluation using extensive individual and wastewater longitudinal SARS-CoV-2 WGS datasets in Northern Ireland (NI). Methods: WGS of SARS-CoV-2 was performed on >4k WW samples and >23k individuals across NI from 14th November 2021 to 11th March 2023. SARS-CoV-2 RNA was amplified using the ARTIC nCov-2019 protocol and sequenced on an Illumina MiSeq. Wastewater data were analysed using Freyja to determine variant compositions, which were compared to individual data through time series and correlation analyses. Inter-programme agreements were evaluated by mean absolute error (MAE) calculations. WW treatment plant (WWTP) performances were ranked by mean MAE. Volatile periods were identified using numerical derivative analyses. Geospatial spreading patterns were determined by horizontal curve shifting. Findings: Strong concordance was observed between wastewater and individual variant compositions and distributions, influenced by sequencing rate and variant diversity. Overall variant compositions derived from individual sequences and each WWTP were regionally clustered rather than dominated by local population size. Both individual and WW sequencing detected common nucleotide substitutions across many variants and complementary additional substitutions. Conserved spreading patterns were identified using both approaches. Interpretation: Both individual and wastewater WGS effectively monitor SARS-CoV-2 variant dynamics. Combining these approaches enhances confidence in predicting the composition and spread of major variants, particularly with higher sequencing rates. Each method detects unique mutations, and their integration improves overall genome surveillance. ### Competing Interest Statement JWM and DFG are directors of BioSeer Ltd (NI701440), a UK company that has the potential to offer wastewater testing services. ### Funding Statement Funding: Individual sequencing was funded via the Belfast Health and Social Care Trust (Department of Health for Northern Ireland) and the COVID-19 Genomics UK (COG-UK) consortium, which was supported by the Medical Research Council (MRC), UK Research and Innovation (UKRI), the National Institute for Health Research (NIHR), the Department of Health and Social Care (DHSC), and the Wellcome Sanger Institute. The NI Wastewater Surveillance Programme was funded by the Department of Health for Northern Ireland. EPT was supported through the COG-UK Early Career Funding Scheme. ### Author Declarations I confirm all relevant ethical guidelines have been followed, and any necessary IRB and/or ethics committee approvals have been obtained. Yes The details of the IRB/oversight body that provided approval or exemption for the research described are given below: The COG-UK study protocol was approved by the Public Health England Research Ethics Governance Group. (Reference: R&D NR0195) I confirm that all necessary patient/participant consent has been obtained and the appropriate institutional forms have been archived, and that any patient/participant/sample identifiers included were not known to anyone (e.g., hospital staff, patients or participants themselves) outside the research group so cannot be used to identify individuals. Yes I understand that all clinical trials and any other prospective interventional studies must be registered with an ICMJE-approved registry, such as ClinicalTrials.gov. I confirm that any such study reported in the manuscript has been registered and the trial registration ID is provided (note: if posting a prospective study registered retrospectively, please provide a statement in the trial ID field explaining why the study was not registered in advance). Yes I have followed all appropriate research reporting guidelines, such as any relevant EQUATOR Network research reporting checklist(s) and other pertinent material, if applicable. Yes All data produced in the present study are available upon reasonable request to the authors.
The rumen represents a dynamic microbial ecosystem where fermentation metabolites and microbial concentrations change over time in response to dietary changes. The integration of microbial genomic knowledge and dynamic modelling can enhance our system-level understanding of rumen ecosystem’s function. However, such an integration between dynamic models and rumen microbiota data is lacking. The objective of this work was to integrate rumen microbiota time series determined by 16S rRNA gene amplicon sequencing into a dynamic modelling framework to link microbial data to the dynamics of the volatile fatty acids (VFA) production during fermentation. For that, we used the theory of state observers to develop a model that estimates the dynamics of VFA from the data of microbial functional proxies associated with the specific production of each VFA. We determined the microbial proxies using CowPi to infer the functional potential of the rumen microbiota and extrapolate their functional modules from KEGG (Kyoto Encyclopedia of Genes and Genomes). The approach was challenged using data from an in vitro RUSITEC experiment and from an in vivo experiment with four cows. The model performance was evaluated by the coefficient of variation of the root mean square error (CRMSE). For the in vitro case study, the mean CVRMSE were 9.8% for acetate, 14% for butyrate and 14.5% for propionate. For the in vivo case study, the mean CVRMSE were 16.4% for acetate, 15.8% for butyrate and 19.8% for propionate. The mean CVRMSE for the VFA molar fractions were 3.1% for acetate, 3.8% for butyrate and 8.9% for propionate. Ours results show the promising application of state observers integrated with microbiota time series data for predicting rumen microbial metabolism.
A deeper understanding of the relationship between the antimicrobial resistance (AMR) gene carriage and phenotype is necessary to develop effective response strategies against this global burden. AMR phenotype is often a result of multi-gene interactions; therefore, we need approaches that go beyond current simple AMR gene identification tools. Machine-learning (ML) methods may meet this challenge and allow the development of rapid computational approaches for AMR phenotype classification. To examine this, we applied multiple ML techniques to 16,950 bacterial genomes across 28 genera, with corresponding MICs for 23 antibiotics with the aim of training models to accurately determine the AMR phenotype from sequenced genomes. This resulted in a >1.5-fold increase in AMR phenotype prediction accuracy over AMR gene identification alone. Furthermore, we revealed 528 unique (often species-specific) genomic routes to antibiotic resistance, including genes not previously linked to the AMR phenotype. Our study demonstrates the utility of ML in predicting AMR phenotypes across diverse clinically relevant organisms and antibiotics. This research proposes a rapid computational method to support laboratory-based identification of the AMR phenotype in pathogens.
The rumen ecosystem harbours a galaxy of microbes working in syntrophy to carry out a metabolic cascade of hydrolytic and fermentative reactions. This fermentation process allows ruminants to harvest nutrients from a wide range of feedstuff otherwise inaccessible to the host. The interconnection between the ruminant and its rumen microbiota shapes key animal phenotypes such as feed efficiency and methane emissions and suggests the potential of reducing methane emissions and enhancing feed conversion into animal products by manipulating the rumen microbiota. Whilst significant technological progress in omics techniques has increased our knowledge of the rumen microbiota and its genome (microbiome), translating omics knowledge into effective microbial manipulation strategies remains a great challenge. This challenge can be addressed by modelling approaches integrating causality principles and thus going beyond current correlation-based approaches applied to analyse rumen microbial genomic data. However, existing rumen models are not yet adapted to capitalise on microbial genomic information. This gap between the rumen microbiota available omics data and the way microbial metabolism is represented in the existing rumen models needs to be filled to enhance rumen understanding and produce better predictive models with capabilities for guiding nutritional strategies. To fill this gap, the integration of computational biology tools and mathematical modelling frameworks is needed to translate the information of the metabolic potential of the rumen microbes (inferred from their genomes) into a mathematical object. In this paper, we aim to discuss the potential use of two modelling approaches for the integration of microbial genomic information into dynamic models. The first modelling approach explores the theory of state observers to integrate microbial time series data into rumen fermentation models. The second approach is based on the genome-scale network reconstructions of rumen microbes. For a given microorganism, the network reconstruction produces a stoichiometry matrix of the metabolism. This matrix is the core of the so-called genome-scale metabolic models which can be exploited by a plethora of methods comprised within the constraint-based reconstruction and analysis approaches. We will discuss how these methods can be used to produce the next-generation models of the rumen microbiome.
A molecular level perspective on how novel phenotypes evolve is contingent on our understanding of how genomes evolve through time, and of particular interest is how novel elements emerge or are lost. Mechanisms of protein evolution such as gene duplication have been well established. Studies of gene fusion events show they often generate novel functions and adaptive benefits. Identifying gene fusion and fission events on a genome scale allows us to establish the mode and tempo of emergence of composite genes across the animal tree of life, and allows us to test the repeatability of evolution in terms of determining how often composite genes can arise independently. Here we show that ∼5% of all animal gene families are composite, and their phylogenetic distribution suggests an abrupt, rather than gradual, emergence during animal evolution. We find that gene fusion occurs at a higher rate than fission (73.3% vs 25.4%) in animal composite genes, but many gene fusions (79% of the 73.3%) have more complex patterns including subsequent fission or loss. We demonstrate that nodes such as Bilateria, Euteleostomi, and Eutheria, have significantly higher rates of accumulation of composite genes. We observe that in general deuterostomes have a greater amount of composite genes as compared to protostomes. Intriguingly, up to 41% of composite gene families have evolved independently in different clades showing that the same solutions to protein innovation have evolved time and again in animals. Significance statement New genes emerge and are lost from genomes over time. Mechanisms that can produce new genes include, but are not limited to, gene duplication, retrotransposition, de novo gene genesis, and gene fusion/fission. In this work, we show that new genes formed by fusing distinct homologous gene families together comprise a significant portion of the animal proteome. Their pattern of emergence through time is not gradual throughout the animal phylogeny - it is intensified on nodes of major transition in animal phylogeny. Interestingly, we see that evolution replays the tape frequently in these genes with 41% of gene fusion/fission events occurring independently throughout animal evolution.
Honey bees use plant material to manufacture their own food. These insect pollinators visit flowers repeatedly to collect nectar and pollen, which are shared with other hive bees to produce honey and beebread. While producing these products, beehives accumulate a tremendous amount of microbes, including bacteria that derive from plants and different parts of the honey bees’ body. In this study, we conducted 16S rDNA metataxonomic analysis on honey and beebread samples that were collected from 15 beehives in the southeast of England in order to quantify the bacteria associated with beehives. The results highlighted that honeybee products carry a significant variety of bacterial groups that comprise bee commensals, environmental bacteria and pathogens of plants and animals. Remarkably, this bacterial diversity differs amongst the beehives, suggesting a defined fingerprint that is affected, not only by the nectar and pollen gathered from local plants, but also from other environmental sources. In summary, our results show that every hive possesses their own distinct microbiome, and that honeybee products are valuable indicators of the bacteria present in the beehives and their surrounding environment.
Large regions of prokaryotic genomes are currently without any annotation, in part due to well-established limitations of annotation tools. For example, it is routine for genes using alternative start codons to be misreported or completely omitted. Therefore, we present StORF-Reporter, a tool that takes an annotated genome and returns regions that may contain missing CDS genes from unannotated regions. StORF-Reporter consists of two parts. The first begins with the extraction of unannotated regions from an annotated genome. Next, Stop-ORFs (StORFs) are identified in these unannotated regions. StORFs are open reading frames that are delimited by stop codons and thus can capture those genes most often missing in genome annotations. We show this methodology recovers genes missing from canonical genome annotations. We inspect the results of the genomes of model organisms, the pangenome of Escherichia coli, and a set of 5109 prokaryotic genomes of 247 genera from the Ensembl Bacteria database. StORF-Reporter extended the core, soft-core and accessory gene collections, identified novel gene families and extended families into additional genera. The high levels of sequence conservation observed between genera suggest that many of these StORFs are likely to be functional genes that should now be considered for inclusion in canonical annotations.
Abstract Background Antimicrobial resistance (AMR) genes are found to be ubiquitous within the microbiome, even when antimicrobial usage is absent. To identify the AMR phenotype, the most common method is to use a laboratory-based assay. Yet, when dealing with samples from the microbiome, many species are difficult to culture within the laboratory. The vast quantity of strains would be time-consuming to culture. To avoid this, a computational approach may be a more favourable choice. AMR gene finder tools are efficient at determining the AMR genotype. Despite this, how a genotype relates to the AMR phenotype is still an open question. Methods To evaluate the relationship between the AMR phenotype and the AMR genotype, 16 950 genomes from BV-BRC which had corresponding MIC values were analysed. Using Weka’s J48 decision tree model, the relationship between the AMR phenotype and the AMR genotype was analysed. The role of accessory genes in relation to the AMR phenotype was analysed in the same way. Results The J48 models could predict the AMR phenotype accurately using AMR genes and accessory genes; the average accuracy was 91.7% and 92.2%, respectively. The results found that gene co-occurrence, presence and absence of genes are key factors to analyse when identifying the AMR phenotype from genomic data. These factors are not evaluated by commonly used AMR gene finder tools, which could miss vital information to determine the correct phenotype. Conclusions Our results highlight why we should continue to research the relationship between the AMR phenotype and genomic data.
There is conflicting evidence as to whether Porifera (sponges) or Ctenophora (comb jellies) comprise the root of the animal phylogeny. Support for either a Porifera-sister or Ctenophore-sister tree has been extensively examined in the context of model selection, taxon sampling, and outgroup selection. The influence of dataset construction is comparatively understudied. We re-examine five animal phylogeny datasets that have supported either root hypothesis using an approach designed to enrich orthologous signal in phylogenomic datasets. We find that many component orthogroups in animal datasets fail to recover major lineages as monophyletic with the exception of Ctenophora, regardless of the supported root. Enriching these datasets to retain orthogroups recovering ≥3 major lineages reduces dataset size by up to 50% while retaining underlying phylogenetic information and taxon sampling. Site-heterogeneous phylogenomic analysis of these enriched datasets recovers both Porifera-sister and Ctenophora-sister positions, even with additional constraints on outgroup sampling. Two datasets which previously supported Ctenophora-sister support Porifera-sister upon enrichment. All enriched datasets display improved model fitness under posterior predictive analysis. While not conclusively rooting animals at either Porifera or Ctenophora, we do see an increase in signal for Porifera-sister and a decrease in signal for Ctenophore-sister when data are filtered for orthologous signal. Our results indicate that dataset size and construction as well as model fit influence animal root inference.
Background Manipulating the rhizosphere microbial community through beneficial microorganism inoculation has gained interest in improving crop productivity and stress resistance. Synthetic microbial communities, known as SynComs, mimic natural microbial compositions while reducing the number of components. However, achieving this goal requires a comprehensive understanding of natural microbial communities and carefully selecting compatible microorganisms with colonization traits, which still pose challenges. In this study, we employed multi-genome metabolic modeling of 270 previously described metagenome-assembled genomes from Campos rupestres to design a synthetic microbial community to improve the yield of important crop plants. Results We used a targeted approach to select a minimal community (MinCom) encompassing essential compounds for microbial metabolism and compounds relevant to plant interactions. This resulted in a reduction of the initial community size by approximately 4.5-fold. Notably, the MinCom retained crucial genes associated with essential plant growth-promoting traits, such as iron acquisition, exopolysaccharide production, potassium solubilization, nitrogen fixation, GABA production, and IAA-related tryptophan metabolism. Furthermore, our in-silico selection for the SymComs, based on a comprehensive understanding of microbe-microbe-plant interactions, yielded a set of six hub species that displayed notable taxonomic novelty, including members of the Eremiobacterota and Verrucomicrobiota phyla. Conclusion Overall, the study contributes to the growing body of research on synthetic microbial communities and their potential to enhance agricultural practices. The insights gained from our in-silico approach and the selection of hub species pave the way for further investigations into the development of tailored microbial communities that can optimize crop productivity and improve stress resilience in agricultural systems.
Microbiomes are rife for biotechnological exploitation, particularly the rumen microbiome, due to their complexicity and diversity. In this study, antimicrobial peptides (AMPs) from the rumen microbiome (Lynronne 1, 2, 3 and P15s) were assessed for their therapeutic potential against seven clinical strains of Pseudomonas aeruginosa . All AMPs exhibited antimicrobial activity against all strains, with minimum inhibitory concentrations (MICs) ranging from 4–512 µg/mL. Time-kill kinetics of all AMPs at 3× MIC values against strains PAO1 and LES431 showed complete kill within 10 min to 4 h, although P15s was not bactericidal against PAO1. All AMPs significantly inhibited biofilm formation by strains PAO1 and LES431, and induction of resistance assays showed no decrease in activity against these strains. AMP cytotoxicity against human lung cells was also minimal. In terms of mechanism of action, the AMPs showed affinity towards PAO1 and LES431 bacterial membrane lipids, efficiently permeabilising the P. aeruginosa membrane. Transcriptome and metabolome analysis revealed increased catalytic activity at the cell membrane and promotion of β-oxidation of fatty acids. Finally, tests performed with the Galleria mellonella infection model showed that Lynronne 1 and 2 were efficacious in vivo, with a 100% survival rate following treatment at 32 mg/kg and 128 mg/kg, respectively. This study illustrates the therapeutic potential of microbiome-derived AMPs against P. aeruginosa infections.
With an increasing human population access to ruminant products is an important factor in global food supply. While ruminants contribute to climate change, climate change could also affect ruminant production. Here we investigated how the plant response to climate change affects forage quality and subsequent rumen fermentation. Models of near future climate change (2050) predict increases in temperature, CO 2 , precipitation and altered weather systems which will produce stress responses in field crops. We hypothesised that pre-exposure to altered climate conditions causes compositional changes and also primes plant cells such that their post-ingestion metabolic response to the rumen is altered. This “stress memory” effect was investigated by screening ten forage grass varieties in five differing climate scenarios, including current climate (2020), future climate (2050), or future climate plus flooding, drought or heat shock. While varietal differences in fermentation were detected in terms of gas production, there was little effect of elevated temperature or CO 2 compared with controls (2020). All varieties consistently showed decreased digestibility linked to decreased methane production as a result of drought or an acute flood treatment. These results indicate that efforts to breed future forage varieties should target tolerance of acute stress rather than long term climate.
Bacteriophages (phages) are viruses that target bacteria, with the ability to lyse and kill host bacterial cells. Due to this, they have been of some interest as a therapeutic since their discovery in the early 1900s, but with the recent increase in antibiotic resistance, phages have seen a resurgence in attention. Current methods of isolation and purification of phages can be long and tedious, with caesium chloride concentration gradients the gold standard for purifying a phage fraction. Isolation of novel phages requires centrifugation and ultrafiltration of mixed samples, such as water sources, effluent or faecal samples etc, to prepare phage filtrates for further testing. We propose countercurrent chromatography as a novel and alternative approach to use when studying phages, as a scalable and high-yield method for obtaining phage fractions. However, the full extent of the usefulness and resolution of separation with this technique has not been researched; it requires optimization and ample testing before this can be revealed. Here we present an initial study to determine survivability of two phages, T4 and ϕX174, using only water as a mobile phase in a Spectrum Series 20 HPCCC. Both phages were found to remain active once eluted from the column. Phages do not fully elute from the column and sodium hydroxide is necessary to flush the column between runs to deactivate remaining phages.
Conflicting studies place a group of bilaterian invertebrates containing xenoturbellids and acoelomorphs, the Xenacoelomorpha, as either the primary emerging bilaterian phylum(1-6) or within Deuterostomia, sister to Am-bulacraria.(7-11) Although their placement as sister to the rest of Bilateria supports relatively simple morphology in the ancestral bilaterian, their alternative placement within Deuterostomia suggests a morphologically com-plex ancestral bilaterian along with extensive loss of major phenotypic traits in the Xenacoelomorpha. Recent studies have questioned whether Deuterostomia should be considered monophyletic at all.(10,12,13) Hidden pa-ralogy and poor phylogenetic signal present a major challenge for reconstructing species phylogenies.(14-18) Here, we assess whether these issues have contributed to the conflict over the placement of Xenacoelomor-pha. We reanalyzed published datasets, enriching for orthogroups whose gene trees support well-resolved clans elsewhere in the animal tree.(16) We find that most genes in previously published datasets violate incon-testable clans, suggesting that hidden paralogy and low phylogenetic signal affect the ability to reconstruct branching patterns at deep nodes in the animal tree. We demonstrate that removing orthogroups that cannot recapitulate incontestable relationships alters the final topology that is inferred, while simultaneously improving the fit of the model to the data. We discover increased, but ultimately not conclusive, support for the existence of Xenambulacraria in our set of filtered orthogroups. At a time when we are progressing toward sequencing all life on the planet, we argue that long-standing contentious issues in the tree of life will be resolved using smaller amounts of better quality data that can be modeled adequately.(19)
Here we report two antimicrobial peptides (AMPs), HG2 and HG4 identified from a rumen microbiome metagenomic dataset, with activity against multidrug-resistant (MDR) bacteria, especially methicillin-resistant Staphylococcus aureus (MRSA) strains, a major hospital and community-acquired pathogen. We employed the classifier model design to analyse, visualise, and interpret AMP activities. This approach allowed in silico discrimination of promising lead AMP candidates for experimental evaluation. The lead AMPs, HG2 and HG4, are fast-acting and show anti-biofilm and anti-inflammatory activities in vitro and demonstrated little toxicity to human primary cell lines. The peptides were effective in vivo within a Galleria mellonella model of MRSA USA300 infection. In terms of mechanism of action, HG2 and HG4 appear to interact with the cytoplasmic membrane of target cells and may inhibit other cellular processes, whilst preferentially binding to bacterial lipids over human cell lipids. Therefore, these AMPs may offer additional therapeutic templates for MDR bacterial infections.
Amanda Clare合作论文数Department of Computer Science,
Aberystwyth University10
Tobias Doerks合作论文数EMBL5