Microbiome studies reveal the taxonomic and functional composition of microbial communities inhabiting many diverse environments. Comprehensive microbiome repositories, such as MGnify, organize data into studies, each consisting of multiple sequencing runs or assemblies and accompanying metadata. This structure enables integrative, large-scale, cross-study analyses, leading to broader insights across ecosystems, hosts, and experimental contexts. Despite extensive microbiome research, methods for defining similarity between studies and validating those similarity metrics, remain insufficiently established, especially for large-scale analyses. To address this, we evaluate whether taxonomic and functional similarities from MGnify can serve as reliable indicators of study relatedness between study pairs, testing multiple metrics against conceptual relatedness (e.g., shared environments, goals, or methods). To scale validation, we introduce a framework that applies a Large Language Model (LLM) to study descriptions, categorizing study pairs by relatedness. Our results show that functional similarity correlates more strongly with LLM-inferred study relatedness than taxonomic similarity, highlighting both the promise and limitations of current metrics. Via the above, we demonstrate the value of combining microbial profiles with LLM-driven semantic reasoning to navigate the expanding landscape of metagenomic research.
Plasmids are extrachromosomal DNA molecules that enable horizontal gene transfer in bacteria, often conferring advantages such as antibiotic resistance. Despite their importance, plasmids are underrepresented in genomic databases because of challenges in assembling them, caused by mosaicism and microdiversity. Current plasmid assemblers rely on detecting circular paths in single-sample assembly graphs but face limitations because of graph fragmentation, entanglement and low coverage. We introduce PlasMAAG (plasmid and organism metagenomic binning using assembly-alignment graphs), a method to recover plasmids and cellular genomes from metagenomic samples. PlasMAAG complements assembly graph signals across samples by generating an 'assembly-alignment graph', which is used alongside common binning features for improved plasmid reconstruction. On synthetic benchmark datasets, PlasMAAG reconstructed 50-121% more near-complete plasmids than competing methods and improved the Matthews correlation coefficient of geNomad contig classification by 28-106%. On hospital sewage samples, PlasMAAG outperformed competing methods, reconstructing 33% more plasmid sequences. PlasMAAG enables the study of organism-plasmid associations and intraplasmid diversity across samples.
Topologically associating domains (TADs) are generally considered as a homogeneous basic units of genome folding, which is critical for transcriptional regulation. However, recent studies indicate that both the TAD domain structures and the boundaries between them are not as homogeneous as originally recognized. Here, we address the heterogeneity of the TAD boundaries in the human genome at a large scale, which varies between active and inactive chromatin and across cell lines and tissues. To address this, based on the well-annotated TAD boundaries extracted from multiple cell lines and tissues, we examine their nucleotide content, resulting in two main clusters, one GC-rich and one AT-rich, which are mainly distributed in active and inactive chromatin, respectively. Also, they contain different types of repetitive sequences and have different epigenetic patterns, with more CTCF binding motifs in the GC-rich cluster. Hence, our observations of the TAD boundary content provide novel insights into TAD genomic architecture. In addition, we find that cell- or tissue-specific boundaries are less evolutionarily conserved than other boundaries. We highlight the importance of TAD boundary diversity in different functional contexts and discuss the importance of the different types of repetitive sequences and epigenetic patterns in the two main types of boundaries.
Inflammation is a complex biological process that upregulates numerous genes compared to homeostasis. These transcriptional changes can dominate the statistical signal in differential expression analyses and confound the detection of disease-specific effects. Here, we exploit inflammatory signals in public clinical transcriptomic datasets to derive the inflammatome, a comprehensive set of 2,000 genes consistently upregulated in various inflammatory diseases vs. healthy controls. Within this set, we define the inflammation signature, a high-confidence subset of 100 genes capturing the most consistent changes. We demonstrate how these gene sets can be used to identify and filter inflammation-associated changes in gene expression and protein abundance. Additionally, we introduce a sample-wise inflammation score derived from the inflammation signature and show its correlation with clinical disease severity. To support broad usability, we provide these functionalities in a user-friendly Shiny app. This resource will enable users to assess the effect of inflammation in future transcriptomic and proteomic studies.
Climate Change (CC) is reshaping all ecosystem processes and structures. Microbial data provide valuable insights into how microbial processes contribute to CC and how CC, in turn, alters microbial communities. However, the growing volume of environmental genomics data makes identifying CC-related records challenging. The Climate Change Metagenomic Record Index (CCMRI) has been developed to harvest metagenomic/microbiome records pertaining to CC and to provide researchers with a curated database of CC-related microbiome studies (https://ccmri.hcmr.gr). To guide interpretation, the database’s 169 metagenomic studies have been labelled according to their relation to CC as CC-caused, CC-causing, and CC-mitigating. They have also been annotated with the CC phenomena they explore, like methane production, temperature rise, permafrost thawing, greenhouse gas emission, methanotrophy, and ocean acidification. To ease navigation, they have also been classified according to their biome as aquatic, terrestrial, host-associated, and engineered. The CCMRI database was initially constructed through manual curation of all aquatic and terrestrial studies in the MGnify resource. It was then expanded with the help of the CCMRI curation-assistant system. This leveraged Large Language Models to scan the remaining MGnify studies, filtered them for relevance, and proposed candidates for inclusion. With a recall greater than 90%, the system achieved high accuracy in identifying CC-related studies. The final decisions on CC-relatedness and categorization were performed by a human curator. This approach combines the efficiency of automation with human oversight and greatly reduces the curation effort, ensuring sustainability and scalability.
A fundamental understanding of genome organization relies on accurately annotating topologically associating domains (TADs) and their boundaries. This is crucial for understanding how cis-regulatory elements regulate gene expression. To go beyond calling TADs and boundaries from Hi-C data, several machine learning-based methods have been proposed to go the step further and predict TAD boundaries from genomic sequences. As the growing evidence of TADs and their boundaries, TADs have been proved exhibiting diverse properties, such as differences in replication timing and epigenetic patterns. However, existing methods do not take this heterogeneity into account. To address this, we propose a method called TADBpred for TAD boundary prediction in a large genomic context in humans. TADBpred focuses on TAD boundaries in active and inactive chromatin across cell-lines and tissues, which are GC-rich and AT-rich, respectively. By integrating genomic elements and sequence composition, we designed two models for GC-rich and AT-rich boundaries, respectively. When testing the performance on respective independent held-out datasets, we obtain AUC scores of 0.91 and 0.80. Our results indicate that TADBpred excels in TAD boundary prediction. Additionally, feature importance analysis highlights the essential features for different classes of TAD boundaries, thereby enhancing our understanding of these TAD boundaries.
Identifying disease-relevant proteins and pathways remains a fundamental challenge in understanding disease mechanisms and supporting therapeutic development. While omics analyses can provide valuable insights, they typically consider each gene/protein separately rather than at the level of biological systems. This can be addressed by combining the omics data with protein networks. We integrate disease-specific omics data with a universal functional association network from STRING, which we represent using node2vec embedding. This way, we constructed disease maps for seven diseases spanning inflammatory, oncological, neurological, and vascular diseases based on genetics, transcriptomics, somatic mutation, and proteomics data. Compared to omics analysis alone, the use of a simple linear model on top of network embedding enabled us to identify 2–4 times as many known disease-relevant proteins at the same specificity. Clustering of the resulting disease maps revealed both functional modules shared by many diseases, such as inflammatory pathways and cancer hallmarks, and disease-specific modules, such as keratinization in atopic dermatitis and extracellular matrix remodeling in aortic aneurysm. Together, these results highlight the value of protein network embedding when analyzing omics data to understand diseases.
Linkage-specific ubiquitin chains govern the outcome of numerous critical ubiquitin-dependent signaling processes, but their targets and functional impacts remain incompletely understood due to a paucity of tools for their specific detection and manipulation. Here, we applied a cell-based ubiquitin replacement strategy enabling targeted conditional abrogation of each of the seven lysine-based ubiquitin linkages in human cells to profile system-wide impacts of disabling formation of individual chain types. This revealed proteins and processes regulated by each of these poly-ubiquitin topologies and indispensable roles of K48-, K63- and K27-linkages in cell proliferation. We show that K29-linked ubiquitylation is strongly associated with chromosome biology, and that the H3K9me3 methyltransferase SUV39H1 is a prominent cellular target of this modification. K29-linked ubiquitylation catalyzed by TRIP12 and reversed by TRABID constitutes the essential degradation signal for SUV39H1 and is primed and extended by Cullin-RING ubiquitin ligase activity. Preventing K29-linkage-dependent SUV39H1 turnover deregulates H3K9me3 homeostasis but not other histone modifications. Collectively, these data resources illuminate cellular functions of linkage-specific ubiquitin chains and establish a key role of K29-linked ubiquitylation in epigenome integrity.
MOTIVATION:Representation learning has revolutionized sequence-based prediction of protein function and subcellular localization. Protein networks are an important source of information complementary to sequences, but the use of protein networks has proven to be challenging in the context of machine learning, especially in a cross-species setting. RESULTS:We leveraged the STRING database of protein networks and orthology relations for 1322 eukaryotes to generate network-based cross-species protein embeddings. We did this by first creating species-specific network embeddings and subsequently aligning them based on orthology relations to facilitate direct cross-species comparisons. We show that these aligned network embeddings ensure consistency across species without sacrificing quality compared to species-specific network embeddings. We also show that the aligned network embeddings are complementary to sequence embedding techniques, despite the use of sequence-based orthology relations in the alignment process. Finally, we validated the embeddings by using them for two well-established tasks: subcellular localization prediction and protein function prediction. Training logistic regression classifiers on aligned network embeddings and sequence embeddings improved the accuracy over using sequence alone, reaching performance numbers close to state-of-the-art deep-learning methods. AVAILABILITY AND IMPLEMENTATION:The source code and scripts for generating the network-based cross-species protein embeddings are available at https://github.com/deweihu96/SPACE. Precomputed network embeddings and sequence embeddings for all eukaryotic proteins are included in STRING version 12.0 (https://string-db.org/cgi/download).
Impaired gut barrier function may lead to progression of liver fibrosis in people with alcohol-related liver disease. The postbiotic ReFerm® can lower gut barrier permeability and may thereby reduce fibrosis formation. Here, we report the results from an open-labelled, single centre randomized controlled trial where 56 patients with advanced, compensated, alcohol-related liver disease were assigned 1:1 to receive either ReFerm® (n = 28) or standard nutritional support (Fresubin®, n = 28) for 24 weeks. The primary outcome was a ≥ 10% reduction of the fibrosis formation marker alpha-smooth muscle actin in liver biopsies, assessed by a blinded pathologist using automated digital imaging analysis. Paired liver biopsies meeting quality criteria for the primary outcome were available for 40 participants (ReFerm®, n = 21 and Fresubin®, n = 19). This reduction was observed in 29% of patients receiving ReFerm®, compared to 14% with Fresubin® (OR = 2.40; 95% CI 0.63 to 9.16; p = 0.200). No treatment-related serious adverse events occurred. Our findings suggest that ReFerm® may reduce liver fibrosis by enhancing gut barrier function, potentially preventing the progression of alcohol-related liver disease.
Mass spectrometry (MS)-based proteomics is a well-established strategy for analyzing complex biological mixtures. Many MS instruments and data acquisition strategies are available, and the data they acquire differ substantially, thus requiring tailored analysis algorithms. Hence, many dedicated bioinformatics workflows are developed. These are in constant evolution, and the community lacks a centralized platform for comparing their performance. Here, we propose ProteoBench, a single platform that brings together software developers and software users to provide an ever-evolving comparison of state-of-the-art proteomics data processing tools. ProteoBench is an open-source resource that enables the community to evaluate data analysis workflows, develop benchmarking modules dedicated to specific comparisons, and discuss the best methods to compare software tools. The platform ensures that the benchmark evolves alongside advances in proteomics data analysis workflows. ProteoBench guides researchers towards the best-suited tool and parameters for their specific project and data according to their needs, and developers can test their newly developed tools or workflows privately, before adding them as public references. This community-driven effort will increase transparency and reproducibility between MS data analysis workflows, as well as facilitate the development and publication of software workflows in the field. ### Competing Interest Statement The authors have declared no competing interest.
Lifestyle factors (LSFs) are increasingly recognized as instrumental in both the development and control of diseases. Despite their importance, there is a lack of methods to extract relations between LSFs and diseases from the literature, a step necessary to consolidate the currently available knowledge into a structured form. As simple co-occurrence-based relation extraction (RE) approaches are unable to distinguish between the different types of LSF-disease relations, context-aware models such as transformers are required to extract and classify these relations into specific relation types. However, no comprehensive LSF-disease RE system existed, nor a corpus suitable for developing one. We present LSD600 (available at https://zenodo.org/records/13952449), the first corpus specifically designed for LSF-disease RE, comprising 600 abstracts with 1900 relations of eight distinct types between 5027 diseases and 6930 LSF entities. We evaluated LSD600's quality by training a RoBERTa model on the corpus, achieving an F-score of 68.5% for the multilabel RE task on the held-out test set. We further validated LSD600 by using the trained model on the two Nutrition-Disease and FoodDisease datasets, where it achieved F-scores of 70.7% and 80.7%, respectively. Building on these performance results, LSD600 and the RE system trained on it can be valuable resources to fill the existing gap in this area and pave the way for downstream applications.Database URL: https://zenodo.org/records/13952449
CRISPR-derived base editors (BE) enable precise single nucleotide substitution without introducing double-stranded DNA breaks. Apart from the base editing enzymes, efficient base editing strongly depends on both the CRISPR guide RNA (gRNA) efficiency and the edited position. Here, we show that the accuracy of BE gRNA design can be significantly improved by generating more data and by introducing deep neural networks trained on multiple different datasets simultaneously. Generating ~20,000 gRNAs for A•T to G•C and C•G to T•A conversions, we present such deep learning models, which also allow users to do dataset-aware predictions. The methods are available online and as stand-alone software.
Giant tortoises exhibit exceptional longevity, often exceeding the human lifespan. To understand the genomic and epigenomic basis of their longevity, we analyzed the DNA sequence and methylome of Jonathan, an Aldabra giant tortoise (Aldabrachelys gigantea), estimated to be 192 years old. Relative to other giant tortoises (Aldabrachelys gigantea and Chelonoidis abingdonii), we found Jonathan has gene variants in pathways associated with aging, including DNA repair and telomere regulation. Consistent with his advanced age, Jonathan has significant age-related changes in DNA methylation and methylation entropy, compared with a 5-year-old Aldabra individual. Notably, we found that low entropy regions in Jonathan's methylome were enriched for genes involved in the electron transport chain. This suggests that high-fidelity transcription of these genes may be crucial for extreme longevity. With this data, we propose a model for aging, that links efficient mitochondrial energy production with nuclear maintenance of low methylation entropy. ### Competing Interest Statement The Regents of the University of California are the sole owner of patents and patent applications directed at epigenetic biomarkers for which Steve Horvath is a named inventor; SH is a founder and paid consultant of the non-profit Epigenetic Clock Development Foundation that licenses these patents. SH is a Principal Investigator at the Altos Labs, Cambridge Institute of Science. The other authors declare no competing interests.
Mass spectrometry-based proteomics allows the quantification of thousands of proteins, protein variants, and their modifications, in many biological samples. These are derived from the measurement of peptide relative quantities, and it is not always possible to distinguish proteins with similar sequences due to the absence of protein- specific peptides. In such cases, peptide signals are reported in protein groups that can correspond to several genes. Here, we show that multi-gene protein groups have a limited impact on GO-term enrichment, but selecting only one gene per group affects network analysis. We thus present the Cytoscape app Proteo Visualizer (https://apps. cytoscape.org/apps/ProteoVisualizer) that is designed for retrieving protein interaction networks from STRING using protein groups as input and thus allows visualization and network analysis of bottom-up MS-based proteomics data sets.
MOTIVATION:Dictionary-based named entity recognition (NER) allows terms to be detected in a corpus and normalized to biomedical databases and ontologies. However, adaptation to different entity types requires new high-quality dictionaries and associated lists of blocked names for each type. The latter are so far created by identifying cases that cause many false positives through manual inspection of individual names, a process that scales poorly. RESULTS:In this work, we aim to improve block list s by automatically identifying names to block, based on the context in which they appear. By comparing results of three well-established biomedical NER methods, we generated a dataset of over 12.5 million text spans where the methods agree on the boundaries and type of entity tagged. These were used to generate positive and negative examples of contexts for four entity types (genes, diseases, species, and chemicals), which were used to train a Transformer-based model (BioBERT) to perform entity type classification. Application of the best model (F1-score = 96.7%) allowed us to generate a list of problematic names that should be blocked. Introducing this into our system doubled the size of the previous list of corpus-wide blocked names. In addition, we generated a document-specific list that allows ambiguous names to be blocked in specific documents. These changes boosted text mining precision by ∼5.5% on average, and over 8.5% for chemical and 7.5% for gene names, positively affecting several biological databases utilizing this NER system, like the STRING database, with only a minor drop in recall (0.6%). AVAILABILITY AND IMPLEMENTATION:All resources are available through Zenodo https://doi.org/10.5281/zenodo.11243139 and GitHub https://doi.org/10.5281/zenodo.10289360.
We read with great interest the review by Tranah et al, which highlighted that alcohol consumption, alterations in the gut microbiome and impairment of the gut barrier function were linked to the development of alcoholrelated liver disease (ALD). However, the impact of acute alcohol consumption on the circulating microbiome in patients with ALD remains unclear. To address this gap, we conducted a controlled acute alcohol intervention in healthy controls (n=8), individuals with ALD (n=14) and nonalcoholic fatty liver disease (NAFLD) (n=14). Ethanol (2.5 mL of 40% EtOH per kg body weight) was instilled via a nasogastric tube over 30 min by infusion pump. To sample hepatic and systemic venous blood simultaneously, we placed a catheter in a hepatic vein via a transjugular access and another catheter in the right internal jugular vein. Blood was sampled at eight time points over 3 hours (figure 1A). 3 We performed 16S rRNA gene amplicon sequencing (16S sequencing) and measurement of 16S rRNA gene copy number using qPCR (16S qPCR) in 576 blood samples. Circulating microbial DNA quantities were significantly higher in hepatic/systemic circulation compared with negative controls, indicating that microbial contamination was well contained and had a negligible impact on the results (figure 1B). Alcohol concentration was significantly higher in the hepatic venous blood until 90 min after the intervention compared with the systemic circulation, indicating Letter
The Knowledge Management Center (KMC) for the Illuminating the Druggable Genome (IDG) project aims to aggregate, update, and articulate protein-centric data knowledge for the entire human proteome, with emphasis on the understudied proteins from the three IDG protein families. KMC collates and analyzes data from over 70 resources to compile the Target Central Resource Database (TCRD), which is the web-based informatics platform (Pharos). These data include experimental, computational, and text-mined information on protein structures, compound interactions, and disease and phenotype associations. Based on this knowledge, proteins are classified into different Target Development Levels (TDLs) for identification of understudied targets. Additional work by the KMC focuses on enriching target knowledge and producing DrugCentral and other data visualization tools for expanding investigation of understudied targets.
Søren Brunak合作论文数Rigshospitalet;Novo Nordisk Foundation Center for Protein Research, University of Copenhagen;Department of Systems Biology, Technical University of Denmark53
Alexander Roth合作论文数Spezifikation und Entwicklung universitarer Lern- und Arbeitsumgebungen8
Tobias Doerks合作论文数EMBL8