Motivation Spurious protein sequences, resulting from gene prediction errors, theoretically should not yield folded structures. AlphaFold2 was previously shown to predict short spurious sequences with high pLDDT scores and was therefore unlikely to distinguish between real proteins and spurious proteins which are usually short. We evaluate whether newer structure prediction methods (ESMFold and AlphaFold3) similarly predict short sequences with high pLDDT or if they better discriminate between spurious and real proteins.Results All three structure prediction methods (ESMFold, AlphaFold2, and AlphaFold3) predict short spurious sequences from AntiFam with unexpectedly high pLDDT scores, however the discrimination between spurious and real proteins improves beyond 100 amino acids. By analysing sequences with disparate pTM and pLDDT scores, we identified two potentially novel spurious shadow ORFs in Swiss-Prot and one potentially non-spurious AntiFam entry. Using the structure prediction scores, we developed a Gaussian Process Model and evaluated its performance on AlphaFold DB, identifying potential spurious proteins at scale. While limited on its own, this model can increase confidence in spurious protein identification when combined with other methods.Availability and implementation Structure predictions are available at https://doi.org/10.5281/zenodo.20426908. Model implementation and figure generation code are available at https://github.com/0rra/fold_unfold2.
Expression Atlas (https://www.ebi.ac.uk/gxa/home) is EMBL-EBI's comprehensive knowledgebase for gene and protein expression across tissues, cell types, conditions, and multiple species. Since our last update, Expression Atlas has expanded substantially in both content and functionality, now comprising >4500 studies from 67 species, with increased proteomics coverage and updated Genotype-Tissue Expression (GTEx) tissue profiles. The resource also includes hundreds of single-cell RNA-seq experiments spanning 21 species, among them externally analysed community datasets such as Tabula Sapiens and GTEx single-nucleus profiles, allowing exploration of curated atlases while maintaining their original analytical framework. Key methodological advances include a new marker gene analysis module for bulk baseline experiments, alongside workflow updates that improve reproducibility. Expression Atlas data are integrated into EMBL-EBI resources such as Ensembl, UniProt, and Europe PMC and disseminated through collaboration with model organism communities such as FlyBase and Gramene. The resource also supports translational research through the European Diagnostic Transcriptomic Library and integration with the Open Targets platform. Future directions include modernizing analysis pipelines, enhancing programmatic access, and delivering AI-ready data formats, strengthening Expression Atlas as a findable, accessible, interoperable, and reusable (FAIR) community-driven resource for both fundamental and translational discovery.
Motivation:InterProScan is a widely used software package for large-scale protein function classification and an essential component of UniProt, Ensembl and MGnify Genomes large-scale annotation pipelines. For the past 12 years, InterProScan 5 has provided robust analysis but faced challenges in scalability, installation, data management, and integration with modern workflow systems. We present InterProScan 6, a complete reimplementation as a Nextflow pipeline that enhances scalability, portability, and reproducibility across diverse computational environments, including local, HPC, and cloud platforms. Key improvements include decoupling of application code from signature data, native container support, on-demand data management, integration of modern deep-learning-based predictors, and a redesigned Matches API for retrieving pre-computed annotations in JSON format. Results:Across nine reference proteomes ranging from bacteria to complex eukaryotes, InterProScan 6 consistently reduced wall-clock runtime relative to InterProScan 5, with speedups of approximately two-fold on large eukaryotic proteomes. When pre-computed annotations were available via the Matches API, runtimes were reduced to minutes. Comparison of annotations generated from all Swiss-Prot sequences shows that InterProScan 6 reproduces InterProScan 5 results with near-identical precision and sensitivity across all InterPro member databases. InterProScan 6 therefore enables more efficient, flexible, and reproducible genome-scale protein function annotation. Availability and implementation:InterProScan 6 is distributed under the Apache 2.0 license. Its documentation is hosted on ReadTheDocs (https://interproscan6.readthedocs.io/) and its source code is available on GitHub (https://github.com/ebi-pf-team/interproscan6).
Motivation: Continuing advances in genome and metagenome sequencing expand the number of identified conserved protein families that remain functionally uncharacterized and contain domains of unknown function (DUFs). Functional-association resources such as STRING provide biological context, but mostly do not distinguish indirect association from physical interaction. We assessed whether AlphaFold 3 complex prediction, combined with STRING evidence and domain-level analysis of interfaces and interaction partners, can help identify and characterize DUF-containing proteins. Results: We generated four structural-prediction cohorts from STRING associations involving DUF-containing proteins and evaluated the predicted complexes using interface ipSAE, average pLDDT and buried surface area. An L2-regularized logistic regression model was trained on an initial cohort of predictions from high-confidence STRING associations to prioritize DUF-containing candidates likely to produce structurally confident AlphaFold 3 complexes. The model was then applied across all 12,535 organisms represented in STRING v12.0, followed by grouping into DUF-family and partner-architecture modules, covering 2,076 unique DUF families. The final L2-model screen contained 12,298 successfully modelled protein pairs, including 1,208 (9.82%) complexes meeting a strict-confidence criterion and 2,433 (19.78%) meeting a more liberal confidence criterion. Two examples suggest roles for DUF4130 in nucleic-acid-associated radical-SAM biology and DUF5819 in a bacterial system related to vitamin-K-dependent carboxylation. Availability and implementation: Predicted structures and associated metadata are available through Zenodo at https://doi.org/10.5281/zenodo.21875362. The model implementation and code used to generate the analyses and figures are available at https://github.com/linoriep/Proteome-scale-structure-prediction-of-DUF-containing-protein-protein-interactions.
MOTIVATION:The exponential growth of non-coding RNA research-with over 230 000 papers published since 2000-has created an urgent knowledge management crisis in molecular biology. Despite their crucial regulatory roles, microRNAs (miRNAs) face a significant curation bottleneck, with only 1400 articles manually curated to the Gene Ontology (GO) knowledgebase over a decade. This highlights the critical need for automated systems that can accelerate biocuration while maintaining high-quality standards. RESULTS:We present GOFlowLLM, an automated curation pipeline powered by reasoning-enabled Large Language Models (LLMs) that follows established GO curation flowcharts to extract and structure miRNA-mediated gene silencing data at scale. When evaluated on existing curation, GOFlowLLM selects the correct GO term in 90% of cases, with curators agreeing with 95% of the system's reasoning steps and 90% of the evidence selected. Applied to 6996 previously uncurated articles using the Qwen QwQ-32B model, our system identified 2538 new candidate GO annotations on 1785 articles in just 58 hours-potentially doubling the available miRNA GO curation. Manual review shows curators agreed with the selected term in 87% of cases, the model's reasoning in 92% of cases, and the extracted evidence in 93%. The integration of reasoning traces provides transparent justification for annotations that can be reviewed by human curators, addressing a key challenge in adopting AI for scientific curation. AVAILABILITY AND IMPLEMENTATION:GOFlowLLM is implemented as an automated pipeline that follows expert-designed reasoning frameworks to maintain curation quality. The system is available on GitHub: https://github.com/RNAcentral/GO_Flow_LLM.
Data-driven biology relies on structured knowledge generated by expert biocurators, yet this work remains largely unrecognized in traditional academic assessments. To bridge this gap, we present the updated APICURON platform, a credit-attribution infrastructure that formally acknowledges these scientific contributions. Rather than relying on delayed batch reporting, the system captures curation events as they happen and transforms them into verifiable units of work. This design allows independent resources to define and update their own recognition models while preserving the historical record of each contribution. For researchers, APICURON highlights recent activity alongside lifetime achievements and connects verified activities to persistent academic profiles via ORCID. APICURON has been successfully integrated across biological knowledgebases and data resources, demonstrating its application to diverse workflows. Extending beyond biodata resources, it also supports recognition of non-traditional research artefacts, including training materials and research software, without imposing a rigid definition of contribution.
Many bacteria, including the important human pathogen Pseudomonas aeruginosa , are naturally found in antibiotic-tolerant, multicellular biofilms. Cell-cell interactions within P. aeruginosa biofilms are mediated by a large fibrillar adhesin called CdrA in an extracellular polysaccharide-dependent manner. Here, we report an electron cryomicroscopy structure of the 60 kDa CdrA adhesive N-terminus, which combined with electron cryotomography of focused-ion beam milled specimens, allows us to derive a complete in situ model of the native adhesin. Our structure reveals a small adhesive domain (called ADEPT) at the distal tip of CdrA that is nearly perfectly conserved across the P. aeruginosa pangenome, with structural similarity to previously reported sugar-binding domains in multiple bacterial species. Inhibitory nanobodies targeting CdrA that reduce biofilm formation bind to epitopes in, or close to, the ADEPT on bacterial cells. Furthermore, structure-guided mutagenesis of residues within the ADEPT abolishes bacterial aggregation, and genomic deletion of the whole ADEPT leads to strong attenuation of biofilm formation. Our data forms a rational basis for future targeted inhibition of pathogenic P. aeruginosa biofilms and elucidates the mechanism of biofilm formation mediated by fibrillar adhesins that are widespread in bacteria.
Rfam is a comprehensive database of non-coding RNA (ncRNA) families providing curated sequence alignments, consensus secondary structures, and covariance models for thousands of RNA families. The database is essential for identifying structured non-coding RNAs in newly sequenced genomes and understanding RNA structure-function relationships. Here we present computational protocols for automated ncRNA annotation of viral genomes, and for programmatic interaction with Rfam through its RESTful API. We showcase genome-wide RNA structure visualization from a genome sequence and from a multiple sequence alignment by generating comprehensive 2D structure diagrams using newly developed features in R2DT. We also present practical examples for retrieving family metadata, downloading alignments, accessing secondary structures, and searching user sequences from the Rfam API. These methods enable researchers in virology and RNA biology to integrate Rfam data into custom bioinformatics pipelines, comparative analyses, and machine learning workflows.
Curation of literature in life sciences is a growing challenge. The continued increase in the rate of publication, coupled with the relatively fixed number of curators worldwide, presents a major challenge to developers of biomedical knowledgebases. Very few knowledgebases have resources to scale to the whole relevant literature and all have to prioritize their efforts. In this work, we take a first step to alleviating the lack of curator time in RNA science by generating summaries of literature for noncoding RNAs using large language models (LLMs). We demonstrate that high-quality, factually accurate summaries with accurate references can be automatically generated from the literature using a commercial LLM and a chain of prompts and checks. Manual assessment was carried out for a subset of summaries, with the majority being rated extremely high quality. We apply our tool to a selection of >4600 ncRNAs and make the generated summaries available via the RNAcentral resource. We conclude that automated literature summarization is feasible with the current generation of LLMs, provided that careful prompting and automated checking are applied. Database URL: https://rnacentral.org/
RNAcentral was founded in 2014 to serve as a comprehensive database of non-coding RNA sequences. It began by providing a single unified interface to more specialised resources, and now contains 45 million sequences. It has grown beyond providing a single interface to many specialised resources and now provides several services and analyses. These include secondary structure prediction with R2DT, sequence search, and analysis with Rfam. Since its last publication in 2021, RNAcentral has developed two major features. First, literature integration with the development of LitScan and LitSumm. LitScan automatically identifies and links relevant publications to RNA entries, while LitSumm uses natural language processing to generate functional summaries from the literature. Together, these tools address the critical challenge of connecting sequence data with scattered functional knowledge across thousands of publications. Secondly, RNAcentral has created gene level entries. Gene level entries represent a large structural change to RNAcentral. While RNAcentral previously organized data exclusively at the sequence level, we now group related transcripts into gene-centric views. This allows researchers to explore all isoforms, splice variants, and related sequences for a gene in a unified interface, better reflecting biological organization and facilitating comparative analyses. RNAcentral is freely available at: https://rnacentral.org .
Many proteins harbor covalent intramolecular bonds that enhance their stability and resistance to thermal, mechanical, and proteolytic insults. Intramolecular isopeptide bonds represent one such covalent interaction, yet their distribution across protein domains and organisms has been largely unexplored. Here, we sought to address this by employing a large-scale prediction of intramolecular isopeptide bonds in the AlphaFold database using the structural template-based software Isopeptor. Our findings reveal an extensive phyletic distribution in bacterial and archaeal surface proteins resembling fibrillar adhesins and pilins. All identified intramolecular isopeptide bonds are found in two structurally distinct folds, CnaA-like or CnaB-like, from a relatively small set of related Pfam families, including 10 novel families that we predict to contain intramolecular isopeptide bonds. One CnaA-like domain of unknown function, DUF11 (renamed here to "CLIPPER") is broadly distributed in cell-surface proteins from Gram-positive bacteria, Gram-negative bacteria, and archaea, and is structurally and biophysically characterized in this work. Using x-ray crystallography, we resolve a CLIPPER domain from a Gram-negative fibrillar adhesin that contains an intramolecular isopeptide bond and further demonstrate that it imparts thermostability and resistance to proteolysis. Our findings demonstrate the extensive distribution of intramolecular isopeptide bond-containing protein domains in nature and structurally resolve the previously cryptic CLIPPER domain.
Motivation:Intramolecular isopeptide bonds contribute to the structural stability of proteins, and have primarily been identified in domains of bacterial fibrillar adhesins and pili. At present, there is no systematic method available to detect them in newly determined molecular structures. This can result in mis-annotations and incorrect modeling. Results:Here, we present Isopeptor, a computational tool designed to predict the presence of intramolecular isopeptide bonds in experimentally determined structures. Isopeptor utilizes structure-guided template matching via the Jess software, combined with a logistic regression classifier that incorporates root mean square deviation and relative solvent accessible area as key features. The tool demonstrates a precision of 1.0 and a recall of 0.947 when tested on a Protein Data Bank subset of domains known to contain intramolecular isopeptide bonds that have been deposited with incorrectly modeled geometries. Availability and implementation:Isopeptor's Python-based implementation supports integration into bioinformatics workflows and can be accessed via the command line, through a Python API or via a Google Colaboratory implementation (https://colab.research.google.com/github/FranceCosta/Isopeptor_development/blob/main/notebooks/Isopeptide_finder.ipynb). Source code is hosted on GitHub (https://github.com/FranceCosta/isopeptor) and can be installed via the Python package installation manager PIP.
InterPro (https://www.ebi.ac.uk/interpro) is a freely accessible resource for the classification of protein sequences into families. It integrates predictive models, known as signatures, from multiple member databases to classify sequences into families and predict the presence of domains and significant sites. The InterPro database provides annotations for over 200 million sequences, ensuring extensive coverage of UniProtKB, the standard repository of protein sequences, and includes mappings to several other major resources, such as Gene Ontology (GO), Protein Data Bank in Europe (PDBe) and the AlphaFold Protein Structure Database. In this publication, we report on the status of InterPro (version 101.0), detailing new developments in the database, associated web interface and software. Notable updates include the increased integration of structures predicted by AlphaFold and the enhanced description of protein families using artificial intelligence. Over the past two years, more than 5000 new InterPro entries have been created. The InterPro website now offers access to 85 000 protein families and domains from its member databases and serves as a long-term archive for retired databases. InterPro data, software and tools are freely available. [GRAPHICS] .
The classification of novel protein folds remains a central challenge in structural bioinformatics, particularly as deep learning models like AlphaFold2 dramatically expand the universe of predicted protein structures. In this study, we investigated 664 candidate novel fold (CNF) domains from the TED database that both TED and DPAM methods had classified with low confidence. These CNFs span a structurally diverse and largely non-redundant set of domains, most of which lack clear sequence or structural similarity to known folds. Many CNFs appear as insertions into known transmembrane or enzymatic domains, while others occur in modular architectures, co-occurring with interaction or catalytic folds such as β-barrels, zinc fingers, or Rossmann-like domains. Although some CNFs resemble known folds that have undergone topological rearrangements or circular permutations, others result from errors in domain boundary prediction, often due to truncated sequences or tightly packed domain duplications. Our analyses led to the creation of 190 new Pfam families, many classified as domains of unknown function (DUFs), and revealed intriguing cases of zinc-binding and disulfide-rich architectures that contribute to fold space expansion. A small subset of CNFs helped define new superfamilies by linking previously unclassified but structurally related domains. Taken together, this work underscores the importance of integrating structural, evolutionary, and contextual information to resolve challenging fold assignments and provides a roadmap for extending protein classification frameworks into previously uncharted structural territory.
Classification of protein domains based on homology and structural similarity serves as a fundamental tool to gain biological insights into protein function. Recent advancements in protein structure prediction, exemplified by AlphaFold, have revolutionized the availability of protein structural data. We focus on classifying about 9000 Pfam families into ECOD (Evolutionary Classification of Domains) by using predicted AlphaFold models and the DPAM (Domain Parser for AlphaFold Models) tool. Our results offer insights into their homologous relationships and domain boundaries. More than half of these Pfam families contain DPAM domains that can be confidently assigned to the ECOD hierarchy. Most assigned domains belong to highly populated folds such as Immunoglobulin-like (IgL), Armadillo (ARM), helix-turn-helix (HTH), and Src homology 3 (SH3). A large fraction of DPAM domains, however, cannot be confidently assigned to ECOD homologous groups. These unassigned domains exhibit statistically different characteristics, including shorter average length, fewer secondary structure elements, and more abundant transmembrane segments. They could potentially define novel families remotely related to domains with known structures or novel superfamilies and folds. Manual scrutiny of a subset of these domains revealed an abundance of internal duplications and recurring structural motifs. Exploring sequence and structural features such as disulfide bond patterns, metal-binding sites, and enzyme active sites helped uncover novel structural folds as well as remote evolutionary relationships. By bridging the gap between sequence-based Pfam and structure-based ECOD domain classifications, our study contributes to a more comprehensive understanding of the protein universe by providing structural and functional insights into previously uncharacterized proteins.
The evolutionary classification of protein domains (ECOD) classifies protein domains using a combination of sequence and structural data (http://prodata.swmed.edu/ecod). Here we present the culmination of our previous efforts at classifying domains from predicted structures, principally from the AlphaFold Database (AFDB), by integrating these domains with our existing classification of PDB structures. This combined classification includes both domains from our previous, purely experimental, classification of domains as well as domains from our provisional classification of 48 proteomes in AFDB predicted from model organisms and organisms of concern to global health. ECOD classifies over 1.8 M domains from over 1000 000 proteins collectively deposited in the PDB and AFDB. Additionally, we have changed the F-group classification reference used for ECOD, deprecating our original ECODf library and instead relying on direct collaboration with the Pfam sequence family database to inform our classification. Pfam provides similar coverage of ECOD with family classification while being more accurate and less redundant. By eliminating duplication of effort, we can improve both classifications. Finally, we discuss the initial deployment of DrugDomain, a database of domain-ligand interactions, on ECOD and discuss future plans.
Motivation:Data reuse is a common and vital practice in molecular biology and enables the knowledge gathered over recent decades to drive discovery and innovation in the life sciences. Much of this knowledge has been collated into molecular biology databases, such as UniProtKB, and these resources derive enormous value from sharing data among themselves. However, quantifying and documenting this kind of data reuse remains a challenge.Results:The article reports on a one-day virtual workshop hosted by the UniProt Consortium in March 2023, attended by representatives from biodata resources, experts in data management, and NIH program managers. Workshop discussions focused on strategies for tracking data reuse, best practices for reusing data, and the challenges associated with data reuse and tracking. Surveys and discussions showed that data reuse is widespread, but critical information for reproducibility is sometimes lacking. Challenges include costs of tracking data reuse, tensions between tracking data and open sharing, restrictive licenses, and difficulties in tracking commercial data use. Recommendations that emerged from the discussion include: development of standardized formats for documenting data reuse, education about the obstacles posed by restrictive licenses, and continued recognition by funding agencies that data management is a critical activity that requires dedicated resources.Availability and implementation:Summaries of survey results are available at: https://docs.google.com/forms/d/1j-VU2ifEKb9C-sW6l3ATB79dgHdRk5v_lESv2hawnso/viewanalytics (survey of data providers) and https://docs.google.com/forms/d/18WbJFutUd7qiZoEzbOytFYXSfWFT61hVce0vjvIwIjk/viewanalytics (survey of users).
Motivation:High confidence structure prediction models have become available for nearly all protein sequences. More than 200 million AlphaFold2 models are now publicly available. We observe that there can be significant variability in the prediction confidence as judged by plDDT scores across a protein family. We have explored whether the predictions with lower plDDT in a family can be improved by the use of higher plDDT templates from the family as template structures in AlphaFold2. Results:Our work shows that about one-third of the time structures with a low plDDT can be "rescued," moved from low to reasonable confidence. We also find that surprisingly in many cases we get a higher plDDT model when we switch off the multiple sequence alignment (MSA) option in AlphaFold2 and solely rely on a high-quality template. However, we find the best overall strategy is to make predictions both with and without the MSA information and select the model with the highest average plDDT. We also find that using high plDDT models as templates can increase the speed of AlphaFold2 as implemented in ColabFold. Additionally, we try to demonstrate that as well as having increased overall plDDT, the models are likely to have higher quality structures as judged by two metrics. Availability and implementation:We have implemented our pipeline in NextFlow and it is available in GitHub: https://github.com/FranceCosta/AF2Fix.
The Pfam protein families database is a comprehensive collection of protein domains and families used for genome annotation and protein structure and function analysis (https://www.ebi.ac.uk/interpro/). This update describes major developments in Pfam since 2020, including decommissioning the Pfam website and integration with InterPro, harmonization with the ECOD structural classification, and expanded curation of metagenomic, microprotein and repeat-containing families. We highlight how AlphaFold structure predictions are being leveraged to refine domain boundaries and identify new domains. New families discovered through large-scale sequence similarity analysis of AlphaFold models are described. We also detail the development of Pfam-N, which uses deep learning to expand family coverage, achieving an 8.8% increase in UniProtKB coverage compared to standard Pfam. We discuss plans for more frequent Pfam releases integrated with InterPro and the potential for artificial intelligence to further assist curation. Despite recent advances, many protein families remain to be classified, and Pfam continues working toward comprehensive coverage of the protein universe.
David Binns合作论文数European Bioinformatics Institute8