
The Ensembl project (https://www.ensembl.org) is a public and open resource providing access to genomes, annotations, high-quality tools, and methods applicable to species from across the tree of life. This year has witnessed nearly a doubling in our rate of annotation and genome release, with 1927 new genomes released, with the total number of genomes now standing at 37 546. This includes expanded support for the human and barley pangenomes. We also present two new interfaces providing improved mechanisms to explore and interrogate genome regulation annotations. As our focus remains on sustainable scaling, we have archived Ensembl Rapid Release and accelerated the move to the new Ensembl platform. Ensembl release 116 (Q1-2026) will be the last release on the current platform.
Thousands of short open reading frames (sORFs) are translated outside of annotated coding sequences. Recent studies have pioneered searching for sORF-encoded microproteins in mass spectrometry (MS)-based proteomics and peptidomics datasets. Here, we assessed literature-reported MS-based identifications of unannotated human proteins. We find that studies vary by three orders of magnitude in the number of unannotated proteins they report. Of nearly 10,000 reported sORF-encoded peptides, 96% were unique to a single study, and 12% mapped to annotated proteins or proteoforms. Manual curation of a benchmark dataset of 406 manually evaluated spectra from 204 sORF-encoded proteins revealed large variation in peptide-spectrum match (PSM) quality between studies, with immunopeptidomics studies generally reporting higher quality PSMs than conventional enzymatic digests of whole cell lysates. We estimate that 65% of predicted sORF-encoded protein detections in immunopeptidomics studies were supported by high-quality PSMs versus 7.8% in non-immunopeptidomics datasets. Our work stresses the need for standardized protocols and analysis workflows to guide future advancements in microprotein detection by MS towards uncovering how many human microproteins exist.
The European Nucleotide Archive (ENA; https://www.ebi.ac.uk/ena), hosted at the European Molecular Biology Laboratory's European Bioinformatics Institute (EMBL-EBI), remains a global, open-access platform for the submission, archiving, dissemination, and reuse of nucleotide sequence data. In 2025, ENA continues to advance its mission of fostering FAIR (findable, accessible, interoperable, reusable) data principles through innovations in interoperability, scalability, and global engagement, providing infrastructure for a rapidly growing volume of data across diverse domains. This article highlights the key developments in 2025, including the progress of the technical transformation, enhanced support for large-scale biodiversity projects, and the implementation of the International Nucleotide Sequence Database Collaboration Global Participation Initiative. We also discuss infrastructure enhancements to handle exponential data growth and improve user experiences and data discovery.
Gramene (gramene.org) is a comprehensive reference database for comparative plant genomics and pathway analysis, integrating functional annotations, evidence-based curated pathways and their projections, and multi-omics datasets. Since our last report, Gramene has added crop-specific pan-genome portals for maize, sorghum, rice, and grapevine. These pan-genome portals host population-scale datasets and multiple assembled genomes per species, all anchored by shared reference genomes. Importantly, these portals now adopt standardized rsIDs for genetic variants, advancing FAIR data principles and enabling cross-database interoperability. The main site is now Gramene Plants, emphasizing its broad genome coverage. Release 69 features 233 reference genomes, curated pathways for 139 species, expression data from 1026 studies across 27 species, and genetic variation data mapped to 27 genomes from 19 species. Key updates to the integrated search functionality include embedded expression viewers from the Bio-Analytic Resource for Plant Biology and EMBL-EBI Expression Atlas, a literature-curated catalog of gene functions, and a new Germplasm tab linking accessions with loss-of-function alleles to seed repositories. These advances reinforce Gramene as a comprehensive platform for exploring plant genomic diversity, gene function, and evolutionary conservation across the Green Tree of Life and within key agricultural species.
This is version 1.3.2 of the image CIF dictionary (imgCIF) and crystallographic binary file (CBF) dictionary. Use of this dictionary is described in Chapter 3.7. The CBF format is described in Chapter 2.3 and Chapter 5.6 describes a software library for manipulating image data. The data names defined here extend the macromolecular CIF dictionary (Chapter 4.5). Keywords: crystallography; CIF; CIF dictionaries; Crystallographic Information File; imgCIF; image-supporting Crystallographic Information File; data names; data categories; DDL2