Advances in artificial intelligence (AI)-driven bioinformatics promise democratized discovery, yet major inequities persist. Equitable adoption of bioinformatics tools will require sustained investment in infrastructure, training, institutions, and global communities, not just access.
The AlphaFold Protein Structure Database (AFDB; https://alphafold.ebi.ac.uk), developed by EMBL-EBI and Google DeepMind, provides open access to hundreds of millions of high-accuracy protein structure predictions, transforming research in structural biology and the wider life sciences. Since its launch, AFDB has become a widely used bioinformatics resource, integrated into major databases, visualization platforms, and analysis pipelines. Here, we report the update of the database to align with the UniProt 2025_03 release, along with a comprehensive redesign of the entry page to enhance usability, accessibility, and structural interpretation. The new design integrates annotations directly with an interactive 3D viewer and introduces dedicated domains and summary tabs. Structural coverage has also been updated to include isoforms plus underlying multiple sequence alignments. Data are available through the website, FTP, Google Cloud, and updated APIs. Together, these advances reinforce AFDB as a sustainable resource for exploring protein sequence-structure relationships.
The AlphaFold Protein Structure Database (https://alphafold.ebi.ac.uk/) has made significant strides in enhancing its utility and accessibility for the life science research community. The recent integration of AlphaMissense predictions enables access to the pathogenicity of human protein missense variants, with an innovative and interactive heatmap and 3D visualisation that display variant data at the residue level. Users can now toggle between structure model quality (pLDDT) and average pathogenicity scores, providing insights into the implications of specific residue changes. The Foldseek integration offers a rapid and accurate method for protein structure searches and comparisons. Bulk data download options further facilitate comprehensive data analysis and integration with other computational tools. The 3D-Beacons framework (https://www.ebi.ac.uk/pdbe/pdbe-kb/3dbeacons/) has also been enhanced with detailed annotation endpoints (such as AlphaMissense data) and integrates LevyLab’s dataset of homomeric AlphaFold 2 models. These advancements significantly improve the functionality and accessibility of these resources, enabling discoveries using structure data.
The AlphaFold Education Summit (AFES) was a collaborative initiative designed to enable researchers and educators globally to fully harness the transformative potential of AlphaFold. This event provided a platform to engage with the possibilities and challenges of integrating AlphaFold into scientific research and education. Special focus was given to addressing the unique needs and limitations faced by educators and researchers in under-resourced regions, aiming to enhance both accessibility and practical application in diverse contexts. By providing access to bespoke training, resources, and a supportive community, we aimed to empower scientists to advance their research and address local, regional and global scientific challenges. The goal of the summit was to train and empower researchers, educators and others to upskill and train researchers working in under-resourced regions in the use and application of AlphaFold to their research.
The rapid advancement of automatic structure solution methods, driven by the availability of high-quality predicted structures from AlphaFold and the growing adoption of multi-crystal and serial experiments, has created a pressing need for streamlining routine operations, automating structure solution projects, and efficiently handling large volumes of data. Modern software solutions must be both robust and user-friendly, supporting manual workflows while enabling high-throughput operations to keep pace with the high data collection rates of modern beamlines. Here, we present new developments in CCP4 Cloud that address these challenges by providing predefined and customizable automatic workflows, which can be seamlessly integrated with experimental facilities, offering a powerful solution for modern macromolecular crystallography. CCP4 Cloud is available as a public service at https://cloud.ccp4.ac.uk.
Two years on from the initial release of AlphaFold, we have seen its widespread adoption as a structure prediction tool. Here, we discuss some of the latest work based on AlphaFold, with a particular focus on its use within the structural biology community. This encompasses use cases like speeding up structure determination itself, enabling new computational studies, and building new tools and workflows. We also look at the ongoing validation of AlphaFold, as its predictions continue to be compared against large numbers of experimental structures to further delineate the model’s capabilities and limitations.
Abstract The AlphaFold Database Protein Structure Database (AlphaFold DB, https://alphafold.ebi.ac.uk) has significantly impacted structural biology by amassing over 214 million predicted protein structures, expanding from the initial 300k structures released in 2021. Enabled by the groundbreaking AlphaFold2 artificial intelligence (AI) system, the predictions archived in AlphaFold DB have been integrated into primary data resources such as PDB, UniProt, Ensembl, InterPro and MobiDB. Our manuscript details subsequent enhancements in data archiving, covering successive releases encompassing model organisms, global health proteomes, Swiss-Prot integration, and a host of curated protein datasets. We detail the data access mechanisms of AlphaFold DB, from direct file access via FTP to advanced queries using Google Cloud Public Datasets and the programmatic access endpoints of the database. We also discuss the improvements and services added since its initial release, including enhancements to the Predicted Aligned Error viewer, customisation options for the 3D viewer, and improvements in the search engine of AlphaFold DB.
The Collaborative Computational Project No. 4 (CCP4) is a UK-led international collective with a mission to develop, test, distribute and promote software for macromolecular crystallography. The CCP4 suite is a multiplatform collection of programs brought together by familiar execution routines, a set of common libraries and graphical interfaces. The CCP4 suite has experienced several considerable changes since its last reference article, involving new infrastructure, original programs and graphical interfaces. This article, which is intended as a general literature citation for the use of the CCP4 software suite in structure determination, will guide the reader through such transformations, offering a general overview of the new features and outlining future developments. As such, it aims to highlight the individual programs that comprise the suite and to provide the latest references to them for perusal by crystallographers around the world.
Nowadays, progress in the determination of three-dimensional macromolecular structures from diffraction images is achieved partly at the cost of increasing data volumes. This is due to the deployment of modern high-speed, high-resolution detectors, the increased complexity and variety of crystallographic software, the use of extensive databases and high-performance computing. This limits what can be accomplished with personal, offline, computing equipment in terms of both productivity and maintainability. There is also an issue of long-term data maintenance and availability of structure-solution projects as the links between experimental observations and the final results deposited in the PDB. In this article, CCP4 Cloud, a new front-end of the CCP4 software suite, is presented which mitigates these effects by providing an online, cloud-based environment for crystallographic computation. CCP4 Cloud was developed for the efficient delivery of computing power, database services and seamless integration with web resources. It provides a rich graphical user interface that allows project sharing and long-term storage for structure-solution projects, and can be linked to data-producing facilities. The system is distributed with the CCP4 software suite version 7.1 and higher, and an online publicly available instance of CCP4 Cloud is provided by CCP4.
Structure solution in macromolecular crystallography is not always a straightforward process and it may be rather difficult for structural biologists without advanced training.A trained crystallographer exploits an extended set of approaches and tricks, based on the analysis of several indicators, general assessment of the case, and developed strategies for dealing with a particular class of problems.If such an approach, used by an expert, can be formalised in terms of an algorithm, then it can be implemented as a computer program or automated user advice system to help users solving structures quicker and with a higher success rate.It is not surprising then, that programs for macromolecular crystallography are moving towards full automation, taking off burden from researchers and lowering the entry barriers for novice users.During the last few years, substantial progress has been made towards automation of the whole macromolecular structure determination process.There are a number of examples of successful automatic solutions for various stages of structure determination, such as molecular replacement, experimental phasing, and refinement [1][2][3][4][5].In this communication, we report two novel automation features implemented in CCP4 Cloud [6], the new system for solving macromolecular structures online, released with CCP4 Software Suite 7.1 in 2020.Automated user advice framework, named Verdicts, provides simple graphical representation of results quality with detailed analysis of points for improvement as part of every task (e.g., refinement) report.The analysis includes suggestions on what could be done in order to improve the result (i.e., which parameters could be optimised).Then, the task can be re-run with the suggested parameters, which can be further adjusted by the user as appropriate.Another automation feature, Workflows, was designed for unfolding structure solution Projects, or their parts, automatically using user-supplied data.Such automatically initiated and unfolded Projects may include a number of tasks, arranged in branching Project Trees as if this were done by the user themselves.In common cases without complications, this may result in structure solved, and if not, then a starting Project is offered to the user for analysis and further manipulations, where simple, first-order structure solution attempts are already performed.Any task or branch of the starting Project may be cloned and re-run with optimized parameters, and new tasks may be added as needed.Workflows combine automation and human expert skills, and, therefore, represent an excellent starting point for users with different level of expertise, ranging from novices to experienced crystallographers.Workflows are particularly useful in a common case of processing large sets of isomorphous crystals, because, once structure is solved in one crystal, the process is well-repeatable in systems with moderate modifications.
Regularisers stabilise macromolecular refinement, ensuring consistency between derived models and prior knowledge. Restraints representing chemical information help local structure adopt chemically reasonable conformations. At lower resolutions, supplementary restraints are used: encouraging consistency with models of homologous structures, formation of hydrogen bonding networks, nucleotide base pairing and stacking. Jelly-body restraints stabilise model refinement without injecting externally derived information. An anharmonic penalty function is used in order to control robustness to outliers due to inconsistencies between data and prior. Additional restraints and robust estimation are also useful for model building, increasing real space refinement convergence radius and stability.
The Collaborative Computational Project Number 4 in Protein Crystallography (CCP4) exists to maintain, develop and provide world-class software that allows researchers to determine macromolecular structures by X-ray crystallography and other biophysical techniques.Over 40 years of existence, CCP4 Software was assembled and distributed as an integrated Suite of programs, traditionally operated via CCP4i (2) GUI in Linux, OSX and Windows platforms.Modern trends in computing suggest a fast-growing interest to mobile platforms and cloud solutions for data management and operations in practically all areas.Answering to these trends, CCP4 releases beta version of CCP4 Cloud, developed for essentially remote and distributed deployment of CCP4 Software.CCP4 Cloud allows a user to keep all necessary data and projects in the cloud and perform all scope of crystallographic computations, from image processing to final refinement, ligand fitting and deposition, remotely via a common web-browser, optionally complemented with CCP4 Cloud Client for interactive model building with Coot.The talk will present architectural solutions and key features of CCP4 Cloud such as ability to seamlessly import data from 3rd party sites (e.g., synchrotrons), high scalability of computational background, convenient (big) data management for multiple users, rich graphical interface with built-in molecular graphics, enhanced data and structure solution pathway provenance.CCP4 Cloud may be used from CCP4 Web portal with any device running a modern web-browser (including tablets and smartphones).All the source code is open and freely available for installation elsewhere to serve local researchers in a lab, or institution, or pharma, or a synchrotron.
Recent advances in instrumentation and software have resulted in cryo-EM rapidly becoming the method of choice for structural biologists, especially for those studying the three-dimensional structures of very large macromolecular complexes. In this contribution, the tools available for macromolecular structure refinement into cryo-EM reconstructions that are available via CCP-EM are reviewed, specifically focusing on REFMAC5 and related tools. Whilst originally designed with a view to refinement against X-ray diffraction data, some of these tools have been able to be repurposed for cryo-EM owing to the same principles being applicable to refinement against cryo-EM maps. Since both techniques are used to elucidate macromolecular structures, tools encapsulating prior knowledge about macromolecules can easily be transferred. However, there are some significant qualitative differences that must be acknowledged and accounted for; relevant differences between these techniques are highlighted. The importance of phases is considered and the potential utility of replacing inaccurate amplitudes with their expectations is justified. More pragmatically, an upper bound on the correlation between observed and calculated Fourier coefficients, expressed in terms of the Fourier shell correlation between half-maps, is demonstrated. The importance of selecting appropriate levels of map blurring/sharpening is emphasized, which may be facilitated by considering the behaviour of the average map amplitude at different resolutions, as well as the utility of simultaneously viewing multiple blurred/sharpened maps. Features that are important for the purposes of computational efficiency are discussed, notably the Divide and Conquer pipeline for the parallel refinement of large macromolecular complexes. Techniques that have recently been developed or improved in Coot to facilitate and expedite the building, fitting and refinement of atomic models into cryo-EM maps are summarized. Finally, a tool for symmetry identification from a given map or coordinate set, ProSHADE, which can identify the point group of a map and thus may be used during deposition as well as during molecular visualization, is introduced.
Refinement is a process that involves bringing into agreement the structural model, available prior knowledge and experimental data. To achieve this, the refinement procedure optimizes a posterior conditional probability distribution of model parameters, including atomic coordinates, atomic displacement parameters (B factors), scale factors, parameters of the solvent model and twin fractions in the case of twinned crystals, given observed data such as observed amplitudes or intensities of structure factors. A library of chemical restraints is typically used to ensure consistency between the model and the prior knowledge of stereochemistry. If the observation-to-parameter ratio is small, for example when diffraction data only extend to low resolution, the Bayesian framework implemented in REFMAC5 uses external restraints to inject additional information extracted from structures of homologous proteins, prior knowledge about secondary-structure formation and even data obtained using different experimental methods, for example NMR. The refinement procedure also generates the 'best' weighted electron-density maps, which are useful for further model (re) building. Here, the refinement of macromolecular structures using REFMAC5 and related tools distributed as part of the CCP4 suite is discussed.
This review describes some of the problems encountered during low-resolution refinement and map calculation. Refinement is considered as an application of Bayes' theorem, allowing combination of information from various sources including crystallographic experimental data and prior chemical and structural knowledge. The sources of prior knowledge relevant to macromolecules include basic chemical information such as bonds and angles, structural information from reference models of known homologs, knowledge about secondary structures, hydrogen bonding patterns, and similarity of non-crystallographically related copies of a molecule. Additionally, prior information encapsulating local conformational conservation is exploited, keeping local interatomic distances similar to those in the starting atomic model. The importance of designing an accurate likelihood function-the only link between model parameters and observed data-is emphasized. The review also reemphasizes the importance of phases, and describes how the use of raw observed amplitudes could give a better correlation between the calculated and "true" maps. It is shown that very noisy or absent observations can be replaced by calculated structure factors, weighted according to the accuracy of the atomic model. This approach helps to smoothen the map. However, such replacement should be used sparingly, as the bias toward errors in the model could be too much to avoid. It is in general recommended that, whenever a new map is calculated, map quality should be judged by inspection of the parts of the map where there is no atomic model. It is also noted that it is advisable to work with multiple blurred and sharpened maps, as different parts of a crystal may exhibit different degrees of mobility. Doing so can allow accurate building of atomic models, accounting for overall shape as well as finer structural details. Some of the results described in this review have been implemented in the programs REFMAC5, ProSMART and LORESTR, which are available as part of the CCP4 software suite.
Poor diffraction quality of macromolecular crystals is a common problem: various types of disorder result in the weakening of high-resolution observations, anisotropic diffraction and other problems.Such datasets have low information content; hence the ratio of the number of observations to adjustable model parameters is small.This ratio could be improved by using complementary information -our prior knowledge about macromolecular structures.Restraints produced based on known related structures could be used as such an additional source of information; ProSMART is one of the programs that can generate restraints using reference protein/RNA/DNA structures as well as for secondary structure elements.These restraints are used by REFMAC5 to stabilise refinement of an atomic model against low-resolution diffraction data.We have tested various refinement strategies and different REFMAC5 [1] and ProSMART [2] parameters on a test set of more than a hundred structures with resolution below 3.0 Å, for which structures of high-resolution homologues are available.We found that refinement with external restraints is sensitive to the selection of homologous structures for restraint generation; some homologues of the same target structure may improve refinement much better than others.The best-performing refinement protocols have been implemented in LORESTR: an automated pipeline for structure refinement at low resolution [3], distributed as part of the CCP4 suite.The pipeline facilitates the fully-automated selection of optimal external restraints from ProSMART for structure refinement by REFMAC5.It can automatically run a BLAST search to identify homologues, and download the corresponding models from the PDB.It automatically detects twinning, and finds the optimal scaling method and parameters for solvent modeling.The pipeline runs a number of refinement protocols in order to find the best protocol for each particular case.In our tests, LORESTR was able to produce substantially better quality models in the vast majority of cases, improving both R-factors and model geometry for 94% of test cases.The dramatic improvement in R-factors and the geometric quality of low-resolution models observed when using the fullyautomated mode of the pipeline demonstrates its potential use for researchers working with low-resolution cases, especially during the initial stages of refinement, or when unable to further progress with refinement.Currently we are working on multi-crystal refinement: treatment of the special case where several low-resolution X-ray diffraction datasets and models are available for a particular protein.We have designed a procedure to co-refine all structures simultaneously, executing multiple concurrent REFMAC5 refinements, generating external restraints for each model using all others, and iterating until convergence.This technique allows information transfer between the structures, which could potentially improve refinement and thus the quality of resulting models.
Lin28A is a post-transcriptional regulator of gene expression that interacts with and negatively regulates the biogenesis of let-7 family miRNAs. Recent data suggested that Lin28A also binds the putative tumor suppressor miR-363, a member of the 106~363 cluster of miRNAs. Affinity for this miRNA and the stoichiometry of the protein-RNA complex are unknown. Characterization of human Lin28's interaction with RNA has been complicated by difficulties in producing stable RNA-free protein. We have engineered a maltose binding protein fusion with Lin28, which binds let-7 miRNA with a Kd of 54.1 ± 4.2 nM, in agreement with previous data on a murine homologue. We show that human Lin28A binds miR-363 with a 1:1 stoichiometry and with a similar, if not higher, affinity (Kd = 16.6 ± 1.9 nM). Further analysis suggests that the interaction of the N-terminal cold shock domain of Lin28A with RNA is salt-dependent, supporting a model in which the cold shock domain allows the protein to sample RNA substrates through transient electrostatic interactions.
Since the ratio of the number of observations to adjustable parameters is small at low resolution, it is necessary to use complementary information for the analysis of such data. ProSMART is a program that can generate restraints for macromolecules using homologous structures, as well as generic restraints for the stabilization of secondary structures. These restraints are used by REFMAC5 to stabilize the refinement of an atomic model. However, the optimal refinement protocol varies from case to case, and it is not always obvious how to select appropriate homologous structure(s), or other sources of prior information, for restraint generation. After running extensive tests on a large data set of low-resolution models, the best-performing refinement protocols and strategies for the selection of homologous structures have been identified. These strategies and protocols have been implemented in the Low-Resolution Structure Refinement (LORESTR) pipeline. The pipeline performs auto-detection of twinning and selects the optimal scaling method and solvent parameters. LORESTR can either use user-supplied homologous structures, or run an automated BLAST search and download homologues from the PDB. The pipeline executes multiple model-refinement instances using different parameters in order to find the best protocol. Tests show that the automated pipeline improves R factors, geometry and Ramachandran statistics for 94% of the low-resolution cases from the PDB included in the test set.
Background: Lin28 proteins are post‐transcriptional regulators of gene expression with multiple roles in development and the regulation of pluripotency in stem cells. Much attention has focussed on Lin28 proteins as negative regulators of let‐7 miRNA biogenesis; a function that is conserved in several animal groups and in multiple processes. However, there is increasing evidence that Lin28 proteins have additional roles, distinct from regulation of let‐7 abundance. We have previously demonstrated that lin28 proteins have functions associated with the regulation of early cell lineage specification in Xenopus embryos, independent of a lin28/let‐7 regulatory axis. However, the nature of lin28 targets in Xenopus development remains obscure. Results: Here, we show that mir‐17∼92 and mir‐106∼363 cluster miRNAs are down‐regulated in response to lin28 knockdown, and RNAs from these clusters are co‐expressed with lin28 genes during germ layer specification. Mature miRNAs derived from pre‐mir‐363 are most sensitive to lin28 inhibition. We demonstrate that lin28a binds to the terminal loop of pre‐mir‐363 with an affinity similar to that of let‐7, and that this high affinity interaction requires to conserved a GGAG motif. Conclusions: Our data suggest a novel function for amphibian lin28 proteins as positive regulators of mir‐17∼92 family miRNAs. Developmental Dynamics 245:34–46, 2016. © 2015 Wiley Periodicals, Inc.