CRISPR-Cas9 is a powerful genome-editing tool1, but genome-wide off-target activity can hinder therapeutic applications. Negative supercoiling ((-)SC) has been implicated in off-target activity, but a molecular-level understanding is lacking. Here, using (-)SC DNA minicircles, we observe supercoiling-driven structural defects in the DNA that are resolved by Cas9 binding. Cryo-electron microscopy structures of Cas9 bound in both the on-target and off-target configurations highlight that the Cas9 HNH domain is poised in a more catalytically competent conformation. New DNA-RNA mismatch geometries are accommodated across the protospacer and structural plasticity in the protospacer adjacent motif distal region of the protospacer is topology dependent. Together, our study reveals the molecular basis for (-)SC-induced Cas9 targeting and provides a framework for the design of next-generation high-fidelity CRISPR effectors with topological context.
The cryogenic sample-electron microscopy (cryoEM) field has generated significant amounts of 3D Electron Microscopy (3DEM) volumetric data and associated metadata, now comprehensively archived in the Electron Microscopy Data Bank (EMDB - www.emdatabank.org ) and the Electron Microscopy Public Image Archive (EMPIAR - www.empiar.org ). Harnessing the full potential of these resources requires robust, flexible, and publicly accessible tools for data exploration, analysis and retrieval. Here, we present Chart Builder, an interactive web-based platform that enables researchers to create customizable, publication-quality visualizations directly from archival metadata, validation assessments, and cross-reference annotations. Chart Builder integrates the same query-driven and flexible Solr search system as EMDB search, into a user interface with tools to assist users to filter, group, and compare data without programming expertise. It supports multiple chart types (including line, bar, area, scatter (2D and 3D), histogram, bubble, pie, geographic and Venn diagrams) with customizable axes, data series, and statistical operators. Users can apply global filters, define temporal, categorical, or custom query-based axes, and explore multi-dimensional relationships interactively. Chart data-points are linked to their underlying datasets, such that visualisation interaction opens entry-level or archive-level search results for inspection and datasets from charts may be exported in several ways. Findability, accessibility, interoperability and reusability of data are facilitated by these direct access and export mechanisms, including HTML embedding, persistent URL sharing and chart/data download options. By combining interactivity and ease of use with up-to-date access to the EMDB and EMPIAR archive metadata, both computational and experimental communities may explore and visualize current metadata and export to formats for further analysis or as publication-ready figures. Chart Builder promotes community-driven data analysis and empowers users to evaluate trends in the biological 3DEM field. Chart Builder is freely accessible and fully integrated into the EMDB website at https://www.ebi.ac.uk/emdb/statistics/builder/ .
Effectively communicating knowledge related to molecular structures and their associated data remains a challenge, as traditional static figures limit interactivity and professional visualization tools often require substantial expertise. MolViewStories addresses these limitations by providing an open-source, web-based platform for creating and sharing interactive, narrative-driven molecular visualizations. The platform uses the MolViewSpec standard for reproducible scene specification, extended to support animations, interactive descriptions, and synchronized audio commentary. Visualization is powered by the Mol* Viewer, which leverages Web Graphics Library (WebGL) for efficient 3D rendering and WebXR for immersive virtual and augmented reality experiences. Users can construct molecular narratives through an intuitive graphical interface or a command-line workflow, enabling both exploratory and automated use. Completed stories can be shared online, exported locally, or distributed as self-contained packages that remain functional indefinitely. Each story is assigned a persistent uniform resource locator (URL) and can be modified or reused as a template, promoting collaboration and community-driven content creation. We demonstrate the capabilities of MolViewStories through a diverse set of narratives illustrating its broad applicability to research communication, education, and public outreach. Together, these examples highlight how interactive, web-based storytelling can make molecular data more accessible, reproducible, and engaging. MolViewStories is freely available at https://molstar.org/mol-view-stories with open-source code accessible at https://github.com/molstar/mol-view-stories.
PDB-IHM is a branch of the Protein Data Bank (PDB), a Worldwide Protein Data Bank (wwPDB) Core Archive, that expands its scope by allowing for additional biomolecular structure representations and types of experimental information (i.e., integrative/hybrid structure models). As of October 2025, PDB-IHM contained 374 entries, benefitting from multi-scale and multi-state representations and 17 types of experimental data. These structure models are assigned PDB accession codes and are archived alongside other experimental structures in the PDB. Rigorous interpretation of a structure model requires assessment of underlying data quality, consistency with the input data, and estimates of positional uncertainty of its components. Herein, we present the IHMValidation pipeline (https://validate.pdb-ihm.org; https://github.com/salilab/IHMValidation) based on recommendations from the wwPDB Integrative Methods Task Force plus the small-angle scattering (SAS), chemical crosslinking mass spectrometry (crosslinking-MS), and cryo-electron microscopy and tomography (3DEM) communities. The IHMValidation report (available in both PDF and HTML formats) comprises six sections: (i) overview; (ii) model details; (iii) data quality assessments; (iv) local geometry assessments (i.e., model quality); (v) fit of the model to the data used to generate it; and (vi) fit of the model to the data used for validation. Future expansions of the IHMValidation pipeline will: (i) reflect recommendations coming from additional experimental communities, including Förster resonance energy transfer (FRET) and hydrogen/deuterium exchange MS (HDX-MS); (ii) include other validation criteria, such as Bayesian likelihoods for the data; and (iii) represent estimates of structure model uncertainty based on the variation among alternative models satisfying input data.
Structures of many large biomolecular assemblies are now being determined using integrative approaches. In these approaches, information derived from multiple experimental and computational methods is combined to compute three-dimensional structures of multi-protein complexes and other macromolecular machines. A standalone prototype data resource for integrative structures called PDB-Dev was built, based on recommendations of the Integrative and Hybrid Methods (IHM) Task Force of the Worldwide Protein Data Bank (wwPDB). This effort included developing data standards and software tools for collecting, curating, validating, visualizing, archiving, and disseminating integrative structures that span diverse spatiotemporal scales and conformational states. Mechanisms have been created to validate integrative structures based on the experimental data underpinning them. Building upon this foundational framework, PDB-Dev has been further expanded to handle large dynamic macromolecular systems and integrative structures that combine, for example, experimental restraints with atomic coordinates computed by machine learning algorithms. Data standards and supporting tools have also been extended to capture information about biomolecular dynamics, such as conformational transitions and related kinetic data derived from biophysical methods. Recently, PDB-Dev was unified with the PDB archive and rebranded as PDB-IHM (pdb-ihm.org), further promoting FAIR (Findable, Accessible, Interoperable, and Reusable) principles of data stewardship for integrative structural biology.
Atomic coordinate models are important for the interpretation of 3D maps produced with cryoEM and cryoET (3D electron microscopy; 3DEM). In addition to visual inspection of such maps and models, quantitative metrics can inform about the reliability of the atomic coordinates, in particular how well the model is supported by the experimentally determined 3DEM map. A recently introduced metric, Q-score, was shown to correlate well with the reported resolution of the map for well fitted models. Here, we present new statistical analyses of Q-score based on its application to ∼10 000 maps and models archived in the EMDB (Electron Microscopy Data Bank) and PDB (Protein Data Bank). Further, we introduce two new metrics based on Q-score to represent each map and model relative to all entries in the EMDB and those with similar resolution. We explore through illustrative examples of proteins, nucleic acids and small molecules how Q-scores can indicate whether the atomic coordinates are well fitted to 3DEM maps and also whether some parts of a map may be poorly resolved due to factors such as molecular flexibility, radiation damage and/or conformational heterogeneity. These examples and statistical analyses provide a basis for how Q-scores can be interpreted effectively in order to evaluate 3DEM maps and atomic coordinate models prior to publication and archiving.
The Protein Data Bank (PDB) is the global repository for public-domain experimentally determined 3D biomolecular structural information. The archival nature of the PDB presents certain challenges pertaining to updating or adding associated annotations from trusted external biodata resources. While each Worldwide PDB (wwPDB) partner has made best efforts to provide up-to-date external annotations, accessing and integrating information from disparate wwPDB data centers can be an involved process. To address this issue, the wwPDB has established the PDB Next Generation (or NextGen) Archive, developed to centralize and streamline access to enriched structural annotations from wwPDB partners and trusted external sources. At present, the NextGen Archive provides mappings between experimentally determined 3D structures of proteins and UniProt amino acid sequences, domain annotations from Pfam, SCOP2 and CATH databases and intra-molecular connectivity information. Since launch, the PDB NextGen Archive has seen substantial user engagement with over 3.5 million data file downloads, ensuring researchers have access to accurate, up-to-date and easily accessible structural annotations. Database URL: http://www.wwpdb.org/ftp/pdb-nextgen-archive-site.
Electron cryo-microscopy image-processing workflows are typically composed of elements that may, broadly speaking, be categorized as high-throughput workloads which transition to high-performance workloads as preprocessed data are aggregated. The high-throughput elements are of particular importance in the context of live processing, where an optimal response is highly coupled to the temporal profile of the data collection. In other words, each movie should be processed as quickly as possible at the earliest opportunity. The high level of disconnected parallelization in the high-throughput problem directly allows a completely scalable solution across a distributed computer system, with the only technical obstacle being an efficient and reliable implementation. The cloud computing frameworks primarily developed for the deployment of high-availability web applications provide an environment with a number of appealing features for such high-throughput processing tasks. Here, an implementation of an early-stage processing pipeline for electron cryotomography experiments using a service-based architecture deployed on a Kubernetes cluster is discussed in order to demonstrate the benefits of this approach and how it may be extended to scenarios of considerably increased complexity.
In January 2020, a workshop was held at EMBL-EBI (Hinxton, UK) to discuss data requirements for the deposition and validation of cryoEM structures, with a focus on single-particle analysis. The meeting was attended by 47 experts in data processing, model building and refinement, validation, and archiving of such structures. This report describes the workshop's motivation and history, the topics discussed, and the resulting consensus recommendations. Some challenges for future methods-development efforts in this area are also highlighted, as is the implementation to date of some of the recommendations.
The widespread adoption of cryoEM technologies for structural biology has pushed the discipline to new frontiers. A significant worldwide effort has refined the single-particle analysis (SPA) workflow into a reasonably standardized procedure. Significant investments of development time have been made, particularly in sample preparation, microscope data-collection efficiency, pipeline analyses and data archiving. The widespread adoption of specific commercial microscopes, software for controlling them and best practices developed at facilities worldwide has also begun to establish a degree of standardization to data structures coming from the SPA workflow. There is opportunity to capitalize on this moment in the maturation of the field, to capture metadata from SPA experiments and correlate the metadata with experimental outcomes, which is presented here in a set of programs called EMinsight. This tool aims to prototype the framework and types of analyses that could lead to new insights into optimal microscope configurations as well as to define methods for metadata capture to assist with the archiving of cryoEM SPA data. It is also envisaged that this tool will be useful to microscope operators and facilities looking to rapidly generate reports on SPA data-collection and screening sessions.
IHMCIF (github.com/ihmwg/IHMCIF) is a data information framework that supports archiving and disseminating macromolecular structures determined by integrative or hybrid modeling (IHM), and making them Findable, Accessible, Interoperable, and Reusable (FAIR). IHMCIF is an extension of the Protein Data Bank Exchange/macromolecular Crystallographic Information Framework (PDBx/mmCIF) that serves as the framework for the Protein Data Bank (PDB) to archive experimentally determined atomic structures of biological macromolecules and their complexes with one another and small molecule ligands (e.g., enzyme cofactors and drugs). IHMCIF serves as the foundational data standard for the PDB-Dev prototype system, developed for archiving and disseminating integrative structures. It utilizes a flexible data representation to describe integrative structures that span multiple spatiotemporal scales and structural states with definitions for restraints from a variety of experimental methods contributing to integrative structural biology. The IHMCIF extension was created with the benefit of considerable community input and recommendations gathered by the Worldwide Protein Data Bank (wwPDB) Task Force for Integrative or Hybrid Methods (wwpdb.org/task/hybrid). Herein, we describe the development of IHMCIF to support evolving methodologies and ongoing advancements in integrative structural biology. Ultimately, IHMCIF will facilitate the unification of PDB-Dev data and tools with the PDB archive so that integrative structures can be archived and disseminated through PDB.
CRISPR Cas9 is a powerful tool used for genome editing, however spurious off-targeting raise implications for safe use in therapeutic genome-editing applications. Recently, off-targets have been captured through X-ray and Cryo-EM but by using short linear substrates they miss the topological constraint that Cas9 would experience within the human cell. To overcome this, we utilise minicircle DNA (mcDNA) to capture, characterise and visualise Cas9 under topological constraint and under the influence of negative supercoiling.
Clathrins are self-assembling cytoplasmic proteins that serve to mediate membrane trafficking. At intracellular membranes, individual clathrin subunits surround the invaginating membrane to form a protein coat that assists in cargo capture and vesicle formation. The protein coat is a polyhedral, multimeric assembly of clathrin triskelia (3-legged structures partly composed of 3 clathrin heavy chains (CHC)). There are two forms of CHC in vertebrates CHC17 and CHC22 which have distinct cellular functions. CHC17 is implicated in receptor-mediated endocytosis at the plasma membrane and organelle biogenesis at the trans-Golgi network. During clathrin-mediated endocytosis, CHC17 cannot recognise membrane or cargo and so an adaptor protein binds the membrane, selects the cargo, and associates with clathrin leading to pit formation. Several adaptor proteins have clathrin binding sites and colocalize with clathrin structures in cells. Our structural knowledge of these adaptor-clathrin interactions, and their functional importance, is unclear. We analysed the cryo-EM structure of CHC17 cages assembled in the presence of the clathrin-binding subunit (β2-appendage) of assembly polypeptide-2 (AP2) (the adaptor protein that is thought to primarily initiate clathrin recruitment). We found that the β2-appendage binds in at least two positions in the cage. We propose that β2-appendage binding to more than one triskelion is a key feature of the system and likely explains why clathrin assembly is driven by AP2. These data then led us to ask: is multi-modal binding a fundamental property of clathrin-adaptor interactions? CHC22 acts to sequester the GLUT4 glucose transporter in an insulin-responsive compartment - a behaviour critical to controlling blood sugar levels. Understanding how clathrin self-assembles into basket-like structures to facilitate such cellular functions will significantly advance our understanding of clathrin biology. To this end, we endeavor to map the molecular structure of CHC22.
Cryo-electron tomography (cryo-ET) has been gaining momentum in recent years, especially since the introduction of direct electron detectors, improved automated acquisition strategies, preparative techniques that expand the possibilities of what the electron microscope can image at high-resolution using cryo-ET and new subtomogram averaging software. Additionally, data acquisition has become increasingly streamlined, making it more accessible to many users. The SARS-CoV-2 pandemic has further accelerated remote cryo-electron microscopy (cryo-EM) data collection, especially for single-particle cryo-EM, in many facilities globally, providing uninterrupted user access to state-of-the-art instruments during the pandemic. With the recent advances in Tomo5 (software for 3D electron tomography), remote cryo-ET data collection has become robust and easy to handle from anywhere in the world. This article aims to provide a detailed walk-through, starting from the data collection setup in the tomography software for the process of a (remote) cryo-ET data collection session with detailed troubleshooting. The (remote) data collection protocol is further complemented with the workflow for structure determination at near-atomic resolution by subtomogram averaging with emClarity, using apoferritin as an example.
Balanced proliferation-quiescence decisions are vital during normal development and in tissue homeostasis, and their dysregulation underlies tumorigenesis. Entry into proliferative cycles is driven by Cyclin/Cyclin-dependent kinases (Cdks). Conserved Cdk inhibitors (CKIs) p21Cip1/Waf1, p27Kip1, and p57Kip2 bind to Cyclin/Cdks and inhibit Cdk activity. p27 tyrosine phosphorylation, in response to mitogenic signaling, promotes activation of CyclinD/Cdk4 and CyclinA/Cdk2. Tyrosine phosphorylation is conserved in p21 and p57, although the number of sites differs. We use molecular-dynamics simulations to compare the structural changes in Cyclin/Cdk/CKI trimers induced by single and multiple tyrosine phosphorylation in CKIs and their impact on CyclinD/Cdk4 and CyclinA/Cdk2 activity. Despite shared structural features, CKI binding induces distinct structural responses in Cyclin/Cdks and the predicted effects of CKI tyrosine phosphorylation on Cdk activity are not conserved across CKIs. Our analyses suggest how CKIs may have evolved to be sensitive to different inputs to give context-dependent control of Cdk activity.
Clathrin-coated pits are formed by the recognition of membrane and cargo by the heterotetrameric AP2 complex and the subsequent recruitment of clathrin triskelia. A potential role for AP2 in coated-pit assembly beyond initial clathrin recruitment has not been explored. Clathrin binds the β2 subunit of AP2, and several binding sites on β2 and on the clathrin heavy chain have been identified, but our structural knowledge of these interactions is incomplete and their functional importance during endocytosis is unclear. Here, we analysed the cryo-EM structure of clathrin cages assembled in the presence of β2 hinge and appendage (β2HA) domains. We find that the β2-appendage binds in at least two positions in the cage, demonstrating that multi-modal binding is a fundamental property of clathrin-AP2 interactions. In one position, β2-appendage cross-links two adjacent terminal domains from different triskelia below the vertex. Functional analysis of β2HA-clathrin interactions reveals that endocytosis requires two clathrin interaction sites: a clathrin-box motif on the hinge and the “sandwich site” on the appendage, with the appendage “platform site” having less importance. From these studies and the work of others, we propose that β2-appendage binding to more than one clathrin triskelion is a key feature of the system and likely explains why clathrin assembly is driven by AP2.
MiDAC is one of seven distinct, large multi-protein complexes that recruit class I histone deacetylases to the genome to regulate gene expression. Despite implications of involvement in cell cycle regulation and in several cancers, surprisingly little is known about the function or structure of MiDAC. Here we show that MiDAC is important for chromosome alignment during mitosis in cancer cell lines. Mice lacking the MiDAC proteins, DNTTIP1 or MIDEAS, die with identical phenotypes during late embryogenesis due to perturbations in gene expression that result in heart malformation and haematopoietic failure. This suggests that MiDAC has an essential and unique function that cannot be compensated by other HDAC complexes. Consistent with this, the cryoEM structure of MiDAC reveals a unique and distinctive mode of assembly. Four copies of HDAC1 are positioned at the periphery with outward-facing active sites suggesting that the complex may target multiple nucleosomes implying a processive deacetylase function.