We present a tour de force of atomistic molecular dynamics simulations involving the coordinated effort of 14 research groups of the Ascona B-DNA Consortium (ABC). This initiative provides a complete characterization of the 2080 DNA hexamers embedded in 190 carefully selected 20-mer duplexes, each simulated in replicate for at least 10 microseconds in explicit solvent. The consortium generates 0.25 petabytes of data, capturing millisecond-scale ensembles at the oligomer level and dynamics up to 10-1 seconds at the base-pair level. Analysis yields a comprehensive description of sequence-dependent DNA properties, including rare events such as backbone transitions, reversible base-pair changes, and partial unfolding. Processing these atomistic ensembles reveals a hidden physical code of DNA, helping explain rules of genome composition and evolution beyond coding regions. This community effort delivers unprecedented, validated FAIR data to support coarse-grained and AI models of DNA at cellular scale.
Linear peptides play essential roles in biology and drug discovery, frequently mediating protein-protein interactions through short, flexible motifs. However, their structural plasticity─ranging from disordered to context-dependent folding─makes them challenging targets for molecular simulations. In this work, we benchmark the performance of 11 popular and emerging fixed-charge force fields across a curated set of 12 peptides spanning structured miniproteins, context-sensitive epitopes, and disordered sequences. Each peptide was simulated from both folded (200 ns) and extended (10 μs) states to assess stability, folding behavior, and force field biases. Our analysis reveals consistent trends: some force fields exhibit strong structural bias, others allow reversible fluctuations, and no single model performs optimally across all systems. The study highlights limitations in current force fields' ability to balance disorder and secondary structure, particularly when modeling conformational selection. These results offer practical guidance for peptide modeling and establish a benchmark framework for future force field development and validation in peptide-relevant regimes.
The NMR Exchange Format (NEF) is a community-driven standard for representing NMR experimental data in a consistent, interoperable, and machine-readable form. Built on the STAR syntax, NEF provides a structured framework for storing and exchanging chemical shifts, peak lists, various types of structural restraints, and related metadata, thus allowing for data exchange across software platforms. By enabling direct, lossless transfer of information, NEF simplifies multi-software workflows, improves reproducibility, and supports FAIR (Findable, Accessible, Interoperable, Reusable) data principles. We describe the NEF specification, its current implementation across commonly used NMR software packages, and its application in areas including biomolecular structure determination, metabolomics, and ligand screening. Testing demonstrates that NEF can be used to exchange complete datasets between programs without loss of information or functionality. We also outline recent developments and future directions, such as inclusion of NMR relaxation data and support for non-standard residue topologies. NEFs growing adoption highlights its potential as a unifying standard for NMR data, enabling more efficient, transparent and collaborative research.
Physics-based approaches rely on accurate force fields and efficient sampling to provide mechanistic insight into biomolecular systems. Recent AI-driven advances are transforming this landscape, replacing empirically fitted force fields-built from manually curated atom types-with machine-learned models. Achieving interoperability between these new force fields and established sampling strategies-and validating them on challenging benchmark sets that extend far beyond near-native states-is essential for progress in the field. To this end, we introduce the FoldBind benchmark set, a collection of 18 systems encompassing 14 protein-folding cases and 4 peptide-protein complexes that undergo folding upon binding. This suite expands existing validation efforts by probing both conformational transitions and binding-induced folding, offering a rigorous test for sampling methods and force-field accuracy alike. To explore these systems, we employ the Modeling Employing Limited Data (MELD) framework as the sampling engine. MELD accelerates conformational exploration by integrating ambiguous or noisy physical restraints-for example, the general expectation that proteins form hydrophobic cores-within a Bayesian inference formalism. By balancing exploration (broad conformational search) and exploitation (stabilization of structures consistent with physics and data), MELD efficiently accesses native-like states that are otherwise inaccessible to conventional molecular dynamics. Under identical data conditions, the quality of the force field determines which states are stabilized and whether the correct native basin emerges. Furthermore, the ability of a force field to consistently stabilize the native basin among multiple data-compatible states provides an additional measure of its physical realism. Together, this FoldBind benchmark, along with the information used in MELD, can be used to test and distinguish future force field development efforts.
The exact biological role of mitochondrial supercomplexes remains debated, particularly their role in guiding redox proteins across membranes during energy conversion. We integrate multiscale modeling and single particle cryo-electron microscopy (cryo-EM) to examine electron transfer in mitochondrial supercomplexes composed of complexes III and IV (CIII and CIV). Using bioinformatic and entropy-based methods, we generated structural ensembles capturing conformations of CIII's disordered QCR6 hinge within the yeast CIII2CIV2 supercomplex. Molecular and Brownian Dynamics simulations reveal that these negatively charged hinge states electrostatically couple with redox proteins, promoting their binding and directional diffusion across the membrane on millisecond timescales. Rather than hindering transfer, disorder lowers the diffusion barrier. Anionic lipids reinforce this recognition by retaining a membrane pool of redox proteins when hinge length is critical. Cryo-EM models of ΔQCR6 show large rearrangements, yet maintain a robust electrostatic environment enabling surface-mediated transfer despite reduced charge. Overall, electron carriers confined on bioenergetic membranes follow a refolding-guided diffusion mechanism that enhances supercomplex energy conversion efficiency by nearly 30%.
Bacterial ribonuclease P (RNase P) is an essential ribonucleoprotein enzyme that catalyzes 5' leader removal from precursor tRNAs (ptRNAs) using a catalytic RNA subunit (P RNA) and an essential protein cofactor (RnpA). RnpA binds near the P RNA active site, facilitating catalysis and enhancing ptRNA binding by contacting 5' leader sequences. However, key information regarding how sequence variation, particularly among bacterial pathogens, influences folding, dynamics, P RNA activation, and catalytic function is lacking. Using sequence similarity network (SSN) analyses of >1800 RnpA sequences, we identify two major subfamilies, a Bacilli-specific class (RnpA-1) that associates with divergent Type B P RNAs, and a broader class (RnpA-2) that associates with ancestral Type A P RNAs, with species-specific variation concentrated at the N- and C-termini. Computational and biophysical studies show that RnpA-1 proteins, including those from Staphylococcus aureus and Enterococcus faecium, exhibit greater conformational dynamics and reduced thermal stability relative to RnpA-2 proteins from representative Gram-negative pathogens. Despite these differences, both families comparably enhance binding of ptRNA to their cognate Type A or B P RNA. In contrast, kinetic studies reveal higher kcat values for Type B RNase Ps and rate limiting product release, while Type A enzymes are limited by precatalytic steps. Reconstitution with noncognate subunits produces selective defects in either kcat or KM demonstrating that RnpA identity makes distinct contributions to substrate binding and catalytic activation. These results further define the conserved and divergent structural, dynamic and functional features of RnpA proteins and establish a foundation for understanding biological function and inhibitor targeting.
Correction for ‘Hybrid AI/physics pipeline for miniprotein binder prioritization: application to the BRD3 ET domain’ by Jokent Gaza et al. , Chem. Commun. , 2025, 61 , 19028–19031, https://doi.org/10.1039/D5CC05032D.
The bromodomain and extraterminal domain (BET) family uses its conserved ET domains to recognize diverse peptide motifs, yet exhibits paralog-specific binding preferences whose structural origins remain poorly understood. Using extensive molecular dynamics (MD) simulations of BRD3-ET and BRD4-ET and experimental data for the unbound and peptide-bound states, we show that paralog selectivity arises not from large structural rearrangements but from subtle differences in the dynamics of the α2-α3 loop. Two divergent residues at positions 35 and 36 in this loop modulate the formation of flanking helices (η1 and η2), which in turn control the opening of the peptide-binding cavity and determine how each paralog accommodates distinct binding modes. These sequence-encoded dynamical differences shape the number, stability, and geometry of accessible binding modes and provide a structural rationale for paralog-specific targeting of BET proteins.
Macrocyclic peptides present a challenging regime for molecular modeling in which minimal chemical modifications can strongly reweight conformational ensembles while leaving binding function largely intact. The combination of topological constraint, residual flexibility, and noncanonical linkages complicates both conformational sampling and force-field parametrization, particularly when differences in binding arise from ensemble redistribution rather than distinct bound geometries.
We previously showed that AlphaFold2 can be used to screen for peptide-binding epitopes targeting the extraterminal (ET) domain of Bromodomain and Extraterminal (BET) proteins from candidate protein partners identified in pull-down experiments. However, such approaches require large numbers of AlphaFold2 calculations, making exhaustive screening impractical for larger datasets, such as viral proteomes that may target the ET domain. In many cases, identifying a substantial fraction of binders-even without exhaustive coverage-would already provide valuable biological insight into these interaction networks. Here, we show that an active learning strategy based on Thompson sampling (TS) can efficiently explore peptide sequence space. Using a library derived from BRD3 pull-down experiments, TS recovers 50% of all binders using 15% of the queries required by exhaustive sampling (3.3 times improvement over random sampling). Moreover, TS consistently identifies experimentally known binding epitopes with substantially fewer queries. Because the approach relies only on binary labels, it is readily transferable to other protein-peptide systems where AF-based binding classification is applicable, as well as to peptide-property predictors for properties such as solubility or aggregation propensity.
DNA exhibits local conformational preferences that affect its ability to adopt biologically relevant conformations, such as those required for binding proteins. Traditional methods, like Markov state models and molecular dynamics (MD) simulations, have advanced our understanding but often struggle to capture these rare conformational states due to high computational demands. Here, we introduce a novel AI framework based on dynamical graphical models (DGMs), a generative machine learning approach trained on equilibrium MD data, to predict DNA conformational transitions that are never seen in the MD ensembles. By leveraging local DNA interactions, DGMs generate a comprehensive transition matrix that captures both thermodynamic and kinetic properties of unsampled states, enabling accurate predictions of rare global conformations without the need for extensive sampling. Applying this model to the B→A transition, we demonstrate that DGMs can efficiently predict sequence-dependent A-DNA preferences, achieving results that align closely with replica exchange umbrella sampling simulations. DGMs provide new insights into DNA sequence-structure relationships, paving the way for applications in DNA sequence design and optimization.
AlphaFold2 (AF2) revolutionized protein structure prediction, yet it is often conflated with the protein folding problem. Structure prediction seeks a static conformation, whereas folding concerns the dynamic process of structure formation. We challenge the current status quo, showing that AF2 has implicitly learned some folding principles. Its learned biophysical energy function, though imperfect, enables rapid discovery of folding pathways within minutes. Operating AF2 without multiple sequence alignments (MSAs) or templates forces sampling across its entire energy landscape, akin to ab initio modeling. Among over 7000 proteins, a fraction folds from sequence alone, highlighting the smoothness of AF2's learned surface. Iterating and recycling predictions uncover intermediate structures consistent with experiments, suggesting a "local-first, global-later" mechanism. For designed proteins with optimized local interactions, AF2's landscape becomes too smooth to reveal intermediates. These findings illuminate what AF2 has learned and open avenues for probing protein folding mechanisms and experimental intermediates.
We review MELD, an accelerator of Molecular Dynamics simulations of biomolecules. MELD (Modeling Employing Limited Data) integrates molecular dynamics (MD) with a variety of types of structural information through Bayesian inference, generating ensembles of protein and DNA structures having proper Boltzmann populations. MELD minimizes the computational sampling of irrelevant regions of phase space by applying energetic penalties to areas that conflict with the available data. MELD is effective in refining protein structures using NMR or cryo-EM data or predicting protein-ligand binding poses. As a plugin for OpenMM, MELD is interoperable with other enhanced sampling methods, offering a versatile tool for structural determination in computational chemistry and biophysics.
AI-based protein design can rapidly generate thousands of candidate binders, but most fail to fold or bind productively, creating a critical need for robust prioritization. We present a generalizable hybrid pipeline that integrates deep-learning design and physics-based simulations to filter large libraries down to a handful of high-confidence candidates.
In the Big Data era, a change of paradigm in the use of molecular dynamics is required. Trajectories should be stored under FAIR (findable, accessible, interoperable and reusable) requirements to favor its reuse by the community under an open science paradigm.