Anton 3 is the newest member in a family of supercomputers specially designed for atomic-level simulation of molecules relevant to biology (e.g., DNA, proteins, and drug molecules). Anton 3 achieves order-of-magnitude improvements in time-to-solution over its predecessor, Anton 2 (the current state of the art), and is over 100-fold faster than any other currently available supercomputer, thereby enabling broad new avenues of research on critical questions in biology and drug discovery. This speedup means that a 512-node Anton 3 simulates a million atoms at over 100 microseconds per day. Furthermore, Anton 3 attains this performance while consuming an order of magnitude less energy per simulated microsecond than any other machine. Like its predecessors, Anton 3 was designed from the ground up around a new custom chip to best exploit the capabilities offered by new technologies. We present here the main architectural and algorithmic developments that were necessary to achieve such significant advances.
The evaluation of electrostatic energy for a set of point charges in a periodic lattice is a computationally expensive part of molecular dynamics simulations (and other applications) because of the long-range nature of the Coulomb interaction. A standard approach is to decompose the Coulomb potential into a near part, typically evaluated by direct summation up to a cutoff radius, and a far part, typically evaluated in Fourier space. In practice, all decomposition approaches involve approximations-such as cutting off the near-part direct sum-but it may be possible to find new decompositions with improved trade-offs between accuracy and performance. Here, we present the u-series, a new decomposition of the Coulomb potential that is more accurate than the standard (Ewald) decomposition for a given amount of computational effort and achieves the same accuracy as the Ewald decomposition with approximately half the computational effort. These improvements, which we demonstrate numerically using a lipid membrane system, arise because the u-series is smooth on the entire real axis and exact up to the cutoff radius. Additional performance improvements over the Ewald decomposition may be possible in certain situations because the far part of the u-series is a sum of Gaussians and can thus be evaluated using algorithms that require a separable convolution kernel; we describe one such algorithm that reduces communication latency at the expense of communication bandwidth and computation, a trade-off that may be advantageous on modern massively parallel supercomputers.
Molecular dynamics (MD) simulation has become a powerful tool for characterizing at an atomic level of detail the conformational changes undergone by proteins. The application of such simulations to RNA structures, however, has proven more challenging, due in large part to the fact that the physical models ("force fields") available for MD simulations of RNA molecules are substantially less accurate in many respects than those currently available for proteins. Here, we introduce an extensive revision of a widely used RNA force field in which the parameters have been modified, based on quantum mechanical calculations and existing experimental information, to more accurately reflect the fundamental forces that stabilize RNA structures. We evaluate these revised parameters through long-timescale MD simulations of a set of RNA molecules that covers a wide range of structural complexity, including singlestranded RNAs, RNA duplexes, RNA hairpins, and riboswitches. The structural and thermodynamic properties measured in these simulations exhibited dramatically improved agreementwith experimentally determined values. Based on the comparisons we performed, this RNA force field appears to achieve a level of accuracy comparable to that of state-of-the-art protein force fields, thus significantly advancing the utility of MD simulation as a tool for elucidating the structural dynamics and function of RNA molecules and RNA-containing biological assemblies.
We describe a framework for designing the sequences of multiple nucleic acid strands intended to hybridize in solution via a prescribed reaction pathway. Sequence design is formulated as a multistate optimization problem using a set of target test tubes to represent reactant, intermediate, and product states of the system, as well as to model crosstalk between components. Each target test tube contains a set of desired "on -target" complexes, each with a target secondary structure and target concentration, and a set of undesired "off -target" complexes, each with vanishing target concentration. Optimization of the equilibrium ensemble properties of the target test tubes implements both a positive design paradigm, explicitly designing for on pathway elementary steps, and a negative design paradigm, explicitly designing against off -pathway crosstalk. Sequence design is performed subject to diverse user -specified sequence constraints including composition constraints, complementarity constraints, pattern prevention constraints, and biological constraints. Constrained multistate sequence design facilitates nucleic acid reaction pathway engineering for diverse applications in molecular programming and synthetic biology. Design jobs can be run online via the NUPACK web application.
NEUROSCIENCE Correction for “Feature attention evokes task-specific pattern selectivity in V4 neurons,” by Anna E. Ipata, Angela L. Gee, and Michael E. Goldberg, which appeared in issue 42, October 16, 2012, of Proc Natl Acad Sci USA (109:16778–16785; first published October 5, 2012; 10.1073/pnas.1215402109). The authors note that Fig. 3 and its corresponding legend appeared incorrectly. The corrected figure and its corrected legend appear below. This error does not affect the conclusions of the article.
The use of molecular dynamics simulations to provide atomic-level descriptions of biological processes tends to be computationally demanding, and a number of approximations are thus commonly employed to improve computational efficiency. In the past, the effect of these approximations on macromolecular structure and stability has been evaluated mostly through quantitative studies of small-molecule systems or qualitative observations of short-timescale simulations of biological macromolecules. Here we present a quantitative evaluation of two commonly employed approximations, using a test system that has been the subject of a number of previous protein folding studies--the villin headpiece. In particular, we examined the effect of (i) the use of a cutoff-based force-shifting technique rather than an Ewald summation for the treatment of electrostatic interactions, and (ii) the length of the cutoff used to determine how many pairwise interactions are included in the calculation of both electrostatic and van der Waals forces. Our results show that the free energy of folding is relatively insensitive to the choice of cutoff beyond 9 Å, and to whether an Ewald method is used to account for long-range electrostatic interactions. In contrast, we find that the structural properties of the unfolded state depend more strongly on the two approximations examined here.
Molecular dynamics simulations capture the behavior of biological macromolecules in full atomic detail, but their computational demands, combined with the challenge of appropriately modeling the relevant physics, have historically restricted their length and accuracy. Dramatic recent improvements in achievable simulation speed and the underlying physical models have enabled atomic-level simulations on timescales as long as milliseconds that capture key biochemical processes such as protein folding, drug binding, membrane transport, and the conformational changes critical to protein function. Such simulation may serve as a computational microscope, revealing biomolecular mechanisms at spatial and temporal scales that are difficult to observe experimentally. We describe the rapidly evolving state of the art for atomic-level biomolecular simulation, illustrate the types of biological discoveries that can now be made through simulation, and discuss challenges motivating continued innovation in this field.
We previously introduced hybridization chain reactions (HCR), in which metastable DNA hairpins undergo conditional self-assembly to form long nicked double-stranded ‘polymers’ in the presence of a DNA initiator molecule (RM Dirks and NA Pierce, PNAS 2004, 10, 15275). HCR systems have been engineered to function as orthogonal in situ amplifiers for multiplexed bioimaging (HMT Choi et al., Nat Biotech, in press) and as programmable mechanical transducers for selectively killing cultured human cancer cells (S Venkataraman et al., PNAS 2010, 107, 16777). Here, we model the equilibrium and kinetic properties of HCR, revealing sources of non-ideal behavior and methods for controlling system performance. Our results demonstrate that HCR is accurately modeled as a living alternating copolymerization.
Molecular simulations aim to sample all of the thermodynamically important states; when the sampling is inadequate, inaccuracy follows. A widely used technique to enhance sampling in simulations is Hamiltonian exchange. This technique introduces auxiliary Hamiltonians under which sampling is computationally efficient and attempts to exchange the molecular states among the auxiliary and the original Hamiltonians. The effectiveness of Hamiltonian exchange depends in part on the probability that the trial exchanges can be accepted, which involves good choices of auxiliary Hamiltonians and a good method of generating the trial exchanges. In this paper, we investigate nonequilibrium simulations as trial exchange generators and develop a theoretical model for the efficiency of Hamiltonian exchange and an algorithm to better configure such simulations. We show that properly configured nonequilibrium simulations can modestly increase the overall efficiency of Hamiltonian exchange.
Cancer cells are characterized by genetic mutations that deregulate cell proliferation and suppress cell death. To arrest the uncontrolled replication of malignant cells, conventional chemotherapies systemically disrupt cell division, causing diverse and often severe side effects as a result of collateral damage to normal cells. Seeking to address this shortcoming, we pursue therapeutic regulation that is conditional, activating selectively in cancer cells. This functionality is achieved using small conditional RNAs that interact and change conformation to mechanically transduce between detection of a cancer mutation and activation of a therapeutic pathway. Here, we describe small conditional RNAs that undergo hybridization chain reactions (HCR) to induce cell death via an innate immune response if and only if a cognate mRNA cancer marker is detected within a cell. The sequences of the small conditional RNAs can be designed to accept different mRNA markers as inputs to HCR transduction, providing a programmable framework for selective killing of diverse cancer cells. In cultured human cancer cells (glioblastoma, prostate carcinoma, Ewing's sarcoma), HCR transduction mediates cell death with striking efficacy and selectivity, yielding a 20- to 100-fold reduction in population for cells containing a cognate marker, and no measurable reduction otherwise. Our results indicate that programmable mechanical transduction with small conditional RNAs represents a fundamental principle for exploring therapeutic conditional regulation in living cells.
The Nucleic Acid Package (NUPACK) is a growing software suite for the analysis and design of nucleic acid systems. The NUPACK web server (http://www.nupack.org) currently enables: Analysis: thermodynamic analysis of dilute solutions of interacting nucleic acid strands. Design: sequence design for complexes of nucleic acid strands intended to adopt a target secondary structure at equilibrium. Utilities: evaluation, display, and annotation of equilibrium properties of a complex of nucleic acid strands. NUPACK algorithms are formulated in terms of nucleic acid secondary structure. In most cases, pseudoknots are excluded from the structural ensemble. © 2010 Wiley Periodicals, Inc. J Comput Chem, 2010
We present a synthetic molecular motor capable of autonomous nanoscale transport in solution. Inspired by bacterial pathogens such as Rickettsia rickettsii, which locomote by inducing the polymerization of the protein actin at their surfaces to form 'comet tails' 1, the motor operates by polymerizing a double-helical DNA tail(2). DNA strands are propelled processively at the living end of the growing polymers, demonstrating autonomous locomotion powered by the free energy of DNA hybridization.
Motivated by the analysis of natural and engineered DNA and RNA systems, we present the first algorithm for calculating the partition function of an unpseudoknotted complex of multiple interacting nucleic acid strands. This dynamic program is based on a rigorous extension of secondary structure models to the multistranded case, addressing representation and distinguishability issues that do not arise for single-stranded structures. We then derive the form of the partition function for a fixed volume containing a dilute solution of nucleic acid complexes. This expression can be evaluated explicitly for small numbers of strands, allowing the calculation of the equilibrium population distribution for each species of complex. Alternatively, for large systems (e.g., a test tube), we show that the unique complex concentrations corresponding to thermodynamic equilibrium can be obtained by solving a convex programming problem. Partition function and concentration information can then be used to calculate equilibrium base-pairing observables. The underlying physics and mathematical formulation of these problems lead to an interesting blend of approaches, including ideas from graph theory, group theory, dynamic programming, combinatorics, convex optimization, and Lagrange duality.
L'invention porte sur l'utilisation d'une reaction en chaine d'hybridation (HCR) pour former des polymeres d'ARN double brin en presence d'une cible telle qu'un acide nucleique associe a une maladie ou a un trouble. Les polymeres d'ARN sont de preference capables d'activer la kinase PKR dependante de l'ARN. Cette activation via ARN-HCR peut servir a traiter une grande variete de maladies et de troubles, par ciblage specifique de cellules malades.