
CRISPR-based transcriptional regulation technologies, including CRISPR activation (CRISPRa) and CRISPR interference (CRISPRi), offer precise and programmable control over gene expression, representing a major advance in gene and epigenetic therapy. CRISPRa uses nuclease-inactive Cas proteins fused to transcriptional activators to upregulate target genes, while CRISPRi employs repressor domains for gene silencing. Preclinical studies have demonstrated the efficacy of CRISPRa/i in models of metabolic, neurological, muscular, and oncological diseases. Notably, CRISPRi-based therapies have entered clinical trials for conditions like hepatitis B and muscular dystrophy, showing encouraging safety and efficacy profiles. Despite ongoing challenges related to delivery efficiency, immunogenicity, and off-target activity, innovations in protein engineering and guide RNA design are rapidly enhancing the precision and safety of these technologies. Overall, CRISPRa and CRISPRi are poised to transform the treatment of genetic and epigenetic disorders, with continued optimization expected to accelerate their clinical adoption and broaden their therapeutic impact.
In this paper we present a theoretical design for tile-based self-assembling systems where individual tile sets that are universal for various tasks (e.g. building shapes or encoding arbitrary data for algorithmic systems) can be provided their input solely using sequences of variations in temperatures. Although there is prior theoretical work in so-called "temperature programming," the results presented here are based upon recent experimental work with "blocked tiles" that provides plausible evidence for their practical implementation. We develop and present an abstract mathematical model, the BlockTAM, for such systems that is based upon the experimental work, and provide constructions within it for universal number encoding systems and a universal shape building construction. These results mirror previous results in temperature programming, covering the two ends of the spectrum that seek to balance the scale factors of assemblies with the number of temperature phases required, while leveraging the features of the BlockTAM that we hope provide a pathway for future experimental implementations.
For general discrete Chemical Reaction Networks (CRNs), the fundamental problem of reachability - the question of whether a target configuration can be produced from a given initial configuration - was recently shown to be Ackermann-complete. However, many open questions remain about which features of the CRN model drive this complexity. We study a restricted class of CRNs with void rules, reactions that only decrease species counts. We further examine this regime in the motivated model of step CRNs, which allow additional species to be introduced in discrete stages. With and without steps, we characterize the complexity of the reachability problem for CRNs with void rules. We show that, without steps, reachability remains polynomial-time solvable for bimolecular systems but becomes NP-complete for larger reactions. Conversely, with just a single step, reachability becomes NP-complete even for bimolecular systems. Our results provide a nearly complete classification of void-rule reachability problems into tractable and intractable cases, with only a single exception.
DNA strand displacement, a collective name for certain behaviors of short strands of DNA, has been used to build many interesting molecular devices over the past few decades. Among those devices are general implementation schemes for Chemical Reaction Networks, suggesting a place in an abstraction hierarchy for complex molecular programming. However, the possibilities of DNA strand displacement are far from fully explored. On a theoretical level, most DNA strand displacement systems are built out of a few simple motifs, with the space of possible motifs otherwise unexplored. On a practical level, the desire for general, large-scale DNA strand displacement systems is not fulfilled. Those systems that are scalable are not general, and those that are general don't scale up well. We have recently been exploring the space of possibilities for DNA strand displacement systems where all input complexes are made out of at most two strands of DNA. As a test case, we've had an open question of whether such systems can implement general Chemical Reaction Networks, in a way that has a certain set of other desirable properties - reversible, systematic, O(1) toeholds, bimolecular reactions, and correct according to CRN bisimulation - that the state-of-the-art implementations with more than 2-stranded inputs have. Until now we've had a few results that have all but one of those desirable properties, including one based on a novel mechanism we called coupled reconfiguration, but that depended on the physically questionable assumption that pseudoknots cannot occur. We wondered whether the same type of mechanism could be done in a pseudoknot-robust way. In this work we show that in fact, coupled reconfiguration can be done in a pseudoknot-robust way, and this mechanism can implement general Chemical Reaction Networks with all inputs being single strands of DNA. Going further, the same motifs used in this mechanism can implement stacks and surface-based bimolecular reactions. Those have been previously studied as part of polymer extensions of the Chemical Reaction Network model, and on an abstract model level, the resulting extensions are Turing-complete in ways the base Chemical Reaction Network model is not. Our mechanisms are significantly different from previously tested DNA strand displacement systems, which raises questions about their ability to be implemented experimentally, but we have some reasons to believe the challenges are solvable. So we present the pseudoknot-robust coupled reconfiguration mechanism and its use for general Chemical Reaction Network implementations; we present the extensions of the mechanism to stack and surface reactions; and we discuss the possible obstacles and solutions to experimental implementation, as well as the theoretical implications of this mechanism.
A grand challenge facing molecular programmers is the rational development of fast, robust, and isothermal architectures akin to "chemical central processing units" that can sense (bio-)chemical signals from their environment, perform complex computation, and orchestrate a physical response in situ. DNA strand displacement systems (DSDs) remain a compelling candidate, but are hampered by spurious reaction pathways that lead to incorrect output. DSDs that utilize the systematic leakless motif can be made arbitrarily robust at the cost of increasing redundancy and network size (scaling), and thus a degradation of kinetic performance. Another class of architectures utilize DNA hybridization, extension, and signal production of entirely sequestered outputs via strand-displacing polymerases (SDPs) that have resulted in impressive demonstrations; however, they face similar challenges of aberrant behavior such as mis-priming by incorrect signals. Our work introduces a unified polymerase-dependent toehold-mediated strand displacement (PD-TMSD) architecture that integrates the programmed specificity of DSDs with the unique advantages of SDPs. This unification enables systems that can be made arbitrarily robust, at any concentration range, without increasing network size. We propose a number of gate designs and composition rules to compute arbitrary Boolean functions, emulate arbitrary chemical reaction networks, and explore time-bounded probabilistic computation made possible by certain classes of SDPs. Our theoretical exploration is backed by preliminary experimental demonstrations. This contribution was inspired by the belief that molecular programming can meet or exceed the complexity exhibited in biology if we embrace its best understood molecular machinery and couple it with systematic design principles built upon a strong theoretical foundation.
A fundamental problem in crystallisation, and in molecular tile-based self-assembly in particular, is how to simultaneously control its two main constituent processes: seeded growth and spontaneous nucleation. Often, we desire out-of-equilibrium growth without spontaneous nucleation, which can be achieved through careful calibration of temperature, concentration and experimental time-scale a laborious and overly-sensitive approach. Another technique is to find alternative nucleation-resistant tile designs [Minev et al, 2001]. Rogers, Evans and Woods [In prep] propose blockers: short DNA strands designed to dynamically block DNA tile sides, altering self-assembly dynamics. Experiments showed independent and tunable control on nucleation and growth rates. Here, we provide a theoretical explanation for these surprising results. We formally define the kBlock model where blockers bind to tiles at thermodynamic equilibrium in solution and stochastic kinetics allow self-assembly of a tiled structure. In an intentionally simplified mathematical setting we show that blockers permit reasonable seeded growth rates, akin to a non-blocked tile system at lower tile concentration, crucially giving nucleation rates that are exponentially suppressed. We then implement the kBlock model in a stochastic simulator, with results showing remarkable alignment with oversimplified theory. We provide evidence of blocker-induced tile buffering, where a large reservoir of blocked tiles slowly feeds a small unblocked tile subpopulation which acts like a regular, non-blocked, low tile concentration system, yet is capable of long-term buffered assembly. Finally, and perhaps most satisfyingly, theory and simulations align remarkably well with DNA self-assembly experiments over a wide range of concentrations and temperatures, matching the size of growth temperature windows to within 12%. Blockers are a straightforward solution to the challenging problem of simultaneously and independently controlling growth and nucleation, using a motif compatible with many DNA tile systems.
Many molecular systems are best understood in terms of prototypical species and reactions. The central dogma and related biochemistry are rife with examples: gene i is transcribed into RNA i, which is translated into protein i; kinase n phosphorylates substrate m; protein p dimerizes with protein q. Engineered nucleic acid systems also often have this form: oligonucleotide i hybridizes to complementary oligonucleotide j; signal strand n displaces the output of seesaw gate m; hairpin p triggers the opening of target q. When there are many variants of a small number of prototypes, it can be conceptually cleaner and computationally more efficient to represent the full system in terms of indexed species (e.g. for dimerization, M-p, D-pq) and indexed reactions (M-p + M-q -> D-pq). Here, we formalize the Indexed Chemical Reaction Network (ICRN) model and describe a Python software package designed to simulate such systems in the well-mixed and reaction-diffusion settings, using a differentiable programming framework originally developed for large-scale neural network models, taking advantage of GPU acceleration when available. Notably, this framework makes it straightforward to train the models' initial conditions and rate constants to optimize a target behavior, such as matching experimental data, performing a computation, or exhibiting spatial pattern formation. The natural map of indexed chemical reaction networks onto neural network formalisms provides a tangible yet general perspective for translating concepts and techniques from the theory and practice of neural computation into the design of biomolecular systems.
In this paper we study the relationship between mathematical models of tile-based self-assembly which differ in terms of the synchronicity of tile additions. In the standard abstract Tile Assembly Model (aTAM), each step of assembly consists of a single tile being added to an assembly. At any given time, each location on the perimeter of an assembly to which a tile can legally bind is called a frontier location, and for each step of assembly one frontier location is randomly selected and a tile is added. In the Synchronous Tile Assembly Model (syncTAM), at each step of assembly every frontier location simultaneously receives a tile. Our results show that while directed, non-cooperative syncTAM systems are capable of universal computation (while directed, non-cooperative aTAM systems are known not to be), and they are capable of building shapes that can't be built within the aTAM, the non-cooperative aTAM is also capable of building shapes that can't be built within the syncTAM even cooperatively. We show a variety of results that demonstrate the similarities and differences between these two models.
We address the task of secondary structure design for de novo 3D RNA origami wireframe structures in a way that takes into account the specifics of a cotranscriptional folding setting. We consider two issues: firstly, avoiding the topological obstacle of "polymerase trapping", where some helical domain cannot be hybridised due to a closed kissing-loop pair blocking the winding of the strand relative to the polymerase-DNA-template complex; and secondly, minimising the number of distinct kissing-loop designs needed, by reusing KL pairs that have already been hybridised in the folding process. For the first task, we present an efficient strand-routing method that guarantees the absence of polymerase traps for any 3D wireframe model, and for the second task, we provide a graph-theoretic formulation of the minimisation problem, show that it is NP-complete in the general case, and outline a branch-and-bound type enumerative approach to solving it. Key concepts in both cases are depth-first search in graphs and the ensuing DFS spanning trees. Both algorithms have been implemented in the DNAforge design tool (https://dnaforge.org) and we present some examples of the results.
Computing equilibrium concentrations of molecular complexes is generally analytically intractable and requires numerical approaches. In this work we focus on the polymer-monomer level, where indivisible molecules (monomers) combine to form complexes (polymers). Rather than employing free-energy parameters for each polymer, we focus on the athermic setting where all interactions preserve enthalpy. This setting aligns with the strongly bonded (domain-based) regime in DNA nanotechnology when strands can bind in different ways, but always with maximum overall bonding -- and is consistent with the saturated configurations in the Thermodynamic Binding Networks (TBNs) model. Within this context, we develop an iterative algorithm for assigning polymer concentrations to satisfy detailed-balance, where on-target (desired) polymers are in high concentrations and off-target (undesired) polymers are in low. Even if not directly executed, our algorithm provides effective insights into upper bounds on concentration of off-target polymers, connecting combinatorial arguments about discrete configurations such as those in the TBN model to real-valued concentrations. We conclude with an application of our method to decreasing leak in DNA logic and signal propagation. Our results offer a new framework for design and verification of equilibrium concentrations when configurations are distinguished by entropic forces.
RNA co-transcriptionality, where RNA is spliced or folded during transcription from DNA templates, offers promising potential for molecular programming. It enables programmable folding of nano-scale RNA structures and has recently been shown to be Turing universal. While post-transcriptional splicing is well studied, co-transcriptional splicing is gaining attention for its efficiency, though its unpredictability still remains a challenge. In this paper, we focus on engineering co-transcriptional splicing, not only as a natural phenomenon but as a programmable mechanism for generating specific RNA target sequences from DNA templates. The problem we address is whether we can encode a set of RNA sequences for a given system onto a DNA template word, ensuring that all the sequences are generated through co-transcriptional splicing. Given that finding the optimal encoding has been shown to be NP-complete under the various energy models considered, we propose a practical alternative approach under the logarithmic energy model. More specifically, we provide a construction that encodes an arbitrary nondeterministic finite automaton (NFA) into a circular DNA template from which co-transcriptional splicing produces all sequences accepted by the NFA. As all finite languages can be efficiently encoded as NFA, this framework solves the problem of finding small DNA templates for arbitrary target sets of RNA sequences. The quest to obtain the smallest possible such templates naturally leads us to consider the problem of minimizing NFA and certain practically motivated variants of it, but as we show, those minimization problems are computationally intractable.
To understand and engineer biological and artificial nucleic acid systems, algorithms are employed for prediction of secondary structures at thermodynamic equilibrium. Dynamic programming algorithms are used to compute the most favoured, or Minimum Free Energy (MFE), structure, and the Partition Function (PF), a tool for assigning a probability to any structure. However, in some situations, such as when there are large numbers of strands, or pseudoknoted systems, NP-hardness results show that such algorithms are unlikely, but only for MFE. Curiously, algorithmic hardness results were not shown for PF, leaving two open questions on the complexity of PF for multiple strands and single strands with pseudoknots. The challenge is that while the MFE problem cares only about one, or a few structures, PF is a summation over the entire secondary structure space, giving theorists the vibe that computing PF should not only be as hard as MFE, but should be even harder. We answer both questions. First, we show that computing PF is #P-hard for systems with an unbounded number of strands, answering a question of Condon Hajiaghayi, and Thachuk [DNA27]. Second, for even a single strand, but allowing pseudoknots, we find that PF is #P-hard. Our proof relies on a novel magnification trick that leads to a tightly-woven set of reductions between five key thermodynamic problems: MFE, PF, their decision versions, and #SSEL that counts structures of a given energy. Our reductions show these five problems are fundamentally related for any energy model amenable to magnification. That general classification clarifies the mathematical landscape of nucleic acid energy models and yields several open questions.
Background/Objective: Oligodeoxynucleotides (ODNs) containing base-labile modifications such as N4-acetyldeoxycytidine (4acC), N6-acetyladenosine (6acA), N2-acetylguanosine (2acG), and N4-methyoxycarbonyldeoxycytidine (4mcC) are highly challenging to synthesize because standard ODN synthesis methods require deprotection and cleavage under strongly basic and nucleophilic conditions, and there is a lack of ideal alternative methods to solve the problem. The objective of this work is to explore the capability of the recently developed 1,3-dithian-2-yl-methoxycarbonyl (Dmoc) method for the incorporation of multiple 4acC modifications into a single ODN molecule and the feasibility of using the method for the incorporation of the 6acA, 2acG and 4mcC modifications into ODNs. Methods: The sensitive ODNs were synthesized on an automated solid phase synthesizer using the Dmoc group as the linker and the methyl Dmoc (meDmoc) group for the protection of the exo-amino groups of nucleobases. Deprotection and cleavage were achieved under non-nucleophilic and weakly basic conditions. Results: The 4acC, 6acA, 2acG, and 4mcC were all found to be stable under the mild ODN deprotection and cleavage conditions. Up to four 4acC modifications were able to be incorporated into a single 19-mer ODN molecule. ODNs containing the 6acA, 2acG, and 4mcC modifications were also successfully synthesized. The ODNs were characterized using RP HPLC, capillary electrophoresis, gel electrophoresis and MALDI MS. Conclusions: Among the modified nucleotides, 4acC has been found in nature and proven beneficial to DNA duplex stability. A method for the synthesis of ODNs containing multiple 4acC modifications is expected to find applications in biological studies involving 4acC. Although 6acA, 2acG, and 4mcC have not been found in nature, a synthetic route to ODNs containing them is expected to facilitate projects aimed at studying their biophysical properties as well as their potential for antisense, RNAi, CRISPR, and mRNA therapeutic applications.
Background/Objectives: Culex species mosquitoes are globally distributed and transmit several pathogens that impact animal and public health, including West Nile virus, Usutu virus, and Plasmodium relictum. Despite their relevance, Culex species are less widely studied than Aedes and Anopheles mosquitoes. To expand the genetic tools used to study Culex mosquitoes, we previously developed an optimized plasmid for transient Cas9 and single-guide RNA (sgRNA) expression in Culex quinquefasciatus cells to generate gene knockouts. Here, we established a monoclonal cell line that consistently expresses Cas9 and can be used for screens to determine gene function or antiviral activity. Methods: We used this system to perform the successful gene editing of seven genes and subsequent testing for potential antiviral effects, using a simple single-guide RNA (sgRNA) transfection and subsequent virus infection. Results: We were able to show antiviral effects for the Cx. quinquefasciatus genes dicer-2, argonaute-2b, vago, piwi5, piwi6a, and cullin4a. In comparison to the RNAi-mediated gene silencing of dicer-2, argonaute-2b, and piwi5, our Cas9/sgRNA approach showed an enhanced ability to detect antiviral effects. Conclusions: We propose that this cell line offers a new tool for studying gene function in Cx. quinquefasciatus mosquitoes that avoids the use of RNAi. This short study also serves as a proof-of-concept for future gene knock-ins in these cells. Our cell line expands the molecular resources available for vector competence research and will support the design of future research strategies to reduce the transmission of mosquito-borne diseases.
Living systems are capable on the one hand of eliciting a coordinated response to changing environments (also known as adaptation), and on the other hand, they are capable of reproducing themselves. Notably, adaptation to environmental change requires the monitoring of the surroundings, while reproduction requires monitoring oneself. These two tasks appear separate and make use of different sources of information. Yet, both the process of adaptation as well as that of reproduction are inextricably coupled to alterations in genomic DNA expression, while a cell behaves as an indivisible unity in which apparently independent processes and mechanisms are both integrated and coordinated. We argue that at the most basic level, this integration is enabled by the unique property of the DNA to act as a double coding device harboring two logically distinct types of information. We review biological systems of different complexities and infer that the inter-conversion of these two distinct types of DNA information represents a fundamental self-referential device underlying both systemic integration and coordinated adaptive responses.
We present a general design technique for rendering any 3D wireframe model, that is any connected graph linearly embedded in 3D space, as an RNA origami nanostructure with a minimum number of kissing loops. The design algorithm, which applies some ideas and methods from topological graph theory, produces renderings that contain at most one kissing-loop pair for many interesting model families, including for instance all fully triangulated wireframes and the wireframes of all Platonic solids. The design method is already implemented and available for use in the design tool DNAforge (https://dnaforge.org).
Localized molecular devices are a powerful tool for engineering complex information-processing circuits and molecular robots. Their practical advantages include speed and scalability of interactions between components tethered near to each other on an underlying nanostructure, and the ability to restrict interactions between more distant components. The latter is a critical feature that must be factored into computational tools for the design and simulation of localized molecular devices: unlike in solution-phase systems, the geometries of molecular interactions must be accounted for when attempting to determine the network of possible reactions in a tethered molecular system. This work aims to address that challenge by integrating, for the first time, automated approaches to analysis of molecular geometry with reaction enumeration algorithms for DNA strand displacement reaction networks that can be applied to tethered molecular systems. By adapting a simple approach to solving the biophysical constraints inherent in molecular interactions to be applicable to tethered systems, we produce a localized reaction enumeration system that enhances previous approaches to reaction enumeration in tethered system by not requiring users to explicitly specify the subsets of components that are capable of interacting. This greatly simplifies the user's task and could also be used as the basis of future systems for automated placement or routing of signal-transmission and logical processing in molecular devices. We apply this system to several published example systems from the literature, including both tethered molecular logic systems and molecular robots.
Life is chemical intelligence. What is the source of intelligent behavior in molecular systems? Here we illustrate how, in contrast to the common belief that energy use in non-equilibrium reactions is essential, the detailed balance equilibrium properties of multicomponent liquid interactions are sufficient for sophisticated information processing. Our approach derives from the classical Boltzmann machine model for probabilistic neural networks, inheriting key principles such as representing probability distributions via quadratic energy functions, clamping input variables to infer conditional probability distributions, accommodating omnidirectional computation, and learning energy parameters via a wake phase / sleep phase algorithm that performs gradient descent on the relative entropy with respect to the target distribution. While the cubic lattice model of multicomponent liquids is standard, the behaviors exhibited by the trained molecules capture both previously-observed phenomena such as core-shell condensate architectures as well as novel phenomena such as an analog of Hopfield associative memories that perform recall by contact with a patterned surface. Our final example demonstrates equilibrium classification of MNIST digits. Experimental implementation using DNA nanostar liquids is conceptually straightforward.
Molecular programmers and nanostructure engineers use domain-level design to abstract away messy DNA/RNA sequence, chemical and geometric details. Such domain-level abstractions are enforced by sequence design principles and provide a key principle that allows scaling up of complex multistranded DNA/RNA programs and structures. Determining the most favoured secondary structure, or Minimum Free Energy (MFE), of a set of strands, is typically studied at the sequence level but has seen limited domain-level work. We analyse the computational complexity of MFE for multistranded systems in a simple setting were we allow only 1 or 2 domains per strand. On the one hand, with 2-domain strands, we find that the MFE decision problem is NP-complete, even without pseudoknots, and requires exponential time algorithms assuming SAT does. On the other hand, in the simplest case of 1-domain strands there are efficient MFE algorithms for various binding modes. However, even in this single-domain case, MFE is P-hard for promiscuous binding, where one domain may bind to multiple as experimentally used by Nikitin [Nat Chem., 2023], which in turn implies that strands consisting of a single domain efficiently implement arbitrary Boolean circuits.
Oxidative stress-mediated biomolecular damage is a characteristic feature of ionizing radiation (IR) injury, leading to genomic instability and chronic health implications. Specifically, a dose- and linear energy transfer (LET)-dependent persistent increase in oxidative DNA damage has been reported in many tissues and biofluids months after IR exposure. Contrary to low-LET photon radiation, high-LET IR exposure is known to cause significantly higher accumulations of DNA damage, even at sublethal doses, compared to low-LET IR. High-LET IR is prevalent in the deep space environment (i.e., beyond Earth's magnetosphere), and its exposure could potentially impair astronauts' health. Therefore, the development of biomarkers to assess and monitor the levels of oxidative DNA damage can aid in the early detection of health risks and would also allow timely intervention. Among the recognized biomarkers of oxidative DNA damage, 8-oxo-7,8-dihydro-2'-deoxyguanosine (8-OxodG) has emerged as a promising candidate, indicative of chronic oxidative stress. It has been reported to exhibit differing levels following equivalent doses of low- and high-LET IR. This review discusses 8-OxodG as a potential biomarker of high-LET radiation-induced chronic stress, with special emphasis on its potential sources, formation, repair mechanisms, and detection methods. Furthermore, this review addresses the pathobiological implications of high-LET IR exposure and its association with 8-OxodG. Understanding the association between high-LET IR exposure-induced chronic oxidative stress, systemic levels of 8-OxodG, and their potential health risks can provide a framework for developing a comprehensive health monitoring biomarker system to safeguard the well-being of astronauts during space missions and optimize long-term health outcomes.