R6K plasmids are commonly used for a wide range of genome engineering applications due to their ability to support transient delivery of genetic cargos in many hosts. The maintenance of R6K plasmids requires specific strains. Unfortunately, many of these have obscure backgrounds, limited availability and were not built for efficient cloning. To address this issue, we present the construction and characterization of a series of Pir E. coli strains called SHARK that are built from the DH10B derivative, Marionette-Clo. All SHARK strains have a genome encoded pir gene for stable R6K plasmid maintenance and a λCI gene for tight unconditional repression of specific genes on plasmids. We show that SHARK strains are >100-fold more efficient than a commercial Pir strain when transformed with large and complex cloning reactions. SHARK is intended to help facilitate the cloning of R6K plasmids for challenging genome engineering projects, with all strains and genetic tools for their assembly being made publicly available.
Many microorganisms alter their movement in response to light. These responses can drive collective behaviours like photoaccumulation and photodispersion, which play a key role in broader biological functions like photosynthesis. Our understanding of these emergent phenomena is severely limited by difficulties in obtaining the data needed to establish accurate models that can serve as a basis for multi-scale analyses. Here, we address this issue by developing an integrated experimental and computational platform to collect large temporal imaging datasets that allow for the inference of 'digital twins'-mathematically precise computational models that accurately mirror the behaviour of individual microorganisms-and show that they can replicate the light response of diverse microorganisms in silico. We demonstrate that a generalized phenomenological model capable of simultaneously capturing dynamic speed variations and multiple light responses can be effectively parametrized from experimental data to capture key behavioural traits of two commonly studied photo-responsive microorganisms (Euglena gracilis and Volvox aureus). We also show our model's ability to accurately reproduce patterns of movement for individuals and populations in response to dynamic and spatially varying light patterns. This work takes steps towards the automated phenotyping of multi-scale behaviours in biology and unlocks new opportunities for the design of spatial control algorithms to guide collective microorganism behaviour.
Abstract Science is losing knowledge it cannot afford to lose. Negative results go unpublished, hard-won expertise walks out the door with departing researchers, and preservation efforts remain fragmented. The consequences are wasted resources, duplicated effort, and missed discoveries. In this perspective, we argue that the research community can act now by embracing alternative dissemination channels, improving documentation best practices, and building sustainable digital infrastructure. We envision moderated platforms for sharing null results and practical know-how, community-driven standards, and AI-powered tools that lower barriers to implementation. With coordinated effort, science can become more open, efficient, and resilient for future generations.
DNA polymerases are complex molecular machines capable of replicating genetic material using a template-driven process. While the copying function of these enzymes is well established, their ability to perform untemplated DNA synthesis is less well characterized. Here, we explore the ability of DNA polymerases to synthesize DNA fragments in the absence of a template. We use long-read nanopore sequencing, real-time fluorescence assays, and atomic force microscopy to observe the synthesis and physical structure of pools of DNA products derived from a diverse set of natural and engineered DNA polymerases across varying temperatures and buffer compositions. We detail the features of the DNA fragments generated, enrichment of select sequence motifs, and demonstrate that the sequence composition of the synthesized DNA can be altered by modifying environmental conditions. This work provides extensive data to better discern the process of untemplated DNA polymerase activity and may support its potential repurposing as a technology for the guided synthesis of DNA sequences on the kilobase-scale and beyond.
Guidelines for managing scientific data have been established under the FAIR principles requiring that data be Findable, Accessible, Interoperable, and Reusable. In many scientific disciplines, especially computational biology, both data and models are key to progress. For this reason, and recognizing that such models are a very special type of 'data', we argue that computational models, especially mechanistic models prevalent in medicine, physiology and systems biology, deserve a complementary set of guidelines. We propose the CURE principles, emphasizing that models should be Credible, Understandable, Reproducible, and Extensible. We delve into each principle, discussing verification, validation, and uncertainty quantification for model credibility; the clarity of model descriptions and annotations for understandability; adherence to standards and open science practices for reproducibility; and the use of open standards and modular code for extensibility and reuse. We outline recommended and baseline requirements for each aspect of CURE, aiming to enhance the impact and trustworthiness of computational models, particularly in biomedical applications where credibility is paramount. Our perspective underscores the need for a more disciplined approach to modeling, aligning with emerging trends such as Digital Twins and emphasizing the importance of data and modeling standards for interoperability and reuse. Finally, we emphasize that given the non-trivial effort required to implement the guidelines, the community moves to automate as many of the guidelines as possible.
The insertion of large genetic circuits and metabolic pathways into bacterial genomes is becoming increasingly common within the field of synthetic biology due to the improved robustness and stability that come with genome integration. CRISPR-associated transposases (CASTs) enable RNA-guided DNA insertion without introducing double-stranded breaks and have been shown to function across diverse bacterial species. Here, we present an improved tool called pSPIN-GG and supporting protocols for simplified CAST-based genome engineering. The pSPIN-GG system includes Golden Gate-compatible promoter, guide, and cargo modules for simple assembly, a green fluorescent protein dropout cassette for rapid verification of guide replacement, and a set of tested sites within the Escherichia coli BL21 chromosome to enable gene dosing of genetic cargoes. These refinements support accelerated library construction, reduce assembly and screening burden, and expand the accessibility of CAST systems for multiplexed bacterial genome engineering.
Plasmids are central to modern biotechnology, especially therapeutic development, yet their propagation in Escherichia coli remains difficult to predict. Although expression-induced burden is well understood and can be mitigated, the impact of foreign DNA segments that do not function in bacteria on plasmid propagation and stability remains largely unknown. Here we developed a pooled, sequencing-based framework that performs quantitative profiling across millions of bases, enabling high-resolution assessment of plasmid fitness at scale and revealing cryptic, sequence-encoded interactions between foreign DNA elements and bacterial hosts. Promoter-like motifs, transcription factor binding site homology, and recombination-prone architectures emerge as major determinants of propagation efficiency, with context-dependent effects demonstrating that plasmid behaviour arises from higher-order interactions between parts rather than isolated elements. Extending this framework, we introduce TRACE, a neural-network model trained on degenerate sequence libraries that predicts plasmid propagation directly from sequence. TRACE generalises across plasmid architectures and can be fine-tuned on experimental datasets to improve predictions of manufacturability and host compatibility. These advances establish a generalisable, data-driven framework for understanding and designing host-aware plasmids, transforming plasmid production from an empirical process into a predictable property of DNA sequence.
Standards play a crucial role in ensuring consistency, interoperability, and efficiency of communication across various disciplines. In the field of synthetic biology, the Synthetic Biology Open Language (SBOL) Visual standard was introduced in 2013 to establish a structured framework for visually representing genetic designs. Over the past decade, SBOL Visual has evolved from a simple set of 21 glyphs into a comprehensive diagrammatic language for biological designs. This perspective reflects on the first ten years of SBOL Visual, tracing its evolution from inception to version 3.0. We examine the standard's adoption over time, highlighting its growing use in scientific publications, the development of supporting visualization tools, and ongoing efforts to enhance clarity and accessibility in communicating genetic design information. While trends in adoption show steady increases, achieving full compliance and use of best practices will require additional efforts. Looking ahead, the continued refinement of SBOL Visual and broader community engagement will be essential to ensuring its long-term value as the field of synthetic biology develops.
The ability to precisely insert DNA payloads into a genome enables the comprehensive engineering of cellular phenotypes and the creation of new biotechnologies. To achieve such modifications, the most widely used techniques rely on a host cell's native DNA repair mechanisms like homologous recombination, which hampers their broader use in organisms lacking these capabilities. Here, we explore the current landscape of genome integration systems with a particular focus on those that function in bacteria and are precise, self-contained, and portable, placing minimal requirements on the host cell. Through a historical analysis, we observe long-term use of recombineering technologies, a recent rise in the use of CRISPR-guided systems that consist of associated integrase machinery, and growing efforts to modify non-model organisms. Looking forward, we highlight some of the remaining challenges and how synthetic genomics may offer a way to create bacterial strains optimized for extensive long-term modification. As the field of synthetic biology sets its sights on real-world impact, the effective engineering of genomes will be critical to shaping the robust phenotypes that applications demand.
Whole-cell models (WCMs) are multi-scale computational models that aim to simulate the function of all genes and processes within a cell. This approach is promising for designing genomes tailored for specific tasks. However, a limitation of WCMs is their long runtime. Here, we show how machine learning (ML) surrogates can be used to address this limitation by training them on WCM data to accurately predict cell division. Our ML surrogate achieves a 95% reduction in computational time compared with the original WCM. We then show that the surrogate and a genome-design algorithm can generate an in silico-reduced E. coli cell, where 40% of the genes included in the WCM were removed. The reduced genome is validated using the WCM and interpreted biologically using Gene Ontology analysis. This approach illustrates how the holistic understanding gained from a WCM can be leveraged for synthetic biology tasks while reducing runtime. A record of this paper’s transparent peer review process is included in the supplemental information.
Recombinases are versatile enzymes able to perform the precise insertion, deletion, and rearrangement of DNA and can act as a foundation for programmable genetic logic and memory. Fundamental to their use are accurate measurements of function. However, these are often laborious, time-consuming, and costly to collect. To address this, we developed a semi-automated workflow that combines low-cost liquid handling robotics, multiplexed long-read nanopore sequencing, and a supporting computational analysis tool to enable the high-throughput and detailed characterization of recombinase parts and circuits when used in a variety of contexts and organisms. Our approach overcomes the limitations of typically used fluorescence-based assays and is able to monitor temporal dynamics, observe structural changes at a nucleotide resolution, and unravel the internal workings of complex multi-state circuits. The ability to scale-up and automate genetic circuit characterization is an essential step towards more rigorous biological metrology that can support the construction of predictive models for efficiently engineering biology. ### Competing Interest Statement The authors have declared no competing interest. Biotechnology and Biological Sciences Research Council, https://ror.org/00cwqg982, BB/W013959/1, BB/Y007638/1 Engineering and Physical Sciences Research Council, EP/N510129/1 Royal Society, https://ror.org/03wnrjx87, URF\R\221008 National Science Foundation, https://ror.org/021nxhr62, 2340175 BWF, Career Award at the Scientific Interface CZ Biohub, San Francisco Investigatorship
Protein-protein conjugation systems are a powerful way of creating fusion proteins and enable the dynamic combination of protein domains with diverse functionalities. However, the insertion of these systems into enzymes is often performed with little consideration of the structural impact they might have. This is particularly relevant when modifying complex molecular machines that transition between numerous conformational states. Here, we address this issue by developing SIMPLIFE, a computational workflow that supports the design of optimal insertion sites for conjugation tags based on the structure of the proteins involved and performs localised residue redesign where needed. We demonstrate how SIMPLIFE can be used to effectively augment the function of T7 RNA polymerase using the DogCatcher-DogTag system, enabling diverse and dynamically varying mutations within a targeted region of DNA. This work demonstrates the power of combining biophysical and machine learning based approaches for protein structure prediction to efficiently augment the function of molecular machines, accelerating our ability to combine complex biochemical functionalities in new ways. ### Competing Interest Statement The authors have declared no competing interest. Biotechnology and Biological Sciences Research Council, https://ror.org/00cwqg982, BB/W012448/1, BB/W013959/1 Engineering and Physical Sciences Research Council, EP/Y014073/1, EP/N510129/1, EP/S017542/2 Royal Society, URF/R/221008
Metabarcoding is a valuable tool for characterizing the communities that underpin the functioning of ecosystems. However, current methods often rely on polymerase chain reaction (PCR) amplification for enrichment of marker genes. PCR can introduce significant biases that affect quantification and is typically restricted to one target loci at a time, limiting the diversity that can be captured in a single reaction. Here, we address these issues by using Cas9 to enrich marker genes for long-read nanopore sequencing directly from a DNA sample, removing the need for PCR. We show that this approach can effectively isolate a 4.5 kb region covering partial 18S and 28S rRNA genes and the ITS region in a mixed nematode community, and further adapt our approach for characterizing a diverse microbial community. We demonstrate the ability for Cas9-based enrichment to support multiplexed targeting of several different DNA regions simultaneously, enabling optimal marker gene selection for different clades of interest within a sample. We also find a strong correlation between input DNA concentrations and output read proportions for mixed-species samples, demonstrating the ability for quantification of relative species abundance. This study lays a foundation for targeted long-read sequencing to more fully capture the diversity of organisms present in complex environments.
Both Japan and the UK have recognized the growing importance of synthetic and engineering biology for transforming life science research and transitioning toward a sustainable biobased economy. Such a shift will require extensive international cooperation and collaboration. In this viewpoint, we provide a summary of the recent "Japan-UK Synthetic Biology Conference, Spring 2025" that aimed to facilitate new links between researchers across the broad field of synthetic biology. We cover the core scientific topics discussed, distill some of the emerging trends, and outline the remaining challenges that are hampering progress. We end by highlighting some of the ways in which international collaborations may help address these issues through a combination of sharing expertise, national infrastructures, and aligned funding.
Biofilms are responsible for most chronic infections and are highly resistant to antibiotic treatments. Previous studies have demonstrated that periodic dosing of antibiotics can help sensitize persistent subpopulations and reduce the overall dosage required for treatment. Because the dynamics and mechanisms of biofilm growth and the formation of persister cells are diverse and are affected by environmental conditions, it remains a challenge to design optimal periodic dosing regimens. Here, we develop a computational agent-based model to streamline this process and determine key parameters for effective treatment. We used our model to test a broad range of persistence switching dynamics and found that if periodic antibiotic dosing was tuned to biofilm dynamics, the dose required for effective treatment could be reduced by nearly 77%. The biofilm architecture and its response to antibiotics were found to depend on the dynamics of persister cells. Despite some differences in the response of biofilm governed by different persister switching rates, we found that a general optimized periodic treatment was still effective in significantly reducing the required antibiotic dose. As persistence becomes better quantified and understood, our model has the potential to act as a foundation for more effective strategies to target bacterial infections.
ABSTRACT Computational protein design has emerged as a powerful tool for creating proteins with novel functionalities. However, most existing methods ignore structural dynamics even though they are known to play a central role in many protein functions. Furthermore, methods like molecular dynamics that are able to simulate protein movements are computationally demanding and do not scale for the design of even moderately sized proteins. Here, we develop a probabilistic coarse-grained model to overcome these limitations and support the design of the structural dynamics of modular repeat proteins. Our model allows us to rapidly calculate the probability distribution of structural conformations of large modular proteins, enabling efficient screening of design candidates based on features of their dynamics. We demonstrate this capability by exploring the design landscape of 4–6 module repeat proteins. We assess the flexibility, curvature and multi-state potential of over 65,000 protein variants and identify the roles that particular modules play in controlling these features. Although our focus here is on protein design, the methods developed are easily generalised to any modular structure (e.g., DNA origami), offering a means to incorporate dynamics into diverse biological design workflows.
Synthetic biology projects increasingly use modular DNA assembly or synthetic in vivo recombination to generate diverse combinatorial libraries of genetic constructs for testing. But as these designs expand to multigene systems it becomes challenging to sequence these in a cost-effective way that reveals the genotype to phenotype relationships in the libraries. Here, we introduce a new quick, low-cost method designed for assessing combinational designs of genome-integrated multigene constructs that we call Pool of Long Amplified Reads (POLAR) sequencing. POLAR-seq takes genomic DNA isolated from library pools and uses long range PCR to amplify target genomic regions up to 35 kb long containing combinatorial designs. The pool of long amplicons is then directly read by nanopore sequencing with full length reads then used to identify the gene content and structural variation of individual genotypes in the library and read count indicating how abundant a genotype is within the pool. Using yeast cells with loxP-containing synthetic gene clusters that rearrange in vivo in the presence of Cre recombinase, we demonstrate how POLAR-seq can be used to identify global patterns from combinatorial experiments, find the most abundant genotypes in a pool and also be adapted to sequence-verify gene clusters from isolated strains. ### Competing Interest Statement K.C. is now an employee of Oxford Nanopore Technologies but was solely employed by Imperial College London during generation of the data included in this paper. All other authors declare no conflicts of interest.
Keywords: biological biofabrication, engineered living materials, biohybrid materials, biofabrication, functionally graded biomaterials, synthetic biology, bioprogrammable materials, manufacture of biological systems