Yarrowia lipolytica is a promising host for a range of biotechnological applications. However, the costs and sustainability of common feedstocks limit the commercial viability of current bioproduction. Consolidated bioprocessing using plant biomass offers a desirable alternative to reduce costs and improve sustainability, though such approaches are often constrained in single-organism systems due to metabolic burden. Here, we have investigated the use of division of labour (DOL) in Y. lipolytica consortia for the degradation of starch, a common component of plant biomass. We engineered a panel of strains expressing α-amylase and/or glucoamylase, with varying promoter and signal peptide combinations. We found that stronger expression led to higher metabolic burden, and that expressing both amylases imposed a greater burden than expressing either enzyme alone. When combining both amylase strains, we found that some consortia could achieve faster growth rates than the equivalent monoculture. The fastest growing consortium identified from the panel was scaled-up into flasks, where the consortium achieved the highest final biomass, although the fastest growth rate was achieved by the α-amylase strain alone. These findings suggest that DOL is a promising strategy for degrading complex macromolecules in plant biomass, offering a potential route to more sustainable bioprocessing.
Periodic behavior is a widespread biological phenomenon that occurs on multiple spatiotemporal scales, where upstream stimuli are encoded into dynamic intracellular signals. Restoring rhythmic dynamics after perturbations, or inducing them in otherwise steady-state systems, typically requires redesigning the underlying network, limiting scalability. Here we propose a reference-free molecular feedback controller built from the Incoherent Feedforward Loop (IFFL) motif that destabilizes the steady state of a non-oscillating system into sustained oscillations. We apply the controller to a self-inhibiting gene and show, through a Hopf bifurcation analysis, that sustained oscillations are induced in the fast sequestration regime when the indirect production branch of the IFFL motif dominates the direct one, and the actuation gain is sufficiently large relative to the process. Numerical simulations validate these conditions and characterize the robustness of the resulting oscillations, revealing that the period is robust to parametric perturbations while the amplitude is sensitive. Ultimately, we explore whether the same IFFL architecture extends to higher-dimensional processes by applying it, as a case study, to a negative feedback system in its non-oscillatory regime.
Engineered microbial consortia are emerging as programmable systems capable of sensing and responding to their environment. However, maintaining defined community composition over time remains challenging, particularly in bioprocesses where growth conditions and metabolic burdens continuously shift. Here, we develop a burden-aware multicellular RNA-based feedback control system that stabilises coculture composition by coupling gene expression burden to growth regulation. The system integrates three modules: quorum sensing-based communication, an RNA-based comparator computing deviations from a target ratio, and tuneable growth regulation via heterologous expression burden or CRISPRi-mediated knockdowns. In a two-strain E. coli coculture, this architecture maintains stable coculture ratios over 24-hour batch cultures, recovers growth rates by up to 90% following burden-induced growth reduction, and increases protein production yields by up to 81% in the slower-growing strain. We achieve tuneability by adjusting RNA binding strength and quorum-sensing signal production. This work demonstrates that burden-driven growth control can be used to stabilise and tune synthetic microbial consortia.
Information propagation by sequence-specific, template-catalysed molecular assembly is a key process facilitating life's biochemical complexity, yielding thousands of sequence-defined proteins from only 20 distinct building blocks. However, exploitation of catalytic templating is rare in non-biological contexts, particularly in enzyme-free environments, where even the template-catalysed formation of dimers is challenging. Typically, product inhibition-the tendency of products to bind to templates more strongly than individual monomers-prevents catalytic turnover. Here we present a rationally designed enzyme-free system in which a DNA template catalyses, with weak product inhibition, the production of sequence-specific DNA dimers. We demonstrate selective templating of nine different dimers with high specificity and catalytic turnover, then we show that the products can participate in downstream reactions, and finally that the dimerization can be coupled to covalent bond formation. Most importantly, our mechanism demonstrates a design principle for constructing synthetic molecular templating systems, a first step towards applying this powerful motif in non-biological contexts to construct many complex molecules and materials from a small number of building blocks.
Dynamic DNA nanotechnology creates programmable reaction networks and nanodevices by using DNA strands. The key reaction in dynamic DNA nanotechnology is the exchange of DNA strands between different molecular species, achieved through three- and four-way strand exchange reactions. While both reactions have been widely used, the four-way exchange reaction has traditionally been slower and less efficient than the three-way reaction. In this paper, we describe a new mechanism to optimize the kinetics of the four-way strand exchange reaction by adding bulges to the toeholds of the four-way DNA complexes involved in the reaction. These bulges facilitate an alternative branch migration mechanism and destabilize the four-way DNA junction, increasing the four-way strand exchange rate by an order of magnitude. This advancement has the potential to expand the field of dynamic DNA nanotechnology by enabling efficient four-way strand exchange reactions for in vivo applications.
Large Language Models (LLMs) demonstrate remarkable generalizability across diverse tasks, yet genomic foundation models (GFMs) still require separate finetuning for each downstream application, creating significant overhead as model sizes grow. Moreover, existing GFMs are constrained by rigid output formats, limiting their applicability to various genomic tasks. In this work, we revisit the transformer-based auto-regressive models and introduce Omni-DNA, a family of cross-modal multi-task models ranging from 20 million to 1 billion parameters. Our approach consists of two stages: (i) pretraining on DNA sequences with next token prediction objective, and (ii) expanding the multi-modal task-specific tokens and finetuning for multiple downstream tasks simultaneously. When evaluated on the Nucleotide Transformer and GB benchmarks, Omni-DNA achieves state-of-the-art performance on 18 out of 26 tasks. Through multi-task finetuning, Omni-DNA addresses 10 acetylation and methylation tasks at once, surpassing models trained on each task individually. Finally, we design two complex genomic tasks, DNA2Function and Needle-in-DNA, which map DNA sequences to textual functional descriptions and images, respectively, indicating Omni-DNA's cross-modal capabilities to broaden the scope of genomic applications. All the models are available through https://huggingface.co/collections/zehui127
With advancements in synthetic biology and metabolic engineering, microorganisms can now be engineered to perform increasingly complex functions, which may be limited by the resources available in individual cells. Introducing heterologous metabolic pathways introduces both genetic burden due to the competition for cellular transcription and translational machinery, as well as metabolic burden due to the redirection of metabolic flux from the native metabolic pathways. Division of labor in synthetic microbial communities offers a promising approach to enhance metabolic efficiency and resilience in bioproduction. By distributing complex metabolic pathways across multiple subpopulations, the resource competition and metabolic burden imposed on an individual cell are reduced, potentially enabling more efficient production of target compounds. Violacein is a high-value pigment with antitumor properties that exemplifies such a challenge due to its complex bioproduction pathway, imposing a significant metabolic burden on host cells. In this study, we investigated the benefits of division of labor for violacein production by splitting the violacein bioproduction pathway between two subpopulations of Escherichia coli-based synthetic communities. We tested several pathway splitting strategies and reported that splitting the pathway into two subpopulations expressing VioABE and VioDC at a final composition of 60:40 yields a 2.5-fold increase in violacein production as compared to a monoculture. We demonstrated that the coculture outperforms the monoculture when both subpopulations exhibit similar metabolic burden levels, resulting in comparable growth rates, and when both subpopulations are present in sufficiently high proportions.
Graph augmentation methods play a crucial role in improving the performance and enhancing generalisation capabilities in Graph Neural Networks (GNNs). Existing graph augmentation methods mainly perturb the graph structures, and are usually limited to pairwise node relations. These methods cannot fully address the complexities of real-world large-scale networks, which often involve higher-order node relations beyond only being pairwise. Meanwhile, real-world graph datasets are predominantly modelled as simple graphs, due to the scarcity of data that can be used to form higher-order edges. Therefore, reconfiguring the higher-order edges as an integration into graph augmentation strategies lights up a promising research path to address the aforementioned issues. In this paper, we present Topological Augmentation (TopoAug), a novel graph augmentation method that builds a combinatorial complex from the original graph by constructing virtual hyperedges directly from the raw data. TopoAug then produces auxiliary node features by extracting information from the combinatorial complex, which are used for enhancing GNN performances on downstream tasks. We design three diverse virtual hyperedge construction strategies to accompany the construction of combinatorial complexes: (1) via graph statistics, (2) from multiple data perspectives, and (3) utilising multi-modality. Furthermore, to facilitate TopoAug evaluation, we provide 23 novel real-world graph datasets across various domains including social media, biology, and e-commerce. Our empirical study shows that TopoAug consistently and significantly outperforms GNN baselines and other graph augmentation methods, across a variety of application contexts, which clearly indicates that it can effectively incorporate higher-order node relations into the graph augmentation for real-world complex networks.
Within a cell, synthetic and native genes compete for expression machinery, influencing cellular process dynamics through resource couplings. Models that simplify competitive resource binding kinetics can guide the design of strategies for countering these couplings. However, in bacteria resource availability and cell growth rate are interlinked, which complicates resource-aware biocircuit design. Capturing this interdependence requires coarse-grained bacterial cell models that balance accurate representation of metabolic regulation against simplicity and interpretability. We propose a coarse-grained E. coli cell model that combines the ease of simplified resource coupling analysis with appreciation of bacterial growth regulation mechanisms and the processes relevant for biocircuit design. Reliably capturing known growth phenomena, it provides a unifying explanation to disparate empirical relations between growth and synthetic gene expression. Considering a biomolecular controller that makes cell-wide ribosome availability robust to perturbations, we showcase our model’s usefulness in numerically prototyping biocircuits and deriving analytical relations for design guidance.
Spores are highly resistant dormant cells, adapted for survival and dispersal, that can withstand unfavourable environmental conditions for extended periods of time and later reactivate. Understanding the germination process of microbial spores is important in numerous areas including agriculture, food safety and health, and other sectors of biotechnology. Microfluidics combined with high-resolution microscopy allows to study spore germination at the single-cell level, revealing behaviours that would be hidden in standard population-level studies. Here, we present a microfluidic platform – the so-called four-conditions microfluidic chemostat (4CMC) – for germination studies where spores are confined to monolayers inside microchambers, allowing the testing of four growth conditions in parallel. This platform can be used with multiple species, including non-model organisms, and is compatible with existing image analysis software. In this study, we focused on three soil dwellers, two bacteria and one fungus, and revealed new insights into their germination. We studied endospores of the model bacterium Bacillus subtilis and demonstrated a correlation between spore density and germination in rich media. We then investigated the germination of the obligate-oxalotrophic environmental bacterium Ammoniphilus oxalaticus in a concentration gradient of potassium oxalate, showing that lower concentrations result in more spores germinating compared to higher concentrations. We also used this microfluidic platform to study the soil beneficial filamentous fungus Trichoderma rossicum, showing for the first time that the size of the spores and hyphae increase in response to increased nutrient availability, while germination times remain the same. Our platform allows to better understand microbial behaviour at the single-cell level, under a variety of controlled conditions. While we used it to decipher the responsiveness of soil dwellers’ spores, it would also be suitable for other spores from bacteria or filamentous fungi, but also vegetative cells and yeast, and even microbial communities.
Advances in synthetic biology depend on our ability to predictably engineer robust biomolecular systems in living cells. The functioning of these synthetic biomolecular systems requires the consumption of shared cellular resources, which imposes a gene expression burden that may impact the performance of the cell and the synthetic system. In this paper, we show the effect of resource constraints on quantitative and qualitative aspects of gene expression in multiple circuits. We utilise a resource-aware modelling framework to show that stabilization can be achieved in a class of integral controllers. The results open possibilities for the design of lean biomolecular controllers.
Graphs are widely used to encapsulate a variety of data formats, but real-world networks often involve complex node relations beyond only being pairwise. While hypergraphs and hierarchical graphs have been developed and employed to account for the complex node relations, they cannot fully represent these complexities in practice. Additionally, though many Graph Neural Networks (GNNs) have been proposed for representation learning on higher-order graphs, they are usually only evaluated on simple graph datasets. Therefore, there is a need for a unified modelling of higher-order graphs, and a collection of comprehensive datasets with an accessible evaluation framework to fully understand the performance of these algorithms on complex graphs. In this paper, we introduce the concept of hybrid graphs, a unified definition for higher-order graphs, and present the Hybrid Graph Benchmark (HGB). HGB contains 23 real-world hybrid graph datasets across various domains such as biology, social media, and e-commerce. Furthermore, we provide an extensible evaluation framework and a supporting codebase to facilitate the training and evaluation of GNNs on HGB. Our empirical study of existing GNNs on HGB reveals various research opportunities and gaps, including (1) evaluating the actual performance improvement of hypergraph GNNs over simple graph GNNs; (2) comparing the impact of different sampling strategies on hybrid graph learning methods; and (3) exploring ways to integrate simple graph and hypergraph information. We make our source code and full datasets publicly available at https://zehui127.github.io/hybrid-graph-benchmark/.
Competition for intracellular resources, also known as gene expression burden, induces coupling between independently co-expressed genes, a detrimental effect on predictability and reliability of gene circuits in mammalian cells. We recently showed that microRNA (miRNA)-mediated target downregulation correlates with the upregulation of a co-expressed gene, and by exploiting miRNAs-based incoherent-feed-forward loops (iFFLs) we stabilise a gene of interest against burden. Considering these findings, we speculate that miRNA-mediated gene downregulation causes cellular resource redistribution. Despite the extensive use of miRNA in synthetic circuits regulation, this indirect effect was never reported before. Here we developed a synthetic genetic system that embeds miRNA regulation, and a mathematical model, MIRELLA, to unravel the miRNA (MI) RolE on intracellular resource aLLocAtion. We report that the link between miRNA-gene downregulation and independent genes upregulation is a result of the concerted action of ribosome redistribution and 'queueing-effect' on the RNA degradation pathway. Taken together, our results provide for the first time insights into the hidden regulatory interaction of miRNA-based synthetic networks, potentially relevant also in endogenous gene regulation. Our observations allow to define rules for complexity- and context-aware design of genetic circuits, in which transgenes co-expression can be modulated by tuning resource availability via number and location of miRNA target sites.
Microbial consortia have been utilised for centuries to produce fermented foods and have great potential in applications such as therapeutics, biomaterials, fertilisers, and biobased production. Working together, microbes become specialized and perform complex tasks more efficiently, strengthening both cooperation and stability of the microbial community. However, imbalanced proportions of microbial community members can lead to unoptimized and diminished yields in biotechnology. To address this, we developed a burden-aware RNA-based multicellular feedback control system that stabilises and tunes coculture compositions. The system consists of three modules: a quorum sensing-based communication module to provide information about the densities of cocultured strains, an RNA-based comparator module to compare the ratio of densities of both strains to a pre-set desired ratio, and a customisable growth module that relies either on heterologous gene expression or on CRISPRi knockdowns to tune growth rates. We demonstrated that heterologous expression burden could be used to stabilise composition in a two-member E. coli coculture. This is the first coculture composition controller that does not rely on toxins or syntrophy for growth regulation and uses RNA sequestration to stabilise and control coculture composition. This work provides a fundamental basis to explore burden-aware multicellular feedback control strategies for robust stabilisation of synthetic community compositions.
Predicting the evolution of engineered cell populations is a highly sought-after goal in biotechnology. While models of evolutionary dynamics are far from new, their application to synthetic systems is scarce where the vast combination of genetic parts and regulatory elements creates a unique challenge. To address this gap, we here-in present a framework that allows one to connect the DNA design of varied genetic devices with mutation spread in a growing cell population. Users can specify the functional parts of their system and the degree of mutation heterogeneity to explore, after which our model generates host-aware transition dynamics between different mutation phenotypes over time. We show how our framework can be used to generate insightful hypotheses across broad applications, from how a device's components can be tweaked to optimise long-term protein yield and genetic shelf life, to generating new design paradigms for gene regulatory networks that improve their functionality.
Laboratory automation and mathematical optimisation are key to improving the efficiency of synthetic biology research. While there are algorithms optimising the construct designs and synthesis strategies for DNA assembly, the optimisation of how DNA assembly reaction mixes are prepared remains largely unexplored. Here, we focus on reducing the pipette tip consumption of a liquid-handling robot as it delivers DNA parts across a multi-well plate where several constructs are being assembled in parallel. We propose a linear programming formulation of this problem based on the capacitated vehicle routing problem, along with an algorithm which applies a linear programming solver to our formulation, hence providing a strategy to prepare a given set of DNA assembly mixes using fewer pipette tips. The algorithm performed well in randomly generated and real-life scenarios concerning several modular DNA assembly standards, proving capable of reducing the pipette tip consumption by up to 61% in large-scale cases. Combining automatic process optimisation and robotic liquid-handling, our strategy promises to greatly improve the efficiency of DNA assembly, either used alone or in combination with other algorithmic methods.
Gene expression depends on the cellular con-text. One major contributor to gene expression variability is competition for limited transcriptional and translational re-sources, which may induce indirect couplings among otherwise independently-regulated genes. Here, we apply control theoretical concepts and tools to design an incoherent feedforward loop (iFFL) biomolecular controller operating in mammalian cells using translational-resource competition couplings. Harnessing a resource-aware mathematical model, we demonstrate analytically and computationally that our resource-aware design can achieve near-constant set-point regulation of gene expression whilst ensuring robustness to plasmid uptake variation. We also provide an analytical condition on the model parameters to guide the design of the resource-aware iFFL controller ensuring robustness and performance in set-point regulation. Our theoretical design based on translational-resource competition couplings represents a promising approach to build more sophisticated resource-aware control circuits operating at the host-cell level.
While inter-lab calibration standards are approaching mainstream usage in synthetic biology, such calibrations are not in fact sufficient for absolute protein quantification required for modelling synthetic circuits. Fluorescein-based calibration of plate reader and flow cytometry instruments allows the measurement of green fluorescent protein (GFP) in synthetic cells to graduate from arbitrary units to calibrated units, but retains important caveats. Fluorescein is only a good calibrant for green FPs, leaving other FPs uncalibrated, and only provides conversions to units of brightness, not to molecule numbers. Ideal assay calibrants in molecular biology consist of the same molecule as the one to be measured – in this case, a purified preparation of the appropriate fluorescent protein. Here we show that by using purified FP calibrants, all protein species in a synthetic circuit can be quantified in absolute terms using no advanced instrumentation. We develop a SEVA (Standardised European Vector Architecture)-based expression vector that allows the high-level production of soluble protein and describe a straightforward 2-day protocol for the purification of micrograms of FP, followed by a calibration that relates fluorescence activity to protein mass. We validate this protocol by calculating conversion factors for a panel of commonly-used FPs including superfolder GFP, mCherry, mScarlet-I and mTagBFP2 on multiple laboratory instruments, and use these calibrations to debug synthetic circuits. We also demonstrate that the suspected bias of the presence of mCherry on OD600 measurements is real, but in practice requires extreme overexpression (~100,000 proteins per cell) to have a meaningful impact on cell density estimates.
Abstract Despite advances in bacterial genome engineering, delivery of large synthetic constructs remains challenging in practice. In this study, we propose a straightforward and robust approach for the markerless integration of DNA fragments encoding whole metabolic pathways into the genome. This approach relies on the replacement of a counterselection marker with cargo DNA cassettes via λRed recombineering. We employed a counterselection strategy involving a genetic circuit based on the CI repressor of λ phage. Our design ensures elimination of most spontaneous mutants, and thus provides a counterselection stringency close to the maximum possible. We improved the efficiency of integrating long PCR-generated cassettes by exploiting the Ocr antirestriction function of T7 phage, which completely prevents degradation of unmethylated DNA by restriction endonucleases in wild-type bacteria. The employment of highly restrictive counterselection and ocr-assisted λRed recombineering allowed markerless integration of operon-sized cassettes into arbitrary genomic loci of four enterobacterial species with an efficiency of 50–100%. In the case of Escherichia coli, our strategy ensures simple combination of markerless mutations in a single strain via P1 transduction. Overall, the proposed approach can serve as a general tool for synthetic biology and metabolic engineering in a range of bacterial hosts.
BACKGROUND:Low-cost sustainable feedstocks are essential for commercially viable biotechnologies. These feedstocks, often derived from plant or food waste, contain a multitude of different complex biomolecules which require multiple enzymes to hydrolyse and metabolise. Current standard biotechnology uses monocultures in which a single host expresses all the proteins required for the consolidated bioprocess. However, these hosts have limited capacity for expressing proteins before growth is impacted. This limitation may be overcome by utilising division of labour (DOL) in a consortium, where each member expresses a single protein of a longer degradation pathway.RESULTS:Here, we model a two-strain consortium, with one strain expressing an endohydrolase and a second strain expressing an exohydrolase, for cooperative degradation of a complex substrate. Our results suggest that there is a balance between increasing expression to enhance degradation versus the burden that higher expression causes. Once a threshold of burden is reached, the consortium will consistently perform better than an equivalent single-cell monoculture.CONCLUSIONS:We demonstrate that resource-aware whole-cell models can be used to predict the benefits and limitations of using consortia systems to overcome burden. Our model predicts the region of expression where DOL would be beneficial for growth on starch, which will assist in making informed design choices for this, and other, complex-substrate degradation pathways.