Ramkrishna, Kompala, and Tsao proposed the cybernetic model of microbial growth, in which cells allocate enzyme synthesis resources according to a matching rule that mimics rational decision-making. The matching rule was later shown to be optimal under general assumptions about the underlying return-on-investment structure, yet the specific objective the cell maximizes and the constraints bounding that choice were never written down as an explicit economic decision. Here we supply that missing decision, recasting cybernetic enzyme-synthesis control as a consumer choice problem from microeconomic theory: the cell allocates a limited proteome budget among competing catabolic enzymes as a linear program (LP), maximizing a linear growth utility subject to a linear proteome budget constraint. Because the utility is linear, the LP's solution is geometric: whenever the iso-utility line's slope differs from the budget constraint's, the optimum is a corner, and the entire proteome budget is allocated to the enzyme for the single most profitable substrate. Corner solutions correspond to diauxic growth, and sequential substrate consumption follows from the choice of corner rather than a distinct regulatory mechanism. Only when the two slopes coincide does the optimum spread across the entire budget line instead of concentrating at a single corner; this degenerate case is where simultaneous substrate use becomes admissible. Using kinetic parameters from single-substrate experiments together with a small set of model-level parameters set in this study, the LP-derived cybernetic variables reproduced the diauxic and triauxic batch growth of Klebsiella oxytoca on glucose-xylose and glucose-xylose-lactose mixtures, achieving a fit comparable to the classical matching law. Thus, sequential substrate use is the generic outcome of growth-maximizing specialization under perfect substitutability, and co-utilization is the degenerate case of equal profitability.
In silico tools are important for generating novel hypotheses and exploring alternatives in de novo metabolic pathway design. However, while many computational frameworks have been proposed for retrobiosynthesis, few successful examples of algorithm-guided xenobiotic biochemical retrosynthesis have been reported in the literature. Deep learning has improved the quality of synthesis and retrosynthesis in organic chemistry applications. Inspired by this progress, we explored combining deep learning of biochemical transformations with the traditional retrobiosynthetic workflow to improve in silico synthetic metabolic pathway designs. To develop our computational biosynthetic pathway design framework, we assembled metabolic reaction and enzymatic template data from public databases. A data augmentation procedure, adapted from literature, was carried out to enrich the assembled reaction dataset with artificial metabolic reactions generated by enzymatic reaction templates. Two neural network-based pathway ranking models were trained as binary classifiers to distinguish assembled reactions from artificial counterparts; each model output a scalar quantifying the plausibility of a 1-step or 2-step pathway. Combining these two models with enzymatic templates, we built a multistep retrobiosynthesis pipeline and validated it by reproducing some natural and non-natural pathways computationally.
We present BSTModelKit.jl, an open-source Julia package for constructing, solving, and analyzing Biochemical Systems Theory (BST) models of biochemical networks. The package implements S-system representations, a canonical power-law formalism for modeling metabolic and regulatory networks. BSTModelKit.jl provides a declarative model specification format, dynamic simulation via ordinary differential equation (ODE) integration, steady-state computation, and global sensitivity analysis using the Morris and Sobol methods. The package leverages the Julia scientific computing ecosystem, in particular the SciML suite of differential equation solvers, to provide efficient and flexible model analysis tools. We describe the mathematical formulation, software design, and demonstrate the package capabilities with illustrative examples.
Small longitudinal clinical cohorts, common in maternal health, rare diseases, and early-phase trials, limit computational modeling: too few patients to train reliable models, yet too costly and slow to expand through additional enrollment. We present multiplicity-weighted Stochastic Attention (SA), a generative framework based on modern Hopfield network theory that addresses this gap. SA embeds real patient profiles as memory patterns in a continuous energy landscape and generates novel synthetic patients via Langevin dynamics that interpolate between stored patterns while preserving the geometry of the original cohort. Per-pattern multiplicity weights enable targeted amplification of rare clinical subgroups at inference time without retraining. We applied SA to a longitudinal coagulation dataset from 23 pregnant patients spanning 72 biochemical features across 3 visits (pre-pregnancy baseline, first trimester, and third trimester), including rare subgroups such as polycystic ovary syndrome and preeclampsia. Synthetic patients generated by SA were statistically, structurally, and mechanistically indistinguishable from their real counterparts across multiple independent validation tests, including an ordinary differential equation model of the coagulation cascade. A downstream utility test further showed that a mechanistic model calibrated entirely on synthetic patients predicted held-out real patient outcomes as well as one calibrated on real data. These results demonstrate that SA can produce clinically useful synthetic cohorts from very small longitudinal datasets, enabling data-augmented modeling in small-cohort settings.
Most protein families have fewer than 100 known members, a regime where deep generative models overfit or collapse. We propose stochastic attention (SA), a training-free sampler that treats the modern Hopfield energy over a protein alignment as a Boltzmann distribution and draws samples via Langevin dynamics. The score function is a closed-form softmax attention operation requiring no training, no pretraining data, and no GPU, with cost linear in alignment size. Across eight Pfam families, SA generates sequences with low amino acid compositional divergence, substantial novelty, and structural plausibility confirmed by ESMFold and AlphaFold2. Generated sequences fold more faithfully to canonical family structures than natural members in six of eight families. Against profile HMMs, EvoDiff, and the MSA Transformer, which produce sequences that drift far outside the family, SA maintains 51 to 66 percent identity while remaining novel, in seconds on a laptop. The critical temperature governing generation is predicted from PCA dimensionality alone, enabling fully automatic operation. Controls confirm SA encodes correlated substitution patterns, not just per-position amino acid frequencies.
Mathematical models of natural and man-made systems often have many adjustable parameters that must be estimated from multiple, potentially conflicting datasets. Rather than reporting a single best-fit parameter vector, it is often more informative to generate an ensemble of parameter sets that collectively map out the trade-offs among competing objectives. This paper presents ParetoEnsembles.jl, an open-source Julia package that generates such ensembles using Pareto Optimal Ensemble Techniques (POETs), a simulated-annealing-based algorithm that requires no gradient information. The implementation corrects the original dominance relation from weak to strict Pareto dominance, reduces the per-iteration ranking cost from O n 2 m to O ( n m ) through an incremental update scheme, and adds multi-chain parallel execution for improved front coverage. We demonstrate the package on a cell-free gene expression model fitted to experimental data and a blood coagulation cascade model with ten estimated rate constants and three objectives. A controlled synthetic-data study reveals parameter identifiability structure, with individual rate constants off by several-fold yet model predictions accurate to 6-7%. A five-replicate coverage analysis confirms that timing features are reliably covered while peak amplitude is systematically overconfident. Validation against published experimental thrombin generation data demonstrates that the ensemble predicts held-out conditions to within 1-10% despite inherent model approximation error. By making ensemble generation lightweight and accessible, ParetoEnsembles.jl aims to lower the barrier to routine uncertainty characterization in mechanistic modeling.
Protein sequence generation via stochastic attention produces plausible family members from small alignments without training, but treats all stored sequences equally and cannot direct generation toward a functional subset of interest. We show that a single scalar parameter, added as a bias to the sampler's attention logits, continuously shifts generation from the full family toward a user-specified subset, with no retraining and no change to the model architecture. A practitioner supplies a small set of sequences (for example, hits from a binding screen) and a multiplicity ratio that controls how strongly generation favors them. The method is agnostic to what the subset represents: binding, stability, specificity, or any other property. We find that the conditioning is exact at the level of the sampler's internal representation, but that the decoded sequence phenotype can fall short because the dimensionality reduction used to encode sequences does not always preserve the residue-level variation that defines the functional split. We term this discrepancy the calibration gap and show that it is predicted by a simple geometric measure of how well the encoding separates the functional subset from the rest of the family. Experiments on five Pfam families (Kunitz, SH3, WW, Homeobox, and Forkhead domains) confirm the monotonic relationship between separation and gap across a fourfold range of geometries. Applied to omega-conotoxin peptides targeting a calcium channel involved in pain signaling, curated seeding from 23 characterized binders produces over a thousand candidates that preserve the primary pharmacophore and all experimentally identified binding determinants. These results show that stochastic attention enables practitioners to expand a handful of experimentally characterized sequences into diverse candidate libraries without retraining a generative model.
This work presents a generative pre-trained transformer (GPT) designed for modeling financial time series. The GPT functions as an order generation engine within a discrete event simulator, enabling realistic replication of limit order book dynamics. Our model leverages recent advancements in large language models to produce long sequences of order messages in a steaming manner. Our results demonstrate that the model successfully reproduces key features of order flow data, even when the initial order flow prompt is no longer present within the model's context window. Moreover, evaluations reveal that the model captures several statistical properties, or 'stylized facts', characteristic of real financial markets and broader macro-scale data distributions. Collectively, this work marks a significant step toward creating high-fidelity, interactive market simulations.
In this study, we developed a computational framework for simulating large-scale agent-based financial markets. Our platform supports trading multiple simultaneous assets and leverages distributed computing to scale the number and complexity of simulated agents. Heterogeneous agents make decisions in parallel, and their orders are processed through a realistic, continuous double auction matching engine. We present a baseline model implementation and show that it captures several known statistical properties of real financial markets (i.e., stylized facts). Further, we demonstrate these results without fitting models to historical financial data. Thus, this framework could be used for direct applications such as human-in-the-loop machine learning or to explore theoretically exciting questions about market microstructure's role in forming the statistical regularities of real markets. To the best of our knowledge, this study is the first to implement multiple assets, parallel agent decision-making, a continuous double auction mechanism, and intelligent agent types in a scalable real-time environment.
Amphiphilic copolymers (AP) represent a class of novel antibiofouling materials whose chemistry and composition can be tuned to optimize their performance. However, the enormous chemistry‐composition design space associated with AP makes their performance optimization laborious; it is not experimentally feasible to assess and validate all possible AP compositions even with the use of rapid screening methodologies. To address this constraint, a robust model development paradigm is reported, yielding a versatile machine learning approach that accurately predicts biofilm formation by Pseudomonas aeruginosa on a library of AP. The model excels in extracting underlying patterns in a “pooled” dataset from various experimental sources, thereby expanding the design space accessible to the model to a much larger selection of AP chemistries and compositions. The model is used to screen virtual libraries of AP for identification of best‐performing candidates for experimental validation. Initiated chemical vapor deposition is used for the precision synthesis of the model‐selected AP chemistries and compositions for validation at solid–liquid interface (often used in conventional antifouling studies) as well as the air–liquid–solid triple interface. Despite the vastly different growth conditions, the model successfully identifies the best‐performing AP for biofilm inhibition at the triple interface.
Cell-free synthetic systems are composed of the parts required for transcription and translation processes in a buffered solution. Thus, unlike living cells, cell-free systems are amenable to rapid adjustment of the reaction composition and easy sampling. Further, because cellular growth and maintenance requirements are absent, all resources can go toward synthesizing the product of interest. Recent improvement in key performance metrics, such as yield, reaction duration, and portability, has increased the space of possible applications open to cell-free systems and lowered the time required to design-build-test new circuitry. One promising application area is biosensing. This study describes developing and modeling a D-gluconate biosensor circuit operating in a reconstituted cell-free system. Model parameters were estimated using time-resolved measurements of the mRNA and protein concentration with and without the addition of D-gluconate. Sensor performance was predicted using the model for D-gluconate concentrations not used in model training. The model predicted the transcription and translation kinetics and the dose response of the circuit over several orders of magnitude of D-gluconate concentration. Global sensitivity analysis of the model parameters gave detailed insight into the operation of the sensor circuit. Taken together, this study reported an in-depth, systems-level analysis of a D-gluconate biosensor circuit operating in a reconstituted cell-free system. This circuit could be used directly to estimate D-gluconate or as a subsystem in a more extensive synthetic gene expression program.
Breast cancer metastasis is initiated by invasion of tumor cells into the collagen type I-rich stroma to reach adjacent blood vessels. Prior work has identified that metabolic plasticity is a key requirement of tumor cell invasion into collagen. However, it remains largely unclear how blood vessels affect this relationship. Here, we developed a microfluidic platform to analyze how tumor cells invade collagen in the presence and absence of a microvascular channel. We demonstrate that endothelial cells secrete pro-migratory factors that direct tumor cell invasion toward the microvessel. Analysis of tumor cell metabolism using metabolic imaging, metabolomics, and computational flux balance analysis revealed that these changes are accompanied by increased rates of glycolysis and oxygen consumption caused by broad alterations of glucose metabolism. Indeed, restricting glucose availability decreased endothelial cell-induced tumor cell invasion. Our results suggest that endothelial cells promote tumor invasion into the stroma due, in part, to reprogramming tumor cell metabolism.
Cell-free protein expression has become a widely used research tool in systems and synthetic biology and a promising technology for protein biomanufacturing. Cell-free protein synthesis relies on in-vitro transcription and translation processes to produce a protein of interest. However, transcription and translation depend upon the operation of complex metabolic pathways for precursor and energy regeneration. Toward understanding the role of metabolism in a cell-free system, we developed a dynamic constraint-based simulation of protein production in the myTXTL E. coli cell-free system with and without electron transport chain inhibitors. Time-resolved absolute metabolite measurements for â"³ = 63 metabolites, along with absolute concentration measurements of the mRNA and protein abundance and measurements of enzyme activity, were integrated with kinetic and enzyme abundance information to simulate the time evolution of metabolic flux and protein production with and without inhibitors. The metabolic flux distribution estimated by the model, along with the experimental metabolite and enzyme activity data, suggested that the myTXTL cell-free system has an active central carbon metabolism with glutamate powering the TCA cycle. Further, the electron transport chain inhibitor studies suggested the presence of oxidative phosphorylation activity in the myTXTL cell-free system; the oxidative phosphorylation inhibitors provided biochemical evidence that myTXTL relied, at least partially, on oxidative phosphorylation to generate the energy required to sustain transcription and translation for a 16-hour batch reaction.
Cell-free systems for gene expression have gained attention as platforms for the facile study of genetic circuits and as highly effective tools for teaching. Despite recent progress, the technology remains inaccessible for many in low- and middle-income countries due to the expensive reagents required for its manufacturing, as well as specialized equipment required for distribution and storage. To address these challenges, we deconstructed processes required for cell-free mixture preparation and developed a set of alternative low-cost strategies for easy production and sharing of extracts. First, we explored the stability of cell-free reactions dried through a low-cost device based on silica beads, as an alternative to commercial automated freeze dryers. Second, we report the positive effect of lactose as an additive for increasing protein synthesis in maltodextrin-based cell-free reactions using either circular or linear DNA templates. The modifications were used to produce active amounts of two high-value reagents: the isothermal polymerase Bst and the restriction enzyme BsaI. Third, we demonstrated the endogenous regeneration of nucleoside triphosphates and synthesis of pyruvate in cell-free systems (CFSs) based on phosphoenol pyruvate (PEP) and maltodextrin (MDX). We exploited this novel finding to demonstrate the use of a cell-free mixture completely free of any exogenous nucleotide triphosphates (NTPs) to generate high yields of sfGFP expression. Together, these modifications can produce desiccated extracts that are 203-424-fold cheaper than commercial versions. These improvements will facilitate wider use of CFS for research and education purposes.
Cell-free systems are a widely used research tool in systems and synthetic biology and a promising platform for manufacturing of proteins and chemicals. In the past, cell-free biology was primarily used to better understand fundamental biochemical processes. Notably, E. coli cell-free extracts were used in the 1960s to decipher the sequencing of the genetic code. Since then, the transcription and translation capabilities of cell-free systems have been repeatedly optimized to improve energy efficiency and product yield. Today, cell-free systems, in combination with the rise of synthetic biology, have taken on a new role as a promising technology for just-in-time manufacturing of therapeutically important biologics and high-value small molecules. They have also been implemented at an industrial scale for the production of antibodies and cytokines. In this review, we discuss the evolution of cell-free technologies, in particular advancements in extract preparation, cell-free protein synthesis, and cell-free metabolic engineering applications. We then conclude with a discussion of the mathematical modeling of cell-free systems. Mathematical modeling of cell-free processes could be critical to addressing performance bottlenecks and estimating the costs of cell-free manufactured products.
A major objective of synthetic glycobiology is to re-engineer existing cellular glycosylation pathways from the top down or construct non-natural ones from the bottom up for new and useful purposes. Here, we have developed a set of orthogonal pathways for eukaryotic O-linked protein glycosylation in Escherichia coli that installed the cancer-associated mucin-type glycans Tn, T, sialyl-Tn and sialyl-T onto serine residues in acceptor motifs derived from different human O-glycoproteins. These same glycoengineered bacteria were used to supply crude cell extracts enriched with glycosylation machinery that permitted cell-free construction of O-glycoproteins in a one-pot reaction. In addition, O-glycosylation-competent bacteria were able to generate an antigenically authentic Tn-MUC1 glycoform that exhibited reactivity with antibody 5E5, which specifically recognizes cancer-associated glycoforms of MUC1. We anticipate that the orthogonal glycoprotein biosynthesis pathways developed here will provide facile access to structurally diverse O-glycoforms for a range of important scientific and therapeutic applications.
Transcription and translation are at the heart of metabolism and signal transduction. In this study, we developed an effective biophysical modeling approach to simulate transcription and translation processes. The model, composed of coupled ordinary differential equations, was tested by comparing simulations of two cell free synthetic circuits with experimental measurements generated in this study. First, we considered a simple circuit in which sigma factor 70 induced the expression of green fluorescent protein. This relatively simple case was then followed by a more complex negative feedback circuit in which two control genes were coupled to the expression of a third reporter gene, green fluorescent protein. Many of the model parameters were estimated from previous biophysical studies in the literature, while the remaining unknown model parameters for each circuit were estimated by minimizing the difference between model simulations and messenger RNA (mRNA) and protein measurements generated in this study. In particular, either parameter estimates from published studies were used directly, or characteristic values found in the literature were used to establish feasible ranges for the parameter estimation problem. In order to perform a detailed analysis of the influence of individual model parameters on the expression dynamics of each circuit, global sensitivity analysis was used. Taken together, the effective biophysical modeling approach captured the expression dynamics, including the transcription dynamics, for the two synthetic cell free circuits. While, we considered only two circuits here, this approach could potentially be extended to simulate other genetic circuits in both cell free and whole cell biomolecular applications as the equations governing the regulatory control functions are modular and easily modifiable. The model code, parameters, and analysis scripts are available for download under an MIT software license from the Varnerlab GitHub repository.
Many markup languages can be used to encode biological networks, each with strengths and weaknesses. Model specifications written in these languages can then used, in conjunction with proprietary software packages e.g., MATLAB, or open community alternatives, to simulate the behavior of biological systems. In this study, we present the Simplified English Modeling Language (SEML) and associated compiler, as an alternative to existing approaches. SEML supports the specification of biological reaction systems in a simple natural language like syntax. Models encoded in SEML are transformed into executable code using a compiler written in the open-source Julia programming language. The compiler performs a sequence of operations, including tokenization, syntactic and semantic error checking, to convert SEML into an intermediate representation (IR). From the intermediate representation, the compiler then generates executable code in one of three programming languages: Julia, Python or MATLAB. Currently, SEML supports both kinetic and constraint based model generation for signal transduction and metabolic modeling. In this study, we demonstrate SEML by modeling two proof-of-concept prototypical networks: a constraint-based model solved using flux balance analysis (FBA) and a kinetic model encoded as Ordinary Differential Equations (ODEs). SEML is a promising tool for encoding and sharing human-readable biological models, however it is still in its infancy. With further development, SEML has the potential to handle more unstructured natural language inputs, generate more complex models types and convert its natural language markup to currently used model interchange formats such systems biology markup language.
Clinical studies have shown that all-trans retinoic acid (RA), which is often used in treatment of cancer patients, improves hemostatic parameters and bleeding complications such as disseminated intravascular coagulation (DIC). However, the mechanisms underlying this improvement have yet to be elucidated. In vitro studies have reported that RA upregulates thrombomodulin (TM) expression on the endothelial cell surface. The objective of this study was to investigate how and to what extent the TM concentration changes after RA treatment in cancer patients, and how this variation influences the blood coagulation cascade. In this study, we introduced an ordinary differential equation (ODE) model of gene expression for the RA-induced upregulation of TM concentration. Coupling the gene expression model with a two-compartment pharmacokinetic model of RA, we obtained the time-dependent changes in TM and thrombomodulin-mRNA (TMR) concentrations following oral administration of RA. Our results indicated that the TM concentration reached its peak level almost 14 h after taking a single oral dose (110 $$ \frac{mg}{m^2} $$ ) of RA. Continuous treatment with RA resulted in oscillatory expression of TM on the endothelial cell surface. We then coupled the gene expression model with a mechanistic model of the coagulation cascade, and showed that the elevated levels of TM over the course of RA therapy with a single daily oral dose (110 $$ \frac{mg}{m^2} $$ ) of RA, reduced the peak thrombin levels and endogenous thrombin potential (ETP) up to 50 and 49%, respectively. We showed that progressive reductions in plasma levels of RA, observed in continuous RA therapy with a once-daily oral dose (110 $$ \frac{mg}{m^2} $$ ) of RA, did not affect TM-mediated reduction of thrombin generation significantly. This finding prompts the hypothesis that continuous RA treatment has more consistent therapeutic effects on coagulation disorders than on cancer. Our results indicate that the oscillatory upregulation of TM expression on the endothelial cells over the course of RA therapy could potentially contribute to the treatment of coagulation abnormalities in cancer patients. Further studies on the impacts of RA therapy on the procoagulant activity of cancer cells are needed to better elucidate the mechanisms by which RA therapy improves hemostatic abnormalities in cancer.
In non-acute promyelotic leukemia (APL)- non myelocytic leukemia (AML), identification of a signaling signature would predict potentially actionable targets to enhance differentiation effects of all-trans-retinoic acid (RA) and make combination differentiation therapy realizable. Components of such a signaling machine/signalsome found to drive RA-induced differentiation discerned in a FAB M2 cell line/model (HL-60) were further characterized and then compared against AML patient expression profiles. FICZ, known to enhance RA-induced differentiation, was used to experimentally augment signaling for analysis. FRET revealed novel signalsome protein associations: CD38 with pS376SLP76 and caveolin-1 with CD38 and AhR. The signaling molecules driving differentiation in HL-60 cluster in non-APL AML de novo samples, too. Pearson correlation coefficients for this molecular ensemble are nearer 1 in the FAB M2 subtype than in non-APL AML. SLP76 correlation to RXR alpha and p47phox were conserved in FAB M2 model and patient subtype but not in general non-APL AML. The signalsome ergo identifies potential actionable targets in AML.