Ribonucleic acid (RNA) molecules play informational, structural, and metabolic roles in all living cells. RNAs are chains of nucleotides containing bases {A, C, G, U} that interact via base pairings to determine higher order structure and functionality. The RNA folding problem is to predict one or more secondary RNA structures from a given primary sequence of bases. From a mathematical modeling perspective, solutions to the RNA folding problem come from minimizing the thermodynamic free energy of a structure by selecting which bases will be paired, subject to a set of constraints. Here we report on a Quadratic Unconstrained Binary Optimization (QUBO) modeling paradigm that fits naturally with the parameters and constraints required for RNA folding prediction. Three QUBO models are presented along with a hybrid metaheuristic algorithm. Extensive testing results show a strong positive correlation with benchmark results.
This synthetic biology study explores if BCDs (Bicistronic designs) would be able to operate as effectively in vitro as they do in vivo, to produce a predictable level of protein.
Synthetic biologists construct parts, devices and systems to engineer cells for applications in medicine, biofuels, chemical commodities, and the environment. A common problem with projects that focus on the synthesis of natural and engineered proteins is unpredictable translation directed by ribosomal binding sites (RBS). Researchers have hypothesized that base pairing interactions between RBSs and downstream coding sequences form secondary structures that affect the efficiency of translation initiation. Bicistronic designs (BCDs) were developed in 2013 by Drew Endy and his colleagues to enable more predictable and reliable in vivo translation. BCDs have an RBS that leads to the production of a leader polypeptide with no functionality and a second RBS that directs the production of a protein of interest. BCDs are better at producing predictable protein levels in bacterial cells than simple RBSs, thereby improving the ability of BCDs to function reliably and predictably as classic synthetic biology parts. In a companion study, we used GFP expression to compare the translational efficiencies of 19 BCDs during cell‐free protein synthesis (CFPS) to their efficiencies in vivo and found a rank correlation of 0.88. This report describes our efforts to expand the toolkit for CFPS protein production to microfluidic droplets. We used a microfluidic flow control system to rapidly produce nanoliter‐scale droplets and characterized them with microscopy. Image analysis of the droplet fluorescence intensities allowed us to quantify and compare BCD‐directed protein synthesis within droplets. We automated this process for many images to accelerate the workflow. We focused on three BCDs for CFPS of GFP in droplets and documented significant increases in observed fluorescence compared to batch CFPS outside of droplets. We also found that the rank order function for the three BCDs was the same in droplets as it was in batch CFPS. Our results support the use of BCDs to gain predictability of protein production in CFPS. These results set the stage for partitioning large libraries of gene regulatory combinations into microfluidic droplets to find the best combinations suited for a given synthetic biology application.Support or Funding InformationNSF RUI MCB‐1613281 to Missouri Western State University, NSF RUI MCB‐1613203 to Davidson College
Synthetic biology integrates molecular biology tools and an engineering mindset to address challenges in medicine, agriculture, bioremediation, and biomanufacturing. A persistent problem in synthetic biology has been designing genetic circuits that produce predictable levels of protein. In 2013, Mutalik and colleagues developed bicistronic designs (BCDs) that make protein production more predicable in bacterial cells (in vivo). With the growing interest in producing proteins outside of cells (in vitro), we wanted to know if BCDs would work as predictably in cell-free protein synthesis (CFPS) as they do in E. coli cells. We tested 20 BCDs in CFPS and found they performed very similarly in vitro and in vivo. As a step toward developing methods for protein production in artificial cells, we also tested 3 BCDs inside nanoliter-scaled microfluidic droplets. The BCDs worked well in the microfluidic droplets, but their relative protein production levels were not as predictable as expected. These results suggest that the conditions under which gene expression happens in droplets result in a different relationship between genetic control elements such as BCDs and protein production than exists in batch CFPS or in cells. KEYWORDS: Bicistronic Design; Synthetic Biology; Cell-Free Protein Synthesis; Microfluidics
Synthetic biology takes an engineering approach to biological systems for the construction of parts, devices and systems that address challenges associated with medicine, biofuels, the environment, and more. Although the production of natural and engineered proteins is fundamental to this effort, a persistent challenge to protein production has been unreliable translation directed by a given ribosomal binding site (RBS) depending on which gene of interest is downstream of it. Based on many observations, investigators have hypothesized that the 5’ UTR containing an RBS and early codons can bind to downstream coding mRNA to form base paired secondary structures that decrease the efficiency of translation initiation. In 2013, Drew Endy and colleagues developed bicistronic designs (BCDs) for predictable and reliable in vivo translation. BCDs contain two RBSs and two protein‐encoding cistrons. The first RBS leads to the production of a 16 amino acid leader polypeptide with no functionality, whereas the second RBS directs the production of a protein encoded by a gene of interest. A key feature of BCDs is that the stop codon for the first cistron (TAA) shares its terminal adenine with the first base of the start codon (ATG) for the second cistron. In other words, TAATG is both a stop codon in one reading frame of cistron 1 and a start codon in a shifted reading frame for cistron 2. BCDs were found to be better at producing predictable protein levels in vivo than simple RBSs, thereby improving the ability of BCDs to function reliably and predictably as classic synthetic biology parts in bacterial cells. We wanted to learn if BCDs worked equally well for in vitro protein production during cell‐free protein synthesis (CFPS). CFPS employs transcriptional and translational machinery extracted from cells producing the viral T7 RNA polymerase. We built GFP expression cassettes using 19 BCDs and tested them with CFPS. We compared the rank order of the 19 BCDs in terms of in vitro CFPS GFP production with the in vivo results published by the Endy group. We calculated a rank correlation of 0.88 between the in vitro and in vivo data, showing that BCD control of protein production is comparable in both contexts. The significance of our work is that synthetic biologists using CFPS are well advised to employ BCDs so they can produce predictable amounts of protein for widespread applications in vitro.Support or Funding InformationNSF RUI MCB‐1613281 to Missouri Western State University; NSF RUI MCB‐1613203 to Davidson College
Our undergraduate researchers work in the field of metabolic engineering to harness metabolic pathways for the production of useful metabolites with applications in the biofuel, pharmaceutical, or industrial chemical industries. We published a metabolic engineering approach called programmed evolution that introduces variation in regulatory elements for the expression of enzymes that control orthogonal metabolic pathways. Programmed evolution employs a riboswitch‐based fitness module to transduce the production of a desired metabolite into antibiotic resistance. We validated programmed evolution by optimizing the conversion of caffeine to theophylline by caffeine demethylase, but we recognized several limitations to the widespread use of our approach, including crosstalk between orthogonal and native bacterial metabolism, potential toxicity of metabolites, and poor representation of large variation spaces. To avoid these limitations, we are developing an in vitro approach to programmed evolution. For cell‐free protein synthesis (CFPS), we found supercoiled plasmid was a much better template than linear PCR amplicon. We designed, constructed, and have begun testing a novel fitness module that responds to the production of theophylline produced by the action of caffeine demethylase produced by CFPS. The fitness module encodes a fusion protein composed of three functional units. The amino portion contains a Gal4 DNA binding domain that will target binding sites on a library of DNA regulatory elements for the expression of caffeine demethylase. The carboxyl terminus includes streptavidin which will bind to avidin beads for physical separation of selected DNA regulatory elements. Sandwiched in between these two domains is GFP, which serves as a biosensor. Using this in vitro selection strategy, we plan to pursue the optimization of CFPS for the discovery of efficient metabolic pathways of interest.Support or Funding InformationNSF RUI grant MCB‐1613203 to Davidson College and MCB‐1613281 to Missouri Western State University, Davidson College Martin Genomics ProgramThis abstract is from the Experimental Biology 2019 Meeting. There is no full text article associated with this abstract published in The FASEB Journal.
Objective The purpose of this project was to use an in vivo method to discover riboswitches that are activated by new ligands. We employed phage-assisted continuous evolution (PACE) to evolve new riboswitches in vivo. We started with one translational riboswitch and one transcriptional riboswitch, both of which were activated by theophylline. We used xanthine as the new target ligand during positive selection followed by negative selection using theophylline. The goal was to generate very large M13 phage populations that contained unknown mutations, some of which would result in new aptamer specificity. We discovered side products of three new theophylline translational riboswitches with different levels of protein production. Results We used next generation sequencing to identify M13 phage that carried riboswitch mutations. We cloned and characterized the most abundant riboswitch mutants and discovered three variants that produce different levels of translational output while retaining their theophylline specificity. Although we were unable to demonstrate evolution of new riboswitch ligand specificity using PACE, we recommend careful design of recombinant M13 phage to avoid evolution of “cheaters” that short circuit the intended selection pressure.
rClone Red is a low-cost and student-friendly research tool that has been used successfully in undergraduate teaching laboratories. It enables students to perform original research within the financial and time constraints of a typical undergraduate environment. Students can strengthen their understanding of the initiation of bacterial translation by cloning ribosomal binding sites of their own design and using a red fluorescent protein reporter to measure translation efficiency. Online microbial genome sequences and the mFold website enable students to explore homologous rRNA gene sequences and RNA folding, respectively. In this report, we described how students in a genetics course who were given the opportunity to use rClone Red demonstrated significant learning gains on 16 of 20 concepts, and made original discoveries about the function of ribosome binding sites. By combining the highly successful cloning method of golden gate assembly with the dual reporter proteins of green fluorescent protein and red fluorescent protein, rClone Red enables novice undergraduates to make new discoveries about the mechanisms of translational initiation, while learning the core concepts of genetic information flow in bacteria.
The premise of biological modularity is an ontological claim that appears to come out of practice.We understand that the biological world is modular because we can manipulate different parts of organisms in ways that would only work if there were discrete parts that were interchangeable.This is the foundation of the BioBrick assembly method widely used in synthetic biology.It is one of a number of methods that allows practitioners to construct and reconstruct biological pathways and devices using DNA libraries of standardized parts with known functions.In this paper, we investigate how the practice of synthetic biology reconfigures biological understanding of the key concepts of modularity and evolvability.We illustrate how this practice approach takes engineering knowledge and uses it to try to understand biological organization by showing how the construction of functional parts and processes can be used in synthetic experimental evolution.We introduce a new approach within synthetic biology that uses the premise of a parts-based ontology together with that of organismal self-organization to optimize orthogonal metabolic pathways in E. coli.We then use this and other examples to help characterize semisynthetic categories of modularity, parthood, and evolvability within the discipline.
Integration of research experience into classroom is an important and vital experience for all undergraduates. These course‐based undergraduate research experiences (CUREs) have grown from independent instructor lead projects to large consortium driven experiences. The impact and importance of CUREs on students at all levels in biochemistry was the focus of a National Science Foundation funded think tank. The state of biochemistry CUREs and suggestions for moving biochemistry forward as well as a practical guide (supplementary material) are reported here. © 2016 by The International Union of Biochemistry and Molecular Biology, 45(1):7–12, 2017.
In this paper, we introduce two special classes of digraphs.A limited outdegree grid (LOG) directed graph is a digraph derived from an n × n grid graph by removing some edges and replacing some edges with arcs such that no vertex has outdegree greater than 1.A greatest increase grid (GIG) directed graph is a LOG digraph whose vertices can be labeled with distinct labels such that each arc represents the direction of greatest increase in the underlying grid graph.We enumerate both GIG and LOG digraphs for the 3×3 case.
Students often memorize the definition of a transcriptional promoter but fail to fully understand the critical role promoters play in gene expression. This laboratory lesson allows students to conduct original research by identifying and characterizing promoters found in prokaryotes. Students start with primary literature, design and clone a short promoter, and test how well their promoter works. This laboratory lesson is an easy way for faculty with limited time and budgets to give their students access to real research in the context of traditional teaching labs that meet once a week for under three hours. The pClone Red Introductory Biology lesson uses synthetic biology methods and makes cloning so simple that we have 100% success rates with first year students. Students use a database to archive their promoter sequences and the performance of the promoter under standard conditions. The database permits synthetic biology researchers around the world to find a promoter that suits their needs and compare relative levels of transcription. The core methodology in this lesson is identical to the core methodology in the companion Genetics Lesson by Eckdahl and Campbell. The methods are reproduced in both lessons for the benefit of readers. The two CourseSource lessons provide the detailed information needed to reproduce the pedagogical research results published in CBE - Life Sciences Education by Campbell et al., 2014.
The Muller F element (4.2 Mb, ~80 protein-coding genes) is an unusual autosome of Drosophila melanogaster; it is mostly heterochromatic with a low recombination rate. To investigate how these properties impact the evolution of repeats and genes, we manually improved the sequence and annotated the genes on the D. erecta, D. mojavensis, and D. grimshawi F elements and euchromatic domains from the Muller D element. We find that F elements have greater transposon density (25–50%) than euchromatic reference regions (3–11%). Among the F elements, D. grimshawi has the lowest transposon density (particularly DINE-1: 2% vs. 11–27%). F element genes have larger coding spans, more coding exons, larger introns, and lower codon bias. Comparison of the Effective Number of Codons with the Codon Adaptation Index shows that, in contrast to the other species, codon bias in D. grimshawi F element genes can be attributed primarily to selection instead of mutational biases, suggesting that density and types of transposons affect the degree of local heterochromatin formation. F element genes have lower estimated DNA melting temperatures than D element genes, potentially facilitating transcription through heterochromatin. Most F element genes (~90%) have remained on that element, but the F element has smaller syntenic blocks than genome averages (3.4–3.6 vs. 8.4–8.8 genes per block), indicating greater rates of inversion despite lower rates of recombination. Overall, the F element has maintained characteristics that are distinct from other autosomes in the Drosophila lineage, illuminating the constraints imposed by a heterochromatic milieu.
Current use of microbes for metabolic engineering suffers from loss of metabolic output due to natural selection. Rather than combat the evolution of bacterial populations, we chose to embrace what makes biological engineering unique among engineering fields - evolving materials. We harnessed bacteria to compute solutions to the biological problem of metabolic pathway optimization. Our approach is called Programmed Evolution to capture two concepts. First, a population of cells is programmed with DNA code to enable it to compute solutions to a chosen optimization problem. As analog computers, bacteria process known and unknown inputs and direct the output of their biochemical hardware. Second, the system employs the evolution of bacteria toward an optimal metabolic solution by imposing fitness defined by metabolic output. The current study is a proof-of-concept for Programmed Evolution applied to the optimization of a metabolic pathway for the conversion of caffeine to theophylline in E. coli. Introduced genotype variations included strength of the promoter and ribosome binding site, plasmid copy number, and chaperone proteins. We constructed 24 strains using all combinations of the genetic variables. We used a theophylline riboswitch and a tetracycline resistance gene to link theophylline production to fitness. After subjecting the mixed population to selection, we measured a change in the distribution of genotypes in the population and an increased conversion of caffeine to theophylline among the most fit strains, demonstrating Programmed Evolution. Programmed Evolution inverts the standard paradigm in metabolic engineering by harnessing evolution instead of fighting it. Our modular system enables researchers to program bacteria and use evolution to determine the combination of genetic control elements that optimizes catabolic or anabolic output and to maintain it in a population of cells. Programmed Evolution could be used for applications in energy, pharmaceuticals, chemical commodities, biomining, and bioremediation.
Students often memorize the definition of a transcriptional promoter but fail to fully understand the critical role promoters play in gene expression. This laboratory lesson allows students to conduct original research by characterizing functional regions within known prokaryotic promoters. Students begin the lesson by learning the properties of transcriptional promoter DNA sequences. They design mutations for a constitutive promoter and discuss their designs as a class to choose which mutations to clone and characterize. This lesson provides an easy way for faculty with limited time and budgets to give their students access to real research in the context of traditional teaching labs that meet once a week for under three hours. The pClone Red Genetics lesson uses synthetic biology methods and makes cloning so simple that we have 100% success rates with sophomores taking Genetics. Students archive promoter sequences and their performances under standard conditions. The database permits synthetic biology researchers around the world to find a promoter that suits their needs and compare relative levels of transcription. The core methodology in this lesson is identical to the core methodology in the companion Introductory Biology Lesson by Campbell and Eckdahl. The methods are reproduced in both lessons for the benefit of readers. The two CourseSource lessons provide the detailed information needed to reproduce the pedagogical research results published in CBE – Life Sciences Education by Campbell et al., 2014.
We are developing Programmed Evolution as a system to optimize metabolic pathways in E. coli. The system allows introduction of variation in the regulatory elements of a genetic circuit encoding enzymes that control the desired orthogonal metabolism. Bacteria with variations that contribute best to metabolic output are selected through a fitness module. We reasoned that variation in plasmid copy number (PCN) of vectors containing the genetic circuit would affect metabolic output. We designed and cloned mutations in the pMB1 origin of replication to produce a change in PCN. We mutated the promoter regions of the RNA I and RNA II genes. PCN was measured by the use of Real Time Quantitative Polymerase Chain Reaction (RT‐qPCR) by comparison of five serial dilutions of chromosomal and plasmid DNA. Analysis of Ct values of plasmid and chromosomal DNA suggested altered PCN in some of the mutants. PCN measurements from RT‐qPCR were corroborated with measurements of DNA yield from plasmid preps of bacteria carrying the mutated origins. The results provide a collection of origins with several PCNs that can be used to introduce variation that is expected to affect metabolic output during Programmed Evolution.
There is widespread agreement that science, technology, engineering, and mathematics programs should provide undergraduates with research experience. Practical issues and limited resources, however, make this a challenge. We have developed a bioinformatics project that provides a course-based research experience for students at a diverse group of schools and offers the opportunity to tailor this experience to local curriculum and institution-specific student needs. We assessed both attitude and knowledge gains, looking for insights into how students respond given this wide range of curricular and institutional variables. While different approaches all appear to result in learning gains, we find that a significant investment of course time is required to enable students to show gains commensurate to a summer research experience. An alumni survey revealed that time spent on a research project is also a significant factor in the value former students assign to the experience one or more years later. We conclude: 1) implementation of a bioinformatics project within the biology curriculum provides a mechanism for successfully engaging large numbers of students in undergraduate research; 2) benefits to students are achievable at a wide variety of academic institutions; and 3) successful implementation of course-based research experiences requires significant investment of instructional time for students to gain full benefit.
The Vision and Change report recommended genuine research experiences for undergraduate biology students. Authentic research improves science education, increases the number of scientifically literate citizens, and encourages students to pursue research. Synthetic biology is well suited for undergraduate research and is a growing area of science. We developed a laboratory module called pClone that empowers students to use advances in molecular cloning methods to discover new promoters for use by synthetic biologists. Our educational goals are consistent with Vision and Change and emphasize core concepts and competencies. pClone is a family of three plasmids that students use to clone a new transcriptional promoter or mutate a canonical promoter and measure promoter activity in Escherichia coli. We also developed the Registry of Functional Promoters, an open-access database of student promoter research results. Using pre- and posttests, we measured significant learning gains among students using pClone in introductory biology and genetics classes. Student posttest scores were significantly better than scores of students who did not use pClone. pClone is an easy and affordable mechanism for large-enrollment labs to meet the high standards of Vision and Change.