The Synthetic Biology Open Language (SBOL) is a community-driven and machine-readable standard for creating and sharing biological designs. SBOL has evolved over the years, resulting in three major versions, each introducing new features. Although SBOL3 is the latest community-adopted version, it has not been fully integrated into software tools and repositories due to the lack of essential libraries to facilitate this transition. In this paper, we present the SBOL Converter, a backward-compatibility tool that can update legacy designs to the latest SBOL version. Designs can also be converted to previous versions to allow legacy tools to be used in ongoing projects. Moreover, it incorporates existing SBOL tools, providing a unified interface to orchestrate the validation of designs and checking against various compliance rules and best practices. The converter can also process other popular biological formats such as GenBank and FASTA. As a lightweight tool, it can be integrated into engineering biology design workflows to accelerate tool development and enable the transition to the latest versions as they are released, facilitating the reusability and reproducibility of biological designs.
Abstract Large language models have transformed software engineering practices. However, generated artefacts are not always developer-friendly and may partially meet complex requirements. As the need to standardise, integrate, and develop tools in engineering biology increases, novel approaches are needed to create and maintain intuitive software sustainably. Here, we present an ontology-driven approach using large language models to create user-facing software libraries for knowledge graphs. We introduce an ontology-to-language framework to systematically map domain terms and graph structures. We then demonstrate this approach by creating an ontology for the latest Synthetic Biology Open Language standard and generating the sbol-script software library, which can be used within browsers or to develop applications with native web support. This ontology-driven software engineering approach and these resources are essential for the community and to facilitate the development of sustainable software projects. The SBOL3 Ontology and the sbol-script library are available from https://github.com/SynBioDex/sbol-owl3 and https://github.com/SynBioDex/sbol-script .
Synthetic nucleic acids are a key input to modern biotechnology, yet they represent dual-use materials that require robust screening to mitigate biosecurity risks. The prevailing screening paradigm, which identifies sequences of concern (SoCs) through sequence similarity to controlled pathogens and toxins, may not fully capture risks posed by AI tools that can decouple biomolecular function from reliance on known sequences. Rapidly advancing biodesign capabilities enable the generation of genes and proteins that might evade sequence-based detection. We highlight the critical need for function-based screening approaches that can detect sequences capable of hazardous biological functions, regardless of similarity to known SoCs. We examine the feasibility of function-based screening with an initial focus on proteins, arguing that, while protein sequence space is vast, biologically functional proteins are significantly constrained by biophysical and biochemical requirements that can be learned and modeled. We propose a concrete implementation framework organized along a continuum of complexity, starting with toxins as the most tractable targets before expanding to more complex pathogenic functions. We then discuss open challenges and describe a research and development strategy to address them.
Readily available nucleic acid synthesis is both critical for the bioeconomy and an increasingly pressing security concern due to the potential for accidental or deliberate misuse. While biosecurity experts broadly agree that nucleic acid providers should screen orders for potential “sequences of concern,” there has previously been no agreed standard for how to define and recognize such sequences. To address this gap, we first organized a collection of test sets containing 1.1 million sequences from pathogens and toxins on the Australia Group Common Control Lists and their non-controlled relatives, along with model organisms and synthetic constructs. An initial categorization of sequences as to whether or not they were sequences of concern was produced by comparing the results of four biosecurity screening systems for each of these sequences, finding that these systems already agreed on the categorization of more than 80% of sequences. We then refined these results through a science-based stakeholder review process to define a rubric for determining whether a sequence should be flagged as a potential sequence of concern, then applied this rubric to improve the categorization of sequences in test sets. The result is a rubric that identifies sequences of concern with respect to human pandemic-potential viruses, key classes of low-risk genes, and controlled toxins. Applying this rubric to the test set collection has, to date, reduced the number of test sequences with disputed categorization by 44.3% for controlled viruses and 10.7% across the collection of test sets as a whole. Together, the rubric and the test sets provide a concrete “sequence of concern” definition that can be used as a foundation for development of biosecurity screening standards and policy and is also continuing to be refined in ongoing work.
Rapid advancements in AI have enabled significant progress in protein and nucleic acid design, but they also pose biosecurity challenges. We examine the vulnerabilities of biosecurity screening software (BSS) to AI-reformulated synthetic homologs of proteins of concern (POCs) that have been fragmented into smaller segments. We evaluate four BSS tools that were recently patched to enhance their AI resiliency. Without any further modification, we found that two of the four tools were capable of robustly detecting fragments as short as 50 nucleotides, demonstrating screening capabilities that exceed those requested in the United States Framework for Nucleic Acid Synthesis. Upgraded versions of the other two tools improved performance. Although our findings confirm the effectiveness of the tested BSS tools, at the same time, they emphasize the urgency of developing alternate BSS approaches to counter evolving AI-enabled biosecurity risks.
Standards play a crucial role in ensuring consistency, interoperability, and efficiency of communication across various disciplines. In the field of synthetic biology, the Synthetic Biology Open Language (SBOL) Visual standard was introduced in 2013 to establish a structured framework for visually representing genetic designs. Over the past decade, SBOL Visual has evolved from a simple set of 21 glyphs into a comprehensive diagrammatic language for biological designs. This perspective reflects on the first ten years of SBOL Visual, tracing its evolution from inception to version 3.0. We examine the standard's adoption over time, highlighting its growing use in scientific publications, the development of supporting visualization tools, and ongoing efforts to enhance clarity and accessibility in communicating genetic design information. While trends in adoption show steady increases, achieving full compliance and use of best practices will require additional efforts. Looking ahead, the continued refinement of SBOL Visual and broader community engagement will be essential to ensuring its long-term value as the field of synthetic biology develops.
Plant synthetic biologists have been working to adapt the CRISPRa and CRISPRi promoter regulation methods for applications such as improving crops or installing other valuable pathways. With other organisms, strong transcriptional control has typically required multiple gRNA target sites, which poses a critical engineering choice between heterogeneous sites, which allow each gRNA to target existing locations in a promoter, and identical sites, which typically require modification of the promoter. Here, we investigate the consequences of this choice for CRISPRi plant promoter regulation via simulation-based analysis, using model parameters based on single gRNA regulation and constitutive promoters in Nicotiana benthamiana and Arabidopsis thaliana. Using models of 2-6 gRNA target sites to compare heterogeneous versus identical sites for tunability, sensitivity to parameter values, and sensitivity to cell-to-cell variation, we find that identical gRNA target sites are predicted to yield far more effective transcriptional repression than heterogeneous sites.
Introduction: Nucleic acid synthesis is a dual-use technology that can benefit fields such as biology, medicine, and information storage. However, synthetic nucleic acids could also potentially be used negligently and ultimately cause harm, or be used with malicious intent to cause harm. Thus, this technology needs to be appropriately safeguarded. Sequence screening is one component of a biosecurity protocol for preventing such harm and consists of identifying Sequences of Concern (SOCs). There exist many fit-for-purpose tools that have been developed for nucleic acid synthesis sequence screening. However, questions remain regarding their performance with respect to the consistency of screening. Methods: To aid in determining if screening tools are harmonized in regard to baseline sequence screening (which represents a minimum acceptable level of performance), the National Institute of Standards and Technology (NIST) constructed a test dataset based on current screening recommendations. NIST then sent blinded datasets to sequence screening tool developers for testing. Results: Overall, there was a general agreement between the tools and NIST labels given to the sequences, and all tools had a baseline performance of >95% sensitivity and >97% accuracy. Disagreement on specific sequences largely arose from single tools and could be traced to differences in defining a SOC and/or methodological differences in screening algorithms.
Flow cytometry is a powerful quantitative assay supporting high-throughput collection of single-cell data with a high dynamic range. For flow cytometry to yield reproducible data with a quantitative relationship to the underlying biology, however, requires that 1) appropriate process controls are collected along with experimental samples, 2) these process controls are used for unit calibration and quality control, and 3) data is analyzed using appropriate statistics. To this end, this article describes methods for quantitative flow cytometry through addition of process controls and analyses, thereby enabling better development, modeling, and debugging of engineered biological organisms. The methods described here have specifically been developed in the context of transient transfections in mammalian cells, but may in many cases be adaptable to other categories of transfection and other types of cells.
Fast-moving advances in AI-assisted protein engineering are enabling breakthroughs in the life sciences that promise numerous beneficial applications. At the same time, these new capabilities are creating potential biosecurity challenges by providing new pathways to intentional or accidental synthesis of genes that encode hazardous proteins. The synthesis of nucleic acids is a key choke point in the AI-assisted protein engineering pipeline as it is where digital designs are transformed into physical instructions that can produce potentially harmful proteins. Thus, one focus for efforts to enhance biosecurity in the face of new AI-enabled capabilities is on bolstering the screening of orders by nucleic acid synthesis providers. We describe a multistakeholder, cross-sector effort to address biosecurity challenges with uses of AI-powered biological design tools to reformulate naturally occurring proteins of concern to create synthetic homologs that have low sequence identity to the wild-type proteins. We evaluated the abilities of traditional nucleic acid biosecurity screening tools to detect these synthetic homologs and found that, of tools tested, not all could previously detect such AI-redesigned sequences reliably. However, as we report, patches were built and deployed to improve detection rates over the course of the project, resulting in a final mean detection rate over tools of 97% of the synthetic homologs that were determined, using in-silico metrics, to be more likely to retain wild-type-like function. Finally, we make recommendations on approaches for studying and addressing the rising risk of adversarial AI-assisted protein engineering attacks like the one we identified and worked to mitigate. ### Competing Interest Statement B.J.W and E.H. are based at an organization that is engaged with research, development, and fielding of AI technologies, including AI-assisted protein engineering technologies. A.C and J.D. are based at DNA synthesis companies. T.A., C.B., J.B., K.F., B.G., T.M., S.T.M., and N.W. are affiliated with institutions that build and deploy biosecurity screening software.
Synthetic biology is an interdisciplinary field that brings together engineering and biology concepts alongside the arts and social sciences to develop solutions to pressing problems in our world. The education of students entering this field has relied on a diverse set of pedagogical methods to accomplish this goal. One non-profit group, iGEM-the International Genetically Engineered Machine competition, has been a driver of students' awareness of synthetic biology for the last 20 years giving many young researchers their first experience in the field of synthetic biology. Dissemination of synthetic biology concepts by iGEM has occurred through several programs including a webinar series started during the 2020 COVID pandemic. The iGEM webinar series successfully engaged students by taking inspiration from synthetic biology programs in Europe, North America, and Asia that had themselves evolved alongside iGEM. The webinar designers modeled the content after their experiences in iGEM as well as their academic courses, pedagogy, and mentoring experiences. This series has produced globally accessible pedagogy for both technical synthetic biology knowledge and the communication skills necessary to build and communicate synthetic biology projects. The hope is that this series functions as a lasting blueprint that can be used by future educators in synthetic biology and other disciplines to reduce barriers that students face when attempting to enter cutting edge fields.
Objective: DNA synthesis companies screen orders to detect controlled sequences with misuse risks. Assessing screening accuracy is challenging owing to the breadth of biological risks and ambiguities in risk definitions. Here, we detail an International Gene Synthesis Consortium working group's rationale and process to develop a prototype DNA synthesis screening test dataset, aiming to establish a baseline of screening system accuracy to compare with various screening approaches.Methodology: Construction of the prototype test dataset involved four tool developers screening nucleic acid sequences from three taxonomic clusters of controlled organisms (Orbivirus, Francisella tularensis, and Coccidioides). Results were mapped onto predefined, comparable categories, checking for consensus or conflicts. Conflicts were grouped based on gene annotation and resolved through discussion.Results: The process highlighted several long-standing challenges in DNA synthesis screening, including the qualitative differences in approaches taken by screening tools. Our findings highlight the lack of clarity in assessing pathogen sequences with respect to regulatory control language, compounded by scientific uncertainty. We illustrate the current degree of consensus and existing challenges using classification statistics and specific examples.Conclusions and Next Steps: This prototype underscores the necessity of expert-regulator coordination in assessing gene-associated risks, offering a template for creating test sets across all taxonomic groups on international control lists. Expanding the working group would enrich dataset comprehensiveness, enabling a transition from species-focused to function-focused regulatory controls. This sets the foundation for quality control, certification, and improved risk assessment in DNA synthesis screening.
Laboratory protocols are critical to biological research and development, yet difficult to communicate and reproduce across projects, investigators, and organizations. While many attempts have been made to address this challenge, there is currently no available protocol representation that is unambiguous enough for precise interpretation and automation, yet simultaneously “human friendly” and abstract enough to enable reuse and adaptation. The Laboratory Open Protocol language (LabOP) is a free and open protocol representation aiming to address this gap, building on a foundation of UML, Autoprotocol, Aquarium, SBOL RDF, and the Provenance Ontology. LabOP provides a linked-data representation both for protocols and for records of their execution and the resulting data, as well as a framework for exporting from LabOP for execution by either humans or laboratory automation. LabOP is currently implemented in the form of an RDF knowledge representation, specification document, and Python library, and supports execution as manual “paper protocols,” by Autoprotocol or by Opentrons. From this initial implementation, LabOP is being further developed as an open community effort.
Spreading informationthrough a network of devices is a core activity for most distributed systems. Self-stabilizing algorithms for information spreading are one of the key building blocks enabling aggregate computing to provide resilient coordination in open complex distributed systems. This article improves a general spreading block in the aggregate computing literature by making it resilient to network perturbations, establishes its global uniform asymptotic stability, and proves that it is ultimately bounded under persistent disturbances. The ultimate bounds depend only on the magnitude of the largest perturbation and the network diameter, and three design parameters trading off competing aspects of performance. For example, as in many dynamical systems, values leading to greater resilience to network perturbations slow convergence and vice versa.
Synthetic biologists have made great progress over the past decade in developing methods for modular assembly of genetic sequences and in engineering biological systems with a wide variety of functions in various contexts and organisms. However, current paradigms in the field entangle sequence and functionality in a manner that makes abstraction difficult, reduces engineering flexibility and impairs predictability and design reuse. Functional Synthetic Biology aims to overcome these impediments by focusing the design of biological systems on function, rather than on sequence. This reorientation will decouple the engineering of biological devices from the specifics of how those devices are put to use, requiring both conceptual and organizational change, as well as supporting software tooling. Realizing this vision of Functional Synthetic Biology will allow more flexibility in how devices are used, more opportunity for reuse of devices and data, improvements in predictability and reductions in technical risk and cost.
Synthetic biology builds upon genetics, molecular biology, and metabolic engineering by applying engineering principles to the design of biological systems. When designing a synthetic system, synthetic biologists need to exchange information about multiple types of molecules, the intended behavior of the system, and actual experimental measurements. The Synthetic Biology Open Language (SBOL) has been developed as a standard to support the specification and exchange of biological design information in synthetic biology, following an open community process involving both wet bench scientists and dry scientific modelers and software developers, across academia, industry, and other institutions. This document describes SBOL 3.0.0, which condenses and simplifies previous versions of SBOL based on experiences in deployment across a variety of scientific and industrial settings. In particular, SBOL 3.0.0, (1) separates sequence features from part/sub-part relationships, (2) renames Component Definition/Component to Component/Sub-Component, (3) merges Component and Module classes, (4) ensures consistency between data model and ontology terms, (5) extends the means to define and reference Sub-Components, (6) refines requirements on object URIs, (7) enables graph-based serialization, (8) moves Systems Biology Ontology (SBO) for Component types, (9) makes all sequence associations explicit, (10) makes interfaces explicit, (11) generalizes Sequence Constraints into a general structural Constraint class, and (12) expands the set of allowed constraints.
Standards support synthetic biology research by enabling the exchange of component information. However, using formal representations, such as the Synthetic Biology Open Language (SBOL), typically requires either a thorough understanding of these standards or a suite of tools developed in concurrence with the ontologies. Since these tools may be a barrier for use by many practitioners, the Excel-SBOL Converter was developed to facilitate the use of SBOL and integration into existing workflows. The converter consists of two Python libraries: one that converts Excel templates to SBOL and another that converts SBOL to an Excel workbook. Both libraries can be used either directly or via a SynBioHub plugin.
Abstract Synthetic biology builds upon genetics, molecular biology, and metabolic engineering by applying engineering principles to the design of biological systems. When designing a synthetic system, synthetic biologists need to exchange information about multiple types of molecules, the intended behavior of the system, and actual experimental measurements. The Synthetic Biology Open Language (SBOL) has been developed as a standard to support the specification and exchange of biological design information in synthetic biology, following an open community process involving both bench scientists and scientific modelers and software developers, across academia, industry, and other institutions. This document describes SBOL 3.1.0, which improves on version 3.0.0 by including a number of corrections and clarifications as well as several other updates and enhancements. First, this version includes a complete set of validation rules for checking whether documents are valid SBOL 3. Second, the best practices section has been moved to an online repository that allows for more rapid and interactive of sharing these conventions. Third, it includes updates based upon six community approved enhancement proposals. Two enhancement proposals are related to the representation of an object’s namespace. In particular, the Namespace class has been removed and replaced with a namespace property on each class. Another enhancement is the generalization of the CombinatorialDeriviation class to allow direct use of Features and Measures. Next, the Participation class now allow Interactions to be participants to describe higher-order interactions. Another change is the use of Sequence Ontology terms for Feature orientation. Finally, this version of SBOL has generalized from using Unique Reference Identifiers (URIs) to Internationalized Resource Identifiers (IRIs) to support international character sets.
As synthetic biology becomes increasingly capable and accessible, it is likewise increasingly critical to be able to make accurate biosecurity determinations regarding the pathogenicity or toxicity of particular nucleic acid or amino acid sequences. At present, this is typically done using the BLAST algorithm to determine the best match with sequences in the NCBI nucleic acid and protein databases. Neither BLAST nor any of the NCBI databases, however, are actually designed for biosafety determination. Critically, taxonomic errors or ambiguities in the NCBI nucleic acid and protein databases can also cause errors in BLAST-based taxonomic categorization. With heavily studied taxa and frequently used biotechnology tools, even low frequency taxonomic categorization issues can lead to high rates of errors in biosecurity decision-making. Here we focus on the implications for false positives, finding that BLAST against NCBI's protein database will now incorrectly categorize a number of commonly used biotechnology tool sequences as the pathogens or toxins with which they have been used. Paradoxically, this implies that problems are expected to be most acute for the pathogens and toxins of highest interest and for the most widely used biotechnology tools. We thus conclude that biosecurity tools should shift away from BLAST against general purpose databases and towards new methods that are specifically tailored for biosafety purposes.
The design and construction of genetic systems, in silico, in vitro, or in vivo, often involve the handling of various pieces of DNA that exist in different forms across an assembly process: as a standalone "part" sequence, as an insert into a carrier vector, as a digested fragment, etc. Communication about these different forms of a part and their relationships is often confusing, however, because of a lack of standardized terms. Here, we present a systematic terminology and an associated set of practices for representing genetic parts at various stages of design, synthesis, and assembly. These practices are intended to represent any of the wide array of approaches based on embedding parts in carrier vectors, such as BioBricks or Type IIS methods (e.g., GoldenGate, MoClo, GoldenBraid, and PhytoBricks), and have been successfully used as a basis for cross-institutional coordination and software tooling in the iGEM Engineering Committee.