Nanobodies are a class of small, monomeric camelid antibody fragments that can bind target antigens with high affinity and specificity. Their small size, structural simplicity, and limited reliance on disulfide bonding makes them attractive for intracellular expression for labeling and perturbing cellular processes in live cells. However, screening campaigns carried out exclusively in vitro often yield antigen binders that fail to perform well in live cells due to low expression, misfolding, or mistargeting. We demonstrate that traditional in vitro screening of a nanobody library combined with an intracellular bioluminescence resonant energy transfer (BRET) proximity sensor approach for sequence down-selection can yield strong in vitro binders that also perform well as intrabodies, in this case capable of binding to, and inhibiting the enzymatic activity of, ITCH E3 ubiquitin ligase in human cells. This strategy allows a more direct and scalable path toward intrabody discovery.
Motivation:Generating phylogenomic trees from the genomic data is essential in understanding biological systems. Each step of this complex process has received extensive attention and has been significantly streamlined over the years. Given the public availability of data, obtaining genomes for a wide selection of species is straightforward. However, analyzing that data to generate a phylogenomic tree is a multistep process with legitimate scientific and technical challenges, often requiring a significant input from a domain-area scientist. Results:We present Poplar, a new, streamlined computational pipeline, to address the computational logistical issues that arise when constructing the phylogenomic trees. It provides a framework that runs state-of-the-art software for essential steps in the phylogenomic pipeline, beginning from a genome with or without an annotation, and resulting in a species tree. Running Poplar requires no external databases. In the execution, it enables parallelism for execution for clusters and cloud computing. The trees generated by Poplar match closely with state-of-the-art published trees. The usage and performance of Poplar is far simpler and quicker than manually running a phylogenomic pipeline. Availability and implementation:Freely available on GitHub at https://github.com/sandialabs/poplar. Implemented using Python and supported on Linux.
Understanding and controlling gene expression in organisms is essential for optimizing biological processes, whether in service of bioeconomic processes, human health, or environmental regulation. Epigenetic modifications play a significant role in regulating gene expression by altering chromatin structure, DNA accessibility and protein binding. While a significant amount is known about the combinatorial effects of epigenetics on gene expression, our understanding of the degree to which the orchestration of these mechanisms is conserved in gene expression regulation across species, particularly for non-model organisms, remains limited. In this study, we aim to predict gene expression levels based on epigenetic modifications in chromatin across different fungal species, to enable transferring information about well characterized species to poorly understood species. We developed a custom hybrid deep learning model, EAGLE (Evolutionary distance-Adaptable Gene expression Learned from Epigenomics), which combines convolutional layers and multi-head attention mechanisms to capture both local and global dependencies in epigenetic data. We demonstrate the cross-species performance of EAGLE across fungi, a kingdom containing both pathogens and biomanufacturing chassis and where understanding epigenetic regulation in under-characterized species would be transformative for bioeconomic, environmental, and biomedical applications. EAGLE outperformed shallow learning models and a modified transformer benchmarking model, achieving up to 80% accuracy and 89% AUROC for intra-species validation and 77% accuracy and 83% AUROC in cross-species prediction tasks. SHAP analysis revealed that EAGLE identifies important epigenetic features that drive gene expression, providing insights for experimental design and potential future epigenome engineering work. Our findings demonstrate the potential of EAGLE to generalize across fungal species, offering a versatile tool for optimizing fungal gene expression in multiple sectors. In addition, our architecture can be adapted for cross-species tasks across the tree of life where detailed molecular and genetic information can be scarce. ### Competing Interest Statement The authors have declared no competing interest.
While CRISPRi was previously established in Synechococcus sp. PCC 7002 (hereafter 7002), the design principles for guide RNA (gRNA) effectiveness remain largely unknown. Here, 76 strains of 7002 were constructed with gRNAs targeting three reporter systems to evaluate features that impact gRNA efficiency. Correlation analysis of the data revealed that important features of gRNA design include the position relative to the start codon, GC content, protospacer adjacent motif (PAM) site, minimum free energy, and targeted DNA strand. Unexpectedly, some gRNAs targeting upstream of the promoter region showed small but significant increases in reporter expression, and gRNAs targeting the terminator region showed greater repression than gRNAs targeting the 3 ' end of the coding sequence. Machine learning algorithms enabled prediction of gRNA effectiveness, with Random Forest having the best performance across all training sets. This study demonstrates that high-density gRNA data and machine learning can improve gRNA design for tuning gene expression in 7002.
Immune checkpoint immunotherapy (ICI) can re-activate immune reactions against neoantigens, leading to remarkable remission in cancer patients. Nevertheless, only a minority of patients are responsive to ICI, and approaches for prediction of responsiveness are needed to improve the success of cancer treatments. While the tumor mutational burden (TMB) correlates positively with responsiveness and survival of patients undergoing ICI, the influence of the subcellular localizations of the neoantigens remains unclear. Here, we demonstrate in both a mouse melanoma model and human clinical datasets of 1,722 ICI-treated patients that a high proportion of membrane-localized neoantigens, particularly at the plasma membrane, correlate with responsiveness to ICI therapy and improved overall survival across multiple cancer types. We further show that combining membrane localization and TMB analyses can enhance the predictability of cancer patient response to ICI. Our results may have important implications for establishing future clinical guidelines to direct the choice of treatment toward ICI.
Emerging concepts for neuromorphic computing, bioelectronics, and brain-computer interfacing inspire new research avenues aimed at understanding the relationship between oxidation state and conductivity in unexplored materials. This report expands the materials playground for neuromorphic devices to include a mixed valence inorganic 3D coordination framework, a ruthenium Prussian blue analog (RuPBA), for flexible and biocompatible artificial synapses that reversibly switch conductance by more than four orders of magnitude based on electrochemically tunable oxidation state. The electrochemically tunable degree of mixed valency and electronic coupling between N-coordinated Ru sites controls the carrier concentration and mobility, as supported by density functional theory computations and application of electron transfer theory to in situ spectroscopy of intervalence charge transfer. Retention of programmed states is improved by nearly two orders of magnitude compared to extensively studied organic polymers, thus reducing the frequency, complexity, and energy costs associated with error correction schemes. This report demonstrates dopamine-mediated plasticity of RuPBA synapses and biocompatibility of RuPBA with neuronal cells, evoking prospective application for brain-computer interfacing.
This project was broadly motivated by the need for new hardware that can process information such as images and sounds right at the point of where the information is sensed (e.g. edge computing). The project was further motivated by recent discoveries by g roup demonstrating that while certain organic polymer blends can be used to fabricate elements of such hardware, the need to mix ionic and electronic conducting phases imposed limits on performance, dimensional scalability and the degree of fundamental und erstanding of how such devices operated. As an alternative to blended polymers containing distinct ionic and electronic conducting phases, in this LDRD project we have discovered that a family of mixed valence coordination compounds called Prussian blue an alogue (PBAs), with an open framework structure and ability to conduct both ionic and electronic charge, can be used for inkjet - printed flexible artificial synapses that reversibly switch conductance by more than four orders of magnitude based on electrochemically tunable oxidation state. Retention of programmed states is improved by nearly two orders of magnitude compar ed to the extensively studied organic polymers, thus enabling in - memory compute and avoiding energy costly off - chip access during training. We demonstrate dopamine detection using PBA synapses and biocompatibility with living neurons, evoking prospective a pplication for brain - computer interfacing. By application of electron transfer theory to in - situ spectroscopic probing of intervalence charge transfer, we elucidate a switching mechanism whereby the degree of mixed valency between N - coordinated Ru sites co ntrols the carrier concentration and mobility, as supported by density functional theory (DFT) .
Reversible electrochemical doping of ruthenium hexacyanoruthenate, a type of Prussian blue analogue (PBA), enables the on-demand tuning of electronic conductivity by more than four orders of magnitude. Inkjet-printed electrochemical random access memory (ECRAM) devices based on Ru-PBA and lithium- or proton-conducting ionogel electrolytes exhibit excellent switching efficiency and long-term memory retention, important characteristics for analog artificial synapses in neuromorphic circuits. We also demonstrate excellent biocompatibility with live neurons and the use of Ru-PBA ECRAM devices to detect dopamine, promising first steps toward connecting artificial and biological neural networks. In-situ probing of metal-metal charge transfer by UV/Vis/NIR absorption spectroscopy reveals a switching mechanism whereby electrochemically tunable valence mixing between N-coordinated Ru sites controls the carrier concentration and mobility, as independently supported by both Marcus-Hush electron transfer theory and more conventional band structure predictions from DFT. The experimental agreement achieved by both theoretical approaches supports a general mechanistic picture that intramolecular charge transfer reactions, more commonly studied in polynuclear mixed valence small molecules, are central to electronic conductivity in extended coordination frameworks.
Operon prediction in prokaryotes is critical not only for understanding the regulation of endogenous gene expression, but also for exogenous targeting of genes using newly developed tools such as CRISPR-based gene modulation. A number of methods have used transcriptomics data to predict operons, based on the premise that contiguous genes in an operon will be expressed at similar levels. While promising results have been observed using these methods, most of them do not address uncertainty caused by technical variability between experiments, which is especially relevant when the amount of data available is small. In addition, many existing methods do not provide the flexibility to determine the stringency with which genes should be evaluated for being in an operon pair. We present OperonSEQer, a set of machine learning algorithms that uses the statistic and p-value from a non-parametric analysis of variance test (Kruskal-Wallis) to determine the likelihood that two adjacent genes are expressed from the same RNA molecule. We implement a voting system to allow users to choose the stringency of operon calls depending on whether your priority is high recall or high specificity. In addition, we provide the code so that users can retrain the algorithm and re-establish hyperparameters based on any data they choose, allowing for this method to be expanded as additional data is generated. We show that our approach detects operon pairs that are missed by current methods by comparing our predictions to publicly available long-read sequencing data. OperonSEQer therefore improves on existing methods in terms of accuracy, flexibility, and adaptability.
Mesenchymal stromal cells (MSCs) have broad-ranging therapeutic properties, including the ability to inhibit bacterial growth and resolve infection. However, the genetic mechanisms regulating these antibacterial properties in MSCs are largely unknown. Here, we utilized a systems-based approach to compare MSCs from different genetic backgrounds that displayed differences in antibacterial activity. Although both MSCs satisfied traditional MSC-defining criteria, comparative transcriptomics and quantitative membrane proteomics revealed two unique molecular profiles. The antibacterial MSCs responded rapidly to bacterial lipopolysaccharide (LPS) and had elevated levels of the LPS co-receptor CD14. CRISPR-mediated overexpression of endogenous CD14 in MSCs resulted in faster LPS response and enhanced antibacterial activity. Single-cell RNA sequencing of CD14-upregulated MSCs revealed a shift in transcriptional ground state and a more uniform LPS-induced response. Our results highlight the impact of genetic background on MSC phenotypic diversity and demonstrate that overexpression of CD14 can prime these cells to be more responsive to bacterial challenge.
The goal of this work was to pioneer a novel, low-overhead protocol for simultaneously assaying cell-surface markers and intracellular gene expression in a single mammalian cell. The purpose of developing such a method is to be able to understand the mechanisms by which pathogens engage with individual mammalian cells, depending on their cell surface proteins, and how both host and pathogen gene expression changes are reflective of these mechanisms. The knowledge gained from such analyses of single cells will ultimately lead to more robust pathogen detection and countermeasures. Our method was aimed at streamlining both the upstream cell sample preparation using microfluidic methods, as well as the actual library making protocol. Specifically, we wanted to implement a random hexamer-based reverse transcription of all RNA within a single cell (as opposed to oligo dT-based which would only capture polyadenylated transcripts), and then use a CRISPR-based method called scDash to deplete ribosomal DNAs (since ribosomal RNAs make up the majority of the RNA in a mammalian cell). After significant troubleshooting, we demonstrate that we are able to prepare cDNA from RNA using the random hexamer primer, and perform the rDNA depletion. We also show that we can visualize individually stained cells, setting up the pipeline for connecting surface markers to RNA-sequencing profiles. Finally, we test a number of devices for various parts of the pipeline, including bead generation, optical barcoding and cell dispensing, and demonstrate that while some of these have potential, more work is needed to optimize this part of the pipeline.
This report describes research conducted to use data science and machine learning methods to distinguish targeted genome editing versus natural mutation and sequencer machine noise. Genome editing capabilities have been around for more than 20 years, and the efficiencies of these techniques has improved dramatically in the last 5+ years, notably with the rise of CRISPR-Cas technology. Whether or not a specific genome has been the target of an edit is concern for U.S. national security. The research detailed in this report provides first steps to address this concern. A large amount of data is necessary in our research, thus we invested considerable time collecting and processing it. We use an ensemble of decision tree and deep neural network machine learning methods as well as anomaly detection to detect genome edits given either whole exome or genome DNA reads. The edit detection results we obtained with our algorithms tested against samples held out during training of our methods are significantly better than random guessing, achieving high F1 and recall scores as well as with precision overall.
We describe efforts in generating synthetic malware samples that have specified behaviors that can then be used to train a machine learning (ML) algorithm to detect behaviors in malware. The idea behind detecting behaviors is that a set of core behaviors exists that are often shared in many malware variants and that being able to detect behaviors will improve the detection of novel malware. However, empirically the multi-label task of detecting behaviors is significantly more difficult than malware classification, only achieving on average 84% accuracy across all behaviors as opposed to the greater than 95% multi-class or binary accuracy reported in many malware detection studies. One of the difficulties in identifying behaviors is that while there are ample malware samples, most data sources do not include behavioral labels, which means that generally there is insufficient training data for behavior identification. Inspired by the success of generative models in improving image processing techniques, we examine and extend a 1) conditional variational auto-encoder and 2) a flow-based generative model for malware generation with behavior labels. Initial experiments indicate that synthetic data is able to capture behavioral information and increase the recall of behaviors in novel malware from 32% to 45% without increasing false positives and to 52% with increased false positives.
Mesenchymal stromal cells (MSCs) have broad-ranging therapeutic capabilities, however MSC use is confounded by cell-to-cell heterogeneity, and source-to-source phenotypic inconsistencies. We utilized a systems-based approach to compare MSCs which displayed different capacity for antibacterial activity. Although MSCs from both sources satisfied traditional MSC-defining criteria, comparative transcriptomics and quantitative membrane proteomics demonstrated two unique molecular profiles. The antibacterial MSCs respond rapidly to bacterial lipopolysaccharide (LPS) and have elevated levels of the LPS co-receptor CD14. CRISPR-mediated overexpression of endogenous CD14 in non-antibacterial MSCs resulted in faster LPS response and enhanced antimicrobial activity. Single-cell transcriptome profiling of CD14-activated MSCs revealed uniform enhancement of LPS response kinetics, and a shift in the ground state of these MSCs. Our results demonstrate that systems-level analysis can reveal critical molecular targets to optimize desirable properties in MSCs, and that overexpression of CD14 in these cells can shift their state to be more responsive to future bacterial challenge.
develop cyanobacterial strains that are optimized for growth or metabolite production under a wide range of environmental conditions. The optimization of cyanobacterial strains will directly advance U.S. energy and climate security by enabling domestic biofuel production while simultaneously mitigating atmospheric greenhouse gases through photoautotrophic fixation of carbon dioxide.
Sandia National Laboratories currently has 27 COVID-related Laboratory Directed Research & Development (LDRD) projects focused on helping the nation during the pandemic. These LDRD projects cross many disciplines including bioscience, computing & information sciences, engineering science, materials science, nanodevices & microsystems, and radiation effects & high energy density science.
Integrative genetic elements (IGEs) are mobile multigene DNA units that integrate into and excise from host bacterial chromosomes. Each IGE usually targets a specific site within a conserved host gene, integrating in a manner that preserves target gene function. However, a small number of bacterial genes are known to be inactivated upon IGE integration and reactivated upon excision, regulating phenotypes of virulence, mutation rate, and terminal differentiation in multicellular bacteria. The list of regulated gene integrity (RGI) cases has been slow-growing because IGEs have been challenging to precisely and comprehensively locate in genomes. We present software (TIGER) that maps IGEs with unprecedented precision and without attB site bias. TIGER uses a comparative genomic, ping-pong BLAST approach, based on the principle that the IGE integration module (i.e., its int-attP region) is cohesive. The resultant IGEs, along with integrase phylogenetic analysis and gene inactivation tests, revealed 19 new cases of genes whose integrity is regulated by IGEs (including dut, eccCa1, gntT, hrpB, merA, ompN, prkA, tqsA, traG, yifB, yfaT and ynfE ), as well as recovering previously known cases (in sigK, spsM, comK, mlrA , and hlb genes). It also recovered known clades of site-promiscuous integrases and identified possible new ones.