MOTIVATION:High-throughput sequencing (HTS) is a modern sequencing technology used to profile microbiomes by sequencing thousands of short genomic fragments from the microorganisms within a given sample. This technology presents a unique opportunity for artificial intelligence to comprehend the underlying functional relationships of microbial communities. However, due to the unstructured nature of HTS data, nearly all computational models are limited to processing DNA sequences individually. This limitation causes them to miss out on key interactions between microorganisms, significantly hindering our understanding of how these interactions influence the microbial communities as a whole. Furthermore, most computational methods rely on post-processing of samples which could inadvertently introduce unintentional protocol-specific bias. RESULTS:Addressing these concerns, we present SetBERT, a robust pre-training methodology for creating generalized deep learning models for processing HTS data to produce contextualized embeddings and be fine-tuned for downstream tasks with explainable predictions. By leveraging sequence interactions, we show that SetBERT significantly outperforms other models in taxonomic classification with genus-level classification accuracy of 95%. Furthermore, we demonstrate that SetBERT is able to accurately explain its predictions autonomously by confirming the biological-relevance of taxa identified by the model. AVAILABILITY AND IMPLEMENTATION:All source code is available at https://github.com/DLii-Research/setbert. SetBERT may be used through the q2-deepdna QIIME 2 plugin whose source code is available at https://github.com/DLii-Research/q2-deepdna.
Shifts in agricultural land use over the past 200 years have led to a loss of nearly 50% of existing wetlands in the USA, and agricultural activities contribute up to 65% of the nutrients that reach the Mississippi River Basin, directly contributing to biological disasters such as the hypoxic Gulf of Mexico "Dead" Zone. Federal efforts to construct and restore wetland habitats have been employed to mitigate the detrimental effects of eutrophication, with an emphasis on the restoration of ecosystem services such as nutrient cycling and retention. Soil microbial assemblages drive biogeochemical cycles and offer a unique and sensitive framework for the accurate evaluation, restoration, and management of ecosystem services. The purpose of this study was to elucidate patterns of soil bacteria within and among wetlands by developing diversity profiles from high-throughput sequencing data, link functional gene copy number of nitrogen cycling genes to measured nutrient flux rates collected from flow-through incubation cores, and predict nutrient flux using microbial assemblage composition. Soil microbial assemblages showed fine-scale turnover in soil cores collected across the topsoil horizon (0-5 cm; top vs bottom partitions) and were structured by restoration practices on the easements (tree planting, shallow water, remnant forest). Connections between soil assemblage composition, functional gene copy number, and nutrient flux rates show the potential for soil bacterial assemblages to be used as bioindicators for nutrient cycling on the landscape. In addition, the predictive accuracy of flux rates was improved when implementing deep learning models that paired connected samples across time.
Emerging infectious diseases are increasingly recognized as a significant threat to global biodiversity conservation. Elucidating the relationship between pathogens and the host microbiome could lead to novel approaches for mitigating disease impacts. Pathogens can alter the host microbiome by inducing dysbiosis, an ecological state characterized by a reduction in bacterial alpha diversity, an increase in pathobionts, or a shift in beta diversity. We used the snake fungal disease (SFD; ophidiomycosis), system to examine how an emerging pathogen may induce dysbiosis across two experimental scales. We used quantitative polymerase chain reaction, bacterial amplicon sequencing, and a deep learning neural network to characterize the skin microbiome of free-ranging snakes across a broad phylogenetic and spatial extent. Habitat suitability models were used to find variables associated with fungal presence on the landscape. We also conducted a laboratory study of northern watersnakes to examine temporal changes in the skin microbiome following inoculation with Ophidiomyces ophidiicola . Patterns characteristic of dysbiosis were found at both scales, as were nonlinear changes in alpha and alterations in beta diversity, although structural-level and dispersion changes differed between field and laboratory contexts. The neural network was far more accurate (99.8% positive predictive value [PPV]) in predicting disease state than other analytic techniques (36.4% PPV). The genus Pseudomonas was characteristic of disease-negative microbiomes, whereas, positive snakes were characterized by the pathobionts Chryseobacterium , Paracoccus , and Sphingobacterium . Geographic regions suitable for O. ophidiicola had high pathogen loads (>0.66 maximum sensitivity + specificity). We found that pathogen-induced dysbiosis of the microbiome followed predictable trends, that disease state could be classified with neural network analyses, and that habitat suitability models predicted habitat for the SFD pathogen.
Bitcoin is known for its high volatility, which makes it challenging to accurately predict future prices. In this study, we aim to forecast Bitcoin prices for a month by incorporating exogenous variables, specifically the interest rate and recession probability. Our primary objective is to explore whether these variables have a positive impact on the prediction of Bitcoin prices. We used two popular time series forecasting models: Long Short-Term Memory (LSTM) and Facebook Prophet. Our approach involves exploring the impact of these exogenous variables on the performance of the models and comparing their results through plots and cross-validation. We trained the models using historical Bitcoin price data along with exogenous variables and evaluated their performance on a test dataset. Our results indicate that LSTM outperforms Facebook Prophet in terms of Bitcoin price prediction accuracy. This is because, while Facebook Prophet is optimized for statistical forecasting modeling, LSTM has the capability to learn intricate patterns and relationships given the right architecture with sufficient neurons. Importantly, we demonstrate that incorporating interest rates and recession probabilities significantly enhances the predictive capability of our models. Our findings suggest that changes in interest rates and recession probabilities have an impact on Bitcoin prices, and our models perform better when equipped with this valuable information.
Transformers, although first designed for sequence processing, can also handle unordered sets like point cloud data. Additionally, contrastive pretraining has emerged as a successful technique in image processing but remains unexplored for point cloud data. We develop and integrate a new point cloud pretraining technique inspired by the Simple Framework for Contrastive Learning (SimCLR) into the Set Transformer (ST) and Point Cloud Transformer (PCT) architectures and explore model performance using a novel 3D body scan dataset and the canonical datasets ShapeNet and ModelNet. For the 3D body scan dataset, this integration boosts initial training performance and maintains overall higher performance for classification tasks, and demonstrates better stability/convergence for regression tasks in comparison to non-pretrained (Naive) counterparts. Furthermore, experiments examining strong generalization (relative performance on previously unseen classes) show improvement for pretrained models compared to Naive models. Consistent benefits across tasks and data sets are observed based on additional experiments performed on the ShapeNet core dataset. Overall, we show how contrastive pretraining for point cloud data is a viable strategy for improving the performance of Transformers on downstream tasks and accelerating the training process.
Neuroscience provides a rich source of inspiration for new types of algorithms and architectures to employ when building AI and the resulting biologically-plausible approaches that provide formal, testable models of brain function. The working memory toolkit (WMtk), was developed to assist the integration of an artificial neural network (ANN)-based computational neuroscience model of working memory into reinforcement learning (RL) agents, mitigating the details of ANN design and providing a simple symbolic encoding interface. While the WMtk allows RL agents to perform well in partially-observable domains, it requires prefiltering of sensory information by the programmer: a task often delegated to dimensional attention mechanisms in other cognitive architectures. To fill this gap, we develop and test a biologically-plausible dimensional attention filter for the WMtk and validate model performance using a partially-observable 1D maze task. We show that the attention filter improves learning behavior in two ways by: 1) speeding up learning in the short-term, early in training and 2) developing emergent alternative strategies which optimize performance over the long-term.
Understanding how pH modulates underlying protein-protein interaction requires in-depth analysis across combinations of HIV proteins and broadly neutralizing antibodies (bnAbs). Traditional labs are not practical to screen all sequence variations and titrate of pH, highlighting the need to analyze such interactions theoretically. To fill this need, we generate homology models of predetermined HIV-1 gp120 in complex with bnAbs, through computational simulations to observe the binding energies as environmental pH varies. We compare simulation data to experimental data of the same structures at pH 5.5 and pH 7.4 specifically to present an 83.3% agreement between the two approaches to support the hypothesis that binding is stronger at lower pH. We then make observations of binding energy predictions across broad spectrum pH to provide insight into factors limiting the effectiveness of bnAbs in vivo. We conclude that theoretical models, and simulations involving them, provide data that closely mimic those of laboratory results.
Neurobiologically-inspired working memory models demonstrate human/animal capabilities to rapidly adapt and alter responses to the environment via context-switching and error monitoring. However, the application of these models outside of reinforcement learning problems has been relatively unexplored. We present a new framework compatible with Tensorflow/Keras enabling the integration of working memory-inspired mechanisms into typical neural network architectures. These mechanisms allow models to autonomously learn multiple tasks, statically or dynamically allocated. We also examine the generalization of the framework across a variety of multi-context supervised learning and reinforcement learning tasks. The resulting experiments successfully integrate these mechanisms with multi-layer and convolutional neural network architectures and the diversity of problems solved demonstrates the framework's generalizability across a variety of architectures and tasks.
SARS-CoV-2 is a novel virus that crossed over into humans in 2019 and declared a pandemic in early 2020. To understand how the virus infects a new host, we need to understand the mechanistic functions involved with the binding process. To address this need, we generate homology models of SARS-CoV-2 spikes as monomer and trimer to determine the feasibility of reduced computational requirements by using monomer structures. We further generate homology models of the conserved region of SARS-CoV-2 spike subunit s1 noted as the receptor binding domain (RBD) and the human angiotensin-converting enzyme 2 (ACE2). To determine functional breadth of spike monomer, trimer and RBD in relation with ACE2, we apply Coulombs Law to determine an electric force between combinations with ACE2 across the range of pH from 3.0 to 9.0 in 0.1 increments. The results indicate that spike trimer should be used to determine mechanistic binding function and these data indicate that variations of spike sequence influence breadth of function. Our results also indicate the RBD has a broader range of function across pH compared to spike trimer, but is influenced by the range of function presented by the spike trimer.
Social media platforms can struggle to enforce rules preventing online abuse and hate speech due to the large amount of content that must be manually reviewed. Machine learning approaches have been proposed in the literature as a way to automate much of these labors, but social content in multiple languages further complicates this issue. Past work has focused on first building word embeddings in the target language which limits the application of such embeddings to other languages. We use the Google Neural Machine Translator (NMT) to identify and translate Non-English text to English to make the system language agnostic. We can therefore use already available pre-trained word embeddings, instead of training our models and word embeddings in different languages. We have experimented with different word-embedding and classifier pairs as we aimed to assess whether translated English data gives us accuracy comparable to an untranslated English dataset. Our best performing model, SVM with TF-IDF, gave us a 10-fold accuracy of 95.56 percent followed by the BERT model with a 10-fold accuracy of 94.66 percent on the translated data. This accuracy is close to the accuracy of the untranslated English dataset and far better than the accuracy of the untranslated Hindi dataset.
Transfer learning allows for knowledge to generalize across tasks, resulting in increased learning speed and/or performance. These tasks must have commonalities that allow for knowledge to be transferred. The main goal of transfer learning in the reinforcement learning domain is to train and learn on one or more source tasks in order to learn a target task that exhibits better performance than if transfer was not used (Taylor and Stone 2009). Furthermore, the use of output-gated neural network models of working memory has been shown to increase generalization for supervised learning tasks (Kriete and Noelle 2011; Kriete et al. 2013). We propose that working memory-based generalization plays a significant role in a model's ability to transfer knowledge successfully across tasks. Thus, we extended the Holographic Working Memory Toolkit (HWMtk) (Dubois and Phillips 2017; Phillips and Noelle 2005) to utilize the generalization benefits of output gating within a working memory system. Finally. the model's utility was tested on a temporally extended, partially observable 5x5 2D grid-world maze task that required the agent to learn 3 tasks over the duration of the training period. The results indicate that the addition of output gating increases the initial learning performance of an agent in target tasks and decreases the learning time required to reach a fixed performance threshold.
An integral function of fully autonomous robots and humans is the ability to focus attention on a few relevant percepts to reach a certain goal while disregarding irrelevant percepts. Humans and animals rely on the interactions between the Pre-Frontal Cortex (PFC) and the Basal Ganglia (BG) to achieve this focus called Working Memory (WM). The Working Memory Toolkit (WMtk) was developed based on a computational neuroscience model of this phenomenon with Temporal Difference (TD) Learning for autonomous systems. Recent adaptations of the toolkit either utilize Abstract Task Representations (ATRs) to solve Feedback-Based (FB) tasks or storage of past input features to solve Sensory-Based (SB) tasks, but not both. We propose a new model, SBFBWMtk, which combines both approaches, ATRs and input storage, with a static or dynamic number of ATRs. The results of our experiments show that SBFBWMtk performs effectively for tasks that exhibit SB, FB, or both properties.
An integral function of fully autonomous robots and humans is the ability to focus attention on a few relevant percepts to reach a certain goal while disregarding irrelevant percepts. Humans and animals rely on the interactions between the Pre-Frontal Cortex (PFC) and the Basal Ganglia (BG) to achieve this focus called Working Memory (WM). The Working Memory Toolkit (WMtk) was developed based on a computational neuroscience model of this phenomenon with Temporal Difference (TD) Learning for autonomous systems. Recent adaptations of the toolkit either utilize Abstract Task Representations (ATRs) to solve Non-Observable (NO) tasks or storage of past input features to solve Partially-Observable (PO) tasks, but not both. We propose a new model, PONOWMtk, which combines both approaches, ATRs and input storage, with a static or dynamic number of ATRs. The results of our experiments show that PONOWMtk performs effectively for tasks that exhibit PO, NO, or both properties.
The protein folding problem has been studied in the field of molecular biophysics and biochemistry for many years. Even small changes in folding patterns may lead to serious diseases such as Alzheimer's or Parkinson's where proteins are folded either too quickly or too slowly. Molecular dynamics (MD) is one of the tools used to understand how proteins fold into native conformations. While it captures sequences of conformations that lead over time to the folded state, limitations in simulation timescales remain problematic. Although many approaches have been suggested to speed up the simulation process using rapid changes in temperature or pressure, we propose a rational approach, Greedy-proximal A* (GPA*), derived from path finding algorithms to explore the supposed shortest path folding pathway from the unfolded to a given folded conformation. We introduce several new protein structure comparison metrics based on the contact map distance to help mitigate the challenges faced by "standard" metrics. We test our approach on proteins which represent the two main types of secondary structure: (a) the Trp-cage miniprotein construct TC5b (1L2Y) which is a short, fast-folding protein that represents an α-helical secondary structure formed because of a locked tryptophan in the middle, (b) the immunoglobulin binding domain of the streptococcal protein G (1GB1), containing an α-helix and several β-sheets, and (c) the chicken villin subdomain HP-35, N68H protein (1YRF)-one of the fastest folding proteins which forms three α-helices. We compare our algorithm to replica-exchange MD and steered MD methods which represent the main algorithms used for accelerating folding proteins with MD. We find that GPA* not only reduces the computational time needed to obtain the folded conformation without adding artificial energy bias but also makes it possible to generate trajectories which contain minimal motions needed for the folding transition.
Decades of research has yet to provide a vaccine for HIV, the virus which causes AIDS. Recent theoretical research has turned attention to mucosa pH levels over systemic pH levels. Previous research in this field developed a computational approach for determining pH sensitivity that indicated higher potential for transmission at mucosa pH levels present during intercourse. The process was extended to incorporate a principal component analysis (PCA)-based machine learning technique for classification of gp120 proteins against a known transmitted variant called Biomolecular Electro-Static Indexing (BESI). The original process has since been extended to the residue level by a process we termed Electrostatic Variance Masking (EVM) and used in conjunction with BESI to determine structural differences present among various subspecies across Clades A1 and C. Results indicate that structures outside of the core selected by EVM may be responsible for binding affinity observed in many other studies and that pH modulation of select substructures indicated by EVM may influence specific regions of the viral envelope protein (Env) involved in protein-protein interactions.
Flexible, rapid, and predictive approaches that do not require the use of large numbers of vertebrate test animals are needed because the chemical universe remains largely untested for potential hazards. Development of robust new approach methodologies and nontesting approaches requires the use of existing information via curated, integrated data sets. The ecological threshold of toxicological concern (ecoTTC) represents one such new approach methodology that can predict a conservative de minimis toxicity value for chemicals with little or no information available. For the creation of an ecoTTC tool, a large, diverse environmental data set was developed from multiple sources, with harmonization, characterization, and information quality assessment steps to ensure that the information could be effectively organized and mined. The resulting EnviroTox database contains 91 217 aquatic toxicity records representing 1563 species and 4016 unique Chemical Abstracts Service numbers and is a robust, curated database containing high‐quality aquatic toxicity studies that are traceable to the original information source. Chemical‐specific information is also linked to each record and includes physico‐chemical information, chemical descriptors, and mode of action classifications. Toxicity data are associated with the physico‐chemical data, mode of action classifications, and curated taxonomic information for the organisms tested. The EnviroTox platform also includes 3 analysis tools: a predicted‐no‐effect concentration calculator, an ecoTTC distribution tool, and a chemical toxicity distribution tool. Although the EnviroTox database and tools were originally developed to support ecoTTC analysis and development, they have broader applicability to the field of ecological risk assessment. Environ Toxicol Chem 2019;9999:1–12. © 2019 The Authors. Environmental Toxicology and Chemistry published by Wiley Periodicals, Inc. on behalf of SETAC.