Robotic systems are no longer simply built and designed to perform sequential repetitive tasks primarily in a static manufacturing environment. Systems such as autonomous vehicles make use of intricate machine learning algorithms to adapt their behavior to dynamic conditions in their operating environment. These machine learning algorithms provide an additional attack surface for an adversary to exploit in order to perform a cyberattack. Since an attack on robotic systems such as autonomous vehicles have the potential to cause great damage and harm to humans, it is essential that detection and defenses of these attacks be explored. This paper discusses the plausibility of direct and indirect cyberattacks on a machine learning model through the use of a virtual autonomous vehicle operating in a simulation environment using a machine learning model for control. Using this vehicle, this paper proposes various methods of detection of cyberattacks on its machine learning model and discusses possible defense mechanisms to prevent such attacks.
The human NEIL1 DNA glycosylase is one of 11 mammalian glycosylases that initiate base excision repair. While substrate preference, catalytic mechanism, and structural information of NEIL1's ordered residues are available, limited information on its subcellular localization, compounded by relatively low endogenous expression levels, have impeded our understanding of NEIL1. Here, we employed a previously developed computational framework to optimize the mitochondrial localization signal of NEIL1, enabling the visualization of its specific targeting to the mitochondrion via confocal microscopy. While we observed clear mitochondrial localization and increased glycosylase/lyase activity in mitochondrial extracts from low-moderate NEIL1 expression, high NEIL1 mitochondrial expression levels proved harmful, potentially leading to cell death.
We have developed a computational framework for constructing synthetic signal peptides from a base set of protein sequences. A large number of structured "building blocks", represented as m-step ordered pairs of amino acids, are extracted from the base sequences. Using a straightforward procedure, the building blocks enable the construction of a diverse set of synthetic signal peptides and targeting sequences that have the potential for industrial and therapeutic purposes. We have validated the proposed framework using several state-of-the-art sequence prediction platforms such as Signal-BLAST, SignalP-5.0, MULocDeep, and DeepMito. Experimental results show the computational framework can successfully generate synthetic signal peptides and targeting sequences and transform non-signaling sequences into synthetic signal peptides.
The concept of equipment maintenance is older than the industrial revolution. The mode, medium, and timing of maintenance during equipment life cycle have evolved from reactive maintenance to predictive maintenance to prescriptive maintenance. Prescriptive maintenance, which incorporates the Internet of Things, digitization, and artificial intelligence, has the potential to greatly improve upon proactive maintenance. The growth in applications based on prescriptive maintenance has been exponential. However existing solutions are piecemeal and lack complete solutions to keep equipment operating at optimal cost. We propose a Holistic end-to-end Prescriptive Maintenance Framework (HeePMF) that uses maintenance needs analysis, equipment, and operational data with predictive technologies and feedback to generate actionable insights. Other features implemented are personnel scheduling, supply chain improvement, field-replaceable unit (FRU) management, process improvement, and knowledge management. The working of framework (HeePMF) is demonstrated using datasets from the 2019 HACKtheMACHINE Data Science Competition. The implementation demonstrates data integration, selection of few critical discriminants using feature reduction, missing dataset computation, and noise removal. It ulilizes and demonstrates predictive algorithms to determine sub-components for impending failures at individual equipment and fleet levels, and then providing mechanism for complete repair solutions (FRUs order, service personnel scheduling, equipment downtime management). Steps like service personnel optimal deployment are implemented using a simulated dataset. This HeePMF is extensible and makes it truly prescriptive, resulting in a much reduced unplanned equipment downtime at optimal cost. This paper successfully defines extensible end-to-end holistic prescriptive equipment maintenance framework with optimal cost and demonstrates with a very thorough case study.
ABSTRACT We developed a computational method for constructing synthetic signal peptides from a base set of signal peptides (SPs) and non-SP sequences. A large number of structured “building blocks”, represented as m -step ordered pairs of amino acids, are extracted from the base. Using a straightforward procedure, the building blocks enable the construction of a diverse set of synthetic SPs that could be utilized for industrial and therapeutic purposes. We have validated the proposed methodology using existing sequence prediction platforms such as Signal-BLAST and MULocDeep. In one experiment, 9,555 protein sequences were generated from a large randomly selected set of “building blocks”. Signal-BLAST identified 8,444 (88%) of the sequences as signal peptides. In addition, the Signal-BLAST tool predicted that the generated synthetic sequences belonged to 854 distinct eukaryotic organisms. Here, we provide detailed descriptions and results from various experiments illustrating the potential usefulness of the methodology in generating signal peptide protein sequences.
Flexible fine-grained weather forecasting is a problem of national importance due to its stark impacts on economic development and human livelihoods. It remains challenging for such forecasting, given the limitation of currently employed statistical models, that usually involve the complex simulation governed by atmosphere physical equations. To address such a challenge, we develop a deep learning-based prediction model, called Micro-Macro, aiming to precisely forecast weather conditions in the fine temporal resolution (i.e., multiple consecutive short time horizons) based on both the atmospheric numerical output of WRF-HRRR (the weather research and forecasting model with high-resolution rapid refresh) and the ground observation of Mesonet stations. It includes: 1) an Encoder which leverages a set of LSTM units to process the past measurements sequentially in the temporal domain, arriving at a final dense vector that can capture the sequential temporal patterns; 2) a Periodical Mapper which is designed to extract the periodical patterns from past measurements; and 3) a Decoder which employs multiple LSTM units sequentially to forecast a set of weather parameters in the next few short time horizons. Our solution permits temporal scaling in weather parameter predictions flexibly, yielding precise weather forecasting in desirable temporal resolutions. It resorts to a number of Micro-Macro model instances, called modelets, one for each weather parameter per Mesonet station site, to collectively predict a target region precisely. Extensive experiments are conducted to forecast four important weather parameters at two Mesonet station sites. The results exhibit that our Micro-Macro model can achieve high prediction accuracy, outperforming almost all compared counterparts on four parameters of interest.
Action rule mining seeks to generate rules that indicate what changes can be made to move an object from one class (state) to another class. An action rule is composed of the changes, known as actions, that correlate with the change in the class value. The current work in action rule mining focuses on frequent, or highly occurring, action rules. The working assumption is that the end user is interested in class transitions that move a large number of objects from the initial class to the new (final) class with a high degree of confidence. Currently, very little work focuses on rare action rules; that is, classes which occur infrequently. In this paper, we provide a definition for a rare action rule and then propose a consequent-constraint based algorithm for generating these rare action rules.
In this paper, we conduct a systematic study for the very first time on the poisoning attack to neural collaborative filtering-based recommender systems, exploring both availability and target attacks with their respective goals of distorting recommended results and promoting specific targets. The key challenge arises on how to perform effective poisoning attacks by an attacker with limited manipulations to reduce expense, while achieving the maximum attack objectives. With an extensive study for exploring the characteristics of neural collaborative filterings, we develop a rigorous model for specifying the constraints of attacks, and then define different objective functions to capture the essential goals for availability attack and target attack. Formulated into optimization problems which are in the complex forms of non-convex programming, these attack models are effectively solved by our delicately designed algorithms. Our proposed poisoning attack solutions are evaluated on datasets from different web platforms, e.g., Amazon, Twitter, and MovieLens. Experimental results have demonstrated that both of them are effective, soundly outperforming the baseline methods.
Detecting anomalies, in many cases, is only the first step. Often, it is merely the first step in a process that seeks to recover from an anomalous situation so that the system is back to a normal state. Manual analysis is a possible approach; however, this can be time consuming and error prone. Ideally, it would be desirable to have an automated means for resolving anomalies, which is invoked whenever one or more anomalies are detected. In this paper, we present a tool, TADS (Transformation of Anomalies in Data Streams), that creates automated recommendations based upon the observed streaming data. This is accomplished by utilizing the recently introduced concept of dynamic action rules with the Massive Online Analysis (MOA) data streaming platform. In experimental results, we demonstrate that TADS is able to correctly generate recommendations 100% of the time for over half the experiments and over 90% for all but two experimental conditions. In addition, only seconds are required to generate the personalized recommendation for each anomaly. Hence, the results indicate that TADS is a viable approach for correcting anomalies within a streaming environment.
Optimal inventory management always requires the right number and types of parts to be available at the right time at the right place. However, this has become even more critical in the age of prescriptive maintenance, which seeks to predict an impending failure along with what part(s) are about to fail. To take advantage of these predictions, it is essential that the parts, personnel and time frame are made available in a timely fashion to keep the operation going uninterrupted. Usually, the level of parts maintained in a part warehouse is determined using static thresholding where quantities of each part is determined and is fixed for a given part. In literature, such methods are known as static thresholding. While work has been done to accommodate unpredictability of parts usage in supply chain, there is lack of attention that has been paid to consider the demand of parts in future due to prediction of impending failure. To address this, this paper introduces a noble method and system, entitled dynamic thresholding, to ensure the part inventory adapts to predicted usage while maintaining minimum inventory management cost. The feasibility of the concept and system is demonstrated using simulated datasets.
The concept of equipment maintenance is as old, if not older, as industrial revolution. However, the mode, medium and timing of maintenance during equipment life cycle have evolved from reactive maintenance to prescriptive maintenance. Reactive maintenance was the main approach practiced until 1980s. The proactive maintenance was introduced in 1990s to insulate customers from equipment failures by using equipment data and fixing problems remotely. Prescriptive maintenance, which incorporates the Internet of Things, digitization and artificial intelligence, especially machine learning, has gained importance as in intelligent approach to equipment maintenance and has the potential to greatly improve upon proactive maintenance. However, existing prescriptive maintenance solutions are piece-meal and lack complete solutions to keep equipment in working conditions most of the time at optimal cost. We propose a Holistic End to End Prescriptive Maintenance (HeePMF) framework resulting in much reduced unplanned equipment downtime at optimal cost. This framework provides solution by integrating maintenance needs analysis, equipment and operational data, predictive technologies with feedback, personnel scheduling, supply chain improvement, part management, process improvement and knowledge management. The working of framework is demonstrated by providing a case study.
Action rule mining develops rules that describe which attributes should be changed in order to move an object from an undesired state to a desired state, with the understanding that some attributes cannot be changed. While such rules can be very useful for end-users, a limitation in prior work is the underlying assumption that the attributes of a dataset are discrete in nature. To address this limitation, we propose a method for generating action rules for objects described by continuously valued data. As part of the process, we developed a model for determining the effectiveness of the change, which permits more tailored recommendations for how to modify objects. Experimental results indicate that we can successfully create action rules for continuous valued data, and the use of automated tuning reduces the number of changes that must be performed to move an object from an undesired state to a desired state.
Action rules are rules that describe how to transition a decision attribute from an undesired state to a desired state, with the understanding that some attributes are stable and others are flexible. Stable attributes, such as "age", may not be changed, whereas flexible attributes, such as "interest rate", may be changed. Action rules have great potential in data mining, as they output easily interpretable rules which can immediately be useful to a decision maker. However, at present, the methods to generate all valid action rules are computationally expensive. To address this, methods have been proposed that prune swaths of the search space as rules are generated; this results in computational efficiency, at the expense of potentially not discovering many useful rules. In this work, a method, called Multi-Objective Evolutionary Action Rule (MOEAR) mining, is introduced. MOEAR optimizes the discovery of action rules using standard evolutionary algorithm principles. Experimental results show that MOEAR is able to generate a large number of potentially interesting action rules, including those rules that could be categorized as "rare", while achieving good computational performance.
With computer software becoming more important and prolific in today's world, malicious software (malware) continues to be one of its greatest security threats. Alongside this trend, smartphones and mobile devices have become the prominent method for accessing the Internet and its vast resources of information and business applications. With the amount and variety of Android based devices increasing daily, the need for better and more accurate malware detection approaches for the Android platform also increases. In this paper, we explore whether a data mining technique originally developed to detect malware on a Windows operating system can be utilized to detect malware in Android mobile devices. In addition, we propose a novel algorithm for detecting malware on Android that relies on step sizes and a simplified multi-layer vector space (MLVS) model. We compare the effectiveness of these two techniques, with the goal of determining optimal step sizes for our modified MLVS (MMLVS) approach to detect Android malware. Our results show that the two methods are able to correctly classify the samples as malware or uninfected with strong accuracy. In addition, we identify key elements that need to be address to permit further improvement within Android environments.
Analysing and classifying sequences based on similarities and differences is a mathematical problem of escalating relevance and importance in many scientific disciplines. One of the primary challenges in applying machine learning algorithms to sequential data, such as biological sequences, is the extraction and representation of significant features from the data. To address this problem, we have recently developed a representation, entitled Multi-Layered Vector Spaces (MLVS), which is a simple mathematical model that maps sequences into a set of MLVS. We demonstrate the usefulness of the model by applying it to the problem of identifying signal peptides. MLVS feature vectors are generated from a collection of protein sequences and the resulting vectors are used to create support vector machine classifiers. Experiments show that the MLVS-based classifiers are able to outperform or perform on par with several existing methods that are specifically designed for the purpose of identifying signal peptides.
This chapter describes a comprehensive granular model for decision making with complex data. This granular model first uses information decomposition to form a horizontal set of granules for each of the data instances. Each granule is a partial view of the corresponding data instance; and aggregately all the partial views of that data instance provide a complete representation for the instance. Then, the decision making based on the original data can be divided and distributed to decision making on the collection of each partial view. The decisions made on all partial views will then be aggregated to form a final global decision. Moreover, on each partial view, a sequential M+1 way decision making (a simple extension of Yao's 3-way decision making) can be carried out to reach a local decision. This chapter further categorizes stock price predication problem using the proposed decision model and incorporates the MLVS model for biological sequence classification into the proposed decision model. It is suggested that the proposed model provide a general framework to address the complexity and volume challenges in big data analytics.
Clinical trials for interventions that seek to delay the onset of Alzheimer’s disease (AD) are hampered by inadequate methods for selecting study subjects who are at risk, and who may therefore benefit from the interventions being studied. Automated monitoring tools may facilitate clinical research and thereby reduce the impact of AD on individuals, caregivers, society at large, and government healthcare infrastructure. We studied the 18F-deoxyglucose positron emission tomography (FDG-PET) scans of research subjects from the Alzheimer’s Disease Neuroimaging Initiative (ADNI), using a Machine Learning technique. Three hundred ninety-four FDG-PET scans were obtained from the ADNI database. An automated procedure was used to extract measurements from 31 regions of each PET surface projection. These data points were used to evaluate the sensitivity and specificity of support vector machine (SVM) classifiers and to compare both Linear and Radial-Basis SVM techniques against a classic thresholding method used in earlier work.
Analyzing and classifying sequence data based on structural similarities and differences is a mathematical problem of escalating relevance. Indeed, a primary challenge in designing machine learning algorithms to analyzing sequence data is the extraction and representation of significant features. This paper introduces a generalized sequence feature extraction model, referred to as the Generalized Multi-Layered Vector Spaces (GMLVS) model. Unlike most models that represent sequence data based on subsequences frequency, the GMLVS model represents a given sequence as a collection of features, where each individual feature captures the spatial relationships between two subsequences and can be mapped into a feature vector. The utility of this approach is demonstrated via two special cases of the GMLVS model, namely, Lossless Decomposition (LD) and the Multi-Layered Vector Spaces (MLVS). Experimental evaluation show the GMLVS inspired models generated feature vectors that, combined with basic machine learning techniques, are able to achieve high classification performance.
The task of learning action rules aims to provide recommendations to analysts seeking to achieve a specific change. An action rule is constructed as a series of changes, or actions, which can be made to the flexible characteristics of a given object that ultimately triggers the desired change. Existing action rule discovery methods utilize a generate-and-test approach in which candidate action rules are generated and those that satisfy the user-defined thresholds are returned. A shortcoming of this operational model is there is no guarantee all objects are covered by the generated action rules. In this paper, we define a new methodology referred to as Targeted Action Rule Discovery (TARD). This methodology represents an object driven approach in which an action rule is explicitly discovered per target object. A TARD method is proposed that effectively discovers action rules through the iterative construction of multiple decision trees. Experiments show the proposed method is able to provide higher quality rules than the well-known Association Action Rule (AAR) method.