
The concept of equipment maintenance is older than the industrial revolution. The mode, medium, and timing of maintenance during equipment life cycle have evolved from reactive maintenance to predictive maintenance to prescriptive maintenance. Prescriptive maintenance, which incorporates the Internet of Things, digitization, and artificial intelligence, has the potential to greatly improve upon proactive maintenance. The growth in applications based on prescriptive maintenance has been exponential. However existing solutions are piecemeal and lack complete solutions to keep equipment operating at optimal cost. We propose a Holistic end-to-end Prescriptive Maintenance Framework (HeePMF) that uses maintenance needs analysis, equipment, and operational data with predictive technologies and feedback to generate actionable insights. Other features implemented are personnel scheduling, supply chain improvement, field-replaceable unit (FRU) management, process improvement, and knowledge management. The working of framework (HeePMF) is demonstrated using datasets from the 2019 HACKtheMACHINE Data Science Competition. The implementation demonstrates data integration, selection of few critical discriminants using feature reduction, missing dataset computation, and noise removal. It ulilizes and demonstrates predictive algorithms to determine sub-components for impending failures at individual equipment and fleet levels, and then providing mechanism for complete repair solutions (FRUs order, service personnel scheduling, equipment downtime management). Steps like service personnel optimal deployment are implemented using a simulated dataset. This HeePMF is extensible and makes it truly prescriptive, resulting in a much reduced unplanned equipment downtime at optimal cost. This paper successfully defines extensible end-to-end holistic prescriptive equipment maintenance framework with optimal cost and demonstrates with a very thorough case study.
The technological revolution enabled by the Industry 4.0 revolution is fusing the physical and digital worlds through the confluence of various technologies (i.e., Internet of Things (IoT), Artificial Intelligence, Cyber-physical Systems, and smart factories) [1] [2].Industrial systems in several domains ranging from manufacturing, transportation, energy, defense, automotive, and buildings operate in an environment that is highly dynamic, safety-critical, and uncertain [3].Bringing automation and connectivity into these industrial systems introduces system complexity inherent to systemlevel integration and operation.It is our endeavor to engineer future industrial systems that not only augment automation technologies, but are also safe and dependable.Cyber-physical systems are engineered systems built from seamless integration of computation (i.e., sensing, computing, and networking) and physical components.There are many data-centric challenges related to the implementation, operation, control, and optimization of these systems with a high degree of complexity.These challenges arise from data, model, or system-level integration [4].Data related issues arise from inherent nature of sensory data, that includes noisy, uncertain, partially informative, and dynamic data.The modeling-related challenges range from how to improve the robustness of machine learning models, primarily when used
Cyber-physical systems (CPS) are finding increasing application in many domains. CPS are composed of sensors, actuators, a central decision-making unit, and a network connecting all of these components. The design of CPS involves the selection of these hardware and software components, and this design process could be limited by a cost constraint. This study assumes that the central decision-making unit is a binary classifier, and casts the design problem as a feature selection problem for the binary classifier where each feature has an associated cost. Receiver operating characteristic (ROC) curves are a useful tool for comparing and selecting binary classifiers; however, ROC curves only consider the misclassification cost of the classifier and ignore other costs such as the cost of the features. The authors previously proposed a method called ROC Convex Hull with Cost (ROCCHC) that is used to select ROC optimal classifiers when cost is a factor. ROCCHC extends the widely used ROC Convex Hull (ROCCH) method by combining it with the Pareto analysis for cost optimization. This paper proposes using the ROCCHC analysis as the evaluation function for feature selection search methods without requiring an exhaustive search over the feature space. This analysis is performed on 6 real-world data sets, including a diagnostic cyber-physical system for hydraulic actuators. The ROCCHC analysis is demonstrated using sequential forward and backward search. The results are compared with the ROCCH selection method and a popular Pareto selection method that uses classification accuracy and feature cost.
The complexity metric is an effective tool to evaluate the behavioral dynamics in systems with high level of nonlinearity and interconnectivity. This paper aims to evaluate the effectiveness of using the entropic complexity as a feature to enhance situational awareness in dynamical systems. In fact, the complexity measurement aims to detect dynamical changes and the pattern recognition tool discerns a particular dynamical change from other types of dynamics. In this study, parameters associated with the permutation entropy and also the complexity are determined in real time. The test system is a small-scale microgrid in which a solid-state transformer (SST) is operating under different dynamical conditions. The complexity measurement unit provides datasets that will be used for detection and identification of particular dynamics in the system that can be potentially used in real-time decision-making. Results show the effectiveness of the proposed approach in detection and recognition of certain dynamics in microgrids.
Predictive maintenance applications for a wide variety of industrial and commercial components are increasingly utilizing imaging-based sensors along with AI (artificial intelligence)/ML (machine learning) based analytics to determine wear of components. Credibility of the analytics, especially for component health, is strongly dependent on explainability. We initially introduce an explainable framework involving a novel light transmission image processing–based methodology utilizing statistical distance metrics (e.g., Wasserstein distance (WD), Kolmogorov-Smirnov statistic) for discriminative classification of unstructured images combined with Bayesian inference/regression to estimate wear level for an air filter application. Subsequently, we incorporate neural network–based models into this framework to develop an AI framework retaining a high level of explainability. The explainable elements of this novel AI model include generation of a statistical distance pseudometric with a feedforward neural network as a discriminative classifier, a spatial block bootstrapping approach to generate synthetic training data, and the use of this discriminant classifier as a predictor in a Bayesian inference/regression model to predict wear levels with 95% prediction and credible intervals. This explainable AI framework can be extended to other families of applications utilizing synthetically generated unstructured and structured images in predictive maintenance and health monitoring.
Journey by aircraft is the only option for long-distance transportation and also one of the frequently used modes of transportation of passengers. As a result, safety of passengers and efficiency of the aircraft depend on maintaining efficient running conditions. Although many safety standards are followed in the design of the aircraft, and thus there are fewer accidents, it is necessary to perform a thorough analysis to avoid risks that may occur during flight time. In the present work, we propose a maintenance strategy, Failing And Not Falling (F&!F), based on the Federal Aviation Administration (FAA) data in the USA. We work with the dataset of Boeing 737. The data consists of 72 features with 137,236 records which describe an aircraft accident or incident. These features are used to predict whether an incident will be identified during aircraft maintenance or during aircraft operation and what specific type of incident will occur. The prediction method is based on the integration of a decision tree and a unique neural network at each node of the decision tree. The results obtained using different architectures show how deep the neural networks should be, how to identify the relevant features, and the success of combining decision trees and neural networks. Moreover, the neural networks and the decision tree approach also successfully identified the important features of maintenance. This method can be used for the maintenance of any data in multiple domains.
Ever-increasing amounts of data and requirements to process them in real time lead to more and more analytics platforms and software systems designed according to the concept of stream processing. A common area of application is processing continuous data streams from sensors, for example, IoT devices or performance monitoring tools. In addition to analyzing pure sensor data, analyses of data for entire groups of sensors often need to be performed. Therefore, data streams of the individual sensors have to be continuously aggregated to a data stream for a group. Motivated by a real-world application scenario of analyzing power consumption in Industry 4.0 environments, we propose that such a stream aggregation approach has to allow for aggregating sensors in hierarchical groups, support multiple such hierarchies in parallel, provide reconfiguration at runtime, and preserve the scalability and reliability qualities of stream processing techniques. We propose a stream processing architecture fulfilling these requirements, which can be integrated into existing big data architectures. As all state-of-the-art stream processing frameworks have to handle a trade-off between latency, resource-efficiency, and correctness, our proposed architecture can be configured for low latency and resource-efficient computation or for always ensuring correct results. To assist adopters in choosing appropriate configuration options, we provide an experimental comparison. We present a pilot implementation of our proposed architecture and show how it is used in industry. Furthermore, in experimental evaluations we show that our solution scales linearly with the amount of sensors and provides adequate reliability in the presence of faults.
The aim of this paper is to identify the co-expressed potential genes that may serve for the development of the portions of normal or tumor. This paper differentiates the co-expressed genes into normal samples and tumor samples from gene expression dataset GSE25066. Since the dataset has vague boundaries and having common characteristics between the clusters, identifying the subgroups contain similar gene expression is really a tricky task one. Therefore, this paper introduces an effective fuzzy iterative clustering algorithm by incorporating kernel function, possibilistic c-means, fuzzy memberships, neighborhood information, median of neighboring objects and penalty term. The performances of the proposed clustering techniques have been shown through the succession experimental works on GSE25066. The effects of clustering results have been proved through comparing the resulted classes with ground truth.
Acrylic polymer composite, also known as “bone cement,” is widely used in orthopedic surgery to anchor hip arthroplasty. The distribution of the random damage over the various modes changes as the stress increases, and it occurs in the form of microscopic events with specific characteristics. Evidence of the random damage is the continuous occurrence of randomly generated microscopic damage events as observed by many authors. Knowledge is lacking, however, the quantification of these random damage events. The occurrence of random damage events followed primarily a Gaussian probability distribution model, $$ f(y)={a}_0{e}^{-{\left(\frac{y-{b}_0}{c_0}\right)}^2} $$ for the cases under tension loads, where ao, bo, and co are the magnification factor (the highest probability when the scale of an event is equal to bo), and bo and co are related to the centroid, and peak distribution. An exponential distribution model $$ f(y)={a}_0{e}^{-{b}_0y} $$ was found for the cases under bending, where ao is, again, the magnification factor, and bo is the damage event scale constant. We found a fact that the occurrence of random damage events was relatively independent of the applied stress in certain loading stages such that there were not significant changes of random damage events when the applied stress increased. This fact suggested that the use of damage event accumulation may be a better indicator to reveal the integrity of this material than that of the applied stress alone. Our empirical models may be used as a ground work to reveal and describe the occurrence of multiscale random damage events in brittle and semi-brittle materials.
Advances in computer vision technology have expanded the possibilities to facilitate complex task automation for integration into large-scale data processing solutions. Despite these advances, however, there is still a need to develop simple and efficient algorithms for image feature extraction and classification to enable easier and faster implementation into real-world applications. Here, a new method is described to extract features from images that can be used for image classification. It uses a fuzzy c-means (FCM) clustering-based approach that allows for unique object patterns to be spatially re-mapped onto a binary sparse matrix with which principles from recurrence quantification analysis statistics (RQAS) can be applied. RQAS are computationally efficient and can be used to create a short feature vector for effective binary and multi-class image classification. The utility of this method is demonstrated using both simulated and real datasets that include objects embedded in complex backgrounds, and is compared with another widely used and highly effective thresholding feature extraction method (local binary patterns (LBP)). Results show that the FCM-RQAS method described here can perform as well or better than LBP and supports the use and further development of RQAS-based image feature extraction for computer vision applications.
This study examines acoustic emission data obtained during intermittent static tension loading of progressively fatigued 4340 steel and 7075-T651 aluminum specimens with the aim of inferring fatigue damage information from static tension testing. Acoustic emission data were collected using a novel loading procedure based on the Dunegan corollary. Results from 4340 steel testing showed a moderate correlation between total acoustic emission energy parameter and the number of cyclic loading cycles. Results from 7075-T651 aluminum testing showed a moderate correlation between the information entropy parameter and loading cycles. A supervised neural network was assessed to be 54.0 ± 19.1% accurate in predicting cyclic loading cycles for 4340 steel specimens and 52.0 ± 19.2% accurate for 7075-T651 aluminum specimens. Overall, results showed that limited but potentially useful fatigue damage information from 4340 steel or 7075-T651 aluminum is contained within acoustic emission signals collected during elastic tension loading.
With the rise in science and technology, the publishing and sharing of knowledge have become an undeniable factor for the growth of an academic venue. The worth of any publication depends upon how frequently that article is being cited and by which prestige. In this work, the inter-field citation index (IFC) technique is proposed which highlights and differentiates between interdisciplinary and intradisciplinary citations and their role in the research community. The proposed technique can perform well for the researchers for their overall impact and ranking. Further, to explain and calculate the prestige of an article as per the proposed idea, a comparison between the proposed work and h-index and global citation impact (GCI) has also been anticipated.
The aim of data science is to catch up with the data-intensive life style as well as the demand for decision support, which becomes common in various domains such as medical, education, and other smart solutions. As such, high quality of data analysis is greatly desired for accurate and effective downstreaming exploitations. Specific to data clustering, vast amounts of works have concentrated on modeling a distance metric and a clustering algorithm, with the assumption of a complete data. However, this might not always be the case as missing values can occur in the dataset under examination. Instead of filling in these values using an imputation method, a recent study successfully makes use of the consensus clustering to overcome the problem without committing an explicit imputation procedure. This paper extends the previous framework to link-based consensus clustering that provides a more refined summarization of cluster ensemble, hence the resulting data partition. It exhibits a promising performance on several benchmark data collections obtained from UCI repository.
Classification is one of the supervised learning models, and enhancing the performance of a classification model has been a challenging research problem in the fields of machine learning (ML) and data mining. The goal of ML is to produce or build a model that can be used to perform classification. It is important to achieve superior performance of the classification model. Obtaining a better performance is important for almost all fields including healthcare. Researchers have been using different ML techniques to obtain better performance of their models; ensemble techniques are also used to combine multiple base learner models. The ML technique called super learning or stacked-ensemble is an ensemble method that finds the optimal weighted average of diverse learning models. In this paper, we have used super learning or stacked-ensemble achieving better performance on four benchmark data sets that are related to healthcare. Experimental results show that super learning has a better performance compared to the individual base learners and the baseline ensemble.
This article presents a robust predictive model using parametric copula-based regression. We show that copula selection test procedures and predictive conditional distributions can be used to assess model adequacy and predictive validity. We offer simulation experiments to demonstrate the ability of our diagnostic procedure to correctly identify the true data generating process. Finally, we apply our methodology on a well-known insurance claims dataset to produce the distribution profile of allocated loss adjustment expense for given pre-specified indemnity payments information. The availability of this entire expense distribution will provide greater insight to the decision-makers before allocating resources for a given insurance claim.
It is challenging to determine the damage type or mechanism under different stress states because of the high degree of disorder of the microdamage inside the metal material. The application of many nondestructive testing methods makes the realization of this subject possible. In this work, acoustic emission (AE) was implemented to test the microdamage evolution process of two aluminum alloy materials (1060 and 6063). Two types of notched specimens (shear and tensile) had been used. AE signatures acquired during testing were used to construct the multicomponent variate D A damage matrix. The multicomponent variate D A matrix and probabilistic entropy were applied to analyze the diversity of the microdamage evolution of different materials under different stress states. And the fracture surfaces were observed by a scanning electron microscope (SEM) to verify the correctness of the analysis results. Consequently, the probabilistic entropy results show that AE data characterization can effectively distinguish the significant difference of microdamage evolution of the aluminum alloys under two kinds of stress state. It is also shown that the microdamage evolution mechanism represented by the probabilistic entropy of aluminum alloys under the same stress state is consistent.
Deep learning recently attracts considerable attention thanks to its powerful computational capacities in image processing and natural language processing. More and more real estate brokers provide online “Deep” expert systems to help clients with their inquiry of targeted properties before deciding on the transaction of properties. The real estate appraisal is one of the most significant concerns for the clients. In the appraisal process, the estimation of house price depends not only on its attributes but also their neighbors. The influence from neighbors is known as peer-dependence, which is not directly measurable. Thus, real estate appraisal can be improved if the valuation includes the peer-dependence measurement. In this paper, we propose a peer-dependence valuation model (PDVM), which is capable of converting the peer-dependence-based valuation problem into a sequence prediction problem. In the proposed model, we first develop a method, K-nearest similar house sampling (KNSHS), to generate sequences from the to-be-value house and nearby houses. Secondly, the bidirectional long short-term memory (B-LSTM) layers extract the features of sequences. Finally, the fully connected (FC) layer estimates the house price based on the features. The experimental results indicate that our model outperforms the other state-of-the-art machine learning models being used for real estate appraisal.
Structural health monitoring (SHM) will be pivotal for safe and economic operation of wind turbines. Timely discovery of changes within the structure and means of prediction of required maintenance will reduce production costs of electricity and catastrophic failures. Long-term structural acceleration recording can support damage detection on turbine towers and document progression of fatigue. Conventional acceleration recordings are based on wired sensor nodes at fixed positions with privileged accessibility and electric power supply. However, such positions might be near vibration nodes and not necessarily experience the maximum vibration amplitude. Shifts in eigenfrequencies can be an indicator of changes in structural stiffness, hence damage, but also be caused by environmental effects, e.g., temperature. Damages generate local effects while the structure’s vibration spectrum is a global evaluation. If a sensor is close to the location of damage, the probability of detection is increased. Wireless sensors powered by batteries are advantageous for this task as they are independent of cabling for power supply and data transmission. Such monitoring of turbine tower structures is not common in practice and requires new data-enabled techniques to discover deviations from the optimal way of wind turbine operation. This paper proposes a new approach using wireless high-resolution acceleration measurement sensor nodes, exploiting the vibration response of wind turbine towers. Influences of acceleration resolution and sensor node locations onto the accuracy of eigenfrequency determination are demonstrated. A comparison between acceleration recordings by wireless sensor nodes and their wired counterparts is presented to prove the equivalence of the wireless sensing method. Finally, new data compression techniques used with the sensor nodes are discussed to reduce wireless transmission to a minimum.
This paper aims to present optimal clustering techniques for analyzing high-dimensional cancer databases with missing attributes and overlapped objects. Analyzing the high-dimensional database with missing values is considered as most difficult task, and so far, there is no optimal cluster technique available for clustering the cancer database. Therefore, this paper develops the effective fuzzy clustering techniques that incorporate Cauchy kernel induced distance, rudimentary centroids, possibilistic memberships, fuzzy memberships, and prototype equation. To reduce the computing time of algorithms, this paper introduces a method for finding reasonable initial cluster centers. Experimental results indicate that the proposed methods are suitable for the breast cancer databases with missing attributes, and the results indicate that the methods outperform in clustering the databases into available subclasses.
One-size-fits-all has never succeeded profoundly and this is the reason why personalization has become so important. Many fatal diseases face such problem due to intense and risky medication procedures with immense side effects. Personalized medication through precision medicine is the new big thing that will provide better treatment solution through analysis and artificial intelligence. In this work, we have introduced for the first time a multi-modular training approach (MMSG+DRRL/ERRL) for generalized representation for cancer drug prediction based on gene expression for personalized treatment through the use of different system-genetics information like subclinical medical information to determine their suitability. We have shown how different modular information can be used computationally to enhance prediction behavior using the genetic and subcategory information of cancer diseases. Subcategorization analysis can decode many useful information of the disease and can help in better treatment modeling. We have shown for the first time how we can use the subcategory information to extract sensitivity of cancer cell lines for different drugs, a procedure never tried before. We have utilized the best practices of the feature training for deep learning to determine the interacting genes for particular drugs.