In the Industrial Internet of Things (IIoT), data analytics plays an important role in the management and planning of production processes. In particular, forecasting future states helps to ensure the operation of the equipment, optimize its performance and detect anomalies and failures in a timely manner. We propose a framework called ForecaState for IIoT state forecasting. ForecaState utilizes variations of recurrent neural networks (RNNs), including long short-term memory (LSTM) and gated recurrent units (GRUs), to capture complex patterns and dependencies in industrial data. These advanced RNN architectures have proven to be highly effective in capturing complex temporal dependencies and patterns in data. We also integrate popular optimization techniques such as grid search, random search, Bayesian optimization, and HyperBand algorithms into our framework to fine-tune the hyperparameters of RNN models. This ensures that models can extract the most relevant information from the data and provide accurate predictions. We evaluate our framework for short-term and long-term forecasting problems using the SWaT (Safe Water Treatment) and ETT (Electricity Transformer Temperature) datasets. These datasets reflect industrial scenarios and serve as a benchmark for comparing different forecasting approaches. The experimental results show that the proposed framework is not inferior or even superior to relevant approaches in terms of the quality of industrial data forecasting. We achieved an improvement in the quality of forecasting by more than 78% compared to the analogs to predict the parameters of the following production processes: water purification and electricity supply.
Protecting the Industrial Internet of Things against computer attacks is currently becoming an important issue. Using wavelet analysis to detect malicious intrusions into a computer network, which may be caused by computer attacks, is a rather interesting approach to solving this problem. This approach allows one to quickly detect abnormal changes in network traffic caused by security threats. The paper considers a network intrusion detection technique based on using methods of wavelet and statistical analysis. Models are proposed for statistical evaluation of wavelets selected for detecting network intrusions, for which the most preferable wavelet is selected by testing statistical hypotheses about the equality of average values, variances, and distribution laws in the reference and noisy (intrusion-prone) network traffic samples. The technique for detecting network computer attacks is based on analyzing the energy spectrum of the signal, reconstructed from the coefficients of the wavelet decomposition using the most preferable wavelet. The sensitivity of network intrusion detection to the frequency range of the reconstructed signal is estimated. Evaluation of the proposed approach based on the results of the experiments confirms its efficiency.
The objective of the study is to develop models of adversarial attacks against machine learning components of intrusion detection systems, such as the Fast Gradient Sign Method and Boiling frog attacks. The research methods consist of modeling the attack impacts in Python. For poisoning attacks, malicious traffic is mixed into the training data. Evasion attacks are modeled by adding noise. The F-measure, Precision, and Recall metrics were used for assessment of effectiveness of the detection models under attacks. The experiments were conducted on three different intrusion detection systems based on different classification models: random forest, multilayer perceptron, deep machine learning, and operant vector machine learning. As a result of the study, evasion and poisoning attacks against machine learning components were modeled. As a result of the modeling, low stability of all studied classifier models to adversarial attacks was revealed. Further research will be devoted to studying methods of protection against attacks of these classes.
The utilization of machine learning (ML) techniques for intrusion detection systems (IDS) in cybersecurity has become increasingly prevalent, demonstrating substantial advancements and effectiveness. This survey systematically reviews the use of ML techniques in IDS for cybersecurity, highlighting both advancements and associated challenges. By examining 130 recent studies, this survey systematically reviews the use of ML techniques in IDS, categorizing them into traditional ML-based, single-task deep learning (DL)-based, and multi-task DL-based approaches. Among cited works, traditional ML models like the decision tree and Gaussian mixture have achieved accuracies of 99.96
Due to the increasing number of attacks on web applications, ensuring the security of web applications is an important area in the field of information security. The article presents a combined approach for detecting attacks on web applications. The approach uses convolutional neural networks combined with multilayer perceptrons to detect single attacks. The random forest is also used to detect attacks on web applications based on a dataset that includes multiple attacks. The random forest classifies attacks based on additional parameters such as method, the content type, connection, length, and content. The test results showed that the presented approach performed better than others when detecting the same type of attacks from Kaggle datasets, and together with the modification using a random forest, it was better than others to identify multiple attacks. The proposed approach can be used to improve the security systems of web applications, which will reduce the risk of successful attacks.
When deploying complex, highly loaded end-to-end solutions, container systems are becoming an increasingly popular choice as an alternative to virtual operating systems. These technologies can significantly reduce the cost of computing resources, and also provide flexible and fast scaling of container infrastructure. However, attacks and vulnerabilities that grow proportionally to the technologies being implemented leave the question on the security of such systems open. In order to improve the efficiency of anomalous behavior detection in container systems, this paper presents an approach based on a new method for generating a training dataset based on creating and generating sequences of process histograms. The process of creating histograms is based on tracing and counting all system calls executed within a process. After the histograms are created, they undergo a normalization stage to “remember” significant information about the sequences of executed processes. After that, the generated dataset is passed to the unsupervised Autoencoder neural network model for training or subsequent detection. The process of anomalous behavior detection is based on calculating the reconstruction error of the input test data vector in relation to a given threshold value. Exceeding the specified threshold indicates the presence of abnormal values in the test data vector. The paper describes in detail the stages of creating a prototype of the proposed solution: data collection and normalization, training and detection. Estimates obtained during a number of experiments show that the proposed method demonstrates an accuracy of more than 92%.
Industrial Internet of Things (IIoT) presents a range of benefits but also introduces security vulnerabilities. This paper systematically compares different deep learning model structures for IIoT intrusion detection. Four Convolutional Neural Network (CNN) classifiers are implemented, including hybrid CNN-GRU, 1D Xception, and 1D Resnet models. The evaluation focuses on robustness and generalizability across two datasets to assess overall performance and detection rates. To handle imbalanced data, preprocessing involves PCA feature selection and hybrid resampling techniques. Furthermore, we design a distributed training process for massive datasets and continuous learning, enabling efficient large-scale processing. For the experimental evaluation, we make use of two datasets with multi-label classification tasks and analyze performance metrics of different models in terms of their detection abilities and efficiency. The proposed methods demonstrate strong performance, successfully identifying even rare attacks. In addition, we conduct a performance comparison with existing publications that utilize the same datasets, confirming that our models are on par with state-of-the-art IIoT models.
The paper proposes an intelligent methodology for user-centric privacy risk assessment that could be used in expert systems designed to assess and manage privacy risks. It is based on semantic modeling of the scenarios for using personal data and analysis of the privacy policies written in natural language. It also introduces a novel approach to calculating numerical privacy risk scores. This risk calculation allows users and organizations to understand the impact on privacy that may arise from implementation of the privacy policy specified in a text and enhance the privacy risk management procedures accordingly. Thus, the motivation of our research arises from the need for objective and clear privacy risk scores demonstrating possible impacts on the privacy of the users and potential financial losses of the organizations on the one hand, and the absence of the end-to-end methodology for assessment of privacy risks arising from the privacy policies on the other hand. The paper details the ontology construction process including the expert-based definition of the ontology competency questions and the privacy risk calculation algorithm. Examples of ontology fragments for the selected privacy policy are given. The experiments were conducted using the OPP-115 corpus of privacy policies comprising 115 privacy policies and demonstrated the applicability of the proposed approach to calculate and explain the risk score. The results of the experiments are validated by the experts.
The object of the study is a new methodological approach to assessing and predicting the level of protection of complex cyber-physical objects, as well as to the construction and application of models and methods of catastro-phe theory, as a new mathematical and methodological tool for increasing the reliability of the assessment and forecast of the level of protection of critical infrastructure objects (CIO) using the Internet of Things technology (ITT) in the interests of timely warning of danger and taking preventive measures to improve their information security. The proposed approach is based on well-known methods of catastrophe theory, in particular, on the methods of studying bifurcations that allow implementing the assessment and forecasting of the level of protection of objects of this class using algorithms for identifying and verifying the boundary and potentially catastrophic state (level) of protection with smoothly increasing changes in the parameters of external conditions, for example, smooth changes in the intensity of detected signs of computer attacks. In this case, the algorithm for assessing and predicting the security level, from the point of view of the mathematics of catastrophe theory and the theory of state spaces, is considered as an analysis of the process of transition of the security level of a critical infrastructure object from state to state. A detailed analysis of the distinctive features of this approach is made, determining the feasibility and conditions of its application for assessing and predicting a potentially dangerous, alarming level of security of CIO using ITT. A sequence of calculations and analytical expressions for calculating the estimated values of the security state (level) for various categories of signs of potential computer attacks are developed and described in detail. The results of experimental calculations are given for an example of assessing and predicting the security state (level) taking into account the intensity of receipt (detection) of various signs of computer attacks on CIO using ITT.
Today, one of the means of protecting network infrastructure from cyberattacks is intrusion detection systems. Digitalization requires the use of tools that can cope not only with known types of attacks, but also with previously undescribed ones. Machine learning can be used to protect against such threats. The paper presents models and algorithms for protecting against evasion attacks on machine learning components of intrusion detection systems. The novelty is that for the first time, a simulation of the use of a protection subsystem based on long-short-term memory autoencoders during a fast gradient sign attack was carried out. The methodology consists in simulating adversarial attacks with an assessment of the effectiveness of protection using classical metrics: accuracy, recall, F-measure. The results of the study showed the effectiveness of the proposed subsystem for protecting machine learning components of intrusion detection systems from evasion attacks. The detection indicators were restored almost to their original values.
In industrial Internet of Things (IoT) systems, explaining anomalies plays a crucial role in identifying bottlenecks and optimizing processes. This paper proposes an approach to anomaly detection using an autoencoder and its explana-tion based on the SHAP method. The purpose of the anomaly explanation is to provide a set of data features in indus-trial IoT systems that most significantly influence anomaly detection. The novelty of this approach lies in its ability to quantify the contribution of individual features for specific data samples and to calculate an average contribution across the dataset, providing a feature importance ranking. The proposed approach is tested on Industrial IoT datasets with varying feature counts and data volumes. The anomaly detection achieves an F-measure of 88-93%, outperform-ing the comparable methods discussed. The study demonstrates how explainable artificial intelligence can identify the causes of anomalies in both individual samples and datasets as a whole. The theoretical importance of the proposed approach lies in its ability to shed light on the workings of intelligent detection models, enabling the identification of factors influencing their outcomes and uncovering previously unnoticed patterns. In practice, this method enhances security system operators' understanding of ongoing processes, aiding in threat identification and error detection within data.
An approach to verification of functional and structural specifications implemented in custom integrated circuits based on invasive research methods is presented. The relevance of this research is determined by the necessity of verification of functional-structural specifications supplied by third-party implementers of hardware implementations of information security algorithms, the difficulty of detecting modifications of these algorithms and undocumented capabilities implemented at the hardware level, and the lack of uniform, universal or standardized methods for solving this problem. The mathematical formulation of the research problem is specified; its essence is to verify the equality of the values of the declared specification parameters and their values restored by the reverse engineering method. The results of the application of the verification technique of functional and structural specifications are presented using examples of its adaptation to the study of hardware-implemented DES and AES encryption algorithms. The restored functional and structural blocks of the algorithms (in particular, the substitution block) were successfully verified.
The paper considers the issues of using a decision-making model with a linguistic representation and approximation for fuzzy assessment of the reliability level to identify the signs of malicious activity in the infrastructure of an industrial Smart City. Such an approach allows one to formally linguistically describe the observed and measured signs of such suspicious activity. In its mathematical essence, it is a method for constructing a process of multi-factorial identification of malicious activity signs to make decisions on the reliability level of identifying these signs. This approach is focused on working with a linguistic representation (comparison) of variables and allows one to effectively apply linguistic analysis of these many fuzzy variables using linguistic approximation. The results of experimental calculations are described, allowing, as an example, to judge the practical applicability of the proposed mathematical mechanisms of fuzzy assessment.
Anomalies in the work of data center users can be caused by both Structured Query Language (SQL) injection attacks and user attempts to make unauthorized access to data. The paper explores various machine learning models to detect such anomalies. The peculiarity of the problem being solved is its focus on the university data centers, whose databases have a non-normalized structure. In this case, the problem of reducing the feature space arises. The paper proposes an algorithm for generating a dataset based on typing the data table names. The experimental results obtained on supervised, unsupervised and semi-supervised machine learning models confirmed the high efficiency of the proposed approach. They showed that the support vector machine, random forest, Gaussian Naive Bayes, and neural network models are the most effective in detecting known SQL injections, and the local outlier factor semi-supervised learning model is the most effective in detecting unknown SQL injections and unauthorized access attempts.
Disruption of the normal functioning of distributed computing systems due to computer attacks, the impact of malicious software and other methods of implementing unauthorized access to the information processed in them can lead to significant damage, sometimes not measurable in monetary terms. The problem of information protection in distributed computing systems related to critical infrastructure objects is especially relevant. Achieving the required level of information protection necessitates assessing the security of information in distributed computing systems at all stages of the life cycle. The paper presents an approach to quantitative assessment of the level of protection from computer attacks in in distributed computing systems, which ensures increased efficiency of security management by using a complex indicator that takes into account both the characteristics of the security breach process and the characteristics of the security process, as well as the use of a security graph that takes into account the real structure of in distributed computing system.
This paper presents an intelligent system to automate the process of detecting and analysing cyber-attacks using Suricata logs, graph neural networks (GNNs), and large language models (LLMs). The proposed approach is based on several key components: collecting and preprocessing network events from Suricata, building an ontological model of attacks using MITRE ATT&CK, applying graph neural networks to identify relationships between events, and finally integrating a language model for dialogue interaction with the operator and generating attack hypotheses. Experimental results demonstrate high accuracy in detecting anomalous network patterns and operator friendliness and indicate the potential for further development of the system for use in high-load and distributed infrastructures.
The Internet of Things (IoT) has revolutionised technology within intelligent urban environments; however, this has concurrently given rise to security and privacy risks, including the proliferation of various types of malware, which can lead to detrimental consequences. This paper presents a GAN-inspired approach for the classification of malware imagery, employing an autoencoder (AE) as the synthetic data generator and leveraging transfer learning for the discriminator. This framework is designed to identify various malware threats that target IoT networks through the use of RGB images collected directly from malware samples. The generator is specifically constructed for effective data reconstruction, incorporating different AE architectures and denoising techniques, while the discriminator utilises a pre-trained convolutional neural network (CNN)-based model to maximise performance. Furthermore, to address data imbalance in the multi-label classification task, we introduced a self-adjustive oversampling technique to augment the sample volume from minority classes. The proposed method was evaluated on several multi-label malware-based imagery datasets to assess its robustness. Comparative performance analysis was conducted using well-established image classification models, including VGG19, MobileNet, and Xception, which were integrated into the discriminator model as a pre-trained block. The results demonstrate that the variational AE-GAN is highly implementable and scalable for the malware classification task, exhibiting commendable detection performance and generalisability.
Представлен подход к верификации функционально-структурных спецификаций, реализованных в заказных интегральных схемах, основанный на инвазивных методах исследования. Актуальность проведённого исследования обусловлена необходимостью проведения верификации функционально-структурных спецификаций, поставляемых сторонними исполнителями аппаратных реализаций алгоритмов обеспечения информационной безопасности, сложностью выявления на аппаратном уровне модификаций этих алгоритмов и внедрённых в них недокументированных возможностей и отсутствием единых универсальных или стандартизированных методов решения этой задачи. Сформулирована математическая постановка задачи исследования, суть которой состоит в проверке равенства значений параметров заявленной спецификации с их значениями, восстановленными методом обратного проектирования. Представлены результаты применения предложенного подхода к верификации функционально-структурных спецификаций на примерах аппаратно-реализованных алгоритмов шифрования DES и AES. Восстановленные функционально-структурные блоки алгоритмов (в частности – блок подстановок) были успешно верифицированы.
The growing complexity of cyber threats requires innovative machine learning techniques, and imagebased malware classification opens up new possibilities. Meanwhile, existing research has largely overlooked the impact of noise and obfuscation techniques commonly employed by malware authors to evade detection, and there is a critical gap in using noise simulation as a means of replicating real-world malware obfuscation techniques and adopting denoising framework to counteract these challenges. This study introduces an image denoising technique based on a U-Net combined with a GAN framework to address noise interference and obfuscation challenges in image-based malware analysis. The proposed methodology addresses existing classification limitations by introducing noise addition, which simulates obfuscated malware, and denoising strategies to restore robust image representations. To evaluate the approach, we used multiple CNN-based classifiers to assess noise resistance across architectures and datasets, measuring across two multi-class public datasets, MALIMG and BIG-15. For example, the MALIMG classification accuracy improved from 23.73% to 88.84% with denoising applied after Gaussian noise injection, demonstrating robustness. This approach contributes to improving malware detection by offering a robust framework for noise-resilient classification in noisy conditions.
The article explores an effective approach to detecting anomalies in cryptocurrency transactions using neural network models, including convolutional, deep, and gated recurrent units (GRUs), and compares their performance with other existing methods for identifying illicit transactions in cryptocurrency networks. A research of relevant studies is con-ducted in the fields of transaction data analysis in the cryptocurrency network, data visualization for transaction anal-ysis, and the use of computer vision techniques for detecting anomalous behavior. The subject area of the study is de-fined. The problem of detecting anomalies in cryptocurrency transactions is based on the fact that these transactions are pseudonymous, i.e. there are no direct indications of the identity of the sender and recipient. The relevance and contribution of this work lie in the development of a method capable of identifying anomalous transactions with high accuracy in near real-time. Experimental studies were conducted using a dataset of cryptocurrency transactions, apply-ing both neural and non-neural classifiers. The results are compared against existing approaches in the field. The exper-iments demonstrated that gated recurrent units outperformed other neural models in this task, achieving an accuracy of 0.94, precision of 0.95, recall of 0.93, and F1-score of 0.94, indicating the high effectiveness of the proposed model. Nonetheless, this approach showed slightly lower performance compared to traditional machine learning algorithms, such as optimized distributed gradient boosting. The novelty of the proposed approach lies in its use of statistical char-acteristics derived from the transaction graph, combined with deep learning and gradient boosting techniques. The ap-proach can be applied in the development of software tools for detecting illicit cryptocurrency transactions within in-formation security systems and digital forensics.