Accurate cloud resource forecasting is essential for proactive resource provisioning, maintaining Quality of Service (QoS), and reducing operational costs in dynamic cloud environments. The existing forecasting approaches predominantly estimate future CPU workload directly from historical resource traces, which often overlook the relationship between customer service demand and subsequent resource consumption. This study proposes a two-stage integrated forecasting model that explicitly models this dependency by first forecasting customer service requests, expressed as Transactions Per Second (TPS), and subsequently estimating future CPU workload from the TPS forecast. Both the forecasting component and resource prediction component employed the XGBoost model within a cascaded learning architecture, complemented by adaptive online retraining using an expanding-window strategy to address concept drift in continuously evolving cloud workloads. The proposed work was evaluated using real-world traces collected from a private cloud environment comprising ten applications. Experimental results demonstrate robust forecasting performance by achieving Symmetric Mean Absolute Percentage Error (SMAPE) below 7% for most applications, with the best-performing application achieving an MAE of 0.7372, RMSE of 1.1866, SMAPE of 3.57%, and an R2 of 0.9185. Horizon-wise drift analysis confirmed stable recursive forecasting behavior with controlled error accumulation across a 60-step prediction horizon. Compared with the conventional direct CPU forecasting method, the proposed two-stage integrated model gives improved forecasting robustness, computational efficiency, and interpretability, making it well-suited for proactive resource management and intelligent auto-scaling in cloud computing environments.
Variable workloads and dynamic Quality of Service (QoS) demands make it difficult to manage cloud resources efficiently, especially in terms of CPU utilization. Over-provisioning of cloud resources leads to unnecessary energy consumption and costs, while under-provisioning decreases QoS and affects customer satisfaction. This study explores lightweight machine learning models to predict CPU utilization in cloud environments based on the services requested by cloud customers. Previous works mainly focus on traditional time-series-based CPU workload prediction, mostly using deep learning models, and did not address the problem of future CPU workload prediction for heterogeneous applications running in cloud environments. This study presents a predictive framework for CPU usage across multiple applications running in a private cloud. The proposed framework consists of several steps, including pre-processing, handling missing values and standardizing data, refining the target variable (CPU) through a rolling window approach, and evaluating various lightweight machine learning models (Bayes Ridge, Elastic Net, Lasso, Ridge, and MLP) on a dataset with 10 heterogeneous cloud applications. The experimental results demonstrate that the Ridge model shows robust performance in predicting CPU loads for most of the applications compared to the rest of the models, supporting intelligent resource management in cloud computing environments.
We evaluate Forward Error Correction (FEC) codes in the context of a novel routing protocol HDARP+ for airborne networks. HDARP+ uses directional antennas and dynamic FEC coding to avoid detection by adversaries. The use of FEC coding is dynamic in the sense that different FEC codes, or no FEC code, will be used depending on the relative position of friendly and adversary aircraft. Due to the real-time restrictions in airborne networks, encoding and decoding must be fast and done through table lookup. Since we use table lookup, the FEC codes must be short. We evaluate two types of short FEC codes: Reed-Solomon (RS) codes, and FEC codes found using greedy search (called GS codes). The results show that the RS codes are better than the GS codes at handling error bursts. However, the GS codes are more flexible when it comes to finding attractive trade-offs between the code's ability to increase the number of the cases when hostile detection can be avoided (related to the coding gain), the code rate and the amount of memory required for implementing lookup tables.
With the rapid growth of internet technologies, IT businesses are transferring to cloud-based systems, and cloud-based services are in high demand among internet users. Therefore, appropriate allocation of resources in cloud computing environments is essential. The companies can reduce costs by saving energy by dynamically scaling up or down the number of active servers. In this context, this study presents a machine learning-based model for accurate prediction of CPU utilization. Previous studies employed timestamp-based data to predict CPU utilization in cloud computing, while the proposed work uses incoming user requests to predict CPU workload so that a timely decision can be made to scale up or scale down the servers in a cloud computing environment. The proposed model is based on several machine learning algorithms that are stacked into a single model called the stacking model for CPU workload prediction. The effectiveness of the proposed stacking model was tested on several evaluation metrics to validate its performance. Furthermore, the performance of the proposed stacking model is also compared with other state-of-the-art machine learning models such as support vector machines (SVM), decision trees (DT), random forests (RF), gradient boosting, and extreme gradient boosting (XGBoost).
We present a method, including tool support, for bibliometric mining of trends in large and dynamic research areas. The method is applied to the machine learning research area for the years 2013 to 2022. A total number of 398,782 documents from Scopus were analyzed. A taxonomy containing 26 research directions within machine learning was defined by four experts with the help of a Python program and existing taxonomies. The trends in terms of productivity, growth rate, and citations were analyzed for the research directions in the taxonomy. Our results show that the two directions, Applications and Algorithms, are the largest, and that the direction Convolutional Neural Networks is the one that grows the fastest and has the highest average number of citations per document. It also turns out that there is a clear correlation between the growth rate and the average number of citations per document, i.e., documents in fast-growing research directions have more citations. The trends for machine learning research in four geographic regions (North America, Europe, the BRICS countries, and The Rest of the World) were also analyzed. The number of documents during the time period considered is approximately the same for all regions. BRICS has the highest growth rate, and, on average, North America has the highest number of citations per document. Using our tool and method, we expect that one could perform a similar study in some other large and dynamic research area in a relatively short time.
Using a novel method and tool in the form of a Python program, we present a bibliometric study based oil 46,937 documents related to smart cities from the Scopus database. The study identifies important research directions and trends during the time period 2014 to 2023. We also present the growth of smart city research for five geographic regions. Citation analysis for research directions and regions is also performed. The results show that smart city research in general stopped growing around 2019. However, some research directions are still growing, e.g., smart city research related to machine learning and AI. India is the only geographic region where smart city research still is growing. We also see that the number of citations of a smart city document from North America is on average a factor 3.74 larger than the number of citations to a document from India.
Here we present a novel routing protocol HDARP+ for airborne tactical networks that use directional antennas. HDARP+ extends the existing protocol HDARP (Hostile-Direction Aware Routing Protocol) by reducing the risk for detection by adversary aircraft even further. Compared to HDARP, the extension in HDARP+ introduces dynamic Forward Error Correction (FEC) coding. The FEC code is dynamic in the sense that different FEC codes, or no FEC code, will be used depending on the relative position of the receiver and adversary aircraft. We evaluate three different Reed-Solomon FEC codes based on three criteria: the ability to transmit in the presence of adversaries without being detected, the reduction of the effective communication bandwidth, and the implementation cost in terms of the sizes of lookup tables for encoding and decoding. We argue that (variations of) HDARP+ will be implemented in future airborne tactical networks. This paper was originally presented at the NATO Science and Technology Organization Symposium (ICMCIS) organized by the Information Systems Technology (IST) Panel, IST-205-RSY - the ICMCIS, held in Koblenz, Germany, 23–24 April 2024.
Contemporary airborne radio networks are usually implemented using omnidirectional antennas. Unfortunately, such networks suffer from disadvantages such as easy detection by hostile aircraft and potential information leakage. In this paper, we present a novel mobile ad hoc network (MANET) routing protocol based on directional antennas and situation awareness data that utilizes adaptive multihop routing to avoid sending information in directions where hostile nodes are present. Our protocol is implemented in the OMNEST simulator and evaluated using two realistic flight scenarios involving 8 and 24 aircraft, respectively. The results show that our protocol has significantly fewer leaked packets than comparative protocols, but at a slightly higher cost in terms of longer packet lifetime.
In this paper a program and methodology for bibliometric mining of research trends and directions is presented. The method is applied to the research area Big Data for the time period 2012 to 2022, using the Scopus database. It turns out that the 10 most important research directions in Big Data are Machine learning, Deep learning and neural networks, Internet of things, Data mining, Cloud computing, Artificial intelligence, Healthcare, Security and privacy, Review, and Manufacturing. The role of Big Data research in different fields of science and technology is also analysed. For four geographic regions (North America, European Union, China, and The Rest of the World) different activity levels in Big Data during different parts of the time period are analysed. North America was the most active region during the first part of the time period. During the last years China is the most active region. The citation scores for documents from different regions and from different research directions within Big Data are also compared. North America has the highest average citation score among the geographic regions and the research direction Review has the highest average citation score among the research directions. The program and methodology for bibliometric mining developed in this study can be used also for other large research areas. Now that the program and methodology have been developed, it is expected that one could perform a similar study in some other research area in a couple of days.
During the last decade, we have witnessed a rapid development of extended reality (XR) technologies such as augmented reality (AR) and virtual reality (VR). Further, there have been tremendous advancements in artificial intelligence (AI) and machine learning (ML). These two trends will have a significant impact on future digital societies. The vision of an immersive, ubiquitous, and intelligent virtual space opens up new opportunities for creating an enhanced digital world in which the users are at the center of the development process, so-called intelligent realities (IRs). The “Human-Centered Intelligent Realities” (HINTS) profile project will develop concepts, principles, methods, algorithms, and tools for human-centered IRs, thus leading the way for future immersive, user-aware, and intelligent interactive digital environments. The HINTS project is centered around an ecosystem combining XR and communication paradigms to form novel intelligent digital systems. HINTS will provide users with new ways to understand, collaborate with, and control digital systems. These novel ways will be based on visual and data-driven platforms which enable tangible, immersive cognitive interactions within real and virtual realities. Thus, exploiting digital systems in a more efficient, effective, engaging, and resource-aware condition. Moreover, the systems will be equipped with cognitive features based on AI and ML, which allow users to engage with digital realities and data in novel forms. This paper describes the HINTS profile project and its initial results.
The availability of large amounts of data in combination with Big Data analytics has transformed many application domains. In this paper, we provide insights into how the area has developed in the last decade. First, we identify seven major application areas and six groups of important enabling technologies for Big Data applications and systems. Then, using bibliometrics and an extensive literature review of more than 80 papers, we identify the most important research trends in these areas. In addition, our bibliometric analysis also includes trends in different geographical regions. Our results indicate that manufacturing and agriculture or forestry are the two application areas with the fastest growth. Furthermore, our bibliometric study shows that deep learning and edge or fog computing are the enabling technologies increasing the most. We believe that the data presented in this paper provide a good overview of the current research trends in Big Data and that this kind of information is very useful when setting strategic agendas for Big Data research.
Advances in 5G and the Internet of Things (IoT) have to cater to the diverse and varying needs of different stakeholders, devices, sensors, applications, networks, and access technologies that come together for a dedicated IoT network for a synergistic purpose. Therefore, there is a need for a solution that can assimilate the various requirements and policies to dynamically and intelligently orchestrate them in the dedicated IoT network. Thus we identify and describe a representative industry-relevant use case for such a smart and adaptive environment through interviews with experts from a leading telecommunication vendor. We further propose and evaluate candidate architectures to achieve dynamic and intelligent orchestration in such a smart environment using a systematic approach for architecture design and by engaging six senior domain and IoT experts. The candidate architecture with an adaptive and intelligent element ("Smart AAA agent") was found superior for modifiability, scalability, and performance in the assessments. This architecture also explores the enhanced role of authentication, authorization, and accounting (AAA) and makes the base for complete orchestration. The results indicate that the proposed architecture can meet the requirements for a dedicated IoT network, which may be used in further research or as a reference for industry solutions.
Network anomaly detection for critical infrastructure supervisory control and data acquisition (SCADA) systems is the first line of defense against cyber-attacks. Often hybrid methods, such as machine learning with signature-based intrusion detection methods, are employed to improve the detection results. Here an attempt is made to enhance the support vector-based outlier detection method by leveraging behavioural attribute extension of the network nodes. The network nodes are modeled as graph vertices to construct related attributes that enhance network characterisation and potentially improve unsupervised anomaly detection ability for SCADA network. IEC 104 SCADA protocol communication data with good domain fidelity is utilised for empirical testing. The results demonstrate that the proposed approach achieves significant improvements over the baseline approach (average F_1 score increased from 0.6 to 0.9, and Matthews correlation coefficient (MCC) from 0.3 to 0.8). The achieved outcome also surpasses the unsupervised scores of related literature. For critical networks, the identification of attacks is indispensable. The result shows an insignificant missed-alert rate ( 0.3% on average), the lowest among related works. The gathered results show that the proposed approach can expose rouge SCADA nodes reasonably and assist in further pruning the identified unusual instances.
Large amount of data are generated from in-situ monitoring of additive manufacturing (AM) processes which is later used in prediction modelling for defect classification to speed up quality inspection of products. A high volume of this process data is defect-free (majority class) and a lower volume of this data has defects (minority class) which result in the class-imbalance issue. Using imbalanced datasets, classifiers often provide sub-optimal classification results, i.e. better performance on the majority class than the minority class. However, it is important for process engineers that models classify defects more accurately than the class with no defects since this is crucial for quality inspection. Hence, we address the class-imbalance issue in manufacturing process data to support in-situ quality control of additive manufactured components. For this, we propose cluster-based adaptive data augmentation (CADA) for oversampling to address the class-imbalance problem. Quantitative experiments are conducted to evaluate the performance of the proposed method and to compare with other selected oversampling methods using AM datasets from an aerospace industry and a publicly available casting manufacturing dataset. The results show that CADA outperformed random oversampling and the SMOTE method and is similar to random data augmentation and cluster-based oversampling. Furthermore, the results of the statistical significance test show that there is a significant difference between the studied methods. As such, the CADA method can be considered as an alternative method for oversampling to improve the performance of models on the minority class.
The power grid is a build-up of a mesh of thousands of sensors, embedded devices, and terminal units that communicate over different media. The heterogeneity of modern and legacy equipment calls for attention towards diverse network security measures. The critical infrastructure employs different security measures to detect and prevent adversaries, e.g., through signature-based tools. These approaches lack the potential to identify unknown attacks. Machine learning has the prospective to address novel attack vectors. This paper systematically evaluates the efficacy of learning algorithms from different families for intrusion detection in IEC 60870-5-104 protocol. One-class SVM and k-Nearest Neighbour unsupervised learning models show small potential when being tested on the IEC 104 unseen dataset with Area Under the Curve score 0.64 and 0.59, in the same order; and Matthews Correlation Coefficient value 0.3 and 0.2, respectively. The experimental results suggest little feasibility of the evaluated unsupervised learning approaches for anomaly detection in IEC 104 communication and recommend coupling it with other anomaly detection techniques.
This paper aims to address data labelling issues in process data to support in-situ process monitoring of additive manufactured components. For this, we adopted an active learning (AL) approach to minimise the manual effort for data labelling for classification models. In this study, we present an approach that utilises pre-trained models to extract deep features from images, and clustering and query by committee sampling to select the representative samples to build defect classification models. We conduct quantitative experiments to evaluate the proposed method’s performance and compare it with other selected state-of-the-art AL approaches using a dataset of additive manufacturing (AM) and a publicly available dataset. The experimental results show that the proposed approach outperforms AL with committee based sampling, and AL with clustering and random sampling. The results of the statistical significance test show that there is a significant difference between the studied AL approaches. Hence, the proposed AL approach can be considered an alternative method to reduce labelling costs when building defects classification models, whose generalizability is most likely plausible.
Context: In the light of the swift and iterative nature of Agile Software Development (ASD) practices, establishing deeper insights into capability measurement within the context of team formation is crucial, as the capability of individuals and teams can affect team performance and productivity. Although a former Systematic Literature Review (SLR) synthesized the state of the art in relation to capability measurement in ASD with a focus on selecting individuals to agile teams, and capabilities related to team performance and success, determining to what degree the SLR's results apply to practice can provide progressive insights to both research and practice. Objective: Our study investigates how agile practitioners perceive the relevance of individual and team level measures for characterizing the capability of an agile team and its members. Furthermore, to scrutinize variations in practitioners' perceptions, our study further analyzes perceptions across stratified demographic groups. Method: We undertook a Web-based survey using a questionnaire built based on the capability measures identified from a previously conducted SLR. Results: Our survey responses (60) indicate that 127 individual and 28 team capability measures were considered as relevant by the majority of practitioners. We also identified seven individual and one team capability measure that have not been previously characterized by our SLR. The surveyed practitioners suggested that an agile team member's responsibility and questioning skills significantly represent the member's capability. Conclusion: Results from our survey align with our SLR's findings. Measures associated with social aspects were observed to be dominant compared to technical and innovative aspects. Our results can support agile practitioners in their team composition decisions.
In railway traffic systems, it is essential to achieve a high punctuality to satisfy the goals of the involved stakeholders. Thus, whenever disturbances occur, it is important to effectively reschedule trains while considering the perspectives of various stakeholders. This typically involves solving a multi-objective train rescheduling problem, which is much more complex than its single-objective counterpart. Solving such a problem in real time for practically relevant problem sizes is computationally challenging. The reason is that the rescheduling solution(s) of interest are dispersed across a large search tree. The tree needs to be navigated fast while pruning off branches leading to undesirable solutions and exploring branches leading to potentially desirable solutions. The use of parallel computing enables such a fast navigation of the tree. This article presents a heuristic parallel algorithm to solve the multi-objective train rescheduling problem. The parallel algorithm combines a depth-first search with simultaneous breadth-wise tree exploration while searching the tree for solutions. An existing parallel algorithm for single-objective train rescheduling has been redesigned, primarily, by (i) pruning based on multiple metrics, and (ii) maintaining a set of upper bounds. The redesign improved the quality of the obtained rescheduling solutions and showed better speedups for several disturbance scenarios.