Every company conducts evaluations to ensure the quality of its products and services, often utilizing multivariate simultaneous control charts to monitor the process mean and variability concurrently. The objective of this study is to overcome a significant limitation in the Maximum Half-Normal Multivariate Control Chart (Max-Half-Mchart): its vulnerability to outliers, which can trigger masking and swamping effects and lead to inaccurate process monitoring. The primary scientific contribution is the development of two robust versions of the Max-Half-Mchart by integrating the fast minimum covariance determinant (Fast-MCD) and deterministic minimum covariance determinant (Det-MCD) estimators into the chart’s statistical framework. The evaluation criteria for these methods include the average run length (ARL) to assess process shift detection speed and classification accuracy, false positive (FP) rate, false negative (FN) rate, and area under the curve (AUC) to measure outlier detection performance. Simulation results indicate that, while both robust charts effectively detect process shifts, the Det-MCD-based robust Max-Half-Mchart is particularly superior for lower contamination levels (5–20%), whereas the Fast-MCD-based chart performs best at higher contamination levels (30%). An illustrative application to ordinary Portland cement (OPC) quality data confirms the practical superiority of the Det-MCD approach, which detected six out-of-control signals compared with only two identified by conventional methods. These results suggest that the proposed robust charts are highly sensitive tools for maintaining quality in the presence of contaminated data.
Control charts are widely used in the industrial world to monitor the average and variability of production processes. Max-Half-Mchart is a multivariate control chart that is not particularly effective in handling many outliers. This research aims to develop a control chart that is more resistant to outliers by using Minimum Regularized Covariance Determinant (MRCD). MRCD is a development of the MCD method, which is better at dealing with 'fat data," namely, situations in which the number of variables is greater than the number of observations. The performance of a robust Max-Half-Mchart control chart based on MRCD was evaluated using the average run length (ARL) against shifts in the process mean, process variance, and simultaneous shifts. A comparison was also made of the outlier detection accuracy between the robust Max-Half-Mchart based on MRCD and the standard Max-Half-Mchart. Simulation results demonstrated that the MRCD-based robust chart is most sensitive to simultaneous shifts in the mean and variance, significantly outperforming the conventional method in "de-masking" process deviations. The robust framework maintains higher accuracy and AUC levels even at extreme contamination stages of 30% to 40% outliers, where traditional charts typically fail. A practical application to cement quality data further substantiated these findings, as the robust chart successfully identified 14 out-of-control signals (comprising the mean, variability, and simultaneous shifts), whereas the conventional chart detected none. These results indicate that the MRCD-based Max-Half-Mchart offers a more reliable and responsive quality monitoring system for complex industrial datasets.
Modern manufacturing systems generate increasingly complex and high-dimensional data that pose significant challenges for traditional pattern recognition and process monitoring. Although Support Vector Data Description (SVDD) is a robust one-class classification technique for high-dimensional data, it can be sensitive to volatile process shifts. To address this, this study proposes a novel D'AEWMA control chart that integrates the SVDD with an Adaptive Exponentially Weighted Moving Average (AEWMA) mechanism. By dynamically adjusting the smoothing parameters based on the magnitude of the error, the D'AEWMA chart effectively identifies process mean shifts across varying magnitudes while minimizing false alarms. Performance evaluation using Average Run Length (ARL) metrics demonstrates that the D'AEWMA chart slightly outperforms the traditional D'EWMA chart, particularly in detecting small shifts and when the correlation between quality characteristics is low. Application to synthetic datasets and real-world clinker quality characteristics confirmed the proposed chart’s superior sensitivity and its ability to maintain a stable in-control state without triggering premature false alarms.
Forecasting train passenger demand is essential for supporting strategic decision-making and optimizing resource allocation in the transportation industry. This study aimed to develop a predictive model for the number of passengers on the Surabaya-Jakarta train route using the Extreme Gradient Boosting (XGBoost) algorithm. Owing to the non-linear nature of the historical count time series data (January 2019 to December 2024) and the significant disruptive impact of the COVID-19 pandemic, traditional linear models such as ARIMA were considered less appropriate. To optimize the XGBoost model, we comparatively evaluated two distinct input approaches: significant Partial Autocorrelation Function (PACF) lag and the sliding window method. Hyperparameter tuning was conducted via grid search, and the models were rigorously evaluated using Time Series Cross-Validation to prevent information leakage. Furthermore, the study compared recursive and direct multi-step forecasting strategies to project passenger volumes for the next 12 months. The analysis revealed that the sliding window approach with a window size of 4 yielded the best performance on the testing data, achieving a Mean Absolute Percentage Error (MAPE) of 10.94% and significantly outperforming the PACF lag method, which was prone to overfitting. Additionally, recursive forecasting is more rational and effective at capturing complex seasonal patterns and short-term fluctuations than direct forecasting. The final 12-month projection for 2025 indicates clear seasonal fluctuations, with a low in March (10,319 passengers) and a peak in November (20,932 passengers), providing a data-driven foundation for the train company to proactively optimize capacity planning, operational scheduling, and human resource management in the future.
Statistical Process Control (SPC) often struggles to efficiently monitor both minor and major process shifts within a single framework. To address this challenge, this study introduces a Dynamic Huber-Weighted Adaptive Exponentially Weighted Moving Average Max-Multivariate (AEWMA Max-M) statistic designed for simultaneous multivariate process monitoring. By integrating a Huber-based scoring function, the proposed scheme dynamically adjusts its smoothing parameters based on the magnitude of observed deviations. For subtle changes, the framework heavily weights historical data to enhance sensitivity, whereas for significant errors, it reduces the influence of past data to rapidly detect major structural shifts. The diagnostic performance of this chart was rigorously evaluated using Average Run Length (ARL) simulations. The findings demonstrate that the proposed dynamic framework systematically outperforms the traditional EWMA Max-M and basic Max-M models in efficiently pinpointing isolated and concurrent shifts in the process mean and variance. Furthermore, the practical efficacy of the AEWMA Max-M scheme was validated using multivariate industrial quality data from cement clinker production across five distinct parameters. The empirical application confirmed its superior responsiveness, successfully flagging a maximum of ninety-eight distinct out-of-control instances under high-sensitivity configurations. Ultimately, this adaptive statistic provides a robust and versatile tool for complex quality surveillance.
This study addresses the challenge of frequency mismatches between covariates and responses, known as mixed-frequency data, by using low-frequency variable information to predict high-frequency GEV-distributed variables with the GEVReg-MIDAS model. Furthermore, using an unlimited number of covariates increases computational load and model complexity. To address this issue, a covariate selection process based on Bayesian principles is integrated into the GEVReg-MIDAS model to form the GEVReg-MIDAS-SSVS. This proposed integrated model is tested on two classification datasets of clothing import prices, designed to produce price ranges (interval forecasts) as an early warning system for tax avoidance. The results show that the GEVReg-MIDAS-SSVS model outperforms the GEVReg-MIDAS model without feature selection, as evaluated by the ES and QRMSE metrics. GEVReg-MIDAS-SSVS also demonstrates the flexibility of the information contained in the model, where each modeling mismatch is more informative, representing the influence of its covariates. The import price range forecasting results indicate a price range that aligns with the training and testing data intervals. The results of this study can serve as a basis for policy-making and for methods to be applied in the evaluation of imported goods prices.
This study investigates the utilization of Google Maps reviews to assess hospital service quality. Patient-generated reviews were analyzed using a sentiment analysis framework incorporating the Bidirectional Encoder Representations from Transformers (BERT) classification model. The p control chart was employed to monitor the distribution of negative sentiment. The results of the sentiment analysis revealed a predominance of positive reviews over negative ones. The BERT classifier achieved excellent performance, with AUC values of 99.95% and 93.72% for training and testing data, respectively. However, the p control chart indicated that the hospital's performance still requires improvement, as several observations fell outside the statistically controlled range. Common patient complaints centered on lengthy wait times and queues, highlighting areas for targeted quality enhancement initiatives. This research demonstrates the potential of leveraging patient feedback to inform hospital quality improvement efforts.
Ensuring product quality in manufacturing is crucial for customer satisfaction and cost minimization. Traditional Max-M control charts, while foundational for multivariate individual data, are less sensitive to small process shifts. This study introduces an enhanced Exponentially Weighted Moving Average Max Multivariate (EWMA Max-M) control chart, designed to improve the early detection of process anomalies. By integrating the EWMA process with the Max-M chart, the proposed methodology performs better in detecting small shifts, as demonstrated through Average Run Length (ARL) comparisons. Key factors such as the number of quality characteristics (p), the correlation between characteristics (ρ), and the mean vector shift (λ) were considered. Simulation and real-world applications, including cement data analysis, confirm the effectiveness of the EWMA Max-M chart in enhancing process monitoring.
Indonesia, as the largest Muslim-majority country, has significant potential to enhance its Shariah financial sector, which has been growing rapidly, around 7.43% from 2023 to 2024, and contributing to the national economy. However, political and natural disasters have influenced the economy and Shariah-compliant stocks. This study focuses on forecasting Shariah-compliant stock prices using Generalized Autoregressive Conditional Heteroscedasticity (GARCH) models and estimating investment risks via Value at Risk (VaR) for four Islamic banks listed in IDX: BRIS, BTPS, BANK, and PNBS. The findings indicate that GARCH models effectively capture stock price dynamics and provide accurate 10-day forecasts. Additionally, the models reliably predict VaR, validated through backtesting at various confidence levels. These insights are valuable for financial regulators and risk managers, aiding in policy design to ensure market stability by enabling the implementation of measures such as stricter capital reserve requirements for institutions with high-risk exposure and mandatory adoption of advanced risk management techniques like dynamic stress testing. Such policies not only mitigate systemic risks during periods of financial volatility but also enhance the overall resilience and robustness of the financial system. For investors, accurate risk predictions support informed decision-making, enhance portfolio protection, and optimize risk management.
Electricity load forecasting is crucial for effective energy management, particularly in minimizing energy production and distribution costs. Traditional models like SARIMA and Singular Spectrum Analysis (SSA) have been widely used but often need to capture complex nonlinear patterns and deal with data uncertainties. This study aims to develop and evaluate a hybrid forecasting model that combines the Prophet model with a Neural Network Autoregressive (Prophet-NAR) model, referred to as PropNAR. The objective is to enhance the accuracy of hourly electricity load forecasting in Malaysia by addressing the limitations of existing models. The proposed PropNAR integrated the strengths of the Prophet model in capturing deterministic structures, such as trends, seasonality, and holiday effects with the NAR model’s ability to handle nonlinear stochastic relationships. Additionally, SSA-based bagging is employed to manage data uncertainties, and ensemble techniques are applied to further refine the forecasting accuracy. The hybrid PropNAR model demonstrated significant improvements in forecasting accuracy. Specifically, it reduced Mean Absolute Error (MAE), Root Mean Square Error (RMSE), and Mean Absolute Percentage Error (MAPE) by approximately 21%-86% compared to the standalone Prophet model. On the Malaysian electricity load datasets, the PropNAR model achieved MAE values ranging from 527.26 to 1023.78, RMSE values from 752.70 to 1498.54, and MAPE values from 1.12% to 2.04%. These results indicate a substantial enhancement over SARIMA and SSA-NAR in handling outliers and data variability. The proposed hybrid PropNAR model offers a robust solution for Malaysian short-term electricity load forecasting, outperforming conventional models in accuracy and reliability.
This study proposes novel framework to enhance statistical process control (SPC) of water quality by addressing the pervasive issue of autocorrelation in time-series data. We investigate the characteristics of pH, turbidity, and KMnO₄ in Surabaya city's water, revealing significant autocorrelation that compromises statistical independence assumption crucial for reliable SPC. To overcome this, Generative Adversarial Network (GAN) model was developed to generate decorrelated residual time-series. The efficacy of GAN model in reducing autocorrelation was quantitatively validated, achieving Mean Squared Error (MSE) of 0.0054, Root Mean Squared Error (RMSE) of 0.0738, and Mean Absolute Error (MAE) of 0.0556. Subsequently, these GAN-derived residuals were integrated into Multivariate Exponentially Weighted Moving Average (MEWMA) control chart for process monitoring. Phase I analysis detected 33 out-of-control signals; after identifying and removing outliers, process was brought under statistical control with no further out-of-control signals detected. However, subsequent Phase II online monitoring detected eight statistically significant out-of-control signals, indicating a potential loss of process stability over time. Our findings underscore the significant utility of GAN-based residual analysis as a robust strategy for mitigating autocorrelation effects in environmental water quality data. This approach leads to improved process monitoring and enables early anomaly detection, crucial for proactive water quality management.
PM2.5 air pollution poses significant health risks, particularly in urban areas such as Jakarta, where concentrations frequently surpass acceptable levels due to rapid urbanization. This study addresses autocorrelation in air quality data and evaluates the monitoring performance of XGBoost and Support Vector Regression (SVR) models using Individual and Exponentially Weighted Moving Average (EWMA) Charts. PM2.5 levels were obtained from Jakarta's Air Quality Index. The findings reveal that the SVR model effectively manages autocorrelation, while the combination of XGBoost and the EWMA chart yielded superior monitoring performance. Specifically, this approach detected only one out-of-control (OOC) point in Phase II and none in Phase I, with identified shifts ranging from moderate to large. Overall, the XGBoost and EWMA chart integration offers a robust solution for precise air quality monitoring and minimizes false alarms. The identification of OOC points provides actionable insights by highlighting significant deviations in air quality data that may require immediate intervention.Key points: • SVR and XGBoost model regression was introduced to enhance forecasting accuracy. • EWMA chart based on XGBoost residuals has better monitoring results.
Abstract In recent years, modern manufacturing systems have become increasingly complex and generate a wide variety of data. This poses challenges for finding patterns and effective monitoring processes. One-class classification (OCC) is a machine learning technique suitable for finding data patterns. OCC algorithms learn a model of normal process behavior and flag deviations. This helps to detect outliers and process shifts. Support Vector Data Description (SVDD), an OCC type, was used to monitor high-dimensional data. It constructs a hypersphere around normal data and flag outliers. However, SVDD is sensitive to volatile process shifts. Adaptive Exponentially Weighted Moving Average (AEWMA) control charts address this issue. They dynamically adjust the control limits based on the process-mean changes. This study proposes a new AEWMA-based SVDD control chart (D'AEWMA) for monitoring multivariate processes. The D'AEWMA chart outperformed the traditional D'EWMA chart in detecting out-of-control signals and minimizing false alarms.
Intrusion detection systems (IDS) are crucial in safeguarding network security by identifying unauthorized access attempts through various techniques. Statistical Process Control (SPC), particularly Hotelling’s T2 control charts, is noted for monitoring network traffic against known attack patterns or anomaly detection. This research advances the domain by incorporating robust statistical estimators—namely, the Fast-MCD and MRCD (Minimum Regularized Covariance Determinant) estimators—into bootstrap-enhanced Hotelling’s T2 control charts. These enhanced charts aim to strengthen detection accuracy by offering improved resistance to outlier contamination, a prevalent challenge in intrusion detection. The methodology emphasizes the MRCD estimator’s robustness in overcoming the limitations of traditional T2 charts, especially in environments with a high incidence of outliers. Applying the proposed bootstrap-based robust T2 charts to the UNSW-NB15 dataset illustrates a marked enhancement in intrusion detection performance. Results indicate superior performance of the proposed method over conventional T2 and Fast-MCD-based T2 charts in detection accuracy, even in varied levels of outlier contamination. Despite increasing execution time, the precision and reliability in detecting intrusions present a justified trade-off. The findings underscore the significant potential of integrating robust statistical methods to enhance IDS effectiveness.
In this work, the mixed multivariate T2 control chart’s detailed performance evaluation based on PCA mix is explored. The control limit of the proposed control chart is calculated using the kernel density approach. Through simulation studies, the proposed chart’s performance is assessed in terms of its capacity to identify outliers and process shifts. When 30% more outliers are included in the data, the proposed chart provides a consistent accuracy rate for identifying mixed outliers. For the balanced percentage of attribute qualities, misdetection happens because of the high false alarm rate. For unbalanced attribute qualities and excessive proportions, the masking effect is the key issue. The proposed chart shows the improved performance for the shift in identifying the shift in the process.
Technological advancements have profoundly influenced various sectors, including tourism, by simplifying international travel and increasing the demand for passports. In this context, public feedback on the Surabaya Immigration Office's services, gathered from Google Maps reviews, presents an invaluable dataset for analysis. This study introduces a novel approach to quality management systems in immigration services by applying the Convolutional Long Short-Term Memory (Co-LSTM) for sentiment analysis of public reviews and p-attribute control charts for statistical process control. Focusing on the Surabaya Immigration Office, we analyze public feedback from Google Maps to assess service quality. Our methodology preprocesses and classifies opinion data into sentiment classes, employing Co-LSTM for its superior accuracy over traditional models. The sentiment analysis reveals a predominance of positive over negative reviews, with classification accuracy demonstrating Area Under Curve (AUC) values of 98.61% for training and 85.66% for testing data. Furthermore, the p-attribute control charts are utilized to monitor service defects, identifying areas of uncontrolled variability and suggesting the necessity for service improvement interventions. The study uncovers that the primary public grievance relates to the perceived lack of friendliness and politeness from office staff. By integrating sentiment analysis with statistical process control, this research offers a comprehensive approach for immigration service providers to enhance service quality, respond proactively to public sentiment, and ensure customer satisfaction.
Intrusion detection is generally carried out by matching network traffic patterns with known attack patterns or by identifying abnormal network traffic patterns. One statistical methodological approach used in intrusion detection is Statistical Process Control (SPC) by constructing a control chart. Hotelling’s T2 control chart is a multivariate control chart commonly used to monitor the mean process. The performance of the T2 chart in monitoring mean shifts can be increased if a robust estimator is utilized. Based on previous research, T2 based on the Fast-MCD estimator has good performance in monitoring low to medium outlier contaminated data. Therefore, the MRCD estimator can be used to detect intrusion. On the other hand, this research focuses on developing a bootstrap-based robust Hotelling’s T2 charts with Fast-MCD and MRCD estimators for evaluating performance in detecting intrusion on intrusion detection datasets. Based application of UNSW-NB15, the proposed chart has better performance than the conventional T2 and Fast-MCD-based T2 despite the longer execution time.
Multivariate control charts have been applied in many sectors. One of the sectors that employ this method is network intrusion detection. However, the issue arises when the conventional control chart faces difficulty monitoring the network-traffic data that do not follow a normal distribution as required. Consequently, more false alarms will be found when inspecting network traffic data. To settle this problem, support vector data description (SVDD) is suggested. The control chart based on the SVDD distance can be applied for the non-normal distribution, even the unknown distributions. Kernel density estimation (KDE) is the nonparametric approach that can be applied in estimating the control limit of the non-parametric control charts. Based on these facts, a multivariate chart based on the integrated SVDD and KDE (SVDD-KDE) is proposed to monitor the network's anomaly. Simulation using the synthetic dataset is performed to examine the performance of the SVDD-KDE chart in detecting multivariate data shifts and outliers. Based on the simulation results, the proposed method produces better performance in detecting shifts and higher accuracy in detecting outliers. Further, the proposed method is applied in the intrusion detection system (IDS) to monitor network attacks. The NSL-KDD data is analyzed as the benchmark dataset. A comparison between the SVDD-KDE chart with the other IDS-based-control chart and the machine learning algorithms is executed. Although the it has high computational cost, the results show that the IDS based on the SVDD-KDE chart produces a high accuracy at 0.917 and AUC at 0.915 with a low false positive rate compared to several algorithms.
Indonesia is battling the COVID-19 pandemic. One of the government's strategies to break the virus's transmission chain is to track digital contacts in Indonesia using the PeduliLindungi application. The Google Play comment section is where users can express their opinions about the app. User opinions discovered on Google Play can be used to perform sentiment analysis and quality evaluation. The Naïve Bayes classification can be used to identify how user opinions contain positive, neutral, or negative sentiments in user reviews of the PeduliLindungi app on Google Play. The p and Laney p' charts can be used for quality evaluation. Laney p' control chart is an attribute chart used to monitor the proportion of defects with large and varied sample sizes. The data used in this study is from April 1, 2020, to March 31, 2022. According to the sentiment analysis results of user reviews of the PeduliLindungi app on Google Play, there are more negative reviews than positive classes. The classification accuracy has an Area Under Curve (AUC) value of 89.05%. This result shows that the test data has good classification. The monitoring results using p and Laney p' charts based on ratings and user reviews of the PeduliLindungi app show that the processes are still not statistically controlled. These findings indicate that the app developer still needs to make improvements.