
Missing data refer to situations in which part of a dataset is unavailable due to unrecorded information during the data collection process. Such issues must be addressed because they may affect analytical accuracy. This study evaluates the performance of two univariate imputation methods: Kalman Smoothing and Seasonal Trend Decomposition using LOESS (STL Decomposition), in which LOESS stands for Locally Estimated Scatterplot Smoothing. Both methods were applied to time series data collected from a weather station in Lampung Province, Indonesia, covering the period 2001-2024. The analysis incorporates missing data mechanisms, namely MCAR (Missing Completely at Random), MAR (Missing at Random), and MNAR (Missing Not at Random), with missing rates of 5%, 10%, 20%, 30%, 40%, and 50%. The variables examined include average temperature, average relative humidity, total precipitation, average solar radiation, and average wind speed. The results indicate that STL Decomposition outperforms Kalman Smoothing in terms of precision, consistently yielding lower RMSE values, particularly at missing data levels of 5%, 10%, and 20%, and across all missing data mechanisms. Although Kalman Smoothing provides relatively greater stability in preserving temporal dependencies, STL Decomposition demonstrates superior accuracy in imputing missing values, as reflected in its consistently lower RMSE across various missing data scenarios.
This paper presents the BerTG distribution, an innovative three-parameter discrete probability model developed by convolving a Bernoulli random variable with an independently distributed transmuted geometric random variable. The suggested distribution constitutes a significant and adaptable generalization that en- compasses various established count distributions as specific instances, thereby offering a cohesive framework for modeling a range of count data formats. The BerTG distribution is notable for its exceptional ability to handle many types of dispersion, including as overdispersion, underdispersion, and equidispersion, com- monly found in real-world count data, thereby overcoming a significant weakness of numerous conventional discrete models. A thorough examination of the distributional and structural characteristics of the BerTG model is conducted, including the probability mass function, cumulative distribution function, moments, moment generating function, factorial moments, probability generating function, and index of dispersion, among other aspects. Special emphasis is placed on reliability-theoretic attributes, encompassing the hazard rate function, survival function, reverse hazard rate function, and conditional expectation, which are metic- ulously generated and examined. Additionally, essential actuarial metrics, including the stop-loss premium, value-at-risk, and tail value-at-risk, are analysed to illustrate the model’s appropriateness for risk-theoretic applications. Model parameters are estimated by maximum likelihood estimation, and the asymptotic prop- erties of the resultant estimators are determined. A comprehensive simulation analysis is performed to assess the finite-sample performance of the estimators concerning bias, mean squared error, and consistency across diverse parameter configurations and sample sizes, thereby validating the reliability and accuracy of the estimation technique. The practical applicability of the BerTG distribution is evidenced through real-world data applications, wherein the model is applied to several empirical count datasets displaying diverse dispersion traits. Comparative analyses with various competing discrete distributions demonstrate that the BerTG model consistently attains superior goodness-of-fit performance, as indicated by standard model selection criteria such as the Akaike information criterion, Bayesian information criterion, and chi- square goodness-of-fit statistics. The results combined demonstrate that the BerTG distribution is a very competitive, manageable, and adaptable instrument for the statistical modelling of count data across several applicable fields.
This study develops double and multiple (three-stage) acceptance sampling plans based on the X-gamma distribution, a flexible model widely used in reliability engineering. The plans are built around truncated life tests, meaning testing stops either after a predetermined time or once a set number of failures occur. For the two-stage (double) sampling approach, we determine the smallest practical sample sizes and calculate the average sample number (ASN). We then expand the framework to a three-stage plan, outlining how it works and specifying the sample sizes needed at each step. To evaluate performance, we derive operating characteristic (OC) curves for all plan types, showing how well they distinguish between good and poor-quality lots across different quality levels. We also identify the minimum ratio of actual to specified mean life required to keep producer’s risk within acceptable limits. Additional efficiency metrics like average total inspection (ATI) and average outgoing quality (AOQ) are calculated to give a complete picture of how each plan performs in practice. Real-life numerical examples are included to walk through how these plans guide lot-acceptance decisions using the X-gamma model. Ultimately, this work gives quality control professionals reliable, distribution-specific tools for designing smarter, more efficient acceptance sampling procedures.
Recently, the practice of employing control charts to monitor manufacturing processes has gained significant interest in the field of statistical process control (SPC). A properly designed control chart is capable of promptly detecting any shifts in the process. In this article, a novel two-sided CUSUM control chart with repetitive sampling is introduced for monitoring the process mean. To evaluate the performance of the proposed CUSUM control chart, various statistical measures such as the average, standard deviation, and percentiles (including the median) of run lengths are utilized. These measures are assessed under different distribution scenarios. Furthermore, a comparison is made between the performance of the proposed chart and several existing control charts. By presenting the new two-sided CUSUM control chart with repetitive sampling, this article contributes to the advancement of process monitoring techniques in SPC. The evaluation reveals that while the proposed chart is highly effective for detecting very small process shifts, there is a trade-off in efficiency as the shift magnitude increases. These findings provide insights into the specific scenarios where the CUSUM-RS chart is most beneficial compared to existing tools.
This paper introduces the logarithmic Topp-Leone-G (LTL-G) family, a novel generalized class engineered to significantly enhance the modeling capacity of baseline distributions for complex empirical data exhibiting pronounced skewness, heavy tails, and non-monotonic hazard structures. We rigorously establish the theoretical foundations of the proposed specification, deriving explicit linear expansions for the probability density function, closed-form expressions for ordinary and incomplete moments, and comprehensive distributional characterizations based on truncated moments, reverse hazard functions, and conditional expectations. To critically evaluate inferential reliability, we designed an extensive Monte Carlo simulation protocol comparing six competing estimation methodologies including the maximum likelihood (MLE), ordinary least squares (OLS), Cram & eacute;r-von Mises (CVM), Anderson-Darling (ADE), right-tail ADE (RTADE), and left-tail ADE (LTADE) across systematically varied sample sizes and challenging parameter configurations. The finite-sample performance is rigorously quantified through bias, root mean squared error, and Kolmogorov-Smirnov diagnostics, with dedicated attention to the convergence behavior of key risk indicators (KRIs). We compute Value-at-Risk (VaR), Tail Value-at-Risk (TVaR), Tail Variance (TV), Tail Mean Variance (TMV), and expected shortfall (ELq) to demonstrate the framework's superior capacity for extreme-tail quantification under data scarcity. Empirical validation is conducted using two high-stake real-life datasets: Social Security Administration (SSA) disability beneficiary records and a UK motor non-comprehensive claims development triangle. The analytical results consistently reveal that the new family specification accurately accommodates extreme dispersion and temporal claim dependencies, while delivering a statistically rigorous foundation for modern actuarial reserving and evidence-based capital allocation.
In this paper, we have discussed the reliability estimation for a multicomponent stress-strength (MCSS) model when stress and strength follow the Kumaraswamy inverse Weibull distribution, given by Shahbaz et al. (2012). We have obtained the maximum likelihood estimate of the reliability alongside the asymptotic distribution of the parameters involved. Also, the asymptotic confidence intervals have been obtained. An extensive simulation study has been conducted to assess the performance of the estimates. A real data application has also been given. It is found that the reliability increases with an increase in one of the shape parameters of stress distribution.
Lifetime distributions play a key role in statistical modeling, with extensive applications across biostatistics, reliability engineering, and survival analysis. This paper introduces a novel and flexible bivariate lifetime model, termed the Bi-variate Cubic Transmuted Weibull Distribution (BCTWD), which extends the transmuted Weibull framework proposed by Alsalafi et al. (2025) by incorporating a cubic transmutation mechanism to enhance modeling flexibility and capture complex dependence structures. Existing bivariate Weibull models cannot simultaneously accommodate flexible marginal tail behavior and complex dependence structures, limiting their applicability in scenarios with heterogeneous failure patterns. The theoretical foundations of the proposed BCTWD are rigorously developed, including its joint and marginal probability density and cumulative distribution functions, along with essential statistical and reliability properties. Parameter estimation is performed using both the Maximum Likelihood (ML) and Inference Functions for Margins (IFM) methods, whose performances are systematically evaluated through simulation experiments. The simulation outcomes indicate that the estimators are, For n=200, the maximum absolute bias for shape parameters are 0.048, and the maximum MSE is 0.29, indicating satisfactory finite-sample performance, particularly for the shape parameters under heavy-tailed scenarios. An empirical application to bilateral eye failure time data from a diabetic retinopathy study demonstrates the practical utility of the proposed model. Based on the maximum likelihood estimates and model selection criteria, including AIC, AICc, and BIC, the BCTWD achieves superior goodness-of-fit compared with the Bivariate Transmuted Weibull (BTW) and Bivariate Weibull (BW) distributions. While the BCTWD exhibits slightly greater parameter variability due to its added flexibility, it provides the most accurate representation of the data, confirming its effectiveness in modeling dependent lifetimes. Overall, the BCTWD enriches the family of multivariate lifetime distributions by offering enhanced adaptability and interpretability, making it a valuable tool for applications in reliability analysis, biostatistics, and survival modeling.
Autism spectrum disorder (ASD) is a complex neurodevelopmental condition that typically emerges in early childhood and persists throughout life, making early and objective detection crucial. This study integrates graph theoretic approach with anomaly detection approaches to identify atypical functional brain regions in children aged up to 5 years with ASD. Resting-state fMRI data of 53 ASD and 53 healthy subjects were used to construct functional connectivity matrices were across 32 functional regions of interest, from which graph theoretic features were extracted. The one-class SVM achieved an AUC of 0.733 in identifying atypical regions. Atypicality was observed in the salience network , specifically, in the supramarginal gyrus, anterior cingulate cortex and left anterior insula. In the visual network, medial occipital and laterl occipital regions were identified as atypical. The language network showed atypical regions in the right inferior frontal gyrus and the left posterior superior temporal gyrus. The dorsal attention network exhibited atypicality in the right frontal eye field region. Graph-theoretic analysis to regional atypicality highlighted disruptions in integration, segregation, and hub-related characteristics.
Cokriging is a multivariate spatial method used to predict the observed value for a primary variable in an unknown location with the help of a spatially correlated secondary variable. The existence of two or more nonlinear secondary variables in predicting spatial data usually arises, especially in cokriging. Therefore, a method that can improve the model's predictive power by adding the interaction of variables is proposed. The proposed method can be effectively used, especially when the primary and secondary variables have a nonlinear relationship. By transforming the nonlinear variables, a higher correlation can be attained. This study used principal component analysis with interaction (PCAI) method among secondary variables to reduce two or more secondary variables into one dimension as a secondary variable in the cokriging technique. The proposed method was tested and verified through simulation and real data using the 2015 South Korea Air Pollution dataset, a dataset known for its complex spatial patterns and high variability, to prove its validity and usefulness. The predicted residual error sum of squares (PRESS) statistic was used for cross-validation. Computations were done using the R Project for Statistical Computing software. PCAI as a secondary variable gives the lowest PRESS value compared to only one secondary variable or principal component analysis (PCA). Considering the criterion, the lowest value of PRESS indicates the best model. Thus, PCAI cokriging outperformed PCA cokriging. Using PCAI as a secondary variable may be a better method than PCA for with nonlinear multicovariates.
This paper introduces a new extension of the Chen distribution, designed to better model extreme low-flow events in hydrology and rare events in the medical field. The proposed model incorporates asymmetrical and heavy-tailed behavior, making it particularly useful for analyzing extreme values in complex real datasets. We derive the mathematical properties of the BGC distribution and apply two advanced analytical techniques: the Mean-of-Order-P (MOOP) method to determine the optimal value of P (referred to as Opt-P), and the Peaks Over Threshold Value-at-Risk (PORT-VaR) approach to identify and assess critical extreme events. These methods are applied to real datasets including relief times, minimum river flow data from the Cuiab & aacute; River, and U.S. indemnity losses from general liability claims. The MOOP analysis shows that increasing the order P leads to reduced Mean Squared Error (MSE) and Bias, indicating improved estimation accuracy. For example, in the relief times dataset, MSE decreases from 0.64 at P=1 to 0.3844 at P=5. Similarly, for the minimum flow data, MSE drops from 4402.88 to 3684.27 with increasing P, highlighting the benefits of higher-order statistics in capturing central tendencies. Using PORT-VaR, we analyze extreme peaks under varying confidence levels (50%, 70%, 90%, and 99%) and compute key risk indicators such as Value-at-Risk (VaR) , Tail Value-at-Risk (TVaR) , Mean Excess Loss (MEXL) , Tail Variance (TV) , and Tail Mean Variance (TMV) . In the relief times dataset, VaR increases from 1.70 at 50% confidence to 3.055 at 99% confidence, demonstrating growing risk exposure at higher confidence levels. For the minimum flow data, VaR rises from 115.925 at 50% to 157.169 at 99%, underscoring the importance of adaptive risk thresholds in managing water scarcity and dam safety. A financial case study using U.S. indemnity loss data further validates the robustness of the BGC model in capturing tail behavior and estimating extreme risks. At the 99% confidence level, VaR reaches 170400 (in thousands of USD), and MEXL is 203411, illustrating the nonlinear growth of risk in heavy-tailed insurance claims. Finally, a comparative study under a historical financial claims data through an application.
This paper presents a novel exponential model with two parameters, placing particular attention on its practical applications to skewed data as the central area of investigation. The mathematical characteristics of this atypical distribution are established, in a lucid and succinct manner, by the discoveries made in this investigation. Furthermore, it is worth noting that there exist three distinct approaches to describing the distribution. The process of estimating the parameters of the novel model involves employing a range of established methodologies, including the Bayesian technique. When confronted with censored data, the maximum likelihood technique is commonly considered as a viable approach. Pitman's closeness criteria are employed as the comparative tool when assessing the probability estimate in relation to Bayesian estimation approaches. During the computation of Bayesian estimations, three distinct loss functions, namely generalized quadratic, Linex, and entropy, are employed. A multitude of simulated experiments are conducted to assess the efficacy of various estimation methodologies. The BB algorithm is employed to facilitate the comparison and contrast between the Bayesian technique and the censored maximum likelihood strategy. The Nikulin-Rao-Robson (NKRR) statistic was derived by conducting two empirical studies using real-world data sets characterized by skewed distributions, along with simulation research conducted in an unfiltered environment. Furthermore, this paper delineates two other uses within the same context. The study's findings illustrate the efficacy of the approaches presented for the purposes of distribution and estimation.
This paper examines the characterizations of the five recent univariate continuous probability distributions (2022-2025) that were proposed relatively recently. These characterizations are based on: (i) a simple relationship between two truncated moments; (ii) reverse hazard function. It should be mentioned that for the characterization (i) the cumulative distribution function need not have a closed form and depends on the solution of a first order differential equation, which provides a bridge between probability and differential equations.
Background: Pakistan has witnessed concerning shifts in HIV epidemic especially in Punjab, where HIV and AIDS incidence continues to rise. This study compares the predictive accuracy of the Prophet model, machine learning model with classical ARIMA configurations for monthly HIV and AIDS case forecasting in Punjab. Methods: Monthly surveillance data (January 2020-October 2025) from Punjab AIDS Control Program (PACP) was used to train and validate Prophet and multiple ARIMA models. The modelling performance was assessed using RMSE, MAE, MAPE, BIC and also Ljung Box Q tests. Forward forecasts were generated for HIV reactive and AIDS (CD4 < 200) cases through 2026. Results: Machine learning model (Prophet) outperformed all ARIMA models in forecasting HIV reactive cases by achieving the lowest RMSE (132.6) and MAPE (16.4%), for AIDS cases projection, all models exhibited high error rates (Prophet MAPE > 300%) with ARIMA (0,1,0)(0,1,1)12 better performance (MAPE similar to 174%). Forecasted outputs estimates approximately 8,490 new HIV cases in 2026 with uncertainty bounds reaching nearly 15,000 cases, indicating a continued upward trajectory and for AIDS the count in 2026 may rise to 25,596 new cases, thou, forecasting AIDS remains a challenge. The results demonstrate superior ability of Prophet model to capture non-linear trends and seasonality in HIV surveillance data. Conclusion: Prophet model superior performance reflects its ability to model nonlinear and seasonally irregular HIV surveillance data. Integration of machine learning techniques such as Prophet model into provincial HIV programs can enhance planning and accelerate progress toward achieving UNAIDS 95-95-95 targets.
In this article, we will derive a closed-form estimator for the probability R = P(X < Y ) based on a ranked set sampling (RSS) scheme when the he random variables X and Y are assumed to follow the Lehmann Type-II (L-II) family of distributions. Estimating R through the maximum likelihood (ML) method within the RSS framework does not yield an analytical solution because of the non-linear components present in the likelihood equations. In this context, we employ a modified maximum likelihood (MML) estimation approach to derive a closed-form estimator for R. Estimates of R under both ML and MML techniques along with their corresponding asymptotic confidence intervals are determined and compared in a simulation study under one of the distributions of the L-II family called the inverse Topp-Leone distribution. At the end, the simulation results are strengthened using a real example in the field of agriculture.
This manuscript presents a new extension of Rayleigh distribution by employing the concept of Kth order equilibrium method. The introduced model is termed as the Kth-order equilibrium Rayleigh distribution (KERD). Various statistical properties of the new distribution, including its aging behavior and stochastic ordering relations are analyzed. Explicit expressions are derived for moments, conditional moments, incomplete moments, the mean residual function, the mean waiting function, entropy measures and order statistics. Distribution characterization has been examined. Maximum likelihood estimation method is used to estimate the parameters. A simulation study using the Anderson-Darling test statistic is carried out to analyze the asymptotic behavior of maximum likelihood estimators. The behaviors of bias and mean square error are observed with the increase in sample size. The applications of new distribution are demonstrated using two different real life datasets. Ultimately, a comparison is conducted among KERD and its sub-models regarding their fit using information criterion tools.
The bivariate compound zero-truncated Poisson-gamma distribution represents the sum of a random number of bivariate Gamma variables, with the count governed by a zero-truncated Poisson distribution. This formulation makes the model particularly suitable for applications in actuarial science, climatology, and reliability engineering, where zero outcomes are structurally absent. However, owing to the intractable form of the probability density function, which involves an infinite series, direct maximum likelihood estimation becomes computationally demanding. In this study, we use standard (exact) maximum likelihood estimation when event counts are observed (complete data and Scenario A) and employ the saddle-point approximation only when counts are latent (Scenario B). We developed a stable maximum likelihood estimation based on the saddle-point approximation. We derived the cumulative distribution function from the cumulant generating function and obtained the probability density function using numerical differentiation. Detailed derivations, implementation guidelines in the R programming language, and a parameter initialization strategy using the method of moments estimation are provided. A simulation study using various sample sizes demonstrated the accuracy, consistency, and superiority of this method over the moment-based estimators. Computational challenges and limitations are discussed, along with potential extensions to model the dependence structures using copulas. In addition, we develop a likelihood ratio test and a formal symmetry test (for example, H-0 : alpha(1 )= alpha(2), beta(1) = beta(2)) to compare nested specifications, enabling principled inference on symmetry and overall model adequacy.
The Weibull distribution, widely utilized due to its flexibility, often requires generalization to improve its fit to real-world data. The Transmuted Weibull Distribution offers enhanced flexibility by incorporating a transmutation parameter. Metaheuristic algorithms have emerged as robust tools for parameter estimation, particularly for probability distributions with complex likelihood functions. This study compares the performance of four metaheuristic algorithms: Genetic Algorithm (GA), Particle Swarm Optimization (PSO), Differential Evolution (DE), and Artificial Bee Colony (ABC) against the traditional Newton-Raphson (NR) algorithm for estimating parameters of the Transmuted Weibull Distribution (TWD). Extensive Monte Carlo simulations evaluated the algorithms' efficiencies using metrics like log-likelihood values, bias, mean squared error (MSE), and deficiency. Additionally, the methods are applied to real-world datasets to compare their practical utility. Both simulation and real data application results revealed that metaheuristic algorithms outperformed traditional Newton-Raphson (NR) optimization.
This study proposes a new and versatile family of continuous probability models known as the log-exponential generated (LEG) distributions, with particular emphasis on the log-exponential generated Weibull (LEGW) model as its prominent representative. By introducing an additional layer of parameterization, the family offers improved adaptability in shaping distributional forms, especially regarding skewness and heavy-tailed behavior. The LEGW formulation proves especially relevant for reliability data and for capturing rare but impactful events where asymmetry plays a major role. The work details the theoretical framework of the family through explicit expressions for its cumulative distribution function (CDF) and probability density function (PDF), alongside the corresponding hazard rate function (HRF). Several analytical characteristics are also investigated, including series representations and behavior in the extreme tail. To demonstrate practical value, the paper conducts risk evaluations employing sophisticated key risk indicators (KRIs) such as Value-at-Risk (VaR), Tail Value-at-Risk (TVaR), and tail mean-variance measure (TMVq) across multiple quantile levels. Parameter estimation is addressed using several techniques, including maximum likelihood estimation (MLE), the Cram & eacute;r-von Mises approach (CVM), and the Anderson-Darling estimator (ADE), in addition to their right-tail adjusted (RTADE) and left-tail adjusted variants (LTADE) to better capture extreme behaviors. Comparative performance analyses are carried out using both controlled simulation scenarios and real data from the insurance and housing sectors to test robustness under heavy-tail conditions. The findings highlight the effectiveness of the LEGW model in applied risk assessment, supported by evidence from insurance claims and economic datasets.
Anemia continues to be a significant public health issue, particularly impacting women aged 15 to 49. To improve the modeling of anemia prevalence, this study introduces the proposed distribution, offering enhanced flexibility for capturing skewed and heavy-tailed data structures. The model is applied to country-level data from Pakistan, with global trends from World Bank data serving as a comparative backdrop. The TLEG-E distribution demonstrates superior fit and interpretability compared to traditional models, effectively highlighting a declining trend in anemia among Pakistani women, potentially reflecting the impact of health policy reforms and improved nutritional access. While global prevalence varies widely across regions, the emphasis here lies in the methodological advancement and its utility for health data modeling. The proposed framework provides a robust statistical foundation for tracking anemia trends and can support more targeted policy interventions. Its adaptability makes it suitable for broader applications in epidemiological research, enabling more precise assessments of public health initiatives across diverse populations.
Heteroscedasticity is a well-known violation of an assumption in parametric regression analysis. In such cases, to handle this problem, a generalized least squares method is used. In this article, we have manifested the robustness of nonparametric regression in the case of heteroscedastic errors. Nonparametric regression is a robust method that proceeds without requiring inflexible assumptions from the model. We empirically compared the performance of the generalized least squares method with multivariate nonparametric kernel regression. Multivariate nonparametric kernel regression is used with a Gaussian kernel and six bandwidths on China's per capita consumption expenditure. The performance of nonparametric regression with Bayesian bandwidth was found better on the basis of mean squared error. Simulation results are also presented, with their graphical representation, where nonparametric regression with different bandwidths at different heteroscedastic levels is observed, and we found that our proposed method performed best in both presence and absence of homoscedasticity.