
This study develops a robust framework for modeling dynamic volatility, asymmetry, and tail dependence in financial returns, focusing on the daily returns of Natural Resource Index (NRI) and the Oil and Gas Index (OGI). Standard multivariate volatility models, such as DCC-GARCH, often fail to adequately capture extreme co-movements, asymmetric dependence, and tail risk under non-normality, potentially producing biased variance and correlation estimates. To address these limitations, we propose a copula-based eGARCH model with multivariate Student-t (MVT) marginals, which jointly captures log-volatility dynamics and flexible tail dependence. Among 18 candidate models, the Copula-eGARCH-MVT model with eGARCH marginals and a multivariate Student-t copula provided the best fit, effectively accounting for heavy tails, volatility clustering, asymmetry, and dynamic dependence. Monte-Carlo simulations based on the empirical estimation confirmed that this framework reduces variance estimation bias, improves coverage probabilities, and more accurately models dynamic correlations and tail risk compared to standard DCC-GARCH models. This Monte-Carlo study showed that the coverage probability achieved by the eGARCH model is closer to the target 95
This study investigates the estimation of a sensitive population mean within a two-occasion successive sampling framework by employing calibration estimators in the presence of measurement error. Several calibration-based estimators are proposed by integrating Compulsory and Optional Randomized Response Techniques at both sampling occasions while explicitly accounting for measurement error. The efficiency of the proposed estimators is examined and compared with their corresponding direct estimators in order to assess the impact of measurement error on estimation accuracy. To illustrate the practical applicability of the suggested methodology, a simulation study based on a natural population associated with COVID-19 infection is conducted, demonstrating the implementation of Compulsory and Optional Randomized Response Techniques across both sampling occasions.
The most critical component of predictive analytics is parameter estimation, which aims to enhance the performance and interpretability of predictions. This work proposes a novel hybrid methodology combining Linear Mixed Models—which enable random effects estimation—with a fixed-effects model based on Extreme Learning Machines. A key challenge lies in balancing the interpretability of linear models with the complexity required to capture intricate data patterns. The proposed framework uses Linear Mixed Models to account for variability arising from random effects, such as subject-specific differences, while Extreme Learning Machines efficiently model complex fixed-effects relationships through their rapid learning and modeling capabilities. This dual approach captures both structured random variation and sophisticated fixed-effect patterns, thereby improving the accuracy and robustness of parameter estimates and enhancing the model’s adaptability to diverse data structures. Experimental results demonstrate that the hybrid model significantly outperforms conventional models in both predictive performance and computational speed. The proposed approach is broadly applicable across multiple disciplines, including healthcare, finance, and social sciences, which commonly involve complex, hierarchical data.
This study proposes a comprehensive statistical framework for reliability analysis under step-stress accelerated life testing (SSALT), accounting for multiple competing failure modes and progressive Type-I hybrid censoring. The framework combines the Tampered Random Variable (TRV) model to represent stress transition effects with the Inverted Topp–Leone (ITL) distribution to model cause-specific lifetimes. Maximum likelihood and Bayesian estimation methods are developed for parameter inference with incomplete data. Monte Carlo simulations are used to assess estimator performance, showing that the Bayesian approach provides improved accuracy and stability, particularly in small samples and under heavy censoring. Convergence of the Bayesian estimates is confirmed using standard diagnostic measures. An optimal test plan is also examined to identify efficient experimental designs under progressive hybrid censoring. The results indicate that larger sample sizes and balanced censoring schemes enhance estimation efficiency and inferential stability. Overall, the proposed framework provides a practical and reliable tool for analyzing complex engineering systems operating under varying stress conditions.
Forecasting exchange rate volatility remains complex because of nonlinear dynamics, structural breaks, and regime-dependent behavior in foreign exchange (FX) markets. This study proposes an early-warning framework that integrates Dynamic Time Warping (DTW) and K-Nearest Neighbors (KNN) with a Global Calibration Interval (GCI) for detecting volatility regime transitions in NZD/USD. Volatility is defined as the five-day rolling standard deviation of daily logarithmic returns, and candidate fluctuation events are identified using a dependence-adjusted Hoeffding criterion. Event-bounded intervals are converted into cumulative analytical prefixes after a 30-observation grace period and updated in five-observation increments. This procedure yields 2679 analytical prediction instances from 6574 raw daily observations. The instances are split chronologically into 911 training, 884 validation, and 884 test cases. For each instance, the mean of the five nearest-neighbor DTW distances is min–max normalized using parameters estimated from the combined training–validation data. A 95
In this paper, we develop a Bernstein polynomial framework for estimating cumulative distribution functions and quantiles in the presence of random right-censoring and dependent observations. Exploiting the smoothness and shape-preserving properties of Bernstein polynomials, we establish uniform strong consistency, derive explicit asymptotic bias and variance expansions, and prove asymptotic normality for the proposed estimators. We further investigate their finite-sample performance through mean squared error (MSE) and mean integrated squared error (MISE) criteria and construct a quantile estimator by inverting the proposed Bernstein-based distribution function estimator. Our results extend previous Bernstein-based approaches, including those of [1] and [21], to a unified framework accommodating both random right-censoring and α -mixing dependence.
This paper introduces the Log–Exponential Fréchet (LEF) distribution as a flexible generator-based extension of the classical Fréchet model for modeling right-skewed and heavy-tailed data. The proposed model enhances tail adaptability while preserving the fundamental characteristics of the Fréchet baseline. Parameter estimation is developed under the balanced ranked set sampling (BRSS) framework, which improves sampling efficiency in situations where measurements are costly or difficult to obtain. A unified inferential framework is established to accommodate several estimation approaches, including maximum likelihood, least squares, maximum product of spacings, Cramér–von Mises, Anderson–Darling, minimum spacing distance, and Linex-based spacing estimators. The framework is based a transformed-uniform representations, allowing the derivation of unified objective functions and associated gradient expressions. The performance of the estimators is evaluated through an extensive Monte Carlo simulation study under different sample sizes. The results indicate that estimation accuracy improves as the sample size increases, with the Cramér–von Mises and Anderson–Darling estimators demonstrating superior stability and accuracy. The practical applicability of the proposed model is illustrated using two real datasets, supported by goodness-of-fit comparisons with competing distributions. Overall, the proposed LEF distribution and the associated inferential framework provide an effective and flexible approach for modeling and analyzing heavy-tailed data in reliability and engineering applications.
This paper develops a comprehensive statistical framework for analyzing a progressive-stress accelerated life test model under progressive censoring, when transformed latent failure times follow a unit inverse Weibull distribution. The progressive stress is assumed to be proportional to time, and a cumulative exposure model is adopted to account for the effect of changing stress levels. We develop the maximum product of spacing estimators as an alternative to frequently used maximum likelihood estimators. Appropriate prior specifications are implemented to derive Bayes estimates under different loss functions. Interval estimation is also considered, and bootstrap, asymptotic, and credible intervals for the parameters are constructed. We have conducted a Monte Carlo simulation study to evaluate the efficiency and precision of the proposed estimators. Finally, the model’s practical applicability is demonstrated using a real-world data set, highlighting its effectiveness in reliability analysis under progressively changing stress conditions. Optimal plans are discussed by implementing various optimality criteria under different schemes.
Simulation-based evaluation is used to demonstrate computational efficiency and predictive accuracy gains from emulator-based order selection. Order identification using estimation method for large multivariate time series models presents substantial computational challenges, primarily at the estimation stage. Existing recommendations for model order selection in high-dimensional time series have largely focused on identifying models that best fit the observed data. When the goal of the analysis is prediction, alternative selection criteria may be more appropriate. This paper proposes an efficient, prediction-based approach to autoregressive model order identification for big multivariate time series. The simulation results demonstrate that the proposed method can significantly reduce computational time while still yielding accurate and reliable model orders for forecasting purposes. The method is illustrated with a big multivariate time series of weekly initial unemployment claims across 20 U.S. states.
Spatial clustering is important for identifying regions with similar spatial patterns in spatial datasets. This study focuses on selecting the optimal tuning parameter for the generalized lasso in spatial clustering analysis. Common approaches for selecting the tuning parameter in the generalized lasso include generalized cross-validation (GCV) and approximate leave-one-out cross-validation (ALOCV). However, these methods often produce substantially different tuning parameter values, which may lead to inconsistent clustering results and misinterpretation. In general, ALOCV tends to select larger tuning parameters, whereas GCV tends to select smaller ones. To address this issue, we propose an ensemble learning cross-validation (ELCV) approach that combines the validation errors from ALOCV and GCV using arithmetic, geometric, and harmonic means to obtain a more balanced tuning parameter selection. In addition, an analytical justification of the proposed ensemble framework is provided to demonstrate its theoretical relationship with ALOCV and GCV. A simulation study was conducted under four spatial clustering scenarios, namely three separated clusters, three connected clusters, five separated clusters, and five connected clusters, combined with three noise standard deviation levels to evaluate the robustness of the proposed methods. The Index of Edge Detection Accuracy (IEDA) was used as the primary criterion for assessing clustering performance. The simulation results showed that the proposed methods based on the geometric mean, and the harmonic mean, consistently achieved better and more stable performance across different scenarios and noise levels, as indicated by higher IEDA values and lower estimation errors compared to ALOCV and GCV. Finally, the proposed methods were applied to cluster the productivity of oil palm fresh fruit bunches (FFB) across several planting blocks in oil palm concessions in Kalimantan, Indonesia.
In an observational study using data from Kaiser Permanente Northern California (KPNC), we evaluated the effectiveness of respiratory syncytial virus (RSV) immunoprophylaxis palivizumab, a monthly injection that eligible infants receive during the winter RSV season, on RSV-related morbidity, particularly recurrent bronchiolitis in infancy. Traditional methods such as extended Cox Proportional Hazards (CPH) models are commonly used to analyze these recurrent event data. We propose a new approach based on the nonhomogeneous Poisson process that explicitly models the impact of a past event episode and a time-varying treatment on the likelihood of future episodes of the event. Two models, the common-hazard and distinct-hazard models, are developed. These models and the extended CPH model are evaluated in simulation studies and applied to the KPNC dataset.
This paper investigates the sequential Bayesian estimation of the sum of two failure rates in a series system, where the component lifetimes follow exponential distributions. Traditional fixed sampling allocations, based on prior engineering judgment, are often suboptimal under limited testing budgets or poor prior knowledge. We develop a fully adaptive two-stage sequential procedure that dynamically allocates tests between components to recover from prior mismatch. Theoretical analysis establishes the first-order asymptotic optimality of the sequential scheme. Extensive simulations demonstrate that sequential allocation consistently reduces mean squared error compared to best fixed allocation, while remaining robust across a variety of priors and budget constraints. The results offer a practical and theoretically grounded strategy for resource-efficient reliability testing and risk assessment in engineering systems.
Suppose Y_1,… ,Y_n are observations from an AR(p) process with mean μ , e.g., Y_t=μ +ϕ _1(Y_t-1-μ )+⋯ +ϕ _p(Y_t-p-μ )+Z_t , where {Z_t} is an IID sequence with mean zero, variance σ ^2 , and common distribution function F_0(z) . Indexing by quantile q=F_0(z) , it will be shown that if the parameters are estimated via maximum likelihood (MLE) using only the first half of the observations ( Y_1,… ,Y_⌊ n/2⌋ ), then the empirical process based on all of the residuals will converge in distribution to B(q), where B is a standard Brownian bridge on [0, 1]. If all the data is used to estimate the parameters using MLE or least squares (LS), then the empirical process of the residuals will converge in distribution to B(q) plus a correction term. This is the content of [9] in the Gaussian case and of [8] in the more general case. Nevertheless, this correction term disappears when using the first half of the observations to estimate the parameters. These results may be viewed as extensions of the half-sample device of [6] and also connect with the results of [5] for the sample ACF.
The aim of this paper is to present a discussion on some known approximations for the price of Basket options and at the same time introduce new approximations for the Greeks based in those approximations. The analysis centers on the Greeks Delta, Vega, and Cega, with a special focus on Cega, which reflects the sensitivity to the correlation coefficient, a parameter specific to options on multiple underlying assets. Python programming was used to study the effects of correlation and other parameters that influence the price of an option. Since basket options lack an explicit pricing formula, Monte Carlo simulation and other methods were employed to calculate their price and derivatives. To analyze the results for the price and Greeks of this kind of options, Python codes where developed and implemented and a discussion based in two different scenarios is presented. It was observed that the discussed alternative methods for approximating the price and Greeks of basket options give good results when compared with Monte Carlo in some scenarios, being in that situation more efficient, due to the latter’s longer computational time.
UniLasso and the smoothly clipped absolute deviation (SCAD) penalty method are compared for identifying active factors in two-level supersaturated designs. A two-stage UniLasso procedure is proposed in which an initial UniLasso fit serves as a screening step and the selected factors are subsequently refined by stepwise forward selection to improve model parsimony. Using benchmark simulation settings previously employed in the literature, the methods are evaluated under varying levels of sparsity and heterogeneous effect sizes. Across a wide range of scenarios, UniLasso exhibits stronger screening performance and frequently achieves higher true-model recovery rates than SCAD, particularly in more challenging settings involving larger numbers of active factors and smaller effects. Although the initial UniLasso stage tends to select larger models, the proposed refinement substantially reduces model size while retaining most of the identification advantage. Analyses of several real-data examples show that the two-stage procedure generally yields more stable and parsimonious models than SCAD, while also illustrating the trade-off between recovering active effects and maintaining model parsimony. Overall, the results suggest that second-stage refinement substantially improves the practical utility of UniLasso.
A classical problem in the field of extreme value theory is the estimation of the extreme value index (EVI), which enables modeling of the maximum (or minimum) of a random sample from an unknown distribution. The EVI classifies the underlying probability distributions into three categories for the maximum domain of attraction: positive, negative, or zero. The Hill and Pickands estimators for the EVI are well known in the frequentist literature. In this paper, we propose a Bayesian method for inference on the EVI. We define a sequence of functionals of the distribution that approximates the EVI. The constructed functionals are different depending on whether the EVI is positive, negative, or zero. We assign a Dirichlet process prior on the distribution of the observations and use the induced posterior distribution of the sequence of functionals to make inferences on the EVI. We establish a Bernstein-von Mises Theorem for the posterior distribution of the EVI, which gives rise to the posterior contraction rate and asymptotic frequentist coverage of a Bayesian (1-α ) -credible interval. Further, to test the hypothesis that the EVI is zero, we propose a test based on a modified Bayes factor. We use a power of the empirical likelihood and compute the marginal likelihood under the null hypothesis by the Hamiltonian Monte Carlo method. We show the effectiveness of the proposed Bayesian methods through a comprehensive simulation study.
DNA methylation is an important epigenetic event associated with cancers. Different genomic sites tend to be co-methylated. It is unclear which correlation metrics should be used to study co-methylation and how they perform in identifying highly co-methylated (HCM) sites. The impact of different features of the data is also unclear, e.g., outlier, low variance, and data transformation (from B to M = logit(B)). We therefore conducted comparative analyses of six metrics, Pearson, Spearman, Kendall, Hoeffding, Distance, and Maximal Information Coefficient (MIC). Key findings are summarized below. First, the numbers of HCM sites identified by the six metrics were very different when using a fixed cutoff value. Pearson and Distance identified more HCM sites and had strong similarities. However, these metrics were susceptible to outliers and data transformation. They identified more HCM sites when using B values, but these sites tended to have outliers and lower variance. Second, Kendall’s and Hoeffding’s scores were significantly lower than those of the other metrics, leading to fewer HCM pairs being identified. Third, MIC required a large sample size to perform effectively. Although it may detect unique correlation patterns, it is difficult to interpret these patterns biologically. In summary, considering many factors together (e.g., outliers, low variance, data transformation, cutoffs, and runtime), researchers should carefully evaluate the distinct strengths and limitations of each method to select the most appropriate one for their methylation data and the goal of their analysis.
Human papillomavirus (HPV) infection, particularly with HPV types 16 and 18, is the most prevalent viral carcinogen among head and neck cancers. Notably, HPV-positive (HPV+) status is associated with improved survival and treatment response compared to HPV-negative (HPV−) cancers. However, comprehensive and rigorous analyses of HPV serology remain limited by substantial heterogeneity and numerous outliers. In this study, we analyzed 454 serum samples from 151 patients with squamous cell carcinoma of the head and neck, collected at a single institution between 2007 and 2020. Total IgG antibody titers against the E7 oncoproteins of HPV16 and HPV18 were measured using ELISA (Khanal et al., 2015). The 95