
Copula-based approaches have been increasingly employed to model the dependency among soil parameters. Nevertheless, existing applications are predominantly confined to conventional low-order bivariate copulas, which limits their capability to capture complex multivariate dependency structures. To overcome this limitation, this study introduces the use of vine copula models for advanced dependency modelling in geotechnical engineering. To demonstrate the applicability of various vine copula structures, a slope reliability analysis framework was developed by integrating Monte Carlo simulation with a Multilayer Perceptron regression model. The influence of different vine copula configurations on the estimated probability of failure was systematically examined, and their performance was rigorously evaluated. The results reveal that the choice of vine copula structure has a significant impact on the predicted failure probability. Moreover, the proposed vine copula models exhibit closer agreement with the observed statistical characteristics of soil parameters compared to the traditional multivariate normal distribution approach. Overall, this study provides practical insights into the selection and construction of appropriate vine copula models for high-dimensional dependency modelling, particularly in geotechnical applications. The proposed framework enables more reliable estimation of slope failure probabilities. Journal of Statistical Research 2026, Vol. 60, No. 1, pp. 213-230.
In this article, we introduce a one parameter asymmetric version of the standard Laplace distribution, named “the alternate asymmetric Laplace distribution (AALD)”. The proposed model extends the classical Laplace law through the inclusion of a single shape parameter ρ. While most of the existing asymmetric models rely on multiple shape parameters and lack closed-form expressions for higher-order characteristics, the AALD, with its single parameter structure, offers analytical tractability and flexibility. Specifically, closed-form expressions are derived for the cumulative distribution function, moment generating function, quantiles, entropy, and order statistics, underscoring the theoretical richness of the model. The location-scale extension of AALD, named “the extended alternate asymmetric Laplace distribution (EAALD)” is introduced and studied in detail. The estimation and testing procedures for the parameters of EAALD are discussed and the performance of the obtained estimators is examined through a simulation study. Through applications to two real-world data sets from distinct domains, EAALD consistently provides a superior fit compared to the established asymmetric models, highlighting both its methodological novelty and practical relevance. Journal of Statistical Research 2026, Vol. 60, No. 1, pp. 1-31.
The Systematic sampling method has been proven effective in finite population for estimation of the population characteristics. However, its application to an infinite population requires further exploration. This article aims to extend the systematic sampling method to infinite populations and presents several approximation techniques. The performance of each method is evaluated based on different characteristics of the population distribution. Specifically, we focus on important distributions and investigate the case of Gamma distribution here. Three approaches are considered, with the first method showing promising results. This method involves assigning Gamma probabilities to all possible samples, avoiding the need to truncate the infinite population into a finite one. Also, it is seen that estimates are closed to original one if all samples include mode or neighbourhood of mode. The findings of this study shed light on the applicability of systematic sampling for infinite populations and contribute towards the estimation of population characteristics by this kind of work. Journal of Statistical Research 2025, Vol. 59, No. 2, pp. 305-323
Accurate modelling of tropical storm tracks and intensities requires statistical methods that respect the intrinsic directional geometry of storm motion. Conventional approaches frequently rely on Euclidean approximations, which may distort inference, bias parameter estimation, and mischaracterize predictive uncertainty on spherical manifolds. This paper develops a geometry-aware statistical framework for analysing tropical storm dynamics by explicitly incorporating circular and spherical structures. The proposed methodology integrates circular descriptive analysis, regime-specific movement modelling, circular regression for wind intensity, spherical state-space filtering for trajectory evolution, and direction-dependent extreme value modelling for severe events. Storm headings (ϕ) and turning behaviour (θ) are modelled using von Mises and wrapped Cauchy distributions, while extreme wind intensities are characterized within a generalized Pareto framework with direction-dependent parameters (ξ, σ). The practical utility of the approach is demonstrated through comprehensive empirical benchmarking against Euclidean baselines, homogeneous directional models, and machine-learning alternatives, alongside Monte Carlo simulation studies assessing estimator bias, variance stability, and regime recovery. Results indicate systematic improvements in trajectory prediction, wind intensity forecasting, model parsimony, and uncertainty calibration (coverage). These statistical gains translate into more reliable preparedness decisions within a cost–loss framework. Overall, the findings demonstrate that explicitly respecting circular and spherical geometry yields tangible inferential, predictive, and decision-theoretic advantages for tropical cyclone risk assessment. Journal of Statistical Research 2026, Vol. 60, No. 1, pp. 113-131.
A new five parameter lifetime distribution called the Harris Extended modified Weibull (HEMW) distribution is proposed, various statistical properties including moments, generating function, incomplete moments, mean deviations, Bonferroni and Lorenz Curves, residual life and reversed residual functions are investigated. Estimation of the model parameters by the method of maximum likelihood is discussed. Applications to some real data sets which motivate the flexibility and potentiality of the new model are provided. For the proposed model an acceptance sampling plan is also derived. Journal of Statistical Research 2026, Vol. 60, No. 1, pp. 175-196.
A new distribution referred to as the reflected, shifted, truncated, generalized exponential (RSTGE) distribution is proposed to model negatively skewed data. We estimate model parameters using a hybrid method of percentile and maximum likelihood estimation. The performance of the proposed hybrid method is evaluated through Monte Carlo simulations. Estimators are evaluated using root mean squared error and average deviation. We compare the RSTGE distribution to the exponential, generalized F, generalized gamma, Gompertz, log-logistic, log-normal, Rayleigh, and Weibull distributions in negatively skewed real data sets with complete, right censored and interval censored observations. Our study suggests that the new distribution performs better than the eight distributions mentioned above when modeling complete, mildly right-censored, or interval-censored data, and is comparable with those that appropriately model heavily right-censored data. Journal of Statistical Research 2026, Vol. 60, No. 1, pp. 51-77.
This study presents a comprehensive exploratory and directional analysis of 1,000 georeferenced earthquake events from the seismically active Fiji–Tonga region, based on the classical quakes dataset in R. Integrating both multivariate and circular statistical methods, the analysis investigates the spatial, depth-dependent, and directional structure of seismicity along a complex subduction zone. Classical techniques—including kernel-based density estimation, spatial point process modeling, and correlation analysis—reveal a distinctly bimodal depth distribution and a strong linear relationship between magnitude and the number of reporting stations. The estimated spatial intensity surface highlights seismic hotspots consistent with subduction interfaces and tectonic boundaries. To capture directional behavior, a unified circular statistical framework is introduced, incorporating bearing computation, uniformity tests (Rayleigh, Kuiper, Watson), circular-linear regression, and finite mixture modeling using mixed von Mises distributions. This enables decomposition of complex directional patterns into interpretable fault-related clusters and identification of depth-direction coupling. Applied to the Fiji trench, the method detects SW–NE bimodality, a dominant orientation near 127◦, a secondary cluster near 292◦, and significant regression effects ( ˆβ = 0.036, p < 0.001). The findings underscore the utility of classical datasets for modern geostatistical workflows and highlight the value of open-access seismic data in understanding global tectonic processes. All analyses were performed using reproducible code. Journal of Statistical Research 2025, Vol. 59, No. 2, pp. 145-166.
This study introduces an intelligent hybrid estimation framework for modeling the time-dependent failure behavior of repairable systems. The proposed approach embeds the Inverse Weibull Process (IWP) within a Non-Homogeneous Poisson Process (NHPP) structure and compares the performance of the traditional Maximum Likelihood Estimation (MLE) method with a hybrid algorithm that integrates Artificial Neural Networks and the Artificial Bee Colony optimization technique (ANN-ABC). Both simulation experiments and an application to real clinical data from 299 patients with heart failure demonstrate that the ANN–ABC estimator achieves lower Root Mean Squared Error (RMSE) and Bayesian Information Criterion (BIC) values, particularly for moderate to large datasets. These findings highlight the potential of hybrid intelligent methods as robust alternatives to conventional estimation techniques, offer improved precision in modeling failure intensities, and support predictive maintenance as well as intelligent healthcare monitoring systems. Journal of Statistical Research 2026, Vol. 60, No. 1, pp. 33-50.
In this paper, we consider the problem of the Gamma regression model under left-censored data with covariates. The method investigated consists of solving left-censored maximum likelihood estimating equations. We show that the resulting estimates are asymptotically normal. A simulation study assesses the proposed parameters’ finite-sample properties and the root mean square error estimates. An application using car insurance data is presented to estimate the covariates coefficients in calculating the provisions for claims to be paid. We examine the effect of the censoring variable on the calculation of provisions. We will employ a machine learning algorithm called Random Forest to show the impact of the presence of the censoring variable. Finally, we address financial risk management that considers the Value at Risk (VaR), the Expected Shortfall (ES), and the backtesting of the VaR. Journal of Statistical Research 2025, Vol. 59, No. 2, pp. 221-247.
This study presents a new approach for smooth estimation of distribution and density functions by leveraging beta regression and generalized additive models (GAM). The approach estimates both functions by smoothing the first derivative of left mean absolute deviation (MAD) function under the condition of nondecreasing distribution function, and nonnegativity of the density. This is achieved by using beta regression and generalized additive models with various link functions (logit, probit, cloglog, and cauchit) that are applied to a polynomial function where the degree is selected based on minimum mean absolute regression error and positivity of the first derivative. Additionally, the confidence limits for the distribution function are derived based on normal approximation and the beta distribution to assess the precision of estimates. The proposed method is evaluated on simulated datasets featuring unimodal, multimodal and real datasets. The results suggest that the proposed method exhibits superior performance compared to the kernel-based estimators, particularly in terms of smoothness and accuracy for small sample sizes. Journal of Statistical Research 2026, Vol. 60, No. 1, pp. 133-155.
In this article, we investigate methods for estimating the residual entropy function of a length-biased sample. We show that the proposed estimators are strongly consistent and asymptotically normal under suitable regularity conditions. We also conduct simulations to evaluate the performance of the proposed estimation techniques and provide an example to show how these techniques can be applied to real data. Journal of Statistical Research 2026, Vol. 60, No. 1, pp. 197-211.
Survey sampling heavily relies on the availability of a robust sampling frame, which is often difficult to obtain, leading to undercoverage bias and misleading results. In the present era, an incomplete sampling frame poses a significant challenge, affecting survey results and findings. Literature offers widely used solutions to this pervasive issue that include multiple framework approach and post-survey adjustments. The predecessor-successor (PS) method proposed by Hansen et al. (1963) provides a way to uniquely identify the units not listed in the sampling frame by linking them to the existing units. Despite its potential to mitigate the bias arising from incomplete sampling frames, the P-S method was not explored extensively until it was later formulated mathematically by Singh (1983), Singh (1989) and others. Building on this legacy and addressing the limitations of simple random sampling in diverse populations, this paper introduces an estimator for the population mean within a stratified sampling framework. The properties of the proposed estimator are discussed and the findings are further supported by an empirical study as well as a simulation study. Journal of Statistical Research 2026, Vol. 60, No. 1, pp. 79-90.
One of the dimensions of a pension system’s performance is coverage. In most countries, pension systems require a minimum number of years of contribution as part of their eligibility requirements. For example, in Argentina, a minimum of 30 years of contributions is required. Therefore, it is not enough to study what percentage of the target population is covered at a given time; it is also necessary to study employment histories over a considerable period of time. Sometimes, developing countries do not have sufficient information for this purpose, but they incorporate new information as they digitize their records. The Bayesian approach can be useful in these cases where information is limited but regularly updated. The objective of this paper is to demonstrate the usefulness of the sequential Bayesian approach for estimating the proportion of workers eligible to retire in Argentina year after year. It is observed that as more information is incorporated, the proportion of people who remain active contributors (and, therefore, eligible for a pension) decreases. This implies that, for any individual, the probability of meeting the contribution requirement decreases as time passes from the first month of contribution. At the end of the process (once all available information has been incorporated), the Bayesian proportion estimate is different from that obtained with a frequentist approach, which is explained by the importance of a priori information provided by prior knowledge about the phenomenon. This type of sequential estimation exercise may be of interest to social security decision-makers. Journal of Statistical Research 2025, Vol. 59, No. 1, pp. 249-257
This article obtains locally R-optimal designs for the Poisson regression model using the square-root link function. In the generalized linear model (GLM) configuration, the information matrix is determined by the model’s unknown parameters. In such cases, an experimenter must use the strategy of discovering local optimum designs, which entails first guessing the best value for the parameters and then calculating the optimal designs. The R-optimality criterion has been proposed in the literature as an alternative to the most frequently used D-optimality criterion when the experimenter wishes to minimize the volume of the confidence region for unknown parameters based on Bonferroni t-intervals. The necessary and sufficient conditions of this optimality criterion are verified through the equivalence theorem. All numerical computations were performed using Mathematica 7.0. Journal of Statistical Research 2026, Vol. 60, No. 1, pp. 157-173.
Estimating density derivatives is a powerful technique in statistical data analysis. It has diverse applications in machine learning, signal processing, and statistical analysis. The kernel method is one of the most popular methods in nonparametric density derivative estimation, but this estimator is biased and is not consistent when the data are near the endpoints of the support. This paper investigates the challenge of estimating the first-order derivative of an unknown probability density function defined on the interval [0, 1]. We focus our study near the right boundary. The asymptotic properties are derived. A Monte Carlo study and real data example are provided to illustrate the finite sample performance of the proposed estimator. Journal of Statistical Research 2025, Vol. 59, No. 2, pp. 183-201
Understanding the timing of the first antenatal care (ANC) visit requires statistical methods capable of handling bounded, skewed, and heterogeneous count outcomes subject to truncation. Conventional count models such as the Poisson and Negative Binomial are often inadequate in this setting due to their limited ability to accommodate simultaneous truncation and dispersion effects. This paper introduces the Truncated Beta-Geometric (TBG) distribution as a flexible modelling framework for truncated discrete data. The TBG model, derived by truncating the standard Beta-Geometric distribution, accommodates overdispersion, underdispersion, and bounded count outcomes, addressing heterogeneity in healthseeking behavior. Key statistical properties of the TBG distribution, including its probability mass function, mean, and variance, are presented along with a numerical parameter estimation method using maximum likelihood. A Monte Carlo simulation study evaluated the performance of the model under different sample sizes, truncation intervals, and parameter settings, demonstrating that estimator bias stems more from structural features than sample size. The model is applied to data from the 2000 Oman National Health Survey, where timing to first ANC visit was recorded across 1,299 women having ANC visits. The results demonstrate that TBG outperformed the untruncated Beta-Geometric, in terms of AIC, BIC, and goodness-of-fit. The model is further extended to a regression framework, allowing covariate inclusion. The results demonstrate significant associations of ANC timing with maternal age, education, parity, and urban residence. The TBG regression model demonstrated superior flexibility and interpretability, establishing it as a robust tool for modeling truncated count outcomes in public health research. Journal of Statistical Research 2026, Vol. 60, No. 1, pp. 91-111.
Trimmed mean has been considered a robust estimator of location parameters over the last five decades. The issue of outlier detection has been considered using analytical and graphical statistical tools. This article proposes a graphical device based on the trimmed mean to check whether data has any outliers or not, and in the presence of outliers, the proposed graphical device enables the estimation of the proportion of the outliers as well. The extension of the methodology to the high dimensional data is also outlined. Furthermore, the proposed visualization toolkit is implemented on economic data and gives us an idea of the presence/absence of influential observations/outliers. Journal of Statistical Research 2025, Vol. 59, No. 1, pp. 9-30
In mathematics a random walk (also known as drunkard’s walk) is a succession of random steps. In 1905 Karl Pearson introduced the term “random walk”. A Bernoulli random walk is the random walk on the integer number line Z which starts at 0 and at each step moves +1 or −1 with equal probability. A Pearsonian random walk is a walk in the plane that starts at the origin 0 and consists of length 1 taken in uniformly random direction. In this paper several known and new results of Bernoulli and Pearsonian walks will be presented. Journal of Statistical Research 2025, Vol. 59, No. 1, pp. 47-64
The estimation of finite population mean is always of interest for different sampling techniques and it is the basic measure to find from sample to estimate one the most applicable central tendency. In literature, under simple random sampling without replacement people used auxiliary variable, its rank or empirical distribution function in different estimation approaches such as regression, ratio, exponential or combination of these to improve the efficiency of the estimator. In literature, either rank or empirical distribution function have been used while constructing the estimator because both cannot be used due to the fact that empirical distribution function of a variable is based on its rank, therefore, both are perfectly correlated. In this paper, our argument is that the dual use is not effective rather an additional independent auxiliary variable may be effective for efficiency improvement. To investigate this, we proposed difference-cum-exponential estimator using two auxiliary variables and also the dual use of one of the auxiliary variable in the form of its empirical distribution function. We also deduced some special cases of the proposed estimator. These special cases will help us to investigate the argument. The mean square errors of the proposed estimator and its special cases are derived. The proposed estimator, its special cases and potential existing estimators are compared using empirical study based on real life population for numerical investigation of the argument. The simulation study is also conducted for symmetric and skewed populations to asses the sampling stability of the competitive estimators using empirical mean square error and also it will help to further investigate the argument. Journal of Statistical Research 2025, Vol. 59, No. 1, pp. 81-97
Concentration inequalities involving tail behavior of distributions have currently become very popular owing to their applications in the growing machine learning literature. The paper offers one such inequality for the less known, but equally important inverse Gaussian distribution. Journal of Statistical Research 2025, Vol. 59, No. 1, pp. 5-8