
Finite population (FP :ℱ ) total prediction using a survey sample (SS :s^* ) taken from the ℱ using a suitable sampling design (D_s^*) has been an important topic of research interest in survey sampling over the last five decades. In general a model based approach is followed for such a prediction where the ℱ based hypothetical data involving the responses and covariates are assumed to follow a super-population (SP :𝒮 ) regression model, and the non-sampled (ns) response total is predicted by its 𝒮 model based expectation. Next, as the expected total contains unknown model parameters such as regression effects, an optimal prediction estimate is supposed to be obtained by optimally estimating the 𝒮 regression parameters using the SS (s^* ⊂ℱ) taken from the FP. For the purpose, the decades long existing studies, have been using the sample ( s^* ) based conditional (using conditioning on the sample) OLS/GLS (ordinary/generalized least square) estimates for the 𝒮 regression parameters in the independence and correlation setups, respectively. However as we demonstrate in this paper in the context of a binary population total (equivalent to proportion) prediction, in contrary to the claims made in the existing studies, the aforementioned OLS/GLS estimates fail to be model unbiased for the 𝒮 regression parameters, and hence they fail to produce valid MU prediction for the total. More specifically, it is shown that the OLS estimates are not MU, rather they are design cum model unbiased (DCMU); whereas the GLS estimates are neither MU nor DCMU. As a remedy, the sampling weighted OLS (SWOLS) estimates are used (as they are DCMU) for optimal prediction for independent binary data, whereas a new survey sample based doubly weighted (SSDW) regression estimates are derived which are DCMU under the correlation setup and the produce DCMU prediction. This DCM based principle based on the sampling sequence (s^* ⊂ℱ⊂𝒮) is further implemented to obtain variance of the regression effects estimators and subsequently the variance of the DCMU predictor for the total. The asymptotic properties such as consistency of the DCMU predictor is also established.
We study nonlinear processes for which quadratic predictors provide superior forecasts to using linear predictors alone. For this class, designated as second order forecastable, we develop theoretical results for quadratic prediction, offering three main contributions: (1) a new theory for infinite arrays and their factorization is developed, generalizing multivariate spectral factorization theory; (2) two formulas for the quadratic h-step ahead forecast filter (based on an infinite past) are derived; (3) a necessary and sufficient condition involving the bispectrum is derived, for discerning whether quadratic prediction offers a benefit over linear prediction. For the two quadratic forecast filters, the first is an exact expression whereas the second involves some implicitly defined quantities, and is more practical for computation. The new results are numerically illustrated on several nonlinear processes. Finally, the quadratic forecasting filter is applied to survey time series data of the U.S. Census Bureau’s Annual Retail Trade Survey, and the forecasts are used to detect survey discontinuity potentially arising from reclassification.
For clinical trials that address questions of non-inferiority or superiority, the null hypothesis specifies a non-zero difference between response probabilities p_0 ( p_1 ) under control(treatment) regimens. A confidence interval for p_0 under this hypothesis provides a summary of our knowledge of the baseline probability p_0 , which can easily be converted into an interval for the treatment probability p_1 . When the hypothesised difference between p_0 and p_1 is zero, there is an exact interval for p_0 based on the total number of responses known as the Clopper-Pearson interval. Unfortunately, this construction does not extend to the hypothesis of non-zero difference. There is no quickly computed accurate method currently available. In this paper, we propose and investigate the properties of a conceptually simple and computationally feasible interval and demonstrate that it has reasonable one-sided and excellent two-sided coverage properties. This is not only useful as a summary of post-experimental knowledge, but is a key input into the testing procedure of Berger and Boos (1994).
We investigate the problem of nonparametric density estimation on ℝ_+ in the presence of incomplete responses under a Missing At Random (MAR) mechanism. Since standard symmetric kernel estimators are known to suffer from severe boundary effects on positive supports, we develop an inverse-probability-weighted gamma kernel methodology specifically designed to accommodate simultaneously the support constraint and the missingness structure. The proposed procedure is built from an ideal reweighted estimator involving the unknown observation probabilities and from a feasible estimator obtained by replacing these propensity scores with a nonparametric Nadaraya–Watson-type estimate based on fully observed auxiliary covariates. Under suitable regularity and smoothing assumptions, we derive a detailed asymptotic theory for both estimators. In particular, we establish explicit pointwise expansions for the bias and variance, obtain the corresponding mean squared error and mean integrated squared error, and prove asymptotic normality in the interior region. These results extend the classical gamma-kernel framework for complete data to the substantially more delicate setting of MAR incompleteness, and they quantify the additional stochastic cost induced by propensity-score estimation. The finite-sample behavior of the method is assessed through an extensive Monte Carlo study covering several positive-supported models, varying dependence levels between the study variable and the auxiliary covariates, and multiple missingness rates. The numerical evidence shows that the proposed estimator retains the boundary-adaptive advantages of gamma kernels while providing stable and accurate recovery of the target density under incomplete observation. A real-data analysis further illustrates its practical relevance for distributional inference with positive-valued variables subject to missingness.
Let N=∏ _i=1^kp_i^a_i be a fixed positive integer with known prime factorization. We investigate bootstrap resampling from the prime-factor multiset of N and study two associated statistics, namely the log-product of the resampled integer and the number of distinct observed primes appearing in the resample. It is shown that the log-product has explicit mean and variance determined by the empirical law of the logarithms of the observed primes, while bootstrap- ω is an occupancy sum with explicit mean, variance, and negative pairwise covariance structure. In the square-free case, exact formulas are obtained and asymptotic constants are identified, yielding a benchmark on the Erdős–Kac scale in random-like regimes. Numerical experiments on a synthetic corpus of factorizations illustrate the formulas, the square-free benchmark, and the behavior of the diagnostics across smooth, square-free, spiky, and mixed factorization regimes.
This paper extends the generalized linear model framework to composite responses of the form Z = X + Y , where X and Y follow different distributions. The proposed approach enhances flexibility in modeling composite response variables arising from heterogeneous components. We study the natural exponential family associated with the composite response and show that the corresponding link and variance functions depend on a first-order nonlinear differential equation. Explicit solutions are obtained only in the Poisson–normal and exponential–normal cases. For the general case, we propose approximation and estimation methods based on Runge–Kutta approach, nonparametric techniques, and the Lagrange inversion formula. Simulation results demonstrate satisfactory performance in estimating regression coefficients. An application to insurance data illustrates the practical relevance of the proposed methodology.
Estimating the parameters of the generalized logistic distribution has posed a challenge for many inference approaches, particularly in the frequentist approach. Surprisingly, there have been very few attempts to study the objective Bayesian analysis approach for this problem. In our study, we investigate objective Bayesian analysis using Jeffreys prior, reference priors, matching priors, and the maximal data information prior. We demonstrate that using these non-informative priors does not result in proper posterior distributions. Additionally, we develop a Bayesian analysis based on reference priors with partial information, which yields proper posterior distributions. To evaluate the performance of these priors, we conduct a small Markov Chain Monte Carlo (MCMC) study examining three of these prior distributions. The results demonstrate strong performance in terms of mean squared error and coverage probability. Finally, we utilize these priors to obtain estimations and credible sets for the distribution parameters in a specific example.
The Kolmogorov’s 0-1 law is so powerful and fundamental in Probability theory that the result of the law of large numbers is its corollary. The latter is the theoretical bedrock of modern day statistical machine learning. Expected resurgence in the interest of the law’s proof is challenged by its measure theoretic nature. By observing 0-1 law as False-Truth law like in logic, we show here that the proof of the law is simplified. This simplicity comes from the fact that our approach hardly needs any measure theory beyond defining an obvious probability space. A novel contribution via this article is the use of logic within the proof of a fundamental result in the theory of probability.
This paper presents an investigation of the asymptotic behavior within natural exponential families, with particular focus on the limiting distributions of scaled family. We examine the scaled random variable Y_t = X_t/t , where X_t belongs to a natural exponential family with mean parameter tm ( m > 0 ) as t → 0^+ . Under the fundamental assumption that the second derivative of the variance function V extends continuously to zero with V”(0) 0 , we establish that Y_t converges in distribution to a Gamma law with explicitly determined shape and rate parameters α = 2/V”(0) and β = 2/m V”(0) , respectively. The proposed proof integrates tools from convex analysis, probability theory, and functional analysis, including the Arzelà-Ascoli theorem. The paper also provides a detailed illustrative example using the uniform distribution on [0, 1], where the limiting distribution simplifies to an exponential law.
Bayrakdar, Kuş and Ng [Journal of Quality Technology, 2026, 1–18, doi: 10.1080/00224065.2025.2612362] introduced the shape-extended continuous Bernoulli distribution and derived series representations for its rth raw moment, its moment generating function and, under the restriction γ _1 =γ _2 , its stress-strength reliability θ= (X < Y) . Each of these representations involves one or more infinite sums that must be truncated for practical use, raising questions of convergence, accuracy and computational efficiency. In this note, we resolve all three issues simultaneously by expressing each quantity in terms of well-known special functions. Specifically, we show that the rth raw moment and the moment generating function admit exact closed forms involving the confluent hypergeometric function _1F_1 and the Srivastava-Daoust generalized Kampé de Fériet function 𝒮 , respectively, both of which are implemented to arbitrary precision in all major computer-algebra systems. We also derive a closed form expression for the stress-strength reliability θ = (X < Y) that places no restriction on the shape parameters γ _1 and γ _2 , thereby extending the result of Bayrakdar et al. (2026).
Likelihood-free inference for simulator-based statistical models has recently attracted an interest, both in the machine learning and statistics. The primary focus of has been to approximate the posterior distribution of model parameters, either by various types of Monte Carlo sampling algorithms or deep neural network -based surrogate models. Frequentist inference for simulator-based models has been given much less attention to date, despite that it would be particularly amenable to applications, where implicit asymptotic approximation of the likelihood is expected to be accurate and can leverage computationally efficient strategies. Here we derive a set of results to enable estimation, hypothesis testing and construction of confidence intervals for model parameters using asymptotic properties of the Jensen–Shannon divergence. Such asymptotic approximation offers a rapid alternative to more computation-intensive approaches and can be attractive for diverse applications of simulator-based models.
In this paper, we construct a new family of infinitely divisible probability measures, concentrated on a proper cone, with a flexible parametrization. This construction permits to define from any infinitely divisible distribution on the real line, a new multivariate cone-valued probability measure. A characterization of a sub-class of these distributions satisfying a sufficient condition in terms of the Lévy measure, is established. Moreover, several properties, such as the explicit form of the Lévy measure and the closedness under scaling and convolution, are provided. These results extend previous constructions of cone valued gamma and the multivariate stable distributions. The proposed approach is used to build a generalized cone valued inverse-Gaussian distribution, for which, a numerical illustration is performed.
The five-parameter bivariate Birnbaum-Saunders and bivariate lognormal distributions are two models used to study positive lifetime bivariate data. These distributions exhibit striking similarities, from the surface plot of their joint density functions to the nature of their reliability functions and hazard gradients within certain parameter spaces. Additionally, their marginal distributions share similar shapes. This paper aims to discriminate between these two distributions. To achieve this, we use the difference between the maximized log-likelihood functions as a discrimination statistic. The asymptotic distribution of the test statistic is derived to compute the probability of correct selection (PCS). Moreover, we have also studied the effects of model misspecification on the mode of the marginal hazard function as well as on the components of the hazard gradient. Due to the limitations of the usual likelihood approach, we also propose a suitable modification for the decision rule. Finally, we do a real data analysis.
We study the maximum likelihood estimator of the location parameter of the Pearson Type VII distribution with known scale. We rigorously establish precise asymptotic properties such as strong consistency, asymptotic normality, Bahadur efficiency and asymptotic variance of the maximum likelihood estimator. Our focus is the heavy-tailed case, including the Cauchy distribution. The main difficulty lies in the fact that the likelihood equation may have multiple roots; nevertheless, the maximum likelihood estimator performs well for large samples.
We consider the modulation of data given by random vectors X_n ∈ℝ^d_n , n ∈ℕ . For each X_n , one chooses an independent modulating random vector Ξ _n ∈ℝ^d_n and forms the projection Y_n = Ξ _n'X_n . It is shown, under regularity conditions on X_n and Ξ _n , that Y_n|Ξ _n converges weakly in probability to a normal distribution. More broadly, the conditional joint distribution of a family of projections constructed from random samples from X_n and Ξ _n is shown to converge weakly to a matrix normal distribution. We derive, via G. Pólya’s characterization of the normal distribution, a necessary and sufficient condition on Y_n for Ξ _n to be normally distributed, and we show that our results motivate generalizations of Pólya’s theorem. When Ξ _n has a spherically symmetric distribution we deduce, through I. J. Schoenberg’s characterization of the spherically symmetric characteristic functions on Hilbert spaces, that the probability density function of Y_n|Ξ _n converges pointwise in certain pth means to a mixture of normal densities and a rate of convergence is quantified, resulting in uniform convergence. The cumulative distribution function of Y_n|Ξ _n is shown to converge uniformly in those pth means to the distribution function of the same mixture, and a Lipschitz property is obtained. Examples of distributions for X_n that satisfy our results include the Bingham distributions on hyperspheres of random radii, uniform distributions on hyperspheres and hypercubes of random volumes, and multivariate normal distributions; and examples of such Ξ _n include the multivariate t-, multivariate Laplace, and spherically symmetric stable distributions.
We introduce a nonasymptotic framework for sub-Poisson distributions with moment generating function dominated by that of a Poisson distribution. At its core is a new notion of optimal sub-Poisson variance proxy, analogous to the variance parameter in the sub-Gaussian setting. This framework allows us to derive a Bennett-type concentration inequality without boundedness assumptions and to show that the sub-Poisson property is closed under key operations including independent sums and convex combinations, but not under all linear operations such as scalar multiplication. We derive bounds relating the sub-Poisson variance proxy to sub-Gaussian and sub-exponential Orlicz norms. Taken together, these results unify the treatment of Bernoulli and Poisson random variables and their signed versions in their natural tail regime.
Likelihood-free inference for simulator-based statistical models has recently attracted an interest, both in the machine learning and statistics. The primary focus of has been to approximate the posterior distribution of model parameters, either by various types of Monte Carlo sampling algorithms or deep neural network -based surrogate models. Frequentist inference for simulator-based models has been given much less attention to date, despite that it would be particularly amenable to applications, where implicit asymptotic approximation of the likelihood is expected to be accurate and can leverage computationally efficient strategies. Here we derive a set of results to enable estimation, hypothesis testing and construction of confidence intervals for model parameters using asymptotic properties of the Jensen–Shannon divergence. Such asymptotic approximation offers a rapid alternative to more computation-intensive approaches and can be attractive for diverse applications of simulator-based models.
For a univariate monotone regression function, the location where a specific value is attained is called a regression quantile. We study the coverage of a Bayesian credible interval for a regression quantile in a nonparametric monotone regression model, assuming that the quantile is unique and the regression function has a positive derivative. We consider piecewise constant functions with equal intervals and put independent normal priors on the step heights. To comply with the monotonicity constraint for the regression function, we induce a “projection-posterior” by imposing the monotonicity constraint on samples from the posterior distribution of the step-heights. We demonstrate two different interesting phenomena in this context. First, we show that the asymptotic coverage of a credible interval is higher than the credibility, the opposite of a phenomenon observed for credible regions for smooth functions. Further, targeted asymptotic coverage may be obtained using an appropriate lower credibility level. Next, we show that the posterior contraction rate for the regression quantile can be improved from n^-1/3 to n^-1/2 by sampling in two stages, sparing a fraction of the sampling budget to sample later from a credible interval obtained in the first stage. We also show that the coverage of a second-stage credible interval agrees with its credibility. We study the finite sample performance of the methods through a simulation study.