We examine Markov's inequality in the light of renewal theory and reliability theory. Suppose the non negative random variable (rv) X has cumulative distribution function (cdf) F with survival function (F) over bar := Pr (X > x) = 1 - F and left-continuous version of the survival function (F) over bar (x(-)) := Pr(X >= x), x >= 0. We determine the points, if any, such that x >= mu and (F) over bar (x(-)) = mu/x. We offer an alternative proof of Markov's inequality by observing that, if some collection of events {A(t) : t >= 0} exists such that Pr(A(t)) = t (F) over bar (t(-))/mu, then because t (F) over bar (t(-))/mu equals a probability, it must satisfy t (F) over bar (t(-))/mu <= 1, which is equivalent to Markov's inequality. We choose events connected to stationary renewal processes. When we know only the sample size n and the sample average (x) over bar (n) of n non negative observations x(1), ..., x(n), we establish an upper bound on the left-continuous version of the empirical survival function that improves Markov's inequality. We show that an upper bound of Markov type for the survival function is sharp when F is "new better than used in expectation" (NBUE) or has "decreasing mean residual life" (DMRL).
We present a general framework for prediction in which a prediction is in the form of a distribution function, called ‘predictive distribution function’. This predictive distribution function is well suited for prescribing the notion of confidence under the frequentist interpretation and providing meaningful answers for prediction-related questions. Its very form of a distribution function also makes itself a useful tool for quantifying uncertainty in prediction. This general framework is formulated and illustrated using the so-called confidence distributions (CDs). This CD-based prediction approach inherits many desirable properties of CD, including its capacity to serve as a common platform for directly connecting the existing procedures of predictive inference in Bayesian, fiducial and frequentist paradigms. We discuss the theory underlying the CD-based predictive distribution and related efficiency and optimality. We also propose a simple yet broadly applicable Monte-Carlo algorithm for implementing the proposed approach. This concrete algorithm together with the proposed definition and associated theoretical development provide a comprehensive statistical inference framework for prediction. Finally, the approach is illustrated by simulation studies and a real project on predicting the application submissions to a government agency. The latter shows the applicability of the proposed approach to even dependent data settings
SignificanceMany quantities are extremely large extremely rarely. Examples include income, wealth, financial returns, insurance losses, firm size, and city population size; earthquake magnitude, hurricane energy, tornado outbreaks, precipitation, and flooding; and pest outbreaks, infectious epidemics, and forest fires. When such a quantity is modeled as a nonnegative random variable with a heavy upper tail, the probability of an observation larger than some threshold falls as a small power (the “tail index”) of the threshold. When the tail index is small enough, the mean and all higher moments of the random quantity are infinite. Surprisingly, the sample mean and the sample higher moments obey orderly scaling laws, which we prove and apply to estimating the tail index.
Many cryptocurrencies including Bitcoin are susceptible to a so-called double-spend attack, where someone dishonestly attempts to reverse a recently confirmed transaction. The duration and likelihood of success of such an attack depends on the recency of the transaction and the computational power of the attacker, and these can be related to the distribution of time for counts from one Poisson process to exceed counts from another by some desired amount. We derive an exact expression for this distribution and show how it can be used to obtain efficient simulation estimators. We also give closed-form analytic approximations and illustrate their accuracy.
Let F be an NWUE distribution with mean 1 and G be the stationary renewal distribution of F. We would expect G to converge in distribution to the unit exponential distribution as its mean goes to 1. In this paper, we derive sharp bounds for the Kolmogorov distance between G and the unit exponential distribution, as well as between G and an exponential distribution with the same mean as G. We apply the bounds to geometric convolutions and to first passage times.
In a family, parameterized by θ, of non-negative random variables with finite, positive second moment, Taylor's law (TL) asserts that the population variance is proportional to a power of the population mean as θ varies: σ 2 (θ) = a [μ(θ)] b , a > 0. TL, sometimes called fluctuation scaling, holds widely in science, probability theory, and stochastic processes. Here we report diverse examples of TL with b = 2 (equivalent to a constant coefficient of variation) arising from a difference of random variables in normed vector spaces of dimension 1 and larger. In these examples, we compute a exactly using, in some cases, a simple, new technique. These examples may prove useful in future models that involve differences of random variables, including models of the spatial distribution and migration of human populations.
In this paper we provide an overview as well as new (definitive) results of an approach to boundary crossing. The first published results in this direction appeared in de la Peña and Giné (1999) book on decoupling. They include order of magnitude bounds for the first hitting time of the norm of continuous Banach-Space valued processes with independent increments. One of our main results is a sharp lower bound for the first hitting time of càdlàg real-valued processes X(t), where X(0)=0 with arbitrary dependence structure: ETrγ≥∫01{a−1(rα)}γdα, where Tr=inf{t>0:X(t)≥r},a(t)=E{sup0≤s≤tX(s)} and γ>0. Under certain extra conditions, we also obtain an upper bound for ETrγ. As the main text suggests, although Tr is defined as the hitting time of X(t) hitting a level boundary, the bounds developed can be extended to more general processes and boundaries. We shall illustrate applications of the bounds derived for additive processes, Gaussian Processes, Bessel Processes, Bessel bridges among others. By considering the non-random function a(t), we can show that in various situations, ETr≈a−1(r).
One problem of wide interest involves estimating expected crossing-times. Several tools have been developed to solve this problem beginning with the works of Wald and the theory of sequential analysis. Deriving the explicit close form solution for the expected crossing times may be difficult. In this paper, we provide a framework that can be used to estimate expected crossing times of arbitrary stochastic processes. Our key assumption is the knowledge of the average behavior of the supremum of the process. Our results include a universal sharp lower bound on the expected crossing times. Furthermore, for a wide class of time-homogeneous, Markov processes, including Bessel processes, we are able to derive an upper boundE[a(Tr)]≤2r, which implies that supr>0|((E[a(Tr)]−r)/r)|≤1, wherea(t)=E[suptXt] with {Xt}t≥0be a non-negative, measurable process. This inequality motivates our claim thata(t) can be viewed as a natural clock for all such processes. The cases of multidimensional processes, non-symmetric and random boundaries are handled as well. We also present applications of these bounds on renewal processes in Example 10 and other stochastic processes.
Suppose we have n objects of different weights. We randomly sample pairs of objects, and for each sampled pair use a balance scale to determine which of the two objects is heavier. It is assumed that the sequence of sampled pairs is iid, each selection uniformly distributed on the set of n(n−1)/2 pairs. We continue sampling until the first time that we can definitively identify the heaviest of the n objects. The problem of interest is to compute the expected number of selected pairs.
Let F be a new better than used in expectation (NBUE) distribution function with mean μ . In a previous paper (Brown in Probab. Eng. Inf. Sci. 20:195–230, 2006 ), the author derived the following bound. For any t ≥ μ , F(t) = 𝑃𝑟(X ≥ t) ≤ e^-[tμ-1]. The main result of this paper is to show that this bound is sharp. Other sharp bounds for NBUE distributions are also derived.
A random walk that is skip-free to the left can only move down one level at a time but can skip up several levels. Such random walk features prominently in many places in applied probability including queuing theory and the theory of branching processes. This article exploits the special structure in this class of random walk to obtain a number of simplified derivations for results that are much more difficult in general cases. Although some of the results in this article have appeared elsewhere, our proof approach is different.
The LRMC method for predicting NCAA Tournament results from regular-season game outcomes is a two-part process consisting of a logistic regression model to estimate head-to-head differences in team strength, followed by a Markov chain model to combine those differences into an overall ranking We consider replacing each of the two parts of LRMC with alternative models, empirical Bayes and ordinary least squares, that attempt to accomplish the same goal. Computational results show that replacing the logistic regression with either of two empirical Bayes models yields a statistically-significant improvement when the probabilities are jointly conditioned.
We consider order statistics corresponding to X1, …, Xn, where $X_{i} \sim \lambda_{i}^{-1} \cal{E}$, i = 1, …, n, ℰ1, …, ℰn are independent and identically distributed exponentials with mean 1, and λ1, …, λn are possibly dependent, possibly nonidentically distributed, positive random variables, with $E\lambda_{i}^{-1} < \infty$. Thus, λ1, …, λn, can be interpreted as random failure rates, and their dependency might be due to common environmental factors.
We consider the classical coupon collector's problem in which each new coupon collected is type i with probability p i ; ∑ i =1 n p i =1. We derive some formulas concerning N , the number of coupons needed to have a complete set of at least one of each type, that are computationally useful when n is not too large. We also present efficient simulation procedures for determining P ( N > k ), as well as analytic bounds for this probability.
Consider a continuous nonnegative random variable X with mean μ and hazard function h. Assume further that a≤h(t)≤b for all t≥0. Under these constraints, we obtain sharp two-sided bounds for $\bar{F}(t) = \hbox{Pr}(X \gt t)$. An application to birth and death processes is discussed.
We study a model arising in chemistry where n elements numbered 1, 2, …, n are randomly permuted and if i is immediately to the left of i + 1 then they become stuck together to form a cluster. The resulting clusters are then numbered and considered as elements, and this process keeps repeating until only a single cluster is remaining. In this article we study properties of the distribution of the number of permutations required.
We consider the transformation T that takes a distribution F into the distribution of the length of the interval covering a fixed point in the stationary renewal process corresponding to F. This transformation has been referred to as size-biasing, length-biasing, the renewal length transformation, and the stationary lifetime operator. We review and develop properties of this transformation and apply it to diverse areas.
For a normal sample with unknown mean, the almost universally used estimator of the variance, σ2, is “the sample variance.” This estimator is the minumum variance unbiased estimator of σ2, but it is inadmissible under square error loss. It is dominated by the maximum likelihood estimator, which is also inadmissible. We consider a class of estimators and compare these estimators under a class of loss functions which we call “log symmetric.”