In a data set of non-metallic inclusion sizes in samples from engineering steel, common order statistics fail to serve as a suitable model for ascendingly ordered measurements within single samples. Therefore, a flexible model of ordered random variables is proposed, which allows for changes of distributions described by model parameters. Joint maximum likelihood estimation of these parameters and the shape parameter of an underlying left-truncated Weibull distribution is considered, and a model test is developed for the null-hypothesis of common order statistics being an adequate model. To overcome small data situations, a link-function approach is examined in order to reduce the number of involved model parameters as well as to propose to use a link-function parameter as a material indicator. An asymptotic test is provided to check for the presence of a linear link function, and tests for hypotheses about two link-function parameters are studied. Moreover, the construction of simultaneous confidence regions for the link-function parameters as well as of confidence bands for the entire graph of the link function are presented. Throughout, the findings are applied to the real metallurgical data set. Similar problems and data structures arise in other fields of material science and applications such as geology.
Load-sharing systems arise in many different reliability applications, for instance, when modeling tensile strength of fibrous composites in textile industry or lifetimes of redundant technical systems in engineering. Sequential order statistics serve as a flexible model for the ordered component failure times of such systems and allow the residual lifetime distribution of the components to change after each component failure. In a proportional hazard rate setting, the model consists of some baseline distribution function and several model parameters describing successive adjustments of the hazard rates of the operating components. This work provides nonparametric confidence bands for the baseline distribution function, where the model parameters may be known or unknown. In case of known model parameters, we show how to construct exact confidence bands based on Kolmogorov-Smirnov type statistics, which are distribution-free with respect to the baseline distribution. If the model parameters are unknown, finite sample inference turns out to be infeasible, and asymptotic confidence bands for the baseline distribution function are derived. As a technical tool, we extend the existing asymptotic theory of semiparametric estimators based on the profile-likelihood approach.
As a flexible extension of the common Poisson model, the Conway–Maxwell–Poisson distribution allows for describing under- and overdispersion in count data via an additional parameter. Estimation methods for two Conway–Maxwell–Poisson parameters are then required to specify the model. In this work, two characterization results are provided related to maximum likelihood estimation of the Conway–Maxwell–Poisson parameters. The first states that maximum likelihood estimation fails if and only if the range of the observations is less than two. Assuming that the maximum likelihood estimate exists, the second result then comprises a simple necessary and sufficient condition for the maximum likelihood estimate to be a solution of the likelihood equation; otherwise it lies on the boundary of the parameter set. A simulation study is carried out to investigate the accuracy of the maximum likelihood estimate in dependence of the range of the underlying observations.
The Conway-Maxwell-Poisson (CMP) distribution has been introduced as a generalization of the common Poisson distribution and is widely used in applications, where underdispersed or overdispersed count data are recorded. To fit the model to a given data set, inferential methods are then used to assess two CMP parameters. In this paper, uniformly most powerful unbiased (UMPU) tests are developed for two-sided hypotheses and interval hypotheses about one of the CMP parameters, where the other CMP parameter is assumed to be unknown, too. These tests are exact (non-asymptotic) and do not involve joint estimation of the CMP parameters. Finding the critical values of the test statistics requires some computational effort and turns out to be more complex than for existing UMPU tests for one-sided hypotheses, but they can be obtained by solving certain optimization problems. As a prominent example, the tests can be applied for a model check in the sense whether the common Poisson distribution is adequate for describing the data or has to be rejected in favour of the CMP distribution. In this context, a simulation study is carried out to compare the power of the UMPU tests to that of previous testing methods from the literature. Optimal confidence regions for the CMP parameters are derived, which have minimum coverage probabilities of false parameters among all equal-levelled confidence sets. Moreover, simultaneous confidence regions for the CMP parameters are constructed and utilized to establish confidence intervals for the approximated mean and variance of the CMP distribution in case of overdispersion. Applications of the statistical procedures to two real data sets are included.
As a well-known and important extension of the common Poisson model with an additional parameter, Conway-Maxwell-Poisson (CMP) distributions allow for describing under- and overdispersion in discrete data. Constituting a two-parameter exponential family, CMP distributions possess useful structural and statistical properties. However, the exponential family is not steep and maximum likelihood estimation may fail even for non-trivial data sets, which is different from the Poisson case, where maximum likelihood estimation only fails if all data outcomes are zero. Conditions are examined for existence and non-existence of maximum likelihood estimates in the full family as well as in subfamilies of CMP distributions, and several figures illustrate the problem.
Uniformly most powerful unbiased tests for one-sided hypotheses about the dispersion parameter of the Conway–Maxwell–Poisson distribution are derived by utilizing the exponential family structure, and it is shown how to obtain the critical values of the relevant test statistics via simulation. The tests, which do not require parameter estimators, are applied to several real data sets from the literature to statistically confirm under- or overdispersion, giving evidence against the common Poisson model for count data.
Based on a progressively type-II censored sample from the exponential distribution with unknown location and scale parameter, confidence bands are proposed for the underlying distribution function by using confidence regions for the parameters and Kolmogorov-Smirnov type statistics. Simple explicit representations for the boundaries and for the coverage probabilities of the confidence bands are analytically derived, and the performance of the bands is compared in terms of band width and area by means of a data example. As a by-product, a novel confidence region for the location-scale parameter is obtained. Extensions of the results to related models for ordered data, such as sequential order statistics, as well as to other underlying location-scale families of distributions are discussed.
A formal definition of an exponential family is the starting point of a widely self-contained, mathematically rigorous and concise course on exponential families in a systematic structure. The definition is followed by a variety of detailed examples of one- and multiparameter, uni- and multivariate probability distributions. Basic notions and representations of exponential families, such as minimal and canonical representation, natural parametrization, and mean value parametrization are introduced and structural results are presented, which are illustrated by examples.
With a focus on statistical inference, distributional and statistical properties of multivariate and multiparameter exponential families are studied, such as generating functions, marginal and conditional distributions, and product measures. Sufficiency and completeness of statistics in exponential families is examined, and representations of the score statistic and the Fisher information matrix are shown. Many examples illustrate the results. Several divergence and distance measures are considered, and representations in terms of the cumulant function and the mean value function are given.
In step-stress experiments, test units are successively exposed to higher usually increasing levels of stress to cause earlier failures and to shorten the duration of the experiment. When parameters are associated with the stress levels, one problem is to estimate the parameter corresponding to normal operating conditions based on failure data obtained under higher stress levels. For this purpose, a link function connecting parameters and stress levels is usually assumed, the validity of which is often at the discretion of the experimenter. In a general step-stress model based on multiple samples of sequential order statistics, we provide exact statistical tests to decide whether the assumption of some link function is adequate. The null hypothesis of a proportional, linear, power or log-linear link function is considered in detail, and associated inferential results are stated. In any case, except for the linear link function, the test statistics derived are shown to have only one distribution under the null hypothesis, which simplifies the computation of (exact) critical values. Asymptotic results are addressed, and a power study is performed for testing on a log-linear link function. Some improvements of the tests in terms of power are discussed.
In a one-parameter exponential family, uniformly most powerful tests are stated for particular, but important hypotheses. Moreover, uniformly most powerful unbiased tests are shown in other cases and in multiparameter exponential families. Statistical tests for several parameters simultaneously are considered with a focus on likelihood-ratio tests, where representations and asymptotic properties are obtained. Throughout, proofs and derivations are given, and several examples illustrate the results.
Parametric families of probability distributions and their properties are extensively studied in the literature on statistical modeling and inference. Exponential families of distributions comprise density functions of a particular form, which enables general assertions and leads to nice features. With a focus on statistical inference, a motivation is given to examine exponential families in a general and unified approach. Moreover, aims and outline are addressed along with a literature overview on books in the field of Mathematical Statistics, which discuss exponential families and their properties.
Estimation of model parameters of sequential order statistics under linear and nonlinear link function assumptions is considered. Utilizing the arising curved exponential family structure, conditions for existence and uniqueness as well as the validity of asymptotic properties of maximum likelihood estimators are stated. Minimal sufficiency and completeness of the associated canonical statistics are discussed.
Within exponential families, which may consist of multi-parameter and multivariate distributions, a variety of divergence measures, such as the Kullback–Leibler divergence, the Cressie–Read divergence, the Rényi divergence, and the Hellinger metric, can be explicitly expressed in terms of the respective cumulant function and mean value function. Moreover, the same applies to related entropy and affinity measures. We compile representations scattered in the literature and present a unified approach to the derivation in exponential families. As a statistical application, we highlight their use in the construction of confidence regions in a multi-sample setup.
The usefulness of results for exponential families is demonstrated by means of three multivariate examples of different kind. As a discrete probability distribution, the family of negative multinomial distributions is considered, and results in parameter estimation and statistical testing are shown. The Dirichlet distribution serves as an example of continuous distributions, where maximum likelihood estimation, as well as the likelihood ratio test and the Wald test for a simple null hypothesis are addressed. Moreover, findings in exponential families are applied to generalized order statistics, which form a unifying approach to various models of ordered random variables.
As a specific proportional hazard rates model, sequential order statistics can be used to describe the lifetimes of load-sharing systems. Inference for these systems needs to account for small sample sizes, which are prevalent in reliability applications. By exploiting the probabilistic structure of sequential order statistics, in this article, we derive exact finite-sample inference procedures to test for the load-sharing parameters and for the nonparametrically specified baseline distribution, treating the respective other part as a nuisance quantity. This improves upon previous approaches for the model, which either assume a fully parametric specification or rely on asymptotic results. Simulations show that the tests derived are able to detect deviations from the null hypothesis at small sample sizes. Critical values for a prominent case are tabulated.
Independent samples from different exponential distributions are available in many statistical applications, for example, in queuing or reliability models, where the random variables describe waiting times and lifetimes, respectively. When two samples are observed, inference is then carried out for two location and two scale parameters. In multivariate setups including common location and common scale parameter assumptions, we provide confidence regions for the parameters of interest, which have minimum Lebesgue measure among all those based on the usual pivotal quantities and with the same or higher confidence level; in particular, they improve in terms of area (volume) upon the standard 'trapezoidal' confidence regions being constructed by combining independent univariate pivot statistics. The proposed confidence regions do not require any factorization of the overall confidence level, and their calculations need simple Monte Carlo simulations, only. Although focusing on two complete samples, generalizations of the results to more than two and doubly type-II censored samples are possible.
A multi-sample set-up of sequential order statistics from Weibull distribution functions with known scale parameters and a common unknown shape parameter is considered. The respective likelihood equation may have multiple roots even in the single-sample case, which is demonstrated by a simple example and illustrated with a simulation study. Uniqueness of the root of the likelihood equation and of the maximum likelihood estimator is examined with respect to different models of ordered data, sufficient conditions for uniqueness are shown, and the distribution of the number of roots of the likelihood equation is seen to be independent of the unknown shape parameter.
Characterizing relations via Renyi entropy of m-generalized order statistics are considered along with examples and related stochastic orderings. Previous results for common order statistics are included.