Sample quantiles, such as the median, are often better suited than the sample mean for summarising location characteristics of a data set. Similarly, linear combinations of sample quantiles and ratios of such linear combinations, e.g. the interquartile range and quantile-based skewness measures, are often used to quantify characteristics such as spread and skew. While often reported, it is uncommon to accompany quantile estimates with confidence intervals or standard errors. The rquest package provides a simple way to conduct hypothesis tests and derive confidence intervals for quantiles, linear combinations of quantiles, ratios of dependent linear combinations (e.g., Bowley's measure of skewness) and differences and ratios of all of the above for comparisons between independent samples. Many commonly used measures based on quantiles are included, although it is also very simple for users to define their own. Additionally, quantile-based measures of inequality are also considered. The methods are based on recent research showing that reliable distribution-free confidence intervals can be obtained, even for moderate sample sizes. Several examples are provided herein.
Motivation Meta-analysis methods widely used for combining metabolomics data do not account for correlation between metabolites or missing values. Within- and between-study variability are also often overlooked. These can give results with inferior statistical properties, leading to misidentification of biomarkers.Results We propose a multivariate meta-analysis model for high-dimensional metabolomics data (MetaHD), which accommodates the correlation between metabolites, within- and between-study variances, and missing values. MetaHD can be used for integrating and collectively analysing individual-level metabolomics data generated from multiple studies as well as for combining summary estimates. We show that MetaHD leads to lower root mean square error compared to the existing approaches. Furthermore, we demonstrate that MetaHD, which exploits the borrowing strength between metabolites, could be particularly useful in the presence of missing data compared with univariate meta-analysis methods, which can return biased estimates in the presence of data missing at random.Availability and implementation The MetaHD R package can be downloaded through Comprehensive R Archive Network (CRAN) repository. A detailed vignette with example datasets and code to prepare data and analyses are available on https://bookdown.org/a2delivera/MetaHD/.
Estimation of the four generalized lambda distribution parameters is not straightforward, and available estimators that perform best have large computation times. In this paper, we introduce a simple two-step estimator of the parameters that is comparatively very quick to compute and performs well when compared with other methods. This computational efficiency makes the use of bootstrapping to obtain interval estimators for the parameters possible. Simulations are used to assess the performance of the new estimators and applications to several data sets are included.
We examine a commonly used relative poverty measure called the headcount ratio ( $$H_{p}$$ ), defined to be the proportion of incomes falling below the relative poverty line, which is defined to be a fraction p of the median income. We do this by considering this concept for theoretical income populations, and its potential for determining actual changes following transfer of incomes from the wealthy to those whose incomes fall below the relative poverty line. In the process we derive and evaluate the performance of large sample confidence intervals for $$H_{p}$$ . Finally, we illustrate the estimators on real income data sets.
Chi-squared tests for lack of fit are traditionally employed to find evidence against a hypothesized model, with the model accepted if the Karl Pearson statistic comparing observed and expected numbers of observations falling within cells is not significantly large. However, if one really wants evidence for goodness of fit, it is better to adopt an equivalence testing approach in which small values of the chi-squared statistic are evidence for the desired model. This method requires one to define what is meant by equivalence to the desired model, and guidelines are proposed. Then a simple extension of the classical normalizing transformation for the non-central chi-squared distribution places these values on a simple to interpret calibration scale for evidence. It is shown that the evidence can distinguish between normal and nearby models, as well between the Poisson and over-dispersed models. Applications to evaluation of random number generators and to uniformity of the digits of pi are included. Sample sizes required to obtain a desired expected evidence for goodness of fit are also provided.
The coefficient of variation (CV) is commonly used to measure relative dispersion. However, since it is based on the sample mean and standard deviation, outliers can adversely affect it. Additionally, for skewed distributions the mean and standard deviation may be difficult to interpret and, consequently, that may also be the case for the CV . Here we investigate the extent to which quantile-based measures of relative dispersion can provide appropriate summary information as an alternative to the CV. In particular, we investigate two measures, the first being the interquartile range (in lieu of the standard deviation), divided by the median (in lieu of the mean), and the second being the median absolute deviation, divided by the median, as robust estimators of relative dispersion. In addition to comparing the influence functions of the competing estimators and their asymptotic biases and variances, we compare interval estimators using simulation studies to assess coverage.
The quantile ratio index is a simple and effective measure of relative inequality for income data that is resistant to outliers. A useful property of this index is investigated here: given a partition of the income distribution into a union of sets of symmetric quantiles, one can find the inequality for each set and readily combine them in a weighted average to obtain the index for the entire population. When applied to data for various years, one can track how these contributions to inequality vary over time, as illustrated here for Australian Bureau of Statistics income and wealth data.
The four-parameter Generalized Lambda distribution (GLD) can be used to approximate many probability distributions. We present a simple and efficient two-stage process for finding optimal GLD parameters to approximate a specified distribution. The probability density quantile function is first used to find the best GLD shape parameters. Given those shape parameters, it is then straightforward to find the best location and scale parameters. We highlight the excellent performance of our approach with comparisons to two existing and popular methods for a wide choice of distributions. Finally, we show that this is method can be used with other distributions by providing applications also to the Generalized Beta distribution.
We demonstrate that questions of convergence and divergence regarding shapes of distributions can be carried out in a location- and scale-free environment. This environment is the class of probability density quantiles (pdQs), obtained by normalizing the composition of the density with the associated quantile function. It has earlier been shown that the pdQ is representative of a location-scale family and carries essential information regarding shape and tail behavior of the family. The class of pdQs are densities of continuous distributions with common domain, the unit interval, facilitating metric and semi-metric comparisons. The Kullback–Leibler divergences from uniformity of these pdQs are mapped to illustrate their relative positions with respect to uniformity. To gain more insight into the information that is conserved under the pdQ mapping, we repeatedly apply the pdQ mapping and find that further applications of it are quite generally entropy increasing so convergence to the uniform distribution is investigated. New fixed point theorems are established with elementary probabilistic arguments and illustrated by examples.
Ratios of quantiles are often computed for income distributions as rough measures of inequality, and inference for such ratios has recently become available. The special case when the quantiles are symmetrically chosen; that is, when the p/2 quantile is divided by the (1 − p/2) quantile, is of special interest because the graph of such ratios, plotted as a function of p over the unit interval, yields an informative inequality curve. The area above the curve and less than the horizontal line at one is an easily interpretable measure of inequality. The advantages of these concepts over the traditional Lorenz curve and Gini coefficient are numerous: they are defined for all positive income distributions, they can be robustly estimated and large sample confidence intervals for the inequality coefficient are easily found. Moreover, the inequality curves satisfy a median-based transference principle and are convex for many commonly assumed income distributions.
Many measures of peakedness, heavy-tailedness and kurtosis have been proposed in the literature, mainly because kurtosis, as originally defined, is a complex combination of the other two concepts. Insight into all three concepts can be gained by studying Ruppert's ratios of interquantile ranges. They are not only monotone in Horn's measure of peakedness when applied to the central portion of the population, but also monotone in the practical tail-index of Morgenthaler and Tukey, when applied to the tails. Distribution-free confidence intervals are found for Ruppert's ratios, and sample sizes required to obtain such intervals for a pre-specified relative width and level are provided. In addition, the empirical power of distribution-free tests for peakedness and bimodality are found for symmetric beta families and mixtures of $t$ distributions. An R script that computes the confidence intervals is provided in online supplementary material.
For every discrete or continuous location-scale family having a square-integrable density, there is a unique continuous probability distribution on the unit interval that is determined by the density-quantile composition introduced by Parzen in 1979. These probability density quantiles (pdQs) only differ in shape, and can be usefully compared with the Hellinger distance or Kullback-Leibler divergences. Convergent empirical estimates of these pdQs are provided, which leads to a robust global fitting procedure of shape families to data. Asymmetry can be measured in terms of distance or divergence of pdQs from the symmetric class. Further, a precise classification of shapes by tail behavior can be defined simply in terms of pdQ boundary derivatives.
Ratios of sample percentiles or of quantiles based on a single sample are often published for skewed income data to illustrate aspects of income inequality, but distribution-free confidence intervals for such ratios are not available in the literature. Here we derive and compare two large-sample methods for obtaining such intervals. They both require good distribution-free estimates of the quantile density at the quantiles of interest, and such estimates have recently become available. Simulation studies for various sample sizes are carried out for Pareto, lognormal and exponential distributions, as well as fitted generalized lambda distributions, to determine the coverage probabilities and widths of the intervals. Robustness of the estimators to contamination or a positive proportion of zero incomes is examined via influence functions and simulations. The motivating example is Australian household income data where ratios of quantiles measure inequality, but of course these results apply equally to data from other countries.
A standard approach to confidence intervals for quantiles requires good estimates of the quantile density. The optimal bandwidth for kernel estimation of the quantile density depends on an underlying location‐scale family only through the quantile optimality ratio (QOR), which is the starting point for our results. While the QOR is not distribution‐free, it turns out that what is optimal for one family often works quite well for families having similar shape. This allows one to rely on a single representative QOR if one has a rough idea of the distributional shape. Another option that we explore assumes the data can be modelled by the highly flexible generalized lambda distribution (GLD), already studied by others, and we show that using the QOR for the estimated GLD can lead to more than competitive intervals. Effective confidence intervals for the difference between quantiles from independent populations is a byproduct. Copyright © 2016 John Wiley & Sons, Ltd.
The classical Lorenz curve is often used to depict inequality in a population of incomes, and the associated Gini coefficient is relied upon to make comparisons between different countries and other groups. The sample estimates of these moment-based concepts are sensitive to outliers and so we investigate the extent to which quantile-based definitions can capture income inequality and lead to more robust procedures. Distribution-free estimates of the corresponding coefficients of inequality are obtained, as well as sample sizes required to estimate them to a given accuracy. Convexity, transference and robustness of the measures are examined and illustrated.
Some equivalence tests are based on two one-sided tests, where in many applications the test statistics are approximately normal. We define and find evidence for equivalence in Z-tests and then one- and two-sample binomial tests as well as for t-tests. Multivariate equivalence tests are typically based on statistics with non-central chi-squared or non-central F distributions in which the non-centrality parameter λ is a measure of heterogeneity of several groups. Classical tests of the null λ ≥ λ 0 versus the equivalence alternative λ < λ 0 are available, but simple formulae for power functions are not. In these tests, the equivalence limit λ 0 is typically chosen by context. We provide extensions of classical variance stabilizing transformations for the non-central chi-squared and F distributions that are easy to implement and which lead to indicators of evidence for equivalence. Approximate power functions are also obtained via simple expressions for the expected evidence in these equivalence tests.
When conducting a meta-analysis of standardized mean differences (SMDs), it is common to assume equal variances in the two arms of each study. This leads to Cohen's d estimates for which interpretation is simple. However, this simplicity should not be used as a justification for the assumption of equal variances in situations where evidence may suggest that it is incorrect. Until now, researchers have either used an F-test for each individual study as a justification for the equality of variances or perhaps even conveniently ignored such tools altogether. In this paper we propose using a meta-analysis of F-test statistics to estimate the ratio of variances prior to the combination of SMD's. This procedure allows some studies to be included that might otherwise be omitted by individual fixed level tests for unequal variances, sometimes occur even when the assumption of equal variances holds. The estimated ratio of variances, as well as associated confidence intervals, can be used as guidance as to whether the assumption of equal variances is violated. The estimators considered include variance stabilization transformations (VST) of the F-test statistics as well as MLE estimators. The VST approaches enable the use of QQ-plots to visually inspect for violations of equal variances while the MLE estimator easily allows for the introduction of a random effect. When there is evidence of unequal variances, this work provides a means to formally justify the use of less common methods such as log ratio of means when studies are measured on a different scale.