We propose a new method to address the nonparametric Behrens-Fisher problem, allowing for unequal distribution functions across the two samples. The procedure tests the null hypothesis H 0 : θ = 1 / 2 $\mathcal {H}_0: \theta = \nicefrac {1}{2}$ , where θ = P ( X < Y ) + 1 / 2 P ( X = Y ) $\theta = \text{P}(X<Y) + \nicefrac {1}{2}\text{P}(X=Y)$ denotes the Mann-Whitney effect. Apart from the trivial case of one-point distributions, no restrictions are imposed on the underlying data distribution. The test is derived by evaluating the ratio of the true variance σ N 2 $\sigma _N^2$ of the Mann-Whitney effect estimator θ ̂ N $\widehat{\theta }_N$ to its theoretical maximum, as derived from the Birnbaum-Klose inequality. Through simulations, we demonstrate that the proposed test effectively controls the type-I error rate under various conditions, including small and unbalanced sample sizes, and different data-generating mechanisms. Notably, it provides better control of the type-I error rate than the widely used Brunner-Munzel test, particularly at small significance levels such as α = 0.005 $\alpha = 0.005$ . We further construct range-preserving compatible confidence intervals and show that they exhibit improved coverage compared to the confidence intervals compatible to the Brunner-Munzel test. Finally, we illustrate the application of the method in a clinical trial example.
We propose a new test to address the nonparametric Behrens-Fisher problem involving different distribution functions in the two samples. Our procedure tests the null hypothesis ℋ_0: θ = 1/2, where θ = P(X<Y) + 1/2P(X=Y) denotes the Mann-Whitney effect. No restrictions on the underlying distributions of the data are imposed with the trivial exception of one-point distributions. The method is based on evaluating the ratio of the variance σ_N^2 of the Mann-Whitney effect estimator θ to its theoretical maximum, as derived from the Birnbaum-Klose inequality. Through simulations, we demonstrate that the proposed test effectively controls the type-I error rate under various conditions, including small sample sizes, unbalanced designs, and different data-generating mechanisms. Notably, it provides better control of the type-1 error rate compared to the widely used Brunner-Munzel test, particularly at small significance levels such as α∈{0.01, 0.005}. Additionally, we derive range-preserving compatible confidence intervals, showing that they offer improved coverage over those compatible to the Brunner-Munzel test. Finally, we illustrate the application of our method in a clinical trial example.
Many estimators of the variance of the well-known unbiased and uniform most powerful estimator of the Mann–Whitney effect, are considered in the literature. Some of these estimators are only valid in cases of no ties or are biased in small sample sizes where the amount of bias is not discussed. Here, we derive an unbiased estimator based on different rankings, the so-called ’placements’ (Orban and Wolfe in Commun Stat Theory Methods 9:883–904, 1980), which is therefore easy to compute. This estimator does not require the assumption of continuous distribution functions and is also valid in the case of ties. Moreover, it is shown that this estimator is non-negative and has a sharp upper bound, which may be considered an empirical version of the well-known Birnbaum–Klose inequality. The derivation of this estimator provides an option to compute the biases of some commonly used estimators in the literature. Simulations demonstrate that, for small sample sizes, the biases of these estimators depend on the underlying distribution functions and thus are not under control. This means that in the case of a biased estimator, simulation results for the type-I error of a test or the coverage probability of a confidence interval do not only depend on the quality of the approximation of by a normal distribution but also an additional unknown bias caused by the variance estimator. Finally, it is shown that this estimator is L_2 -consistent.
Rank methods are well-established tools for comparing two or multiple (independent) groups. Statistical planning methods for the computing the required sample size(s) to detect a specific alternative with predefined power are lacking. In the present paper, we develop numerical algorithms for sample size planning of pseudo-rank-based multiple contrast tests. We discuss the treatment effects and different ways to approximate variance parameters within the estimation scheme. We further compare pairwise with global rank methods in detail. Extensive simulation studies show that the sample size estimators are accurate. A real data example illustrates the application of the methods.
A time-to-first-event composite endpoint analysis has well-known shortcomings in evaluating a treatment effect in cardiovascular clinical trials. It does not fully describe the clinical benefit of therapy because the severity of the events, events repeated over time, and clinically relevant nonsurvival outcomes cannot be considered. The generalized pairwise comparisons (GPC) method adds flexibility in defining the primary endpoint by including any number and type of outcomes that best capture the clinical benefit of a therapy as compared with standard of care. Clinically important outcomes, including bleeding severity, number of interventions, and quality of life, can easily be integrated in a single analysis. The treatment effect in GPC can be expressed by the net treatment benefit, the success odds, or the win ratio. This review provides guidance on the use of GPC and the choice of treatment effect measures for the analysis and reporting of cardiovascular trials.
Many experiments can be modeled by a factorial design which allows statistical analysis of main factors and their interactions. A plethora of parametric inference procedures have been developed, for instance based on normality and additivity of the effects. However, often, it is not reasonable to assume a parametric model, or even normality, and effects may not be expressed well in terms of location shifts. In these situations, the use of a fully nonparametric model may be advisable. Nevertheless, until very recently, the straightforward application of nonparametric methods in complex designs has been hampered by the lack of a comprehensive R package. This gap has now been closed by the novel R-package rankFD that implements current state of the art nonparametric ranking methods for the analysis of factorial designs. In this paper, we describe its use, along with detailed interpretations of the results.
While there appears to be a general consensus in the literature on the definition of the estimand and estimator associated with the Wilcoxon-Mann-Whitney test, it seems somewhat less clear as to how best to estimate the variance. In addition to the Wilcoxon-Mann-Whitney test, we review different proposals of variance estimators consistent under both the null hypothesis and the alternative. Moreover, in case of small sample sizes, an approximation of the distribution of the test statistic based on the t-distribution, a logit transformation and a permutation approach have been proposed. Focussing as well on different estimators of the degrees of freedom as regards the t-approximation, we carried out simulations for a range of scenarios, with results indicating that the performance of different variance estimators in terms of controlling the type I error rate largely depends on the heteroskedasticity pattern and the sample size allocation ratio, not on the specific type of distributions employed. By and large, a particular t-approximation together with Perme and Manevski's variance estimator best maintains the nominal significance level
Clinical translation from bench to bedside often remains challenging even despite promising preclinical evidence. Among many drivers like biological complexity or poorly understood disease pathology, preclinical evidence often lacks desired robustness. Reasons include low sample sizes, selective reporting, publication bias, and consequently inflated effect sizes. In this context, there is growing consensus that confirmatory multicenter studies -by weeding out false positives- represent an important step in strengthening and generating preclinical evidence before moving on to clinical research. However, there is little guidance on what such a preclinical confirmatory study entails and when it should be conducted in the research trajectory. To close this gap, we organized a workshop to bring together statisticians, clinicians, preclinical scientists, and meta-researcher to discuss and develop recommendations that are solution-oriented and feasible for practitioners. Herein, we summarize and review current approaches and outline strategies that provide decision-critical guidance on when to start and subsequently how to plan a confirmatory study. We define a set of minimum criteria and strategies to strengthen validity before engaging in a confirmatory preclinical trial, including sample size considerations that take the inherent uncertainty of initial (exploratory) studies into account. Beyond this specific guidance, we highlight knowledge gaps that require further research and discuss the role of confirmatory studies in translational biomedical research. In conclusion, this workshop report highlights the need for close interaction and open and honest debate between statisticians, preclinical scientists, meta-researchers (that conduct research on research), and clinicians already at an early stage of a given preclinical research trajectory.
Rank-based inference methods are applied in various disciplines, typically when procedures relying on standard normal theory are not justifiable. Various specific rank-based methods have been developed for two and more samples and also for general factorial designs (e.g. Kruskal-Wallis test or Akritas-Arnold-Brunner test). It is the aim of the present paper (1) to demonstrate that traditional rank procedures for several samples or general factorial designs may lead to surprising results in case of unequal sample sizes as compared with equal sample sizes, (2) to explain why this is the case and (3) to provide a way to overcome these disadvantages. Theoretical investigations show that the surprising results can be explained by considering the non-centralities of the test statistics, which may be non-zero for the usual rank-based procedures in case of unequal sample sizes, while they may be equal to 0 in case of equal sample sizes. A simple solution is to consider unweighted relative effects instead of weighted relative effects. The former effects are estimated by means of the so-called pseudo-ranks, while the usual ranks naturally lead to the latter effects. A real data example illustrates the practical meaning of the theoretical discussions.
Rank-based methods are frequently used in the life sciences, and in the empirical sciences in general. Among the best-known examples of nonparametric rank-based tests are the Wilcoxon-Mann-Whitney test and the Kruskal–Wallis test. However, recently, potential pitfalls and paradoxical results pertaining to the use of traditional rank-based procedures for more than two samples have been highlighted, and the so-called pseudo-ranks have been proposed as a remedy for this type of problems. The aim of the present article is twofold: First, we show that pseudo-ranks might also behave counterintuitively when splitting up groups. Second, since the use of pseudo-ranks leads to a slightly different interpretation of the results, we provide some guidance regarding the decision for one or the other approach, in particular with respect to interpretability and generalizability of the findings. It turns out that the choice of the reference distribution, to which the individual groups are compared, is crucial. The practically relevant implications of these aspects are illustrated by a discussion of a dataset from epilepsy research. Summing up, one should decide based on thorough case-by-case considerations whether ranks or pseudo-ranks are appropriate.
The win ratio, a recently proposed measure for comparing the benefit of two treatment groups, allows ties in the data but ignores ties in the inference. In this article, we highlight some difficulties that this can lead to, and we propose to focus on the win odds instead, a modification of the win ratio which takes ties into account. We construct hypothesis tests and confidence intervals for the win odds, and we investigate their properties through simulations and in a case study. We conclude that the win odds should be preferred over the win ratio.
Many popular nonparametric inferential methods are based on ranks. Among the most commonly used and most famous tests are for example the Wilcoxon-Mann-Whitney test for two independent samples, and the Kruskal-Wallis test for multiple independent groups. However, recently, it has become clear that the use of ranks may lead to paradoxical results in case of more than two groups. Luckily, these problems can be avoided simply by using pseudo-ranks instead of ranks. These pseudo-ranks, however, suffer from being (a) at first less intuitive and not as straightforward in their interpretation, (b) computationally much more expensive to calculate. The computational cost has been prohibitive, for example, for large-scale simulative evaluations or application of resampling-based pseudorank procedures. In this paper, we provide different algorithms to calculate pseudo-ranks efficiently in order to solve problem (b) and thus render it possible to overcome the current limitations of procedures based on pseudo-ranks.
This book explains how to analyze independent data which originates from factorial designs and provides clear explanations of the modern rank-based inference methodology and numerous illustrations with real data examples as well as the necessary R/SAS code.
This chapter provides an introduction into basic statistical terminology regarding different data types, measurement scales, variables, factors, and study designs, illustrated with several examples. Good scientific practice requires research reproducibility. This includes sound statistical modeling and informed choice of appropriate statistical methods for inference. Choosing valid statistical methods requires a firm understanding of the basic terminology and concepts presented in this chapter. Readers will be able to differentiate between the different data types encountered in practice, and understand why this is important. Further, readers will gain familiarity with concepts and notation of experimental design, so that they can choose appropriate designs and valid models for many typical situations themselves, or evaluate correct use by others.
In this chapter, general results are derived on which the particular lemmas, theorems, and results stated in the previous chapters are based. We mainly intend to present here a closed theory for models involving fixed effects. Readers only interested in applications or in procedures for special designs might skip this chapter. But those readers who are interested in the background of the results and procedures will find the respective proofs and derivations in the following pages. The presentation of these derivations assumes about a year of Master's level coursework in statistics, with the respective mathematical understanding.