Multiple Hypotheses, Simes' Test of† Y. Hochberg, Y. HochbergSearch for more papers by this authorG. Hommel, G. HommelSearch for more papers by this author Y. Hochberg, Y. HochbergSearch for more papers by this authorG. Hommel, G. HommelSearch for more papers by this author First published: 29 September 2014 https://doi.org/10.1002/9781118445112.stat01510 †This article was originally published online in 2006 in Encyclopedia of Statistical Sciences, © John Wiley & Sons, Inc. and republished in Wiley StatsRef: Statistics Reference Online, 2014. Read the full textAboutPDF ToolsRequest permissionExport citationAdd to favoritesTrack citation ShareShare Give accessShare full text accessShare full-text accessPlease review our Terms and Conditions of Use and check box below to share full-text version of article.I have read and accept the Wiley Online Library Terms and Conditions of UseShareable LinkUse the link below to share a full-text version of this article with your friends and colleagues. Learn more.Copy URL Share a linkShare onFacebookTwitterLinked InRedditWechat Wiley StatsRef: Statistics Reference OnlineBrowse other articles of this reference work:BROWSE BY TOPICBROWSE A-Z RelatedInformation
We discuss Bayesian attitudes towards adjusting inferences for multiplicities. In the simplest Bayesian view, there is no need for adjustments and the Bayesian perspective is similar to that of the frequentist who makes inferences on a per-comparison basis. However, as we explain, Bayesian thinking can lead to making adjustments that are in the same spirit as those made by frequentists who subscribe to preserving the familywise error rate. We describe the differences between assuming independent prior distributions and hierarchical prior distributions. As an example of the latter, we illustrate the use of a Dirichlet process prior distribution in the context of multiplicities. We also discuss some quasi-Bayesian procedures which combine Bayesian and frequentist ideas. This shows the potential of Bayesian methodology to yield procedures that can be evaluated using “objective” criteria. Finally, we comment on the role of subjectivity in Bayesian approaches to the complex realm of multiple comparisons problems, and on robust vs. informative priors.
The problem of identifying the lowest dose level for which the mean response differs from that at the zero dose level is considered. A general framework for stepwise testing procedures that use contrasts among the dose level means is proposed. Using this framework, several new procedures are derived. These and some existing procedures, including that of Williams (1971, Biometrics 27, 103-117; 1972, Biometrics 28, 519-531), are compared analytically and by an extensive simulation study for the normal theory balanced one-way layout case. It is pointed out that the procedures based on the so-called step and basin contrasts proposed by Ruberg (1989, Journal of American Statistical Association 84, 816-822) have excessively high type I familywise error rates (FWEs) and, hence, they should not be used. Some findings of the simulation study are as follows: For monotone dose mean configurations, Williams' procedure and two step-down test procedures based on Helmert and linear contrasts offer the best performance. For nonmonotone dose mean configurations, the performance of Williams' procedure does degrade somewhat, but the other two procedures are still the best. For more complex designs, a simple step-down test procedure that uses any alpha-level tests (not necessarily t-tests) to compare each dose level with the zero dose level controls the FWE and is the only alternative available, but its power is rather low, especially under nonmonotone configurations. Step-up procedures are generally dominated by step-down procedures when the same contrasts are used although the differences are not great.
Shaffer (J. Amer. Statist. Assoc. 81 (1986) 826–831) gave simple and more powerful modifications of Holm's (Scand. J. Statist. 6 (1979) 65–70) procedure to families of Logically Related Hypotheses (LRH). Since then, several procedures more powerful than Holm's (based on Simes (Biometrika 73 (1986) 751–754)) have been introduced, with partial or no discussion of their improvements in cases of LRH. In this article, simple and more powerful modifications of these procedures (analogous to Shaffer's modifications) for general problems of LRH are given. These modified procedures are more powerful than the modified Holm's procedure in general problems of LRH. The various modifications are demonstrated and their superiority over existing procedures is indicated for some LRH. Another extension concerns one-sided tests with normal correlated test statistics. Some analysis of the type-I error-rate associated with Simes' type procedures in this case indicates satisfactory control. A class of problems suitable for applications of a Simes' type procedure in view of the given results, is indicated.
SUMMARY The common approach to the multiplicity problem calls for controlling the familywise error rate (FWER). This approach, though, has faults, and we point out a few. A different approach to problems of multiple significance testing is presented. It calls for controlling the expected proportion of falsely rejected hypotheses -the false discovery rate. This error rate is equivalent to the FWER when all hypotheses are true but is smaller otherwise. Therefore, in problems where the control of the false discovery rate rather than that of the FWER is desired, there is potential for a gain in power. A simple sequential Bonferronitype procedure is proved to control the false discovery rate for independent test statistics, and a simulation study shows that the gain in power is substantial. The use of the new procedure and the appropriateness of the criterion are illustrated with examples.
On the basis of a retrospective study of 71 children followed for 24 months after diagnosis of type I insulin dependent diabetes a fitted mathematical model was constructed for the prediction of the course of beta cell function from the time of diagnosis. Two equations were derived, one for the maximal basal (B-max) and the other for the maximal i.v. glucagon stimulated peak C-peptide (P-max) levels reached during the remission period. The prognostic variables selected for analysis were: peak C-peptide levels at diagnosis (Po), age, sex, degree of obesity, pubertal rating, the presence of islet cell antibodies (ICA) and levels of GHb. Multivariate analysis of the data showed that Po (p = 0.0006), puberty (p = 0.041), obesity (p = 0.0021), sex (p = 0.031), ICA (p = 0.0045) and GHb (p = 0.0066) significantly contributed to the prediction formula obtained for B-max whereas the contribution of the above variables for P-max were: Po (p = 0.0019), puberty (p = 0.0187), obesity (p = 0.0058), sex (p = 0.0598), ICA (p = 0.0187) and GHb (p = 0.0027). The residuals of the observed values from the values fitted by the predicted equations served to define two separate groups demonstrating distinct differences in the natural course of beta cell function in type I diabetes. This fitted model may thus be useful in distinguishing between newly diagnosed young patients who will undergo remission, requiring lower insulin doses, and those who have little chance for remission. It might also be helpful in the selection of patients most likely to benefit from immunosuppression or modulation, to maximize the benefit to risk ratio for such patients.
The problem of multiple comparisons is discussed in the context of medical research. The need for more powerful procedures than classical multiple comparison procedures is indicated. To this end some new, general and simple procedures are discussed and demonstrated by two examples from the medical literature: the neuropsychologic effects of unidentified childhood exposure to lead, and the sleep patterns of sober chronic alcoholics.
The therapeutic effect of bopindolol (a new long-acting β-adrenoreceptor antagonist with mild intrinsic sympathomimetic activity) in the treatment of angina pectoris was compared with that of nifedipine in 16 patients (all male; average age 57 years, range 51–66), all of whom were experiencing at least three attacks of typical anginal pain daily, following a previous myocardial infarction. In addition, all showed at least a 1 -mm ST segment depression during a standard exercise test (Bruce protocol). A single-blind crossover design was used with two 2-week placebo periods and two 3-week active treatment periods. The patients were randomized to treatment with either 2 mg bopindolol once a day or 20 mg nifedipine three times a day. Judged by the number of anginal attacks daily, the therapeutic effect of bopindolol was greater than that of nifedipine, but both were clearly more effective than the initial placebo. The number of anginal attacks during the intermediate placebo phase, although greater than during either active treatment phase, was still markedly lower than during the initial placebo phase. The increase in the heart rate and blood pressure seen during the exercise tests was reduced more by bopindolol than by nifedipine. Adverse effects were more common during nifedipine treatment than during bopindolol treatment. Six patients reported adverse effects while taking nifedipine (palpitations, headache, oedema), three being unable to continue therapy and one being admitted to hospital with hypotension. During bopindolol treatment one patient reported weakness and headaches. The study shows that bopindolol is at least as effective as nifedipine in the treatment of angina pectoris and that the incidence of adverse effects is lower.
PROCEDURES BASED ON CLASSICAL APPROACHES FOR FIXED--EFFECTS LINEAR MODELS WITH NORMAL HOMOSCEDASTIC INDEPENDENT ERRORS. Some Theory of Multiple Comparisons Procedure Fixed--effects Linear Models. Single--step Procedures for Pairwise and More General Comparisons Among All Treatments. Stepwise Procedures for Pairwise and More General Comparisons Among All Treatments. Procedures for Some Other Nonhierarchical Finite Families of Comparisons. Designing Experiments for Multiple Comparisons. PROCEDURES FOR OTHER MODELS AND PROBLEMS, AND PROCEDURES BASED ON ALTERNATIVE APPROACHES. Procedures for One--way Layouts with Unequal Variances. Procedures for Some Mixed--effects Models. Distribution--free and Robust Procedures. Some Miscellaneous Multiple Comparison Problems. Optimal Procedures Using Decision--theoretic, Bayesian, and Other Approaches. Appendixes. Tables. References. Index.
AbstractThe lichen Caloplaca aurantia, when growing on concrete roof tiles, accumulates high concentrations of heavy metals. The uptake correlates with the ambient metal concentration. The results of lichen and substratum analyses of samples from a heavily polluted area indicate that the concentrations of metals in the lichen are many times higher than in the substratum. The relative content of the eight metals, Fe, Mn, Zn, Cu, Pb, Cr, Ni and Cd was determined. The content of Fe, Cr, Ni and Cd was not significantly different in the various tile zones sampled. Fe, Mn, Zn, Cu, Cr, Ni and Cd were mainly taken up by the lichen directly from the atmosphere. The uptake of these metals from the tile is almost negligible. The maximum uptake of Pb by C. aurantia from the tile glaze was 10.95°° of the total amount in the lichen. The presence of the lichen on the tiles prevents the penetration of Pb, Zn and Mn into the substratum.
The problem of deciding the signs of $k$ parameters $(\theta_1, \cdots, \theta_k) \equiv \mathbf{theta}$ based on $(\hat{\theta}_1, \cdots, \hat{\theta}_k) \sim N(\mathbf{\theta,\Sigma})$ such that $p_\mathbf{\theta} \{$\text{any error$\} \leq \alpha \forall \mathbf{\theta}$ is discussed by Bohrer and Schervish (1980). They characterize a desirable class of procedures called locally optimal. For the case $k = 2, \mathbf{\Sigma = I}$, and $\alpha \leq \frac{1}{3}$, they present a particular rule from this class called the double cross. In this paper, we address the problem of selecting a best rule from among all locally optimal rules when $k = 2$ and $\mathbf{\Sigma = I}$. When $\alpha \leq \frac{1}{3}$, the double cross is shown to be an attractive choice. Other rules are obtained for higher values of $\alpha$. We also examine a more general optimization criterion than the one used by Bohrer and Schervish and obtain different optimal rules for several classes of problems. The optimal rule corresponding to one of these classes has no two-decision region. A modification of the formulation is offered under which a well-known rule (with two decision regions) emerges as the unique optimal procedure.
The subject of this paper is the admissibility of a popular procedure (which will be labeled ‘SB’ after Spjøtvoll and Bohrer) for classifying signs of k ≥ 2 unknowwn scalar parameters based on normally distributed estimators with known covariance matrix. First, a wide class of procedures, all of which involve 2k separate α-level tests of each of the possible sign configurations of the k parameters and all control the probability of any misclassification at level α under any parameter vector, is introduced. Under a particular criterion it is shown that the SB rule is uniformally dominated by some members of that class. On the other hand, it is shown that a generalized SB rule is the only admissible rule in another class of simpler procedures.
Consider two independent random variables X and Y. The functional R = Pr(X less than Y) [or gamma = Pr(X less than Y) - Pr(Y less than X)] is of practical importance in many situations, including clinical trials, genetics, and reliability. In this paper several approaches to estimation of gamma when X and Y are presented in discretized (categorical) form are analyzed and compared. Asymptotic formulas for the variances of the estimators are derived; use of the bootstrap to estimate variances is also discussed. Computer simulations indicate that the choice of the best estimator depends on the value of gamma, the underlying distribution, and the sparseness of the data. It is shown that the bootstrap provides a robust estimate of variance. Several examples are treated.
Procedures for multiple comparisons among treatment means, in analysis of covariance with a random concomitant, are considered. A Tukey-Kramer (TK)-type procedure is introduced for a conditional analysis and is compared with Thigpen and Paulson's (1974) unconditional procedure. It is first established (by simulation) that the TK procedure controls the unconditional familywise error rate at the nominal level set for the conditional procedure. In the second part of this work, the confidence interval lengths obtained by the two procedures are compared. It is found that the expected ratio of the confidence interval length obtained by the conditional method to that obtained by the unconditional method is generally less than 1. Moreover, the probability that the length of a TK interval will be larger than that of the Thigpen and Paulson procedure is (roughly) between .2 and .3. Key Words: Random concomitantConditional and unconditional proceduresTukey-Kramer
Serum from young normal BALB/c mice was found to contain IgM antibodies able to mediate complement-dependent lysis of certain syngeneic or allogeneic tumor target cells. The titer of such naturally occurring antitumor antibodies ( NATA ) was found to increase with aging. A longitudinal serological study comparing the cytotoxicity potential of NATA from normal and from urethan-treated BALB/c mice was performed. It was found that urethan-treated mice that did not develop primary lung-adenomas within the duration of the experiment had significantly lower NATA titers, against one out of 4 target cells assayed, than urethan-treated animals that developed lung adenomas. This difference was evident in two independent experiments. The results suggested that the lower NATA activity of the urethan-treated mice that did not develop tumors existed even before exposure to the carcinogenic insult. This raises the possibility that certain populations could be segregated according to their natural antibody profile into those individuals which will develop primary tumors within a certain period if exposed to a subthreshold amount of carcinogen, and those which will not.
Two new estimators for calibrating unknowns from dose-response curves, in a system of quality-controlled assays, are examined. In contrast with the conventional estimator which uses only the results of the one assay in which the response of the unknown dose is measured, the new estimators also utilize the results of all other assays through the replications of the control samples in the system. The first estimator is based on maximizing the likelihood of the given system (with respect to the different dose-response parameters, the levels of the control samples and the levels of the unknowns) when response errors are normally distributed. The second estimator is a regression-like estimator obtained by subtracting from the conventional estimator its estimated regression on the deviation of the calibrated control levels in the given assay from their average values in the system. Evaluations of the reductions in bias and variance attained by the new estimators show when substantial reductions in mean square error can be expected. The new estimators are illustrated with a system of 22 hFSH radioimmunoassays.