We provide a method for finding the optimal double sampling plan for estimating the mean value of a continuous outcome. It is assumed that the fallible and true outcome data are related by a multivariate linear regression model where only some of the explanatory variables are sampled. Conditions under which double sampling is preferred over standard sampling plans are determined. An application of the method to a well-known data set on air pollution is presented.
We test the hypothesis that the treatment of drug addicts reduces consumption of hard drugs by comparing the change in drug consumption of treated addicts before and after treatment with the change in drug consumption with demographically similar untreated addicts. If treated drug addicts are positively self-selected, naïve treatment evaluations that ignore self-selectivity into treatment will be over-optimistic on the benefits of treatment. Therefore, naïve evaluations that find "evidence" of a treatment effect are not informative. However, naïve evaluations that find no evidence of a treatment effect are informative because they are inconsistent with the hypothesis that treatment is beneficial. We refer to this as the “partial corroboration” methodology. We apply the partial corroboration methodology using a large longitudinal sample of treated and untreated addicts in Israel. Our main result is that the change in drug use frequency among the treated addicts is not significantly different from the change among their untreated counterparts. This naïve evaluation is inconsistent with the hypothesis that treatment reduces the drug consumption of addicts. We suggest that social security benefit for drug addicts should be regarded as an invalid benefit and not be made conditional on receiving treatment.
The authors consider the Bayesian analysis of multinomial data in the presence of misclassification. Misclassification of the multinomial cell entries leads to problems of identifiability which are categorized into two types. The first type, referred to as the permutation-type nonidentifiabilities, may be handled with constraints that are suggested by the structure of the problem. Problems of identifiability of the second type are addressed with informative prior information via Dirichlet distributions. Computations are carried out using a Gibbs sampling algorithm.
We use panel data on Israeli courts to estimate the “production function” for case dispositions. Our results show that the number of case dispositions is independent of the number of serving judges, and that “productivity”, as measured by completed cases per judge, varies directly with the caseload per judge. These results suggest that the productivity of judges is endogenous; for the same caseload judges complete more cases under pressure, and complete less when new judges are appointed. They also suggest that the practice of determining the number of judges by fixed “Leontieff” input–output coefficients is not appropriate.
I. Identification with Incomplete Observations, Data Mining.- Bounding Entries in Multi-way Contingency Tables Given aSet of Marginal Totals.- Identification and Estimation with Incomplete Data.- Computational Information Retrieval.- Studying Treatment Response to Inform Treatment Choice.- II. Bayesian Methods and Modelling.- Some Interactive Decision Problems Emerging in Statistical Games.- Probabilistic Modelling: An Historical andPhilosophical Digression.- A Bayesian View on Sampling the 2 x 2 Table.- Bayesian Designs for Binomial Experiments.- On the Second Order Minimax Improvement of the Sample Mean in the Estimation of a Mean Value of the Exponential Dispersion Family.- Bayesian Analysis of Cell Migration !*Linking Experimental Data and Theoretical Models.- III. Testing, Goodness of Fit and Randomness.- Sequential Bayes Detection of Trend Changes.- Box!*Cox Transformation for Semiparametric Comparison of Two Samples.- Minimax Nonparametric Goodness-of-Fit Testing.- Testing Randomness on the Basis of theNumber of Different Patterns.- The 7r* Index as a New Alternative for Assessing Goodness of Fit of Logistic Regression.- IV. Statistics of Stationary Processes.- Consistent Estimation of Early and Frequent Change Points.- Asymptotic Behaviour of Estimators of the Parameters of Nearly Unstable INAR(1) Models.- Guessing the Output of a Stationary Binary Time Series.- Asymptotic Expansions for Long-MemoryStationary Gaussian Processes.- Contributors.
This volume is a collection of papers presented at a conference held in Shoresh Holiday Resort near Jerusalem, Israel, in December 2000 organized by the Israeli Ministry of Science, Culture and Sport.
This research is concerned with the determination of the demand for “lotto” in Israel. While an important focus of our research is upon the effects on the demand for lotto of ticket pricing and jackpot announcements, we also investigate several empirical phenomena that are apparently inconsistent with expected utility theory. These include an effect we call “lottomania” which is induced by rollover, and “prize fatigue” when the jackpot does not increase. Another aberration from expected utility theory is that the underlying odds of winning have no measurable effect on sales.
The payout rate on lotto is normally fixed. We show that such a policy is generally suboptimal from the lotto authorities' point of view. The payout rate should be allowed to vary according to the number of rollovers that have occurred. To illustrate our argument, we simulate and optimize an econometric model of the lotto market in Israel. We also consider whether it is profitable to increase the frequency of lotto from once to twice a week.
We provide a method for finding the optimal double sampling plan for estimating the mean value of a continuous outcome. It is assumed that the fallible and true outcome data are related by a simple linear regression model. The design parameters are the total sample size, N, and the number of doubly sampled units n. We show that under certain conditions the efficiency gains, relative to standard sampling plans are considerable.
Using data for 1974–1990 we estimate an econometric model of the Israeli housing market. Novel features of the model include the treatment of the starts–completions nexus, the implications of capital market imperfections for both building contractors and households, and the effects of public sector housing construction on the dynamics of the housing market. We find that the price elasticity of demand for housing is small, as is the supply elasticity. This implies that shocks to the housing market continue to affect housing prices and new building for many years.
A methodology is proposed for estimating `status quo effects' (resistance to change) and `loss aversion' (willingness to pay is less than willingness to accept) from consumer survey data. Conjoint data on Israeli households' attitudes to electricity outages reveal that both of these effects are pronounced and that they depend on the interviewees' characteristics. Ordered conditional logit estimation indicates that while the estimates are rank sensitive the suggested methodology continues to be applicable. Finally, the conjoint estimates are compared with those obtained by applying the method of contingent valuation.
Cross-section data on investment in back-up generators and uninterruptable power supplies (UPS) are used to infer the implied cost of electricity outages in the business and public sectors in Israel. Two-limit tobit models of the demand for back-up are estimated and used to simulate the mitigated and unmitigated cost of power outages. These "revealed preference estimates of outage costs are then compared with estimates based on the method of subjective evaluation.
A conditional resampling methodology is proposed for improving the estimators of multinomial classification probabilities in the presence of a fallible classifier. The (unbiased) estimator requires knowledge of the misclassification error rates, which are obtained by resampling from the initial sample and applying an infallible classifier. Tenenbein's “double sampling” methodology resamples randomly from the initial sample. We show that substantial gains in efficiency are obtainable if, instead of simple random resampling, different resampling rates are allowed for the fallibly determined classes in the initial sample. In most situations, this resampling methodology is neither more difficult nor more expensive than the double sampling methodology. A very promising area of application of the proposed methodology is sampling inspection in quality control, where products can either be classified quickly by inspectors or by a more thorough, perhaps destructive, inspection.
SUMMARY Consider a binary random variable having outcomes subject to misclassification errors where the probabilities of misclassification errors depend on outcome. The asymptotic relative efficiencies are derived for comparing two binomial populations, comparing stratified binomial populations and for paired comparisons. The probabilities of misclassification can vary with strata or be different for each pair. The criterion for optimum experimental designs is shown to be the ratio of the asymptotic relative efficiency to expected cost per observation. This criterion implies that mixtures of different experiments can never be optimal. A number of experimental strategies is discussed.
Journal Article Locally optimal design for comparing two probabilities from binomial data subject to misclassification Get access HERMAN CHERNOFF, HERMAN CHERNOFF Department of Statistics, Harvard UniversityCambridge, Massachusetts 02138, U.S.A. Search for other works by this author on: Oxford Academic Google Scholar YOEL HAITOVSKY YOEL HAITOVSKY Departments of Economics and Statistics, Hebrew UniversityJerusalem 91905, Israel Search for other works by this author on: Oxford Academic Google Scholar Biometrika, Volume 77, Issue 4, December 1990, Pages 797–805, https://doi.org/10.1093/biomet/77.4.797 Published: 01 December 1990 Article history Received: 01 April 1989 Revision received: 01 April 1990 Published: 01 December 1990
A multivariate linear regression model with q responses as a linear function of p independent variables is considered with a p × q parameter matrix B. The least-squares or normal-theory maximum likelihood estimate of B is deficient in that it takes no account of the ‘across regression’ correlations, and ignores the Stein effect. A remedy was offered by Brown & Zidek (1980) in the form of a multivariate ridge estimator. A richer class of estimators is obtained here by casting the model in a linear hierarchical framework, obtaining the Brown & Zidek multivariate ridge estimates, Efron & Morris's estimates of several normal mean vectors and Fearn's Bayesian estimates of growth curves as special cases. The unknown covariance case results in an identifiability problem, which can be overcome by a Bayesian approach using conjugate priors for the unidentified covariance matrices.
In recent years the Israeli government has somewhat relaxed the previously severe restrictions on charter flights to Israel. This study suggests that liberalisation will benefit the balance of payments and the national economy, but that some restriction is desirable.
This article is concerned with hierarchical prior distributions and the effect of replacing the distribution of a component in the hierarchy with a diffuse distribution where all nondiffuse distributions are multivariate normal. Let f denote the posterior density function and g = gm, the approximation to f obtained by truncating the hierarchy at stage m. The Kullback-Leibler information index, I(f, g) = ∫ f log(fg), will be used to measure the accuracy of g to avoid declaring specific objectives such as estimation or prediction. It is intuitively plausible that g will be increasingly more accurate as m increases; we show by theorems and two examples that this is sometimes but not always true. In the second example the behavior of I(f, g) depends on both the data and f's parameters and it may reach a minimum at an intermediate value of m which would then be optimal.
The ridge estimator of the usual linear model is generalized by the introduction of an a priori vector r and an associated positive semidefinite matrix S. It is then shown that the generalized ridge estimator can be justified in two ways: (a) by the minimization of the residual sum of squares subject to a constraint on the length, in the metric S, of the vector of differences between r and the estimated linear model coefficients, (b) by incorporating prior knowledge, r playing the role of the vector of means and S proportional to the precision matrix. Both a Bayesian and an Aitken generalized least squares frameworks are used for the latter. The properties of the new estimator are derived and compared to the ordinary least squares estimator. The new method is illustrated with different assumptions on the form of the S matrix.