Prediction for very large data sets is typically carried out in two stages, variable selection and pattern recognition. Ordinarily variable selection involves seeing how well individual explanatory variables are correlated with the dependent variable. This practice neglects the possible interactions among the variables. Simulations have shown that a statistic I, that we used for variable selection is much better correlated with predictivity than significance levels. We explain this by defining theoretical predictivity and show how I is related to predictivity. We calculate the biases of the overoptimistic training estimate of predictivity and of the pessimistic out of sample estimate. Corrections for the bias lead to improved estimates of the potential predictivity using small groups of possibly interacting variables. These results support the use of I in the variable selection phase of prediction for data sets such as in GWAS (Genome wide association studies) where there are very many explanatory variables and modest sample sizes. Reference is made to another publication using I, which led to a reduction in the error rate of prediction from 30% to 8%, for a data set with, 4,918 variables and 97 subjects. This data set had been previously studied by scientists for over 10 years.
A nonparametric estimate ,B* is presented for the slope of a regression line Y = 83oX + V subject to the truncation Y c yo. This model is relevant to a cosmological controversy which concerns Hubble's Law in Astronomy. The estimate /3 * corresponds to the zero-crossing of a random function S, (/3), which for each /8 is a Mann-Whitney type of statistic designed to measure heterogeneity among the calculated residuals Y - /3X. The asymptotic distribution of P * is derived making extensive use of U-statistics to show that Sn (/3o) is asymptotically normal and then showing that S,, (/3) behaves like Sn (/3o) plus a deterministic term which is locally linear. Results on asymptotic efficiency are compared with finite sample size results by simulation.
Significance Good prediction, especially in the context of big data, is important. Common approaches to prediction include using a significance-based criterion for evaluating variables to use in models and evaluating variables and models simultaneously for prediction using cross-validation or independent test data. The first approach can lead to choosing less-predictive variables, because significance does not imply predictivity. The second approach can be improved through considering a variable’s predictivity as a parameter to be estimated. The literature currently lacks measures that do this. We suggest a measure that evaluates variables’ abilities to predict, the I -score. The I -score is effective in differentiating between noisy and predictive variables in big data and can be related to a lower bound for the correct prediction rate.
Thus far, genome-wide association studies (GWAS) have been disappointing in the inability of investigators to use the results of identified, statistically significant variants in complex diseases to make predictions useful for personalized medicine. Why are significant variables not leading to good prediction of outcomes? We point out that this problem is prevalent in simple as well as complex data, in the sciences as well as the social sciences. We offer a brief explanation and some statistical insights on why higher significance cannot automatically imply stronger predictivity and illustrate through simulations and a real breast cancer example. We also demonstrate that highly predictive variables do not necessarily appear as highly significant, thus evading the researcher using significance-based methods. We point out that what makes variables good for prediction versus significance depends on different properties of the underlying distributions. If prediction is the goal, we must lay aside significance as the only selection standard. We suggest that progress in prediction requires efforts toward a new research agenda of searching for a novel criterion to retrieve highly predictive variables rather than highly significant variables. We offer an alternative approach that was not designed for significance, the partition retention method, which was very effective predicting on a long-studied breast cancer data set, by reducing the classification error rate from 30% to 8%.
At the early age of 15, I graduated from Townsend Harris high school in New York and made the daring decision to study mathematics at the City College of New York (CCNY) during the depression, rather than some practical subject like accounting. The Mathematics faculty of CCNY was of mixed quality, but the mathematics majors were exceptionally good. Years later, one of the graduate students in statistics at Stanford found a copy of the 1939 yearbook with a picture of the Math Club. He posted it with a sign "Know your Faculty." At CCNY we had an excellent training in undergraduate mathematics, but since there was no graduate program, there was no opportunity to take courses in the advanced subjects of modern research. I was too immature to understand whether my innocent attempts to do original research were meaningful or not. This gave me an appetite for applied research where successfully confronting a real problem that was not trivial had to be useful.
A trend in all scientific disciplines, based on advances in technology, is the increasing availability of high dimensional data in which are buried important information. A current urgent challenge to statisticians is to develop effective methods of finding the useful information from the vast amounts of messy and noisy data available, most of which are noninformative. This paper presents a general computer intensive approach, based on a method pioneered by Lo and Zheng for detecting which, of many potential explanatory variables, have an influence on a dependent variable Y. This approach is suited to detect influential variables, where causal effects depend on the confluence of values of several variables. It has the advantage of avoiding a difficult direct analysis, involving possibly thousands of variables, by dealing with many randomly selected small subsets from which smaller subsets are selected, guided by a measure of influence I. The main objective is to discover the influential variables, rather than to measure their effects. Once they are detected, the problem of dealing with a much smaller group of influential variables should be vulnerable to appropriate analysis. In a sense, we are confining our attention to locating a few needles in a haystack.
The usual test that a sample comes from a distribution of given form is performed by counting the number of observations falling into specified cells and applying the $\chi^2$ test to these frequencies. In estimating the parameters for this test, one may use the maximum likelihood (or equivalent) estimate based (1) on the cell frequencies, or (2) on the original observations. This paper shows that in (2), unlike the well known result for (1), the test statistic does not have a limiting $\chi^2$-distribution, but that it is stochastically larger than would be expected under the $\chi^2$ theory. The limiting distribution is obtained and some examples are computed. These indicate that the error is not serious in the case of fitting a Poisson distribution, but may be so for the fitting of a normal.
Analysis of a subset of case-control sporadic breast cancer data, [from the National Cancer Institute's Cancer Genetic Markers of Susceptibility (CGEMS) initiative], focusing on 18 breast cancer-related genes with 304 SNPs, indicates that there are many interesting interactions that form two- and three-way networks in which BRCA1 plays a dominant and central role. The apparent interactions of BRCA1 with many other genes suggests the conjecture that BRCA1 serves as a protective gene and that some mutations in it or in related genes may prevent it from carrying out this protective function even if the patients are not carriers of known cancer-predisposing BRCA1 mutations. The method of analysis features the evaluation of the effect of a gene by averaging the effects of the SNPs covered by that gene. Marginal methods that test one gene at a time fail to show any effect. That may be related to the fact that each of these 18 genes adds very little to the risk of cancer. Analysis that relates the ratio of interactions to the maximum of the first-order effects discovers significant gene pairs and triplets.
(2007). The Effect of Fasting during Ramadan on Traffic Accidents in Turkey. CHANCE: Vol. 20, Interview with a Centennial Chart, pp. 10-18.
In his reply, Thomas asserts [1] that Linsker et al. [2] “relies more on assumptions than on data,” in coming to our conclusion that he erred in his paper [3]. In that paper he concluded that the alleged gunshot from the Grassy Knoll was contemporaneous with the assassination of President Kennedy. However, it is Thomas’ argument that relies more on assumptions than on data. The principal issues raised by Thomas’ reply [1] concern the use of the dispatcher's time annotations, and the question of whether the utterance “I’ll check it” (denoted CHECK) is a valid crosstalk. Regarding the first of these issues, Thomas continues to draw conclusions based on the dispatcher's annotations. These are too unreliable to support a meaningful inference. In fact we used them merely to show that our preferred time line was consistent with them. At no time did we base any conclusions on them. Regarding CHECK, Thomas has misunderstood or misrepresented our analysis, and wrongly claims that the results of our pattern cross-correlation (PCC) tests support his conclusion that CHECK is a crosstalk. In this rebuttal we (a) address both of these issues, (b) show, by straightforward spectrographic measurements that can be performed by any reader, that CHECK is not a crosstalk, and (c) address other issues raised by [1].
An improved understanding of cellular responses during normal anterior cruciate ligament (ACL) function or repair is essential for clinical assessments, understanding ligament biology, and the implementation of tissue engineering strategies. The present study utilized quantitative real-time RT-PCR combined with univariate and multivariate statistical analyses to establish a quantitative database of marker transcript expression that can provide a "blueprint" of ACL wound healing. Selected markers (collagen types I and III, biglycan, decorin, MMP-1, MMP-2, MMP-9, and TIMP-1) were assessed from 33 torn ACLs harvested during reconstructive surgery. Trends were observed between postinjury period and marker expressions. Significant correlations between marker expression existed and were most prominent between collagen types I and III. Canonical correlation analysis established a relationship between patient demographics and a combination of all marker expressions. The currently observed trends and correlations may assist in identifying appropriate tissue samples and provide a baseline information of marker expression level that can support in vitro optimization of environmental cues for ligament tissue engineering application.
We have revisited the acoustic evidence in the Kennedy assassination--recordings of the two Dallas police radio channels upon which our original NRC report (Ramsey NF et al., Report of the Committee on Ballistic Acoustics. National Research Council (US). Washington: National Academy Press, 1982. Posted at http://www.nap.edu/catalog/10264.html) was based--in response to the assertion by DB Thomas (Echo correlation analysis and the acoustic evidence in the Kennedy assassination revisited. Science and Justice 2001; 41: 21-32) that alleged gunshot sounds (on Channel 1), apparently recorded from a motorcycle officer's stuck-open microphone, occur at the exact time of the assassination (as established by emergency communications on Channel 2). We have critically reviewed these two publications, and have performed additional analyses. In particular we have used recorded 60 Hz hum and correlation methods to obtain accurate speed calibrations for recordings made on both channels, cepstral analysis to seek instances of repeated segments during playback of Channel 2 (which could result from groove jumping), and spectrographic and correlation methods to analyze instances of putative crosstalk used to synchronize the two channels. This paper identifies serious errors in the Thomas paper and corrects errors in the NRC report. We reaffirm the earlier conclusion of the NRC report that the alleged "shot" sounds were recorded approximately one minute after the assassination.
With very large sample sizes the conventional calculations for tests of the equality of two probabilities can lead to very small P-values. In those cases, the large deviation effects make inappropriate the asymptotic normality approximations on which those calculations are based. While reasonable interpretations of the data would tend to reject the hypothesis in those cases, it is desireable to have conservative estimates which don’t underestimate the P-value. The calculation of such estimates is presented here.
One justification offered for the use of the P-value of the Fisher Exact Test for comparing two probabilities, which involves conditioning on the margins, is that the margins carry very little if any relevant information. While the margins are conceded to be “almost” ancillary, no one has measured how much information is in the margins. We compare the total information available with the information in the margins. While the total is proportional to the sample size, the information in the margins is almost independent of the sample size, and if anything, tends to decrease as the sample size increases.
Utilizing a two-dimensional tissue culture plastic screening system and a fractional factorial design, specific media formulations and growth factor combinations were determined that support human bone marrow stromal cell (BMSC) differentiation toward fibroblast characteristics for utilization in tissue engineering, specifically cell morphology and alignment, metabolic activity, abundant expression of collagen types I and III, and negligible expression of other tissue-specific markers. BMSCs were cultured for up to 14 days on tissue culture plastic, supplemented with Dulbecco's Minimal Essential Medium (DMEM)/10% FBS or Advanced DMEM(ADMEM)/5% FBS. Each medium base was supplemented with one of nine possible growth factor combinations and ascorbate-2-phosphate (Asc-2-P) for the duration of culture. ADMEM supported comparable cell viability with half the serum content of the DMEM formulation. Asc-2-P was potent in promoting BMSC proliferation, in the absence of a mitogen, supporting significant increases in cell activity over 14 days of culture. DMEM promoted significant increases in cell viability for 7 of 9 growth factor groups when compared to their ADMEM counterparts. ADMEM, however, promoted increased cell transcript and protein expression, as 5 of 9 growth factor combinations induced a 200% increase in collagen type I versus equivalent DMEM cultures. Cell morphology and collagen type I immunostaining, when assessed in context of MTT and RNA results, identified 3 growth factor and medium combinations that supported fibroblast differentiation for future development of ligament tissue in vitro.