ABSTRACT In statistics, samples are drawn from a population in a data‐generating process (DGP). Standard errors measure the uncertainty in estimates of population parameters. In science, evidence is generated to test hypotheses in an evidence‐generating process (EGP). We claim that EGP variation across researchers adds uncertainty—nonstandard errors (NSEs). We study NSEs by letting 164 teams test the same hypotheses on the same data. NSEs turn out to be sizable, but smaller for more reproducible or higher rated research. Adding peer‐review stages reduces NSEs. We further find that this type of uncertainty is underestimated by participants.
This paper studies the trading behavior of different types of traders in commodity futures and their impact on liquidity consumption/provision as well as price discovery in the market. CME classifies each trade by its Customer Type Indicator (CTI) into four groups: a local trader who trades for his own account (CTI1), a commercial clearing member for his proprietary accounts (CTI2), an exchange member for his own account though a local trader (CTI3), and the general public (non-members) (CTI4). We find that non-members (CTI4) consume most of the short-term (intraday) liquidity while local traders as market makers are its main provider. Such a liquidity provision yields a substantial Sharpe ratio for the latter and constitutes most of the intraday volume. Most of the interday trading and position taking come from groups CTI2 and CTI3, reflecting their longer term needs for hedging and speculation. We also find that the imbalance in demand and supply in the market can explain a significant part of the daily price movements. In addition, changes in the overnight positions of the general public and clearing members contribute mostly to daily price changes. Moreover, we find that daily changes in the positions of CTI3 group can forecast future price movements, reflecting possible information advantage they may possess.
This paper proposes a two-step method for an omnibus misspecification test for constant parameters in the volatility equation of stochastic volatility models. The proposed test has a well-known null asymptotic distribution free of nuisance parameters. It is easy to implement and has low computational cost. Monte Carlo simulations support the relevance of the proposed method, evaluate the performance of the procedure, and highlight its small computational load. An empirical application shows the relevance of the procedure.
This paper proposes a two-step method for an omnibus misspecification test for constant parameters in nonlinear models. The procedure is easy to implement and has a low computational cost. The asymptotic distribution and the consistency of the procedure are derived. Monte Carlo simulations support the relevance of the proposed method, evaluate the performance of the procedure, and highlight its small computational load. An empirical application illustrates the relevance of the procedure.
Este trabajo estudia las relaciones dinamicas entre el consumo de electricidad y el Producto Interior Bruto en Espana. Utilizando la funcion de correlacion cruzada se constata que el PIB causa al consumo de electricidad, con un efecto instantaneo (elasticidad) de 0.95 y un efecto a largo plazo de 0.42. Tecnicas de remuestreo y contrastes de estabilidad endogenos permiten aceptar la significacion de la relacion estimada y no rechazan la hipotesis de estabilidad en el periodo considerado. This paper studies the dynamic relationships between energy consumption and GDP in Spain. The cross correlation function shows that the GDP causes the electricity consumption. The impact multiplier is 0.95 and the long run effect is 0.42. Bootstrap techniques and structural change test asses the significativeness and the stability along the considered sample of the estimated model.
We examine the dynamic relation between return and volume of individual stocks. Using a simple model in which investors trade to share risk or speculate on private information, we show that returns generated by risk-sharing trades tend to reverse themselves while returns generated by speculative trades tend to continue themselves. We test this theoretical prediction by analyzing the relation between daily volume and first-order return autocorrelation for individual stocks listed on the NYSE and AMEX. We find that the cross-sectional variation in the relation between volume and return autocorrelation is related to the extent of informed trading in a manner consistent with the theoretical prediction.
Testing asset pricing models is closely related to specification searchanalysis in quantitative economics. Most specification search processes selectmodels based on some goodness of fit statistic (such as R2 orrelated F). The effects of the sequential search on the statistical testsshould be taken into account when looking for the maximum goodness of fit. Toavoid misspecified models it is useful to study the selected models based bothon the full sample and along the sample. This paper presents a conditionalsequential procedure for the specification search process with linearregression models that minimizes data snooping or data mining. It is acombined test that first considers the search for the `best' set of regressorsand, conditional on this set, studies its significance and/or stability alongthe sample. The characteristics of the conditional tests are presented. Itsefficacy is illustrated with a model of future returns as a function of pastvolume and returns.
This papers proposes a two-step procedure for testing for constant parameters in GARCH models using recursive tests. We test the stability of the parameters conditional on consistent estimates obtained with the full sample. In the first step, we identify and estimate the best GARCH model using the full sample. In the second, using the first-step results and noting that the conditional variance is consistent under the null of constant parameters, recursive statistics are used to test for parameter constancy. The tests are applied to the S&P 500 stock index and to series of the exchange rates of the dollar for the mark, pound, and yen.
Recursive estimates can be useful for diagnostic purposes, but algorithms for estimating dynamic models recursively with autocorrelated perturbations can be computationally complicated. Thus, we propose a Conditional Recursive Least Squares algorithm (CRLS): given initial full-sample consistent estimates obtained from a correctly specified model, the model is linearized to obtain recursive consistent estimators along the full sample. These may in turn be used to compute statistics to test for structural breaks with unknown break dates. This procedure is illustrated with the Gas-Furnace data.
Specification analysis precedes model selection for structural analysis or forecasting. To explain a variable, one chooses an optimal subset of k predictors among m indicated variables, often maximizing some goodness of fit or R^2 (or F ). Without such a process, one has potentially misleading data mining. Foster et al. (1997) use maximum R^2 to for this purpose. They feel proper cut-off points of the R^2 distribution require consideration of the selection procedure and hence the use of the distribution function of the maximal R^2 . This difficult function must either be simulated by Monte Carlo or approximated as in Foster et al. with Bonferroni or Rencher and Pun bounds. White (1997) proposes using a 'Reality Check,' comparing forecasting performance of the candidate against a benchmark. Out-of-sample prediction is a good performance test, but choosing the benchmark model is more difficult. Surprisingly the full sample is not often exploited in testing for data mining. We argue that testing with both full sample and recursive estimation along the sample reduces data mining problems. Before accepting a model with significant global R^2 , it is of use to test for coefficient stability and significance of R^2 along the full sample. A sound theoretical model should remain valid if estimated and tested recursively. Foster et al. use R^2 estimated with the full sample. But models may comply with maximal R^2 statistics and be spurious (nonconstant coefficients). We propose to consider the information from the recursive estimations to detect this situation. We add to the processes of model selection and data mining possible parameter variation, which can bias the choice of benchmark model or the specification search among the m variables. Time-varying parameters (TVP) that are assumed constant produce misspecification error, possibly contaminating subsequent analyses. Thus, del Hoyo and Llorente (1998a) study the improvement in forecasting arising by considering non constant parameters. We consider both means (discrimination and stability) for decreasing biases in choosing a model. The first stage uses the R^2 or R^2_{max} to select the optimal explanatory variables. The second stage tests stability and constancy of the relationship. The conditional distributions of the recursive statistics are tabulated, conditional on the discrimination stage. The innovation here is the sequential consideration of both procedures. Section 1 introduces the problem. Section 2 tabulates the distributions of the relevant statistics, and their size and power are considered. Section 3 introduces the sequential procedure described above. The conditional distributions are studied. Section 5 gives an illustration with a model proposed by Campbell, Grossman and Wang (1993). Section 6 concludes.
This paper tests for the informational content of trading volume {\em per se} as postulated by Blume, Easley, and O'Hara (1994) and about its identification function in price reversals as postulated in Campbell, Grossman, and Wang (1993) using different trading strategy schemes. We find evidence about the information impounded in volume and its difference from that of return information. We also find evidence of price reversal for highly traded stocks. For low traded stocks we do not find price trending. All of these results are robust to the stock size, to different time horizons, and to different criterions for defining low and high volume.
Most speci cation search processes select models based on some goodness of t statistic (i.e. R or related F ). The e ects of the sequential search on the statistical tests should be taken into account when looking for the maximum goodness of t. To avoid misspeci ed models it is useful to study the selected models based on the full sample, and along the sample. This paper presents a conditional sequential procedure to be used in the speci cation search process of linear regression models, as a way to minimize data snooping or data mining. It is a combined test, rst it considers the search for the \best" set of regressors, and conditional on this set, it studies their signi cance along the sample. The characteristics of the conditional test are presented. Its usefulness is considered with one application.
Forecasting with unstable relationships under the assumption that they are constant, might cause inefficient forecasts. Thus, before embarking upon any forecast exercise, it is convenient to first verify the sequential significance of the relationship, and then its stability. This paper studies recursive and rolling estimators, and related sequential tests to test for significant and constant coefficients. If constancy is rejected, the estimations may suggest dynamic models, that jointly with the considered model, could improve the results of the forecasting exercise. These procedures are applied to the price volume relationship in the stock market as postulated in Campbell, Grossman and Wang (1993). The significance but non stability of the relationship is shown. There are gains in the forecast performance when considering the empirical model for the rolling estimators, jointly with the initial structural model. Key Words: Recursive Estimators, Rolling Estimators, Recursive Sequential Test, Monte Carlo Methods, Campbell-Grossman-Wang model.