This paper examines the use of Hidden Markov Models (HMMs) to investigate patterns in US bank failures. HMMs are statistical tools designed for handling sequential data and uncovering hidden states, making them particularly useful for studying systemic financial events. The paper provides an application of HMMs to U.S. bank failures data and presents results of an empirical study that analyzes historical data on banking crises in the United States. Four HMMs are employed for the analysis of U.S. bank failures data, and their relative performance is studied with respect to model fitting and forecasting. Empirical results are presented for each of the HMMs employed. Global decoding is employed to predict the most likely state sequence of large bank failure events via the Viterbi algorithm. The findings demonstrate the ability of HMMs to uncover unobservable economic conditions and enhance predictive capabilities for financial stability evaluations.
Dynamic linear regression models are used widely in applied econometric research. Most applications employ linear autoregressive (AR) models, distributed lag (DL) models or autoregressive distributed lag (ARDL) models. These models, however, perform poorly for data sets with unknown, complex nonlinear patterns. This paper studies nonlinear and semiparametric extensions of the dynamic linear regression model and explores the autoregressive (AR) extensions of two semiparametric techniques to allow unknown forms of nonlinearities in the regression function. The autoregressive GAM (GAM-AR) and autoregressive multivariate adaptive regression splines (MARS-AR) studied in the paper automatically discover and incorporate nonlinearities in autoregressive (AR) models. Performance comparisons among these semiparametric AR models and the linear AR model are carried out via their application to Australian data on growth in GDP and unemployment using RMSE and GCV measures. Â
While a variety of specification tests are routinely employed to test for misspecification in linear regression model, such tests and their applications to the truncated and censored regression models are uncommon. This paper develops a regression error specification test (RESET) for the truncated regression model as an extension of the popular RESET for the linear regression model (Ramsey (1969)). The two proposed extensions TRESET1 and TRESET2 developed in the paper are applied to labor force participation data from Mroz (1987). The paper studies the empirical size and power properties of the proposed tests via Monte Carlo experiments. Our simulation results suggest that both TRESET tests have reasonably good size and power properties for the truncated regression model in medium to large samples. However, TRESET2 consistently outperforms TRESET1 both in terms of empirical size and power in our experiments. Â
The paper studies various response transformation models for discrete choice and categorical data. These response transformation models are fitted to binary response data on beverage choice. Several models are compared, and the best model is selected using AICs and deviances. The transformations include extensions of the widely used Box-Cox transformation to Normality for continuous data to categorical data. The econometric techniques employed in the paper are widely applicable to the analysis of count, binary response, and duration types of data encountered in business and economics.
The paper presents applications of a class of semi-parametric models called generalized additive models (GAMs) to several business and economic datasets. Applications include analysis of wage-education relationship, brand choice, and number of trips to a doctors office. The dependent variable may be continuous, categorical or count. These semi-parametric models are flexible and robust extensions of Logit, Poisson, Negative Binomial and other generalized linear models. The GAMs are represented using penalized regression splines and are estimated by penalized regression methods. The degree of smoothness for the unknown functions in the linear predictor part of the GAM is estimated using cross validation. The GAMs allow us to build a regression surface as a sum of lower-dimensional nonparametric terms circumventing the curse of dimensionality: the slow convergence of an estimator to the true value in high dimensions. For each application studied in the paper, several GAMs are compared and the best model is selected using AIC, UBRE score, deviances, and R-sq (adjusted). The econometric techniques utilized in the paper are widely applicable to the analysis of count, binary response and duration types of data encountered in business and economics.
Principal Component Analysis (PCA) is a very versatile technique for dimension reduction in multivariate data. Classical PCA is very sensitive to outliers and can lead to misleading conclusions in the presence of outliers. This article studies the merits of robust PCA relative to classical PCA when outliers are present. An algorithm due to Filzmoser et al. (2006) based on a modification of the projection pursuit algorithm of Croux and Ruiz-Gazen (2005) is used for robust PCA computations for a financial data set as well as simulated data sets. Our simulation results indicate that robust PCA generally leads to greater reduction in model dimension than classical PCA in data sets with outliers.
This article presents robust J and encompassing tests for testing nonnested hypotheses in the presence of outliers in the data. The proposed tests are based on least absolute deviations (LAD) and M-estimators unlike the conventional J and encompassing tests, which are based on least squares or maximum likelihood estimators. These tests can lead to more reliable inference in the presence of outliers than tests based on nonrobust estimators. The tests are illustrated with applications to two economic data sets and an artificially generated data set and compared with their nonrobust counterparts.
Exogeneity testing is studied in the presence of outliers in response variables. Robust tests based on least absolute deviations (LAD) and M estimators are proposed and illustrated with an application to Mroz (1987) data. Our simulation results show that the proposed robust tests outperform the traditional Hausman test for exogeneity in terms of empirical power in the presence of outliers in response variables. Nevertheless, unlike the conventional Hausman test, which is undersized, the empirical size of the LAD-based exogeneity test exceeds its nominal size.
Generalized linear models (GLMs) are generalizations of linear regression models, which allow fitting regression models to response data that follow a general exponential family. GLMs are used widely in social sciences for fitting regression models to count data, qualitative response data and duration data. While a variety of specification tests have been developed for the linear regression model and are routinely applied for testing for misspecification of functional form, omitted variables, and the normality assumption, such tests and their applications to GLMs are uncommon. This paper develops a regression error specification test (RESET) for GLMs as an extension of the popular RESET for the linear regression model (Ramsey (1969)). Applications of the RESET to three economic data sets are presented and the finite sample power properties are studied via a Monte Carlo experiment.
Deriving the observed information matrix in ordered probit and logit models using the complete-data likelihood function—solution.
Common econometric estimators such as least squares, least absolute deviations (LAD), instrumental variables, maximum likelihood, and semiparametric estimators are non-robust against data contamination. Despite the known superiority of high-breakdown point (HBP) estimators in such situations, the HBP estimators have rarely been used in economics. This article presents some applications of an HBP estimator called the S-estimator (Rousseeuw and Yohai, Robust and Nonlinear Time Series Analysis (Eds) W. H. Franke and R. D. Martin, Springer-Verlag, NY, pp. 256-72, 1984) to the estimation of a linear regression model and compares the results with those obtained by ordinary least squares (OLS) and LAD methods. It is found that significance of variables as well as signs of coefficient estimates can be quite different under HBP estimation than under OLS and LAD estimation.
Pre-test estimation has been studied extensively for linear regression and simultaneous equation models. Recently attention has turned to pre-test estimation in non-linear models. This article studies pre-test maximum likelihood estimation in Poisson regression model. It presents its risk characteristics and compare them with those of restricted and unrestricted maximum likelihood estimators based on squared error loss function in a Monte Carlo experiment.
The EM algorithm is a widely used technique for finding maximum likelihood (ML) estimates when the data are not fully observed. Despite its popularity for computing ML estimates in unrestricted problems and the need for simplified computations for problems with equality and inequality restrictions, there have been few applications of the algorithm to restricted ML estimation. The EM algorithm is presented for restricted ML estimation and provides its applications to the probit model under equality and inequality restrictions using two small data sets.
The maximum likelihood estimator for the Probit model can be substantially biased in small samples. This paper proposes a bias-corrected jackknife maximum likelihood estimator (JMLE) for the Probit model which corrects bias up to O(1/n-squared) unlike the ordinary MLE which corrects bias up to O(1/n). An application of the JMLE to Spector and Mazzeo (1980) data for analysing the effectiveness of a new method of teaching economics is also presented.
The problem of unidentifiability of parameters of the complete data pdf due to truncation is studied in the single sample case. It is shown that if the support of the complete data pdf depends on one of the parameters of the pdf and the pdf takes a special multiplicative form, then the parameter on which the support of the pdf depends is not identifiable from the truncated pdf. Some examples of the pdf's for which unidentifiability can occur due to truncation are also given.
In a recent paper in this journal, Manning showed that the ordinary least squares (OLS) estimator of regression coefficients in a logit model with continuous and bounded dependent variable measured with error is unbiased but inefficient. This note shows that, contrary to the author's claim, the OLS estimator is biased in general. Manning's form of heteroscedasticity is also shown to be incorrect. It is shown that Manning's assumptions on the distribution of measurement error are inadequate for determining the form of heteroscedasticity.
Robinson (1982a) presented a general approach to serial correlation in limited dependent variable models and proved the strong consistency and asymptotic normality of the quasi-maximum likelihood estimator (QMLE) for the Tobit model with serial correlation, obtained under the assumption of independent errors. This paper proves the strong consistency and asymptotic normality of the QMLE based on independent errors for the truncated regression model with serial correlation and gives consistent estimators for the limiting covariance matrix of the QMLE. Keywords: tobit modeltruncated regression modelserial correlationquasi-maximum likelihood estimator
A relationship between the logit model, normal discriminant analysis, and mixtures of multivariate normal distributions is discussed. It is shown that the likelihood equations for multivariate normal mixtures can be obtained from the likelihood equations for the normal discriminant analysis model by simply replacing the index variables with weights that are logistic probabilities, and that all three models use the same linear discriminant function to classify observations into different populations. Some implications of these relationships for data analysis are also discussed.