This article proposes and studies two Huber-type estimation approaches, namely, the Huber instrumental variable (IV) estimation and the Huber generalized method of moments (GMM) estimation, for a spatial autoregressive model. We establish the consistency, asymptotic distributions, finite sample breakdown points, and influence functions of these estimators. Simulation studies show that compared to the corresponding traditional estimators (the two-stage least squares estimator, the best IV estimator, and the GMM estimator), our estimators are more robust when the unknown disturbances are long-tailed, and our estimators only lose a little efficiency when the disturbances are short-tailed. Moreover, the Huber GMM estimator also outperforms several robust estimators in the literature. Finally, we apply our estimation method to investigate the impact of the urban heat island effect on housing prices. A package is published on GitHub for practitioners to use in their empirical studies.
This article proposes an efficient two-step estimator for a spatial autoregressive (SAR) model with SAR disturbances (SARAR). By leveraging the residual-adjusted estimation framework of Hatanaka (1974, 1976) and Dhrymes (1974), our estimator achieves asymptotic efficiency comparable to the quasi-maximum likelihood estimator (QMLE) or the best generalized method of moments estimator (BGMME) in Liu, Lee, and Bollingerm (2010), while significantly reducing computational complexity and execution time. Monte Carlo simulations demonstrate the superior performance of our numerical procedure across both small and relatively large sample sizes. An empirical application to U.S. county-level homicide data reveals significant positive spatial spillover effects, highlighting the critical need for multi-regional collaboration in crime prevention and economic development policies to reduce homicide rates.
Summary We provide a novel analytic procedure to construct best linear and quadratic moments of the generalized method of moments estimation for a large class of cross‐sectional network and spatial econometric models. These moments generate an estimator that is asymptotically more efficient than the quasi‐maximum likelihood estimator when the disturbances follow a non‐normal and unknown distribution. We apply this procedure to a high‐order spatial autoregressive model with spatial errors, where the disturbances are heteroskedastic. Two normality tests of disturbances are developed. We apply the model to employment data in US counties, which demonstrates spatial interdependence patterns of regional employment growth.
We introduce a spatial autoregressive hurdle model for nonnegative origin–destination flows yN,ij. The model incorporates a hurdle formulation to elucidate the different data-generating processes for zero and positive flows. Our model specifies three types of spatial influences on flow yN,ij that quantify the impact of third-party characteristics on the flow yN,ij: (i) the effect of outflows from origin j, (ii) the effect of inflows to destination i, and (iii) the effect of flows among third-party units. We account for two-way fixed effects in the model to capture the inherent characteristics of both origins and destinations. We employ maximum likelihood estimation to estimate the model parameters. To address statistical inference issues, we analyze the asymptotic properties of the ML estimator using the spatial near-epoch dependence concept. We confirm the presence of an asymptotic bias that arises from the fixed effects, whose dimensions grow with the sample size. Applying our model to migration flows among U.S. states, we estimate significant spatial influences, particularly from inflows to destinations and outflows from origins. Our findings support the notion that zero and positive flow formations are distinct. Consequently, our proposed model outperforms the spatial autoregressive Tobit specification for origin–destination flows, thus providing a better fit to the data.
This article investigates QML and GMM estimation of spatial autoregressive (SAR) models in which the column sums of the spatial weights matrix might not be uniformly bounded. We develop a central limit theorem in which the number of columns with unbounded sums can be finite or infinite and the magnitude of their column sums can be O(n(delta)) if delta < 1. Asymptotic distributions of QML and GMM estimators are derived under this setting, including the GMM estimators with the best linear and quadratic moments when the disturbances are not normally distributed. The Monte Carlo experiments show that these QML and GMM estimators have satisfactory finite sample performances, while cases with a column sums magnitude of O(n) might not have satisfactory performance. An empirical application with growth convergence in which the trade flow network has the feature of dominant units is provided. Supplementary materials for this article are available online.
This article investigates asymptotic properties of quasi-maximum likelihood (QML) estimates for flow data on the dual gravity model in international trade with spatial interactions (dependence). The dual gravity model has a well-established economic foundation, and it takes the form of a spatial autoregressive (SAR) model. The dual gravity model originates from Behrens et al., but the spatial weights matrix motivated by their economic theory has a feature that violates existing regularity conditions for asymptotic econometrics analysis. By overcoming the limitations of existing asymptotic theory, we show that QML estimates are consistent and asymptotically normal. The simulation results show the satisfactory finite sample performance of the estimates. We illustrate the usefulness of the model by investigating the McCallum "border puzzle" in the gravity literature.
We extend LeSage and Pace (2008)'s spatial autoregressive model for origin–destination flows by accommodating two-way fixed effects. A partial likelihood approach is used for estimation by applying an orthogonal transformation to remove fixed effects in the model. The quasi-maximum likelihood (QML) estimator of the partial log-likelihood function is consistent and asymptotically centered normal. Monte Carlo experiments verify this advantage in finite samples. From the U.S. migration flows, significant spatial influences are captured with smaller magnitudes than those from the model without fixed effects.
When studying the consistency of an estimator without a closed-form solution for a spatial econometric model, we usually assume that the parameter space is compact. However, compactness assumptions are restrictive as we need to know the boundaries of parameter spaces. We establish a consistency theorem for concave objective functions. We apply this result to rebuild the consistency of the quasi maximum likelihood estimator (QMLE) of a spatial autoregressive (SAR) model and a SAR Tobit model. Their log-likelihood functions are not concave, but they can be concave after proper reparameterization as in Olsen (1978).
This paper studies asymptotic properties of a posterior probability density and Bayesian estimators of spatial econometric models in the classical statistical framework. We focus on the high-order spatial autoregressive model with spatial autoregressive disturbance terms, due to a computational advantage of Bayesian estimation. We also study the asymptotic properties of Bayesian estimation of the spatial autoregressive Tobit model, as an example of nonlinear spatial models. Simulation studies show that even when the sample size is small or moderate, the posterior distribution of parameters is well approximated by a normal distribution, and Bayesian estimators have satisfactory performance, as classical large sample theory predicts.
This paper introduces a spatial panel data model describing local agents’ intertemporal decision-making with their network interactions and network evolution. Our model’s purpose is to give a tool to analyze local governments’ behaviors when there exist network interactions among them with endogenously changing networks. To provide a theoretical foundation of our model, we establish a network interaction model for forward-looking agents. An agent’s current action can affect his future network links via time-varying economic indicators. To estimate the model’s parameters, we consider a GMM estimation method based on first-order conditions of agents’ approximated lifetime problems. Asymptotic properties of the GMM estimator are studied for statistical inferences. Using our model, we find evidence of positive spillovers among U.S. states’ public welfare expenditures with coevolution between them and spatial-economic networks.
This paper introduces dynamic panel spatial vector autoregressive models. We study features of dynamics and spatial interactions that an SVAR model can generate and classify the model into stable or unstable cases by partitioning parameter spaces. For stable, spatial cointegration, and mixed cointegration cases, we investigate identification and QML estimation of the models to take into account simultaneity and correlated relationships. Asymptotic properties and bias-corrected estimators are presented. To detect unknown cointegration relationships, we introduce a sequential likelihood ratio testing procedure. Simulations show the advantage of QMLEs on bias reduction and efficiency gains. The empirical application provides evidences on ancient China’s market integration.
This paper considers generalized method of moments (GMM) and sequential GMM (SGMM) estimation of dynamic short panel data models. The efficient GMM motivated from the quasi maximum likelihood (QML) can avoid the use of many instrument variables (IV) for estimation. It can be asymptotically efficient as maximum likelihood estimators (MLE) when disturbances are normal, and can be more efficient than QML estimators when disturbances are not normal. The SGMM, which also incorporates many IVs, generalizes the minimum distance estimation originated in Hsiao et al. . By focusing on the estimation of parameters of interest, the SGMM saves computational burden caused by nuisance parameters such as variances of disturbances. It is asymptotically as efficient as the corresponding GMM. In particular, the SGMM based on QML scores can generate a closed-form root estimator for the dynamic parameter, which is asymptotically as efficient as the QML estimator. Nuisance parameters can also be estimated efficiently by an additional SGMM step if they are of interest.
This article presents a methodology for empirically identifying the key player, whose removal from the network leads to the optimal change in aggregate activity level in equilibrium [Ballester, C., Calvo-Armengol, A., and Zenou, Y. (2006), "Who's Who in Networks. Wanted: The Key Player," Econometrica, 74: 1403-1417], allowing the network links to rewire after the removal of the key player. First, we propose an IV-based estimation strategy for the social-interaction effect, which is needed to determine the equilibrium activity level of a network, taking into account the potential network endogeneity. Next, to simulate the network evolution process after the removal of the key player, we adopt the general network formation model in Mele [(2017), "A Structural Model of Dense Network Formation," Econometrica, 85: 825-850] and extend it to incorporate the unobserved individual heterogeneity in link formation decisions. We illustrate the methodology by providing the key player rankings in juvenile delinquency using information on friendship networks among U.S. teenagers. We find that the key player is not necessarily the most active delinquent or the delinquent who ranks the highest in standard (not microfounded) centrality measures. We also find that, compared to a policy that removes the most active delinquent from the network, a key-player-targeted policy leads to a much higher delinquency reduction.
This paper considers two-step generalized empirical likelihood (GEL) estimation and tests with martingale differences when there is a computationally simple $\sqrt n$ -consistent estimator of nuisance parameters or the nuisance parameters can be eliminated with an estimating function of parameters of interest. As an initial estimate might have asymptotic impact on final estimates, we propose general $C(\alpha )$ -type transformed moments to eliminate the impact, and use them in the GEL framework to construct estimation and tests robust to initial estimates. This two-step approach can save computational burden as the numbers of moments and parameters are reduced. A properly constructed two-step GEL (TGEL) estimator of parameters of interest is asymptotically as efficient as the corresponding joint GEL estimator. TGEL removes several higher-order bias terms of a corresponding two-step generalized method of moments. Our moment functions at the true parameters are martingales, thus they cover some spatial and time series models. We investigate tests for parameter restrictions in the TGEL framework, which are locally as powerful as those in the joint GEL framework when the two-step estimator is efficient.
This paper investigates the first difference (FD) estimation of spatial dynamic panel data (SDPD) models with fixed effects using quasi-maximum likelihood (QML) approach, where both n and T are large. We show that the QML estimation for the SDPD with FD can be reduced to the direct estimation of individual effects, except for the estimation of variance parameter. After bias correction, these two approaches would yield asymptotically equivalent estimates for all parameters including the variance parameter. Our results extend the equivalence of LSDV estimate and GLS estimate of FD equation in the panel regression model to the spatial dynamic panel model, which includes the conventional dynamic panel as a special case. Our analysis highlights the importance of initial values rather than the many fixed effects in spatial panel models.
This paper considers closed-form root estimators for spatial autoregressive models with spatial autoregressive disturbances (SARAR model). We first derive a simple consistent closed-form estimator. Then we construct feasible moment conditions that are quadratic in the spatial lag and spatial error dependence parameters separately, which generate root estimators with closed forms. We consider both the cases with homoskedastic and unknown heteroskedastic disturbances. In the homoskedastic case, the root estimator can be asymptotically as efficient as the quasi-maximum likelihood estimator (QMLE); in the heteroskedastic case, it can be asymptotically as efficient as a method of moments estimator (MME) that sets adjusted quasi-maximum likelihood scores to zero, where the adjusted scores have zero means at the true parameters. The root estimators and their associated standard errors can avoid the computation of any matrix determinants or inverses, so they are computationally simple without iterations, especially valuable for big data where the sample size is large.
This paper investigates the quasi-maximum likelihood estimation of short dynamic panel data models. We consider their estimation on both fixed effects and random effects specifications and propose a Hausman test when exogenous variables are present. For a dynamic panel model, initial conditions play important roles in model structure and estimation, and they give rise to a between equation under the random effects framework. With the between equation properly defined, we show that the random effects model can be decomposed into a within equation and a between equation; hence, the random effects estimate is a pooling of the within and between estimates. Thus, our paper extends the pooling in the static panel data model (Maddala, 1971a) to the setting of dynamic panel data. This decomposition of a dynamic panel data model is revealing and valuable for estimation and the formulation of a Hausman test to test the possible correlation of individual effects with included regressors. Monte Carlo experiments are conducted to investigate the finite sample performance of estimators and the Hausman test. An empirical application of growth convergence in OECD countries is provided.
This paper studies the estimation of a cross-sectional spatial autoregressive (SAR) model with spatial weights constructed by bilateral variables like the trade or investment between regions. We model the possible endogeneity in spatial weights due to the correlation between the error term in the SAR model and unobserved interactive fixed effects in bilateral variables. Using a control function approach, we propose two-stage estimation methods and establish their consistency and asymptotic normality. Finite sample properties are investigated by a Monte Carlo study. We further apply our method to an empirical study of interactions among different US industries through production networks.
We model network formation and interactions under a unified framework by considering that individuals anticipate the effect of network structure on the utility of network interactions when choosing links. There are two advantages of this modeling approach: first, we can evaluate whether network interactions drive friendship formation or not. Second, we can control for the friendship selection bias on estimated interaction effects. We provide microfoundations of this statistical model based on the subgame perfect equilibrium of a two‐stage game and propose a Bayesian MCMC approach for estimating the model. We apply the model to study American high school students' friendship networks using the Add Health dataset. From two interaction variables, GPA and smoking frequency, we find that the utility of interactions in academic learning is important for friendship formation, whereas the utility of interactions in smoking is not. However, both GPA and smoking frequency are subject to significant peer effects.