Economic analysis based on text data has expanded rapidly, yet the ultra-high dimensionality of text data presents substantial challenges for term selection and estimation. We propose the Information-Adaptive Lasso and Information-Adaptive Spike-and-Slab Lasso as novel frequentist and Bayesian approaches to address these challenges. A key theoretical contribution of this study is the establishment of the rate of convergence in term-selection consistency, which we show to be faster than those achieved in the existing Lasso literature. Applying our methods to Federal Open Market Committee (FOMC) statements, we identify and estimate high-impact terms driving fluctuations in monetary policy uncertainty.
This article introduces and analyzes a framework that accommodates general heterogeneity in regression modeling. It demonstrates that regression models with fixed or time-varying parameters can be estimated using the ordinary least squares (OLS) and time-varying OLS methods, respectively, across a broad class of regressors and noise processes not covered by existing theory. The proposed setting facilitates the development of asymptotic theory and the estimation of robust standard errors. The robust confidence interval estimators accommodate substantial heterogeneity in both regressors and noise. The resulting robust standard error estimates coincide with White's (1980, Econometrica 48, 817-838) heteroskedasticity-consistent estimator but are applicable to a broader range of conditions, including models with missing data. They are computationally simple and perform well in Monte Carlo simulations. Their robustness, generality, and ease of implementation make them highly suitable for empirical applications. Finally, the article provides a brief empirical illustration.
The large heterogeneous panel data models are extended to the setting where the heterogenous coefficients are changing over time and the regressors are endogenous. Kernel-based non-parametric time-varying parameter instrumental variable mean group (TVP-IV-MG) estimator is proposed for the time-varying cross-sectional mean coefficients. The uniform consistency is shown and the pointwise asymptotic normality of the proposed estimator is derived. A data-driven bandwidth selection procedure is also proposed. The finite sample performance of the proposed estimator is investigated through a Monte Carlo study and an empirical application on multi-country Phillips curve with time-varying parameters.
In this paper, we propose a deep pooled estimator, motivated by the universal approximation property of neural networks, to capture nonlinear relationships between predictors and targets when modeling and forecasting with panel data. The approach is flexible, accommodating different penalty functions and potentially high-dimensional predictors. It allows for nonlinear cross-sectional dependencies. To evidence the utility of the proposed estimator when forecasting, we apply it in two different applications. First, we forecast the progression of new COVID-19 cases across G7 countries. Second, we forecast inflation in the G7. In both applications, our method delivers significant forecasting gains over both linear panel and nonlinear time-series (unit-specific) models that do not pool data across countries. These results highlight the importance when forecasting of pooling cross-country information via a flexible nonlinear model. Examining partial derivatives from our model provides interpretable insights: school closures and workplace restrictions show declining effectiveness as COVID-19 immunity strengthened, while the inflation-unemployment relationship proves highly unstable across both countries and time periods, particularly during the post-pandemic inflation surge.
We propose a regularized OLS-type estimator for variable selection and estimation in high-dimensional linear regression models. Our approach builds on recent advances in sparse covariance and precision matrix estimation, employing a thresholding-based regularization scheme that jointly incorporates all available regressors. This allows for a theoretically grounded estimation framework that extends beyond existing procedures, and accommodates more general assumptions on the regressors, including dependent stochastic processes and fat tails. We establish asymptotic validity of the estimator and develop an enhanced procedure based on a Bonferroni-type testing rule, which improves the identification of true regressors. We assess the finite sample performance of our approach by an extensive set of Monte Carlo experiments, also comparing it with other alternatives proposed in the literature. Finally, we illustrate the practical relevance of the methodology in an empirical macroeconomic forecasting application.
This paper studies model selection in high-dimensional regression settings. We propose a Boosting with Multiple Testing (BMT) approach designed to address proxy-signal contamination in environments with strongly correlated regressors, including settings driven by latent common factors. At each stage, a single regressor is selected conditional on those already included, while a family-wise multiple testing filter is applied to the remaining candidates. Under weak dependence conditions, we show that BMT enjoys oracle-type properties relative to an approximating model that contains all true signals and excludes pure noise variables. In particular, the probability of selecting any noise regressor converges to zero, while all true signals are retained with probability approaching one. Additional novel results establish conditions under which BMT recovers the exact true model andavoids the selection of proxy regressors. A theoretical comparison shows that BMT remains consistent in settings where penalised regression methods, such as the Lasso, fail due to restrictive global design conditions. Monte Carlo evidence confirms that BMT delivers higher model selection accuracy and lower estimation error, particularly under strong multicollinearity. Empirical illustrations using macro–financial data further demonstrate that BMT yields sparse, interpretable models with favourable out-of-sample performance.
We introduce a multiscale measure of network reconfiguration based on the joint use of Detrended Cross-Correlation Analysis (DCCA) and Minimum Spanning Tree (MST) filtering. The proposed metric – the Elastic Detrended Cross-Correlation Ratio (Elastic DCCR) – is defined as a finite-difference measure of the logarithmic sensitivity of the average MST length to the observation scale. It captures how the structure of cross-correlation networks deforms across different investment horizons. When applied to a network of global equity indices, the Elastic DCCR exhibits sharp deviations during episodes of financial stress, reflecting abrupt reorganisations of the scale-dependent geometry of cross-market correlations; the sign of the deviation distinguishes episodes of increasing from decreasing multiscale coherence. The measure reveals scale-dependent reconfigurations in network topology that are not visible in single-scale analyses, and highlights clear differences between periods of elevated market stress and periods of relative calm. The approach does not assume covariance stationarity and relies only on scale-dependent detrended correlations; as a result, it is broadly applicable to other complex systems in which interaction strength varies with scale.
This paper analyzes how public-health–induced panics reshape price discovery between the overnight and trading-hours windows. Using COVID-19 as a natural experiment, we construct text-based measures that disentangle pandemic-related sentiment from conventional economic pessimism and decompose returns into overnight and intraday legs. Fiscal outlays and vaccine developments are included in our analysis as countermeasures. Our results suggest that sentiment derived from news integrating pandemic and economic information displaced conventional economic pessimism as the dominant short-horizon driver. During the Omicron wave, such news generated sizable opening gaps with limited subsequent reversals. Investors reacted positively to announcements about fiscal recovery programs, vaccine breakthroughs, and distribution milestones; the marginal impact of this information diminished once policies were implemented and vaccines became widely available, consistent with swift information assimilation. Overall, we provide a general mechanism for market recoveries from public-health rare disasters in which panic advances and concentrates price discovery in the overnight window, and the contributions of fiscal and vaccine announcements to price discovery exhibit diminishing marginal effects as implementation proceeds.
This paper examines how relationships between parent-firm characteristics and facility-level toxic releases evolve over time. Using 238,304 observations for 7,447 U.S. manufacturing facilities from 1992 to 2023, we link on-site releases from the Toxics Release Inventory to financial, managerial, and macroeconomic data. A time-varying mean-group estimator accommodates changes in average coefficients over time and heterogeneity across facilities. The main result is a persistent contrast between operating scale and investment-related adjustment: sales is positively associated with release growth, whereas investment intensity is negatively associated. This pattern remains in joint specifications, parsimonious models, and the main robustness exercises. Other financial and managerial associations are identified as well. These findings suggest that environmental policy evaluation may benefit from distinguishing release changes associated with production expansion from those associated with investment-related adjustment and from allowing for variation across periods and production settings. For corporate managers, the same distinction can inform how emissions considerations are incorporated into capital budgeting and production planning.
We develop a novel methodology for estimation and inference in high-dimensional panel network models with latent dual structures. The framework allows outcomes to be affected simultaneously by positive and negative interaction channels, accommodating settings in which some interactions reinforce outcomes while others generate competition and displacement effects. The proposed method identifies and estimates the network directly from the structural model using observed data without the need to pre-specify the network. Network recovery is achieved through a sequential instrumental-variable screening procedure. We establish exact support recovery and oracle-equivalent post-selection inference. An application to U.S. corporate leverage data reveals the coexistence of reinforcing and displacement interactions in firms' financial decisions.
This paper proposes a nonlinear boosting with multiple testing (BMT) approach to variable selection in high-dimensional generalised linear models with binary responses. At each stage of the BMT procedure, the model is updated by adding only the most significant covariate, conditional on those already selected in previous stages, while taking into account the multiple testing nature of the problem. It is shown that, under the stated conditions, the BMT procedure selects all covariates whose true coefficients are nonzero, and no other covariates, with probability tending to one. Furthermore, the procedure enjoys an oracle property, in the sense that the post-BMT maximum likelihood estimator of the parameters of the model is asymptotically equivalent to an oracle estimator that knows the correct sparse model in advance. Monte Carlo experiments demonstrate that BMT outperforms competing methods, delivering high covariate-selection accuracy and low parameter estimation error. An empirical example illustrates that BMT delivers a predictive model for the probability that U.S. inflation exceeds a given threshold over a 12-month horizon which has very good out-of-sample performance.
We study spillover effects in corporate toxic emissions using a heterogeneous panel network of U.S. industrial facilities from 2000-2023. Rather than imposing a network structure a priori, we uncover an unobserved web of influence directly from the data using recent advances in high-dimensional network econometrics. Indirect effects transmitted through the estimated network account for about 28
High-dimensional regression specification and analysis is a complex and active area of research in statistics, machine learning, and econometrics. This paper proposes a new approach, Boosting with Multiple Testing (BMT), which combines forward stepwise variable selection with the multiple testing framework of Chudik et al (2018). At each stage, the model is updated by adding only the most significant regressor conditional on those already included, while a family-wise multiple testing filter is applied to the remaining candidates. In this way, the method retains the strong screening properties of Chudik et al (2018) while operating in a less greedy manner with respect to proxy and noise variables. Using sharp probability inequalities for heterogeneous strongly mixing processes from Dendramis et al (2022), we show that BMT enjoys oracle type properties relative to an approximating model that includes all true signals and excludes pure noise variables: this model is selected with probability tending to one, and the resulting estimator achieves standard parametric rates for prediction error and coefficient estimation. Additional results establish conditions under which BMT recovers the exact true model and avoids selection of proxy signals. Monte Carlo experiments indicate that BMT performs very well relative to OCMT and Lasso type procedures, delivering higher model selection accuracy and smaller RMSE for the estimated coefficients, especially under strong multicollinearity of the regressors. Two empirical illustrations based on a large set of macro-financial indicators as covariates, show that BMT yields sparse, interpretable specifications with favourable out-of-sample performance.
We introduce a multiscale measure of network instability based on the joint use of Detrended Cross-Correlation Analysis (DCCA) and Minimum Spanning Tree (MST) filtering. The proposed metric, the Elastic Detrended Cross-Correlation Ratio (Elastic DCCR), is defined as a finite-difference measure of the logarithmic sensitivity of the average MST length to the observation scale. It captures how the structure of cross-correlation networks deforms across different investment horizons. When applied to a network of global equity indices, the Elastic DCCR rises sharply during episodes of financial stress, reflecting increased short-term coordination among investors and a contraction of correlation distances. The measure reveals scale-dependent reconfigurations in network topology that are not visible in single-scale analyses, and highlights clear differences between stressed and stable market regimes. The approach does not assume covariance stationarity and relies only on scale-dependent detrended correlations; as a result, it is broadly applicable to other complex systems in which interaction strength varies with scale.
We estimate the fiscal (spending) multiplier using quarterly US data, 1981Q3-2024Q4. We define government spending shocks as actual minus expected expenditure growth, the latter obtained from the Survey of Professional Forecasters. We employ the Jord & agrave; local projections method, coupled with state dependence of parameters, with smooth transition between states. A key testable hypothesis is that the positive and negative spending shocks have numerically (as well as qualitatively) different effects. We find that multipliers of shocks differ qualitatively (in terms of cyclicality) as well as quantitatively. Multipliers are almost always above unity (and often well above). Importantly, we uncover evidence that negative shocks have stronger effects over longer periods of time in the case of the FEC multipliers and likely with the PVIR-FM multipliers, too. Pooled-shock estimation can seriously bias results. Additionally, there is strong evidence that the two types of shock produce almost uniformly significantly different estimated coefficients of our key estimable equation.
In this paper we examine the existence of heterogeneity within a group, in panels with latent grouping structure. The assumption of within group homogeneity is prevalent in this literature, implying that the formation of groups alleviates cross-sectional heterogeneity, regardless of the prior knowledge of groups. While the latter hypothesis makes inference powerful, it can be often restrictive. We allow for models with richer heterogeneity that can be found both in the cross-section and within a group, without imposing the simple assumption that all groups must be heterogeneous. We further contribute to the method proposed by \cite{su2016identifying}, by showing that the model parameters can be consistently estimated and the groups, while unknown, can be identifiable in the presence of different types of heterogeneity. Within the same framework we consider the validity of assuming both cross-sectional and within group homogeneity, using testing procedures. Simulations demonstrate good finite-sample performance of the approach in both classification and estimation, while empirical applications across several datasets provide evidence of multiple clusters, as well as reject the hypothesis of within group homogeneity.
Durbin regressions are found to be remarkably successful in terms of estimating static regression coefficients in the presence of regression errors which may contain long memory or non linear components. The paper extends Baillie, Diebold, Kapetanios, Kim and Mora (2025) which focuses on weakly stationary AR errors to this wider context. The results suggest that Durbin regressions should the preferred approach for estimation of time series regressions. The paper also documents the poor performance of OLS-HAC inference when applied to these regressions.
We investigate whether the Balanced Labour Market Act (WAB) of 2020, intended to reduce the disparity between permanent and temporary employees in The Netherlands, has achieved its desired aim. Using a synthetic control method, we find that the introduction of the WAB led to a substantial reduction in the number of temporary contracts, whereas the number of permanent workers increased. As of the WAB's announcement in May 2019, strong anticipatory effects were evident. Inconclusive evidence of unintended side-effects, such as the substitution of temporary workers by self-employment work schemes, marks an avenue for future research.
The Themed Issue Machine Learning for Economic Policy consists of 12 papers at the intersection of machine learning, nontraditional data sources and economic policymaking. We will introduce the Themed Issue and review its contributions.