We show how to identify the distributions of the latent components in the two-way dyadic model for bipartite networks y(i,l) = alpha(i)+eta(l)+epsilon(i,l). This is achieved by a repeated application of the extension of the classical lemma of Kotlarski (1967) in Evdokimov and White (2012). We provide two separate sets of assumptions under which all the latent distributions are identified. Both rely on some of the latent components being identically distributed.
We derive optimal statistical decision rules for discrete choice problems when payoffs depend on a partially-identified parameter $\theta$ and the decision maker can use a point-identified parameter $P$ to deduce restrictions on $\theta$. Leading examples include optimal treatment choice under partial identification and optimal pricing with rich unobserved heterogeneity. Our optimal decision rules minimize the maximum risk or regret over the identified set of payoffs conditional on $P$ and use the data efficiently to learn about $P$. We discuss implementation of optimal decision rules via the bootstrap and Bayesian methods, in both parametric and semiparametric models. We provide detailed applications to treatment choice and optimal pricing. Using a limits of experiments framework, we show that our optimal decision rules can dominate seemingly natural alternatives. Our asymptotic approach is well suited for realistic empirical settings in which the derivation of finite-sample optimal rules is intractable.
This paper studies a simple dynamic linear panel regression model with interactive fixed effects in which the variable of interest is measured with error. To estimate the dynamic coefficient, we consider the least-squares minimum distance (LS-MD) estimation method.
We prove a central limit theorem for network formation models with strategic interactions and homophilous agents. Since data often consists of observations on a single large network, we consider an asymptotic framework in which the network size diverges. We argue that a modification of "stabilization" conditions from the literature on geometric graphs provides a useful high-level formulation of weak dependence which we utilize to establish an abstract central limit theorem. Using results in branching process theory, we derive interpretable primitive conditions for stabilization. The main conditions restrict the strength of strategic interactions and equilibrium selection mechanism. We discuss practical inference procedures justified by our results.
In this paper we investigate panel regression models with interactive fixed effects. We propose two new estimation methods that are based on minimizing convex objective functions. The first method minimizes the sum of squared residuals with a nuclear (trace) norm regularization. The second method minimizes the nuclear norm of the residuals. We establish the consistency of the two resulting estimators. Those estimators have a very important computational advantage compared to the existing least squares (LS) estimator, in that they are defined as minimizers of a convex objective function. In addition, the nuclear norm penalization helps to resolve a potential identification problem for interactive fixed effect models, in particular when the regressors are low-rank and the number of the factors is unknown. We also show how to construct estimators that are asymptotically equivalent to the least squares (LS) estimator in Bai (2009) and Moon and Weidner (2017) by using our nuclear norm regularized or minimized estimators as initial values for a finite number of LS minimizing iteration steps. This iteration avoids any non-convex minimization, while the original LS estimation problem is generally non-convex, and can have multiple local minima.
We consider a generalized method of moments framework in which a part of the data vector is missing for some units in a completely unrestricted, potentially endogenous way. In this setup, the parameters of interest are usually only partially identified. We characterize the identified set for such parameters using the support function of the convex set of moment predictions consistent with the data. This identified set is sharp, valid for both continuous and discrete data, and straightforward to estimate. We also propose a statistic for testing hypotheses and constructing confidence regions for the true parameter, show that standard nonparametric bootstrap may not be valid, and suggest a fix using the bootstrap for directionally differentiable functionals of Fang and Santos (2019). A set of Monte Carlo simulations demonstrates that both our estimator and the confidence region perform well when samples are moderately large and the data have bounded supports.
We develop a Lagrange Multiplier (LM) test of neglected heterogeneity in dyadic models. The test statistic is derived by modifying Breusch and Pagan (1980)’s test. We establish the asymptotic distribution of the test statistic under the null using a novel martingale construction. We also consider the power of the LM test in generic panel models. Even though the test is motivated by random effects, we show that it has a power for detecting fixed effects as well. Finally, we examine how the estimation noise of the maximum likelihood estimator affects the asymptotic distribution of the test under the null, and show that such a noise may be ignored in large samples.
This article addresses the robust estimation of linear regression models in the presence of potentially endogenous outliers. Through Monte Carlo simulations, we demonstrate that existing methods using L1-regularization on case-specific parameters, including the Huber estimator and the least absolute deviation (LAD) estimator, exhibit significant bias when outliers are endogenous. Motivated by this finding, we investigate L0-regularized estimation methods. We propose systematic heuristic algorithms, notably a local combinatorial search refinement based on the iterative hard-thresholding solution, to solve the combinatorial optimization problem of the L0-regularized estimation efficiently. Our Monte Carlo simulations yield two key results: (i) The local combinatorial search algorithm substantially improves solution quality compared to the initial projection-based hard-thresholding algorithm while offering greater computational efficiency than directly solving the original optimization problem; (ii) The L0-regularized estimator demonstrates superior performance in terms of bias reduction, estimation accuracy, and out-of-sample prediction errors compared to L1-regularized alternatives. In the stock return forecasting application, our method identifies the crisis periods across rolling windows and improves the prediction accuracy over baseline methods. An accompanying R package is provided for practitioners.
This study provides an econometric methodology to test a linear structural relationship among economic variables. We propose the so-called distance-difference (DD) test and show that it has omnibus power against arbitrary nonlinear structural relationships. If the DD-test rejects the linear model hypothesis, a sequential testing procedure assisted by the DD-test can consistently estimate the degree of a polynomial function that arbitrarily approximates the nonlinear structural equation. Using extensive Monte Carlo simulations, we confirm the DD-test’s finite sample properties and compare its performance with the sequential testing procedure assisted by the J-test and moment selection criteria. Finally, through investigation, we empirically illustrate the relationship between the value-added and its production factors using firm-level data from the United States. We demonstrate that the production function has exhibited a factor-biased technological change instead of Hicks-neutral technology presumed by the Cobb–Douglas production function.
Background: The TeloVac study indicated GV1001 did not improve the survival of advanced pancreatic ductal adenocarcinoma (PDAC). However, the cytokine examinations suggested that high serum eotaxin levels may predict responses to GV1001. This Phase III trial assessed the efficacy of GV1001 with gemcitabine/capecitabine for eotaxin-high patients with untreated advanced PDAC.. MethodsPatients recruited from 16 hospitals received gemcitabine (1000 mg/m(2), D 1, 8, and 15)/capecitabine (830 mg/m(2) BID for 21 days) per month either with (GV1001 group) or without (control group) GV1001 (0.56 mg; D 1, 3, and 5, once on week 2-4, 6, then monthly thereafter) at random in a 1:1 ratio. The primary endpoint was overall survival (OS) and secondary end points included time to progression (TTP), objective response rate, and safety. Results: Total 148 patients were randomly assigned to the GV1001 (n = 75) and control groups (n = 73). The GV1001 group showed improved median OS (11.3 vs. 7.5 months, P = 0.021) and TTP (7.3 vs. 4.5 months, P = 0.021) compared to the control group. Grade >3 adverse events were reported in 77.3% and 73.1% in the GV1001 and control groups (P = 0.562), respectively. Conclusions: GV1001 plus gemcitabine/capecitabine improved OS and TTP compared to gemcitabine/capecitabine alone in eotaxin-high patients with advanced PDAC.Clinical trial registrationNCT02854072.
In this article, we consider a survival function estimation method that may be suitable for analyses of clinical trials of cancer treatments whose prognosis is known to be poor such as pancreatic cancer treatment. Typically, these kinds of trials are not double-blind, and patients in the control group may drop out in more significant numbers than in the treatment group if their disease progresses (DP). If disease progression is associated with a higher risk of death, then censoring becomes dependent. To estimate the survival function with dependent censoring, we use copula-graphic estimation, where a parametric copula function is used to model the dependence in the joint survival function of the event and censoring time. In this article, we propose a novel method that one can use in choosing the copula parameter. As an application example, we estimate the survival function of the overall survival time of the KG4/2015 study, the phase 3 clinical trial of the efficacy of GV1001 as a treatment for pancreatic cancer. We provide both statistical and clinical pieces of evidence that support the violation of independent censoring. Applying the estimation method with dependent censoring, we obtain that the estimates of the median survival times are 339 days in the treatment group and 225.5 days in the control group. We also find that the estimated difference of the medians is 113.5 days, and the difference is statistically significant at the one-sided level with size 2.5 % .
We incorporate a version of a spike and slab prior, comprising a pointmass at zero ("spike") and a Normal distribution around zero ("slab") into a dynamic panel data framework to model coefficient heterogeneity. In addition to homogeneity and full heterogeneity, our specification can also capture sparse heterogeneity, that is, there is a core group of units that share common parameters and a set of deviators with idiosyncratic parameters. We fit a model with unobserved components to income data from the Panel Study of Income Dynamics. We find evidence for sparse heterogeneity for balanced panels composed of individuals with long employment histories.
We use a dynamic panel Tobit model with heteroskedasticity to generate forecasts for a large cross-section of short time series of censored observations. Our fully Bayesian approach allows us to flexibly estimate the cross-sectional distribution of heterogeneous coefficients and then implicitly use this distribution as prior to construct Bayes forecasts for the individual time series. In addition to density forecasts, we construct set forecasts that explicitly target the average coverage probability for the cross-section. We present a novel application in which we forecast bank-level loan charge-off rates for small banks.
For an $N \times T$ random matrix $X(\beta )$ with weakly dependent uniformly sub-Gaussian entries $x_{it}(\beta )$ that may depend on a possibly infinite-dimensional parameter $\beta \in \mathbf {B}$ , we obtain a uniform bound on its operator norm of the form $\mathbb {E} \sup _{\beta \in \mathbf {B}} ||X(\beta )|| \leq CK \left (\sqrt {\max (N,T)} + \gamma _2(\mathbf {B},d_{\mathbf {B}})\right )$ , where C is an absolute constant, K controls the tail behavior of (the increments of) $x_{it}(\cdot )$ , and $\gamma _2(\mathbf {B},d_{\mathbf {B}})$ is Talagrand’s functional, a measure of multiscale complexity of the metric space $(\mathbf {B},d_{\mathbf {B}})$ . We illustrate how this result may be used for estimation that seeks to minimize the operator norm of moment conditions as well as for estimation of the maximal number of factors with functional data.
We propose a method of estimating the linear-in-means model of peer effects in which the peer group, defined by a social network, is endogenous in the outcome equation for peer effects. Endogeneity is due to unobservable individual characteristics that influence both link formation in the network and the outcome of interest. We propose two estimators of the peer effect equation that control for the endogeneity of the social connections using a control function approach. We leave the functional form of the control function unspecified and treat it as unknown. To estimate the model, we use a sieve semiparametric approach, and we establish asymptotics of the semiparametric estimator.
In this paper, we investigate seemingly unrelated regression (SUR) models that allow the number of equations (N) to be large, and to be comparable to the number of the observations in each equation (T). It is well known in the literature that the conventional SUR estimator, for example, the generalized least squares (GLS) estimator of Zellner (1962) does not perform well. As the main contribution of the paper, we propose a new feasible GLS estimator called the feasible graphical lasso (FGLasso) estimator. For a feasible implementation of the GLS estimator, we use the graphical lasso estimation of the precision matrix (the inverse of the covariance matrix of the equation system errors) assuming that the underlying unknown precision matrix is sparse. We derive asymptotic theories of the new estimator and investigate its finite sample properties via Monte-Carlo simulations.
We provide an alternative econometrics methodology to estimate a standard heterogeneous income profiles (HIP) model. Our alternative setup allows for the HIP coefficients to be fixed in the sense that they can be arbitrarily correlated with the explanatory variables of the HIP equation and can be treated as parameters to be estimated. As an empirical application, we analyse the extent to which different sources of labour income shocks account for the persistence and variation of composite income shocks based on the HIP model. Our estimation results from the Panel Study of Income Dynamics (PSID) suggest that the random effect assumption of no correlation can lead to biases. We also find that job displacements account for the persistence of income shocks the most, especially for high school educated individuals.
Peter Phillips has had a tremendous impact on econometric theory and practice [...]
We use a decision-theoretic framework to study the problem of forecasting discrete outcomes when the forecaster is unable to discriminate among a set of plausible forecast distributions because of partial identification or concerns about model misspecification or structural breaks. We derive "robust" forecasts which minimize maximum risk or regret over the set of forecast distributions. We show that for a large class of models including semiparametric panel data models for dynamic discrete choice, the robust forecasts depend in a natural way on a small number of convex optimization problems which can be simplified using duality methods. Finally, we derive "efficient robust" forecasts to deal with the problem of first having to estimate the set of forecast distributions and develop a suitable asymptotic efficiency theory. Forecasts obtained by replacing nuisance parameters that characterize the set of forecast distributions with efficient first-stage estimators can be strictly dominated by our efficient robust forecasts.