Nonresponse in probability sampling presents a long-standing challenge in survey sampling, often necessitating simultaneous adjustments to address sampling and selection biases. We develop a statistical framework that explicitly models sampling weights as random variables and establish the semiparametric efficiency bound for the parameter of interest under nonresponse. This study investigates strategies for eliminating bias and effectively utilizing available information, extending beyond nonresponse issues to data integration with external summary statistics. The proposed estimators are characterized by their efficiency and double robustness. However, realizing full efficiency hinges on the accurate specification of underlying models. To enhance robustness against potential model misspecification, we expand double robustness to multiple robustness through a novel two-step empirical likelihood approach. A numerical study evaluates the finite-sample performance of our methods. Additionally, we apply these methods to a dataset from the National Health and Nutrition Examination Survey, effectively integrating summary statistics from the National Health Interview Survey.
更多
查看译文
关键词
data integration,empirical likelihood,missing data,semiparametric efficiency,survey sampling