We consider the estimation of linear functionals of the mixing distribution in a nonparametric empirical Bayes framework. Our main interest is in situations in which the mixing distribution is only partially identifiable, as may arise in complex sampling situations with nonresponse. We argue that estimating the functional by applying it to the semiparametric maximum likelihood estimator of the mixing distribution is an efficient tool, even when the maximum likelihood estimator is not unique.
In non-randomized treatment allocation models, treatment is assigned to a unit based on a score, e.g., scholarship is allocated based on the score obtained in a merit test, while antihypertensive treatments are allocated based on blood pressure level. In this paper, we present a new model coined SCENTS: Score Explained Non-Randomized Treatment Systems that utilizes the dependency of the score on the explanatory variables to permit efficient estimation. We derive an estimator of the treatment effect which is n consistent, asymptotically normal, and achieves semiparametric efficiency under normal errors. The analysis is extended to ultra-high dimensional vectors of covariates, where a n consistent and asymptotically normal debiased estimator is proposed. We analyze two real data sets via our method and compare our results with those obtained by using previous approaches like regression discontinuity design. Some possible extensions are discussed.
In many practical situations, randomly assigning treatments to subjects is uncommon due to feasibility constraints. For example, economic aid programs and merit-based scholarships are often restricted to those meeting specific income or exam score thresholds. In these scenarios, traditional approaches to estimating treatment effects typically focus solely on observations near the cutoff point, thereby excluding a significant portion of the sample and potentially leading to information loss. Moreover, these methods generally achieve a non-parametric convergence rate. While some approaches, e.g., Mukherjee et al. (2021), attempt to tackle these issues, they commonly assume that treatment effects are constant across individuals, an assumption that is often unrealistic in practice. In this study, we propose a differencing and matching-based estimator of the average treatment effect on the treated (ATT) in the presence of heterogeneous treatment effects, utilizing all available observations. We establish the asymptotic normality of our estimator and illustrate its effectiveness through various synthetic and real data analyses. Additionally, we demonstrate that our method yields non-parametric estimates of the conditional average treatment effect (CATE) and individual treatment effect (ITE) as a byproduct.
Modern large language model (LLM) alignment techniques rely on human feedback, but it is unclear whether these techniques fundamentally limit the capabilities of aligned LLMs. In particular, it is unknown if it is possible to align (stronger) LLMs with superhuman capabilities with (weaker) human feedback without degrading their capabilities. This is an instance of the weak-to-strong generalization problem: using feedback from a weaker (less capable) model to train a stronger (more capable) model. We prove that weak-to-strong generalization is possible by eliciting latent knowledge from pre-trained LLMs. In particular, we cast the weak-to-strong generalization problem as a transfer learning problem in which we wish to transfer a latent concept prior from a weak model to a strong pre-trained model. We prove that a naive fine-tuning approach suffers from fundamental limitations, but an alternative refinement-based approach suggested by the problem structure provably overcomes the limitations of fine-tuning. Finally, we demonstrate the practical applicability of the refinement approach in multiple LLM alignment tasks.
Standard techniques for aligning large language models (LLMs) utilize human-produced data, which could limit the capability of any aligned LLM to human level. Label refinement and weak training have emerged as promising strategies to address this superalignment problem. In this work, we adopt probabilistic assumptions commonly used to study label refinement and analyze whether refinement can be outperformed by alternative approaches, including computationally intractable oracle methods. We show that both weak training and label refinement suffer from irreducible error, leaving a performance gap between label refinement and the oracle. These results motivate future research into developing alternative methods for weak to strong generalization that synthesize the practicality of label refinement or weak training and the optimality of the oracle procedure.
We consider the estimation of the mixing distribution of a normal distribution where both the shift and scale are unobserved random variables. We argue that in general, the model is not identifiable. We give an elegant non-constructive proof that the model is identifiable if the shift parameter is bounded by a known value. However, we argue that the generalized maximum likelihood estimator is inconsistent even if the shift parameter is bounded and the shift and scale parameters are independent. The mixing distribution, however, is identifiable if we have more than one observation per any realization of the latent shift and scale.
We introduce Sequential Probability Ratio Bisection (SPRB), a novel stochastic approximation algorithm that adapts to the local behavior of the (regression) function of interest around its root. We establish theoretical guarantees for SPRB's asymptotic performance, showing that it achieves the optimal convergence rate and minimal asymptotic variance even when the target function's derivative at the root is small (at most half the step size), a regime where the classical Robbins-Monro procedure typically suffers reduced convergence rates. Further, we show that if the regression function is discontinuous at the root, Robbins-Monro converges at a rate of 1/n whilst SPRB attains exponential convergence. If the regression function has vanishing first-order derivative, SPRB attains a faster rate of convergence compared to stochastic approximation. As part of our analysis, we derive a nonasymptotic bound on the expected sample size and establish a generalized Central Limit Theorem under random stopping times. Remarkably, SPRB automatically provides nonasymptotic time-uniform confidence sequences that do not explicitly require knowledge of the convergence rate. We demonstrate the practical effectiveness of SPRB through simulation results.
We argue that the Bayesian paradigm, of a prior which represents the beliefs of the statistician before observing the data, is not feasible in ultra-high-dimensional models. We claim that natural priors that represent the a priori beliefs fail in unpredictable ways under values of the parameters that cannot be honestly ignored. We do not claim that the frequentist estimators we present cannot be mimicked by Bayesian procedures, but that these Bayesian procedures do not represent beliefs. They were created with the frequentist analysis in mind, and in most cases, they cannot represent a consistent set of beliefs about the parameters (for example, since they depend on the loss function, the particular functional of interest, and not only on the a priori knowledge, different priors should be used for different analyses of the same data set). In a way, these are frequentist procedures using a Bayesian technique. The paper presents different examples where the subjective point of view fails. It is argued that the arguments based on Wald's and Savage's seminal works are not relevant to the validity of the subjective Bayesian paradigm. The discussion tries to deal with the fundamentals, but the argument is based on a firm mathematical proofs.
Motivated by equilibrium models of labor markets, we develop a formulation of causal strategic classification in which strategic agents can directly manipulate their outcomes. As an application, we consider employers that seek to anticipate the strategic response of a labor force when developing a hiring policy. We show theoretically that employers with performatively optimal hiring policies improve employer reward, labor force skill level, and labor force equity (compared to employers that do not anticipate the strategic labor force response) in the classic Coate-Loury labor market model. Empirically, we show that these desirable properties of performative hiring policies do generalize to our own formulation of a general equilibrium labor market. On the other hand, we also observe that the benefits of performatively optimal hiring policies are brittle in some aspects. We demonstrate that in our formulation a performative employer both harms workers by reducing their aggregate welfare and fails to prevent discrimination when more sophisticated wage and cost structures are introduced.
Reported coefficients by relative position for all cancer (with 95% confidence interval)
Supplementary Figure 1 shows a map of U.S. counties contained in our analysis of composite cancer incidence according to their respective time zones (n = 2853).
Supplementary Figure 4 shows the relative weights of 607 counties studied by Gu et al. (2017)
We discuss the asymptotics of the nonparametric maximum likelihood estimator (NPMLE) in the normal mixture model. We then prove the convergence rate of the NPMLE decision in the empirical Bayes problem with normal observations. We point to (and use) the connection between the NPMLE decision and Stein unbiased risk estimator (SURE). Next, we prove that the same solution is optimal in the compound decision problem where the unobserved parameters are not assumed to be random. Similar results are usually claimed using an oracle-based argument. However, we contend that the standard oracle argument is not valid. It was only partially proved that it can be fixed, and the existing proofs of these partial results are tedious. Our approach, on the other hand, is straightforward and short.
Output of natural splines for most prevalent cancers. Figure shows the output of natural splines conducted on incidence by longitude and time zone for four of the most prevalent cancers, with 95% bootstrap confidence band. A, Liver and bile duct cancer (without MST, n = 1,063). B, Lung and bronchus cancer (n = 2,441). C, Colon and rectum cancer (n = 1,955). D, Pancreas cancer (n = 1,191).
Abstract The debate over daylight saving time (DST) has surged, with interests in the effects of sunlight exposure on health. Prior studies simulated DST and standard time conditions by analyzing different locations within time zones and neighboring areas across time zone borders. We analyzed cancer incidence rates from various longitudinal positions within time zones and at time zone borders in the contiguous United States. Using data from State Cancer Profiles (2016–2020), we analyzed total cancer of 19 types and specific rates for eight cancers, adjusted for age and includes all demographics. log-linear regression is used to replicate a previous study, and spatial regression models are employed to explore discontinuities at borders. Cancer rate differences lack statistical significance within time zones and near borders for total cancer and most individual cancers. Exceptions included breast, prostate, and liver and bile duct cancers, which exhibited significant relationships with relative position at the 95% significance level. Breast and liver and bile duct cancers saw decreases, while prostate cancer incidence increased from west to east within time zones. Relative position does not have a significant impact on cancer incidence, hence cancer development in general. Isolated exceptions may warrant further investigation as more data become available. Our findings challenge prior research, revealing numerous inconsistencies. These disparities urge a reconsideration of the potential disparities in human health associated with DST and standard time. They offer insights contribute to the ongoing discussion surrounding the retention or abandonment of DST. Significance: In this article, we investigate the relation between the epidemiology of cancer incidence in the United States and time zone–related longitudinal positions. Our results differ from previous research, which were based on a subset of our data, and show that the time zone effect on cancer incidence rate is not significant. Our research provides implications on the implementation of DST by suggesting that there is no cancer-risk associated reason to prefer one time over the other. Our study also uses regression discontinuity design using natural splines, a more advanced statistical method, to increase robustness of our result. Our findings challenge prior research, revealing numerous inconsistencies. These disparities urge a reconsideration of the potential disparities in human health associated with DST and standard time. They offer insights contribute to the ongoing discussion surrounding the retention or abandonment of DST.
Output of natural splines for total cancers. Figure shows the output of natural splines conducted on incidence by longitude and time zone for total cancer, with 95% bootstrap confidence band (n = 2,853). Pointwise Confidence Interval (CI) (A), uniform CI (B) restricted to counties in overlapping regions.
In many prediction problems, the predictive model affects the distribution of the prediction target. This phenomenon is known as performativity and is often caused by the behavior of individuals with vested interests in the outcome of the predictive model. Although performativity is generally problematic because it manifests as distribution shifts, we develop algorithmic fairness practices that leverage performativity to achieve stronger group fairness guarantees in social classification problems (compared to what is achievable in non-performative settings). In particular, we leverage the policymaker’s ability to steer the population to remedy inequities in the long term. A crucial benefit of this approach is that it is possible to resolve the incompatibilities between conflicting group fairness definitions.
Output of linear approximation for total cancers. Figure shows linear approximation results of total cancer incidence rate by relative position with 95% bootstrap confidence band with 607 counties (A) and with 2,853 counties (B).
The debate over daylight saving time (DST) has surged, with interests in the effects of sunlight exposure on health. Prior studies simulated DST and standard time conditions by analyzing different locations within time zones and neighboring areas across time zone borders. We analyzed cancer incidence rates from various longitudinal positions within time zones and at time zone borders in the contiguous United States. Using data from State Cancer Profiles (2016-2020), we analyzed total cancer of 19 types and specific rates for eight cancers, adjusted for age and includes all demographics. log-linear regression is used to replicate a previous study, and spatial regression models are employed to explore discontinuities at borders. Cancer rate differences lack statistical significance within time zones and near borders for total cancer and most individual cancers. Exceptions included breast, prostate, and liver and bile duct cancers, which exhibited significant relationships with relative position at the 95% significance level. Breast and liver and bile duct cancers saw decreases, while prostate cancer incidence increased from west to east within time zones. Relative position does not have a significant impact on cancer incidence, hence cancer development in general. Isolated exceptions may warrant further investigation as more data become available. Our findings challenge prior research, revealing numerous inconsistencies. These disparities urge a reconsideration of the potential disparities in human health associated with DST and standard time. They offer insights contribute to the ongoing discussion surrounding the retention or abandonment of DST. SIGNIFICANCE:In this article, we investigate the relation between the epidemiology of cancer incidence in the United States and time zone-related longitudinal positions. Our results differ from previous research, which were based on a subset of our data, and show that the time zone effect on cancer incidence rate is not significant. Our research provides implications on the implementation of DST by suggesting that there is no cancer-risk associated reason to prefer one time over the other. Our study also uses regression discontinuity design using natural splines, a more advanced statistical method, to increase robustness of our result. Our findings challenge prior research, revealing numerous inconsistencies. These disparities urge a reconsideration of the potential disparities in human health associated with DST and standard time. They offer insights contribute to the ongoing discussion surrounding the retention or abandonment of DST.