The density ratio model (DRM) is a semiparametric model that relates the distributions from multiple samples to a nonparametrically defined reference distribution via exponential tilting, with finite-dimensional parameters governing their differences in shape. When multiple types of partially observed (censored/truncated) failure time data are collected in an observational study, the DRM can be utilized to conduct a single unified analysis of the combined data. In this paper, we extend the methodology for censored length-biased/truncated data to the DRM framework and formulate the inference using empirical likelihood. We develop an EM algorithm to compute the DRM-based maximum empirical likelihood estimators of the model parameters and survival function, and assess its performance through extensive simulations under correct model specification, overspecification, and misspecification, across a range of failure-time distributions and censoring proportions. We also illustrate the efficacy of our method by analyzing the duration of time spent from admission to discharge in a Montreal-area hospital in Canada. The R code that implements our method is available on GitHub at \href{https://github.com/gozhang/DRM-combined-survival}{DRM-combined-survival}.
In causal inference, an important problem is to quantify the effects of interventions or treatments. Many studies focus on estimating the mean causal effects; however, these estimands may offer limited insight since two distributions can share the same mean yet exhibit significant differences. Examining the causal effects from a distributional perspective provides a more thorough understanding. In this paper, we employ a semiparametric density ratio model (DRM) to characterize the counterfactual distributions, introducing a framework that assumes a latent structure shared by these distributions. Our model offers flexibility by avoiding strict parametric assumptions on the counterfactual distributions. Specifically, the DRM incorporates a nonparametric component that can be estimated through the method of empirical likelihood (EL), using the data from all the groups stemming from multiple interventions. Consequently, the EL-DRM framework enables inference of the counterfactual distribution functions and their functionals, facilitating direct and transparent causal inference from a distributional perspective. Numerical studies on both synthetic and real-world data validate the effectiveness of our approach.
Gaussian mixtures are widely used for approximating density functions in various applications such as density estimation, belief propagation, and Bayesian filtering. These applications often utilize Gaussian mixtures as initial approximations that are updated recursively. A key challenge in these recursive processes stems from the exponential increase in the mixture's order, resulting in intractable inference. To overcome the difficulty, the Gaussian mixture reduction (GMR), which approximates a high order Gaussian mixture by one with a lower order, can be used. Although existing clustering-based methods are known for their satisfactory performance and computational efficiency, their convergence properties and optimal targets remain unknown. In this paper, we propose a novel optimization-based GMR method based on composite transportation divergence (CTD). We develop a majorization-minimization algorithm for computing the reduced mixture and establish its theoretical convergence under general conditions. Furthermore, we demonstrate that many existing clustering-based methods are special cases of ours, effectively bridging the gap between optimization-based and clustering-based techniques. Our unified framework empowers users to select the most appropriate cost function in CTD to achieve superior performance in their specific applications. Through extensive empirical experiments, we demonstrate the efficiency and effectiveness of our proposed method, showcasing its potential in various domains.
In many statistical and econometric applications, we gather individual samples from various interconnected populations that undeniably exhibit common latent structures. Utilizing a model that incorporates these latent structures for such data enhances the efficiency of inferences. Recently, many researchers have been adopting the semiparametric density ratio model (DRM) to address the presence of latent structures. The DRM enables estimation of each population distribution using pooled data, resulting in statistically more efficient estimations in contrast to nonparametric methods that analyze each sample in isolation. In this article, we investigate the limit of the efficiency improvement attainable through the DRM. We focus on situations where one population's sample size significantly exceeds those of the other populations. In such scenarios, we demonstrate that the DRM-based inferences for populations with smaller sample sizes achieve the highest attainable asymptotic efficiency as if a parametric model is assumed. The estimands we consider include the model parameters, distribution functions, and quantiles. We use simulation experiments to support the theoretical findings with a specific focus on quantile estimation. Additionally, we provide an analysis of real revenue data from U.S. collegiate sports to illustrate the efficacy of our contribution.
In many applications, we collect samples from multiple interconnected populations. These population distributions share some latent structure, so it is advantageous to jointly analyze the samples to make efficient inferences on the multiple distributions and their functionals. One effective way to connect the distributions is the density ratio model (DRM). A key ingredient of the DRM is that the log density ratios are linear combinations of prespecified functions; the vector formed by these functions is called the basis function. The benefit of DRM relies on correctly specifying the basis function to a large degree. In applications, the user may not have a complete knowledge to enable a suitable choice of the basis function, and many discussions have been devoted to this topic. In this article, we consider the still open problem of a data-adaptive choice of the basis function that can alleviate the risk of severe model misspecification. We propose a data-adaptive approach to the choice of basis function based on functional principal component analysis. Under some conditions, we show that this approach leads to consistent basis function estimation. Our simulation results show that the proposed adaptive choice can achieve an efficiency gain. We use a real-data example from economics to demonstrate the efficiency gain and the ease of our approach.
Population quantiles are important parameters in many applications. Enthusiasm for the development of effective statistical inference procedures for quantiles and their functions has been high for the past decade. In this article, we study inference methods for quantiles when multiple samples from linked populations are available. The research problems we consider have a wide range of applications. For example, to study the evolution of the economic status of a country, economists monitor changes in the quantiles of annual household incomes, based on multiple survey datasets collected annually. Even with multiple samples, a routine approach would estimate the quantiles of different populations separately. Such approaches ignore the fact that these populations are linked and share some intrinsic latent structure. Recently, many researchers have advocated the use of the density ratio model (DRM) to account for this latent structure and have developed more efficient procedures based on pooled data. The nonparametric empirical likelihood (EL) is subsequently employed. Interestingly, there has been no discussion in this context of the EL-based likelihood ratio test (ELRT) for population quantiles. We explore the use of the ELRT for hypotheses concerning quantiles and confidence regions under the DRM. We show that the ELRT statistic has a chi-square limiting distribution under the null hypothesis. Simulation experiments show that the chi-square distributions approximate the finite-sample distributions well and lead to accurate tests and confidence regions. The DRM helps to improve statistical efficiency. We also give a real-data example to illustrate the efficiency of the proposed method.