
In this study, we focus on high-dimensional time-to-event data in the situation of multiple treatment options. Employing the A-learning framework, we formulate A-learning-based estimating equations and introduce a loss function that can be integrated with regularization penalties. When either the baseline covariate model or the propensity score model is correctly specified, we achieve variable selection for key covariates to optimize treatment decision regimes. Extensive simulation experiments, along with an analysis of the ACTG 175 clinical trial involving HIV-infected patients, validate the effectiveness of the proposed methodology.
Gaussian processes (GP) are known for their outstanding smoothing properties and flexibility in modeling complex nonlinear relationships among variables. Recently, they have been utilized in actuarial science to estimate potential insurance claims. However, in insurance practice, claims data often exhibits non-negative support and asymmetry, resulting in significant model errors when the GP model is employed for claim prediction. This paper intends to replace the Gaussian distribution in GP with a log-Gaussian distribution to construct a log-Gaussian process (LGP) model. The properties and training methods of LGP are discussed. Furthermore, the LGP regression model is employed for estimating claims reserve. It is shown that the LGP regression model offers higher prediction accuracy and reliability compared to other models, particularly the GP regression model, making it more suitable for claims reserving.
Conditional Value-at-Risk (CVaR) is a widely used coherent risk measure in the finance and machine learning communities to ensure robustness and fairness. Focusing on CVaR portfolio optimization problems in the expanding global markets, this paper analyzes sparsity-induced portfolio strategies. We first analyze the equivalence of the regularized CVaR minimization problem and its l(0)-constrained counterpart in the norm ball, which suggests the validity of regularizers for sparsity-induced CVaR minimization. Furthermore, we establish the finite-sample statistical estimation error rate for the proposed method in the ultrahigh-dimensional scenario, when the number of assets under management can scale exponentially with the sample size of the historical data. The numerical experiments demonstrate the risk aversion preference and robustness of the proposed portfolio strategy on both synthetic data and the S&P 500 dataset. the to forced mittee foundations, mization hedging ippi of a markets quately they the ceed the U.S.
In longitudinal follow-up studies, recurrent event data with a terminal event are frequently encountered. In certain situations, some subjects drop out of the study for a period of time for various reasons before returning to the study again and this may happen more than once. Disregarding these gaps and treating the data as a regular recurrent event dataset may lead to biased estimations and potentially misleading conclusions. In this paper, we propose a flexible additive-multiplicative rates model for the analysis of recurrent event data with a terminal event and multiple intermittent gaps. This model allows for both additive and multiplicative effects of covariates, while leaving the dependence structure among the recurrent and terminal events unspecified. To infer the parameters of interest in the proposed model, we construct unbiased estimating equations and establish the asymptotic properties of the resulting estimators. Simulation studies demonstrate that considering intermittent gaps yields more accurate estimations in finite samples compared to the naive method. An application to a medical cost study of chronic heart failure is provided for illustration.
Panel data offers valuable insights into both temporal and cross-sectional variations, making it essential for analyzing complex problems. To address dimensionality reduction challenges in high-dimensional varying coefficient panel data models with fixed effects, we propose a novel variable selection method that combines basis function approximation with non-convex group penalties (SCAD/MCP). Using auxiliary regression and orthogonal projection, we isolate fixed effects to prevent their interference in variable selection. Basis function expansion approximates nonlinear coefficients, enabling simultaneous identification of the true model structure and estimation of non-zero coefficients. Under some regularity conditions, our method consistently identifies the true model, with the estimator exhibiting the oracle property. We employ a group descent algorithm to solve the penalized objective function. Simulation studies demonstrate that SCAD and MCP outperform Lasso in variable selection accuracy, reducing false positives and enhancing coefficient estimation. When applied to real data, our method effectively identifies important variables and delivers superior predictive performance.
In the era of precision medicine, developing accurate predictive models for cancer prognosis can guide clinical treatments and improve patient survival. Multi-omics integration can improve the performance and interpretability of prognostic models. However, multi-omics integration models have limitations owing to the heterogeneity of data and the complex regulatory relationships among different platforms. This study proposes a novel Bayesian Cox proportional hazards model with a structural equation model framework (BSEMsurvCox) to integrate three omics platforms. The No U-turn sampling (NUTS) algorithm was used to fit our model. Extensive simulation studies have shown that our model is superior to two non-integrated models including Bayesian Cox and Traditional Cox and an integration model named Block Forest Cox. Furthermore, two real datasets were used to demonstrate the superiority of the proposed model. The study findings indicate that the BSEMsurvCox model provides a higher predictive performance and biological interpretability than non-integrated and other integration models.
Effective risk stratification is essential for providing tailored therapies, improving patient outcomes, and optimizing healthcare resources by identifying sub-populations with similar health risks. However, accurate risk ranking is challenging in the presence of heterogeneous subgroups. Under these instances, subgroup-level information can be leveraged to refine the overall risk ranking. We propose a novel approach that integrates within-subgroup risk ranking percentiles to enhance the overall cohort risk stratification. This method uses both a global model and subgroup-specific models along with optimized weights to improve discriminatory performance across the entire cohort. The proposed method is validated through extensive simulations and applied to a study of end-stage renal disease patients awaiting kidney transplantation.
Length-biased interval-censored data occur frequently many areas such as health science research and many methods for their analyses have been proposed. In particular, some methods have been developed under the mixture cure model. However, most of these use logistic regression model the cured subgroup, which may be limited or reasonable in some cases. Corresponding to this, we propose an estimation procedure under the proportional hazards cure model that makes use of the generalized extreme value regression. A computationally simple and stable EM algorithm through introducing three layers of data augmentation is developed, and the consistency and asymptotic normality of the proposed estimators are established. extensive simulation study is conducted and indicates that the proposed procedure works well in practice. In addition, the method is applied to a set of real data.
A key task in microbiome data analysis is to estimate the dependence pattern among different microbial taxa. However, microbiome datasets from the real world impose a great challenge on standard correlation analysis due to various factors, such as sum-to-one constraint and heavy tails. To handle this challenge, this paper proposes a robust precision matrix estimation by assuming that the log-basis random vector follows a class of continuous elliptical distributions. The proposed Kendall's tau statistics based estimation procedure enjoys several desirable properties. Theoretically, we derive the convergence rate of the proposed estimator under the spectral norm. The sign consistency is also established via a thresholding step. Computationally, the proposed estimator can be computed by linear programming. So we can employ it to large-scale datasets. Empirically, simulation studies and an application to a mouse skin microbiome dataset are conducted to corroborate the superiority of the proposed method.
As a useful semiparametric learning method, varying coefficient models mitigate the "curse of dimensionality" of full nonparametric models while retain interpretability and flexibility in modelling. In certain applications, monotone coefficient functions are needed in the varying coefficient models. In this paper, we propose a monotone estimation procedure for the coefficient functions in the varying coefficient models. The proposed method not only ensures monotone estimates, but also reduces mean squared errors in estimation for such coefficient functions. Furthermore, this method is computationally efficient as it does not require constrained optimization. The asymptotic normality of the monotone coefficient function estimator is established, and numerical studies are conducted to demonstrate the practicality and advantages of the proposed method.
This paper introduces a generalized functional additive model (G-FAM) that accommodates responses generated from various distributions within the exponential family, including normal, binomial, and Poisson distributions. We minimize a penalized negative log-likelihood function within the reproducing kernel Hilbert space (RKHS) framework to estimate the unknown functional coefficients. We further establish the optimal convergence rate of the proposed estimator under mild conditions. The empirical performance of our method is demonstrated through simulation studies and an application to real data. This paper introduces a generalized functional additive model (G-FAM) that accommodates responses generated from various distributions within the exponential family, including normal, binomial, and Poisson distributions. We minimize a penalized negative log-likelihood function within the reproducing kernel Hilbert space (RKHS) framework to estimate the unknown functional coefficients. We further establish the optimal convergence rate of the proposed estimator under mild conditions. The empirical performance of our method is demonstrated through simulation studies and an application to real data.
Adherence in behavioral modification studies is critical for ensuring reliable results, but it is often difficult to monitor or understand factors that may contribute to it. This study aims to investigate what factors may be associated with adherence and whether adherence is associated with the trial's primary outcome, using the Fitbit actimetry data. We used "Improving Reproductive Fitness through Pretreatment with Lifestyle Modification in Obese Women with Unexplained Infertility (FIT-PLESE)" dataset (Clinicaltrials.gov: NCT02432209), including 358 women with a total of 57,496 observations. We defined an objective adherence score as the adherence to the increased exercise and performed correlation and regression analyses to examine the associations between adherence scores, baseline variables, and healthy live births (primary outcome). The overall adherence score was significantly associated with factors like baseline steps, education level, and history of prior conception. Notably, a higher adherence score was also associated with increased odds of a healthy live birth. These results showed that exercise adherence was associated with participants' characteristics as well as the trial's primary outcome. These findings also pointed to a broader focus beyond weight-loss treatment, suggesting that maintaining a high level of exercise adherence may be the key to healthy live births.
We propose a novel transfer learning method for high-dimensional linear regression that integrates transferability detection into a unified framework. By selectively leveraging informative source datasets, the proposed method improves estimation efficiency for the target data. It effectively identifies whether each source dataset is transferable and provides asymptotically confidence intervals. To assess the effectiveness of the transferability detection mechanism, we conduct extensive simulation studies comparing our method with existing approaches. The efficiency and robustness of the proposed framework are further demonstrated through comprehensive simulations and a real-data analysis using the heterogeneous CHARLS dataset.
High-throughput sequencing technology enables a quantitative examination of microbial communities, enhancing the ability to explore connections between the human microbiome and various diseases. Due to the complex nature of microbiome data, which is characterized by high dimensionality, zero inflation, and overdispersion, we propose a zero-inflated negative binomial factor analysis (ZINBFA) model to address these challenges. This model assumes that the sequencing read counts follow a zero-inflated negative binomial (ZINB) distribution, and constructs link functions to analyze the potential low-rank structure of the mean parameter and zero-inflation parameter. To determine the number of latent factors, an information criterion is employed, while the alternating maximum likelihood algorithm is utilized to estimate the unknown parameters within the ZINBFA model. The proposed ZINBFA model is demonstrated to exhibit superior performance and advantages through comprehensive simulation studies and real data applications.
This paper is focused on testing the covariance between high dimensional random vectors. A bootstrap test procedure has been developed, and its theoretical properties have been studied. We show that the asymptotic null distribution is a random variable combining a finite chi-squared-type mixture with a normal approximation, which implies that the normal approximation can be seen as a special case of our theoretical results. We have adopted the power enhancement technique to improve the empirical power of the bootstrap-based test, especially under sparse alternatives. Three propositions have been established to provide an intuitive understanding of the abstract assumptions. The finite-sample performance of the proposed test procedure demonstrates that it better controls empirical size than some existing tests, even when the normal approximation is invalid. Also, it shows higher empirical power, especially under sparse alternatives. Real data analysis is provided to illustrate our proposed test.
Genetic pleiotropy, where a single gene influences multiple phenotypic traits, is critical for understanding genetic functions and disease mechanisms. However, many methods for detecting pleiotropy overlook the issue of missing data, common in biological studies. In this paper, we assume the response is missing at random (MaR), which is commonly used in statistics analysis. The inverse probability weighting (IPW) method is used for parameter estimation and an integrated decision procedure is applied for genetic pleiotropy test. Simulation studies demonstrate the method's efficacy, and applications to real data illustrates its practical utility.
Although subgroup recovery in unsupervised learning for heterogeneous regression data has garnered considerable attention, the dynamics of heterogeneous analysis in the absence of a well-defined subgroup structure are not well understood. This paper investigates the simultaneous estimation of regression effects and subgroup structures within a newly developed fuzzy subgroup framework, where regression effects are variable. Conformal inference is commonly employed for addressing predictive challenges and has demonstrated substantial empirical success. However, it is predominantly effective within homogeneous regression models and often struggles when applied outside this framework. In response, our study introduces an innovative conformal machine learning approach tailored for heterogeneous regression analysis, which employs regression residuals as a nonconformity measure. Key advantages of this method include the absence of an objective function and the elimination of optimization for estimating regression parameters, which facilitates its application to large-scale regression data. Importantly, this approach does not rely on the assumption of distinct subgroup patterns. We provide empirical validation through simulation studies and demonstrate the applicability of our method using datasets on body measurements and house prices. Code is available at: https://github.com/DongYiii/Conformalized-Fuzzy-Subgroup-Analysis.
The stock network is a financial knowledge graph where nodes represent stocks and edges capture their relationships, forming a weighted network. Community detection in such networks is essential for sector division, enabling investors to identify market trends and optimize investment strategies. Traditional approaches typically rely on disjoint community detection, assigning each stock to a single sector, or conventional overlapping community detection, assigning stocks to multiple sectors without quantifying their degree of association with each. To address these limitations, this study introduces the Mixed-SCORE method for identifying overlapping communities in weighted stock networks, termed the Weighted Mixed-SCORE method. Unlike traditional methods, this approach distinguishes between pure nodes (exclusive to one community) and mixed nodes (associated with multiple sectors with varying weights). By leveraging network weights, the method offers a more nuanced understanding of stock relationships and sector structures. Using data from 469 S&P 500 stocks between 2018 and 2022, we demonstrate the effectiveness of the Weighted Mixed-SCORE method in uncovering overlapping community structures. Our analysis reveals the central role of pure nodes within their respective communities and examines how mixed nodes bridge multiple sectors, providing valuable insights for market analysis and investment decision-making. This study not only advances the application of community detection in financial networks but also offers a robust tool for sector analysis and portfolio optimization.
In risk management, expectile is a widely applied tail risk measure as well as Value-at-Risk. It is of great theoretical interest to estimate extreme conditional expectiles, because classic statistical approaches may introduce substantial bias when estimating extreme quantiles or expectiles due to the data sparsity on the tail region. Xu, Hou and Li (2022) introduces an approach for this estimation based on a tail equivalence transition relationship between the quantile and expectile, but without any theoretical study. In this paper, we first develop the theoretical results for the estimation based on the tail equivalence transition. Second, we propose another novel estimation method and study its asymptotic properties. Simulation studies show that both methods based on the tail equivalence transition perform well. Empirical studies applied to S&P 500 index return illustrate that our two estimation methods provide robust risk measurement tools for studying extreme tail risk.
This paper introduces a highly scalable tuning-free algorithm for variable selection in logistic regression using Polya-Gamma data augmentation. The proposed method is both theoretically consistent and robust to potential misspecification of the tuning parameter, achieved through a hierarchical approach. Existing works suitable for high-dimensional settings primarily rely on t-approximation of the logistic density, which is not based on the original likelihood. The proposed method not only builds upon the exact logistic likelihood, offering superior empirical performance, but is also more computationally efficient, particularly in cases involving highly correlated covariates, as demonstrated in a comprehensive simulation study. We apply our method to a gene expression PCR dataset from mice and an RNA-seq dataset from asthma studies in humans. By comparing its performance to existing frequentist and Bayesian methods in variable selection, we demonstrate the competitive predictive capabilities of the Polya-Gamma-based approach. Our results indicate that this method enhances the accuracy of variable selection and improves the robustness of predictions in complex, high-dimensional datasets.