
This paper studies the performance of difference-in-differences (DID) estimators when outcome data for untreated units are entirely unobserved and pseudo-controls are constructed via matching with external donors. This study aims to compare the finite-sample performance of several matching-based DID estimators. In doing so, we additionally formulate inverse probability weighting (IPW) and doubly robust (DR) versions within the matching framework, extending existing approaches beyond the commonly used two-way fixed effects (TWFE) and regression adjustment (REG) estimators. Monte Carlo simulations compare the four estimators under varying covariate specifications, matching quality, and violations of the parallel trends assumption. Results show that well-specified matching improves estimation accuracy and robustness. The reliability of DID estimates in treated-only contexts depends critically on the quality of matching.
Breast cancer is the most diagnosed cancer and the second most common cause of morbidity and mortality among women in the United States. Despite the substantial progress in reducing breast cancer mortality in the US over the past decades, disparities are still exit, especially among Black women. Continued surveillance and the study of breast cancer are essential to monitor progress, evaluate policies, and inform cancer control strategies. However, a complete and clean data set is sometimes unavailable for analysis despite every effort, posing challenges in the study and negatively affecting decision making in the related areas. The breast cancer mortality rates for the US females by race are obtained from 1990 to 2023 from National Cancer Institute (NCI). In this study, we implement functional data analysis approach to estimate the missing values. The method can estimate the values even when one or few observations are available with high degrees of confidence and the residulas are statistically significant at 0.05 level. The model estimates a continuous decline in the breast cancer mortality rates for Asian Pacific Islander women, while the rates remain stable for black and American Indian/Alaska Native women.
This study proposes a two-sample test statistic for evaluating the homogeneity of survival functions in the presence of a non-susceptible fraction under cross-sectional survival data composed of a single examination time and the event status. A linear rank statistic is proposed to simultaneously assess the homogeneity of both the incidence proportion and the latency distribution. Estimation of the null distribution is carried out through the EM algorithm combined with an adjusted pooled-adjacent-violators algorithm within a mixture cure modeling framework. The performance of the proposed procedures is investigated through extensive simulation studies and an IOL(Intraocular lenses) calcification data set is analyzed as a real data example.
The kernel estimator of density was studied by Rosenblatt (1956) and Parzen (1962), followed by numerous subsequent studies. Most of these studies focus on cases where the density is continuous. Huh (2002) studied the phenomenon that when the density or its derivatives have a discontinuity point, the estimation accuracy of the kernel estimator decreases near the discontinuity point. Otsu, Xu, and Matsushita (2013) and Mynbaev, Martins-Filho, and Henderson (2022) developed methods for inferring discontinuities in density estimates when the location of the discontinuity is known. Loader (1996a) estimated the density using a local polynomial estimator based on a log-likelihood function weighted by a kernel function. Based on Loader's estimation method, this study aims to propose an approach for estimating a discontinuity in density or its derivatives. We will conduct numerical experiments on the proposed estimators using simulation studies.
This paper studies robust learning methods for deep neural networks in the presence of outliers. While conventional training based on mean squared error (MSE) is optimal under normality assumptions, it is highly sensitive to anomalous observations commonly encountered in real-world data. To address this limitation, we adopt the minimum density power divergence framework, which enables a flexible trade-off between robustness and statistical efficiency through a tuning parameter. This paper extends the framework to univariate time series settings and shows that the resulting loss function down-weight the contribution of observations with large residuals to the gradient of model parameters during training. In addition, we integrate an outlier detection procedure based on standardized residuals and tail probability estimation. A data-driven strategy for selecting the tuning parameter is also provided. Simulation studies demonstrate the effectiveness of the proposed method in achieving robust estimation and reliable outlier detection.
This research introduces an advanced image classification model for the early detection and precise identification of diseases and pests affecting tomato and rose plants. The study evaluates three CNN-based architectures: a custom baseline CNN specifically developed for this investigation, and fine-tuned versions of two established pretrained models, ResNet18 and GoogLeNet. Plant image datasets sourced from AIHub provided standardized, comprehensive training and evaluation data. Performance was rigorously assessed using accuracy, macro F1-score, and weighted F1-score, offering a robust evaluation under class-imbalanced conditions typical in agricultural contexts. GoogLeNet achieved the best performance on the tomato dataset (accuracy: 0.9941, macro F1-score: 0.9615, weighted F1-score: 0.9941), while ResNet18 excelled on the rose dataset (accuracy: 0.9966, macro F1-score: 0.9726, weighted F1-score: 0.9966). The findings provide valuable guidance on selecting effective deep learning models for early detection and classification of agricultural diseases and pests, laying a foundation for intelligent diagnostic systems in smart farming environments.
Food security is a crucial indicator of a country's development, influenced by various social, economic, and health factors. This study aims to analyze the factors affecting the food security index (FSI) in Indonesia by considering spatial variations using eigenvector spatial filtering (ESF) approach. This method provides more stable parameter estimates compared to global regression and geographically weighted regression (GWR), which often suffer from multicollinearity and autocorrelation issues. The study utilizes data from 514 districts/cities in Indonesia with eight explanatory variables, including life expectancy, poverty rate, and population growth. The results indicate that the random effect spatial filtering varying coefficient (RE ESF-VC) model achieves a higher goodness-of-fit compared to the GWR, and random effect eigenvector spatial filtering (RE ESF) models, with an adjusted R2 of 0.801. The model identifies that the key factors influencing FSI vary spatially, with the Percentage of Poor Population being the most locally significant factor. These findings provide valuable insights for local governments in designing more effective and region-specific food security policies.
This study aims to model and quantify long-term care demands for dementia patients by analyzing their health state transitions and survival durations. Utilizing data from Korean Long-Term Care Insurance (K-LTCI) system provided by the National Health Insurance Service, we applied a multi-state modeling framework to examine the transition dynamics among care levels for individuals diagnosed with dementia. Transition probability matrices and state-specific life expectancies were estimated by sex, age group, and care level. The results indicate that individuals in relatively mild conditions exhibit a strong tendency to remain in the same state, while those in more severe states show lower stability and a markedly higher probability of transitioning to death. Furthermore, even within the same age and care level, substantial differences in life expectancy were observed by sex, with female beneficiaries consistently exhibiting longer survival durations. The estimated transition probabilities and life expectancy metrics derived from this study have practical implications for rate setting of dementia-specific insurance products, forecasting care demands for public LTCI schemes, and the design of personalized care strategies. Future research may extend this framework by integrating both care utilization and cost components, as well as differentiating between types of care (In-Home vs. Facility), thereby enabling the development of more refined predictive models.
Accurate estimation under privacy constraints is a critical challenge in modern data analysis, particularly when responses exhibit complex dependence structures and heavy-tailed noise. In this study, we investigate stratified randomized response mechanisms combined with differential privacy (DP), incorporating copula-based modeling to capture dependence between multiple noise components. We consider Laplace and Gaussian noise distributions and examine the effects of copula family, dependence parameter (theta), privacy level (E), and stratumspecific randomization probabilities on bias and variance. Through extensive simulations, we demonstrate that Laplace noise amplifies variability and is highly sensitive to copula dependence, while Gaussian noise provides more stable estimates. Clayton and Gaussian copulas offer robust performance under moderate dependence, whereas Gumbel copulas exhibit extreme bias and variance under high dependence. Real data analysis using the mtcars dataset corroborates these findings, showing that copula-based modeling improves estimation stability compared to ignoring dependence. Our results provide practical guidance for selecting copula families, noise types, and privacy parameters, highlighting the trade-offs between privacy and statistical accuracy in privacypreserving data collection.
Numerous methods have been studied for appropriately handling nonresponse, and recent attention has focused on missing not at random (MNAR) nonresponse, where the response probability depends on the variable of interest. Lee and Shin (2022) and Lee and Shin (2024) proposed imputation methods to appropriately handle such non-ignorable nonresponse. In both studies, bias derived from known response probability models was estimated and removed from the imputation estimators, thereby improving the accuracy of the estimation. However, the application of known response probability models is often limited in practice. This study proposes a bias-corrected imputation method that estimates theoretical bias under an arbitrary response probability model and an arbitrary outcome model, and applies this correction to the imputation estimator. Furthermore, the validity of the proposed method is verified through simulation studies.
In the modern quality control segment, the skip-lot sampling plan is still significant among all others plans due to rising production volumes and the demand for cost-effective inspection methods that will yield high-quality outputs. Unlike other sampling plans, while inspecting a submitted lot, a skip-lot plan is economically advantageous and ensures high quality. The skip-lot sampling plan utilizes single sampling plan (SSP) or double sampling plan (DSP) as the reference plan during both normal and skipping inspections. However, using these plans as the reference can lead to bias, favouring either the producer or the consumer. In this paper, a novel approach is illustrated where the skip-lot sampling plan of type 3 is having the provision of two different reference plans in the normal and skipping phases. The proposed plan is termed as the Multi-Reference Skip-lot Sampling Plan of type 3 (MR-SkSP-3). The plan is then compared with the help of performance measures such as operational characteristic (OC) function and average sample number (ASN). The comparison is done between the proposed plan and existing skip-lot sampling plans which use single sampling plan or double sampling plan as reference plan in both inspection phases. The comparison is made based on performance measures with graphical and tabulated illustrations. The comparative analysis proves that the proposed plan successfully balances the satisfaction of both producers and consumers. By leveraging the strengths of conventional skip-lot sampling plans that use single reference plans, it achieves superior performance.
This study proposes a spatial modeling framework for analyzing traffic accident risks using six years of traffic data (2018-2023) from Cheongju City, South Korea. To address the limitations of traditional statistical approaches that fail to capture spatial continuity and neighborhood interactions, we employ Graph Trend Filtering (GTF), a regularization-based technique that smooths data across a graph structure reflecting subdistrict-level adjacency. We further enhance the model by introducing edge-adaptive penalty weights based on observed accident differentials, enabling locally adaptive estimation. The proposed method effectively identifies spatially heterogeneous risk patterns and discontinuities, revealing high-risk zones in both urban cores and peripheral industrial areas. Visualization and clustering of smoothed accident risks offer actionable insights into regional disparities in traffic safety. This work contributes a robust, data-driven methodology for localized policy design and provides empirical evidence to support targeted accident prevention strategies at the subdistrict level.
In this study, we are motivated to study the asymptotic power of Kendall's tau within the framework of a generalized partially linear regression model, given by Y = b1X1 + & centerdot;& centerdot; & centerdot;+ bpXp + phi(V1, ... , Vq) + e where (X1, ... , Xp) represents parametric regressors p, (V1, ... , Vq) are nonparametric regressors q and e denotes the error term. The model incorporates p unknown parameters b1,... , bp and an unknown smooth function phi(& centerdot;, ... , & centerdot;). Our main objective is to test for dependence between the combined set of regressors S = (X1, ... , Xp, V1, ... , Vq)' and the error term e. Under the null hypothesis of independence, the general-order differences between observed responses and their estimated counterparts remain independent. To construct an appropriate test, we introduce statistics based on V-statistics, derived from bivariate observations formed using these differences. These statistics act as empirical analogs of Kendall's tau, as originally proposed in Kendall (1938). Additionally, we consider the test statistics based on the nonparametric measures tau & lowast; and dCov in analogous manner to carry out a comparative power analysis under a sequence of local alternatives. By specifying different conditional distributions for the error term e, we assess the asymptotic power through the limiting distribution of a non-degenerate V-statistic. For a real dataset, we compare their effectiveness by evaluating p-values and finite sample powers.
Survival analysis is widely used to model event times and plays a crucial role in various fields such as medicine, life sciences, and engineering. However, conventional survival models often focus on single-event scenarios, making them inadequate when multiple events exist. To address this limitation, competing risks models have been introduced and have gained increasing attention in recent studies. Despite this progress, research that simultaneously considers cluster effects, model misspecification, and informative censoring in competing risks settings remains limited. In this study, we compare various survival models, including the Fine and Gray (FG) model, Cause-Specific Hazard (CSH) model, Event-Free Survival (EFS) model, Random Survival Forest (RSF), Cox regression with frailty (CF), Mixed Effect Cox model (CoxME), Competing Risks Regression for Stratified and Clustered Data (CRRSC), penalized regression models (LASSO, MCP, SCAD), boosting methods (mboost, CoxBoost), and Direct Binomial Regression (DBR). Through simulation studies, we evaluate model performance under varying cluster effects, covariate structures, and informative censoring criteria. Model performance is assessed by evaluating discrimination, prediction error, and calibration metrics. Based on the results, we analyze the performance of each model under different data characteristics and provide guidelines for selecting the most appropriate model in competing risks analysis under given circumstance.
This paper considers testing the mixture hypothesis of Poisson regression models using the likelihood ratio (LR) test. The main motivation of the mixture hypothesis is the unobserved heterogeneity. The null hypothesis of interest is that there is no unobserved heterogeneity in the data. Due to the nonstandard conditions described in the text, the LR test does not weakly converge to the standard chi-squared random variable under the null. We derive its limiting distribution by assuming popular Poisson regression models in terms of a Gaussian process. Furthermore, we discuss how to obtain the asymptotic critical values of the LR test consistently. Finally, we implement various Monte Carlo experiments and compare its power with the specification test proposed by Lee (1986). The simulation shows that the LR test is more powerful than the specification test when its small sample size distortion is adjusted.
This paper proposed a Bayesian hurdle discrete Weibull regression model for clustered count data. Modeling based on discrete Weibull model can handle various dispersion type which is an attractive feature. Also in the presence of many zeros we focus on hurdle model instead of zero inflated model which is also widely used for excess zero count data. We use a Bayesian mixed effect approach for analyzing such data. MCMC Gibbs sampling algorithm was used in order to obtain the estimators within a Bayesian mixed regression model. Furthermore, the proposed model is applied to analyze the 2020 Korean medical panel (KMP) annual data to assess the factors influencing the number of emergency services used per year in 17 different regions. We compared the results with our proposed model with those of hurdle Poisson and hurdle negative binomial regression models. The result shows that using hurdle discrete Weibull regression model gives the best fit. We also held a simulation study to investigate the performance and the robustness of hurdle discrete Weibull regression model. It was shown that our proposed method gives accurate results and reveals that it can flexibly model clustered count data with various dispersion and data with many zeros.
Interest rate-linked insurance contracts with an embedded Guaranteed Minimum Surrender Benefit (GMSB) option are exposed to market risks, such as interest rate fluctuations. To prevent significant losses caused by policyholders exercising their surrender options during economic downturns and low-interest-rate environments, insurance companies must hold available capital exceeding the required capital. To calculate this required capital for GMSB, which generally does not have a closed-form solution, a scenario-based approach can be used. This involves valuing the GMSB payout under different real-world outer scenarios. A nested stochastic model refers to a framework that values GMSB using inner scenarios embedded within outer scenarios. This study applies the nested stochastic modeling methodology to the calculation of required capital for GMSB guarantees, empirically analyzing the probability distribution of future GMSB values. Additionally, the study evaluates GMSB valuations under various conditions. The analysis reveals that GMSB values are negatively skewed and inversely correlated with interest rates. As the time horizon lengthens, both guarantee costs and required capital increase. Furthermore, as the policyholder's entry age increases, uncertainty rises, leading to a higher required capital ratio.
In this paper, a single-server Markovian queuing model under multiple-vacation policy is analyzed, with a focus on server state-dependent customer impatience manifested through balking and reneging. Observations indicate that customers arriving to find long queues or extended waiting times may exhibit impatience through balking or reneging during both vacation and busy periods of the server, respectively. Notably, customer impatience intensifies when the server is in vacation state. During this state, implementing specific retention mechanisms can effectively reduce customer reneging. This research derives closed-form expressions for traditional and newly designed performance measures. In addition, sensitivity analysis is conducted to examine the impact of the system parameters on different performance metrics. This investigation contributes to understanding the customer behavior in queuing systems with intermittently available services.
Ensemble learning methods such as stacking models have shown significant improvements in predictive performance by combining multiple heterogeneous base learners. However, the complexity of the layered architecture and the diversity of base learners in stacking models often come at the cost of interpretability. Shapley Additive exPlanations (SHAP) provide an axiomatic framework, derived from cooperative game theory, to quantify feature contributions. This paper introduces the concept of partial Shapley values for stacking models and proposes a method to decompose the Shapley values of stacking models into a sum of partial Shapley values propagated through base learners. We illustrate the approach using a stacked generalization example on a pharmacogenomics dataset and visualize both local and global partial Shapley values.