
The hidden Markov model (HMM) is an unsupervised statistical learning method capable of modeling sequential data by varying the number of hidden states. This paper investigates an HMM-based monitoring procedure for autocorrelated processes, focusing on how the number of hidden states influences its performance. The performance is evaluated using the average run length (ARL) and the relative mean index (RMI). The proposed approach is also benchmarked against a deep learning-based RNN residual chart. The results indicate that the optimal number of HMM states depends on both the autocorrelation structure of the process and the type of process change. Notably, for autoregressive (AR) models with strong first-order autocorrelation, the HMM-based procedure demonstrated superior overall performance compared to the RNN-based procedure.
Radical cystectomy is the standard treatment for muscle-invasive bladder cancer. With the growing adoption of robot-assisted radical cystectomy (RARC) since the early 2000s, comparative studies have shown inconsistent results, often because of methodological biases such as immortal time bias and selection bias. This study aimed to compare overall survival (OS) between open radical cystectomy (ORC) and RARC in patients with bladder cancer using a target trial emulation framework to minimize these biases. Using 45,595 patients identified from the Korean National Health Insurance Service database, we applied a cloning-censoring-weighting (CCW) procedure with inverse probability of censoring weights and compared the results with conventional approaches, including immortal-time-included, immortal-time-excluded, and landmark analyses. The conventional analyses yielded inflated hazard ratios suggesting a survival advantage for ORC, whereas the CCW-based analysis showed no significant difference in OS between the two approaches (hazard ratio 1.02, 95% CI 0.94-1.11). These results suggest that the previously reported survival advantage of ORC may largely reflect methodological bias rather than a genuine clinical benefit.
The purpose of this study is to examine the robustness of the Yuen test in controlling Type I error rates under extreme conditions that frequently arise in applied science research. An analysis of variance (ANOVA) was conducted using four factors: population distribution, total sample size, differences in sample size between groups, and heteroscedasticity between groups. The results showed that the actual Type I error rate of the Yuen test tended to decrease under long-tailed distributions and small sample conditions. Moreover, greater imbalance in group sample sizes was associated with smaller Type I error rates in non-normal conditions. When the assumption of homogeneity of variances was violated, the Type I error rate varied depending on the combination of total sample size and group sample sizes. The findings indicate that the Yuen test was more affected by total and group sample sizes than by distribution shape or variance equality. Therefore, it is recommended to use the t-test when the assumption of homogeneity of variances is satisfied, and to use the Welch test when sample sizes are very small to avoid sample loss due to trimming.
Interest in and potential use of marijuana has been increasing among adolescents. However, data on adolescent marijuana use still show most respondents having no experience using marijuana. Therefore, there is a high possibility that interpretation distortion will occur when fitting simple regression due to the excessive zero values. To address the issue, this study applied Zero-inflated poisson (ZIP) regression model to identify the relationship between the number of marijuana uses or the probability to use it and the home environment, leisure activities, and psychological characteristics. The ZIP model allows for separate estimates of the probability of marijuana use and the frequency of use. The results showed that social environments, family background, and individual behavioral patterns affected the number of marijuana use. These results can serve as the basis for designing youth drug prevention policies and education programs and suggest the direction of establishing a comprehensive prevention system.
The discrimination of large language models (LLMs) against humans is an emerging issue in the real-world application of LLMs. The hallucination and collapse phenomenon is a crucial problem. The starting point for addressing these problems is to discriminate between human and machine texts. In this paper, we investigate the data construction, statistical analysis, including Latent Dirichlet Allocation (LDA), for the features, and the construction of the classifier. Through this process, we validate the possibility of certain features that are useful in discriminating the source of LLMs. We can construct a more elaborate classifier for discriminating between humans and LLMs, and furthermore, various LLMs. It can contribute to classifying the various LLMs and detecting their usage and their behaviors.
This paper considers U-statistics with the Block Wild Bootstrap (BWB) to detect change points in the mean of time series. While the conventional CUSUM test is widely used, its performance deteriorates in the presence of autocorrelation, non-normality or outliers, often resulting in substantial size distortion. In contrast, the proposed BWB Wilcoxon type U-statistic test is more robust to non-normality and outliers. It maintains stable size control and high statistical power across various challenging settings, including those with strong autocorrelation and extreme observations. In particular, when strong outliers are present, the BWB procedure alleviates the conservative tendency of the classical U-statistic and substantially improves its finite-sample performance. This study offers a practical alternative to existing methods for detecting change points in time series with complex dependence structures or data irregularities.
The Pearson correlation coefficient is a standard tool for measuring linear relationships, but it fails to detect the complex nonlinear dependencies often found in real-world data. To overcome this limitation, this study proposes and validates a novel methodology that combines the Alternating Conditional Expectation (ACE) algorithm, a non-parametric transformation technique, with distance correlation (dCor), a robust measure of dependence regardless of its form. We demonstrate through extensive simulations-encompassing various nonlinear functions and non-normal error environments-that the ACE transformation significantly enhances the detection power of dCor. The improvement is particularly pronounced in the presence of outliers, where the error term follows a heavy-tailed or skewed distribution. Furthermore, our analysis reveals that the improvement metric (Delta dCor) acts as a powerful diagnostic tool; strong, symmetric nonlinear patterns, initially masked in the raw data, become clearly identifiable through this metric after the ACE transformation. This suggests that the proposed ACE-dCor framework can be utilized not just for testing dependence, but for diagnosing the intrinsic nature of the relationship between variables.
Sentence order critically affects textual coherence and comprehension, yet real-world data often exhibit disrupted ordering. This study investigates Korean context-aware sentence ordering with Large Language Models, comparing three approaches-Pairwise, Sequence, and Global-through fine-tuning of pretrained models such as KLUE-BERT, KoELECTRA, KLUE-RoBERTa, and T5. Experiments were conducted on the DACON Context-Aware Sentence Ordering AI Competition dataset, comprising 7,350 training and 1,780 test samples. The Pair-wise approach effectively captured local sentence relations but failed to model global coherence. The Sequence approach provided an intuitive framework, yet its performance degraded with longer inputs due to overfitting. By contrast, the Global approach, formulated as a classification problem over all permutations, exhibited the most consistent and superior results. Notably, the KLUE-RoBERTa-based Global model achieved the highest score of 83.71% on the private leaderboard.
Competing risks analysis accounts for situations where mutually exclusive events may occur, providing a more accurate estimation of event probabilities than traditional survival methods. With the increasing interest in high-dimensional data analysis, this study compares the statistical models of CSH and Fine-Gray, the Debiased casebase, the machine learning model RSF, and the deep learning model DeepHit. Extensive simulation studies were conducted under five risk structures and two covariate dimensions, and additionally evaluated under a representative setting with a large sample size. A real-world analysis was performed using colon cancer data from the SEER database. Model performance was assessed using the time-dependent concordance index and the integrated Brier score. The findings of this study are expected to provide practical guidance for selecting appropriate models in competing risks analysis.
This study refines a seasonality procedure commonly applied in Theta-based forecasting models and assesses its impact on the forecasting performance of the DOTM(dynamic optimized theta model), one of the representative Theta models. Existing Theta-based models typically diagnose seasonality using lag-m autocorrelation and apply classical multiplicative decomposition when seasonality is present. We introduce the Kruskal-Wallis test and the Modified QS test for more reliable seasonality identification, and employ sophisticated seasonal-adjustment methods based on the exponential smoothing model and STL decomposition. Using quarterly and monthly time series from the M3 and M4 datasets, we conduct extensive forecasting experiments. The results show that the DOTM adopting the proposed seasonality procedure achieves better forecasting accuracy than the conventional approach, particularly in terms of symmetric mean absolute percentage and mean absolute percentage errors.
Despite their success across domains, Transformer models face challenges in time-series forecasting due to their permutation-invariant attention mechanism, which neglects positional dependencies. Traditional positional encoding alleviates this issue, but its element-wise addition to input embeddings often causes interference between semantic and positional information. This study investigates the structural and theoretical characteristics of concatenation-based positional encoding in comparison with the conventional addition-based approach. By analyzing the attention score formulation, we show that concatenation-based encoding enables semantic and positional components to be processed in independent subspaces, thereby introducing a different inductive bias in the attention mechanism. Attention map visualizations are further employed to qualitatively examine how positional information is reflected under each encoding structure, providing insights into their interpretability. Empirical evaluations are conducted on both time-series forecasting and natural language processing tasks to examine performance and interpretability. The results show that concatenation-based positional encoding yields improved performance compared to the addition-based approach across both task domains, with more noticeable gains in scenarios where positional information plays a critical role. Through a systematic analysis of structural behavior and task-dependent effects, this work contributes to a clearer understanding of how positional information is handled in Transformer models.
As data complexity grows, deep learning models like Long Short-Term Memory (LSTM) have become fundamental for time series analysis. However, despite the emergence of newer architectures, a systematic analysis of the functional roles of LSTM's internal gates is lacking. This study addresses this gap by quantifying the importance of each gate to develop a more efficient model. We employ two analytical approaches: ablation studies to assess component necessity and a novel gate attribution analysis to measure output sensitivity to temporal gate perturbations. Our findings reveal that the forget gate is universally critical for learning long-term dependencies, while the output gate's attribution is concentrated in local intervals preceding prediction. Based on these insights, we propose BiNO-LSTM, an architecture that removes the output gate and incorporates bidirectional context to improve both efficiency and performance. Across five benchmarks, BiNO-LSTM performs comparably to or better than other RNN-based models on real-world datasets, while the best overall architecture varies by task (e.g., TCN on S&P/JSB and Transformer on Adding/Copying). This work contributes a new framework for interpreting gate dynamics and a lightweight yet powerful model structure.
In time series analysis, the coexistence of structural mean shifts and transient anomalies can distort statistical inference, making it essential to distinguish and estimate both phenomena simultaneously. However, existing methods typically handle change points and outliers separately or focus exclusively on only one type of variation, failing to provide an integrated solution. In this paper, we extend a standard mean-shift change point model to concurrently detect both change points and outliers by incorporating either an & ell;(0)(hard-thresholding) or an & ell;(1)(soft-thresholding) penalty for outlier identification. We compare their empirical performance through simulations across varying noise levels and outlier frequencies, evaluating change point detection accuracy and outlier identification rates. We also analyze how each penalty's theoretical properties are reflected in practice. Finally, we apply our method to real accelerometer sensor data, demonstrating its practical utility in accurately localizing both activity transitions and transient sensor spikes within a unified framework.
This study aims to quantitatively estimate the relative infection risk of the COVID-19 Delta variant compared to pre-Delta strains and examine regional differences in transmission dynamics, thereby providing policy-relevant evidence for future variant responses. Seventeen metropolitan and rural regions in South Korea were analyzed using an SIR-based compartmental model combined with maximum likelihood estimation. Infection rates were estimated for the pre-Delta period (January 2020-June 2021) and the Delta-dominant period (July-December 2021), from which the Delta variant's relative infection risk was calculated. The analysis revealed that the relative infection risk of the Delta variant was approximately 1.6 to 3.2 times higher than that of pre-Delta strains. Predicted confirmed cases, derived from estimated infection rates, closely matched actual case data, confirming the model's overall goodness of fit. Correlation analysis between vaccination rates and relative infection risk during the Delta period yielded a Pearson correlation coefficient of-0.034, suggesting that the spread of the Delta variant was influenced more by regional structural factors than by vaccination rates. This study provides empirical evidence that quantifies the Delta variant's relative infection risk across regions, offering a robust scientific basis for targeted intervention strategies and scenario planning for future variants.
This study investigates a diagnostic pipeline for panoramic dental X-ray analysis that integrates tooth enumeration and disease detection, and quantitatively compares the diagnostic performance of various combinations of segmentation and detection models applied within this pipeline. In the tooth localization stage, detection (DINO) and segmentation models (SE U-Net, Mask2Former, OneFormer) were jointly utilized to enhance boundary level precision. For disease detection, three versions of YOLO models were adopted to evaluate structural differences in performance. Using the MICCAI Dentex Challenge 2023 dataset, nine combinations of segmentation and detection models were evaluated. OneFormer was the most accurate segmentation model and YOLOv9 the most accurate disease detector when assessed individually. In the integrated pipeline, the combination of OneFormer and YOLOv8 achieved the highest average precision of 0.411, while the combination of SE U-Net and YOLOv9 showed the highest average recall of 0.622. Interestingly, detection models had a greater influence on overall performance, while segmentation models contributed meaningfully once they achieved sufficient localization quality. With the baseline fusion weight, YOLOv8 achieved the highest precision and YOLOv9 achieved the highest recall. Increasing the YOLO weight led to YOLOv9 delivering the best overall performance.
Data may not be fully observed due to various reasons. When the missing data mechanism is missing at random, the bias in imputation can be reduced by forming imputation classes based on observed variables. Imputation classes can be generated by utilizing other observed variables, however, the quantity of imputation classes can significantly increase when a substantial number of variables are taken into account. To choose only relevant small number of variables, it has been suggested to create imputation classes by utilizing tree-based algorithms. Nevertheless, when dealing with high-dimensional data, these techniques can still lead to an excessive number of imputation classes. Therefore, this study proposes to form imputation classes by employing semi-supervised clustering algorithms. Simulations based on generated data and real data were conducted to compare the proposed techniques with complete-case analysis, imputation without considering imputation classes, and imputation by utilizing tree-based algorithms. The simulation indicates that the proposed method can reduce the bias in imputed data compared to other methods, and it is feasible to effectively manage the number of imputation classes.
The HAR model has the advantage of effectively capturing the characteristics of strongly dependent time series by linearly combining daily, weekly, and monthly lag structures. However, in cases like cryptocurrencies, where the concept of trading days differs from traditional financial markets, using a fixed lag structure may fail to adequately reflect the nature of the data. To address this limitation, this study establishes the structure of the LsHAR model and proposes a method to dynamically estimate the lag structure using the least squares method and 1 step-ahead out-of-sample forecasting. Through simulation experiments, it is shown that both methods converge toward the true lag structure as the sample size grows. In the empirical analysis, the predictive performance of the LsHAR model and the HAR model was compared using realized volatilities of national stock indices and cryptocurrencies. The results showed that the LsHAR model, by estimating the lag structure, outperformed the traditional HAR model in most indices and cryptocurrencies.
This study decomposes Korean unemployment into frictional, structural, and cyclical components using micro-data from the economically active population survey (2005-2024). To address unobserved heterogeneity, the paper employs a finite mixture model estimated with a semi-supervised expectation-maximization (EM) algorithm. The model's identification is anchored by partial labels from self-reported job separation reasons, which ensures an economically meaningful classification and mitigates the arbitrariness of purely data-driven clustering. The framework also incorporates right-censoring to correct for top-coding biases and uses a multinomial logit model with individual covariates to classify unlabeled observations. Robustness is assessed by comparing exponential (constant hazard), log-normal (non-monotonic hazard), and Weibull (monotonic hazard) distributions. The results confirm that the cyclical share of unemployment is strongly pro-cyclical, while the duration of structural unemployment shows a persistent upward trend. Critically, the log-normal model, which provides a superior fit based on BIC, estimates mean durations more than 50% longer than the exponential specification. This finding highlights that models lacking flexible duration dependence risk systematically understating unemployment persistence.
Genome-wide association studies have played a significant role in identifying genetic variants associated with diseases. However, single SNP analyses have shown limitations in explaining the heritability of complex traits. Increasing attention has been directed toward exploring gene-gene interactions to address these challenges. Studies on survival data with censoring remain relatively scarce compared to those focusing on binary or continuous traits. This study compares methodologies for detecting gene-gene interactions in survival data, focusing on multi-locus dimension reduction techniques and machine learning-based approaches. Representative methods were introduced, and their statistical power was evaluated through simulation studies under various realistic scenarios. Based on the results, this study proposes suitable methodologies for different scenarios, providing a practical guideline for effectively identifying gene-gene interaction effects in survival data.
Variational Autoencoder (VAE) is a foundational method used in generative deep neural networks that has significantly contributed to recent advances in artificial intelligence. However, VAE is challenging to understand since its theoretical underpinnings involve complex statistical concepts. This paper elucidates how VAE operates by providing a systematic and accessible overview of the statistical foundations of VAE. It presents VAE as a generalization of reduced-rank regression and factor regression, and revisits the EM algorithm to interpret the meaning of ELBO which is the objective function of VAE. It concludes by discussing variational inference, amortized inference, the architecture of VAE, and implementation strategies to provide deeper insights into VAE.