
ABSTRACT When a sampling distribution is discrete, the coverage of a confidence interval follows a sequence of peaks and troughs when plotted against the value of the parameter of interest. Then methods of forming a confidence interval must either be conservative, with a coverage that is almost always above the nominal confidence level, or give a coverage that is sometimes below the nominal level. Many methods have been proposed that adopt the latter policy, so the requirements of a confidence interval are often tempered. Garthwaite et al. (2024) suggest that the definition of a confidence interval should be relaxed in a clearly defined way if the definition is not strictly adhered to. For one‐sided intervals, they define locally correct confidence intervals as intervals whose average coverage between consecutive peaks is no smaller than the nominal level. Two‐sided locally correct confidence intervals are formed from two one‐sided intervals. They propose a method for forming locally correct intervals for a binomial proportion that minimizes the average length of such intervals. In this article, we generalize the method so that it can be used with other discrete sampling models. The method is illustrated through examples and compared with alternative methods.
ABSTRACT Intraday transaction‐level data on financial assets constitute high‐frequency, irregularly spaced time series, posing challenges for volatility modeling. Traditional stochastic volatility models, designed for regularly spaced time series, are not directly applicable in this setting. This paper introduces two hierarchical Bayesian models—the irregular basic multivariate stochastic volatility autoregressive conditional duration model and the irregular dynamic multivariate stochastic volatility autoregressive conditional duration model to address these challenges and effectively model volatility in irregular intra‐day data. These models incorporate the autoregressive conditional duration model to account for irregular gaps between transactions, improving the understanding of market microstructure and future dynamics. The analysis is conducted within a Bayesian framework using the Hamiltonian Monte Carlo algorithm with No‐U‐turn sampler in R via the cmdstanr package. We demonstrate the efficacy of our methodology through simulation studies and an illustrative example using intra‐day prices of healthcare stocks traded on the New York Stock Exchange at the microsecond level. Utilizing the refresh time sampling technique, we synchronize transactions, compute synchronized log‐returns and gaps, and use these for modeling purposes.
We study the admissibility of generalized Bayes estimators for regression coefficients in normal linear models under hierarchical -priors. While admissibility results are available for canonical normal means problems with unknown variance, their applicability to the regression problem is nontrivial. We explore the admissibility of regression-induced generalized Bayes estimators under scaled quadratic loss with arbitrary positive definite weight matrices by carefully transporting recent normal means results to the regression parameter. Explicit conditions on the hyperparameters governing the prior on are derived. Our results provide a decision-theoretic justification for generalized Bayes estimators of regression coefficients under hierarchical g-priors.
This paper studies the asymptotic distribution of a constrained lasso-type estimator for denoising signals defined on the nodes of a graph, where the underlying structure encodes relationships between variables. We show that, under suitable assumptions on the penalization parameters, the limiting distribution of the estimator is obtained by applying the corresponding constrained procedure to the asymptotic distribution of the unrestricted estimator. Thus, the constrained estimator shares the same convergence rate as the unrestricted estimator. Without the fusion penalty, the limiting distribution is obtained by applying the individual nearly isotonic estimators to the corresponding sub-vectors of the unrestricted estimator's asymptotic distribution, similarly to the limit behavior of isotonic regression.
ABSTRACT We propose a new omnibus goodness‐of‐fit test based on trigonometric moments of probability‐integral‐transformed data. The test builds on the framework of the LK test introduced by Langholz and Kronmal [J. Amer. Statist. Assoc. 86 (1991), 1077–1084], but fully exploits the covariance structure of the associated trigonometric statistics. As a result, our test statistic converges under the null hypothesis to a distribution, even in the presence of nuisance parameters, yielding a well‐calibrated rejection region. We derive the exact asymptotic covariance matrix required for normalization and propose a unified approach to computing the LK normalizing scalar. The applicability of both the proposed test and the LK test is substantially expanded by providing implementation details for 11 families of continuous distributions, covering most commonly used parametric models. Simulation studies demonstrate accurate empirical size, close to the nominal level, and strong power properties, yielding fully plug‐and‐play procedures. Further insight is provided by an analysis under local alternatives. The methodology is illustrated using surface temperature forecast errors from a numerical weather prediction model.
Markov chains are fundamental models for stochastic dynamics, with applications in a wide range of areas such as population dynamics, queueing systems, reinforcement learning, and Monte Carlo methods. Estimating the transition matrix and stationary distribution from observed sample paths is a core statistical challenge, particularly when multiple independent trajectories are available. While classical theory typically assumes identical chains with known stationary distributions, real-world data often arise from heterogeneous chains whose transition kernels and stationary measures might differ from a common target. We analyze empirical estimators for such parallel Markov processes and establish sharp concentration inequalities that generalize Bernstein-type bounds from standard time averages to ensemble-time averages. Our results provide nonasymptotic error bounds and consistency guarantees in high-dimensional regimes, accommodating sparse or weakly mixing chains, model mismatch, nonstationary initializations, and partially corrupted data. These findings offer rigorous foundations for statistical inference in heterogeneous Markov chain settings common in modern computational applications.
We propose an enhanced estimation method for the Box-Cox transformation (BCT) cure rate model parameters by introducing a generic maximum likelihood estimation algorithm, the sequential quadratic Hamiltonian (SQH) scheme, which is based on a gradient-free approach. We apply the SQH algorithm to the BCT cure model and, through an extensive simulation study, compare its model fitting results with those obtained using the recently developed non-linear conjugate gradient (NCG) algorithm. Since the NCG method has already been shown to outperform the well-known expectation maximization algorithm, our focus is on demonstrating the superiority of the SQH algorithm over NCG. First, we show that the SQH algorithm produces estimates with smaller bias and root mean square error for all BCT cure model parameters, resulting in more accurate and precise cure rate estimates. We then demonstrate that, being gradient-free, the SQH algorithm requires less CPU time to generate estimates compared to the NCG algorithm, which only computes the gradient and not the Hessian. These advantages indicate that the SQH algorithm performs favorably compared with the NCG method for estimation in the BCT cure model. Finally, we apply the SQH algorithm to analyze a well-known melanoma dataset and present the results.
Risk prediction models are widely used to guide real-world decision-making in areas such as healthcare and economics, and they also play a key role in estimating nuisance parameters in semiparametric inference. The super learner is a machine learning framework that combines a library of prediction algorithms into a meta-learner using cross-validated loss. In the context of right-censored data, careful consideration must be given to both the choice of the loss function and the estimation of expected loss. Moreover, estimators such as the inverse probability of censoring weighting requires accurate modeling and an estimator of the censoring distribution. We propose a novel approach to super learning for survival analysis that jointly evaluates the candidate learners for both the event-time distribution and the censoring distribution. Our method imposes no restrictions on the algorithms included in the library, accommodates competing risks, and does not rely on a single prespecified estimator of the censoring distribution. We establish a finite-sample bound on the average price we pay for using cross-validation, and show that this price vanishes asymptotically, up to poly-logarithmic terms, provided that the size of the library does not grow faster than at a polynomial rate in the sample size. We demonstrate the practical utility of our method using prostate cancer data and compare it to existing super learner algorithms for survival analysis using synthesized data.
In active surveillance of prostate cancer, cancer progression is interval-censored and the examination to detect progression is subject to misclassification, usually false negatives. Meanwhile, patients may initiate early treatment before progression detection, constituting a competing risk. We developed the Misclassification-Corrected Interval-censored Cause-specific Joint Model (MCICJM) to estimate the association between longitudinal biomarkers and cancer progression in this setting. The sensitivity of the examination is considered in the likelihood of this model via a parameter that may be set to a specific value if the sensitivity is known, or for which a prior distribution can be specified if the sensitivity is unknown. Our simulation results show that misspecification of the sensitivity parameter or ignoring it entirely impacts the model parameters, especially the parameter uncertainty and the baseline hazards. Moreover, specification of a prior distribution for the sensitivity parameter may reduce the risk of misspecification in settings where the exact sensitivity is unknown, but may cause identifiability issues. Thus, imposing restrictions on the baseline hazards is recommended. A trade-off between modelling with a sensitivity constant at the risk of misspecification and a sensitivity prior at the cost of flexibility needs to be decided.
In this paper, we address the adequacy testing problem for high-dimensional factor-augmented regression models. Existing methods often require the estimation of high-dimensional precision matrices or regression coefficients, resulting in significant computational complexity. To overcome this challenge, we refine the bootstrap testing procedure based on the maximum-type test statistic of the estimated score function as introduced by Beyhum and Striaukas (2024b). Our method avoids the estimation of both precision matrices and high-dimensional regression coefficients, thereby substantially reducing computational complexity and simplifying implementation. Theoretically, we establish the validity of the proposed procedure under the null hypothesis and demonstrate its power to detect signals under suitable conditions. We demonstrate the finite-sample performance of the proposed method through comprehensive simulations and an empirical illustration using a macroeconomic dataset.
Diffusion processes and more generally, stochastic differential equations (SDEs), are widely used to model natural and financial systems. However, accurately simulating them remains challenging due to the limitations of discretization methods. We propose a recursive algorithm to approximate the transition density of scalar diffusion processes using Hermite polynomial expansions. Unlike standard numerical schemes, our method uses an expansion in Hermite polynomials to approximate the transition density without requiring an arbitrarily small discretization step. This approximation is then used to simulate diffusion paths with high fidelity. Numerical experiments, including the Vasicek and CIR processes, confirm the effectiveness and efficiency of the method.
We consider a sparse vector autoregressive model with divergent lag order. As a linear model, all its explanatory variables are lagged responses such that there may be high correlation between them. Hence, the broken adaptive ridge procedure is employed for its iterative algorithm, which starts with a ridge estimator as the initial one. We obtained parameter estimation and model selection simultaneously by the procedure named VBAR in this paper. Theoretically, we established that the VBAR procedure is consistent for model selection and an oracle for parameter estimation. Simulations demonstrate the superiority of the VBAR procedure over Lasso, Adaptive Lasso, and SCAD procedures. Additionally, the Google Flu Trends data are analyzed by the VBAR procedure, which gives a more sparse model and more accurate predictions compared with other procedures.
A family of multivariate maximum-entropy distributions with general discrete support is derived. Members of the family are distinguished from each other by different specified sets of joint non-central moments. One of the members can be viewed as a discrete analogue of the continuous multivariate normal distribution. In an illustrative example, maximum likelihood estimation is used in fitting several bivariate discrete maximum-entropy distributions to empirical data.
This article proposes three new goodness-of-fit tests for the zero-altered Poisson distribution, or equivalently for the positive Poisson distribution, based on positive data, that is, data truncated at 0, whose test statistic is built using a characterization of this law. It is shown that the proposed tests are consistent against any fixed alternative and that the parametric bootstrap method can accurately approximate their null distributions. The power of these tests is investigated through a large simulation study, where it is also compared with some existing tests, showing a very competitive behavior. Several applications to real datasets illustrate the usefulness of the tests.
In this article, we propose a consistent estimator for the asymptotic variance matrix of the least squares estimator in Fractionally AutoRegressive Integrated Moving-Average (FARIMA) models. Our approach allows for error terms that are uncorrelated, but not necessarily independent. Modified versions of the Wald, Lagrange Multiplier, and Likelihood Ratio tests are proposed for testing linear restrictions on the parameters of these models, particularly for assessing the presence or absence of long-memory characteristics. Simulation studies are conducted to support the theoretical findings. Additionally, an application to Nikkei stock returns and monthly temperature data demonstrates the practical relevance of the theoretical results.
Logistic regression has been a standard multivariate analysis method for binary outcomes in clinical and epidemiological studies; however, the odds ratios cannot be interpreted as effect measures directly. The modified Poisson and least-squares regressions are alternative effective methods to provide risk ratio and risk difference estimates. However, their ordinary Wald-type inference methods using the sandwich variance estimator seriously underestimate the statistical errors under small or moderate sample settings. In this article, we develop alternative likelihood-ratio-type inference methods for these regression analyses based on Wedderburn's quasi-likelihood theory. An advantage of the proposed methods is that we have correct information for the true models (i.e., the binomial log-linear and linear models). Using this modeling information, we develop an effective parametric bootstrap algorithm for accurate inferences. In particular, we propose the Bartlett-type mean calibration approach and bootstrap test-based approach for the quasi-likelihood ratio statistic. In addition, we propose another computationally efficient modified approximate quasi-likelihood ratio statistic whose large sample distribution can be approximated by the distribution and its bootstrap inference method. In numerical studies by simulations, the new bootstrap-based methods outperformed the current standard Wald-type confidence interval. We applied these methods to a clinical study of epilepsy.
Causal mediation analysis (CMA) is commonly used to investigate the extent to which an intermediary variable mediates the effect of the exposure on the outcome. The naive estimators related to CMA can be severely biased when a measurement error is in the common confounder for the exposure-outcome, exposure-mediator, and mediator-outcome relationship. This paper quantifies classical, non-differential measurement error bias on a common confounder in estimating natural effects. It also proposes three methods based on the method of moments (MoM), regression calibration, and SIMEX for estimating effects associated with CMA in the presence of measurement error in a common confounder. Additionally, the article derives using observed data to estimate the ranges of the extent of measurement error, enabling one to implement sensitivity analyses and assess the robustness of effect estimates even without prior information on the measurement error. The performance of the proposed corrected methods of estimating natural effects is compared using a simulation study. The proposed MoM method performs better in the context of bias and standard error than the other methods considered in such situations.
This paper presents the exact analytical solution to the long‐standing problem of determining the maximum likelihood estimator (MLE) for the parameter of the logarithmic distribution. Despite having been introduced more than 80 years ago, no one has derived an explicit MLE for the logarithmic distribution parameter. In addition to this primary result, we also derive the MLE asymptotic variance and provide other related findings that further contribute to the understanding of this distribution.
Detecting influential observations and verifying their impact on model fitting and parameter estimation are essential steps in statistical modeling. Different approaches can be utilized to that end, including the case-deletion method, which evaluates the individual impact on the estimation process, and the local influence approach, which investigates the model sensitivity under some perturbation. This paper introduces case-deletion measures and influence diagnosis for a flexible class of longitudinal models, the scale mixture of skew-normal linear mixed models. This approach includes skewed and heavier-than-normal-tailed distributions while accounting for useful within-subject dependence structures. The method's capability of detecting atypical observations under repeated measurements and the impact of outliers on parameter estimation from models accounting for different distributions are evaluated in simulation studies and a real data illustration.
The availability of both right‐censored and left‐truncated right‐censored failure time data commonly occurs when failure/censoring times are collected by cross‐sectioning the population and acquiring new failure/censoring times during a follow‐up period. Under a frequentist paradigm, the Kaplan‐Meier estimator can be adjusted to account for both types of data separately when estimating the failure time survival function. In this note, we review the analogous Bayesian nonparametric survival function estimators that are applicable to the individual types of data, separately, and propose two Bayesian nonparametric survival function estimators using both the right‐censored and left‐truncated right‐censored failure time data, simultaneously.