
Given an exogenous binary treatment D, covariates X, and any outcome Y (binary, count, continuous, …), Ordinary Least Squares (OLS) of Y on D−E(D|X) is consistent to the “overlap-weight” average of heterogeneous effects, where E(D|X) is the propensity score (PS). When D is endogenous with an instrument Z, the Instrumental Variable Estimator (IVE) of Y on D with Z−E(Z|X) as the instrument yields an analogous finding. The conventional approach uses probit/logit for the PS or “instrument score (IS)” E(Z|X), which, however, may be misspecified. In this paper, we use Y−E(Y|X) instead of Y to make the IVE (and OLS) “double-debiasing”, and estimate the nuisance functions {E(Y|X),PS,IS} with machine learning methods. Since this “brute force” nonparametric approach may suffer from nonparametric dimensionality and computational issues, we also explore a middle ground (between the conventional and nonparametric approaches) using Y−E(Y|PS,IS) instead of Y and Y−E(Y|X).
A statistic to test MANOVA hypothesis is proposed for increasing number of independent samples. The statistic, defined as linear combination of U-statistics, is robust to normality and homoscedasticity assumptions, and can be used for balanced or unbalanced models. Whereas the focus is on the number of samples to be large, the limit also incorporates large sample sizes, where the dimension is kept fixed, although may exceed the number of samples as well as the number of replications per sample. Accuracy of the tests statistic is shown through simulations.
Testing cross-sectional independence in high-dimensional sparse panel data is critical. We derive the Berry–Esseen bound for the normalized maximum correlation test statistic, quantifying its convergence rate.
The good lattice point (GLP) method is a classical technique in quasi-Monte Carlo integration and a powerful tool for constructing uniform designs in computer experiments. Despite its widespread use, the rank of the matrix generated by the GLP method has remained an open problem for about 70 years (since its inception in 1959). While an upper bound was established in 2021 by Elsawah et al. the exact rank remained unresolved. This paper provides the first rigorous determination of this rank for doubly even sample sizes. Using number-theoretic properties of coprime integers and the Euler function, together with a novel induction argument on auxiliary vectors, we prove that for any GLP matrix with doubly even row size n, the rank is exactly 12φ(n)+1, where φ is the Euler function. We also derive necessary and sufficient conditions for constructing a full-rank GLP matrix. These results close a long-standing gap in the theory of GLP sets and provide a solid foundation for future work on the general case of any row size.
For a probability measure μ with unbounded support, W∞(μ,μn)=∞ almost surely. This makes the untruncated empirical W∞ problem degenerate. We therefore truncate the population law to μDn≔μ(⋅∣Dn) and study W∞(μDn,μn), thereby retaining all empirical atoms. For Gibbs measures μ(dx)=Z−1e−V(x)dx on Rd, d≥2, with smooth moving level sets Dn={V≤tn}, we prove that, in the critical truncation regime and under uniform boundary-regularity and multiscale core-matching conditions, W∞(μDn,μn)=ΘP(anlog(Rn/an)), where an and Rn are the normal and tangential boundary scales, respectively. The logarithmic factor reflects the balance between local expected occupancy and the covering complexity of the moving boundary, rather than fluctuations of the largest observed potential level; this boundary-entropy amplification is absent in one dimension.
Objective priors are commonly used in Bayesian analysis when prior information is not available or when the likelihood should dominate the posterior distribution. For lifetime data, censoring changes the observed experiment but does not change the latent failure-time distribution. In this paper, we propose an objective prior obtained from local variations of cumulative probabilities. The prior is induced by a Cramér–von Mises type discrepancy between distribution functions. It is invariant under smooth reparameterizations and monotone transformations of time, and is coherent with parameter-free reductions such as non-informative censoring. The results indicate that cumulative probabilities provide a natural scale for constructing objective priors in censored survival models.
In this paper, we derive recursions for compound distributions in cases where the claim number distributions satisfy Panjer’s recursion pn=pn−1(α+β/n) and where the derivative of the probability generating function (pgf) of both the univariate and multivariate classes of aggregate claim amounts can be written as a ratio of two polynomials. Finally, examples of discrete claim size distributions belonging to these classes are provided.
We propose a fused variational estimator to recover the effective number of communities in stochastic block models (SBMs) by deliberately overfitting the model and subsequently merging fitted blocks that exhibit nearly identical connectivity profiles. Specifically, we introduce a novel SBM-specific discrepancy measure and a fusion penalty that encourages redundant blocks to merge, along with a barrier term ensuring all block proportions remain positive. The resulting estimator is efficiently computed using a variational EM-type algorithm featuring closed-form updates. Simulation studies and an application to a real-world political alliance network demonstrate that the proposed method accurately recovers the true number of communities and yields highly interpretable community structures.
This paper investigates $L^p$-quantile regression estimation for varying coefficient partially functional linear models. The infinite-dimensional slope function and the varying coefficients are estimated using principal component basis and B-spline basis, respectively. The resulting estimators are shown to enjoy desirable asymptotic properties under some regularity conditions. Numerical studies demonstrate the advantages of the proposed method over existing approaches.
The p-rotor walk on Z is a self-interacting walk that interpolates between the simple random walk and the deterministic rotor walk. While the weak convergence of this model to a perturbed Brownian motion is known, its almost sure asymptotic boundaries have not been characterized. In this paper, we establish the exact Law of the Iterated Logarithm (LIL) for the p-rotor walk. Utilizing the decomposition of the walk into a martingale perturbed by its running extrema, we obtain first a functional Law of the Iterated Logarithm for the linearly interpolated paths of the p-walk. We then obtain the classical LIL constants by solving a calculus of variations problem over the perturbed Strassen set.
We develop a spectral based coefficient of determination to measure how well the spectral density of a stationary process is represented by the class of MA(q) models. Using periodogram-based estimators, we establish asymptotic normality, derive tests for the MA(q) hypothesis, and construct procedures for determining the smallest order q achieving a prescribed approximation quality.
This short note explores the maximum-entropy walk on the unit interval that is a median-martingale. That is, the median of its next state is equal to its current state. The stationary distribution of this walk is shown to be the arcsine distribution, and we provide a proof that elucidates the connection to two classical arcsine laws for Brownian motion. The notion of a martingale is further generalized, and a larger class of walks is considered and similarly characterized.
In many analyses the object reported at the end is not fixed in advance, but is chosen after a preliminary search over variables, subgroups, transformations, models or contrasts. Classical selective-inference methods are most effective when this search can be written as an explicit selection event. This note treats the less structured case in which the selection rule is a black box and inference is required for the target indexed by the selected object. We show that, for any fixed-target confidence procedure, selected-target noncoverage is bounded by the nominal fixed-target noncoverage plus the average total variation distance between the marginal law of the inferential data and its conditional law given the selected object. A mutual-information bound follows immediately. The result recovers sample splitting as the zero-leakage case and gives explicit guarantees for noisy screening through a Gaussian information bound. Thus the inferential cost of black-box selection is quantified by the information that the selected object carries about the inferential sample.
We propose a unified mixture sampler (UMS) that provides a universal estimation framework for nonlinear state–space models with ‘exp-exp’ likelihood kernels. Unlike existing methods that require deriving new mixture approximations for each specific distribution, our approach dynamically adapts the standard ten-component mixture from Omori et al. (2007) through a deterministic re-centering and rescaling algorithm. Applying this to the stochastic conditional duration (SCD) model, we demonstrate that the proposed sampler can efficiently handle unknown shape parameters — such as those in Weibull or Gamma distributions — by updating mixture components near-instantaneously during MCMC iterations. The UMS not only simplifies implementation but also ensures exact inference via a lightweight Metropolis–Hastings step. Numerical examples show that our method substantially outperforms a conventional slice sampling approach, significantly reducing autocorrelation in MCMC samples while maintaining high computational efficiency. This unified framework encompasses a wide range of applications, including logit and Poisson models, as well as various SCD model specifications, providing a highly efficient alternative to model-specific samplers.
We study a random field formed by summing contributions from the points of a marked Poisson point process. Such fields can be used to model a range of physical phenomena. For instance, each point may correspond to a galaxy contributing light to an observed image, or to a neuron firing and generating an electrical signal. In both cases, the contribution from each point depends on its associated mark. Under mild conditions, including uniform bounds on the individual contributions, we show that the rate at which points occur and the distribution of their marks can be recovered from the law of the observed field.
At focus-relevant boundaries, a hard-selected focus may lack a finite influence function despite bounded component influence. We derive local mixture limits, show nonuniform gross-error sensitivity under exponential weighting, and analyze candidate models fitted by robust estimating equations.