
A limit order book (LOB) queueing system is considered in which limit orders are generated by Markov modulated Poisson processes (MMPP) to capture the clustered nature of order arrivals. The queueing dynamics are represented by a multidimensional birth-death type Markov chain, and the probability distributions of the state variables are obtained through matrix computing procedures for Markov chains. By modifying the Markov chain structure and assigning selected states as absorbing states, two conditional probabilities relevant to high-frequency trading are computed: the probability of a midprice increase and the probability of order execution before a midprice change. Numerical results show that clustered MMPP arrivals substantially influence both probabilities, indicating that clustering in the limit order arrival process is an important factor in limit order book queueing models.
Model averaging serves as a robust and effective framework for improving predictive accuracy in survival analysis. Despite its potential, the majority of existing research has primarily concentrated on right-censored data. Unlike right-censored data, which permits precise observation of certain failure times, interval-censored data offers significantly less information, as exact failure times remain unobserved. To address this challenge, an optimal model averaging estimation method tailored specifically for interval-censored data is introduced. This approach leverages sieve maximum likelihood estimation to estimate the optimal weights of the candidate models. Extensive simulation studies underscore the superior performance of the proposed method across diverse evaluation metrics, demonstrating its flexibility and robustness under varying error distributions. Moreover, the asymptotic properties of the weight estimators are established. Finally, the practical utility of the method is validated through its application to real-world data from a cardiac allograft vasculopathy (CAV) study.
Estimating the parameters of max-stable parametric models poses significant challenges, particularly when some parameters lie on the boundary of the parameter space. This situation arises when a subset of variables exhibits extreme values simultaneously, while the remaining variables do not—a phenomenon commonly referred to as an extreme direction. A novel estimator is proposed for the parameters of a general parametric mixture model, incorporating a threshold exceedances approach based on a pseudo-norm penalization. The latter plays a crucial role in accurately identifying parameters at the boundary of the parameter space. Additionally, the estimator comes with a data-driven algorithm to detect groups of variables corresponding to extreme directions. The performance of the estimator is assessed in terms of both parameter estimation and the identification of extreme directions through extensive simulation studies. Finally, the method is applied to two real-world datasets: discharge measurements at stations along the Danube river, and financial portfolio losses from stocks listed on the NYSE, AMEX, and NASDAQ. In both applications, the sets of variables that can become large simultaneously are identified.
A novel data-driven methodology is presented for the joint selection of prior parameters for both fixed and random effects in Linear Mixed Models (LMMs). This approach facilitates the estimation of complex random-effects structures, as well as potentially high-dimensional data. Although Bayesian frameworks require the specification of informative prior parameters, such values are often unavailable a priori - especially for random-effect covariances. The proposed method automates this selection through an Empirical Bayes framework, which maximizes the marginal likelihood using an efficient Laplace approximation. Numerical simulations demonstrate that this methodology significantly enhances parameter estimation accuracy and predictive performance. Finally, an application to a real-world air pollution and health dataset illustrates how the method enables the use of more sophisticated and statistically appropriate models to improve predictive outcomes.
Large spatial datasets with non-Gaussian responses are increasingly common in environmental monitoring, ecology, and remote sensing, yet scalable Bayesian inference for such data remains challenging. Markov chain Monte Carlo (MCMC) methods are often prohibitive for large datasets, and existing variational Bayes methods rely on conjugacy or strong approximations that limit their applicability and can underestimate posterior variances. A scalable variational framework that incorporates semi-implicit variational inference (SIVI) with basis representations of spatial generalized linear mixed models (SGLMMs), which may not have conjugacy, is proposed. The proposed framework accommodates gamma, negative binomial, Poisson, Bernoulli, and Gaussian responses on continuous spatial domains. Across 20 simulation scenarios with 50,000 locations, SIVI achieves predictive accuracy and posterior distributions comparable to Metropolis–Hastings and Hamiltonian Monte Carlo while providing notable computational speedups. Applications to MODIS land surface temperature and Blue Jay abundance further demonstrate the utility of the approach for large non-Gaussian spatial datasets.
In many settings, a data curator or data analyst links records from two files to produce an integrated dataset. These linked data are then used to estimate regression models of interest. This two-stage approach does not necessarily account for the uncertainty in the model parameters that results from uncertainty in the linkages. Further, it does not leverage the relationships among the non-linking variables in the two files to help identify incorrect linkages. A multiple imputation framework is proposed to address these shortcomings. First, a bipartite Bayesian record linkage model is used to generate multiple plausible linked datasets. This model does not use the non-linking variables. Second, each linked file is presumed to comprise a mixture of true links and false links. The mixture model is estimated using an EM algorithm that leverages the information provided by the non-linking variables. Point and variance estimates of the regression parameters in each plausible linked file are combined via multiple imputation inferences. Using simulation studies of linear regressions, it is demonstrated that the mixture modeling approach can have desirable repeated sampling properties. The mixture modeling approach is illustrated using Bayesian record linkage of data from the Survey of Household Income and Wealth, examining a regression involving the persistence of income.
Multi-state models are commonly used for intermittent observations of a state over time, but these are generally based on the Markov assumption, that transition rates are independent of the time spent in current and previous states. In a semi-Markov model, the rates can depend on the time spent in the current state, though available methods for this are either restricted to specific state structures or lack general software. One approach uses a “phase-type” distribution for the sojourn time in a state, which expresses a semi-Markov model as a hidden Markov model, allowing the likelihood to be calculated easily for any state structure. While this approach involves a proliferation of latent parameters, identifiability can be improved by restricting the phase-type family to one with similar properties to a simpler distribution such as the Gamma or Weibull. A moment-matching method to obtain this family is proposed, making general semi-Markov models for intermittent data accessible in software for the first time. The method is implemented in a new R package, msmbayes, which implements Bayesian or maximum likelihood estimation for multi-state models with general state structures and covariates. The software is tested through simulation-based calibration, and an application to cognitive function decline illustrates the use of the method in a typical modelling workflow.
Multivariate extreme value analysis quantifies the probability and magnitude of joint extreme events. Classical multivariate models, such as max-stable or multivariate generalised Pareto distributions, generally have a high computational cost of fitting, which limits their application. To overcome this, models based on the asymptotically dependent multivariate Pareto distribution have recently incorporated graphical models to induce sparsity and reduce the dimension of the parameter space. While this approach is computationally efficient, the assumption of asymptotic dependence is inappropriate for many applications. The conditional multivariate extreme value model (CMEVM) is a popular model for which the asymptotic dependence assumption is not required. Unfortunately, inference for this model is semi-parametric, and consequently, it has poor predictive performance in high dimensions. An extension of the CMEVM that allows both the incorporation and selection of sparse dependence structures, and fully parametric prediction is proposed. The approach fills a current gap in statistical methodology by extending graphical models to asymptotically independent multivariate extreme value models. To support inference in high dimensions, a stepwise inference procedure that is computationally efficient and loses no information or predictive power is proposed. Simulation studies show the model is highly flexible, and an application to discharges in the upper Danube River basin provides promising results.
For multivariate time series streams, each sequence grows indefinitely over time and exhibits complex interdependencies with others. Such data have become increasingly prevalent across various fields. These streams are processed sequentially in the order of generation, with each data point discarded immediately after processing. The one-pass processing constraint, combined with the lack of historical data, poses significant challenges that traditional methods struggle to overcome effectively. To address these issues, a comprehensive framework for online estimation and dynamic forecasting based on the vector autoregressive (VAR) model is proposed. Beyond online updating, the framework also dynamically estimates the lag order, accounting for the resulting changes in the dimensionality of the summary statistics and controls error accumulation in the covariance matrix estimation. Theoretical results show that the proposed online estimators are asymptotically equivalent to their oracle counterparts, while the resulting dynamic forecasts exhibit stable predictive performance throughout the online updating process. Finally, extensive numerical experiments and real-world data analyses demonstrate the practical effectiveness and robustness of the proposed approach, highlighting its suitability for real-time applications under memory and computational constraints.
Connectivity between human brain regions has been proved to be highly related to phenotypical characteristics. Based on the graphical model, a novel nonparametric statistical method is proposed to estimate such dynamic connectivity among human brain regions. In the proposed model, the nodes of the graph are random functions rather than random variables, placing them within the scope of functional data analysis. Statistical operators are then exploited to capture the interdependence between these functions, which are allowed to vary with external covariates. Specifically, these random functions and operators are assumed to reside in nested Hilbert spaces, and a new precision operator-termed the nonparametric conditional additive precision operator-is constructed to capture dynamic interdependence. This operator serves as a functional counterpart to the precision matrix in the traditional Gaussian graphical model. The approach is also applicable when external covariates are high-dimensional or random functions. Furthermore, theoretical analysis establishes the consistency and optimal convergence rates for the proposed estimators. Simulation studies demonstrate that the proposed method significantly outperforms existing competitors in estimating graphs with nonlinear edges and accurately recovering how the strength of edges changes with external covariates. Finally, the framework is successfully applied to a real functional magnetic resonance imaging (fMRI) dataset.
Function registration (or alignment) is a fundamental challenge in functional data analysis, where the goal is to identify compositional variability and remove such “noise” before conducting in-depth analysis such as component methods and regressions. While significant progress has been made in this area over the past few decades, the registration of observations with both compositional and additive noises remains understudied. Most state-of-the-art methods, including the well-known Fisher-Rao approach, rely on derivatives and often cannot adequately account for additive noise. A Bayesian framework is proposed that overcomes these limitations by directly modeling likelihoods in the original function space, enhancing robustness to additive noise. Moreover, the centered log-ratio transformation is employed to represent time-warping functions, mapping them to a square-integrable subspace that is well-suited for stochastic process priors. The framework is extended to conduct signal estimation in the presence of compositional noise, additive noise, constant translation, and constant scaling, achieving desirable results. Through extensive simulations and real-data applications, the superior performance and robustness of the new method over state-of-the-art approaches are demonstrated.
Boundary detection in two-dimensional spatial data under heavy-tailed errors is investigated, with particular emphasis on structural changes in nonconvex, curved, and multi-region settings. Traditional boundary detection methods based on local mean differences can be unstable in the presence of heavy-tailed noise and outliers. A robust local scanning procedure is developed by integrating bounded M-score transformation with quadrant-wise local block contrasts. The resulting statistic uses bounded score-transformed residuals to reduce the influence of extreme observations while retaining the interpretability of local scanning. Asymptotic properties under the null hypothesis and local alternatives are established, and the role of the robust transformation in mitigating heavy-tailed perturbations is clarified. Simulation studies compare the proposed method with a least-squares-type local statistic under several parameter settings. The results indicate greater stability in hypothesis testing and improved boundary recovery when heavy-tailed perturbations are strong. A semi-synthetic housing price example further demonstrates the applicability of the procedure to complex spatial boundary scenarios.
A simultaneous confidence region (SCR) is constructed for the long-run covariance function of functional time series when the data is collected with error contamination. First, the asymptotic properties of a kernel-weighted estimator under C[0,1]2 topology is derived, leading to simultaneous inference. Then, a B-spline estimator is proposed, which enjoys the corresponding “oracle” efficiency, i.e., it is asymptotically equivalent to the estimator derived from fully observed trajectories without measurement errors. The developed methodology can be extended to independence testing and relevant difference testing. Extensive numerical experiments support the theoretical findings. Electroencephalogram (EEG) data are used to illustrate the proposed method.
High-dimensional statistical inference on linear functionals of regression parameters is challenging. Traditional methods are based on penalized estimates for regression parameters. An existing approach sidesteps estimation of regression parameters by transforming linear functionals into one part of coefficients within a restructured model with orthogonal features, assuming known predictor covariance matrix. However, when predictor covariance is unknown, that method faults to an identity matrix, imposing sparsity on the precision matrix. The proposed method introduces a data-splitting strategy to maintain feature orthogonality by plugging in the precision matrix estimated from a subset of data. The statistic is constructed using the other part of the data, ensuring the test statistic's asymptotic normality under null and local alternatives. The method's strength lies in its ability to handle non-sparse original regression parameters or loadings, as well as strong correlations among predictor variables. The proposed cross-fitting method enhances statistical power under mild constraints, supported by theory and validated by numerical studies.
Relational event models (REMs) can infer the generative properties of longitudinally observed social networks with instantaneous edges. They assume conditional independence of edges given sufficient network statistics formed over the past event sequence. A popular specification in REMs is to subject these statistics to exponential temporal decay with a fixed half-life parameter to attribute higher importance to more recent edge events in the formation of network statistics. Assuming a fixed half-life parameter may cause biased estimates and obfuscates the temporal horizon over which network effects operate in empirical social systems. These limitations are addressed by proposing fully Bayesian estimation of REMs and designating the half-life parameter as an estimable quantity. A “pre-computation” strategy is devised to speed up calculations for practical feasibility of the sampling procedure. The approach is adapted to discourse network analysis, which models political actors’ statements about their preferred policy beliefs as dynamic networks. An application to the policy debate on reforming the German public pension system illustrates how temporal decay for inertia, actor activity, belief popularity, and actor homophily can be estimated alongside the main coefficients. Convergence diagnostics and an illustration of bias correction relative to fixed parameters are provided, stabilisation using hyper-parameters and computational complexity are discussed, and the approach is extended to include both Breslow’s and Efron’s methods for breaking ties in the event sequence.
With modern data increasingly collected and stored across computing nodes by design, such as by time, geography, or client, distributed quantile regression must confront non-randomly stored data, non-smooth objectives, and communication bottlenecks all at once. We propose two communication-efficient distributed quantile regression (QR) estimators that address these challenges jointly. First, we implement convolution smoothing to render the check loss differentiable and construct a Poisson subsampling-based estimator within the smoothed QR framework. Theoretically, we establish its convergence rate and asymptotic normality and derive L-optimal sampling probabilities by minimizing the trace of the asymptotic covariance, ensuring statistical efficiency. Building on this foundation, we introduce a Distributed Smoothed Quantile Regression estimator with Poisson subsampling (DSQR-P), in which each worker transmits only a small Poisson subsample and the associated gradient information to the master node, achieving substantial communication reduction and computational scalability while preserving near full-sample accuracy. For high-dimensional data, we further develop a regularized distributed estimator (DSQRH-P) incorporating sparsity-inducing penalties. The resulting optimization problem is efficiently solved via the Local Adaptive Majorization-Minimization (LAMM) algorithm. We also establish its oracle property under standard sparsity conditions. Extensive simulation studies and two real data applications demonstrate that the proposed methods achieve excellent scalability, robustness, and statistical efficiency while substantially reducing computational and communication costs.