
We introduce a novel finite mixture model for biclustering ordinal data matrices, simultaneously clustering rows and columns within the Underlying Response Variable (URV) framework. Ordinal responses are modeled as discretizations of latent Gaussian variables, while component-specific covariance matrices are parsimoniously parameterized via a factor analytic structure to capture complex dependence patterns efficiently. To address computational challenges arising from the high-dimensional likelihood evaluation inherent to ordinal data, model parameters are estimated using a Composite Likelihood (CL) approach, which offers significant computational advantages with negligible loss of efficiency. Extensive simulation studies demonstrate the effectiveness of the proposed method in terms of accurate cluster recovery and reliable parameter estimation. We further illustrate the applicability of the methodology through an empirical analysis of two real-world ordinal datasets, highlighting its potential for uncovering meaningful latent bicluster structures.
We develop quantile ratio regression for panel data, to model covariate effects on ratios of upper and lower conditional quantiles. The proposed estimator is based on semi-parametric estimation of conditional quantiles, which are then linked to covariates via a linear model. Dependence within subjects is accommodated by a ridge-type penalty on unit-specific intercepts, which shrinks individual effects while allowing for heterogeneity. We also introduce a computationally efficient one-step empirical Bayes procedure for selecting the penalty parameter. A simulation study under different dependence structures and outcome distributions shows that the ridge-regularized estimator reduces mean squared error and improves out-of-sample prediction relative to unpenalized and naïve alternatives. An application to a four-year panel of about twenty thousand European households illustrates how the method can be used to analyze income inequality through conditional income quantile ratios.
High-frequency pairs trading requires robust decisions on pair selection, rebalancing suitability, and portfolio weight allocation under volatile market conditions. To address this issue, this study proposes a multi-stage machine learning framework for portfolio selection and weight optimization in high-frequency crypto-asset pairs trading. This framework decomposes the allocation problem into economically meaningful stages and allows each decision layer to be modeled according to its own statistical structure. First, cointegration analysis is used to identify statistically meaningful asset pairs among Bitcoin-denominated crypto-assets. Second, machine learning models distinguish portfolios suitable for active rebalancing from those more consistent with buy-and-hold behavior. Finally, the optimal portfolio weight is predicted through quantile-based multi-class classification, where allocation classes are constructed according to information ratio-maximizing weights. Using minute-level data for 40 Bitcoin-denominated crypto-assets over the 2022–2025 period, the proposed approach is evaluated under hourly, daily, and weekly rebalancing frequencies. The empirical findings show that ensemble-based models, particularly Random Forest and gradient boosting methods, provide superior predictive performance across different classification settings. The results further reveal a trade-off between decision granularity and predictive accuracy, as well as a non-monotonic relationship between rebalancing frequency and model performance. The findings highlight the importance of aligning model complexity and rebalancing design with the structure of the allocation problem, and provide practical insights for the development of data-driven pairs trading strategies in high-frequency markets.
In probabilistic record linkage, analysts seek to link records in two data files using error-prone identifiers like names or demographic variables. One approach is to treat the linkage structure as a random variable in a Bayesian model specification. The estimation process typically results in multiple plausible linked files that can be used for downstream analyses or to estimate uncertainty in the linkages. Often, however, analysts are interested in a point estimate of the linkage structure. Typically, they use a Bayes estimator that, in the usual implementation, declares record pairs links only if the model assigns them a posterior probability of being true links that exceeds 0.5. This estimator can miss links that have high posterior probability but do not cross the majority threshold. We introduce a strategy to augment the Bayes estimator with the goal of catching some of these missed links. The basic idea is to specify a loss function for adding links to the Bayes estimator. The loss parameters can be tuned to be more or less aggressive in declaring pairs as links. We illustrate the strategy using a variety of simulation studies and discuss when it can be expected, or not, to enhance the Bayes estimator.
Many modern data sets contain sequences of short regression profiles collected over time, where departures from stability are small, persistent, and difficult to separate from routine variation. This paper proposes a mixed memory procedure for monitoring linear profiles, denoted MEC_3 , that jointly tracks the intercept, the slope, and the error variance in a model with a coded explanatory variable. The method combines exponential smoothing with simple one sided accumulation recursions, which improves sensitivity to gradual changes while keeping componentwise interpretation. The main contribution is a simulation-based design strategy that calibrates control limits to a target in control average run length using projected Robbins Monro updates. This calibration step is a stochastic approximation routine that works with Monte Carlo estimates and adapts the limits through stable, low-cost iterations, which makes the design practical when closed-form limits are unavailable. A unified scaling places all charting statistics on a common unit control limit, which simplifies comparison across components and across competing designs. Performance is studied using run length summaries and domain averaged efficiency criteria over ranges of location and scale shifts, with comparisons to standard profile charts based on Shewhart, exponentially weighted moving average, and cumulative sum ideas under matched false alarm protection. An interactive Shiny application (Application available at https://linearprofiles.shinyapps.io/MEC3/ ) implements the full workflow, including calibration, simulation experiments, and exportable reports for reproducible use. A data example with 225 profiles illustrates earlier signaling under small sustained departures while maintaining the intended false alarm control.
In this paper, we propose a distributed two-sample test of high-dimensional mean vectors based on the random integration of difference (RID) technique to handle the large-scale high-dimensional data. Compared with the full-sample test, the proposed distributed RID test overcomes storage cost and computing time limitations of a single machine, thereby enhancing testing efficiency. Furthermore, under some conditions, we establish the asymptotic properties of the distributed RID statistic and show that the power expressions for the distributed RID test and the full-sample RID test are asymptotically identical. In particular, the proposed distributed RID test is powerful when the non-zero signals have almost the same sign and are weakly dense, or when the non-zero signals are dense under the alternative hypothesis. Numerical simulations and a real-data analysis demonstrate the superior performance of the proposed method.
In this paper we present a novel approach for optimizing betting strategies in sports gambling and evaluate it using data from the English Premier League. Our model integrates deep learning techniques, combining neural network with portfolio optimization. We have achieved remarkable profits of 35.8
This paper develops a deterministic inference framework for asymmetric stochastic volatility models. The method combines a nonlinear quadrature filter and smoother (NQFS), based on Gauss-Hermite rules, with a generalised EM (GEM) algorithm for maximum-likelihood estimation. The latent state is propagated and smoothed using deterministic quadrature, and the resulting approximate Gaussian moments are used in closed-form E-steps for the ASV parameters. From a computational-statistics perspective, the NQFS–GEM framework provides a deterministic and reproducible alternative to particle-based and fully simulation-based approaches. A quadrature-point sensitivity study shows that only a small Gauss-Hermite grid (about five to seven nodes) is sufficient to stabilise the approximate likelihood, while the computational cost grows only linearly in the number of points. In a speed–accuracy comparison with quasi-maximum-likelihood (QML) and a simple particle filter smoother (PFS), NQFS is only modestly slower than QML in benign settings, yet it is much more stable and substantially more accurate than the PFS baseline. The framework is extended to handle both Gaussian and Student-t observation densities through a minor modification of the measurement update. On daily equity index returns, the likelihood strongly favours the Gaussian specification, and the parameter estimates are economically plausible with well-behaved residual diagnostics. Overall, the results indicate that the proposed NQFS–GEM framework can be used as a practical reference implementation for likelihood-based inference and model comparison in nonlinear state-space models of stochastic volatility.
Empirical studies in the social sciences often include covariates that are compositional in nature—vectors of shares that sum to a constant—alongside noncompositional covariates. A common approach to handling compositional data in linear regression is to omit one component to avoid perfect multicollinearity. I demonstrate why these coefficients (along with standard errors and t-statistics) can be highly sensitive to the choice of omitted category. Transforming the compositional data using additive logarithmic ratios (ALR) yields permutation-invariant regressions and permits counterfactual changes in the implied composition that remain within the simplex. Furthermore, I demonstrate that applying a simple scale factor to the ALR coefficients generates the coefficients and standard errors associated with isometric logarithmic ratios (ILR) for the variables of interest. Finally, using log-ratios does not exacerbate inherent problems of multicollinearity associated with compositional data. Economic growth regressions incorporating compositional and noncompositional covariates are used to illustrate.
Functional data have attracted considerable attention across a wide range of scientific and industrial applications, including economics, chemistry, climatology, and engineering. Among existing scalar-on-function regression methods, functional principal component regression (FPCR) provides a well-established and theoretically grounded framework, relying on a linear modeling structure in the principal component score space. In this paper, we propose Functional Kolmogorov–Arnold Networks (F-KAN), a nonlinear extension of the FPCR framework that integrates the Kolmogorov–Arnold network architecture with functional principal component representations. We further prove the consistency of the proposed F-KAN estimator. Furthermore, extensive simulation studies and experiments on four real-world spectroscopy datasets demonstrate that F-KAN achieves competitive predictive performance compared with a range of existing functional regression approaches. Beyond prediction accuracy, we explore the learned functional-response relationships and flexibility of the proposed model, assess its robustness under different noise levels, and quantify predictive uncertainty using an adaptive conformal prediction framework. Overall, the results suggest that F-KAN provides a flexible and theoretically grounded framework for modeling and exploring complex functional relationships in scalar-on-function regression.
In this paper, we consider the feature screening problem of ultra-high dimensional longitudinal heterogeneous data, which significantly extends the existing frameworks on ultra-high dimensional heterogeneous data and ultra-high dimensional longitudinal data focusing only on mean regression. A quantile adaptive feature screening approach is proposed by integrating independence screening with quadratic inference functions (QIF). This framework offers two distinctive features: (1) it takes into account the within-subject dependency and is more efficient than that of ignoring the correlation and assuming independence for each subject; (2) it allows the set of active variables to vary with different quantiles, thereby providing a more comprehensive description for the real data and flexibility to accommodate heterogeneity. The sure screening property is shown under some regularity conditions. Some simulation studies and a real data analysis are conducted to assess the effectiveness of the proposed screening method. The numerical results indicate that the proposed method is an effective tool to handle with the ultra-high dimensional longitudinal data.
Simultaneous estimation and variable selection become difficult in linear models when the design is high-dimensional and predictors are correlated. Although the Lasso and related shrinkage estimators control model complexity, their selection can be inconsistent under certain conditions. To address this, we propose the triple shrinkage adaptive GO estimator, which extends the GO framework with adaptive, coefficient-specific weights. This multi-level shrinkage produces flexible penalization and achieves oracle properties asymptotically, yielding performance comparable to methods that effectively know the true support. The new estimator preserves the grouping effect, a key characteristic of the adaptive ElasticNet, such that coefficients of highly correlated predictors are shrunk toward one another. An efficient algorithm compatible with existing Lasso solutions makes this estimator computationally viable. The proposed approach, therefore, offers a robust alternative for improving estimation accuracy and support recovery in sparse linear models with dependent predictors.
Bayesian analysis provides a robust way to incorporate prior knowledge into statistical models, but eliciting diverse subjective priors for parameters in the unit interval [0, 1] remains lacking. These priors are vital for modelling probabilities, proportions, and success rates in many real-world applications. Beta distribution, although popular for its simplicity and conjugacy to the binomial models, may not always capture the true prior beliefs in certain real-world applications. This paper explores eliciting some alternative informative priors for binomial sampling models, beyond the conventional usage of the Beta distribution. We have developed an interactive Shiny R application to support the elicitation process of 14 different prior distributions. This tool helps users to visualise, compare priors and see their characteristics and perform a full prior-to-posterior Bayesian analysis for binomial models. In three important examples, we demonstrate elicitation processes for estimating the presence of a plant species, Fritillaria meleagris, in a meadow ecosystem; the probability of Drosophila melanogaster egg hatching under mutation and of passive transfer success in neonatal foals.
Estimation of treatment effects is a central focus in epidemiological, economic, sociological, and related disciplines. The strong ignorability assumption is typically imposed to identify treatment effects. The conditional mean of potential outcomes can be derived by integrating quantile regression. In this paper, we propose using deep neural networks to approximate the nonparametric component in a partially linear quantile regression model and to estimate the propensity score model, thereby constructing a doubly robust estimator for the average treatment effect. We rigorously prove that the proposed estimator is consistent and asymptotically normal. Simulation studies demonstrate that when the potential outcome follows a linear model, the proposed estimator performs comparably to linear quantile regression estimators while outperforming other methods. For nonlinear potential outcome models with sufficient sample sizes, our estimator exhibits superior finite-sample performance compared to existing approaches. Finally, we apply the proposed method to analyze the impact of the presence or absence of rainfall on PM2.5 concentrations in a real-world environmental study.
In the classification task, high dimensionality and class imbalance pose significant challenges to enhancing both efficiency and accuracy. To tackle the high-dimensionality issue, we have devised the Outer-Product-of-Gradient (OPG) method for Sufficient Dimension Reduction (SDR), specifically utilizing a Kernel based on Random Forests (KeRF). Additionally, to address class imbalance, we integrate multiple class-imbalance learning techniques, including XGBoost, cost-sensitive SVM (CS-SVM), and fuzzy SVM for class imbalance learning (FSVM-CIL). Our experiments were conducted on datasets encompassing both balanced and unbalanced sample distributions. The results reveal that the OPG method utilizing KeRF compares favorably in computational speed to the original OPG method, which employs the straightforward ROT law for selecting the bandwidth of the Gaussian kernel. In terms of classification results, for balanced datasets, both OPG methods exhibit comparable performance and outperform other SDR methods. For imbalanced datasets, the OPG-KeRF method combined with class-imbalance learners achieves superior performance, attaining higher average PR-AUC values with smaller standard deviations compared to alternative SDR methods.
This study develops a Vector of Probability Density Functions (VPDF) framework and an automatic fuzzy clustering algorithm for image data. Unlike conventional PDF-based clustering methods that represent each object by a single PDF, the proposed approach treats each image as a vector-valued PDF object and reformulates the clustering process in the VPDF space, including similarity measurement, cluster representation, membership estimation, and convergence analysis. Image features are extracted and represented as VPDFs, which are then used to automatically determine the optimal number of clusters, identify cluster elements, and estimate the membership probability of each image. The proposed algorithm is rigorously formulated, and its convergence is analyzed at each phase. Experimental results on image datasets show that the proposed method outperforms several widely used clustering approaches, demonstrating its potential for image recognition and related real-world applications.
For data after the logarithmic transformation, we compared performances of two types of estimators. The transformed data (TD) estimator is simply the kernel density estimator (KDE) computed from the log-transformed data. The plug-in (PI) estimator uses KDE to replace the unknown density f of the original sample in the formula for the density g of the log-transformed sample. We developed the large sample theory and proposed practical implementations for the PI estimator. We found that a single-bandwidth PI estimator tends to produce density estimates that are bumpy in the right tail, whereas the adaptive PI estimators based on the estimated mean squared error (MSE) optimal bandwidth function overcome this disadvantage. Our best performing adaptive PI estimator is based on replacing g by the normal reference density. Such an estimator is referred to as the adaptive normal PI estimator. Even though the TD estimator is difficult to beat in the considered setting, still the adaptive normal PI estimator showed superior performance around the modes of g and, overall, in the multimodal settings.
We introduce BIG-SAVE and BIG-DR, novel sufficient dimension reduction methods specifically designed to handle massive datasets, that challenge conventional computation. Our approaches extend traditional sufficient dimension reduction by adopting a divide-and-conquer strategy: the full dataset is partitioned into manageable chunks, sufficient dimension reduction is performed independently on each subset, and the resulting subspaces are integrated using a proximity-based aggregation method. Unlike existing sufficient dimension reduction techniques, BIG-SAVE and BIG-DR are optimized for parallel computing and memory mapping, allowing them to process data that exceed the limits of available random access memory. Moreover, our approach ensures robust estimation stability across chunks and maintains exhaustiveness, ensuring full recovery of the central subspace under mild conditions. We demonstrate the practical performance of our proposed methods through extensive simulation studies and real-world data applications. Our results show that BIG-SAVE and BIG-DR achieve competitive estimation accuracy, significant computational efficiency, and scalability making them suitable for modern large-scale data analysis tasks in scientific and industrial domains.
In this article, we consider nonparametric estimation of the cumulative incidence function (CIF) for left-truncated and interval-censored competing risks (LT-ICC) data. We propose two estimators and show that they are asymptotically equivalent and consistent. The first estimator is the non-parametric maximum likelihood estimator (NPMLE) derived from the likelihood functions for LT-ICC data. In particular, we correctly formulate the innermost intervals for developing self-consistent equations and establish consistency of the NPMLE. The second estimator, called the pseudo-likelihood estimator (PLE), is obtained by implementing a suitable constraint on the procedure from Hudgens et al. (2001), leading to the asymptotic equivalence of the PLE and the NPMLE. Simulation studies show that both estimators perform well.
The goal of this paper is to propose supervised learning methods for three-dimensional landmark data. We introduce novel adaptations of Linear Discriminant Analysis (SLDA), Quadratic Discriminant Analysis (SQDA), and K-Nearest Neighbors (Shape-KNN), specifically modified to operate within the non-Euclidean geometry of Kendall’s pre-shape sphere and its corresponding tangent space. To systematically evaluate the performance of the proposed methods, we conducted Monte Carlo simulations, sampling data directly from the manifold using the von Mises-Fisher distribution. The numerical results reveal that SLDA demonstrates superior robustness and stability in high-dimensional, noisy scenarios, whereas SQDA is highly susceptible to the dimensionality when sample sizes are limited. Moreover, we evaluate the models on a real-world 3D dataset. In this application, SQDA achieved the highest accuracy, showcasing its ability to effectively model class-specific covariance structures. Ultimately, this study provides a comprehensive framework and practical guidelines for classifying 3D morphological data in statistical shape analysis.