
There has been limited research on confidence intervals for the relative difference in matched-pair designs. Previous studies have generally reported that the logarithmic transformation method (LCI) outperforms the Fieller method (FCI). However, the LCI is based on large-sample theory, and its reliability in small-sample studies has long been questioned. To address this issue, this paper introduces several new confidence interval estimators for the relative difference in matched-pair designs. We conducted Monte Carlo simulations to compare the proposed methods with the existing LCI and FCI methods, in terms of coverage probability and expected confidence width. The results show that the SPA method (the saddlepoint approximation method) significantly outperforms all other methods across all parameter configurations, particularly in small-sample or extreme settings, where it exhibits the narrowest intervals and most stable coverage. Nevertheless, while the SPA method performs well, it is computationally intensive and fails to provide a closed-form expression for the interval, thus limiting its practical applicability. The WSM (based on the Wilson score and MOVER) achieves a good balance between interval width and coverage; its performance closely approaches that of SPA when the sample size is large. Overall, the SPA method is recommended when computational resources permit and high precision is required, while the WSM method serves as an efficient and practical alternative, especially in large-sample applications. Finally, all methods are illustrated using two real examples.
This paper introduces an investment management strategy that integrates clustering techniques with advanced portfolio optimization models to tackle the limitations of traditional approaches in high-dimensional and highly correlated financial environments. The proposed framework consists of two key components: a clustering algorithm that defines stock similarity based on autocorrelation coefficients, and a general expected-utility portfolio optimization model equipped with an elastic-net regularizer that combines l1 and l2 norms to promote sparse yet stable allocations in high-dimensional and highly correlated settings. Applying the clustering algorithm, the framework groups stocks into clusters with low internal similarity, thereby increasing portfolio diversity. To reduce portfolio complexity, two decision rules are introduced to select representative stocks from each cluster, ensuring that the final investment portfolio remains compact without sacrificing heterogeneity. We employ a shrinkage-based optimization approach for portfolio weight allocation using a linear combination of l1 and l2 norms. This design eliminates the reliance on direct estimation of the covariance matrix, which is often unstable in sparse settings, and allows the model to accommodate strong inter-stock correlations better. Empirical studies conducted on historical S&P 500 data demonstrate that the proposed method enhances average returns and improves the overall stability and robustness of the investment portfolio, outperforming both the unclustered baseline and other benchmark portfolio strategies.
The reinforcing efficacy of a product, characterized by its demand curve, is an important construct for understanding a product's potential to cause addiction. Hypothetical purchase tasks (HPTs) are used to measure demand for addictive products because they are cost-efficient, ethical, and safe to administer. However, statistical methods used to estimate demand curves from HPTs are often flawed. We introduce a Bayesian hierarchical modeling strategy for estimating the intensity and elasticity of demand from HPT data collected in an experimental setting. The approach addresses several recurring challenges in the analysis of HPT data by accommodating count outcomes, repeated measures, and individual-level heterogeneity, while allowing demand curves and treatment effects to be estimated jointly within a single inferential framework. When applicable, the hierarchical model allows us to simultaneously estimate baseline demand quantities and include them as covariates to better inform the relationship of treatments to demand. Further, uncertainty in the span of the population demand curve is explicitly modeled, rather than treated as a fixed quantity. We fit the model using Bayesian techniques supported by simulations. The approach is applied to data from a randomized crossover study comparing the intensity and elasticity of demand for flavored little cigars/cigarillos (LCCs) to unflavored LCCs.
We establish a novel framework for minimum deviation portfolio optimization by directly connecting a risk-averse stochastic problem (RASP) to linear regression models. By leveraging a set of coherent risk measures and scoring functions, our methodology generalizes classical models while enabling alternative formulations that account for tail risk and asymmetric scoring functions. Using S&P 100 stock data, we empirically illustrate our approach. Nonparametric hypothesis tests indicate significant differences in risk and risk-adjusted performance across RASP specifications.
Ecological count data often contain structural zeros and positive-count distributions that are more concentrated than the Poisson benchmark. In such settings, standard Poisson and negative-binomial regressions may be too restrictive, especially when underdispersion persists after accounting for covariates and effort. We propose a likelihood-based framework that combines a hurdle specification for structural zeros with a grouped-Poisson mechanism for the positive counts. The positive-count component is modeled by a shared-intensity Poisson- Qm mixture, in which both components depend on the same intensity parameter and a mixing weight controls the contribution of the grouped component. This yields an explicit finite-sum likelihood while preserving a clear intensity interpretation. The grouping index m is treated as a discrete model index and selected over a finite candidate set by information criteria or held-out log-score. We also place the total-count model within a sum-and-split framework for hierarchical taxonomic counts. We establish targeted theoretical properties, including local identifiability for fixed m, aggregation coherence of the split layer, and consistency of BIC-based selection over finite candidate sets. Simulations and a NEON small-mammal illustration show that the proposed extension is conservative in near-Poisson regimes and useful when the positive-count distribution is genuinely underdispersed.
By combining Tiku's modified maximum likelihood (MML) procedure with Welch's (W) test, the parametric bootstrap (PB) test, and the generalized F (GF) test, three different robust tests are proposed. These tests for the equality of group means are used when the underlying distribution is long-tailed symmetric (LTS) and the variances are unknown and arbitrary. The proposed tests are compared with the traditional W, PB, and GF tests based on least squares (LS) estimators in terms of their Type I error rates and powers via an extensive Monte Carlo simulation study. The simulation results indicate that the proposed tests generally perform better than the existing tests. In addition to evaluating their Type I error rates and powers, the robustness properties of the proposed tests against deviations from the assumed distribution are also investigated. Moreover, an R package called RobustANOVA, which includes the implementation of the proposed estimators and test statistics, has been made available on CRAN.
We present a Bayesian vector autoregressive (BVAR) model designed for panel data. In small-area applications, traditional vector autoregressive (VAR) models quickly become overparameterized, and standard BVARs often rely on aggressive global shrinkage, risking over-shrinking meaningful regional dynamics. To address these challenges, we develop a spatial BVAR-CAR framework that combines global regularization with flexible local shrinkage. We introduce a conditional autoregressive prior on region-specific intercepts to capture spatial dependance and a hierarchical shrinkage (horseshoe-type) prior on autoregressive coefficients to borrow strength across regions, stabilizing estimation in high-dimensional settings. This setup eliminates redundant parameters while retaining heterogeneous dynamics across local areas. We evaluate forecasting performance using two annual panel datasets: average hourly earnings of production employes in California Metropolitan Statistical Areas (small areas) and gender unemployment gaps in African countries. Our BVAR-CAR model outperforms three benchmarks - a univariate AR(1) model, a restricted VAR model with shared hyperparameters, and an unrestricted BVAR without spatial and global-local shrinkage priors - in both panels. The results highlight the benefits of spatial pooling in Bayesian models when time series are short. By integrating cross-sectional structure with temporal dependance, our approach provides a flexible and interpretable solution for forecasting and small-area estimation in regional economic analysis.
This study contributes to the ongoing evaluation of heavy metal contamination, specifically arsenic, mercury, chromium, lead, and cadmium concentrations in the soil. The article focuses on developing a suitable methodology for comparing random samples to monitor the occurrence of heavy metals over three years. This need arises from the impossibility of applying classical statistical tests due to unfulfilled assumptions for their use. The results provide important insights for environmental scientists tracking hazardous polymetals, but also reveal challenges in using standard statistical methods for this type of data. In Chile, monitoring heavy metal contamination in soil is crucial, also due to extensive mining activities. This study examines the changes in soil contamination by chemical elements in Arica, Chile in years 2013, 2014, and 2015. We conducted a detailed analysis of individual chemical elements to compare the levels of arsenic, mercury, chromium, lead, and cadmium across the three consecutive years. Our findings highlight the heavy-tailed distribution of several contaminants and offer valuable insights for comparing complex data sets using administrative data. Understanding homoscedasticity is crucial for testing mean equality in datasets and analyzing soil contamination. It ensures reliability and accuracy of statistical models, leading to more precise monitoring of contamination development.
This paper provides a new method to construct and harmonize partial seriations from binary matrices. There have been many methods proposed to perform a seriation on a matrix of binary data, in which rows and columns are permuted to arrive at an optimal ordering of the matrix elements. However, no one method is assured to produce the best result, as routines are highly susceptible to the particular distribution of 1s in the matrix. In certain cases, the omission of rows and columns may yield in an improved seriation, but this leaves investigators with the tedious job of slicing matrices and re-running analyses. Working with several partial seriations raises the problem of how to combine such 'strands' (subsets of the data that have been seriated) to obtain a consensus out of potentially discrepant orderings. This paper therefore proposes an agglomerative routine involving simple linear regression, called Lakhesis. According to optimality measures, Lakhesis has the capacity to outperform conventional approaches. The R lakhesis package also provides a graphical interface. Investigators can explore binary matrices and seriate partial subsets using correspondence analysis (both simple CA and Procrustes-fit CA). The partial seriations are then 'lakhesized' into a single consensus seriation.
We employed the Least absolute shrinkage and selection operator (Lasso) for shrinkage estimation and variable selection, but it is sensitive to outliers. On the other hand, the Least absolute deviation (LAD) regression is robust but lacks uniqueness in high-dimensional modeling. Therefore, there is a clear need for robust regression methods tailored to high-dimensional data. One effective solution is to combine LAD regression with Lasso methods, creating the LAD-Lasso regression. Recent research has shown that the Liu-Lasso estimator outperforms LASSO, as it integrates the penalization functions of the Liu estimator with those of Lasso. Studies indicate that Liu-Lasso automatically selects variables and demonstrates significant predictive performance with minimal mean squared error across sparse and non-sparse data structures. Building on this, we propose a robust version of Liu-Lasso called LAD-Liu-Lasso, inspired by the LAD-Lasso approach. We compared LAD-Liu-Lasso with the Least Squares method, Liu estimator, LAD regression, and LAD-Lasso. Additionally, LAD-Liu-Lasso benefits from easily estimated tuning parameters and maintains the same asymptotic efficiency as the unpenalized LAD estimator. Extensive simulation studies confirm the satisfactory finite-sample performance of LAD-Liu-Lasso, and we provide two real examples for illustration.
In order to deal with the inference about complex data in which the variables present high dimension and high correlations within blocks, the response variable presents skewness and long tails. We propose a Bayesian latent factor analysis for quantile (BFQ) to model these datasets. Our model incorporates the shared latent factors in the explanatory variables with high correlations and the response variable. We employ the asymmetric Laplace distribution(ALD) to formulate the quantile in the response variable model. In order to fit nonlinear effects of some explanatory variables to the response variable, the single-index formulation is used. Simulation results indicate that our BFQ model outperforms traditional models when dealing with non-normally distributed data. Furthermore, from the real data application, we can see our BFQ can fit breast cancer dataset well which is from the Surveillance, Epidemiology, and End Results Program(SEER).
Reference intervals are data-based intervals that contain a pre-specified proportion of the healthy population. These intervals are known to be among the most widely used decision-making tools in laboratory medicine. These intervals are obtained by taking the central 95% part of the measurements of the reference population. Thus, the 2.5th and 97.5th percentiles of the population are the endpoints of the interval. Since these population percentiles are unknown in practice, the reference interval is estimated from available data, either from a single study or preferably from a meta-analysis of multiple studies. In meta-analysis studies, the reference interval is typically estimated by taking the confidence interval for the pooled mean. However, this approach does not incorporate the natural variation across different studies. Thus we develop a methodology to compute the central tolerance interval, which is then used as the estimate for the reference interval in a one-way random effects model. This approach reflects both sampling variation and the variability in the distribution of a new subject. Numerical results show that the proposed methodology is accurate, with coverage probabilities close to the nominal value. We apply the proposed methods to real-life data from a study to compute reference ranges for wake-after-sleep onset.