
ABSTRACT This paper investigates stochastic versions of the classical Bass diffusion model, with a focus on the development and comparison of parameter estimation procedures. Two stochastic processes that generalize the classical Bass model, which is commonly used to describe sales dynamics in emerging markets, are considered. For these processes, parameter estimation procedures that identify the potential market size, the innovation and imitation rates, and the amplitude of random fluctuations of the phenomenon under study are developed. The first estimation procedure is based on a discretization of the original process and employs a quasi‐maximum likelihood approach. The second procedure relies on a new parametrization of the Bass curve and leads to a maximum likelihood estimation method. An extensive simulation study is presented to assess the statistical accuracy and robustness of the proposed procedures and to provide a comparative evaluation of their performance under different scenarios. Finally, the applicability of the proposed inferential framework is illustrated through an empirical study on mobile social networking adoption.
Detecting structural changes in multivariate functional data is difficult when strong dependence and non-linear relationships obscure persistent shifts. Existing statistical process control and multivariate functional monitoring methods often lose power in such settings due to variance inflation induced by cross-correlation. This study focuses on improving practical changepoint detection through dependence-aware preprocessing within a functional monitoring framework. We propose a change point detection method that integrates nonparametric conditional distribution transformations with functional principal component analysis (FPCA). The conditional transformation reduces inter-variable dependence by standardizing each functional component relative to its conditional distribution, allowing localized dependence changes to emerge more clearly after common variation is removed. FPCA then captures dominant temporal variation, enabling sustained mean and variance shifts to be detected through low-dimensional score processes. Method effectiveness is assessed through simulation studies with controlled change points under varying dependence structures. Performance is evaluated using detection delay, false alarm rates, and localization accuracy, and compared against established multivariate functional monitoring methods without conditional preprocessing. Both simulations and financial applications show that the conditional framework detects structural shifts missed by conventional marginal monitoring approaches, particularly in strongly correlated environments. The proposed approach consistently achieves faster detection and improved stability in highly dependent settings. An application to high-frequency financial index data further demonstrates that the framework can identify stock-specific regime changes obscured by broader market movements.
Reliability practitioners are very interested in which factors affect product reliability. Constrained by time and cost, they hope to terminate the experiment early without affecting the identification results of important factors. In design of experiments (DOE), practitioners usually construct the relationship between parameter and factor effects. However, the true model is usually complicated. Parametric models are often mis-specified. In this article, we proposed a semiparametric framework to terminate the experiment and identify important factors. First, we introduce a loss function to describe the total loss of the manufacturers and the customers, then the response is the type of smaller-the-better. Second, we assume the lifetime data follow a Weibull distribution and get the estimates of parameters using the Bayesian method, then we compute the total loss. Next, we calculate the sum of squares (SS) gotten from different experimental levels and terminate the experiment if all the weighted relative rates of changes of SS are less than a threshold value. Finally, the classification and regression tree (CART) method is introduced to drop some unimportant factors and identify important factors through the analysis of variance (ANOVA). Additionally, we confirm the proposed method through a real example and find that the important factors obtained from the proposed terminated experiment are the same as those found from the original terminated experiment.
Integrating energy islands into the European electricity market is a key challenge for the energy transition. This study investigates the impact of the 'Sorgente-Rizziconi' interconnector, activated on 28 May 2016, on electricity price volatility in the Sicilian market zone, which had previously been fully isolated from mainland Italy. Using daily data from 2015 to 2018, the analysis applies a semi-parametric GARCH model with a logistic intervention function to estimate changes in conditional price variance. A fully non-parametric additive model is employed as a robustness check, allowing the data to shape volatility dynamics without imposing a predefined structure. The results reveal that the interconnector significantly increased price volatility in Sicily, without reducing average price levels. No significant effects were observed in other Italian market zones. These findings highlight the context-dependent nature of infrastructure impacts and suggest that physical integration alone does not guarantee price stability. The results have important implications for energy policy, investment planning and risk management in electricity markets.
Accurately forecasting how innovations are adopted is crucial for launching new products and long-term performance. While traditional diffusion models like the Bass model have been widely used to study adoption trends, they often assume deterministic behavior and overlook the randomness inherent in real-world markets. Stochastic differential equation (SDE)-based models offer a more flexible framework by incorporating random fluctuations, typically through additive noise. However, these models rarely account for multiplicative noise, where the intensity of uncertainty increases with the number of adopters. To address this gap, we propose two SDE-based innovation diffusion models that explicitly include multiplicative noise and consider constant and logistic time-dependent adoption rates. These models are evaluated using real-world sales data for technological products, and their performance is compared using established metrics. By incorporating multiplicative uncertainty, the proposed models offer a more realistic representation of adoption dynamics, making them valuable tools for understanding diffusion in uncertain and rapidly evolving markets.
Consider a financial or insurance system with a finite number of individual components. The Tail Moment (TM) risk measure is defined as the conditional moment of some individual risk given that a system crisis occurs, and it is one of the systemic risk measures that have garnered significant scholarly attention in recent years. This paper is devoted to study the asymptotic behavior of the TM under a general framework, in which the individual risks are modelled by real-valued random variables and they follow gamma-like distributions with scale parameters not necessarily the same. Some precise asymptotic estimates of the TM are obtained when the individual risks are mutually independent or have a dependence structure of the Farlie-Gumbel-Morgenstern (FGM) type. When the system contains two individual risks with the FGM joint distribution, we discuss in detail the critical case of extremely negative dependence, which is rarely considered in the existing works. Moreover, a numerical study is given to illustrate the approximation performance of our main asymptotic results. We also provide an application in capital allocation based on real insurance data.
Based on the work of Choi et al. (2023) and Han et al. (2023), this paper develops an insider trading model with trading constraints in a transparent market, where regulatory securities laws mandate public disclosure of insider trading volume. Utilizing dynamic programming principles and filtering theory, we derive a closed-form market equilibrium characterized by a mixed trading strategy and distinct market prices in the trading and updating phases. Our analysis demonstrates that, in contrast to an opaque market, at the equilibrium of this transparent market, disclosure rules and the mixed strategy adopted by the insider adding a noise term (with volatility consistent with that of noise traders) together reduce market liquidity, slow down the speed of information disclosure, and result in a higher remaining variance that decreases gradually. Meanwhile, this mechanism significantly compresses insiders' profits and effectively protects the interests of noise traders.
Bootstrapping is a statistical inference method based on resampling a data set, with replacement. Befitting bootstrap analysis is a method of bootstrapping data that accounts for the data generation process. In this paper, we examine with case studies the effect of BBA in unbalanced data where the data generation is producing groups of unequal size. We provide a sensitivity analysis of bootstrapping and befitting bootstrap analysis and show how befitting bootstrap analysis is affected by imbalance, large or small. The Python code used in the analysis is available in a GitHub repository. The first section is an introduction. Section 2 is a methodological background on bootstrap analysis. Section 3 reviews befitting bootstrap analysis and befitting cross validation used in predictive analytics. In Section 4 there are two case studies based on a simulation and a dental adhesive strength study. Section 5 concludes the paper with a discussion and lessons learned from the case studies.
As the oil and gas industry faces increasing scrutiny over its climate impact, it becomes essential to adopt effective strategies to monitor and reduce greenhouse gas (GHG) emissions. During the extraction phase of hydrocarbons, the generation of energy by burning gas is the primary contributor to emissions. In this paper, we propose a data-driven methodology for estimating fuel gas in relation to production at treatment plant level where the extraction process takes place. The proposed approach is designed with a pragmatic perspective, considering both the industrial setting and the constraints imposed by available empirical data. Given that this analysis relies on administrative data, extensive preprocessing has been implemented to effectively analyze and model the phenomenon. To enhance the analysis and identify key variables, various clustering techniques were used to group treatment plants exhibiting similar behavior patterns. Despite the comprehensive preliminary analysis, inherent challenges persisted, including the presence of highly correlated numerical variables, which resulted in outcomes that were misaligned with real-world phenomena. In order to address these issues, Principal Component Analysis (PCA) was adopted to mitigate the effects of confounding variables. This approach, combined with an unsupervised random forest algorithm, facilitated the categorization into four distinct clusters. These clustered observations were then used for a split-panel regression analysis.
Regular two-level fractional factorial designs are widely used for factor screening experiments, where the objective is to efficiently identify the set of active factors from a larger initial group of factors. The 16-run designs are very popular for screening because they can accommodate a reasonably large number of factors, and for 6-8 factors, they are Resolution IV, while for larger numbers of factors they are Resolution III. Assuming that 3-factor and higher-order interactions are negligible, the Resolution IV designs provide a clear estimate of the main effects while aliasing all 2-factor interactions with each other, and the Resolution III designs alias main effects and 2-factor interactions. Because of the aliasing, follow-up experiments are often required to obtain complete information about main effects and 2-factor interactions. However, there are many situations where follow-up experimentation isn't possible. Nonregular fractional designs that do not have complete aliasing involving main effects and 2-factor interactions can be a good alternative for these situations. However, analysis methods for these designs are an ongoing area of research. We investigate analysis methods for a class of non-regular 2-level fractional factorials for 9-14 factors in 16 runs. In these designs, there is no complete aliasing between the main factors and the two-factor interactions, so these designs are useful alternatives to the regular Resolution III fractions. The analysis methods we consider are forward stepwise regression, the least absolute shrinkage and selection operator (LASSO), and the Dantzig selector method. We show that in most cases, for large and medium underlying model regression coefficients, stepwise regression and the LASSO outperform the Dantzig selector in correctly identifying the set of active factors for situations where the number of active factors does not exceed approximately half of the number of degrees of freedom for the design.
This study provides a comprehensive bibliometric analysis of the development of Explainable Artificial Intelligence (XAI) research from 1993 to 2024. The objective is to explore key contributors, thematic trends, and the evolution of methodologies within the field. By employing network analysis, Multiple Correspondence Analysis, and co-citation techniques, the study identifies major research clusters, global collaboration patterns, and the most influential keywords. The results highlight the dual focus of XAI research: the development of technical methods to improve model interpretability and their applications across diverse domains such as healthcare, risk management, and climate science. Furthermore, the historiographic network captures the progression of XAI from foundational concepts to specialized applications, emphasizing its interdisciplinary growth. This analysis offers valuable insights into the trajectory of XAI research, aiming to guide future advancements and promoting further collaboration in this critical area.
A Bayesian framework quantifies market efficiency under -stable return distributions. Measure-theoretic foundations establish the existence and regularity of hierarchical posteriors for . A predictive mutual information score is introduced, satisfying convexity and diffeomorphism invariance under regular variation. Geometric ergodicity is established for Metropolis-within-Gibbs samplers. These samplers target -stable posteriors, with effective sample size bounds extended to heavy-tailed targets. When applied to major financial indices from 2008 to 2023, the framework discriminates heavy-tailed assets from predictable light-tailed ones. It also detects efficiency breakdowns during financial crises.
In this article, we present a dynamic version of the integer autoregressive (INAR) processes for count data. The proposed Bayesian model provides a unification of the previously considered models to describe temporal correlations in univariate time series of counts. We develop Bayesian inference for the proposed class of models via MCMC and introduce a particle filtering (PF) algorithm for sequential inference. The new class of models are compared with their static counterparts using actual count series and additional insights provided by the new models are discussed.
This paper contributes to the existing literature on cluster analysis presenting a new method to identify the best partition, that is, the optimal number of clusters, when the -means clustering algorithm is adopted. Clustering is an unsupervised learning technique aimed at constructing partitions such that items within the same cluster are similar, while those in different clusters exhibit distinct differences. As an exploratory analysis, the effectiveness of clustering algorithms, particularly non-hierarchical ones, depends heavily on key decisions, such as determining the number of clusters in the final partition. In the literature, various cluster validity indices are available to assist in identifying the optimal partition. However, these indices often produce conflicting recommendations, which may not provide clear guidance to researchers and practitioners. To address this gap, the Multivariate Permutation Cluster Validity (MPCV) test has been introduced to compare partitions across multiple clustering performance metrics. This approach synthesizes information from widely used clustering indices, including the Silhouette, Dunn, Calinski-Harabasz, and Davies-Bouldin indices, aiming to achieve an optimal balance between internal homogeneity and external separability. Our approach is presented and discussed using both simulated and real data, highlighting its main advantages.
The reliability and behavior of coherent systems have been widely explored in the literature. In this study, we focus on the problem of parameter estimation for a coherent system whose exact configuration is unknown. We assume that component lifetimes are independent and identically distributed, following a Weibull distribution. The available data consist of the system lifetimes and the number of component failures observed at the time of each system failure. We derive maximum likelihood estimators (MLEs) for the model parameters and, using Fisher's conditionality principle, establish their conditional consistency and asymptotic normality, including the joint (bivariate) asymptotic distribution. The performance of the proposed estimators is assessed through simulation studies and further demonstrated using a real-world data set on aircraft air conditioning systems modeled with a 2-out-of-4 reliability structure. Both conditional and unconditional approaches are evaluated, and model adequacy is supported through goodness-of-fit tests.
Traditional short-rate models introduce volatility directly into the instantaneous rate via Brownian shocks. However, empirical data suggest that short-term interest rates exhibit smoother behavior than such models imply. We propose a two-factor Gaussian short-rate model in which the short rate is a deterministic exponential filter of a stochastic mean-reverting latent mean. The randomness enters exclusively through a latent Ornstein-Uhlenbeck process representing long-term expectations, yielding a two-dimensional affine Markov structure in the state variables, while the short rate itself admits a non-Markovian Volterra-type representation. A closed-form exponential-affine characteristic function for the integrated rate is derived, allowing for analytical bond pricing and Fourier-based derivative valuation. Empirical results on U.S. Treasury data show that the model outperforms the Vasicek benchmark in both in-sample fit and out-of-sample forecasting, with a reduction in RMSE. The model yields stable and economically interpretable parameters while requiring only a single source of randomness. Its mathematical tractability offers a compelling alternative for interest rate modeling, monetary policy analysis, and risk management.
Graphs and charts have been used for hundreds of years for data visualization, perhaps thousands of years if we count pictographs and symbols in early written languages. Creating these graphs and charts was very time-consuming when done by hand, and journals and book publishers were reluctant to include them because of the cost. Text and tables were far easier to print. This mindset began to change about 50 years ago with the advent of computer software that made creating graphs and charts easier, as well as a series of books by Edward Tufte, John Tukey, and others that explained how the exploration of large amounts of data benefited from visualizations. Over the past decade, the development of new software and applications has enabled dynamic, interactive data visualizations, making sophisticated presentations accessible to modern researchers across various fields. No longer are these command-and-control centers with dynamic interactive displays limited to well-funded organizations such as nuclear power plants, chemical company dashboards, and NASA; they are now becoming more commonplace on websites, smartphone displays, automobile dashboards, and many other business and industrial applications. In this paper, we provide numerous examples, guidelines for creating dynamic and interactive data visualizations, and offer some cautions regarding data quality.
Statistical process control (SPC) charts are widely used in various fields to detect distributional changes in sequential processes. Traditional SPC charts are primarily designed to identify abrupt changes in process parameters, such as sudden shifts in the mean or variance. However, many real-world applications involve gradual changes over time, commonly referred to as drifts. This paper develops three change-point detection charts for identifying linear drifts in the mean vector, the covariance matrix, and both simultaneously in a multivariate process. The proposed charts are constructed based on the generalized likelihood ratio statistic and change-point detection techniques. These methods do not require pre-specification of procedure parameters and provide an estimate of the change-point location once a signal is given. Numerical studies demonstrate that the proposed charts are more effective in detecting linear drifts in the process mean and/or covariance matrix compared to conventional control charts designed for detecting abrupt changes.