Latent Class Analysis (LCA) is widely used to identify unobserved subgroups in social and behavioural sciences. A long-standing challenge for LCA is the interpretability of the latent classes, due to the high complexity of the estimated item response probability matrix. To address this, we propose a computationally efficient post-estimation refinement procedure that enhances model interpretability by a sparse model estimate. The method begins by estimating a classical, unrestricted, latent class model and determining the number of classes using the Bayesian information criterion (BIC). It is followed by a refinement step that further performs model selection on the item-specific response probabilities based on the initial estimate. This refinement penalises the number of distinct response probability levels per item, collapsing redundant levels to yield a sparse matrix that is significantly easier to interpret than those produced by classical LCA. We provide asymptotic theory showing that the proposed procedure consistently recovers the sparse pattern of the item response probabilities for each item, and further validate its performance through extensive simulations. The practical power of the proposed method is further illustrated via an application to survey data on social role performance, where it provides a parsimonious and clear characterisation of the resulting latent classes. The code for implementing the proposed method is publicly available at https://github.com/florence07/Sparse-LCA-Refinement.
The weighted cumulative information generating function (wcigf) is introduced and interrelations with various weighted and non-weighted measures of uncertainty are explored in detail. It turns out that the wcigf represents the generating function of several weighted cumulative versions of the Shannon entropy. Additionally, several properties and representations of wcigf are shown, and connections with order statistics are studied. Further applications arise by defining weighted efficiency Gini-based indices for systems in terms of ratios of wcigf’s.
A heterogeneous step-stress accelerated life testing (hSSALT) model for Type-I censored data is developed, assuming exponential lifetimes and allowing for inhomogeneous aging patterns among testing items. The underlying heterogeneity is captured using a mixture model approach, in which group memberships are unknown. The model is formulated under the cumulative exposure assumption, while both continuous and interval monitoring schemes are considered. In the interval-monitoring setting, exact failure times are unobserved; instead, only the numbers of failures occurring within the time intervals defined by predetermined inspection points are available. The inspection points include the stress-level change points as well as additional intermediate time points. An expectation–maximization (EM) algorithm is adapted for maximum likelihood estimation of the model parameters for both monitoring schemes. Asymptotic and bootstrap confidence intervals are also derived. Numerical examples demonstrate the performance of the proposed hSSALT model across different scenarios and confirm its validity when inspection intervals of appropriate length are used.
Indirect Hard Modeling (IHM) is a physics-based evaluation method for the quantitative analysis of fluid compositions using spectroscopic techniques such as Raman spectroscopy. In this approach, mixture spectra are represented as a superposition of pure substance models, with each component described by a sum of parameterized peak functions. Nevertheless, the accuracy of the compositions prediction depends critically on user decisions regarding both the number of peak functions and the specific parameter adjustments employed. In this work, we apply an expectation-maximization (EM) based algorithm for generating spectral reconstructions of pure substance models that does not require the pre-specification of the number of peaks or any initial values. The efficient and fast performance of the used EM algorithm enables the fit of a given spectrum for an unknown number of peaks, based on a model selection criterion. In simulation studies, we demonstrate that this approach can recognize the true underlying function in settings of high noise, peak overlapping and background signals, yielding reliable results. In a validation study, the algorithm was tested using experimental data. It was integrated into an Indirect Hard Modeling framework and applied to three chemical test systems. The quality of the obtained results were in the range of other automated IHM model generating approaches while significantly reducing both time and computational effort.
In this paper, we propose an automatic, user-independent algorithm based on an expectation-maximization (EM) approach to recover the structure of spectral data emerging from Raman spectroscopy - a well-established method used, among others, for the identification of substances in materials. More precisely, the goal of this work is to represent a given Raman signal through a suitable statistical model to identify the unknown substance from which this signal emerges from. Commonly used techniques in the field of Raman spectroscopy are based on least-squares estimation and they highly depend on the initial values specified by the user, leading to inconsistent and non-reproducible results. In contrast, the presented EM algorithm does not require any user-specified initial values, as a peak birth strategy is used, which effectively resolves these issues. Furthermore, the presented approach enables the fit of a given spectrum for an unknown number of underlying peaks by employing the BIC model selection criterion. The accuracy and robustness of the proposed algorithm is demonstrated in simulation studies considering different settings, such as high noise, strong baselines and low data availability.
Common step-stress accelerated life testing (SSALT) models assume that all testing items are sampled from a homogeneous population. However, this is often not the case in practice. Practitioners observe inhomogeneous aging patterns among items of the same production batch. This work proposes a simple SSALT model with exponentially distributed, Type II censored lifetimes that accounts for underlying heterogeneity in aging. To capture the inhomogeneity, a mixture model is introduced and an expectation-maximization (EM) algorithm for censored data is constructed for approximating the maximum likelihood estimates of the model's parameters. The validity of the suggested model and its advantage over the SSALT model in the presence of heterogeneity are demonstrated via simulation studies. Additionally, the log-link function, used to extrapolate the inferential results to normal operating condition (NOC), is adjusted to accommodate the heterogeneous setup. For log-link models, it is demonstrated that in presence of heterogeneity, the common model always overestimates the lifetime under NOC. In contrast, the proposed model, accounting for heterogeneity, reduces the bias in estimation and extrapolation.
The concepts of entropy and divergence, along with their past, residual, and interval variants are revisited in a reliability theory context and generalized families of them that are based on phi-functions are discussed. Special emphasis is given in the parametric family of entropies and divergences of Cressie and Read. For non-negative and absolutely continuous random variables, the dual to Shannon entropy measure of uncertainty, the extropy, is considered and its link to a specific member of the phi-entropies family is shown. A number of examples demonstrate the implementation of the generalized entropies and divergences, exhibiting their utility.
We present a Bayesian approach for the analysis of rating data when a scaling component is taken into account, thus incorporating a specific form of heteroskedasticity. Model-based probability effect measures for comparing distributions of several groups, adjusted for explanatory variables affecting both location and scale components, are proposed. Markov Chain Monte Carlo techniques are implemented to obtain parameter estimates of the fitted model and the associated effect measures. An analysis on students’ evaluation of a university curriculum counselling service is carried out to assess the performance of the method and demonstrate its valuable support for the decision-making process.
For the battery industry, quick determination of the ageing behaviour of lithium-ion batteries is important both for the evaluation of existing designs as well as for R&D on future technologies. However, the target battery lifetime is 8–10 years, which implies low ageing rates that lead to an unacceptably long ageing test duration under real operation conditions. Therefore, ageing characterisation tests need to be accelerated to obtain ageing patterns in a period ranging from a few weeks to a few months. Known strategies, such as increasing the severity of stress factors, for example, temperature, current, and taking measurements with particularly high precision, need care in application to achieve meaningful results. We observe that this challenge does not receive enough attention in typical ageing studies. Therefore, this review introduces the definition and challenge of accelerated ageing along existing methods to accelerate the characterisation of battery ageing and lifetime modelling. We systematically discuss approaches along the existing literature. In this context, several test conditions and feasible acceleration strategies are highlighted, and the underlying modelling and statistical perspective is provided. This makes the review valuable for all who set up ageing tests, interpret ageing data, or rely on ageing data to predict battery lifetime.
A maximum likelihood (ML) method is known to produce inconsistent estimators if the likelihood function is unbounded from above, e.g., in models with heavy-tailed distribution or unknown shift parameter. Such examples can be found in accelerated life testing (ALT) experiments, commonly applied in reliability studies on extremely durable products. In such cases, the maximum product of spacings (MPS) approach can be used as a more robust alternative that leads to asymptotically efficient estimators. Here, the MPS method is adjusted for Type-I censored samples in order to address complications that may arise in estimation. Moreover, the asymptotic theory of MPS estimators is adapted for the framework of Type-I censored simple step-stress ALT (SSALT) experiments. As an application, a failure-rate-based simple SSALT model with Weibull lifetimes, sharing a common shape parameter on both the stress levels, is considered. It is shown that the MPS estimator exists in situations where the ML fails to produce parameter estimations. Furthermore, the ML and MPS approaches are compared via a simulation study and applied to a real-life data example.
Step-stress is a special type of accelerated life-testing procedure that allows the experimenter to test the units of interest under various stress conditions changed (usually increased) at different intermediate time points. In this paper, we study the problem of testing hypothesis for the scale parameter of a simple step-stress model with exponential lifetimes and under Type-II censoring. We consider several modifications of the log-likelihood ratio statistic and eliminate the distributional dependence on the unknown lifetime parameters by exploiting the scale invariant properties of the normalized failure spacings. The presented results and the ratio statistic are further generalized to the multilevel step-stress case under the log-link assumption. We compare the power performance of the proposed tests via Monte Carlo simulations. As an illustration, the described procedures are applied to a real data example from the literature.