Fuzzy random variables were introduced to model uncertain quantities being simultaneously random and imprecise. Convergence of sequences and series of random variables is an important issue within theoretical considerations and practical applications. For fuzzy random variables, several types of convergence have been defined. In this paper, we focus on the almost sure convergence, the convergence in probability, and the convergence in distribution of series of independent fuzzy random variables taking values in the space F-cocp(no)(R-d) of the fuzzy subsets of R-d with nonvoid convex compact alpha-cuts and support functions being integrable of order p for a fixed positive integer d and p >= 1. For these series, we formulate and prove a counterpart of the It & ocirc;-Nisio Theorem, characterizing convergence in a separable Banach space. In the case of p= 2, we also obtain a fuzzy counterpart of the Three Series Theorem. Finally, we prove some theorems concerning the convergence in the q-th mean as well as the q-th and the exponential moments of series of independent F-cocp(no)(R-d)-valued random variables for q> 0.
Fuzzy linguistic summaries provide compact, human-readable descriptions of complex data. In explainable artificial intelligence (XAI), they have been used to transform numerical model explanations into natural-language statements. Traditional quality criteria, such as the degree of truth and the degree of support, are exclusively data-driven and do not necessarily capture domain knowledge. In clinical settings, experts often reason contrastively and relative to baseline expectations: a pattern is informative if it differentially characterizes one diagnostic state versus others and if it increases the concentration of that state under the observed pattern. In this paper, we introduce two new quality criteria, contrast and surprise, for assessing fuzzy linguistic summaries in the classification contexts. In the current formulation, the proposed criteria are defined for the qualifier-summarizer pattern and do not directly depend on the quantifier, so they are intended to complement, rather than replace, summary-level criteria such as the degree of truth and degree of support. We provide illustrative experiments on real-life data to show their complementarity with the truth value and support.
Linguistic summaries are an intuitive tool for obtaining analysis and data mining results that are easy to use, even for novice users. Until now, linguistic summarization has been used primarily to describe and facilitate the interpretation of large data sets. This work aims to develop methods enabling the construction of linguistically quantified sentences reflecting both the sequence of observations of a time series as well as the estimated parameters of hidden Markov models. The resulting fuzzy linguistic summaries with hidden Markov models (HMMs) may be exemplified as follows: "For most observations around 1.1, we have a high exact match rate". Preliminary results illustrate the effectiveness of the proposed approach using simulation methods.
Abstract The monitoring of inhomogeneous and non-stationary processes composed of segments and subsegments is considered. The structure of this segmentation is typical for medical data, describing voice characteristics of Bipolar Disorder (BD) psychiatric patients, calculated from their recorded smartphone calls. Data from subsegments are described by different probability distributions and are represented by histograms. Then, data from subsegments belonging to the same segment are aggregated using probability boxes (p-boxes) methodology and a simple probabilistic method. Finally, the mean value of each of the aggregated segments is described by a fuzzy triangular number. Therefore, the stream of consecutive segments is represented by the stream of fuzzy numbers. Several control charts for such fuzzy data are proposed. Their statistical properties are evaluated using simulated synthetic data. The simulation model is related to the real-life data obtained from the monitoring of BD patients. The results of simulations demonstrate the applicability of the proposed procedure for monitoring of BD patients.
Monitoring of inhomogeneousHryniewicz, O. Kaczmarek-Majer, K. non-stationary processes has been considered. The application of well-known statistical methods for the analysis of such processes may be questionable, and in the case of long streams of data even infeasible. In the paper, we consider processes consisting of segments and subsegments. The data from subsegments belonging to respective segments are represented by histograms. For consecutive segments, they are aggregated using probability boxes (p-boxes) and a simple probabilistic method. As a result of this aggregation, consecutive segments of the monitored process are represented by triangular fuzzy numbers. These fuzzy numbers may be used for process monitoring using statistical process control (SPC) methods, such as, e.g., control charts, for fuzzy data.
Distinguishing between web traffic generated by bots and humans is an important task in the evaluation of online marketing campaigns. One of the main challenges is related to only partial availability of the performance metrics: although some users can be unambiguously classified as bots, the correct label is uncertain in many cases. This calls for the use of classifiers capable of explaining their decisions. This paper demonstrates two such mechanisms based on features carefully engineered from web logs. The first is a man-made rule-based system. The second is a hierarchical model that first performs clustering and next classification using human-centred, interpretable methods. The stability of the proposed methods is analyzed and a minimal set of features that convey the class-discriminating information is selected. The proposed data processing and analysis methodology are successfully applied to real-world data sets from online publishers.
Controlling the impact of partial supervision on the outcomes of modeling is of uttermost importance in semisupervised fuzzy clustering. Semi-Supervised Fuzzy C-Means (SSFCMeans), a specific model we consider, uses a single hyperparameter called a scaling factor α to weigh the impact of partially labeled data. This concept became widespread and was reused directly in many works building on SSFCMeans, or even applied to other fuzzy clustering algorithms such as Possibilistic C-Means. However, none of the works challenged the original interpretation of α which suggests that the impact of partial supervision is directly proportional to the scaling factor. We fill the above research gap and thoroughly analyze this relationship. We provide novel explanations of the scaling factor α in terms of the key element of fuzzy clustering - the membership values. We prove that the impact of partial supervision is a non-linear function of α. Our approach is rooted in the explainability framework, which distinguishes interpretation from an explanation and treats the latter as superior. Explaining the scaling factor leads to an explainable impact of partial supervision and enables greater control of it. Finally, built on the novel explanations, we propose a unified, analytically justified framework for selecting the value of the hyperparameter α that is based on the crossvalidation approach. We illustrate that the proposed framework enables an extensive analysis of the impact of partial supervision in SSFCMeans with a simulation experiment.
INTRODUCTION:Voice features could be a sensitive marker of affective state in bipolar disorder (BD). Smartphone apps offer an excellent opportunity to collect voice data in the natural setting and become a useful tool in phase prediction in BD. AIMS OF THE STUDY:We investigate the relations between the symptoms of BD, evaluated by psychiatrists, and patients' voice characteristics. A smartphone app extracted acoustic parameters from the daily phone calls of n = 51 patients. We show how the prosodic, spectral, and voice quality features correlate with clinically assessed affective states and explore their usefulness in predicting the BD phase. METHODS:A smartphone app (BDmon) was developed to collect the voice signal and extract its physical features. BD patients used the application on average for 208 days. Psychiatrists assessed the severity of BD symptoms using the Hamilton depression rating scale -17 and the Young Mania rating scale. We analyze the relations between acoustic features of speech and patients' mental states using linear generalized mixed-effect models. RESULTS:The prosodic, spectral, and voice quality parameters, are valid markers in assessing the severity of manic and depressive symptoms. The accuracy of the predictive generalized mixed-effect model is 70.9%-71.4%. Significant differences in the effect sizes and directions are observed between female and male subgroups. The greater the severity of mania in males, the louder (β = 1.6) and higher the tone of voice (β = 0.71), more clearly (β = 1.35), and more sharply they speak (β = 0.95), and their conversations are longer (β = 1.64). For females, the observations are either exactly the opposite-the greater the severity of mania, the quieter (β = -0.27) and lower the tone of voice (β = -0.21) and less clearly (β = -0.25) they speak - or no correlations are found (length of speech). On the other hand, the greater the severity of bipolar depression in males, the quieter (β = -1.07) and less clearly they speak (β = -1.00). In females, no distinct correlations between the severity of depressive symptoms and the change in voice parameters are found. CONCLUSIONS:Speech analysis provides physiological markers of affective symptoms in BD and acoustic features extracted from speech are effective in predicting BD phases. This could personalize monitoring and care for BD patients, helping to decide whether a specialist should be consulted.
Semi-Supervised Fuzzy C-Means (SSFCMeans) model enables inclusion of additional knowledge about the true class of a part of the training data. With this partial supervision, there comes a new possibility to use this model as a classifier. The main goal should be thus to minimize the classification error, just as in the fully supervised setting. However, the typical problems with minimizing the training error, test error, and avoiding the phenomenon of overfitting must be carefully considered with respect to the characteristics of the SSFCMeans model. In this work, we fill the identified research gap and analyze the way of handling partial supervision in Semi-Supervised Fuzzy C-Means and its impact on the aforementioned issues. We investigate this relationship experimentally using artificially simulated data. We show that the training error for the training phase is directly related to the scaling factor α and is deterministically assured to be equal to 0 in some cases. We further illustrate our main findings for real-life partially labeled data collected from smartphones of patients with bipolar disorder in a problem of predicting the phase of the disease.
Abstract Life expectancy is an essential indicator of economic development and health status. However, the related databases describing the overall life expectancy are relatively large and created under conditions of uncertainty, particularly regarding adjustments related to redistributing deaths of unknown age or splitting data into finer age categories. In parallel, comprehension of general trends related to the life expectancy indicator is a crucial topic from the perspective of both the private and public sectors. It supports long-term decision-making about social policies. The key question addressed in this study is the relation between life expectancy and income inequality. Fuzzy linguistic summarization is applied to explore the inequity measured with the Gini coefficient and life expectancy. We show that the outcomes of this intelligent linguistic analysis reveal new information that complements the traditional correlation analysis. The fuzzy summarization approach enables capturing and explaining, in a human-consistent way, the relations between the considered indicators. The experimental results are presented for yearly data from 24 European countries observed from 1995 to 2021. The results are promising and show the usefulness of the linguistic summarization approach for explaining the relation between life expectancy and income inequality. In particular, although the length of life varies according to gender, the relationship between life expectancy and inequality follows a similar pattern for females and males. The relationship may seem intuitive, but previous research does not confirm it unequivocally. Furthermore, our study shows a level of inequality for which changes in income distribution do not significantly impact life expectancy. JEL Classification: C0 , J1
This book presents ample, richly illustrated account on results and experience from the analysis of data concerning behavior patterns on the Web.
We start with the description of general context and with formulation of the problem that we address. In this manner we set a framework for both the particular issues that we deal with on a technical level, described in the consecutive parts of the book, and for the potential implications thereof, some of them forwarded as more general conclusions or hypotheses.
The case study presented in this book highlights the properties and challenges of distinguishing bot and human traffic using weblogs and compares several solutions to this task. We present here, first, the observations related to the methodological side of the potential problem solution.
This chapter is devoted to presentation of results, which were obtained for the data considered with the use of clustering algorithms. Clustering consists in grouping of similar observations, while separating the dissimilar ones. (The groups obtained therefrom are referred to, exactly, as clusters.) Hence, we might suspect that the observations we analyse should fall into different clusters, and that first of all depending upon whether they represent “bot” or “human” behavior, based on the apparently obvious assumption that bots are more similar to (at least some) bots than to humans and that humans are more similar to (at least some) humans than to bots. Therefrom the attempts we report on in the present chapter.
Classification is a decision-making problem, in which we aim at the assignment of a correct class label to an observation. In the typical scenario, the set of available class labels is fixed beforehand and remains unchanged. Advanced data processing streams allow for a more flexible definition of this task. Typically, they admit the existence of a “novelty” class, to which some observations may be assigned.
Having characterised the problem that we address and the way of acquiring data that we use, we now turn to the features, which can be extracted from the raw data, and the choice of the possibly good selection of a subset of these variables from the point of view of the problem at hand.
When dealing with challenging, real-world datasets, one may turn to hybrid data processing approaches that join several kinds of data analysis algorithms. The hybrid data analysis techniques are popular in the literature, where we may find fusions of optimization and classification algorithms, clustering and classification algorithms.
Analyzing data from the web is now one of the primary tasks, understood in a variety of manners and solved for a very wide variety of purposes. The talk describes the experience from a project, devoted to analyzing such data while drawing some more general conclusions. The project was aimed at distinguishing artificial ad-related traffic from the genuine one. The rationale is simple: The flow of money depends upon the number of clicks on/views of an ad. If so, fake clicking changes the market to the benefit of some, and to the loss of the other ones. The talk describes the problem and its conceptual framing, as well as a number of technical details, involving the issues and techniques of (1) variable analysis and choice; (2) clustering; (3) classification/classifiers; (4) potential hybrid techniques, along with citations of the most interesting results. These often imply definite general conclusions, some of them quite surprising.
In this chapter, we will look in a more detail at the ways in which data were acquired and processed in the framework of the project in question. We will do so against the background of observations and examples already provided in the preceding section, starting with the first section of the present chapter.