
This paper introduces WECM, a novel evidential and subspace clustering algorithm. It is based on the Evidential c-means, a variant of the k-means designed to produce a credal partition, allowing a better representation of the partial knowledge regarding the class membership of objects. The WECM algorithm integrates weights on features and clusters to enhance the clustering separability and interpretability. Experiments conducted on synthetic and real data show the positive effects of the weights on the clustering performances.
We propose a novel feature selection (FS) method based on peculiar fuzzy set generation, aggregation, and ordering. In particular, here we propose to elicit the fuzzy membership by a proper probability-possibility transformation of frequencies stemming from the bootstrap application of different filter FS methods. At the same time, we aggregate such vague scores of each feature via the recently introduced SMART-or fuzzy aggregation operator. Finally, to rank the features for the selection proposal we adopt Yager's ordering. Empirical results on benchmark databases show an overperformance of our approach with respect to different generation techniques, or aggregation functions and orderings.
The ongoing digital era is significantly impacted by the dissemination of fake news and disinformation. Addressing this issue involves delving into the realm of Artificial Intelligence, where Natural Language Processing (NLP) stands out as one of the most active areas capable of contributing to solutions. In this paper, we present an analysis grounded in argument mining, with a specific focus on understanding the creation of fake pieces of content in social media. Our research reveals that deceptive narratives in social media often incorporate a substantial component of arguments. This finding not only sheds light on the intricacies of misinformation but also provides valuable insights for future research in combating this pervasive issue.
One of the forms to fight the spread of hate speech is to reply to such utterances with counter speech. In this paper, we present models to identify counter speech, based on an annotated set of comments to Youtube videos spoken in Portuguese. We leverage the sequence of replies to comments, to form a corpus of pairs of comments, where a target is labelled as neutral, hate speech or counter speech, relative to a context, which corresponds to a preceding comment. To the best of our knowledge, this is the first corpus with counter speech examples in Portuguese. Using such corpus, we compute models by fine-tuning pre-trained models based on Transformers, and experiment with both multilingual and Portuguese pre-trained models. Our approach follows a recent work for English, both in corpus design and experimental setup, and we obtain similar performance results in Portuguese. Warning: This work contains offensive and hateful text that some might find upsetting. It does not represent the views of the authors.
In the paper a convex game with discontinuous payoff functions is considered. This paper introduces the concept of quasi-Nash equilibrium, which allows us to analyze games with discontinuous payoff functions by approximating them with continuous and concave functions. We show that if the approximation is close enough, then the quasi-Nash equilibrium is close to the true Nash equilibrium (if it exists).
We consider the marginal problem in Dempster-Shafer theory, investigating the structure of a suitable set of bivariate joint belief functions having fixed marginals, by relying on copula theory. Next, we formulate two Kantorovich-like optimal transport problems, either seeking to minimize the Choquet integral of a given cost function with respect to the reference set of joint belief functions or its dual functional. We finally give a noticeable application by choosing a metric as cost function: this permits to define pessimistic and optimistic Choquet-Wasserstein pseudo-distances, that can be used to compare belief functions on the same space.
The classical Efron's bootstrap is a widely used tool in statistical inference. However, because of its disadvantages, many other resampling algorithms were proposed in the literature, especially for the real-valued data. In this paper, we consider three resampling methods for the special case of the interval real-valued data. They were inspired by the smoothed bootstrap and special algorithms known for fuzzy numbers. Using numerical simulations and statistical tools, the introduced methods are compared with the Efron's bootstrap. It seems that these new algorithms produce samples that can be considered as "similar, but not exactly the same" as the initial data, which is an important aim in the case of resampling methods.
The information disorder phenomenon represents one of the main challenges for the current society, that researchers of a huge variety of scientific areas are trying to solve. To date, the majority of the studies carried out in this context are focused on Machine Learning-based Fake News detectors trained on textual data. Despite the numerous attempts available in the literature and the promising results of such models, unfortunately, the expectations are not truly met considering their use in real-world scenarios. The main limitations are directly related to the nature of the phenomenon since the model trained on past events can't understand and classify novel contents and breaking news. In the current work, we study if sentiment discrepancy between news and its associated evidence can determine the possible presence of fake content. In fact checking activities, the evidence are inferred claims or information used to accept or reject shared news [2]. To do that we have conducted a preliminary study in which the outcome is a metric called Negativity Score. This metric can be used as a feature in Fake News detection models and automated fact-checking activities. For completeness, we have also provided a framework that uses the proposed metric. The results of preliminary experiments highlight the possibility of exploiting sentiment aspects in addition to more common and well-known approaches.
This paper focuses on the development of two prediction models for a solar photovoltaic system that is part of a multimachine industrial manufacturing plant. These models are part of the set of models that form the digital twin of their physical counterparts, which will be used to perform control and optimization strategies to maximize the use of renewable energy sources within a Digital Twin (DT) architecture. The first model is based on a fuzzy neural network and the second one is a Gaussian regression model. The obtained models present a good performance in the prediction of the nonlinear dynamic over the entire operating range in the system.
Detecting bots on social media platforms is a major challenge, as these automated entities are constantly evolving to evade detection. In this study, we investigate the main features that contribute to the difficulty of bot detection. Leveraging the TwiBot-20 dataset, we analyze the characteristics of misclassified accounts and explore the reasons behind their erroneous classification. Our approach combines feature engineering, Machine Learning with Random Forest, and the interpretation of model predictions using SHAP (SHapley Additive exPlanations) values. We employ clustering techniques to identify patterns in feature contributions and provide insights into the complexities of distinguishing between human and automated accounts. Our findings highlight the nature of bot detection and the need for advanced methods to address the problem of social media manipulation.
Rule-based approximate reasoning systems are an important decision-making tool in many application problems. The use of expert knowledge or machine learning techniques to create rules does not exhaust the problems of representing data and decision dependencies, therefore we propose a hybrid/mixed technique for creating a set of rules while effectively modeling uncertainty through interval-valued fuzzy representation in the problem of detecting falls of elderly people. The obtained prediction confirms the correctness of the choice of diagnostic methodology.
We consider a dynamic portfolio selection problem in a finite horizon binomial market model, composed of a non-dividend-paying risky stock and a risk-free bond. We assume that the investor's behavior distinguishes between gains and losses, as in the classical cumulative prospect theory (CPT). This is achieved by considering preferences that are represented by a CPT-like functional, depending on an S-shaped utility function. At the same time, we model investor's beliefs on gains and losses through two different epsilon-contaminations of the "real-world" probability measure. We formulate the portfolio selection problem in terms of the final wealth and reduce it to an iterative search problem over the set of optimal solutions of a family of non-linear optimization problems.
The DE-MCZ algorithm is an improvement of the DE-MC method, which joins the differential evolution with the theory of the Markov chains. It aims to ensure the numerical effectiveness and the convergence speed of the special variant of the Metropolis-Hastings algorithm with the help of an additional, self-adapting initial matrix. In this paper, we add the modes detection procedures to the DE-MCZ algorithm to increase its abilities in sampling from multimodal target densities. As our numerical experiments suggest, the obtained DE-MCmodes algorithm provides results that give a better fit to the desired target density than the classical approaches.
In this paper, we propose a new approach for detecting and removing impulse noise. The method uses the new framework of cloud filtering for detecting noise locations. This new framework use efficient mathematical tools to filter with sets of filters rather than with a single one. Once a (very) noisy pixel is detected, its illumination is estimated by an extension of the median filtering applied on the neighborhood defined by the cloud. Experiments on various images demonstrate the capacity of the algorithm to identify noisy pixels (especially at low noise rate) while well preserving image edges.
This paper presents the identification of a desalination plant model using fuzzy inference techniques, as well as their comparison with the Linear Parameter Variation (LPV) experimental identification. Identification of the plant model has been carried out using the fuzzy C-means clustering (FCM) technique. The identified model was then validated, and the estimated output was compared with the measured output. Both models were obtained with experimental data by running the plant in three different scenarios, with the only variation in the operating point of the waste reuse valve, although the differences are minimal. The results obtained show that the FCM presents the lowest variability in the estimates, the lowest discrepancy between the predicted and observed values.
Linguistic summaries are an intuitive tool for obtaining analysis and data mining results that are easy to use, even for novice users. Until now, linguistic summarization has been used primarily to describe and facilitate the interpretation of large data sets. This work aims to develop methods enabling the construction of linguistically quantified sentences reflecting both the sequence of observations of a time series as well as the estimated parameters of hidden Markov models. The resulting fuzzy linguistic summaries with hidden Markov models (HMMs) may be exemplified as follows: "For most observations around 1.1, we have a high exact match rate". Preliminary results illustrate the effectiveness of the proposed approach using simulation methods.
Three-valued logics became a classical topic for logicians and not surprisingly, they were extended to partial fuzzy logics that allow modeling distinct types of undefined truth-values, i.e., values, that are neither true nor false, even in the graded sense. Such logics and related algebras may model reasoning with non-denoting terms, missing or unknown values, and other interesting cases. However, in order to be able to model real cases, the algebraic models need to be mirrored in applied tools such as inference systems. Therefore, the investigation of partial fuzzy relational equations that question the most natural property of such systems is a straightforward step. This step has been already made, however, satisfactory results were obtained only for the direct product inference (compositional rule of inference). Intuitively, this was due to the application of partial algebras that employ the so-called lower boundary strategy. This article introduces upper boundary algebraic strategy and shows, that equally satisfactory results may be obtained also for the Bandler-Kohout subproduct.
The Human Development Index (HDI) is a widely recognized measure, designed to assess the overall well-being and development of nations. This article explores, in the context of Multiple Criteria Decision Analysis, the integration of the discrete 2-additive Choquet Integral into the computation of the HDI, by analyzing the interactions among its dimensions. We show that the HDI formula is equivalent to an additive model and then these interactions can only be interpreted as the possible interactions.
While food significantly impacts our daily lives, the health implications of dietary choices are equally crucial. Recipe adaptation systems emerge here as a helpful practice for automatically modifying food to recipes. They offer users diverse options to substitute ingredients while fulfilling specific needs. However, users may lack expertise in nutrition, making health-conscious decisions more challenging. In this study, we propose to use fuzzy linguistic variables to offer easily understandable nutritional information for final users. We show the process of incorporating fuzzy nutritional labels into the final interface of a food system, demonstrating the practical implementation of our approach. This user-friendly interface empowers individuals with understandable dietary information to make informed choices, thus improving the nutritional quality of their food selections.
This paper introduces a novel fuzzy association rule mining algorithm explicitly developed for federated environments. With exponential growth in datasets and increasing data privacy concerns, solutions such as federated learning have become at the forefront of secure and efficient data analysis. However, efficiently finding meaningful and relevant patterns in data across decentralized databases remains challenging. To address this, we propose integrating fuzzy logic with association rule mining in a federated setting. The ability of fuzzy logic to handle uncertainty and nuance in data combined with the distributed data mining process of federated systems, creates an efficient, secure, and powerful tool for pattern discovery. Our proposed algorithm respects data privacy and effectively manages communication overhead, an innate challenge in federated systems. Experimental results demonstrate the efficacy of the proposed algorithm. The system has significant implications for the healthcare sector, where data volume and privacy concerns are paramount.