
Partially observable environments present increased decisionmaking complexity for Reinforcement Learning agents. This paper explores methods for improving automated decision-making in partially observable environments by utilizing Variational Auto-Encoders (VAEs) trained on expert demonstrations to generate intrinsic rewards for reinforcement learning (RL) agents. Specifically, a VAE is pre-trained on expert demonstrations to construct a latent representation of successful decision-making. This latent representation is utilized to generate intrinsic rewards via KL-divergence, augmenting the extrinsic reward signal during Proximal Policy Optimization (PPO) training. Experiments are conducted using a symbolic-matching navigation task within a Unity MLAgents environment, which requires memory formation and hierarchical reasoning. Results indicate that incorporating demonstration-based intrinsic rewards improves the learning efficiency and convergence of PPO agents compared to baseline models without intrinsic rewards. The findings suggest that latent-space representations from demonstrations can effectively guide exploration in challenging RL scenarios, and that certain behaviors are retained from demonstrations.
Similarity functions between sets and in particular, between intervals, are traditionally assumed to verify a monotonicity condition. It requires that if three sets (intervals) are ordered by inclusion, the similarity between the greatest and the smallest must be smaller than the similarity between the middle one and any of the other two. We find this property coherent but too weak since it does not require anything from triplets of sets that are not perfectly ordered. A condition introduced in the last years that involves the concept of infimum and supremum applies to every triplet of intervals and generalizes the classical axiom of monotonicity for similarities. Despite it is a logical and intuitive condition, it is easy to check that it is not satisfied for every similarity. In this contribution we focus on similarities defined from embeddings and more particularly, from two families of embeddings and we study if the two families of similarity measures obtained comply with the general condition of monotonicity introduced two years ago. A positive answer is obtained in all the cases studied.
ChessFormer introduces a novel searchless chess engine leveraging transformer architecture to approximate human decision-making in chess. Trained on a vast dataset of 3 billion chess positions, our model learns its entire decision-making process directly from training data. Evaluations show an improvement in human move-matching accuracy over prior models in high-Elo ranges and the model's ability to distinguish between human and algorithmic decision-making, offering potential applications in chess analysis or cheat detection.
Choquet integrals are a class of non-additive aggregation operators that generalize the weighted arithmetic means and order statistics (for example, the sample median). It is a prevalent example of fuzzy integrals that has been extended in various ways. In particular, replacing the naive difference between assessments in the Choquet integral by restricted dissimilarity functions produces d-Choquet integrals. We define and study differentially private d-Choquet integrals. They are designed to satisfy differential privacy with the introduction of Lagrangian noise calibrated by their sensitivity. In this way, a new additive noise differential privacy mechanism is proposed that generalizes differentially private Choquet integrals. The later case already gave valuable information about the sensitivity of linear combination of order statistics and weighted means, for which the literature on privacy preservation is abundant. To meet the needs of practical implementation, we compute the sensitivity of d-Choquet integrals associated with several types of fuzzy measures and restricted dissimilarity functions. Other theoretical results link these achievements with knowledge about the particular case of differentially private Choquet integrals.
Decision-making in complex systems often involves dealing with imprecise or uncertain information, frequently represented using fuzzy sets, particularly Triangular Fuzzy Numbers (TFNs). A crucial aspect of many fuzzy methods is the quantification of distance between TFNs. Many distance measures assume that all values are in the same scale, requiring a preliminary normalization stage when applied to heterogeneous attributes with different scales or units. This paper proposes the Triangular Fuzzy Rescaling Distance ( d_TR ), a metric designed to address this challenge. The d_TR uniquely integrates Linear Rescaling (LRE) directly into the distance calculation, ensuring normalization during the comparison of fuzzy numbers. We formally prove that d_TR satisfies the properties of a metric, including non-negativity, identity, symmetry, and the triangle inequality. Furthermore, we demonstrate that d_TR is bounded, scale-invariant, and origin-invariant. These properties, combined with a weighting vector for prioritizing dimensions, make d_TR suitable for applications involving heterogeneous fuzzy data, such as the construction of synthetic indicators, distance-based machine learning algorithms or multicriteria-decision aiding.
This paper aims to evaluate the impact of the various optimization algorithms used in methods based on logistic regression fitting in positive and unlabeled classification under the SCAR assumption. In our work, we consider the joint algorithm based on logistic regression. The impact of non-linear optimization algorithms and simulated annealing technique is examined on eleven real data sets from machine learning repositories and one synthetic set. To assess this impact, we evaluate the AUC, accuracy, recall, precision and F1 score of the classification for the naive and the joint method.
This article considers effective procedures for clustering customers in E-commerce, according to the importance ascribed by a firm to various aspects of the clustering. These procedures are based on data on the behavior of customers online and enable websites to be adapted to differing groups of customers. The effectiveness of a clustering procedure can be assessed according to various aspects, e.g. the clarity of the clusters, their relative sizes and the predictive value of the clustering. There is no single variable that definitively describes the effectiveness of an algorithm according to each aspect. Hence, using principal component analysis, we define standardized synthetic variables that describe the attractiveness of a clustering algorithm according to each aspect. This reduces the dimension of the resulting multi-criteria assessment. Weights can then be ascribed by an expert to a small number of synthetic variables that have a clear interpretation. We use such a procedure to compare the effectiveness of forty clustering procedures. These are based on a combination of five algorithms and eight different numbers of clusters. Although the data used are problem-specific, the results illustrate some of the strengths and weaknesses of the procedures used.
Payment channel networks like the Lightning Network (LN) boost cryptocurrency scalability but might have some privacy risks, as establishing a direct channel reveals user or entity relationships. This paper introduces a novel method using edge local differential privacy (LDP) applied during public channel creation to protect channel existence privacy. Our approach adds statistical uncertainty to the network graph, providing plausible deniability about direct links between specific users while ensuring channels remain usable for routing payments. We formally introduce edge LDP for this context, define relevant utility metrics focused on multi-hop payment feasibility, cost, and node centrality, and evaluate the privacy versus utility trade-offs using a snapshot of the Bitcoin LN.
The study proposes a fuzzy clustering algorithm referred to as the Sharma–Mittal divergence-regularized fuzzy c-means, which adopts the Sharma–Mittal divergence (SM-divergence) as the regularizer, where the SM-divergence is a generalization of the Tsallis and Rényi divergences. The algorithm is shown to reduce to two conventional algorithms, the Rényi and Tsallis divergence-based algorithms, by adjusting the parameter values at both the objective function and algorithmic levels. This study also examines the differences between the use of divergence in the Rényi and SM-divergence-based algorithms and object-wise divergences in the Kullback–Leibler and Tsallis divergence-based algorithms. Subsequently, this study proposed two fuzzy clustering algorithms by applying object-wise Rényi and SM-divergences. Numerical experiments on an artificial dataset verified the theoretical findings and highlighted the impact of the fuzzification parameter on clustering results.
Polypharmacy, patients suffering with extremely many kinds of drugs for particular diseases, has been critical issues in aging society. It is not trivial to explore primary factors associated with polypharmacy because it requires huge amount of confidential personal healthcare data. In this study, we propose a new methodology applying local perturbation ensuring differential privacy to the healthcare data to identify risk factors in polypharmacy. With a large-scale, anonymized 390,000 health insurance claims for over-60-years-old individuals, we apply two-dimensional randomized responses to reveal relative risks of suffering polypharmacy for each of potential factors.
For quantitative portfolio optimization problems, predicting the various asset returns accurately is essential. Transformer architectures, which have achieved state-of-the-art performance on a broad variety of sequence modeling tasks, offer significant potential for capturing the complex temporal dynamics inherent in financial markets, going beyond traditional limitations. However, financial markets present unique challenges due to non-stationarity and complex relationships between stocks. Also, while numerous variants of the Transformer have been proposed for general time series forecasting (TSF) with a variety of mechanisms, their relative performance and best suitability for ranking-based portfolio selection are barely examined. We close this gap by adapting and evaluating several Transformer architectures (Vanilla Transformer, CrossFormer, adapted MASTER and iTransformer) for daily return forecasting to enable ranking-based portfolio selection on S&P500 using stock market data. Focusing on their ability to model both temporal patterns in individual stocks in the market and valuable relationships between them. Our contribution is a benchmark comparing these adapted architectures in a realistic ranking task. The analysis shows how the different mechanisms capture temporal and cross-sectional information, identifying the strengths and weaknesses of each model in generating profitable rankings and offering practical insight into the choice of suitable transformers for quantitative trading.
In this paper, we introduce a dual-focus training approach that explicitly teaches a model what it should predict and then extends this knowledge by teaching it what not to predict. Our core contribution is a custom loss function that not only penalizes deviations from the ground truth label but also imposes additional costs when the prediction drifts toward a false positive region. The proposed solution is effective for regressions and classifications of one or two types of objects and is based on adding additional surrogate classes of negative data that teaches the models to discriminate features specific to false classes from the positive ones. We evaluated the method on multiple types of problems: regression-based tasks (glint detection dataset) and detection tasks (brain tumors, red blood cell detection, and parking slot occupancy). Empirical results demonstrate substantial reductions in false negatives and false positives, highlighting the practical value of explicitly encoding negative-avoidance signals in the training objective. Results indicate significant improvements in both recall and precision, with notable IoU gains (up to 14.79
Introductory statistics courses require instructors to convey mathematical and analytical concepts clearly, engagingly, and incrementally. Traditionally, -generated slides are widely used due to their precision and readability, yet these presentations often struggle to maintain student attention and replicate the natural cognitive pacing provided by handwritten blackboard explanations. To bridge this pedagogical gap, we propose a rule based method to animate statistical lecture slides by closely emulating the incremental writing sequence found in traditional handwriting.
We focus on the inference process in a probabilistic-fuzzy IFTHEN rule system, where antecedents are modeled as fuzzy sets and consequents as weighted quantile functions, offering insights into the probability distribution of response data. While the weighted arithmetic mean has been the conventional inference method, other suitable approaches exist. This work compares it with three alternatives: the weighted geometric mean, weighted L1 minimization, and a mixture distributionbased method. Through experiments on both simulated and real-world data, we evaluate their performance using standard statistical measures.
Sentiment analysis is needed for businesses to understand and analyze customer feedback. Transformer models like RoBERTa and ALBERT represent the state-of-the-art in sentiment analysis. Their key limitation is that their standard usage typically relies only on the maximum predicted probability, discarding key information embedded within the full output probability distribution. This paper investigates enhancing transformer-based sentiment analysis by integrating Adaptive NeuroFuzzy Inference Systems (ANFIS) to model predictions based on six scores derived from probability distributions. We fine-tune RoBERTa and ALBERT on the yelp review dataset. Then, we extract six features characterizing prediction confidence, entropy, and potential error indicators from their probability outputs. These features serve as inputs to ANFIS models, which are trained to predict the likelihood of the baseline prediction being correct. We evaluate two strategies for combining the baseline predictions using the ANFIS trust scores: a threshold-based selection method and a trust-weighted averaging method. Evaluations compare the hybrid system against the individual baseline models, showing that ANFIS improves performance. This work contributes a novel methodology for integrating transformer probability outputs with ANFIS for trust assessment, a hybrid architecture enhancing sentiment analysis robustness, and an empirical validation with performance improvements.
Rule-based systems, particularly those using "IF antecedents THEN consequents" rules, play a crucial role in mathematical modeling and decision-making. This work focuses on a specific type of rule-based system: the Probabilistic-Fuzzy Inference System. Here, antecedents are represented as fuzzy sets, while consequents are modeled as probability distributions using quantile functions. The inference process relies on the L1-Fuzzy transform, also known as the Quantile Fuzzy transform. For each fuzzy antecedent, a weighted quantile of order p is computed, and the inverse quantile transform derives the empirical quantile function for any input value. Initially, weighted quantiles are modeled as scalar values. To better capture dependencies between input and output data, we extend this model to a piecewise linear function, developed in a structured manner with a comprehensive computational algorithm. Its effectiveness is demonstrated through an illustrative example.
Distributions are ubiquitous, with applications extending far beyond probability and statistics into nearly all scientific disciplines. In the field of machine learning, finite probability distributions form the softmax output layers in Convolutional Neural Networks (CNNs) and Large Language Models (LLMs), where they represent class probabilities in image classification and token probabilities in language generation. However, when dealing with very large distributions, this incurs significant storage and computational costs. In our previous work, we introduced generalizations of Shannon entropy derived from f-divergences, using majorization as a reference framework for comparing homogeneity. Here, we extend these entropies as tools for dimensionality reduction while integrating them into the concept of Shannon's information channel.
When we humans observe objects, we take views and perspectives. These views can present different features depending on the location observed, which are oriented or symmetric. This paper presents a Qualitative Object Descriptor (QOD) to model objects as embedded in a 3D cube. The QOD extracts a spatial descriptor of an object which allows human understanding. This paper also presents a visual similarity measure for QOD (SimVis) which has been used to analyze the similarity between scenes that contain pairs of objects as those in the Cube Comparison Test (CCT) which has been extensively used in the literature to measure spatial skills in people.
Ecological inference (EI) encompasses diverse methods to address the frequent lack of individual-level data in various areas, particularly in decision-making processes. Considering only aggregated information, these methods allow inferences to be drawn about individual behaviors and preferences. The difficulty in EI stems from the potential for the ecological fallacy, where inferences drawn from aggregated data do not accurately reflect relationships at the individual level. Numerous statistical techniques and computational tools have been developed to address this challenge. This study focuses on comparing and contrasting several prominent R packages specifically designed for ecological inference: lphom, eiPack, eiCircle, ecolRxC and eiopt2. The comparison delves into the distinct computational approaches employed by each package, examines the underlying statistical models that form their foundation, and critically evaluates how each handles the inherent uncertainty associated with ecological inference. This includes considering how each package addresses potential biases and limitations inherent in aggregated data. To provide a robust empirical assessment, this study leveraged paired survey data from the Spanish Center for Sociological Research (CIS) focusing on pre- and post-election studies related to the 2015 Spanish General Election.
Within the Ontology Based Data Access (OBDA) framework, users can query relational data sources using an ontology to which the source is linked via declarative mappings. In a world where data sharing is widespread, ensuring privacy while managing data poses a significant challenge. Controlled Query Evaluation (CQE) is a privacy preserving query answering framework in the presence of ontologies, where policies representing confidential information are used to devise suitable censors that enforce data protection. The integration of CQE within OBDA was recently proposed through the PolicyProtected OBDA (PPOBDA) framework, which is based on embedding policies into mappings. Such framework is essentially theoretical, and the effectiveness with which PPOBDA policies are able to capture real-world privacy requirements has not been assessed so far. In this work, we carry out such an evaluation, utilizing the well-known MIMIC-III hospital dataset, which recently has been mapped, by adopting the OBDA framework, to the Fast Healthcare Interoperability Resources (FHIR) ontology. We identify relevant privacy requirements by analyzing the legal regulations on data sharing expressed in HIPAA of US Federal Law and GDPR of the EU, show how they can be expressed via PPOBA policies, and analyze the impact of these policies on the answers to a set of representative queries. Our analysis exposes both strengths and weaknesses of the PPOBA framework in relation to these practically relevant privacy regulations. Furthermore, we perform a performance evaluation of the OBDA framework implemented over the MIMIC-III dataset via the FHIR ontology, assessing the overhead introduced by the PPOBDA policies and its implications on such realworld use case.