Fraud detection systems (FDS) must contend with several intertwined challenges: class imbalance, robustness against adaptive adversaries, black-box constraints, and need for explainability. While each of these challenges has been addressed individually in the literature, few works offer an integrated approach capable of operating under realistic adversarial settings. These solutions often rely on Logistic Regression to predict fraudulent behavior, and on data augmentation techniques in an indirect or ad hoc manner, aiming to approximate real-world conditions without fully capturing their complexity. In this study, we introduce JDA-NC-FDS, a novel fraud detection model that tackles these 4 issues. First, we assess the comparative performance of Logistic Regression and XGBoost within an iterative adversarial learning setup. Second, we propose to rely on the Jacobian Data Augmentation instead of the traditional SMOTE. In addition, we enhance JDA to incorporate categorical features, a critical yet underexplored type of dimension in fraud detection. Experimental results demonstrate that JDA-NC-FDS improves model performance across multiple adversarial training rounds, especially in terms of PR-AUC, while preserving explainability and adaptability.
Awarding gaps have been commonly observed between different socio-demographic categories of students, especially in the domains of sociology and learning science. Recent research has shown that using Learning Analytics models could be exploited to reduce these gaps, and therefore contribute to making the learning process more inclusive and equitable. This demonstration paper presents CERSEI, a new web-based learning prototype that aims to enhance inclusiveness by exploiting two Learning Analytics models: a cognitive effort model and an activity recommender built upon the cognitive effort model. Previous research has indeed shown a strong interplay between socio-economic status, effort and motivation, e.g., families from higher socio-economic status tend to mobilize more resources to prevent their children from falling down the social ladder. Some categories of students might therefore have fewer sources of motivation and exert less effort, or a higher tendency to exert effort on specific activities that are not the most relevant for succeeding. CERSEI allows students to track their effort by assigning ratings on their activities using the RSME scale and to receive engaging recommendations of learning activities. This will allow us to collect the relevant data to better understand how effort is exerted by different categories of students and how recommendations can impact them. Based on the outcomes of the related analysis, we will then aim at creating better Learning Analytics models. We expect that these models will help to provide more inclusive and equitable learning.
BackgroundEarly literacy and numeracy skills are developed during early childhood. Among the many factors that influence the development of such skills, the literature shows that the executive functions, especially the response inhibition (RI)-that is the capability to block out or to tune out what can be considered irrelevant information or action to the learning task-is one of the most essential functions. There are specific tests used to appraise these children's inhibition skills, but these tests are generally time-consuming, and demand specialized human resources. ObjectivesWe present a computational approach to model children's RI behaviour through the analysis of educational traces left in an educational app. This modelling allows the automatic and instant identification of the RI level of children without the need of a human-conducted test. MethodsOur modelling is based on two definitions of RI found in the literature, from which we derived a mathematical formalism of three variables we used to query the traces dataset and isolate the RI behaviour of each student from the learning traces generated in the app. The sample population is composed of children from diverse socioeconomic backgrounds. The model is then assessed by comparing it to a traditional human-conducted RI test suitable for kindergarten children, the Head-Toes-Knees-Shoulder (HTKS) task. Results and conclusionsThe results show that our RI model can explain an important part of the HTKS variance (up to 0.45 according to the adjusted R-2) when taking the HTKS results as a dependent variable for a multiple regression model. In practice, our model can be integrated in a learning app and become a powerful tool for instant preliminary identification of dysfunctional RI behaviour, especially in the early stages of children's education. Once students are identified by our model as having a dysfunctional RI behaviour, teachers can rapidly act to help them. Besides, the proposed model requires only very simple data to work, which means it can be easily integrated into different learning apps.
The ability to initiate gait involves a complex coordination between posture and movement, known as anticipatory postural adjustments (APAs). The emotional context in which gait initiation occurs can impact several spatio-temporal parameters, particularly the duration of APAs. While previous studies have used biologically relevant stimuli to induce emotions, such as images of pleasant or unpleasant scenes, to the best of our knowledge, the impact of the emotional context induced by music on gait initiation has not been explored yet. This paper presents a new dataset collected to study this impact. Objective biomechanical and physiological data were collected from participants during and after music listening, and subjective emotional responses were assessed using questionnaires. We also focused on two factors, liking judgment and familiarity, known to modulate emotions. Our preliminary analyses shows the impact of the emotional context induced by music on gait initiation, and confirms the strong importance of liking judgment and familiarity on the emotional context.
Students’ effort is often considered to be a key element in the learning process. As such, it can be a relevant element to integrate in learning analytics tools, such as dashboards, intelligent tutoring systems, adaptive hypermedia systems, and recommendation systems. A prerequisite to do so is to measure and predict it from learning data, which poses some challenges. We propose to rely on the cognitive load theory to infer the students’ perceived effort using subjective, performance, behavioral and physiological data collected from 120 seventh grade students. We also estimate students’ effort in future tasks using the data from previous tasks. Our results show a high relevance of interaction data to measure students’ effort, especially when compared to physiological data. Moreover, we also found that using the data collected on previous tasks allows us to achieve slightly higher accuracy values than the data collected during the task execution. Finally, this approach also allowed us to predict students’ perceived effort in future tasks, which, to the best of our knowledge, is one of the first attempts towards this goal.
In this paper, we rely on the Cognitive Load Theory and explore how multimodal data can be used to measure students' effort at the task level. Different from what we expected, the subjective effort ratings have a higher correlation with the students' scores, while the behavioral and physiological data have higher correlations with the scores than with the effort ratings. Moreover, we found that, in the context of our study, ability had a stronger influence on students' success than the prior knowledge, while none of these variables had an influence on the effort ratings. Finally, we propose a new effort model based on students' activity.
Early literacy and numeracy skills are developed during childhood at kindergarten level. Among the many factors that influence the development of such skills, the literature shows that the executive function of inhibition – i.e. the blocking out or tuning out of information or action that is irrelevant to the learning task – is one of the most important. There are many tests to assess children’s inhibition skills; however, such tests are generally time-consuming and have a short lifespan. In this context, we propose a computational approach to model children’s inhibition skills by using only student traces from a learning app as input. We propose a mathematical formalization of three related inhibition features, which could be used as input to classification algorithms.
The digitization of music, the emergence of online streaming platforms and mobile apps have dramatically changed the ways we consume music. Today, much of the music that we listen to is organized in some form of a playlist, and many users of modern music platforms create playlists for themselves or to share them with others. The manual creation of such playlists can however be demanding, in particular due to the huge amount of possible tracks that are available online. To help users in this task, music platforms like Spotify provide users with interactive tools for playlist creation. These tools usually recommend additional songs to include given a playlist title or some initial tracks. Interestingly, little is known so far about the effects of providing such a recommendation functionality. We therefore conducted a user study involving 270 subjects, where one half of the participants—the treatment group—were provided with automated recommendations when performing a playlist construction task. We then analyzed to what extent such recommendations are adopted by users and how they influence their choices. Our results, among other aspects, show that about two thirds of the treatment group made active use of the recommendations. Further analyses provide additional insights about the underlying reasons why users selected certain recommendations. Finally, our study also reveals that the mere presence of the recommendations impacts the choices of the participants, even in cases when none of the recommendations was actually chosen.
Decades of studies have shown that studentu0027s success is strongly dependent on their effort. Recently, this concept made its way into the domain of Learning Analytics. One of the major difficulties of these works is to correctly define the effort and to find relevant means of measuring it. Our approach is based on the Cognitive Load Theory, which provides a theoretical background issued from Learning Sciences, desired by the Learning Analytics domain. The cognitive load is a multidimensional construct that represents the load that performing a given task imposes on the cognitive system, and is often considered by researchers as being equivalent to mental effort. The cognitive load has long been studied in educational sciences, and several types of measures have been proposed that can be classified into four categories: (1) subjective measures, i.e., studentsu0027 perceived effort, (2) performance measures, e.g., the outcome of student work assessments, (3) physiological measures, such as pupil dilation and heart rate, and (4) behavioral measures, such as points of fixations, and keyboard and mouse usage. In an exploratory work, we proposed a new cognitive load measurement model based on behavioral data. Our data consisted in keyboard and mouse usage, as well as page views and fixation points from an eye tracker, and were collected in the context of an online Esperanto course. Our results showed that eye tracking data provided a better indication of effort than keyboard, mouse and page view data, and that a slight complementarity exists between these two types of information. In the same spirit, Larmuseau et al. (2019) investigated the correlation between the cognitive load and two physiological measures from smart watches: skin conductance and skin temperature. The participants were future school teachers taking a course as part of their training. One of their main findings is a moderate correlation between effort and skin conductance. However, both these last approaches are preliminary and only focused on small samples (less than 20 participants).
Purpose The purpose of this paper is to present the METAL project, a French open learning analytics (LA) project for secondary school, that aims at improving the quality of teaching. The originality of METAL is that it relies on research through exploratory activities and focuses on all the aspects of a learning analytics environment. Design/methodology/approach This work introduces the different concerns of the project: collection and storage of multi-source data owned by a variety of stakeholders, selection and promotion of standards, design of an open-source LRS, conception of dashboards with their final users, trust, usability, design of explainable multi-source data-mining algorithms. Findings All the dimensions of METAL are presented, as well as the way they are approached: data sources, data storage, through the implementation of an LRS, design of dashboards for secondary school, based on co-design sessions data mining algorithms and experiments, in line with privacy and ethics concerns. Originality/value The issue of a global dissemination of LA at an institution level or at a broader level such as a territory or a study level is still a hot topic in the literature, and is one of the focus and originality of this paper, associated with the large spectrum of different concerns.
Today’s online music services like Spotify provide their listeners with different types of music recommendations, e.g., in the form of weekly recommendations or personalized radio stations. Such recommendations are often based, at least in parts, on collaborative filtering techniques. In this chapter, we first review the different types of music recommendations that can be found in practice and discuss the specific challenges of the domain. Next, we discuss technical approaches for the problems of music discovery and next-track recommendation in more depth, with a specific focus on their practical application at Spotify. Finally, we further elaborate on open challenges in the field and revisit the specific problems of evaluating music recommendation systems in academic environments.
Modern music platforms like Spotify support users to create new playlists through interactive tools. Given an empty or initial playlist, these tools often recommend additional songs, which could be included in the playlist based, e.g., on the title of the playlist or the set of tracks that are already in the playlist. In this work, we analyze in which ways the recommendations of such playlist construction support tools influence the behavior of users and the characteristics of the resulting playlists. We report the results of a between-subjects user study involving 123 subjects. Our analysis shows that users provided with recommendation support were more engaged and explored more alternatives than the control group. Presumably influenced by the recommender, they also picked significantly less popular items, which leads to a higher potential for discovery. The effort required to browse the additional alternatives, however, increased the users' perceived difficulty of the process.
HAL is a multi-disciplinary open access archive for the deposit and dissemination of scientific research documents, whether they are published or not. The documents may come from teaching and research institutions in France or abroad, or from public or private research centers. L’archive ouverte pluridisciplinaire HAL, est destinée au dépôt et à la diffusion de documents scientifiques de niveau recherche, publiés ou non, émanant des établissements d’enseignement et de recherche français ou étrangers, des laboratoires publics ou privés. Projet METAL : Plan de collecte de données Azim Roussanaly, Thomas Toulotte, Laura Infante Blanco, Anne Boyer, Armelle Brun, Geoffray Bonnin
Die passende Musik fur einen gewunschten Anwendungszweck auszuwahlen, etwa fur eine Wiedergabeliste fur Hintergrundmusik oder als Untermalung in einem Werbespot, ist aufgrund von verschiedensten Anforderungen und der schieren Menge an verfugbaren Stucken ein aufwandiger Prozess. Es existieren zahlreiche Kriterien, beispielsweise Metadaten, aber auch die Beschaffenheit der Musik selbst, anhand derer ein Stuck charakterisiert werden kann. Mithilfe von Empfehlungssystemen – speziellen Algorithmen, die Elemente anhand festgelegter Kriterien auswahlen konnen – lasst sich dieser Prozess vereinfachen und teilweise automatisieren. Ihre Daten beziehen solche Systeme oft aus sogenannten Musikdatenbanken, die Informationen uber Musikstucke aggregieren und kategorisieren, und damit die Moglichkeit bieten, Titel nach verschiedenen Kriterien zu finden, dem Anwendungszweck gemas auszuwahlen und oft auch direkt zu erwerben oder abzuspielen. In diesem Kapitel wird das Problem der automatisierten Erstellung von Wiedergabelisten charakterisiert sowie algorithmische Ansatze im Uberblick vorgestellt. Anschliesend wird eine Ubersicht uber aktuelle Online-Musikdatenbanken gegeben.
In recent years, a number of approaches have been developed for the automatic recognition of music genres, but also more specific categories (styles, moods, personal preferences, etc.). Among the different sources for building classification models, features extracted from the audio signal play an important role in the literature. Although such features can be extracted from any digitised music piece independently of the availability of other information sources, their extraction can require considerable computational costs and the audio alone does not always contain enough information for the identification of the distinctive properties of a musical category. In this work we consider playlists that are created and shared by music listeners as another interesting source for feature extraction and music categorisation. The main idea is that the tracks of a playlist are often from the same artist or belong to the same category, e.g. they have the same genre or style, which allows us to exploit their co-occurrences for classification tasks. In the paper, we evaluate strategies for better genre and style classification based on the analysis of larger collections of user-provided playlists and compare them to a recent classification technique from the literature. Our first results indicate that an already comparably simple playlist-based classifiers can in some cases outperform an advanced audio-based classification technique.