
In this work, we propose a novel, drift-aware aggregation method for ensemble systems that dynamically adapts to changing data conditions while maintaining computational efficiency. Our approach leverages an adaptive error smoothing mechanism, a time-adaptive correlation penalty, and a dynamic mixing parameter to integrate a drift-aware weighting scheme that emphasizes recent performance and preserves model diversity. Comprehensive experiments on real-world datasets demonstrate that our method significantly outperforms traditional aggregation techniques and state-of-the-art algorithms, particularly in environments characterized by recurring concept drift and data scarcity. These results underscore the potential of our approach as a practical and scalable solution for adaptive ensemble aggregation in dynamic, resource-constrained settings.
As e-commerce continues to grow rapidly, the threat of online fraud has become a critical concern, posing significant risks to businesses and consumers alike. Traditional fraud detection systems struggle to keep pace with the evolving strategies of fraudsters, especially given the challenges of high-dimensional data, class imbalance, and limited interpretability. In this work, we introduce LEMON, a Light and Efficient fraMework for e-cOmmerce Fraud DetectioN, that uses Discriminating Code Set (DCS) in graph theory to attain an intrinsic set of features for enhancing the performance and efficiency of fraud detection models. We evaluate LEMON on the IEEE-CIS Fraud Detection dataset, comparing it with baseline techniques including manual feature engineering (Kaggle Top-1), XGBoost-based feature importance, and a hybrid LEMON+XGBoost model. Our results show that the approach not only achieves the highest ROC-AUC (0.976) and F1-score (0.78) but also significantly reduces training time. These findings highlight the effectiveness of DCS in selecting a compact, highly discriminative feature subset that improves fraud detection in real-world, large-scale e-commerce systems.
Mining frequent closed patterns is a fundamental task in both frequent pattern mining and association rule mining. In this research, we propose two efficient algorithms, tTRK-CloITP and dTRK-CloITP, for mining top‑rank‑k frequent closed inter‑transaction patterns (FCITPs) using tidset and diffset structures, respectively. Our approach offers several key contributions. First, the database is scanned once to generate 1‑patterns along with their corresponding tidsets. From this set of 1‑patterns, inter‑transaction 2‑patterns are generated using either the tidset or diffset structure without requiring another scan of the entire mega‑transaction database. This enables efficient inter‑transaction pattern generation starting from the 2‑pattern level. A depth‑first search (DFS) traversal is then used to iteratively generate (l + 1)‑patterns from l‑patterns. Second, we introduce additional pruning techniques based on either the tidset or diffset structure to efficiently identify closed patterns when inserting them into the top‑rank‑k table, which contains the complete set of top‑rank‑k frequent closed inter‑transaction patterns. Finally, we conduct extensive experiments on benchmark datasets from the FIMI repository ( http://fimi.uantwerpen.be/data/ ) to demonstrate the effectiveness and efficiency of the proposed algorithms in terms of both runtime and memory usage.
This paper describes a system in which an agent can find successful action strategies in the Wumpus World—a classic environment for testing logical reasoning. Despite such successful strategies can be fully described and implemented using first-order logic, it is challenging for reinforcement learning algorithms to automatically find them because of partially observable states, sparse rewards and the logic-based nature of the problem, which characterizes the Wumpus World environment. Our solution consists of the design of sensation maps, where partial observations are accumulated, shaping the reward function, custom two-stage ϵ -greedy action selection strategy, and curriculum learning. With these components, we were able to train agents using just a basic DQN algorithm—the pioneer of deep reinforcement learning. Our experiments confirm the good performance of the developed method improving the results described in the literature. Our solution brings reinforcement learning-based approaches closer to the complete solution of the problem designed to be solved using first-order logic.
In this work we tackle the problem of the data quality and labeling in the machine learning task of detecting fake news and disinformation. The major contribution of this paper is the new proposition to use large language models as an additional annotation mechanism in order to enrich the dataset and model with the new information. In this manner, we are able to faster annotate new content, and limit so called aging effect of the models. Hereby, we also evaluate our approach and provide promising results.
Continual learning faces the critical challenge of catastrophic forgetting when models learn new knowledge. In this paper, we introduce a method that is based on mechanics (CLM), an approach that applies conservation principles from mechanics to balance the learning dynamics between new and previously acquired information. By establishing analogies between model parameters and physical quantities (mapping Fisher Information to mass and parameter gradients to velocity), our method enables more effective parameter integration across sequential learning tasks. We conducted experiments using the ViT-B/16 architecture on CIFAR-100 and ImageNet-R datasets, where CLM demonstrates comparative performance with some existing related techniques. Our approach achieves a 0.31
Temporal Knowledge Graphs (TKGs) represent dynamic knowledge structures that encode time-sensitive relationships between entities, enabling systems to understand how facts evolve over time. Recent approaches to Temporal Knowledge Graph Reasoning (TKGR) have leveraged Large Language Models (LLMs), but face significant limitations: they often rely solely on first-order historical information, struggle with heavy information loads, and have yet to fully utilize LLMs’ potential for reasoning with semantically similar information. Additionally, current methods either lack interpretability or struggle with effective temporal rule learning. We present MSKGen (Multi-Source Knowledge-Based Generation), a novel query-aware approach for TKGR that integrates multiple knowledge sources to generate accurate predictions. By integrating rule-based facts with semantically retrieved facts, MSKGen maintains interpretability while maximizing LLMs’ semantic capabilities, addressing the information load challenges faced by current LLM implementations and offering significant advancements in combining structured temporal reasoning with semantic understanding for knowledge graph reasoning tasks. Experimental results across several common datasets demonstrate MSKGen’s superior performance, achieving significant improvements over state-of-the-art methods, confirming the effectiveness of our multi-source knowledge integration approach for temporal knowledge graph reasoning tasks.
Excessive central airway collapse (ECAC) is a pathological condition marked by significant narrowing of the central airway during expiration. Current diagnostic methods rely on visual estimation of the luminal area during dynamic bronchoscopy, which often leads to interobserver variability, affecting both diagnosis and treatment decisions. Such variability can be reduced by computationally estimating areas from a segmentation of airways. In order to have reliable measures, segmentation methods should be able to adapt to the varying illumination conditions of intra-operative videos using a limited and sparse amount of annotated frames. While machine learning (ML) models can be trained on few data, they do not have the adaptability of deep learning (DL) approaches trained on larger datasets. In this study, we introduce a system for endoscopic dynamics assessment that leverages a ML approach to a DL model for the dynamic assessment of luminal area in bronchoscopy videos. Our DL system is based on a U-Net model trained using a semi-automated annotated database built upon segmentations obtained from a ML method trained using only 100 manually annotated images. The U-Net is compared to the ML method in terms of temporal continuity in 10 videos and quality of the segmentations of the 100 annotated images. The adaptability of Unet to illumination conditions is assessed on an independent set of 24 videos. Results demonstrate that U-Net has higher generalization, establishing it as a more effective tool for dynamic video segmentation and, thus, the clinical assessment of ECAC.
Healthcare time series data is vital for monitoring patient activity but often contains noise and missing values due to various reasons such as sensor errors or data interruptions. Imputation, i.e., filling in the missing values, is a common way to deal with this issue. In this study, we propose folding the time series so that univariate time series can be turned into tabular data and using Multiple Imputation with Random Forest (MICE-RF) for imputation. Next, we compare this imputation strategy with state-of-the-art deep learning approaches (SAITS, BRITS, Transformer) for noisy, missing time series data in terms of MAE, F1-score, AUC, and MCC across missing data rates (10
In this study, the problem of motion-based artifact reduction in magnetic resonance imaging scans is considered. Motion-based artifacts caused by the patient’s body movements during the examination. The solution of the problem proposed in this work is based on deep learning methods directly applied to scans. The specific model used in our investigation was a Wesserstein generative adversarial network. The first module of the applied model was a discriminator that was used to classify real or fake scans. The second module serves as a generator and is used to produce fake scans. These models are trained iteratively with the use of ADAM optimizer which is an extension of stochastic gradient descent. The model was trained with 100 motion-free T1 scans with simulated motion artifacts. The corrected scans were evaluated using structural image similarity (SSIM), mean square error (MSE), and peak signal-to-noise ratio (PSNR), and the measures were compared with the scans without motion artifacts. In our results, we observed an improvement in the corrected scans for each of the measures used.
Phishing websites are designed to steal sensitive user information, causing financial damage. Detecting and blocking these websites before users provide personal data is a critical task in cybersecurity. This paper proposes HUGPhish, a phishing URL detection method using URL and HTML feature extraction. This method extracts the top potential n-gram features of URLs combined with handcrafted features extracted from them. HTML content is also analyzed and represented as graphs based on HTML tags. HUGPhish uses Graph Neural Networks (GNNs) to extract embeddings from HTML graphs. These embeddings, along with handcrafted HTML features and extracted URL features, are integrated into a comprehensive feature set, which is then classified using a LightGBM model for accurate phishing detection. The experimental results clearly demonstrate that HUGPhish outperforms all of the compared methods. With an F1 score of 90.92
The growing volume of data demands intelligent systems to process complex information, yet in business contexts, an excess of options can overwhelm customers. Existing recommendation systems address this through various approaches, notably Topic Modeling with Latent Dirichlet Allocation (LDA) [1] and neural network-based sequential recommendation models like Time Interval-aware Self-Attention Sequential Recommendation (TiSASRec) [2]. LDA can typically process customer reviews to generate topic-based user groups and in this way, we can recommend products within similar clusters. However, it is primarily designed for clustering rather than direct recommendation. In contrast, TiSASRec captures sequential user preferences by incorporating time intervals between interactions but often struggles with ranking relevant items. As each technique has its limitations, this research proposes a hybrid model that integrates the strengths of both approaches to improve recommendation accuracy in e-commerce transactions.
Traditional methods for creating robust machine learning models typically require extensive annotation, leading to a growing interest in leveraging semi-supervised learning for vision tasks. We hereby present a novel video action detection architecture capable of achieving high accuracy with minimal label requirements. Initially, the Video Swin Transformer employs local self-attention to extract meaningful video features. These features are then fed into a decoder pipeline inspired by UNet3+, effectively enhancing feature aggregation. Furthermore, skip connections between the encoder and decoder are reinforced by a transformer block. We evaluate our approach on two benchmarks, UCF101-24 and JHMDB-21, under annotation levels of 20
Understanding group communication dynamics is essential in fields such as psychotherapy, education, and organizational development. This paper introduces a novel method for automatically analyzing verbal interactions in group settings using large language models (LLMs). Drawing on Foulkes’ theory of group analysis, we define six semantic dimensions of communication and apply multivariate scoring to utterances. The method is validated on a real therapeutic session. By aggregating these scores over time and extracting dynamic indicators we characterize group development and compare model outputs to expert human assessments. Results show that while GPT-4o and Gemini Flash 2.0 demonstrate reasonable agreement with human ratings, they differ in temporal sensitivity and responsiveness. The study underscores the potential of LLMs in supporting group dynamics analysis.
Synthetic lethality (SL) is a genetic interaction in which the simultaneous perturbation of two genes leads to cell death, whereas perturbation of either gene alone is viable. Challenging in predicting synthetic lethality pairs usually comes from a highly biased source of data (only SL positive or only SL negative pairs) and insufficient data background (research data are usually single-omic). In this work, we create a large multi-omics dataset coming from diverse settings, including transcriptomics profiles, genetic perturbations (i.e., RNA interference, Clustered Regularly Interspaced Short Palindromic Repeats), and genomics data (i.e., nucleotide sequences and amino acid sequences). Here, we also propose a multimodal model containing different modules specifically designed to capture biological insights of each type of omic data. Quantitative results show that our method has a top-notch specificity in predicting synthetic lethality despite its simplicity compared to other advanced techniques like graph-based models with a specificity of 99.75
Deep Neural Networks (DNNs) are being developed and applied in various image-related tasks. However, their predictions can be easily manipulated by adding small perturbations to the original images, creating adversarial images that are considered difficult for humans to distinguish. This paper proposes an adversarial attack method for black-box models, implemented in two attack phases. The first phase is a full-image attack (L2-loss adversarial attack) to approximate the gradient, using Principal Component Analysis (PCA) or Truncated SVD (Singular Value Decomposition) to reduce the number of queries. The attacked images from this phase are then used to identify important image regions that influence the model’s decision. The second phase attacks the already adversarial images from the first phase, but focuses only on the important regions, aiming to reduce L2-loss. The results on the CIFAR-10 dataset with the ResNet model (92.31
Currently, we live in a world shaped by AI and ML solutions. It is no longer surprising that, in our everyday online activities, various algorithms tailor the content we browse to our preferences. These are often incomplete or inaccurately inferred from our past interactions. This situation often leads to the problem of the information bubble, as AI-driven content personalization reinforces our existing views and limits access to diverse perspectives. The goal of this paper is to analyze existing mechanisms that address this problem. We approach the issue from multiple angles, considering algorithmic aspects as well as fairness and ethical concerns.
Current air quality monitoring systems struggle to integrate heterogeneous historical data into AI algorithms that could provide deeper insights into severe pollution impacts. We present PANORAMA, a knowledge graph approach that predicts links between pollutant exposure and health outcomes by integrating diverse air quality and health datasets. Using an Extract, Transform, and Load process, we incorporated discrete data into a knowledge graph and evaluated its inferential capabilities through embedding models. Our case study in France’s Gironde department yielded promising results (MR: 11.0, MRR: 0.403, Hits@1/3/10: 0.326/0.447/0.499) using the Rotate model, demonstrating the potential of semantic structuring and knowledge graph technology for environmental public health applications.
Topic modeling of Vietnamese legal documents presents unique challenges due to the complex hierarchical structure of Vietnam’s legal system and the distinct linguistic characteristics of the Vietnamese language. In this paper, we propose a novel ensemble framework that combines the strengths of two complementary approaches: Topic Variational Autoencoders (Topic VAE) and Sentence-BERT (SBERT) as tokenizer. Our framework addresses domain-specific challenges through specialized Vietnamese text preprocessing and a weighted integration mechanism that balances probabilistic modeling with contextual semantics. Experiments conducted on a dataset of 32 Vietnamese legal documents comprising 9,443 articles demonstrate that our ensemble approach outperforms traditional methods, achieving superior coherence and diversity scores compared to Latent Dirichlet Allocation (LDA) and BERTopic baselines. The proposed system offers practical benefits for Vietnam’s ongoing digital transformation initiatives by improving document organization, search functionality, and information accessibility within legal institutions.
Computed tomography (CT) imaging is essential in medical diagnosis, with image quality heavily dependent on the selection of reconstruction kernels. Sharp kernels improve spatial resolution but increase noise, while soft kernels diminish noise at the expense of edge definition. This paper introduces a unique Variational Mode Decomposition with Quaternion Bilateral Filtering (VMD-QBF) method to convert sharp-kernel CT images into soft-kernel versions while maintaining critical structural information. The suggested method is assessed in comparison to conventional denoising techniques, such as Non-Local Means, Anisotropic Diffusion, Bilateral Filtering, and Quaternion Bilateral Filtering (QBF), utilizing various reconstruction kernels (B50, B46, B41, B36, B35, B31). The evaluation is performed with Mean Squared Error (MSE), Structural Similarity Index (SSIM), Multiscale SSIM (MS-SSIM), and Peak Signal-to-Noise Ratio (PSNR). Experimental findings indicate that VMD-QBF surpasses traditional filtering methods, attaining minimal MSE and maximal PSNR, while preserving enhanced structural similarity across all evaluated kernels. The results validate the efficacy of the suggested strategy in reducing noise while maintaining essential image characteristics, positioning it as a viable solution for post-reconstruction CT image enhancement.