
Temporal information is pivotal to real-world events, where events, roles, and relationships are inherently dynamic over time. Neglecting temporal structures in retrieval-augmented systems leads to chronological misalignment of evidence, ultimately compromising the accuracy of fact-checking for time-sensitive data. We present T-HiGra, a temporal graph RAG framework that treats time as a first-class signal across hierarchical knowledge graphs and the retrieval pipeline for open-domain, multi-hop question answering. T-HiGra (i) attaches structured temporal information data and provenance to entities, relations, passages and edges; (ii) preserves historical fidelity via alias validity windows during entity merging; and (iii) applies an adaptive time-aware ranking method that jointly combines temporal proximity and semantic similarity. We evaluate T-HiGra on a benchmark for time-sensitive questions, and report multi-facet RAG diagnostics using LLM-as-judge evaluators alongside traditional QA metrics. Against strong baselines (BM25, HippoRAG,...), we achieve the highest correctness (76.21
Predictive Process Monitoring (PPM) leverages historical event data recorded during historical business process executions, to forecast future scenarios of new ongoing process executions. Traditional PPM methods assume that business processes remain in a steady state over time. However, this is often not the real-world case due to concept drifts. This paper presents a PPM method, named ATLAS, which uses an online learning strategy coupled with a LLM, to develop a LSTM model for predicting the next-activity of ongoing process executions (running cases). It keeps the PPM model continuously updated each time a new event is observed. The evaluation assesses the performance the proposed approach in terms of accuracy and computation time also compared to several related online PPM approaches.
Distributed Acoustic Sensing (DAS) technology has emerged as a powerful tool for large-scale acoustic monitoring, transforming standard fiber optic cables into dense arrays of virtual microphones. When combined with artificial intelligence, particularly deep learning, DAS enables scalable and automated detection of acoustic events, making it a promising solution for whale monitoring across vast marine environments. In the field of marine bioacoustics, DAS provides significant advantages in terms of spatial coverage and robustness compared to traditional hydrophone arrays. This paper presents a novel deep learning-based approach to detect whale vocalizations from DAS data. The proposed method leverages a hybrid architecture combining Convolutional Neural Networks (CNN) and Bidirectional Long Short-Term Memory (BiLSTM) networks to extract both spatial and temporal features from waterfall diagrams derived from DAS recordings. A multi-stage preprocessing pipeline including frequency filtering and frequency-wavenumber (f-k) filtering is applied to enhance signal quality and isolate whale calls from background noise. Each DAS recording is treated as a spatio-temporal matrix, and sequences of such matrices are sequentially analyzed capturing temporal dependencies. Experimental evaluations on a public dataset from the Ocean Observatories Initiative (OOI) RCA North Cable demonstrate that the CNN-BiLSTM model outperforms CNN and CNN-LSTM baselines, achieving a F1-score of 96
Aviation safety organizations accumulate large repositories of incident reports whose operational knowledge remains difficult to extract systematically at scale. Standard topic modeling applied to raw documents surfaces distributional patterns anchored to surface form rather than the causal structures relevant to safety analysis, and industrial constraints further limit viable approaches. We propose a pipeline decoupling representation shaping from structure discovery: an open-weight LLM summarizes the cause of each incident before topic discovery on the resulting representations, with a weighted multi-dimensional quality filter acting as an automated safeguard against hallucination and aspect drift. Applied to a large corpus of NTSB incident reports, cause-focused representations outperform raw-document baselines in topic coherence, and alignment with expert taxonomy. These results establish representation shaping as the primary lever for aligning discovered topic structures with operational safety objectives, opening concrete deployment perspectives for aviation knowledge management.
Weather radar mosaics provide high-resolution precipitation estimates but are often degraded by beam-blockage artefacts caused by terrain or infrastructure. These artefacts introduce systematic attenuation of radar reflectivity and distort precipitation patterns, particularly in complex topographic environments. Traditional correction approaches typically rely on geometric modeling or multi-radar fusion and often treat detection and correction as separate problems. In this work, we propose a guided conditional diffusion framework for correcting radar beam-blockage artefacts. The model reconstructs the precipitation field through a conditional reverse diffusion process guided by the corrupted radar observation and its associated quality map. During training, beam-blockage patterns are simulated on clean radar sectors to generate corrupted observations while preserving the original precipitation fields as references. This enables the model to learn both the spatial characteristics of blockage and the structure of precipitation fields. The method is evaluated on real radar mosaics containing both simulated and naturally occurring beam-blockage effects. Quantitative results show that the proposed approach reduces reconstruction errors within blocked regions while preserving precipitation values in unaffected areas. Qualitative analysis further demonstrates the recovery of coherent precipitation structures in degraded regions.
Accurate classification of electroencephalographic (EEG) signals in multi-subject Brain–Computer Interfaces (BCIs) is challenged by high inter-subject variability. When data from different individuals are aggregated, distribution shifts and subject-specific patterns may introduce noise and reduce the ability of machine learning models to learn robust representations. This issue is particularly critical in Steady-State Visually Evoked Potential (SSVEP)-based BCIs, where precise frequency-specific responses must be reliably detected. In this work, we propose a Mixture-of-Experts (MoE) framework for SSVEP classification. Each expert is a Multi-Level ResNet trained on one or multiple subjects, while a gating network learns to combine experts’ predictions. The framework is evaluated on a real-world multi-subject SSVEP dataset acquired using a Bitbrain Diadem IoT device. Results show improved classification performance and computational efficiency compatible with real-time deployment on devices with limited processing capabilities.
Evolving a malware into a family is an effective technique to hinder detection mechanisms. A recent trend exploits generative models to support threat actors in the creation of “mutations” for rapidly preparing malware families. However, artificial intelligence can also be used to develop effective countermeasures. To this end, we propose MalARN, a deep learning-based solution for creating synthetic representations of malware variants to make detectors more robust and facilitate spotting never-seen threats. MalARN takes advantage of a pre-trained large language model to map both malicious and benign binary samples into embeddings. To bypass the requirement for executable binaries, an adversarial reconstruction network is used to operate directly in the embedding space. Evaluated against four real malware families, MalARN outperforms the baseline solution in terms of specific metrics for unbalanced scenarios.
When Retrieval-Augmented Generation (RAG) systems are used in educational settings, ensuring reliability and transparency becomes essential. However, these systems can sometimes generate information that is not properly grounded in the underlying knowledge sources, and the reasons behind these failures remain insufficiently understood. In this paper, we introduce a diagnostic framework designed to analyze hallucination phenomena in RAG-based educational recommendation systems. Our approach distinguishes between two main sources of errors: those originating from pipeline-level issues, such as generation artifacts or catalogue traceability problems, and those related to retrieval grounding. To better understand these behaviors, we evaluate system performance under different retrieval depths, which allows us to identify two distinct phenomena: structural hallucinations and depth-sensitive generalization. Experiments conducted on a real educational catalogue show that evaluation protocols and retrieval depth strongly influence how hallucinations are diagnosed. These results highlight the importance of careful evaluation when deploying RAG systems in educational environments. Our framework provides practical methodological guidelines for analyzing groundedness and improving the reliability of educational AI systems.
Accurate indoor radio coverage mapping is essential for coordinating multi-agent robotic systems, like MARS (Multi-Agent Robotic Systems), and ensuring reliable wireless communication in complex environments. This paper presents a method to validate and improve cellular signal coverage maps predicted via physics-based ray tracing using robot-collected measurements. A mobile robot performs Simultaneous Localization and Mapping (SLAM) while recording signal metrics such as Received Signal Strength Indicator (RSSI) and Signal-to-Interference-plus-Noise Ratio (SINR). An initial coverage map is generated using a differentiable ray tracing simulator (Sionna RT), and residuals are defined as the difference between measured signals and simulated predictions. We observe that these residuals correlate with simple geometric features. Therefore, we propose a lightweight, structure-aware model that predicts the residuals based on wall count and distance, and integrates them into the ray-tracing output to obtain a corrected map. Experiments conducted in multiple indoor environments using a 5G small-cell transmitter demonstrate that the proposed approach significantly improves prediction accuracy compared to ray tracing alone.
Ocode develops a digital maintenance log system based on a low-energy blockchain infrastructure, ensuring traceability and secure data sharing for assets such as vehicles, bicycles, housing, and high-value goods. With the emergence of generative AI, these logbooks are evolving into intelligent assistants capable of extracting relevant information, detecting inconsistencies, and supporting proactive asset management. The current implementation relies on cloud-based Large Language Models (LLMs) combined with Retrieval-Augmented Generation (RAG). Although effective, this architecture raises concerns regarding data privacy and sovereignty, dependence on external providers, and the environmental cost of LLM inference. To address these challenges, we explore the use of Small Language Models (SLMs), which can be deployed in controlled environments, offering improved data governance and a lower carbon footprint while maintaining competitive performance. However, the adoption of SLMs is hindered by the lack of standardized evaluation methodologies. This paper introduces a carbon-aware evaluation framework based on two key metrics: response accuracy and carbon footprint per request. The methodology supports informed industrial decision-making when choosing between local SLMs and external LLM services, and is validated using the IFEval benchmark to ensure reproducibility. Our results highlight the trade-offs between performance and environmental impact, demonstrating that mid-sized models offer a strong balance between instruction-following capabilities and operational sustainability, paving the way for more frugal and trustworthy generative AI systems in industrial maintenance contexts.
Cancer patients undergoing immunotherapy experience diverse and evolving Quality of Life (QoL) trajectories that can significantly impact treatment outcomes. This study presents a machine learning framework for predicting QoL deterioration by integrating patient clustering with symptom classification over an 18-month follow-up period. We analyzed longitudinal data from the QUALITOP real-world cohort, comprising cancer patients from two European hospital partners. Our methodology introduces several key innovations: a temporal aggregation approach consolidating variable-length trajectories into predictive features; direct analysis of patient-reported FACT-G item responses rather than composite scores, enhancing model explainability while maintaining high accuracy performances; dual-perspective analysis clustering patients into distinct QoL trajectories while identifying symptom patterns simultaneously; and empirical validation through ensemble comparison of multiple clustering algorithms and classification methods. Our findings demonstrate that patient-centric, temporal-aggregation-based analysis of QoL data can effectively support clinical decision-making in oncology.
Lifelong machine learning systems continuously learn in dynamic environments, using experience from previous tasks to improve future learning and prediction. In multi-label lifelong learning, the model assigns multiple labels to each instance and updates its knowledge progressively as new tasks arrive. The aim is to maintain accurate predictions for earlier tasks while integrating new ones. However, combining lifelong learning with multi-label prediction presents a particular challenge in retaining past knowledge due to the high diversity and overlap in label semantics. To address this, we introduce Memory-based Diffusion Replay (ReMemDiff), a framework that learns from multi-label image tasks while reducing catastrophic forgetting. Our approach uses stable diffusion, a generative model that synthesises high-quality images via a progressive denoising process. It also incorporates variational autoencoders (VAEs) for efficient latent space operations, as well as a U-Net for precise denoising in the reverse diffusion process. This is guided by textual embeddings that are generated using the bootstrapping language image pre-training (BLIP) model. The ReMemDiff framework uses a convolutional neural network (CNN) to assign multiple labels to images. Selected image samples are converted into textual representations and stored in a buffer. These are then used to generate synthetic images interleaved with real images from the current task. This ensures that the generated samples reflect past tasks and reinforce knowledge retention. Experimental evaluations on two multi-task, multi-label image datasets validate the effectiveness of our approach.
Business Process Deviance (BPD) refers to the phenomenon of business process executions (traces) that deviate with respect to the desirable outcome. BPD detection is a Predictive Process Monitoring (PPM), which covers predictive methods designed to distinguish deviant traces from non-deviant traces. In this paper, we focus on the task of predicting BPD cases without waiting for the completion of process executions. To this purpose, we describe a PPM methodology, named FIREFOX, to monitor ongoing traces of a business process, to early identify process executions that are expected to deviate from the desirable outcome. FIREFOX equips BPD predictions with counterfactual recommendations of actions to perform in the near future, to avoid the detected deviance risk. A preliminary evaluation explores the effectiveness of the proposed methodology on traces recorded for some business processes.
Patients in oncology face challenges regarding their rehabilitation during and after the treatment. Development in the field of Clinical Decision Support Systems (CDSS) offer nowadays new opportunities to transform complex medical guidelines into personalized, and intuitive physical activity recommendations. In this paper, we present a hybrid neuro-symbolic approach that combines LLM-RAG with Knowledge Graph, to recommend adapted physical activity (APA) to patients in oncology. Our solution integrates a recommendation module that leverages a fine-tuned LLM algorithm for generating tailored APA recommendations using a Retrieval Augmented Generation (RAG) module, to ground the generation process in a clinical context, thereby mitigating the risk of LLM hallucination, and a Knowledge graph integrating the SNOMED-CT ontology to validate the LLM output. The framework is built in a modular way following Service Oriented Architectures (SOA) to allow multiple sources data integration, and was implemented in a mobile application prototype for validation. The recommendation process responds dynamically to patient data and context changes, demonstrating a workflow that continuously adapts to patient profiles, preferences, and other relevant information (This work is being carried out in the context of the regional project SQVALD: http://oncocentre.org/wp-content/uploads/4_SQVALD-08122022-Journ .).
Humans possess a unique ability to recognize and categorize geometric patterns. Cognitive neuroscience suggests this sensitivity may reflect a structured symbolic “Language of Thought.” Recent findings challenge this symbolic-only account: similar sensitivity can emerge in large neural networks trained on massive visual data. We propose a neuro-symbolic model showing that human sensitivity to geometric regularity can be explained by a small set of grounded geometric primitives, without invoking either a full symbolic grammar or large-scale statistical learning. In a series of three simulations, we show that our model allows to explore which set of fundamental geometric features best explains the patterns of human behavior in a Geometric Intruder Task.
Foundational time-series models, that is potentially large and pre-trained architectures trained on diverse temporal data, offer a promising new paradigm for building energy management systems (BEMS), where accurate load forecasting is essential yet often constrained by limited appliance-level data. This study investigates the zero-shot forecasting capabilities of several state-of-the-art foundational models when applied to appliance-level load forecasting across heterogeneous appliances and buildings. Without any task-specific fine-tuning, these models are evaluated on their ability to generalize to unseen appliances, varying consumption patterns, and diverse operational contexts. The analysis spans multiple building types and device categories, reflecting realistic BEMS deployment scenarios. Results highlight the extent to which broad temporal priors embedded in foundational models can substitute for traditional, data-hungry forecasting pipelines. The findings reveal both the strengths and current limitations of zero-shot approaches, offering insights into their practical viability and outlining pathways for integrating foundational models into next-generation energy management systems.
One of the most enduring and difficult issues in data analysis and machine learning is missing data. A missing value corresponds to an unobserved entry in a dataset that would otherwise contain valuable information for statistical inference or predictive modeling. The presence of missing values in a dataset affects both the performance of machine learning models and the reliability of interpretation methods. In this work, we propose a Local Reliability Index (LRI) that measures the local reliability of model predictions by combining the SHAP importance of each feature with a confidence score on observed or imputed data. Missing values are imputed using missForest, a robust non-parametric method based on random forests. In its initial version, LRI uses a constant mask for imputed values, limiting interpretability. We then introduce an improved uncertainty-aware version, where the mask for imputed values is determined in a data-driven manner. While the uncertainty-aware approach does not necessarily increase the average LRI dramatically, it produces a score that is more robust, interpretable, and consistent with the actual uncertainty level, particularly for high missing rates. The proposed LRI offers a practical tool to guide the analysis and interpretation of machine learning models on incomplete datasets.
When building a process model from an event log, behavioural process querying methods often focus on analysing an execution’s control flow. This is because event logs typically do not encode data flow through a process explicitly. Domain knowledge about the effects of activities on objects and relations between objects can be modelled using condition-effect rules, and exploiting path querying in such rules enables us to determine the values of dynamic object attributes that are indirectly related to events. This allows us to leverage temporally local knowledge of a system’s evolving data state to generate timelines from process instances. Within this setting, our main contribution is two-fold: (i) We enrich previous models with a new form of rules that allows for dynamic attribute value changes to be computed and recorded as part of the timeline. (ii) We provide a formalization of rule semantics using a transition system, in which states represent data configurations, and transitions represent the updates to those configurations resulting from the rules’ effects. Finally, we provide a formal construction that generates a timeline that is correct with respect to the aforementioned rule semantics.
With the rapid advancement of modern technologies, exponential data growth has become an unavoidable reality, which presents major obstacles to knowledge extraction. Gradual patterns enable the discovery of meaningful relationships within such data, expressed in the form “more/less X, more/less Y,” where X and Y denote data attributes. These patterns play a crucial role in supporting data-driven decision-making processes. Although several approaches for gradual pattern discovery have been proposed in the literature, the computational cost, particularly in terms of execution time and memory consumption, remains a major challenge, even with the availability of parallel solutions. In this paper, we propose the use of Principal Component Analysis (PCA) as a preprocessing technique to accelerate gradual pattern discovery by reducing data dimensionality. The proposed approach is evaluated on multiple datasets from diverse application domains and compared against the GRAANK and GRAANK-based metaheuristic algorithms (PSO, GA, PRS). Experimental results demonstrate a substantial reduction in both execution time and memory usage.
Extracting fine-grained technical entities from unstructured automotive service reports is critical for data-driven fault diagnosis, yet remains challenging due to jargon-heavy language and limited labeled data. Zero-shot Named Entity Recognition (NER) models such as GLiNER enable rapid deployment without domain-specific training, but their effectiveness in specialized technical domains is unclear. We hypothesize that while zero-shot models perform well on broad-domain benchmarks, they fail in low-resource domain-specific settings where domain-adapted fine-tuning is required. To evaluate this, we compare zero-shot GLiNER with fine-tuned RoBERTa and WG-BERT. Models are evaluated on clean synthetic data, noisy synthetic data, and a gold set of manually annotated real-world service reports from a Swedish truck manufacturer. Results show that GLiNER enables fast prototyping but achieves substantially lower F1-scores, while WG-BERT fine-tuned on synthetic data consistently outperforms RoBERTa and general-purpose zero-shot models. To the best of our knowledge, no publicly available NER model is currently tailored to fine-grained automotive technical entities. Consequently, we release our datasets and source code to facilitate further research in this domain(Project code and datasets are available at https://github.com/Adeelzafar/From-Zero-Shot-to-Domain-Precision .).