
This paper presents KMagent, a multi-agent Large Language Model-based knowledge management platform designed to accelerate product innovation within industrial Research Development environments. The system integrates structured Knowledge Graph construction, LLM-driven Question Answering, and a comprehensive evaluation framework under a Retrieval-Augmented Generation architecture. By combining LLM reasoning with domain-specific KGs, KMagent produces accurate and context-aware responses to complex queries. To assess performance, we introduce a dual-layered evaluation framework covering both the structural integrity of the KG and the factual quality of generated responses using reference-based metrics such as BLEU, ROUGE, METEOR, BERTScore and reference-free metrics such as UniEval and G-Eval. Using polymer degradation as a case study, we demonstrate the system’s effectiveness across knowledge extraction, reasoning, and evaluation.
The complexity of combinatorial optimization problems typically leads to a steep performance decrease of solving approaches with increasing problem size. We explore the usage of machine learning (ML) to substantially reduce the search space size of large problem instances by predicting and removing unpromising parts in order to accelerate the subsequent solving process. More specifically, we explore graph sparsening techniques for the electric autonomous dial-a-ride problem (E-ADARP), where self-driving electric vehicles are used to provide an efficient and sustainable ride-sharing service by serving customer transportation requests between pickup and drop-off locations within specified time windows. Approaches utilizing support vector machines as well as gradient boosted trees are compared to and also combined with a common k-nearest neighbor heuristic. Our goal is to boost the performance of a state-of-the-art large neighborhood search for the E-ADARP to make it well applicable to instances with up to 5200 requests and 260 vehicles. Our ML models are trained on representative instances and close-to-optimal solutions obtained from excessively long runs in a weakly supervised fashion. We uncover challenges, a fundamental limitation, and benefits of this approach and are able to achieve the intended scalability.
This paper proposes a smart solution for on-street parking based on reinforcement learning. As part of a collaborative parking system, we show the ability of reinforcement learning (RL) techniques to find optimal strategies for guiding drivers toward available spots, especially in areas where free parking is scarce. We implement a proof of concept by first training an agent using a Q-learning method in a reduced-dimensional virtual environment to validate the feasibility of our approach. To simulate realistic urban dynamics, we design a multi-agent system comprising a city sub-model based on the Manhattan grid, partitioned in several areas, and a mobility model that reflects typical traffic flows and parking behaviors. The agent-based simulations are conducted using the NetLogo platform. From the resulting synthetic data, which capture the collaborative interactions between agents in search of parking, we build a dataset to train a deep reinforcement learning (DRL) model. Our experiments demonstrate the potential of offline DRL methods to learn efficient policies from simulated collaborative parking scenarios, contributing to the development of intelligent, data-driven urban mobility solutions.
To reduce the environmental impact of civil aviation, future aircraft configurations will be equipped with Very High Aspect Ratio Wings. Such wings enable drag reduction but their structural sizing aiming at weight minimization is a challenge due to high structural constraints. To reduce these constraints, different load alleviation strategies can be employed to reduce the loads for critical sizing cases. Positive and negative gusts are two of these critical cases but they require the measurement of the flow field in front of the aircraft in order to apply relevant load alleviation strategies. This can be done using a wind lidar. To determine the 3D wind field, the lidar is addressed along different axes to obtain the projections of the wind along each of them. In this paper, we present a methodology to incorporate a priori knowledge of the turbulence structure into the estimation procedure.
This paper proposes an advanced approach focused on the automatic tuning of fuzzy rules in a Self-Tunable Fuzzy Inference System (STFIS) network. By combining neural network principles with a zero-order Takagi-Sugeno-type FIS (Sugeno FIS), the approach enables continuous adaptation to external disturbances and system changes. The mini-JEAN control architecture optimizes the STFIS rule base in real-time, enabling efficient control of an underactuated, coupled system: the inverted pendulum. A backpropagation gradient descent learning algorithm is proposed to determine the STFIS network's weight values in real-time, supervised by an adaptive mechanism. A clustering algorithm is then applied to organize these values and extract decision rules from the weight data. The performance of the STFIS controller is compared with that of a fixed Rule-Based Fuzzy System (FPID controller). The results demonstrate that the STFIS controller outperforms the FPID controller in terms of efficiency and robustness, particularly when subjected to external disturbances. These findings confirm that the STFIS controller is a promising solution for real-time control applications.
With the increasing prevalence of dementia, timely access to assessments, diagnosis, and support is vital. However, the shortage of dementia specialists often leads to significant delays in a patient accessing time-sensitive treatments. To address this challenge, we proposed MemoryChat, a conversational dementia assessment tool designed to collect crucial information on the patient’s condition before the first clinical consultation. The Large Language Model-powered pipeline interacts with an informant to conduct a dynamic, multi-turn dialogue, collecting details. This interaction is transformed into a clinical summary and a preliminary diagnosis, serving as a reference for the clinician. Additionally, we introduce a novel clinical question dataset, 𝐌𝐞𝐦𝐨𝐫𝐲𝐂𝐡𝐚𝐭𝐐_𝐁𝐚𝐧𝐤 , which provides content to facilitate a comprehensive and standardised assessment of a person with suspected dementia. Our evaluation framework systematically assesses the clinical viability of MemoryChat by measuring the conversational, summarisation, and diagnostic capabilities. We compared different strong-performing Large Language Models, including Mistral:7b, LLaMA3.1:8b, and Qwen3:14b. We found that the configuration of LLaMA3.1-8b for interactivity and summarisation, alongside Qwen3-14b for diagnostic prediction, would formulate a potential option for MemoryChat to be used in a clinical setting.
Anomaly detection in cloud infrastructure faces significant challenges due to the massive scale of data and complex interdependencies among metrics. While traditional research has primarily focused on single-metric detection, modern cloud systems require understanding relationships across multiple time series. The scarcity of labeled anomalies at scale makes supervised approaches impractical, creating a significant gap between academic solutions and industry requirements. In this work, we evaluate state-of-the-art unsupervised multivariate anomaly detection algorithms on established benchmarks and identify their practical limitations in real-world deployments. To bridge the research-practice gap, we introduce a comprehensive dataset comprising 74.5 million data points across 1311 metrics in multivariate time series. Our dataset can support multiple tasks other than anomaly detection like forecasting, classification, explainability, and root cause analysis, serving both academic research and practical cloud monitoring needs. Through extensive empirical evaluation and practical insights, we demonstrate how theoretical approaches can be effectively adapted for production cloud environments.
In many real-world scenarios, an optimal solution for a decision maker depends on the response of another decision maker. This is known as bilevel optimization, which involves two levels of optimization. In this approach, the lower-level problem (the follower) appears as a constraint of the upper level problem (the leader). In practical applications, the follower may have multiple global optima, so the leader must consider the various assumptions about the decision of the follower. This uncertainty caused by the follower’s multimodality poses a significant challenge for the leader in making risk-averse decisions. To address this challenge, we propose an approach for the leader in a bilevel optimization problem. Our approach involves calculating the leader’s uncertainty for a given decision-making problem using a metric that quantifies the level of risk involved. This uncertainty is then used to determine the leader’s optimism and pessimism probabilities. To ensure that the leader is aware of the current risk situation, we impose constraints on their decision-making process. These constraints help the leader make informed choices that minimize the impact of uncertainty. To solve the bilevel optimization problem, we implement a black-box approach that utilizes Bayesian optimization (BO). This approach significantly improves the efficiency of the solution and allows the leader to make more informed decisions in the face of uncertainty. The performance of our approach is evaluated on two test benchmark problems, and we demonstrate a significant improvement in the leader’s uncertainty in these scenarios. Our approach provides a practical solution for decision-makers facing uncertainty caused by the multimodality of the follower problem.
The increasing demand for clean energy production has driven scientists and researchers to develop innovative and eco-friendly solutions that promote greener practices in the industry. Lithium-ion batteries have played a crucial role in energy storage due to their high efficiency, reliability, and seamless integration with renewable energy sources, making them a key component in sustainable power generation. Among these, lithium iron phosphate batteries stand out for their enhanced safety, longer lifespan, thermal stability, and environmental friendliness, making them an ideal choice for applications such as electric vehicles and grid storage. As artificial intelligence becomes increasingly integrated into the energy sector, its predictive capabilities are being leveraged to analyze and optimize lithium iron phosphate battery performance. A key aspect of this optimization is the estimation of State of Health, which is essential for accurate lifetime prediction and effective battery management. However, a major challenge with artificial intelligence, particularly deep learning models such as neural networks, is their black-box nature, making them difficult to interpret. Physics-informed neural networks offer a promising solution by embedding fundamental physical laws into the learning process, enhancing model interpretability while ensuring consistency with known battery dynamics. This paper proposes an interpretable Physics-informed neural network framework with a tailored loss function specifically designed for lithium iron phosphate batteries, enabling more accurate State of Health estimation, predictive maintenance, and performance optimization while maintaining transparency in decision-making. The proposed Physics-informed approach aligns with the global effort toward cleaner energy storage and responsible resource utilization.
This study proposes the application of three methods from the statistical literature for reconstructing the position of an occluded marker in Motion Capture (MoCap) data. Specifically, we investigate the use of Gaussian Process Regression (GPR), the Synthetic Control Method (SCM), and a Matrix Completion (MC) algorithm. Unlike traditional gap-filling techniques, which rely solely on the occluded marker’s time series, these methods exploit relationships between observed and missing markers to improve reconstruction accuracy. To assess their effectiveness, we apply these methods to a MoCap dataset comprising 3656 frames of a 3D human body performing approximately 17 distinct movements. A nested k-fold cross-validation framework is implemented using a selection of markers covering all major body parts. Performance is evaluated using the Mean Absolute Error (MAE). Our findings indicate that all three methods achieve satisfactory reconstruction accuracy, although computational efficiency varies significantly. The MC algorithm produces the most accurate results while also being the fastest method, whereas SCM exhibits the lowest accuracy. Finally, the methods are applied to real occlusions in the available dataset.
This paper addresses some limitations of the Cross-Entropy (CE) method in solving high-dimensional and complex optimization problems, particularly when using Gaussian distributions. We propose an alternative law-smooth update scheme to overcome these challenges and enhance convergence. This approach is well suited to the use of a variety of sampling laws, the laws of the exponential family, including Dirichlet distributions, which can be used for the optimization of permutation-invariant criteria. Specifically, we focus on permutation-invariant optimization problems, which often present difficulties due to multiple global optima and complex criterion variations. To resolve these issues, we introduce two approaches: one uses sampling laws to generate unique representatives for equivalent elements within permutations, while the other utilizes a balanced bijective mapping from a vector space to element classes within permutations. These methods are explored theoretically. Applications to benchmark functions (Rosenbrock and permutation-modified Rosenbrock) is presented. Finally is presented a real-world application to multispectral band selection for anomaly detection.
In this paper we show how leveraging the structure of the solution space of the Traveling Salesman Problem (TSP) can lead to a dramatic improvement of the performance of state of the art diffusion-based neural solvers. Building on recent approaches of DIFUSCO and T2TCO which pipeline a diffusion-based solution generation with a local search procedure, we propose IDEQ (constrained Inverse Diffusion and EQuivalence class-based training of diffusion models for combinatorial optimization). IDEQ improves the quality of the solutions by leveraging the constrained structure of the TSP state space. Indeed, the solution space consists of locally optimal Hamiltonian tours which is a much smaller space than the space of adjacency matrices used in previous works. Also, the local search procedure defines an equivalence class of Hamiltonian tours: all elements of this equivalence class reach the same local optimum after the application of the local search. This should be aligned with the supervised training objective of the diffusion. IDEQ addresses these two points. Our experiments show that IDEQ achieves 0.3
Graph Neural Networks (GNNs) have emerged as powerful tools for learning representations of structured data. An essential element of these models is neighborhood aggregation, where a node’s representation is updated based on its context (neighbors). Some variants of GNNs consider only local neighbor information during node updating. Ignoring global structural details leads to inadequate learning and differentiation of graph structures. To address these challenges, we introduce GraphSAGEnES (GSnES), a new graph neural network which employs a different aggregation mechanism based on similarity and entropy. Empirical results conducted on the Stochastic Block Model, a random graph model with distinct vertex communities, demonstrate that GSnES proves to be an effective method in classification tasks—particularly in graphs with low feature separability—demonstrating statistically significant improved performance in terms of accuracy, balanced accuracy, F1 Score, and Matthews correlation coefficient. This analysis contributes to developing more efficient aggregation mechanisms, potentially improving GNN architectures for various applications. In addition to evaluating these performance indicators, the model quantitatively assesses the respective contributions of similarity and entropy to the outcome. In this sense, the aim is to mitigate one of the main limitations of GNN models, which is their lack of explainability.
Non-invasive brain-computer interfaces (BCIs) can improve quality of life for individuals with motor disabilities. However, input data non-stationarity often leads to performance degradation over time. Adaptive BCIs aim to address this challenge. Recent studies have proposed leveraging error-related potentials (ErrPs) - neural signals elicited during self-made or observed errors - as a natural feedback mechanism for closed-loop systems. While these approaches demonstrate potential performance gains, their effectiveness relies heavily on accurate ErrP detection, which remains challenging. In a previous study, we introduced a novel reinforcement learning-based BCI framework incorporating ErrPs as reward signals. We systematically examine how misclassification rates in ErrP detection in terms of false positives (FPs) and false negatives (FNs) influence closed-loop performance. We use both synthetic and real datasets to evaluate two contextual bandit algorithms (LinUCB and NeuralUCB), trained to map motor imagery-related time-frequency modulations to actions in a binary task. Firstly, our findings show that the sensitivity of agents to FPs and FNs depends on baseline accuracies. In some conditions, such as insufficient exploration, false negatives might be more detrimental than false positives, but this needs further investigation. Proper agent parametrization and enough data samples can compensate the negative effects to some extent.
The field of Relation Extraction and its downstream applications are experiencing a considerable surge of attention due to the emergence of Large Language Models (LLMs), shifting the focus from Sentence-level to Document-level Relation Extraction. Despite this, employing LLMs introduces new challenges, such as the need for defining a prompting strategy and measures for controlling hallucinations. Moreover, limited input lengths and performance degradation on large text inputs, even when considerably below the limit, require careful considerations when prompting the model. This work aims to show how different prompting strategies, small changes to text processing, and even the choice of model and nature of the corpora, can yield vastly different results when performing this complex NLP task. In doing so, we highlight the necessity of establishing, assessing, and managing these elements in LLM-based Relation Extraction processes, which would eventually enable defining a baseline system applicable to both general and domain-specific corpora. Furthermore, by understanding these factors, we can enhance the reliability and accuracy of LLM-based Relation Extraction systems, leading to more robust applications across various domains.
Knowledge Base Population (KBP) aims to populate structured databases with facts extracted from text, encompassing tasks such as named entity recognition, coreference resolution, relation extraction, and entity linking. Traditional pipeline-based approaches sequentially chain modular components, leading to error propagation and unidirectional information flow. Additionally, black-box components often lack transparency and interpretability. In this paper, we propose a probabilistic pipeline framework for joint inference in end-to-end KBP. Our approach enables globally consistent decision-making by integrating local component feedback and external background knowledge. A key advantage is its ability to seamlessly incorporate knowledge about pipeline components, ontology constraints, linguistic patterns, and corpus characteristics. We evaluate our framework on two core KBP tasks: exhaustive relation extraction and entity linking.
Efficient energy management is critical for industrial consumers aiming to minimize operational costs while maintaining reliability. The electricity costs are crucial for companies’ competitiveness, and their reduction is important for the sector’s sustainability. This paper presents an optimization framework for determining the optimal contracted power for industrial energy consumers in Spain, leveraging historical smart meter data and available supply contract data. The proposed model evaluates historical consumption and billing structures to optimize contracted power, balancing fixed contracted power costs and penalization costs due to missing the optimal point. The study highlights significant cost savings, up to around 7
We consider a version of the Dynamic Electric Autonomous Dial-a-Ride Problem (DynEADARP) recently proposed and approached by means of a Genetic Programming (GP) hyperheuristic. This problem integrates the challenges of the static dial-a-ride problem with those of considering the charging of the fleet of electric vehicles and the dynamic nature of customer requests, i.e., the online aspect. The objective is to minimize the total travel time for serving all requests in a given time horizon together with penalties for late request pickups. We consider an optimization-based solution framework that utilizes a Large Neighborhood Search (LNS) for the static variant of the problem. This LNS is (re-)applied whenever new requests arrive and always considers the current state of the vehicles and all available, not yet served requests. Analyzing results of a first version in which the LNS performs in a pure myopic way that does not consider expected future requests, we observe two major weaknesses: (a) charging is often done much too late, and (b) it is not always good to head to planned pickup locations as early as possible. To address these weaknesses, we utilize reinforcement learning (RL) for learning two functions offline that are used within the LNS to incentivize earlier charging and waiting dependent on features of the current state and the expected number of future requests in a sensible way. An experimental evaluation shows that (a) the LNS-based approach can scale well to large instances with up to 10000 requests, (b) it outperforms the former GP hyperheuristic with a gap of up to 100
Requirements Engineering (RE) is a critical yet time-consuming phase in software development, often hindered by inconsistencies and inaccuracies. We propose a novel framework that combines Foundation Models (FMs) and Multi-Agent Systems (MAS) to enhance RE efficiency, accuracy, and quality. Our framework aims to automate routine tasks and provide intelligent assistance throughout all RE phases. In a preliminary Proof-of-Concept (PoC) experiment, we explored the capabilities of Large Language Models (LLMs) in RE, evaluating their performance on various interlinked tasks across multiple phases. Our results showed that LLMs can achieve human-like accuracy in assessing requirements deliverables, highlighting the importance of suitable LLMs for each phase. We discuss the implications of our findings and outline a research agenda to address the limitations of FM-based MAS in RE, focusing on key objectives such as data availability, model calibration, and human-AI collaboration. Our goal is to create a more efficient, collaborative, and reliable RE process.
The digitization of biological collections plays a crucial role in making scientific data more accessible and enabling long-term studies on biodiversity trends. While millions of herbarium specimens have been digitized worldwide, the extraction of metadata from handwritten labels remains a significant challenge. Traditionally, this task has been performed by experts or citizen scientists, ensuring high-quality data at the cost of substantial time and resources. In this paper, a system was developed that automates metadata extraction from herbarium specimen images. The system combines supervised deep learning techniques with a database-driven validation step and an open-source Large Language Model (LLM) to classify images of specific biological collections by collector, country, and year. It integrates an out-of-distribution detection mechanism using confidence scores, followed by a deep learning classification pipeline, and a database-supported LLM agent that refines the final predictions based on existing knowledge. The evaluation results demonstrate that a structured AI pipeline can streamline and improve the digitization of herbarium specimens by extracting metadata more efficiently than manual methods, all while preserving the level of accuracy required for scientific research.