Designing data services in Cyber-Physical Production Networks (CPPN) involves managing service composition across organisational boundaries, supporting both vertical and horizontal integration, while ensuring data sovereignty, access control, and regulatory compliance. In this paper, we present D3SD (Three-layer Data Services Designer), a web-based tool for designing data services using a three-layered service model that distinguishes between different levels of the production network. D3SD provides process designers in CPPN with a BPMN-based graphical environment for modelling atomic data services that operate on data within individual supply chain actors. It also supports their composition at the smart factory and supply chain levels by leveraging a variant of the Process-to-Services (P2S) approach to suggest groupings of atomic data services that maximise internal cohesion while minimising coupling between aggregations. To ensure privacy-by-design and data sovereignty, the tool integrates Role-Based Access Control (RBAC) and GDPR responsibilities directly into the design process. We demonstrate D3SD through a multi-actor supply chain scenario derived from an industrial use case in a research project.
Enterprise information systems increasingly rely on LLM-based pipelines, yet they offer no systematic guarantees on the consistency of generated outputs with domain knowledge, constraint compliance, or execution provenance. This paper presents OntoLLM, a conceptual framework that structures LLM-based pipelines as governed and traceable processes through three layers: a Conceptual Layer grounding execution in ontologies and constraints; a Procedural Layer defining step types, executors, templates and trust requirements; and a Monitoring and Trust Layer assessing trust requirements at each execution step. We illustrate its potential through two representative scenarios, namely knowledge graph creation and personalised recommendation, showing how the architecture is conceived to enforce trust requirements at each step of an LLM-based pipeline. OntoLLM is positioned as a conceptual foundation for engineering trustworthy and semantically governed LLM pipelines.
Big Data Analytics (BDA) techniques have proven effective in managing the volume, velocity, and heterogeneity of industrial data streams, enabling fault detection and anomaly monitoring through scalable processing of high-frequency sensor measurements. Nevertheless, current approaches offer limited interpretability of BDA outcomes and overlook the valuable knowledge embedded in maintenance reports, logs, and technical documentation, thereby hindering the integration of datadriven insights with domain knowledge. Large Language Models (LLMs), with their capabilities in extracting knowledge from unstructured text and generating context-aware explanations, open the opportunity to rethink how quantitative analytics and semantic knowledge can be combined in industrial BDA. This paper presents a vision for a four-tier BDA architecture that addresses both limitations through the integration of ontologyguided Knowledge Graph construction and quantitative stream analytics. We motivate the design through a real Smart Manufacturing scenario, illustrate its analytical potential via three exploration scenarios, and identify the open challenges that a mature realization of this vision must confront.
In the context of Smart Manufacturing and the Internet of Production, data service discovery plays a central role in enabling cross-organizational collaboration and data-driven innovation. Nevertheless, effective discovery and composition of data services often require deep technical knowledge, limiting the autonomy of domain experts and R D managers in designing analytics workflows. This paper presents a cooperative approach to data service discovery that combines Large Language Models (LLMs) with Retrieval-Augmented Generation (RAG), leveraging a conceptual model of data services and analytics scenarios. On top of this model, a set of prompting strategies are designed to support different levels of user expertise and interaction goals. These strategies leverage the cooperative nature of the approach, enabling domain experts and R D managers to incrementally build, extend, and refine analytics data service pipelines through the interaction with LLMs. We describe how these prompting strategies are tightly integrated with the RAG components to inject contextual knowledge derived from a catalog of data services and analytics scenarios. The system is implemented using open-source technologies and evaluated extensively in a real-world smart factory case study. Our evaluation includes both quantitative metrics (precision, recall, faithfulness, factual correctness) and a qualitative user study, demonstrating the effectiveness of prompting strategies and the feasibility of LLM-supported data service discovery in cooperative industrial settings.
Food recommender systems must integrate heterogeneous, multi-level knowledge, including nutritional, environmental, and sensory aspects. However, sensory perception, despite its key role in user satisfaction, remains underinvestigated. In this work, we propose a web-based food advisory architecture that includes a Sensory Analysis Ontology, aligned with ISO standards, and embedded within a modular Retrieval-Augmented Generation (RAG) pipeline. Our main contribution is the design and engineering of a web-based system that combines symbolic and sub-symbolic components to enable: (i) ontology-driven semantic indexing of food-related data, (ii) query rewriting using domain-aligned vocabulary, and (iii) a hybrid retrieval layer combining dense vector search with Text-to-SQL over structured sensory databases. This design enables recommendations that match user’s taste preferences based on certified panelists evaluations. Experimental evaluation, performed using a benchmark of real-world, multi-domain queries across nutrition, sustainability, and sensory domains, shows that full integration of ontological knowledge across indexing, retrieval, and generation phases improves response quality, achieving up to +34.6
System logs represent a valuable source of cyber threat intelligence (CTI), capturing attacker behaviors, exploited vulnerabilities, and traces of malicious activity. Yet their utility is often limited by a lack of structure, semantic inconsistency, and fragmentation across devices and sessions. Extracting actionable CTI from logs, therefore, requires approaches that can reconcile noisy, heterogeneous data into coherent and interoperable representations. We introduce OntoLogX, an autonomous AI agent that leverages large language models (LLMs) to transform raw logs into ontology‐grounded knowledge graphs (KGs). OntoLogX integrates a lightweight log ontology with retrieval augmented generation and iterative correction steps, ensuring that generated KGs are syntactically and semantically valid. Beyond event‐level analysis, the system aggregates KGs into sessions and employs a LLM to predict MITRE ATT&CK tactics, linking low‐level log evidence to higher‐level adversarial objectives. We evaluate OntoLogX on both public and real‐world honeypot datasets, demonstrating robust KG generation across multiple LLMs backends and accurate mapping of adversarial activity to MITRE ATT&CK tactics. Results highlight the effectiveness of the methodology in constructing ontology‐compliant KGs, along with their value in extracting actionable CTI.
The rapid digital transformation of industrial ecosystems is turning manufacturing environments into interconnected Cyber-Physical Production Networks (CPPN), where coordination across multiple actors is driven by continuous data exchange. In this context, Service-Oriented Architectures provide an effective paradigm for exposing modular, reusable, and interoperable services that encapsulate data access, processing, and sharing capabilities across organizational boundaries. Despite their central role, identifying and structuring such services remains challenging in multi-actor CPPN due to data sovereignty constraints, heterogeneous systems, and distributed processes. This paper presents an interactive visual tool for data service identification in CPPN, which implements and extends the P2S (Process-to-Services) methodology to explicitly account for data dependencies and organizational boundaries of CPPN. Starting from Business Process Model and Notation (BPMN) process descriptions, the tool combines hierarchical clustering with interactive visual representations to generate and explore multiple service aggregations, enabling designers to reason about tradeoffs between cohesion, coupling, and data governance constraints. By supporting the interactive exploration of alternative solutions, the tool instantiates a visual knowledge management paradigm that bridges implicit process knowledge embedded in BPMN models and explicit architectural decisions, enabling human-in-the-loop decision-making with reduced cognitive complexity.
Agri-food supply chains are complex and distributed ecosystems involving heterogeneous actors, from primary producers to retailers and consumers. Ensuring information consistency and coordination among these actors is essential to improve traceability, efficiency, and transparency. In this context, Blockchain Technology (BCT) has emerged as a promising enabler for trustworthy and decentralized data management. Its ability to provide immutable and transparent records makes it particularly suitable for supporting traceability and accountability in agri-food processes. However, challenges emerge in complex, intertwined supply chains, where different chains may adopt heterogeneous blockchain platforms, each characterized by its own technological infrastructure, smart contract language, and approach to on-chain/off-chain data management. These differences hinder interoperability, scalability, and cost-effective data sharing. This issue is particularly relevant in agri-food ecosystems, where it is common for the same actor to participate in multiple intertwined supply chains. Despite its importance, the considered problem is still underexplored in the literature. We propose MAIA, a model-based methodological approach for blockchain Integration in intertwined Agri-food supply chains, covering the full lifecycle, from requirements elicitation to technology-agnostic modeling of resources and services, and finally to blockchain-specific implementation. Based on a batch-oriented model, the approach ensures a resource-oriented perspective and scalable integration across heterogeneous platforms. A proof-of-concept case study, validated on both Ethereum and Hyperledger Fabric, demonstrates its effectiveness in addressing integration challenges and, at the instantiation level, optimizing resource management in terms of costs, execution time, and scalability.
Supporting consumers in making autonomous food choices that are sustainable and nutritionally complete is an increasingly complex task that must take into account several needs to foster eating habits of health-conscious consumers, while reducing food waste and environmental impact. While Generative AI and Large Language Models (LLMs) show promising results in this domain due to their natural language processing capabilities, they suffer from critical limitations, including hallucinations, knowledge gaps, and limited ability to handle factual information. To mitigate such limitations, Retrieval-Augmented Generation (RAG), which retrieves relevant information from external sources to enhance the capabilities of LLMs, has shown effectiveness in many domains. However, existing RAG approaches typically operate on unstructured text that lacks sophisticated symbolic representations of complex domain knowledge. This work proposes an ontology-enhanced conversational food advisory system that integrates a modular ontology, named FoCOSA (Food Consumer-Oriented Sustainability-Aware), within several key tasks of a RAG-based system, enhancing LLM reasoning with domain knowledge, while simultaneously improving the interpretation of user requests, thus improving retrieval effectiveness and interaction fluidity. Experimental evaluations demonstrate the efficacy of the approach, and the study concludes with guidelines for selecting appropriate settings for food recommendation scenarios considering the complexity of natural language queries and other contextual factors.
Service composition is a key application area for LLM-based multi-agent systems, where dedicated agents, coordinated by LLMs, handle distinct phases of the service composition process. Assigning the same LLM to all agents is often inefficient and economically unsustainable, as different phases have varying requirements in terms of capabilities, deployment modalities, and cost. This paper proposes a decision-support framework based on Fuzzy Cognitive Maps (FCMs) to guide the allocation of LLM categories to agents in multi-agent service composition systems. The framework models causal relationships among contextual factors, phase selection, and LLM category assignment, providing interpretable recommendations to support decision-making. The target users are practitioners and architects configuring multi-agent systems who have a background in service-oriented principles, but may lack deep expertise in LLM evaluation or benchmarking, and therefore require structured guidance to navigate trade-offs before deployment. To validate the approach, a proof-of-concept web-based tool has been developed, enabling users to simulate different operational scenarios and observe how contextual conditions affect system configuration and LLM allocation.
In modern smart factories, supply chains are no longer isolated; instead, they are evolving into interconnected and dynamic networks, where intertwined supply chains enable real-time collaboration and data sharing for adaptive decision-making across multiple stakeholders. By harnessing data from sensors and connected devices, data-driven decisions can be made to optimize the entire supply chain, and to provide novel and customer-friendly products and services. Cyber-Physical Systems form the foundation of Cyber-Physical Production Systems (CPPS) by enabling real-time data exchange and intelligent automation at the factory level, while horizontal integration connects CPPS across different production facilities to enhance supply chain coordination, thus forming the so-called Cyber-Physical Production Networks (CPPN). In CPPN, the Internet of Services (IoS) paradigm, in combination with the Internet of Things (IoT), plays a crucial role in facilitating horizontal integration and seamless collaboration between intertwined supply chains. Since the IoS paradigm has to enable data sharing and processing within individual smart factories and across factory borders, there is a need to design service-oriented architectures specifically tailored to data governance in both CPPS and CPPN. However, existing service-oriented approaches for CPPS primarily focus on deployment layers (e.g., fog/edge computing or IT/production levels) while neglecting data-oriented aspects, limiting modularity and effective data service design across CPPS and CPPN levels. To bridge this gap, in this paper, we propose a multi-layered service-oriented model for CPPN focused on data services, which includes atomic services for data collection and processing, and composite services for governing the data flow within smart factories and throughout the supply chains they participate in. One of the significant advantages of the multi-layered approach is a clear separation of concerns in service design, with the ability to address issues of modularity, scalability, data sovereignty and data access, by distinguishing between CPPS and CPPN levels. In the paper, we critically evaluate different strategies for the management of a service ecosystem that is compliant with the proposed model.
The Internet of Services paradigm promotes using data services to accomplish diverse analytics tasks, enhancing collaboration amongst the actors of a production network. While domain experts and R&D managers are familiar with the observed system and can identify the data needed for analytics and digital innovation purposes, IT specialists typically handle data service discovery and their combination into analytics pipelines. Recently, Large Language Models (LLMs) have been recognised as valuable tools to support domain experts and R&D managers in specifying service needs and designing preliminary analytics pipelines, which can then be implemented by IT specialists, thus bridging the gap between domain and technical expertise. However, constructing effective prompts tailored to domain experts for interacting with LLM-based systems still demands advanced technical skills, as well as extensive knowledge of the catalog of available services and how they can be combined into analytics pipelines. To address this challenge, we propose an LLM-based approach for data service discovery within the Internet of Production context, which leverages Retrieval-Augmented Generation (RAG) applied to a catalog of atomic data services and relevant analytics pipelines. Preliminary experiments evaluating the effectiveness of the approach were conducted in a real-world case study within a Smart Factory research project.
Nowadays, workflows in judiciary systems are undergoing rapid transformation, stimulated by the technological opportunities offered by Generative AI solutions, including Large Language Models (LLMs). These technologies offer promising tools for addressing the inefficiencies and accessibility challenges inherent in traditional judicial workflows, which have long resisted digital modernization. By automating repetitive and time-intensive tasks such as text summarization and document analysis, LLMs can assist humans, thus enhancing operational effectiveness. This paper presents an exploratory study conducted in the scope of a collaboration between researchers and IT experts from the University of Brescia and the Prosecutor General’s Office at the Court of Appeal of Brescia. Adopting a Design Science Research methodology, the paper describes the design and evaluation of a Proof-of-Concept Web application that leverages LLMs and prompt engineering to support text summarization and analysis tasks. The prototype is intended to assist legal professionals with domain expertise, but without advanced IT skills, in interacting effectively with Generative AI technologies. The study highlights the potential of LLMs to streamline human effort, reduce manual overhead, and support decision-making, while also pointing out key challenges for their adoption in judicial workflows.
Effective Cyber Threat Intelligence (CTI) relies upon accurately structured and semantically enriched information extracted from cybersecurity system logs. However, current methodologies often struggle to identify and interpret malicious events reliably and transparently, particularly in cases involving unstructured or ambiguous log entries. In this work, we propose a novel methodology that combines ontology-driven structured outputs with Large Language Models (LLMs), to build an Artificial Intelligence (AI) agent that improves the accuracy and explainability of information extraction from cybersecurity logs. Central to our approach is the integration of domain ontologies and SHACL-based constraints to guide the language model's output structure and enforce semantic validity over the resulting graph. Extracted information is organized into an ontology-enriched graph database, enabling future semantic analysis and querying. The design of our methodology is motivated by the analytical requirements associated with honeypot log data, which typically comprises predominantly malicious activity. While our case study illustrates the relevance of this scenario, the experimental evaluation is conducted using publicly available datasets. Results demonstrate that our method achieves higher accuracy in information extraction compared to traditional prompt-only approaches, with a deliberate focus on extraction quality rather than processing speed.
In the context of the Internet of Services paradigm for Industry 4.0, data services can be discovered and composed to accomplish different data analytics scenarios amongst the actors of a production network. Recently, Large Language Models (LLMs) have been increasingly considered for service discovery and composition as a promising alternative to previous approaches that often require substantial effort to produce formal service descriptions and/or annotations. In this paper, we introduce some exploratory experiments on the use of an LLM-based system for the discovery of data services to fulfil data analysis scenarios. First, a data service model, that represents in a declarative way data service operations, is provided. Then, we propose prompt templates for the interaction with the LLM-based system, that leverages the service model, aimed at reducing trial-and-error interactions for identifying potential service candidates. The effectiveness of the approach is being assessed in a real-world case study of a research project.
Monitoring business processes within complex supply chains demands efficient data collection and analytics tailored to diverse phenomena. Traditional centralized solutions face limitations in adapting to the dynamic nature of supply chains. This calls for distributed solutions which break the usual architectural assumption to have a central entity in charge of collecting, integrating and offering tools for the analysis. This project, embedded in a larger initiative called MICS, proposes an inno-vative distributed monitoring solution integrating blockchain for a trustworthy and efficient data analytics strategy that preserves data sovereignty in complex collaborative environments. Leveraging the cloud -edge continuum, the solution aims to ensure secure data exchange, adherence to agreements, and real-time analytics. Expected outcomes include an innovative federated architecture, 5G slice management solutions, an adversarial analysis of supply chain security, and a proof-of-concept implementation of the blockchain-based data flow tracking system. These developments aim to enhance the reliability, security, and efficiency of supply chain monitoring in dynamic industrial environments.
Cinzia Cappiello合作论文数Polytechnic University of Milan,Department of Electronics, Information and Bioengineering7
Andrea Maurino合作论文数Politecnico di Milano;Dipartimento di Elettronica ed informazione3