With digital transformation, industrial companies today are facing the challenges to change and innovate their business, by leveraging digital technologies and tools to support their processes and their operations. One of their main challenges is the management of the company knowledge, especially when tacit and owned by industry workers. In this paper, we illustrate how knowledge graphs can be the turning point to allow industry workers digitize and exploit the knowledge about the “what”, the “how” and the “why” of their everyday activities.In particular, we focus on the “how” by illustrating the challenges related to procedural knowledge management, i.e., the knowledge about processes and workflows that employees need to follow, and comply with, to correctly execute their tasks, in order to improve efficiency and effectiveness, to reduce risks and human errors and to optimize operations. We also explain the relationship in this context between knowledge graphs and sub-symbolic AI approaches.
Processes, workflows and guidelines are core to ensure the correct functioning of industrial companies: for the successful operations of factory lines, machinery or services, often industry operators rely on their past experience and know-how. The effect is that this Procedural Knowledge (PK) remains tacit and, as such, difficult to exploit efficiently and effectively. This paper presents PKO, the Procedural Knowledge Ontology, which enables the explicit modeling of procedures and their executions, by reusing and extending existing ontologies. PKO is built on requirements collected from three heterogeneous industrial use cases and can be exploited by any AI and data-driven tools that rely on a shared and interoperable representation to support the governance of PK throughout its life cycle. We describe its structure and design methodology, and outline its relevance, quality, and impact by discussing applications leveraging PKO for PK elicitation and exploitation.
Procedural Knowledge is the know-how expressed in the form of sequences of steps needed to perform some tasks. Procedures are usually described by means of natural language texts, such as recipes or maintenance manuals, possibly spread across different documents and systems, and their interpretation and subsequent execution is often left to the reader. Representing such procedures in a Knowledge Graph (KG) can be the basis to build digital tools to support those users who need to apply or execute them. In this paper, we leverage Large Language Model (LLM) capabilities and propose a prompt engineering approach to extract steps, actions, objects, equipment and temporal information from a textual procedure, in order to populate a Procedural KG according to a pre-defined ontology. We evaluate the KG extraction results by means of a user study, in order to qualitatively and quantitatively assess the perceived quality and usefulness of the LLM-extracted procedural knowledge. We show that LLMs can produce outputs of acceptable quality and we assess the subjective perception of AI by human evaluators.
The ability to extract valuable information from documents and convert it into knowledge is crucial for driving technological innovation across industries. While adding metadata to manuals enhances their searchability, the real knowledge is still hidden in the procedural information they contain, which offers vital guidance for operators. Therefore, the approach of extracting and transforming unstructured human-readable information into machine-interpretable data is fundamental for establishing cutting-edge digital knowledge-based platforms. This paper presents a methodology tailored to the specific requirements of users who are seeking support in extracting and representing procedural knowledge from documents. We introduce a tool designed to support users in manually annotating procedures within PDF documents and generating a corresponding procedural knowledge graph. We assess the tool in real-world scenarios, aimed at evaluating its effectiveness in accomplishing various tasks. Finally, we generate a procedural knowledge graph that can facilitate knowledge discovery.
Digitalization is entering the industrial sector and different needs are emerging to support shop floor operators; in particular, they need to retrieve information to support their operations (e.g., during maintenance activities), from structured and unstructured sources, as well as from other people’s experience. Sharing knowledge and making it accessible to industrial workers is therefore a key challenge that Semantic Web technologies are able to address and solve. In this paper, we present a modular ontology that we engineered in order to support the collection, extraction and structuring of relevant information for industrial operators in a “knowledge hub” (K-Hub). In particular, our K-Hub ontology covers several aspects, from document annotation/retrieval to procedure support, from manufacturing domain concepts to company-specific information. We discuss its engineering process, extensibility and availability, as well as its current and future application scenarios to support industrial workers.
The fragmentation of data sources and data catalogues in the transportation domain is currently limiting the possibility of combining multi-source data for the definition of innovative data-driven mobility scenarios. In the European context, the implementation of National Access Points by Member States is promoting the publication of transportation data, but the adoption of not standardised, proprietary and not machine-readable metadata profiles currently prevents data sources to be easily found across National Access Points and to be interconnected this way. This paper describes the efforts towards the development of a napDCAT-AP specification for the harmonisation of metadata in transportation data catalogues. Starting from the analysis of the European recommendations for the interoperability of data catalogues, a roadmap for the design, implementation and publication of napDCAT-AP is described reporting guidelines and best practices. A consolidated list of requirements, collected from experts and transportation stakeholders, is discussed as the first step towards napDCAT-AP.
Born in the context of the financial market, “continuous reporting” is an activity that many municipalities must perform, to address novel issues such as ecology. In this paper, we study the case of “environmental reporting” by adopting the vision provided by the RADAR Framework. The paper shows how the innovative approach provided by the framework for continuous reporting can be beneficial for municipalities, to address the long-lasting activity of environmental monitoring, provided that data and knowledge about data are collected by the framework, which supports design and continued generation of reports.
The capability of extracting useful information from documents and further transferring it into knowledge is essential for advancing technology innovations in industries. Procedures described within the service manuals provide guidelines as unstructured human-readable documents. Although annotating manuals with metadata makes them searchable, the real knowledge is still hidden in the procedural information which provides essential guidance for the operators. Therefore, there is a need to develop data curation techniques in order to build such procedural knowledge. However, creating this knowledge automatically can be hard to explicitly articulate as it refers to abilities and skills that may be hard to explain and describe. Still, manuals and other documentation often can include the description of procedures in terms of steps of a process or predefined plans. In this paper, we provide an overview of the state-of-the-art approaches based on manual and automatic annotations with or without human-in-the-loop involvement. We will discuss the challenges and the opportunities based on representative state-of-the-art work related to the annotation of documents as well as their semantic representation that may support knowledge curation of procedures in the manuals.
Many organizations must produce many reports for various reasons. Although this activity could appear simple to carry out, this fact is not at all true: indeed, generating reports requires the collection of possibly large and heterogeneous data sets. Furthermore, different professional figures are involved in the process, possibly with different skills (database technicians, domain experts, employees): the lack of common knowledge and of a unifying framework significantly obstructs the effective and efficient definition and continuous generation of reports. This paper presents a novel framework named RADAR, which is the acronym for “Resilient Application for Dependable Aided Reporting”: the framework has been devised to be a ”bridge” between data and employees in charge of generating reports. Specifically, it builds a common knowledge base in which database administrators and domain experts describe their knowledge about the application domain and the gathered data; this knowledge can be browsed by employees to find out the relevant data to aggregate and insert into reports, while designing report layouts; the framework assists the overall process from data definition to report generation. The paper presents the application scenario and the vision by means of a running example, defines the data model and presents the architecture of the framework.
Highly-heterogeneous and fast-arriving large amounts of data, otherwise said Big Data, induced the development of novel Data Management technologies. In this paper, the members of the IFIP Working Group 2.6 share their expertise in some of these technologies, focusing on: recent advancements in data integration, metadata management, data quality, graph management, as well as data stream and fog computing are discussed.
The value of product lifecycle management systems (PLMS) is more and more recognised by companies and its use current has enormously increased. It is mainly used during the product design when different roles collaborate for sharing models, take review decisions, and approve or reject preliminary results. Often, companies have a general 'picture' about the processes involving PLMS (who performs an activity, when it is performed, what is done) but this knowledge can be reinforced, improved and modified using process mining. Here the knowledge is extracted from the event logs, and model-aware analytics are generated to evaluate the modelled, known and executed process. The business rules filter the logs and verify the impact on the process mining metrics to minimise the divergences between modelled and actual processes and improve the resulting quality metrics. The results help business users to identify lines of investigation for deviations from expected behaviour and propose improvement measures.
In the last decade, students facing a PhD course in Europe find terrible difficulties in reaching a permanent position in the academy. The situation gets worse when graduated PhDs have to migrate to public/private organisations that are not always ready to understand and improve the research experience. In such a situation, one of the most critical aspects is encountered immediately in the recruitment phase, since the keywords used in job offers portals are based on the employers' vocabulary and usually do not match the words that a researcher would use to describe her/his experience. Therefore, it is widely recognised that there is a need to define a system that can support a recruiters team in recruiting PhDs. The approach presented in this paper aims at designing a decision support tool able to guide the choices of recruiter of any company in the evaluation of profiles of candidates with PhD.
Abstract Background From the last decade, data mining techniques, employed in particular in customer relationship management, have assumed a key role in the profitability and operations of companies. To support small and medium companies (SMEs), several innovative and continuously improving tools have been developed that allow SMEs to utilize the internal and external data sources to increase their competitiveness. Objectives In this paper, an analysis of the impact of digitalization, and in particular data mining techniques, in the context of SMEs development is presented. Methods/Approach A review of various sources has been conducted, with the focus on open source tools, since in the context of the Italian economy they are used by SMEs the most. Results First, the analysis presents a brief review of the data mining techniques available and shows how they are practically employed in small companies. Second, an economical review of investments in data mining projects in Italy is presented. Conclusions The review indicates that data mining techniques can boost a company in the market. However, the awareness of data mining as a company asset is still not strong in Italian SMEs and most investments in Italy are still carried out by large companies.
This paper proposes a new tool in the field of telemedicine, defined as a specific branch where IT supports medicine, in case distance impairs the proper care to be delivered to a patient. All the information contained into medical texts, if properly extracted, may be suitable for searching, classification, or statistical analysis. For this reason, in order to reduce errors and improve quality control, a proper information extraction tool may be useful. In this direction, this work presents a Machine Learning Multi-Label approach for the classification of the information extracted from the pathology reports into relevant categories. The aim is to integrate automatic classifiers to improve the current workflow of medical experts, by defining a Multi- Label approach, able to consider all the features of a model, together with their relationships.
This work proposes the design of a decision support tool able to guide the choices of any company HR manager in the evaluation of the profiles of PhD candidates. This paper is part of an ongoing research in the field of PhD profiling. The novelty here is an evolutionary fuzzy model, based on the Membership Functions (MFs) optimization, used to obtain the soft skills candidate profiles. The general aim of the project is the definition of a set of fuzzy rules that are very similar to those that a HR expert would otherwise have to calculate each time for each selected profile and for each individual skill.
Large companies and organizations periodically feed their information systems with large data flows. Apart from theclassical operational activities, they are called to prepare aggregated reports to send to institutions and rating
Today, large organisations and regulated markets are subject to the control of external audit associations, which require the submission of a huge amount of information in the form of predefined and rigidly structured reports. The compilation of these reports requires the extraction, transformation and integration of data from different heterogeneous operational databases. This task is usually performed by developing a software ad hoc for each report, or by adopting a data warehouse and analysis tools, which are now established technologies. Unfortunately, the data warehousing process is notoriously long and error prone, and is therefore particularly inefficient when the output of the data warehousing is represented by a limited number of reports. This article presents “MMBR”, an approach that can generate a multidimensional model from the structure of expected reports as data warehouse output. The approach is able to generate the multidimensional model and populate the data warehouse by defining a knowledge base specific to the domain. Although the use of semantic information in data storage is not new, the novel contribution of our approach is represented by the idea of simplifying the design phase of the data warehouse, making it more efficient, by using an industry-specific knowledge base and a report-based approach.
The increasing volume of data created and exchanged in distributed architectures has made databases a critical asset to ensure availability and reliability of business operations. For this reason, a new family of databases, called NoSQL, has been proposed. To better understand the impact this evolution can have on organizations it is useful to focus on the notion of Online Analytical Processing (OLAP). This approach identifies techniques to interactively analyze multidimensional data from multiple perspectives and is today essential for supporting Business Intelligence. The objective of this paper is to benchmark OLAP queries on relational and graph databases containing the same sample of data. In particular, the relational model has been implemented by using MySQL while the graph model has been realized thanks to the Neo4j graph database. Our results, confirm previous experiments that registered better performances for graph databases when re-aggregation of data is required.
Irene Celino合作论文数CEFRIEL - Politecnico di Milano8
Andrea Maurino合作论文数Politecnico di Milano;Dipartimento di Elettronica ed informazione2
Stefano Cagnoni合作论文数Department of Engineering and Architecture, University of Parma2