
Knowledge reuse can increase the quality and efficiency of different activities, as well as reduce risks. The digital transformation of a city is a complex task that requires new approaches and management technologies. Knowledge reuse can reinforce city digital transformation initiatives, e.g. data integration and modeling activities can be improved via ontology reuse, design and development of digital solutions can be enhanced via reference architectures and reuse of existing methods and systems. Such reuse can speed up the work of city development managers, digital architects, and other ICT specialists, as well as decrease costs and risks. Although there are platforms for knowledge sharing and reuse within the smart city domain, they apply a traditional document-oriented approach to knowledge representation. This limits knowledge integration and the creation of intelligent services. So, knowledge graph technology was selected for collecting and curating digital city planning reusable content. The Open Research Knowledge Graph (ORKG) provided an opportunity for piloting this approach. The current paper will overview the digital city planning observatory within the ORKG, including main knowledge representation mechanisms and representative content examples, and provide the link between content and application scenarios.
Knowledge Extraction (KE) techniques are used to convert unstructured information present in texts to Knowledge Graphs (KGs) which can be queried and explored. Despite their potential for cultural heritage domains, such as Art History, these techniques often encounter limitations if applied to domain-specific data. In this paper we present the main challenges that KE has to face on art-historical texts, by using as case study Giorgio Vasari’s The Lives of The Artists . This paper discusses the following NLP tasks for art-historical texts, namely entity recognition and linking, coreference resolution, time extraction, motif extraction and artwork extraction. Several strategies to annotate art-historical data for these tasks and evaluate NLP models are also proposed.
Semantic Storytelling is the (semi)-automatic generation of new content (storylines) based on information extracted from document collections presented with helpful visualisation techniques. This paper summarises our previous work and describes the Semantic Storytelling vision and technical approach. We describe experiments that focus on the identification of relations between text segments extracted from documents written by different authors; discourse relations have, so far, primarily been researched within single documents only. The results confirm our intuition that discourse parsing is difficult to apply to the inter-text level, but they are encouraging as a first step. Similarly, the identification of inter-text relations using pairwise document classification yields promising results. Lastly, we show the effectiveness of paragraph ordering for coherent story generation.
This paper analyses the correlation between online news articles’ credibility and compliance with journalistic standards and ethics. Our main hypothesis is that news sources with low credibility adhere less to journalistic standards and ethics than high credibility sources. Additionally, the influence of web page layout factors is analysed. Our definition of metrics is based on the existing literature regarding journalistic best practices, as well as the list of credibility signals provided by the W3C Credible Web Community Group. News article credibility is assessed based on a score, which is calculated as a sum of linear subscores for the identified metrics divided in four credibility signal categories: Formality, Neutrality, Transparency and Layout. A curated dataset of 250 recent news items from known fake news sources, as well as 200 news items from established real news sources, is used in the testing phase. Although the comparison between real news and false news shows that, on average, the credibility score of genuine news is only 5% higher than that of fake news, the results do show more significant differences for certain subcategories and signals.
Digital Maktaba (DM) is an interdisciplinary project to create a digital library of texts in non-Latin alphabets (Arabic, Persian, Azerbaijani). The dataset is made available by the digital library heritage of the ”La Pira” library in the history and doctrines of Islam based in Palermo, which is the hub of the Foundation for Religious Sciences (FSCIRE, Bologna). Establishing protocols for the creation, maintenance and cataloguing of historical content in non-Latin alphabets is the long-term goal of DM. The first step of this project was to create an innovative workflow for automatic extraction of information and metadata from title pages of Arabic script texts. The Optical Character Recognition (OCR) tool uses various recognition systems, text processing techniques and corpora in order to provide accurate extraction and metadata of document content. In this paper we address the ongoing development of this novel tool and, for the first time, we present a demo of the current version that we have designed for the extraction and cataloguing process by showing a use case on an Arabic book frontispiece. In particular, we delve into the details of the tool workflow for automatically converting and uploading PDFs from the digital library, for the automatic extraction of cataloguing metadata and the semiautomatic (at the current stage) process of cataloguing. We also shortly discuss future prospects and the many additional features that we are planning to develop.
In this work we introduce Coreon Multilingual Knowledge System, a visual tool for concept-based data modeling and curation. Coreon environment is built on the visual paradigm; its goal is to offer a comprehensible and user-friendly solution for domain experts, well-versed in concept modelling dialects, as well as ad-hoc users, who are new to curation of semantic content. We describe the mechanism of the tool, which, being machine-readable, is powered by the language-agnostic knowledge graph, capably embedding the non-deterministic phenomena of the human language.
The capability of extracting useful information from documents and further transferring it into knowledge is essential for advancing technology innovations in industries. Procedures described within the service manuals provide guidelines as unstructured human-readable documents. Although annotating manuals with metadata makes them searchable, the real knowledge is still hidden in the procedural information which provides essential guidance for the operators. Therefore, there is a need to develop data curation techniques in order to build such procedural knowledge. However, creating this knowledge automatically can be hard to explicitly articulate as it refers to abilities and skills that may be hard to explain and describe. Still, manuals and other documentation often can include the description of procedures in terms of steps of a process or predefined plans. In this paper, we provide an overview of the state-of-the-art approaches based on manual and automatic annotations with or without human-in-the-loop involvement. We will discuss the challenges and the opportunities based on representative state-of-the-art work related to the annotation of documents as well as their semantic representation that may support knowledge curation of procedures in the manuals.
Artificial Intelligence and Machine Learning offer enormous potential for applications in the digitization and digital curation of cultural heritage. But cultural heritage institutions have also produced large amounts of digital data that can be suitable to improve AI methods and models. At the same time there are problems and issues with data used in AI industry and research, which frequently lack quality curation and introduce or reinforce biases. What are the main obstacles for reuse of digitized cultural heritage as data for AI, and what can libraries with their quality awareness and long established practices and competencies of curation contribute to the field of AI?
. Most question answering tasks are oriented towards open domain factoid questions. In comparison, much less work has studied both factoid and open ended questions in closed domains. We have chosen a current state-of-art BERT model for our question answering exper-iment, and investigate the effectiveness of the BERT model for both factoid and open-ended questions in the museum domain, in a realistic setting. We conducted a web based experiment where we collected 285 questions relating to museum pictures. We manually determined the answers from the description texts of the pictures and classified them into answerable/un-answerable and factoid/open-ended. We passed the questions through a BERT model and evaluated their performance with our created dataset. Matching our expectations, BERT performed better for factoid questions, while it was only able to answer 36% of the open-ended questions. Further analysis showed that questions that can be answered from a single sentence or two are easier for the BERT model. We have also found that the individual picture and description text have some implications for the performance of the BERT model. Finally, we pro-pose how to overcome the current limitations of out of the box question answering solutions in realistic settings and point out important factors for designing the context for getting a better question answering model using BERT.
We present the conceptual design of a language technology (LT) system that enables enhanced document curation and processing of different documents types by providing customized NLP workflows that respond and adapt to the extracted characteristics of the input documents. To optimize document and text understanding, the processing steps will not only incorporate textual features but also layout and document type related features like document structure, and the communicative function of specific parts or constituents of a document (e. g., header, subtitle, paragraph, footer). We tackle the lack of standardized representation formats for many of these document features by presenting the first draft of an ontology (QOntology) we plan to incorporate into the overall workflow manager. Since the work is still in progress, we present the theoretical background and conceptual design decisions of the approach which will be the basis of experiments in future work.
The field of quantum computing is developing rapidly. As a result, a variety of quantum hardware, software development kits, and quantum algorithms have been developed in recent years. However, knowledge about these artifacts is either not available or spread among different sources. Thus, to analyze, compare, and evaluate knowledge on quantum computing an integrated knowledge base is required. In this paper, we introduce key concepts of an ontology for quantum algorithms and their implementations. The presented ontology serves as basis for a collaborative platform for researchers and practitioners to support collection and development of knowledge on the field of quantum computing.
This paper presents a corpus annotated for the task of directspeech extraction in Croatian.The paper focuses on the annotation of the quotation, co-reference resolution, and sentiment annotation in SETimes news corpus in Croatian and on the analysis of its language-specific differences compared to English. From this, a list of the phenomena that require special attention when performing these annotations is derived. The generated corpus with quotation features annotations can be used for multiple tasks in the field of Natural Language Processing.
. Attorneys and legal practitioners spend inordinate amounts of time reading case-law documents, trying to find relevant, precedential or exemplary decisions that support particular patterns of claims made in adjudicatory matters on behalf of their clients. To ameliorate this ubiquitous problem, we have crafted a Legal and Regulatory domain-specific ontology that works in tandem with our enterprise upper ontology. Up until recently, attorneys have relied and trusted books (including digital books) over more modern ways of consuming information. Since the start of the COVID-19 pandemic, there has been an erosion of the belief in the need for information to be delivered in book form [2]. As legal professionals are used to retrieving granular information in their personal search engine of choice (most start with Google), they come to expect similar capabilities from their legal search engine, which requires extracting more domain-specific insights from jurisprudence when using research products in their daily activities. Specifically, the metadata that is traditionally represented in our content is descriptive of, and specifically focused on, the document as the canonical subject. Rather than focusing on the document, we instead employ an abstract, information-centric representation of the specific semantic units that are realized in the text of jurisprudence documents (such as claims made by the litigants, facts of the case, etc.). We will describe some of the challenges we have faced, and lessons learned, in moving from a more “traditional” document-based mantra of enrichment to more domain-specific semantics.
An important part in European cultural identity relies on European cities and in particular on their histories and cultural heritage. Nuremberg, the home of important artists such as Albrecht Dürer and Hans Sachs developed into the epitome of German and European culture already during the Middle Ages. Throughout history, the city experienced a number of transformations, especially with its almost complete destruction during World War 2. This position paper presents TRANSRAZ, a project with the goal to recreate Nuremberg by means of an interactive 3D tool to explore the city’s architecture and culture ranging from the 17th to the 21st century. The goal of this position paper is to discuss the ongoing work of connecting heterogeneous historical data from various sources previously hidden in archives to the 3D model using knowledge graphs for a scientifically accurate interactive exploration on the Web.
. Book digitization is being increasingly enhanced, as it facilitates not only the dissemination and preservation of cultural heritage but also the analysis of large amounts of textual data as well as the extraction and discovery of knowledge in a faster, dynamic and interactive way. Quite often, OCR, as the core technology of book digitization, has to address major dif-ficulties related to the condition of the primary source or to scanning issues. The main contribution of this paper is to provide an extensive study on Tesseract, an open-source OCR system, including image pre-processing and text post-processing methods, that overcome a variety of image handling problems. Additionally, a re-trained Greek language model, based on individual fonts training plus pairs of image-text training, is being provided. Finally, this paper proposes a pipeline of methods, including text line detection, that result in enhanced accuracy for Greek Literature documents, even when they consist of distorted pages, due to scanning issues or damaged physical material.
. Between 1973 and 1985, a civic-military dictatorship ruled in Uruguay. Systematic violations of human rights marked this period. Project Cruzar.uy aims to develop tools and methodologies to analyze historical documents from that period. We present the advances in this ongoing project. We describe a set of tools to automatize the extraction and organization of information from the archives using computational tools including image processing, machine learning, natural language processing, information extraction and integration.
. In order to help improve the quality, coverage and performance of automated translation solutions for current and future Connecting Europe Facility (CEF) digital services, the European Language Resource Coordination (ELRC) was set up in 2015 through a service contract operating under the European Commission’s CEF SMART 2014/1074 programme . Since then, ELRC initiated a number of actions to support the collection of Language Resources (LRs) within the public sector in EU member and CEF-affiliated countries. All resources shared by the contributors were gathered and curated in the ELRC-SHARE Re-pository, after having passed the validation process developed by ELRC. This paper provides insights into the overall data collection and curation process (including both technical and legal validation of resources) employed within ELRC. The ELRC Helpdesk provides both technical and legal guidance (e.g. Intellectual Property Rights (IPR) clearance support) to potential data contributors, thus enabling the sustainable sharing of language data.
. This paper addresses one of the largest and most complex data curation workflows in existence: Wikipedia and Wikidata, with a high number of users and curators adding factual information from exter-nal sources via a non-systematic Wiki workflow to Wikipedia’s infoboxes and Wikidata items. We present high-level analyses of the current state, the challenges and limitations in this workflow and supplement it with a quantitative and semantic analysis of the resulting data spaces by deploying DBpedia’s integration and extraction capabilities. Based on an analysis of millions of references from Wikipedia infoboxes in different languages, we can find the most important sources which can be used to enrich other knowledge bases with information of better quality. An initial tool is presented, the GlobalFactSync browser, as a prototype to discuss further measures to develop a more systematic approach for data curation in the WikiVerse.
The global coronavirus pandemic has brought another ongoing crisis into the spotlight: that of digital misinformation. While society at a global scale is facing challenges that demand scientific solutions as never before, trust in experts and scientific expertise is falling, and conspiracy theories abound. At the same time, science itself is not without challenges, such as the reproducibility crisis across multiple domains. A contributing factor to misinformation is the way that scientific research is undertaken and reported in isolated and conflicting units, rather than as a holistic aggregate of information. In this position paper, I will argue that scientific ontologies and digital curation will be essential tools for transforming how scientific research is conducted and reported to address the problem of misinformation.
. We present the concept of extending a multilingual verb lexicon also to include German. In this lexicon, verbs are grouped by meaning and by semantic properties (following frame semantics) to form multilingual classes, linking Czech and English verbs. Entries are further linked to external lexical resources like VerbNet and PropBank. In this paper, we present our plan also to include German verbs, by experimenting with word alignments to obtain candidates linked to existing English entries, and identify possible approaches to obtain semantic role information. In addition, we identify German-specific lexical resources to link to. This small-scale pilot study aims to provide a blueprint for extending a lexical resource with a new language.