
This paper focuses on Digital Libraries that geovisualize cultural data, highlighting the need to define them as a separate category termed “Cultural Mapping Libraries”, based on the culture-location connection to represent cultural data on a map. An exploratory analysis of Digital Libraries conforming to that definition brought forward the observation that existing Digital Libraries fail to geovisualize the entirety of cultural data per point of interest, resulting in a phenomenon we termed Cultural Snapshot. This phenomenon was confirmed by the results of a systematic bibliographic research. To address it, this paper proposes the use of the Semantic Web principles to efficiently interconnect spatial cultural data through time, per geographic location. This way, points of interest transform into scenery where culture evolves over time. This evolution is expressed as events over time, in an event-oriented manner, endorsed by the CIDOC Conceptual Reference Model (CRM). We posit the use of CIDOC CRM as the baseline for defining the logic of Cultural Mapping Libraries as part of the Culture Domain in accordance with the Digital Library Reference Model, in order to define the rules of cultural data management by the system. Our future goal is to transform this conceptual definition in to inferencing rules that resolve the Cultural Snapshot, supporting a more complete geovisualization of cultural data.
Usually, the focus of metadata annotation is on the research output rather than the context in which products were generated. The objective of this project was to develop a framework for contextual metadata, involving six research infrastructures (RIs) from two different domains. As a first step semi-structured interviews were performed to assess the current handling of contextual metadata. Then, these results were put into perspective with the main entities of research processes in general, leading to a framework for contextual metadata. From the discussion with the RIs and in alignment with the referenced literature, basic entities related to contextual metadata are defined and organised in a framework. In summary, a considerable amount of contextual metadata information is already covered by the RIs, however, not always explicit but implicit within text descriptions. The RIs involved see contextual metadata as necessary to improve replicability and reliability of research and FAIRness of data.
The FAIR (Findable, Accessible, Interoperable, Reusable) principles and their adoption as best practices in data management have increased the publication of research data and associated metadata. The relevance of clinical care data for further in-depth research on the virus and its consequences has been brought to light by the COVID-19 pandemic and the global efforts to combat it. This paper presents a practical FAIRification workflow for the transformation and publication of clinical research data into FAIR (meta)data. It is an extended version of a practical FAIRification workflow approach validated in the Virus Outbreak Data Network Brazil (VODAN BR) pilot. Furthermore, it discusses how the FAIRification process contributes to reproducible clinical research data and presents lessons learned with the development of this process.
Data lake metadata management is crucial for clearly describing stored data and ensuring efficient search query results, especially for semi-structured and unstructured data. Moreover, high-quality metadata provides the necessary information for decision-making. Therefore, we propose a new approach called Metadata Management for Data Lakes (MMDL) based on metadata quality and classification. Additionally, we introduce a generic model adapted to the context of a data lake by including data lake zones to identify data location during its lifecycle. The model aligns with the proposed approach and enables the handling of metrics used to evaluate metadata quality and metadata sources for provenance classification.
Paper retractions are rising due to the absence of reliable tools to detect and identify fraudulent articles before publication. This paper proposes an ontology-based decision support system to prevent compromised peer review by carefully selecting qualified reviewers and avoiding potential conflicts of interest. There are three main contributions: (i) formulating the criteria for the selection of qualified reviewers as well as for recognising potential conflicts of interest; (ii) designing the ontology-based decision support system; (iii) designing and performing the methodology for validation. A pilot test with 30 computer science experts is conducted to determine qualified reviewers and conflict criteria, including the design of ontologies structure. Subsequently, three selected computer science experts and journal editors are requested to evaluate a set of test data as the ground truth. Overall results show that the proposed solution achieves 91% accuracy in qualified reviewer selection and 94% accuracy in conflict-of-interest detection.
Meteorological data, essential in a variety of applications, has been made available as open data through different portals, either governmental, associative or private ones. Making this data fully findable and reusable for experts from other domains than meteorology requires considerable efforts to guarantee compliance to the FAIR principles. Nowadays, most efforts in data FAIRification are limited to semantic metadata describing the overall features of data sets. However, such a description is not enough to fully address data interoperability and reusability by other scientific communities. This paper addresses this weakness by proposing a semantic model to represent different kinds of metadata, describing the data schema and the internal structure of a data set distribution, together with domain-specific definitions. This model is used to provide a reusable schema of the SYNOP data set, a largely used governmental meteorological data set in France. The impact of using the proposed model for improving FAIRness was evaluated.
Companies have migrated their operational activities from paper documents to automated processes with fully digital storage. This management trend is positive, but printed documents, in most cases, cannot be discarded for administrative or legal reasons. This research used data extraction to enrich the database of a Non-Governmental Organisation (NGO) that monitors the use of public financial resources in counties. The implementation analysed the digital files containing official documents and identified the words with the highest occurrence according to algorithms presented in the research results. The solution created in the research added metadata to improve the search for documents in the database and improve the procedural follow-up of administrative and judicial actions. The results were positive with success in the extraction of the keywords in each document and presented with examples in the results section, showing the steps used to add metadata in the documents.
This study evaluates the performance of four image recognition tools (Amazon Rekognition, Clarifai, Imagga and Google Cloud Vision API) for automatic image metadata. The experiment was conducted on various image categories, including human, animal, plant and flower, view and landscape, vegetable and fruit, food, vehicle, tourist landmark, art and culture and old book cover and posters. Semantic and label-based analysis was used to evaluate the performance of each tool. Results indicate that each tool performed differently across categories, demonstrating the importance of selecting the appropriate tool for specific tasks. Clarifai was found to perform best for human, animal and food image tagging, while Amazon Rekognition was best for vegetable-fruit and vehicle images. Imagga performed best for plant and flower, art and culture and old book cover and posters image recognition, while Google Cloud Vision API performed best for view and landscape and tourist landmark recognition.
The goal of the current work is to improve and assess the existing metadata for clinical narrative information. The metadata schema for clinical narrative information borrows from the domain of narrative information and clinical narration. The elements corresponding to the narrative information were from the elements of narration, narrative theories and entities in the ontology-based narrative models. The elements of clinical narration are identified from the patient stories. This study found that the narrative in general and medical narratives in specific lack a metadata schema. The work developed a metadata framework for medical narrative information. Following the development of this framework, an evaluation by experts was conducted through the Delphi method.
General-purpose Knowledge Bases (KBs) have been used for various applications. An essential step for leveraging the content of KBs on domain-specific tasks is to discover their schema. In this paper, we propose ANCHOR, an end-to-end pipeline for schema discovery from general-purpose KB in an automated way. ANCHOR identifies a domain of interest based on category mapping from KB. Next, it learns representations of entities in this domain based on the entity-category mappings and uses these representations to identify the entities' topics within this domain. Finally, ANCHOR generates a profile for each topic using a strategy based on attributes co-occurrence. We have evaluated ANCHOR on four domains. The results show that: (1) the learned entity representation effectively produces better entity clusters than some traditional and embedding-based baselines; (2) our solution produces a high-quality profile for the discovered topics.
Cross-border data transfers and their legal aspects have created a daunting landscape for application and service providers, in which rules and regulations need to be constantly monitored and addressed, especially in dynamic scenarios such as cloud brokerage or cloud/edge operations. The aim of this work is to semantically model several concepts surrounding international data transfers based on the current changes and formulate them around a newly defined ontology (CIDaTa). The work exploits 23 existing ontologies, as dictated by the Linked Data paradigm, and introduces 54 links between them. This aids the IT professional in answering questions regarding the legality of a transfer or the necessary steps needed to achieve it. Example questions set to the framework are demonstrated that can enhance the understanding of the implications of a data transfer, enabling future additions that can lead to more automated management of these transfers.
This article proposes OntoAthena, an ontology to represent knowledge in the domain of intelligent services for e-learning. The purpose of the proposed ontology is to support educational practices and provide intelligent services, helping managers, teachers and students in the teaching and learning process. OntoAthena was specified through Ontology Development 101 and implemented in Protégé 5.5.0. The ontology has 22 classes, 70 sub-classes, 41 object properties and 30 data properties. Four competency questions were used for assessment. The results obtained demonstrated the correct classification of the instances in the OntoAthena. As scientific contribution, this study presents the first ontology to support intelligent services for e-learning.
Folklore culture is the summary and embodiment of a way of life, a normative form for people, covering local life in rural communities. A folklore museum typically displays historical objects used as part of people's everyday lives. The representation of information in the digital realm is an important aspect of cultural heritage institutions, with museums taking over a considerable proportion. Having taken into consideration requirements set by the existing literature and examining relevant research and approaches, the authors developed an analysis framework including 16 popular folklore museums around the world, in order to examine various aspects of their digital presence. More specifically, the authors followed a two-fold methodological approach conducting a functionality and semantic analysis of the chosen museums with the purpose of determining the current approaches and tendencies with respect to representation schemata for folklore museum artifacts and their digital representation.
The partitioning of RDF data on a large scale allows generating a set of RDF data subgraphs. METIS is a graph partitioning technique that minimises the cost of partitioning. METIS applies, among other things, to RDF graphs. However, the semantics introduced in the description of RDF data is not taken into account in the partitioning process in METIS. For this, we propose in this paper a step of pre-processing RDF data before partitioning these data. The objective of this step is to improve the quality of semantic partitioning of RDF graphs. The evaluation of the RDF pre-processing step for METIS was performed on real and synthetic data.
Providing research data to be readable, accurate and understandable by human and autonomous computational agents is challenging, primarily if published on the web. We present Nano-PROV, a workflow-based approach that aims at semantic enrichment of data and provenance control of published research data sets. The workflow uses the nanopublications for data transformation, a reliable format for dynamically publishing research outputs. Further, Nano-PROV adopts UN-PROV, a unified provenance guideline centred on nanopublication for identifying and controlling data and workflow provenance. In this paper, we developed computational experiments to evaluate the workflow by generating a nano-pub data model based on the genomic scenario, showing how the proposal may circumvent various issues regarded with data reusability, interoperability, and discoverability issues. Compared with related works, our results demonstrated the feasibility of Nano-PROV to enhance the semantic expressivity of research data and its metadata annotations.
Currently, most of the repositories implement Dublin Core (DC) as metadata standard, allowing the application of Protocol for Metadata Harvesting (OAI-PMH: Open Archives Initiative - Protocol for Metadata Harvesting). However, DC is not the most appropriate standard for the description of learning objects, which makes necessary resort to other standards. The Learning Object Metadata (LOM) standard emerges as the most suitable for the description of learning objects. Also, other standards such as Common European Research Information Format (CERIF) Metadata Object Description Schema (MODS) arise among others. This variety of standards makes the interoperability between repositories will become increasingly complex. Most solutions, until now, propose to adopt a metadata standard and include the necessary metadata to be harvested. This paper presents a solution based on ontologies for interoperability between repositories that use different metadata in the description of its objects.