Kadaster, the Dutch National Land Registry and Mapping Agency, has been actively publishing their base registries as linked (open) spatial data for several years. To date, a number of these base registers as well as a number of external datasets have been successfully published as linked data and are publicly available. Increasing demand for linked data products and the availability of new linked data technologies have highlighted the need for a new, innovative approach to linked data publication within the organisation in the interest of reducing the time and costs associated with said publication. The new approach to linked data publication is novel in both its approach to dataset modelling, transformation, and publication architecture. In modelling whole datasets, a clear distinction is made between the Information Model and the Knowledge Model to capture both the organisation-specific requirements and to support external, community standards in the publication process. The publication architecture consists of several steps where instance data are loaded from their source as GML and transformed using an Enhancer and published in the triple store. Both the modelling and publication architecture form part of Kadaster’s larger vision for the development of the Kadaster Knowledge Graph through the integration of the various linked datasets.
In an era of ever-increasing scientific publications available, scientists struggle to keep pace with the literature, interpret research results and identify new research hypotheses to falsify. This is particularly in fields such as the social sciences, where automated support for scientific discovery is still widely unavailable and unimplemented. In this work, we introduce an automated system that supports social scientists in identifying new research hypotheses. With the idea that knowledge graphs help modeling domain-specific information, and that machine learning can be used to identify the most relevant facts therein, we frame the problem of hypothesis discovery as a link prediction task, where the ComplEx model is used to predict new relationships between entities of a knowledge graph representing scientific papers and their experimental details. The final output consists in fully formulated hypotheses including the newly discovered triples (hypothesis statement), along with supporting statements from the knowledge graph (hypothesis evidence and hypothesis history). A quantitative and qualitative evaluation is carried using experts in the field. Encouraging results show that a simple combination of machine learning and knowledge graph methods can serve as a basis for automated scientific discovery.
In Chapter 24, a co-author listed on the Consent to Publish form was inadvertently forgotten. This mistake has been corrected and the forgotten co-author has been added.
Graph-based traversal is an important navigation paradigm for the Semantic Web, where datasets are interlinked to provide context. While following links may result in the discovery of complementary data sources and enriched query results, it is widely recognized that traversing the LOD Cloud indiscriminately results in low quality answers. Over the years, approaches have been published that help to determine whether links are trustworthy or not, based on certain criteria. While such approaches are often useful for specific datasets and/or in specific applications, they are not yet widely used in practice or at the scale of the entire LOD Cloud. This paper introduces a new resource called MetaLink. MetaLink is a dataset that contains metadata for a very large set of owl : sameAs links that are crawled from the LOD Cloud. MetaLink encodes a previously published error metric for each of these links. MetaLink is published in combination with LOD-a-lot, a dataset that is based on a large crawl of a subset of the LOD Cloud. By combining MetaLink and LOD-a-lot, applications are able to make informed decisions about whether or not to follow specific links on the LOD Cloud. This paper describes our approach for creating the MetaLink dataset. It describes the vocabulary that it uses and provides an overview of multiple real-world use cases in which the MetaLink dataset can solve non-trivial research and application challenges that were not addressed before.
In the absence of a central naming authority on the Semantic Web, it is common for different data sets to refer to the same thing by different names. Whenever multiple names are used to denote the same thing, owl:sameAs statements are needed in order to link the data and foster reuse. Studies that date back as far as 2009, observed that the owl:sameAs property is sometimes used incorrectly. In our previous work, we presented an identity graph containing over 500 million explicit and 35 billion implied owl:sameAs statements, and presented a scalable approach for automatically calculating an error degree for each identity statement. In this paper, we generate subgraphs of the overall identity graph that correspond to certain error degrees. We show that even though the Semantic Web contains many erroneous owl:sameAs statements, it is still possible to use Semantic Web data while at the same time minimising the adverse effects of misusing owl:sameAs.
Linked Data is an innovative approach for publishing heterogeneous data sources on the web. As such, it can transcend the traditional confines of separate databases, as well as the confines of separate institutions. At the same time, businesses and governmental organizations alike are trying to cope with ever-increasing quantities of heterogeneous data that must be used across multiple departments, manufacturing locations, and governmental bodies. Linked Data would, therefore, be a great technological solution for today's organizational problems. However, we observe that a serious gap exists between Linked Data research and business research. While Linked Data research is almost exclusively technologically oriented, the business research literature has not devoted much attention to the use of Linked Data solutions yet. In this paper, we seek to bridge this gap, by introducing a real-world use case where Linked Data technologies are applied in large-scale government settings. We argue in detail that Linked Data provides a major contribution to the business vision of a modern governmental institution based on the experience of the Netherlands' Cadastre Land Registry and Mapping Agency (Kadaster), so far, the largest implementation of Linked Data in the governments of the Netherlands.
The field of geographic information science has grown exponentially over the last few decades and, particularly within the context of the pervasiveness of the internet, bears witness to a rapid transition of its associated technologies from stand-alone systems to increasingly networked and distributed systems as geospatial information becomes increasingly available online. With its long-standing history for innovation, the field has adopted many disruptive technologies from the fields of computer and information sciences through this transition towards web geographic information systems (GIS); most interestingly in the context of this research is the limited uptake of semantic web technologies by the field and its associated technologies, the lack of which has resulted in a technological disjoint between these fields. As the field seeks to make geospatial information more accessible to more users and in more contexts through 'self-service' applications, the use of these technologies is imperative to support the interoperability between distributed data sources. This paper aims to provide insight into what linked data tooling already exists, and based on the features of these, what may be possible for the achievement of self-service GIS. Findings include what visualisation, interactivity, analytics and usability features could be included in the realisation of self-service GIS, pointing to the opportunities that exist in bringing GIS technologies closer to the user.
Search tasks provide a medium for the evaluation of system performance and the underlying analytical aspects of IR systems. Researchers have recently developed new interfaces or mechanisms to support vague information needs and struggling search. However, little attention has been paid to the generation of a unified task set for evaluation and comparison of search engine improvements for struggling search. Generation of such tasks is inherently difficult, as each task is supposed to trigger struggling and exploring user behavior rather than simple search behavior. Moreover, the everchanging landscape of information needs would render old task sets less ideal if not unusable for system evaluation. In this paper, we propose a task generation method and develop a crowd-powered platform called TaskGenie to generate struggling search tasks online. Our experiments and analysis show that the generated tasks are qualified to emulate struggling search behaviors consisting of ‘repeated similar queries’ and ‘quick-back clicks’, etc. – tasks of diverse topics, high quality and difficulty can be created using this framework. For the benefit of the community, we publicly released the platform, a task set containing 80 topically diverse struggling search tasks generated and examined in this work, and the corresponding anonymized user behavior logs.
After more than a decade, the supply-driven approach to publishing public (open) data has resulted in an ever-growing number of data silos. Hundreds of thousands of datasets have been catalogued and can be accessed at data portals at different administrative levels. However, usually, users do not think in terms of datasets when they search for information. Instead, they are interested in information that is most likely scattered across several datasets. In the world of proprietary in-company data, organizations invest heavily in connecting data in knowledge graphs and/or store data in data lakes with the intention of having an integrated view of the data for analysis. With the rise of machine learning, it is a common belief that governments can improve their services, for example, by allowing citizens to get answers related to government information from virtual assistants like Alexa or Siri. To provide high-quality answers, these systems need to be fed with knowledge graphs. In this paper, we share our experience of constructing and using the first open government knowledge graph in the Netherlands. Based on the developed demonstrators, we elaborate on the value of having such a graph and demonstrate its use in the context of improved data browsing, multicriteria analysis for urban planning, and the development of location-aware chat bots.
In order to conduct large-scale semantic analyses, it is necessary to calculate the deductive closure of very large hierarchical structures. Unfortunately, contemporary reasoners cannot be applied at this scale, unless they rely on expensive hardware such as a multi-node in-memory cluster. In order to handle large-scale semantic analyses on commodity hardware such as regular laptops we introduced [1] a novel data structure called Equivalence Set Graph (ESG). An ESG allows to specify compact views of large RDF graphs thus easing the accomplishment of statistical observations like the number of concepts defined in a graph, the shape of ontological hierarchies etc. ESGs are built by a procedure presented in [1] that delivers graphs as a set of maps storing nodes and edges. In this demo paper (i) we show how facts entailed by an ESG and the graph itself can be specified in RDF following a novel introduced ontology; and, (ii) we present two datasets resulting from the triplification of two ESG graphs (one for classes and one for properties).
MetaLink is a dataset that contains metadata for a very large set of owl:sameAs links that are crawled from the LOD Cloud. MetaLink encodes a previously published error metric for each of these links [Raad et al., 2018]. This error degree ranges from 0.0 (most likely correct) till 1.0 (most likely incorrect). The idea is that the more an owl:sameAs link is isolated in the network (of all owl:sameAs links), the higher error degree this link will have. Experiments shows that discarding the 1M owl:sameAs links with an error degree >0.99 can significantly increase the quality of the transitive closure. Also by keeping only the 400M owl:sameAs links with error degree <= 0.4, the resulting closure is 100% precise in several manually evaluated cases. The resulted equivalence classes from these different closures are publicly available online. MetaLink is published in combination with LOD-a-lot, a dataset that is based on a very large crawl of a subset of the LOD Cloud. By combining MetaLink and LOD-a-lot, applications are able to make informed decisions about whether or not to follow specific links on the LOD Cloud. This dataset contains 4,352,602,452 unique triples, and is available in HDT (Header Dictionary Triples) format. It can be navigated online using the TriplyDB Linked Data hosting platform: https://krr.triply.cc/krr/metalink. A figure describing the vocabulary of the MetaLink dataset can be found here. Classes are displayed by circles and properties are displayed by arcs. The MetaLink-specific classes and properties are displayed in red, the blue classes and properties are reused from existing vocabularies.
Open Governmental Data publishing has had mixed success. While many governmental bodies are publishing an increasing number of datasets online, the potential usefulness is rather low. This paper describes action research conducted within the context of the Dutch Cadastre's open data platform. We start by observing contemporary (Dutch) Open Data platforms and observe that dataset reuse is not always realized. We introduce Linked Open Data, which promises to deliver solutions to the lack of Open Data reuse. In the process of implementing Linked Data in practice, we observe that users face a knowledge and skill and that contemporary Linked Open Data tooling is often unable to properly advertise the usefulness of datasets to potential users, thereby hampering reuse. We therefore develop four components for Linked Data viewing to enhance the current situation, making it easier to observe what a dataset is about and which potential use cases it could serve.
In a decentralised knowledge representation system such as the Web of Data, it is common and indeed desirable for different knowledge graphs to overlap. Whenever multiple names are used to denote the same thing, owl:sameAs statements are needed in order to link the data and foster reuse. Whilst the deductive value of such identity statements can be extremely useful in enhancing various knowledge-based systems, incorrect use of identity can have wide-ranging effects in a global knowledge space like the Web of Data. With several works already proven that identity in the Web is broken, this survey investigates the current state of this sameAs problem. An open discussion highlights the main weaknesses suffered by solutions in the literature, and draws open challenges to be faced in the future.
This paper presents an empirical study aiming at understanding the modeling style and the overall semantic structure of Linked Open Data. We observe how classes, properties and individuals are used in practice. We also investigate how hierarchies of concepts are structured, and how much they are linked. In addition to discussing the results, this paper contributes (i) a conceptual framework, including a set of metrics, which generalises over the observable constructs; (ii) an open source implementation that facilitates its application to other Linked Data knowledge graphs.
In the absence of a central naming authority on the Semantic Web, it is common for different datasets to refer to the same thing by different IRIs. Whenever multiple names are used to denote the same thing, owl :sameAs statements are needed in order to link the data and foster reuse. Studies that date back as far as 2009, have observed that the owl : sameAs property is sometimes used incorrectly. In this paper, we show how network metrics such as the community structure of the owl: sameAs graph can be used in order to detect such possibly erroneous statements. One benefit of the here presented approach is that it can be applied to the network of owl : sameAs links itself, and does not rely on any additional knowledge. In order to illustrate its ability to scale, the approach is evaluated on the largest collection of identity links to date, containing over 558M owl : sameAs links scraped from the LOD Cloud.