Sharing data between researchers, whether openly or not, requires effort, particularly concerning its metadata. What is the minimum metadata needed to aid discovery? Once data has been discovered, what metadata is needed in order to be able to evaluate its usefulness? And, since it's not realistic to expect everyone to use the same metadata standard to describe data, how can different systems interoperate with the metadata that is commonly provided? These topics and more were discussed in Amsterdam in late 2016.
e BigDataEurope (BDE) project is developing exactly the kind of computing infrastructure that European stakeholders need when handling large volumes of data in a variety of formats; the results are open-source and their use is completely free. Coordinated by Fraunhofer IAIS, BDE is working directly with partners that represent the seven Societal Challenges identied by the European Commission (Health, Food, Energy, Transport, Climate, Social Sciences and Security). For each community, a pilot that makes use of BDEs technology stack to address the Big Data needs identied by these challenges is well under way. 1 THE BIG DATA INTEGRATOR PLATFORM BDE’s Integrator Platform (BDI) makes the processing of big data simpler, cheaper and more exible than ever before. It oers basic building blocks to get started with common big data technologies and makes integration of dierent technologies and applications easy. Components such as Apache Spark, Hadoop HDFS, Apache Flink, Apache Flume and Apache Kaa can be built into a pipeline through a simple graphical UI. ose components can help handle the velocity and volume dimensions, but BDI is also leading the way in tackling that third big data problem: variety. is is done through BDI’s Semantic Data Lake and components like SANSA1 which performs analytics on semantically structured RDF data by providing out-of-the-box scalable algorithms for massive datasets. BDI is an open source platform based on Docker, today’s virtualisation technique of choice. It works on a local machine or on hundreds of nodes using Docker Swarm, and can run in-house, or within an external cloud environment (not provided by BDE). BDE applications are provided as docker containers, making their installation and set-up a 10-minute job. With the help of latest Docker features, BDI oers: • Swarm-based networking • Load Balancing • Service Discovery • Multi-host networking with integrated KV-Store • Fault tolerance Docker Compose helps to create multiple containers on multiple nodes using a single command and a single compose le. Docker Compose V2 and Docker Swarm aim to implement full integration, 1hp://sansa-stack.net/ which means that it is feasible to point a Compose app at a Swarm cluster and make its use possible in the same manner as if a single Docker host is being used. It is notable that the latest Docker components provide greater resemblance to Kubernetes in terms of orchestration features, and Swarm presents a beer choice in terms of shiing from a local/development environment to a cluster. e BDE Team provides baseline Docker images for Apache Hadoop, Spark, Flink and many others. Components were selected based on the requirements gathered from the seven Societal Challenges. us, the Platform makes it feasible to perform a variety of big data tasks, including message passing (Kaa, Flume), storage (Hive, Cassandra). e platform is able to handle RDF triples at scale using components like FOX, SemaGrow and 4Store; with particular emphasis on the triplication of geospatial data using GeoTriples, Sextant and Strabon. BDI has enriched the Docker platform, a high-level depiction of which is shown in Figure 1, with a layer of supporting services, helping in the setup, maintenance and monitoring of the pipeline and workows: • e Init daemon allows to dene workows by monitoring the start-up status of inter-dependent Docker components. • e Pipeline Service and Builder are developed to support the creation of workows. • e Pipeline Monitor front-end demonstrates the current status of the Docker components. • e Integrator UI integrates the dierent ocialWeb UIs of select pipeline components under one Integrated and personalised view. Furthermore, the Swarm UI visualises the status of a swarm cluster and allows to scale and monitor the cluster services. Figure 1: BDI platform’s high-level modular architecture For BDI platform progress updates please refer to the dedicated page2; or try it out or engage with our community3.
In this paper, we present the results of the user requirements and interface design phase for a prototype system, designed to enhance interaction with cultural heritage collections online through means of a pathway metaphor. We present a single user interaction model that supports various work and information seeking tasks undertaken by both expert and non-expert users within the context of collection exploration and path creation. The user interaction model is shown to enable seamless movement between interaction modes, with the potential over time to encourage deeper engagement and learning.
The POWDER protocol is a Semantic Web technology that takes advantage of natural groupings of URIs to annotate all the resources in a regular expression-delineated sub-space of the URI space. POWDER is a mechanism for accreditation, trustmarking and resource discovery, emphasising the publishing of attributed metadata by third parties and trusted authorities. Demonstrating its versatility, it has also been deployed in unforeseen use cases, such as repository compression. In this paper, we present the POWDER protocol, explain its position in the Semantic Web architecture, expose and discuss current implementations and use cases and future directions.
Digitisation of the cultural heritage means that a significant amount of material is now available through online digital library portals. However, the vast quantity of cultural heritage material can also be overwhelming for many users who lack knowledge of the collections, subject knowledge and the specialist language used to describe this content. Search portals often provide little or no guidance on how to find and interpret this information. The situation is very different in museums and galleries where collections are organized in exhibitions which offer themes and stories that visitors can explore. The PATHS project, which is funded under the European Commission's FP7 programme, is developing a system that explores the familiar metaphor of a trail (pathway) to enhance the discovery and use of the content made available in digital libraries. This paper will report on the findings of the user requirements analysis and the specifications for the first prototype of the PATHS system which is based on contents from Europeana and the Alinari Archives.
In this paper we present and discuss the implementation and deployment of the Protocol for Web Description Resources (POWDER) W3C Recommendation for a large RDF repository containing millions of triples. POWDER enables taking advantage of natural groupings of URIs and their reflection on the denoted things' properties; our application implements a POWDER service that intercepts the API between the RDF store and the inference layer above it and provides annotations that appear as explicit statements to the inference service. The approach is tested on a multi million-triple store of news documents and events, where it achieves dramatic savings on storage space without impacting querying time.
The QUATRO Plus project, a follow on from the original QUATRO Project, aims to balance the wisdom of the crowds with the knowledge of the experts. It uses a mixture of authenticated data sources and the opinions of end users expressed through social networking software to build a dataset that is authoritative and trustworthy. The dataset describes online resources using RDF with the upcoming W3C Recommendation, POWDER, as the underlying transport and storage mechanism. Data can be added to or queried through a variety of tools provided by the project, some of which are described in detail in this paper.
I engineered and built the Sony Walkman Pro tape delay in summer 1983. I still use it in performance, with the same cassettes and Walkman recorders, some 24 years later. It became really tiresome to arrange for two stereo tape recorders at every gig, and I was determined to find a way to make this delay system. The Walkman was perfect for this purpose because it had very accurate speed control (simply leave the speed tune knob to OFF) and the recorders themselves needed no modification. I know
QUATRO is an on-going EC-funded project which aims to provide a common vocabulary and machine readable schema for quality labeling of Web content, as well as ways to automatically show the contents of the label(s) found in a Web resource, and functionalities for checking the validity of these labels. The paper presents the QUATRO processes for label validation and user notification, and outlines the architecture of QUATRO system.
As the number of medical websites in various languages increases, it is increasingly necessary to establish specific criteria and control measures that give consumers some guarantee that the health websites they are visiting meet a minimum level of quality standards. Further, reassurance is needed that the professionals offering the information are suitably qualified. The paper briefly presents the current mechanisms for labelling medical web content and introduces the work done in the EC-funded project Quatro. This has defined a vocabulary for quality labels and a schema to deliver them in a machine-processable format. In addition, the paper proposes the development of a labelling platform that will assist the work of medical labelling agencies in automating, up to a certain level, the retrieval of unlabelled medical websites and their labelling, and the monitoring of labelled websites as to whether they are still satisfying the criteria.
The semantic web is a truly exciting development and has the potential to revolutionise the way people access online data. This is obvious to those involved with its development. However, conveying that potential and sharing the enthusiasm with policy makers can be very difficult . The semantic web must be repackaged as a commercial sales tools.
Andrea Perego合作论文数European Commission, Joint Research Centre2