Integration of data on coin finds into the ARIADNE portal presents a number of challenges, partly related to the use of the Getty AAT as a controlled vocabulary, in particular its incomplete coverage and hierarchical structure. Further problems arise from the fact that the majority of the data are being provided not from primarily numismatic projects and institutions, but rather from disparate archaeological resources that integrate a wide range of artefacts and records. As a result, they often focus less on the peculiarities of using the AAT for coin finds and the difficulties that arise from its use. This article illustrates how this can lead to quite disparate and inconsistent mappings of data to the ARIADNE portal. Potential solutions such as aligning the AAT to the established vocabulary of Nomisma.org, or even implementing the latter in the portal, as well as implementing a standard mapping for coin finds, are discussed. Also addressed are the possibilities for complex, granular searches using the SPARQL endpoint in the ARIADNElab Virtual Research Area.
The ethical and societal implications of artificial intelligence systems raise concerns. In this article, we outline a novel process based on applied ethics, namely, Z-Inspection®, to assess if an AI system is trustworthy. We use the definition of trustworthy AI given by the high-level European Commission’s expert group on AI. Z-Inspection® is a general inspection process that can be applied to a variety of domains where AI systems are used, such as business, healthcare, and public sector, among many others. To the best of our knowledge, Z-Inspection® is the first process to assess trustworthy AI in practice.
Corpora, thing-editions, are a central epistemic instrument of knowledge generation for archaeology. Due to digitisation attention has turned once more to various strategies and the politics of representation, and in the course of the debate on (post-)factualism and cultures of knowledge and data we are striving to render the production of knowledge more visible and comprehensible. Following B. Latour’s concept of circulating reference, editions are products of a praxeological connection between the world and representations. This is a report on the ongoing discussion of central questions and objectives for digital corpora and research data management following the FAIR (findable, accessible, inter-operable, re-usable) principle.
zur Konferenz Digital Humanities im deutschsprachigen Raum 2019 Nomisma.org: Numismatik und das Semantic Web
Iconographic representations on ancient artifacts are described in many existing databases and literature as human readable text. We applied Natural Language Processing (NLP) approaches in order to extract the semantics out of these textual descriptions and in this way enable semantic searches over them. This allows more sophisticated requests compared to the common existing keyword searches. As we show in our experiments based on numismatic datasets, the approach is generic in the sense that once the system is trained on one dataset, it can be applied without any further manual work also to datasets that have similar content. Of course, additional adaptions would further improve the results. Since the approach requires manual work only during the training phase, it can easily be applied to huge datasets without manual work and therefore without major extra costs. In fact, in our experience bigger datasets generate even better results because there is more data for training. Since our approach is not bound to a certain domain and the numismatic datasets are just an example, it could serve as a blueprint for many other areas. It could also help to build bridges between disciplines since textual iconographic descriptions are to be found also for pottery, sculpture and elsewhere.
One of the application fields for data analysis techniques and technologies gaining momentum is the area of social good or “common good”, covering cases related to humanitarian crises, global health care, or ecology and environmental issues, among others. The promotion of data-driven projects in this field aims at increasing the efficacy and efficiency of social initiatives, improving the way these actions help humanity in general and people in need in particular. This application field, however, poses its own barriers and challenges when developing data-driven projects, lagging behind in comparison with other scenarios. These challenges derive from aspects such as the scope and scale of the social issue to solve, cultural and political barriers, the skills of main stakeholders and the technological resources available, the motivation to be engaged in such projects, or the ethical and legal issues related to sensitive data. This paper analyzes the application of data projects in the field of social good, reviewing its current state and noteworthy initiatives, and presenting a framework covering the key aspects to analyze in such projects. The goal is to provide guidelines to understand the main challenges and opportunities for this type of data project, as well as identifying the main differential issues compared to “classical” data projects in general. A case study is presented on the initial steps and stakeholder analysis of a data project for the inclusion of refugees in the city of Frankfurt, Germany, in order to empirically confront the framework with a real example. Keywords—Data-Driven projects, humanitarian operations, personal and sensitive data, social good, stakeholders analysis.
In the first part of this chapter we illustrate how a big data project can be set up and optimized. We explain the general value of big data analytics for the enterprise and how value can be derived by analyzing big data. We go on to introduce the characteristics of big data projects and how such projects can be set up, optimized and managed. Two exemplary real word use cases of big data projects are described at the end of the first part. To be able to choose the optimal big data tools for given requirements, the relevant technologies for handling big data are outlined in the second part of this chapter. This part includes technologies such as NoSQL and NewSQL systems, in-memory databases, analytical platforms and Hadoop based solutions. Finally, the chapter is concluded with an overview over big data benchmarks that allow for performance optimization and evaluation of big data technologies. Especially with the new big data applications, there are requirements that make the platforms more complex and more heterogeneous. The relevant benchmarks designed for big data technologies are categorized in the last part.
The interest to become a data scientist or related professions in data science domain is rapidly growing. To meet such a demand, we propose a novel educational service that aims to provide a tailored learning paths for data science. Our target user is one who aims to be an expert in data science. Our approach is to analyze the background of the practitioner and match the learning units. A critical feature is that we use gamification to reinforce the practitioner engagement. We believe that our work provides a practical guideline for those who wants to learn data science.
To be able to handle big data workloads, modern NoSQL database management systems like Cassandra are designed to scale well over multiple machines. However, with each additional machine in a cluster, the likelihood for hardware failure increases. In order to still achieve high availability and fault tolerance, the data needs to be replicated within the cluster. Predictable and stable response times are required by many applications even in the case of a node failure. While Cassandra guarantees high availability, the influence of a node failure on the system performance is still unclear.
Natascha Hoebel合作论文数4