In order to perform cutting-edge research like AI model training, a large amount of data needs to be accessed. However, data providers are often reluctant to share their data with researchers as these might contain personal data and thereby sharing may introduce serious risks with significant personal, institutional or societal impacts. Apart from the need to control these risks, data providers must also comply with regulations like GDPR, which creates an additional overhead that makes data sharing even less appealing to data providers. Technologies like anonymization can play a critical role when sharing data that may contain personal information by offering privacy preservation measures like face or license plate anonymization. Therefore, we propose a framework to support data sharing of personal data for research by integrating anonymization, risk assessment and automatic licence agreement generation. The framework offers a practical and efficient solution for organisations seeking to enhance data-sharing practices without compromising information security.
The publication focuses on exploring ways of integrating smart contracts within a data space dedicated to the public security domain. This data space shall facilitate involving various stakeholders, such as law enforcement agencies (LEAs), research facilities and universities as well as and research-oriented companies by taking their needs and requirements into account. On the one hand, we discuss the benefits of smart contracts in this context, particularly how blockchain-backed processes can contribute to building a trustworthy data sharing ecosystem. On the other hand, we also investigate challenges and costs associated with this approach.
Given the numerous opportunities provided by rapidly evolving digital innovations, we need to address and assess the social and political risks that come with naively applying AI algorithms, especially in high-risk sectors. We argue that though the existing guidelines and regulations are a good starting point, we still need to implement effective solutions that can be integrated into the current workflow of developing ethical AI applications. We introduce the idea of AI Ethics Labs as institutionalised "spaces for doubt" providing platforms for a frequent and intensive collaboration between developers and social scientists, thus reducing the potential risks of developed algorithms.
In this paper, we present deep learning frameworks for audio-visual scene classification (SC) and indicate how individual visual and audio features as well as their combination affect SC performance.Our extensive experiments, which are conducted on DCASE (IEEE AASP Challenge on Detection and Classification of Acoustic Scenes and Events) Task 1B development dataset, achieve the best classification accuracy of 82.2\%, 91.1\%, and 93.9\% with audio input only, visual input only, and both audio-visual input, respectively.The highest classification accuracy of 93.9\%, obtained from an ensemble of audio-based and visual-based frameworks, shows an improvement of 16.5\% compared with DCASE baseline.
Extraction of event causality and especially implicit causality from text data is a challenging task. Causality is often treated as a specific relation type and can be considered as a part of relation extraction or relation classification task. Many causality identification-related tasks are designed to select the most plausible alternative of a set of possible causes and consider multiple-choice classification settings. Since there are powerful Question Answering (QA) systems pretrained on large text corpora, we investigated a zero-shot QA-based approach for event causality extraction using a Wikipedia-based dataset containing event descriptions (articles) and annotated causes. We aimed to evaluate to what extent reading comprehension ability of the QA-pipeline can be used for event-related causality extraction from plain text without any additional training. Some evaluation challenges and limitations of the data were discussed. We compared the performance of a two-step pipeline consisting of passage retrieval and extractive QA with QA-only pipeline on event-associated articles and mixed ones. Our systems achieved average cosine semantic similarity scores of 44 45% in different settings.
Sexism has become an increasingly major problem on social networks during the last years. The first shared task on sEXism Identification in Social neTworks (EXIST) at IberLEF 2021 is an international competition in the field of Natural Language Processing (NLP) with the aim to automatically identify sexism in social media content by applying machine learning methods. Thereby sexism detection is formulated as a coarse (binary) classification problem and a fine-grained classification task that distinguishes multiple types of sexist content (e.g., dominance, stereotyping, and objectification). This paper presents the contribution of the AIT_FHSTP team at the EXIST2021 benchmark for both tasks. To solve the tasks we applied two multilingual transformer models, one based on multilingual BERT and one based on XLM-R. Our approach uses two different strategies to adapt the transformers to the detection of sexist content: first, unsupervised pre-training with additional data and second, supervised fine-tuning with additional and augmented data. For both tasks our best model is XLM-R with unsupervised pre-training on the EXIST data and additional datasets and fine-tuning on the provided dataset. The best run for the binary classification (task 1) achieves a macro F1-score of 0.7752 and scores 5th rank in the benchmark; for the multiclass classification (task 2) our best submission scores 6th rank with a macro F1-score of 0.5589.
This report shows a deep learning framework for audio-visual scene classification (SC). Our extensive experiments, which are conducted on DCASE Task 1B development dataset, achieve the best classification accuracy of 82.2%, 91.1%, and 93.9% with audio input only, visual input only, and both audiovisual input, respectively.
We present DreamDrug, a crowdsourced dataset for detecting mentions of drugs in noisy user-generated item listings from darknet markets. Our dataset contains nearly 15,000 manually annotated drug entities in over 3,500 item listings scraped from the darknet market platform “DreamMarket” in 2017. We also train and evaluate baseline models for detecting these entities, using contextual language models fine-tuned in a few-shot setting and on the full dataset, and examine the effect of pretraining on in-domain unannotated corpora.
Over the past decade, the darknet has created unprecedented opportunities for trafficking in illicit goods, such as weapons and drugs, and it has provided new ways to offer crime as a service. Natural language processing techniques can be applied to find the types of goods that are traded in these markets. In this paper we present the results of evaluating state-of-the-art machine learning methods for the classification of darknet market offers. Several embeddings, such as GloVe embeddings [20], Fasttext [15], Tensor Flow Universal Sentence Encoder [7], Flair’s contextual string embedding [2] and term-frequency inverse-document-frequency (TF-IDF), as well as our domain-specific darknet embedding have been evaluated with a series of machine learning models, such as Random Forest, SVM, Naïve Bayes and Multilayer Perceptron. To find the best combination of feature set and machine learning model for this task, the performance was evaluated on a publicly available collection covering 13 darknet markets with more than 10 million product offers [6]. After extracting unique advertisements from the corpus, the classifier was trained on a subset with those advertisements that contain strings related to weapons. The purpose was to determine how well the classifier can distinguish between different types of advertisements which seem all to be related to weapons according to the keywords they contain. The best performance for this classification task was achieved using the Copyright © 2020 for this paper by its authors. Use permitted under Creative Commons License Attribution 4.0 International (CC BY 4.0). Proceedings of the 11th International Conference on Applied Informatics Eger, Hungary, January 29–31, 2020, published at http://ceur-ws.org
ABSTRACT This article sets in context Data Warehouses (DWs) and Online Analytical Processing (OLAP) against the backdrop of databases and Big Data and shows how data warehouses and OLAP were incorporated into the digital archiving specifications work in the E-ARK project. Both theoretical and practical aspects are thoroughly considered, and the article ends with a detailed use case of the steps needed to produce a Dissemination Information Package (DIP) for an OLAP representation for a data warehouse.
The Data Market Austria (DMA) is an ecosystem of federated data and service infrastructures. It aims at establishing a market platform where data assets can be made accessible and offered for purchase. To support this, the DMA offers a central portal with catalogue and search services as well as a set of microservices for metadata mapping or enrichment, data quality assessment, data set submission, storage, management, and dissemination. The DMA relies on blockchain technology to allow a network of DMA members sharing information about the provenance and trading of datasets and services. This paper describes the blockchain application scenarios and the implementation of the blockchain-based distributed setup of the DMA.
The Data Market Austria (DMA) is an ecosystem of federated data and service infrastructures. It aims at making data from various data providers accessible and interoperable by allowing the submission, storage, management and dissemination of static datasets or streaming data services. By creating a metadata vocabulary, standardizing the ingest of data and ensuring the quality and completeness of metadata, it lays the ground to enable participants to share or consume datasets residing in different infrastructures. This demo focuses on the mapping services used in the DMA to standardize data from different sources using a modified version of the DCAT metadata schema. We present tools that enable inter organizational integration of datasets, in a manner that is both user-friendly and powerful enough to handle vast amounts of data.
Conversational systems allow us to interact with computational and robotic systems. Such approaches are often deliberately limited to the context of a given task. We apply audio analysis to either broaden or to adaptively set this context based on identified surrounding acoustic scenes or events.