Deep neural networks, especially the Convolutional Neural Network (CNN) models, have shown promising results in multivariate time series data analysis. However, the predictions of these data-driven black-box models are tough to interpret from a human perspective, making it questionable to trust and rely on the predictions made by these models, specifically for time series data with the append-only feature. This paper proposes a new approach to interpret the CNN outputs by extracting and clustering the activated time series sequences learned from a trained network. These sequences show the representative features for each output label and form interpretable representations from the original time series data. Our approach is the first framework to identify each signal’s role and dependencies, consider all possible combinations of signals in the multivariate time-series input, and visualize the data representative features. Our experiments on the Baydogan’s archive indicate remarkable improvements in the interpretability of the network predictions and relation identification of each input signal to the output label and the channels of the network layers. Furthermore, the conducted experiments confirm that the extracted patterns are representative of the multivariate input and changing them results in a drastic reduction in the prediction accuracy.
Temporal web graphs have been attracting much attention recently due to their important applications in web search, data mining, and social network analysis. Accumulated over long periods, those graphs have grown gigantic in size and rich in temporal evolution, which poses tough challenges for data storage and management. Though a few temporal graph management systems were previously proposed, none of them can simultaneously satisfy both essential requirements when retrieving on temporal web graphs: very large data scalability and very low querying latency.In this work, we address the above gap in existing works by developing a highly efficient temporal graph management system which is dedicated to web graphs. To this end, we greatly extend the most efficient framework for managing large static web graphs to handle temporal information using the property matrix while preserving most of the outstanding features of the base framework. Ultimately, our proposed system can achieve a nearly instant response for vertex-centric temporal retrieval while still being scalable to huge datasets. Experiments on a real-world dataset with more than 43B nodes and 317B links show that using a small non-dedicated cluster, our system can reach a reduction of data storage space up to 88% of raw data size and reduce the retrieval time by 20%, compared to the baselines. We also demonstrate that our system also yields a significant reduction of computational costs for many graph ranking algorithms.
Criminal investigations contain sensitive and confidential material and are nonpublic by nature. Access to investigation data is very limited and restricted to only selected groups of individuals. Even for research purposes, data typically cannot be accessed freely. Within criminal investigations, data is still processed manually to a large extent. Solutions provided for automation of this processing — or even of individual processing steps — can be assumed to have a significant impact on the work of Law Enforcement Agencies (LEAs). Automation may effectively be key to handle large and complex amounts of data in an efficient manner under the typical operating conditions of LEAs. This paper introduces the ROXANNE Simulated Dataset (ROXSD), a dataset with unique properties prepared by the ROXANNE Project with assistance from several LEAs, to facilitate the development and evaluation of novel tools and technologies for criminal investigations. ROXSD consists of a set of simulated intercepted telephone conversations in a variety of languages. The story follows a realistic setting and includes the conditions and constraints of a real investigation. The network topology corresponding to the conversations was created by partner LEAs to reflect various typical organized crime groups. Conversations have been transcribed carefully and annotated in the original language and in English. The dataset is expected to provide a sound basis for further research and is available to download for researchers under signed agreement.
Data have an important role in evaluating the performance of NILM algorithms. The best performance of NILM algorithms is achieved with high-quality evaluation data. However, many existing real-world data sets come with a low sampling quality, and often with gaps, lacking data for some recording periods. As a result, in such data, NILM algorithms can hardly recognize devices and estimate their power consumption properly. An important step towards improving the performance of these energy disaggregation methods is to improve the quality of the data sets. In this paper, we carry out experiments using several methods to increase the sampling rate of low sampling rate data. Our results show that augmentation of low-frequency data can support the considered NILM algorithms in estimating appliances’ consumption with a higher F-score measurement.
Search analytics and trends data is widely used by media, politicians, economists, and scientists in various decision-making processes. The data providers often use sampling when calculating the request results, due to the huge data volume that would need to be processed otherwise. The representativity of such samples is typically assured by the providers. Often, limited or no information about the reliability and validity of the service or the sampling confidence are provided by the services and, as a consequence, the data quality has to be assured by the users themselves, before using it for further analysis. In this paper, we develop an experimental setup to estimate and measure possible variation in service results for the example of Google Trends. Our work demonstrates that the inconsistencies in Google Trends Data and the resulting contradictions in analyses and predictions are systematic and particularly large when analyzing timespans of eights months or less. In our experiments, the representativity claimed by the service was disproved in many cases. We found that beyond search volume and timespan, there are additional factors for the deviations that can only be explained by Google itself. When working with Google Trends data, users must be aware of the marked risks associated with the inconsistencies in the samples.
Node embedding has recently shown state-of-the-art performance in various network analysis tasks. However, most of the existing node embedding methods do not consider the uncertainty of the input data, which is often the case in practice. This work offers an empirical evaluation of the typical node embedding methods when applied on uncertain networks. Precisely, we examine the performance of embedding vectors obtained by these methods in a set of downstream tasks. To this end, we employ a wide range of uncertain networks and traditional prepossessing techniques for dealing with uncertainty. Our findings suggest that the existing node embedding methods perform practically well on networks with uncertainty once the network data is appropriately prepossessed.
Trying to comprehend the structure and content of large text corpora can be a daunting and often time consuming task. In this paper, we introduce a novel tool that exploits the structural properties for extracting and visualizing the underlying topics in a given dataset. To this end, we make use of a combination of latent topic analysis, discriminative feature selection applied on top of the category structure of corpora, and various ranking methods in order to extract the most representative topics for a given corpus. The visual moniker to depict the outcome of these methods can be chosen based on the context. Such visual representations can be useful for depicting trends, identifying "hot" topics, and discovering interesting patterns in the underlying data. As applications, we create example representations for a variety of corpora obtained from conference proceedings, movie summaries, and newsgroup postings. Our user experiments demonstrate the viability of our approach, with a flower-like visualization inspired by the "wheel of emotion", for generating high quality representative topics and for unearthing hidden structures and connections in large document corpora.
Online citizen science projects have been increasingly used in a variety of disciplines and contexts to enable large-scale scientific research. The successes of such projects have encouraged the development of customisable platforms to enable anyone to run their own citizen science project. However, the process of designing and building a citizen science project remains complex, with projects requiring both human computation and social aspects to sustain user motivation and achieve project goals. In this paper, we conduct a systematic survey of 48 citizen science projects to identify common features and functionality. Supported by online community literature, we use structured walk-throughs to identify different mechanisms used to encourage volunteer contributions across four dimensions: task visibility, goals, feedback, and rewards. Our findings contribute to the ongoing discussion on citizen science design and the relationship between community and microtask design for achieving successful outcomes.
Street art is public art; it's accessible; it's of the people; it's an urban voice; it's on public view; it's on-the-street. Nonetheless, the World Wide Web has been party responsible for street art becoming both recognised and popular. As a result, this study is investigating how street art is represented on the open photo-sharing platform, Flickr. It is a social network site offering a large portfolio of photographs showing a wide range of images, which have been categorised and classified using tags'. By using a visual content analysis based on theoretically determined categories to examine the uploaded photographs, this investigation will shed light on what a selection of Flickr users recognise as street art and how they record and index it.
Social graph construction from various sources has been of interest to researchers due to its application potential and the broad range of technical challenges involved. The World Wide Web provides a huge amount of continuously updated data and information on a wide range of topics created by a variety of content providers, and makes the study of extracted people networks and their temporal evolution valuable for social as well as computer scientists. In this paper we present SocGraph - an extraction and exploration system for social relations from the content of around 2 billion web pages collected by the Internet Archive over the 17 years time period between 1996 and 2013. We describe methods for constructing large social graphs from extracted relations and introduce an interface to study their temporal evolution.
Many modern data analytics applications in areas such as crisis management, stock trading, and healthcare, rely on components capable of nearly real-time processing of streaming data produced at varying rates. In addition to automatic processing methods, many tasks involved in those applications require further human assessment and analysis. However, current crowdsourcing platforms and systems do not support stream processing with variable loads. In this paper, we investigate how incentive mechanisms in competition based crowdsourcing can be employed in such scenarios. More specifically, we explore techniques for stimulating workers to dynamically adapt to both anticipated and sudden changes in data volume and processing demand, and we analyze effects such as data processing throughput, peak-to-average ratios, and saturation effects. To this end, we study a wide range of incentive schemes and utility functions inspired by real world applications. Our large-scale experimental evaluation with more than 900 participants and more than 6200 hours of work spent by crowd workers demonstrates that our competition based mechanisms are capable of adjusting the throughput of online workers and lead to substantial on-demand performance boosts.
Many modern data analytics applications in areas such as crisis management, stock trading, and healthcare, rely on components capable of nearly real-time processing of streaming data produced at varying rates. In addition to automatic processing methods, many tasks involved in those applications require further human assessment and analysis. However, current crowdsourcing platforms and systems do not support stream processing with variable loads. In this paper, we investigate how incentive mechanisms in competition based crowdsourcing can be employed in such scenarios. More specifically, we explore techniques for stimulating workers to dynamically adapt to both anticipated and sudden changes in data volume and processing demand, and we analyze effects such as data processing throughput, peak-to-average ratios, and saturation effects. To this end, we study a wide range of incentive schemes and utility functions inspired by real world applications. Our large-scale experimental evaluation with more than 900 participants and more than 6200 hours of work spent by crowd workers demonstrates that our competition based mechanisms are capable of adjusting the throughput of online workers and lead to substantial on-demand performance boosts.
Social network analysis is leveraged in a variety of applications such as identifying influential entities, detecting communities with special interests, and determining the flow of information and innovations. However, existing approaches for extracting social networks from unstructured Web content do not scale well and are only feasible for small graphs. In this paper, we introduce novel methodologies for query-based search engine mining, enabling efficient extraction of social networks from large amounts of Web data. To this end, we use patterns in phrase queries for retrieving entity connections, and employ a bootstrapping approach for iteratively expanding the pattern set. Our experimental evaluation in different domains demonstrates that our algorithms provide high quality results and allow for scalable and efficient construction of social graphs.
Obwohl es kaum denkbar ist, dass jemand private Informationen wie das Geburtsdatum oder eine private Fotosammlung einer unbekannten Person auf der Straße mitteilt, werden dennoch im Internet solche persönlichen Daten Tag für Tag von Benutzern öffentlich zugänglich gemacht. Nach Bekanntgabe solcher Informationen hat der Benutzer weder Einfluss darauf, wo und wie lange sie gespeichert werden, noch Kenntnis darüber, wer Zugang zu den Daten hat. Besonders viele Daten dieser Art werden in sozialen Netzen geteilt. In dieser Arbeit beschäftigen wir uns damit, Verfahren und Modelle zu entwickeln, die auf der einen Seite dem Benutzer erlauben, vertrauliche Daten sicher und effizient zu indizieren und danach zu suchen, auf der anderen Seite den Benutzer automatisch auf die Vertraulichkeit der Daten aufmerksam machen.
Many data processing tasks such as semantic annotation of images, translation of texts in foreign languages, and labeling of training data for machine learning models require human input, and, on a large scale, can only be accurately solved using crowd based online work. Recent work shows that frameworks where crowd workers compete against each other can drastically reduce crowdsourcing costs, and outperform conventional reward schemes where the payment of online workers is proportional to the number of accomplished tasks ("pay-per-task"). In this paper, we investigate how team mechanisms can be leveraged to further improve the cost efficiency of crowdsourcing competitions. To this end, we introduce strategies for team based crowdsourcing, ranging from team formation processes where workers are randomly assigned to competing teams, over strategies involving self-organization where workers actively participate in team building, to combinations of team and individual competitions. Our large-scale experimental evaluation with more than 1,100 participants and overall 5,400 hours of work spent by crowd workers demonstrates that our team based crowdsourcing mechanisms are well accepted by online workers and lead to substantial performance boosts.
Im vergangenen Jahrzehnt ist das Interesse an der wissenschaftlichen Nachnutzung des Materials aus früheren qualitativen sozialwissenschaftlichen Studien gestiegen. Im Kontrast dazu steht das bisher bescheidene Volumen an für sekundäranalytische Forschung archivierten und genutzten Datensätzen. Dieses Papier gibt einen Überblick über die dahinterstehenden Daten- und Archivierungsprobleme und die Möglichkeiten, über eine Kombination aus manuellen und IT-Werkzeugen eine wesentliche Hürde des Forschungsprozesses - die Auswahl geeigneten Forschungsmaterials - zu überwinden.
Most of the modern Web platforms have recognized the demand for resource sharing among their users and provide powerful tools for upload and sharing of user-generated content and personal documents. The need for effective information sharing and search within the growing amount of documents pushed forward further development of search infrastructures for enterprise data management systems. Applications for creation of the real, as well as, of the virtual communities are developed and connect people that are working on similar topics or share similar interests. The power of such tools and the easiness of their use, often make users careless about possible threats with respects to confidentiality of the information that is disclosed during the sharing. Privacypreserving document exchange among collaboration groups in social Web and across enterprises requires efficient techniques for sharing and search of access-controlled information in largely untrusted environments. Thus, privacy preservation mechanisms need to be, by design, an inseparable part in sharing systems and tools. In this thesis we propose solutions for confident information storage, indexing and sharing. In particular, we address: (1) Efficient privacy-preserving indexing and retrieval of shared documents; (2) Scalable content diversity analysis for large document corpora suitable for supporting efficient access control for an inverted index; and (3) Privacy-oriented content analysis of shared content to automatically distinguish public from potentially sensitive content. First, we introduce ZERBER+R a facility, which enable indexing sensitive documents in a way that the information leakage from the index is bounded by a tunable parameter. This facility is able to efficiently and securely locate and select top-k relevant documents the user is allowed to access without leaking information about contained documents and their term distribution statistics. To provide tunable resistance to statistical attacks, ZERBER+R employs a novel term merging scheme that has minimal impact on the index lookup costs and relies on largely untrusted index servers. Our experiments have shown that ZERBER+R makes economical use of network bandwidth, requires minimal key management, and answers queries almost as fast as an ordinary inverted index. Second, we develop two novel, efficient algorithms for content analysis of large document collections to estimate their topical diversity with probabilistic guarantees. Among many application scenarios, this measure is potentially useful to improve access control properties by supporting the index partitioning. Whilst state-of-the-art diversity measurement algorithms require quadratic computational time with respect to the dataset size, our methods, called TrackDJ and SampleDJ, can be executed in linear time and constant time, respectively. Finally, we present a framework for privacy degree estimation and corresponding classification of photographs as they are shared and published by the users of social web applications. The system supports users in selecting adequate privacy settings. To this end we trained classification models on a large-scale dataset obtained in a social annotation game. We analysed the usefulness of textual and visual information in the context of privacy-oriented image classification. In addition, we developed privacy-oriented search and privacy-based search result diversification techniques. Our classification and ranking experiments have shown excellent performance for textual, and visual information and their combinations. The results presented in this thesis can be directly applied in various privacy-aware data sharing scenarios and Web applications such as Flickr, Facebook or Twitter and due to high efficiency the developed algorithms can be integrated in large public and enterprise search engines. Additionally, our theoretical findings open a number of research directions such as studying the efficiency of secured indexes, intelligent index partitioning according to access rights, and sensitivity degree estimation of personal text and multimedia objects like videos or emails.
Social network analysis is leveraged in a variety of applications such as identifying influential entities, detecting communities with special interests, and determining the flow of information and innovations. However, existing approaches for extracting social networks from unstructured Web content do not scale well and are only feasible for small graphs. In this paper, we introduce novel methodologies for query-based search engine mining, enabling efficient extraction of social networks from large amounts of Web data. To this end, we use patterns in phrase queries for retrieving entity connections, and employ a bootstrapping approach for iteratively expanding the pattern set. Our experimental evaluation in different domains demonstrates that our algorithms provide high quality results and allow for scalable and efficient construction of social graphs.
Many data processing tasks such as semantic annotation of images, translation of texts in foreign languages, and labeling of training data for machine learning models require human input, and, on a large scale, can only be accurately solved using crowd based online work. Recent work shows that frameworks where crowd workers compete against each other can drastically reduce crowdsourcing costs, and outperform conventional reward schemes where the payment of online workers is proportional to the number of accomplished tasks ("pay-per-task"). In this paper, we investigate how team mechanisms can be leveraged to further improve the cost efficiency of crowdsourcing competitions. To this end, we introduce strategies for team based crowdsourcing, ranging from team formation processes where workers are randomly assigned to competing teams, over strategies involving self-organization where workers actively participate in team building, to combinations of team and individual competitions. Our large-scale experimental evaluation with more than 1,100 participants and overall 5,400 hours of work spent by crowd workers demonstrates that our team based crowdsourcing mechanisms are well accepted by online workers and lead to substantial performance boosts.
Crowd based online work is leveraged in a variety of applications such as semantic annotation of images, translation of texts in foreign languages, and labeling of training data for machine learning models. However, annotating large amounts of data through crowdsourcing can be slow and costly. In order to improve both cost and time efficiency of crowdsourcing we examine alternative reward mechanisms compared to the "Pay-per-HIT" scheme commonly used in platforms such as Amazon Mechanical Turk. To this end, we explore a wide range of monetary reward schemes that are inspired by the success of competitions, lotteries, and games of luck. Our large-scale experimental evaluation with an overall budget of more than 1,000 USD and with 2,700 hours of work spent by crowd workers demonstrates that our alternative reward mechanisms are well accepted by online workers and lead to substantial performance boosts.