
Developing machine learning models can be seen as a process similar to the one established for traditional software development. A key difference between the two lies in the strong dependency between the quality of a machine learning model and the quality of the data used to train or perform evaluations. In this work, we demonstrate how different aspects of data quality propagate through various stages of machine learning development. By performing a joint analysis of the impact of well-known data quality dimensions and the downstream machine learning process, we show that different components of a typical MLOps pipeline can be efficiently designed, providing both a technical and theoretical perspective.
We are a group of volunteers — researchers, software engineers, privacy and public health experts — who have developed a privacy-preserving mobile app intervention to reduce the spread of COVID-19. Our mobile app performs automatic decentralized contact tracing using bluetooth proximity networks. Our volunteers care strongly about preserving human life and human rights. All data we could collect is voluntary and fully anonymized. All code is transparent. It is open source [1] and could be easily reviewed, reproduced and used anywhere on the planet. The app could be installed by anyone with a bluetooth-capable smartphone, alerting them to their risk of having been in contact with a confirmed case of COVID-19, and helping them to protect themselves and their friends, families, and other contacts altruistically. We believe scalable measures like an app are especially helpful in communities where contact tracing resources are too limited to match the scope of the pandemic. We’re building this app to provide components and tools that public health agencies can use to supplement their pre-existing efforts to fight COVID-19, assisted by voluntary public action.
Contact tracing is a critical part of reopening society. While manual contact tracing by medical professionals is essential, there’s growing acknowledgement that supplementing this with digital approaches might make a significant difference in the speed with which we can reopen. Safe Paths is open source, standards-based, privacy-first framework that works closely with public health entities. Our approach is to roll out apps, SDKs, privacy-preserving network backbones, and interoperable protocols, so that any developer can build experiences that perform contact-tracing and related activities in safe, easy to use ways. In this paper, we’ll compare some of the existing technologies that can be used to aid contact tracing and exposure notification efforts, and make the case for using a holistic, multi-modal approach rather than relying exclusively on a single technology.
The global health threat from COVID-19 has been controlled in a number of instances by large-scale testing and contact tracing efforts. We created this document to suggest three functionalities on how we might best harness computing technologies to supporting the goals of public health organizations in minimizing morbidity and mortality associated with the spread of COVID-19, while protecting the civil liberties of individuals. In particular, this work advocates for a third-party free approach to assisted mobile contact tracing, because such an approach mitigates the security and privacy risks of requiring a trusted third party. We also explicitly consider the inferential risks involved in any contract tracing system, where any alert to a user could itself give rise to de-anonymizing information. More generally, we hope to participate in bringing together colleagues in industry, academia, and civil society to discuss and converge on ideas around a critical issue rising with attempts to mitigate the COVID-19 pandemic.
The COVID-19 pandemic is a global health crisis of our time that has significantly affected almost every single person on earth in just several months. Even worse, we do not know when it will slow down and how long it will last. Analogous to fighting the COVID-19 pandemic in the physical world, data scientists need to deal with the infodemic of COVID-19 data to discover useful insight in order to guide wise and informative decisions, where the COVID-19 infodemic refers to all (messy) data relevant to COVID-19. In this paper, we present D EEP E YE , an end-to-end data science system for monitoring and exploring COVID-19 data, which ranges from (task-driven) data preparation, (descriptive, diagnostic, and prescriptive) data analytics, user interactions through (linked) spatio-temporal data visualizations, and applications in different use cases.
Our society currently faces the most profound and deeply disruptive public health crisis in modern history. As communities across the world grapple with the COVID-19 pandemic, scientific advances spanning biochemistry and epidemiology to manufacturing and data engineering offer hope—and a spectrum of guidance is unfolding in an effort to respond to monumental shifts in our daily lives. The rising demand for data and the emerging efforts to responsibly collect, share, and analyze information across traditional boundaries play a vital role in our next steps. From facilitating dialogue and decision making, to underscoring the importance of a shared, honest assessment of where we are in our collective fight against the pandemic, data are now more important than ever. Our computational capacity and the value that we as a community can generate through data science are fundamentally crucial to our resilience. As we look to transition from crisis response to longer-term recovery, we have an unprecedented opportunity to re-imagine a new data-enabled future.
Scientific data management has become an increasingly difficult challenge. Modern experiments and instruments are generating unprecedented volumes of data and their accompanying dataflows are becoming more complex. Straightforward approaches are no longer applicable at the required scales and there are few reports on long-term operational experiences. This article reports on our experiences in this field: we illustrate challenges in operating the exascale dataflows of the high energy physics experiment ATLAS, we detail the concepts and architecture of the data management system Rucio that was purposely built to take up these challenges, we describe how other experiments evaluated, modified, and adopted the Rucio system for their own needs, and we show how Rucio will evolve to cope with the ever increasing needs of the community.
Remote Procedure Call (RPC) has long been an inherent component of parallel file systems and I/O forwarding middleware in high-performance computing (HPC). RPCs are used in this environment to issue I/O operations and transfer data from compute nodes to gateway and server storage nodes. With HPC systems becoming more heterogeneous, data volumes reaching new thresholds, and I/O standing as the main bottleneck, there is a growing need in the HPC community to build distributed services and adopt new workflows that are, nonetheless, no longer dictated by monolithic parallel file systems. These include specialized storage, data analysis, and telemetry services that can be adapted to fit application needs. Parallel file system RPC facilities have never been exposed to service or middleware developers, however, leaving them with two choices: MPI or the low-level fabric network protocol. In this article, we show how an independent RPC framework can be used as a building block for developing user-level data services at exascale. We identify the design choices that must be considered in terms of both performance and resilience for HPC data services, and we discuss the directions taken to palliate current HPC system constraints.
Through the years, the database community has periodically looked at developments in technology and engaged in hand-wringing over the idea that we are becoming irrelevant. The cry “have we missed the boat – again” is common; e.g., here is a panel I served on several years ago [8]. My goal in this essay is to argue that the database field and the techniques that have come from this research are still essential for “data science,” that is, for the exploitation of data to solve problems of importance in application fields – science, commerce, medicine and such. I believe, as I assume most readers of this article believe, that the field of database systems has always had at its core the study of how to deal with the largest amounts of data possible at the time, whether that be megabytes of corporate payroll data, terabytes of genomic information, or petabytes of satellite output. Thus whatever study of data is necessary at the time – that’s our job. To advance this argument, I want to look at three issues:
The Adaptable I/O System (ADIOS) represents the culmination of substantial investment in Scientific Data Management, and it has demonstrated success for several important extreme-scale science cases. However, looking towards the exascale and beyond, we see the development of yet more stringent data management requirements that require new abstractions. Therefore, there is an opportunity to attempt to connect the traditional realms of HPC I/O optimization with the Database / Data Management community. In this paper, we offer some specific examples from our ongoing work in managing data structures, services, and performance at the extreme scale for scientific computing. Using the publish/subscribe model afforded by ADIOS, we demonstrate a set of services that connect data format, metadata, queries, data reduction, and high-performance delivery. The resulting publish/subscribe framework facilitates connection to on-line workflow systems to enable the dynamic capabilities that will be required for ex-
The combination of high-performance computing towards Exascale power and numerical techniques enables exploring complex physical phenomena using large-scale spatio-temporal modeling and simulation. The improvements on the fidelity of phenomena simulation require more sophisticated uncertainty quantification analysis, leaving behind measurements restricted to low order statistical moments and moving towards more expressive probability density functions models of uncertainty. In this paper, we consider the problem of answering uncertainty quantification queries over large spatio-temporal simulation results. We propose the SU Q 2 method based on the Generalized Lambda Distribution (GLD) function. GLD fitting is an embarrassingly parallel process that scales linearly to the number of available cores on the number of simulation points. Furthermore, the answer of queries is entirely based on computed GLDs and the corresponding clusters, which enables trading the huge amount of simulation output data by 4 values in the GLD parametrization per simulation point. The methodology presented in this paper becomes an important ingredient in converging simulations improvements to the Exascale computational power.
Contact tracing is an important method to control the spread of an infectious disease such as COVID-19. However, existing contact tracing methods alone cannot provide sufficient coverage and do not successfully address privacy concerns of the participating entities. Current solutions do not utilize the huge volume of data stored in business databases and individual digital devices. This information is typically stored in data silos and cannot be used due to regulations in place. To successfully unlock the potential of contact tracing, we need to consider both data utilization from multiple sources and the privacy of the participating parties. To this end, we propose BeeTrace, a unified platform that breaks data silos and deploys state-of-the-art cryptographic protocols to guarantee privacy goals.
Governments around the world have become increasingly frustrated with tech giants dictating public health policy. The software created by Apple and Google enables individuals to track their own potential exposure through collated exposure notifications. However, the same software prohibits location tracking, denying key information needed by public health officials for robust contract tracing. This information is needed to treat and isolate COVID-19 positive people, identify transmission hotspots, and protect against continued spread of infection. In this article, we present two simple ideas: the lighthouse and the covid-commons that address the needs of public health authorities while preserving the privacy-sensitive goals of the Apple and google exposure notification protocols.
This document describes and analyzes a system for secure and privacy-preserving proximity tracing at large scale. This system, referred to as DP3T, provides a technological foundation to help slow the spread of SARS-CoV-2 by simplifying and accelerating the process of notifying people who might have been exposed to the virus so that they can take appropriate measures to break its transmission chain. The system aims to minimise privacy and security risks for individuals and communities and guarantee the highest level of data protection. The goal of our proximity tracing system is to determine who has been in close physical proximity to a COVID-19 positive person and thus exposed to the virus, without revealing the contact's identity or where the contact occurred. To achieve this goal, users run a smartphone app that continually broadcasts an ephemeral, pseudo-random ID representing the user's phone and also records the pseudo-random IDs observed from smartphones in close proximity. When a patient is diagnosed with COVID-19, she can upload pseudo-random IDs previously broadcast from her phone to a central server. Prior to the upload, all data remains exclusively on the user's phone. Other users' apps can use data from the server to locally estimate whether the device's owner was exposed to the virus through close-range physical proximity to a COVID-19 positive person who has uploaded their data. In case the app detects a high risk, it will inform the user.
Contact tracing is an essential tool in containing infectious diseases such as COVID-19. Many countries and research groups have launched or announced mobile apps to facilitate contact tracing by recording contacts between users with some privacy considerations. Most of the focus has been on using random tokens, which are exchanged during encounters and stored locally on users' phones. Prior systems allow users to search over released tokens in order to learn if they have recently been in the proximity of a user that has since been diagnosed with the disease. However, prior approaches do not provide end-to-end privacy in the collection and querying of tokens. In particular, these approaches are vulnerable to either linkage attacks by users using token metadata, linkage attacks by the server, or false reporting by users. In this work, we introduce Epione, a lightweight system for contact tracing with strong privacy protections. Epione alerts users directly if any of their contacts have been diagnosed with the disease, while protecting the privacy of users' contacts from both central services and other users, and provides protection against false reporting. As a key building block, we present a new cryptographic tool for secure two-party private set intersection cardinality (PSI-CA), which allows two parties, each holding a set of items, to learn the intersection size of two private sets without revealing intersection items. We specifically tailor it to the case of large-scale contact tracing where clients have small input sets and the server's database of tokens is much larger.
Fairness is increasingly recognized as a critical component of machine learning systems. However, it is the underlying data on which these systems are trained that often reflects discrimination, suggesting a data management problem. In this paper, we first make a distinction between associational and causal definitions of fairness in the literature and argue that the concept of fairness requires causal reasoning. We then review existing works and identify future opportunities for applying data management techniques to causal algorithmic fairness.
At MIT, we have been collaborating on two real-world projects dealing with Machine Leaning (ML) and large amounts of data. The first project deals with performing ML on a 30T data set of EEG (electroencephalogram) sleep study data, in collaboration with Massachusetts General Hospital (MGH). They have collected EEG data on about 2000 patients with time durations varying from a couple of hours to more than 24 hours. In all, they have recorded 21 channels of data, with a total volume of about 30T. Furthermore, they have decomposed this data into about 450M segments, each 2 seconds long. Their project goal is to use ML to classify each 2-second segment “snippet” into one of 12 categories that make sense to doctors. Two papers on this project were presented at the recent VLDB conference [1, 6]. The second project is also in conjunction with MGH and deals with the spread of infections in the hospital. MGH has made 11 years of patient data available, including exactly which room each patient occupied from admission to discharge. For infected patients, they wish to infer how an infection was spread, i.e. by sharing a nurse, by sequentially inhabiting the same room, etc. Because MGH changed patient software in 2016, we first have a data integration problem, which we are addressing using MIT-built ML integration tools [1, 7]. Then, we propose to infer infection pathways from their integrated data. Finally, one of us has been a principal at an ML data integration company (Tamr) that has done more than 300 ML projects primarily for Fortune 2000 customers. Tamr is in the business of data integration at scale. Typically their projects deal with integrating several-to-many data sets by normalizing the data to a common format, correcting data errors, merging and deduplicating the result, and creating “golden records” for each cluster of records thought to represent the same entity. Typically, input data sets represent parts, suppliers, customers or other entities, and this end-to-end process is usually called “mastering” for a particular entity. The guts of the Tamr system is a collection of ML algorithms to perform schema matching, deduplication and golden record construction. This paper summarizes our thoughts, based on our experience with these use cases. The remainder of this paper is divided into sections containing our observations.