In the Big Data era, both the academic community and industry agree that a crucial point to obtain the maximum benefits from the explosive data growth is integrating information from different sources, and also combining methodologies to analyze and process it. For this reason, sharing data so that third parties can build new applications or services based on it is nowadays a trend. Although most data sharing initiatives are based on public data, the ability to reuse data generated by private companies is starting to gain importance as some of them (such as Google, Twitter, BBC or New York Times) are providing access to part of their data. However, current solutions for sharing data with third parties are not fully convenient to either or both data owners and data consumers. Therefore we present dataClay, a distributed data store designed to share data with external players in a secure and flexible way based on the concepts of identity and encapsulation. We also prove that dataClay is comparable in terms of performance with trendy NoSQL technologies while providing extra functionality, and resolves impedance mismatch issues based on the Object Oriented paradigm for data representation. (C) 2017 Elsevier Inc. All rights reserved.
Data sharing and especially enabling third parties to build new services using large amounts of shared data is clearly a trend for the future and a main driver for innovation. However, sharing data is a challenging and involved process today: The owner of the data wants to maintain full and immediate control on what can be done with it, while users are interested in offering new services which may involve arbitrary and complex processing over large volumes of data. Currently, flexibility in building applications can only be achieved with public or non-sensitive data, which is released without restrictions. In contrast, if the data provider wants to impose conditions on how data is used, access to data is centralized and only predefined functions are provided to the users. We advocate for an alternative that takes the best of both worlds: distributing control on data among the data itself to provide flexibility to consumers. To this end, we exploit the well-known concept of object, an abstraction that couples data and code, and make it act and react according to the circumstances.
Current Data as a Service solutions present a lack of flexibility in terms of allowing users to customize the underlying data models by including new concepts or functionalities. Data providers either publish global APIs to make data available, or "sell" and transfer data to clients so they can do whatever they want with it. Thereby, collaboration and B2B becomes limited and sometimes is not even feasible. Our technology implements the necessary mechanisms for data providers to enable their clients to enrich data models both with additional concepts and with new methods that can be executed and, in turn, published as new services.
T. Cortes合作论文数Computer Architecture Department (DAC)
Universitat Polit??cnica de Catalunya (UPC)4