Today, Big Data, IoT and Analytics are driving and making the differences in key performing top organizations. The interplay of these three areas can be instrumental for the future development of research, complex systems and enterprises. IoT will be estimated to rise to billions of devices connected by 2020 [1]. This has huge implications for research, businesses and future activities for mankind. The development of this vision is pivotal to support and foster IoTBDS (iotbds.org). Sensors are used to extract an unprecedented amount of data, which can be filtered and processed in the IoT by machine networks, automated analysis, tools and systems. This can become a Big Data layer abstraction from the captured data if the process is well-organized. Hereby, complex information systems can obtain the benefits of collecting, processing and analyzing highly valuable data. This is one of the most important scopes of COMPLEXIS (complexis.org). Therefore, starting from the pillars of IoT, Analytics andBig Data, we can build complex information architectures by using tools like social media, predictive modeling, insight analysis and sentiment analysis. Eventually, we can build three additional layers: complex data, information and knowledge, and offer service related to these three valued-added layers. The application fields are widespread in fields like social networking, financial services, or biological research. This special issue (SI) call is focused on contributions that explore and demonstrate related areas, approaches and recommendations for both IoTBDS and COMPLEXIS. In particular, research work and projects that offer solutions to the underlying prob-
The paper describes conceptual and technological principles of the human-computer cloud, that allows to deploy and run human-based applications. It also presents two ways to build decision support services on top of the proposed cloud environment for problems where workflows are not (or cannot be) defined in advance. The first extension is represented by a decision support service leveraging task ontology to build the missing workflow, the second utilizes the idea of human-machine collective intelligence environment, where the workflow is defined in the process of a (sometimes, guided) collaboration of the participants.
: With usage of big data and development of advanced technology, smart wearable devices have come to market with multifunction. The wearables include smart watches, smart bands, smart glasses and so on. The big data and data analysis can provide more benefits for users and companies. As for users, smart wearables can help to provide a more convenient and healthier lifestyle. Besides, the big data can provide more support for decision making of companies and may encourage creativity. However, ethical problems also appear with the development of smart wearables. First, vulnerabilities of system can be found, which may pose potential threat to users. Second, misused data will also cause bad influence and the companies should be more transparent on the usage of personal data. Third, some smart wearable companies take priority of multifunction of devices rather than data safety. Based on these problems, different entities should take on responsibilities. Users should improve their awareness and knowledge to protect their personal data. As for companies, some principles of personal data released by OECD can be applied when companies deal with personal information. When it comes to society, supervisions of the industry and support for advanced technology should be carried out.
Cloud Computing has reached maturity in software architecture, methods and technologies. The research and development work has moved from the context of exploration and formalization to the application. Nowadays, Cloud Computing offers unprecedented possibilities in a wide range of new computation areas, becoming a key topic in the academia and industry, not only contributing to the critical questions of the How but also opening new scientific questions needing the foundation for the What and Why matters.
Geo-distributed data centres (DCs) that recently established due to the increasing use of on-demand cloud services have increasingly attracted cloud providers as well as researchers attention. Energy and data transmission cost are two significant problems that degrades the cloud provider net profit. However, increasing awareness about CO2 emissions leads to a greater demand for cleaner products and services. Most of the proposed approaches tackle these problems separately. This paper proposes green approach for joint management of virtual machine (VM) and data placement that results in less energy consumption, less CO2 emission, and less access latency towards large-scale cloud providers operational cost minimization. To advance the performance of the proposed model, a novel machine-learning model was constructed. Extensive simulation using synthetic and real data are conducted using the CloudSim simulator to validate the effectiveness of the proposed model. The promising results approve the efficacy of the CELA model compared to other competing models in reducing network latency, energy consumption, CO2 emission and total cloud provider operational cost.
The elasticity feature of cloud computing has been proved as pertinent for parallel applications, since users do not need to take care about the best choice for the number of processes/resources beforehand. To accomplish this, the most common approaches use threshold-based reactive elasticity or time-consuming proactive elasticity. However, both present at least one problem related to: the need of a previous user experience, lack on handling load peaks, completion of parameters or design for a specific infrastructure and workload setting. In this regard, we developed a hybrid elasticity service for parallel applications named SelfElastic. As parameterless model, SelfElastic presents a closed control loop elasticity architecture that adapts at runtime the values of lower and upper thresholds. Besides presenting SelfElastic, our purpose is to provide a comparison with our previous work on reactive elasticity called AutoElastic. The results present the SelfElastic’s lightweight feature, besides highlighting its performance competitiveness in terms of application time and cost metrics.
Data produced within marine and terrestrial biodiversity research projects that evaluate and monitor Good Environmental Status, have a high potential for use by stakeholders involved in environmental management. However, environmental data, especially in ecology, are not readily accessible to various users. The specific scientific goals and the logics of project organization and information gathering lead to a decentralized data distribution. In such a heterogeneous system across different organizations and data formats, it is difficult to efficiently harmonize the outputs. Few tools are available to assist. For instance standards and specific protocols can be applied to interconnect databases. Such semantic approaches greatly increase data interoperability. This communication present the recent results and the consortium IndexMEED (Indexing for Mining Ecological and Environmental Data) activity that aims to build new approaches to investigate complex research questions, and support the emergence of new scientific hypotheses based on graph theory Auber et al. 2014). Current developments in data mining based on graphs, as well as the potential for relevant contributions to environmental research, particularly about strategic decision-making, and new ways of organizing data will be presented (David et al. 2015). In particular, the consortium makes decisions on how i) to analyze heterogeneous distributed data spread throughout different databases combining molecular and habitat characteristics data [3], ii) to create matches and incorporate some approximations, iii) to identify statistical relationships between observed data and the emergence of contextual patterns using a calculation library and distributed calculation center at the European level, iv) to encourage openness and sharing data while complying with the general principles of FAIR (Findable, Accessible, Interoperable, Re-usable and citable) in order to enhance data value and their utilization. IndexMEED participants are now exploring the ability of two scientific communities (ecology sensu lato and computer sciences) to work together, using several studies cases. The ECOSCOPE project aims to meet the need to access structured and complementary omics-datasets to better understand biodiversity state and its dynamics. Indeed, the ECOSCOPE case study targets to visualize, through the graph approach, links between datasets and databases from genetics to ecosystems. Another case study, displaying anthropology fossils and omics on the same graph, will also be presented. DEVOTES (DEVelopment Of innovative Tools for understanding marine biodiversity and assessing good Environmental Status) and CIGESMED (Coralligenous based Indicators to evaluate and monitor the "Good Environmental Status" of the MEDiterranean coastal water) European projects, conducted by IMBE, are focused on photo quadrats, cartography and omics data of the marine hard bottom in order to discover context patterns helpful to build decision support system building. Study case “65 Millions d’observateurs” French project is testing AskOmics to provide a graph-based querying interface using RDF (Resource Description Framework) and SPARQL technologies. Scientific questions can be resolved by the new data mining approaches that offer new ways to investigate heterogeneous environmental data with graph mining (Munoz et al. 2017). The uses of data from biodiversity research demonstrate the prototype functionalities (David et al. 2016) and introduce new perspectives to analyze environmental and societal responses including decision-making at large scale, both at the information system level and the observing system level than at the observed system level.
Extisting biodiversity databases contain an abundance of information.To turn such information into knowledge, it is necessary to address several information-model issues.Biodiversity data are collected for various scientific objectives, often even without clear preliminary objectives, may follow different taxonomy standards and organization logic, and be held in multiple file formats and utilising a variety of database technologies.This paper presents a graph catalogue model for the metadata management of biodiversity databases.It explores the possible operation of data mining and visualization to guide the analysis of heterogeneous biodiversity data.In particular, we would propose contributions to the problems of (1) the analysis of heterogeneous distributed data found across different databases, (2) the identification of matches and approximations between data sets, and (3) the identificaton of relationships between various databases.This paper describes a proof of concept of an infrastructure testbed and its basic operations, presenting an evaluation of the resulting system in comparison with the ideal expectations of the ecologist.
A representative set of workflows found in bioinformatics pipelines must deal with large data sets. Most scientific workflows are defined as Direct Acyclic Graphs (DAGs). Despite DAGs are useful to understand dependence relationships, they do not provide any information about input, output and temporal data files. This information about the location of files of data intensive applications helps to avoid performance issues.This paper presents a multiworkflow store-aware scheduler in a cluster environment called Critical Path File Location (CPFL) policy where the access time to disk is more relevant than network, as an extension of the classical list scheduling policies. Our purpose is to find the best location of data files in a hierarchical storage system.The resulting algorithm is tested in an HPC cluster and in a simulated cluster scenario with bioinformatics synthetic workflows, and largely used benchmarks like Montage and Epigenomics. The resulting simulator is tuned and validated with the first test results from the real infrastructure. The evaluation of our proposal shows promising results up to 70% on benchmarks in real HPC clusters using 128 cores and up to 69% of makespan improvement on simulated 512 cores clusters with a deviation between 0.9% and 3% regarding the real HPC cluster. (C) 2017 The Authors. Published by Elsevier B.V.
The CLOSER 2018 proceedings focus on cloud computing and service science, cloud application portability, cloud computing architecture, cloud interoperability, business process management and web services, business services realized by IT services, cloud brokering, etc.
Data produced by biodiversity research projects that evaluate and monitor Good Environmental Status have a high potential for use by stakeholders involved in [marine] environmental management. The lack of specific scientific objectives, poor organizational logic, and a characteristically disorganized collection of information leads to a decentralized data distribution, hampering environmental research. In such a heterogeneous system across different organizations and data formats, it is difficult to efficiently harmonize the outputs. There are few tools available to assist. The task of the newly created consortium of IndexMeed is to index biodiversity data (and to provide an index of qualified existing open datasets) and make it possible to build graphs to assist in the analysis and development of new ways to mine data. Standards (including TDWG recommendations) and specific protocols can be applied to interconnect databases. Such semantic approaches greatly increase data interoperability. The aim of this poster is to present the 2016 IndexMed workshop results (https://indexmed2016.sciencesconf.org) and recent actions of the consortium (renamed IndexMeed - Indexing for Mining Ecological and Environmental Data): new approaches to investigate complex research questions and support the emergence of new scientific hypotheses. With one day of plenary sessions and two days of practical workshops, this event was dedicated to the sharing of experience and expertise, the acquisition of practical methods to construct graphs and value data through metadata and ”data papers”. Recent developments in data mining based on graphs, the …
Data produced by biodiversity research projects that evaluate and monitor Good Environmental Status have a high potential for use by stakeholders involved in [marine] environmental management. The lack of specific scientific objectives, poor organizational logic, and a characteristically disorganized collection of information leads to a decentralized data distribution, hampering environmental research. In such a heterogeneous system across different organizations and data formats, it is difficult to efficiently harmonize the outputs. There are few tools available to assist. The task of the newly created consortium of IndexMeed is to index biodiversity data (and to provide an index of qualified existing open datasets) and make it possible to build graphs to assist in the analysis and development of new ways to mine data. Standards (including TDWG recommendations) and specific protocols can be applied to interconnect databases. Such semantic approaches greatly increase data interoperability. The aim of this poster is to present the 2016 IndexMed workshop results (https://indexmed2016.sciencesconf.org) and recent actions of the consortium (renamed IndexMeed - Indexing for Mining Ecological and Environmental Data): new approaches to investigate complex research questions and support the emergence of new scientific hypotheses. With one day of plenary sessions and two days of practical workshops, this event was dedicated to the sharing of experience and expertise, the acquisition of practical methods to construct graphs and value data through metadata and ”data papers”. Recent developments in data mining based on graphs, the potential for important contributions to environmental research, particularly about strategic decision-making, and new ways of organizing data were also discussed at the workshop. In particular, this workshop promoted decisions on how (i) to analyze heterogeneous distributed data spread in different databases, (ii) to create matches and incorporate some approximations, (iii) to identify statistical relationships between observed data and the emergence of contextual patterns, and (iv) to encourage openness and the sharing of data, in order to value data and their utilization. The IndexMeed project participants are now exploring the ability of two scientific communities (ecology sensu lato and computer sciences) to work together. The uses of data from biodiversity research demonstrate the prototype functionalities and introduce new perspectives to analyze environmental and societal responses including decision-making. Output of the seminar lists scientific questions that can be resolved by the new data mining approaches and proposes new ways to investigate heterogeneous environmental data with graph mining.
Victor Fernandez Albor1, Marcos Seco1, Victor Mendez Munoz2, Tomas Fernandez Pena3, Juan Saborido Silva1 and Ricardo Graciani Diaz4 1 Physics department, Santiago de Compostela University Av Ciencias sn, Santiago de Compostela, Spain E-mail: {victormanuel.fernandez,marcos.seco,juan.saborido}@usc.es 2 Computer Architecture and Operating Systems (CAOS),Universidad Autonoma de Barcelona E-mail: victor.mendez@uab.es 3 Research Center in Information Technologies (CiTIUS), Santiago de Compostela University Av Ciencias sn, Santiago de Compostela, Spain E-mail: tf.pena@usc.es 4 Departamento de Estructura y Constituyentes de la Materia,Barcelona University E-mail: Ricardo.graciani@ecm.ub.es
Data produced by biodiversity research projects that evaluate and monitor ‘Good Environmental Status’ have a high potential for use by stakeholders involved in [marine] environmental management. The lack of specific scientific objectives, poor organizational logic, and a characteristically disorganized collection of information leads to a decentralized data distribution, hampering environmental research. In such a heterogeneous system across different organizations and data formats, it is difficult to efficiently harmonize the outputs. There are few tools available to assist.
One of the benefits of OCCI stems from simplifying the life of developers aiming to integrate multiple cloud managers. It provides them with a single protocol to abstract the differences between cloud service implementations used on sites run by different providers. This comes particularly handy in federated clouds, such as the EGI Federated Cloud Platform, which bring together providers who run different cloud management platforms on their sites: most notably OpenNebula, OpenStack, or Synnefo. Thanks to the wealth of approaches and tools now available to developers of virtual resource management solutions, different paths may be chosen, ranging from a small-scale use of an existing command line client or single-user graphical interface, to libraries ready for integration with large workload management frameworks and job submission portals relied on by large science communities across Europe. From lone wolves in the long-tail of science to virtual organizations counting thousands of users, OCCI simplifies their life through standardization, unification, and simplification. Hence cloud applications based on OCCI can focus on user specifications, saving cost and reaching a robust development life-cycle. To demonstrate this, the paper shows several EGI Federated Cloud experiences, demonstrating the possible approaches and design principles.
The OCCI standard has been in use for half a decade, with multiple server-side and client-side implementations in use across the world in heterogeneous cloud environments. The real-world experience uncovered certain peculiarities or even deficiencies which had to be addressed either with workarounds, agreements between implementers, or with updates to the standard. This article sums up implementersâ experience with the standard, evaluating its maturity and discussing in detail some of the issues arising during development and use of OCCI-compliant interfaces. It shows how particular issues were tackled at different levels, and what the motivation was for some of the most recent changes introduced in the OCCI 1.2 specification.
The internet of things (IoT) is potentially interconnecting unprecedented amounts of raw data, opening countless possibilities by two main logical layers: become data into information, then turn information into knowledge. The former is about filtering the significance in the appropriate format, while the latter provides emerging categories of the whole domain. This path of the data is a bottom-up flow. On the other hand, the path of the process is a top-down flow, starting at the strategic level of business and scientific institutions. Today, the path of the process treasures a sizeable amount of well-known methods, architectures and technologies: the so called Big Data. On the top, Big Data analytics aims variable association (e-commerce), data mining (predictive behaviour) or clustering (marketing segmentation). Digging the Big Data architecture there are a myriad of enabling technologies for data taking, storage and management. However the strategic aim is to enhance knowledge with the appropriate information, which does need of data, but not vice versa. In the way, the magnitude of upcoming data from the IoT will disrupt the data centres. To cope with the extreme scale is a matter of moving the computing services towards the data sources. This paper explores the possibilities of providing many of the IoT services which are currently hosted in monolithic cloud centres, moving these computing services into nano data centres (NaDa). Particularly, data-information processes, which usually are performing at sub-problem domains. NaDa distributes computing power over the already present machines of the IP provides, like gateways or wireless routers to overcome latency, storage cost and alleviate transmissions. Large scale questionnaires have been taken for 300 IT professionals to validate the points of view for IoT adoption. Considering IoT is by definition connected to the Internet, NaDa may be used to implement the logical low layer architecture of the services. Obviously, such distributed NaDa send results on a logical high layer in charge of the information-knowledge turn. This layer requires the whole picture of the domain to enable those processes of Big Data analytics on the top. (C) 2016 Elsevier Ltd. All rights reserved.
on Cloud Computing and Services Science), held in