
This paper presents a novel approach to writing TOSCA templates for application reusability and portability in a modular auto-scaling and orchestration framework (MiCADO). The approach defines cloud resources as well as application containers in a flexible and generic way, and allows for those definitions to be extended with specific properties related to a desired container orchestrator chosen at deployment time. The approach is demonstrated in a proof-of-concept where only a minor change was required to a previously used application template in order to achieve the successful deployment and lifecycle management of the popular web authoring tool Wordpress on a new realization of the MiCADO framework featuring a different container orchestrator.
Researchers and scientists use aggregations of data from a diverse combination of sources, including partners, open data providers and commercial data suppliers. As the complexity of such data ecosystems increases, and in turn leads to the generation of new reusable assets, it becomes ever more difficult to track data usage, and to maintain a clear view on where data in a system has originated and makes onward contributions. Reliable traceability on data usage is needed for accountability, both in demonstrating the right to use data, and having assurance that the data is as it is claimed to be. Society is demanding more accountability in data-driven and artificial intelligence systems deployed and used commercially and in the public sector. This paper introduces the conceptual design of a model for data traceability based on a Bill of Materials scheme, widely used for supply chain traceability in manufacturing industries, and presents details of the architecture and implementation of a gateway built upon the model. Use of the gateway is illustrated through a case study, which demonstrates how data and artifacts used in an experiment would be defined and instantiated to achieve the desired traceability goals, and how blockchain technology can facilitate accurate recordings of transactions between contributors.
nanoHUB annually serves 17,000+ registered users with over 1 million simulations. In the past, we have used data analytics to demonstrate that nanoHUB can be a powerful scientific knowledge sharing platform. We used retrospective data analytics to show how simulation tools were used in structured education and how simulation tools were used in novel research. With the use of such retrospective analytics, we have made strategic decisions in terms of tool and content developments and justified continued nanoHUB investments by the US National Science Foundation (NSF). As we migrate towards a sustainable nanoHUB we must embrace similar processes pursued by in similar platforms such as Uber or AirBnB: we need to create actionable data analytics that can rapidly support user experience and help grow the supply in the two-sided market platform – we need to improve the experience of providers as well as end-users. This paper describes some aspects on how we pursue user behavior analysis inside the virtual worlds of nanotechnology simulation tools. From such user behavior we plan to derive actionable analytics that influence user behaviors as they interact with nanoHUB. Keywords— nanoHUB; HUBzero; science gateways; user behavior; analytics; cluster; meander; education INTRODUCTION AND BACKGROUND nanoHUB is a scientific knowledge platform that has enabled over 3,500 researchers and educators to share 500+ research simulation tools and models as well as 6,000+ lectures and tutorials globally through a novel cyberinfrastructure. nanoHUB annually serves 17,000+ registered users with over 1 million simulations in an end-to-end user-oriented scientific computing cloud. Over 1.5 million visitors access the openly available web content items annually. These might be considered impressive summative numbers, but they do not address if the site has any impact or what these users are doing. Understanding these numbers requires some background on the original intentions and cyberinfrastructure developments around nanoHUB. Fundamental issues raised by peerreviewers were the perceived ability of a University project to provide a stable, national-level infrastructure, provide support for the offered services, and provide compute cycles for an ever-growing user base. From the very beginning in 1996 [1], the predecessor to nanoHUB called Purdue Network Computing Hub (PUNCH) was created to enable researchers to share their code without re-writes through novel web interfaces with end-users in education and research. PUNCH was so novel that even the web-server had to be created within the team. By 2004 the standard web-form-interfaces were antiquated and did not inspire the interactive exploration of simulation results with rapid “What If?” questions that users might have. Users had to download their simulation data to manipulate them in a form where they can be truly used. nanoHUB was not an end-toend usage platform. It became clear that the system had to be revamped to enable the hosting of user-friendly engineeringuse inspired interactive applications. Such interactive sessions had to be hosted in a reliable, scalable middleware that was running in production mode, not as a research paper demonstration. 3D dataset exploration had to be supported on remote, dedicated GPUs that deliver the results to end users. RAPPTURE, the Rapid APPlication infrastrucTURE toolkit [2] enabled researchers, who typically did not have any graphical user interfaces to their codes to describe the input and outputs of their codes in XML and to generate a GUI. New middleware [3] enabled 1,000+ users to be hosted simultaneously on a moderate cluster of about 20 compute nodes. A novel remote GPU-based visualization system [4] supported hundreds of simultaneous sessions. nanoHUB established the first community accounts on TeraGrid and OSG which would execute heavy-lifting nanoHUB simulation jobs completely transparently on behalf of users who had no accounts on these grid platforms [5]. We developed processes [6] to continually test the reliability of these remote grid services to ensure smooth user services. For application support we developed policies and operational infrastructure that enabled tool contributors to support and improve their tools through question & answer forums and through wishlists. As this novel infrastructure emerged in 2005 we observed rapid growth in the simulation user base from the historical numbers of 500 annual users to over 10,000 in a few years. As 11th International Workshop on Science Gateways (IWSG 2019), 12-14 June 2019 questions of technical feasibility were addressed new questions as to actual and potential impact emerged. Early-on our peer reviewers raised fundamental questions whether such research-based simulation tools could be used by other researchers at all and if these tools could be used in education without specific customizations. The nanoHUB team developed analytics that documented nanoHUB use research through reference and citation searches in the scientific literature. Today we can document over 2,200 papers that cite nanoHUB and we keep track of the used resources and tools, to provide attribution to the published tools. When we showed the first 200 formal citations our peers remained unconvinced that this could be good research. We then began to track secondary citations, which today sum to over 30,000 resulting in an h-index of 82. Our peers had a similarly strong opinion that research tools could not be used in education. We therefore developed novel clustering algorithms [7] that documented systematic used of simulation tools in formal education settings. Today we can show that over 35,000 students in over 1,800 classes at over 180 institutions have used nanoHUB in formalized education settings. We could also measure the time-to-adoption between tool publication and first-time systematic use in a classroom. The median time was determined to be less than 6 months. From the analysis of research use and education use we can begin to qualify the attributes of the underlying simulation tools. We found significant use in education and in research for many of the nanoHUB tools. These research and education impact studies are documented in detail in Nature Nanotechnology [8]. We used retrospective data analytics to show how simulation tools were used in structured education and how simulation tools were used in novel research. We showed that the transition from research tool publication to adoption in the classroom is happening rapidly in typically less than six months and demonstrated through longitudinal data how research tools migrate into education. With the use of these retrospective analytics, we have made strategic decisions in terms of tool and content developments and justified continued investments by NSF into nanoHUB. As we migrate towards a sustainable nanoHUB we must embrace similar processes pursued by in similar platforms such as Uber or AirBnB: we need to create actionable data analytics that can rapidly support user experience and help grow the supply in the two-sided market platform – we need to improve the experience of providers as well as end-users. II RESEARCH QUESTIONS Beyond raw numbers of users and simulations, we have over the years continued to ask ourselves: How do users behave in their virtual world of a simulation tool? More specifically: How do they “travel” through the design/exploration world? How many individual simulations do they run within one session? How many parameters do users change? How different do researchers, classroom users, and selfstudy users behave? How different do different classes behave? Does different class instruction material / scaffolding make a difference? Can we provide feedback to instructors on their classrooms? Given certain usage patterns inside the tool: Can we improve the tools and provide feedback to the developers? There are a variety of different requirements that need to be met to address some of these questions in a scalable infrastructure such as: Storage/availability of individual simulation runs within user sessions A data description language that is shared across different tools A large set of simulation runs and participants Other user data such as classroom participation, or researcher identification, geolocation, etc. In the next Sections we describe some of our first results that begin to address some of these questions. For our initial study presented here we focus on the user behavior for PN Junction Lab [9] which is consistently one of the top 10 nanoHUB tools [10] within any year. Despite our codename pntoy the tool is powered by an industrial strength semiconductor device modeling tool called PADRE [11]. Instead of learning the complex PADRE input language that involves gridding, geometry, material and environmental specifications, users can easily ask “What if?” questions in a toy-like fashion. II SEARCHERS AND WILDCATTERS RAPPTURE provides a rather generic description of simulation tool inputs and outputs. Over 90% of the 500+ nanoHUB simulation tools utilize RAPPTURE as their data description language. With existing simulation logs we can now begin to study the user behavior inside simulation tools. Each simulation tool typically consists of 10 to 50 parameters that are exposed to the users. Most of these parameters are freeform numbers such as length, doping, effective mass, dielectric constant, temperature etc. with their specific units, while there is also a significant set of discrete options such as model or geometry choices. Assuming that each parameter might have just 10 reasonable choices, then each tool spans a configurational design space of at least 1010 to 1050. The dimensionality of these tools is clearly too large to be intuitively understood. We developed a visualization methodology [12] to flatten an N-dimensional space into 2 dimensions. Figure 1 shows the conceptual mapping and shows two significantly differen
Reference architectures for big data and machine learning include not only interconnected building blocks but important considerations (among others) for scalability, manageability and usability issues as well. Leveraging on such reference architectures, the automated deployment of distributed toolsets and frameworks on various clouds is still challenging due to the diversity of technologies and protocols. The paper focuses particularly on the widespread Apache Spark cluster with Jupyter as the particularly addressed framework, and the Occopus cloud‐agnostic orchestrator tool for automating its deployment and maintenance stages. The presented approach has been demonstrated and validated with a new, promising text classification application on the Hungarian academic research infrastructure, the OpenStack‐based MTA Cloud. The paper explains the concept, the applied components, and illustrates their usage with real use‐case measurements.
The International Technology Alliance in Distributed Analytics and Information Science (DAIS-ITA) is a research program conducting fundamental research across a consortium of organizations from academia, industry and government in the US and UK. A key output of the program are the academic publications, and from these a rich social and topical network between authors and organizations emerges. To capture and convey this we have created the publicly available Science Library as a user-centric, interactive portal. The Science Library is a Controlled Natural Language (CNL) driven research gateway, which allows the community to explore and query the publications and networks through an open, webbased application. The data is represented through interactive visualizations along with the ability for users to query the model using a natural language conversational interface. The CNL based approach models the data through concepts, properties and relationships which are defined using CNL and are therefore both human readable and directly machine processable. This captures complex semantics in a simple format and enables nontechnical users to participate in the continuous improvement of the data model behind the Science Library application. This paper presents the features, implementation and design considerations of the Science Library and the underlying CNL implementation. Keywords—Controlled English; Human-Computer Interaction; Science Gateways; Science Library; Visualization.
When executing scientific workflows in a distributed environment, anomalies of the workflow behavior are often caused by a mixture of different issues, e.g., careless design of the workflow logic, buggy workflow components, unexpected performance bottlenecks or resource failure at the underlying infrastructure. The provenance information only defines data evolution at the workflow level, which does not have an explicit connection with the system logs provided by the underlying infrastructure. Analyzing provenance information and apposite system metrics requires expertise and a considerable amount of manual effort. Moreover, it is often time-consuming to aggregate this information and correlate events occurring at different levels in the infrastructure. In this paper, we propose an architecture to automate the integration among the workflow provenance information with the performance information collected from infrastructure nodes running workflow tasks. Our architecture enables workflow developers or domain scientists to effectively browse workflow execution information together with the system metrics, and analyze contextual information for possible anomalies.
— The ESGF Dashboard is a key component of the Earth System Grid Federation (ESGF). It provides a distributed and scalable software infrastructure responsible for capturing a comprehensive set of data usage and data archive metrics both at the single site and federation level. The data usage information is related to the number of downloads and successful downloads and the number of distinct downloaded files, grouped by variable, model, experiment, etc. On the other hand, the data archive information is related to the total number of published datasets, total data volume and CMIP5 models and modelling institutes. All the above metrics relate to both cross and specific projects that are very notable in the climate community, such as, CMIP5, CMIP6, Obs4MIPs and CORDEX. From a Science Gateway perspective, the ESGF Dashboard presents the collected metrics through its Community Gateway (ESGF Dashboard User Interface). 1
—The importance of software and web services is changing. So it sufficed to state the software used during scientific data processing, when publishing results in journals. This also applies to web services, which can be methods used in scientific work as well as results of this work. Awareness for the evanescence of tools and services, which hinders the reproducibility of published work, increases. Similar to data repositories, which archive datasets and make them citable, a service for software artifacts would enhance sustainability of science. So we present the CiTAR (Citing & Archiving Research) service in this paper, which enables re-searches to preserve computational environments and make them citable. In contrast to pure data repositories, CiTAR guarantees the executability of archived environments by providing generic runtimes.
—Applications that make use of Internet of Things (IoT) capture an enormous amount of raw data from sensors and actuators, which is frequently transmitted towards the cloud data centres for processing and analysis. However, due to varying and unpredictable data generation rates and network latency, sending the data towards a cloud data centre can lead to a performance bottleneck. With the emergence of Fog and Edge computing hosted microservices, data processing could be moved towards the network edge. We propose a novel Pareto-based approach that makes use of a multi-criteria bin packing optimisation for efficient and optimal distributed deployment of microservices – along edge, fog/cloudlet and cloud tiers. This optimisation takes account of non-functional requirements, such as operational cost, compute resource utilisation, service availability, response time, latency and similar. The results show that the present approach provides an optimal and sustainable consumption of compute resources and improves Quality of Service of the application during its runtime. The approach can also be integrated into software engineering workbenches for the creation and deployment of cloud-native applications, enabling partitioning of an application across the multiple infrastructure tiers outlined above.
The enhancements in IT solutions and the open science movement are injecting changes in the practices dealing with data collection, collation, processing and analytics, and publishing in all the domains, including agri-food. However, in implementing these changes one of the major issues faced by the agri-food researchers is the fragmentation of the “assets” to be exploited when performing research tasks, e.g. data of interest are heterogeneous and scattered across several repositories, the tools modellers rely on are diverse and often make use of limited computing capacity, the publishing practices are various and rarely aim at making available the “whole story” with datasets, processes, workflows. This paper presents the AGINFRA PLUS endeavour to overcome these limitations by providing researchers in three designated communities with Virtual Research Environments facilitating the use of the “assets” of interest and promote collaboration. Keywords—Virtual research environment; Agroclimatic modeling; Food safety risks assessment; Food security
Sustainability of science gateways and continuous funding for their developer teams is a major concern that many projects face. The HUBzero® project and its science gateway framework have evolved to be self-sustained via diversifying funding resources, extending outreach measures to further communities and targeting sustainability from different angles in concrete instances. nanoHUB, PURR and OneSciencePlace are examples of how the HUBzero® team and platform build science gateways and take their specific services into account to address sustainability beyond securing funding and outreach activities. They have been integrating additional procedures and concepts for sustainability: nanoHUB invests into reliability of the over 500 simulation tools and high quality lecture and tutorial content to keep the trust of the large community with over 1.5 million users; PURR developed policies and methods for preserving research output in a sensible and sustainable way and OneSciencePlace addresses the concern of projects that have a lack of continuous funding for maintaining a science gateway by offering a solution to keep science gateways available to their communities. The paper goes into detail for measures for sustainability for HUBzero® and especially for nanoHUB, PURR and OneSciencePlace. Keywords—HUBzero®; nanoHUB; PURR; OneSciencePlace; science gateways; sustainability; research content; research frameworks INTRODUCTION The importance of sustainability of research software in general and thus of science gateways as subgroup has been recognized by various researchers, funding bodies and organizations evident in funded projects such as the Science Gateways Community Institute (SGCI) [1] and the UK Software Sustainability Institute (SSI) [2] as well as initiatives such as the Research Software Engineer (RSE) Association in the UK [3] and the RSE Communities in Germany and in the US [4]. There are many definitions for sustainability of software available, i.e. SSI states in their manifesto “Sustainability means that the software you use today will be available and continue to be improved and supported in the future.” [5]. C.C. Venters et al. [6] define software sustainability as a composite, nonfunctional requirement which is “a measure of a systems extensibility, interoperability, maintainability, portability, reusability, scalability, and usability”. Most definitions consider maintainability, fulfilling its purpose over time and surviving uncertainty as essential characteristics for sustainable software. Achieving sustainability based on these three characteristics requires continuous effort and a variety of actions by a project and/or group developing software or a science gateway, respectively. In the remainder of the paper we focus on science gateways and the science gateway landscape to define the variations of actionable items. Science gateways are created for specific communities and are embedded in the science gateway landscape with similar Copyright © 2021 for this paper by its authors. Use permitted under Creative Commons License Attribution 4.0 International (CC BY 4.0). 11 International Workshop on Science Gateways (IWSG 2019), 12-14 June 2019 and/or competing science gateways. Existing mature frameworks and APIs such as HUBzero® [7], Galaxy [8] and the Agave Platform [9] allow for creating science gateways more efficiently and support developers on focusing on a specific gateway while offering features such as connecting to distributed computing out of the box. The services of science gateways vary from offering simulations tools to data collections to computational workflows with different requirements on the user interface and the underlying research infrastructure. The services and the target communities of various science gateways might be very different from each other, but actionable items can be determined in a similar way. We distinguish four key variations for actionable items: 1. a technical area, 2. a community area, 3. a science gateway landscape area and 4. a stakeholder or funding area. Examples for action items for the areas include 1. Use of well-defined software engineering practices to support extensibility, interoperability, maintainability, portability, reusability, scalability, usability, reliability and security 2. Support measures and extension of features and/or technologies in a science gateway driven by the needs of a community 3. Outreach and expansion to new communities 4. Diversifying funding The areas are not isolated from each other but influence and overlap with each other. For example, after analyzing the science gateway landscape and reaching out to a new promising community, the development of a novel science gateway necessitates the technical implementation based on gathering requirements from the community. The definition of concrete actionable items is a mixture of performing analyses and tasks in all four areas. The HUBzero® project has been achieving sustainability for its science gateway framework and the team via multiple measures. The science gateway framework started in 1996 as online platform PUNCH [10] for nanoelectronic research and teaching. It was horizontally expanded for more simulation tools to nanoHUB [11] and vertically to HUBzero® to serve as generic science gateway framework for a variety of communities. These expansions led to novel developments on technical side. Reaching out to communities includes the participation in conferences and workshops as presenters and/or sponsor, social media such as Twitter and a yearly event that offers the opportunity to clients to interact with the HUBzero® team face-to-face. The financial independence of the developer team from funding provided by the Purdue University was a major step. It has been attained by diversifying funding resources with participation in grants, offering hosting services and offering memberships in the HUBzero® foundation, which allow supporting instances with a limited number of development time and consultancy for usability and community outreach measures specifically for the instance. Examples for action items in nanoHUB, PURR, and OneSciencePlace are described in detail in the sections III – V after presenting the background for activities to reach sustainability for science gateways.
Data processing in data intensive scientific fields like bioinformatics is automated to a great extent. Among others, automation is achieved with workflow engines that execute an explicitly stated sequence of computations. Scientists can use these workflows through science gateways or they develop them by their own. In both cases they may have to preprocess their raw data and also may want to further process the workflow output. The scientist has to take care about provenance of the whole data processing pipeline. This is not a trivial task due to the diverse set of computational tools and environments used during the transformation of raw data to the final results. Thus we created a metadata schema to provide provenance for data processing pipelines and implemented a tool that creates this metadata during the execution of typical scientific computations. Provenance, Reproducibility, Workflows, Science Gateways—
Many applications process large quantities of data that takes significant time and requires big amount of computational resources. Optimising the execution of such applications in a cloud computing environment by keeping costs at minimum but still completing the task by a set deadline has paramount importance. As container-based technologies are becoming more widespread, support for job-queuing and auto-scaling in such environments is becoming important. Current container technologies, such as Docker or Kubernetes provide limited support in this area. This paper presents JQueuer and CAutoScaler, a couple of cloud-independent solutions that offer job-queuing and automated scalability at the level of containers. Applying these solutions leads to more cloud-aware applications providing transparent auto-scaling for end-users and optimising execution time and costs. Business and science gateways will benefit from using an orchestrator combined with JQueuer and CAutoScaler since it will provide the layers needed to auto-scale the containers and to batch/sweep the jobs from a queue depending on a userdefined policy. Keywords—cloud computing, container technologies, Docker Swarm, JQueuer, autoscaler, MiCADO
—Science gateways have been developed over the last twenty years and have grown into a large community of practice, as evidenced by international workshops and conferences. Because of the diversity of approaches to creating science gateways and the always changing landscape of technologies, the community lacks a common definition for the term “science gateway” itself and common terminology for describing the common components of a gateway architecture. Instead, a wide range of definitions and understandings exist and are used in different communities; this is evident, for example, in discussions whether science gateways are the same as virtual research environments. This paper attempts to address these issues by focusing on how science gateways support scientific research and considering the consequences on cyberinfrastructure. Keywords—science gateways, cyberinfrastructure
This paper presents IoT-Hub a new scalable, elastic, efficient, and portable Internet of Things (IoT) data-platform based on microservices for monitoring and analysing largescale sensor data in real-time. IoT-Hub allows us to collect, process, and store large amounts of data from multiple sensors in distributed locations—which could be deployed as a backend for Virtual Research Environments (VRE) or Science Gateways. In the proposed data-platform, all required software, which involves a variety of state-of-the-art open-source middleware, is packed into containers and deployed in a cloud environment. As a result, the engineering and computational time and costs for deployment and execution is significantly reduced. Keywords—IoT, Science Gateway, Virtual Research Environment, Data-Frameworks, Containers, Data Science, Microservices
With an estimated 20 billion connected devices by 2020 generating enormous amounts of data, more data-centric ways of working are needed to cope with the dynamic load and reconfigurability of on-demand computing. There is a growing range of complex, specialised means by which this flexibility can be achieved, e.g. Software-defined networking (SDN). Specification of Quality of Service (QoS) constraints for time-critical characteristics, such as network availability and bandwidth, will be needed, in the same way that compute requirements can be specified in today's infrastructures. This is the motivation for SWITCH -- an EU-funded H2020 project addressing the entire lifecycle of time-critical, self-adaptive cloud applications by developing new middleware and tools for interactive specification of such applications. This paper presents a user-facing perspective on SWITCH by discussing the SWITCH Interactive Development Environment (SIDE) Workbench. SIDE provides a programmable and dynamic graphical modeling environment for cloud applications that ensures efficient use of compute and network resources while satisfying time-critical QoS requirements. SIDE enables a user to specify the software components, properties and requirements, QoS parameters, machine requirements and their composition into a fully operational, multi-tier cloud application. In order to enable SIDE to represent the software and infrastructure constraints and to communicate them to other SWITCH components, we have defined a co-programming model using TOSCA that is capable of representing the application's state during the entire lifecycle of the application. We show how the SIDE Web GUI, along with TOSCA and the other subsystems, can support three use cases and provide a walk-through of one of these use cases to illustrate the power of such an approach.