Today’s problems require a plethora of analytics tasks to be conducted to tackle state-of-the-art computational challenges posed in society impacting many areas including health care, automotive, banking, natural language processing, image detection, and many more data analytics-related tasks. Sharing existing analytics functions allows reuse and reduces overall effort. However, integrating deployment frameworks in the age of cloud computing are often out of reach for domain experts. Simple frameworks are needed that allow even non-experts to deploy and host services in the cloud. To avoid vendor lock-in, we require a generalized composable analytics service framework that allows users to integrate their services and those offered in clouds, not only by one, but by many cloud compute and service providers.We report on work that we conducted to provide a service integration framework for composing generalized analytics frame-works on multi-cloud providers that we call our Generalized AI Service (GAS) Generator. We demonstrate the framework’s usability by showcasing useful analytics workflows on various cloud providers, including AWS, Azure, and Google, and edge computing IoT devices. The examples are based on Scikit learn so they can be used in educational settings, replicated, and expanded upon. Benchmarks are used to compare the different services and showcase general replicability.
Mass spectrometers are tools in the field of proteomics. Each mass spectrometer manufacturer uses its individual data format and software tools, making the creation of additional software tools and databases difficult and incompatible with one another. The Mass Spectrum I/O Project (MSIOP) addresses this problem, allowing for storage and analysis of mass spectrometer data from multiple manufacturers across various platforms while providing a framework upon which mass spectrometry software tools can be constructed.
Designing low latency applications that can process large volumes data with higher efficiency is a challenging problem. With the limited time to process data, usage of online algorithms are becoming important in the big-data applications. Stream processing is a well-known area that has been studied for a long time. In this research, our objective is to use state of the art big-data analytic engines to implement online algorithms and compare the strengths and weaknesses in each system. We use a streaming version of Support Vector Machines (SVM) and KMeans to do the analysis. Apache Flink, Apache Storm and Twister2 streaming frameworks are used to implement these algorithms. Our study focuses on the efficiency of online training of these algorithms and the results show higher performance in Twister2 framework for these algorithms.
Project CH-818664, KVM: This paper reports on the experience that we gained form using chameleon cloud as part of a course on Big Data and Open Source Software. The course had 62 registered users and used 91% of its allocation with 18195 SUs. The course studies software used in many commercial activities related to Big Data. The backdrop for course contains more than 370 software subsystems from which the students can select. Chameleon cloud was used by the students to conduct a course project that included the creation of a reproducible big data computational infrastructure with DevOps, as well as the execution of an application in this infrastructure. A rich variety of 31 projects were conducted. The result of the project with its 31 projects is published in an online class report. We report on success and challenges of chameleon cloud as used in classes. Chameleon Cloud provided the compute resources to make this class possible.
The Technology Audit Service has developed, XDMoD, a resource management tool. This paper utilizes XDMoD and the XDMoD data warehouse that it draws from to provide a broad overview of several aspects of XSEDE users and their usage. Some important trends include: 1) in spite of a large yearly turnover, there is a core of users persisting over many years, 2) user job submission has changed from primarily faculty members to students and postdocs, 3) increases in usage in Molecular Biosciences and Materials Research has outstripped that of other fields of science, 4) the distribution of user external funding is bimodal with one group having a large ratio of external funding to internal XSEDE funding (ie, CPU cycles) and a second group having a small ratio of external to internal (CPU cycle) funding, 5) user job efficiency is also bimodal with a group of presumably new users running mainly small inefficient jobs and another group of users running larger more efficient jobs, 6) finally, based on an analysis of citations of published papers, the scientific impact of XSEDE coupled with the service providers is demonstrated in the statistically significant advantage it provides to the research of its users.
Without doubt the analysis of data from Polar Regions is an important aspect of identifying environmental impact by humans. The calculation of automatic techniques for determining ice and snow layer boundaries in radar echograms remains a hard problem because of two reasons. First, the data volume and the associated calculations to retrieve results are large requiring considerable supercomputing time. Second, the data itself provides challenges based on the high degree of noise, the often faint layer boundaries, and confusing linear structures caused by signal reflections and clutter. Thus it is necessary to optimize existing and future algorithms while improving the performance, but also introduce new techniques that combine together weak image cues, reasoning explicitly about uncertainty in both the evidence and the resulting layer boundary estimates. In this paper we report on two important findings. First the performance improvement of the Multi-look time processor program, and second the introduction of a new analysis technique improving our image analysis.
Heterogeneous parallel system with multi processors and accelerators are becoming ubiquitous due to better cost-performance and energy-efficiency. These heterogeneous processor architectures have different instruction sets and are optimized for either task-latency or throughput purposes. Challenges occur in regard to programmability and performance when executing SPMD computations on heterogeneous architectures simultaneously. In order to meet these challenges, we implemented a MapReduce runtime system to co-process SPMD job on GPUs and CPUs on shared memory system. We are proposing a heterogeneous MapReduce programming interface for the developer and leverage the two-level scheduling approach in order to efficiently schedule tasks with heterogeneous granularities on the GPUs and CPUs. Experimental results of C-means clustering, matrix multiplication and word count indicate that using all CPU cores increase the GPU performance by 11.5%, 5.1%, and 41.9% respectively. Keyword: MapReduce, SPMD, GPU, CUDA, Multi-Level-
1.6 Performance increases much faster than performance per watt of energy consumed.................. 15 1.7 Possible energy to performance trade-off.......... 17
We present a workflow-based algorithm for identifying threads to an urban water management system. Through Grid computing we provide the necessary high-performance computing resources to deliver quickly solutions to the problem. We prototyped a new middleware called cyberaide, that enables easy access to Grid resources through portals or the command line. A workflow system is used to manage resources in fault tolerant fashion. In addition, we contrast the architecture with a Hadoop implementation. Resources from TeraGrid and FutureGrid are used to test the feasibility of using the toolkit for a scientific application.
FutureGrid provides novel computing capabilities that enable reproducible experiments while simultaneously supporting dynamic provisioning. This paper describes the FutureGrid experiment management framework to create and execute large scale scientific experiments for researchers around the globe. The experiments executed are performed by the various users of FutureGrid ranging from administrators to software developers and end users. The Experiment management framework will consist of software tools that record user and system actions to generate a reproducible set of tasks and resource configurations. Additionally, the experiment management framework can be used to share not only the experiment setup, but also performance information for the specific instantiation of the experiment. This makes it possible to compare a variety of experiment setups and analyze the impact Grid and Cloud software stacks have.
Cyberinfrastructure offers a vision of advanced knowledge infrastructure for research and education. It integrates diverse resources across geographically distributed resources and human communities. Cyberaide is a service oriented architecture and abstraction framework that integrates a large number of available commodity libraries and allows users to access cyberinfrastructure through Web 2.0 technologies. This paper describes the Cyberaide virtual appliance, a solution of on-demand deployment of cyberinfrastructure middleware, i.e. Cyberaide. The proposed solution is based on an open and free technology and software — Cyberaide JavaScript, a service oriented architecture (SOA) and grid abstraction framework that allows users to access the grid infrastructures through JavaScript. The Cyberaide virtual appliance is built by installing and configuring Cyberaide JavaScript in a virtual machine. Established Cyberaide virtual appliances can then be used via a Web browser, allowing users to create, distribute and maintain cyberinfrastructure related software more easily even without the need to do the “tricky” installation process on their own. We argue that our solution of providing Cyberaide virtual appliance can make users easy to access cyberinfrastructure, manage their work and build user organizations.
Advanced IT solutions allow users to create, organize and share user services and computing resources across a wide range of heterogeneous platforms. This makes managing (updating and deploying) the IT systems considerably more complex. Management tasks are so entangled within the existing components that it requires a IT professional with advanced training and intimate knowledge of the existing configuration to complete them. In this chapter, we will present the deployment solution for a live complex IT system, the Emergency Services Directory (ESD). Many different technologies exist that attempt automate the application deployment process by allowing a user to describe the dependencies, provide the configuration parameters, and specify the application’s required technologies, such as the operating system, database or Web server technology. ESD’s deployment solution takes advantage of the benefits provided by virtualization, and virtual appliances in particular. Each of ESD’s components are wrapped and contained with in a virtual appliance image. To automatically distribute and deploy the virtual appliance image on-demand, we utilize the Cyberaide Creative tool, which is a tool being developed in the on-going Cyberaide project.
In order to satisfy the need for sophisticated experiment and simulation management solutions for the scientific user community, various frameworks must be provided. Such frameworks include APIs, services, templates, patterns, GUIs, command line tools, and workflow systems that are specifically addressed towards the goal of assisting in the complex process of experiment and simulation management. Workflow by itself is just one of the ingredients to a successful experiment and simulation management tool. The Java CoG Kit provides an extensive framework that helps in the creation of process management frameworks for Grid and non Grid resource environments. Hence, process management in the Java CoG Kit can be defined using a Java API providing task sets, queues, and Direct Acyclic Graphs (DAGs). An alternate solution is provided in a parallel extensible scripting language with an XML syntax (a native syntax is also simultaneously supported). Visualization and monitoring interfaces are provided for both solutions, with plans for developing more sophisticated but simple to use editors. However, in this chapter we will mostly focus on our workflow solutions. The Java CoG Kit workflow solutions are developed around an abstract, high level, asynchronous task library which integrates the common Grid tasks: job submission, file transfer, and file operations. The chapter is structured as follows. First, we provide an overview of the Java CoG Kit and its evolution that led to an integrated approach to Grid Computing. We present the task abstractions library which is necessary for a flexible Grid workflow system. Next, we provide an overview of the different workflow solutions that are supported by the Java CoG Kit. Our main section focuses on only one of these solutions, in the form of a parallel scripting language which supports an XML syntax for easy integration with other tools, as well as a native, more human-oriented syntax. Additionally, a workflow repository of components is also presented, which allows sharing of workflows between multiple participants, and dynamic modification of workflows. We exemplify the use of the workflow system with a simple, conceptual application. We conclude the chapter with ongoing research activities.
Proteomic researchers who study mass spectrometry data have expressed a need for an accurate public database of empirically derived curated mass spectrum information. Lack of such a database limits proteomic researchers in their ability to identify and study proteins. Until recently, storage space and computing power has been the limiting factor in developing tools to handle the vast amount of mass spectrometry information. Now, the resources are available to store, organize, and analyze mass spectrometry information. The Illinois BioGrid Mass Spectrometry Database is a database of empirically derived tandem mass spectra of peptides created to provide researchers with an organized and searchable database of curated spectrum information to allow more accurate protein identification. This paper will discuss the methods used to import the mass spectra into the Illinois Bio-Grid Mass Spectrometry Database, as well as the database requirements, motivation, use cases, design, and results.
As part of the Java CoG Kit we have defined a sophisticated workflow framework. This workflow frame-work projects an integrated approach towards executing tasks in Grid and non-Grid environments. One of the services needed is a convenient service to store, retrieve, and modify workflow components defined by the community similar to systems such as the comprehensive perl archive network. The availability of such a service will not only allow the definition of components useful for the greater grid community, but it will also be possible that it can be reused to support dynamically changing workflows managed by collaborative groups. In this paper we present a simple extensible framework to design, build, and deploy a workflow repository service. This repository is intended to be used in ad-hoc Grids or in community Grids.
This paper describes an ad hoc Grid security infrastructure developed as a part of the Java CoG Kit project. It supports several requirements specific to the sporadic nature of ad hoc Grids. It focuses on identity management, identity verification, and authorization control in spontaneous Grid collaborations without pre-established policies or environments. It adopts established community standards, with modifications where needed. This paper also discusses the integration of the ad hoc Grid security infrastructure in an ad hoc Grid implementation. The implementation supports secure collaboration in ad hoc Grids using commodity technologies such as the Java CoG Kit, JXTA, GSI, and XACML.
Geoffrey Fox合作论文数Department of Physics, College of Arts and Sciences, Indiana University;Department of Intelligent Systems Engineering, Indiana University;Community Grid Laboratory, Indiana University;Digital Science Center of Pervasive Technology Institute;School of Engineering and Applied Science, University of Virginia15