Dramatic growth in the use of computing in educational practices and research workflows encourages two-year institutions to adopt large-scale computing practices rapidly. Advanced cyberinfrastructure (CI) resources are required to support educators, students, and researchers in cloud computing, data sciences, and related cross-cutting fields such as smart manufacturing. While community colleges will differ in enrollment, geographical location, business models, and programs offered, the underlying institutional needs to store, access, manage, analyze, compute, and curate their data remain the same. Working with regional efforts to advance computing and networking, we discuss collaborative models to develop appropriate adoption models with them.
Since 2015, the XSEDE Cyberinfrastructure Resource Integration (XCRI) team has engaged in a series of on-site visits of various types with institutions wishing to implement research computing resources in a way that gains from the lessons learned by larger XSEDE Resource Providers. The team originally developed a variety of flexible toolkits for such institutions, with the intent that they could be picked up at will and used to bootstrap local research computing programs by interested folks. In practice, the real value provided by the XCRI has proven to be the deep side-by-side work, mentorship, and focus on relationship building that comes along with carrying out a site visit. While the in-person aspect of these visits ceased of necessity beginning in 2020, activities continued in a distributed fashion, which proved to be a boon for several reasons. This paper provides an overview of the site visit process, and an exploration of the deep benefits of these remote collaborations, which are an efficient way to spread RCD knowledge to under-resourced institutions.
The XSEDE Data Transfer Services (DTS) group focuses on streamlining and improving the data transfer experiences of the national academic research community, while also buttressing and future-proofing the underlying networks that support these transfers. In this paper, the DTS group shares how network and data transfer technologies have evolved over the past six years, with the backdrop of the Distributed Terascale Facility (DTF) and TeraGrid projects that served the national community before the advent of XSEDE. We delve into improvements, challenges, and trends in network and data transfer technologies, and the uses of these technologies in academic institutions across the country, which today translate into 100s of users of CI moving many terabytes each month. We also review the key lessons learned while serving the community in this regard, and what the future holds for academic networking and data transfer.
Authorship analysis is an important subject in the field of natural language processing. It allows the detection of the most likely writer of articles, news, books, or messages. This technique has multiple uses in tasks related to authorship attribution, detection of plagiarism, style analysis, sources of misinformation, etc. The focus of this paper is to explore the limitations and sensitiveness of established approaches to adversarial manipulations of inputs. To this end, and using those established techniques, we first developed an experimental frame-work for author detection and input perturbations. Next, we experimentally evaluated the performance of the authorship detection model to a collection of semantic-preserving adversarial perturbations of input narratives. Finally, we compare and analyze the effects of different perturbation strategies, input and model configurations, and the effects of these on the author detection model.
Though they are the backbone of a center’s infrastructure and store some of its most vital information, often times, the databases responsible for tracking project allocations, job accounting, storage commitments, and other core center information aren’t designed with the proper attention and thought to scale and adapt with the center as it grows and changes. Off-the-shelf solutions for these data stores are limited, and come with their own assumptions. As the field of research computing and thus the demands on the center continue to evolve, so to must the accounting solutions be able to accommodate the changing mission and goals. Many centers develop their own solutions in house, which requires additional effort to build and maintain. In this work, we present a data model which attempts to provide a more generalized solution to this problem, as well as a feature-rich set of tools provided by leveraging the powerful and widely-used Django web framework and ORM, which we call OpenAcct. We explore other solutions to this data management problem including both vendor-provided and in-house developed options. We then discuss the details of the data model, before demonstrating the capabilities of the application layer. Finally, future plans for the project are discussed.
The Evaluating and Enhancing the eXtreme Digital Cyberinfrastructure for Maximum Usability and Science Impact project, known as the Technology Investigation Service (TIS), was a collaboration between University of Illinois at Urbana-Champaign National Center for Supercomputing Applications, Pittsburgh Supercomputing Center, The University of Texas at Austin Texas Advanced Computing Center, University of Tennessee National Institute for Computational Sciences, and University of Virginia which identified and evaluated potential technologies to close the gap between the XSEDE (http://www.xsede.org) service offerings and the needs of XSEDE users. This project was funded by the NSF Division of Advanced Cyberinfrastructure (award ACI 09-46505) in response to the "Technology Audit and Insertion Service" component of the "TeraGrid Phase III: eXtreme Digital Resources for Science and Engineering (XD)" solicitation (NSF 08-571). Over the project lifetime the two major goals of TIS were: 1) identifying, tracking, evaluating and making recommendations of new technologies to XSEDE for consideration of adoption and 2) raising awareness of TIS to XSEDE and other stakeholders to solicit their input on technologies for consideration for evaluation. In accomplishing the goals, the following four significant outcomes from TIS resulted: the development and deployment of the XSEDE Technology Evaluation Database; the development of the significantly improved XSEDE software search capability; the technology evaluation process; and the evaluations performed along with their corresponding technology adoption recommendations to XSEDE. This paper highlights the life-cycle of the TIS project, including lessons learned and project outcomes.
Authentication for HPC resources has always been a double edged issue. On one hand, HPC facilities would like users to login as easily as possible, but with the increase and complexity of system exploits, HPC centers would like to protect their systems to the highest degree possible, which often leads to complicated login mechanisms. While solutions like Two Factor authentication are optimal from a system administration view point, user buy in hasn't been as vociferous as one would hope. In this paper, we discuss the implementation of an alternate solution, CILogon, at NICS. We start with a brief overview of CILogon and then delve into the implementation details. We discuss how we incorporated CILogon to enable users to use their campus credentials to login to the NICS user portal, as well as our compute resource Darter. We discuss the issues we faced during implementation and the strategies we implemented to overcome them.
In this paper we give a brief overview of the three projects that were chosen for XSEDE-PRACE collaboration in 2014. We begin this paper with an introduction of the XSEDE and PRACE organizations and the motivation for a collaborative effort between these two organizations. We then talk about the three projects that were involved in this collaboration. We provide an overview of the projects themselves and what was in scope for this collaboration. We also outline the hurdles and issues faced during this unique collaborative effort and also discuss the benefits the projects derived from this collaboration. We finally outline the future steps envisioned for XSEDE-PRACE collaborative efforts going forward.
In this paper we discuss the implementation of UNICORE in XSEDE. UNICORE is a Grid middleware tool that was identified by XSEDE to further the areas of remote job submission, campus bridging and workflows. We talk about the overall architecture of UNICORE, a typical HPC environment at XSEDE and why UNICORE is a good fit for this environment. We also discuss the initial efforts made by the UNICORE development team as well as XSEDE's Software development team to integrate UNICORE into the XSEDE landscape. We detail how UNICORE went through the XSEDE engineering process and highlight deployment details at XSEDE. We touch upon how UNICORE is beneficial to the HPC user community. In our final section we talk about future efforts to better integrate UNICORE within XSEDE.
In this paper we examine the process of designing and deploying a replacement for the old XSEDE ticket system. We look at each step, from initial concept to final deployment. The impact of XSEDE and service provider-specific policies and needs is outlined along with the final implemented global policies. We review the software and technologies under consideration along with a detailed analysis of the chosen solution. Deployed in May, 2013, the new ticket system is based on Request Tracker by Best Practical. We utilized distributed configuration management to speed deployment and arrange for rapid failover transitions providing increased uptime and ease of management between the National Institute for Computational Sciences and the Texas Advanced Computing Center. Novel high availability practices will deliver a near 100% uptime for the deployed system. We discuss future plans to deploy a federated interface that will allow the various Service Providers to update the central XSEDE ticket system directly from their local ticket systems. This federated interface will provide added functionality and will allow the XSEDE ticket system to tie into existing workflows at each Service Provider thus making it both a critical and agile service. Finally, we examine future directions the project can take in order to provide for the changing needs of the wider XSEDE community.
In this paper we examine the process of designing and deploying a replacement for the old XSEDE ticket system. We look at each step, from initial concept to final deployment. The impact of XSEDE and service provider-specific policies and needs is outlined along with the final implemented global policies. We review the software and technologies under consideration along with a detailed analysis of the chosen solution. Deployed in May, 2013, the new ticket system is based on Request Tracker by Best Practical. We utilized distributed configuration management to speed deployment and arrange for rapid failover transitions providing increased uptime and ease of management between the National Institute for Computational Sciences and the Texas Advanced Computing Center. Novel high availability practices will deliver a near 100% uptime for the deployed system. We discuss future plans to deploy a federated interface that will allow the various Service Providers to update the central XSEDE ticket system directly from their local ticket systems. This federated interface will provide added functionality and will allow the XSEDE ticket system to tie into existing workflows at each Service Provider thus making it both a critical and agile service. Finally, we examine future directions the project can take in order to provide for the changing needs of the wider XSEDE community.
In late 2010, The Georgia Institute of Technology along with its partners - the Oak Ridge National Lab, the University of Tennessee-Knoxville, and the National Institute for Computational Sciences, deployed the Keeneland Initial Delivery System (KIDS) - a 201 Teraflop, 120node HP SL390 system with 240 Intel Xeon CPUs and 360 NVIDIA Fermi graphics processors as a part of the Keeneland Project. The Keeneland Project is a five-year Track 2D cooperative agreement awarded by the National Science Foundation (NSF) in 2009 for the deployment of an innovative high performance computing system in order to bring emerging architectures to the open science community, KIDS is being used to develop programming tools and libraries in order to ensure that the project can productively accelerate important scientific and engineering applications. Until late 2011, there was no formal mechanism in place for quantifying the efficiency of GPU usage on the Keeneland system because most applications did not have the appropriate administrative tools and vendor support. GPU administration has largely been an afterthought as vendors in this space are focused on gaming and video applications. There is a compelling need to monitor GPU utilization on Keeneland for the purposes of proper system administration and future planning for Keeneland Final System, which is expected to be in production in July 2012. With the release of CUDA 4.1, NVIDIA added enhanced functionality to the nvidia-system management interface (nvidia-smi) tool, which is a management and monitoring command line utility that leverages the NVIDIA Management Library (NVML). NVML is a C-based API for monitoring and managing various states of the NVIDIA GPU devices. It provides a direct access to the queries and commands exposed via nvidia-smi. Using nvidia-smi, a monitoring tool was built for KIDS, to monitor utilization and memory usage on the GPUs. In this paper, we discuss the development of the GPU Utilization tool in depth, and its implementation details on KIDS. We also provide an analysis of the utilization statistics generated by this tool. For example, we identify utilization trends across jobs submitted on KIDS - such as overall GPU utilization as compared to CPU utilization. We also examine GPU utilization from the perspective of software - which packages are most frequently used, and how do they compare with respect to GPU utilization and memory usage. Collection and analysis of this data is essential for facilitating heterogeneous computing on the Keeneland Initial Delivery System. Future direction for the usage of these statistics is to provide insights on overall usage of the system, determine appropriate ratios for jobs (CPU to GPU, GPU to host memory), assist in scheduling policy management, and determine software utilization. These statistics become even more relevant as the center prepares for the deployment of the Keeneland Final System. As heterogeneous computing appears to be more and more common, and is quickly becoming the standard, this information will help greatly in delivering consistent high uptime and assist software developers in writing more efficient code for the majority of the codebases aimed at heterogeneous systems.
High performance computing resources attract a wide range of computational users and corresponding job widths and lengths. For example, on the petaflop Cray XT5 machine, Kraken, users submit jobs ranging from a few hundred cores (capacity computing) to over hundred thousand cores (capability computing). Traditionally it has been difficult to maintain high utilization while juggling such a diverse job mix. This paper explores four unique approaches to achieve our scheduling goals of maximizing utilization on four distinct resources at the National Institute for Computational Sciences. The resources include the petaflop machine, Kraken, Athena — a 166 TF Cray XT4, a 4 TB shared memory NUMA machine called Nautilus, and a GPU cluster called Keeneland.
In late 2009, the National Institute for Computational Sciences placed in production the world‘s fastest academic supercomputer (third overall, November 2009 Top500), a Cray XT5 named Kraken. Currently Kraken provides 1.17 Petaflops through 112,896 compute nodes accounting for over 60% of the total cycles available to the National Science Foundation users via the TeraGrid. Kraken has two missions that have proven difficult to simultaneously reconcile: providing the maximum number of total cycles to the community, while enabling full machine runs for “hero” users. Historically, this has been attempted by allowing schedulers to choose the time for the beginning of large jobs, with a concomitant reduction in utilization. This paper outlines a novel approach implemented at NICS, whereby the “draining” of the system is forced on a weekly basis, followed by consecutive full machine runs. As previous simulation results suggested, this led to utilization of over 90% (the equivalent of a 300+ Teraflop supercomputer, or several million dollars of compute time per year) with a 92% average over the
Increasingly large datasets acquired by NASA for global climate studies demand larger computation memory and higher CPU speed to mine out useful and revealing information. While boosting the CPU frequency is getting harder, clustering multiple lower performance computers thus becomes increasingly popular. This prompts a trend of parallelizing the existing algorithms and methods by mathematicians and computer scientists. In this paper, we take on the task of parallelizing the Nonnegative Tensor Factorization (NTF) method, with the purposes of distributing large datasets into each cluster node and thus reducing the demand on a single node, blocking and localizing the computation at the maximal degree, and finally minimizing the memory use for storing matrices or tensors by exploiting their structural relationships. Numerical experiments were performed on a NASA global sea surface temperature dataset and result factors were analyzed and discussed.
Phil Andrews合作论文数San Diego Supercomputer Center, UCSD1
Shava Smallen合作论文数San Diego Supercomputer Center1