This article details a correction to: Gandhi, S., Diggs, S., Córdoba, M.A., Bezuidenhout, L., Cobe, R., El Jadid, S., Peterson, B., Quick, R., Shanahan, H., Venkataraman, S., Okorafor, E. and Van den Eynden, V. (2026) ‘Building Responsible and Sustainable Open Data Literacy Skills for Early Career Researchers: A Decade of the SoRDS Programme’, Data Science Journal, 25(1), p. 12. https://doi.org/10.5334/dsj-2026-012.
Recent advancements in Artificial Intelligence (AI), including generative systems like ChatGPT, have motivated a reevaluation of how AI methods can revolutionize research computing with possible enhancements in accessibility, resource allocation, and cybersecurity. Science gateways have long served as solutions for facilitating research by simplifying complexities associated with using advanced computing resources and promoting collaboration. Despite their shared domain in research computing, the AI and science gateway communities have a slight overlap compared to more significant potential. To bridge this gap, the NSF Center of Excellence for Science Gateways (SGX3) has started a strategic initiative called a Blueprint Factory, to map the future of science gateways in an AI-driven ecosystem. Blueprint Factories are intensive endeavors spanning 18 months and aim to identify the technical capabilities required to support evolving scientific domains and computing resources. As the initial event of the AI Blueprint Factory, a Birds-of-a-Feather (BoF) session was organized at PEARC’23. The session focused on emerging AI needs in research computing and their implications for science gateways. The session involved interactive discussions to explore novel modes of computing, data sharing considerations, and the potential for broadening access to compute resources through AI integration. The outcomes of this session provided valuable insights into key themes, opportunities, and challenges surrounding the integration of AI into science gateways. This paper details the methods used and the results obtained by this interactive session. It then provides an outlook on the next steps of the Blueprint Factory.
Contemporary research, particularly when addressing the most significant transdisciplinary research challenges, cannot effectively be done without a range of skills relating to data management, data analysis, and cyberinfrastructure (CI). These data and CI skills are common to all disciplines that conduct data-centric research. Research Data Science acts as a vital component of the scientific process. In a grassroots attempt to address this gap, the CODATA-RDA Schools of Research Data Science (SoRDS) was founded in 2016 to provide instruction on foundational data science and open research concepts to early career researchers in low and middle-income countries. This partnership between international collaborators has since 2016 provided this training to over 1000 early career researchers in 24 events in 10 countries worldwide. The most recent event was held at Georgia Institute of Technology and focused on health equity and included researchers from minority-serving institutions in the southeast United States. This paper covers the background of the SoRDS project along with organization and curriculum details. It also covers the transition of the events from a non-domain-centric curriculum to spotlighting biological and social health equity data and what we learned to make future health-related events more engaging and valuable to the attendees. It also looks toward future events that will serve international students studying health informatics and other data-centric disciplines.
This paper gives a summary of implementation activities in the realm of FAIR Digital Objects (FDO). It gives an idea which software components are robust and used for many years, which components
We summarize student-led work to build a science gateway for RNAMake, which is software for modeling the three-dimensional structure of RNA molecules. The gateway uses Apache Airavata, which has been extended to support HTCondor submissions. The students also extended the Airavata Django Portal to provide customized user interfaces. In the process, the students learned open source software and open governance practices.
Data science skills are rapidly becoming a necessity in modern science. In response to this need, institutions and organizations around the world are developing research data science curricula to teach the programming and computational skills that are needed to build and maintain data infrastructures and maximize the use of available data. To date, however, few of these courses have included an explicit ethics component, and developing such components can be challenging. This paper describes a novel approach to teaching data ethics on short courses developed for the CODATA-RDA Schools for Research Data Science. The ethics content of these schools is centred on the concept of open and responsible (data) science citizenship that draws on virtue ethics to promote ethics of practice. Despite having little formal teaching time, this concept of citizenship is made central to the course by distributing ethics content across technical modules. Ethics instruction consists of a wide range of techniques, including stand-alone lectures, group discussions and mini-exercises linked to technical modules. This multi-level approach enables students to develop an understanding both of “responsible and open (data) science citizenship”, and of how such responsibilities are implemented in daily research practices within their home environment. This approach successfully locates ethics within daily data science practice, and allows students to see how small actions build into larger ethical concerns. This emphasises that ethics are not something “removed from daily research” or the remit of data generators/end users, but rather are a vital concern for all data scientists.
Research data currently face a huge increase of data objects with an increasing variety of types(data types, formats) and variety of workflows by which objects need to be managed across their lifecycle by data infrastructures. Researchers desire to shorten the workflows from data generation to analysis and publication, and the full workflow needs to become transparent to multiple stakeholders, including research administrators and funders. This poses challenges for research infrastructures and user-oriented data services in terms of not only making data and workflows findable, accessible, interoperable and reusable, but also doing so in a way that leverages machine support for better efficiency. One primary need to be addressed is that of findability, and achieving better findability has benefits for other aspects of data and workflow management. In this article, we describe how machine capabilities can be extended to make workflows more findable, in particular by leveraging the Digital Object Architecture, common object operations and machine learning techniques.
Global middleware infrastructure is insufficient for robust data identification, discovery, and use. While infrastructure is emerging within sub-ecosystems such as the DOI ecosystem of services purposed for data and literature objects (i.e., DataCite, CHORUS, CrossRef), in general the layers of abstraction that have made the Internet so easy to build on, is lacking for data especially for computer (machine) automated services. The goal of the PID Kernel Information recommendation is to advance a small change to middleware infrastructure by injecting a tiny amount of carefully selected metadata into a Persistent ID (PID) record. This carefully chosen and placed information has the potential to stimulate development of an entire ecosystem of third party services that can process the anticipated billions of PIDs and do so with more information at hand about a digital object (no need for costly link following) than just its globally persistent ID (PID). The key challenge of the PID Kernel Information working group was to determine which from amongst thousands of relevant metadata elements are suitable to embed in the PID record. This recommendation lays out principles to guide in the identification of information suitable for inclusion in the PID record. The information contained in a PID record is represented by a PID Kernel Information profile which must be publicly and globally available. For PID Kernel Information to be effective in stimulating an ecosystem of data services, the number of different profiles of PID Kernel Information must be small and their content stable. The recommendation includes a draft profile with illustrating examples and cases for adoption in practice. 1 RDA Recommendation on PID Kernel Information 1. Background and Scope The primary purpose of the guiding principles is to help profile developers determine which information should (and should not) be included in a PID Kernel Information profile. What is a PID Kernel Information profile? PID Kernel Information is a set of attributes stored within the PID record, i.e., information stored at a global or local PID registry and accessible by a resolver. The primary purpose of PID Kernel Information is in support of smart machine actionable decisions that can be accomplished through inspection of the PID record alone. PID Kernel Information profiles are registered schemas for interpreting PID KI records. PID records may be created according to specific profiles and checked for conformance against them. In other words, PID KI records are concrete instantiations of profiles, comparable to how we consider objects as instantiations of classes in object-oriented programming. Scope and applicability The principles apply to profiles for PIDs that reference (point to) data objects which have a single or canonical digital manifestation. The object itself can be digital data, code, metadata or the digital representation of a physical object, etc. These principles are geared toward PID systems with the following attributes: 1) they register, store, and retrieve a small amount of metadata, and 2) there is a globally discoverable service available through which information about the PID Kernel Information profiles can be retrieved, which are referenced in the PID records. Current systems that potentially meet these requirements include the Archive Resource Key (ARK) and the Handle service (and Data Type Registry) and conceptually any URN-based service with a managed global resolver. We expect other systems do as well. The WG used the Handle service as a model or test case in the development of these principles and seeks feedback on how they apply to similar systems. The WG also builds on the existing RDA Recommendation for a Data Type Registry as a 1 globally available (and distributed) type registry for both data types, such as data format definitions, and PID Kernel Information profiles. Whether the same registry should take both roles at the same time needs further exploration. 1 DOI: dx.doi.org/10.15497/A5BCD108-ECC4-41BE-91A7-20112FF77458 2 RDA Recommendation on PID Kernel Information
Data Intensive science, a Fourth Dimension of Science as introduced by J. Gray [1] is a fact in many large institutions, however, participation is mainly limited to data scientists in those institutions, i.e. many scientists are excluded, and many relevant data driven scientific questions can simply not be taken up, since the costs are too high. Similarly we are far away from a flourishing data economy where many companies can participate, in particular innovative small enterprises and start-ups. There are a number of different reasons for this situation including uncertainties with respect to legal issues. Here, we want to address an issue which is conceptual as well as technical.
Throughout the last decade the Open Science Grid (OSG) has been fielding requests from user communities, resource owners, and funding agencies to provide information about utilization of OSG resources. Requested data include traditional accounting - core-hours utilized - as well as users certificate Distinguished Name, their affiliations, and field of science. The OSG accounting service, Gratia, developed in 2006, is able to provide this information and much more. However, with the rapid expansion and transformation of the OSG resources and access to them, we are faced with several challenges in adapting and maintaining the current accounting service. The newest changes include, but are not limited to, acceptance of users from numerous university campuses, whose jobs are flocking to OSG resources, expansion into new types of resources (public and private clouds, allocation-based HPC resources, and GPU farms), migration to pilot-based systems, and migration to multicore environments. In order to have a scalable, sustainable and expandable accounting service for the next few years, we are embarking on the development of the next-generation OSG accounting service, GRACC, that will be based on open-source technology and will be compatible with the existing system. It will consist of swappable, independent components, such as Logstash, Elasticsearch, Grafana, and RabbitMQ, that communicate through a data exchange. GRACC will continue to interface EGI and XSEDE accounting services and provide information in accordance with existing agreements. We will present the current architecture and working prototype.
The Open Science Grid (OSG) relies upon the network as a critical part of the distributed infrastructures it enables. In 2012, OSG added a new focus area in networking with a goal of becoming the primary source of network information for its members and collaborators. This includes gathering, organizing, and providing network metrics to guarantee effective network usage and prompt detection and resolution of any network issues, including connection failures, congestion, and traffic routing.
The Structural Protein-Ligand Interactome (SPLINTER) project predicts the interaction of thousands of small molecules with thousands of proteins.These interactions are predicted using the three-dimensional structure of the bound complex between each pair of protein and compound that is predicted by molecular docking.These docking runs consist of millions of individual short jobs each lasting only minutes.However, computing resources to execute these jobs (which cumulatively take tens of millions of CPU hours) are not readily or easily available in a cost effective manner.By looking to National Cyberinfrastructure resources, and specifically the Open Science Grid (OSG), we have been able to harness CPU power for researchers at the Indiana University School of Medicine to provide a quick and efficient solution to their unmet computing needs.Using the job submission infrastructure provided by the OSG, the docking data and simulation executable was sent to more than 100 universities and research centers worldwide.These opportunistic resources provided millions of CPU hours in a matter of days, greatly reducing time docking simulation time for the research group.The overall impact of this approach allows researchers to identify small molecule candidates for individual proteins, or new protein targets for existing FDA-approved drugs and biologically active compounds.
Over the course of 2012-13 the Open Science Grid (OSG) transitioned the identity management system for its science user community from the DOE Grids public key infrastructure (PKI) to a new OSG PKI.This transition was significant in its scope, touching on nearly all aspects of the OSG infrastructure and community.The transition also entailed the adoption of a commercial certificate service as a key component of OSG's PKI.This transition offers a rare opportunity to better understand identity management and how to prepare for and implement changes in an identity management system.In this paper, we describe OSG's transition and lessons learned from it.We discuss the overall project management approach, including a division of the project into planning, piloting, design, development, implementation and transition phases.We discuss the considered alternatives, both for implementations of the OSG PKI as well as alternatives to a PKI such as federated identity, as well as the criteria we used to make our decision.We conclude with a set of lessons learned from both implementation and in retrospect, and a set of recommendations for other identity systems.
The Open Science Grid encourages the concept of software portability: a user's scientific application should be able to run at as many sites as possible. It is necessary to provide a mechanism for OSG Virtual Organizations to install software at sites. Since its initial release, the OSG Compute Element has provided an application software installation directory to Virtual Organizations, where they can create their own sub-directory, install software into that sub-directory, and have the directory shared on the worker nodes at that site. The current model has shortcomings with regard to permissions, policies, versioning, and the lack of a unified, collective procedure or toolset for deploying software across all sites. Therefore, a new mechanism for data and software distributing is desirable. The architecture for the OSG Application Software Installation Service (OASIS) is a server-client model: the software and data are installed only once in a single place, and are automatically distributed to all client sites simultaneously. Central file distribution offers other advantages, including server-side authentication and authorization, activity records, quota management, data validation and inspection, and well-defined versioning and deletion policies. The architecture, as well as a complete analysis of the current implementation, will be described in this paper.
This material is based upon work supported by the National Science Foundation under Grant No. ABI-1062432, Craig Stewart, PI. William Barnett, Matthew Hahn, and Michael Lynch, co-PIs. This work was supported in part by the Lilly Endowment, Inc. and the Indiana University Pervasive Technology Institute. Any opinions presented here are those of the presenter(s) and do not necessarily represent the opinions of the National Science Foundation or any other funding agencies
Eric A. Wernert合作论文数Advanced Visualization Laboratory2