Incorporating user experience (UX) testing when creating research cyberinfrastructure is often overlooked, but if left too late, the cost of retrofitting is considerable, and the very clients the cyberinfrastructure was built to serve may be lost. Successfully integrating UX testing into the product development cycle can be difficult but rewarding. This paper describes how UX evaluations were incorporated over ten years of operation of DataONE (www.dataone.org), a multi-sector science research cyberinfrastructure project created to support the discovery, access, and sustainability of data about life on Earth and the environment that sustains it. The diverse stakeholders in DataONE include data creators and users such as researchers and government workers across the broad scope of the earth and environmental sciences as well as those who hold and manage data such as libraries and data repositories. Between 2009 and 2019 DataONE members designed and constructed data management tools and services to fulfill the DataONE objectives. To assist in achieving its goals, a participatory design approach was used by establishing several largely volunteer and stakeholder-representative working groups, including the Usability and Assessment Working Group. This Working Group conducted over forty UX evaluations to assess the usability of DataONE products and websites at various stages of the development process. In addition to improving the usability of DataONE products, the UX evaluations fostered community involvement by building trust and engagement with the products being developed. The DataONE UX experience yields several important lessons which will improve the success of other projects. It is our conclusion that UX testing should be a mandatory part of the design of any cyberinfrastructure project.
With the mass adoption of data analysis in several scientific fields such as climatology, medicine, astronomy and astrophysics, the availability of an appropriate analytics infrastructure has become a necessity increasingly recognized by the scientific community. However, appropriate tools and applications are required to process the large volume of data collected and generated by researchers. One of the biggest challenges lies in the fact that these tools need to be gathered to be applied in specific domains. The area of bioclimatic data is a scientific field that still has much to improve in this matter. It is a field of study that lacks great efforts in the direction to provide methodologies and tools to facilitate the understanding of the complex phenomena involved in the influence that environmental variables have on biodiversity on the planet. Thus, the purpose of this work is to propose a big data analytics architecture that presents an ecosystem that systematizes and facilitates the task of the scientists to deal with the complexity in the bioclimatic data analysis, providing tools for storage, management, analysis using machine learning algorithms and data mining, and visualization tools. The methodological approach of this work was to make a thorough bibliographical study to verify the most used tools and the suitability of each one to the purpose of the work. In addition, the literature provided indications of software ecosystem implementations methodologies that served as a guide in the architecture design. Within the architecture, we attempted to gather a set of bioclimatic data based on a subset of data obtained from the Atmospheric Radiation Measurement (ARM) data repository for climatic data, and the Brazilian Biodiversity Portal for biodiversity data. As a result, we were able to gather a series of tools to access data such as Cassandra, distribution of processing such as Spark, programming interface represented by Jupyter Notebook, system modules for data format conversion, machine learning algorithms libraries and software for data visualization. This research discuss the importance of a domain purpose design of a data analysis architecture for bioclimatic data. We concluded that this type of ecosystem is imperative to facilitate the research process and increase the quality of the results.
In 2013, the Office of Science and Technology Policy (OSTP) issued a memorandum directing Federal agencies with over $100 million in annual research and development expenditures to develop a plan to support increased access to federally funded research results. In response, the US Geological Survey developed a Public Access Plan and published four new data management policies. The policies focus on review and approval of scientific data supporting scholarly conclusions, requirements for metadata, preservation, and data management planning. The new policies, in conjunction with the Public Access Plan, represent a shift in culture in how the USGS manages and provides access to its science data. The USGS recognizes that successful implementation of these new policies requires multiple pillars of support, from USGS leadership and staff buy-in, to effective tools. Active community engagement in the Bureau is stimulated through the Community for Data Integration (CDI), an open forum for community discussion and engagement, and an important component creating buy-in and contributing to the success of the new policies. Also critical are a suite of tools available to scientists to ensure their ability to implement the policies. Finally, support from leadership that manifests in the Fundamental Science Practices Advisory Council (FSPAC), a committee of representatives from across the Bureau who preside over policies and guidance is a critical component. While far from complete, the USGS has shifted its approach to science data management by engaging the community, offering tools to support policy, and providing leadership support for the quality and scientific integrity of USGS science data.
In 2013, the Office of Science and Technology Policy (OSTP) issued a memorandum directing Federal agencies with over $100 million in annual research and development expenditures to develop a plan to support increased access to federally funded research results. In response, the US Geological Survey developed a Public Access Plan and published four new data management policies. The policies focus on review and approval of scientific data supporting scholarly conclusions, requirements for metadata, preservation, and data management planning. The new policies, in conjunction with the Public Access Plan, represent a shift in culture in how the USGS manages and provides access to its science data. The USGS recognizes that successful implementation of these new policies requires multiple pillars of support, from USGS leadership and staff buy-in, to effective tools. Active community engagement in the Bureau is stimulated through the Community for Data Integration (CDI), an open forum for community discussion and engagement, and an important component creating buy-in and contributing to the success of the new policies. Also critical are a suite of tools available to scientists to ensure their ability to implement the policies. Finally, support from leadership that manifests in the Fundamental Science Practices Advisory Council (FSPAC), a committee of representatives from across the Bureau who preside over policies and guidance is a critical component. While far from complete, the USGS has shifted its approach to science data management by engaging the community, offering tools to support policy, and providing leadership support for the quality and scientific integrity of USGS science data.
In recent years, concern about the misuse of natural resources has been increasing. It is essential to know in detail the biodiversity of an ecosystem to understand and analyze the impact of human activities on nature, as well as to promote the economic growth of a country. To achieve these goals, public and private institutions are aggregating and sharing biological data around the world by means of biodiversity data portals. The main purpose of those portals is to provide a set of tools that help users and institutions catalog, analyze, and publish raw data about different species in a manner that is open and freely available to any interested party. Normally the process of choosing the best software solution is not straightforward. This paper proposes a methodology to evaluate a collection of data portals to establish a clear and consistent selection process that analyzes a collection of requirements and research purposes. The proposed approach is based on three strategies: the use of software engineering techniques to identify the desired group of features to be available in the data portal; the application of the Kano Satisfaction Model to score each requirement according to a preset weight of importance; and the use of tree-maps to visualize the requirements based on their implementation priority, to establish a portal deployment road-map. The proposed methodology is broadly applicable to portal analyses for many communities of practice.
In order to better understand the current state of data management education in multiple fields of science, this study surveyed scientists, including information scientists, about their data management education practices, including at what levels they are teaching data management, which topics they covering, and what barriers they experience in teaching these topics. We found that a handful of scientists are teaching data management in undergraduate, graduate, and other types of courses, as well as outside of classroom settings. Commonly taught data management topics included quality control, protecting data, and management planning. However, few instructors felt they were covering data management topics thoroughly, and respondents cited barriers such as lack of time, lack of necessary expertise, and lack of information for teaching data management. We offer some potential explanations for the existing state of data management education and suggest areas for further research.
Usability refers to the ease and accessibility of a system. Usability testing seeks to study how users interact with a system in order to improve the users' experience and satisfaction in achieving their objectives with the system. Usability testing is an important metric for improving a library's online services, including research data services. Libraries can help make research data available by providing repositories and data curation services for researchers to house their collected data. Providing services throughout the science data life cycle (i.e. plan, collect, share, and preserve) is important for producing higher quality research, expanding its impact, and data reuse. The Data Observation Network for Earth (DataONE) is supported by the US National Science Foundation and seeks to provide the framework and cyber-infrastructure to meet the needs of the science community to provide constant and secure access to Earth observational data. The DataONE network has heavily invested and implemented a comprehensive Usability Program to ensure user-centric software and components are made available to the variety of DataONE stakeholders. DataONE's ONEMercury is a search tool for scientific data, and the ONEDrive is a mounted workspace on the user's computer that works with ONEMercury. In 2012, a usability test was performed of the DataONE's ONEMercury tool to evaluate how scientists engage with its content and information. Twenty-six participants performed a series of tasks using the tool. MORAE software recorded the sessions, including screen display, keystrokes, and mouse movements. Participants were also asked to think aloud as they completed the tasks. The results were analyzed by observation, think aloud, time on task, and number of errors. Another usability test was performed of the DataONE's ONEDrive to assess user impressions as the tool was in development. Six participants were shown a wireframe of the tool and asked for their feedback. This paper proposes to examine the results from the ONEMercury and ONEDrive tests and draw implications for libraries
Objectives: The primary objectives of this study are to gauge the various levels of Research Data Service academic libraries provide based on demographic factors, gauging RDS growth since 2011, and what obstacles may prevent expansion or growth of services. Methods: Survey of academic institutions through stratified random sample of ACRL library directors across the U.S. and Canada. Frequencies and chi-square analysis were applied, with some responses grouped into broader categories for analysis. Results: Minimal to no change for what services were offered between survey years, and interviews with library directors were conducted to help explain this lack of change. Conclusion: Further analysis is forthcoming for a librarians study to help explain possible discrepancies in organizational objectives and librarian sentiments of RDS.
The incorporation of data sharing into the research lifecycle is an important part of modern scholarly debate. In this study, the DataONE Usability and Assessment working group addresses two primary goals: To examine the current state of data sharing and reuse perceptions and practices among research scientists as they compare to the 2009/2010 baseline study, and to examine differences in practices and perceptions across age groups, geographic regions, and subject disciplines. We distributed surveys to a multinational sample of scientific researchers at two different time periods (October 2009 to July 2010 and October 2013 to March 2014) to observe current states of data sharing and to see what, if any, changes have occurred in the past 3-4 years. We also looked at differences across age, geographic, and discipline-based groups as they currently exist in the 2013/2014 survey. Results point to increased acceptance of and willingness to engage in data sharing, as well as an increase in actual data sharing behaviors. However, there is also increased perceived risk associated with data sharing, and specific barriers to data sharing persist. There are also differences across age groups, with younger respondents feeling more favorably toward data sharing and reuse, yet making less of their data available than older respondents. Geographic differences exist as well, which can in part be understood in terms of collectivist and individualist cultural differences. An examination of subject disciplines shows that the constraints and enablers of data sharing and reuse manifest differently across disciplines. Implications of these findings include the continued need to build infrastructure that promotes data sharing while recognizing the needs of different research communities. Moving into the future, organizations such as DataONE will continue to assess, monitor, educate, and provide the infrastructure necessary to support such complex grand science challenges.
Biodiversity information is essential for understanding and managing the environment. However, identifying and providing the forms and types of biodiversity information most needed for research and decision-making is a significant challenge. While research needs and data gaps within particular topics or regions have received substantial attention, other information aspects such as data formats, sources, metadata, and information tools have received little. Focusing on the US southeast, a region of global biodiversity importance, this paper assesses the biodiversity information needs of environmental researchers, managers, and decision makers. Survey results of biodiversity information users’ information needs, information-seeking behaviors and preferred information source attributes support previous conclusions that useful biodiversity information must be easily and quickly accessible, available in forms that allow integration and visualization and appropriately matched to users’ needs. Survey results concerning additional information aspects suggest successful participation in both the creation and provision of biodiversity information include an increased focus on information search and other tools for data management, discovery, and description.
Nobody is better suited to describe data than the scientist who created it. This description about a data is called Metadata. In general terms, Metadata represents the who, what, when, where, why and how of the dataset [1]. eXtensible Markup Language (XML) is the preferred output format for metadata, as it makes it portable and, more importantly, suitable for system discoverability. The newly developed ORNL Metadata Editor (OME) is a Web-based tool that allows users to create and maintain XML files containing key information, or metadata, about the research. Metadata include information about the specific projects, parameters, time periods, and locations associated with the data. Such information helps put the research findings in context. In addition, the metadata produced using OME will allow other researchers to find these data via Metadata clearinghouses like Mercury [2][4]. OME is part of ORNL s Mercury software fleet [2][3]. It was jointly developed to support projects funded by the United States Geological Survey (USGS), U.S. Department of Energy (DOE), National Aeronautics and Space Administration (NASA) and National Oceanic and Atmospheric Administration (NOAA). OME s architecture provides a customizable interface to support project-specific requirements. Using this new architecture, the ORNL team developed OME instances formore » USGS s Core Science Analytics, Synthesis, and Libraries (CSAS&L), DOE s Next Generation Ecosystem Experiments (NGEE) and Atmospheric Radiation Measurement (ARM) Program, and the international Surface Ocean Carbon Dioxide ATlas (SOCAT). Researchers simply use the ORNL Metadata Editor to enter relevant metadata into a Web-based form. From the information on the form, the Metadata Editor can create an XML file on the server that the editor is installed or to the user s personal computer. Researchers can also use the ORNL Metadata Editor to modify existing XML metadata files. As an example, an NGEE Arctic scientist use OME to register their datasets to the NGEE data archive and allows the NGEE archive to publish these datasets via a data search portal (http://ngee.ornl.gov/data). These highly descriptive metadata created using OME allows the Archive to enable advanced data search options using keyword, geo-spatial, temporal and ontology filters. Similarly, ARM OME allows scientists or principal investigators (PIs) to submit their data products to the ARM data archive. How would OME help Big Data Centers like the Oak Ridge National Laboratory Distributed Active Archive Center (ORNL DAAC)? The ORNL DAAC is one of NASA s Earth Observing System Data and Information System (EOSDIS) data centers managed by the Earth Science Data and Information System (ESDIS) Project. The ORNL DAAC archives data produced by NASA's Terrestrial Ecology Program. The DAAC provides data and information relevant to biogeochemical dynamics, ecological data, and environmental processes, critical for understanding the dynamics relating to the biological, geological, and chemical components of the Earth's environment. Typically data produced, archived and analyzed is at a scale of multiple petabytes, which makes the discoverability of the data very challenging. Without proper metadata associated with the data, it is difficult to find the data you are looking for and equally difficult to use and understand the data. OME will allow data centers like the NGEE and ORNL DAAC to produce meaningful, high quality, standards-based, descriptive information about their data products in-turn helping with the data discoverability and interoperability. Useful Links: USGS OME: http://mercury.ornl.gov/OME/ NGEE OME: http://ngee-arctic.ornl.gov/ngeemetadata/ ARM OME: http://archive2.ornl.gov/armome/ Contact: Ranjeet Devarakonda (devarakondar@ornl.gov) References: [1] Federal Geographic Data Committee. Content standard for digital geospatial metadata. Federal Geographic Data Committee, 1998. [2] Devarakonda, Ranjeet, et al. Mercury: reusable metadata management, data discovery and access system. Earth Science Informatics 3.1-2 (2010): 87-94. [3] Wilson, B. E., Palanisamy, G., Devarakonda, R., Rhyne, B. T., Lindsley, C., & Green, J. (2010). Mercury Toolset for Spatiotemporal Metadata. [4] Pouchard, L. C., Branstetter, M. L., Cook, R. B., Devarakonda, R., Green, J., Palanisamy, G., ... & Noy, N. F. (2013). A Linked Science investigation: enhancing climate change data discovery with semantic technologies. Earth science informatics, 6(3), 175-185.« less