Despite growing recognition of the importance of public data to the modern economy and to scientific progress, long-term investment in the repositories that manage and disseminate scientific data in easily accessible-ways remains elusive. Repositories are asked to demonstrate that there is a net value of their data and services to justify continued funding or attract new funding sources. Here, representatives from a number of environmental and Earth science repositories evaluate approaches for assessing the costs and benefits of publishing scientific data in their repositories, identifying various metrics that repositories typically use to report on the impact and value of their data products and services, plus additional metrics that would be useful but are not typically measured. We rated each metric by (a) the difficulty of implementation by our specific repositories and (b) its importance for value determination. As managers of environmental data repositories, we find that some of the most easily obtainable data-use metrics (such as data downloads and page views) may be less indicative of value than metrics that relate to discoverability and broader use. Other intangible but equally important metrics (e.g., laws or regulations impacted, lives saved, new proposals generated), will require considerable additional research to describe and develop, plus resources to implement at scale. As value can only be determined from the point of view of a stakeholder, it is likely that multiple sets of metrics will be needed, tailored to specific stakeholder needs. Moreover, economically based analyses or the use of specialists in the field are expensive and can happen only as resources permit.
Significant progress has been made in the past few years in the development of recommendations, policies, and procedures for creating and promoting citations to data sets, software, and other research infrastructures like computing facilities. Open questions remain, however, about the extent to which referencing practices of authors of scholarly publications are changing in ways desired by these initiatives. This paper uses four focused case studies to evaluate whether research infrastructures are being increasingly identified and referenced in the research literature via persistent citable identifiers. The findings of the case studies show that references to such resources are increasing, but that the patterns of these increases are variable. In addition, the study suggests that citation practices for data sets may change more slowly than citation practices for software and research facilities, due to the inertia of existing practices for referencing the use of data. Similarly, existing practices for acknowledging computing support may slow the adoption of formal citations for computing resources.
Recent policy shifts on the part of funding agencies and journal publishers are causing changes in the acknowledgment and citation behaviors of scholars. A growing emphasis on open science and reproducibility is changing how authors cite and acknowledge "research infrastructures"-entities that are used as inputs to or as underlying foundations for scholarly research, including data sets, software packages, computational models, observational platforms, and computing facilities. At the same time, stakeholder interest in quantitative understanding of impact is spurring increased collection and analysis of metrics related to use of research infrastructures. This article reviews work spanning several decades on tracing and assessing the outcomes and impacts from these kinds of research infrastructures. We discuss how research infrastructures are identified and referenced by scholars in the research literature and how those references are being collected and analyzed for the purposes of evaluating impact. Synthesizing common features of a wide range of studies, we identify notable challenges that impede the analysis of impact metrics for research infrastructures and outline key open research questions that can guide future research and applications related to such metrics.
This research investigates the usage distribution of instructional resources shared among educators in an online learning community. The usage of a resource is defined by the number of unique educators who use (click on) it. We explored what the usage distribution of these resources looks like and we investigated what underlying mechanisms may have generated the observed distribution. Our results indicate that the usage distribution of resources follows a power law. Furthermore, our results also suggest that an educator’s decision to use a resource may be influenced by the prior decisions of others. 82.6% of 2500 simulations of an information cascade model developed to model the resource selection process of educators resulted in a power law distribution as observed in our data. Information cascades provide a natural way of understanding how individuals may imitate the decisions of others even when such decisions do not align with their perssonal preferences.
The EarthCollab project is using the VIVO Semantic Web software suite to support the discovery of information, data, and potential collaborators within the geodesy and polar science communities. This paper discusses the ontology selection, consolidation, and reuse efforts of EarthCollab. EarthCollab’s ontology design approach heavily emphasizes ontology reuse, bringing together existing ontologies to support diverse use cases related to the discovery of geoscience information and resources. We developed a small local ontology to tie these existing ontologies together and to build appropriate geoscience-relevant connections. Five key ontology decision drivers are presented to outline EarthCollab’s ontology design process and decision points: use cases, existing systems and metadata, semantic application dependencies, external ontology characteristics, and community recommendations for good ontological modeling practices.
This paper presents a computational approach to understanding and predicting the behavior of Earth Science educators using an online curriculum planning tool incorporating digital library resources. It expands on prior work on understanding educators' adoption and use of digital library resources [2] by introducing a methodology for characterizing user behaviors and understanding the trends and frequent patterns of use that are observable from these behaviors.
Engaging young science learners today requires a plethora of tools that oftentimes leverages technology in novel ways. This paper describes the use of several 21st century technologies to engage science learners in locally relevant climate science research projects and the presentation of these projects in an entirely online virtual student conference. Case studies demonstrating the use of and effectiveness of 21st century technologies and GLOBE protocols are also included. Through technology, students were able to find out more about distant locations and their own environments, talk to scientists, and make studying climate science personally relevant. Finally, the implementation and structure of the GLOBE Virtual Student Conference is described.
As technology continues to disrupt education at nearly all levels from K–12 to college and beyond, the challenges of understanding the impact technology has on teaching continue to mount. One critical area that yet remains open, is examining teachers’ usage of technology by specifically collecting detailed data of their technology use, developing techniques to analyze that data and then finding meaningful connections that may show the value of that technology. In this research, we will present a model for predicting test score gains using data points drawn from typical educational data sources such as teacher experience, student demographics and classroom dynamics, as well as from the online usage behaviors of teachers. Building upon prior work in developing a usage typology of teachers using an online curriculum planning system, the Curriculum Customization Service (CCS), to assist in the development of their instruction and planning for an Earth systems curriculum, we apply the results of this typology to add new information to a model for predicting test score gains on a district-level Earth systems subject area exam. Using both multinomial logistic regression and Naive Bayes algorithms on the proposed model, we show that even with a simplification of the highly complex tapestry of variables that go into teacher and student performance, teacher usage of the CCS proved valuable to the predictive capability in average and above average test score gains cases.
Recommender systems have become part of the standard toolkit of web personalization. These same tools and techniques are now making their way into educational and adaptive e-learning systems. In this chapter, we will discuss aspects of a prototype system, the Customized Learning Service for Concept Knowledge (CLICK), an application designed to provide digital library resources recommendations based on user’s concept knowledge demonstrated through automated evaluation and approximation of their knowledge state from essay writing. We present the underlying concepts behind recommender systems, review learner models as they are designed within the CLICK environment, and review the lessons learned. We will discuss aspects of how CLICK supports intentional learning as well as extensions to the existing technology to improve such support. Future challenges and directions for CLICK and related technologies are also discussed.
Dean B. Krafft合作论文数Cornell University Library6