Effective discovery and integration of ecological data within data management systems requires rich semantic information that can describe and relate the types of information contained within disparate data sets. Within the Semtools project, we have developed approaches for expressing and representing semantic annotations of data sets for supplementing attribute and data-level metadata with terms drawn from domain-specific ontologies. Annotations provide a formal mechanism that can be used together with reasoning systems to enhance existing data discovery and integration approaches. We describe extensions to the Ecological Metadata Language (EML) and associated tools for storing and using semantic annotations. Specifically, we describe new user interface components implemented within the Morpho metadata editor for capturing user-supplied semantic annotations, extensions to the Metacat system for storing and accessing annotations and corresponding OWL-DL ontologies, and a new API within Metacat that uses annotation metadata to provide concept-based search and integration of data sets. Keywords—ontologies; annotation; data discovery and integration
In many scientific disciplines, including ecology, hydrology, and earth science, scientific analysis requires access to a broad range of observational data. However, because of the amount and heterogeneity both in the structure and semantics) of observational data, approaches are needed that allow scientists to easily discover and analyze them. To address this issue, we describe a framework for accessing sobservational data. This framework combines a core observational model, domain-specific ontologies compatible with the core model, and a semantic annotation language. The annotation language provides a formal bridge between the core model and the underlying data to enable queries and analysis over annotations. The framework has been implemented to take advantage of ontology and web-based standards, and has also been integrated within a popular metadata tool for managing ecological datasets.
The Ecological Metadata Language is an effective specification for describing data for long-term storage and interpretation. When used in conjunction with a metadata repository such as Metacat, and a metadata editing tool such as Morpho, the Ecological Metadata Language allows a large community of researchers to access and to share their data. Although the Ecological Metadata Language/Morpho/Metacat toolkit provides a rich data documentation mechanism, current methods for retrieving metadata-described data can be laborious and time consuming. Moreover, the structural and semantic heterogeneity of ecological data sets makes the development of custom solutions for integrating and querying these data prohibitively costly for large-scale synthesis. The Data Manager Library leverages the Ecological Metadata Language to provide automated data processing features that allow efficient data access, querying, and manipulation without custom development. The library can be used for many data management tasks and was designed to be immediately useful as well as extensible and easy to incorporate within existing applications. In this paper we describe the motivation for developing the Data Manager Library, provide an overview of its implementation, illustrate ideas for potential use by describing several planned and existing deployments, and describe future work to extend the library.
One of the challenges of research in science education is storing, managing and querying the large amounts of diverse student assessment data that are typically collected in many Science, Technology, Engineering and Mathematics (STEM) courses. Furthermore, longitudinal studies across courses and ABET accreditation necessitate tracking students throughout their academic programs in which each course will have different types of data. Researchers need to manage, assign metadata to, merge, sort, and query all of these data to support instructional decisions, research and accreditation. To address these needs we have constructed a database to support both data-driven instructional decision making and research in STEM education. We have built upon existing metadata standards to define an extensible Educational Metadata Language (EdML) that enables assessments to be tagged based on taxonomies, standard psychometrics such as difficulty and discrimination, and other data to facilitate cross-study analyses. Once a collection of assessment data are available, faculty can examine their assessment data to evaluate historical trends, analyze the effectiveness of pedagogical techniques and strategies, or compare the performance of different teaching and assessment techniques within their course or across institutions.