Knowledge of how personal experience with climate and air quality influence personal attitudes, concerns, and actions about environmental issues is increasingly important. A solid foundation for such studies is to combine interview or survey data on respondents’ attitudes and beliefs with indices created from independent meteorological or environmental monitoring data matched to the respondents location. In this project, indicators of climate and air pollution were integrated with data from the European Social Survey for a selection of large European urban regions. A prototype provenance description application was also developed for describing the workflow for creating indicators and integrating data. Our main focus was on creating indicators that represent regional anomalies in local air quality and weather for a range of time windows up-to-and-including the dates of the interviews. The goal is to facilitate investigation of relationships between urban citizen’s attitudes and behaviors as represented in the survey responses and the conditions in their local environment.
The Climate Neutral and Smart Cities project is part of the EOSC Future WP 6.3, exploring the best approaches for sharing data within cross-domain research projects. This paper looks at the implications for metadata exchange in a cross-domain research project, as explored in the project prototype. Cross-domain standard metadata is meeded to support collaborative research teams combining a mix of expertise, particularly around data lineage.
For research institutes, data libraries, and data archives, validating RDF data according to predefined constraints is a much sought-after feature, particularly as this is taken for granted in the XML world. Based on our work in two international working groups on RDF validation and jointly identified requirements to formulate constraints and validate RDF data, we have published 81 types of constraints that are required by various stakeholders for data applications. In this paper, we evaluate the usability of identified constraint types for assessing RDF data quality by (1) collecting and classifying 115 constraints on vocabularies commonly used in the social, behavioral, and economic sciences, either from the vocabularies themselves or from domain experts, and (2) validating 15,694 data sets (4.26 billion triples) of research data against these constraints. We classify each constraint according to (1) the severity of occurring violations and (2) based on which types of constraint languages are able to express its constraint type. Based on the large-scale evaluation, we formulate several findings to direct the further development of constraint languages.
DDI as a Common Format for Export and Import for Statistical Packages
For research institutes, data libraries, and data archives, RDF data validation according to predefined constraints is a much sought-after feature, particularly as this is taken for granted in the XML world. Based on our work in the DCMI RDF Application Profiles Task Group and in cooperation with the W3C Data Shapes Working Group, we identified and published by today 81 types of constraints that are required by various stakeholders for data applications. In this paper, in collaboration with several domain experts we formulate 115 constraints on three different vocabularies (DDI-RDF, QB, and SKOS) and classify them according to (1) the severity of an occurring violation and (2) the complexity of the constraint expression in common constraint languages. We evaluate the data quality of 15,694 data sets (4.26 billion triples) of research data for the social, behavioral, and economic sciences obtained from 33 SPARQL endpoints. Based on the results, we formulate several findings to direct the further development of constraint languages.
To ensure high quality of and trust in both metadata and data, their representation in RDF must satisfy certain criteria - specified in terms of RDF constraints. From 2012 to 2015 together with other Linked Data community members and experts from the social, behavioral, and economic sciences (SBE), we developed diverse vocabularies to represent SBE metadata and rectangular data in RDF. The DDI-RDF Discovery Vocabulary (DDI-RDF) is designed to support the dissemination, management, and reuse of unit-record data, i.e., data about individuals, households, and businesses, collected in form of responses to studies and archived for research purposes. The RDF Data Cube Vocabulary (QB) is a W3C recommendation for expressing data cubes, i.e. multi-dimensional aggregate data and its metadata. Physical Data Description (PHDD) is a vocabulary to model data in rectangular format, i.e., tabular data. The data could either be represented in records with character-separated values (CSV) or fixed length. The Simple Knowledge Organization System (SKOS) is a vocabulary to build knowledge organization systems such as thesauri, classification schemes, and taxonomies. XKOS is a SKOS extension to describe formal statistical classifications. In this paper, we describe RDF constraints to validate metadata on unit-record data (DDI-RDF), aggregated data (QB), thesauri (SKOS), and statistical classifications (XKOS) and to validate tabular data (PHDD) - all of them represented in RDF. We classified these constraints according to the severity of occurring constraint violations. This technical report is updated continuously as modifying, adding, and deleting constraints remains ongoing work.
From 2012 to 2015 together with other Linked Data community members and experts from the social, behavioural, and economic sciences (SBE ), we developed diverse vocabularies to represent SBE metadata and rectangular data in RDF. The DDI-RDF Discovery Vocabulary (Disco) is designed to support the dissemination, management, and reuse of person-level data, i.e., data about individuals, households, and businesses, collected in form of responses to studies and archived for research purposes. The RDF Data Cube Vocabulary (Data Cube) is a W3C recommendation for expressing data cubes, i.e. multi-dimensional aggregate data. Physical Data Description (PHDD) is a vocabulary to model data in rectangular format. The data could either be represented in records with character-separated values (CSV ) or fixed length. The Simple Knowledge Organization System (SKOS) is a vocabulary to build knowledge organization systems such as thesauri, classification schemes, and taxonomies. XKOS is a SKOS extension to describe formal statistical classifications. To ensure high quality of and trust in both metadata and data, their representation in RDF must satisfy certain criteria specified in terms of RDF constraints. In this paper, we evaluated the metadata and data quality of large real world aggregated (QB), person-level (Disco), thesauri (SKOS), rectangular (PHDD), and statistical classification (XKOS) data sets by means of RDF constraints. RDF Constraints are instances of RDF constraint types either corresponding to RDF validation requirements or to data model specific constraint types. We validated more than 4.2 billion triples and 15 thousand data sets using the RDF Validator, a validation environment which is available at http://purl.org/net/rdfval-demo.
DDI-RDF Discovery – A Discovery Model for Microdata
To ensure high quality of and trust in both metadata and data, their representation in RDF must satisfy certain criteria - specified in terms of RDF constraints. From 2012 to 2015 together with other Linked Data community members and experts from the social, behavioural, and economic sciences (SBE), we developed diverse vocabularies to represent SBE metadata and rectangular data in RDF. The DDI-RDF Discovery Vocabulary (Disco) is designed to support the dissemination, management, and reuse of person-level data, i.e., data about individuals, households, and businesses, collected in form of responses to studies and archived for research purposes. The RDF Data Cube Vocabulary (Data Cube) is a W3C recommendation for expressing data cubes, i.e. multi-dimensional aggregate data. Physical Data Description (PHDD) is a vocabulary to model data in rectangular format. The data could either be represented in records with character-separated values (CSV) or fixed length. The Simple Knowledge Organization System (SKOS) is a vocabulary to build knowledge organization systems such as thesauri, classification schemes, and taxonomies. XKOS is a SKOS extension to describe formal statistical classifications. In this paper, we describe RDF constraints to validate metadata on person-level data (Disco), aggregated data (Data Cube), thesauri (SKOS), and statistical classifications (XKOS) and to validate rectangular data (PHDD). We assign RDF constraints to RDF constraint types either corresponding to RDF validation requirements or to data model specific constraint types. This technical report is updated continuously as modifying, adding, and deleting constraints remains ongoing work.
The Qualitative Data Model Working Group was established in January 2010 with the charge “To develop a robust XML- based schema for qualitative data exchange (compliant with DDI) and encourage tools development based upon these needs.” This report describes the preliminary model developed by that group via online meetings, and working meetings in Gothenburg (2011) and Bergen (2012). This model, described in UML, was developed to cover three main scenarios: Qualitative data collections needing metadata at the object level only Qualitative data collections where segments of objects need to be delineated and described and where segments of different physical representations of the same logical objects possibly need to be linked Qualitative data collections as in the second case where related quantitative data have been generated through techniques such as text mining
The Linked Data and the Social Science data communities developed the DDI-RDF Discovery Vocabulary, an ontology of the Data Documentation Initiative, in order to support the discovery of person-level data and its metadata. The Data Documentation Initiative (DDI) is an acknowledged international standard for the documentation and management of data from the social, behavioral, and economic sciences. Within the context of DDI-RDF Discovery Vocabulary, we reuse well elaborated and accepted vocabularies to a large extend. Vocabularies like DCMI, FOAF, ORG, ADMS, PROV-O, SKOS and XKOS, DCAT, and Data Cube. This paper focuses on the description of how other vocabularies are reused reasonably and on the description of use cases which are associated with the usage of the DDI-RDF Discovery Vocabulary.