This paper describes a group of online services which are designed to support social survey research and the production of statistical results. The 'Grid Enabled Specialist Data Environment' (GESDE) services constitute three related systems which offer facilities to search for, extract and exploit supplementary data and metadata concerned with the measurement and operationalisation of survey variables. The services also offer users the opportunity to deposit and distribute their own supplementary data resources for the benefit of dissemination and replication of the details of their own analysis. The GESDE services focus upon three application areas: specialist data relating to the measurement of occupations; educational qualifications; and ethnicity (including nationality, language, religion, national identity). They identify information resources related to the operationalisation of variables which seek to measure each of these concepts - examples include coding frames, crosswalk and translation files, and standardisation and harmonisation recommendations. These resources constitute important supplementary data which can be usefully exploited in the analysis of survey data. The GESDE services work by collecting together as much of this supplementary data as possible, and making it searchable and retrievable to others. This paper discusses the current features of the GESDE services (which have been designed as part of a wider programme of ‘e-Science’ research in the UK), and considers ongoing challenges in providing effective support for variable-oriented statistical analysis in the social sciences.
The last two decades have seen substantially increased potential for quantitative social science research. This has been made possible by the significant expansion of publicly available social science datasets, the development of new analytical methodologies, such as microsimulation, and increases in computing power. These rich resources do, however, bring with them substantial challenges associated with organizing and using data. These processes are often referred to as ‘data management’. The Data Management through e-Social Science (DAMES) project is working to support activities of data management for social science research. This paper describes the DAMES infrastructure, focusing on the data-fusion process that is central to the project approach. It covers: the background and requirements for provision of resources by DAMES; the use of grid technologies to provide easy-to-use tools and user front-ends for several common social science data-management tasks such as data fusion; the approach taken to solve problems related to data resources and metadata relevant to social science applications; and the implementation of the architecture that has been designed to achieve this infrastructure.
Metadata Creation, Transformation and Discovery for Social Science Data Management: The DAMES Project Infrastructure
This article discusses how quantitative data analysis in the social sciences can engage with and exploit an e-Infrastructure. We highlight how a number of activities that are central to quantitative data analysis, referred to as "data management,'' can benefit from e-Infrastructural support. We conclude by discussing how these issues are relevant to the Data Management through e-Social Science (DAMES) research Node, an ongoing project that aims to develop e-Infrastructural resources for quantitative data analysis in the social sciences.
Quantitative social survey research struggles to reconcile the widespread collection of categorical measures, and the common desire to conduct analyses which are not well suited to categorical data (involving comparing standardized, relative positions). By describing approaches taken in the DAMES NCeSS research Node, which is developing facilities to assist in the management of categorical data on occupations, educational qualifications, and ethnicity, this paper argues that tools and practices associated with e-Social Science offer an opportunity to raise standards in the analysis of categorical data.
Grid computing is moving from its original focus on the physical sciences to other disciplines such as the social sciences. The orientation of these newer applications is on data management rather than processing. This paper describes how the DAMES project (Data Management through E-Social Science) is developing grid-based solutions for handling data in a distributed environment. The paper describes the approach being taken to meet key challenges: metadata for effective use of datasets, and data-oriented workflows for e-social science.