In 2012 the Australian Bureau of Meteorology published a dataset, ACORN-SAT, containing the homogenised daily temperature observations of 112 locations throughout Australia for the last 100 years. The dataset employs the latest analysis techniques and takes advantage of newly digitised observational data to monitor climate variability and change in Australia. The observations in ACORN-SAT were initially published only as comma separated values, whereas the metadata was published in a PDF report. In 2013 we converted the metadata and the observation data into RDF and published the result as Linked Open Data, accessible online via a pilot government linked data service built on the Linked Data API. In this article we describe the process of transforming the original tabular data into a Linked Sensor Data Cube (in: Proc. of the 5th International Workshop on Semantic Sensor Networks, SSN12, CEUR-WS.org, 2012, pp. 1-16) based on the W3C Semantic Sensor Network ontology (Web Semantics: Science, Services and Agents on the World Wide Web 17 (2012), 25-32) and the W3C RDF Data Cube vocabulary (The RDF Data Cube Vocabulary, W3C Recommendation, 16 January 2014). We further discuss how the dataset has since been used and interlinked with near-real time weather observations for the 112 sensing locations of the ACORN-SAT that are published by the Bureau of Meteorology. Both the original ACORN-SAT dataset and the weather observation data are accessible online at lab.environment.data.gov.au.
The Bureau of Meteorology publishes a range of observational data from rainfall to river levels. The primary way of accessing these data is through HTML pages or a range of CSV and JSON documents. There are a number of challenges in accessing these data: formats and access mechanisms are different across each data product; there are no explicit relationships between concepts within the data (e.g. whether the specific phenomenon being measured is the same in two different data sets) and no explicit relationships between data sets (e.g. manual daily rainfall observations relate to other rainfall observation types). Linked data is an emerging data-publishing concept that promotes publishing data with links between concepts described in a consistent way. It has potential to assist with issues of identity of objects on the web (e.g. monitoring stations), relationships between data, and standardising access mechanisms. We are publishing near real-time weather observations as Linked Data to allow connections to be made between current conditions and long-term averages. This builds on previous work that published the Australian Climate Observations Reference Network - Surface Air Temperature (ACORN-SAT) data as Linked Data. The published observations allow navigation and query across the two data sets, allowing dynamic comparison of current observations with long-term averages. We describe the vocabularies used to model these data, an approach for linking between related concepts, and the prototype APIs to access the data.
We compare experiences in modelling geographic information in OWL by two very different methods. Both methods have a root in OGC standards, but one followed a semantic web development style and the other a UML-flavoured Model Driven Architecture style. We ask how much of the OGC UML modelling should be preserved in the transition from UML modelling to OWL modelling and from XML to open linked data.
The effort over more than a decade to establish the semantic web [Berners-Lee et. al., 2001] has received a major boost in recent years through the Open Government movement. Governments around the world are seeking technical solutions to enable more open and transparent access to Public Sector Information (PSI) they hold. Existing technical protocols and data standards tend to be domain specific, and so limit the ability to publish and integrate data across domains (health, environment, statistics, education, etc.).
Traditionally, the formal scientific output in most fields of natural science has been limited to peer-reviewed academic journal publications, with less attention paid to the chain of intermediate data results and their associated metadata, including provenance. In effect, this has constrained the representation and verification of the data provenance to the confines of the related publications. Detailed knowledge of a dataset’s provenance is essential to establish the pedigree of the data for its effective re-use, and to avoid redundant re-enactment of the experiment or computation involved. It is increasingly important for open-access data to determine their authenticity and quality, especially considering the growing volumes of datasets appearing in the public domain. To address these issues, we present an approach that combines the Digital Object Identifier (DOI) – a widely adopted citation technique – with existing, widely adopted climate science data standards to formally publish detailed provenance of a climate research dataset as an associated scientific workflow. This is integrated with linked-data compliant data re-use standards (e.g. OAI-ORE) to enable a seamless link between a publication and the complete trail of lineage of the corresponding dataset, including the dataset itself.
The Australian Bureau of Meteorology (BOM) has recently published a homogenised daily temperature dataset, ACORN-SAT, for the monitoring of climate variability and change in Australia. The dataset employs the latest analysis techniques and takes advantage of newly digitised observational data to provide a daily temperature record over the last 100 years. In this paper, we present a case-study to publish the ACORN-SAT as Linked Data. We use the Semantic Sensor Network ontology to deliver the publicly available metadata about the BOM weather stations and their deployment history as linked data. Additionally, for concepts that are not covered by existing vocabularies, we have developed domain ontologies to define the adjusted aggregate variables and associated parameters for the ACORN-SAT homogenised observation data, the BOM weather stations and the BOM Rainfall districts. We use the RDF Data Cube Vocabulary to publish the originally released tabular time series data and structure it into slices to support multiple views and query endpoints. We further describe how these linked open vocabularies have been used and combined in the context of this project to make this dataset linkable to existing or future linked open data resources. We also discuss the versatility of the new service for the consumers of the ACORNSAT dataset and uncover some issues which are specific to such long term climate data time series. The resulting Linked Sensor Data Cube is now accessible online via a pilot government linked data service built on the Linked Data API at lab.environment.data.gov.au.
Long-term sustainability of spatial data infrastructures: a metadata framework and principles of geo-archiving
Linked data offers a novel and more flexible means of sharing complex geospatial datasets by breaking away from the traditional domainspecific technologies used for accessing and integrating geospatial data with heterogeneous sources and disparate formats. In 2010, the UK Cabinet Office released a set of draft guidelines for exposing geospatial data as linked-data in support of the UK Open Data initiative. These draft guidelines have been proposed under the UK Location Strategy in specific recognition of the importance of geospatial data, and also with a view to promote linked-data within the EU INSPIRE community. This paper presents a customisable opensource linked-data framework developed by the GeoTOD-II project that implements these guidelines. The framework provides an efficient means for exposing both existing and new data sources in the linked-data form. We also attempt to articulate and address a number of issues and hidden assumptions with these guidelines identified during the development of the framework.
With growing concerns about environmental problems, and an exponential increase in computing capabilities over the last decade, the geospatial community has been producing increasingly voluminous and diverse environmental datasets. Long-term preservation of these environmental data exposed through uniform and interoperable Spatial Data Infrastructures (SDIs) is not typically addressed, but is highly important for applications that require continued access to both current and historical data, e.g., for monitoring climate change. The work presented in this article investigates the requirements for ensuring sustained access to environmental data from the perspective of a preservation-aware SDI. We take INSPIRE as an exemplar for our analysis and model development. In addition, we present an implementation approach in the form of a Geo-Portal that incorporates a preservation profile of the ISO 19115 metadata standard.
The Climate Science Modelling Language (CSML) was originally developed as part of the NERC Data Grid (NDG) project in the UK. It was one of the first Geography Markup Language (GML) application schemas describing complex feature types for the metocean domain. CSML feature types can be used to describe typical climate products such as model runs or atmospheric profiles. CSML has been successfully used within NDG to provide harmonised access to a number of different data sources. For example, meteorological observations held in heterogeneous databases by the British Atmospheric Data Centre (BADC) and Centre for Ecology and Hydrology (CEH) were served uniformly as CSML features via Web Feature Service. CSML has now been substantially revised to harmonise it with the latest developments in OGC and ISO conceptual modelling for geographic information. In particular, CSML is now aligned with the near-final ISO 19156 Observations & Measurements (O&M) standard. CSML combines the O&M concept of ’sampling features’ together with an observation result based on the coverage model (ISO 19123). This general pattern is specialised for particular data types of interest, classified on the basis of sampling geometry and topology. In parallel work, the OGC Met Ocean Domain Working Group has established a conceptual modelling ac- tivity. This is a cross-organisational effort aimed at reaching consensus on a common core data model that could be re-used in a number of met-related application areas: operational meteorology, aviation meteorology, climate studies, and the research community. It is significant to note that this group has also identified sampling geometry and topology as a key classification axis for data types. Using the Model Driven Architecture (MDA) approach as adopted by INSPIRE we demonstrate how the CSML application schema is derived from a formal UML conceptual model based on the ISO TC211 framework. By employing MDA tools which map consistently between UML and GML we can treat the formal UML model as the primary governed artefact and automatically produce the GML schema as a secondary output. Finally we describe how increased convergence between CSML and Scientific Feature Types in the Unidata Commmon Data Model may assist with bridging the implementation gap between OGC/ISO services and the CF-NetCDF binary data management community. This improved agreement at the conceptual (feature type) level is important to enable better interoperability at the data exchange and service levels.
The Metadata Objects for Linking Environmental Sciences (MOLES) model has been developed within the Natural Environment Research Council (NERC) DataGrid project [NERC DataGrid] to fill a missing part of the ‘metadata spectrum’. It is a framework within which to encode the relationships between the tools used to obtain data, the activities which organised their use, and the datasets produced. MOLES is primarily of use to consumers of data, especially in an interdisciplinary context, to allow them to establish details of provenance, and to compare and contrast such information without recourse to discipline-specific metadata or private communications with the original investigators [Lawrence et al 2009]. MOLES is also of use to the custodians of data, providing an organising paradigm for the data and metadata.
The use of a semantically rich registry containing a Feature Type Catalogue (FTC) to represent the semantics of geographic feature types including operations, attributes and relationships between feature types is required to realise the benefits of Spatial Data Infrastructures (SDIs). Specifically, such information provides a more complete representation of the semantics of the concepts used in the SDI, and enables advanced navigation, discovery and utilisation of discovered resources. The presented approach creates an FTC implementation in which attributes, associations and operations for a given feature type are encapsulated within the FTC, and these conceptual representations are separated from the implementation aspects of the web services that may realise the operations in the FTC. This differs from previous approaches that combine the implementation and conceptual aspects of behaviour in a web service ontology, but separate the behavioural aspects from the static aspects of the semantics of the concept or feature type. These principles are demonstrated by the implementation of such a registry using open standards. The ebXML Registry Information Model (ebRIM) was used to incorporate the FTC described in ISO 19110 by extending the Open Geospatial Consortium ebRIM Profile for the Web Catalogue Service (CSW) and adding a number of stored queries to allow the FTC component of the standards‐compliant registry to be interrogated. The registry was populated with feature types from the marine domain, incorporating objects that conform to both the object and field views of the world. The implemented registry demonstrates the benefits of inheritance of feature type operations, attributes and associations, the ability to navigate around the FTC and the advantages of separating the conceptual from the implementation aspects of the FTC. Further work is required to formalise the model and include axioms to allow enhanced semantic expressiveness and the development of reasoning capabilities.
The Heterogeneous Missions Accessibility (HMA) project is a joint activity of the European and Canadian Space Agencies lead by ESA through its Ground Segment Coordination Body (GSCB). It aims to provide a seamless and harmonised access to heterogeneous Earth observation (EO) datasets from multiple mission ground segments. To achieve this goal of interoperability, the HMA project is developing standardised metadata descriptions at collectionand product-level, as well as standardised network service interfaces for data discovery, ordering, planning, user management, and data access. These interfaces will be implemented in the EO Data Access and Integration Layer (DAIL), providing an integrated, harmonised access across multiple mission ground segments.
There is remarkable agreement in expectations today for vastly improved ocean data management a decade from now --capabilities that will help to bring significant benefits to ocean research and to society.Advancing data management to such a degree, however, will require cultural and policy changes that are slow to effect.The technological foundations upon which data management systems are built are certain to continue advancing rapidly in parallel.These considerations argue for adopting attitudes of pragmatism and realism when planning data management strategies.In this paper we adopt those attitudes as we outline opportunities for progress in ocean data management.We begin with a synopsis of expectations for integrated ocean data management a decade from now.We discuss factors that should be considered by those evaluating candidate "standards".We highlight challenges and opportunities in a number of technical areas, including "Web 2.0" applications, data modeling, data discovery and metadata, real-time operational data, archival of data, biological data management and satellite data management.We discuss the importance of investments in the development of software toolkits to accelerate progress.We conclude the paper by recommending a few specific, short term targets for implementation, that we believe to be both significant and achievable, and calling for action by community leadership to effect these advancements.
As a second step, data specifications for the first set of themes has been developed based on the modelling framework. The themes include addresses, transport networks, protected sites, hydrography, administrative areas and others. The data specifications were developed by selected experts nominated by stakeholders from all over Europe. For each theme a working group was established in early 2008 working on their specific theme and collaborating with the other working groups on cross-theme issues. After a public review of the draft specifications starting in December 2008, an open testing process and thorough comment resolution process, the draft technical implementing rules for these themes have been approved by the INSPIRE Committee. After they enter into force they become part of the legal framework and European Member States have to implement these rules.
The ultimate goal of much current research in earth science informatics is to enable more efficient discovery and use of environmental data. Large-scale efforts are underway at regional and global levels. For instance the European INSPIRE Directive (2007/2/EC) and international GEOSS initiative will both provide unprecedented catalogues of earth observation and environmental data, with links to online services providing direct access to digital data repositories. While the motivation for these emerging infrastructures is clear (e.g. understanding global change), it is less obvious how they might be implemented. Standards will play a major role and considerable effort is currently being devoted to their development by bodies like the International Organisation for Standardisation and the Open Geospatial Consortium. Internet search engines are amongst the most popular websites visited today. Using the metaphor of a web search portal, we review the potential of new geospatial standards to provide an advanced, user-friendly approach to discovery and use of climate-science data.