Meteorological data, essential in a variety of applications, has been made available as open data through different portals, either governmental, associative or private ones. Making this data fully findable and reusable for experts from other domains than meteorology requires considerable efforts to guarantee compliance to the FAIR principles. Nowadays, most efforts in data FAIRification are limited to semantic metadata describing the overall features of data sets. However, such a description is not enough to fully address data interoperability and reusability by other scientific communities. This paper addresses this weakness by proposing a semantic model to represent different kinds of metadata, describing the data schema and the internal structure of a data set distribution, together with domain-specific definitions. This model is used to provide a reusable schema of the SYNOP data set, a largely used governmental meteorological data set in France. The impact of using the proposed model for improving FAIRness was evaluated.
Résumé Rendre les jeux de données météorologiques FAIR est un enjeu crucial pour la recherche scientifique. Le mo-dèle dmo-core permet, pour les données tabulaires, d’expliciter la sémantique des colonnes en les reliant à des concepts d’ontologies de domaine. Ces concepts étant généralement peu ou pas documentés, nous pro-posons (1) d’enrichir dmo-core afin de pouvoir as-socier aux colonnes leurs définitions issues d’une res-source lexicale, et (2) de générer une telle ressource, à partir du Vocabulaire Météorologique International de l’Organisation Mondiale de Météorologie (INMEVO).
Open data is exposed in several formats, including tabular format. However, the meaning of columns, that can also be seen as dimensions, is not always explicit what makes difficult the reuse of this data for data consumers. This paper presents the FAIRification process of tabular and multidimensional datasets that relies on a (FAIR) core semantic model that is able to represent different kinds of metadata, including the data schema and the internal structure of a dataset. We describe how the instantiation of such a model offers in addition the possibility to describe the semantics of columns using domain ontologies. Once instantiated, this model forms a set of formal metadata that documents the dataset and facilitates understanding by data consumers. This process is then applied to three metereological datasets, for which the degree of improvement of the FAIRness (“I” and “R”) has been evaluated.
Meteorological institutions produce a valuable amount of data as a direct or side product of their activities, which can be potentially explored in diverse applications. However, making this data fully reusable requires considerable efforts in order to guarantee compliance to the FAIR principles. While most efforts in data FAIRification are limited to describing data with semantic metadata, such a description is not enough to fully address interoperability and reusability. We tackle this weakness by proposing a rich ontological model to represent both metadata and data schema of meteorological data. We apply the proposed model on a largely used meteorological dataset, the "SYNOP" dataset of Météo-France and show how the proposed model improves FAIRness.
Tabular format is a common format in open data. However, the meaning of columns is not always explicit which makes if difficult for non-domain experts to reuse the data. While most efforts in making data FAIR are limited to semantic metadata describing the overall features of datasets, such a description is not enough to ensure data interoperability and reusability. This paper proposes to reduce this weakness thanks to a (FAIR) core semantic model that is able to represent different kinds of metadata, including the data schema and the internal structure of a dataset. This model can then be linked to domain-specific definitions to provide domain understanding to data consumers.
Résumé Rendre les données météorologiques FAIR pour faciliter leur réutilisation est un enjeu stratégique car ce sont des données essentielles à la recherche scientifique dans de nombreux domaines. Cet article propose un modèle sémantique associant un modèle de métadonnées et un modèle de données pour décrire les données météorologiques d’observation. En effet, la modélisation des (méta)données est une étape essentielle vers leur FAIRisation. Nous utilisons le jeu de données "SYNOP" de Météo-France pour illustrer les difficultés liées à l’accès et à la compréhension de ce type de données, et pour montrer comment le modèle proposé améliore leur adhésion aux principes "F", "I", et "R". Mots-clés Données météorologiques, principes FAIR, métadonnées sémantiques. Abstract Making meteorological data FAIR in order to ease its reuse is a strategic issue because this data is essential to advance research in many fields. This work proposes a semantic model which combines a metadata model and a data model for describing meteorological observation data. Indeed, mode-ling (meta)data is an essential step towards their FAIRification. We use the SYNOP open dataset made available by Météo-France to illustrate how difficult data access and un-derstanding can be, and how the use of the proposed model to represent meteorological data improves their compliance with the "F", "I" and "R" principles.
SYNOP dataset is one of the open meteorological datasets privided by Meteo-France on its data portal. The dataset includes observation data from international surface observation messages circulating on the Global Telecommunication System (GTS) of the World Meteorological Organization (WMO) . The choice of this dataset is motivated by the fact that these data are open and free, they concern several atmospheric parameters measured (temperature, humidity, wind direction and force, atmospheric pressure, precipitation height, etc.). These parameters are important for many scientific studies, particularly because SYNOP dataset are published by all WMO member states. SYNOP data is published as open data.
Business process (BP) modelling is an active area of research due to its multiple applications.For systems that support/monitor operators to perform their tasks (i.e., tasks of a given BP), a formal representation is essential.Various BP ontologies are available to formally represent BP.In this paper, we review and compare a set of nine BP ontologies according to their ability to represent process specification and process execution in a finegrained way to enable task monitoring.The comparison shows that, on the one hand, ontologies developed from scratch establish a clear distinction between process specification and process execution, but do not allow to represent workflow constraints required for process execution.On the other hand, most of the ontologies, that are ontological versions of existing BP modeling languages, focus only on process specifications but do not represent process execution, or mix the representation of BP specification and execution.
Any industrial company has its own business processes, which is a number of related tasks that have to be executed to reach well-defined goals. In order to analyze, improve, simulate and automate these processes, it is essential to represent them in a formal way. The activity of representing business processes is known as Business Process Modelling (BPM); it is an active research area that attracts more and more attention with the emergence of Industry 4.0. Semantic Web technologies, especially ontologies, are promising means to advance BPM and to realize the Industry 4.0 vision. In this scope, we developed the BBO (BPMN 2.0 Based Ontology) ontology for business process representation, by reusing existing ontologies and meta-models like BPMN 2.0, the state-of-the-art meta-model for business process representation. We evaluated BBO using schema metrics, which showed that it was a deep and rich ontology with a variety of relationships. Thanks to a use case, we illustrated the ability of BBO to represent real business processes in a fine-grained way and to express and answer the competency questions identified at the specification stage.
Any industrial company has its own business processes, which is a number of related tasks that have to be executed to reach well-defined goals. In order to analyze, improve, simulate and automate these processes, it is essential to represent them in a formal way. The activity of representing business processes is known as Business Process Modelling (BPM); it is an active research area that attracts more and more attention with the emergence of Industry 4.0. Semantic Web technologies, especially ontologies, are promising means to advance BPM and to realize the Industry 4.0 vision. In this scope, we developed the BBO (BPMN 2.0 Based Ontology) ontology for business process representation, by reusing existing ontologies and meta-models like BPMN 2.0, the state-of-the-art meta-model for business process representation. We evaluated BBO using schema metrics, which showed that it was a deep and rich ontology with a variety of relationships. Thanks to a use case, we illustrated the ability of BBO to represent real business processes in a fine-grained way and to express and answer the competency questions identified at the specification stage.
Automatic construction of semantic resources at large scale usually relies on general purpose corpora as Wikipedia. This resource, by nature rich in encyclopedic knowledge, exposes part of this knowledge with strongly structured elements (infoboxes, categories, etc.). Several extractors have targeted these structures in order to enrich or to populate semantic resources as DBpedia, YAGO or BabelNet. The remain semi-structured textual structures, such as vertical enumerative structures (those using typographic and dispositional layout) have been however under-exploited. However, frequent in corpora, they are rich sources of specific semantic relations, such as hypernyms. This paper presents a distant learning approach for extracting hypernym relations from vertical enumerative structures of Wikipedia, with the aim of enriching DBpedia. Our relation extraction approach achieves an overall precision of 62%, and 99% of the extracted relations can enrich DBpedia, with respect to a reference corpus.
Extraire des relations d'hyperonymie a partir des textes est une des etapes cles de la construction automatique d'ontologies et du peuplement de bases de connaissances. Plusieurs types de methodes (linguistiques, statistiques, combinees) ont ete exploites par une variete de propositions dans la litterature. Les apports respectifs et la complementarite de ces methodes sont cependant encore mal identifies pour optimiser leur combinaison. Dans cet article, nous nous interessons a la complementarite de deux methodes de nature differente, l'une basee sur les patrons linguistiques, l'autre sur l'apprentissage supervise, pour identifier la relation d'hyperonymie a travers differents modes d'expression. Nous avons applique ces methodes a un sous-corpus de Wikipedia en francais, compose des pages de desambiguisation. Ce corpus se prete bien a la mise en oeuvre des deux approches retenues car ces textes sont particulierement riches en relations d'hyperonymie, et contiennent a la fois des formulations redigees et d'autres syntaxiquement pauvres. Nous avons compare les resultats des deux methodes prises independamment afin d'etablir leurs performances respectives, et de les comparer avec le resultat des deux methodes appliquees ensemble. Les meilleurs resultats obtenus correspondent a ce dernier cas de figure avec une F-mesure de 0.68. De plus, l'extracteur Wikipedia issu de ce travail permet d'enrichir la ressource semantique DBPedia en francais : 55% des relations identifiees par notre extracteur ne sont pas deja presentes dans DBPedia.
Nathalie Aussenac-Gilles合作论文数CNRS - IRIT33
Nathalie Hernandez合作论文数Departement de Mathematiques-Informatique, Universite Toulouse le Mirail1