Trade classifications are a necessary prerequisite for the compilation of trade statistics, and they should – beyond that – be regarded as a valuable base for the definition of shared controlled vocabularies for linked business data that deal with import, export etc. The Standard International Trade Classification (SITC) provided by the UN Statistics Division is a widely used classification mostly applied for scientific and analytical purposes. SITC – as most other trade classifications – is available today only in text or spreadsheet formats. These formats reveal the inner hierarchical structure of SITC to the human reader, because SITC trade codes are built according to the decimal classification scheme, but unfortunately, SITC’s inner structure is opaque to computer applications in text and spreadsheet formats. The paper discusses an approach to set up an OWL-2 ontology for SITC that states subsumption relations between classes of goods. This kind of semantic underpinning of SITC is suited to ease both checking and extending SITC and to derive from it a shared controlled vocabulary for business linked data. Some problems of today’s SITC (among them missing inner nodes of the trade code hierarchy) are carefully discussed, and the paper motivates several decisions that were taken for ontology design. Finally, the study introduces the semantic reasoner as a tool for the (at least partial) automatic derivation of structural information for SITC from the trade code building rule. The paper reports on reasoner runtimes observed for different versions of the SITC ontology and for different versions of the Pellet reasoner.
research-article Share on Enhancing Human-Transcribed Records by Using OCR Authors: Jesper Zedlitz Dept. of Computer Science, Christian-Albrechts-Universität, Kiel, Germany Dept. of Computer Science, Christian-Albrechts-Universität, Kiel, GermanyView Profile , Norbert Luttenberger Dept. of Computer Science, Christian-Albrechts-Universität, Kiel, Germany Dept. of Computer Science, Christian-Albrechts-Universität, Kiel, GermanyView Profile Authors Info & Claims DATeCH2017: Proceedings of the 2nd International Conference on Digital Access to Textual Cultural HeritageJune 2017 Pages 21–26https://doi.org/10.1145/3078081.3078094Published:01 June 2017Publication History 0citation50DownloadsMetricsTotal Citations0Total Downloads50Last 12 Months4Last 6 weeks1 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access
research-article Share on 750 Volunteers Transcribing 31,000 Pages with 8.5 million Entries Online: an Evaluation Authors: Jesper Zedlitz Dept. of Computer Science, Christian-Albrechts-Universität, Kiel, Germany Dept. of Computer Science, Christian-Albrechts-Universität, Kiel, GermanyView Profile , Norbert Luttenberger Dept. of Computer Science, Christian-Albrechts-Universität, Kiel, Germany Dept. of Computer Science, Christian-Albrechts-Universität, Kiel, GermanyView Profile Authors Info & Claims DATeCH2017: Proceedings of the 2nd International Conference on Digital Access to Textual Cultural HeritageJune 2017 Pages 3–8https://doi.org/10.1145/3078081.3078086Published:01 June 2017Publication History 0citation68DownloadsMetricsTotal Citations0Total Downloads68Last 12 Months7Last 6 weeks3 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access
Trade classifications are a necessary prerequisite for the compilation of trade statistics, and they should—beyond that—be regarded as valuable base for the definition of shared controlled vocabularies for linked business data that deal with import, export etc. The Standard International Trade Classification (SITC) provided by the UN Statistics Division is a widely used classification mostly applied for scientific and analytical purposes. SITC—as most other trade classifications—is available today only in text or spreadsheet formats. These formats reveal the inner hierarchical structure of SITC to the human reader, because SITC trade codes are built according to the decimal classification scheme, but unfortunately, SITC's inner structure is opaque to computer applications in text and spreadsheet formats. In this paper, we discuss an approach to set up an OWL-2 ontology for SITC that states subsumption relations between classes of goods. This kind of semantic underpinning of SITC is suited to ease both checking and extending SITC and to derive from it a shared controlled vocabulary for business linked data. We carefully discuss some problems of today's SITC (among them missing inner nodes of the trade code hierarchy), and we motivate several decisions that we took for ontology design. Finally, we introduce the semantic reasoner as a tool for the (at least partial) automatic derivation of structural information for SITC from the trade code building rule. We report on reasoner runtimes observed for different version of the SITC ontology and for different versions of the Pellet reasoner.
We present a web-based collaborative system to transcribe serial historic sources to structured data. The web-based system runs completely in the web browser without additional plug-ins. Its key feature is the fact that data entry is performed directly on scanned images: The scan is used as background image of the browser window; it is overlaid by the entry mask as well as text boxes with already transcribed data. The system has been successfully used to transcribe more than 31,000 pages from the German WW1 casualty lists to structured data resulting in more than 8.5 million entries.
These data are the contents of the document 55177 in the City Archives of Kiel "Todesopfer durch Luftangriffe". The document lists civilian casualties of the bombing on Kiel during WW2 . The document was brought in electronic form by volunteers using the data-entry system (DES) of the Verein für Computergenealogie (=society for computer Genealogy) . As part of a master's thesis at the Institute of computer science of the University of Kiel coordinates were determined for addresses. The following fields are included in the data: page family name given name(s) day of death remarks day of birth occupation place of death place of residence The following details are not in the source. They were either calculated (age), entered manually (gender) or are the result of geocoding (coordinates) and subsequent matching with district boundaries. age gender place of residence longitude place of residence latitude place of death longitude place of death latitude place of residence district place of death district
Identifying and referencing places is important for many fields of research. Very different approaches of how to represent administrative structures on the Semantic Web can be found. This survey attempts to provide a broad overview of systems that work on (historic) administrative information. We present a classification for such systems, with special attention to the difference that arise from the processing of historic data. We also describe a sample of systems which approach the problem in very different ways. We conclude by evaluating which of the presented characteristics make a system universal and futureproof. Keywords—conceptual modelling; administrative affiliation; semantic web; linked open data; historical information
Both OWL-2 and UML static class diagrams lend themselves very well for conceptual modelling of complex information systems. Both languages have their advantages. In order to benefit from the advantages and software tools of both languages, it is usually necessary to repeat the modelling process for each language. We have investigated whether and how conceptual models written in one language can be automatically transformed into models written in the other language. For this purpose we investigated differences and similarities of various model elements (such as element type, data types, relationship types) in static UML data models and OWL-2 ontologies. We provide a transformation for similar elements. Keywords—UML; OWL; conceptual modelling; model transformation; meta modelling
Both OWL-2 and UML static class diagrams lend themselves very well for conceptual modeling of complex information systems. To ease the choice between either of these languages it worthwhile to clarify the differences and similarities in the representation of different kinds of datatypes (primitive types, enumerations, complex datatypes, and generalization of datatypes) in static UML data models and OWL-2 ontologies. Where similarities allow a transformation of datatypes from one language into the other, we describe a possible transformation. Keywords—UML; OWL; datatypes; conceptual modeling
Für eine erfolgreiche semantische Suche in großen Datenbeständen ist in den meisten Fällen ein umfangreiches Hintergrundwissen erforderlich. Dieses Hintergrundwissen bezieht sich– grob gesehen– auf die in einer domain of discourse verwendeten Konzepte („Klassen“), ihren Zusammenhang zueinander („Axiome“ zur Beschreibung von Klasseneigenschaften einschließlich der Beschreibung von Unterklassen-Beziehungen), die Eigenschaften der Klasseninstanzen („properties“) und – in vielen Fällen – auf eine Menge von namentlich bezeichneten Individuen, die im betrachteten Gegenstandsbereich eine besondere Rolle spielen. Dieses Hintergrundwissen wird in vielen Diskursgemeinschaften in formalenOntologien zusammengefasst, die z. B. in der Web Ontology Language 2 (OWL-2) aufgeschrieben sind. Dank der Abstützung auf formaler Logik (speziell: auf description logics) kann in solchen Ontologien implizit enthaltenes Wissen durch automatisches Schlussfolgern („reasoning“) sichtbar gemacht werden. Vor diesem Hintergrund hat sich eine neue Disziplin herausgebildet, die als „Ontology Engineering“ bezeichnet wird. Es geht um die Konstruktion von Ontologien mit ingenieurmäßigen Methoden. In diesem Aufsatz wollen wir im Sinne des Ontology Engineering zeigen, wie sich formale, OWL-2-basierte Ontologien mit Hilfe bestimmter Metriken charakterisieren lassen. Wir gehen dabei nur auf den terminologischen Teil von Ontologien („Schema“) ein, da die in einer Ontologie enthaltenen Individuen nur wenig über die Konstruktionsprinzipien einer Ontologie aussagen können. Der Aufsatz ist wie folgt aufgebaut: Zunächst zeigen wir, für wen und welche Anwendungsfälle OntologieMetriken sinnvoll sind. Es folgt eine Zusammenstellung, welche Artikel bisher zum Thema Ontologie-Metriken veröffentlicht wurden. Der anschließende Hauptteil stellt bekannte und neue Metriken vor. Dabei wird auf Schwierigkeiten eingegangen, die sich u. a. bei der Anwendung existierender Ontologie-Metriken für OWL-2 ergeben. Im folgenden Abschnitt wenden wir einige der vorgestellten Metriken auf eine Menge von aktuellen Ontologien an, vergleichen diese wenn möglich mit älteren Ergebnissen und weisen auf interessante Messergebnisse hin.
"A childs year of birth is always greater than the year of birth of its parents." - it is not easily possible to code this simple knowledge into a pure OWL ontology, i.e. without using any additional rule languages. Therefore it is not easy in OWL to detect semantic violations in this kind of statements. The two challenges are putting two orders ("greater" and "parent") into relation and representing integers as individuals allowing a reasoner to infer knowledge about the "greater" relation. In the first part of this contribution we show a pattern for putting two transitive and asymmetric orders into a relation, such that conflicting information results in an inconsistent ontology. In the second part we present a pattern for expressing integers using their binary code. Due to the special construction a reasoner can infer knowledge about the relation between all integers in the ontology. By combining the two patterns we are able to represent the initial statement in an ontology.
In this paper we present a transformation between UML class diagrams and OWL 2 ontologies. We specify the transformation on the M2 level using the QVT transformation language and the meta-models of UML and OWL 2. For this purpose we analyze similarities and differences between UML and OWL 2 and identify incompatible language features.
The ISO 19103 standard—defining rules and guidelines for conceptual modeling in the geographic domain—has deliberately chosen the Unified Modeling Language (UML) as “conceptual schema language” for geographic information systems. From today’s perspective—i.e. when taking into account today’s mature semantic web technology—another language might also be envisioned as language for specifying applicationoriented conceptual models, namely the Web Ontology Language OWL 2. Both language definitions refer to comparable meta-models laid down in terms of OMG’s Meta Object Facility, but in contrast to UML, OWL 2 is fully built upon formal logic which allows logical reasoning on OWL 2 ontologies. In this paper, we investigate language similarities and differences by specifying and implementing the transformation on the meta-model level using the QVT transformation language.
IEEE-1451[1] and OGC Sensor Web Enablement (OGC SWE)[2] define standard protocols to operate instruments, including methods to calibrate, configure, trigger data acquisition, and retrieve instrument data based on specified temporal and geospatial criteria. These standards also provide standard ways to describe instrument capabilities, properties, and data structures produced by the instrument. These standard operational protocols and descriptions enable observing systems to manage very diverse instruments as well as to acquire, process, and interpret their data in a uniform and automated manner. We refer to this property as “instrument interoperability”. This paper describes integration and evaluation of MBARI PUCK protocol [3] within different observatories including OBSEA [4,5] in Spain, the ESONET test-bed in Germany, and the SmartBay observatory in Canada.
Many sensor networks have been deployed to monitor Earth's environment, and more are planned for the future. Environmental sensors have continuously improved by becoming smaller, cheaper, more intelligent, and more reliable. But due to the large number of sensor manufacturers and accompanying protocols, integrating diverse sensors into observing systems is not straightforward, requiring development of driver software and manual tedious configuration. Use of standard protocols and formats can improve and automate the process of sensor installation, operation, and data processing. The Open Geospatial Consortium's Sensor Web Enablement (SWE) initiative defines standards which make sensors available over the Web through standardized formats and Web Service interfaces by hiding the heterogeneity of sensor protocols from the application layer. Current SWE standards do not deal with actual sensor protocols, and the connection between sensors and SWE services is usually established by manually adapting the internals of the SWE service implementation to the specific sensor interface. Such sensor drivers have to be built for each kind of sensor interface, which leads to extensive efforts in developing large-scale systems. To tackle this issue we have developed a model for Sensor Interface Descriptors (SID) which enables the declarative description of sensor interfaces, including the definition of the communication protocol, sensor commands, processing steps and metadata association. The model is designed as a profile and extension of OGC SWE's Sensor Model Language standard. In this model, a SID is defined in XML for each kind of sensor protocol. SID instances for particular sensor types can be reused in different scenarios and can be shared among user communities. A SID interpreter can be built which translates between various sensor protocols and SWE protocols, hence closing the described interoperability gap. The SID interpreter is independent of any particular sensor technology, and can communicate with any sensor whose protocol can be described by a SID. The SID interpreter transfers retrieved sensor data to a Sensor Observation Service, and transforms tasks submitted to a Sensor Planning Service to actual sensor commands. The proposed SWE PUCK protocol complements SID by providing a standard way to associate a sensor with a SID, thereby completely automating the sensor integration process. PUCK protocol is implemented in sensor firmware, and provides a means to retrieve a universally unique identifer, metadata and other information from the device itself through its communication interface. Thus the SID interpreter can retrieve a SID directly from the sensor through PUCK protocol. Alternatively the interpreter can retrieve the sensor’s SID from an external source, based on the unique sensor ID provided by PUCK protocol. In this presentation, we describe the end-to-end integration of several commercial oceanographic instruments into a sensor network using PUCK, SID and SWE services. We also present a user-friendly, graphical tool to generate SIDs and tools to visualize sensor data