Archaeological data repositories usually manage excavation data collections as project-level entities with restricted capacities to facilitate search or aggregation of excavation data at the sub-collection level (trenches, finds, season reports or excavation diaries etc.). More granular access to excavation data collections would enable layered querying across their informational content. In the past decade, several attempts to adapt CIDOC CRM in order to provide more explicit descriptions of the excavation universe have resulted in the use of domain-specific model extensions (e.g. CRMarchaeo, CRMsci, CRMba). Each focuses on corresponding aspects of the excavation research process, while their combined usage has potential to support expressive data mappings at the sub-collection level. As part of the ARIADNEplus project, several CIDOC CRM developers and domain experts have collaborated to undertake conceptual mapping exercises, to address the practicalities of bringing excavation data descriptions together and to link these to our overall aspirations in terms of excavation data discoverability and reusability. In this contribution, we discuss the current state and future directions of the field of semantic representation of archaeological excavation data and consider several issues that constrain the applicability of existing solutions. We identify five key enabling technologies or research areas (Conceptual models and semantic data structures, Conceptual modelling patterns, Data mapping workflows and tools, Learning technologies and Semantic queries) and assign readiness levels to assess their level of technological maturity. Our research demonstrates that while the existing models and domain-specific extensions are deemed adequate, there is a need for more user-friendly methods and tools to structure data in meaningful and interoperable ways. The next steps involve consolidating relevant semantic structures, improving modelling implementation guidance, adhering to consistent workflows, developing engaging curricula, and documenting real-case examples to demonstrate the benefits and results of semantic data integration.
Language is a common good and a common property. Access to information about language should be fast, easy, and intuitive. The electronic dictionary should therefore be a knowledge base with language as its access point, and with simple, yet rich access to (combinations of) linguistic and non-linguistic facts. One query frame and basic reading and writing skills must be enough to get meaningful results. This solution presupposes (1) a fine grained and systematic database format for dictionary storage and linkage to materials, and (2) a query system offering ease of access for inexperienced users. At the same time, lexicography must be able to prove itself trustworthy by offering access to sources both for usage and for normative decisions. The system described here is used for one academic multivolume dictionary and for standard monolingual students' dictionaries. It is suited to lexicographical projects where source documentation has priority. The focus is on dictionaries integrated with other language resources and produced for the Web.
This paper describes the background and methods for the production of CIDOC-CRM compliant data sets from diverse collections of source data. The construction of such data sets is based on data in column format, typically exported for databases, as well as free text, typically created through scanning and OCR processing or transcription.
The content in information systems and virtual reconstructions in the cultural heritage sector is to a large degree directly based on information deduced from the study of texts. In many cases, even if the texts are available electronically, the links from the deduced facts to the original texts are not available and in many cases very costly to re-establish. Reproducibility of results is a core concept in text-based research as in all research. Thus, such links should be expressed explicitly in the systems and in accordance with the data standards developed in the fields of text encoding and conceptual modelling. To do this it is necessary to create a combined understanding of text encoding represented by the TEI guidelines and the understanding of conceptual models represented by initiatives like the CIDOC CRM and FRBRoo. In this article, we study a part of this complex by comparing the expressive power of the real world descriptions TEI P5 by mapping central parts of the CIDOC CRM onto TEI P5. It is clear that the TEI P5 has moved a great step in the direction towards an event-oriented model compared with TEI P4. Our use of CIDOC CRM as a yardstick shows that the expressiveness of TEI P5 can be greatly improved by extending the scope of very restricted elements like the relation element and adding a few new elements to the TEI.
In historically oriented research like archaeology, the determination of the chronology of events in the past plays an important role. For example, a fire of a house can seal off the layers physically below and give a partial relative dating of these. A well known tool in this area is the Harris’ Matrix used to systematize the contexts and layers found in an excavation. In this paper we will discuss a related but more general tool for documenting and analysing temporal entities like events. This tool is developed as a module of a four dimensional event-oriented documentation database based on the conceptual model CIDOC-CRM (ISO21127). The database is developed for an archaeological excavation project in Western Norway. In addition to places, events and actors the database is designed to contain texts, images and maps used to document such entities. In use the system will contain a dataset of events, their time-spans and relations between events. The system can detect conflicting dating, increase precision of starts, ends and durations of events and finally display a spatial and chronological overview. Given a time and a place within the dataset, the system can display all possible chronologies for the events in the set. So far, this tool has shown a great potential being used in projects involving large amounts of archive material as preparation for new excavations. Further development includes the possibility to use other temporal constraints, such as durations and exploring the potential of adding spatial constraints and constraints on actors.
The tutorial first addresses requirements and semantic problems to integrate digital information into large scale, meaningful networks of knowledge that support not only access to source documents but also use and reuse of integrated information. The pros and cons of developing global ontologies are discussed. It is argued that core ontologies of relationships are fundamental to schema integration and play a completely different role to that of specialist terminologies in practical knowledge management. The CIDOC Conceptual Reference Model (CRM) is presented as an example of such a global model. It is a core ontology and new ISO standard (ISO 21127, accepted September 2006), originally designed for the semantic integration of information from museums, libraries, and archives. It is a product of re-engineering the dominant underlying common concepts from representative data structures. It is not prescriptive, but provides a controlled language to describe common high-level semantics that allow for information integration at the schema level. The tutorial addresses part of the technology needed for information aggregation and integration in the global information environment, namely the question to which extent and in which form global schema integration is feasible. The ability of the CRM to support integration has been demonstrated in a large range of different domains including cultural heritage, e-science and biodiversity. Conceptual modeling by specializing such a well-tested core ontology not only reduces drastically development time and improves system quality, but provides basic semantic interoperability more or less for free. The tutorial will present characteristic applications.
Collection management systems have been developed and used in memory institutions like museums in many decades. In the last decade there has been an increasing understanding for the necessity of interconnecting such systems both on the local and the global levels. At the international level one finds initiatives like the European Digitial Library (EDL) and in many countries there are similar initiatives on the national or regional level, see GAUSDAL (2006). From a technical point of view there are few problems connected to the implementation of search portals enabling simulations searches in many cultural heritage databases and many such portals have been launched in the recent years. However, the actual interconnection of cultural heritage databases often reveals more serious problems. Even inside homogenous language areas integration is made difficult due to the use of different classification systems, and even when the same systems are used, they are often used differently. In addition to the more organised classification systems and nomenclatures, one has the large problem of different names denoting the same real world entity and identical names denoting different entities. Thus integration is non trivial both on the micro and macro level and is a well know head ache for all museum documentalists. The solution to the name part of the integration problem is co-reference, where pairs of expressions, e.g. names, are specified to point to the same real-world thing. Co-referencing is a very difficult task to solve automatically, and a very time-consuming manual task. But scholars, researchers, museum curators and library cataloguers continuously trace co-references in their daily work. This extremely time-consuming work can be made more efficient by sharing the knowledge. Such sharing would complement centralized approaches found in present Digital Library Systems. This concern led CIDOC to establish a working group for Co-reference in 2007. A good example on the development of co-reference links between a text collection and a museum database can be found in recent work from the Perseus Project (Babeu 2007). In this paper we discuss the name and identifier part of integration problem (co-referencing) and some practical solution a part based on experiences from the Nordic countries. Our discussion is mainly based on experiences drawn from the distributed search pilot of large Swedish KMM project (Knowledge in Management in Museums) which has revealed most of the integration problems mentioned above. The system is CIDOC CRM bases and is now extended to 5 Nordic countriesThe paper will present experience from a test implementation of a distributed co-reference assistant. Since the cultural history of the Nordic countries is highly interwoven and information
n the last couple of years, there has been a growing interesttowards including into TEI documents information aboutthe world rather than information concerning the text of thedocument to be encoded only .W e see examples of this throughrecent additions to the TEI standard, e.g. the person element(TEI P5, sec. 20.4.2), as well as through the work in theOntologies SIG since itwas established in 2004 (TEI OntologySIG WIKI). In the SIG, the topic of discussion is how toorganise this kind of information about the world according tospecific ontologies.One particularly promising ontology in this context is the CRM(CIDOC 2003). It has been used together with T opic Maps toorganize information from TEI documents (T uohy 2006) andas an attempt to find a solution to the so-called exhibitionproblem (Eide 2006). Further , attempts have been made toformalise a way to connect TEI and CRM documents (Ore2006).In this paper ,we propose a method for automatic generation ofCRM conforming models based on TEI documents. W e willdiscuss limitations to this approach, as well as ways these maybe overcome.