The work in this paper is motivated from two different perspectives: First, gazetteers as an important data source for Geographic Information Retrieval (GIR) applications often lack historic place name information. More focused historic gazetteers are a far cry from being complete and often specialize only on certain geographic regions or time periods. Second, research on historic route descriptions---so called itineraries---is an important task in many research disciplines such as geography, linguistics, history, religion, or even medicine. This research on historic itineraries is characterized by manual, time-consuming work with only minimalistic IT support through gazetteers and map services. We address both perspectives and present a depth-first branch-and-bound (DFBnB) algorithm for deducing historic place names and thus the stops of ancient travel routes from itinerary tables. Multiple phonetic and character-based string distances are evaluated when resolving parts of an itinerary first published in 1563.
Teaching an information retrieval (IR) course for three different target groups is challenging. These target groups are i) resident students at the University of Bamberg, Germany ii) remote full-time students from other Bavarian universities, and iii) remote but part-time students from all parts of Germany which usually study the course besides their day-to-day work relationship. As a consequence, we have participants with heterogeneous previous knowledge and potentially different expectations with respect to the course content. In this paper we will only briefly describe the didactic aspects and challenges of the course. The clear focus of the paper lies on an in-depth quantitative evaluation how various IR topics presented in the lectures and tutorials fit the needs of the different target groups. We evaluate interestingness and importance of the course topics according to the students' impressions. The goal of the evaluation is to get hints for future refinements of IR course content---in general but also with respect to the needs of different target groups.
Die Itinerarforschung beschäftigt sich mit der Erschließung historischer Straßennetze. Sie differenziert zwei Itinerararten. Als Itinerare werden einerseits historische Reisewege hochstehender Personen und Herrscher bezeichnet. Diese Reisewege wurden und werden meist anhand historischer Dokumente und Urkunden rekonstruiert, die Auskunft darüber geben, zu welcher Zeit sich gewisse Personen an bestimmten Orten aufgehalten haben. Andererseits bezeichnen Itinerare auch Reisewegverläufe, einzeln oder in Form von Sammlungen, die unmittelbar als solche zusammengetragen wurden (Szabó 2009: 85). Die Erforschung historischer Itinerare ist ein wichtiger Arbeitsschwerpunkt in verschiedenen Wissenschaftsdisziplinen. Dies ist in Teilen dadurch bedingt, dass historische Personen, die die Itinerare entweder direkt erstellt haben oder auf Basis deren Vita die Itinerare durch Dritte erstellt wurden, häufig in verschiedenen Rollen unterwegs waren. So tritt etwa Hieronymus Münzer auf seiner Spanienund Frankreichreise gleichzeitig als Arzt, Historiker, Kaufmann, Pilger und Geograph in Erscheinung (Hurtienne 2009: 268). Nicht zuletzt deshalb ist die „Altwegeforschung“ ein stark interdisziplinäres Forschungsfeld (Veling 2014). Charakteristisch für die Itinerarforschung ist eine manuell geprägte und zeitaufwändige Arbeitsweise. Ein wesentlicher Aspekt bei der Erschließung ist etwa die Identifizierung der in den Itineraren genannten Orte (Hurtienne 2009: 269). Der Ansatz, der in dieser Arbeit beschrieben wird, versucht Werkzeuge zu entwerfen, die Forscher_innen in der Itinerarund Altwegeforschung in verschiedenen Wissenschaftsbereichen unterstützen. Ziel ist es, die zeitaufwändige, manuelle Erschließung der Itinerare effizienter zu gestalten und später auch den Vergleich verschiedener Itinerare im großen Stile zu fördern. Außerdem soll es ermöglicht werden, leichter Fehler und Inkonsistenzen in den Itinerarquellen zu identifizieren. Darüber hinaus soll die Erweiterung von Ortsverzeichnissen, insbesondere um historische Informationen, erleichtert werden. Ortsverzeichnisse, sog. Gazetteers, sind häufig unvollständig und lückenhaft, insbesondere wenn es um historische Informationen geht. Ferner beschränken sich historische Gazetteers häufig auf bestimmte geografische Gebiete und/oder Zeitperioden. Für die Anreicherung der Gazetteers stellen Itinerare eine wesentliche Datenbasis dar, aus der sich computerunterstützt mit Hilfe des hier skizzierten Ansatzes wichtige Informationen ableiten lassen. Während Blank und Henrich (2015) bereits die grundlegende Idee und eine Abgrenzung gegenüber verwandten, technischen Arbeiten im Geografischen Information Retrieval adressieren, beleuchtet die vorliegende Arbeit insbesondere die Anwendbarkeit und Einsatzmöglichkeiten des Ansatzes.
This paper is concerned with automatic georeferencing of river networks from raster images such as aerial photos or maps. Determining reasonable assignments between a given network of rivers derived from a textual description and an image is subject to high combinatorial complexity and uncertainty. We investigate the application of spatial reasoning in automatic georeferencing.
Identifying a competent service provider or contractor is not an easy task. This is especially true in the IT sector with its short-lived trends and products. In this context an"IT Company Atlas" giving substantiated information about the IT companies in a certain region seems to be a promising idea. However, a B2B directory administered manually for this purpose is expensive to establish and to maintain. An alternative approach discussed in this paper is to apply expert search techniques and provide a search engine searching for companies instead of documents. We propose a system searching for companies with expertise in a given field sketched by a keyword query. The system exploits the websites of the companies and covers all aspects: determining and representing the expertise of the companies, query processing and retrieval models, as well as query formulation and result presentation. In addition to the theoretical background and a system description we present experimental results comparing the effectiveness of different retrieval models adopted from expert search in the scenario we call "company search".
Summarization is an important means to cope with the challenges of big data. Summaries can help to achieve a first overview, they can be used to characterize subsets, they allow for the targeted access to data, and they build the basis for visualization techniques. In the present article, we point out the role of summaries as well as potential application scenarios. As examples, summarization techniques for spatial data (as an example for specific low dimensional techniques) and for general metric spaces (as a generic example with a broad spectrum of applications) are described. Furthermore, their use for resource selection and resource visualization in large distributed scenarios is outlined.
Usually retrieval systems search for documents relevant to a certain query or|more general|information need. However, in some situations the user is not interested in documents but other types of entities. In the paper at hand, we will propose a system searching for companies with expertise in a given eld sketched by a keyword query. The system covers all aspects: determining and representing the expertise of the companies, query processing and retrieval models, as well as query formulation and result presentation.
Many gazetteers contain only a small amount of historic place name information and spelling variants of places. Even more focused historic gazetteers are far from being complete and often specialize on certain geographic regions or particular time periods. On the other hand, there are huge amounts of historic route descriptions, so called itineraries. They represent massive knowledge sources from which historic place names and spelling variants can be deduced. The analysis and geocoding of those route descriptions---often done by hand---is an important task in the humanities. To cope with these problems, we present preliminary ideas how to automatically deduce historic place names and thus travel routes from historic route descriptions.
Toponyms in texts and search queries are often used figuratively and do not directly refer to the locations they reference in their literal sense. Different usage kinds and stylistic devices characterize toponym usages in texts. It is thus crucial for a Geographic Information Retrieval (GIR) system to precisely distinguish these different toponym usages at indexing and at query time in order to best address a given information need and the geospatial footprint of a document. For that purpose, we analyze which of the classic stylistic devices such as allegories, metaphors, or metonymies are used together with toponyms. We use these categories as a foundation for a systematic approach towards the characterization of toponym usages in texts which we believe is necessary to further boost retrieval effectiveness of future GIR systems. A prototype implements this characterization exemplary for texts written in German. We evaluate the effectiveness of our approach against a reference corpus to show the general feasibility. Our approach provides a basis for a wide range of more sophisticated applications such as for example text genre detection.
The amount of media items on the web is increasing tremendously, especially regarding personal media items. To effectively collaborate over and share these massive amounts of media objects, there is a strong need for adequate indexing and search techniques. Trends like social networks, large-storage mobile devices and high-bandwidth networks make peer-to-peer (P2P) information retrieval systems of deep interest. Hence, resource selection based on compact resource descriptions is used to efficiently determine promising peers w.r.t. a query. To design effective media search applications, multiple search criteria need to be addressed. Subsequently, besides text or visual media content, geospatial data is frequently used. We propose techniques to summarize and select collections of georeferenced media items in P2P systems. Generally, these summarization techniques can be divided into geometric and space partitioning approaches. This paper presents and evaluates techniques of a third category, hybrid approaches that combine features of geometric and space partitioning techniques.
The notion of quality in its broadest sense is central to information retrieval (IR) where a user’s information need is to be fulfilled as good as possible. A user searching for cars on sale in Bamberg might be interested in car dealers geographically close to Bamberg with high user ratings. The buyer might already know or trust a person who trusts the particular dealer. Furthermore, the cars which are sold by the dealer should offer a high quality on different levels–the type of car in general as well as the car to be bought. If the buyer can only travel to Bamberg on weekends, availability of the car dealer becomes another important factor. As this example shows, the integration of various quality aspects in IR is challenging but essential. Thus, there is a need for scalable and efficient indexing and retrieval techniques which can cope with such search situations. Here, metric space access methods (MAMs) present a flexible indexing paradigm.We will briefly review these techniques and show how they can be applied in the context of qualityaware IR. Furthermore, we will present IF4MI which is purely based on the inverted file concept and thus inherently provides a multi-feature MAM. It can make use of extensive knowledge in the field of inverted file-based indexing and represents a versatile indexing technique for quality-aware IR.
Similarity search in general metric spaces is a key aspect in many application fields. Metric space indexing provides a flexible indexing paradigm and is solely based on the use of a distance metric. No assumption is made about the representation of the database objects. Nowadays, ever-increasing data volumes require large-scale distributed retrieval architectures. Here, local and global indexing schemes are distinguished. In the local indexing approach, every resource administers a set of documents and indexes them locally. Resource descriptions providing the basis for resource selection can be disseminated to avoid all resources being contacted when answering a query. On the other hand, global indexing schemes are based on a single index which is distributed so that every resource is responsible for a certain part of the index. For local indexing, only few exact approaches have been proposed which support general metric space indexing. In this paper, we introduce RS4MI—an exact resource selection approach for general metric space indexing. We compare RS4MI with approaches presented in literature based on a peer-to-peer scenario when searching for similar images by image content. RS4MI can outperform two exact general metric space resource selection schemes in case of range queries. Fewer resources are contacted by RS4MI with—at the same time—more space efficient resource descriptions.
In this chapter, the authors outline how collections of georeferenced media items can be indexed and searched in P2P IR systems. They discuss different types of P2P IR systems and focus in detail on an approach based on collection description and selection techniques. This approach tries to adequately describe and select collections of georeferenced media items. Finally, the authors discuss its broad applicability in various application fields.
The ever-increasing amount of media items on the World Wide Web and on private devices leads to a strong need for adequate indexing and search techniques. Trends such as personal media archives, social networks, mobile devices with huge storage space, and networks with high bandwidth capacities make distributed solutions and in particular Peer-to-Peer (P2P) Information Retrieval (IR) systems attractive. On the other hand, when designing effective media search applications, various search criteria have to be addressed. Hereby, geospatial information is frequently used as well as other criteria, such as text, audio or visual media content, and date and time information. In this chapter, the authors outline how collections of georeferenced media items can be indexed and searched in P2P IR systems. They discuss different types of P2P IR systems and focus in detail on an approach based on collection description and selection techniques. This approach tries to adequately describe and select collections of georeferenced media items. Finally, the authors discuss its broad applicability in various application fields.
Content-based similarity search is an important task in multimedia information retrieval (IR). Here, metric space access methods (MAMs) can be applied. They are purely based on the use of a metric distance. No assumption is made about the representation of the feature objects. On the one hand, approximate MAMs have been proposed relying on the inverted file---the de facto standard index structure for text retrieval. On the other hand, there are many exact hierarchical and multi-step MAMs. We present IF4MI (Inverted Files for Metric Indexing), the first exact metric access method (MAM) based on the inverted file concept. IF4MI can outperform existing MAMs such as the M-tree and the PM-tree. In addition, the pruning power of current state-of-the-art techniques---namely the Metric Index---can be brought to inverted files without relying on an additional mechanism which maps feature objects to one-dimensional values for storing them in adequate data structures such as a B + -tree. IF4MI is conceptually appealing since it can make use of extensive knowledge in the field of inverted file-based indexing. As one example, we show how the efficient processing of textual filter queries---an important task in multimedia IR---is inherently supported.
Die stete Zunahme der Menge an Medienobjek- ten und -kollektionen sowohl im WWW als auch auf privaten Endgerfzu einem starken Bedarf nach ad¨ aquaten Indexierungs- und Such- technologien. Trends wie persMedienar- chive, soziale Netzwerke und mobile Ger¨ ate mit groser Speicherkapazit¨ at sowie Netzwerke mit hoher Bandbreite machen in diesem Kontext ver- teilte L¨ osungen und Peer-to-Peer (P2P) Techno- logien interessant. Hier dient eine Ressourcen- selektion, die auf kompakten Beschreibungen der von den Ressourcen verwalteten Inhalte basiert, dazu, f¨ ur eine bestimmte Anfrage vielverspre- chende Ressourcen zu ermitteln. Neben z.B. tex- tuellen Informationen k¨ onnen diese Zusammen- fassungen auch geographische Informationen der Medienobjekte eines Archivs beschreiben (z.B. wo verwaltete Bilder aufgenommen wurden). Diese Arbeit prund evaluiert verschie- dene Techniken zur Ressourcenauswahl, die auf der Beschreibung der geographischen Daten persMedienarchive basieren. Dabei wird nach in der Neines bestimmten Anfrage- punktes liegenden Medienobjekten gesucht.