Subject indexing is vital for discovery but hard to sustain at scale and across languages. We release a large bilingual (English/German) corpus of catalog records annotated with the Integrated Authority File (GND), plus a machine-actionable GND taxonomy. The resource enables ontology-aware multi-label classification, mapping text to authority terms, and agent-assisted cataloging with reproducible, authority-grounded evaluation. We provide a brief statistical profile and qualitative error analyses of three systems. We invite the community to assess not only accuracy but usefulness and transparency, toward authority-anchored AI co-pilots that amplify catalogers' work.
With a large controlled vocabulary as the label set, the task of automated subject indexing in a library can be understood as a multi-label classification task. If the set of subject terms is large, the problem fits the Extreme Multi-Label Classification (XMLC) objective. In this study, we apply a selection of specialised supervised XMLC methods to the test case of subject indexing contemporary German scientific literature, collected at the German National Library (DNB). We contrast these results by including a classical lexical matching baseline and three of our own recently developed LLM-based methods into the benchmark. Algorithms are evaluated and compared in several metrics. This includes binary relevance comparisons with previously indexed material, as well as graded relevance ratings by professional subject librarians. A challenge for all methods is to reliably make suggestions from the long tail of the subject vocabulary. We find that supervised XMLC algorithms relying on transformer-based dense features give best results in terms of overall binary relevance metrics. However, focusing on graded relevance and performance in the long tail of our subject vocabulary, the LLM-based generative methods give better results, making them a promising alternative for future productive use.
The motivation for this paper is the need of the German National Library to massively increase its web archiving activities as the 12,000 snapshots collected per year are considered not to be enough to reflect the German web appropriately. As collecting everything which could be defined as the German web is impossible, the concept of selecting an “exemplary diversity” is introduced and substantiated by the guiding principles of relevance and diversity. The long-term goal is to find methods to (semi-)automatise the selection process and the prerequisite for this are well-defined selection criteria and sources of relevant as well as diverse topics that can be exploited. First steps to operationalise the guiding principles are outlined and supplemented by technological challenges and requirements to support the selection process as well as in the light of the advancement of web technology. Finally, the paper proposes a mix of curational and automatised measures and sources of information that should be taken into consideration to cater for a good coverage of relevance and diversity, supplemented by technological aspects regarding alternative representations of web content in web archives to be discussed to bring the content of web archives closer to the original user experience.
Virtual architectural reconstructions have been developed for nearly four decades now, with the London Charter providing a framework for their scholarly use, particularly emphasising the need for transparent and evaluable documentation. Despite this, documentation remains uncommon due to the absence of binding standards, limited incentives, and the considerable effort required. Crucially, experience shows that documentation cannot realistically be produced retrospectively and must be created alongside with the reconstruction process. To address this gap, IDOVIR, a platform supported by the German Research Foundation (DFG), aims to streamline documentation workflows and improve communication among researchers by providing an efficient, web-based framework for recording and sharing information. More than ten years of experience with such approaches have both refined existing methods and highlighted the need for broader standardisation. The community is increasingly seeking to define general principles, criteria, and best practices for high-quality documentation beyond project-specific solutions. This paper contributes to that effort by outlining key requirements, proposing structural guidelines, and compiling relevant source types for reconstructions, with the goal of fostering consensus and improving the transparency, understanding, and evaluation of virtual reconstruction projects. With our proposal, we want to intensify the necessary communication process in the community to establish coordinated standardisation, which is lacking at the moment.
Virtual reconstructions in the fields of architecture, archaeology, and cultural heritage are based on a variety of sources and arguments that should be documented in a comprehensible manner (paradata). A comprehensive, standardized classification of sources tailored explicitly for virtual reconstructions would greatly facilitate the creation of comparable and objective documentation. To address both, the DFG-funded project IDOVIR is developing a freely accessible online tool for the structured documentation of reconstruction processes. Despite some proposed vocabularies, a comprehensive source classification still does not exist. To this end, a hierarchical classification scheme of sources with seven primary outline levels has been developed, which is based on existing vocabularies and experiences from real projects. Each source can be further described using objective criteria (e.g., source type, context of origin, scale). According to Linked Open Data principles, the terms are linked to controlled vocabularies such as Getty AAT, GND, or Wikidata to enable interoperability and integration into research infrastructures (e.g., NFDI, EOSC). The classification system and the tool are currently brought to discussion within the community and continuously developed. Another key feature of IDOVIR is comparative visualization, which allows sources and 3D reconstructions to be viewed side by side, overlaid, and analysed. This supports the critical evaluation of reconstruction proposals during the development process.