
This paper presents a structured, machine-readable dataset of approximately 13,000 Byzantine book epigrams, mostly dating from the 11th to 15th centuries and written in Byzantine Greek. The dataset is derived from the Database of Byzantine Book Epigrams, an ongoing Ghent University project that provides a web interface for exploring the epigrams but does not support large-scale exports or cross-referencing of metadata. Our pipeline extracts and normalizes information from Elasticsearch and PostgreSQL, stores it in SQLite, and exports the results to Zenodo. By structuring the data, it becomes easier to use in computational and AI-driven research and allows intuitive linking of metadata with information from other sources.
The Techno-Optimism Archive Dataset is a curated collection of historical press examples of technological optimism and techno-solutionist thought. Drawn from dispersed digital archives of newspapers and magazines, it captures recurring promises that framed new technologies as remedies for major social, political, moral, and economic problems. The dataset consists of project-authored metadata and summaries describing individual records and was compiled primarily through archival research across digital repositories in order to bring into a reusable form materials that are publicly accessible yet often difficult to discover. It is deposited on Zenodo with accompanying documentation. The dataset can support research in the history of technology, media history, science and technology studies, and digital humanities.
The SCIROS project dataset was produced during the process of conducting a systematic literature review that addressed critical gaps in the understanding of open science. The review followed PRISMA guidelines. A cross-platform query in Scopus, Web of Science, and OpenAlex, identified the initial set of records, which was then further deduplicated and screened for relevance. The dataset is available on Zenodo as a Zotero library (RIS format) and as a spreadsheet. The Zotero library can be reused for full-text consultation and corpus analyses. The spreadsheet can be reused in network analyses (e.g., co-authorship and co-citation analyses). The Python script can be replicated for cross-platform bibliographic queries.
This discussion paper introduces a reusable digital research infrastructure, currently under construction at Ghent University, for the multidisciplinary annotation historical texts in all their material manifestations. The framework interlinks IIIF (International Image Interoperability Framework) manifests, semi-automatically produced transcriptions, and layered historical, linguistic, literary, and material annotations within a Linked Data environment. The infrastructure enables seamless collaboration across disciplines and supports the interoperable reuse of data that are typically siloed or mutually incommensurable. We illustrate the concept with a prototype corpus; however, the model is applicable to any historical text corpus regardless of origin. The result is an environment in which machine-readable annotations and metadata can be combined to explore the richness of variation in the material record across time, space, and cultural as well as linguistic contexts, thereby unlocking richer, more transparent analyses and maximising the reuse potential existing and newly produced annotations.
This paper presents ongoing research on the reuse of the AlpiLinK corpus within a broader effort to adapt Automatic Speech Recognition (ASR) technology to the linguistic challenges posed by South Tyrolean dialect, a cluster of Upper German varieties spoken in the multilingual province of South Tyrol. While Standard German dominates written communication, everyday speech frequently occurs in dialect, whose phonological, morphological and lexical divergence from the standard limits the performance of mainstream ASR systems. Within this context, we develop fine-tuned models built upon openly available ASR architectures and trained on domain-specific data, with Standard German as the target written output. Central to this work is the repurposing of the only publicly available dataset for this language pair, the AlpiLinK Corpus, which was originally created for dialectological research rather than for machine-learning applications. We combine AlpiLinK with additional sources, including online media, partner-contributed recordings and in-house material, within an ongoing data collection and refinement process. We describe the characteristics, strengths and limitations of AlpiLinK from the perspective of its reuse for ASR fine-tuning, alongside the methodological steps required for data preparation and model training in this setting. Preliminary results indicate that the current fine-tuned model substantially accelerates transcription workflows, while still exhibiting weaknesses in dialect-specific forms, punctuation, named entities and multilingual interference. We conclude by outlining recommendations for improving datasets intended for low-resource dialectal ASR, with particular attention to audio quality, licensing considerations, metadata stability and explicit consent for machine-learning-related reuse.