
Breve storia, dalla sua nascita a tutt'oggi, del Latinitatis Italicae Medii Aevi Lexicon, secolare e rilevante impresa scientifica lessicografica, relativa alla latinità italiana del Medioevo.
The proliferation of digital linguistic resources for Latin has created significant research opportunities, yet their potential is often constrained by a lack of interoperability. Developed in isolation, many corpora, dictionaries, and lexicons use heterogeneous formats, query languages, annotation criteria and tag sets, hindering integrated data analysis. The LiLa (Linking Latin) project addresses this challenge by creating a Knowledge Base of interconnected resources built on Linked Open Data principles. At its core is a lemma-based architecture, where a central Lemma Bank harmonizes divergent lemmatization practices across different sources, enabling seamless data integration. This paper introduces the fundamental structure of the LiLa Knowledge Base, which employs standard ontologies like OntoLex-Lemon and models data using RDF. We demonstrate the practical value of this interoperable ecosystem through a series of ready-to-use SPARQL queries. These use cases showcase how researchers can perform complex, cross-resource analyses, such as comparing lexical inventories between Classical and Medieval texts, examining word formation, or tracing semantic concepts across multiple corpora and dictionaries. By linking previously fragmented data, LiLa not only streamlines scholarly inquiry but also establishes a new paradigm for creating comprehensive, interconnected digital ecosystems for (not only) historical languages, making Latin a model for linguistic resource interoperability.
The Latin Text Archive (LTA) is an online platform hosted by the Berlin-Brandenburg Academy of Sciences (BBAW) since 2020 (https://LTA.bbaw.de). Its primary objective is to facilitate computer-assisted semantic analysis of Latin texts and corpora spanning various epochs and genres. The LTA collaborates with prominent text providers and related projects in this field. Its core activities center on post-philological editorial text preparation, which is essential for implementing text mining techniques in corpus-based historical semantics. The archive lemmatizes and stores Latin texts, augments them with relevant metadata, and organizes them within thematic or genre-specific corpora. These texts can be also read online and downloaded in various formats. Currently in a beta version, the LTA offers already 12,960 curated texts authored by 1,280 identified individuals, amounting to 54 million words. Furthermore, the LTA supplies access to its morphological lexicon, which supports the lemmatization process. Through the 'Latin Universe', users may also access additional texts not yet fully curated. Both texts and corpora are searchable via third-party tools such as 'Voyant-Tools' or through integrated functionalities like the 'Time series query' — which allows for diachronic comparison of keywords and lemmas — and 'Diacollo', which analyses co-occurring lemmas over time.
This article presents the main features of Hyperbase Web based on the Latin databases of L.A.S.L.A. This software stands out above all for its tools that enrich the reading by statistical measurements on word distribution throughout texts, as well as its deep learning algorithm applied to text analysis. After describing the L.A.S.L.A. corpus implemented in the various Latin databases, we will present the documentary and static functions accessible from the Hyperbase Web interface. Each main feature will be illustrated with machine output and enriched with linguistic interpretation suggesting complementary reading paths.
L’emendazione dei testi letterari antichi rappresenta una delle sfide più complesse della filologia classica. I modelli esistenti per semplificare questo task (Latin BERT e Logion) adottano un approccio di tipo fill-mask che presenta alcuni limiti significativi. Questo contributo introduce Philo-L1, un LLM di tipo seq2seq di circa 297 milioni di parametri basato sull'architettura T5, che tratta l'emendatio dei testi letterari latini come un task di generazione di testo con denoising dell’input del modello, e Ianus AI, la piattaforma web pensata per il suo utilizzo. Philo-L1, ottenuto dal fine-tuning di Philo-1-preview (a sua volta, risultato del fine-tuning di PhilTa), è stato addestrato su un dataset sintetico di circa 5 milioni di coppie di frasi contenenti nove classi di corruttele: errori paleografici, di pronuncia, di divisio, di inversione, di eco, saut du même au même, errori da integrazione con parola-segnale, aplografie e dittografie. In fase di valutazione, il modello ha raggiunto un’exact match accuracy (EMA) del 74.01%, una perplexity di 1.17 e un BLEU score di 94.51. Il confronto diretto con Latin BERT conferma la validità dell'approccio proposto (EMA: 77.96% vs 0.50%). In futuro, si prevede di ampliare le funzionalità del modello e di integrare tecniche di chain of thought ed Explainable AI.
Created by Yves Ouvrard, Collatinus is a lemmatizer and a morphological analyzer for Latin. Its initial aim was to help teachers prepare texts together with their vocabulary list, and to assist beginners in reading authentic Latin texts independently. In addition to the short translation proposed, Collatinus allows users to consult dictionaries, both digital and image-based, in order to access full lexical information. Collatinus relies on a lexical base of more than 80,000 lemmas and on tables of word-endings. The lexical base includes a short translation of the lemmas, mainly in English and in French. Collatinus’ ability to split an inflected form in its stem and ending enables it to provide its full morphological analysis. As the quantities of each syllable are given in its databases, Collatinus is also able to scan texts metrically or to indicate accentuation. Several external tools have been developed to extend these core possibilities. More recently, Collatinus has been coupled to AI techniques (Latin-BERT) to choose the more appropriate lemmatization and analysis for each word in context. The AI model was trained on the annotated LASLA corpus, and the resulting tagger produces output files in the same APN format used by LASLA.
Il contributo presenta il sito SERICA e le sue opportunità di utilizzo e di sviluppo per le ricerche linguistiche e per le indagini testuali relative a un corpus di testi in progressivo incremento e contenente prevalentemente scritti latino, ma anche qualche caso di testi spagnoli e russi.
While the field of Digital Humanities has successfully established robust infrastructures for the textual analysis of Latin, the auditory dimension of the language is still largely undeveloped. Specifically, the complex quantitative rhythm and intonation of classical poetry cannot be accurately replicated by Text-to-Speech models. This paper presents a computational workflow designed to bridge this gap, leveraging verified metrical data to produce high-fidelity and prosodically accurate audio recordings. By using the structured XML scansions of the Pedecerto project, the proposed pipeline employs a rule-based pre-processing routine to convert standard orthography into a phonetic script optimized for acoustic modelling. These adapted texts are then fed into a multimodal Large Language Model, which is steered via in-context prompt engineering to observe syllable quantity, ictus placement, pauses, and eventual elision. The technical architecture of this system is detailed, analyzing the specific orthographic interventions and prompts required to overcome the stress-timed bias of contemporary AI models. Finally, the implications of this tool for the wider Digital Humanities ecosystem are discussed, with particular attention to its potential to democratize access to Latin learning, support accessibility, and add new audio layers to existing digital projects and infrastructures.
Il presente contributo contiene una panoramica sulle varie componenti del progetto DigilibLT (https://digiliblt.uniupo.it/), biblioteca digitale nata nel 2010 presso l’Università del Piemonte Orientale, contenente testi latini in prosa di argomento secolare della Tarda Antichità. Vengono presentati le principali funzioni del sito web, le collaborazioni, i progetti nati come spin off di DigilibLT (TBL, SpaceLat) e gli sviluppi futuri, legati all’imminente aggiornamento del sito.
La banca dati Usuarium, dedicata allo studio delle varianti liturgiche medievali, può essere uno strumento utile per approfondire la conoscenza di alcuni testi mediolatini. Grazie alla possibilità di compiere ricerche sulla base del testo oppure del giorno, all’elenco completo delle fonti utilizzate (centinaia di messali) e alla visualizzazione di mappe, Usuarium può rendere più agevole lo studio di opere che presentano spesso dei riferimenti alla liturgia, come accade per la raccolta di introduzioni ai sermoni Proprietates rerum naturalium adaptatae sermonibus de tempore (seconda metà del XIV secolo). Ognuna di queste introduzioni si conclude con il Vangelo del giorno e spesso contiene anche i riferimenti all’introitus, alla collecta e all’epistola del giorno. Il database Usuarium è stato utilizzato per individuare i riferimenti liturgici. Si è quindi cercato di valutare se la scelta di certe letture possa portare degli indizi sul luogo di origine dell’opera. Questa operazione è stata resa più complessa dalla forte mobilità della tradizione delle Proprietates (ad esempio, si notano spesso differenze nella scelta delle Collectae). Infine, si offre una sintesi dei principali limiti e vantaggi dell’utilizzo di Usuarium.
RAMMSES (Augmented Reality for Medieval Music in Siena and Sienese) is a project funded in 2020-22 by Regione Toscana. Manuscripts and musical fragments of the Sienese area, dating from the 11th to 13th centuries, have been digitized, transcribed (both texts and melodies) and accompanied by a series of metadata, regarding genre, book, and liturgical occasions. On the website (www.rammses.unisi.it), each manuscript is accompanied by a paleographic description, the images uploaded using IIIF technology, an explanation of its origins – enriched by analysis of the feasts and songs contained within, as well as a comparison with similar manuscripts and repertoires – and a notice on peculiarities in the formulary. The records of the chants specific or, otherwise, representative of the Sienese area were performed by the Siena Cathedral Choir, led by Maestro Donati. The content of the AR experience, created from these chants, also allows you to view the pages of the codex containing the piece in question, along with information designed for a non-specialist audience. You can reach this experience through the website, or via QR-codes in the places where these pieces were performed centuries ago, or, finally, through a specially created virtual exhibition. All the data contained in the database can be accessed via a search interface that allows for more complex or narrow searches.
The Eurasian Latin Archive (ELA) is a digital platform designed to collect and analyse Latin and multilingual texts relating to East Asia from the medieval and early modern periods. Initially developed within the DAS-MeMo project (2018-2020) and subsequently expanded through the SERICA framework, ELA offers textual data and computational tools for corpus-based research. This article presents ELA as a research environment, outlining its data model, technical architecture, and NLP workflow. Through selected examples, it shows how users can query, filter and compare documents. Beyond a functional overview, the article reflects on some methodological challenges in modelling multilingual Latin corpora and sketches possible directions for future extensions of the project.
Compared to its two more prestigious sisters, Classical and Romance, Medieval Latin metrics still suffers a significant quantitative and qualitative gap in scholarly production and critical tools. Considering manuals alone, while there are dozens and dozens of products dedicated to Ancient and Romance metrics, only two exist on Medieval latin one, both now over fifty years old. The situation regarding metrical repertoires is even worse: while numerous exist for the two sisters, none for Medieval Latin metrics. This paper aims to propose a (digital) tool for cataloging Medieval Latin poetic production. A Medieval Latin Metricological Repertory in the form of a database, containing essential historical-literary, and basic metrical information, to provide a metrical framework for individual texts of Medieval Latin verse.
Si intendono illustrare le modalità e i risultati dell’impiego delle due banche-dati digitali Poetria Nova 2 e Corpus Corporum nell’analisi dello stile e della paternità delle inscriptiones metriche di Alcuino, a partire dalle ricerche di Hans-Dieter Burghardt, che, nella sua tesi dottorale inedita discussa a Heidelberg nel 1960, aveva effettuato un primo tentativo di analisi attributiva estesa a tutti i componimenti creduti composti da Alcuino contenuti nel perduto codice di Saint-Bertin, edito da André Duchesne nel 1617. Si intende esporre, in primo luogo, gli elementi di cui si è tenuto conto per la determinazione dell’autenticità di un testo epigrafico a partire dalla ricerca dei paralleli intratestuali; in secondo luogo, le modalità concrete con cui è stato interrogato il testo (tipi di ricerche effettuate sulle banche-dati; risultati attesi/registrati, etc.) attraverso lo strumento digitale Poetria Nova 2; in terzo luogo, si vogliono illustrare le modalità di impiego del database Corpus Corporum nella risoluzione di alcuni casi in cui, basandosi solo sugli elementi considerati (previamente esposti), la verifica dell’attribuzione risulti più problematica e la paternità alcuiniana più difficilmente determinabile. Il contributo si chiude tracciando un sintetico bilancio dei casi in cui Poetria Nova 2 e Corpus Corporum si sono rivelati decisivi per l’attribuzione della paternità.
This article surveys the development of ALIM (Archivio della Latinità Italiana del Medioevo) from its origins in the 1990s to the present. Established under the auspices of the Unione Accademica Nazionale to collect Italian-area Latin texts for the prospective compilation of a dictionary of European Medieval Latin, ALIM has received sustained institutional support, including PRIN funding. The paper highlights the project’s defining features, including the integration of documentary and literary texts, curated Collections, and a sophisticated search syntax, as well as subsequent enhancements such as ALIM Plus, which enables users to include or exclude texts beyond the original chronological (7th–14th centuries) and geographical (Italian area) scope. As a digital library, ALIM provides open-access texts in multiple formats, including TEI-encoded XML, while also supporting advanced linguistic analysis. In particular, the Lexicon tool, designed by F. Stella and developed by L. Tessarolo, facilitates investigations of lexical overlap and exclusivity across corpora. Closely associated with the ALIM project is EVT (dev. R. Rosselli Del Turco), a tool that has become widely adopted for the creation of digital critical editions.
Il Vocabolario Dantesco Latino (VDL) è uno strumento digitale che prevede la schedatura lessicografica integrale e sistematica delle opere latine di Dante, liberamente consultabile online e strettamente connesso con il parallelo Vocabolario Dantesco volgare (VD), che scheda le voci della Commedia. Il portale del VDL è stato inaugurato in occasione dell’ultimo centenario dantesco del 2021, con la pubblicazione delle prime voci, ed è oggi un cantiere in piena attività, in grado di offrire alla comunità scientifica già più di settecento voci lessicografiche. L’universo linguistico bilingue dantesco richiede uno strumento di ricerca sistematico e dialogante con altre risorse digitali: in questa prospettiva, i due rami del Vocabolario Dantesco bilingue, autonomi ma paralleli e interattivi, consentono di studiare l’intero patrimonio lessicale contenuto nelle opere di Dante, sia volgari che latine. L’articolo si concentra sullo sviluppo digitale, le caratteristiche tecniche e le applicazioni del Vocabolario Dantesco Latino (VDL), illustrando nello specifico il trattamento lessicografico delle voci latine dantesche, che si distingue per il suo metodo interdisciplinare, dato che, oltre a considerare aspetti linguistici, tiene conto anche delle implicazioni filologico-ecdotiche ed ermeneutiche dei lemmi che ricorrono in luoghi testuali controversi.
Il contributo illustra i fondamenti teorici e metodologici per la realizzazione di un nuovo Ernesti, inteso come riedizione critica e digitalmente implementata dei lessici tecnici della retorica greca e latina pubblicati nel 1795 e nel 1797. Dopo aver discusso i limiti dell’opera originaria, vengono esaminati strumenti lessicografici successivi, evidenziandone punti di forza e criticità. Come caso di studio, si analizza il lemma syncrisis, al fine di mostrare come un approccio integrato e comparativo alle tradizioni greca e latina permetta di ricostruire reti concettuali di maggiore complessità. Infine, si delineano i requisiti per una versione semanticamente arricchita, basata su codifica XML-TEI e su ontologie in grado di gestire l’eterogeneità delle classificazioni e di supportare interrogazioni avanzate. Uno strumento di questo tipo consentirebbe indagini sincroniche e diacroniche ad alta granularità, rendendo possibile tracciare con precisione la fortuna e l’evoluzione dei termini retorici attraverso autori, opere, generi e periodi.
The Corpus Corporum project hosted by the University of Zurich is the largest structured digital collection of Latin texts. The texts span from antiquity to the twentieth century, currently totalling approximately 226 million words across thirty corpora. Conceived as an open-access research infrastructure, it provides philologists, linguists, historians, and scholars of Latin with a unified environment for reading, searching, and analysing texts encoded in standardised TEI XML format. Important Latin dictionaries are integrated into the site. The platform, built on open-source technologies including BaseX, Sphinx, and TreeTagger, maintains a distinction between corpus, author, work, and edition levels, and integrates persistent identifiers (VIAF, Wikidata) and external resources such as geschichtsquellen.de. Recent advancements are discussed in the article, especially two major new analytical tools. The Text Reuse module enables configurable intertextual analysis based on k-skip-n-gram algorithms, while the Metrical Analysis module automatically identifies Latin poetic metres. These innovations allow large-scale, reproducible investigations of textual transmission and poetic structure. An example concerning the sources of Isidore of Seville’s Etymologiae is briefly discussed. Future developments envision AI-assisted translation, semantic indexing, and synonym-based search, thereby enhancing the platform’s potential as a comprehensive, interoperable resource for digital Latin philology and the broader field of computational humanities.
Sono presentate le attività seminariali e di laboratorio che hanno coinvolto studenti dell’Università di Macerata, con particolare attenzione all’applicazione di tecniche di Layout Analysis, Handwritten Text Recognition e correzione manuale su campioni estratti da due manoscritti virgiliani. Viene discusso inoltre il contesto di collaborazione fra Università, CNR e Infrastrutture di Ricerca.
Il contributo intende illustrare caratteristiche e potenzialità del portale Sidoniana, sviluppato nell’ambito di un percorso di dottorato presso l’Università Ca’ Foscari di Venezia e finalizzato all’allestimento di una nuova edizione digitale dell’epistolario di Sidonio Apollinare. La risorsa digitale, ancora ampiamente in via di sviluppo, ospita come testo-base dell’ottavo libro delle epistole – adoperato come case-study – quello dell’edizione del 1887 convenzionalmente attribuita a Christian Lütjohann, corredato di un apparato critico frutto di una nuova recensio della tradizione manoscritta e di un’aggiornata messa appunto dei principali aspetti problematici della vicenda trasmissiva di quest’opera. Il portale include inoltre alcuni strumenti pensati per favorire una prosecuzione degli studi sulla storia del testo di Sidonio.