Proceedings of the 2nd International Conference on Digital Tools & Uses Congress(2020)
Universités de Neuchâtel et de Genève
被引用1|浏览0
摘要
We investigate the creation of a 17th c. French literary corpus. We present the main options regarding available standards, the training data we created and the efficiency of the models produced for OCR, spelling normalisation and lemmatisation - always with open-source solutions. We also present our encoding choices and the global logic of a corpus designed as a virtuous circle, enhancing automatically the tools that are used for its construction.