In the last decade, significant, largely-governmental funding has been applied to the automatic transcription of handwritten documents. Uses for this kind of technology are somewhat limited given that the numbers of handwritten documents are on the decline. However, certain types of handwritten historical records can be crucial for genealogical research in that they identify key vital facts. In recent years, organizations like FamilySearch have exhausted huge efforts to identify, digitize, and transcribe these kinds of genealogically-rich records. Until now, such transcription has largely been done through massive crowd-sourced labor. We believe handwriting recognition technology is only a few years away from profitable application to genealogical documents. To test this hypothesis, we developed an evaluation paradigm for measuring handwriting recognition performance on four data collections of differing genres and languages. We invited research organizations to participate in the evaluation and compared performance to the outcome of human annotation. In this paper, we provide the details of this paradigm, including the guidelines, corpora and evaluation tools. Then we illustrate the exciting system results which suggest that the state-of-the-art is very close to providing real-world benefit to the automatic transcription of genealogically-rich documents.
In this paper, we illustrate the creation of a completely new Treebank which has significant application to the genealogical space and, to the best of our knowledge, has never been described before. We document the creation of a Personal Name Treebank (PNTB) which, though still a work in progress, already contains over 150,000 name structure classifications for people names derived from all the cultures, time frames, and writing scripts that are observed in our 800-million-name Common Pedigree at new.familysearch.org. The Common Pedigree includes names from various millennia, name from all countries of the world, and names rendered not only in Latin, but also in scripts such as Cyrillic and CJK. We describe the PNTB and its components, and we give a number of examples where this is particularly beneficial to genealogical search.
Personal names are the most critical elements for discovery and compilation of one’s heritage. However, historical and multilingual names are subject to many alterations and variations which make searching and matching of such names a great challenge. Consequently, we have created a name matching corpus and evaluation which capture most of these variations and provide a means whereby different name matching systems can be compared as to their effectiveness on matching these historical and cross-lingual names. It is our plan to make this corpus and evaluation available to the public in order to provide a means for wide-scale improvements in historical name matching. We here describe the formation of the corpus, the evaluation methodology and metrics. Lastly we show the performance of a number of a name matching systems and identify potential directions for future enhancements.