Abstract Over the last nine years, a forty-seven-person digital humanities project has explored the feasibility of computer assisted paleography for Syriac. That is, could one use big data, visual analytics, and recent advances in the digital analysis of handwriting to better understand Syriac manuscripts and Syriac manuscript culture? Last June we launched the public-facing part of our project titled the “Digital Analysis of Syriac Handwriting” or DASH (dash.Stanford.edu). DASH consists of digital page images from 90% of the world’s surviving Syriac manuscripts securely dated to before the twelfth century, as well as 90,000 individually trimmed letters from these manuscripts. In addition to introducing this tool to scholars of Syriac, this article presents a series of novel visualization tools that can also serve as a starting point for digital paleography projects in other linguistic traditions.
Scholars have traditionally categorised early Syriac manuscripts as either Estrangela or Serto. The same categories dominate the prevailing narrative of how Syriac script is thought to have developed. Most see Estrangela as the earliest strata of Syriac and Serto as a later development. More recent scholarship explores how early manuscripts support neither this stark division between script styles nor a sequential development. Of particular challenge to this paradigm are a series of securely dated colophons and notes which use a script style different than the main part of the text. But previous work has looked at only five examples of this phenomenon. By expanding this investigation to 30 examples and drawing upon a recent compiled digital corpus of over 100,000 early Syriac letter forms, the present article explores how large data sets, digital analysis, and visual analytics can help one better understand the development of Aramaic scripts.
As part of a larger digital paleography project, our team has assembled a database of tens of thousands of individual Syriac letters and letter data from 96% of extant early Syriac manuscripts that have a secure composition date. Long term, such data can help scholars develop more accurate ways to classify Syriac scripts. In the present article we use this data to illustrate just how frequently the most common way of categorizing Syriac scripts as either Estrangela or Serto does not accurately convey the ways early scribes actually wrote. In addition to challenging this “Standard Model” of Syriac scripts, the project illustrates how large data sets, digital analysis, and visual analytics can help researchers address key philological and historical problems. Ruben Duval’s 1881 Traite Syriaque begins with a series of charts outlining the development of Syriac script. These charts divide the language into three scripts, a much later “Nestorian” script and—the focus of the present paper—two earlier scripts. According to Duval, the original Syriac script was Estrangela and, starting in the eighth century, there appeared a derived script that Duval termed “Jacobite” but is more commonly known as Serto. Duval was not the first scholar to divide early Syriac into these mutually exclusive script styles; similar divisions are extant among tenth-century Arabic writers and also appear in the thirteenth-century Syriac writings of Bar Hebreaus. This terminology was later adopted by most scholars of Syriac and, by the late nineteenth century, this Gabrielle Lachtrup, Laura Larson, Audrey Lehrer, Sam Miller, Breanna Murphy, Bianca Ng, Paige “Gigi” Zeiler, Carmen Paul, Isabelle Pequignot, Caitlin Rajala, Siddhi Shah, Becca Shofar, Julia Spector, Sara Therrien, Renee Wah, Stephanie Xie, Alice Yang, and Kira Yates. Inquiries about this article should be directed to Michael Penn at mppenn@stanford.edu. 2 George Anton Kiraz, A Grammar of the Syriac Language. Volume 1. Orthography (Piscataway, NJ: Gorgias Press, 2012), 215. Challenging the Estrangela / Serto Divide 45 schema was so well known that even the Victorian novelist George Eliot referred to Estrangela in her notebooks. Since the nineteenth century, our knowledge of the Syriac language has advanced immensely. But, when it comes to our categories of Syriac script, these have remained essentially the same. Consider, for example, Figure 1. On the left is the Syriac section of the script chart Duval published in 1881. To the right appears a 2016 script chart found in the most recently published Syriac text book. In almost all details these are identical. They show three scripts, Estrangela, Serto (aka. Jacobite), and the much later East Syrian (aka. Nestorian). For the focus of this paper, Estrangela and Serto, there is little morphological variation between certain letters (e.g. zayn, nun). For a number of letters, however, there is substantial variation between the two scripts (i.e. olaph, dolath, he, waw, mim, rish, taw, and final lomahd). As both the 1881 and 2016 script charts suggest, what makes these categories so appealing is how easily one can differentiate them from each other. For those letters that show variance, an Estrangela document will 3 Jane Irwin, George Eliot's Daniel Deronda Notebooks (Cambridge: Cambridge University Press, 1996), 406, 438. 4 Steven C. Hallam, Basics of Classical Syriac: Complete Grammar, Workbook, and Lexicon (Grand Rapids: Zondervan, 2016). 5 Final ayn could also be added to this list. But, because a final ayn appears so infrequently in Syriac, it was not feasible for us to identify a final ayn in every document. Our preliminary analysis suggests, however, that a scribe that uses an E final lomadh also uses an E final ayn and a scribe who uses an S final lomadh also uses an S final lomadh. The mim undergoes at least two substantial changes over time: 1) the earliest book hands usually have an open form of the mim in which they maintain a small opening on the baseline (just as they do a waw); 2) long after a closed form of the mim becomes popular, it further changes shape with a loop on the right and a left arm that meets the baseline making a v-shape on top. Because of our interest in earlier manuscripts, we have focused on the first of these changes and are defining an E mim as having an opening on the base line. There are also several letters, particularly gomal, teth, qoph, and shin that develop a substantially rounder form over time. But among securely dated manuscripts these more rounded forms do not clearly appear until the twelfth century and therefore are not the focus of this paper.
This paper describes a set of hand-isolated character samples selected from securely dated manuscripts written in Syriac between 300 and 1300 C.E., which are being made available for research purposes. The collection can be used for a number of applications, including ground truth for character segmentation and form analysis for paleographical dating. Several applications based upon convolutional neural networks demonstrate the possibilities of the data set.
Acknowledgments Prologue: The Year 630 Introduction Account ad 637 Chronicle ad 640 Letters, Isho'yahb III Apocalypse of Pseudo-Ephrem Khuzistan Chronicle Maronite Chronicle Syriac Life of Maximus the Confessor Canons, George I Colophon of British Library Additional 14,666 Letter, Athanasius of Balad Book of Main Points, John bar Penkaye Apocalypse of Pseudo-Methodius Edessene Apocalypse Exegesis of the Pericopes of the Gospel, Hnanisho' I Life of Theodute Colophon of British Library Additional 14,448 Apocalypse of John the Little Chronicle ad 705 Letters, Jacob of Edessa Chronicle, Jacob of Edessa Scholia, Jacob of Edessa Against the Armenians, Jacob of Edessa Kamed Inscriptions Chronicle of Disasters Chronicle ad 724 Disputation of John and the Emir Exegetical Homilies, Mar Abba II Disputation of Bet Hale Bibliography Index