We present results of a long-term team collaboration of mathematicians and biologists. We focus on building a mathematical framework for the shape space constituted by a collection of homologous bones or teeth from many species. The biological application is to quantitative morphological understanding of the evolutionary history of primates in particular, and mammals more generally. Similar to the practice of biologists, we leverage the power of the whole collection for results that are more robust than can be obtained by only pairwise comparisons, using tools from differential geometry and machine learning. This paper concentrates on the mathematical framework. We review methods for comparing anatomical surfaces, discuss the problem of registration and alignment, and address the computation of different distances. Next, we cover broader questions related to cross-dataset landmark selection, shape segmentation, and shape classification analysis. This paper summarizes the work of many team members other than the authors; in this paper that unites (for the first time) all their results in one joint context, space restrictions prevent a full description of the mathematical details, which are thoroughly covered in the original articles. Although our application is to the study of anatomical surfaces, we believe our approach has much wider applicability.
In this study, we consider the question of repairing and recovering a low-dimensional manifold embedded in high-dimensional space from noisy scattered data. Given a noisy point cloud sampled from a low-dimensional manifold, suppose that part of the scattered data is missing, which results in holes. In these settings, the main goal is to accurately and efficiently reconstruct the information within these gaps. While in three-dimensions the problem has been extensively studied, the challenge of reinstating missing information for high-dimensional manifolds remains open. In this paper, we propose a new approach named Manifold Repairing via Locally Optimal Projection (R-MLOP). The method is defined as a minimization problem with three terms. First, we leverage the spatial proximity to the holes to balance between denoising the data and preserving geometric continuity among points situated along the hole’s boundaries. In addition, a penalty term is added to guarantee a quasi-uniform sampling of the unknown manifold. We prove that the suggested solution recovers missing information inside the hole, with an approximation order that is controlled by the density of the given scattered data as well as the size of the amended hole. The effectiveness of our approach is demonstrated by considering different manifold topologies, for single and multiple-hole repairing, in low and high dimensions.
The Bible is the product of a complex process of oral and written transmissions that stretched across centuries and traditions. This implies ongoing revision of the "original" or oldest textual layers over the course of hundreds of years. Although critical scholarship recognizes this fact, debates abound regarding the reconstruction of the different layers, their date of composition and their historical backgrounds. Traditional methodologies have grappled with these challenges through textual and diachronic criticism, employing linguistic, stylistic, inner-biblical, archaeological and historical criteria. In this study, we use computer-assisted methods to address the question of authorship of biblical texts by employing statistical analysis that is particularly sensitive to deviations in word frequencies. Here, the term "word" may be generalized to "n-gram" (a sequence of words) or other countable text features. This paper consists of two parts. In the first part, we focus on differentiating between three distinct scribal corpora across numerous chapters in the Enneateuch, the first nine books of the Bible. Specifically, we examine 50 chapters labeled according to biblical exegesis considerations into three corpora: the old layer in Deuteronomy (D), texts belonging to the "Deuteronomistic History" in Joshua-to-Kings (DtrH), and the Priestly writings (P). For pragmatic reasons, we chose entire chapters, in which the number of verses potentially attributed to different authors or redactors is negligible. Without prior assumptions about author identity, our approach leverages subtle differences in word frequencies to distinguish among the three corpora and identify author-dependent linguistic properties. Our analysis indicated that the first two scribal corpora - (D, the oldest layers of Deuteronomy, and DtrH, the so-called Deuteronomistic History) - are much more closely related to each other than they are to the third, (P). This observation aligns with scholarly consensus. In addition, we attained high accuracy in attributing authorship by evaluating the similarity of each chapter to the reference corpora. In the second part of the paper, we report on our use of the three corpora as ground truth to examine other biblical texts whose authorship is disputed by biblical experts. Here, we demonstrate the potential contribution of insights achieved in the first part. Our paper sheds new light on the question of authorship of biblical texts by offering interpretable, statistically significant evidence of the existence of linguistic characteristics in the writing of biblical authors/redactors, that can be identified automatically. Our methodology thus provides a new tool to address disputed matters in biblical studies.
We consider a problem of great practical interest: the repairing and recovery of a low-dimensional manifold embedded in high-dimensional space from noisy scattered data. Suppose that we observe a point cloud sampled from the low-dimensional manifold, with noise, and let us assume that there are holes in the data. Can we recover missing information inside the holes? While in low-dimension the problem was extensively studied, manifold repairing in high dimension is still an open problem. We introduce a new approach, called Repairing Manifold Locally Optimal Projection (R-MLOP), that expands the MLOP method introduced by Faigenbaum-Golovin et al. in 2020, to cope with manifold repairing in low and high-dimensional cases. The proposed method can deal with multiple holes in a manifold. We prove the validity of the proposed method, and demonstrate the effectiveness of our approach by considering different manifold topologies, for single and multiple holes repairing, in low and high dimensions.
In the desire to quantify the success of neural networks in deep learning and other applications, there is a great interest in understanding which functions are efficiently approximated by the outputs of neural networks. By now, there exists a variety of results which show that a wide range of functions can be approximated with sometimes surprising accuracy by these outputs. For example, it is known that the set of functions that can be approximated with exponential accuracy (in terms of the number of parameters used) includes, on one hand, very smooth functions such as polynomials and analytic functions and, on the other hand, very rough functions such as the Weierstrass function, which is nowhere differentiable. In this paper, we add to the latter class of rough functions by showing that it also includes refinable functions. Namely, we show that refinable functions are approximated by the outputs of deep ReLU neural networks with a fixed width and increasing depth with accuracy exponential in terms of their number of parameters. Our results apply to functions used in the standard construction of wavelets as well as to functions constructed via subdivision algorithms in Computer Aided Geometric Design.
Over the years, various algorithms were developed, attempting to imitate the Human Visual System (HVS), and evaluate the perceptual image quality. However, for certain image distortions, the functionality of the HVS continues to be an enigma, and echoing its behavior remains a challenge (especially for ill-defined distortions). In this paper, we learn to compare the image quality of two registered images, with respect to a chosen distortion. Our method takes advantage of the fact that at times, simulating image distortion and later evaluating its relative image quality, is easier than assessing its absolute value. Thus, given a pair of images, we look for an optimal dimensional reduction function that will map each image to a numerical score, so that the scores will reflect the image quality relation (i.e., a less distorted image will receive a lower score). We look for an optimal dimensional reduction mapping in the form of a Deep Neural Network which minimizes the violation of image quality order. Subsequently, we extend the method to order a set of images by utilizing the predicted level of the chosen distortion. We demonstrate the validity of our method on Latent Chromatic Aberration and Moire distortions, on synthetic and real datasets.
Most surviving biblical period Hebrew inscriptions are ostraca (ink-on-clay texts). They are poorly preserved and might fade rapidly once unearthed. Their proper and timely documentation is therefore essential. Our study of numerous Hebrew ostraca has demonstrated that multispectral imaging has the potential to reveal letters on ostraca otherwise invisible to the naked eye. In the case of Arad Ostracon No. 16 from Judah, dated to ca. 600 BCE, we unveiled three lines of text on its supposedly blank reverse side. This surprising outcome led us to question how many ostraca we might be discarding during excavations simply because the sherds look blank. To tackle the problem, we propose a preliminary excavation protocol for screening ceramic sherds prior to disposal. The protocol is based on our limited experience rather than fully supported statistical test experiments. Here we demonstrate the application of this procedure on recently unearthed pottery sherds from the excavation at Kiriath-jearim near Jerusalem.
Ancient texts are unique evidence providing a glimpse into the thoughts, day-to-day life, and culture of people of long-gone eras. Paleography, the study of writing, aims at documenting the inscriptions, transliterating the texts, reconstructing their historical context, and studying the evolution of writing itself. The digital revolution gave rise to computational paleography, introducing new tools of data acquisition, image processing, and machine learning. Herein, we will provide an introduction to the emerging field of computational paleography through the lens of ancient Hebrew inscriptions, dating from the Iron Age through the Middle Ages. The years that passed since their composition had a great effect on their preservation level, including blurs, stains, and erosions; moreover, some documents tend to fade in the years after their discovery. Therefore, it is of paramount importance to promptly document ancient inscriptions using the most suitable imaging techniques, such as visible, infra-red, or multispectral imaging. Image analysis and processing techniques, such as binarizations, letter segmentation, and letters’ prior estimation are valuable in their own right or may serve as a stage for subsequent tasks. We will also discuss automatic handwriting analysis and writers’ identification, which could shed light on the historical background of the inscriptions.
This article deals with the question of literacy in Israel and Judah. We deploy algorithmic and forensic methods to reveal the number of authors in two corpora of ostraca: Arad in Judah (ca. 600 BCE) and Samaria in Israel (eighth century BCE). In Judah, literacy disseminated down the military system to the quarter-master of the remote fort of Arad and possibly to his assistant. At Samaria our algorithmic work revealed two scribes over a minimum period of seven years, which seems to represent palace administration apparatus.
A highly discussed issue in the fields of Hebrew epigraphy and biblical research is the level of literacy in the Iron Age kingdoms of Israel and Judah (Rollston 2010; Davies and Romer 2013; Schmidt 2015). Treating this topic using biblical texts, for example, the references to scribes at the time of a given monarch, may lead to circular argumentation: The reality behind a given account may reflect the time of the authors, who could have lived centuries later and retrojected their own situation back onto earlier history. A preferable methodology is to consider the material evidence-the corpora of Iron Age Hebrew ostraca from archaeological excavations. The idea is to use algorithmic and forensic methods to distinguish between handwritings and thus the number of authors in a given corpus.
Past excavations in Samaria, capital of biblical Israel, yielded a corpus of Hebrew ink on clay inscriptions (ostraca) that documents wine and oil shipments to the palace from surrounding localities. Many questions regarding these early 8th century BCE texts, in particular the location of their composition, have been debated. Authorship in countryside villages or estates would attest to widespread literacy in a relatively early phase of ancient Israel's history. Here we report an algorithmic investigation of 31 of the inscriptions. Our study establishes that they were most likely written by two scribes who recorded the shipments in Samaria. We achieved our results through a method comprised of image processing and newly developed statistical learning techniques. These outcomes contrast with our previous results, which indicated widespread literacy in the kingdom of Judah a century and half to two centuries later, ca. 600 BCE.
Arad is a well preserved desert fort on the southern frontier of the biblical kingdom of Judah. Excavation of the site yielded over 100 Hebrew ostraca (ink inscriptions on potsherds) dated to ca. 600 BCE, the eve of Nebuchadnezzar’s destruction of Jerusalem. Due to the site’s isolation, small size and texts that were written in a short time span, the Arad corpus holds important keys to understanding dissemination of literacy in Judah. Here we present the handwriting analysis of 18 Arad inscriptions, including more than 150 pair-wise assessments of writer’s identity. The examination was performed by two new algorithmic handwriting analysis methods and independently by a professional forensic document examiner. To the best of our knowledge, no such large-scale pair-wise assessments of ancient documents by a forensic expert has previously been published. Comparison of forensic examination with algorithmic analysis is also unique. Our study demonstrates substantial agreement between the results of these independent methods of investigation. Remarkably, the forensic examination reveals a high probability of at least 12 writers within the analyzed corpus. This is a major increment over the previously published algorithmic estimations, which revealed 4–7 writers for the same assemblage. The high literacy rate detected within the small Arad stronghold, estimated (using broadly-accepted paleo-demographic coefficients) to have accommodated 20–30 soldiers, demonstrates widespread literacy in the late 7th century BCE Judahite military and administration apparatuses, with the ability to compose biblical texts during this period a possible by-product.
In this paper, we consider the fundamental problem of approximation of functions on a low-dimensional manifold embedded in a high-dimensional space, with noise affecting both in the data and values of the functions. Due to the curse of dimensionality, as well as to the presence of noise, the classical approximation methods applicable in low dimensions are less effective in the high-dimensional case. We propose a new approximation method that leverages the advantages of the Manifold Locally Optimal Projection (MLOP) method (introduced by Faigenbaum-Golovin and Levin in 2020) and the strengths of the method of Radial Basis Functions (RBF). The method is parametrization free, requires no knowledge regarding the manifold's intrinsic dimension, can handle noise and outliers in both the function values and in the location of the data, and is applied directly in the high dimensions. We show that the complexity of the method is linear in the dimension of the manifold and squared-logarithmic in the dimension of the codomain of the function. Subsequently, we demonstrate the effectiveness of our approach by considering different manifold topologies and show the robustness of the method to various noise levels.
This article deals with Ostracon 24 from Arad, Side A. It has three parts, written by different authors: an introduction to the multispectral imaging of the ostracon, which made this study possible, followed by two alternative decipherments of the inscription.
Three Hebrew ostraca, found near Khirbet Zanu’ (Ḥorvat Zanoaḥ) and published by Milevski and Naveh in 2005, were re-imaged using a high-end multispectral imaging technique. The re-imaging yielded dozens of changed or added characters and resulted in renewed, larger and improved readings, hereby published. In addition, we interpret the texts of the ostraca and place them in the context of the economy and administration of Judah in the seventh century BCE.
Handwriting is considered a unique "fingerprint" that characterizes a scribe (it is even used as evidence in modern forensics). In paleography (the study of ancient writing), it is presumed that each writer has a one prototype for each letter in the alphabet. Commonly, for ancient inscriptions, letters are organized into paleographic tables (where the rows are the alphabet letters, and the columns represent the examined inscriptions). These tables play a significant role in dating inscriptions based on their resemblance to columns in the table. In this paper, we argue that each scribe "fingerprint" is not represented by a single character prototype, but in fact by a distribution of characters. We introduce a framework for automatically identifying the writer style and constructing paleographic tables based on character histograms. Subsequently, we propose a method for comparing short documents utilizing letter distribution. We demonstrate the validity of the methods on two handwritten datasets: Modern and Ancient Hebrew pertaining to the First Temple period. Our methodology on the ancient dataset enables us to provide additional evidence concerning the level of literacy in the kingdom of Judah ca. 600 BCE.
Visible, near-infrared and shortwave-infrared (VNIR-SWIR) spectroscopy is an efficient approach for predicting soil properties because it reduces the time and cost of analyses. However, its advantages are hampered by the presence of soil moisture, which masks the major spectral absorptions of the soil and distorts the overall spectral shape. Hence, developing a procedure that skips the drying process for soil properties assessment directly from wet soil samples could save invaluable time. The goal of this study was twofold: proposing two approaches, partial least squares (PLS) and nearest neighbor spectral correction (NNSC), for dry spectral prediction and utilizing those spectra to demonstrate the ability to predict soil clay content. For these purposes, we measured 830 samples taken from eight common soil types in Israel that were sampled at 66 different locations. The dry spectrum accuracy was measured using the spectral angle mapper (SAM) and the average sum of deviations squared (ASDS), which resulted in low prediction errors of less than 8% and 14%, respectively. Later, our hypothesis was tested using the predicted dry soil spectra to predict the clay content, which resulted in R2 of 0.69 and 0.58 in the PLS and NNSC methods, respectively. Finally, our results were compared to those obtained by external parameter orthogonalization (EPO) and direct standardization (DS). This study demonstrates the ability to evaluate the dry spectral fingerprint of a wet soil sample, which can be utilized in various pedological aspects such as soil monitoring, soil classification, and soil properties assessment.
Most surviving biblical period Hebrew inscriptions are ostraca—ink-on-clay texts. They are poorly preserved and once unearthed, fade rapidly. Therefore, proper and timely documentation of ostraca is essential. Here we show a striking example of a hitherto invisible text on the back side of an ostracon revealed via multispectral imaging. This ostracon, found at the desert fortress of Arad and dated to ca. 600 BCE (the eve of Judah’s destruction by Nebuchadnezzar), has been on display for half a century. Its front side has been thoroughly studied, while its back side was considered blank. Our research revealed three lines of text on the supposedly blank side and four "new" lines on the front side. Our results demonstrate the need for multispectral image acquisition for both sides of all ancient ink ostraca. Moreover, in certain cases we recommend employing multispectral techniques for screening newly unearthed ceramic potsherds prior to disposal.
Arad Ostracon 16 is part of the Elyashiv Archive, dated to ca. 600 B.C. It was published as bearing an inscription on the recto only. New multispectral images of the ostracon have enabled us to reveal a hitherto invisible inscription on the verso, as well as additional letters, words, and complete lines on the recto. We present here the new images and offer our new reading and reinterpretation of the ostracon.