Language contact is a pervasive phenomenon reflected in the borrowing of words from donor to recipient languages. Most computational approaches to borrowing detection treat all languages under study as equally important, even though dominant languages have a stronger impact on heritage languages than vice versa. We test new methods for lexical borrowing detection in contact situations where dominant languages play an important role, applying two classical sequence comparison methods and one machine learning method to a sample of seven Latin American languages which have all borrowed extensively from Spanish. All methods perform well, with the supervised machine learning system outperforming the classical systems. A review of detection errors shows that borrowing detection could be substantially improved by taking into account donor words with divergent meanings from recipient words.
We represent the complexity of Yine (Arawak) morphology with a finite state transducer (FST) based morphological analyzer. Yine is a low-resource indigenous polysynthetic Peruvian language spoken by approximately 3,000 people and is classified as ‘definitely endangered’ by UNESCO. We review Yine morphology focusing on morphophonology, possessive constructions and verbal predicates. Then we develop FSTs to model these components proposing techniques to solve challenging problems such as complex patterns of incorporating open and closed category arguments. This is a work in progress and we still have more to do in the development and verification of our analyzer. Our analyzer will serve both as a tool to better document the Yine language and as a component of natural language processing (NLP) applications such as spell checking and correction.
Identification of lexical borrowings, transfer of words between languages, is an essential practice of historical linguistics and a vital tool in analysis of language contact and cultural events in general. We seek to improve tools for automatic detection of lexical borrowings, focusing here on detecting borrowed words from monolingual wordlists. Starting with a recurrent neural lexical language model and competing entropies approach, we incorporate a more current Transformer based lexical model. From there we experiment with several different models and approaches including a lexical donor model with augmented wordlist. The Transformer model reduces execution time and minimally improves borrowing detection. The augmented donor model shows some promise. A substantive change in approach or model is needed to make significant gains in identification of lexical borrowings.
Lexical borrowing, the transfer of words from one language to another, is one of the most frequent processes in language evolution. In order to detect borrowings, linguists make use of various strategies, combining evidence from various sources. Despite the increasing popularity of computational approaches in comparative linguistics, automated approaches to lexical borrowing detection are still in their infancy, disregarding many aspects of the evidence that is routinely considered by human experts. One example for this kind of evidence are phonological and phonotactic clues that are especially useful for the detection of recent borrowings that have not yet been adapted to the structure of their recipient languages. In this study, we test how these clues can be exploited in automated frameworks for borrowing detection. By modeling phonology and phonotactics with the support of Support Vector Machines, Markov models, and recurrent neural networks, we propose a framework for the supervised detection of borrowings in mono-lingual wordlists. Based on a substantially revised dataset in which lexical borrowings have been thoroughly annotated for 41 different languages from different families, featuring a large typological diversity, we use these models to conduct a series of experiments to investigate their performance in mono-lingual borrowing detection. While the general results appear largely unsatisfying at a first glance, further tests show that the performance of our models improves with increasing amounts of attested borrowings and in those cases where most borrowings were introduced by one donor language alone. Our results show that phonological and phonotactic clues derived from monolingual language data alone are often not sufficient to detect borrowings when using them in isolation. Based on our detailed findings, however, we express hope that they could prove to be useful in integrated approaches that take multi-lingual information into account.
Siguiendo los métodos propuestos y las herramientas desarrolladas por Hammarström, Castermans, Forkel et al. (2018) para la visualización simultánea de índices de vitalidad lingüística y descripción gramatical, el presente artículo ofrece un análisis cuantitativo y cualitativo de los logros alcanzados y los desafíos pendientes en materia de documentación y descripción de la diversidad lingüística peruana. Se busca contribuir a determinar las verdaderas dimensiones de nuestro conocimiento sobre la diversidad lingüística de nuestro país y proponer algunas prioridades para una futura política para la diversidad lingüística peruana en la que descripción, documentación y revitalización se entiendan como tareas indesligables.
Alonso Vasquez, Renzo Ego Aguirre, Candy Angulo, John Miller, Claudia Villanueva, Željko Agić, Roberto Zariquiey, Arturo Oncevay. Proceedings of the Second Workshop on Universal Dependencies (UDW 2018). 2018.
We envisioned responsive generic hierarchical text summarization with summaries organized by section and paragraph based on hierarchical structure topic models. But we had to be sure that topic models were stable for the sampled corpora. To that end we developed a methodology for aligning multiple hierarchical structure topic models run over the same corpus under similar conditions, calculating a representative centroid model, and reporting stability of the centroid model. We ran stability experiments for standard corpora and a development corpus of Global Warming articles. We found flat and hierarchical structures of two levels plus the root offer stable centroid models, but hierarchical structures of three levels plus the root didn’t seem stable enough for use in hierarchical summarization.
Part of speech tagging is a fundamental component in many NLP systems. When taggers developed in one domain are used in another domain, the performance can degrade considerably. We present a method for developing taggers for new domains without requiring POS annotated text in the new domain. Our method involves using raw domain text and identifying related words to form a domain specific lexicon. This lexicon provides the initial lexical probabilities for EM training of an HMM model. We evaluate the method by applying it in the Biology domain and show that we achieve results that are comparable with some taggers developed for this domain.
Part of Speech (POS) tagging is often a prerequisite for tasks such as partial parsing and information extraction. However, when a POS tagger is simply ported to another domain the tagger's accuracy drops. This problem can be addressed through hand annotation of a corpus in the new domain and supervised training of a new tagger. In our methodology, we use existing raw text and a generic POS annotated corpus to develop taggers for new domains without hand annotation or supervised training. We focus in particular on out-of-vocabulary words since they reduce accuracy (Lease and Charniak. 2005; Smith et al. 2005).
Part-of-speech (POS) tagging is a fundamental component for performing natural language tasks such as parsing, information extraction, and question answering. When POS taggers are trained in one domain and applied in significantly different domains, their performance can degrade dramatically. We present a methodology for rapid adaptation of POS taggers to new domains. Our technique is unsupervised in that a manually annotated corpus for the new domain is not necessary. We use suffix information gathered from large amounts of raw text as well as orthographic information to increase the lexical coverage. We present an experiment in the Biological domain where our POS tagger achieves results comparable to POS taggers specifically trained to this domain.
An abstract is not available for this content so a preview has been provided. Please use the Get access link above for information on how to access this content.
An inexpensive inclined-screen smolt trap was designed and constructed for use in rivers having highly variable flow regimes. The trap included a pontoon-supported floating catch barge and an adjustable inclined screen made of parallel aluminum rods that effectively strained large volumes ofwater, transported smolts without injury, and was highly resistant to debris buildup and easily cleaned. The inclined screen was supported by a movable carriage within a stationary frame that permitted the screen to be deployed at a wide range of depths and angles depending on flow conditions and amount of water-borne debris. The trap was operated for three field seasons in the Bois Brule River, a large Wisconsin tributary to western Lake Superior, to capture and retain parr and smolts of steelhead Oncorhynchus mykiss, coho salmon O. kisutch, chinook salmon O. tshawytscha, and brown trout Salmo trutta, ranging from 45 to 300 mm in total length. The trap remained operational in flows ranging from 2.1 to 17.3 m3/s and through large variations in debris content without sustaining damage or requiring excessive maintenance. The design could be adapted to most locations where a low-head dam exists or can be established.
We have previously described the relationships of serum and saliva concentrations of caffeine, theophylline and total methylxanthines in premature infants. We have now extended the previous studies in order to validate the previously derived relationships with new data. The new serum to saliva relationships, derived using regression and ratio models, are cross validated against the relationships from the previous data, and vice versa. A good cross validation was observed for caffeine and total methylxanthine concentrations in infants treated with caffeine. In the theophylline treatment group, theophylline concentrations did not cross validate well, whereas the total methylxanthine concentrations did. Since in vivo conversion of theophylline to caffeine and vice versa may affect the individual methylxanthine relationships, the total methylxanthine equations are recommended for predicting serum concentrations from salivary concentrations.
We examined the clinical significance of noninvasive intracranial pressure measurements and pulsatility indices in 74 infants with confirmed IC-IVh. The intracranial pressure measurements were obtained using the applanation principle, and the pulsatility indices were calculated from the Doppler flow velocity tracings of the anterior cerebral artery. Fifty-three infants (71.6%) who died had a significantly lower birth weight and gestational age than those who survived. Survival rate decreased significantly with increased intracranial pressure (P less than 0.0002) and increased pulsatility indices (P less than 0.0001). We found no significant relationship between outcome and the size of IC-IVH demonstrated by CT scan. Birth weight, intracranial pressure measurements, and cerebral arterial pulsatile flow changes appear to be major prognostic indicators in neonatal IC-IVH.
The relationships of serum (Se) to saliva (Sa) concentrations of Caffeine (Caf) and total methylxanthines (Mx) (Mx = Caf + Theophylline (Theo)) have been established in neonates treated with Caf for apnea (J.Pediatr.96:494,1980). Serum to Sa relationships derived with additional studies have been crossvalidated against the old data. The standard deviations (SD) and correlation coefficients (R) for the Se to Sa relationships are presented below. Data for Mx is presented because of Caf to Theo conversion in neonates.From the high R values for both Caf and Mx it appears that the newly derived relationships crossvalidate well with the old data.
Normal, human diploid skin fibroblasts (GM-10) were grown in culture and the changes in the mean cellular, nuclear and nucleolar dry mass and area were monitored simultaneously at regular intervals until the last passage. The results which were analyzed statistically showed that all measures increased during aging in vitro with a most striking surge for all values from passage 29 to 31. The area measures of the cell, nucleus and nucleosis were well correlated with the measures of dry mass for the same structures. The results are discussed and possible reasons for the findings are offered.