The past few years have led to the widespread recognition that morphology is an independent domain of language functioning in dynamic interdependence with more familiar domains such as phonology and syntax. This has permitted nuanced research into the organization of morphological systems as well as the development of hypotheses concerning factors responsible for such organization. In this chapter we compare two classes of hypotheses adaptive explanations and neutral ones - for attested differences in morphological complexity claimed to correspond with sociocultural and demographic factors. While both examine language change as a (cultural) evolutionary process, we argue that much recent work on adaptive hypotheses for morphological complexity has been uncritically adaptationist, neglecting key results and lessons from population genetics about how to study evolutionary systems. Finally, we argue that neutral explanations are presently more likely explanations for the apparent association of morphological complexity and smaller, historically more isolated populations and should a priori be preferred over adaptive explanations unless and until a high evidential burden has been met.
Default unification operations combine strict information with information from one or more defeasible feature structures. Many such operations require finding the maximal subsets of a set of atomic constraints that are consistent with each other and with the strict feature structure, where a subset is maximally consistent with respect to the subsumption ordering if no constraint can be added to it without creating an inconsistency. Although this problem is NP-complete, there are a number of heuristic optimizations that can be used to substantially reduce the size of the search space. In this article, we propose a novel optimization, leaf pruning, which in some cases yields an improvement in running time of several orders of magnitude over previously described algorithms. This makes default unification efficient enough to be practical for a wide range of problems and applications.
Several studies have recently applied sentiment-based lexicons to Twitter to gauge local sentiment to understand health behaviors and outcomes for local areas. While this research has demonstrated the vast potential of this approach, lingering questions remain regarding the validity of Twitter mining and surveillance in local health research. First, how well does this approach predict health outcomes at very local scales, such as neighborhoods? Second, how robust are the findings garnered from sentiment signals when accounting for spatial effects? To evaluate these questions, we link 2,076,025 tweets from 66,219 distinct users in the city of San Diego over the period of 2014-12-06 to 2017-05-24 to the 500 Cities Project data and 2010-2014 American Community Survey data. We determine how well sentiment predicts self-rated mental health, sleep quality, and heart disease at a census tract level, controlling for neighborhood characteristics and spatial autocorrelation. We find that sentiment is related to some outcomes on its own, but these relationships are not present when controlling for other neighborhood factors. Evaluating our encoding strategy more closely, we discuss the limitations of existing measures of neighborhood sentiment, calling for more attention to how race/ethnicity and socio-economic status play into inferences drawn from such measures.
There has been a broad resurgence in word-based approaches and the reconceptualization of classical ‘word and paradigm’ (WP) approaches as general models of morphological analysis. WP models are well adapted to the description and analysis of complex morphological patterns, most transparently clear in inflection. Modern WP models demonstrate how morphological organization is fundamentally implicational: the central role of words (and paradigms) reflects their predictive value in a morphological system. Understanding the nature of morphological organization, within and across languages, requires exploration of the fundamental elements of implicational relations. Descriptively this involves identifying the internal structure of words and the ways this structure facilitates an external organization into patterns of relatedness. Theoretically, it is necessary to identify analytic tools appropriate for specifying and quantifying word-internal and word-external organization. This type of analytic approach encourages the investigation of the types of learning theories that may play a role in determining the patterns observed to occur and thereby help to explain their learnability.
Previous work demonstrates that a word's status as morphologically-simple or complex may be reflected in its phonetic realisation. One possible source for these effects is phonetic paradigm uniformity, in which an intended word's phonetic realisation is influenced by its morphological relatives. For example, the realisation of the inflected word frees should be influenced by the phonological plan for free, and thus be non-homophonous with the morphologically-simple word freeze. We test this prediction by analysing productions of forty such inflected/simple word pairs, embedded in pseudo-conversational speech structured to avoid metalinguistic task effects, and balanced for frequency, orthography, as well as segmental and prosodic context. We find that stem and suffix durations are significantly longer by about 4–7% in fricative-final inflected words (frees, laps) compared to their simple counterparts (freeze, lapse), while we find a null effect for stop-final words. The result suggests that wordforms influence production of their relatives.
In traditional word-and-paradigm models of morphology, an inflectional system is represented via a set of exemplary paradigms. Novel wordforms are produced by analogy with previously encountered forms. This paper describes a recurrent neural network which can use this strategy to learn the paradigms of a morphologically complex language based on incomplete and randomized input. Results are given which show good performance for a range of typologically diverse languages.
This paper examines the syntactic and semantic behavior of object arguments in Moro, a Kordofanian language spoken in central Sudan. In particular, we focus on multiple object constructions (ditransitives, applicatives, and causatives) and show that these objects exhibit symmetrical syntactic behavior; e.g., any object can passivize or be realized as an object marker, and all can do so simultaneously. Moreover, we demonstrate that each object can bear any of the non-agentive roles in a verb's semantic role inventory and that the resulting ambiguities are an entailment of symmetrical object constructions of the type found in Moro. Previous treatments of symmetrical languages have assumed a syntactic asymmetry between multiple objects and have developed theoretical analyses that treat symmetrical behaviors as departures from an asymmetrical basic organization of clausal syntax. We take a different approach: we develop a Head-Driven Phrase Structure Grammar account that allows a partial ordering of the argument structure (ARG-ST) list. The guiding idea is that languages differ with respect to the organization of their ARG-ST lists and their consequences for grammatical function realization: there is no privileged encoding, but there is large variation within the parameters defined by ARG-ST organization. This accounts directly for the symmetrical behaviors of multiple objects. We also show how this approach can be extended to account for certain asymmetrical behaviors in Moro.
Abstract Inheritance plays several distinct but crucial roles in the representation of linguistic information in Head-Driven Phrase Structure Grammar and other constraint-based frameworks. The primary use of inheritance is in the definition of the type signature, a specification of what counts as a well-formed linguistic object. In more recent developments of the theory, inheritance is also used to express substantive linguistic generalizations. This latter use is quite different from the original problem that inheritance was introduced into HPSG to solve, and there are other knowledge representation devices that might be more appropriate. In particular, delegation can also be used to express linguistic generalizations. The use of delegation can simplify HPSG analyses of some phenomena and can also clarify some of the issues that arise in the use default unification.
OBJECTIVE:To characterize the current language that is used in describing and defining gout, its symptoms, and its treatment by reviewing recent publications in rheumatology and determining how word choice may, or may not, be reflective of recent scientific developments in gout specifically.METHODS:This was a computational linguistics study, using collocations analyses and concordance analyses on a database of scientific literature related to gout. The final data set for analysis included 2,590 articles, all relating to gout and published between May 2003 and May 2013 and amounting to 12,101,036 tokens (sentence segments). Analysis was conducted by a team of linguists and social scientists.RESULTS:Our primary finding is that current disease language in gout is marked by ambiguity and imprecision, as evidenced by numerous terms that have similar but distinct meanings, but are nevertheless used interchangeably, therefore blending the slight but significant distinctions between these words. Whereas treatment language is characterized by a multitude of terms to describe a therapeutic mechanism of action, there is a relative void of terms and phrases used to describe success (treating to target) in gout.CONCLUSION:The data suggest that the language used to describe gout could be improved and updated. A transformation from an antiquated and insufficiently descript terminological set to one that reflects the recent scientific and clinical advancements made in the category would maximize opportunities for patient and physician understanding.
Recent developments in the Word and Pattern approach to complex morphology have argued that words and the patterned relations between words are primary objects of morphological analysis. The primacy of words has two part/whole dimensions: the nature of their internal structure and the nature of their external relations to one another. Words consist of constitutive parts and words themselves are parts of larger patterns of systemic relatedness. We argue that internal structure is essentially discriminative, rather than morphemic, i. e., what is crucial for morphological organization is the ability to discriminate (patterns of) words from one another and all types of internal distinctions suffice to facilitate the necessary discriminability to establish patterns of words. The value and the operation of a discriminative perspective on the internal structure of words is also evident in the analysis of an entirely different phenomenon. Greenberg's (1963) Universal 34 states that "No language has a trial number unless it has a dual. No language has a dual unless it has a plural." We present an associative model of the acquisition of grammatical number based on the Rescorla-Wagner learning theory Rescorla & Wagner (1972) that predicts this generalization. Number as a real-world category is inherently structured: higher numerosity sets are mentioned less frequently than lower numerosity sets, and higher numerosity sets always contain lower numerosity sets. Using simulations, we demonstrate that these facts, along with general principles of probabilistic learning, lead to the emergence of Greenberg's Number Hierarchy. The value of a discriminative perspective for language analysis (Ramscar & Yarlett 2007, Ramscar et al. 2010, 2013) becomes clear in both word-based morphology and its explanatory role addressing a typological conundrum.
Beyond caricatures: Commentary on Evans 2014 Farrell Ackerman and Robert Malouf The language myth: Why language is not an instinct. By Vyvyan Evans. Cambridge: Cambridge University Press, 2014. Pp. xi, 304. ISBN 9781107619753. $29.99. ‘In order to understand what another person is saying you must assume it is true, and try to imagine what it might be true of.’ —George Miller, quoted by Ross (1982:8) As will be familiar to a Language audience, many of the central questions of language study were identified by luminaries of language—Wilhelm von Humboldt, Hermann Paul, Ferdinand de Saussure, Mikołaj Kruszewski, William Dwight Whitney, Edward Sapir, Leonard Bloomfield, John R. Firth, Roman Jakobson, and Charles Hockett, among others—who preceded both mainstream generative grammar (MGG) and modern cognitive/functional approaches.1 These earlier generations of linguists were both awed and challenged by the intriguing properties of human languages. How can they be understood as biological, physiological, cognitive, and/or social objects? What is the relation between the synchronic and diachronic aspects of language, and how does this impact the definition of the linguistic objects to be explained? What is the nature of the different (sub)systems that make up language, and how are they learned so quickly, without explicit instruction? Are there unifying properties underlying the obvious diversity of languages? Can an understanding of other animal communication systems provide insight into human language? How does language development, both in communities and in individuals, relate to issues concerning phylogenetic and ontogenetic development of other traits and behaviors in different complex adaptive systems? Today, theoretical linguists largely agree that these and other big mysteries present challenges for the scientific study of human language. Divergences among modern approaches largely come down to the different bets they make about the best ways to address these questions. Disagreements concern different hypotheses about appropriate analytic assumptions, relevant methodologies of inquiry, effective theoretical constructs, the relevance of empirical crosslinguistic data to favorite theoretical assumptions and theory construction, and the relations between linguistic inquiry and research in other disciplines that explore the phylogeny and ontogeny of other natural complex phenomena. So, one could argue, it is not the fundamental questions that distinguish different approaches, but rather the variety of hypotheses marshaled to address them. In this context, theoretical linguistics should exhibit vigorous, substantive cross-theoretical debate about both [End Page 189] analyses of particular phenomena and the general assumptions and methodologies that guide competing analyses. If this was the state of discourse among the more commanding voices in the field, we doubt that Evans’s The language myth (TLM) would have been written. The book reflects the combat rather than the convergences (both acknowledged and not) concerning ideas and methodologies that have begun to characterize modern research in grammar. Many linguists are collaborating and synthesizing across research traditions and theories that have conventionally ignored one another or paid just enough attention to disparage each other, largely by misrepresentation and triumphal dismissal. The new pluralistic and interdisciplinary research is moving beyond the caricatures of familiar theoretical paradigms. TLM and the responses from MGG that it has generated, however, seem set on perpetuating the culture of caricature. This diminishes a focus on the necessarily multidisciplinary qualitative and quantitative study of language and, thereby, impoverishes efforts to popularize its real mysteries and results. We suspect that many readers will find the tone of TLM to be baitingly belligerent and occasionally obnoxious: for example, E employs sometimes recurring dismissive phrases (‘the language-as-instinct crowd’, ‘Chomsky and co.’, ‘self-dubbed evolutionary psychologists’, ‘swathes of fervent followers’, etc.) and makes irritating references to ‘supermodels’ in language examples throughout. He intends to deliver a drubbing and does so in a way that likely would leave a critical but naive reader wondering what is wrong with this unfamiliar field called linguistics. It is a jeremiad against perceived inequity, if not iniquity, in the house of language analysis.2 E places a vigorous focus on ‘debunking the myths’, as he understands them: While I, and a great many other professional linguists, now think the old view is wrong, nevertheless, the old view—Universal Grammar [(UG)]: the eponymous ‘language myth’—still lingers; despite being completely wrong, it is...
A summary is not available for this content so a preview has been provided. Please use the Get access link above for information on how to access this content.
Greenberg’s (1963) Universal 34 states that “No language has a trial number unless it has a dual. No language has a dual unless it has a plural.” We present an associative model of the acquisition of grammatical number based on the RescorlaWagner learning theory (Rescorla & Wagner, 1972) that predicts this generalization. Number as a real-world category is inherently structured: higher numerosity sets are mentioned less frequently than lower numerosity sets, and higher numerosity sets always contain lower numerosity sets. Using simulations, we demonstrate that these facts, along with general principles of probabilistic learning, lead to the emergence of Greenberg’s Number Hierarchy.
One dimension of this task is segmentation. For example, how can learners separate the stems from the affixes that signal a particular morphosyntactic property? If this information was all that learners had, they might hypothesize that zavod has the lexical meaning ‘factory’, and the –ov suffix signals genitive plural. A large body of experimental research addresses this syntagmatic, structural challenge of identifying recurrent partial forms (e.g., Saffran et al. 1996; Finley and Newport 2011; Aslin and Newport 2012). However, there is also a paradigmatic aspect to the problem. For example, the table below shows some alternative possibilities for how a plural form might be realized with different cases in Russian (Baerman et al. 2009).
Crosslinguistically, inflectional morphology exhibits a spectacular range of complexity in both the structure of individual words and the organization of systems that words participate in. We distinguish two dimensions in the analysis of morphological complexity. Enumerative complexity (E-complexity) reflects the number of morphosyntactic distinctions that languages make and the strategies employed to encode them, concerning either the internal composition of words or the arrangement of classes of words into inflection classes. This, we argue, is constrained by integrative complexity (I-complexity). The I-complexity of an inflectional system reflects the difficulty that a paradigmatic system poses for language users (rather than lexicographers) in information-theoretic terms. This becomes clear by distinguishing average paradigm entropy from average conditional entropy. The average entropy of a paradigm is the uncertainty in guessing the realization for a particular cell of the paradigm of a particular lexeme (given knowledge of the possible exponents). This gives one a measure of the complexity of a morphological system—systems with more exponents and more inflection classes will in general have higher average paradigm entropy— but it presupposes a problem that adult native speakers will never encounter. In order to know that a lexeme exists, the speaker must have heard at least one word form, so in the worst case a speaker will be faced with predicting a word form based on knowledge of one other word form of that lexeme. Thus, a better measure of morphological complexity is the average conditional entropy, the average uncertainty in guessing the realization of one randomly selected cell in the paradigm of a lexeme given the realization of one other randomly selected cell. This is the I-complexity of paradigm organization. Viewed from this information-theoretic perspective, languages that appear to differ greatly in their E-complexity—the number of exponents, inflectional classes, and principal parts— can actually be quite similar in terms of the challenge they pose for a language user who already knows how the system works. We adduce evidence for this hypothesis from three sources: a comparison between languages of varying degrees of E-complexity, a case study from the particularly challenging conjugational system of Chiquihuitlán Mazatec, and a Monte Carlo simulation modeling the encoding of morphosyntactic properties into formal expressions. The results of these analyses provide evidence for the crucial status of words and paradigms for understanding morphological organization.*
Crosslinguistically, inflectional morphology exhibits a spectacular range of complexity in both the structure of individual words and the organization of systems that words participate in. We distinguish two dimensions in the analysis of morphological complexity. ENUMERATIVE COMPLEXITY (E-complexity) reflects the number of morphosyntactic distinctions that languages make and the strategies employed to encode them, concerning either the internal composition of words or the arrangement of classes of words into inflection classes. This, we argue, is constrained by INTEGRATIVE COMPLEXITY (I-complexity). The I-complexity of an inflectional system reflects the difficulty that a paradigmatic system poses for language users (rather than lexicographers) in information-theoretic terms. This becomes clear by distinguishing AVERAGE PARADIGM ENTROPY from AVERAGE CONDITIONAL ENTROPY. The average entropy of a paradigm is the uncertainty in guessing the realization for a particular cell of the paradigm of a particular lexeme (given knowledge of the possible exponents). This gives one a measure of the complexity of a morphological system systems with more exponents and more inflection classes will in general have higher average paradigm entropy but it presupposes a problem that adult native speakers will never encounter. In order to know that a lexeme exists, the speaker must have heard at least one word form, so in the worst case a speaker will be faced with predicting a word form based on knowledge of one other word form of that lexeme. Thus, a better measure of morphological complexity is the average conditional entropy, the average uncertainty in guessing the realization of one randomly selected cell in the paradigm of a lexeme given the realization of one other randomly selected cell. This is the I-complexity of paradigm organization. Viewed from this information-theoretic perspective, languages that appear to differ greatly in their E-complexity the number of exponents, inflectional classes, and principal parts can actually be quite similar in terms of the challenge they pose for a language user who already knows how the system works. We adduce evidence for this hypothesis from three sources: a comparison between languages of varying degrees of E-complexity, a case study from the particularly challenging conjugational system of Chiquihuitlan Mazatec, and a Monte Carlo simulation modeling the encoding of morphosyntactic properties into formal expressions. The results of these analyses provide evidence for the crucial status of words and paradigms for understanding morphological organization.*
Daniel Flickinger合作论文数School of Humanities and Sciences, Stanford University2
Bernd Kiefer合作论文数Language Technology Lab, DFKI GmbH1