In this work, we introduce the NavProc 1.0 Corpus - a medium-scale, annotated corpus of procedural texts within the naval domain - for use as a first step in modeling procedural structures derived from real-world data sources. In particular, we have rigorously produced annotations of frame semantics (i.e., PropBank-inspired trigger/role links) across verbal, nominal, and adjectival frames. Furthermore, we have annotated 21 distinct types of semantic markers and structural links between textual elements (e.g., frame triggers, entities, modifiers) which, taken together, result in a text-focused graph of semantic elements. Such a graph can be used to derive a more complex procedure structure for use in personnel training, simulation, or collaborative procedure execution. Altogether, this annotation effort has encompassed 158 procedural units composed of 2,316 sentences, 44,459 tokens, and 48,137 distinct span annotations. Furthermore, we describe and report LLM-based extraction scores for use as a baseline in future research using this dataset.
In this work, we introduce the NavProc 1.0 Corpus – a medium-scale, annotated corpus of procedural texts within the naval domain – for use as a first step in modeling procedural structures derived from real-world data sources. In particular, we have rigorously produced annotations of frame semantics (i.e., PropBank-inspired trigger/role links) across verbal, nominal, and adjectival frames. Furthermore, we have annotated 21 distinct types of semantic markers and structural links between textual elements (e.g., frame triggers, entities, modifiers) which, taken together, result in a text-focused graph of semantic elements. Such a graph can be used to derive a more complex procedure structure for use in personnel training, simulation, or collaborative procedure execution. Altogether, this annotation effort has encompassed 158 procedural units composed of 2,316 sentences, 44,459 tokens, and 48,137 distinct span annotations. Furthermore, we describe and report LLM-based extraction scores for use as a baseline in future research using this dataset.
Here we present a method for detecting an individual's level of conscientiousness based on an analysis of the content of their Facebook status updates. Our model is based on the identification of semantic evidence of facets related to conscientiousness; an individual's belief of their control over events around them and their goal orientation. The model achieves a correlation of r=.27 on a subset of the Facebook data published for the myPersonality workshop, with an accuracy of 58.13% for detecting if an individual is above or below the median and 68.03% for those outside of one standard deviation. While we take a narrow approach and identify only one personality trait, the general methodology of directly looking for evidence of traits in an individual's utterances is applicable to discovering models for all of the personality traits.
The voice of the customer is never more loudly heard than through social media. These online comments and reviews provide the insights marketers need to better build, design, and clarify the message around their products and services. Current approaches to mining these insights mainly focus on the volume and trend of sentiment. However, sentiment is not enough to discover actionable insights from these valuable social data. In this paper, we outline a four-factor model (Attitudinal, Sociocultural, Personal, and Behavioral) for mining consumer insights from social data that combines research in consumer and social psychology, discourse processing, and sentiment analysis. We present our current efforts in the automatic identification of a subset of the components making up these factors. In particular, we identify beliefs toward and about products and experiences, social actions in the form of recommendations, and intentions in the form of promises.
One way in which marketers gain insights about consumers is by identifying the occasions in which consumers use their products and which are invoked by their products.Identifying occasions helps in consumer segmentation, answering why consumers purchase a product, and where and when they use it.Additionally, the types of occasions a consumer participates in and the social settings surrounding those occasions provide insights into the consumer's personality and sociocultural self.Insights such as these are required for understanding consumer behavior, which marketers need to better design and sell their products.In this paper, we describe a methodology for extracting and categorizing occasions from product reviews, product descriptions, and forum posts.We examine using a maximum entropy markov model (MEMM) and a linear chain conditional random field (CRF) for extraction and find the CRF results in a 72.4% F1-measure.Extracted occasions are categorized as one of six high-level types (Celebratory, Special, Seasonal, Temporal, Weather-Related, and Other) using a support vector machine with an 88.5% macroaveraged F1-measure.
An individual‘s ability to produce quality work is a function of their current motivation, their control over the results of their work, and the social influences of other individuals. All of these factors can be identified in the language that individuals use to discuss their work with their peers. Previous approaches to modeling motivation have relied on social-network and time-series analysis to predict the popularity of a contribution to user-generated content site. In contrast, we show how an individual’s use of language can reflect their level of motivation and can be used to predict their future performance. We compare our results to an analysis of motivation based on utility theory. We show that an understanding of the language contained in comments on user generated content sites provides significant insight into an author’s level of motivation and the potential quality of their future work.
In this paper we present a generative model entitled the Author Perspective Model for the classification of deontic modality in event mentions. In the model modals, adverbials, and predicates associated with an event mention are generated by either a topic or author perspective where the author perspective is one of the three high level categories of deontic modality. We train the model with data gathered by a small set of seed phrases for each of the deontic modality categories. Our results show that we are able to classify the category of deontic modality with a micro-averaged F-Measure of 67.3%.
In this work, we present two complementary methods for the expansion of psycholinguistics norms. The first method is a random-traversal spreading activation approach which transfers existing norms onto semantically related terms using notions of synonymy, hypernymy, and pertainymy to approach full coverage of the English language. The second method makes use of recent advances in distributional similarity representation to transfer existing norms to their closest neighbors in a high-dimensional vector space. These two methods (along with a naive hybrid approach combining the two) have been shown to significantly outperform a state-of-the-art resource expansion system at our pilot task of imageability expansion. We have evaluated these systems in a cross-validation experiment using 8,188 norms found in existing pscholinguistics literature. We have also validated the quality of these combined norms by performing a small study using Amazon Mechanical Turk (AMT).
In this paper we semi-automatically construct a multilingual lexicon for Maslow’s seven categories of needs. We then use the semi-automatically constructed lexicons and a metaphor recognition system to analyze the change in rate of the expression of needs in the presence of metaphor. We examine four languages, English, Farsi, Russian, and Spanish, and focus on metaphors whose target concept is related to poverty or taxation.
We present a tiered-approach to the recognition of metaphor. The first tier is made up of highly precise expert-driven lexico-syntactic patterns which are automatically expended on in the second tier using lexical and dependency transformations. The final tier utilizes an SVM classifier using a variety of syntactic, semantic, and psycholinguistic features to determine if an expression is metaphoric. We focus on the recognition of metaphors in which the target is associated with the concept of "Economic Inequality" and examine the effectiveness of our approach for metaphors expressed in English, Farsi, Russian, and Spanish. Through experimental analysis we show that the proposed approach is capable of achieving 67.4% to 77.8% F-Measure depending on the language.
The intersection of psychology and computational linguistics is capable of providing novel automated insight into the language of everyday cognition through analysis of micro-blogs. While Twitter is often seen as banal or focused only on thewho,what,when orwhere tweets can actually serve as a source for learning about the language people use to express complex cogntive states and their cultural identity. In this contribution we introduce a novel model which captures latent cultural dimensions through an individual’s expressions of intentionality. We then show how these latent cultures can be used to create a culturally-sensitive model which provides enahnced detection of signals of intentionality in tweets. Finally, we demonstrate how these models reveal interesting cross-cultural differences in the goals and motivations of individuals from different cultures.
We present a novel approach to the problem of multilingual conceptual metaphor recognition. Our approach extends recent work in conceptual metaphor discovery by combining a complex methodology for facet-based concept induction with a distributional vector space model of linguistic and conceptual metaphor. In the evaluation of our system in English, Spanish, Russian, and Farsi, we experiment with several state-of-the-art vector space models and demonstrate a clear benefit to the fine-grained concept representation that forms the basis of our methodology for conceptual metaphor recognition.
Our everyday language reflects our psychological and cognitive state and effects the states of other individuals. In this contribution we look at the intersection between motivational state and language. We create a set of hashtags, which are annotated for the degree to which they are used by individuals to mark-up language that is indicative of a collection of factors that interact with an individual’s motivational state. We look for tags that reflect a goal mention, reward, or a perception of control. Finally, we present results for a language-model based classifier which is able to predict the presence of one of these factors in a tweet with between 69\% and 80\% accuracy on a balanced testing set. Our approach suggests that hashtags can be used to understand, not just the language of topics, but the deeper psychological and social meaning of a tweet.
Metaphor is a pervasive feature of human language that enables us to conceptualize and communicate abstract concepts using more concrete terminology. Unfortunately, computational models of natural language understanding - including systems for question answering, textual entailment, lexical substitution, and word-sense disambiguation - are unable to appropriately grasp the semantic content of metaphor and other forms of figurative language. In particular, we address the problem of understanding metaphoric language in the context of entailment (or paraphrase) detection. We build upon our existing state-of-the-art textual entailment system to specifically address issues of lexical entailment within a metaphoric context and have performed an in-depth experimental analysis to determine which techniques are most effective at interpreting metaphorical text. Our results suggest that a machine learning system trained on metaphor-rich data can achieve an accuracy above 90% for verbal metaphors using a combination of lexical, semantic, and contextual measures of term similarity.
ii Introduction Characteristic to all areas of human activity (from poetic to ordinary to scientific) and, thus, to all types of discourse, metaphor becomes an important problem for natural language processing. Its ubiquity in language has been established in a number of corpus studies and the role it plays in human reasoning has been confirmed in psychological experiments. This makes metaphor an important research area for computational and cognitive linguistics, and its automatic identification and interpretation indispensable for any semantics-oriented NLP application. The work on metaphor in NLP and AI started in the 1980s, providing us with a wealth of ideas on the structure and mechanisms of the phenomenon. The last decade witnessed a technological leap in natural language computation, whereby manually crafted rules gradually give way to more robust corpus-based statistical methods. This is also the case for metaphor research. In the recent years, the problem of metaphor modeling has been steadily gaining interest within the NLP community, with a growing number of approaches exploiting statistical techniques. Compared to more traditional approaches based on hand-coded knowledge, these more recent methods tend to have a wider coverage, as well as be more efficient, accurate and robust. However, even the statistical metaphor processing approaches so far often focused on a limited domain or a subset of phenomena. At the same time, recent work on computational lexical semantics and lexical acquisition techniques, as well as a wide range of NLP methods applying machine learning to open-domain semantic tasks, open many new avenues for creation of large-scale robust tools for recognition and interpretation of metaphor. This workshop is the first one focused on modelling of metaphor using NLP techniques. Recent related events include workshops on Computational Approaches to Figurative Language (NAACL 2007) and on Computational Approaches to Linguistic Creativity (NAACL 2009, NAACL 2010). We received 14 submissions and accepted 10. Each paper was carefully reviewed by at least 3 members of the Program Committee. The selected papers offer explorations into the following directions: (1) creation of metaphor-annotated datasets; (2) identification of new features that are useful for metaphor identification; (3) cross-lingual metaphor identification. The papers represent a variety of approaches to utilization and creation of datasets. While existing annotated corpora were used in some papers (Dunn, Tsvetkov et al), most papers describe creation of new annotated materials. Along with annotation guidelines adapted from the MIP and MIPVU procedures (Badryzlova et al), more intuitive …
The emergence of discussion and debate on social media necessitates the development of new models for processing dialogue. Vitally important to inferring the social implicatures of dialogue on social media is to understand the social goals and desires of the participants. Thus, to infer social implicatures methods for capturing the social goals and intentions of the participants must be first developed. In this paper, we propose a set of fifteen social acts to infer the social goals of dialogue participants. Social acts capture the complex social actions individuals signal through their utterances. We present a semi-supervised algorithm called the Social Act Conversation Model (SACM) for the fifteen social acts. The algorithm is based on the premise that linguistic expressions in social dialogue relate directly to the topic being discussed or to the social actions of the participants. We show that incorporating the social acts identified by the SACM with existing pattern-based identification can increase the performance in inferring social implicatures (adversarial behavior, pursuing power, and leadership) for online dialogue communicated in English, Arabic, and Chinese.
We present a method of constructing the semantic signatures of target concepts expressed in metaphoric expressions as well as a method to determine the conceptual space of a metaphor using the constructed semantic signatures and a semantic expansion. We evaluate our methodology by focusing on metaphors where the target concept is Governance. Using the semantic signature constructed for this concept, we show that the conceptual spaces generated by our method are judged to be highly acceptable by humans.
In this paper, we investigate whether the social roles of dialogue participants can be recognized through the social actions performed by the participant in their interactions with others in the group. Specifically we focus on determining if a participant is the leader of the group. We decompose the problem into identifying the social goals for participant discourse segments. These social goals are represented through a set of eleven psychologically-motivated social acts. We then model leadership using a sociological-inspired model called social rank which takes into account the social capital accumulated by the participant over the course of a single dialogue. We explore these models in task-oriented dialogues communicated in English, Arabic, and Chinese and show that the incorporation of social rank can improve precision of detecting the leader by 14% in English, 8% in Arabic, and 4% in Chinese.