In identifying the print colors of words when some combinations of color and word occur more frequently than others, people quickly show evidence of learning these associations. This contingency learning effect is evident in faster and more accurate responses to high-contingency combinations than to low-contingency combinations. Across four experiments, we systematically varied the number of response-irrelevant word stimuli connected to response-relevant colors. In each experiment, one group experienced the typical contingency learning paradigm with three colors linked to three words; other groups saw more words (six or twelve) linked to the same three colors. All four experiments disconfirmed a central prediction derived from the Parallel Episodic Processing (PEP 2.0) model (Schmidt et al., 2016)—that the magnitude of the contingency learning effect should remain stable as more words are added to the response-irrelevant dimension, as long as the color-word contingency ratios are maintained. Responses to high-contingency items did slow down numerically as the number of words increased between groups, consistent with the prediction from PEP 2.0, but these changes were unreliable. Inconsistent with PEP 2.0, however, overall response time did not slow down and responses to low-contingency items actually sped up as the number of words increased across groups. These findings suggest that the PEP 2.0 model should be modified to incorporate response interference caused by high-probability associations when responding to low-probability combinations.
The production effect-that reading aloud leads to better memory than does reading silently-has been defined narrowly with reference to memory; it has been explored largely using word lists as the material to be read and remembered. But might the benefit of production extend beyond memory and beyond individual words? In a series of four experiments, passages from reading comprehension tests served as the study material. Participants read some passages aloud and others silently. After each passage, they completed multiple-choice questions about that passage. Separating the multiple-choice questions into memory-focused versus comprehension-focused questions, we observed a consistent production benefit only for the memory-focused questions. Production clearly improves memory for text, not just for individual words, and also extends to multiple-choice testing. The overall pattern of findings fits with the distinctiveness account of production-that information read aloud stands out at study and at test from information read silently. Only when the tested information is a very close match to the studied information, as is the case for memory questions but not for comprehension questions, does production improve accuracy.
Personas are quantifiable and describable ways of grouping people based on their behaviours. They are valuable to businesses as it enables them to better understand their customer base. The creation of personas from survey data requires establishing the client requirements, building a quantifiable personality scale, developing personality questions for a survey, and human subjective analysis. In this work, we have utilised clustering to automate the persona development process. We have developed a real-world survey for children (from 17 countries) which included 25 personality-based questions (based on the OCEAN model), 22 questions that captured purchase behaviour, and other general features from the children's land-scape. There were 63,969 completed questionnaires with a high proportion of categorical features, which were preprocessed to allow different segmentation methods to be tested. Preliminary results with simple K-means and a Euclidean distance function demonstrated that this was inappropriate for the survey data set. A novel distance function for K-means clustering has been developed, which can handle a mixture of feature types and to allow the importance of each feature to be varied, using a linearly weighted distance method. The function also incorporates the haversine distance function to provide a distance between two locations, enabling potential cultural differences to be exam-ined.We have also implemented Gaussian Mixture Model on the same feature set to compare the results and see the limitations of Gaussian Models Our novel approach generated clusters based on a combination of features including personality, consumer behaviour and location which has demonstrated key cultural differences across the globe. Results from our novel approach show that location distance is one of the key features when constructing personas as culture (location) has a significant effect on the way children answer survey questions.
We propose a novel phenomenon, attention contagion, defined as the spread of attentive (or inattentive) states among members of a group. We examined attention contagion in a learning environment in which pairs of undergraduate students watched a lecture video. Each pair consisted of a participant and a confederate trained to exhibit attentive behaviors (e.g., leaning forward) or inattentive behaviors (e.g., slouching). In Experiment 1, confederates sat in front of participants and could be seen. Relative to participants who watched the lecture with an inattentive confederate, participants with an attentive confederate: (a) self-reported higher levels of attentiveness, (b) behaved more attentively (e.g., took more notes), and (c) had better memory for lecture content. In Experiment 2, confederates sat behind participants. Despite confederates not being visible, participants were still aware of whether confederates were acting attentively or inattentively, and participants were still susceptible to attention contagion. Our findings suggest that distraction is one factor that contributes to the spread of inattentiveness (Experiment 1), but this phenomenon apparently can still occur in the absence of distraction (Experiment 2). We propose an account of how (in)attentiveness spreads across students and discuss practical implications regarding how learning is affected in the classroom. (PsycInfo Database Record (c) 2021 APA, all rights reserved).
Eyes in a schematic face and arrows presented at fixation can each cue an upcoming lateralized target such that responses to the target are faster to a valid than an invalid cue (sometimes claimed to reflect “automatic” orienting). One test of an automatic process concerns the extent to which it can be interfered with by another process. The present experiment investigates the ability of eyes and arrows to cue an upcoming target when both cues are present at the same time. On some trials they are congruent (both cues signal the same direction); on other trials they are incongruent (the two cues signal opposite directions). When the cues are congruent a valid cue produced faster response times than an invalid cue. In the incongruent case arrows are resistant to interference from eyes, whereas an incongruent arrow eliminates a cueing effect for eyes. The discussion elaborates briefly on the theoretical implications.
A widely held account asserts that single words are automatically identified in the absence of an intent to process them in the form of identifying a task set, and implementing it. We provide novel evidence that there is no fixed relation between intention and visual word identification. Subjects were randomly cued on a trial-by-trial basis as to whether to read aloud a single target word (Go) or not (No-go). When the Go-No Go probability was 50% (Experiment 1) the effect of stimulus quality (bright vs. dim targets) was the same size as in a separate block of 100% Go trials. In Experiment 2, where the Go-No Go probability was 80% in the cued condition, the stimulus quality effect was smaller than in the block of all Go trials. These results can be understood in terms of Go trial probability moderating whether subjects (i) hold off beginning to process the target until an intention in the form of a Task Set has been implemented, or (ii) begin to identify the target during the time taken to implement a Task Set. The additivity of stimulus quality and cueing conditions in Experiment 1 support the view that target processing only begins when a Task Set is in place, whereas the under-additivity of stimulus quality and cueing condition in Experiment 2 supports the interpretation that target identification can start during the time that a Task Set is being implemented. Taken together with other results, we conclude that there is no fixed relation between an intention and word identification; context is everything.
It is a widely held view that the determination of eye gaze direction is "automatic" in various senses (e.g., innate; informationally encapsulated; triggered without intent). The determination of arrow direction is also held to be automatic (following a certain amount of learning) despite not being innate. The present experiments evaluate the automaticity assumption of both eyes and arrows in terms of an interference criterion. The results of 10 experiments support the inference that explicit judgements of eye gaze direction, when participants respond with a lateralized key press, are (a) neither automatic in the strong sense (they are interfered with by an uninformative, incongruent arrow in the display) and (b) nor are they are automatic in a weaker sense (uninformative, incongruent arrows interfere more strongly with the determination of eye gaze direction than uninformative, incongruent eyes interfere with the arrow direction task). However, the determination of arrow direction is also not strongly automatic, given that it is interfered with by irrelevant eyes. At least with respect to an interference criterion, the determination of eye gaze direction appears less prepotent than the determination of arrow direction, which itself is only weakly automatic. (PsycInfo Database Record (c) 2021 APA, all rights reserved).
Measurement of the semantic and syntactic similarity of human utterances is essential in allowing machines to understand dialogue with users. However, human language is complex, and the semantic meaning of an utterance is usually dependent upon the context at a given time and learnt experience of the meaning of the words that are used. This is particularly challenging when automatically understanding the meaning of social media, such as tweets, which can contain non-standard language. Short Text Semantic Similarity measures can be adapted to measure the degree of similarity of a pair of tweets. This work presents a new Semantic and Syntactic Similarity Measure (TSSSM) for political tweets. The approach uses word embeddings to determine semantic similarity and extracts syntactic features to overcome the limitations of current measures which may miss identical sequences of words. A large dataset of tweets focusing on the political domain were collected, pre-processed and used to train the word embedding model, with various experiments performed to determine the optimal model and parameters. A selection of tweet pairs were evaluated by humans for semantic equivalence and correlated against the measure. The new measure can be used in a variety of applications, including for identifying and analyzing political narratives. Experiments on three diverse human-labelled test datasets demonstrate that the measure outperforms an existing measure, performs well on tweets from the political domain and may also generalize outside the political domain.
Distributed word representations have recently contributed to significant improvements in many natural language processing (NLP) tasks. Distributional semantics have become amongst the important trends in machine learning (ML) applications. Word embeddings are distributed representations of words that learn semantic relationships from a large corpus of text. In the social context, the distributed representation of a word is likely to be different from general text word embeddings. This is relatively due to the unique lexical semantic features and morphological structure of social media text such as tweets, which implies different word vector representations. In this paper, we collect and present a political social dataset that consists of over four million English tweets. An artificial neural network (NN) is trained to learn word co-occurrence and generate word vectors from the political corpus of tweets. The model is 136MB and includes word representations for a vocabulary of over 86K unique words and phrases. The learned model shall contribute to the success of many ML and NLP applications in microblogging Social Network Analysis (OSN), such as semantic similarity and cluster analysis tasks.
Measuring textual semantic similarity has been a subject of intense discussion in NLP and AI for many years. A new area of research has emerged that applies semantic similarity measures within Twitter. However, the development of these measures for the semantic analysis of tweets imposes fundamental challenges. The sparsity, ambiguity, and informality present in social media are hampering the performance of traditional textual similarity measures as "tweets", have special syntactic and semantic characteristics. This paper reviews and evaluates the performance of topological, statistical, and hybrid similarity measures, in the context of Twitter analysis. Furthermore, the performance of each measure is compared against a naïve keyword-based similarity computation method to assess the significance of semantic computation in capturing the meaning in tweets. An experiment is designed and conducted to evaluate the different measures through examining various metrics, including correlation, error rates, and statistical tests on a benchmark dataset. The potential weaknesses of semantic similarity measures in relation to Twitter applications of textual similarity assessment and the research contributions are discussed. This research highlights challenges and potential improvement areas for the semantic similarity of tweets, a resource for researchers and practitioners.
Short text similarity measures have lots of applications in online social networks (OSN), as they are being integrated in machine learning algorithms. However, the data quality is a major challenge in most OSNs, particularly Twitter. The sparse, ambiguous, informal, and unstructured nature of the medium impose difficulties to capture the underlying semantics of the text. Therefore, text pre-processing is a crucial phase in similarity identification applications, such as clustering and classification. This is because selecting the appropriate data processing methods contributes to the increase in correlations of the similarity measure. This research proposes a novel heuristic-driven pre-processing methodology for enhancing the performance of similarity measures in the context of Twitter tweets. The components of the proposed pre-processing methodology are discussed and evaluated on an annotated dataset that was published as part of SemEval-2014 shared task. An experimental analysis was conducted using the cosine angle as a similarity measure to assess the effect of our method against a baseline (C-Method). Experimental results indicate that our approach outperforms the baseline in terms of correlations and error rates.
Twitter, a microblogging online social network (OSN), has quickly gained prominence as it provides people with the opportunity to communicate and share posts and topics. Tremendous value lies in automated analysing and reasoning about such data in order to derive meaningful insights, which carries potential opportunities for businesses, users, and consumers. However, the sheer volume, noise, and dynamism of Twitter, imposes challenges that hinder the efficacy of observing clusters with high intra-cluster (i. e. minimum variance) and low inter-cluster similarities. This review focuses on research that has used various clustering algorithms to analyse Twitter data streams and identify hidden patterns in tweets where text is highly unstructured. This paper performs a comparative analysis on approaches of unsupervised learning in order to determine whether empirical findings support the enhancement of decision support and pattern recognition applications. A review of the literature identified 13 studies that implemented different clustering methods. A comparison including clustering methods, algorithms, number of clusters, dataset(s) size, distance measure, clustering features, evaluation methods, and results was conducted. The conclusion reports that the use of unsupervised learning in mining social media data has several weaknesses. Success criteria and future directions for research and practice to the research community are discussed.
Our willingness to persist in problem solving is often held up as a critical component in being successful. Allied against this ability, however, are a number of situational factors that undermine our persistence. In the present investigation, the authors examine 1 such factor-knowing that the answers to a problem are easily accessible. Does having answers to a problem available reduce our willingness to persist in solving it ourselves? Across 4 experiments, participants (university students from a large Canadian University) solved multisolution anagrams and were either provided the answers after giving up (and knew they would receive the answers) or not. Results demonstrated that individuals persisted for less time in the former condition. In addition, participants did not seem to be aware of the effect that answers had on their decisions to quit. Implications for our understanding of the role that access to answers has on persistence across a number of domains (e.g., education, Internet) are discussed.
The ease with which individuals can access information has changed drastically with the advent of the Internet. Understanding how this change in our information landscape influences thinking represents an important question for psychological science. Research has demonstrated that we have a fairly accurate sense of the relative availability of internal information – a feeling-of-knowing. Here we examine the extent to which individuals have developed a sense of the relative availability of information stored on the Internet (i.e., externally) – a feeling-of-findability. Results demonstrate that when individuals do not know the answer to a question their feeling-of-findability accurately predicts the amount of time it takes them to locate the answer on the Internet. Furthermore, this feeling-of-findability, when individuals do not know the answer to a question, is unrelated to individuals' feeling-of-knowing, despite the fact that the latter is also demonstrated to predict search times. Instead, feeling-of-findability appears to be predicted by intuitions about how difficult it will be to generate a successful search query and the popularity of the type of information sought.
Recent technological advances have given rise to an information-gathering tool unparalleled by any in human history-the Internet. Understanding how access to such a powerful informational tool influences how we think represents an important question for psychological science. In the present investigation we examined the impact of access to the Internet on the metacognitive processes that govern our decisions about what we "know" and "don't know." Results demonstrated that access to the Internet influenced individuals' willingness to volunteer answers, which led to fewer correct answers overall but greater accuracy when an answer was offered. Critically, access to the Internet also influenced feeling-of-knowing, and this accounted for some (but not all) of the effect on willingness to volunteer answers. These findings demonstrate that access to the Internet can influence metacognitive processes, and contribute novel insights into the operation of the transactive memory system formed by people and the Internet.
Fuzzy sentence semantic similarity measures are designed to be applied to real world problems where a computer system is required to assess the similarity between human natural language and words or prototype sentences stored within a knowledge base. Such measures are often developed for a specific corpus/domain where a limited set of words and sentences are evaluated. As new "fuzzy" measures are developed the research challenge is on how to evaluate them. Traditional approaches have involved rigorous and complex human involvement in compiling benchmark datasets and obtaining human similarity measures. Existing datasets often contain limited fuzzy words and do allow the fuzzy measures to be exhaustively tested. This paper presents an automatic method for the generation of a Multiple Fuzzy Word Dataset (MFWD) from a corpus. A Fuzzy Sentence Pairing Algorithm is used to extract and augment high, medium and low similarity sentence pairs with multiple fuzzy words. Human ratings are collected through crowdsourcing and the MFWD is evaluated using both fuzzy and traditional sentence similarity measures. The results indicated that fuzzy measures returned a higher correlation with human ratings compared with traditional measures.
This paper details the development of a novel and practical Conversational Agent for the Arabic language called ArabChat. A conversational Agent is a computer program that attempts to simulate conversations between machine and human. In this paper, the term `conversation' or `utterance' refers to real-time chat exchange between machine and human. The proposed framework for developing the Arabic Conversational Agent (ArabChat) is based on Pattern Matching approach to handle users' conversations. The Pattern Matching approach is based on the matching process between a user's utterance and pre-scripted patterns that represents different topics organized through novel scripting structure. A real experiment has been done in Applied Science University in Jordan as an information point advisor for their native Arabic students to evaluate the ArabChat.
Short text semantic similarity (STSS) measures are algorithms designed to compare short texts and return a level of similarity between them. However, until recently such measures have ignored perception or fuzzy based words (i.e. very hot, cold less cold) in calculations of both word and sentence similarity. Evaluation of such measures is usually achieved through the use of benchmark data sets comprising of a set of rigorously collected sentence pairs which have been evaluated by human participants. A weakness of these datasets is that the sentences pairs include limited, if any, fuzzy based words that makes them impractical for evaluating fuzzy sentence similarity measures. In this paper, a method is presented for the creation of a new benchmark dataset known as SFWD (Single Fuzzy Word Dataset). After creation the data set is then used in the evaluation of FAST, an ontology based fuzzy algorithm for semantic similarity testing that uses concepts of fuzzy and computing with words to allow for the accurate representation of fuzzy based words. The SFWD is then used to undertake a comparative analysis of other established STSS measures.
The focus of computerised learning has shifted from content delivery towards personalised online learning with Intelligent Tutoring Systems (ITS). Oscar Conversational ITS (CITS) is a sophisticated ITS that uses a natural language interface to enable learners to construct their own knowledge through discussion. Oscar CITS aims to mimic a human tutor by dynamically detecting and adapting to an individual's learning styles whilst directing the conversational tutorial. Oscar CITS is currently live and being successfully used to support learning by university students. The major contribution of this paper is the development of the novel Oscar CITS adaptation algorithm and its application to the Felder–Silverman learning styles model. The generic Oscar CITS adaptation algorithm uniquely combines the strength of an individual's learning style preference with the available adaptive tutoring material for each tutorial question to decide the best fitting adaptation. A case study is described, where Oscar CITS is implemented to deliver an adaptive SQL tutorial. Two experiments are reported which empirically test the Oscar CITS adaptation algorithm with students in a real teaching/learning environment. The results show that learners experiencing a conversational tutorial personalised to their learning styles performed significantly better during the tutorial than those with an unmatched tutorial.
This paper considers AI problems concerning reasoning in multi-agent environment. We introduce and study multi-agents’ non-linear temporal logic TS4^U_K_n based on arbitrary (in particular, non-linear, finite or infinite) frames with reflexive and transitive accessibility relations, and individual symmetric accessibility relations R i for agents. Main accent of our paper is modeling of logical uncertainty for statements via interaction of agents (passing knowledge). Conception of interacting agents is implemented via arbitrary finite paths of transitions by agents accessibility relations. We address problems decidability and satisfiability for TS4^U_K_n . It is proved that TS4^U_K_n is decidable (and, in particular, the satisfiability problem for it is also decidable). We suggest an algorithm for checking satisfiability based on computation possibility of refutation special inference rues in finite models of effectively bounded size.