We describe an efficient method to accurately estimate the effectiveness of a previously trained deep learning model for use in a new learning task. We use this method, "Predict To Learn" (P2L), to predict the most likely "source" dataset to produce effective transfer for training on a "target" dataset. We validate our approach extensively across 21 tasks, including image classification tasks and semantic relationship prediction tasks in the linguistic domain. The P2L approach selects the best transfer learning model on 62% of the tasks, compared with a baseline of 48% of cases when using a heuristic of selecting the largest source dataset and 52% of cases when using a distance measure between source and target datasets. Further, our work results in an 8% reduction in error rate. Finally, we also show that a model trained from merging multiple source model datasets does not necessarily result in improved transfer learning. This suggests that performance of the target model depends upon the relative composition of the source dataset as well as their absolute scale, as measured by our novel method we term 'P2L'.
Declarative memory is supported by distributed brain networks in which the medial-temporal lobes (MTLs) and pFC serve as important hubs. Identifying the unique and shared contributions of these regions to successful memory performance is an active area of research, and a growing literature suggests that these structures often work together to support declarative memory. Here, we present data from a context-dependent relational memory task in which participants learned that individuals belonged in a single room in each of two buildings. Room assignment was consistent with an underlying contextual rule structure in which male and female participants were assigned to opposite sides of a building and the side assignment switched between buildings. In two experiments, neural correlates of performance on this task were evaluated using multiple neuroimaging tools: diffusion tensor imaging (Experiment 1), magnetic resonance elastography (Experiment 1), and functional MRI (Experiment 2). Structural and functional data from each individual modality provided complementary and consistent evidence that the hippocampus and the adjacent white matter tract (i.e., fornix) supported relational memory, whereas the ventromedial pFC/OFC (vmPFC/OFC) and the white matter tract connecting vmPFC/OFC to MTL (i.e., uncinate fasciculus) supported memory-guided rule use. Together, these data suggest that MTL and pFC structures differentially contribute to and support contextually guided relational memory.
Conversational dialog systems are well known to be an effective tool for learning. Modern approaches to natural language processing and machine learning have enabled various enhancements to conversational systems but they mostly rely on text- or speech-only interactions, which puts limits on how learners can express and explore their knowledge. We introduce a novel method that addresses such limitations by adopting a visualization that is coordinated with a text-based conversational interface. This allows learners to seamlessly perceive and express knowledge through language and visual representations.
IBM and Pearson have partnered to develop dialogue-based intelligent tutoring systems at an unprecedented scale. We leveraged the decades-long research in intelligent tutoring systems (specifically dialogue based tutoring systems) and advances in machine learning and natural language processing to create a Watson dialogue-based tutor (WDBT). WDBT is currently being used by hundreds of students across multiple institutions. This paper describes our plans for preliminary evaluations of WDBT. Our formal evaluations begin in Spring 2018 and we will present findings shortly thereafter.
In 2016, IBM and Pearson announced a partnership to deliver a next generation learning service in the form of a dialogue-based tutor. Dialogue-based tutoring systems have demonstrated efficacy, but they are difficult to design and scale across domains. We have developed a framework for enabling digital courseware with a dialogue-based tutoring experience that can be applied to new domains with additional domain-specific content, but without re-design of the conversation flow or use case. The framework uses a content model that’s consistent across domains, which enables a general dialogue-based tutoring strategy. We identify several challenges to this approach, as well as recommendations for future work. The Pearson-IBM partnership In 2016, IBM and Pearson announced a partnership to deliver next generation learning services in the form of a digital tutoring system. The aim of this partnership was to create learning experiences powered by Pearson’s highquality content and IBM’s Watson technologies. While there have been many successful intelligent tutoring systems (ITSs), our challenge was scaling the process across multiple disciplines and titles. This paper is organized into 4 parts. First, we briefly review the value of dialogue-based tutoring systems. Second, we describe what we call the Watson dialogue-based tutor. Third, we will describe some of our approaches to scaling and evaluating the Watson dialogue-based tutor. Finally, describe the limitations of our approach and recommendations for future work. Dialogue-based tutoring systems Dialogue-based tutoring systems (DBTs) are an approach to ITSs that create a learning experience driven by natural language dialogue and classification of student natural language responses (e.g., Graesser, 2011). DBT conversations can be described as Socratic because the tutor guides the student through concepts via dialogue moves, which can include questions, hints, and other prompts. DBTs have been claimed to support a variety of learning principles and strategies, like encouraging constructive behaviors and self-explanations (M. T. H. Chi, 2009), deep reasoning questions (e.g., Graesser & Person, 1994), and conceptual understanding through scaffolding (e.g., VanLehn, 2011). DBTs require students to construct natural language responses, which can have a positive impact on memory and comprehension of source text (e.g., McNamara, 1992). DBTs can also provide immediate feedback to facilitate learning (Shute, 2008). For example, AutoTutor is a DBT, i.e., an ITS that initiates discourse with a student. The discourse patterns of the earliest AutoTutor were inspired by analyses of approximately 100 hours of non-expert human tutoring interactions (Graesser, 2011), which showed that students in need of tutoring are not active, self-regulated learners, and are not aware of their knowledge deficits. This affects how students converse: they do not effectively take command of the tutorial agenda and typically ask only 6-8 genuine information-seeking questions per hour. In contrast, tutors set 100% of the agenda, introduced 93% of the topics, presented 82% of examples, and asked 80% of the questions. Tutors do this by invoking a curriculum script of topics, problems, questions, and examples to drive a Socratic tutoring dialogue with students. Based on this tutor analysis, AutoTutor was designed to control the conversation through an expectation-misconception discourse model of tutoring (Graesser, 2011). This consists of a set of anticipated correct ideal answers (expectations) and a set of invalid answers frequently expressed by students (misconceptions). AutoTutor follows this design in a five-step tutoring framework: (1) tutor poses a question/problem, (2) student attempts to answer, (3) tutor provides brief evaluation as feedback, (4) collaborative interaction to improve the answer, (5) tutor checks if student understands. Efficacy of DBTs Steenbergen-Hu and Cooper (2014) conducted a meta-analysis of 39 studies evaluating the use of ITSs (including DBTs) in higher education. The researchers found an overall, moderate, positive effect (g = .35) favoring the use of ITS over other instructional conditions. When compared specifically to alternatives that were either “self-reliant learning activities” or no-treatment conditions, the use of ITSs appeared to offer a large advantage (g =.86). AutoTutor in particular has shown significant learning gains over non-interactive learning materials in a variety of math and science domains: computer literacy, physics, biology, and critical thinking (Graesser, 2011). Typically, higher gains were found for more complex questions, such as “how” and “why” questions, versus shallow questions, such as “who” or “what” questions (Nye, Graesser and Hu, 2014). Scalability of DBTs While DBTs have demonstrated a wide range of possible behaviors and pedagogical strategies, building a DBT for a new domain, course, or textbook is a non-trivial task. Even when the use case and learning goals are clearly defined, creating the necessary domain models for these tutoring systems can be very challenging for domain experts. This is a general problem for ITSs, which is why the researchers behind the most widely adopted tutoring systems have also developed authoring tools (e.g., Aleven, McLaren, Sewall, & Koedinger, 2009). Watson DBT Watson dialogue-based tutor (WDBT) follows the design of AutoTutor in many respects, but with a content creation and iterative design cycle to support application to new domains. WDBT begins with deep reasoning questions and then provides hints to assist students to give a response that matches a set of assertions or knowledge components. An example transcript from an interaction with WDBT is shown in Table 1. WDBT is made up of 6 components to achieve this functionality: (1) Domain Model, (2) Dialogue Content, (3) Natural Language Response Classification, (4) Question Answering, (5) Learner Modeling, and (6) Dialogue Management. Domain Model The domain model defines the knowledge and skills we want students to learn. We create a domain model for a specific title (i.e. textbook) that breaks down the knowledge into educational objectives consisting of learning objectives and enabling objectives. Learning objectives are broad learning goal statements e.g., “Analyze physical changes that occur in middle adulthood.” Enabling objectives are more granular learning goal statements that support the learning objective e.g., “Identify the physical benchmarks of change in middle adulthood.” Educational objectives serve as the foundation for creating content and assessment in Pearson, so the same framework was used to enable WDBT. The domain model also contains misconception statements that are aligned to educational objectives. Domain experts create the domain models, choosing learning and enabling objectives that are particularly difficult for students or that would benefit from a conversational learning experience. For each title, domain models typically contain 24 learning objectives and 110 enabling objectives (about 5 enabling objectives per learning objective). Because the learning objectives follow the general organization of the title, the domain model acts as an extension of the core content in the title. Dialogue Content The dialogue content for WDBT consists of a mix of content created manually by domain experts, and automatically extracted content. Subject matter experts (SMEs) are responsible for authoring: • Main question and answer pairs aligned to both learning and enabling objectives. The questions are broad, high-level conceptual questions, to assess deep comprehension of source text. WDBT delivers these questions to students. The corresponding answers represent ideal responses and are used to evaluate student answer attempts. • Hint questions and assertions for every main question and answer pair. Assertions are independent knowledge components that make up a main answer. Hints are intended to elicit an assertion from a student. WDBT delivers these hint questions to students. • Misconception statements and corresponding feedback messages for every applicable learning and enabling objective. In the event that response classification detects that a student is expressing a misconception, WDBT delivers the corresponding feedback message. Other forms of dialogue content are automatically or semi-automatically extracted: • Question-answer pairs based on key terms and other sources within the textbook. • Concept groupings based on key concepts found in the unstructured text. These groupings are automatically extracted (1) but validated by a subject matter expert. • Fill in the blank prompts that are dynamically created within the dialogue using an assertion and response classification gap analysis. This approach to content organization is applied across titles to facilitate faster adaptation to new titles, but it is not without its limitations. There is still a considerable amount of human effort required for this process, although it is much less than what would be required if new conversational scripts (and associated content) were created for each domain. This points to the inherent trade-off between human effort savings and dialogue richness that can be addressed by better automatic content extraction techniques. Table 1: An excerpt of an interaction between a learner and Watson Dialogue-Based Tutor (WDBT). Dialogue moves and responses from WDBT are left aligned. Learner responses are right aligned. What is the nature of gender differences? its is common for gender to be defined by sex Looks like I have something different... How do sociologists define sex? biological differneces Well said. Sex refers to the biological differences between men and women. There are 2 key facts left in this review. What is gender? idk No problem. I’ll help you get this! Gender refers to socially created differences between men and women. How are gender diff
The availability of open educational resources (OER) has enabled educators and researchers to access a variety of learning assessments online. OER communities are particularly useful for gathering multiple choice questions (MCQs), which are easy to grade, but difficult to design well. To account for this, OERs often rely on crowd-sourced data to validate the quality of MCQs. However, because crowds contain many non-experts, and are susceptible to question framing effects, they may produce ratings driven by guessing on the basis of surface-level linguistic features, rather than deep topic knowledge. Consumers of OER multiple choice questions (and authors of original multiple choice questions) would benefit from a tool that automatically provided feedback on assessment quality, and assessed the degree to which OER MCQs are susceptible to framing effects. This paper describes a model that is trained to use domain-naive strategies to guess which multiple choice answer is correct. The extent to which this model can predict the correct answer to an MCQ is an indicator that the MCQ is a poor measure of domain-specific knowledge. We describe an integration of this model with a front-end visualizer and MCQ authoring tool.
The potential impact of brain training methods for enhancing human cognition in healthy and clinical populations has motivated increasing public interest and scientific scrutiny. At issue is the merits of intervention modalities, such as computer-based cognitive training, physical exercise training, and non-invasive brain stimulation, and whether such interventions synergistically enhance cognition. To investigate this issue, we conducted a comprehensive 4-month randomized controlled trial in which 318 healthy, young adults were enrolled in one of five interventions: (1) Computer-based cognitive training on six adaptive tests of executive function; (2) Cognitive and physical exercise training; (3) Cognitive training combined with non-invasive brain stimulation and physical exercise training; (4) Active control training in adaptive visual search and change detection tasks; and (5) Passive control. Our findings demonstrate that multimodal training significantly enhanced learning (relative to computer-based cognitive training alone) and provided an effective method to promote skill learning across multiple cognitive domains, spanning executive functions, working memory, and planning and problem solving. These results help to establish the beneficial effects of multimodal intervention and identify key areas for future research in the continued effort to improve human cognition.
Recent advances in artificial intelligence and natural language processing greatly enhance the capabilities of intelligent tutoring systems. However, gathering a subject-appropriate corpus of training data remains challenging. In order to address this issue, we present a system based on a hybrid Wizard-of-Oz technique, which enables cognitive systems to work in tandem with a human operator (the "wizard"), to enhance collection of dialog variants.
Although the hippocampus experiences age-related anatomical and functional deterioration, the effects of aging vary across hippocampal-dependent cognitive processes. In particular, whether or not the hippocampus is known to be required for a spatial memory process is not an accurate predictor on its own of whether aging will affect performance. Therefore, the primary objective of this study was to compare the effects of healthy aging on a test of spatial pattern separation and a test of spatial relational processing, which are two aspects of spatial memory that uniquely emphasize the use of multiple hippocampal-dependent processes. Spatial pattern separation supports spatial memory by preserving unique representations for distinct locations. Spatial relational processing forms relational representations of objects to locations or between objects and other objects in space. To test our primary objective, 30 young (18-30 years; 21F) and 30 older participants (60-80 years; 21F) all completed a spatial pattern separation task and a task designed to require spatial relational processing through spatial reconstruction. To ensure aging effects were not due to inadequate time to develop optimal strategies or become comfortable with the testing devices, a subset of participants had extended practice across three sessions on each task. Results showed that older adults performed more poorly than young on the spatial reconstruction task that emphasized the use of spatial relational processing, and that age effects persisted even after controlling for pattern separation performance. Further, older adults performed more poorly on spatial reconstruction than young adults even after three testing sessions each separated by 7-10 days, suggesting effects of aging are resistant to extended practice and likely reflect genuine decline in hippocampal memory abilities.
OBJECTIVE:Subjective memory concerns (SMCs) in healthy older adults are associated with future decline and can indicate preclinical dementia. However, SMCs may be multiply determined, and often correlate with affective or psychosocial variables rather than with performance on memory tests. Our objective was to identify sensitive and selective methods to disentangle the underlying causes of SMCs.METHOD:Because preclinical dementia pathology targets the hippocampus, we hypothesized that performance on hippocampally dependent relational memory tests would correlate with SMCs. We thus administered a series of memory tasks with varying dependence on relational memory processing to 91 older adults, along with questionnaires assessing depression, anxiety, and memory self-efficacy. We used correlational, regression, and mediation analyses to compare the variance in SMCs accounted for by these measures.RESULTS:Performance on the task most dependent on relational memory processing showed a stronger negative association with SMCs than did other memory performance metrics. SMCs were also negatively associated with memory self-efficacy. These 2 measures, along with age and education, accounted for 40% of the variance in SMCs. Self-efficacy and relational memory were uncorrelated and independent predictors of SMCs. Moreover, self-efficacy statistically mediated the relationship between SMCs and depression and anxiety, which can be detrimental to cognitive aging.CONCLUSIONS:These data identify multiple mechanisms that can contribute to SMCs, and suggest that SMCs can both cause and be caused by age-related cognitive decline. Relational memory measures may be effective assays of objective memory difficulties, while assessing self-efficacy could identify detrimental affective responses to cognitive aging. (PsycINFO Database Record
Mnemonic processing engages multiple systems that cooperate and compete to support task performance. Exploring these systems' interaction requires memory tasks that produce rich data with multiple patterns of performance sensitive to different processing sub-components. Here we present a novel context-dependent relational memory paradigm designed to engage multiple learning and memory systems. In this task, participants learned unique face-room associations in two distinct contexts (i.e., different colored buildings). Faces occupied rooms as determined by an implicit gender-by-side rule structure (e.g., male faces on the left and female faces on the right) and all faces were seen in both contexts. In two experiments, we use behavioral and eye-tracking measures to investigate interactions among different memory representations in both younger and older adult populations; furthermore we link these representations to volumetric variations in hippocampus and ventromedial PFC among older adults. Overall, performance was very accurate. Successful face placement into a studied room systematically varied with hippocampal volume. Selecting the studied room in the wrong context was the most typical error. The proportion of these errors to correct responses positively correlated with ventromedial prefrontal volume. This novel task provides a powerful tool for investigating both the unique and interacting contributions of these systems in support of relational memory.
Counterfactual reasoning is a hallmark of human thought, enabling the capacity to shift from perceiving the immediate environment to an alternative, imagined perspective. Mental representations of counterfactual possibilities (e.g., imagined past events or future outcomes not yet at hand) provide the basis for learning from past experience, enable planning and prediction, support creativity and insight, and give rise to emotions and social attributions (e.g., regret and blame). Yet remarkably little is known about the psychological and neural foundations of counterfactual reasoning. In this review, we survey recent findings from psychology and neuroscience indicating that counterfactual thought depends on an integrative network of systems for affective processing, mental simulation, and cognitive control. We review evidence to elucidate how these mechanisms are systematically altered through psychiatric illness and neurological disease. We propose that counterfactual thinking depends on the coordination of multiple information processing systems that together enable adaptive behavior and goal-directed decision making and make recommendations for the study of counterfactual inference in health, aging, and disease.
Seizure detection is rapidly improving thanks to novel computational approaches [1] and public databases of electrocorticography (ECoG) data [2]. But these approaches rarely model the high-frequency oscillations that underlie seizure pathology [3]. In the current work, we employ a phase-locked loop (PLL) neural network [4] to emulate a mesoscale circuit undergoing high-frequency oscillation. We phase-lock the nodes of the emulator to raw voltages recorded from chronically implanted ECoG electrodes in a canine model of epilepsy, and demonstrate that the emulator experiences uncontrolled phase oscillation far in advance of either behaviorally observed seizure activity or fluctuations in ECoG voltages. Using distance weighting to train the epilepsy network emulator, we localize the ECoG electrodes responsible for destabilization, and present measures sensitive to these phase disruptions to establish a forecasting period for ictal activity. We discuss how phase oscillations from real-time epilepsy network emulation could serve as a closed-loop feedback control signal to interrupt ictal activity.
The hippocampus has been implicated in a diverse set of cognitive domains and paradigms, including cognitive mapping, long-term memory, and relational memory, at long or short study-test intervals. Despite the diversity of these areas, their association with the hippocampus may rely on an underlying commonality of relational memory processing shared among them. Most studies assess hippocampal memory within just one of these domains, making it difficult to know whether these paradigms all assess a similar underlying cognitive construct tied to the hippocampus. Here we directly tested the commonality among disparate tasks linked to the hippocampus by using PCA on performance from a battery of 12 cognitive tasks that included two traditional, long-delay neuropsychological tests of memory and two laboratory tests of relational memory (one of spatial and one of visual object associations) that imposed only short delays between study and test. Also included were different tests of memory, executive function, and processing speed. Structural MRI scans from a subset of participants were used to quantify the volume of the hippocampus and other subcortical regions. Results revealed that the 12 tasks clustered into four components; critically, the two neuropsychological tasks of long-term verbal memory and the two laboratory tests of relational memory loaded onto one component. Moreover, bilateral hippocampal volume was strongly tied to performance on this component. Taken together, these data emphasize the important contribution the hippocampus makes to relational memory processing across a broad range of tasks that span multiple domains.
Successful behavior requires actively acquiring and representing information about the environment and people, and manipulating and using those acquired representations flexibly to optimally act in and on the world. The frontal lobes have figured prominently in most accounts of flexible or goal-directed behavior, as evidenced by often-reported behavioral inflexibility in individuals with frontal lobe dysfunction. Here, we propose that the hippocampus also plays a critical role by forming and reconstructing relational memory representations that underlie flexible cognition and social behavior. There is mounting evidence that damage to the hippocampus can produce inflexible and maladaptive behavior when such behavior places high demands on the generation, recombination, and flexible use of information. This is seen in abilities as diverse as memory, navigation, exploration, imagination, creativity, decision-making, character judgments, establishing and maintaining social bonds, empathy, social discourse, and language use. Thus, the hippocampus, together with its extensive interconnections with other neural systems, supports the flexible use of information in general. Further, we suggest that this understanding has important clinical implications. Hippocampal abnormalities can produce profound deficits in real-world situations, which typically place high demands on the flexible use of information, but are not always obvious on diagnostic tools tuned to frontal lobe function. This review documents the role of the hippocampus in supporting flexible representations and aims to expand our understanding of the dynamic networks that operate as we move through and create meaning of our world.
Hippocampal damage causes profound yet circumscribed memory impairment across diverse stimulus types and testing formats. Here, within a single test format involving a single class of stimuli, we identified different performance errors to better characterize the specifics of the underlying deficit. The task involved study and reconstruction of object arrays across brief retention intervals. The most striking feature of patients' with hippocampal damage performance was that they tended to reverse the relative positions of item pairs within arrays of any size, effectively "swapping" pairs of objects. These "swap errors" were the primary error type in amnesia, almost never occurred in healthy comparison participants, and actually contributed to poor performance on more traditional metrics (such as distance between studied and reconstructed location). Patients made swap errors even in trials involving only a single pair of objects. The selectivity and severity of this particular deficit creates serious challenges for theories of memory and hippocampus. © 2013 Wiley Periodicals, Inc.