Everyone agrees that real cognition requires much more than static pattern recognition. In particular, it requires the ability to learn sequences of patterns (or actions) But learning sequences really means being able to learn multiple sequences, one after the other, wi thout the most recently learned ones erasing the previously learned ones. But if catastrophic interference is a problem for the sequential learning of individual patterns, the problem is amplified many times over when multiple sequences of patterns have to be learned consecutively, because each new sequence consists of many linked patterns. In this paper we will present a connectionist architecture that would seem to solve the problem of multiple sequence learning using pseudopatterns.
A neural network consisting of a sensorimotor module associated with a dynamic memory (DM) is presented. The sensorimotor module is made of a Kohonen layer where exteroceptive and proprioceptive sensory information are combined on a functional map. The sensory layer controls a motor layer which drives the effectors of either a numerically simulated arm or an artificial jointed arm. After a learning phase, the arm is able to perform goal-directed movements. The dynamic memory is designed to be able to learn, memorize and execute temporal sequences. One of the main features of DM is that the repetition of learned sequences is triggered by the occurrence of an item relating to its identity, and the rhythm at which the model produces the response sequence can be controlled and freely modulated by a sub-system mimicking attcntional processes. The association of both modules give rise to a neural network that can learn spatial shapes and translate them in motor terms by means of a simulated or robotic jointed arm, at any point in its working space (translation invariance) and at any suitable amplitude scale (size invariance).
According to predictions made by ACV98 connectionist model of reading, behavioral experiments have shown that syllabic length affects naming latencies for pseudo-words but not for words. ACV98 postulates the existence of two successive reading procedures. According to it, any orthographical object (word or pseudo-word) is first submitted to a global processing which aims to find a "lexical familiarity" computed from previously experienced words. If it is found, the corresponding phonological form of the stimulus is activated and reading is performed. If it is not found, a subsequent analytical processing is performed within the input stimulus using the same route, in order to extract familiar orthographic components (typically, syllables) as well as their corresponding phonological forms. The syllabic phonological forms are successively and temporarily maintained in a phonological buffer, before being assembled into a whole phonological form. Overall, this model predicts that words are read by using the global procedure, while the pseudo-words are read by using the analytical procedure. The present event-related fMRI study aimed to assess the effect of syllabic length on cerebral activity during reading, in order to obtain additional anatomo-functional information as support to this model as well as to behavioral results. Based on ACV98 predictions and in terms of cerebral network, we hypothesized (1) a lexicality effect: there should be no differences between words and pseudo-words, they are processed within a common network of cerebral regions, and (2) a syllable length effect: the length influences cerebral activity only during pseudo-words reading because only long pseudo-words are processed following an analytical procedure involving supplementary visual analysis, visuo-spatial attention and working memory processes. Eight right-handed volunteers performed a silent reading task on French printed words and pseudo-words. A pseudo-randomized event-related. paradigm with 6 types (words=W and pseudo-words=PW composed of 1, 2 and 3 syllables) of stimuli was used. Data processing was performed using SPM'99. Our results have shown that (1) Words and pseudo-words involved common mechanisms such as visuo-orthographic, phonological, attentional and motor processes. No process was significantly more involved for one or the other of two types of stimuli (word or pseudo-word). This result is a priori more in agreement with "single way models" and particularly with ACV'98 than with dual-route models. Nonetheless the dual-route model cannot be excluded; (2) A length effect on cerebral activity was obtained only during pseudo-word reading suggesting an analytical procedure involvement during reading of long pseudo-words, as ACV'98 predicts.
Attractors of nonlinear neural systems are at the core of the memory self-refreshing mechanism of human memory models that suppose memories are dynamically maintained in a distributed network [Ans, B., and Rousset, S. (1997), ‘Avoiding Catastrophic Forgetting by Coupling Two Reverberating Neural Networks’ Comptes Rendus de l'Académie des Sciences Paris, Life Sciences, 320, 989–997; Ans, B., and Rousset, S. (2000), ‘Neural Networks with a Self-Refreshing Memory: Knowledge Transfer in Sequential Learning Tasks Without Catastrophic Forgetting’, Connection Science, 12, 1–19; Ans, B., Rousset, S., French, R.M., and Musca, S.C. (2002), ‘Preventing Catastrophic Interference in Multiple-Sequence Learning Using Coupled Reverberating Elman Networks’, in Proceedings of the 24th Annual Meeting of the Cognitive Science Society, eds. W.D. Gray and C.D. Schunn, Mahwah, NJ: Lawrence Erlbaum Associates, pp. 71–76; Ans, B., Rousset, S., French, R.M., and Musca, S.C. (2004), ‘Self-Refreshing Memory in Artificial Neural Networks: Learning Temporal Sequences Without Catastrophic Forgetting’, Connection Science, 16, 71–99; Ans, B. (2004), ‘Sequential Learning in Distributed Neural Networks Without Catastrophic Forgetting: A Single and Realistic Self-Refreshing Memory can do it’, Neural Information Processing-Letters and Reviews, 4, 27–32]. Are humans able to learn never seen items from attractor patterns generated by a highly distributed artificial neural network? First, an opposition method was implemented to ensure that the attractors are not the items used to train the network, the source items: attractors were selected to be more similar (both at the exemplar and the centroïd level) to some control items than to the source items. In spite of this very severe selection, blank networks trained only on selected attractors performed better at test on the never seen source items than on the never seen control items. The results of two behavioural experiments using the opposition method show that humans exhibit more familiarity with the never seen source items than with the never seen control items, just as networks do. Thus, humans are sensitive to the particular type of information that allows distributed artificial neural networks to dynamically maintain their memory, and this information does not amount to the exemplars used to train the network that produced the attractors.
The cognitive mechanisms involved in polysyllabic pseudo-word processing—and their neurobiological correlates—were studied through the analysis of length effects on French words and pseudo-words in reading and lexical decision. Connectionist simulations conducted on the ACV98 network (Ans, B., Carbonnel, S., Valdois, S., 1998. A connectionist multiple-trace memory model for polysyllabic word reading. Psychol. Rev. 105, 678–723) paralleled the behavioral data in showing a strong length effect on naming latencies for pseudo-words only and the absence of length effect for both words and pseudo-words in lexical decision. Length effects in reading were characterized at the neurobiological level by a significant and specific activity increase for pseudo-words as compared to words in the right lingual gyrus (BA 19), the left superior parietal lobule and precuneus (BA7), the left middle temporal gyrus (BA21) and the left cerebellum. The behavioral results suggest that polysyllabic pseudo-word reading mainly relies on an analytic procedure. At the biological level, additional activations in visual and visual attentional brain areas during long pseudo-word reading emphasize the role of visual and visual attentional processes in pseudo-word reading. The present findings place important constraints on theories of reading in suggesting the involvement of a serial mechanism based on visual attentional processing in pseudo-word reading.
A travers l'étude des effets de longueur en lecture et décision lexicale selon la nature des items présentés (mots ou pseudo-mots), cette étude tente d'évaluer la nature des procédures impliquées dans le traitement des pseudo-mots longs. Les résultats expérimentaux montrent, en lecture, l'existence d'une interaction lexicalité x longueur qui n'est pas retrouvée en décision lexicale. De plus, l'effet massif de longueur observé sur les pseudo-mots en lecture immédiate ne peut être dû au processus de génération articulatoire puisque cet effet disparaît en situation de lecture différée. Les simulations effectuées dans le cadre du modèle multitrace de lecture (Ans, Carbonnel & Valdois, 1998) sont largement compatibles avec les résultats expérimentaux suggérant que la lecture des pseudo-mots repose sur une procédure analytique n'intervenant ni en lecture de mots ni en décision lexicale.
Creating False Memories in Humans with an Artificial Neural Network: Implications for Theories of Memory Consolidation Serban C. Musca (Serban.Musca@upmf-grenoble.fr) Stephane Rousset (Stephane.Rousset@upmf-grenoble.fr) Bernard Ans (Bernard.Ans@upmf-grenoble.fr) Psychology and NeuroCognition Lab - CNRS UMR 5105, University Pierre Mendes France (Grenoble 2), 1251 Avenue Centrale, BP 47, 38040 Grenoble cedex 9, France Abstract Building on the human memory model that consider LTM to be similar to a distributed network (McClelland, McNaughton & O'Reilly, 1995), and informed by the recent solutions to catastrophic forgetting that suppose memories are dynamically maintained in a dual architecture through a memory self-refreshing mechanism (Ans & Rousset, 1997, 2000; Ans et al., 2002, 2004; French, 1997), we checked whether false memories of never seen (target) items can be created in humans by exposure to pseudo-patterns generated from random input in an artificial neural network (previously trained on the target items). In a behavioral experiment using an opposition method it is shown that the answer is yes: Though the pseudo-patterns presented to the participants were selected so as to resemble (both at the exemplar and the prototype level) more the control items than the target items, the participants exhibited more familiarity for the target items previously learned by the artificial neural network. This behavioral result analogous to the one found in simulations indicates that humans, like distributed neural networks, are able to make use of the information the memory self-refreshing mechanism is based upon. The implications of these findings are discussed in the framework of memory consolidation. Keywords: distributed information; neural networks; human memory; representation format in human memory; memory self-refreshing; exemplar theory; prototype theory; indirect memory test; familiarity; perceptual fluency. Introduction Can information be transported between a GDN–a multi-layered network trained by a gradient descent learning procedure–and humans? This question is central to models of human memory that suppose that LTM is similar to a distributed network (e.g. McClelland, McNaughton & O'Reilly, 1995) and that memories are dynamically maintained in a dual architecture through a memory self- refreshing mechanism (Ans & Rousset, 1997, 2000; Ans et al., 2002, 2004; French, 1997). GDN's memory gradually emerges as a result of the processing of the training exemplars: the connection weights between the processing units reach values that allow the network to perform correctly. Thus the memory of a given trained GDN can be conceived as the particular set of connections weights between its processing units. Of course, when trained on a new set of exemplars S 2 , the connection weights of a network previously trained on a set of exemplars S 1 change in order to allow the network to perform correctly on S 2 , and this new connection weight set does not allow the network to perform correctly on S 1 any more (catastrophic interference or catastrophic forgetting: McCloskey & Cohen, 1989; Ratcliff, 1990). To get round this obstacle, an obvious solution is to train S 1 and S 2 concurrently, thus transforming sequential learning (i.e. first S 1 , then S 2 ) into concurrent learning (i.e. S 1 , and S 2 at the same time). However, concurrent learning relies on the assumption that S 1 is still available when S 2 is to be learned, an unreasonable assumption when GDNs are used to simulate human memory phenomena: Every old exemplar is not available nor is it learned anew each time some new exemplars are learned (Blackmon et al., 2004). A first step towards a more plausible solution to the problem of catastrophic forgetting in GDNs in sequential learning tasks is due to Robins (1995): Once a network has been trained on S 1 , its memory is sampled, thus generating random input-computed output pairs (or pseudo-items) that are stored in a non-neuromimetic memory; then, instead of training the network on S 2 only, it is trained both on S 2 and on the stored pseudo-items. If this solution makes it possible to reduce catastrophic forgetting in the absence of S 1 exemplars, it also resorts to an implausible copy-paste procedure in order to store the pseudo-items before they are used as training material. The next solutions (Ans & Rousset, 1997; French, 1997) avoid the copy-paste procedure by having recourse to GDN architectures that are able to learn on the fly the random input-computed output pairs. For instance, Ans & Rousset's (1997) architecture is made of two separate GDNs, NET1 and NET2; once trained on S 1 , NET1 generates reverberated random input-computed output pairs (called pseudo-patterns, PPs) that are used to train NET2. Then, when NET1 is to learn a new set of exemplars S 2 , NET1 is not only trained on S 2 but also on PPs generated this time in NET2 (and conveying information on S 1 ). Were a third new training set S 3 to be learned, NET1's memory would first be transmitted to NET2 through PPs, then NET1 would be trained on both S 3 and PPs generated in NET2 (now conveying information on both S 1 and S 2 ). To sum up, this architecture is very efficient in avoiding catastrophic
While retroactive interference (RI) is a well-known phenomenon in humans, the differential effect of the structure of the learning material was only seldom addressed. Mirman and Spivey (2001, Connection Science, 13: 257-275) reported on behavioural results that show more RI for the subjects exposed to 'Structured' items than for those exposed to 'Unstructured' items. These authors claimed that two complementary memory systems functioning on radically different neural mechanisms are required to account for the behavioural results they reported. Using the same paradigm but controlling for proactive interference, we found the opposite pattern of results, that is, more R1 for subjects exposed to 'Unstructured' items than for those exposed to 'Structured' items (experiment 1). Two additional experiments showed that this structure effect on R1 is a genuine one. Experiment 2 confirmed that the design of experiment I forced the subjects from the 'Structured' condition to learn the items at the exemplar level, thus allowing for a close match between the two to-be-compared conditions (as 'Unstructured' condition items can be learned only at the exemplar level). Experiment 3 verified that the subjects from the 'Structured' condition could generalize to novel items. Simulations conducted with a three-layer neural network, that is, a single-memory system, produced a pattern of results that mirrors the structure effect reported here. By construction, Mirman and Spivey's architecture cannot simulate this behavioural structure effect. The results are discussed within the framework of catastrophic interference in distributed neural networks, with an emphasis on the relevance of these networks to the modelling of human memory.
In sequential learning tasks artificial distributed neural networks forget catastrophically, that is, new learned information most often erases the one previously learned. This major weakness is not only cognitively implausible, as human gradually forget, but disastrous for most practical applications. An efficient solution to catastrophic forgetting has been recently proposed for backpropagation networks, the reverberating self-refreshing mechanism: when new external events are learned they have to be interleaved with internally-generated pseudo-events (from simple random activations) reflecting the previously learned information. Since self-generated patterns cannot be learned by a same backpropagation network, because desired targets are lacking, this solution used two complementary networks. In the present paper it is proposed a new self-refreshing mechanism based on a single-network architecture that can learn its own production reflecting its history (i.e., a self-learning ability). In addition, in place of backpropagation, widely considered to be not biologically realistic, a more plausible learning rule is used: the deterministic version of the Contrastive Hebbian Learning algorithm, or CHL. Simulations of sequential learning tasks show that the proposed single self-refreshing memory has the ability to avoid catastrophic forgetting.
Following Mirman and Spivey's investigation [12], Musca, Rousset and Ans conducted a study on the influence of the nature of the to-be-learned material on retroactive interference (RI) in humans [13]. More RI was found for unstructured than for structured material, a result opposed to that of Mirman and Spivey [12]. This paper first presents two simulations. The first, using a three-layer backpropagation hetero-associator produced a pattern of RI results that mirrored qualitatively the structure effect on RI found in humans [13]. However the level of RI was high. In the second simulation the Dual Reverberant memory Self-Refreshing neural network model (DRSR) of Ans and Rousset [1, 2] was used. As expected, the global level of RI was reduced and the structure effect on RI was still present. We further investigated the functioning of DRSR in this situation. A proactive interference (PI) was observed, and also a structure effect on PI. Furthermore, the structure effect on RI and the structure effect on PI were negatively correlated. This trade-off between structure effect on RI and structure effect on PI found in simulation points to an interesting potential phenomenon to be investigated in humans.
While humans forget gradually, highly distributed connectionist networks forget catastrophically: newly learned information often completely erases previously learned information. This is not just implausible cognitively, but disastrous practically. However, it is not easy in connectionist cognitive modelling to keep away from highly distributed neural networks, if only because of their ability to generalize. A realistic and effective system that solves the problem of catastrophic interference in sequential learning of ‘static’ (i.e. non-temporally ordered) patterns has been proposed recently (Robins 1995, Connection Science, 7: 123–146, 1996, Connection Science, 8: 259–275, Ans and Rousset 1997, CR Académie des Sciences Paris, Life Sciences, 320: 989–997, French 1997, Connection Science, 9: 353–379, 1999, Trends in Cognitive Sciences, 3: 128–135, Ans and Rousset 2000, Connection Science, 12: 1–19). The basic principle is to learn new external patterns interleaved with internally generated ‘pseudopatterns’ (generated from random activation) that reflect the previously learned information. However, to be credible, this self-refreshing mechanism for static learning has to encompass our human ability to learn serially many temporal sequences of patterns without catastrophic forgetting. Temporal sequence learning is arguably more important than static pattern learning in the real world. In this paper, we develop a dual-network architecture in which self-generated pseudopatterns reflect (non-temporally) all the sequences of temporally ordered items previously learned. Using these pseudopatterns, several self-refreshing mechanisms that eliminate catastrophic forgetting in sequence learning are described and their efficiency is demonstrated through simulations. Finally, an experiment is presented that evidences a close similarity between human and simulated behaviour.
The present study describes two Frenchteenagers with developmental reading andwriting impairments whose performance wascompared to that of chronological age andreading age matched non-dyslexic participants.Laurent conforms to the pattern of phonologicaldyslexia: he exhibits a poor performance inpseudo-word reading and spelling, producesphonologically inaccurate misspellings butreads most exception words accurately. Nicolas,in contrast, is poor in reading and spelling ofexception words but is quite good atpseudo-word spelling, suggesting that hesuffers from surface dyslexia and dysgraphia.The two participants were submitted to anextensive battery of metaphonological tasks andto two visual attentional tasks. Laurentdemonstrated poor phonemic awareness skills butgood visual processing abilities, while Nicolasshowed the reverse pattern with severedifficulties in the visual attentional tasksbut good phonemic awareness. The presentresults suggest that a visual attentionaldisorder might be found to be associated withthe pattern of developmental surface dyslexia.The present findings further show thatphonological and visual processing deficits candissociate in developmental dyslexia.
This event-related fMRI study has assessed the cerebral activations obtained during the reading of words and of pseudo-words of varying length (one, two and three syllables). Eight right-handed volunteers were examined. A pseudo-randomized fMRI paradigm with six types of stimuli was applied and the SPM'99 software was used for data processing. The number of cerebral regions involved in reading increased with the stimulus length, both for words and for pseudo-words. This concerned not only regions related to sensory-motor aspects of reading but also regions related to more "central" language processes. Independently of the length, the reading of words and of pseudo-words activated the same regions, an observation consistent with the connectionist models for reading. Considering the length of the stimuli, we obtained significant differences, in terms of cerebral regions, only between polysyllabic words and pseudo-words, not between monosyllabic words and pseudo-words. These results are in line with the connectionist view of reading, especially with the multi-trace connectionist model of Ans et al (3).
Everyone agrees that real cognition requires much more than static pattern recognition. In particular, it requires the ability to learn sequences of patterns (or actions) But learning sequences really means being able to learn multiple sequences, one after the other, wi thout the most recently learned ones erasing the previously learned ones. But if catastrophic interference is a problem for the sequential learning of individual patterns, the problem is amplified many times over when multiple sequences of patterns have to be learned consecutively, because each new sequence consists of many linked patterns. In this paper we will present a connectionist architecture that would seem to solve the problem of multiple sequence learning using pseudopatterns. Introduction Building a robot that could unfailingly recognize and respond to hundreds of objects in the world – apples, mice, telephones and paper napkins, among them – would unquestionably constitute a major artificial intelligence tour de force. But everyone agrees that real cognition requires much more than static pattern recognition. In particular, it requires the ability to learn sequences of patterns (or actions). This was the primary reason for the development of the simple recurrent network (SRN, Elman, 1990) and the many variants of this architecture. But learning sequences means more than being able to learn a single, isolated sequence of patterns: it means being able to learn multiple sequences, one after the other, without the most recently learned ones erasing the previously learned ones. But if catastrophic interference – the phenomenon whereby new learning completely erases old learning – is a problem with static pattern learning (McCloskey & Cohen, 1989; Ratcliff, 1990), the problem is amplified many times over when multiple sequences of patterns have to be learned consecutively, because each sequence consists of many new linked patterns. What hope is there for a previously learned sequence of patterns to survive after the network has learned a new sequence consisting of many individual patterns? In this paper, we will present a connectionist architecture that solves the problem of multiple sequence learning. Catastrophic interference The problem of catastrophic interference (or forgetting) has been with the connectionist community for well over a decade now (McCloskey & Cohen, 1989; Ratcliff, 1990; for a review see Sharkey & Sharkey, 1995). Catastrophic forgetting occurs when newly learned information suddenly and completely erases information that was previously learned by the network, a phenomenon that is not only implausible cognitively, but disastrous for most practical applications. The problem has been studied by numerous authors over the past decade (see French, 1999 for a review). The problem is that the very property – a single set of weights to encode information – that gives connectionist networks their remarkable abilities of generalization and graceful degradation in the presence of incomplete information are also the root cause of catastrophic interference (see, for example, French, 1992). Various authors (Ans & Rousset, 1997, 2000; French, 1997; Robins, 1995) have developed systems that rehearse on pseudo-episodes (or pseudopatterns), rather than on the real items that were previously learned. The basic principle of this mechanism is when learning new external patterns to interleave them with internally-generated pseudopatterns. These latter patterns, self-generated by the network from random activation, reflect (but are not identical to) the previously learned information. It has now been established that this pseudopattern rehearsal method effectively eliminates catastrophic forgetting. A serious problem remains, however, and that is this: cognition involves more than being able to sequentially learn a series of "static" (non-temporal) patterns without interference. It is of equal importance to be able to serially learn many of temporal sequences of patterns. We will propose an pseudopattern-based architecture that can effectively learn multiple temporal pat terns consecutively. The key insight of this paper is this: Once an SRN has learned a particular sequence, each pseudopattern generated by that network reflects the entire sequence (or set of sequences) that has been learned .
The present study describes two French teenagers with developmental reading and writing impairments whose performance was compared to that of chronological age and read- ing age matched non-dyslexic participants. Laurent conforms to the pattern of phonological dyslexia: he exhibits a poor performance in pseudo-word reading and spelling, produces phonologically inaccurate misspellings but reads most exception words accurately. Nicolas, in contrast, is poor in reading and spelling of exception words but is quite good at pseudo- word spelling, suggesting that he suffers from surface dyslexia and dysgraphia. The two participants were submitted to an extensive battery of metaphonological tasks and to two visual attentional tasks. Laurent demonstrated poor phonemic awareness skills but good visual processing abilities, while Nicolas showed the reverse pattern with severe difficulties in the visual attentional tasks but good phonemic awareness. The present results suggest that a visual attentional disorder might be found to be associated with the pattern of developmental surface dyslexia. The present findings further show that phonological and visual processing deficits can dissociate in developmental dyslexia. Summary 3D-Quantitative structure-activity relationships of amino acid and thioxo amino acid pyrrolidides and thiazolidides as inhibitors of bacterial (F. meningosepticum) and human placenta prolyl oligopeptidase were investigated using CoMFA and COMSIA approaches. For the bacterial POP, the best CoMFA model obtained from 16 inhibitors is a five-component model with the following statistics: R2 cv = 0.852 and RMSEcv = 0.213 for the cross-validation, and R2 = 0.954 and RMSE = 0.119 for the fitted. The best COMSIA model obtained is a two-component model with the following statistics: R2 cv = 0.826 and RMSEcv = 0.203 for the cross-validation, and R2 = 0.915 and RMSE = 0.142 for the fitted. For the human placenta POP, the best CoMFA model obtained is again a five-component model. This model has the following statistics: R2 cv = 0.885 and RMSEcv = 0.470 for the cross-validation, and R2 = 0.966 and RMSE = 0.256 for the fitted. The best COMSIA model obtained is a two- component model with the following statistics: R2cv = 0.922 and RMSEcv = 0.339 for the cross-validation, and R2 = 0.959 and RMSE = 0.247 for the fitted. A comparison of the homology models of the human and the bacterial POP reveals similarities and differences in the binding sites of the two proteins and provides some clues for the observed selectivity.
Monica Baciu1, Olivier David2, Mathilde Pachot-Clouard3, Serge Carbonnel1, Bernard Ans1, Christoph Segebarth3 1Laboratoire de Psychologie Experimentale, Universite Pierre Mendes-France, BP 47, Grenoble Cedex 9, France; 2CNRS UPR 640 LENA, Hopital de la Salpetriere, 47, Boulevard de l'Hopital, Paris Cedex 13, France; 3INSERM U438, Centre Hospitalier Universitaire, Pavillon B, Grenoble Cedex 9, France;