Prével and colleagues reported excitatory learning with a backward conditioned stimulus (CS) in a conditioned reinforcement preparation. Their results add to existing evidence of backward CSs sometimes being excitatory and were viewed as challenging the view that learning is driven by prediction error reduction, which assumes that only predictive (i.e., forward) relationships are learned. The results instead were consistent with the assumptions of both Miller's Temporal Coding Hypothesis and Wagner's Sometimes Opponent Processes (SOP) model. The present experiment extended the conditioned reinforcement preparation developed by Prével et al. to a backward second-order conditioning preparation, with the aim of discriminating between these two accounts. We tested whether a second-order CS can serve as an effective conditioned reinforcer, even when the first-order CS with which it was paired is a backward CS that elicits no responding. Evidence of conditioned reinforcement was found, despite no conditioned response (CR) being elicited by the first-order backward CS. The evidence of second-order conditioning in the absence of excitatory conditioning to the first-order CS is interpreted as a challenge to SOP. In contrast, the present results are consistent with the Temporal Coding Hypothesis and constitute a conceptual replication in humans of previous reports of excitatory second-order conditioning in rodents with a backward CS. The proposal is made that learning is driven by "discrepancy" with prior experience as opposed to " prediction error."
In the present study, excitatory backward conditioning was assessed in a conditioned reinforcement paradigm. The experiment was conducted with human subjects and consisted of five conditions. In all conditions, US reinforcing value (i.e. time reduction of a timer) was assessed in phase 1 using a concurrent FR schedule, with one response key leading to US presentation and the other key leading to no-US. In phase 2, two discrete stimuli, S+ and S−, were paired with US and no-US respectively using an operant contingency. For three groups, backward contingencies were arranged, and two of these were designed to rule out a trace (forward) conditioning interpretation of the results. The two other groups served as control conditions (forward and neutral conditions). Finally, in phase 3 for all groups the CSs were delivered in a concurrent FR schedule similar to phase 1, but with no US. Responding during phase 3 showed conditioned reinforcement effects and hence excitatory backward conditioning. Implications of the results for conditioned reinforcement models are discussed.
A review of the literature concerning space perception shows that a recurrent question is about specification. This question refers to the modalities of environmental properties (distal stimulus) representation by proximal stimulus (Heider, 1926). Two different popular theories known as ecological and inferential have been proposed, while they still share the same informational interpretation of this problematic. An alternative interpretation is therefore proposed in terms of reinforcement contingencies rather than information. So, the need of reconsidering the function of perceptive activities is shown, just as the functional similarity between this kind of activity and others behaviors of organisms. Moreover, the possibility to shift research from optical and acoustical questions towards new experimental and theoretical investigations is highlighted. Finally, it is argued that this interpretation offers naturally a paradigm which provides experimental advantages over more classical one. This paradigm is that of schedules of reinforcemen.
We propose an operant approach to the emergence of cooperation in the iterated Prisoners' Oilemma (IPO). The approach yields to the design of reinforcementlearning agents whose behavioral repertoire includes not only cooperation-related behaviors, but also controlling behaviors that may influence the behavior of the other player. The task of an agent is to learn to coordinate its own cooperation- and controlrelated behaviors with those of the other agents. It is suggested that this situation is closer to natural cooperative situations than the classical approaches to the IPO.
This paper deals with an extension of behavioral principies to the study 01 social situations, In order to understand how individual contingencies are structured in a collective situation, we propose to investigate social situations using experiments with humans, in conjunction with simulations with behavioral artificial agents, In the first part, we present results obtained with humans in a minimal social situation. In this kind of situation, participants unknowingly interact by reinforcing and punishing each other. We observed that cooperation increased despite the fact that participants were unaware of the consequences of their behaviors, for they were not informed that they were in a social situation. The second part describes the implementation of five reinforcementlearning strategies in a computer simulation, whose performances were compared to the one observed in humans in an analogous situation, The Staddon-Zhang strategy was the best one to optimize cooperation and model human performance.
Saccade and smooth pursuit are the eye movements used by primates to shift gaze. In this article we review evidence for the effects of reinforcement on several dimensions of these responses such as their latencies, velocities or amplitudes. We propose that these responses are operant behaviours controlled by their consequences on performance of visually guided tasks. Studying the conditions under which particular eye movement patterns might emerge from the cumulative effects of reinforcement provides critical insights about how motor responses are attuned to environmental exigencies.
The purpose of this study was to evaluate the effectiveness of a high-probability (high-p) request sequence as a means of increasing compliance with medical examination tasks. Participants were children who had been diagnosed with autism and who exhibited noncompliance during general medical examinations. The inclusion of the high-p request sequence effectively increased compliance with medical examination tasks. In addition, the procedure was efficient, could be implemented by parents and medical professionals, and did not involve aversive procedures.
Justification of effort is a form of cognitive dissonance in which the subjective value of an outcome is directly related to the effort that went into obtaining it. However, it is likely that in social contexts (such as the requirements for joining a group) an inference can be made (perhaps incorrectly) that an outcome that requires greater effort to obtain in fact has greater value. Here we present evidence that a cognitive dissonance effect can be found in children under conditions that offer better control for the social value of the outcome. This effect is quite similar to contrast effects that recently have been studied in animals. We suggest that contrast between the effort required to obtain the outcome and the outcome itself provides a more parsimonious account of this phenomenon and perhaps other related cognitive dissonance phenomena as well. Research will be needed to identify cognitive dissonance processes that are different from contrast effects of this kind. nt]mis|This research was facilitated by a fellowship from the Fulbright Scholar Program and the Nord-Pas de Calais Regional Council, as well as by a visiting professorship at the University of Lille III for T.R.Z. Preparation of the article was facilitated by National Institute of Mental Health Grant MH 63726 to T.R.Z.
Humans prefer (conditioned) rewards that follow greater effort (Aronson & Mills, 1959). This phenomenon can be interpreted as evidence for cognitive dissonance (or as justification of effort) but may also result from (1) the contrast between the relatively greater effort and the signal for reinforcement or (2) the delay reduction signaled by the conditioned reinforcer. In the present study, we examined the effect of prior force and prior time to produce stimuli associated with equal reinforcement. As expected, pressing with greater force or for a longer time was less preferred than pressing with less force or for a shorter time. However, participants preferred the conditioned reinforcer that followed greater force and more time. Furthermore, participants preferred a long duration with no force requirement over a shorter duration with a high force requirement and, consistent with the contrast account but not with the delay reduction account, they preferred the conditioned stimulus that followed the less preferred, shorter duration, high-force event. Within-trial contrast provides a more parsimonious account than justification of effort, and a more complete account than delay reduction.
Une revue de la littérature concernant la perception de l'espace permet de mettre en évidence qu'une question récurrente est celle de la «spécification». Cette interrogation renvoie aux modalités de représentation des propriétés de l'environnement (stimulus distaux) par les stimulus proximaux (Heider, 1926). Différentes réponses ont ainsi été apportées sous les noms de théories inférentielles (Knill, 2001; Marr, 1982) et écologiques (Gibson, 1966), tout en partageant une même interprétation informationnelle de cette problématique. Ces débats étant toujours d'actualité, une réinterprétation de ce problème de la spécification en termes de contingences de renforcement (Skinner, 1938) est ici proposée. Ainsi, la nécessité de reconsidérer la fonction des activités perceptives est rendue apparente, de même que la similarité fonctionnelle entre ces activités et les autres comportements de l'organisme. La possibilité de déplacer l'attention des problèmes d'optique ou d'acoustique vers d'autres questions expérimentales et théoriques est de plus soulignée. Il est enfin soutenu que cette interprétation fournie avec elle un paradigme offrant un meilleur contrôle expérimental que les approches classiques. Ce paradigme est celui des programmes de renforcement.
A critique of the operant procedures used to study the ontogenesis of temporal regulation in infants and children is presented. The main thesis is that there is a transition in such regulation from nonhuman-like contingency-governed operant behavior to verbally-governed behavior in humans. Some studies have shown that responding of infants and young children during fixed-interval (FI) and differential-reinforcement-of-low rate (DRL) schedules is typical of the behavior of nonhumans under such schedules, but other studies have yielded different results. These inconsistent data may be explained by procedural differences between the experiments. The understanding and modelling of the ontogenesis of temporal regulation in humans require further experimental analyses that take into consideration the methodological differences between human and nonhuman animal studies outlined in this review.
We propose to verify the relevance of a selectionist approach to the spatial organization of behaviour. Reinforcement is retained as the principle of selection. Some experimental data show that auditory stimulus influences infant reaching behaviour. In the same way, reinforcement of the leg positioning with visual or tactile stimuli has been exhibited. Thus, the generality of this principle is confirmed. We propose using reinforcement algorihms to formalise the shaping of reaching behaviour.
In this paper, we present MAABAC, a generic model for building adaptive agents: they learn new behaviors by interacting with their environment. These agents adapt their behavior by way of reinforcement learning, namely temporal difference methods. MAABAC is presented in its generality and then, different instantiations of the generic model are presented and experiments are reported. These experiments show the strength of this way of learning.
Markovian decision problems are a kind of optimization problems in which an agent must learn how to optimize the amount of reward it can collect during its interaction with its environment. We use them to analyze the task faced by an animal in random and variable schedules of reinforcement. Predictions of the model derived from this analysis are compared to three sets of data obtained in men, rats and pigeons and are contrasted with the ones of its main challenger in psychology, Herrnstein's equation. This reveals the existence of two response strategies in ratio schedules, one which corresponds to our model, the other which is closer to Herrnstein's equation.
Since the end of the XIXth century, the influence of learning on natural selection has been considered. More recently, this influence has been investigated using computer simulations. However, it has not yet been shown how the ability of learning can be the product of natural selection. This point is precisely the subject of this paper.
The law of effect is a very simple law which relates the probability of emission of a behavior by a living being to the consequences of the emission of this behavior by this living being in the past. As such, this law models very basic learning. This law can be considered as an experimental fact as far as it has been observed for a whole range of living beings including human beings. In this paper, we first show that this general law can be the result of a selection process such as natural selection. Then, we show that the implementation of this law can lead to the design of adaptive systems which can mimic very closely the way a new-born develops coordinated movements. To sum-up, we show that the ability to learn such coordinated movements and exhibit adaptive behaviors can result from a multi-stage process of selection.
The law of effect is a very simple law which relates the probability of emission of a behavior by a living being to the consequences of the emission of this behavior by this living being in the past. As such, this law models very basic learning. This law can be considered as an experimental fact as far as it has been observed for a whole range of living beings including human beings. In this paper, we first show that this general law can be the result of a selection process such as natural selection. Then, we show that the implementation of this law can lead to the design of adaptive systems which can mimic very closely the way a new-born develops coordinated movements. To sum-up, we show that the ability to learn such coordinated movements and exhibit adaptive behaviors can result from a multi-stage process of selection.
La modélisation et l'analyse des données dans les programmes temporels opérants tient une place essentielle dans l'analyse du comportement. Les caractéristiques importantes de ces modèles seront présentées en distinguant deux approches majeures. La première se donne pour objectif de modéliser les propriétés molaires des comportements et insiste particulièrement sur les processus cognitifs mis en jeu, La seconde approche modélise les états stables du comportement mais aussi l'acquisition de ceuxci. Ces modèles, dynamiques, insistent davantage sur les propriétés des liens entre réponse et renforcement. L'accent sera particulièrement porté sur l'approche de la variabilité comportementale sous-jacente au modèledynamique nonlinéaireet lamise enperspectivedecette approcheen psychologie.