
Understanding how outcome affective value influences Pavlovian learning is critical for elucidating the mechanisms underlying human emotional learning. In this study, participants completed a Pavlovian conditioning task involving outcomes varying in their level of aversiveness: a tactile (nonaversive) or painful (aversive) shock, and no-shock. Skin conductance responses (SCRs), explicit pleasantness ratings of each CS, and CS-outcome contingency ratings were recorded. Although both painful and tactile outcomes elicited unconditioned SCRs, with greater responses for the former, only the painful outcome triggered anticipatory conditioned SCRs and robust prediction error-related SCRs following its unexpected omission. Computational modeling, through a Rescorla-Wagner model, revealed that prediction errors estimated by the model significantly predicted SCRs following painful outcome omissions, confirming their role as prediction error signals. In addition, the model-derived outcome sensitivity parameter, reflecting the affective value assigned to the different outcomes by each participant, was higher for painful than tactile stimuli, effectively discriminating between them, whereas learning rates did not differ across outcomes. These findings refine our understanding of emotional learning by showing how the affective value of the outcome shapes the predictive mechanisms underlying conditioned responding, as revealed through converging autonomic, explicit, and computational modeling evidence. (PsycInfo Database Record (c) 2026 APA, all rights reserved).
Bastos and Krupenye (2026, Science, 391: 583-586) present an innovative series of studies in which they explore the capacity of a single enculturated bonobo, Kanzi, to represent pretend objects-in other words-"imagination." Their experiments involved pantomimed actions of pouring and emptying juice or placing and dumping out grapes from transparent cups or bowls and asking Kanzi to indicate where the juice or grape would remain, indicating that he was tracking an imagined object, but they failed to account for cuing or alternative explanations. (PsycInfo Database Record (c) 2026 APA, all rights reserved).
Perceptual learning can be defined as a relatively permanent change in discrimination performance as a result of experience or exposure. One key index of perceptual learning is the intermixed-blocked effect in which exposure to two ambiguous or perceptually similar stimuli (i.e., AX and BX) exposed in an intermixed fashion (AX, BX, AX, BX, AX, BX) produces enhanced discrimination performance compared with blocked exposure (AX, AX, AX, BX, BX, BX). Previous imaging data have implicated multiple brain regions in the intermixed-blocked effect. In the present study, transcranial direct current stimulation was used to explore the causal relationship of two regions, the dorsolateral prefrontal cortex (DLPFC) and the posterior parietal cortex (PPC), to performance. A mixed, double-blind, design administered two online sessions of transcranial direct current stimulation (active and sham) to 48 participants, although they viewed pairs of similar stimuli. Half of the participants received active stimulation to the DLPFC, and the other half to the PPC. Anodal stimulation in either the DLPFC or the PPC provided no modulation of discrimination performance relative to sham stimulation. These results potentially question the generality of the interpretation of other studies in which stimulation of these areas does impact on other indices of perceptual learning. (PsycInfo Database Record (c) 2026 APA, all rights reserved).
Rewards strengthen behaviors they follow. This principle, articulated in Thorndike's Law of Effect, is foundational to behavioral science-but it obscures two different ways behaviors may be strengthened. First, and axiomatic, is that the frequency of rewarded behaviors increases. Second, but less well established, is that the variability of rewarded behaviors decreases. Indeed, since Thorndike's pioneering research, investigations into behavioral variability have posed a challenging theoretical puzzle: Some indicate a winnowing of behavioral variants, whereas others suggest an ongoing waxing and waning of even dominant behavioral variants throughout rewarded training. We devised a new experimental task which allowed us to monitor pigeons' sequential pecking of five visually distinctive touchscreen buttons, with all sequences delivering food reward. The 120 possible five-peck sequences we monitored over 250 daily sessions produced a rich set of data to help solve this puzzle. We found that pigeons did decrease the diversity of the sequences they performed. Nevertheless, pigeons continued to perform many different sequences-with even the most dominant sequences frequently rising and falling. These findings and others in diverse realms of behavioral science, neuroscience, and computer science suggest a far-reaching solution to this puzzle: The form and frequency of consistently rewarded behavior involve a dynamic adaptive balance between stability produced by the Law of Effect and variability induced by a persistent exploratory predisposition. Called "the edge of chaos," this balance point preserves responses that reliably secure rewards and engenders behavioral flexibility in the event that the prevailing contingencies change or better options arise. (PsycInfo Database Record (c) 2026 APA, all rights reserved).
Temporal patterns of experiences dictate the depths and rates of learning, forgetting, and extinction but the biological processes involved are fundamentally unknown. Kukushkin et al. (2024) demonstrate that nonneuronal cells in culture can provide a molecular read out of temporal differences in stimulation patterns and therefore provide a tractable model system with which to rigorously investigate how cells time. (PsycInfo Database Record (c) 2026 APA, all rights reserved).
The expression of an association between a conditioned stimulus (CS) and an aversive unconditioned stimulus (US) can be weakened by presenting the CS by itself (extinction [Ext]), pairing it with an appetitive US (counterconditioning [CC]), or pairing it with a neutral stimulus (novelty-facilitated extinction [NFE]). The present research tested whether NFE is less susceptible to ABC renewal than Ext and CC. In two experiments, participants viewed streams of rapid trials. After each stream, participants rated how likely it was that the target CS would be followed by the target US (i.e., predictive learning) as well as the valence of the target CS (i.e., evaluative conditioning). A stream was composed of two phases: Phase 1 established an association between the target CS and target US while Phase 2 aimed at disrupting the expression of this association through Ext, CC, or NFE. Phase 1 occurred in Context A while Phase 2 occurred in Context B. Prediction and valence ratings occurred in either Context A, B, or C. Neither Experiment 1 nor Experiment 2 found differences across interference conditions with predictive testing, regardless of test context. In Experiment 2, better controlled for context effect, CC and NFE altered the CS valence (CC more than NFE) when testing occurred in B, but the difference disappeared when testing occurred in either A or C. The present data do not support the hypothesis that NFE is less susceptible to ABC renewal than either Ext or CC. (PsycInfo Database Record (c) 2026 APA, all rights reserved).
Sequential information processing by animals is a fundamental component of understanding cognition in nonhuman species. Auditory processing is especially important given its implications for acoustic communication, language, and music evolution. In two experiments using an auditory go/no-go procedure, we examined how pigeons processed same/different (S/D) sound sequences. Experiment 1 tested three pigeons with blocked, recurring, or cyclic organizations of two, three, four, or six elements within sequential presentations of 12 1.5-s sounds. Experiment 2 tested four pigeons and examined how midsequence S/D transitions affected ongoing discrimination using contrasting priming strings of varying lengths. Both experiments revealed the birds' ongoing sensitivity to auditory S/D presentations, with the strongest influence exerted by more recently experienced items. The experimental data and subsequent computational modeling suggested pigeons used working memory limited to the last three to four sounds, spanning 4-6 s, to judge the S/D property. These results have important implications for different hypotheses regarding the structure and processing of cross- and multimodal S/D discriminations. (PsycInfo Database Record (c) 2026 APA, all rights reserved).
The peak-interval procedure is one of the most widely used paradigms to assess interval timing in animal models, and disruptions in performance are commonly interpreted as disruptions in the timing mechanism. The time-sharing model (TSM) accounts for such changes by proposing that attentional resources are shared between temporal processing and the processing of novel events, producing pulse loss in the accumulator and leading to rightward shifts in peak time and decreases in response rate. According to the TSM, sustained changes in the temporal cue should induce continuous pulse loss and a marked degradation of temporal control. The present study tested these predictions by exposing pigeons to a peak-interval procedure in which the intensity of the temporal cue was parametrically varied throughout the trial. In Experiment 1, increases in signal intensity preserved response distributions centered near the trained interval, albeit with increased dispersion, whereas decreases in intensity produced irregular distributions, greater divergence from baseline performance, and longer response latencies. Experiment 2 showed that pigeons can acquire and maintain temporal control under low-intensity signals when explicitly trained, while preserving the same asymmetry between intensity increments and decrements. Together, these findings are not fully consistent with the predictions of the TSM and indicate that parametric variations in signal intensity primarily modulate the likelihood and consistency with which timing behavior is expressed. These results highlight the importance of considering nontiming behaviors, motivational factors, and variability when interpreting performance in peak-interval procedures and suggest the consideration of a dissociative modular hypothesis of performance as a complementary framework. (PsycInfo Database Record (c) 2026 APA, all rights reserved).
In an operant feature-positive discrimination, a response during a target stimulus (A) is reinforced when presented with a feature stimulus (X), but not when presented alone (XA+/A-). When the feature and the target occur simultaneously, direct control of the response by X is typically observed, whereas serial pairings produce occasion setting. The present experiment evaluated the effects of the temporal arrangement and spatial stability of the target in a spatial task with pigeons. On separate trials, Features W and X (display color) were simultaneously presented with Landmarks A and B (visual icons) displayed at the same location (out of a row of eight locations) across trials, and pecks at the same response box were reinforced on every trial (goal; simultaneous/static). On other trials, Features Y and Z preceded the onset of Landmarks C and D, which varied in their location across trials, and the reinforced response box was positioned relative to the landmark (serial/dynamic). Transfer tests compared responding to features and landmarks with similar (i.e., W:B, X:A, Y→D, Z→C; Test 1) or different (i.e., W:D, X:C, Y→B, Z→A; Test 2) training histories. Test 1 revealed evidence more consistent with occasion setting during serial/dynamic transfer, whereas evidence more consistent with direct control was observed during simultaneous/static transfer. Test 2 revealed asymmetries in transfer when the features and the targets from different training histories were tested. The results are largely consistent with past research, but testing within a spatial task allowed for an analysis of whether and where responses occurred. (PsycInfo Database Record (c) 2026 APA, all rights reserved).
Goal-driven spatial reorientation in rectangular arenas, with or without landmarks, relies on metric attributes and environmental cues aligned with left-right orientation. Although the phenomenon has been documented in several vertebrates, its mechanisms remain little explored; in reptiles, however, neither the phenomenon nor its mechanisms have been investigated. This study tested whether spatial geometry and exploratory strategies shape reorientation in Hermann's tortoise (Testudo hermanni). Individuals were trained in a rectangular arena to locate a food reward by selecting either two geometrically symmetric corners ("No Blue Wall") or a single corner marked by a blue landmark positioned nearby ("Blue Wall Near") or at a greater distance ("Blue Wall Far"). Tortoises successfully learned to reorient using both geometric and landmark information, adopting consistent strategies to reach the goal. Qualitative observations revealed lateralized wall-following routines and a preference for perimeter exploration. Quantitatively, tortoises in the "Blue Wall Near" condition executed fewer turns, moved faster, and reached the target more quickly than those in the "Blue Wall Far" or "No Blue Wall" conditions, indicating that close proximity to a visual cue enhances spatial efficiency. In contrast, when landmarks were absent, individuals relied more heavily on geometric properties of the arena, following longer and less direct paths. Moreover, the "Blue Wall Near" condition promoted edge-following behavior, suggesting that nearby boundaries further support navigation. These findings demonstrate that reptiles can flexibly integrate geometric and landmark cues during spatial reorientation, underscoring the role of movement in shaping spatial strategies and providing novel insights into the cognitive processes underlying navigation. (PsycInfo Database Record (c) 2026 APA, all rights reserved).
The partial reinforcement extinction effect (PREE) occurs when conditioned responses that were acquired with partial (inconsistent) reinforcement (PRf) extinguish more slowly than responses acquired with continuous (consistent) reinforcement. Sequential theory (Capaldi, 1966, 1994) offers the most popular account of the PREE-that, under PRf, subjects learn that they are reinforced after experiencing sequences of nonreinforcement, and this promotes continued responding during extinction. Recently, Jiao and Harris (2024) showed that rats could learn to anticipate reinforcement after a sequence of nonreinforced trials, but for these rats, being able to predict reinforcement reduced the PREE. However, it is possible that those rats learned to anticipate reinforcement based on time intervals rather than trial sequences. This issue was addressed in the present study. Four groups of rats were trained with both a continuous (consistent) reinforcement and PRf stimulus, and the timing and sequences of reinforcement were systematically manipulated between groups. The results showed that rats used temporal intervals, and not sequences, to predict reinforcement of the PRf stimulus, which calls into question previous findings for sequence learning and directly challenges the main premise of sequential theory. Moreover, during subsequent extinction, a PREE was only observed in groups that had not been able to predict reinforcement. The findings suggest that the PREE depends on uncertainty about reinforcement, and thus increased certainty about reinforcement makes the absence of reinforcement more informative, which promotes faster extinction learning. (PsycInfo Database Record (c) 2026 APA, all rights reserved).
Six experiments with human participants explored the effects of a reinforcer devaluation procedure on the microstructure of free-operant schedule performance, which is theoretically and empirically novel. The goal was to test whether different aspects of the microstructure of free-operant responding would be differentially affected by the manipulation to test the view that bout-initiation responses can be regarded as stimulus-driven habits both susceptible to devaluation, and within-bout responses as goal-directed actions that are susceptible to such effects. Humans responded to a fictitious investment game in which responses made investments in two countries with different currencies. These currencies had an exchange rate into Great British Pound (GBP), which could be manipulated to devalue the currency (reinforcer). In Experiments 1a and 2, random ratio schedules showed overall devaluation effects, which were limited to within-bout rates and more pronounced on longer ratios. In Experiments 1b and 3, random interval schedules showed limited devaluation effects, overall, except for with shorter intervals. This effect was seen on some schedules for within-bout responding. Experiments 4 and 5 manipulated the response-reinforcer feedback function (Experiment 4) and participants' ability to experience contingency variations by pretraining to be variable or static in response rates (Experiment 5). Only when there was a stronger response-reinforcer relationship (Experiment 4), which could be experienced due to variable responding (Experiment 5), did the devaluation effect occur, and these effects were limited to within-bout responses. These data suggest that an account of free-operant responding based on there being, at least, two classes of response-bout initiations that are stimulus-driven habits, and within-bout responses that are goal-directed actions-may explain many schedule effects. (PsycInfo Database Record (c) 2026 APA, all rights reserved).
Two experiments examined the hypothesis that reinstatement, the recovery of an extinguished response to a cue with the introduction of the unsignaled outcome used in conditioning, can result from context-outcome associations changing between extinction and testing resulting in a renewal of responding to the extinguished cue. The experiments manipulated whether context-outcome associations at testing were the same as those in conditioning or extinction. Using a space-shooter video game, participants learned to press a key in response to a spaceship (O1) which was signaled by a sensor cue (R) in conditioning and then received extinction where R appeared alone. In both phases, context-outcome associations were manipulated by unsignaled presentations of another spaceship (O2, associated with a different key), or its absence. In a third phase, context-outcome associations in the absence of R and O1 were manipulated to be the same as those in conditioning or extinction. When O2 was present during conditioning and immediately before test, but absent in extinction, a strong recovery of extinguished R-O1 responding was observed. When O2 was absent during conditioning and immediately before testing, but present in extinction, a smaller recovery of R-O1 was observed. When a new spaceship (O3) was introduced immediately before testing, a recovery of R-O1 was observed only if O2 had been present in conditioning, suggesting that the effects observed were not simply disinhibition of responding. (PsycInfo Database Record (c) 2026 APA, all rights reserved).
Value-modulated attentional capture (VMAC) refers to the tendency for highly valued cues to capture attention even when counterproductive to an individual's goals. VMAC is thought to measure incentive salience and has been shown to correlate with multiple measures of compulsivity. The current study demonstrates the first successful back-translation of VMAC to rodents. To achieve this, mice (Experiment 1) or rats (Experiment 2) were trained to sign-track to a high-value lever signaling three pellets and a low-value lever signaling one pellet, which would later serve as distractors. Rodents were next trained to nose-poke an illuminated port for a pellet, which would serve as the "target" response. The key VMAC test phase compared nose-poke performance in the presence of high- versus low-value lever distractors. Rodents made significantly more omissions in the presence of the high- than the low-value lever distractor despite losing three times as many pellets on these trials. This replicates findings from human VMAC tasks, in which participants are consistently impaired on high-value relative to low-value distractor trials despite greater reward loss. In Experiment 2, we showed that this effect persisted despite devaluation of the pellet outcome by conditioned taste aversion, even when the disliked outcome was presented, suggesting a compulsion-like mechanism. Together, these data show that VMAC can be observed in both mice and rats, which opens new avenues for investigation of its behavioral and neural underpinnings, with implications for understanding and treating compulsive disorders. (PsycInfo Database Record (c) 2026 APA, all rights reserved).
In the ephemeral reward task, animals are presented with two choice alternatives, one optimal, the other suboptimal. Choosing the suboptimal alternative delivers one immediate reward and ends the trial, whereas choosing the optimal alternative also yields one immediate reward but allows subsequent access to the reward associated with the suboptimal alternative. While species such as cleaner wrasse and grey parrots excel at this task, others-including pigeons, primates, and rats-struggle, raising questions about the factors influencing success. This study investigated these factors by examining performance in starlings under standard and modified task conditions. Across two experiments, starlings successfully learned to prefer the optimal option. In these experiments, we occasionally included single-option trials, which allowed birds to experience the outcomes of each choice in isolation. In Experiment 2, we also manipulated the delay between the two sequential rewards to test its effect on performance. Preference for the optimal option declined as the delay increased, suggesting that shorter delays facilitate credit assignment to the initial choice. We hypothesize that shorter delays facilitate the association between initial choices and subsequent rewards and that differences in apparatus, intertrial intervals, and the rate of memory decay may also influence performance in the task. Overall, our results highlight the complexity of the ephemeral reward task and suggest the potential interplay of ecological relevance and task design. (PsycInfo Database Record (c) 2025 APA, all rights reserved).
Food-caching chickadees are known to cache thousands of food items and retrieve these caches using, at least in part, spatial memory. New research shows memory recall is associated with remote activation of hippocampal place cells by gaze using two peaks in neuronal firing: an early peak predicts the gaze direction and a later peak reflects the gaze. (PsycInfo Database Record (c) 2025 APA, all rights reserved).
Habituation and dishabituation are fundamental adaptive processes that govern how animals respond to repeated stimuli. Habituation is defined as a decline in response to irrelevant stimuli, and dishabituation reactivates this response upon qualitatively different stimulation. Here, we explored these processes in bumblebees (Bombus terrestris) by exposing freely foraging individuals to a repeated overhead looming stimulus, followed by a distinct vibration. We identified three defensive responses-flight, disturbance leg-lift response, and startle-and found that only flight probability showed robust habituation and dishabituation. Disturbance leg-lift response remained consistently frequent, whereas startle initially increased and later declined when flight was reinstated. Our findings demonstrate clear habituation and dishabituation of defensive responses in bumblebees within a novel free-flying testing paradigm, providing initial support for response-specific plasticity mechanisms. The results underscore the importance of differentiating among multiple defensive responses to better understand the mechanisms driving habituation and dishabituation, suggesting that bumblebee defense strategies are finely tuned across multiple stimulus-response pathways. (PsycInfo Database Record (c) 2025 APA, all rights reserved).
Townrow and Krupenye (2025) show that bonobos will point more in a cooperative task when their partner is ignorant of the location of the desired food. While their study convincingly shows that bonobos can track ignorance, one can question whether it provides evidence that they can represent it as such. (PsycInfo Database Record (c) 2025 APA, all rights reserved).
Valid predictors of an outcome attract more attention than stimuli that are nonpredictive. Furthermore, stimuli that have a probabilistic association with an outcome attract more attention than stimuli that have a deterministic association with an outcome. Two experiments investigated whether predictive validity and outcome uncertainty resulted in the establishment of a more accurate stimulus representation, in which accuracy was measured as the strength of associations between different elements of a compound stimulus. In Experiment 1, pairs of stimuli were established as outcome predictive (always followed by the same outcome) and presented in conjunction with nonpredictive pairs of stimuli (equally likely to be followed by two different outcomes). Outcome uncertainty was also manipulated, between groups, by establishing either a deterministic (100%) or probabilistic (80%) contingency between the predictive pairs and their outcomes. The test trials revealed more accurate recognition for which predictive stimuli were paired together relative to nonpredictive stimuli; however, there was no effect of outcome uncertainty. Experiment 2 reproduced the effect observed in the deterministic group from Experiment 1 and also demonstrated that the superior performance to the predictive stimuli over the nonpredictive stimuli was only evident when, at test, the choice stimuli had predicted different outcomes during training. These results were interpreted as the consequence of two pathways to accurate stimulus representation: direct (within-compound associations) and indirect (mediated through the activation of the outcome) and are discussed in the context of attentional theories of associative learning. (PsycInfo Database Record (c) 2025 APA, all rights reserved).
In the Ephemeral Reward Task, a subject is presented with a choice between two stimuli, A and B. If it chooses A, it gets a reward and the trial is over. If it chooses B, it gets a reward and it can then respond to A, to obtain a second reward. Wrasse (cleaner fish) and parrots learn to choose B optimally within 100 trials, primates may also learn, whereas pigeons and rats do not. We attempted to determine why pigeons have difficulty learning their task. First, we tested the hypothesis that pigeons fail because the outcome after choice of A is similar to the outcome after a response to A given choice of B. For group AC, after the choice of B, stimulus A changed to stimulus C. For group BC, after the choice of stimulus B, stimulus B changed to stimulus C. For group BB, after the choice of stimulus B, stimulus B remained for a second reward. None of the three groups learned to choose optimally. In Experiment 2, the probability of reward for choice of stimulus A or B was reduced to 50%. Pigeons learned to choose optimally. We suggest that the difference in value between one and two rewards may not be as great as the difference in value between 0.5 and one reward. (PsycInfo Database Record (c) 2025 APA, all rights reserved).