Abstract The midsession reversal (MSR) task is frequently used to study behavioral flexibility and decision strategies in animals. In a typical version of the task, subjects complete 80 trials in which they choose between two simultaneously presented stimuli, S1 and S2. During the first 40 trials, responses to S1 are reinforced, whereas responses to S2 are not. The contingencies then reverse without warning: From trial 41 to 80, only responses to S2 are reinforced. In birds, performance in this task is often characterized by anticipatory and perseverative errors around the reversal point, suggesting a reliance on elapsed time since the session began. In contrast, rats tested in operant conditioning chambers typically show near-optimal performance with few errors, a pattern often interpreted as evidence that rats rely primarily on local reinforcement cues rather than temporal information. The present study investigated whether rats exclusively rely on local cues in the MSR task or whether temporal information also contributes to the decision process. Two groups of rats were trained with different intertrial intervals (ITIs; 5 s or 10 s) while the reversal point remained fixed at Trial 41. During acquisition, both groups diplayed similar learning rates and near-optimal steady-state performance with minimal anticipatory or perseverative errors. However, when the ITI was manipulated in probe sessions, systematic shifts in switching behavior emerged. Rats adjusted their choices according to the temporal midpoint experienced during training rather than the nominal trial number of the reversal. These results suggest that rats rely on a mixed strategy that integrates local reinforcement cues with global timing information. Temporal control may therefore be present even when it is not expressed during standard training conditions.
To assess the degree of temporal control in the midsession reversal task, pigeons learned a simultaneous discrimination with two stimuli, S1 and S2. Choices of S1 were reinforced on the first 40 trials and choices of S2 on the last 40. Variable intertrial intervals (ITIs) separated the trials. The pigeons learned to reverse preference from S1 to S2 near trial 40. To see if pigeons had learned a temporal discrimination, we either doubled or halved the mean of the ITIs during test sessions. If choices were based on temporal cues, preference should reverse at the same time in the session but on half as many trials into the session when the mean ITI doubled and twice as many trials later when the mean ITI halved. Variable ITIs should reduce generalization decrement and thereby reveal temporal control more clearly. The reversal trial changed in the directions predicted by timing, but the magnitude of the change was smaller than predicted. Most individual choice patterns were consistent with temporal control, but a few were consistent with control by trial number or by the events of the previous trial. Variable ITIs seem to reduce the weight of temporal cues relative to the weight of nontemporal cues.
Cognitive (or behavioral) flexibility is considered an executive function characterized by patterns of behavioral adjustment in response to changes in environmental demands, which tends to decline with aging. The simple discrimination reversal task is a useful way to evaluate this function, as it directly measures processes related to performance change, such as sensitivity to consequences, learning set formation, and concept formation. Few studies on aging have employed this task, and those that have did not examine its component processes or include middle-aged adults. This study aimed to evaluate cognitive flexibility and its component processes through a simple discrimination reversal task, applied to 100 participants divided into four age groups: emerging adults, younger adults, middle-aged adults, and older adults. After learning three simple simultaneous visual discriminations, the function of the positive and negative stimuli was reversed three times, with participants needing to meet a performance criterion each time. Older participants were more likely to fail to meet the performance criterion in some of the reversals, a pattern consistent with reduced sensitivity to consequences and failure in class formation. Moreover, older individuals who succeeded in the task learned the new function assigned to stimuli more slowly during reversals and were less likely to form classes in the second reversal. However, all participants who met the criterion across the three reversals showed evidence of learning-set formation, regardless of age.
In the ephemeral reward task, animals are presented with two choice alternatives, one optimal, the other suboptimal. Choosing the suboptimal alternative delivers one immediate reward and ends the trial, whereas choosing the optimal alternative also yields one immediate reward but allows subsequent access to the reward associated with the suboptimal alternative. While species such as cleaner wrasse and grey parrots excel at this task, others-including pigeons, primates, and rats-struggle, raising questions about the factors influencing success. This study investigated these factors by examining performance in starlings under standard and modified task conditions. Across two experiments, starlings successfully learned to prefer the optimal option. In these experiments, we occasionally included single-option trials, which allowed birds to experience the outcomes of each choice in isolation. In Experiment 2, we also manipulated the delay between the two sequential rewards to test its effect on performance. Preference for the optimal option declined as the delay increased, suggesting that shorter delays facilitate credit assignment to the initial choice. We hypothesize that shorter delays facilitate the association between initial choices and subsequent rewards and that differences in apparatus, intertrial intervals, and the rate of memory decay may also influence performance in the task. Overall, our results highlight the complexity of the ephemeral reward task and suggest the potential interplay of ecological relevance and task design. (PsycInfo Database Record (c) 2025 APA, all rights reserved).
This study investigates the mechanisms that underlie pigeons' performance in the number‐left task. After producing x light flashes, pigeons had to choose between a standard option that delivered reinforcement after a fixed number of additional flashes, S = 4, and a number‐left option that delivered reinforcement after a variable number of additional flashes, L = 8 − x . In Experiment 1, pigeons were trained with forced and choice trials with 1 ≤ x ≤ 7. During testing, the number of choice trials was simply increased. In Experiment 2, pigeons were trained only with the anchor numerosities x = 1 and x = 7 and during testing unreinforced probe trials introduced the intermediate numerosities, x = 2, 3, 4, 5, and 6. Performance was similar in both experiments and consistent with a computational mechanism. To test whether performance in the previous experiments was due to the substantial overlap in the induced generalization gradients around the anchor numerosities, in Experiments 3a and 3b, we selected anchor numerosities that were farther apart ( x = 5 and x = 50, with S = 12 and L = 53 − x ). Yet, pigeons' performance remained similar. We discuss the implications of these findings for the mechanisms that underlie numerosity discrimination.
Given a choice between a simple option offering a preferred-food item (e.g., a grape, G) and a combo option offering the same preferred-food item plus a less-preferred food item (e.g., a grape + a slice of cucumber, GC), animals often behave suboptimally by either being indifferent between the two options or by preferring the simple option—the “less-is-better” effect. To explain indifference, the selective-value hypothesis assumes that, in the choice context, the subjective value (V) of the less-preferred food is zero (i.e., VGC = VG + 0 = VG). To explain the less-is-better effect, the average-quality hypothesis assumes that the value of the combo equals the average of its components’ values [i.e., VGC = Average(VG,VC) < VG]. No unified account explains both sets of experimental findings. To test these hypotheses further, we presented six capuchin monkeys (Sapajus sp.) with a variety of choice tests, some simple (GC vs. G) and some complex (2G1C vs. 3G1C; or 2C1G vs. 3C1G), with a final sample of four individuals per test. The results confirm that capuchin monkeys also behave suboptimally, revealing either indifference or the less-is-better effect. Crucially, our findings suggest that the value of the less-preferred food, C, may become negative through a contrast-like effect. By expanding the selective-value hypothesis to accommodate situations where the less-preferred item’s value may be reduced to zero or become negative, we suggest a unified, process-based account of the two sets of research findings.
This study examined how starlings (Sturnus unicolor) adapt to a serial learning task with a predictable reversal in the reinforcement contingencies at midsession. The birds learned a simultaneous discrimination between two options, S1 and S2 (red and green key light colors). Choices of S1 were rewarded during the first 40 trials and choices of S2 were rewarded during the last 40 trials, with variable exponentially distributed ITIs separating the trials. Then, to test the hypothesis that starlings anticipate the midsession reversal based on time into the session, we changed the average of the ITIs during a test session. The hypothesis predicted that with ITIs twice as short during testing, preference would shift from S1 to S2 twice as many trials later than in training, and with ITIs twice as long during testing, preference would shift twice as many trials earlier than in training. Results showed that preference shifted in the predicted direction, but the shifts were smaller in magnitude than predicted. Cumulative difference records plotting choices across time- or trial-into-the-session revealed a variety of adjusting strategies, some consistent with the use of temporal cues, others consistent with the use of local or numerical cues. The variability of strategies occurred both between and within subjects and suggests that multiple cues combine to control behavior in the midsession reversal task.
In a variety of laboratory preparations, several animal species prefer signaled over unsignaled outcomes. Here we examine whether pigeons prefer options that signal the delay to reward over options that do not and how this preference changes with the ratio of the delays. We offered pigeons repeated choices between two alternatives leading to a short or a long delay to reward. For one alternative ( informative ), the short and long delays were reliably signaled by different stimuli (e.g., S S for short delays, S L for long delays). For the other ( non-informative ), the delays were not reliably signaled by the stimuli presented ( S 1 and S 2 ). Across conditions, we varied the durations of the short and long delays, hence their ratio, while keeping the average delay to reward constant. Pigeons preferred the informative over the non-informative option and this preference became stronger as the ratio of the long to the short delay increased. A modified version of the Δ – Σ hypothesis (González et al., J Exp Anal Behav 113(3):591–608. https://doi.org/10.1002/jeab.595 , 2020a) incorporating a contrast-like process between the immediacies to reward signaled by each stimulus accounted well for our findings. Functionally, we argue that a preference for signaled delays hinges on the potential instrumental advantage typically conveyed by information.
Under certain conditions, pigeons prefer information about whether food will be forthcoming at the end of an interval to a higher chance of obtaining the food. In the typical protocol, choosing one option (Informative) is followed by one of two 10-s long terminal-link stimuli: S-G always ending in food or S (R) never ending in food, with S-G occurring only 20% of the trials. The other option (Non-informative) is also followed by one of two 10-s long terminal-link stimuli: S-B or S-Y, both ending in food 50% of the trials. Although the Informative option yields food with a lower probability than the Non-informative (0.2 vs. 0.5), pigeons prefer it. To determine whether such preference occurs because S-G and S (R) disambiguate the trial outcome immediately upon choice, we delayed the moment the disambiguation took place in two experiments. In Experiment 1, when the Informative option was chosen, S-G always ensued for t seconds of the terminal-link, and then the standard contingencies followed. Experiment 2 was similar, except that S (R) always ensued for t seconds. Across conditions, t varied from 0 to 10 s. In both experiments, preference for the Informative option decreased with t, but the effect was stronger in Experiment 1. We discuss the implication of these findings for functional and mechanistic models of suboptimal choice.
In the Mid-Session Reversal task (MSR), an animal chooses between two options, S1 and S2. Rewards follow S1 but not S2 from trials 1-40, and S2 but not S1 from trials 41-80. With pigeons, the psychometric function relating S1 choice proportion to trial number starts close to 1 and ends close to 0, with indifference (PSE) close to trial 40. Surprisingly, pigeons make anticipatory errors, choosing S2 before trial 41, and perseverative errors, choosing S1 after trial 40. These errors suggest that they use time into the session as the preference reversal cue. We tested this timing hypothesis with 10 Spotless starlings. After learning the MSR task with a T-s Inter-Trial Interval (ITI), they were exposed to either 2 T or T/2 ITIs during testing. Doubling the ITI should shift the psychometric function to the left and halve its PSE, whereas halving the ITI should shift the function to the right and double its PSE. When the starlings received one pellet per reward, the ITI manipulation was effective: The psychometric functions shifted in the direction and by the amount predicted by the timing hypothesis. However, non-temporal cues also influenced choice.
The midsession reversal task involves a simultaneous discrimination between stimuli S1 and S2. Choice of S1 but not S2 is reinforced during the first 40 trials, and choice of S2 but not S1 is reinforced during the last 40 trials. Trials are separated by a constant intertrial interval (ITI). Pigeons learn the task seemingly by timing the moment of the reversal trial. Hence, most of their errors occur around trial 40 (S2 choices before trial 41 and S1 choices after trial 40). It has been found that when the ITI is doubled on a test session, the reversal trial is halved, a result consistent with timing. However, inconsistent with timing, halving the ITI on a test session did not double the reversal trial. The asymmetry of ITI effects could be due to the intrusion of novel cues during testing, cues that preempt the timing cue. To test this hypothesis, we ran two types of tests after the regular training in the midsession reversal task, one with S1 and S2 choices always reinforced, and another with S1 always reinforced but S2 reinforced only after 20 trials when the ITI doubled or 40 trials when the ITI halved. For most pigeons, performance was consistent with timing both when the ITI doubled and when it was halved, but some pigeons appeared to follow strategies based on counting or on reinforcement contingencies.
To study how multiple stimuli may control discriminative behavior, we exposed fifteen pigeons to a symbolic matching-to-sample task with three samples that differed only in duration (2, 6, and 18s) and two keylight colors as comparisons. The pigeons learned to choose one comparison after the shortest sample, and the other comparison after the intermediate and longest samples. A 30-s intertrial interval (ITI), illuminated with the houselight, separated the trials. Previous data has suggested that, in this arrangement, both sample keylight and the ITI houselight influence choice. To assess this joint stimulus control, we introduced two tests. In the no-sample test, the keylight was not illuminated and the comparisons followed the ITI immediately; in the dark-ITI test, the houselight was not illuminated. Results confirmed that both stimuli influenced choice, with an apparent trade-off between them: The more a pigeon relied on one stimulus, the less it seemed to rely on the other. We discuss potential models of joint stimulus in temporal discrimination tasks.
In the study of animal timing over the last 100 years, we identify three different periods, each characterized by a distinct activity. In the first period, researchers brought timing into the laboratory and explored its multiple expressions empirically. In the second period, the growing body of empirical findings inspired researchers to develop a plethora of timing models that vary in theoretical orientation, scope, depth, and quantitative explicitness. We argue that it is now the time to advance towards a third period, wherein researchers select models by comparing them with one another and with data. We make our case by contrasting how the scalar expectancy theory and the learning-to-time model conceive of temporal memory and learning both in concurrent timing tasks and in retrospective timing tasks. We identify four problems related to the structure of temporal memory and to the rules of temporal learning that challenge these models and that should drive the next steps in modeling the timing abilities of animals. (PsycInfo Database Record (c) 2022 APA, all rights reserved).
A Correction to this paper has been published:
Journal of the Experimental Analysis of BehaviorVolume 115, Issue 2 p. 596-603 Theoretical Article Dissolving the molar–molecular controversy Armando Machado, Corresponding Author Armando Machado [email protected] orcid.org/0000-0003-4380-3711 University of Aveiro, Portugal Address correspondence to: Armando Machado, Email: [email protected]Search for more papers by this authorMarco Vasconcelos, Marco Vasconcelos University of Aveiro, PortugalSearch for more papers by this author Armando Machado, Corresponding Author Armando Machado [email protected] orcid.org/0000-0003-4380-3711 University of Aveiro, Portugal Address correspondence to: Armando Machado, Email: [email protected]Search for more papers by this authorMarco Vasconcelos, Marco Vasconcelos University of Aveiro, PortugalSearch for more papers by this author First published: 26 January 2021 https://doi.org/10.1002/jeab.675 AM and MV were supported by the Portuguese Foundation for Science and Technology (UIDB/04810/2020) Read the full textAboutPDF ToolsRequest permissionExport citationAdd to favoritesTrack citation ShareShare Give accessShare full text accessShare full-text accessPlease review our Terms and Conditions of Use and check box below to share full-text version of article.I have read and accept the Wiley Online Library Terms and Conditions of UseShareable LinkUse the link below to share a full-text version of this article with your friends and colleagues. Learn more.Copy URL Share a linkShare onEmailFacebookTwitterLinkedInRedditWechat References Bachelard, G. (1984). The new scientific spirit. Beacon Press. (Originally published in 1934.) Google Scholar Baer, D. M. (1981). The imposition of structure on behavior and the demolition of behavioral structures. In H. E. Howe (Ed.), Nebraska symposium on motivation, 2 (pp. 217–254). University of Nebraska Press. Google Scholar Blough, D. S. (1966). The reinforcement of least-frequent interresponse times. Journal of the Experimental Analysis of Behavior, 9(5), 581-591. https://doi.org/10.1901/jeab.1966.9-581 10.1901/jeab.1966.9-581 CASPubMedWeb of Science®Google Scholar Bryant, D., & Church, R. (1974). The determinants of random choice. Animal Learning & Behavior, 2, 245-248. https://doi.org/10.3758/BF03199188 10.3758/BF03199188 Web of Science®Google Scholar Donahoe, J., Palmer, D., & Burgos, J. (1997). The unit of selection: What do reinforcers reinforce? Journal of the Experimental Analysis of Behavior, 67(2), 259–273. https://doi.org/10.1901/jeab.1997.67-259 10.1901/jeab.1997.67-259 CASPubMedWeb of Science®Google Scholar Donahoe, J. W. (2012). Origins of the molar–molecular debate. European Journal of Behavior Analysis, 13(2), 195-200. https://doi.org/10.1080/15021149.2012.11434422 10.1080/15021149.2012.11434422 Google Scholar Herrnstein, R. J. (1979). Derivatives of matching. Psychological Review, 86(5), 486-495. https://doi.org/10.1037/0033-295X.86.5.486 10.1037/0033-295X.86.5.486 Web of Science®Google Scholar Heyman, G. M. (1979). A Markov model description of changeover probabilities on concurrent variable interval schedules. Journal of the Experimental Analysis of Behavior, 31(1), 41-51. https://doi.org/10.1901/jeab.1979.31-41 10.1901/jeab.1979.31-41 CASPubMedWeb of Science®Google Scholar Hinson, J. M., & Staddon, J. E. R. (1983). Hill-climbing in pigeons. Journal of the Experimental Analysis of Behavior, 39(1), 25-47. https://doi.org/10.1901/jeab.1983.39-25 10.1901/jeab.1983.39-25 CASPubMedWeb of Science®Google Scholar Hiraoka, K. (1984). Discrete-trial probability learning in rats: Effects of local contingencies of reinforcement. Animal Learning and Behavior, 12, 343-349. https://doi.org/10.3758/BF03199978 10.3758/BF03199978 Web of Science®Google Scholar Krebs, J. R., Kacelnik, A., & Taylor, P. (1978). Test of Optimal Sampling by Foraging Great Tits. Nature, 275(5675), 27-31. https://doi.org/10.1038/275027a0 10.1038/275027a0 Web of Science®Google Scholar Machado, A. (1989). Operant conditioning of behavioral variability using a percentile reinforcement schedule. Journal of the Experimental Analysis of Behavior, 52(2), 155-166. https://doi.org/10.1901/jeab.1989.52-155 10.1901/jeab.1989.52-155 CASPubMedWeb of Science®Google Scholar Machado, A. (1992). Behavioral variability and frequency-dependent selection. Journal of the Experimental Analysis of Behavior, 58(2), 241-263. https://doi.org/10.1901/jeab.1992.58-241 10.1901/jeab.1992.58-241 CASPubMedWeb of Science®Google Scholar Machado, A. (1993). Learning variable and stereotypical sequences of response: some data and a new model. Behavioral Processes, 30(2), 103-130. https://doi.org/10.1016/0376-6357(93)90002-9 10.1016/0376-6357(93)90002-9 CASPubMedWeb of Science®Google Scholar Machado, A., & Keen, R. (1999). The learning of response patterns in choice situations. Animal Learning & behavior, 27, 251-271. https://doi.org/10.3758/BF03199724 10.3758/BF03199724 Web of Science®Google Scholar Machado, A., Keen, R., & Macaux, E., (2008). Making analogies work: a selectionist model of choice. In Nancy K. Innis (Ed.), Reflections on adaptive behavior: Essays in honor of J. E. R. Staddon (pp. 23-50). The MIT Press. Google Scholar Machado, A., & Tonneau, F. (2012) Operant Variability: Procedures and Processes. The Behavior Analyst, 35(2), 249–255. https://doi.org/10.1007/BF03392284 10.1007/BF03392284 PubMedWeb of Science®Google Scholar Nergaard, S. K., & Holth, P. (2020). A critical review of the support for variability as an operant dimension. Perspectives on Behavior Science, 43(3), 579–603. https://doi.org/10.1007/s40614-020-00262-y 10.1007/s40614-020-00262-y PubMedWeb of Science®Google Scholar Neuringer, A., Jensen, G. (2012). The predictably unpredictable operant. Comparative Cognition & Behavior Reviews, 7, 55-84. https://doi.org/10.3819/ccbr.2012.7004 10.3819/ccbr.2012.70004 Google Scholar Nevin, J. A. (1969). Interval reinforcement of choice behavior in discrete trials. Journal of the Experimental Analysis of Behavior, 12(6), 875-885. https://doi.org/10.1901/jeab.1969.12-875 10.1901/jeab.1969.12-875 CASPubMedWeb of Science®Google Scholar Nevin, J. A. (1979). Overall matching versus momentary maximizing: Nevin (1969) revisited. Journal of Experimental Psychology: Animal Behavior Processes, 5(3). 300-306. https://doi.org/10.1037/0097-7403.5.3.300 10.1037/0097-7403.5.3.300 Web of Science®Google Scholar Page, S., & Neuringer, A. (1985). Variability is an operant. Journal of Experimental Psychology: Animal Behavior Processes, 11(3), 429-452. https://doi.org/10.1037/0097-7403.11.3.429 10.1037/0097-7403.11.3.429 Web of Science®Google Scholar Shimp, C. P. (1966). Probabilistically reinforced choice behavior in pigeons. Journal of the Experimental Analysis of Behavior, 9(4), 443-455. https://doi.org/10.1901/jeab.1966.9-443 10.1901/jeab.1966.9-443 CASPubMedWeb of Science®Google Scholar Shimp, C. P. (1967). Reinforcement of least-frequent sequences of choices. Journal of the Experimental Analysis of Behavior, 10(1), 57-65. https://doi.org/10.1901/jeab.1967.10-57 10.1901/jeab.1967.10-57 CASPubMedWeb of Science®Google Scholar Shimp, C. P. (1969). Optimal behavior in free-operant experiments. Psychological Review, 76(2), 97-112. https://doi.org/10.1037/h0027311 10.1037/h0027311 Web of Science®Google Scholar Shimp, C. P. (1976). Short-term memory in the pigeon: The previously reinforced response. Journal of the Experimental Analysis of Behavior, 26(3), 487-493. https://doi.org/10.1901/jeab.1976.26-487 10.1901/jeab.1976.26-487 CASPubMedWeb of Science®Google Scholar Shimp, C. P. (1982). Choice and behavioral patterning. Journal of the Experimental Analysis of Behavior, 37, 157–169. https://doi.org/10.1901/jeab.1982.37-157. 10.1901/jeab.1982.37-157 CASPubMedWeb of Science®Google Scholar Shimp, C. P. (2020). Molecular (moment-to-moment) and molar (aggregate) analyses of behavior. Journal of the Experimental Analysis of Behavior, 114(3), 394-429. https://doi.org/10.1002/jeab.626 10.1002/jeab.626 PubMedWeb of Science®Google Scholar Silberberg, A., Hamilton, B., Ziriax, J. M., & Casey, J. (1978). The structure of choice. Journal of Experimental Psychology: Animal Behavior Processes, 4(4), 292-301. https://doi.org/10.1037/0097-7403.4.4.368 10.1037/0097-7403.4.4.368 Web of Science®Google Scholar Silberberg, A. & Williams, D. (1974). Choice behavior in discrete trials: A demonstration of the occurrence of a response strategy. Journal of the Experimental Analysis of Behavior, 21(2), 315-322. https://doi.org/10.1901/jeab.1974.21-315 10.1901/jeab.1974.21-315 CASPubMedWeb of Science®Google Scholar Silberberg, A., & Ziriax, J. M. (1982). The interchangeover time as a molecular dependent variable in concurrent schedules. In M. L. Commons, R. J. Herrnstein, & H. Rachlin (Eds.), Quantitative Analyses of Behavior, Vol. II: Matching and maximizing accounts (pp. 131-151). Ballinger. Google Scholar Silberberg, A. & Ziriax, J. M. (1985). Molecular maximizing characterizes choice on Vaughan's (1981) procedure. Journal of the Experimental Analysis of Behavior, 43(1), 83-96. https://doi.org/10.1901/jeab.1985.43-83 10.1901/jeab.1985.43-83 CASPubMedWeb of Science®Google Scholar Thomas, G., Kacelnik, A., & Van der Meulen, J. (1985). The three-spined stickleback and the two-armed bandit. Behaviour, 93(1-4), 227–240. https://doi.org/10.1163/156853986X00900 10.1163/156853986X00900 Web of Science®Google Scholar Timberlake, W. (1999). Biological behaviorism. In W. O'Donohue, R. Kitchener (Eds.), Handbook of behaviorism (p. 243–284). Academic Press. 10.1016/B978-012524190-8/50011-5 Google Scholar Vasconcelos, M., Fortes, I., & Kacelnik, A. (2017). On the Structure and Role of Optimality Models in the Study of Behavior. In J. Call (Ed.), APA Handbook of Comparative Psychology (Vol. 2, pp. 287-307). American Psychological Association. 10.1037/0000012-014 Google Scholar Williams, B. A. (1972). Probability learning as a function of momentarily reinforcement probability. Journal of the Experimental Analysis of Behavior, 17(3), 363-368. https://doi.org/10.1901/jeab.1972.17-363 10.1901/jeab.1972.17-363 CASPubMedWeb of Science®Google Scholar Williams, B. A. (1983). Effects of intertrial interval on momentary maximizing. Behaviour Analysis Letters, 3, 35-42. Web of Science®Google Scholar Williams, B. A. (1988). Reinforcement, choice, and response strength. In R. C. Atkinson, R. J. Herrnstein, G. Lindzey, & R. D. Luce (Eds.), Steven's handbook of experimental psychology. ( 2nd Ed., Vol. 2, pp. 167-244). New York: Wiley. Google Scholar Williams, B. A. (1991). Choice as a function of local versus molar reinforcement contingencies. Journal of the Experimental Analysis of Behavior, 56(3), 455-473. https://doi.org/10.1901/jeab.1991.56-455 10.1901/jeab.1991.56-455 CASPubMedWeb of Science®Google Scholar Zeiler, M. D. (1987). On Optimal Choice Strategies. Journal of Experimental Psychology: Animal Behavior Processes, 13(1), 31-39. https://doi.org/10.1037/0097-7403.13.1.31 10.1037/0097-7403.13.1.31 Web of Science®Google Scholar Volume115, Issue2March 2021Pages 596-603 ReferencesRelatedInformation
We used a midsession reversal task to investigate how temporal and situational cues may combine to determine choice in frequently changing environments. Pigeons learned a simultaneous discrimination with 2 stimuli: S1 and S2. Choices of S1 were reinforced only during the first trials, and choices of S2 were reinforced only during the last trials of the session, that is, the reinforcement contingencies reversed once during the session. To weaken the temporal cue (time into the session) that signaled the reversal trial, we varied the location of reversal trial randomly across sessions; to weaken the situational cue (the outcome of the previous trials that might support a win-stay/loose shift strategy), we varied the payoff probabilities associated with S1 and S2. Performance was consistent neither with the exclusive use of a timing strategy, nor with the exclusive use of a situational, win-stay/lose-shift strategy. Instead, choice seemed to be under joint control of both cues. The relative influence of these cues was dynamic: When payoff was higher for S1 than S2, behavior was less time-controlled than when the payoff was higher for S2 that S1, or when they were equal. We advance a descriptive mixture model of joint control for the midsession reversal task. (PsycInfo Database Record (c) 2021 APA, all rights reserved).
We investigated how base rates affect temporal discrimination. In a temporal bisection task, pigeons learned to choose one key after a short sample and another key after a long sample. When presented with a range of intermediate samples they produced a psychometric function characterized by a bias and a scale parameter. When one of the trained samples was more frequent than the other, only the location parameter changed, with the pigeons biasing their choices toward the key associated with the most frequent sample. We then reproduced the bisection task in a long operant chamber, with choice keys far apart, and tracked the pigeons' motion patterns during the sample. Pigeons learned to approach the short key following sample onset, wait on the "short side" for a few seconds, and then, when the sample continued, depart toward the long key. This time-place curve was affected by sample base rate: The probability of pigeons going directly to the long side after sample onset increased when long samples were more frequent than short samples, indicating a decrease of temporal control. We found no evidence of changes in temporal sensitivity. The results are most consistent with models of timing that take into account biasing effects and competition for stimulus control. (PsycInfo Database Record (c) 2021 APA, all rights reserved).
Eckard and Lattal (2020) summarized the behavioristic view of hypothetical constructs and theories, and then, in a novel and timely manner, applied this view to a critique of internal clock models of temporal control. In our three-part commentary, we aim to contribute to the authors' discussion by first expanding upon their view of the positive contributions afforded by constructs and theories. We then refine and question their view of the perils of reifying constructs and assigning them causal properties. Finally, we suggest to behavior analysts four rules of conduct for dealing with mediational theories: tolerate constructs proposed with sufficient reason; consider them seriously, both empirically and conceptually; develop alternative, behavior-analytic models with overlapping empirical domains; and contrast the various models. Through variation and selection, behavioral science will evolve.
In the suboptimal-choice task, birds systematically choose the leaner but informative option (suboptimal) over the richer but non-informative option (optimal). The task has two variations. In the standard task, the optimal option includes two terminal link stimuli. In the original task, it includes a single terminal link stimulus. Two models, the temporal information account (Cunningham and Shahan, J Exp Psychol Anim Learn Cogn 44:1–22, 2018) and the ∆-∑ hypothesis (González et al., J Exp Anal Behav 113:591–608, 2020), presuppose that these procedures are equivalent, but no formal comparison is available. Here we test whether or not these procedures are functionally equivalent. One group of pigeons was trained with the standard procedure, another group with the original procedure, and a third group was trained with a hybrid of the other two (i.e., the two options were the optimal links of the standard and original procedures). Our findings indicate that the number of terminal link stimuli in the optimal option is inconsequential vis-à-vis choice. Moreover, our findings also indicate that latencies to respond are a sensitive metric of value and choice. As predicted by the Sequential Choice Model, we were able to predict simultaneous choices from the latencies of sequential choices and observed a substantial shortening of latencies during simultaneous choices.