A classic distinction in perceptual information processing is whether stimuli are composed of separable dimensions, which are highly analyzable, or integral dimensions, which are processed holistically. Previous tests of a set of logical-rule models of classification have shown that separable-dimension stimuli are processed serially if the dimensions are spatially separated and as a mixture of serial and parallel processes if the dimensions are spatially overlapping (Fifić, Little, & Nosofsky, 2010; Little, Nosofsky, & Denton, 2011). In the current research, the logical-rule models are applied to predict response-time (RT) data from participants trained to classify integral-dimension color stimuli into rule-based categories. In dramatic contrast to the previous results for separable-dimension stimuli, analysis of the current data indicated that processing was best captured by a single-channel coactive model. The results converge with previous operations that suggest holistic processing of integral-dimension stimuli and demonstrate considerable generality for the application of the logical-rule models to predicting RT data from rule-based classification experiments.
Studies of incidental category learning support the hypothesis of an implicit prototype-extraction system that is distinct from explicit memory (Smith, 2008). In those studies, patients with explicit-memory impairments due to damage to the medial-temporal lobe performed normally in implicit categorization tasks (Bozoki, Grossman, & Smith, 2006; Knowlton & Squire, 1993). However, alternative interpretations are that (a) even people with impairments to a single memory system have sufficient resources to succeed on the particular categorization tasks that have been tested (Nosofsky & Zaki, 1998; Zaki & Nosofsky, 2001) and (b) working memory can be used at time of test to learn the categories (Palmeri & Flanery, 1999). In the present experiments, patients with amnestic mild cognitive impairment or early Alzheimer's disease were tested in prototype-extraction tasks to examine these possibilities. In a categorization task involving discrete-feature stimuli, the majority of subjects relied on memories for exceedingly few features, even when the task structure strongly encouraged reliance on broad-based prototypes. In a dot-pattern categorization task, even the memory-impaired patients were able to use working memory at time of test to extract the category structure (at least for the stimulus set used in past work). We argue that the results weaken the past case made in favor of a separate system of implicit prototype extraction.
Visual identification of briefly presented target words is affected by the presence of nondiagnostic prime words that immediately precede the target, flanker words simultaneously presented adjacent to the target, and visual masks that immediately follow the target in the same location. Priming is duration dependent: In a forced choice target identification task, brief primes produce a strong preference to choose the primed alternative, whereas long primes have the opposite effect. The ROUSE model (Huber, Shiffrin, Lyle, & Ruys, Psychological Review 108:149–182, 2001) predicts this interaction by positing that prime features are confused with target features and that evidence regarding the prime features is discounted less for short primes and more for long primes, when both are compared with the optimal level. In the present study, we augmented the typical short-term priming experiment by adding flankers that appeared simultaneously with the target and remained for a short or long duration. In the experiment, we replicated previous priming effects and produced novel effects of flanker duration. ROUSE accounted for both the priming and flanker findings with the previously posited processes, but with different quantitative parameters for flankers: Relative to optimal levels of discounting, all flanker features were underdiscounted, but longer flankers were discounted more than short flankers.
A recent resurgence in logical-rule theories of categorization has motivated the development of a class of models that predict not only choice probabilities but also categorization response times (RTs; Fifić, Little, & Nosofsky, 2010). The new models combine mental-architecture and random-walk approaches within an integrated framework and predict detailed RT-distribution data at the level of individual participants and individual stimuli. To date, however, tests of the models have been limited to validation tests in which participants were provided with explicit instructions to adopt particular processing strategies for implementing the rules. In the present research, we test conditions in which categories are learned via induction over training exemplars and in which participants are free to adopt whatever classification strategy they choose. In addition, we explore how variations in stimulus formats, involving either spatially separated or overlapping dimensions, influence processing modes in rule-based classification tasks. In conditions involving spatially separated dimensions, strong evidence is obtained for application of logical-rule strategies operating in a serial-self-terminating processing mode. In conditions involving spatially overlapping dimensions, preliminary evidence is obtained that a mixture of serial and parallel processing underlies the application of rule-based classification strategies. The logical-rule models fare considerably better than major extant alternative models in accounting for the categorization RTs.
The phenomenon in associative learning called ‘blocking’ has been a central target for theories that aim to explain learning and attention. The training phases of the blocking procedure, described in detail below, can be run in backward order, yet still produce the blocking effect (e.g., Shanks,1985; Dickinson and Burke,1996; Kruschke and Blair, 2000). Explaining forward and backward blocking, along with other associative learning phenomena, has proven to be challenging. Some accounts of forward blocking include an attentional mechanism that modulates the influence of the cues on learning and/or responding (e.g., Kruschke, 2001; Mackintosh, 1975). Some accounts of backward blocking employ a Bayesian framework in which different combinations of associative weights are considered simultaneously, with more belief allocated to the combination that is most consistent with training items (e.g., Dayan and Kakade, 2001; Tenenbaum and Griffths, 2003). A theoretical framework that is able to combine the attentional and Bayesian approaches is called ‘locally Bayesian learning’ (Kruschke, 2006b). The framework is based on the idea that a learning system may consist of a sequence of subsystems in a feed-forward chain, each of which is a locally Bayesian learner. The argument for locally-learning layers was as follows. First, Bayesian learning is very attractive for explaining retrospective revaluation effects such as backward blocking, among many other phenomena (Chater, Tenenbaum, and Yuille, 2006). Second, globally Bayesian learning may also be unattractive for a number of reasons. In a large learning system, there are too many combinations of parameters to keep track of in a monolithic joint parameter space. Furthermore, many globally Bayesian models do not explain learning phenomena (such as ‘highlighting’, Kruschke, 2010) that depend on training order, because the models treat all training items as equally representative of
Exploring Active Learning in a Bayesian Framework Stephen Denton Indiana University, Bloomington John Kruschke Indiana University, Bloomington Abstract: Bayesian approaches provide a framework for models of active learning–learning in which stimuli are actively probed to disambiguate potential beliefs regarding outcomes. Within a Bayesian framework, uncertainty across beliefs is inherently represented and expected uncertainty reductions for candidate stimuli can be evaluated. Bayesian active learning models offer the prediction that an active learner would select the stimuli for which the expected uncertainty across all hypotheses is minimized. This research contrasts four possible hypothesis spaces for active learning consisting of two simple cue-combination models and two possible priors. An automated search of associative learning structures for which the models make maximally different predictions was performed. Participants were tested on these same structures in an allergy diagnosis context and were asked which cues they would find the most informative to learn about; i.e., their active learning preferences were assessed. Model and prior combinations that best mimic human active learning are discussed.
The authors conducted a short-term repetition priming experiment (using a visual, forced-choice word identification task) that compared a standard priming condition, where prime and target words appeared in the same spatial location, with an experimental condition in which prime and target words were spatially separated enough to necessitate an eye movement. Prime presentation duration was manipulated and, within both eye movement conditions, it was found that short primes produced a preference to choose a primed alternative, whereas for longer duration primes this preference was absent. Based on the similarity between eye movement conditions, it is argued that prime and target features from separate fixations are still confusable and that evidence regarding prime feature must still be discounted. A computation model that includes these offsetting components of source confusion and discounting provides an excellent account of our results.
Short-Term Word Priming Across Eye Movements Stephen E. Denton (sedenton@umail.iu.edu) and Richard M. Shiffrin (shiffrin@indiana.edu) Department of Psychological and Brain Sciences, and Cognitive Science Program, Indiana University 1101 E. 10th Street, Bloomington, IN 47405 USA Abstract or the foil. The corresponding prime conditions are termed target-primed, foil-primed, and neither-primed, respectively. The preference effects revealed by Huber et al. (2001) proved to be readily manipulatable, changing in magnitude and direction as a function of prime saliency (e.g., Huber, Shiffrin, Quach, & Lyle, 2002; Weidemann et al., 2005; Wei- demann, Huber, & Shiffrin, 2008). Thus, when the prime is made more salient, either through a longer presentation du- ration or though an active task that required responding to the prime, the conventional priming effect diminishes or in some cases reverses. Prime saliency manipulations result in reduced facilitation (or slight deficits) in target-primed con- ditions, while leading to increased accuracy in foil-primed conditions. The authors conducted a short-term repetition priming exper- iment (using a visual, forced-choice word identification task) that compared a standard priming condition, where prime and target words appeared in the same spatial location, with an ex- perimental condition in which prime and target words were spatially separated enough to necessitate an eye movement. Prime presentation duration was manipulated and, within both eye movement conditions, it was found that short primes pro- duced a preference to choose a primed alternative, whereas for longer duration primes this preference was absent. Based on the similarity between eye movement conditions, it is argued that prime and target features from separate fixations are still confusable and that evidence regarding prime feature must still be discounted. A computation model that includes these offset- ting components of source confusion and discounting provides an excellent account of our results. Keywords: short-term priming; immediate priming; repetition priming; perceptual identification The ROUSE Model The term priming refers to a well-known information pro- cessing effect wherein one stimulus (a prime) influences a similar or related stimulus (a target) presented at a later time. The influence the prime has on the target is usually one of facilitation. In the typical priming task trial, the prime is pre- sented first followed by a target that is briefly flashed and masked. This paradigm is referred to as short-term or im- mediate priming because primes are presented immediately prior to the targets, with inter-stimulus intervals generally less than a second. The first stimuli is called a prime because it is thought to “prime the pump” for a related target, yielding faster and more accurate responding. However, priming does not simply result in facilitation. Previous research has indi- cated that there is an intricate set priming effects that occur across various situations and experimental conditions (e.g., Huber, Shiffrin, Lyle, & Ruys, 2001; Weidemann, Huber, & Shiffrin, 2005). Some experiments find facilitation, while other fail to find facilitation or even find reliable deficits due to priming. Huber et al. (2001) began to systematically examine the effects of priming using a two-alternative forced-choice (2- AFC) testing procedure with the goal of separating the per- ceptual and decisional components involved in priming. Ob- servers were asked to correctly identify a previously flashed target word when given a choice between it and an incorrect foil word. This study indicated that priming largely arises from preferences to choose whatever has been primed. For example, in repetition priming (priming in which the prime can be the same word as either the target or foil), it was found was that with short prime presentation durations, prim- ing with the target increases accuracy and priming with the foil decreases accuracy, when both are compared to a control condition in which the prime is unrelated to either the target To account for a range of findings within the 2-AFC iden- tification paradigm, Huber et al. (2001) developed a feature- based Bayesian model of short-term priming. The responding optimally to unknown sources of evidence (ROUSE) model accounted for experimental priming data by incorporating the two offsetting components of feature source confusion and discounting. The source confusion portion of the model posits that features can be carried over from the prime to the target percept, without source information. Thus, when a choice word is presented, feature activations could be due to the prime presentation, the target presentation, and/or noise without the source of the activation being known to the sys- tem. This factor alone can produce the standard priming ef- fect, as it causes a preference toward prime-related choice words. The discounting mechanism in the model can coun- teract this preference, because this component posits that per- ceived features are assigned evidence and feature evidence is discounted when known to have come from the prime. A Bayesian decision process then arrives at an optimal response given it is operating on this noisy and discounted evidence. The implication here is that making the prime more salient (e.g., long presentation duration) leads to increased discount- ing of prime feature evidence. Thus, discounting mechanisms can explain a lack of, or a reversal in, the typical priming pref- erence. In the ROUSE model, choice words are represented as a feature vector, typically consisting of 20 binary fea- tures (Huber et al., 2001). Features can be independently activated by the prime, with a probability α, by the target, with a probability β, or be activated due to noise, with a prob- ability γ. The system is assumed to only have access to what features are active and not the source of their activation (i.e., source confusion), therefore the probabilistic effect of α, if
Erickson and Kruschke (1998, 2002) demonstrated that in rule-plus-exception categorization, people generalize category knowledge by extrapolating in a rule-like fashion, even when they are presented with a novel stimulus that is most similar to a known exception. Although exemplar models have been found to be deficient in explaining rule-based extrapolation, Rodrigues and Murre (2007) offered a variation of an exemplar model that was better able to account for such performance. Here, we present the results of a new rule-plus-exception experiment that yields rule-like extrapolation similar to that of previous experiments, and yet the data are not accounted for by Rodrigues and Murre’s augmented exemplar model. Further, a hybrid rule-and-exemplar model is shown to better describe the data. Thus, we maintain that rule-plus-exception categorization continues to be a challenge for exemplar-only models.
The associative learning effect called blocking has previously been found in many cue-competition paradigms where all cues are of equal salience. Previous research by Hall, Mackintosh, Goodall,and dal Martello (1977) found that, in animals, salient cues were less likely to be blocked. Crucially, they also found that when the to-be-blocked cue was highly salient, the blocking cue would lose some control over responding. The present article extends these findings to humans and suggests that shifts in attention can explain the apparent loss of control by the previously learned cue. A connectionist model that implements attentional learning is shown to fit the main trends in the data. Model comparisons suggest that mere forgetting, implemented as weight decay, cannot explain the results.