
Visual information typically dominates auditory cues in spatial localization, particularly in depth, leading to the proximity image effect (PIE) in which displaced sounds are perceived closer to plausible visual objects than they actually are (visual capture). It is unclear whether fundamental properties of audiovisual integration generalize to virtual environments where spatial perception is systematically distorted. Across two experiments, we tested whether real-world sound sources are visually captured by virtual objects in immersive virtual reality (VR). Participants completed a perceptual-matching task in which they indicated when a real-world sound source aligned with the depth of a virtual object. In Experiment 1, visual capture reliably occurred; however, the effect was opposite to real-world findings, with stronger capture happening when the sound originated in front of the visual object rather than behind it. Experiment 2 replicated this pattern and revealed significant visual distance underestimation measured with a blind-walking task. When judgments of auditory-visual co-location were interpreted relative to the participants' perceived visual distance rather than the actual physical distance, the expected real-world symmetry reemerged. These results demonstrate that core principles of multisensory integration are preserved in VR, but VR-specific biases, such as the visual underestimation of distance, may alter the direction of the established effects.
Interpersonal distance is a fundamental social cue, yet it remains unclear whether repeated exposure to gaze-triggered approach or recede animations is associated with subsequent gaze allocation. This study examined whether such gaze-triggered changes in apparent interpersonal distance were associated with changes in dwell-time allocation toward face stimuli. Fifty adults completed two within-subject blocks (Approach vs. Recede), each consisting of Baseline, Training, and Test phases. During Training, fixating a designated face for a gaze-duration threshold triggered a predefined approach or recede animation that proceeded independently of subsequent gaze behavior, while the comparison face remained static. Gaze allocation was operationalized as the proportion of dwell time directed toward the face associated with the gaze-triggered animation and was assessed during the static Baseline and Test phases. Linear mixed-effects analyses revealed a significant Condition × Phase interaction: dwell-time proportion decreased from Baseline to Test following gaze-triggered approach animations, whereas no reliable group-level change was observed following recede animations. These findings indicate that repeated exposure to gaze-triggered changes in apparent interpersonal distance is associated with systematic changes in subsequent gaze allocation. The approach-related reduction suggests that gaze-triggered distance changes may redistribute later gaze allocation rather than simply increasing looking toward the stimulus associated with approach. The observed effects do not imply subjective evaluation, value acquisition, or a specific learning mechanism. This paradigm provides a controlled approach for examining how gaze-triggered interpersonal-distance changes are associated with later attentional allocation. De-identified trial-level data and analysis code for this study are available via the Open Science Framework: https://doi.org/10.17605/OSF.IO/EG4B6 . This study was not preregistered.
Responses to targets are often slower when distractors are incongruent rather than congruent with the required response. Past research indicates that alerting cues can magnify this congruency effect, such that responses are faster when the target is cued, but distractors have a greater negative influence. This alerting-congruency interaction is reliably found in the arrow-version of Eriksen’s flanker task but not in the Stroop task, leading some to suggest that the interaction only occurs with stimuli that have pre-existing directional associations. The current experiments test the possibility that the alerting-congruency interaction depends on the timing of distractor processing, rather than directional associations, by varying whether the distractor appears before the target or simultaneously with the target. In Experiment 1 (N = 158) using Stroop stimuli, and Experiment 2 (N = 157) using the arrow-version of the Eriksen flanker task, results replicate past findings: no alerting-congruency interaction in the Stroop task and a robust interaction in the arrow-flanker task when distractors and targets appear simultaneously. Importantly, previewing the distractor reverses both of these patterns. Delta plots showing congruency effects across the reaction-time distribution support the conclusion that the timing of distractor activation may be critical. Discussion focuses on implications for theories of the alerting-congruency interaction. Subject-level data for all experiments are publicly available at https://osf.io/uxebp/?view_only=0db3638a66454684aee9fc41550a3a48 . The studies reported in this article were not preregistered.
Recent research demonstrates that if two identical squares moving uniformly towards each other suddenly disappear when their leading edges contact, participants will have an illusory perception of the two squares as non-contact before their disappearance. Inspired by the “sound-induced visual segmentation” hypothesis that when two approaching 2D objects reach their closest state, a coincident sound will lead to an overestimation of their distance, the current study explored whether a task-irrelevant sound delivered when the two squares make contact can enhance this non-contact illusion and examined its underlying mechanisms. Across Experiments 1 and 2, the probability of the non-contact response was consistently higher when the sound was present than when it was absent. This sound-induced increase is inherently free from the collision-based response bias—if the sound had suggested a collision between two objects, it would have led to more contact, rather than non-contact, responses. Moreover, the non-contact response was never slower than the contact response in the sound-present condition, which contradicts the attentional distraction explanation because it would predict slower non-contact responses in that condition. Meanwhile, the high time-resolution event-related potential (ERP) data in Experiment 2 showed that the early cross-modal PD170 component (125–175 ms after sound onset) was significantly larger for the sound-present non-contact response than for the sound-present contact response, which indicates that low-level audiovisual cross-modal interactions underlie the sound-induced enhancement of the non-contact illusion, substantiating its perceptual nature. Together, the current findings establish this sound-induced enhancement as a promising new paradigm for studying multisensory processing. The trial-level data, subject-level data, analysis scripts and experimental files for both experiments are available at https://osf.io/zt92d/overview . This study was not preregistered.
We investigated the effects of periodic performance feedback across three experiments. In Experiment 1, participants were assigned to either a feedback condition, in which there was periodic performance feedback (n = 302), or to a control condition, which received no feedback (n = 293). Whereas the control condition demonstrated a large vigilance decrement, even over a brief task duration, participants in the feedback condition performed both better overall and demonstrated a significantly shallower decrement. Experiment 2 (n = 103) was a preregistered replication with an extended task duration, and we directly replicated the findings. Experiment 3 was another preregistered experiment (n = 283), in which we examined three conditions with the extended duration: feedback for the entire task, feedback for only the first half of the task, and no feedback for the entire duration. The conditions did not differ in performance. However, in a post hoc cross-experimental analysis, we discovered significant moderation of the feedback effect via fluid intelligence. The findings have implications for how we view vigilance as a cognitive construct, why it typically wanes with time, and how it may be stabilized. All data and materials are publicly available on the Open Science Framework ( https://osf.io/fzd2h/ )
Crossmodal correspondences (CMCs: systematic associations between feature dimensions across sensory modalities) are known to facilitate performance in simple detection and discrimination tasks and can bias reports of ambiguous figures. The present study tested whether a lightness/pitch CMC biases reported dominance in binocular rivalry. Across three experiments, participants viewed dichoptically presented black and white circles accompanied by relatively high- or low-pitched tones. Dominance reports were coded as congruent when they matched the lightness/pitch correspondence (white/high, black/low) and incongruent when they did not (white/low, black/high). In two experiments, participants reported whether the dominant percept was black or white (direct task); in a third, they judged line orientation while lightness was irrelevant (indirect task). Congruency effects were observed only in the direct task, and only at longer timescales: in intermittent presentation, at 5-s and 8-s exposures; and in continuous presentation, for 5-s sound alternations with responses following tone changes by 1.6–3.2 s, peaking at about 2.6 s. No reliable effects were found in the indirect task. Dominance response durations averaged approximately 2.5 s in the direct task but were significantly longer, approximately 3.5 s, in the indirect task. These results suggest that a pre-existing, untrained lightness/pitch correspondence can bias reported rivalry dominance, but only when the relevant feature is attended and during temporal windows aligned with natural rivalry dynamics. Because dominance was measured by report, the findings are best interpreted as effects on reported dominance rather than as definitive evidence for purely perceptual modulation; nevertheless, they are consistent with subtle crossmodal influences during phases of rivalry instability. The data for this study are publicly available on the Open Science Framework (OSF) at https://osf.io/u3ycm/overview?view_only=9fc10ddf4d6c4af8a7d70f2e06c2aef6 . The study was not preregistered.
Absolute pitch is the ability to label the pitch chroma of isolated tones. Recent work has demonstrated rapid improvements in pitch identification using an incidental contingency learning task. In the present work, we used a similar approach to assess whether knowledge about the pitch chroma (rather than pitch heights) of tones is possible in our task. In Experiment 1 with nonmusicians, seven notes in two octaves were played most often with their corresponding written note names and less often with the others. Tones from a third octave were played equally often with all note names (i.e., no contingency manipulation) to test for chroma-based generalization. We found significant but very small learning and generalization effects, revealing the relative difficulty of focusing on the pitch chroma during learning. The same task was used again in a sample of musicians in Experiment 2 with the thinking that musicians would have enhanced pitch chroma perception. As expected, we found more compelling evidence for learning and generalization, albeit only in explicit identification performance. In Experiment 3, we assessed whether it is fundamentally impossible to acquire knowledge based on pitch chroma by training participants with Shepard tones, which are largely devoid of pitch height cues. Interestingly, we found robust learning effects for these tones. However, we did not find evidence for transfer toward piano tones. We conclude that the major reason why absolute pitch is difficult to learn in adulthood is due to difficulty in perceiving pitch chroma due to the saliency of pitch height cues.
Previous research has shown that production of an action toward a visual stimulus may prioritize features of that stimulus for a brief period of time. One recent report has suggested that such an action effect specifically influences ensemble processing – the extraction of summary statistics from a display of multiple items. In the present study we replicate the results that led to that conclusion by having participants assess the mean size of elements in an ensemble after having made an action to either a small or large object. The size of the acted-on object biased the ensemble judgment, but merely viewing the same object produced no such bias. In two subsequent experiments we found the same pattern of results when participants were asked to judge the size of an individual object – not an ensemble. Because the same pattern occurred even when ensemble processing was not required, the action appears to affect some aspect of the task other than ensemble processing (such as perception of visual stimuli in general, or processes related to memory, decision making or response selection), limiting the conclusions that are possible about ensemble processing, but revealing an important new effect of action.
When navigating the visual world, we have considerable control over what stimuli to select for deeper processing, and we can often resist capture by physically salient but task-irrelevant distractors. A debate has emerged regarding whether distractor avoidance is accomplished by active goal-driven attentional control, or by passive exclusion of distractors from the narrow attentional focus of a highly serial search, as the attentional window account claims. In the present study, in order to explore the issue, we measured the extent to which attentional avoidance of salient distractors is affected by the seriality of the search. In three experiments, we invoked different levels of seriality by separately manipulating the similarity among search elements, the stimulus pattern complexity, and the conjunctive features that defined the target. To assess attentional allocation, probe trials were interleaved with the search trials. While the results showed significant avoidance of color singleton distractors in all experiments, the strength of the avoidance effect, as measured by both search latency and probe report, did not differ as a function of the seriality of the search. An additional control experiment verified the item-by-item serial search processes invoked by most of our methods. The results are inconsistent with the passive filtering proposed by the attentional window account and instead imply that active top-down control is a more likely explanation for distractor avoidance. Trial-level data and analysis code for all experiments have been made publicly available online ( https://osf.io/njvu2 ), and none of the experiments was preregistered.
Music is the synthesis of a multitude of components: the combination of spectrotemporal structures emerging from different acoustic or electronic instruments, carefully orchestrated to form a connected whole. From these complex textures, our auditory system groups components through a process called musical scene analysis (MSA). In this review, we aim to address several questions about MSA. What are the perceptual principles underpinning musical scene perception? How are these principles being probed with different experimental paradigms? What stimulus and listener factors shape MSA? And what are future perspectives that further drive this field of research? We will find that Gestalt principles from psychology help us organize a musical scene through primitive and schema-based grouping cues. Such principles are being investigated through both minimalist and ecological experimental designs that give rise to several factors affecting MSA. We will present studies that have investigated blend and segregation, salience, and complexity, as stimulus-based aspects, and both peripheral and cognitive factors, which explain inter-individual differences among listeners. We conclude by exploring different perspectives on how music is attuned to MSA and the aesthetic affordances of multi-source music – perspectives that might inform future research in the field.
Previous research has shown that people are sensitive to statistical regularities and will implicitly learn to bias attention toward locations and features frequently associated with search targets. However, most prior work has involved a single biasing contingency. In Experiment 1, we investigate whether individuals can simultaneously learn and implement two distinct contingencies: one based on location and another based on color. The results suggest that both contingencies are learned implicitly and exert independent effects on attentional allocation. Experiments 2 and 3 examine whether shifting one feature to the volitional system via endogenous cues affects the implicit learning and implementation of contingencies for the other feature. The findings indicate that endogenous cues for one feature do not block implicit learning of contingencies in the other. However, the influence of implicit contingencies on attentional allocation depends on the validity of the volitional cues: they are effective when cues are neutral or valid, but not when invalid. This pattern suggests that the implicit and volitional systems operate hierarchically, with the volitional system taking precedence. Efficient attentional guidance is key to effective visual search. While both volitional control and statistical learning can guide attention, little is known about how attention is influenced when multiple guidance mechanisms compete. Using a task with competing statistical regularities, one for target location and one for color, we show that both statistical learning biases independently influence attention. Shifting a feature to the volitional control system does not disrupt the statistical learning of the other, but invalid volitional cues eliminate statistical learning effects, suggesting a hierarchical relationship. These findings offer new insight into how multiple guidance systems shape attentional allocation. Subject level data and SPSS analysis syntax available on the Open Science Framework at: https://osf.io/5vt74/overview . The experiments were not preregistered, but the data, and analysis syntax are publicly available at https://osf.io/5vt74/overview .
The visual system can quickly extract statistical properties of multiple objects (ensembles). Observers can explicitly access and report the distribution of low-level features, such as color. Here, we investigate whether explicit access extends to the distributions of high-level features (e.g., emotional expressions of a group of faces). In Experiment 1, we presented observers with ensembles that had a Gaussian, uniform, or bimodal distribution of emotional expressions (or colors as a low-level baseline condition). The task was to report the frequency of a randomly chosen feature value. Observers’ responses closely followed the underlying distribution for colors but were much noisier for the emotional expressions. Observers could distinguish all the three color distributions, while they could discern only the most contrasting distributions of emotional expressions (Gaussian and bimodal). Additionally, modeling showed that observers’ performance in both conditions was based on the integration of global information rather than the subsampling of a few objects. Experiment 2 showed that disrupting holistic face processing via face inversion manipulation does not change the observer’s performance for ensembles of emotional expressions. This suggests that holistic processing does not substantially contribute to forming distributional representations of emotional expressions. Instead, the visual system relies on low-level and mid-level correlates (e.g., mouth curvature, brows’ tilt) to build a distributional representation of multiple facial expressions. Our findings help to explain how people can quickly obtain emotional information from many faces in a crowd and have a rich perceptual experience despite severe capacity limitations of visual attention and working memory. The trial-level data, stimuli, stimuli presentation code and analysis code from all experiments reported in this article can be accessed online ( https://osf.io/cfuy3/ ). The studies were not preregistered.
Color singleton distractors’ interference is known to be modulated by rejection mechanisms based on the distractors’ probability of occurrence at different locations. Here, to further address color-singleton distractor rejection, we used a modified version of the additional-singleton paradigm, where four consecutive displays were presented in rapid succession on each trial. In Experiment 1 the irrelevant singleton was either presented in the last display (single location, 40
Distortions in time perception under pain have been widely reported, yet the mechanisms underlying these effects remain unclear. The Attentional Gate Model and prior empirical research indicate that attention and arousal are associated with pain and time perception, yet their roles remain unexamined in a unified framework. The study aimed to investigate the mediating roles of arousal, attention control, and attention bias in the relationship between pain and time perception. A systematic search of China National Knowledge Infrastructure, Wanfang, Web of Science, and PubMed identified studies reporting associations among pain, arousal, attention control, attention bias, and time perception. A two-stage meta-analytic structural equation modeling approach was used to test mediation, with multilevel meta-regression analyses examining moderation by age, gender, pain type, and target duration. A total of 154 articles (176 independent studies) were included. Pain was positively correlated with time perception, arousal, and attention bias. Both arousal and attention bias were positively correlated with time perception, while arousal was negatively correlated with attention control and attention bias. Mediation analysis showed that only arousal had a significant mediating effect. Moderation analysis indicated that age, gender, and target duration significantly moderated the relationships between attentional processes and time perception, with age additionally moderating the relationship between attention control and attention bias. Overall, arousal showed a mediating role between pain and time perception, while attentional effects are context-dependent. This study extends the Attentional Gate Model to pain and clarifies the roles of arousal and attentional mechanisms in temporal experience under pain. This study was preregistered on the Open Science Framework ( https://osf.io/suzhn/ ). The data and code are publicly available online ( https://osf.io/3ucvb/ ).
Research suggests that items from similar spatiotemporal contexts are more likely to be retrieved together. Here, we sought to test whether preserving temporal order from encoding to test might lead to a retrieval benefit. Then, we tested the participants using test items that were either presented in the same order as during encoding, or the order of items was different due to it being randomized. Across two experiments, we found that maintaining the same temporal order did not improve overall memory performance in the visual old/new recognition task we used. Next, we show that conditional-response probability plots demonstrate temporal grouping from the study order, but the strength of that grouping was unmodulated by whether the same order or a random order was used during recognition testing. Lastly, we used a two-alternative forced-choice task and again found that presenting old items in their original order, without intervening new items, did not enhance recognition memory performance. Thus, we find that for visually presented common objects tested with recognition, we do not see that the same order boosts participants’ ability to retrieve a sequence of visual objects from memory.
Sustained attention is notoriously difficult to maintain over time, with attentional lapses emerging rapidly during prolonged tasks. Motivation has been identified as a key factor in reducing these lapses and enhancing task engagement. According to recent accounts such as the goal-competition hypothesis, the perceived benefit associated with a goal helps sustain its active representation in working memory. These benefits can arise from the prospect of gaining a reward, avoiding a penalty, or both. However, it remains unclear whether avoiding a penalty – or the combination of penalty and reward – improved sustained goal maintenance to the same extent as pursuing a reward alone. To address this question, we recruited 30 participants to complete a “continuous performance task” (CPT) in which the type of benefit associated with the goal varied every 20 trials (reward, penalty, both, or no benefit). Replicating prior findings, we observed that rewards reduced attentional lapses compared to a no-benefit baseline. Avoiding penalties and combining both incentives also improved sustained attention in CPT, though the penalty condition was less effective than the combined condition. This suggests that while any perceived benefit can support goal maintenance, the motivational mechanisms may differ depending on the type of benefit – particularly when avoiding penalties alone. The data and codes are available on the Open Science Framework repository at: https://osf.io/scmht/?view_only=df4ed739b7df40b19cc1ba00097450c5 . It includes trial-level data, processed subject-level data, experiment task code, and analysis code. The study was not preregistered.
Auditory scene analysis enables listeners to segregate and group a mixture of sounds into coherent streams, but it remains controversial to what extent this process depends on the availability of central attention. While previous research suggests that the guidance of visual attention may be independent of central capacity limitations, it remains unclear whether auditory selective attention can operate in parallel with central attention. Across two online experiments, it was tested whether attentional search in complex auditory scenes is limited in capacity and subject to the central processing bottleneck using a dual-task paradigm with the locus-of-slack method. Participants were first asked to identify target sounds within auditory scenes of varying set sizes (1-8 sounds), either as a single task or together with a visual discrimination task at variable stimulus onset asynchronies (SOAs). Results from single-task blocks revealed robust auditory set-size effects on response times and error rates, indicating capacity limitations and serial processing during auditory scene analysis. In the dual-task paradigm, auditory search times increased at short SOAs, demonstrating a psychological refractory period. Critically, the auditory set-size effects persisted, but were not attenuated at short SOAs, suggesting that auditory search time has not been absorbed into the slack due to the central processing bottleneck. In contrast to some evidence from the visual attention literature, these findings indicate that auditory scene analysis and central attention depend on shared capacity limitations, thus highlighting the importance of attentional control in complex listening situations.
When processing hierarchical compound figures with local elements in a global shape (often called the "Navon task"), an observer may perceive information from the local or the global level more quickly and/or accurately. This difference is referred to as local-global bias. This basic principle - measuring a difference between the local and global level - is conceptually very broad and can be implemented with many different task designs and figure designs. Importantly, however, there exists a pervasive notion of this bias being a unitary construct or even representative of a stable perceptual or cognitive style. However, few studies have examined the various facets of local-global bias across different task designs or assessed the retest reliability of these measures. Our study aimed to deepen the understanding of local-global bias by using three different hierarchical compound figure tasks: the Navon task and two Kimchi-Palmer task variations (one preference-based, one performance-based). We tested 75 participants across two sessions 1 week apart. Our findings revealed that metrics from the Navon task are not test-retest reliable, whereas two metrics from the Kimchi-Palmer tasks are. Some metrics were correlated between the Kimchi-Palmer tasks, but not between the Kimchi-Palmer task and the Navon task. These results suggest that local-global bias is not a unitary construct, and that different tasks measure different underlying mechanisms. We discuss possible reasons for these differences and suggest that future research should use multiple local-global tasks or otherwise prioritize Kimchi-Palmer tasks for studying individual traits and possibly Navon tasks for temporary and group-level effects.
The impact of target prevalence (the proportion of trials in which a target appears) has been widely studied within visual search. The low prevalence effect is the finding that rare targets are often missed. Mind wandering, where one's attention drifts away from a primary goal, impairs performance across various tasks, but its influence in low-prevalence search remains unclear. We recruited 60 participants from the general population and examined their mind-wandering rates as they searched for either low- (10%) or high- (50%) prevalence targets. When participants were engaged in mind wandering, response accuracy was reduced and target-absent response times increased. Low-prevalence targets were missed more than high-prevalence targets, but to our surprise, we discovered that most of the low-prevalence effect could be attributed to mind wandering in the majority our analyses. This marks a substantial shift in our understanding of the prevalence effect.