Accurate and efficient objective assessment methods for multidimensional quality of experience (QoE) of stereoscopic 3D content are in urgent demand. Especially, depth perception quantization for source 3D content is necessary to optimize the comprehensive QoE. A few studies evaluate the depth quality degradation due to compression or transmission, while the objective assessment of perceptual depth intensity has not been deeply investigated. In this paper, we propose a depth intensity metric (DIM) to simulate subjective evaluation on uncompressed source 3D images. Considering the multi-cue aggregation mechanism in depth perception as well as visual characteristic of binocular fusion, perceptual features are extracted from binocular and monocular cues to generate the depth intensity perception score. The effectiveness of DIM is verified by conducting subjective experiment on publicly available stereoscopic image databases. Experimental results demonstrate that DIM can achieve higher accuracy compared to the state-of-the-art depth intensity metric.
Functional relations can be directly “seen” and used to organize perceptual representations, but interactive positions are prerequisites for such perceptual grouping. The current study examined whether visual working memory (VWM) could automatically take advantage of functional relations in a more flexible way. In three experiments, participants were required to memorize a set of objects while ignoring the underlying relations between them. Results showed that although functional grouping was task-irrelevant, functional relations between objects could be extracted and used to enhance memory performance. More interestingly, functionally related objects could be grouped in the working memory phase even when they were not spatial interactive (Experiment 2). It was different from the perceptual grouping effect (Experiment 1) and reveals a unique memory grouping mode. Moreover, such functional grouping could still happen when object pairs entered VWM sequentially (Experiment 3), suggesting an active modification according to functional relations inside VWM. These findings suggest that VWM representations can be automatically structured through functional grouping, and grouping takes place in a flexible manner that can break the spatiotemporal constraints of perception.
The limited capacity of visual working memory (VWM) requires an efficient information selection mechanism. However, an uneconomical object-based encoding manner of VWM has been consistently found in previous studies. That is, not only the target feature, but also the task-irrelevant features from the same object are extracted into VWM. Besides the totally task-irrelevant feature which is never useful throughout the whole task, there is another kind of “irrelevant feature” that is necessary initially but no longer useful for the left task, termed as key feature in Chen & Wyble (2015)’s Attribute Amnesia studies. Previous studies showed that despite participants could not explicitly report the key feature in a surprise memory test, they had some memory traces for this unreportable information that could still produce an inter-trial priming effect. The current study sought to investigate the status of memory representation of a key feature by directly comparing it with the memory representation of a totally irrelevant feature from the same object (which serves as a baseline). In a series of experiments, participants were asked to memorize one feature of a single object, and then performed a visual search task in which there was either one of the distractors matching the task-irrelevant feature or key feature of the memory item, or there was no match between the memory and search display. Surprisingly, the results convergingly showed that despite a reliable WM-driven attentional bias effect (i.e., longer search time in the match condition than the no-match neutral condition) was generated by a task-irrelevant feature of an object, there was no such an effect triggered by the key feature. These findings suggested that participants might tend to actively inhibit just used information (i.e., key feature), resulting in even weaker memory representation of such information as compared to that of completely task-irrelevant features.
A visual scene can be recursively parsed into parts and sub-parts. We present a causal model of scene parsing that can synthesize and identify parse trees, predict perceptual complexity, and pass a (limited) Turing test. It describes the causal process of generating a scene as splitting it recursively with vertical and horizontal cuts, modulated by two parameters: (a) Splitting Factor (SF) – a large SF favors splitting a part into more sub-parts, generating a wide and shallow tree; (b) Part Similarity (PS) – a large PS favors evenly splitting a part into sub-parts. Given a scene, it infers the best parsing tree by evaluating the number of parts and their similarities at each partition. It inspires three human experiments beyond RT/accuracy measurements. (1) "Just cut it". Participants freely cut a blank scene into 6 rectangles recursively. With human-generated images, the model removed all free parameters by estimating human SF and PS priors. (2) "Complexity comparison". Some scenes are immediately perceived as more complex than others. Scene complexity can be quantified as "information content", determined by the probability of generating that scene (higher probability carries less information). Participants ranked the complexities of 20 scenes via paired comparisons. The model ranked the same images by computing their information contents. The result revealed a strong correlation between human and model rankings (r2 = 0.85). (3) Turing test. Each scene can be interpreted by multiple parsing trees. Participants viewed a scene and one parsing tree, then reported whether the tree was generated by a human or machine. Two baseline models were introduced: (a) the causal model with non-informative SF and PS priors; (b) a model uniformly sampling a tree from valid ones. Only the causal model with human priors passed Turing test. These results demonstrate how to formalize human scene parsing with a causal model. Meeting abstract presented at VSS 2018
Agents are not omnipotent. Instead, their motions are driven by both their internal intentions and outside constraints (such as an impulsive dog resisting its leash). In many cases, agents are even pulled away from their goals. How does vision interpret such intention-motion dissociations? One answer is that it simply can't, as suggested by studies showing that vision has little tolerance for deviations from perfect goal-directed motions (e.g. Gao et al., 2009). Alternatively, vision can detect intentions from deviated motions, provided it can explain away deviations as causal constraints imposed by the environment. We tested this hypothesis with the Search-For-Chasing task, in which a "wolf" pursues a "sheep" among distractors. The wolf's sheep-directed motion is compromised either by Causal or Non-Causal deviations. Causal deviations are generated by introducing the classic "motion hierarchy" phenomenon (Johansson, 1950) into a chasing display. The hierarchy puts the wolf on a virtual "leash", dragging the wolf away from the sheep-direction by composing the wolf's pursuit with the motion of its superior in the hierarchy. Non-Causal deviations are created by keeping the magnitude of the deviations while destroying the motion hierarchy. In Expt.1, both the wolf and sheep are constrained by a single superior, producing an averaged 45° deviation in the wolf's pursuit. Chasing detection is 21% higher with Causal deviations, compared to Non-Causal deviations. In Expt.2, the wolf is the subordinate of a distractor while the sheep flees freely, producing an averaged 65° deviation. The performance is again 20% higher with Causal-deviations. These results demonstrate that perceived animacy is more robust and flexible than previously suggested, provided the display can be explained by a causal hierarchical structure. We summarize it as a "leash resistance" effect, in which vision intelligently interprets intention-motion dissociations by jointly inferring the agent's internal intentions and outside causal constraints. Meeting abstract presented at VSS 2018
Walking direction is an important attribute of biological motion because it carries key information, such as the specific intention of the walker. Although it is known that spatial attention is guided by walking direction, it remains unclear whether this attentional shift is reflexive (i.e., constantly shifts to the walking direction) or not. A richer interpretation of this effect is that attention is guided to seek the information that is necessary to understand the motion. To investigate this issue, we examined how backward-walking biological motion orients attention because the intention of walking backward is usually to avoid something that walking forward would encounter. The results showed that attention was oriented to the walking-away direction of biological motion instead of the walking-toward direction (Experiment 1), and this effect was not due to the gaze direction of biological motion (Experiment 2). Our findings suggest that the attentional shift triggered by walking direction is not reflexive, thus providing support for the rich interpretation of these attentional effects.
The "working" function of visual working memory (VWM) has been highlighted by recent studies. Findings demonstrated that the sequentially presented visual elements would be involuntarily integrated into visual objects or figures inside VWM, providing evidence that VWM functions as a buffer serving perceptual processes by storing the intermediate perceptual representations for further processing. In those studies, the number of visual elements was usually controlled within the capacity of VWM; however, the realistic environment we live in is so rich and complex that the visual system has to constantly deal with massive visual information. How is such enormous amount of information actually processed with limited VWM capacity? Notwithstanding researchers know that the visual system can extract statistical properties of crowds of objects to form ensemble representations, it is largely unclear whether and how ensemble representations integrate inside VWM. This issue was investigated in the present study. Participants viewed two temporally separated groups of discs, after a short time, they reported the memorized mean size of either one of the groups or the whole (i.e., all the discs in the two groups) by adjusting a probe disc. The results indicated that participants were able to report accurate mean size of each group and the whole set of discs, respectively. More importantly, the reported mean size of the whole could be predicted by the pooled mean calculated based on the reported means of two individual groups. This result suggested that the temporally separated ensemble representations stored in VWM are able to be integrated into a higher-level ensemble representation, using the perceived statistics of the crowds of objects. Thus, when the amount of objects exceeds the capacity of VWM, the visual system will chose to store the necessary statistics for describing the ensemble and supporting further statistical computation. Meeting abstract presented at VSS 2017
It has been suggested that visual working memory (VWM) is involved in integrating the sampled discrete information into a coherent visual percept. However, how this integration takes place in VWM for the sequentially processed information remains unclear. We recently demonstrated that VWM can realize and use potential Gestalt principles within the sequentially encoded representations: The closure and similarity cues among the sequentially presented objects significantly enhanced VWM performance relative to conditions without gestalt cues (Gao, Gao, Tang, Shui, & Shen, 2016, Organization principles in visual working memory: Evidence from sequential stimulus display. Cognition, 146, 277-288). In the current study, we examined (1) whether the VWM organization is an automatic process regardless of attention, (2) if VWM organization is a voluntary process, which type of attention plays a pivotal role. To this end, we displayed the to-be-memorized stimuli sequentially and in half of trials there were gestalt cues; critically, an attention-consuming task (visual search task for space-based attention or mental rotation task for object-based attention) was added into the maintenance phase of VWM. We predicted that if the VWM organization was an automatic process, the secondary task should affect the organization effect. If the VWM organization was a voluntary process, we predicted that the secondary task would erase the organization effect. Experiments 1 and 2 tested the role of attention underlying closure, and found that visual search task erased the organization effect while mental rotation task did not. Experiments 3 and 4 tested the role of attention underlying similarity, and found that mental rotation task erased the organization effect while visual search task did not. Together, we suggest that VWM organization is a voluntary process, yet the key attention is determined by the nature of the memorized stimuli. Meeting abstract presented at VSS 2017
Biological motion (BM) broadly refers to the movements of animate entities. It contains abundant social information, therefore, recognizing and understanding biological motion is of great survival significance to human beings. A few recent studies have begun to reveal the underlying mechanisms of holding BM information in working memory, for instance, by showing that BM information is stored independently from color, shape, and location. However, no study so far has investigated the interaction between affect and holding BM in working memory. The current study explored this issue by exploring the impact of happy, neutral, and negative affect in holding BM in working memory. In Experiment 1, we required participants to remember 2-5 BM stimuli in a change-detection task after inducing different affective states. We found that working memory capacity of BM was significantly dropped in the negative affect condition than in the happy or neutral affect condition, and no difference was found between the latter two. The reduction of BM capacity led by the negative affect was further confirmed in Experiment 2, by using an EEG index of mu-suppression which could track the load of BM information stored in working memory. Finally, in Experiment 3 we examined whether negative affect would improve the precision of BM in working memory, and found that the BM precision was kept stable between neutral and negative affect. Taking together, the current study suggests that negative affect reduces working memory capacity of BM; yet the precision of BM in working memory was not affected. Meeting abstract presented at VSS 2016
Action prediction, a crucial ability to support social activities, is sensitive to the individual goals of expected actions. This article reports a novel finding that the predictions of observed actions for a temporarily invisible agent are influenced, and even enhanced, when this agent has a joint/collective goal to implement coordinated actions with others (i.e., with coordination information). Specifically, we manipulated the coordination information by presenting two chasers and one common target to perform coordinated or individual chases, and subjects were required to predict the expected action (i.e., position) for one chaser after it became momentarily invisible. To control for possible low-level physical properties, we also established some intense paired controls for each type of chase, such as backward replay (Experiment 1), making the chasing target invisible (Experiment 2) and a direct manipulation of the goal-directedness of one chaser's movements to disrupt coordination information (Experiment 3). The results show that the prediction error for invisible chasers depends on whether the second chaser is coordinated with the first, and this effect vanishes when the chasers behaves with exactly the same motions, but without coordination information between them; furthermore, this influence results in enhancing the performance of action prediction. These findings extend the influential factors of action prediction to the level of observed coordination information, implying that the functional characteristic of mutual constraints of coordinated actions can be utilized by vision.
Visual working memory (VWM) adopts a specific manner of object-based encoding (OBE) to extract perceptual information: Whenever one feature-dimension is selected for entry into VWM, the others are also extracted. Currently most studies revealing OBE probed an 'irrelevant-change distracting effect', where changes of irrelevant-features dramatically affected the performance of the target feature. However, the existence of irrelevant-feature change may affect participants' processing manner, leading to a false-positive result. The current study conducted a strict examination of OBE in VWM, by probing whether irrelevant-features guided the deployment of attention in visual search. The participants memorized an object's colour yet ignored shape and concurrently performed a visual-search task. They searched for a target line among distractor lines, each embedded within a different object. One object in the search display could match the shape, colour, or both dimensions of the memory item, but this object never contained the target line. Relative to a neutral baseline, where there was no match between the memory and search displays, search time was significantly prolonged in all match conditions, regardless of whether the memory item was displayed for 100 or 1000 ms. These results suggest that task-irrelevant shape was extracted into VWM, supporting OBE in VWM.
It has been suggested that visual working memory (VWM) adopts an object-based encoding (OBE) manner to extract perceptual information into VWM. That is, whenever even one feature-dimension is selected for entry into VWM, the others are also extracted automatically. Almost all extant studies revealing OBE were conducted by probing an "irrelevant-change distracting effect", in which a change of stored irrelevant-feature dramatically affects the change detection performance of the target feature. However, the existence of irrelevant feature change may affect the participants' processing manner, leading to a false positive result. In the current study, we conducted a strict examination of OBE in VWM, by keeping the irrelevant feature of the memory item constant while probing whether task-irrelevant feature can guide the early deployment of attention in visual search. In particular, we required the participants to memorize an object's color while ignoring shape. Critically, we inserted a visual search task into the maintenance phase of VWM, and the participants searched for a target line among distractor lines, each embedded within a different object. One object in the search display could match the shape, the color or both dimensions of the memory item, but this object never contained the target line. Relative to a neutral baseline (no match between the memory and the search displays), we found that the search time was significantly prolonged in all the three match conditions. Moreover, this pattern was not modulated by the exposure time of memory array (100 or 1000ms), suggesting that similar to the task-relevant feature color, the task-irrelevant shape was also extracted into VWM, and hence affected the search task in a top-down manner. Therefore, the OBE exists in VWM. Meeting abstract presented at VSS 2016
We report on how visual working memory (VWM) forms intact perceptual representations of visual objects using sub-object elements. Specifically, when objects were divided into fragments and sequentially encoded into VWM, the fragments were involuntarily integrated into objects in VWM, as evidenced by the occurrence of both positive and negative object-based attention effects: In Experiment 1, when subjects' attention was cued to a location occupied by the VWM object, the target presented at the location of that object was perceived as occurring earlier than that presented at the location of a different object. In Experiment 2, responses to a target were significantly slower when a distractor was presented at the same location as the cued object (Experiment 2). These results suggest that object fragments can be integrated into objects within VWM in a manner similar to that of visual perception.
Human beings are social in nature. Interacting with other people using human body is one of the most important activities and abilities in our daily life. To have a coherent visual perception of dynamic actions and engaging in normal social interaction, we have to store these interactive movements into working memory (WM). However, the WM mechanism in processing these interactive biological movements (BMs) remains unknown. In the current study, we explored the representation format of interactive BMs stored in WM, by testing two distinct hypotheses: (1) each interactive BM can be stored as one unit in WM (integrated unit hypothesis); (b) constituents of interactive BMs are stored separately in WM (individual storage hypothesis). In a change detection task, we required the participants to memorize interactive BMs in which two person were in beating, dancing, dashing, drinking, talking, or conversation. We found that there was no difference between memorizing four interactive BMs (containing eight individual BMs) and four individual BMs, and both performances were significantly lower than remembering two interactive BMs (Experiment 1). In Experiment 2, we further test whether spatial proximity but not social interaction resulted in the results of Experiment 1, by introducing a random-pair condition in which the social interaction was destroyed but spatial proximity still existed. We found that participants remembered equally well between two interactive BMs and two individual BMs, and both performances were significantly better than remembering two random-pair BMs; there was no difference between memorizing two random-pair BMs and four individual BMs. Together, these results suggest that an interactive biological movement containing two individual moments are stored as one unit in WM. Meeting abstract presented at VSS 2016
Although the mechanisms of visual working memory (VWM) have been studied extensively in recent years, the active property of VWM has received less attention. In the current study, we examined how VWM integrates sequentially presented stimuli by focusing on the role of Gestalt principles, which are important organizing principles in perceptual integration. We manipulated the level of Gestalt cues among three or four sequentially presented objects that were memorized. The Gestalt principle could not emerge unless all the objects appeared together. We distinguished two hypotheses: a perception-alike hypothesis and an encoding-specificity hypothesis. The former predicts that the Gestalt cue will play a role in information integration within VWM; the latter predicts that the Gestalt cue will not operate within VWM. In four experiments, we demonstrated that collinearity (Experiment 1) and closure (Experiment 2) cues significantly improved VWM performance, and this facilitation was not affected by the testing manner (Experiment 3) or by adding extra colors to the memorized objects (Experiment 4). Finally, we re-established the Gestalt cue benefit with similarity cues (Experiment 5). These findings together suggest that VWM realizes and uses potential Gestalt principles within the stored representations, supporting a perception-alike hypothesis.
Visual working memory is highly sensitive to global configurations in addition to features of each individual object. When objects are moving, their configuration varies correspondingly. Here we explore the geometric rules governing the maintenance of such a dynamic configuration in visual working memory. Our investigation is guided by the Erlangen program, which is a hierarchy of geometric stability, including affine, projective and topological invariants. The configuration here was defined as the virtual polygon with the four dots being its vertices. In all cases, the boundary of the virtual polygon overlapped with the convex hull of the four dots. The shape of the virtual polygon was gradually transformed to a new one by the motion of each dot in the memory display. In a change detection task, this memory displays were categorized by which geometry invariance was violated by the objects' motions. The results show that (a) there was no decrement of memory performance until the projective invariance was violated; (b) more dramatic changes (such as a topological change) cannot further enlarge the decrement; (c) objects causing the violation of projective invariance were better encoded in memory. These results collectively demonstrate that projective invariance is the only geometric property determining the maintenance of a dynamic configuration in visual working memory. (Acknowledgement: This research is supported by the National Natural Science Foundation of China (No. 31170974; 31170975).) Meeting abstract presented at VSS 2015
Visual working memory is highly sensitive to global configurations in addition to the features of each object. When objects move, their configuration varies correspondingly. In this study, we explored the geometric rules governing the maintenance of a dynamic configuration in visual working memory. Our investigation is guided by Klein's Erlangen program, a hierarchy of geometric stability that includes affine, projective, and topological invariants. In a change-detection task, memory displays were categorized by which geometric invariance was violated by the objects' motions. The results showed that (a) there was no decrement in memory performance until the projective invariance was violated, (b) more dramatic changes (such as a topological change) did not further enlarge the decrement, and (c) objects causing the violation of projective invariance were better encoded into memory. These results collectively demonstrate that projective invariance is the only geometric property determining the maintenance of a dynamic configuration in visual working memory.
以往有关异族效应的研究主要聚焦在其在知觉阶段的加工机制,尚未有研究探讨异族效应对工作记忆加工过程的影响.本研究采用延迟匹配范式,以反映比较过程中表征冲突的前额N2为指标,对异族效应影响工作记忆比较过程的机制进行了探讨.结果发现,异族效应调节人脸在工作记忆中的比较过程:本族人脸较异族人脸可诱发更负的前额N2.该研究为异族效应的发生机制提供了新的可能解释.