
The motivation for use of biosensors in audiovisual media is made by highlighting problem of signal loss due to wide variability in playback devices. A metadata system that allows creatives to steer signal modifications as a function of audience emotion and cognition as determined by biosensor analysis.
Attentional stances are particular ways in which the location of the source or seat of attention is situated within bodily experience. The goal of the study is to use this untapped cognitive resource to refine the construct of mindfulness by providing objective measurement and experimental control of these aspects of mindfulness. Location of the seat of attention, also described as self-location or egocenter , has been shown to have measurable impact on cognitive skill, emotional temperament, and self-construal, as well as social and moral attitudes. Our recent study has shown that the seat of attention can be volitionally self-regulated into various internal attentional stances that are stably associated with distinct patterns of cortical activation as measured with EEG. These results suggest that control of attentional stance should provide direct management of specific cognitive and emotional resources. Two specific global cognitive states associated with mindfulness are the positive emotional states of vipaśyanā and śamatha . The current study of the correlations of 11 attentional stances for with positive emotions reveal the association of vipaśyanā with a diffuse attentional stance centered on the abdomen, and śamatha with a focused attentional stance situated along the midline of the body. In addition to advancing denotation and differentiation of these two distinct elements of mindfulness, these results provide the opportunity for more efficient mindfulness training. Such opportunity would benefit anyone needing stress-reducing mindfulness training, and in particular, for underserved individuals in society who are most in need of stress management.
Lightness perception is a long-standing topic in research on human vision, but very few image-computable models of lightness have been formulated. Recent work in computer vision has used artifical neural networks and deep learning to estimate surface reflectance and other intrinsic image properties. Here we investigate whether such networks are useful as models of human lightness perception. We train a standard deep learning architecture on a novel image set that consists of simple geometric objects with a few different surface reflectance patterns. We find that the model performs well on this image set, generalizes well across small variations, and outperforms three other computational models. The network has partial lightness constancy, much like human observers, in that illumination changes have a systematic but moderate effect on its reflectance estimates. However, the network generalizes poorly beyond the type of images in its training set: it fails on a lightness matching task with unfamiliar stimuli, and does not account for several lightness illusions experienced by human observers.
Changes in the footballing world’s approach to technology and innovation contributed to the decision by the International Football Association Board (IFAB) to introduce Video Assistant Referees (VAR). The change meant that under strict protocols referees could use video replays to review decisions in the event of a “clear and obvious error” or a “serious missed incident”. This led to the need by Fédération Internationale de Football Association (FIFA) to develop methods for quality control of the VAR-systems, which was done in collaboration with RISE Research Institutes of Sweden AB. One of the important aspects is the video quality. The novelty of this study is that it has performed a user study specifically targeting video experts i.e., to measure the perceived quality of video professionals working with video production as their main occupation. An experiment was performed involving 25 video experts. In addition, six video quality models have been benchmarked against the user data and evaluated to show which of the models could provide the best predictions of perceived quality for this application. Video Quality Metric for variable frame delay (VQM_VFD) had the best performance for both formats, followed by Video Multimethod Assessment Fusion (VMAF) and VQM General model.
Accurate models of the electroretinogram are important both for understanding the multifold processes of light transduction to ecologically useful signals by the retina, but also its diagnostic capabilities for the identification of the array of retinal diseases. The present neuroanalytic model of the human rod ERG is elaborated from the same general principles as that of Hood & Birch (1992), but incorporates the more recent understanding of the early stages of ERG generation by Robson & Frishman (2014). As a result, it provides a significantly better match in six different waveform features of the canonical ERG flash intensity series than previous models of rod responses.
The population of low vision people increases continuously with the acceleration of aging society.As reported by the World Health Organization (WHO), most of this population is over the age of 50 years and 81% were not concerned by any visual problem before.A visual deficiency can dramatically affect the quality of life and challenge the preservation of a safe independent existence.This study presents a LED-based lighting approach to assist people facing an age-related visual impairment.The research procedure is based on a psychophysical experiment consisting in the ordering of standard color samples.Volunteers wearing low vision simulation goggles performed such an ordering under different illumination conditions produced by a 24-channel multispectral lighting system.A filtering technique using color rendering indices coupled with color measurements allowed to objectively determine the lighting conditions providing the best scores in terms of color discrimination.Experimental results demonstrated that white light obtained by a special mixing of three selected channels can improve the color perception of low vision people in comparison to white LEDs nowadays available on the market for general lighting.Even if additional studies are required to go further, these first results give hope for the design of smart lighting devices that might adapt to the visual needs of the visually impaired.
How is the cortical navigation network reorganized by the Likova Cognitive-Kinesthetic Navigation Training? We measured Granger-causal connectivity of the frontal-hippocampal-insular-retrosplenial-V1 network of cortical areas before and after this one-week training in the blind. Primarily top-down influences were seen during two tasks of drawing-from-memory (drawing complex maps and drawing the shortest path between designated map locations), with the dominant role being congruent influences from the egocentric insular to the allocentric spatial retrosplenial cortex and the amodal-spatial sketchpad of V1, with concomitant influences of the frontal cortex on these areas. After training, and during planning-from-memory of the best on-demand path, the hippocampus played a much stronger role, with the V1 sketchpad feeding information forward to the retrosplenial region. The inverse causal influences among these regions generally followed a recursive feedback model of the opposite pattern to a subset of congruent influences. Thus, this navigational network reorganized its pattern of causal influences with task demands and the navigation training, which produced marked enhancement of the navigational skills.
The critical flicker fusion (CFF) is the frequency of changes at which a temporally periodic light will begin to appear com-pletely steady to an observer. This value is affected by several visual factors, such as the luminance of the stimulus or its loca-tion on the retina. With new high dynamic range (HDR) displays, operating at higher luminance levels, and virtual reality (VR) displays, presenting at wide fields-of-view, the effective CFF may change significantly from values expected for traditional presen-tation. In this work we use a prototype HDR VR display capable of luminances up to 20,000cd/m 2 to gather a novel set of CFF measurements for never before examined levels of luminance, eccentricity, and size. Our data is useful to study the temporal be-havior of the visual system at high luminance levels, as well as setting useful thresholds for display engineering.
Spatial and temporal contrast sensitivity is typically measured using different stimuli. Gabor patterns are used to measure spatial contrast sensitivity and flickering discs are used for temporal contrast sensitivity. The data from both types of studies is difficult to compare as there is no well-established relationship between the sensitivity to disc and Gabor patterns. The goal of this work is to propose a model that can predict the contrast sensitivity of a disc using the more commonly available data and models for Gabors. To that end, we measured the contrast sensitivity for discs of different sizes, shown at different luminance levels, and for both achromatic and chromatic (isoluminant) contrast. We used this data to compare 6 different models, each of which tested a different hypothesis on the detection and integration mechanisms of disc contrast. The results indicate that multiple detectors contribute to the perception of disc stimuli, and each can be modelled either using an energy model, or the peak spatial frequency of the contrast sensitivity function.
Both natural scene statistics and ground surfaces have been shown to play important roles in visual perception, in particular, in the perception of distance. Yet, there have been surprisingly few studies looking at the natural statistics of distances to the ground, and the studies that have been done used a loose definition of ground. Additionally, perception studies investigating the role of the ground surface typically use artificial scenes containing perfectly flat ground surfaces with relatively few non-ground objects present, whereas ground surfaces in natural scenes are typically non-planar and have a large number of non-ground objects occluding the ground. Our study investigates the distance statistics of many natural scenes across three datasets, with the goal of separately analyzing the ground surface and non-ground objects. We used a recent filtering method to partition LiDAR-acquired 3D point clouds into ground points and non-ground points. We then examined the way in which distance distributions depend on distance, viewing elevation angle, and simulated viewing height. We found, first, that the distance distribution of ground points shares some similarities with that of a perfectly flat plane, namely with a sharp peak at a near distance that depends on viewing height, but also some differences. Second, we also found that the distribution of non-ground points is flatter and did not vary with viewing height. Third, we found that the proportion of non-ground points increases with viewing elevation angle. Our findings provide further insight into the statistical information available for distance perception in natural scenes, and suggest that studies of distance perception should consider a broader range of ground surfaces and object distributions than what has been used in the past in order to better reflect the statistics of natural scenes.
Chapter 27 Video for Change Tina Askanius, Tina AskaniusSearch for more papers by this author Tina Askanius, Tina AskaniusSearch for more papers by this author Book Editor(s):Karin Gwinn Wilkins, Karin Gwinn WilkinsSearch for more papers by this authorThomas Tufte, Thomas TufteSearch for more papers by this authorRafael Obregon, Rafael ObregonSearch for more papers by this author First published: 07 March 2014 https://doi.org/10.1002/9781118505328.ch27Citations: 12 AboutPDFPDF ToolsRequest permissionExport citationAdd to favoritesTrack citation ShareShareShare a linkShare onFacebookTwitterLinked InRedditWechat Summary This chapter argues that in order to understand contemporary forms of video activism, we need to extend our analytical scope beyond the confinements of the strategic work of social movement actors. It offers a schematic map of the various fields and disciplines in which video for social change has been given analytical attention pointing towards three key approaches and theoretical horizons in the literature. The chapter provides a brief history of video activism by drawing parallels between the early days of analog video cassettes during the so called "Portapak revolution" to the emergence of digital video and online sharing platforms such as YouTube. It provides an introduction to the notion of "radical online video" as a label for identifying and analyzing a broad range of video for change in contemporary online environments. Finally, the chapter broaches a discussion of some of the challenges facing video activism today. Citing Literature The Handbook of Development Communication and Social Change RelatedInformation
Olfaction is ingrained into the fabric of our daily lives and constitutes an integral part of our perceptual reality. Within this reality, there are crossmodal interactions and sensory expectations; understanding how olfaction interacts with other sensory modalities is crucial for augmenting interactive experiences with more advanced multisensorial capabilities. This knowledge will eventually lead to better designs, more engaging experiences, and enhancing the perceived quality of experience. Toward this end, the authors investigated a range of crossmodal correspondences between ten olfactory stimuli and different modalities (angularity of shapes, smoothness of texture, pleasantness, pitch, colors, musical genres, and emotional dimensions) using a sample of 68 observers. Consistent crossmodal correspondences were obtained in all cases, including our novel modality (the smoothness of texture). These associations are most likely mediated by both the knowledge of an odor’s identity and the underlying hedonic ratings: the knowledge of an odor’s identity plays a role when judging the emotional and musical dimensions but not for the angularity of shapes, smoothness of texture, perceived pleasantness, or pitch. Overall, hedonics was the most dominant mediator of crossmodal correspondences.
Last year at HVEI, I presented a computational model of lightness perception inspired by data from primate neurophysiology. That model first encodes local spatially-directed local contrast in the image, then integrates the resulting local contrast signals across space to compute lightness (Rudd, J Percept Imaging, 2020; HVEI Proceedings, 2021). Here I computer simulate the lightness model and generalize it to model color perception by including in the model local color contrast detectors that have properties similar to those of cortical “double-opponent” (DO) neurons. DO neurons make local spatial comparisons between the activities of L vs M and S vs (L + M) cones and half-wave rectify these local color contrast comparisons to produce psychophysical channels that encode, roughly, amounts of ‘red,’ ‘green’, ‘blue’, and ‘yellow.”
In this study, we developed a learning model to discriminate between skilled and novice users in visual inspection by visualizing their skills in the direction of gaze using an eye tracker. This model enabled us to analyze the difference in skill between skilled and novice users.
Remote operation and Augmented Telepresence are fields of interest for novel industrial applications in e.g., construction and mining. In this study, we report on an ongoing investigation of the Quality of Experience aspects of an Augmented Telepresence system for remote operation. The system can achieve view augmentation with selective content removal and Novel Perspective view generation. Two formal subjective studies have been performed with test participants scoring their experience while using the system with different levels of view augmentation. The participants also gave free-form feedback on the system and their experiences. The first experiment focused on the effects of in-view augmentations and interface distributions on wall patterns perception. The second one focused on the effects of augmentations on the depth and 3D environment understanding. The participants’ feedback from experiment 1 showed that the majority of participants preferred to use the original camera views and the Disocclusion Augmentation view instead of the Novel Perspective views. Moreover, the Disocclusion Augmentation, that was shown in combination with other views seemed beneficial. When the views were isolated in experiment 2, the impact of the Disocclusion Augmentation view was found to be lower than the Novel Perspective views.
Subjective evaluations are necessary to learn how expected viewers perceive the quality of a system. Traditionally, non-expert subjective tests are preferred rather than expert tests. In this study, we conducted subjective evaluation experiments for non-experts and experts on compressed 8K videos using the double stimulus impairment scale (DSIS) method and analyzed the experimental results expressed in terms of the mean opinion score (MOS), which is the average of individual scores. Furthermore, we investigated the differences between non-experts and experts by considering a new method in P.913 that estimates an improved MOS and a new experimental method using experts, called expert viewing protocol (EVP). Our contribution shows advantages of conducting expert subjective tests, such as EVP: expert tests allow to perform experiments with fewer subjects, to distinguish between original and distorted images, to determine a lower threshold for the image quality, to distribute scores in an appropriate range, and to constantly gain MOS values equal to improved MOS values.
Deep learning technology has made a significant improvement in image recognition performance.Unlike single-labeled training and inference, multi-labeled classification tasks hardly characterize individual label in the training of deep neural networks due to co-occurrence of the labels.Training data contains few samples of separated single label and the networks learn diverse compositions of labels from the data.Contextual bias caused by the co-occurrence of labels disturbs multi-label classification.We propose Identical and Disparate Feature Decomposition (INDeeD) from multi-label data that explicitly learn the characteristics of individual label.By training a backbone network combined with Identical and Disparate blocks on the instance pairs of partially common and contrastable labels, the network is generalized to decompose and learn individual label features.Proposed INDeeD scheme can be simply incorporated in any type of networks.We use ML-MNIST, ML-CIFAR-10, VOC-2007, and MS-COCO datasets to evaluate the performance of IN-DeeD showing improved mAP over baseline.
This exploratory study was designed to examine the effects of visual experience and specific texture parameters on both discriminative and aesthetic aspects of tactile perception. To this end, the authors conducted two experiments using a novel behavioral (ranking) approach in blind and (blindfolded) sighted individuals. Groups of congenitally blind, late blind, and (blindfolded) sighted participants made relative stimulus preference, aesthetic appreciation, and smoothness or softness judgment of two-dimensional (2D) or three-dimensional (3D) tactile surfaces through active touch. In both experiments, the aesthetic judgment was assessed on three affective dimensions, Relaxation, Hedonics, and Arousal, hypothesized to underlie visual aesthetics in a prior study. Results demonstrated that none of these behavioral judgments significantly varied as a function of visual experience in either experiment. However, irrespective of visual experience, significant differences were identified in all these behavioral judgments across the physical levels of smoothness or softness. In general, 2D smoothness or 3D softness discrimination was proportional to the level of physical smoothness or softness. Second, the smoother or softer tactile stimuli were preferred over the rougher or harder tactile stimuli. Third, the 3D affective structure of visual aesthetics appeared to be amodal and applicable to tactile aesthetics. However, analysis of the aesthetic profile across the affective dimensions revealed some striking differences between the forms of appreciation of smoothness and softness, uncovering unanticipated substructures in the nascent field of tactile aesthetics. While the physically softer 3D stimuli received higher ranks on all three affective dimensions, the physically smoother 2D stimuli received higher ranks on the Relaxation and Hedonics but lower ranks on the Arousal dimension. Moreover, the Relaxation and Hedonics ranks accurately overlapped with one another across all the physical levels of softness/hardness, but not across the physical levels of smoothness/roughness. These findings suggest that physical texture parameters not only affect basic tactile discrimination but differentially mediate tactile preferences, and aesthetic appreciation. The theoretical and practical implications of these novel findings are discussed.
Although chromatic adaptation eases us to adopt our vision to nuanced whites, viewing more than two substantially different white balances costs perceptual workload and appeals to poor quality control. This study proposed a method for evaluating the color tolerance of light modules using a uniformity analyzer focusing on the instrument panels in passenger cars, two premium line-up vehicles from Hyundai and Mercedes Benz. Using a luminance uniformity analyzer, we captured three main lighting regions in their instrument panels: clusters, steering wheel, and center console. Based on u’ and v’ values, we identified and compared the chromaticity coordinates of the white lighting components. The measurement-based judgment supports the manufacturer in achieving the quality objectively and consistently.
This paper introduces a new framework to predict visual attention of omnidirectional images. The key setup of our architecture is the simultaneous prediction of the saliency map and a corresponding scanpath for a given stimulus. The framework implements a fully encoder-decoder convolutional neural network augmented by an attention module to generate representative saliency maps. In addition, an auxiliary network is employed to generate probable viewport center fixation points through the SoftArgMax function. The latter allows to derive fixation points from feature maps. To take advantage of the scanpath prediction, an adaptive joint probability distribution model is then applied to construct the final unbiased saliency map by leveraging the encoder decoder-based saliency map and the scanpath-based saliency heatmap. The proposed framework was evaluated in terms of saliency and scanpath prediction, and the results were compared to state-of-the-art methods on Salient360! dataset. The results showed the relevance of our framework and the benefits of such architecture for further omnidirectional visual attention prediction tasks.