Visual shape perception is central to many everyday tasks, from object recognition to grasping and handling tools.1-10 Yet how shape is encoded in the visual system remains poorly understood. Here, we probed shape representations using visual aftereffects-perceptual distortions that occur following extended exposure to a stimulus.11-17 Such effects are thought to be caused by adaptation in neural populations that encode both simple, low-level stimulus characteristics17-20 and more abstract, high-level object features.21-23 To tease these two contributions apart, we used machine -learning methods to synthesize novel shapes in a multidimensional shape space, derived from a large database of natural shapes.24 Stimuli were carefully selected such that low-level and high-level adaptation models made distinct predictions about the shapes that observers would perceive following adaptation. We found that adaptation along vector trajectories in the high-level shape space predicted shape aftereffects better than simple low-level processes. Our findings reveal the central role of high-level statistical features in the visual representation of shape. The findings also hint that human vision is attuned to the distribution of shapes experienced in the natural environment.
To grasp an object successfully, we must select appropriate contact regions for our hands on the surface of the object. However, identifying such regions is challenging. This paper describes a workflow to estimate the contact regions from marker-based tracking data. Participants grasp real objects, while we track the 3D position of both the objects and the hand, including the fingers' joints. We first determine the joint Euler angles from a selection of tracked markers positioned on the back of the hand. Then, we use state-of-the-art hand mesh reconstruction algorithms to generate a mesh model of the participant's hand in the current pose and the 3D position.Using objects that were either 3D printed or 3D scanned-and are, thus, available as both real objects and mesh data-allows the hand and object meshes to be co -registered. In turn, this allows the estimation of approximate contact regions by calculating the intersections between the hand mesh and the co-registered 3D object mesh. The method may be used to estimate where and how humans grasp objects under a variety of conditions. Therefore, the method could be of interest to researchers studying visual and haptic perception, motor control, human-computer interaction in virtual and augmented reality, and robotics.
Shape perception is essential for numerous everyday behaviors from object recognition to grasping and handling objects. Yet how the brain encodes shape remains poorly understood. Here, we probed shape representations using visual aftereffects—perceptual distortions that occur following extended exposure to a stimulus—to resolve a long-standing debate about shape encoding. We implemented contrasting low-level and high-level computational models of neural adaptation, which made precise and distinct predictions about the illusory shape distortions the observers experience following adaptation. Directly pitting the predictions of the two models against one another revealed that the perceptual distortions are driven by high-level shape attributes derived from the statistics of natural shapes. Our findings suggest that the diverse shape attributes thought to underlie shape encoding (e.g., curvature distributions, ‘skeletons’, aspect ratio) are the result of a visual system that learns to encode natural shape geometries based on observing many objects.
Virtual Reality (VR) is a powerful tool for studying perception, allowing researchers to present stimuli whose appearance and physical behaviour deviate from reality. Here, we used VR to investigate visually-guided grasping in mixed reality environments. By aligning real and virtual environments, participants could view virtual versions of objects while touching their real counterparts at the same location. We asked participants to grasp the objects, while we tracked markers attached to their hands and to the objects using 8 Qualisys Miqus M5 motion-tracking cameras. Additionally, we recorded temporally aligned videos using 6 Qualisys Miqus Video cameras. Marker positions were streamed into a virtual environment containing a realistic replica of the real object, which matched the motion of the real object. To verify the alignment between real and virtual scenes, we rendered the virtual environment from the same viewpoints as the video cameras. For visualization purposes, we overlaid the virtual object renderings onto the video recordings, thus generating a new, mixed reality depiction of the scene. To quantify the alignment overlap, we computed the intersection over union of the object segmented out of video and rendered scenes, achieving an average alignment of 84±3 %. We also quantified the alignment accuracy as the distance in image space between real and rendered tracked object markers. We achieved an average accuracy of 6.5±0.5 pixels, which translated to a misalignment of 1.7±0.1 mm and 5.1±0.3 minutes of visual angle for an average viewing distance of 115 cm, validating the effectiveness of our alignment procedure. By rendering the aligned virtual objects with different textures, we can vary the appearance of the objects’ material composition. This allows us to perform novel explorations investigating how the visuomotor system adapts to conflicts between the visual appearance and physical properties of graspable objects.
Humans use vision to select where and how to grasp objects. How this is accomplished remains unexplored in many respects. For example, previous research has often focused on the visual selection of digit contact points (e.g., during precision-grip grasping). Yet, the whole surface of our fingers and palm may come into contact with objects during natural, unconstrained grasping. Here, we investigated how finger contact areas varied when participants grasped objects of different materials. In a first experiment, we asked 6 participants to freely grasp and lift wood (50g) and brass (700g) bars (2.5 cm grip width) while wearing a data glove. We recorded the hand pose (joint angles of each finger) selected by participants to grasp the objects. Participants employed distinct hand postures: they selected a precision grip when grasping light wooden objects, but employed multiple digits when grasping heavy brass objects (p<.05). In a second experiment, we asked 5 participants to grasp a single bar (2.5 cm grip width) with either a precision grip or a multi-digit grasp. We coated the stimulus with thermochromic paint—which changes colour at body temperature—allowing us to estimate the hand contact regions on the objects. As expected, total contact area increased when participants used multi-digit grasps (p<.01). Interestingly however, the average contact area across fingers decreased with multi-digit grasps (p<.05), possibly because contact surfaces were distributed across multiple fingers, thus requiring smaller grip forces per finger to lift the object. Therefore, when visually selecting how to grasp heavier objects, participants increased the number of digits they employed to increase the total grip surface area while maintaining relatively low grip forces. Our findings demonstrate how humans modify their grasping behaviours to achieve comfortable and effective grasps as perceived shape and material properties vary.
The discovery of mental rotation was one of the most significant landmarks in experimental psychology, leading to the ongoing assumption that to visually compare objects from different three-dimensional viewpoints, we use explicit internal simulations of object rotations, to 'mentally adjust' one object until it matches the other1. These rotations are thought to be performed on three-dimensional representations of the object, by literal analogy to physical rotations. In particular, it is thought that an imagined object is continuously adjusted at a constant three-dimensional angular rotation rate from its initial orientation to the final orientation through all intervening viewpoints2. While qualitative theories have tried to account for this phenomenon3, to date there has been no explicit, image-computable model of the underlying processes. As a result, there is no quantitative account of why some object viewpoints appear more similar to one another than others when the three-dimensional angular difference between them is the same4,5. We reasoned that the specific pattern of non-uniformities in the perception of viewpoints can reveal the visual computations underlying mental rotation. We therefore compared human viewpoint perception with a model based on the kind of two-dimensional 'optical flow' computations that are thought to underlie motion perception in biological vision6, finding that the model reproduces the specific errors that participants make. This suggests that mental rotation involves simulating the two-dimensional retinal image change that would occur when rotating objects. When we compare objects, we do not do so in a distal three-dimensional representation as previously assumed, but by measuring how much the proximal stimulus would change if we watched the object rotate, capturing perspectival appearance changes7.
Our world is filled with complex objects. To make perceptual inferences, decisions, and action plans, humans must often judge the three-dimensional shape and orientation of objects they encounter. Despite the importance of view-invariant object perception, the computations underlying the 3D rotational discriminability of objects remain unclear. Here, we aimed to devise a metric that predicts viewpoint similarity for rotated novel objects, based on the projected position shifts of surface points as the object rotates. We created six 3D Shepard-Metzler-like objects of varying complexity, and rendered each from nineteen equally-spaced viewpoints rotated independently around the horizontal or vertical axes. For each object and viewpoint, participants were shown a standard pair of views of the object, along with a test pair, and used a mouse to rotate one of the test pair objects so that the test views were the same perceived rotational distance apart as the standard pair. If the reported distance between the test pair was smaller than the standard pair, this viewpoint was taken to be more rotationally discriminable, and vice versa. For n = 29 participants, our findings reveal substantial and consistent variations in perceived viewpoint similarity across different object orientations. We then developed a metric that predicted these variations, based on the sum of the projected displacement vectors of visible surface points as the object incrementally rotated from one viewpoint to the next. This metric predicted human rotational discriminability results for both horizontal and vertical rotations with striking accuracy [R = 0.60]. This suggests that to judge viewpoint similarity, observers mentally rotate the object and estimate the projected position shifts of points in the imagined view relative to the seen view. This metric provides a computational correlate linking theories of viewpoint discrimination and mental rotation.
When we grasp an object, we must select appropriate contact points on the surface of the object to accomplish a comfortable and stable grasp. Previous studies have identified multiple key factors that influence the selection of grasp points during precision grip grasps. The natural grip axis is one such constraint that human participants are typically unwilling to oppose. In this study, we aimed to challenge the preference for the natural grip axis by selecting objects for which grasps aligned with the natural grip axis would result in differing levels of torque. We tracked participants’ index finger and thumb while they lifted L-shaped cuboids that were presented in various orientations (-45°, 0°, 45°, 90° or 180° counter-clockwise) with respect to the natural grasp axis. The cuboids were either made purely out of wood (light weight) or brass and wood (high weight) such that the objects’ weight distribution would produce high torques if the grasp axis and surface normals at the contact points with the object were misaligned. Analyses of the grasp patterns showed that participants had strong perceptual intuitions regarding the torque produced by different grasps. Specifically, for the brass object participants consistently counter-balanced torque by selecting grasp points such that the thumb and index finger placements required low forces to lift and manipulate the objects. These results suggest that humans are able to make physically accurate predictions about the behavior of complex objects.
In order to grasp an object successfully, we must select appropriate contact regions for our hands on the surface of the object. However, identifying such regions is challenging. Here, we describe a workflow to estimate contact regions from marker-based tracking data. Participants grasp real objects, while we track the 3D position of both the objects and the hand including the fingers’ joints. We first determine joint Euler angles from a selection of tracked markers positioned on the back of the hand. Then, we use state-of-the-art hand mesh reconstruction algorithms to generate a mesh model of the participant’s hand in the current pose and 3D position. Using objects that were either 3D printed, or 3D scanned—and are thus available as both real objects and mesh data—allows us to co-register the hand and object meshes. In turn, this allows us to estimate approximate contact regions by calculating intersections between the hand mesh and the co-registered 3D object mesh. The method may be used to estimate where and how humans grasp objects under a variety of conditions. Therefore, the method could be of interest to researchers studying visual and haptic perception, motor control, human-computer interaction in virtual and augmented reality, and robotics. SUMMARY When we grasp an object, multiple regions of the fingers and hand typically make contact with the object’s surface. Reconstructing such contact regions is challenging. Here, we present a method for approximately estimating contact regions, by combining marker-based motion capture with existing deep learning-based hand mesh reconstruction.
A widely-used psychophysical tool for inferring visual mechanisms is adaptation. Perceptual distortions known as aftereffects arise following extended visual exposure to a stimulus, including complex patterns and shapes. Some researchers have argued that shape aftereffects reveal adaptation of mechanisms sensitive to global shape properties, while others propose they can be explained by localized adaptation to simpler properties such as tilt. Here, we investigate methods to tease these hypotheses apart. Most previous works used simpler and/or familiar forms (e.g., radial frequency patterns, geometric shapes, faces). Here we use complex but naturalistic shapes synthesized by Generative Adversarial Networks (GANs) trained on >25,000 animal silhouettes. Drawing samples from the GAN’s latent space allows us to synthesize novel 2D shapes that transition smoothly between one another, and are complex yet systematically related. Observers adapted to individual shapes from this generative shape space, then judged the appearance of nearby shapes in a two-alternative forced-choice task. Their responses demonstrated robust and systematic perceptual distortions of shape using such stimuli. Indeed, a given shape could predictably be made to look like specific other shapes depending on which adaptor stimulus was used. To tease apart the relative role of local vs. global adaptation, we simulated the effects of two variants of tilt and positional aftereffects: the local model assumes aftereffects exaggerate differences between test and adaptor within localized image regions, while the global model assumes such distortions exaggerate differences between ‘corresponding’ parts of shapes, after completing high-level inferences regarding part-correspondence. Initial findings show that positional adaptation has a larger contribution than tilt adaptation to novel-shape aftereffects. We then show how to use the generative shape networks to synthesize tailored shape sets that can tease apart the predictions of local vs global models of adaptation. Our findings provide new methods for probing and modeling adaptation to complex visual stimuli.
A 3D object seen from different viewpoints can elicit vastly different retinal images. Differences between views depend on object geometry and initial pose, rendering relative pose estimation computationally challenging. Still, humans can easily judge object identity across views and estimate the relative pose between them. Here, we sought to measure how accurately observers can estimate pose similarity for 3D objects, and how these judgements are influenced by object geometry and changes in its retinal projection. We first mapped out human judgements of relative viewpoints using a multi-arrangement task. On each trial, observers (N=16) were asked to spatially arrange 31 views of one of three novel or three familiar 3D objects by viewpoint similarity. The resulting arrangements broadly matched ground-truth viewpoint differences with deviations that were consistent across observers (i.e. representational similarity analysis revealed correlations with ground-truth below the noise ceiling across objects). We implemented several candidate computational models, based on 2D image features or object geometry, and evaluated their ability to predict human judgements. Strategies using 2D features failed to account for human data. However, a metric based on the union and intersection of visible surface area across views (‘Surface IoU’) predicted human judgments on par with ground-truth. In order to maximise our power to differentiate between candidate strategies, we selected triads of viewpoints for individual objects over which pairs of models strongly disagreed (e.g. where similar changes in viewing angle produced very different changes in image pixels). We presented these triads in a two-alternative forced-choice experiment in which participants judged which of two views appeared closest to a target view. Across triad judgements and free arrangements, we gathered a rich dataset of human viewpoint perception for many objects and viewpoints that allows us to evaluate the ability of computational models to predict human strategies for judging relative viewpoint.
bioRxiv - the preprint server for biology, operated by Cold Spring Harbor Laboratory, a research and educational institution.
Shape is a defining feature of objects, and human observers can effortlessly compare shapes to determine how similar they are. Yet, to date, no image-computable model can predict how visually similar or different shapes appear. Such a model would be an invaluable tool for neuroscientists and could provide insights into computations underlying human shape perception. To address this need, we developed a model ('ShapeComp'), based on over 100 shape features (e.g., area, compactness, Fourier descriptors). When trained to capture the variance in a database of >25,000 animal silhouettes, ShapeComp accurately predicts human shape similarity judgments between pairs of shapes without fitting any parameters to human data. To test the model, we created carefully selected arrays of complex novel shapes using a Generative Adversarial Network trained on the animal silhouettes, which we presented to observers in a wide range of tasks. Our findings show that incorporating multiple ShapeComp dimensions facilitates the prediction of human shape similarity across a small number of shapes, and also captures much of the variance in the multiple arrangements of many shapes. ShapeComp outperforms both conventional pixel-based metrics and state-of-the-art convolutional neural networks, and can also be used to generate perceptually uniform stimulus sets, making it a powerful tool for investigating shape and object representations in the human brain.
When people judge the temporal order (TOJ task) of two tactile stimuli at the two hands while their hands are crossed, performance is much worse than with uncrossed hands [1]. This crossed-hands deficit is widely considered to indicate interferences of external spatial coordinates with body-centered coordinates in the localization of touch [2]. Similar deficits have also been observed when people are only about to move their hands towards a crossed position [3]-[5], suggesting a predictive update of external spatial coordinates. Here, we extend the investigation of the dynamics of external coordinates during hand movement. Participants performed a TOJ task while they executed an uncrossing or a crossing movement, and during presentation of the TOJ stimuli the present posture of the hands was crossed, uncrossed or in-between. Present, future and past crossed-hands postures decreased performance in the TOJ task, suggesting that the update of external spatial coordinates of touch includes both predictive processes and processes that preserve the recent past. In addition, our data corroborate the flip model of crossed-hands deficits [1], and suggest that more pronounced deficits come along with higher time requirements to resolve interferences.
When people judge the temporal order (TOJ task) of two tactile stimuli at the two hands while their hands are crossed, performance is much worse than with uncrossed hands [1]. This crossed-hands deficit is widely considered to indicate interferences of external spatial coordinates with body-centered coordinates in the localization of touch [2]. Similar deficits have also been observed when people are only about to move their hands towards a crossed position [3-5], suggesting a predictive update of external spatial coordinates. Here, we extend the investigation of the dynamics of external coordinates during hand movement. Participants performed a TOJ task while they executed an uncrossing or a crossing movement, and during presentation of the TOJ stimuli the present posture of the hands was crossed, uncrossed or in-between. Present, future and past crossed-hands postures decreased performance in the TOJ task, suggesting that the update of external spatial coordinates of touch includes both predictive processes and processes that preserve the recent past. In addition, our data corroborate the flip model of crossed-hands deficits [1], and suggest that more pronounced deficits come along with higher time requirements to resolve interferences.