There is increasing interest in real-time brain-computer interfaces (BCIs) for the passive monitoring of human cognitive state, including cognitive workload. Too often, however, effective BCIs based on machine learning techniques may function as “black boxes” that are difficult to analyze or interpret. In an effort toward more interpretable BCIs, we studied a family of N-back working memory tasks using a machine learning model, Gaussian Process Regression (GPR), which was both powerful and amenable to analysis. Participants performed the N-back task with three stimulus variants, auditory-verbal, visual-spatial, and visual-numeric, each at three working memory loads. GPR models were trained and tested on EEG data from all three task variants combined, in an effort to identify a model that could be predictive of mental workload demand regardless of stimulus modality. To provide a comparison for GPR performance, a model was additionally trained using multiple linear regression (MLR). The GPR model was effective when trained on individual participant EEG data, resulting in an average standardized mean squared error (sMSE) between true and predicted N-back levels of 0.44. In comparison, the MLR model using the same data resulted in an average sMSE of 0.55. We additionally demonstrate how GPR can be used to identify which EEG features are relevant for prediction of cognitive workload in an individual participant. A fraction of EEG features accounted for the majority of the model’s predictive power; using only the top 25% of features performed nearly as well as using 100% of features. Subsets of features identified by linear models (ANOVA) were not as efficient as subsets identified by GPR. This raises the possibility of BCIs that require fewer model features while capturing all of the information needed to achieve high predictive accuracy.
Current machine learning algorithms identify statistical regularities in complex data sets and are regularly used across a range of application domains, but they lack the robustness and generalizability associated with human learning. If machine learning techniques could enable computers to learn from fewer examples, transfer knowledge between tasks, and adapt to changing contexts and environments, the results would have very broad scientific and societal impacts. Increased processing and memory resources have enabled larger, more capable learning models, but there is growing recognition that even greater computing resources would not be sufficient to yield algorithms capable of learning from a few examples and generalizing beyond initial training sets. This paper presents perspectives on feature selection, representation schemes and interpretability, transfer learning, continuous learning, and learning and adaptation in time-varying contexts and environments, five key areas that are essential for advancing machine learning capabilities. Appropriate learning tasks that require these capabilities can demonstrate the strengths of novel machine learning approaches that could address these challenges.
Purpose. We measured the separate influences of monocular and binocular cues on planning and executing arm movements aimed at picking up an object. Method. Subjects viewed a large coin (7 cm diameter, 0.95 cm thickness) that was suspended in a virtual environment binocularly under central fixation. A robot arm positioned an unseen physical coin so that it was co-aligned with the virtual coin in the workspace. The coin's slant (orientation in depth) varied randomly across trials, and we introduced cue conflicts of 0–10° between the monocular cues (contour and texture) and binocular information (disparity) that defined the coin's orientation either at the beginning of each trial or following a mask that appeared upon movement initiation to prevent subjects from becoming aware of changes. Subjects reached for and picked up the real coin by the edges and from the side using a precision grip, and we optically tracked the positions of each subject's thumb and index finger throughout each trial. We used the vector between these fingers as an indicator of the subject's estimate of the coin's orientation. No visual feedback about the positions of the fingers was available until the fingertips made contact with the coin. Results. Subjects were very accurate about how they positioned their fingers to pick up the coin since the mean orientation of the vector between the fingers typically matched the orientation of the coin with a standard deviation of 3–4° on cue-consistent, unperturbed trials. As in a previous study involving object placement (Greenwald et al., 2005), binocular information dominated the execution phase of the movements (the normalized binocular weight was 0.6–0.7). Conclusion. The cue weights we measured during this grasping task were similar to those found for object placement, suggesting that our previous findings generalize to other visuomotor tasks.
Purpose. We quantified the information provided by the outputs of biologically-plausible mechanisms sensitive to orientation disparity to determine the usefulness of this cue for estimating surface slant. Method. Using a modified disparity energy model tuned to orientation disparity, we computed responses to simulated surfaces textured with random mean-vertical or broadband noise and slanted away from the viewer about the horizontal axis. Based on the responses of our cell population, we estimated the slant of each surface using a Naïve Bayes decoder that compared the mean response levels with expected activity based on distributions for surfaces at a wide range of slants, and we compared this performance to estimates of the separability of different slants using linear discriminant analysis with gradient descent. To verify that the information used in the estimates were based on orientation disparity and not purely local orientation, we repeated the Naïve Bayes analysis using an energy model with monocular input. Results. The Naïve Bayes estimator and the linear discriminant analysis produced similar results with standard deviations ranging from 2–6 degrees, which is within the range of normal acuity for slant from stereo. The estimates from the monocular control were uninformative for the broadband noise textures and produced standard deviations that were more than 4 times larger for the oriented noise textures, indicating that the information was carried by orientation disparity. Conclusion. Given the performance of our model, orientation disparity appears to be a plausible source of information for estimating 3D surface orientation.
We assessed the usefulness of stereopsis across the visual field by quantifying how retinal eccentricity and distance from the horopter affect humans' relative dependence on monocular and binocular cues about 3D orientation. The reliabilities of monocular and binocular cues both decline with eccentricity, but the reliability of binocular information decreases more rapidly. Binocular cue reliability also declines with increasing distance from the horopter, whereas the reliability of monocular cues is virtually unaffected. We measured how subjects integrated these cues to orient their hands when grasping oriented discs at different eccentricities and distances from the horopter. Subjects relied increasingly less on binocular disparity as targets' retinal eccentricity and distance from the horopter increased. The measured cue influences were consistent with what would be predicted from the relative cue reliabilities at the various target locations. Our results showed that relative reliability affects how cues influence motor control and that stereopsis is of limited use in the periphery and away from the horopter because monocular cues are more reliable in these regions.
Visual cue integration strategies are known to depend on cue reliability and how rapidly the visual system processes incoming information. We investigated whether these strategies also depend on differences in the information demands for different natural tasks. Using two common goal-oriented tasks, prehension and object placement, we determined whether monocular and binocular information influence estimates of three-dimensional (3D) orientation differently depending on task demands. Both tasks rely on accurate 3D orientation estimates, but 3D position is potentially more important for grasping. Subjects placed an object on or picked up a disc in a virtual environment. On some trials, the monocular cues (aspect ratio and texture compression) and binocular cues (e.g., binocular disparity) suggested slightly different 3D orientations for the disc; these conflicts either were present upon initial stimulus presentation or were introduced after movement initiation, which allowed us to quantify how information from the cues accumulated over time. We analyzed the time-varying orientations of subjects' fingers in the grasping task and those of the object in the object placement task to quantify how different visual cues influenced motor control. In the first experiment, different subjects performed each task, and those performing the grasping task relied on binocular information more when orienting their hands than those performing the object placement task. When subjects in the second experiment performed both tasks in interleaved sessions, binocular cues were still more influential during grasping than object placement, and the different cue integration strategies observed for each task in isolation were maintained. In both experiments, the temporal analyses showed that subjects processed binocular information faster than monocular information, but task demands did not affect the time course of cue processing. How one uses visual cues for motor control depends on the task being performed, although how quickly the information is processed appears to be task invariant.
Orientation disparity, the difference in orientation that results when a texture element on a slanted surface is projected to the two eyes, has been proposed as a binocular cue for 3D orientation. Since orientation disparity is confounded with position disparity, neither behavioral nor neurophysiological experiments have successfully isolated its contribution to slant estimates or established whether the visual system uses it. Using a modified disparity energy model, we simulated a population of binocular visual cortical neurons tuned to orientation disparity and measured the amount of Fisher information contained in the activity patterns. We evaluated the potential contribution of orientation disparity to 3D orientation estimation and delimited the stimulus conditions under which it is a reliable cue. Our results suggest that orientation disparity is an efficient source of information about 3D orientation and that it is plausible that the visual system could have mechanisms that are sensitive to it. Although orientation disparity is neither necessary nor sufficient for estimating slant, it appears that it could be useful when combined with estimates from position disparity gradients and monocular perspective cues.
The visual system continuously integrates multiple sensory cues to help plan and control everyday motor tasks. We quantified how subjects integrated monocular cues (contour and texture) and binocular cues (disparity and vergence) about 3D surface orientation throughout an object placement task and found that binocular cues contributed more to online control than planning. A temporal analysis of corrective responses to stimulus perturbations revealed that the visuomotor system processes binocular cues faster than monocular cues. This suggests that binocular cues dominated online control because they were available sooner, thus affecting a larger proportion of the movement. This was consistent with our finding that the relative influence of binocular information was higher for short-duration movements than long-duration movements. A motor control model that optimally integrates cues with different delays accounts for our findings and shows that cue integration for motor control depends in part on the time course of cue processing.
Purpose. We measured the relative contributions of monocular and binocular depth cues to planning and online control of a reaching task. Method. Subjects placed a cylinder onto a circular target surface that varied randomly in slant (orientation in depth) between trials while viewing a binocular image of the surface in a 3D virtual environment. ±5 conflicts between monocular (e.g. texture) and binocular cues were either (1) present in the stimulus at initial presentation (unperturbed trials), affecting both planning and online control, or (2) added after movement initiation (perturbation trials), affecting only the online control component of the movement. A robot arm aligned a real target surface with the virtual surface image to provide haptic feedback at the end of each trial. An optical tracking device measured the position and orientation of the cylinder throughout the movement, and the moving, oriented cylinder was also rendered in the virtual environment. On all trials, the screen flashed repeatedly for 167 ms upon movement initiation to mask the motion signal created by cue changes added in the perturbation trials. Results. In the perturbation trials, binocular cues influenced final cylinder orientation more than did monocular cues. For unperturbed trials, binocular cues remained more influential but to a lesser degree. A temporal analysis of cue weights revealed that the influence of binocular information on subjects' movements accrued faster than that of monocular information. This was consistent with the finding that final cylinder orientation was more strongly correlated with the binocular cue for fast than for slow movements. Conclusions. Humans appear to give more weight to binocular cues for online control than for planning. This can be explained by our finding that binocular cues are processed more quickly than monocular cues, effectively giving them more influence over the course of what are relatively short movements.