Purpose:We aim to assess the perceptual tasks in which convolutional neural networks (CNNs) might be better tools than commonly used linear model observers (LMOs) to evaluate medical image quality. Approach:We compared the LMOs (channelized Hotelling [CHO] and frequency convolution channels observers [FCO]) and CNN detection accuracies for tasks with a few possible signal locations (location known exactly) and for the search for mass and microcalcification signals embedded in 2D/3D breast tomosynthesis phantoms. We also compared the LMOs and CNN accuracies to those of radiologists in the search tasks. We analyzed radiologists' eye position to assess whether they fixate longer at locations considered suspicious by the LMOs or those by the CNN. Results:LMOs resulted in similar detection accuracies [area under the receiver operating characteristic curve (AUC)] to the CNN for tasks with up to 100 signal locations but lower accuracies in the search task for microcalcification and mass 3D images. Radiologists' AUC was significantly higher ( p < 1 e - 4 ) than that of LMOs for the microcalcification 2D search (CHO, FCO) and 3D mass search ( p < 0.05 , CHO) but was not higher than the CNN's AUC. For both signal types, radiologists fixated longer on the locations of the highest response scores of the CNN than those of the LMOs but only reached statistical significance for the mass (masses: p = 0.009 versus CHO and p = 0.004 versus FCO). Conclusion:We show that CNNs are a more suitable model observer for search tasks. Like radiologists but not traditional LMOs, CNNs can discount false positives arising from anatomical backgrounds.
Efforts to restore vision via neural implants have outpaced the ability to predict what users will perceive, leaving patients and clinicians without reliable tools for surgical planning or device selection. To bridge this critical gap, we introduce a computational virtual patient (CVP) pipeline that integrates anatomically grounded phosphene simulation with task-optimized deep neural networks (DNNs) to forecast patient perceptual capabilities across diverse prosthetic designs and tasks. We evaluate performance across six visual tasks, six electrode configurations, and two artificial vision models, positioning our CVP approach as a scalable pre-implantation method. Several chosen tasks align with the Functional Low-Vision Observer Rated Assessment (FLORA), revealing correspondence between model-predicted difficulty and real-world patient outcomes. Further, DNNs exhibited strong correspondence with psychophysical data collected from normally sighted subjects viewing phosphene simulations, capturing both overall task difficulty and performance variation across implant configurations. While performance was generally aligned, DNNs sometimes diverged from humans in which specific stimuli were misclassified, reflecting differences in underlying decision strategies between artificial agents and human observers. The findings position CVP as a scientific tool for probing perception under prosthetic vision, an engine to inform device development, and a clinically relevant framework for pre-surgical forecasting. ### Competing Interest Statement The authors have declared no competing interest. UC Noyce Initiative
Human eye movements while identifying a face, searching for targets, and executing motor actions are directed to regions contributing to task accuracy. However, what humans look at and do when free-viewing a scene without a specific task is not well understood. We show that observers’ free-viewing fixations are similar to fixations of observers instructed to describe the scenes and dissimilar to fixations of observers counting objects or searching for specific objects. Small visual alterations to images that change a scene’s understanding but not the most salient or its meaning map alter where humans most frequently fixate. Free viewing fixations are more frequently directed to objects critical to the understanding of a scene (objects that, when erased from the scene, maximally alter the scene’s description) rather than the most salient, most meaningfully judged scene region (meaning map), or the object perceived to be grasped or gazed at. By having observers describe the scene while maintaining fixation on objects relevant or irrelevant to scene understanding, we show that eye movements during free viewing are functionally important to understand scenes accurately. The theoretical framework also explains the high frequency of fixations on people in scenes because, when people are erased, scene descriptions are maximally altered. Thus, we conclude that an important default task during free viewing for the human brain is understanding scenes, reflected by frequent eye movements toward people and objects that maximize accurate scene understanding.
We introduce INTERLACE, a novel framework that prunes redundant layers in VLMs while maintaining performance through sample-efficient finetuning. Existing layer pruning methods lead to significant performance drop when applied to VLMs. Instead, we analyze triplets of consecutive layers to identify local redundancy, removing the most redundant of the first two layers, finetune the remaining layer to compensate for the lost capacity, and freeze the third layer to serve as a stable anchor during finetuning. We found that this interleaved finetune-freeze design enables rapid convergence with minimal data after pruning. By finetuning only a subset of layers on just 1
Although models exist that predict human response times (RTs) in tasks such as target search and visual discrimination, the development of image-computable predictors for scene understanding time remains an open challenge. Recent advances in vision-language models (VLMs), which can generate scene descriptions for arbitrary images, combined with the availability of quantitative metrics for comparing linguistic descriptions, offer a new opportunity to model human scene understanding. We hypothesize that the primary bottleneck in human scene understanding and the driving source of variability in response times across scenes is the interaction between the foveated nature of the human visual system and the spatial distribution of task-relevant visual information within an image. Based on this assumption, we propose a novel image-computable model that integrates foveated vision with VLMs to produce a spatially resolved map of scene understanding as a function of fixation location (Foveated Scene Understanding Map, or F-SUM), along with an aggregate F-SUM score. This metric correlates with average (N=17) human RTs (r=0.47) and number of saccades (r=0.51) required to comprehend a scene (across 277 scenes). The F-SUM score also correlates with average (N=16) human description accuracy (r=-0.56) in time-limited presentations. These correlations significantly exceed those of standard image-based metrics such as clutter, visual complexity, and scene ambiguity based on language entropy. Together, our work introduces a new image-computable metric for predicting human response times in scene understanding and demonstrates the importance of foveated visual processing in shaping comprehension difficulty.
The ability to quickly and precisely follow another person's gaze reflects critical evolutionary mechanisms underlying social interactions, such as attention modulation and the prediction of others' future actions. Recent studies show that observers use another person's gaze direction and peripheral scene information to make anticipatory saccades toward the gaze goal. However, it remains unclear how these eye movements are influenced by complex features of natural scenes, such as a foveal gazer, multiple peripheral gaze goals, and the relative distance between gazer and goal. We presented dynamic stimuli (videos) of real-world scenes with or without a gazer shifting their head to gaze at other individuals (gaze goals). Participants were instructed to search for a specific target individual in the videos while their eye movements were recorded. We measured the accuracy of the first saccade in locating the gaze goal. First, we found that the absence of a foveal gazer significantly increased saccade error, but only when the goal was at least approximately 9 degrees of visual angle from the initial fixation. First saccade amplitude and onset latency were higher in the gazer-present condition. Second, when there were multiple potential gaze goals in the periphery, the first saccade was directed to the individual closer to the initial fixation (gazer) location. Finally, the presence of multiple peripheral gaze goals shortened saccade latencies and increased the frequency of anticipatory saccades made before the gazer completed their head movement. These findings extend our understanding of gaze following in complex, naturalistic scenes and inform theories of attention and real-world decision-making.
Visual complexity prediction is a fundamental problem in computer vision with applications in image compression, retrieval, and classification. Understanding what makes humans perceive an image as complex is also a long-standing question in cognitive science. Recent approaches have leveraged multimodal models that combine visual and linguistic representations, but it remains unclear whether language information is necessary for this task. We propose DReX (DINO-ResNet Fusion), a vision-only model that fuses self-supervised and convolutional representations through a learnable attention mechanism to predict image complexity. Our architecture integrates multi-scale hierarchical features from ResNet-50 with semantically rich representations from DINOv3 ViT-S/16, enabling the model to capture both low-level texture patterns and high-level semantic structure. DReX achieves state-of-the-art performance on the IC9600 benchmark (Pearson r = 0.9581), surpassing previous methods–including those trained on multimodal image-text data–while using approximately 21.5x fewer learnable parameters. Furthermore, DReX generalizes robustly across multiple datasets and metrics, achieving superior results on Pearson and Spearman correlation, Root Mean Square Error (RMSE), and Mean Absolute Error (MAE). Ablation and attention analyses confirm that DReX leverages complementary cues from both backbones, with the DINOv3 [CLS] token enhancing sensitivity to visual complexity. Our findings suggest that visual features alone can be sufficient for human-aligned complexity prediction and that, when properly fused, self-supervised transformers and supervised deep convolutional neural networks offer complementary and synergistic benefits for this task.
Humans consistently land their first saccade to a face a preferred fixation location (PFL). Humans also typically process faces as wholes, as evidenced by perceptual effects such as the composite face effect (CFE). However, not known is whether an individual's tendency to process faces as wholes varies with their gaze patterns on the face. Here, we investigated variation of the CFE with the PFL. We compared the strength of the CFE for two groups of observers who were screened to have their PFLs either higher up, closer to the eyes, or lower on the face, closer to the tip of the nose. During the task, observers maintained their gaze at either their own group's mean PFL or at the other group's mean PFL. We found that the top half of the face elicits a stronger CFE than the bottom half. Further, the strength of the CFE was modulated by the distance of the PFL from the eyes, such that individuals with a PFL closer to the eyes had stronger CFE than those with a PFL closer to the mouth. Finally, the top-half CFE for both upper-lookers and lower-lookers was abolished when they fixated at a non-preferred location on the face. Our findings show that the CFE relies on internal face representations shaped by the long-term use of a consistent oculomotor strategy to view faces.
Double reading of screening mammograms, a feature of many breast cancer screening programs, is impacted by interactions between the two image readers. In this work, we describe how the bivariate binormal (BVBN) model, originally developed for statistical analysis of reader studies, can be used to analyze double reading of screening mammograms. The model posits two bivariate normal distributions that describe the distribution of latent decision variables of the two readers for cancer and non-cancer cases. The BVBN allows for the estimation of correlation coefficients between the decision variables of two readers, independent of performance and the threshold for recall. We contend that these correlation coefficients are a useful way to characterize interactions between readers because they characterize associations at the level of the perceptual response in a way that is consistent with Signal Detection Theory. We describe the BVBN model and show how parameters can be estimated from count data under an assumed multinomial distribution. The analysis presented focuses on two aspects of the BVBN model. For implementation using binary data, an equal-variance assumption on latent decision variables is required. Otherwise, the model is over-parameterized. We characterize and discuss the consequence of this assumption. We also show how disagreement rates, an alternative measure of reader interactions, suffer from base-rate effects making them more difficult to interpret than the correlation coefficients of the BVBN model.
Purpose: Radiologists are tasked with visually scrutinizing large amounts of data produced by 3D volumetric imaging modalities. Small signals can go unnoticed during the 3d search because they are hard to detect in the visual periphery. Recent advances in machine learning and computer vision have led to effective computer-aided detection (CADe) support systems with the potential to mitigate perceptual errors. Approach: Sixteen non-expert observers searched through digital breast tomosynthesis (DBT) phantoms and single cross-sectional slices of the DBT phantoms. The 3D/2D searches occurred with and without a convolutional neural network (CNN)-based CADe support system. The model provided observers with bounding boxes superimposed on the image stimuli while they looked for a small microcalcification signal and a large mass signal. Eye gaze positions were recorded and correlated with changes in the area under the ROC curve (AUC). Results: The CNN-CADe improved the 3D search for the small microcalcification signal (delta AUC = 0.098, p = 0.0002) and the 2D search for the large mass signal (delta AUC = 0.076, p = 0.002). The CNN-CADe benefit in 3D for the small signal was markedly greater than in 2D (delta delta AUC = 0.066, p = 0.035). Analysis of individual differences suggests that those who explored the least with eye movements benefited the most from the CNN-CADe (r = -0.528, p = 0.036). However, for the large signal, the 2D benefit was not significantly greater than the 3D benefit (delta delta AUC = 0.033, p = 0.133). Conclusion: The CNN-CADe brings unique performance benefits to the 3D (vs. 2D) search of small signals by reducing errors caused by the under-exploration of the volumetric data.
The rapid progress in Multimodal Large Language Models (MLLMs) has significantly advanced their ability to process and understand complex visual and textual information. However, the integration of multiple images and extensive textual contexts remains a challenge due to the inherent limitation of the models' capacity to handle long input sequences efficiently. In this paper, we introduce SEEKER, a multimodal large language model designed to tackle this issue. SEEKER aims to optimize the compact encoding of long text by compressing the text sequence into the visual pixel space via images, enabling the model to handle long text within a fixed token-length budget efficiently. Our empirical experiments on six long-context multimodal tasks demonstrate that SEEKER can leverage fewer image tokens to convey the same amount of textual information compared with the OCR-based approach, and is more efficient in understanding long-form multimodal input and generating long-form textual output, outperforming all existing proprietary and open-source MLLMs by large margins.