Humans appear to represent objects for intuitive physics with coarse, volumetric bodies” that smooth concavities - trading fine visual details for efficient physical predictions - yet their internal structure is largely unknown. Segmentation models, in contrast, optimize pixel-accurate masks that may misalign with such bodies. We ask whether and when these models nonetheless acquire human-like bodies. Using a time-to-collision (TTC) behavioral paradigm, we introduce a comparison pipeline and alignment metric, then vary model training time, size, and effective capacity via pruning. Across all manipulations, alignment with human behavior follows an inverse U-shaped curve: small/briefly trained/pruned models under-segment into blobs; large/fully trained models over-segment with boundary wiggles; and an intermediate ideal body granularity” best matches humans. This suggests human-like coarse bodies emerge from resource constraints rather than bespoke biases, and points to simple knobs - early checkpoints, modest architectures, light pruning - for eliciting physics-efficient representations. We situate these results within resource-rational accounts balancing recognition detail against physical affordances.
Learning grounded word meaning from natural experience requires resolving two ambiguities in infant-view recordings: when the named referent appears and where it is in a cluttered frame. In SAYCam-style data, caregiver speech is sparse and weakly synchronized with egocentric video, so single-frame contrastive pairing yields noisy positives in which the intended object is absent or entangled with distractors. We propose BabyMind, an object-first bias for child-view contrastive learning under sparse, noisy supervision. BabyMind extracts candidate object embeddings using an offline mask-based region interface, links candidates across a short utterance-centered window into lightweight object files via tracking, and aligns utterances to bags of object files with a prototype-space multiple-instance contrastive objective. Track-coherence and global-object agreement regularizers stabilize learning and transfer object-file structure into the global frame embedding used at evaluation. On SAYCam-S, BabyMind improves Labeled-S 15 forced-choice accuracy by +2.6 points over CVCL and yields consistent gains on in-vocabulary out-of-distribution benchmarks. Code is available at https://github.com/sathiiii/BabyMind.
Large multimodal models (LMMs) typically process visual inputs with uniform resolution across the entire field of view, leading to inefficiencies when non-critical image regions are processed as precisely as key areas. Inspired by the human visual system’s foveated approach, we apply a sampling method to leading architectures such as MDETR, BLIP2, InstructBLIP, LLaVA, and ViLT, and evaluate their performance with variable (foveated) resolution inputs. Results show that foveated sampling boosts accuracy in visual tasks like question answering and object detection under tight pixel budgets, improving performance by up to 2.7% on the GQA dataset, 2.1% on SEED-Bench, and 2.0% on VQAv2 compared to uniform sampling. Furthermore, we show that indiscriminate resolution increases yield diminishing returns, with models achieving up to 80% of their full capability using just 3% of the pixels, even on complex tasks. Foveated sampling prompts more human-like processing within models, such as neuronal selectivity and globally acting self-attention in vision transformers. This paper provides a foundational analysis of foveated sampling’s impact on existing models, suggesting that more efficient architectural adaptations, mimicking human visual processing, are a promising research venue for the community. Potential applications of our findings center low power minimal bandwidth devices (such as UAVs and edge devices), where compact and efficient vision is critical.
AI-based systems for visual scene understanding benefit from a large field of view (FOV). Multiple camera systems extend the FOV, but larger and higher-quality images strain acquisition, communication, and computing resources. Sub-sampling the FOV effectively addresses this challenge without compromising performance on complex tasks that require fine visual cues and contextual information. We demonstrate that a variable sampling scheme, inspired by human vision, outperforms uniform sampling in several visual question answering (VQA) tasks with a limited sample budget (3
Interferon Regulatory Factor 1 (IRF1) plays a pivotal role in interferon (IFN) signaling, yet its context-dependent regulatory functions remain incompletely understood. Here, we dissect the impact of IRF1 on gene regulation in HeLa cells, by targeted knockout (KO) or overexpression (OE) of IRF1. IRF1 KO did not impair interferon stimulated gene (ISG) expression regulation upon IFN-β stimulation, but partially diminished IFN-ψ induced gene regulation. IRF1 KO did show a homeostatic role in basal gene abundance, including increasing the abundance of some antiviral genes. RNA-seq analysis showed altered expression of both ISGs and immune signaling genes, implicating IRF1 as a dual regulator that fine-tunes gene abundance through both activation and repression. IRF1 OE induced potent antiviral protection in the absence of exogenous IFN, mediated by type I IFN secretion, particularly of IFN-α subtypes. This paracrine effect was confirmed by transcriptomics, cytokine profiling, and mass spectrometry, and was functional even in JAK1-deficient or Ruxolitinib-treated cells but not type I IFN receptor KO cells, suggesting the involvement of non-canonical signaling pathways. Hierarchical clustering of RNA-seq data revealed distinct IFN-independent gene clusters activated or repressed by IRF1, including pathways related to adaptive immunity and T cell function. Using protein-binding microarrays and predictive modeling, we mapped IRF1 binding across promoters and validated functional motifs in the IFIT2 gene promoter by a reporter assay. Our integrative approach establishes IRF1 as a central regulator of antiviral immunity, capable of shaping gene expression both through cytokine signaling and direct promoter binding.
Human physical reasoning relies on internal "body" representations - coarse, volumetric approximations that capture an object's extent and support intuitive predictions about motion and physics. While psychophysical evidence suggests humans use such coarse representations, their internal structure remains largely unknown. Here we test whether vision models trained for segmentation develop comparable representations. We adapt a psychophysical experiment conducted with 50 human participants to a semantic segmentation task and test a family of seven segmentation networks, varying in size. We find that smaller models naturally form human-like coarse body representations, whereas larger models tend toward overly detailed, fine-grain encodings. Our results demonstrate that coarse representations can emerge under limited computational resources, and that machine representations can provide a scalable path toward understanding the structure of physical reasoning in the brain.
The capability of Deep Neural Networks (DNNs) to recognize objects in orientations outside the distribution of the training data is not well understood. We present evidence that DNNs are capable of generalizing to objects in novel orientations by disseminating orientation-invariance obtained from familiar objects seen from many viewpoints. This capability strengthens when training the DNN with an increasing number of familiar objects, but only in orientations that involve 2D rotations of familiar orientations. We show that this dissemination is achieved via neurons tuned to common features between familiar and unfamiliar objects. These results implicate brain-like neural mechanisms for generalization.
Interferon regulatory factor 1 (IRF1) plays a pivotal role in interferon (IFN) signaling. Here, we dissect the impact of IRF1 on gene transcription regulation in HeLa cells, by targeted knockout (KO) or overexpression of IRF1. IRF1 KO partially diminished IFN-γ but not IFN-β induced gene regulation. IRF1 KO did show a homeostatic role in basal transcript abundance, including increasing the abundance of antiviral gene transcripts, apparently through increased expression of other IRF genes. IRF1 overexpression induced potent antiviral protection, which is mediated by secretion of type I IFN proteins, particularly of IFN-α subtypes, which expression is driven by IRF1. This paracrine effect was confirmed by transcriptomics, cytokine profiling, and mass spectrometry. Surprisingly, antiviral protection was observed also in JAK1 KO or ruxolitinib-treated cells but not in type I IFN receptor KO cells, suggesting the involvement of noncanonical signaling pathways. Hierarchical clustering of RNA-seq data revealed distinct IFN-independent gene clusters activated or repressed by IRF1, including pathways related to adaptive immunity and T cell function. Using protein-binding microarrays and predictive modeling, we generated an energy-normalized binding matrix for IRF1, enabling sequence-specific prediction of promoter-binding affinities beyond classical consensus motifs. This approach allows estimation of IRF1-binding potential across diverse genomic contexts as validated for the IFIT2 gene promoter by a reporter assay. Evaluating the biological significance of our study, we show that IRF1 abundance varies by 10,000-fold between cell lines, with positive correlations of IRF1 with the abundance of gene transcripts involved in antiviral and immune-driving activities.
Large multi-modal models (LMMs) show increasing performance in realistic visual tasks for images and, more recently, for videos. For example, given a video sequence, such models are able to describe in detail objects, the surroundings and dynamic actions. In this study, we explored the extent to which these models ground their semantic understanding in the actual visual input. Specifically, given sequences of hands interacting with objects, we asked models when and where the interaction begins or ends. For this purpose, we introduce a first of its kind, large-scale dataset with more than 20K annotated interactions on videos from the Something-Something-V2 dataset. 250 AMTurk human annotators labeled core interaction events, particularly when and where objects and agents become attached (`contact') or detached (`release'). We asked SoTA LMMs, including GPT, Gemini and Qwen to locate these events in short videos, each with a single event. The results show that while models reliably name target objects and identify actions, they exhibit a form of `shortcut learning' where semantic success masks a failure in physical grounding. Specifically, they consistently fail to identify the frame where the interaction begins or ends and poorly localize the physical event within the scene. This disconnect suggests that while LMMs excel at System 1 intuitive pattern recognition (naming the action and objects), they lack the System 2 cognitive foundations required to reason about physical primitives like `contact' and `release', hence truly ground dynamic scenes in physical reality.
Artificial intelligence (AI) scene understanding systems can benefit from utilizing a large visual field of view (FOV). Some existing systems already employ multiple cameras to extend their FOV, however, increasing image size and quality presents an overwhelming challenge to the acquisition and computing resources for such systems. An effective solution is to sub-sample the FOV, without impairing the model's performance on complex visual tasks. In this paper, we show that a variable sampling scheme, inspired by human vision, remarkably outperforms a uniform sampling scheme by 2% accuracy (65% vs. 63%) in the challenging task of scene visual question answering (VQA), under a limited samples budget (3% of the full resolution baseline). The improvement is achieved without any image scanning, and the variable resolution peaks at an arbitrarily chosen fixed image location. Our study also compared basic visual sub-tasks, in particular image classification and object detection. Comparing the variable and uniform models revealed differences in the representations learned by the different models which yield a consistently improved performance of the variable resolution models. We show that the variable sampling scheme allows the models to benefit in low resolution areas, by propagating information from the finer resolution areas, and at the same time higher resolution areas benefit from contextual information at lower resolution in the periphery. The results show the potential of the biologically-inspired image representation to improve the design of visual acquisition and processing models in future AI-based systems.
Diabetic Retinopathy (DR) is a common complication of diabetes that, in severe cases, can result in blindness. Accurate clinical treatment is imperative to prevent these cases and relies considerably on an exact diagnosis of the various symptoms of DR. We aim to advance DR diagnosis by providing a practical tool to automatically classify Optical Coherence Tomography (OCT) scans for DR and to identify and localize DR-related morphological features within the scans. Our system obtains raw OCT input and only sparse clinical annotations at the volume level, which can be obtained automatically from routine electronic medical records. We developed a novel neural network architecture, OCT-Transformer, that obtains state-of-the-art classification results compared to previous models and does so with limited training data. We base our architecture on an attention mechanism and show this to be the driving factor for the boost in performance. We additionally use our model to locate pixels within the input scans that explain its classification.
IntroductionChronic lymphocytic leukemia (CLL) is characterized by an aberrant cytokine network that can support tumor growth by triggering janus kinase (JAK)/STAT pathways. Targeting cytokine-signaling should then be a rational therapeutic strategy but the JAK inhibitor ruxolitinib failed to control and seemingly accelerated the disease in clinical trials.MethodsThe effect of ruxolitinib on primary human CLL cells was studied in vitro and in vivo.ResultsRuxolitinib increased phosphorylation of IRAK4, an important toll-like receptor (TLR)- signaling intermediate, in circulating CLL cells in vitro. It also enhanced p38 and NFKB1 phosphorylation while lowering STAT3 phosphorylation in CLL cells activated with TLR-7/8 agonists and IL-2. Among the cytokines made by activated CLL cells, high levels of IL-10 contributed strongly to STAT3 phosphorylation and inhibited TLR7 activity. Ruxolitinib limited TLR-mediated IL10 transcription and markedly reduced IL-10 production in vitro. It also decreased blood levels of IL-10 while increasing TNFα along with phospho-p38 expression and gene sets associated with TLR-activation in CLL cells in vivo. The bruton's tyrosine kinase inhibitor ibrutinib decreased IL-10 production in vitro but, in contrast to ruxolitinib, blocked initial IL10 transcription induced by TLR-signaling in vitro, decreased TNFα production, and deactivates CLL cells in vivo.DiscussionThese findings suggest the possible benefits of inhibiting growth factors with JAK inhibitors in CLL are outweighed by negative effects on potential tumor suppressors such as IL-10 that allow unrestrained activation of NFκB by drivers such as TLRs. Specific inhibition of growth-promoting cytokines with blocking antibodies or infusing suppressive cytokines like IL-10 might be better strategies to manipulate cytokines in CLL.
In an ideal human-robot collaboration, autonomous robots work side-by-side with humans in a joint workspace, often performing complementary tasks to the humans. A robotic ability to infer human intention and goals directly from human behavior will facilitate the collaboration and maximize its efficiency. In this paper, we focus on inferring which object the human wants picked up next, based on what the human is looking at, by visually following the human gaze and head orientation. We develop a coordination protocol for a team of aerial robots to extract effective human head and gaze cues. The aerial robots are controlled to navigate around the human and collect data that improves the detection of the human's gaze and hence the intended object to be picked up. The effectiveness of the approach is shown using simulations in AirSim, a photo-realistic simulator.
Visual scene understanding involves processing and integration from different levels of visual tasks, including recognition of objects, actions and interactions. Here we study the dynamics of scene understanding over time. In particular, we study the time trajectory of scene interpretation, by controlling the exposure time with perceptual masking. 140 MTurk participants were instructed to provide a detailed free-recall description to 14 stimuli images portraying various interactions between animate agents (humans and pets) and other agents and objects. They were instructed to report the type of objects and agents in the image with their properties and inter-relations. For each image, subjects were assigned to one of seven exposure conditions: 50, 75, 100, 125, 200, 500 and 2000ms followed by a mask. A fixation cross at the center of the image frame appeared prior to image display. Participants had 15 minutes for task completion. Evaluation of the subjects’ responses was conducted by 4 scorers, who followed a detailed analysis protocol, which minimized subjective judgements. Preliminary results indicate consistent trends in the time evolution of scene perception: (i) human agents are reported earlier than objects and global scene description, even when objects appear at the center of fixation (e.g. ‘two men’ before ‘a park bench’); (ii) actions are reported earlier than the acted upon objects (e.g. ‘drinking’ before ‘cup’); (iii) for human agents, the number of agents is reported early, followed by age, and gender is reported on the average later (e.g. ‘two people’, before ‘two kids’, and then ‘two boys’). These findings are interesting from a modeling perspective since they do not fit the common scene understanding paradigm in computer vision, where objects are first detected and only then their inter-relations are processed. We will consider scene perception schemes that are more consistent with human dynamics of scene perception than current approaches.
Delivering medication to the lungs via nebulization of pharmaceuticals is a noninvasive and efficient therapy route, particularly for respiratory diseases. The recent worldwide severe acute respiratory syndrome coronavirus type 2 (SARS-CoV-2) pandemic urges the development of such therapies as an effective alternative to vaccines. The main difficulties in using inhalation therapy are the development of effective medicine and methods to stabilize the biological molecules and transfer them to the lungs efficiently following nebulization. We have developed a high-affinity angiotensin-converting enzyme 2 (ACE2) receptor-binding domain (RBD-62) that can be used as a medication to inhibit infection with SARS-CoV-2 and its variants. In this study, we established a nebulization protocol for drug delivery by inhalation using two commercial vibrating mesh (VM) nebulizers (Aerogen Solo and PARI eFlow) that generate similar mist size distribution in a size range that allows efficient deposition in the small respiratory airway. In a series of experiments, we show the high activity of RBD-62, interferon-α2 (IFN-α2), and other proteins following nebulization. The addition of gelatin significantly stabilizes the proteins and enhances the fractions of active proteins after nebulization, minimizing the medication dosage. Furthermore, hamster inhalation experiments verified the feasibility of the protocol in pulmonary drug delivery. In short, the gelatin-modified RBD-62 formulation in coordination with VM nebulizer can be used as a therapy to cure SARS-CoV-2.
Gaze understanding-a suggested precursor for understanding others' intentions-requires recovery of gaze direction from the observed person's head and eye position. This challenging computation is naturally acquired at infancy without explicit external guidance, but can it be learned later if vision is extremely poor throughout early childhood? We addressed this question by studying gaze following in Ethiopian patients with early bilateral congenital cataracts diagnosed and treated by us only at late childhood. This sight restoration provided a unique opportunity to directly address basic issues on the roles of "nature" and "nurture" in development, as it caused a selective perturbation to the natural process, eliminating some gaze-direction cues while leaving others still available. Following surgery, the patients' visual acuity typically improved substantially, allowing discrimination of pupil position in the eye. Yet, the patients failed to show eye gaze-following effects and fixated less than controls on the eyes-two spontaneous behaviors typically seen in controls. Our model for unsupervised learning of gaze direction explains how head-based gaze following can develop under severe image blur, resembling preoperative conditions. It also suggests why, despite acquiring sufficient resolution to extract eye position, automatic eye gaze following is not established after surgery due to lack of detailed early visual experience. We suggest that visual skills acquired in infancy in an unsupervised manner will be difficult or impossible to acquire when internal guidance is no longer available, even when sufficient image resolution for the task is restored. This creates fundamental barriers to spontaneous vision recovery following prolonged deprivation in early age.
[This corrects the article DOI: 10.1371/journal.ppat.1009800.].