
This paper introduces a novel framework for generating high-quality images from "visual sentences" extracted from video sequences. By combining a lightweight autoregressive model with a Vector Quantized Generative Adversarial Network (VQGAN), our approach achieves a favorable trade-off between computational efficiency and image fidelity. Unlike conventional methods that require substantial resources, the proposed framework efficiently captures sequential patterns in partially annotated frames and synthesizes coherent, contextually accurate images. Empirical results demonstrate that our method not only attains state-of-the-art performance on various benchmarks but also reduces inference overhead, making it well-suited for real-time and resource-constrained environments. Furthermore, we explore its applicability to medical image analysis, showcasing robust denoising, brightness adjustment, and segmentation capabilities. Overall, our contributions highlight an effective balance between performance and efficiency, paving the way for scalable and adaptive image generation across diverse multimedia domains.
Recent advancements in AI models, like ChatGPT, are structured to retain user interactions, which could inadvertently include sensitive healthcare data. In the healthcare field, particularly when radiologists use AI-driven diagnostic tools hosted on online platforms, there is a risk that medical imaging data may be repurposed for future AI training without explicit consent, spotlighting critical privacy and intellectual property concerns around healthcare data usage. Addressing these privacy challenges, a novel approach known as Unlearnable Examples (UEs) has been introduced, aiming to make data unlearnable to deep learning models. A prominent method within this area, called Unlearnable Clustering (UC), has shown improved UE performance with larger batch sizes but was previously limited by computational resources (e.g., a single workstation). To push the boundaries of UE performance with theoretically unlimited resources, we scaled up UC learning across various datasets using Distributed Data Parallel (DDP) training on the Summit supercomputer. Our goal was to examine UE efficacy at high-performance computing (HPC) levels to prevent unauthorized learning and enhance data security, particularly exploring the impact of batch size on UE's unlearnability. Utilizing the robust computational capabilities of the Summit, extensive experiments were conducted on diverse datasets such as Pets, MedMNist, Flowers, and Flowers102. Our findings reveal that both overly large and overly small batch sizes can lead to performance instability and affect accuracy. However, the relationship between batch size and unlearnability varied across datasets, highlighting the necessity for tailored batch size strategies to achieve optimal data protection. The use of Summit's high-performance GPUs, along with the efficiency of the DDP framework, facilitated rapid updates of model parameters and consistent training across nodes. Our results underscore the critical role of selecting appropriate batch sizes based on the specific characteristics of each dataset to prevent learning and ensure data security in deep learning applications. The source code is publicly available at https://github.com/hrlblab/UE_HPC.
Whole Slide Image (WSI) analysis plays a crucial role in modern digital pathology, enabling large-scale feature extraction from tissue samples. However, traditional feature extraction pipelines based on tools like CellProfiler often involve lengthy workflows, requiring WSI segmentation into patches, feature extraction at the patch level, and subsequent mapping back to the original WSI. To address these challenges, we present PySpatial, a high-speed pathomics toolkit specifically designed for WSI-level analysis. PySpatial streamlines the conventional pipeline by directly operating on computational regions of interest, reducing redundant processing steps. Utilizing rtree-based spatial indexing and matrix-based computation, PySpatial efficiently maps and processes computational regions, significantly accelerating feature extraction while maintaining high accuracy. Our experiments on two datasets-Perivascular Epithelioid Cell (PEC) and data from the Kidney Precision Medicine Project (KPMP)-demonstrate substantial performance improvements. For smaller and sparse objects in PEC datasets, PySpatial achieves nearly a 10-fold speedup compared to standard CellProfiler pipelines. For larger objects, such as glomeruli and arteries in KPMP datasets, PySpatial achieves a 2-fold speedup. These results highlight PySpatial's potential to handle large-scale WSI analysis with enhanced efficiency and accuracy, paving the way for broader applications in digital pathology.
Multi-modal learning adeptly integrates visual and textual data, but its application to histopathology image and text analysis remains challenging, particularly with large, high-resolution images like gigapixel Whole Slide Images (WSIs). Current methods typically rely on manual region labeling or multi-stage learning to assemble local representations (e.g., patch-level) into global features (e.g., slide-level). However, there is no effective way to integrate multi-scale image representations with text data in a seamless end-to-end process. In this study, we introduce Multi-Level Text-Guided Representation End-to-End Learning (mTREE). This novel text-guided approach effectively captures multi-scale WSI representations by utilizing information from accompanying textual pathology information. mTREE innovatively combines - the localization of key areas (global-to-local) and the development of a WSI-level image-text representation (local-to-global) - into a unified, end-to-end learning framework. In this model, textual information serves a dual purpose: firstly, functioning as an attention map to accurately identify key areas, and secondly, acting as a conduit for integrating textual features into the comprehensive representation of the image. Our study demonstrates the effectiveness of mTREE through quantitative analyses in two image-related tasks: classification and survival prediction, showcasing its remarkable superiority over baselines.
The segment anything model (SAM) was released as a foundation model for image segmentation. The promptable segmentation model was trained by over 1 billion masks on 11M licensed and privacy-respecting images. The model supports zero-shot image segmentation with various segmentation prompts (e.g., points, boxes, masks). It makes the SAM attractive for medical image analysis, especially for digital pathology where the training data are rare. In this study, we evaluate the zero-shot segmentation performance of SAM model on representative segmentation tasks on whole slide imaging (WSI), including (1) tumor segmentation, (2) non-tumor tissue segmentation, (3) cell nuclei segmentation. Core Results: The results suggest that the zero-shot SAM model achieves remarkable segmentation performance for large connected objects. However, it does not consistently achieve satisfying performance for dense instance object segmentation, even with 20 prompts (clicks/boxes) on each image. We also summarized the identified limitations for digital pathology: (1) image resolution, (2) multiple scales, (3) prompt selection, and (4) model fine-tuning. In the future, the few-shot fine-tuning with images from downstream pathological segmentation tasks might help the model to achieve better performance in dense object segmentation.
Avoiding person-to-person collisions is critical for visual field loss patients. Any intervention claiming to improve the safety of such patients should empirically demonstrate its efficacy. To design a VR mobility testing platform presenting multiple pedestrians, a distinction between colliding and non-colliding pedestrians must be clearly defined. We measured nine normally sighted subjects' collision envelopes (CE; an egocentric boundary distinguishing collision and non-collision) and found it changes based on the approaching pedestrian's bearing angle and speed. For person-to-person collision events for the VR mobility testing platform, non-colliding pedestrians should not evade the CE.
Scientific user facilities present a unique set of challenges for image processing due to the large volume of data generated from experiments and simulations. Furthermore, developing and implementing algorithms for real-time processing and analysis while correcting for any artifacts or distortions in images remains a complex task, given the computational requirements of the processing algorithms. In a collaborative effort across multiple Department of Energy national laboratories, the "MLExchange" project is focused on addressing these challenges. MLExchange is a Machine Learning framework deploying interactive web interfaces to enhance and accelerate data analysis. The platform allows users to easily upload, visualize, label, and train networks. The resulting models can be deployed on real data while both results and models could be shared with the scientists. The MLExchange web-based application for image segmentation allows for training, testing, and evaluating multiple machine learning models on hand-labeled tomography data. This environment provides users with an intuitive interface for segmenting images using a variety of machine learning algorithms and deep-learning neural networks. Additionally, these tools have the potential to overcome limitations in traditional image segmentation techniques, particularly for complex and low-contrast images.
How is the cortical navigation network reorganized by the Likova Cognitive-Kinesthetic Navigation Training? We measured Granger-causal connectivity of the frontal-hippocampal-insular-retrosplenial-V1 network of cortical areas before and after this one-week training in the blind. Primarily top-down influences were seen during two tasks of drawing-from-memory (drawing complex maps and drawing the shortest path between designated map locations), with the dominant role being congruent influences from the egocentric insular to the allocentric spatial retrosplenial cortex and the amodal-spatial sketchpad of V1, with concomitant influences of the frontal cortex on these areas. After training, and during planning-from-memory of the best on-demand path, the hippocampus played a much stronger role, with the V1 sketchpad feeding information forward to the retrosplenial region. The inverse causal influences among these regions generally followed a recursive feedback model of the opposite pattern to a subset of congruent influences. Thus, this navigational network reorganized its pattern of causal influences with task demands and the navigation training, which produced marked enhancement of the navigational skills.
Radiologists and pathologists frequently make highly consequential perceptual decisions. For example, visually searching for a tumor and recognizing whether it is malignant can have a life-changing impact on a patient. Unfortunately, all human perceivers-even radiologists-have perceptual biases. Because human perceivers (medical doctors) will, for the foreseeable future, be the final judges of whether a tumor is malignant, understanding and mitigating human perceptual biases is important. While there has been research on perceptual biases in medical image perception tasks, the stimuli used for these studies were highly artificial and often critiqued. Realistic stimuli have not been used because it has not been possible to generate or control them for psychophysical experiments. Here, we propose to use Generative Adversarial Networks (GAN) to create vivid and realistic medical image stimuli that can be used in psychophysical and computer vision studies of medical image perception. Our model can generate tumor-like stimuli with specified shapes and realistic textures in a controlled manner. Various experiments showed the authenticity of our GAN-generated stimuli and the controllability of our model.
Mass Spectrometry Imaging (MSI), using traditional rectilinear scanning, takes hours to days for high spatial resolution acquisitions. Given that most pixels within a sample's field of view are often neither relevant to underlying biological structures nor chemically informative, MSI presents as a prime candidate for integration with sparse and dynamic sampling algorithms. During a scan, stochastic models determine which locations probabilistically contain information critical to the generation of low-error reconstructions. Decreasing the number of required physical measurements thereby minimizes overall acquisition times. A Deep Learning Approach for Dynamic Sampling (DLADS), utilizing a Convolutional Neural Network (CNN) and encapsulating molecular mass intensity distributions within a third dimension, demonstrates a simulated 70% throughput improvement for Nanospray Desorption Electrospray Ionization (nano-DESI) MSI tissues. Evaluations are conducted between DLADS, a Supervised Learning Approach for Dynamic Sampling, with Least-Squares regression (SLADS-LS), and a Multi-Layer Perceptron (MLP) network (SLADS-Net). When compared with SLADS-LS, limited to a single m/z channel, as well as multichannel SLADS-LS and SLADS-Net, DLADS respectively improves regression performance by 36.7%, 7.0%, and 6.2%, resulting in gains to reconstruction quality of 6.0%, 2.1%, and 3.4% for acquisition of targeted m/z.
We describe the development of a multipurpose haptic stimulus delivery and spatiomotor recording system with tactile map-overlays for electronic processing This innovative multipurpose spatiomotor capture system will serve a wide range of functions in the training and behavioral assessment of spatial memory and precise motor control for blindness rehabilitation, both for STEM learning and for navigation training and map reading. Capacitive coupling through the map-overlays to the touch-tablet screen below them allows precise recording i) of hand movements during haptic exploration of tactile raised-line images on one tablet and ii) of line-drawing trajectories on the other, for analysis of navigational errors, speed, time elapsed, etc. Thus, this system will provide for the first time in an integrated and automated manner quantitative assessments of the whole 'perception-cognition-action' loop - from non-visual exploration strategies, spatial memory, precise spatiomotor control and coordination, drawing performance, and navigation capabilities, as well as of haptic and movement planning and control. The accuracy of memory encoding, in particular, can be assessed by the memory-drawing operation of the capture system. Importantly, this system allows for both remote and in-person operation. Although the focus is on visually impaired populations, the system is designed to equally serve training and assessments in the normally sighted as well.
In order to better understand how our visual system processes information, we must understand the underlying brain connectivity architecture, and how it can get reorganized under visual deprivation. The full extent to which visual development and visual loss affect connectivity is not well known. To investigate the effect of the onset of blindness on structural connectivity both at the whole-brain voxel-wise level and at the level of all major white-matter tracts, we applied two complementary Diffusion-Tension Imaging (DTI) methods, TBSS and AFQ. Diffusion-weighted brain images were collected from three groups of participants: congenitally blind (CB), acquired blind (AB), and fully sighted controls. The differences between these groups were evaluated on a voxel-wise scale with Tract-Based Spatial Statistics (TBSS) method, and on larger-scale with Automated Fiber Quantification (AFQ), a method that allows for between-group comparisons at the level of the major fiber tracts. TBSS revealed that both blind groups tended to have higher FA than sighted controls in the central structures of the brain. AFQ revealed that, where the three groups differed, congenitally blind participants tended to be more similar to sighted controls than to those participants who had acquired blindness later in life. These differences were specifically manifested in the left uncinated fasciculus, the right corticospinal fasciculus, and the left superior longitudinal fasciculus, areas broadly associated with a range of higher-level cognitive systems.
To address the longstanding questions of whether the blind-from-birth have an innate face-schema, what plasticity mechanisms underlie non-visual face learning, and whether there are interhemispheric face processing differences in face processing in the blind, we used a unique non-visual drawing-based training in congenitally blind (CB), late-blind (LB) and blindfolded-sighted (BF) groups of adults. This Cognitive-Kinesthetic Drawing approach previously developed by Likova (e.g., 2010, 2012, 2013) enabled us to rapidly train and study training-driven neuroplasticity in both the blind and sighted groups. The five-day two-hour training taught participants to haptically explore, recognize, memorize raised-line images, and draw them free-hand from memory, in detail, including the fine facial characteristics of the face stimuli. Such drawings represent an externalization of the formed memory. Functional MRI was run before and after the training. Tactile-face perception activated the occipito-temporal cortex in all groups. However, the training led to a strong, predominantly left-hemispheric reorganization in the two blind groups, in contrast to right-hemispheric in blindfolded-sighted, i.e., the post-training response-change was stronger in the left hemisphere in the blind, but in the right in the blindfolded. This is the first study to discover interhemispheric differences in non-visual face processing. Remarkably, for face perception this learning-based change was positive in the CB and BF groups, but negative in the LB-group. Both the lateralization and inversed-sign learning effects were specific to face perception, but absent for the control nonface categories of small objects and houses. The unexpected inversed-sign training effect in CB vs LB suggests different stages of brain plasticity in the ventral pathway specific to the face category. Importantly, the fact that only after a very few days of our training, the totally-blind-from-birth CB manifested a very good (haptic) face perception, and even developed strong empathy to the explored faces, implies a preexisting face schema that can be "unmasked" and "tuned up" by a proper learning procedure. The Likova Cognitive-Kinesthetic Training is a powerful tool for driving brain plasticity, and providing deeper insights into non-visual learning, including emergence of perceptual categories. A rebound learning model and a neuro-Bayesian economy principle are proposed to explain the multidimensional learning effects. The results provide new insights into the Nature-vs-Nurture interplay in rapid brain plasticity and neurorehabilitation.
Rapidly evolving technologies like data analysis, smartphone and web-based applications, and the Internet of things have been increasingly used for healthy living, fitness and well-being. These technologies are being utilized by various research studies to reduce obesity. This paper demonstrates design and development of a dataflow protocol that integrates several applications. After registration of a user, activity, nutrition and other lifestyle data from participants are retrieved in a centralized cloud dedicated for health promotion. In addition, users are provided accounts in an e-Learning environment from which learning outcomes can be retrieved. Using the proposed system, health promotion campaigners have the ability to provide feedback to the participants using a dedicated messaging system. Participants authorize the system to use their activity data for the program participation. The implemented system and servicing protocol minimize personnel overhead of large-scale health promotion campaigns and are scalable to assist automated interventions, from automated data retrieval to automated messaging feedback. This paper describes end-to -end workflow of the proposed system. The case study tests are carried with Fitbit Flex2 activity trackers, Withings Scale, Verizon Android-based tablets, Moodle learning management system, and Articulate RISE for learning content development.
Understanding perception and aesthetic appeal of arts and environmental objects, what is appreciated, liked, or preferred, and why, is of prime importance for improving the functional capacity of the blind and visually impaired and the ergonomic design for their environment, which however so far, has been examined only in sighted individuals. This paper provides a general overview of the first experimental study of tactile aesthetics as a function of visual experience and level of visual deprivation, using both behavioral and brain imaging techniques. We investigated how blind people perceive 3D tactile objects, how they characterize them, and whether the tactile perception, and tactile shape preference (liking or disliking) and tactile aesthetic appreciation (judging tactile qualities of an object, such as pleasantness, comfortableness etc.) of 3D tactile objects can be affected by the level of visual experience. The study employed innovative behavioral measures, such as new forms of aesthetic preference-appreciation and perceptual discrimination questionnaires, in combination with advanced functional Magnetic Resonance Imaging (fMRI) techniques, and compared congenitally blind, late-onset blind and blindfolded (sighted) participants. Behavioral results demonstrated that both blind and blindfolded-sighted participants assessed curved or rounded 3D tactile objects as significantly more pleasing than sharp 3D tactile objects, and symmetric 3D tactile objects as significantly more pleasing than asymmetric 3D tactile objects. However, as compared to the sighted, blind people showed better skills in tactile discrimination as demonstrated by accuracy and speed of discrimination. Functional MRI results demonstrated that there was a large overlap and characteristic differences in the aesthetic appreciation brain networks in the blind and the sighted. As demonstrated both populations commonly recruited the somatosensory and motor areas of the brain, but with stronger activations in the blind as compared to the sighted. Secondly, sighted people recruited more frontal regions whereas blind people, in particular, the congenitally blind, paradoxically recruited more 'visual' areas of the brain. These differences were more pronounced between the sighted and the congenitally blind rather than between the sighted and the late-onset blind, indicating the key influence of the onset time of visual deprivation. Understanding of the underlying brain mechanisms should have a wide range of important implications for a generalized cross-sensory theory and practice in the rapidly evolving field of neuroaesthetics, as well as for 'cutting-edge' rehabilitation technologies for the blind and the visually impaired.
A supervised learning approach for dynamic sampling (SLADS) was developed to reduce X-ray exposure prior to data collection in protein structure determination. Implementation of this algorithm allowed reduction of the X-ray dose to the central core of the crystal by up to 20-fold compared to current raster scanning approaches. This dose reduction corresponds directly to a reduction on X-ray damage to the protein crystals prior to data collection for structure determination. Implementation at a beamline at Argonne National Laboratory suggests promise for the use of the SLADS approach to aid in the analysis of X-ray labile crystals. The potential benefits match a growing need for improvements in automated approaches for microcrystal positioning.
Conceptual knowledge allows us to comprehend the multisensory stimulation impinging on our senses. Its representation in the anterior temporal lobe is a subject of considerable debate, with the "enigmatic" temporal pole (TP) being at the center of that debate. The controversial models of the organization of knowledge representation in TP range from unilateral to fully unified bilateral representational systems. To address the multitude of mutually exclusive options, we developed a novel cross-modal approach in a multifactorial brain imaging study of the blind, manipulating the modality (verbal vs pictorial) of both the reception source (reading text/verbal vs images/pictorial) and the expression (writing text/verbal vs drawing/pictorial) of conceptual knowledge. Furthermore, we also varied the level of familiarity. This study is the first to investigate the functional organization of (amodal) conceptual knowledge in TP in the blind, as well as, the first study of drawing based on the conceptual knowledge from memory of sentences delivered through Braille reading. Through this paradigm, we were able to functionally identify two novel subdivisions of the temporal pole - the TPa, at the apex, and the TPdm - dorso-medially. Their response characteristics revealed a complex interplay of non-visual specializations within the temporal pole, with a diversity of excitatory/inhibitory inversions as a function of hemisphere, task-domain and familiarity, which motivate an expanded neurocognitive analysis of conceptual knowledge. The interplay of inter-hemispheric specializations found here accounts for the variety of seemingly conflicting models in previous research for conceptual knowledge representation, reconciling them through the set of factors we have investigated: the two main knowledge domains (verbal and pictorial/sensory-motor) and the two main knowledge processing modes (receptive and expressive), including the level of familiarity as a modifier. Furthermore, the interplay of these factors allowed us to also reveal for the first time a system of complementary symmetries, asymmetries and unexpected anti-symmetries in the TP organization. Thus, taken together these results constitute a unifying explanation of the conflicting models in previous research on conceptual knowledge representation.
Contrast sensitivity (CS) quantifies an observer's ability to detect the smallest (threshold) luminance difference between a target and its surrounding. In clinical settings, printed letter contrast charts are commonly used, and the contrast of the letter stimuli is specified by the Weber contrast definition. Those paper-printed charts use negative polarity contrast (NP, dark letters on bright background) and are not available with positive polarity contrast (PP, bright letters on dark background), as needed in a number of applications. We implemented a mobile CS measuring app supporting both NP and PP contrast stimuli that mimic the paper charts for NP. A novel modified Weber definition was developed to specify the contrast of PP letters. The validity of the app is established in comparison with the paper chart. We found that our app generates more accurate and a wider range of contrast stimuli than the paper chart (especially at the critical high CS, low contrast range), and found a clear difference between NP and PP CS measures (CSNP>CSPP) despite the symmetry afforded by the modified Weber contrast definition. Our app provides a convenient way to measure CS in both lighted and dark environments.
Fundamental forms of high-order cognition, such as reading and writing, are usually studied in the context of one modality - vision. People without sight, however, use the kinesthetic-based Braille writing, and haptic-based Braille reading. We asked whether the cognitive and motor control mechanisms underlying writing and reading are modality-specific or supramodal. While a number of previous functional Magnetic Resonance Imaging (fMRI) studies have investigated the brain network for Braille reading in the blind, such studies on Braille writing are lacking. Consequently, no comparative network analysis of Braille writing vs. reading exists. Here, we report the first study of Braille writing, and a comparison of the brain organization for Braille writing vs Braille reading. FMRI was conducted in a Siemens 3T Trio scanner. Our custom MRI-compatible drawing/writing lectern was further modified to provide for Braille reading and writing. Each of five paragraphs of novel Braille text describing objects, faces and navigation sequences was read, then reproduced twice by Braille writing from memory, then read a second time. During Braille reading, the haptic-sensing of the Braille letters strongly activated not only the early visual area V1 and V2, but some highly specialized areas, such as the classical visual grapheme area and the Exner motor grapheme area. Braille-writing-from-memory, engaged a significantly more extensive network in dorsal motor, somatosensory/kinesthetic, dorsal parietal and prefrontal cortex. However, in contrast to the largely extended V1 activation in drawing-from-memory in the blind after training (Likova, 2012), Braille writing from memory generated focal activation restricted to the most foveal part of V1, presumably reflecting topographically the focal demands of such a "pin-pricking" task.
Edges derived from abrupt luminance changes in images carry essential information for object recognition. Typical binary edge images (black edges on white background or white edges on black background) have been used to represent features (edges and cusps) in scenes. However, the polarity of cusps and edges may contain important depth information (depth from shading) which is lost in the binary edge representation. This depth information may be restored, to some degree, using bipolar edges. We compared recognition rates of 16 binary edge images, or bipolar features, by 26 subjects. Object recognition rates were higher with bipolar edges and the improvement was significant in scenes with complex backgrounds.