
Deep learning techniques have revolutionized the fields of image restoration and image quality assessment in recent years. While image restoration methods typically utilize synthetically distorted training data for training, deep quality assessment models often require expensive labeled subjective data. However, recent studies have shown that activations of deep neural networks trained for visual modeling tasks can also be used for perceptual quality assessment of images. Following this intuition, we propose a novel attention-based convolutional neural network capable of simultaneously performing both image restoration and quality assessment. We achieve this by training a JPEG deblocking network augmented with "quality attention" maps and demonstrating state-of-the-art deblocking accuracy, achieving a high correlation of predicted quality with human opinion scores.
Instructional activity recognition is an analytical tool for the observation of classroom education. One of the primary challenges in this domain is dealing with the intricate and heterogeneous interactions between teachers, students, and instructional objects. To address these complex dynamics, we present an innovative activity recognition pipeline designed explicitly for instructional videos, leveraging a multi-semantic attention mechanism. Our novel pipeline uses a transformer network that incorporates several types of instructional semantic attention, including teacher-to-students, students-to-students, teacher-to-object, and students-to-object relationships. This comprehensive approach allows us to classify various interactive activity labels effectively. The effectiveness of our proposed algorithm is demonstrated through its evaluation on our annotated instructional activity dataset.
Of the 50,000 known spider species that taxonomists have identified, a subset of them are considered to be particularly significant due to the harmful physiological effects their venom has on humans, which often demands prompt and precise identification of the spider in emergency scenarios. Traditional spider identification relies on expert knowledge of morphological characteristics for identification, but in critical scenarios this may be inadequate due to time or knowledge constraints. Thanks to the rise of machine learning, we have developed an effective solution to this problem through the testing of powerful deep learning models. In this paper, we utilize various proven image classification models as a backbone, then fine-tune them on a curated dataset of spider images from the citizen science platform iNaturalist with an emphasis on spiders that are particularly harmful to humans. Experimental results are favorable and indicate that modern image classification models perform well on the task of spider species identification. Our highest performing model is a ConvNeXtV2 backbone model which achieves 91.2% accuracy on our testing set. Compared to previous related works, our fine-tuned model is able to achieve higher classification accuracy while handling a much larger number of spider species.
We evaluate the object detection capabilities of deep learning based CNNs on midwave/longwave dual-band infrared (DBIR) video sequences for the first time. The characterization of CNN object detection performance on DBIR data, and in particular comparative analysis of the performance of DBIR systems relative to single-band longwave infrared (LWIR) and midwave infrared (MWIR) systems, has not been reported previously in the open literature. This is due at least in part to a general lack of labeled, publicly available DBIR data sets. In this paper, we apply a well-known, state-of-the-art CNN to DBIR data for the first time. A new labeled DBIR data set was generated comprising multiple classes of vehicles, people, airplanes, and birds. YOLOv4, pre-trained on the MS COCO dataset, was used for inference on the MWIR and LWIR channels of the DBIR sensor independently. The resulting detections from the two bands were considered both separately and jointly. The labeled objects of this DBIR data set were grouped into small, medium, and large classes. Detection performance on the medium and large objects was comparable to YOLOv4 performance reported previously in the open literature for visible wavelength objects in terms of average precision and average recall. Recall performance on small objects showed a significant size-dependent advantage for DBIR over LWIR or MWIR alone.
When viewing the brain as a sophisticated, nonlinear dynamic system, employing complexity measures offers a valuable way to measure the intricate and dynamic aspects of spontaneous psychotic brain activity. These measures can help us identify irregularities and patterns in complex systems. In our study, we utilized fuzzy recurrence plots and sample entropy to evaluate the dynamic characteristics of psychiatric disorders. This assessment focused on understanding the temporal and spatial neural activity patterns, and more specifically, we applied complexity measures to investigate the functional connectivity within the psychotic brain. This involves understanding how different brain regions synchronize their activity, and complexity measures can reveal the patterns of these connections. It provides a means to understand how different brain regions interact and communicate under resting-state abnormal conditions. This study offers evidence demonstrating that fuzzy recurrence plots can serve as descriptors for functional connectivity and discusses their relevance to sample entropy in the context of the psychotic brain. In summary, complexity measures offer valuable insights that enrich our comprehension of atypical brain activity and the complexities present in the psychotic brain.
In this study, we fine-tuned a stable diffusion model to synthesize high resolution chest X-ray images (512x512) with bilateral lung edema caused by COVID-19 pneumonia using the class-specific prior preservation strategy. 300 positive images were selected from the MIDRC dataset as subject instances with an additional 400 negative images for class prior preservation. We synthesized images respectively using the new technique and the conventional technique for comparison. The synthetic images by the stable diffusion fine-tuned by the prior preservation technique have the Frechet inception distance (FID) of 9.2158 and kernel inception distance (KID) 0.0818 computed with the real positive images, which is superior to the synthetic images using the conventional methods such as WGAN and DDIM. The classification accuracy is 0.9975 with precision of 1.0 and recall of 0.9950 when the synthetic positive images with the real negative images were classified by a trained vision transformer (ViT). We conclude that the stable diffusion model can synthesize high-quality and high-resolution chest x-ray images using the prior preservation strategy with a small number of real images as subject instances and text prompt as guidance for the designated patterns.
Unsupervised person re-identification (re-ID) aims to learn identity information from a source domain (e.g. one surveillance system) and apply it to a target domain (e.g. a different surveillance system). This is challenging due to occlusion, viewpoint, and illumination variations between the different domains (i.e. systems). In this paper, we propose a neural network architecture, known as Synthetic Model Bank (SMB), to address illumination variation in unsupervised person re-ID. The basic idea of SMB is to use synthetic data for training different re-ID models for different illumination conditions. From our experiments, the proposed SMB outperforms other synthetic augmentation methods on several re-ID benchmarks.
Person re-identification (re-ID) has wide applications in surveillance and security. It is also challenging due to viewpoint, occlusion and illumination variations across different cameras. One solution to unsupervised person re-ID problems is synthetic data augmentation. Generative neural networks have been used to translate images from the source domain into the target domain. In this paper, we introduce a new virtual-human image dataset that can be used as the source domain for person re-ID. This new dataset has images labeled by person identity, background, viewpoint and illumination intensity. We also explore GAN-based and Diffusion-based generative methods for unpaired image-to-image translation and provide qualitative and quantitative evaluation for the synthetic results.
The high dimensionality and complexity of neuroimaging data necessitate large datasets to develop robust and high-performing deep learning models. However, the neuroimaging field is notably hampered by the scarcity of such datasets. In this work, we proposed a data augmentation and validation framework that utilizes dynamic forecasting with Long Short-Term Memory (LSTM) networks to enrich datasets. We extended multivariate time series data by predicting the time courses of independent component networks (ICNs) in both one-step and recursive configurations. The effectiveness of these augmented datasets was then compared with the original data using various deep learning models designed for chronological age prediction tasks. The results suggest that our approach improves model performance, providing a robust solution to overcome the challenges presented by the limited size of neuroimaging datasets.
Neural network training data is often corrupted by equipment malfunction or noise leading to red blurry and incomplete data. This paper proposes a combination of a reconstruction technique and a neural network to deal with data corruption in a machine vision task. Specifically, we consider minimizing the tensor nuclear norm for low-rank data completion and denoising and demonstrate the method’s effectiveness using a convolutional neural network (CNN) for image classification. We conduct classification experiments on 3 datasets, showing consistently that training on reconstructed images achieves improved accuracy ranging from 7-25% over training using corrupted data.
Swim pose estimation and recognition is a challenging problem in machine learning and artificial intelligence as the body of the swimmer is continuously submerged under the water. The objective of this paper is to enhance existing ML models for estimating swim poses and for recognizing strokes to aid swimmers in pursuing a more perfect technique. We developed a novel methodology augmenting raw video data and adjusting a YOLOv7 base model to enhance swim pose estimation. We found the standard multi-class classification using a CNN to be insufficient for stroke recognition due to the similarity between strokes, so we designed a hierarchical binary classification tree using multiple ensembles of multilayer perceptron (MLP), CNN, and residual network (ResNet) models. Through these optimizations, the confidence level of pose estimation has increased by over $30 \%$, and the ensembles of our recognition model have achieved approximately $80 \%$ accuracy. Fine-tuning of our recognition models and research combining joint keypoint coordinates with angle measurements as inputs could further increase the accuracy of our models.
Hyperspectral imagery comprises a rich source of remote sensing data which can be used for various analysis tasks such as target identification. Machine learning techniques allow analysts to build models that can be trained to perform material identification to high accuracy. Yet key to implementing trained classifier models is understanding on which spectral features the model relies for making decisions. Harnessing explainability methodology along with self-supervised models such as autoencoders, we can begin to probe the limits of what a classification model outputs for end users. In this work, we demonstrate the use of an autoencoder models and alternate spectral representations for contrastive explanations as an explainability method for material classification in hyperspectral imagery data.
Schizophrenia (SZ) is a complex neuropsychiatric disorder characterized by disrupted integration among distributed brain areas. Currently, there are no objective diagnostic tests for schizophrenia. The evaluation of brain network entropy can provide insights into understanding pathological connectomic anomalies in schizophrenia patients, and together with other cross-sectional data potentially suggest biomarkers for diagnosing this complex disease. Our study analyzed resting-state functional magnetic resonance imaging (fMRI) data from 314 subjects, including 153 with schizophrenia patients and 163 age-and gender-matched healthy controls. We focused on 47 functionally relevant intrinsic connectivity brain networks obtained by group independent component analysis (GICA) in a previous study. We evaluated static and dynamic connectivity entropies and found 22 intrinsic connectivity networks with significant differences in heterogeneity of connectivity levels across available networks between SZ patients and healthy controls. These networks are associated with subcortical (SC), auditory (AUD), visual (VIS), somatomotor (SM), cognitive control (CC), and cerebellar (CB) functional brain domains.
While the analysis of temporal signal fluctuations and co-fluctuations has long been a fixture of blood oxygenation-level dependent (BOLD) functional magnetic resonance imaging (fMRI) research, the role and implications of spatial propagation within the 4D neurovascular BOLD signal has been almost entirely neglected. As part of a larger research program aimed at capturing and analyzing spatially propagative dynamics in BOLD fMRI, we report here a method that exposes large-scale functional attractors of flow processes defined via Markov processes defined at the voxel level. The brainwide stationary distributions of these voxel-level Markov processes represent patterns of signal accumulation toward which the brain exerts a probabilistic propagative undertow. These probabilistic propagative attractors are spatially structured and organized interpretably over functional regions. They also differ significantly between schizophrenia patients and healthy controls.
The alarming global statistics concerning traffic fatalities have prompted society and regulatory bodies to act to enhance road safety. The diverse array of cameras situated along roadways, each with its own set of advantages and limitations, creates a rich landscape of options while also making the development of detection and tracking algorithms more intricate and challenging. To achieve dynamic and reliable tracking, semantic and contextual information from the images is indispensable for forming tracking lines. This paper aims to develop a system for analyzing moving vehicles, encompassing vehicle detection and tracking. We employ the YoloV7 algorithm for initial object detection. In the tracking phase, we leverage semantic information from the vehicles via a series of feature extractors. Finally, the convex hull algorithm allows us to focus on areas of interest while tracking moving vehicles. The results demonstrate that the proposed solution performs comparably to state-of-the-art approaches, which predominantly rely on deep learning techniques.
Diabetic peripheral neuropathy (DPN) is a complication of diabetes that causes severe foot pain and frequently leads to amputation. Early detection is critical to saving patients from foot ulcers and amputation. This paper presents an automatic system to identify thermal biomarkers associated with DPN. Research Design and Methods: 141 subjects diagnosed with diabetes mellitus (DM) were enrolled in the study. Subjects were categorized as those with DM but without DPN, those with DPN based on a positive nerve conduction study, and those without DPN. A support vector machine (SVM) was used to classify the subjects having DPN based on thermal parameters related to temperature recovery after a cold provocation. Results: The classifier produced a sensitivity/specificity of 0.78/0.89 in identifying DPN. Conclusions: The SVM classifier can identify patients with the large fiber form of DPN. A different reference standard had to be used to detect small fiber neuropathy.
Hyperspectral unmixing (HU) in Hyperspectral Imaging (HSI) analyzes and breaks down the constituent components within a pixel, providing insights into the object composition within subpixel regions. The availability of diverse methodologies for endmember extraction (EE) are available for HU. Recent work shows the potential value of HU in Very High Spatial Resolution Hyperspectral Images (VHSR-HSI), enabling spectral signature extraction, capturing spectral variability and extracting material spatial distribution. In this study, we aim to study the effectiveness of traditional endmember extraction algorithms for HU within the realm of VHSR-HSI. Preliminary results using and comparing multiple endmember extraction techniques are presented, placing specific emphasis on the efficacy of N-FINDR, Pixel Purity Index (PPI) and a modified version of PPI developed by the authors. Our approach is applied to hyperspectral images obtained from close-range observations using a standoff hyperspectral imager. Results show the effectiveness of NFINDR and modified PPI to extract the spectral signature of the different classes in the image.
Information-theoretic image quality assessment (IQA) models such as Visual Information Fidelity (VIF) and Spatio-temporal Reduced Reference Entropic Differences (ST-RRED) have enjoyed great success by seamlessly integrating natural scene statistics (NSS) with information theory. The Gaussian Scale Mixture (GSM) model that governs the wavelet subband coefficients of natural images forms the foundation for these algorithms. However, the explosion of user-generated content on social media, which is typically distorted by one or more of many possible unknown impairments, has revealed the limitations of NSS-based IQA models that rely on the simple GSM model. Here, we seek to elaborate the VIF index by deriving useful properties of the Multivariate Generalized Gaussian Distribution (MGGD), and using them to study the behavior of VIF under a Generalized GSM (GGSM) model.
An echocardiogram is a video sequence of a human heart captured using ultrasound imaging, which helps in diagnosis of cardiovascular diseases. Deep learning methods, which require large amounts of training data, have shown success in using echocardiograms to detect cardiovascular disorders. Large datasets of echocardiograms that can be used for machine learning training are scarce. This problem can be addressed by generating synthetic echocardiograms that can be used for machine learning training. In this paper, we propose a video diffusion method for echocardiograms generation. We show that our method generates better echocardiograms with higher resolution as compared to existing methods.
Improved motion navigation with efficient data usage and robustness remains a challenge for free-breathing MRI. This paper presents a new 2D auto-navigation technique called VAR-NAV developed for stack-of-stars acquisitions to address this challenge. VAR-NAV uses 2D projection images and multiscale AM-FM demodulation to directly estimate the position of the liver dome with 100% data efficiency. Results can be retrospectively corrected, similar to the VAR system in football video assistant referees. VAR-NAV is used to continuously sort acquired data and reconstruct motion-resolved 4D images, outperforming conventional ID PCA-based navigation without requiring additional acquisition. This efficient and accurate method of clinical 4D MRI offers a promising solution to the challenge of improved motion navigation.