Vision Transformers have excelled in computer vision but their attention mechanisms operate independently across layers, limiting information flow and feature learning. We propose an effective cross-layer attention propagation method that preserves and integrates historical attention matrices across encoder layers, offering a principled refinement of inter-layer information flow in Vision Transformers. This approach enables progressive refinement of attention patterns throughout the transformer hierarchy, enhancing feature acquisition and optimization dynamics. The method requires minimal architectural changes, adding only attention matrix storage and blending operations. Comprehensive experiments on CIFAR-100 and TinyImageNet demonstrate consistent accuracy improvements, with ViT performance increasing from 75.74
Anomaly detection in medical imaging, particularly chest X-rays, is critical in early disease diagnosis and monitoring. Data privacy concerns and complex collection processes lead to limited biomedical data availability. The traditional anomaly detection methods often require large amounts of labeled data, making them impractical in many medical settings. These challenges become particularly acute during epidemics, where rapid model development is essential, and positive samples may be scarce. This paper presents Zero-Anomaly Learning Framework(ZALF) for anomaly detection using Laplacian of Gaussian Generative Adversarial Networks (LoG-GANs) to address these limitations by using a fully unsupervised learning paradigm for effectively learning from normal chest X-ray images only. We propose an unsupervised anomaly detector based on GAN that has only been trained on normal images. It makes use of LoG features for multi-scale structure, a Fourier Loss for global similarity and shift robustness and attention to focus synthesis on important regions. This decreases false positives and enhances detection accuracy. On SARS-Cov-2 CXR dataset, ZALF achieves 0.9888 ROC_AUC, 0.9629 F1 score, 0.9139 Sensitivity and 0.9402 Specificity. These results highlight the potential of our proposed unsupervised method offering scalable and robust solutions for future healthcare needs.
Convolutional Neural Networks (CNNs) have made remarkable strides; however, they remain susceptible to vulnerabilities, particularly in the face of minor image perturbations that humans can easily recognize. This weakness, often termed as 'attacks', underscores the limited robustness of CNNs and the need for research into fortifying their resistance against such manipulations. This study introduces a novel Non-Uniform Illumination (NUI) attack technique, where images are subtly altered using varying NUI masks. Extensive experiments are conducted on widely-accepted datasets including CIFAR10, TinyImageNet, and CalTech256, focusing on image classification with 12 different NUI attack models. The resilience of VGG, ResNet, MobilenetV3-small and InceptionV3 models against NUI attacks are evaluated. Our results show a substantial decline in the CNN models' classification accuracy when subjected to NUI attacks, indicating their vulnerability under non-uniform illumination. To mitigate this, a defense strategy is proposed, including NUI-attacked images, generated through the new NUI transformation, into the training set. The results demonstrate a significant enhancement in CNN model performance when confronted with perturbed images affected by NUI attacks. This strategy seeks to bolster CNN models' resilience against NUI attacks.
We introduce a learning-based CNN architecture for both pre- and post-processing in conjunction with the standard JPEG codec to enhance compression artifact removal. Our approach advances previous methods that use compression-decompression networks by incorporating: (a) an edge-aware loss function designed to reduce blurring, a common issue in earlier works, alongside a super-resolution CNN for post-processing, which improves rate-distortion performance at low bit rates; (b) distinct implementations for compact and full-resolution image representations, tailored for low and high bit rates, respectively. Compared to JPEG, JPEG2000, BPG, VVC-Intra, and recent CNN-based methods, our algorithm demonstrates significant improvements in PSNR, with gains of up to 21% and 25% at low (0.2-0.3 bpp) and high bit rates (0.5-0.7 bpp), respectively. Additionally, the MS-SSIM gain reaches approximately 71% and 65% at low and high bit rates, respectively. This timely method has substantial potential to engage multimedia industries, researchers, and standard-setting agencies.
In recent years, the combined use of Vision Transformer (ViT) and Convolutional Neural Network (CNN) has shown promising results in tasks related to satellite imagery. In our study, we propose a 3D-CmT (3D-CNN meets Transformer) model for Hyperspectral Image Classification. This model leverages the unique capabilities of both 3D-CNN and ViT to effectively classify images captured by hyperspectral imaging. To learn the local features of the narrow and contiguous electromagnetic spectrum of the hyperspectral images, we utilize a 3D-CNN under the spectral feature extraction (SFE) module. Subsequently, a transformer encoder (TE) module is applied on top of the 3D-CNN to incorporate global attention and model long-range dependencies for spatial information in the images. We conducted experiments using commonly used hyperspectral image datasets and performed various ablation studies, such as evaluating the impact of image patch size and different percentages of training samples. The performance of our proposed model is comparable to that of other CNN-based, transformer-based, and hybrid CNN-Transformer-based models in terms of model parameters and accuracy. In addition, we conducted quantitative and qualitative analyses to assess the performance of our model.
Vision transformers (ViT) are deep learning architecture that uses the self-attention mechanism of transformers to analyze and process visual data, such as images. This process involves dividing an input image into smaller patches, which allows each patch to focus on the other. This approach enables high efficiency in a wide range of computer vision applications. Similar to other deep learning models, the ViT model is also sensitive to the color distribution present in images. Color and contrast variations could affect how well the model generalizes to datasets or real-world situations. This research studies the impact of the Color Channel Perturbation (CCP) Attack on Vision Transformer models. In CCP attacks, we get a new transformed dataset used for testing. We conduct the experiments on three well-known datasets, CIFAR10, Caltech256, and TinyImageNet, commonly used in computer vision for image classification tasks. We use three state-of-the-art vision transformer architectures, namely ViT, Pooling-based Vision Transformer (PiT), and Class-Attention in Image Transformers (CaiT). The result shows that the CCP attack is able to fool the vision transformer models heavily on the Caltech256 and TinyImageNet datasets. This study shows a need to investigate the defense mechanism for vision transformer models against CCP attack.
Vision Transformers have proven their mettle across a variety of computer vision problems, however, their reliance on pretraining with very large-scale datasets such as JET-300M is also no secret, as large amounts of data is very conductive to effective feature learning. In this paper, however, we propose a novel Vision Transformer architecture PAEViT that aims to effectively learn generalized features from small-sized datasets, usually having only a few hundred images per class at most. By refining the attention scores based on patch -level interactions and modulating it to enhance the trained model's ability to focus more on task relevant patches, PAEViT is able to significantly improve upon the performance of the regular ViT model, as well as other variants of the same, when the model is being trained solely on small -sized datasets. In data -constrained situations and visual recognition tasks that do not conform well with the existing large-scale datasets, PAEViT can be used to create effective and scalable solutions with all the features and attention scores being based only on relevant data. Our code is publicly available at https://github.com/AkashVermaIN/PAEViT.
Titanium alloy Ti6Al4V forgings with different microstructures were prepared using different combinations of thermo-mechanical processing (TMP) cycles. Ultrasonic inspection, a key non-destructive testing method for inspection of soundness of forgings. This study investigates the relationship between microstructures and ultrasonic responses in Ti6Al4V forgings made as per different TMP cycles. It specifically looks at how different combinations of forging and heat treatment impact the microstructure, and how these changes in turn affect ultrasonic inspection sensitivity. Due to a combination of thermo-mechanical working followed by thermal treatment, it is possible to achieve superior sensitivity in ultrasonic inspection. Results have shown that microstructural features like ratio/volume fraction of phases, phase morphologies, non-uniform microstructures, size of the phases, colonies of grains and micro-texture significantly alter the ultrasonic response. An attempt has been made to establish the correlation between the various types of microstructures, micro-textures obtained by different combinations of forging, heat treatment operations and the resultant ultrasonic response. Morphological changes in regions with micro-texture in the α + β two phase region due to forging and heat treatment were analyzed by EBSD. Notably, the acicular α’, particularly in colony form, can reduce ultrasonic sensitivity due to increased scatter and attenuation. Conversely, fine acicular α′, with random texture that is transformed from primary alpha (αp), results in an improved ultrasonic response of class AA with S/N ratio > 12 dB, regardless of prior thermo-mechanical history of the forgings.
Image super-resolution (SR) aims to reconstruct high-quality images from low-resolution inputs, a task particularly challenging in face-related applications due to extreme degradations and modality differences (e.g., visible, low-resolution, near-infrared). Conventional convolutional neural networks (CNNs) and GANbased approaches have achieved notable success; however, they often struggle with preserving identity and fine structural details at high upscaling factors. In this work, we introduce UpAttTrans, a novel attention mechanism that connects original and upsampled features for better detail recovery based on vision transformer for SR. The core generator leverages a custom UpAttTrans module that translates input image patches into embeddings, processes them through transformer layers enhanced with connector-up attention, and reconstructs high-resolution outputs with improved detail retention. We evaluate our model on the CelebA dataset across multiple upscaling factors (4x, 8x, 16x, 32x, and 64x). UpAttTrans achieves a 24.63% increase in PSNR, 21.56% in SSIM, and 19.61% reduction in FID for 4x and 8x SR, outperforming state-of-the-art baselines. Additionally, for higher magnification levels, our model maintains strong performance, with average gains of 6.20% in PSNR and 21.49% in SSIM, indicating its robustness in extreme SR settings. These findings suggest that UpAttTrans holds significant promise for real-world applications such as face recognition in surveillance, forensic image enhancement, and cross-spectral matching, where high-quality reconstruction from severely degraded inputs is critical.
Urban roads face major challenges during rainfall, impacting visibility and road texture. Rain degrades the performance of vision-based systems in traffic safety and monitoring. Real-world rainy road datasets are scarce and often lack diversity. This limits the training of models for image translation under rainy conditions. Effective rainy image synthesis is essential for advancing robust autonomous driving systems. This paper introduces a novel methodology to bridge this gap by generating synthetic datasets simulating adverse weather conditions, particularly rainfall, using Generative Adversarial Networks (GANs). We introduce a newly curated dataset of Kathmandu road scenes to provide diverse, real-world clear-weather imagery. Leveraging these datasets, we employ advanced deep learning techniques to synthesize high-quality rainy road scenes. In experiments, RFDETR achieved a mean Average Precision of 0.405 (averaged over IoU $0.50-0.95$) on the RainGAN-augmented dataset 0.605 at IoU 0.50 and 0.416 at IoU 0.75 demonstrating that our augmentation preserves detection performance under simulated rainfall. Proposed GAN also outperforms existing GAN variants, achieving an FID of 8.34 and LPIPS of 0.12 (versus DCGAN 18.45/0.29, Pix2Pix 15.32/0.24, and StyleGAN 10.57/0.18), indicating realism in the generated imagery Generated outputs are evaluated using Fréchet Inception Distance (FID) and Learned Perceptual Image Patch Similarity (LPIPS) metrics to ensure realism and perceptual fidelity. The paper outlines the data creation pipeline, model implementation, evaluation strategy, and broader implications, offering a robust framework for weatheraffected scene synthesis and road safety research.
Image super-resolution aims to synthesize high-resolution image from a low-resolution image. It is an active area to overcome the resolution limitations in several applications like low-resolution object-recognition, medical image enhancement, etc. The generative adversarial network (GAN) based methods have been the state-of-the-art for image super-resolution by utilizing the convolutional neural networks (CNNs) based generator and discriminator networks. However, the CNNs are not able to exploit the global information very effectively in contrast to the transformers, which are the recent breakthrough in deep learning by exploiting the self-attention mechanism. Motivated from the success of transformers in language and vision applications, we propose a SRTransGAN for image super-resolution using transformer based GAN. Specifically, we propose a novel transformer-based encoder-decoder network as a generator to generate 2x images and 4x images. We design the discriminator network using vision transformer which uses the image as sequence of patches and hence useful for binary classification between synthesized and real high-resolution images. The proposed SRTransGAN outperforms the existing methods by 4.38 % on an average of PSNR and SSIM scores. We also analyze the saliency map to understand the learning ability of the proposed method.
Because of the synergistic characteristics of the data they provide, thermal imagers are becoming more and more common as secondary data-collecting modules to complement classical optical imagers. Their application is especially beneficial when ambient elements like light, smoke, or other particulates need to be handled. As a result, fusion-based detection technology, like assistive driving systems, is trying to incorporate thermal imagers. In this work, we demonstrate that it is possible to create a mask for pedestrians in thermal imaging by building a basic Denoising Diffusion Probabilistic Model and using the results from the Segment Anything Model. Moreover, we show how the anticipated mask can fuse RGB and thermal pictures into a 3-channel output. Finally, we demonstrate how the loss function optimizes the end-to-end analysis in this field. Results from both the quantitative and qualitative analyses validate our methodology.
Typical machine learning models typically demand a substantial volume of data for effective training, ensuring optimal performance during testing. However, these models often fail to specify the extent of data required—how big data is big enough to begin with? The “label few, classify many” paradigm is interesting in such a scenario where data is limited, and this holds true for healthcare informatics like others. To address this challenge, we explore the application of n-shot GANs in generating synthetic images, particularly in the context of chest X-rays for detecting COVID-19. Starting with one-shot GAN, our findings reveal that training on just 160 healthy chest X-rays achieves outstanding results, with maximum AUCs of 97.2 and 97.4 from ROC and PR curves on a test dataset comprising approximately 3600 chest X-rays. These results demonstrate comparable performance with state-of-the-art approaches in the field.
Image super-resolution generation aims to generate a high-resolution image from its low-resolution image. However, more complex neural networks bring high computational costs and memory usage. It is still an active area that promises to overcome resolution limitations in many applications. Transformers have made significant progress in computer vision tasks in recent years due to their robust self-attention mechanism. However, recent works on the transformer for image super-resolution also contain convolution operations. We propose a patch translator for image super-resolution (PTSR) to address this problem. The proposed PTSR is a transformer-based GAN network with no convolution operations. We introduce a novel patch translator module for regenerating the improved patches utilizing multi-head attention, which is further using by the generator to generate the 2× and 4× super-resolution images. The experiments use benchmark datasets, including DIV2K, Set5, Set14, and BSD100. The results of the proposed model show an improvement on average for 4× super-resolution by 21.66% in PSNR score and 11.59% in SSIM score, compared to the best competitive models. We also analyze the proposed loss and saliency map to show the effectiveness of the proposed method. The code used in the paper will be publicly available at https://github.com/nbaghel777/PTSR.
Object detection on roadways is crucial for autonomous driving and advanced driver assistance systems. However, adverse weather conditions, notably rain, significantly degrade the performance of these systems. This paper presents a novel approach to improve the detection of road objects in rainy weather scenarios by applying YOLOv11 model. This includes specialized data augmentation techniques to simulate rainy conditions, adjustments in network architecture to improve resilience to rain-induced noise, and optimized training strategies to enhance model performance. The study leverages BDD100K, Cityscapes, and DAWN-Rainy datasets of various road scenarios under different rain intensities. We systematically augment these datasets to ensure the model learns to identify objects obscured by rain streaks and reflections. Such enhancements enable better handling of occlusions and reduced visibility in the feature extraction layers. Also, to ensure the model’s efficiency and suitability for real-time applications, we apply a network pruning technique, which reduces the model size and computational requirements without sacrificing performance. Extensive experiments demonstrate that our model has a comparable mean Average Precision with the baseline YOLOv11 but at a 2x compression ratio under rainy conditions. This research contributes to the field of autonomous driving by providing a more reliable object detection system for adverse weather conditions, improving overall road safety.
This paper presents a novel transformer-guided encoder-decoder network for facial Image translation by using multi-layer perceptron and various depth levels. The proposed network exploits the advantage of multi-head self-attention mechanisms for efficient learning. Inspired by the generative capabilities of the generative adversarial network (GAN) for image-to-image translation, we employed the proposed Transformer-Inception based Encoder in the GAN framework. The proposed encoder-decoder network is used as a generator network. The proposed TranscoderGAN model facilitates the intense cyclic training for image-to-image translation efficiently by utilizing inception-based learning in residual networks. The proposed TranscoderGAN is used for facial image synthesis over the sketch in the CUHK dataset as well as the thermal faces of the WHU-IIP dataset. The proposed TranscoderGAN outperforms the state-of-the-art GAN methods for facial image translation. In quantitative results analysis, on an average { 10.36 } improvement is found over Structural Similarity Index Measure (SSIM) and Visual Information Fidelity (VIF), respectively, while { 55.38 } reduction over the Fréchet Inception Distance (Fid) for the CUHK dataset. For WHU-IIP thermal face dataset, the average improvement by the proposed model is { 9.63 %, 1.14 %} over SSIM, VIF, and { 89.25 %} reduction in Fid metrics, respectively.
Firearm-related violence poses a persistent threat to public safety, often challenging for law enforcement agencies. We introduce ThreatNet, a multimodal firearm threat assessment network that detects weapons in unconstrained environments and assesses the threat levels posed by the scene. Our ThreatNet integrates a YOLOv10-based object detector to precisely localize the weapons with a dual-branched transformer-based image captioner to generate textual scene descriptions and classify the threat levels. We present YoutubeGDD caption dataset, an extension of YoutubeGDD, featuring real-world weapon images with five captions per image to support multimodal analysis. We finetune the MS-COCO pretrained YOLOv10 detector on YoutubeGDD dataset for weapon recognition, while we train the captioner and threat classifier on both YoutubeGDD caption and MS-COCO caption datasets. We evaluate the model's performance using standard captioning metrics and threat classification accuracy, and benchmark YOLOv10 variants on YoutubeGDD dataset and captioner on MS-COCO caption dataset.
Face recognition in selfie images is a challenging yet crucial task in computer vision with applications ranging from security systems to social media platforms. In the rapidly evolving field of computer vision, face recognition in uncontrolled settings, such as selfies taken with smartphone cameras, poses unique challenges. This paper examines the efficacy of Vision Transformers (ViT) and its variants, including Class Attention in Transformers (CaiT), CrossViT, MaxViT, MobileViT, Pooling-based Vision Transformer (PiT), and Swin Transformer for face recognition in selfie images. The experiments are performed on the Wild Selfie Dataset (WSD). The dataset features 45,424 selfie images from 42 individuals, characterized by significant variability in expression, lighting, and occlusion, making it a robust testbed for advanced face recognition algorithms. Our study demonstrates that ViTs, leveraging their attention-based mechanisms, surpass the benchmarks set by Convolutional Neural Networks (CNNs) such as VGGFace, VGGFace2, and FaceNet. This paper delineates the performance improvements each ViT variant offers and discusses their potential to streamline face recognition tasks in real-world applications.
The paper explores the idea of joint image encryption and lossless compression. Based upon an exhaustive literature survey, a research gap is identified and addressed using the novel prime product encoding technique. The encoding is a security centric image encoding scheme based upon prime numbers. An integrated multi-image lightweight encryption and lossless compression framework is presented in this paper that uses state of the art cryptosystem for key exchange and block permutation over encoded domain, based approach for ensuring image secrecy. The encoding itself may be tweaked as per the levels of security and compression required. For encryption, only block permutation is used as the input image is already encrypted being in compressed domain. Performance analysis is discussed using noise analysis, key space analysis, PSNR, SSIM, correlation along with compression performance are discussed. The scheme maintains original image resolution while decreasing bit-rate for the blocks.
Metal additive manufacturing (AM) route has gained attraction during recent times owing to its capability to realize complex geometries with fewer processing steps and shorter lead times. Laser powder bed fusion (LPBF) is the most adopted AM technique for producing intricate parts with a fine resolution for engineering applications in general and aerospace applications in particular. Fluid feedline inlet elbows and support clamp components required for a rocket liquid engine were printed with aluminium alloy AlSi10Mg material through LPBF process. These components were previously made from a cast aluminium alloy AS7G in T6 temper condition. AlSi10Mg is characterized by good strength and hardness compared to AS7G and also exhibits high dynamic load-bearing capacity. Hence, it is typically used for 3D printing of parts with thin walls and complex geometries subjected to high loads. AlSi10Mg was selected for this system because it is a reliable alloy that can be 3D printed and is readily available in powder form. It shares many of the same physical, chemical, and mechanical properties as AS7G cast alloy in the T6 condition. A product-specific production and qualification strategy was developed for these parts’ processing and characterisation utilising the LPBF method in line with applicable international standards. These parts have been qualified for intended aerospace applications and those aspects are discussed in this work. This involves powder characterization, measurement of physical and mechanical properties, dimensional inspection, and computed tomography non-destructive testing. The parts were found to be meeting the functional requirements of liquid rocket engines in all respects.