Traditional methods for evaluating crop ripeness are critiqued for their inefficiency and potential harm to produce. The use of image-processing and deep-learning techniques can solve these issues as a trend in non-destructive methods. However, an overfitting problem arises when optimization and generalization are used to estimate the parameters of the next epoch. In this paper, we develop specialized models with a high volume of training images for a single type of crop to achieve the goal of 100% accuracy for both test and validation datasets. This development contributes insights into leveraging deep learning for crop assessment, emphasizing its potential application in diverse agricultural scenarios. Experimental results show that the proposed models are superior to several existing available methods.
Recently, self-supervised learning has drawn lots of attention from researchers. CLIP is a vision-language model that performs cross-modality contrastive pre-training. In this paper, we propose a novel method of prompt tuning by optimal transport to improve zero-shot generalization of the CLIP pre-trained model. Existing entropy-based approaches fail to consider the global structure of output distribution, and cannot align distributions effectively across domains. We develop the Optimal Transport-Test Time Prompt Tuning, named OT-TPT, to resolve this issue. With the help of optimal transport, it can directly align distributions to provide a global regularization effect, and therefore improve robustness against noise and distribution shifts. Moreover, a Sinkhorn regularization term is adopted to provide an efficient and smooth approximation that reduces distribution shifts while improving zero-shot generalization. Experimental results show that the proposed OT-TPT can achieve higher classification accuracies over existing state-of-the-art approaches.
Although contrastive learning has been playing a critical role in pattern recognition, how to optimize positive pairs through data transformation is still not well developed up to now. In this paper, we propose a novel Adaptive Data Transformation, named ADTrans, to identify an optimal sequence of data transformations, which enables generating high-quality positive pairs adaptively during contrastive training. Extensive experiments on benchmark datasets have shown that ADTrans can improve the performance of representation learning on downstream tasks significantly, including image classification, instance segmentation, and object detection. It can achieve a classification accuracy of 12% and 9% higher than the existing MOCO v2, SimSiam, and BYOL on the STL 10 and TinyImageNet datasets, respectively, with the ResNet-18 backbone. Moreover, it outperforms MOCO v2 on COCO instance segmentation, object detection, and Pascal VOC instance segmentation.
Both a deep understanding of visual cues and their contextual importance are demanded by effective image captioning. However, seamlessly integrating balanced contextual information continues to be a substantial challenge. In this paper, we present FreConCap, a novel Frequency-guided Con-textual Image Captioning framework, to overcome the challenge using high-frequency and background features, along with object-level region features. We transform grid features into frequency domain and filter out low-frequency components by a cutoff ratio that enhances fine details critical for detailed visual understanding. Multi-Stream Cross Attention is developed to reduce the modality gap between vision and language, and to capture the interaction of text features with high-frequency local features, objects, context, and their relationships. Our experiments on the MS COCO image captioning benchmark show the superiority of our approach as compared with existing methods for enhanced image captions with more contextual information.
In medical imaging, accurate diagnosis heavily relies on effective image enhancement techniques, particularly for X-ray images. Existing methods often suffer from various challenges such as sacrificing global image characteristics over local image characteristics or vice versa. In this paper, we present a novel approach, called G-CLAHE (Global-Contrast Limited Adaptive Histogram Equalization), which perfectly suits medical imaging with a focus on X-rays. This method adapts from Global Histogram Equalization (GHE) and Contrast Limited Adaptive Histogram Equalization (CLAHE) to take both advantages and avoid weakness to preserve local and global characteristics. Experimental results show that it can significantly improve current state-of-the-art algorithms to effectively address their limitations and enhance the contrast and quality of X-ray images for diagnostic accuracy.
In order to safeguard image copyrights, zero-watermarking technology extracts robust features and generates watermarks without altering the original image. Traditional zero-watermarking methods rely on handcrafted feature descriptors to enhance their performance. With the advancement of deep learning, this paper introduces “ZWNet”, an end-to-end zero-watermarking scheme that obviates the necessity for specialized knowledge in image features and is exclusively composed of artificial neural networks. The architecture of ZWNet synergistically incorporates ConvNeXt and LK-PAN to augment the extraction of local features while accounting for the global context. A key aspect of ZWNet is its watermark block, as the network head part, which fulfills functions such as feature optimization, identifier output, encryption, and copyright fusion. The training strategy addresses the challenge of simultaneously enhancing robustness and discriminability by producing the same identifier for attacked images and distinct identifiers for different images. Experimental validation of ZWNet’s performance has been conducted, demonstrating its robustness with the normalized coefficient of the zero-watermark consistently exceeding 0.97 against rotation, noise, crop, and blur attacks. Regarding discriminability, the Hamming distance of the generated watermarks exceeds 88 for images with the same copyright but different content. Furthermore, the efficiency of watermark generation is affirmed, with an average processing time of 96 ms. These experimental results substantiate the superiority of the proposed scheme over existing zero-watermarking methods.
Counting small pixel-sized vehicles and crowds in unmanned aerial vehicles (UAV) images is crucial across diverse fields, including geographic information collection, traffic monitoring, item delivery, communication network relay stations, as well as target segmentation, detection, and tracking. This task poses significant challenges due to factors such as varying view angles, non-fixed drone cameras, small object sizes, changing illumination, object occlusion, and image jitter. In this paper, we introduce a novel multi-data-augmentation and multi-deep-learning framework designed for counting small vehicles and crowds in UAV images. The framework harnesses the strengths of specific deep-learning detection models, coupled with the convolutional block attention module and data augmentation techniques. Additionally, we present a new method for detecting cars, motorcycles, and persons with small pixel sizes. Our proposed method undergoes evaluation on the test dataset v2 of the 2022 AI Cup competition, where we secured the first place on the private leaderboard by achieving the highest harmonic mean. Subsequent experimental results demonstrate that our framework outperforms the existing YOLOv7-E6E model. We also conducted comparative experiments using the publicly available VisDrone datasets, and the results show that our model outperforms the other models with the highest AP50 score of 52%.
Mathematical morphology and convolution operators are two different methods to extract the characteristics and structures of images. Over the past decades, Deep Convolutional Neural Networks (DCNN) have been proven to be more powerful than traditional image-processing approaches. In this paper, we propose a novel structure called Deep Hybrid Neural Network (DHNN) by taking advantage of the convolution and morphological neural layers. Its practical application to polyp detection in medical images is illustrated. For experimental completeness, we adopt nine polyp image datasets, including publicly available data and our own collected data. For performance comparisons, we select three backbone models. Experimental results show that our DHNN achieves the best performance in comparisons in terms of computational complexity and accurate performance.
Drug property prediction, especially toxicity, helps reduce risks in a range of real-world applications. In this paper, we aim to apply various machine-learning models for solving the drug toxicity prediction problem. Among various machine-learning approaches, we select five suitable representatives: random forest, multi-layer perceptron, logistic regression, graph convolutional neural network, and graph isomorphism network (GIN) for conducting experiments on six datasets for toxicity prediction, including Tox 21, ClinTox, ToxCast, SIDER, HIV, and BACE. We design the GIN with four hidden layers and select the Adam optimizer with the learning rate [Formula: see text] and the batch size [Formula: see text]. Furthermore, we use a batch norm layer inside each of the GIN hidden layers. Experimental results show that the designed GIN model is most efficient in distinguishing between safe and toxic drugs and outperforms the others under the supervision of ROC AUC score and recall.
The automatic detection and recognition for motorcycle license plates present a very challenging task since they appear more compact and versatile than vehicle license plates. In this paper, we present an efficient detection and recognition system for motorcycle license plates based on decision tree and deep learning. It can be successfully carried out under various conditions, such as frontal, horizontally or vertically skewed, blurry, poor illumination, large viewing distances or angles, distortions, multiple license plates in an image, at night or interfered with brake lights, and headlights. Experimental results show that our system performs the best when testing with multiple license plates images under different conditions as compared against six state-of-the-art methods. Furthermore, our detection and recognition system have shown more accurate results than three commercial automatic license plate recognition systems in evaluation using accuracy, precision, recall, and F1 rates.
Incorporating geometric transformations that reflect the relative position changes between an observer and an object into computer vision and deep learning models has attracted much attention in recent years. However, the existing proposals mainly focus on the affine transformation that is insufficient to reflect such geometric position changes. Furthermore, current solutions often apply a neural network module to learn a single transformation matrix, which not only ignores the importance of multi-view analysis but also includes extra training parameters from the module apart from the transformation matrix parameters that increase the model complexity. In this paper, a perspective transformation layer is proposed in the context of deep learning. The proposed layer can learn homography, therefore reflecting the geometric positions between observers and objects. In addition, by directly training its transformation matrices, a single proposed layer can learn an adjustable number of multiple viewpoints without considering module parameters. The experiments and evaluations confirm the superiority of the proposed layer.
Digital image watermarking is the process of embedding and extracting a watermark covertly on a cover-image. To dynamically adapt image watermarking algorithms, deep learning-based image watermarking schemes have attracted increased attention during recent years. However, existing deep learning-based watermarking methods neither fully apply the fitting ability to learn and automate the embedding and extracting algorithms, nor achieve the properties of robustness and blindness simultaneously. In this paper, a robust and blind image watermarking scheme based on deep learning neural networks is proposed. To minimize the requirement of domain knowledge, the fitting ability of deep neural networks is exploited to learn and generalize an automated image watermarking algorithm. A deep learning architecture is specially designed for image watermarking tasks, which will be trained in an unsupervised manner to avoid human intervention and annotation. To facilitate flexible applications, the robustness of the proposed scheme is achieved without requiring any prior knowledge or adversarial examples of possible attacks. A challenging case of watermark extraction from phone camera-captured images demonstrates the robustness and practicality of the proposal. The experiments, evaluation, and application cases confirm the superiority of the proposed scheme.
The chest X-ray images are difficult to classify for the radiologists due to the noisy nature. The existing models based on convolutional neural networks contain a giant number of parameters, and thus require multi-advanced GPUs to deploy. In this paper, we are the first to develop the adaptive morphological neural networks to classify chest X-ray images, such as pneumonia and COVID-19. A novel structure, which can self-learn morphological dilation and erosion, is proposed to determine the most suitable depth of the adaptive layer. Experimental results on the chest X-ray and the COVID-19 datasets show that the proposed model can achieve the highest classification rate as compared against the existing models. Moreover, it can significantly reduce the computational parameters of the existing models by 97%. The advantage makes the developed model more attractive than others to deploy in the internet and other device platforms.
Image segmentation as a clustering problem is to identify pixel groups on an image without any preliminary labels available. It remains a challenge in machine vision because of the variations in size and shape of image segments. Furthermore, determining the segment number in an image is NP-hard without prior knowledge of the image content. This paper presents an automatic color image pixel clustering scheme based on mussels wandering optimization. By applying an activation variable to determine the number of clusters along with the cluster centers optimization, an image is segmented with minimal prior knowledge and human intervention. By revising the within- and between-class sum of squares ratio for random natural image contents, we provide a novel fitness function for image pixel clustering tasks. Comprehensive empirical studies of the proposed scheme against other state-of-the-art competitors on synthetic data and the ASD dataset have demonstrated the promising performance of the proposed scheme.
Remote sensing techniques have been developed over the past decades to acquire data without being in contact of the target object or data source. Their application on land-cover image segmentation has attracted significant attention during recent years. With the help of satellites, scientists and researchers can collect and store high-resolution image data that can be further be processed, segmented, and classified. However, these research results have not yet been synthesized to provide coherent guidance on the effect of variant land-cover segmentation processes. In this paper, we present a novel model that augments segmentation using smaller networks to segment individual classes. The combined network is trained on the same data but with the masks, combined and trained using categorical cross entropy. Experimental results show that the proposed method produces the highest mean IoU (Intersection of Union) as compared against several existing state-of-the-art models on the DeepGlobe dataset.
Chest X-ray images are notoriously difficult to analyze due to the noisy nature. Automatic identification of pneumonia on medical images has attracted intensive study recently. In this paper, a novel joint-task architecture that can learn pneumonia classification and segmentation simultaneously is presented. Two modules, including an image preprocessing module and an attention module, are developed to improve both the classification and segmentation accuracies. Results from the experiments performed on the massive dataset of the Radiology Society of North America have confirmed its superiority over the other existing methods. The classification test accuracy is improved from 0.89 to 0.95, and the segmentation model achieves an improved mean precision result of 0.58–0.78. Finally, two weakly supervised learning methods, class-saliency map and Grad-CAM, are used to highlight the corresponding pixels or areas which have significant influence on the classification model, such that the refined segmentation can focus on the correct areas with high confidence.
We examine to what extent the GICS sector categorization of equity securities may be systematically reconstructed from historical quarterly firm fundamental data using gradient boosted tree classification. Model complexity and performance tradeoffs are examined and relative feature importance is described. Potential extensions are outlined including ideas to improve feature engineering, validating internal consistency and integrating additional data sources to further improve classification accuracy.
This paper presents a novel deep learning classification technique applied on optical coherence tomography (OCT) retinal images. We propose the deep neural networks based on Vgg16 pretrained network model. The OCT retinal image dataset consists of four classes, including three most common retina diseases and one normal retina scan. Because the scale of training data is not sufficiently large, we use the transfer learning technique. Since the convolutional neural networks are sensitive to a little data change, we use data augmentation to analyze the classified results on retinal images. The input grayscale OCT scan images are converted to RGB images using colormaps. We have evaluated different types of classifiers with variant parameters in training the network architecture. Experimental results show that testing accuracy of 99.48% can be obtained as combined on all the classes.
Adversarial attacks in medical AI imaging systems can lead to misdiagnosis and insurance fraud as recently highlighted by Finlayson et. al. in Science 2019. They can also be carried out on widely used ECG time-series data as shown in Han et. al. in Nature Medicine 2020. At the heart of adversarial attacks are imperceptible distortions that are visually and statistically undetectable but cause the machine learning model to misclassify data. Recent empirical studies have shown that a gradient-free trained sign activation neural network ensemble model requires a larger distortion than state of the art models. We apply them on medical data in this study as a potential solution to detect and deter adversarial attacks. We show on chest X-ray and histopathology images, and on two ECG datasets that this model requires a greater distortion to be fooled than full-precision, binary, and convolutional neural networks, and random forests. We show that adversaries targeting the gradient-free sign networks are visually distinguishable from the original data and thus likely to be detected by human inspection. Since the sign network distortions are higher we expect an automated method could be developed to detect and deter attacks in advance. Our work here is a significant step towards safe and secure medical machine learning.