traffic sign detection is critical to autonomous driving and intelligent transportation systems. However, accurately detecting small and distant signs against complex backgrounds remains challenging, often resulting in low accuracy and high miss rates. To address these challenges, we present LRTSD-DETR, a Lightweight Real-Time Traffic Sign Detection Transformer. This network significantly enhances performance through our proposed Multi-Scale Channel-Aware Fusion (MSCA-Fusion) module, which employs a scale alignment strategy and a learnable channel-wise weighting mechanism to strengthen cross-scale feature representation. Furthermore, our novel Partial Convolution Enhanced Block (PCE Block) combines partial and pointwise convolutions to reduce computational complexity while preserving high-quality feature representations. Finally, we adopt the Wise-IoU v3 regression loss, which dynamically prioritizes high-quality samples to improve both localization accuracy and detection stability. Results from extensive experiments on the TT100K dataset demonstrate that LRTSD-DETR outperforms RT-DETR by achieving improvements of 5.36%, 1.98%, 3.57%, and 3.58% in precision, recall, F1-score, and mAP@0.5, respectively. Additionally, it realized a 29% reduction in parameters and a 22.4% drop in FLOPs, resulting in a well-balanced trade-off between accuracy and efficiency without compromising realtime detection. Further evaluations on the CCTSDB dataset validate its strong generalization capability, confirming its effectiveness across different traffic sign datasets.
The rapid growth of social media has resulted in an explosion of online news content, leading to a significant increase in the spread of misleading or false information. While machine learning techniques have been widely applied to detect fake news, the scarcity of labeled datasets remains a critical challenge. Misinformation frequently appears as paired text and images, where a news article or headline is accompanied by a related visuals. In this paper, we introduce a self-learning multimodal model for fake news classification. The model leverages contrastive learning, a robust method for feature extraction that operates without requiring labeled data, and integrates the strengths of Large Language Models (LLMs) to jointly analyze both text and image features. LLMs are excel at this task due to their ability to process diverse linguistic data drawn from extensive training corpora. Our experimental results on a public dataset demonstrate that the proposed model outperforms several state-of-the-art classification approaches, achieving over 85% accuracy, precision, recall, and F1-score. These findings highlight the model's effectiveness in tackling the challenges of multimodal fake news detection.
Visual anomaly detection is a highly challenging task, often categorized as a one-class classification and segmentation problem. Recent studies have demonstrated that the student-teacher (S-T) framework effectively addresses this challenge. However, most S-T frameworks rely solely on pre-trained teacher networks to guide student networks in learning multi-scale similar features, overlooking the potential of the student networks to enhance learning through multi-scale feature fusion. In this study, we propose a novel model named PFADSeg, which integrates a pre-trained teacher network, a denoising student network with multi-scale feature fusion, and a guided anomaly segmentation network into a unified framework. By adopting a unique teacher-encoder and student-decoder denoising mode, the model improves the student network's ability to learn from teacher network features. Furthermore, an adaptive feature fusion mechanism is introduced to train a self-supervised segmentation network that synthesizes anomaly masks autonomously, significantly increasing detection performance. Rigorous evaluations on the widely-used MVTec AD dataset demonstrate that PFADSeg exhibits excellent performance, achieving an image-level AUC of 98.9
The innovative generation of vector graphics with fine-grained images using Artificial Intelligence has become an important task in edge extraction. In this paper, we take Qiang embroidery image as an example due to its containing fine-grained edges, which is more suitable for the study of image processing and pattern recognition. We firstly adopt appropriate pre-processing methods, improved adaptive median filtering (IAMF) for the image to reduce image noise. Then, the Xception based on convolutional neural networks is used for edge detection and extraction. Results show that Qiang embroidery images, after denoising and edge extraction, can be clearly identified the shape characteristics of the images. Based on this approach, it can be converted into vector graphics for digital preservation and further artistic reinterpretation. The use of the Xception effectively addresses Qiang embroidery extraction in two-dimensional vector images, offering a practical reference for preserving related intangible cultural heritage.
Accurate and efficient question-answering systems are essential for high-quality patient care in the medical field. While Large Language Models (LLMs) have made remarkable strides across various domains, they still face challenges in medical question answering, particularly in understanding domain-specific terminology and performing complex reasoning, limiting their effectiveness in critical applications. To address this, we propose a multi-agent medical question-answering (MedQA) system incorporating similar case generation. We leverage the Llama3.1:70B model in a multi-agent architecture to enhance enhance zero-shot classification on the MedQA dataset, utilizing the model’s inherent medical knowledge and reasoning capabilities without additional training data. Experimental results show substantial gains over existing benchmark models, with improvements of 7% in both accuracy and F1-score across various medical QA tasks. Furthermore, we examine the model’s interpretability and reliability in addressing complex medical queries. This research not only offers a robust solution for medical question answering but also establishes a foundation for broader applications of LLMs in the medical domain.
Harnessing the power of artificial intelligence(AI) approaches to innovatively generating the vector graphics of fine-grained patterns has become an important task in image edge extraction, particularly on the domain of intangible cultural heritage (ICH) images where they are typically fine-grained and having the complex edges. With higher autonomy, the machine learning algorithms are able to accurately extract the image information, understand and convey the concept contained in it. In this paper, we take Qiang embroidery patterns as an example due to containing fine-grained patterns, which is more suitable for the study of image processing and pattern recognition techniques. We firstly adopt appropriate pre-processing methods, improved adaptive median filtering(IAMF) and non-local mean for the two different types of Qiang embroidery patterns to reduce image noise. Then, the Xception algorithm based on convolutional neural networks(CNNs) is used for edge detection and extraction to generate vector graphics of the patterns. Experimental results show that Qiang embroidery patterns, after denoising and edge extraction, can be clearly identified the shape characteristics of the patterns. Based on this approach, the images can be converted into vector graphics for the digital preservation and further artistic reinterpretation. The use of the Xception algorithm effectively solves the problem of extraction of Qiang embroidery in two-dimensional vectorial images. In addition, our proposed method provides a reliable practical reference for the preservation of other related ICH images.
Accurate cloud classification plays a crucial role in aviation safety, climate monitoring, and localized weather forecasting. Current research has been focusing on machine learning techniques, particularly deep learning based model, for the types identification. However, traditional approaches such as convolutional neural networks (CNNs) encounter difficulties in capturing global contextual information. In addition, they are computationally expensive, which restricts their usability in resource-limited environments. To tackle these issues, we present the Cloud Vision Transformer (CloudViT), a lightweight model that integrates CNNs with Transformers. The integration enables an effective balance between local and global feature extraction. To be specific, CloudViT comprises two innovative modules: Feature Extraction (E_Module) and Downsampling (D_Module). These modules are able to significantly reduce the number of model parameters and computational complexity while maintaining translation invariance and enhancing contextual comprehension. Overall, the CloudViT includes 0.93 x 10(6) parameters, which decreases more than ten times compared to the SOTA (State-of-the-Art) model CloudNet. Comprehensive evaluations conducted on the HBMCD and SWIMCAT datasets showcase the outstanding performance of CloudViT. It achieves classification accuracies of 98.45% and 100%, respectively. Moreover, the efficiency and scalability of CloudViT make it an ideal candidate for deployment in mobile cloud observation systems, enabling real-time cloud image classification. The proposed hybrid architecture of CloudViT offers a promising approach for advancing ground-based cloud image classification. It holds significant potential for both optimizing performance and facilitating practical deployment scenarios.
With the development of social media, the amount of fake news has risen significantly and had a great impact on both individuals and society. The restrictions imposed by censors make the objective reporting of news difficult. Most studies use supervised methods, relying on a large amount of labeled data for fake news detection, which hinders the effectiveness of the detection. Meanwhile, the focus of these studies is on the detection of fake news in a single modality, either text or images, but actual fake news is more often in the form of text–image pairs. In this paper, we introduce a self-supervised model grounded in contrastive learning. This model facilitates simultaneous feature extraction for both text and images by employing dot product graphic matching. Through contrastive learning, it augments the extraction capability of image features, leading to a robust visual feature extraction ability with reduced training data requirements. The model’s effectiveness was assessed against the baseline using the COSMOS fake news dataset. The experiments reveal that, when detecting fake news with mismatched text–image pairs, only approximately 3% of the data are used for training. The model achieves an accuracy of 80%, equivalent to 95% of the original model’s performance using full-size data for training. Notably, replacing the text encoding layer enhances experimental stability, providing a substantial advantage over the original model, specifically on the COSMOS dataset.
Providing explanations within the recommendation system would boost user satisfaction and foster trust, especially by elaborating on the reasons for selecting recommended items tailored to the user. The predominant approach in this domain revolves around generating text-based explanations, with a notable emphasis on applying large language models (LLMs). However, refining LLMs for explainable recommendations proves impractical due to time constraints and computing resource limitations. As an alternative, the current approach involves training the prompt rather than the LLM. In this study, we developed a model that utilizes the ID vectors of user and item inputs as prompts for GPT-2. We employed a joint training mechanism within a multi-task learning framework to optimize both the recommendation task and explanation task. This strategy enables a more effective exploration of users’ interests, improving recommendation effectiveness and user satisfaction. Through the experiments, our method achieving 1.59 DIV, 0.57 USR and 0.41 FCR on the Yelp, TripAdvisor and Amazon dataset respectively, demonstrates superior performance over four SOTA methods in terms of explainability evaluation metric. In addition, we identified that the proposed model is able to ensure stable textual quality on the three public datasets.
Generative adversarial networks (GANs) have remarkably advanced in diverse domains, especially image generation and editing. However, the misuse of GANs for generating deceptive images, such as face replacement, raises significant security concerns, which have gained widespread attention. Therefore, it is urgent to develop effective detection methods to distinguish between real and fake images. Current research centers around the application of transfer learning. Nevertheless, it encounters challenges such as knowledge forgetting from the original dataset and inadequate performance when dealing with imbalanced data during training. To alleviate this issue, this paper introduces a novel GAN-generated image detection algorithm called X-Transfer, which enhances transfer learning by utilizing two neural networks that employ interleaved parallel gradient transmission. In addition, we combine AUC loss and cross-entropy loss to improve the model's performance. We carry out comprehensive experiments on multiple facial image datasets. The results show that our model outperforms the general transferring approach, and the best metric achieves 99.04%, which is increased by approximately 10%. Furthermore, we demonstrate excellent performance on non-face datasets, validating its generality and broader application prospects.
In recent years, style transfer techniques have emerged as a powerful tool for infusing life into static images, finding applications across various domains. One notable area is the realm of cultural dissemination, where style transfer can enhance interactivity for the public audience. However, conventional approaches often rely on deep learning models, necessitating time-consuming training processes and making them less practical for lightweight devices. In this paper, we explore the application of FaceBlit in the cultural domain. Using the example of terracotta warriors, we demonstrate style transfer applied to human faces, creating engaging images with distinct characteristics. This method offers the advantage of minimal training requirements, with just one photo needed to introduce the desired style. Experimental validation underscores the effectiveness of the FaceBlit in style transfer, yielding $i^{1}$ mpressive results in both images and videos. This study introduces an innovative approach to cultural dissemination, promising exciting possibilities for interactive cultural content.
Precipitation nowcasting plays an important role in mitigating the damage caused by severe weather. The objective of precipitation nowcasting is to forecast the weather conditions 0–2 h ahead. Traditional models based on numerical weather prediction and radar echo extrapolation obtain relatively better results. In recent years, models based on deep learning have also been applied to precipitation nowcasting and have shown improvement. However, the forecast accuracy is decreased with longer forecast times and higher intensities. To mitigate the shortcomings of existing models for precipitation nowcasting, we propose a novel model that fuses spatiotemporal features for precipitation nowcasting. The proposed model uses an encoder–forecaster framework that is similar to U-Net. First, in the encoder, we propose a spatial and temporal multi-head squared attention module based on MaxPool and AveragePool to capture every independent sequence feature, as well as a global spatial and temporal feedforward network, to learn the global and long-distance relationships between whole spatiotemporal sequences. Second, we propose a cross-feature fusion strategy to enhance the interactions between features. This strategy is applied to the components of the forecaster. Based on the cross-feature fusion strategy, we constructed a novel multi-head squared cross-feature fusion attention module and cross-feature fusion feedforward network in the forecaster. Comprehensive experimental results demonstrated that the proposed model more effectively forecasted high-intensity levels than other models. These results prove the effectiveness of the proposed model in terms of predicting convective weather. This indicates that our proposed model provides a feasible solution for precipitation nowcasting. Extensive experiments also proved the effectiveness of the components of the proposed model.
Internet public opinion is closely related to our life in social network. The wanton growth of some negative public opinions has an extremely serious impact on the social stability and national security. After the guidance of government manually, some negative public opinion is well controlled and people’s life gain more positive energy. How to use Internet technology to automatically and promptly guide public opinion events and reduce the harm to society is currently challenging research. Therefore, in this paper, we propose a positive public opinion guidance model based on dual learning for negative Internet public opinion, hereinafter denoted to as the dual-PPOG model. Firstly, we use the Fast Unfolding algorithm to divide social networks into the public opinion guidance communities. In these communities, we detect the positive opinion guider and negative opinion receiver by our defined PageRank (PR) variant. Secondly, inspired by dual learning, we construct the public opinion guidance model and evaluate whether the guidance is successful through the feedback signal. Through the repeated guidance of the positive opinion guider to the negative opinion receiver, the public opinion guidance is successful. This is the main process for the dual positive public opinion mechanism. Finally, we guide the remaining nodes based on the opinion dynamics. The experiment demonstrates beneficial effects of our proposed model of dual-PPOG. Experimental results on three real-world datasets intercepted from Twitter, E-mail and Facebook show that the model of dual-PPOG can capture useful information in the network topology. Compared with the methods of HK, AE, Random and AIA on the three datasets from small to large in scale, the percentage of positive opinion increased by 4%, 6.9%, and 2.7% respectively, which shows our approach achieve significant improvements and effectiveness compared to all baselines.
As growing usage of social media websites in the recent decades, the amount of news articles spreading online rapidly, resulting in an unprecedented scale of potentially fraudulent information. Although a plenty of studies have applied the supervised machine learning approaches to detect such content, the lack of gold standard training data has hindered the development. Analysing the single data format, either fake text description or fake image, is the mainstream direction for the current research. However, the misinformation in real-world scenario is commonly formed as a text-image pair where the news article/news title is described as text content, and usually followed by the related image. Given the strong ability of learning features without labelled data, contrastive learning, as a self-learning approach, has emerged and achieved success on the computer vision. In this paper, our goal is to explore the constrastive learning in the domain of misinformation identification. We developed a self-learning model and carried out the comprehensive experiments on a public data set named COSMOS. Comparing to the baseline classifier, our model shows the superior performance of non-matched image-text pair detection (approximately 10%) when the training data is insufficient. In addition, we observed the stability for contrsative learning and suggested the use of it offers large reductions in the number of training data, whilst maintaining comparable classification results.
There are significant background changes and complex spatial correspondences between multi-modal remote sensing images, and it is difficult for existing methods to extract common features between images effectively, leading to poor matching results. In order to improve the matching effect, features with high robustness are extracted; this paper proposes a multi-temporal remote sensing matching algorithm CMRM (CNN multi-modal remote sensing matching) based on deformable convolution and cross-attention. First, based on the VGG16 backbone network, Deformable VGG16 (DeVgg) is constructed by introducing deformable convolutions to adapt to significant geometric distortions in remote sensing images of different shapes and scales; second, the features extracted from DeVgg are input to the cross-attention module to better capture the spatial correspondence of images with background changes; and finally, the key points and corresponding descriptors are extracted from the output feature map. In the feature matching stage, in order to solve the problem of poor matching quality of feature points, BFMatcher is used for rough registration, and then the RANSAC algorithm with adaptive threshold is used for constraint. The proposed algorithm in this paper performs well on the public dataset HPatches, with MMA values of 0.672, 0.710, and 0.785 when the threshold is selected as 3–5. The results show that compared to existing methods, our method improves the matching accuracy of multi-modal remote sensing images.
Snow and ice crystals are important components of hydrometeors in clouds. Different shapes of snow and ice crystals are determined by their different physical formation and growth processes. Accurate observation of the shapes of snow and ice crystals is an important prerequisite for revealing the microphysical structure and precipitation mechanism of clouds. This paper summarizes studies and methods of observing crystals of ice and snow in the past half-century. The development of crystal measurement and shape classification technology for snow and ice is reviewed. The new progress of snow and ice crystal observation and shape classification and recognition technology is analyzed and summarized. The aim is to provide a reference of images for further studies of cloud microphysical structure and precipitation mechanism in China.
Public opinion in online social network is currently an important topic related to our life closely, however, the research on public opinion guidance lacks a theoretical system and leveraging machine learning to deal with public opinion guidance is even minimal, so it is critical to study the control, guidance and the evolution of public opinion and to use machine learning to study the guidance of public opinion. In the paper, we proposed a public opinion guidance model based on dual learning, hereinafter referred to as the dual-POG model. Firstly, we use the Girvan and Newman (GN) algorithm to divide communities in social network and detect the opinion leaders. Secondly, apply the dual learning to construct the public opinion guidance model which is the main idea of the dual guidance mechanism we proposed to guide the leaders. Finally, we guide the remaining nodes based on opinion dynamics. This study applies machine learning to guide the online public opinion, which has an important significance for current public opinion development. The experiments demonstrate beneficial effects of the model of dual-POG. Comparison experiments indicate that the proposed approach outperforms the other methods.
Deformable medical image registration plays a vital role in medical image applications, such as placing different temporal images at the same time point or different modality images into the same coordinate system. Various strategies have been developed to satisfy the increasing needs of deformable medical image registration. One popular registration method is estimating the displacement field by computing the optical flow between two images. The motion field (flow field) is computed based on either gray-value or handcrafted descriptors such as the scale-invariant feature transform (SIFT). These methods assume that illumination is constant between images. However, medical images may not always satisfy this assumption. In this study, we propose a metric learning-based motion estimation method called Siamese Flow for deformable medical image registration. We train metric learners using a Siamese network, which produces an image patch descriptor that guarantees a smaller feature distance in two similar anatomical structures and a larger feature distance in two dissimilar anatomical structures. In the proposed registration framework, the flow field is computed based on such features and is close to the real deformation field due to the excellent feature representation ability of the Siamese network. Experimental results demonstrate that the proposed method outperforms the Demons, SIFT Flow, Elastix, and VoxelMorph networks regarding registration accuracy and robustness, particularly with large deformations.
It is of great significance to make full use of the complementary advantages of different modality imaging information for improving the accuracy of tumor segmentation and formulating precise radiotherapy plans. This paper proposed a multi-tasking parallel training method, which combined the attention mechanism of specific tasks to mine the effective information of different modals. It has three parallel learning networks based on parameter sharing, including CT segmentation network, MRI segmentation network, and the joint learning network of similarity measurement between CT and MRI images. CT and MRI segmentation networks learned their specific task features, and used the attention module of specific tasks to enhance the utilization of effective features while learning shared features. The similarity measurement learning network jointly learned the similarity between CT and MRI images, and combined the specific task features shared by CT and MRI segmentation networks to segment multimodal tumor images. Comparing the results of single-modal and multi-modal tumor image segmentation, it is proved that multi-modal segmentation can provide more abundant features and effectively locate the tumor location, especially in the fuzzy adhesion region of the tumor boundary. In addition, other multi-modal image segmentation methods were compared, and the results also prove that the multi-task learning method is suitable for multi-modal image segmentation and has achieved better segmentation results.
The atmospheric ozone layer plays an important role in the interaction of extraterrestrial environmental systems. Obtaining complete ozone data can help relevant scientific researchers better analyze the state of the ozone layer and predict its impact on global or local climate. Due to the operating orbit and other factors, the polar weather satellite may lose some data when collecting data. We use an encoder-decoder convolutional neural network to repair missing data, and use a discriminant network to judge the quality of the data. The model showed amazing performance on a test set that was not related to the training set.