
Image deraining is an image restoration task that aims at decomposing a rainy image into the clear scene layer and the rain layer. Most existing deraining methods use supervised learning and training on synthetic rainy-clear pairs under the assumption that the only difference between rain and non-rain images is the rain effect. Since scenes constantly change due to vehicle movement and other activities, this assumption nearly fails when applied to traffic data. In this work, we introduce a de-raining framework using contrastive learning based on Vision Transformer. The proposed framework utilizes weakly paired images, in which the difference between images is not only caused by rain. In addition, we collect a real-world deraining dataset from traffic surveillance cameras. This is the first real-world deraining dataset in the traffic domain. The experiments show that our method achieves competitive performance with other existing methods while taking less time in the inference process.
This paper presents a new hybrid deep-learning model, U-ReACH, for detecting ships over satellite images. Specifically, the new U-ReACH model combines two well-known deep learning models, Unet and ResNet. However, applying a deep-learning model straight into satellite imagery-related tasks usually leads to poor performance due to data imbalance. An Auxiliary Classification Head is added at the end of the feature extraction stage to overcome the issue. Besides, a new loss function is used to handle the problem of imbalanced data. The experiment results were executed on the Airbus Ship dataset, including 192,556 satellite images, which showed the promise of the new approach.
In the digital age, identifying forgery in images has become a significant challenge for image analysis systems and applications. In this study, we approach this issue from a different perspective compared to existing methods for this problem, proposing an attack method to forgery detection models. Instead of the traditional approach of randomly modifying images, we combine a range of diverse strategies, ranging from the utilization of diffusion models to the implementation of camera trace erasing techniques aimed at eliminating the remnants left by forgers, consequently leading to the deception of forensic methods. This is not a simple problem, it requires a strategy to choose which models and in what order is most appropriate to both attack effectively and ensure that image quality is not reduced. Experimental results demonstrate that our method not only surpasses modern counter-forensics methods but also preserves the naturalness and semantics of the processed images. In this way, our research provides a valuable contribution to understanding and responding to the threats faced by modern image analysis systems.
In the digital age, the authenticity of photographs is increasingly questioned due to the rise of sophisticated AI and human manipulation techniques. This paper presents the development of an AI-assisted system designed to aid users in identifying and verifying the genuineness of photos. Unlike fully automated solutions, our approach emphasizes user empowerment, providing tools to enhance human expertise in photo authentication. The system integrates advanced algorithms to detect potential manipulations and localize regions of interest, directing users’ attention to specific areas that warrant closer examination. Users can conduct more efficient and effective authenticity checks without scrutinizing the entire image by focusing on these targeted regions. Our goal is to augment, not replace, the critical role of human judgment in verifying photo authenticity. The system’s design and implementation are discussed, along with case studies demonstrating its practical application in various scenarios. This research aims to contribute to the growing field of digital forensics by providing a robust, user-friendly tool that bridges the gap between AI capabilities and human expertise in photo verification.
Brain tumors pose a significant medical challenge, necessitating precise and rapid diagnosis for effective treatment and improved patient outcomes. This paper introduces knowledge distillation, which has the potential to revolutionize brain tumor diagnosis by enabling early identification from medical imaging data. Using a sophisticated teacher’ model to capture intricate patterns, we distill this knowledge into a more efficient "student’ model, aiming for comparable accuracy with reduced memory usage and improved inference times. Our method, based on a dataset of 357 MRI scans, demonstrated the potential of knowledge distillation in brain tumor diagnosis, offering a promising avenue for advancing patient care. The proposed model serves as a vital tool for healthcare practitioners, providing accurate and efficient support in detecting brain tumors and contributing to advancements in healthcare technology. The evaluation results indicate the effectiveness of our technique, achieving an impressive accuracy of 98.10
Text-based Person Search (TBPS) has emerged as a significant research topic in information retrieval domain, garnering considerable attention and development thanks to its wide potential applications. While some existing research on TBPS has achieved notable milestones in the English language, many proposed models often exhibit low efficiency when extended to other low-resource languages, such as Vietnamese due to the difference of language characteristics. In this paper for the first time, the role of language characteristics and data enrichment on the effectiveness of TBPS is fully assessed. To this end, two state of the art models for TBPS in English ViTAA and IRRA are chosen and adapted to TBPS in Vietnamese. Two benchmark datasets CUHK-PEDES for English and 300VnPersonsearch for Vietnamese, are employed in our experiments. Experimental results show that evaluating systems within the same domain yields significantly higher results than evaluating across different domains, increasing by 30.66% for the ViTAA model and 22.03% for the IRRA model at R@1. The obtained results provide valuable insights into the influence of language characteristics and data enrichment on the effectiveness of TBPS models.
Evolutionary algorithms have demonstrated effectiveness in neural architecture search (NAS). However, previous studies mainly focus on the clean performance of evolutionary neural architecture search algorithms (ENAS), neglecting attention to their robustness against adversarial attacks. Additionally, the training-free performance metrics show their promise in NAS by offering a means of estimating network performance at trivial costs. This paper comprehensively evaluates the robustness of ENAS algorithms under various settings. Specifically, we implement both single-objective ENAS and multi-objective ENAS algorithms. Different training-based and training-free metrics are employed for each algorithm. The results obtained from extensive experiments conducted on well-known NAS benchmarks, including NAS-Bench-201, NAS-Bench-Suite-Zero, and the Robustness dataset, yield two insightful findings. First, utilizing training-free performance metrics for finding robust network architectures is more efficient and promising compared to utilizing training-based ones. Second, conducting NAS runs with multi-objective ENAS allows for figuring out multiple networks that exhibit diversity in both robustness and characteristics.
The websites’ privacy policies reflect how their users’ personal data is collected, processed, stored, and used. This paper proposes a risk analysis method to assist users in understanding and identifying potential privacy risks in these policies. The method consists of two phases. First, we propose a transfer learning process to refine an existing large language model (LLM) to enable it to understand and answer legal questions effectively. Second, we introduce a retrieval-augmented generation (RAG) technique applied to the refined LLM to analyze risks in privacy policies. Additionally, we define a checklist to assess the compliance of privacy policies with legal regulations. This checklist is mapped to 20 queries, which are used to build prompts for the refined LLM. This allows our refined LLM to provide explanatory answers to users regarding the overall compliance level of privacy policies with relevant regulations, stating the potential risks in these privacy policies. The transfer learning process has been conducted on two existing LLMs. Experiments were performed on 200 random legal questions to compare these fine-tuned LLMs with the originals, demonstrating the efficiency of our proposed method in terms of the cosine similarity index and the F1 score.
In ovarian cancer diagnosis and treatment, ultrasound is widely used as a screening method thanks to its low cost. Accurately segmenting ovarian tumors from ultrasound images is an important step for further investigation. However, due to the heterogeneity of tumors and the low quality of ultrasound images, the segmentation of ovarian tumors is a challenging task. This paper presents a method for ovarian tumor segmentation from ultrasound images based on Segment Anything Model (SAM). - a transformer based model trained on large-scale datasets. SAM has been proved to be effective for many segmentation tasks. However, SAM traditionally optimizes loss functions based on regions without considering structural similarity constraints between the actual and predicted regions. To enhance model learning, we incorporate IoU, SSIM, and Focal loss functions in SAM model. Furthermore, we utilize two prompts methods: manual prompts and automatic prompts based on the detection results of YOLOv5. Experimental results on OTU2D dataset show that the proposed method outperforms many state of the art methods and the baseline model with 95.12% of Precision when using manual prompts and 91.39% for automatic prompts. Additionally, to evaluate the generalization of the proposed method, we collect a new dataset named OvaTUS dataset and perform cross dataset evaluation.
Scene understanding at the instance level is an essential task in computer vision to support modern Advanced Driver Assistance Systems. Solutions have been proposed with abundant annotated training data. However, the annotation at the instance level is high-cost due to huge manual efforts. In this work, we solve this problem by introducing InstSynth, an advanced framework leveraging instance-wise annotations as conditions to enrich the training data. Existing methods focused on semantic segmentation via using prompts to synthesize image-annotation pairs, facing an unrealistic manner. Our proposals utilize the strength of such large generative models to synthesize instance data with prompt-guided and mask-based mechanisms to boost the performance of the instance-level scene understanding models. We empirically improve the performance of the latest instance segmentation architectures of FastInst and OneFormer by 14.49% and 11.59% AP, respectively, evaluated on the Cityscapes benchmark. Accordingly, we construct an instance-level synthesized dataset, dubbed IS-Cityscapes, with over a 4× larger number of instances in comparison with the vanilla Cityscapes. Code can be found at https://github.com/danhntd/InstSynth.
In digital age, the widespread accessibility of information dissemination has inadvertently accelerated the spread of misinformation, highlighting the critical necessity for robust factchecking systems, particularly for low-resourced languages. Our study propose the ViNSV system, a pioneering approach tailored for the Vietnamese language, utilizing the ISE-DSC01 dataset. It employs a novel multi-stage text ranking method, integrating advanced natural language processing techniques such as BM25, Sentence-BERT for in-depth semantic analysis, and XLM-R for comprehensive factual verification. Our extensive evaluation reveals that ViNSV achieves a breakthrough in accuracy, with a Strict Accuracy of 76.33%, thereby setting a promising benchmark in the field of Vietnamese Fact Extraction and Verification. The ViNSV system not only enhances the reliability of digital content but also offers a scalable model for adaptation to other languages with similar challenges, representing a significant contribution to ensuring global information integrity.
Few-shot object detection (FSOD) models detect new objects with scarce training samples on novel data using pre-trained models on large-scale base data. Recent FSOD works use a stable diffusion model to generate new synthetic objects for novel classes. However, their synthetic objects lack diversity due to fixed sizes, shapes, or uncontrollable occlusions. Besides, most previous FSOD methods forget to effectively exploit base classes, which cannot leverage data diversity to create distinguished features and enhance the generalization.A recent method [1] exploits and shows potential results while it provides more base classes. The simple method only uses fake classes, which leads to the ambiguity between the fake and real base classes and then reduces the performance when the model gets more data. To solve the above issues, we introduce HBC using Fine-grained Data generation using Diffusion (FDD) and Three-Stage Training (TST) to create and leverage large-scale data. Specifically, FDD constructs hierarchical information from the base and extended real classes in ImageNet to find new fine-grained classes. The new objects of fine-grained classes that are added to the base image using a controllable diffusion model have different sizes, shapes, and positions based on the base bounding boxes. The method creates real fine-grained classes with rich synthetic samples and minimizes the ambiguity between these classes and the original base classes. Then, TST uses the generated data to enhance the distinguished ability of features and improve the detector without negatively affecting the performance of base classes. In the end, we provide comprehensive results to demonstrate the effectiveness of our proposed method in FSOD settings. Our method improves averaged 4.5% and 2.8% nAP50 compared to the baseline and previous work, respectively.
Currently, the problem of automatic question answering (Q/A) with applications in chatbot services has become essential, addressing various issues in daily life. There are many approaches to solving this problem, but each presents its own challenges. The generative AI approach has recently gained widespread attention due to its ability to provide smooth, human-like responses. However, it has the drawback of producing coherent-sounding but inaccurate or fabricated content, known as "hallucinations". In this study, we aim to find a solution for the question-answering task in the context of the Vietnamese economy. We propose the Generative Vietnamese Economy Chatbot (GVEC), based on the Vietnam Economy Information Database (VEID) retrieved from VnEconomy systems. Our proposition is to apply Retrieval-Augmented Generation (RAG) to various LLM systems and test them on our specially designed benchmark, called the Vietnamese Numeric Economy Information Question/Answer Datasets (VNEIQAD), generated by our Question/Answers Generating with Numerical Information (QAGwNI) algorithm. This will help better evaluate the solution in the economic context.
Massive Open Online Courses (MOOCs) have become a pioneer in providing access to knowledge for everyone around the world. These courses transcend geographical and linguistic barriers, allowing anyone, anywhere, to learn and enhance their knowledge. Although previous research on recommendation models has shown promising results in course recommendations, building such systems for MOOC platforms still presents significant challenges. Implicit user feedback often lacks explicit negative signals, making it difficult to accurately model user preferences. Additionally, the diversity and sparsity of data, especially for new or niche courses, further hinder traditional methods. This creates a pressing need for new and improved solutions in this field. In this study, we propose H-BERT4Rec, an enhancement of the BERT4Rec model, which leverages Heterogeneous Information Networks (HINs) to address these challenges. HINs integrate diverse data sources and capture complex relationships between entities such as courses, videos, and users. This not only enhances the understanding of user preferences but also strengthens the ability to recommend suitable courses. H-BERT4Rec utilizes Heterogeneous Network Embedding(HNE)-node embedding generation method that leverages HIN to create Pre-train Embeddings, and then improves the BERT4Rec architecture, leading to more accurate and personalized recommendations. We conduct experiments on a real-world MOOC dataset to demonstrate the superior performance of H-BERT4Rec compared to baseline models, achieving an improvement of up to 55,04%. This study contributes a promising new approach for personalized course recommendations in MOOCs, enhancing the learning experience for millions of learners worldwide. This improvement not only promises significant benefits for learners but also opens up new directions for research and development in the field of online education.
Despite ongoing research efforts in recent years, machine translation in low-resource settings remains challenging. This paper explores recent methods to enhance low-resource machine translation (LoResMT), especially for both-side low-resource language pairs like the Lao-Vietnamese language pair. In this case, our experiments highlight the importance of translation network’s hyper-parameter optimization for LoResMT. Back-translation methods and fine-tuning multilingual pre-trained models also gain positive results. With the small size of the Lao-Vietnamese dataset from the VLSP 2023 MT challenge shared task, our hyper-parameter optimization achieved 24.61 BLEU score, an improvement of 22 BLEU score over the default model settings. Using back-translation method boosts the performance to 27.79 BLEU score. Additionally, fine-tuning the pre-trained model mT5 on the training set yields a result of 28.05 BLEU score. The best achieved BLEU score of our Lao-Vietnamese translation system is comparable to the best score announced in the related works with less data.
Video retrieval has become an important task in computer vision, with video content uploaded to the Internet every hour. Along with retrieving the relevant visual content, users may also want to perform several post-processing steps such as visual editing, understanding or video summarizing. However, to our knowledge, there is no such integrated system that enables users to perform downstream visual understanding and editing tasks via text prompts. In this work, we propose VISA framework, which combines a visual programming module with a video search system. Specifically, our interactive framework offers fundamental video retrieval with semantic search, text search and audio search with descriptive inputs summarized by a large language model (LLM). After obtaining the video frame results, users can provide natural language instructions as guidance for image understanding and editing tasks. Having the in-context learning capability of LLMs, our visual programming module generates high-level and interpretable pseudocodes from the given instructions. The corresponding Python programs are then executed to achieve the desired results. We evaluate our VISA framework on the 2023 Ho Chi Minh City AI Challenge dataset and the image editing component on the MagicBrush benchmark.
The increasing complexity of software systems has necessitated more sophisticated security measures, particularly in the domain of vulnerability detection. Traditional machine learning (ML) and deep learning (DL) techniques often fall short when source code is treated merely as text, prompting a shift toward graph learning methods that leverage specific graph representations of code to enhance detection capabilities. These representations, including Abstract Syntax Tree (AST), Control Flow Graph (CFG), Data Flow Graph (DFG), Program Dependence Graph (PDG), and Code Property Graph (CPG), encapsulate the structural and semantic intricacies of programming code, offering a robust framework for identifying vulnerabilities. Despite advances in this field, a comprehensive understanding of the impact that each graph representation has on the effectiveness of vulnerability detection is still lacking. Our paper introduces a general architecture for a graph-based vulnerability detection system and conducts empirical studies on two real-world datasets, BigVul and FUNDED. This research systematically assesses how variations in graph representations—AST, CFG, DFG, PDG, and CPG—affect the efficacy of software vulnerability detection, providing pivotal insights that could guide future research and enhance practical applications in cybersecurity.
Class activation maps (CAMs) are a crucial tool for visualizing the regions within an image that a classifier uses to identify a given object category. Recent studies have leveraged CAMs for weakly supervised object localization and have achieved promising results. However, they are limited to scenarios with a single object class for each image. This paper introduces a novel framework that extends CAMs for weakly supervised object detection. Experiments on the PASCAL VOC 2007 and 2012 datasets with two backbones, ResNet and VGG, have demonstrated that our method achieves promising results.
The exponential growth of multimodal content across social media platforms, comprising text, images, audio, and video, has catalyzed substantial interest in artificial intelligence, particularly in multi-modal sentiment analysis (MSA). This study presents a comprehensive survey of 30 research papers published between 2020 and 2024 by eminent publishers such as Elsevier, ACM, IEEE, Springer, and others indexed in Google Scholar. Our analysis primarily focuses on exploring multimodal fusion techniques and features, with specific emphasis on the integration of text and image data. Additionally, the article offers an overview of the evolution, definition, and historical context of MSA. It delves into the current challenges and potential advantages of MSA, investigating recent datasets and sophisticated models. Furthermore, the study provides insights into prospective research directions. Notably, this review offers valuable recommendations for advancing research and developing more robust MSA models, thus serving as a valuable resource for both academic and industry researchers engaged in this burgeoning field.
This paper presents our work in building a Vietnamese dataset for command and speaker recognition problems. We built a website that allows users of mobile devices or personal computers to be able to provide their voice samples easily. We collected more than fifteen thousand utterances of nineteen Vietnamese commands from more than two hundred volunteers. The commands are primarily used for applications on edge devices that interact with users via voice. Models of command recognition are often much smaller than those of speech recognition and, thus, are preferable on edge devices. We then fine-tuned and evaluated recent methods of recognizing voice commands and verifying speakers with our dataset. The selected methods rely on a moderately deep learning model that tries to vectorize input speech signals represented by their spectrograms and performs classification on the vectors with a multilayer perceptron network. The experimental results showed that the models fit very well on the dataset and are promising for applying to command and speaker verification in Vietnamese. The fine-tuned model of Vietnamese command recognition achieved an F1-score of 99%. The fine-tuned model of Vietnamese speaker verification had equal error rates of 4.49% and 2.27% with text-independent and text-dependent test sets respectively. Our solution can be applied to enhance security for smart interaction systems that need to authenticate users for critical functionalities.