
We introduce a novel framework, termed RMGANs, that merges Relativistic Generative Adversarial Networks (RGANs) with Margin Losses. This framework capitalizes on the strengths of both RGANs’ discriminators, which assess the realism of real and fake data, and the advantages of Margin Losses, well-known for their ability to create distinct class separation. We apply this combination to the problem of semi-supervised learning (SSL). Our study delves into the architecture of RMGANs, provides a mathematical analysis to assess the impact of margin utilization on RMGAN losses, and offers guidance on choosing hyper-parameters. We also conduct experiments across the MNIST and CIFAR-10 datasets. The empirical results clearly demonstrate RMGANs’ effectiveness in achieving higher accuracy compared to the state-of-the-art work (MarginGAN) in the SSL fashion.
Recent blockchain-based systems for managing credentials show advantages over paper-based procedures. However, issuing credentials with blockchain could conflict with current management rules and policies. One of the possible conflicts is the auditability. Most blockchain-based systems for credentials focus on security, efficiency and privacy while ignoring the auditability of the system. In this paper, we propose a new system IU-TransCert for issuing, verifying, and auditing academic credentials. The system uses a new data structure named the Auditable Merkle Tree that enables credential issuance and built-in auditing capabilities. The auditable data fields can be customized to meet regulations. Credentials are published to the blockchain in the root node of the Auditable Merkle Tree, allowing access for auditors while preserving privacy. The system provides automated and transparent auditing processes for educational authorities to independently verify credentials without involving issuers. We also present a prototype to demonstrate feasibility, and a security analysis to examine protections against threats. The analysis and discussion shows that the proposed system could enhance credential privacy, efficiency, integrity, and auditability across the university ecosystem.
In the engineering domain, representing real-world objects using a body of data, called a digital twin, which is frequently updated by "live" measurements, has shown various advantages over traditional modelling and simulation techniques. Consequently, urban planners have a strong interest in digital twin technology, since it provides them with a laboratory for experimenting with data before making far-reaching decisions. Realizing these decisions involves the work of professionals in the architecture, engineering and construction (AEC) domain who nowadays collaborate via the methodology of building information modeling (BIM). At the same time, the citizen plays an integral role both in the data acquisition phase, while also being a beneficiary of the improved resource management strategies. In this paper, we present a prototype for a "digital energy twin" platform we designed in cooperation with the city of Regensburg. We show how our extensible platform design can satisfy the various requirements of multiple user groups through a series of data processing solutions and visualizations, indicating valuable design and implementation guidelines for future projects. In particular, we focus on two example use cases concerning building electricity monitoring and BIM. By implementing a flexible data processing architecture we can involve citizens in the data acquisition process, meeting the demands of modern users regarding maximum transparency in the handling of their data.
The demand for high-resolution automotive radar operating in the 77GHz band is on the rise, especially as we approach the practical implementation of autonomous driving technology. With the increasing prevalence of in-vehicle Chirp Sequence (CS) radars in the future, there is a growing concern about the potential for broadband inter-radar interference. This interference poses a significant risk, potentially leading to a higher likelihood of undetected targets. To address this problem, various algorithm-based and learning-based schemes have been proposed to suppress the inter-radar interference, e.g., iterative threshold based zero suppression method and RNN (Recurrent Neural Network) based interference suppression method. However, they only demonstrate their effectiveness based on simulation results. In this paper, we conducted a multi-radar interference experiment with up to four interference sources and compared the performance of different schemes by using the collected real data.
Video retrieval is the process of finding specific video content in a large database. This is a crucial challenge in the age of digital multimedia. This article proposes a new approach to video retrieval using advanced deep learning models to extract features and perform retrieval tasks based on those features. Our method combines multiple feature extraction methods, including keyframe extraction, OpenAI CLIP [7] feature extraction, object detection, and automatic speech recognition (ASR). We use BERT [3] embeddings to encode these transcripts and store them in JSON and binary file formats. Our system achieves remarkable results in indexing and retrieving videos based on their visual, audio, textual, and contextual attributes. Our system can also retrieve videos based on either a single text description or multiple text descriptions of a sequence of events. We conducted extensive tests on diverse video data from Ho Chi Minh City AI Challenge 2023 competition organizers to validate the effectiveness of our approach. The results demonstrate that our proposed system is superior to other methods in terms of both retrieval accuracy and speed, making it highly suitable for real-time applications.
Long-range (LoRa) technology has become the mainstream of low-power wireless communication technology for long propagation distances. Developing mobile nodes with low-power technology is a crucial aspect; thus, energy efficiency (EE) has been considered in LoRa network designs. This area has been the subject of extensive research over the years, with a primary focus on enhancing signal transmission quality and minimizing noise components during communication. In this research, we propose a novel approach to optimize EE for LoRa networks, leveraging the principles of Q-Learning. Our proposed method entails dynamically adjusting the transmission power based on the evolving state of the network environment, ultimately striving to attain the highest EE possibility. To capture the optimality of our learned transmission power, the Gradient Descent method is utilized to point out the progress. Numerical results reveal a good convergence of learning-based and gradient-based approaches and describe the trade-off between performance and inference time. Our proposed algorithm showcases remarkable outcomes, achieving an impressive 91% of the gradient-based solution while reducing the computation time by an average of 30%. These findings emphasize the substantial potential of reinforcement learning techniques, particularly Q-Learning for LoRa networks.
In an era defined by the proliferation of digital content and a growing reliance on video data, the need for effective anomaly detection systems has never been more pressing. This paper introduces a sophisticated system architecture designed to address the complex challenges associated with acquiring, processing, and presenting anomaly videos. At its core, our architecture prioritizes openness and modularity, allowing for seamless upgrades and customization. This approach ensures adaptability to evolving technology trends and user preferences. We emphasize the crucial aspects of component interfacing and user interaction, highlighting the integration of feedback mechanisms for ongoing system refinement. Additionally, we contribute significantly to the research community by extending an established anomaly event dataset using proven methods and techniques. This extension enhances the dataset’s breadth and depth, providing a valuable resource for training and evaluating anomaly event retrieval systems. Our paper presents a forward-looking system architecture poised to meet the demands of anomaly video detection while also enriching the available resources for anomaly event research. Subsequent sections will delve into architecture components and methodologies, showcasing its potential to revolutionize modern anomaly detection systems.
The non-fungible token (NFT) market has experienced significant growth in recent years. Similar to real estate and artwork, NFT is illiquid and it does not have a market price. Thus, there arises a need for a comprehensive scoring system to analyze the value of NFTs. Using a single score is often inadequate, as NFT value is influenced by numerous factors and is assessed based on personal objectives. Addressing the issue, this study explores critical factors influencing NFT value and develops three scores to rank NFTs including rarity, return on investment (ROI), and reputation score. The scores serve to evaluate both NFTs and NFT collections. (i) The rarity score evaluates the level of scarcity, and it applies to NFTs within a collection. (ii) The reputation score evaluates the interest of communities in NFT projects, considering the number of followers and the interaction rate; the score is for NFT collections. (iii) The ROI score assesses the profit generated by NFTs; it applies to both NFTs and NFT collections. Our empirical results show that a well-distributed rarity score of NFTs enhances their demand and profit; high-rarity NFTs are typically associated with a low number of sale transactions. NFT collections with high ROI and high reputation generally draw more attention and yield a positive return. In addition to the three scores, this paper presents a system to collect and process a huge amount of blockchain and social network data for NFT evaluation.
Multimodal learning tries to increase generalization performance by leveraging information from several data modalities. However, effectively integrating such multi-modality data remains a complex endeavor, particularly when dealing with incomplete data. In real-world scenarios, modalities are often not entirely absent but rather incomplete due to various external or internal factors. For instance, audio data can be corrupted by noise, and text data may suffer from inaccuracies stemming from automatic speech recognition errors. To address these challenges, we propose a novel framework for incomplete multimodal learning in conversational contexts, named GAT-FP. Our GAT-FP model incorporates two key graph neural network-based modules, namely, "Feature Propagation" and "Graph Attention Network". These modules are designed to estimate missing features and discern the significance of interactions among incomplete feature nodes within the graph structure. To validate the effectiveness of our approach, we conduct extensive experiments on two well-established benchmark multimodal conversational datasets: IEMOCAP and MELD. The experimental results demonstrate that our GAT-FP model surpasses existing state-of-the-art methods in the realm of incomplete multimodal learning.
Colonoscopy is widely acknowledged as the most efficient screening method for detecting colorectal cancer and its early stages, such as polyps. However, the procedure faces challenges with high miss rates due to the heterogeneity of polyps and the dependence on individual observers. Therefore, several deep learning systems have been proposed considering the criticality of polyp detection and segmentation in clinical practices. While existing approaches have shown advancements in their results, they still possess important limitations. Convolutional Neural Network - based methods have a restricted ability to leverage long-range semantic dependencies. On the other hand, transformer-based methods struggle to learn the local relationships among pixels. To address this issue, we introduce ConvTransNet, a novel deep neural network that combines the hierarchical representation of vision transformers with comprehensive features extracted from a convolutional backbone. In particular, we leverage the features extracted from two powerful backbones, ConvNeXt as CNN-based and Dual Attention Vision Transformer as transformer-based. By incorporating multi-stage features through residual blocks, ConvTransNet effectively captures both global and local relationships within the image. Through extensive experiments, ConvTransNet demonstrates impressive performance on the Kvasir-SEG dataset by achieving a Dice coefficient of 0.928 and an IOU score of 0.882. Additionally, when compared to previous methods on various datasets, ConvTransNet consistently achieves competitive results, showcasing its effectiveness and potential.
Along with the success of deep learning are extensive models and huge amounts of data. Therefore, efficiency in deep learning is one of the most concerning problems. Many methods have been proposed to reduce the complexity of models and have achieved promising results. In this paper, we look at this problem from the data perspective. By leveraging the strengths of Tensor methods in data processing as well as the efficiency of deep learning models, we aim to reduce costs in many aspects such as storage and computation. We present a data-driven deep learning approach for high-dimensional data classification problems. Specifically, we use Tucker Decomposition, a tensor decomposition method, to factorize the large, complex structure raw data into small factors and use them as the input of lightweight deep model architecture. We evaluate our approach on high-dimensional, complex datasets such as video classification on the Jester dataset and 3D object classification on the ModelNet dataset. Our proposal achieves competitive results at a reasonable cost.
Our study focuses on the potential for modifications of Inception-like architecture within the electrocardiogram (ECG) domain. To this end, we introduce IncepSE, a novel network characterized by strategic architectural incorporation that leverages the strengths of both InceptionTime and channel attention mechanisms. Furthermore, we propose a training setup that employs stabilization techniques that are aimed at tackling the formidable challenges of severe imbalance dataset PTB-XL and gradient corruption. By this means, we manage to set a new height for deep learning model in a supervised learning manner across the majority of tasks. Our model consistently surpasses InceptionTime by substantial margins compared to other state-of-the-arts in this domain, noticeably 0.013 AUROC score improvement in the "all" task, while also mitigating the inherent dataset fluctuations during training.
Evolutionary Reinforcement Learning (ERL) combines the sample-efficiency property of Reinforcement Learning and exploration capabilities from the population-based search of Evolutionary Computation. These methods have shown promising performance on many continuous control tasks. However, one could observe the instability that may occur from such methods. Several works have shown that the experiences coming from the population individuals lead the state distribution shift in the RL policy updating process. A vanilla remedy method has been proposed to alleviate this issue by separating the experience transitions into two distinct replay buffers for the RL policy and the population and mixing the samples from the two buffers with a fixed ratio to update the RL policy. The effectiveness of this approach has been shown empirically on an ERL method where Evolution Strategies (ES) assists an external RL agent. Nevertheless, there has not been any thorough investigation on Genetic Algorithm (GA) based ERL to understand how this method performs on these ERL approaches. In this paper, we analyze the influence of off-policy data coming from the GA population to the RL policy and how the mixing method performs on a state-of-the-art ERL method, namely Proximal Distilled Evolutionary Reinforcement Learning (PDERL).
Differential expression gene (DEG) analysis of transcriptomic data allows for a comprehensive examination of the regulation in gene expression profiles related to specific biological states. The result of this analysis typically consists of an extensive record of genes that display varying levels of expression among two or more groups. A portion of these genes with altered expression could potentially function as candidate biomarkers, chosen through either existing biological insights or data-driven techniques. In diagnosing sepsis, a life-threatening health problem, our work proposes a novel approach using immune-related gene data to identify the optimal gene combination as signature biomarkers to improve the diagnosis performance. Our proposed method involves sequential gene selection procedures, including the DEG analysis and the machine learning-based importance assessment, and a Recursive Feature Elimination (RFE) process supported by Principal Component Analysis (PCA). The selected gene combination, which consists of twelve immune-related genes, shows remarkable cross-validation results with an accuracy of 99.35%, AUC score of 99.56%, Sensitivity and a Specificity of 99.44% and 90.00%, respectively. Besides, the proposed 12 gene markers combined with the XGBoost algorithm were also tested in three individual cohorts with appropriately significant results, demonstrating the effectiveness of our developed method in different cohorts and the reliability of the proposed gene selection procedure.
The advent of smart homes has revolutionized residential living, integrating advanced technologies and intelligent devices to create secure, comfortable, and efficient environments. However, this integration of diverse smart devices has brought significant cybersecurity challenges. Detecting and analyzing abnormal network packets have become paramount, signifying potential intrusions, malicious activities, or system errors and ensuring the security and stability of smart home systems. Machine learning techniques, such as Decision Trees, Support Vector Machines (SVM), Convolutional Neural Networks (CNN), K-Nearest Neighbors (KNN), Recurrent Neural Networks (RNN), and Random Forests, have shown promise in addressing these challenges. However, most research has concentrated on anomaly detection rather than malicious activity in smart homes. The vast datasets collected from various scenarios pose methodological and algorithmic challenges for applying machine learning techniques. To fill these research gaps, our study introduces traditional machine-learning methods for detecting abnormal network packets in smart homes using the IoT-23 dataset. It involves preprocessing the dataset, extracting relevant features, and training various machine learning models. The correlation matrix helps validate the feature selection of the best models based on performance metrics like precision, F1-score, recall, accuracy ratio, training score, and training time cost. Additionally, the study classifies 12 types of malicious malware across different machine learning models, considering performance within the context of smart home devices. This study implements real-time anomaly detection on the Raspberry Pi using packet captures and Zeek flowmeter methods. The findings contribute insights into models suitable for smart home security. In addition, our research enhances the understanding and application of machine learning methods for bolstering security in smart homes.
Nowadays, security and safety issues are very complex and tend to be related to criminals using weapons to commit crimes, posing many potential risks to society. Recently, Deep Learning models have been researched and applied to Computer Vision problems. This article focuses on training a weapon detection system using the YOLO - 5, 7, 8 model and the Swin Transformer model combining Mask R-CNN, Cascade Mask R-CNN, Mask RepPoints V2 for detecting weapon in the images and providing early warning. Object recognition solution The research focuses on 3 main objects of weapons: Pistols, Rifles, and Knives. Due to limited available data, the WeaponData_VN dataset was built and described.
It is undeniable that high-quality data plays a crucial role in achieving good outcomes. However, obtaining such a dataset is always challenging, particularly in multi-label classification, where a fully labeled dataset is required in traditional approaches. This challenge has led to the emergence of several effective learning techniques called single positive multi-label learning (SPML), which utilize multi-label training images annotated with only one single positive label. In our work, we propose an effective method that improves the cutting-edge BoostLU baseline by leveraging efficient label-weighted loss and prior knowledge based regularization strategies. Firstly, we present a novel approach for reweighting the contribution of each label to the total loss, giving higher weight to the true positive label while eliminating unreliable pseudo negative labels by assigning them zero weight. This is achieved through the exploration of large per-label losses. Additionally, we introduce an auxiliary loss to regulate the expected value of positive labels, encouraging the model to predict a reasonable number of positive labels per image based on prior knowledge about the dataset. Experimental results across several benchmark datasets showcase the superior performance of our method compared to both the baseline and other state-of-the-art methods.
Event retrieval from large collections of TV news videos is crucial for efficient information access, enabling researchers, journalists, and the general public to quickly locate and analyze relevant content amidst the vast sea of news coverage, facilitating informed decision-making and a comprehensive understanding of significant events. This paper presents an overview of the AI-driven video retrieval task in Ho Chi Minh City AI Challenge 2023. The competition draws inspiration from internationally recognized competitions, namely the Video Browser Showdown (VBS) and the Lifelog Search Challenge (LSC). Participants are tasked with developing AI models to retrieve specific video segments from a diverse dataset from reputable news channels. The dataset comprises a vast collection of videos, keyframes, object detections, CLIP features, and metadata. It is divided into three packs with a total of 1,270 videos, spanning approximately 360 hours of content. The challenge comprises two groups. Group A is open to students, researchers, and practitioners in artificial intelligence and information retrieval, emphasizing substantial knowledge and experience. Group B is tailored for high school students, focusing on nurturing interest, learning, and engagement among the next generation of AI enthusiasts. The wide variation in the content of queries challenged participants to demonstrate their adaptability and creativity in effectively retrieving diverse events from the extensive TV news video dataset. The winning teams showcased promising solutions by effectively harnessing artificial intelligence and information retrieval techniques to excel in event retrieval from a vast collection of TV news videos.
Skeleton-based action recognition is a challenging problem due to the high dimensionality and noisy nature of skeleton data. Graph convolution networks (GCNs), which use graph topology to extract representative features, have been effective for skeleton-based action recognition in recent years. However, effectively learning and aggregating topology information is challenging problem. In this work, we propose a strategy to construct topology representation to skeleton-based action recognition that combines language knowledge to learn the topology. Specifically, borrows the idea from Language Knowledge-Assisted Representation Learning (LAGCN) [20], which uses a large-scale language model (LLM) to learn a priori global relationship (GPR) topology that captures the global structural relationships between the joints and a priori category relationship (CPR) topology between nodes in the skeleton graph to capture the category-specific relationships between the joints. We propose to apply the GPR topology as a prior topology, provide significant momentum to learn the model along with the CPR which is used to learn the class-distinguishable features into the Channel-wise Topology Refinement Graph Convolution (CTRGCN) [4]. The proposed approach is evaluated on the NTU RGB+D and NW-UCLA datasets. The results show that the proposed approach achieves promising results with 96.76% on the NW-UCLA dataset and 97% on the NTU dataset with the cross-view benchmark along with 92.8% on NTU cross-subject benchmark.