Variational autoencoder (VAE)-based frameworks possess a natural advantage in modeling the shared and private information inherent in multimodal data. However, current models focus on improving the quality of shared representations from the reconstruction perspective, lacking explicit mechanisms to model their underlying semantic structure. In this paper, we propose the multimodal Gaussian mixture variational autoencoder with consistency regularizations, which introduces a Gaussian mixture prior over the shared latent space to enhance its semantic structure and encourage the formation of cluster-aware latent representations. To address the cross-modal inconsistency problem under missing modality conditions, we propose a cluster-guided regularization strategy that enforces the cross-modal consistency using the pseudo-category labels from unsupervised clustering. Additionally, we design a self-supervised contrastive regularization strategy to align semantically similar representations across modalities. Extensive experiments on MNIST-SVHN and MNIST-CDCB datasets demonstrate that our method significantly outperforms prior state-of-the-art models in generation, classification, and retrieval tasks.
Depth completion aims to generate dense and accurate depth maps from sparse depth maps. Existing RGB-guided approaches fail to handle depth discontinuities at object boundaries and are prone to texture confusion. To address these issues, this paper proposes a Semantic-Aware based Depth Completion Network (SADCN), which enhances depth reconstruction and geometric consistency through a synergistic mechanism of semantic prior guidance and precise cross-modal feature alignment. SADCN introduces two innovative modules: the Semantic Prior Module (SPM), which extracts high-level semantic representations for depth completion, and the Cross-Modal Alignment Fusion module (CAF), which aligns semantic and RGB-D information to generate semantic-aware features. SADCN is evaluated on the KITTI and NYU Depth V2 datasets, and the results are compared against CSPN, LRRU, and DeepLiDAR. Experimental results demonstrate that SADCN outperforms these approaches, while preserving sharp object boundaries and structural integrity in complex scenes.
Reinforcement learning for fermentation control is often constrained by inefficient online exploration and delayed state acquisition. To address these challenges, this paper proposes a collaborative offline-online reinforcement learning framework for fermentation control (COORL-FC). The framework integrates offline policy pre-training via Conservative Q-Learning, simulator-assisted policy validation and hyperparameter selection, and online fine-tuning using the Asynchronous Advantage Actor-Critic algorithm. In addition, a near-infrared spectroscopy-based soft sensor is introduced to provide real-time estimates of key biochemical variables, thereby enabling closed-loop decision-making. The proposed method is evaluated on a high-cell-density Saccharomyces cerevisiae fermentation process for glucose-feeding optimization. Experimental results show that the proposed approach improves the final OD600 by 37.6
Screening potential formulations from numerous medicines is a critically important task in the field of Traditional Chinese Medicine (TCM), which is typically completed by experienced physicians based on the patient's symptoms and the properties of different Medicines. However, the complex mechanism of action of TCM makes this task very challenging. To overcome these hurdles, research on TCM formulations has shifted towards target-based methods inspired by evaluation methods used in Western medicine, aiming for gaining a deeper understanding of TCM and its potential efficacy. Nevertheless, TCM has more action targets compared to Western medicine, which often leads to computational bottlenecks. Traditional machine learning-based methods can significantly reduce computational time, but they are less interpretable and more prone to overfitting. To this end, this paper proposes an efficient and accurate framework for screening TCM prescriptions. Specifically, we start by identifying the key targets for the specific disease and analyzing the interaction relationships between these targets. We then utilize a Graph Convolutional Network to extract community relationships between the targets and build a trustworthy hypergraph based on this information. Using this structure, we obtain a prescription representation and train a prescription evaluation network to learn the merits of existing TCM prescriptions. Finally, we evaluate our proposed method on two common chronic diseases in clinical practice, Parkinson's disease and chronic gastritis, and the results demonstrate the effectiveness of proposed method in TCM prescription screening and evaluation.
Despite the significant advances of Convolutional neural networks (CNNs) and Transformers in image deraining, they either suffer from limited receptive fields or incur quadratic complexity, leading to an imbalance between performance and efficiency. Recently, state space models (SSMs) have demonstrated significant potential in modeling long-range dependencies while maintaining linear complexity. However, existing Mamba-based approaches lack the exploration of useful complementary information from multiple image scales, which could be beneficial for facilitating rain removal. In this paper, we propose an effective multi-scale state-space model-based framework (MS-RainMamba) to explore richer scale-space information for better image deraining. Specifically, we design a local-enhanced state space module to better aggregate rich local and global information. In contrast to existing methods that adopt fixed-scale scanning for feature extraction, we develop a multi-scale hierarchical 2D scanning technique to better help image restoration. Experimental results on six benchmarks show that the proposed method performs favorably against state-of-the-art models.
Classifying noisy data streams presents significant challenges due to the intertwined effects of attribute noise and label noise, which are rarely addressed simultaneously in existing methods. While concept drift and noise handling are critical in data stream learning, prior works often neglect the joint impact of attribute and label noise, leading to degraded model performance. In this paper, we propose ECIFC (Ensemble Classifier via Integration of Filter and Correction), a simple yet robust framework that tackles both attribute and label noise in a unified manner. Unlike complex approaches, ECIFC integrates noise filtering, label correction, and ensemble learning into a single framework with theoretical guarantees. To demonstrate the effectiveness of our method, we simply use the EM algorithm to estimate attribute and label log probabilities, use K-means to automatically determine noise thresholds, and derive ensemble classifier weights. Crucially, we rigorously prove the convergence of ECIFC, ensuring its stability in dynamic environments. Experiments on synthetic and real-world data streams demonstrate that ECIFC effectively identifies and mitigates both types of noise, achieving superior classification accuracy compared to state-of-the-art methods. The simplicity of our framework, combined with its theoretical foundation and empirical efficacy, makes ECIFC a practical and scalable solution for real-world noisy data stream classification.
Generative Question Answering (GQA) has spread across various industries. However, the potential of GQA during Question Answering (QA) interactions remains underutilized. Consequently, this paper introduces AugSBertChat, a GQA model that integrates user feedback for enhanced performance and utility. This method can be divided into two parts: predicting the probability of liking a reply and generating the reply. In order to make better use of user feedback to improve the quality of replies, we first formulate the QA task as the Semantic Text Similarity (STS) task, using Sentence-RoBERTa to obtain the similarity of QA pairs in high-dimensional space. In particular, in-domain symmetric semantic search is used to enhance our model performance. Subsequently, we construct some prompts that are more suitable for XiaoAi's QA scene, and employ P-tuning v2 to efficiently fine-tune the ChatGLM-6B parameters. Finally, we conducted experiments in the NLPCC-2023-Shared-Task-9 User Feedback Prediction and Response Generation (UFPRG) and achieved good results, placing third among all teams, which demonstrates the effectiveness of our proposed method.
Fine-grained visual classification (FGVC) is a highly challenging task due to the inherently subtle inter-class differences and the large intra-class differences. Researchers have attempted to address this challenge through Convolutional Neural Network (CNN)-based and Transformer-based methods, each of which has its own unique advantages. In order to share the metrics of both CNN-based and Transformer-based methods simultaneously and learn the latent features of fine-grained images efficiently. We introduce knowledge distillation to the field of FGVC for the first time and propose a novel method of Part-Selection based Data-efficient image Transformer (PS-DeiT), which incorporated the strengths of both CNN and Transformer models. More specifically, we propose the Part Selection Module to select the most discriminative image regions and exclude irrelevant regions, and employ a contrastive loss function that measures the similarity of images to distinguish the confusable classes in the task. Finally, we demonstrate the effectiveness of the proposed method PS-DeiT on four popular fine-grained datasets, i.e., CUB-200–2011, Stanford Cars, Stanford Dogs and NABirds, which achieves the accuracy of 90.8
Time series reconstruction is a crucial data processing step in time domain as-tronomy and serves as the foundation for fitting light curves and conducting time domain analysis.For many large-field time domain surveys,it is necessary to complete this com-putational process within a single exposure cycle.With the rapid increase in astronomical data,existing methods for astronomical data processing struggle to simultaneously meet the accuracy and efficiency requirements of time-series reconstruction.The memory-based computing general-purpose distributed framework,Spark,holds the potential to improve the efficiency of this process.However,applying Spark directly often encounters issues.MapRe-duce distributed models like Hadoop and Spark require relatively independent tasks among distributed cluster nodes and minimal data transfer across nodes during execution.Oth-erwise,frequent communication becomes an efficiency bottleneck for the application of the model.However,due to the presence of boundary problems in cross-matching,it is inevitable to transmit newly added data at the boundaries,severely restricting the concurrency of the model and reducing the acceleration ratio in practical parallel model applications.There-fore,we propose a non-blocking asynchronous execution flow,where each distributed process handles continuous processing exclusively for independent sky regions.The delayed batch appending of additional identification tasks from block-edge newly added celestial bodies in other nodes is determined based on the progress of each process.This ensures that identification calculations are not omitted,thereby improving concurrent efficiency while maintaining algorithm accuracy.Additionally,a research study was conducted on different join strategies between two tables,examining them from both theoretical and experimental perspectives.Furthermore,a join-free strategy was proposed.Finally,the design of an effi-cient time-series reconstruction system based on the Spark distributed framework validates the aforementioned research.Experimental results demonstrate a significant improvement in the efficiency of the proposed time-series reconstruction algorithm compared to previ-ous research,laying a solid foundation for the analysis of astronomical time-series data in time-domain astronomy.
The medical dialogue system aims to create a smart consultation platform for diagnosing diseases. Prior research uses doctor-patient dialogue history for responses, neglectingmedical clues guidance. This oversight can lead to inconsistencies between responses and crucial medical clues in context. To solve this problem, we propose Entity Perception and Reasoning forMedical Dialogue System (EPR), which is built on two components, i.e., Entity-PerceptionModule and Local Entity Attention Module. Entity-Perception Module first predicts medical entities (e.g., symptoms, diseases, and medicines) included in the next response through dialogue history as explicit clues to simulate the diagnosis process of real doctors, then Local Entity Attention Module detects the corresponding relevance between medical entities in dialogue history and medical entities in the next response as implicit clues to reason internal transferability between medical entities. Finally, EPR aggregates above medical clues and guide dialogue history to achieve the consistency of medical response and contextual reasoning logic. Experimental results show that these methods effectively improve entity-based metrics on MedDG.
In the field of computer vision, many perception methods rely on depth information captured by depth cameras. However, the integrity of depth maps is hindered by the reflection and refraction of light on transparent objects. Existing methods of completing depth map are usually impractical due to depth estimation error or unacceptably slow inference speeds. To address this challenge, we propose a lightweight depth completion model based on the Mobile Vision Transformer (LDCM-MViT), which uses a Mobile Guide Block (MGB). The MGB can efficiently fuse features from RGB and depth maps with limited parameters. Furthermore, we provide two types of fusion strategies to process RGB and depth features to get final depth map. Finally, we demonstrate the performance of LDCM MViT compared with the DDC-SRGBD model and GuideFormer model on Matterport3D and KITTI datasets. Experimental results show that our model has a higher accuracy in comparison to the traditional methods with limited parameters, especially on edge devices.
Automatic classification of stellar spectra contributes to the study of the structure and evolution of the Milky Way and star formation. Currently available methods exhibit unsatisfactory spectral classification accuracy. This study investigates a method called DSRL, which is primarily used for automated and accurate classification of LAMOST stellar spectra based on MK classification criteria. The method utilizes discrete wavelet transform to decompose the spectra into high-frequency and low-frequency information, and combines residual networks and long short-term memory networks to extract both high-frequency and low-frequency features. By introducing self-distillation (DSRL-1, DSRL-2, and DSRL-3), the classification accuracy is improved. DSRL-3 demonstrates superior performance across multiple metrics compared to existing methods. In both three-class(F ,G ,K) and ten-class(A0, A5, F0, F5, G0, G5, K0, K5, M0, M5) experiments, DSRL-3 achieves impressive accuracy, precision, recall, and F1-Score results. Specifically, the accuracy performance reaches 94.50
In our paper, we propose the Adaptive Attention-based Generative Adversarial Network (AAGAN) for text to image generation, and the modal combines the multi-layer GANs and Adaptive Attention Mechanisms to control the fine-grained image generation process at different levels. The core components of AAGAN include the Adaptive Attention Module (AAM) and Spatial-Channel Instance Normalization (SCIN). AAM can dynamically improve the attention weights depending on instructive attention standard based on the paired text-image features in both spatial and channel dimensions. In addition, we propose the loss function for spatial and channel respectively to constrain the proximity of the correlation to the instructive standard. SCIN ensures that the training process is not influenced by samples within the same batch. According to our experimental results, AAGAN can achieve high-quality image generation based on natural language descriptions.
Stock market forecasting remains a significant challenge within the financial sector. The employment of nonlinear predictive techniques has become increasingly prevalent in this domain. This paper introduces a novel approach to stock prediction that leverages a latent space to address the complexities of high-dimensional and intricate stock data. Our methodology integrates Variational Auto-Encoders (VAE) with Long Short-Term Memory networks (LSTM) to first embed the original data into a latent space via VAE, followed by training an LSTM model within this space. This approach not only reduces the model's complexity but also enhances the efficiency of the training process. The proposed model's effectiveness is empirically validated through its application to the S&P 500 dataset, which represents the United States stock market.
Policy search is an efficient learning method in the field of deep reinforcement learning (DRL), which is capable of solving large-scale problems with continuous state and action spaces and widely used in real-world problems. However, such method usually requires a large number of trajectory samples and extensive training time, and may suffer from poor generalization ability, making it difficult to generalize the learned policy model to seemingly small changes in the environment. In order to solve the above problems, this paper proposes a policy search DRL method based on latent space. Specifically, this paper extends the idea of state representation learning to action representation learning, i.e. learning a policy in the latent space of action representations, and then mapping the action representations to the real action space. With the introduction of representation learning models, this paper abandons the traditional end-to-end training manner in DRL and divides the whole task into two stages: large-scale representation model learning and the small-scale policy model learning, where unsupervised learning methods are employed to learn the representation models and policy search methods are used to learn the small-scale policy model. Large-scale representation models can ensure the capacity for generalization and expressiveness, while small-scale policy model can reduce the burden of policy learning, thus alleviating the issues of low sample utilization, low learning efficiency and weak generalization of action selection in DRL to some extent. Finally, the effectiveness of introducing the latent state and action representations is demonstrated by the intelligent control task CarRacing and Cheetah.
State representations considerably accelerate learning speed and improve data efficiency for deep reinforcement learning (DRL), especially for visual tasks. Task-relevant state representations could focus on features relevant to the task, filter out irrelevant elements, and thus further improve performance. However, task-relevant representations are typically obtained through model-based DRL methods, which involves the challenging task of learning a transition function. Moreover, inaccuracies in the learned transition function can potentially lead to performance degradation and negatively impact the learning of the policy. In this paper, to address the above issue, we propose a novel method of explainable task-relevant state representation (ETrSR) for model-free DRL that is direct, robust, and without any requirement of learning of a transition model. More specifically, the proposed ETrSR first disentangles the features from the states based on the beta variational autoencoder (β-VAE). Then, a reward prediction model is employed to bootstrap these features to be relevant to the task, and the explainable states can be obtained by decoding the task-related features. Finally, we validate our proposed method on the CarRacing environment and various tasks in the DeepMind control suite (DMC), which demonstrates the explainability for better understanding of the decision-making process and the outstanding performance of the proposed method even in environments with strong distractions.
Multi-scenario text generation is an essential task in natural language generation because of the multi -scene interlaced property of real-world problems. Traditional methods typically train the multi-scenario text generation models based on maximum likelihood estimation, which may suffer from the problem of exposure bias. Reinforcement learning (RL) based text generation methods could mitigate the exposure bias problem to some extent. However, the RL-based text generation methods are limited to the single -scenario tasks, which cannot be straightforwardly generalized to new scenario tasks. To address this prob-lem, in this paper, we propose a multi-scenario text generation method based on meta RL (MetaRL-TG), which implements the method of model-agnostic meta-learning (MAML) in the framework of RL-based text generation. The proposed MetaRL-TG method first learns the initial parameters from multiple train-ing tasks, then fine-tunes them in the target task. Thus, the proposed method is expected to efficiently achieve high-quality generated text in the new scenario. Finally, the effectiveness and generalization ca-pability of the proposed method are demonstrated for eight scenarios through English test datasets.(c) 2022 Elsevier B.V. All rights reserved.
Deep reinforcement learning (DRL) is an effective learning paradigm to realize general artificial intelligence, and has achieved remarkable achievements in a series of real-world applications. However, Deep reinforcement learning has some challenges, such as generalization capability and sample efficiency. Representation learning based on deep neural networks can effectively alleviate the above problems by learning the underlying structure of the environment. Therefore, latent space based deep reinforcement learning has become the popular method in this field. This paper systematically summarizes the research progress of latent space based representation learning in DRL, divides it into state representation, action representation and dynamic model in latent space, and analyzes the existing methods of DRL based on latent space. Finally, we summarize the successful applications of latent space based DRL, and discuss the future directions.
基于深度学习的解耦表示学习可以通过数据生成的方式解耦数据内部多维度、多层次的潜在生成因素,并解释其内在规律,提高模型对数据的自主探索能力.传统基于结构化先验的解耦模型只能实现各个层次之间的解耦,不能实现层次内部的解耦,如变分层次自编码(variational ladder auto-encoders,VLAE)模型.本文提出全相关约束下的变分层次自编码(variational ladder auto-encoder based on total correlation,TC-VLAE)模型,该模型以变分层次自编码模型为基础,对多层次模型结构中的每一层都加入非结构化先验的全相关项作为正则化项,促进此层内部隐空间中各维度之间的相互独立,使模型实现层次内部的解耦,提高整个模型的解耦表示学习能力.在模型训练时采用渐进式训练方式优化模型训练,充分发挥多层次模型结构的优势.本文最后在常用解耦数据集 3Dshapes 数据集、3Dchairs 数据集、CelebA人脸数据集和dSprites数据集上设计对比实验,验证了TC-VLAE模型在解耦表示学习方面有明显的优势.
The clustering method based on deep learning can automatically learn the latent features of data, and can be easily generalized to large-scale datasets with high-dimension. Traditional deep clustering methods pay more at-tention to extracting hidden layer features of data through deep neural networks to improve clustering accuracy, and less analyze the determinism of data categories in clustering tasks. At the same time, there is a lack of analysis of the discrete latent vector distribution after imposing constraints. This paper proposes a variational deep generative clustering model under entropy regularizations (VDGC-ER), which uses the variational auto-encoder as the basic framework and introduces the Gaussian mixture model as prior of the latent variables. This paper first proposes the sample entropy regularization term to the discrete latent vector of Gaussian mixture model to improve the clustering accuracy of the model. Further, this paper defines the aggregated sample entropy regularization term on the discrete latent vector to reduce the clustering imbalance, so that it can avoid local optimization and improve the generative diversity. Then, this paper uses the Monte Carlo sampling and re-parameterization strategies to estimate the optimi-zation objective of VDGC-ER model, and uses the stochastic gradient descent method to calculate the model para-meters. Finally, this paper designs the comparison experiments on MNIST, REUTERS, REUTERS-10K and HHAR datasets to demonstrate the performance of the VDGC-ER model. Experimental results show that the model can not only generate high quality samples, but also present high accuracy clustering.