In recent years, video recognition models have witnessed the rapid development of Deep Neural Networks (DNNs). However, these models remain not robust to adversarial examples that are created by adding imperceptible perturbations to clean samples. Recent studies indicate that generating adversarial examples in the hard-label black-box setting is particularly challenging yet highly practical. Compared to image recognition models, there are few hard-label black-box adversarial example generation algorithms for video recognition models. To this end, we propose a hard-label black-box video adversarial example generation algorithm, referred to as Dynamic Black-box Algorithm (DBA). First, DBA uses the binary search algorithm to find the boundary video between two original videos; then, the sampling-based algorithm is used to estimate the gradient on the boundary video; finally, with a dynamic step size adjustment strategy, DBA moves the boundary video towards the direction of the estimated gradient to generate the adversarial video. Additionally, we designed another strategy to skip invalid samples generated during the adversarial example generation process. Experiments demonstrate that DBA attains a superior trade-off between the magnitude of perturbations and query efficiency. Specifically, DBA outperforms state-of-the-art algorithms, achieving an average reduction in Mean Squared Error (MSE) of over 50%.
We propose a symbolic model checking framework for Linear Temporal Dynamic Logic (LTDL) via compositional testers, extending Linear Dynamic Logic (LDL) with bidirectional LTL operators. This enables past-time specifications (unattainable in pure LDL) while preserving LTL’s conciseness and verification efficiency, without increasing theoretical expressiveness. We make four contributions: (1) algorithms for constructing and optimizing regular grammars for LTDL path expressions to eliminate redundant variables and productions; (2) an algorithm for eliminating zero-delay cycles in regular grammars to avoid algebraic loops; (3) a method for generating compositional testers from LTDL; and (4) LTDLTester, an open-source tool for producing SMV-compatible LTDL testers. This reduces verification to checking the initial states of a system-tester synchronous parallel composition against output assertions, allowing direct application of standard symbolic checkers (e.g., NuSMV). Experiments show our approach significantly outperforms MCMAS _LDLK , the first native LDL model checker.
Adaptive Traffic Signal Control (ATSC) is a pivotal research area within intelligent transportation systems, aiming to enhance transportation efficiency and alleviate traffic congestion at signalized intersections. While multi-agent deep reinforcement learning has been extensively applied to ATSC, existing approaches commonly frame it as a fully cooperative problem, presupposing that all agents are committed to pursuing a collective optimal solution. However, achieving such altruistic cooperation is often impractical. Furthermore, as the number of agents escalates, challenges such as the curse of dimensionality and non-stationarity arise, complicating the learning process. To address these issues, we propose a novel perspective by framing ATSC as a competitive-cooperative game trade-off scenario and design a multi-agent framework, termed Neighborhood Coordinated and Holistic Optimized Actor-Critic (NcHo-AC). Specifically, we introduce a novel traffic state representation, design a sophisticated feature extraction network, develop a robust training algorithm, and leverage mean field approximation to model population-level agent interactions. These designs foster neighborhood-level cooperation and communication, facilitate the learning of the desired Nash equilibrium, and mitigate the noise caused by agents’ exploratory behaviors, thereby alleviating non-stationarity and the curse of dimensionality, while enhancing scalability to large-scale traffic networks. Comprehensive experiments conducted on both synthetic and real-world datasets demonstrate that NcHo-AC significantly outperforms state-of-the-art baselines across four key metrics: average travel time, average queue length, delay, and throughput, along with improved convergence, robustness, and interpretability.
Metaheuristic algorithms (MH) and machine learning (ML) are important components of artificial intelligence (AI). The synergy between MH’s optimization search capabilities and ML’s data analysis strengths has proven to be highly effective, providing a powerful combination for delivering high-quality solutions across diverse fields. Particularly in real-world applications such as autonomous driving and healthcare, the integration of MH and ML can significantly enhance the intelligence level and decision-making efficiency of systems, addressing the urgent societal and industrial demands for high-efficiency and high-precision solutions. This paper clarifies the logic behind the surge in research on MH-ML hybrid algorithms and addresses the gaps in current reviews regarding their timeliness and breadth of perspective. We begin by elucidating the fundamental concepts underpinning MH and ML, followed by a comprehensive classification framework that categorizes and synthesizes the latest research findings systematically. The paper concludes with an exploration of the challenges inherent in MH-ML hybrid algorithms and proposes future research directions. The analysis of the collected literature demonstrates that the integration of MH and ML generally enhances the performance of algorithms in specific problems. Despite progress from simple combinations to deeper integrations, challenges such as theoretical lag and interpretability remain. Future research will focus on solving these challenges by exploring further integration, using open-source tools, and adapting across diverse domains to expand the use of MH-ML hybrid algorithms.
For speech enhancement tasks, spectrum utilization in the time-frequency domain is crucial, as it enhances the effectiveness of audio feature extraction while reducing computational consumption. Among current speech enhancement methods in the time-frequency domain, DenseBlock and the dual-path transformer have demonstrated promising results. In this paper, to further improve the performance of speech enhancement, we optimize these two modules and propose a novel mapping neural network, DDP-Unet, which comprises three components: the encoder, the decoder, and the bottleneck. Firstly, we introduce a lightweight module, depth-point convolutional layer (DPCL), which employs point-wise and depth-wise convolutions. DPCL is then integrated into our novel DCdenseBlock, expanding DenseBlock's receptive and enhancing feature fusion in the encoder and decoder stages. Additionally, to increase the breadth and depth of feature fusion in the dual-path transformer, we implement a dual-path transformer as the bottleneck. DDP-Unet is then evaluated on two public datasets, VCTK + DEMAND and DNS Challenge 2020. Experimental results demonstrate that DDP-Unet outperforms most existing models, achieving state-of-the-art performances on STOI, PESQ, Si-SDR metrics.
The Minimum Weight Vertex Cover (MinWVC) problem has applications in clustering aggregation, electronic auctions, map labeling, etc. Given a simple undirected graph, a vertex cover is a vertex subset such that any edge in this graph has endpoints in it. In MinWVC, every vertex has a weight, and the task is to obtain a vertex cover and minimize its total vertex weight, i.e., the MinWVW problem looks for a lightest cover.The reduction approach checks a subgraph and determines whether certain vertices lie in an optimal cover in low complexity. In reality, reductions usually face a trade-off between effectiveness and efficiency. More specifically, the more cases in which they need to examine vertices’ belonging to an optimal solution, the exponentially more run time they consume.This paper presents a general theorem that guides us in finding effective and efficient reduction rules. These rules are not only able to address cases that have already been handled by current rules but are also usable in some other cases, which illustrate the advantages of our approach.
The goal of generalized planning (GP) is to find a generalized solution for a class of planning problems. One of effective means to solve GP is to transform a GP problem into an abstract planning problem, which can be easily solved. Recently, Lin et al. proposed a novel abstract model for GP, namely generalized linear integer numeric planning (GLINP), whose solution is an algorithmic-like structure called a planning program. They also developed an inductive approach to generating planning programs for GLINP. However, it has no theoretical guarantee that the generated planning program holds for infinitely many problem instances. To address this defect, we propose an automatic approach to verify whether the planning program works for infinitely many problem instances in this paper. We translate the planning program into a set of trace axioms finitely represented by linear integer arithmetic with uninterpreted predicate and function symbols (LIAUPF), and reduce the problem to the entailment problem of LIAUPF. Due to the undecidability of entailment problem in LIAUPF, we identify a class of planning programs whose trace axioms can be simplified in linear integer arithmetic (LIA), that is, a decidable fragment of LIAUPF, when reasoning about only the input and output of planning programs. As a result, the correctness verification of this class of programs becomes decidable.
Cormputed tomography (CT) scanning is an effective medical imaging modality widely used in clinical medicine for diagnosing various conditions. CT can generate three-dimensional images, thus providing more information than traditional two-dimensional radiographs. However, this comes at a cost, as it involves higher radiation, increased expense, and more time consumption. With the advancement of artificial intelligence, particularly the rise of computer vision, researchers have explored using deep convolutional neural networks and generative adversarial networks for low-radiation CT reconstruction tasks. Yet, existing CT reconstruction methods from X-ray images often focus on pixel-level difference metrics, neglecting the perceptual differences considered by the human visual system, especially in accurately restoring internal details in medical images. In response to this, this paper introduces the auto-encoder-based generative adversarial network (AECT-GAN) model, which integrates an auto-encoder structure and Sobel Gradient Guidance (SGG) mechanism within the discriminator, aiming to enhance the fidelity of image detail reproduction. Experimental validation on the LIDC-IDRI lung CT dataset has demonstrated our AECT-GAN method’s superiority in qualitative and quantitative evaluations, notably achieving significant improvements in preserving fine contours and textures in reconstructed images. Furthermore, applying this model to the IXI brain MRI dataset conclusively proves its widespread applicability and outstanding performance in the medical imaging domain.
Two analogous tasks have emerged in multimodal dialogue research: multimodal dialogue response generation and multimodal task-oriented dialogue. Both tasks share the goal of response multi-round, interactive content based on the multimodal dialogue history, but the latter focuses on accomplishing specific objectives which can be viewed as the former fine-tuned. The fine-tuning strategy may cause catastrophic forgetting and overfitting on few well-annotated data. Despite considerable progress in both areas, many existing works rely on retrieval-based approaches and additional auxiliary knowledge bases. To address these issues, we propose a Hierarchical Fusion Framework (HFF) for multimodal dialogue response generation. HFF blends these two tasks to learn a generation model from a data-driven perspective by introducing multi-dataset learning scheme, achieving a balance between generalization and expertise. In this work, multi-dataset learning is cast as a multi-objective optimization problem due to potential conflicts between datasets, necessitating a trade-off based on data distribution during training. Hierarchical fusion is performed sequentially between modalities and datasets, which could efficiently establish clear cross-modal relationships and integrate knowledge from multi-dataset. Specifically, HFF aligns extracted unimodal features (image and text) before fusing them through cross-modal attention and integrates them into multimodal encoder-decoder for generating responses. By optimizing the fusion between corpora from multi-dataset as conflicting objectives to satisfy Pareto optimality, our approach effectively facilitates both multimodal task-oriented and task-unoriented dialogues. Experimental results demonstrate the effectiveness of HFF and its comparable performance with all baselines.
The Centralized Training and Decentralized Execution (CTDE) paradigm, where a centralized critic is allowed to access global information during the training phase while maintaining the learned policies executed with only local information in a decentralized way, has achieved great progress in recent years. Despite the progress, CTDE may suffer from the issue of Centralized-Decentralized Mismatch (CDM): the suboptimality of one agent's policy can exacerbate policy learning of other agents through the centralized joint critic. In contrast to centralized learning, the cooperative model that most closely resembles the way humans cooperate in nature is fully decentralized, i.e. Independent Learning (IL). However, there are still two issues that need to be addressed before agents coordinate through IL: (1) how agents are aware of the presence of other agents, and (2) how to coordinate with other agents to improve joint policy under IL. In this paper, we propose an inference-based coordinated MARL method: Deep Motor System (DMS). DMS first presents the idea of individual intention inference where agents are allowed to disentangle other agents from their environment. Secondly, causal inference was introduced to enhance coordination by reasoning each agent's effect on others' behavior. The proposed model was extensively experimented on a series of Multi-Agent MuJoCo and StarCraftII tasks. Results show that the proposed method outperforms independent learning algorithms and the coordination behavior among agents can be learned even without the CTDE paradigm compared to the state-of-the-art baselines including IPPO and HAPPO.
In recent years, the transfer learning method of replacing acoustic features with phonetic features has become a new paradigm for end-to-end spoken language recognition. However, these larger transfer learning models always encode too much redundant information. In this paper, we propose a lightweight language recognition decoder based on a phonetic learnable dictionary encoding (PLDE) layer, which is more suitable for phonetic features and achieves better recognition performances while significantly reducing the number of parameters. The lightweight decoder consists of three main parts: (1) a phonetic learnable dictionary with ghost clusters, which improves the traditional LDE pooling layer and enhances the model’s ability to model noise with ghost clusters; (2) coarse-grained chunk-level pooling, which can highlight the phone sequence and suppress noise around ghost clusters, and hence reduce their influence to the subsequent network; (3) fine-grained chunk-level projection, which enables the discriminative network to obtain more linguistic information and hence improve the model’s modelling ability. These three parts simplify the language recognition decoder into a PLDE pooling layer, reducing the parameter size of the decoder by at least one order of magnitude while achieving better recognition performances. In experiments on the OLR2020 dataset, the Cavg of the proposed method exceeds that of the current state-of-the-art language recognition system, achieving 24.68% and 42.24% improvements on the cross-channel test set and unknown noise test set, respectively. Furthermore, experimental results on the OLR2021 dataset also demonstrate the effectiveness of PLDE.
Alzheimer’s disease (AD) is a neurodegenerative disease and there is by far no effective treatment for it, especially in its late stage. Circular RNAs (circRNAs), known as a class of non-coding RNAs are widely observed in eukaryotic transcriptomes, and are reported to play an important role in neurodegenerative diseases including AD. circRNAs usually act as microRNA (miRNA) inhibitors or «sponges» to regulate the function of miRNAs, leading to subsequent changes in protein activities and functions. Accumulating evidence indicates that circRNAs can serve as potential biomarker in AD early prediction. The functional roles of circRNAs are very versatile including miRNAs binding - thus affecting downstream gene expression, generating abnormally translated protein peptides, and affecting epigenetic modifications which subsequently affect AD related gene expressions. Therefore, identifying AD-related circRNAs can contribute to AD early diagnosis and intervention. In this work, we collected and curated an AD-related circRNA dataset; by exploring the circRNAs’ corresponding DNA loci distribution in chromatin 3D conformation (3D genome) and utilize the such 3D genome information, we were able to selected a concise yet predictively effective circRNA panel, based on which, significantly better AD prediction machine learning models were achieved.
Pointer analysis or points-to analysis (PTA) is a static program analysis for variables in a program, which determines a set of heap objects that individual variables may refer to at run time. In the literature, various types of context-sensitive analyses have been applied to improve the precision of PTA. In this article, we propose a framework that unifies existing context-sensitive PTA methods, under which we further explore more efficient ways for points-to calculation. In particular, we propose a call-graph-based context generation algorithm that combines the object-sensitive PTA and parameter-sensitive PTA approaches, and we implement the algorithm in the Soot compiler framework. Our new algorithm generates contexts for methods in a more complete and effective way, and it has been shown to achieve better precision with fewer generated contexts and less execution time than some of the known state-of-the-art context-sensitive approaches for PTA when tested with a selection of benchmarks from the DaCapo suite.
Interventional therapy is the main treatment for intermediate and advanced hepatocellular carcinoma (HCC) and liver injury is one of the common adverse effects after interventional therapy. This study aimed to compare the liver injury after transcatheter arterial chemoembolization (TACE) combined with immune checkpoint inhibitors (ICIs) and tyrosine kinase inhibitors (TKIs) and hepatic arterial infusion chemotherapy (HAIC) combined with ICIs and TKIs for intermediate and advanced HCC. We retrospectively enrolled patients with intermediate and advanced HCC who received TACE/HAIC combined with ICIs and TKIs from January 2019 to November 2021, in Nanfang Hospital. The liver function indexs within 1 week before treatment and on the first day after intervention were recorded. The degree of postoperative liver injury was assessed according to the Common Terminology Criteria for Adverse Events version 5.0 (CTCAE 5.0). 82 patients were treated with HAIC combined with ICIs and TKIs and another 77 patients received TACE combined with ICIs and TKIs. There were no significant differences in gender, age, cirrhosis, Child-Pugh grade, ALBI grade and combined ICIs and TKIs regimen at baseline between the two groups. The patients had more advanced tumor stage, heavier tumor load, and poorer liver reserve function in the HAIC group. Patients who underwent HAIC had larger proportion of BCLC stage C HCC(81.7% vs 63.6%,p=0.01),greater percentage of tumor with a maximum diameter greater than or equal to 10 centimeters(64.6% vs 32.5%,p<0.001)than patients in the TACE group. The frequencies of elevated ALT (28.0% vs 63.6%; P<0.001), elevated AST (54.9% vs 85.7%; p<0.001), and hyperbilirubinemia (40.2% vs 55.8%,p=0.049) were significantly lower in the HAIC group than in the TACE group. There were more patients having ALBI score of grade III after therapy in the TACE group than in the HAIC group(16.9% vs 6.1%,p=0.032). The effectiveness of HAIC with ICIs and TKIs were comparable to TACE with ICIs and TKIs. The incidence of liver injury events was higher in the TACE group than in the HAIC group.
Lipid droplet deposition in the kidney induces oxidative stress, which can worsen kidney function in diabetes. Scavenger receptor CD36 and fatty acid binding protein 4 (FABP4) are highly expressed in renal proximal tubules (RPTs) in diabetes and mediate lipid deposition by uptake free-fatty acid (FFA). The role of nuclear factor erythroid 2-related factor 2 (Nrf2) in diabetic kidney disease (DKD) progression is controversial. To investigate whether genetic deletion of Nrf2 could attenuate lipid droplet deposition and oxidative stress, as well as ameliorate tubular injury in db/db mice via down-regulation of CD36 and FABP4 expression in RPTCs.
While intensive therapy with insulin and renin-angiotensin system (RAS) blockers are effective in retarding diabetic kidney disease (DKD) progression, they do not cure it. Angiotensinogen (AGT, the sole precursor of all angiotensin) expression is up-regulated in renal proximal tubules (RPTs) of animals and patients with diabetes. However, the role of intrarenal RAS in DKD progression is not well understood.
The Minimum Vertex Weighted Coloring (MinVWC) problem is an important generalization of the classic Minimum Vertex Coloring (MinVC) problem which is NP-hard. Given a simple undirected graph G=(V,E), the MinVC problem is to find a coloring s.t. any pair of adjacent vertices are assigned different colors and the number of colors used is minimized. The MinVWC problem associates each vertex with a positive weight and defines the weight of a color to be the weight of its heaviest vertices, then the goal is the find a coloring that minimizes the sum of weights over all colors. Among various approaches, reduction is an effective one. It tries to obtain a subgraph whose optimal solutions can conveniently be extended into optimal ones for the whole graph, without costly branching. In this paper, we propose a reduction algorithm based on maximal clique enumeration. More specifically our algorithm utilizes a certain proportion of maximal cliques and obtains lower bounds in order to perform reductions. It alternates between clique sampling and graph reductions and consists of three successive procedures: promising clique reductions, better bound reductions and post reductions. Experimental results show that our algorithm returns considerably smaller subgraphs for numerous large benchmark graphs, compared to the most recent method named RedLS. Also, we evaluate individual impacts and some practical properties of our algorithm. Furthermore, we have a theorem which indicates that the reduction effects of our algorithm are equivalent to that of a counterpart which enumerates all maximal cliques in the whole graph if the run time is sufficiently long.
Thanks to Network Function Virtualization (NFV), Internet Service Providers (ISPs) can improve network resource utilization with significantly reduced capital and operational expenditures. To dig deeper into the potential of NFV, an important challenge is the resource allocation problem in NFV (NFV-RA), which can be divided into three stages: VNFs chain composition, VNF forwarding graph embedding, and VNFs scheduling. The key to the NFV-RA problem is to design an effective and coordinated resource allocation algorithm for the three stages. Besides, the NFV-RA problem has been proved to be NP-Hard, and thus most existing approaches focus on heuristic and meta-heuristic algorithms. In this paper, we propose an NFV online coordinated resource allocation framework (OCRA) that completes the three stages simultaneously in a coordinated manner by combining parallel Multi-Agent Deep Reinforcement Learning with novel neural networks and RL training techniques. The extensive experimental results show that compared with the state-of-the-art solutions, OCRA is highly-efficient in terms of time, with up to 50% and 10.8% improvement on resource overhead and acceptance ratio, respectively.
Named entity recognition (NER) is the localization and classification of entities with specific meanings in text data, usually used for applications such as relation extraction, question answering, etc. Chinese is a language with Chinese characters as the basic unit, but a Chinese named entity is normally a word containing several characters, so both the relationships between words and those between characters play an important role in Chinese NER. At present, a large number of studies have demonstrated that reasonable word information can effectively improve deep learning models for Chinese NER. Besides, graph convolution can help deep learning models perform better for sequence labeling. Therefore, in this article, we combine word information and graph convolution and propose our Lattice-Transformer-Graph (LTG) deep learning model for Chinese NER. The proposed model pays more attention to additional word information through position-attention, and therefore can learn relationships between characters by using lattice-transformer. Moreover, the adapted graph convolutional layer enables the model to learn both richer character relationships and word relationships and hence helps to recognize Chinese named entities better. Our experiments show that compared with 12 other state-of-the-art models, LTG achieves the best results on the public datasets of Microsoft Research Asia, Resume, and WeiboNER, with the F1 score of 95.89%, 96.81%, and 72.32%, respectively.
Computed Tomography (CT) is a medical imaging modality that can generate more informative 3D images than 2D X-rays. However, this advantage comes at the expense of more radiation exposure, higher costs, and longer acquisition time. Hence, the reconstruction of 3D CT images using a limited number of 2D X-rays has gained significant importance as an economical alternative. Nevertheless, existing methods primarily prioritize minimizing pixel/voxel-level intensity discrepancies, often neglecting the preservation of textural details in the synthesized images. This oversight directly impacts the quality of the reconstructed images and thus affects the clinical diagnosis. To address the deficits, this paper presents a new self-driven generative adversarial network model (SdCT-GAN), which is motivated to pay more attention to image details by introducing a novel auto-encoder structure in the discriminator. In addition, a Sobel Gradient Guider (SGG) idea is applied throughout the model, where the edge information from the 2D X-ray image at the input can be integrated. Moreover, LPIPS (Learned Perceptual Image Patch Similarity) evaluation metric is adopted that can quantitatively evaluate the fine contours and textures of reconstructed images better than the existing ones. Finally, the qualitative and quantitative results of the empirical studies justify the power of the proposed model compared to mainstream state-of-the-art baselines.