The automatic generation of code comments is a crucial aspect in software engineering as it allows the description of source code functions in natural language. In recent years, researchers have utilized advanced machine translation models, particularly transformer-based models, to achieve remarkable results in the task of code summarization. Despite these advancements, current methods still face some challenges. First, existing models fail to incorporate the rich structural information of the code well, resulting in inadequate summary statements. Second, most generative models are too sensitive to the editing of code text, which results in insufficient generalization ability of the model. To address these issues, we proposed a new model named SPA-Trans(Structure Position-aware Attention Transformer-based model). SPA-Trans uses a distance matrix based on the relative distance of nodes in the AST (Abstract Syntax Tree) to represent the structural correlation between each token, enhancing the model's ability to capture structural information of the source code. Additionally, in order to reduce the sensitivity of the model to code editing, we used the adversarial training method in the embedding layer to simulate code editing, which improves the generalization of the model. Our experiments on real-world datasets in Java and Python validate the effectiveness of our proposed method.
Data heterogeneity hinders clinical deployment of medical image analysis models, and generative data augmentation helps mitigate this issue. However, recent diffusion-based methods that synthesize image-mask pairs often ignore distribution shifts between generated and real images across scenarios, and such mismatches can markedly degrade downstream performance. To address this issue, we propose AlignFlow, a flow matching model that aligns with the target reference image distribution via differentiable reward fine-tuning, and remains effective even when only a small number of reference images are provided. Specifically, we divide the training of the flow matching model into two stages: in the first stage, the model fits the training data to generate plausible images; Then, we introduce a distribution alignment mechanism and employ differentiable reward to steer the generated images toward the distribution of the given samples from the target domain. In addition, to enhance the diversity of generated masks, we also design a flow matching based mask generation to complement the diversity in regions of interest. Extensive experiments demonstrate the effectiveness of our approach, i.e., performance improvement by 3.5-4.0
In autonomous driving, real-time decision-making requires semantic segmentation models to accurately recognize known objects while reliably identifying unexpected anomalies (e.g., road debris, animals, sudden obstacles). Most existing approaches operate under a closed-set assumption, where all test categories are assumed to have been seen during training, limiting their ability to handle unknown objects in real-world scenarios. Recently, semantic image synthesis has emerged as a promising paradigm by generating images from semantic layouts and comparing them with real observations to reveal inconsistencies. Existing methods often rely on pixel-level differences that may not effectively capture meaningful discrepancies. To address this limitation, we propose a novel discrepancy-aware framework based on semantic image synthesis. Instead of modeling unknown objects in label space, the method introduces controlled semantic perturbations as a proxy to induce semantic-visual inconsistencies, which are transferred from label space to image space and further reflected in deep feature space. The model then learns to identify potential anomalous regions by capturing discrepancy patterns between real and synthesized images through multi-scale feature fusion. Extensive experiments on multiple benchmark datasets demonstrate that the method consistently achieves superior accuracy and robustness over state-of-the-art approaches, highlighting its effectiveness for safety-critical autonomous driving applications. Code is available at: https://github.com/ah-ke/SYN-UOD.
Embodied agents in visual navigation tasks can infer the object location based on commonsense knowledge about unknown environments. Most existing methods primarily focus on constructing 2D scene graphs using panoramic observations or encoding semantic information in 2D representations using large language models. However, these methods hinder to perceive 3D scene geometry, prone to ignore the surrounding details by ambiguous selection of observations and lack grounding interaction with the environment, which generalizes poorly to novel objects or unseen environments. We propose BevNav framework to solve these issues from three aspects: (i) We introduce a novel Bird's Eye View (BEV) scene graph (BevSG) that utilizes multi-view 2D information transformed into 3D under the supervision of 3D detection to encode scene layouts and geometric clues. It can distinguish multi-view semantically similar objects and make plans in this graph. (ii) We propose BEV-BLIP based contrastive learning that aligns the BEV and language grounding inputs. It transfers constrained commonsense knowledge in pre-trained models without other training in the environments. (iii) We design BEV-based view search navigation strategy, which encourages representations that encode the semantics, relationships, and positional information of objects. This policy leverages the topological relations of locally collected BEV representations to infer invisible objects. Utilizing BevSG, the agent can predict a BEV graph decision score, for more accurate action prediction, thus improving exploration efficiency. Extensive experiments demonstrate that BevNav shows promising results on Gibson, HM3D, and ProcTHOR, which exhibits higher success rates than existing graph-based and LLM-based methods, indicating the feasibility of BevSG and commonsense knowledge from language models, leading efficient semantic exploration.
Semantic image synthesis aims to generate target images conditioned on given semantic labels, but existing methods often struggle with maintaining high visual quality and accurate semantic alignment. To address these challenges, we propose VD-GAN, a novel framework that integrates advanced architectural and functional innovations. Our variational generator, built on an enhanced U-Net architecture combining a pre-trained Swin transformer and CNN, captures both global and local semantic features, generating high-quality images. To further boost performance, we design two innovative modules: the Conditional Residual Attention Module (CRAM) for dimensionality reduction modulation and the Channel and Spatial Attention Mechanism (CSAM) for extracting key semantic relationships across channel and spatial dimensions. Additionally, we introduce a dual-function discriminator that not only distinguishes real and synthesized images, but also performs multi-class segmentation on synthesized images, guided by a redefined class-balanced cross-entropy loss to ensure semantic consistency. Extensive experiments show that VD-GAN outperforms the latest supervised methods, with improvements of (FID, mIoU, Acc) by (5.40%, 4.37%, 1.48%) and increases in auxiliary metrics (LPIPS, TOPIQ) by (2.45%, 23.52%). The code will be available at https://github.com/ah-ke/VD-GAN.git.
The task of object goal navigation (ObjNav) requires the agent to locate the given target object within a complex dynamic scene. To successfully accomplish the task, the agent needs to well understand the scenes, make executable decisions with less steps, avoid collisions, and successfully navigate to the target. As a result, efficient environmental perception and scene graph-inspired path planning is important to successfully accomplish the ObjNav task. In this paper, we present a hierarchical scene graph (HSG) contrastive learning, which consists of (1) a multimodal graph mixer that aligns the visual and textual information using open-vocabulary detector with GLIP. It can be regarded as an "eagle eye" to perceive target-related frontiers and suppress irrelevant information, and (2) a graph constructer that takes observed RGBD images to incrementally build a hierarchical scene graph. It acts as the "brain" that memorizes the common scene layout, (3) an action control contrastive learning that takes the graph contextual relationships as input to predict optimal actions to the target. It is treated as the "limbs" of the agent, coordinating and correcting incorrect movements. On the task of ObjNav, experiments on Gibson, HM3D, MP3D, and ProcTHOR demonstrate that navigation plans from the HSG framework achieve significantly higher success rates than existing map-based method, indicating the feasibility of executing navigation utilizing commonsense knowledge from language models leading efficient semantic exploration. Code is available at https://github.com/luosword/HSG4VN.
Mimicking natural bio-surfaces with specific microstructures that possess inherent antibacterial ability provides a promising bacteriostatic strategy. However, current methods for biomimetic surface antisepsis can only be implemented on planar substrates made of limited materials. Herein, we developed a 3D printing-based strategy to facilely customize biomimetic microstructures on specified surfaces to successfully enable bacteriostatic function. Two-photon polymerization was employed to fabricate elaborate bacteriostatic microstructures on either flat or curved surfaces regardless of their material (daily-used glass or plastics). Streptococcus mutans were cultured on microstructured surfaces as model bacteria to validate surface bacteriostatic performance. As an example, the strategy was adopted to endow clear aligner substrate surface with bacteriostatic ability, which would enhance the applicability and convenience of these orthodontic devices for maintaining personal oral hygiene. Our strategy will inspire daily or industrial applications aiming at surface sterilization in a customized, convenient, and safe manner.
Combining artificial intelligence with static analysis is an effective method for classifying malicious code. Due to the development of anti-analysis techniques, malicious code commonly employs obfuscation methods like packing, which result in garbled assembly code and the loss of original semantics. Consequently, existing pre-trained code language models are rendered ineffective in such scenarios. Current research addresses this issue by converting malicious bytecode into grayscale images and extracting visual features for classification. However, this process truncates the original sequence, compromising its coherence and structure. Furthermore, the image dimensions undergo compression and cropping based on the model’s input requirements, leading to the loss of intricate details. Our solution is a lossless encoding method for the visual structure of code, enabling unrestricted processing of malicious code images of any size. We convert bytecode files into semantically lossless images with proportional width. Then, we use image interleaving encoding to address semantic truncation issues caused by traditional image preprocessing methods. This method also prevents the loss of original code information due to image cropping or compression. For feature extraction, our goal is to combine the lossless encoding results with both local receptive field features and global contextual features. For local features, we achieve uniform embedding of variably sized input samples into equally sized feature maps using a multi-scale feature extraction module. For global contextual features, we reframe the feature maps along the row dimension, treating them as long-text sequences embedded in a matrix. We segment the feature maps into multiple row patch blocks and modify the Transformer’s input components to cache and merge the hidden states of each block. Comparative experiments on various malware datasets demonstrate the effectiveness of our method, consistently achieving outstanding performance across classification metrics.
Semantic image synthesis approaches has been dominated by the modelling of Convolutional Neural Networks (CNN). Due to the limitations of local perception, their performance improvement seems to have plateaued in recent years. To tackle this issue, we propose the SC-UNet model, which is a UNet-like network fused Swin Transformer and CNN for semantic image synthesis. Photorealistic image synthesis conditional on the given semantic layout depends on the high-level semantics and the low-level positions. To improve the synthesis performance, we design a novel conditional residual fusion module for the model decoder to efficiently fuse the hierarchical feature maps extracted at different scales. Moreover, this module combines the opposition-based learning mechanism and the weight assignment mechanism for enhancing and attending the semantic information. Compared to pure CNN-based models, our SC-UNet combines the local and global perceptions to better extract high- and low-level features and better fuse multi-scale features. We have conducted an extensive amount of comparison experiments, both in quantitative and qualitative terms, to validate the effectiveness of our proposed SC-UNet model for semantic image synthesis. The outcomes illustrate that SC-UNet distinctively outperforms the state-of-the-art model on three benchmark datasets (Citysacpes, ADE20K, and COCO-Stuff) including numerous real-scene images.
The task of visual navigation (VN) is steering the agent find target object only using visual perceptions. Previous works largely exploit multimodal information (e.g. visual and training memory) to improve the environmental perception ability, while making less effort to leverage interchange information. Besides, multimodal fusion tends to ignore the data dependencies (prefer a part of the modal data) as well as the supervision of the action.In this work, we present a novel multimodal graph learning (MGL) structure for VN, which consists of three parts. (1) the multimodal fusion exploits the rich information across spatial, RGB, and depth information about objects’ place, as well as semantic information about their categories, (2) adaptive relation graph (ARG) is dynamically built using object detectors, which encodes multimodal fusion and adapt to a novel environment. It embeds its navigation history and other useful task-oriented structural information, thus make the agent own the association ability and make advisable informed decisions and (3) action boost module (ABM) aims to assist the agent make intelligent decisions, which predicts more accurate action using beneficial training experience. Our agent can foresight what the goal state may look like and how to get closer towards that state. These combinations of the “what” and the “how” allow the agent to navigate to the target object effectively. We validate our approach on the AI2-THOR dataset. It reports 24.2% and 23.7% increase in SPL(Success weighted by Per Length) and SR(Success Rate) compared with baselines, respectively. Code and datasets can be found in https://github.com/luosword/ABM_VN.
Major depression is a severe psychological disorder typically diagnosed using scale tests and through the subjective assessment of medical professionals. Along with the continuous development of machine learning techniques, computer technology has been increasingly employed to identify depression in recent years. Traditional methods of automatic depression recognition rely on using the patient's physiological data, such as facial expressions, voice, electroencephalography (EEG), and magnetic resonance imaging (MRI) as input. However, the acquisition cost of these data is relatively high, making it unsuitable for large-scale depression screening. Thus, we explore the possibility of utilizing a house-tree-person (HTP) drawing to automatically detect major depression without requiring the patient's physiological data. The dataset we used for this study consisted of 309 drawings depicting individuals at risk of major depression and 290 drawings depicting individuals without depression risk. We classified the eight features extracted from HTP sketches using four machine-learning models and used multiple cross-validations to calculate recognition rates. The best classification accuracy rate among these models reached 97.2%. Additionally, we conducted ablation experiments to analyze the association between features and information on depression pathology. The results of Wilcoxon rank-sum tests showed that seven of the eight features significantly differed between the major depression group and the regular group. We demonstrated significant differences in HTP drawings between patients with severe depression and everyday individuals, and using HTP sketches to identify depression automatically is feasible, providing a new approach for automatic identification and large-scale screening of depression.
Since the requirements of cross-domain/layer communications, the Space-ground Integrated Information Network (SIIN) becomes a strategic research area. To improve the service sustainability and reduce the latency of data transmission, literature works focus on evaluating the status of channels between ground stations and satellites, but underestimate the power of dynamic data allocation for handover management. This paper explores the relationship of data allocation and seamless handover in SIIN to provide high-reliability and service sustainability. We propose a Channel Perceiving-based Handover Management (CPHM) strategy to optimize the utilization of channels and dynamically adjust the data allocation strategy. Specifically, CPHM perceives the motion status of satellites to accurately evaluate their service time and reconstruct connectivities, e.g., altitude, velocity, motion direction, and location. Furthermore, CPHM evaluates the service capability of satellites to generate the strategy of data allocation and dynamically adjust this strategy. Then, to improve utilization of channels, CPHM manages transmission queues according the strategy of data allocation and length of queues. Extensive simulation results show that CPHM outperforms other baseline algorithms in terms of delivery ratio, average delivery latency, and interruption ratio.
With the development of Vehicular Ad-hoc Networks (VANETs), several data security challenges are revealed, such as data hijacking and interception. Although vehicles are authorized, malicious behaviors still be carried out. Security lapses may lead to potential accidents, which emphasizes the importance of laying a solid security foundation for VANETs. Thanks to the base security layer provided by cryptography technologies, security problems can be solved in VANETs to avoid accidents. However, trust management focuses on the analysis and identification of misbehavior, to ensure secure interactions among vehicles, and preserve data integrity against security issues. This paper explores trust assessments that consider the transmission path of message as a novel indicator, to provide a comprehensive and accurate trust assessment. We propose a Multidimensional trust Evidence Fusion and Path-Backtracking mechanism for trust management scheme (MEFPB) in VANETs. MEFPB integrates the multidimensional trust evidence fusion and path-backtracking mechanism. Specifically, MEFPB utilizes the Dempster-Shafer theory to fuse multi-dimensional indicators (direct trust, indirect trust, and transmission path of message) for evaluating the trustworthiness of vehicles. The direct and indirect trust are supplied by the message-sending vehicle and its neighbors (i.e., other vehicles). The transmission path of message is provided by roadside units. Furthermore, the path-backtracking mechanism identifies and traces malicious behaviors based on the transmission path of message. Moreover, extensive experiments demonstrate that our scheme significantly outperforms other baseline schemes, exhibiting a high malicious behavior detection rate within VANETs.
In the field of code intelligence, pretrained models exhibit impressive performance. However, it is imperative to modify all model parameters and maintain full copies for various tasks. Moreover, the effectiveness of fine-tuning a pretrained model depends on the availability of data, which can be constrained in practical settings. Prefix-tuning, a novel approach in NLP for addressing the aforementioned issue, has demonstrated promising results in numerous tasks. This paper aims to investigate the potential of prefix-tuning in the field of code intelligence, an area that has been relatively understudied. We also acknowledge the critical role of initialization in the performance of prefixtuning, which has been neglected in prior research. To address this gap, we introduce a novel method termed "adaptive prefixtuning". This approach entails training the prefix parameters on adapting tasks, followed by subsequent tuning on downstream tasks. We perform prefix-tuning and adaptive prefix-tuning using well-known pretrained models, CodeT5 and CodeGPT, and conduct experiments on four code intelligence tasks, namely, defect detection, code completion, code summarization, and code translation. Our experiments revealed that, in situations with restricted data availability, adaptive prefix-tuning yields substantial performance enhancements compared to fine-tuning, whereas prefix-tuning does not exhibit evident advantages. Furthermore, adaptive prefix-tuning has demonstrated superior performance relative to prefix-tuning in scenarios with comprehensive datasets. In certain cases, it has even outperformed traditional fine-tuning methods. Our findings indicate that within the field of code intelligence, adaptive prefix-tuning can serve as an effective substitute for fine-tuning, especially in situations with constrained data availability.
The accurate detection of malware is of paramount importance in today’s society, where the number of computers is vast. Traditional methods for classifying malicious code face three main challenges. Firstly, previous deep learning models encounter difficulties in handling excessively long input sequences due to their structural and computational resource limitations. Secondly, the rapid changes and rapid propagation of malicious samples pose certain challenges in collecting a sufficient number of samples and ensuring the accuracy of sample labels. Thirdly, in scenarios with limited sample availability, prior approaches failed to consider the augmentation of the model’s generalization performance when dealing with a limited number of samples. To address the first challenge, we propose a novel method that segments malicious samples into subroutines to extract opcode sequences. These sequences are then converted into two-dimensional opcode sequences and fed into a pretrained model. Additionally, we transform the relevant features of malicious sample opcodes, registers, and assembly language comments into one-dimensional vectors as prefixes. This approach guides the model to better learn sample feature representations, ultimately enhancing the classification performance. To solve the latter two challenges, we introduce a combination of supervised contrastive learning and cross-entropy loss, utilizing support set computation to generate prototype vectors. This enables the model to effectively tackle emerging variations of malicious code and enhance its performance in few-shot scenarios. Experimental results demonstrate that our method outperforms existing malicious code classification models and achieves outstanding performance in few-shot experiments.
Recently, with the continuous advancement of deep learning techniques, research on sketch synthesis has been progressing. However, existing methods still face challenges in generating human-like freehand sketches from real-world natural images at both object and scene levels. To address this, we propose SketchDiffusion, a text-guided freehand sketch synthesis method based on conditional stable diffusion. In SketchDiffusion, we design a novel image enhancing module to efficiently extract high-quality image features. Moreover, we utilize additional guidance from global and local features extracted by a U-shaped diffusion guidance network to control the noise addition and denoising process of the diffusion model, thereby significantly improving controllability and performance in freehand sketch synthesis. Beyond the model architecture, we leverage the designed BLIP-based text generation method to create 70,280 text prompts for foreground, background, and panorama sketch synthesis in the extensive SketchyCOCO dataset, thereby improving the overall effectiveness of model training. Compared to the state-of-the-art methods, our proposed SketchDiffusion has shown an average improvement of over 16.4%, 16.75%, and 12.8% on three quantitative metrics (sketch recognition, sketch-based retrieval, and user perceptual study), respectively. Furthermore, our approach not only excels in synthesizing freehand sketches containing multiple abstract objects but also has multiple applications in supporting human–computer interaction.
Efficiently searching and reusing code from expansive codebases is pivotal for enhancing developers’ productivity. In recent times, the emergence of deep learning-driven neural ranking models, characterized by their vast dimensions and intricate interaction mechanisms, has been noteworthy. Yet, these models, in real-world scenarios, pose computational challenges due to their high dimensionality. Moreover, models rooted in interaction necessitate querying every piece of code within a voluminous corpus. While these methodologies offer superior accuracy, their online retrieval process is considerably more time-consuming compared to traditional Information Retrieval (IR) techniques. Addressing this, we introduce “ExCS”, an innovative code search tool designed to expedite the code search process without compromising on accuracy. ExCS innovatively employs code expansion in its offline phase, leveraging predictions on potential queries for specific codes, thereby enriching the code’s semantic depth. During online retrieval, ExCS prioritizes IR-based methods to pinpoint a concise set of persuasive candidates. Our evaluations, conducted on the Java dataset from CodeSearchNet, reveal that ExCS achieves a remarkable 90% reduction in retrieval duration while maintaining an impressive 99% retrieval accuracy.
Flexible perovskite light-emitting diodes (PeLEDs) constitute an emerging technology opening new opportunities in the fields of lighting and display for portable and wearable electronics. Poly(3,4-ethylenedioxythiophene):poly(stryrenesulfonate) (PEDOT:PSS) as one of the most promising flexible electrode materials has attracted extensive attention. However, the patterning and conductivity issues of PEDOT:PSS electrodes should be addressed primarily. Here, a photopolymerizable additive is proposed to endow the PEDOT:PSS electrodes with photopatternability. Moreover, this additive can also improve the conductivity of the PEDOT:PSS electrode from 0.16 to 627 S/cm because of the phase separation between PEDOT and PSS components and conformation transition of PEDOT chains. Eventually, highly conductive PEDOT:PSS electrodes with various patterns are applied in flexible PeLEDs, demonstrating a high luminance of 25972 cd/m2 and a current efficiency of 25.1 cd/A. This work provides a facile and effective method of patterning and improving the conductivity of PEDOT:PSS electrodes simultaneously, demonstrating the great potential of PEDOT:PSS electrodes in flexible perovskite optoelectronics.
Development of high‐performance and stable pure red perovskite light‐emitting devices (LEDs) promotes commercialization of perovskite. Benefiting from the quantum confinement effect, pure red CsPbI 3 quantum dots (QDs) can avoid the color drift issue for mixed halide systems. Herein, a facile temperature‐dependent hot‐injection method combined with low temperature gradient centrifugation is employed to prepare 5.60 ± 0.05 nm pure red QDs. The QDs exhibit 642 nm photoluminescence (PL) with 37 nm full width at half maximum, and their Commission Internationale de l'Eclairage (CIE) coordinates located at (0.708, 0.292) match well with standard pure red for Rec. 2020. The PL half‐lifetime for pure red QDs solution is 210 days under ambient conditions, and both the X and Y coordinates of CIE only fluctuate within (0.002, 0.002). Finally, a simple NH 4 I post‐treatment to decrease ligand length further boosts the QDs conductivity, and the pure red LED exhibits 9.4% maximum external quantum efficiency. The electroluminescence spectra exhibits a sharp peak at 645 nm, and their CIE coordinates are located at (0.707, 0.290). Furthermore, the pure red LED exhibits both good working and color stability. Its half‐lifetime is 25.1 min and color drifts are only (0.002, 0.002) after 30.0 min of continuous working at 60 cd m −2 with a constant current density of 3.5 mA cm −2 .
Large-scale code search is a crucial task in software engineering, yet existing deep learning based models often embed Abstract Syntax Trees (ASTs) and code sequences separately, limiting their ability to learn the correlation between structural and textual features. To address this limitation, we propose a novel code search model that automatically generates Token-Level Information Flow Graphs (TL-IFGs) from aligned AST nodes and source code tokens. Our model includes an aligner that establishes a one-to-one correspondence between AST leaves and code tokens, which we make publicly available, along with a processed dataset to facilitate further research. The model automatically generates a TL-IFG for each code snippet from the aligned datas by predicting the information flow at the token-level, which ensures that structural and textual features are highly correlated during the embedding process. We also generate TL-IFGs for descriptions and embed them using a similar process. Experimental results demonstrate that our model outperforms state-of-the-art code search models, indicating the effectiveness of our approach. Furthermore, an ablation study shows that the generated TL-IFGs for both code and description positively impact model performance.