Existing Graph-RAG approaches struggle to align complex, multi-hop queries with knowledge graph (KG) structures, often due to disjoint entity and relation retrieval or over-reliance on LLM-generated reasoning plans. To bridge this semantic and structural gap, we propose PRIME, a training-free, structure-aware framework that unifies entities and relations into subgraph-level intermediate representations and enables incremental, semantics-guided exploration. This design drastically reduces the search space while preserving structural coherence. On KGQA benchmarks, PRIME achieves 91.6% Hit@1 on WebQSP and 78.3% F1 on CWQ, outperforming the best baselines by 2.8% and 4.5%, respectively-even when using LLaMA2-7B. It also runs in 11.5 seconds on CWQ, over 1,000 & times; faster than training-based methods, demonstrating superior accuracy and efficiency in deep multi-hop reasoning scenarios.
Efficient and effective modeling of feature interactions is key to large-scale Click-Through Rate (CTR) prediction. Although existing feature interaction methods have improved the model accuracy, their computational consumption still increase exponentially with the number of feature fields and become severe efficiency bottleneck in real-world industrial scenarios. To address the issues, we propose an E fficient and E ffective NET work for large-scale CTR prediction named EENet . EENet presents a new alternating stacking architecture of implicit and explicit interaction layers, and each implicit layer in EENet can reduce both local computational and parameter load remarkably. EENet also designs a unified explicit interaction operation which can only use simple matrix multiplication to capture field-wise patterns. Moreover, the order of multiplications in EENet is rearranged to further decrease the computational complexity from quadratic to linear with respect to the number of feature fields. EENet thus can support the high efficiency in real-practice industrial scenarios with hundreds of feature fields. A set of extensive experiments is performed on two public datasets and one industrial dataset for effectiveness evaluation, and five larger-scale synthetic datasets for efficiency evaluation. The results highlight that our EENet can significantly outperform the state-of-the-art models in terms of both efficiency and scalability, while also maintaining superior effectiveness. Compared with DCNv2 and FiBiNet, EENet achieves 8.06 \(\times\) and 36.72 \(\times\) efficiency improvements in training, and 2.02 \(\times\) and 48.88 \(\times\) improvements in inference, respectively. Our solution and source code are available at https://github.com/Yeedzhi/EENet .
Sequential recommendation aims to model users' evolving preferences based on their historical interactions. Recent advances leverage Transformer-based architectures to capture global dependencies, but existing methods often suffer from high computational overhead, primarily due to discontinuous memory access in temporal encoding and dense attention over long sequences. To address these limitations, we propose FuXi-gamma a novel sequential recommendation framework that improves both effectiveness and efficiency through principled architectural design. FuXi-gamma adopts a decoder-only Transformer structure and introduces two key innovations: (1) An exponential-power temporal encoder that encodes relative temporal intervals using a tunable exponential decay function inspired by the Ebbinghaus forgetting curve. This encoder enables flexible modeling of both short-term and long-term preferences while maintaining high efficiency through continuous memory access and pure matrix operations. (2) A diagonal-sparse positional mechanism that prunes low-contribution attention blocks using a diagonal-sliding strategy guided by the persymmetry of Toeplitz matrix. Extensive experiments on four real-world datasets demonstrate that FuXi-.. achieves state-of-the-art performance in recommendation quality, while accelerating training by up to 4.74x and inference by up to 6.18x, making it a practical and scalable solution for long-sequence recommendation. Code: https://github.com/Yeedzhi/FuXi-gamma.
Deep learning technology has enhanced the ability of Click-through rate (CTR) prediction models to learn features and improve prediction accuracy. However, it is challenging to deploy CTR models on GPU smoothly and perform inference efficiently, because there is a huge mismatch between the serial computational pattern and the parallel model structure. In this paper, we propose DPIFrame, the first dual parallelizable framework to accelerate CTR model inference. In DPIFrame, a) a dual parallelizable architecture is proposed to perform parallel CTR model inference in both intra-module and inter-module; b) an efficient multi-table lookup algorithm is presented for embedding operations through anticipating the whole workload in advance; c) a breadth-first stream scheduling strategy is designed for fine-grained management of parallel computation on GPU to further supporting the dual parallel execution. Extensive experiments are conducted on two real-world datasets, and the results highlight that DPIFrame can reduce the embedding latency efficiently by 23.0× compared to PyTorch. Compared with PyTorch, TorchRec, HugeCTR, and OneFlow, DPIFrame can achieve state-of-the-art inference performance on GPU with speedups of 5.83×, 4.29×, 2.15×, and 2.0×, respectively.
AI-driven network security relies increasingly on Large Language Models (LLMs) to detect sophisticated threats; however, their deployment on resource-constrained edge devices is severely hindered by immense parameter scales. While unstructured pruning offers a theoretical reduction in model size, commodity Graphics Processing Unit (GPU) architectures fail to efficiently leverage element-wise sparsity due to the mismatch between fine-grained pruning patterns and the coarse-grained parallelism of Tensor Cores, leading to latency bottlenecks that compromise real-time analysis of high-volume security telemetry. To bridge this gap, we propose SPARTA (Sparse Parallel Architecture for Real-Time Threat Analysis), an algorithm–architecture co-design framework. Specifically, we integrate a hardware-based address remapping interface to enable flexible row-offset access. This mechanism facilitates a novel graph-based column vector merging strategy that aligns sparse data with Tensor Core parallelism, complemented by a pipelined execution scheme to mask decoding latencies. Evaluations on Llama2-7B and Llama2-13B benchmarks demonstrate that SPARTA achieves an average speedup of 2.35× compared to Flash-LLM, with peak speedups reaching 5.05×. These findings indicate that hardware-aware microarchitectural adaptations can effectively mitigate the penalties of unstructured sparsity, providing a viable pathway for efficient deployment in resource-constrained edge security.
State access is a critical part of smart contract execution which seriously affects the efficiency of smart contract execution in the mainstream Ethereum blockchain. To reduce state access latency, existing studies typically require manual source code modifications, which in practice may deliver limited performance gains and shift the burden to developers. In this paper, we propose CROSC to reduce state access latency and improve smart contract execution efficiency by a compilation-runtime joint optimization approach. CROSC consists of three key parts: 1) a runtime memory management mechanism named Fast State Memory (FastSM) to fully utilize the working memory and provide the context for the contract compiler; 2) a State Variable Address Relocation (SVAR) strategy to minimize costly persistent storage operations by precisely redirecting state variable access targets during compilation; 3) a one-shot unpacking design that eliminates frequent decoding overhead for low-bitwidth state variables. Preliminary experimental results highlight that, compared with the baseline compilation and runtime system of Ethereum, CROSC can achieve 2.5$\times$x and 7.5$\times$x speedups for single state load and store operations, respectively. CROSC reduces state access latency by up to 81.3%, and overall contract execution latency by 32.9% on average across 14 typical types of smart contracts. Extended evaluations on ERC20 and ERC721 token standard contracts show that CROSC delivers significant benefits in critical areas while remaining unobtrusive for less intensive state operations.
Most manipulators require extensive operational space; however, in environments where space is limited, these devices must be compact during periods of inactivity. To address this challenge, a redundant rigid-flexible coupling deployable manipulator has been developed that optimizes space utilization and enhances operational capabilities. This development is informed by a detailed examination of the structure and motion performance of the Kresling origami unit. Equivalence principles for the mechanism are proposed, and an optimal rigid-flexible coupling equivalent mechanism unit is selected by integrating motion feasibility analysis with the significance of flexible structures. A 3RUU mechanism unit is chosen, and six such units are serially connected to construct a deployable manipulator. The workspace and mechanical properties of the manipulator are characterized, and principles for implementing reach-point motion are proposed to ensure superior overall performance. Experimental results show that the designed manipulator achieves a folding ratio of 2.58, supports a maximum load of 2611.1 g, and exhibits high flexibility and excellent overall performance in reach-point motion. These findings provide a solid foundation for the broader application of this type of manipulator.
Post-training quantization (PTQ) is critical for deploying Vision Transformers (ViTs) on resource-constrained devices. However, outliers clustered in activation channels tend to dominate the quantization range and induce significant accuracy degradation. To address the outlier challenge, this paper proposes an outlier resilient PTQ method through outlier decomposition, namely ORQ-ViT. Its core idea is to decompose outliers clustered in outlier channels into isolated outliers, so that they can be easily excluded, thus alleviating their adverse impact. Specifically, we decompose activations along the patch token dimension. Since there are very few outlier channels, decomposed rows after outlier decomposition usually cover several isolated outliers, which can be easily identified and filtered. We further design an adaptive quantization range determination strategy during quantization parameters initialization to prevent outliers from serving as boundary values of the quantization range. ORQ-ViT can improve quantization levels utilization to generate activations with higher quantization resolution, thereby achieving higher accuracy. Additionally, ORQ-ViT supports pure integer matrix multiplications to ensure the inference efficiency of quantized ViTs on edge hardware. Extensive experiments demonstrate that our method achieves state-of-the-art (SOTA) accuracy across various ViT variants under multiple low bit-width scenarios. For image classification, the top-1 accuracy of ORQ-ViT outperforms that of the SOTA methods by an average of 2.26% at W4A4. Even for object detection and instance segmentation, ORQ-ViT also delivers highly competitive results. We also evaluate the inference efficiency of pure integer matrix multiplications and the results show that our method can achieve up to 2.1x speedup.
Smart contract vulnerability detection has attracted increasing attention due to billions of economic losses caused by vulnerabilities. Existing smart contract vulnerability detection methods have high false negative and high false positive rates. To address these issues, we present ByteEye, a bytecode level smart contract vulnerability detection framework with Graph Neural Networks (GNNs). ByteEye first constructs an edge-enhanced Control Flow Graph (CFG) to maintain rich information from the low-level bytecode with low latency. ByteEye also designs and incorporates both general information and vulnerability-specific information into its detection method as bytecode level features. Furthermore, ByteEye flexibly supports machine/deep learning models, especially with graph neural networks, which can facilitate vulnerability detection precisely. The extensive experimental results highlight that ByteEye outperforms the state-of-the-art approaches on all three types of vulnerability detection. ByteEye can achieve an average of 35.29
Exploring the expected quantizing scheme with suitable mixed-precision policy is the key to compress deep neural networks (DNNs) in high efficiency and accuracy. This exploration implies heavy workloads for domain experts, and an automatic compression method is needed. However, the huge search space of the automatic method introduces plenty of computing budgets that make the automatic process challenging to be applied in real scenarios. In this paper, we propose an end-to-end framework named AutoQNN, for automatically quantizing different layers utilizing different schemes and bitwidths without any human labor. AutoQNN can seek desirable quantizing schemes and mixed-precision policies for mainstream DNN models efficiently by involving three techniques: quantizing scheme search (QSS), quantizing precision learning (QPL), and quantized architecture generation (QAG). QSS introduces five quantizing schemes and defines three new schemes as a candidate set for scheme search, and then uses the Differentiable Neural Architecture Search (DNAS) algorithm to seek the layer- or model-desired scheme from the set. QPL is the first method to learn mixed-precision policies by reparameterizing the bitwidths of quantizing schemes, to the best of our knowledge. QPL optimizes both classification loss and precision loss of DNNs efficiently and obtains the relatively optimal mixed-precision model within limited model size and memory footprint. QAG is designed to convert arbitrary architectures into corresponding quantized ones without manual intervention, to facilitate end-to-end neural network quantization. We have implemented AutoQNN and integrated it into Keras. Extensive experiments demonstrate that AutoQNN can consistently outperform state-of-the-art quantization. For 2-bit weight and activation of AlexNet and ResNet18, AutoQNN can achieve the accuracy results of 59.75% and 68.86%, respectively, and obtain accuracy improvements by up to 1.65% and 1.74%, respectively, compared with state-of-the-art methods. Especially, compared with the full-precision AlexNet and ResNet18, the 2-bit models only slightly incur accuracy degradation by 0.26% and 0.76%, respectively, which can fulfill practical application demands.
U-rib butt weld cracking is one of the most typical fatigue failure modes of orthotropic steel decks (OSDs), which directly endangers the safety of bridges. Since the on-site quality of butt welds is difficult to guarantee, manufacturing defects are an important factor affecting fatigue performance. A full-scale OSD section specimen was manufactured with a real process, and the force characteristics and fatigue stress history of U-rib butt welds were investigated, and the fatigue crack growth characteristics were analyzed. A finite element model (FEM) was developed to extract the structural stress components of typical failure modes, and the equivalent structural stress method was used to quantify the effect of manufacturing defect sizes on the failure modes and fatigue life of U-rib butt welds. The results showed that the fatigue strength of U-rib butt welds was evaluated to be 125.1 MPa (corresponding to 4.6 million load cycles) using the equivalent structural stress method. When the radius of defect R < 1.25 mm, the fatigue cracking mode is weld toe cracking; when R > 1.25 mm, the fatigue cracking mode is internal defect cracking, which has a significant effect on the fatigue life.
Real-time monitoring of cable tension is essential in detecting damage and assessing the service performance of bridges. Vibration frequency-based methods for cable tension identification have been increasingly investigated in bridge engineering over the last decade. However, cable tension estimation using vibration frequency-based methods can only obtain the average cable tension over a specified time interval. Therefore, identification of instantaneous variation of cable tension still remains to be a pending issue. To tackle this problem, this study proposes an algorithm defined as Synchrosqueezing Wave-packet-based Instantaneous Frequency Tracking (SWIFT) to identify the instantaneous change of cable tension by monitoring its acceleration responses. The proposed algorithm can extract the time-varying instantaneous frequency of the tested cable and further identify the variation of instantaneous cable tension by time-frequency analysis of the transversal motion. Moreover, the identification accuracy and noise robustness are numerically verified using a synthetic time-varying multi-frequencies signal. Finally, experimental validation is conducted on a prototype cable and an actual cable of the Sutong Yangtze River bridge, respectively. The theoretical, numerical, and experimental results show that the proposed method can precisely identify instantaneous cable tension under noisy conditions.
Lightweight convolutional neural networks (CNNs) enable lower inference latency and data traffic, facilitating deployment on resource-constrained edge devices such as field-programmable gate arrays (FPGAs). However, CNNs inference requires access to off-chip synchronous dynamic random-access memory (SDRAM), which significantly degrades inference speed and system power efficiency. In this paper, we propose an adaptive dataflow scheduling method for lightweight CNN accelerator on FPGAs named ADS-CNN. The key idea of ADS-CNN is to efficiently utilize on-chip resources and reduce the amount of SDRAM access. To achieve the reuse of logical resources, we design a time division multiplexing calculation engine to be integrated in ADS-CNN. We implement a configurable module for the convolution controller to adapt to the data reuse of different convolution layers, thus reducing the off-chip access. Furthermore, we exploit on-chip memory blocks as buffers based on the configuration of different layers in lightweight CNNs. On the resource-constrained Intel CycloneV SoC 5CSEBA6 FPGA platform, we evaluated six common lightweight CNN models to demonstrate the performance advantages of ADS-CNN. The evaluation results indicate that, compared with accelerators that use traditional tiling strategy dataflow, our ADS-CNN can achieve up to 1.29× speedup with the overall dataflow scale compression of 23.7%.
SummaryRecently, Kubernetes is widely used to manage and schedule the resources of microservices in cloud‐native distributed applications, as the most famous container orchestration framework. However, Kubernetes preferentially schedules microservices to nodes with rich and balanced CPU and memory resources on a single node. The native scheduler of Kubernetes, called Kube‐scheduler, may cause resource fragmentation and decrease resource utilization. In this paper, we propose a deep reinforcement learning enhanced Kubernetes scheduler named DRS. We initially frame the Kubernetes scheduling problem as a Markov decision process with intricately designed state, action, and reward structures in an effort to increase resource usage and decrease load imbalance. Then, we design and implement DRS mointor to perceive six parameters concerning resource utilization and create a thorough picture of all available resources globally. Finally, DRS can automatically learn the scheduling policy through interaction with the Kubernetes cluster, without relying on expert knowledge about workload and cluster status. We implement a prototype of DRS in a Kubernetes cluster with five nodes and evaluate its performance. Experimental results highlight that DRS overcomes the shortcomings of Kube‐scheduler and achieves the expected scheduling target with three workloads. With only 3.27% CPU overhead and 0.648% communication delay, DRS outperforms Kube‐scheduler by 27.29% in terms of resource utilization and reduces load imbalance by 2.90 times on average.
Current application of intrinsic piezoresistivity of cementitious materials filled with nano conductive substances is limited in sensing quasi-static signals featuring low-frequency and high-amplitude. Its self-sensing capacity for low-amplitude and high-frequency vibro-acousto signals is not fully investigated. In this study, different doses of carbon black nanoparticles (CBN) are added in cementitious materials and their conductivity is measured to find the percolation zone where the tunnelling effect dominates the conduction mechanism, followed by the characterization of their piezoresistivity. Consequently, impact forces and crack-related energy release, which generate dynamic signals in the range of vibrational and acoustic frequencies, are individually utilized to test the behaviour of such materials for sensing dynamic vibro-acousto signals. The results show that CBN-modified cementitious composites which show excellent piezoresistivity behaviour are capable of locally self-sensing dynamic signals as sensors. Comparing with piezoelectrical sensors, the developed sensor shows higher sensitivity to crack initiation.
An accurate and timely cracking assessment, including the presence, location and crack geometric feature measurement, is crucial for evaluating concrete wind towers. Therefore, the early identification of cracks is a critical procedure in promptly evaluating structural integrity. This study proposed an ad-hoc encoder–decoder network based on DeepLabv3+ with depth separable convolutions to automatically segment cracks from real-world images captured from various concrete wind towers. The combined advantages of the improved DeepLabv3+ and the lightweight MobileNet v2 are suitable as a benchmark due to their high performance and universality. Four experiments were conducted to determine the model design choice and crack feature measurement capability: (1) six parametric tests using various pre-trained base networks and algorithm optimisers, (2) the influence of complex background noise (i.e., handwriting script) on crack segmentation performance, (3) comparative studies with cutting-edge pixel-wise segmentation models and (4) crack feature measurement (i.e., length and width). The research outcome demonstrated that DeepLabv3+ with MobileNet v2 can potentially be applied for efficient and accurate crack segmentation in concrete wind towers with complex backgrounds.
As one of the high-strength structures of the aerospace composite panels, the bonding of the stiffened skins affects the mechanical properties of the overall structure. A nonlinear detection method based on subharmonic modulation is proposed in order to detect the debonding damage of composite stiffened structures in the present paper. The debonding damage interface is simplified to a two-degree-of-freedom nonlinear model using interface contact theory. The multi-scale method is employed to analyse the generation mechanism of subharmonic modulation. The experiment is performed on the composite stiffened plate. Then based on the frequency response function characteristics of in-situ piezoelectric actuator/sensor, both high- and low-frequency excitation signals are applied to the composite stiffened plate, and the debonding damage of stiffeners is identified by the subharmonic modulation component in the response signal spectrum. It is shown, both theoretically and experimentally, that the subharmonic modulation method can effectively detect the interfacial debonding damage of composite stiffened structures and is insusceptible to the interference of environmental noise and the inherent nonlinearity of materials.
Multi-exit network is a promising architecture for efficient model inference by sharing backbone networks and weights among multiple exits. However, the gradient conflict of the shared weights results in sub-optimal accuracy. This paper introduces Deep Feature Surgery (DFS), which consists of feature partitioning and feature referencing approaches to resolve gradient conflict issues during the training of multi-exit networks. The feature partitioning separates shared features along the depth axis among all exits to alleviate gradient conflict while simultaneously promoting joint optimization for each exit. Subsequently, feature referencing enhances multi-scale features for distinct exits across varying depths to improve the model accuracy. Furthermore, DFS reduces the training operations with the reduced complexity of backpropagation. Experimental results on Cifar100 and ImageNet datasets exhibit that DFS provides up to a 50.00% reduction in training time and attains up to a 6.94% enhancement in accuracy when contrasted with baseline methods across diverse models and tasks. Budgeted batch classification evaluation on MSDNet demonstrates that DFS uses about 2x fewer average FLOPs per image to achieve the same classification accuracy as baseline methods on Cifar100.
The performance bottleneck of blockchain has shifted from consensus to serial smart contract execution in transaction validation. Previous works predominantly focus on inter-contract parallel execution, but they fail to address the inherent limitations of each smart contract execution performance. In this paper, we propose PaVM, the first smart contract virtual machine that supports both inter-contract and intra-contract parallel execution to accelerate the validation process. PaVM consists of (1) key instructions for precisely recording entire runtime information at the instruction level, (2) a runtime system with a re-designed machine state and thread management to facilitate parallel execution, and (3) a read/write-operation-based receipt generation method to ensure both the correctness of operations and the consistency of blockchain data. We evaluate PaVM on the Ethereum testnet, demonstrating that it can outperform the mainstream blockchain client Geth. Our evaluation results reveal that PaVM speeds up overall validation performance by 33.4×, and enhances validation throughput by up to 46×.
Convolutional Neural Networks (CNNs) can benefit from the computational reductions provided by the Winograd minimal filtering algorithm and weight pruning. However, harnessing the potential of both methods simultaneously introduces complexity in designing pruning algorithms and accelerators. Prior studies aimed to establish regular sparsity patterns in the Winograd domain, but they were primarily suited for small tiles, with domain transformation dictating the sparsity ratio. The irregularities in data access and domain transformation pose challenges in accelerator design, especially for larger Winograd tiles. This paper introduces ”Winols,” an innovative algorithm-hardware co-design strategy that emphasizes the strengths of the large-tiling Winograd algorithm. Through a spatial-to-Winograd relevance degree evaluation, we extensively explore domain transformation and propose a cross-domain pruning technique that retains sparsity across both spatial and Winograd domains. To compress pruned weight matrices, we invent a relative column encoding scheme. We further design an FPGA-based accelerator for CNN models with large Winograd tiles and sparse matrix-vector operations. Evaluations indicate our pruning method achieves up to 80% weight tile sparsity in the Winograd domain without compromising accuracy. Our Winols accelerator outperforms dense accelerator by a factor of 31.7 × in inference latency. When compared with prevailing sparse Winograd accelerators, Winols reduces latency by an average of 10.9 ×, and improves DSP and energy efficiencies by over 5.6 × and 5.7 ×, respectively. When compared with the CPU and GPU platform, Winols accelerator with tile size 8 × 8 achieves 24.6 × and 2.84 × energy efficiency improvements, respectively.