
The rapid expansion of the Internet of Things (IoT) across industries such as healthcare, manufacturing, transportation, and smart cities has made these networks prime targets for cyber attacks. Due to their distributed nature, device diversity, and resource constraints, traditional cyber security solutions alone are insufficient to protect against evolving threats. In addition, the increasing complexity of managing numerous interconnected devices and the limitations of realtime threat detection heighten the risk of cyber breaches. As a result, researchers and engineers are shifting beyond purely defensive cyber security approaches and focusing instead on recoverability and adaptability through cyber resilience mechanisms. The primary objective of cyber resilience in IoT networks is to go beyond conventional protective layers, ensuring long-term sustainability and strengthening resilience against persistent and sophisticated cyber threats. This survey analyses the cyber resilience concept and its steps in IoT networks and outlines challenges in providing cyber resilience in these networks. We review existing definitions of cyber resilience, highlighting their limitations in the IoT context. Also, the relationship between the key security features of IoT networks and cyber resilience is examined. We categorise proposed cyber resilience mechanisms according to their operational layers within the IoT architecture and evaluate them across multiple dimensions, including resilience phases, alignment with IoT requirements, the application domain and employed techniques. Furthermore, this survey examines several directions for future research by highlighting the diverse challenges posed by the various facets of IoT networks within this research domain. The findings of this research contribute to the existing body of knowledge on IoT security and cyber resilience while laying the groundwork for future research and development. Ultimately, this survey seeks to support the development of effective and sustainable strategies to ensure the security and resilience of IoT networks in the face of evolving cyber threats.
Despite the remarkable achievements of deep neural networks (DNNs) in numerous fields, the growing number of parameters and computational complexity severely limit their deployment feasibility on edge devices. Against this backdrop, lightweight DNNs have not only become a hot topic in academic research but also a key technological pathway to promote the democratization and implementation of AI. This article focuses on reviewing the methods of designing lightweight DNN architectures to achieve model lightweighting, aiming to provide researchers with effective solutions for designing lightweight model architectures. The article distinguishes between convolutional-based and Transformer-based frameworks for model design and delves into several typical lightweight model structural designs, development paths, and their pros and cons. By experimentally comparing the lightweighting metrics of different models, this article points out that model selection needs to be closely integrated with the constraints of specific application scenarios. Finally, the article observes that future breakthroughs may lie in exploring the lightweighting of hybrid architectures that combine convolution and Transformer, to integrate the advantages of local perception and global modeling, and further enhance model expressiveness while maintaining efficiency. In summary, this article not only provides a comprehensive review of lightweight model structural design but also emphasizes its practical guidance and development direction in promoting the implementation of edge intelligence.
Retrieval-Augmented Generation (RAG) enhances AI-generated content by integrating external knowledge, improving relevance, and reducing hallucinations. However, RAG also introduces risks related to reliability, safety, privacy, fairness, explainability, and accountability, which impact trustworthiness. While various methods aim to address these concerns, a unified framework is lacking. This survey bridges that gap by presenting a comprehensive roadmap for trustworthy RAG systems. We provide a structured analysis of key challenges, existing solutions, and future directions across these aspects. We further organize the field around how the retrieval and generation stages and the six trustworthiness dimensions interact, demonstrating their synergies and trade-offs through a controlled cross-dimensional analysis. Additionally, we highlight downstream applications where trustworthy RAG can make a significant impact, encouraging further research and adoption in real-world AI systems.
Due to their widespread applications in decentralized and privacy preserving technologies, commitment schemes have become increasingly important cryptographic primitives. With a wide variety of applications, many new constructions have been proposed, each enjoying different features and security guarantees. In this article, we systematize the designs, features, properties, and applications of vector commitments (VCs). We define vector, polynomial, and functional commitments and we discuss the relationships shared between these types of commitment schemes. We first provide an overview of the definitions of the commitment schemes we will consider, as well as their security notions and various properties they can have. We proceed to compare popular constructions, taking into account the properties each one enjoys, their proof/update information sizes, and their proof/commitment complexities. We also consider their effectiveness in various decentralized and privacy preserving applications. Finally, we conclude by discussing some potential directions for future work.
As next-generation systems become increasingly complex and interconnected, they face the burden of new and numerous critical security challenges. Capability Hardware Enhanced RISC Instructions (CHERI) has emerged as a promising technology to mitigate memory safety vulnerabilities, one of the major security threats in current computing systems. CHERI introduces a hardware-enforced, fine-grained memory protection model that integrates a fat-pointer representation, combining memory bounds and permissions to ensure safer memory access and mitigate common memory-related exploits. Despite its potential, information about CHERI evolution, design principles, and applications is scattered among multiple sources, making it difficult to fully grasp and correlate its features and applications. This article aims to fill such gap by providing a detailed survey of CHERI technology, covering its historical evolution, key concepts, and current hardware and software implementations. We explore the ongoing challenges of its adoption in industry, the research efforts driving its growth as well as speculating and discussing future research directions.
This is a corrigendum for the article "40 Years of Designing Code Comprehension Experiments: A Systematic Mapping Study" published in ACM Comput. Surv. 56, 4, Article 106 (November 2023), 42 pages.
The Industrial Internet of Things (IIoT) is characterized by the generation of vast amounts of time-series data. Modern IIoT systems enable efficient collection, storage, and querying of massive industrial time-series data, making the processing and analysis of such data a key enabler for data-driven decision-making in modern manufacturing. To provide researchers and practitioners with comprehensive guidance on industrial time series data analysis, this paper presents a systematic review of state-of-the-art methods—spanning statistical approaches, machine learning (ML), deep learning (DL), and cutting-edge large models—along with their applications in industrial decision-making. It details the application status of these methods in key equipment condition monitoring, manufacturing process supervision, and energy network management. Additionally, the paper discusses existing gaps between methods and real-world applications, as well as future trends and challenges, such as optimizing data structures for cost-sensitive learning, exploring causality and time-series-oriented model architectures, and developing cascaded/hybrid pipelines for end-to-end industrial use cases. Ultimately, this review aims to inspire innovations in realizing data-driven intelligent decision-making for next-generation IIoT systems.
Transformers have revolutionized Natural Language Processing (NLP) by capturing complex patterns in data, but they still face challenges in modeling long-range dependency, especially in tasks like Intelligent Document Understanding (IDU), where maintaining coherence across lengthy and varied documents is crucial. This survey reviews Transformer-based models for addressing long-range dependency challenges in IDU. It introduces a taxonomy of key approaches and discusses representative models within memory-focused, attention-focused, and augmentation-focused strategies. Then it explores the types of input modalities and datasets used in IDU, as well as practical applications in generation, text classification, information extraction, multimodal document understanding and sequence modeling. This survey also discusses practical deployment challenges that arise when applying these models to long and complex documents in real-world document processing pipelines. Finally, we outline future research directions, including enhancing memory efficiency, optimizing attention mechanisms, improving long-sequence processing, and advancing multimodal integration. We also emphasize the importance of developing models that can effectively handle complex multimodal documents, incorporate domain-specific knowledge (e.g., legal, medical) and generalize across domains for broader applicability. Overall, this paper highlights advancements in Transformer-based models for long-range dependency handling in IDU and identifies key research directions for building more scalable and structurally aware document understanding systems.
Quantum noise poses a significant challenge for current near-term quantum computing. Quantum error mitigation (QEM) has therefore emerged as a key strategy, offering a practical and effective solution to reduce error impacts and enhance the performance of near-term variational quantum circuits (VQC). Given the lack of a comprehensive survey on this important topic in the literature, this article provides a dedicated overview of QEM, including both during and after the training of VQC. Specifically, during VQC training, we explore and discuss key QEM techniques such as optimal control and dynamical decoupling. For the post-processing stage of VQC, we will examine and discuss mitigation techniques such as zero-noise extrapolation (ZNE) and probabilistic error cancellation (PEC). For each of these QEM techniques, we will investigate the fundamentals, mitigation concepts, and recent advances. Subsequently, we also explore research toolboxes, including Mitiq and Qiskit Aer, as well as provide a case study that demonstrates contextual multi-armed bandit-guided ZNE. Finally, we discuss ongoing problems and future research initiatives, including ML-assisted mitigation and integration with error correction. This survey paper aims to synthesize the state-of-the-art in QEM, offer organized insights across approaches, and propose potential paths towards error-resilient quantum computing.
Code generation represents a critical intersection of Software Engineering (SE) and Artificial Intelligence (AI). Within this broader landscape, Verilog, as a representative hardware description language (HDL), is fundamental to Electronic Design Automation (EDA), recent research has increasingly focused on leveraging Large Language Models (LLMs) to automate Verilog code generation, particularly at the Register Transfer Level (RTL) design. Despite growing interest, a comprehensive survey of this domain remains absent. This review addresses this gap by providing a systematic literature review of LLM-based Verilog code generation, analyzing 102 papers (70 published and 32 high-quality preprints) from SE, AI, and EDA venues. We structure our analysis around four key research questions: (1) identifying the LLMs utilized, (2) examining evaluation datasets and metrics, (3) categorizing generation techniques, and (4) analyzing alignment approaches. Furthermore, we synthesize findings to identify critical limitations in current studies regarding effectiveness and integration. Finally, we outline a roadmap highlighting potential opportunities for future research in LLM-assisted hardware design.
This tutorial provides an accessible and implementation-oriented introduction to Gaussian process learning-based model predictive control (GP-MPC), which combines probabilistic residual modeling with receding-horizon control. Its central tutorial contribution is a detailed, step-by-step derivation of multi-step mean and covariance propagation for GP-augmented prediction models. The derivation shows how commonly used propagation formulas follow from the laws of total expectation and total covariance, and clarifies the roles of the GP posterior mean, GP posterior covariance, uncertain inputs, and query–output cross-covariances. Building on this foundation, the article distinguishes mean-only unconstrained GP predictive control from uncertainty-aware constrained GP-MPC, clarifies regulation and output-tracking formulations, and provides concise implementation and computational-complexity guidance, including practical software-tool references and discussion of the scope of closed-loop guarantees. Mobile-robot path-following examples illustrate mean-only unconstrained implementations, whereas mixed-vehicle platooning illustrates uncertainty-aware constrained GP-MPC with uncertainty-dependent safety-distance tightening. The tutorial is intended to help researchers and practitioners understand, implement, and critically evaluate GP-MPC designs for robotic and autonomous systems operating under modeling uncertainty.
Recent advances at the intersection of reinforcement learning (RL) and Multimodal Foundation Models have enabled agents that not only perceive complex visual scenes but also reason, generate, and act within them. This survey offers a critical and up-to-date synthesis of the field. We first formalize visual RL problems and trace the evolution of policy-optimization strategies from RLHF to verifiable reward paradigms, and from Proximal Policy Optimization to Group Relative Policy Optimization. We then organize more than 200 representative works into four thematic pillars: multi-modal large language models, visual generation, unified model frameworks, and vision-language-action models. For each pillar we examine algorithmic design, reward engineering, benchmark progress, and we distill trends such as curriculum-driven training, preference-aligned diffusion, and unified reward modeling. Finally, we review evaluation protocols spanning policy-level, trajectory-level preference, and training diagnostic stability, and we identify open challenges that include sample efficiency, generalization, and safe deployment. Our goal is to provide researchers and practitioners with a coherent map of the rapidly expanding landscape of visual RL and to highlight promising directions for future inquiry. Resources are available at: https://github.com/weijiawu/Awesome-RL-for-Multimodal-Foundation-Models.
Through analyzing and mining the relationships among different objects, graph processing is playing an increasingly important role in various application domains, such as social network analysis, product recommendation, and traffic planning. Unfortunately, real-world graphs often exhibit enormous sizes (i.e., trillions of vertices and edges) and complex structures, which makes large-scale in-memory graph processing extremely challenging, if not impractical, and necessitates out-of-core approaches. Therefore, numerous out-of-core graph processing systems have been developed in recent years to efficiently store and process these large graphs. By exploiting the low-cost HDD-/SSD-based external storage and designing disk-friendly graph data placement and execution models, these systems can achieve relatively good performance with low hardware costs, making them a cost-effective solution for large-scale graph analytics. In this article, we conduct a survey on the designs and implementations of out-of-core graph processing systems. Specifically, we review the key techniques in different dimensions of optimization for out-of-core graph processing systems, including graph preprocessing, graph algorithm execution, utilization of emerging storage devices, and miscellaneous optimizations. For each dimension, we analyze the technical challenges and provide critical insights. Furthermore, we explore and discuss the opportunities for future research of out-of-core graph processing systems. This survey will help researchers better understand and gain useful insights into the large and complex design space of out-of-core graph processing.
Rapid AI development across industries raises pressing security and privacy risks. This work presents a unified comparison of large language models, AI agents, and embodied agents, introducing a taxonomy of risks spanning data, models, systems, content, and applications, alongside a catalog of 24 specific threats. We contrast attack surfaces and methods across the three system types to reveal common patterns and distinctive vulnerabilities. We also survey mainstream AI security assessment frameworks and evaluate how relevant laws and regulations currently address these risks. Finally, we outline concrete directions for future research and practice aimed at building robust, secure agent ecosystems.
Compute centers have passed through several major evolutions during the past 30 years. From the storage systems perspective, the evolution has fundamentally changed the way data is structured, represented, stored and accessed, moving from supercomputer attached file systems to massively distributed services. Throughout this period, new hardware technologies, whether for computing or for data storage, have in turn brought major changes, even broken new ground, and posed real scientific challenges. To address these challenges, innovative concepts and paradigms have been introduced with various implementations and integrated into large-scale solutions which remain in continuous evolution, reconciling hardware advances, application requirements, and contextual constraints. This article retraces the milestones of this technological journey in a chronological order, starting with the Petascale period and the introduction of IO-Proxies to meet the need for increasing storage capacity and reduced access latency, exploring a wealth of data structures, protocols, and other software components to build parallel, efficient and scalable solutions. The challenges of the Exascale era were even greater and innovations were introduced to address them, especially in terms of scheduling policies to match the complexity of the associated workflows and dataflows. Then, the IO-Proxies evolved to a more versatile paradigm implementing the Ephemeral Services.
Recently, with advances in Large Language Models(LLMs), robot navigation models have demonstrated superior generalization capabilities across environment perception, decision-making, reasoning, planning, instruction understanding, and human-robot interaction. In this article, we systematically review recent LLM-based robot navigation research articles and categorize them into a novel taxonomy comprising perception, planning, control, interaction, and coordination. We also present an overview of the principal datasets, simulations, and metrics used in robot navigation, analyzing the distinctive characteristics of the datasets and the performance of the main LLM-based methods. Furthermore, we discuss the challenges hindering the integration of LLMs into robot navigation and provide opportunities and potential directions for future development.
Recommender systems have become integral to personalized content delivery, with deep learning (DL) techniques substantially improving their accuracy and scalability. We examine the integration of generative models into context-aware recommender systems, addressing challenges related to dynamic, partially observable, and latent user contexts. Moreover, we present foundational definitions of context, strategies for incorporating contextual cues, and the role of sequential and interactive generative systems in enhancing recommendation quality. Finally, we explore how generative models enable richer representations of user behaviors and facilitate generating and evaluating contextually relevant recommendations.
Scan chains inserted into the circuit can be exploited as side channels for attackers. This has led to a growing need for secure scan design. The research in the secure scan design started from primitive techniques that involve adding specific gates in the circuit to hinder the interpretation of output data. Subsequently, various attack methods emerged, leading to the differentiation of secure scan design into various categories. In this survey article, the technical goal of each approach is introduced first. Subsequently, research evolution of secure scan designs is analyzed, aiming to present the directions for future studies.
The digital asset market is an emerging market of blockchain-based cryptographic assets, characterized by extreme volatility, non-linear dependencies, and a rapidly evolving ecosystem. Studies on optimizing digital-asset-only portfolios are still sparse, fragmented, and undocumented. This survey provides a comprehensive examination of portfolio optimization methods tailored to this market. Drawing on 119 publications discovered via a systematic literature review spanning from 2017 to 2025, we present a detailed bibliometric analysis that illuminates research trends, publication patterns, and thematic gaps. We also provide an overview of this space, focusing on the most commonly used portfolio optimization approaches, evaluation metrics, and digital assets involved. The relevant literature is organized into four primary categories: (a) traditional and statistical methods, (b) evolutionary algorithms and swarm intelligence, (c) machine learning and deep learning, and (d) reinforcement learning. Within each category, we describe the most frequently used portfolio optimization methods, while highlighting representative works that illustrate the strengths and weaknesses of these specific approaches in digital-asset-only portfolios. By consolidating previously fragmented literature, this survey fosters a holistic understanding of digital asset portfolio optimization, providing a roadmap for future investigation, while serving as a point of reference offering guidance to researchers and practitioners navigating this evolving field.
Entity Resolution (ER) is a fundamental data integration task aimed at identifying records that refer to the same real-world entity. While traditional methods relied on brittle, handcrafted features, the recent shift to Deep Learning (DL) enables automatic feature learning via semantic embeddings, dramatically improving performance, particularly on noisy and unstructured textual data. This survey provides a structured overview of this rapidly advancing field. We chart the evolution of the embedding models that underpin modern ER, from static word vectors to context-aware Transformers. Following the canonical blocking and matching pipeline, we analyze state-of-the-art DL techniques for both stages, categorizing them by their learning paradigms and architectures. We also examine the emerging frontier of using Large Language Models (LLMs) in few-shot and chain-of-thought settings. Finally, we synthesize our findings, evaluate the limitations and inherent difficulty of existing benchmark datasets, discuss critical challenges like fairness and explainability, and outline key directions for future research.