
The Smart Hat is an advanced wearable assistive system designed to improve mobility and independence for visually impaired users. It integrates voice-guided navigation, real-time object detection, ultrasonic obstacle sensing, and contextual feedback to enable safe and independent movement. Built on a Raspberry Pi 5 platform with edge computing capabilities, the system ensures low-latency processing while utilizing secure cloud services (Firebase, Firestore, Google Cloud Platform) for data storage and analysis. Real-time remote monitoring is facilitated through Ngrok. Experimental validation demonstrates robust performance in obstacle avoidance, accurate navigation, and responsive feedback across diverse environments. By enhancing user autonomy and confidence, this innovative solution highlights the transformative potential of wearable assistive technologies in improving accessibility for the visually impaired.
Cyber attacks on AI systems are recognized as a significant threat. In particular, adversarial attacks, which digitally perturb input data to cause AI to make incorrect predictions, are a serious concern. Recently, such attacks have been pointed out to be applied in the physical world. For instance, attackers place transparent stickers with adversarial perturbations on a camera lens to disrupt AI image perception. When AI-based factory automation (FA) systems are exposed to such threats, the AI systems can, for example, fail to detect defective products or internal system faults. These failures can lead to economic losses and system disruptions. In this paper, we implement the physical adversarial attacks and confirm their feasibility against AI-based FA systems through two experiments: (1) simulation experiments using the open-source FA-specific dataset (MVTec AD), and (2) a real-world experiment where physical stickers perturb camera lenses. As a result, in both evaluations, it is quantitatively shown that our attacks successfully caused malfunction in AI-based FA systems and that images perturbed by our method maintained high similarity to the original images. These experimental results reveal that adversarial attacks are feasible and pose a serious threat to FA systems.
Neural processing units (NPUs) have become essential in modern client and edge platforms, offering unparalleled efficiency by delivering high throughput at low power. This is critical to improve the TOPS/W of the NPU, leading to longer battery life. While NPUs were initially designed to efficiently execute computer vision (CV) workloads such as CNNs, the rising demand to run transformer-based large language models (LLMs) locally now calls for significant architectural and software adaptation. This paper presents LLM-NPU, a comprehensive software-hardware co-optimization framework that enables scalable, power-efficient LLM deployment on NPUs under tight compute and memory budgets. We present software solutions such as vertical and horizontal operator fusion, quantization-aware weight compression, hybrid key-value (KV) quantization, eviction strategies, and static-shape inference that target memory bottlenecks and compute inefficiencies in LLM execution. On the hardware side, we explore domain-specialized NPU enhancements, including processing-in-memory architectures, extended input channel accumulation, structured sparsity acceleration, GEMM engine optimizations, mixed precision, microscaling format support, and fusion-aware execution pipelines. These co-designed innovations can collectively improve the energy, throughput, and latency of NPUs for LLM workloads.
Small Modular Reactors (SMRs) are an emerging nuclear energy technology offering the potential to reduce industrial loads on the electric grid to improve frequency stability and power quality. These compact nuclear facilities provide reliable energy, especially in remote areas. However, integrating SMRs into the bulk electric system introduces new cyber-physical vulnerabilities, including increased exposure to cyberattacks and potential cascading disruptions. This paper introduces a Cyber-Informed Design framework tailored to Cyber-Physical Energy Systems incorporating SMRs. The framework employs attack trees to model threat scenarios and identify vulnerabilities across four functional layers of SMR stations. We apply the threat modeling approach to a reference architecture and map the attack paths to the MITRE ATT&CK for Industrial Control Systems to provide insights into improving the resilience of cyber-physical energy systems.
Caching systems require not only a high hit rate but also adaptability and scalability to effectively manage diverse workloads. This study explores the possible directions of using machine learning for adaptive caching. We explore the implementation of several caching strategies from recent literature, including SIEVE, S3-FIFO, and TinyLFU, highlighting the limitations of traditional methods. We leverage the development of these new caching strategies in combination with machine learning to improve adaptability at small cache sizes.We test LeCaR, a model-driven approach that switches between LRU and LFU, on a real-world trace from Microsoft Azure. We build on it by incorporating multiple combinations of traditional and modern methods. We note that modified_LeCaR, with LRU, LFU, and S3-FIFO, tested on real-world traces, demonstrates improved performance over the original LeCaR, particularly in scenarios with smaller cache sizes. modified_LeCaR’s strong performance on smaller cache sizes makes it well-suited for resource-constrained systems (e.g., embedded systems). Especially because we do not use resource-heavy algorithms (e.g., neural networks).
Non-Terrestrial Networks (NTNs) are rapidly becoming a foundational element of global connectivity, extending coverage and capacity through the integration of terrestrial, airborne, and spaceborne communication platforms. The Open Radio Access Network (O-RAN) architecture, grounded in disaggregation, open interfaces, and intelligent control, provides a promising foundation for orchestrating heterogeneous and dynamic communication systems. However, applying O-RAN principles to NTNs introduces distinct challenges, including limited onboard compute and energy resources, highly dynamic topologies, and latency-sensitive control across spatially distributed domains. This paper introduces O2-RAN, a unified architectural framework designed to support multi-layer, multi-segment NTN deployments. We examine the placement and coordination of key O-RAN components across terrestrial, aerial (e.g., HAPS, UAVs), and orbital (LEO, GEO) segments, and analyze the associated trade-offs, particularly in terms of latency, control distribution, and resource constraints. The discussion highlights architectural strategies for enabling resilient, adaptive, and resource-efficient orchestration across NTN infrastructures, paving the way for scalable and autonomous O-RAN deployments in future NTNs.
Machine learning models are increasingly utilized in critical areas such as finance, hiring, and criminal justice, yet they often inherit or amplify societal biases, leading to unfair outcomes. Addressing algorithmic fairness is no longer optional but essential for building trustworthy systems. This paper proposes a novel framework that integrates fairness as a primary optimization goal during the hyperparameter tuning phase of model development. Using the FLASH algorithm, a fast sequential model-based optimization technique, we demonstrate that it is possible to simultaneously optimize for predictive accuracy and fairness metrics. Our experiments across multiple real-world datasets reveal that incorporating fairness constraints during model optimization significantly reduces bias without substantially compromising performance. Furthermore, the proposed approach outperforms several established bias mitigation techniques. These findings highlight the critical role of software engineers in embedding fairness into the machine learning lifecycle and present a practical, scalable path toward more equitable AI systems.
Conventional robots, such as wheeled or legged designs, offer numerous advantages for specific tasks. However, their rigid structures, limited flexibility, and difficulty in navigating tight spaces or varied terrain present major drawbacks in environments where high adaptability is required. In contrast, snake-shaped robots are inherently more flexible and maneuverable, making them better suited for complex and variable environments where traditional robots face limitations, such as search and rescue, industrial inspection, and exploration of unknown terrain. We developed a modular snake robot featuring a novel active-skin drive system designed to achieve greater flexibility and support complex locomotion. To enable multiple degrees of freedom, a mecanum-inspired active-skin mechanism was implemented, giving each segment omnidirectional movement capabilities. A modular wireless control system, integrated with visual feedback and AI capability, was developed to enhance the robot’s adaptability in dynamic environments and facilitate autonomous operation. The robot demonstrated its capabilities in a range of specialized tasks, including obstacle capture and removal, gap traversal, and ramp climbing, achieving speeds of up to 10 cm/s across various surfaces. Two AI models, one using YOLO object detection and the other employing multimodal AI-based image analysis were evaluated, highlighting their respective strengths and limitations. Future investigation of the active skin locomotion system and AI control methods may allow a more efficient, responsive, and versatile robot.
Network-on-Chip (NoC) facilitates seamless intercommunication in in-memory computing (IMC)-based deep neural network (DNN) accelerators. However, these NoCs are vulnerable to flooding attacks, which generate useless packets during communication to result in significant bandwidth reduction. In this paper, for the first time, we propose CONCEAL, a novel flooding attack pertaining to NoCs in IMC-based DNN accelerators. By devising adversarially crafted stealth packets and injecting them through NoC during inter-tile communication in the accelerator, CONCEAL introduces substantial deterioration in the performance of the NoC, along with reduced computational efficiency encompassing latency and dynamic energy consumption in the accelerator. Our novel attack framework's effectiveness is derived from the fact that it only slightly reduces the application-level classification accuracy of the accelerator. The negligible degradation in accuracy makes CONCEAL difficult to detect and renders the attack to remain stealthy. Our experimental evaluations demonstrate that the proposed CONCEAL framework can furnish up to 425.83% and 257.61% performance degradation in terms of energy and latency respectively (including both communication and computation), while only exhibiting <= 5% (in some cases < 1%) reduction in the classification accuracy of the IMC-based DNN accelerator.
Deep Neural Networks (DNN) are being used in several applications on edge devices in real time. Moreover, multiple DNNs rather than a single DNN are executed at the same time to perform different modalities of a task. State-of-the-art edge devices typically consist of multiple processing units on-chip. Therefore, DNNs must be intelligently scheduled at runtime to harness the capability of multiple processing elements at the fullest. Communication between different processing elements and the memory constitute significant portion of the total latency of DNN execution. Although there are several state-of-the-art DNN scheduling and mapping policies which aim to reduce latency of DNN execution, they do not take communication into account to minimize their execution time. Therefore, in this work we present communication-aware task scheduling for DNNs on multicore systems. Specifically, we profile different DNN layers to obtain communication latency between them and employ a greedy algorithm at runtime to map the DNN layers onto available processing elements on a system. Extensive experimental evaluations with various DNNs used at real time show that our proposed technique is able to minimize the execution latency by up to 46.60% with respect to state-of-the-art DNN mapping technique.
Recommender systems are essential for tailoring user experiences on online platforms, mostly depending on user-content interactions to enhance engagement metrics like click-through rate (CTR). Conventional recommendation models frequently fail to identify higher-order relationships and intricate user behavior patterns, constraining their capacity to deliver precise and varied recommendations. This study examines the utilization of Graph Neural Networks (GNNs) to improve content ranking systems through the generation of task-specific user and content embeddings.
Large Language Models (LLMs) have revolutionized the fields of language understanding and generation, positioning themselves as transformative tools across a wide array of applications as general-purpose task solvers. However, the immense size of LLMs introduces significant challenges for deployment on traditional computing systems, largely due to the inherent von Neumann bottleneck. Recently, Computein-Memory (CiM) architectures have emerged as a groundbreaking solution to accelerate Deep Learning (DL) models, seamlessly integrating computation within the memory fabric to overcome the memory wall bottleneck. While there has been recent progress in mapping generative AIs and LLMs to CiM architectures, the deployment of these models with hundreds of billions of parameters remains very challenging, particularly for edge computing environments. To address the specific challenges of deploying LLMs in memory-centric architectures, we coin the term Language-in-Memory (LiM) to describe a new class of models specifically optimized for such systems. In this context, we present, for the first time, LiM Pruner, a suite of LiM compression techniques specifically designed to induce sparsity in pretrained LLMs. Additionally, we investigate data-path architectures to enhance support for pruned LiM models by addressing potential dislocation challenges. Our comprehensive evaluation of the LiM Pruner on the LLaMA-2 model demonstrates its superior performance compared to traditional pruning techniques.
This study introduces a novel Quantized Convolutional Spiking Neural Network (QCSNN) architecture for real-time arrhythmia detection from electrocardiogram (ECG) signals, designed for deployment on memory-constrained wearable healthcare devices. Unlike conventional models that rely on cloud-based processing or require substantial computational resources, our approach enables on-device inference through a two-stage, fully quantized spiking neural network trained with surrogate gradient descent. The proposed QCSNN classifies ECG signals into four clinically relevant categories—Normal, Supraventricular Ectopic Beat (SVEB), Ventricular Ectopic Beat (VEB), and Fusion (F)—achieving sensitivities of 88.97%, 83.51%, 96.4%, and 82.06%, respectively, and outperforming existing benchmarks while reducing memory usage by 76.2%. Trained with quantization-aware learning and leaky integrate-and-fire (LIF) spiking neurons, this architecture enables accurate, energy-efficient arrhythmia detection within the power and memory limits of smartwatches and portable monitors. This work advances scalable edge artificial intelligence (AI) for digital health, supporting continuous, real-time cardiac monitoring and early intervention.
In this work, we investigate the vulnerability of Static Random Access Memory (SRAM) to data imprinting attacks and propose a secure bit-cell architecture that integrates memristor-based control for conditional data toggling. The design extends the conventional 6T SRAM by incorporating a memristor-assisted path capable of performing XOR operations, thereby enabling in-situ logic and enhanced security. The proposed bit-cell simultaneously supports three key functionalities: logical implication (NOT), XOR-based computation, and secure data toggling. Additionally, we extend this architecture to an array-level configuration that enables simultaneous toggling of multiple bits within a single clock cycle. This array-wide control significantly improves resistance to data imprinting while maintaining computational efficiency. The design was implemented using 22nm GF FD-SOI technology with TiO2-based memristor models. Extensive simulations confirm the functional correctness and robustness of the proposed architecture under various attack scenarios. The results demonstrate that the design offers strong protection against data imprinting with minimal impact on speed and power, making it a promising solution for secure memory in cryptographic and AIoT systems.
Explainable artificial intelligence improves the interpretability of machine learning models, which is essential for reliable decision-making in healthcare. Integrating informationtheoretic (IT) approaches into feature selection (FS) facilitates a more rational assessment of feature importance (FI) by capturing redundant and synergistic contributions. This study proposes a novel combined FI-FS method that extends the High-order Interactions Feature Importance (Hi-Fi) framework by replacing variance-based with IT metrics for improved quantification of high-order feature interactions. The method adapts the Leave One Covariate Out metric to identify feature subsets that maximise or minimise Conditional Mutual Information (CMI), prioritising features that increase synergy and reduce redundancy. Feature interpretation is further supported by analysing how selected variables interact, revealing both individual and joint predictive relevance. The proposed framework is applied to cardiovascular time series from 127 young, healthy individuals recorded at rest and under stress conditions (postural and mental stress). The analysed features are extracted from beat-to-beat electrocardiographic RR intervals, pulse-pulse intervals (PP), and systolic and diastolic blood pressure (SBP, DBP) time series and are computed in the time, frequency, and information domains. FI results show that features reflecting variability, spectral content, and entropy, especially from RR, DBP, and PP, are more informative than static measures. RR features consistently show the highest unique information, while PP contributes mainly through synergy. The FS process reduces feature count by +/- 20%, and a Support Vector Machine trained on the selected features achieves +/- 80% accuracy in multi-class stress classification. Overall, this IT-based Hi-Fi framework captures individual informativeness and complex signal interactions, improving dependency detection, interpretability, and physiological insight.
Sensors- and Sensing-based Human Activity Recognition (HAR) systems play a crucial role in healthcare, assistive technologies, and human-computer interaction. In such a context, recognizing hand-based micro activities remains a significant challenge due to their complexity, variability, and limited availability of labeled data. This paper introduces a novel Few-Shot Learning (FSL) approach for recognizing 24 distinct hand-based Activities of Daily Living (ADLs) using data from a wrist-worn inertial sensor. Our method leverages Prototypical Networks enhanced with hard negative mining to improve class separability and generalization in low-data scenarios. The proposed approach is evaluated on datasets containing data from 43 different subjects, demonstrating its effectiveness in learning from minimal labeled examples (i.e., only 35% of the data). The experimental results demonstrate the robustness of our methodology, achieving competitive accuracy using only four samples x subject x class compared to the entire full dataset for micro-ADL recognition. These findings highlight the potential of FSL in real-world applications where data collection is both costly and time-consuming.
The increasing research efforts in 3D ICs and chiplet technology, driven by the limitations of traditional scaling and the slowdown of Moore's Law, have established Through-Silicon Vias (TSVs) as a pivotal component in VLSI design. In this paper, we design triangular TSVs for space-efficient structuring to enhance integration density and conduct transmission and cross-talk analysis. The key objective of this work is to leverage the offset hexagram arrangement of triangles in TSVs to enhance area efficiency, making it suitable for smaller technology nodes. Furthermore, transmission and crosstalk analysis are performed to demonstrate its capability in ultra-wideband and efficient chipto-chip communication. For 10 TSVs arranged in a row, the triangular offset hexagram pattern occupies 31.96% less area compared to cylindrical and square TSVs, and the area reduction will improve further as the number of TSVs in a row increases. Moreover, the architecture exhibits a -10 dB S-11 bandwidth from 0.01 to 55.51 GHz and demonstrates negligible crosstalk between ports. In addition, we evaluate performance under variations in side length, dielectric thickness, height, and edge distance to ensure a rigorous analysis. Collectively, triangular TSVs show strong potential to remarkably enhance interconnect density while maintaining excellent transmission performance, positioning them as promising candidates for next-generation 3D ICs.
The term "closing the loop" effectively captures the concept of creating a complete, automated security cycle that addresses the current gap between security threat identification and mitigation in hardware design. Current research shows that hardware security verification processes are predominantly manual, labor-intensive, and struggle to scale with increasing design complexity. This work aims to address this challenge by envisioning LLMs to automate both the detection and resolution phases, creating a true closed-loop system for secure hardware design. While providing a survey of LLM-driven hardware security frameworks, this work aims to highlight the critical concept of security loop closure leveraging the power of LLMs and security feedback to drive security closure of hardware designs. The framework outlines the techniques to incorporate comprehensive feedback mechanisms for quality assurance, including automated evaluation metrics, human expert validation, and cross-verification with multiple LLM agents to ensure reliability and accuracy of security assessments.
For effective embedded system design, system architects model, configure, and simulate transaction-level models to explore the design space to find the optimal design candidate for final implementation. Due to the exponential growth in the design complexity of embedded systems and the Internet of Things (IoT), running many TLM simulations has become time-consuming, inefficient, and power-demanding. In this work, we propose a machine learning model that captures and learns the complexities of the SystemC TLM-2.0 loosely-timed contention-aware (LT-CA) models, which consider the critical effect of memory and interconnect contention in system-level design and performance estimation. Our experimental results on the TLM-2.0 LT-CA models of two representative DNNs, GoogLeNet and ResNet, show high accuracy of our proposed model with a mean absolute percentage error (MAPE) of 3.12%. Using the enhanced predictive model, we effectively explore the design space to search and identify the optimal sets of design configurations derived through Pareto analysis supported by specialized performance-evaluating functions. Given the expeditious predictive model, the Pareto analysis shows 8 optimal design candidates out of 1,000 experimental candidates with 3 orders of magnitude speed-up. The proposed framework emphasizes the reduction of time and cost constraints by saving hundreds of hours of TLM simulation, enhancing the overall efficiency of system-level modeling and simulation.
In this paper, we introduce a machine learning-based performance modeling framework that accurately predicts both inference latency and energy consumption of neural networks deployed on edge GPUs, specifically targeting the NVIDIA Jetson Nano platform. To support this effort, we construct a comprehensive benchmark dataset consisting of latency and energy measurements for a wide range of deep learning models, spanning from lightweight architectures with 100 million multiply-accumulate (MAC) operations to large models with up to 50 billion MACs. Our performance modeler demonstrates high predictive accuracy across a diverse set of well-known architectures and generalizes effectively to unseen models. We further integrate our predictor into a hardware-aware neural architecture search (NAS) framework, showing that it accelerates the NAS process from days to hours without sacrificing the quality of the selected architectures. This work highlights the potential of learning-based performance estimation to enable fast, efficient, and hardware-aware deep learning model design for edge deployment.