We derive unified lower bounds on the mean squared error (MSE) of distributed quantum sensor fusion under Byzantine faults and decoherence. Building on the classical Brooks-Iyengar overlap function and its vector extension, the predictive outlier model for virtual sensor tracking, and SPOTLESS spatial-temporal verification, we establish a two-parameter family of bounds indexed by entanglement visibility V and fault fraction f/M. For M quantum sensors with N atoms each and sensitivity eta, the MSE of any estimator satisfies MSE >= (1-V^2)/(4*N*eta^2*M_eff) + V^2/(4*N*eta^2*M_eff^2), where M_eff = M-2f under Brooks-Iyengar Byzantine fault tolerance and M_eff = M-f when predictive outlier detection successfully identifies faulty sensors. The bound interpolates continuously between the standard quantum limit (V=0, scaling as 1/sqrt(M_eff)) and the Heisenberg limit (V=1, scaling as 1/M_eff). Monte Carlo simulations with up to 64 sensors validate the theoretical scaling laws. Validation on the Intel Berkeley Lab 54-mote dataset with spatial clustering demonstrates 20-27 dB SNR improvement from entanglement per cluster, and reveals that missing classical sensor data degrades fusion agreement in the same pattern as quantum decoherence. The framework bridges quantum metrology with classical stream-processing architectures including Data-Cleaning Trees and the 80-20 Power Law for scale-invariant clustering.
Large Language Models (LLMs) such as Gemma-2B have shown strong performance in various natural language processing tasks. However, general-purpose models often lack the domain expertise required for cybersecurity applications. This work presents a methodology to fine-tune the Gemma-2B model into a domain-specific cybersecurity LLM. We detail the processes of dataset preparation, fine-tuning, and synthetic data generation, along with implications for real-world applications in threat detection, forensic investigation, and attack analysis. Experiments highlight challenges in prompt length distribution during domain-specific fine-tuning. Uneven prompt lengths limit the model's effective use of the context window, constraining local inference to 200-400 tokens despite hardware support for longer sequences. Chain-of-thought styled prompts, paired with quantized weights, yielded the best performance under these constraints. To address context limitations, we employed a hybrid strategy using cloud LLMs for synthetic data generation and local fine-tuning for deployment efficiency. To extend the evaluation, we introduce a Retrieval-Augmented Generation (RAG) pipeline and graph-based reasoning framework. This approach enables structured alignment with MITRE ATT CK techniques through STIX-based threat intelligence, enhancing recall in multi-hop and long-context scenarios. Graph modules encode entity-neighborhood context and tactic chains, helping mitigate the constraints of short prompt windows. Results demonstrate improved model alignment with tactic, technique, and procedure (TTP) coverage, validating the utility of graph-augmented LLMs in cybersecurity threat intelligence applications.
Fault Injection attack is a type of side-channel attack on the Physical Unclonable Function (PUF) module that can induce faults in the PUF response by manipulating the PUF circuit behavior through voltage glitches, laser attacks, temperature manipulations, or any other attacks potentially leading to information loss or security system failure. This type of attack exposes the physical characteristics of PUFs that can be analyzed to predict or compromise the unique challenge-response pairs (CRPs) reducing the security and reliability of the PUF. Mitigation strategies against such attacks typically include adding noise to the PUF output, using error-correcting codes, or enhanced cryptographic protocols that obscure physical side-channel attacks. In this research, we propose a Generative Adversarial Network (GAN) based security model, that monitors the PUF behavior and detects the variations in PUF response. The model can detect glitches in the PUF response and generate alerts to take mitigation measures.
This article presents a novel hardware-assisted distributed ledger-based solution for simultaneous device and data security in smart healthcare. This article presents a novel architecture that integrates PUF, blockchain, and Tangle for Security-by-Design (SbD) of healthcare cyber–physical systems (H-CPSs). Healthcare systems around the world have undergone massive technological transformation and have seen growing adoption with the advancement of Internet-of-Medical Things (IoMT). The technological transformation of healthcare systems to telemedicine, e-health, connected health, and remote health is being made possible with the sophisticated integration of IoMT with machine learning, big data, artificial intelligence (AI), and other technologies. As healthcare systems are becoming more accessible and advanced, security and privacy have become pivotal for the smooth integration and functioning of various systems in H-CPSs. In this work, we present a novel approach that integrates PUF with IOTA Tangle and blockchain and works by storing the PUF keys of a patient’s Body Area Network (BAN) inside blockchain to access, store, and share globally. Each patient has a network of smart wearables and a gateway to obtain the physiological sensor data securely. To facilitate communication among various stakeholders in healthcare systems, IOTA Tangle’s Masked Authentication Messaging (MAM) communication protocol has been used, which securely enables patients to communicate, share, and store data on Tangle. The MAM channel works in the restricted mode in the proposed architecture, which can be accessed using the patient’s gateway PUF key. Furthermore, the successful verification of PUF enables patients to securely send and share physiological sensor data from various wearable and implantable medical devices embedded with PUF. Finally, healthcare system entities like physicians, hospital admin networks, and remote monitoring systems can securely establish communication with patients using MAM and retrieve the patient’s BAN PUF keys from the blockchain securely. Our experimental analysis shows that the proposed approach successfully integrates three security primitives, PUF, blockchain, and Tangle, providing decentralized access control and security in H-CPS with minimal energy requirements, data storage, and response time.
The rapid adoption of Internet-of-Medical-Things (IoMT) has revolutionized e-health systems, particularly in remote patient monitoring. With the growing adoption of Internet-of-Medical-Things (IoMT) in delivering technologically advanced health services, the security of Medtronic devices is pivotal as the security and privacy of data from these devices are directly related to patient safety. PUF has been the most widely adopted hardware security primitive which has been successfully integrated with various Internet-of-Things (IoT) based applications, particularly in smart healthcare for facilitating device security. To facilitate security and access control to IoMT devices, this work proposes a novel cybersecurity solution using PUF for facilitating global access to IoMT devices. The proposed framework presents an approach that enables the patient’s body area network devices supported by PUF to be securely accessible and controllable globally. The proposed cybersecurity solution has been experimentally validated using state-of-the-art SRAM PUF, a delay based PUF, and a trusted platform module (TPM) primitive.
The scope of Smart electronics and its increasing market worldwide has made cybersecurity an important challenge. The Security-by-Design (SbD) principle, an emerging cybersecurity area, focuses on building security/privacy-enabled primitives at the design stage of an electronic system. This paper proposes a novel Physical Unclonable Function (PUF) based Trusted Platform Module (TPM) for SbD primitive. The proposed SbD primitive works by performing secure verification of the PUF key using TPM’s Encryption and Decryption engine. The securely verified PUF Key is then bound to TPM using Platform Configuration Registers (PCR). PCRs in TPM facilitate a secure boot process and effective access control to TPM’s NonVolatile memory through an enhanced authorization policy. By binding PUF with PCR in TPM, a novel PUF-based access control policy can be defined, bringing in a new security ecosystem for the emerging Internet-of-Everything era. The proposed SbD approach has been experimentally validated by successfully integrating various PUF topologies with Hardware TPM.
This work presents a sustainable cybersecurity solution using Physical Unclonable Functions (PUF), Trusted Platform Module (TPM), and Tangle Distributed Ledger Technology (DLT) for sustainable device and data security. Security-by-Design (SbD) or Hardware- Assisted Security (HAS) solutions have gained much prominence due to the requirement of tamper-proof storage for hardwareassisted cryptography solutions. Designing complex security mechanisms can impact their efficiency as IoT applications are more decentralized. In the proposed architecture, we presented a novel TPM-enabled PUF-based security mechanism with effective integration of PUF with TPM. The proposed mechanism is based on the process of sealing the PUF key in the TPM, which cannot be accessed outside the TPM and can only be unsealed by the TPM itself. A specified NV-index is assigned to each IoT node for sealing the PUF key to TPM using the Media Access Control (MAC) address. Access to the TPM's Non-Volatile Random Access Memory (NVRAM) is defined by the TPM's Enhanced Authorization policies as specified by the Trust Computing Group (TCG). The proposed architecture uses Tangle for sustainable data security and storage in decentralized IoT systems through a Masked Authentication Messaging (MAM) scheme for efficient and secure access control to Tangle. We validated the proposed approach through experimental analysis and implementation, which substantiates the potential of the presented PUFchain 4.0 for decentralized IoT-driven security solutions.
Online tracking is a feature of many state-of-the-art object trackers. When learning online, the data is limited, so the tracker learns a sketch of the object's features. For a tracker to successfully re-identify the same object in the future frames in many different contexts, including occlusions, the tracker has to keep meta-data over time. In multiobjective inferences, this can exponentially increase the costs and is an ill-posed problem. This paper introduces a model-based framework that combines an ensemble of offline pre-trained models cascaded with domain-specific context for spatial tracking. Our method is efficient in reidentifying objects detected by any camera detector as there is minimal online computation. The second model uses a cosine similarity ranking of the label detected by the first model to find its corresponding set of raw images from the domain training set. A high score means model one has previously seen the object, and a low score amounts to a new detection. By using a two-stage AI-trained ensemble at the edge device, we show that the proposed tracker can perform 10 times faster with its precise detection, and the reidentification at the second stage is accurate, avoiding ID flipping for longer durations on video streams.
With many data breaches and spoofing attacks on our networks, it becomes imperative to provide a reliable method for verifying the integrity of the source. Blockchain location-based proof-of-origin is explored for tracking trucks and vehicles. Blockchain applications that support quick authentication with these non-mutable ledger properties: consensus and implemented as smart contracts at the edge. This Blockchain application will now be known as the POWTracker platform, gathering data from multiple cameras. POWTracker is based on an existing GPS-based blockchain ledger and runs on an edge device that uses AI consensus and multiple cameras. By using GPS algorithms, we present a novel mining algorithm that rewards POW miners, providing a trustworthy, verifiable proof-of-location system.
To incorporate object locations in a multi-target detection model, we assume that a close duplicate cannot be learned by the model efficiently. So, we use a region-based approach which uses more object location compared to the ground truth locations to localize the targets. The proposed model is able to learn a similarity metric with respect to the ground truth locations which is robust (low false positives) enough for varying images conditions, small aerial target sizes and using few training samples. We report preliminary results on how transfer learning of meta-data affects small aerial target localization accuracies. Quality ranking from Intersection-over-Union (IOU) in region segmentation models on the aerial ground truth data using pre-trained models from ImageNet, AlexNet, and CIFAR-10 and initialization with three aerial datasets such as the satellite imagery XView2.
The purpose of this paper is to study data fusion applications in traditional, spatial, and aerial video stream applications which addresses the processing of data from multiple sources using co-occurrence information and uses a common semantic metric. Use of co-occurrence information to infer semantic relations between measurements avoids the need to make use of such external information, such as labels. Many of the current Vector Space Models (VSM) do not preserve the co-occurrence information, leading to a less than useful similarity metric. We propose a proximity matrix embedding part of the learning metric representation which has entries showing the relations between co-occurrence frequency observed in input sets. First, we show an implicit spatial sensor proximity matrix calculation using Jaccard similarity for an array of sensor measurements and compare with the state-of-the-art kernel PCA learning from feature space proximity representation; it relates to a k-radius ball of nearest neighbors. Finally, we extend the class co-occurrence boosting of our unsupervised model using pre-trained multi-modal reuse.
Traditional event detection from video frames are based on a batch or offline based algorithms: it is assumed that a single event is present within each video, and videos are processed, typically via a pre-processing algorithm which requires enormous amounts of computation and takes lots of CPU time to complete the task. While this can be suitable for tasks which have specified training and testing phases where time is not critical, it is entirely unacceptable for some real-world applications which require a prompt, real-time event interpretation on time. With the recent success of using multiple models for learning features such as generative adversarial autoencoder (GANS), we propose a two-model approach for real-time detection. Like GANs which learns the generative model of the dataset and further optimizes by using the discriminator which learn per sample difference between generated images. The proposed architecture uses a pre-trained model with a large dataset which is used to boost weekly labeled instances in parallel with deep-layers for the small aerial targets with a fraction of the computation time for training and detection with high accuracy. We emphasize previous work on unsupervised learning due to overheads in training labeled data in the sensor domain.
Ensemble Stream Modeling and Data-cleaning are sensor information processing systems have different training and testing methods by which their goals are cross-validated. This research examines a mechanism, which seeks to extract novel patterns by generating ensembles from data. The main goal of label-less stream processing is to process the sensed events to eliminate the noises that are uncorrelated, and choose the most likely model without over fitting thus obtaining higher model confidence. Higher quality streams can be realized by combining many short streams into an ensemble which has the desired quality. The framework for the investigation is an existing data mining tool. First, to accommodate feature extraction such as a bush or natural forest-fire event we make an assumption of the burnt area (BA*), sensed ground truth as our target variable obtained from logs. Even though this is an obvious model choice the results are disappointing. The reasons for this are two: One, the histogram of fire activity is highly skewed. Two, the measured sensor parameters are highly correlated. Since using non descriptive features does not yield good results, we resort to temporal features. By doing so we carefully eliminate the averaging effects; the resulting histogram is more satisfactory and conceptual knowledge is learned from sensor streams. Second is the process of feature induction by cross-validating attributes with single or multi-target variables to minimize training error. We use F-measure score, which combines precision and accuracy to determine the false alarm rate of fire events. The multi-target data-cleaning trees use information purity of the target leaf-nodes to learn higher order features. A sensitive variance measure such as ƒ-test is performed during each node's split to select the best attribute. Ensemble stream model approach proved to improve when using complicated features with a simpler tree classifier. The ensemble framework for data-cleaning and the enhancements to quantify quality of fitness (30% spatial, 10% temporal, and 90% mobility reduction) of sensor led to the formation of streams for sensor-enabled applications. Which further motivates the novelty of stream quality labeling and its importance in solving vast amounts of real-time mobile streams generated today.
A general nonparametric technique is proposed for the analysis of multi-resolution and multivariate feature space to isolate faulty sensors. The basic overlap function of the technique is an existing one-dimensional fault-detection Brooks-Iyengar algorithm which uses weighted precision and accuracy for static data. We prove the dual of the existing overlap function can isolate the measurement intervals in the multi-dimensional feature space for both labelled and unlabeled publicly available datasets. It is shown that computable complexity of learning the feature space increases linearly with the size of the input. The experimental results showed that by using mean average precision of all sensors using ensemble model for dynamic events. The proposed algorithm performed well in the presence of noise across many static and dynamic action recognition datasets.
Rare event learning has not been actively researched since lately due to the unavailability of algorithms which deal with big samples. The research addresses spatio-temporal streams from multi-resolution sensors to find actionable items from a perspective of real-time algorithms. This computing framework is independent of the number of input samples, application domain, labelled or label-less streams. A sampling overlap algorithm such as Brooks-Iyengar is used for dealing with noisy sensor streams. We extend the existing noise pre-processing algorithms using Data-Cleaning trees. Pre-processing using ensemble of trees using bagging and multi-target regression showed robustness to random noise and missing data. As spatio-temporal streams are highly statistically correlated, we prove that a temporal window based sampling from sensor data streams converges after n samples using Hoeffding bounds. Which can be used for fast prediction of new samples in real-time. The Data-cleaning tree model uses a nonparametric node splitting technique, which can be learned in an iterative way which scales linearly in memory consumption for any size input stream. The improved task based ensemble extraction is compared with non-linear computation models using various SVM kernels for speed and accuracy. We show using empirical datasets the explicit rule learning computation is linear in time and is only dependent on the number of leafs present in the tree ensemble. The use of unpruned trees (t) in our proposed ensemble always yields minimum number (m) of leafs keeping pre-processing computation to n × t log m compared to N2 for Gram Matrix. We also show that the task based feature induction yields higher Qualify of Data (QoD) in the feature space compared to kernel methods using Gram Matrix.
The process of inversion, estimation and reconstruction of the sensor quality matrix, allows modeling the precision and accuracy, and in general the reliability of the model. When the sensor data ranges are not known a priori, current systems do not train on new data samples, rather they approximate based on the parameter's global average value, losing most of the spatial and temporal features. The proposed model, which we call SPOTLESS, checks the spatial integrity and temporal plausibility of streams generated by mobility patterns due to varying channel conditions. We define a minimum quality of the measured sensor data as local stream (QoD) requirements to give high precision by using distributed labeled training. In our SPOTLESS data-cleaning steps, to account for packet errors due to varying channel conditions, a soft-phy based decoding is selected for various Bit Error Rates (BER), minimizing packet loss at the mobile receiver. Numerical experiments for Rayleigh fading channels and mobile BER model examples are compared with large deployment of ground sensor collecting static data streams and Data MULE collecting multi-hop temporal data from the sensor to provide hypothetical parameter accuracy. Our results were obtained in the context of provisioning a minimum precision and accuracy stream (QoD) required for 802.15.4 mobile services. SPOTLESS data-cleaning algorithm coding provides 90% precision for static streams, and increases the plausible relevance of multi-hop mobile streams by 85% for task-based learning.
The baseline discrete parameters capture only the sensor ranges, making event prediction function hard to train with a Gaussian density function, without specific temporal understanding of the datasets. The dynamic features present in a sequence of patterns are localized and used to predict events, which otherwise may be an attributing feature to the static data mining algorithm. The machine learning repository provides collection of supervised databases that are used for the empirical analysis of event prediction algorithms with unsupervised datasets from distributed wireless sensor networks. A measure which combines precision and recall for a small dataset is F-measure and is the weighted harmonic mean of precision and relevance. The availability of such a system is expected to allow more flexible modeling approaches and much more rapid model turnaround for exploratory analysis. From statistical point of view, if the attributes have similar values then it creates high bias creating what is called over-fitting error during learning.
We measure reliability in sensor networks which are dependent on limited resources of individual sensor nodes such has battery capacity, transmission range and channel interference due to simultaneous wireless transmissions. From the initial simulation it is estimated that the routing errors using a distributed algorithm for a large network is less susceptible to failures when compared to using a table driven routing algorithm. To further address other influencing factors which are not related to resource allocation or routing of the sensor network we study the correlated issues, which makes sensor network unique to the categories of wireless network applications. The simulation results show that due to 1-bit-mask accuracy and the CDF codes used to represent measured values in the decoder buffer is fault-tolerant and also increases the communication rate by 70% due to information redundancy within a sensor cluster. Keywords—Sensor Data Reliability, Slepian & Wolf Coding, Cosets, Huffman Trees, BER, Baysian Error.