6G networks are expected to be AI-native, intent-driven, and economically programmable, requiring fundamentally new approaches to network slice orchestration. Existing slicing frameworks, largely designed for 5G, rely on static policies and manual workflows and are ill-suited for the dynamic, multi-domain, and service-centric nature of emerging 6G environments. In this paper, we propose an agentic AI control plane architecture for 6G network slice orchestration, monitoring, and trading that treats orchestration as a holistic control function encompassing slice planning, deployment, continuous monitoring, and economically informed decision-making. The proposed control plane is realized as a layered architecture in which multiple cooperating AI agents. To support flexible and on-demand slice utilization, the control plane incorporates market-aware orchestration capabilities, allowing slice requirements, pricing, and availability to be jointly considered during orchestration decisions. A natural language interface, implemented using the Model Context Protocol (MCP), enables users and applications to interact with control-plane functions through intent-based queries while enforcing safety and policy constraints. To ensure responsible and explainable autonomy, the control plane integrates fine-tuned large language models organized as a multi-model consortium, governed by a dedicated reasoning model. The proposed approach is evaluated using a real-world testbed with multiple mobile core instances (e.g Open5GS) integrated with Ericsson's RAN infrastructure. The results demonstrate that combining agentic autonomy, closed-loop SLA assurance, market-aware orchestration, and natural language control enables a scalable and adaptive 6G-native control plane for network slice management, highlighting the potential of agentic AI as a foundational mechanism for future 6G networks.
AI agent-based systems are becoming increasingly integral to modern software architectures, enabling autonomous decision-making, dynamic task execution, and multimodal interactions through large language models (LLMs). However, these systems introduce novel and evolving security challenges, including prompt injection attacks, context poisoning, model manipulation, and opaque agent-to-agent communication, that are not effectively captured by traditional threat modeling frameworks. In this paper, we introduce ASTRIDE, an automated threat modeling platform purpose-built for AI agent-based systems. ASTRIDE extends the classical STRIDE framework by introducing a new threat category, A for AI Agent-Specific Attacks, which encompasses emerging vulnerabilities such as prompt injection, unsafe tool invocation, and reasoning subversion, unique to agent-based applications. To automate threat modeling, ASTRIDE combines a consortium of fine-tuned vision-language models (VLMs) with the OpenAI-gpt-oss reasoning LLM to perform end-to-end analysis directly from visual agent architecture diagrams, such as data flow diagrams(DFDs). LLM agents orchestrate the end-to-end threat modeling automation process by coordinating interactions between the VLM consortium and the reasoning LLM. Our evaluations demonstrate that ASTRIDE provides accurate, scalable, and explainable threat modeling for next-generation intelligent systems. To the best of our knowledge, ASTRIDE is the first framework to both extend STRIDE with AI-specific threats and integrate fine-tuned VLMs with a reasoning LLM to fully automate diagram-driven threat modeling in AI agent-based applications.
The rise of 5G networks has improved connectivity, speed, and low-latency communication but also introduces new security risks, particularly in the radio frequency (RF) domain. Traditional security methods and packet-based intrusion detection systems often fail to detect real-time RF attacks such as jamming, spoofing, or unauthorized signal access. Emerging AI-based RF classifiers are typically black-box models, limiting their use in safety-critical 5G environments that require explainable and accountable decisions. To address these challenges, we propose Deep-RF, an Agentic AI framework for RF signal classification that combines consortium of fine-tuned vision-language models (VLMs) with a reasoning large language model (OpenAI-gpt-oss LLM) for real-time attack detection. The VLMs are trained on labeled RF signal images to detect anomalies, and their predictions are refined through a consensus-driven reasoning process by the LLM, which enforces policies and provides traceable explanations. AI agents manage interactions between the VLMs and LLM, enabling automated, reliable, and transparent decision-making. Experimental results show that this approach delivers an accurate, robust, and explainable system for detecting RF threats in 5G networks. To our knowledge, this is the first application of a VLM consortium integrated with a reasoning LLM for RF signal classification, offering a responsible AI foundation for next-generation 5G security.
The World Wide Web was built on an assumption held for three decades: the primary consumer of web content is a human being. This permeates every layer; its access model presumes human visitors, its economics rest on human attention, and its content targets human perception. The rapid emergence of AI agents as intermediaries between humans and web content invalidates this assumption. Yet the web resists agents through blanket blocking, CAPTCHA-based exclusion, and economic models that treat agent access as extraction rather than legitimate interaction. This paper proposes a principled redesign across three layers. At the access layer, agents acting for humans should inherit equivalent access rights, governed by rate limiting and agent identification metadata in HTTP requests, analogous to browser headers, alongside a dual-layer architecture serving human-readable and agent-optimized content from the same domain. At the economic layer, we propose an intent-based tier framework grounded in the agent-as-human-proxy principle: an agent's economic obligation mirrors that of the human it represents. A token-based subscription model meters content in tokens rather than pageviews, alongside a commissioned content economy anchoring AI content production in human intentionality. At the content layer, we identify epistemic recursion, the self-referential loop in which AI-generated content is consumed by agents to produce further content, progressively detaching web knowledge from human ground truth. We propose the Agent Text Markup Language (ATML), a four-level human supervision tier model, and a cryptographic provenance chain to counter this threat. Together these constitute ten design principles for an agent-first internet, one in which agents are first-class citizens whose integration requires renegotiating the web's foundational social contract across access, economics, and content.
The accelerating adoption of large language models, retrieval-augmented generation pipelines, and multi-agent AI workflows has created a structural governance crisis. Organizations cannot govern what they cannot see, and existing compliance methodologies built for deterministic web applications provide no mechanism for discovering or continuously validating AI systems that emerge across engineering teams without formal oversight. The result is a widening trust gap between what regulators demand as proof of AI governance maturity and what organizations can demonstrate. This paper proposes AI Trust OS, a governance architecture for continuous, autonomous AI observability and zero-trust compliance. AI Trust OS reconceptualizes compliance as an always-on, telemetry-driven operating layer in which AI systems are discovered through observability signals, control assertions are collected by automated probes, and trust artifacts are synthesized continuously. The framework rests on four principles: proactive discovery, telemetry evidence over manual attestation, continuous posture over point-in-time audit, and architecture-backed proof over policy-document trust. The framework operates through a zero-trust telemetry boundary in which ephemeral read-only probes validate structural metadata without ingressing source code or payload-level PII. An AI Observability Extractor Agent scans LangSmith and Datadog LLM telemetry, automatically registering undocumented AI systems and shifting governance from organizational self-report to empirical machine observation. Evaluated across ISO 42001, the EU AI Act, SOC 2, GDPR, and HIPAA, the paper argues that telemetry-first AI governance represents a categorical architectural shift in how enterprise trust is produced and demonstrated.
Post-Traumatic Stress Disorder (PTSD) is fundamentally a neuroplastic problem traumatic contact events encode over-reactive neural pathways through Hebbian long-term potentiation, producing hair-triggered amygdala-HPA stress cascades that fire before conscious awareness can intercept them. Existing therapeutic approaches, prolonged exposure, EMDR, cognitive behavioural therapy, operate predominantly downstream of the reactive cascade, teaching patients to tolerate or reframe distress after it has arisen. While clinically valuable, these suppression-based approaches do not produce the upstream pathway dissolution that constitutes lasting structural neural reorganisation. This paper proposes MindGap, a privacy-preserving on-device conversational AI framework that delivers structured neuroplastic rehabilitation for PTSD through the practice of dependent origination, a Buddhist psychological framework that identifies the precise moment between the pre-cognitive affective signal and the reactive elaboration that follows as the site of therapeutic intervention. MindGap guides patients through three progressive layers of observation at this feeling tone gap: noticing the bare affective signal before reactive elaboration, recognising it as self-arising rather than caused by the stimulus, and recognising the conditioned implicit belief beneath the feeling. Each layer corresponds to progressively deeper prefrontal regulatory engagement and progressively deeper long-term depression-mediated weakening of the reactive pathway, producing genuine upstream dissolution rather than downstream suppression. Running entirely on-device with no data egress, MindGap delivers daily calibrated exposure sessions through a fine-tuned lightweight large language model, making it deployable in sensitive clinical and military contexts where cloud-based solutions are not permitted.
Retail supply chain operations in supermarket chains involve continuous, high-volume manual workflows spanning demand forecasting, procurement, supplier coordination, and inventory replenishment, processes that are repetitive, decision-intensive, and difficult to scale without significant human effort. Despite growing investment in data analytics, the decision-making and coordination layers of these workflows remain predominantly manual, reactive, and fragmented across outlets, distribution centers, and supplier networks. This paper introduces Flowr, a novel agentic AI framework for automating end-to-end retail supply chain workflows in large-scale supermarket operations. Flowr systematically decomposes manual supply chain operations into specialized AI agents, each responsible for a clearly defined cognitive role, enabling automation of processes previously dependent on continuous human coordination. To ensure task accuracy and adherence to responsible AI principles, the framework employs a consortium of fine-tuned, domain-specialized large language models coordinated by a central reasoning LLM. Central to the framework is a human-in-the-loop orchestration model in which supply chain managers supervise and intervene across workflow stages via a Model Context Protocol (MCP)-enabled interface, preserving accountability and organizational control. Evaluation demonstrates that Flowr significantly reduces manual coordination overhead, improves demand-supply alignment, and enables proactive exception handling at a scale unachievable through manual processes. The framework was validated in collaboration with a large-scale supermarket chain and is domain-independent, offering a generalizable blueprint for agentic AI-driven supply chain automation across large-scale enterprise settings.
Quantum walks offer promising advantages for search algorithms over graphs. Among these, Grover's quantum search provides a quadratic speedup with a time complexity of 0 (VN) for unstructured search problems. Nevertheless, Grover search performs poorly on cyclic graphs due to the dynamics of the Grover walk. This study explores the behavior of Grover's quantum walk on cyclic graphs (CN), analyzing the probability distribution of finding the marked vertex. To analyze this behavior, we extend each vertex of the cycle by attaching semi-infinite-length paths (tails). We develop a direct analytical approach to obtain the transition matrix T, whose elements are independent of N. For small cycles C3-CS, we observe success probabilities exceeding 0.5. In contrast, with large cycles, the effectiveness drops rapidly. We observe that the success probability approaches zero (P(u*)-* 0) as the size of the graph increases (N-* 00), indicating the limitations of the quantum search in cyclic structures.
Modern 5G networks offer a network-sliced infrastructure where each network slice contains a dedicated 5G core software service layer. The 5G core software services in each slice shares common core network resources to meet specific customer needs. A primary challenge in 5G network slicing involves resource sharing and efficient network slice orchestration. Container-based methodologies, including tools like Docker and Kubernetes, have become popular for orchestrating 5G network slice services and managing configurations in microservices-based cloud-native service deployment. However, despite their utility, these tools present significant challenges. Their complexity often necessitates dedicated DevOps teams for effective management, while configuration management can prove arduous, and end-to-end supply chain oversight is lacking. To address these challenges, this paper introduces “Llama-Recipe,” a cloud-native 5G-core service deployment and orchestration platform integrating Generative AI, SBOM, PBOM and NFT. 5G-core service configurations across different network slices are represented as “HOCON (Human-Optimized Config Object Notation)” config objects adhering to the GitOps paradigm. Leveraging custom-trained Meta's Llama2 LLM, Llama-Recipe generates the Kubernetes manifests for network-sliced 5G-core services based on the defined HOCON configurations. The generated Kubernetes manifests of the 5G-core services are deployed in designated Kubernetes clusters utilizing GitOps tools (e.g., ArgoCD), ensuring seamless and automated deployment processes. Additionally, Llama-Recipe introduced a novel mechanism to handle end-to-end supply chain verification of 5G-core software services using Software-Bill of Materials (SBOM) and Pipeline-Bill of Materials (PBOM). SBOMs track all the dependencies and PBOMs facilitate the comprehensive tracking of end-to-end supply chain data for 5G-core software services, enhancing transparency and security. These PBOMs are also generated using the fine-tuned Meta's Llama-2 LLM and are encoded as NFT tokens with a novel NFT token schema. This schema enables easy verification and validation of supply-chain data during deployments, thus helping to prevent various supply-chain attacks. To fine-tune the Meta's Llama2 LLM, we've undertaken a meticulous training process, collaborating with Qlora to transform a 4-bit quantized pre-trained language model into Low-Rank Adapters(LoRA). The effectiveness of the Llama-Recipe is demonstrated through a real-world test-bed deployment in a sliced network scenario, utilizing multiple 5G cores (i.e., Open5GS) across Ericsson's new Radio Access Network (RAN).
In recent years, blockchain has experienced widespread adoption across various industries, becoming integral to numerous enterprise applications. Concurrently, the rise of generative AI and LLMs has transformed human-computer interactions, offering advanced capabilities in understanding and generating human-like text. The introduction of the MCP has further enhanced AI integration by standardizing communication between AI systems and external data sources. Despite these advancements, there is still no standardized method for seamlessly integrating LLM applications and blockchain. To address this concern, we propose "MCC: Model Context Contracts" a novel framework that enables LLMs to interact directly with blockchain smart contracts through MCP-like protocol. This integration allows AI agents to invoke blockchain smart contracts, facilitating more dynamic and context-aware interactions between users and blockchain networks. Essentially, it empowers users to interact with blockchain systems and perform transactions using queries in natural language. Within this proposed architecture, blockchain smart contracts can function as intelligent agents capable of recognizing user input in natural language and executing the corresponding transactions. To ensure that the LLM accurately interprets natural language inputs and maps them to the appropriate MCP functions, the LLM was fine-tuned using a custom dataset comprising user inputs paired with their corresponding MCP server functions. This fine-tuning process significantly improved the platform's performance and accuracy. To validate the effectiveness of MCC, we have developed an end-to-end prototype implemented on the Rahasak blockchain with the fine-tuned Llama-4 LLM. To the best of our knowledge, this research represents the first approach to using the concept of Model Context Protocol to integrate LLMs with blockchain.
Software services are crucial for reliable communication and networking; therefore, Site Reliability Engineering (SRE) is important to ensure these systems stay reliable and perform well in cloud-native environments. SRE leverages tools like Prometheus and Grafana to monitor system metrics, defining critical Service Level Indicators (SLIs) and Service Level Objectives (SLOs) for maintaining high service standards. However, a significant challenge arises as many developers often lack in-depth understanding of these tools and the intricacies involved in defining appropriate SLIs and SLOs. To bridge this gap, we propose a novel SRE platform, called SRE-Llama, enhanced by Generative-AI, Federated Learning, Blockchain, and Non-Fungible Tokens (NFTs). This platform aims to automate and simplify the process of monitoring, SLI/SLO generation, and alert management, offering ease in accessibility and efficy for developers. The system operates by capturing metrics from cloud-native services and storing them in a time-series database, like Prometheus and Mimir. Utilizing this stored data, our platform employs Federated Learning models to identify the most relevant and impactful SLI metrics for different services and SLOs, addressing concerns around data privacy. Subsequently, fine-tuned Meta's Llama-3 LLM is adopted to intelligently generate SLIs, SLOs, error budgets, and associated alerting mechanisms based on these identified SLI metrics. A unique aspect of our platform is the encoding of generated SLIs and SLOs as NFT objects, which are then stored on a Blockchain. This feature provides immutable record-keeping and facilitates easy verification and auditing of the SRE metrics and objectives. The automation of the proposed platform is governed by the blockchain smart contracts. The proposed SRE-Llama platform prototype has been implemented with a use case featuring a customized Open5GS 5G Core.
Modern processors tend to incorporate multiple CPU cores. These multiple CPU cores, running at the same or different clock frequencies, enable the effective distribution of workload and efficiency in energy consumption. Although Electromagnetic Side-Channel Analysis (EM-SCA) has been shown to be an effective and non-invasive method to acquire forensic insights from smartphones and Internet of Things (IoT) devices, the presence of multiple CPU cores has the potential to cause disruptions in this process. This research focuses on analysing the impact of multi-core CPU emissions — specifically the iPhone 13 and iPhone 14 Pro — on the EM-SCA-based forensic insights acquisition procedure. To achieve this, we developed a novel multi-core EM-SCA model specifically for iPhone models by integrating electromagnetic (EM) radiation traces captured from different core clusters of a single device. The developed multi-core model is then subjected to three transfer learning processes: inductive learning, feature extraction, and fine-tuning. The model is tested using individual single-core datasets collected at specific system-clock frequencies of the device. The findings of both smartphones indicate that inductive transfer learning consistently yields poor results, ranging between 5% and 20%, regardless of the core cluster. Although feature extraction provides moderate accuracy for certain datasets — around 50% to 70% for the iPhone 13 and 20% to 92% for the iPhone 14 Pro — it is the fine-tuning process that proves to be the most effective. Fine-tuning supports a wide range of datasets across different system-clock frequencies, achieving classification accuracy as high as 99%. This highlights fine-tuning as the most reliable transfer learning technique for multi-core forensic investigations. We also tested for catastrophic forgetting to evaluate the robustness of the multi-core model when using single-core datasets from the same devices. The results demonstrate that the accuracy of the multi-core model remains unchanged, even after the transfer learning process across various datasets.
Examinations are fundamental to education, yet conducting secure computer-based exams in disrupted environments presents significant challenges. This research introduces a Secure by Design protocol leveraging Delay Tolerant Networks (DTN) to overcome connectivity gaps in remote and resource-constrained areas. The proposed solution integrates physical, administrative, and technical controls to ensure the confidentiality, integrity, and availability of examination data. Through an iterative action research approach, the system evolved from a centralized Moodle server to standalone local servers, enabling offline functionality and enhanced resilience. Tested across over 180,000 candidates in Sri Lanka’s largest computer-based examination, the framework effectively addressed power outages, internet disruptions, and logistical constraints. The findings demonstrate the protocol’s effectiveness in promoting equitable and reliable access to education, ensuring examination continuity despite adverse conditions.
Accurate assessment of neuromuscular reflexes, such as the H-reflex, plays a critical role in sports science, rehabilitation, and clinical neurology. Traditional analysis of H-reflex EMG waveforms is subject to variability and interpretation bias among clinicians and researchers, limiting reliability and standardization. To address these challenges, we propose a Fine-Tuned Vision-Language Model (VLM) Consortium and a reasoning Large-Language Model (LLM)-enabled Decision Support System for automated H-reflex waveform interpretation and diagnosis. Our approach leverages multiple VLMs, each fine-tuned on curated datasets of H-reflex EMG waveform images annotated with clinical observations, recovery timelines, and athlete metadata. These models are capable of extracting key electrophysiological features and predicting neuromuscular states, including fatigue, injury, and recovery, directly from EMG images and contextual metadata. Diagnostic outputs from the VLM consortium are aggregated using a consensus-based method and refined by a specialized reasoning LLM, which ensures robust, transparent, and explainable decision support for clinicians and sports scientists. The end-to-end platform orchestrates seamless communication between the VLM ensemble and the reasoning LLM, integrating prompt engineering strategies and automated reasoning workflows using LLM Agents. Experimental results demonstrate that this hybrid system delivers highly accurate, consistent, and interpretable H-reflex assessments, significantly advancing the automation and standardization of neuromuscular diagnostics. To our knowledge, this work represents the first integration of a fine-tuned VLM consortium with a reasoning LLM for image-based H-reflex analysis, laying the foundation for next-generation AI-assisted neuromuscular assessment and athlete monitoring platforms.
Recent years have seen many industrial implementations and much scholastic research, i.e., prototypes and theoretical frameworks, in Decentralized Identity Management Systems (DIDMS). It is safe to say that Attestation-Based Attribute-Based Decentralized IDM (ABABDIDM) has not received anywhere near the same level of attention in the literature as general Attribute-Based DIDMs (ABDIDM), i.e, decentralized Attribute-Based Access Control (ABAC). The use of decentralization, i.e., DIDM, is to improve upon the security and privacy-related issues of centralized Identity Management Systems (IDM) and Attribute-Based IDMs (ABIDM). And blockchain is the framework used for decentralization in all these schemes. Many DIDMs - even ABDIDMs - have been defined on popular blockchains such as Hyperledger, Ethereum, and Bitcoin. However, despite the characteristics of Ripple that makes it appealing for an ABIDM, there is a lack of research to develop an Identity Management System (IDMS) on Ripple in literature. We have attempted to conceptualize an ABABDIDM on Ripple.
Multipath Transmission Control Protocol (MPTCP) creates multiple subflows using the available network interfaces of the device to provide a single network connection between the two hosts. Each subflow is a TCP connection and MPTCP uses TCP options field to communicate the control signals related to the MPTCP operations. MPTCP exposes single TCP connection to the application and multiplexes traffic to the subflows based on their path characteristics. Therefore, the subflow characteristics are interrelated. We hypothesise that by observing one subflow, the network related information of the unseen subflow can be inferred. A set of experiments were conducted to test this hypothesis on MPTCP connections with two subflows using the Mininet network emulator. We demonstrated that there is a strong correlation between the data flow characteristics of the two subflows. Therefore, it is possible to elicit flow information of a subflow by observing the other subflow.
Agentic AI marks a major shift in how autonomous systems reason, plan, and execute multi-step tasks. Unlike traditional single model prompting, agentic workflows integrate multiple specialized agents with different Large Language Models(LLMs), tool-augmented capabilities, orchestration logic, and external system interactions to form dynamic pipelines capable of autonomous decision-making and action. As adoption accelerates across industry and research, organizations face a central challenge: how to design, engineer, and operate production-grade agentic AI workflows that are reliable, observable, maintainable, and aligned with safety and governance requirements. This paper provides a practical, end-to-end guide for designing, developing, and deploying production-quality agentic AI systems. We introduce a structured engineering lifecycle encompassing workflow decomposition, multi-agent design patterns, Model Context Protocol(MCP), and tool integration, deterministic orchestration, Responsible-AI considerations, and environment-aware deployment strategies. We then present nine core best practices for engineering production-grade agentic AI workflows, including tool-first design over MCP, pure-function invocation, single-tool and single-responsibility agents, externalized prompt management, Responsible-AI-aligned model-consortium design, clean separation between workflow logic and MCP servers, containerized deployment for scalable operations, and adherence to the Keep it Simple, Stupid (KISS) principle to maintain simplicity and robustness. To demonstrate these principles in practice, we present a comprehensive case study: a multimodal news-analysis and media-generation workflow. By combining architectural guidance, operational patterns, and practical implementation insights, this paper offers a foundational reference to build robust, extensible, and production-ready agentic AI workflows.
The emergence of Agentic AI is fundamentally transforming how software is designed, developed, and maintained. Traditional software development methodologies such as Agile, Kanban, ShapeUp, etc, were originally designed for human-centric teams and are increasingly inadequate in environments where autonomous AI agents contribute to planning, coding, testing, and continuous learning. To address this methodological gap, we present "Agentsway" a novel software development framework designed for ecosystems where AI agents operate as first-class collaborators. Agentsway introduces a structured lifecycle centered on human orchestration, and privacy-preserving collaboration among specialized AI agents. The framework defines distinct roles for planning, prompting, coding, testing, and fine-tuning agents, each contributing to iterative improvement and adaptive learning throughout the development process. By integrating fine-tuned LLMs that leverage outputs and feedback from different agents throughout the development cycle as part of a retrospective learning process, Agentsway enhances domain-specific reasoning, and explainable decision-making across the entire software development lifecycle. Responsible AI principles are further embedded across the agents through the coordinated use of multiple fine-tuned LLMs and advanced reasoning models, ensuring balanced, transparent, and accountable decision-making. This work advances software engineering by formalizing agent-centric collaboration, integrating privacy-by-design principles, and defining measurable metrics for productivity and trust. Agentsway represents a foundational step toward the next generation of AI-native, self-improving software development methodologies. To the best of our knowledge, this is the first research effort to introduce a dedicated methodology explicitly designed for AI agent-based software engineering teams.
Current wind energy data platforms face significant challenges in securing and managing extensive data from both offshore and onshore wind farms. These challenges include vulnerabilities to cyber-attacks, data tampering, breaches, complex data-sharing issues due to privacy concerns and regulatory compliance, and a lack of scalability and flexibility in analytical tools for real-time data processing. This paper proposes a novel multilayered data security architecture, termed "VindSec-Llama," to address these challenges. It integrates Generative AI, blockchain, federated learning, and Pipeline Bill of Materials (PBOM) to enhance data analytics, model development, and security across several layers, including Infrastructure, Data Lake, Federated Learning, MLOps, Data Provenance, and LLM. Each layer is designed to meet specific functional requirements, such as handling large datasets, facilitating secure federated learning, automating risk management, and ensuring data provenance and traceability. The platform, deployable in server environments (cloud or on-premises), complies with the Risk Management Framework (RMF) guidelines and security standards. It features a blockchain-enabled, coordinator-less federated learning system to enhance data privacy and security by enabling the development of privacy-preserving machine learning models with data from different wind farms. Automation plays a pivotal role throughout VindSec-Llama, with Meta’s custom-trained Llama-3 LLM used for generating remediation scripts in the Infrastructure Layer and for producing PPBOM in the MLOps Layer. The Llama-3 LLM has been quantized and fine-tuned using Qlora to ensure optimal performance on consumer-grade hardware. The MLOps pipeline setup, a critical functionality of VindSec-Llama, ensures seamless integration and deployment of machine learning models, embodying best practices in continuous integration and delivery. This setup is geared towards maximizing security, compliance, and operational efficiency. A prototype of the platform has been implemented within a wind-energy testbed with the collaboration of Department of Energy US, illustrating its practical applications and benefits.
Agentic AI represents a major shift in how autonomous systems reason, plan, and execute multi-step tasks through the coordination of Large Language Models (LLMs), Vision Language Models (VLMs), tools, and external services. While these systems enable powerful new capabilities, increasing autonomy introduces critical challenges related to explainability, accountability, robustness, and governance, especially when agent outputs influence downstream actions or decisions. Existing agentic AI implementations often emphasize functionality and scalability, yet provide limited mechanisms for understanding decision rationale or enforcing responsibility across agent interactions. This paper presents a Responsible(RAI) and Explainable(XAI) AI Agent Architecture for production-grade agentic workflows based on multi-model consensus and reasoning-layer governance. In the proposed design, a consortium of heterogeneous LLM and VLM agents independently generates candidate outputs from a shared input context, explicitly exposing uncertainty, disagreement, and alternative interpretations. A dedicated reasoning agent then performs structured consolidation across these outputs, enforcing safety and policy constraints, mitigating hallucinations and bias, and producing auditable, evidence-backed decisions. Explainability is achieved through explicit cross-model comparison and preserved intermediate outputs, while responsibility is enforced through centralized reasoning-layer control and agent-level constraints. We evaluate the architecture across multiple real-world agentic AI workflows, demonstrating that consensus-driven reasoning improves robustness, transparency, and operational trust across diverse application domains. This work provides practical guidance for designing agentic AI systems that are autonomous and scalable, yet responsible and explainable by construction.