6G networks aim to further improve QoS while optimizing environmental sustainability metrics, and integrate increased levels of intelligence for automated control and performance optimization. However, directly applying AI/ML-driven decisions can pose risks if the models are not adequately trained. Additionally, balancing performance with energy efficiency is essential due to the high energy demands of complex AI/ML models. Finally, the use of AI/ML introduces additional vulnerabilities that must be mitigated in addition to the need for explainable decisions to ensure trustworthiness. This article aims to address these aspects by jointly integrating intelligence, energy efficiency, and trustworthiness into M&O for 6G. The proposed framework combines AI/ML-based solutions for dynamic network and service orchestration that optimize performance and energy metrics, and Network Digital Twin management to support AI/ML processes. Additionally, the framework introduces distributed and sustainable MLOps to monitor and further minimize the energy consumption of AI/ ML processes, as well as security and explainability mechanisms to enhance the system's trustworthiness. Evaluations of the framework's components demonstrate its feasibility and performance.
The rapid development and scaling of mobile telecommunications networks, together with related domains such as the edge-cloud continuum have raised significant concerns regarding energy consumption and environmental sustainability. Addressing these concerns requires a focus on CPU energy consumption, as CPUs are among the largest energy consumers in these systems. This paper investigates existing techniques, with a focus on CPU idle states (C-states), performance states (P-states), and frequency scaling governors implemented at both hardware and software levels. These mechanisms enable the dynamic adjustment of CPU parameters, providing opportunities to optimize power consumption, frequency, voltage, and overall system performance. In this regard, three CPUs with different architectures from well-known manufacturers, Intel® and AMD®, are thoroughly examined. A comprehensive dataset, collected under three load scenarios (idle, medium, and high), is used to support the analysis, reflect realistic runtime conditions, and enable a comparison of the technological differences in how these parameters are exposed and utilized.
Unmanned Aerial Vehicles (UAVs), particularly multi-UAV systems, are increasingly deployed to provide communication-enabled services to ground users while operating in shared airspace, where safety and performance must be jointly addressed under environmental uncertainty and limited information exchange. Motivated by this challenge, we model each agent as a UAV together with its associated ground users, allowing each agent to potentially have distinct performance objectives, and develop a safe decentralized planning framework that leverages refreshed sensing information to update uncertainty descriptions and support per-agent decision-making over a finite horizon. Safety is enforced through collision avoidance constraints, and we analyze conditions and planning schemes under which feasibility can be maintained over time under uncertain neighbor motion and time-varying local conditions, such as channel states and energy availability. Simulations in a multi-agent Non-Orthogonal Multiple Access (NOMA)-enabled UAV relaying setting with a shrinking-horizon Model Predictive Control (MPC) implementation demonstrate the advantages of the proposed approach over baselines, enabled by its adaptation to evolving conditions using updated uncertainty information.
A paradigm shift is underway in 6G from the design of reactive systems to autonomous and adaptive solutions that take advantage of agentic Artificial Intelligence (AI). Agentic AI solutions emerge to inject distributed intelligence into 6G management and orchestration, where multiple agents can plan steps and collaborate to achieve common goals. In the current article, we detail an Agentic AI approach for managing the intent for distributed applications deployed over 6G infrastructure. Distributed applications are provided on computing resources that span from the radio access to the edge to the core part of a 6G-enabled computing continuum. The proposed approach leverages Knowledge Graph and Generative AI technologies, enabling agents to reason over explicit relationships and dynamically generate explainable and reliable plans. A prototype implementation is detailed, accompanied by evaluation results that highlight the efficiency of the proposed approach to improve automation and efficiency on the intent lifecycle management.
Open RAN (O-RAN) enables intelligent and programmable radio access network control, but continuously exposing fine-grained per-UE telemetry to near-real-time (near-RT) controllers is costly and may also raise privacy concerns. Network Tomography (NT) can offer an indirect monitoring alternative by inferring hidden UE-level states from coarse aggregate measurements and UE-cell association information. In this work, we bring NT into O-RAN by estimating fine-grained downlink throughput for UE-cell pairs from aggregate per-cell throughput observations. We formulate this task as an ill-posed linear inverse problem and develop a modular pipeline consisting of an Activity Detector based on per-cell decision trees, a topology-aware, LISTA-inspired unfolded Throughput Estimator, and a Post-Processing step. Evaluation in a simulated environment with UE mobility and multi-cell connectivity shows that accurate inference is feasible and identifies the total number of UEs and their mobility as the main factors affecting performance.
The application of Social and Emotional Learning (SEL) methodologies within classrooms has been widely adopted and examined, highlighting its positive impact on the academic, emotional, and social sectors at both the individual and classroom levels. Through this wide applicability, a need is identified to focus more on the development of SEL approaches that support the integration of the various parts of a SEL methodology (e.g., planning, implementation, monitoring and assessment). Such approaches should not underestimate the value of the SEL assessment. It is necessary to engage both the developers and the practitioners of SEL methodologies and provide tools to assist teachers to better interpret the collected data and results and plan targeted SEL activities. Motivated by such needs, we detail the EduCardia SEL methodology, which focuses on the provision of a set of guidelines and tools to assess and improve the social and emotional competencies of students in primary and secondary educational levels, taking advantage of Information and Communication Technologies (ICT). The EduCardia SEL methodology considers the integration of the various parts of a SEL methodological approach and supports continuous assessment processes while offering user-friendly interfaces to interpret the produced results. Evaluation results are provided based on the application of the EduCardia SEL methodology in K12 schools in Greece for a short period of time, considering three different age groups. The targeted competencies are shown to be improved, while teachers provide insights on the applicability and efficiency of the proposed approach.
The development of modern distributed applications has been widely based on microservices-based paradigms, where distributed application graphs with strict Quality of Service requirements are deployed over programmable infrastructure. In such deployments, a set of objectives and constraints for the efficient provision of the distributed application has to be considered. A high-level description of these objectives and constraints is mentioned with the term intent. To satisfy the intent of a distributed application, intent-driven orchestration approaches emerge where high-level goals have to be translated to deployment and operational policies. In the current manuscript, we detail an approach that supports the intent lifecycle management from the intent generation towards its validation, translation into deployment plans, management through policies enforcement, and reporting regarding its fulfillment. Evaluation results are provided based on the development of a simulation environment and the examination of the performance impact of the activation of the various phases of the intent lifecycle.
Artificial intelligence (AI) has captivated research in network resource orchestration, driving advancements and innovations, including the much-anticipated integration of terrestrial networks (TNs) and non-terrestrial networks (NTNs). Although AI has already been incorporated into standards to complement or replace network functions, its reliability in maintaining the required quality of service (QoS) in highly dynamic and complex environments is often overlooked. This article highlights the significant contributions of AI to unified resource management on the ground and in various orbits while addressing its reliability challenges in complex integrated TN/NTN systems. In this context, the different phases of the AI lifecycle are analyzed and their key processes and steps are identified to introduce mechanisms that ensure reliability at each stage. A principled reliable AI framework is ultimately designed to meet the stringent reliability requirements of highly volatile environments like the integrated TNs/NTNs. To shed light on the practical application of the framework, a comprehensive example is provided for the problem of computation task offloading in integrated TNs/NTNs. This proof of concept emphasizes the design of fully decentralized and scalable solutions while examining the significance and impact of mechanisms that respond to unseen states and maintain the required QoS level safely. Overall, this article provides a guide for tackling the challenges of ensuring reliability in AI-enabled TN/NTN functions that can serve as a powerful catalyst for ongoing standardization efforts.
Intent-driven orchestration approaches emerged recently to abstract the complexity of the management of distributed applications and services from the end users. Such approaches are innovative, however they lack in terms of abstraction and applicability in different scenarios and in cases where multiple stakeholders are involved in decision-making. In this paper, an approach for intent lifecycle management is detailed, considering the deployment of applications over resources in the distributed computing continuum. An emerging ecosystem is examined where infrastructure from multiple providers may have to be used for efficient provision of distributed applications. The approach includes the conceptualization of a Knowledge Graph to assist semantic validation functionalities and decision-making in the various developed control loops, as well as the creation of knowledge from the information collected from the intent assessment processes. Based on the development of a simulation kit, evaluation results are provided for scenarios that focus on performance and energy efficiency.
The increasing complexity and scale of cloud-native and edge-enabled systems have placed observability pipelines under pressure, especially in 6G-ready infrastructures where bandwidth, energy, and responsiveness are critical. Traditional telemetry components are often static and unaware of system-level goals, limiting their ability to support data-driven network and service management. In this paper, we propose a hierarchical multi-agent reinforcement learning approach for adaptive telemetry data collection within an Observability-as-a-Service framework. The proposed mechanism introduces a two-level control plane architecture where local agents autonomously adjust data collection parameters, while a global agent ensures that their actions align with end-to-end objectives. We develop a prototype implementation and conduct a comprehensive evaluation against several baselines. The results demonstrate significant improvements in Observability Service Level Agreement satisfaction, fairness, and adaptability under dynamic system conditions.
This article aims to explore the requirements and challenges that Intent-Based Management (IBM) systems should look towards to deliver proper automation management in multi-stakeholder 6G scenarios. To do so, the evolution of the telecommunications actors is presented to identify how they will interact with the use of intents and what IBM systems must face to fulfil with the tenants expectations while interacting among other IBM domains (aggregation vs. federation models). In a second step, this article presents how the IBM systems functionalities should be organized (based on standards research and comparative) and a set of enablers using them to properly deliver the automation needed to accomplish the requests done by the tenants. On a third step, a set of End-to-End (E2E) intent-based services are described to illustrate how the different enablers (and the functionalities they use) may work to reach the E2E service goal. Finally, a set of experimental results related to how intent conflicts should be manage is presented, followed by the conclusions.
The emergence of 6G mobile communications promises transformative advancements powered by Artificial Intelligence (AI) and Machine Learning (ML). However, these capabilities introduce significant challenges, mainly due to the heterogeneous nature of 6G-enabled environments. This paper presents an AI-driven orchestration framework designed to address the complexities of future 6G infrastructures. The proposed framework orchestrates microservice-based applications across diverse layers of the computing continuum and performs key tasks such as service placement, computation offloading, horizontal scaling, and live migration, leveraging AI to enhance performance, energy efficiency, and automation. Additionally, it incorporates Distributed Ledger Technology (DLT) to enable federated service management when centralized control is impractical. We validate this approach through experimental trials with a latency-sensitive 6G application, analyzing intra- and inter-domain orchestration, proactive extreme edge migration, and inter-cluster communication technologies. A set of representative workflows are provided as open-source implementations.
Modern heterogeneous networks and distributed applications deployed across the computing continuum pose significant monitoring and management challenges due to their complexity and dynamic nature. Cloud-native observability has emerged as a powerful paradigm to overcome these challenges, enabling the extraction of knowledge and operational insights by combining information from metrics, logs, and traces. At the same time, advanced Artificial Intelligence (AI) techniques drive network automation toward Zero-Touch Network and Service Management (ZSM). In this work, we present FUSION, an integrated observability and analytics framework that unifies heterogeneous telemetry data and leverages AI-powered intelligence to extract knowledge, identify behavioral patterns, and increase automation. A conceptual architecture is detailed, standardizing heterogeneous data collection and fusion in a unified observability layer, while supporting the integration of advanced AI-based analytics pipelines. A prototype implementation demonstrates the use of representative AI-based methods for anomaly detection, root cause analysis, and intelligent orchestration.
Transparency, collaboration, security, and digital sovereignty are the open source values that are also important to many European countries and institutions. Open source software offers practical advantages, cost savings, and opportunities for innovation that contribute to the region's technological advancement and competitiveness, as well as fosters collaboration among developers and encourages innovation. On top of these values, several European projects try to build a sustainable and innovative future. In this paper the authors present two examples of an open source applications of meta-operating system (metaOS). The first NExt generation Meta Operating systems (NEMO) builds the future of the Artificial Intelligence of Things(AIoT)-edge-cloud continuum by introducing an open source, modular and cybersecure metaOS. The second regards the open source solutions provided by the NEPHELE project to assist orchestration of distributed applications across resources in the computing continuum. In this matter, the open source components of the metaOS projects are presented for each functional layer in the architecture, demonstrating the value of open source for the metaOS sustainability in a specific use case.
Network Slicing (NS) and network function disaggregation are two enablers that jointly lead to the evolution of communications to the latest and future generations. In this work, we focus on NS for the disaggregated RAN architecture called ORAN and tackle jointly the problems of slice admission control and placement over compute and network resources on edge or regional clouds. We employ a Reinforcement Learning (RL) approach to ensure adaptability to evolving network conditions and slice request patterns in an efficient way without the need of traffic forecasts. Contrary to similar approaches, we optimize the split of the slice Service Function Chain (SFC), while considering a general slice type. We have evaluated the generalization properties of the trained agents in testing data that deviate from those in the training ones. In addition, we have performed comparative evaluations with RL-based approaches that perform static split to show the significant improvements brought by our proposed optimal function split.
The emergence of new technologies of 5G/6G networks and the Internet of Things (IoT) drives the transition from traditional Cloud Computing systems to the Edge Cloud Continuum — an interconnected distributed computing environment. Deploying modern applications in such a complex setting poses significant challenges for efficient dynamic resource management. Besides their several benefits, current orchestration platforms disregard aspects such as the dynamic behavior of applications’ demands, heterogeneity of the infrastructure’s resources, and the overall complexity when dealing with interdependent resource allocation decision parameters. The Digital Twin concept envisions assisting application deployments not only by providing offline simulations for experimental assessment in multi-cluster settings but also by actively guiding the orchestration process. In this paper, we aim to provide a modeling and simulation framework to optimize the performance of the underlying infrastructure in terms of resilience and sustainability. We investigate the application of automata theory to model such systems by analyzing their possible states, specifically, using Petri Nets, a mathematical framework for representing discrete event systems, as the primary modeling tool. Therefore, a comprehensive modeling approach is presented to simulate the resource scheduling decisions of an established multi-cluster framework, namely Karmada. Moreover, through this Petri Net modeling approach, we can efficiently optimize the performance of the orchestration process considering the power consumption and workload load balancing of a multi-cluster topology. Extensive evaluation indicates the efficacy of the proposed framework in accurately approximating Karmada’s behavior for various scheduling policies. Also, the proposed framework is capable of assessing the performance of several scheduling policies and guiding the system towards efficient resource management in complex scenarios, exploiting the polynomial complexity of the Petri Net to identify scheduling states.
The advent of sixth-generation networks (6G) introduces transformative capabilities in terms of ultra-reliable low-latency communications, massive connectivity, and edge cloud integration, that promise to enable time-sensitive applications and immersive services. These ambitious capabilities, however, pose significant challenges for resilient and dynamic resource allocation, requiring adaptation of orchestration frameworks towards efficient and sustainable resources management. In this paper, we propose a Model Predictive Control (MPC)-based framework for 6G network slicing, focusing on minimizing power consumption, end-to-end delay, and slice reconfiguration costs. The framework is designed to operate as an optimization mechanism integrated within a 6G network slice management architecture, aligned with emerging virtualization, monitoring and analytics functionalities. The proposed MPC-based framework leverages future demand forecasts to optimize resource allocation over a receding time horizon. Additionally, we introduce a proxy-based adaptation of the MPC to enhance scalability and decision speed, ensuring efficient operations even in large-scale 6G settings. Through simulations, we demonstrate the efficacy of the proposed solution compared to other approaches.
The rapid growth of Internet of Things (IoT) devices and emerging technologies, along with the growing demands of edge-deployed applications, has led to a complex paradigm where computation often shifts dynamically acrooss the IoT-Edge-Cloud continuum. The NEPHELE project addresses these complexities by enabling seamless orchestration across a diverse spectrum of computing resources, spanning multi-cloud environments to the far-Edge. In this paper, we present NEPHELE's multi-cloud infrastructure, built to overcome key orchestration challenges within cloud and edge environments. We discuss the core components and architectural decisions, focusing on multi-cluster resource orchestration mechanisms, integrated monitoring for local and multi-cloud systems, inter- and intra-cluster scaling, and networking capabilities. Experimental results demonstrate the efficiency of our infrastructure, highlighting overhead management in service deployment, migration, networking, and scaling scenarios, thus highlighting NEPHELE's robustness in handling distributed applications across heterogeneous environments.
Nikolaos Konstantinou合作论文数National Technical University of Athens9