In this paper, we propose a novel method for generating realistic multi-step counterfactual explanations for machine learning-controlled robots using 2D Light Detection and Ranging (LiDAR) data. Machine learning algorithms based on black-box models such as neural networks suffer from a lack of interpretability. This lack of interpretability is critical when such models are used to control robots, which are often deployed in safety-critical environments. Robotic policies also act in a closed loop with the environment and rely on high-dimensional sensor input, which makes their behavior difficult to explain from a single observation. Our method extends our previous single-step approach by generating explanations based on forward simulations of the robot’s dynamics. The proposed approach explains the behavior of agents over a time horizon, making it suitable for reinforcement learning (RL)-based policies. To support these explanations, we introduce new cost functions that enable users to pose trajectory-level questions such as ‘‘why does the robot not visit this area instead?’’. We evaluate our method on a TurtleBot3 platform controlled by an RL policy trained to navigate in unknown environments. We compare the proposed multi-step method with the single-step approach and show that multi-step counterfactuals provide more informative explanations. In particular, the method shows how alternative obstacle configurations could affect the robot’s trajectory over several time steps, revealing behavior that is missed by single-step explanations.
Autonomous marine vessels require robust control systems capable of handling model uncertainties and environmental disturbances during precision manoeuvrers. This work presents three main contributions to address these challenges: First, a novel overactuated Reinforcement Learning-based Nonlinear Model Predictive Control (RL-NMPC) on the 4-azimuth thruster milliAmpere 1 research ferry with control allocation directly integrated in the optimal control problem. A combination of Q-learning and model-learning is employed to simultaneously improve the control goal of the NMPC and adapt the model online. The integrated allocation in the optimal control problem avoids allocator-controller conflict and improves robustness under wind disturbances and parameter drift. Second, a hybrid reference filter that solves coordinated surge, sway and yaw trajectory profiling without weight distortion between Degrees-of-Freedom (DoF). The filter combines trapezoidal velocity generation with body-frame velocity saturation and third-order filtering to provide smooth, predictable trajectory references for multi-DoF manoeuvres. Third, a comparative study of RL-NMPC, NMPC, and dynamic positioning control systems under 40% modelling errors and 4.0 m/s wind disturbances on the established 8-Corner Test trajectory. Simulation results demonstrate that RL-NMPC achieves superior tracking performance, reduced energy consumption, and lower actuator wear compared to both NMPC and DP controllers, while maintaining robust control under environmental disturbances.
This paper presents a fully integrated autonomous docking system validated through closed-loop sea trials on the milliAmpere1 research ferry operating in a live maritime harbour with moving vessels. Real harbour environments require continuous situational awareness and adaptive decision-making under dynamic traffic conditions. The proposed architecture combines cartographic land masking, LiDAR-based clustering, probabilistic multi-target tracking (JIPDA), dynamic footprint estimation, adaptive docking pose selection, and real-time path replanning within a finite state machine framework. Rather than introducing new algorithms, the contribution lies in system-level integration and operational validation of a complete perception-to-control pipeline under realistic maritime constraints. The system is demonstrated in multiple closed-loop experiments including collision avoidance and adaptive docking with moving obstacles. Results highlight both performance characteristics and practical deployment considerations, including runtime behaviour, sensor limitations, and integration trade-offs. The work provides empirical evidence that robust autonomous docking in dynamic harbour environments can be achieved through carefully engineered integration of established methods.
This work proposes a COLREGs-compliant collision-avoidance (COLAV) framework for autonomous docking transits based on Control Barrier Functions (CBFs). Starting with the definition of a safety set using the target-ship domain, we introduce a higher-order CBF (HOCBF) to strictly enforce domain exclusion, together with a scenario-dependent weighting of the CBF constraint to handle thrust saturation while preserving feasibility. The CBF acts as a reactive safety filter on top of conventional guidance and control, producing minimally modified inputs. We validate the approach in the high-fidelity milliAmpere1 simulator across COLREGs Rules 13–16 and in field trials (overtaking, head-on, give-way). Our results demonstrate reliable, rule-compliant avoidance and robust handling of time-critical obstacles, while also revealing tuning interactions with the heading/docking controller.
Large Language Models (LLMs) are increasingly used for decision-making and reasoning tasks, yet their potential as controllers for physical systems remains largely unexplored. This work investigates whether LLMs can function as interpretable controllers for a dynamic thermal environment, examining their ability to follow setpoints, interpret natural-language commands, reason about actuator effects, and incorporate prior model-based knowledge. Five LLMs of varying scales are evaluated under multiple scenarios, including settings with penalties on heater or fan usage and cases where the models have access to a physics-based prediction tool. The results show that control performance depends on model complexity: while low- and mid-scale models frequently misinterpret actuator dynamics or generate inconsistent reasoning, high-complexity models such as Qwen-3~14B and GPT-4o achieve accurate temperature tracking, stable actuator usage, and coherent explanations aligned with physical principles. Incorporating a physics-based model significantly improves control smoothness and energy efficiency by enabling anticipatory decision-making. A detailed reasoning taxonomy further reveals a clear progression from causal misinterpretation in smaller models to cohesive and temporally aware reasoning in larger ones. The findings demonstrate that LLMs can act as interpretable controllers when sufficiently capable and appropriately grounded in domain knowledge, highlighting promising opportunities for hybrid model-based and language-driven control strategies that can provide plausible explanations.
Deep reinforcement learning (DRL) has the potential to provide an operationally efficient and rapidly deployable alternative to classical ship dynamic-positioning (DP) controllers; however, the limited transparency of neural network-based policies currently hinders their adoption on regulated vessels. This paper investigates whether a Proximal Policy Optimization (PPO) controller, augmented with live Shapley Additive Explanations (SHAP), can be trained, evaluated, and deployed on a physical autonomous vessel to perform the DP task while keeping human operators “in the loop”. To achieve this, we developed a ROS-based Gymnasium environment around a 3-DOF ship simulator with a four-thruster action space, which decomposes azimuth forces into surge and sway components to preserve SHAP additivity. The actor-critic network, consisting of two 64-unit tanh layers, was trained for 500,000 time steps (approximately 35 h) using a shaped reward that balances pose error, velocity, and thruster wear. Default Stable-Baselines3 hyperparameters were retained to ensure robustness. Sea trials in Trondheim, Norway confirmed that the proposed controller: (i) held station inside the mandated error of 0.25 m and 5 degrees, with an action space [-10,10] m (ii) followed straight-line and spline paths with metre-level accuracy, due to not integrator and constantly updating target, and (iii) matched simulated behavior except for mild oscillations attributed to velocity calculations and disturbances. During experimental runs, a four-window dashboard streamed SHAP bar plots and force arrows at 4 Hz. Operators consistently observed that the dominant feature, such as pose error or sway velocity, correctly drove the corresponding thruster pair, providing validation of the explanation pipeline. To the author’s knowledge, this is the first end-to-end demonstration of DRL combined with live explainable AI (XAI) on a real marine craft.
Autonomous electric ferries offer an eco-friendly solution for sustainable and intelligent urban mobility. Currently, these ferries are undergoing successful trial operations on short routes for passenger transportation in European and Scandinavian cities. This paper presents experimental results on model identification, thrust allocation (TA), and dynamic positioning (DP) system for the milliAmpere1 (mA1) autonomous electric ferry prototype. The mA1 is a small prototype ferry used as a test platform in several research projects at the Norwegian University of Science and Technology (NTNU). It recently went through a system upgrade and is now equipped with four symmetrically distributed azimuth thrusters. These new systems require a new mathematical model and a reliable DP and TA system to enable autonomous operation. The DP system is responsible for calculating the total force and moment required to keep position or move between waypoints. At the same time, the TA algorithm coordinates the thrusters to ensure that the resulting force they generate matches the request from the DP control algorithm. The DP is designed with a horizontal plane control law, using a feedforward plus feedback control. The TA is formulated as a quadratic programming problem, which includes cost functions for power consumption. A third-order reference guidance system is implemented to ensure smooth and continuous signals for the DP desired position, velocity, and acceleration. Additionally, system model identification was performed to determine the parameters of the actuator dynamics. Full-scale trials were conducted to evaluate the performance of the proposed system in a sheltered basin situated in Trondheim, Norway, under low-wind conditions. The field trial results indicate that the proposed approach successfully stabilized and maintained a desired trajectory.
Autonomous ferries for urban mobility have gained significant attention in recent years, with notable advances in short route passenger transportation in Scandinavian countries. This paper presents the hardware and software onboard the milliAmpere1 (mA1) autonomous ferry, a small ferry [1] utilized as a test platform for several research projects at the Norwegian University of Science and Technology (NTNU), Norway. The mA1 is equipped with situational awareness, docking, guidance, navigation, and dynamic positioning systems, enabling it to perform autonomous navigation. The system architecture onboard the mA1 is divided into several subsystems: (1) propulsion system, composed of 4 azimuthal electric thrusters connected through a CAN bus network, (2) navigation system, consisting of a GNSS station receiver with RTK correction and an IMU, (3) power supply system, composed of 8 absorbent glass mat batteries that supply power to the whole system, including 24V DC and 220V AC, (4) onboard control and monitoring system, which includes a fanless industrial computer running Linux and ROS-based software, and (5) situational awareness system, which features a LiDAR, maritime radar, 5 EO cameras, and 5 IR cameras. Additionally, the mA1 is equipped with a radio joystick controller and emergency stop buttons to ensure the safe operation. After an upgrade of the thrusters and electrical systems in 2024, preliminary full-scale experiments were conducted to assess the performance of the hardware and software, and the correct functionality and operation of all onboard systems were confirmed while operating under moderate weather conditions.
This paper presents a novel method for generating realistic counterfactual explanations (CFEs) in machine learning (ML)-based control for mobile robots using 2D LiDAR. ML models, especially artificial neural networks (ANNs), can provide advanced decision-making and control capabilities by learning from data. However, they often function as black boxes, making it challenging to interpret them. This is especially a problem in safety-critical control applications. To generate realistic CFEs, we parameterize the LiDAR space with simple shapes such as circles and rectangles, whose parameters are chosen by a genetic algorithm, and the configurations are transformed into LiDAR data by raycasting. Our model-agnostic approach generates CFEs in the form of synthetic LiDAR data that resembles a base LiDAR state but is modified to produce a pre-defined ML model control output based on a query from the user. We demonstrate our method on a mobile robot, the TurtleBot3, controlled using deep reinforcement learning (DRL) in real-world and simulated scenarios. Our method generates logical and realistic CFEs, which helps to interpret the DRL agent's decision making. This paper contributes towards advancing explainable AI in mobile robotics, and our method could be a tool for understanding, debugging, and improving ML-based autonomous control.
This paper presents a novel approach to solving the docking problem for Autonomous Surface Vessels (ASVs) while considering the Convention on the International Regulations for Preventing Collisions at Sea (COLREGs). The proposed method combines computational geometry, numerical optimal control, and a collision avoidance (COLAV) system to generate a collision-free path from an initial location to the docking point. The proposed approach first performs Voronoi partitioning of the non-convex area to create waypoints, forming a roadmap that ensures safe distances from harbors and obstacles. A search algorithm then finds the waypoint sequence corresponding to the shortest collision-free path. Subsequently, a nonlinear optimal control problem (NOCP) further refines the optimal trajectory and generates control inputs corresponding to this trajectory. The efficiency of our approach is demonstrated by simulations. Our results indicate that the proposed approach is suitable for transit and docking applications, where intricate and safe maneuvering at low speeds is often a necessity.
Deploying multi-robot systems for Inspection and Maintenance (IM) operations offers several advantages, including covering a large mission area and providing redundancy against individual failure. These features are particularly relevant for IM operations in offshore oil and gas platforms due to the extensive number of sensors and equipment involved, as well as the high level of fault-tolerant resilience required. This paper presents preliminary experimental results of a multi-robot mission planning system designed for performing autonomous inspection at Equinor’s K-Lab Test Centre in Haugesund, Norway. The high-level action planner is based on the Simultaneous Task Planning (STP) algorithm. The STP planner has the capability to compute plans for multi-robot systems, as it allows for concurrent actions. The system can take into account high-level actions and their expected durations, such as "visiting a specific location", "inspecting a sensor", or "taking a picture of a specified component". Additionally, STP can replan in real-time in case of events such as the need to revisit a waypoint or low battery status. Communication with the robots is achieved through Equinor’s Integration and Supervisory Control of Autonomous Robots (ISAR) framework, ensuring a fast and secure Wi-Fi connection. The collision avoidance, guidance, and control system is handled at the low-level layer of each robot. The proposed technique was validated through field experiments involving two unmanned ground robots performing an inspection round, which includes visiting a sequence of predefined waypoints and incorporates replanning for low battery situations. Two types of experiments were conducted: one scenario where robots collaborated by overlapping capabilities; and another scenario where robots collaborated by combining capabilities.
Following the recent upgrade of the propulsion plant configuration for the milliAmpere1 passenger ferry prototype, a model-based controller pipeline suitable for all the milliAmpere1 speed ranges is proposed and evaluated. Some well-known methods for force allocation and reference model systems are implemented together with a nonlinear model-based motion controller, and a comparison study for different combinations is carried out to define the best solution for this application. Considering low-speed operations, the efficiency of three force allocation solutions and three reference models with two possible thruster configurations are investigated. The resulting controllers are evaluated with five performance metrics. For higher-speed operations, a solution with a reference model and a force allocation is presented and investigated. The controllers have been tested via numerical simulations, and based on the performance metrics, indications for the best design options are provided.
In this work, we use optimal control to change the behavior of a deep reinforcement learning policy by optimizing directly in the policy's latent space. We hypothesize that distinct behavioral patterns, termed behavioral modes, can be identified within certain regions of a deep reinforcement learning policy's latent space, meaning that specific actions or strategies are preferred within these regions. We identify these behavioral modes using latent space dimension-reduction with Pairwise Controlled Manifold Approximation Projection (PaCMAP). Using the actions generated by the optimal control procedure, we move the system from one behavioral mode to another. We subsequently utilize these actions as a filter for interpreting the neural network policy. The results show that this approach can impose desired behavioral modes in the policy, demonstrated by showing how a failed episode can be made successful and vice versa using the lunar lander reinforcement learning environment.
Mission planning constitutes an important feature of autonomy for Maritime Autonomous Surface Ships (MASS). Nevertheless, this research topic remains largely unexplored, as the majority of academic and industry projects primarily focus on developing low-level systems, such as control, collision avoidance, and situational awareness. The main contribution of this paper is to address this problem by developing a high-level decision-making system capable of generating an efficient and feasible temporal sequence of high-level actions, which is then sent to the ship control systems responsible for execution. The mission planner is based on the simultaneous temporal planner (STP), which in our case considers temporal actions related to, for example, moving to a specific location, activating docking mode or starting the process of container (un)loading, which are then executed by their respective control systems. Contrary to classical artificial intelligence (AI) planning algorithms, Temporal AI planning algorithms, such as STP, can consider duration of actions, which allows more realistic representation of the mission. We connect the high-level mission planner with the ship's guidance, navigation and control (GNC) system, which has path-planning, path-following control and fuzzy logic-based collision avoidance capabilities. The efficiency of our approach is demonstrated through a series of simulations of a MASS operating in a realistic marine environment including other ships and static obstacles.
This study addresses the challenges of maneuvering a large container ship in confined waters under the influence of wind and currents. The proposed guidance, navigation, and control system presents a novel combination of modeling methodologies, control allocation techniques, and nonlinear control systems theory. We consider a docking scenario through a simulation of a realistic, narrow harbor accommodating area-specific currents and global wind forces. Control allocation is solved by utilizing detailed models of the actuators and hydrodynamics in a mixed-integer-like solution for optimal thrust allocation, and an inverse mapping to the control inputs. Furthermore, our approach introduces a modified Wageningen B-Series model, that is capable of modeling propellers with negative pitch dynamics. The complexity of safely docking in environments like the Antwerp harbor, where wind and currents pose significant risks, is tackled using a nonlinear PID control system integrated with Line-Of-Sight (LOS) guidance. This enables precise maneuvering of a fully actuated Roll-On/Roll-Of vessel to its docking position and orientation. Given that few studies on automated docking holistically address the intricate challenges of complex harbor environments and the concurrent impact of multiple external forces, our work proposes a comprehensive solution based on an integrated combination of guidance, navigation, and control systems. Copyright (c) 2024 The Authors. This is an open access article under the CC BY-NC-ND license (https://creativecommons.org/licenses/by-nc-nd/4.0/)
Understanding the behavior of deep reinforcement learning (DRL) agents is crucial for improving their performance and reliability. However, the complexity of their policies often makes them challenging to understand. In this paper, we introduce a new approach for investigating the behavior modes of DRL policies, which involves utilizing dimensionality reduction and trajectory clustering in the latent space of neural networks. Specifically, we use Pairwise Controlled Manifold Approximation Projection (PaCMAP) for dimensionality reduction and TRACLUS for trajectory clustering to analyze the latent space of a DRL policy trained on the Mountain Car control task. Our methodology helps identify diverse behavior patterns and suboptimal choices by the policy, thus allowing for targeted improvements. We demonstrate how our approach, combined with domain knowledge, can enhance a policy's performance in specific regions of the state space.
Inspection and maintenance (IM) operations in offshore oil& gas platforms involve significant challenges due to the harsh conditions in such remote environments. Therefore, there is a need for robotic solutions that can accomplish IM missions autonomously, a goal that has become more feasible in the last years due to advances in instrumentation, perception and artificial intelligence (AI). In this paper, we present initial experimental results of the Taurob unmanned ground vehicle (UGV) performing autonomous inspection at Equinor's K-Lab Test Centre, located in Haugesund, Norway. The proposed solution employs the simultaneous task planning (STP) algorithm, introduced by Furelos-Blanco in [1], as a high-level action planner. STP can take into account durative high-level actions, such as "visit a specific location", "inspect a sensor", or "take a picture of a specified component" and, in addition, can replan online in case of events such as low battery status, or the need to revisit a location. Communication with the robot is achieved via Equinor's Integration and Supervisory Control of Autonomous Robots (ISAR) framework, which guarantees a fast and secure wi-fi connection. Two types of experiments were carried out: 1) A basic inspection round, which involves visiting a sequence of predefined waypoints; 2) A replanning scenario, which also takes into account the battery status. Our preliminary tests, where fast and efficient plans were computed and executed in real time, including replanning, showed that connecting the guidance and control system of a UGV with a high-level planner like STP is a promising approach to increased autonomy in IM missions.
The increasing use of complex and uninterpretable Artificial Intelligence (AI) models has led to a growing demand for AI model transparency. In response, the research field of Explainable AI (XAI) is growing, intending to increase trust in black box models by explaining the decisions made. In this paper, we apply the XAI methods SHAP, decision trees, and ProtoDash to Deep Reinforcement Learning (DRL) policies trained in software environments simulating road vehicle traffic cases of different complexity. The SHAP algorithms Deep SHAP and Kernel SHAP allowed us to assess the rationality of the DRL agents' policy by assigning feature attributions to their decisions. Limitations were found in manipulating the input of Deep SHAP, while Kernel SHAP demonstrated more flexibility. Decision trees provided insight into the agents' behavior by splitting and classifying policy decisions, revealing rational tendencies in their overall behavior. ProtoDash was used to select states representative of each action, which highlighted weaknesses in the trained policy. The capabilities these XAI methods demonstrated in terms of insight and transparency can be used to increase trust in black-box decision-making, facilitating for industrial adaptation of AI models.
Recent rapid advances within autonomy for unmanned aerial vehicles (UAVs) make it possible to consider deploying them for critical operations in challenging environments, such as the ocean. Going even further, the combined operation of UAVs and autonomous surface vessels (ASVs) could be a strong enabler for exploration and search and rescue missions at sea, to name a few examples. The main contribution of this paper involves designing the perception, planning and control modules for the fundamental phase of such a collaborative operation, that is, the automatic drone landing of an UAV on an ASV. The perception system utilizes an Extended Kalman Filter that fuses the estimates from two computer vision systems, one based on deep learning and one on traditional edge detection approaches. The control system is a cascaded structure of position, velocity, and attitude control. The high-level planning is based on the GraphPlan algorithm and generates action sequences to solve three missions of different complexity based on a set of defined domain variables. The efficiency of the proposed solution is demonstrated via field trials involving a Parrot Anafi quadcopter landing on a helipad installed on DNV's Revolt ASV. Our preliminary experimental results demonstrate that the system is capable of landing the UAV on the Revolt at sea, in conditions that were almost stationary, or involved forced rolling motion.