Modern AI systems achieve remarkable performance through fundamentally stochastic processes—machine learning models that function as high-dimensional probability density functions, outputting the most likely predictions given training data. While these systems can match or exceed human performance on average, their methodology produces fundamentally different failure modes than human reasoning, leading to errors that appear nonsensical from a human perspective but are predictable given their probabilistic nature. This has critical implications for high-consequence environments such as military applications where decisions cannot be reversed and may affect lives and material assets definitively. Through detailed analysis of contemporary AI’s working mechanisms—particularly how knowledge is acquired through statistical pattern recognition rather than causal reasoning—this paper demonstrates why AI systems inherit biases, cannot distinguish plausibility from factual correctness, and exhibit confident behaviour even when wrong. Written to provide guidance for non-technical stakeholders, specifically but not exclusively in the military domain, it posits that for effective deployment of AI in high-consequence scenarios, processes need to be implemented that make sure all human stakeholders are aware of these facts, develop adequate scepticism of the AI system, and remain actively involved in the decision-making. For military applications specifically, this understanding reveals that effective human-AI collaboration requires more than oversight: it demands co-learning frameworks that maintain meaningful human control through bidirectional information flow, and behavioural and functional awareness on the human side. We give an outlook to decentralized, co-learned AI system tailored to specific teams in dedicated co-learning labs to mitigate power concentration risks while preserving essential human capacities, including moral judgment to exercise mercy.
The development of autonomous aerial robots capable of safely navigating complex real-world environments without or with little human intervention represents a major milestone in robotics and artificial intelligence (AI). While rapid advances in AI-enabled decision-making, sensing, and control systems are unlocking new capabilities for unmanned aerial vehicles (UAVs), their translation into safe and scalable real-life applications remains a major challenge. In this Perspective, we examine key AI technologies relevant to aerial autonomy and discuss early application scenarios in unmanned aviation and airspace management, with a focus on their assurance-relevant properties. We analyze regulatory obstacles that limit deployment, particularly for AI-enabled and beyond visual line of sight (BVLOS) operations, and highlight why traditional risk assessment and certification approaches are need to be updated to account for adaptive, data-driven systems. Building on this analysis, we argue that testing infrastructure must be understood as a core scientific instrument, enabling systematic evidence generation under realistic and safety-critical conditions, validating autonomous functions, ensuring safety, and building trust among regulators and the public. As a concrete example, we introduce LINA, a scientifically-grounded, integrated experimentation and validation platform in Switzerland designed to support iterative, regulator-aware development of autonomous systems across technology readiness levels. We highlight how LINA function as sandbox for system-level science, regulatory learning, and trust building, thereby enabling the responsible and societally acceptable integration of autonomous aerial systems and strengthening Switzerland’s role in advancing aerial robotics research and innovation.
Graph pooling is a fundamental operation in Graph Neural Networks (GNNs), designed to simplify graphs by reducing the number of nodes and edges while preserving essential structural information for classification tasks. However, most existing pooling methods tend to overlook edge weights and rely on a single-view pooling strategy that focuses either on local or global topological information, failing to capture the full structural context of the graph. To address these limitations, this study introduces a novel Dominant Set Multi-View Pooling (DSMVPool) method featuring two main contributions. First, we propose a dominant-set cluster pooling approach that analyzes the overall graph architecture and connectivity patterns, identifies potential clusters using edge weight information, and generates a coarser graph view. In addition, we create two complementary pooled views by selecting the most representative nodes based on local topology and node features. Second, we design a fusion-view attention layer that integrates the coarser graph structure with the pooled graph views, enabling our method to simultaneously capture and combine global and local structural information and node features. Extensive experiments on four graph classification benchmarks, covering computer vision, chemical, biological, and social networks, demonstrate that DSMVPool achieves superior performance compared to state-of-the-art methods.
Many document types use intrinsic, convention-driven structures that serve to encode precise and structured information, such as the conventions governing engineering drawings. However, state-of-the-art approaches treat document recognition as a mere computer vision problem, neglecting these underlying document-type-specific structural properties, making them dependent on sub-optimal heuristic post-processing and rendering many less frequent or more complicated document types inaccessible to modern document recognition. We suggest a novel perspective that frames document recognition as a transcription task from a document to a record. This implies a natural grouping of documents based on the intrinsic structure inherent in their transcription, where related document types can be treated (and learned) similarly. We propose a method to design structure-specific inductive biases for the underlying machine-learned end-to-end document recognition systems, and a respective base transformer architecture that we successfully adapt to different structures. We demonstrate the effectiveness of the so-found inductive biases in extensive experiments with progressively complex record structures from monophonic sheet music, shape drawings, and simplified engineering drawings. By integrating an inductive bias for unrestricted graph structures, we train the first-ever successful end-to-end model to transcribe engineering drawings to their inherently interlinked information. Our approach is relevant to inform the design of document recognition systems for document types that are less well understood than standard OCR, OMR, etc., and serves as a guide to unify the design of future document foundation models.
Background: Agents for computer use (ACUs) are systems that execute complex tasks on digital devices-such as personal computers or mobile phones-given instructions in natural language. These agents automate tasks by controlling software through low-level actions like mouse clicks and touchscreen gestures. However, despite rapid progress, ACUs are not yet mature for everyday use. Objectives: This survey examines the current state-of-the-art, identifies trends, and points out research gaps in the development of practical ACUs. The goal is to provide a comprehensive review and analysis that helps advance general-purpose, robust, and scalable agents for real-world computer use. Methods: We introduce a multifaceted taxonomy of ACUs across three dimensions: (I) the domain perspective, characterizing the contexts in which agents operate; (II) the interaction perspective, describing observation modalities (e.g., screenshots, HTML) and action modalities (e.g., mouse, keyboard, code execution); and (III) the agent perspective, detailing how agents perceive, reason, and learn. We review 87 original research papers about ACUs and 33 relevant datasets, covering both foundation model-based and specialized approaches. Results: Our taxonomy comprehensively structures state-of-the-art approaches and establishes the groundwork for guiding future ACU research. We found that the field is transitioning from specialized agents toward foundation-model-based agents, a shift from text to image-based observation space, and an increasing adoption of behavior cloning methodologies. Furthermore, we identify six key research gaps: insufficient generalization, inefficient learning, limited planning, low task complexity in benchmarks, non-standardized evaluation, and a disconnect between research and practical conditions. Conclusions: To continue rapid improvements in the field, we recommend focusing on: (a) vision-based observations and low-level control to enhance generalization; (b) adaptive learning beyond static prompting; (c) effective planning and reasoning capabilities; (d) realistic, high-complexity benchmarks; (e) standardized evaluation criteria based on task success; and (f) aligning agent design with real-world deployment constraints. Collectively, our findings and proposed directions help develop more general-purpose agents for everyday digital tasks.
Graph augmentations effectively enhance the robustness and generalization of Graph Neural Networks (GNNs), particularly for graph classification tasks. However, existing augmentation methods, like NodeDrop, randomly drop a certain portion of nodes to generate augmented graphs without preserving the essential topological structures of the original graph, potentially modifying label information. To address this issue, we introduce a novel Node-Dropping Augmentation (NDAUG) method for graph classification tasks. Our method leverages node degree as a criterion to selectively drop less important nodes (low-degree) and preserve essential graph structures, generating diverse and informative augmented graphs. Further, in the case of isolated nodes, we develop a structure learning method to reconnect these isolated nodes by learning attention-based relationships between nodes. Experiments demonstrate that combining the proposed NDAUG with existing GNN models yields an average improvement of 2–5% accuracy on eight graph classification benchmarks compared to the state-of-the-art baselines.
In this study, we present a vendor-agnostic, deep learning-based system for the automated analysis of transthoracic pulsed-wave tissue Doppler imaging (TDI), which decouples image acquisition from interpretation and enables centralized, fleet-wide analysis across devices. The model ingests standard TDI from heterogeneous ultrasound systems and automatically extracts key diagnostic markerspeak systolic velocity $(S^{\prime}$), early diastolic velocity $(e^{\prime}$), and late diastolic/atrial contraction velocity $(a^{\prime}$) using a single, unified pipeline. Conceptually, this harmonizes measurements across vendors and sites, improving consistency, comparability, and longitudinal tracking without devicespecific calibration or tooling. Procedurally, a central inference service supports asynchronous batch processing and human-in-the-loop review, thereby shifting analysis off-console, allowing ultrasound scanners to remain fully available for acquisition. In our clinical dataset, which spans two ultrasound vendors and diverse cardiac cycles, the system correctly identified more than 93% of tissuevelocity landmarks. In 50% of studies, all automated detections matched expert annotations, eliminating the need for manual edits. This approach streamlines offline TDI analysis, accelerates turnaround, and supports scalable, standardized cardiac assessments.
We introduce the cooperative network architecture (CNA), a model that represents sensory signals using structured, recurrently connected networks of neurons, termed "nets." Nets are dynamically assembled from overlapping net fragments, which are learned based on statistical regularities in sensory input. This architecture offers robustness to noise, deformation, and generalization to out-of-distribution data, addressing challenges in current vision systems from a novel perspective. We demonstrate that net fragments can be learned without supervision and flexibly recombined to encode novel patterns, enabling figure completion and resilience to noise. Our findings establish CNA as a promising paradigm for developing neural representations that integrate local feature processing with global structure formation, providing a foundation for future research on invariant object recognition.
We present the methodology and results of the Deep Retrieval team for subtask 4b of the CLEF CheckThat! 2025 competition, which focuses on retrieving relevant scientific literature for given social media posts. To address this task, we propose a hybrid retrieval pipeline that combines lexical precision, semantic generalization, and deep contextual re-ranking, enabling robust retrieval that bridges the informal-to-formal language gap. Specifically, we combine BM25-based keyword matching with a FAISS vector store using a fine-tuned INF-Retriever-v1 model for dense semantic retrieval. BM25 returns the top 30 candidates, and semantic search yields 100 candidates, which are then merged and re-ranked via a large language model (LLM)-based cross-encoder. Our approach achieves a mean reciprocal rank at 5 (MRR@5) of 76.46
Artificial Intelligence (AI) is transforming every aspect of modern society. It demonstrates a high potential to contribute to more flexible operations of safety-critical network infrastructures under deep transformation to tackle global challenges, such as climate change, energy transition, efficiency, and digital transformation, including increasing infrastructure resilience to natural and human-made hazards. The widespread adoption of AI creates the conditions for a new and inevitable interaction between humans and AI-based decision systems. In such a scenario, creating an ecosystem in which humans and AI interact healthily, where the roles and positions of both actors are well-defined, is a critical challenge for research and industry in the coming years. This perspective article outlines the challenges and requirements for effective human-AI interaction by taking an interdisciplinary point of view that merges computer science, decision-making sciences, psychological constructs, and industrial practices. The work focuses on three emblematic safety-critical scenarios from two different domains: energy (power grids) and mobility (railway networks and air traffic management).
In recent years, Graph Neural Networks (GNNs) have demonstrated significant influence on the analysis of graph structures by leveraging message-passing mechanisms to aggregate neighborhood information and perform various graph-related tasks from node classification to link prediction. Recently, GNNs have mostly been developed to deal with different types of graph structures, such as homophily (similar labels among connected nodes) and heterophily (dissimilar labels among connected nodes). However, existing methods lack the ability to combine node features and graph topology optimally to deal with heterophily. This paper proposes a Community-HOP-based GNN model for dealing with homophilic and heterophilic graph structures. Specifically, we incorporate valuable insights from the graph community structure to guide the feature aggregation process of the GNN layer to learn diverse graph properties and improve performance on node-level tasks. Extensive experiments on six node-level datasets under standard metrics demonstrate that the Community-HOP method surpasses existing baselines.
To go from (passive) process monitoring to active process control, an effective AI system must learn about the behavior of the complex system from very limited training data, forming an ad-hoc digital twin with respect to process inputs and outputs that captures the consequences of actions on the process's world. We propose a novel methodology based on learning world models that disentangles process parameters in the learned latent representation, allowing for fine-grained control. Representation learning is driven by the latent factors influencing the processes through contrastive learning within a joint embedding predictive architecture. This makes changes in representations predictable from changes in inputs and vice versa, facilitating interpretability of key factors responsible for process variations, paving the way for effective control actions to keep the process within operational bounds. The effectiveness of our method is validated on the example of plastic injection molding, demonstrating practical relevance in proposing specific control actions for a notoriously unstable process.
The 3D Master method streamlines the transfer of product information from design to production, utilizing 3D model files containing Product Manufacturing Information. This approach facilitates direct access to crucial data like materials, geometric dimensions, and tolerances for each step of the metal additive manufacturing (MAM) of parts. By leveraging this data, the 3D Master method enables the automation of accurate cost evaluation. This contribution introduces a method leveraging the 3D Master to automate a precise manufacturing cost calculation for MAM parts using powder-based fusion processes. It proposes a frame based on the quality level of the data provided by the customer to quantify the accuracy of the estimated cost, thanks to a performance index (KPI). A build cost model based on an optimal volumetric energy density calculation achieved through a theoretical and statistical approach is also provided. The study, conducted on 20 reference MAM parts of varying geometrical complexities, demonstrates a relative deviation of normalized actual and calculated cost difference below 10
Automated tissue segmentation in medical imaging plays a critical role in clinical AI-assisted decision-making, and particularly in the assessment of body composition from CT scans. However, acquiring data and annotations of sufficient quality to train deep-learning models is expensive and timeconsuming. In this work, we propose a novel approach to improve data efficiency and model accuracy by leveraging domain knowledge about biologically relevant tissue-specific Hounsfield unit (HU) ranges as an inductive bias for learning. Specifically, we extend the input representation of deep learning-based segmentation models with binary masks indicating potential tissue types, where each binary mask is created from thresholds derived from medical literature. Our method not only enhances segmentation performance by up to 5% for intramuscular adipose tissue but surpasses the performance of the baseline model with 50% of the training data. Our easy-to-apply method thus improves data efficiency and facilitates the development and use of segmentation models in resource-constrained clinical settings.
Humans and animals recognize objects irrespective of the beholder's point of view, which may drastically change their appearance. Artificial pattern recognizers strive to also achieve this, e.g., through translational invariance in convolutional neural networks (CNNs). However, CNNs and vision transformers (ViTs) both perform poorly on rotated inputs. Here we present AMR (artificial mental rotation), a method for dealing with in-plane rotations focusing on large datasets and architectural flexibility, our simple AMR implementation works with all common CNN and ViT architectures. We test it on randomly rotated versions of ImageNet, Stanford Cars, and Oxford Pet. With a top-1 error (averaged across datasets and architectures) of 0.743, AMR outperforms rotational data augmentation (average top-1 error of 0.626) by 19%. We also easily transfer a trained AMR module to a downstream task to improve the performance of a pre-trained semantic segmentation model on rotated CoCo from 32.7 to 55.2 IoU.
Self-attention is essential to Transformer architectures, yet how information is embedded in the self-attention matrices and how different objective functions impact this process remains unclear. We present a mathematical framework to analyze self-attention matrices by deriving the structures governing their weight updates. Using this framework, we demonstrate that bidirectional training induces symmetry in the weight matrices, while autoregressive training results in directionality and column dominance. Our theoretical findings are validated across multiple Transformer models — including ModernBERT, GPT, LLaMA3, and Mistral — and input modalities like text, vision, and audio. Finally, we apply these insights by showing that symmetric initialization improves the performance of encoder-only models on language tasks. This mathematical analysis offers a novel theoretical perspective on how information is embedded through self-attention, thereby improving the interpretability of Transformer models.
At least since the introduction of ChatGPT, the abilities of generative large language models (LLMs), sometimes called GPTs, are at the center of the attention of AI researchers, entrepreneurs, and others. However, for many applications, it is not possible to call an existing LLM service via an API due to data protection concerns or when no task-appropriate LLM exists. On the other hand, deploying or training a private LLM is often prohibitively computationally expensive. In this paper, we give an overview of the most important recent methodologies that help reduce the computational footprint of LLMs. We further present extensive benchmarks for seven methods from two of the most important areas of recent progress: model quantization and low-rank adapters, showcasing how it is possible to leverage state-of-the-art LLMs with limited resources. Our benchmarks include resource consumption metrics (e.g. GPU memory usage), a state-of-the-art quantitative performance evaluation as well as a qualitative performance study conducted by eight individual human raters. Our evaluations show that quantization has a profound effect on GPU memory requirements. However, we also show that these quantization methods, contrary to how they are advertised, cause a noticeable loss in text quality. We further show that low-rank adapters allow effective model fine-tuning with moderate compute resources. For methods that require less than 16 GB of GPU memory, we provide easy-to-use Jupyter notebooks that allow anyone to deploy and fine-tune state-of-the-art LLMs on the Google Colab free tier within minutes without any prior experience or infrastructure.
Online process monitoring is essential to detect failures and respond promptly in automated industrial processes such as injection molding. Traditional systems rely on experienced operators manually defining operational boundaries around a reference signal. We propose a data-driven representation that auto-tunes the sensitivity to a pre-set specificity threshold and automatically detects anomalies alongside interpretable indices that help identify root causes. Our automated system achieved an average AUC of 0.998 and detected 100 percent of the anomalies with the proposed dynamic calibration of the data-driven embedding method. The dynamic calibration, which accounted for drift, boosts the average specificity from 0.362 to 0.869. The outputs also indicate the direction and relative magnitude of characteristic deviations caused by machine parameters, including holding pressure, mold temperature, and injection speed. The AI-derived process boundaries are superior to manual annotation in tested real-world production environments.
Including uncertainty is essential for accurate decision-making in underground applications. We propose a novel approach to consider structural uncertainty in two enhanced geothermal systems (EGSs) using machine learning (ML) models. The results of numerical simulations show that a small change in the structural model can cause a significant variation in the tracer breakthrough curves (BTCs). To develop a more robust method for including structural uncertainty, we train three different ML models: decision tree regression (DTR), random forest regression (RFR), and gradient boosting regression (GBR). DTR and RFR predict the entire BTC at once, but they are susceptible to overfitting and underfitting. In contrast, GBR predicts each time step of the BTC as a separate target variable, considering the possible correlation between consecutive time steps. This approach is implemented using a chain of regression models. The chain model achieves an acceptable increase in RMSE from train to test data, confirming its ability to capture both the general trend and small-scale heterogeneities of the BTCs. Additionally, using the ML model instead of the numerical solver reduces the computational time by six orders of magnitude. This time efficiency allows us to calculate BTCs for 2′000 different reservoir models, enabling a more comprehensive structural uncertainty quantification for EGS cases. The chain model is particularly promising, as it is robust to overfitting and underfitting and can generate BTCs for a large number of structural models efficiently.