Type 2 diabetes (T2D) is a chronic disease currently affecting around 500 million people worldwide with often severe health consequences. Yet, histopathological analyses are still inadequate to infer the glycaemic state of a person based on morphological alterations linked to impaired insulin secretion and β-cell failure in T2D. Giga-pixel microscopy can capture subtle morphological changes, but data complexity exceeds human analysis capabilities. In response, we generate a dataset of pancreas whole-slide images from living donors with multiple chromogenic and multiplex immunofluorescence stainings and train deep learning models to predict the T2D status. Using explainable AI, we make the learned relationships interpretable, quantify them as biomarkers, and assess their association with T2D. Remarkably, the highest prediction performance is achieved by simultaneously focusing on islet α- and δ-cells and neuronal axons, alongside subtle pancreatic alterations in T2D donors such as larger adipocyte clusters, altered islet-adipocyte proximity and smaller islets. This data-driven approach provides a foundation for future research into relevant diagnostic and therapeutic targets, refining several hypotheses regarding tissue alterations associated with T2D.
This article introduces a dataset designed for the detection of partial discharges in transmission power lines using covered conductors through a contact galvanic method, sourced from real environments across 23 different power lines in various locations. Though partially introduced in a Kaggle competition (only 3% of data), its full extent is disclosed here for the first time. The dataset is distinguished by its rich, imbalanced distribution across seven classes, derived from signals processed via a sophisticated voltage-based method, and supplemented with extracted features to aid analysis. Its scale, detailed labeling, and real-world basis offer unparalleled opportunities for developing machine learning algorithms aimed at fault detection. This contribution holds vast potential for reuse in electrical engineering research focused on enhancing power distribution network reliability and safety, particularly in the context of predictive maintenance and understanding partial discharge behaviors.
Can a low-cost software-defined radio (SDR) stack provide usable signatures for GIS-like partial discharge (PD) screening? This micro-study offers an initial confirmation using an SDRplay RSP1A receiver paired with a transient earth voltage (TEV) sensor and a heatmap-first post-processing workflow. A controlled protocol spans alternating PD/no-PD runs, fixed-gain (autogain-off) ablation, focused-frequency operation, and long-capture two-pass measurements. Evidence is intentionally lightweight: qualitative heatmap separability, supported by tail-sensitive spectral metrics on matched scenario pairs. Fixed-gain and long-capture conditions produce the most consistent PD-enhanced signatures, whereas autogain reduces separability in alternating wideband runs. Overall, the results suggest SDR+TEV monitoring can serve as a practical, IoT-oriented early-warning pipeline, and motivate further research on repeatability, calibration, and field deployment constraints.
RoofeNet-Multiis a rooftop photovoltaic (PV) interpretation dataset designed to preserve label uncertainty in nadir aerial imagery by releasing multiple independent annotations per image. It contains 5,359 building-centered orthophoto tiles from the Czech Republic. The release provides 7,402 human annotation sets collected via a web-based platform (mean 1.38 sets per image) and one matched vision–language model (VLM) annotation set per image (5,359 sets), generated with Google Gemini (model identifier gemini-2.5-pro-preview-03-25; see Methods). Among the 5,359 images, 3,857 have exactly one human set and 1,502 have two or more human sets. Each annotation set comprises axis-aligned bounding boxes for (a) roof planes (flat vs. pitched) and (b) rooftop obstructions relevant to PV placement (e.g., chimneys, vents, skylights). Unlike conventional benchmarks that publish a single “ground truth”, RoofeNet-Multisupports uncertainty-aware training and evaluation, studies of human–AI agreement under a shared schema, and preliminary PV potential estimation workflows that can account for ambiguity in roof interpretation.
Contactless partial discharge (PD) monitoring on covered conductors needs more than high offline accuracy: the evaluation protocol must avoid leakage between correlated acquisitions, and the selected representation must be plausible for low-power field stations. We present a leakage-safe grouped benchmark for antenna-based PD detection on a public covered-conductor dataset collected across multiple monitoring contexts. The benchmark compares rich wavelet features, compact infinite impulse response (IIR) edge representations, a handcrafted baseline, and modern deep or time-series comparators under one frozen grouped protocol. Grouped evaluation changes the apparent ranking of methods. The rich wavelet family provides the strongest offline accuracy, reaching the highest grouped Matthews correlation coefficient (MCC), but it is also the most expensive feature path. The enhanced edge IIR model is the most useful field-screening compromise: it reduces false positives relative to the handcrafted baseline while keeping extraction cost far below the wavelet path. A concise ablation shows that selected windows, IIR subbands, cross-band descriptors, envelope features, and classifier capacity each affect the grouped edge result. We also include a short hardware-facing continuation: the edge path can be rewritten as a streamable field-programmable gate array (FPGA)-oriented feature transport, but that result is treated only as simulation-level feasibility rather than board validation. Overall, the study supports a practical two-stage monitoring view: use the grouped benchmark to choose a trusted edge screen, and reserve the richer wavelet path for offline or less-constrained analysis.
Active learning (AL) has the potential to drastically reduce annotation costs in 3D biomedical image segmentation, where expert labeling of volumetric data is both time-consuming and expensive. Yet, existing AL methods are unable to consistently outperform improved random sampling baselines adapted to 3D data, leaving the field without a reliable solution. We introduce Class-stratified Scheduled Power Predictive Entropy (ClaSP PE), a simple and effective query strategy that addresses two key limitations of standard uncertainty-based AL methods: class imbalance and redundancy in early selections. ClaSP PE combines class-stratified querying to ensure coverage of underrepresented structures and log-scale power noising with a decaying schedule to enforce query diversity in early-stage AL and encourage exploitation later. In our evaluation on 24 experimental settings using four 3D biomedical datasets within the comprehensive nnActive benchmark, ClaSP PE is the only method that generally outperforms improved random baselines in terms of both segmentation quality with statistically significant gains, whilst remaining annotation efficient. Furthermore, we explicitly simulate the real-world application by testing our method on four previously unseen datasets without manual adaptation, where all experiment parameters are set according to predefined guidelines. The results confirm that ClaSP PE robustly generalizes to novel tasks without requiring dataset-specific tuning. Within the nnActive framework, we present compelling evidence that an AL method can consistently outperform random baselines adapted to 3D segmentation, in terms of both performance and annotation efficiency in a realistic, close-to-production scenario. Our open-source implementation and clear deployment guidelines make it readily applicable in practice. Code is at https://github.com/MIC-DKFZ/nnActive.
Contactless antenna measurements can make partial-discharge (PD) screening practical for covered-conductor networks, but deployment is a station-shift problem rather than a static classifier-ranking problem. A model trained on previous stations meets new electromagnetic backgrounds, sensor placements, and local event morphologies at a newly monitored station. On 128,671 public Dataset B measurements, a repeated signal-level benchmark reaches mean Matthews correlation coefficient (MCC) 0.609, while source-only stationheld-out transfer drops to approximately 0.153. The gap shows why mixed-station evaluation is not enough for operational PD screening. We study station commissioning: source-model triage, expert review, target-regime audit, and local calibration using a small number of confirmed target positives and target background windows. Station 52007 exposes the key mechanism. At k = 20, cluster-diverse calibration reaches diagnostic MCC 0.120, while random, dominant-first, and regime-proportional diagnostic coverage reach 0.303, 0.334, and 0.361. At k = 50, regime-proportional coverage reaches 0.504. In a ten-seed six-station k = 20 repeat, random and regime-proportional target-positive selection are effectively tied (mean MCC 0.487 and 0.486) and both exceed cluster-diverse selection (0.441). The result is a stationcommissioning benchmark for converting limited expert-confirmed local evidence into auditable station-specific PD screening.
Accurate differentiation between partial discharges (PD) and corona discharges in XLPE-covered conductors is crucial for power system diagnostics, yet remains limited by the lack of specialized, high-fidelity datasets for machine learning (ML) model development. This paper presents a high-resolution dataset (107 samples per 20 ms) acquired using a contactless dual-antenna system under controlled laboratory conditions simulating medium-voltage overhead distribution lines. The dataset includes 100 labeled measurements per class across five discharge types (PD, corona, mixed states, and high-impedance variants) and two background conditions (with and without high voltage), collected over a two-day campaign. By providing experimentally isolated signal types, this resource enables the development and benchmarking of ML models specifically tailored to the PD–corona classification challenge. Key applications include lightweight classification models for edge devices, synthetic data generation to augment limited training sets, and investigations into noise robustness, real-time monitoring, and explainable diagnostics. Through a controlled yet realistic acquisition design, the dataset supports the creation of advanced ML-based tools for non-invasive fault identification—enhancing diagnostic accuracy, mitigating insulation risks, and improving safety in critical power infrastructure.
Semantic segmentation is crucial for various biomedical applications, yet its reliance on large annotated datasets presents a bottleneck due to the high cost and specialized expertise required for manual labeling. Active Learning (AL) aims to mitigate this challenge by querying only the most informative samples, thereby reducing annotation effort. However, in the domain of 3D biomedical imaging, there is no consensus on whether AL consistently outperforms Random sampling. Four evaluation pitfalls hinder the current methodological assessment. These are (1) restriction to too few datasets and annotation budgets, (2) using 2D models on 3D images without partial annotations, (3) Random baseline not being adapted to the task, and (4) measuring annotation cost only in voxels. In this work, we introduce nnActive, an open-source AL framework that overcomes these pitfalls by (1) means of a large scale study spanning four biomedical imaging datasets and three label regimes, (2) extending nnU-Net by using partial annotations for training with 3D patch-based query selection, (3) proposing Foreground Aware Random sampling strategies tackling the foreground-background class imbalance of medical images and (4) propose the foreground efficiency metric, which captures the low annotation cost of background-regions. We reveal the following findings: (A) while all AL methods outperform standard Random sampling, none reliably surpasses an improved Foreground Aware Random sampling; (B) benefits of AL depend on task specific parameters; (C) Predictive Entropy is overall the best performing AL method, but likely requires the most annotation effort; (D) AL performance can be improved with more compute intensive design choices. As a holistic, open-source framework, nnActive can serve as a catalyst for research and application of AL in 3D biomedical imaging. Code is at: https://github.com/MIC-DKFZ/nnActive
Identifying predictive covariates, which forecast individual treatment effectiveness, is crucial for decision-making across different disciplines such as personalized medicine. These covariates, referred to as biomarkers, are extracted from pretreatment data, often within randomized controlled trials, and should be distinguished from prognostic biomarkers, which are independent of treatment assignment. Our study focuses on discovering predictive imaging biomarkers, specific image features, by leveraging pretreatment images to uncover new causal relationships. Unlike laborintensive approaches relying on handcrafted features prone to bias, we present a novel task of directly learning predictive features from images. We propose an evaluation protocol to assess a model's ability to identify predictive imaging biomarkers and differentiate them from purely prognostic ones by employing statistical testing and a comprehensive analysis of image feature attribution. We explore the suitability of deep learning models originally developed for estimating the conditional average treatment effect (CATE) for this task, which have been assessed primarily for their precision of CATE estimation while overlooking the evaluation of imaging biomarker discovery. Our proof-of-concept analysis demonstrates the feasibility and potential of our approach in discovering and validating predictive imaging biomarkers from synthetic outcomes and real-world image datasets. Our code is available at https://github.com/MIC-DKFZ/predictive_image_biomarker_analysis.
Partial discharges (PDs) in XLPE-covered conductors are critical precursors to insulation failure in medium-voltage networks. Despite advancements in radiometric PD detection using deep learning, the classification of multiple fault types remains underexplored, and model interpretability challenges hinder practical deployment. This study presents a spectrogram-based deep learning framework for classifying 12 PD fault types and background conditions using data from a BONI-WHIP antenna. By converting time-domain signals into spectrograms and employing a ResNet50V2 classifier, the framework achieves high multi-class classification accuracy. To enhance interpretability, Gradient-weighted Class Activation Mapping (Grad-CAM) and SHapley Additive Explanations (SHAP) identify the spectral features influencing predictions, aligning with known PD phenomena such as high-frequency emissions during discharges. The results demonstrate the potential for explainable AI in condition monitoring, with further validation under field conditions recommended to confirm its applicability.
Accurate detection of partial discharges (PDs) in medium-voltage overhead transmission lines is critical for preemptive maintenance and avoiding costly outages, yet it is challenged by scarce labeled data and pervasive electromagnetic interference. This paper investigates a hybrid simulation-and-data-driven framework in which synthetically generated PD signals are used to pretrain deep neural networks and are subsequently fine-tuned on a limited set of real overhead-line measurements. The synthetic pipeline systematically varies PD repetition rates, amplitude distributions, vegetation-contact scenarios, and noise conditions, producing diverse time-series and spectrogram-like representations that approximate real operating environments. We conduct a comprehensive ablation study across multiple architectures-Convolutional Neural Networks (CNNs), a Vision Transformer (ViT), and a Long Short-Term Memory (LSTM) network-and analyze their sensitivity to granular sweeps of synthetic-data parameters. CNN-based models decisively outperform ViT and LSTM counterparts on the spectrogram-based classification task, while ViT and LSTM fail to learn meaningful representation. For the successful CNNs, pretraining on carefully parameterized synthetic datasets-particularly those reflecting higher PD activity, such as our Datasets 3 and 4-consistently improves downstream performance on real data, boosting the Matthews Correlation Coefficient (MCC) on imbalanced, cost-sensitive test sets by roughly 10-20% compared with training from scratch. At the same time, we show that poorly aligned synthetic data can degrade generalization, underscoring the need for accurate noise calibration and domain-aligned simulation. Overall, the results confirm that (i) architectural choice is pivotal for PD detection in overhead lines and (ii) well-designed synthetic data is a powerful, practical lever for achieving reliable and cost-effective PD monitoring when real labeled data are limited.
Partial discharges (PD) in cross-linked polyethylene insulated covered conductors (CCs) present a challenge to power system reliability, particularly in areas where vegetation clearance is restricted. While antenna-based PD detection offers a non-contact solution, the scarcity of positive samples and inherent signal noise create a significantly imbalanced dataset, hindering traditional classification approaches. Furthermore, the lack of prior research on Conditional Generative Adversarial Networks (cGANs) for PD detection in CCs makes direct performance evaluation difficult. To address these limitations, this study explores the potential of cGANs in mitigating data scarcity and enhancing PD detection in CCs. We propose a novel hyperparameter tuning methodology that optimizes cGANs based on classification performance using the Matthews Correlation Coefficient as a metric. This approach allows us to indirectly gauge the cGAN's ability to generate realistic, balanced synthetic PD data, that helps classification. Results suggest that a well-tuned cGAN can successfully generate synthetic data to augment limited real-world samples. This expanded dataset significantly enhances the accuracy of subsequent PD classification tasks. Additionally, the method facilitates system adaptability in the event of hardware upgrades (e.g., antennas, ADCs) by reducing the need for extensive new data collection. This study demonstrates the potential of cGANs as a valuable tool for improving PD detection in CCs, leading to enhanced power system reliability and proactive maintenance.
The various limitations of Generative AI, such as hallucinations and model failures, have made it crucial to understand the role of different modalities in Visual Language Model (VLM) predictions. Our work investigates how the integration of information from image and text modalities influences the performance and behavior of VLMs in visual question answering (VQA) and reasoning tasks. We measure this effect through answer accuracy, reasoning quality, model uncertainty, and modality relevance. We study the interplay between text and image modalities in different configurations where visual content is essential for solving the VQA task. Our contributions include (1) the Semantic Interventions (SI)-VQA dataset, (2) a benchmark study of various VLM architectures under different modality configurations, and (3) the Interactive Semantic Interventions (ISI) tool. The SI-VQA dataset serves as the foundation for the benchmark, while the ISI tool provides an interface to test and apply semantic interventions in image and text inputs, enabling more fine-grained analysis. Our results show that complementary information between modalities improves answer and reasoning quality, while contradictory information harms model performance and confidence. Image text annotations have minimal impact on accuracy and uncertainty, slightly increasing image relevance. Attention analysis confirms the dominant role of image inputs over text in VQA tasks. In this study, we evaluate state-of-the-art VLMs that allow us to extract attention coefficients for each modality. A key finding is PaliGemma's harmful overconfidence, which poses a higher risk of silent failures compared to the LLaVA models. This work sets the foundation for rigorous analysis of modality integration, supported by datasets specifically designed for this purpose.
This abstract presents a dataset for the detection of fault types in XLPE-covered conductors utilized in 22 kV medium voltage power distribution systems. We employed an antenna-based approach for detecting partial discharges. The dataset encompasses 12 distinct fault categories, ranging from ground phase faults to inter-phase faults, and no-fault case with steel or covered conductor as fault. We also used three different antennas. Each sample is a single measurement from antenna, consisting of 106 data points as floating numbers. The utilization of the antenna-based method offers the potential for a more cost-effective and straightforward installation for the detection of partial discharges. The objective of dataset is to enhance the identification of fault types, thereby promoting broader adoption of covered conductors in overhead power distribution lines. Such adoption proves particularly beneficial in confined areas, including natural parks, where safety is a prime concern. It is noteworthy that this dataset represents an original contribution, as no prior publication has addressed detection for this specific range of fault types and method of detection.