Deep neural networks are among the most successful algorithms in terms of performance and scalability across different domains. However, since these networks are black boxes, their usability is severely restricted due to a lack of interpretability. Existing interpretability methods do not address the analysis of time-series-based networks specifically enough. This paper shows that an analysis in the frequency domain can not only highlight relevant areas in the input signal better than existing methods but is also more robust to fluctuations in the signal. In this paper, FreqAtt is presented - a framework that enables post-hoc interpretation of time-series analysis. To achieve this, the relevant frequencies are evaluated, and the signal is either filtered or the relevant input data is marked. FreqAtt is evaluated using a wide range of statistical metrics to provide a broad overview of its performance. The results show that using frequency-based attribution, especially in combination with traditional attribution on top of the frequency-optimized signal, provides strong performance across different metrics.
BACKGROUND:Therapeutic decisions in clinical oncology are commonly established through interdisciplinary consensus in multidisciplinary cancer conferences (MCC). Artificial intelligence (AI) may support these processes by generating data-driven treatment recommendations (TR). We developed and evaluated an explainable AI system designed to reproduce MCC-based treatment decisions for metastatic and non-metastatic prostate cancer (PC). METHODS:Clinical data from patients with histologically confirmed PC discussed in MCC between 2015 and 2022 were transformed into structured datasets. A hierarchical modeling framework was implemented to first predict overarching treatment categories and subsequently specify therapeutic strategies. Multiple machine learning and deep learning algorithms were trained to replicate MCC recommendations. Model performance was assessed using F1-scores. RESULTS:A total of 5478 MCC cases including 76 clinical input variables and 23 treatment output parameters were analyzed. The AI system generated automated TR with high predictive accuracy across both hierarchical levels. For high-level categories, F1-scores reached 0.89 for surgery and 0.81 for radiation therapy. For detailed recommendations, F1-scores reached 0.99 for prostatectomy and 0.98 for PSMA-ligand therapy. Lower performance in anti-cancer drug categories likely reflects smaller sample sizes. Feature importance analyses ensured model transparency and interpretability. CONCLUSION:To our knowledge, this study presents one of the first large-scale explainable AI system capable of generating MCC-aligned treatment recommendations for metastatic and non-metastatic PC within a multi-target framework. It incorporates the largest reported number of clinical input and treatment output parameters in this setting. Strong predictive performance and interpretability support its potential as a scalable decision-support tool in multidisciplinary oncology. Prospective validation is warranted.
Dieser Artikel beschreibt Schlüsselanwendungen von künstlicher Intelligenz (KI) in der klinischen Medizin. Bildanalyse: KI optimiert beispielsweise die Zystoskopie bei Blasentumoren und unterstützt die Erkennung von Pneumonie auf Röntgenbildern mit hoher Sensitivität und Spezifität, auch die Detektion von Basalzellkarzinomen in der Dermatologie, und verbesserte Erkennung diabetischer Retinopathie durch Deep Learning wird hervorgehoben. KI-Prädiktion mittels genomischer Analysen (Multi-OMICs): KI kann Therapieansprechen und Überlebenswahrscheinlichkeit vorhersagen, beispielsweise den Nutzen einer Paclitaxel-Chemotherapie bei Magenkarzinom oder das Ansprechen auf neoadjuvante Chemotherapie beim muskelinvasivem Blasenkarzinom. Echtzeit-Monitoring: KI-Algorithmen erfassen kontinuierlich Daten (z. B. über „Wearables“) zur frühzeitigen Erkennung von Komplikationen oder zur Erstellung prädiktiver Modelle für das Therapieansprechen auf Intensivstationen. Klinische Entscheidungsunterstützungssysteme (CDSS) mit KI werden in Zukunft unerlässlich, um die Qualität evidenzbasierter Therapieempfehlungen zu steigern und den klinischen Arbeitsaufwand zu reduzieren. Ein Beispielprojekt ist KITTU, ein KI-gestütztes System zur Generierung evidenzbasierter Therapieempfehlungen für das Prostatakarzinom, Urothelkarzinom (UC) und das Nierenzellkarzinom (RCC) in multidisziplinären Tumorkonferenzen (MCC). Das System transformiert retrospektive Patientendaten sowie Wissen aus klinischen Studien und Leitlinien in softwaretaugliche Formate, um erklärbare KI-Empfehlungen zu erstellen. KITTU nutzt einen zweistufigen Klassifikationsansatz: zuerst übergeordnete High-Level-Empfehlungen (z. B. „Operation“, „medikamentöse Therapie“), gefolgt von spezifischeren Low-Level-Empfehlungen (z. B. „Zystektomie“, „Pembrolizumab“). Die Erklärbarkeit der Empfehlungen wird durch SHAP(SHapley Additive exPlanations)-Werte gewährleistet, die zeigen, welche Patientendaten die Vorhersage positiv oder negativ beeinflusst haben. Diese Informationen werden auf einem interaktiven Dashboard für medizinisches Personal visualisiert.
BACKGROUND:Decisions on the best available treatment in clinical oncology are based on expert opinions in multidisciplinary cancer conferences (MCC). Artificial intelligence (AI) could increase evidence-based treatment by generating additional treatment recommendations (TR). We aimed to develop such an AI system for urothelial carcinoma (UC) and renal cell carcinoma (RCC). METHODS:Comprehensive data of patients with histologically confirmed UC and RCC who received MCC recommendations in the years 2015 - 2022 were transformed into machine readable representations. Development of a two-step process to train a classifier to mimic TR was followed by identification of superordinate and detailed categories of TR. Machine learning (CatBoost, XGBoost, Random Forest) and deep learning (TabPFN, TabNet, SoftOrdering CNN, FCN) techniques were trained. Results were measured by F1-scores for accuracy weights. RESULTS:AI training was performed with 1617 (UC) and 880 (RCC) MCC recommendations (77 and 76 patient input parameters). The AI system generated fully automated TR with excellent F1-scores for UC (e.g. 'Surgery' 0.81, 'Anti-cancer drug' 0.83, 'Gemcitabine/Cisplatin' 0.88) and RCC (e.g. 'Anti-cancer drug' 0.92 'Nivolumab' 0.78, 'Pembrolizumab/Axitinib' 0.89). Explainability is provided by clinical features and their importance score. Finally, TR and explainability were visualized on a dashboard. CONCLUSION:This study demonstrates for the first time AI-generated, explainable TR in UC and RCC with excellent performance results as a potential support tool for high-quality, evidence-based TR in MCC. The comprehensive technical and clinical development sets global reference standards for future AI developments in MCC recommendations in clinical oncology. Next, prospective validation of the results is mandatory.
465 Background: The expert panel in multidisciplinary cancer conferences (MCC) decides on the best available treatment for the individual cancer patient. To support these complex evidence-based decisions, an artificial intelligence (AI) was developed to generate treatment recommendations for renal cell cancer (RCC) patients to support decision-making in MCC. Methods: We have transformed comprehensive patient data (99 individual machine-readable features) of 880 MCC recommendations for RCC from the years 2015 - 2022 into representations that can be used in software development. We developed a two-step process in order to train classifiers to mimic MMC recommendations. First, we identified superordinate categories of the recommendations. Afterwards, we specified the detailed recommendation. For this purpose, we used different machine learning (CatBoost, XGBoost, Random Forest) and deep learning (SoftOrdering-1d-CNN) approaches with 787 training cases and 93 test set cases. Accuracy weights are determined by F1-Score. Results: The KITTU-AI is able to generate fully automated treatment recommendations for patients with histologically confirmed RCC in MCC. First, the AI can decide which kind of superordinate recommendation should be applied, e.g. surgery or anticancer-drugs (Table). Second, our AI system is able to suggest the specific surgical treatment as well as the correct drugs (Table). The AI-generated recommendation is presented explainable based on the clinical features and their importance score. Conclusions: To our knowledge, we present the first time data for fully automated AI-based treatment recommendations for MMC in RCC with promising accuracy rates.Our selected AI architecture is able to learn and to generate medically comprehensible and explainable treatment recommendations. Small numbers of recommended therapies hamper AI training but this will improve with increasing numbers over time. Next, clinical trial data will be implemented to enable a higher level of explainability. Meanwhile, the first prospective validation is ongoing. Accuracy rates for AI-generated treatment recommendations of renal cell cancer based on F1-scores (test set of 93 cases). Task (number) F1-Score ↑ #Classes Class (number) F1-Score High Level Prediction (93) 0.76 (CatBoost) 5 Surgery (10)Medication (45)Aftercare (22)Best supportive care (4)Radiotherapy (12) 0.450.920.880.000.47 Low Level Surgical Prediction (10) 0.81(Soft Ordering) 3 Primary tumor resection (4)Resection of recurrent tumor (1)Metastases resection (5) 0.670.000.73 Low LevelDrug Prediction (45) 0.73 (XGBoost) 8 Sunitinib (8)Nivolumab (7)Cabozantinib (6)Pembrolizumab/Axitinib (5)Nivolumab/Ipilimumab (4)Pazopanib (3)Pembrolizumab (3)Other (9) 0.520.780.400.890.601.001.000.67
Artificial intelligence (AI) promises to be the next revolutionary step in modern society. Yet, its role in all fields of industry and science need to be determined. One very promising field is represented by AI-based decision-making tools in clinical oncology leading to more comprehensive, personalized therapy approaches. In this review, the authors provide an overview on all relevant technical applications of AI in oncology, which are required to understand the future challenges and realistic perspectives for decision-making tools. In recent years, various applications of AI in medicine have been developed focusing on the analysis of radiological and pathological images. AI applications encompass large amounts of complex data supporting clinical decision-making and reducing errors by objectively quantifying all aspects of the data collected. In clinical oncology, almost all patients receive a treatment recommendation in a multidisciplinary cancer conference at the beginning and during their treatment periods. These highly complex decisions are based on a large amount of information (of the patients and of the various treatment options), which need to be analyzed and correctly classified in a short time. In this review, the authors describe the technical and medical requirements of AI to address these scientific challenges in a multidisciplinary manner. Major challenges in the use of AI in oncology and decision-making tools are data security, data representation, and explainability of AI-based outcome predictions, in particular for decision-making processes in multidisciplinary cancer conferences. Finally, limitations and potential solutions are described and compared for current and future research attempts.
As data-driven AI systems become increasingly integrated into industry, concerns have recently arisen regarding potential privacy breaches and the inadvertent leakage of sensitive user data through the exploitation of these systems. In this paper, we explore the intersection of data privacy and AI-powered document analysis systems, presenting a comprehensive benchmark of well-known privacy-preserving methods for the task of document image classification. In particular, we investigate four different privacy methods---Differential Privacy (DP), Federated Learning (FL), Differentially Private Federated Learning (DP-FL), and Secure Multi-Party Computation (SMPC)---on two well-known document benchmark datasets, namely RVL-CDIP and Tobacco3482. Furthermore, we investigate the performance of each method under a variety of configurations for thorough benchmarking. Finally, the privacy strength of each approach is assessed by subjecting the private models to well-known membership inference attacks. Our results demonstrate that, with sufficient tuning of hyperparameters, Differential Privacy (DP) can achieve reasonable performance on the task of document image classification while also ensuring rigorous privacy constraints, both in standalone and federated learning setups. On the other hand, while FL-based approaches present less implementation complexity and incur little to no loss in performance on the task, they do not offer sufficient protection against privacy attacks. By rigorously benchmarking various privacy approaches, our study paves the way for integrating deep document classification models into industrial pipelines while meeting regulatory and ethical standards, including GDPR and the AI Act 2022.
Deep learning has proven to be successful in various domains and for different tasks. However, when it comes to private data several restrictions are making it difficult to use deep learning approaches in these application fields. Recent approaches try to generate data privately instead of applying a privacy-preserving mechanism directly, on top of the classifier. The solution is to create public data from private data in a manner that preserves the privacy of the data. In this work, two very prominent GAN-based architectures were evaluated in the context of private time series classification. In contrast to previous work, mostly limited to the image domain, the scope of this benchmark was the time series domain. The experiments show that especially GSWGAN performs well across a variety of public datasets outperforming the competitor DPWGAN. An analysis of the generated datasets further validates the superiority of GSWGAN in the context of time series generation.
Since the advent of deep learning (DL), the field has witnessed a continuous stream of innovations. However, the translation of these advancements into practical applications has not kept pace, particularly in safety-critical domains where artificial intelligence (AI) must meet stringent regulatory and ethical standards. This is underscored by the ongoing research in eXplainable AI (XAI) and privacy-preserving machine learning (PPML), which seek to address some limitations associated with these opaque and data-intensive models. Despite brisk research activity in both fields, little attention has been paid to their interaction. This work is the first to thoroughly investigate the effects of privacy-preserving techniques on explanations generated by common XAI methods for DL models. A detailed experimental analysis is conducted to quantify the impact of private training on the explanations provided by DL models, applied to six image datasets and five time series datasets across various domains. The analysis comprises three privacy techniques, nine XAI methods, and seven model architectures. The findings suggest non-negligible changes in explanations through the implementation of privacy measures. Apart from reporting individual effects of PPML on XAI, the paper gives clear recommendations for the choice of techniques in real applications. By unveiling the interdependencies of these pivotal technologies, this research marks an initial step toward resolving the challenges that hinder the deployment of AI in safety-critical settings.
With the advent of machine learning in applications of critical infrastructure such as healthcare and energy, privacy is a growing concern in the minds of stakeholders. It is pivotal to ensure that neither the model nor the data can be used to extract sensitive information used by attackers against individuals or to harm whole societies through the exploitation of critical infrastructure. The applicability of machine learning in these domains is mostly limited due to a lack of trust regarding the transparency and the privacy constraints. Various safety-critical use cases (mostly relying on time-series data) are currently underrepresented in privacy-related considerations. By evaluating several privacy-preserving methods regarding their applicability on time-series data, we validated the inefficacy of encryption for deep learning, the strong dataset dependence of differential privacy, and the broad applicability of federated methods.
Citations are generally analyzed using only quantitative measures while excluding qualitative aspects such as sentiment and intent. However, qualitative aspects provide deeper insights into the impact of a scientific research artifact and make it possible to focus on relevant literature free from bias associated with quantitative aspects. Therefore, it is possible to rank and categorize papers based on their sentiment and intent. For this purpose, larger citation sentiment datasets are required. However, from a time and cost perspective, curating a large citation sentiment dataset is a challenging task. Particularly, citation sentiment analysis suffers from both data scarcity and tremendous costs for dataset annotation. To overcome the bottleneck of data scarcity in the citation analysis domain we explore the impact of out-domain data during training to enhance the model performance. Our results emphasize the use of different scheduling methods based on the use case. We empirically found that a model trained using sequential data scheduling is more suitable for domain-specific usecases. Conversely, shuffled data feeding achieves better performance on a cross-domain task. Based on our findings, we propose an end-to-end trainable multi-task model that covers the sentiment and intent analysis that utilizes out-domain datasets to overcome the data scarcity.
Privacy-preservation is of key importance for the transition of modern deep learning algorithms into everyday applications dealing with sensitive data, such as healthcare, finance and several other domains of critical infrastructure. One major impediment of research in computer science is the considerable time investment required to set up experiments and their evaluation. In the domain of privacy-preserving deep learning, this is aggravated by the dispersion of implementations throughout frameworks and libraries. This work introduces and documents PPML-TSA, a versatile framework for privacy-preserving time series classification. Our framework was initially used to evaluate privacy-preserving methods across different model architectures and datasets. PPML-TSA offers a modular design suitable for performing classification on all datasets from the entire UCR & UEA repository. Its modular implementation offers quick and easy adaptation and extension to increase the number of supported model architectures and datasets. The code supports a variety of model architectures (such as AlexNet, FCN, FDN, LSTM, and LeNet) as well as privacy-preserving deep learning methods (Differential Privacy, Federated Learning, their combination and Homomorphic Encryption) out-of-the-box. We believe that our framework facilitates further research on privacy-preserving deep learning, resulting in accelerated innovation and disruption in the field.
In the last decade neural network have made huge impact both in industry and research due to their ability to extract meaningful features from imprecise or complex data, and by achieving super human performance in several domains. However, due to the lack of transparency the use of these networks is hampered in the areas with safety critical areas. In safety-critical areas, this is necessary by law. Recently several methods have been proposed to uncover this black box by providing interpreation of predictions made by these models. The paper focuses on time series analysis and benchmark several state-of-the-art attribution methods which compute explanations for convolutional classifiers. The presented experiments involve gradient-based and perturbation-based attribution methods. A detailed analysis shows that perturbation-based approaches are superior concerning the Sensitivity and occlusion game. These methods tend to produce explanations with higher continuity. Contrarily, the gradient-based techniques are superb in runtime and Infidelity. In addition, a validation the dependence of the methods on the trained model, feasible application domains, and individual characteristics is attached. The findings accentuate that choosing the best-suited attribution method is strongly correlated with the desired use case. Neither category of attribution methods nor a single approach has shown outstanding performance across all aspects.
Since the mid-10s, the era of Deep Learning (DL) has continued to this day, bringing forth new superlatives and innovations each year. Nevertheless, the speed with which these innovations translate into real applications lags behind this fast pace. Safety-critical applications, in particular, underlie strict regulatory and ethical requirements which need to be taken care of and are still active areas of debate. eXplainable AI (XAI) and privacy-preserving machine learning (PPML) are both crucial research fields, aiming at mitigating some of the drawbacks of prevailing data-hungry black-box models in DL. Despite brisk research activity in the respective fields, no attention has yet been paid to their interaction. This work is the first to investigate the impact of private learning techniques on generated explanations for DL-based models. In an extensive experimental analysis covering various image and time series datasets from multiple domains, as well as varying privacy techniques, XAI methods, and model architectures, the effects of private training on generated explanations are studied. The findings suggest non-negligible changes in explanations through the introduction of privacy. Apart from reporting individual effects of PPML on XAI, the paper gives clear recommendations for the choice of techniques in real applications. By unveiling the interdependencies of these pivotal technologies, this work is a first step towards overcoming the remaining hurdles for practically applicable AI in safety-critical domains.
Classification of time-series data is pivotal for a wide range of applications and comes with many challenges. Although the amount of publicly available datasets increases rapidly, deep neural models are only partially exploited in contrast to the traditional methods. These methods get preferred in safety-critical, financial, or medical fields because of their interpretable results. However, their performance and scalability are limited, and finding suitable explanations for time-series classification tasks is challenging due to the intrinsic nature of concealed concepts in the time-series data. Visual analysis of complete time-series comes with an extensive cognitive overload, as it is difficult to perceive and leads to confusion. Therefore, we believe that patch-wise processing of the data results in a more interpretable representation. To bridge this gap, and to reduce the cognitive overload for interpretation of time series data, we propose a novel hybrid approach that utilizes deep neural networks and traditional machine learning algorithms for an interpretable and scale-able time-series classification approach. Both quantitively and qualitatively PatchX shows superiority to its counterparts with an edge of interoperability.
Citations play a vital role in understanding the impact of scientific literature. Generally, citations are analyzed quantitatively whereas qualitative analysis of citations can reveal deeper insights into the impact of a scientific artifact in the community. Therefore, citation impact analysis including sentiment and intent classification enables us to quantify the quality of the citations which can eventually assist us in the estimation of ranking and impact. The contribution of this paper is three-fold. First, we provide ImpactCite, which is an XLNet-based method for citation impact analysis. Second, we propose a clean and reliable dataset for citation sentiment analysis. Third, we benchmark the well-known language models like BERT and ALBERT along with our proposed approach for both tasks of sentiment and intent classification. All evaluations are performed on a set of publicly available citation analysis datasets. Evaluation results reveal that ImpactCite achieves a new state-of-the-art performance for both citation intent and sentiment classification by outperforming the existing approaches by 3.44% and 1.33% in F1-score. Therefore, the evaluation results suggest that ImpactCite is a single solution for both sentiment and intent analysis to better understand the impact of a citation.
Deep learning methods have shown great success in several domains as they process a large amount of data efficiently, capable of solving complex classification, forecast, segmentation, and other tasks. However, they come with the inherent drawback of inexplicability limiting their applicability and trustworthiness. Although there exists work addressing this perspective, most of the existing approaches are limited to the image modality due to the intuitive and prominent concepts. Conversely, the concepts in the time-series domain are more complex and non-comprehensive but these and an explanation for the network decision are pivotal in critical domains like medical, financial, or industry. Addressing the need for an explainable approach, we propose a novel interpretable network scheme, designed to inherently use an explainable reasoning process inspired by the human cognition without the need of additional post-hoc explainability methods. Therefore, class-specific patches are used as they cover local concepts relevant to the classification to reveal similarities with samples of the same class. In addition, we introduce a novel loss concerning interpretability and accuracy that constraints P2ExNet to provide viable explanations of the data including relevant patches, their position, class similarities, and comparison methods without compromising accuracy. Analysis of the results on eight publicly available time-series datasets reveals that P2ExNet reaches comparable performance when compared to its counterparts while inherently providing understandable and traceable decisions.