Deep neural networks are a promising tool for Audio Event Classification. In contrast to other data like natural images, there are many sensible and non-obvious representations for audio data, which could serve as input to these models. Due to their black-box nature, the effect of different input representations has so far mostly been investigated by measuring classification performance. In this work, we leverage eXplainable AI (XAI), to understand the underlying classification strategies of models trained on different input representations. Specifically, we compare two model architectures with regard to relevant input features used for Audio Event Detection: one directly processes the signal as the raw waveform, and the other takes its time-frequency spectrogram representation as input. We show how relevance heatmaps obtained via Layer-wise Relevance Propagation uncover representation-dependent decision strategies. With these insights, we can make a well-informed decision about the best model and input representation in terms of robustness and representativity. Further, we can test whether the model’s classification strategies align with human requirements.
Deep neural networks are a promising tool for Audio Event Classification. In contrast to other data like natural images, there are many sensible and non-obvious representations for audio data, which could serve as input to these models. Due to their black-box nature, the effect of different input representations has so far mostly been investigated by measuring classification performance. In this work, we leverage eXplainable AI (XAI), to understand the underlying classification strategies of models trained on different input representations. Specifically, we compare two model architectures with regard to relevant input features used for Audio Event Detection: one directly processes the signal as the raw waveform, and the other takes in its time-frequency spectrogram representation. We show how relevance heatmaps obtained via "Siren"Layer-wise Relevance Propagation uncover representation-dependent decision strategies. With these insights, we can make a well-informed decision about the best input representation in terms of robustness and representativity and confirm that the model's classification strategies align with human requirements.
While machine learning is currently transforming the field of histopathology, the domain lacks a comprehensive evaluation of state-of-the-art models based on essential but complementary quality requirements beyond a mere classification accuracy. In order to fill this gap, we developed a new methodology to extensively evaluate a wide range of classification models, including recent vision transformers, and convolutional neural networks such as: ConvNeXt, ResNet (BiT), Inception, ViT and Swin transformer, with and without supervised or self-supervised pretraining. We thoroughly tested the models on five widely used histopathology datasets containing whole slide images of breast, gastric, and colorectal cancer and developed a novel approach using an image-to-image translation model to assess the robustness of a cancer classification model against stain variations. Further, we extended existing interpretability methods to previously unstudied models and systematically reveal insights of the models' classification strategies that allow for plausibility checks and systematic comparisons. The study resulted in specific model recommendations for practitioners as well as putting forward a general methodology to quantify a model's quality according to complementary requirements that can be transferred to future model architectures.
—While machine learning is currently transforming the field of histopathology, the domain lacks a comprehensive evaluation of state-of-the-art models based on essential but complementary quality requirements beyond a mere classification accuracy. In order to fill this gap, we conducted an extensive evaluation by benchmarking a wide range of classification models, including recent vision transformers, convolutional neural networks and hybrid models comprising transformer and convolutional models. We thoroughly tested the models on five widely used histopathology datasets containing whole slide images of breast, gastric, and colorectal cancer and developed a novel approach using an image-to-image translation model to assess the robustness of a cancer classification model against stain variations. Further, we extended existing interpretability methods to previously unstudied models and systematically reveal insights of the models’ classification strategies that allow for plausibility checks and systematic comparisons. The study resulted in specific model recommendations for practitioners as well as putting forward a general methodology to quantify a model’s quality according to complementary requirements that can be transferred to future model architectures.
1 Is truth dead in the information age? A president who entitles nuclear energy to be renewable and emission-free and congratulates the army for taking over airports during the American revolutionary war; an enormous search engine using big data to suggest the fastest route home from work or to apparently know your consumer habits better than yourself; a filter bubble of flat-earthers who conspire against the so-called lie-spreading government which sprays population-controlling chemtrails into the sky; welcome to the 21st century. Since these uncanny and twisted examples are abstracted from our everyday lives, it is possible to draw parallels between the fictional world of the dystopian novel "Nineteen Eighty-Four" (1949) by George Orwell and the distorted reality fabricated by political spokesmen and social media. For instance, the surveillance through technology in the form of television screens observing the viewer are replaced by mobile devices continuously tracking your digital footprint. Moreover, "Newspeak", a language invented by the totalitarian party meant to prevent free thinking, gives way to political correctness in our society forcing public figures as well as anyone setting foot on the territory of the world wide web to mind their words. Therefore, it is no coincidence that the aforesaid manipulation tools and the emergence of buzzwords like ‘fake news' and ‘alternative facts' lead to the deduction that we live in an Orwellian world. Although some of the stated theories are reasonable, is it this straightforward to announce a truth decay as the result of emotions replacing facts? Is truth dead in the information age? Does George Orwell himself refer to the current situation when making a statement such as "The very concept of objective truth is fading out of the world. Lies will pass into history.". Firstly, and most importantly, some preliminary work to get a better grasp of the entire topic is required namely, to answer the following questions: What in particular is defined by the ubiquitous but vague term ‘truth'? Why is the pursuit of wisdom and truth so deeply rooted in the subconsciousness of our species and do we even truly want to know the truth? How did humanity handle the issue of knowledge sharing before the emergence of the data processing machine we call computers? Which circumstances and real-life examples lead to this bold and striking claim that truth is dead? Which arguments dispel it? Rather than a thorough and in-depth analysis, this essay is primarily meant to be food for thought covering different subject areas ranging from physics, through politics, to philosophy.