This study delves into the largely uncharted domain of biases in photoacoustic imaging, spotlighting potential shortcut learning as a key issue in reliable machine learning. Our focus is on hardware variation biases. We identify device-specific traits that create detectable fingerprints in photoacoustic images, demonstrate machine learning's capability to use these discrepancies to determine the device that acquired the image, and highlight their potential impact on machine learning model predictions in downstream tasks, such as disease classification.
Intelligent systems in interventional healthcare depend on the reliable perception of the environment. In this context, photoacoustic tomography (PAT) has emerged as a non-invasive, functional imaging modality with great clinical potential. Current research focuses on converting the high-dimensional, not human-interpretable spectral data into the underlying functional information, specifically the blood oxygenation. One of the largely unexplored issues stalling clinical advances is the fact that the quantification problem is ambiguous, i.e. that radically different tissue parameter configurations could lead to almost identical photoacoustic spectra. In the present work, we tackle this problem with conditional Invertible Neural Networks (cINNs). Going beyond traditional point estimates, our network is used to compute an approximation of the conditional posterior density of tissue parameters given the photoacoustic spectrum. To this end, an automatic mode detection algorithm extracts the plausible solution from the sample-based posterior. According to a comprehensive validation study based on both synthetic and real images, our approach is well-suited for exploring ambiguity in quantitative PAT.
Synthetic medical image generation has evolved as a key technique for neural network training and validation. A core challenge, however, remains in the domain gap between simulations and real data. While deep learning-based domain transfer using Cycle Generative Adversarial Networks and similar architectures has led to substantial progress in the field, there are use cases in which state-of-the-art approaches still fail to generate training images that produce convincing results on relevant downstream tasks. Here, we address this issue with a domain transfer approach based on conditional invertible neural networks (cINNs). As a particular advantage, our method inherently guarantees cycle consistency through its invertible architecture, and network training can efficiently be conducted with maximum likelihood training. To showcase our method's generic applicability, we apply it to two spectral imaging modalities at different scales, namely hyperspectral imaging (pixel-level) and photoacoustic tomography (image-level). According to comprehensive experiments, our method enables the generation of realistic spectral data and outperforms the state of the art on two downstream classification tasks (binary and multi-class). cINN-based domain transfer could thus evolve as an important method for realistic synthetic data generation in the field of spectral imaging and beyond.
Surgical scene understanding is a key prerequisite for contextaware decision support in the operating room. While deep learning-based approaches have already reached or even surpassed human performance in various fields, the task of surgical action recognition remains a major challenge. With this contribution, we are the first to investigate the concept of self-distillation as a means of addressing class imbalance and potential label ambiguity in surgical video analysis. Our proposed method is a heterogeneous ensemble of three models that use Swin Transfomers as backbone and the concepts of self-distillation and multi-task learning as core design choices. According to ablation studies performed with the CholecT45 challenge data via cross-validation, the biggest performance boost is achieved by the usage of soft labels obtained by self-distillation. External validation of our method on an independent test set was achieved by providing a Docker container of our inference model to the challenge organizers. According to their analysis, our method outperforms all other solutions submitted to the latest challenge in the field. Our approach thus shows the potential of self-distillation for becoming an important tool in medical image analysis applications.
Formalizing surgical activities as triplets of the used instruments, actions performed, and target anatomies is becoming a gold standard approach for surgical activity modeling. The benefit is that this formalization helps to obtain a more detailed understanding of tool-tissue interaction which can be used to develop better Artificial Intelligence assistance for image-guided surgery. Earlier efforts and the CholecTriplet challenge introduced in 2021 have put together techniques aimed at recognizing these triplets from surgical footage. Estimating also the spatial locations of the triplets would offer a more precise intraoperative context-aware decision support for computer-assisted intervention. This paper presents the CholecTriplet2022 challenge, which extends surgical action triplet modeling from recognition to detection. It includes weakly-supervised bounding box localization of every visible surgical instrument (or tool), as the key actors, and the modeling of each tool-activity in the form of ‹instrument, verb, target› triplet. The paper describes a baseline method and 10 new deep learning algorithms presented at the challenge to solve the task. It also provides thorough methodological comparisons of the methods, an in-depth analysis of the obtained results across multiple metrics, visual and procedural challenges; their significance, and useful insights for future research directions and applications in surgery.
Peripheral artery disease (PAD) is widespread among the elderly population where narrowing arteries in lower limbs are causing a lack of perfusion. This work explores the benefit of volumetric photoacoustic imaging (v-PAI) over conventional 2D PAI for PAD diagnosis and monitoring. To this end, we leverage the recently proposed approach of Tattoo tomography, which generates a v-PAI representation from a set of 2D PAI slices. Preliminary results of the ongoing study indicate that v-PAI can increase the sensitivity of early-stage PAD detection. Conclusively our Tattoo approach has the potential to become a valuable tool in PAD diagnostics.
Photoakustische Tomographie (PAT) ist eine neuartige Bildgebung, die es ermöglicht, morphologische und funktionelle Gewebeeigenschaften wie zum Beispiel die Sauerstoffsättigung in Echtzeit und räumlich aufgelöst darzustellen. Obwohl dies ein vielversprechender Ansatz für Diagnose, Therapie und Verlaufskontrolle verschiedener Erkrankungen ist, lassen aktuelle photakustische Sonden lediglich die Aufnahme im eingeschränkten Bereich der zweidimensionalen (2D) Bildebene zu. Wir stellen daher in den neuartigen Ansatz der Tattoo-Tomographie vor, welcher ohne externe Trackinghardware eine 3D-Rekonstruktion von PAT Bildern ermöglicht. Zentrales Element ist hierbei ein optisches Muster (Tattoo), welches vor der PAT-Messung auf der Haut oberhalb der Zielstruktur platziert wird. Das klar definierte Design des Musters ermöglicht es, die räumliche Lage der PAT-Sonde aus einzelnen Schnittbildern relativ zum Koordinatensystem des Musters zu bestimmen und daraus ein 3D-Volumen zu rekonstruieren. In einer Erweiterung kann das Tattoo-Konzept zudem zur Bildfusion von PAT mit anderen Bildgebungsverfahren, wie zum Beispiel CT und MRT, verwendet werden. Die Integration von entsprechenden Markern der gewünschten Fusionsmodalität in die Geometrie des optischen Musters ermöglicht hier die Bestimmung der Fusionstransformation. Der Tattoo-Ansatz wurde in Phantom- und in vivo Experimenten evaluiert. Die Ergebnisse zeigen, dass Tattoo-Tomographie eine präzise 3D-Rekonstruktion mit Submillimetergenauigkeit sowie eine Bildfusion zwischen PAT und CT/MRT mit einem Targetregistrierungsfehler < 3mm (Phantom) ermöglicht. Im Gegensatz zu bisherigen Verfahren kommt Tattoo-Tomographie ohne komplexe externe Hardware oder aufwändiges, maßgeschneidertes Training neuronaler Netze aus. Durch seine einfache Anwendbarkeit und Präzision hat der Ansatz somit das Potential, sich zu einem wertvollen Instrument für klinische 3D-Photoakustik zu entwickeln [1].
Photoacoustic imaging potentially allows for the real-time visualization of functional human tissue parameters such as oxygenation but is subject to a challenging underlying quantification problem. While in silico studies have revealed the great potential of deep learning (DL) methodology in solving this problem, the inherent lack of an efficient gold standard method for model training and validation remains a grand challenge. This work investigates whether DL can be leveraged to accurately and efficiently simulate photon propagation in biological tissue, enabling photoacoustic image synthesis. Our approach is based on estimating the initial pressure distribution of the photoacoustic waves from the underlying optical properties using a back-propagatable neural network trained on synthetic data. In proof-of-concept studies, we validated the performance of two complementary neural network architectures, namely a conventional U-Net-like model and a Fourier Neural Operator (FNO) network. Our in silico validation on multispectral human forearm images shows that DL methods can speed up image generation by a factor of 100 when compared to Monte Carlo simulations with 5×108 photons. While the FNO is slightly more accurate than the U-Net, when compared to Monte Carlo simulations performed with a reduced number of photons (5×106), both neural network architectures achieve equivalent accuracy. In contrast to Monte Carlo simulations, the proposed DL models can be used as inherently differentiable surrogate models in the photoacoustic image synthesis pipeline, allowing for back-propagation of the synthesis error and gradient-based optimization over the entire pipeline. Due to their efficiency, they have the potential to enable large-scale training data generation that can expedite the clinical application of photoacoustic imaging.
Optical and acoustic imaging techniques enable noninvasive visualization of structural and functional tissue properties. Data-driven approaches for quantification of these properties are promising, but they rely on highly accurate simulations due to the lack of ground truth knowledge. We recently introduced the open-source simulation and image processing for photonics and acoustics (SIMPA) Python toolkit that has quickly been adopted by the community in the context of the IPASC consortium for standardized reconstruction. We present new developments in the toolkit including e.g. improved tissue and device modeling and provide an outlook on future directions aiming at improving the realism of simulations.
SIGNIFICANCE:Optical and acoustic imaging techniques enable noninvasive visualisation of structural and functional properties of tissue. The quantification of measurements, however, remains challenging due to the inverse problems that must be solved. Emerging data-driven approaches are promising, but they rely heavily on the presence of high-quality simulations across a range of wavelengths due to the lack of ground truth knowledge of tissue acoustical and optical properties in realistic settings. AIM:To facilitate this process, we present the open-source simulation and image processing for photonics and acoustics (SIMPA) Python toolkit. SIMPA is being developed according to modern software design standards. APPROACH:SIMPA enables the use of computational forward models, data processing algorithms, and digital device twins to simulate realistic images within a single pipeline. SIMPA's module implementations can be seamlessly exchanged as SIMPA abstracts from the concrete implementation of each forward model and builds the simulation pipeline in a modular fashion. Furthermore, SIMPA provides comprehensive libraries of biological structures, such as vessels, as well as optical and acoustic properties and other functionalities for the generation of realistic tissue models. RESULTS:To showcase the capabilities of SIMPA, we show examples in the context of photoacoustic imaging: the diversity of creatable tissue models, the customisability of a simulation pipeline, and the degree of realism of the simulations. CONCLUSIONS:SIMPA is an open-source toolkit that can be used to simulate optical and acoustic imaging modalities. The code is available at: https://github.com/IMSY-DKFZ/simpa, and all of the examples and experiments in this paper can be reproduced using the code available at: https://github.com/IMSY-DKFZ/simpa_paper_experiments.
Photoacoustic tomography (PAT) has the potential to recover morphological and functional tissue properties with high spatial resolution. However, previous attempts to solve the optical inverse problem with supervised machine learning were hampered by the absence of labeled reference data. While this bottleneck has been tackled by simulating training data, the domain gap between real and simulated images remains an unsolved challenge. We propose a novel approach to PAT image synthesis that involves subdividing the challenge of generating plausible simulations into two disjoint problems: (1) Probabilistic generation of realistic tissue morphology, and (2) pixel-wise assignment of corresponding optical and acoustic properties. The former is achieved with Generative Adversarial Networks (GANs) trained on semantically annotated medical imaging data. According to a validation study on a downstream task our approach yields more realistic synthetic images than the traditional model-based approach and could therefore become a fundamental step for deep learning-based quantitative PAT (qPAT).
Photoacoustic (PA) imaging has the potential to revolutionize functional medical imaging in healthcare due to the valuable information on tissue physiology contained in multispectral photoacoustic measurements. Clinical translation of the technology requires conversion of the high-dimensional acquired data into clinically relevant and interpretable information. In this work, we present a deep learning-based approach to semantic segmentation of multispectral photoacoustic images to facilitate image interpretability. Manually annotated photoacoustic and ultrasound imaging data are used as reference and enable the training of a deep learning-based segmentation algorithm in a supervised manner. Based on a validation study with experimentally acquired data from 16 healthy human volunteers, we show that automatic tissue segmentation can be used to create powerful analyses and visualizations of multispectral photoacoustic images. Due to the intuitive representation of high-dimensional information, such a preprocessing algorithm could be a valuable means to facilitate the clinical translation of photoacoustic imaging.
In this work, we present the open source “Simulation and Image Processing for Photoacoustic Imaging (SIMPA)” toolkit that facilitates simulation of multispectral photoacoustic images by streamlining the use of state-of-the-art frameworks that numerically approximate the respective forward models. SIMPA provides modules for all the relevant steps for photoacoustic forward simulation: tissue modelling, optical forward modelling, acoustic modelling, noise modelling, as well as image reconstruction. We demonstrate the capabilities of SIMPA by performing image simulation using MCX and k-Wave for the optical and acoustic forward modelling, as well as an experimentally determined noise model and a custom tissue model.
Previous work on 3D freehand photoacoustic imaging has focused on the development of specialized hardware or the use of tracking devices. In this work, we present a novel approach towards 3D volume compounding using an optical pattern attached to the skin. By design, the pattern allows context-aware calculation of the PA image pose in a pattern reference frame, enabling 3D reconstruction while also making the method robust against patient motion. Due to its easy handling optical pattern-enabled context-aware PA imaging could be a promising approach for 3D PA in a clinical environment.
Photoacoustic tomography (PAT) has the potential to recover morphological and functional tissue properties such as blood oxygenation with high spatial resolution and in an interventional setting. However, decades of research invested in solving the inverse problem of recovering clinically relevant tissue properties from spectral measurements have failed to produce solutions that can quantify tissue parameters robustly in a clinical setting. Previous attempts to address the limitations of model-based approaches with machine learning were hampered by the absence of labeled reference data needed for supervised algorithm training. While this bottleneck has been tackled by simulating training data, the domain gap between real and simulated images remains a huge unsolved challenge. As a first step to address this bottleneck, we propose a novel approach to PAT data simulation, which we refer to as "learning to simulate". Our approach involves subdividing the challenge of generating plausible simulations into two disjoint problems: (1) Probabilistic generation of realistic tissue morphology, represented by semantic segmentation maps and (2) pixel-wise assignment of corresponding optical and acoustic properties. In the present work, we focus on the first challenge. Specifically, we leverage the concept of Generative Adversarial Networks (GANs) trained on semantically annotated medical imaging data to generate plausible tissue geometries. According to an initial in silico feasibility study our approach is well-suited for contributing to realistic PAT image synthesis and could thus become a fundamental step for deep learning-based quantitative PAT.
Photoacoustic imaging (PAI) is an emerging medical imaging modality that provides high contrast and spatial resolution. A core unsolved problem to effectively support interventional healthcare is the accurate quantification of the optical tissue properties, such as the absorption and scattering coefficients. The contribution of this work is two-fold. We demonstrate the strong dependence of deep learning-based approaches on the chosen training data and we present a novel approach to generating simulated training data. According to initial in silico results, our method could serve as an important first step related to generating adequate training data for PAI applications.