In the era of artificial intelligence, machines are demonstrating an unprecedented capacity to learn from massive amounts of real-world data to perform human-like cognitive processes, enabling them to recognize environments, objects, and conditions and make critical decisions more accurately than ever. In the medical field, the potential to generate realistic, privacy-preserving, unbiased synthetic data can be the key to unlocking the potential of artificial intelligence in medicine and overcoming the current barriers such as data privacy concerns and high data curation costs. Advanced data-driven solutions could lead towards more robust clinical decision support systems and enhanced clinical training. This Perspective critically examines current and emerging advances in synthetic data generation, and highlights its anticipated transformational effect for early and efficient prevention, diagnosis and treatment of gastrointestinal diseases. Research challenges and directions are identified for leveraging the benefits of synthetic data as well as translating and adopting them in clinical workflows.
Effective shadow removal is pivotal in enhancing visual image quality in applications such as computer vision and digital photography. During the last decades physics-based and machine learning methodologies have been proposed; however, most of them have limited capacity in capturing complex shadow patterns due to restrictive model assumptions, neglecting the fact that shadows usually appear at different scales. Existing benchmarking shadow removal datasets have a limited number of images with simple scenes containing uniform shadows cast by single objects, whereas only a few of them include both manual shadow annotations and paired shadow-free images. To address these limitations, in natural and urban scenes with complex shadows, the contribution of this study is twofold, it: a) proposes a novel deep learning architecture, named Soft-Hard Attention U-net (SHAU), focusing on multiscale shadow removal; b) provides a novel synthetic dataset, named Multiscale Shadow Removal Dataset (MSRD), containing complex shadow patterns of multiple scales, serving as a privacy-preserving benchmark dataset. SHAU incorporates both soft and hard attention modules, along with multiscale feature extraction blocks, enable effective shadow removal of different scales. Experimental results show that SHAU outperforms state-of-the-art methods, improving Peak Signal-to-Noise Ratio and Root Mean Square Error in shadow areas by 18.7% and 55.1%, respectively.
Synthetic Data Generation (SDG) based on Artificial Intelligence (AI) can transform the way clinical medicine is delivered by overcoming privacy barriers that currently render clinical data sharing difficult. This is the key to accelerating the development of digital tools contributing to enhanced patient safety. Such tools include robust data-driven clinical decision support systems, and example-based digital training tools that will enable healthcare professionals to improve their diagnostic performance for enhanced patient safety. This study focuses on the clinical evaluation of medical SDG, with a proof-of-concept investigation on diagnosing Inflammatory Bowel Disease (IBD) using Wireless Capsule Endoscopy (WCE) images. Its scientific contributions include (a) a novel protocol for the systematic Clinical Evaluation of Medical Image Synthesis (CEMIS); (b) a novel variational autoencoder-based model, named TIDE-II, which enhances its predecessor model, TIDE (This Intestine Does not Exist), for the generation of high-resolution synthetic WCE images; and (c) a comprehensive evaluation of the synthetic images using the CEMIS protocol by 10 international WCE specialists, in terms of image quality, diversity, and realism, as well as their utility for clinical decision-making. The results show that TIDE-II generates clinically plausible, very realistic WCE images, of improved quality compared to relevant state-of-the-art generative models. Concludingly, CEMIS can serve as a reference for future research on medical image-generation techniques, while the adaptation/extension of the architecture of TIDE-II to other imaging domains can be promising.
The segmentation of anatomical structures in medical images and particularly in MRI scans, is essential for clinical diagnosis and monitoring disease progression. While Deep Learning (DL) architectures, such as U-Net and its extensions are very effective in medical image segmentation tasks, they often struggle with preserving fine-grained details and global contextual information. This is especially challenging for MRI data segmentation, where anatomical structures are characterized by irregular boundaries and variations in shape, contrast, and scale. To address this challenge, we propose a novel DL architecture for MRI segmentation across different anatomical structures. Specifically, the architecture introduces a module, named Multi-Resolution Feature Fusion (MRFF), that can be easily integrated into any U-Net-like architecture. The MRFF is integrated in all levels of an encode-decoder structure, along with attention mechanisms and skip connections to extract features at multiple resolutions, enabling the model to capture both fine-grained details and global contextual information. We evaluate the MRFFU-Net on two publicly available benchmark MRI datasets of different anatomical targets; one for Cerebrospinal Fluid (CSF) segmentation in spinal MR scans, and one for left atrium cardiac segmentation, from the Medical Segmentation Decathlon (MSD) challenge. Experimental results indicate that MRFFU-Net outperforms state-of-the-art models across multiple evaluation metrics, demonstrating its effectiveness in MRI segmentation.
Gastrointestinal (GI) imaging via Wireless Capsule Endoscopy (WCE) generates a large number of images requiring manual screening. Deep learning-based Clinical Decision Support (CDS) systems can assist screening, yet their performance relies on the existence of large, diverse, training medical datasets. However, the scarcity of such data, due to privacy constraints and annotation costs, hinders CDS development. Generative machine learning offers a viable solution to combat this limitation. While current Synthetic Data Generation (SDG) methods, such as Generative Adversarial Networks (GANs) and Variational Autoencoders (VAEs) have been explored, they often face challenges with training stability and capturing sufficient visual diversity, especially when synthesizing abnormal findings. This work introduces a novel VAE-based methodology for medical image synthesis and presents its application for the generation of WCE images. The novel contributions of this work include a) multiscale extension of the Vector Quantized VAE model, named as Multiscale Vector Quantized Variational Autoencoder (MSVQ-VAE); b) unlike other VAE-based SDG models for WCE image generation, MSVQ-VAE is used to seamlessly introduce abnormalities into normal WCE images; c) it enables conditional generation of synthetic images, enabling the introduction of different types of abnormalities into the normal WCE images; d) it performs experiments with a variety of abnormality types, including polyps, vascular and inflammatory conditions. The utility of the generated images for CDS is assessed via image classification. Comparative experiments demonstrate that training a CDS classifier using the abnormal images generated by the proposed methodology yield comparable results with a classifier trained with only real data. The generality of the proposed methodology promises its applicability to various domains related to medical multimedia.
Medical image synthesis has emerged as a promising solution to address the limited availability of annotated medical data needed for training machine learning algorithms in the context of image-based Clinical Decision Support (CDS) systems. To this end, Generative Adversarial Networks (GANs) have been mainly applied to support the algorithm training process by generating synthetic images for data augmentation. However, in the field of Wireless Capsule Endoscopy (WCE), the limited content diversity and size of existing publicly available annotated datasets adversely affect both the training stability and synthesis performance of GANs. In this paper a novel Variational Autoencoder (VAE) architecture is proposed for WCE image synthesis, namely 'This Intestine Does not Exist' (TIDE). This is the first VAE architecture comprising multiscale feature extraction convolutional blocks and residual connections. Its advantage is that it enables the generation of high-quality and diverse datasets even with a limited number of training images. Contrary to the current approaches, which are oriented towards the augmentation of the available datasets, this study demonstrates that using TIDE, real WCE datasets can be fully substituted by artificially generated ones, without compromising classification performance of CDS. It performs a spherical experimental evaluation study that covers both quantitative and qualitative aspects, including a user evaluation study performed by WCE specialists, which validate from a medical viewpoint that both the normal and abnormal WCE images synthesized by TIDE are sufficiently realistic. The quantitative results obtained by comparative experiments validate that the proposed architecture outperforms the state-of-the-art.
Visual impairment affects a significantly large number of individuals and in many cases it may disturb their daily life. This study presents a novel prototype wearable device which aims to assist them in tasks such as navigation within outdoor sites of cultural interest, while it may be used to enhance their experience. It presents an overview of the system architecture, which includes a wearable component, a cloud infrastructure, and a remote assistant service. The wearable components is a set of smart glasses, equipped with a stereoscopic camera, serving as a link between the cloud infrastructure and the remote assisting service of the system. User evaluation trials of the system have been performed, focusing on the obstacle detection capabilities of the system, and preliminary results are reported. The results provide initial evidence about the usability of the system, and indicate its potential as a future assistive technology for people with impairments.
The assistive navigation of visually impaired individuals requires the development of different algorithms for obstacle detection, recognition, avoidance, and path planning. The assessment and optimization of such algorithms in the real world is a painstaking process that requires repetitive measurements under stable conditions, which is usually difficult to achieve and costly. To this end, digital twin environments can be used to replicate relevant real-life situations, enabling the evaluation and optimization of algorithms through adjustable and cost-effective simulations. This chapter presents a digital twin framework for the simulation and evaluation of assistive navigation systems, and its application in the context of a camera-based wearable system for visually impaired individuals in an outdoor cultural space. The system incorporates an obstacle avoidance algorithm based on fuzzy logic. The utility and the effectiveness of this framework are demonstrated with an indicative simulation study.
The generalization performance of deep learning models is closely associated with the number and diversity of data available upon training. While in many applications there is a large number of data available in public, in domains such as medical image analysis, the data availability is limited. This can be largely attributed to data privacy legislations, including the General Data Protection Regulation (GDPR), and the cost of data annotation by experts. Aiming to address this issue, data augmentation approaches employing deep generative models have emerged. Existing augmentation techniques are primarily based on Generative Adversarial Networks (GANs). However, ill-posed training issues of GANs such as nonconvergence, mode collapse and instability in conjunction with their demand for large scale training datasets, complicate their use in medical imaging modalities. Motivated by these issues, this paper investigates the performance of alternative generative models i.e., Variational Autoencoders (VAEs) in endoscopic image synthesis tasks. Contrary to the conventional GAN-based approaches that aiming at augmenting the existing endoscopic datasets the proposed methodology constitutes feasible the complete substitution of medical imaging datasets from real individuals with artificially generated ones. The experimental results obtained validate the effectiveness of the proposed methodology over the state-of-art.
Machine Learning (ML) applications are growing in an unprecedented scale. The development of easy-to-use machine-learning application frameworks has enabled the development of advanced artificial intelligence (AI) applications with only a few lines of self-explanatory code. As a result, ML-based AI is becoming approachable by mainstream developers and small businesses. However, the deployment of ML algorithms for remote high throughput ML task execution, involving complex data-processing pipelines can still be challenging, especially with respect to production ML use cases. To cope with this issue, in this paper we propose a novel system architecture that enables Algorithm-agnostic, Scalable ML (ASML) task execution for high throughput applications. It aims to provide an answer to the research question of how to design and implement an abstraction framework, suitable for the deployment of end-to-end ML pipelines in a generic and standard way. The proposed ASML architecture manages horizontal scaling, task scheduling, reporting, monitoring and execution of multi-client ML tasks using modular, extensible components that abstract the execution details of the underlying algorithms. Experiments in the context of obstacle detection and recognition, as well as in the context of abnormality detection in medical image streams, demonstrate its capacity for parallel, mission critical, task execution.
Convolutional neural networks (CNNs) are artificial learning systems typically based on two operations: convolution, which implements feature extraction through filtering, and pooling, which implements dimensionality reduction. The impact of pooling in the classification performance of the CNNs has been highlighted in several previous works, and a variety of alternative pooling operators have been proposed. However, only a few of them tackle with the uncertainty that is naturally propagated from the input layer to the feature maps of the hidden layers through convolutions. In this article we present a novel pooling operation based on (type-1) fuzzy sets to cope with the local imprecision of the feature maps, and we investigate its performance in the context of image classification. Fuzzy pooling is performed by fuzzification, aggregation, and defuzzification of feature map neighborhoods. It is used for the construction of a fuzzy pooling layer that can be applied as a drop-in replacement of the current, crisp, pooling layers of CNN architectures. Several experiments using publicly available datasets show that the proposed approach can enhance the classification performance of a CNN. A comparative evaluation shows that it outperforms state-of-the-art pooling approaches.
Neural network-based solutions are under development to alleviate physicians from the tedious task of small-bowel capsule endoscopy reviewing. Computer-assisted detection is a critical step, aiming to reduce reading times while maintaining accuracy. Weakly supervised solutions have shown promising results; however, video-level evaluations are scarce, and no prospective studies have been conducted yet. Automated characterization (in terms of diagnosis and pertinence) by supervised machine learning solutions is the next step. It relies on large, thoroughly labeled databases, for which preliminary "ground truth" definitions by experts are of tremendous importance. Other developments are under ways, to assist physicians in localizing anatomical landmarks and findings in the small bowel, in measuring lesions, and in rating bowel cleanliness. It is still questioned whether artificial intelligence will enter the market with proprietary, built-in or plug-in software, or with a universal cloud-based service, and how it will be accepted by physicians and patients.
Bone metastasis is among the most frequent in diseases to patients suffering from metastatic cancer, such as breast or prostate cancer. A popular diagnostic method is bone scintigraphy where the whole body of the patient is scanned. However, hot spots that are presented in the scanned image can be misleading, making the accurate and reliable diagnosis of bone metastasis a challenge. Artificial intelligence can play a crucial role as a decision support tool to alleviate the burden of generating manual annotations on images and therefore prevent oversights by medical experts. So far, several state-of-the-art convolutional neural networks (CNN) have been employed to address bone metastasis diagnosis as a binary or multiclass classification problem achieving adequate accuracy (higher than 90%). However, due to their increased complexity (number of layers and free parameters), these networks are severely dependent on the number of available training images that are typically limited within the medical domain. Our study was dedicated to the use of a new deep learning architecture that overcomes the computational burden by using a convolutional neural network with a significantly lower number of floating-point operations (FLOPs) and free parameters. The proposed lightweight look-behind fully convolutional neural network was implemented and compared with several well-known powerful CNNs, such as ResNet50, VGG16, Inception V3, Xception, and MobileNet on an imaging dataset of moderate size (778 images from male subjects with prostate cancer). The results prove the superiority of the proposed methodology over the current state-of-the-art on identifying bone metastasis. The proposed methodology demonstrates a unique potential to revolutionize image-based diagnostics enabling new possibilities for enhanced cancer metastasis monitoring and treatment.
Every day, visually challenged people (VCP) face mobility restrictions and accessibility limitations. A short walk to a nearby destination, which for other individuals is taken for granted, becomes a challenge. To tackle this problem, we propose a novel visual perception system for outdoor navigation that can be evolved into an everyday visual aid for VCP. The proposed methodology is integrated in a wearable visual perception system (VPS). The proposed approach efficiently incorporates deep learning, object recognition models, along with an obstacle detection methodology based on human eye fixation prediction using Generative Adversarial Networks. An uncertainty-aware modeling of the obstacle risk assessment and spatial localization has been employed, following a fuzzy logic approach, for robust obstacle detection. The above combination can translate the position and the type of detected obstacles into descriptive linguistic expressions, allowing the users to easily understand their location in the environment and avoid them. The performance and capabilities of the proposed method are investigated in the context of safe navigation of VCP in outdoor environments of cultural interest through obstacle recognition and detection. Additionally, a comparison between the proposed system and relevant state-of-the-art systems for the safe navigation of VCP, focused on design and user-requirements satisfaction, is performed.
Visual impairment restricts everyday mobility and limits the accessibility of places, which for the non-visually impaired is taken for granted. A short walk to a close destination, such as a market or a school becomes an everyday challenge. In this chapter, we present a novel solution to this problem that can evolve into an everyday visual aid for people with limited sight or total blindness. The proposed solution is a digital system, wearable like smart-glasses, equipped with cameras. An intelligent system module, incorporating efficient deep learning and uncertainty-aware decision-making algorithms, interprets the video scenes, translates them into speech, and describes them to the user through audio. The user can almost naturally interact with the system via a speech-based user interface, which is also capable of understanding the user’s emotions. The capabilities of this system are investigated in the context of accessibility and guidance to outdoor environments of cultural interest, such as the historic triangle of Athens. A survey of relevant state-of-the-art systems, technologies and services is performed, identifying critical system components that better adapt to the goals of the system, user needs and requirements, toward a user-centered architecture design.
In this paper, we propose a novel Fully Convolutional Neural Network (FCN) architecture aiming to aid the detection of abnormalities, such as polyps, ulcers and blood, in gastrointestinal (GI) endoscopy images. The proposed architecture, named Look-Behind FCN (LB-FCN), is capable of extracting multi-scale image features by using blocks of parallel convolutional layers with different filter sizes. These blocks are connected by Look-Behind (LB) connections, so that the features they produce are combined with features extracted from behind layers, thus preserving the respective information. Furthermore, it has a smaller number of free parameters than conventional Convolutional Neural Network (CNN) architectures, which makes it suitable for training with smaller datasets. This is particularly useful in medical image analysis, since data availability is usually limited due to ethicolegal constraints. The performance of LB-FCN is evaluated on both flexible and wireless capsule endoscopy datasets, reaching 99.72% and 93.50%, in terms of Area Under receiving operating Characteristic (AUC) respectively. (C) 2018 Elsevier Ltd. All rights reserved.
Staircase detection in natural images has several applications in the context of robotics and visually impaired navigation. Previous works are mainly based on handcrafted feature extraction and supervised learning using fully annotated images. In this work we address the problem of staircase detection in weakly labeled natural images, using a novel Fully Convolutional neural Network (FCN), named LB-FCN light. The proposed network is an enhanced version of our recent Look-Behind FCN (LB-FCN), suitable for deployment on mobile and embedded devices. Its architecture features multi-scale feature extraction, depthwise separable convolutions and residual learning. To evaluate its computational and classification performance, we have created a weakly-labeled benchmark dataset from publicly available images. The results from the experimental evaluation of LB-FCN light indicate its advantageous performance over the relevant state-of-the-art architectures.
The generalization performance in deep learning is linked to the size and the variations of the samples available during training. This is apparent in the domain of computeraided gastrointestinal tract abnormality detection, where the lesions can vary a lot from each other and the number of available samples is limited, mainly due to personal data protection legislations. In this work we present a novel approach of tackling the problem of limited training data availability by making use of artificially generated images. More specifically we trained a Generative Adversarial Network (GAN) using Wireless Capsule Endoscopy (WCE) images to generate fake but realistic images from the small bowel. The generated images were then used to train a Convolutional Neural Network (CNN) to identify inflammatory conditions on real WCE images. To evaluate the performance of our approach, in our experiments we compare the generalization performance of the same CNN architecture trained separately with real and fake images, obtaining 90.9% and 79.1% Area Under Receiver Operating Characteristic (AUC), respectively. The results show that training using solely artificially generated data can be effective in cases where real training data are inaccessible.
The detection of abnormalities in endoscopic video frames can contribute in the early and more accurate detection of pathologic conditions. In this paper we present a novel Convolutional Neural Network (CNN) architecture for automatic detection of abnormal images in endoscopic video sequences. It features multiscale representation of the endoscopic images in its structure, and peephole connections contributing in enhanced generalization with less computational requirements. An important aspect of the proposed architecture is that it enables weakly-supervised learning, using only semantically annotated images. A novel cross-dataset experimental study is performed to investigate its generalization performance on various publicly available datasets. The results validate that the proposed architecture outperforms recent approaches, with results reaching up to 90.66% in terms of the area under the receiver operating characteristic.
Artur Krukowski合作论文数Intracom S. A. Telecom Solution
Telco Business Software Division
R&D Unit1