[This corrects the article DOI: 10.3389/fonc.2024.1417862.].
Solutions to vision tasks in gastrointestinal endoscopy (GIE) conventionally use image encoders pretrained in a supervised manner with ImageNet-1k as backbones. However, the use of modern self-supervised pretraining algorithms and a recent dataset of 100k unlabelled GIE images (Hyperkvasir-unlabelled) may allow for improvements. In this work, we study the fine-tuned performance of models with ResNet50 and ViT-B backbones pretrained in self-supervised and supervised manners with ImageNet-1k and Hyperkvasir-unlabelled (self-supervised only) in a range of GIE vision tasks. In addition to identifying the most suitable pretraining pipeline and backbone architecture for each task, out of those considered, our results suggest three general principles. Firstly, that self-supervised pretraining generally produces more suitable backbones for GIE vision tasks than supervised pretraining. Secondly, that self-supervised pretraining with ImageNet-1k is typically more suitable than pretraining with Hyperkvasir-unlabelled, with the notable exception of monocular depth estimation in colonoscopy. Thirdly, that ViT-Bs are more suitable in polyp segmentation and monocular depth estimation in colonoscopy, ResNet50s are more suitable in polyp detection, and both architectures perform similarly in anatomical landmark recognition and pathological finding characterisation. We hope this work draws attention to the complexity of pretraining for GIE vision tasks, informs this development of more suitable approaches than the convention, and inspires further research on this topic to help advance this development. Code available: https://www.github.com/ESandML/SSL4GIE.
IntroductionColorectal cancer (CRC) is one of the main causes of deaths worldwide. Early detection and diagnosis of its precursor lesion, the polyp, is key to reduce its mortality and to improve procedure efficiency. During the last two decades, several computational methods have been proposed to assist clinicians in detection, segmentation and classification tasks but the lack of a common public validation framework makes it difficult to determine which of them is ready to be deployed in the exploration room.MethodsThis study presents a complete validation framework and we compare several methodologies for each of the polyp characterization tasks.ResultsResults show that the majority of the approaches are able to provide good performance for the detection and segmentation task, but that there is room for improvement regarding polyp classification.DiscussionWhile studied show promising results in the assistance of polyp detection and segmentation tasks, further research should be done in classification task to obtain reliable results to assist the clinicians during the procedure. The presented framework provides a standarized method for evaluating and comparing different approaches, which could facilitate the identification of clinically prepared assisting methods.
Estimating rigid objects' poses is one of the fundamental problems in computer vision, with a range of applications across automation and augmented reality. Most existing approaches adopt one network per object class strategy, depend heavily on objects' 3D models, depth data, and employ a time-consuming iterative refinement, which could be impractical for some applications. This paper presents a novel approach, CVAM-Pose, for multi-object monocular pose estimation that addresses these limitations. The CVAM-Pose method employs a label-embedded conditional variational autoencoder network, to implicitly abstract regularised representations of multiple objects in a single low-dimensional latent space. This autoencoding process uses only images captured by a projective camera and is robust to objects' occlusion and scene clutter. The classes of objects are one-hot encoded and embedded throughout the network. The proposed label-embedded pose regression strategy interprets the learnt latent space representations utilising continuous pose representations. Ablation tests and systematic evaluations demonstrate the scalability and efficiency of the CVAM-Pose method for multi-object scenarios. The proposed CVAM-Pose outperforms competing latent space approaches. For example, it is respectively 25 using the AR_VSD metric on the Linemod-Occluded dataset. It also achieves results somewhat comparable to methods reliant on 3D models reported in BOP challenges. Code available: https://github.com/JZhao12/CVAM-Pose
Polyp segmentation within colonoscopy video frames using deep learning models has the potential to automate the workflow of clinicians. This could help improve the early detection rate and characterization of polyps which could progress to colorectal cancer. Recent state-of-the-art deep learning polyp segmentation models have combined the outputs of Fully Convolutional Network architectures and Transformer Network architectures which work in parallel. In this paper we propose modifications to the current state-of-the-art polyp segmentation model FCBFormer. The transformer architecture of the FCBFormer is replaced with a SwinV2 Transformer-UNET and minor changes to the Fully Convolutional Network architecture are made to create the FCB-SwinV2 Transformer. The performance of the FCB-SwinV2 Transformer is evaluated on the popular colonoscopy segmentation bench-marking datasets Kvasir-SEG and CVC-ClinicDB. Generalizability tests are also conducted. The FCB-SwinV2 Transformer is able to consistently achieve higher mDice scores across all tests conducted and therefore represents new state-of-the-art performance. Issues found with how colonoscopy segmentation model performance is evaluated within literature are also re-ported and discussed. One of the most important issues identified is that when evaluating performance on the CVC-ClinicDB dataset it would be preferable to ensure no data leakage from video sequences occurs during the training/validation/test data partition.
Colorectal cancer is one of the most common cancers in the world. While colonoscopy is an effective screening technique, navigating an endoscope through the colon to detect polyps is challenging. A 3D map of the observed surfaces could enhance the identification of unscreened colon tissue and serve as a training platform. However, reconstructing the colon from video footage remains difficult. Learning-based approaches hold promise as robust alternatives, but necessitate extensive datasets. Establishing a benchmark dataset, the 2022 EndoVis sub-challenge SimCol3D aimed to facilitate data-driven depth and pose prediction during colonoscopy. The challenge was hosted as part of MICCAI 2022 in Singapore. Six teams from around the world and representatives from academia and industry participated in the three sub-challenges: synthetic depth prediction, synthetic pose prediction, and real pose prediction. This paper describes the challenge, the submitted methods, and their results. We show that depth prediction from synthetic colonoscopy images is robustly solvable, while pose estimation remains an open research question.
It has recently been demonstrated that pretraining backbones in a self-supervised manner generally provides better fine-tuned polyp segmentation performance, and that models with ViT-B backbones typically perform better than models with ResNet50 backbones. In this paper, we extend this recent work to consider generalisability. I.e., we assess the performance of models on a different dataset to that used for fine-tuning, accounting for variation in network architecture and pretraining pipeline (algorithm and dataset). This reveals how well models with different pretrained backbones generalise to data of a somewhat different distribution to the training data, which will likely arise in deployment due to different cameras and demographics of patients, amongst other factors. We observe that the previous findings, regarding pretraining pipelines for polyp segmentation, hold true when considering generalisability. However, our results imply that models with ResNet50 backbones typically generalise better, despite being outperformed by models with ViT-B backbones in evaluation on the test set from the same dataset used for fine-tuning.
Amongst all future developments it is the electrification of heat that is anticipated to have the largest impact on seasonal and interannual electricity demand. There is therefore a need to accurately quantify and assess this impact. This paper uses a combination of existing advanced techniques to modify the historic electricity demand to incorporate the impact of heat pumps alone for long-term historic weather data using Great Britain as an example. The methods for generating time series were compared and extensively validated. This includes comparisons with measured data that have not been used previously for this purpose. The research reveals that for predicted 2050 heat pump penetration levels the monthly demand for electricity doubles in winter. This leads to an increase of approximately 30 TWh for each winter month and a 37 % increase in year-to-year variability of electricity demand due to weather. Peak electricity demand is very sensitive to the method of generating heat demand and the assumptions on hourly heat pump operating profiles, suggesting inaccuracies of 25 % in esti-mates of future peak demand. This work, rather than just assessing the impact of projected changes provides a reference case for policy makers to guide the decision process and planning for future scenarios.
This paper presents a novel development of a real-time evaluation system for facial asymmetry. Using this tool, it is expected possibly to assess the severity of facial asymmetry based on local and global comparisons of facial components. While the local one measures distances between the landmarks inside each facial region, the global one focuses on the overall difference of all facial regions about the sagittal plane. This reported preliminary work is focused on assessing the suitability of the existing image landmark detection methods for a real-time evaluation. Three commonly used deep learning-based methodologies have been implemented and tested in order to identify robust facial landmark detection under various challenging conditions. It is hypothesized, that the proposed system will be able to provide a more accurate measurement of facial asymmetry. The key is the proposed use of the geodesic distance calculated based on the geometry of human faces, with the help of the state-of-art depth camera.
Most vision-based 3D pose estimation approaches typically rely on knowledge of object’s 3D model, depth measurements, and often require time-consuming iterative refinement to improve accuracy. However, these can be seen as limiting factors for broader real-life applications. The main motivation for this paper is to address these limitations. To solve this, a novel Convolutional Variational Auto-Encoder based Multi-Level Network for object 3D pose estimation (CVML-Pose) method is proposed. Unlike most other methods, the proposed CVML-Pose implicitly learns an object’s 3D pose from only RGB images encoded in its latent space without knowing the object’s 3D model, depth information, or performing a post-refinement. CVML-Pose consists of two main modules: (i) CVML-AE representing convolutional variational autoencoder, whose role is to extract features from RGB images, (ii) Multi-Layer Perceptron and K-Nearest Neighbor regressors mapping the latent variables to object 3D pose including, respectively, rotation and translation. The proposed CVML-Pose has been evaluated on the LineMod and LineMod-Occlusion benchmark datasets. It has been shown to outperform other methods based on latent representations and achieves comparable results to the state-of-the-art, but without use of a 3D model or depth measurements. Utilizing the t-Distributed Stochastic Neighbor Embedding algorithm, the CVML-Pose latent space is shown to successfully represent objects’ category and topology. This opens up a prospect of integrated estimation of pose and other attributes (possibly also including surface finish or shape variations), which, with real-time processing due to the absence of iterative refinement, can facilitate various robotic applications. Code available: https://github.com/JZhao12/CVML-Pose.
For computer vision systems based on artificial neural networks, increasing the resolution of images typically improves the performance of the network. However, ImageNet pre-trained Vision Transformer (ViT) models are typically only openly available for 2242 and 3842 image resolutions. To determine the impact of using higher resolution images with ViT systems the performance differences between ViT-B/16 models (designed for 3842 and 5442 image resolutions) were evaluated. The multi-label classification RANZCR CLiP challenge dataset, which contains over 30,000 high resolution labelled chest X-ray images, was used throughout this investigation. The performance of the ViT 3842 and ViT 5442 models with no ImageNet pre-training (i.e. models were only trained using RANZCR data) was firstly compared to see if using higher resolution images increases performance. After this, a multi-resolution fine-tuning approach was investigated for transfer learning. This approach was achieved by transferring learned parameters from ImageNet pre-trained ViT 3842 models, which had undergone further training on the 3842 RANZCR data, to ViT 5442 models which were then trained on the 5442 RANZCR data. Learned parameters were transferred via a tensor slice copying technique. The results obtained provide evidence that using larger image resolutions positively impacts ViT network performance and that multi-resolution fine-tuning can lead to performance gains. The multi-resolution fine-tuning approach used in this investigation could potentially improve the performance of other computer vision systems which use ViT based networks. The results of this investigation may also warrant the development of new ViT variants optimized to work with high resolution image datasets.
Colonoscopy is widely recognised as the gold standard procedure for the early detection of colorectal cancer (CRC). Segmentation is valuable for two significant clinical applications, namely lesion detection and classification, providing means to improve accuracy and robustness. The manual segmentation of polyps in colonoscopy images is time-consuming. As a result, the use of deep learning (DL) for automation of polyp segmentation has become important. However, DL-based solutions can be vulnerable to overfitting and the resulting inability to generalise to images captured by different colonoscopes. Recent transformer-based architectures for semantic segmentation both achieve higher performance and generalise better than alternatives, however typically predict a segmentation map of h/4×w/4 spatial dimensions for a h× w input image. To this end, we propose a new architecture for full-size segmentation which leverages the strengths of a transformer in extracting the most important features for segmentation in a primary branch, while compensating for its limitations in full-size prediction with a secondary fully convolutional branch. The resulting features from both branches are then fused for final prediction of a h× w segmentation map. We demonstrate our method’s state-of-the-art performance with respect to the mDice, mIoU, mPrecision, and mRecall metrics, on both the Kvasir-SEG and CVC-ClinicDB dataset benchmarks. Additionally, we train the model on each of these datasets and evaluate on the other to demonstrate its superior generalisation performance. Code available: https://github.com/CVML-UCLan/FCBFormer .
In this paper, we introduce mREAL-GAN, a generative adversarial network (GAN) for the parallel generation of multiple residential electrical appliance load (mREAL) profiles. mREAL-GAN is intended for use in community-scale low-voltage network analysis, and represents a departure from previous methods for this purpose, which break the generation of appliance load profiles into several steps and largely model each appliance independently. Instead, mREAL-GAN models appliance load profiles in an end-to-end manner, and generates multiple appliance load profiles in parallel in a way that captures inter-dependencies. We show that mREAL-GAN generates load profiles for individual appliance-types with greater fidelity than a popular example of previous methods, and demonstrate its ability to capture inter-dependencies between appliances.
The Endoscopy Computer Vision Challenge (EndoCV) is a crowd-sourcing initiative to address eminent problems in developing reliable computer aided detection and diagnosis endoscopy systems and suggest a pathway for clinical translation of technologies. Whilst endoscopy is a widely used diagnostic and treatment tool for hollow-organs, there are several core challenges often faced by endoscopists, mainly: 1) presence of multi-class artefacts that hinder their visual interpretation, and 2) difficulty in identifying subtle precancerous precursors and cancer abnormalities. Artefacts often affect the robustness of deep learning methods applied to the gastrointestinal tract organs as they can be confused with tissue of interest. EndoCV2020 challenges are designed to address research questions in these remits. In this paper, we present a summary of methods developed by the top 17 teams and provide an objective comparison of state-of-the-art methods and methods designed by the participants for two sub-challenges: i) artefact detection and segmentation (EAD2020), and ii) disease detection and segmentation (EDD2020). Multi-center, multi-organ, multi-class, and multi-modal clinical endoscopy datasets were compiled for both EAD2020 and EDD2020 sub-challenges. The out-of-sample generalization ability of detection algorithms was also evaluated. Whilst most teams focused on accuracy improvements, only a few methods hold credibility for clinical usability. The best performing teams provided solutions to tackle class imbalance, and variabilities in size, origin, modality and occurrences by exploring data augmentation, data fusion, and optimal class thresholding techniques.
National heat demand time series are important inputs into national energy system models. Although time series for primary fuel such as gas might be available, heat demand is not and measuring heat demand is only possible for individual buildings. Four different methods are used in this work to generate daily heat demand time series for Great Britain for 2016–2018 from temperature and windspeed and are validated against heat demand derived from national grid gas demand. All seem to model heat demand well.
Automated analysis of endoscopic images is becoming increasingly significant for an early detection of numerous cancers and minimally invasive surgical procedures. The paper briefly describes the methodology adopted for the 2020 Endoscopy Artefact Detection and Segmentation (EAD2020) challenge1. A number of novel variants of the DeepLab V3+\r\nencoder-decoder architecture have been investigated, implemented and tested for the segmentation sub-challenge. Modifications were introduced to improve: selection of image futures, segmentation of small objects, and use of the encoder output information. The proposed methods achieved competitive segmentation score results on both release-I and releaseII test datasets. For the detection sub-challenge three off-theshelf deep detection networks have been optimised and evaluated on the EAD data.
This paper reports on a new CT volume registration method, using 3D Convolutional Neural Networks (CNN). The proposed method uses the Least Square Generative Adversarial Network (LSGAN) model consisting of the Contraction-Expansion registration network as the LSGAN’s generator and a deep 3D CNN classification network as the LSGAN’s discriminator. The training of the generator is performed first on its own, using Charbonnier and smoothness loss functions, with progressive weights update moving from lower to higher resolution layers of the Expander. Subsequently, the complete network (Contraction-Expansion with the Discriminator) is trained as a LSGAN network. For the training, CREATIS and COPDgene datasets have been used in a self-supervised paradigm, using 3D warping of the moving volume to estimate the error with respect to the reference volume. The input to the network has 256 × 256 × 128 × 2 voxels and the output is displacement field of 128 × 128 × 64 × 3 voxels. The Contraction-Expansion registration network, on its own, achieves mean error of 1.30 mm with 1.70 standard deviation (SD) on the DIR-LAB dataset. When the whole proposed LSGAN network is used, the mean error is further reduced to 1.13 mm with 0.67 (SD). Therefore, the use of the GAN paradigm reduces the mean error by approximately 15
This paper presents a comparison of bottom up models that generate appliance load profiles. The comparison is based on their ability to accurately distribute load over time-of-day. This is a key feature of model performance if the model is used to assess the impact of low carbon technologies and practices on the network. No work has yet assessed models on this basis. In this work, the temporal characteristics of load are captured using histograms, and similarity between the histogram representations of measured and generated data is assessed using the Wasserstein distance. This is then applied to compare the results of three models, which were developed here by adopting approaches used in previous research. One is based on occupant presence, one on occupant activity, and one on empirical data. Typical statistical tests showed that the comparison method is robust and can be used for this purpose.
Analyses of polyp images play an important role in an early detection of colorectal cancer. An automated polyp segmentation is seen as one of the methods that could improve the accuracy of the colonoscopic examination. The paper describes evaluation study of a segmentation method developed for the Endoscopic Vision Gastrointestinal Image ANAlysis – (GIANA) polyp segmentation sub-challenges. The proposed polyp segmentation algorithm is based on a fully convolutional network (FCN) model. The paper describes cross-validation results on the training GIANA dataset. Various tests have been evaluated, including network configuration, effects of data augmentation, and performance of the method as a function of polyp characteristics. The proposed method delivers state-of-the-art results. It secured the first place for the image segmentation tasks at the 2017 GIANA challenge and the second place for the SD images at the 2018 GIANA challenge.