Chronic obstructive pulmonary disease (COPD) is the third leading cause of death worldwide, and emphysema is present in the majority of affected patients and can be identified on computed tomography (CT). This study investigated whether radiomic features derived from automatically and adaptively segmented low-attenuation lung regions can capture distinct imaging characteristics of COPD beyond conventional emphysema measures. Radiomic features were extracted from 6078 chest CT scans of 2243 participants from the COPDGene cohort. Emphysematous regions were segmented using the MimSeg method based on Gaussian mixture modelling with patient-adjusted thresholding, and radiomic features were computed for individual lesion clusters and aggregated per patient using summary statistics, yielding 780 features per subject. Uniform Manifold Approximation and Projection (UMAP) was used to generate a low-dimensional embedding, and feature contributions were evaluated using SHAP analysis and statistical testing. The resulting embedding demonstrated structured patterns broadly aligned with Global Initiative for Chronic Obstructive Lung Disease (GOLD) stages, with greater overlap among GOLD 0–2 and more consolidated groupings for GOLD 3 and 4, reflecting differences in disease severity. The most influential features were predominantly derived from Grey Level Run Length Matrix measures, capturing textural heterogeneity and spatial organisation of emphysematous changes that are not directly described by standard density-based metrics. These findings suggest that radiomic analysis of adaptively segmented CT data may provide complementary and structurally distinct information relative to conventional emphysema measures, supporting a more nuanced characterisation of emphysema patterns in COPD.
Background Lung cancer remains the deadliest cancer worldwide because it is often diagnosed too late. Effective treatment depends on detection at an early screening stage. However, the growing number of patients and the limited number of radiologists lead to prolonged diagnostic waiting times. In very early stage lung cancer, nodule visibility is further reduced by adjacent blood vessels and airway walls, because nodules are often connected to or supplied by these structures. Task-specific analysis of the bronchovascular bundle is therefore important for efficient nodule detection, and its removal can increase the diagnostic potential of lung cancer screening. Materials and Methods To assess the efficacy of the proposed method, we used series from widely utilized LDCT datasets, including the Duke Lung Cancer Screening (DLCS) dataset and the Pilot Pomeranian Lung Cancer Screening Program. The proposed bronchovascular bundle segmentation pipeline, RONALD, operates on computed tomography images and returns binary masks of vessels and bronchi located in the lung parenchyma. The method includes a preprocessing stage with lung, lobe, and mediastinum segmentation, followed by separate vessel and bronchial tree segmentation. Results The proposed pipeline segmented the bronchovascular bundle in low-dose computed tomography scans while improving nodule retention compared with other segmentation methods: from 93.98 Conclusion The resulting segmentations can improve lung nodule detection in the very early stages of lung cancer.
Quantitative assessment of muscle MRI is crucial for monitoring neuromuscular disorders (NMD). This study introduces an automated radiomic phenotyping framework based on original features engineered across five main architectural domains: quantitative morphometry, spatial distribution, geometric shape, interactions between progressive fat replacement stages, and graph-based topology. Utilizing 1184 MRI scans from the CoMPaSS-NMD project, we map the complex 3D architecture of heterogeneous intramuscular lipodegeneration into objective, morphologically interpretable biomarkers. We introduce a graph-based skeletonization of fat infiltrates to quantify muscle architectural changes, establishing a multi-dimensional extension of traditional, spatially-agnostic volume metrics by mapping topological networks across the entire 3D muscle volume. Statistical screening via non-parametric Kruskal-Wallis analysis confirmed the discriminative power of these novel descriptors across the genetic hierarchy. Notably, topological network metrics (e.g., SF1_Skel_Nodes, ε^2 = 0.2656) and interface dynamics metrics (e.g., SF2_To_SF1_Dist_Min, ε^2 = 0.2092) demonstrated substantial effect sizes, providing deeper structural insights than classical volumetric assessments. Post-hoc pairwise evaluations and UMAP projections further indicated the capability of these topological and 3D geometric invariants to capture disease-specific macroscopic infiltration patterns. These results demonstrate that global architectural features represent a highly promising class of biomarkers for differential diagnosis, offering new avenues for tracking longitudinal disease dynamics in neuromuscular diagnostics. The developed automated feature extraction pipeline is integrated and available within the MUSCAT (MUSCle fAt Topology) library.
GRAP-MOT is a new approach for solving the person MOT problem dedicated to videos of closed areas with overlapping multi-camera views, where person occlusion frequently occurs. Our novel graph-weighted solution updates a person's identification label online based on tracks and the person's characteristic features. To find the best solution, we deeply investigated all elements of the MOT process, including feature extraction, tracking, and community search. Furthermore, GRAP-MOT is equipped with a person's position estimation module, which gives additional key information to the MOT method, ensuring better results than methods without position data. We tested GRAP-MOT on recordings acquired in a closed-area model and on publicly available real datasets that fulfil the requirement of a highly congested space, showing the superiority of our proposition. Finally, we analyzed existing metrics used to compare MOT algorithms and concluded that IDF1 is more adequate than MOTA in such comparisons. We made our code, along with the acquired dataset, publicly available.
The advancement of deep learning methods across various applications has forced the creation of enormous training datasets. However, obtaining suitable real-world datasets is often challenging for various reasons. Consequently, numerous studies have emerged focusing on the generation and utilization of synthetic data in the training process. Hence, there is no universal formula for preparing synthetic data and leveraging it in network training to maximize the effectiveness of various detection methods. This work provides a comprehensive overview of several synthetic data generation techniques, followed by a thorough investigation into the impact of training methods and the selection of synthetic data quantities. The outcomes of this research enable the formulation of conclusions regarding the recipe for developing synthetic data with high efficacy in enhancing detection methods. The main conclusion for the synthetic data generation methods is to ensure maximum diversity at a high level of photorealism, which allows improving the classification quality by more than 5% to even 19% for different detection metrics.
This study proposes a methodology for the automated localisation of lower limb bones on T1 weighted MRI scans, employing a deep learning (DL) approach. The primary objective is to facilitate precise identification of skeletal structures, thereby supporting radiomics based diagnostics of neuromuscular disorders. The developed framework is not confined to the recognition of lower limb bones. A dataset of 1,243 MRI scans was used, with a subset of 29 manually labelled bone segmentations of six key lower limb bone classes. Axial slices were divided into training (2,283), validation (300), and hold-out (378) sets. A two part segmentation pipeline was developed using a combination of U-Net and ResNet architectures, with a custom cost function to handle variable label presence across slices. A novel method for obtaining precise bone-related slice location on the MRI volume was developed. Segmentation quality for Tibia and Femur was high, achieving 86.04% and 86.97% Dice score on the hold-out subset. The ResNet classifier correctly identified the defined regions on the volume, achieving AUCs over 97% on the hold-out subset for most leg fragments except for the knee label. This work introduces a new method for anatomical spatial localisation estimation in MRI scans. Unlike previous studies, which could only recognise body parts, the proposed method also estimates the precise bone-related slice location within the scanned volume, providing an added layer of anatomical context. The developed pipeline can be retrained for other body fragments. This solution exhibits a strong potential for use in clinical workflows, especially for studies involving musculoskeletal diseases.
Chest X-rays (CXRs) are widely used for diagnosing respiratory diseases, including the recent example of COVID-19. Supervised deep learning techniques can help detect cases faster and monitor disease progression. However, they are usually developed using coarser data annotations, which may insufficiently capture the heterogeneous disease portrait. We propose the pipeline called CIRCA ( https://circa.aei.polsl.pl ) for a CXR-based screening support system, developed using 6 diverse datasets. Our tool includes lung segmentation, quantitative assessment of data heterogeneity, and a hierarchical three-class decision system using a convolutional network and radiomic features. Lung segmentation showed an accuracy of ~ 94% in the validation and test sets, while classification accuracy was equal 86%, 83%, and 72% for normal, COVID-19, and other pneumonia classes in the independent test set. Three radiomically distinct subtypes were identified per class. In the hold-out set, the classification subtype-specific cross-dataset NPV ranged from 95 to 100%, with PPV from 86 to 100% for all subtypes except N3 (early stage or convalescent) and both C3 and P3 (probable co-occurrence of COVID-19). Using an independent test set gave similar results. The dataset-specific subtype proportions combined with various predictive qualities of subtypes partly explain the widely reported poor generalization of AI-based prediction systems.
The application of image analysis methods to calculate the distance from the camera to the object allows the replacement of specialized hardware devices for distance estimation. In the case of public transport, estimation of the exact position of the passenger gives also the option to determine whether the passenger is inside or outside the vehicle. In the presented work, several distance estimation methods based on typical analytical models and machine learning (ML) methods were tested using recordings from three cameras located in the minibus model. Human head detection was used instead of the entire passenger body to avoid occlusion problems. The analytical method showed worse performance than ML methods in all cases. The difference in the performance of ML models between cameras was negligible and there was no best method found. The computational time for ML models ranges from 0.35 to 100.57 ms, which should result in successful real-world applications. The developed approach can be used not only in public transport but also in all closed areas for the calculation of people or crowd density.
The outbreak of the SARS-CoV-2 pandemic has put healthcare systems worldwide to their limits, resulting in increased waiting time for diagnosis and required medical assistance. With chest radiographs (CXR) being one of the most common COVID-19 diagnosis methods, many artificial intelligence tools for image-based COVID-19 detection have been developed, often trained on a small number of images from COVID-19-positive patients. Thus, the need for high-quality and well-annotated CXR image databases increased. This paper introduces POLCOVID dataset, containing chest X-ray (CXR) images of patients with COVID-19 or other-type pneumonia, and healthy individuals gathered from 15 Polish hospitals. The original radiographs are accompanied by the preprocessed images limited to the lung area and the corresponding lung masks obtained with the segmentation model. Moreover, the manually created lung masks are provided for a part of POLCOVID dataset and the other four publicly available CXR image collections. POLCOVID dataset can help in pneumonia or COVID-19 diagnosis, while the set of matched images and lung masks may serve for the development of lung segmentation solutions.
Due to its predominantly asymptomatic or mildly symptomatic progression, lung cancer is often diagnosed in advanced stages, resulting in poorer survival rates for patients. As with other cancers, early detection significantly improves the chances of successful treatment. Early diagnosis can be facilitated through screening programs designed to detect lung tissue tumors when they are still small, typically around 3mm in size. However, the analysis of extensive screening program data is hampered by limited access to medical experts. In this study, we developed a procedure for identifying potential malignant neoplastic lesions within lung parenchyma. The system leverages machine learning (ML) techniques applied to two types of measurements: low-dose Computed Tomography-based radiomics and metabolomics. Using data from two Polish screening programs, two ML algorithms were tested, along with various integration methods, to create a final model that combines both modalities to support lung cancer screening.
When the COVID-19 pandemic commenced in 2020, scientists assisted medical specialists with diagnostic algorithm development. One scientific research area related to COVID-19 diagnosis was medical imaging and its potential to support molecular tests. Unfortunately, several systems reported high accuracy in development but did not fare well in clinical application. The reason was poor generalization, a long-standing issue in AI development. Researchers found many causes of this issue and decided to refer to them as confounders, meaning a set of artefacts and methodological errors associated with the method. We aim to contribute to this steed by highlighting an undiscussed confounder related to image resolution. 20 216 chest X-ray images (CXR) from worldwide centres were analyzed. The CXRs were bijectively projected into the 2D domain by performing Uniform Manifold Approximation and Projection (UMAP) embedding on the radiomic features (rUMAP) or CNN-based neural features (nUMAP) from the pre-last layer of the pre-trained classification neural network. Additional 44 339 thorax CXRs were used for validation. The comprehensive analysis of the multimodality of the density distribution in rUMAP/nUMAP domains and its relation to the original image properties was used to identify the main confounders. nUMAP revealed a hidden bias of neural networks towards the image resolution, which the regular up-sampling procedure cannot compensate for. The issue appears regardless of the network architecture and is not observed in a high-resolution dataset. The impact of the resolution heterogeneity can be partially diminished by applying advanced deep-learning-based super-resolution networks. rUMAP and nUMAP are great tools for image homogeneity analysis and bias discovery, as demonstrated by applying them to COVID-19 image data. Nonetheless, nUMAP could be applied to any type of data for which a deep neural network could be constructed. Advanced image super-resolution solutions are needed to reduce the impact of the resolution diversity on the classification network decision.
Segmentation of the bronchovascular bundle within the lung parenchyma is a key step for the proper analysis and planning of many pulmonary diseases. It might also be considered the preprocessing step when the goal is to segment the nodules from the lung parenchyma. We propose a segmentation pipeline for the bronchovascular bundle based on the Computed Tomography images, returning either binary or labelled masks of vessels and bronchi situated in the lung parenchyma. The method consists of two modules, modeling of the bronchial tree and vessels. The core revolves around a similar pipeline, the determination of the initial perimeter by the GMM method, skeletonization, and hierarchical analysis of the created graph. We tested our method on both low-dose CT and standard-dose CT, with various pathologies, reconstructed with various slice thicknesses, and acquired from various machines. We conclude that the method is invariant with respect to the origin and parameters of the CT series. Our pipeline is best suited for studies with healthy patients, patients with lung nodules, and patients with emphysema.
Due to the large accumulation of patients requiring hospitalization, the COVID-19 pandemic disease caused a high overload of health systems, even in developed countries. Deep learning techniques based on medical imaging data can help in the faster detection of COVID-19 cases and monitoring of disease progression. Regardless of the numerous proposed solutions for lung X-rays, none of them is a product that can be used in the clinic. Five different datasets (POLCOVID, AIforCOVID, COVIDx, NIH, and artificially generated data) were used to construct a representative dataset of 23 799 CXRs for model training; 1 050 images were used as a hold-out test set, and 44 247 as independent test set (BIMCV database). A U-Net-based model was developed to identify a clinically relevant region of the CXR. Each image class (normal, pneumonia, and COVID-19) was divided into 3 subtypes using a 2D Gaussian mixture model. A decision tree was used to aggregate predictions from the InceptionV3 network based on processed CXRs and a dense neural network on radiomic features. The lung segmentation model gave the Sorensen-Dice coefficient of 94.86% in the validation dataset, and 93.36% in the testing dataset. In 5-fold cross-validation, the accuracy for all classes ranged from 91% to 93%, keeping slightly higher specificity than sensitivity and NPV than PPV. In the hold-out test set, the balanced accuracy ranged between 68% and 100%. The highest performance was obtained for the subtypes N1, P1, and C1. A similar performance was obtained on the independent dataset for normal and COVID-19 class subtypes. Seventy-six percent of COVID-19 patients wrongly classified as normal cases were annotated by radiologists as with no signs of disease. Finally, we developed the online service (https://circa.aei.polsl.pl) to provide access to fast diagnosis support tools.