Using unmanned aerial vehicles (UAVs) to track multiple individuals simultaneously in their natural environment is a powerful approach for better understanding the collective behavior of primates. Previous studies have demonstrated the feasibility of automating primate behavior classification from video data, but these studies have been carried out in captivity or from ground-based cameras. However, to understand group behavior and the self-organization of a collective, the whole troop needs to be seen at a scale where behavior can be seen in relation to the natural environment in which ecological decisions are made. To tackle this challenge, this study presents a novel dataset for baboon detection, tracking, and behavior recognition from drone videos where troops are observed on-the-move in their natural environment as they move to and from their sleeping sites. Videos were captured from drones at Mpala Research Centre, a research station located in Laikipia County, in central Kenya. The baboon detection dataset was created by manually annotating all baboons in drone videos with bounding boxes. A tiling method was subsequently applied to create a pyramid of images at various scales from the original 5.3K resolution images, resulting in approximately 30K images used for baboon detection. The baboon tracking dataset is derived from the baboon detection dataset, where bounding boxes are consistently assigned the same ID throughout the video. This process resulted in half an hour of dense tracking data. The baboon behavior recognition dataset was generated by converting tracks into mini-scenes, a video subregion centered on each animal. These mini-scenes were annotated with 12 distinct behavior types and one additional category for occlusion, resulting in over 20 hours of data. Benchmark results show mean average precision (mAP) of 92.62 https://baboonland.xyz .
We present a novel dataset for animal behavior recognition collected in-situ using video from drones flown over the Mpala Research Centre in Kenya. Videos from DJI Mavic 2S drones flown in January 2023 were acquired at 5.4K resolution in accordance with IACUC protocols, and processed to detect and track each animal in the frames. An image subregion centered on each animal was extracted and combined in sequence to form a “mini-scene”. Be-haviors were then manually labeled for each frame of each mini-scene by a team of annotators overseen by an expert behavioral ecologist. The resulting labeled mini-scenes form our resulting behavior dataset, consisting of more than 10 hours of annotated videos of reticulated gi-raffes, plains zebras, and Grevy's zebras, and encompassing seven types of animal behavior and an additional category for occlusions. Benchmark results for state-of-the-art behavioral recognition architectures show labeling accu-racy of 61.9% for macro-average (per class), and 86.7% for micro-average (per instance). Our dataset complements recent larger, more diverse animal behavior sets and smaller, more specialized ones by being collected in-situ and from drones, both important considerations for the future of an-imal behavior research. The dataset can be accessed at https://dirtmaxim.github.io/kabr.
Access to large image volumes through camera traps and crowdsourcing provides novel possibilities for animal monitoring and conservation. It calls for automatic methods for analysis, in particular, when re-identifying individual animals from the images. Most existing re-identification methods rely on either hand-crafted local features or end-to-end learning of fur pattern similarity. The former does not need labeled training data, while the latter, although very data-hungry typically outperforms the former when enough training data is available. We propose a novel re-identification pipeline that combines the strengths of both approaches by utilizing modern learnable local features and feature aggregation. This creates representative pattern feature embeddings that provide high re-identification accuracy while allowing us to apply the method to small datasets by using pre-trained feature descriptors. We report a comprehensive comparison of different modern local features and demonstrate the advantages of the proposed pipeline on two very different species.
In this paper, we extend the dataset statistics, model benchmarks, and performance analysis for the recently published KABR dataset, an in situ dataset for ungulate behavior recognition using aerial footage from the Mpala Research Centre in Kenya. The dataset comprises video footage of reticulated giraffes (lat. Giraffa reticulata), Plains zebras (lat. Equus quagga), and Grévy’s zebras (lat. Equus grevyi) captured using a DJI Mavic 2S drone. It includes both spatiotemporal (i.e., mini-scenes) and behavior annotations provided by an expert behavioral ecologist. In total, KABR has more than 10 hours of annotated video. We extend the previous work in four key areas by: (i) providing comprehensive dataset statistics to reveal new insights into the data distribution across behavior classes and species; (ii) extending the set of existing benchmark models to include a new state-of-the-art transformer; (iii) investigating weight initialization strategies and exploring whether pretraining on human action recognition datasets is transferable to in situ animal behavior recognition directly (i.e., zero-shot) or as initialization for end-to-end model training; and (iv) performing a detailed statistical analysis of the performance of these models across species, behavior, and formally defined segments of the long-tailed distribution. The KABR dataset addresses the limitations of previous datasets sourced from controlled environments, offering a more authentic representation of natural animal behaviors. This work marks a significant advancement in the automatic analysis of wildlife behavior, leveraging drone technology to overcome traditional observational challenges and enabling a more nuanced understanding of animal interactions in their natural habitats. The dataset is available at https://kabrdata.xyz
In situ imageomics leverages machine learning techniques to infer biological traits from images collected in the field, or in situ, to study individuals organisms, groups of wildlife, and whole ecosystems. Such datasets provide real-time social and environmental context to inferred biological traits, which can enable new, data-driven conservation and ecosystem management. The development of machine learning techniques to extract biological traits from images are impeded by the volume and quality data required to train these models. Autonomous, unmanned aerial vehicles (UAVs), are well suited to collect in situ imageomics data as they can traverse remote terrain quickly to collect large volumes of data with greater consistency and reliability compared to manually piloted UAV missions. However, little guidance exists on optimizing autonomous UAV missions for the purposes of remote sensing for conservation and biodiversity monitoring. The UAV video dataset curated by KABR: In-Situ Dataset for Kenyan Animal Behavior Recognition from Drone Videos required three weeks to collect, a time-consuming and expensive endeavor. Our analysis of KABR revealed that a third of the videos gathered were unusable for the purposes of inferring wildlife behavior. We analyzed the flight telemetry data from portions of UAV videos that were usable for inferring wildlife behavior, and demonstrate how these insights can be integrated into an autonomous remote sensing system to track wildlife in real time. Our autonomous remote sensing system optimizes the UAV's actions to increase the yield of usable data, and matches the flight path of an expert pilot with an 87 representing an 18.2
In situ imageomics is a new approach to study ecological, biological and evolutionary systems wherein large image and video data sets are captured in the wild and machine learning methods are used to infer biological traits of individual organisms, animal social groups, species, and even whole ecosystems. Monitoring biological traits over large spaces and long periods of time could enable new, data-driven approaches to wildlife conservation, biodiversity, and sustainable ecosystem management. However, to accurately infer biological traits, machine learning methods for images require voluminous and high quality data. Adaptive, data-driven approaches are hamstrung by the speed at which data can be captured and processed. Camera traps and unmanned aerial vehicles (UAVs) produce voluminous data, but lose track of individuals over large areas, fail to capture social dynamics, and waste time and storage on images with poor lighting and view angles. In this vision paper, we make the case for a research agenda for in situ imageomics that depends on significant advances in autonomic and self-aware computing. Precisely, we seek autonomous data collection that manages camera angles, aircraft positioning, conflicting actions for multiple traits of interest, energy availability, and cost factors. Given the tools to detect object and identify individuals, we propose a research challenge: Which optimization model should the data collection system employ to accurately identify, characterize, and draw inferences from biological traits while respecting a budget? Using zebra and giraffe behavioral data collected over three weeks at the Mpala Research Centre in Laikipia County, Kenya, we quantify the volume and quality of data collected using existing approaches. Our proposed autonomic navigation policy for in situ imageomics collection has an F1 score of 82% compared to an expert pilot, and provides greater safety and consistency, suggesting great potential for state-of-the-art autonomic approaches if they can be scaled up to fully address the problem.
Around 60-80% of radiological errors are attributed to overlooked abnormalities, the rate of which increases at the end of work shifts. In this study, we run an experiment to investigate if artificial intelligence (AI) can assist in detecting radiologists' gaze patterns that correlate with fatigue. A retrospective database of lung X-ray images with the reference diagnoses was used. The X-ray images were acquired from 400 subjects with a mean age of 49 ± 17, and 61% men. Four practicing radiologists read these images while their eye movements were recorded. The radiologists passed a series of concentration tests at prearranged breaks of the experiment. A U-Net neural network was adapted to annotate lung anatomy on X-rays and calculate coverage and information gain features from the radiologists' eye movements over lung fields. The lung coverage, information gain, and eye tracker-based features were compared with the cumulative work done (CDW) label for each radiologist. The gaze-traveled distance, X-ray coverage, and lung coverage statistically significantly (p < 0.01) deteriorated with cumulative work done (CWD) for three out of four radiologists. The reading time and information gain over lungs statistically significantly deteriorated for all four radiologists. We discovered a novel AI-based metric blending reading time, speed, and organ coverage, which can be used to predict changes in the fatigue-related image reading patterns.
Detection of lung diseases from chest X-rays has been of great interest from the research community during the last decade. Despite the existence of large annotated public databases, computer-aided diagnostic solutions still fail on challenging rare abnormality cases. In this study, we investigated the paradigm of combining the analysis of chest X-rays and physician gaze patterns during the analysis of these X-rays to improve the computerized diagnostic accuracy. Tobii Eye Tracker 4C has been mounted to a physician workstation and his eye movements were recorded during the analysis of 400 chest X-rays in two days of work. The X-rays have been sampled from CheXpert, RSNA, and SIIM-ACR public databases labeled with 14 different pathology types. The task was formulated as a binary classification problem. A ResNet34-based neural network has been trained to map the input chest X-ray with the output physician gaze map and binary pathology label. The proposed network improved the diagnostic accuracy to 0.714 of the area under receiving operator curve (AUC) from 0.681 AUC obtained for the same ResNet34 trained to generate binary pathology labels alone. The proposed study has demonstrated the potential benefits of using gaze information in computerized diagnostic solutions.
Radiologist-AI interaction is a novel area of research of potentially great impact. It has been observed in the literature that the radiologists' performance deteriorates towards the shift ends and there is a visual change in their gaze patterns. However, the quantitative features in these patterns that would be predictive of fatigue have not yet been discovered. A radiologist was recruited to read chest X-rays, while his eye movements were recorded. His fatigue was measured using the target concentration test and Stroop test having the number of analyzed X-rays being the reference fatigue metric. A framework with two convolutional neural networks based on UNet and ResNeXt50 architectures was developed for the segmentation of lung fields. This segmentation was used to analyze radiologist's gaze patterns. With a correlation coefficient of 0.82, the eye gaze features extracted lung segmentation exhibited the strongest fatigue predictive powers in contrast to alternative features.
Lung cancer constitutes more than 20% of all cancer deaths in the Russian Federation. About 34.2% of these cases were diagnosed late, which significantly reduces the life expectancy of patients. Chest X-rays are the main screening method for lung cancer in Russia. The small size of nodules and difficult localizations are the reasons why nodules are often missed during routine scanning. The aim of this study is to employ modern deep learning tools for the localization of nodules in chest X-Rays. We assembled a database of chest X-rays, where each X-ray with nodules was accompanied by a mask with Gaussian-shaped nodule heatmaps. We developed a U-net like network for the reconstruction of the nodules heatmaps from original X-rays. Considering the fact that nodules occupy an extremely small part of X-ray, and are often hidden behind clavicles or heart, we also developed a new cost function that computes the mean absolute error between nodule heatmaps and the model output giving more value to area in close proximity to the nodule. The model was applied to detect nodules from X-rays from a publicly-available JSRT database. The database contains 154 images with nodules. All training samples were heavily augmented using spatial augmentations such as random scale, rotations, flips as well as pixel-level augmentations such as blur and contrast change. We generate candidate nodule locations by extracting connected components from the network output. А 67.8% detection rate with an average of 4 false positives per correctly recognized nodule has been achieved. Our model also achieves a 7 mm error in terms of average distance, between centers of a candidate and a true nodule. Our numbers are comparable or superior to the values reported in the literature. We demonstrated the potential of deep neural networks to address an important topic of detection of lung nodules from routine X-ray scans for assisting radiologists in the diagnosis of lung cancer.
The potential of deep learning to advance lung nodule detection from chest X-rays is significantly compromised by the lack of large annotated databases and noisy labels in the existing databases. The aim of this study is to investigate the applicability of the novel Confident Learning approach for chest X-ray database cleaning and nodule detection improving. We took a subset of the NIH Chest X-ray Dataset of 14 Common Thorax Disease Categories that contains only chest X-ray images with the presence of nodules and the same amount of chest X-ray images of healthy lungs. Next, we split the obtained dataset into train and test sets. In turn, the train set was split into 4-folds to train models using a cross-validation procedure. After that, we trained an Xception (Convolutional Neural Network) model for each fold to classify chest X-ray images with nodules. We calculated probabilities for the whole train set using a cross-validation approach and evaluated the performance of trained models on the test set. To obtain noisy labels, we have to apply a family of theory and algorithms called Confident Learning with provable guarantees of exact noise estimation and label error finding. The algorithm takes noisy labels and predicted probabilities as input and returns found label errors ordered by the likelihood of being an error. We took 5% of the noisiest samples from the list provided by the algorithm and eliminated them from our train set. Then, we repeated the training pipeline but using the clean version of our train set and evaluated it on the test set. Originally, our classification pipeline gives an accuracy of 71.49%, while after applying the Confident Learning to prune noisy samples, we improved the accuracy to 72.4%. We also brought in a professional radiologist to interpret found label errors by the Confident Learning algorithm. We provided our radiologist with 100 clean chest X-ray images and asked him to classify them. In most cases, radiologist results and known dataset labels for clean X-ray images matched (72%: 36 TP, 36 TN, 18 FP, 10 FN). In this case, false-positive and false-negative predictions by the radiologist can be explained by the fact that the original dataset contains many instances where several pathologies are present in the radiogram at the same time and the radiologist most likely referred these instances to another pathology (not nodular formations). In the next experiment, we gave radiologist 100 noisy chest X-ray images found by the Confident Learning algorithm. It turned out that results obtained from the radiologist and known dataset labels for these noisy X-ray images were very different (only 39% matched: 19 TP, 20 TN, 34 FP, 27 FN). Our experiments showed that cleaning the datasets can improve the performance of deep learning algorithms on the example of detection of X-rays with lung nodules. Experiments with a radiologist showed that noisy samples are most likely incorrectly labeled or contain atypical cases of nodular formations.
Bone suppression in chest x-rays is an important processing step that can often improve visual detection of lung pathologies hidden under ribs or clavicle shadows. Current diagnostic imaging protocol does not include hardware-based bone suppression, hence the need for a software-based solution. This paper evaluates various deep learning models adapted for bone suppression task, namely, we implemented several state-of-the-art deep learning architectures: convolution autoencoder, U-net, FPN, cGAN; augmented them with domain-specific denoising techniques, such as wavelet decomposition, with the aim to identify the optimal solution for chest x-ray analysis. Our results show that wavelet decomposition does not improve the rib suppression, “skip connections” modification outperforms baseline autoencoder approach with and without the usage of the wavelet decomposition, the residual models are trained faster than plain models and achieve higher validation scores.
Pneumothorax is potentially a life-threatening disease that requires urgent diagnosis and treatment. The chest X-ray is the diagnostic modality of choice when pneumothorax is suspected. The computer-aided diagnosis of pneumothorax has received a dramatic boost in the last few years due to deep learning advances and the first public pneumothorax diagnosis competition with 15257 chest X-rays manually annotated by a team of 19 radiologists. This paper describes one of the top frameworks that participated in the competition. The framework investigates the benefits of combining the Unet convolutional neural network with various backbones, namely ResNet34, SE-ResNext50, SE-ResNext101, and DenseNet121. The paper presents a step-by-step instruction for the framework application, including data augmentation, and different pre- and post-processing steps. The performance of the framework was of 0.8574 measured in terms of the Dice coefficient. The second contribution of the paper is the comparison of the deep learning framework against three experienced radiologists on the pneumothorax detection and segmentation on challenging X-rays. We also evaluated how diagnostic confidence of radiologists affects the accuracy of the diagnosis and observed that the deep learning framework and radiologists find the same X-rays to be easy/difficult to analyze (p-value <1e4). Finally, the methodology of all top-performing teams from the competition leaderboard was analyzed to find the consistent methodological patterns of accurate pneumothorax detection and segmentation.
PURPOSE:Segmentation of organs from chest X-ray images is an essential task for an accurate and reliable diagnosis of lung diseases and chest organ morphometry. In this study, we investigated the benefits of augmenting state-of-the-art deep convolutional neural networks (CNNs) for image segmentation with organ contour information and evaluated the performance of such augmentation on segmentation of lung fields, heart, and clavicles from chest X-ray images.METHODS:Three state-of-the-art CNNs were augmented, namely the UNet and LinkNet architecture with the ResNeXt feature extraction backbone, and the Tiramisu architecture with the DenseNet. All CNN architectures were trained on ground-truth segmentation masks and additionally on the corresponding contours. The contribution of such contour-based augmentation was evaluated against the contour-free architectures, and 20 existing algorithms for lung field segmentation.RESULTS:The proposed contour-aware segmentation improved the segmentation performance, and when compared against existing algorithms on the same publicly available database of 247 chest X-ray images, the UNet architecture with the ResNeXt50 encoder combined with the contour-aware approach resulted in the best overall segmentation performance, achieving a Jaccard overlap coefficient of 0.971, 0.933, and 0.903 for the lung fields, heart, and clavicles, respectively.CONCLUSION:In this study, we proposed to augment CNN architectures for CXR segmentation with organ contour information and were able to significantly improve segmentation accuracy and outperform all existing solution using a public chest X-ray database.
Pneumonia is a bacterial, viral, or fungal infection of one or both sides of the lungs that causes lung alveoli to fill up with fluid or pus, which is usually diagnosed with chest x-rays. This work investigates opportunities for applying machine learning solutions for automated detection and localization of pneumonia on chest x-ray images. We propose an ensemble of two convolutional neural networks, namely RetinaNet and Mask R-CNN for pneumonia detection and localization. We validated our solution on a recently released dataset of 26,684 images from Kaggle Pneumonia Detection Challenge and were score among the top 3% of submitted solutions. With 0.793 recall, we developed a reliable solution for automated pneumonia diagnosis and validated it on the largest clinical database publicity available to date. Some of the challenging cases were additionally examined by a team of physicians, who helped us to interpret the obtained results and confirm their practical applicability.
Low-rank tensor approximations are very promising for compression of deep neural networks. We propose a new simple and efficient iterative approach, which alternates low-rank factorization with smart rank selection and fine-tuning. We demonstrate the efficiency of our method comparing to non-iterative ones. Our approach improves the compression rate while maintaining the accuracy for a variety of tasks.
Diagnosis of lung pathologies from CXRs is one of the main tasks in modern image-based diagnosis. Automation of lung pathology diagnosis is greatly facilitated by recent developments in deep learning-based clinical decision making. The performance of deep learning solutions has the tendency to improve with the growing number of training X-rays, which can be artificially increased by augmentation of training X-rays. Commonly, different augmentation approaches are greedily applied to the available training data without investigating the necessity and actual contribution of individual augmentation. Our work aims to fill this gap in computerized lung pathology diagnosis and evaluate the contribution of different data augmentation approaches by leveraging the publicly available ChestX-ray14 dataset.
The low-rank tensor approximation is very promising for the compression of deep neural networks. We propose a new simple and efficient iterative approach, which alternates low-rank factorization with a smart rank selection and fine-tuning. We demonstrate the efficiency of our method comparing to non-iterative ones. Our approach improves the compression rate while maintaining the accuracy for a variety of tasks.
Deep convolutional neural networks contain tens of millions of parameters, making them impossible to work efficiently on embedded devices. We propose iterative approach of applying low-rank approximation to compress deep convolutional neural networks. Since classification and object detection are the most favored tasks for embedded devices, we demonstrate the effectiveness of our approach by compressing AlexNet, VGG-16, YOLOv2 and Tiny YOLO networks. Our results show the superiority of the proposed method compared to non-repetitive ones. We demonstrate higher compression ratio providing less accuracy loss.
Christopher Stewart合作论文数Department of Computer Science and Engineering The Ohio State University1