Microscopy is a key technique to visualize and understand biology. Electron microscopy (EM) facilitates the investigation of cellular ultrastructure at biomolecular resolution. Cellular EM was recently revolutionized by automation and digitalisation allowing routine capture of large areas and volumes at nanoscale resolution. Analysis, however, is hampered by the greyscale nature of electron images and their large data volume, often requiring laborious manual annotation. Here we demonstrate unsupervised and automated extraction of biomolecular assemblies in conventionally processed tissues using large-scale hyperspectral energy-dispersive X-ray (EDX) imaging. First, we discriminated biological features in the context of tissue based on selected elemental maps. Next, we designed a data-driven workflow based on dimensionality reduction and spectral mixture analysis, allowing the visualization and isolation of subcellular features with minimal manual intervention. Broad implementations of the presented methodology will accelerate the understanding of biological ultrastructure.
Instance segmentation is crucial for insightful analysis in the increasing use of large-scale electron microscopy (EM) to gain a better understanding of disease causes or progression. Instance segmentation is a more granular version of semantic segmentation, as it identifies and distinguishes individual object instances, whereas semantic segmentation only identifies object classes. In this study, we introduce a two-stage unsupervised approach called COFI, which stands for Coarse-Semantic to Fine-Instance segmentation, for the application of mitochondria segmentation in large-scale 2D EM images. In its first stage, it produces a rough region mask by clustering image patches and prompting a user to select the regions of interest. This is followed by a boundary delineation method based on the brain-inspired COSFIRE filter which is augmented by an inhibition component that makes it robust to image texture and noise. The effectiveness of the proposed COFI approach is evaluated on an EM dataset of the heart muscle of a mouse tissue, which consisted of four tiles of 16384 × 16384 pixels, containing a total of 2287 instances of mitochondria among other subcellular structures. It consistently achieved panoptic quality measures that are substantially superior to competing supervised methodologies. Besides its elevated effectiveness, the proposed COFI approach is conceptually simple and sufficiently versatile as the structure of interest is not intrinsic to the method.
Electron microscopy (EM) enables high-resolution imaging of tissues and cells based on 2D and 3D imaging techniques. Due to the laborious and time-consuming nature of manual segmentation of large-scale EM datasets, automated segmentation approaches are crucial. This review focuses on the progress of deep learning-based segmentation techniques in large-scale cellular EM throughout the last six years, during which significant progress has been made in both semantic and instance segmentation. A detailed account is given for the key datasets that contributed to the proliferation of deep learning in 2D and 3D EM segmentation. The review covers supervised, unsupervised, and self-supervised learning methods and examines how these algorithms were adapted to the task of segmenting cellular and sub-cellular structures in EM images. The special challenges posed by such images, like heterogeneity and spatial complexity, and the network architectures that overcame some of them are described. Moreover, an overview of the evaluation measures used to benchmark EM datasets in various segmentation tasks is provided. Finally, an outlook of current trends and future prospects of EM segmentation is given, especially with large-scale models and unlabeled images to learn generic features across EM datasets.
The fseval Python package allows benchmarking Feature Selection and Feature Ranking algorithms on a large scale, and facilitates the comparison of multiple algorithms in a systematic way.In particular, fseval enables users to run experiments in parallel and distributed over multiple machines, and export the results to an SQL database.The execution of an experiment can be fully determined by a configuration file, which means the experiment results can be reproduced easily, given only the configuration file.fseval has high test coverage, continuous integration, and rich documentation.The package is open source and can be installed through PyPI.
As dimensions of datasets in predictive modelling continue to grow, feature selection becomes increasingly practical. Datasets with complex feature interactions and high levels of redundancy still present a challenge to existing feature selection methods. We propose a novel framework for feature selection that relies on boosting, or sample re-weighting, to select sets of informative features in classification problems. The method uses as its basis the feature rankings derived from fast and scalable tree-boosting models, such as XGBoost. We compare the proposed method to standard feature selection algorithms on 9 benchmark datasets. We show that the proposed approach reaches higher accuracies with fewer features on most of the tested datasets, and that the selected features have lower redundancy.
Dystocia or difficult calving in cattle is detrimental to the health of the afflicted cows and has a negative economic impact on the dairy industry. The goal of this study was to create a data-driven tool for predicting the calving difficulty of non-heifer cows using input variables that are known prior to the moment of insemination. Compared to past studies, we excluded input variables that can only be known during or after insemination, such as birth weight and gestation length. This makes the model suitable for informing mating decisions that could reduce the incidence of difficult calvings or mitigate their consequences. We used a dataset consisting of 131,527 calving records of Holstein cattle, from which we derived a total of 274 phenotypic features and estimated breeding values. The distribution of classes in the dataset was 96.7 % normal calvings, and 3.3 % difficult calvings. We used a gradient boosted trees (XGBoost) as the learning model and a bagging ensemble approach to deal with the extreme class imbalance. The model achieved an average area under the ROC curve of 0.73 on unseen test data. Using feature importance analysis, we identified a number of features that have a high discriminatory value for calving difficulty, including maternal and paternal breeding values, and past phenotypic measurements of the cow.
The muscle grading of livestock is a primary component of valuation in the meat industry. In pigs, the muscularity of a live animal is traditionally estimated by visual and tactile inspection from an experienced assessor. In addition to being a time-consuming process, scoring of this kind suffers from inconsistencies inherent to the subjectivity of human assessment. On the other hand, accurate, computer-driven methods for carcass composition estimation, such as magnetic resonance imaging (MRI) and computed tomography scans (CT-scans), are expensive and cumbersome to both the animals and their handlers. In this paper, we propose a method that is fast, inexpensive, and non-invasive for estimating the muscularity of live pigs, using RGB-D computer vision and machine learning. We used morphological features extracted from the depth images of pigs to train a classifier that estimates the muscle scores that are likely to be given by a human assessor. The depth images were obtained from a Kinect v1 camera which was placed over an aisle through which the pigs passed freely. The data came from 3246 pigs, each having 20 depth images, and a muscle score from 1 to 7 (reduced later to 5 scores) assigned by an experienced assessor. The classification based on morphological features of the pig's body shape-using a gradient boosted classifier-resulted in a mean absolute error of 0.65 in tenfold cross-validation. Notably, the majority of the errors corresponded to pigs being classified as having muscle scores adjacent to the groundtruth labels given by the assessor. According to the end users of this application, the proposed approach could be used to replace expert assessors at the farm.
Domestic pigs vary in the age at which they reach slaughter weight even under the controlled conditions of modern pig farming. Early and accurate estimates of when a pig will reach slaughter weight can lead to logistic efficiency in farms. In this study, we compare four methods in predicting the age at which a pig reaches slaughter weight (120 kg). Namely, we compare the following regression tree-based ensemble methods: random forest (RF), extremely randomized trees (ET), gradient boosted machines (GBM), and XGBoost. Data from 32979 pigs is used, comprising a combination of phenotypic features and estimated breeding values (EBV). We found that the boosting ensemble methods, GBM and XGBoost, achieve lower prediction errors than the parallel ensembles methods, RF and ET. On the other hand, RF and ET have fewer parameters to tune, and perform adequately well with default parameter settings.
Muscle grading of livestock is a crucial component of valuation in the meat industry. In pigs, the muscularity of a live animal is traditionally estimated by visual and tactile inspection from an experienced assessor. In addition to being a time-consuming process, scoring of this kind suffers from inconsistencies inherent to ...
The weight of a pig and the rate of its growth are key elements in pig production. In particular, predicting future growth is extremely useful, since it can help in determining feed costs, pen space requirements, and the age at which a pig reaches a desired slaughter weight. However, making these predictions is challenging, due to the natural variation in how individual pigs grow, and the different causes of this variation. In this paper, we used machine learning, namely random forest (RF) regression, for predicting the age at which the slaughter weight of 120 kg is reached. Additionally, we used the variable importance score from RF to quantify the importance of different types of input data for that prediction. Data of 32,979 purebred Large White pigs were provided by Topigs Norsvin, consisting of phenotypic data, estimated breeding values (EBVs), along with pedigree and pedigree-genetic relationships. Moreover, we presented a 2-step data reduction procedure, based on random projections (RPs) and principal component analysis (PCA), to extract features from the pedigree and genetic similarity matrices for use as inputs in the prediction models. Our results showed that relevant phenotypic features were the most effective in predicting the output (age at 120 kg), explaining approximately 62% of its variance (i.e., R2 = 0.62). Estimated breeding value, pedigree, or pedigree-genetic features interchangeably explain 2% of additional variance when added to the phenotypic features, while explaining, respectively, 38%, 39%, and 34% of the variance when used separately.