Background and Objective: The classification of bone marrow (BM) cells by light microscopy is an important cornerstone of hematological diagnosis, performed thousands of times a day by highly trained specialists in laboratories worldwide. As the manual evaluation of blood or BM smears is very time-consuming and prone to inter-observer variation, new reliable automated systems are needed. Methods: We aim to improve the automatic classification performance of hematological cell types. Therefore, we evaluate four state-of-the-art Convolutional Neural Network (CNN) architectures on a dataset of 171, 374 microscopic cytological single-cell images obtained from BM smears from 945 patients diagnosed with a variety of hematological diseases. We further evaluate the effect of an in-domain vs. out-of-domain pre-training, and assess whether class activation maps provide human-interpretable explanations for the models' predictions. Results: The best performing pre-trained model (Regnet_y_32gf) yields a mean precision, recall, and F1 scores of 0.787 +/- 0.060, 0.755 +/- 0.061, and 0.762 +/- 0.050, respectively. This is a 53.5% improvement in precision and 7.3% improvement in recall over previous results with CNNs (ResNeXt-50) that were trained from scratch. The out-of-domain pre-training apparently yields general feature extractors/filters that apply very well to the BM cell classification use case. The class activation maps on cell types with characteristic morphological features were found to be consistent with the explanations of a human domain expert. For example, the Auer rods in the cytoplasm were the predictive cellular feature for correctly classified images of faggot cells. Conclusions: Our study provides data that can help hematology laboratories to choose the optimal training strategy for blood cell classification deep learning models to improve computer-assisted blood and bone marrow cell identification. It also highlights the need for more specific training data, i.e. images of difficult-to-classify classes, including cells labeled with disease information.
The development of data science, the increase of computational power, the availability of the internet infrastructure for data exchange and the urgency for an understanding of complex systems require a responsible and ethical use of computational models in science, communication and decision-making. Starting with a discussion of the width of different purposes of computational models, we first investigate the process of model construction as an interplay of theory and experimentation. We emphasise the different aspects of the tension between model variables and experimentally measurable observables. The resolution of this tension is a prerequisite for the responsible use of models and an instrumental part of using models in the scientific processes. We then discuss the impact of models and the responsibility that results from the fact that models support and may also guide experimentation. Further, we investigate the difference between computational modelling in an interdisciplinary science project and computational models as tools in transdisciplinary decision support. We regard the communication of model structures and modelling results as essential; however, this communication cannot happen in a technical manner, but model structures and modelling results must be translated into a “narrative.” We discuss the role of concepts from disciplines such as literary theory, communication science, and cultural studies and the potential gains that a broader approach can obtain. Considering concepts from the liberal arts, we conclude that there is, besides the responsibility of the model author, also a responsibility of the user/reader of the modelling results.
We present a model for the spread, transmission and competition of skills with an emphasis on the role of spatial mobility of individuals. From a methodological point of view, we seek mathematical and computational simplicity in the sense of a minimal model. This minimalism lets us use a infinite dimensional simplex space and not a Euclidean space as underlying structure. Such a simplex captures the essentials of spatial heterogeneity without the mathematical difficulties of neighborhood structures. In the presented model, individuals may have no skill or either skill A or B. Individuals are born unskilled and may acquire skills by learning from a skilled individual. Skill A results in a small reproductive advantage and is easy to transmit (teaching happens at high rate), whereas skill B is harder to teach but results in a high benefit. The model exhibits a rich behavior; after an initial transient, the system settles to a fix point (constant distribution of skills), whereby the distribution of skills depends on a mobility parameter m. We observe different regimes, and as the main result, we conclude that for some settings of the system parameters, the spread of the (harder to learn but more beneficial) skill B is only possible within a specific range of the mobility parameter. From a technical point of view, this paper presents the application of the PRESS-method (probability reduced evolution of spatially resolved species) that enables the study of spatial effects in a very efficient manner. We analyze the consequences of spatial organization and argue that we can study aspects of social dynamics in an infinite dimensional simplex space. In spite of this maybe daunting name, the dynamics on such a structure is comparably easy to implement. The model we present is far from reflecting all the details of human interaction. On the contrary, we deliberately tailored the model to be as simple as possible from a mathematical point of view (but still reflecting central properties of spatial organization). This approach is guided by physics, where seemingly simple models which obviously don't reflect the true physical behavior of a system (such as the Ising model) are nevertheless suited to reveal fundamental aspects and limiting cases of the real world.