This work proposes a new software system ASSISLT to support speech therapy for children and adults using deep learning approaches. The application offers an adjustable set of exercises recommended by a speech therapist and aims to motivate and help with regular home practice. Augmented reality is employed to lead the exercise moves and to appraise the effort. The core of the system is automatically evaluating exercises using a webcam and developing image processing and neural network methods for face, lips, teeth, and tongue detection. The pipeline is shown together with solutions of subtasks and with demonstrations of the functionality. The statistical validation of the ASSISLT is provided, comparing the performance of speech therapy specialists and the software.
Deep learning-based methods for classification and segmentation require large training sets. Generating training data is often a tedious and expensive task. In industrial applications, such as automated visual inspection of products in an assemble line, objects for classification are well defined yet labeled data are difficult to obtain. To alleviate the problem of manual labeling, we propose to train a convolutional neural network with an automatically generated training set using a naive classifier with handcrafted features. We show that when the naive classifier has high precision then the trained network has both high precision and recall despite the low recall of the naive classifier. We demonstrate the proposed methodology on real scenario of detecting a car coolant tank. However, the proposed methodology facilitates collection of train data for a wider type of CNN based methods such as near-duplicate image detection or segmenting tampered areas of images.
The novel human computer interface is introduced, based on tongue and lips movements and using video data from a commercially available camera. The size and direction of the movements are extracted and can be used for setting cursor actions or to other relevant activities. The movement detection is based on convolutional neural networks. The applicability of the proposed solution is shown on the ASSISLT system [1], aimed to support speech therapy for adults and children with inborn and acquired motor speech disorders. The system focuses on individual treatment using exercises that improve tongue motion and thus articulation. The system offers an adjustable set of exercises which proper performance is motivated using augmented reality. Automatic evaluation of the performance of therapeutic movements allows the therapist to objectively follow the progress of the treatment.
Breast ultrasonography (US) presents an alternative to mammography in young asymptomatic individuals and a complementary examination in screening of women with dense breasts. Handheld US is the standard-of-care, yet when used in whole-breast examination, no effort has been devoted to monitoring breast coverage and missed regions, which is the purpose of this study. We introduce a computer-aided system assisting radiologists and US technologists in covering the whole breast with minimum alteration to the standard workflow. The proposed system comprises a standard US device, proprietary electromagnetic 3D tracking technology and software that combines US visual and tracking data to estimate a probe trajectory, total time spent in different breast segments, and a map of missed regions. A case study, which involved four radiologists (two junior and two senior) performing whole-breast ultrasound in 75 asymptomatic patients, was conducted to test the importance and relevance of the system. The mean process time per breast was $$74\pm 22\,{\mathrm {s}}$$ , with no statistically significant difference between the left and the right sides, and slightly longer examination time of junior radiologists. The process time density shows that central parts of the breast have better coverage compared to the periphery. Within the central part, missed regions of minimum detectable size of $$0.09\,{\mathrm {cm}}^2$$ occur in $$8\%$$ of examinations, and non-negligible $$1\,{\mathrm {cm}}^2$$ regions occur in $$3\%$$ of cases. The results of the case study indicate that missed regions are present in handheld whole-breast US, which renders the proposed system for tracking the probe position during examination a valuable tool for monitoring coverage.
The presented paper proposes a new method for unique automatic evaluation of speech therapy exercises, one part of the future software system for speech therapy support. The method is based on the detection of the lips and tongue movements, which will help to evaluate the quality of the exercise implementation. Four different types of exercises are introduced and the corresponding features, capturing the quality of the movements, are shown. The method was tested using manually annotated data and the proposed features were evaluated and analyzed. At the second part, the tongue detection is proposed based on the convolutional neural network approach and preliminary results were shown.
Ultrasound examination plays an important role in both breast cancer screening and diagnostics. One of the drawbacks of the US examination is the uncertainty whether the whole breast was scanned. The proposed paper addresses the methodology how the completeness of the examination can be efficiently evaluated. We propose an affordable solution for simultaneously tracking and grabbing a video from a free-hand 2D ultrasound transducer during standard breast examinations by means of the probe motion tracking. From the recorded data we calculate duration in seconds, for which every part of the examined region has been captured and perform algorithmically local 3D reconstruction. Thus the system can inform the specialist performing the exam about regions that were insufficiently examined and minimize the risk of not detecting developing harmful lesions. The measure for the evaluation and comparison of the individual examinations is proposed. The functionality of the method is illustrated.
We present a new method for segmentation of phase-contrast microscopic images of cells. The algorithm is based on the variational formulation of the level set method, i.e. minimizing of a functional, which describes the level set function. The functional is minimized by a gradient flow described by an evolutionary partial differential equation. The most significant new ideas are initialization using thresholding and the introduction of a new term based on local variance that speeds up convergence and achieves more accurate results. The proposed algorithm is applied on real data and compared with another algorithm. Our method yields an average gain in accuracy of 2 %.
Barbara Zitová合作论文数Department of Image Processing;Academy of Sciences of the Czech Republic;Institute of Information Theory and Automation5