Videokymographic (VKG) images of the human larynx are often used for automatic vibratory feature extraction for diagnostic purposes. One of the most challenging parameters to evaluate is the mucosal wave's presence and its lateral peaks' sharpness. Although these features can be clinically helpful and give an insight into the health and pliability of vocal fold mucosa, the identification and visual estimation of the sharpness can be challenging for human examiners and even more so for an automatic process. This work aims to create and validate a method that can automatically quantify the lateral peak sharpness from the VKG images using a convolutional neural network.
This work proposes a new software system ASSISLT to support speech therapy for children and adults using deep learning approaches. The application offers an adjustable set of exercises recommended by a speech therapist and aims to motivate and help with regular home practice. Augmented reality is employed to lead the exercise moves and to appraise the effort. The core of the system is automatically evaluating exercises using a webcam and developing image processing and neural network methods for face, lips, teeth, and tongue detection. The pipeline is shown together with solutions of subtasks and with demonstrations of the functionality. The statistical validation of the ASSISLT is provided, comparing the performance of speech therapy specialists and the software.
The novel human computer interface is introduced, based on tongue and lips movements and using video data from a commercially available camera. The size and direction of the movements are extracted and can be used for setting cursor actions or to other relevant activities. The movement detection is based on convolutional neural networks. The applicability of the proposed solution is shown on the ASSISLT system [1], aimed to support speech therapy for adults and children with inborn and acquired motor speech disorders. The system focuses on individual treatment using exercises that improve tongue motion and thus articulation. The system offers an adjustable set of exercises which proper performance is motivated using augmented reality. Automatic evaluation of the performance of therapeutic movements allows the therapist to objectively follow the progress of the treatment.
The presented paper proposes a new method for unique automatic evaluation of speech therapy exercises, one part of the future software system for speech therapy support. The method is based on the detection of the lips and tongue movements, which will help to evaluate the quality of the exercise implementation. Four different types of exercises are introduced and the corresponding features, capturing the quality of the movements, are shown. The method was tested using manually annotated data and the proposed features were evaluated and analyzed. At the second part, the tongue detection is proposed based on the convolutional neural network approach and preliminary results were shown.
Barbara Zitová合作论文数Department of Image Processing;Academy of Sciences of the Czech Republic;Institute of Information Theory and Automation4