The rising prevalence of autism spectrum disorder, coupled with limited professional resources, highlights the urgency of developing efficient diagnostic tools. While standardized assessments exist, identifying subtle communication deficits, especially during multimodal interactions, remains time-consuming and prone to human error. To address this, we propose an automated behavior analysis framework that aims to support clinicians by accurately detecting both verbal and non-verbal communication markers. Specifically, we put forth a composite artificial intelligence framework that integrates various deep learning algorithms to analyze information from body and hand poses, object detection, tracking and manipulation, and speech. By combining these features with a rule-based system, we can identify events within the Autism Diagnostic Observation Schedule second edition, Construction Task, where participants initiate requests. These requests can be verbal, non-verbal or a combination of both resulting in multimodal interactions. Building on our prior work, this paper introduces a smart glass technology component, integrating gaze and blinking analysis, which are challenging for clinicians to monitor, given the multi-task nature of their role. These additions enable the detection of eye contact, a crucial social cue. Our approach allows us to recognize gestures, identify hand object manipulations, detect eye contact, and understand the natural language in clinician-participant interactions. We achieve 94% and 73% F-1 score, on verbal and non-verbal request detection, respectively, which may improve, as deep learning advances.
Automatic personality trait assessment is essential for high-quality human-machine interactions. Systems capable of human behavior analysis could be used for self-driving cars, medical research, and surveillance, among many others. We present a multimodal deep neural network with a Siamese extension for apparent personality trait prediction trained on short video recordings and exploiting modality invariant embeddings. Acoustic, visual, and textual information are utilized to reach high-performance solutions in this task. Due to the highly centralized target distribution of the analyzed dataset, the changes in the third digit are relevant. Our proposed method addresses the challenge of under-represented extreme values, achieves 0.0033 MAE average improvement, and shows a clear advantage over the baseline multimodal DNN without the introduced module.
This work presents BlinkLinMulT, a transformer-based framework for eye blink detection. While most existing approaches rely on frame-wise eye state classification, recent advancements in transformer-based sequence models have not been explored in the blink detection literature. Our approach effectively combines low- and high-level feature sequences with linear complexity cross-modal attention mechanisms and addresses challenges such as lighting changes and a wide range of head poses. Our work is the first to leverage the transformer architecture for blink presence detection and eye state recognition while successfully implementing an efficient fusion of input features. In our experiments, we utilized several publicly available benchmark datasets (CEW, ZJU, MRL Eye, RT-BENE, EyeBlink8, Researcher’s Night, and TalkingFace) to extensively show the state-of-the-art performance and generalization capability of our trained model. We hope the proposed method can serve as a new baseline for further research.
Dyadic and small group collaboration is an evolutionary advantageous behaviour and the need for such collaboration is a regular occurrence in day to day life. In this paper we estimate the perceived personality traits of individuals in dyadic and small groups over thin-slices of interaction on four multimodal datasets. We find that our transformer based predictive model performs similarly to human annotators tasked with predicting the perceived big-five personality traits of participants. Using this model we analyse the estimated perceived personality traits of individuals performing tasks in small groups and dyads. Permutation analysis shows that in the case of small groups undergoing collaborative tasks, the perceived personality of group members clusters, this is also observed for dyads in a collaborative problem solving task, but not in dyads under non-collaborative task settings. Additionally, we find that the group level average perceived personality traits provide a better predictor of group performance than the group level average self-reported personality traits.
Human-machine, human-robot interaction, and collaboration appear in diverse fields, from homecare to Cyber-Physical Systems. Technological development is fast, whereas real-time methods for social communication analysis that can measure small changes in sentiment and personality states, including visual, acoustic and language modalities are lagging, par-ticularly when the goal is to build robust, appearance invariant, and fair methods. We study and compare methods capable of fusing modalities while satisfying real-time and invariant appearance conditions. We compare state-of-the-art transformer architectures in sentiment estimation and introduce them in the much less explored field of personality perception. We show that the architectures perform differently on automatic sentiment and personality perception, suggesting that each task may be better captured/modeled by a particular method. Our work calls attention to the attractive properties of the linear versions of the transformer architectures. In particular, we show that the best results are achieved by fusing the different architectures’ preprocessing methods. However, they pose extreme conditions in computation power and energy consumption for real-time computations for quadratic transformers due to their memory requirements. In turn, linear transformers pave the way for quantifying small changes in sentiment estimation and personality perception for real-time social communications for machines and robots.
Cloud-based speech services are powerful practical tools but the privacy of the speakers raises important legal concerns when exposed to the Internet. We propose a deep neural network solution that removes personal characteristics from human speech by converting it to the voice of a Text-to-Speech (TTS) system before sending the utterance to the cloud. The network learns to transcode sequences of vocoder parameters, delta and delta-delta features of human speech to those of the TTS engine. We evaluated several TTS systems, vocoders and audio alignment techniques. We measured the performance of our method by (i) comparing the result of speech recognition on the de-identified utterances with the original texts, (ii) computing the Mel-Cepstral Distortion of the aligned TTS and the transcoded sequences, and (iii) questioning human participants in A-not-B, 2AFC and 6AFC tasks. Our approach achieves the level required by diverse applications.
•A new feature selection approach is proposed, which combines and utilizes multiple individual methods in order to achieve a more generalized solution.•A hybrid solution is proposed in the paper which combines the given, available (supervised, state-of-the-art) feature selection techniques that have their own specific, but fixed feature evaluation measures/metrics.•Adaptivity of the proposed algorithm is realized in such a way that at an individual step of the feature selection algorithm it iterates not only in the space of the variables but in the space of available features selection techniques, too.•Different directions of tests were applied: linear and non-linear dependencies with varying data distribution, noise and outliers; using benchmark datasets of the UCI Machine Learning Repository and also own real-life datasets; comparison to recent state-of-the-art feature selection methods.•TThe proposed AHFS nearly doubles the accuracy (resulting in around half value for the related modeling error) compared to the individual methods, making it a superior feature selection algorithm.
We estimated the contribution of different factors in segmentation tasks by means of deep neural networks. Results indicated that texture and optical flow have similar power, but they seem not to add up. In turn, we decided to study the `Common Fate Principle' of the 100 years gestaltism suggesting that elements that move together belong together. We developed a simple, fast, and efficient episodic segmentation method that - to some extent - resembles the `how system' of the visual processing: we dropped every piece of information except motion, and started from pure optical flow estimations on 2D videos. For the sake of segmentation, we used a parallel and fast hierarchical supervoxel algorithm. We studied (i) grid topology in space and time, (ii) 2D grid in space and topology dictated by the optical flow in time, and (iii) added deep network based depth estimation from 2D images. We measure performances on episodic foreground-background segmentation task of the Davis benchmark videos. Results are competitive to state-of-the-art segmentation techniques.
Medial axis representation (a.k.a. shape skeleton) seems to be present in visual processing, but its relevance has remained unclear. Here, we show the potentials of the medial axis transformation in the temporal propagation of superpixels. We combine (i) state-of-the-art deep neural network `sensors' for optical flow and for depth estimation and (ii) a superpixel algorithm with (iii) the medial axis transformation to obtain frame-to-frame propagation of visual objects. We study the precision of this deep learning facilitated superpixel temporal propagation. We discuss the advantages of the method compared to the temporal propagation of the superpixels themselves.
The paper introduces a methodology to define production trend classes and also the results to serve with trend prognosis in a given manufacturing situation. The prognosis is valid for one, selected production measure (e.g. a quality dimension of one product, like diameters, angles, surface roughness, pressure, basis position, etc.) but the applied model takes into account the past values of many other, related production data collected typically on the shop-floor, too. Consequently, it is useful in batch or (customized) mass production environments. The proposed solution is applicable to realize production control inside the tolerance limits to proactively avoid the production process going outside from the given upper and lower tolerance limits. The solution was developed and validated on real data collected on the shop-floor; the paper also summarizes the validated application results of the proposed methodology.