Emotions are central to interactive systems, shaping user experience, decision-making, and engagement. In educational contexts, emotionally expressive avatars can foster empathy, motivation, and social belonging, but existing solutions face critical limitations. This article proposes an emotion-aware multidimensional framework that integrates technical, pedagogical, sociocultural, and operational perspectives to guide the equitable design of interactive educational systems. Grounded in a systematic review of 127 studies (2014–2024), our results show that current approaches face three key challenges: (1) realism-accessibility trade-offs (e.g., diffusion models’ F1=0.91 vs. GANs’ 34ms latency), (2) cultural bias in emotion recognition (89% Western dominance reduced to 20% using our adaptation protocols), and (3) limited pedagogical integration of affective features (with gains of 23% in retention when aligned with Bloom’s Taxonomy). The review highlights fragmentation across technical, pedagogical, and cultural domains, underscoring the need for integration. To address this, we propose an Emotion-Aware Multidimensional Framework that unifies these perspectives into actionable design and evaluation protocols. By situating emotions as a core dimension of system design, the study contributes not only to educational applications but also to the broader field of interactive systems where affective engagement is critical.
Generative Adversarial Networks (GANs) have made significant advancements in generating high-quality facial animations, transforming static images into dynamic expressions. However, a persistent challenge in facial animation is ensuring temporal coherence-smooth and consistent transitions between frames. This issue often leads to visual artifacts and unnatural motion, hindering the realism of animated facial expressions. In this paper, we introduce a novel method to improve temporal coherence in GAN-based facial animation. By incorporating a specialized temporal consistency module and a recurrent loss function, our approach reduces abrupt transitions and enhances the fluidity of facial expression synthesis. We present both quantitative and qualitative evaluations of our method, which reduces temporal artifacts by 58% (TCM) and improves PSNR by 3.3 dB over baselines. We validate our approach across three facial expression datasets, demonstrating perceptual gains in temporal coherence, essential for robust emotion-driven character animation pipelines.
Learning environments can benefit from the enactive concept which considers human and technological aspects as coupled together so that interaction is seen as a cycle of perceptually guided actions. Although existing literature reviews have been conducted to address the use of technologies in enactive learning environments, we observe they do not address the relations of technology with core concepts of enactive systems and its social aspects. This paper presents a systematic literature review (SLR) to investigate if and how core concepts of enactive systems are being addressed by learning environments. The SLR was based on the PRISMA protocol, and considered relevant sources of studies (ACM DL, IEEE Xplore, SpringerLink, Scopus, Scielo). We based our analysis on a set of 104 selected articles (from 4514 raised), scrutinized with regards to contextual information, and a proposed set of categories of data related to subjects identified in the studies (e.g., embodiment, technology, social, among others). Our work reveals the types of technologies being used, the types of enactive interactions promoted by the technology-based environments, and the main challenges for a research agenda in the field. Our discussion raised open challenges related to a need for a common vocabulary, a framework to organize concepts in the field, and missing points that deserve further research, such as social aspects of enactive system.
Technological devices integrate people’s lives, actively connecting the Physical and the Digital worlds. Through ubiquitous computing, we can build a network of objects that, in addition to exchanging information, can perceive and act in a certain environment. The network of interconnected objects that is part of Human daily life can and should be called the Internet of Human Things (IoHT) as it involves re-configurations in living environments, in which objects (physical things) start to interact with each other and with people, often without human awareness. For this reason, the design of an IoT environment requires a good understanding of the problem and an evolution of ideas towards a possible solution in which People and other Physical things are linked to the Digital and can be considered as a single (Social, Physical and Digital) System. This article presents a Socially Aware Design process which evolves a solution starting from the understanding of the problem by the interested parties who act as co-designers. The process is illustrated with three workshops held with children, accompanied by their families in a hospital environment, whose objective was the evolution of an IoHT-based (Internet of Human Things) scenario. From the understanding of the design problem to the closing of the third workshop, the maturation of a design process for IoT environments with people is presented and discussed.
Os conceitos de modelagem espaço-temporais estão bem definidos nos campos da geometria e matemática, com representações tridimensionais consolidadas como o Tesseract, ou hipercubo. Embora estas representações funcionem em ambientes baseados em coordenadas espaciais, mesmo na computação e matemática, um objeto quadridimensional ainda varia de contexto e aplicação. Ainda mais desafiador é imaginarmos uma realidade com estas características aplicado a experiências visuais e a navegação em um ambiente deste tipo. Este trabalho investiga as representações espaço-temporais observadas nas artes e literatura contemporânea aplicadas na representação ou descrição de realidade, analisando seu paralelo com as modelagens 4D computacionais apontando como este tipo de ambiente pode ser corretamente descrito e as limitações narrativas aplicadas a estes objetos. A discussão apresentada neste trabalho tem como objetivo apontar como os conceitos de objetos quadridimensionais se apresentam na cultura contemporânea, e como estas descrições podem caminhar com as descobertas computacionais.
New ubiquitous technologies and interaction paradigms can play a key role in facilitating social interaction processes among children. A class of systems named "enactive" aims to support fluid interaction between technology and people via feedback cycles - the effect of technology on the human agent is fed back by the human action on the technology - based on the use of data sensors. We have extended this concept in a long-term project on socioenactive systems, by emphasizing social aspects in the enactive phenomena. In this article, we investigate how a robotbased experience can promote social behavior among children in an educational context. The study was conducted in workshops with 26 children (4-5 years old), organized in two groups. The system scenario used a narrative based on an adaptation of the Little Red Riding Hood tale to investigate children's interaction in playing embodied-based situations. Data captured from video-recorded workshop sessions were analyzed post hoc using the Grounded Theory methods. In total, 26 interaction coding were identified, with a high interrater reliability assessed by Cohen's Kappa (k = 0.84). Our findings indicate that 50.5% of the children's actions were the result of children-children and group-robot interactions, compared to the 38% of children interacting individually with the robot. This indicates children's high degree of embodied peer collaboration and initiative to accomplish the tasks in the proposed scenario. These results contribute to inform the design and construction of future socioenactive systems.
Facial expressions are important data to understand how systems in social environments impact people in it. The presence of new technologies and new coupled forms of interaction with the ubiquity of computing and social networks, present challenges that require the consideration of new factors as emotional. Socioenactive systems represent a complex scenario that requires the treatment of technological aspects in which the consideration of the social dynamic, enhanced by concepts such as affective computing and enactive systems. This work presents a proposal for facial recognition in the wild applied to outputs of socioenative systems. These results reinforce how the design of socioenactive systems can promote positive changes in the emotional state of children in an educational context and promote social interactions.
Constructionism has different meanings for what the learner is constructing, a concrete object or an idea, and whether this construction is done through the use of digital technology. In the embodied-based environment created in this study, children carry out or "construct" a series of actions, a performance, which allows them to solve the task of directing an already programmed robot to a particular target. The activity was based on an adaptation of the Little Red Riding Hood narrative. Children played the role of rangers (instead of hunters) who had to coordinate their actions in order to help a robot (an mBot characterized as a Robot-Wolf) find Grandma's laboratory, so she could fix its GPS. The children wore boots, which were used to interact with the Robot-Wolf. The main question we addressed was how to create this embodied-based environment for kindergarten children, and how to identify the actions children performed and the concepts they used in this construction. The research was based on the Socially-Aware Design, and the study was conducted in a school setting for kindergarten students. 26 children (11F, 15 M), between 4 and 5 years old, participated in this study. The children's activities were recorded and analysed using the Grounded Theory methodology. The results show that the children's performative sequence of actions is distinct from the hands-on, heads-in processes that are part of classical constructionist environments. For the children, the process and the product of their (body) "dance" were underlying their movement (perceiving and acting), working together so as to accomplish the given task in the environment. Practitioner notes What is already known about this topic Constructionism has shown as a way for learning via materialization of tangible objects or construction of ideas. There are very few examples of constructionist learning environments for kindergarten children to interact with digital technology. What this paper adds The creation of the constructionist embodied-based environment for supporting kindergarten children interacting with digital technologies. The constructionist embodied-based environment created was effective in supporting kindergarten children interacting with an already programmed robot. Children's performative sequence of actions reveals no separation of the hands-on, heads-in processes as in classical constructionist environments. The creation of a constructionist environment in which children can act and collaborate to construct a series of actions which allow them to solve a specific challenge. Implications for practice and/or policy Concepts like leadership and anticipation which were embodied are expressed by the children help to understand how the robot-based activity can help children to construct knowledge related to, for example, computational thinking.
Computational systems based on ubiquitous and pervasive technology present several challenges related to the interaction of people with scenarios constituted by sensors and actuators, changing the mindset of what we used to understand as interaction with a computer. This also has influence in the ways of considering the design of systems based on contemporary technology for the educational context. To cope with the challenges of ubiquitous computing, the concept of socioenactive system is being constructed as a system in which human and technological aspects are coupled together in a cycle of perceptually guided actions of people interacting with elements of the physical environment and with other people in the same scenario. In this work we address the design of a socioenactive system as an evolution of two previous systems designed and experimented with 5-year-old children in an educational context. The contribution of this paper is twofold: 1. We present an analysis of two different systems tested in educational scenarios, pointing out the lack of elements that should be present in a complete cycle of socioenactive systems, suggesting requirements for a third system; 2. We present an architecture for the third system and a simulation of its usage. Results of the third system and its simulation inform the next activities of bringing it to real life in a practice proposed for the same audience and context as the previous systems.
3D Avatars are an efficient solution to complement the representation of sign languages in computational environments. A challenge, however, is to generate facial expressions realistically, without high computational cost in the synthesis process. This work synthesizes facial expressions with precise control through spatio-temporal parameters automatically. With parameters compatible with the gesture synthesis models for 3D avatars, it is possible to build complex expressions and interpolations of emotions through the model presented. The built method uses independent regions that allow the optimization of the animation synthesis process, reducing the computational cost and allowing independent control of the main facial regions. This work contributes to the definition of non-manual markers for 3D Avatar facia expression and its synthesis process. Also, a dataset with the base expressions was built where 4D information of the geometric control points of the avatar built for the experiments presented is found. The results of the generated outputs are validated in comparison with other expression classification approaches using Spatio-temporal data and machine learning, presenting superior accuracy for the base expressions. The rating is reinforced by evaluations conducted with the deaf community showing a positive acceptance of the facial expressions and synthesized emotions.
Contemporary computer systems fully integrate people’s daily lives in the most diversified contexts. The adequate construction of those systems demands well-defined patterns for including a whole community of users encompassing people with disabilities. Current interaction scenarios with pervasive computing technologies present design challenges for the applicability of standards and guidelines such as those of the W3C Web Content Accessibility Guidelines (WCAG), suited for Web systems. Such scenarios require further studies to deal with the endeavor of supporting designers to project pervasive interactive systems that consider all people, regardless of their condition and ability, allowing different types of interactions to arise. In this article, we conduct an exploratory study to investigate the potential and limitations of current accessibility guidelines for pervasive systems. We explore the last decade of literature studies presenting solutions and challenges for accessibility found for pervasive and ubiquitous computing contexts. The results of this research indicate that it is desirable to define accessibility models for pervasive systems if we consider an accessibility layer that will indicate principles and solutions that should be generic enough to support different categories of impairments and contexts.
Emotions can be synthesized in virtual environments through spatial calculations that define regions and intensities of displacement, using landmark controllers that consider spatio-temporal variations. This work presents a proposal for calculating spatio-temporal mesh for 3D objects that can be used in the realistic synthesis of emotions in virtual environments. his technique is based on calculating centroids by facial region and uses classification by machine learning to define the positions of geometric controllers making the animations realistic.
Systems that use virtual environments with avatars for information communication are of fundamental importance in contemporary life. They are even more relevant in the context of supporting sign language communication for accessibility purposes. Although facial expressions provide message context and define part of the information transmitted, e.g., irony or sarcasm, facial expressions are usually considered as a static background feature in a primarily gestural language in computational systems. This article proposes a novel parametric model of facial expression synthesis through a 3D avatar representing complex facial expressions leveraging emotion context. Our technique explores interpolation of the base expressions in the geometric animation through centroids control and Spatio-temporal data. The proposed method automatically generates complex facial expressions with controllers that use region parameterization as in manual models used for sign language representation. Our approach to the generation of facial expressions adds emotion to representation, which is a determining factor in defining the tone of a message. This work contributes with the definition of non-manual markers for Sign Languages 3D Avatar and the refinement of the synthesized message in sign languages, proposing a complete model for facial parameters and synthesis using geometric centroid regions interpolation. A dataset with facial expressions was generated using the proposed model and validated using machine learning algorithms. In addition, evaluations conducted with the deaf community showed a positive acceptance of the facial expressions and synthesized emotions.
The mobile robots popularization, in terrestrial, aerial or aquatic context, opens the possibility of a new research niche: the use of these robots for activities such as monitoring, search/rescue, cleaning services, as well their use in precision agriculture. As many times the area of operation of these robots is very large, it becomes unfeasible the fulfillment of the activity using a single robot, but even using many robots, there is the limitation of the battery autonomy. With the popularization of autonomous robots use, the search for mechanisms to make these robots recharge their battery autonomously are impulsed, however, the choice of the best place to insert the charging stations considering the displacement time saved and use of battery to move the robots is a challenge. The present work shows a Java Desktop graphical application that makes use of the JST library that requests data to the user, such as the map of the place to be explored, quantity of charging stations, if these stations must be in static places, perhaps due to electrical outlet locations restriction, or if they could be dynamically placed. Based on these data an optimal solution to insert the charging stations as well as the definition of each robot activity area are presented. Algorithms such as Voronoi, Viktor Grabarchuk and Centroid position are used in this process. The Voronoi Algorithm allowed the balanced division of the action area into a group of robots considering static recharge positions. The combination of the Viktor Grabarchk and Centroid Position algorithms allowed a balanced division of the operating area for different robots, and also, the definition of a central position to allocate the bases of recharges, which reduces the time of displacement of the robot to the base when necessary.
Facial expressions and associates emotions important factors defining the message tone and context information in both spoken and sign language communication. Sign language virtual systems use structural models that define control values specifying body configurations in the animation process, where facial parameters are generally relegated to simple templates or completely neglected. In this work, a facial expression parametrization for avatars is proposed through an procedure that aims to identify the most relevant facial landmarks and emotions in the context of sign languages, in order to enhance automatic sign synthesis systems. An analysis of the influence of landmarks on a geometric mesh based on MPEG-4 model of human face is performed in this research, aiming to identify the principal components and their relationships, in order to allow further optimization of the animation process, supporting faster and lighter avatar animation.