In this paper, we propose a new annotation scheme to classify different types of clauses in Terms-and-Conditions contracts with the ultimate goal of supporting legal experts to quickly identify and assess problematic issues in this type of legal documents. To this end, we built a small corpus of Terms-and-Conditions contracts and finalized an annotation scheme of 14 categories, eventually reaching an inter-annotator agreement of 0.92. Then, for 11 of them, we experimented with binary classification tasks using few-shot prompting with a multilingual T5 and two fine-tuned versions of two BERT-based LLMs for Italian. Our experiments showed the feasibility of automatic classification of our categories by reaching accuracies ranging from .79 to .95 on validation tasks.
This paper introduces a no-code platform for modular prompt engineering, designed to democratize access to generative AI for nondevelopers. By integrating advanced technologies such as Node.js, Express, MongoDB, and Azure OpenAI services, the platform provides a robust and flexible environment for creating and managing AI-driven tasks. The intuitive frontend, built with React and TypeScript, enables users with minimal coding expertise to design, execute, and evaluate complex AI workflows. A key feature of the platform is its extensible plugin system, which allows users to easily incorporate additional functionalities to meet their specific needs. This no-code approach empowers a broader audience to harness the power of generative AI, fostering innovation and enabling diverse applications across various fields. By lowering the technical barriers, the platform paves the way for widespread adoption of AI technologies, driving the future of AI-enhanced solutions.
This paper presents key aspects and trade-offs that designers and Human-Computer Interaction practitioners might encounter when designing multimodal interaction for older adults. The paper gathers literature on multimodal interaction and assistive technology, and describes a set of design challenges specific for older users. Building on these main design challenges, four trade-offs in the design of multimodal technology for this target group are presented and discussed. To highlight the relevance of the trade-offs in the design process of multimodal technology for older adults, two of the four reported trade-offs are illustrated with two user studies that explored mid-air and speech-based interaction with a tablet device. The first study investigates the design trade-offs related to redundant multimodal commands in older, middle-aged and younger adults, whereas the second one investigates the design choices related to the definition of a set of mid-air one-hand gestures and voice input commands. Further reflections highlight the design trade-offs that such considerations bring in the process, presenting an overview of the design choices involved and of their potential consequences.
The paper presents the design of an assistive reading tool that integrates read-aloud technology with eye tracking to regulate the speed of reading and support struggling readers in following the text while listening to it. The paper describes the design rationale of this approach, following the theory of auditory-visual integration, in terms of an automatic self-adaptable technique based on the reader's gaze that provides an individualized interaction experience. This tool has been assessed in a controlled experiment with 20 children (aged 8-10 years) with a diagnosis of dyslexia and a control group of 20 children with typical reading abilities. The results show that children with reading difficulties improved their comprehension scores by 24% measured on a standardized instrument for the assessment of reading comprehension and that children with more inaccurate reading (N = 9) tended to benefit more. The findings are discussed in terms of a better integration between audio and visual text information, paving the way to improve standard read-aloud technology with gaze-contingency and self-adaptable techniques to personalize the reading experience.
Multimodal human–computer interaction has been sought to provide not only more compelling interactive experiences, but also more accessible interfaces to mobile devices. With the advance in mobile technology and in affordable sensors, multimodal research that leverages and combines multiple interaction modalities (such as speech, touch, vision, and gesture) has become more and more prominent. This article provides a framework for the key aspects in mid-air gesture and speech-based interaction for older adults. It explores the literature on multimodal interaction and older adults as technology users and summarises the main findings for this type of users. Building on these findings, a number of crucial factors to take into consideration when designing multimodal mobile technology for older adults are described. The aim of this work is to promote the usefulness and potential of multimodal technologies based on mid-air gestures and voice input for making older adults' interaction with mobile devices more accessible and inclusive.
JVRB, 16(2019), no. 2. - This paper presents two early studies aimed at investigating issues concerning the design of multimodal interaction based on voice commands and one-hand mid-air gestures - with mobile technology specifically designed for visually impaired and elderly users. These studies were carried out on a new device allowing enhanced speech recognition (including lip movement analysis) and mid-air gesture interaction on Android operating system (smartphone and tablet PC). We discuss the initial findings and challenges raised by these novel interaction modalities, and in particular the issues regarding the design of feedback and feedforward, the problem of false positives, and the correct orientation and distance of the hand and the device during the interaction. Finally, we present a set of feedback and feedforward solutions designed to overcome the main issues highlighted.
This paper presents two early studies aimed at investigating issues concerning the design of multimodal interaction based on voice commands and one-hand mid-air gestures with mobile technology specifically designed for visually impaired and elderly users. These studies were carried out on a new device allowing enhanced speech recognition (including lip movement analysis) and mid-air gesture interaction on Android operating system (smartphone and tablet PC). We discuss the initial findings and challenges raised by these novel interaction modalities, and in particular the issues regarding the design of feedback and feedDigital Peer Publishing Licence Any party may pass on this Work by electronic means and make it available for download under the terms and conditions of the current version of the Digital Peer Publishing Licence (DPPL). The text of the licence may be accessed and retrieved via Internet at http://www.dipp.nrw.de/. First presented at the 3rd International Conference on Human-Computer Interaction Theory and Applications (HUCAPP) 2018, extended and revised for JVRB forward, the problem of false positives, and the correct orientation and distance of the hand and the device during the interaction. Finally, we present a set of feedback and feedforward solutions designed to overcome the main issues highlighted.
Assessing reading skills is an important task teachers have to perform at the beginning of a new scholastic year to evaluate the starting level of the class and properly plan next learning activities. Digital tools based on automatic speech recognition (ASR) may be really useful to support teachers in this task, currently very time consuming and prone to human errors. This paper presents a web application for automatically assessing fluency and accuracy of oral reading in children attending Italian primary and lower secondary schools. Our system, based on ASR technology, implements the Cornoldi’s MT battery, which is a well-known Italian test to assess reading skills. The front-end of the system has been designed following the participatory design approach by involving end users from the beginning of the creation process. Teachers may use our system to both test student’s reading skills and monitor their performance over time. In fact, the system offers an effective graphical visualization of the assessment results for both individual students and entire class. The paper also presents the results of a pilot study to evaluate the system usability with teachers.
Designing and developing mobile technology that is able to meet the needs of older adults is fundamental to improve their independent living and expand their social inclusion. However, although mobile technology is nowadays widely present in our every-day activities, older adults continue to lag in its adoption. While exploring what hinders older adults in adopting mobile technology and questioning about how to increase their accessibility to it, the paper presents the ECOMODE project, whose technology based on the Event-Driven Compressive (EDC) paradigm is a possible answer. First, to contextualize our study, the paper describes the ECOMODE technology based on multimodal interaction, i.e. mid-air gestures combined with voice commands. Then, it details the process followed to design the interaction based on the ECOMODE technology, which aims to increase accessibility and usability of mobile devices for older adults.
This paper presents two early studies aimed at investigating issues concerning the design of multimodal interaction based on voice commands and mid-air gestures with mobile technology specifically designed for visually impaired and elderly users. These studies have been carried out on a new device allowing enhanced speech recognition (interpreting lip movements) and mid-air gesture interaction on Android devices (smartphone and tablet PC). The initial findings and challenges raised by these novel interaction modalities are discussed. These mainly centre on issues of feedback and feedforward, the avoidance of false positives and point of reference or orientation issues regarding the device and the mid-air gestures.
In this paper we describe a system for analyzing the reading errors made by children of the primary and middle schools. To assess the reading skills of children in terms of reading accuracy and speed, a standard reading achievement test, developed by educational psychologists and named “Prove MT” (MT reading test), is used in the Italian schools. This test is based on a set of texts specific for different ages, from 7 to 13 years old. At present, during the test, children are asked to read aloud short stories, while teachers manually write down the reading errors on a sheet and then compute a total score based on several measures, such as duration of the whole reading, number of read syllables per second, number and type of errors, etc. The system we have developed is aimed to support the teachers in this task by automatically detecting the reading errors and estimating the needed measures. To do this we use an automatic speech-totext transcription system that employs a language model (LM) trained over the texts containing the stories to read. In addition, we embed in the LM an error model that allows to take into account typical reading errors, mostly consisting in pronunciation errors, substitutions of syllables or words, word truncation, etc. To evaluate the performance of our system we collected 20 audio recordings, uttered by 8-13 years old children, reading a novel belonging to “Prove MT” set. It is worth mentioning that the error model proposed in this paper for assessing the reading capabilities of children performs closely to an “oracle” error model obtained from manual transcriptions of the readings themselves.
The goal of this half-day workshop is to investigate effective ways to leverage the recent advances in the automatic recognition of mid-air gestures and speech commands, with a special focus on (but not limited to) interaction with mobile technology and inclusive design for older adults and people with special needs. In particular, the workshop aims to investigate these issues from a multidisciplinary prospective by bringing together experts on interaction design, user experience, usability, accessibility, innovative hardware and software solutions that are related to technology based on mid-air gestures, speech and multi-modal interaction.
Since they can integrate a wide range of interactive modalities, multimodal interfaces are considered to improve accessibility for a variety of users, including older adults. However, only few works have actually explored how older adults approach multimodal interaction outside specific contexts and have done so mainly in comparison to much younger users. This study explores how older (65+ years old), middle-aged (55-65 years old) and younger adults (25-35 years old) use mobile multimodal interaction in an everyday activity (i.e. taking photos with a tablet) by using midair gestures and voice commands, and investigates the differences and similarities between the considered age groups. Preliminary findings from a video-analysis show that all groups easily combine the proposed modalities when interacting with a tablet device. Furthermore, compared to younger adults, older and middle-aged adults show similarities in the way they perform gesture and voice commands.
Incorporating research on personality recognition into computers, both from a cognitive as well as an engineering perspective, would facilitate the interactions between humans and machines. Previous attempts on personality recognition have focused on a variety of different corpora (ranging from text to audiovisual data), scenarios (interviews, meetings), channels of communication (audio, video, text), and different subsets of personality traits (out of the five ones from the Big Five Model). Our study uses simple acoustic and visual nonverbal features extracted from multimodal data, which have been recorded in previously uninvestigated scenarios, and consider all five personality traits and not just a subset. First, we look at the human-machine interaction scenario, where we introduce the display of different “collaboration levels.” Second, we look at the contribution of the human-human interaction (HHI) scenario on the emergence of personality traits. Investigating the HHI scenario creates a stronger basis for future human-agents interactions. Our goal is to study, from a computational approach, the emergence degree of the five personality traits in these two scenarios. The results demonstrate the relevance of each of the two scenarios when it comes to the degree of emergence of certain traits and the feasibility to automatically recognize personality under different conditions.
This paper aims to present the evaluation of eSchooling, an ICT system supporting competence-based education. The eSchooling team ran an experiment involving ten high schools throughout an entire school year, so as to cover all the teaching stages (from activities planning to the final students evaluations). In order to guarantee objectivity and independence, external experts performed monitoring, validation and evaluation of the experimental results. These experts analysed different aspects: the logs of the eSchooling system, the response of a user community, the logs of the interactive e-book system and the outputs of users focus groups. The main indication derived from the analysis of the collected data is to involve entire groups of teachers of the same class, rather than isolated ones, when evaluating such kinds of educative tools. Another suggestion is to move in the direction of a tighter integration with other ICT tools such as the electronic board for recording activity, even when it is not competence-based.
Children with reading difficulties face several obstacles in learning to fluently read written material. Multimedia applications integrating text-to-speech (TTS) synthesisers are valuable tools for supporting reading activities. The paper presents GARY, an application that combines TTS synthesis with eye tracking. GARY is meant to be used on a tablet device coupled with an eye tracker. Making use of the information from reader's eye movement, the system allows users to adapt the speed rate of the synthesised voice to their pace of reading. The paper describes the system, its functioning and future steps in designing a tool for supporting readers' ability in making the connection between the sounds heard and the letters read.
The paper presents the preliminary studies carried out in ECOMODE (Event-Driven Compressive Vision for Multimodal Interaction with Mobile Devices), an EU project aiming to design a multimodal interaction suitable for elderly people using mobile technology. The project exploits EDC (Event Driven Compressive) paradigm to realize a new generation of low-power multimodal human-computer interfaces for mobile devices. In our studies we investigated (1) how older adults interact with mobile devices and (2) how they use applications based on mid-air gesture interaction. Our studies basically confirm the findings of previous works about usability and accessibility issues, specifically characterizing elderly users' interaction with mobile technology.
The paper presents the preliminary studies carried out in ECOMODE (Event-Driven Compressive Vision for Multimodal Interaction with Mobile Devices), an EU project aiming to design a multimodal interaction suitable for elderly people using mobile technology. The project exploits EDC (Event Driven Compressive) paradigm to realize a new generation of low-power multimodal human-computer interfaces for mobile devices. In our studies we investigated (1) how older adults interact with mobile devices and (2) how they use applications based on mid-air gesture interaction. Our studies basically confirm the findings of previous works about usability and accessibility issues, specifically characterizing elderly users' interaction with mobile technology.
This paper presents a preliminary study aiming to define a list of guidelines for designing effective software tools to support dyslexic children while reading an e-text. We start our work with a literature review, which main result is that, to our knowledge, such guidelines do not exist. After introducing the main difficulties met by dyslexic children in reading, we highlight two categories of actions that should be inserted in an effective tool for dyslexics: personalized text visualization and supported reading. Then, we describe an application we developed as a design example. Finally, we conclude with some preliminary considerations about how to evaluate the designed applications.
Roberto Basili合作论文数Department of Computer Science;University of Rome "Tor Vergata"4