Reassembling real-world archaeological artifacts from fragmented pieces poses significant challenges due to erosion, missing regions, irregular shapes, and large-scale ambiguity. Traditional jigsaw puzzle solvers, often designed for clean synthetic scenarios, struggle under these conditions, especially when the number of fragments grows into the thousands, as in the RePAIR benchmark. In this paper, we propose a human-in-the-loop (HIL) puzzle solving framework designed to address the complexity and scale of real-world cultural heritage reconstruction. Our approach integrates an automatic relaxation-labeling solver with interactive human guidance, allowing users to iteratively lock verified placements, correct errors, and guide the system toward semantically and geometrically coherent assemblies. We introduce two complementary interaction strategies, Iterative Anchoring and Continuous Interactive Refinement, which support scalable reconstruction across varying levels of ambiguity and puzzle size. Experiments on several RePAIR groups demonstrate that our hybrid approach substantially outperforms both fully automatic and manual baselines in accuracy and efficiency, offering a practical solution for large-scale expert-in-the-loop artifact reassembly.
Collaborative problem solving (CPS) encourages children to communicate effectively and fosters creativity and critical thinking. New digital technologies can potentially provide opportunities for innovative ways to promote CPS at an early age. However, one key challenge of such technologies is to be engaging and effective while also pedagogically adapting to developmental needs of young learners. In this paper, we propose that a multimodal conversational AI system, Kid Space, providing interactive learning experiences could enrich CPS and promote positive educational outcomes. Kid Space combines multimodal classroom sensing with a projected 3D virtual environment and an animated conversational agent to engage with children via spoken dialogue and physical interactions. Within an exploratory case study, we formatively evaluated CPS behaviors of a set of children engaged with Kid Space to inform design modifications for the type of CPS learning analytics educators could benefit from and improved support of CPS behaviors in conversational AI-mediated learning environments. Our quantitative results indicate that there were significant correlations between CPS behaviors and (1) joint engagement of students working together and (2) pedagogical interventions provided by the instructional assistant facilitating the learning experience. The results imply that, with necessary design augmentations, Kid Space could potentially play a valuable role in supporting CPS processes of young learners.
Reconstructing accurate 3D models of archaeological artifacts, particularly painted ones like frescoes, presents significant challenges due to their complex geometries and the need to capture detailed pictorial information. Traditional 3D scanning technologies excel in geometry but fall short in capturing texture quality, making RGB cameras a better choice for painted artifacts. This paper proposes a comprehensive pipeline that integrates automated segmentation with photogrammetric tools to address these issues. The semi-automatic segmentation phase minimizes manual effort while ensuring accurate mask generation for intricate artifacts, and the reconstruction process efficiently produces dense 3D models with high geometric and textural fidelity. The final registration of models ensures precise alignment, resulting in an accurate and detailed reconstruction. Tested on fragmented frescoes as part of the EU Horizon 2020 RePAIR project, our pipeline demonstrates a scalable and computationally efficient solution for large-scale CH digitization and restoration, aiding the reassembly of fragmented artifacts with minimal manual input.
Reassembling 3D broken objects is a challenging task. A robust solution that generalizes well must deal with diverse patterns associated with different types of broken objects. We propose a method that tackles the pairwise assembly of 3D point clouds, that is agnostic on the type of object, and that relies solely on their geometrical information, without any prior information on the shape of the reconstructed object. The method receives two point clouds as input and segments them into regions using detected closed boundary contours, known as breaking curves. Possible alignment combinations of the regions of each broken object are evaluated and the best one is selected as the final alignment. Experiments were carried out both on available 3D scanned objects and on a recent benchmark for synthetic broken objects. Results show that our solution performs well in reassembling different kinds of broken objects.
Jigsaw puzzle solving is a challenging task for computer vision since it requires high-level spatial and semantic reasoning. To solve the problem, existing approaches invariably use color and/or shape information but in many real-world scenarios, such as in archaeological fresco reconstruction, this kind of clues is often unreliable due to severe physical and pictorial deterioration of the individual fragments. This makes state-of-the-art approaches entirely unusable in practice. On the other hand, in such cases, simple geometrical patterns such as lines or curves offer a powerful yet unexplored clue. In an attempt to fill in this gap, in this paper we introduce a new challenging version of the puzzle solving problem in which one deliberately ignores conventional color and shape features and relies solely on the presence of linear geometrical patterns. The reconstruction process is then only driven by one of the most fundamental principles of Gestalt perceptual organization, namely Wertheimer's law of good continuation. In order to tackle this problem, we formulate the puzzle solving problem as the problem of finding a Nash equilibrium of a (noncooperative) multiplayer game and use classical multi-population replicator dynamics to solve it. The proposed approach is general and allows us to deal with pieces of arbitrary shape, size and orientation. We evaluate our approach on both synthetic and real-world data and compare it with state-of-the-art algorithms. The results show the intrinsic complexity of our purely line-based puzzle problem as well as the relative effectiveness of our game-theoretic formulation.
Educational technology research has found that parents of young children widely share concerns about extended screen time, lack of physical activity, and lack of social interaction. Kid Space was developed to address these concerns by enabling multi-modal and immersive collaborative play-based learning. Kid Space utilizes multiple sensing technologies with an immersive physical space through a human-scale wall projection and incorporates a conversational AI agent to interact with children, understand individual progress, and personalize learning experiences in a blended physical and digital environment. To evaluate Kid Space in the wild, we conducted a multi-method user study involving a quasi-experimental design and exploratory case study with 14 students and three educators in an elementary school. Mixed methods for data collection and analysis were used to understand the students' and educators' perceptions of Kid Space and its impact on the students’ educational outcomes (learning engagement, experience, and performance). The findings showed (1) positive perceptions toward Kid Space, (2) high levels of engagement - with decreased screen time (41% of the time), increased physical activity (99.3% of the time), and increased social interactions with conversational AI agent and the other collaborating student (52% of the time), and (3) significant learning gains after experiencing Kid Space (24% gain, paired t-test: p < 0.01). These positive results are accompanied by critical user insights for improving future iterations of Kid Space to validate long-term educational outcomes.
Previous research showed that the parents acknowledged the technology's benefits for their young children's learning, however, they are still worried about the extended screen time, lack of physical activity and lack of social interactions. To address these concerns, we developed Kid Space to enable pedagogically appropriate technology use for children in early childhood education by combining various sensing technologies with a multi-modal conversational artificial intelligence system that can interact with children, understand individual progress and provide personalised learning experiences. To understand the impact of Kid Space on the parents' initial concerns about technology use by their young children, we conducted a multi-method user study: (1) a quasi-experimental design and (2) formative research method using an exploratory case study with a set of children and their parents experiencing Kid Space in their homes. The results show that after experiencing Kid Space with their children, the parents felt significantly less concerned about screen time, social interactions and physical activity and reported positive perceptions towards pedagogical value of Kid Space. Detailed analysis on the multi-modal data quantitatively and qualitatively validated why Kid Space alleviated these concerns. Future research is needed to validate long-term educational value of Kid Space and generate insights for improvement for next iterations.Practitioner notesWhat is already known about this topicPlay-based learning is critical for young children's education, but digital games create major concerns around extended screen time, lack of physical activity and lack of social interactions.Blending digital and physical spaces could support pedagogically appropriate technology use for young children. Towards this end, there are some exemplary studies in the state-of-the-art reporting positive educational outcomes as an effect of utilising such spaces. However, none of these studies supported children's most natural mode of communication in their interactions with the systems-speaking.Pedagogical conversational agents (PCAs) are promising, but they are tricky when it comes to young children's speech because of unique technical challenges resulted from how children use language and communicate with digital systems.What this paper addsTo our best knowledge, Kid Space is one of the earliest implementations of a PCA with a multi-modal artificial intelligence (AI) system utilising physical and digital learning manipulatives for maths learning with a focus on early childhood education. The key contributions of this paper are (1) the design and development of an end-to-end multi-modal system enabling Wizard-of-Oz experimentation for initial evaluations with users, (2) the creation of a multi-modal, in-the-wild labelled dataset with children-agent, children-parent and children-physical/digital space interactions enabling advancements for AI components for later evaluations with users and (3) the generation of rich insights from an initial research study on user perceptions and engagement as well as actionable findings to improve Kid Space experiences for next iterations and inform key design features for similar systems. Implications for practice and/or policyThe results of the study implied a set of areas for improvement-or design features-for Kid Space and other similar pedagogical conversational systems developed for children's home usages: (1) easier setup and usage with optimised setup size addressing diverse space limitations at homes, (2) minimised latency between Oscar (the conversational pedagogical agent) and child interactions (eg, adding multimodal dialogue system to reduce the need for a human wizard), (3) more advanced personalisation, social (including more verbal interactions) and pedagogical skills for Oscar with increased contextual awareness (eg, sending children's engagement), (4) scalability and higher visual quality of content with diverse games and learning outcomes, (5) parental control features over Kid Space platform and Oscar (eg, time limit, content, etc.) and (6) accessibility features (eg, captions turned on for multilingual children) and support for neurodiversity.
This paper proposes the RePAIR dataset that represents a challenging benchmark to test modern computational and data driven methods for puzzle-solving and reassembly tasks. Our dataset has unique properties that are uncommon to current benchmarks for 2D and 3D puzzle solving. The fragments and fractures are realistic, caused by a collapse of a fresco during a World War II bombing at the Pompeii archaeological park. The fragments are also eroded and have missing pieces with irregular shapes and different dimensions, challenging further the reassembly algorithms. The dataset is multi-modal providing high resolution images with characteristic pictorial elements, detailed 3D scans of the fragments and meta-data annotated by the archaeologists. Ground truth has been generated through several years of unceasing fieldwork, including the excavation and cleaning of each fragment, followed by manual puzzle solving by archaeologists of a subset of approx. 1000 pieces among the 16000 available. After digitizing all the fragments in 3D, a benchmark was prepared to challenge current reassembly and puzzle-solving methods that often solve more simplistic synthetic scenarios. The tested baselines show that there clearly exists a gap to fill in solving this computationally complex problem.
Lagoons are highly valued coastal environments providing unique ecosystem services. However, they are fragile and vulnerable to natural processes and anthropogenic activities. Concurrently, climate change pressures, are likely to lead to severe ecological impacts on lagoon ecosystems. Among these, direct effects are mainly through changes in temperature and associated physico-chemical alterations, whereas indirect ones, mediated through processes such as extreme weather events in the catchment, include the alteration of nutrient loading patterns among others that can, in turn, modify the trophic states leading to depletion or to eutrophication. This phenomenon can lead, under certain circumstances, to harmful algal blooms events, anoxia, and mortality of aquatic flora and fauna, or to the reduction of primary production, with cascading effects on the whole trophic web with dramatic consequences for aquaculture, fishery, and recreational activities. The complexity of eutrophication processes, characterized by compounding and interconnected pressures, highlights the importance of adequate sophisticated methods to estimate future ecological impacts on fragile lagoon environments. In this context, a novel framework combining Machine Learning (ML) and biogeochemical models is proposed, leveraging the potential offered by both approaches to unravel and modelling environmental systems featured by compounding pressures. Multi-Layer Perceptron (MLP) and Random Forest (RF) models are used (trained, validated, and tested) within the Venice Lagoon case study to assimilate historical heterogenous WQ data (i.e., water temperature, salinity, and dissolved oxygen) and spatio-temporal information (i.e., monitoring station location and month), and to predict changes in chlorophyll-a (Chl-a) conditions. Then, projections from the biogeochemical model SHYFEM-BFM for 2049, and 2099 timeframes under RCP 8.5 are integrated to evaluate Chl-a variations under future bio-geochemical conditions forced by climate change projections. Annual and seasonal Chl-a predictions were performed out by classes based on two classification modes established on the descriptive statistics computed on baseline data: i) binary classification of Chl-a values under and over the median value, ii) multi-class classification defined by Chl-a quartiles. Results from the case study showed as the RF successfully classifies Chl-a under the baseline scenario with an overall model accuracy of about 80% for the median classification mode, and 61% for the quartile classification mode. Overall, a decreasing trend for the lowest Chl-a values (below the first quartile, i.e. 0.85 µg/l) can be observed, with an opposite rising fashion for the highest Chl-a values (above the fourth quartile, i.e. 2.78 µg/l). On the seasonal level, summer remains the season with the highest Chl-a values in all scenarios, although in 2099 a strong increase in Chl-a is also expected during the spring one. The proposed novel framework represents a valuable approach to strengthen both eutrophication modelling and scenarios analysis, by placing artificial intelligence-based models alongside biogeochemical models.
Archaeological fragment processing is crucial to support the analysis of pictorial contents of broken artifacts. In this paper, we focus on the unexplored task of semantic segmentation of fresco fragments. This task enables the extraction of semantic information from a fragment, facilitating subsequent tasks like fragment classification or reassembly. We introduce a semantic segmentation dataset of fresco fragments acquired at the Pompeii Archeological Site, accompanied by baseline models. Additionally, we introduce a supplementary task of fragment cleaning, providing a dataset with the detection of manual annotations of archaeological marks that require restoration before further analysis. Our experiments, using standard metrics and state-of-the-art baselines, demonstrate that semantic segmentation of fresco fragments is feasible, paving the way toward more complex activities that require a semantic understanding of fragmented artifacts. Dataset with annotations, and code will be released at https://repairproject.github.io/fragment-restoration/
Drowsiness is a state of fatigue or sleepiness characterized by a strong urge to sleep. It is correlated with a progressive decline in response time, compromised processing of available information, more errors in short -term memory, and reduced vigilance behaviors. The electroencephalogram (EEG), a recording of the brain's electrical activities, has demonstrated the most robust association with drowsiness. As a result, EEG is widely recognized as a dependable indicator for evaluating drowsiness, fatigue, and performance levels. In this survey paper, we thoroughly investigate the application of shallow and deep neural network approaches utilizing EEG signals for the detection of fatigue and drowsiness. As far as our knowledge extends, this is the pioneering survey paper dedicated to exploring this specific research domain. The paper presents a comprehensive overview of the diverse EEG features utilized in the detection of fatigue and drowsiness, the different types of neural networks, and the reported performance of these methods in the literature. Additionally, the paper thoroughly examines the challenges and limitations associated with EEG-based fatigue and drowsiness detection and highlights directions for future research. The survey aims to offer a comprehensive overview of the existing methods in EEG-based fatigue and drowsiness detection, serving as a valuable resource for researchers and practitioners working in the respective field.
People in remote meetings in open spaces might choose to speak with a restrained voice due to concerns around privacy or disturbing others. Research shows that persons prefer to use soft voice (voice with lower amplitude and pitch, but with harmonic tones in its spectrum) over whispered voice (voice with the lowest amplitude, and no harmonics at all) to avoid being overheard during such calls. We present a lightweight classifier based in a simple feed-forward neural network, which uses normalized Log-Mel spectrum of voice captured by a headset as input, and can detect if the person is using soft voice. This allows to enhance soft voice with more precision and responsiveness than regular amplitude compensation ("auto-gain") systems. In this show and tell, we present a real-time demo of the voice classifier. Viewers will see our algorithm detect in real-time soft voice vs other voice types, in a regular PC, with voice captured with a headset.
This paper reports steps in probing the artistic methods of figurative painters through computational algorithms.We explore a comparative method that investigates the relation between the source of a painting, typically a photograph or an earlier painting, and the painting itself.A first crucial step in this process is to find the source and to crop, standardize and align it to the painting so that a comparison becomes possible.The next step is to apply different low-level algorithms to construct difference maps for color, edges, texture, brightness, etc. From this basis, various subsequent operations become possible to detect and compare features of the image, such as facial action units and the emotions they signify.This paper demonstrates a pipeline we have built and tested using paintings by a renowned contemporary painter Luc Tuymans.We focus in this paper particularly on the alignment process, on edge difference maps, and on the utility of the comparative method for bringing out the semantic significance of a painting.
Prior research has shown that computer-mediated communication introduces barriers to effective non-verbal communication. This creates potential challenges as workers increasingly collaborate remotely in a post-pandemic world. To further investigate these barriers, we conducted an exploratory study with five remote workers featuring qualitative user testing of concepts designed to facilitate non-verbal feedback using multi-modal AI technologies. Prior to the study, a team of UX designers created low-fidelity concept prototypes using different variables important for participant feedback (that is, modality, type, user control, visibility, granularity, timing, and categories). Using these prototypes, we conducted semi-structured interviews (90-min long) with scenario-based questions. The results showed that the participants perceived the prototypes for multi-modal feedback positively, especially for the multi-tasking scenarios. However, this perception differed based on different roles: Presenters versus listeners. In general, the participants indicated that the presenters would be more interested in feedback than the listeners, as long as that feedback was uncluttered and actionable. By contrast, when commenting as listeners (that is, as individuals whose feedback would be monitored), the participants expressed concerns over the accuracy of AI models, a sense of being monitored (privacy concerns), and potential social risks to provide uncomfortable feedback. As a future work, we plan to iterate on the current prototypes to address these concerns, and deploy a larger study with refined and more functional prototypes.
COVID‐19 has precipitated a massive social experiment – the sudden shift of millions of knowledge workers from their traditional offices to homes or other remote work locations. This has inspired heated debates and new ways of imagining the future of work. This paper hopes to contribute to a better understanding of these changes by reporting on the results of several dozen in‐depth interviews with remote workers from a variety of geographies, industries and professions. We focus in particular on their experiences of remote meetings, with special attention to complaints workers have with their current implementation. As we learned, workers' complaints tended to be driven by social – rather than productivity or technical – concerns. We explore this social dimension in depth, propose a framework for thinking about meetings as rituals, and suggest how this emphasis might inform the design of technology to support remote collaboration.
Since the spring of 2020, many early childhood education programs (pre-K, K, 1st, and 2nd grades) had to close as governments around the world took serious measures to slow down the transmission of COVID-19. As a result, the pandemic forced many early childhood teachers to start teaching online and continue supporting their students remotely. Unfortunately, there were few lessons that these teachers could learn from experience to cope with this change since online learning in early childhood settings had been scarce until the outbreak of the pandemic. In response, the goal of this interview study was to investigate how early childhood teachers in public and private schools implemented online learning during the pandemic, the challenges they encountered when teaching online, and their suggestions to address these challenges. The results showed that the teachers did not sit still and patiently wait for the re-opening of the schools. Instead, they took assorted initiatives to support their students’ learning and development remotely. They faced several challenges on the way but also suggested various methods to address these challenges through developmentally appropriate technology use. The results of this study have implications for teachers when early childhood programs return to normal. The study creates opportunities for future research to gain greater understanding of the design and implementation of online learning activities with young learners.
Recognizing the emotion an image evokes in the observer has long attracted the interest of the community for its many potential applications. However, it is a challenging task mainly due to the inherent complexity and subjectivity of human feelings. Such a difficulty is exacerbated in the domain of visual arts, mainly because of their abstract nature. In this work, we propose a new version of the artistic knowledge graph we were working on, namely Artgraph, obtained by integrating the emotion labels provided by the ArtEmis dataset. The proposed graph enables emotion-based information retrieval and knowledge discovery even without training a learning model. In addition, we propose an artwork emotion classification system that jointly exploits visual features and knowledge graph-embeddings. Experimental evaluation revealed that while improvements in emotion classification depend mainly on the use of visual features, the prediction of style, genre and emotion can benefit from the simultaneous exploitation of visual and contextual features and can assist each other in a synergistic way.
HCI tends to treat the humble office computer as a solved problem, yet most office workers still experience frustration when IT helpdesks need to be called. Why does this apparently “solved” problem persist? The software/hardware stack on a standard enterprise computer involves an astounding variety of possible drivers and application versions that can conflict with one another, leading to greater opportunity for breakdown, regardless of skills or resources of IT organizations. This circumstance lends itself to the use of telemetry and artificial intelligence (AI) for problem diagnosis and stokes aspirations of fully automating enterprise PC maintenance. To explore the human and organizational factors at work in applying data and AI to this problem, we designed a series of exploratory studies at a large technology company in the United States: (1) remote diary study with semi-structured interviews (n = 30), (2) quasi-experimental study with pretest-posttest design (n = 11), and (3) ethnographic study with open-ended interviews (n = 8). The results show that user frustration with malfunctioning PCs persisted because of the sociotechnical dynamic between employees, PCs, and IT support. Feedback loops between employees and IT played a central role in dialing up or tamping down frustration that accumulated over the long term. The results also indicate that telemetry and AI could provide new opportunities to tamp down user frustration when data were treated as a communication medium between employees and IT support. The results suggest three major design recommendations for preventing frustration buildup: (1) Redesigning PC telemetry data and transparency mechanisms to support two-way communication between IT and users, including shared analysis of malfunctioning data; (2) considering users’ buildup of frustration, not just the quality of any single interaction, when designing any IT service solutions; (3) incorporating uses of technology that embrace human-AI collaboration technologies not to automate IT troubleshooting work but to support the human creativity necessary for troubleshooting. Utilizing the design principles we identified in this study, there is a need for further research and development to explore novel feedback systems between enterprise PC users and IT.
Parents recognize the potential benefits of technology for their young children but are wary of too much screen time and its potential deficits in terms of social engagement and physical activity. To address these concerns, related literature suggests technology usages with a blend of digital and physical learning experiences. Towards this end, we developed Kid Space, incorporating immersive computing experiences designed to engage children more actively in physical movement and social collaboration during play-based learning. The technology features an animated peer learner, Oscar, who aims to understand and respond to children's actions and utterances using extensive multimodal sensing and sensemaking technologies. To investigate student engagement during Kid Space learning experiences, an exploratory case study was designed using a formative research method with eight first-grade students. Multimodal data (audio and video) along with observational, interview, and questionnaire data were collected and analyzed. The results show that the students demonstrated high levels of engagement, less attention focused on the screen (projected wall), and more physical activity. In addition to these promising results, the study also enabled us to understand actionable insights to improve Kid Space for future deployments (e.g., the need for real-time personalization). We plan to incorporate the lessons learned from this preliminary study and deploy Kid Space with real-time personalization features for longer periods with more students.