Player engagement is crucial for understanding and optimizing gaming experiences, yet the research community lacks comprehensive multimodal datasets with reliable engagement annotations. We present a dataset combining six synchronized data streams—EEG, eye tracking, heart rate, user inputs, webcam footage, and gameplay frames—collected from 39 participants playing popular games across varying difficulty levels. Our dataset’s distinctive feature lies in its temporal precision, achieved through strategic integration of engagement surveys during natural game pauses, minimizing both recall bias and gameplay disruption. The dataset includes 900 annotated gameplay sessions with four psychological metrics (engagement, interest, stress, excitement). Initial analyses revealed surprising findings: human judges achieved only 0.48 F1-score in engagement assessment from webcam footage, while a flow theory-based model reached 0.60 F1-score using difficulty and player experience. Our multimodal neural model combining EEG, eye tracking, and facial features demonstrated the dataset’s potential with a 0.51 F1-score despite class imbalance. This comprehensive dataset enables various research directions in engagement measurement and modeling, supporting the development of more robust real-time engagement detection systems.
Player engagement is crucial for the success of modern video games, yet its real-time measurement remains challenging due to the intrusive nature of traditional measurement methods. In this article, we present a novel framework for nonintrusive, real-time, and indirect measurement of engagement in multiplayer online games based on flow theory. Our approach combines graph convolutional networks for modeling player interactions with Transformer networks for temporal processing, enabling indirect measurement of both player skill and game challenge, which in turn are used to classify player engagement. Using playerunknown’s battlegrounds (PUBGs) as a case study, we demonstrate that our framework can effectively measure phase-specific engagement using one minute of gameplay telemetry data. Our framework achieves 73% accuracy and 0.83 ROC-AUC in engagement classification, matching the performance of traditional survey-based methods while operating nonintrusively and in real time. Further cross-domain validation of the framework, as is and without transfer learning, with the games FIFA’23 and Street Fighter V, leads to 66% accuracy, demonstrating the model’s stable performance despite the significant differences in the test domains. Interestingly, our results suggest that objective gameplay metrics may better reflect engagement than subjective player assessments, with skill estimates showing significant correlation with self-reports.
This article presents a review on the process of estimating player engagement in video gaming. To stay ahead of their competitors in entertainment, game developers need to understand, estimate, and maximize player engagement. We address the multidimensional nature of engagement, encompassing cognitive, emotional, and behavioral aspects across various gaming domains. We present a taxonomy of the diverse modalities for quantifying engagement, including physiological signals, observable behaviors, and gameplay data. We identify the challenges of conducting representative subjective studies in this domain and summarize various methods for establishing ground truth measurements. By synthesizing existing research, we provide insights into modeling techniques, highlight research gaps, and offer practical guidelines for implementing engagement measurement strategies. This review aims to aid researchers and industry professionals in navigating the complexities of player engagement estimation, ultimately contributing to enhanced game design, marketing, and user retention in the competitive gaming landscape.
Adopting Artificial Intelligence (AI) systems in measurement instruments and systems entails a necessity to predict the error contributed by the AI model to the measured value, especially on out-of-sample data. However, reporting aggregated error estimates, such as model accuracy or Root Mean Square Error (RMSE) as is customary in AI, cannot quantify the error of the AI model for a single measurement instance, which is what we need in measurement. In this paper, we propose a novel method to estimate the AI model’s error for a single measurement. Our goal is to predict the error and use it to correct the predicted measurand’s quantity. To do so, in the first step we use an existing dataset to train an AI model that predicts measurement values, which is the usual approach for designing AI-assisted measurement systems. Our contribution is in the second step, where we create a secondary dataset that consists of the errors between the ground truth and the predicted values, and we use this dataset to train a secondary AI model that predicts the errors. We then adjust the predicted measurement values with the predicted error values. Our performance evaluations on the well-known California Housing dataset shows that our approach lowers the measurement predictions’ Mean Absolute Percentage Error from 21% to 16%, resulting in more accurate measurements.
On June 24, 2018, Turkey conducted a highly consequential election in which the Turkish people elected their president and parliament in the first election under a new presidential system. During the election period, the Turkish people extensively shared their political opinions on Twitter. One aspect of polarization among the electorate was support for or opposition to the reelection of Recep Tayyip Erdoğan. In this paper, we present an unsupervised method for target-specific stance detection in a polarized setting, specifically Turkish politics, achieving 90% precision in identifying user stances, while maintaining more than 80% recall. The method involves representing users in an embedding space using Google's Convolutional Neural Network (CNN) based multilingual universal sentence encoder. The representations are then projected onto a lower dimensional space in a manner that reflects similarities and are consequently clustered. We show the effectiveness of our method in properly clustering users of divergent groups across multiple targets that include political figures, different groups, and parties. We perform our analysis on a large dataset of 108M Turkish election-related tweets along with the timeline tweets of 168k Turkish users, who authored 213M tweets. Given the resultant user stances, we are able to observe correlations between topics and compute topic polarization.
Detecting offensive language on Twitter has many applications ranging from detecting/predicting bullying to measuring polarization. In this paper, we focus on building a large Arabic offensive tweet dataset. We introduce a method for building a dataset that is not biased by topic, dialect, or target. We produce the largest Arabic dataset to date with special tags for vulgarity and hate speech. We thoroughly analyze the dataset to determine which topics, dialects, and gender are most associated with offensive tweets and how Arabic speakers use offensive language. Lastly, we conduct many experiments to produce strong results (F1 = 83.2) on the dataset using SOTA techniques.
On June 24, 2018, Turkey conducted a highly-consequential election in which the Turkish people elected their president and parliament in the first election under a new presidential system. During the election period, the Turkish people extensively shared their political opinions on Twitter. One access of polarization among the electorate was support for or opposition to the reelection of Recep Tayyip Erdogan. In this paper, we explore the polarization between the two groups on their political opinions and lifestyle, and examine whether polarization had increased in the lead up to the election. We conduct our analysis on two collected datasets covering the time periods before and during the election period that we split into pro- and anti-Erdogan groups. For the pro and anti splits of both datasets, we generate separate word embedding models, and then use the four generated models to contrast the neighborhood (in the embedding space) of the political leaders, political issues, and lifestyle choices (e.g., beverages, food, and vacation). Our analysis shows that the two groups agree on some topics, such as terrorism and organizations threatening the country, but disagree on others, such as refugees and lifestyle choices. Polarization towards party leaders is more pronounced, and polarization further increased during the election time.
We study the evolution of our university’s social networks over time, capturing direct, contextual, and latent changes in these networks. With the assumption of our university’s social dynamics being embodied in the networks we construct, we continuously monitor these networks in order to gain an understanding of the changes they go through and their evolution. Our system has three main components: (i) crawling the web for collecting data, (ii) networked data analysis, and (iii) data storytelling. Our goal is to render the social development of our university as a community in a lucid and insightful manner.
Speech impediment affecting children with hearing difficulties and speech disorders requires speech therapy and much practice to overcome. To motivate the children to practice more, serious games can be used because children are more inclined to play games. In this paper, we have designed and implemented a serious game in which children can learn to speak specific words that they are expected to know before the age of 7. The game consists of an avatar controlled by the child through speech, with the objective of moving the avatar around the environment to earn coins. The avatar is controlled by voice commands such as Jump, Ahead, Back, Left, Right. Children will be guided by an arrow during the game instead of a getting help from a therapist or a teacher to guide the child to the next coin. This allows the child to practice longer hours, compared to clinical approaches under the supervision of a therapist, which are time-limited.
Visual Sequential Memory (VSM) allows a person to perform tasks such as remembering letters, numbers, objects or shapes in the correct order. Its deficit can lead to challenges in one's personal life, including dyslexia and dyscalculia. Detecting Visual Sequential Memory Deficit (VSMD) is essential for those who suffer from its related consequences. But current clinical methods don't have a high rate of diagnosis, and also for treatment are limited to the few hours the person spends in the clinic. In this paper, we propose an Origami based Serious Game, called Memori, for the diagnosis and treatment of children with VSMD. We illustrate the rationale behind using Origami, the design process of our game, and its implementation.
When performing tasks such as remembering letters, numbers, objects, or shapes, a persons Visual Sequential Memory (VSM) plays a crucial role, especially when the order of the tasks is important. Lack of VSM makes the persons life more challenging, possibly leading to dyslexia and dyscalculia. As such, it is important to detect and treat Visual Sequential Memory Deficit (VSMD). But current clinical methods have a low rate of diagnosis, and also offer limited hours to persons being treated in the clinics. In this paper, we propose an Origami based Serious Game, called Memori, as a synthetic instrument for the diagnosis, performance measurement, and treatment of people with VSMD. We illustrate the rationale behind using Origami, the design process of our game, and its implementation. Our preliminarily performance evaluations with 24 adults reveal a 13% improvement of memory and 1.00 score increase in performance while a slight decrease occurred in attentiveness from 2.33 to 2.02 for people who use our tool.
In this paper, we describe our efforts at OSACT Shared Task on Offensive Language Detection. The shared task consists of two subtasks: offensive language detection (Subtask A) and hate speech detection (Subtask B). For offensive language detection, a system combination of Support Vector Machines (SVMs) and Deep Neural Networks (DNNs) achieved the best results on development set, which ranked 1st in the official results for Subtask A with F1-score of 90.51% on the test set. For hate speech detection, DNNs were less effective and a system combination of multiple SVMs with different parameters achieved the best results on development set, which ranked 4th in official results for Subtask B with F1-macro score of 80.63% on the test set.
Shervin Shirmohammadi合作论文数University of Ottawa;School of Information Technology and Engineering (SITE)7