This paper investigates how ChatGPT, an AI chatbot developed by OpenAI, can be introduced to STEM education, specifically, an "Introduction to Cognitive Neuroscience" class. Using mixed-method research, the study conducted an experiment to collect students' performance scores and their feedback to examine the potential impacts from ChatGPT on their critical thinking skills, long-term retention of knowledge, and group learning interactions through collaborative projects. The results demonstrate significant disparities between AI-generated (ChatGPT) and human input. Upon analyzing the grade fluctuations before and after receiving input, it was shown that students who received feedback from ChatGPT encountered a more significant decrease (median = -12) compared to those who received feedback from humans (median = -5). This result indicates potential shortcomings in the effectiveness of AI feedback. Human feedback has significantly higher Retention Proxy scores than that of ChatGPT feedback, suggesting that human feedback can be potentially more effective in fostering long-term retention of course material. Investigation into collaboration dynamics of learning found that feedback given by humans tends to be more positive (63 vs. 40 in sentiment score) and also more focused on improvement and understanding. The theme of feedback is different between two conditions. That is, human feedback emphasizes scientific and detailed approaches while ChatGPT feedback emphasizes educational aspects and cognitive functions. These results indicate that while ChatGPT has potential benefits on educational settings, human feedback is superior in many ways. This study contributes to the continuing discourse on the use of AI in STEM education and emphasizes the significance of maintaining a harmonious combination of AI and human involvement in delivering educational feedback.
Adult education faces complex challenges beyond those in higher education. This preliminary study investigates the specific challenges perceived by adult learners and adult educators through a linguistic analysis. Using Natural Language Processing (NLP) techniques, including Linguistic Inquiry and Word Count (LIWC), we analyzed video-based interviews with 23 adult learners and 4 adult educators in an US institution. Our findings revealed that adult learners emphasize practical concerns such as learning methods, technology access, and time management, while adult educators focus on systemic issues such as policy and institutional support. The findings also indicated that adult educators showed greater emotional variability regarding systemic concerns, whereas adult learners expressed consistent emotional patterns toward immediate concerns. Therefore, this study highlights the importance of integrating adult learners' practical needs with adult educators' systemic insights to develop responsive adult education programs.
Ad designers often use sequences of shots in video ads, where frames are similar within a shot but vary across shots. These visual variations, along with changes in auditory and narrative cues, can interrupt viewers' attention. In this paper, we address the underexplored task of applying multimodal feature extraction techniques to marketing problems. We introduce the "AttInfaForAd" dataset, containing 111 baby product video ads with visual ground truth labels indicating points of interest in the first, middle, and last frames of each shot, identified by 75 shoppers. We propose attention interruption measures and use multimodal techniques to extract visual, auditory, and linguistic features from video ads. Our feature-infused model achieved the lowest mean absolute error and highest R-square among various machine learning algorithms in predicting shopper attention interruption. We highlight the significance of these features in driving attention interruption. By open-sourcing the dataset and model code, we aim to encourage further research in this crucial area. (Dataset and model code available at https://github.com/ostadabbas/Baby-Product-Video-Ads).
The Resilience Evaluation Scale (RES) is a newly developed measure of resilience written in both English and Dutch languages. To date, there have not been comprehensive psychometric evaluations of the RES’ performance, including validity for use in non-Western cultural populations and languages. In our attempt to address this void, we conducted a psychometric evaluation of the RES utilizing a Western, sample of U.S. college students and non-Western sample of Chinese college students. Our psychometric evaluation of the RES in a Western, English-speaking sample of U.S. college students indicated mixed results on the construct validity of the RES for measuring resilience. We also found that the factor structure of the RES lacked configural invariance across U.S. college student and Chinese college student samples. Results suggested that additional research is needed to assess whether the RES appropriately measures internal factors of resilience or requires modification. We also highlight the need for continued development of cross-culturally valid measures, and possibly different conceptualizations, of resilience across cultural and linguistic groups.
Automated human action recognition, a burgeoning field within computer vision, boasts diverse applications spanning surveillance, security, human-computer interaction, tele-health, and sports analysis. Precise action recognition in infants serves a multitude of pivotal purposes, encompassing safety monitoring, developmental milestone tracking, early intervention for developmental delays, fostering parent-infant bonds, advancing computer-aided diagnostics, and contributing to the scientific comprehension of child development. This paper delves into the intricacies of infant action recognition, a domain that has remained relatively uncharted despite the accomplishments in adult action recognition. In this study, we introduce a groundbreaking dataset called ``InfActPrimitive'', encompassing five significant infant milestone action categories, and we incorporate specialized preprocessing for infant data. We conducted an extensive comparative analysis employing cutting-edge skeleton-based action recognition models using this dataset. Our findings reveal that, although the PoseC3D model achieves the highest accuracy at approximately 71%, the remaining models struggle to accurately capture the dynamics of infant actions. This highlights a substantial knowledge gap between infant and adult action recognition domains and the urgent need for data-efficient pipeline models.
The emergence of artificial intelligence has incited a paradigm shift across the spectrum of human endeavors, with ChatGPT serving as a catalyst for the transformation of various established domains, including but not limited to education, journalism, security, and ethics. In the post-pandemic era, the widespread adoption of remote work has prompted the educational sector to reassess conventional pedagogical methods. This paper is to scrutinize the underlying psychological principles of ChatGPT, delve into the factors that captivate user attention, and implicate its ramifications on the future of learning. The ultimate objective of this study is to instigate a scholarly discourse on the interplay between technological advancements in education and the evolution of human learning patterns, raising the question of whether technology is driving human evolution or vice versa.
Automatic detection of infant actions from home videos could aid medical and behavioral specialists in the early detection of motor impairments in infancy. However, most computer vision approaches for action recognition are centered around adult subjects, following datasets and benchmarks in the field. In this work, we present a data-efficient pipeline for infant action recognition based on the idea of modeling an action as a time sequence consisting of two different stable postures with a transition period between them. The postures are detected frame-wise from the estimated 2D and 3D infant body poses and the action sequence is segmented based on the posture-driven low-dimensional features of each frame. To spur further research in the field, we also created and release the first-of-its-kind infant action dataset—InfAct—consisting of 200 fully annotated home videos representing a wide range of common infant actions, intended as a public benchmark. Among the ten more common classes of infant actions, our action recognition model achieved 78.0% accuracy when tested on InfAct, highlighting the promise of video-based infant action recognition as a viable monitoring tool for infant motor development 1 .
Cutting (2021) argues that the narrational complexity of fiction film can be quantified similarly to computational measures of text complexity. Narrational complexity refers to the structure that arises from how a story is told. This article expands upon Cutting's proposal by taking inspiration from contemporary approaches for measuring text complexity. These approaches reject the notion that complexity can be measured via a limited set of indices as Cutting proposed for narrational complexity. Similarly, we argue that narrational complexity for fiction films should be multi-dimension and include indices that are associated with events, characters, and the rules that govern the fictional world. We discuss the viability of using computational approaches to analyze video and natural language processing to develop approaches to measuring narrational complexity.
Bilateral postural symmetry plays a key role as a potential risk marker for autism spectrum disorder (ASD) and as a symptom of congenital muscular torticollis (CMT) in infants, but current methods of assessing symmetry require laborious clinical expert assessments. In this paper, we develop a computer vision based infant symmetry assessment system, leveraging 3D human pose estimation for infants. Evaluation and calibration of our system against ground truth assessments is complicated by our findings from a survey of human ratings of angle and symmetry, that such ratings exhibit low inter-rater reliability. To rectify this, we develop a Bayesian estimator of the ground truth derived from a probabilistic graphical model of fallible human raters. We show that the 3D infant pose estimation model can achieve 68% area under the receiver operating characteristic curve performance in predicting the Bayesian aggregate labels, compared to only 61% from a 2D infant pose estimation model and 60% from a 3D adult pose estimation model, highlighting the importance of 3D poses and infant domain knowledge in assessing infant body symmetry. Our survey analysis also suggests that human ratings are susceptible to higher levels of bias and inconsistency, and hence our final 3D pose-based symmetry assessment system is calibrated but not directly supervised by Bayesian aggregate human ratings, yielding higher levels of consistency and lower levels of inter-limb assessment bias 1 .
STEM learning aims to prepare students with hands-on and problem-based learning. However, teacher-centered instruction has been the predominant course delivery technique in STEM education regardless face-to-face or online learning context. Using both quantitative and qualitative research methods, this study explores the expectations of effective online courses based on Moore’s three types of interactions among Chinese STEM college students taking synchronous teacher-centered lecture-based online courses. A total of 175 undergraduate STEM students were recruited at one Chinese university. Results indicate that these students expect their instructors to integrate activities to motivate interactions with their instructor, peers, and the learning content. Students’ perceptions of the advantages and challenges of taking synchronous lecture-based courses are also discussed. It is expected that the findings would enlighten professionals of higher education in China to adjust teacher-centered instruction and to adequately prepare and train online instructors to foster an active online learning environment in STEM fields.
We lay the groundwork for research in the algorithmic comprehension of infant faces, in anticipation of applications from healthcare to psychology, especially in the early prediction of developmental disorders. Specifically, we introduce the first-ever dataset of infant faces annotated with facial landmark coordinates and pose attributes, demonstrate the inadequacies of existing facial landmark estimation algorithms in the infant domain, and train new state-of-the-art models that significantly improve upon those algorithms using domain adaptation techniques. We touch on the closely related task of facial detection for infants, and also on a challenging case study of infrared baby monitor images gathered by our lab as part of in-field research into the aforementioned developmental issues 1
This paper investigates how interdisciplinary research impacts the film industry in research and practice by introducing psychological concepts. Psychology, especially neural and cognitive science, provides a distinct advantage when examining humans’ audio-visual processing mechanisms and esthetics questions regarding the film. By introducing psychology, film researchers and filmmakers could rethink and evaluate the current research paradigm from a broader point of view. This paper consists of three parts: (1) a discussion on the nature of film using an interdisciplinary approach; (2) a discussion on the characteristics and attributes of film; (3) an introduction of the psychological concept of “affordance” to film studies and practice. Although the film interdisciplinary research paradigm is still under development, we argue that introducing the other subjects is innovating the field of film research, providing us with a new angle to examine the intersections of ubiquitous but complex human esthetics activities.
暑期档对于电影行业是非常重要的档期.因为它不仅仅是全年中时间跨度最长的档期,也是优秀电影问世以及票房创新的档期.本文试图从跨学科角度,将电影认知科学和人工智能中的自然语言处理结合,分析2021年暑假档中10部电影的表现,解读对应的原因.
近二十年来,由大卫·波德维尔发起的带有跨学科性质的电影感知研究,将电影视听研究推向了一个新的实证研究领域,从过去单纯的电影视听研究转到动态连锁研究,即,视听设计如何影响观众的心理以及喜好.目前该领域的大部分研究者都来自于心理学和计算机科学,由于缺乏电影理论和实践经验,对于电影视听的设计研究部分没有触及到视听语言的核心.本文利用电影感知研究的核心理论——事件分割研究和记忆研究,采用实验方法,对李安的五部成功作品从镜头长度、事件分割、叙事结构等方面进行观众反馈数据收集、分析和研究.实验结果表明,影片的镜头长度和事件分割符合人类感知的特性和记忆规律,在叙事结构方面,电影镜头长度和事件长度分配合理,容易被观众理解,进而产生情感上的共鸣.本实验结果对电影视听特性的实证研究和电影感知理论在实践中的应用具有一定意义.
面对经济和文化等方面的冲击,中国电影研究需要顺势而为:从传统的理论性研究汇入创作实践与理论体系交融的大方向,从单一的学科研究转向到跨学科的探索.电影感知研究正是解决电影跨学科问题的最佳研究方法,它将传统电影研究与其他学科的现实经验相结合,可以从根本上推进中国电影研究以及实践在世界电影产业舞台上的位置.
电影感知研究, 是欧美近20年来兴起的跨学科实证研究.电影感知研究主要以心理学的感知科学为基础, 并且结合电影学、传播学、计算机科学、统计学等学科, 以实证的方法探讨电影的构建与观众解读机制之间的联系.文章通过对电影感知研究历史的梳理和总结, 探讨了传统电影研究中忽略的问题, 讨论了电影感知研究的学科性质和应用前景.