To support the integration of AI in education, this empirical study investigated what lessons college students learned from using Generative AI for writing. We recruited 47 students in the United States from a university writing course. Students completed an assignment in which they used Generative AI tools (e.g., ChatGPT) to draft an application letter or personal statement. Data were collected using a survey of five open-ended questions about their writing process, what worked, what did not work, how to better write with AI, and general lessons learned. We applied thematic analysis and sentiment analysis methods to analyze students’ responses. Results show that (1) students went through multiple rounds of prompting; (2) students identified strengths of AI, such as connection to topic, template generation, and sentence quality; (3) the weaknesses of AI included general language, robotic tone and lacking emotion, lacking personal voice, and lacking critical thinking; (4) students wished to improve AI-generated writing by adding personal stories, connections to posting, feelings and thoughts, and deleting repetitive language; and (5) their overall attitudes toward AI tool were positive. We believe our findings can help relieve some concerns about cheating with AI. We also suggested strategies to regulate the use of AI.
Assessing teams and providing feedback on scenario-based training typically requires human observers or scenario-specific metrics crafted by experts, due to the complexity of general-purpose automated tools to assess team performance. Machine learning can help infer team performance patterns, but labeled data for a specific training scenario is often sparse. To address this issue, the Semi-Supervised Learning for Assessing Team Simulations (SLATS) project investigated the feasibility of semi-supervised learning and transfer learning which leverages training data from related scenarios to classify performance on a target scenario with the same metrics but a different terrain context. To this approach, we analyzed performance of teams in the first-person shooter Team Fortress 2 (TF2). TF2 teams for the “Capture Point” mode were classified into archetypes based on the performance of the team and the performance of individual members of the team across the corpus: novice, weak link, team of experts, and expert team. To investigate the feasibility of transfer learning, we isolated matches from two of the most frequent maps/terrains. Results found that leveraging data from the source map always improved classification F1-scores compared to relying solely upon target (test) map training data. The greatest benefits were observed when target data was limited (0 to 42 target examples). While further research is required to explore the effectiveness of transfer learning across training scenarios that are more dissimilar (e.g., different simulations, rather than just different maps), these results offer a promising direction to help bootstrap team assessments on new training scenarios by leveraging data from earlier, comparable scenarios. However, efficiently calculating reusable metrics for model features based on low-level scenario events and logs remains a challenge that requires further research.
An important prerequisite to trend-aware authoring is that scenarios be authorable and inspectable by instructors but also machine-readable such that authoring tools can assist with integrating real-world patterns into training. In this research, we use a semi-structured approach to authoring flight training scenarios in which textual descriptions of related scenario elements (i.e., happening at roughly the same time) are grouped together and assigned training objectives and phases of flight. This same representation can be used to represent real-world emergencies allowing their integration into scenarios for more realistic training. Such a representation is sufficient to support a recommender that ranks possible insertion points for real-world emergencies using constraints (i.e., the phase of flight of the emergency must match the phase of flight of the insertion point) and a ranking score. Our ranking score is currently based on matching training objectives associated with the emergency with training objectives in the scenario (i.e., training the same skills but using a more realistic example). The recommender is integrated into the scenario editor such that instructors can see the ranked injection points and modify the scenario by selecting one of these points.
This report is an invitation for educators, policymakers, technologists, and learners to consider how generative AI can contribute to the future of education?. It aims to lay down a foundation upon which we can start building an educational ecosystem that is dynamic, inclusive, and profoundly human, despite being significantly aided by artificial intelligence?.
The emergence of widely-used artificial intelligence (AI) has created a critical need for AI expertise, not just as a research area but for workers in the wide variety of careers and roles that AI disrupts. While AI is still an area of research for new processing, application, and development – it continues to partially automate, augment, or replace many of the tasks which are performed through active use of human hands. While recently publicized items such as ChatGPT and MidJourney have made press in their adjustment to writing and image generation technology, the basic workflow of copyeditors and digital artists was completely transformed, inside of the year, to a combination of partially automated or fully automated AI tasks. While some blame AI as part of the “problem”, it is naturally part of the “solution” – AI tools to help workers develop AI competencies. The paper describes an array of strategies which the DoD and its ICT UARC are using to address the fundamental problem of quickly upskilling the DoD workforce of over 2 million adult learners.
Results from Randomized Controlled Trials (RCTs) establish the comparative effectiveness of interventions, and are in turn critical inputs for evidence-based care. However, results from RCTs are presented in (often unstructured) natural language articles describing the design, execution, and outcomes of trials; clinicians must manually extract findings pertaining to interventions and outcomes of interest from such articles. This onerous manual process has motivated work on (semi-)automating extraction of structured evidence from trial reports. In this work we propose and evaluate a text-to-text model built on instruction-tuned Large Language Models (LLMs) to jointly extract Interventions, Outcomes, and Comparators (ICO elements) from clinical abstracts, and infer the associated results reported. Manual (expert) and automated evaluations indicate that framing evidence extraction as a conditional generation task and fine-tuning LLMs for this purpose realizes considerable ($\sim$20 point absolute F1 score) gains over the previous SOTA. We perform ablations and error analyses to assess aspects that contribute to model performance, and to highlight potential directions for further improvements. We apply our model to a collection of published RCTs through mid-2022, and release a searchable database of structured findings: http://ico-relations.ebm-nlp.com
Mentoring promotes underserved students' persistence in STEM but is difficult to scale up. Conversational virtual agents can help address this problem by conveying a mentor's experiences to larger audiences. The present study examined college students' $$(N = 138)$$ utilization of CareerFair.ai, an online platform featuring virtual agent-mentors that were self-recorded by sixteen real-life mentors and built using principles from the earlier MentorPal framework. Participants completed a single-session study which included 30 min of active interaction with CareerFair.ai, sandwiched between pre-test and post-test surveys. Students' user experience and learning gains were examined, both for the overall sample and with a lens of diversity and equity across different, potentially underserved demographic groups. Findings included positive pre/post changes in intent to pursue STEM coursework and high user acceptance ratings (e.g., expected benefit, ease of use), with under-represented minority (URM) students giving significantly higher ratings on average than non-URM students. Self-reported learning gains of interest, actual content viewed on the CareerFair.ai platform, and actual learning gains were associated with one another, suggesting that the platform may be a useful resource in meeting a wide range of career exploration needs. Overall, the CareerFair.ai platform shows promise in scaling up aspects of mentoring to serve the needs of diverse groups of college students.
Paleoart is an important medium that communicates scientific understanding about prehistoric life to both the public and researchers. However, despite its broad influence, the scientific and aesthetic decisions that go into paleoart are rarely described in formal academic literature or subjected to peer review. This is unfortunate, as paleoart can easily create and perpetuate misconceptions that are carried through generations of iterative popular media. As an example of what we hope will become a standard article type in paleontological journals, we describe the process and latest scientific research used to develop 13 new paleoart reconstructions of Ice Age animals found in the La Brea Tar Pits, including the saber-toothed cat, dire wolf, and teratorn. We adopted a stylized low polygon aesthetic for these three-dimensional (3D), animated virtual models both to support learning objectives and to optimize performance for smartphone based augmented reality (AR) experiences. We encourage all researchers to follow the example of this article by publishing paleoart descriptions for any major new work that, at a minimum, reference the aesthetic and scientific reasoning behind general posture and proportions, gross appearance of soft tissues, coloration, and behavior.
Informal learning environments, such as museums, provide unique opportunities for science learning. They are deliberately designed to impact public understanding of science and shape visitors' attitudes and behaviors. As a developing technology, augmented reality (AR) offers the transformative potential to support museums' educational missions by enhancing visitors' experience, thereby creating effective conditions for learning and personalized interactions with science. We implemented an AR-enhanced exhibit at the La Brea Tar Pits (LBTP) to reduce scientific misconceptions and explore the role of interest and emotions around science and AR technology as it related to learning and knowledge revision. Using a pretest-posttest design, 62 adults completed an AR experience that addressed two scientific misconceptions related to the consistency of tar and frequency of large animal entrapment. We found that participants had significantly fewer misconceptions at posttest than at pretest. Participants also reported higher levels of interest in science content than AR technology and discriminated between emotions they experienced with regard to science content and AR technology. Feelings of curiosity predicted knowledge revision and interest in both science content and AR technology. These findings may be useful for museums and other science communicators seeking to create AR interventions that support learning and conceptual change.
Despite strong evidence that dialog-based intelligent tutoring systems (ITS) can increase learning gains, few courses include these tutors. In this research, we posit that existing dialog-based tutoring systems are not widely used because they are too complex and unfamiliar for a typical teacher to adapt or augment. OpenTutor is an open-source research project intended to scale up dialog-based tutoring by enabling ordinary teachers to rapidly author and improve dialog-based ITS, where authoring is presented through familiar tasks such as assessment item creation and grading. Formative usability results from a set of five non-CS educators are presented, which indicate that the OpenTutor system was relatively easy to use but that teachers would closely consider the cost benefit for time vs. student outcomes. Specifically, while OpenTutor grading was faster than expected, teachers reported that they would only spend any additional time (compared to a multiple choice) if the content required deeper learning. To decrease time to train answer classifiers, OpenTutor is investigating ways to reduce cold-start problems for tutoring dialogs.
Games and simulations can be more engaging than other educational tools (e.g., textbooks, videos, problem sets), and this engagement can lead to improved short- and long-term learning. However, engagement in game-based learning is not automatic, and instead requires iterative design. In this work, we explore and compare metrics from research on learning sciences and from game design, considering different time scales of human action, ranging from biological engagement (e.g., eye gaze) up to lasting social ties (e.g., community building). Certain game-design approaches used for commercial games may be useful for game-based learning, such as establishing bottom-line metrics aligned to why the game was built or analyzing engagement in terms of facets or archetypes rather than on a unidirectional scale. Further research is required to study the interaction between engagement at different time scales, particularly for cases where higher long-term engagement is indicated by lower short-term engagement (e.g., skipping easy content).
Despite the critical role of teachers in the educational process, few advanced learning technologies have been developed to support teacher-instruction or professional development. This lack of support is particularly acute for middle school math teachers, where only 37% felt well prepared to scaffold instruction to address the needs of diverse students in a national sample. To address this gap, the Advancing Middle School Teachers’ Understanding of Proportional Reasoning project is researching techniques to apply pedagogical virtual agents and dialog-based tutoring to enhance teachers' content knowledge and pedagogical content knowledge. This paper describes the design of a conversational, agent-based intelligent tutoring system to support teachers' professional development. Pedagogical strategies are presented that leverage a virtual human facilitator to tutor pedagogical content knowledge (how to teach proportions to students), as opposed to content knowledge (understanding proportions). The roles for different virtual facilitator capabilities are presented, including embedding actions into virtual agent dialog, open-response versus choice-based tutoring, ungraded pop-up sub-activities (e.g. whiteboard, calculator, note-taking). Usability feedback for a small cohort of instructors pursuing graduate studies was collected. In this feedback, teachers rated the system ease of use and perceived usefulness moderately well, but also reported confusion about what to expect from the system in terms of flow between lessons and support by the facilitator.
Adapting training in real time can be challenging for instructors. Real-time simulation can present rapid sequences of events, making it difficult for an instructor to attribute errors or omissions to specific underling gaps in skills and knowledge. Monitoring multiple students simultaneously imposes additional attentional workload on an instructor. This challenge can be further exacerbated when an instructor’s view of the student is obscured by virtual reality (VR) equipment. To support instructors’ ability to adapt training, Eduworks and USC’s Institute for Creative Technologies are developing machine learning (ML) models that can measure user engagement during training simulations and offer recommendations for restoring lapses in engagement. We have created a system, called the Observational Motivation and Engagement Generalized Appliance (OMEGA), which we tested in the context of a new U.S. Air Force approach to Specialized Undergraduate Pilot Training (SUPT) called Pilot Training Next (PTN). PTN integrates traditional flying sorties with VR-enabled ground-based training devices to achieve training efficiencies, improve readiness, and increase throughput. The virtual environment provides a rich source of raw data that machine learning models can use to associate user activity with user engagement. We created a testbed for data capture to construct the ML models, based on theoretical foundations we developed previously. Our research explores OMEGA’s potential to help alert an instructor pilot (IP) to student distraction by flagging attention and engagement lapses. Our hypothesis is that OMEGA could help an IP adapt learning, and potentially manage multiple students at the same time, with alerts of lapsed attention and recommendations for restoring engagement. To test this hypothesis, we ran pilots through multiple PTN scenarios to create data for training the model. In this paper, we report on work to create machine learning models using three different techniques, and present model performance data using standard machine learning metrics. We discuss the modeling approach used to generate instructor recommendations. Future work will present results from a formative evaluation using instructor pilots. These early findings provide preliminary validation for the use of ML models for learning to detect engagement from the rich data sources characteristic of virtual environments. These findings will be applicable across a broad range of conventional and VR training applications.
Introduction Ideally, health conditions causing the greatest global disease burden should attract increased research attention. We conducted a comprehensive global study investigating the number of randomised controlled trials (RCTs) published on different health conditions, and how this compares with the global disease burden that they impose.Methods We use machine learning to monitor PubMed daily, and find and analyse RCT reports. We assessed RCTs investigating the leading causes of morbidity and mortality from the Global Burden of Disease study. Using regression models, we compared numbers of actual RCTs in different health conditions to numbers predicted from their global disease burden (disability-adjusted life years (DALYs)). We investigated whether RCT numbers differed for conditions disproportionately affecting countries with lower socioeconomic development.Results We estimate 463 000 articles describing RCTs (95% prediction interval 439 000 to 485 000) were published from 1990 to July 2020. RCTs recruited a median of 72 participants (IQR 32–195). 82% of RCTs were conducted by researchers in the top fifth of countries by socio-economic development. As DALYs increased for a particular health condition by 10%, the number of RCTs in the same year increased by 5% (3.2%–6.9%), but the association was weak (adjusted R2=0.13). Conditions disproportionately affecting countries with lower socioeconomic development, including respiratory infections and tuberculosis (7000 RCTs below predicted) and enteric infections (9700 RCTs below predicted), appear relatively under-researched for their disease burden. Each 10% shift in DALYs towards countries with low and middle socioeconomic development was associated with a 4% reduction in RCTs (3.7%–4.9%). These disparities have not changed substantially over time.Conclusion Research priorities are not well optimised to reduce the global burden of disease. Most RCTs are produced by highly developed countries, and the health needs of these countries have been, on average, favoured.