Generative AI (GenAI), specifically the Large Language Model (LLM), focuses on creating content like text, images, audio, or other data types similar to that created by humans. Advancements in various LLMs have accelerated AI development, potentially impacting various fields, including education. Can LLMs produce educational content with similar readability features as human-generated reading passages? This study compared readability measures such as reading speed, comprehension, and qualitative characteristics (familiarity, interest, and perceived quality) of 300-word 8th-grade level text passages authored by humans and ChatGPT3.5. We found that ChatGPT3.5-generated passages can be read faster with better comprehension than conventionally human-authored passages. ChatGPT3.5 passages were also rated higher-quality passages and had comparable ratings for interest and familiarity compared to human-authored passages. These findings implicate that the LLM enhances educational materials’ usability and effectiveness. More importantly, LLM offers scalability and consistency, catering to diverse learning needs for educators and students.
In our age of ubiquitous digital displays, adults often read in short, opportunistic interludes. In this context of Interlude Reading , we consider if manipulating font choice can improve adult readers’ reading outcomes. Our studies normalize font size by human perception and use hundreds of crowdsourced participants to provide a foundation for understanding, which fonts people prefer and which fonts make them more effective readers. Participants’ reading speeds (measured in words-per-minute (WPM)) increased by 35% when comparing fastest and slowest fonts without affecting reading comprehension. High WPM variability across fonts suggests that one font does not fit all. We provide font recommendations related to higher reading speed and discuss the need for individuation, allowing digital devices to match their readers’ needs in the moment. We provide recommendations from one of the most significant online reading efforts to date. To complement this, we release our materials and tools with this article.
Our prior work shows that individual typeface selection per reader can significantly increase reading speed while maintaining comprehension. The present work extends that work to consider the influence of character and word spacing on adult reading performance in a large online study. In Study I, 102 remote Amazon Mechanical Turk participants (41 W, 61 M; ages 18--72, mean = 38.4), read passages with 11 different character spacings from -0.1em to 0.4em (increments of 0.05em). Each participant read all passages in a single font, randomly selected from a pool of six. In Study II, 113 participants (51 W, 60 M, 2 O; ages 21--72, mean = 35.2) followed similar study procedures to read passages with 8 different inter-word spacings (-0.2, -0.1, 0, 0.1, 0.2, 0.5, 1, 1.5 em). Across both studies, participants read 8th-grade level passages of 160--178 words, split across two screens. Consistent with similar remote studies, the second screen of each passage was read faster than the first (mean differences of 38 WPM and 29 WPM in Study I & II, respectively). Overall, character spacing had stronger effects on reading speed than word spacing, which may support the word-level theory of reading. However, the minimum negative and maximum positive spacings (character -0.1em and 0.4em; word -0.2em and 1.5em) showed significant speed, comprehension, and self-reported preference drop-offs. Faster readers suffered greater slowdowns at wide spacing (r = -0.38, p < .001). Character spacing significantly impacted reading speed on the second screen of reading, but not the first (X2(10) = 32.7, p < 0.001 and X2(10) = 9.7, N.S., respectively ). Both findings point to increased spacing hindering faster reading, complementing prior research showing wider spacing may help struggling readers. Together, these findings join a growing body of work supporting readability interventions at the individual level.
The amount of text people need to read and understand grows daily. Software defaults, designers, or publishers often choose the fonts people read in. However, matching individuals with a faster font could help them cope with information overload. We collaborated with typographers to (1) select eight fonts designed for digital reading to systematically compare their effectiveness and to (2) understand how font and reader characteristics affect reading speed. We collected font preferences, reading speeds, and characteristics from 252 crowdsourced participants in a remote readability study. We use font and reader characteristics to train FontMART, a learning to rank model that automatically orders a set of eight fonts per participant by predicted reading speed. FontMART’s fastest font prediction shows an average increase of 14–25 WPM compared to other font defaults, without hindering comprehension. This encouraging evidence provides motivation for adding our personalized font recommendation to future interactive systems.
Voice interfaces reduce visual demand compared with visual-manual interfaces, but the extent depends on design. This study compared visual demand during baseline driving with driving while using voice or manual inputs to place calls with Chevrolet MyLink, Volvo Sensus, or a smartphone. Mean glance duration and total eyes-off-road-time increased when using manual input compared with baseline driving; only eyes off road time increased with voice input. Confusion matrices developed with hidden Markov modelling characterise the similarity of glance sequences during baseline driving and while making phone calls. Glance sequences with the MyLink voice interface were misclassified as baseline driving more frequently than the other voice interfaces. Conversely, glance sequences with the Sensus and smartphone voice interfaces were more often misclassified as manual phone calling. Thus, the MyLink voice interface not only reduced the overall visual demand of placing calls, but produced glance patterns more similar to driving without another task. Practitioner Summary: The attention map and confusion matrix methodologies provide ways of characterising similarities and differences in glance behaviour across secondary task conditions, complementing traditional temporally based metrics (e.g. mean glance duration, long duration glances) while addressing some of the limitations of total-eyes-off-road-time (TEORT) for comparing secondary task behaviour to baseline driving.
The quantity of information an individual absorbs each day is rapidly expanding. We consider how everyday people adapt to these demands, reading to maximize speed or comprehension, and the resultant speed – comprehension trade-offs. In a large-scale Interlude Reading experiment, 445 crowdworkers read 12 short passages set to a 12th-grade reading level. They read passages in Times and 5 other randomly selected fonts from a group of 26 total fonts. They read 2 passages per font and answered 3 comprehension questions after reading each passage. We divided each passage into 2 short screens to measure reading speed. We found a significant inverse correlation between reading speed and comprehension per-trial (R = -0.27, p < 0.001), consistent with a speed-accuracy trade-off. We found a similar correlation when aggregating per font, suggesting that font may mediate the effect (R = -0.39, p < = 0.041). Font also significantly affected reading speed (p = 0.018). Confirmatory analyses on each font's relative speed (ranked speeds per participant) demonstrated a strengthened relationship (R = -0.48, p = 0.011), suggesting the fonts' design characteristics mediated the trade-off within and between participants. Reading speed rises by 31 WPM over the session, and second screens are consistently read faster than first screens by 31 WPM (both p < 0.001). Conversely, passage order and the font did not significantly affect comprehension, nor did they indicate that additional practice could mediate the trade-off. Interestingly, practice effects appear insufficient to overcome an intrinsic speed – comprehension trade-off. Fonts demonstrated the same trade-off in an intra-participant analysis; our results suggest that leveraging typeface design could minimize individuals' trade-offs. Indeed, finding the right font for the reader could improve individual performance in the context of their reading goal: comprehension or speed.
Readability is on the cusp of a revolution. Fixed text is becoming fluid as a proliferation of digital reading devices rewrite what a document can do. As past constraints make way for more flexible opportunities, there is great need to understand how reading formats can be tuned to the situation and the individual. We aim to provide a firm foundation for readability research, a comprehensive framework for modern, multi-disciplinary readability research. Readability refers to aspects of visual information design which impact information flow from the page to the reader. Readability can be enhanced by changes to the set of typographical characteristics of a text. These aspects can be modified on-demand, instantly improving the ease with which a reader can process and derive meaning from text. We call on a multi-disciplinary research community to take up these challenges to elevate reading outcomes and provide the tools to do so effectively.
Modern digital interfaces display typeface in ways new to the 500 year old art of typography, driving a shift in reading from primarily long-form to increasingly short-form. In safety-critical settings, such at-a-glance reading competes with the need to understand the environment. To keep both type and the environment legible, a variety of 'middle layer' approaches are employed. But what is the best approach to presenting type over complex backgrounds so as to preserve legibility? This work tests and ranks middle layers in three studies. In the first study, Gaussian blur and semi-transparent 'scrim' middle layer techniques best maximise legibility. In the second, an optimal combination of the two is identified. In the third, letter-localised middle layers are tested, with results favouring drop-shadows. These results, discussed in mixed reality (MR) including overlays, virtual reality (VR), and augmented reality (AR), considers a future in which glanceable reading amidst complex backgrounds is common. Practitioner summary: Typography over complex backgrounds, meant to be read and understood at a glance, was once niche but today is a growing design challenge for graphical user interface HCI. We provide a technique, evidence-based strategies, and illuminating results for maximising legibility of glanceable typography over complex backgrounds.
Typography plays an increasingly important role in today's dynamic digital interfaces. Graphic designers and interface engineers have more typographic options than ever before. Sorting through this maze of design choices can be a daunting task. Here we present the results of an experiment comparing differences in glance-based legibility between eight popular sans-serif typefaces. The results show typography to be more than a matter of taste, especially in safety critical contexts such as in-vehicle interfaces. Our work provides both a method and rationale for using glanceable typefaces, as well as actionable information to guide design decisions for optimised usability in the fast-paced mobile world in which information is increasingly consumed in a few short glances. Practitioner summary: There is presently no accepted scientific method for comparing font legibility under time-pressure, in 'glanceable' interfaces such as automotive displays and smartphone notifications. A 'bake-off' method is demonstrated with eight popular sans-serif typefaces. The results produce actionable information to guide design decisions when information must be consumed at-a-glance. Abbreviations: DOT: department of transportation; FAA: Federal Aviation Administration; GHz: gigahertz; Hz: hertz; IEC: International Electrotechnical Commission; ISO: International Organization for Standardization; LCD: liquid crystal display; MIT: Massachusetts Institute of Technology; ms: milliseconds; OS: operating system.
The impact of using a smartwatch to initiate phone calls on driver workload, attention, and performance was compared to smartphone visual-manual (VM) and auditory-vocal (AV) interfaces. In a driving simulator, 36 participants placed calls using each method. While task time and number of glances were greater for AV calling on the smartwatch vs. smartphone, remote detection task (R-DRT) responsiveness, mean single glance duration, percentage of long duration off-road glances, total off-road glance time, and percent time looking off-road were similar; the later metrics were all significantly higher for the VM interface vs. AV methods. Heart rate and skin conductance were higher during phone calling tasks than "just driving", but did not consistently differentiate calling method. Participants exhibited more erratic driving behavior (lane position and major steering wheel reversals) for smartphone VM calling compared to both AV methods. Workload ratings were lower for AV calling on both devices vs. VM calling.
Reading at a glance, once a relatively infrequent mode of reading, is becoming common. Mobile interaction paradigms increasingly dominate the way in which users obtain information about the world, which often requires reading at a glance, whether from a smartphone, wearable device, or in-vehicle interface. Recent research in these areas has shown that a number of factors can affect text legibility when words are briefly presented in isolation. Here we expand upon this work by examining how legibility is affected by more crowded presentations. Word arrays were combined with a lexical decision task, in which the size of the text elements and the inter-line spacing (leading) between individual items were manipulated to gauge their relative impacts on text legibility. In addition, a single-word presentation condition that randomized the location of presentation was compared with previous work that held position constant. Results show that larger text was more legible than smaller text. Wider leading significantly enhanced legibility as well, but contrary to expectations, wider leading did not fully counteract decrements in legibility at smaller text sizes. Single-word stimuli presented with random positioning were more difficult to read than stationary counterparts from earlier studies. Finally, crowded displays required much greater processing time compared to single-word displays. These results have implications for modern interface design, which often present interactions in the form of scrollable and/or selectable lists. The present findings are of practical interest to the wide community of graphic designers and interface engineers responsible for developing our interfaces of daily use.
Allergen information on food labels is not standardized, making allergen avoidance difficult for consumers. This study investigated the speed and accuracy of allergen identification on commercial packaging across different types of warning labels. The results identified packaging label characteristics significantly correlated with faster and more accurate identification of allergens. Standardizing warning and safe-to-consume labels may reduce risk of accidental allergen exposure for consumers managing food allergies.
Interest in leveraging smartphone technology for scientific data collection has increased significantly in recent years. Mobile platforms have now been employed to investigate a variety of physiological and behavioral phenomena. Here we add to this rapidly growing body of work, using a specially designed mobile application to collect data on text legibility using a paradigm that mirrors established laboratory methods. Lexical decision data (an established proxy for text legibility) were collected from a smartphone platform over the course of several months, ultimately resulting in a sample of trials equivalent to a moderately sized lab experiment. Two typefaces (Frutiger and Eurostile) were tested in both positive and negative polarities. Results suggest that the participant sample was highly motivated and willing to participate in periodic task probes during an engagement period of 1-2 weeks. Consistent with previous work, positive polarity text was read more easily than negative polarity, and response accuracy rose with display duration. However, no significant effects were observed for typeface (i.e., comparing Eurostile to Frutiger) under the testing conditions. Although the application successfully collected activity state and illumination for a majority of trials, sampling rates were insufficient to make comparisons along these dimensions. There was also some concern that the application framework may not present stimuli with reliable timing. Given the relatively small sample size of this initial investigation and the uncontrolled experiment setting, the results suggest that there is substantive potential for this approach as a viable platform for experimental data collection. The trade-offs inherent in a mobile data collection are substantial, and are discussed in detail for this project.
Drivers adapt their glance behavior when using automation, which may detract attention from their surroundings. Glance behavior during parallel parking maneuvers performed with and without automated steering was compared. Drivers directed a smaller proportion of their glances toward the parking space and spent less time looking at it when using automation than when not using automation. The proportion of glances and time spent looking at the instrument cluster containing information from the automation increased significantly. Drivers also spent a significantly larger proportion of time looking at the instrument cluster and a smaller proportion looking forward and rearward when using automation while approaching a parking space. The system selected the parking space in the approach phase, which may have drawn attention to the instrument cluster. In conclusion, when using automated steering during parallel parking drivers monitored their surroundings less and looked at system displays more presumably to supervise the automation. The safety implications of these changes in glance behavior should be explored in future research. (C) 2018 Elsevier Ltd. All rights reserved.
In-vehicle information systems that allow drivers to use a single voice command to complete a task rather than multiple commands better keep drivers' attention toward the road especially compared with when drivers complete the task manually. However, single voice commands are longer and more complex and may be difficult for older drivers to use. The current study examined the glance behavior, workload, and driving performance of driverse age 20-66 years when they placed a call using their hands or voice with the Chevrolet MyLink or Volvo Sensus information system during highway driving. In general, as age increased, drivers took longer to complete phone calls, reported greater workload when using voice commands, and made significantly more off-road glances lasting longer than two seconds when placing calls relative to younger drivers. Both the voice-command systems of MyLink and Sensus increased the proportion of time that drivers were looking at the road when calling compared with manual phone calling, but the relative increase was greater when using MyLink's single-voice-command system compared with the multiple-command system of Sensus, and this advantage grew as drivers aged. The findings indicate that placing calls while driving using voice commands helps drivers of all ages keep their attention toward the road better than doing so manually, and that, contrary to expectation, using a single-command system like MyLink's worked better than a multiple-command system like Sensus for older drivers as well as younger ones. (C) 2018 Published by Elsevier Ltd.
When designers typographically tweak fonts to make an interface look ‘cool,’ they do so amid a rich design tradition, albeit one that is little-studied in regards to the rapid ‘at a glance’ reading afforded by many modern electronic displays. Such glanceable reading is routinely performed during human-machine interactions where accessing text competes with attention to crucial operational environments. There, adverse events of significant consequence can materialize in milliseconds. As such, the present study set out to test the lower threshold of time needed to read and process text modified with three common typographic manipulations: letter height, width, and case. Results showed significant penalties for the smaller size. Lowercase and condensed width text also decreased performance, especially when presented at a smaller size. These results have important implications for the types of design decisions commonly faced by interface professionals, and underscore the importance of typographic research into the human performance impact of seemingly “aesthetic” design decisions. The cost of “cool” design may be quite steep in high-risk contexts.
The Alliance of Automobile Manufacturers and the National Highway Traffic Safety Administration have each developed a set of guidelines intended to help developers of embedded in-vehicle systems minimize the visual demand placed on a driver interacting with the visual-manual interface of the system. Though based on similar precepts, the guidelines differ in the evaluation methodologies and the criteria used to define safe levels of visual demand. The current study compared the pass/fail conclusions from applying the two guidelines. Four visual-manual tasks were evaluated using two embedded in-vehicle systems (Volvo Sensus, Chevrolet MyLink) during highway driving. Only a preset radio tuning task met the threshold for acceptable visual demand in both guidelines. The pass/fail conclusions for three of the four tasks [manual radio tuning (fail), preset radio tuning (pass), easy contact calling (fail)] performed using either system were the same for both guidelines; calling a contact with multiple possible numbers using MyLink failed both guidelines, and with Sensus the task passed the Alliance guidelines but not NHTSA's. Exploratory analyses suggested that broadening the age range of the participant sample specified in the Alliance guidelines beyond 45-65 year olds did not change pass/fail conclusions. Results from a Monte Carlo simulation suggested that relying on data from a single trial per the NHTSA guidelines may reduce the repeatability of pass/fail conclusions. Interestingly, the manual radio tuning task failed to pass both sets of guidelines, even though the organizations used it as a reference task for setting acceptable levels of visual demand. Perhaps this indicates that radios have become more difficult to tune than the ones that provided the basis for the guidelines; however, naturalistic driving studies have not indicated increased risk from tuning more modern radios. Analysis of glance behavior during naturalistic driving may provide opportunities to further refine the acceptable thresholds for visual demand. (C) 2017 Elsevier Ltd. All rights reserved.