The development of large datasets of natural images has galvanized progress in psychology, neuroscience, and computer science. Notably, the THINGS database constitutes a collective effort towards understanding of human visual knowledge by accumulating rich data on a shared set of visual object concepts across several studies. In this paper, we introduce Drawing of THINGS ( DoT ), a novel dataset of 28,627 human drawings of 1854 diverse object concepts, sampled systematically from concrete picturable and nameable nouns in the American English language, mirroring the structure of the THINGS image database. In addition to data on drawings’ stroke history, we further collected fine-grained recognition data for each drawing, along with metadata on participant demographics, drawing ability, and mental imagery. We characterize people’s ability to communicate and recognize semantic information encoded in drawings and compare this ability to their ability to recognize real-world images of the same visual objects. We also explore the relationship between drawing understanding and the memorability and typicality of the objects contained in THINGS. In sum, we envision DoT as a powerful tool that builds on the THINGS database to advance understanding of how humans express knowledge about visual concepts.
Most people learn to draw casually at an early age, but learning to create accurate observational drawings usually require extensive practice. Why is accurate drawing so hard? And what distinguishes people skilled at accurate drawing from those who are not? Empirical research has primarily studied these questions for the task of copying a source picture. This article argues that the difficulty of copying tasks can be explained by limitations of visual memory and peripheral vision. In addition, variation in the control of eye movements could explain mixed findings regarding the relationship between perceptual skill and drawing skill, and regarding the role of prior knowledge in drawing. We explore the hypothesis that artists aiming to create accurate drawings from observation employ a broad suite of strategies for coordinating their eye movements with drawing actions, enabling them to transfer information from the source image to their drawing within the limitations of visual memory and peripheral vision. We advocate for further study of these strategies in naturalistic scenarios that extend beyond copying tasks.
Human behavior is fundamentally generative: People create pictures, write stories, compose music, and engage in conversation. Traditional approaches in psychology and cognitive science have not focused on this open-endedness, instead favoring more constrained task settings that admit a limited set of outcomes. Although those approaches have been fruitful, new approaches might be needed to develop a unified understanding of the generative, open-ended behaviors that are so emblematic of human cognition. This article demonstrates the value of generative behaviors as targets for cognitive modeling by providing rich behavioral data that reveal how multiple cognitive processes coordinate. Drawing production serves as a case study illustrating this approach, showing how perception, memory, social inference, and motor control coordinate flexibly on the basis of communicative context. Recent advances in generative artificial intelligence offer both new tools for modeling open-ended human behavior and new comparative targets for understanding similarities and differences between human and machine intelligence. However, applying these tools effectively might require new experimental paradigms, larger data sets, and careful consideration of what mechanistic correspondence between models and human cognition is necessary for scientific progress. Embracing the open-ended nature of human thought and behavior poses methodological challenges but offers a promising path toward understanding the most distinctive aspects of human intelligence.
Language models can use verifiable rewards to improve at a wide variety of reasoning tasks. However, both parametric (e.g. RLVR) and non-parametric (e.g. prompt optimization) approaches to doing so typically require hundreds of training samples and thousands of model rollouts, making them expensive in the best case and intractable in the worst. To address this challenge, we introduce Contrastive Reflection (CORE), a non-parametric learning algorithm that compares past reasoning traces to generate insights: short natural-language descriptions of reasoning strategies and constraints that capture differences between successful and unsuccessful problem attempts. Across four reasoning tasks, we demonstrate that CORE enables more rapid improvement than both parametric (GRPO) and non-parametric (GEPA, episodic RAG, and MemRL) methods, while using fewer rollouts. Under fixed rollout budgets with as few as five training samples, we then show that CORE also achieves comparable or greater performance gains than each baseline. Finally, we highlight how CORE is also substantially more context-efficient than non-parametric baselines, requiring fewer prompt tokens while storing learned knowledge as compact, interpretable natural-language insights. Our results therefore suggest that distilling contrasts between successful and unsuccessful reasoning traces into abstract and useful insights can provide a more efficient and interpretable route to model self-improvement than weight updates, prompt optimization, or direct reuse of stored reasoning traces.
A quintessential feature of human intelligence is the ability to create ad hoc conventions over time to achieve shared goals efficiently. We investigate how communication strategies evolve through repeated collaboration as people coordinate on shared procedural abstractions. To this end, we conducted an online unimodal study (n = 98) using natural language to probe abstraction hierarchies. In a follow-up lab study (n = 40), we examined how multimodal communication (speech and gestures) changed during physical collaboration. Pairs used augmented reality to isolate their partner's hand and voice; one participant viewed a 3D virtual tower and sent instructions to the other, who built the physical tower. Participants became faster and more accurate by establishing linguistic and gestural abstractions and using cross-modal redundancy to emphasize key changes from previous interactions. Based on these findings, we extend probabilistic models of convention formation to multimodal settings, capturing shifts in modality preferences. Our findings and model provide building blocks for designing convention-aware intelligent agents situated in the physical world.
How might messages about large language models (LLMs) found in public discourse influence the way people think about and interact with these models? To explore this question, we randomly assigned participants (N = 470) to watch short informational videos presenting LLMs as either machines, tools, or companions—or to watch no video. We then assessed how strongly they believed LLMs to possess various mental capacities, such as the ability to have intentions or remember things. We found that participants who watched video messages presenting LLMs as companions reported believing that LLMs more fully possessed these capacities than did participants in other groups. In a follow-up study (N = 604), we replicated these findings and found nuanced effects on how these videos also impact people’s reliance on LLM-generated responses when seeking out factual information. Together, these studies suggest that messages about LLMs—beyond technical advances—may shape what people believe about these systems and how they rely on LLM-generated responses.
Many problems seem to require a flash of insight to solve. What form do these sudden insights take, and what impact do they have on how people approach similar problems in the future? In this work, we prompted participants (N = 189) to talk aloud as they attempted to solve a sequence of five "matchstick-arithmetic" problems. These problems either all relied on the same kind of non-obvious solution (Same group) or a different kind each time (Different group). We found that Same participants improved more rapidly than Different participants, and as they improved, they talked more and talked about different things when solving later problems. Specifically, they were more likely to spontaneously categorize the problem they were working on. Taken together, these findings suggest that a hallmark of transferable insights is their accessibility for verbal report, even if the underlying precursors of insight remain difficult to articulate.
Many innovations have come from people working together as partners in thought. These partnerships, however, are not restricted to single encounters. Some of the most meaningful collaborations evolve over weeks, months, or even lifetimes. What are the core computations that enable long-term thought partnerships? Prior work in cognitive science has made initial progress by investigating how people construct mental models of their partners on the fly, establish common ground using language and other modalities, and generate joint plans that lead to successful outcomes. However, it remains unknown what cognitive mechanisms enable such interactions to evolve into genuine partnerships over longer timescales, especially under measures of success that extend beyond task performance. Theoretical and empirical progress on these issues could be instrumental for defining and designing AI systems that may even be capable of establishing long-term thought partnerships with humans. This article outlines several promising avenues for leveraging approaches from cognitive science and AI to study enriching intellectual partnerships.
People often seek out ways to watch others perform complex action sequences (e.g., sports). What makes some sequences more enjoyable to watch than others? We generated 24 video clips of gameplay from a Flappy Bird-style video game. Clips varied in difficulty (how often players succeeded on average) and in moment-to-moment uncertainty (how likely the player was to crash at any given step). Participants (N=864) rated each video on one of three dimensions: how much they enjoyed it, how difficult the level appeared, or how dangerous the player's trajectory appeared. We found that participants preferred videos where the player seemed to be completing more difficult obstacle courses, but dangerousness did not predict enjoyment ratings. These findings show how procedurally generated stimuli can isolate the factors that affect how enjoyable an action sequence is to watch.
Data literacy is becoming a foundational skillset in our increasingly data-driven society. Fields that rely heavily on data, such as data visualization, cognitive psychology, and artificial intelligence, each contribute unique perspectives in studying data literacy. Yet, these efforts remain siloed. Researchers studying visualization literacy and AI literacy have recently run separate workshops at ACM CHI, each independently calling for interdisciplinary conversations. Such conversations would not only advance the study of data literacy but also inform research agendas within each individual discipline. Building on ACM CHI’s track record of attracting diverse researchers across many disciplines, we propose a workshop that connects these literacies through the common ground of data literacy. We invite interdisciplinary perspectives by bringing researchers, practitioners, and students together with expertise from fields including HCI, data visualization, psychology, AI, and learning sciences to develop a holistic framework for understanding, measuring, and teaching data literacy that is grounded in real-world applications.
When first meeting somebody, we’re faced with the challenge of “getting to know them.” Why do some questions seem to enable this better than others? In Experiment 1, participants (N=185) evaluated a large bank of conversational questions. We found that questions varied along a reliable latent dimension of interpersonal depth ranging from “small talk” to “deep” questions. In Experiment 2 (N=188), participants answered a subset of these questions along with a number of self-report personality scales. Using a language model to estimate how informative participants’ free responses were, we find that individualized personality predictions were more accurate when incorporating free responses; furthermore, responses to deeper questions supported more accurate personality inferences than small talk. Taken together, results suggest not only that responses contained the statistical information necessary to make abstract social inferences, but also that people have accurate intuitions about which conversational topics enable learning about and connecting with others.
In collaborative creation tasks, people steer artifacts towards specific goals by _refining_ them with _multimodal_ communication over multiple rounds of interaction. In contrast, generative AI excels at creating artifacts in a single turn but can struggle to make precise refinements that match our design intent. To close this gap, we present mrCAD, a dataset of multi-turn interactions in which pairs of humans iteratively created and refined computer-aided designs (CADs). In each game, a _Designer sent instructions to a _Maker_, explaining how to create and subsequently refine a CAD to match a target design that only the _Designer_ could see. mrCAD consists of 6,082 communication games, 15,163 instruction-execution rounds, played between 1,092 pairs of human players. Crucially, _Designers_ had access to two communication modalities – text and drawing. Analysis finds that players relied more on text in refinement than in initial generation instructions, and used different linguistic elements for refinement than for generation. We also find that state-of-the-art VLMs are better at following generation instructions than refinement instructions. These results lay the foundation for modeling multi-turn, multimodal communication not captured in prior datasets.
As large language models (LLMs) become increasingly popular and prevalent in media and daily conversations, individuals encounter different portrayals of LLMs from various sources. It is important to understand how these portrayals can shape their beliefs about LLMs as this can have downstream impacts on adoption and usage behaviors. In this work, we investigate what mental capacities individuals attribute to LLMs after being exposed to short videos adopting one of three portrayals: mechanistic (LLMs as machines), functional (LLMs as tools), and intentional (LLMs as companions). We find that the intentional portrayal increases the attribution of mental capacities to LLMs, and that individuals tend to attribute mind-related capacities the most, followed by heart- then body-related capacities. We discuss the implications of these findings, provide recommendations on how to portray LLMs, and outline directions for future research.
A key feature of human collaboration is the ability to iteratively refine the concepts we have communicated. In contrast, while generative AI excels at the generation of content, it often struggles to make specific language-guided modifications of its prior outputs. To bridge the gap between how humans and machines perform edits, we present mrCAD, a dataset of multimodal instructions in a communication game. In each game, players created computer aided designs (CADs) and refined them over several rounds to match specific target designs. Only one player, the Designer, could see the target, and they must instruct the other player, the Maker, using text, drawing, or a combination of modalities. mrCAD consists of 6,082 communication games, 15,163 instruction-execution rounds, played between 1,092 pairs of human players. We analyze the dataset and find that generation and refinement instructions differ in their composition of drawing and text. Using the mrCAD task as a benchmark, we find that state-of-the-art VLMs are better at following generation instructions than refinement instructions. These results lay a foundation for analyzing and modeling a multimodal language of refinement that is not represented in previous datasets.
Sketching serves as a versatile tool for externalizing ideas, enabling rapid exploration and visual communication that spans various disciplines. While artificial systems have driven substantial advances in content creation and human-computer interaction, capturing the dynamic and abstract nature of human sketching remains challenging. In this work, we introduce SketchAgent, a language-driven, sequential sketch generation method that enables users to create, modify, and refine sketches through dynamic, conversational interactions. Our approach requires no training or fine-tuning. Instead, we leverage the sequential nature and rich prior knowledge of off-the-shelf multimodal large language models (LLMs). We present an intuitive sketching language, introduced to the model through in-context examples, enabling it to "draw" using string-based actions. These are processed into vector graphics and then rendered to create a sketch on a pixel canvas, which can be accessed again for further tasks. By drawing stroke by stroke, our agent captures the evolving, dynamic qualities intrinsic to sketching. We demonstrate that SketchAgent can generate sketches from diverse prompts, engage in dialogue-driven drawing, and collaborate meaningfully with human users.