Conversational partners align the meanings of their words over the course of interaction to coordinate and communicate. One process of alignment is lexical entrainment, whereby partners mirror and abbreviate their word usage to converge on shared terms for referents relevant to the conversation. However, lexical entrainment may result in inefficient mimicry that does not add new information, suggesting that task-oriented communication may favor alignment through other means. The present study investigates the process of alignment in Danish conversations in which dyads learned to categorize unfamiliar "aliens" using trial-and-error feedback. Performance improved as dyad communication became less verbose, measured as a decrease in the entropy of word usage. Word usage also diverged between partners as measured by Jensen-Shannon Divergence, which indicates that alignment was not achieved through lexical entrainment. A computational model of dyadic communication is shown to account for the alien game results in terms of joint least effort. The model shows that alignment of partner referents can increase as a result of minimizing both the joint entropy of dyadic word usage and the conditional entropy of individual referents given the joint signal distribution. We conclude that the principle of least effort, originally proposed to shape language evolution, may also support alignment in task-oriented communication.
Vocal responses from caregivers are believed to promote more frequent and more advanced infant vocalizations. However, studies that examine this relationship typically do not account for the fact that infant and adult vocalizations are distributed in hierarchical clusters over the course of the day. These bursts and lulls create a challenge for accurately detecting the effects of adult input at immediate turn-by-turn timescales within real-world behavior, as adult responses tend to happen during already occurring bursts of infant vocalizations. Analyzing daylong audio recordings of real-world vocal communication between human infants (ages 3, 6, 9, and 18 months) and their adult caregivers, we first show that both infant and caregiver vocalization events are clustered in time, as evidenced by positive correlations between successive inter-event intervals (IEIs). We propose an approach informed by flight time analyses in foraging studies to assess whether the timing of a vocal agent's next vocalization is modified by inputs from another vocal agent, controlling for the first agent's previous IEI. For both infants and adults, receiving a social response predicts that the individual will vocalize again sooner than they would have in the absence of a response. Overall, our results are consistent with a view of infant-caregiver vocal interactions as an 'interpersonal foraging' process with inherent multi-scale dynamics wherein social responses are among the resources the individuals are foraging for. The analytic approaches introduced here have broad utility to study communication in other modalities, contexts, and species.
Neural network modeling has played a central role in psycholinguistic studies of lexical processing, but the recent advent of large language models (LLMs) offers a different approach that may yield new insights into the mental lexicon. Four LLMs were prompted across three experiments to test how they generate psycholinguistic ratings of words in comparison with humans. LLM ratings, averaged across varying list contexts, were found to be highly correlated with human ratings, and differences in correlation strengths were partly explained by differences in rating ambiguity. LLM context manipulations strengthened correlations with human ratings through better calibration, and variability in LLM ratings was correlated with human inter-rater variability. Additional results from testing LLM generation of word naming latencies showed functional deviations from factors that underlie human word naming, indicating that lexical function assembly in LLMs is currently limited by patterns of co-occurrence in textual data. Patterns at finer-grained timescales are needed in the training data to model online lexical processes. We conclude that LLMs used context to guide the assembly of generalized lexical functions, rather than recalling ratings and latencies from training data.
Traditional theories of the human mental lexicon posit dedicated mechanisms of processing that develop as sustained functions of brain and mind. Large Language Models (LLMs) provide a new approach in which lexical functions emerge from the learning and processing of sequences in contexts. We prompted lexical functions in ChatGPT and compared numeric responses with averaged human data for a sample of 390 words for a range of lexical variables, some derived from corpus analyses and some from Likert ratings. ChatGPT responses were moderately to highly correlated with mean values, more so for GPT-4 versus GPT-3.5, and responses were sensitive to context and human inter-rater reliability. We argue that responses were not recalled from memorized training data but were instead soft-assembled from more general-purpose representations. Emergent functions in LLMs offer a new approach to modeling language and cognitive processes.
Evolutionary theories of foraging hypothesize that foraging strategies evolve to maximize search efficiency. Many studies have investigated the central trade-off between explore–exploit and how individual foragers manage it under various conditions. For foragers in groups, this trade-off can be affected by the social environment, influencing the evolution of individual search strategies. Previous work has shown that when learning socially, explorative search strategies can optimize group search efficiency. However, social learning can cause discrepancies in strategies that benefit the group versus an individual. We model the evolution of explorative and exploitative strategies using Lévy exponents under different levels of social learning and investigate their effect on individual and group search efficiencies. We show that reliance on social learning can lead to the evolution of mixed groups that are not optimally efficient. Exploiters can have a selective advantage in scrounging findings by explorers, but too many exploiters can diminish group efficiencies. However, greater opportunities for social learning can increase the benefits of explorative strategies. Finally, we show that area-restricted search can help individuals balance exploration and exploitation, and make groups more efficient. Our results demonstrate how exploration and exploitation must be balanced at both individual and collective levels for efficient search.
Discussion of AI alignment (alignment between humans and AI systems) has focused on value alignment, broadly referring to creating AI systems that share human values. We argue that before we can even attempt to align values, it is imperative that AI systems and humans align the concepts they use to understand the world. We integrate ideas from philosophy, cognitive science, and deep learning to explain the need for concept alignment, not just value alignment, between humans and machines. We summarize existing accounts of how humans and machines currently learn concepts, and we outline opportunities and challenges in the path towards shared concepts. Finally, we explain how we can leverage the tools already being developed in cognitive science and AI research to accelerate progress towards concept alignment.
Language is intrinsically multimodal. Speakers use gestures, prosody, gaze, and facial expressions as cues that complement and expand the meaning expressed in their words. These varied signals operate in remarkably flexible coordination, constantly adapting to the conversational partners and topics as they change over time. We argue that an ecological approach to multimodal behavior offers a promising account of natural conversation as it takes place both in experimental contexts, and in natural ones outside the lab. After reviewing major historical themes in the study of language and communication, we describe how this ecological perspective situates future work, especially work that seeks to quantify these processes. We describe a quantitative hypothesis that multimodal signals are projected on manifolds of lower dimension that can be described in terms of dynamical systems. We refer to these lower dimensional patterns as "pragmatic modes," and compare this idea to a number of prior theoretical proposals. We describe how the notion of pragmatic mode frames a quantitative basis to supplement and extend prior research with explicitly quantitative goals. The paper concludes with an outline to link quantitative descriptions of multimodality with more abstract, qualitative theories of the past few decades, and describe how future research might explore pragmatic modes, how they change over the course of conversation, and relate to our understanding of human communication.
In 2020, the combination of police killings of unarmed Black people, including George Floyd, Breonna Taylor, and Ahmaud Arbery, and the Coronavirus Disease 2019 (COVID-19) pandemic brought about public outrage over long-standing inequalities in society. The events of 2020 ignited global attention to systemic racism and racial inequalities, including the lack of diversity, equity, and inclusion in the academy and especially in science, technology, engineering, mathematics, and medicine (STEMM) fields. Racial and ethnic diversity in graduate programs in particular warrants special attention as graduate students of color report experiencing alarming rates of racism, discrimination, microaggressions, and other exclusionary behaviors. As part of the Graduate Dean's Advisory Council on Diversity (GDACD) at the University of California Merced, the authors of this manuscript held a year-long discussion on these issues and ways to take meaningful action to address these persistent issues of injustices. We have outlined 10 rules to help graduate programs develop antiracist practices to promote racial and ethnic justice, equity, diversity, and inclusion (JEDI) in the academy. We focus on efforts to address systemic causes of the underrepresentation and attrition of students from minoritized communities. The 10 rules are developed to allow graduate groups to formulate and implement rules and policies to address root causes of underrepresentation of minoritized students in graduate education.
Information retrieval from brain responses to auditory and visual stimuli has shown success through classification of song names and image classes presented to participants while recording EEG signals. Information retrieval in the form of reconstructing auditory stimuli has also shown some success, but here we improve on previous methods by reconstructing music stimuli well enough to be perceived and identified independently. Furthermore, deep learning models were trained on time-aligned music stimuli spectrum for each corresponding one-second window of EEG recording, which greatly reduces feature extraction steps needed when compared to prior studies. The NMED-Tempo and NMED-Hindi datasets of participants passively listening to full length songs were used to train and validate Convolutional Neural Network (CNN) regressors. The efficacy of raw voltage versus power spectrum inputs and linear versus mel spectrogram outputs were tested, and all inputs and outputs were converted into 2D images. The quality of reconstructed spectrograms was assessed by training classifiers which showed 81% accuracy for mel-spectrograms and 72% for linear spectrograms (10% chance accuracy). Lastly, reconstructions of auditory music stimuli were discriminated by listeners at an 85% success rate (50% chance) in a two-alternative match-to-sample task.
Timing is critical to successful social interactions. The temporal structure of dyadic vocal interactions emerges from the rhythm, timing, and frequency of each individual’s vocalizations and reflects how the dyad dynamically organizes and adapts during an interaction. This study investigated the temporal structure of vocal interactions longitudinally in parent-child dyads of typically developing (TD) infants (n=49; 9-18 months; 48% male) and toddlers with ASD (n=23; 27.2±5.0 months; 91.3% male) to identify how developing language and social skills impact the temporal dynamics of the interaction. Acoustic hierarchical temporal structure (HTS), a measure of the nested clustering of acoustic events across multiple timescales, was measured in free play interactions using Allan Factor. HTS reflects a signal’s temporal complexity and variability, with greater HTS indicating reduced flexibility of the dyadic system. Child expressive language significantly predicted HTS (ß=-0.2) longitudinally across TD infants, with greater dyadic HTS associated with lower child language skills. ASD dyads exhibited greater HTS (i.e., more rigid temporal structure) than nonverbal matched (d=0.41) and expressive language matched TD dyads (d=0.28). Increased HTS in ASD dyads occurred at timescales > 1 second, suggesting greater structuring of pragmatic aspects of interaction. Results provide a new window into how language development and social reciprocity serve as constraints to shape parent-child interaction dynamics and showcase a novel automated approach to characterizing vocal interactions across multiple timescales during early childhood.
During the first years of life, infant vocalizations change considerably, as infants develop the vocalization skills that enable them to produce speech sounds. Characterizations based on specific acoustic features, protophone categories, or phonetic transcription are able to provide a representation of the sounds infants make at different ages and in different contexts but do not fully describe how sounds are perceived by listeners, can be inefficient to obtain at large scales, and are difficult to visualize in two dimensions without additional statistical processing. Machine-learning-based approaches provide the opportunity to complement these characterizations with purely data-driven representations of infant sounds. Here, we use spectral features extraction and unsupervised machine learning, specifically Uniform Manifold Approximation (UMAP), to obtain a novel 2-dimensional spatial representation of infant and caregiver vocalizations extracted from day-long home recordings. UMAP yields a continuous and well-distributed space conducive to certain analyses of infant vocal development. For instance, we found that the dispersion of infant vocalization acoustics within the 2-D space over a day increased from 3 to 9 months, and then decreased from 9 to 18 months. The method also permits analysis of similarity between infant and adult vocalizations, which also shows changes with infant age.
Classifying EEG responses to naturalistic acoustic stimuli is of theoretical and practical importance, but standard approaches are limited by processing individual channels separately on very short sound segments (a few seconds or less). Recent developments have shown classification for music stimuli (~2 mins) by extracting spectral components from EEG and using convolutional neural networks (CNNs). This paper proposes an efficient method to map raw EEG signals to individual songs listened for end-to-end classification. EEG channels are treated as a dimension of a [Channel x Sample] image tile, and images are classified using CNNs. Our experimental results (88.7%) compete with state-of-the-art methods (85.0%), yet our classification task is more challenging by processing longer stimuli that were similar to each other in perceptual quality, and were unfamiliar to participants. We also adopt a transfer learning scheme using a pre-trained ResNet-50, confirming the effectiveness of transfer learning despite image domains unrelated from each other.
Search requires balancing exploring for more options and exploiting the ones previously found. Individuals foraging in a group face another trade-off: whether to engage in social learning to exploit the solutions found by others or to solitarily search for unexplored solutions. Social learning can better exploit learned information and decrease the costs of finding new resources, but excessive social learning can lead to over-exploitation and too little exploration for new solutions. We study how these two trade-offs interact to influence search efficiency in a model of collective foraging under conditions of varying resource abundance, resource density and group size. We modelled individual search strategies as Lévy walks, where a power-law exponent (μ) controlled the trade-off between exploitative and explorative movements in individual search. We modulated the trade-off between individual search and social learning using a selectivity parameter that determined how agents responded to social cues in terms of distance and likely opportunity costs. Our results show that social learning is favoured in rich and clustered environments, but also that the benefits of exploiting social information are maximized by engaging in high levels of individual exploration. We show that selective use of social information can modulate the disadvantages of excessive social learning, especially in larger groups and when individual exploration is limited. Finally, we found that the optimal combination of individual exploration and social learning gave rise to trajectories with μ ≈ 2 and provide support for the general optimality of such patterns in search. Our work sheds light on the interplay between individual search and social learning, and has broader implications for collective search and problem-solving.
It is now widely accepted that the brunt of animal communication is conducted via several modalities, e.g. acoustic and visual, either simultaneously or sequentially. This is a laudable multimodal turn relative to traditional accounts of temporal aspects of animal communication which have focused on a single modality at a time. However, the fields that are currently contributing to the study of multimodal communication are highly varied, and still largely disconnected given their sole focus on a particular level of description or their particular concern with human or non-human animals. Here, we provide an integrative overview of converging findings that show how multimodal processes occurring at neural, bodily, as well as social interactional levels each contribute uniquely to the complex rhythms that characterize communication in human and non-human animals. Though we address findings for each of these levels independently, we conclude that the most important challenge in this field is to identify how processes at these different levels connect. This article is part of the theme issue 'Synchrony and rhythm interaction: from the brain to behavioural ecology'.
The majority of existing studies investigating characteristics of overt social behavior in individuals with autism spectrum disorder (ASD) relied on informants' evaluation through questionnaires and behavioral coding techniques. As a novelty, this study aimed to quantify the complex movements produced during social interactions in order to test differences in ASD movement dynamics and their convergence, or lack thereof, during social interactions. Twenty children with ASD and twenty-three children with typical development (TD) were videotaped while engaged in a face-to-face conversation with an interviewer. An image differencing technique was utilized to extract the movement time series. Spectral analyses were conducted to quantify the average power of movement, and the fractal scaling of movement. The degree of complexity matching was calculated to capture the level of behavioral coordination between the interviewer and children. Results demonstrated that the average power was significantly higher (p < 0.01), and the fractal scaling was steeper (p < 0.05) in children with ASD, suggesting excessive and less complex movement as compared to the TD peers. Complexity matching occurred between children and interviewers, but there was no reliable difference in the strength of matching between the ASD and TD children. Descriptive trends in the interviewer's behavior suggest that her movements adapted to match both ASD and TD movements equally well. The findings of our study might shed light on seeking novel behavioral markers of ASD, and on developing automatic ASD screening techniques during daily social interactions. LAY SUMMARY: By implementing an objective behavioral quantifying technique, our study demonstrated that children with autism had more body movement during face-to-face conversation, and they moved in a less complex way. The current diagnosis of autism heavily relies on doctor's experiences. These findings suggest a potential that autism might be automatically screened during daily social interactions.
The landscape of graduate science education is changing as efforts to diversify the professoriate have increased because academic faculty jobs at universities have grown scarce and more competitive. With this context as a backdrop, the present research examines the perceptions and career goals of advisors and advisees through surveys of PhD students (Study 1, N = 195) and faculty mentors (Study 2, N = 272) in science, technology, engineering, and math disciplines. Study 1 examined actual preferences and career goals of PhD students among three options: research careers, teaching careers, and non-academic careers in industry, and compared the actual preferences of students with what they perceived as being the normative preferences of faculty. Overall, students had mixed preferences but perceived that their advisors had a strong normative preference for research careers for them. Moreover, students who ranked research positions as most desirable felt the most belonging in their academic departments. Further analyses revealed no differences in career preferences as a function of underrepresented minority (URM) student status or first-generation (FG) status, but URM and FG students felt less belonging in their academic departments. Study 2 examined faculty preferences for different careers for their advisees, both in general and for current students in particular. While faculty advisors preferred students to go into research in general, when focusing on specific students, they saw their preferences as being closely aligned with the career preference of each PhD student. Faculty advisors did not perceive any difference in belonging between their students as a function of their URM status. Discrepancies between student and faculty perceptions may occur, in part, because faculty and students do not engage in sufficient discussions about the wider range of career options beyond academic research. Supporting this possibility, PhD students and faculty advisors reported feeling more comfortable discussing research careers with each other than either non-academic industry positions or teaching positions. Discussion centers on the implications of these findings for interpersonal and institutional efforts to foster diversity in the professoriate and to create open communication about career development.
Efficient foraging depends on decisions that account for the costs and benefits of various activities like movement, perception, and planning. We conducted a virtual foraging experiment set in the foothills of the Himalayas to examine how time and energy are expended to forage efficiently, and how foraging changes when constrained to a home range. Two hundred players foraged the human-scale landscape with simulated energy expenditure in search of naturally distributed resources. Results showed that efficient foragers produced periods of locomotion interleaved with perception and planning that approached theoretical expectations for Lévy walks, regardless of the home-range constraint. Despite this constancy, efficient home-range foraging trajectories were less diffusive by virtue of restricting locomotive search and spending more time instead scanning the environment to plan movement and detect far-away resources. Altogether, results demonstrate that humans can forage efficiently by arranging and adjusting Lévy-distributed search activities in response to environmental and task constraints.
Humans and other complex organisms exhibit intelligent behaviors as individual agents and as groups of coordinated agents. They can switch between independent and collective modes of behavior, and flexible switching can be advantageous for adapting to ongoing changes in conditions. In the present study, we investigated the flexibility between independent and collective modes of behavior in a simulated social foraging task designed to benefit from both modes: distancing among ten foraging agents promoted faster detection of resources, whereas flocking promoted faster consumption. There was a tradeoff between faster detection versus faster consumption, but both factors contributed to foraging success. Results showed that group foraging performance among simulated agents was enhanced by loose coupling that balanced distancing and flocking among agents and enabled them to fluidly switch among a variety of groupings. We also examined the effects of more sophisticated cognitive capacities by studying how human players improve performance when they control one of the search agents. Results showed that human intervention further enhanced group performance with loosely coupled agents, and human foragers performed better when coordinating with loosely coupled agents. Humans players adapted their balance of independent versus collective search modes in response to the dynamics of simulated agents, thereby demonstrating the importance of adaptive flexibility in social foraging.