Pedestrian dynamics are characterized as complex physical systems constrained by social norms, like personal space. Yet, prior research has ignored the spatial norms imposed by others' social interactions, such as conversations. Unlike physical obstacles, the boundaries of social interactions are latent — even more so than those of personal space — as if they are invisible walls, and must be inferred from signs of interactional involvement. In four field experiments with 4,911 participants, we show that pedestrians integrate others' gaze, proximity, body orientation, and talk to avoid interrupting possible interactions. However, pedestrians also collectively violate these spatial norms by walking through an interaction if other pedestrians had already done so. These results demonstrate how physical mobility depends on pedestrians' social computations.
Understanding non-human primate behavior is essential for advancing animal welfare and uncovering the roots of human sociality. However, automated analysis remains limited. Existing methods, many of which are human-centric, are typically task-specific, handling detection, tracking, or behavior recognition in isolation. We present AlphaChimp, an end-to-end and unified framework for chimpanzee detection, tracking, and spatiotemporal behavior recognition. To address the unique challenges of chimpanzee video analysis, such as frequent occlusions and social interactions, we build upon a DETR-based architecture with crucial modifications. Our model integrates multi-resolution temporal features to capture long-term contextual cues and employs attention mechanisms to model spatial relationships between individuals. This unique design allows AlphaChimp to jointly capture individual and social interactions. Evaluated on the ChimpACT dataset, the only benchmark with fine-grained spatiotemporal annotations of chimpanzee behavior, our method achieves state-of-the-art performance, with a 10 project page and hope this will facilitate future research in animal social dynamics.
Past research on interspecies communication has shown that animals can be trained to use Augmentative Interspecies Communication (AIC) devices, such as soundboards, to make simple requests of their caretakers. The recent uptake in AIC devices by hundreds of pet owners around the world offers a novel opportunity to investigate whether AIC is possible with owner-trained family dogs. To answer this question, we carried out two studies to test pet dogs’ ability to recognise and respond appropriately to food-related, play-related, and outside-related words on their soundboards. One study was conducted by researchers, and the other by citizen scientists who followed the same procedure. Further, we investigated whether these behaviours depended on the identity of the person presenting the word (unfamiliar person or dog’s owner) and the mode of its presentation (spoken or produced by a pressed button). We find that dogs produced contextually appropriate behaviours for both play-related and outside-related words regardless of the identity of the person producing them and the mode in which they were produced. Therefore, pet dogs can be successfully taught by their owners to associate words recorded onto soundboard buttons to their outcomes in the real world, and they respond appropriately to these words even when they are presented in the absence of any other cues, such as the owner’s body language.
Understanding non-human primate behavior is crucial for improving animal welfare, modeling social behavior, and gaining insights into both distinctly human and shared behaviors. Despite recent advances in computer vision, automated analysis of primate behavior remains challenging due to the complexity of their social interactions and the lack of specialized algorithms. Existing methods often struggle with the nuanced behaviors and frequent occlusions characteristic of primate social dynamics. This study aims to develop an effective method for automated detection, tracking, and recognition of chimpanzee behaviors in video footage. Here we show that our proposed method, AlphaChimp, an end-to-end approach that simultaneously detects chimpanzee positions and estimates behavior categories from videos, significantly outperforms existing methods in behavior recognition. AlphaChimp achieves approximately 10 behavior recognition compared to state-of-the-art methods, particularly excelling in the recognition of social behaviors. This superior performance stems from AlphaChimp's innovative architecture, which integrates temporal feature fusion with a Transformer-based self-attention mechanism, enabling more effective capture and interpretation of complex social interactions among chimpanzees. Our approach bridges the gap between computer vision and primatology, enhancing technical capabilities and deepening our understanding of primate communication and sociality. We release our code and models and hope this will facilitate future research in animal social dynamics. This work contributes to ethology, cognitive science, and artificial intelligence, offering new perspectives on social intelligence.
Understanding the behavior of non-human primates is crucial for improving animal welfare, modeling social behavior, and gaining insights into distinctively human and phylogenetically shared behaviors. However, the lack of datasets on non-human primate behavior hinders in-depth exploration of primate social interactions, posing challenges to research on our closest living relatives. To address these limitations, we present ChimpACT, a comprehensive dataset for quantifying the longitudinal behavior and social relations of chimpanzees within a social group. Spanning from 2015 to 2018, ChimpACT features videos of a group of over 20 chimpanzees residing at the Leipzig Zoo, Germany, with a particular focus on documenting the developmental trajectory of one young male, Azibo. ChimpACT is both comprehensive and challenging, consisting of 163 videos with a cumulative 160,500 frames, each richly annotated with detection, identification, pose estimation, and fine-grained spatiotemporal behavior labels. We benchmark representative methods of three tracks on ChimpACT: (i) tracking and identification, (ii) pose estimation, and (iii) spatiotemporal action detection of the chimpanzees. Our experiments reveal that ChimpACT offers ample opportunities for both devising new methods and adapting existing ones to solve fundamental computer vision tasks applied to chimpanzee groups, such as detection, pose estimation, and behavior analysis, ultimately deepening our comprehension of communication and sociality in non-human primates.
Non-intrusive, real-time analysis of the dynamics of the eye region allows us to monitor humans’ visual attention allocation and estimate their mental state during the performance of real-world tasks, which can potentially benefit a wide range of human-computer interaction (HCI) applications. While commercial eye-tracking devices have been frequently employed, the difficulty of customizing these devices places unnecessary constraints on the exploration of more efficient, end-to-end models of eye dynamics. In this work, we propose CLERA, a unified model for Cognitive Load and Eye Region Analysis, which achieves precise keypoint detection and spatiotemporal tracking in a joint-learning framework. Our method demonstrates significant efficiency and outperforms prior work on tasks including cognitive load estimation, eye landmark detection, and blink estimation. We also introduce a large-scale dataset of 30 k human faces with joint pupil, eye-openness, and landmark annotation, which aims at supporting future HCI research on human factors and eye-related analysis.
Semantic scene segmentation has primarily been addressed by forming high-level visual representations of single images. The problem of semantic segmentation in dynamic scenes has begun to receive attention with the video object segmentation and tracking problem. While there has been some recent work attempt to use deep learning models on the video level, what is not known is how the temporal dynamics information is contributing to the full scene segmentation. Moreover, most existing datasets only provide full scene annotation on non-consecutive images to ensure the variability of scenes, making it even harder to explore novel methods on video-level modeling. To address the above issues, our work takes steps to explore the behavior of modern spatiotemporal modeling approaches by: 1) constructing the MIT DriveSeg dataset, a large-scale video driving scene segmentation dataset, densely annotated for pixel-level semantic classes with 5000 consecutive video frames, and 2) proposing a joint-learning framework that reveals the contribution of temporal dynamics information in regard to different semantic classes in the driving scene. This work is intended to help assess current methods and support further exploration of the value of temporal dynamics information in video-level scene segmentation.
The Interaction Engine Hypothesis postulates that humans have a unique ability and motivation for social interaction. A crucial juncture in the ontogeny of the interaction engine could be around 2-4 years of age, but observational studies of children in natural contexts are limited. These data appear critical also for comparison with non-human primates. Here, we report on focal observations on 31 children aged 2- and 4-years old in four preschools (10 h per child). Children interact with a wide range of partners, many infrequently, but with one or two close friends. Four-year olds engage in cooperative social interactions more often than 2-year olds and fight less than 2-year olds. Conversations and playing with objects are the most frequent social interaction types in both age groups. Children engage in social interactions with peers frequently (on average 13 distinct social interactions per hour) and briefly (28 s on average) and shorter than those of great apes in comparable studies. Their social interactions feature entry and exit phases about two-thirds of the time, less frequently than great apes. The results support the Interaction Engine Hypothesis, as young children manifest a remarkable motivation and ability for fast-paced interactions with multiple partners. This article is part of the theme issue 'Revisiting the human 'interaction engine': comparative approaches to social action coordination'.
This technical report summarizes and provides detailed information about the MIT DriveSeg (Semi-auto) dataset, including technical aspects in data collection, annotation, and potential research directions.
This technical report summarizes and provides more detailed information about the MIT DriveSeg dataset [3], including technical aspects in data collection, annotation, and potential research directions.
We use an immersive virtual reality environment to explore the intricate social cues that underlie non-verbal communication involved in a pedestrian’s crossing decision. We “hack” non-verbal communication between pedestrian and vehicle by engineering a set of 15 vehicle trajectories, some of which follow social conventions and some that break them. By subverting social expectations of vehicle behavior we show that pedestrians may use vehicle kinematics to infer social intentions and not merely as the state of a moving object. We investigate human behavior in this virtual world by conducting a study of 22 subjects, with each subject experiencing and responding to each of the trajectories by moving their body, legs, arms, and head in both the physical and the virtual world. Both quantitative and qualitative responses are collected and analyzed, showing that, in fact, social cues can be engineered through vehicle trajectory manipulation. In addition, we demonstrate that immersive virtual worlds which allow the pedestrian to move around freely, provide a powerful way to understand both the mechanisms of human perception and the social signaling involved in pedestrian-vehicle interaction.
Today, and possibly for a long time to come, the full driving task is too complex an activity to be fully formalized as a sensing-acting robotics system that can be explicitly solved through model-based and learning-based approaches in order to achieve full unconstrained vehicle autonomy. Localization, mapping, scene perception, vehicle control, trajectory optimization, and higher-level planning decisions associated with autonomous vehicle development remain full of open challenges. This is especially true for unconstrained, real-world operation where the margin of allowable error is extremely small and the number of edge-cases is extremely large. Until these problems are solved, human beings will remain an integral part of the driving task, monitoring the AI system as it performs anywhere from just over 0% to just under 100% of the driving. The governing objectives of the MIT Advanced Vehicle Technology (MIT-AVT) study are to 1) undertake large-scale real-world driving data collection that includes high-definition video to fuel the development of deep learning-based internal and external perception systems; 2) gain a holistic understanding of how human beings interact with vehicle automation technology by integrating video data with vehicle state data, driver characteristics, mental models, and self-reported experiences with technology; and 3) identify how technology and other factors related to automation adoption and use can be improved in ways that save lives. In pursuing these objectives, we have instrumented 23 Tesla Model S and Model X vehicles, 2 Volvo S90 vehicles, 2 Range Rover Evoque, and 2 Cadillac CT6 vehicles for both long-term (over a year per driver) and medium-term (one month per driver) naturalistic driving data collection. Furthermore, we are continually developing new methods for the analysis of the massive-scale dataset collected from the instrumented vehicle fleet. The recorded data streams include IMU, GPS, and CAN messages, and high-definition video streams of the driver's face, the driver cabin, the forward roadway, and the instrument cluster (on select vehicles). The study is on-going and growing. To date, we have 122 participants, 15610 days of participation, 511638 mi, and 7.1 billion video frames. This paper presents the design of the study, the data collection hardware, the processing of the data, and the computer vision algorithms currently being used to extract actionable knowledge from the data.
When asked, a majority of people believe that, as pedestrians, they make eye contact with the driver of an approaching vehicle when making their crossing decisions. This work presents evidence that this widely held belief is false. We do so by showing that, in majority of cases where conflict is possible, pedestrians begin crossing long before they are able to see the driver through the windshield. In other words, we are able to circumvent the very difficult question of whether pedestrians choose to make eye contact with drivers, by showing that whether they think they do or not, they can't. Specifically, we show that over 90\% of people in representative lighting conditions cannot determine the gaze of the driver at 15m and see the driver at all at 30m. This means that, for example, that given the common city speed limit of 25mph, more than 99% of pedestrians would have begun crossing before being able to see either the driver or the driver's gaze. In other words, from the perspective of the pedestrian, in most situations involving an approaching vehicle, the crossing decision is made by the pedestrian solely based on the kinematics of the vehicle without needing to determine that eye contact was made by explicitly detecting the eyes of the driver.
Humans, as both pedestrians and drivers, generally skillfully navigate traffic intersections. Despite the uncertainty, danger, and the non-verbal nature of communication commonly found in these interactions, there are surprisingly few collisions considering the total number of interactions. As the role of automation technology in vehicles grows, it becomes increasingly critical to understand the relationship between pedestrian and driver behavior: how pedestrians perceive the actions of a vehicle/driver and how pedestrians make crossing decisions. The relationship between time-to-arrival (TTA) and pedestrian gap acceptance (i.e., whether a pedestrian chooses to cross under a given window of time to cross) has been extensively investigated. However, the dynamic nature of vehicle trajectories in the context of non-verbal communication has not been systematically explored. Our work provides evidence that trajectory dynamics, such as changes in TTA, can be powerful signals in the non-verbal communication between drivers and pedestrians. Moreover, we investigate these effects in both simulated and realworld datasets, both larger than have previously been considered in literature to the best of our knowledge.
We present a micro-traffic simulation (named DeepTraffic) where the perception, control, and planning systems for one of the cars are all handled by a single neural network as part of a model-free, off-policy reinforcement learning process. The primary goal of DeepTraffic is to make the hands-on study of deep reinforcement learning accessible to thousands of students, educators, and researchers in order to inspire and fuel the exploration and evaluation of DQN variants and hyperparameter configurations through large-scale, open competition. This paper investigates the crowd-sourced hyperparameter tuning of the policy network that resulted from the first iteration of the DeepTraffic competition where thousands of participants actively searched through the hyperparameter space with the objective of their neural network submission to make it onto the top-10 leaderboard.
We present a traffic simulation named DeepTraffic where the planning systems for a subset of the vehicles are handled by a neural network as part of a model-free, off-policy reinforcement learning process. The primary goal of DeepTraffic is to make the hands-on study of deep reinforcement learning accessible to thousands of students, educators, and researchers in order to inspire and fuel the exploration and evaluation of deep Q-learning network variants and hyperparameter configurations through large-scale, open competition. This paper investigates the crowd-sourced hyperparameter tuning of the policy network that resulted from the first iteration of the DeepTraffic competition where thousands of participants actively searched through the hyperparameter space.