User profiling refers to inferring people's attributes of interest (AoIs) like gender and occupation, which enables various applications ranging from personalized services to collective analyses. Massive nonlinguistic audio data brings a novel opportunity for user profiling due to the prevalence of studying spontaneous face-to-face communication. In this poster, we are the first to build a user profiling system to infer gender and personality based on nonlinguistic audio. Instead of linguistic or acoustic features which are unable to extract, we focus on conversational features that could reflect AoIs. We firstly develop an adaptive voice activity detection algorithm that could address individual differences in voice and false-positive voice activities caused by people nearby. Secondly, we propose a gender-assisted multi-task learning method to combat dynamics in human behavior by integrating gender differences and the correlation of personality traits. The experimental evaluation of 100 people in 273 meetings indicates the superiority of the proposed method in gender identification and personality recognition respectively.
We use electronics badges to measure in-person communication across companies from an accelerator program, and analyze its relationship with their performance. Our analysis shows that both subjective and objective performance correlates with the amount communication exhibited by early stage companies. In general, more communication correlates with better performance, though too much communication with other teams seems harmful. Lower internal communication entropy correlates with higher performance. Companies that spent more time with the program mentors do better. Large companies reported higher levels of satisfaction compared to small companies.
. We use nonlinguistic audio analysis to investigate differences in speaking style of men and women in study groups discussions. Our results show that women took shorter turns than men and that women interrupted men more than men interrupted women. We found no significant differences in other turn taking characteristics.
Group gender is essential in understanding social interaction and group dynamics. With the increasing privacy concerns of studying face-to-face communication in natural settings, many participants are not open to raw audio recording. Existing voice-based gender identification methods rely on acoustic characteristics caused by physiological differences and phonetic differences. However, these methods might become ineffective with privacy-sensitive audio for two main reasons. First, compared to raw audio, privacy-sensitive audio contains significantly fewer acoustic features. Moreover, natural settings generate various uncertainties in the audio data. In this paper, we make the first attempt to identify group gender using privacy-sensitive audio. Instead of extracting acoustic features from privacy-sensitive audio, we focus on conversational features including turn-taking behaviors and interruption patterns. However, conversational behaviors are unstable in gender identification as human behaviors are affected by many factors like emotion and environment. We utilize ensemble feature selection and a two-stage classification to improve the effectiveness and robustness of our approach. Ensemble feature selection could reduce the risk of choosing an unstable subset of features by aggregating the outputs of multiple feature selectors. In the first stage, we infer the gender composition (mixed-gender or same-gender) of a group which is used as an additional input feature for identifying group gender in the second stage. The estimated gender composition significantly improves the performance as it could partially account for the dynamics in conversational behaviors. According to the experimental evaluation of 100 people in 273 meetings, the proposed method outperforms baseline approaches and achieves an F1-score of 0.77 using linear SVM.
To understand and manage complex organizations, we must develop tools capable of measuring human social interaction accurately and uniformly. Current technologies that measure face-to-face communication do not measure interaction in a unified manner and often ignore remote interaction, an increasingly common communication modality. In this article we present Rhythm, a platform that combines wearable electronic badges and online applications to capture team-level and network-level interaction patterns in organizations. The platform measures conversation time, turn-taking behavior, and the physical proximity of both co-located and distributed members. Our goal is to empower organizations and researchers to measure formal and informal social interaction across teams, divisions, and locations. We describe two pilot studies that use this platform and discuss how measurement systems like Rhythm may further the fields of computational social science and organizational design.
Today's age of data holds high potential to enhance the way we pursue and monitor progress in the fields of development and humanitarian action. We study the relation between data utility and privacy risk in large-scale behavioral data, focusing on mobile phone metadata as paradigmatic domain. To measure utility, we survey experts about the value of mobile phone metadata at various spatial and temporal granularity levels. To measure privacy, we propose a formal and intuitive measure of reidentification risk$\unicode{x2014}$the information ratio$\unicode{x2014}$and compute it at each granularity level. Our results confirm the existence of a stark tradeoff between data utility and reidentifiability, where the most valuable datasets are also most prone to reidentification. When data is specified at ZIP-code and hourly levels, outside knowledge of only 7% of a person's data suffices for reidentification and retrieval of the remaining 93%. In contrast, in the least valuable dataset, specified at municipality and daily levels, reidentification requires on average outside knowledge of 51%, or 31 data points, of a person's data to retrieve the remaining 49%. Overall, our findings show that coarsening data directly erodes its value, and highlight the need for using data-coarsening, not as stand-alone mechanism, but in combination with data-sharing models that provide adjustable degrees of accountability and security.
We present Open Badges, an open-source framework an toolkit for measuring and shaping face-to-face social interactions using either custom hardware devices or smart phones, and real-time web-based visualizations. Open Badges is a modular system that allows researchers to monitor and collect interaction data from people engaged in real-life social settings. In this paper we describe the technical aspects of the Open Badges project and the motivation for its creation.
We present Breakout, a group interaction platform for online courses that enables the creation and measurement of face-to-face peer learning groups in online settings. Breakout is designed to help students easily engage in synchronous, video breakout session based peer learning in settings that otherwise force students to rely on asynchronous text-based communication. The platform also offers data collection and intervention tools for studying the communication patterns inherent in online learning environments. The goals of the system are twofold: to enhance student engagement in online learning settings and to create a platform for research into the relationship between distributed group interaction patterns and learning outcomes.
Optimizing the use of available resources is one of the key challenges in activities that consist of interactions with a large number of “target individuals”, with the ultimate goal of affecting as many of them as possible, such as in marketing, service provision and political campaigns. Typically, the cost of interactions is monotonically increasing such that a method for maximizing the performance of these campaigns is required. This chapter proposes a mathematical model to compute an optimized campaign by automatically determining the number of interacting units and their type, and how they should be allocated to different geographical regions in order to maximize the campaign's performance. The proposed model is validated using real world mobility data.
Optimizing the use of available resources is one of the key challenges in activities that consist of interactions with a large number of “target individuals,” with the ultimate goal of “winning” as many of them as possible, such as in marketing, service provision, political campaigns, or homeland security. Typically, the cost of interactions is monotonically increasing such that a method for maximizing the performance of these campaigns iPs required. In this paper, we propose a mathematical model to compute an optimized campaign by automatically determining the number of interacting units and their type, and how they should be allocated to different geographical regions in order to maximize the campaign's performance. We validate our proposed model using real world mobility data.