Depression is a serious and prevalent mental illness. The ubiquitous adoption of smartphones have enabled new opportunities for depression screening. Recently studies have used physical location and activity information automatically collected on smartphones for depression prediction. Social interactions also play a vital role in the overall health and well-being of individuals. In this work, we explore the feasibility of using social interaction data, specifically SMS and phone call logs, collected on smartphones for predicting depression. We extract a comprehensive set of features from such data. In addition, we construct a family of machine learning models by using these features for depression prediction. Using the social interaction data collected via an Android phone app from college-age students, we compare the characteristics of SMS and phone call usage patterns between depressed and non-depressed participants. We find that they exhibit more distinguishing behaviors in outgoing SMS messages and phone calls, which are initiated by the users, than incoming SMS messages and phone calls. Our results also demonstrate that social interaction data can be used to predict depression effectively, with F 1 score as high as 0.82.
Recent studies have demonstrated that geographic location features collected using smartphones can be a powerful predictor for depression. While location information can be conveniently gathered by GPS, typical datasets suffer from significant periods of missing data due to various factors (e.g., phone power dynamics, limitations of GPS). A common approach is to remove the time periods with significant missing data before data analysis. In this paper, we develop an approach that fuses location data collected from two sources: GPS and WiFi association records. Our evaluation demonstrates that our data fusion approach leads to significantly more complete data, which improves feature extraction and depression screening.
Depression is a serious mental health problem. Recently, researchers have proposed novel approaches that use sensing data collected passively on smartphones for automatic depression screening. While these studies have explored several types of sensing data (e.g., location, activity, conversation), none of them has leveraged Internet traffic of smartphones, which can be collected with little energy consumption and the data is insensitive to phone hardware. In this paper, we explore using coarse-grained meta-data of Internet traffic on smartphones for depression screening. We develop techniques to identify Internet usage sessions (i.e., time periods when a user is online) and extract a novel set of features based on usage sessions from the Internet traffic meta-data. Our results demonstrate that Internet usage features can reflect the different behavioral characteristics between depressed and non-depressed participants, confirming findings in psychological sciences, which have relied on surveys or questionnaires instead of real Internet traffic as in our study. Furthermore, we develop machine learning based prediction models that use these features to predict depression. Our evaluation shows that Internet usage features can be used for effective depression prediction, leading to F 1 score as high as 0.80.
Depression is a serious mental illness. The symptoms associated with depression are both behavioral (in appetite, energy level, sleep) and cognitive (in interests, mood, concentration). Currently, survey instruments are commonly used to keep track of depression symptoms, which are burdensome and difficult to execute on a continuous basis. In this paper, we explore the feasibility of predicting all major categories of depressive symptoms automatically using smartphone data. Specifically, we consider two types of smartphone data, one collected passively on smartphones (through an app running on the phones) and the other collected from an institution's WiFi infrastructure (that does not require direct data capture on the phones), and construct a family of machine learning based models for the prediction. Both scenarios require no efforts from the users, and can provide objective assessment on depressive symptoms. Using smartphone data collected from 182 college students in a two-phase study, our results demonstrate that smartphone data can be used to predict both behavioral and cognitive symptoms effectively, with F1 score as high as 0.86. Our study makes a significant step forward over existing studies that only focus on predicting overall depression status (i.e., whether one is depressed or not).
Depression is a serious public health problem. Current diagnosis techniques rely on physician-administered or patient self-administered interview tools, which are burdensome and suffer from recall bias. Recent studies have proposed new approaches that use sensing data collected on smartphones to serve as "human sensors" for automatic depression screening. These approaches, however, require running an app on the phones for continuous data collection. We explore a novel approach that uses data collected from WiFi infrastructure for large-scale automatic depression screening. Specifically, when smartphones connect to a WiFi network, their locations (and hence the locations of the users) can be determined by the access points that they associate with; the location information over time provides important insights into the behavior of the users, which can be used for depression screening. To investigate the feasibility of this approach, we have analyzed two datasets, each collected over several months, involving tens of participants recruited from a university. Our results demonstrate that WiFi meta-data is effective for passive depression screening: the F1 scores are as high as 0.85 for predicting depression, comparable to those obtained by using sensing data collected directly from smartphones.
Depression is a common mood disorder that causes severe medical problems and interferes negatively with daily life. Identifying human behavior patterns that are predictive or indicative of depressive disorder is important. Clinical diagnosis of depression relies on costly clinician assessment using survey instruments which may not objectively reflect the fluctuation of daily behavior. Self-administered surveys, such as the Quick Inventory of Depressive Symptomatology (QIDS) commonly used to monitor depression, may show disparities from clinical decision. Smartphones provide easy access to many behavioral parameters, and Fitbit wrist bands are becoming another important tool to assess variables such as heart rates and sleep efficiency that are complementary to smartphone sensors. However, data used to identify depression indicators have been limited to a single platform either iPhone, or Android, or Fitbit alone due to the variation in their methods of data collection. The present work represents a large-scale effort to collect and integrate data from mobile phones, wearable devices, and self reports in depression analysis by designing a new machine learning approach. This approach constructs sparse mappings from sensing variables collected by various tools to two separate targets: self-reported QIDS scores and clinical assessment of depression severity. We propose a so-called heterogeneous multi-task feature learning method that jointly builds inference models for related tasks but of different types including classification and regression tasks. The proposed method was evaluated using data collected from 103 college students and could predict the QIDS score with an R2 reaching 0.44 and depression severity with an F1-score as high as 0.77. By imposing appropriate regularizers, our approach identified strong depression indicators such as time staying at home and total time asleep.
Wireless network traces have been widely used to understand human behaviors and provide value-added services. Sanitization based techniques have been shown to be severely lacking in protecting sensitive user information embedded in such traces. In this paper, we take an encryption based approach that provides much stronger protection of user privacy. One challenge in encrypting wireless network traces is how to encrypt time range while maintaining the utility of the traces. We propose two practical encryption techniques to support queries that involve time range. These two techniques provide much stronger security guarantee than existing order preserving encryption schemes, and present different tradeoffs in complexity, as well as storage and network bandwidth requirement. Last, we quantify the performance of the proposed approach using a smart campus prototype. The results show that our approach only leads to moderate increase in storage, network bandwidth and computation overhead, demonstrating the practicality of our approach.
Depression is a serious health disorder. In this study, we investigate the feasibility of depression screening using sensor data collected from smartphones. We extract various behavioral features from smartphone sensing data and investigate the efficacy of various machine learning tools to predict clinical diagnoses and PHQ-9 scores (a quantitative tool for aiding depression screening in practice). A notable feature of our study is that we leverage a dataset that includes clinical ground truth. We find that behavioral data from smartphones can predict clinical depression with good accuracy. In addition, combining behavioral data and PHQ-9 scores can provide prediction accuracy significantly exceeding each in isolation, indicating that behavioral data captures relevant features that are not reflected by PHQ-9 scores. Finally, we develop multi-feature regression models for PHQ-9 scores that achieve significantly improved accuracy compared to direct regression models based on single features.
Depression is a major public health issue with direct and significant effects on both physical and mental health. In this study, we analyze smartphone sensing data to find differential behavioral features that are correlated with depression measures such as patient health questionnaire (PHQ-9). Our approach uses an innovative multi-view bi-clustering algorithm. It takes multiple views of sensing data as input to identify homogeneous behavioral groups and simultaneously the key sensing features that characterize the different groups. Using a publicly available dataset, we discover that these behavioral groups with differential sensing features are highly discriminative of PHQ-9 scores that are self reported by the study subjects. For instance, the group comprising less active users in the sensed activities corresponds to overall higher PHQ-9 scores. We then employ the key sensing features that distinguish the different groups to create predictive models to predict the group assignment of individuals. We verify the generalizability of these models using the support vector machine classifier. Cross validation studies show that our classifiers can classify individuals into the correct subgroups with an overall accuracy of 87%.
Online service providers often use challenge questions (a.k.a. knowledge‐based authentication) to facilitate resetting of passwords or to provide an extra layer of security for authentication. While prior schemes explored both static and dynamic challenge questions to improve security, they do not systematically investigate the problem of designing challenge questions and its effect on user recall performance. Interestingly, as answering different styles of questions may require different amount of cognitive effort and evoke different reactions among users, we argue that the style of challenge questions itself can have a significant effect on user recall performance and usability of such systems. To address this void and investigate the effect of question types on user performance, this paper explores location‐based challenge question generation schemes where different types of questions are generated based on users’ locations tracked by smartphones and presented to users. For evaluation, we deployed our location tracking application on users’ smartphones and conducted two real‐life studies using four different kinds of challenge questions. Each study was approximately 30 days long and had 14 and 15 users respectively. Our findings suggest that the question type can have a significant effect on user performance. Finally, as individual users may vary in terms of performance and recall rate, we investigate and present a Bayesian classifier based authentication algorithm that can authenticate legitimate users with high accuracy by leveraging individual response patterns while reducing the success rate of adversaries.
This paper investigates a location-based authentication system where authentication questions are generated based on users' locations tracked by smartphones. More specifically, the system builds a location profile for a user based on periodically logged Wi-Fi access point beacons over time, and leverages this location profile to generate authentication questions. To evaluate the various aspects of this location-based authentication approach, we deployed the application on users' smartphones and conducted a real-life study for one month with 14 users. To simulate various kinds of adversaries (e.g., Naive vs. Knowledgeable), in our study, we recruited volunteers in pairs (e.g., Friends), in addition to single participants. Over the course of the experiment, each user is periodically presented with two sets of authentication questions. The first set is generated based on a user's own data. The second set is generated based on a randomly selected user's data. Additionally, in cases of paired participants, each user is presented with a third set of questions which is generated based on the user's friend's data. In each case, three different kinds of questions of varying difficulty levels are generated and presented to the user. Finally, we present a Bayesian classifier based authentication algorithm that can authenticate legitimate users with high accuracy by leveraging individual response patterns. We also discuss various aspects of location-based authentication mechanisms based on our findings in this paper.
Integration of NFC radios into smartphones is expediting the adoption of mobile devices as the preferred method for accessing physical locations, bank accounts, and other valuable resources. The pervasive nature of authentication using these mobile devices, however, comes with increased security considerations stemming from the possibility of physical loss of the device. To minimize the risk caused by stolen devices, this paper introduces a method for confirming the identity of a device's user based on her recent macroscopic behavior over space and time. The user's behavior is continuously recorded by a set of devices embedded in the environment (e.g., Wi-Fi Access Point) and used to train a probabilistic n-gram model. Subsequently, deviations caused by stolen devices can be detected by comparing the user's recent behavior against the trained model. Our first evaluation results demonstrate the ability of the proposed approach to detect anomalies in the user's behavior without generating a significant number of false alarms.
Despite much progress in emergency management, effective techniques for real-time tracking of emergency events are still lacking. We envision a promising direction to achieve real-time emergency tracking is through widely adopted smartphones. In this paper, we explore the first step in achieving this goal, namely, locating emergency in real time using smartphones. Our main contribution is a novel approach that locates emergencies by analyzing AP (access point) association events collected from a campus Wi-Fi network. It is motivated by the observation that human behavior and mobility pattern are significantly altered in the face of emergency, which is reflected in how their smartphones associate with the APs in the network. More specifically, our approach locates emergency by discovering APs with abnormal association patterns using Extreme Value Theory (EVT). Preliminary evaluation using real data collected from a university campus network demonstrates the effectiveness of our approach.
The thesis of this dissertation is that by analyzing the temporal properties of sensor measurements, patterns generated by macroscopic behaviors can be discovered and used to form virtual sensors that can convert low-level sensor events into actionable knowledge. Macroscopic behaviors are defined as human activities and routines that evolve over large spaces and extended periods of time and can thus be learned only in an unsupervised manner. Our work is based on two key observations: (1) most human behaviors are sequences of very primitive actions or events that can either be sensed directly by sensors or indirectly by pre-processing the sensor data; and (2) the same type of human activity will often trigger a similar pattern of sensor events in space and time. Based on these observations, this dissertation introduces three common types of macroscopic behaviors and proposes methods for their recognition. The first concerns the problem of identifying periodic human activities. The second focuses on discovering classes of frequent Spatio-Temporal Activities (STAs) from location traces, namely areas that people consistently spend time at approximately the same time intervals every day. The third problem deals with the detection of simple group behaviors specified in the form of sequences of interactions. The ability to define virtual sensors for these macroscopic behaviors is in the core of the BehaviorScope system — a human-centric sensing system aiming to offer real-world services as assisted living and power efficiency in large buildings. In the latter case we show how our system can potentially reduce the electricity consumption of an office building by up to 33.8%.
Given the ongoing widespread deployment of low frequency electricity sub-metering devices at residential and commercial buildings, fine-grained usage information of end-loads can bring a new powerful sensing modality in Cyber-Physical Systems (CPS). Motivated by the opportunity, this paper describes an algorithm of estimating the ON/OFF sequences for typical household end-loads in close-to-real-time using an off-the-shelf power meter. Unlike previous algorithms that lacks in scalability to support diverse applications in CPS our algorithm is designed to provide control knobs to support various trade-offs between accuracy and computation load or delay to satisfy the different application requirements. We experimentally verify the proposed algorithm using a collection of home appliances. Our experiment result shows that our algorithm is able to detect ON/OFF sequences of 7 appliances nearly without error and 3 appliances with moderate error rate less than 6% among 12 typical household appliances.
Integration of Near Field Communication (NFC) sensors into mobile devices has enabled their use for authentication. The ubiquitous nature of authentication using mobile devices comes though with increased security considerations. In addition to the risk of being stolen, mobile devices are increasingly susceptible to different types of software attacks. User-provided passwords such as Personal Identification Numbers (PINs) are often employed to ameliorate these limitations. The use of passwords though, has its own vulnerabilities that are mainly caused by the passwords' static nature and low entropy. To eliminate the security threats caused by untrusted devices and static passwords, we propose the development of new types of biometric authentication, based on macroscopic human behavior.
This poster paper describes a method for estimating the ON/OFF usage profile for appliances inside a home using an off-the-Unlike previous approaches that put their emphasis on the detection of individual events using high frequency sampling, our approach aims to reliably detect sequences of ON/OFF events from traces of less reliable ON/OFF events identified on data from low-frequency sampling. The proposed algorithm is evaluated experimentally using a collection of home appliances.
A large collection of mobile sensing applications depend on the knowledge of the user's whereabouts and are heavily based on GPS location measurements. Although knowledge of location is very desirable, in many mobile applications excessive GPS sampling is very energy taxing thus posing a barrier to application sustainability. To mitigate this problem, in this paper we examine how to reduce GPS sensing redundancies by extracting the state of a person and using it to drive GPS sampling on mobile phones. Using a GPS dataset we first describe how to extract the spatio-temporal states of the user. We then use the knowledge of the user's state to reduce GPS sampling rate, helping to make mobile applications more sustainable.
This paper presents an automated methodology for extracting the spatiotemporal activity model of a person using a wireless sensor network deployed inside a home. The sensor network is modeled as a source of spatiotemporal symbols whose output is triggered by the monitored person’s motion over space and time. Using this stream of symbols, the problem of human activity modeling is formulated as a spatiotemporal pattern-matching problem on top of the sequence of symbolic information the sensor network produces, and is solved using an exhaustive search algorithm. The effectiveness of the proposed methodology is demonstrated on a real 30-day dataset extracted from an ongoing deployment of a sensor network inside a home monitoring an elder. The developed algorithm examines the person’s data over these 30 days and automatically extracts the person’s daily pattern.
Alexander Russell合作论文数Department of Computer Science & Engineering;University of Connecticut8
Dimitrios Lymberopoulos合作论文数Microsoft Research7