We present a case study of productive flyby users (PFB users) on a recommendation website. These users exhibit counterintuitive behavior: they input a large amount of data during their first visit but never return. This phenomenon can have both positive and negative impacts on the system. On the positive side, their high productivity contributes a substantial amount of data. On the negative side, they may input inappropriate ratings that violate the assumptions of recommendation algorithms, potentially undermining system performance. To better understand the nature and causes of this behavior, we investigated their motivations, expectations, reasons for leaving, and the potential risks associated with their ratings using a mixed-methods approach. Specifically, we conducted interviews with 11 users, surveyed 41 users, and analyzed the impact of 1,000 PFB users on the performance of recommendation algorithms for regular users. Our findings revealed diverse motivations among PFB users. Some engaged with the system merely to pass the time, while others had unrealistic expectations of the recommender system. Regarding rating quality, 27% of surveyed users admitted to rating movies they had not seen, citing reasons such as browsing too quickly or attempting to manipulate the algorithm. Notably, users who reported leaving because they were "just killing time and forgot about the website" were the most likely to rate unseen movies. Overall, PFB users significantly influence recommendation algorithms and their performance for regular users. While some subgroups negatively affect prediction accuracy, others provide valuable data contributions. We discuss strategies for recommendation websites to better support these users or mitigate the impact of inappropriate ratings.
An increasingly important aspect of designing recommender systems involves considering how recommendations will influence consumer choices. This paper addresses this issue by introducing a method for collecting user beliefs about un-experienced goods - a critical predictor of choice behavior. We implemented this method on the MovieLens platform, resulting in a rich dataset that combines user ratings, beliefs, and observed recommendations. We document challenges to such data collection, including selection bias in response and limited coverage of the product space. This unique resource empowers researchers to delve deeper into user behavior and analyze user choices absent recommendations, measure the effectiveness of recommendations, and prototype algorithms that leverage user belief data, ultimately leading to more impactful recommender systems. The dataset can be found at https://grouplens.org/datasets/movielens/ml_belief_2024/.
We conduct a 6 month field experiment on a movie-recommendation platform to identify if and how recommendation systems affect consumption. We use within-consumer randomization at the good level and elicit beliefs about unconsumed goods to disentangle exposure from informational effects. We have three experimental groups: (a) control, (b) exposed, and (c) recommended + exposed goods where only goods in (c) are recommended and we elicit beliefs about goods in (b) and (c). Comparing across these treatment arms we find recommendations increase consumption beyond its role in exposing goods to consumers. We provide support for an informational mechanism: recommendations affect consumers' beliefs, which in turn explain consumption. Recommendations reduce uncertainty about goods consumers are most uncertain about and induce information acquisition. Finally, we find evidence for spatial correlation in beliefs.
Collaborative filtering algorithms find useful patterns in rating and consumption data and exploit these patterns to guide users to good items. Many of these patterns reflect important real-world phenomena driving interactions between the various users and items; other patterns may be irrelevant or reflect undesired discrimination, such as discrimination in publishing or purchasing against authors who are women or ethnic minorities. In this work, we examine the response of collaborative filtering recommender algorithms to the distribution of their input data with respect to one dimension of social concern, namely content creator gender. Using publicly available book ratings data, we measure the distribution of the genders of the authors of books in user rating profiles and recommendation lists produced from this data. We find that common collaborative filtering algorithms tend to propagate at least some of each user’s tendency to rate or read male or female authors into their resulting recommendations, although they differ in both the strength of this propagation and the variance in the gender balance of the recommendation lists they produce. The data, experimental design, and statistical methods are designed to be reusable for studying potentially discriminatory social dimensions of recommendations in other domains and settings as well.
Online technologies offer great promise to expand models of delivery for therapeutic interventions to help users cope with increasingly common mental illnesses like anxiety and depression. For example, "cognitive reappraisal" is a skill that involves changing one's perspective on negative thoughts in order to improve one's emotional state. In this work, we present Flip*Doubt, a novel crowd-powered web application that provides users with cognitive reappraisals ("reframes") of negative thoughts. A one-month field deployment of Flip*Doubt with 13 graduate students yielded a data set of negative thoughts paired with positive reframes, as well as rich interview data about how participants interacted with the system. Through this deployment, our work contributes: (1) an in-depth qualitative understanding of how participants used a crowd-powered cognitive reappraisal system in the wild; and (2) detailed codebooks that capture informative context about negative input thoughts and reframes. Our results surface data-derived hypotheses that may help to explain what types of reframes are helpful for users, while also providing guidance to future researchers and developers interested in building collaborative systems for mental health. In our discussion, we outline implications for systems research to leverage peer training and support, as well as opportunities to integrate AI/ML-based algorithms to support the cognitive reappraisal task. (Note: This paper includes potentially triggering mentions of mental health issues and suicide.)
This article contains a set of datasets used to demonstrate a strong regularity in inter-activity time. See the paper: User Session Identification Based on Strong Regularities in Inter-activity Time http://arxiv.org/abs/1411.2878 Abstract Session identification is a common strategy used to develop metrics for web analytics and behavioral analyses of user-facing systems. Past work has argued that session identification strategies based on an inactivity threshold is inherently arbitrary or advocated that thresholds be set at about 30 minutes. In this work, we demonstrate a strong regularity in the temporal rhythms of user initiated events across several different domains of online activity (incl. video gaming, search, page views and volunteer contributions). We describe a methodology for identifying clusters of user activity and argue that regularity with which these activity clusters appear implies a good rule-of-thumb inactivity threshold of about 1 hour. We conclude with implications that these temporal rhythms may have for system design based on our observations and theories of goal-directed human activity.
University of Minnesota Ph.D. dissertation. August 2018. Major: Computer Science. Advisor: Joseph Konstan. 1 computer file (PDF); x, 274
Recommender systems help users find information by recommending content that a user might not know about, but will hopefully like. Rating-based collaborative filtering recommender systems do this by finding patterns that are consistent across the ratings of other users. These patterns can be used on their own, or in conjunction with other forms of social information access to identify and recommend content that a user might like. This chapter reviews the concepts, algorithms, and means of evaluation that are at the core of collaborative filtering research and practice. While there are many recommendation algorithms, the ones we cover serve as the basis for much of past and present algorithm development. After presenting these algorithms we present examples of two more recent directions in recommendation algorithms: learning-to-rank and ensemble recommendation algorithms. We finish by describing how collaborative filtering algorithms can be evaluated, and listing available resources and datasets to support further experimentation. The goal of this chapter is to provide the basis of knowledge needed for readers to explore more advanced topics in recommendation.
Emoji are popular in digital communication, but they are rendered differently on different viewing platforms (e.g., iOS, Android). It is unknown how many people are aware that emoji have multiple renderings, or whether they would change their emoji-bearing messages if they could see how these messages render on recipients' devices. We developed software to expose the multi-rendering nature of emoji and explored whether this increased visibility would affect how people communicate with emoji. Through a survey of 710 Twitter users who recently posted an emoji-bearing tweet, we found that at least 25% of respondents were unaware that the emoji they posted could appear differently to their followers. Additionally, after being shown how one of their tweets rendered across platforms, 20% of respondents reported that they would have edited or not sent the tweet. These statistics reflect millions of potentially regretful tweets shared per day because people cannot see emoji rendering differences across platforms. Our results motivate the development of tools that increase the visibility of emoji rendering differences across platforms, and we contribute our cross-platform emoji rendering software to facilitate this effort.
As the gig economy continues to grow and freelance work moves online, five-star reputation systems are becoming more and more common. At the same time, there are increasing accounts of race and gender bias in evaluations of gig workers, with negative impacts for those workers. We report on a series of four Mechanical Turk-based studies in which participants who rated simulated gig work did not show race- or gender bias, while manipulation checks showed they reliably distinguished between low- and high-quality work. Given prior research, this was a striking result. To explore further, we used a Bayesian approach to verify absence of ratings bias (as opposed to merely not detecting bias). This Bayesian test let us identify an upper- bound: if any bias did exist in our studies, it was below an average of 0.2 stars on a five-star scale. We discuss possible interpretations of our results and outline future work to better understand the results.
Recent studies have found that people interpret emoji characters inconsistently, creating significant potential for miscommunication. However, this research examined emoji in isolation, without consideration of any surrounding text. Prior work has hypothesized that examining emoji in their natural textual contexts would substantially reduce potential for miscommunication. To investigate this hypothesis, we carried out a controlled study with 2,482 participants who interpreted emoji both in isolation and in multiple textual contexts. After comparing the variability of emoji interpretation in each condition, we found that our results do not support the hypothesis in prior work: when emoji are interpreted in textual contexts, the potential for miscommunication appears to be roughly the same. We also identify directions for future research to better understand the interplay between emoji and textual context.
Recommender systems are not one-size-fits-all; different algorithms and data sources have different strengths, making them a better or worse fit for different users and use cases. As one way of taking advantage of the relative merits of different algorithms, we gave users the ability to change the algorithm providing their movie recommendations and studied how they make use of this power. We conducted our study with the launch of a new version of the MovieLens movie recommender that supports multiple recommender algorithms and allows users to choose the algorithm they want to provide their recommendations. We examine log data from user interactions with this new feature to under-stand whether and how users switch among recommender algorithms, and select a final algorithm to use. We also look at the properties of the algorithms as they were experienced by users and examine their relationships to user behavior. We found that a substantial portion of our user base (25%) used the recommender-switching feature. The majority of users who used the control only switched algorithms a few times, trying a few out and settling down on an algorithm that they would leave alone. The largest number of users prefer a matrix factorization algorithm, followed closely by item-item collaborative filtering; users selected both of these algorithms much more often than they chose a non-personalized mean recommender. The algorithms did produce measurably different recommender lists for the users in the study, but these differences were not directly predictive of user choice.
Session identification is a common strategy used to develop metrics for web analytics and perform behavioral analyses of user-facing systems. Past work has argued that session identification strategies based on an inactivity threshold is inherently arbitrary or has advocated that thresholds be set at about 30 minutes. In this work, we demonstrate a strong regularity in the temporal rhythms of user initiated events across several different domains of online activity (incl. video gaming, search, page views and volunteer contributions). We describe a methodology for identifying clusters of user activity and argue that the regularity with which these activity clusters appear implies a good rule-of-thumb inactivity threshold of about 1 hour. We conclude with implications that these temporal rhythms may have for system design based on our observations and theories of goal-directed human activity.
The new user experience is one of the important problems in recommender systems. Past work on recommending for new users has focused on the process of gathering information from the user. Our work focuses on how different algorithms behave for new users. We describe a methodology that we use to compare representatives of three common families of algorithms along eleven different metrics. We find that for the first few ratings a baseline algorithm performs better than three common collaborative filtering algorithms. Once we have a few ratings, we find that Funk's SVD algorithm has the best overall performance. We also find that ItemItem, a very commonly deployed algorithm, performs very poorly for new users. Our results can inform the design of interfaces and algorithms for new users.
GroupLens Research-Department of Computer Science, University of Minnesota, Twin Cities, USA The authors thank Ting-Yu Wang for his contribution in developing the user interface. Authors would also like to thank Shilad Sen and the rest of GroupLens for their feed- back on the ideas presented in this paper. We also thank the MovieLens community for their responses, feedback and suggestions. This paper is funded by National Science Foundation grant IIS 09-64695.
One of the challenges for recommender systems is that users struggle to accurately map their internal preferences to external measures of quality such as ratings. We study two methods for supporting the mapping process: (i) reminding the user of characteristics of items by providing personalized tags and (ii) relating rating decisions to prior rating decisions using exemplars. In our study, we introduce interfaces that provide these methods of support. We also present a set of methodologies to evaluate the efficacy of the new interfaces via a user experiment. Our results suggest that presenting exemplars during the rating process helps users rate more consistently, and increases the quality of the data.
Most recommender systems assume user ratings accurately represent user preferences. However, prior research shows that user ratings are imperfect and noisy. Moreover, this noise limits the measurable predictive power of any recommender system. We propose an information theoretic framework for quantifying the preference information contained in ratings and predictions. We computationally explore the properties of our model and apply our framework to estimate the efficiency of different rating scales for real world datasets. We then estimate how the amount of information predictions give to users is related to the scale ratings are collected on. Our findings suggest a tradeoff in rating scale granularity: while previous research indicates that coarse scales (such as thumbs up / thumbs down) take less time, we find that ratings with these scales provide less predictive value to users. We introduce a new measure, preference bits per second, to quantitatively reconcile this tradeoff.
Plant-population recovery across large disturbance areas is often seed-limited. An understanding of seed dispersal patterns is fundamental for determining natural-regeneration potential. However, forecasting seed dispersal rates across heterogeneous landscapes remains a challenge. Our objectives were to determine (i) the landscape patterning of post-disturbance seed dispersal, and underlying sources of variation and the scale at which they operate, and (ii) how the natural seed dispersal patterns relate to a seed augmentation strategy. Vertical seed trapping experiments were replicated across 2 years and five burned and/or managed landscapes in sagebrush steppe. Multi-scale sampling and hierarch- ical Bayesian models were used to determine the scale of spatial variation in seed dispersal. We then integrated an empirical and mechanistic dispersal kernel for wind-dispersed species to project rates of seed dispersal and compared natural seed arrival to typical post-fire aerial seeding rates. Seeds were captured across the range of tested dispersal distances, up to a maximum distance of 26 m from seed-source plants, although dispersal to the furthest traps was variable. Seed dispersal was better explained by transect heterogeneity than by patch or site hetero- geneity (transects were nested within patch within site). The number of seeds captured varied from a modelled mean of ~13 m −2 adjacent to patches of seed-producing plants, to nearly none at 10 m from patches, standardized over a 49-day period. Maximum seed dispersal distances on average were estimated to be 16 m according to a novel modelling approach using a ‘latent’ variable for dispersal distance based on seed trapping heights. Surprisingly, statistical representation of wind did not improve model fit and seed rain was not related to the large variation in total available seed of adjacent patches. The models predicted severe seed limitations were likely on typical burned areas, especially compared to the mean 95–250 seeds per m 2 that previous literature suggested were required to generate sagebrush recovery. More broadly, our Bayesian data fusion approach could be applied to other cases that require quantitative estimates of long-distance seed dispersal across heterogeneous landscapes.
Joseph A. Konstan合作论文数Department of Computer Science and Engineering, College of Science and Engineering, University of Minnesota7
John Riedl合作论文数Department of Computer Science and Engineering, College of Science and Engineering, University of Minnesota3