As AI-generated content (e.g., "slop") becomes more prevalent online, people are developing strategies to attempt to identify it (or, conversely, to gain confidence that something is not AI-generated). What strategies are people using, and how are they changing over time as generative AI models themselves change? In this work, we catalog and analyze 2 years and 8 months of the AI detection strategies discussed by users of two popular Reddit communities (r/isthisAI and r/RealOrAI) that use the wisdom of crowds to identify AI-generated media. Through a mixed-method analysis of 13,098 posts and 222,060 comments within these communities, we catalog and analyze the prevalence of 12 AI-detection strategies, including examining fine-grained physical details, recognizing trends in AI-created content, and the assumptions people make about what models are capable of producing. Furthermore, we find that these strategies and mental models shift over time in accordance with changing AI capabilities and in response to online social trends. By systematically cataloging users' AI detection strategies, we lay the groundwork for user-facing guidance and future research.
Synthetic nonconsensual explicit imagery, also referred to as “deepfake nudes”, is becoming faster and easier to generate. In the last year, synthetic nonconsensual explicit imagery was reported in at least ten US middle and high schools, generated by students of other students. Teachers are at the front lines of this new form of image abuse and have a valuable perspective on threat models in this context. We interviewed 17 US teachers to understand their opinions and concerns about synthetic nonconsensual explicit imagery in schools. No teachers knew of it happening at their schools, but most expected it to be a growing issue. Teachers proposed many interventions, such as improving reporting mechanisms, focusing on consent in sex education, and updating technology policies. However, teachers disagreed about appropriate consequences for students who create such images. We unpack our findings relative to differing models of justice, sexual violence, and sociopolitical challenges within schools.
Ads are often designed visually, with images and videos conveying information. In this work, we study the accessibility of ads on the web to users of screen readers. We approach this in two ways: first, we conducted a measurement and analysis of 90 websites over a month, collecting ads and auditing their behavior against a subset of best practices established by the Web Content Accessibility Guidelines (WCAG). Then, to put our measurement findings in context, we interviewed 13 blind participants who navigate the web with a screen reader to understand their experiences with (in)accessible ads. We find that the overall web ad ecosystem is fairly inaccessible in multiple ways: many images are missing alt-text, unlabeled links make it confusing for folks to navigate, and closing ads can be tricky. But, there are straightforward ways to improve: because only a few large companies dominate the ad ecosystem, making small changes to the way they enforce accessibility standards can make a large difference.
Online ads are a major source of information on the web. The mass reach of online advertising is often leveraged for information dissemination, at times with an objective to influence public opinion (e.g., election misinformation). We hypothesized that online advertising, due to its reach and potential, might have been used to spread information around the 2022 Russian invasion of Ukraine. Thus, to understand the online ad ecosystem during this conflict, we conducted a five-month long large-scale measurement study of online advertising in Ukraine, Russia, and the US. We studied advertising trends of ad platforms that delivered ads in Ukraine, Russia, and the US and conducted an in-depth qualitative analysis of the conflict-related ad content. We found that prominent US-based advertisers continued to support Russian websites, and a portion of online ads were used to spread conflict-related information, including protesting the invasion, and spreading awareness, which might have otherwise potentially been censored in Russia.
In addition to being a health and fitness band, the Amazon Halo offers users information about how their voices sound, i.e., their ‘tones’. The Halo’s tone analysis capability leverages machine learning, which can lead to potentially biased inferences. We develop an auditing framework to evaluate the Amazon Halo’s tone analysis capabilities for gender biases. Our results show that the Halo exhibits statistically significant gender biases, when the same emotion is conveyed by professional women and men actors through their recorded voices. For example, we find that over 75% of the words used by the Halo to describe men’s emotions are positive whereas fewer than 50% of the words used by the Halo to describe women’s voices are positive. The Halo describes women as being ‘angry’, ‘disappointed’, ‘uncomfortable’, and ‘annoyed’ more often than men (adjectives with negative valence). The Halo describes men as being ‘knowledgeable’, ‘confident’, and ‘focused’ more often than women (adjectives with positive valence). Overall, our findings underscore that even commercially deployed ML models for day-to-day consumer use exhibit strong biases.
—Microtask platforms aim to pair employers and workers to complete small tasks for modest pay, but the purposes of these tasks are not always benign. Researchers have identified problematic tasks on popular microtask platforms like Amazon Mechanical Turk, and such tasks may be relatively more common on smaller platforms. Recent work examining these smaller alternative platforms is limited, and the nature of the work on these platforms evolves with the desires of employers and the practices of the platforms. We provide an up-to-date view of the work available via alternative microtask platforms. To do so, we collected details from three alternative platforms over approximately a month, categorizing the available work. We find that potentially abusive work persists in well-known categories like search engine optimization, but we also uncovered new and emerging categories of work, such as tasks that may manipulate spam filters. We comprehensively explore these categories and discuss potential mitigation approaches.
If a security feature requires user data, concerns over secondary uses of that data may influence user adoption of the feature. We explore secondary uses of phone numbers that users share for two-factor authentication. Some companies have reused these numbers for purposes unrelated to security, such as targeted advertising. Focusing on top sites, we assessed user-observable secondary uses of phone numbers in two ways. First, we examined web traffic for evidence that sites share numbers with third parties when the user enrolls in two-factor authentication. Second, we monitored calls, voicemail, and text messages to the phone numbers over a two-month period after enrollment. We observed neither form of secondary use in our analysis. Our results suggest a consistent norm against these secondary uses, with potential implications for companies considering practices that deviate from these norms.
Web skimming poses an increasingly serious threat to online shoppers, with recent cases reportedly occurring on prominent websites like British Airways and Newegg. While prior work largely measures the prevalence of these attacks and explores their technical details, we examine the use of credentials stolen via web skimming. We identified 50 sites apparently compromised to host payment card skimming code. For each site, we attempted to purchase items using a unique payment card. Over an eleven-month study period, we monitored the payment cards for signs of abuse. We observed attempted misuse of 15 of the 50 payment cards. With a single exception, the time from exposure of a payment card until observed misuse was at least 50 days. Thieves tried to use the 15 cards at least 45 times. These attempted payments ranged from $0.10 to $122.44 and totaled $1,342.91. To place our findings in context, we compare these results to a separate study in which we exposed payment data directly via an online paste site. Our observations suggest that the impact of web skimming may not be apparent for an extended period following an incident.
AbstractInternet advertising and analytics technology companies are increasingly trying to find ways to link behavior across the various devices consumers own. Thiscross-device trackingcan provide a more complete view into a consumer’s behavior and can be valuable for a range of purposes, including ad targeting, research, and conversion attribution. However, consumers may not be aware of how and how often their behavior is tracked across different devices. We designed this study to try to assess what information about cross-device tracking (including data flows and policy disclosures) is observable from the perspective of the end user. Our paper demonstrates how data that is routinely collected and shared online could be used by online third parties to track consumers across devices.