Malvertising is a malicious activity that leverages advertising to distribute various forms of malware. Because advertising is the key revenue generator for numerous Internet companies, large ad networks, such as Google, Yahoo and Microsoft, invest a lot of effort to mitigate malicious ads from their ad networks. This drives adversaries to look for alternative methods to deploy malvertising. In this paper, we show that browser extensions that use ads as their monetization strategy often facilitate the deployment of malvertising. Moreover, while some extensions simply serve ads from ad networks that support malvertising, other extensions maliciously alter the content of visited webpages to force users into installing malware. To measure the extent of these behaviors we developed Expector, a system that automatically inspects and identifies browser extensions that inject ads, and then classifies these ads as malicious or benign based on their landing pages. Using Expector, we automatically inspected over 18,000 Chrome browser extensions. We found 292 extensions that inject ads, and detected 56 extensions that participate in malvertising using 16 different ad networks and with a total user base of 602,417.
HTTP Adaptive Streaming dominates most of the traffic on the Internet today and a large fraction is driven by video consumption on mobile phones and tablet devices. The client player implementations from different popular commercial services make different parameter and architectural choices, which result from heterogeneous APIs and device limitations. As a result, they exhibit quite different behaviors even for the same service, with very different traffic consumption characteristics. In this paper, we examine three major streaming services -- Netflix, Youtube, and Hulu, over the two dominant mobile platforms -- iOS and Android, in order to understand the impact of these design choices. We infer detailed session behavior based on passively collected packet traces over a large set of experiments across the providers and network types. We discover varying amounts of "redundant" traffic in the presence of bandwidth adaptation across the services, which negatively impacts network resources. We also find these design choices lead to unfairness in bandwidth consumption on shared networks across different platforms. In particular, we find the Android Netflix player is able to take a larger fraction of shared bandwidth when competing with the iOS implementation.
We present a new ad fraud mechanism that enables publishers to increase their ad revenue by deceiving the ad exchange and advertisers to target higher paying ads at users visiting the publisher's site. Our attack is based on polluting users' online interest profile by issuing requests to content not explicitly requested by the user, such that it influences the ad selection process. We address several challenges involved in setting up the attack for the two most commonly used ad targeting mechanisms -- re-marketing and behavioral targeting. We validate the attack for one of the largest ad exchanges and empirically measure the monetary gains of the publisher by emulating the attack using web traces of 619 real users. Our results show that the attack is effective in biasing ads towards the desired higher-paying advertisers; the polluter can influence up to 74% and 12% of the total ad impressions for re-marketing and behavioral pollution, respectively. The attack is robust to diverse browsing patterns and online interests of users. Finally, the attack is lucrative and on average the attack can increase revenue of fraudlent publishers by as much as 33%.
Data privacy audits that ensure policy compliant usage of personal data are increasingly enforced (internally or externally) on service providers that amass and process user data. Existing approaches to privacy auditing fall short of addressing the challenges introduced by modern personalized services that analyze user data using large scale machine learning and data mining (MLDM) algorithms, and provide users custom privacy controls. In this paper, we present GraphAudit, an auditing framework for large-scale graph mining platforms that can check the compliance of a wide range of expressive privacy policies. GraphAudit achieves this by reconstructing the runtime context of data use by MLDM algorithms by logging data accesses and tracing data flows. We implement GraphAudit over GraphLab, a popular distributed graph mining framework and evaluate its performance using commonly used MLDM algorithms and real worlds graph datasets. Our evaluation shows that GraphAudit performance overheads are moderate and exhibits scaling properties similar to GraphLab. Moreover, the overheads can be reduced by amortizing the I/O overhead of logging across distributed machines in the cluster. This results in overheads relative to GraphLab that are as low as 5.76x, with execution times that can automate privacy auditing compared to the otherwise manual and error prone approaches currently used.
Today's smartphone operating systems frequently fail to provide users with adequate control over and visibility into how third-party applications use their privacy-sensitive data. We address these shortcomings with TaintDroid, an efficient, systemwide dynamic taint tracking and analysis system capable of simultaneously tracking multiple sources of sensitive data. TaintDroid provides real-time analysis by leveraging Android's virtualized execution environment. Using TaintDroid to monitor the behavior of 30 popular third-party Android applications, we found 68 instances of misappropriation of users' location and device identification information across 20 applications. Monitoring sensitive data with TaintDroid provides informed use of third-party applications for phone users and valuable input for smartphone security service firms seeking to identify misbehaving applications.
To address the pressing need to provide transparency into the online targeted advertising ecosystem, we present AdReveal, a practical measurement and analysis framework, that provides a first look at the prevalence of different ad targeting mechanisms. We design and implement a browser based tool that provides detailed measurements of online display ads, and develop analysis techniques to characterize the contextual, behavioral and re-marketing based targeting mechanisms used by advertisers. Our analysis is based on a large dataset consisting of measurements from 103K webpages and 139K display ads. Our results show that advertisers frequently target users based on their online interests; almost half of the ad categories employ behavioral targeting. Ads related to Insurance, Real Estate and Travel and Tourism make extensive use of behavioral targeting. Furthermore, up to 65% of ad categories received by users are behaviorally targeted. Finally, our analysis of re-marketing shows that it is adopted by a wide range of websites and the most commonly targeted re-marketing based ads are from the Travel and Tourism and Shopping categories.
We explore the feasibility of identifying users from the unique patterns they exhibit when interacting with an individual electrical appliance in the home. We evaluate the effectiveness of a supervised learning based approach for user identification from a dataset of appliance usage collected across five users and three kitchen appliances over a period of eight weeks. Our results show that using appliance usage information alone provides a moderate average accuracy of 32% for group sizes of up to five users in the home. However augmenting usage information with hints about user presence can improve accuracy by 15-20%.
The multi-billion dollar online advertising ecosystem is driven by the ubiquitous tracking of end users’ online activities. Targeted ads are matched to individual user interests mined by online tracking services. Remarkably, an end user today – the presumed beneficiary – has little transparency into this process, cannot reason about how their online behavior is being used, and lacks influence in determining the types of ads delivered to them. In this paper, we address this pressing need for transparency into online ads by developing mechanisms and simple metrics that allow users to understand the extent to which they are targeted, and the level to which different ad categories are focused on selective parts of a their interest profiles. We implement these mechanisms as a client-side browser tool, AdReveal, and evaluate these metrics with web browsing data derived from the public search logs of 20 individuals. Our results show that AdReveal can effectively distinguish between different targeting mechanisms employed by advertisers and also provide fine-grained transparency into the ad targeting mechanisms. Finally, using the building blocks in AdReveal, we propose a practical and novel browser based control mechanism that lets users opt-out of specific ad categories and prevent these ads from being displayed to the user.
FWe present a workshop proposal focusing on future digital home infrastructures and systems needed to support ubiquitous computing applications and services. The workshop aims to be the premiere venue that brings together researchers and practitioners across the disciplines of Systems and Networking, Ubiquitous Computing and HCI to elaborate ways in which the current infrastructure in the digital home can be reshaped to meet the needs of users.
—Video search has become a very important tool, with the ever-growing size of multimedia collections. This work introduces our Video Semantic Indexing system. Our experiments show that Residual Vectors provide an efficient way of aggregating local descriptors, with complementary gain with respect to BoVW. Also, we show that systems using a limited number of descriptors and machine learning techniques can still be quite effective. Our first participation at the TRECVID evaluation has been very fruitful: our team was ranked 6 th in the light version of the Semantic Indexing task.
We are pleased to announce the release of a tool that records detailed measurements of the wireless channel along with received 802.11 packet traces. It runs on a commodity 802.11n NIC, and records Channel State Information (CSI) based on the 802.11 standard. Unlike Receive Signal Strength Indicator (RSSI) values, which merely capture the total power received at the listener, the CSI contains information about the channel between sender and receiver at the level of individual data subcarriers, for each pair of transmit and receive antennas. Our toolkit uses the Intel WiFi Link 5300 wireless NIC with 3 antennas. It works on up-to-date Linux operating systems: in our testbed we use Ubuntu 10.04 LTS with the 2.6.36 kernel. The measurement setup comprises our customized versions of Intel's close-source firmware and open-source iwlwifi wireless driver, userspace tools to enable these measurements, access point functionality for controlling both ends of the link, and Matlab (or Octave) scripts for data analysis. We are releasing the binary of the modified firmware, and the source code to all the other components.
Increasingly, mobile devices equipped with 802.11n interfaces are being used for a wide variety of applications including bandwidth-intensive HD video streaming. Recent work has shown that 802.11n interfaces are power-hungry, so energy management is an important challenge. 802.11n implementations have additional power states relative to earlier generations of 802.11 technology, so energy management challenges for 802.11n are qualitatively different compared to that faced by prior work. In this paper, we describe the design and implementation of Snooze, an energy management technique for 802.11n which uses two novel and inter-dependent mechanisms: client micro-sleeps and antenna configuration management . In Snooze, the APmonitors traffic on the WLAN and directs client sleep times and durations as well as antenna configurations, without significantly affecting throughput or delay. Snooze achieves 30~85% energy-savings over CAM across workloads ranging from VoIP and video streaming to file downloads and chats.
As more services have come to rely on sensor data such as audio and photos collected by mobile phone users, verifying the authenticity of this data has become critical for service correctness. At the same time, clients require the flexibility to tradeoff the fidelity of the data they contribute for resource efficiency or privacy. This paper describes YouProve, a partnership between a mobile device's trusted hardware and software that allows untrusted client applications to directly control the fidelity of data they upload and services to verify that the meaning of source data is preserved. The key to our approach is trusted analysis of derived data, which generates statements comparing the content of a derived data item to its source. Experiments with a prototype implementation for Android demonstrate that YouProve is feasible. Our photo analyzer is over 99% accurate at identifying regions changed only through meaning-preserving modifications such as cropping, compression, and scaling. Our audio analyzer is similarly accurate at detecting which sub-clips of a source audio clip are present in a derived version, even in the face of compression, normalization, splicing, and other modifications. Finally, performance and power costs are reasonable, with analyzers having little noticeable effect on interactive applications and CPU-intensive analysis completing asynchronously in under 70 seconds for 5-minute audio clips and under 30 seconds for 5-megapixel photos.
IEEE 802.11 wireless networks become increasing more complex and interesting inside homes. A number of home automation, home security, and entertainment products rely on wireless technologies for easy deployment without the need for wiring. Moreover, a number of such applications are fundamentally changing the traffic mix of a home wireless network, resulting in uplink traffic that is not only triggered by the users but that could potentially be nearly continuous in nature, such as wireless home security products, where each individual camera is likely to stream large amounts of data in high traffic areas. Given the diversity of traffic sources and their importance to the user, wireless home APs today can ship with Wireless Multimedia (WMM) support that prioritizes VoIP and video traffic for better user experience. In this paper, however, we note that the type and importance of applications to a home user may be much more diverse than 4 traffic classes could accommodate. In response, we survey the landscape of possible solution in particular when it comes to pacing traffic sources inside the network. We discuss the tradeoffs that such a design space exposes and test the performance of several solutions using ns3 simulations. Finally, we note that instead of a strict prioritization of traffic streams, a simple mechanism by which the user can pace traffic to provision more resource to the traffic of importance may be sufficient.
We are pleased to announce the release of a tool that records detailed measurements of the wireless channel along with received 802.11 packet traces. It runs on a commodity 802.11n NIC, and records Channel State Information (CSI) based on the 802.11 standard. Unlike Receive Signal Strength Indicator (RSSI) values, which merely capture the total power received at the listener, the CSI contains information about the channel between sender and receiver at the level of individual data subcarriers, for each pair of transmit and receive antennas. Our toolkit uses the Intel WiFi Link 5300 wireless NIC with 3 antennas. It works on up-to-date Linux operating systems: in our testbed we use Ubuntu 10.04 LTS with the 2.6.36 kernel. The measurement setup comprises our customized versions of Intel's close-source firmware and open-source iwlwifi wireless driver, userspace tools to enable these measurements, access point functionality for controlling both ends of the link, and Matlab (or Octave) scripts for data analysis. We are releasing the binary of the modified firmware, and the source code to all the other components.