Machine learning–based network intrusion detection systems (NIDS) monitor network traffic to identify malicious activity. While NIDS perform well in controlled environments, their effectiveness deteriorates in real-world deployments due to concept drift. As the underlying distribution evolves and new attacks emerge, benign traffic nevertheless remains dominant, so rare and emerging attacks are sparsely observed. This severe class imbalance makes it difficult to identify which data should be prioritized for model adaptation and often leads to high labeling cost, as identifying the few informative attacks requires extensive manual inspection. We present NetGuard, a generative active adaptation framework that enables robust and label-efficient adaptation of NIDS under dynamic network conditions. NetGuard combines density-aware active sampling to identify informative drifted traffic for annotation with deep generative modeling to synthesize diverse, minority-class samples with high-fidelity to expand coverage of rare attacks. By jointly selecting what to label and how to augment scarce data, NetGuard mitigates performance degradation caused by both distribution shift and long-tailed attack distributions. We carry out a comprehensive evaluation of NetGuard across intrusion detection datasets including a large-scale real-world ISP dataset. On average, NetGuard improves overall detection performance by 2×, while achieving 8–15× higher F1 and 4–7× lower false-negative rates for rare and emerging attacks compared to state-of-the-art sample selection baselines, using only 1% labeled data. NetGuard enhances rare attack detection while reducing labeling costs, rendering it scalable and practical for intrusion detection.
AI agents are rapidly becoming more capable and widely deployed, promising substantial gains in productivity and enabling new classes of applications. However, their growing autonomy also introduces significant privacy and security risks. Existing defenses are predominantly agent-centric, relying on the agent itself to detect threats and enforce privacy and security policies. This approach is fundamentally limited because it entrusts policy enforcement to AI agents whose LLM-driven behavior is inherently nondeterministic and vulnerable to manipulation through attacks such as prompt injection. As a result, current defenses cannot reliably prevent privacy and security threats, highlighting a critical need for a new solution to securing AI agent systems. The networking community has long grappled with similar challenges and offers insightful principles we can borrow to design a more secure AI agent system. These include centralized control with distributed enforcement, capability-based access for mediating requests to sensitive resources, and least privilege through zero-trust enforcement. Historically, these principles have provided strong deterministic guarantees for networked systems. However, these principles alone are insufficient for AI agents because the safety and appropriateness of an agent's actions often depend on semantic context beyond the expressiveness of static rules. Building on these principles, we advocate for a systematic approach to AI agent security that combines deterministic enforcement mechanisms, which provide strong security guarantees, with semantic, context-aware policies that enable nuanced decision-making. We then present a reference architecture and identify key research questions and future directions to guide the design of secure and privacy-preserving AI agent systems.
Access to raw network traffic data is essential for many computer networking tasks, from traffic modeling to performance evaluation. Unfortunately, this data is scarce due to high collection costs and governance rules. and lack robust evaluations tied to real-world utility. We propose a new method based on state space models called NetSSM that generates raw network traffic at the packet-level granularity. Our approach captures interactions between multiple, interleaved flows -- an objective unexplored in prior work -- and effectively reasons about flow-state in sessions to capture traffic characteristics. NetSSM accomplishes this by training with a context window more than 8× longer, and produces traces up to 78× longer than existing transformer-based raw packet generators. compliance with standard protocol requirements and flow and session-level traffic characteristics.
LLM inference has become a global-scale, heterogeneous workload spanning agents, retrieval, tool-use, code execution and multi-modal reasoning. These workloads naturally enable context reuse from overlapping inputs, creating a major opportunity to store and reuse the contexts' KV Caches instead of recomputing them. However, model-side advances that shrink the KV Cache and system-side advances that reduce compute, storage, and transfer costs are evolve independently within legacy cloud boundaries. We argue that future inference infrastructure should allow decoupling of compute and KV Cache storage across cloud and datacenters. The network becomes an active distribution channel; bandwidth, latency and pricing directly determines how the KV Cache should be managed. We propose a vision for an Internet for the KV Cache, with KV Cache management working as a content-distribution system. In this view, KV Cache storage and recompute decisions are driven by model, infrastructure, and application metrics, to enable adaptive, content-driven decisions for minimizing latency and cost.
Large Language Models (LLMs) such as ChatGPT can infer personal attributes from seemingly innocuous text, raising privacy risks beyond memorized data leakage. While prior work has demonstrated these risks, little is known about how users estimate and respond. We conducted a survey with 240 U.S. participants who judged text snippets for inference risks, reported concern levels, and attempted rewrites to block inference. We compared their rewrites with those generated by ChatGPT and Rescriber, a state-of-the-art sanitization tool. Results show that participants struggled to anticipate inference, performing a little better than chance. User rewrites were effective in just 28% of cases - better than Rescriber but worse than ChatGPT. We examined our participants' rewriting strategies, and observed that while paraphrasing was the most common strategy it is also the least effective; instead abstraction and adding ambiguity were more successful. Our work highlights the importance of inference-aware design in LLM interactions.
Aligned language models often exhibit a recognizable AI-like style, yet its connection to post-training and internal representations remains poorly understood. In this work, we study whether post-training introduces or amplifies AI-like stylistic regularities and whether these regularities have a localized internal signature. To this end, we compare human text, base-model generations, and aligned-model generations under matched human-source prefixes. Aligned generations show lower human-corpus affinity and higher AI-detection rates than base generations, suggesting that post-training shifts generated text away from human-corpus style and toward detector-visible AI-like text. We then introduce PASTA (Post-training Alignment Signature Targeted Ablation), a training-free method that estimates a post-training alignment signature from aligned-base residual contrasts and ablates the corresponding direction during decoding. Across 11 aligned models and 6 AI detectors, PASTA lowers the detection rate for most aligned models; this effect transfers well across detectors and is not reproduced by random directions. Qualitative analysis suggests that PASTA generations remain relevant and coherent while exhibiting greater stylistic variation. Together, these results show that AI-like stylistic effects of post-training can be measured, localized, and causally tested through activation ablation.
The Low Latency, Low Loss, Scalable Throughput (L4S) architecture promises to reduce queuing delay while sustaining high throughput. Prior work has largely evaluated L4S in synthetic environments or controlled testbeds, leaving its real-world performance underexplored. In this study, we measure L4S performance specifically on Apple services delivered over Comcast residential networks. We deploy 83 Raspberry Pi devices across Comcast subscriber households and conduct over 120000 controlled experiments comparing L4S to traditional congestion control. Our results show that L4S reduces tail latency by up to 25
Latency anomalies, defined as persistent or transient increases in round-trip time (RTT), are common in residential Internet performance. When multiple users observe anomalies to the same destination, this may reflect shared infrastructure, routing behavior, or congestion. Inferring such shared behavior is challenging because anomaly magnitudes vary widely across devices, even within the same ISP and geographic area, and detailed network topology information is often unavailable. We study whether devices experiencing a shared latency anomaly observe similar changes in RTT magnitude using a topology-agnostic approach. Using four months of high-frequency RTT measurements from 99 residential probes in Chicago, we detect shared anomalies and analyze their consistency in amplitude and duration without relying on traceroutes or explicit path information. Building on prior change-point detection techniques, we find that many shared anomalies exhibit similar amplitude across users, particularly within the same ISP. Motivated by this observation, we design a sampling algorithm that reduces redundancy by selecting representative devices under user-defined constraints. Our approach captures 95 percent of aggregate anomaly impact using fewer than half of the deployed probes. Compared to two baselines, it identifies significantly more unique anomalies at comparable coverage levels. We further show that geographic diversity remains important when selecting probes within a single ISP, even at city scale. Overall, our results demonstrate that anomaly amplitude and duration provide effective topology-independent signals for scalable monitoring, troubleshooting, and cost-efficient sampling in residential Internet measurement.
We test a social media conversational agent for canvassing on the topic of anti-transgender prejudice, replicating and benchmarking treatment effects. In-person deep canvassing is the gold standard for durably changing attitudes on polarizing topics. However, door-to-door canvassing is costly, and many populations may not be feasibly reached in this manner. Campaigns are already conducting outreach using digital tools, including text messages and social media. If appropriately trained agents messaging over social media can achieve a fraction of the effect of in-person canvassing, canvassing may be scaled up to achieve large overall impacts at lower costs. Scripts used in this application are based on those used by transgender allies in the original study. To personalize messaging, the conversational agent uses natural language processing to detect conversational topics, and shares relevant pre-scripted messages of information and third-person experiences, encouraging respondents to engage in perspective-taking with respect to an outgroup. This study demonstrates the potential of automated social media messaging for deep canvassing, with possible applications by governments, public health agencies, and political organizations. Estimated effects are positive and significant under covariate adjustment and reweighting; due to important differential attrition, partial-identification bounds are also reported and include zero.
IoT environments such as smart homes are susceptible to privacy inference attacks, where attackers can analyze patterns of encrypted network traffic to infer the state of devices and even the activities of people. While most existing attacks exploit ML techniques for discovering such traffic patterns, they underperform on wireless traffic, especially Wi-Fi, due to its heavy noise and packet losses of wireless sniffing. In addition, these approaches commonly target at distinguishing chunked IoT event traffic samples, and they failed at effectively tracking multiple events simultaneously. In this work, we propose WiFinger, a fine-grained multi-IoT event fingerprinting approach against noisy traffic. WiFinger turns the traffic pattern classification task into a subsequence matching problem and introduces novel techniques to account for the high time complexity while maintaining high accuracy. Experiments demonstrate that our method outperforms existing approaches on Wi-Fi traffic, achieving an average recall of 85
Online platforms are seeing increasing amounts of AI-generated content—text and other forms of media that are made or co-created with generative AI. This trend suggests platforms may need to establish governance frameworks, including policies and enforcement strategies for how users create, post, share, and engage with such content to encourage responsible use. We investigate the governance of AI-generated content across 40 popular social media platforms. Just over two-thirds explicitly describe governance of AI-generated content spanning six themes. Most platforms focus on moderating AI-generated content that violates established content rules and discloses AI-generated content. Fewer platforms—those that are focused on creativity and knowledge-sharing—address other issues such as ownership and monetization. Based on these findings, we suggest stakeholders and policymakers develop more direct, comprehensive, and forward-looking AI-generated content governance, as well as tools and education for users about the use of such content.
Application-layer filters are crucial for network traffic analysis tasks such as telemetry, Quality of Experience (QoE) monitoring, and intrusion detection. Unlike general traffic classification, which assigns applications to flows without strict real-time requirements, filters must operate inline where balancing accuracy and performance is paramount. Traditional classification techniques based on pattern-matching (e.g., port numbers, IP addresses, or TLS SNI) offer low latency but face significant accuracy challenges as traffic becomes increasingly encrypted. While Machine Learning (ML) models achieve high classification accuracy on encrypted traffic, their computational overhead limits their applicability as inline filters. This paper introduces LoFi, a hybrid approach that combines the efficiency of pattern-matching with the accuracy of ML-based classification. By selectively applying ML models only when necessary, LoFi reduces packet loss from $\mathbf{3 8. 3 3 \%}$ (ML-only baseline) to $\mathbf{1. 1 7 \%}$ on CAIDA traces while maintaining low computational overhead.
Low Latency, Low Loss, and Scalable Throughput (L4S), as an emerging router-queue management technique, has seen steady deployment in the industry. An L4S-enabled router assigns each packet to the queue based on the packet header marking. Currently, L4S employs per-flow queue selection, i.e. all packets of a flow are marked the same way and thus use the same queues, even though each packet is marked separately. However, this may hurt tail latency and latency-sensitive applications because transient congestion and queue buildups may only affect a fraction of packets in a flow. We present SwiftQueue, a new L4S queue-selection strategy in which a sender uses a novel per-packet latency predictor to pinpoint which packets likely have latency spikes or drops. The insight is that many packet-level latency variations result from complex interactions among recent packets at shared router queues. Yet, these intricate packet-level latency patterns are hard to learn efficiently by traditional models. Instead, SwiftQueue uses a custom Transformer, which is well-studied for its expressiveness on sequential patterns, to predict the next packet's latency based on the latencies of recently received ACKs. Based on the predicted latency of each outgoing packet, SwiftQueue's sender dynamically marks the L4S packet header to assign packets to potentially different queues, even within the same flow. Using real network traces, we show that SwiftQueue is 45-65
Traditional tools for verifying network configurations: (1) require high manual effort or lack semantic depth, and (2) lack interpretable, human-readable explanations to support diagnosis. Large Language Models (LLMs) present a promising alternative, but simple approaches like full-file, partition-based, or chain-of-thought prompting fail to handle the large, interconnected nature of configuration files. We introduce the Context-Aware Prompting (CAP) framework, which enables an LLM to reason about network configurations like a human expert. CAP first identifies all potentially relevant neighboring, similar, and referenced configuration segments. CAP then engages the LLM in a structured dialogue, where the model requests the specific context it needs before performing a focused analysis. This on-demand approach provides the necessary context for deep reasoning while avoiding overload. Our evaluation on production network configurations demonstrates that CAP reliably detects a wide range of known and previously unknown configuration errors. CAP outperforms existing LLM-based baselines, and achieves performance comparable to state-of-the-art non-LLM tools.
To protect consumer privacy, the California Consumer Privacy Act (CCPA) requires businesses to provide consumers with a straightforward way to opt out of the sale and sharing of their personal information. However, the control that businesses enjoy over the opt-out process allows them to impose hurdles on consumers aiming to opt out, including by employing dark patterns. Motivated by the enactment of the California Privacy Rights Act (CPRA), which strengthens the CCPA and explicitly forbids certain dark patterns in the opt-out process, we investigate how dark patterns are used in opt-out processes and assess their compliance with CCPA regulations. Our research on 330 CCPA-subject websites reveals that these websites employ a variety of dark patterns. Some of these patterns are explicitly prohibited under the CCPA; others seem to take advantage of legal loopholes.
Despite extensive research measuring broadband quality, limited work has considered the interaction between residential Internet throughput and IP protocol. Public speed tests used to determine Internet quality do not, by and large, control for IP version; results, analysis, and recommendations are made using a hybrid of IPv4 and IPv6 data. Given the range of protocol designs, software stacks, and network infrastructure differences between the protocols, combined with the recent significant increase in IPv6 adoption, it is critical to understand what role the IP protocol plays in measuring Internet speeds. In this work, we systematically compare IPv4 and IPv6 speeds in residential access networks, examining differences in throughput experienced by households depending on IP version. Our findings demonstrate that IPv4 and IPv6 throughput differ in many instances and motivate a large-scale re-evaluation of our assumptions on IP version in future speed test analysis. Specifically, we find that IPv4 and IPv6 speeds differ in a significant number of cases, with up to 18.3% of our measurements differing by over 5%. Our findings indicate that substantial speed differences between IP versions can be driven by provider-specific factors such as speed tiers. Furthermore, we observe differences in IPv4 and IPv6 data depending on the speed-test software and testing infrastructure. Thus, this work guides future research about Internet speeds on how to consider and control for IP version.
This study, authored by the Broadband Internet Technical Advisory Group (BITAG), examines the development and standardization of machine-readable broadband Internet access service labels as required under the U.S. Federal Communications Commission’s (FCC) Empowering Broadband Consumers Through Transparency Order. The report analyzes the technical, regulatory, and market implications of implementing a uniform, machine-readable format—specifically recommending the use of a Comma Separated Value (CSV) schema—to enhance data interoperability, transparency, and consumer empowerment. The research identifies the benefits of a common data structure, including improved efficiency in broadband market analysis, enhanced regulatory compliance, and facilitation of third-party data aggregation for consumer-facing comparison tools. It further explores privacy and security considerations, geographic and tax variability challenges, and the technical mechanisms for unified discovery, such as the use of well-known URIs registered with IANA. The report emphasizes backward compatibility, extensibility of schema versions, and best practices for encryption, validation, and authenticity verification through Transport Layer Security (TLS) and digital signatures. By articulating a detailed technical schema and accompanying CSV data dictionary, this work contributes to the broader discourse on open data standards, network transparency, and the intersection of regulation and digital infrastructure. It provides a foundational framework for scholars, policymakers, and technologists analyzing the interoperability of broadband information systems and the implications of machine-readable regulatory disclosures for market accountability and consumer rights.
Intimate Partner Infiltration (IPI)–a type of Intimate Partner Violence (IPV) that typically requires physical access to a victim's device–is a pervasive concern in the United States, often manifesting through digital surveillance, control, and monitoring. Unlike conventional cyberattacks, IPI perpetrators leverage close proximity and personal knowledge to circumvent standard protections, underscoring the need for targeted interventions. While security clinics and other human-centered approaches effectively tailor solutions for survivors, their scalability remains constrained by resource limitations and the need for specialized counseling. In this paper, we present AID, an Automated IPI Detection system that continuously monitors for unauthorized access and suspicious behaviors on smartphones. AID employs a two-stage architecture to process multimodal signals stealthily and preserve user privacy. A brief calibration phase upon installation enables AID to adapt to each user's behavioral patterns, achieving high accuracy with minimal false alarms. Our 27-participant user study demonstrates that AID achieves highly accurate detection of non-owner access and fine-grained IPI-related activities, attaining an end-to-end top-3 F1 score of 0.981 with a false positive rate of 4 security clinics, scaling their ability to identify IPI tactics and deliver personalized, far-reaching support to survivors.
Accurate and efficient inference on network traffic through machine learning models is important for many management tasks, from traffic prioritization to anomaly detection. Existing ML inference pipelines differ primarily in their feature design: those based on summary flow statistics (e.g., packet sizes, inter-arrival times) are lightweight and efficient, though they may be less accurate for fine-grained classification, whereas pipelines that consume features directly from raw packet capture data can achieve higher accuracy but at significantly greater computational and resource cost. In this paper, we develop Just-in-Time Traffic Inference(JITI), a model serving system to support fast and accurate network traffic inference in raw packet-capture-based machine learning inference pipelines. Offline, JITI builds a curated pool of diverse trained models with varied feature and performance requirements. Online, JITI responds to traffic fluctuations via an adaptive scheduler that selects the model from the pool that offers the highest accuracy-to-efficiency ratio within system resource limits, thereby providing inference accuracy comparable to the more complex and resource-intensive packet-capture-based methods, with minimal efficiency compromise. Using traffic application inference as an example task, our evaluation shows that JITI improves inference performance by 18% over flow-statistics-based methods; when benchmarked against state-of-the-art packet-capture-based methods, JITI results in a worst-case drop in F1-Score of only 12.3%, while reducing the average inference decision time by ~127x.