Anonymous networks facilitate privacy-preserving information transmission by ensuring data confidentiality and communication anonymity. While encryption effectively secures content, anonymity remains vulnerable to traffic analysis, which can expose users' network identities and communication relationships. Consequently, anonymity alone is insufficient, necessitating integrated technologies to further strengthen communication security. This paper introduces covert communication to anonymous networks by proposing a novel Covert Information Transmission Framework over Anonymous Networks, designed to realize both covert and anonymous data transfer. To the best of our knowledge, this is the first work to propose covert communication schemes specifically within anonymous networks. In our framework, we establish covert channels in the Tor network and transmit covert information based on Tor's hidden service protocol. Utilizing Tor's anonymous circuits, we strategically select suitable protocol data fields as information carriers and propose three covert communication schemes with distinct characteristics. These schemes enable secure data exchanges between clients and hidden services, catering to diverse communication requirements. Finally, we implement the proposed framework in a real-world Tor network and conduct extensive experiments to evaluate its feasibility and performance.
We present DarkEE, a framework that extracts document-level event intelligence from Dark Web pages to support the Cyber Threat Intelligence (CTI) pipeline. It is evaluated on DarkEvents, a manually annotated dataset distilled from 15.4 million Tor pages. Containing 164 full-page documents of up to 20,000 words, it preserves the noise and length of real-world data. Our schema defines 11 dark-web-specific event categories (e.g., Hacking, Counterfeits) along with 16 semantic argument roles (e.g., Payment Method, Communication Channel) tailored to the domain. To handle these challenging inputs, DarkEE employs a clues-enhanced two-stage framework utilizing Large Language Models. It incorporates eventrelevant textual clues as in-context demonstrations to guide reasoning chains, enabling robust extraction from uncurated text. We have conducted experiments on the overall Document-level Event Extraction (DEE) task and its two subtasks corresponding to our two-stage pipeline: Documentlevel Event Detection (DED) and Document-level Event Argument Extraction (DEAE). The results demonstrate that our framework achieves significant improvements in Precision, Recall, and F1score across all three tasks compared to strong baselines. We contribute our DarkEvents to the community to promote its advancement through further research: https://anonymous.4open.science/ r/dw darkevents-7E6F.
Freenet is a well-known peer-to-peer network for anonymous file sharing. It preserves user anonymity by concealing the originating node, either an uploader or a downloader, among the relay nodes that form the multi-hop routing path. Prior work has examined protocol-level behavioral differences between originating and relay nodes, while leaving the underlying data transmission mechanism largely unexplored. In this paper, we identify a critical node role leakage during data transmission. Specifically, an uploader sends the entire block from its local storage in pieces, whereas a relay node forwards only the block pieces it has received. Although both uploaders and relay nodes encapsulate piece messages into fixed-size packets, a relay node may emit relay-specific under-filled packets when subsequent pieces are unavailable. Leveraging this role leakage, we develop a deanonymization attack that identifies uploaders by actively triggering relay-specific packets during the handling of block upload requests. Real-world experiments demonstrate that, by aggregating 10 block-upload observations, the attacker can reduce the false positive rate (FPR) of uploader identification to close to 0.
Freenet is a widely used anonymous communication system designed for file sharing. It preserves anonymity for both uploaders and downloaders via hop-by-hop routing and the enforcement of uniform protocols across nodes, preventing identification of the originating node along the routing path. Previous work has shown that the originating node can be deanonymized based on observable differences in interaction behaviors between nodes, while overlooking unobservable internal differences. In this paper, we identify a fundamental distinction between uploaders and relay nodes in their patterns of inserting application-layer messages into in-memory message queues. Although this difference is unobservable to a malicious node, we show that the FIFO (First-In, First-Out) property of the message queue allows the internal message insertion state to be mirrored in transmitted messages by triggering a beacon message. This insight enables a novel deanonymization attack against uploaders. We further address two challenges in conducting the attack: preventing the internal difference from being undermined and ensuring that it is effectively mirrored in transmitted messages. Real-world experiments demonstrate the feasibility and effectiveness of our attack, achieving a nearly 100% true positive rate with a maximum 4.17% false positive rate. Our work demonstrates that even unobservable internal differences can be potential threats to Freenet.
Tor is a widely used network for anonymous communication, designed to safeguard the privacy and anonymity of both users and service providers through multi-layered encryption and multi-hop routing. Despite its robust architecture, Tor remains susceptible to various attacks, including denial-of-service (DoS) and deanonymization. However, these attacks are constrained by high resource requirements, scalability limitations, and the defenses implemented within the Tor network. Furthermore, they are particularly ineffective in identifying and circumventing the Guard nodes that protect onion services. In this paper, we uncover a novel vulnerability in Tor’s circuit construction process and bandwidth scheduling, termed the Circuit Circle vulnerability. Exploiting this flaw, attackers can create circular circuits, leading to bandwidth contention and overloading a Tor node. To demonstrate the severity of this vulnerability, we propose Duplicate-Node Attack, a three-phase strategy that identifies Guard nodes, exhausts their bandwidth, and performs coarse geolocation inference of onion services. Unlike conventional methods, our Guard node identification technique bypasses existing Tor Vanguard defenses without requiring control over any relays. Our extensive experimental results confirm that Duplicate-Node Attack can reliably identify Guard nodes, exhaust their bandwidth with minimal cost, and infer the geolocation of onion services, with an average latency bias of 35.4 ms.
The Tor network, while offering anonymity through traffic routing across volunteer-operated nodes, remains vulnerable to attacks that aim to deanonymize users by correlating traffic patterns between colluded Entry and Exit nodes in circuits. This paper presents a novel approach for detecting anomalous circuits in the Tor network, and for the first time provides a more comprehensive identification of potential malicious accomplice nodes in Tor by taking roles of nodes in anomalous circuits into consideration. Our method strategically utilizes modified Middle nodes to capture traffic data, followed by a novel circuit classification based on traffic patterns to pinpoint concerned circuits. Two kinds of anomalies are identified: routing anomalies and usage anomalies, that respectively represent the anomalies with explicit or implicit violation of Tor's circuit construction guidelines. This leads to a successful revealing of totally 1,960 anomalous nodes in Tor. Furthermore, we apply clustering analysis with considering corresponding anomalous circuits and other key characteristics to the detected anomalous nodes, revealing potential hidden organizations behind these nodes that can threaten the network's security. Our findings highlight the necessity for the Tor project to adopt targeted mitigation strategies to enhance overall network security and privacy.
Unmanned intelligent vehicles are exposed to high risks of network attack,hardware attack,operating system attack and software attack.They are susceptible to physical or remote security attacks,causing it to deviate from the de-livery trajectory and fail the delivery task,or even be manipulated to disrupt normal operation of the factory.To address this problem,a dual-verified secure localization method for unmanned intelligent vehicles was proposed.The existing Wi-Fi network infrastructure was utilized by the vehicles for fingerprinting localization and a feature fusion strategy was designed to realize the dynamic fusion of Wi-Fi and magnetic field fingerprints.Multiple environmental monitoring points were deployed to collect the sound signals made by vehicles to calculate the position based on time difference of arrival and spatial segmentation method.Then the location reported by the vehicle was compared with the result of moni-toring points for verification.Once an abnormal position was detected,an alert would be issued,ensuring the normal op-eration of the unmanned intelligent vehicles.The experimental results in the real indoor scenarios show that the proposed method can effectively track the positions of the target unmanned intelligent vehicle,and the positioning accuracy is bet-ter than existing benchmark algorithms.
Freenet is a well-known anonymous communication system that enables file sharing among users. It employs a probabilistic hops-to-live (HTL) decrement approach to hide the originator among nodes in a multi-hop path. Therefore, all nodes shall exhibit identical behaviors to preserve anonymity. However, we discover that the path folding mechanism in Freenet violates this principle due to behavior discrepancy between downloaders and intermediate nodes. The path folding mechanism is designed to optimize the network topology of Freenet. A delayed path folding message by a successor node may incur a timeout event at its predecessor, and an intermediate node reacts differently to such timeout with a downloader. Therefore, malicious nodes can deliberately trigger the timeout event to identify downloaders. The complex implementation of the path folding timeout detection mechanism in Freenet complicates our de-anonymization attack. We thoroughly analyze the underlying cause and develop three strategies to manipulate three types of messages respectively at the malicious node, minimizing the false positive rate. We conduct extensive real-world experiments to verify the feasibility and effectiveness of our attack. They show that our attack achieves a true positive rate of 100% and false positive rate of near 0% under two different Freenet download modes.
The onion service is the most important mechanism of the Tor network which enables service providers to publish anonymously various TCP services, such as web services. To access the target onion services, clients first know the 56-byte onion addresses. However, randomly generated onion addresses are difficult to memorize and can be easily used by attackers to generate phishing sites with similar onion addresses. In this paper, we propose a correlated onion address generation approach which is capable of generating a unique onion address via a customized string and a root onion address. This approach enables the generated onion addresses to be computed by clients using a human-memorable string, resulting in easier access to onion services. Based on this approach, we design and implement a Tor Domain Name System (TorDNS) that allows different service providers to register anonymously and clients to access anonymous services quickly through human-memorable pseudo-onion addresses. TorDNS is compatible with existing onion service mechanism and does not introduce additional privacy and security issues. In addition, similarity detection of pseudo-onion addresses can effectively reduce the risk of phishing sites on the Tor network.
IFTTT is one of the most popular Trigger-Action Programming platforms. The rules generated in IFTTT are named IoT Applets. Despite the powerful programming interface provided by IFTTT, establishing an Applet requires technical skills and is not convenient enough for most users. To address this problem, we propose a gesture based programming method to help end users establish and manage IoT Applets in a convenient way. It requires employment of an RGB-D camera, and recognizes users’ pointing rays and hand actions. The obtained information is interpreted to certain devices and device events for Applet management. An experiment involving 20 participants validates the performance of our proposed method.
The rapid development of Internet of Things technology has allowed a massive number of devices to be connected, resulting in a lot of private data being transmitted over the network. To protect the security and privacy, anonymous communication technologies such as Tor are widely deployed. They use complex mechanisms such as encryption and multi-hop forwarding to hide the communication relationship and content. However, much existing work on the website fingerprinting attack has proven that there is still a risk of privacy disclosure. Website fingerprinting attacks can extract side channel information from encrypted traffic to form a fingerprint that identifies the victim’s destination website. But most work is conducted in the ideal environments, assuming the absence of background traffic, which raises doubts about the effectiveness in the real world. In this paper, we relax the strong assumptions of the attack model and propose a practical multi-tab website fingerprinting attack. We first constructed a CNN model to distinguish whether the unknown traffic is generated by accessing a single web page or multiple web pages. Then we design a split point recognition method based on the BalanceCascade algorithm to separate the overlapping traffic. Furthermore, we build recognition models based on ResNet and multi-head self-attention mechanism to identify the tail-missing and head-overlapping traffic sequences respectively. Extensive experiments are carried out in the real-world scenario to evaluate the proposed method. When the time interval is randomly selected within 5–15s, we achieve an accuracy of 77.34% in split point recognition. And the accuracy of identifying the two pages is 88.19% and 63.02% respectively.
Previous research has shown that Tor traffic can be easily identified, making Tor connections frequently blocked. In order to access the Tor network successfully, some censorship circumvention tools such as Shadowsocks and OpenVPN are utilized as front-proxy to connect to Tor entry nodes. However, the distinguishability of Tor traffic over these censorship circumvention tools has not yet been fully evaluated. By analyzing the equal-size segmentation mechanism of Tor and the transmission mechanisms of circumvention tools, we find that the payload length distribution of Tor traffic encrypted and encapsulated through these tools displays a distinct pattern, which makes such Tor traffic retain distinguishable from regular encrypted traffic. To verify this finding, we develop an automated, large-scale Tor traffic collection system to capture Tor traffic forwarded by various circumvention tools, and then design corresponding algorithms to extract traffic features in terms of payload length distribution. Finally, we perform the evaluation on the distin-guishability between the captured Tor traffic and the normal non-Tor traffic through extracted features. The F1-Score can achieve 0.99 with the false positive rate close to 0 when using Support Vector Machines for training and classification. The experimental results prove that circumvention tools cannot mask the inherent features of Tor traffic, and thus the Tor traffic forwarded by these tools can still be clearly distinguished from normal non-Tor traffic.
The development of home internet of things (H-IoT) devices brings convenience but poses significant privacy and security risks. Existing research minimizes data uploaded to the cloud but fails to process data locally, resulting in a trade-off between privacy and functionality. In this paper, we propose a privacy-preserving method that identifies and processes sensitive data sent from H-IoT devices at the edge side, ensuring functionality while preserving privacy. Our method applies different identification strategies to packets with different features, making it applicable to most H-IoT devices and scenarios. We validate our approach through experiments on a prototype system that monitors multiple cameras, demonstrating its effectiveness in preserving privacy while maintaining functionality.
The Internet of Vehicles (IoV) has witnessed a substantial surge in its expansion during recent years, with the integration of edge servers emerging as an increasingly prevalent practice to cater to content requests that demand a superior quality of service. Nonetheless, conventional cache policies often prove inadequate for IoV applications, primarily due to their incapacity to effectively handle a diverse array of content requests, the high-speed mobility of vehicles, and the inherent instability of network connections. Within the scope of this study, we commence by bifurcating various content requests into two distinct categories: latency-sensitive content request and bandwidth-sensitive content request. Subsequently, we establish a model to evaluate the service quality under the joint consideration of different types of content requests, the mobility of vehicles and storage capacity of edge nodes. Furthermore, we transform the quality-of-service (QoS) penalty function into a system reward function. This adaption enables us to propose an innovative edge cache scheme founded upon the Deep Deterministic Policy Gradient (DDPG) algorithm of Reinforcement Learning (RL) method, which empowers dynamic adjustments in response to the evolving IoV environment. To validate the effectiveness of our proposed approach, we harness the Simulation of Urban Mobility (SUMO) traffic simulation software and construct a traffic road scenario based on a specific part segment of Nanjing Beltway. A wide-ranging set of contrast experiments was performed to ensure the improved performance of the DDPG based deep reinforcement learning method. Simulation experimental results show that the proposed algorithm converges quickly and outperforms existing algorithms in terms of service quality-hit ratio for latency-sensitive content and transmission speed for bandwidth-sensitive content.
One of the features of network traffic in Internet of Things (IoT) environments is that various IoT devices periodically communicate with their vendor services by sending and receiving packets with unique characteristics through private protocols. This paper investigates semantic attacks in IoT environments. An IoT semantic attack is active, covert, and more dangerous in comparison with traditional semantic attacks. A compromised IoT server actively establishes and maintains a communication channel with its device, and covertly injects fingerprints into the communicated packets. Most importantly, this server not only de-anonymizes other IPs, but also infers the machine states of other devices (IPs). Traditional traffic anonymization techniques, e.g., Crypto-PAn and Multi-View, either cannot ensure data utility or is vulnerable to semantic attacks. To address this problem, this paper proposes a prefix- and distribution-preserving traffic anonymization method named PD-PAn, which generates multiple anonymized views of the original traffic log to defend against semantic attacks. The prefix relationship is preserved in the real view to ensure data utility, while the IP distribution characteristic is preserved in all the views to ensure privacy. Intensive experiments verify the vulnerability of the state-of-the-art techniques and effectiveness of PD-PAn.
The traditional live video transmission optimization mechanism is deployed on the server side,which cannot quickly respond to the dynamic changes of the user's wireless network environment.To address this problem,a live video transmission optimization mechanism based on edge intelligence called S-Edge was proposed.It was deployed on the OpenWrt-based wireless access point,and comprehensively utilized the wireless channel state information such as airtime utilization and signal-to-noise ratio to make intelligent decisions on terminal priority and transmission rate based on fuzzy logic theory.Furthermore,the active queue management with hierarchical token bucket and service demand-driven wire-less transmission rate adaptive control technologies were introduced to realize the real-time scheduling of live video data.In order to verify the effectiveness and performance of the proposed mechanism,a high client-density environment was built through user clusters based on multi-radio interfaces in the real-world scenario.Experimental results show that S-Edge can significantly reduce the average delay and packet loss rate,which meets QoS requirements of live video transmission services in the high client-density environment.
Website Fingerprinting (WF) attacks can extract side channel information from encrypted traffic to form a fingerprint that identifies the victim's destination website, even if traffic is sophisticatedly anonymized by Tor. Many offline defenses have been proposed and claimed to have achieved good effectiveness. However, such work is more of a theoretical optimization study than a technology that can be applied to real-time traffic in the practical scenario. Because defenders generate optimized defense schemes only if the complete traffic traces are obtained. The practicality and effectiveness are doubtful. In this paper, we provide an in-depth analysis of the difficulties faced in porting existing offline defenses to the online scenarios. And then the online WF defense based on the non-targeted adversarial patch is proposed. To reduce the overhead, we use the Gradient-weighted Class Activation Mapping (Grad-CAM) algorithm to identify critical segments that have high contribution to the classification. In addition, we optimize the adversarial patch generation process by splitting patches and limiting the values, so that the pre-trained patches can be injected and discarded in real-time traffic. Extensive experiments are carried out to evaluate the effectiveness of our defense. When bandwidth overhead is set to 20%, the accuracies of the two state-of-the-art attacks, DF and Var-CNN, drop to 10.83% and 15.49%, respectively. Furthermore, we implement the real-time patch traffic injection based on WFPadTools framework in the online scenario, and achieve a defense accuracy of 95.50% with 12.57% time overhead.
With the extensive implementation of the strong public health interventions in China, many models proposed to predict COVID-19 epidemic are no longer applicable to the current epidemic development. In this paper, a COVID-19 prediction method is proposed based on a staging SEITR model with consideration of strong public health interventions in China. The method simulates preventive and control measures such as mass nucleic acid testing and quarantine of close contacts by introducing the role of Isolates and the transformation of Exposed to Isolated. The experimental evaluation uses real epidemic data from six cities including Nanjing, Yangzhou, and etc. The accuracy of prediction for total number of infections reaches 95.8% with the data of the first 15 days of the outbreak. In addition, the prediction accuracy of the end of the pandemic is 95.07%. These show that the proposed method can effectively predict the course of the epidemic and it is practical for relevant departments to formulate reasonable prevention and control measures.
Website fingerprinting (WF) attacks allow an attacker to eavesdrop on the encrypted network traffic between a victim and an anonymous communication system so as to infer the real destination websites visited by a victim. Recently, the deep learning (DL) based WF attacks are proposed to extract high level features by DL algorithms to achieve better performance than that of the traditional WF attacks and defeat the existing defense techniques. To mitigate this issue, we propose a-genetic-programming-based variant cover traffic search technique to generate defense strategies for effectively injecting dummy Tor cells into the raw Tor traffic. We randomly perform mutation operations on labeled original traffic traces by injecting dummy Tor cells into the traces to derive variant cover traffic. A high level feature distance based fitness function is designed to improve the mutation rate to discover successful variant traffic traces that can fool the DL-based WF classifiers. Then the dummy Tor cell injection patterns in the successful variant traces are extracted as defense strategies that can be applied to the Tor traffic. Extensive experiments demonstrate that we can introduce 8.1% of bandwidth overhead to significantly decrease the accuracy rate below 0.4% in the realistic open-world setting.
Cybercrime is significantly growing as the development of internet technology. To mitigate this issue, the law enforcement adopts network surveillance technology to track a suspect and derive the online profile. However, the traditional network surveillance using the single-device tracking method can only acquire part of a suspect's online activities. With the emergence of different types of devices (e.g., personal computers, mobile phones, and smart wearable devices) in the mobile edge computing (MEC) environment, one suspect can employ multiple devices to launch a cybercrime. In this paper, we investigate a novel cross-device tracking approach which is able to correlate one suspect's different devices so as to help the law enforcement monitor a suspect's online activities more comprehensively. Our approach is based on the network traffic analysis of instant messaging (IM) applications, which are typical commercial service providers (CSPs) in the MEC environment. We notice a new habit of using IM applications, that is, one individual logs in the same account on multiple devices. This habit brings about devices' receiving sync messages, which can be utilized to correlate devices. We choose five popular apps (i.e., WhatsApp, Facebook Messenger, WeChat, QQ, and Skype) to prove our approach's effectiveness. The experimental results show that our approach can identify IM messages with highF1-scores (e.g., QQ's PC message is 0.966, and QQ's phone message is 0.924) and achieve an average correlating accuracy of 89.58% of five apps in an 8-people experiment, with the fastest correlation speed achieved in 100 s.