Transactional Memory (TM) is one kind of approach to maximize parallel performance for multicore systems. There are conflicts When two or more parallel transactions access the same location and at least one access is a write. Contention management(CM) refers to the mechanisms used to guarantee forward—to avoid performance pathology, and to promote throughput. In this paper, we introduce a new CM police. We remitted six of seven performance pathologies summered by Bobba. Our result shows high performance for large transactions, while get moderate improvement or little slowdown for small transactions. The performance of the systems used this policies combined with other policy are steady.
Hardware Transactional Memory (HTM) is a promising Transactional Memory (TM) implementation because of its strong atomicity and high performance. Unfortunately, most contention management approaches in HTMs are dedicated to specific transaction conflict scenarios and it is hard to choose a universal strategy for different workloads. In addition, HTM performance degrades sharply when there are severe transaction conflicts. In this paper, we present a Global Contention Management Scheme (GCMS) to resolve severe transaction conflicts in HTMs. Our scheme depends on a Deadlock and Livelock Detection Mechanism (DLDM) and a Global Contention Manager (GCM) to resolve severe transaction conflicts. This scheme is orthogonal to the rest of the contention management policies. We have incorporated GCMS into different HTMs and compared the performance of the enhanced systems with that of the original HTMs with the STAMP benchmark suite. The results demonstrate that the performance of the enhanced HTMs is improved.
Image processing tasks in remote sensing and computer vision require an enormous amount of computation, especially in practical real-time applications. An array computer for low-and intermediate-level image processing is designed and implemented based on FPGA. To improve the system usability and the portability of application programs, a parallel image processing software environment is proposed based on the image algebra theory presented by G. X. Ritter. To program the parallel architecture, users need only describe their algorithms with image algebra operations provided in the environment. The environment can automatically select the optimal or approximately optimal parallel codes from the parallel implementation codes library according to the algorithm description and a cost model, and execute the application program. The parallel execution of application programs and the hardware details of the parallel architecture are transparent to users.
Providing service differentiation by using distributed control at the MAC layer is an effective mechanism for supporting Quality of Service (QoS) in wireless ad hoc networks. However, existing QoS-capable MAC protocols have been designed in single channel environments. Multichannel technology has been demonstrated to significantly increase network capacity, reduce chances of contention and collision of data transmission. So, QoS guarantee is easier to achieve by using multichannel technology. In this paper, we devise a supporting service differentiation multichannel MAC protocol. Two interfaces are used, an interface staying at common control channel is responsible for assigning data channel and another can dynamically switch channel for data transmission. In the common control channel, different priority flows adopt different backoff parameters to access channel to acquire data channel assignment chance. Simulation results show the protocol outperforms Enhanced Distributed Coordination Function (EDCF) of IEEE 802.11e.
With the development of multimedia technology, SIP (Session Initiation Protocol), as a simple, flexible and extensible protocol, has become the research focus of the NGN. In this case, the security issues of SIP become a very critical problem simultaneously. Through studying the security of SIP, this paper validates five attack ways in practical circumstances, including Registration Hijacking, INVITE attack, re-INVITE attack, Tearing Down Sessions, and DoS. Finally, through synthetic analysis and experiments, proposes four available measures to enhance the security of SIP: Improved Identity Authentication for HTTP Digest, Encryption with Hop-by-Hop, Forbidding the Lawless Third Part Register, Forbid the rewriting in 'From' Field and 'To' Field.
Network processor (NP) is optimized to performnetwork tasks. It uses massive parallel processing architecture to achieve high performance. Ad hoc network is an exciting research aspect due to the characters of self-organization, dynamically topology and temporary network life. However, all the characters make the security problem more serious. Denial-of-Service (DoS) attack is the main puzzle in the security of Ad hoc network. A novel NP-based security scheme is proposed to combat the attack. Security agent is established by a hardware thread in NP. Agent can update itself at some interval by the trustworthiness of the neighbor nodes. Agent can trace the RREQ and RREP messages stream to aggregate the key information and analyze them by intrusion detection algorithm. NS2 simulator is expanded to validate the security scheme. Simulation results show that NP-based security scheme is effective to detect DoS attack.
String matching is a very important component of many network applications. Persistent increase of network bandwidth needs high performance string matching algorithms. Traditional software algorithms cannot fulfill requirement of content filter in high-speed network. Design of high-speed string matching based on servos' array in FPGA is presented by dynamic adjusting servos to obtain powerful parallel process performance. Through simulations and implemented on FPGA, feasibility and rationality are validated. With improving performance of automata algorithm, structure of storage and filter level, it can be improved further.
This paper presents the design and implementation of a novel mechanism for secure DNS dynamic updates based on an overlay network, which inherits the scalability, self-organization, and fault tolerance from the overlay network. By employing a lightweight gossip-based multicast mechanism instead of traditional caching, the dynamic updates can be distributed quickly and maintain an excellent consistency, which prevent the stale records from poisoning the domain space. The experimental results show that the proposed mechanism provides enhanced resilience to single point of failures and DDoS attacks.
Exponential growth in the number of on-chip transistors with smaller size, make each generation of embedded microprocessors capable to supply more processing ability. In this paper a microarchitecture approach is proposed to make a simultaneous multithreading extension on ARM ISA processor. By exploiting both Instruction Level Parallelism and Thread Level Parallelism, the architecture can be expected to achieve better tradeoff between performance and hardware cost. The organization of the architecture is described. Detailed simulations of microarchitecture show IPC is improved with the multithreading extension of the simplesim-arm architecture greatly, especially for some benchmark combination.
Cooperation of multi-domain massively parallel processor systems in computing grid environment provides new opportunities for multisite job scheduling. At the same time, in the area of co-allocation, heterogeneity, network adaptability and scalability raise the challenge for the international design of multisite job scheduling models and algorithms. It presents multisite job scheduling schema through the introduction of multisite job scheduling model and the performance model under the grid environment. It introduces two job multisite and cooperative scheduling models and algorithms with the core of the optimal and greedy-heuristic resource selection strategies. Meanwhile, compared with single and multisite cooperative scheduling models and algorithms introduced by Sabin, Yahyapour and other persons, the validity and advance of the scheduling model and the performance model herein are proved.
The penalty associated with data cache misses is one of the obstacles to the performance of SMT trace processors. The increased latency is not only required to resolve the missing data, the miss will also have negative impact on the PE resources utilization rate of the SMT trace processors. When data cache miss occurs in SMT trace processors, all the completed traces following the data-miss-trace (a trace with at least one data cache miss) will be delayed to commit for the data cache miss event. PE resources occupied by those traces can not be released until traces are committed, which wastes the PE execution resources and hampers the performance of SMT trace processors.In this paper, we propose several data cache miss sensitive thread scheduling mechanisms with the aim to tolerate the penalties of data cache misses. By choosing the thread wisely in trace dispatch and trace commit stages, the SMT trace processors performance can be improved further. Simulation results show that when using L1-L2 thread scheduling mechanism, performance will be improved by 2.8% (2-thread SMT trace processors), 8.0% (4-thread SMT trace processors) and 18.2% (8-thread SMT trace processors) with 8-PE. 512 KB L2 cache configuration. (C) 2006 Elsevier B.V. All rights reserved.
With the development of multimedia technology, Session Initiation Protocol (SIP), which is simple, flexible and extensible, has become focus of research for the next generation network. Along with this, the security aspects of SIP have become a very critical problem simultaneously. This paper reveals five SIP security vulnerabilities through analysis and experimentation, which are Registration Hijacking, INVITE package attack, Re-INVITE attack, TDS (Tearing Down Sessions) and DoS(Deny of Service and Amplification).
This paper proposed an efficient hardware architecture to satisfy the computation requirement in the motion estimation of AVC/H. 264 standard. It supports the variable block size motion estimation, and has the high pipeline efficiency and high performance-price rate. Also, the PE and the add tree in the architecture are very flexible which allows this architecture has better trade-off between the area and the computation capability. Experimental result shows that this architecture has the powerful computation capability of coding 720 × 576 picture size at 48 fps with the search range of 65 × 65.
To improve the efficiency and scalability of conflict detection for multi-dimensional classifiers, a new algorithm, based on grid of trie (GoT) algorithm, was proposed. The new algorithm uses Patricia trie, constricts the length of Internet protocol (IP) prefix in order to use Hash technology, and improves the performance of the algorithm by adding ingress and egress of firewall for each filter.
The performance of trace cache processor rests with trace cache efficiency to a great extent. Higher trace cache miss rate will reduce performance significantly because only a low fetch-bandwidth can be maintained by conventional instruction cache. Unfortunately, with the ever increasing conventional application scale, higher trace cache miss rate is inevitable for the relative small capacity of trace cache, which will become the bottleneck of performance improvement. In this paper, we proposed trace cache hierarchy to remedy the limited capacity of 1-level trace cache. 2-level trace cache is incorporated in trace processors. But the simulation results show that only augmenting 2-level trace cache can not bring significant performance improvement for the long access latency. So we propose a path-based trace prefetch mechanism to reduce the latency of 2-level trace cache access further. Path-based trace prefetch mechanism is developed on top of next N trace prediction mechanism. By predicting the next N trace from current and prefetching it into trace prefetch buffer from 2-level trace cache, the access latency of 2-level trace cache can be reduced. The simulation results show that augmenting an 8K-Entry, eight-way 2-level trace cache, an 16-Entry trace prefetch buffer and prefetch distance set to 3, the average IPC improvement is 12.0% for eight SPECint95 benchmarks.
Distributed denial of service DDoS is a major threat to the availability of Internet services. As one of the most difficult problems in network security, it has received considerable attention from the mass media and the research community. In this paper, we design an effective and practical countermeasure which allows a general-purpose TCP-based public server to sustain high availability even during severe DDoS attacks. A novel microeconomic framework based on Generalized Vickrey auction GVA is proposed. By adopting this mechanism, not only the availability of services is improved, but also the total utility of legitimate clients can be maximized. Initial simulations have shown that this mechanism is highly effective in preferentially dropping attacker traffic over legitimate client traffic, and the protected server can remain operational under various system loads and severely attacked conditions. The results indicate that it is a promising approach to countering DDoS attacks.
This paper seeks to quantitatively understand the nature of the current threat towards the common name servers. A new tracking technique based on statistical model is proposed to locate the anomalous name servers by analyzing the real-world DNS traffic. After summarizing the attacks towards DNS, the detection method based on associative feature analysis is presented. Experiments are conducted which highlighting both the payload anomaly and the data flow anomaly, and the experimental results reveal the efficiency of our method in detecting the anomalous behaviors of name servers.
With the increase of Internet bandwidth and the development of Internet applications, gigabit exchange devices are used widely. The reasonable design of high-speed data buffer is a key to break throughput rate necklace. We provide a new design of multi-level buffer structure based on Field Programmable Gates Array (FPGA). Parallel Schedule algorithm increases packets transmission speed. By improving pipeline the structure can be applied for ten-gigabit-rate environments.
Efficient multisite job scheduling facilitates the cooperation of multi-domain massively parallel processor systems in a computing grid environment. However, co-allocation, heterogeneity, adaptability, and scalability emerge as tough challenges for the design of multisite job scheduling models and algorithms. This paper presents a new multisite job scheduling schema based on the multisite job scheduling model and the performance model for a heterogeneous grid environment. There are three key components: resource selection, reservation, and backfilling. The optimal and greedy-heuristic adaptive resource selection strategies are introduced. The conservative and easy backfilling are incorporated into the backfilling procedure. Experiments indicate that the scheduler and the algorithm are effective and perform better than a non-adaptive algorithm.