String matching algorithms are critical to several scientific fields. Beside text processing and databases, emerging applications such as DNA protein sequence analysis, data mining, information security software, antivirus, ma- chine learning, all exploit string matching algorithms [3]. All these applica- tions usually process large quantity of textual data, require high performance and/or predictable execution times. Among all the string matching algorithms, one of the most studied, especially for text processing and security applica- tions, is the Aho-Corasick algorithm. 1 2 Book title goes here Aho-Corasick is an exact, multi-pattern string matching algorithm which performs the search in a time linearly proportional to the length of the input text independently from pattern set size. However, depending on the imple- mentation, when the number of patterns increase, the memory occupation may raise drastically. In turn, this can lead to significant variability in the performance, due to the memory access times and the caching effects. This is a significant concern for many mission critical applications and modern high performance architectures. For example, security applications such as Network Intrusion Detection Systems (NIDS), must be able to scan network traffic against very large dictionaries in real time. Modern Ethernet links reach up to 10 Gbps,more » and malicious threats are already well over 1 million, and expo- nentially growing [28]. When performing the search, a NIDS should not slow down the network, or let network packets pass unchecked. Nevertheless, on the current state-of-the-art cache based processors, there may be a large per- formance variability when dealing with big dictionaries and inputs that have different frequencies of matching patterns. In particular, when few patterns are matched and they are all in the cache, the procedure is fast. Instead, when they are not in the cache, often because many patterns are matched and the caches are continuously thrashed, they should be retrieved from the system memory and the procedure is slowed down by the increased latency. Efficient implementations of string matching algorithms have been the fo- cus of several works, targeting Field Programmable Gate Arrays [4, 25, 15, 5], highly multi-threaded solutions like the Cray XMT [34], multicore proces- sors [19] or heterogeneous processors like the Cell Broadband Engine [35, 22]. Recently, several researchers have also started to investigate the use Graphic Processing Units (GPUs) for string matching algorithms in security applica- tions [20, 10, 32, 33]. Most of these approaches mainly focus on reaching high peak performance, or try to optimize the memory occupation, rather than looking at performance stability. However, hardware solutions supports only small dictionary sizes due to lack of memory and are difficult to customize, while platforms such as the Cell/B.E. are very complex to program.« less
Irregular applications present unpredictable memory-access patterns, data-dependent control flow, and fine-grained data transfers. Only a holistic view spanning all layers of the hardware and software stack can provide effective solutions to address these challenges.
With computing systems becoming ubiquitous, numerous data sets of extremely large size are becoming available foranalysis. Often the data collected have complex, graph based structures, which makes them difficult to process with traditional tools. Moreover, the irregularities in the data sets, and in the analysis algorithms, hamper the scaling of performance in large distributedhigh-performance systems, optimized for locality exploitation and regular data structures. In this paper we present an approach tosystem design that enable efficient execution of applications with irregular memory patterns on a distributed, many-core architecture, based on off-the-shelf cores. We introduce a set of hardware and software components, which provide a distributed global address space, fine-grained synchronization and latency hiding of remote accesses with multithreading. An FPGA prototype has been implemented to explore the design with a set of typical irregular kernels. We finally present an analytical model that highlights the benefits of the approach and helps identifying the bottlenecks in the prototype. The experimental evaluation on graph basedapplications demonstrates the scalability of the architecture for different configurations of the whole system.
This paper concerns the use of two double layer asphalts employed in Florence to reduce traffic noise in two different urban streets. In this case, tyre/road interaction is negligible because of the low speed of vehicles. For this reason, it was expected that porous pavements couldn’t give a great contribution in the reduction of traffic noise. Measurements have been carried out during two years after the laying down of the new asphalts, taking sound pressure level at the border of the road, together with fluxes of vehicles. A statistical analysis has been implemented on the base of a model that permits to evaluate the actual contribution of the asphalts. Results point out a quite significant reduction of sound levels and long-term durability, even in urban streets with low vehicle speed. 1 INTRODUCTION Porous asphalts have been used mainly to improve the security of the streets, by means of their drainage properties. Moreover, the benefit achievable with such products in terms of acoustic gain is well known, but until now tests and controls on their acoustical properties have been carried out for high velocity of the vehicles, as highways or extra-urban roads. Such studies have underlined that the excitement of tire vibrations by the road roughness and aerodynamic noise, caused by air-pumping, dominate the overall noise. Otherwise, it seems difficult to extend these results to urban area; in fact, the mechanism of noise generation may be different in this case, as here the speed is generally lower than 50 km/h and queues of vehicles are possible. The first experiments on the efficacy of porous asphalts laid down in urban area, have been carried out in France, starting from the second half of the eighties. Acoustical gains comparable with those obtained in highways or extra-urban roads, have been measured [1]. These results have been confirmed by more recent experiments carried out in Italy, in an urban street of Modena [2]. Nevertheless, such previous studies, involving single layer porous asphalts, have shown that the acoustical performances decrease in the course of the first year of lifetime. Such short-term durability depends on the clogging of porosity, due to many factors, as dust, oil, materials from tyre deterioration, etc... Acoustical duration of surfaces is by far the most important problem to optimize costs and benefits, as well as to find real alternative solutions to traditional asphalts. Short-term durability in terms of drainage and acoustical properties, has led to develop a new generation of pavements, made by two porous layers with different granulometry, that are superimposed. The fine grain of the higher layer works as a filter for dust and deposits which are removed by means of the cleaning effect due to vehicles passages. On the other hand, the coarse grain of the lower layer is projected to drain away rainwater. Copyright SFA InterNoise 2000 2 Two porous double layer asphalts were laid in Florence, within an experimental project to reduce traffic noise in dense urban areas. Measurements have been aimed to analyse several aspects: 1) the reduction of sound pressure levels in relation to the amount of traffic in the streets, which is the object of the present work; 2) the changing in the propagation of sound [3]; 3) the absorption coefficient of the asphalts [4, 5]. 2 THE STREETS AND THE ASPHALTS UNDER TEST Both streets present typical features of a city with ancient plant. One of these (A) is located in the centre of Florence. It is approximately 12 m wide and it has buildings on both sides, 18-20 m high. The other street (B) is in a small centre nearby Florence. It is very narrow (6 m in width) and it has low buildings only on one side (Fig. 1). Figure 1: Street A, on the left; street B, in the two photos, on the right. The two streets present markedly urban features, as they are characterised by low speeds of vehicles and intense traffic, with different composition of vehicles. One street (A) presents a high percentage of scooters (60%) and an appreciable portion of buses (10%). In the other case (B) cars are predominant in the total traffic flow, while scooters and heavy vehicles are almost negligible. Here speeds are quite low (about 30 km/h) as the street is very narrow and queues are possible. Both pavements are composed of two draining layers with different granulometry, as shown in Fig. 2. Figure 2: Thickness and granulometry of the two asphalts (A, on the left, and B on the right). 3 RESULTS Noise level in urban streets typically shows a large variability, induced by changes in traffic flow. Such fluctuation has almost the same magnitude of the quantities to be evaluated. It means that fluctuations on measurements of a few dBA, do not allow to estimate the variation due to the asphalts, as it is valuable of the same magnitude. To appreciate the contribution of the pavements, it is necessary to evaluate noise Copyright SFA InterNoise 2000 3 levels in standard conditions of traffic, that it means fixed volume and proportion of vehicles. Therefore, fluxes for each type of vehicle have been measured and correlated with corresponding sound levels, by means of a mathematical model. A non-linear law has been used, whose unknown parameters have been adjusted to fit measured data. The shape of the chosen function derives from simple physical considerations: LAeq = A + 10 · log10 (car + B · scooter + C · heavy) (1) Car, scooter and heavy represent fluxes of the corresponding type of vehicles. Heavy is rather a generic class, but, in the two studied streets, it is mainly composed of buses. The other terms, A, B and C are the regression variables. A set of values for them has been computed in each measurement campaign. Fixing the same conditions of noise emission (i.e., the logarithmic argument in Eq. 1) it has been possible to compare homogeneous data, even if relative to different measurement campaigns. Fig. 3 shows sound levels obtained using this procedure, at each measurement campaign (months are counted from the date in which the asphalts were laid down; such point is marked with 0 in the x-axis). Figure 3: Calculated sound pressure levels vs. months, by means of the correlation procedure. The error bar that appears in Fig. 3, has been estimated comparing the standard error of the regression with the uncertainty coming from the phonometric measurements of the sound levels. In the latter case, it has been assumed an indicative error of 0.5 dBA, as the same technical devices conforming to the specifications of class 1, have been used all the times. Moreover, the calibration of the instruments has been checked sometimes, by means of a pistonphone of class 0. As the statistic error has resulted less or equal to the uncertainty of the measurements, the error of 0.5 dBA has been attributed also to the noise level computed by the fitted function. Initially, both asphalts present an acoustic gain of about 3.5 dBA. In the case of A, such value is substantially unchanged after two years. On the other hand, the asphalt B shows a worse trend, during a shorter lifetime. Different performances of the two pavements may be due not only to the granulometry, but also to many factors, such as volume and composition of traffic, as well as other environmental factors. 4 CONCLUSION Noise reduction due to a double layer asphalt amounts to approximately 3.5 dBA. It is generally emphasized the role of porous surfaces in relation to the so called ”rolling noise” which is predominant in high velocity context. Nevertheless, our results show that double layer asphalts can reduce traffic noise also when employed in urban streets. Here, speeds of vehicles are generally not so high and the tyre/road interaction is not always the dominant component of the global noise emitted by vehicles. In one of the streets under test, the large portion of scooters in the total traffic flow, enforces the idea that rolling noise is negligible. As a consequence of these evaluations a different action mechanism must be hypothesised. Copyright SFA InterNoise 2000 4 Until now long-term durability of porous surfaces has been considered crucial, and not yet proved to be sufficient. Our results point out that double layer porous asphalts maintain acoustic properties for a longer time than a single layer in the same operative conditions. Acoustic duration seems to be comparable with their lifetime.
We propose an intermediate approach between full custom hardware systems and full-software tools. Figure 1 shows the overview of the proposed architecture. We start from an off-the-shelf architecture composed of simple, in-order cores and an on-chip interconnection. The onchip interconnection interfaces the processing core with the memory controller for the external memory (DDR3) and the shared I/O peripherals. We add three custom components: the Global Memory Access Scheduler (GMAS), the Global Network Interface (GNI) and the Global SYNChronization module (GSYNC). The GMAS enables support for the scrambled address space. It also implements part of the support latency tolerance, storing remote memory operations, and acts as a scheduler for lightweight software multithreading.
Due to current and future technology issues, multi-core processing systems are required to provide support for adaptivity to an ever increasing extent. This requirement may descend from demands of fault-tolerance as well as from dynamic Quality-of-Service (QoS) management strategies, depending on the targeted application and power budget. This paper presents a Network-on-Chip (NoC)-based Multi-Processor System-on-Chip (MPSoC) platform for video decoding applications that provides system adaptivity and reduced power consumption. The platform specifically targets execution of Polyhedral Process Network (PPN) streaming applications. System adaptivity is achieved through support for runtime migration of PPN processes between different tiles, while the power consumption is reduced at runtime through clock gating of inactive processing tiles. The details of how the migration process and clock gating mechanisms are implemented in the platform, both in hardware and middleware, will be presented, along with a characterization of the introduced overhead. In its standard operating mode, the adaptive platform executes a PPN implementation of an H.264 decoder on a stream of video packets coming from a network connection. The network packets are analyzed through a deep packet inspection kernel, OpenDPI, to distinguish between video and special reconfiguration packets. Upon reception of a reconfiguration packet from the network, the adaptive platform performs an on-line reconfiguration that employs runtime PPN process migration to modify the amount of computational resources allocated to execution of the H.264 decoder application. The results demonstrate the feasibility of the approach and its possible applicability to the broader class of PPN streaming applications.
The recent emergence of large-scale knowledge discovery, data mining and social network analysis, irregular applications have gained renewed interest. Cache-based architectures do not provide optimal performances with such workloads, mainly due to the low spatial and temporal locality of their control and memory access patterns. This paper presents a multi-node, multi-core, multi-threaded shared-memory system architecture designed for the execution of large-scale irregular applications, and built on top of three pillars that support these workloads. First, transparent hardware support for Partitioned Global Address Space (PGAS) provides a large globally-shared address space with no software library overhead. Second, multi-threaded multi-core processing nodes achieve the necessary latency tolerance required when accessing physically distributed global memory. Third, hardware support is provided for inter-thread synchronization on the global address space. An analytical performance model that accounts for the main architecture and application characteristics is presented. The hardware design of the proposed custom architectural building blocks is then described. Finally, a multi-board FPGA prototype of the proposed system with typical irregular kernels and benchmarks is presented. The experimental evaluation demonstrates the architecture performance scalability for different configurations of the whole system.
The use of FPGA platforms developed with off-the-shelf soft cores has recently emerged as one of the most promising fast prototyping approaches to design, evaluate and validate new architectural components for multi- and many-core processors. The approach appears to provide valuable benefits: optimizations to complex designs can be evaluated directly in hardware, at speeds hundreds of times faster than simulation, with efforts apparently limited only to the development of the new components. However, current FPGA toolchains that allow quick deployment of system-on-chip designs still have troubles when implementing multiprocessor designs. Often, a significant effort is also required to address the limitations of these toolchains. In this paper we discuss the design of a multi-node FPGA prototype, developed with the Xilinx toolchain, for exploring components to optimize multi- and many-core processors for the execution of irregular applications. Irregular applications, such as data-mining and social network analysis, employ large, pointer-based data structures (graphs, unbalanced trees, unstructured grids) that present poor locality and are very difficult to partition. Commodity clusters, which integrate powerful multi-core cache-based processors, are optimized for locality and employ distributed memory programming models. Developing irregular applications on them is complex, and often it does not provide performance scaling. We designed a set of hardware/software components that can potentially enhance commodity processors for efficiently executing irregular applications on multi-node systems, and we have integrated and validated them by exploiting FPGA rapid prototyping. We present the components and the prototype, highlighting the benefits and challenges in using such approach for architectural studies. We present an initial study on the tradeoffs of the platform, showing how prototyping can be effective, but also underlining the aspects that still need to be improved in the toolchain to allow better and deeper analysis.
We present an efficient implementation of the radix sort algorithm for the Tilera TILEPro64 processor. The TILEPro64 is one of the first successful commercial manycore processors. It is composed of 64 tiles interconnected through multiple fast Networks-on-chip and features a fully coherent, shared distributed cache. The architecture has a large degree of flexibility, and allows various optimization strategies. We describe how we mapped the algorithm to this architecture. We present an in-depth analysis of the optimizations for each phase of the algorithm with respect to the processor's sustained performance. We discuss the overall throughput reached by our radix sort implementation (up to 132 MK/s) and show that it provides comparable or better performance-per-watt with respect to state-of-the art implementations on x86 processors and graphic processing units.
Massively multithreaded architectures like the Cray XMT address the needs of irregular data-intensive applications better than commodity clusters. A proposed evolution of the XMT integrates multicore processors and next-generation interconnects, along with memory reference aggregation to optimize network utilization.
This paper presents an architecture for high performance computing systems specifically targeted to irregular applications. We show how a multi-core paradigm can benefit from next-generation memories and networks, while still resorting to fine-grained multi-threading for latency tolerance. At the same time, we also show how such an architecture template must employ specific techniques to optimize bandwidth utilization and achieve better scalability, proposing a mechanism based on remote memory references aggregation. We explore the proposed architecture template, using a custom simulation infrastructure, and validate its performance with three typical irregular applications. Our experimental results show the benefits provided by the multi-core approach, in terms of improved scalability, and by the reference aggregation technique, in terms of contention reduction and bandwidth optimization. For a configuration with 32 nodes, 8 cores and 2 memory controllers per node, the proposed bandwidth optimization technique with the best parameters achieves from 1.20 to 2.15 times higher performance and a reduction of network traffic up to 34.7% with the considered applications.
Application Specific Instruction-set Processors (ASIPs) expose to the designer a large number of degrees of freedom. Accurate and rapid simulation tools are needed to explore the design space. To this aim, FPGA-based emulators have recently been proposed as an alternative to pure software cycle-accurate simulator. However, the advantages of on-hardware emulation are reduced by the overhead of the RTL synthesis process that needs to be run for each configuration to be emulated. The work presented in this paper aims at mitigating this overhead, exploiting a form of software-driven platform runtime reconfiguration. We present a complete emulation toolchain that, given a set of candidate ASIP configurations, identifies and builds an overdimensioned architecture capable of being reconfigured via software at runtime, emulating all the design space points under evaluation. The approach has been validated against two different case studies, a filtering kernel and an M-JPEG encoding kernel. Moreover, the presented emulation toolchain couples FPGA emulation with activity-based physical modeling to extract area and power/energy consumption figures. We show how the adoption of the presented toolchain reduces significantly the design space exploration time, while introducing an overhead lower than 10% for the FPGA resources and lower than 0.5% in terms of operating frequency.
Irregular applications, such as data mining or graph-based computations, show unpredictable memory/network access patterns and control structures. Massively multithreaded architectures with large processor counts, like the Cray MTA-1, MTA-2, and XMT, appear to address irregular application requirements better than commodity clusters. However, the research on massively multithreaded systems is currently limited by the lack of adequate architectural simulation infrastructures due to issues such as size of the machines, memory footprint, simulation speed, accuracy, and customization. At the same time, Shared Memory MultiProcessors (SMPs) with multicore processors have become an attractive platform to simulate large-scale systems. This paper introduces a cycle-level simulator of the massively multithreaded Cray XMT supercomputer. The simulator runs unmodified XMT applications. We discuss how we tackled the challenges posed by its development, detailing the techniques implemented to obtain high-simulation speed while maintaining a high accuracy. By mapping XMT processors (ThreadStorm with 128 hardware threads) to host computing cores, the simulation speed remains constant as the number of simulated processors increases, up to the number of available host cores. The simulator supports zero-overhead switching among different accuracy levels at runtime and includes a parametric network and memory model that takes into account contention and hot spotting. On a modern 48-core SMP host, the proposed infrastructure simulates a large set of irregular applications 500 to 2,000 times slower than real time when compared to a 128-processor XMT, with an accuracy error under 10 percent. Emulation is only from 25 to 200 times slower than real time. The paper also presents a case study, where the simulation infrastructure is used to identify bottlenecks in the current XMT architecture and to estimate the performance scaling of a possible multicore design with next generation memory and network interconnect.
This chapter contains sections titled: Introduction State of the Art A Tool for Energy-Aware FPGA-Based Emulation: The MADNESS Project Experience Enabling FPGA-Based DSE: Runtime-Reconfigurable Emulators Use Cases References
This workshop, this year in its second edition, aims at bringing together scientists with all these different backgrounds to discuss, define and design methods and technologies for efficiently supporting irregular applications on current and future machines. The call for papers attracted 10 submissions from Asia, Europe and the United States. The program committee accepted 6 papers, organized in three sessions: architectures for irregular algorithms, using GPUs for Irregular Applications, and Programming Models for Irregular Applications. The program included two keynote speeches: "Big Data: OPs vs FLOPs" by Steve Wallach and "Irregular Applications and Their Architectural Challenges" by Dr. Pradeep Dubey. Finally, an outstanding panel of speakers (Umit Catalyurek, Pradeep Dubey, David Gleich, Alex Ramirez, Vinod Tipparaju, Steve Wallach) discussing current challenges and possible solutions for the efficient execution of irregular applications concluded the program. We hope that these proceedings will serve as a valuable reference for researchers and developers in the field of irregular applications.