
In this paper, we present an all-optical network architecture and a routing protocol for it. The coloured sparse optical torus network (CSOT) consists of an n × n torus for which n = b2 for some b in {2, 3, 4, ...}. Processors of the network are deployed at nodes for which (i + j) mod b = 0, where i and j are row and column indices of a node and b is the block size and the number of wavelengths used. The number of processors is P = b3. Routing is based on scheduled transmission of packets and wavelength-division multiplexing. The routing protocol ensures that no electro-optical conversionis needed at the intermediate nodes and all the packets injected into the routing machinery reach their targets without collisions. A work-optimal routing of h-relations is achieved for a reasonable size of h in (P log P).
This paper investigates the potential of General Purpose Graphic Processing Unit (GPGPU) for the serve rand HMI parts of Energy Management System (EMS). TheHMI investigation focuses on the applicability and performance improvement of GPGPU for scattered data interpolation algorithms typically used to visually represent the overall state of a power network. The server side investigation focuses onfine grain parallelization of EMS applications by targeting the sparse linear solver. The different performance evaluations show the high potential of GPGPU for the HMI part with a speedup factor up to 100 at the cost of acceptable approximations, while the benefit on the server side varies from a speedup factor of up to 300 to 0 depending on the application.
Registration of partial scan data sets is still a challenge for today's CAD systems and CAD system users. Many of the known methods rely on user interaction or feature recognition. For non-regular users this is too time consuming and error prone. The paper describes a method to register partial scan data by fitting a large fat tetrahedron (LFT) in the target point cloud. The method is computational intensive and in its CPU implementation not fit for interactive use. The independency of the points in the data-sets makes massive parallel computing applicable. The paper describes the implementation of the method on a GTX260 GPU using the CUDA programming environment. A performance gain of 10 times compared to a conventional CPU implementation was achieved, which can be further improved by implementing a pre selection method of the result, bringing interactive use in range after further optimization.
The opportunities that new technologies offers businesses cannot be ignored as they stimulate the growth of new markets for goods and services that are widely sought after. The concept of green technologies has fuelled the increasing demand for environmentally, economically and socially sustainable goods and services, thus businesses are developing their products in greener ways, offering greener services which enhance their sustainability efforts in any economy. Green technology has been defined by the authors as the application of technological expertise to change the processes, methods and techniques of producing goods, from energy consuming, environmentally unfriendly resources, to energy saving, economically viable and environmentally friendly products.
In this paper we have explored different possibilities for partitioning the tasks between hardware, software and locality for the implementation of the vision sensor node, used in wireless vision sensor network. Wireless vision sensor network is an emerging field which combines image sensor, on board computation and communication links. Compared to the traditional wireless sensor networks which operate on one dimensional data, wireless vision sensor networks operate on two dimensional data which requires higher processing power and communication bandwidth. The research focus within the field of wireless vision sensor networks have been on two different assumptions involving either sending raw data to the central base station without local processing or conducting all processing locally at the sensor node and transmitting only the final results. Our research work focus on determining an optimal point of hardware/software partitioning as well as partitioning between local and central processing, based on minimum energy consumption for vision processing operation. The lifetime of the vision sensor node is predicted by evaluating the energy requirement of the embedded platform with a combination of FPGA and micro controller for the implementation of the vision sensor node. Our results show that sending compressed images after pixel based tasks will result in a longer battery life time with reasonable hardware cost for the vision sensor node.
The ongoing move of hardware platforms to many-core processor challenges the traditional software design methodology. It is critical to develop new programming paradigms and efficient ways to port legacy applications. This paper analyzed a typical packet processing application and also the cache hierarchy and behavior of Raw architecture many-core processor. It presented an easy to implement run-time dynamic core grouping approach to improve the system performance. This approach reduced the cache swap latency by grouping neighbor cores attached to the mesh network. It optimized the scale of group by experimental data got beforehand. The test results showed this approach can improve the Deep Packet Inspection (DPI) system performance around 10% with very minor code change.
This paper presents our parallelization and implementation of the ORTHOMIN solver on the Cell Broadband Engine. The solution of linear systems of equation sis one of the most central processing unit-intensive steps in many engineering and simulation applications and can greatly benefit from the multitude of SIMD-capable synergistic processor element (SPE) cores in the Cell processor. We report the serial ORTHOMIN implementation on the Cell's PowerPCprocessor element (PPE), and the parallelization and performance analysis of ORTHOMIN across 8 SPEs forTridiagonal (1-D reservoir grid) and Heptadiagonal (3-Dreservoir grid) matrices. Our implementation is shown to scale well with data size, and grid dimensionality.
As the vision of ubiquitous computing becomes reality, there is a possibility for user interfaces that follow the user through the physical world by jumping between local display devices. This paper presents previous attempts at this functionality and identifies their limitations. Justified by these limitations, we then present our novel method that uses new features of web-based technology. This work has been recently developed and deployed as part of a joint European ATRACO project in our purpose built iSpace living lab.
Variable economic conditions, new forms of customer buying behavior and in particular new technology is likely to cause the emergence of new or growing existing hospitality markets. In developed European economies, all the more attention is paid to studying the role of new technologies in the field of hospitality marketing. Research focus will be directed towards the usage of technological systems in marketing of hospitality. Internet, as one of the most significant technological phenomena of our time, provides hotel managers completely new competitive opportunities, of which the most significant opportunity to provide immediate and always open access to information throughout the world. According to latest researches, Croatian hotel managers usually use these technological systems: Own web site, System to access the Internet, Points of sale (POS), System of calculating the telephone calls, Local area network (LAN), Uniform system of accounts for lodging Industry (USALI), Intranet system, Central reservation system (CRS), Marketing information system (MIS) and. Object management system (PMS). Modern hotel managers want to achieve automation of marketing strategies and interactive communication with potential guests. However, certain hotels are still not developed a marketing strategy through the Internet pages. Adjustment trends must be timely and possible, so it is necessary to intensify investment in that direction. As a result, it is assumed that the possibility of information and purchase via the Internet to stimulate the future purchase of tourists and change their previous spending habits. To make hotel companies improve their e-marketing strategy, the research shows insufficient usage of the following Web 2.0 tools in Croatian hospitality market: Instant Messaging, Internet Relay Chat, Internet Forums, Social Network Services, Social Guides, Social Bookmarking, Social Reputation Network, Web logs, Social Citations, Peer-to-peer Social Networks, Virtual Presence, Virtual Worlds & Massively Multiplayer Online Games, VOIP -- Internet Telephony and Mobile Internet. Establishing relationships and technological innovations significantly affect the shortening of business cycles and thus the importance of creating competitive advantages in the hospitality industry of Croatia.
This paper presents a goal-oriented approach for adaptive resource management at the OS level. The approach is based on fuzzy logic component. The main aim of the paper is the evaluation of the benefits of using heuristics for the above purpose versus previous work based on building a knowledge base for evaluating alternative system parameter configurations. The paper also presents a prototype which aims to adapt the system, with the goal of improving the level of performance offered to the users of the system, by adjusting different settings accessible through the Workload Manager. The prototype uses the build into OS resource-oriented mechanism and is doing a translation into goal-oriented approach to optimize the usage of resources available while executing a set of applications with different needs.
Codes on graphs have become the most important way to reach channel capacity. A new problem for High Performance Computing has been constructed with this work. The graph search problem as posed in this work is the coarsest version of the problem which corresponds to the exhaustive search case. The coarse grain graph search (CGGS) problem chooses an optimal parity-check matrix, based on minimum value of BER objective function. Based upon the specified range of parity-check matrices and range of signal to noise ratio (SNR), the CGGS generates the parity-check matrices using the Modified PEG algorithm, performs LDPC encoding, adds noise to the encoder output, performs LDPC decoding, computes Bi terror rate (BER), and thus the BER objective function. TheCGGS problem in its current implementation runs in an MPIimplementation on the T2K supercomputer at The University Of Tokyo. Two scheduling strategies have been proposed, Generalized Max-Min pairing (GMMP) and random pairing(RP) scheduling algorithms have been proposed, and good parallel performance of CGGS over an MPI implementation onT2K has been achieved.
NP Complete problems are one of the most complex problems in computer science but their vast applications in real world always pushes the scientists to explore new ways to solve them. We extended the original problem definition of Boolean Satisfiability Problem to finding all satisfiable solutions of a given problem instance and used massively parallel architecture of CUDA (Compute Unified Device Architecture) for reducing its run time. Certainly, exponential time complexity of this NP Complete problem can't be so easily reduced to polynomial time but through our implementation it can be well challenged for sufficiently large number of inputs. Our fully optimized implementation attained speedup ranging from 19x to 188x on Nvidia Ge Force GTX 260 for different type of SAT problems.
The emergence of the third dimension in Network-on-Chip (NoC) design as a quest to improve the quality of service (QoS) of on-chip communication has evolved with enormous interest. However the underlying router architecture of 3D NoCs have more area footprint than 2D routers. In this paper, we investigate heterogeneous 3D NoC topologies with the focus on finding a balance between the manufacturing cost and the QoS by employing the area and performance benefits provided by 2D routers and 3D NoC-bus hybrid router architectures in 3D mesh topology. Experimental results show a negligible penalty in throughput of 3D mesh with homogeneous distribution of 3D NoC-bus hybrid routers. The heterogeneity however provides superiority in area efficiency of the NoC resources.
Due to stagnating CPU cycles, future performance gains in automation firmware are unlikely to be achieved without parallelization for multi-core architectures. However, for a sophisticated system comprising millions of lines of code, this process induces significant effort, especially when having to keep real time and safety conditions. As efficiency matters in corporate software development, obtaining maximum speedup by spending no more implementation effort than necessary is intended. Thus, the design of a parallel firmware is recommended to base on the results of a model-based exploration and evaluation of efficient parallelization alternatives. For this purpose, we developed the EEEPA tool chain, that starts with graph-based firmware modeling on basis of dynamic event logs.
The Internet has become a dominant part of the learning process of students in advanced learning. Many of these students currently do not represent a considerable portion of the consumer market in terms of buying power but given their continuous exposure to the internet and their eventual adoption of Internet banking, in the future, when they enter the work force, will represent that significant group of individual that have influential purchasing power in the financial market.
The minimization of power losses in the medium voltage (MV) grid requires adjustment of network of power sources. This problem is particularly important for renewable energy sources, for example for the farms of wind generators. Their placement and nominal power should be selected according to the configuration of the network and the largest loads. The presented problem is solved using a genetic algorithm (GA). The formulation of the GA algorithm and its performance for different numbers of power sources is analyzed. The optimal placements of wind generators were computed for some case problems. The algorithm is validated with a medium sized electrical grid. The formulation of the parallel version of the genetic algorithm is presented. Its properties are verified on the cluster of workstations environment.
State Machines (ASM) are mathematically defined environment for high-level system design, verification and analysis. This paper presents a definition of the hybrid approach to the specification, analysis and testing of stateful grid services using ASM. This approach allows an easy integration of created specification of developed middle ware with existing components of grid systems. The important advantage of this approach is an automatic testing of the implementation, following the model-based testing approach. This allows a smooth transition from the specification to implementation stage, as well as investigation of features of specification and implementation, at every stage of their development. Also, a software environment has been developed which implements the defined approach.
We consider both point-wise and block-wise versions for solving a linear triangular system on homogeneous machines. With p identical processors and for a problem of size N where N/r=n=2pq+1 (q≥3 and r being the block-size), we design an optimal parallel algorithm. We show its optimality in terms of both computing and communication costs. Finally, we determine the optimal value of the block size which minimizes the parallel execution time. A series of experimentations confirm the theoretical results.
The Infobright database engine stores data in column wise chunks in a highly compressed form. Access to data is slowed down by the necessity of fetching required chunks from disk and decompressing them. A prefetching mechanism is meant to perform these tasks shortly before a running query needs the data for processing. The prefetching implementation in Infobright is described and evaluated. An average speedup of 1.7 is reported. The speedup reaches 2.4 for some queries and generally is higher for long lasting queries for which the data fetching and decompression cost dominated the execution time. The limitations of the prefetching are presented on both system level and SQL query level.
A new framework for designing evolved program execution control in distributed programs is discussed in the paper. The framework provides an infrastructure for designing distributed program control based on monitoring of global application states. Global control constructs are proposed which logically bind distributed program modules and define the flow of control dependent on the monitoring of global application states. Such control can be organized in programs at the process and thread levels. Special processes and threads called synchronizers collect state information from application modules, construct strongly consistent global states and evaluate control predicates on global states. Based on this evaluation control signals are sent to processes and threads to define inter module flow of control and to influence the internal module behavior. The proposed constructs are incorporated into the framework as a graphical API which is compiled into C/C++ programs with the MPI2, pthreads and Open Mp libraries for communication.