With the increase in the non-linearity and complexity of the driving system’s environment, developing and optimizing related applications is becoming more crucial and remains an open challenge for researchers and automotive companies alike. Model predictive control (MPC) is a well-known classic control strategy used to solve online optimization problems. MPC is computationally expensive and resource-consuming. Recently, machine learning has become an effective alternative to classical control systems. This paper provides a developed deep neural network (DNN)-based control strategy for automated steering deployed on FPGA. The DNN model was designed and trained based on the behavior of the traditional MPC controller. The performance of the DNN model is evaluated compared to the performance of the designed MPC which already proved its merit in automated driving task. A new automatic intellectual property generator based on the Xilinx system generator (XSG) has been developed, not only to perform the deployment but also to optimize it. The performance was evaluated based on the ability of the controllers to drive the lateral deviation and yaw angle of the vehicle to be as close as possible to zero. The DNN model was implemented on FPGA using two different data types, fixed-point and floating-point, in order to evaluate the efficiency in the terms of performance and resource consumption. The obtained results show that the suggested DNN model provided a satisfactory performance and successfully imitated the behavior of the traditional MPC with a very small root mean square error (RMSE = 0.011228 rad). Additionally, the results show that the deployments using fixed-point data greatly reduced resource consumption compared to the floating-point data type while maintaining satisfactory performance and meeting the safety conditions
Embedded systems often use FPGA (Field Programmable Gate Array) chips for their implementation. Soft processors are optimized for their implementation in such circuits. Embedded system processor load is crucial when real-time applications are implemented. Processor load can be analyzed in several ways. The application described in this paper demonstrates a possible method to analyze the processor behavior and a model was developed, which is able to classify processor load. The model was implemented in FPGA without any modification of the original embedded system. The results of the analysis mark the time periods when the processor is overloaded.
Az egyetemi oktatásban is jelentős szerepe van a gyakorlati oktatásnak. A műszaki képzések tantárgyainak gyakorlatain a hallgatók méréseket végeznek, tervezési feladatokat oldanak meg. A munkájukat dokumentálniuk kell mérési jegyzőkönyvek, tervezési dokumentáció, vagy szoftverek formájában. A Miskolci Egyetem Automatizálási és Infokommunikációs Intézetében kifejlesztett egyedi informatikai rendszer a gyakorlati oktatás minden fázisát támogatja, beleértve a feladatok rendszerezését, a beosztások elkészítését, a feladat kiadást, a dokumentumok összegyűjtését és értékelését.
Since the beginning of using parallelized computing units it is known that the actual performance is less than the possible nominal performance: the unproductive part of the computing performance remains “dark”. It is also known that the amount of dark performance strongly depends on the number of parallelly working processing units, so its role must be crucial for supercomputers, where in the coming exa-scale models millions of processors are utilized, as well as for the exa-scale applications they are running, like brain simulation and Earth simulation. Although the effects affecting parallel performance are known from the beginning, their relative weights have been considerably changed with the development of the field and strongly depend on the type of application. For large computer systems the “dark performance” represents a new major obstacle, in addition the former ones, like “heat wall”, “memory wall”, “dark silicon”, etc. The careful reconsideration discovers that in contrast with the general belief, supercomputer performance has an upper limit, and reaching that limit explains some strange and mysterious events, like canceling projects immediately before their target date or that the special-purpose brain simulator cannot outperform the many-thread simulator running on a general-purpose supercomputer.
This article explores the possibilities of new computation architecture. The paper tries to present some modifications in multiprocessor architectures in order to obtain performance increase in computation speed by parallel memory access. In the presented example it is shown how an interrupt can be served skipping the interrupt service routine usual steps.
As all other computing laws [1], the computing performance shows a “logistic curve”-like behavior, rather than unlimited growth. Since cca. 2000, the single-processor performance increased only marginally [2], unlike the need for more computing power in the everyday tasks. The stalling forced computer experts to look for alternative methods. The preferred way seems to be to continue the traditions of the single-processor approach: to assemble systems comprising several segregated processors, connected in various ways.
The today computers, as taught in the elementary computer architecture courses, are based essentially on the principles, suggested by von Neumann in a 1946 report. Since that, only minor improvements have been carried out, although the impressive technical developments seem to hide this fact. In the meantime, also a whole software industry has grown up, and masses of society are using computers, sometimes in quite unexpected places. After several decades of developments, the architecture suggested by Neumann (and its consequences) is heavily criticized. Most people are hoping some kind of solution from reconfigurable computing (RC). The present paper attempts to review in the light of the technical background and recent developments, how much of the original assumptions are still valid and also suggests some alternative ideas for improving the widely used computer architectures using RC, to make them more effective. The ideas suggested here are ”patches” to the single-processor architecture, and considerable speed gain is expected from those architectural changes in a multi-tasking execution environment.
Although clusters usually represent a single large resource for grid computing, most of the time they are utilized as individual resources for desktop grids. This chapter presents a lightweight, fault-tolerant approach for connecting firewalled already deployed clusters for desktop grid computing. It allows to easily offer cluster resources for desktop grids without requiring a large administrative overhead of connecting them separately. Our approach uses Condor for cluster management and an extended version of BOINC to connect those resources across the Internet or wide area networks. We describe security considerations and present our security model for the system and sharing resources. We present running projects where our solution is being used to utilize already deployed clusters requiring low effort from administrators.
This paper describes structure and internal behaviorof GRAPNEL programming environment. Developmentof the GRAPNEL environment was supported byEC within COPERNICUS Research Project No. 5383.This project has implemented a graphical programmingenvironment for parallel programming. The developedenvironment helps programmers to develop, write, execute,and verify parallel programs for distributed environmentsmore quickly.1. IntroductionGRAPNEL is a graphical programming languagewhich has...
In this paper we present a workflow solution to support graphically the design, execution, monitoring, and performance visualisation of complex grid applications. The described workflow concept can provide interoperability among different types of legacy applications on heterogeneous computational platforms, such as Condor or Globus based grids. The major design and implementation issues concerning the integration of Condor tools, Mercury grid monitoring infrastructure, PROVE performance visualisation tool, and the new workflow layer of P-GRADE are discussed in two scenarios. The integrated version of P-GRADE represents the thick client concept, while the portal version needs only a thin client and can be accessed by a standard web browser. To illustrate the application of our approach in the grid, an ultra-short range weather prediction system is presented that can be executed in a grid testbed and visualised not only at workflow level but at the level of individual parallel jobs, too.
This paper describes behavior of source code generator of GRADE system which is a graphical programming environment to help programmers to develop, write, execute, and verify parallel programs for distributed environments. The described tool is a translator which automatically converts graphical representation of the developed program into C source code.
To provide high-level graphical support for developing message passing programs, an integrated programming environment (GRADE) is being developed. GRADE provides tools to construct, execute, debug, monitor and visualise message-passing based parallel programs. GRADE provides a general graphical interface that hides low-level details of the underlying message-passing system thus, it allows the user to concentrate on really important aspects of parallel program development such as task decomposition. The current paper describes the translation mechanism that is applied in GRADE to generate the executable message-passing code from the high-level graphical description of the user application.
Róbert Lovas合作论文数Laboratory of Parallel and Distributed Systems1