In the context of factory automation, faulttolerant sensor systems play an important role in enhancing the safety of Networked Control Systems (NCSs). Using incorrect sensor data can have disastrous consequences. This paper proposes a flexible fault-tolerant sensor system based on the self-purging architecture. It is implemented on the same SRAM-based FPGA housing the NCS's controller. The same circuit can accommodate four, five, or six sensors depending on system requirements. It is fail-safe; it activates an error signal when it detects an error in its own circuitry, which leads to an immediate system shutdown before a faulty output propagates. It also activates a signal to indicate when the number of remaining sensors is not enough for its voter to produce a correct result. Finally, it uses biased voting to enable the user to connect a mix of original and functionally identical lower-end sensors in order to reduce system cost.
Reliable sensors are essential in Networked Control Systems (NCSs). Harsh industrial environments can significantly reduce sensor lifetime. Using fault tolerance at the sensor level is a potential solution to this problem. This paper proposes a fault-tolerant self-purging architecture with four functionally identical sensors, not necessarily of the same quality. The Error Recovery Circuit (ERC) is implemented on the SRAM-based FPGA housing the system’s controller. A design is proposed to protect the ERC from soft errors, namely single and multiple event upsets. A modified alternating logic technique is used to detect and recover from these soft errors, whether in the ERC’s configuration or user bits. The design is implemented targeting an AMD FPGA and validated using Vivado in various faulting scenarios.
This paper focuses on SRAM-based Field Programmable Gate Array (FPGA) systems in space applications. These systems are vulnerable to Single Event Upsets (SEUs) which can lead to catastrophic failures. Fault tolerance is often used to mitigate these effects; however, this may require redundant components. If these components require relatively high FPGA resources, replicating them may not be feasible. Therefore, this paper proposes a fault-tolerant architecture based on the reconfigurable duplication one in the literature, which only requires one extra copy of the module. Using the multi-die technology, where the effect of SEUs is different on each die, a modified reconfigurable duplication architecture is proposed, requiring less area resources than the conventional one. Using Continuous Time Markov Chains (CTMCs), it is proven that, according to the system parameters, the proposed architecture may have a higher reliability than the conventional one. Furthermore, it is shown that the failure rate of one of the two redundant modules does not affect architecture reliability. The coverage of the error detection mechanism is also investigated, and it is shown that a small change in its value may have a significant effect on reliability.
Due to the extensive use of sensors in factory automation, their reliability is an important research topic. This paper revisits the self-purging fault-tolerant architecture in the context of Field Programmable Gate Array (FPGA)-based workcell controllers. Both the error detection mechanism of this architecture and the controller are housed on the same FPGA. Since FPGAs are affected by Single Event Upsets (SEUs), this paper proposes a different design for the error detection mechanism to mitigate the effects of these SEUs. The design detects SEUs in both configuration and user bits without the need for redundant copies. The error detection mechanism was successfully implemented on an AMD Zybo Z7-10 FPGA and tested in various faulty scenarios demonstrating its ability to recover from SEUs in both user and configuration bits.
Advanced Driver Assistance Systems (ADAS) are being thoroughly researched nowadays especially in the field of autonomous vehicles. However, motorcycles and Advanced Rider Assistance Systems (ARAS) have not yet received the same level of attention. This paper is an attempt to address some of the problems encountered by motorcycle riders in urban areas, namely obstacle detection and alerting the rider if vehicles behind him/her are travelling at an illegally high speed. The proposed system uses a fusion algorithm that combines readings from two LiDAR sensors, one on the motorcycle and the other in the rider’s helmet. Additionally, two cameras are fixed on the motorcycle, one facing forward and the other facing backwards. YOLO, Ultra-Fast Lane Detection network and sensor front/rear sensor fusion are then used to reach a reliable decision to alert the rider of a potential danger and to advise him/her regarding the most appropriate maneuver to avoid the danger.
In the context of factory automation and reliable workcells, this paper proposes a fail-safe low-cost redundant sensor system. It is based on the self-purging fault-tolerant architecture. The system uses sensors of different qualities to reduce cost while guaranteeing that the workcell only uses data from an original high-quality sensor. Since some workcell controllers are implemented on Static Random Access Memory (SRAM)-based Field-Programmable Gate Arrays (FPGAs), the proposed system is implemented on a 7 -series AMD FPGA. However, since FPGAs are vulnerable to Single Event Upsets (SEUs), the proposed system uses spatial and temporal redundancy to mitigate their effects, whether in the user or configuration bits. It has three outputs, the voted sensor output, an error signal which is activated when a problem is detected in the proposed system itself or when sensor redundancy is exhausted, and a warning signal indicating that the system is operating with only one original sensor. A reliability analysis is then conducted to show that the decrease in lifetime of the proposed system can be minimal compared to a conventional self-purging system.
This paper studies the introduction of fault tolerance into Flexible Manufacturing System (FMS) design. The focus is on workcell controller failures. The proposed solution does not require any additional hardware. Riverbed simulations are used to prove that the proposed solution satisfies the required timing constraints. To keep the packet delay within acceptable limits, it is necessary to operate some workcells at a reduced speed. To quantify the effect of speed reduction (in case of a controller failure) on the production rate, a performability model is developed which takes into account controller lifetime and the speed of operation of the different workcells in the FMS over time.
This paper proposes a fault-tolerant flexible manufacturing system (FMS) that features a dual-level fault tolerance mechanism at both the workcell and system levels to enhance reliability. The workcell controller was implemented on a Field Programmable Gate Array (FPGA). Reconfigurable duplication was used as the first level of fault tolerance at the workcell level. It was shown how to detect and recover from FPGA faults such as Single Event Upsets (SEUs), hard faults, and Single Event Functional Interrupts (SEFIs). The prototype of the workcell controller was successfully implemented using two Zybo Z7-20 AMD boards and an Arduino DUE. Petri Nets were used to prove that controller reliability increased by 346% after 1440 operational hours. The second level of fault tolerance was at the FMS level; the Supervisor (SUP) took over the responsibilities of any malfunctioning workcell controller. Riverbed software was used to prove that the system successfully met the end-to-end delay requirements. Finally, Matlab showed that there is a further increase in performability.
This paper presents an enhanced fault-tolerant architecture for industrial Networked Control Systems (NCS) that simultaneously supports supervision video transmission and maintains strict real-time constraints. Building upon previous work that optimized video and control data flows through static traffic management techniques, this study introduces approaches to address fault tolerance and scalability challenges. First, supervision video enhanced quality operating up to the standard 50 frames per second (fps) was tested. Second the use of Field Programmable Gate Array (FPGA) resources embedded within the core network switch to handle controller fault recovery locally is proposed and tested. The use of this FPGA-enhanced architecture effectively isolates control and video traffic, ensuring real-time responsiveness even with supervisory cameras operating at high frame rates. Simulation results confirm maximum delays in faulty scenarios, well within the real-time requirements. The proposed solutions demonstrate possible scalable methods for integrating fault tolerance and high-quality video supervision in modern industrial NCS, with the FPGA-based design offering latency improvements of 7.58% over previously studied designs.
This paper focuses on prolonging the lifetime of pacemakers. Instead of just using a battery, a fault-tolerant energy harvesting system is designed; it is based on helical devices attached to the pacemaker leads. The design takes into consideration that the energy harvesting circuitry and a smaller battery will occupy the same area as the original battery. It is shown that the proposed circuit is able to supply the necessary power at 2.2V and tolerate the failure of any single component in the energy harvesting system in addition to several multiple failures. The small non-rechargeable battery will only supply current when the entire energy harvesting system fails.
Precision agriculture is a growing field worldwide. The quality of the monitoring process in greenhouses has an important effect on the quality of the crops. This paper studies a 200m × 40m greenhouse with wireless sensors, actuators, a controller, and five Access Points (APs). The greenhouse has two different crops, one more sensitive than the other and requiring a more dependable monitoring process. A Performability model is developed (based on Markov models) to help greenhouse management make appropriate decisions regarding the best sensor data routing scheme in case of one or more AP failures. The model considers the location of the crops in the greenhouse as well as their different monitoring requirements and profit margins. A case study case is presented, to illustrate the use of the model, which produced counter intuitive results.
Fault-Tolerant (FT) architectures are often used in the context of Industrial Automation, especially in safety-critical systems. Coverage is a very important parameter; it measures the ability of these systems to detect and recover from expected failures. This paper investigates the effect of changes in the value of the Coverage parameter on the steady state availability of several commonly used FT architectures, namely Triple Modular Redundancy (TMR), Sift-Out and Reconfigurable Duplication. When comparing TMR and Sift-Out, it is found that, at relatively low values of coverage, adding redundancy will reduce availability and increase downtime; this is a counter intuitive result. For Reconfigurable Duplication, it is found that, at low Mean Time To Repair (MTTR), the effect of small changes in coverage has little effect on architecture availability. Hence, it is preferable to invest in the quality of the system's redundant modules instead of the quality of the error detection and recovery mechanisms. In contrast, at higher MTTRs, a small increase in the coverage affects availability; hence, investing in the quality of error detection will lower downtime and decrease profit loss.
This paper addresses two issues. First, in the context of factory automation, a low-cost fault-tolerant sensor system is proposed along with its FPGA-based voter. Second, the voter itself is designed to recover from failures affecting its own circuitry. The sensor system consists of one high-quality sensor used for measurement along with two functionally identical, but lower cost sensors used for monitoring purposes. A biased voter circuit is designed and implemented on the same Field Programmable Gate Array (FPGA) used to house the controller. It is shown that the voter can detect and recover from Single Event Upsets (SEUs) which often occur in the harsh factory environments. Both FPGA configuration and user bits are protected. It was shown that the proposed design requires less resources and power than some other designs presented in the literature. Focusing next on recovering from SEUs in both user and configuration FPGA bits, another use case is investigated, namely an I2C master; it is shown that the proposed design again requires less resources than some other commonly used techniques. Both proposed designs were implemented on the Xilinx Zynq-7000 (xc7z020clg400-1).
Electronic components in spacecrafts operate in very harsh environments. Furthermore, radiation levels may change when moving from one orbit to another. SRAM-based FPGA systems are affected by radiation which induce soft and hard errors. This paper proposes multi-die FPGAs with dice of different technology nodes to take advantage of the higher reliability of older nodes. Several fault-tolerant architectures are analyzed, and it is shown how to choose the appropriate one with modules on one or both dice. Biased voters are designed to increase architecture reliability. A case study illustrates the use of this technique on a Xilinx FPGA.
Unmanned Aerial Vehicles (UAVs) are quickly becoming a very important component in many applications, such as search and rescue. In many of these applications, Deep Neural Networks (DNNs) are used, especially for obstacle detection and avoidance. Furthermore, online retraining may be required for adaptation to new environments. This paper shows how to enable UAVs equipped with an FPGA-based controller to achieve the required accuracy, whether in the inference phase or in the retraining phase, with minimal impact on power and performance. Dynamic Function eXchange (DFX) is used to download the retraining module bitstream utilizing Posit16 multipliers when the UAV navigates a new environment. Otherwise, a Posit8 multiplier (lower in precision) based implementation is used for the inference procedure that requires less precision than the retraining procedure. An implementation on the EP4CE22F17C6 FPGA is presented to show that a Posit16 multiplier has very close power utilization and performance to that of a Posit8 multiplier.
Motors are extensively used in factory automation. A correct determination of the motor shaft position is crucial for the success of any process. Rotary Gray encoders are often used to determine motor shaft position because of the reliability of their readings. In this paper, two error detection and masking mechanisms are developed to ensure the correctness of the data produced by the encoder for applications requiring either a small number of codewords or a larger number of codewords. These mechanisms rely on analyzing data from three identical encoders, such that the correct reading is produced even if one of the encoders fails and outputs a non-codeword or an incorrect codeword. Both mechanisms are validated on the Intel Cyclone IV E EP4CE22F17C6 FPGA, where exhaustive testing was carried out successfully to prove their error detection and masking capabilities.
Sensors are essential components in factory automation. Their reliability can be increased by using fault-tolerant architectures such as Triple Modular Redundancy (TMR). Furthermore, error detection and diagnosis can significantly reduce downtime. This paper first proposes an Analog Error Detection and Masking (AEDM) circuit that produces a correct analog reading as long as at least two of the three analog sensors are operational even if the readings are not identical. This circuit operates in the analog domain. It is then combined with a Digital Error Detection and Masking (DEDM) circuit from the literature, and it is proven that the combined mechanism can detect up to one failure while identifying the failure type (sensor failure or Analog to Digital converter (ADC) failure). This mechanism is then slightly adapted by adding two multiplexers to detect up to two failures simultaneously (one sensor and one ADC). Both mechanisms can propagate a correct output to the rest of the system while also identifying the erroneous modules for easier diagnosis and less repair time. The AEDM was successfully simulated on ELDO.
This paper focuses on the issue of FPGA-based controller reliability in the context of Flexible Manufacturing Systems (FMS). Due to the industrial harsh environment, FPGA-based controllers are susceptible to both failures and aging effects. A framework that integrates Fault-Tolerant (FT) FPGAs with aging mitigation capabilities is proposed to improve FPGA-based industrial controllers' reliability and extend their lifetime. The fault set includes soft faults represented by both Single Event Upsets (SEUs) and Multiple Event Upsets (MEUs), hard faults, and aging effects. The design relies on a customized Triple Modular Redundancy (TMR) technique for fault detection and masking with Dynamic Function eXchange (DFX) capability for partial reconfiguration. In most cases, the reconfiguration does not require any system interruption. The design is also able to differentiate between hard, soft and aging failures. Finally, a simple case study is described to illustrate the system's behavior upon infection by hard fault or aging. The Zybo Z7-20 board is used in this case study.
Smart greenhouses enable the cultivation of various crops in different environmental conditions. Wireless Sensor Networks (WSNs) and Networked Control Systems (NCSs) could be utilized within greenhouses to access, control and monitor different environmental parameters. This paper studies a greenhouse with a special focus on the wireless NCS. It first shows how to use the 2.4GHz frequency band while minimizing Access Point (AP) cost and maintaining an acceptable Packet Loss Rate (PLR). Second, fault-tolerant AP architectures are studied, and a metric is proposed which incorporates AP failure and repair rates, as well as NCS information efficiency to help management select an appropriate architecture, to minimize profit loss. This metric is shown to produce some counter-intuitive results. Finally, Riverbed simulations are used to obtain the minimum sensor transmission power to produce the minimum acceptable PLR in the context of the 2.4GHz band. This power is much lower than that for the 5GHz band; this will prolong sensor battery lifetime.
This paper addresses the issue of analog sensor reliability in the context of Advanced Driver Assistance Systems (ADAS) in the automotive industry. Nowadays, many vehicles use two sensors for error detection. Two fail-safe and fault-tolerant architectures are proposed to increase sensor lifetime at a low cost. One of these architectures is proven to be more reliable and of a lower cost than a conventional Triple Modular Redundant (TMR) system. Both designs rely on a tunable TMR detection mechanism to process the analog outputs of three functionally equivalent sensors taking into account the unavoidable discrepancies after the Analog-to-digital conversion. This hardware error detection mechanism will reduce the processing time of the sensor data compared to a software scheme. The mechanism is implemented on Intel Cyclone IV GX EP4CGX150DF31I7AD FPGA; circuit delay and resource utilization are obtained. Finally, this circuit delay is compared to the execution time of a functionally equivalent software implemented on a NIOS II processor.