在需要长时间可靠运行的软件系统中,由于持续运行时间和任务响应速度的要求增加,工作组件在被探测到失效后将被冗余组件实时替换.但现有可靠性优化研究通常假设冷备份冗余在所有积极冗余组件失效后才使用.针对支持实时替换的混合冗余策略,对其冗余度优化分配进行研究.该策略不仅能够保障系统可靠性,而且能够保障系统性能,故选用实时可用性和任务完成效率两类约束条件,建立冗余配置代价最小化模型.基于马尔可夫链理论对可靠性及性能两类系统指标进行定量分析;采用数值计算方法对非线性的状态分析模型进行计算;改进二元组编码遗传算法对上述优化问题进行求解.采用实例对串并联系统中实时可用性及任务完成效率的分析进行了说明,并对优化冗余分配模型进行了验证.实验结果表明,在相同冗余度下,支持实时替换的混合冗余策略在任务完成效率方面优于传统的混合冗余策略.所以,在相同约束条件下不同混合冗余策略需要采用不同的冗余优化配置方案.
While joint redundancy and maintenance strategies are used to maintain system reliability, optimization is often conducted to choose appropriate configuration parameters for each strategy. Existing research mainly deals with imperfect preventive maintenance strategy optimization and ignores the impact of inspection and detection interval before maintenance occurs. So, this paper aims at a joint redundancy and inspection-based maintenance strategy which is widely used in computing systems. Following the existing Markovchain based evaluation method, an optimization model is built to choose appropriate redundancy for system structure and inspection interval for maintenance. This model is constructed to achieve best system performance under certain reliability constraint, whereas the reliability and performance models are built according to component redundancy and inspection interval. Since there is no closed-form formula of this optimization model, a greedy iterative search algorithm is used to get optimal solutions for inspection rate under each redundancy value. Empirical studies show the process of building the optimization model and calculating the optimal parameters from the model. The results indicate that this optimization method could find optimal redundancy as well as inspection rate.
Mixed redundancy strategies are generally used in cloud-based systems, with different node switch-mechanisms from traditional fault-tolerant strategies. Existing studies often concentrate on optimizing a single strategy in cloud computing environment and ignore the impact of mixed redundancy strategies. Therefore, a model is proposed to evaluate and optimize the reliability and performance of cloud-based degraded systems subject to a mixed active and cold standby redundancy strategy. In this strategy, node switching is triggered by a continual monitoring and detection mechanism when active nodes fail. To evaluate the transient availability and the expected job completion rate of systems with such kind of strategy, a continuous-time Markov chain model is built on the state transition process and a numerical method is used to solve the model. To choose the optimal redundancy for the mixed strategy under system constraints, a greedy search algorithm is proposed after sensitivity analysis. Illustrative examples were presented to explain the process of calculating the transient probability of each system state and in turn, the availability and performance of the whole system. It was shown that the near-optimal redundancy solution could be obtained using the optimization method. The comparison with optimization of the traditional mixed redundancy strategy proved that the system behavior was different using different kinds of mixed strategies and less redundancy was assigned for the new type of mixed strategy under the same system constraint.
The threshold value used in receiver autonomous integrity monitoring algorithms to identify faults has a significant impact on positioning integrity and GPS/GNSS availability. The value is usually selected empirically or under certain distribution assumptions; its calculation for a non-Gaussian test statistic has not been solved. For fault detection methods using a particle filter, a new heuristic method is proposed to select an appropriate fault detection threshold value using an optimization model. In this method, a non-Gaussian cumulative log likelihood ratio (LLR) value is used as the test statistic. Its threshold is determined using an integrity risk minimization problem with an availability constraint. Since there is no closed form for this optimization model, a genetic algorithm with a local search strategy is adopted to find a near-optimal solution. Experimental results show that this method can be used to compute the non-Gaussian fault detection threshold value subject to different availability constraints. Comparisons with empirical and distribution-based methods indicate that while meeting the same probability for false-alert constraint, the probability of missed detection in the optimized approach is much lower than for other methods, especially for small numbers of errors. Since the cumulative LLR value does not exhibit obvious statistical features for any distribution, the performance of our optimized approach is stable for different test cases and satellite data sets.
大数据环境中监控和冗余混合策略的采用引起资源优化配置模型的状态空间膨胀,进化搜索算法在整型与非整型变量结合的解空间中的搜索效率有待提高,为此提出了基于搜索邻域分析的三元组模因算法.在分析了监控频率等参数变化对组件及系统可靠性增长影响的基础上,针对监控频率提出了基于变长邻域的近邻生成方法,针对策略选项提出了与组件关联的近邻生成方法,采用模因算法框架并改进了局部搜索算子,通过组件间的迭代搜索在保持个体优势的同时增大搜索范围,该算法能够用于求解混合策略下的组件保障措施选项及相应优化配置参数;与现有多策略搜索算法相比,在相同可靠性约束下,该算法能够得到消耗更低的资源配置结果;局部搜索策略对算法稳定性未造成明显影响.
While researchers have concentrated on the optimization of joint redundancy and maintenance mechanism, maintenance in computing systems is quite different from that in traditional systems. Considering a routine monitoring and inspection mechanisms is conducted to detect component status and trigger repair process, this paper pays attention to the optimization problem of joint redundancy and inspection-based maintenance mechanism. After conducting steady state analysis on subsystems using inspection-based maintenance, shared repair facility and component redundancy, optimization model is built to search appropriate system structure and maintenance policy which maximizes system performance while meeting availability and cost constraints. Due to the complexity of uncertain optimization model, genetic algorithm is used to search optimal solution, using triple-element encoding mechanism and specifically designed operators. Illustrative examples are conducted to show that the optimization model and corresponding solution technique could be used to search optimal system configuration under given constraints and different cost constraints would lead to different optimization result while meeting availability constraints.
Mixed redundancy strategy is generally used in cloud-based systems, with different node switch mechanism from traditional mixed strategy. However, related researches often concentrates on traditional mixed redundancy strategy in which cold standby components is working only after all active nodes fail. So a model is developed to evaluate the reliability and performance of cloud-based degraded system subjected to mixed active and cold standby redundancy strategy with continual monitoring and detection mechanism. It is assumed that the node switching process is triggered once some active nodes fail and there are available standby nodes. A continuous-time Markov chain is built on top of the state transition process and both transient and steady state availability and expected job completion rate are used to evaluate system metrics with or with repair facilities. A numerical method is used to solve the model and sensitivity analysis is conducted on different redundancy strategy. Illustrative examples using real-world data were presented to explain the process of calculating the probability of each state and the different kinds of availability and performance. The comparison with traditional mixed redundancy strategy proved that the system behavior was different using different kinds of mixed strategy and the analysis model for traditional strategy was not suitable for strategies in cloud-bases system.
After analyzing the communication requirements of emergency rescue management-decisional support system in Whole-Treatment-Chain (WTC) and characteristics of wireless communication technologies, an integrated wireless communication platform based on 3G and satellites is proposed, with the dynamic switching method for different application scenarios. The implementation of this platform is presented afterwards followed by several experiments. Results analysis shows that this platform could maintain relatively stable commutation for both shelter hospital and ambulance but does not perform extremely well on medical train due to signal blocking. In a word, it enhances the application of management-decisional support system during emergency medical rescue.
With the rapid development of global navigation satellite system, Receiver Autonomous Integrity Monitoring (RAIM) has attracted attention from many researchers and there still are a lot of unsolved problems in this field. Traditional RAIM algorithms are built upon Kalman filter under the assumption that satellite signal noise follows Gaussian distribution. However, with the impact of ionospheric delay error and other factors, the measurement noise usually doesn’t follow Gaussian distribution. Under strong interference and harsh environmental conditions, particle filter is often employed to improve the efficiency of RAIM with non-Gaussian distribution errors. But those algorithms on particle filter are poor in convergence accuracy and stability because of particle degeneracy and this paper aims to provide a solution for particle degeneracy. Using the idea of approximate probability, this paper also takes use of particle filter in RAIM for fault detection and employs the idea of genetic operations to avoid particle degrading too early. Based on the classic particle filter procedure, simulation binary crossover operator, polynomial mutation operator and roulette wheeling selection method are used to produce new generation of particles from old generation in resampling process to increase particle diversity in state space as well as keeping good performance particles. The modified particle filter is then used in the fault detection and isolation process of RAIM with cumulative log likelihood ratio test. Finally, experiment are conducted using IGS tracking station observation data and the results showed that, our algorithm could be used to detect and isolate faults under non-Gaussian noise environment, which proved the efficiency of our algorithm in RAIM. Comparing with RAIM algorithms based on traditional particle filters, our algorithm could improve the fault detection accuracy and convergence rate under non-Gaussian noise conditions as well as avoiding particle degeneracy.